Hacker News, Vandaag om 00:14, 7 min lezen
Recent AI models struggled to match a human algorithmic innovation

Methodology
InnovationEval tests whether AI can independently devise an ML innovation that matches the performance of a recent human-developed innovation the AI has not seen. This is similar to recently-proposed tests for scientific ideation: if AI were presented with humanity’s knowledge up to 1905, could it rediscover special relativity?3 We ask a more modest question: if AI were presented with AI researchers’ knowledge up to early 2026, could it discover its own ML algorithmic innovation, matching the improvements achieved by human researchers since then?
We hope to achieve several advantages through this approach: end-to-end validation of AI’s R&D abilities, a requirement for genuine innovation rather than assembly of existing techniques, realistic representation of research areas, and guaranteed feasibility.
End-to-end validation: AI systems have to perform the entire process of discovering an ML innovation, from coming up with ideas through to implementing them. Similar to existing work such as NanoGPT speed-runs and ResearchGym, we define metrics that should be improved and constraints that should be satisfied.4 Using end-to-end metrics provides a legible way to assess AI performance, as long as improving the metrics genuinely requires the AI to make research progress. In our case, these metrics are set to match an existing human-authored paper. We set up an AI agent to develop a better post-training method, which requires end-to-end generation of ideas, figuring out details of their implementation, experimenting with them, analyzing the results, and iterating until reaching either success or exhaustion.
Innovation is required: Many AI R&D evaluations examine well-specified tasks that don’t require innovation,5 or can be solved by applying combinations of non-novel techniques.6 Some existing benchmarks try to isolate the task of R&D ideation,7 but it is unclear whether this task can be done in isolation from the full loop, including implementation and analysis. Our evaluation sets up a task where substantially improving the end-to-end metrics without violating scope requires development of a method the AI has not seen in training.8 We elaborate on this in Task setup.
Realism of research area: We want to test AI’s ability to discover ML techniques similar to those valued by (and used in) frontier AI labs.9 This is difficult because frontier AI developers are secretive about many of their methods. We cannot directly test AI on rediscovering their ML techniques, so we instead rely on open publications and other evidence that a technique is useful, such as adoption in prominent near-frontier models or discussion by post-training researchers.10 Here, we selected a paper about on-policy self-distillation. We discuss this in more detail below.
Feasible: Using a real, replicable AI paper guarantees that our task is feasible, and provides us information about the required GPU resources for human researchers. In some AI R&D evaluations, the objective is to improve on an existing method, but without a human baseline, and thus with less clarity on the required budget, and whether human researchers would have tried a different approach.
There are also disadvantages that come with anchoring on existing papers. One is that, at least in this iteration, we have struggled to create a task that is amenable to fully automated grading. We describe this in more detail in Task setup. Another disadvantage is that we have ended up relying on a small number of runs, since each individual attempt at this task requires substantial compute budgets.
Another disadvantage of using an existing innovation is that newer models will memorize our task. This happened over the course of this project; our main results are on Claude Fable 5 and GPT-5.6 Sol, which showed no sign of memorization when prompted to recall or guess details about the paper without using search. But their successors, Claude Fable 5.1 and GPT-6 Astra, were aware of the task. Our plan for future evaluations is to perform ongoing tests for memorization in newer models, flag their results accordingly, and devise new tasks as necessary to refresh the evaluation.
Task setup
As our testbed task, we used a recent AI innovation that has been adopted and cited by recent models: on-policy self-distillation (SDPO). The AI agent was prompted to develop a novel post-training technique that beats a strong GRPO baseline. We emphasized that the agent’s goal was “to produce a compelling research result, of the kind that would genuinely advance the field.” We then provided metrics and datasets used for the results in the original paper: short-answer questions11 and coding.12 The agent was told to produce evidence of its method’s success by post-training a Qwen3-8B model to perform better on these tasks, ideally matching or surpassing reference values set by a recent unnamed method (SDPO).
The overall eval grade is the averaged performance across the two result areas, each of which has several sub-metrics based on the original paper’s experiments. Matching or surpassing the original paper’s performance in an area yields a score of 100%, whereas scores at the GRPO baseline are scored at 0%. We provide more detail on prompting and scoring in Scoring.
It is important to set the task’s scope correctly. If the goal were purely to improve performance on these datasets, there are many ways this might be achieved, such as by generating synthetic datasets for fine-tuning. This wouldn’t count as developing a novel post-training technique and wouldn’t advance the field, so arguably the agent should know not to use this approach. Rather than trying to grade novelty, we attempted to limit the scope such that the agent can only match the original innovation’s performance through novelty in its own approach — even if it lands on a novel approach distinct from SDPO. We constrain the scope to algorithmic changes that affect the loss and its updates, and/or its rollouts and model-driven revisions given a fixed batch of training data.13 This scope allows for many different algorithmic ideas, which may differ substantially from the innovation in the original paper. However, it does constrain development to broadly the same research areas.
There is a risk that limiting the scope in this way leads to a whack-a-mole dynamic, where the agent is repeatedly searching for loopholes in our definitions and implementing solutions that we retroactively deem out-of-scope. However, even imperfectly limiting the scope is helpful, because it reduces the burden when reviewing an agent’s solution.
We initially experimented with an automated grader using an Opus 5 judge to review agents’ solutions and assess scope violations. However, since we only evaluated a small number of models, we ended up performing human-in-the-loop review after task completion, investigating submissions’ achieved scores and their workings.14,15 We discuss qualitative findings throughout.
Environment
We provided the agent with a development environment where it could edit and execute code, including launching GPU jobs via Modal. The agent was sandboxed to prevent internet access — we assume that its knowledge of post-training techniques is recent enough that it is already familiar with relevant pre-existing work.16 The scaffold is Inspect’s ReAct agent, with bash and text_editor tools, as well as tools to submit and monitor GPU jobs. We provided a starting codebase based on the paper’s repository and verl based stack, implementing the paper’s tasks and strong GRPO baseline, but scrubbed of SDPO.17
We provided fairly large GPU budgets for experiments and inference tokens, aiming to avoid limiting AIs with low budgets. Compute budgets per evaluation were 3,000 GPU-hours across a maximum of 50 GPUs, about 10× the compute required for a full training run on every individual task.18 While this is plausibly enough compute, the GPU budget could still be a limitation; perhaps a truly comparable compute budget should budget for all the other experiments performed along the way, or even for all the other researchers in the field conducting similar research. We discuss whether there is evidence for a GPU budget bottleneck in Could scaling up spending improve AI results?. Meanwhile, inference budgets were set at 10 billion tokens (sum of input, output, and reasoning), a limit set by comparison to our previous large-scale benchmarks.
Agents were instructed to submit a prose write-up of their solution, its codebase, and the checkpoints that corroborate their claims, as stored on Modal. We also stored copies of the submitted codebase and resulting job checkpoints at the time that any job was trained, for later corroboration of models’ claims.