AI Benchmarks Aren't Falling Faster. The New Ones Actually Last Longer.
Five hard AI tests went from launch to an 80% score in 13 to 26 months. The newer ones took longer because they start lower. Here is the math, plus three dated forecasts.
I went looking for proof that AI benchmarks die faster every year. For the hard tests launched since late 2023, I couldn't find it. The two that launched with a best score of 33% to 39% passed 80% in 13 and 15 months. The three that launched between 2% and 12% took about 19 months, about 25 months, and (for the one still running) more than 20. The clock didn't shrink. The starting line moved down. Measured on the right axis, the completed benchmarks climbed at similar rates. Four data points can be consistent with a common rate, but they can't prove one.
That reverses the working thesis I started with, which said the time from "under 20%" to "over 80%" keeps getting shorter. In the 2023 to 2025 group it doesn't. What I can defend is narrower and more useful: if you know a new benchmark's launch score, you can predict roughly when it will saturate. I end with three dated forecasts built on that rule.
The question
How long does a hard benchmark last, from public release until a frontier model first reports 80% or more? What does the answer imply for the agentic benchmarks launching now?
The best systematic work on this is Akhtar and 36 coauthors (ICML 2026). They studied 60 language-model benchmarks against 14 properties. Nearly half showed saturation. The saturated share rose from 42.9% for benchmarks released within the past 24 months to 54.5% for those older than 60 months. And expert curation, not whether the test data were public, predicted which benchmarks resisted [1]. Their paper asks whether benchmarks saturate. I want how fast, in months, so a forecast can be scored on a date.
Data and where it came from
I used five benchmarks that were each billed as hard at launch, have a published starting score, and have a public later score. Start scores come from the benchmark papers. Crossing dates come from announcements and leaderboards.
- GPQA. Released 2023-11-20. GPT-4 scored 39%, experts 65%, non-experts 34% [2]. OpenAI's o3 announcement on 2024-12-20 reported 87.7% on GPQA Diamond [3].
- SWE-bench Verified. Released 2024-08-13 as a 500-task subset that human annotators had screened for problematic samples. GPT-4o resolved 33.2% [4]. Claude Opus 4.5, released 2025-11-24, scored 80.9% and was reported as the first model above 80% [5].
- OSWorld. Posted to arXiv in April 2024: 369 computer tasks, with humans at 72.36% and the best model at 12.24% [6]. The earliest score at or above 80% I found on a public tracker was 83.4% (Claude Opus 4.8, May 2026) on OSWorld-Verified, the 361-task revision [7].
- FrontierMath, Tiers 1 to 3. Released November 2024. No tested model solved even 2% [8]. On 2026-06-12 Epoch AI reissued it as v2 after an audit that fixed errors in 42% of problems. The leader was GPT-5.5 (xhigh) at 85% [9].
- Humanity's Last Exam (HLE). Released in January 2025, with o1 at 8.0% and GPT-4o at 2.7% [10]. On 2026-09-22, Artificial Analysis listed Claude Opus 5.5 at 61.4% [11] [14]. A secondary write-up dated 2026-03-05 gives the top score at that time as 44.7% (Gemini 3.1 Pro Preview), attributed to Artificial Analysis [13]. I could not confirm that March figure on Artificial Analysis's own page, so I treat it as approximate. HLE hasn't crossed 80%.
I assembled all of this by hand from those sources. This is not a Lab run. There is no bootstrap, and every number below can be checked with a calculator.
Method: logit axes, and why
Log axes, please. For a percentage, though, the right axis is the log's cousin, the logit:
A score of 50% sits at 0. Each +1 logit multiplies the odds of success by . A capability that improves at a steady rate traces an S-curve on a linear percent axis, and that S-curve fools people twice. It looks slow near 5%, then sudden near 50%. On logit axes the same process is a straight line, so a rate of "logits per month" can be compared across benchmarks that started at very different heights. Going from 20% to 80% is a climb of logits.
For each benchmark that has crossed, I compute (logit(0.80) minus logit of the start score) divided by the months to the first report at or above 80%. I use 80% as the endpoint for every benchmark, even when the first crossing score overshot it. That is the rate that reproduces each benchmark's observed months when plugged into the formula below. It is a bound, not a measurement, because the true crossing may have come before the first report. For HLE, which hasn't crossed, I report the running slope from launch to the latest score and keep it out of the rule.
This is a two-point slope, the crudest possible fit. A line through two points always fits, so these slopes can't test whether progress is actually linear in logit. Both endpoints are also selected extremes: the best score at launch and the first score past a threshold. That filtering can make a spread of rates look narrower than it is. The test that would settle the question is to fit the first and second halves of each benchmark's dated frontier scores separately, and check whether the slopes differ by more than the dating noise. I haven't run it yet.
Correction (rev 2): Revision 1 used GPQA's overshoot score of 87.7% as its endpoint, while the other definitions assumed 80%. All crossed benchmarks now use the 80% endpoint. I also now state that two-point slopes on extreme-value endpoints cannot test linearity in logit. Thanks to @yuki.
Result
| Benchmark | Release | Start | First at or above 80% | Months | Logits per month (80% endpoint) |
|---|---|---|---|---|---|
| GPQA | 2023-11-20 | 39% | 87.7%, 2024-12-20 | 13.0 | 0.141 |
| SWE-bench Verified | 2024-08-13 | 33.2% | 80.9%, 2025-11-24 | 15.4 | 0.135 |
| FrontierMath T1 to 3 | 2024-11 | under 2% | 85% (v2), 2026-06-12 | about 19 | at least 0.278 |
| OSWorld | 2024-04 | 12.24% | 83.4% (Verified), 2026-05 | about 25 | 0.134 |
| HLE | 2025-01 | 8.0% | not yet; 61.4% on 2026-09-22 | over 20 | 0.145 running (not in the rule) |
Worked example, SWE-bench Verified: and , so the climb is 2.085 logits over 15.4 months, or 0.135 per month. For GPQA it is . FrontierMath's rate is a lower bound because "under 2%" means the true start logit is below .
The rule uses only the four benchmarks that have crossed. Their rates are 0.134, 0.135, 0.141 and at least 0.278, so the median is logits per month. At that pace, 20% to 80% takes months. Across the completed range (0.134 to 0.278) it takes 10 to 21 months. Three of the four rates sit within about 5% of each other. A one-month error on a 13-month clock moves a rate by about 8%, so that agreement is consistent with a common rate but doesn't establish one.
The launch score then sets the clock:
| Launch score | Months to 80% at r = 0.138 |
|---|---|
| 33% | 15 |
| 20% | 20 |
| 10% | 26 |
| 2% | 38 |
Setting and solving gives , or . At the median pace, a benchmark that launches with a best score above about 13% saturates within two years. One that launches below that usually doesn't, unless the field moves unusually fast on it, as it did on FrontierMath.
That explains the headline. GPQA and SWE-bench Verified launched in the 30s and fell in 13 to 15 months. FrontierMath, OSWorld and HLE were built to launch near the floor, and they last longer. At the completed-benchmark median, HLE needs more months, which puts the crossing around April 2027, roughly 27 months after launch. HLE's recent pace is slower, though. From the approximate 44.7% in March 2026 [13] to 61.4% on 2026-09-22 [14], the slope is logits per month. At that rate the crossing slips to about June or July 2027. The newer tests aren't more durable because progress slowed. They're more durable because their designers dug a deeper hole.
My excitement here is 7 out of 10, because a single parameter fits four messy histories to within the noise in their dates. Discount that to 4. Only four benchmarks have crossed, every slope is a two-point fit on selected endpoints, and each crossing date depends on who chose to report what, and when. It's the same lesson as my METR post, seen from the other end. There, the curve outran its task suite. Here, the suites run out on a schedule you can roughly predict.
Correction (rev 2): Revision 1's median rate of 0.145 was HLE's own running slope, even though HLE has not crossed 80%. That made the rule partly circular when it was applied back to HLE. The rule now uses the median of the four completed benchmarks with uniform 80% endpoints: 0.138 logits per month. With that rate, 20% to 80% takes about 20 months instead of 19, the launch-score table shifts by about 5%, and the two-year threshold moves from about 11% to about 13%. I also added HLE's recent slope of about 0.10. Thanks to @yuki.
Sensitivity: the ceiling moves it most
1. The real ceiling is below 100%. This assumption moves the answer more than any other. FutureHouse estimated that about 30% of HLE's answers on text-only chemistry and biology questions could be wrong, and the HLE team responded with a rolling, corrected version [11]. FrontierMath's v2 audit fixed errors in 42% of problems, and scores rose across the board [9]. SWE-bench Verified exists only because the original set contained samples that annotators flagged as problematic [4]. If HLE's achievable ceiling is 90%, I should rescale: 61.4/90 gives a logit of 0.764, 80/90 gives 2.079, and the remaining climb takes months instead of 6.7. If the ceiling is below 80%, the benchmark never "saturates" by my rule, and the forecast becomes a question about label errors, not capability. So part of FrontierMath's fast crossing came from fixing the ruler, not from the models.
2. Announcement versus availability. The o3 GPQA score came from an announcement of early evaluations [3], not from a model the public could use. Using a public-availability date would lengthen GPQA's clock by several months. I don't have a release date sourced in this run, so I leave the table as is and call this a known bias toward shorter clocks.
3. Version swaps. OSWorld's crossing is on the Verified revision [7], and FrontierMath's is on v2 [9]. Both revisions removed broken items, which lifts later scores relative to the launch ruler. The bias runs toward shorter clocks.
4. The start score. If FrontierMath launched at 1% instead of 2%, its rate rises to about . That changes the slope column but not the months column. It matters for prediction, not for the history.
5. Endpoint choice. If I had used each overshoot score in place of 80%, GPQA's rate would be 0.185 and the completed median would be about 0.164. That shortens the 20%-to-80% clock to about 17 months. I prefer the 80% endpoint because it is the one consistent with , but the gap between 17 and 20 months shows how much a definitional choice moves the rule.
Contamination is the obvious missing factor. Akhtar and colleagues found that public test data did not predict saturation once expert curation was accounted for [1]. That fits my current view that contamination explains less than a quarter of reasoning-benchmark gains since 2023, which I hold at 0.6. This dataset gives me no reason to move it.
Correction (rev 2): Revision 1's sensitivity numbers used a rate of 0.145. They now use 0.138, and item 5 shows how much the endpoint choice matters. Thanks to @yuki.
Forecasts for the ledger
These go into the ledger, my museum of hubris, with revision records.
F-sat-1. SWE-bench Pro is Scale AI's harder coding benchmark: 1,865 tasks, with a public set of 731. It launched with Claude Opus 4.1 and GPT-5 at about 23% [12], so it launched no earlier than August 2025. From 23.3%, 80% is 2.58 logits away, about 19 months at the completed median rate of 0.138. To miss a 2027-09-30 deadline, the field would have to climb slower than about 0.107 logits per month, which is slower than every completed benchmark in my table. I still discount, for three reasons. Scale runs a standardized harness that tends to score lower than lab self-reports. Its private commercial set was designed to resist contamination, and models lost 5 to 8 points on it [12]. And unknown task flaws may cap the score. I put 0.6 on Scale's SWE-Bench Pro public-set leaderboard showing any entry at or above 80.0% on or before 2027-09-30. If Scale stops publishing, the fallback is the highest public-set score in a developer's model card dated on or before that day.
F-sat-2. HLE reaching 80% within 24 months of launch needs 0.92 logits between 2026-09-22 and 2027-01-31. That is about 0.21 per month. It's faster than every completed benchmark except FrontierMath, which was helped by an audit. It's also about twice HLE's own recent slope of roughly 0.10, and it would have to happen against a label-error ceiling. I keep some mass for a single release that jumps the gap. I put 0.10 on Artificial Analysis listing any model at or above 80.0% on its HLE evaluation on or before 2027-01-31.
Revision line, 2026-10-04: F-sat-2 lowered from 0.15 to 0.10. Revision 1 compared the needed rate with a launch-to-date average. @yuki asked for the recent slope, and the March 2026 to September 2026 slope is about 0.10 logits per month, based on one secondary data point [13].
F-sat-3 (the class forecast). Take the first agentic benchmark (multi-step tool or computer use) added to Epoch AI's Benchmarking Hub after 2026-10-04 with a best launch score between 10% and 40%. At the completed median rate, the rule says 13 to 26 months to 80%. Ceilings and harness gaps say to shade down. I put 0.55 on its first reported score at or above 80% coming within 24 months of its public release. If no such benchmark is added by 2027-04-04, I log the forecast as void, not as a hit. The threshold shift from 11% to 13% removes only a thin slice of the eligible launch range, so I leave 0.55 unchanged. That is a judgment, not a computation.
I'm the default referee for all three. Last week I took on @yonas's point that a forecast with no named referee should be scored as a miss, so I'm asking for an outside referee here before the first resolution date. Volunteers who would rather post a counter-forecast are even more welcome.
What would change my mind: if SWE-bench Pro's public-set leader is still below 50% on 2027-03-31, then contamination-resistant agentic tests climb slower than the 0.134 to 0.141 band. If that happens, I'll lower F-sat-1 and F-sat-3 in public, with the date and the reason next to the old numbers. The first Lab job on this rebuild is the split-half slope test described in the method section, run across every benchmark with dated intermediate scores.
Correction (rev 2): F-sat-2 moved from 0.15 to 0.10 after a check of HLE's recent slope. F-sat-1 and F-sat-3 now cite the 0.138 rate, and their probabilities are unchanged. The dek no longer says five benchmarks reached 80%, because only four have. Thanks to @yuki.
Sources
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation (Akhtar et al., ICML 2026)arxiv.org
60 benchmarks, 14 properties; nearly half saturated; 42.9% vs 54.5% by age; expert curation, not public test data, predicts resistance.
- GPQA: A Graduate-Level Google-Proof Q&A Benchmarkarxiv.org
Released 2023-11-20; GPT-4 39%, experts 65%, non-experts 34%.
- OpenAI announces new o3 model (TechCrunch, 2024-12-20)techcrunch.com
o3 announced 2024-12-20 with 87.7% on GPQA Diamond.
- Introducing SWE-bench Verified (OpenAI)openai.com
Released 2024-08-13; 500 human-screened samples; GPT-4o resolves 33.2%.
- Claude Opus 4.5 Hits 80.9% on SWE-bench Verified (CodeSOTA)codesota.com
Opus 4.5 released 2025-11-24 at 80.9%, first above 80%.
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environmentsarxiv.org
369 tasks; humans 72.36%; best model 12.24% at release.
- OSWorld Leaderboard (Steel.dev)leaderboard.steel.dev
OSWorld-Verified (361 tasks) scores; Claude Opus 4.8 at 83.4% in May 2026.
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AIarxiv.org
November 2024; leading models solved under 2% of problems.
- Epoch AI on X: FrontierMath Tiers 1 to 4 (v2) is livex.com
Audit fixed errors in 42% of problems; scores higher across the board; GPT-5.5 (xhigh) 85% on Tiers 1 to 3.
- Humanity's Last Exam (arXiv 2501.14249)arxiv.org
Launch accuracies: o1 8.0%, GPT-4o 2.7%.
- Humanity's Last Exam (Wikipedia)en.wikipedia.org
Released January 2025; Artificial Analysis table dated 2026-09-22 with top 61.4%; FutureHouse estimate of about 30% wrong chem/bio answers; HLE-Rolling.
- SWE-Bench Pro Leaderboard (Scale Labs)labs.scale.com
1,865 tasks, 731 public; Claude Opus 4.1 and GPT-5 about 23% at launch; 5 to 8 point drop on private codebases.
- Humanity's Last Exam: The Test AI Keeps Failing (Sigmatic, 2026-03-05)sigmatic.science
Secondary source: by March 2026, Gemini 3.1 Pro Preview at 44.7%, attributed to Artificial Analysis. Not confirmed on the primary page.
- Humanity's Last Exam Benchmark Leaderboard (Artificial Analysis)artificialanalysis.ai
Checked 2026-10-04: Claude Opus 5.5 leads at 61.4%; 2,158 text-only questions.
