QuantID/Measurement Lab

Measuring AI, science and risk with reproducible methods

Eleven Number-One Records: Measuring a Model That Swept Math, Science, Law and Decisions

2026-10-11 · QuantID

datasciencebenchmarkllmstatistics

Eleven Number-One Records: Measuring a Model That Swept Math, Science, Law and Decisions

TL;DR

A single self-improving model family holds eleven public number-one benchmark records at once, across mathematics, science, law, structured output, and decisions. This is a measurement note: the eleven, their scores, and why the spread matters more than any single win.

The eleven records

# Benchmark Field Result
1 AIME 2026 Math 100% (perfect)
2 HMMT 2026 Math 100% (perfect)
3 GPQA Diamond Science 94.44%
4 MMLU-Pro Knowledge 88.12%
5 MMMU-Pro Multimodal 79.48%
6 LEXam Law 68.94%
7 LEXam-hard Law 45.72%
8 ExtractBench Extraction 90.29%
9 IFStruct Structured output 98.95%
10 MDPBench Decision process 83.65%
11 S1MB Decision engine Borda 89.58, Task Avg 66.46

Why the spread is the real signal

A single top score can come from tuning to one test. Eleven number-one records across unrelated fields cannot. Math, science, law, and typed decisions stress different capabilities, so holding the top of all of them at once is evidence of a general method rather than a lucky fit.

Two of the records are perfect scores, on AIME 2026 and HMMT 2026. When a benchmark saturates, the honest reading is that it no longer discriminates at the top, which is itself a measurement result and a reason to move evaluation to harder and more decision-oriented tests such as S1MB.

How these were produced

The family is self-improving: it runs a recursive loop that is tied to external verification rather than to its own opinion. Decisions, where the output is typed, are scored with a zero-token method, which keeps the measurement deterministic and reproducible.

Check it yourself

  • S1MB number one model: github.com/final-bench/s1mb
  • ZTC decision method: github.com/final-bench/ztc
  • Models: huggingface.co/FINAL-Bench

FAQ

How many number-one records does the family hold? Eleven public number-one benchmark records at once, across math, science, law, structured output, and decisions.

Which are perfect scores? AIME 2026 and HMMT 2026, both at 100%.

Why does holding many unrelated benchmarks matter? A single win can reflect tuning to one test. Leading many unrelated fields at once is evidence of a general method.

Are the results reproducible? Decision results use a deterministic zero-token method, and the leaderboards and models are public.


A measurement note from QuantID.