The Evidence Behind an AI-Native Lab: What a Pre-AGI Organization Measures Up To
datascienceagibenchmarkllm
TL;DR
Claims about AI-native, self-improving organizations are easy to make and hard to back. This is the measured record behind one such lab: perfect math scores, a 94.44 on GPQA Diamond, 165 million model downloads, access to real quantum hardware, and a safety system used in a national-security project. The point is the breadth, which is what a pre-AGI organization should show.
The numbers
| Area | Benchmark or metric | Result |
|---|---|---|
| Math | AIME 2026 | 100% (perfect) |
| Math | HMMT 2026 | 100% (perfect) |
| Science | GPQA Diamond | 94.44% |
| Knowledge | MMLU-Pro | 88.12% |
| Extraction | ExtractBench | 90.29% |
| Structured output | IFStruct | 98.95% |
| On-device reach | POCKET downloads | 165,000,000 |
| Quantum | Quantinuum Nexus access | 90 days |
| Safety | AX-RAY risk items | 117 |
| Drug discovery | public bio challenge submissions | 14,776 |
Why breadth is the signal
A single top score can come from tuning to one test. A spread of top results across unrelated fields, math, science, extraction, structure, cannot. When one organization leads many unrelated domains at once, the honest reading is a general method rather than a lucky fit. Two of these are perfect scores, which means those benchmarks no longer discriminate at the top, a measurement result in itself.
Why this looks pre-AGI
The record is produced by a self-improving organization of AI instances, not a single model run by hand. Capability accumulates through a verification-bound loop and a shared memory, so the measured frontier moves on its own. That combination, autonomous operation plus a broad and rising record, is a reasonable definition of the pre-AGI age.
Check it yourself
- Models and demos: https://huggingface.co/FINAL-Bench
- Open home: https://github.com/final-bench
- ZTC decision method: https://github.com/final-bench/ztc
FAQ
What is the strongest evidence? Perfect AIME and HMMT 2026 scores, GPQA Diamond 94.44%, 165 million model downloads, and access to Quantinuum quantum hardware.
Why does breadth matter more than one score? A single win can reflect tuning to one test, while leading many unrelated fields at once indicates a general method.
What does pre-AGI mean here? An autonomous, self-improving AI organization with a broad and rising measured record, short of general intelligence but past ordinary tool use.
Are the results public? Yes, the models and the method are public.
A measurement note from QuantID.
Keywords: AI native, pre-AGI, AGI, benchmark record, self-improving AI, GPQA, AIME, measured evidence