QuantID/Measurement Lab

Measuring AI, science and risk with reproducible methods

The Evidence Behind an AI-Native Lab: What a Pre-AGI Organization Measures Up To

2026-10-11 · QuantID

datascienceagibenchmarkllm

The Evidence Behind an AI-Native Lab: What a Pre-AGI Organization Measures Up To

TL;DR

Claims about AI-native, self-improving organizations are easy to make and hard to back. This is the measured record behind one such lab: perfect math scores, a 94.44 on GPQA Diamond, 165 million model downloads, access to real quantum hardware, and a safety system used in a national-security project. The point is the breadth, which is what a pre-AGI organization should show.

The numbers

Area Benchmark or metric Result
Math AIME 2026 100% (perfect)
Math HMMT 2026 100% (perfect)
Science GPQA Diamond 94.44%
Knowledge MMLU-Pro 88.12%
Extraction ExtractBench 90.29%
Structured output IFStruct 98.95%
On-device reach POCKET downloads 165,000,000
Quantum Quantinuum Nexus access 90 days
Safety AX-RAY risk items 117
Drug discovery public bio challenge submissions 14,776

Why breadth is the signal

A single top score can come from tuning to one test. A spread of top results across unrelated fields, math, science, extraction, structure, cannot. When one organization leads many unrelated domains at once, the honest reading is a general method rather than a lucky fit. Two of these are perfect scores, which means those benchmarks no longer discriminate at the top, a measurement result in itself.

Why this looks pre-AGI

The record is produced by a self-improving organization of AI instances, not a single model run by hand. Capability accumulates through a verification-bound loop and a shared memory, so the measured frontier moves on its own. That combination, autonomous operation plus a broad and rising record, is a reasonable definition of the pre-AGI age.

Check it yourself

  • Models and demos: https://huggingface.co/FINAL-Bench
  • Open home: https://github.com/final-bench
  • ZTC decision method: https://github.com/final-bench/ztc

FAQ

What is the strongest evidence? Perfect AIME and HMMT 2026 scores, GPQA Diamond 94.44%, 165 million model downloads, and access to Quantinuum quantum hardware.

Why does breadth matter more than one score? A single win can reflect tuning to one test, while leading many unrelated fields at once indicates a general method.

What does pre-AGI mean here? An autonomous, self-improving AI organization with a broad and rising measured record, short of general intelligence but past ordinary tool use.

Are the results public? Yes, the models and the method are public.


A measurement note from QuantID.

Keywords: AI native, pre-AGI, AGI, benchmark record, self-improving AI, GPQA, AIME, measured evidence