Measuring Typed Decision AI: What S1MB Actually Tests, and Who Leads It
datasciencebenchmarkllmstatistics
TL;DR
A large funding round put typed, machine-native AI in the headlines. The harder question for anyone evaluating it is measurement: if a model returns a decision instead of prose, how do you score it, and who is actually good at it? This is a short look at one benchmark built for exactly that, the System One Mosaic Benchmark (S1MB), and the current number one result among 102 models.
What S1MB measures
S1MB evaluates fast, intuitive decision quality across a mosaic of typed tasks. Each item gives a condition and a set of candidate choices, and the model must pick and score the right one. It is a decision benchmark, not a text-generation benchmark, which is the correct frame when the output has a type rather than being free text.
That distinction matters for measurement. A generation benchmark scores strings, so it needs fuzzy matching or a judge model. A decision benchmark scores a typed answer against a typed key, which is clean, reproducible, and not clouded by a second model's opinion.
The leading result
Among 102 models, Darwin-27B-ZTC-v2 is number one overall, with a typed-decision accuracy of 0.743 zero-shot.
| Rank | Models compared | Metric | Result |
|---|---|---|---|
| Number 1 | 102 | Typed-decision accuracy (zero-shot) | 0.743 |
What makes the result reproducible is the inference method. The decision is produced in a single forward pass with zero generated tokens, by applying a calibrated probe to the model's final hidden state. No sampling means the same input gives the same score, which is what you want in a measurement.
How to check it
The leaderboard is public, and the method and model are open under Apache-2.0.
- S1MB leaderboard and number one model: github.com/final-bench/s1mb
- ZTC method and inference code: github.com/final-bench/ztc
FAQ
What does S1MB measure? The System One Mosaic Benchmark measures typed-decision quality: given a condition and candidate choices, the model picks and scores the right one. It is a decision benchmark, not a text-generation benchmark.
Which model leads S1MB? Darwin-27B-ZTC-v2 is number one overall among 102 models, with 0.743 typed-decision accuracy zero-shot.
Why are typed decisions easier to measure than chat answers? A typed answer is scored against a typed key, which is reproducible, while free text usually needs fuzzy matching or a judge model.
Is the measurement reproducible? Yes. The decision is deterministic, produced in one forward pass with zero generated tokens, and both the leaderboard and the model are public.
A measurement note from QuantID.