How to read the overall score
The score shows a model's relative position across the current eligible evidence set. It is useful for screening, but it is not IQ and a score of 97 is not seven percent more capable than a score of 90.
Scoring & Data
Start with the aggregate view, then review capability rankings, original sources, price, and your own use case.
Different sources are made comparable while their original meaning remains visible.
The score shows a model's relative position across the current eligible evidence set. It is useful for screening, but it is not IQ and a score of 97 is not seven percent more capable than a score of 90.
Missing evaluations are not treated as zero. Coverage is shown alongside the score, and limited evidence reduces ranking confidence so models cannot lead on just a few results.
Closely related evaluations are organized into benchmark families so one capability does not gain extra influence simply because it has many similar tests.
A model can use different names across sources. Confirmed aliases map to a specific model and version; records that cannot be matched reliably remain separate.
Coverage shows how much relevant capability evidence is available. Broad coverage usually makes a rank more stable; limited coverage is better read as an early signal.
Coding, reasoning, math, knowledge, and instruction following each sort by their relevant capability measures. They are not renamed copies of the overall ranking.
Status describes evidence confidence, not whether a model is useful.
Model selection does not have one universal number. Non-capability signals deserve their own comparisons.
Output speed, catalog latency, observed TTFT, throughput, and reliability are published in a standalone runtime view. None of these fields are folded into capability scores.