Compare current model versions across independent benchmarks, community preference data, and official pricing. Scores show relative capability, not model IQ or an absolute performance gap.
| 01 | AN | 97.4 | Leading | Reasoning, Knowledge | $5 | $25 | Moderate 60% | ||
| 02 | AN | 95.8 | Leading | Knowledge, Coding | $10 | $50 | Limited 40% | ||
| 03 | OP | 95.1 | Leading | Math, Reasoning | $5 | $30 | High 80% | ||
| 04 | AN | 93.6 | Leading | Math, Reasoning | N/A | N/A | High 80% | N/A | |
| 05 | MA | 93.0 | Leading | Reasoning, Knowledge | $3 | $15 | Moderate 60% | ||
| 06 | AC | 89.3 | Strong | Instruction Following, Reasoning | $2 | $6 | High 80% | ||
| 07 | ME | 88.3 | Strong | Knowledge, Coding | $1.25 | $4.25 | Limited 40% | ||
| 08 | AN | 87.9 | Strong | Reasoning, Knowledge | $5 | $25 | Moderate 60% | ||
| 09 | OP | 87.6 | Strong | Reasoning, Math | $2 | $12 | High 80% | ||
| 10 | AN | 84.2 | Strong | Reasoning, Knowledge | $5 | $25 | Moderate 60% | ||
| 11 | AN | 84.0 | Strong | Knowledge, Coding | $2 | $10 | Moderate 60% | ||
| 12 | OP | 84.0 | Strong | Reasoning, Knowledge | $5 | $30 | High 80% | ||
| 13 | ME | 82.3 | Provisional | Coding | $1.25 | $4.25 | Limited 20% | ||
| 14 | XA | 82.2 | Provisional | Reasoning, Coding | $2 | $6 | Limited 40% | ||
| 15 | AN | 79.6 | Strong | Math, Knowledge | $5 | $25 | High 80% | ||
| 16 | GO | 79.3 | Provisional | Coding | $1.5 | $7.5 | Limited 20% | ||
| 17 | ZA | 79.3 | Strong | Reasoning, Math | $1.4 | $4.4 | High 80% | ||
| 18 | GO | 79.2 | Strong | Reasoning, Knowledge | $1.5 | $9 | Moderate 60% | ||
| 19 | AC | 78.8 | Strong | Math, Instruction Following | $2.5 | $7.5 | High 100% | ||
| 20 | BY | 78.7 | Strong | Reasoning, Knowledge | N/A | N/A | High 80% | ||
| 21 | DE | 78.3 | Provisional | Coding | $0.14 | $0.28 | Limited 20% | ||
| 22 | GO | 77.5 | Strong | Reasoning, Knowledge | $2 | $12 | Moderate 60% | ||
| 23 | OP | 77.1 | Strong | Reasoning, Math | $2.5 | $15 | High 80% | ||
| 24 | OP | 76.4 | Strong | Reasoning, Math | $0.20 | $1.2 | High 80% | ||
| 25 | OP | 75.8 | Provisional | Knowledge, Math | $30 | $180 | Limited 40% | ||
| 26 | MA | 75.5 | Strong | Reasoning, Math | $0.95 | $4 | High 80% | ||
| 27 | DE | 73.8 | Competitive | Reasoning, Knowledge | $0.435 | $0.87 | High 80% | ||
| 28 | AC | 73.8 | Competitive | Instruction Following, Reasoning | $0.50 | $3 | High 100% | ||
| 29 | ME | 73.2 | Competitive | Knowledge, Reasoning | N/A | N/A | Moderate 60% | ||
| 30 | BY | 73.1 | Competitive | Reasoning, Knowledge | N/A | N/A | High 80% |
Each leaderboard sorts by its own capability measure, not by the overall score.
A focused set of broad-coverage, current, score-eligible evaluations.
These selection signals are shown together but remain separate. Price and runtime never change the capability score.
Only models with positive official pay-as-you-go token prices are shown.
Each scored version uses its fastest tracked model-provider record. Output Speed is generated tokens received per second.
Runtime data is provider-specific and remains independent from capability, evidence coverage, and price. Open the full runtime page for catalog latency, TTFT, observed throughput, and reliability.