Benchmark category
Math evaluations, model coverage and rankings. Each benchmark keeps its original scale and methodology.
Data as of 2026-08-11
Select a benchmark to inspect model results, evidence fields and scoring direction.
| MMLU-Pro | MMLU-Pro | text | Score | 129 | featured | B | Yes |
| AIME 2025 | AIME 2025 | text | Score | 114 | featured | C | Yes |
| MMLU | MMLU | text | Score | 100 | featured | B | Yes |
| Humanity's Last Exam | Humanity's Last Exam | multimodal | Score | 93 | featured | B | Yes |
| MATH | MATH | text | Score | 71 | featured | B | Yes |
| AIME 2024 | AIME 2024 | text | Score | 53 | featured | B | Yes |
| MMMLU | MMMLU | text | Score | 49 | featured | B | Yes |
| GSM8k | GSM8k | text | Score | 48 | featured | B | Yes |
| MMLU-Redux | MMLU-Redux | text | Score | 48 | featured | B | Yes |
| MathVista | MathVista | multimodal | Score | 39 | featured | B | Yes |
| LiveBench | LiveBench | text | Score | 38 | featured | B | Yes |
| SuperGPQA | SuperGPQA | text | Score | 34 | featured | B | Yes |
| HMMT 2025 | HMMT 2025 | text | Score | 33 | featured | B | Yes |
| MATH-500 | MATH-500 | text | Score | 32 | featured | C | Yes |
| MathVision | MathVision | multimodal | Score | 32 | featured | B | Yes |
| MMLU-ProX | MMLU-ProX | text | Score | 32 | featured | B | Yes |
| MGSM | MGSM | text | Score | 31 | featured | B | Yes |
| DROP | DROP | text | Score | 30 | featured | B | Yes |
| HMMT25 | HMMT25 | text | Score | 25 | featured | B | Yes |
| MathVista-Mini | MathVista-Mini | multimodal | Score | 23 | featured | B | Yes |
| PolyMATH | PolyMATH | multimodal | Score | 23 | featured | B | Yes |
| BIG-Bench Hard | BIG-Bench Hard | text | Score | 21 | featured | B | Yes |
| IMO-AnswerBench | IMO-AnswerBench | text | Score | 19 | featured | B | Yes |
| SciCode | SciCode | text | Score | 19 | featured | B | Yes |
| AIME 2026 | AIME 2026 | text | Score | 18 | featured | B | Yes |
| FrontierMath | FrontierMath | text | Score | 17 | featured | B | Yes |
| CodeForces | CodeForces | text | Normalized rating | 16 | featured | B | Yes |
| LiveBench 20241125 | LiveBench 20241125 | text | Score | 14 | featured | B | Yes |
| HiddenMath | HiddenMath | text | Score | 13 | featured | B | Yes |
| BBH | BBH | text | Score | 12 | featured | B | Yes |
| HMMT Feb 26 | HMMT Feb 26 | text | Score | 11 | featured | B | Yes |
| AGIEval | AGIEval | text | Score | 10 | featured | B | Yes |
Leading results from high-coverage benchmarks with at least two models.
How this category is assembled on llmboard.ai.
This page groups benchmarks whose primary or display category matches math. It does not average incompatible metrics into a new category score.
Open an individual benchmark to inspect score direction, evidence level, participant count and model results.
Benchmarks are grouped by their primary and display categories.