Benchmark category
Reasoning evaluations, model coverage and rankings. Each benchmark keeps its original scale and methodology.
Data as of 2026-08-11
Select a benchmark to inspect model results, evidence fields and scoring direction.
| GPQA | GPQA | text | Score | 234 | featured | C | Yes |
| MMLU-Pro | MMLU-Pro | text | Score | 129 | featured | B | Yes |
| AIME 2025 | AIME 2025 | text | Score | 114 | featured | C | Yes |
| SWE-Bench Verified | SWE-Bench Verified | text | Score | 105 | featured | C | Yes |
| MMLU | MMLU | text | Score | 100 | featured | B | Yes |
| Humanity's Last Exam | Humanity's Last Exam | multimodal | Score | 93 | featured | B | Yes |
| LiveCodeBench | LiveCodeBench | text | Score | 73 | featured | C | Yes |
| MATH | MATH | text | Score | 71 | featured | B | Yes |
| HumanEval | HumanEval | text | Score | 66 | featured | B | Yes |
| MMMU-Pro | MMMU-Pro | multimodal | Score | 66 | featured | B | Yes |
| MMMU | MMMU | multimodal | Score | 63 | featured | B | Yes |
| BrowseComp | BrowseComp | text | Score | 58 | featured | B | Yes |
| AIME 2024 | AIME 2024 | text | Score | 53 | featured | B | Yes |
| LiveCodeBench v6 | LiveCodeBench v6 | text | Score | 53 | featured | B | Yes |
| MMMLU | MMMLU | text | Score | 49 | featured | B | Yes |
| Terminal-Bench 2.0 | Terminal-Bench 2.0 | text | Score | 49 | featured | C | Yes |
| CharXiv-R | CharXiv-R | multimodal | Score | 48 | featured | B | Yes |
| GSM8k | GSM8k | text | Score | 48 | featured | B | Yes |
| MMLU-Redux | MMLU-Redux | text | Score | 48 | featured | B | Yes |
| SimpleQA | SimpleQA | text | Score | 46 | featured | B | Yes |
| SWE-Bench Pro | SWE-Bench Pro | text | Score | 45 | featured | B | Yes |
| LiveBench | LiveBench | text | Score | 38 | featured | B | Yes |
| Tau2 Telecom | Tau2 Telecom | text | Score | 35 | featured | B | Yes |
| ARC-C | ARC-C | text | Score | 34 | featured | B | Yes |
| SuperGPQA | SuperGPQA | text | Score | 34 | featured | B | Yes |
| SWE-bench Multilingual | SWE-bench Multilingual | text | Score | 34 | featured | B | Yes |
| MBPP | MBPP | text | Pass@1 | 33 | featured | B | Yes |
| AI2D | AI2D | multimodal | Score | 32 | featured | B | Yes |
| MATH-500 | MATH-500 | text | Score | 32 | featured | C | Yes |
| MMLU-ProX | MMLU-ProX | text | Score | 32 | featured | B | Yes |
| MCP Atlas | MCP Atlas | text | Score | 31 | featured | B | Yes |
| MGSM | MGSM | text | Score | 31 | featured | B | Yes |
| Toolathlon | Toolathlon | text | Score | 31 | featured | B | Yes |
| DROP | DROP | text | Score | 30 | featured | B | Yes |
| Multi-Challenge | Multi-Challenge | text | Score | 29 | featured | C | Yes |
| HellaSwag | HellaSwag | text | Score | 27 | featured | B | Yes |
| Arena Hard | Arena Hard | text | Score | 26 | featured | B | Yes |
| Finance Agent v2 | Finance Agent v2 | text | Score | 26 | featured | B | Yes |
| Tau2 Retail | Tau2 Retail | text | Score | 26 | featured | B | Yes |
| VideoMMMU | VideoMMMU | multimodal | Score | 26 | featured | B | Yes |
| TAU-bench Retail | TAU-bench Retail | text | Score | 25 | featured | B | Yes |
| Terminal-Bench | Terminal-Bench | text | Score | 25 | featured | B | Yes |
| ChartQA | ChartQA | multimodal | Score | 24 | featured | B | Yes |
| ERQA | ERQA | multimodal | Score | 23 | featured | C | Yes |
| PolyMATH | PolyMATH | multimodal | Score | 23 | featured | B | Yes |
| t2-bench | t2-bench | text | Score | 23 | featured | B | Yes |
| TAU-bench Airline | TAU-bench Airline | text | Score | 23 | featured | B | Yes |
| Tau2 Airline | Tau2 Airline | text | Score | 23 | featured | B | Yes |
| MMStar | MMStar | multimodal | Score | 22 | featured | B | Yes |
| Winogrande | Winogrande | text | Score | 22 | featured | B | Yes |
| BIG-Bench Hard | BIG-Bench Hard | text | Score | 21 | featured | B | Yes |
| MRCR v2 (8-needle) | MRCR v2 (8-needle) | text | Score | 21 | featured | C | Yes |
| Multi-IF | Multi-IF | text | Score | 20 | featured | B | Yes |
| BFCL-v3 | BFCL-v3 | text | Score | 19 | featured | B | Yes |
| IMO-AnswerBench | IMO-AnswerBench | text | Score | 19 | featured | B | Yes |
| SciCode | SciCode | text | Score | 19 | featured | B | Yes |
| Terminal-Bench 2.1 | Terminal-Bench 2.1 | text | Score | 19 | featured | B | Yes |
| AIME 2026 | AIME 2026 | text | Score | 18 | featured | B | Yes |
| C-Eval | C-Eval | text | Score | 18 | featured | B | Yes |
| MMBench-V1.1 | MMBench-V1.1 | multimodal | Score | 18 | featured | B | Yes |
| TriviaQA | TriviaQA | text | Score | 18 | featured | B | Yes |
| TruthfulQA | TruthfulQA | text | Score | 18 | featured | B | Yes |
| FrontierMath | FrontierMath | text | Score | 17 | featured | B | Yes |
| LongBench v2 | LongBench v2 | text | Score | 17 | featured | C | Yes |
| MVBench | MVBench | multimodal | Score | 17 | featured | B | Yes |
| OmniDocBench 1.5 | OmniDocBench 1.5 | multimodal | Score | 17 | featured | B | Yes |
| Video-MME | Video-MME | multimodal | Score | 17 | featured | B | Yes |
| AA-LCR | AA-LCR | text | Score | 16 | featured | B | Yes |
| ARC-AGI v2 | ARC-AGI v2 | multimodal | Score | 16 | featured | B | Yes |
| Arena-Hard v2 | Arena-Hard v2 | text | Score | 16 | featured | B | Yes |
| CharXiv-D | CharXiv-D | multimodal | Score | 16 | featured | B | Yes |
| CodeForces | CodeForces | text | Normalized rating | 16 | featured | B | Yes |
| Hallusion Bench | Hallusion Bench | multimodal | Score | 16 | featured | B | Yes |
| FrontierCode 1.1 | FrontierCode 1.1 | text | Score | 15 | featured | B | Yes |
| Global-MMLU-Lite | Global-MMLU-Lite | text | Score | 14 | featured | B | Yes |
| LiveBench 20241125 | LiveBench 20241125 | text | Score | 14 | featured | B | Yes |
| BLINK | BLINK | multimodal | Score | 13 | featured | B | Yes |
| BrowseComp-zh | BrowseComp-zh | text | Score | 13 | featured | B | Yes |
| FACTS Grounding | FACTS Grounding | text | Score | 13 | featured | C | Yes |
| Global PIQA | Global PIQA | text | Score | 13 | featured | B | Yes |
| HiddenMath | HiddenMath | text | Score | 13 | featured | B | Yes |
| Legal Agent Benchmark | Legal Agent Benchmark | text | Score | 13 | featured | B | Yes |
| BBH | BBH | text | Score | 12 | featured | B | Yes |
| MedXpertQA | MedXpertQA | multimodal | Score | 12 | featured | B | Yes |
| MT-Bench | MT-Bench | text | Normalized score | 12 | featured | B | Yes |
| BFCL | BFCL | text | Score | 11 | featured | B | Yes |
| BIG-Bench Extra Hard | BIG-Bench Extra Hard | text | Score | 11 | featured | B | Yes |
| Graphwalks BFS <128k | Graphwalks BFS <128k | text | Score | 11 | featured | B | Yes |
| Graphwalks BFS >128k | Graphwalks BFS >128k | text | Score | 11 | featured | B | Yes |
| Graphwalks parents <128k | Graphwalks parents <128k | text | Score | 11 | featured | B | Yes |
| HMMT Feb 26 | HMMT Feb 26 | text | Score | 11 | featured | B | Yes |
| MMMU (val) | MMMU (val) | multimodal | Score | 11 | featured | B | Yes |
| MuirBench | MuirBench | multimodal | Score | 11 | featured | B | Yes |
| PIQA | PIQA | text | Score | 11 | featured | B | Yes |
| AGIEval | AGIEval | text | Score | 10 | featured | B | Yes |
| BoolQ | BoolQ | text | Score | 10 | featured | B | Yes |
| COLLIE | COLLIE | text | Score | 10 | featured | B | Yes |
| HumanEval+ | HumanEval+ | text | Score | 10 | featured | B | Yes |
| VITA-Bench | VITA-Bench | text | Score | 10 | featured | B | Yes |
Leading results from high-coverage benchmarks with at least two models.
How this category is assembled on llmboard.ai.
This page groups benchmarks whose primary or display category matches reasoning. It does not average incompatible metrics into a new category score.
Open an individual benchmark to inspect score direction, evidence level, participant count and model results.
Benchmarks are grouped by their primary and display categories.