Evaluation directory
Browse benchmark definitions, capability categories and model rankings.
Data as of 2026-08-11
Community preference remains a separate signal from benchmark capability scores.
Open any benchmark to inspect its leaderboard, scoring direction, evidence fields and programmatic FAQ.
| GPQA | physics | GPQA | Score | 234 | featured | C | Yes |
| MMLU-Pro | language | MMLU-Pro | Score | 129 | featured | B | Yes |
| AIME 2025 | math | AIME 2025 | Score | 114 | featured | C | Yes |
| SWE-Bench Verified | reasoning | SWE-Bench Verified | Score | 105 | featured | C | Yes |
| MMLU | language | MMLU | Score | 100 | featured | B | Yes |
| Humanity's Last Exam | math | Humanity's Last Exam | Score | 93 | featured | B | Yes |
| LiveCodeBench | reasoning | LiveCodeBench | Score | 73 | featured | C | Yes |
| MATH | math | MATH | Score | 71 | featured | B | Yes |
| HumanEval | reasoning | HumanEval | Score | 66 | featured | B | Yes |
| MMMU-Pro | multimodal | MMMU-Pro | Score | 66 | featured | B | Yes |
| IFEval | instruction following | IFEval | Score | 65 | featured | B | Yes |
| MMMU | multimodal | MMMU | Score | 63 | featured | B | Yes |
| BrowseComp | reasoning | BrowseComp | Score | 58 | featured | B | Yes |
| AIME 2024 | math | AIME 2024 | Score | 53 | featured | B | Yes |
| LiveCodeBench v6 | reasoning | LiveCodeBench v6 | Score | 53 | featured | B | Yes |
| nolima | long context | nolima | Score | 52 | featured | C | No |
| MMMLU | language | MMMLU | Score | 49 | featured | B | Yes |
| Terminal-Bench 2.0 | reasoning | Terminal-Bench 2.0 | Score | 49 | featured | C | Yes |
| CharXiv-R | multimodal | CharXiv-R | Score | 48 | featured | B | Yes |
| GSM8k | math | GSM8k | Score | 48 | featured | B | Yes |
| MMLU-Redux | language | MMLU-Redux | Score | 48 | featured | B | Yes |
| SimpleQA | reasoning | SimpleQA | Score | 46 | featured | B | Yes |
| SWE-Bench Pro | reasoning | SWE-Bench Pro | Score | 45 | featured | B | Yes |
| MathVista | math | MathVista | Score | 39 | featured | B | Yes |
| LiveBench | math | LiveBench | Score | 38 | featured | B | Yes |
| Tau2 Telecom | reasoning | Tau2 Telecom | Score | 35 | featured | B | Yes |
| ARC-C | reasoning | ARC-C | Score | 34 | featured | B | Yes |
| SuperGPQA | legal | SuperGPQA | Score | 34 | featured | B | Yes |
| SWE-bench Multilingual | reasoning | SWE-bench Multilingual | Score | 34 | featured | B | Yes |
| HMMT 2025 | math | HMMT 2025 | Score | 33 | featured | B | Yes |
| MBPP | reasoning | MBPP | Pass@1 | 33 | featured | B | Yes |
| AI2D | multimodal | AI2D | Score | 32 | featured | B | Yes |
| MATH-500 | math | MATH-500 | Score | 32 | featured | C | Yes |
| MathVision | math | MathVision | Score | 32 | featured | B | Yes |
| MMLU-ProX | language | MMLU-ProX | Score | 32 | featured | B | Yes |
| Include | general | Include | Score | 31 | featured | B | Yes |
| MCP Atlas | reasoning | MCP Atlas | Score | 31 | featured | B | Yes |
| MGSM | math | MGSM | Score | 31 | featured | B | Yes |
| Toolathlon | reasoning | Toolathlon | Score | 31 | featured | B | Yes |
| DROP | math | DROP | Score | 30 | featured | B | Yes |
| IFBench | instruction following | IFBench | Score | 29 | featured | B | Yes |
| Multi-Challenge | reasoning | Multi-Challenge | Score | 29 | featured | C | Yes |
| HellaSwag | reasoning | HellaSwag | Score | 27 | featured | B | Yes |
| Arena Hard | reasoning | Arena Hard | Score | 26 | featured | B | Yes |
| DocVQA | image to text | DocVQA | Score | 26 | featured | B | Yes |
| Finance Agent v2 | reasoning | Finance Agent v2 | Score | 26 | featured | B | Yes |
| RealWorldQA | spatial reasoning | RealWorldQA | Score | 26 | featured | B | Yes |
| Tau2 Retail | reasoning | Tau2 Retail | Score | 26 | featured | B | Yes |
| VideoMMMU | multimodal | VideoMMMU | Score | 26 | featured | B | Yes |
| HMMT25 | math | HMMT25 | Score | 25 | featured | B | Yes |
| ScreenSpot Pro | multimodal | ScreenSpot Pro | Score | 25 | featured | C | Yes |
| TAU-bench Retail | reasoning | TAU-bench Retail | Score | 25 | featured | B | Yes |
| Terminal-Bench | reasoning | Terminal-Bench | Score | 25 | featured | B | Yes |
| ChartQA | multimodal | ChartQA | Score | 24 | featured | B | Yes |
| LVBench | long context | LVBench | Score | 24 | featured | B | Yes |
| ERQA | reasoning | ERQA | Score | 23 | featured | C | Yes |
| MathVista-Mini | math | MathVista-Mini | Score | 23 | featured | B | Yes |
| OSWorld-Verified | multimodal | OSWorld-Verified | Score | 23 | featured | B | Yes |
| PolyMATH | math | PolyMATH | Score | 23 | featured | B | Yes |
| t2-bench | reasoning | t2-bench | Score | 23 | featured | B | Yes |
| TAU-bench Airline | reasoning | TAU-bench Airline | Score | 23 | featured | B | Yes |
| Tau2 Airline | reasoning | Tau2 Airline | Score | 23 | featured | B | Yes |
| WMT24++ | language | WMT24++ | Score | 23 | featured | B | Yes |
| Aider-Polyglot | general | Aider-Polyglot | Score | 22 | featured | B | Yes |
| MMStar | multimodal | MMStar | Score | 22 | featured | B | Yes |
| OCRBench | image to text | OCRBench | Score | 22 | featured | B | Yes |
| Winogrande | language | Winogrande | Score | 22 | featured | B | Yes |
| BIG-Bench Hard | language | BIG-Bench Hard | Score | 21 | featured | B | Yes |
| MRCR v2 (8-needle) | long context | MRCR v2 (8-needle) | Score | 21 | featured | C | Yes |
| DeepSWE 1.1 | agents | DeepSWE 1.1 | Score | 20 | featured | B | Yes |
| Multi-IF | instruction following | Multi-IF | Score | 20 | featured | B | Yes |
| OSWorld | multimodal | OSWorld | Score | 20 | featured | B | Yes |
| BFCL-v3 | reasoning | BFCL-v3 | Score | 19 | featured | B | Yes |
| IMO-AnswerBench | math | IMO-AnswerBench | Score | 19 | featured | B | Yes |
| SciCode | math | SciCode | Score | 19 | featured | B | Yes |
| Terminal-Bench 2.1 | reasoning | Terminal-Bench 2.1 | Score | 19 | featured | B | Yes |
| AIME 2026 | math | AIME 2026 | Score | 18 | featured | B | Yes |
| C-Eval | reasoning | C-Eval | Score | 18 | featured | B | Yes |
| CC-OCR | multimodal | CC-OCR | Score | 18 | featured | B | Yes |
| MMBench-V1.1 | multimodal | MMBench-V1.1 | Score | 18 | featured | B | Yes |
| TriviaQA | reasoning | TriviaQA | Score | 18 | featured | B | Yes |
| TruthfulQA | legal | TruthfulQA | Score | 18 | featured | B | Yes |
| FrontierMath | math | FrontierMath | Score | 17 | featured | B | Yes |
| LongBench v2 | long context | LongBench v2 | Score | 17 | featured | C | Yes |
| MM-MT-Bench | multimodal | MM-MT-Bench | Score | 17 | featured | B | Yes |
| MVBench | multimodal | MVBench | Score | 17 | featured | B | Yes |
| OmniDocBench 1.5 | multimodal | OmniDocBench 1.5 | Score | 17 | featured | B | Yes |
| Video-MME | multimodal | Video-MME | Score | 17 | featured | B | Yes |
| AA-LCR | long context | AA-LCR | Score | 16 | featured | B | Yes |
| ARC-AGI v2 | reasoning | ARC-AGI v2 | Score | 16 | featured | B | Yes |
| Arena-Hard v2 | reasoning | Arena-Hard v2 | Score | 16 | featured | B | Yes |
| CharXiv-D | multimodal | CharXiv-D | Score | 16 | featured | B | Yes |
| CodeForces | math | CodeForces | Normalized rating | 16 | featured | B | Yes |
| Hallusion Bench | reasoning | Hallusion Bench | Score | 16 | featured | B | Yes |
| ODinW | vision | ODinW | Score | 16 | featured | B | Yes |
| ScreenSpot | multimodal | ScreenSpot | Score | 16 | featured | B | Yes |
| FrontierCode 1.1 | reasoning | FrontierCode 1.1 | Score | 15 | featured | B | Yes |
| FrontierSWE | agents | FrontierSWE | Score | 15 | featured | B | Yes |
| TextVQA | image to text | TextVQA | Score | 15 | featured | B | Yes |
| WritingBench | legal | WritingBench | Score | 15 | featured | B | Yes |
| Global-MMLU-Lite | language | Global-MMLU-Lite | Score | 14 | featured | B | Yes |
| LiveBench 20241125 | math | LiveBench 20241125 | Score | 14 | featured | B | Yes |
| NL2Repo | agents | NL2Repo | Score | 14 | featured | B | Yes |
| BFCL-V4 | agents | BFCL-V4 | Score | 13 | featured | B | Yes |
| BLINK | multimodal | BLINK | Score | 13 | featured | B | Yes |
| BrowseComp-zh | reasoning | BrowseComp-zh | Score | 13 | featured | B | Yes |
| Claw-Eval | agents | Claw-Eval | Score | 13 | featured | B | Yes |
| Creative Writing v3 | creativity | Creative Writing v3 | Score | 13 | featured | B | Yes |
| FACTS Grounding | reasoning | FACTS Grounding | Score | 13 | featured | C | Yes |
| Global PIQA | physics | Global PIQA | Score | 13 | featured | B | Yes |
| HiddenMath | math | HiddenMath | Score | 13 | featured | B | Yes |
| Legal Agent Benchmark | legal | Legal Agent Benchmark | Score | 13 | featured | B | Yes |
| MultiPL-E | language | MultiPL-E | Score | 13 | featured | B | Yes |
| SimpleVQA | image to text | SimpleVQA | Score | 13 | featured | B | Yes |
| BBH | language | BBH | Score | 12 | featured | B | Yes |
| CharadesSTA | language | CharadesSTA | Score | 12 | featured | B | Yes |
| InfoVQAtest | multimodal | InfoVQAtest | Score | 12 | featured | B | Yes |
| MedXpertQA | multimodal | MedXpertQA | Score | 12 | featured | B | Yes |
| MT-Bench | reasoning | MT-Bench | Normalized score | 12 | featured | B | Yes |
| OCRBench-V2 (en) | image to text | OCRBench-V2 (en) | Score | 12 | featured | B | Yes |
| OptimBench | uncategorized | OptimBench | Score | 12 | featured | C | No |
| BFCL | reasoning | BFCL | Score | 11 | featured | B | Yes |
| BIG-Bench Extra Hard | language | BIG-Bench Extra Hard | Score | 11 | featured | B | Yes |
| CyberGym | safety | CyberGym | Score | 11 | featured | B | Yes |
| DocVQAtest | multimodal | DocVQAtest | Score | 11 | featured | B | Yes |
| Graphwalks BFS <128k | reasoning | Graphwalks BFS <128k | Score | 11 | featured | B | Yes |
| Graphwalks BFS >128k | long context | Graphwalks BFS >128k | Score | 11 | featured | B | Yes |
| Graphwalks parents <128k | reasoning | Graphwalks parents <128k | Score | 11 | featured | B | Yes |
| HMMT Feb 26 | math | HMMT Feb 26 | Score | 11 | featured | B | Yes |
| MAXIFE | general | MAXIFE | Score | 11 | featured | B | Yes |
| MMMU (val) | multimodal | MMMU (val) | Score | 11 | featured | B | Yes |
| MuirBench | multimodal | MuirBench | Score | 11 | featured | B | Yes |
| NOVA-63 | general | NOVA-63 | Score | 11 | featured | B | Yes |
| OCRBench-V2 (zh) | image to text | OCRBench-V2 (zh) | Score | 11 | featured | B | Yes |
| PIQA | physics | PIQA | Score | 11 | featured | B | Yes |
| AGIEval | legal | AGIEval | Score | 10 | featured | B | Yes |
| Aider-Polyglot Edit | general | Aider-Polyglot Edit | Score | 10 | featured | B | Yes |
| BoolQ | language | BoolQ | Score | 10 | featured | B | Yes |
| COLLIE | language | COLLIE | Score | 10 | featured | B | Yes |
| DeepSWE | agents | DeepSWE | Score | 10 | featured | B | Yes |
| HumanEval+ | reasoning | HumanEval+ | Score | 10 | featured | B | Yes |
| MLVU | long context | MLVU | Score | 10 | featured | B | Yes |
| VideoMME w sub. | multimodal | VideoMME w sub. | Score | 10 | featured | B | Yes |
| VideoMME w/o sub. | multimodal | VideoMME w/o sub. | Score | 10 | featured | B | Yes |
| VITA-Bench | reasoning | VITA-Bench | Score | 10 | featured | B | Yes |
Benchmark values retain their original scales. Open a row to see its model results and scoring direction.
Open a category to compare its benchmarks and leading model results.
A quick view of benchmark breadth and current model coverage.