A more robust and challenging multi-task language understanding benchmark that extends MMLU by expanding multiple-choice options from 4 to 10, eliminating trivial questions, and focusing on reasoning-intensive tasks. Features over 12,000 curated questions across 14 domains and causes a 16-33% accuracy drop compared to original MMLU.
A more robust and challenging multi-task language understanding benchmark that extends MMLU by expanding multiple-choice options from 4 to 10, eliminating trivial questions, and focusing on reasoning-intensive tasks. Features over 12,000 curated questions across 14 domains and causes a 16-33% accuracy drop compared to original MMLU.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Family
MMLU-Pro
Modality
text
Primary category
language
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
mmlu-pro|llm-stats-current
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
FAQ
Common questions about MMLU-Pro.
Which model scores highest on MMLU-Pro?
Qwen3.7 Max is currently ranked first with 89.6%.
What does MMLU-Pro measure?
A more robust and challenging multi-task language understanding benchmark that extends MMLU by expanding multiple-choice options from 4 to 10, eliminating trivial questions, and focusing on reasoning-intensive tasks. Features over 12,000 curated questions across 14 domains and causes a 16-33% accuracy drop compared to original MMLU.
Is a higher score better?
Yes. Higher values rank better for this benchmark.
How many models are compared?
100 model results are currently shown.
Does this benchmark affect the overall score?
Yes. This benchmark can contribute to the current LLMBoard capability score.