math benchmark
STEM-focused subset of the Massive Multitask Language Understanding benchmark, evaluating language models on science, technology, engineering, and mathematics topics including physics, chemistry, mathematics, and other technical subjects.
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | AC | 80.9% | 100.0% | 2 | C | |
| 02 | AC | 76.4% | 0.0% | 2 | C |
A closer view of the leading scores on this benchmark.
The leading models and scores on this benchmark.
What MMLU-STEM measures and how its scores work.
STEM-focused subset of the Massive Multitask Language Understanding benchmark, evaluating language models on science, technology, engineering, and mathematics topics including physics, chemistry, mathematics, and other technical subjects.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about MMLU-STEM.
Qwen2.5 32B Instruct is currently ranked first with 80.9%.
STEM-focused subset of the Massive Multitask Language Understanding benchmark, evaluating language models on science, technology, engineering, and mathematics topics including physics, chemistry, mathematics, and other technical subjects.
Yes. Higher values rank better for this benchmark.
2 model results are currently shown.
No. This benchmark is shown for reference but does not contribute to the overall score.