llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

math benchmark

AIME 2025

All 30 problems from the 2025 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

Updated Aug 11, 2026

Models100
Model coverage114
MetricScore
EvidenceC

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

AIME 2025 Ranking

Higher score ranks better on this benchmark.

100 rows
Columns

Show columns

01GOGemini 3 ProGoogle100.0%100.0%114CAug 11, 2026
02OPGPT-5.2OpenAI100.0%99.1%114CAug 11, 2026
03OPGPT-5.2 ProOpenAI100.0%98.2%114CAug 11, 2026
04XAGrok-4 HeavyxAI100.0%97.3%114CAug 11, 2026
05MAKimi K2-Thinking-0905Moonshot AI100.0%96.5%114CAug 11, 2026
06ANClaude Opus 4.6Anthropic99.8%95.6%114CAug 11, 2026
07GOGemini 3 FlashGoogle99.7%94.7%114CAug 11, 2026
08OPGPT-5.1 HighOpenAI99.6%93.8%114CAug 11, 2026
09MELongCat-Flash-Thinking-2601Meituan99.6%92.9%114CAug 11, 2026
10NVNemotron 3 Nano (30B A3B)NVIDIA99.2%92.0%114CAug 11, 2026
11OPGPT OSS 20B HighOpenAI98.7%91.2%114CAug 11, 2026
12OPGPT-5.1 MediumOpenAI98.4%90.3%114CAug 11, 2026
13BYSeed 2.0 ProByteDance98.3%89.4%114CAug 11, 2026
14STStep-3.5-FlashStepFun97.3%88.5%114CAug 11, 2026
15MIMAI-Thinking-1Microsoft97.0%87.6%114CAug 11, 2026
16OPGPT-5.1 Codex HighOpenAI96.7%86.7%114CAug 11, 2026
17SASarvam-105BSarvam AI96.7%85.8%114CAug 11, 2026
18SASarvam-30BSarvam AI96.7%85.0%114CAug 11, 2026
19MAKimi K2.5Moonshot AI96.1%84.1%114CAug 11, 2026
20DEDeepSeek-V3.2-SpecialeDeepSeek96.0%83.2%114CAug 11, 2026
21ZAGLM-4.7Zhipu AI95.7%82.3%114CAug 11, 2026
22OPGPT-5OpenAI94.6%81.4%114CAug 11, 2026
23OPGPT-5 HighOpenAI94.6%80.5%114CAug 11, 2026
24XIMiMo-V2-FlashXiaomi94.1%79.7%114CAug 11, 2026
25OPGPT-5.1OpenAI94.0%78.8%114CAug 11, 2026
26OPGPT-5.1 InstantOpenAI94.0%77.9%114CAug 11, 2026
27OPGPT-5.1 ThinkingOpenAI94.0%77.0%114CAug 11, 2026
28ZAGLM-4.6Zhipu AI93.9%76.1%114CAug 11, 2026
29XAGrok-3xAI93.3%75.2%114CAug 11, 2026
30DEDeepSeek-V3.2 (Thinking)DeepSeek93.1%74.3%114CAug 11, 2026
31DEDeepSeek-V3.2DeepSeek93.1%73.5%114CAug 11, 2026
32BYSeed 2.0 LiteByteDance93.0%72.6%114CAug 11, 2026
33LAK-EXAONE-236B-A23BLG AI Research92.8%71.7%114CAug 11, 2026
34OPo4-miniOpenAI92.7%70.8%114CAug 11, 2026
35OPGPT OSS 120B HighOpenAI92.5%69.9%114CAug 11, 2026
36AMNova 2 ProAmazon92.3%69.0%114CAug 11, 2026
37ACQwen3-235B-A22B-Thinking-2507Alibaba Cloud / Qwen Team92.3%68.1%114CAug 11, 2026
38AMNova 2 OmniAmazon92.1%67.3%114CAug 11, 2026
39XAGrok 4 FastxAI92.0%66.4%114CAug 11, 2026
40XAGrok-4xAI91.7%65.5%114CAug 11, 2026
41ZAGLM-4.7-FlashZhipu AI91.6%64.6%114CAug 11, 2026
42OPGPT-5 miniOpenAI91.1%63.7%114CAug 11, 2026
43INMercury 2Inception91.1%62.8%114CAug 11, 2026
44AMNova 2 LiteAmazon91.0%62.0%114CAug 11, 2026
45XAGrok-3 MinixAI90.8%61.1%114CAug 11, 2026
46MELongCat-Flash-ThinkingMeituan90.6%60.2%114CAug 11, 2026
47NVNemotron 3 Super (120B A12B)NVIDIA90.2%59.3%114CAug 11, 2026
48COCommand A+Cohere90.0%58.4%114CAug 11, 2026
49ACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team89.7%57.5%114CAug 11, 2026
50DEDeepSeek-V3.2-ExpDeepSeek89.3%56.6%114CAug 11, 2026
51OPGPT-5 MediumOpenAI88.9%55.8%114CAug 11, 2026
52GOGemini 2.5 Pro Preview 06-05Google88.0%54.9%114CAug 11, 2026
53ACQwen3-Next-80B-A3B-ThinkingAlibaba Cloud / Qwen Team87.8%54.0%114CAug 11, 2026
54STStep3-VL-10BStepFun87.7%53.1%114CAug 11, 2026
55DEDeepSeek-R1-0528DeepSeek87.5%52.2%114CAug 11, 2026
56ANClaude Sonnet 4.5Anthropic87.0%51.3%114CAug 11, 2026
57BAERNIE 5.0Baidu87.0%50.4%114CAug 11, 2026
58OPo3OpenAI86.4%49.6%114CAug 11, 2026
59MAMistral Medium 3.5Mistral AI86.3%48.7%114CAug 11, 2026
60OPGPT-5 nanoOpenAI85.2%47.8%114CAug 11, 2026
61MAMinistral 3 (14B Reasoning 2512)Mistral AI85.0%46.9%114CAug 11, 2026
62MAMistral Small 4Mistral AI83.8%46.0%114CAug 11, 2026
63ACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team83.7%45.1%114CAug 11, 2026
64ACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team83.1%44.3%114CAug 11, 2026
65GOGemini 2.5 ProGoogle83.0%43.4%114CAug 11, 2026
66ACQwen3 MaxAlibaba Cloud / Qwen Team81.6%42.5%114CAug 11, 2026
67ACQwen3 235B A22BAlibaba Cloud / Qwen Team81.5%41.6%114CAug 11, 2026
68OPGPT-5.5 InstantOpenAI81.2%40.7%114CAug 11, 2026
69MIMiniMax M2.1MiniMax81.0%39.8%114CAug 11, 2026
70ANClaude Haiku 4.5Anthropic80.7%38.9%114CAug 11, 2026
71ACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team80.3%38.0%114CAug 11, 2026
72MAMinistral 3 (8B Reasoning 2512)Mistral AI78.7%37.2%114CAug 11, 2026
73OPMiniCPM-SALAOpenBMB78.3%36.3%114CAug 11, 2026
74ANClaude Opus 4.1Anthropic78.0%35.4%114CAug 11, 2026
75MIMiniMax M2MiniMax78.0%34.5%114CAug 11, 2026
76MIPhi 4 Reasoning PlusMicrosoft78.0%33.6%114CAug 11, 2026
77MIMiniMax M1 80KMiniMax76.9%32.7%114CAug 11, 2026
78ANClaude Opus 4Anthropic75.5%31.9%114CAug 11, 2026
79ACQwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team74.7%31.0%114CAug 11, 2026
80MIMiniMax M1 40KMiniMax74.6%30.1%114CAug 11, 2026
81ACQwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team74.5%29.2%114CAug 11, 2026
82ACQwen3 32BAlibaba Cloud / Qwen Team72.9%28.3%114CAug 11, 2026
83NVLlama 3.1 Nemotron Ultra 253B v1NVIDIA72.5%27.4%114CAug 11, 2026
84MAMin istral 3 (3B Reasoning 2512)Mistral AI72.1%26.6%114CAug 11, 2026
85NVNemotron Nano 9B v2NVIDIA72.1%25.7%114CAug 11, 2026
86GOGemini 2.5 FlashGoogle72.0%24.8%114CAug 11, 2026
87ACQwen3 30B A3BAlibaba Cloud / Qwen Team70.9%23.9%114CAug 11, 2026
88ANClaude Sonnet 4Anthropic70.5%23.0%114CAug 11, 2026
89ACQwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team70.3%22.1%114CAug 11, 2026
90ACQwen3-Next-80B-A3B-InstructAlibaba Cloud / Qwen Team69.5%21.2%114CAug 11, 2026
91ACQwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team69.3%20.4%114CAug 11, 2026
92ACQwen3 VL 32B InstructAlibaba Cloud / Qwen Team66.2%19.5%114CAug 11, 2026
93MAMagistral MediumMistral AI64.9%18.6%114CAug 11, 2026
94MELongCat-Flash-LiteMeituan63.2%17.7%114CAug 11, 2026
95MIPhi 4 ReasoningMicrosoft62.9%16.8%114CAug 11, 2026
96MAMagistral Small 2506Mistral AI62.8%15.9%114CAug 11, 2026
97MELongCat-Flash-ChatMeituan61.3%15.0%114CAug 11, 2026
98NVLlama-3.3 Nemotron Super 49B v1NVIDIA58.4%14.2%114CAug 11, 2026
99ANClaude 3.7 SonnetAnthropic54.8%13.3%114CAug 11, 2026
100DEDeepSeek-V3.1DeepSeek49.8%12.4%114CAug 11, 2026

AIME 2025 Score Distribution

A closer view of the leading scores on this benchmark.

AIME 2025

AIME 2025 Highlights

The leading models and scores on this benchmark.

Rank #1Gemini 3 Pro100.0%Rank #2GPT-5.2100.0%Rank #3GPT-5.2 Pro100.0%Rank #4Grok-4 Heavy100.0%

What is AIME 2025?

What AIME 2025 measures and how its scores work.

All 30 problems from the 2025 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of C.

Family
AIME 2025
Modality
text
Primary category
math
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
aime-2025|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about AIME 2025.

Which model scores highest on AIME 2025?

Gemini 3 Pro is currently ranked first with 100.0%.

What does AIME 2025 measure?

All 30 problems from the 2025 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

100 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.