llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

Capability ranking

Best AI for Reasoning

Best Reasoning Models

Updated 2026-08-11

Ranked models99
Vendors17
Open-weight models17
Official prices40

On this page

  • Overview
  • Leaders
  • Capability
  • Ranking
  • Value
  • Composition
  • Compare
  • FAQ

What this ranking says

A quick reading of the current reasoning ranking before the full data table.

Claude Opus 4.8 leads this page at 99.0, 0.0 points ahead of GPT-5.5.

Gemma 4 31B is the highest-ranked open-weight option at #7. Nova Micro has the lowest official input price among ranked models.

Top reasoning modelClaude Opus 4.899.0 reasoning scoreBest open-weight modelGemma 4 31BRank #7Lowest official inputNova Micro$0.04 / 1MLargest contextGemini 1.5 Pro2.1M tokens

Reasoning Ranking Leaders

The leading reasoning scores, with the axis focused on the competitive range rather than zero.

Top reasoning models

Reasoning Capability Profile

Compare reasoning strength with each model's broader capability score.

ModelReasoningKnowledgeMathCodingInstruction
Claude Opus 4.8Anthropic99.094.6N/A85.7N/A
GPT-5.5OpenAI99.083.762.574.3N/A
Qwen3.8 MaxAlibaba Cloud / Qwen Team97.670.7N/A80.1100.0
Gemini 3.1 ProGoogle96.290.3N/A54.0N/A
Qwen3.7 MaxAlibaba Cloud / Qwen Team91.589.193.185.692.2
19.24162.784.5106.3-7.118.744.670.596.4Reasoning scoreOverall score
80 models with comparable dataOpen a model by selecting its logo
GPQA94.3%Gemini 3.1 Pro85 compared models
MMLU-Pro89.6%Qwen3.7 Max65 compared models
AIME 2025100.0%Gemini 3 Pro28 compared models
SWE-Bench Verified88.6%Claude Opus 4.826 compared models
MMLU90.6%Qwen3 VL 235B A22B Thinking55 compared models
Humanity's Last Exam58.4%Muse Spark30 compared models
LiveCodeBench79.0%Grok-413 compared models
MATH89.0%Gemma 3 27B35 compared models

Reasoning Model Ranking

Models are ordered by their reasoning score.

99 rows
Columns

Show columns

01ANClaude Opus 4.8AnthropicClaude Opus 4.899.087.994.6N/A85.7N/A60.0%1MClosed weights$5$25
02OPGPT-5.5OpenAIGPT-5.599.084.083.762.574.3N/A80.0%1.1MClosed weights$5$30
03ACQwen3.8 MaxAlibaba Cloud / Qwen TeamQwen3.8 Max97.689.370.7N/A80.1100.080.0%1MClosed weights$2$6
04GOGemini 3.1 ProGoogleGemini 3.1 Pro96.277.590.3N/A54.0N/A60.0%1MClosed weights$2$12
05ACQwen3.7 MaxAlibaba Cloud / Qwen TeamQwen3.7 Max91.578.889.193.185.692.2100.0%1MClosed weights$2.5$7.5
06OPGPT-5.4OpenAIGPT-5.491.277.163.081.355.2N/A80.0%1.1MClosed weights$2.5$15
07GOGemma 4 31BGoogleGemma 4 31B87.660.167.435.361.5N/A80.0%262.1KOpen weightsN/AN/A
08ACQwen3.7 PlusAlibaba Cloud / Qwen TeamQwen3.7-Plus84.773.883.566.961.687.8100.0%1MClosed weights$0.50$3
09XIMiMo V2.5 ProXiaomiMiMo-V2.5-Pro84.046.762.892.849.7N/A80.0%1MOpen weights$0.435$0.87
10ANClaude Opus 4.6AnthropicClaude Opus 4.683.279.687.795.674.7N/A80.0%1MClosed weights$5$25
11ACQwen3.5 397B A17BAlibaba Cloud / Qwen TeamQwen3.5-397B-A17B81.366.382.850.157.481.9100.0%262.1KOpen weights$0.60$3.6
12ACQwen3.6 PlusAlibaba Cloud / Qwen TeamQwen3.6 Plus80.967.587.373.351.872.7100.0%1MClosed weights$0.50$3
13GOGemma 4 26B A4BGoogleGemma 4 26B-A4B79.252.652.826.557.7N/A80.0%262.1KOpen weightsN/AN/A
14OPGPT-5.2-ProOpenAIGPT-5.2 Pro78.571.659.899.1N/AN/A60.0%400KClosed weights$21$168
15ANClaude Sonnet 4.6AnthropicClaude Sonnet 4.678.272.677.7N/A48.7N/A60.0%1MClosed weights$3$15
16ANClaude Sonnet 3.5AnthropicClaude 3.5 Sonnet77.832.152.079.955.7N/A80.0%200KClosed weightsN/AN/A
17OPGPT-5.2OpenAIGPT-5.274.071.469.990.087.5N/A80.0%400KClosed weights$1.75$14
18ANClaude Opus 3AnthropicClaude 3 Opus73.816.730.955.446.1N/A80.0%200KClosed weightsN/AN/A
19GOGemini 3 ProGoogleGemini 3 Pro73.467.374.598.254.4N/A80.0%1MClosed weightsN/AN/A
20GOGemini 3 FlashGoogleGemini 3 Flash72.164.869.094.753.4N/A80.0%1MClosed weights$0.50$3
21ACQwen3.6 27BAlibaba Cloud / Qwen TeamQwen3.6-27B71.860.571.449.949.0N/A80.0%262.1KOpen weights$0.60$3.6
22ACQwen3 235B A22B ThinkingAlibaba Cloud / Qwen TeamQwen3-235B-A22B-Thinking-250771.448.561.963.553.984.9100.0%262.1KClosed weightsN/AN/A
23COCommand R+CohereCommand R+71.20.030.810.6N/AN/A60.0%128KClosed weights$0.15$0.60
24GOGemma 4 12BGoogleGemma 4 12B70.338.230.217.629.336.7100.0%N/AClosed weightsN/AN/A
25AMNova ProAmazonNova Pro68.620.965.269.975.486.7100.0%300KClosed weights$0.80$3.2
26GOGemma 2 27BGoogleGemma 2 27B68.60.059.410.78.6N/A80.0%N/AClosed weightsN/AN/A
27MEMuse SparkMetaMuse Spark68.073.297.8N/A47.2N/A60.0%N/AClosed weightsN/AN/A
28OPGPT-4OpenAIGPT-467.64.670.27.110.8N/A80.0%8.2KClosed weights$30$60
29ACQwen3.5 122B A10BAlibaba Cloud / Qwen TeamQwen3.5-122B-A10B67.459.284.656.354.269.8100.0%262.1KOpen weights$0.40$3.2
30MELlama 3.1 405BMetaLlama 3.1 405B Instruct67.224.541.475.075.462.5100.0%128KClosed weightsN/AN/A
31ACQwen3.5 27BAlibaba Cloud / Qwen TeamQwen3.5-27B62.557.278.059.447.970.6100.0%262.1KOpen weights$0.30$2.4
32GOGemini 1.5 ProGoogleGemini 1.5 Pro62.020.546.166.941.575.7100.0%2.1MClosed weightsN/AN/A
33ACQwen3.5 35B A3BAlibaba Cloud / Qwen TeamQwen3.5-35B-A3B61.852.071.243.844.254.4100.0%262.1KOpen weights$0.25$2
34MIPhi 3.5 MoEMicrosoftPhi-3.5-MoE-instruct61.84.950.634.339.775.0100.0%N/AClosed weightsN/AN/A
35ANClaude Opus 4.5AnthropicClaude Opus 4.561.369.587.5N/A76.2N/A60.0%200KClosed weights$5$25
36ACQwen3 235B A22BAlibaba Cloud / Qwen TeamQwen3-235B-A22B-Instruct-250761.142.470.621.551.272.8100.0%262.1KOpen weights$0.70$2.8
37GODiffusionGemma 26B A4BGoogleDiffusionGemma 26B-A4B59.934.237.911.821.2N/A80.0%N/AClosed weightsN/AN/A
38NRHermes 3 70BNous ResearchHermes 3 70B59.27.239.61.4N/A82.880.0%131.1KClosed weightsN/AN/A
39ACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen TeamQwen3 VL 235B A22B Thinking58.645.855.953.846.177.7100.0%262.1KClosed weightsN/AN/A
40ANClaude Sonnet 3AnthropicClaude 3 Sonnet58.43.518.035.720.0N/A80.0%200KClosed weightsN/AN/A
41ACQwen3 Next 80B A3B ThinkingAlibaba Cloud / Qwen TeamQwen3-Next-80B-A3B-Thinking58.240.965.245.738.562.7100.0%131.1KOpen weights$0.50$6
42NVLlama 3.1 Nemotron 70BNVIDIALlama 3.1 Nemotron 70B Instruct57.33.152.053.2N/A0.080.0%128KOpen weightsN/AN/A
43GOGemini 1.5 FlashGoogleGemini 1.5 Flash55.110.728.152.423.137.7100.0%1MClosed weightsN/AN/A
44ACQwen3 VL 235B A22BAlibaba Cloud / Qwen TeamQwen3 VL 235B A22B Instruct54.043.264.828.053.276.8100.0%262.1KClosed weightsN/AN/A
45OPGPT-4-TurboOpenAIGPT-4 Turbo53.814.671.754.355.4N/A80.0%128KClosed weights$10$30
46ACQwen3 Next 80B A3BAlibaba Cloud / Qwen TeamQwen3-Next-80B-A3B-Instruct53.436.653.018.949.062.7100.0%131.1KOpen weights$0.50$2
47MELlama 3.1 70BMetaLlama 3.1 70B Instruct52.913.224.2100.032.351.6100.0%128KClosed weightsN/AN/A
48XAGrok 4xAIGrok-451.956.564.172.383.3N/A80.0%256KClosed weightsN/AN/A
49AMNova LiteAmazonNova Lite51.413.241.963.748.571.9100.0%300KClosed weights$0.06$0.24
50ACQwen2 72BAlibaba Cloud / Qwen TeamQwen2 72B Instruct50.99.824.941.446.6N/A80.0%N/AClosed weightsN/AN/A
51ACQwen2.5 32BAlibaba Cloud / Qwen TeamQwen2.5 32B Instruct50.614.837.481.768.3N/A80.0%N/AClosed weights$0.70$2.8
52GOGemma 2 9BGoogleGemma 2 9B49.90.039.55.32.3N/A80.0%N/AClosed weightsN/AN/A
53MELongCat Flash ChatMeituanLongCat-Flash-Chat49.336.975.038.249.670.3100.0%128KClosed weightsN/AN/A
54ACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen TeamQwen3 VL 32B Thinking47.940.566.545.133.663.4100.0%N/AClosed weightsN/AN/A
55ACQwen3.5 9BAlibaba Cloud / Qwen TeamQwen3.5-9B47.641.452.028.133.644.6100.0%262.1KOpen weightsN/AN/A
56OPo3OpenAIo347.352.622.337.364.567.9100.0%200KClosed weights$2$8
57MAMistral NeMoMistral AIMistral NeMo Instruct46.90.017.8N/AN/AN/A40.0%128KClosed weights$0.15$0.15
58GOGemma 3 27BGoogleGemma 3 27B46.815.837.989.039.261.1100.0%131.1KClosed weightsN/AN/A
59ANClaude Haiku 3AnthropicClaude 3 Haiku46.40.024.823.127.7N/A80.0%200KClosed weightsN/AN/A
60ACQwen2.5 Coder 32BAlibaba Cloud / Qwen TeamQwen2.5-Coder 32B Instruct46.34.417.239.368.0N/A80.0%128KClosed weights$0.287$0.861
61ALJamba 1.5 LargeAI21 LabsJamba 1.5 Large46.12.733.931.9N/A14.380.0%256KClosed weightsN/AN/A
62GOGemma 4 E4BGoogleGemma 4 E4B45.224.427.15.918.3N/A80.0%131.1KOpen weightsN/AN/A
63ANClaude Haiku 3.5AnthropicClaude 3.5 Haiku42.29.322.742.935.5N/A80.0%200KClosed weightsN/AN/A
64AMNova MicroAmazonNova Micro40.14.329.350.734.647.7100.0%128KClosed weights$0.035$0.14
65GOGemma 3 12BGoogleGemma 3 12B39.411.235.474.832.451.4100.0%131.1KClosed weightsN/AN/A
66ACQwen3 VL 32BAlibaba Cloud / Qwen TeamQwen3 VL 32B Instruct39.335.243.319.59.641.0100.0%N/AClosed weightsN/AN/A
67ACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen TeamQwen3 VL 30B A3B Thinking37.831.843.438.830.839.9100.0%262.1KClosed weightsN/AN/A
68ANClaude Opus 4AnthropicClaude Opus 437.746.468.831.957.2N/A80.0%200KClosed weightsN/AN/A
69ACQwen3.5 4BAlibaba Cloud / Qwen TeamQwen3.5-4B35.832.436.918.825.036.3100.0%N/AClosed weightsN/AN/A
70OPGPT-4o-miniOpenAIGPT-4o mini35.38.150.546.428.5N/A80.0%128KClosed weights$0.15$0.60
71ALJamba 1.5 MiniAI21 LabsJamba 1.5 Mini34.80.011.214.9N/A0.080.0%256.1KClosed weightsN/AN/A
72MAMinistral 8BMistral AIMinistral 8B Instruct34.10.012.927.10.027.3100.0%131.1KOpen weightsN/AN/A
73MIPhi 4MicrosoftPhi 433.89.420.674.338.53.9100.0%16KClosed weights$0.125$0.50
74GOGemma 4 E2BGoogleGemma 4 E2B32.911.614.30.011.5N/A80.0%131.1KOpen weightsN/AN/A
75ACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen TeamQwen3 VL 8B Thinking32.528.547.033.628.941.3100.0%262.1KClosed weightsN/AN/A
76ACQwen3 VL 30B A3BAlibaba Cloud / Qwen TeamQwen3 VL 30B A3B Instruct30.628.841.614.37.737.6100.0%262.1KClosed weightsN/AN/A
77ACQwen3 VL 4B ThinkingAlibaba Cloud / Qwen TeamQwen3 VL 4B Thinking28.423.030.220.913.529.2100.0%262.1KClosed weightsN/AN/A
78MIPhi 3.5 miniMicrosoftPhi-3.5-mini-instruct27.70.042.924.418.7100.0100.0%128KClosed weightsN/AN/A
79IBGranite 3.3 8BIBMGranite 3.3 8B Instruct27.43.848.725.380.810.2100.0%128KClosed weightsN/AN/A
80MELlama 3.1 8BMetaLlama 3.1 8B Instruct26.40.010.20.018.520.3100.0%131.1KClosed weightsN/AN/A
81GOGemma 3n E4B LiteRTGoogleGemma 3n E4B Instructed LiteRT Preview25.70.026.419.916.929.6100.0%N/AClosed weightsN/AN/A
82MIPhi 4 MiniMicrosoftPhi 4 Mini23.80.048.235.6N/AN/A60.0%128KOpen weights$0.075$0.30
83MELlama 3.2 3BMetaLlama 3.2 3B Instruct23.20.04.016.7N/A14.180.0%128KClosed weightsN/AN/A
84ACQwen2.5 Coder 7BAlibaba Cloud / Qwen TeamQwen2.5-Coder 7B Instruct22.80.03.719.950.5N/A80.0%N/AClosed weights$0.144$0.287
85OPGPT-4oOpenAIGPT-4o22.823.935.20.08.517.4100.0%128KClosed weights$2.5$10
86ACQwen3 VL 8BAlibaba Cloud / Qwen TeamQwen3 VL 8B Instruct22.424.825.45.23.936.0100.0%262.1KClosed weightsN/AN/A
87GOGemma 3 4BGoogleGemma 3 4B21.70.013.351.810.850.4100.0%131.1KClosed weightsN/AN/A
88ACQwen2.5 14BAlibaba Cloud / Qwen TeamQwen2.5 14B Instruct20.38.240.073.151.2N/A80.0%N/AClosed weights$0.35$1.4
89OPGPT-3.5-TurboOpenAIGPT-3.5 Turbo18.10.017.210.712.3N/A80.0%16.4KClosed weights$0.50$1.5
90GOGemini DiffusionGoogleGemini Diffusion16.33.646.13.536.8N/A80.0%N/AClosed weightsN/AN/A
91IBIBM Granite 4.0 TinyIBMIBM Granite 4.0 Tiny Preview15.90.025.18.536.93.9100.0%N/AClosed weightsN/AN/A
92ACQwen3.5 2BAlibaba Cloud / Qwen TeamQwen3.5-2B15.12.813.3N/AN/A11.360.0%N/AClosed weightsN/AN/A
93ACQwen3 VL 4BAlibaba Cloud / Qwen TeamQwen3 VL 4B Instruct12.319.333.44.01.916.8100.0%262.1KClosed weightsN/AN/A
94GOGemma 3n E2B LiteRTGoogleGemma 3n E2B Instructed LiteRT (Preview)9.50.010.36.57.011.4100.0%N/AClosed weightsN/AN/A
95GOGemma 3n E4BGoogleGemma 3n E4B Instructed8.00.023.419.916.929.6100.0%32KClosed weightsN/AN/A
96BAERNIE 4.5BaiduERNIE 4.56.50.00.30.00.00.0100.0%128KClosed weightsN/AN/A
97ACQwen3.5 0.8BAlibaba Cloud / Qwen TeamQwen3.5-0.8B2.20.03.3N/AN/A0.660.0%N/AClosed weightsN/AN/A
98GOGemma 3n E2BGoogleGemma 3n E2B Instructed1.80.011.06.57.011.4100.0%N/AClosed weightsN/AN/A
99GOGemma 3 1BGoogleGemma 3 1B0.10.00.66.91.011.7100.0%N/AClosed weightsN/AN/A

Official PAYG prices appear here. Third-party offers remain on the pricing and model detail pages.

Capability at each price point

Official vendor API prices plotted against the reasoning metric used on this page. Price remains a separate decision signal.

020406080100$0.05$0.1$0.2$0.5$1.0$2.0$5.0$10Reasoning scoreOfficial API price blend (8:1 input/output), USD / 1M
Efficient frontier40 models with official PAYG prices

Reasoning Ranking Composition

Vendor concentration and model access among the first 25 ranked products.

Top-model vendor mix

25 models
Alibaba Cloud / Qwen Team7Google6Anthropic5OpenAI4Xiaomi1Cohere1Amazon1

Access model

Open weights
5
Closed weights
20
Official price available
18

Compare the leaders

Open focused comparisons between the current leader and the nearest practical alternatives.

Claude Opus 4.8vsGPT-5.5Claude Opus 4.8vsGemma 4 31BClaude Opus 4.8vsNova Micro
Overall leaderboardCoding leaderboardMath leaderboardKnowledge leaderboardInstruction leaderboard

About this leaderboard

Common questions about Best AI for Reasoning.

What is the best model on Best AI for Reasoning?

Claude Opus 4.8 is currently ranked first with a reasoning score of 99.0.

How is this leaderboard ranked?

Each model appears once using its current scored version. Category pages use the capability named in the title, while the overall page uses the LLMBoard score.

How many models are included?

This page currently ranks 99 unique model products.

Where does pricing come from?

Main price columns use the model vendor's official standard PAYG API rate. Eligible third-party offers appear only in separately labeled columns, and unavailable official prices display as N/A.

Does Arena affect the score?

No. Arena results are displayed as an independent signal and are not included in the current LLMBoard capability score.