llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

reasoning benchmark

SWE-Bench Verified

A verified subset of 500 software engineering problems from real GitHub issues, validated by human annotators for evaluating language models' ability to resolve real-world coding issues by generating patches for Python codebases.

Updated Aug 11, 2026

Models100
Model coverage105
MetricScore
EvidenceC

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

SWE-Bench Verified Ranking

Higher score ranks better on this benchmark.

100 rows
Columns

Show columns

01ANClaude Fable 5Anthropic95.0%100.0%105CAug 11, 2026
02ANClaude Mythos PreviewAnthropic93.9%99.0%105CAug 11, 2026
03ANClaude Opus 4.8Anthropic88.6%98.1%105CAug 11, 2026
04ANClaude Opus 4.7Anthropic87.6%97.1%105CAug 11, 2026
05ANClaude Sonnet 5Anthropic85.2%96.2%105CAug 11, 2026
06ANClaude Opus 4.5Anthropic80.9%95.2%105CAug 11, 2026
07ANClaude Opus 4.6Anthropic80.8%94.2%105CAug 11, 2026
08DEDeepSeek-V4-Pro-MaxDeepSeek80.6%93.3%105CAug 11, 2026
09GOGemini 3.1 ProGoogle80.6%92.3%105CAug 11, 2026
10MIMiniMax M3MiniMax80.5%91.3%105CAug 11, 2026
11ACQwen3.7 MaxAlibaba Cloud / Qwen Team80.4%90.4%105CAug 11, 2026
12MAKimi K2.6Moonshot AI80.2%89.4%105CAug 11, 2026
13MIMiniMax M2.5MiniMax80.2%88.5%105CAug 11, 2026
14OPGPT-5.2OpenAI80.0%87.5%105CAug 11, 2026
15ANClaude Sonnet 4.6Anthropic79.6%86.5%105CAug 11, 2026
16DEDeepSeek-V4-Flash-MaxDeepSeek79.0%85.6%105CAug 11, 2026
17XIMiMo-V2.5-ProXiaomi78.9%84.6%105CAug 11, 2026
18ACQwen3.6 PlusAlibaba Cloud / Qwen Team78.8%83.7%105CAug 11, 2026
19GOGemini 3 FlashGoogle78.0%82.7%105CAug 11, 2026
20TEHy3Tencent78.0%81.7%105CAug 11, 2026
21XIMiMo-V2-ProXiaomi78.0%80.8%105CAug 11, 2026
22ZAGLM-5Zhipu AI77.8%79.8%105CAug 11, 2026
23ACQwen3.7-PlusAlibaba Cloud / Qwen Team77.7%78.8%105CAug 11, 2026
24MAMistral Medium 3.5Mistral AI77.6%77.9%105CAug 11, 2026
25MEMuse SparkMeta77.4%76.9%105CAug 11, 2026
26ACQwen3.6-27BAlibaba Cloud / Qwen Team77.2%76.0%105CAug 11, 2026
27MAKimi K2.5Moonshot AI76.8%75.0%105CAug 11, 2026
28BYSeed 2.0 ProByteDance76.5%74.0%105CAug 11, 2026
29ACQwen3.5-397B-A17BAlibaba Cloud / Qwen Team76.4%73.1%105CAug 11, 2026
30OPGPT-5.1OpenAI76.3%72.1%105CAug 11, 2026
31OPGPT-5.1 InstantOpenAI76.3%71.2%105CAug 11, 2026
32OPGPT-5.1 ThinkingOpenAI76.3%70.2%105CAug 11, 2026
33GOGemini 3 ProGoogle76.2%69.2%105CAug 11, 2026
34MEMuse Glimmer-30BMeta76.0%68.3%105CAug 11, 2026
35OPGPT-5OpenAI74.9%67.3%105CAug 11, 2026
36XIMiMo-V2-OmniXiaomi74.8%66.3%105CAug 11, 2026
37ANClaude Opus 4.1Anthropic74.5%65.4%105CAug 11, 2026
38OPGPT-5 CodexOpenAI74.5%64.4%105CAug 11, 2026
39STStep-3.5-FlashStepFun74.4%63.5%105CAug 11, 2026
40ZAGLM-4.7Zhipu AI73.8%62.5%105CAug 11, 2026
41OPGPT-5.1 CodexOpenAI73.7%61.5%105CAug 11, 2026
42MIMAI-Thinking-1Microsoft73.5%60.6%105CAug 11, 2026
43BYSeed 2.0 LiteByteDance73.5%59.6%105CAug 11, 2026
44XIMiMo-V2-FlashXiaomi73.4%58.6%105CAug 11, 2026
45ACQwen3.6-35B-A3BAlibaba Cloud / Qwen Team73.4%57.7%105CAug 11, 2026
46ANClaude Haiku 4.5Anthropic73.3%56.7%105CAug 11, 2026
47DEDeepSeek-V3.2 (Thinking)DeepSeek73.1%55.8%105CAug 11, 2026
48DEDeepSeek-V3.2DeepSeek73.1%54.8%105CAug 11, 2026
49DEDeepSeek-V3.2-SpecialeDeepSeek73.1%53.9%105CAug 11, 2026
50ANClaude Sonnet 4Anthropic72.7%52.9%105CAug 11, 2026
51ANClaude Opus 4Anthropic72.5%51.9%105CAug 11, 2026
52ACQwen3.5-27BAlibaba Cloud / Qwen Team72.4%51.0%105CAug 11, 2026
53ACQwen3.5-122B-A10BAlibaba Cloud / Qwen Team72.0%50.0%105CAug 11, 2026
54MIMAI-Code-1-FlashMicrosoft71.6%49.0%105CAug 11, 2026
55MAKimi K2-Thinking-0905Moonshot AI71.3%48.1%105CAug 11, 2026
56XAGrok Code Fast 1xAI70.8%47.1%105CAug 11, 2026
57NVNemotron 3 Ultra (550B A55B)NVIDIA70.7%46.1%105CAug 11, 2026
58ANClaude 3.7 SonnetAnthropic70.3%45.2%105CAug 11, 2026
59MELongCat-Flash-Thinking-2601Meituan70.0%44.2%105CAug 11, 2026
60AMNova 2 ProAmazon70.0%43.3%105CAug 11, 2026
61ACQwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen Team69.6%42.3%105CAug 11, 2026
62ACQwen3 MaxAlibaba Cloud / Qwen Team69.6%41.4%105CAug 11, 2026
63MIMiniMax M2MiniMax69.4%40.4%105CAug 11, 2026
64ACQwen3.5-35B-A3BAlibaba Cloud / Qwen Team69.2%39.4%105CAug 11, 2026
65OPo3OpenAI69.1%38.5%105CAug 11, 2026
66OPo4-miniOpenAI68.1%37.5%105CAug 11, 2026
67ZAGLM-4.6Zhipu AI68.0%36.5%105CAug 11, 2026
68DEDeepSeek-V3.2-ExpDeepSeek67.8%35.6%105CAug 11, 2026
69CONorth Mini Code 1.0Cohere67.6%34.6%105CAug 11, 2026
70GOGemini 2.5 Pro Preview 06-05Google67.2%33.6%105CAug 11, 2026
71MIMiniMax M2.1MiniMax67.0%32.7%105CAug 11, 2026
72DEDeepSeek-V3.1DeepSeek66.0%31.7%105CAug 11, 2026
73MAKimi K2-Instruct-0905Moonshot AI65.8%30.8%105CAug 11, 2026
74AMNova 2 LiteAmazon64.5%29.8%105CAug 11, 2026
75ZAGLM-4.5Zhipu AI64.2%28.9%105CAug 11, 2026
76GOGemini 2.5 ProGoogle63.2%27.9%105CAug 11, 2026
77MADevstral MediumMistral AI61.6%26.9%105CAug 11, 2026
78GOGemini 2.5 FlashGoogle60.4%26.0%105CAug 11, 2026
79MELongCat-Flash-ChatMeituan60.4%25.0%105CAug 11, 2026
80MELongCat-Flash-ThinkingMeituan59.4%24.0%105CAug 11, 2026
81ZAGLM-4.7-FlashZhipu AI59.2%23.1%105CAug 11, 2026
82ZAGLM-4.5-AirZhipu AI57.6%22.1%105CAug 11, 2026
83MIMiniMax M1 80KMiniMax56.0%21.1%105CAug 11, 2026
84MIMiniMax M1 40KMiniMax55.6%20.2%105CAug 11, 2026
85OPGPT-4.1OpenAI54.6%19.2%105CAug 11, 2026
86MELongCat-Flash-LiteMeituan54.4%18.3%105CAug 11, 2026
87NVNemotron 3 Super (120B A12B)NVIDIA53.7%17.3%105CAug 11, 2026
88MADevstral Small 1.1Mistral AI53.6%16.4%105CAug 11, 2026
89OPo3-miniOpenAI49.3%15.4%105CAug 11, 2026
90ANClaude 3.5 SonnetAnthropic49.0%14.4%105CAug 11, 2026
91SASarvam-105BSarvam AI45.0%13.5%105CAug 11, 2026
92DEDeepSeek-R1-0528DeepSeek44.6%12.5%105CAug 11, 2026
93DEDeepSeek-V3DeepSeek42.0%11.5%105CAug 11, 2026
94OPo1-previewOpenAI41.3%10.6%105CAug 11, 2026
95OPo1OpenAI41.0%9.6%105CAug 11, 2026
96ANClaude 3.5 HaikuAnthropic40.6%8.7%105CAug 11, 2026
97NVNemotron 3 Nano (30B A3B)NVIDIA38.8%7.7%105CAug 11, 2026
98OPGPT-4.5OpenAI38.0%6.7%105CAug 11, 2026
99SASarvam-30BSarvam AI34.0%5.8%105CAug 11, 2026
100OPGPT-4oOpenAI33.2%4.8%105CAug 11, 2026

SWE-Bench Verified Score Distribution

A closer view of the leading scores on this benchmark.

SWE-Bench Verified

SWE-Bench Verified Highlights

The leading models and scores on this benchmark.

Rank #1Claude Fable 595.0%Rank #2Claude Mythos Preview93.9%Rank #3Claude Opus 4.888.6%Rank #4Claude Opus 4.787.6%

What is SWE-Bench Verified?

What SWE-Bench Verified measures and how its scores work.

A verified subset of 500 software engineering problems from real GitHub issues, validated by human annotators for evaluating language models' ability to resolve real-world coding issues by generating patches for Python codebases.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of C.

Family
SWE-Bench Verified
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
swe-bench-verified|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about SWE-Bench Verified.

Which model scores highest on SWE-Bench Verified?

Claude Fable 5 is currently ranked first with 95.0%.

What does SWE-Bench Verified measure?

A verified subset of 500 software engineering problems from real GitHub issues, validated by human annotators for evaluating language models' ability to resolve real-world coding issues by generating patches for Python codebases.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

100 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.