llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

Evaluation directory

AI Benchmarks

Browse benchmark definitions, capability categories and model rankings.

Data as of 2026-08-11

Benchmarks668
Categories34
Score eligible331
With model coverage665

Independent source ranking

Community preference remains a separate signal from benchmark capability scores.

LM ArenaOriginal preference ratings, votes, and task leaderboards22 boards 254 text results

All benchmarks

Open any benchmark to inspect its leaderboard, scoring direction, evidence fields and programmatic FAQ.

145 rows
Columns

Show columns

GPQAphysicsGPQAScore234featuredCYes
MMLU-ProlanguageMMLU-ProScore129featuredBYes
AIME 2025mathAIME 2025Score114featuredCYes
SWE-Bench VerifiedreasoningSWE-Bench VerifiedScore105featuredCYes
MMLUlanguageMMLUScore100featuredBYes
Humanity's Last ExammathHumanity's Last ExamScore93featuredBYes
LiveCodeBenchreasoningLiveCodeBenchScore73featuredCYes
MATHmathMATHScore71featuredBYes
HumanEvalreasoningHumanEvalScore66featuredBYes
MMMU-PromultimodalMMMU-ProScore66featuredBYes
IFEvalinstruction followingIFEvalScore65featuredBYes
MMMUmultimodalMMMUScore63featuredBYes
BrowseCompreasoningBrowseCompScore58featuredBYes
AIME 2024mathAIME 2024Score53featuredBYes
LiveCodeBench v6reasoningLiveCodeBench v6Score53featuredBYes
nolimalong contextnolimaScore52featuredCNo
MMMLUlanguageMMMLUScore49featuredBYes
Terminal-Bench 2.0reasoningTerminal-Bench 2.0Score49featuredCYes
CharXiv-RmultimodalCharXiv-RScore48featuredBYes
GSM8kmathGSM8kScore48featuredBYes
MMLU-ReduxlanguageMMLU-ReduxScore48featuredBYes
SimpleQAreasoningSimpleQAScore46featuredBYes
SWE-Bench ProreasoningSWE-Bench ProScore45featuredBYes
MathVistamathMathVistaScore39featuredBYes
LiveBenchmathLiveBenchScore38featuredBYes
Tau2 TelecomreasoningTau2 TelecomScore35featuredBYes
ARC-CreasoningARC-CScore34featuredBYes
SuperGPQAlegalSuperGPQAScore34featuredBYes
SWE-bench MultilingualreasoningSWE-bench MultilingualScore34featuredBYes
HMMT 2025mathHMMT 2025Score33featuredBYes
MBPPreasoningMBPPPass@133featuredBYes
AI2DmultimodalAI2DScore32featuredBYes
MATH-500mathMATH-500Score32featuredCYes
MathVisionmathMathVisionScore32featuredBYes
MMLU-ProXlanguageMMLU-ProXScore32featuredBYes
IncludegeneralIncludeScore31featuredBYes
MCP AtlasreasoningMCP AtlasScore31featuredBYes
MGSMmathMGSMScore31featuredBYes
ToolathlonreasoningToolathlonScore31featuredBYes
DROPmathDROPScore30featuredBYes
IFBenchinstruction followingIFBenchScore29featuredBYes
Multi-ChallengereasoningMulti-ChallengeScore29featuredCYes
HellaSwagreasoningHellaSwagScore27featuredBYes
Arena HardreasoningArena HardScore26featuredBYes
DocVQAimage to textDocVQAScore26featuredBYes
Finance Agent v2reasoningFinance Agent v2Score26featuredBYes
RealWorldQAspatial reasoningRealWorldQAScore26featuredBYes
Tau2 RetailreasoningTau2 RetailScore26featuredBYes
VideoMMMUmultimodalVideoMMMUScore26featuredBYes
HMMT25mathHMMT25Score25featuredBYes
ScreenSpot PromultimodalScreenSpot ProScore25featuredCYes
TAU-bench RetailreasoningTAU-bench RetailScore25featuredBYes
Terminal-BenchreasoningTerminal-BenchScore25featuredBYes
ChartQAmultimodalChartQAScore24featuredBYes
LVBenchlong contextLVBenchScore24featuredBYes
ERQAreasoningERQAScore23featuredCYes
MathVista-MinimathMathVista-MiniScore23featuredBYes
OSWorld-VerifiedmultimodalOSWorld-VerifiedScore23featuredBYes
PolyMATHmathPolyMATHScore23featuredBYes
t2-benchreasoningt2-benchScore23featuredBYes
TAU-bench AirlinereasoningTAU-bench AirlineScore23featuredBYes
Tau2 AirlinereasoningTau2 AirlineScore23featuredBYes
WMT24++languageWMT24++Score23featuredBYes
Aider-PolyglotgeneralAider-PolyglotScore22featuredBYes
MMStarmultimodalMMStarScore22featuredBYes
OCRBenchimage to textOCRBenchScore22featuredBYes
WinograndelanguageWinograndeScore22featuredBYes
BIG-Bench HardlanguageBIG-Bench HardScore21featuredBYes
MRCR v2 (8-needle)long contextMRCR v2 (8-needle)Score21featuredCYes
DeepSWE 1.1agentsDeepSWE 1.1Score20featuredBYes
Multi-IFinstruction followingMulti-IFScore20featuredBYes
OSWorldmultimodalOSWorldScore20featuredBYes
BFCL-v3reasoningBFCL-v3Score19featuredBYes
IMO-AnswerBenchmathIMO-AnswerBenchScore19featuredBYes
SciCodemathSciCodeScore19featuredBYes
Terminal-Bench 2.1reasoningTerminal-Bench 2.1Score19featuredBYes
AIME 2026mathAIME 2026Score18featuredBYes
C-EvalreasoningC-EvalScore18featuredBYes
CC-OCRmultimodalCC-OCRScore18featuredBYes
MMBench-V1.1multimodalMMBench-V1.1Score18featuredBYes
TriviaQAreasoningTriviaQAScore18featuredBYes
TruthfulQAlegalTruthfulQAScore18featuredBYes
FrontierMathmathFrontierMathScore17featuredBYes
LongBench v2long contextLongBench v2Score17featuredCYes
MM-MT-BenchmultimodalMM-MT-BenchScore17featuredBYes
MVBenchmultimodalMVBenchScore17featuredBYes
OmniDocBench 1.5multimodalOmniDocBench 1.5Score17featuredBYes
Video-MMEmultimodalVideo-MMEScore17featuredBYes
AA-LCRlong contextAA-LCRScore16featuredBYes
ARC-AGI v2reasoningARC-AGI v2Score16featuredBYes
Arena-Hard v2reasoningArena-Hard v2Score16featuredBYes
CharXiv-DmultimodalCharXiv-DScore16featuredBYes
CodeForcesmathCodeForcesNormalized rating16featuredBYes
Hallusion BenchreasoningHallusion BenchScore16featuredBYes
ODinWvisionODinWScore16featuredBYes
ScreenSpotmultimodalScreenSpotScore16featuredBYes
FrontierCode 1.1reasoningFrontierCode 1.1Score15featuredBYes
FrontierSWEagentsFrontierSWEScore15featuredBYes
TextVQAimage to textTextVQAScore15featuredBYes
WritingBenchlegalWritingBenchScore15featuredBYes
Global-MMLU-LitelanguageGlobal-MMLU-LiteScore14featuredBYes
LiveBench 20241125mathLiveBench 20241125Score14featuredBYes
NL2RepoagentsNL2RepoScore14featuredBYes
BFCL-V4agentsBFCL-V4Score13featuredBYes
BLINKmultimodalBLINKScore13featuredBYes
BrowseComp-zhreasoningBrowseComp-zhScore13featuredBYes
Claw-EvalagentsClaw-EvalScore13featuredBYes
Creative Writing v3creativityCreative Writing v3Score13featuredBYes
FACTS GroundingreasoningFACTS GroundingScore13featuredCYes
Global PIQAphysicsGlobal PIQAScore13featuredBYes
HiddenMathmathHiddenMathScore13featuredBYes
Legal Agent BenchmarklegalLegal Agent BenchmarkScore13featuredBYes
MultiPL-ElanguageMultiPL-EScore13featuredBYes
SimpleVQAimage to textSimpleVQAScore13featuredBYes
BBHlanguageBBHScore12featuredBYes
CharadesSTAlanguageCharadesSTAScore12featuredBYes
InfoVQAtestmultimodalInfoVQAtestScore12featuredBYes
MedXpertQAmultimodalMedXpertQAScore12featuredBYes
MT-BenchreasoningMT-BenchNormalized score12featuredBYes
OCRBench-V2 (en)image to textOCRBench-V2 (en)Score12featuredBYes
OptimBenchuncategorizedOptimBenchScore12featuredCNo
BFCLreasoningBFCLScore11featuredBYes
BIG-Bench Extra HardlanguageBIG-Bench Extra HardScore11featuredBYes
CyberGymsafetyCyberGymScore11featuredBYes
DocVQAtestmultimodalDocVQAtestScore11featuredBYes
Graphwalks BFS <128kreasoningGraphwalks BFS <128kScore11featuredBYes
Graphwalks BFS >128klong contextGraphwalks BFS >128kScore11featuredBYes
Graphwalks parents <128kreasoningGraphwalks parents <128kScore11featuredBYes
HMMT Feb 26mathHMMT Feb 26Score11featuredBYes
MAXIFEgeneralMAXIFEScore11featuredBYes
MMMU (val)multimodalMMMU (val)Score11featuredBYes
MuirBenchmultimodalMuirBenchScore11featuredBYes
NOVA-63generalNOVA-63Score11featuredBYes
OCRBench-V2 (zh)image to textOCRBench-V2 (zh)Score11featuredBYes
PIQAphysicsPIQAScore11featuredBYes
AGIEvallegalAGIEvalScore10featuredBYes
Aider-Polyglot EditgeneralAider-Polyglot EditScore10featuredBYes
BoolQlanguageBoolQScore10featuredBYes
COLLIElanguageCOLLIEScore10featuredBYes
DeepSWEagentsDeepSWEScore10featuredBYes
HumanEval+reasoningHumanEval+Score10featuredBYes
MLVUlong contextMLVUScore10featuredBYes
VideoMME w sub.multimodalVideoMME w sub.Score10featuredBYes
VideoMME w/o sub.multimodalVideoMME w/o sub.Score10featuredBYes
VITA-BenchreasoningVITA-BenchScore10featuredBYes

Benchmark values retain their original scales. Open a row to see its model results and scoring direction.

Capability categories

Open a category to compare its benchmarks and leading model results.

reasoning171 benchmarks
Score eligible
95
Top coverage
105
multimodal130 benchmarks
Score eligible
66
Top coverage
66
math66 benchmarks
Score eligible
39
Top coverage
114
language55 benchmarks
Score eligible
34
Top coverage
129
long context53 benchmarks
Score eligible
20
Top coverage
52
agents50 benchmarks
Score eligible
18
Top coverage
20
safety18 benchmarks
Score eligible
9
Top coverage
11
code15 benchmarks
Score eligible
1
Top coverage
4
general15 benchmarks
Score eligible
9
Top coverage
31
image to text12 benchmarks
Score eligible
10
Top coverage
26
knowledge12 benchmarks
Score eligible
1
Top coverage
3
legal10 benchmarks
Score eligible
6
Top coverage
34
instruction following8 benchmarks
Score eligible
3
Top coverage
65
physics7 benchmarks
Score eligible
3
Top coverage
234
spatial reasoning7 benchmarks
Score eligible
6
Top coverage
26
healthcare6 benchmarks
Score eligible
4
Top coverage
9
productivity6 benchmarks
Score eligible
2
Top coverage
3
uncategorized4 benchmarks
Score eligible
0
Top coverage
12
memory3 benchmarks
Score eligible
0
Top coverage
1
structured output3 benchmarks
Score eligible
1
Top coverage
7
audio2 benchmarks
Score eligible
0
Top coverage
1
finance2 benchmarks
Score eligible
0
Top coverage
1
image-generation2 benchmarks
Score eligible
0
Top coverage
1
3d1 benchmarks
Score eligible
0
Top coverage
1
creativity1 benchmarks
Score eligible
1
Top coverage
13
data analysis1 benchmarks
Score eligible
0
Top coverage
1
factuality1 benchmarks
Score eligible
0
Top coverage
1
frontend development1 benchmarks
Score eligible
1
Top coverage
3
psychology1 benchmarks
Score eligible
1
Top coverage
9
question answering1 benchmarks
Score eligible
0
Top coverage
1
research1 benchmarks
Score eligible
0
Top coverage
1
robotics1 benchmarks
Score eligible
0
Top coverage
1
video1 benchmarks
Score eligible
0
Top coverage
1
vision1 benchmarks
Score eligible
1
Top coverage
16

Benchmark highlights

A quick view of benchmark breadth and current model coverage.

Highest model coverageGPQA234 modelsLargest categoryreasoning171 benchmarks
Scoring inputs331Eligible benchmark definitions
Total benchmarks66834 primary categories