llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

reasoning benchmark

SimpleQA

SimpleQA is a factuality benchmark developed by OpenAI that measures the short-form factual accuracy of large language models. The benchmark contains 4,326 short, fact-seeking questions that are adversarially collected and designed to have single, indisputable answers. Questions cover diverse topics from science and technology to entertainment, and the benchmark also measures model calibration by evaluating whether models know what they know.

Updated Aug 11, 2026

Models46
Model coverage46
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

SimpleQA Ranking

Higher score ranks better on this benchmark.

46 rows
Columns

Show columns

01DEDeepSeek-V3.2-ExpDeepSeek97.1%100.0%46CAug 11, 2026
02XAGrok 4 FastxAI95.0%97.8%46CAug 11, 2026
03DEDeepSeek-V3.1DeepSeek93.4%95.6%46CAug 11, 2026
04DEDeepSeek-R1-0528DeepSeek92.3%93.3%46CAug 11, 2026
05BAERNIE 5.0Baidu75.0%91.1%46CAug 11, 2026
06GOGemini 3 ProGoogle72.1%88.9%46CAug 11, 2026
07GOGemini 3 FlashGoogle68.7%86.7%46CAug 11, 2026
08OPGPT-4.5OpenAI62.5%84.4%46CAug 11, 2026
09DEDeepSeek-V4-Pro-MaxDeepSeek57.9%82.2%46CAug 11, 2026
10ACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team55.4%80.0%46CAug 11, 2026
11ACQwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team54.3%77.8%46CAug 11, 2026
12GOGemini 2.5 Pro Preview 06-05Google54.0%75.6%46CAug 11, 2026
13ACQwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team51.9%73.3%46CAug 11, 2026
14GOGemini 2.5 ProGoogle50.8%71.1%46CAug 11, 2026
15ACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team49.6%68.9%46CAug 11, 2026
16ACQwen3 VL 4B InstructAlibaba Cloud / Qwen Team48.0%66.7%46CAug 11, 2026
17OPo1OpenAI47.0%64.4%46CAug 11, 2026
18ACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team44.4%62.2%46CAug 11, 2026
19GOGemini 3.1 Flash-LiteGoogle43.3%60.0%46CAug 11, 2026
20OPo1-previewOpenAI42.4%57.8%46CAug 11, 2026
21OPGPT-4oOpenAI38.2%55.6%46CAug 11, 2026
22MAKimi K2 BaseMoonshot AI35.3%53.3%46CAug 11, 2026
23DEDeepSeek-V4-Flash-MaxDeepSeek34.1%51.1%46CAug 11, 2026
24MAKimi K2 InstructMoonshot AI31.0%48.9%46CAug 11, 2026
25MAKimi K2-Instruct-0905Moonshot AI31.0%46.7%46CAug 11, 2026
26ACQwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team27.0%44.4%46CAug 11, 2026
27GOGemini 2.5 FlashGoogle26.9%42.2%46CAug 11, 2026
28DEDeepSeek-V3DeepSeek24.9%40.0%46CAug 11, 2026
29ACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team23.9%37.8%46CAug 11, 2026
30MAMistral Large 3 (675B Instruct 2512 Eagle)Mistral AI23.8%35.6%46CAug 11, 2026
31MAMistral Large 3 (675B Base)Mistral AI23.8%33.3%46CAug 11, 2026
32MAMistral Large 3 (675B Instruct 2512 NVFP4)Mistral AI23.8%31.1%46CAug 11, 2026
33MAMistral Large 3 (675B Instruct 2512)Mistral AI23.8%28.9%46CAug 11, 2026
34GOGemini 2.0 Flash-LiteGoogle21.7%26.7%46CAug 11, 2026
35MIMiniMax M1 80KMiniMax18.5%24.4%46CAug 11, 2026
36MIMiniMax M1 40KMiniMax17.9%22.2%46CAug 11, 2026
37OPo3-miniOpenAI15.0%20.0%46CAug 11, 2026
38MAMistral Small 3.2 24B InstructMistral AI12.1%17.8%46CAug 11, 2026
39GOGemini 2.5 Flash-LiteGoogle10.7%15.6%46CAug 11, 2026
40MAMistral Small 3.1 24B InstructMistral AI10.4%13.3%46CAug 11, 2026
41GOGemma 3 27BGoogle10.0%11.1%46CAug 11, 2026
42GOGemma 3 12BGoogle6.3%8.9%46CAug 11, 2026
43GOGemma 3 4BGoogle4.0%6.7%46CAug 11, 2026
44MIPhi 4Microsoft3.0%4.4%46CAug 11, 2026
45GOGemma 3 1BGoogle2.2%2.2%46CAug 11, 2026
46BAERNIE 4.5Baidu1.8%0.0%46CAug 11, 2026

SimpleQA Score Distribution

A closer view of the leading scores on this benchmark.

SimpleQA

SimpleQA Highlights

The leading models and scores on this benchmark.

Rank #1DeepSeek-V3.2-Exp97.1%Rank #2Grok 4 Fast95.0%Rank #3DeepSeek-V3.193.4%Rank #4DeepSeek-R1-052892.3%

What is SimpleQA?

What SimpleQA measures and how its scores work.

SimpleQA is a factuality benchmark developed by OpenAI that measures the short-form factual accuracy of large language models. The benchmark contains 4,326 short, fact-seeking questions that are adversarially collected and designed to have single, indisputable answers. Questions cover diverse topics from science and technology to entertainment, and the benchmark also measures model calibration by evaluating whether models know what they know.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.

Family
SimpleQA
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
simpleqa|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about SimpleQA.

Which model scores highest on SimpleQA?

DeepSeek-V3.2-Exp is currently ranked first with 97.1%.

What does SimpleQA measure?

SimpleQA is a factuality benchmark developed by OpenAI that measures the short-form factual accuracy of large language models. The benchmark contains 4,326 short, fact-seeking questions that are adversarially collected and designed to have single, indisputable answers. Questions cover diverse topics from science and technology to entertainment, and the benchmark also measures model calibration by evaluating whether models know what they know.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

46 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.