llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

reasoning benchmark

LiveCodeBench

LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from programming contests (LeetCode, AtCoder, CodeForces) and evaluates four different scenarios: code generation, self-repair, code execution, and test output prediction. Problems are annotated with release dates to enable evaluation on unseen problems released after a model's training cutoff.

Updated Aug 11, 2026

Models73
Model coverage73
MetricScore
EvidenceC

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

LiveCodeBench Ranking

Higher score ranks better on this benchmark.

73 rows
Columns

Show columns

01DEDeepSeek-V4-Pro-MaxDeepSeek93.5%100.0%73CAug 11, 2026
02DEDeepSeek-V4-Flash-MaxDeepSeek91.6%98.6%73CAug 11, 2026
03DEDeepSeek-V3.2 (Thinking)DeepSeek83.3%97.2%73CAug 11, 2026
04DEDeepSeek-V3.2DeepSeek83.3%95.8%73CAug 11, 2026
05MIMiniMax M2MiniMax83.0%94.4%73CAug 11, 2026
06MELongCat-Flash-Thinking-2601Meituan82.8%93.1%73CAug 11, 2026
07NVNemotron 3 Super (120B A12B)NVIDIA81.2%91.7%73CAug 11, 2026
08XAGrok-3 MinixAI80.4%90.3%73CAug 11, 2026
09XAGrok 4 FastxAI80.0%88.9%73CAug 11, 2026
10XAGrok-3xAI79.4%87.5%73CAug 11, 2026
11XAGrok-4 HeavyxAI79.4%86.1%73CAug 11, 2026
12MELongCat-Flash-ThinkingMeituan79.4%84.7%73CAug 11, 2026
13XAGrok-4xAI79.0%83.3%73CAug 11, 2026
14MIMiniMax M2.1MiniMax78.0%81.9%73CAug 11, 2026
15AMNova 2 ProAmazon74.6%80.6%73CAug 11, 2026
16DEDeepSeek-V3.2-ExpDeepSeek74.1%79.2%73CAug 11, 2026
17DEDeepSeek-R1-0528DeepSeek73.3%77.8%73CAug 11, 2026
18ZAGLM-4.5Zhipu AI72.9%76.4%73CAug 11, 2026
19NVNemotron Nano 9B v2NVIDIA71.1%75.0%73CAug 11, 2026
20AMNova 2 LiteAmazon71.0%73.6%73CAug 11, 2026
21ZAGLM-4.5-AirZhipu AI70.7%72.2%73CAug 11, 2026
22ACQwen3 235B A22BAlibaba Cloud / Qwen Team70.7%70.8%73CAug 11, 2026
23GOGemini 2.5 Pro Preview 06-05Google69.0%69.4%73CAug 11, 2026
24INMercury 2Inception67.0%68.1%73CAug 11, 2026
25NVLlama 3.1 Nemotron Ultra 253B v1NVIDIA66.3%66.7%73CAug 11, 2026
26ACQwen3 32BAlibaba Cloud / Qwen Team65.7%65.3%73CAug 11, 2026
27MIMiniMax M1 80KMiniMax65.0%63.9%73CAug 11, 2026
28MAMinistral 3 (14B Reasoning 2512)Mistral AI64.6%62.5%73CAug 11, 2026
29MAMistral Small 4Mistral AI63.6%61.1%73CAug 11, 2026
30ACQwQ-32BAlibaba Cloud / Qwen Team63.4%59.7%73CAug 11, 2026
31ACQwen3 30B A3BAlibaba Cloud / Qwen Team62.6%58.3%73CAug 11, 2026
32MIMiniMax M1 40KMiniMax62.3%56.9%73CAug 11, 2026
33MAMinistral 3 (8B Reasoning 2512)Mistral AI61.6%55.6%73CAug 11, 2026
34DEDeepSeek R1 Distill Llama 70BDeepSeek57.5%54.2%73CAug 11, 2026
35DEDeepSeek R1 Distill Qwen 32BDeepSeek57.2%52.8%73CAug 11, 2026
36DEDeepSeek-V3.1DeepSeek56.4%51.4%73CAug 11, 2026
37ACQwen2.5 72B InstructAlibaba Cloud / Qwen Team55.5%50.0%73CAug 11, 2026
38MAMin istral 3 (3B Reasoning 2512)Mistral AI54.8%48.6%73CAug 11, 2026
39MIPhi 4 ReasoningMicrosoft53.8%47.2%73CAug 11, 2026
40MAKimi K2-Instruct-0905Moonshot AI53.7%45.8%73CAug 11, 2026
41DEDeepSeek R1 Distill Qwen 14BDeepSeek53.1%44.4%73CAug 11, 2026
42MIPhi 4 Reasoning PlusMicrosoft53.1%43.1%73CAug 11, 2026
43MAMagistral Small 2506Mistral AI51.3%41.7%73CAug 11, 2026
44MAMagistral MediumMistral AI50.3%40.3%73CAug 11, 2026
45DEDeepSeek R1 ZeroDeepSeek50.0%38.9%73CAug 11, 2026
46ACQwQ-32B-PreviewAlibaba Cloud / Qwen Team50.0%37.5%73CAug 11, 2026
47DEDeepSeek-V3 0324DeepSeek49.2%36.1%73CAug 11, 2026
48MELongCat-Flash-ChatMeituan48.0%34.7%73CAug 11, 2026
49MELlama 4 MaverickMeta43.4%33.3%73CAug 11, 2026
50DEDeepSeek R1 Distill Llama 8BDeepSeek39.6%31.9%73CAug 11, 2026
51DEDeepSeek R1 Distill Qwen 7BDeepSeek37.6%30.6%73CAug 11, 2026
52DEDeepSeek-V3DeepSeek37.6%29.2%73CAug 11, 2026
53GOGemini 2.0 FlashGoogle35.1%27.8%73CAug 11, 2026
54MAMistral Large 3 (675B Instruct 2512 Eagle)Mistral AI34.4%26.4%73CAug 11, 2026
55MAMistral Large 3 (675B Base)Mistral AI34.4%25.0%73CAug 11, 2026
56MAMistral Large 3 (675B Instruct 2512 NVFP4)Mistral AI34.4%23.6%73CAug 11, 2026
57MAMistral Large 3 (675B Instruct 2512)Mistral AI34.4%22.2%73CAug 11, 2026
58GOGemini 2.5 Flash-LiteGoogle33.7%20.8%73CAug 11, 2026
59MELlama 4 ScoutMeta32.8%19.4%73CAug 11, 2026
60ACQwen2.5-Coder 32B InstructAlibaba Cloud / Qwen Team31.4%18.1%73CAug 11, 2026
61GOGemini DiffusionGoogle30.9%16.7%73CAug 11, 2026
62GOGemma 3 27BGoogle29.7%15.3%73CAug 11, 2026
63ACQwen2.5 7B InstructAlibaba Cloud / Qwen Team28.7%13.9%73CAug 11, 2026
64ACQwen2 7B InstructAlibaba Cloud / Qwen Team26.6%12.5%73CAug 11, 2026
65GOGemma 3 12BGoogle24.6%11.1%73CAug 11, 2026
66ACQwen2.5-Coder 7B InstructAlibaba Cloud / Qwen Team18.2%9.7%73CAug 11, 2026
67DEDeepSeek R1 Distill Qwen 1.5BDeepSeek16.9%8.3%73CAug 11, 2026
68GOGemma 3n E2B InstructedGoogle13.2%6.9%73CAug 11, 2026
69GOGemma 3n E2B Instructed LiteRT (Preview)Google13.2%5.6%73CAug 11, 2026
70GOGemma 3n E4B InstructedGoogle13.2%4.2%73CAug 11, 2026
71GOGemma 3n E4B Instructed LiteRT PreviewGoogle13.2%2.8%73CAug 11, 2026
72GOGemma 3 4BGoogle12.6%1.4%73CAug 11, 2026
73GOGemma 3 1BGoogle1.9%0.0%73CAug 11, 2026

LiveCodeBench Score Distribution

A closer view of the leading scores on this benchmark.

LiveCodeBench

LiveCodeBench Highlights

The leading models and scores on this benchmark.

Rank #1DeepSeek-V4-Pro-Max93.5%Rank #2DeepSeek-V4-Flash-Max91.6%Rank #3DeepSeek-V3.2 (Thinking)83.3%Rank #4DeepSeek-V3.283.3%

What is LiveCodeBench?

What LiveCodeBench measures and how its scores work.

LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from programming contests (LeetCode, AtCoder, CodeForces) and evaluates four different scenarios: code generation, self-repair, code execution, and test output prediction. Problems are annotated with release dates to enable evaluation on unseen problems released after a model's training cutoff.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of C.

Family
LiveCodeBench
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
livecodebench|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about LiveCodeBench.

Which model scores highest on LiveCodeBench?

DeepSeek-V4-Pro-Max is currently ranked first with 93.5%.

What does LiveCodeBench measure?

LiveCodeBench is a holistic and contamination-free evaluation benchmark for large language models for code. It continuously collects new problems from programming contests (LeetCode, AtCoder, CodeForces) and evaluates four different scenarios: code generation, self-repair, code execution, and test output prediction. Problems are annotated with release dates to enable evaluation on unseen problems released after a model's training cutoff.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

73 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.