llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

reasoning benchmark

ARC-AGI

The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) is a benchmark designed to test general intelligence and abstract reasoning capabilities through visual grid-based transformation tasks. Each task consists of 2-5 demonstration pairs showing input grids transformed into output grids according to underlying rules, with test-takers required to infer these rules and apply them to novel test inputs. The benchmark uses colored grids (up to 30x30) with 10 discrete colors/symbols, designed to measure human-like general fluid intelligence and skill-acquisition efficiency with minimal prior knowledge.

Updated Aug 11, 2026

Models7
Model coverage7
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

ARC-AGI Ranking

Higher score ranks better on this benchmark.

7 rows
Columns

Show columns

01OPGPT-5.5OpenAI95.0%100.0%7CAug 11, 2026
02OPGPT-5.4OpenAI93.7%83.3%7CAug 11, 2026
03OPGPT-5.2 ProOpenAI90.5%66.7%7CAug 11, 2026
04OPo3OpenAI88.0%50.0%7CAug 11, 2026
05OPGPT-5.2OpenAI86.2%33.3%7CAug 11, 2026
06MELongCat-Flash-ThinkingMeituan50.3%16.7%7CAug 11, 2026
07ACQwen3-235B-A22B-Instruct-2507Alibaba Cloud / Qwen Team41.8%0.0%7CAug 11, 2026

ARC-AGI Score Distribution

A closer view of the leading scores on this benchmark.

ARC-AGI

ARC-AGI Highlights

The leading models and scores on this benchmark.

Rank #1GPT-5.595.0%Rank #2GPT-5.493.7%Rank #3GPT-5.2 Pro90.5%Rank #4o388.0%

What is ARC-AGI?

What ARC-AGI measures and how its scores work.

The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) is a benchmark designed to test general intelligence and abstract reasoning capabilities through visual grid-based transformation tasks. Each task consists of 2-5 demonstration pairs showing input grids transformed into output grids according to underlying rules, with test-takers required to infer these rules and apply them to novel test inputs. The benchmark uses colored grids (up to 30x30) with 10 discrete colors/symbols, designed to measure human-like general fluid intelligence and skill-acquisition efficiency with minimal prior knowledge.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.

Family
ARC-AGI
Modality
image
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
arc-agi|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about ARC-AGI.

Which model scores highest on ARC-AGI?

GPT-5.5 is currently ranked first with 95.0%.

What does ARC-AGI measure?

The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI) is a benchmark designed to test general intelligence and abstract reasoning capabilities through visual grid-based transformation tasks. Each task consists of 2-5 demonstration pairs showing input grids transformed into output grids according to underlying rules, with test-takers required to infer these rules and apply them to novel test inputs. The benchmark uses colored grids (up to 30x30) with 10 discrete colors/symbols, designed to measure human-like general fluid intelligence and skill-acquisition efficiency with minimal prior knowledge.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

7 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.