llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

LLMBoard ranking

Agent Model Leaderboard

This agent model leaderboard compares the current ranking from the LLMBoard Agent Score, eligible agent benchmarks, official token pricing and supporting task evidence.

Data as of 2026-08-17

On this page

  • Ranking
  • Overview
  • Agent score
  • Composition
  • Top models
  • FAQ

Agent Model Ranking

This leaderboard ranking orders models by their domain-specific LLMBoard score aggregated from eligible benchmark evidence. Pricing, coverage and the source signal remain separate.

30 of 39 rows
Columns

Show columns

Sort by
Rank
Model
LLMBoard Agent Score
Official input / 1M
Official output / 1M
Score updated
Rank01ModelANClaude Fable 5AnthropicLLMBoard Agent Score87.3Official input / 1M$10Official output / 1M$50Score updatedAug 18, 2026
Rank02ModelANClaude Opus 5AnthropicLLMBoard Agent Score85.8Official input / 1M$5Official output / 1M$25Score updatedAug 18, 2026
Rank03ModelOPGPT-5.6-SolOpenAILLMBoard Agent Score85.7Official input / 1M$5Official output / 1M$30Score updatedAug 18, 2026
Rank04ModelMAKimi K3Moonshot AILLMBoard Agent Score85.3Official input / 1M$3Official output / 1M$15Score updatedAug 18, 2026
Rank05ModelOPGPT-5.5OpenAILLMBoard Agent Score77.7Official input / 1M$5Official output / 1M$30Score updatedAug 18, 2026
Rank06ModelXAGrok 4.5xAILLMBoard Agent Score77.0Official input / 1M$2Official output / 1M$6Score updatedAug 18, 2026
Rank07ModelANClaude Opus 4.6AnthropicLLMBoard Agent Score75.7Official input / 1M$5Official output / 1M$25Score updatedAug 18, 2026
Rank08ModelANClaude Opus 4.7AnthropicLLMBoard Agent Score75.6Official input / 1M$5Official output / 1M$25Score updatedAug 18, 2026
Rank09ModelZAGLM 5.2Zhipu AILLMBoard Agent Score74.6Official input / 1M$1.4Official output / 1M$4.4Score updatedAug 18, 2026
Rank10ModelOPGPT-5.4OpenAILLMBoard Agent Score71.7Official input / 1M$2.5Official output / 1M$15Score updatedAug 18, 2026
Rank11ModelANClaude Sonnet 5AnthropicLLMBoard Agent Score69.6Official input / 1M$2Official output / 1M$10Score updatedAug 18, 2026
Rank12ModelACQwen3.8 MaxAlibaba Cloud / Qwen TeamLLMBoard Agent Score67.0Official input / 1M$2Official output / 1M$6Score updatedAug 18, 2026
Rank13ModelOPGPT-5.6-TerraOpenAILLMBoard Agent Score66.9Official input / 1M$2Official output / 1M$12Score updatedAug 18, 2026
Rank14ModelOPGPT-5.6-LunaOpenAILLMBoard Agent Score66.7Official input / 1M$0.20Official output / 1M$1.2Score updatedAug 18, 2026
Rank15ModelDEDeepSeek-V4-FlashDeepSeekLLMBoard Agent Score63.6Official input / 1M$0.14Official output / 1M$0.28Score updatedAug 18, 2026
Rank16ModelANClaude Opus 4.8AnthropicLLMBoard Agent Score62.9Official input / 1M$5Official output / 1M$25Score updatedAug 18, 2026
Rank17ModelANClaude Sonnet 4.6AnthropicLLMBoard Agent Score58.1Official input / 1M$3Official output / 1M$15Score updatedAug 18, 2026
Rank18ModelMAKimi K2.7 CodeMoonshot AILLMBoard Agent Score55.5Official input / 1M$0.95Official output / 1M$4Score updatedAug 18, 2026
Rank19ModelMEMuse Spark 1.1MetaLLMBoard Agent Score54.3Official input / 1M$1.25Official output / 1M$4.25Score updatedAug 18, 2026
Rank20ModelMAKimi K2.6Moonshot AILLMBoard Agent Score49.4Official input / 1M$0.95Official output / 1M$4Score updatedAug 18, 2026
Rank21ModelGOGemini 3.1 ProGoogleLLMBoard Agent Score48.4Official input / 1M$2Official output / 1M$12Score updatedAug 18, 2026
Rank22ModelACQwen3.7 MaxAlibaba Cloud / Qwen TeamLLMBoard Agent Score45.0Official input / 1M$2.5Official output / 1M$7.5Score updatedAug 18, 2026
Rank23ModelZAGLM 5.1Zhipu AILLMBoard Agent Score43.6Official input / 1M$1.4Official output / 1M$4.4Score updatedAug 18, 2026
Rank24ModelGOGemini 3.5 FlashGoogleLLMBoard Agent Score41.5Official input / 1M$1.5Official output / 1M$9Score updatedAug 18, 2026
Rank25ModelDEDeepSeek-V4-ProDeepSeekLLMBoard Agent Score41.1Official input / 1M$0.435Official output / 1M$0.87Score updatedAug 18, 2026
Rank26ModelACQwen3.7 PlusAlibaba Cloud / Qwen TeamLLMBoard Agent Score35.3Official input / 1M$0.50Official output / 1M$3Score updatedAug 18, 2026
Rank27ModelMIMiniMax M3MiniMaxLLMBoard Agent Score35.2Official input / 1M$0.30Official output / 1M$1.2Score updatedAug 18, 2026
Rank28ModelXIMiMo V2.5 ProXiaomiLLMBoard Agent Score32.0Official input / 1M$0.435Official output / 1M$0.87Score updatedAug 18, 2026
Rank29ModelTEHy3TencentLLMBoard Agent Score32.0Official input / 1MN/AOfficial output / 1MN/AScore updatedAug 18, 2026
Rank30ModelTHInklingThinkingmachinesLLMBoard Agent Score26.0Official input / 1M$1.87Official output / 1M$4.68Score updatedAug 18, 2026
Rank31ModelXAGrok Build 0.1xAILLMBoard Agent Score25.6Official input / 1M$1Official output / 1M$2Score updatedAug 18, 2026
Rank32ModelXAGrok 4.3xAILLMBoard Agent Score23.7Official input / 1M$1.25Official output / 1M$2.5Score updatedAug 18, 2026
Rank33ModelGOGemini 3 FlashGoogleLLMBoard Agent Score20.9Official input / 1M$0.50Official output / 1M$3Score updatedAug 18, 2026
Rank34ModelGOGemma 4 31BGoogleLLMBoard Agent Score20.3Official input / 1MN/AOfficial output / 1MN/AScore updatedAug 18, 2026
Rank35ModelMIMiniMax M2.7MiniMaxLLMBoard Agent Score19.7Official input / 1M$0.30Official output / 1M$1.2Score updatedAug 18, 2026
Rank36ModelUPSolar Pro 4UpstageLLMBoard Agent Score19.2Official input / 1M$0.30Official output / 1M$1.2Score updatedAug 18, 2026
Rank37ModelMAMistral Medium 3.5Mistral AILLMBoard Agent Score19.2Official input / 1M$1.5Official output / 1M$7.5Score updatedAug 18, 2026
Rank38ModelGOGemini 3.5 Flash LiteGoogleLLMBoard Agent Score15.8Official input / 1M$0.30Official output / 1M$2.5Score updatedAug 18, 2026
Rank39ModelNVNemotron 3 UltraNVIDIALLMBoard Agent Score14.4Official input / 1M$0.50Official output / 1M$2.5Score updatedAug 18, 2026

Leaderboard ranking overview

This leaderboard overview connects current LLMBoard leaders with benchmark coverage and supporting source activity.

Ranked models39
SignalLLMBoard Agent Score
Top LLMBoard modelClaude Fable 587.3 score · 100% coverageClosest challengerClaude Opus 51.42 LLMBoard points behindMost observationsGrok Build 0.15,456,874 observations

Agent score and confidence

Use confidence and observed task volume to judge how firmly each leaderboard position is supported; these source proportions remain separate from benchmark coverage.

Claude Fable 5Anthropic
12.01%
Claude Opus 5Anthropic
12.19%
GPT-5.6 SolOpenAI
10.86%
Kimi K3Moonshot AI
10.60%
GPT-5.5OpenAI
8.90%
Grok 4.5xAI
6.19%
Claude Opus 4.6Anthropic
6.89%
Claude Opus 4.7Anthropic
7.66%
GLM-5.2Zhipu AI
6.74%
GPT-5.4OpenAI
4.99%
Claude Sonnet 5Anthropic
7.14%
Qwen3.8 MaxAlibaba Cloud / Qwen Team
7.61%
4.05%Source score with confidence interval14.58%
0%4%8%11%15%120.8K336K934.8K2.6M7.2MTAgent scoreObservations
39 models with comparable dataOpen a model by selecting its logo

Agent Model Leaderboard Vendor Mix

Vendor representation among the top-ranked models in this leaderboard.

Top-entry vendor mix

25 models
Anthropic7OpenAI5Moonshot AI3Zhipu AI2Alibaba Cloud / Qwen Team2DeepSeek2Google2xAI1

The Top AI Models for Agent

A practical summary of the first five entries in this agent AI model leaderboard, with benchmark evidence and pricing kept in context.

Ranking basisThis agent AI model leaderboard uses the domain-specific LLMBoard score, evidence coverage and status. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    AN
    Claude Fable 5Anthropic
    LLMBoard Agent Score
    87.3
    Price
    $10 input / $50 output per 1M tokens
    Speed
    Up to 9.1 tok/s via Anthropic

    Strengths

    • Ranks #1 on the current agent LLMBoard ranking
    • 911,709 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score
  2. 02
    AN
    Claude Opus 5Anthropic
    LLMBoard Agent Score
    85.8
    Price
    $5.0 input / $25 output per 1M tokens
    Speed
    Up to 3.3 tok/s via Anthropic

    Strengths

    • Ranks #2 on the current agent LLMBoard ranking
    • 2,246,557 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score
  3. 03
    OP
    GPT-5.6-SolOpenAI
    LLMBoard Agent Score
    85.7
    Price
    $5.0 input / $30 output per 1M tokens
    Speed
    Up to 27 tok/s via OpenAI

    Strengths

    • Ranks #3 on the current agent LLMBoard ranking
    • 1,599,076 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score
  4. 04
    MA
    Kimi K3Moonshot AI
    LLMBoard Agent Score
    85.3
    Price
    $3.0 input / $15 output per 1M tokens
    Speed
    Up to 26 tok/s via Fireworks

    Strengths

    • Ranks #4 on the current agent LLMBoard ranking
    • 2,431,932 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score
  5. 05
    OP
    GPT-5.5OpenAI
    LLMBoard Agent Score
    77.7
    Price
    $5.0 input / $30 output per 1M tokens

    Strengths

    • Ranks #5 on the current agent LLMBoard ranking
    • 2,028,364 observations support the source signal
    • A matched official token price is available

    Considerations

    • The source Arena signal remains visible separately from the aggregated LLMBoard score

Selection summary

Best AI Models for Agent

Claude Fable 5 currently leads the agent ranking at 87.3. Compare evidence coverage, source activity and price separately before choosing a model for production.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

LLMBoard rank #1Claude Fable 587.3 · $10 input / $50 output per 1M tokensLLMBoard rank #2Claude Opus 585.8 · $5.0 input / $25 output per 1M tokensLLMBoard rank #3GPT-5.6-Sol85.7 · $5.0 input / $30 output per 1M tokens

About this leaderboard ranking

Common questions about the Agent Model Leaderboard.

What is the top model on Agent Model Leaderboard?

Claude Fable 5 is currently ranked first with a LLMBoard Agent Score of 87.3.

How does this leaderboard use benchmarks?

The LLMBoard score converts eligible source categories and benchmark families to comparable dimension scores, applies domain weights and reports coverage separately. The leaderboard ranking follows that score.

How many models are shown?

This page currently compares 39 models.

Why can a model be missing?

A model appears only after LLMBoard has calculated a score for the domain used by this ranking. A source Arena result is optional supporting evidence and does not control whether a scored model appears.