llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

Arena benchmark

LM Arena Agent Praise Complaint Leaderboard

LM Arena Agent Praise Complaint model ratings and results.

Updated Aug 17, 2026. Data as of 2026-08-17

Models40
Categories1
Result typeagent_score

On this page

  • Ranking
  • Highlights
  • Distribution
  • Top models
  • About
  • FAQ

Agent Praise Complaint Ranking

Models are ordered by their Arena rank.

30 of 40 rows
Columns

Show columns

Sort by
Rank
Model
Organization
Rating / score
Votes
Observations
Result date
Rank01ModelANClaude Fable 5AnthropicOrganizationanthropicRating / score0.3VotesN/AObservations5.1KResult dateAug 13, 2026
Rank02ModelOPGPT-5.6 SolOpenAIOrganizationopenaiRating / score0.2VotesN/AObservations4.4KResult dateAug 13, 2026
Rank04ModelMAKimi K3Moonshot AIOrganizationmoonshotRating / score0.2VotesN/AObservations8.9KResult dateAug 13, 2026
Rank05ModelANClaude Opus 5AnthropicOrganizationanthropicRating / score0.2VotesN/AObservations3.8KResult dateAug 13, 2026
Rank07ModelANClaude Sonnet 5AnthropicOrganizationanthropicRating / score0.2VotesN/AObservations5.3KResult dateAug 13, 2026
Rank08ModelOPGPT-5.5OpenAIOrganizationopenaiRating / score0.1VotesN/AObservations11.4KResult dateAug 13, 2026
Rank09ModelANClaude Opus 4.8AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations6.3KResult dateAug 13, 2026
Rank11ModelACQwen3.8 MaxAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score0.1VotesN/AObservations2KResult dateAug 13, 2026
Rank13ModelZAGLM-5.2Zhipu AIOrganizationzaiRating / score0.1VotesN/AObservations10.5KResult dateAug 13, 2026
Rank14ModelANClaude Opus 4.7AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations8KResult dateAug 13, 2026
Rank15ModelANClaude Opus 4.6AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations9.3KResult dateAug 13, 2026
Rank16ModelOPGPT-5.6 LunaOpenAIOrganizationopenaiRating / score0.1VotesN/AObservations2.4KResult dateAug 13, 2026
Rank18ModelXAGrok 4.5xAIOrganizationxaiRating / score0.1VotesN/AObservations6.4KResult dateAug 13, 2026
Rank19ModelOPGPT-5.4OpenAIOrganizationopenaiRating / score0.0VotesN/AObservations15KResult dateAug 13, 2026
Rank20ModelGOGemini 3.1 ProGoogleOrganizationgoogleRating / score0.0VotesN/AObservations23.5KResult dateAug 13, 2026
Rank21ModelOPGPT-5.6 TerraOpenAIOrganizationopenaiRating / score0.0VotesN/AObservations3.9KResult dateAug 13, 2026
Rank22ModelMAKimi K2.7 CodeMoonshot AIOrganizationmoonshotRating / score0.0VotesN/AObservations1.7KResult dateAug 13, 2026
Rank23ModelDEDeepSeek-V4-Flash-0731DeepSeekOrganizationdeepseekRating / score0.0VotesN/AObservations8KResult dateAug 13, 2026
Rank24ModelMAKimi K2.6Moonshot AIOrganizationmoonshotRating / score0.0VotesN/AObservations1.8KResult dateAug 13, 2026
Rank26ModelTEHy3TencentOrganizationtencentRating / score0.0VotesN/AObservations2.6KResult dateAug 13, 2026
Rank27ModelANClaude Sonnet 4.6AnthropicOrganizationanthropicRating / score0.0VotesN/AObservations9.6KResult dateAug 13, 2026
Rank28ModelZAGLM-5.1Zhipu AIOrganizationzaiRating / score0.0VotesN/AObservations15.2KResult dateAug 13, 2026
Rank29ModelGOGemini 3.5 FlashGoogleOrganizationgoogleRating / score-0.0VotesN/AObservations25.8KResult dateAug 13, 2026
Rank30ModelDEDeepSeek V4 ProDeepSeekOrganizationdeepseekRating / score-0.0VotesN/AObservations8.4KResult dateAug 13, 2026
Rank31ModelGOGemma 4 31BGoogleOrganizationgoogleRating / score-0.0VotesN/AObservations10.7KResult dateAug 13, 2026
Rank33ModelACQwen3.7 MaxAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score-0.0VotesN/AObservations8.3KResult dateAug 13, 2026
Rank34ModelMEMuse Spark 1.1MetaOrganizationmetaRating / score-0.0VotesN/AObservations16.8KResult dateAug 13, 2026
Rank36ModelXIMiMo-V2.5-ProXiaomiOrganizationxiaomiRating / score-0.1VotesN/AObservations7.5KResult dateAug 13, 2026
Rank37ModelMIMiniMax M3MiniMaxOrganizationminimaxRating / score-0.1VotesN/AObservations7.9KResult dateAug 13, 2026
Rank38ModelDEDeepSeek V4 FlashDeepSeekOrganizationdeepseekRating / score-0.1VotesN/AObservations6.2KResult dateAug 13, 2026
Rank39ModelACQwen3.7-PlusAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score-0.1VotesN/AObservations4.3KResult dateAug 13, 2026
Rank40ModelMAMistral Medium 3.5Mistral AIOrganizationmistralRating / score-0.1VotesN/AObservations1.1KResult dateAug 13, 2026
Rank41ModelGOGemini 3 FlashGoogleOrganizationgoogleRating / score-0.1VotesN/AObservations22.4KResult dateAug 13, 2026
Rank42ModelXAGrok Build 0.1xAIOrganizationxaiRating / score-0.1VotesN/AObservations14.1KResult dateAug 13, 2026
Rank43ModelXAGrok 4.3xAIOrganizationxaiRating / score-0.1VotesN/AObservations16KResult dateAug 13, 2026
Rank44ModelGOGemini 3.5 Flash-LiteGoogleOrganizationgoogleRating / score-0.1VotesN/AObservations5.2KResult dateAug 13, 2026
Rank45ModelNVNemotron 3 Ultra (550B A55B)NVIDIAOrganizationnvidiaRating / score-0.1VotesN/AObservations924Result dateAug 13, 2026
Rank46ModelMIMiniMax M2.7MiniMaxOrganizationminimaxRating / score-0.2VotesN/AObservations4.4KResult dateAug 13, 2026
Rank47ModelUPSolar Pro 4UpstageOrganizationupstageRating / score-0.2VotesN/AObservations805Result dateAug 13, 2026
Rank49ModelTHInklingThinkingmachinesOrganizationthinkyRating / score-0.2VotesN/AObservations7.7KResult dateAug 13, 2026

Arena highlights

Rank #1Claude Fable 50.3 agent scoreRank #2GPT-5.6 Sol0.2 agent scoreRank #4Kimi K30.2 agent scoreRank #5Claude Opus 50.2 agent score

Agent Praise Complaint Rating Distribution

A closer view of the leading model ratings.

Agent Praise Complaint

The Top AI Models for Agent Praise Complaint

The first five models in this source ranking, with price and output speed included when a canonical model match is available.

Ranking basisThis agent praise complaint AI model leaderboard uses the source agent score and Arena rank. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    AN
    Claude Fable 5Anthropic
    agent score
    0.25
    Price
    $10 input / $50 output per 1M tokens
    Speed
    Up to 9.1 tok/s via Anthropic

    Strengths

    • Ranks #1 on Agent Praise Complaint
    • 5,077 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  2. 02
    OP
    GPT-5.6 SolOpenAI
    agent score
    0.23
    Price
    $5.0 input / $30 output per 1M tokens
    Speed
    Up to 27 tok/s via OpenAI

    Strengths

    • Ranks #2 on Agent Praise Complaint
    • 4,369 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  3. 04
    MA
    Kimi K3Moonshot AI
    agent score
    0.20
    Price
    $3.0 input / $15 output per 1M tokens
    Speed
    Up to 26 tok/s via Fireworks

    Strengths

    • Ranks #4 on Agent Praise Complaint
    • 8,906 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  4. 05
    AN
    Claude Opus 5Anthropic
    agent score
    0.19
    Price
    $5.0 input / $25 output per 1M tokens
    Speed
    Up to 3.3 tok/s via Anthropic

    Strengths

    • Ranks #5 on Agent Praise Complaint
    • 3,781 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  5. 07
    AN
    Claude Sonnet 5Anthropic
    agent score
    0.15
    Price
    $2.0 input / $10 output per 1M tokens
    Speed
    Up to 11 tok/s via Anthropic

    Strengths

    • Ranks #7 on Agent Praise Complaint
    • 5,261 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score

Selection summary

Best AI Models for Agent Praise Complaint

Claude Fable 5 currently leads Agent Praise Complaint at 0.25. The best AI model for this use case may change when price, speed and independent benchmark capability are considered.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

Arena rank #1Claude Fable 5$10 input / $50 output per 1M tokensArena rank #2GPT-5.6 Sol$5.0 input / $30 output per 1M tokensArena rank #4Kimi K3$3.0 input / $15 output per 1M tokens

What is the Agent Praise Complaint Arena?

What Agent Praise Complaint measures and how to read its results.

Agent Praise Complaint reports agent score values and currently defines 1 evaluation category.

Arena results are preference or task-outcome signals. They are not automatically equivalent to capability benchmark scores.

FAQ

Common questions about Agent Praise Complaint.

Which model leads Agent Praise Complaint?

Claude Fable 5 is currently ranked first.

What does Agent Praise Complaint measure?

This Arena reports agent score values across 1 configured evaluation category.

Does this result affect the LLMBoard score?

No. Arena results remain a separate preference or task-outcome signal.

How many models are included?

40 model results are currently shown.