llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

Arena benchmark

LM Arena Agent Bash Recovery Steps Leaderboard

LM Arena Agent Bash Recovery Steps model ratings and results.

Updated Aug 17, 2026. Data as of 2026-08-17

Models40
Categories1
Result typeagent_score

On this page

  • Ranking
  • Highlights
  • Distribution
  • Top models
  • About
  • FAQ

Agent Bash Recovery Steps Ranking

Models are ordered by their Arena rank.

30 of 40 rows
Columns

Show columns

Sort by
Rank
Model
Organization
Rating / score
Votes
Observations
Result date
Rank01ModelOPGPT-5.5OpenAIOrganizationopenaiRating / score0.1VotesN/AObservations36.8KResult dateAug 13, 2026
Rank02ModelANClaude Opus 5AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations38.2KResult dateAug 13, 2026
Rank03ModelANClaude Fable 5AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations18.5KResult dateAug 13, 2026
Rank08ModelOPGPT-5.6 LunaOpenAIOrganizationopenaiRating / score0.1VotesN/AObservations12.5KResult dateAug 13, 2026
Rank09ModelANClaude Opus 4.6AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations24.7KResult dateAug 13, 2026
Rank10ModelXAGrok 4.5xAIOrganizationxaiRating / score0.1VotesN/AObservations25.9KResult dateAug 13, 2026
Rank11ModelANClaude Sonnet 5AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations92.7KResult dateAug 13, 2026
Rank12ModelANClaude Sonnet 4.6AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations24.6KResult dateAug 13, 2026
Rank13ModelANClaude Opus 4.8AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations24.2KResult dateAug 13, 2026
Rank14ModelOPGPT-5.6 SolOpenAIOrganizationopenaiRating / score0.1VotesN/AObservations32KResult dateAug 13, 2026
Rank15ModelANClaude Opus 4.7AnthropicOrganizationanthropicRating / score0.1VotesN/AObservations28.2KResult dateAug 13, 2026
Rank16ModelOPGPT-5.6 TerraOpenAIOrganizationopenaiRating / score0.1VotesN/AObservations20.4KResult dateAug 13, 2026
Rank17ModelOPGPT-5.4OpenAIOrganizationopenaiRating / score0.1VotesN/AObservations58.7KResult dateAug 13, 2026
Rank19ModelACQwen3.8 MaxAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score0.1VotesN/AObservations14.5KResult dateAug 13, 2026
Rank20ModelMAKimi K3Moonshot AIOrganizationmoonshotRating / score0.1VotesN/AObservations66.2KResult dateAug 13, 2026
Rank21ModelTHInklingThinkingmachinesOrganizationthinkyRating / score0.1VotesN/AObservations36.5KResult dateAug 13, 2026
Rank22ModelACQwen3.7-PlusAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score0.1VotesN/AObservations19.7KResult dateAug 13, 2026
Rank23ModelZAGLM-5.2Zhipu AIOrganizationzaiRating / score0.1VotesN/AObservations64.3KResult dateAug 13, 2026
Rank24ModelMIMiniMax M3MiniMaxOrganizationminimaxRating / score0.1VotesN/AObservations57.2KResult dateAug 13, 2026
Rank25ModelMEMuse Spark 1.1MetaOrganizationmetaRating / score0.1VotesN/AObservations46.6KResult dateAug 13, 2026
Rank26ModelACQwen3.7 MaxAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score0.1VotesN/AObservations29.3KResult dateAug 13, 2026
Rank27ModelDEDeepSeek-V4-Flash-0731DeepSeekOrganizationdeepseekRating / score0.0VotesN/AObservations110.2KResult dateAug 13, 2026
Rank28ModelDEDeepSeek V4 ProDeepSeekOrganizationdeepseekRating / score0.0VotesN/AObservations48.7KResult dateAug 13, 2026
Rank29ModelTEHy3TencentOrganizationtencentRating / score0.0VotesN/AObservations16KResult dateAug 13, 2026
Rank30ModelDEDeepSeek V4 FlashDeepSeekOrganizationdeepseekRating / score0.0VotesN/AObservations48.6KResult dateAug 13, 2026
Rank32ModelXIMiMo-V2.5-ProXiaomiOrganizationxiaomiRating / score0.0VotesN/AObservations37.5KResult dateAug 13, 2026
Rank33ModelGOGemini 3.5 FlashGoogleOrganizationgoogleRating / score-0.0VotesN/AObservations20.9KResult dateAug 13, 2026
Rank34ModelZAGLM-5.1Zhipu AIOrganizationzaiRating / score-0.0VotesN/AObservations115.6KResult dateAug 13, 2026
Rank35ModelMAKimi K2.7 CodeMoonshot AIOrganizationmoonshotRating / score-0.0VotesN/AObservations13KResult dateAug 13, 2026
Rank37ModelMAMistral Medium 3.5Mistral AIOrganizationmistralRating / score-0.0VotesN/AObservations5.9KResult dateAug 13, 2026
Rank39ModelMAKimi K2.6Moonshot AIOrganizationmoonshotRating / score-0.1VotesN/AObservations9.8KResult dateAug 13, 2026
Rank40ModelGOGemini 3.1 ProGoogleOrganizationgoogleRating / score-0.1VotesN/AObservations186.4KResult dateAug 13, 2026
Rank41ModelXAGrok 4.3xAIOrganizationxaiRating / score-0.1VotesN/AObservations49KResult dateAug 13, 2026
Rank42ModelUPSolar Pro 4UpstageOrganizationupstageRating / score-0.1VotesN/AObservations8.6KResult dateAug 13, 2026
Rank43ModelGOGemini 3.5 Flash-LiteGoogleOrganizationgoogleRating / score-0.2VotesN/AObservations31KResult dateAug 13, 2026
Rank44ModelMIMiniMax M2.7MiniMaxOrganizationminimaxRating / score-0.2VotesN/AObservations23.3KResult dateAug 13, 2026
Rank45ModelGOGemini 3 FlashGoogleOrganizationgoogleRating / score-0.2VotesN/AObservations83KResult dateAug 13, 2026
Rank46ModelXAGrok Build 0.1xAIOrganizationxaiRating / score-0.2VotesN/AObservations169.1KResult dateAug 13, 2026
Rank47ModelNVNemotron 3 Ultra (550B A55B)NVIDIAOrganizationnvidiaRating / score-0.2VotesN/AObservations8.1KResult dateAug 13, 2026
Rank49ModelGOGemma 4 31BGoogleOrganizationgoogleRating / score-0.5VotesN/AObservations27.1KResult dateAug 13, 2026

Arena highlights

Rank #1GPT-5.50.1 agent scoreRank #2Claude Opus 50.1 agent scoreRank #3Claude Fable 50.1 agent scoreRank #8GPT-5.6 Luna0.1 agent score

Agent Bash Recovery Steps Rating Distribution

A closer view of the leading model ratings.

Agent Bash Recovery Steps

The Top AI Models for Agent Bash Recovery Steps

The first five models in this source ranking, with price and output speed included when a canonical model match is available.

Ranking basisThis agent bash recovery steps AI model leaderboard uses the source agent score and Arena rank. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    OP
    GPT-5.5OpenAI
    agent score
    0.15
    Price
    $5.0 input / $30 output per 1M tokens

    Strengths

    • Ranks #1 on Agent Bash Recovery Steps
    • 36,844 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  2. 02
    AN
    Claude Opus 5Anthropic
    agent score
    0.14
    Price
    $5.0 input / $25 output per 1M tokens
    Speed
    Up to 3.3 tok/s via Anthropic

    Strengths

    • Ranks #2 on Agent Bash Recovery Steps
    • 38,217 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  3. 03
    AN
    Claude Fable 5Anthropic
    agent score
    0.14
    Price
    $10 input / $50 output per 1M tokens
    Speed
    Up to 9.1 tok/s via Anthropic

    Strengths

    • Ranks #3 on Agent Bash Recovery Steps
    • 18,512 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  4. 08
    OP
    GPT-5.6 LunaOpenAI
    agent score
    0.12
    Price
    $0.20 input / $1.2 output per 1M tokens
    Speed
    Up to 49 tok/s via OpenAI

    Strengths

    • Ranks #8 on Agent Bash Recovery Steps
    • 12,503 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  5. 09
    AN
    Claude Opus 4.6Anthropic
    agent score
    0.11
    Price
    $5.0 input / $25 output per 1M tokens
    Speed
    Up to 17 tok/s via Anthropic

    Strengths

    • Ranks #9 on Agent Bash Recovery Steps
    • 24,667 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score

Selection summary

Best AI Models for Agent Bash Recovery Steps

GPT-5.5 currently leads Agent Bash Recovery Steps at 0.15. The best AI model for this use case may change when price, speed and independent benchmark capability are considered.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

Arena rank #1GPT-5.5$5.0 input / $30 output per 1M tokensArena rank #2Claude Opus 5$5.0 input / $25 output per 1M tokensArena rank #3Claude Fable 5$10 input / $50 output per 1M tokens

What is the Agent Bash Recovery Steps Arena?

What Agent Bash Recovery Steps measures and how to read its results.

Agent Bash Recovery Steps reports agent score values and currently defines 1 evaluation category.

Arena results are preference or task-outcome signals. They are not automatically equivalent to capability benchmark scores.

FAQ

Common questions about Agent Bash Recovery Steps.

Which model leads Agent Bash Recovery Steps?

GPT-5.5 is currently ranked first.

What does Agent Bash Recovery Steps measure?

This Arena reports agent score values across 1 configured evaluation category.

Does this result affect the LLMBoard score?

No. Arena results remain a separate preference or task-outcome signal.

How many models are included?

40 model results are currently shown.