llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

Arena benchmark

LM Arena Agent Tool Hallucination Leaderboard

LM Arena Agent Tool Hallucination model ratings and results.

Updated Aug 17, 2026. Data as of 2026-08-17

Models40
Categories1
Result typeagent_score

On this page

  • Ranking
  • Highlights
  • Distribution
  • Top models
  • About
  • FAQ

Agent Tool Hallucination Ranking

Models are ordered by their Arena rank.

30 of 40 rows
Columns

Show columns

Sort by
Rank
Model
Organization
Rating / score
Votes
Observations
Result date
Rank01ModelOPGPT-5.6 TerraOpenAIOrganizationopenaiRating / score0.0VotesN/AObservations693.8KResult dateAug 13, 2026
Rank02ModelMAKimi K3Moonshot AIOrganizationmoonshotRating / score0.0VotesN/AObservations2.3MResult dateAug 13, 2026
Rank03ModelOPGPT-5.5OpenAIOrganizationopenaiRating / score0.0VotesN/AObservations1.5MResult dateAug 13, 2026
Rank04ModelOPGPT-5.4OpenAIOrganizationopenaiRating / score0.0VotesN/AObservations2.9MResult dateAug 13, 2026
Rank05ModelXAGrok 4.5xAIOrganizationxaiRating / score0.0VotesN/AObservations1.3MResult dateAug 13, 2026
Rank06ModelZAGLM-5.2Zhipu AIOrganizationzaiRating / score0.0VotesN/AObservations2.2MResult dateAug 13, 2026
Rank07ModelOPGPT-5.6 SolOpenAIOrganizationopenaiRating / score0.0VotesN/AObservations1.5MResult dateAug 13, 2026
Rank09ModelOPGPT-5.6 LunaOpenAIOrganizationopenaiRating / score0.0VotesN/AObservations496.9KResult dateAug 13, 2026
Rank10ModelMAKimi K2.6Moonshot AIOrganizationmoonshotRating / score0.0VotesN/AObservations297.2KResult dateAug 13, 2026
Rank11ModelMAKimi K2.7 CodeMoonshot AIOrganizationmoonshotRating / score0.0VotesN/AObservations375.3KResult dateAug 13, 2026
Rank12ModelANClaude Fable 5AnthropicOrganizationanthropicRating / score0.0VotesN/AObservations853.8KResult dateAug 13, 2026
Rank13ModelANClaude Opus 4.6AnthropicOrganizationanthropicRating / score0.0VotesN/AObservations1.2MResult dateAug 13, 2026
Rank14ModelDEDeepSeek-V4-Flash-0731DeepSeekOrganizationdeepseekRating / score0.0VotesN/AObservations3MResult dateAug 13, 2026
Rank16ModelMEMuse Spark 1.1MetaOrganizationmetaRating / score0.0VotesN/AObservations2.3MResult dateAug 13, 2026
Rank18ModelANClaude Opus 5AnthropicOrganizationanthropicRating / score0.0VotesN/AObservations1.7MResult dateAug 13, 2026
Rank20ModelANClaude Opus 4.7AnthropicOrganizationanthropicRating / score0.0VotesN/AObservations1.3MResult dateAug 13, 2026
Rank21ModelXAGrok 4.3xAIOrganizationxaiRating / score0.0VotesN/AObservations1.1MResult dateAug 13, 2026
Rank24ModelANClaude Sonnet 4.6AnthropicOrganizationanthropicRating / score0.0VotesN/AObservations1.2MResult dateAug 13, 2026
Rank26ModelANClaude Sonnet 5AnthropicOrganizationanthropicRating / score0.0VotesN/AObservations2.3MResult dateAug 13, 2026
Rank27ModelMIMiniMax M2.7MiniMaxOrganizationminimaxRating / score0.0VotesN/AObservations608.1KResult dateAug 13, 2026
Rank28ModelXAGrok Build 0.1xAIOrganizationxaiRating / score0.0VotesN/AObservations5.2MResult dateAug 13, 2026
Rank29ModelGOGemini 3.1 ProGoogleOrganizationgoogleRating / score0.0VotesN/AObservations2.2MResult dateAug 13, 2026
Rank30ModelMIMiniMax M3MiniMaxOrganizationminimaxRating / score0.0VotesN/AObservations1.9MResult dateAug 13, 2026
Rank31ModelACQwen3.7 MaxAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score0.0VotesN/AObservations1.2MResult dateAug 13, 2026
Rank32ModelTHInklingThinkingmachinesOrganizationthinkyRating / score0.0VotesN/AObservations1.2MResult dateAug 13, 2026
Rank33ModelNVNemotron 3 Ultra (550B A55B)NVIDIAOrganizationnvidiaRating / score0.0VotesN/AObservations142.5KResult dateAug 13, 2026
Rank34ModelUPSolar Pro 4UpstageOrganizationupstageRating / score0.0VotesN/AObservations151KResult dateAug 13, 2026
Rank35ModelGOGemini 3.5 FlashGoogleOrganizationgoogleRating / score0.0VotesN/AObservations460.9KResult dateAug 13, 2026
Rank36ModelDEDeepSeek V4 ProDeepSeekOrganizationdeepseekRating / score0.0VotesN/AObservations1.3MResult dateAug 13, 2026
Rank38ModelACQwen3.7-PlusAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score0.0VotesN/AObservations645.6KResult dateAug 13, 2026
Rank39ModelACQwen3.8 MaxAlibaba Cloud / Qwen TeamOrganizationalibabaRating / score0.0VotesN/AObservations533.2KResult dateAug 13, 2026
Rank40ModelXIMiMo-V2.5-ProXiaomiOrganizationxiaomiRating / score0.0VotesN/AObservations1.2MResult dateAug 13, 2026
Rank41ModelGOGemini 3.5 Flash-LiteGoogleOrganizationgoogleRating / score-0.0VotesN/AObservations481.3KResult dateAug 13, 2026
Rank42ModelZAGLM-5.1Zhipu AIOrganizationzaiRating / score-0.0VotesN/AObservations3MResult dateAug 13, 2026
Rank43ModelGOGemini 3 FlashGoogleOrganizationgoogleRating / score-0.0VotesN/AObservations1.6MResult dateAug 13, 2026
Rank45ModelTEHy3TencentOrganizationtencentRating / score-0.0VotesN/AObservations583.7KResult dateAug 13, 2026
Rank46ModelDEDeepSeek V4 FlashDeepSeekOrganizationdeepseekRating / score-0.0VotesN/AObservations1.1MResult dateAug 13, 2026
Rank47ModelMAMistral Medium 3.5Mistral AIOrganizationmistralRating / score-0.0VotesN/AObservations153.8KResult dateAug 13, 2026
Rank48ModelANClaude Opus 4.8AnthropicOrganizationanthropicRating / score-0.3VotesN/AObservations1.3MResult dateAug 13, 2026
Rank49ModelGOGemma 4 31BGoogleOrganizationgoogleRating / score-0.3VotesN/AObservations592.7KResult dateAug 13, 2026

Arena highlights

Rank #1GPT-5.6 Terra0.0 agent scoreRank #2Kimi K30.0 agent scoreRank #3GPT-5.50.0 agent scoreRank #4GPT-5.40.0 agent score

Agent Tool Hallucination Rating Distribution

A closer view of the leading model ratings.

Agent Tool Hallucination

The Top AI Models for Agent Tool Hallucination

The first five models in this source ranking, with price and output speed included when a canonical model match is available.

Ranking basisThis agent tool hallucination AI model leaderboard uses the source agent score and Arena rank. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    OP
    GPT-5.6 TerraOpenAI
    agent score
    0.01
    Price
    $2.0 input / $12 output per 1M tokens
    Speed
    Up to 35 tok/s via OpenAI

    Strengths

    • Ranks #1 on Agent Tool Hallucination
    • 693,841 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  2. 02
    MA
    Kimi K3Moonshot AI
    agent score
    0.01
    Price
    $3.0 input / $15 output per 1M tokens
    Speed
    Up to 26 tok/s via Fireworks

    Strengths

    • Ranks #2 on Agent Tool Hallucination
    • 2,305,107 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  3. 03
    OP
    GPT-5.5OpenAI
    agent score
    0.01
    Price
    $5.0 input / $30 output per 1M tokens

    Strengths

    • Ranks #3 on Agent Tool Hallucination
    • 1,544,032 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  4. 04
    OP
    GPT-5.4OpenAI
    agent score
    0.01
    Price
    $2.5 input / $15 output per 1M tokens
    Speed
    Up to 9.1 tok/s via OpenAI

    Strengths

    • Ranks #4 on Agent Tool Hallucination
    • 2,855,616 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score
  5. 05
    XA
    Grok 4.5xAI
    agent score
    0.01
    Price
    $2.0 input / $6.0 output per 1M tokens
    Speed
    Up to 4.3 tok/s via xAI

    Strengths

    • Ranks #5 on Agent Tool Hallucination
    • 1,326,811 observations support the result
    • Matched official PAYG token pricing is available

    Considerations

    • Arena results are independent from the LLMBoard capability score

Selection summary

Best AI Models for Agent Tool Hallucination

GPT-5.6 Terra currently leads Agent Tool Hallucination at 0.01. The best AI model for this use case may change when price, speed and independent benchmark capability are considered.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

Arena rank #1GPT-5.6 Terra$2.0 input / $12 output per 1M tokensArena rank #2Kimi K3$3.0 input / $15 output per 1M tokensArena rank #3GPT-5.5$5.0 input / $30 output per 1M tokens

What is the Agent Tool Hallucination Arena?

What Agent Tool Hallucination measures and how to read its results.

Agent Tool Hallucination reports agent score values and currently defines 1 evaluation category.

Arena results are preference or task-outcome signals. They are not automatically equivalent to capability benchmark scores.

FAQ

Common questions about Agent Tool Hallucination.

Which model leads Agent Tool Hallucination?

GPT-5.6 Terra is currently ranked first.

What does Agent Tool Hallucination measure?

This Arena reports agent score values across 1 configured evaluation category.

Does this result affect the LLMBoard score?

No. Arena results remain a separate preference or task-outcome signal.

How many models are included?

40 model results are currently shown.