llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

Capability ranking

Best AI for Instruction Following

Use this instruction-following model leaderboard to compare the current ranking from the LLMBoard Instruction Score after applying benchmark evidence requirements.

Updated 2026-08-17

On this page

  • Ranking
  • Overview
  • Leaders
  • Capability
  • Value
  • Composition
  • Compare
  • Top models
  • FAQ

Instruction Model Leaderboard Ranking

This leaderboard ranking orders models by their LLMBoard Instruction Score aggregated from eligible benchmark evidence.

30 of 53 rows
Columns

Show columns

Sort by
Rank
Model
LLMBoard Instruction Score
Context
Official input / 1M
Official output / 1M
Output speed
Rank01ModelAMNova 2 ProAmazonLLMBoard Instruction Score93.9ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank02ModelACQwen3.7 MaxAlibaba Cloud / Qwen TeamLLMBoard Instruction Score91.5Context1MOfficial input / 1M$2.5Official output / 1M$7.5Output speed5.805 tok/s
Rank03ModelACQwen3.7 PlusAlibaba Cloud / Qwen TeamLLMBoard Instruction Score87.0Context1MOfficial input / 1M$0.50Official output / 1M$3Output speedN/A
Rank04ModelACQwen3 235B A22B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score85.2Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank05ModelNVNemotron 3 UltraNVIDIALLMBoard Instruction Score84.5Context1MOfficial input / 1M$0.50Official output / 1M$2.5Output speedN/A
Rank06ModelNRHermes 3 70BNous ResearchLLMBoard Instruction Score81.8Context131.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank07ModelACQwen3.5 397B A17BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score81.6Context262.1KOfficial input / 1M$0.60Official output / 1M$3.6Output speedN/A
Rank08ModelACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score77.3Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank09ModelACQwen3 VL 235B A22BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score77.1Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank10ModelACQwen3 235B A22BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score72.9Context262.1KOfficial input / 1M$0.70Official output / 1M$2.8Output speed68.17 tok/s
Rank11ModelACQwen3.6 PlusAlibaba Cloud / Qwen TeamLLMBoard Instruction Score72.7Context1MOfficial input / 1M$0.50Official output / 1M$3Output speed15.66 tok/s
Rank12ModelNVNemotron 3 SuperNVIDIALLMBoard Instruction Score71.8Context262.1KOfficial input / 1M$0.20Official output / 1M$0.80Output speedN/A
Rank13ModelAMNova 2 LiteAmazonLLMBoard Instruction Score70.9Context1MOfficial input / 1M$0.33Official output / 1M$2.75Output speedN/A
Rank14ModelACQwen3.5 27BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score70.3Context262.1KOfficial input / 1M$0.30Official output / 1M$2.4Output speed6.736 tok/s
Rank15ModelACQwen3.5 122B A10BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score69.6Context262.1KOfficial input / 1M$0.40Official output / 1M$3.2Output speedN/A
Rank16ModelACQwen2.5 72BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score69.3Context131.1KOfficial input / 1M$1.4Official output / 1M$5.6Output speed100 tok/s
Rank17ModelOPo3 miniOpenAILLMBoard Instruction Score68.8Context200KOfficial input / 1M$1.1Official output / 1M$4.4Output speed115 tok/s
Rank18ModelCOCommand A+CohereLLMBoard Instruction Score65.2Context128KOfficial input / 1M$2.5Official output / 1M$10Output speedN/A
Rank19ModelMAKimi K2Moonshot AILLMBoard Instruction Score63.8ContextN/AOfficial input / 1M$0.60Official output / 1M$2.5Output speed45 tok/s
Rank20ModelACQwen3 Next 80B A3BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score63.2Context131.1KOfficial input / 1M$0.50Official output / 1M$2Output speedN/A
Rank21ModelACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score63.1ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank22ModelACQwen3 Next 80B A3B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score62.5Context131.1KOfficial input / 1M$0.50Official output / 1M$6Output speedN/A
Rank23ModelGOGemma 3 27BGoogleLLMBoard Instruction Score61.4Context131.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed33 tok/s
Rank24ModelAMNova 2 OmniAmazonLLMBoard Instruction Score60.1ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank25ModelNVNemotron 3 NanoNVIDIALLMBoard Instruction Score56.7Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank26ModelLALFM2.5 2.6BLiquid AILLMBoard Instruction Score55.3ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank27ModelACQwen3.5 35B A3BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score54.4Context262.1KOfficial input / 1M$0.25Official output / 1M$2Output speedN/A
Rank28ModelGOGemma 3 12BGoogleLLMBoard Instruction Score51.9Context131.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed33 tok/s
Rank29ModelGOGemma 3 4BGoogleLLMBoard Instruction Score50.8Context131.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed33 tok/s
Rank30ModelACQwen3.5 9BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score45.3Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank31ModelACQwen3 VL 32BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score42.3ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank32ModelACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score42.1Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank33ModelOPGPT-4.5OpenAILLMBoard Instruction Score41.0Context128KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed50 tok/s
Rank34ModelACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score40.7Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank35ModelACQwen3 VL 30B A3BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score39.5Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank36ModelMIMAI Thinking 1MicrosoftLLMBoard Instruction Score39.1ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank37ModelACQwen3.5 4BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score37.0ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank38ModelACQwen3 VL 8BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score36.9Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank39ModelACQwen2.5 7BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score34.9Context131.1KOfficial input / 1M$0.175Official output / 1M$0.70Output speed138 tok/s
Rank40ModelMAMistral Small 3 24BMistral AILLMBoard Instruction Score33.3Context32KOfficial input / 1MN/AOfficial output / 1MN/AOutput speed134 tok/s
Rank41ModelOPGPT-4.1OpenAILLMBoard Instruction Score31.8Context1MOfficial input / 1M$2Official output / 1M$8Output speed100 tok/s
Rank42ModelACQwen3 VL 4B ThinkingAlibaba Cloud / Qwen TeamLLMBoard Instruction Score30.2Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank43ModelOPGPT-4.1-miniOpenAILLMBoard Instruction Score24.0Context1MOfficial input / 1M$0.40Official output / 1M$1.6Output speed68.886 tok/s
Rank44ModelOPGPT-4oOpenAILLMBoard Instruction Score20.5Context128KOfficial input / 1M$2.5Official output / 1M$10Output speed132 tok/s
Rank45ModelNVLlama 3.1 Nemotron Nano 8BNVIDIALLMBoard Instruction Score18.2ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank46ModelACQwen3 VL 4BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score17.3Context262.1KOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank47ModelLALFM2.5 VL 3BLiquid AILLMBoard Instruction Score12.9ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank48ModelGOGemma 3 1BGoogleLLMBoard Instruction Score12.1ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank49ModelACQwen3.5 2BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score11.8ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank50ModelCONorth Micro VisionCohereLLMBoard Instruction Score6.1ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A
Rank51ModelMAPixtral 12BMistral AILLMBoard Instruction Score5.3Context128KOfficial input / 1M$0.15Official output / 1M$0.15Output speed0.1 tok/s
Rank52ModelOPGPT-4.1-nanoOpenAILLMBoard Instruction Score4.0Context1MOfficial input / 1M$0.10Official output / 1M$0.40Output speed200 tok/s
Rank53ModelACQwen3.5 0.8BAlibaba Cloud / Qwen TeamLLMBoard Instruction Score0.6ContextN/AOfficial input / 1MN/AOfficial output / 1MN/AOutput speedN/A

Official PAYG prices appear here. Third-party offers remain on the pricing and model detail pages.

What this leaderboard says

Key findings from the current instruction leaderboard ranking and its supporting benchmark coverage.

Ranked models53
Vendors11
Open-weight models13
Official prices23
Top instruction modelNova 2 Pro93.9 instruction scoreBest open-weight modelNemotron 3 UltraRank #5Lowest official inputGPT-4.1-nano$0.10 / 1MLargest contextGPT-4.11M tokens

Nova 2 Pro leads this page at 93.9, 2.5 points ahead of Qwen3.7 Max.

Nemotron 3 Ultra is the highest-ranked open-weight option at #5. GPT-4.1-nano has the lowest official input price among ranked models.

Instruction Leaderboard Leaders

The leading instruction scores in this benchmark-backed ranking, with the axis focused on the competitive leaderboard range.

Top instruction models

Instruction Benchmark Profile

Compare instruction benchmark strength with each model's broader capability score and leaderboard position.

ModelReasoningKnowledgeMathCodingInstruction
Nova 2 ProAmazon65.866.268.966.693.9
Qwen3.7 MaxAlibaba Cloud / Qwen Team91.688.993.484.791.5
Qwen3.7 PlusAlibaba Cloud / Qwen Team87.983.169.360.987.0
Qwen3 235B A22B ThinkingAlibaba Cloud / Qwen Team70.961.063.652.785.2
Nemotron 3 UltraNVIDIA81.376.7100.048.684.5
-8.719.347.375.3103.3-6.316.439.161.884.6LLMBoard Instruction ScoreLLMBoard Overall Score
53 models with comparable dataOpen a model by selecting its logo
GPQA92.4%Qwen3.7 Max43 compared models
MMLU-Pro89.6%Qwen3.7 Max43 compared models
AIME 202599.2%Nemotron 3 Nano (30B A3B)26 compared models
SWE-Bench Verified80.4%Qwen3.7 Max19 compared models
IFEval95.0%Qwen3.5-27B43 compared models
IFBench81.7%Nemotron 3 Ultra (550B A55B)21 compared models
WMT24++86.7%Nemotron 3 Super (120B A12B)19 compared models
MultiPL-E87.9%Qwen3-235B-A22B-Instruct-25076 compared models

Capability at each price point

Official vendor API prices plotted against the instruction metric used on this page. Price remains a separate decision signal.

020406080100$0.2$0.5$1.0$2.0LLMBoard Instruction ScoreOfficial API price blend (8:1 input/output), USD / 1M
Efficient frontier23 models with official PAYG prices

Instruction Leaderboard Composition

Vendor concentration and model access among the first 25 products in the ranking.

Top-model vendor mix

25 models
Alibaba Cloud / Qwen Team14Amazon3NVIDIA3Nous Research1OpenAI1Cohere1Moonshot AI1Google1

Access model

Open weights
10
Closed weights
15
Official price available
16

Compare the leaders

Open focused comparisons between the current leader and the nearest practical alternatives.

Nova 2 ProvsQwen3.7 MaxNova 2 ProvsNemotron 3 UltraNova 2 ProvsGPT-4.1-nano
Overall leaderboardCoding leaderboardReasoning leaderboardMath leaderboardKnowledge leaderboard

The Top AI Models for Instruction Following

A data-backed look at the first five models in this instruction following leaderboard ranking, including benchmark context, price and output speed where available.

Ranking basisThis instruction following AI model leaderboard uses the instruction capability score shown on this page. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    AM
    Nova 2 ProAmazon
    LLMBoard Instruction Score
    93.9

    Strengths

    • Ranks #1 with 93.9 LLMBoard Instruction Score

    Considerations

    • Weights are not marked as open
  2. 02
    AC
    Qwen3.7 MaxAlibaba Cloud / Qwen Team
    LLMBoard Instruction Score
    91.5
    Price
    $2.5 input / $7.5 output per 1M tokens
    Speed
    Up to 5.8 tok/s via Together

    Strengths

    • Ranks #2 with 91.5 LLMBoard Instruction Score
    • 1M-token context window

    Considerations

    • Weights are not marked as open
  3. 03
    AC
    Qwen3.7 PlusAlibaba Cloud / Qwen Team
    LLMBoard Instruction Score
    87.0
    Price
    $0.50 input / $3.0 output per 1M tokens

    Strengths

    • Ranks #3 with 87.0 LLMBoard Instruction Score
    • Lowest official input price among the current top five
    • 1M-token context window

    Considerations

    • Weights are not marked as open
  4. 04
    AC
    Qwen3 235B A22B ThinkingAlibaba Cloud / Qwen Team
    LLMBoard Instruction Score
    85.2

    Strengths

    • Ranks #4 with 85.2 LLMBoard Instruction Score
    • 262.1K-token context window

    Considerations

    • Weights are not marked as open
  5. 05
    NV
    Nemotron 3 UltraNVIDIA
    LLMBoard Instruction Score
    84.5
    Price
    $0.50 input / $2.5 output per 1M tokens

    Strengths

    • Ranks #5 with 84.5 LLMBoard Instruction Score
    • The scored version is marked as open weight
    • Lowest official input price among the current top five

Selection summary

Best AI Models for Instruction Following

Nova 2 Pro is currently the best-ranked LLM for instruction following with a 93.9 instruction score. The score is a relative ranking signal, so price, speed, context and evidence coverage should still be checked separately.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

Capability rank #1Nova 2 Pro93.9Capability rank #2Qwen3.7 MaxUp to 5.8 tok/s via TogetherCapability rank #3Qwen3.7 Plus$0.50 input / $3.0 output per 1M tokens

About this leaderboard

Common questions about the Best AI for Instruction Following leaderboard, benchmark evidence and ranking method.

What is the best model on Best AI for Instruction Following?

Nova 2 Pro is currently ranked first with a instruction score of 93.9.

How is this leaderboard ranked?

Each model appears once using its current scored version. The leaderboard ranking follows the capability named in the title, while the overall page uses the LLMBoard score aggregated from eligible benchmark evidence.

How many models are included?

This page currently ranks 53 unique model products.

Where does pricing come from?

Main price columns use the model vendor's official standard PAYG API rate. Eligible third-party offers appear only in separately labeled columns, and unavailable official prices display as N/A.

Does Arena affect the score?

No. Arena results are displayed as an independent signal and are not included in the current LLMBoard capability score.