llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

reasoning benchmark

SWE-bench Multilingual Leaderboard

A multilingual benchmark for issue resolving in software engineering that covers Java, TypeScript, JavaScript, Go, Rust, C, and C++. Contains 1,632 high-quality instances carefully annotated from 2,456 candidates by 68 expert annotators, designed to evaluate Large Language Models across diverse software ecosystems beyond Python.

Updated Aug 17, 2026

Models38
Model coverage38
MetricScore
EvidenceB

On this page

  • Ranking
  • Highlights
  • Distribution
  • Top models
  • About
  • FAQ

SWE-bench Multilingual Ranking

Higher score ranks better on this benchmark.

30 of 38 rows
Columns

Show columns

Sort by
Rank
Model
Score
Percentile
Participants
Evidence
Evaluated
Rank01ModelANClaude Mythos PreviewAnthropicScore87.3%Percentile100.0%Participants38EvidenceCEvaluatedAug 17, 2026
Rank02ModelANClaude Opus 4.8AnthropicScore84.4%Percentile97.3%Participants38EvidenceCEvaluatedAug 17, 2026
Rank03ModelPOLaguna S 2.1PoolsideScore78.5%Percentile94.6%Participants38EvidenceCEvaluatedAug 17, 2026
Rank04ModelANClaude Sonnet 5AnthropicScore78.3%Percentile91.9%Participants38EvidenceCEvaluatedAug 17, 2026
Rank05ModelACQwen3.7 MaxAlibaba Cloud / Qwen TeamScore78.3%Percentile89.2%Participants38EvidenceCEvaluatedAug 17, 2026
Rank06ModelANClaude Opus 4.6AnthropicScore77.8%Percentile86.5%Participants38EvidenceCEvaluatedAug 17, 2026
Rank07ModelMAKimi K2.6Moonshot AIScore76.7%Percentile83.8%Participants38EvidenceCEvaluatedAug 17, 2026
Rank08ModelMIMiniMax M2.7MiniMaxScore76.5%Percentile81.1%Participants38EvidenceCEvaluatedAug 17, 2026
Rank09ModelDEDeepSeek-V4-Pro-MaxDeepSeekScore76.2%Percentile78.4%Participants38EvidenceCEvaluatedAug 17, 2026
Rank10ModelTEHy3TencentScore75.8%Percentile75.7%Participants38EvidenceCEvaluatedAug 17, 2026
Rank11ModelACQwen3.7-PlusAlibaba Cloud / Qwen TeamScore75.8%Percentile73.0%Participants38EvidenceCEvaluatedAug 17, 2026
Rank12ModelACQwen3.6 PlusAlibaba Cloud / Qwen TeamScore73.8%Percentile70.3%Participants38EvidenceCEvaluatedAug 17, 2026
Rank13ModelDEDeepSeek-V4-Flash-MaxDeepSeekScore73.3%Percentile67.6%Participants38EvidenceCEvaluatedAug 17, 2026
Rank14ModelMAKimi K2.5Moonshot AIScore73.0%Percentile64.9%Participants38EvidenceCEvaluatedAug 17, 2026
Rank15ModelMIMiniMax M2.1MiniMaxScore72.5%Percentile62.2%Participants38EvidenceCEvaluatedAug 17, 2026
Rank16ModelXIMiMo-V2-FlashXiaomiScore71.7%Percentile59.5%Participants38EvidenceCEvaluatedAug 17, 2026
Rank17ModelXIMiMo-V2-ProXiaomiScore71.7%Percentile56.8%Participants38EvidenceCEvaluatedAug 17, 2026
Rank18ModelACQwen3.6-27BAlibaba Cloud / Qwen TeamScore71.3%Percentile54.0%Participants38EvidenceCEvaluatedAug 17, 2026
Rank19ModelDEDeepSeek-V3.2 (Thinking)DeepSeekScore70.2%Percentile51.4%Participants38EvidenceCEvaluatedAug 17, 2026
Rank20ModelDEDeepSeek-V3.2DeepSeekScore70.2%Percentile48.6%Participants38EvidenceCEvaluatedAug 17, 2026
Rank21ModelDEDeepSeek-V4-Flash-0423DeepSeekScore70.2%Percentile46.0%Participants38EvidenceCEvaluatedAug 17, 2026
Rank22ModelACQwen3.5-397B-A17BAlibaba Cloud / Qwen TeamScore69.3%Percentile43.2%Participants38EvidenceCEvaluatedAug 17, 2026
Rank23ModelNVNemotron 3 Ultra (550B A55B)NVIDIAScore67.7%Percentile40.5%Participants38EvidenceCEvaluatedAug 17, 2026
Rank24ModelACQwen3.6-35B-A3BAlibaba Cloud / Qwen TeamScore67.2%Percentile37.8%Participants38EvidenceCEvaluatedAug 17, 2026
Rank25ModelZAGLM-4.7Zhipu AIScore66.7%Percentile35.1%Participants38EvidenceCEvaluatedAug 17, 2026
Rank26ModelMIMAI-Code-1-FlashMicrosoftScore65.5%Percentile32.4%Participants38EvidenceCEvaluatedAug 17, 2026
Rank27ModelPOLaguna XS 2.1PoolsideScore63.1%Percentile29.7%Participants38EvidenceCEvaluatedAug 17, 2026
Rank28ModelMAKimi K2-Thinking-0905Moonshot AIScore61.1%Percentile27.0%Participants38EvidenceCEvaluatedAug 17, 2026
Rank29ModelDEDeepSeek-V3.2-ExpDeepSeekScore57.9%Percentile24.3%Participants38EvidenceCEvaluatedAug 17, 2026
Rank30ModelMIMiniMax M2MiniMaxScore56.5%Percentile21.6%Participants38EvidenceCEvaluatedAug 17, 2026
Rank31ModelACQwen3-Coder 480B A35B InstructAlibaba Cloud / Qwen TeamScore54.7%Percentile18.9%Participants38EvidenceCEvaluatedAug 17, 2026
Rank32ModelDEDeepSeek-V3.1DeepSeekScore54.5%Percentile16.2%Participants38EvidenceCEvaluatedAug 17, 2026
Rank33ModelMAKimi K2 InstructMoonshot AIScore47.3%Percentile13.5%Participants38EvidenceCEvaluatedAug 17, 2026
Rank34ModelMAKimi K2-Instruct-0905Moonshot AIScore47.3%Percentile10.8%Participants38EvidenceCEvaluatedAug 17, 2026
Rank35ModelNVNemotron 3 Super (120B A12B)NVIDIAScore45.8%Percentile8.1%Participants38EvidenceCEvaluatedAug 17, 2026
Rank36ModelNVNemotron 3.5 Lightning (30B A3B)NVIDIAScore39.3%Percentile5.4%Participants38EvidenceCEvaluatedAug 17, 2026
Rank37ModelMELongCat-Flash-LiteMeituanScore38.1%Percentile2.7%Participants38EvidenceCEvaluatedAug 17, 2026
Rank38ModelDEDeepSeek-R1-0528DeepSeekScore30.5%Percentile0.0%Participants38EvidenceCEvaluatedAug 17, 2026

SWE-bench Multilingual Highlights

The leading models and scores on this benchmark.

Rank #1Claude Mythos Preview87.3%Rank #2Claude Opus 4.884.4%Rank #3Laguna S 2.178.5%Rank #4Claude Sonnet 578.3%

SWE-bench Multilingual Score Distribution

A closer view of the leading scores on this benchmark.

SWE-bench Multilingual

The Top AI Models for SWE-bench Multilingual

The first five results on this benchmark, with official price and output speed added where the model identity can be matched.

Ranking basisThis swe-bench multilingual AI model leaderboard uses descending score in the benchmark's original unit. The leaderboard ranking keeps matched price and speed data separate from benchmark evidence.

  1. 01
    AN
    Claude Mythos PreviewAnthropic
    Score
    87.3%

    Strengths

    • Ranks #1 of 38 compared models
    • 100th percentile on this benchmark
    • C evidence result

    Considerations

    • This result measures SWE-bench Multilingual, not total model capability
  2. 02
    AN
    Claude Opus 4.8Anthropic
    Score
    84.4%
    Price
    $5.0 input / $25 output per 1M tokens
    Speed
    Up to 46 tok/s via Anthropic

    Strengths

    • Ranks #2 of 38 compared models
    • 97th percentile on this benchmark
    • C evidence result

    Considerations

    • This result measures SWE-bench Multilingual, not total model capability
  3. 03
    PO
    Laguna S 2.1Poolside
    Score
    78.5%

    Strengths

    • Ranks #3 of 38 compared models
    • 95th percentile on this benchmark
    • C evidence result

    Considerations

    • This result measures SWE-bench Multilingual, not total model capability
  4. 04
    AN
    Claude Sonnet 5Anthropic
    Score
    78.3%
    Price
    $2.0 input / $10 output per 1M tokens
    Speed
    Up to 11 tok/s via Anthropic

    Strengths

    • Ranks #4 of 38 compared models
    • 92th percentile on this benchmark
    • C evidence result

    Considerations

    • This result measures SWE-bench Multilingual, not total model capability
  5. 05
    AC
    Qwen3.7 MaxAlibaba Cloud / Qwen Team
    Score
    78.3%
    Price
    $2.5 input / $7.5 output per 1M tokens
    Speed
    Up to 5.8 tok/s via Together

    Strengths

    • Ranks #5 of 38 compared models
    • 89th percentile on this benchmark
    • C evidence result

    Considerations

    • This result measures SWE-bench Multilingual, not total model capability

Selection summary

Best AI Models for SWE-bench Multilingual

Claude Mythos Preview currently leads SWE-bench Multilingual with 87.3%. It is the top model on this specific benchmark, while the best LLM for the broader task should also be checked against other benchmarks, price and runtime.

Use this leaderboard with the supporting benchmark results and coverage details above. A leaderboard position summarizes the selected ranking signal; it does not replace workload-specific testing.

Benchmark rank #1Claude Mythos Preview87.3%Benchmark rank #2Claude Opus 4.884.4% · $5.0 input / $25 output per 1M tokensBenchmark rank #3Laguna S 2.178.5%

What is SWE-bench Multilingual?

What SWE-bench Multilingual measures and how its scores work.

A multilingual benchmark for issue resolving in software engineering that covers Java, TypeScript, JavaScript, Go, Rust, C, and C++. Contains 1,632 high-quality instances carefully annotated from 2,456 candidates by 68 expert annotators, designed to evaluate Large Language Models across diverse software ecosystems beyond Python.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.

Family
SWE-bench Multilingual
Modality
text
Primary category
reasoning
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
swe-bench-multilingual|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about SWE-bench Multilingual.

Which model scores highest on SWE-bench Multilingual?

Claude Mythos Preview is currently ranked first with 87.3%.

What does SWE-bench Multilingual measure?

A multilingual benchmark for issue resolving in software engineering that covers Java, TypeScript, JavaScript, Go, Rust, C, and C++. Contains 1,632 high-quality instances carefully annotated from 2,456 candidates by 68 expert annotators, designed to evaluate Large Language Models across diverse software ecosystems beyond Python.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

38 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.