llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

multimodal benchmark

MedXpertQA

A comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning, featuring 4,460 questions spanning 17 specialties and 11 body systems. Includes both text-only and multimodal subsets with expert-level exam questions incorporating diverse medical images and rich clinical information.

Updated Aug 11, 2026

Models12
Model coverage12
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

MedXpertQA Ranking

Higher score ranks better on this benchmark.

12 rows
Columns

Show columns

01MEMuse SparkMeta78.4%100.0%12CAug 11, 2026
02ACQwen3.5-122B-A10BAlibaba Cloud / Qwen Team67.3%90.9%12CAug 11, 2026
03ACQwen3.5-27BAlibaba Cloud / Qwen Team62.4%81.8%12CAug 11, 2026
04ACQwen3.5-35B-A3BAlibaba Cloud / Qwen Team61.4%72.7%12CAug 11, 2026
05GOGemma 4 31BGoogle61.3%63.6%12CAug 11, 2026
06GOGemma 4 26B-A4BGoogle58.1%54.5%12CAug 11, 2026
07GODiffusionGemma 26B-A4BGoogle49.0%45.5%12CAug 11, 2026
08GOGemma 4 12BGoogle48.7%36.4%12CAug 11, 2026
09MIMAI-Thinking-1Microsoft43.0%27.3%12CAug 11, 2026
10GOGemma 4 E4BGoogle28.7%18.2%12CAug 11, 2026
11GOGemma 4 E2BGoogle23.5%9.1%12CAug 11, 2026
12GOMedGemma 4B ITGoogle18.8%0.0%12CAug 11, 2026

MedXpertQA Score Distribution

A closer view of the leading scores on this benchmark.

MedXpertQA

MedXpertQA Highlights

The leading models and scores on this benchmark.

Rank #1Muse Spark78.4%Rank #2Qwen3.5-122B-A10B67.3%Rank #3Qwen3.5-27B62.4%Rank #4Qwen3.5-35B-A3B61.4%

What is MedXpertQA?

What MedXpertQA measures and how its scores work.

A comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning, featuring 4,460 questions spanning 17 specialties and 11 body systems. Includes both text-only and multimodal subsets with expert-level exam questions incorporating diverse medical images and rich clinical information.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.

Family
MedXpertQA
Modality
multimodal
Primary category
multimodal
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
medxpertqa|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about MedXpertQA.

Which model scores highest on MedXpertQA?

Muse Spark is currently ranked first with 78.4%.

What does MedXpertQA measure?

A comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning, featuring 4,460 questions spanning 17 specialties and 11 body systems. Includes both text-only and multimodal subsets with expert-level exam questions incorporating diverse medical images and rich clinical information.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

12 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.