llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

multimodal benchmark

CharXiv-R

CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.

Updated Aug 11, 2026

Models48
Model coverage48
MetricScore
EvidenceB

On this page

  • Ranking
  • Distribution
  • Highlights
  • About
  • FAQ

CharXiv-R Ranking

Higher score ranks better on this benchmark.

48 rows
Columns

Show columns

01ANClaude Mythos PreviewAnthropic93.2%100.0%48CAug 11, 2026
02MAKimi K3Moonshot AI91.3%97.9%48CAug 11, 2026
03ANClaude Opus 4.7Anthropic91.0%95.7%48CAug 11, 2026
04ANClaude Opus 4.8Anthropic89.9%93.6%48CAug 11, 2026
05GOGemini 3.6 FlashGoogle89.4%91.5%48CAug 11, 2026
06MEMuse Spark 1.1Meta88.4%89.4%48CAug 11, 2026
07ANClaude Sonnet 5Anthropic88.3%87.2%48CAug 11, 2026
08MAKimi K2.6Moonshot AI86.7%85.1%48CAug 11, 2026
09MEMuse SparkMeta86.4%83.0%48CAug 11, 2026
10BYSeed 2.1 ProByteDance86.4%80.8%48CAug 11, 2026
11ACQwen3.7-PlusAlibaba Cloud / Qwen Team85.9%78.7%48CAug 11, 2026
12GOGemini 3.5 FlashGoogle84.2%76.6%48CAug 11, 2026
13BYSeed 2.1 TurboByteDance83.6%74.5%48CAug 11, 2026
14OPGPT-5.2OpenAI82.1%72.3%48CAug 11, 2026
15OPGPT-5.5 InstantOpenAI81.6%70.2%48CAug 11, 2026
16ACQwen3.6 PlusAlibaba Cloud / Qwen Team81.5%68.1%48CAug 11, 2026
17GOGemini 3 ProGoogle81.4%66.0%48CAug 11, 2026
18OPGPT-5OpenAI81.1%63.8%48CAug 11, 2026
19XIMiMo-V2.5Xiaomi81.0%61.7%48CAug 11, 2026
20GOGemini 3 FlashGoogle80.3%59.6%48CAug 11, 2026
21ACQwen3.5-27BAlibaba Cloud / Qwen Team79.5%57.5%48CAug 11, 2026
22MEMuse Glimmer-30BMeta78.8%55.3%48CAug 11, 2026
23OPo3OpenAI78.6%53.2%48CAug 11, 2026
24ACQwen3.6-27BAlibaba Cloud / Qwen Team78.4%51.1%48CAug 11, 2026
25ACQwen3.6-35B-A3BAlibaba Cloud / Qwen Team78.0%48.9%48CAug 11, 2026
26MAKimi K2.5Moonshot AI77.5%46.8%48CAug 11, 2026
27ACQwen3.5-35B-A3BAlibaba Cloud / Qwen Team77.5%44.7%48CAug 11, 2026
28ANClaude Opus 4.6Anthropic77.4%42.5%48CAug 11, 2026
29ACQwen3.5-122B-A10BAlibaba Cloud / Qwen Team77.2%40.4%48CAug 11, 2026
30GOGemini 3.5 Flash-LiteGoogle76.5%38.3%48CAug 11, 2026
31GOGemini 3.1 Flash-LiteGoogle73.2%36.2%48CAug 11, 2026
32OPo4-miniOpenAI72.0%34.0%48CAug 11, 2026
33ACQwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team66.1%31.9%48CAug 11, 2026
34ACQwen3 VL 32B ThinkingAlibaba Cloud / Qwen Team65.2%29.8%48CAug 11, 2026
35ACQwen3 VL 32B InstructAlibaba Cloud / Qwen Team62.8%27.7%48CAug 11, 2026
36ACQwen3 VL 235B A22B InstructAlibaba Cloud / Qwen Team62.1%25.5%48CAug 11, 2026
37OPGPT-4oOpenAI58.8%23.4%48CAug 11, 2026
38OPGPT-4.1 miniOpenAI56.8%21.3%48CAug 11, 2026
39OPGPT-4.1OpenAI56.7%19.1%48CAug 11, 2026
40ACQwen3 VL 30B A3B ThinkingAlibaba Cloud / Qwen Team56.6%17.0%48CAug 11, 2026
41OPGPT-4.5OpenAI55.4%14.9%48CAug 11, 2026
42ACQwen3 VL 8B ThinkingAlibaba Cloud / Qwen Team53.0%12.8%48CAug 11, 2026
43COCommand A+Cohere52.7%10.6%48CAug 11, 2026
44ACQwen3 VL 4B ThinkingAlibaba Cloud / Qwen Team50.3%8.5%48CAug 11, 2026
45ACQwen3 VL 30B A3B InstructAlibaba Cloud / Qwen Team48.9%6.4%48CAug 11, 2026
46ACQwen3 VL 8B InstructAlibaba Cloud / Qwen Team46.4%4.3%48CAug 11, 2026
47OPGPT-4.1 nanoOpenAI40.5%2.1%48CAug 11, 2026
48ACQwen3 VL 4B InstructAlibaba Cloud / Qwen Team39.7%0.0%48CAug 11, 2026

CharXiv-R Score Distribution

A closer view of the leading scores on this benchmark.

CharXiv-R

CharXiv-R Highlights

The leading models and scores on this benchmark.

Rank #1Claude Mythos Preview93.2%Rank #2Kimi K391.3%Rank #3Claude Opus 4.791.0%Rank #4Claude Opus 4.889.9%

What is CharXiv-R?

What CharXiv-R measures and how its scores work.

CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.

Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.

Family
CharXiv-R
Modality
multimodal
Primary category
multimodal
Score direction
higher
LLMBoard eligible
Yes
Evaluation key
charxiv-r|llm-stats-current

Benchmark scores retain their original unit. Overall score eligibility is shown separately.

FAQ

Common questions about CharXiv-R.

Which model scores highest on CharXiv-R?

Claude Mythos Preview is currently ranked first with 93.2%.

What does CharXiv-R measure?

CharXiv-R is the reasoning component of the CharXiv benchmark, focusing on complex reasoning questions that require synthesizing information across visual chart elements. It evaluates multimodal large language models on their ability to understand and reason about scientific charts from arXiv papers through various reasoning tasks.

Is a higher score better?

Yes. Higher values rank better for this benchmark.

How many models are compared?

48 model results are currently shown.

Does this benchmark affect the overall score?

Yes. This benchmark can contribute to the current LLMBoard capability score.