multimodal benchmark
Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | BY | 89.2% | 100.0% | 17 | C | |
| 02 | BY | 89.0% | 93.8% | 17 | C | |
| 03 | AC | 88.0% | 87.5% | 17 | C | |
| 04 | XI | 87.7% | 81.3% | 17 | C | |
| 05 | MA | 87.4% | 75.0% | 17 | C | |
| 06 | MI | 85.4% | 68.8% | 17 | C | |
| 07 | GO | 84.8% | 62.5% | 17 | C | |
| 08 | AC | 84.2% | 56.3% | 17 | C | |
| 09 | GO | 78.6% | 50.0% | 17 | C | |
| 10 | AM | 77.9% | 43.8% | 17 | C | |
| 11 | GO | 76.1% | 37.5% | 17 | C | |
| 12 | AC | 74.5% | 31.3% | 17 | C | |
| 13 | AC | 73.3% | 25.0% | 17 | C | |
| 14 | AC | 71.8% | 18.8% | 17 | C | |
| 15 | AC | 71.4% | 12.5% | 17 | C | |
| 16 | GO | 66.2% | 6.3% | 17 | C | |
| 17 | MI | 55.0% | 0.0% | 17 | C |
A closer view of the leading scores on this benchmark.
The leading models and scores on this benchmark.
What Video-MME measures and how its scores work.
Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about Video-MME.
Seed 2.1 Pro is currently ranked first with 89.2%.
Video-MME is the first-ever comprehensive evaluation benchmark of Multi-modal Large Language Models (MLLMs) in video analysis. It features 900 videos totaling 254 hours with 2,700 human-annotated question-answer pairs across 6 primary visual domains (Knowledge, Film & Television, Sports Competition, Life Record, Multilingual, and others) and 30 subfields. The benchmark evaluates models across diverse temporal dimensions (11 seconds to 1 hour), integrates multi-modal inputs including video frames, subtitles, and audio, and uses rigorous manual labeling by expert annotators for precise assessment.
Yes. Higher values rank better for this benchmark.
17 model results are currently shown.
Yes. This benchmark can contribute to the current LLMBoard capability score.