multimodal benchmark
ScreenSpot is the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments. The dataset comprises over 1,200 instructions from iOS, Android, macOS, Windows and Web environments, along with annotated element types (text and icon/widget), designed to evaluate visual GUI agents' ability to accurately locate screen elements based on natural language instructions.
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | AC | 95.8% | 100.0% | 16 | C | |
| 02 | AC | 95.7% | 93.3% | 16 | C | |
| 03 | AC | 95.4% | 86.7% | 16 | C | |
| 04 | AC | 95.4% | 80.0% | 16 | C | |
| 05 | AC | 94.7% | 73.3% | 16 | C | |
| 06 | AC | 94.7% | 66.7% | 16 | C | |
| 07 | AC | 94.4% | 60.0% | 16 | C | |
| 08 | AC | 94.0% | 53.3% | 16 | C | |
| 09 | AC | 93.6% | 46.7% | 16 | C | |
| 10 | AC | 92.9% | 40.0% | 16 | C | |
| 11 | AC | 88.5% | 33.3% | 16 | C | |
| 12 | AM | 88.1% | 26.7% | 16 | C | |
| 13 | AC | 87.1% | 20.0% | 16 | C | |
| 14 | AM | 85.4% | 13.3% | 16 | C | |
| 15 | AC | 84.7% | 6.7% | 16 | C | |
| 16 | AM | 83.3% | 0.0% | 16 | C |
A closer view of the leading scores on this benchmark.
The leading models and scores on this benchmark.
What ScreenSpot measures and how its scores work.
ScreenSpot is the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments. The dataset comprises over 1,200 instructions from iOS, Android, macOS, Windows and Web environments, along with annotated element types (text and icon/widget), designed to evaluate visual GUI agents' ability to accurately locate screen elements based on natural language instructions.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about ScreenSpot.
Qwen3 VL 32B Instruct is currently ranked first with 95.8%.
ScreenSpot is the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments. The dataset comprises over 1,200 instructions from iOS, Android, macOS, Windows and Web environments, along with annotated element types (text and icon/widget), designed to evaluate visual GUI agents' ability to accurately locate screen elements based on natural language instructions.
Yes. Higher values rank better for this benchmark.
16 model results are currently shown.
Yes. This benchmark can contribute to the current LLMBoard capability score.