multimodal benchmark
GDP.pdf is a knowledge-work vision benchmark that evaluates models on economically valuable professional tasks presented as visual documents (PDFs), testing document-based reasoning, chart and table interpretation, and problem solving without tools.
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | AN | 81.6% | 100.0% | 5 | C | |
| 02 | OP | 30.7% | 75.0% | 5 | C | |
| 03 | AN | 29.8% | 50.0% | 5 | C | |
| 04 | OP | 24.7% | 25.0% | 5 | C | |
| 05 | OP | 22.7% | 0.0% | 5 | C |
A closer view of the leading scores on this benchmark.
The leading models and scores on this benchmark.
What GDP.pdf measures and how its scores work.
GDP.pdf is a knowledge-work vision benchmark that evaluates models on economically valuable professional tasks presented as visual documents (PDFs), testing document-based reasoning, chart and table interpretation, and problem solving without tools.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about GDP.pdf.
Claude Sonnet 5 is currently ranked first with 81.6%.
GDP.pdf is a knowledge-work vision benchmark that evaluates models on economically valuable professional tasks presented as visual documents (PDFs), testing document-based reasoning, chart and table interpretation, and problem solving without tools.
Yes. Higher values rank better for this benchmark.
5 model results are currently shown.
Yes. This benchmark can contribute to the current LLMBoard capability score.