long context benchmark
QASPER is a dataset of 5,049 information-seeking questions and answers anchored in 1,585 NLP research papers. Questions are written by NLP practitioners who read only titles and abstracts, while answers require understanding the full paper text and provide supporting evidence. The dataset challenges models with complex reasoning across document sections for academic document question answering. Each question seeks information present in the full text and is answered by a separate set of NLP practitioners who also provide supporting evidence to answers.
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | MI | 41.9% | 100.0% | 2 | C | |
| 02 | MI | 40.0% | 0.0% | 2 | C |
A closer view of the leading scores on this benchmark.
The leading models and scores on this benchmark.
What Qasper measures and how its scores work.
QASPER is a dataset of 5,049 information-seeking questions and answers anchored in 1,585 NLP research papers. Questions are written by NLP practitioners who read only titles and abstracts, while answers require understanding the full paper text and provide supporting evidence. The dataset challenges models with complex reasoning across document sections for academic document question answering. Each question seeks information present in the full text and is answered by a separate set of NLP practitioners who also provide supporting evidence to answers.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about Qasper.
Phi-3.5-mini-instruct is currently ranked first with 41.9%.
QASPER is a dataset of 5,049 information-seeking questions and answers anchored in 1,585 NLP research papers. Questions are written by NLP practitioners who read only titles and abstracts, while answers require understanding the full paper text and provide supporting evidence. The dataset challenges models with complex reasoning across document sections for academic document question answering. Each question seeks information present in the full text and is answered by a separate set of NLP practitioners who also provide supporting evidence to answers.
Yes. Higher values rank better for this benchmark.
2 model results are currently shown.
No. This benchmark is shown for reference but does not contribute to the overall score.