long context benchmark
LongCodeBench evaluates the code understanding and comprehension abilities of large language models at very long context windows, scaling up to 1M tokens. It tests whether models can reason about extensive codebases provided in a single prompt by answering multiple-choice questions about the code.
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | AM | 84.0% | 100.0% | 2 | C | |
| 02 | AM | 84.0% | 0.0% | 2 | C |
A closer view of the leading scores on this benchmark.
The leading models and scores on this benchmark.
What LongCodeBench measures and how its scores work.
LongCodeBench evaluates the code understanding and comprehension abilities of large language models at very long context windows, scaling up to 1M tokens. It tests whether models can reason about extensive codebases provided in a single prompt by answering multiple-choice questions about the code.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about LongCodeBench.
Nova 2 Lite is currently ranked first with 84.0%.
LongCodeBench evaluates the code understanding and comprehension abilities of large language models at very long context windows, scaling up to 1M tokens. It tests whether models can reason about extensive codebases provided in a single prompt by answering multiple-choice questions about the code.
Yes. Higher values rank better for this benchmark.
2 model results are currently shown.
No. This benchmark is shown for reference but does not contribute to the overall score.