reasoning benchmark
A rigorous code synthesis evaluation framework that augments existing datasets with extensive test cases generated by LLM and mutation-based strategies to better assess functional correctness of generated code, including HumanEval+ with 80x more test cases
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | MA | 0.803 points | 100.0% | 4 | C | |
| 02 | AC | 0.79 points | 66.7% | 4 | C | |
| 03 | AC | 0.776 points | 33.3% | 4 | C | |
| 04 | AC | 0.703 points | 0.0% | 4 | C |
A closer view of the leading scores on this benchmark.
The leading models and scores on this benchmark.
What EvalPlus measures and how its scores work.
A rigorous code synthesis evaluation framework that augments existing datasets with extensive test cases generated by LLM and mutation-based strategies to better assess functional correctness of generated code, including HumanEval+ with 80x more test cases
Scores are shown in points. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about EvalPlus.
Kimi K2 Base is currently ranked first with 0.803 points.
A rigorous code synthesis evaluation framework that augments existing datasets with extensive test cases generated by LLM and mutation-based strategies to better assess functional correctness of generated code, including HumanEval+ with 80x more test cases
Yes. Higher values rank better for this benchmark.
4 model results are currently shown.
Yes. This benchmark can contribute to the current LLMBoard capability score.