language benchmark
Big-Bench Hard (BBH) is a suite of 23 challenging tasks selected from BIG-Bench for which prior language model evaluations did not outperform the average human-rater. These tasks require multi-step reasoning across diverse domains including arithmetic, logical reasoning, reading comprehension, and commonsense reasoning. The benchmark was designed to test capabilities believed to be beyond current language models and focuses on evaluating complex reasoning skills including temporal understanding, spatial reasoning, causal understanding, and deductive logical reasoning.
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | AC | 88.9% | 100.0% | 12 | C | |
| 02 | XI | 88.4% | 90.9% | 12 | C | |
| 03 | AM | 86.9% | 81.8% | 12 | C | |
| 04 | AC | 84.5% | 72.7% | 12 | C | |
| 05 | DE | 84.3% | 63.6% | 12 | C | |
| 06 | AM | 82.4% | 54.5% | 12 | C | |
| 07 | AC | 82.4% | 45.5% | 12 | C | |
| 08 | OP | 81.5% | 36.4% | 12 | C | |
| 09 | AM | 79.5% | 27.3% | 12 | C | |
| 10 | AC | 78.2% | 18.2% | 12 | C | |
| 11 | NR | 67.8% | 9.1% | 12 | C | |
| 12 | BA | 30.4% | 0.0% | 12 | C |
A closer view of the leading scores on this benchmark.
The leading models and scores on this benchmark.
What BBH measures and how its scores work.
Big-Bench Hard (BBH) is a suite of 23 challenging tasks selected from BIG-Bench for which prior language model evaluations did not outperform the average human-rater. These tasks require multi-step reasoning across diverse domains including arithmetic, logical reasoning, reading comprehension, and commonsense reasoning. The benchmark was designed to test capabilities believed to be beyond current language models and focuses on evaluating complex reasoning skills including temporal understanding, spatial reasoning, causal understanding, and deductive logical reasoning.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about BBH.
Qwen3 235B A22B is currently ranked first with 88.9%.
Big-Bench Hard (BBH) is a suite of 23 challenging tasks selected from BIG-Bench for which prior language model evaluations did not outperform the average human-rater. These tasks require multi-step reasoning across diverse domains including arithmetic, logical reasoning, reading comprehension, and commonsense reasoning. The benchmark was designed to test capabilities believed to be beyond current language models and focuses on evaluating complex reasoning skills including temporal understanding, spatial reasoning, causal understanding, and deductive logical reasoning.
Yes. Higher values rank better for this benchmark.
12 model results are currently shown.
Yes. This benchmark can contribute to the current LLMBoard capability score.