reasoning benchmark
VoiceBench is the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants, evaluating capabilities including general knowledge, instruction-following, reasoning, and safety using both synthetic and real spoken instruction data with diverse speaker characteristics and environmental conditions.
Updated Aug 11, 2026
Higher score ranks better on this benchmark.
| 01 | AC | 74.1% | 100.0% | 1 | C |
The leading models and scores on this benchmark.
What VoiceBench Avg measures and how its scores work.
VoiceBench is the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants, evaluating capabilities including general knowledge, instruction-following, reasoning, and safety using both synthetic and real spoken instruction data with diverse speaker characteristics and environmental conditions.
Scores are shown in ratio. This benchmark is not independently verified and has an evidence level of B.
Benchmark scores retain their original unit. Overall score eligibility is shown separately.
Common questions about VoiceBench Avg.
Qwen2.5-Omni-7B is currently ranked first with 74.1%.
VoiceBench is the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants, evaluating capabilities including general knowledge, instruction-following, reasoning, and safety using both synthetic and real spoken instruction data with diverse speaker characteristics and environmental conditions.
Yes. Higher values rank better for this benchmark.
1 model results are currently shown.
No. This benchmark is shown for reference but does not contribute to the overall score.