NVIDIA model product
Nemotron 3 Super is a 120B total / 12B active parameter hybrid Mamba-Attention Mixture-of-Experts model optimized for agentic reasoning, coding, planning, tool calling, and long-context analysis.
Updated Aug 12, 2026. Default version: Nemotron 3 Super (120B A12B)
Technical details for the model's default version.
This profile uses the model's current scored version. Arena ratings and prices are shown separately.
Benchmark scores for Nemotron 3 Super (120B A12B).
| WMT24++ | 86.7% | 01 | 23 | 100.0% | C | |
| RULER | 91.8% | 02 | 4 | 66.7% | C | |
| Bird-SQL (dev) | 41.8% | 05 | 7 | 33.3% | C | |
| Arena-Hard v2 | 73.9% | 06 | 16 | 66.7% | C | |
| LiveCodeBench | 81.2% | 07 | 73 | 91.7% | C | |
| HMMT 2025 | 94.7% | 08 | 33 | 78.1% | C | |
| SciCode | 42.0% | 11 | 19 | 44.4% | C | |
| MMLU-ProX | 79.4% | 12 | 32 | 64.5% | C | |
| Multi-Challenge | 55.2% | 12 | 29 | 60.7% | C | |
| AA-LCR | 58.3% | 13 | 16 | 20.0% | C | |
| IFBench | 72.6% | 14 | 29 | 53.6% | C | |
| Tau2 Airline | 56.3% | 18 | 23 | 22.7% | C | |
| Terminal-Bench | 25.8% | 22 | 25 | 12.5% | C | |
| Tau2 Retail | 62.8% | 24 | 26 | 8.0% | C | |
| MMLU-Pro | 83.7% | 29 | 129 | 78.1% | C | |
| Tau2 Telecom | 64.4% | 29 | 35 | 17.6% | C | |
| SWE-bench Multilingual | 45.8% | 32 | 34 | 6.1% | C | |
| AIME 2025 | 90.2% | 47 | 114 | 59.3% | C | |
| Terminal-Bench 2.0 | 31.0% | 49 | 49 | 0.0% | C | |
| Humanity's Last Exam | 22.8% | 53 | 93 | 43.5% | C | |
| BrowseComp | 31.3% | 54 | 58 | 7.0% | C | |
| GPQA | 82.7% | 71 | 234 | 70.0% | C | |
| SWE-Bench Verified | 53.7% | 87 | 105 | 17.3% | C |
Preference and agent-evaluation results for the default version.
| text factuality | overall | 136 | 1389.6 | 7,511 | N/A | |
| text | overall | 149 | 1377.6 | 7,537 | N/A | |
| text style control | overall | 180 | 1360.4 | 7,537 | N/A |
Provider-specific output speed and catalog latency for Nemotron 3 Super (120B A12B). Runtime does not affect the capability score.
| DeepInfra | 19.905 tok/s | N/A | 262.1K | 262.1K |
Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.
Official vendor API pricing appears first, followed by individual provider offers.
| Kenari | nemotron-3-super-120b-a12b | global | N/A | N/A | 262.1K | |
| NanoGPT | nvidia/nemotron-3-super-120b-a12b | global | $0.05 | $0.25 | 262.1K | |
| Kilo Gateway | nvidia/nemotron-3-super-120b-a12b | global | $0.085 | $0.40 | 262.1K | |
| OpenRouter | nvidia/nemotron-3-super-120b-a12b | global | $0.085 | $0.40 | 1M | |
| Pioneer | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8 | global | $0.09 | $0.45 | 1M | |
| Vercel AI Gateway | nvidia/nemotron-3-super-120b-a12b | global | $0.15 | $0.65 | 256K | |
| Nvidia | nvidia/nemotron-3-super-120b-a12b | global | $0.20 | $0.80 | 262.1K | |
| Perplexity Agent | nvidia/nemotron-3-super-120b-a12b | global | $0.25 | $2.5 | 1M | |
| Nebius Token Factory | nvidia/nemotron-3-super-120b-a12b | global | $0.30 | $0.90 | 256K | |
| Synthetic | hf:nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 | global | $0.30 | $1 | 262.1K | |
| TensorX | nvidia/nemotron-3-super-120b-a12b | global | $0.30 | $0.90 | 262.1K |
Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.
Available versions of this model. The score column identifies the version used in the overall ranking.
| Nemotron 3 Super (120B A12B) | 44.8 | 120B | 262.1K | 262.1K | Yes | NVIDIA Open Model License Agreement |
Key information about Nemotron 3 Super and its available data.
Nemotron 3 Super is a 120B total / 12B active parameter hybrid Mamba-Attention Mixture-of-Experts model optimized for agentic reasoning, coding, planning, tool calling, and long-context analysis.
It introduces LatentMoE (projecting tokens into a compressed latent space for expert routing, enabling 4x more experts at the same inference cost), Multi-Token Prediction for native speculative decoding (up to 3x faster generation), and native NVFP4 pretraining on Blackwell.
The hybrid architecture interleaves Mamba-2 layers for linear-time sequence processing with strategically placed Transformer attention layers as global anchors, supporting a 1M-token context window. 2 million rollouts. 2x higher throughput than GPT-OSS-120B while maintaining comparable accuracy.
Data as of 2026-08-11.
Open a comparison with the three ranked models immediately above and below this model.
Recommendations prioritize the same model type and family, then the closest LLMBoard score.
Common questions about Nemotron 3 Super.
Nemotron 3 Super's default version was released on Mar 11, 2026.
Nemotron 3 Super's official API price is $0.20 per million input tokens and $0.80 per million output tokens via Nvidia. The lowest tracked third-party offer starts at $0.05 input and $0.25 output via NanoGPT.
Nemotron 3 Super was created by NVIDIA.
The default version has a 262.1K token context window.
Yes. The default version is marked as open weight under NVIDIA Open Model License Agreement .
11 provider offerings are linked to the default version.
Nearby ranked alternatives include o4 mini, GPT-5-mini, K EXAONE 236B A23B.