llmboard.aiLeaderboard Center
Overall
Overall RankingOpen Models
Tools
Model DirectoryCompare Models
Capabilities
CodingReasoningMathKnowledgeInstruction Following
Price & Efficiency
Price & ValueCapability vs. PriceRuntime Performance
Modalities
Image GenerationVideo GenerationSpeech ModelsEmbeddings
Core Benchmarks
GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-ProView all benchmarks
Methods
Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

NVIDIA model product

Nemotron 3 Ultra

Nemotron 3 Ultra is NVIDIA's frontier-scale open model with 550B total / 55B active parameters, built for agentic reasoning, long-context analysis, tool use, and high-stakes RAG.

Updated Aug 12, 2026. Default version: Nemotron 3 Ultra (550B A55B)

Compare
LLMBoard score62.9Nemotron 3 Ultra (550B A55B)
Coverage100%23 benchmark families
Context window1MTokens
Official input price$0.50Nvidia API

On this page

  • Specification
  • Capability
  • Benchmarks
  • Arena
  • Runtime
  • Pricing
  • Versions
  • About
  • Compare
  • Similar models
  • FAQ

Nemotron 3 Ultra Specifications

Technical details for the model's default version.

Version
Nemotron 3 Ultra (550B A55B)
Released
Jun 4, 2026
Knowledge cutoff
Sep 30, 2025
Parameters
550B
Context window
1M
Max output
128K
Inputs
text
Outputs
text
Open weights
Yes
License
OpenMDW License v1.1

Nemotron 3 Ultra Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Nemotron 3 Ultra (550B A55B) category scores

Nemotron 3 Ultra Benchmark Results

Benchmark scores for Nemotron 3 Ultra (550B A55B).

26 rows
Columns

Show columns

Apex84.8%012100.0%CAug 11, 2026
IMO-AnswerBench92.3%0119100.0%CAug 11, 2026
OmniScience78.7%012100.0%CAug 11, 2026
PinchBench90.0%014100.0%CAug 11, 2026
ProfBench56.0%011100.0%CAug 11, 2026
RULER94.7%014100.0%CAug 11, 2026
IFBench81.7%022996.4%CAug 11, 2026
GDPval46.7%0330.0%CAug 11, 2026
CritPT3.1%0440.0%CAug 11, 2026
LiveCodeBench v689.0%045394.2%CAug 11, 2026
LongBench v261.9%041781.3%CAug 11, 2026
MMLU-ProX83.0%053287.1%CAug 11, 2026
TAU3-Bench22.6%0550.0%CAug 11, 2026
Multi-Challenge63.8%062982.1%CAug 11, 2026
WMT24++83.7%062377.3%CAug 11, 2026
Finance Agent53.7%0880.0%CAug 11, 2026
AA-LCR65.4%091646.7%CAug 11, 2026
MMLU-Pro86.8%0912993.8%CAug 11, 2026
SciCode44.6%091955.6%CAug 11, 2026
Terminal-Bench 2.156.4%171911.1%CAug 11, 2026
Finance Agent v237.5%202624.0%BAug 11, 2026
SWE-bench Multilingual67.7%213439.4%CAug 11, 2026
Humanity's Last Exam37.4%379360.9%CAug 11, 2026
GPQA87.0%4223482.4%CAug 11, 2026
BrowseComp44.4%495815.8%CAug 11, 2026
SWE-Bench Verified70.7%5710546.1%CAug 11, 2026

Nemotron 3 Ultra Arena Results

Preference and agent-evaluation results for the default version.

9 rows
Columns

Show columns

agent tool hallucinationoverall330.0N/A135.6KAug 6, 2026
agent bash recovery stepsoverall43-0.2N/A7.8KAug 6, 2026
agent praise complaintoverall43-0.1N/A878Aug 6, 2026
agentoverall44-0.1N/A152.4KAug 6, 2026
agent steerabilityoverall46-0.2N/A4.6KAug 6, 2026
agent task outcome explicitoverall46-0.2N/A3.6KAug 6, 2026
textoverall491445.010,719N/AAug 10, 2026
text factualityoverall951430.010,719N/AAug 10, 2026
text style controloverall961426.410,719N/AAug 10, 2026

Nemotron 3 Ultra Runtime Performance

Provider-specific output speed and catalog latency for Nemotron 3 Ultra (550B A55B). Runtime does not affect the capability score.

No runtime data

No provider-specific speed or latency record is linked to the default version yet.

Browse runtime rankings

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Nemotron 3 Ultra Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
$0.50 input, $2.5 output per 1M
Official provider
Nvidia
Lowest third-party
From $0.10 input, $0.10 output per 1M via routing.run
Tracked offerings
11
11 rows
Columns

Show columns

Kenarinemotron-3-ultra-550b-a55bglobalN/AN/A1MAug 11, 2026
UnoRouternemotron-3-ultra-550b-a55b:freeglobalN/AN/A1MAug 11, 2026
routing.runnemotron-3-ultraglobal$0.10$0.10131.1KAug 11, 2026
Nvidianvidia/nemotron-3-ultra-550b-a55bglobal$0.50$2.51MAug 11, 2026
NanoGPTnvidia/nemotron-3-ultra-550b-a55bglobal$0.50$2.51MAug 11, 2026
LLM Gatewaynemotron-3-ultra-550bglobal$0.50$2.51MAug 11, 2026
Pioneernvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16global$0.50$2.51MAug 11, 2026
Kilo Gatewaynvidia/nemotron-3-ultra-550b-a55bglobal$0.50$2.2512.3KAug 11, 2026
Together AInvidia/nemotron-3-ultra-550b-a55bglobal$0.60$3.6512.3KAug 11, 2026
Vercel AI Gatewaynvidia/nemotron-3-ultra-550b-a55bglobal$0.60$2.41MAug 11, 2026
OpenRouternvidia/nemotron-3-ultra-550b-a55bglobal$0.60$3.6512.3KAug 11, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Nemotron 3 Ultra Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 rows
Columns

Show columns

Nemotron 3 Ultra (550B A55B)Jun 4, 202662.9550B1M128KYesOpenMDW License v1.1

What is Nemotron 3 Ultra?

Key information about Nemotron 3 Ultra and its available data.

Nemotron 3 Ultra is NVIDIA's frontier-scale open model with 550B total / 55B active parameters, built for agentic reasoning, long-context analysis, tool use, and high-stakes RAG.

It uses a hybrid Latent Mixture-of-Experts (LatentMoE) architecture interleaving Mamba-2, MoE, and select Attention layers, with Multi-Token Prediction (MTP) for native speculative decoding, and is pre-trained on ~20T tokens with an NVFP4 recipe.

Reasoning is configurable on/off (plus a medium-effort mode) via the chat template. It supports up to a 1M-token context and 10 languages (English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, Chinese). 1 license.

Data as of 2026-08-11.

Nemotron 3 Ultra vs nearby models

Open a comparison with the three ranked models immediately above and below this model.

Nemotron 3 UltravsGemini 3 FlashNemotron 3 UltravsMiMo V2 ProNemotron 3 UltravsGPT-5.3-CodexNemotron 3 UltravsGPT-5.1Nemotron 3 UltravsMiMoNemotron 3 UltravsGPT-5.1-Instant

Models similar to Nemotron 3 Ultra

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#109-18.1
NV

Nemotron 3 Super

NVIDIA

44.8 LLMBoard

DetailsCompare
#147-27.7
NV

Nemotron 3 Nano

NVIDIA

35.2 LLMBoard

DetailsCompare
#154-30.5
NV

Nemotron Nano 9B

NVIDIA

32.4 LLMBoard

DetailsCompare
#161-32.4
NV

Llama 3.1 Nemotron Ultra 253B

NVIDIA

30.5 LLMBoard

DetailsCompare
#180-39.9
NV

Llama 3.3 Nemotron Super 49B

NVIDIA

23.0 LLMBoard

DetailsCompare
#220-52.5
NV

Llama 3.1 Nemotron Nano 8B

NVIDIA

10.4 LLMBoard

DetailsCompare

FAQ

Common questions about Nemotron 3 Ultra.

When was Nemotron 3 Ultra released?

Nemotron 3 Ultra's default version was released on Jun 4, 2026.

How much does Nemotron 3 Ultra cost?

Nemotron 3 Ultra's official API price is $0.50 per million input tokens and $2.5 per million output tokens via Nvidia. The lowest tracked third-party offer starts at $0.10 input and $0.10 output via routing.run.

Who created Nemotron 3 Ultra?

Nemotron 3 Ultra was created by NVIDIA.

What is the context window for Nemotron 3 Ultra?

The default version has a 1M token context window.

Is Nemotron 3 Ultra open weight?

Yes. The default version is marked as open weight under OpenMDW License v1.1.

How many API providers offer Nemotron 3 Ultra?

11 provider offerings are linked to the default version.

What models should I compare Nemotron 3 Ultra with?

Nearby ranked alternatives include Gemini 3 Flash, MiMo V2 Pro, GPT-5.3-Codex.