llmboard.aiAI model intelligence
Home

Model Rankings

OverallOpen ModelsAgentCodingReasoningMathKnowledgeInstruction FollowingTextVision
Image GenerationImage Editing
Video GenerationImage to VideoVideo Editing
Text to SpeechSpeech to Text
Embeddings

Efficiency

Chat Token PricingImage PricingVideo PricingAudio Pricing
Chat Speed & LatencyProvider Reliability

Benchmarks

GPQAMMLU-ProAIME 2025SWE-Bench VerifiedMMLUHumanity's Last ExamLiveCodeBenchMATHHumanEvalMMMU-Pro
All Benchmarks

Tools

Model DirectoryCompare Models

Scoring & Data

Scoring & Data
393 models668 benchmarks

Leaderboard Center

Overall RankingCodingCore BenchmarksPrice & ValueRuntime Performance

Modalities

All ModelsImage GenerationImage EditingVideo GenerationImage-to-VideoVideo EditingText-to-SpeechSpeech-to-TextEmbeddings

Data & Methods

Scoring MethodAll BenchmarksReasoningMath

Vendors

All VendorsOpenAIAnthropicGoogle
llmboard.aiCopyright 2026 llmboard.ai

NVIDIA model product

Nemotron 3 Ultra

Nemotron 3 Ultra is NVIDIA's frontier-scale open model with 550B total / 55B active parameters, built for agentic reasoning, long-context analysis, tool use, and high-stakes RAG.

Updated Aug 17, 2026. Default version: Nemotron 3 Ultra (550B A55B)

Compare
LLMBoard Score62.7Nemotron 3 Ultra (550B A55B)
Coverage100%24 benchmark families
Context window1MTokens
Official input price$0.50Nvidia API

On this page

  • Capability
  • Benchmarks
  • Arena
  • Pricing
  • Runtime
  • Specification
  • Versions
  • Compare
  • Similar models
  • About
  • FAQ

Nemotron 3 Ultra Capability Profile

This profile uses the model's current scored version. Arena ratings and prices are shown separately.

Nemotron 3 Ultra (550B A55B) LLMBoard score breakdown

Nemotron 3 Ultra Benchmark Results

Benchmark scores for Nemotron 3 Ultra (550B A55B).

26 rows
Columns

Show columns

Sort by
Benchmark
Score
Rank
Participants
Percentile
Evidence
Evaluated
BenchmarkApexScore84.8%Rank01Participants2Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkIMO-AnswerBenchScore92.3%Rank01Participants20Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkOmniScienceScore78.7%Rank01Participants3Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkPinchBenchScore90.0%Rank01Participants6Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkProfBenchScore56.0%Rank01Participants1Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkRULERScore94.7%Rank01Participants4Percentile100.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkGDPvalScore46.7%Rank03Participants3Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkIFBenchScore81.7%Rank03Participants34Percentile93.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkLongBench v2Score61.9%Rank04Participants17Percentile81.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkCritPTScore3.1%Rank05Participants5Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLU-ProXScore83.0%Rank05Participants32Percentile87.1%EvidenceCEvaluatedAug 17, 2026
BenchmarkTAU3-BenchScore22.6%Rank05Participants5Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkLiveCodeBench v6Score89.0%Rank06Participants56Percentile90.9%EvidenceCEvaluatedAug 17, 2026
BenchmarkMulti-ChallengeScore63.8%Rank06Participants29Percentile82.1%EvidenceCEvaluatedAug 17, 2026
BenchmarkWMT24++Score83.7%Rank06Participants23Percentile77.3%EvidenceCEvaluatedAug 17, 2026
BenchmarkFinance AgentScore53.7%Rank08Participants8Percentile0.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkAA-LCRScore65.4%Rank10Participants18Percentile47.1%EvidenceCEvaluatedAug 17, 2026
BenchmarkMMLU-ProScore86.8%Rank10Participants134Percentile93.2%EvidenceCEvaluatedAug 17, 2026
BenchmarkSciCodeScore44.6%Rank10Participants21Percentile55.0%EvidenceCEvaluatedAug 17, 2026
BenchmarkFinance Agent v2Score37.5%Rank20Participants26Percentile24.0%EvidenceBEvaluatedAug 17, 2026
BenchmarkSWE-bench MultilingualScore67.7%Rank23Participants38Percentile40.5%EvidenceCEvaluatedAug 17, 2026
BenchmarkTerminal-Bench 2.1Score56.4%Rank25Participants28Percentile11.1%EvidenceCEvaluatedAug 17, 2026
BenchmarkHumanity's Last ExamScore37.4%Rank40Participants99Percentile60.2%EvidenceCEvaluatedAug 17, 2026
BenchmarkGPQAScore87.0%Rank46Participants239Percentile81.1%EvidenceCEvaluatedAug 17, 2026
BenchmarkBrowseCompScore44.4%Rank52Participants62Percentile16.4%EvidenceCEvaluatedAug 17, 2026
BenchmarkSWE-Bench VerifiedScore70.7%Rank61Participants111Percentile45.5%EvidenceCEvaluatedAug 17, 2026

Nemotron 3 Ultra Arena Results

Preference and agent-evaluation results for the default version.

9 rows
Columns

Show columns

Sort by
Arena
Category
Rank
Rating / score
Votes
Observations
Result date
Arenaagent tool hallucinationCategoryoverallRank33Rating / score0.0VotesN/AObservations142.5KResult dateAug 13, 2026
Arenaagent praise complaintCategoryoverallRank45Rating / score-0.1VotesN/AObservations924Result dateAug 13, 2026
ArenaagentCategoryoverallRank47Rating / score-0.1VotesN/AObservations160.2KResult dateAug 13, 2026
Arenaagent bash recovery stepsCategoryoverallRank47Rating / score-0.2VotesN/AObservations8.1KResult dateAug 13, 2026
Arenaagent steerabilityCategoryoverallRank49Rating / score-0.2VotesN/AObservations4.8KResult dateAug 13, 2026
Arenaagent task outcome explicitCategoryoverallRank49Rating / score-0.2VotesN/AObservations3.7KResult dateAug 13, 2026
ArenatextCategoryoverallRank50Rating / score1445.2Votes10,720ObservationsN/AResult dateAug 12, 2026
Arenatext factualityCategoryoverallRank96Rating / score1430.2Votes10,720ObservationsN/AResult dateAug 12, 2026
Arenatext style controlCategoryoverallRank97Rating / score1426.5Votes10,720ObservationsN/AResult dateAug 12, 2026

Nemotron 3 Ultra Pricing

Official vendor API pricing appears first, followed by individual provider offers.

Official API
$0.50 input, $2.5 output per 1M
Official provider
Nvidia
Lowest third-party
From $0.10 input, $0.10 output per 1M via routing.run
Tracked offerings
15
15 rows
Columns

Show columns

Sort by
Provider
Provider model ID
Region
Input / 1M
Output / 1M
Context
Updated
ProviderKenariProvider model IDnemotron-3-ultra-550b-a55bRegionglobalInput / 1MN/AOutput / 1MN/AContext1MUpdatedAug 17, 2026
ProviderUnoRouterProvider model IDnemotron-3-ultra-550b-a55b:freeRegionglobalInput / 1MN/AOutput / 1MN/AContext1MUpdatedAug 17, 2026
Providerrouting.runProvider model IDnemotron-3-ultraRegionglobalInput / 1M$0.10Output / 1M$0.10Context131.1KUpdatedAug 17, 2026
ProviderNvidiaProvider model IDnvidia/nemotron-3-ultra-550b-a55bRegionglobalInput / 1M$0.50Output / 1M$2.5Context1MUpdatedAug 17, 2026
ProviderNanoGPTProvider model IDnvidia/nemotron-3-ultra-550b-a55bRegionglobalInput / 1M$0.50Output / 1M$2.5Context1MUpdatedAug 17, 2026
ProviderEden AIProvider model IDdeepinfra/nemotron-3-ultra-550b-a55bRegionglobalInput / 1M$0.50Output / 1M$2.2Context262.1KUpdatedAug 17, 2026
ProviderLLM GatewayProvider model IDnemotron-3-ultra-550bRegionglobalInput / 1M$0.50Output / 1M$2.5Context1MUpdatedAug 17, 2026
ProviderKilo GatewayProvider model IDnvidia/nemotron-3-ultra-550b-a55bRegionglobalInput / 1M$0.50Output / 1M$2.2Context512.3KUpdatedAug 17, 2026
ProviderPioneerProvider model IDnvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16RegionglobalInput / 1M$0.50Output / 1M$2.5Context1MUpdatedAug 17, 2026
ProviderOpenRouterProvider model IDnvidia/nemotron-3-ultra-550b-a55bRegionglobalInput / 1M$0.60Output / 1M$3.6Context512.3KUpdatedAug 17, 2026
ProviderTogether AIProvider model IDnvidia/nemotron-3-ultra-550b-a55bRegionglobalInput / 1M$0.60Output / 1M$3.6Context512.3KUpdatedAug 17, 2026
ProviderVercel AI GatewayProvider model IDnvidia/nemotron-3-ultra-550b-a55bRegionglobalInput / 1M$0.60Output / 1M$2.4Context1MUpdatedAug 17, 2026
ProviderEden AIProvider model IDtogether_ai/nvidia/nemotron-3-ultra-550b-a55bRegionglobalInput / 1M$0.60Output / 1M$3.6Context512.3KUpdatedAug 17, 2026
ProviderFireworks AIProvider model IDaccounts/fireworks/models/nemotron-3-ultra-nvfp4RegionglobalInput / 1M$0.60Output / 1M$2.4Context262.1KUpdatedAug 17, 2026
ProviderEden AIProvider model IDnebius/nvidia/Nemotron-3-Ultra-550b-a55bRegionglobalInput / 1M$1Output / 1M$3Context8KUpdatedAug 17, 2026

Official prices use only the vendor's configured official Provider and positive standard USD PAYG rates. Third-party offers remain explicitly labeled.

Nemotron 3 Ultra Runtime Performance

Provider-specific output speed and catalog latency for Nemotron 3 Ultra (550B A55B). Runtime does not affect the capability score.

No runtime data

No provider-specific speed or latency record is linked to the default version yet.

Browse runtime rankings

Output Speed is generated output tokens received per second. Catalog latency is reported separately from observed provider TTFT.

Nemotron 3 Ultra Specifications

Technical details for the model's default version.

Version
Nemotron 3 Ultra (550B A55B)
Released
Jun 4, 2026
Knowledge cutoff
Sep 30, 2025
Parameters
550B
Context window
1M
Max output
128K
Inputs
text
Outputs
text
Open weights
Yes
License
OpenMDW License v1.1

Nemotron 3 Ultra Versions

Available versions of this model. The score column identifies the version used in the overall ranking.

1 rows
Columns

Show columns

Sort by
Version
Released
LLMBoard
Parameters
Context
Max output
Open weights
License
VersionNemotron 3 Ultra (550B A55B)ReleasedJun 4, 2026LLMBoard62.7Parameters550BContext1MMax output128KOpen weightsYesLicenseOpenMDW License v1.1

Nemotron 3 Ultra vs nearby models

Open a comparison with the three ranked models immediately above and below this model.

Nemotron 3 UltravsGemini 3 FlashNemotron 3 UltravsGPT-5.3-CodexNemotron 3 UltravsMiMo V2 ProNemotron 3 UltravsGPT-5.1Nemotron 3 UltravsMiMoNemotron 3 UltravsGPT-5.1-Instant

Models similar to Nemotron 3 Ultra

Recommendations prioritize the same model type and family, then the closest LLMBoard score.

#119-18.4
NV

Nemotron 3 Super

NVIDIA

44.3 LLMBoard

DetailsCompare
#145-25.4
NV

Nemotron 3.5 Lightning

NVIDIA

37.3 LLMBoard

DetailsCompare
#156-27.8
NV

Nemotron 3 Nano

NVIDIA

34.9 LLMBoard

DetailsCompare
#166-31.0
NV

Nemotron Nano 9B

NVIDIA

31.7 LLMBoard

DetailsCompare
#172-32.7
NV

Llama 3.1 Nemotron Ultra 253B

NVIDIA

30.0 LLMBoard

DetailsCompare
#192-40.4
NV

Llama 3.3 Nemotron Super 49B

NVIDIA

22.3 LLMBoard

DetailsCompare

What is Nemotron 3 Ultra?

Key information about Nemotron 3 Ultra and its available data.

Nemotron 3 Ultra is NVIDIA's frontier-scale open model with 550B total / 55B active parameters, built for agentic reasoning, long-context analysis, tool use, and high-stakes RAG.

It uses a hybrid Latent Mixture-of-Experts (LatentMoE) architecture interleaving Mamba-2, MoE, and select Attention layers, with Multi-Token Prediction (MTP) for native speculative decoding, and is pre-trained on ~20T tokens with an NVFP4 recipe.

Reasoning is configurable on/off (plus a medium-effort mode) via the chat template. It supports up to a 1M-token context and 10 languages (English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, Chinese). 1 license.

Data as of 2026-08-17.

FAQ

Common questions about Nemotron 3 Ultra.

When was Nemotron 3 Ultra released?

Nemotron 3 Ultra's default version was released on Jun 4, 2026.

How much does Nemotron 3 Ultra cost?

Nemotron 3 Ultra's official API price is $0.50 per million input tokens and $2.5 per million output tokens via Nvidia. The lowest tracked third-party offer starts at $0.10 input and $0.10 output via routing.run.

Who created Nemotron 3 Ultra?

Nemotron 3 Ultra was created by NVIDIA.

What is the context window for Nemotron 3 Ultra?

The default version has a 1M token context window.

Is Nemotron 3 Ultra open weight?

Yes. The default version is marked as open weight under OpenMDW License v1.1.

How many API providers offer Nemotron 3 Ultra?

15 provider offerings are linked to the default version.

What models should I compare Nemotron 3 Ultra with?

Nearby ranked alternatives include Gemini 3 Flash, GPT-5.3-Codex, MiMo V2 Pro.