Nemotron Super 120B
How TheRouter serves this differently from the vendor
NVIDIA publishes Nemotron 3 Super as open weights, datasets, recipes, NIM deployment, and a native 1M-token context model. AWS Bedrock serves the same model as nvidia.nemotron-super-3-120b with 256K context, 32K max output, text-only I/O, Chat Completions, Invoke, Converse, response streaming, structured outputs, guardrails, and client-side tool calling.
TheRouter exposes it as OpenAI-compatible /v1/chat/completions under nvidia/nemotron-super-120b. TheRouter's live catalog supplies the customer-facing context, output, pricing, modality, and supported-parameter fields; this page records that the upstream Bedrock card is narrower than NVIDIA's native 1M context and that snippets remain capped until operator-funded verification runs.
API guide
Chat completion
Nemotron 3 Super supports OpenAI-compatible chat completions through TheRouter's API. The model supports tools, tool_choice, response_format (JSON mode), stop sequences, temperature, top_p, and max_tokens.
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-super-120b",
"messages": [
{"role": "system", "content": "You are a helpful AI assistant."},
{"role": "user", "content": "Write a Python function to recursively find all .py files in a directory."}
],
"temperature": 0.3,
"max_tokens": 4096
}'Streaming
Nemotron 3 Super supports streaming for real-time token-by-token responses. Use SSE via the OpenAI-compatible API with stream: true.
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-super-120b",
"messages": [{"role": "user", "content": "Explain how latent MoE works in simple terms"}],
"stream": true
}'Tool calling
Nemotron 3 Super has native tool/function calling support, making it well-suited for agentic workflows that require structured tool orchestration.
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-super-120b",
"messages": [{"role": "user", "content": "What is the weather in Tokyo?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}
}],
"tool_choice": "auto"
}'JSON mode
Nemotron 3 Super supports JSON-structured output via response_format, enabling reliable extraction of structured data from unstructured inputs.
curl https://api.therouter.ai/v1/chat/completions \
-H "Authorization: Bearer $THEROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/nemotron-super-120b",
"messages": [{"role": "user", "content": "Extract name, date, and amount from: Invoice INV-2026 from Acme Corp dated May 15, 2026 for $1,250."}],
"response_format": { "type": "json_object" }
}'Fact ledger β every claim on this page traces here
| source | URL | retrieved | |
|---|---|---|---|
| Release date | developer.nvidia.com β | 2026-05-28 | verified |
| Architecture | research.nvidia.com β | 2026-05-28 | verified |
| Native context length | developer.nvidia.com β | 2026-05-28 | verified |
| Context on AWS Bedrock | docs.aws.amazon.com β | 2026-05-28 | verified |
| Max output (TheRouter) | docs.aws.amazon.com β | 2026-05-28 | verified |
| Training data | build.nvidia.com β | 2026-05-28 | verified |
| Post-training | developer.nvidia.com β | 2026-05-28 | verified |
| License | www.nvidia.com β | 2026-05-28 | verified |
| Supported languages | aws.amazon.com β | 2026-05-28 | verified |
| PinchBench | developer.nvidia.com β | 2026-05-28 | verified |
| AIME 2025 | aws.amazon.com β | 2026-05-28 | unknown |
| MMLU-Pro | NVIDIA NeMo Nemotron 3 Super evaluation recipe β | 2026-08-25 | verified |
| GPQA (with tools) | NVIDIA NeMo Nemotron 3 Super evaluation recipe β | 2026-08-25 | verified |
| Arena-Hard V2 | NVIDIA NeMo Nemotron 3 Super evaluation recipe β | 2026-08-25 | verified |
| RULER | NVIDIA NeMo Nemotron 3 Super evaluation recipe β | 2026-08-25 | verified |
| TerminalBench 2.0 | NVIDIA NeMo Nemotron 3 Super evaluation recipe β | 2026-08-25 | verified |
| SWE-Bench Verified | aws.amazon.com β | 2026-05-28 | unknown |
| NVIDIA releases Nemotron 3 Super β open hybrid MoE for agentic reasoning | developer.nvidia.com β | 2026-05-28 | verified |
| AWS Bedrock launches Nemotron 3 Super with fully managed inference | aws.amazon.com β | 2026-05-28 | verified |
| DeepInfra adds Nemotron 3 Super with rapid inference | deepinfra.com β | 2026-05-28 | verified |
| What makes LatentMoE different from standard MoE? | developer.nvidia.com β | 2026-05-28 | to verify |
| What context window does Nemotron 3 Super actually support? | docs.aws.amazon.com β | 2026-05-28 | to verify |
| How does Nemotron 3 Super compare to NVIDIA's previous models? | developer.nvidia.com β | 2026-05-28 | to verify |
| Is Nemotron 3 Super suitable for single-turn chatbot use? | llm-stats.com β | 2026-05-28 | to verify |
| What tools/params are supported via TheRouter API? | docs.aws.amazon.com β | 2026-05-28 | to verify |