Back to Models

Nemotron Super 120B

nvidianvidia/nemotron-super-120b

How TheRouter serves this differently from the vendor

As the vendor operates it

NVIDIA publishes Nemotron 3 Super as open weights, datasets, recipes, NIM deployment, and a native 1M-token context model. AWS Bedrock serves the same model as nvidia.nemotron-super-3-120b with 256K context, 32K max output, text-only I/O, Chat Completions, Invoke, Converse, response streaming, structured outputs, guardrails, and client-side tool calling.

On TheRouter

TheRouter exposes it as OpenAI-compatible /v1/chat/completions under nvidia/nemotron-super-120b. TheRouter's live catalog supplies the customer-facing context, output, pricing, modality, and supported-parameter fields; this page records that the upstream Bedrock card is narrower than NVIDIA's native 1M context and that snippets remain capped until operator-funded verification runs.

API guide

Chat completion

Nemotron 3 Super supports OpenAI-compatible chat completions through TheRouter's API. The model supports tools, tool_choice, response_format (JSON mode), stop sequences, temperature, top_p, and max_tokens.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/nemotron-super-120b",
    "messages": [
      {"role": "system", "content": "You are a helpful AI assistant."},
      {"role": "user", "content": "Write a Python function to recursively find all .py files in a directory."}
    ],
    "temperature": 0.3,
    "max_tokens": 4096
  }'

Streaming

Nemotron 3 Super supports streaming for real-time token-by-token responses. Use SSE via the OpenAI-compatible API with stream: true.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/nemotron-super-120b",
    "messages": [{"role": "user", "content": "Explain how latent MoE works in simple terms"}],
    "stream": true
  }'

Tool calling

Nemotron 3 Super has native tool/function calling support, making it well-suited for agentic workflows that require structured tool orchestration.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/nemotron-super-120b",
    "messages": [{"role": "user", "content": "What is the weather in Tokyo?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather for a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {"type": "string", "description": "City name"}
          },
          "required": ["location"]
        }
      }
    }],
    "tool_choice": "auto"
  }'

JSON mode

Nemotron 3 Super supports JSON-structured output via response_format, enabling reliable extraction of structured data from unstructured inputs.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/nemotron-super-120b",
    "messages": [{"role": "user", "content": "Extract name, date, and amount from: Invoice INV-2026 from Acme Corp dated May 15, 2026 for $1,250."}],
    "response_format": { "type": "json_object" }
  }'
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release datedeveloper.nvidia.com β†—2026-05-28verified
Architectureresearch.nvidia.com β†—2026-05-28verified
Native context lengthdeveloper.nvidia.com β†—2026-05-28verified
Context on AWS Bedrockdocs.aws.amazon.com β†—2026-05-28verified
Max output (TheRouter)docs.aws.amazon.com β†—2026-05-28verified
Training databuild.nvidia.com β†—2026-05-28verified
Post-trainingdeveloper.nvidia.com β†—2026-05-28verified
Licensewww.nvidia.com β†—2026-05-28verified
Supported languagesaws.amazon.com β†—2026-05-28verified
PinchBenchdeveloper.nvidia.com β†—2026-05-28verified
AIME 2025aws.amazon.com β†—2026-05-28unknown
MMLU-ProNVIDIA NeMo Nemotron 3 Super evaluation recipe β†—2026-08-25verified
GPQA (with tools)NVIDIA NeMo Nemotron 3 Super evaluation recipe β†—2026-08-25verified
Arena-Hard V2NVIDIA NeMo Nemotron 3 Super evaluation recipe β†—2026-08-25verified
RULERNVIDIA NeMo Nemotron 3 Super evaluation recipe β†—2026-08-25verified
TerminalBench 2.0NVIDIA NeMo Nemotron 3 Super evaluation recipe β†—2026-08-25verified
SWE-Bench Verifiedaws.amazon.com β†—2026-05-28unknown
NVIDIA releases Nemotron 3 Super β€” open hybrid MoE for agentic reasoningdeveloper.nvidia.com β†—2026-05-28verified
AWS Bedrock launches Nemotron 3 Super with fully managed inferenceaws.amazon.com β†—2026-05-28verified
DeepInfra adds Nemotron 3 Super with rapid inferencedeepinfra.com β†—2026-05-28verified
What makes LatentMoE different from standard MoE?developer.nvidia.com β†—2026-05-28to verify
What context window does Nemotron 3 Super actually support?docs.aws.amazon.com β†—2026-05-28to verify
How does Nemotron 3 Super compare to NVIDIA's previous models?developer.nvidia.com β†—2026-05-28to verify
Is Nemotron 3 Super suitable for single-turn chatbot use?llm-stats.com β†—2026-05-28to verify
What tools/params are supported via TheRouter API?docs.aws.amazon.com β†—2026-05-28to verify
Help & contact