Back to Models

GLM 4.6V Flash

zhipuzhipu/glm-4.6v-flash

API guide

Vision chat completion (image / video / PDF input)

GLM-4.6V-Flash is accessible through the standard OpenAI /v1/chat/completions endpoint at api.therouter.ai. Pass images as image_url content blocks, exactly as you would with GPT-4o. The model supports text, image, video, and PDF inputs in a single request. Function Calling works via the standard tools / tool_choice parameters β€” images and visual tool outputs are handled natively. The API is free; no credits are consumed.

cURL
# Image understanding β€” free via TheRouter
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zhipu/glm-4.6v-flash",
    "messages": [
      {
        "role": "user",
        "content": [
          {
            "type": "image_url",
            "image_url": { "url": "https://therouter.ai/assets/vision-sample.png" }
          },
          {
            "type": "text",
            "text": "Summarise the key trends shown in this chart."
          }
        ]
      }
    ],
    "max_tokens": 1024
  }'

Multimodal Function Calling

GLM-4.6V-Flash is the first GLM vision model with native multimodal Function Calling. Define tools in the standard OpenAI tools format; the model can pass images directly as tool arguments and interpret visual tool outputs (charts, screenshots, search results) without converting them to text first.

cURL
curl https://api.therouter.ai/v1/chat/completions \
  -H "Authorization: Bearer $THEROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zhipu/glm-4.6v-flash",
    "messages": [
      {
        "role": "user",
        "content": [
          { "type": "image_url", "image_url": { "url": "https://therouter.ai/assets/vision-sample.png" } },
          { "type": "text", "text": "Is the server status healthy? Trigger an alert if not." }
        ]
      }
    ],
    "tools": [{
      "type": "function",
      "function": {
        "name": "trigger_alert",
        "description": "Trigger an ops alert",
        "parameters": {
          "type": "object",
          "properties": {
            "service_name": { "type": "string" },
            "severity": { "type": "string", "enum": ["low","medium","high"] },
            "message": { "type": "string" }
          },
          "required": ["service_name","severity","message"]
        }
      }
    }],
    "tool_choice": "auto",
    "max_tokens": 512
  }'
Fact ledger β€” every claim on this page traces here
sourceURLretrieved
Release datehuggingface.co β†—2026-06-09verified
Parameter counthuggingface.co β†—2026-06-09verified
Architecturehuggingface.co β†—2026-06-09verified
Context windowdocs.z.ai β†—2026-06-09verified
Max output tokenshuggingface.co β†—2026-06-09verified
Input modalitiesdocs.z.ai β†—2026-06-09verified
Output modalitydocs.z.ai β†—2026-06-09verified
Pricingdocs.z.ai β†—2026-06-09verified
Native Function Callinghuggingface.co β†—2026-06-09verified
Training cutoffβ€”β€”unknown
Licensehuggingface.co β†—2026-06-09to verify
Local deployment supporthuggingface.co β†—2026-06-09verified
Recommended decoding parametershuggingface.co β†—2026-06-09verified
Z.ai releases GLM-4.6V series with native multimodal Function Callinghuggingface.co β†—2026-06-09verified
Help & contact