Mistral OCR 4 Adds Structural Blocks and Bounding Boxes: What Changes for AI Teams Routing Document Workloads

Mistral OCR 4 ships bounding boxes, typed-block classification, and per-word confidence scores for PDF and document ingestion pipelines. Here is what the release changes for engineering teams designing RAG, agent, and document-routing workflows.

TheRouter Newsroomvia Mistral AI
Mistral OCR 4 structured document AI pipeline diagram showing block extraction and routing flow

When Mistral released OCR 4 on June 23, 2026, the headline was bounding boxes — the most-requested capability since OCR 1. But the more durable operational shift is what structural block extraction does to the document ingestion layer sitting upstream of every LLM API call. For teams routing document workloads to foundation models, this changes the semantic chunking decision, the citation pipeline, and the endpoint-versioning policy in ways that are not obvious from the benchmark numbers alone.

What happened

Mistral OCR 4 (mistral-ocr-4-0) is now the model behind mistral-ocr-latest. It extracts structured content from PDFs, DOCX, PPTX, and image formats across 170 languages and returns three new layers on top of the plain text that predecessors produced:

  • Bounding boxes: Each page's blocks are localized with top-left and bottom-right pixel coordinates. Downstream systems can highlight or redact specific regions without re-parsing the document.
  • Typed-block classification: Every block carries a structural label — text, title, list, table, image, equation, caption, code, references, aside_text, header, footer, signature — in reading order. The label is set via include_blocks=True on the OCR API call.
  • Inline confidence scores: Per-page and per-word confidence scores arrive alongside extracted content. Teams can gate low-confidence regions for human-in-the-loop review rather than passing noisy content into a generation step.

Pricing is $4 per 1,000 pages for standard API access and $5 per 1,000 pages through Document AI (the no-code Studio path). A 50% Batch API discount reduces the standard rate to $2 per 1,000 pages for bulk processing.

On OlmOCRBench, OCR 4 scored 85.20, ranking first among tested systems. In a separate head-to-head human evaluation over 600+ documents across 12+ languages, annotators preferred OCR 4 over competing outputs in the majority of documents, with an average win rate of 72%.

Why it matters for AI engineering teams

The impact is not primarily about OCR accuracy — it is about what structured block output unlocks further down the routing pipeline.

Semantic chunking becomes block-native. Previous OCR outputs were raw markdown strings that chunking strategies had to re-parse with heuristics (headings, separator lines, paragraph detection). OCR 4 block labels supply the same signal directly: a title block is a section boundary, a table block is a retrieval unit, a code block warrants different embedding treatment. Teams no longer have to maintain a secondary chunking layer that fights the OCR output.

Citations get first-class support. Document AI pipelines often attach source citations to LLM outputs, but citation reliability depends on being able to map a generated claim back to a specific document region. Bounding boxes combined with block labels give retrieval-augmented generation systems the coordinates they need for in-context highlighting and precise references — particularly useful for financial, legal, and compliance document workloads.

Confidence scores enable tiered routing. Where a page has low confidence scores, the pipeline can route that page to a more expensive frontier vision model for re-extraction, or flag it for human review, rather than letting noisy content propagate into the generation step. This is a natural insertion point for routing logic: OCR 4 handles the bulk, frontier models handle the exception path.

mistral-ocr-latest now points to OCR 4. Teams that pinned to the latest alias picked up the new behavior automatically on June 23. The include_blocks parameter is opt-in (default False), so basic text extraction calls are backward-compatible. However, any pipeline that processes the raw page JSON structure should validate that the shape returned by OCR 4 matches expectations — OCR 4 adds new top-level fields on page objects when blocks are enabled.

Self-hosted deployment is available. OCR 4 runs in a single container, making it deployable in air-gapped or data-sovereignty environments. For teams with document-data residency requirements, self-managed deployment (available to enterprise customers through Mistral) means OCR processing stays within the organization's own infrastructure before any content is passed to a cloud LLM API.

The router/operator angle

For teams that route document-extraction workloads across multiple models or providers, OCR 4 clarifies a two-tier architecture:

  1. Extraction tier: Mistral OCR 4 as a dedicated document-processing endpoint. Route all PDF and structured document ingestion through mistral-ocr-4-0 (or mistral-ocr-latest once you have validated the schema change). Use include_blocks=True and confidence_scores_granularity=word for pipelines that need bounding boxes or downstream verification.

  2. Generation tier: A separately routed LLM API call (OpenAI, Anthropic, DashScope, etc.) that receives the structured block output from step 1 as context. Because the extraction output is already typed and sequenced, the system prompt can tell the LLM which block types to weight (e.g., title, table, code) and which to ignore (e.g., header, footer).

Routing decisions to revisit in light of OCR 4:

  • Endpoint pinning: Audit any integration that passes model=mistral-ocr-latest and verify the include_blocks default (currently False). Teams that rely on extra page-level fields should test against mistral-ocr-4-0 explicitly before relying on the latest alias.
  • Batch processing cost: At $2 per 1,000 pages on the Batch API, OCR 4 is cost-competitive for bulk document pipelines. Evaluate whether switching to batch mode for non-real-time ingestion jobs reduces cost materially relative to synchronous calls.
  • Confidence-gated fallback: Add a routing step after OCR 4 that inspects per-page confidence scores. Pages below a threshold (e.g., < 0.7) can be sent to a secondary extraction path — another model, a manual queue, or a higher-fidelity scan — rather than being passed directly to the LLM.
  • Provider dependency: Mistral OCR 4 is currently available via the Mistral API and Document AI Studio. Teams routing through a gateway that provides unified access to multiple providers should confirm that the OCR endpoint is accessible through the gateway's OpenAI-compatible layer, or call Mistral's OCR endpoint directly alongside the main completion routing logic.

What TheRouter users should watch or try

Mistral OCR 4's API is a separate, non-chat endpoint (POST /v1/ocr) that does not flow through the standard OpenAI-compatible /v1/chat/completions path. Teams using TheRouter for completion routing should treat document extraction as a separate upstream step: run OCR 4 to produce structured block output, then route the resulting context into the completion call through TheRouter's normal provider routing.

For teams evaluating Mistral as a provider on TheRouter, the OCR 4 release is a meaningful signal about Mistral's investment in the document-AI and enterprise-RAG stack. The addition of bounding boxes and block labels positions Mistral's document processing layer as a credible alternative to vision-model-based extraction pipelines for PDF-heavy workloads — with significantly lower cost per page than routing a scanned document through a frontier vision model.

Help & contact