Codex Eval in Production: How OpenAI's Self-Improving Agent Loop Turns Traces Into Scoped Eval Targets

Codex eval infrastructure separates a working agent from a self-improving one. OpenAI's Tax AI converted production traces into eval targets and ran Codex-driven iterations automatically. The architecture applies to any team building agents on an API gateway.

Published via OpenAI Engineering

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Editorial diagram of a self-improving agent loop: production traces flow into structured eval targets, which drive a Codex iteration cycle

The gap between a working demo agent and a self-improving production agent is not about which model you chose — it is about what your system captures at runtime and how that evidence flows back into the improvement cycle. OpenAI's engineering team published a detailed post on May 27 describing exactly how they closed that loop for Thrive Holdings' Tax AI, a Codex-powered agent that processed 7,000 tax returns this season. The architectural lesson extends well beyond tax software: any team running agents on top of an API layer needs to answer the same question that this case study addresses head-on — is your production infrastructure designed to generate evidence, or just outputs?

What happened

OpenAI's forward-deployed team co-built Tax AI with Thrive Holdings to automate preparation of 1040 and 1041 tax returns for Crete's network of 30+ accounting firms. The system launched, and then — unlike most agent deployments — it measurably improved on its own.

At launch, only 25% of returns reached 75% correct field completion. Within six weeks, that figure reached 86%. The system handled progressively harder filings — W-2s first, then K-1s, rental schedules, and multi-property reconciliation — each expansion driven by structured feedback from the previous phase, not by engineers manually reviewing failures.

The three-part architecture that made this work:

  1. Expert practitioner feedback as a structured signal. Practitioners corrected Tax AI predictions before filing. Those corrections were captured not just as diffs but as classified evidence: true extraction miss, prompt coverage gap, workflow noise, or unsupported format. This classification is what converts a correction into an actionable training signal rather than noise.

  2. Production traces that preserve full context. The system captured the complete path from source file through field extraction, citation, downstream tax-engine mapping, and practitioner correction. This is the critical design decision: a trace that records only inputs and outputs cannot tell you where in the pipeline the failure occurred. A trace that captures intermediate provenance can.

  3. A Codex-driven iteration loop closing the cycle. Structured findings became tailored evals. Codex used those evals as a hill-climbing target — investigating failures, proposing code changes, validating against regression suites, and shipping improvements faster than a manual cycle. The coding agent was not the product; it was the infrastructure operator for the product's own quality loop.

Why it matters for AI engineering teams

Most production agent deployments fail at the feedback loop, not the model. The failure mode is identical across domains: a practitioner, customer support rep, or end user corrects an agent's output; that correction enters a ticketing system or spreadsheet; an engineer reviews it some weeks later and maybe updates a prompt. The agent learns nothing from production until a human manually carries the signal upstream.

The Tax AI architecture breaks this in two places. First, it forces production to generate evidence by design — every interaction is structured to be inspectable, not just logged. Second, it uses a coding agent to accelerate the engineering side of the loop, so the time from "pattern detected in production" to "fix deployed and regression-tested" shrinks from weeks to days.

The eval targeting strategy is also worth noting. The team did not run an undifferentiated eval suite. They grouped failures by pattern (repeated misses on fair-rental-day fields; confusion across multi-property packages), turned patterns into specific eval targets, and let Codex climb each hill. Precision here matters: broad evals that measure aggregate accuracy can improve globally while a specific failure class gets worse. Scoped evals catch regressions that aggregate metrics hide.

The router/operator angle

For teams running agents through an API gateway, the self-improving pattern surfaces two specific infrastructure decisions that are often underspecified.

Trace design is an operator responsibility, not an afterthought. If your gateway records only model inputs and outputs, you can measure accuracy but you cannot debug where in a multi-step agent pipeline the failure originated. Production-grade agent tracing means capturing intermediate states: tool calls, retrieved documents, intermediate reasoning steps, structured field extractions with source citations. This is not about model observability — it is about pipeline observability, and the gateway layer is where that trace record either gets assembled or gets lost.

Eval infrastructure changes your model routing policy. Once you have scoped evals that measure specific capability slices — rental property extraction accuracy, K-1 field precision, multi-doc reconciliation — you can route different pipeline steps to different models based on demonstrated per-task quality, not just benchmark rank. A model that scores high on a general coding leaderboard may not be the right choice for a specific extraction task once you have the evals to measure it. This is the step from "route by cost" to "route by measured task quality."

A practical checklist for teams that want to adopt this pattern:

  • Log intermediate states, not just I/O. Every tool call, retrieval, extraction step, and validation check should be part of the trace. Budget storage for this.
  • Classify corrections before storing them. A raw practitioner correction is noise. A correction tagged with failure category (extraction miss, prompt gap, ambiguous input, workflow skip) is a training signal.
  • Build scoped evals before you need them. The Tax AI team had eval infrastructure in place before the self-improvement loop could run. Teams that try to add evals in response to failures are always behind.
  • Use your coding agent as an operator, not just a product. Codex-class agents can write and run eval scripts, open PRs with fixes, and validate against regression suites. That loop runs faster than manual iteration — but only if your production traces give it a concrete hill to climb.
  • Watch your routing policy as models improve. If you swap a provider or upgrade a model version, existing scoped evals may reveal regressions in specific capability slices that aggregate benchmarks miss.

What TheRouter users should watch or try

Teams routing agent workflows through TheRouter are well-positioned to add the trace-capture layer at the gateway without changing application code. Request-level logging at the API layer gives you a foundation for intermediate state capture when combined with structured tool-call responses from the model.

The self-improvement pattern also strengthens the case for treating routing as a per-task decision informed by evals rather than a static model selection. When you have scoped eval results per capability slice, you can route document-extraction steps differently from reasoning steps — and update that routing policy as your evals accumulate evidence from production.

The OpenAI post is worth reading in full for the concrete rental-property example, which shows how a single practitioner correction flows from a raw data diff through classification, grouping, eval targeting, and Codex-scoped engineering task. That trace from correction to improvement is the pattern — the domain is replaceable.

Help & contact