OpenAI Codex Goes On-Premises: What the Dell Partnership Means for Enterprise AI Routing

OpenAI and Dell are bringing Codex inside enterprise data centers. When your coding agent runs on-prem, the routing architecture changes: data locality, session affinity, cost accounting, and governance controls all need rethinking.

Published via OpenAI

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Abstract editorial illustration showing on-premises server infrastructure and cloud paths converging into a single enterprise AI routing layer

The decision most engineering teams face when adopting Codex is not "should we use it?" — it is "where does the agent run, and how do we route traffic to it?" OpenAI's partnership with Dell Technologies, announced May 18, forces that question into the open for enterprises that cannot send their most sensitive code to a cloud API.

Codex is moving from a cloud-only product toward a hybrid one. The Dell collaboration connects Codex to the Dell AI Data Platform (on-premises enterprise data governance) and the Dell AI Factory (on-premises AI inference infrastructure). The practical consequence: enterprises can, for the first time, run Codex-powered agents alongside the codebases, documentation, and business systems that live inside their own data centers — without extracting that data to OpenAI's cloud.

This changes the routing calculus significantly. Here is what engineering teams need to think through before this deployment path becomes generally available.

What changed

OpenAI and Dell are collaborating on two integration points:

Dell AI Data Platform integration. Codex will connect with Dell's on-premises data platform — the storage, organization, and governance layer that enterprises already use for codebase access, documentation, and business records. This closes the gap that makes Codex less useful for regulated industries: the agent can reason over internal context without that context ever leaving the enterprise perimeter.

Dell AI Factory integration (exploratory). OpenAI is exploring how Codex (and ChatGPT Enterprise) can interface with Dell's on-prem AI inference infrastructure — meaning the model serving layer itself could move on-premises, not just the data integration. This is still exploratory, but it signals a path to fully air-gapped Codex deployment.

The backdrop for this is Gartner's 2026 Magic Quadrant for Enterprise AI Coding Agents, published the same week, which named OpenAI a Leader and specifically cited enterprise governance, sandboxing, RBAC, approval gates, and auditable workspace governance as strengths. The Dell partnership is the infrastructure answer to those governance requirements — a coding agent that qualifies as enterprise-grade in the cloud now has a credible on-prem deployment path.

More than 4 million developers use Codex every week, and organizations including Cisco, Datadog, Dell, and NVIDIA are already deploying it across the software development lifecycle — from code review and test coverage to incident response and reasoning across large repositories.

Why it matters for AI engineering teams

Data locality changes the agent's utility. Codex's core advantage is reasoning over large codebases. For enterprises with strict data residency requirements — financial services, healthcare, defense, public sector — sending internal repositories to a cloud API has been a showstopper. The Dell integration removes that blocker. Teams that rejected Codex on compliance grounds now have a credible on-prem evaluation path.

Hybrid routing means split traffic. In a world where some Codex workloads run on-prem and others use OpenAI's cloud, your routing layer needs to know which path each request takes. Code completion for a public repository? Cloud is fine. Code review of internal payroll systems? On-prem only. This distinction is not built into most current routing configurations — it requires per-project or per-team routing policies, not a single unified endpoint.

Cost accounting becomes more complex. Cloud Codex bills as API tokens. On-prem Codex runs on Dell infrastructure with its own CapEx or OpEx allocation. A unified cost model that reconciles both requires your billing tooling to merge sources that were never designed to be compared. Teams currently exporting per-token costs from the OpenAI billing API will need to layer infrastructure cost attribution on top.

Session affinity for agent state. Codex agents increasingly maintain state across multi-turn interactions — a local execution environment, a running test suite, a partially modified file tree. If your routing layer load-balances across multiple Codex instances (cloud plus on-prem), it must route a resumed session back to the same instance that holds the agent's state, or risk losing context mid-task. This session-affinity constraint, already familiar from stateful web services, now applies to coding agents.

Governance controls must span both planes. Approval gates, RBAC, and auditable workspace governance only work if they apply uniformly to both cloud and on-prem Codex traffic. If your governance layer lives in a cloud-only proxy, on-prem Codex calls bypass it entirely.

The router/operator angle

The Dell partnership introduces a two-tier architecture for enterprise Codex deployment:

Team request
  │
  ├─ Public/non-sensitive code → routing proxy → OpenAI cloud Codex
  │                              (token billing, cloud compliance)
  │
  └─ Sensitive/regulated code → on-prem routing → Dell AI Factory Codex
                                 (data stays on-prem, CapEx billing, local RBAC)

Building this routing split requires four things your infrastructure team should start planning now:

  1. Data classification policy. Define which projects, repositories, or data categories may leave the perimeter. This becomes the routing rule itself, not just a policy document.
  2. Per-project routing config. Route requests at the project or team level, not just the user level. A developer permitted to use cloud Codex for public repos must be blocked from routing regulated code there.
  3. Unified observability. Logs, latency metrics, and token/session counts need to aggregate across both deployment planes. A cloud-only dashboard will have blind spots for on-prem traffic.
  4. Session state management strategy. Decide upfront whether stateful agent sessions are cloud-only, on-prem-only, or allowed to migrate. Mixing planes mid-session creates hard-to-debug problems; avoid it until dedicated migration tooling exists.

The Dell AI Factory model-serving integration is the more radical option. If it matures into a generally available product, enterprises could run the full Codex stack with no outbound AI API calls — effectively converting Codex from a cloud-hosted service to a self-managed model deployment. The routing implications mirror the BYOK (bring-your-own-key) vs. provider-managed tradeoff, elevated to the entire agent execution layer.

What TheRouter users should watch or try

  • Define your data-tier routing policy now. Even if on-prem Codex is not generally available today, the architectural split it creates is coming. Documenting which workloads are cloud-eligible and which must stay on-prem is easier to do before the deployment path exists than after engineers are already using it.

  • Audit your Codex session handling. If you proxy Codex today, test whether multi-turn agent sessions are stable under your current routing configuration. Session affinity issues that are minor in single-turn use become critical as long-horizon coding tasks grow.

  • Watch for the Dell AI Factory availability announcement. The current announcement is exploratory. Once on-prem model serving is confirmed available, procurement and infrastructure teams need lead time — Dell AI Factory is server-rack hardware, not a SaaS toggle.

  • Revisit token-only cost models. If your current setup exports only per-token costs from the OpenAI API, it will undercount total Codex costs in a hybrid deployment. Design a cost model that can incorporate infrastructure allocation before the on-prem path ships.

TheRouter routes OpenAI-compatible requests across configured providers. When Codex on-prem surfaces a locally reachable endpoint — whether via the Dell AI Data Platform integration or Dell AI Factory model serving — that endpoint can be added as a provider in a routing configuration. The per-project or data-tier routing split described above maps directly to provider-level routing rules: sensitive workloads to the on-prem provider, non-sensitive to the cloud provider, with fallback and observability handled at the gateway layer.

Help & contact