Blank white background with no objects or features visible.

TrueFoundry Named Frost & Sullivan's 2026 Global Transformational Innovation Leader. Read report

Conheça o TrueForge: o agent harness de código aberto e independente de fornecedor. Custo 50% menor. Explorar agora→

OpenAI Codex Governance: Controls a Developer Cannot Switch Off

By Ashish Dubey

Published: September 29, 2026

⚡ TL;DR
  • A coding agent is the hardest thing in your estate to govern: it runs on a laptop you do not control, holds a long-lived credential, reads whole repositories and spends per token.
  • Governance means four things holding without developer cooperation: identity, model control, spend attribution, and audit nobody can disable.
  • Codex routes through the TrueFoundry AI Gateway in two auth modes — API key, or ChatGPT Business/Enterprise with the provider’s API key field deliberately left empty and the credential moved to x-tfy-api-key.
  • wire_api = "responses" is required, not a preference, and Virtual Model slugs must match the real model ID or thinking-token handling breaks.
  • Fleet enforcement is an hourly MDM script that file-locks the managed config and issues short-lived tokens — with one caveat: Windows enforcement only works from Codex v0.156.0.

Why coding agents are their own governance problem

Every other AI workload runs on infrastructure you own — your cluster, your network policy, a credential you can rotate. Codex is a CLI process on a developer’s laptop that reads a TOML file in their home directory and talks straight to a model provider. Five properties make it a different class of problem:

Property Why it breaks normal controls
Runs on an endpoint you do not own No sidecar, no mesh, no egress proxy. Config is a user-writable file.
Holds a long-lived credential A key in ~/.codex/config.toml outlives the need for it and anyone’s memory of it.
Reads the whole repository Context assembly is automatic and broad. The agent decides what to send.
Calls tools MCP servers, shell, web search — outbound actions, not just completions.
Spends money per token Continuous, unbudgeted, attributed to whoever’s key is in the file. Often nobody.

The usual reflex is a policy document. It fails for one reason: everything it asks for lives in a file the developer can edit.

What “governed” actually has to mean

Four properties have to hold on every request:

  • Identity. The request carries the individual developer, not a team key. A shared key makes every downstream control — rate limits, budgets, grants, audit — resolve to one meaningless principal.
  • Model control. Which models a developer may reach is a security and cost decision. Pricing differs by an order of magnitude between tiers; data-handling terms differ between providers.
  • Spend attribution. Per developer, per team, per repository — the usual surprise line item once thirty engineers adopt an agent (AI coding agent pricing).
  • Audit the developer cannot switch off. Logging behind a local flag is not audit; it is a request.

A gateway gives you all four — but only if the client points at it and keeps pointing at it. There are three postures for that, and only one is enforcement:

Posture How it is applied Survives a developer edit? Fits
Documented setup Wiki page; developer edits ~/.codex/config.toml No Pilots, contractors
Managed config Admin-owned file read at higher precedence Yes, until a client release changes precedence Most enterprises
Managed config + lock + short-lived tokens Admin file, immutability flag, scheduled credential refresh Yes, and a stolen credential expires Regulated fleets

Locking the file without rotating the token handles tampering but ignores credential theft; rotating without locking does the reverse. You need both.

Where teams get this wrong

Treating the model provider account as the control point. Identity, SSO, workspace membership and admin-console policy for Codex live in OpenAI’s console, not your gateway — as does prompt and output retention, which “is governed by your agreement with OpenAI.” A gateway controls routing, spend, guardrails and its own logging, not your provider contract.

Assuming one MDM push equals permanent enforcement. Clients change config precedence between releases. Codex v0.149.0 dropped managed_config.toml on Windows entirely — it now ignores the file with a startup warning (Ignoring deprecated managed config file), with no registry equivalent. A fleet configured before that release is unenforced and reports no error.

Shipping a shared API key. The fastest way to get thirty developers working, and it destroys every downstream control at once. You cannot rate-limit a person or answer “who sent that prompt” from a tenant-wide credential.

Want to see what governed coding-agent traffic actually looks like?
Point one Codex install at a gateway and read the trace.

How this works in TrueFoundry

Codex treats the TrueFoundry AI Gateway as a model provider; everything else follows from traffic arriving somewhere you control. The gateway details come from the code snippet in the UI.

TrueFoundry AI Gateway code snippet panel showing the gateway base URL and API key to copy into a Codex provider block
TrueFoundry AI Gateway code snippet panel showing the gateway base URL and API key to copy into a Codex provider block

Auth mode 1 — API key

The single-developer form: the token goes in Authorization and the gateway resolves it to a person.

model = "gpt-5.2-codex"
model_provider = "truefoundry"

[model_providers.truefoundry]
name = "TrueFoundry AI Gateway"
base_url = "{GATEWAY_BASE_URL}"
wire_api = "responses"

[model_providers.truefoundry.http_headers]
Authorization = "Bearer TFY_API_KEY"

Auth mode 2 — ChatGPT Business/Enterprise

The developer’s ChatGPT subscription pays for inference and the gateway still governs the traffic. Two things change: requires_openai_auth = true, and the TrueFoundry credential moves from Authorization into x-tfy-api-key, because Codex needs Authorization for its own ChatGPT OAuth token.

model = "gpt-5.2-codex"
model_provider = "truefoundry"

[model_providers.truefoundry]
name = "TrueFoundry AI Gateway"
base_url = "{GATEWAY_BASE_URL}"
wire_api = "responses"
requires_openai_auth = true

[model_providers.truefoundry.http_headers]
x-tfy-api-key = "TFY_API_KEY"

On the gateway side, create a provider under Integrations → Providers of type OpenAI, base URL https://chatgpt.com/backend-api/codex, and leave the API key field empty.

The empty field is the signal, not an omission. If a stored key is present, the gateway uses it instead of the developer’s ChatGPT credentials — breaking the subscription flow and possibly routing traffic under the wrong account. An admin who “helpfully” fills that field converts thirty individually-attributed developers into one anonymous API-key tenant, and nothing errors.

Left empty, you get dual attribution: TrueFoundry records the developer from x-tfy-api-key, OpenAI attributes usage to their ChatGPT seat. Model IDs look like chatgpt-codex/gpt-5.2-codex.

Two constraints follow from the credential being tied to one ChatGPT account. Only one of a virtual model’s targets can be a ChatGPT-subscription model account; API-key-backed targets coexist alongside it. And a per-developer 401, 403 or 429 does not put the shared target into load-balancer cooldown, so one developer’s exhausted quota cannot affect anyone else.

(The docs disagree on plan eligibility: the enterprise-security page says Business or Enterprise, the CLI page also lists Personal. [VERIFY] before promising a tier.)

Why wire_api = "responses" is required

This is the line teams copy wrong and then debug for an afternoon. The gateway appends the /responses endpoint only for the Responses wire. Set wire_api = "chat" and the request goes to /chat/completions against the same base URL, which is not a valid endpoint for Codex traffic. It is also what makes gpt-5.x-codex thinking tokens behave correctly.

The Codex CLI page frames it more softly — responses for gpt-5.x-codex, chat for other models. They reconcile: for Codex-family models the Responses wire is mandatory.

Virtual Model slugs must match the real model ID

Codex sends thinking tokens to certain models and identifies them by matching the model string. Fully qualified names — openai-main/gpt-5, azure-openai/gpt-5 — do not match, so feature detection fails.

TrueFoundry Virtual Models list showing configured virtual model slugs and their provider targets

TrueFoundry Virtual Models list showing configured virtual model slugs and their provider targets

The rule: use a slug matching the actual model ID, e.g. gpt-5.2-codex. Codex recognises those IDs and enables the right features. Do not put the fully qualified name in the Codex config.

Virtual model detail view showing the slug and its load-balanced provider targets

Virtual model detail view showing the slug and its load-balanced provider targets

Slugs are unique across the tenant and set in Virtual Model Provider Group settings. Both the slug and the full group/model path work in a request body; Codex needs the slug form.

Navigation to the Virtual Model configuration screen inside a provider group

Navigation to the Virtual Model configuration screen inside a provider group

Virtual model slug field in the advanced settings form

Virtual model slug field in the advanced settings form

Behind one slug you can weight targets — openai-main/gpt-5 at 70, azure-openai/gpt-5 at 30 — and change the mix without touching a laptop, as in model deprecations and staged cutovers.

Slug naming is also where clients differ sharply. Cursor picks its request format from slug keywords and switches format silently if the slug contains anthropic, claude, openai, gpt, gemini or vertex. Codex wants the model ID in the slug; Cursor wants those keywords out.

Cursor virtual model slug configuration form

Cursor virtual model slug configuration form

Fleet enforcement over MDM

A per-developer config is a documented setup, not a control. Fleet enforcement uses tfy-local-ai-setup, a binary pushed by your MDM and run as root (Administrator on Windows) on a recommended hourly cadence, from github.com/truefoundry/tfy-local-ai-setup/releases; the Codex scripts pin RELEASE_TAG="v1.3.5". Each run does four things:

  1. Saves or loads config from a root-owned file — macOS /Library/Preferences/com.truefoundry.tfy-local-ai-setup.conf, Linux /etc/tfy/tfy-local-ai-setup.conf, Windows C:\ProgramData\TrueFoundry\tfy-local-ai-setup.conf (macOS: chown root:wheel, chmod 644).
  2. Installs or updates the binary, skipped when the installed tag matches RELEASE_TAG.
  3. Detects the logged-in user and silently refreshes the token from ~/.tf/refresh-token.
  4. Writes and locks the managed config — chflags schg on macOS, chattr +i on Linux, an icacls ACL on the directory on Windows.

Step three matters most. The laptop never holds a long-lived gateway token: the refresh token is valid 30 days from last issue, every successful run rotates it, and the credential in the managed config is freshly fetched and scoped to that user’s identity.

Required config: GATEWAY_URL, CONTROL_PLANE_URL, TENANT_NAME. Optional: CODEX_GATEWAY_URL, CODEX_AUTH_MODE, CODEX_DEFAULT_MODEL, CODEX_SETTINGS_FILE. CODEX_AUTH_MODE="api-key" is the default; "chatgpt-subscription" makes the binary write requires_openai_auth = true and move the token to x-tfy-api-key. Machines without Codex are a clean no-op — the binary skips them and exits.

The precedence chains, and why Windows is different

On macOS and Linux the admin file wins, reapplied on every launch:

/etc/codex/managed_config.toml → ~/.codex/config.toml → CLI flags

model_provider = "truefoundry"
model = "gpt-5.2-codex"

[model_providers.truefoundry]
name     = "TrueFoundry Gateway"
base_url = "https://<your-truefoundry-gateway-url>"
wire_api = "responses"

[model_providers.truefoundry.http_headers]
Authorization = "Bearer <freshly-fetched-tfy-token>"

The top-level model key is written only when CODEX_DEFAULT_MODEL is set; unset, model choice stays with the developer while the provider stays enforced.

On Windows the file, the order and the schema all change.

C:\ProgramData\OpenAI\Codex\requirements.toml → ~/.codex/config.toml → CLI flags → C:\ProgramData\OpenAI\Codex\config.toml

Note the last position: Windows config.toml is the machine default with the lowest precedence of the four, so a developer’s own file beats it. requirements.toml beats everything. The provider block matches macOS; the model pin does not:

# C:\ProgramData\OpenAI\Codex\requirements.toml
model_provider = "truefoundry"
# ... same [model_providers.truefoundry] and http_headers blocks ...

[models.new_thread]
model = "gpt-5.2-codex"

The requirements schema has no top-level model key. CODEX_DEFAULT_MODEL goes under [models.new_thread] instead. Copy the macOS block onto Windows unchanged and the model pin silently does nothing.

Two version numbers decide whether any of this is real on Windows. v0.149.0 dropped managed_config.toml there. v0.156.0 is the minimum for enforcement — Codex honours model_provider in requirements.toml only from that build; older builds fall back to the lowest-precedence config.toml, which a developer overrides from their home directory. Check codex --version before treating Windows as controlled. Windows Codex management in the setup binary arrives in tfy-local-ai-setup v1.3.5.

Locking turns the file into a control:

‍

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Start free
Table of Contents

One Gateway for Every LLM, Agent and MCP Server

Book a 30-min with our AI expert

Book a Demo

The fastest way to build, govern and scale your AI

Book Demo
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Discover More

No items found.
September 29, 2026
|
5 min read

LLM as a Judge, Running Inline as a Gateway Guardrail

No items found.
September 29, 2026
|
5 min read

Langfuse vs LangSmith: Which LLM Observability Platform Fits

No items found.
September 29, 2026
|
5 min read

Datadog LLM Observability Pricing in 2026: What It Actually Costs

No items found.
September 29, 2026
|
5 min read

AI Red Teaming for Agents: Attacks, Campaigns, and Runtime Defense

No items found.
No items found.

Recent Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Frequently asked questions

What does OpenAI Codex governance actually require?

Four controls that hold without developer cooperation: a per-developer identity on every request, admin control over reachable models, spend attributed to a person and a team, and audit produced on the network path. In practice: an admin-owned config file read at higher precedence than the developer’s, a file lock, and short-lived credentials.

Can Codex use a ChatGPT Business or Enterprise subscription through a gateway?

Yes. Set requires_openai_auth = true, move the gateway credential to x-tfy-api-key so Codex can use Authorization for its own ChatGPT OAuth token, and create the gateway provider with the API key field empty — a stored key overrides the developer’s ChatGPT credentials and breaks the flow. You get the developer’s identity on the gateway side and their ChatGPT seat on OpenAI’s.

Is Codex enforceable on Windows?

Only from Codex v0.156.0. v0.149.0 removed managed_config.toml support with no registry equivalent; enforcement moved to requirements.toml, a different schema with no top-level model key. Below v0.156.0 only the lowest-precedence config.toml applies, which a developer’s own config overrides.

Can I deploy TrueFoundry in my own VPC or on-prem?

Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.

Does TrueFoundry support MCP and AI agents generally?

Yes. It includes an MCP Gateway, an Agent Gateway, and an MCP & Agents Registry with tool-level access control. Agents on LangGraph, CrewAI, AutoGen, or a custom framework can all be governed centrally.

Does it integrate with my observability stack?

Yes. The gateway is OpenTelemetry-compliant and plugs into Grafana, Datadog, or Prometheus. Each LLM classifier call produces its own span, so classifier latency is visible separately from the served model’s.

Take a quick product tour
Start Product Tour
Product Tour