Envoy AI Gateway (Now Agent Router): An Honest Review
.png)
Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
The rename, in one paragraph
If you searched for “Envoy AI Gateway” and landed here: you are in the right place, and the product still exists. It is now Agent Router, it lives at theagentrouter.ai, and it is a standalone project in the Agentic AI Foundation (a Series of LF Projects, LLC) rather than an Envoy sub-project under the CNCF. The announcement is dated 9 September 2026; the move took effect the next day.
Nothing technical broke. The namespace is still envoy-ai-gateway-system, the API group still aigateway.envoyproxy.io, the CLI still aigw, the licence still Apache 2.0. Manifests in your repo keep working. What changed is the governance home, the docs domain, the repo path and the name. We use “Agent Router” from here.
What Agent Router actually is
Agent Router is an additive layer on Envoy Gateway — not a fork — that teaches Envoy to speak LLM. Its own framing is the clearest one-liner in the category: “Agent Router configures. Envoy handles the traffic.”
Maintainers come from Tetrate, Bloomberg, Tencent, Netflix, AMD and Nutanix, per the 1.0 announcement. Tetrate is the commercial driver, running both a hosted service and a self-managed enterprise product on top. Weekly community meeting, public design proposals, a Discord.
Version history
Roughly one minor release a quarter — a healthy cadence. v1.0 carries the strongest sentence in the project’s documentation:
“We will never break the APIs unless there is a critical security issue, and we will always provide a migration path in the release notes if we ever must.”
That is a real commitment, written down, and more than most commercial gateways will put in print.
Architecture
Agent Router splits cleanly into control plane and data plane.
Control plane. The Kubernetes API Server is the configuration interface. The AI Gateway Controller manages AI-specific resources and configures the External Processor, fine-tuning xDS through the Envoy Gateway extension server mechanism; the Envoy Gateway Controller handles core proxy config.
Data plane. Envoy Proxy, an External Processor (ext-proc) doing AI-specific transformation and validation, and a Rate Limit Service for token-based limiting. The MCP gateway is a lightweight proxy inside the Agent Router sidecar.
Five CRDs, all served at v1beta1 and all covered by the 1.0 stability guarantee: AIGatewayRoute (routes LLM traffic and selects backends), AIServiceBackend (declares an upstream provider or self-hosted model), BackendSecurityPolicy (centralises upstream credentials), GatewayConfig (gateway-level config) and MCPRoute (MCP multiplexing, tool filtering, authorization).
Configuration is YAML applied with kubectl. Documented field names give a feel for the shape — modelNameOverride on a backend for model virtualization, credentialOverride for per-request upstream credentials, streamIdleTimeout for bounded streaming waits, and on an MCPRoute:
# Illustrative shape only - see the API reference for the full schema
spec:
toolSelector:
includeRegex:
- "issue_.*"
backendSelector: ... # CEL, default action Deny (v1.1)
hostnames: ... # max 16 (v1.1)
You are also managing a version matrix. The v1.1.0 baseline is Go 1.26.4, Envoy Gateway v1.8.1, Envoy Proxy v1.38.1, Gateway API v1.5.1, Inference Extension v1.0.2 and MCP Go SDK v1.7.0. Envoy Gateway upgrades become your dependency.
[SCREENSHOT: Agent Router — docs site architecture diagram showing control plane (AI Gateway Controller + Envoy Gateway Controller) and data plane (Envoy Proxy + External Processor + Rate Limit Service)]
Feature inventory, as documented
Providers — 16 out of the box: OpenAI, Azure OpenAI, Gemini, Vertex AI, Bedrock, Anthropic, Mistral, Cohere, Groq, Together AI, DeepInfra, DeepSeek, Hunyuan, SambaNova, Grok, and Tetrate Agent Router Service. Self-hosted vLLM and Ollama are reachable via the CLI.
Unified API. OpenAI-compatible: chat, completions, embeddings, image generation, audio and the Responses API, plus Anthropic-native /anthropic/v1/messages. Cross-provider translation handles cases like Anthropic Messages to Bedrock Converse.
Traffic handling. Token-aware rate limiting. QuotaPolicy for token budgets across time windows. Provider fallback and automatic failover. Model name virtualization via modelNameOverride, so swapping providers is a config change not a code change. Header and body mutations. In v1.1, streamIdleTimeout with failover — if it fires before the first token, a retry moves to the next backend.
Inference routing and credentials. InferencePool and Endpoint Picker support, via both HTTPRoute + InferencePool and AIGatewayRoute + InferencePool — genuinely useful if you run your own GPUs. BackendSecurityPolicy centralises upstream auth: API keys, AWS SigV4, Azure, and GCP cloud-native identity including Workload Identity and GKE ADC.
MCP support is substantial. Server multiplexing behind one endpoint with tools auto-prefixed by backend (github__issue_read), tool filtering, OAuth with PKCE per the MCP authorization spec, streamable HTTP with Last-Event-ID reconnection. The standout is fine-grained authorization using CEL over JWT claims and MCP context — request.mcp.method, request.mcp.tool, request.mcp.params, request.auth.jwt — with tools/list returning only what the caller may use. Few projects match that depth.
Observability. Prometheus metrics following OpenTelemetry GenAI semantic conventions, with separate reasoning-token accounting. OTel tracing with OpenInference compatibility. Access logs carrying AI metadata. v1.1 adds gen_ai.* span attributes and an example Grafana dashboard in the repo.
Guardrails are not a first-party feature. There is no guardrails page in the docs; the Security capability page is largely links to Envoy Gateway’s own security docs. Tetrate’s comparison lists the OSS row as “guardrail hooks”, with actual AI guardrails as an enterprise product.
The genuine strengths
A credible project deserves a credible account of why.
The data plane is Envoy. Not a new proxy written in Python or Node — the proxy that already carries production traffic at enormous scale. Data-plane operational risk is about as low as it gets.
The stability guarantee is rare. A committed-stable v1beta1 control-plane API with a written promise and a migration-path commitment beats what most vendors offer.
Cross-industry maintainership. Tetrate, Bloomberg, Netflix, AMD, Nutanix and Tencent all have maintainers. Not a single-vendor project wearing an OSS badge; the bus factor is real protection.
No crippled community build. Tetrate states it plainly: “No proprietary data plane, no crippled community build, nothing held out of the project so we can sell it back.” The enterprise value-add is fleet management, not gated data-plane features — an honourable open-core line. The project is also standards-native throughout (Gateway API, Inference Extension, OTel GenAI conventions, OpenInference, MCP), so there is little proprietary surface to get locked into.
It is genuinely easy to try. One command, no Kubernetes, no Docker, on Linux or macOS:
OPENAI_API_KEY=sk-... aigw run
That is an OpenAI-compatible router on localhost:1975. Few open source AI gateway projects reach a working endpoint that fast.
Published adopters. The homepage lists Alan by Comma Soft, Bloomberg, LY Corporation, National Research Platform, Nutanix, Paper Compute Co., Simplifai, Stacklok, Tencent Cloud, Tetrate and Unwrap, and the 1.0 post thanks the first three by name. We found no formal case studies with metrics — logos and a thank-you list are not case studies — but three commercial products build on it (Tetrate Agent Router Service and Enterprise, Nutanix Agent Gateway), which is decent evidence of production-grade code.
[SCREENSHOT: Agent Router — homepage adopters row with the eleven published company logos]
Where teams hit trouble
None of this criticises the code. It is an inventory of what the open-source project does not include, evidenced by absence from its docs or presence on Tetrate’s paid comparison, checked 25 September 2026.
1. There is no UI. The docs navigation contains no admin console page. Every model onboarding, key rotation and quota change is a Kubernetes manifest through your GitOps pipeline; Tetrate sells “The Management Console” as an enterprise component. Fine for a platform team. For the twelve application teams who want to try a new model on Thursday, it is a ticket.
2. No platform RBAC or SSO for gateway administration. Kubernetes RBAC on the CRDs is not role-based control over models, MCP servers and teams. Data-plane request auth (JWT, OIDC) comes from Envoy Gateway, but that is caller authentication, not administrative access control. Tetrate lists SSO with Okta/Entra/OIDC/SAML and scoped consoles as enterprise rows.
3. Cost attribution is metrics, not chargeback. Token metrics land in Prometheus. Turning that into per-team spend with a price catalogue, showback and chargeback is yours to build, including keeping the catalogue current. Tetrate is candid: with mixed self-hosted GPUs and per-token APIs, “Attribution has to normalize both or the number you hand finance is fiction.”
4. Budgets are per-gateway, not fleet-wide. QuotaPolicy gives token budgets across time windows, but Tetrate’s comparison marks the OSS behaviour “per gateway”. Run gateways in three regions and an agent refused in one can succeed in another.
5. No prompt or response logging. Access logs carry model names and token counts; there is no documented prompt/completion capture, retention, search or replay. v0.6 added body redaction — the direction of travel is stripping bodies, not storing them.
6. No human-in-the-loop approval, and guardrails are bring-your-own. MCP authorization is allow/deny at request time; nothing gates a destructive tool call behind an approval. Tetrate sells an agent kill switch, but no per-call approval workflow exists. Nor is there PII detection, prompt-injection filtering, content moderation or secrets scanning in the OSS project.
7. Kubernetes is the production prerequisite. aigw run is for laptops. Production means a cluster, plus Envoy Gateway, plus the AI Gateway controller, plus a Rate Limit Service. If you do not run Kubernetes with the Gateway API today, that is a substantial project before you route your first token.
8. Who is accountable at 3am. Community support is GitHub, Discord, a Slack and a Monday meeting. Tetrate Enterprise sells 24/7 support with an SLA, a named account team, CVE-patched builds on a maintained release train, air-gapped distribution, and a current model catalogue. Their comparison labels the OSS column “you maintain it” — the whole review in four words.
Build vs buy, laid out plainly
If your platform team runs Envoy Gateway and has capacity, the right-hand column is a roadmap you can execute. If not, it is a product you are about to build by accident. Price it the way you price any build-vs-buy decision: headcount-months, not licence fees.
Where TrueFoundry sits
We are not going to pretend Agent Router is bad. The difference is scope and who carries the operational load.
TrueFoundry’s AI Gateway is the same architectural idea — one control point in front of model and tool traffic — with the surrounding platform included rather than assigned to you. It adds roughly 3-4 ms, handles 350+ RPS on 1 vCPU, and fronts 1,000+ LLMs through one OpenAI-compatible API.

A console, and RBAC that means something. Model onboarding, key rotation and quota changes happen in a UI with an audit trail, so teams move without a platform ticket. Subjects are users, teams, virtual accounts or agents; resources span provider accounts, MCP servers, agents, clusters and workspaces, each with its own role family. SSO via OIDC or SAML 2.0, with SCIM.

Cost attribution and budgets finance will accept. Budget Limiting V2 with tenant and team-scoped budgets, scoped by subject, model, provider account or metadata, with warn-only mode. Cost tracking runs off an open-source pricing catalog (github.com/truefoundry/models) with region-wise and tiered rates; attribution flows through X-TFY-METADATA. This is the gap that takes longest to close yourself — see cost attribution and team budgets.
Guardrails as a catalogue, not a hook. Nine built-in (secrets, code safety, SQL sanitizer, regex, prompt injection, PII, content moderation, Cedar, OPA) plus 17 external providers, at four points: LLM input, LLM output, MCP pre-tool and MCP post-tool.

One honest caveat: content moderation, PII and prompt-injection guardrails work only when TrueFoundry hosts the gateway, not on your own infrastructure. The rest work either way.
MCP governance with an approval gate. Per-tool enable/disable, read-only and destructive annotations, and human-in-the-loop approval with named/destructive/all scopes, grant expiry and Email/Slack/PagerDuty/Teams notification. No equivalent in the Agent Router docs.

Observability and deployment, not two more projects. OTEL traces and metrics over HTTP or gRPC, Prometheus scraping self-hosted, metrics APIs for model, MCP, guardrail, cache and routing. Runs as SaaS across 12+ regions and three clouds, or self-hosted in your VPC, on-prem or air-gapped — traffic stays in your infrastructure, TrueFoundry out of the live path.

Head to head
One note: Agent Router’s MCP authorization with CEL is excellent, and Apache 2.0 with no gated data plane is a permanent property no commercial product can offer.
Related reading
- What Is an LLM Gateway — the category, defined
- Best AI Gateway in 2026 — the wider field
- Envoy Proxy Alternatives — the layer underneath
- Total Cost of Ownership for GenAI Infrastructure — pricing the build column
- What Is an MCP Gateway — the tool-traffic half
Conclusion
Agent Router — the project you may still think of as Envoy AI Gateway — is one of the better things to happen to open-source AI infrastructure. The data plane is Envoy. The API is stable and the maintainers said so in writing. Six companies keep it alive. The MCP authorization model beats most commercial equivalents. And the open-core line is honest.
The reason to choose something else is not quality. It is division of responsibility. Agent Router gives you the routing layer. The console, the platform RBAC, the per-team cost attribution, the fleet-wide budgets, the prompt logs, the approval workflows, the guardrail catalogue and the 3am pager are all yours — and they are most of the work. A platform team with strong Kubernetes capability and the headcount to own that stack should run it themselves. That is a legitimate choice, and this post is not an argument against it.
If that is not a project you want to start, buy the gateway with the work already done.
TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
What is Envoy AI Gateway?
An open-source AI gateway built as an additive layer on Envoy Gateway and the Kubernetes Gateway API, with Envoy as the data plane and CRDs as the control plane. Since 10 September 2026 it is called Agent Router and lives in the Agentic AI Foundation.
Is Envoy AI Gateway production ready?
It reached v1.0 GA on 23 June 2026 with a stable v1beta1 API and a written no-breaking-changes promise. Published adopters include Bloomberg, LY Corporation, Nutanix and Tencent Cloud. The production question is less about the code than the operational surface you supply around it.
Envoy AI Gateway vs a commercial AI gateway — how do I choose?
Count what you would build. If you run Kubernetes with the Gateway API and have engineers with capacity for a console, RBAC, cost attribution, budgets, prompt logging and guardrails, Agent Router is a legitimate, cheap-in-licence choice. If that list reads like a roadmap you did not plan for, buy it.
Do old Envoy AI Gateway URLs still work?
Yes. aigateway.envoyproxy.io 301-redirects to theagentrouter.ai, and github.com/envoyproxy/ai-gateway redirects to github.com/theagentrouter/agent-router.
TrueFoundryを自社のVPC内またはオンプレミスで実行できますか?
はい。TrueFoundryは、お客様のVPC、オンプレミス、エアギャップ環境、ハイブリッド環境、または複数のクラウドにまたがって動作し、データがお客様のドメイン外に出ることはありません。これが、規制の厳しい企業がSaaSのみのゲートウェイではなくTrueFoundryを選ぶ主な理由です。
Does TrueFoundry work with my observability stack?
Yes. The gateway exports OTEL traces and metrics over HTTP or gRPC, and exposes Prometheus scraping when self-hosted.










.png)
.png)
.png)
.png)
.png)


.webp)
.webp)


.webp)
.webp)
.webp)






