Blank white background with no objects or features visible.

Meet TrueForge: The open-source, vendor-neutral agent harness. 50% lower cost. Explore Now→

New | Auto Routing in the AI Gateway

Auto Routing: send every request to the optimal model that can handle it

Stop paying frontier prices for routine queries. Optimize model spend by matching each request to the model it actually needs.

Skills Registry gives teams a centralized system for creating, versioning, and managing reusable Agent Skills that can be attached across agents without duplicating prompts or operational logic.

Purple gradient square with white background, shiny surface, and rounded corners in rhombus shape.

Proven cost savings

Book a Demo
arrow1

69%

Cheaper

Measured against an all-Opus 5 baseline.

98%

Quality

at 3.2x cheaper per correct answer.

1.9x

Faster

Mean latency fell from 7.6 seconds to 4.0

Purple pole with diamond shaped top on white background.

Made for Real-World AI at Scale

99.9%

Uptime
Centralized failovers, routing, and guardrails ensure your AI apps stay online, even when model providers don’t.

10B+

Requests processed/month
Scalable, high-throughput inference for production AI.

30%

Average cost optimization
Smart routing, batching, and budget controls reduce token waste.
Complexity Tiers

Complexity-based routing that matches the model to the work

Assign a model to the simple, medium, and complex tiers behind one virtual model name, and the gateway picks the tier per request. Swapping in a cheaper model next quarter is a config edit rather than a release.
Skills registry dashboard with search and list of skills, repositories, and descriptions displayed.
Classification Strategies

Choose between Heuristic or LLM classification to determine complexity

Heuristic classification uses out-of-the-box parameters to instantly classify prompts, adding zero cost or latency. LLM classification calls a small model you nominate to accurately identify prompt complexity.
Skills registry repository creation form with fields for repository, skill name, description, and skill instructions.
Conversation Pinning

Multi-turn conversations that never lose capability

Once a tier answers a turn successfully, pinning keeps later turns of that conversation at the same tier or higher. A short follow-up cannot drop an active conversation onto a cheaper model.
Code editor displaying SKILL.md file with default skill template and structured instructions and capabilities listed.

Enterprise-Ready

Your data and models are securely housed within your cloud / on-prem infrastructure

  • Green circle with a white checkmark symbol inside, indicating confirmation or approval status icon.

    Compliance & Security

    SOC 2, HIPAA, and GDPR standards to ensure robust data protection
  • Green circle with a white checkmark symbol inside, indicating confirmation or approval status icon.

    Governance & Access Control

    SSO + Role-Based Access Control (RBAC) & Audit Logging
  • Green circle with a white checkmark symbol inside, indicating confirmation or approval status icon.

    Enterprise Support & Reliability

    24/7 support with SLA-backed response SLAs
Deploy TrueFoundry in any environment

VPC, on-prem, air-gapped, or across multiple clouds.

No data leaves your domain. Enjoy complete sovereignty, isolation, and enterprise-grade compliance wherever TrueFoundry runs

Cloud deployment options including On-Prem, Multi-Cloud, Air-gapped, and AWS, Google Cloud Platform.

TrueFoundry’s AI Gateway standardized how every team interacts with LLMs, embeddings, and RAG components. Instead of scattered integrations, we now control access, routing policies, and safety guardrails centrally. The ability to optimize for cost or latency without changing applications has been a game-changer. It’s made our AI architecture cleaner, more secure, and far easier to scale.

Smiling man with beard and short dark hair in a collared shirt and blazer.
Indroneel G.
Head of IT, Siemens Healthineers

Frequently asked questions

What is Auto Routing?

Auto Routing is a routing strategy in TrueFoundry's AI Gateway that classifies each incoming request as simple, medium, or complex and sends it to the model you assigned to that tier. Unlike weight, priority, or latency-based routing, the decision depends on the content of the request itself.

How much can I actually save?

In our benchmarks, Auto Routing was 69% cheaper than an all-Opus 5 setup on graded academic datasets and 80% cheaper on production-shaped traffic. Route away from a prior-generation top model instead and graded savings land closer to 50%. The gateway logs the resolved tier for every request, so you can measure your own savings directly.

Will routing to cheaper models hurt quality?

Quality retention was 98% of the baseline under deterministic grading. The known limit is that the free heuristic classifier reads the shape of a prompt rather than its true difficulty, so a short but genuinely hard question can be routed to a cheap tier. Switching that virtual model to the LLM classifier recovers most of that quality for roughly ten points less in savings.

What is the difference between the heuristic and LLM classifiers?

The heuristic classifier scores the prompt against a fixed set of signals inside the gateway process. It is free, fully deterministic, and adds no latency. The LLM classifier calls a small model you nominate to judge each request, which is more accurate on ambiguous prompts but adds one billable model call and some latency per request.

Do I have to change my application code?

No. Auto Routing is a setting on a virtual model, and your application keeps calling that virtual model's name. Changing tiers, swapping models, or switching classification strategy happens in configuration.

Can a conversation switch models partway through?

It can move up and it will not move down. Conversation pinning records the tier that successfully served a turn and keeps later turns at that tier or higher. Pins are automatic, need no configuration, and last through ten minutes of inactivity.

What are the current limitations?

Auto Routing supports exactly three tiers, and the heuristic classifier's signals are fixed rather than customizable. It runs on virtual models only, for the chat, completion, and responses model types. The dashboard form supports one target per tier; multiple targets in a tier require the YAML manifest or the API.
Grey wavy lines on white background, abstract wave pattern with multiple curved lines intersecting smoothly.

GenAI infra- simple, faster, cheaper

Trusted by 30+ enterprises and Fortune 500 companies

Take a quick product tour
Start Product Tour
Product Tour