GPT-6.1 Sol Is Now Live on TrueFoundry AI Gateway

Conçu pour la vitesse : latence d'environ 10 ms, même en cas de charge
Une méthode incroyablement rapide pour créer, suivre et déployer vos modèles !
- Gère plus de 350 RPS sur un seul processeur virtuel, aucun réglage n'est nécessaire
- Prêt pour la production avec un support complet pour les entreprises
GPT-6.1 Sol Is Now Live on TrueFoundry AI Gateway
One week after GPT-6 Sol arrived in the catalog, OpenAI has shipped its successor. GPT-6.1 Sol went out on Tuesday, September 29, and it is already in the TrueFoundry AI Gateway model catalog with pricing, context limits and capabilities filled in.
You call it from the same OpenAI-compatible endpoint you use for everything else. No new SDK, no separate integration, no waiting on a gateway release. Change the model string to gpt-6.1-sol on your OpenAI account (openai/gpt-6.1-sol on OpenRouter) and you are on the new model.
Try GPT-6.1 Sol on the TrueFoundry AI Gateway →
What shipped on September 29
GPT-6.1 Sol replaces GPT-6 Sol at the same sticker price, with cheaper caching and the same limits. The numbers that matter for a production budget:
- $2 per million input tokens and $10 per million output, unchanged from GPT-6 Sol. Cache reads are $0.10 per million, half of GPT-6 Sol's $0.20, and cache writes are $2.50. From 272K tokens the rate steps up to $4 input / $15 output, with cache reads at $0.20 and cache writes at $5. Batch runs at half price, $1 / $5. For agents that resend a long system prompt or codebase context on every turn, the cache rate is the number to watch.
- One-fifth of GPT-6 Astra's list price on both input and output ($10 / $50 for Astra).
- 1.05M-token context, 128K max output, with thinking on. Reasoning effort runs from none through low, medium, high and xhigh to max, with medium as the default.

Astra's tier, at a fifth of the price
The catalog puts GPT-6 Astra at $10 per million input tokens and $50 per million output. GPT-6.1 Sol lists at $2 and $10, exactly one-fifth on both sides, and its cache reads at $0.10 are one-tenth of Astra's $1.
Where it sits on September's price ladder
Put the month's releases side by side: GPT-6 Luna at $0.10 per million input tokens, GPT-6 Sol, GPT-6.1 Sol and Claude Sonnet 5.5 at $2, Claude Opus 5.5 at $4, GPT-6 Astra at $10. GPT-6.1 Sol takes GPT-6 Sol's slot on the ladder at the same price, with the lowest cache-read rate in that $2 row.

That is the argument for treating model choice as gateway config rather than application code. Most production traffic is a large number of easy requests and a small number of hard ones. If the easy majority goes to Luna, the middle goes to GPT-6.1 Sol and only the hard tail goes to Astra, the blended bill looks nothing like an all-Astra bill. That only works when switching is a config change, which is what the AI Gateway is for.
How to turn it on

Step 1: Enable the model in your provider account. In the gateway, open Models → your OpenAI account → Models Selection, search for gpt-6.1-sol and tick it. Pricing shows inline ($2 input / $10 output per 1M tokens) because it comes straight from the catalog. Then set Access Control so the model is scoped to the right teams and virtual keys from day one.
OpenAI direct and OpenRouter (openai/gpt-6.1-sol) are both in the catalog today; Microsoft Foundry and AWS Bedrock entries are landing.
Don't see the model in your list? + Add Model at the bottom of the list adds it by hand.
Step 2: Call it. Open the Playground, pick the model, and copy the generated snippet in whichever SDK you use. The OpenAI Python version:
from openai import OpenAI
client = OpenAI(
api_key="<your-truefoundry-api-key>",
base_url="https://gateway.truefoundry.ai", # or your self-hosted gateway URL
)
# Model IDs are <provider-account>/<model>. Swap the string, keep everything else.
r = client.chat.completions.create(
model="openai-main/gpt-6.1-sol",
messages=[{"role": "user", "content": "Summarise this incident in three bullets: ..."}],
extra_headers={"X-TFY-METADATA": '{"team": "platform", "feature": "incident-summary"}'},
)
print(r.choices[0].message.content)
Model IDs are namespaced as <your-provider-account>/<model>, so openai-main/gpt-6-sol and openai-main/gpt-6.1-sol are a one-line swap. The X-TFY-METADATA header tags each call so spend shows up per team or per feature in the metrics view.
Step 3 (optional): Put it behind a virtual model. Create a virtual model and pick Auto Routing. The gateway classifies each request as simple, medium or complex and sends it to the target you assign to that tier. GPT-6 Luna for simple, GPT-6.1 Sol for medium, GPT-6 Astra for complex, with GPT-6 Sol as a fallback while launch-week rate limits settle, is a sensible starting map. In our benchmarks this kind of routing cut cost 50–70% at about 98% of quality. Your application calls one model name throughout.
Before you move production traffic
- Check long prompts. From 272K tokens GPT-6.1 Sol is billed at $4 input / $15 output per million, with cache reads at $0.20, the same input and output step-up as GPT-6 Sol. Batch requests run at half the standard rate ($1 / $5).
- Pricing is public and versioned. Every rate in this post comes from the open-source catalog at truefoundry.com/models, the same one cost tracking uses.
- Budgets, rate limits and guardrails already apply. This is an ordinary catalog model, so whatever per-team policy you have configured covers it from the first request.
- Microsoft Foundry and AWS Bedrock entries will follow in the catalog as they merge.
Start today
GPT-6.1 Sol is the model to reach for when a request needs more than Luna but Astra pricing is hard to justify, and when agents resend long context that the $0.10 cache rate makes cheap. Running it through TrueFoundry means the decision about which requests those are lives in gateway config, where you can change it without a deploy.
Try GPT-6.1 Sol on the TrueFoundry AI Gateway →
Related reading
- GPT-6 Astra Pricing: Where Caching Still Pays at $10 Input – the top of the price ladder
- Introduction to Auto Routing – how complexity-based routing works
- LLM Router: The Three Things That Name Actually Means
- What is an LLM Gateway? – the architectural primer
TrueFoundry AI Gateway offre une latence d'environ 3 à 4 ms, gère plus de 350 RPS sur 1 processeur virtuel, évolue horizontalement facilement et est prête pour la production, tandis que LiteLM souffre d'une latence élevée, peine à dépasser un RPS modéré, ne dispose pas d'une mise à l'échelle intégrée et convient parfaitement aux charges de travail légères ou aux prototypes.












.png)
.png)
.png)
.png)

.png)
.png)
.png)
.png)
.png)
.png)
.png)


.webp)





