Claude Haiku 5.5 Is Now Live on TrueFoundry AI Gateway

Auf Geschwindigkeit ausgelegt: ~ 10 ms Latenz, auch unter Last
Unglaublich schnelle Methode zum Erstellen, Verfolgen und Bereitstellen Ihrer Modelle!
- Verarbeitet mehr als 350 RPS auf nur 1 vCPU — kein Tuning erforderlich
- Produktionsbereit mit vollem Unternehmenssupport
Claude Haiku 5.5 Is Now Live on TrueFoundry AI Gateway
Anthropic has shipped the third model in the Claude 5.5 family. Claude Haiku 5.5 went out on Wednesday, October 7, and it is already in the TrueFoundry AI Gateway model catalog with pricing, context limits and capabilities filled in.
You call it from the same OpenAI-compatible endpoint you use for everything else. No new SDK, no separate integration, no waiting on a gateway release. Change the model string to claude-haiku-5-5 and you are on the new model.
Try Claude Haiku 5.5 on the TrueFoundry AI Gateway →
What shipped on October 7
Anthropic built Haiku 5.5 for high-volume, latency-sensitive work: classification, routing, extraction, summaries, context compaction and subagent tasks. The numbers that matter for a production budget:
- $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens, a tenth of Haiku 4.5's $1 / $5. Cache reads are $0.01 per million, cache writes are $0.125 for the five-minute window and $0.20 for the one-hour window, and the Batch API halves the base rate to $0.05 / $0.25. Prompts over 100,000 tokens bill at $0.50 / $2.50.
- About 75% cheaper to run than Haiku 4.5 on average, by Anthropic's estimate. The saving is smaller than the list-price gap because the newer tokenizer counts the same text as about 30% more tokens. HubSpot recorded its best result yet on its CRM evaluation suite, 92.8% averaged over three runs, and AlphaSense measured a statistically significant gain over Haiku 4.5 on 400 queries (0.84 vs. 0.76).
- The fastest model Anthropic has released. Asana measured more than 30% lower latency on task completions; Box scored 11 points above Haiku 4.5 at about half the latency.
- 1M-token context, 128K max output, up from 200K and 64K on Haiku 4.5, with text and image input. It is the first Haiku with the effort parameter: adaptive thinking is on by default, and the Claude API defaults to medium effort.
- Available from Anthropic directly and on Amazon Bedrock, Google Cloud and Microsoft Foundry. With Haiku 5.5, all three Claude 5.5 models are now in the catalog.

A small model with mid-tier scores
Haiku upgrades usually trade capability for price. On Anthropic's published results this one gains a lot of capability at a lower price. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% against 0.0% for Haiku 4.5 and 16.4% for GPT-6 Luna, which lists at the same $0.10 / $0.50. On OSWorld 2.1 (offline subset) it reaches 72.4% against 15.7% and 48.9%, and on GDPval-AA v2.1 it scores 1620 Elo to Haiku 4.5's 735 and Luna's 1437.

Two caveats. These are Anthropic's own numbers, and Sonnet 5.5 remains well ahead on agentic coding (70.6% on Terminal-Bench 4.0) and on long knowledge work. Run your own evals on your own traffic before you move a workload.
Where it sits on October's price ladder
Put the current models side by side: GPT-6 Luna and Haiku 5.5 at $0.10 per million input tokens, Haiku 4.5 at $1, Sonnet 5.5 at $2, Opus 5.5 at $4, Claude Fable 5.1 at $10. Haiku 5.5 shares the bottom rung with Luna and, on Anthropic's comparison, scores above it on every benchmark both were run on.

That is the argument for treating model choice as gateway config rather than application code. Most production traffic is a large number of short, simple requests and a small number of hard ones. If the simple majority goes to Haiku 5.5 and the rest goes to Sonnet 5.5 or Opus 5.5, the blended bill looks nothing like an all-Sonnet bill. That only works when switching is a config change, which is what the AI Gateway is for.
How to turn it on

Step 1: Enable the model in your provider account. In the gateway, open Models → your Anthropic account → Models Selection, search for claude-haiku-5-5 and tick it. Pricing shows inline ($0.10 input / $0.50 output per 1M tokens) because it comes straight from the catalog. Then set Access Control so the model is scoped to the right teams and virtual keys from day one.
If you run Claude through a cloud, those paths are in the catalog too: AWS Bedrock (global, US, EU, Japan and Australia profiles), Google Vertex AI, Microsoft Foundry, OpenRouter and Perplexity. Regional and data-zone endpoints list at $0.11 / $0.55. Don't see the model in your list yet? + Add Model at the bottom of the list adds it by hand.
Step 2: Call it. Open the Playground, pick the model, and copy the generated snippet in whichever SDK you use. The OpenAI Python version:
from openai import OpenAI
client = OpenAI(
api_key="<your-truefoundry-api-key>",
base_url="https://gateway.truefoundry.ai", # or your self-hosted gateway URL
)
# Model IDs are <provider-account>/<model>. Swap the string, keep everything else.
r = client.chat.completions.create(
model="anthropic-prod/claude-haiku-5-5",
messages=[{"role": "user", "content": "Classify this ticket as billing, bug or feature request: ..."}],
max_tokens=2048, # thinking tokens count toward this limit
extra_headers={"X-TFY-METADATA": '{"team": "support", "feature": "ticket-triage"}'},
)
print(r.choices[0].message.content)Model IDs are namespaced as <your-provider-account>/<model>, so anthropic-prod/claude-haiku-4-5 and anthropic-prod/claude-haiku-5-5 are a one-line swap. The X-TFY-METADATA header tags each call so spend shows up per team or per feature in the metrics view.
Step 3 (optional): Put it behind a virtual model. Create a virtual model and pick Auto Routing. The gateway classifies each request as simple, medium or complex and sends it to the target you assign to that tier. Haiku 5.5 for simple, Sonnet 5.5 for medium, Opus 5.5 for complex, with Haiku 4.5 as a fallback while launch-week rate limits settle, is a sensible starting map. In our benchmarks this kind of routing cut cost 50–70% at about 98% of quality. Your application calls one model name throughout.
Before you move production traffic
- Watch the 100K-token line. Once a prompt goes over 100,000 tokens, the whole request bills at $0.50 / $2.50, and cache reads and writes count toward the threshold. A 20,000-token prompt with a 1,000-token answer costs about $0.0025; a 150,000-token prompt with the same answer costs about $0.0775. Cost tracking applies the tiered rates from the catalog.
- Give thinking room in max_tokens. Thinking tokens count toward max_tokens, so a small limit that worked for a one-word classification on Haiku 4.5 can now stop after the thinking block, before any text. Raise the limit or lower effort.
- Drop temperature, top_p and top_k. A non-default value for any of them returns a 400 error. Assistant prefill and manual budget_tokens also return errors; Anthropic's migration guide covers the replacements.
- Pricing is public and versioned. Every rate in this post comes from the open-source catalog at truefoundry.com/models, the same one cost tracking uses.
- Budgets, rate limits and guardrails already apply. This is an ordinary catalog model, so whatever per-team policy you have configured covers it from the first request.
Start today
Haiku 5.5 is the model for the bulk of your traffic: the classifications, extractions, routing calls and subagent steps that run thousands of times a day. Running it through TrueFoundry means the decision about which requests it handles lives in gateway config, where you can change it without a deploy.
Try Claude Haiku 5.5 on the TrueFoundry AI Gateway →
Related reading
- Claude Sonnet 5.5 Is Now Live on TrueFoundry AI Gateway – the medium tier, same three-step setup
- Introduction to Auto Routing – how complexity-based routing works
- LLM Router: The Three Things That Name Actually Means
- What is an LLM Gateway? – the architectural primer
TrueFoundry AI Gateway bietet eine Latenz von ~3—4 ms, verarbeitet mehr als 350 RPS auf einer vCPU, skaliert problemlos horizontal und ist produktionsbereit, während LiteLM unter einer hohen Latenz leidet, mit moderaten RPS zu kämpfen hat, keine integrierte Skalierung hat und sich am besten für leichte Workloads oder Prototyp-Workloads eignet.












.png)


.png)
.png)
.png)
.png)




.png)








