Datadog LLM Observability Pricing in 2026: What It Actually Costs
.png)
Auf Geschwindigkeit ausgelegt: ~ 10 ms Latenz, auch unter Last
Unglaublich schnelle Methode zum Erstellen, Verfolgen und Bereitstellen Ihrer Modelle!
- Verarbeitet mehr als 350 RPS auf nur 1 vCPU — kein Tuning erforderlich
- Produktionsbereit mit vollem Unternehmenssupport

TrueFoundry AI Gateway bietet eine Latenz von ~3—4 ms, verarbeitet mehr als 350 RPS auf einer vCPU, skaliert problemlos horizontal und ist produktionsbereit, während LiteLM unter einer hohen Latenz leidet, mit moderaten RPS zu kämpfen hat, keine integrierte Skalierung hat und sich am besten für leichte Workloads oder Prototyp-Workloads eignet.



Steuern, implementieren und verfolgen Sie KI in Ihrer eigenen Infrastruktur
Aktuelle Blogs
Häufig gestellte Fragen
What is Datadog LLM Observability pricing in 2026?
Two tiers under the new name, Agent Observability: Free at $0/month for up to 40K LLM spans, 15-day retention; and Pro at $160/month annual with 100K LLM spans included and $3.50 per additional 10K. Month-to-month is $200 and $4.20 per 10K; on-demand $240 and $5.00 per 10K. Rates are higher in ap1, ap2 and uk1.
Is Datadog LLM Observability the same as Agent Observability?
Yes — Datadog renamed it. /product/llm-observability/ 301-redirects to /products/ai/agent-observability/. The docs path and pricing anchor still use the old name, which is why both appear in search.
What counts as a billable LLM span?
A single call to an LLM provider. Of the seven span kinds, only LLM bills — workflow, agent, tool, task, embedding and retrieval are free.
Does Datadog Agent Observability require an APM subscription?
No — Datadog states it is standalone. It also does not draw on APM allotments, so every span above your tier is net-new spend. Whether LLM traces are separately billed as APM traces is undocumented; ask your account team.
Are AI guardrails included in the price?
No. Agent Observability detects prompt injection as an evaluation. Real-time blocking is AI Guard, a separate product billed per million input tokens evaluated including full conversation context, with no published rate.
How do I cut the bill without losing visibility?
Use DD_LLMOBS_SAMPLE_RATE — sampling does not affect Agent Observability metrics, which stay based on 100% of instrumented traffic.









.png)
.png)

.png)
.png)
.png)
.png)
.png)
.png)
.png)


.webp)
.webp)






