Blank white background with no objects or features visible.

We’re sharing complimentary access to the full Gartner Hype Cycle for AI Governance 2026. Get your copy →

Gateway Tracing and Request Logs: Debug Every LLM Call

By Ashish Dubey

Published: October 9, 2026

Why you need tracing on the gateway

When an LLM call is slow, expensive, or wrong, you need to see what actually happened. Did a guardrail redact the prompt? How long did the model take versus the network? What did the gateway send to the provider? Request logging and tracing on the AI Gateway answer these questions for every call.

Because every request already flows through the gateway, you get this visibility without instrumenting each application by hand.

Controlling what gets logged

The request logging docs let you control which requests are logged. The simplest control is per request, using the X-TFY-LOGGING-CONFIG header with a stringified JSON value. Set enabled to true to log the request or false to skip it.

One gotcha worth noting: the value of that header is stringified JSON, not a raw JSON object, because most SDKs serialize headers as strings. The older global logging modes are being replaced by rule based logging configuration, which lets you control logging per subject, model, or metadata and redact sensitive values.

Viewing request logs

To see logged requests in the UI, go to AI Gateway, then Monitor, then Requests. This gives you the list of all logged calls, which you can open to inspect a single request in detail.

Request logs in the Monitor section of the AI Gateway.

Traces and spans, explained

The tracing docs describe a trace as the complete lifecycle of a request as it flows through the services that interact with the LLM. Usually a trace corresponds to a single API call of an application.

A span is an individual unit of work within that trace, such as a function call, an HTTP request, or a model inference. A trace is a tree of spans with parent and child relationships, where a child span is usually caused by its parent. TrueFoundry provides an OpenTelemetry collector backend that stores these traces and a UI to query and analyze them, and it can receive traces from any OpenTelemetry compatible SDK.

A trace is a tree of spans with parent and child relationships.
Can you explain why your last slow request was slow?
We will open a live trace on your traffic and walk the span tree with you.

Reading a single trace

The trace inspection docs walk through a real example: a chat completion with a PII redaction guardrail, captured as five spans that form a hierarchy.

  • ChatCompletion span (root) represents the full request lifecycle from the client perspective. In the example it lasts about 7 seconds and carries token metrics, cost, input, and output.
  • Guardrail span is a child of the root and represents the PII redaction processing, lasting under half a second.
  • Guardrail network call span is the actual HTTP call to the guardrail service, with the HTTP method and status code.
  • Model span is a sibling of the guardrail span and represents the model inference, lasting most of the request time.
  • Model network call span is the actual HTTP call to the provider, for example a POST to the provider chat completions endpoint.

In that same example the input is visibly redacted from a name to a placeholder, which shows the PII guardrail working end to end inside the trace.

A chat completion trace showing the guardrail, model, and outbound network spans.

From a single trace to fleet-wide metrics

Individual traces are for debugging. For trends, the analytics dashboard aggregates the same data. Tabs cover Overview, Model Metrics, MCP Metrics, Guardrail Metrics, Routing Metrics, and Cache Metrics.

Model metrics can be grouped by models, virtual models, users, virtual accounts, teams, or metadata, and the latency views break down request latency, time to first token, inter token latency, and time per output token. Together with per request traces, this gives you both the microscope and the dashboard.

Model metrics in the analytics dashboard, groupable by team, user, model, and more.
Give every team tracing without the instrumentation work
See request logs, span level traces, and analytics on your own gateway traffic.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Start free
Table of Contents

One Gateway for Every LLM, Agent and MCP Server

Book a 30-min with our AI expert

Book a Demo

The fastest way to build, govern and scale your AI

Book Demo
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Discover More

No items found.
October 9, 2026
|
5 min read

Budgets and Quotas on the AI Gateway: Control AI Spend by Team

No items found.
October 9, 2026
|
5 min read

Gateway Tracing and Request Logs: Debug Every LLM Call

No items found.
October 9, 2026
|
5 min read

How to Configure Guardrails on the AI Gateway

No items found.
October 9, 2026
|
5 min read

OpenRouter BYOK explained: cheaper, often faster, and changing

No items found.
No items found.

Recent Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Frequently asked questions

How do I turn logging on or off for a request?

Send the X-TFY-LOGGING-CONFIG header with a stringified JSON value and set enabled to true or false. Remember that the value is stringified JSON, not a raw object, because most SDKs serialize headers as strings.

Where do I view logs and traces?

Logged requests appear under AI Gateway, then Monitor, then Requests. Opening a request shows its trace, which is the tree of spans covering guardrails, the model, and the outbound provider calls.

‍

What is the difference between a trace and a span?

A trace is the full lifecycle of a single request. A span is one unit of work inside that trace, such as a guardrail check or a model inference. Spans form a parent and child tree within the trace

Can I use my own OpenTelemetry SDK?

Yes. TrueFoundry runs an OpenTelemetry collector backend and can receive traces from any OpenTelemetry compatible SDK. For LLM use cases the docs recommend an LLM focused SDK that captures model specific traces and metrics.

Take a quick product tour
Start Product Tour
Product Tour