tfy-llm-gateway Helm chart ships an optional Vector aggregator that sits between the gateway and the control plane. The gateway sends traces to Vector inside the cluster. Vector writes them to a disk buffer on a PersistentVolume and forwards them to the control plane, retrying until delivery succeeds. You can also add extra sinks so the same traces go to collectors such as Pydantic Logfire or your own OpenTelemetry Collector.
How it works
When you setvector.enabled: true, the chart:
- Deploys Vector as a StatefulSet named
tfy-vector(3 replicas by default), each replica with its own PersistentVolumeClaim for the disk buffer. - Renders the Vector configuration into the
tfy-llm-gateway-vector-configConfigMap. It defines anotel_inOTLP source (gRPC4317, HTTP4318) and atfy_otel_tracessink that forwards to<global.controlPlaneURL>/api/otel/v1/traces. - Points the gateway at Vector by setting
TFY_OTEL_EXPORTER_OTLP_TRACES_ENDPOINTtohttp://tfy-vector.<namespace>.svc.cluster.local:4318/v1/traces(https://whenvector.mtls.enabledis true). If you already setenv.TFY_OTEL_EXPORTER_OTLP_TRACES_ENDPOINTyourself, the chart leaves it unchanged. - Authenticates to the control plane with
TFY_API_KEYfrom thetruefoundry-credssecret that you created during gateway plane installation.
Prerequisites
global.controlPlaneURLis set in your values file. The chart fails to render without it when Vector is enabled.- The
truefoundry-credssecret with theTFY_API_KEYkey exists in the gateway namespace. - A StorageClass that can provision
ReadWriteOncevolumes. With the defaults, Vector requests 3 PVCs of22Gieach. - The gateway and Vector run in the release namespace. Vector cannot be combined with a
global.namespaceOverridethat differs from the release namespace. - Nodes can pull
tfy.jfrog.io/tfy-mirror/timberio/vectorusing thetruefoundry-image-pull-secretimage pull secret (override withvector.image.pullSecretsif needed).
Enable Vector
Turn on Vector in your values file
truefoundry-values.yaml:Upgrade the release
Verify the deployment
http://tfy-vector.truefoundry.svc.cluster.local:4318/v1/traces.Send a request through the gateway and check the Vector logs. A successful delivery to the control plane logs HTTP response. status=200 for the tfy_otel_traces sink:Configuration reference
All keys live undervector in the tfy-llm-gateway values file. The defaults suit most deployments.
tfy-llm-gateway chart values and the upstream Vector Helm chart. Sink-level options such as batching, retries, and TLS are documented in the OpenTelemetry sink reference and the Vector buffering model.
Export traces to other OTLP collectors
Besides the control plane, Vector can send the same traces to any OTLP-compatible backend, such as Pydantic Logfire, Grafana, Datadog, Honeycomb, or your own OpenTelemetry Collector. These exports go directly from your gateway plane to the destination and do not pass through the TrueFoundry control plane.Add collectors with vector.extraSinks
Each entry under vector.extraSinks is one Vector sink. The key is the sink ID and the value is the sink configuration. The chart merges these entries into the Vector configuration next to tfy_otel_traces, so you can add as many collectors as you need.
- Sink IDs must be unique.
tfy_otel_tracesandprom_exporterare reserved for the sinks the chart renders, and the chart fails to render if you reuse them. - Every entry needs a
type. For OTLP backends, usetype: opentelemetrywithprotocol.type: httpand the destination’s/v1/tracesendpoint. Vector’s OpenTelemetry sink does not send OTLP over gRPC. - If an entry does not set
inputsorbuffer, the chart fills them fromvector.extraSinkDefaults:inputs: [otel_in.traces](the gateway traces) and an in-memory buffer of 10,000 events that drops new spans when full. This keeps a slow or unreachable collector from stalling delivery to the control plane and to your other collectors. See Sink options to change it. - Vector reloads its configuration in place about every 30 seconds, so adding, changing, or removing an entry does not restart the pods.
Store collector credentials
Do not write tokens or API keys into your values file. Store them in a Kubernetes Secret and reference them from the sink. Other values in TrueFoundry charts reference secrets as${k8s-secret/<secret>/<key>}. That syntax does not work inside a Vector sink, because Vector reads its configuration directly and does not interpolate environment variables. Vector sinks use Vector’s SECRET[...] references instead. The chart uses the same mechanism for its own control plane token.
The chart sets up everything except the Secret itself:
- You create a Kubernetes Secret named
tfy-vector-extra-sink-credentialsin the gateway namespace, with one key per credential. - The chart mounts that Secret into every Vector pod at
/etc/vector-extra-sink-credentials, where each key becomes a file. The mount is optional, so Vector starts normally before the Secret exists. - The chart declares a Vector secret backend named
extra_sink_credentialsthat reads that directory. - In a sink, you write
SECRET[extra_sink_credentials.<KEY>]. Vector replaces it with the value of that key when it loads the configuration. The value never appears in your values file, the Helm release, or the ConfigMap.
"Bearer SECRET[extra_sink_credentials.GRAFANA_TOKEN]".
- kubectl
- Manifest
Rotate a credential
Rotate a credential
Use an existing Secret with a different name
Use an existing Secret with a different name
extra-sink-credentials volume at your Secret in vector.extraVolumes. Helm replaces this list as a whole, so keep the other default volumes:SECRET[extra_sink_credentials.<KEY>], where <KEY> is a key in your Secret.What happens when a key is missing
What happens when a key is missing
Error while retrieving secret from backend "extra_sink_credentials". A pod that is starting fails to start. A running pod keeps its previous configuration until you add the key or fix the reference.Example: Logfire and an in-cluster collector
This example keeps the control plane sink and adds two collectors:logfiresends traces to Pydantic Logfire and authenticates with a write token from thetfy-vector-extra-sink-credentialsSecret.otel_collectorsends traces to an OpenTelemetry Collector in the same cluster, without authentication.
Create the credentials Secret
https://logfire-us.pydantic.dev or https://logfire-eu.pydantic.dev. A token from one region is rejected by the other.Add the sinks to your values file
Upgrade the release
Verify delivery to each sink
status=200:
Gateway traces in Pydantic Logfire
Stop sending traces to the control plane
To send gateway traces only to your own collectors, setvector.forward.enabled: false. The chart then leaves out the tfy_otel_traces sink. At least one entry in vector.extraSinks is required, and the chart fails to render without one.
tfy_otel_traces uses the disk buffer by default. If none of your sinks uses a disk buffer, you can set vector.persistence.enabled: false to skip creating the PersistentVolumeClaims.
How sinks share Vector
Every sink reads the same traces from theotel_in source and keeps its own copy of each span.
- Each sink has its own buffer, retries, and connection settings. A collector that is slow, unreachable, or rejecting requests fills only its own buffer. The control plane sink and your other collectors keep delivering.
- Disk buffers are separate but share one volume. Vector keeps each disk buffer in its own directory under
vector.dataDir, named after the sink ID, with its ownmax_size. All of them live on the same PersistentVolume, sovector.persistence.sizemust hold the sum of every disk buffer plus some headroom. - Buffers are per replica. With 3 replicas, each sink has 3 independent buffers, and the gateway’s spans are spread across the replicas.
- Sinks do not slow down the gateway. Vector responds to the gateway’s OTLP request without waiting for the sinks to deliver, so an unreachable collector does not add latency to trace export.
- Input. Every sink receives every gateway trace. The chart does not support Vector transforms, so you cannot filter or sample spans for one sink only.
- Vector deployment. The OTLP receiver (ports, mTLS, maximum request size), replicas, resources, scheduling, the PersistentVolume, and the Vector version.
- Backpressure. When a sink with
when_full: blockhas a full buffer, Vector stops reading from the source and every sink stalls. Extra sinks usedrop_newestby default for this reason.
Control plane sink compared with extra sinks
Options you can set per extra sink
For anopentelemetry sink, set inputs, buffer, healthcheck, and proxy at the top level. Everything else goes under protocol (URI, headers, retries, batching, TLS, and so on).
This example gives Logfire a 2 GiB disk buffer and a retry limit:
Sink options
Change the default buffer for extra sinks
Change the default buffer for extra sinks
- On a new install, increase
vector.persistence.sizeby the same amount. - On an existing install, either lower
vector.buffer.maxSizeBytesby the same amount, or grow the volumes as described in Resize the buffer disk.
vector.extraSinkDefaults.buffer. Avoid when_full: block on extra sinks: when that buffer fills up, Vector stops reading from the shared source, which also stalls the control plane sink. See Vector buffering model for details.Headers for other collectors
Headers for other collectors
tfy-vector-extra-sink-credentials Secret and reference it from request.headers:Resize the buffer disk
You cannot changevector.persistence.size on an existing install with helm upgrade alone:
- Vector’s PersistentVolumeClaims come from the StatefulSet’s
volumeClaimTemplates. Kubernetes does not allow that field to change on an existing StatefulSet, so the upgrade fails. - Reinstalling does not help either. The claims (
data-tfy-vector-0,data-tfy-vector-1, and so on) are kept when the StatefulSet or the release is deleted, and a new StatefulSet reuses them at their old size.
- Keep the default memory buffer for the new sink. It uses no disk, and you lose only the spans it holds when a pod restarts.
- Split the existing disk. For example, lower
vector.buffer.maxSizeBytesfrom 20 GiB to 14 GiB and give the new sink a 6 GiB disk buffer, all within the default22Givolume.
22Gi to 44Gi.
allowVolumeExpansion: true. Without that, Kubernetes rejects the resize. Check with kubectl get storageclass <STORAGE_CLASS> -o jsonpath='{.allowVolumeExpansion}' — it must print true. If it does not, ask your cluster administrator to enable volume expansion, or use one of the options above that need no resize.Check that the StorageClass allows expansion
true.Grow each claim
vector.replicas is not 3, adjust the loop to match.Delete the StatefulSet object and keep its pods
--cascade=orphan deletes only the StatefulSet object. The Vector pods and their buffers keep running.Upgrade the release with the new sizes
Monitor Vector
Vector exposes its internal metrics in Prometheus format on port9598 at /metrics. With the Prometheus Operator installed, the chart’s PodMonitor scrapes them automatically. Useful series:
tfy_otel_traces with your sink name, such as logfire, to watch an extra sink. See Monitoring Vector for the full metric list.
Troubleshooting
Vector pods stay in Pending
Vector pods stay in Pending
data-tfy-vector-0, data-tfy-vector-1, and so on) are bound with kubectl get pvc -n truefoundry | grep data-tfy-vector. If they are Pending, set vector.persistence.storageClassName to a StorageClass that exists in the cluster.If the PVCs are bound, check node taints and selectors. Vector uses vector.tolerations, vector.nodeSelector, and vector.affinity, and falls back to the matching global.* values when they are empty.The control plane sink logs status=401 or status=403
The control plane sink logs status=401 or status=403
TFY_API_KEY in the truefoundry-creds secret is missing, wrong, or expired. These responses are not retried, and the affected spans are dropped. Update the secret with a valid gateway token. Vector reads the secret from a mounted volume, so restart the pods after updating it: kubectl rollout restart statefulset tfy-vector -n truefoundry.An extra sink logs status=401
An extra sink logs status=401
${ENV_VAR} or ${k8s-secret/...} is sent as that literal text, because Vector does not interpolate either form. Replace it with SECRET[extra_sink_credentials.<KEY>] as described in Store collector credentials.If the reference is correct, the credential itself is wrong or expired, or it belongs to a different region or account than the endpoint. Update the key in the tfy-vector-extra-sink-credentials Secret and restart Vector.Vector fails to start with a secret error
Vector fails to start with a secret error
SECRET[...] reference does not match a key in the Secret, and logs Configuration error. error=Error while retrieving secret from backend "extra_sink_credentials". Check that the tfy-vector-extra-sink-credentials Secret exists in the gateway namespace and that the key name matches exactly (it is case-sensitive). If you override vector.extraVolumes, check that it still contains the extra-sink-credentials volume.The chart fails to render with a Vector error
The chart fails to render with a Vector error
global.controlPlaneURL is required when vector.enabled is true: setglobal.controlPlaneURL.vector.mtls.enabled requires global.mTLS.enabled: enableglobal.mTLSor turn offvector.mtls.vector extraVolumes truefoundry-mtls secretName ...orcustom-ca configMap.name ...: the names in yourvector.extraVolumesoverride do not matchglobal.mTLSorglobal.customCA. Use the names from the error message.vector.extraSinks.<id> collides with a sink the chart renders: rename the sink.tfy_otel_tracesandprom_exporterare reserved.vector.extraSinks.<id> must be a Vector sink config with a type: add atypeto that entry.vector.forward.enabled is false and vector.extraSinks is empty: add a sink or setvector.forward.enabled: true.
Further reading
- Vector OpenTelemetry sink: all options for OTLP exports, including TLS, batching, and retries.
- Vector OpenTelemetry source: the
otel_insource that receives gateway traces. - Vector sinks catalog: every sink type and its options.
- Vector secrets: the
directorybackend andSECRET[...]syntax. - Vector buffering model: memory and disk buffers and
when_fullbehavior. - Vector Helm chart: upstream values that the
tfy-llm-gatewaychart passes through.