LLM as a Judge, Running Inline as a Gateway Guardrail
.png)
Conçu pour la vitesse : latence d'environ 10 ms, même en cas de charge
Une méthode incroyablement rapide pour créer, suivre et déployer vos modèles !
- Gère plus de 350 RPS sur un seul processeur virtuel, aucun réglage n'est nécessaire
- Prêt pour la production avec un support complet pour les entreprises
What LLM-as-a-judge actually is
A judge is a model given three things: an output to assess, a rubric describing what “good” means, and an instruction to return a machine-readable verdict — pass/fail, a score, or a preference between two candidates.
It exists because the alternative does not scale. Exact-match metrics work when there is one right answer; they are useless for a support reply where five answers are all correct. Humans grade those well and slowly. A judge sits in between: worse than a careful human, far better than a string comparison, fast enough for thousands of samples. Treat it as a cheap, noisy, human-correlated proxy.
Two modes that people constantly conflate
The most common mistake here is treating offline evaluation and online gating as one thing with two deployment targets.
The failure row is the one to internalise. Offline, a wrong score costs a decision you revisit next sprint; online, it is a user staring at an error. So tune an online judge for precision even at the cost of recall. Copying a rubric across modes is how teams end up blocking a few percent of good traffic.
Where judges work, and where they fall over
Judges are good at criteria that are describable but not computable — is this grounded in the retrieved context, is the tone right for a clinical setting. You can write the rule in a paragraph but not as a regex. They fall over in three ways.
Position bias. Given two candidates, judges favour one slot; reverse the order and a meaningful share of verdicts flip. The mitigation is to judge each pair twice with positions swapped and count only agreement, which doubles your cost.
Verbosity bias. Longer answers score higher, roughly independent of whether the extra words added anything. This quietly pushes production prompts toward padding: you optimise against the judge, the judge rewards length, the token bill rises, nobody connects the two.
Self-preference. A model rates text resembling its own generations more highly. If GPT judges GPT output the score is flattered — the strongest argument for a judge from a different family than the model under test.
Judges are also badly calibrated on continuous scales: ask for a score out of 10 and everything clusters at 7 and 8. Binary pass/fail against a tight criterion is far more reliable. And a judge cannot verify a fact it does not have — with no reference material you get its priors, not a fact check.
Choosing a judge, and paying for it
Different family from the model under test, because self-preference is real and free to avoid. Smaller than you think, because judging is classification against a long prompt, not reasoning. Pinned, not floating, because a provider alias that rolls forward moves your quality bar without a code change. And a lower p99 than the thing it guards, or the judge becomes your tail latency.
Then the arithmetic. An inline judge is a second inference; check both prompt and response and it is two. That is not overhead you can engineer away, so the question is which slice of traffic:
- Sample it. Judge 5% for quality telemetry; 100% only for safety checks that must never miss.
- Split the rubric. Cheap deterministic checks on everything; the judge only where those pass and the surface is sensitive.
- Gate on input, sample on output. Input rails can run in parallel with the model call, so they cost tokens but little wall-clock time. Output rails cannot.
When a cheaper classifier wins
If you can express the rule as code, express it as code.
Where teams get this wrong
They ship the offline rubric into the request path. A ten-dimension rubric that works in a nightly run becomes a slow, high-variance blocker in production. Online rubrics should be one criterion, binary, worded tightly enough for a small model to answer consistently.
They never measure the judge. It is a model in production with no evaluation of its own. Hand-label a few hundred real examples and read precision and recall separately. If you cannot state the false-positive rate, you cannot state what fraction of users the guardrail is failing.
They roll straight to blocking. A judge that is 90% accurate blocks one in ten good answers the day you enforce it. Audit first, read a week of traces, then promote.
They forget streamed responses. Stream tokens to the browser and an output judge has nothing to judge until the stream ends. Most gateways, TrueFoundry included, skip output guardrails on streaming requests.
How this works in TrueFoundry
TrueFoundry has no eval product. No eval page, no datasets, no experiment tracking, no scoring UI. Offline evaluation belongs to platforms built for it, and TrueFoundry’s role is exporting traces over OpenTelemetry — traces, not OTEL metrics, from that screen. For a scored dataset and a leaderboard, use Braintrust, Arize or an equivalent and point the exporter at it.
What TrueFoundry implements is the online half: a judge running as a guardrail in the request path, its verdict, latency and token spend on the same trace as the request it guarded.

LLM guardrail execution flow: input mutation, parallel input validation, the model call, then output mutation and validation
Each judge below is a small service you deploy, called by the gateway as a rail through the custom-guardrail contract. That contract has one rule worth memorising: HTTP 2xx means the guardrail ran to completion, and the policy decision lives only in the JSON body. The docs warn that a “blocked” signalled as HTTP 400 may be read as a runtime error and let through. Return 2xx with verdict: false.
NVIDIA NeMo: the JUDGE_MODEL path
NeMo runs as a FastAPI wrapper deployed on TrueFoundry — no native NeMo SDK calls from the gateway, no client SDK changes in your app. Inside, it uses a small judge LLM plus Colang, a domain-specific language, to evaluate prompts and responses through the self_check_input and self_check_output rails. The rubric lives in the Colang bundle, where config/prompts.yml holds few-shot examples that in v1 catch DAN-style role-play, “ignore previous instructions”, system-prompt extraction and policy-bypass markers.
The detail that matters operationally: the judge LLM calls back through the same TrueFoundry gateway. The wrapper’s .env carries TFY_BASE_URL, TFY_API_KEY and JUDGE_MODEL — the doc’s example is openai-main/gpt-4o-mini. Because the judge’s inference is gateway traffic, its tokens, cost and latency land in the same dashboards as everything else, so you can see what guardrails cost. Swapping judges is one line: change JUDGE_MODEL (the doc shows openai-main/gpt-4o) and redeploy.

Registering the NeMo self-check rail as a custom guardrail config in TrueFoundry
Register two Custom Guardrail Configs, both Validate — selectors nemo-self-check/nemo-self-check-input and .../nemo-self-check-output — with Enforce But Ignore On Error recommended. The verdict contract is tiny: HTTP 200 plus {"verdict": bool, "message": Optional[str]}; a genuine failure arrives as 5xx.
The gateway dispatches the input rail call and the model call in parallel, so the judge does not add to time-to-first-token; on a block it cancels the in-flight model call. The output rail runs sequentially, because it must.
The docs are candid about the price. Under known limitations, in their own words: every guarded request adds one or two LLM calls, one per direction. They also note there are no streaming-aware guardrails, because the contract is buffered, and that in-memory state is per-replica.
Patronus: a managed judge with a fixed criteria list
Patronus is the closest thing to a turnkey judge in the catalogue, and it is Validate-only — operation is the literal "validate", and the docs state it cannot mutate.

Patronus guardrail configuration in the TrueFoundry AI Gateway
Config type is integration/guardrail-config/patronus. Set target to request or response, then list evaluators; six types exist — answer-relevance, glider, judge, pii, phi, toxicity. The judge type is the LLM-as-a-judge one, and its criteria field is required. All fifteen documented values:
TrueFoundry AI Gateway offre une latence d'environ 3 à 4 ms, gère plus de 350 RPS sur 1 processeur virtuel, évolue horizontalement facilement et est prête pour la production, tandis que LiteLM souffre d'une latence élevée, peine à dépasser un RPS modéré, ne dispose pas d'une mise à l'échelle intégrée et convient parfaitement aux charges de travail légères ou aux prototypes.



Gouvernez, déployez et suivez l'IA dans votre propre infrastructure
Blogs récents
Questions fréquemment posées
What is LLM as a judge?
Using one language model to score another’s output against a written rubric, returning a machine-readable verdict — pass/fail, a score, or a preference between candidates. It exists because open-ended text has no computable correctness metric, and it is best treated as a cheap, noisy proxy for a human rater.
Is LLM-as-a-judge accurate enough to gate production traffic?
For narrow, binary, tightly worded criteria, often yes; for broad quality rubrics, usually not. Measure precision and recall separately on hand-labelled examples first — a judge that is 90% accurate blocks roughly one in ten legitimate answers the day you enforce it.
How much does an inline judge cost?
At least one extra model call per guarded request, and two if you judge both prompt and response — TrueFoundry’s NeMo docs state that arithmetic explicitly. The only lever is coverage: sample for telemetry, reserve full coverage for the checks that must never miss.
Can I deploy TrueFoundry in my own VPC or on-prem?
Yes. TrueFoundry runs in your VPC, on-prem, air-gapped, or hybrid, so prompts and responses never leave your domain even as you route across many providers.
How much latency does the gateway add?
Roughly 3-4 ms, sustaining 350+ RPS on a single vCPU across 1,000+ LLMs. Guardrail time is separate and reported at P50, P75, P90 and P99.
Does it integrate with my existing observability and eval stack?
Yes. The gateway is OpenTelemetry-compliant and exports traces to external backends, which is how gateway data reaches platforms such as Braintrust, Langfuse or Arize. TrueFoundry does not itself run offline scoring jobs.









.png)
.png)

.png)
.png)
.png)
.png)
.png)
.png)
.png)


.webp)
.webp)






