Blank white background with no objects or features visible.

TrueFoundry kündigt die Übernahme von Seldon AI an und erweitert damit seine Control Plane für Enterprise-KI. Vollständigen Bericht lesen →

ETCLOVG: The Seven-Layer Agent Harness Taxonomy, Mapped to a Production Runtime

von Boyu Wang

Published: July 24, 2026

Agent reliability is a property of both the model and the system around it. A 2026 survey by researchers from Carnegie Mellon, Yale, Amazon, and other institutions proposes ETCLOVG — Execution, Tooling, Context, Lifecycle, Observability, Verification, and Governance — as a taxonomy for that surrounding system, or agent harness.

The authors map more than 170 open-source projects to the seven layers and identify both mature areas and operational gaps. This post summarizes the taxonomy, then maps it to TrueFoundry's Agent Harness — separating native runtime capabilities from adjacent patterns and integrations.

Key Takeaways

  • What ETCLOVG is: a 2026 survey taxonomy — Execution, Tooling, Context, Lifecycle, Observability, Verification, Governance — for analyzing the agent harness, the system around the model, mapped against more than 170 open-source projects.
  • Structural core vs. control plane: E·T·C·L make an agent run; O·V·G make it observable, evaluable, and safe — with the survey particularly elevating Observability and Governance as independent layers with their own tooling and ownership.
  • Cross-layer coupling is the operational argument: failures rarely respect layer boundaries, so diagnosis and repair are system questions — the reason a shared substrate for state, identity, telemetry, and policy matters.
  • What maps natively and what does not: TrueFoundry directly maps to documented capabilities across E, T, C, L, O, and much of G. Verification can be connected to the runtime's traces and metadata but should not be described as a native evaluator today.

Priyanka, a head of platform, had a reliability problem she couldn't name. Her teams ran five production agents, and when one misbehaved the post-mortems never agreed on what to fix. One blamed the model and lobbied for an upgrade. One rewrote the prompt. One added a retry. One wanted a sandbox. Each fix was plausible, each addressed a different layer, and none of them shared a vocabulary for saying which layer had actually failed — so the meetings circled. What Priyanka needed wasn't a better model or a louder opinion; it was a map: a way to point at an agent and say "this failure lives in lifecycle, not in the model" or "we have no verification layer at all." ETCLOVG is that map, and this post is the walk through it — first the taxonomy as its authors built it, then each layer against a production runtime, distinguishing native capabilities from adjacent workflows and integrations whether it calls them by these names or not.

1. Why a Harness Taxonomy Exists at All

The survey's central claim is deflationary toward models and inflationary toward infrastructure: the rapid deployment of LLM agents revealed that execution reliability depends less on the base model than on the harness that wraps it. The evidence is partly historical. The ReAct era of 2022–23 wrapped a single model loop in a while-loop, a prompt template, and a small tool-dispatch table; AutoGPT and BabyAGI then exposed the failure modes of that thin wrapper — execution runaway, context blowout, state loss, unmonitored side effects — as infrastructure problems rather than prompt problems. The years since have been the field slowly building the missing infrastructure and, only recently, naming it.

That naming is the survey's three-phase framing, which locates harness engineering precisely. Prompt engineering (2022–24): the lever is the input text, optimized for a single call. Context engineering (2025): the question shifts to "what should the model see at each step?" — retrieval, compaction, tool-result ranking, window saturation. Harness engineering (2026–): as models grow capable enough for long-running tasks, the focus expands to the whole infrastructure wrapper. Each phase subsumes the previous — harness includes context includes prompt — and they overlap rather than replace. ETCLOVG gives that third phase the architectural boundaries the first two never needed.

2. ETCLOVG, Layer by Layer

The taxonomy splits into a structural core and a control plane. E, T, C, and L are the structural pillars — they make an agent run. O, V, and G are the control plane — they make a running agent observable, evaluable, and safe. The survey's notable move is promoting Observability and Governance to independent layers rather than folding them into lifecycle hooks, on the grounds that in production each has its own tooling ecosystem and is owned by a different team. Here are the seven, paraphrased from the survey's definitions.

E — Execution environment. Where the agent's code and actions actually run, and what bounds them: managed sandboxes, microVMs, code-specialized runtimes, computer-use and browser environments, OS-level permission models. This is the blast-radius layer — what an agent can touch when it runs a command.

T — Tool interface and protocol. How external capabilities are described, discovered, and invoked: protocol standards (MCP, A2A), tool description and selection, tool-augmented training, session management. The model only knows tools through this layer.

C — Context and memory management. What the model can see across short-term, session, and persistent horizons: long-horizon context techniques and mitigations for context drift. The subject of the prior engineering phase, now one pillar among four.

L — Lifecycle and orchestration. The control flow that reads and writes state — from the single-agent inner loop to multi-agent patterns to full issue-to-pull-request pipelines. The loop itself lives here.

O — Observability and operations. Traces, costs, failures, reliability signals — tracing platforms, agent-specific ops tools, cost tracking, unified observability. The survey's open-source map finds this layer comparatively thin in public code.

V — Verification and evaluation. Turning tasks and traces into evaluation, failure attribution, and regression feedback: benchmark grounding, controlled execution, multi-level judgement, deployment-time evaluation loops. The difference between "it ran" and "it was right."

G — Governance and security. Constraints across model, system, and organizational sub-layers: permission models, lifecycle hooks, component hardening, declarative constitutions, audit infrastructure. Also thin in open source, the survey notes, and "more often live inside commercial platforms."

Der schnellste Weg, deine KI zu entwickeln, zu steuern und zu skalieren

Melde dich an
Inhaltsverzeichniss

Steuern, implementieren und verfolgen Sie KI in Ihrer eigenen Infrastruktur

Buchen Sie eine 30-minütige Fahrt mit unserem KI-Experte

Eine Demo buchen

Der schnellste Weg, deine KI zu entwickeln, zu steuern und zu skalieren

Demo buchen
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Entdecke mehr

Keine Artikel gefunden.
July 24, 2026
|
Lesedauer: 5 Minuten

ETCLOVG: The Seven-Layer Agent Harness Taxonomy, Mapped to a Production Runtime

Keine Artikel gefunden.
July 24, 2026
|
Lesedauer: 5 Minuten

Best AI Gateway for Secure Data Routing in 2026

Keine Artikel gefunden.
July 24, 2026
|
Lesedauer: 5 Minuten

Best MCP Gateway for Regulated Industries in 2026

Keine Artikel gefunden.
July 24, 2026
|
Lesedauer: 5 Minuten

LangChain vs LangGraph vs LangSmith: What's the Difference in 2026

Keine Artikel gefunden.
Keine Artikel gefunden.

Aktuelle Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Machen Sie eine kurze Produkttour
Produkttour starten
Produkttour