How to Use Claude Managed Agents: A Step-by-Step Setup Guide

Diseñado para la velocidad: ~ 10 ms de latencia, incluso bajo carga
¡Una forma increíblemente rápida de crear, rastrear e implementar sus modelos!
- Gestiona más de 350 RPS en solo 1 vCPU, sin necesidad de ajustes
- Listo para la producción con soporte empresarial completo
Building an AI agent involves more than making a model call. You need to manage the agent loop, execute tools, provide a runtime environment, maintain session state, handle permissions, and observe what the agent is doing.
Claude Managed Agents takes care of much of this infrastructure for you. Instead of building and operating your own agent runtime, you define the agent, configure where it runs, and create sessions for it to execute tasks. The managed runtime handles the model calls and tool execution while your application interacts with the agent through events.
In this guide, we'll show you how to use Claude Managed Agents step by step from creating your first agent and configuring its environment to running a session, adding tools and MCP servers, and handling agent events.
If you want the conceptual grounding first, see our guide to what Claude Managed Agents are and how the architecture works. This guide assumes you’re ready to build.
Prerequisites
Before creating a Managed Agent, you'll need:
- A Claude Console account
- An Anthropic API key
- The Anthropic SDK
- Python, TypeScript, or another supported language
Step 1: Install the CLI and SDK
For Python, install the SDK:
pip install anthropicThen set your API key as an environment variable:
export ANTHROPIC_API_KEY="your-api-key"Anthropic also provides an ant CLI for managing Managed Agent resources.
The SDK reads the key from your environment, so you don't need to include it directly in your application code.
Enable the Managed Agents API
Claude Managed Agents is currently available through Anthropic's beta API. The Managed Agents beta is identified by the:
managed-agents-2026-04-01beta header.
When you use the official SDK, the SDK handles this header for you when you access the Managed Agents APIs.
Your basic client setup looks like:
from anthropic import Anthropic
client = Anthropic()At this point, your application is ready to work with the Managed Agents API.
The next step is to create the agent itself - defining which Claude model it uses, what instructions it follows, and which tools it can access.
Step 2: Define Your Agent
An agent is a Markdown file. The frontmatter contains the agent's configuration, while the body contains its system prompt.
Here's an example of an agent that drafts release notes from a repository:
---
name: Release Notes Agent
model: claude-opus-5-5
tools:
- type: agent_toolset_20260401
---
The agent_toolset_20260401 toolset gives the agent access to Anthropic's pre-built tools, including capabilities such as shell commands and file operations. You can also configure individual tools if you don't need the full toolset. See the tools reference for the available options.
Save the configuration as release-notes-agent.md, then apply it with the ant CLI:
ant apply release-notes-agent.md
This creates the agent and returns its ID.
The agent is a persistent resource. You create it once and then reference its ID whenever you start a session that uses it—you don't recreate the agent for every task.
Next, you'll create the environment where your agent sessions will execute.
Step 3: Create an Environment
The environment defines the sandbox where your agent sessions execute.
Create an environment.yaml file:
# yaml-language-server: $schema=https://platform.claude.com/schemas/ant/beta/environment.json
name: release-notes-env
config:
type: cloud
networking:
type: unrestrictedThen apply it:
ant apply environment.yaml
You can also apply the agent and environment together:
ant apply release-notes-agent.md environment.yamlThe networking: unrestricted setting gives the sandbox unrestricted network access. For agents that work with internal systems or sensitive data, use a more restrictive networking configuration.
You now have the two resources required to run an agent:
- Agent: defines what the agent is and what it can do.
- Environment: defines where the agent runs.
Next, you'll create a session that connects them.
Step 4: Start a Session
A session is one running instance of your agent working on a specific task.
Create a session by passing your agent and environment IDs:
session = client.beta.sessions.create(
agent=agent.id,
environment_id=environment.id,
title="Release notes for v2.4", )
print(f"Session ID: {session.id}")The session ID identifies this particular run. You'll use it to send messages, stream events, and continue interacting with the agent.
The same agent can have multiple sessions. For example, your Release Notes Agent could run a separate session for every release while keeping the same agent configuration.
Step 5: Send a Message and Stream the Response
With the session running, send the agent a task.
Open the event stream first, then send the user message:
with client.beta.sessions.events.stream(session.id) as stream:
client.beta.sessions.events.send(
session.id,
events=[
{
"type": "user.message",
"content": [
{
"type": "text",
"text": "Read the commits since v2.3 and draft release notes to NOTES.md",
},
],
},
],
)
for event in stream:
match event.type:
case "agent.message":
for block in event.content:
if block.type == "text":
print(block.text, end="")
case "agent.tool_use":
print(f"\n[Using tool: {event.name}]")
case "session.status_idle":
print("\n\nAgent finished.")
breakBehind this call, the managed runtime provisions the sandbox from your environment configuration, starts the agent loop, lets Claude select and execute tools, and streams events back to your application.
The three event types above are the ones you'll typically handle first:
agent.message— the agent sends a response.agent.tool_use— the agent uses a tool.session.status_idle— the agent has finished its current work.
For the complete list of events, including status changes and errors, see the events and streaming reference.
At this point, you have a working Managed Agent: you've defined it, given it an execution environment, started a session, and sent it a task.
Next, we'll look at how to steer a running agent and handle long-running sessions.
Step 6: Steer a Running Agent
You don't have to wait for an agent to finish before changing its direction.
You can send a user.interrupt event to stop the current turn, followed by a user.message with the new instruction:
client.beta.sessions.events.send(
session.id,
events=[
{"type": "user.interrupt"},
{
"type": "user.message",
"content": [
{
"type": "text",
"text": "Instead, focus on fixing the bug in line 42.",
},
],
},
],
)The interrupt stops the model's current response and returns control to the session. If the agent is in the middle of a tool call, the interrupt can take longer to apply while that tool finishes. Once the interrupt is processed, the session returns to idle, and the following user.message starts the next turn.
This is useful for long-running tasks where you realize the agent is heading in the wrong direction. Instead of terminating the session and starting over, you can reuse the same session and redirect it with additional context.
You can also interrupt a running session without immediately sending a new instruction:
client.beta.sessions.events.send(
session.id,
events=[
{"type": "user.interrupt"},
],
)The session stays available after the interruption, so you can inspect what happened and decide what to send next.
Next, you'll want to handle what happens when your application loses its connection to a running session.
Step 7: Handle Disconnects and Resume Sessions
A streaming connection can disconnect while an agent is still working. That doesn't mean the session or its work is lost.
Managed Agents persists the session's event history server-side, so your application can reconnect to the same session instead of starting the task again.
The important thing is to persist the session ID when you create the session:
session = client.beta.sessions.create(
agent=agent.id,
environment_id=environment.id,
title="Release notes for v2.4",
)
session_id = session.idIf your application loses its connection, use the same session_id to reconnect and continue interacting with the session.
This matters for long-running agents. A browser tab closing, a process restarting, or a temporary network failure shouldn't force the agent to start the task from scratch.
For production applications, treat the session ID as durable application state: store it somewhere your application can retrieve it after a process restart, rather than keeping it only in memory.
The event stream is only the connection between your application and the running session. The session itself continues to exist independently of that connection.
Next, we'll look at how to give your agent additional capabilities with MCP servers and tools.
Running It on Your Own Infrastructure
The cloud sandbox is the default, and it's the reason some teams can't ship this. If data residency or compliance rules out Anthropic-hosted execution, self-hosted sandboxes let sessions run on infrastructure you control while Anthropic still operates the loop.
Note what does and doesn't move: execution comes to your infrastructure, orchestration stays with Anthropic.
Scheduling Recurring Runs
For agents that should run on a cadence rather than on demand, scheduled deployments run a session on a cron schedule. A nightly triage pass or a weekly report fits here without you standing up a scheduler.
Keeping the Bill Under Control
Claude Managed Agents bills on two dimensions: model tokens, plus a runtime charge of roughly $0.08 per session-hour that covers container time.
Three things worth knowing early:
Idle time is free, so a session waiting on a human doesn't accrue runtime charges. There's no batch discount, so work you'd normally batch costs the same here. And web search is billed separately from the toolset.
The session-hour charge is usually the smaller number. Tokens dominate, and how many tokens a run consumes depends on how the harness manages context rather than on anything you configure. We broke the full cost structure down in Claude Managed Agents pricing.
Where Teams Hit Limits
Two constraints tend to surface after the prototype works. The model is Anthropic's, so there's no switching to a cheaper one when your volume grows. And the loop is operated by Anthropic, so context strategy is not a dial you control.
If you're evaluating rather than committed, Claude Managed Agents alternatives covers the field.
TrueForge: An Open-Source Alternative to Claude Managed Agents, Up to 75% Cheaper

TrueForge is TrueFoundry's open-source, vendor-neutral agent harness for teams that want more control over how their agents run in production. Instead of tying the runtime to a single model provider, TrueForge lets you bring your own models, MCP servers, and infrastructure while handling the agent loop, tool execution, context management, approvals, and sandboxing.
Where a self-hosted agent like Hermes leaves much of the production infrastructure to the developer, TrueForge provides the surrounding control plane needed to operate agents across teams and environments. This includes centralized MCP access and credentials, human-in-the-loop approvals, sandboxed execution, observability, and governance.
You can run it locally with npx @truefoundry/trueforge or deploy it for a team with Docker Compose or Helm.
TrueForge also separates the agent runtime from the model layer. This means teams can switch models or route workloads to different providers without rebuilding the agent itself.Combined with TrueFoundry's AI Gateway, teams can also apply model routing and cost controls to avoid using expensive frontier models for tasks that don't require them. The result is a middle ground between a fully managed runtime and building the entire agent infrastructure yourself: an open-source agent harness with the operational controls needed to run production agents at scale.
In TrueFoundry's benchmark against Claude Managed Agents, TrueForge reached roughly the same task accuracy while using about 40% as many tokens when both were run on Opus 4.8, making it about 30% cheaper per run. When the same benchmark was run with GLM-5.2 on TrueForge, it achieved the same ~11/14 task score at about $2.90 per run versus $11.80 for Claude Managed Agents, roughly 75% lower cost.
TrueFoundry also provides a second lever for reducing model costs through AI Gateway's Auto Routing. Instead of sending every request to a frontier model, Auto Routing classifies requests by complexity and routes them to different model tiers. In TrueFoundry's benchmark across 550 prompts, routing between Haiku, Sonnet, and Opus reduced cost by 69% while retaining 98% of baseline quality, and mean latency dropped from 7.6 seconds to 4.0 seconds. On production-shaped traffic, the overall cost reduction reached 80%.
These two layers address different parts of the cost equation. TrueForge reduces the overhead of running the agent itself, while model routing lets you avoid using an expensive model when the task doesn't require one. For teams running agents at scale, having both controls can matter more than the price of the model alone.
TrueForge is open source and MIT licensed, while TrueFoundry's AI Gateway adds the operational layer teams need when moving from individual agents to production: centralized model access, budgets, guardrails, credential management, and unified traces.
If you're comparing Claude Managed Agents with a more configurable architecture, TrueFoundry is worth considering when model flexibility, infrastructure control, and agent cost optimization matter alongside the agent runtime itself.
The Layer Managed Agents Doesn't Cover
Managed Agents governs one agent's runtime very well. What it doesn't do is govern a fleet which is the problem that shows up around agent number five, when finance asks what any of this costs and security asks who approved the credentials.
An agent harness controls how an agent works. An AI Gateway controls how that agent accesses models.
For example, an organization might run agents using Claude Managed Agents or TrueForge while routing model requests through an AI Gateway.
TrueFoundry's AI Gateway sits in front of every model and MCP call an agent makes, applying centralized model access, per-team budgets, rate limits, guardrails, credential management and unified traces - across whichever agents and frameworks your teams are running. It supports 1,000+ LLMs through a single OpenAI-compatible API, adds roughly 3–4 ms of overhead, and handles 350+ RPS on a single vCPU, so it can sit in the hot path of a long-running agent without becoming the bottleneck. It runs in your own VPC, on-prem or air-gapped.
This becomes particularly useful when an organization runs agents across multiple models rather than relying on a single model provider.
Agent runtime: manages the agent loop, tools, state, and execution.
AI Gateway: manages model access, routing, policies, observability, and provider abstraction.
This separation lets teams change or add models without having to redesign the agent itself.
FAQ
Q: How do you use Claude Managed Agents?
A: Install the ant CLI and the Anthropic SDK, define an agent in a markdown file with its model, system prompt and tools, declare an environment in YAML for the sandbox, then create a session referencing both and stream its events. Agents and environments are created once and reused; sessions are per task. The whole path takes about ten minutes.
Q: Do you need the ant CLI, or can you do everything from the SDK?
A: The CLI is convenient rather than required. Anthropic documents cURL and native SDK paths for every step in the quickstart, across Python, TypeScript, Java, Go, C#, Ruby and PHP. What the CLI adds is ant apply and claude-lock.json, which keep agent and environment IDs versioned alongside your code instead of pasted into a config somewhere.
Q: Can Claude Managed Agents run on your own infrastructure?
A: Partly. Self-hosted sandboxes let sessions execute on infrastructure you control, which is usually enough for data-residency requirements. The orchestration layer stays with Anthropic either way, so this is not the same as running the whole stack yourself.
Q: Does it integrate with my existing observability stack?
A: TrueFoundry's gateway layer is OpenTelemetry-compliant and plugs into Grafana, Datadog, Prometheus or whatever you already run, tracing each request from prompt through to tool and model execution. That matters once agents from different frameworks need to show up in one place.
Q: How do I govern models and MCP servers across a lot of agents?
A: Through a gateway, once per-team API keys stop scaling. TrueFoundry's AI Gateway puts 1,000+ LLMs behind one OpenAI-compatible API at roughly 3 to 4 ms of added latency and 350+ RPS on a single vCPU, with RBAC, budgets, guardrails and credential rotation. Agents on Claude Managed Agents, LangGraph or anything else get governed the same way.
Related reading
- What Are Claude Managed Agents?: the architecture behind the steps above
- Claude Managed Agents Pricing: the full cost breakdown with a worked example
- Claude Agent SDK vs Claude Managed Agents: which one to run in production
- Claude Managed Agents Alternatives: the field, if you're still evaluating
- Agent Harness Best Practices: ten rules that apply whichever harness you pick
Conclusion
Knowing how to use Claude Managed Agents is mostly knowing four objects and the order you create them in: agent, environment, session, events. The setup is genuinely quick. The parts worth spending time on are the ones the quickstart skims, namely disconnect handling, network configuration on your environment, and understanding which half of your bill you can actually influence.
If you reach the point where model choice and context strategy start mattering more than setup speed, that's the signal to look at an open harness. npx @truefoundry/trueforge runs one locally in about a minute, or book a walkthrough if you'd rather see it against your workload.
TrueFoundry AI Gateway ofrece una latencia de entre 3 y 4 ms, gestiona más de 350 RPS en una vCPU, se escala horizontalmente con facilidad y está listo para la producción, mientras que LitellM presenta una latencia alta, tiene dificultades para superar un RPS moderado, carece de escalado integrado y es ideal para cargas de trabajo ligeras o de prototipos.












.png)



.png)
.png)
.png)

.png)
.png)
.png)
.png)
.png)
.png)
.png)
.png)





