Blank white background with no objects or features visible.

Lernen Sie TrueForge kennen: Das Open-Source- und herstellerneutrale Agent Harness. 50 % geringere Kosten. Jetzt entdecken→

From Token Spend to Workflow Economics: What BCG’s Return-on-AI Framework Requires

von Boyu Wang

Published: September 10, 2026

A token bill can tell you what inference cost. It cannot tell you whether the work was accepted, whether a human had to redo it, or whether the business outcome arrived two weeks later. Managing AI economics requires a new unit of control: the workflow run, joined to its eventual outcome.

Source and Scope Note
Source and scope. This article is TrueFoundry’s technical interpretation of BCG’s July 1, 2026 article, “Return on AI: How CFOs and CIOs Can Manage the Token Meter.” BCG and its authors do not endorse this article or TrueFoundry. Product mappings and implementation recommendations are ours. BCG’s capex, opex, and COGS framing is a management lens, not accounting advice; actual treatment depends on applicable standards and company policy.

BCG makes a useful move: stop treating the token as the unit of management and start measuring the cost of a successful outcome. Its return-on-AI framing puts economic return over the combined cost of human intelligence and tokens. It also argues for a workflow-level operating model that can see spend, shape cost, and prove value—or stop the activity.

The idea is sound. The hard part is implementation.

Model gateways observe requests. Agent runtimes observe turns, tools, retries, and pauses. Finance sees invoices. Business systems observe whether a case closed, a payment settled, a pull request shipped, or an escalation reopened. None of those records, alone, is return on AI.

Missing Infrastructure Callout

The missing infrastructure is not another token dashboard. It is an attribution contract that joins AI activity to a settled business outcome.

1. The token is a billing unit, not a business unit

Tokens remain essential telemetry. They explain input size, output size, cache behavior, and part of the cost of a model call. But they do not have stable business meaning across models or providers. Tokenizers differ. Input, output, cached input, cache creation, reasoning, audio, and other modalities can carry different billing rules. A million units in one product is not automatically equivalent to a million units in another.

Even within one pricing category, where billed cost may scale linearly with token volume, workflow economics need not. A larger context can change model behavior. A cheaper model can trigger more retries. A concise answer can increase human review. A cache hit can remove an inference call but return an answer that is stale for the user’s current state. An agent can enter a tool loop that spends little per call and a great deal per resolved case.

BCG identifies four interacting cost forces: adoption breadth and depth, task intensity, context and loops, and model mix. That framing matters because each force exists above the individual request. A single request is often too small to show the work; a monthly provider invoice is too large to assign responsibility or value.

A diagram showing a model request flowing into a workflow run that includes model, tool, human, and failure costs, then into a validated outcome
Figure 1. The workflow run is the smallest practical unit that can carry both a complete cost envelope and a reference to an authoritative outcome.

2. Define two contracts before building the dashboard

A credible workflow-economics system begins with two explicit contracts: a cost envelope and an outcome contract. The first says which resources count. The second says what success means. If either remains implicit, cost-per-outcome becomes a number with an unstable denominator or a negotiable numerator.

The cost envelope

For each workflow run, capture the costs that materially change the decision. That usually includes model inference, tool and infrastructure usage, human oversight, and failure or rework. The categories should be mutually exclusive or governed by an allocation rule so the same correction time is not counted twice. The goal is not false precision. Human time, for example, may be sampled by workflow and role instead of metered for every click. The requirement is consistency: use the same inclusion rules when comparing policies.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Melde dich an
Inhaltsverzeichniss

Steuern, implementieren und verfolgen Sie KI in Ihrer eigenen Infrastruktur

Buchen Sie eine 30-minütige Fahrt mit unserem KI-Experte

Eine Demo buchen

Der schnellste Weg, deine KI zu entwickeln, zu steuern und zu skalieren

Demo buchen
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Entdecke mehr

Keine Artikel gefunden.
September 10, 2026
|
Lesedauer: 5 Minuten

TrueFoundry Integration mit Smallest AI

Keine Artikel gefunden.
September 10, 2026
|
Lesedauer: 5 Minuten

Trojai-Integration mit TrueFoundry

Keine Artikel gefunden.
September 10, 2026
|
Lesedauer: 5 Minuten

Middleware integration with TrueFoundry AI Gateway

LLM-Werkzeuge
Technik und Produkt
LLM-Terminologie
September 10, 2026
|
Lesedauer: 5 Minuten

Gemini 3.5 Flash ist beeindruckend. Das haben wir tatsächlich herausgefunden.

LLMs und GenAI
Keine Artikel gefunden.

Aktuelle Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Machen Sie eine kurze Produkttour
Produkttour starten
Produkttour