Skip to main content
These are the questions that come up most often in support tickets and onboarding calls. Each answer links to the page with the full setup.

Billing and plans

It depends on which credential reaches Anthropic, and that is decided by the model account on the gateway:
  • Model account has an Anthropic API key → every request is billed per token to that key. The user’s seat is not involved. This is the only option for Claude Desktop / Cowork, and one of two options for Claude Code.
  • Model account has no API key (subscription passthrough) → the gateway forwards the user’s own Claude login token, and usage draws from their seat or subscription. Only Claude Code can do this, because it is the only client with a Claude login flow.
Whether per-token spend on a key issued from your Enterprise organization counts toward your Enterprise contract is a commercial question for your Anthropic account team; the gateway doesn’t change it. See the overview for the side-by-side.
Yes. Create the Anthropic model account with the API key left empty, leave ANTHROPIC_AUTH_TOKEN unset, and put your TrueFoundry key in ANTHROPIC_CUSTOM_HEADERS as x-tfy-api-key. The user signs in with their Claude account when Claude Code starts; the gateway authenticates them with the TrueFoundry key and forwards their Claude token to Anthropic. Usage draws from the subscription, and the gateway still records every request per user.Setup is under Authentication by plan.
No. In third-party inference mode the app never signs in to claude.ai. The “Gateway API key” field takes a TrueFoundry token, inference goes through the gateway’s model account, and Anthropic bills that account’s API key per token. The seat’s usage allowance is never consumed by anything that happens inside Desktop.Two consequences worth planning around: a user who only uses Desktop through the gateway doesn’t need a Claude seat at all, and a user who has a seat for claude.ai or Claude Code pays for that seat plus per-token Desktop usage, with no overlap between the two. If you want seats to pay for Desktop, keep Desktop first-party and govern it with Anthropic Inference Hooks instead, accepting that the gateway then has no cost visibility. See Claude Desktop.
Requests start failing. To avoid that, route Claude Code through a virtual model with the subscription (empty-key) account at priority 0 and a pay-per-token Anthropic account as the fallback. The gateway uses the subscription until it is rate-limited, then routes overflow to the API key. Only one target in a virtual model can be a passthrough account, since there is only one Authorization header to forward.Steps are under Fall back from the subscription to the Anthropic API.
No. The gateway counts the tokens in each request and multiplies by the model’s list price so you can attribute spend to users and teams, enforce budgets, and watch trends. TrueFoundry does not charge for tokens. The only token bill is Anthropic’s (or your cloud provider’s), by whichever route the credential took. On subscription passthrough the Analytics figure is an estimate of what the usage would cost at API rates; the seat is flat, so nobody is invoiced that amount.
Yes. The model account then holds your cloud credentials, the gateway always calls the provider with those, and usage is billed per token to your cloud account. Claude plans and subscription passthrough don’t apply; every interface is the gateway-key pattern.The trade-off is fidelity. On Anthropic direct, anthropic-beta headers and new request fields pass through unfiltered. On other providers the gateway checks each beta against a per-provider allowlist and drops the rest, Anthropic’s server-side tools (built-in web search) don’t run, and model ids differ. Claude Code and Desktop work, but some of their newest features arrive later or not at all. See Forwarding Anthropic beta features.

Authentication and identity

From the TrueFoundry credential on the request, sent either as Authorization: Bearer <token> or in the x-tfy-api-key header. The gateway recognises three kinds: tokens it issued (Personal Access Tokens, Virtual Account tokens, and the short-lived tokens tfy-local-ai-setup mints from a device login), JWTs from an identity provider you have registered, and, in SWG mode, an X-Authenticated-User assertion from a trusted upstream proxy.The identity is as good as the token. One Personal Access Token pasted into a shared managed-settings.json makes the whole fleet look like one user. The MDM binary mints a separate token per user from their own SSO login, which is what gives per-user attribution at scale. See Enforce with MDM.
Because two different credentials may need to travel on the same request. ANTHROPIC_AUTH_TOKEN becomes the Authorization header. On the gateway-key pattern that header carries your TrueFoundry key and nothing else is needed. On subscription passthrough Claude Code reserves Authorization for the user’s Claude login token, so the TrueFoundry key moves to x-tfy-api-key in ANTHROPIC_CUSTOM_HEADERS. The gateway checks both locations.
Only when the model account has no API key. If a key is stored, the gateway authenticates the inbound request, then builds its own outbound request with the stored key; the client’s Authorization header never leaves the gateway. If no key is stored, the gateway forwards the inbound Authorization header to Anthropic as-is.So “empty API key” is an instruction to pass the client’s bearer through. It only works when that bearer is a real Anthropic token, which today means Claude Code after /login. Point Claude Desktop at an empty-key account and Anthropic receives a TrueFoundry token and returns 401 Invalid bearer token.
Yes. The thing that is shared is the Anthropic API key, and it lives on the gateway’s model account, never on devices. What MDM pushes is the TrueFoundry credential, and tfy-local-ai-setup --claude-desktop mints that per user: it runs the device login as the signed-in user, writes that user’s token into the per-user managed preferences, locks the file, and refreshes it on every run. Each Desktop request then carries that user’s identity and appears under their name in Analytics, exactly as Claude Code does. See Claude Desktop → MDM.
No. Anthropic tokens are opaque to TrueFoundry; there is no way to validate one or learn who it belongs to, and on passthrough the gateway forwards it without reading it. Every request needs a credential the gateway can verify, which is why passthrough sends both tokens. If you want to avoid minting TrueFoundry tokens, register your IdP and have clients send its JWTs; the gateway validates those directly. See Identity providers.

Claude Code configuration

Not on an Anthropic-direct route; set it to 0 if you want Claude Code’s newer features. Forward specific betas with "ANTHROPIC_CUSTOM_HEADERS": "x-tfy-anthropic-beta: <value>,<value>". x-tfy-anthropic-beta replaces, rather than merges with, a raw anthropic-beta header.On Anthropic /v1/messages there is no allowlist: beta values, safeguards, tool_reference, defer_loading, and thinking blocks pass through. On Bedrock, Vertex, and Foundry the gateway filters betas against a per-provider allowlist and silently drops unsupported ones, so a feature such as tool search may not run there. To see what was forwarded, check the trace attributes tfy.model.anthropic_betas_requested and tfy.model.anthropic_betas_sent. Note that DISABLE_EXPERIMENTAL_BETAS=1 also overrides ENABLE_TOOL_SEARCH=TRUE. See Forwarding Anthropic beta features.
The slug Claude Code sent doesn’t match a gateway model this user can access. Check three things: the provider-account/model-name values in ANTHROPIC_MODEL and the ANTHROPIC_DEFAULT_*_MODEL variables are spelled exactly as the gateway lists them; none of them is missing (a slug ending in /undefined means an env var is unset); and the user has been granted the model. To populate the picker from the gateway instead of hardcoding, set CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1; the gateway serves /v1/models. On a self-hosted gateway, ANTHROPIC_BASE_URL ends at /api/llm, not /api/llm/v1. See Model configuration.
Append [1m] to the model name in settings.json (for example claude-code/claude-sonnet[1m]). Claude Code strips the suffix before sending and treats the model as the 1M-context variant internally, so the gateway sees the normal model name. Don’t create a model with brackets in its name on the gateway; brackets aren’t allowed in integration names. See the Claude Code FAQ.
First upgrade Claude Code. Several reported differences on the subscription-passthrough route were client bugs fixed in later releases. If a current client still misbehaves, open a ticket with the trace attributes tfy.model.anthropic_betas_requested and tfy.model.anthropic_betas_sent from a failing request, the virtual model configuration, and the settings.json env block.

Claude Desktop and Cowork

Expected. Third-party mode bypasses the claude.ai account UI, and that panel is where Anthropic’s connector directory lives. Connectors still work, just differently: remote MCP servers (HTTP/SSE) are admin-pushed through the managedMcpServers managed preference, each pointing at a server on the MCP Gateway, and users click Connect to sign in as themselves; local stdio MCP servers remain user-addable via claude_desktop_config.json while isLocalDevMcpEnabled is true. Claude Desktop doesn’t support MDM-managed and user-added remote servers at the same time, so push a complete catalog. See Govern MCP traffic.
They aren’t lost. Connecting to a gateway creates a new local profile; the earlier conversations are still in the user’s claude.ai account. To bring them across, enable the claudeAiImport managed key and users get an Import from Claude wizard. On a Team or Enterprise workspace an owner must also allow members to export their own data, and uploaded file contents never come across. See Enforce with MDM.
The 401 is from Anthropic, not the gateway. The Anthropic model account that Desktop’s models route to has no API key (or an invalid one), so the gateway forwarded Desktop’s TrueFoundry bearer to Anthropic, which rejected it. Add a valid API key to the model account. Desktop can’t use subscription passthrough, so the model account it uses must always hold a key.
By default Desktop inlines the full schema of every managed connector into each session. Turn on tool search with the toolSearchEnabled managed preference (a plain boolean, MDM-only, Desktop 1.21459.0+); Desktop then loads a tool’s schema only when it is first used. Setting ENABLE_TOOL_SEARCH yourself has no effect, because Desktop strips it from the session environment while toolSearchEnabled is unset. Pilot with one group first: enabling it adds the tool-search-tool-2025-10-19 beta and tool_reference blocks to requests, which pass through on Anthropic direct but may be filtered on other providers. See Enforce with MDM.

Governance and rollout

Govern, yes; track cost, no. Claude Web has no endpoint setting and inference runs in Anthropic’s cloud, so the gateway never sees token counts. For Enterprise orgs, Anthropic Inference Hooks send every prompt from web, desktop, mobile, Claude Code, and Slack to the gateway for an allow/deny verdict, configured once in Anthropic’s admin console. For other plans, aitori intercepts traffic on the device. Either way you get guardrails and an audit trail on the visible content; per-user usage lives in Anthropic’s console. Many customers also block claude.ai at the firewall and direct users to the native apps, which can be fully governed. See Claude Web.
Re-authentication on every model switch was a limitation of early tfy-local-ai-setup releases; current binaries refresh silently from the token in ~/.tf/refresh-token, so update to the latest release and schedule the script hourly. The browser prompt should then appear only on first run or after 30 days without a refresh.Intune device-assigned deployments run as SYSTEM, which cannot open a browser. Use SYSTEM for the privileged parts (install the binary, write config under C:\ProgramData\TrueFoundry, write and lock managed-settings.json), have each user complete the TrueFoundry login once interactively, and let the hourly run refresh from then on. See Enforce with MDM.