OpenRouter free models: what's actually free and what it costs you

Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU — no tuning needed
- Production-ready with full enterprise support
Nine of the nineteen free models on OpenRouter run on providers whose published policy says they may train on your prompts. The label is OpenRouter's, from the structured data policy it publishes for every provider it routes to.
Somebody is paying for the GPUs. The trade is often reasonable: free inference in exchange for usage data, feedback and adoption. "$0 per token" is the bill. The data policy, the single host, and how long the model stays listed are the rest of what it costs, as of 5 October 2026.
What's actually free on OpenRouter
Free models on OpenRouter mostly carry a :free suffix on the model ID, like nvidia/nemotron-3.5-lightning:free. On 5 October the models API listed 17 of them, plus two more priced at zero without the suffix: inclusionai/ling-3.1-flash, added two days earlier, and stealth/space-bunny-alpha, an anonymous model served by OpenRouter's stealth provider. That's 19 free text models. There's also openrouter/free, a router that picks one of them at random for each request, filtering for the features your request needs.
OpenRouter's pricing page describes the Free plan as "25+ free models" from "4 free providers". The API shows 19 models from 10 providers. The gap is probably catalog churn and different counting rather than anything hidden, but it's a reminder to check the API rather than the marketing.
Grouped by what the host's data policy says:
What it costs you: your prompts
OpenRouter publishes a data policy for each of the 92 providers it lists: whether the provider may train on prompts, whether it retains them, and for how long. Only five of those 92 are marked as possibly training on prompts. Four of the five serve free models: NVIDIA, Thinking Machines, Liquid and the stealth provider. The fifth, DeepSeek, is paid.

So for free models specifically, the split is nine on hosts that may train, six on hosts that keep your prompts without training on them (Google AI Studio for 55 days, Cohere for 30, Poolside and AtlasCloud for a period they don't state), and four with zero retention, three of them on Novita.
Three caveats keep this fair. First, these are each provider's default terms. OpenRouter says an individual endpoint can carry a different policy from its provider's default, and it sometimes negotiates stricter terms, so some free routes may be better than shown. Second, when OpenRouter can't establish a policy, it marks the endpoint as both training and retaining, which errs toward caution. Third, "may train" describes what the terms permit, not what happens to any particular prompt.
The terms explain the split. The ones OpenRouter links for NVIDIA are its API trial terms, and for Thinking Machines its free research tier terms. Model makers offering free inference to learn from usage is a normal arrangement, and it's a perfectly good one for public data and side projects. It's a poor one for customer data, internal code or anything under an NDA.
OpenRouter gives you controls. Your privacy settings have separate switches for paid and free endpoints that may train on inputs, plus a zero-data-retention mode, and requests can set data_collection: "deny". Two things to know about them. OpenRouter's pricing table lists data-policy-based routing as a Standard-plan feature, not part of the Free plan. And none of the 92 providers is currently flagged as able to publish prompts, so the "free endpoints that may publish prompts" switch gates nothing today.
What it costs you: reliability
Every one of the 19 free models has exactly one endpoint. A paid model on OpenRouter can have a dozen hosts and fall back between them. A free model has one host, so when it's busy or down there's nowhere to fall back to.
That interacts badly with the privacy controls above. Switch off "free endpoints that may train" and OpenRouter stops routing to those hosts, as documented. With only one host per model, that doesn't send you somewhere safer. It makes nine of the nineteen models unavailable.
One-day uptime ran from 97% to 100% for most free endpoints on the day we pulled them, with one outlier, Nemotron 3 Nano Omni, at 86%. Then there are the caps: 20 requests a minute on any free model, and 50 a day until you've bought at least 10 credits, then 1,000 a day. Failed requests can still count against the daily allowance. Our post on OpenRouter rate limits covers how those caps work and why they depend on your payment history.
What it costs you: a different model than the name suggests
A free variant isn't always the same deployment as the paid model with the same name. Gemma 4 31B is the clearest example: the free endpoint caps output at 32,768 tokens, while paid hosts for the same model go up to 262,141, and there are 13 of them to choose from.
It cuts both ways. Nemotron 3 Ultra's free endpoint offers a million-token context against 262,144 on its paid hosts, and Inkling's free endpoint doubles its paid context. Free can be larger or smaller than the paid route. The model string alone doesn't tell you which, which is true of OpenRouter in general, as we covered in what OpenRouter is and how it routes requests.
What it costs you: it may not be there next month
Free models on OpenRouter are young. On 5 October the median free model had been listed for 75 days, 8 of the 19 had appeared in the previous 60 days, and the oldest, Nemotron 3 Super, had been there 207 days. Two names are previews by label: one has "preview" in its ID, the other is a stealth alpha.
They also disappear. Two free IDs that a widely shared July guide recommended, qwen/qwen3-coder:free and deepseek/deepseek-r1:free, are no longer listed, while the paid versions of both models still are. The free variant was withdrawn, not the model. Treat a free model as a promotion with an end date you won't be told in advance, and pin a paid fallback if anything depends on it.
The openrouter/free router returns a working endpoint on every request by picking a free model at random, so two requests in the same conversation can be answered by different models from different companies under different data policies.
What the paid version costs
This is the part most free-tier guides skip. For six of the nine free models whose host may train, the same model is also served by DeepInfra, which OpenRouter's data marks as neither training nor retaining prompts.
Take Nemotron 3.5 Lightning. Free on NVIDIA, or $0.06 per million input tokens and $0.16 per million output on DeepInfra. At the free tier's own ceiling of 1,000 requests a day, with 2,000 input and 500 output tokens per request, that's 20 cents a day: $6.00 for 30 days, or $6.33 after OpenRouter's 5.5% credit fee. To get 1,000 free requests a day you have to buy $10 of credits anyway, which on its own would cover more than a month and a half of the paid route.

The range is wide, from $6 a month for Lightning to $124 for Inkling, so "just pay" isn't automatically trivial for the biggest models. But for the three smaller Nemotron models, keeping your prompts out of a training pipeline costs less than $16 a month.
When free models make sense
Free OpenRouter models are a good fit for prototyping and for comparing models before you commit, and for evals on public or synthetic data. Side projects, demos and learning fit too. Anything mildly sensitive can go on the four zero-retention models, and should stay off the rest.
They are a poor fit for customer data, internal code, or anything covered by an NDA or a DPA. The same goes for anything that has to stay up, since there's one host and no fallback, and for anything you'll ship, since the model may be withdrawn without notice. At any real volume the paid route is often a few dollars a month.
How TrueFoundry approaches this
Several of these free models come from families that publish open weights, Gemma and Qwen among them. Running those weights on your own GPUs is how you use the model with no third party in the request path. TrueFoundry deploys them with vLLM, TGI and Triton, behind the same AI Gateway that fronts your hosted providers.
If you're keeping OpenRouter, the gateway can sit in front of it and restrict which models each team can call, so free endpoints stay with development teams while production traffic goes to hosts you've vetted. It exposes 1,000+ LLMs through one OpenAI-compatible API, adds roughly 3 to 4 ms of overhead, and handles 350+ RPS on a single vCPU.
Related reading
- OpenRouter reviews, what users say about the platform and where it stops
- OpenRouter prompt caching, on when caching saves money and when it costs more
- OpenRouter alternatives, comparing the options for production teams
- What is an LLM gateway?, the general architecture
Conclusion
OpenRouter free models cost $0 per token, and they are useful for trying a model before you commit. Nearly half run on hosts that may train on your prompts, all of them run on a single host, and few have been listed for more than a few months. Use them to explore, keep sensitive data on zero-retention or paid routes, and check the endpoint before you trust the name.
To run the open weights with no third party in the path, see how the TrueFoundry AI Gateway serves self-hosted and hosted models together.
TrueFoundry AI Gateway delivers ~3–4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
Are OpenRouter free models really free?
They cost nothing per token. You pay in other ways: nine of the nineteen free models run on hosts whose policy says they may train on your prompts, each free model has only one host and no fallback, and usage is capped at 20 requests a minute and 50 a day, or 1,000 a day once you've bought 10 credits.
Do OpenRouter free models train on my data?
It depends on the host. As of 5 October 2026, nine of the nineteen free models ran on providers that OpenRouter marks as possibly training on prompts, six on providers that retain prompts without training, and four on zero-retention hosts. OpenRouter itself says it doesn't train on API data; the policy that matters is the provider's.
How many free models does OpenRouter have?
On 5 October 2026 OpenRouter's API listed 19 free text models: 17 with a :free suffix and two priced at zero without one, served by 10 providers. Its pricing page says "25+". The list changes often, so query /api/v1/models for the current set.
TrueFoundryを自社のVPCやオンプレミスにデプロイできますか?
はい。TrueFoundryはお客様のVPC、オンプレミス、エアギャップ環境、ハイブリッド環境で動作するため、多数のプロバイダーにまたがってルーティングする場合でも、プロンプトとレスポンスがお客様のドメインの外に出ることはありません。
TrueFoundryは、どのモデルサービングバックエンドに対応していますか?
vLLM、TGI、Tritonなどの高性能バックエンドを通じたあらゆるLLMに加え、OllamaなどのOpenAI互換サーバーに対応しています。いずれも同じAI Gatewayに接続され、同じ方法で統制されます。














.png)
.png)



.png)



.png)
.png)
.png)

.png)
.png)





