Blank white background with no objects or features visible.

「Gartner Hype Cycle for AI Governance 2026」の全編を無料で公開しています。レポートを入手する →

OpenRouter free models: what's actually free and what it costs you

By Kshitij Gupta

Published: October 8, 2026

TL;DR:

OpenRouter's free models cost nothing per token. Nine of the nineteen run on hosts whose published policy says they may train on your prompts, every one of them has a single host with no fallback, and they come and go within months. Running the same model on a host that neither trains nor retains costs from about $6 a month at the free tier's own ceiling.

Nine of the nineteen free models on OpenRouter run on providers whose published policy says they may train on your prompts. The label is OpenRouter's, from the structured data policy it publishes for every provider it routes to.

Somebody is paying for the GPUs. The trade is often reasonable: free inference in exchange for usage data, feedback and adoption. "$0 per token" is the bill. The data policy, the single host, and how long the model stays listed are the rest of what it costs, as of 5 October 2026.

What's actually free on OpenRouter

Free models on OpenRouter mostly carry a :free suffix on the model ID, like nvidia/nemotron-3.5-lightning:free. On 5 October the models API listed 17 of them, plus two more priced at zero without the suffix: inclusionai/ling-3.1-flash, added two days earlier, and stealth/space-bunny-alpha, an anonymous model served by OpenRouter's stealth provider. That's 19 free text models. There's also openrouter/free, a router that picks one of them at random for each request, filtering for the features your request needs.

OpenRouter's pricing page describes the Free plan as "25+ free models" from "4 free providers". The API shows 19 models from 10 providers. The gap is probably catalog churn and different counting rather than anything hidden, but it's a reminder to check the API rather than the marketing.

Grouped by what the host's data policy says:

ModelFree hostHost data policyContextMax output
liquid/lfm-2.5-2.6b:freeLiquidMay train65,5368,192
nvidia/nemotron-3-nano-omni-30b-a3breasoning:freeNVIDIAMay train256,00065,536
nvidia/nemotron-3-super-120b-a12b:freeNVIDIAMay train262,144235,929
nvidia/nemotron-3-ultra-550b-a55b:freeNVIDIAMay train1,000,00065,536
nvidia/nemotron-3.5-content-safety:freeNVIDIAMay train128,0008,192
nvidia/nemotron-3.5-lightning:freeNVIDIAMay train1,000,00065,536
stealth/space-bunny-alphaStealthMay train1,000,000524,288
thinkingmachines/inkling-small:freeThinking MachinesMay train1,048,576262,144
thinkingmachines/inkling:freeThinking MachinesMay train1,048,576262,144
cohere/north-mini-code:freeCohereRetains 30 days256,00064,000
dots-studio/dots-3-note-preview:freeAtlasCloudRetains, period not stated512,000460,800
google/gemma-4-26b-a4b-it:freeGoogle AI StudioRetains 55 days262,14432,768
google/gemma-4-31b-it:freeGoogle AI StudioRetains 55 days262,14432,768
poolside/laguna-s-2.1:freePoolsideRetains, period not stated262,14432,768
poolside/laguna-xs-2.1:freePoolsideRetains, period not stated262,14432,768
apodex/apodex-1.1-mini:freeNovitaZero retention262,144235,929
inclusionai/ling-3.0-flash-sante:freeNovitaZero retention262,14432,768
inclusionai/ling-3.1-flashNovitaZero retention262,14432,768
qwen/qwen3.8-27b:freeModelRunZero retention262,144235,929

‍

Take control of your LLM traffic
Route models, enforce budgets, and monitor every request from your own infrastructure.

What it costs you: your prompts

OpenRouter publishes a data policy for each of the 92 providers it lists: whether the provider may train on prompts, whether it retains them, and for how long. Only five of those 92 are marked as possibly training on prompts. Four of the five serve free models: NVIDIA, Thinking Machines, Liquid and the stealth provider. The fifth, DeepSeek, is paid.

Figure 1: the 19 free models, grouped by what the host's data policy says.

So for free models specifically, the split is nine on hosts that may train, six on hosts that keep your prompts without training on them (Google AI Studio for 55 days, Cohere for 30, Poolside and AtlasCloud for a period they don't state), and four with zero retention, three of them on Novita.

Three caveats keep this fair. First, these are each provider's default terms. OpenRouter says an individual endpoint can carry a different policy from its provider's default, and it sometimes negotiates stricter terms, so some free routes may be better than shown. Second, when OpenRouter can't establish a policy, it marks the endpoint as both training and retaining, which errs toward caution. Third, "may train" describes what the terms permit, not what happens to any particular prompt.

The terms explain the split. The ones OpenRouter links for NVIDIA are its API trial terms, and for Thinking Machines its free research tier terms. Model makers offering free inference to learn from usage is a normal arrangement, and it's a perfectly good one for public data and side projects. It's a poor one for customer data, internal code or anything under an NDA.

OpenRouter gives you controls. Your privacy settings have separate switches for paid and free endpoints that may train on inputs, plus a zero-data-retention mode, and requests can set data_collection: "deny". Two things to know about them. OpenRouter's pricing table lists data-policy-based routing as a Standard-plan feature, not part of the Free plan. And none of the 92 providers is currently flagged as able to publish prompts, so the "free endpoints that may publish prompts" switch gates nothing today.

What it costs you: reliability

Every one of the 19 free models has exactly one endpoint. A paid model on OpenRouter can have a dozen hosts and fall back between them. A free model has one host, so when it's busy or down there's nowhere to fall back to.

That interacts badly with the privacy controls above. Switch off "free endpoints that may train" and OpenRouter stops routing to those hosts, as documented. With only one host per model, that doesn't send you somewhere safer. It makes nine of the nineteen models unavailable.

One-day uptime ran from 97% to 100% for most free endpoints on the day we pulled them, with one outlier, Nemotron 3 Nano Omni, at 86%. Then there are the caps: 20 requests a minute on any free model, and 50 a day until you've bought at least 10 credits, then 1,000 a day. Failed requests can still count against the daily allowance. Our post on OpenRouter rate limits covers how those caps work and why they depend on your payment history.

What it costs you: a different model than the name suggests

A free variant isn't always the same deployment as the paid model with the same name. Gemma 4 31B is the clearest example: the free endpoint caps output at 32,768 tokens, while paid hosts for the same model go up to 262,141, and there are 13 of them to choose from.

It cuts both ways. Nemotron 3 Ultra's free endpoint offers a million-token context against 262,144 on its paid hosts, and Inkling's free endpoint doubles its paid context. Free can be larger or smaller than the paid route. The model string alone doesn't tell you which, which is true of OpenRouter in general, as we covered in what OpenRouter is and how it routes requests.

What it costs you: it may not be there next month

Free models on OpenRouter are young. On 5 October the median free model had been listed for 75 days, 8 of the 19 had appeared in the previous 60 days, and the oldest, Nemotron 3 Super, had been there 207 days. Two names are previews by label: one has "preview" in its ID, the other is a stealth alpha.

They also disappear. Two free IDs that a widely shared July guide recommended, qwen/qwen3-coder:free and deepseek/deepseek-r1:free, are no longer listed, while the paid versions of both models still are. The free variant was withdrawn, not the model. Treat a free model as a promotion with an end date you won't be told in advance, and pin a paid fallback if anything depends on it.

The openrouter/free router returns a working endpoint on every request by picking a free model at random, so two requests in the same conversation can be answered by different models from different companies under different data policies.

What the paid version costs

This is the part most free-tier guides skip. For six of the nine free models whose host may train, the same model is also served by DeepInfra, which OpenRouter's data marks as neither training nor retaining prompts.

Take Nemotron 3.5 Lightning. Free on NVIDIA, or $0.06 per million input tokens and $0.16 per million output on DeepInfra. At the free tier's own ceiling of 1,000 requests a day, with 2,000 input and 500 output tokens per request, that's 20 cents a day: $6.00 for 30 days, or $6.33 after OpenRouter's 5.5% credit fee. To get 1,000 free requests a day you have to buy $10 of credits anyway, which on its own would cover more than a month and a half of the paid route.

Figure 2: monthly cost of the same model on a host that doesn't train or retain, at 1,000 requests a day.

The range is wide, from $6 a month for Lightning to $124 for Inkling, so "just pay" isn't automatically trivial for the biggest models. But for the three smaller Nemotron models, keeping your prompts out of a training pipeline costs less than $16 a month.

When free models make sense

Free OpenRouter models are a good fit for prototyping and for comparing models before you commit, and for evals on public or synthetic data. Side projects, demos and learning fit too. Anything mildly sensitive can go on the four zero-retention models, and should stay off the rest.

They are a poor fit for customer data, internal code, or anything covered by an NDA or a DPA. The same goes for anything that has to stay up, since there's one host and no fallback, and for anything you'll ship, since the model may be withdrawn without notice. At any real volume the paid route is often a few dollars a month.

How TrueFoundry approaches this

Several of these free models come from families that publish open weights, Gemma and Qwen among them. Running those weights on your own GPUs is how you use the model with no third party in the request path. TrueFoundry deploys them with vLLM, TGI and Triton, behind the same AI Gateway that fronts your hosted providers.

If you're keeping OpenRouter, the gateway can sit in front of it and restrict which models each team can call, so free endpoints stay with development teams while production traffic goes to hosts you've vetted. It exposes 1,000+ LLMs through one OpenAI-compatible API, adds roughly 3 to 4 ms of overhead, and handles 350+ RPS on a single vCPU.

Try TrueFoundry AI Gateway
Connect your models and start managing LLM traffic through one API.

Related reading

Conclusion

OpenRouter free models cost $0 per token, and they are useful for trying a model before you commit. Nearly half run on hosts that may train on your prompts, all of them run on a single host, and few have been listed for more than a few months. Use them to explore, keep sensitive data on zero-retention or paid routes, and check the endpoint before you trust the name.

To run the open weights with no third party in the path, see how the TrueFoundry AI Gateway serves self-hosted and hosted models together.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Start free
Table of Contents

One Gateway for Every LLM, Agent and MCP Server

Book a 30-min with our AI expert

Book a Demo

The fastest way to build, govern and scale your AI

Book Demo
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Discover More

No items found.
October 8, 2026
|
5 min read

エージェントセキュリティはシステムの問題である:プロンプトインジェクションからランタイム制御まで

No items found.
October 8, 2026
|
5 min read

MCPにおけるHuman in the Loop:TrueFoundryとKongの比較

比較
October 8, 2026
|
5 min read

AIガバナンスフレームワークとは?

No items found.
October 8, 2026
|
5 min read

エンタープライズグレードでのループエンジニアリング:ラップトップループからガバナンスされたランタイムへ

ソートリーダーシップ
No items found.

Recent Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.

Frequently asked questions

Are OpenRouter free models really free?

They cost nothing per token. You pay in other ways: nine of the nineteen free models run on hosts whose policy says they may train on your prompts, each free model has only one host and no fallback, and usage is capped at 20 requests a minute and 50 a day, or 1,000 a day once you've bought 10 credits.

Do OpenRouter free models train on my data?

It depends on the host. As of 5 October 2026, nine of the nineteen free models ran on providers that OpenRouter marks as possibly training on prompts, six on providers that retain prompts without training, and four on zero-retention hosts. OpenRouter itself says it doesn't train on API data; the policy that matters is the provider's.

How many free models does OpenRouter have?

On 5 October 2026 OpenRouter's API listed 19 free text models: 17 with a :free suffix and two priced at zero without one, served by 10 providers. Its pricing page says "25+". The list changes often, so query /api/v1/models for the current set.

TrueFoundryを自社のVPCやオンプレミスにデプロイできますか?

はい。TrueFoundryはお客様のVPC、オンプレミス、エアギャップ環境、ハイブリッド環境で動作するため、多数のプロバイダーにまたがってルーティングする場合でも、プロンプトとレスポンスがお客様のドメインの外に出ることはありません。

TrueFoundryは、どのモデルサービングバックエンドに対応していますか?

vLLM、TGI、Tritonなどの高性能バックエンドを通じたあらゆるLLMに加え、OllamaなどのOpenAI互換サーバーに対応しています。いずれも同じAI Gatewayに接続され、同じ方法で統制されます。

Take a quick product tour
Start Product Tour
Product Tour