Baseten vs Modal: Pricing, Cold Starts and What the Rate Card Hides

Built for Speed: ~10ms Latency, Even Under Load
Blazingly fast way to build, track and deploy your models!
- Handles 350+ RPS on just 1 vCPU â no tuning needed
- Production-ready with full enterprise support
Both are serverless GPU platforms, both scale to zero by default, both run production traffic for companies you have heard of. This is a fair fight, but they are not the same shape. Baseten built a model-serving platform and extended it. Modal built a general compute platform where inference is one product of six. That shows up in the SDK, in what the rate card counts, and in whether you can run it in your own account.
The most useful thing below is the pricing methodology, because published GPU per-hour rates are not comparable between these two and the error is large enough to flip a decision. Figures checked 25 September 2026.
Baseten: overview
Baseten is production inference for open-source, custom and fine-tuned models, across Dedicated Deployments, per-token Model APIs and Training (pricing).
Founded 2019 by Tuhin Srivastava (CEO), Amir Haghighat (CTO), Phil Howes and Pankaj Gupta. Funding ran through Series B, C, D and a $300M Series E at $5B, to a $1.5B Series F at a $13B valuation on 22 June 2026 (announcement); total raised âmore than $2B.â Revenue reportedly grew 20x year on year; ARR is unpublished [VERIFY]. Customers include Cursor, Notion, Harvey, HubSpot, OpenEvidence and Vercel.
On a number you may have seen: press reports in late September 2026 (Axios 22 Sep, Bloomberg 23 Sep) describe talks around a much larger round near a reported $26B. Baseten has published nothing. $13B is the last confirmed figure and the one we use.
[SCREENSHOT: Baseten â the Dedicated Deployments dashboard showing replica states, autoscaling config and latency percentiles for a live model]
- Pricing. Basic is $0/month pay-as-you-go with SOC 2 Type II and HIPAA included and no platform fee on any published tier; Pro and Enterprise are on quote, Enterprise adding SLAs, self-hosted deployments, data residency and advanced RBAC. Billing is per minute, no charge for idle. Free credits exist, amount unpublished [VERIFY]. Baseten also publishes a per-token rate card â GPT OSS 120B at $0.10 in / $0.50 out per million tokens.
- Cold starts. The Baseten Delivery Network caches weights in tiers â node-local NVMe, an in-cluster peer cache, then a mirrored origin â giving 2-3x faster cold starts at over 2 GB/s to H100 nodes and, commercially the important bit, âweight transfer doesnât consume billable GPU timeâ (BDN post, 9 Apr 2026). No latency-in-seconds figure is in current docs [VERIFY]. Baseten blog âLast updatedâ stamps are edit dates, not publication dates.
- Deployment and SDK. Truss packaging â config.yaml plus an optional model/model.py with load and predict, or your own Docker image â with baseten model push --watch for a live-reload dev deployment and baseten model push for an immutable production one. Concurrency is two-level â concurrency_target for the autoscaler, predict_concurrency inside the container â and multi-node with InfiniBand is supported.
- Self-hosted and multi-cloud â the cleanest differentiator. Baseten Cloud; Self-hosted, workload plane in your VPC, where âinference requests route directly to your workload plane without passing through Baseten infrastructureâ; and Hybrid, with spillover. Above that sits multi-cloud capacity management, marketed as â20+ clouds as one GPU poolâ, with cross-region failover.
- Observability and compliance. Latency percentiles, replica states and GPU metrics, plus engine-native vLLM and SGLang metrics auto-detected with no config; Prometheus, OTLP and OTel tracing. SOC 2 Type II, HIPAA, GDPR, with zero data retention by default for synchronous inference. ISO 27001 [VERIFY]; egress pricing not found [VERIFY].
Baseten also acquired Blaxel, an agent-sandbox and microVM infrastructure company, announced 10 September 2026, terms undisclosed (announcement) â a direct move onto Modalâs Sandboxes turf, two weeks old with nothing shipped to judge yet.
Genuinely better for: regulated buyers needing inference in their own VPC, where Modal has no equivalent; anyone with PHI, since HIPAA sits on the free tier; teams wanting TensorRT-LLM-class optimisation without building it; and region-pinned or small-GPU workloads.
Modal: overview
Modal calls itself âa cloud built for AI. Not a single-purpose GPU cloud, but a platform with the right primitivesâ (Series C post). Products: Inference, Sandboxes, Training, Notebooks, Batch, Core Platform. That list is the argument â inference is one item.
Founded 2021 by Erik Bernhardsson and Akshat Bubna; the founding year comes from TechCrunch, not a Modal page [VERIFY]. The team built its own filesystem, container runtime, scheduler and image builder, which shows in the cold-start numbers. Funding ran seed, Series A and an $87M Series B at $1.1B, to a $355M Series C at $4.65B post on 21 May 2026; total âover $466M.â Modal publishes what Baseten does not â âsurpassing $300 million in annualized revenueâ, and that âSandboxes already drive more than a third of our revenueâ. Customers include Cognition, DoorDash, Suno and Ramp.
[SCREENSHOT: Modal â the app dashboard showing container counts, GPU utilisation and cold-start timing for a deployed Function]
- Pricing. Starter $0/mo plus compute, $30/mo free compute, 1-day log retention. Team $250/mo plus compute, $100/mo free compute, 30-day logs. Enterprise custom, adding audit logs, SAML SSO, HIPAA and RBAC. Billing is per second, and here is the crux: CPU and memory are billed separately from the GPU â CPU at $0.04716 per physical core-hour, memory at $0.007992/GiB/hr. Sandbox and Notebook CPU and memory cost 3x the Function rate. Region pinning is 1.15x broad or 1.75x narrow; non-preemptible is 3x. No per-token rate card.
- Cold starts. Containers boot in roughly a second. Memory Snapshots are the named feature: CPU snapshots GA, GPU snapshots Alpha. Modalâs docs state 3-10x faster and the engineering blog headlines 10x (Parakeet 20s to 2s). Modalâs Series C post used a â100xâ line that its own docs and engineering blog do not support, so 3-10x is the figure to cite. GPU snapshots âdo not speed up model loading from storageâ and âmay even worsen them, by adding overheadâ, are generally incompatible with multi-GPU code, and expire after 7 days.
- Deployment and SDK. Decorator-based Python â @app.function(), @app.cls(), @modal.fastapi_endpoint() â with images defined in Python by method chaining rather than a Dockerfile, and modal run / modal deploy / modal serve for the loop. Containers run under gVisor; JS and Go SDKs are Beta. Regions are broad (us, eu, ap) or narrow (us-west, jp) and strictly pinned; multi-node clusters are Beta. If you like writing infrastructure as Python, this is the nicest developer experience in the category.
- No self-hosted option. No BYOC, on-prem or customer-VPC deployment: the Enterprise tier lists no deployment-location feature, cloud= picks which cloud Modal schedules on rather than your account, and llms.txt has zero occurrences of BYOC, self-host, on-prem or VPC. The residency limit is worth quoting: all logs, all durable storage, Function payloads over 2 MiB and all async payloads are stored in the United States regardless of container region.
- Observability and compliance. GPU utilisation, power, temperature and memory, with Modalâs own caveat that these âcanât be used to directly debug performance issuesâ, plus OTLP export anywhere. SOC 2 Type 2 yes. HIPAA is qualified â BAA required, Enterprise-only, and Volumes v1, Images, Memory Snapshots and user code are out of scope of the BAA. Inputs and outputs are retained up to 7 days. ISO 27001 [VERIFY] â not listed.
[SCREENSHOT: Modal â the pricing page showing the GPU table alongside the separate CPU and memory line items]
Genuinely better for: anything that is not a model server â sandboxes, untrusted code, RL environments, batch, notebooks. Python ergonomics. Large-GPU, non-pinned, spiky workloads. GPU breadth, since L40S, A100 40GB and B300 have no Baseten equivalent. And init-heavy cold starts dominated by JIT and imports rather than weight loading.
Architecture and deployment compared
The honest framing on cold starts: they solve different halves of the same problem. Baseten attacks weight delivery and takes transfer off the billing clock. Modal attacks process state with snapshot restore, which explicitly does not help weight loading. If your cold start is 140GB of weights, Basetenâs approach is the relevant one; if it is imports, JIT and CUDA graph capture, Modalâs is.
GPU pricing: the headline rate and the real rate
Methodology. Baseten publishes per-minute rates for all-in instance SKUs â the GPU ships with fixed vCPU and RAM. Modal publishes per-second rates for the GPU only, with CPU and memory billed separately. To compare like for like, convert both to per-hour (Baseten x60, Modal x3600), then add to Modalâs GPU rate the CPU and memory Baseten bundles, at Modalâs own rates: $0.04716 per physical core-hour (one core = two vCPU) and $0.007992 per GiB-hour. Pinning and non-preemptible multipliers are excluded; adding them makes Modal dearer still. Sources: baseten.co/pricing, docs.baseten.co/deployment/resources, modal.com/pricing.
Headline published rates
Adjusted â Modal matched to Basetenâs bundled CPU and RAM
L40S and A100 40GB have no Baseten SKU to adjust against.
The takeaway: âModal is 40% cheaper on H100â collapses to about 19% once you buy comparable CPU and memory, sits at near parity on A100 80GB, and reverses on small GPUs, where Baseten is 9-22% cheaper. Add pinning or the 3x non-preemptible multiplier and Baseten wins outright on any guaranteed workload.
A gotcha for benchmarkers: Modalâs bare gpu="A100" may be silently upgraded to 80GB, and gpu="H100" may land on an H200. Pin it with gpu="H100!".
Where teams hit trouble
1. The rate card is not the bill. Teams size a budget from a per-hour GPU number, then meet CPU, memory, volumes, region multipliers and a 3x sandbox rate. Basetenâs bundling is easier to forecast; Modalâs itemisation is truer to consumption. Neither is comparable without the arithmetic above.
2. Region selection is not data residency. Modal will pin a container to eu-west and still store your logs, durable storage and any payload over 2 MiB in the US. If a regulator is asking, container region is not the answer.
3. The model server is only part of the system. You also have commercial model APIs, authentication, cost attribution, rate limits, guardrails and an audit trail. Neither platform is that layer, and adopting either still leaves it to you.
TrueFoundryâs position: a different question
Straight answer rather than a manufactured three-way race. Baseten and Modal are hosted inference clouds. TrueFoundry runs in your cloud account. Different buyer, different problem.
They answer âhow do I serve this model without operating GPUs.â TrueFoundry answers âwe already have GPU capacity, committed spend, or a regulator, and we need a platform on our own infrastructure.â If the first question is yours, buy one of them â Basetenâs self-hosted mode is the closest overlap.

TrueFoundry deploys as SaaS across more than 12 regions on three clouds, or self-hosted in your VPC via Helm and OpenTofu/Terraform, or on-prem, or air-gapped. Self-hosted, all LLM traffic stays inside your infrastructure and TrueFoundry is not in the live path. Serving covers vLLM, SGLang, Triton, TorchServe, MLflow and LitServe, plus a model registry, autoscaling on CPU, RPS or cron, and fractional GPUs via time-slicing and MIG â which matters when you own the hardware and a whole H100 per small model is waste.

The other half is the AI Gateway, where self-hosted and hosted inference stop being separate problems. It fronts 1,000+ LLMs behind one OpenAI-compatible API, adds roughly 3-4 ms of latency and handles 350+ RPS on 1 vCPU, so your Baseten endpoint, your Modal endpoint and your OpenAI key sit behind one interface with shared cost attribution, rate limits and audit trail. Auto Routing is the cost lever: across 550 prompts, routing by complexity cut cost 69% while retaining 98% of quality, with mean latency falling from 7.6s to 4.0s.

Honest limit: TrueFoundry is not a serverless GPU marketplace â you bring capacity. If you have no GPUs and no wish to get any, pick Baseten or Modal instead.
Head-to-head
Related reading
- Model Deployment Tools â the wider landscape
- Multi-Cloud GPU Orchestration
- Total Cost of Ownership for GenAI Infrastructure
- Scaling to Zero in Kubernetes
- What Is an AI Control Plane?
Conclusion
Baseten and Modal are both good. The comparison is close on the axis people check first â price â and not close at all on the axes that decide procurement.
If the workload must run in your own VPC, or you handle PHI, or you want TensorRT-LLM without doing the engine work, or you run small and pinned GPUs, Baseten is the answer. If you need sandboxes, batch and notebooks alongside inference, or Python ergonomics matter above all, Modal is.
What we would not do is choose on the headline GPU rate. It is not comparable between these two, and correcting it moves an apparent 40% advantage to 19%, to parity, or to the other vendor.
If the real constraint is that the workload has to stay inside infrastructure you already own and pay for, neither hosted platform is the answer. That is a control-plane problem, and the one we build for.
TrueFoundry AI Gateway delivers ~3â4 ms latency, handles 350+ RPS on 1 vCPU, scales horizontally with ease, and is production-ready, while LiteLLM suffers from high latency, struggles beyond moderate RPS, lacks built-in scaling, and is best for light or prototype workloads.


Recent Blogs
Frequently asked questions
Baseten vs Modal: which is cheaper?
It depends on the GPU, and the published rates mislead. On headline numbers Modal looks 37-40% cheaper on H100 and A100 80GB. Add the CPU and memory Baseten bundles and the H100 gap narrows to about 19%, A100 80GB is parity, and on T4 and A10 Baseten is 9-22% cheaper. Modalâs pinning and non-preemptible multipliers push further toward Baseten, and its $250/mo Team fee is fixed cost Baseten does not charge.
What is Modal pricing, exactly?
Starter $0/mo with $30/mo free compute; Team $250/mo with $100/mo free compute; Enterprise custom. Compute is per second and itemised: GPU at its own rate, CPU at $0.04716 per physical core-hour, memory at $0.007992 per GiB-hour. Sandbox and Notebook compute is 3x the Function rate, and pinning and non-preemptible capacity carry multipliers.
What is Baseten pricing, exactly?
Basic is $0/month pay-as-you-go with no platform fee, including SOC 2 Type II and HIPAA; Pro and Enterprise are on quote. GPU compute bills per minute against all-in instance SKUs bundling vCPU and RAM â $6.4998/hr for an H100 80GB. Per-token Model API rates are published; the free-credit amount is not.
How fast are Modalâs cold starts?
Containers boot in roughly a second, and Memory Snapshots deliver 3-10x faster restores per Modalâs docs. The â100xâ figure in Modalâs Series C post is not supported by its own documentation, so plan against 3-10x. GPU snapshots are Alpha and do not speed up loading weights from storage.
Is Baseten valued at $26 billion?
Not confirmed. The last valuation Baseten has published is $13B, from a $1.5B Series F on 22 June 2026. Press reports in late September 2026 describe talks around a larger round, but that is talks-stage reporting and the company has published nothing.










.png)
.png)
.png)
.png)
.png)
.png)

.webp)
.webp)


.webp)
.webp)
.webp)






