DeepSeek V4-Pro Is GA: What an MIT-Licensed Frontier Model Actually Changes
.png)
Auf Geschwindigkeit ausgelegt: ~ 10 ms Latenz, auch unter Last
Unglaublich schnelle Methode zum Erstellen, Verfolgen und Bereitstellen Ihrer Modelle!
- Verarbeitet mehr als 350 RPS auf nur 1 vCPU — kein Tuning erforderlich
- Produktionsbereit mit vollem Unternehmenssupport

TrueFoundry AI Gateway bietet eine Latenz von ~3—4 ms, verarbeitet mehr als 350 RPS auf einer vCPU, skaliert problemlos horizontal und ist produktionsbereit, während LiteLM unter einer hohen Latenz leidet, mit moderaten RPS zu kämpfen hat, keine integrierte Skalierung hat und sich am besten für leichte Workloads oder Prototyp-Workloads eignet.



Steuern, implementieren und verfolgen Sie KI in Ihrer eigenen Infrastruktur
Aktuelle Blogs
Häufig gestellte Fragen
What is DeepSeek V4 and when was it released?
DeepSeek V4 is DeepSeek’s mixture-of-experts model family. The V4 preview, covering V4-Pro and V4-Flash, launched on 24 April 2026. DeepSeek-V4-Pro reached general availability as DeepSeek-V4-Pro-0813 on 13 August 2026.
Is DeepSeek V4-Pro open source?
The weights and repository are published under the MIT licence on Hugging Face at deepseek-ai/DeepSeek-V4-Pro-0813. That covers the weights and code, not the training data, so “open weight” is the more precise term than “open source”.
How much does the DeepSeek V4 API cost?
As of September 2026: $0.66 per million input tokens off-peak and $1.32 at peak on a cache miss, $0.022 and $0.044 on a cache hit, and $1.98 and $3.96 per million output tokens. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays.
Can I self-host DeepSeek V4-Pro?
Yes, and DeepSeek publishes vLLM and SGLang serving recipes. Be realistic about scale: it is a 1.7T-parameter model and the reference configuration in the model card is a multi-GPU GB300 node.
Does an AI gateway add meaningful latency?
TrueFoundry’s AI Gateway adds roughly 3-4 ms and sustains 350+ RPS on 1 vCPU, which is immaterial next to multi-second generation times.
Do I have to pick one model?
No, and the benchmark spread above is the argument against it. Routing by task type across open and closed models is usually cheaper and better than standardising on one.













.png)
.png)
.png)
.png)
.png)
.png)
.png)
.png)




.png)

.png)





