Blank white background with no objects or features visible.

TrueFoundry kündigt die Übernahme von Seldon AI an und erweitert damit seine Control Plane für Enterprise-KI. Vollständigen Bericht lesen →

We Checked Alibaba's Math: Qwen3.8-Max's Price Advantage Doesn't Survive Contact With Real Tasks

von Amrutha Potluri

Published: August 6, 2026

On August 3, Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter model (95B active per inference) priced at $2/$6 per million input/output tokens, roughly 5x cheaper than GPT-5.6 Solon output. Alongside it, Alibaba published five benchmark scores, including a jump onFrontierSWE from 40.7 to 73.5 versus its predecessor.

We ran an independent check ourselves. Ten self-contained coding tasks, a separate hand-written suite (not a reproduction of FrontierSWE or DeepSWE) spanning easy string manipulation up to a dependency-order resolver and a regex-free glob matcher, through one TrueFoundry AI Gateway endpoint, against Qwen3.8-Max, GPT-5.6 Sol, and Kimi K3. We graded the generated code with real unit tests ourselves, not with another model, and ran the full suite three times per model to check whether any of it was just noise.

Qwen3.8-Max passed every task, and it still ends up the most expensive model per completed task of the three. All three models solved 10/10 tasks in most runs (Kimi K3 missed one, an empty-input edge case on the JSON-flattening task, in one of its three runs). On raw pass rate, Qwen3.8-Max looks like a clean win at a fraction of the price. Normalize by what it actually cost to get a correct answer, though, and the ranking flips. GPT-5.6 Sol landed at $0.0067 to $0.0070 per solved task across all three runs, tight and repeatable. Qwen3.8-Max landed at $0.0144 to $0.0284 per solved task: roughly2.8x GPT-5.6 Sol at the low end and over 4x at the high end, despite a 5x cheaper per-token price. Kimi K3 sat in between at $0.0084 to $0.0101.

Why? Qwen3.8-Max is dramatically more verbose on harder problems, and it occasionally spirals.Averaged across the four "hard" tasks in our suite, its completion length was 5,364 tokens per answer, against 729 for Kimi K3 and 282 for GPT-5.6 Sol: about 19x more tokens than GPT-5.6 Sol to solve the same problems, both correctly. The clearest example is the regex-free glob-matching task. One of Qwen3.8-Max's three runs generated 13,430 completion tokens and took298 seconds to answer a problem GPT-5.6 Sol solved in 3.5 seconds using roughly 200 tokens. Both got it right. That's not a rounding error, that's a five-minute coffee break versus a blink. On the JSON-flattening task, Qwen3.8-Max's completion length actually grew across our three repeated runs (6,104, then 7,752, then 13,965 tokens) while the answer stayed correct each time.The path to the right answer just kept getting longer and more expensive.

That's the piece the per-token pricing comparison misses. Six dollars per million output tokens sounds like a bargain next to GPT-5.6 Sol's thirty, until the model uses 20x more tokens to get there. Cost per correct answer, not cost per token, is the number that should drive a model-swap decision, and it isn't the number Alibaba's launch pricing leads with.

Der schnellste Weg, deine KI zu entwickeln, zu steuern und zu skalieren

Melde dich an
Inhaltsverzeichniss

Steuern, implementieren und verfolgen Sie KI in Ihrer eigenen Infrastruktur

Buchen Sie eine 30-minütige Fahrt mit unserem KI-Experte

Eine Demo buchen

Der schnellste Weg, deine KI zu entwickeln, zu steuern und zu skalieren

Demo buchen
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Entdecke mehr

July 20, 2023
|
Lesedauer: 5 Minuten

LLMops CoE: Die nächste Grenze in der MLOps-Landschaft

April 16, 2024
|
Lesedauer: 5 Minuten

Cognita: Entwicklung modularer Open-Source-RAG-Anwendungen für die Produktion

May 25, 2023
|
Lesedauer: 5 Minuten

Open-Source-LLMs: Umarmen oder untergehen

August 27, 2025
|
Lesedauer: 5 Minuten

Kartierung des KI-Marktes vor Ort: Von Chips bis zu Steuerflugzeugen

August 6, 2026
|
Lesedauer: 5 Minuten

MCP 2026-07-28 Ships: Revisiting Apps, Tasks, and Gateway Governance Under the Largest Protocol Revision Since Launch

Keine Artikel gefunden.
August 6, 2026
|
Lesedauer: 5 Minuten

We Checked Alibaba's Math: Qwen3.8-Max's Price Advantage Doesn't Survive Contact With Real Tasks

LLMs und GenAI
August 5, 2026
|
Lesedauer: 5 Minuten

Self-Evolving Agents, Governed: The Enterprise Playbook for Systems That Rewrite Themselves

Keine Artikel gefunden.
August 3, 2026
|
Lesedauer: 5 Minuten

Claude Code --dangerously-skip-permissions erklärt: Risiken, Anwendungsfälle und sicherere Alternativen

Keine Artikel gefunden.
Keine Artikel gefunden.

Aktuelle Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Machen Sie eine kurze Produkttour
Produkttour starten
Produkttour