Blank white background with no objects or features visible.

TrueFoundry Named Frost & Sullivan's 2026 Global Transformational Innovation Leader. Read report

Lernen Sie TrueForge kennen: Das Open-Source- und herstellerneutrale Agent Harness. 50 % geringere Kosten. Jetzt entdecken→

We Checked Alibaba's Math: Qwen3.8-Max's Price Advantage Doesn't Survive Contact With Real Tasks

von Amrutha Potluri

Published: August 6, 2026

On August 3, Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter model (95B active per inference) priced at $2/$6 per million input/output tokens, roughly 5x cheaper than GPT-5.6 Solon output. Alongside it, Alibaba published five benchmark scores, including a jump onFrontierSWE from 40.7 to 73.5 versus its predecessor.

We ran an independent check ourselves. Ten self-contained coding tasks, a separate hand-written suite (not a reproduction of FrontierSWE or DeepSWE) spanning easy string manipulation up to a dependency-order resolver and a regex-free glob matcher, through one TrueFoundry AI Gateway endpoint, against Qwen3.8-Max, GPT-5.6 Sol, and Kimi K3. We graded the generated code with real unit tests ourselves, not with another model, and ran the full suite three times per model to check whether any of it was just noise.

Qwen3.8-Max passed every task, and it still ends up the most expensive model per completed task of the three. All three models solved 10/10 tasks in most runs (Kimi K3 missed one, an empty-input edge case on the JSON-flattening task, in one of its three runs). On raw pass rate, Qwen3.8-Max looks like a clean win at a fraction of the price. Normalize by what it actually cost to get a correct answer, though, and the ranking flips. GPT-5.6 Sol landed at $0.0067 to $0.0070 per solved task across all three runs, tight and repeatable. Qwen3.8-Max landed at $0.0144 to $0.0284 per solved task: roughly2.8x GPT-5.6 Sol at the low end and over 4x at the high end, despite a 5x cheaper per-token price. Kimi K3 sat in between at $0.0084 to $0.0101.

Why? Qwen3.8-Max is dramatically more verbose on harder problems, and it occasionally spirals.Averaged across the four "hard" tasks in our suite, its completion length was 5,364 tokens per answer, against 729 for Kimi K3 and 282 for GPT-5.6 Sol: about 19x more tokens than GPT-5.6 Sol to solve the same problems, both correctly. The clearest example is the regex-free glob-matching task. One of Qwen3.8-Max's three runs generated 13,430 completion tokens and took298 seconds to answer a problem GPT-5.6 Sol solved in 3.5 seconds using roughly 200 tokens. Both got it right. That's not a rounding error, that's a five-minute coffee break versus a blink. On the JSON-flattening task, Qwen3.8-Max's completion length actually grew across our three repeated runs (6,104, then 7,752, then 13,965 tokens) while the answer stayed correct each time.The path to the right answer just kept getting longer and more expensive.

That's the piece the per-token pricing comparison misses. Six dollars per million output tokens sounds like a bargain next to GPT-5.6 Sol's thirty, until the model uses 20x more tokens to get there. Cost per correct answer, not cost per token, is the number that should drive a model-swap decision, and it isn't the number Alibaba's launch pricing leads with.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Melde dich an
Inhaltsverzeichniss

Steuern, implementieren und verfolgen Sie KI in Ihrer eigenen Infrastruktur

Buchen Sie eine 30-minütige Fahrt mit unserem KI-Experte

Eine Demo buchen

Der schnellste Weg, deine KI zu entwickeln, zu steuern und zu skalieren

Demo buchen
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Entdecke mehr

August 26, 2026
|
Lesedauer: 5 Minuten

Gemini 3 Pro: Benchmarks and How to Use It via Gateway

July 20, 2023
|
Lesedauer: 5 Minuten

LLMops CoE: Die nächste Grenze in der MLOps-Landschaft

April 16, 2024
|
Lesedauer: 5 Minuten

Cognita: Entwicklung modularer Open-Source-RAG-Anwendungen für die Produktion

May 25, 2023
|
Lesedauer: 5 Minuten

Open-Source-LLMs: Umarmen oder untergehen

September 19, 2026
|
Lesedauer: 5 Minuten

MCP Tool Approvals, Explained: From Pending Call to Bounded Human Decision

Keine Artikel gefunden.
September 19, 2026
|
Lesedauer: 5 Minuten

HubSpot MCP Server: Tools, Privacy, and How to Connect It Safely

Keine Artikel gefunden.
September 19, 2026
|
Lesedauer: 5 Minuten

Snowflake MCP Server: Tools, Setup, and Cost Control

Keine Artikel gefunden.
September 18, 2026
|
Lesedauer: 5 Minuten

TypeSafe AI's Jev and "System One Models": What Actually Shipped

Agentische KI
Keine Artikel gefunden.

Aktuelle Blogs

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Machen Sie eine kurze Produkttour
Produkttour starten
Produkttour