Blank white background with no objects or features visible.

TrueFoundry anuncia la adquisición de Seldon AI, ampliando su plataforma de control para IA empresarial. Lea el informe completo →

We Checked Alibaba's Math: Qwen3.8-Max's Price Advantage Doesn't Survive Contact With Real Tasks

Por Amrutha Potluri

Published: August 6, 2026

On August 3, Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter model (95B active per inference) priced at $2/$6 per million input/output tokens, roughly 5x cheaper than GPT-5.6 Solon output. Alongside it, Alibaba published five benchmark scores, including a jump onFrontierSWE from 40.7 to 73.5 versus its predecessor.

We ran an independent check ourselves. Ten self-contained coding tasks, a separate hand-written suite (not a reproduction of FrontierSWE or DeepSWE) spanning easy string manipulation up to a dependency-order resolver and a regex-free glob matcher, through one TrueFoundry AI Gateway endpoint, against Qwen3.8-Max, GPT-5.6 Sol, and Kimi K3. We graded the generated code with real unit tests ourselves, not with another model, and ran the full suite three times per model to check whether any of it was just noise.

Qwen3.8-Max passed every task, and it still ends up the most expensive model per completed task of the three. All three models solved 10/10 tasks in most runs (Kimi K3 missed one, an empty-input edge case on the JSON-flattening task, in one of its three runs). On raw pass rate, Qwen3.8-Max looks like a clean win at a fraction of the price. Normalize by what it actually cost to get a correct answer, though, and the ranking flips. GPT-5.6 Sol landed at $0.0067 to $0.0070 per solved task across all three runs, tight and repeatable. Qwen3.8-Max landed at $0.0144 to $0.0284 per solved task: roughly2.8x GPT-5.6 Sol at the low end and over 4x at the high end, despite a 5x cheaper per-token price. Kimi K3 sat in between at $0.0084 to $0.0101.

Why? Qwen3.8-Max is dramatically more verbose on harder problems, and it occasionally spirals.Averaged across the four "hard" tasks in our suite, its completion length was 5,364 tokens per answer, against 729 for Kimi K3 and 282 for GPT-5.6 Sol: about 19x more tokens than GPT-5.6 Sol to solve the same problems, both correctly. The clearest example is the regex-free glob-matching task. One of Qwen3.8-Max's three runs generated 13,430 completion tokens and took298 seconds to answer a problem GPT-5.6 Sol solved in 3.5 seconds using roughly 200 tokens. Both got it right. That's not a rounding error, that's a five-minute coffee break versus a blink. On the JSON-flattening task, Qwen3.8-Max's completion length actually grew across our three repeated runs (6,104, then 7,752, then 13,965 tokens) while the answer stayed correct each time.The path to the right answer just kept getting longer and more expensive.

That's the piece the per-token pricing comparison misses. Six dollars per million output tokens sounds like a bargain next to GPT-5.6 Sol's thirty, until the model uses 20x more tokens to get there. Cost per correct answer, not cost per token, is the number that should drive a model-swap decision, and it isn't the number Alibaba's launch pricing leads with.

La forma más rápida de crear, gobernar y escalar su IA

Inscríbase
Tabla de contenido

Controle, implemente y rastree la IA en su propia infraestructura

Reserva 30 minutos con nuestro Experto en IA

Reserve una demostración

La forma más rápida de crear, gobernar y escalar su IA

Demo del libro
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Descubra más

July 20, 2023
|
5 minutos de lectura

LLMOps CoE: la próxima frontera en el panorama de los MLOps

April 16, 2024
|
5 minutos de lectura

Cognita: Creación de aplicaciones RAG modulares y de código abierto para la producción

May 25, 2023
|
5 minutos de lectura

LLM de código abierto: abrazar o perecer

August 27, 2025
|
5 minutos de lectura

Mapeando el mercado de la IA local: desde chips hasta aviones de control

August 6, 2026
|
5 minutos de lectura

MCP 2026-07-28 Ships: Revisiting Apps, Tasks, and Gateway Governance Under the Largest Protocol Revision Since Launch

No se ha encontrado ningún artículo.
August 6, 2026
|
5 minutos de lectura

We Checked Alibaba's Math: Qwen3.8-Max's Price Advantage Doesn't Survive Contact With Real Tasks

LLMS y GenAI
August 5, 2026
|
5 minutos de lectura

Self-Evolving Agents, Governed: The Enterprise Playbook for Systems That Rewrite Themselves

No se ha encontrado ningún artículo.
August 3, 2026
|
5 minutos de lectura

Explicación del Código Claude: saltarse peligrosamente los permisos: riesgos, casos de uso y alternativas más seguras

No se ha encontrado ningún artículo.
No se ha encontrado ningún artículo.

Blogs recientes

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Realice un recorrido rápido por el producto
Comience el recorrido por el producto
Visita guiada por el producto