Blank white background with no objects or features visible.

Te presentamos TrueForge: el entorno de agentes de código abierto y neutral respecto a proveedores. Un 50% menos de coste. Explorar ahora→

We Checked Alibaba's Math: Qwen3.8-Max's Price Advantage Doesn't Survive Contact With Real Tasks

Por Amrutha Potluri

Published: August 6, 2026

On August 3, Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter model (95B active per inference) priced at $2/$6 per million input/output tokens, roughly 5x cheaper than GPT-5.6 Solon output. Alongside it, Alibaba published five benchmark scores, including a jump onFrontierSWE from 40.7 to 73.5 versus its predecessor.

We ran an independent check ourselves. Ten self-contained coding tasks, a separate hand-written suite (not a reproduction of FrontierSWE or DeepSWE) spanning easy string manipulation up to a dependency-order resolver and a regex-free glob matcher, through one TrueFoundry AI Gateway endpoint, against Qwen3.8-Max, GPT-5.6 Sol, and Kimi K3. We graded the generated code with real unit tests ourselves, not with another model, and ran the full suite three times per model to check whether any of it was just noise.

Qwen3.8-Max passed every task, and it still ends up the most expensive model per completed task of the three. All three models solved 10/10 tasks in most runs (Kimi K3 missed one, an empty-input edge case on the JSON-flattening task, in one of its three runs). On raw pass rate, Qwen3.8-Max looks like a clean win at a fraction of the price. Normalize by what it actually cost to get a correct answer, though, and the ranking flips. GPT-5.6 Sol landed at $0.0067 to $0.0070 per solved task across all three runs, tight and repeatable. Qwen3.8-Max landed at $0.0144 to $0.0284 per solved task: roughly2.8x GPT-5.6 Sol at the low end and over 4x at the high end, despite a 5x cheaper per-token price. Kimi K3 sat in between at $0.0084 to $0.0101.

Why? Qwen3.8-Max is dramatically more verbose on harder problems, and it occasionally spirals.Averaged across the four "hard" tasks in our suite, its completion length was 5,364 tokens per answer, against 729 for Kimi K3 and 282 for GPT-5.6 Sol: about 19x more tokens than GPT-5.6 Sol to solve the same problems, both correctly. The clearest example is the regex-free glob-matching task. One of Qwen3.8-Max's three runs generated 13,430 completion tokens and took298 seconds to answer a problem GPT-5.6 Sol solved in 3.5 seconds using roughly 200 tokens. Both got it right. That's not a rounding error, that's a five-minute coffee break versus a blink. On the JSON-flattening task, Qwen3.8-Max's completion length actually grew across our three repeated runs (6,104, then 7,752, then 13,965 tokens) while the answer stayed correct each time.The path to the right answer just kept getting longer and more expensive.

That's the piece the per-token pricing comparison misses. Six dollars per million output tokens sounds like a bargain next to GPT-5.6 Sol's thirty, until the model uses 20x more tokens to get there. Cost per correct answer, not cost per token, is the number that should drive a model-swap decision, and it isn't the number Alibaba's launch pricing leads with.

Try now.

One gateway for all your models, MCP servers, and agents.
No credit card needed.

Inscríbase
Tabla de contenido

Controle, implemente y rastree la IA en su propia infraestructura

Reserva 30 minutos con nuestro Experto en IA

Reserve una demostración

La forma más rápida de crear, gobernar y escalar su IA

Demo del libro
Summarize with
ChatGPT logo by OpenAI
Perplexity AI logo
Blurry red snowflake on white background, symmetrical frosty design with soft edges and abstract shape.

Descubra más

August 26, 2026
|
5 minutos de lectura

Gemini 3 Pro: Benchmarks and How to Use It via Gateway

July 20, 2023
|
5 minutos de lectura

LLMOps CoE: la próxima frontera en el panorama de los MLOps

April 16, 2024
|
5 minutos de lectura

Cognita: Creación de aplicaciones RAG modulares y de código abierto para la producción

May 25, 2023
|
5 minutos de lectura

LLM de código abierto: abrazar o perecer

Your guide to understanding key aspects of Bifrost pricing
September 11, 2026
|
5 minutos de lectura

Bifrost Pricing: OSS, Enterprise Costs, and What Teams Should Know

No se ha encontrado ningún artículo.
TrueFoundry AI gateway alternative to Requesty AI pricing
September 11, 2026
|
5 minutos de lectura

Requesty AI Pricing: Cost, Features, and Enterprise Fit Explained

No se ha encontrado ningún artículo.
TrueFoundry AI gateway alternative to Solo.io for enterprises
September 11, 2026
|
5 minutos de lectura

Top 5 Solo.io Competitors and Alternatives for 2026

No se ha encontrado ningún artículo.
September 11, 2026
|
5 minutos de lectura

What EU SOC teams should ask before trusting AI on security logs

LLMS y GenAI
No se ha encontrado ningún artículo.

Blogs recientes

Black left pointing arrow symbol on white background, directional indicator.
Black left pointing arrow symbol on white background, directional indicator.
Realice un recorrido rápido por el producto
Comience el recorrido por el producto
Visita guiada por el producto