📊 Full opportunity report: Second Only To Fable 5: Qwen3.8-Max Finally Shows Its Numbers — And The Claim Gets Complicated on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba’s Qwen3.8-Max has disclosed its benchmark results, confirming it as the second-leading model after Fable 5. The model features 2.4 trillion parameters and excels in multimodal tasks, with open weights expected next week.

Alibaba has officially released the detailed benchmark results for its Qwen3.8-Max model, confirming it as the second most powerful model globally after Fable 5. This development follows weeks of speculation and stealth previews, making it a significant milestone in AI model deployment and transparency.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, revealing it has 2.4 trillion parameters and is built on the Qwen3.5 architecture with sparse mixture-of-experts. The model demonstrates top-tier performance on several benchmarks, including Terminal-Bench 2.1 (86.6), surpassing Claude models and only trailing GPT-5.6 Sol at 88.8. It also leads in PaperBench (93.0) and performs well in multimodal and agentic tasks, such as OSWorld-Verified (86.1) and Parametric CAD Bench (91.5).

Alibaba confirmed that open weights for Qwen3.8-Max will ship next week, alongside a smaller 27-billion-parameter checkpoint, Qwen3.8-27B, optimized for local deployment. The model’s active parameters are approximately 95 billion, with the total size at 2.4 trillion, indicating a model that requires multi-node datacenter infrastructure for hosting. The release of the full benchmark data marks a shift from stealth preview to transparency, with the model now broadly accessible via API and open weights expected soon.

At a glance
updateWhen: announced August 3, 2023; benchmarks re…
The developmentAlibaba announced full benchmark results for Qwen3.8-Max, confirming its position as the second most powerful model after Fable 5, with open weights to follow.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Disclosure

This announcement signifies a major step in AI transparency and competition. Alibaba’s disclosure of detailed benchmark results and open weights demonstrates a move toward greater openness in large language model deployment. The model’s high performance in key benchmarks and agentic capabilities suggests it could influence enterprise AI applications and challenge existing leaders like OpenAI and Anthropic. The release of open weights also expands access for researchers and developers, potentially accelerating innovation and adoption.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Development

Over the past two weeks, Alibaba’s AI model activity was largely stealthy, with the model known as kaleb appearing on leaderboards and later confirmed as Qwen3.8-Max during the World AI Conference in Shanghai. The company had previously teased a 2.4 trillion-parameter model but withheld detailed benchmarks until now. The model's preview was accessible via a paid endpoint, sparking speculation about its capabilities and positioning in the market. The recent disclosure follows a pattern of Alibaba gradually revealing its progress, culminating in today’s comprehensive benchmark release.

"We are committed to open AI development and will ship open weights next week, enabling broader access and innovation."

— Alibaba spokesperson

Amazon

high performance AI server

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Deployment and Licensing

While benchmark results are now confirmed, details about the licensing terms for the open weights remain unpublished. It is also unclear whether the 2.4 trillion-parameter model will be fully open-source or subject to restrictions, as previous Alibaba models have varied licenses. Additionally, the performance of the 27B checkpoint in practical, local deployment settings has yet to be demonstrated, and the long-term stability of agentic capabilities remains to be seen.

Amazon

multimodal AI development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s Model Release and Adoption

Alibaba will release the open weights for Qwen3.8-Max next week, enabling researchers and developers to test and deploy the model locally. Monitoring how the model performs in real-world applications, especially in agentic tasks, will be key. Further benchmark results for the 27B checkpoint are expected, along with potential updates on licensing terms. Industry analysts will also watch for how competitors respond and whether Alibaba’s transparency influences market dynamics.

Amazon

AI model training GPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the key performance metrics of Qwen3.8-Max?

Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1, 93.0 on PaperBench, and 86.1 on OSWorld-Verified, among other benchmarks, placing it second only to Fable 5 in overall performance.

Will the open weights be freely available for download?

Alibaba announced that open weights for Qwen3.8-Max will ship next week, but the licensing terms remain unpublished. It is expected that the weights will be accessible, though restrictions may apply depending on licensing decisions.

How does Qwen3.8-Max compare to other models in agentic tasks?

The model shows significant improvements in agentic benchmarks, with scores rising from 21.6 to 56.6 on DeepSWE and from 40.7 to 73.5 on FrontierSWE, indicating enhanced long-horizon reasoning and task execution capabilities.

What are the implications for AI research and deployment?

The detailed benchmark disclosure and upcoming open weights could accelerate AI research, democratize access, and challenge existing market leaders, fostering more competition and innovation in large language models.

Source: ThorstenMeyerAI.com

You May Also Like

Why Digital Proofing Is Becoming Essential for Fine Art Printing

Digital proofing is becoming essential in fine art printing because it helps…

A Skill Is A Folder, Not A Prompt: What Anthropic Learned Running Hundreds Of Them

Anthropic reveals that effective AI skills are structured as folders containing instructions, scripts, and assets, transforming ad-hoc prompting into institutional capability.

I designed a nibble-oriented CPU in Verilog to build a scientific calculator

A programmer has created a custom FPGA-based scientific calculator using a unique nibble-oriented CPU designed in Verilog, with full hardware verification tools.

Fil-C: Garbage In, Memory Safety Out [Video]

Fil-C showcases a new garbage collection approach to improve memory safety in programming, emphasizing the importance of memory management techniques.