📊 Full opportunity report: Qwen3.8-Max's Newest AI Metrics: How Do They Measure Up? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, a 2.4 trillion-parameter AI model, with detailed benchmark results and open weights set to ship next week. The model shows strong performance in multimodal and agentic tasks but trails in some software engineering benchmarks. Its open release signals a significant step for large-scale AI models.

Alibaba has confirmed the broad availability of Qwen3.8-Max, a 2.4 trillion-parameter AI model, along with detailed benchmark results and plans to ship open weights next week. This marks a significant milestone after weeks of speculation and stealth testing, and it positions Alibaba as a major player in large-scale AI development.

On August 3, Alibaba publicly released the full benchmark table for Qwen3.8-Max, confirming its architecture as a sparse mixture-of-experts model built on Qwen3.5, with approximately 95 billion active parameters per query. The model supports multimodal inputs — text, images, video — and outputs text, demonstrating strong performance in various benchmarks, especially in multimodal and agentic tasks.

The benchmark results show Qwen3.8-Max achieving top-tier scores: 86.6 on Terminal-Bench 2.1, surpassing Claude Fable 5 and only behind GPT-5.6 Sol; 93.0 on PaperBench, and 86.1 on OSWorld-Verified. These scores indicate competitive performance in both language understanding and multimodal applications. However, the model trails in certain deep software engineering benchmarks, notably scoring 67.7 on SWE-bench Pro, compared to Fable 5’s 80.0, and 73.5 on FrontierSWE, versus Fable 5’s 88.8.

Alibaba also revealed that the model significantly improves agentic performance over its predecessor, jumping from 21.6 to 56.6 on DeepSWE, and from 40.7 to 73.5 on FrontierSWE, demonstrating its enhanced ability to perform long-horizon, agent-driven tasks. The open weights, set to be released next week, will be available in a 27B configuration suitable for deployment on individual high-memory machines, although the full 2.4T model remains a datacenter artifact due to its size.

At a glance
updateWhen: announced August 3, 2023; benchmarks an…
The developmentAlibaba has officially released detailed specifications and benchmark results for Qwen3.8-Max, confirming its 2.4 trillion parameters and upcoming open weights, marking a major milestone in AI scaling.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Large-Scale Model Release

The release of Qwen3.8-Max and its benchmark details mark a pivotal moment in AI development, highlighting Alibaba’s capacity to produce a model with 2.4 trillion parameters and strong multimodal capabilities. The detailed benchmarks allow for transparent comparison with other leading models, and the open weights will enable wider access and experimentation, especially for smaller organizations and researchers. However, the model’s limitations in certain software engineering benchmarks suggest ongoing challenges in achieving universal performance across tasks.

This development underscores the increasing scale and sophistication of AI models, with potential impacts on AI deployment, research transparency, and the competitive landscape. The open release could accelerate innovation but also raises questions about licensing, licensing transparency, and the practicalities of deploying such large models.

The AI Elephant Problem: Why Every Expert Is Right and Still Incomplete (The Architecture of Complex Problems Book 7)

The AI Elephant Problem: Why Every Expert Is Right and Still Incomplete (The Architecture of Complex Problems Book 7)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Strategy

Alibaba's AI journey has been marked by stealth and strategic releases, with the company previewing Qwen3.8-Max in July during the World AI Conference in Shanghai. Prior to this, the company had teased its scale with claims of a 2.4 trillion-parameter model, but without detailed benchmarks or open weights. The recent announcement follows a pattern of incremental disclosures, culminating in the full benchmark table and upcoming open weights release.

The model was initially identified indirectly through community recognition of a tokenizer quirk, and its capabilities have been gradually revealed through benchmark tests. Alibaba’s approach contrasts with other major AI labs, which often release models with accompanying papers and licenses. Alibaba’s recent strategy emphasizes performance transparency and open access, at least for the 27B variant, while the full 2.4T model remains a datacenter asset.

"We are committed to transparency and open access, and the upcoming release of the open weights for Qwen3.8-27B will enable broader experimentation and deployment."

— Alibaba spokesperson

Optimizing Large Scale AI Workloads with NVIDIA Blackwell:: A Developer’s Guide to the B100 and GB200 Ecosystem

Optimizing Large Scale AI Workloads with NVIDIA Blackwell:: A Developer’s Guide to the B100 and GB200 Ecosystem

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Aspects of Model Licensing and Deployment

Details about the licensing terms for the 2.4 trillion-parameter weights remain unpublished, raising questions about usage rights and restrictions. It is also unclear whether the full model will be available for open-source deployment or limited to enterprise use. The performance of the 27B open-weight variant in real-world applications, especially regarding agentic capabilities, is still to be validated in broader settings.

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

LAFVIN ESP32-S3 1.69" LCD Development Board with Camera, AI Vision Voice Development Kit, Programmable IoT Board with Mic Speaker for STEM Education

  • Powerful Microcontroller: ESP32-S3 with 16MB Flash and 8MB PSRAM
  • AI Vision & Voice: Camera and audio for AI interactions
  • Multiple Algorithm Support: OpenCV, YOLO for face and pose detection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s AI Model Rollout

Alibaba plans to release the open weights for Qwen3.8-27B next week, enabling developers and researchers to test and deploy the model on high-memory hardware. The company will likely continue benchmarking and testing the 2.4T model’s performance in various environments to assess its practical viability. Monitoring community feedback and usage will be critical to understanding the model’s impact and addressing any licensing or technical issues that arise.

Local AI Engineering with Ollama: Run, understand, customize, fine-tune, and build agentic apps on your own hardware

Local AI Engineering with Ollama: Run, understand, customize, fine-tune, and build agentic apps on your own hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights for the 27B variant are scheduled to be released next week, on August 10, 2023.

How does Qwen3.8-Max compare to other large models like GPT-5.6 or Claude Fable 5?

According to Alibaba’s benchmark results, Qwen3.8-Max outperforms Claude Fable 5 and Claude Opus 4.8 on several tests but trails behind GPT-5.6 at maximum effort. Its strengths lie in multimodal and agentic benchmarks.

What are the main limitations of Qwen3.8-Max?

The model underperforms significantly on certain deep software engineering benchmarks, indicating ongoing challenges in achieving universal task competence. Its full-scale deployment is limited by size and licensing uncertainties.

Will Alibaba’s open weights be fully open-source?

The licensing details remain unpublished. While the 27B variant will be available for broader use, the full 2.4T model’s open-source status is still uncertain.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Retirement Care Planner

A new digital tool is being tested to help adult children plan care and finances for aging parents, addressing a growing demographic challenge.

Apertus. The architectural template.

Apertus, developed by Swiss federal research institutions, introduces a novel open, multilingual, compliance-focused AI model as a blueprint for European sovereignty.

Quiet GPUs for Local AI: Acoustic and Thermal Roundup

An in-depth review of the quietest and coolest GPUs for local AI in 2026, focusing on acoustic performance, thermal management, and optimal configurations.

Top 10 AI Startups to Follow in 2026

A curated list of the most promising AI startups to follow in 2026, highlighting innovation, investment, and industry impact.