Qwen3.8-Max's AI Power: Analyzing The Latest Performance Figures

📊 Full opportunity report: Qwen3.8-Max's AI Power: Analyzing The Latest Performance Figures on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has publicly released comprehensive performance figures for its Qwen3.8-Max model, confirming a 2.4 trillion-parameter size and high benchmark scores. The open weights will be available next week, making it the largest open-weight AI model to date. This development highlights Alibaba’s advancing AI capabilities and strategic focus on open deployment.

Alibaba has officially released detailed benchmark data for its Qwen3.8-Max model, confirming it as the largest open-weight AI model to date with 2.4 trillion parameters. This marks a significant milestone in AI development, demonstrating Alibaba’s advancing capabilities in multimodal and agentic AI performance. The company also announced that open weights will be available next week, enabling broader access for researchers and developers.

Alibaba’s Qwen3.8-Max, built on the Qwen3.5 architecture and employing sparse mixture-of-experts, has been confirmed to contain approximately 2.4 trillion total parameters. The model features a 95 billion active parameter count per query, with multimodal input capabilities including text, images, and videos, and outputs in text. The benchmark results, obtained using Alibaba’s own testing harness, show that Qwen3.8-Max achieves a top score of 86.6 on Terminal-Bench 2.1, surpassing several competitors such as Claude Opus 4.8 and Claude Fable 5, but slightly behind GPT-5.6 Sol at 88.8.

Additional tests indicate strong performance in multimodal and agentic tasks, with scores like 86.1 on OSWorld-Verified and 91.5 on Parametric CAD Bench. Notably, the model demonstrated the ability to reproduce research results and outperform its own previous iterations in long-horizon agent tasks, marking a significant leap in practical AI application. However, it trails notably on software engineering benchmarks, such as SWE-bench Pro and FrontierSWE, with gaps of over 12 points compared to Fable 5, indicating areas for future improvement.

The open weights, due next week, will be a 2.4 trillion-parameter checkpoint, but due to their size, they will require multi-node datacenter infrastructure to run. Alibaba also introduced a smaller 27B version, Qwen3.8-27B, optimized for deployment on high-memory single machines, with performance results still to be published.

At a glance
updateWhen: announced August 3, 2023; benchmarks pu…
The developmentAlibaba announced full benchmark results and confirmed open weights for Qwen3.8-Max, establishing it as a leading large-scale model with significant performance metrics.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Release and Open Model Strategy

This announcement confirms Alibaba's position as a major player in large-scale AI development, with a model that rivals the biggest in the industry in terms of parameters and performance. The full benchmark disclosure provides transparency and sets new standards for multimodal and agentic AI capabilities, especially as the open weights become available next week.

For the AI community and industry, the open release of such a large model could accelerate research, foster innovation, and challenge existing market leaders. It also signals Alibaba’s strategic move toward open deployment, potentially reshaping competitive dynamics in AI development and deployment.

Accelerate Everything with Tensor Cores: A Developer’s Guide to High-Performance AI, Efficient Training, and Scalable Models

Accelerate Everything with Tensor Cores: A Developer’s Guide to High-Performance AI, Efficient Training, and Scalable Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Alibaba’s Large-Scale Model Development

Over the past few weeks, Alibaba has built anticipation around its latest model, initially revealing only a slogan claiming it was 'second only to Fable 5' with 2.4 trillion parameters. The model, later identified as Qwen3.8-Max, was previewed in stealth during the July World AI Conference and briefly appeared on the Code Arena leaderboard as 'kaleb.'

Previous models like Kimi K3 and other large-scale efforts from competitors such as Moonshot's 2.8 trillion-parameter Kimi K3 set the stage for this reveal. Alibaba’s strategic approach involved a staged rollout, culminating in the detailed benchmark publication on August 3, confirming the model’s specifications and performance metrics. The company also announced that the open weights would be released next week, marking a significant milestone in open AI deployment.

"Qwen3.8-Max sets new standards in multimodal AI performance, and the open weights will enable broader innovation across the AI community."

— Alibaba spokesperson

MULTI-TECH SYSTEMS 4PORT V.34 Fax Server in/Out Us/Canada

MULTI-TECH SYSTEMS 4PORT V.34 Fax Server in/Out Us/Canada

  • 4-Port V.34 Fax Server: MultiTech Systems 4-port fax server
  • DTMF Routing: Uses DTMF to route faxes to email
  • File Conversion: Converts faxes to PDF or TIFF

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Licensing and Deployment

It remains unclear what the licensing terms will be for the open weights, as Alibaba has not yet published the license details. Historically, Alibaba’s open models have used the Apache 2.0 license, but the upcoming release might differ, especially given the model’s size and potential commercial implications.

Additionally, the performance of the 27B version on practical deployment scenarios and its ability to retain agentic capabilities after compression are still to be confirmed. The benchmark results for this smaller model are not yet available, leaving some uncertainty about its real-world utility.

Artificial Intelligence: AI Engineer's Cheatsheet: Silicon Edition (Ultra-large scale LLM training and inference)

Artificial Intelligence: AI Engineer's Cheatsheet: Silicon Edition (Ultra-large scale LLM training and inference)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Week’s Open Weight Release and Industry Impact

The primary focus will be the release of the 2.4 trillion-parameter open weights next week, which could significantly influence AI research and deployment strategies. Researchers and developers will evaluate the model’s performance in real-world applications, especially in multimodal and agentic contexts.

Further benchmark results for the 27B version are expected, alongside analyses of licensing terms. Industry observers will also watch for how competitors respond and whether Alibaba’s open approach accelerates AI innovation globally.

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled to be released next week, with the exact date to be announced by Alibaba.

How does Qwen3.8-Max compare to other large models like GPT-5.6 or Claude Fable 5?

Qwen3.8-Max scores higher than Claude Fable 5 on several benchmarks such as Terminal-Bench 2.1 but slightly trails GPT-5.6 at 88.8 on the same test. Its multimodal and agentic capabilities are among the strongest in the industry.

What are the potential uses for the open weights once released?

The open weights will enable researchers and companies to deploy large-scale multimodal and agentic AI models, fostering innovation in areas like research automation, multimodal understanding, and autonomous systems.

Will the licensing terms for the open weights be restrictive?

The licensing details have not yet been published. Historically, Alibaba’s open models have used Apache 2.0, but the upcoming release may have different terms, which could impact how the model is used commercially or in research.

Source: ThorstenMeyerAI.com

You May Also Like

Electric Code Calculator

New electric code calculator aims to streamline NEC calculations for electricians, offering offline access and current code compliance.

A Walk Through Of The DeltaNet Family Of Linear Attention Variants

A detailed overview of DeltaNet’s linear attention variants, highlighting confirmed features, potential benefits, and ongoing developments in efficient neural network models.

Smart Studio Headphones: The AI Revolution In 2026

In 2026, AI-integrated studio headphones revolutionize audio production, offering real-time mixing, noise cancellation, and personalized sound profiles. Confirmed developments and ongoing questions explained.

Software engineering. The canonical case.

Empirical data shows a 40% drop in junior hiring, while senior engineers benefit from AI augmentation. The sector reveals a bifurcated impact of AI on labor.