
Qwen 3.8 Max Review: Alibaba's 2.4T Model, Tested
Quick answer:
Qwen3.8-Max-Preview is Alibaba's newest flagship model, a 2.4-trillion-parameter multimodal system previewed on 19 July 2026 at the World AI Conference in Shanghai, which Alibaba calls "second only to Fable 5" among the models it benchmarked. That ranking rests entirely on Alibaba's own internal evaluation — no benchmark table, model card or independent third-party score has been published. The active-parameter count is undisclosed, open weights are promised only as "coming soon," and access for now runs through Alibaba's Token Plan subscriptions, from $6 for 2,500 credits a week up to $68 for 40,000 credits and 6-8 concurrent agents. Treat every headline figure in this launch as provisional until independent evaluators weigh in.
Two days after Moonshot AI's Kimi K3 rattled global chip stocks, Alibaba walked onto the stage at WAIC in Shanghai with a bigger number of its own. Qwen3.8-Max-Preview, the company said, is a 2.4-trillion-parameter multimodal model that trails only Claude Fable 5. It is a striking claim from a serious lab, made in the same week that a rival Chinese model briefly tipped the Philadelphia Semiconductor Index into a bear market. It is also, as of publication, entirely unverified.
That is not a criticism so much as a description of where things stand: the preview is real and purchasable, the parameter count is Alibaba's own figure, and the ranking against Fable 5 has not been checked by anyone outside Alibaba. This article separates what Alibaba has confirmed from what it has only claimed, walks through what is actually known about the model's architecture and pricing, and places the launch in the context of a Chinese open-weight race that produced three trillion-parameter-class releases in barely a month.
Note: this analysis is based on Alibaba/Qwen's own announcement on X, reporting and interactive fact-checking from MarkTechPost, the-decoder, SCMP, OfficeChai and MLQ.ai, and independent benchmark data from Artificial Analysis for Qwen's prior flagship, Qwen3.7-Max. AI Tools Review has not independently reproduced any benchmark, and no independent scores for Qwen3.8-Max-Preview itself existed at the time of writing. Prices are billed in US dollars.
Executive summary
- 2.4 trillion total parameters — Alibaba's own figure, previewed on 19 July 2026 during the World AI Conference in Shanghai, two days after Moonshot AI's 2.8-trillion-parameter Kimi K3.
- "Second only to Fable 5," per Alibaba's official Qwen account on X — a claim with no accompanying benchmark table, model card, or independent verification at preview time.
- Alibaba's first multimodal model above 1 trillion parameters, according to Qwen developer Shuai Bai, processing text, images, video and documents.
- Active-parameter count undisclosed. For a sparse Mixture-of-Experts model this is the single number that determines real inference cost, and Alibaba has not published it.
- No open weights yet. Alibaba says open-weight release is "coming soon," but its previous two Max-tier flagships also launched closed and have not since been open-sourced.
- Bundled pricing, not per-token pricing. Access runs through Token Plan subscriptions from $6 (2,500 credits/7 days) to $68 (40,000 credits/7 days, 6-8 concurrent agents), at a preview discount of 10% off standard rates.
- Arrives amid a genuine Chinese open-weight surge. GLM 5.2, Kimi K3 and now Qwen 3.8 have each launched within about a month of one another, each claiming a place near the top of the global leaderboard.
From Qwen3-Max to Qwen 3.8: the lineage
Alibaba's Qwen team crossed the trillion-parameter threshold for the first time with Qwen3-Max-Preview in September 2025. Qwen 3.6 Max Preview followed in April 2026, and — notably — stayed closed, hosted only through Qwen Studio and Alibaba Cloud rather than released as open weights on day one, a departure from Alibaba's earlier practice of shipping many Qwen models openly from launch. Qwen3.7-Max arrived in May 2026 as an incremental step on the same closed-weight, reasoning-focused path, adding a 1-million-token context window and posting strong scores on agentic and coding evaluations: 92.4 on GPQA-Diamond, 80.4% on SWE-bench Verified, and 69.7 on Terminal-Bench 2.0-Terminus, all figures Alibaba published and that independent trackers such as Artificial Analysis have broadly corroborated (Qwen3.7-Max reached 56.6 on the Artificial Analysis Intelligence Index, fifth overall at the time).
Qwen3.8-Max-Preview, announced less than three months after Qwen 3.6 Max Preview, is a much bigger jump than the version number suggests: total parameters more than double again, to 2.4 trillion, and Alibaba is pitching it as the team's first multimodal model above 1 trillion parameters — a genuine architectural step, not just a scale-up of the 3.7 generation. For now it is following the same closed-preview pattern as its two immediate predecessors: live and purchasable through the Token Plan, with open weights promised but not delivered.
The timing is not incidental. Alibaba owns roughly 36% of Moonshot AI, the Beijing lab behind Kimi K3, which means Qwen 3.8's launch is as much an in-house rivalry as a competitive response — two of Alibaba's own portfolio companies trading blows for the title of strongest Chinese model within the same week. It also lands one month after Z.AI's GLM 5.2, which drew praise from engineers at Vercel and Box for landing within striking distance of Claude Opus 4.8 on FrontierSWE. Three trillion-parameter-class Chinese releases in barely a month is the real story behind any single one of them.
Architecture and training: what is and isn't known
What Alibaba has actually confirmed about Qwen3.8-Max-Preview is narrower than the headline number suggests. Qwen developer Shuai Bai described it as a sparse Mixture-of-Experts model and the team's first multimodal system above 1 trillion parameters, capable of processing text, images, video and documents. Alibaba says it should outperform Qwen3.7-Max specifically on coding, full-stack development, data analysis and office workflows, and the preview is compatible with both OpenAI and Anthropic API protocols, meaning existing coding agents and harnesses can point at it without being rebuilt.

Beyond that, the specifics that would normally accompany a flagship launch are simply absent. There is no model card, no specification sheet, and — critically for a sparse Mixture-of-Experts design — no disclosed active-parameter count. That figure is not a footnote: it is the number that determines how much compute a single forward pass actually costs, and Qwen's own smaller models show how differently "2.4 trillion parameters" can translate into real serving cost depending on sparsity. Qwen3-235B-A22B, for instance, carries 235 billion total parameters but activates only 22 billion per token; Qwen3-30B-A3B activates roughly 3 billion of its 30 billion. Without an equivalent ratio for the Max-tier 2.4T model, the total parameter count alone says very little about what it costs Alibaba — or would cost anyone else — to run it.
The scale also raises a practical question independent of pricing: if the full 2.4-trillion-parameter model does eventually ship as open weights, loading it would require roughly 1.2 terabytes of storage for the weights alone at 4-bit precision, before accounting for the KV cache and runtime overhead — well beyond what even multiple high-end Nvidia H200 GPUs (141GB of memory each) could hold without a cluster. Whether Alibaba ultimately releases the full model, a smaller activated-parameter variant, a quantised checkpoint, or a distilled sibling remains unknown, and materially changes who could realistically self-host it.
Capabilities deep dive
Multimodal input
Alibaba's headline capability claim is multimodality at scale: Qwen3.8-Max-Preview is described as the Qwen team's first model above 1 trillion parameters that can process images, video and documents alongside text. That positions it against Fable 5 and GPT-5.6 Sol on general multimodal reasoning, though — as with the rest of the launch — no multimodal benchmark scores accompanied the announcement.
Coding, full-stack development and office workflows
Alibaba's stated focus for the improvement over Qwen3.7-Max is squarely practical: coding, full-stack development, data analysis and office productivity tasks, rather than a broad claim of general intelligence gains. Combined with OpenAI- and Anthropic-protocol compatibility, this reads as a direct pitch to teams already running agentic coding workflows on other models, inviting them to swap in Qwen 3.8 without rebuilding their tooling.
The Token Plan's concurrent-agent angle
The most concrete agentic detail Alibaba has published is not a benchmark but a capacity figure: the top Token Plan tier is specified to support six to eight concurrent agents. That is a genuinely useful data point for teams evaluating whether the plan can support a multi-agent coding or automation setup, even though it says nothing about how capable any individual agent run actually is. It is also the kind of infrastructure-level claim that is easy to verify directly, unlike a benchmark ranking that requires Alibaba's own harness to reproduce.
Benchmarks: confirmed vs claimed
This is the section where the Qwen 3.8 launch differs most from a typical frontier-model release, and it is worth being blunt about it: there is no benchmark table. Alibaba's official Qwen account posted on X that the model is "one of the most powerful available today, comparable to leading frontier AI models, second only to Fable 5" — but that sentence is the entire public evidence base for the claim. No scores, no named test suite, and no methodology accompanied it.
| Claim | Status as of 20 July 2026 |
|---|---|
| Preview live and purchasable via Token Plan, Qoder, QoderWork | Confirmed |
| Sparse MoE, first Qwen multimodal model above 1T parameters | Confirmed (Shuai Bai) |
| OpenAI and Anthropic protocol compatibility | Confirmed |
| 2.4 trillion total parameters | Alibaba's own figure — no model card |
| "Second only to Fable 5" | Unverified — no benchmark table published |
| Open weights "coming soon" | No date, licence, or repository |
| Active parameters per token | Undisclosed |
The closest thing to an independent reference point comes from Qwen's own previous flagship. Qwen3.7-Max — closed weights, launched May 2026 — scored 56.6 on Artificial Analysis's Intelligence Index, placing it fifth overall and ahead of Google's Gemini 3.5 Flash at the time, alongside published scores of 92.4 on GPQA-Diamond, 80.4% on SWE-bench Verified and 69.7 on Terminal-Bench 2.0-Terminus. Every one of those is a Qwen3.7-Max number, not a Qwen 3.8 number; until Alibaba publishes its own benchmark table, that is the closest verified baseline available, and it is not evidence for or against the "second only to Fable 5" claim.
This is not an unprecedented pattern. Moonshot AI made a very similar claim when it first announced Kimi K3 days earlier — that it trailed only Fable 5 and GPT-5.6 Sol — and that claim was also unverified at launch. Artificial Analysis subsequently scored K3 at 57 on the Intelligence Index, competitive with but not ahead of several other frontier models, which is a useful reminder that early vendor rankings tend to be optimistic relative to what independent testing later confirms. See our full Kimi K3 review for how that claim held up.
Competitive landscape
Qwen 3.8's launch cannot be read in isolation from the week that preceded it. On 17 July 2026, Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight model that briefly held the title of the largest open-source model ever released, and which beat both Claude Fable 5 and GPT-5.6 Sol outright on Arena.ai's frontend-development leaderboard — the first open-weight model to top that particular ranking. The reaction was not confined to AI commentary: South Korea's KOSPI fell more than 6% and Japan's Nikkei dropped over 4% in the days that followed, the Philadelphia Semiconductor Index fell into bear-market territory (down more than 20% from its late-June peak, its worst week in over fifteen months), and US chip and AI-adjacent names including Intel, Micron, AMD and Marvell all declined. Analysts were careful to note the selloff was "overdetermined" — weak Netflix and TSMC earnings, the Iran war, and broader risk-off sentiment all contributed — but Apollo's Torsten Sloek specifically flagged the risk that cheap or free capable AI from Chinese open models could undercut the roughly $700 billion in annual hyperscaler AI infrastructure spending that markets are pricing in as a return on investment.
Qwen 3.8 arrived directly into that atmosphere, two days later, at Alibaba's home-turf conference. OfficeChai's coverage put it plainly: "Over the last month, the west has been stunned by the performance of GLM 5.2, then by Kimi K3, but it appears that Chinese labs aren't close to being done yet." Z.AI's GLM 5.2, Moonshot's Kimi K3, and now Alibaba's Qwen 3.8 form a genuine pattern rather than an isolated event: three separate Chinese labs, each claiming a place at or near the top of the global leaderboard, each shipping open-weight or promising to, within about a month of one another.

Community reaction to the Qwen 3.8 preview split along familiar lines within hours. On Hacker News, the dominant view was that open-weight competition between Chinese labs benefits everyone, though a number of commenters specifically questioned the "second only to Fable 5" framing and described Qwen as a benchmark specialist relative to genuinely broad frontier rivals. On Reddit's r/LocalLLaMA, discussion centred on the practical serving math — whether a 2.4-trillion-parameter model could realistically be self-hosted by anyone outside a hyperscaler, and hope for a smaller distilled variant. MarkTechPost characterised the overall mood as "cautiously positive": enthusiasm for another capable open-weight contender, checked by fatigue over yet another unverified benchmark claim.
Pricing and availability
| Token Plan tier | Price | Credits (7 days) | Concurrent agents |
|---|---|---|---|
| Lite | $6 | 2,500 | 1-2 |
| Standard | ~$18 | 10,000 | 3-4 |
| Pro | $68 | 40,000 | 6-8 |
Qwen3.8-Max-Preview is not sold as a standalone per-token API product. It is bundled into Alibaba's existing Token Plan subscriptions alongside other models — reporting lists Qwen3.7-Max, GLM 5.2, DeepSeek V4 Pro and Wan2.7-Image-Pro as sharing the same credit pool — at a preview discount of 10% off standard pricing. That structure makes a direct like-for-like comparison with Kimi K3's published per-token rate ($3.00 input / $15.00 output per million tokens, cache miss) difficult: credits are consumed across whichever bundled model a user calls, and Alibaba has not published what a Qwen-3.8-only equivalent per-token rate would work out to.
The plans carry OpenAI- and Anthropic-protocol compatibility, so tools such as Cursor, Cline, Codex and Claude Code can point at Qwen 3.8 through existing integrations. Availability is split into two regional deployments — an international endpoint (qwencloud.com) and a China-domestic one (platform.qianwenai.com) — each requiring separate registration and billing. Open-weight release remains promised but undated; anyone planning to self-host should note that Alibaba's two previous Max-tier flagships, Qwen3-Max-Preview and Qwen 3.6 Max Preview, both launched with the same "open weights coming" framing and neither has since been released openly.
Limitations
This is a genuine, current limitations list, not a caveat buried at the bottom of a launch post — as of 20 July 2026, every item below is still true and worth weighing before routing any real workload to Qwen3.8-Max-Preview.
- No independent benchmarks. Neither Artificial Analysis nor LMArena had scored Qwen3.8-Max-Preview at the time of writing. The "second only to Fable 5" claim rests entirely on Alibaba's own internal evaluation.
- No model card or specification sheet. There is no published document detailing training data, safety evaluation, or intended-use guidance — the kind of disclosure Anthropic, OpenAI and Google DeepMind attach to comparable launches.
- No licence. Because open weights have not shipped, there is no licence to review — which means no clarity yet on commercial-use terms, attribution requirements, or restrictions, unlike Kimi K3's already-published Modified MIT terms.
- Undisclosed active-parameter count. The number that determines real inference cost for a sparse MoE model of this size is simply not public.
- No per-token API pricing. Credit-bundle pricing makes cost forecasting for high-volume use harder than a transparent per-million-token rate.
- No confirmed context window, safety card, or hallucination data for Qwen 3.8 specifically — figures cited elsewhere in this article for context window and benchmark scores belong to the predecessor, Qwen3.7-Max, not to Qwen 3.8 itself.
None of this means the underlying claims are false. It means they are, as of publication, unverifiable by anyone outside Alibaba. MarkTechPost's own advice to readers is a reasonable one to repeat here: do not migrate production workloads on the strength of a teaser. Test Qwen3.8-Max-Preview on your own workload through the official console if you have access, and wait for a benchmark table, an active-parameter count, and independent scoring before treating the "second only to Fable 5" claim as settled.
How Qwen 3.8 Max compares
Against Kimi K3, Qwen 3.8 Max is smaller on paper (2.4T vs 2.8T total parameters) but arguably in a stronger structural position on multimodality, given Alibaba's claim that it is the team's first model above 1 trillion parameters to handle images, video and documents natively. Kimi K3 currently has the edge on transparency: it shipped with a benchmark table, a named licence, and (partial) independent verification from Artificial Analysis, none of which Qwen 3.8 has yet. See our Kimi K3 vs Claude Fable 5 scorecard for how a comparably-claimed Chinese frontier model actually performed once independent testing arrived — a useful reference point for how much of Qwen's own "second only to Fable 5" claim might survive the same process.
Against the Western closed frontier — Claude Fable 5 and GPT-5.6 Sol — no direct comparison is currently possible, because Qwen 3.8 has no independently verified score on any shared benchmark. Alibaba's claim, if it holds up the way Kimi K3's roughly did, would put it in the same tier as Claude Opus 4.8 and GPT-5.5 rather than genuinely ahead of the two current frontier leaders. Against fellow open-weight or soon-to-be-open Chinese models, Qwen 3.8 joins GLM 5.2 and Kimi K3 in a crowded and increasingly credible field; for the full scored picture across frontier models, see our live Benchmarks page, which we will update once Qwen 3.8 receives independent scores.
Who should use it — and who should wait
Worth trying now: teams already inside Alibaba's Token Plan ecosystem, or developers curious about a genuinely new multimodal architecture at this scale, who understand they are testing a preview with no benchmark backing and can absorb the cost of the Lite or Standard tier as a low-stakes evaluation. The OpenAI- and Anthropic-protocol compatibility makes this a low-friction experiment for anyone with an existing coding-agent setup.
Better to wait: anyone choosing a model for production workloads, anyone who needs a firm licence before committing (Qwen 3.8 has none yet), anyone who wants to self-host (there are no open weights and, at 2.4 trillion parameters, no realistic way to run this on typical hardware even once they arrive), and anyone specifically trying to validate the "second only to Fable 5" claim before making a purchasing decision. For all of those cases, the sensible move is to wait for Alibaba's benchmark table and for Artificial Analysis or LMArena to publish independent scores — the same process that, days after Kimi K3's launch, turned an unverified claim into an actual, checkable number.
The bottom line
Qwen3.8-Max-Preview is a serious signal from a serious lab — Alibaba does not casually claim second place behind Claude Fable 5 — but it is a signal, not yet evidence. The preview is real, the multimodal architecture claim is credible given Alibaba's track record, and the Token Plan pricing is concrete and checkable today. Everything that would actually substantiate "2.4 trillion parameters" and "second only to Fable 5," however — a model card, an active-parameter count, a benchmark table, a licence, and independent scoring — remains outstanding.
The more durable story here may not be about any single model at all. Three Chinese labs — Z.AI, Moonshot, and now Alibaba — have each shipped or previewed a trillion-parameter-class model within about a month, each claiming a place near the very top of the global leaderboard, in a week that also sent chip stocks into a bear market. Whatever Qwen 3.8's verified score eventually turns out to be, that pattern — not any one preview post on X — is the part of this story worth paying attention to.
Alibaba's official Qwen announcement on X and MarkTechPost's interactive confirmed-vs-claimed breakdown provide the underlying detail referenced throughout this article.
Last updated: 20 July 2026, the day after Qwen3.8-Max-Preview's announcement at WAIC Shanghai. This article will be revised once Alibaba publishes a benchmark table, an active-parameter count, and/or open model weights.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

Frequently Asked Questions
Is Qwen 3.8 Max open source?
How does Qwen 3.8 Max compare to Kimi K3?
When will independent benchmarks for Qwen 3.8 Max be available?
How much does Qwen 3.8 Max cost?
What is Qwen 3.8 Max's active parameter count?
Related Articles

AI Tools Review Editorial Team Expert Verified
Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.




