AI Tools Review
Qwen 3.8 Max Review: Alibaba's 2.4T Model, Tested

Insights

Qwen 3.8 Max Review: Alibaba's 2.4T Model, Tested

AI Tools Review Editorial Team3 August 2026

    Quick answer:

    Qwen3.8-Max went generally available on 3 August 2026, two weeks after its unverified preview, and this time Alibaba backed it with a real benchmark table, real per-token pricing, and a firm open-weights date. It scores 86.6 on Terminal-Bench 2.1 (ahead of Claude Opus 4.8 and Claude Fable 5's 84.6, behind GPT-5.6 Sol's 88.8), leads PaperBench at 93.0, and posts strong multimodal scores including 86.1 on OSWorld-Verified. Pricing is a straightforward $2.00 input / $6.00 output per million tokens, with a 1-million-token context window. Open weights for both Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint are now promised "next week", Alibaba's first firm timeline for either. The active-parameter count remains undisclosed, and none of these benchmark numbers have yet been independently reproduced by Artificial Analysis or LMArena.

    Two weeks ago, Alibaba walked onto the stage at WAIC in Shanghai and claimed Qwen3.8-Max was "second only to Fable 5", with no benchmark table to back it up. On 3 August 2026, that changed: Alibaba pushed Qwen3.8-Max to general availability with a full published benchmark table, real API pricing, and a confirmed open-weights timeline. The number that once had zero supporting evidence now has dozens of scores attached to it, some of them genuinely strong, several of them clearly behind the frontier leaders it was compared against.

    This update replaces the provisional analysis this article carried at preview time. What follows separates the real GA numbers from the surrounding hype, walks through what changed between preview and release, and places Qwen3.8-Max back into the context of the Chinese open-weight race that produced three trillion-parameter-class releases in barely a month.

    Note: this analysis is based on Alibaba's official Qwen3.8 GA blog post (qwen.ai/blog?id=qwen3.8), the published model page on qwencloud.com, MarkTechPost's GA coverage and benchmark-table screenshot, and independent benchmark data from Artificial Analysis for Qwen's prior flagship, Qwen3.7-Max. AI Tools Review has not independently reproduced any benchmark, and no independent third-party scores for Qwen3.8-Max existed at the time of writing. Prices are billed in US dollars.

    Julian Goldie SEO's coverage of Qwen 3.8 Max shortly after release.

    Summary

    • Now generally available: Alibaba took Qwen3.8-Max from unverified preview to GA on 3 August 2026, with a real published benchmark table replacing the earlier bare claim.
    • 86.6 on Terminal-Bench 2.1, ahead of Claude Opus 4.8 and Claude Fable 5 (both 84.6), behind GPT-5.6 Sol (max) at 88.8. Leads PaperBench at 93.0 and IFBench at 82.8.
    • Clearly behind Fable 5 on core software-engineering benchmarks: 67.7 vs 80.0 on SWE-bench Pro, 73.5 vs 88.8 on FrontierSWE. The gains are strongest in multimodal and agentic tasks, not raw coding.
    • Real per-token pricing at last: $2.00 input / $6.00 output per million tokens, with implicit cache reads at $0.25/M, replacing the earlier opaque Token Plan credit bundles.
    • 1-million-token context window confirmed: 991K max input, 131K max output, 262K max reasoning budget.
    • Active-parameter count still undisclosed. For a sparse Mixture-of-Experts model this is the single number that determines real inference cost, and Alibaba has not published it even at GA.
    • Open weights now dated for the first time: both Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint are promised "next week", though no licence has been named yet.
    • Arrives amid a genuine Chinese open-weight surge. GLM 5.2, Kimi K3 and now Qwen 3.8 have each launched within about a month of one another, each claiming a place near the top of the global leaderboard.

    From Qwen3-Max to Qwen 3.8: the lineage

    Alibaba's Qwen team crossed the trillion-parameter threshold for the first time with Qwen3-Max-Preview in September 2025. Qwen 3.6 Max Preview followed in April 2026, and, notably, stayed closed, hosted only through Qwen Studio and Alibaba Cloud rather than released as open weights on day one, a departure from Alibaba's earlier practice of shipping many Qwen models openly from launch. Qwen3.7-Max arrived in May 2026 as an incremental step on the same closed-weight, reasoning-focused path, adding a 1-million-token context window and posting strong scores on agentic and coding evaluations: 92.4 on GPQA-Diamond, 80.4% on SWE-bench Verified, and 69.7 on Terminal-Bench 2.0-Terminus, all figures Alibaba published and that independent trackers such as Artificial Analysis have broadly corroborated (Qwen3.7-Max reached 56.6 on the Artificial Analysis Intelligence Index, fifth overall at the time).

    Qwen3.8-Max-Preview, announced less than three months after Qwen 3.6 Max Preview, is a much bigger jump than the version number suggests: total parameters more than double again, to 2.4 trillion, and Alibaba is pitching it as the team's first multimodal model above 1 trillion parameters, a genuine architectural step, not just a scale-up of the 3.7 generation. Unlike its two immediate predecessors, which have stayed closed since their own previews, Qwen3.8-Max broke that pattern on 3 August 2026: Alibaba took it to full GA with a real benchmark table and per-token pricing, and for the first time gave a firm (if still undated-by-day) promise that open weights are coming "next week".

    The timing is not incidental. Alibaba owns roughly 36% of Moonshot AI, the Beijing lab behind Kimi K3, which means Qwen 3.8's launch is as much an in-house rivalry as a competitive response, two of Alibaba's own portfolio companies trading blows for the title of strongest Chinese model within the same week. It also lands one month after Z.AI's GLM 5.2, which drew praise from engineers at Vercel and Box for landing within striking distance of Claude Opus 4.8 on FrontierSWE. Three trillion-parameter-class Chinese releases in barely a month is the real story behind any single one of them.

    Architecture and training: what is and isn't known

    What Alibaba has confirmed about Qwen3.8-Max is narrower than the headline parameter count suggests, but it grew substantially at GA. Qwen developer Shuai Bai described it as a sparse Mixture-of-Experts model and the team's first multimodal system above 1 trillion parameters, capable of processing text, image and video input. The GA model page confirms a 1-million-token context window (991K max input, 983K with thinking enabled, 131K max output, 262K max reasoning budget), and the API is OpenAI- and DashScope-compatible, meaning existing coding agents and harnesses can point at it with a base-URL and model-ID change rather than a rebuild.

    Editorial graphic reading 'Alibaba Previews Qwen3.8-Max, 2.4 Trillion Parameters, Multimodal Model Unveiled Days After Moonshot's Kimi K3 Open-Weight Launch'
    Qwen3.8-Max-Preview arrived two days after Moonshot's Kimi K3 and during Alibaba Cloud's home-turf conference, WAIC, in Shanghai. Source: MarkTechPost.

    Beyond that, the specifics that would normally accompany a flagship launch are simply absent. There is no model card, no specification sheet, and, critically for a sparse Mixture-of-Experts design, no disclosed active-parameter count. That figure is not a footnote: it is the number that determines how much compute a single forward pass actually costs, and Qwen's own smaller models show how differently "2.4 trillion parameters" can translate into real serving cost depending on sparsity. Qwen3-235B-A22B, for instance, carries 235 billion total parameters but activates only 22 billion per token; Qwen3-30B-A3B activates roughly 3 billion of its 30 billion. Without an equivalent ratio for the Max-tier 2.4T model, the total parameter count alone says very little about what it costs Alibaba (or would cost anyone else) to run it.

    The scale also raises a practical question independent of pricing: if the full 2.4-trillion-parameter model does eventually ship as open weights, loading it would require roughly 1.2 terabytes of storage for the weights alone at 4-bit precision, before accounting for the KV cache and runtime overhead: well beyond what even multiple high-end Nvidia H200 GPUs (141GB of memory each) could hold without a cluster. Whether Alibaba ultimately releases the full model, a smaller activated-parameter variant, a quantised checkpoint, or a distilled sibling remains unknown, and materially changes who could realistically self-host it.

    Capabilities deep dive

    Multimodal input

    Alibaba's headline capability claim is multimodality at scale: Qwen3.8-Max is the Qwen team's first model above 1 trillion parameters that can process images, video and documents alongside text, and at GA this is now the area with the strongest published scores, topping OSWorld-Verified (86.1), Parametric CAD Bench (91.5) and OmniDocBench 1.5 (92.1). That is a genuinely stronger evidence base than the preview offered, though these table-leading scores are also where the softer Qwen3.7-Plus comparison baseline applies, so the size of the generational jump should be read with that caveat in mind.

    Coding, full-stack development and office workflows

    Alibaba's stated focus for the improvement over Qwen3.7-Max is squarely practical: coding, full-stack development, data analysis and office productivity tasks, rather than a broad claim of general intelligence gains. Combined with OpenAI- and DashScope-protocol compatibility, this reads as a direct pitch to teams already running agentic coding workflows on other models, inviting them to swap in Qwen 3.8 without rebuilding their tooling. The GA feature set maps onto four industries specifically: software engineering, legal and financial document review, media and e-commerce operations, and design, with applications spanning repository-scale coding agents, long-document knowledge bases, long-video indexing, structured data extraction and multi-step research assistants.

    Five built-in tools on the Responses API

    The most concrete agentic detail Alibaba shipped at GA is not a benchmark but a set of five built-in tools available directly on the Responses API: code_interpreter, web_search, web_extractor, t2i_search and i2i_search, alongside function calling, structured outputs, batch processing, prefix completion and fine-tuning support. That is a genuinely useful, easy-to-verify infrastructure claim for teams evaluating whether the API can support a multi-agent coding or research setup, independent of how any individual benchmark score holds up under independent scrutiny.

    Benchmarks: what Alibaba actually published at GA

    This is the section that changed most between preview and release. At preview, there was no benchmark table at all, just a single sentence claiming Qwen3.8-Max was "second only to Fable 5". At GA on 3 August 2026, Alibaba published a full table spanning coding-agent and general-agent categories, benchmarked directly against Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol (max) and its own predecessor, Qwen3.7-Max.

    Table of Qwen3.8-Max GA benchmark scores against Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol and Qwen3.7-Max across coding-agent and general-agent categories
    Alibaba's own GA benchmark table for Qwen3.8-Max. Source: qwen.ai, via MarkTechPost.
    BenchmarkQwen3.8-MaxFable 5GPT-5.6 Sol (max)
    Terminal-Bench 2.186.684.688.8
    SWE-bench Pro67.780.064.6
    FrontierSWE73.588.8--
    PaperBench93.088.890.5
    GPQA Diamond92.6----
    OSWorld-Verified86.1----

    The pattern is consistent across the full table: Qwen3.8-Max is genuinely competitive or leading on agentic, tool-use and multimodal benchmarks (topping PaperBench at 93.0, OSWorld-Verified at 86.1, Parametric CAD Bench at 91.5 and OmniDocBench 1.5 at 92.1), but clearly behind Fable 5 on the two benchmarks that most directly measure real-world software-engineering capability, SWE-bench Pro (67.7 vs 80.0) and FrontierSWE (73.5 vs 88.8). Against its own predecessor, Qwen3.7-Max, the generational jump is large: DeepSWE 1.1 moves from 21.6 to 56.6, FrontierSWE from 40.7 to 73.5, and JobBench from 31.3 to 53.4.

    Two caveats belong in an honest read of this table. First, Alibaba's multimodal comparison benchmarks against Qwen3.7-Plus, not Qwen3.7-Max, which flatters the generational delta on those specific rows. Second, Alibaba's own published RL scaling curve for the model peaks at 0.725 near 4,000 training environments, then declines to 0.719 and 0.689, a detail Alibaba included itself but one that is easy to miss in a table full of green numbers. As with the preview, every score in this table is Alibaba's own internal run; no independent evaluator such as Artificial Analysis or LMArena had reproduced these figures at the time of writing.

    This is not an unprecedented pattern. Moonshot AI made a similar claim when it first announced Kimi K3 weeks earlier, that it trailed only Fable 5 and GPT-5.6 Sol, and that claim was also unverified at launch. Artificial Analysis subsequently scored K3 at 57 on the Intelligence Index, competitive with but not ahead of several other frontier models, a useful reminder that vendor-published tables, even detailed ones, still tend to be optimistic relative to what independent testing later confirms. See our full Kimi K3 review for how that claim held up.

    Competitive landscape

    Qwen 3.8's launch cannot be read in isolation from the week that preceded it. On 17 July 2026, Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight model that briefly held the title of the largest open-source model ever released, and which beat both Claude Fable 5 and GPT-5.6 Sol outright on Arena.ai's frontend-development leaderboard, the first open-weight model to top that particular ranking. The reaction was not confined to AI commentary: South Korea's KOSPI fell more than 6% and Japan's Nikkei dropped over 4% in the days that followed, the Philadelphia Semiconductor Index fell into bear-market territory (down more than 20% from its late-June peak, its worst week in over fifteen months), and US chip and AI-adjacent names including Intel, Micron, AMD and Marvell all declined. Analysts were careful to note the selloff was "overdetermined", weak Netflix and TSMC earnings, the Iran war, and broader risk-off sentiment all contributed, but Apollo's Torsten Sloek specifically flagged the risk that cheap or free capable AI from Chinese open models could undercut the roughly $700 billion in annual hyperscaler AI infrastructure spending that markets are pricing in as a return on investment.

    Qwen 3.8 arrived directly into that atmosphere, two days later, at Alibaba's home-turf conference. OfficeChai's coverage put it plainly: "Over the last month, the west has been stunned by the performance of GLM 5.2, then by Kimi K3, but it appears that Chinese labs aren't close to being done yet." Z.AI's GLM 5.2, Moonshot's Kimi K3, and now Alibaba's Qwen 3.8 form a genuine pattern rather than an isolated event: three separate Chinese labs, each claiming a place at or near the top of the global leaderboard, each shipping open-weight or promising to, within about a month of one another.

    A hand holding a smartphone displaying the Qwen app logo and interface
    The Qwen consumer app. Alibaba positions Qwen 3.8 as both a developer-facing API model and a consumer product. Source: South China Morning Post.

    Community reaction to the Qwen 3.8 preview split along familiar lines within hours. On Hacker News, the dominant view was that open-weight competition between Chinese labs benefits everyone, though a number of commenters specifically questioned the "second only to Fable 5" framing and described Qwen as a benchmark specialist relative to genuinely broad frontier rivals. On Reddit's r/LocalLLaMA, discussion centred on the practical serving math: whether a 2.4-trillion-parameter model could realistically be self-hosted by anyone outside a hyperscaler, and hope for a smaller distilled variant. MarkTechPost characterised the overall mood as "cautiously positive": enthusiasm for another capable open-weight contender, checked by fatigue over yet another unverified benchmark claim.

    Pricing and availability

    Rate (per 1M tokens)Price
    Input$2.00
    Output$6.00
    Implicit cache read$0.25
    Explicit cache creation$2.50
    Explicit cache read$0.17

    This is the biggest structural change between preview and GA: Qwen3.8-Max is no longer bundled into opaque Token Plan credit packs. It now has a real, standalone per-token API price, $2.00 input / $6.00 output per million tokens, directly comparable to Kimi K3's published rate of $3.00 input / $15.00 output. Cached input is roughly eight times cheaper than fresh input, so prefix stability now drives cost more than raw prompt length. Rate limits are 2 million tokens per minute and 15,000 requests per minute. The model also ships with five built-in tools on the Responses API: code_interpreter, web_search, web_extractor, t2i_search and i2i_search, alongside function calling, structured outputs, batching, prefix completion and fine-tuning support.

    The GA model is OpenAI- and DashScope-compatible, so tools such as Cursor, Cline, Codex and Claude Code can point at it with a base-URL and model-ID change rather than a rebuild. The hosted API is deployable today by any company size; the open weights are a different matter. At 2.4 trillion total parameters, the flagship checkpoint is a multi-node datacentre artefact, and with the active-parameter count still undisclosed, its real serving cost cannot yet be modelled. Alibaba has now confirmed weights for both Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint are due "next week", its first firm timeline for either, but the 27B model, not the 2.4T flagship, is the realistic self-hosting path on ordinary on-premise GPU hardware.

    Limitations

    GA closed several of the gaps this article flagged at preview time, but a genuine limitations list remains, as of 3 August 2026, every item below is still true and worth weighing before routing a real workload to Qwen3.8-Max.

    • No independent benchmarks yet. Neither Artificial Analysis nor LMArena had scored the GA release at the time of writing. Every number in Alibaba's table, however detailed, is still Alibaba's own internal run.
    • Clearly behind Fable 5 on core coding benchmarks. SWE-bench Pro (67.7 vs 80.0) and FrontierSWE (73.5 vs 88.8) both show a meaningful gap against Anthropic's frontier model, despite Qwen3.8-Max leading on several agentic and multimodal rows.
    • No licence yet. Open weights are dated for the first time ("next week"), but Alibaba has not named the licence that will govern them, so commercial-use terms, attribution requirements and restrictions remain unknown.
    • Undisclosed active-parameter count. The number that determines real inference cost for a sparse MoE model of this size is still not public, even at GA.
    • A self-reported wrinkle in the RL scaling curve. Alibaba's own chart shows performance peaking at 0.725 near 4,000 training environments before declining to 0.719 and 0.689, worth noting precisely because Alibaba published it rather than hid it.
    • Multimodal comparisons use a softer baseline. The multimodal benchmark table compares against Qwen3.7-Plus rather than Qwen3.7-Max, which flatters the generational improvement on those specific scores.

    None of this means the GA numbers are false, GA is a materially stronger evidence base than the preview's single sentence. It means they remain, as of publication, unverified by anyone outside Alibaba. The sensible approach is the same one MarkTechPost recommended at preview: test Qwen3.8-Max on your own workload through the official console, and treat the table as a strong signal rather than a settled ranking until Artificial Analysis or LMArena publish independent scores.

    How Qwen 3.8 Max compares

    Against Kimi K3, Qwen 3.8 Max is smaller on paper (2.4T vs 2.8T total parameters) but now has the more detailed published benchmark table of the two, plus real per-token pricing. Kimi K3 still has the licence-transparency edge: it shipped with a named Modified MIT licence at launch, while Qwen3.8-Max has only a dated promise of open weights and no named licence yet. See our Kimi K3 vs Claude Fable 5 scorecard for how a comparably-claimed Chinese frontier model performed once independent testing arrived, a useful reference point for how much of Qwen's own table might survive the same process.

    Against the Western closed frontier, Qwen3.8-Max's own published numbers put it ahead of Claude Fable 5 on Terminal-Bench 2.1 (86.6 vs 84.6) and PaperBench (93.0 vs 88.8), but clearly behind on SWE-bench Pro and FrontierSWE, and behind GPT-5.6 Sol on Terminal-Bench 2.1 (88.8). Read together, that is a model that is genuinely strong at agentic and multimodal tasks without yet matching Fable 5 on the software-engineering benchmarks most coding teams actually care about. Against fellow open-weight or soon-to-be-open Chinese models, Qwen 3.8 joins GLM 5.2, Kimi K3 and the newer Tencent Hy4 preview in a crowded and increasingly credible field; for the full scored picture across frontier models, see our live Benchmarks page, which we will update once Qwen 3.8 receives independent scores.

    Who should use it, and who should wait

    Worth trying now: teams building agentic, tool-use or multimodal workflows, where Qwen3.8-Max's published scores are genuinely competitive with or ahead of the frontier leaders, and anyone who wants transparent, checkable per-token pricing rather than a credit-bundle black box. The OpenAI- and DashScope-compatible API makes this a low-friction test for anyone with an existing coding-agent setup, and at $2/$6 per million tokens it is meaningfully cheaper to trial than Fable 5 or GPT-5.6 Sol.

    Better to wait: teams whose primary workload is software engineering rather than agentic or multimodal tasks, where Fable 5 still leads by a clear margin on SWE-bench Pro and FrontierSWE; anyone who needs a firm open-source licence before committing (none has been named yet); and anyone who wants to self-host the full 2.4-trillion-parameter flagship rather than the smaller 27B checkpoint. For all of those cases, the sensible move is to wait for the promised open weights to actually land next week, for a named licence, and for Artificial Analysis or LMArena to publish independent scores.

    The bottom line

    Qwen3.8-Max's move from preview to GA is a genuine improvement in evidence quality: a bare claim has become a full benchmark table, real per-token pricing, and a dated open-weights promise. The results are also more nuanced than the preview-era "second only to Fable 5" framing suggested, Qwen3.8-Max leads on several agentic and multimodal benchmarks but trails Fable 5 clearly on the software-engineering benchmarks that matter most for coding-heavy teams.

    The more durable story here may still not be about any single model at all. Three Chinese labs (Z.AI, Moonshot, and now Alibaba) have each shipped a trillion-parameter-class model within about two weeks of one another, and Qwen3.8-Max is the first of the three to back its launch with a full self-published benchmark table rather than a single claim. Whatever independent evaluators eventually confirm, that shift toward more evidence at launch, not any individual score, is the part of this GA release worth paying attention to.

    Alibaba's official Qwen3.8 GA announcement and MarkTechPost's GA coverage and benchmark breakdown provide the underlying detail referenced throughout this article.

    Last updated: 3 August 2026, the day Qwen3.8-Max went generally available with a full benchmark table, real per-token pricing, and a dated open-weights promise. This article will be revised again once the open weights and licence actually ship, and once independent evaluators publish scores.

    Free Guide

    Get the free guide: Claude vs ChatGPT, Gemini & Grok

    A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.

    Pop your email in to get it free
    Preview of the free guide: Claude vs ChatGPT, Gemini and Grok, 2026 features, pricing and what-you-can-do comparison.

    Frequently Asked Questions

    Is Qwen 3.8 Max open source?
    Not yet, but it is close. Alibaba took Qwen3.8-Max to general availability on 3 August 2026 and confirmed open weights for both Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint are shipping 'next week', its first firm timeline for either artefact. No licence has been named yet, so terms remain unconfirmed until the weights actually land on Hugging Face and ModelScope. The 27B checkpoint, not the 2.4-trillion-parameter flagship, is the one realistic self-hosting path on ordinary GPU hardware.
    How does Qwen 3.8 Max compare to Kimi K3?
    Both are Chinese Mixture-of-Experts models that launched within weeks of each other in mid-2026 and both were pitched as trailing only Claude Fable 5. With Qwen3.8-Max now GA and independently benchmarked by Alibaba's own published table, it scores 86.6 on Terminal-Bench 2.1 versus Fable 5's 84.6 and posts strong multimodal numbers (OSWorld-Verified 86.1), but trails Fable 5 clearly on SWE-bench Pro (67.7 vs 80.0) and FrontierSWE (73.5 vs 88.8). Kimi K3 still has the license-transparency edge, having shipped a named Modified MIT licence at launch, something Qwen3.8-Max has promised but not yet delivered. See our full Kimi K3 review on this site for the comparison in detail.
    What are Qwen 3.8 Max's real benchmark scores?
    Alibaba published a full benchmark table alongside the 3 August 2026 GA release. Headline coding-agent scores: Terminal-Bench 2.1 86.6 (ahead of Opus 4.8 and Fable 5's 84.6, behind GPT-5.6 Sol's 88.8), SWE-bench Pro 67.7, FrontierSWE 73.5, PaperBench 93.0 (a table lead), IFBench 82.8, and GPQA Diamond 92.6. The clearest jump over its own predecessor, Qwen3.7-Max, is multimodal and agentic: DeepSWE 1.1 rose from 21.6 to 56.6 and FrontierSWE from 40.7 to 73.5. Two caveats worth keeping: Alibaba's multimodal comparison table benchmarks against Qwen3.7-Plus rather than Qwen3.7-Max, which flatters the generational delta, and none of these figures have yet been independently reproduced by Artificial Analysis or LMArena.
    How much does Qwen 3.8 Max cost?
    As of the 3 August 2026 GA release, Qwen3.8-Max has a real per-token API price: $2.00 per million input tokens and $6.00 per million output tokens. Implicit cache reads cost $0.25 per million tokens (an 8x discount versus fresh input), explicit cache creation is $2.50 per million tokens, and explicit cache reads are $0.17 per million tokens. Rate limits are 2 million tokens per minute and 15,000 requests per minute. This replaces the earlier preview-era Token Plan credit bundles, which made cost forecasting difficult; the new pricing is directly comparable to competitors like Kimi K3's $3.00 input / $15.00 output per million tokens.
    What is Qwen 3.8 Max's context window and active parameter count?
    The GA model page confirms a 1-million-token context window: maximum input is 991K tokens (983K with thinking enabled), maximum output is 131K tokens, and the maximum reasoning budget is 262K tokens. The active-parameter count for the 2.4-trillion-parameter Mixture-of-Experts model is still undisclosed, which remains the single most-cited gap in Alibaba's technical disclosure, since it is the number that actually determines real inference cost on a sparse MoE architecture.
    AI Tools Review Editorial Team

    AI Tools Review Editorial Team Expert verified

    Our editorial team consists of veteran AI researchers, software engineers, and industry analysts. We spend hundreds of hours benchmarking frontier models natively to provide you with objective, actionable intelligence on agentic AI capabilities and cybersecurity landscapes.