Alibaba used the biggest stage of its year to say the quiet part out loud: Qwen 4 exists, it is training, and it comes in four sizes. What it did not do is ship anything. That gap between announcement and availability matters, because the Qwen family is now one of the most downloaded open-weight model lines in the world, and the 27B tier in particular is the model many people will actually run on their own hardware. This guide separates what Alibaba has officially confirmed from what has been reported, inferred or rumoured, and works through what the lineup is likely to mean in practice.
We cover the confirmed-versus-rumoured picture, each tier's expected role, the architecture preview Alibaba has already released, the Qwen 3.x lineage it builds on, realistic memory maths for running a 27B model on a Mac mini or a gaming GPU, and how Qwen 4 is likely to sit against DeepSeek, Kimi, GLM and the US frontier labs.
Note: Alibaba has published almost no technical detail about Qwen 4 itself. Every figure in this article is attributed to its source. Where we estimate (for example memory requirements), we say so and show the working. We checked Alibaba's Hugging Face organisation on 29/09/2026: no Qwen 4 repository exists.
WorldofAI's weekly round-up covers Alibaba's Qwen 4.0 lineup and the Kimi K3.1 identifier spotted in Moonshot's API, alongside the week's US lab news.
Qwen 4 at a Glance
Qwen 4 is Alibaba's next-generation large language model family, announced at the Apsara Conference in Hangzhou on 22/09/2026 and still in training as of 29/09/2026. Here is the short version:
- Status: "currently in training", in the words of Alibaba's official press release.
- Lineup: Qwen 4 Max, Qwen 4 Flash, Qwen 4 Plus and Qwen 4 27B, previewed on stage according to reports from the keynote. The written press release does not list them.
- Release date, price, context window, benchmarks: none published.
- Architecture preview: Qwen3.8-Flash-Next, released on 26/08/2026, is described by Alibaba as "an experimental preview of the architecture that will underpin Qwen4".
- Roadmap: Qwen 4.5 and Qwen 5 are "projected to scale up to 5 to 10 trillion parameters". The current flagship, Qwen3.8-Max, has 2.4 trillion.
- Also at Apsara: the Zhenwu V900 AI chip (mass production Q1 2027), a 20GW data-centre target for 2032, and a batch of Qwen-branded hardware.
What Is Confirmed and What Is Rumoured?
Only two things about Qwen 4 are confirmed in Alibaba's own written materials: that it is in training, and that Qwen 4.5 and Qwen 5 will follow at 5-10 trillion parameters. Everything else comes from on-stage remarks relayed by press and analysts, or from inference based on earlier Qwen releases. The table below sorts each claim by how solid it is.
| Claim | Status | Source |
|---|---|---|
| Qwen 4 is "currently in training" | Confirmed | Alibaba Cloud press release, 22/09/2026 |
| Qwen 4.5 and Qwen 5 to scale to 5-10 trillion parameters | Confirmed (as a projection) | Alibaba Cloud press release; RTHK; Investing.com |
| Qwen3.8-Flash-Next previews the Qwen 4 architecture | Confirmed | Official Hugging Face model card |
| Four tiers: Max, Flash, Plus, 27B | Reported from the keynote stage | OrcaRouter, Yotta Labs and others; not in the written press release |
| Release "very soon" | Reported, no date attached | Keynote coverage attributing it to the Qwen team |
| Qwen 4 27B will be open weights | Likely, unconfirmed | Precedent (Qwen3.8-27B is Apache 2.0); no licence announced |
| Tier roles (Max flagship, Flash cheap, Plus balanced) | Inferred | Naming conventions of Qwen 3.x generations |
| Enterprise API in November, open weights in December | Rumour | Unsourced third-party speculation; no Alibaba statement |
| A September 2026 launch (from a July leak) | Did not happen | July leak attributed to a YouTube channel, per Geeky Gadgets |
| Pricing, context window, benchmarks, parameter counts | Unknown | Nothing published |
One date quirk is worth flagging. Alibaba's press room dates the announcement to 22/09/2026 in Hangzhou, whilst the Alizila write-up of the same conference carries a 24/09/2026 date. The keynote and the first wave of coverage, including RTHK's report, are from 22/09/2026, so that is the date we use.

When Will Qwen 4 Be Released?
Qwen 4 does not have a release date. Alibaba's official statement stops at "currently in training", and no Qwen 4 model has appeared on Hugging Face, ModelScope or Alibaba's Qwen Cloud API as of 29/09/2026. We checked Alibaba's Hugging Face organisation directly: the newest Qwen releases are Qwen-Image-2.1 variants from 14/09/2026 and 20/09/2026, and the newest language models are Qwen3.8-Flash-Next and the Qwen3.8-27B and 2.4T-A95B checkpoints from August.
What can we reasonably infer? Alibaba's recent cadence gives some clues:
- Architecture previews come months early. Qwen3-Next previewed a new architecture in September 2025, and the Qwen3.5 family followed on 17/02/2026, roughly five months later, as CellCog points out. Applying the same gap to Qwen3.8-Flash-Next (26/08/2026) would point to early 2027.
- But the team is now moving faster. Qwen3.8-Max went GA on 03/08/2026 and its open weights followed within days, whilst the 27B checkpoint landed on 14/08/2026. See our Qwen 3.8 Max review for that timeline.
- "Very soon" has been said before. Treat it as intent, not a schedule.
Our honest read: a hosted Qwen 4 Max or Plus preview in Q4 2026 is plausible, and the 27B weights could follow shortly after, as Qwen3.8-27B did. But nothing Alibaba has published supports a specific month, and anyone quoting a precise date is guessing.
What Are the Four Qwen 4 Models?
The Qwen 4 family was previewed as four tiers: Max, Flash, Plus and 27B. Alibaba gave no specifications for any of them, so the roles below are based on how the same names have been used across Qwen 3.x. They are well-founded expectations, not facts.
| Tier | Expected role | Closest current model | Likely access |
|---|---|---|---|
| Qwen 4 Max | Flagship reasoning, coding and agentic model; Alibaba's answer to the top US and Chinese models | Qwen3.8-Max (2.4T parameters, $2/$6 per million tokens, approx. £1.50/£4.50) | Qwen Cloud API first; open weights uncertain |
| Qwen 4 Plus | Balanced price and capability, historically the tier with the broadest multimodal support | Qwen3.7-Plus and the Qwen Plus API line | API |
| Qwen 4 Flash | Low-latency, low-cost, high-volume work: chat, extraction, routing, sub-agents | Qwen3.8-Flash (production version of Flash-Next, $0.16/$0.47 per million tokens, approx. £0.12/£0.35) | API; a Flash-Next-style open checkpoint is possible |
| Qwen 4 27B | Open-weight model for local deployment, fine-tuning and quantisation | Qwen3.8-27B (dense, Apache 2.0, 262K context) | Expected Hugging Face / ModelScope download |
Qwen 4 Max
Qwen 4 Max is the expected flagship. Its predecessor, Qwen3.8-Max, is Alibaba's first multimodal model above one trillion parameters, and Alibaba used Apsara to say that an updated Qwen3.8-Max ran 33 iterative self-improvement cycles over a month of fully automated runs, lifting its Artificial Analysis score from 40 to 45. That recursive self-improvement work is a strong hint about how the Max tier of Qwen 4 is being trained. If Alibaba keeps Qwen3.8-Max-style pricing, Max would undercut US flagships heavily, but no price has been announced.
Qwen 4 Plus
Qwen 4 Plus is expected to be the middle tier: more capable than Flash, cheaper than Max. The one hard data point here is indirect. The Qwen3.8-Flash-Next write-up claims the preview model delivers better results than Qwen3.7-Plus at roughly one-ninth of the training cost, as The Decoder reported. If the new architecture delivers that kind of efficiency gain, a Qwen 4 Plus should land well above the current Plus line.
Qwen 4 Flash
Qwen 4 Flash is the throughput tier. Qwen3.8-Flash-Next already activates only 6 billion parameters per token out of 125 billion, and the hosted Qwen3.8-Flash costs $0.16 input and $0.47 output per million tokens (approx. £0.12/£0.35), roughly twelve times cheaper than Qwen3.8-Max according to The Decoder. Expect Qwen 4 Flash to compete directly with DeepSeek V4.1 Flash and Google's Gemini Flash models on price per task.
Qwen 4 27B
Qwen 4 27B is the model most people reading this will actually be able to run. Its predecessor, Qwen3.8-27B, is a dense vision-language model released under Apache 2.0 with a 262,144-token native context window, and it runs through Ollama, LM Studio, llama.cpp, vLLM and SGLang. An important caveat: if Qwen 4 27B adopts the new Flash-Next architecture, "27B" may not mean a dense 27-billion-parameter model, and memory requirements could differ from the estimates later in this article.
What Will Qwen 4 Be Built On?
Qwen 4 will be built on the architecture Alibaba previewed with Qwen3.8-Flash-Next on 26/08/2026. The official model card is unambiguous, describing the release as "an experimental preview of the architecture that will underpin Qwen4", and TechNode reported that the early release was meant to let developers prepare for the full Qwen4 family. That makes Flash-Next the most concrete technical evidence available about Qwen 4.
According to the Hugging Face model card, the key ingredients are:
- Hybrid attention with Qwen Sparse Attention (QSA). Three of every four layers use Gated DeltaNet, a linear-attention design that keeps a fixed-size state rather than a growing key-value cache. The fourth uses QSA, which selects context at the micro-block level instead of token by token. Alibaba says this "cuts long-context latency significantly".
- Gated Residual. Widened residual streams with data-dependent read and write gates, intended to add expressiveness without hurting training stability.
- N-gram embeddings. A 51-billion-parameter table of 20 million bigrams and trigrams injected at layer 2. Alibaba's pitch is that embeddings scale parameters with less compute than Mixture-of-Experts and can be offloaded from the accelerator.
- Sparse MoE. 125 billion parameters with 6 billion active, plus the 51B n-gram table and a 4B multi-token-prediction module; 48 layers; 262,144-token native context, extendable to about one million.
- Training recipe. Muon and AdamW applied to different weight categories, and no batch-size warm-up.
Two details matter for local users. First, the Gated DeltaNet layers mean only a quarter of layers keep a traditional growing KV cache, which makes long contexts far cheaper in memory. Second, the n-gram table can live in system RAM rather than GPU memory. If Qwen 4 27B inherits both, it could handle long documents on modest hardware better than a conventional dense 27B would. Note that the Flash-Next weights ship under the qwen-community-1.0 licence, not Apache 2.0, which is a reminder that Alibaba's licensing varies release by release.

The Qwen 3.x Lineage: How Did Qwen Get Here?
Qwen 4 follows the busiest year in Qwen's history. The Qwen 3.x generation ran from the original Qwen3 in April 2025 to the 2.4-trillion-parameter Qwen3.8-Max in August 2026, with open-weight releases at almost every step. Understanding that lineage helps set realistic expectations for Qwen 4.
| Date | Release | Why it matters for Qwen 4 |
|---|---|---|
| 29/04/2025 | Qwen3 | Start of the 3.x generation; mixed dense and MoE open releases |
| September 2025 | Qwen3-Next | First architecture preview, introducing hybrid Gated DeltaNet attention |
| 17/02/2026 | Qwen3.5 family | Full family on the Qwen3-Next ideas, about five months after the preview |
| May 2026 | Qwen3.7-Max | Frontier-class hosted flagship |
| 03/08/2026 | Qwen3.8-Max GA | 2.4T parameters, 1M-token context, $2/$6 per million tokens |
| August 2026 | Qwen3.8-2.4T-A95B open weights | First Max-class open release, under a custom licence |
| 14/08/2026 | Qwen3.8-27B | Dense, Apache 2.0, vision and video; the direct predecessor of Qwen 4 27B |
| 26/08/2026 | Qwen3.8-Flash-Next | Official Qwen 4 architecture preview |
| 22/09/2026 | Qwen 4 announced | In training; four tiers previewed |
Alongside the language models, Alibaba has kept a parallel multimodal line. Qwen-Image-3.0 arrived in July without weights or benchmarks, open Qwen-Image-2.1 variants followed in September, and at Apsara Alibaba said Qwen-Image 3.1 will launch later in 2026. It also announced Qwen3.8-LiveTranslate, which it says cuts simultaneous-interpretation latency from 2.8 to 2.3 seconds, and a Qwen-Audio-3.1-TTS-Next model. Whether any of these fold into Qwen 4 branding has not been said.
The key lesson from the lineage is that Alibaba has become more selective about licensing. Qwen3.8-27B is plain Apache 2.0, but the 2.4T-A95B open checkpoint carries a licence requiring a commercial agreement above $50 million (approx. £37.5 million) in qualifying annual revenue, and Flash-Next uses the qwen-community-1.0 licence. So "open" for Qwen 4 could mean several different things.
Qwen 4.5 and Qwen 5: What Do 5-10 Trillion Parameters Mean?
Alibaba's roadmap says the Qwen 4.5 and Qwen 5 series are "projected to scale up to 5 to 10 trillion parameters". That is roughly two to four times the 2.4 trillion parameters of Qwen3.8-Max, and it applies to the generations after Qwen 4, not Qwen 4 itself.
Parameter counts on their own say little about cost. Every recent Qwen flagship is a sparse Mixture-of-Experts model, so only a fraction of parameters are active for each token: Qwen3.8-Max's open checkpoint is named A95B, meaning about 95 billion active out of 2.4 trillion. A 10-trillion-parameter model with a similar sparsity ratio would still need data-centre clusters just to hold its weights, but the compute per token could remain manageable. The n-gram embedding approach in Flash-Next is another route to adding parameters without adding much compute.
The roadmap was paired with hardware. Alibaba's T-Head unit unveiled the Zhenwu V900, which Investing.com reported offers three times the performance of the previous M890, supports clusters of up to 500,000 chips and enters mass production in Q1 2027. Alibaba also set a target of more than 20GW of global data-centre capacity by 2032. CEO Eddie Wu said "the industry's mid-to-long-term demand far outpaces our supply capabilities", and argued that "the truly groundbreaking products of the Machine Intelligence era have not yet arrived". Alibaba's Hong Kong shares rose 5.1% on the news, according to the same report.
For the wider chip context, including why Chinese labs are designing around domestic accelerators, see our analysis of China's AI chip race.
Bart Slodyczka tests the 16GB and 24GB M6 Mac mini for local AI and explains what memory size means for the models you can run.
What Hardware Will Qwen 4 27B Need?
A 27-billion-parameter model needs roughly 14-17GB of memory for its weights at 3-4-bit quantisation, plus extra for context. Because Qwen 4 27B has no model card yet, the most realistic proxy is its predecessor, Qwen3.8-27B, for which community quantisations from Unsloth give real file sizes. The rule of thumb is simple: memory for weights = parameters × bits per weight ÷ 8. For 27 billion parameters at about 4.85 bits per weight (typical of Q4_K_M), that is 27 × 4.85 ÷ 8 ≈ 16.4GB, close to the real 17.1GB file once the vision encoder and embeddings are included.
| Quantisation | Qwen3.8-27B file size | M6 Mac mini 16GB | M6 Mac mini 24GB | 24GB GPU (RTX 3090/4090) |
|---|---|---|---|---|
| BF16 (full) | 54.7GB | No | No | No |
| Q8_0 | 29GB | No | No | No |
| Q6_K | 22.9GB | No | No | Barely; almost no room for context |
| Q5_K_M | 19.8GB | No | Not practical | Yes, modest context |
| Q4_K_M | 17.1GB | No | Tight: raise GPU limit, short context | Yes, the sweet spot |
| Q3_K_M | 13.8GB | No | Yes | Yes, long context |
| 2-bit | 10.7GB | Only just; quality suffers | Yes | Yes |
File sizes are Unsloth's Qwen3.8-27B quantisations, as covered in our Qwen3.8-27B review. Fit columns are our estimates for Qwen 4 27B, assuming a similar size. They will change if the final model differs.
The M6 Mac mini: 16GB vs 24GB
The M6 Mac mini went on sale on 22/09/2026, the same day as the Apsara keynote. It starts at $899 in the US with 16GB of unified memory (approx. £675 at a straight conversion, though Apple's VAT-inclusive UK price is higher) and is available with 16GB, 24GB or 32GB. 9to5Mac and MacRumors report memory bandwidth of 153GB/s on the 16GB model and 170GB/s on the 24GB and 32GB models, up from 120GB/s on the M4.
The catch is that macOS does not give the GPU all of that memory. By default, Apple Silicon Macs of this size let the GPU use roughly two-thirds of unified memory, the rest being reserved for the system and apps. That leaves roughly:
- 16GB Mac mini: about 10.7GB for the model. Only a 2-bit 27B fits, and with almost no room for context. For this machine, a smaller model (8-14B class) or a sparse MoE with few active parameters is the better choice.
- 24GB Mac mini: about 16GB by default. A Q3_K_M 27B (13.8GB) fits with room for a few thousand tokens of context. Q4_K_M (17.1GB) needs the GPU memory limit raised (advanced users do this with the
iogpu.wired_limit_mbsystem setting) and a short context, and leaves little headroom for anything else. - 32GB Mac mini: about 21GB by default, which runs Q4_K_M comfortably and makes this the configuration we would choose for a 27B model.
Speed: token generation on a Mac is limited mainly by memory bandwidth, because every generated token has to read the active weights once. A useful ceiling is bandwidth divided by model size: 170GB/s ÷ 17.1GB ≈ 10 tokens per second for Q4_K_M on a 24GB M6, and about 12 tokens per second for Q3_K_M. Real-world speeds are usually below that ceiling, though multi-token prediction can push effective speed higher. For comparison, Simon Willison reported 15-30 tokens per second running the Qwen3.8-27B Q4_K_M build on a 128GB Apple Silicon MacBook Pro with far higher bandwidth. These are our estimates, not measured Qwen 4 results.
Context memory: the KV cache grows with context length. For a conventional dense model this can add several gigabytes at 32K tokens. If Qwen 4 27B uses the Flash-Next hybrid design, where only one layer in four keeps a full cache, that overhead would shrink considerably, one reason the new architecture could suit memory-limited machines.
GPUs and other options
On a PC, a single 24GB card (RTX 3090 or 4090) is the sweet spot for a 4-bit 27B model; one commenter on Willison's review reported 59 tokens per second for Qwen3.8-27B on an RTX 3090 with multi-token prediction enabled. A 16GB card will need 3-bit or lower, or partial CPU offload. For full-precision inference or serving several users, you are into 48-80GB workstation or data-centre GPUs. Tools like Ollama and LM Studio will almost certainly add Qwen 4 27B within days of release, as they did for Qwen3.8-27B. If you are weighing a Mac for local AI more generally, our earlier piece on the Mac mini local-AI frenzy and our look at Perplexity's Lily local AI engine are useful background.
How Might Qwen 4 Compare With DeepSeek, Kimi, GLM and US Models?
Qwen 4 cannot be benchmarked yet, so any comparison is with what its rivals have already shipped. The table below sets out where the competition stands on 29/09/2026, which is the bar Qwen 4 will have to clear.
| Model | Status on 29/09/2026 | Size | Open weights | API price per million tokens (in/out) |
|---|---|---|---|---|
| Qwen 4 family | In training | Undisclosed | 27B expected, unconfirmed | Not announced |
| Qwen3.8-Max | GA since 03/08/2026 | 2.4T MoE (~95B active) | Yes, custom licence | approx. £1.50 / £4.50 ($2 / $6) |
| DeepSeek V4.1 Flash | Released 10/09/2026; V4.1 Pro not yet released | Not officially published | Not stated in changelog | Reduced on launch; see DeepSeek pricing page |
| Kimi K3 (Moonshot) | API since 16/07/2026; K3.1 identifier spotted 29/09, unconfirmed | 2.8T MoE, 1M context | Yes, Modified MIT | See our Kimi K3 review |
| GLM 5.3 (Z.ai) | Released 14/08/2026 | 743B MoE (~40B active) | Promised after safety hardening | See our GLM 5.3 review |
| GPT-6 Sol (OpenAI) | Available | Undisclosed | No | approx. £1.50 / £7.50 ($2 / $10) |
| GPT-6 Astra (OpenAI) | Available | Undisclosed | No | approx. £7.50 / £37.50 ($10 / $50) |
| Claude Opus 5.5 (Anthropic) | Launched 22/09/2026 | Undisclosed | No | approx. £3 / £15 ($4 / $20) |
Qwen 4 vs DeepSeek V4.1
DeepSeek is the rival Qwen 4 Flash most directly targets. DeepSeek's changelog describes V4.1-Flash, released on 10/09/2026, as "the smallest model in our new architecture family, with native multimodal visual understanding", which implies bigger V4.1 models to come. DeepSeek has also extended its V4 Pro API beyond 14/09/2026 instead of retiring it. Both labs are now shipping new architectures with small models first, so a Qwen 4 Flash versus DeepSeek V4.1 Flash price war is likely. Our DeepSeek V4.1 Flash review and DeepSeek V4 Pro review have the details.
Qwen 4 vs Kimi K3 and K3.1
Moonshot's Kimi K3 is the open-weight model Qwen 4 Max has to beat on licence transparency: 2.8 trillion parameters, a one-million-token context and a Modified MIT licence. On 29/09/2026, TechNode relayed a PConline report that developers had spotted a kimi-k3-1 identifier in Moonshot's API model registry, with a preview suggesting up to one million tokens of context and several reasoning levels. Moonshot has not confirmed it, published documentation or given a date, so treat K3.1 as unconfirmed. For Moonshot's longer-term plans, see Kimi K4: what's confirmed vs rumoured.
Qwen 4 vs GLM
Z.ai's GLM 5.3 took a different path: it uses the same 743-billion-parameter base as GLM 5.2 and gets its gains purely from post-training. That makes GLM a useful contrast. Alibaba is betting on a new architecture and much larger scale; Z.ai showed how far post-training alone can go. GLM's smaller Flash variant, covered in our GLM-5.3-Flash review, competes in the same space as Qwen 4 Flash.
Qwen 4 vs US frontier models
The US frontier moved again in September. Anthropic launched Claude Opus 5.5 on the same day as the Apsara keynote, and OpenAI's GPT-6 family (Sol, Astra and Luna) sets the price and performance range. The pattern of the past year is that Chinese flagships trail the very best US models by a few months on general intelligence whilst matching or beating them on specific coding and agent benchmarks, at a fraction of the price. Qwen3.8-Max, for example, beat Claude Fable 5 on Terminal-Bench 2.1 (86.6 vs 84.6) but trailed clearly on SWE-bench Pro (67.7 vs 80.0). There is no evidence yet on where Qwen 4 lands. The sensible expectation is that Qwen 4 Max narrows the gap and undercuts on price, rather than overtakes outright.
Where Does Qwen 4 Fit in the US-China AI Race?
Qwen 4 is Alibaba's bid to stay at the front of China's open-weight model race whilst building a full domestic stack underneath it. The Apsara announcements tied together models (Qwen 4, 4.5 and 5), chips (Zhenwu V900 and the Yitian 720 and 730 CPUs due in 2027), cloud capacity (20GW by 2032, with new regions in Türkiye, Finland and the Netherlands) and consumer hardware, from the Qwen Book agentic computer to Qwen-branded glasses and earbuds, as TechNode Global reported.
That vertical integration is partly a response to export controls. Chinese labs have less reliable access to top Nvidia hardware, which is one reason efficiency-focused designs like Flash-Next (6B active parameters, offloadable embeddings) matter so much. It is also why open weights remain central to Chinese labs' strategy: they build global developer adoption that closed Chinese APIs struggle to win in Western markets. For the policy side of this story, read why the US-China AI "race" framing is a mistake, and for the debate over open weights and chips, the open-weights letter and Anthropic's response.
For UK and European organisations the practical question is less geopolitical than operational: an open-weight Qwen 4 27B can be run entirely on your own hardware, with no data leaving your network, which sidesteps many of the data-residency concerns that come with using any Chinese-hosted API. The hosted Max, Plus and Flash tiers are a different decision and deserve the same due diligence as any overseas provider.
Who Should Wait for Qwen 4?
Whether to wait depends on what you need and when.
- Local AI hobbyists with a 24GB GPU or a 32GB Mac: worth waiting for Qwen 4 27B, but there is no reason to stop using Qwen3.8-27B today. Your setup will carry over.
- Anyone buying a Mac mini for local AI now: choose the 24GB model at minimum and the 32GB model if you want 4-bit 27B-class models to run comfortably. The 16GB model is fine for smaller models but not for a 27B.
- Developers building on the Qwen API: keep shipping on Qwen3.8-Max or Qwen3.8-Flash. Design your code so the model ID is a config value, because Qwen 4 tiers will likely drop into the same OpenAI-compatible API.
- Teams choosing an open-weight flagship: Kimi K3 and Qwen3.8-2.4T-A95B are available now; check their licences against your revenue and use case. Waiting for Qwen 4 Max only makes sense if you can tolerate an unknown timeline.
- Enterprises needing the strongest model today: Qwen 4 is not an option yet. Compare current US and Chinese flagships on your own tasks, and see our Gemini 4 guide for another model on the horizon.
If you want a hosted way to try Qwen models without setting up an account with Alibaba, OpenRouter lists several, such as Qwen3 Max and Qwen3-Next 80B, and our directory has a profile of Qwen3.7 Max. For the rival Chinese options, see DeepSeek V4 Pro and GLM 5.2.
Sources
- Alibaba Cloud press room: Alibaba Unveils Roadmap on Full-Stack AI Strategy (22/09/2026)
- Alibaba Cloud blog: 2026 Apsara Conference roadmap and global expansion
- Qwen3.8-Flash-Next official model card (Hugging Face)
- RTHK: Alibaba set for AI leap with Qwen and next-gen chip
- Investing.com: Alibaba plans AI model with 5-10 trillion parameters, unveils new chip
- Pandaily: Alibaba puts Qwen4 family into training
- Vietnam Investment Review: Alibaba targets 10 trillion parameters with Qwen 4
- OrcaRouter: Qwen 4 Max announced at Apsara 2026, the four tiers
- The Decoder: Alibaba releases Qwen3.8-Flash-Next
- TechNode: Qwen to open-source Qwen3.8-Flash-Next, previewing Qwen4 architecture
- TechNode: Kimi K3.1 identifier reportedly surfaces in Moonshot API registry
- DeepSeek API changelog
- 9to5Mac: M6 Mac mini review and MacRumors: 2024 vs 2026 Mac mini buyer's guide
The Bottom Line
Qwen 4 is real, in training, and coming in four tiers, but as of 29/09/2026 it is an announcement rather than a product. Alibaba has confirmed the training status and a 5-10 trillion parameter roadmap for Qwen 4.5 and Qwen 5; the tier names come from the stage; and release timing, pricing, licences and benchmarks are all unknown. The best guide to what is coming is Qwen3.8-Flash-Next, which shows Alibaba optimising hard for efficiency: few active parameters, cheaper long context and offloadable embeddings.
For most readers the model that matters is Qwen 4 27B. If it follows its predecessor, plan on about 17GB for a 4-bit build: easy on a 24GB GPU, workable on a 32GB M6 Mac mini, tight on the 24GB model and out of reach on 16GB. Until it ships, Qwen3.8-27B is a strong stand-in, and nothing you set up today will be wasted.
Last updated: 29/09/2026. We will update this article when Alibaba publishes a Qwen 4 model card, pricing or release date.








