Quick Answer:
Ling 3.1 Flash is InclusionAI (Ant Group)'s hybrid-reasoning Mixture-of-Experts model: about 560B total and 25B active parameters, 262K tokens of served context (1M is a design target), built for coding and tool-using agents. It is free until 13/10/2026 on Vercel AI Gateway, OpenCode and Kilo Code. Post-trial pricing, weights and licence are not yet published, and the benchmark scores are provider-reported.
A 560-billion-parameter model for £0 sounds too good to last, and it is: the free window closes on 13/10/2026. Ling 3.1 Flash arrived on 29/09/2026 with an agent-heavy benchmark card, a free route through Vercel and OpenCode, and very little else in the way of documentation.
This review separates what is confirmed from what is merely repeated: the specifications, the access routes, the numbers, and the long list of things nobody has published yet. Where we could not verify something, including the Hugging Face model card, we say so.
Executive Summary
- Released 29/09/2026 (per Build Fast with AI), with Vercel announcing availability on its AI Gateway on 30/09/2026. It is the newest model from InclusionAI, the AGI research lab of Ant Group.
- Sparse Mixture-of-Experts, hybrid reasoning: roughly 560 billion total parameters with roughly 25 billion active per token, aimed at coding, multi-step analysis and tool-using agents.
- 262,144-token context on gateways (Vercel AI Gateway, Kilo Code), with up to 32,768 output tokens. InclusionAI's launch material describes a design target of up to 1 million tokens, which no public route currently serves.
- Free to use until 13/10/2026 through Vercel AI Gateway (model ID
inclusionai/ling-3.1-flash-free), and listed at £0 per million tokens on OpenRouter and Kilo Code. The free ID stops serving when the promotion ends; the standard ID then starts billing at a rate that has not been published. - Provider-reported benchmarks include CyberGym 87.9%, DRACO 85.5%, SkillsBench 68.7%, HealthBench Professional 65.3%, Finance Agent v2 57.9%, SWE Atlas Codebase QnA 55.9%, AutomationBench 52.5% and Terminal-Bench 4.0 at 40.4%. These come from launch screenshots, not an independent evaluation.
- Not yet open: aggregators report that weights and a licence have not been released, and one comparison says open-sourcing is promised after the trial. We could not read the Hugging Face model card directly, so treat any open-weights claim as unverified.
- Our verdict: a serious free option for experimenting with agent and coding workflows this fortnight, but not something to build a production dependency on until pricing, licence and independent results appear.
Lineage: Ant Group, InclusionAI and the Ling Family
InclusionAI is described by LLM Reference as Ant Group's AGI research lab, based in Hangzhou, China, and founded in 2023. Its Ling line is the language-model branch of a wider family that also includes reasoning models and diffusion-based language models. We have covered several siblings: the small and efficient Ling 3.0 Tiny (7.9B total parameters, 1.3B active, 256K context), and the experimental LLaDA2.2-Flash diffusion agent from the same group.
Ling 3.1 Flash represents a large step up in scale from its immediate predecessor. According to Build Fast with AI, Ling 3.0 Flash was a 124B total, 5.1B active model. Ling 3.1 Flash is therefore roughly 4.5 times larger in total parameters and roughly 5 times larger in active parameters. That ratio matters: active parameters determine how much computation each generated token costs, so a 25B-active model should be slower and more expensive to serve per token than its predecessor, even though it is still a fraction of the cost of running a dense 560B model.
The "Flash" label is slightly misleading if you assume small. In the 2026 open-model landscape, several Chinese labs use Flash for their mid-size agent-oriented tier. For context, the comparison site Threat Frontier lists GLM 5.3 Flash at 320B total and 18B active, and Qwen 3.8 Flash Next at 125B total and 6B active. Ling 3.1 Flash is the largest of the three on both counts. Our own reviews of GLM 5.3, GLM-5.3-Flash and Qwen 3.8 Max give more background on those rivals.
Julian Goldie SEO shows how to select Ling 3.1 Flash in OpenCode, build a website from one prompt, and connect it to the DeepSeek Harness.
Architecture and Training: What Is and Is Not Published
Be careful with architecture claims for this model. What the gateway listings and aggregators agree on is short: a Mixture-of-Experts design, about 560B parameters in total, about 25B activated per token, hybrid reasoning (the model can think before answering, and the OpenCode demo shows selectable thinking modes), text input, function calling, structured JSON outputs and reasoning tokens. Kilo Code's listing confirms function calling, tool choice and JSON-schema structured outputs.
What is not published, as far as we could establish, is the layer-level detail: number of experts, routing scheme, attention design and training-token count. Threat Frontier notes that the architecture is unpublished and that the predecessor used a 5:1 stack of Kimi Delta Attention and gated Mamba-style linear attention. That is a fact about the earlier model only, and we would not assume the same recipe carries over. If you see a detailed architecture diagram attributed to Ling 3.1 Flash, check where it came from.
The long-context story deserves the same caution. InclusionAI describes a design capability of up to one million tokens. Every route we could check serves less. Vercel's listing, routed through the provider Novita, shows 262,144 tokens of context and a 32,768-token maximum output. LLM Reference describes the free trial as capped at 256K and says paid access is planned for the 1M window; its figure differs slightly from Vercel's 262,144, which is the sort of rounding discrepancy you should expect between aggregators. The practical reading is simple: plan around roughly a quarter of a million tokens today, and treat 1M as a roadmap item.
Capabilities Deep Dive
Vercel's changelog describes the model as built for "coding, multi-step analysis, and agents that use tools", particularly where the work involves long documents, code and extended task histories. That is the standard pitch for the 2026 agent-model tier, but the benchmark selection InclusionAI chose to publish does back it up: nearly everything on the list is an agentic or tool-use evaluation rather than a classic exam.
Coding and terminal work
The two coding-flavoured numbers are Terminal-Bench 4.0 at 40.4% and SWE Atlas Codebase QnA at 55.9%. Terminal-Bench tests an agent's ability to complete tasks in a command-line environment; the Codebase QnA benchmark asks questions about real repositories, so it measures code comprehension rather than patch writing. Neither is the SWE-bench-style fix-the-issue result that most readers use to compare coding models, and InclusionAI has not published one that we could find. Threat Frontier states plainly that the only benchmark shared by GLM 5.3 Flash (63.4) and Qwen 3.8 Flash Next (58.7), DeepSWE 1.1, is "not published" for Ling 3.1 Flash. Anyone claiming it beats those rivals at coding is guessing.
Agents and tool use
The agent results are the headline. CyberGym at 87.9% and DRACO at 85.5% are the highest numbers on the card, followed by SkillsBench at 68.7%, Finance Agent v2 at 57.9% and AutomationBench at 52.5%. Kilo Code, a coding-agent platform supporting 500+ models, lists the model with tool choice enabled and no content moderation layer, and the YouTube demo shows it driving OpenCode and the DeepSeek Harness. The presenter's own verdict is balanced: it is not frontier-level in the way Claude Opus is, but it is fast and useful for agents and sub-agents.
Long-document and research work
DRACO is described by Build Fast with AI as research and analysis, and HealthBench Professional (65.3%) is a professional-knowledge test. Combined with the 262K window, that points to a model suited to reading long inputs and producing structured analysis. Again, we have no independent long-context retrieval test (needle-in-a-haystack, RULER or similar), so the 262K figure is a limit, not a measure of quality at that length.
Benchmarks: Real Numbers and Honest Caveats
Here is the complete set of numbers we could find. All were reported by the provider through launch screenshots (BenchLM traces them to posts from @AntLingAGI) and repeated by aggregators. BenchLM covers only 8 of its 645 tracked benchmarks for this model and therefore leaves it unranked overall.

- CyberGym 87.9%: cybersecurity agent tasks. Treat with care: a strong score here is a capability signal for both defenders and attackers (see the safety section below).
- DRACO 85.5%: research and analysis tasks.
- SkillsBench 68.7%: how well an agent uses supplied skills or tool packages.
- HealthBench Professional 65.3%: professional-level health knowledge. Not a basis for any clinical use.
- Finance Agent v2 57.9%: financial analysis workflows.
- SWE Atlas Codebase QnA 55.9%: answering questions about codebases.
- AutomationBench 52.5%: automation workflows.
- Terminal-Bench 4.0 40.4%: terminal-based agent tasks. The lowest on the list, and the most relevant to developer-agent use.
- GDPval-AA v2.1 1,673 Elo: appears in the Build Fast with AI scorecard graphic as a general professional-tasks rating. It was not in the text tables of the other aggregators we read, so we list it as a single-source figure.
Honest caveats. First, there is no model-card text we could verify: the Hugging Face page for inclusionAI/Ling-3.1-flash returned an authorisation error to our fetches, so every number above is second-hand. Second, there are no published comparison columns, so we cannot tell you whether 40.4% on Terminal-Bench 4.0 is good against Claude, GPT or Kimi on the same harness. Third, the benchmark suites are newer and less standard than the SWE-bench family, which makes cross-model comparison hard. Build Fast with AI itself warns that provider-reported benchmarks use different evaluation suites. Fourth, a high score on a handful of hand-picked benchmarks is what every launch post contains. Wait for independent runs before treating this as a ranking.
Speed and Latency
Build Fast with AI reports about 3.1 seconds latency and about 88 tokens per second on the Vercel route served by Novita. These are single-provider, single-moment measurements on a free promotional endpoint, and the free tier is shared by a lot of curious users this week. The YouTube demo supports that caution: Julian Goldie notes free-server errors that forced retries during his build. Expect variable performance until the promotion ends, and do not extrapolate the 88 tokens per second figure to your own workload.
At 25B active parameters, a 560B MoE model needs a large amount of memory even though each token is cheap to compute. All 560B parameters must be resident to serve requests, which is why self-hosting is the territory of multi-GPU nodes rather than a workstation. That is academic today, since no weights are available, but it is a useful sanity check when the open release arrives. For genuinely local work, our Qwen3.8-27B review covers a model you can actually run on a high-end consumer setup.
Safety, Security and Data Handling
We did not find a system card for Ling 3.1 Flash, nor any published RSP-style evaluation of cyber, biological or autonomy risk. That absence is notable because the one benchmark at the top of the list, CyberGym, is a cybersecurity evaluation. A model scoring 87.9% on agentic cyber tasks, served free with content moderation disabled on Kilo Code's listing, deserves more safety documentation than has so far been published. We make no claim about misuse; we simply flag that the evidence needed to judge it is missing.
Data handling is the other practical concern. Free promotional endpoints routed through third-party inference providers (Vercel names Novita for its route) may log prompts, and the free ID is a temporary promotion. Do not send customer data, credentials or proprietary source code to a free endpoint unless you have read the provider's data policy and are comfortable with it. The same guidance applies to any hosted Chinese-lab model, and equally to hosted Western models: know where your data goes.
How to Use It Free Before 13/10/2026
There are three practical routes, all confirmed by at least one source. The Vercel route is the best documented.
- Vercel AI Gateway: model IDs
inclusionai/ling-3.1-flash(standard) andinclusionai/ling-3.1-flash-free(promotional). Vercel's changelog, published 30/09/2026, says the model is free to use through 13/10/2026, the standard ID begins billing once the promotion ends, and the free ID stops serving. Setup for Claude Code and Codex is vianpx vercel ai-gateway setupfollowed by/model. - OpenCode: the video shows the model appearing in the free list after updating OpenCode, with a choice of thinking modes, and then being reused in the DeepSeek Harness by adding the OpenCode API key. If it does not appear, the presenter's advice is to check for an update.
- Kilo Code and OpenRouter: Kilo lists inclusionai/ling-3.1-flash at no cost (creation date 02/10/2026 on its page), and LLM Reference lists OpenRouter and Vercel as the two providers, both at zero price.
For terminal-agent users, our guides to the DeepSeek Harness and our comparison of the best AI coding agents explain the harnesses this model plugs into. For the reasoning-tier alternatives, see Kimi K2.7 Code and MiniMax Code 2.0.
Pricing: Free Now, Unknown Later
Today the cost is £0 per million input and output tokens on every route we checked. The important question is what happens after 13/10/2026. Build Fast with AI states that final token pricing is not established in current provider listings, and Threat Frontier says the same: free trial, paid rate unpublished. Anything you read quoting a post-trial price is speculation.
To put the unknown in context, Threat Frontier lists the paid rates of two rivals: GLM 5.3 Flash at $0.15 input and $0.50 output per million tokens (roughly £0.11 and £0.38 at about £0.75 to the dollar), and Qwen 3.8 Flash Next at $0.16 and $0.47 (roughly £0.12 and £0.35). A 560B model will probably be priced above the smaller ones, but that is our inference, not a fact. Budget for the possibility of a meaningful charge, and do not wire a free-tier model ID into production code. If you must test inside a product, put the model ID in configuration so you can switch in minutes.
Compare it also with our coverage of DeepSeek V4.1 Flash and DeepSeek V4 Flash 0731 if cost per token is your main criterion; those have published prices.
Open Weights, Licensing and the Hugging Face Question
A Hugging Face page for inclusionAI/Ling-3.1-flash appears in search results, but every automated request we made to it was rejected as unauthorised, so we could not read its licence, tags or file list. Meanwhile, the aggregators disagree in tone: LLM Reference says "Proprietary with conditional commercial use" and "weights: not released"; Build Fast with AI says weights and licence are not yet publicly available and that a future release is indicated but unconfirmed; Threat Frontier says open source is "promised after trial"; BenchLM labels it proprietary with no licence information.
The sensible synthesis is this: at the time of writing, the model is hosted-only, a future open release has been hinted but not dated, and the licence terms are unknown. If the weights do arrive, a 560B MoE will be a major open-weights event, comparable in scale to the models we covered in our DeepSeek V4 Pro and Kimi K3 launch pieces. Until then, please do not describe it as an open-source model.
Real-World Use vs Benchmarks
The only hands-on evidence we have is the Julian Goldie demonstration, and it is deliberately modest. He builds a website from a single prompt in OpenCode, hits free-server errors that require retries, and concludes the model is quick and good for agents and sub-agents but not at Opus level. That is a sensible reading: a 25B-active model will usually trail the very largest dense or heavily-activated frontier models on hard reasoning, while being cheap and quick enough to use for the bulk of an agent's routine steps.
This points to the most credible use pattern: use a frontier model as planner or reviewer, and delegate high-volume execution (file edits, test running, log reading, summarising) to a cheap, fast model such as this one. Whether Ling 3.1 Flash is the right executor depends on its tool-call reliability, which none of the published benchmarks measures directly. SkillsBench is the nearest proxy, at 68.7%.
Limitations
- Pricing unknown after 13/10/2026; the free ID stops serving on expiry.
- Context gap: 262K served versus a 1M design target.
- Output cap: 32,768 tokens on the trial route, which limits very long single-shot generations.
- No verified weights or licence: hosted-only for now.
- Benchmarks are provider-reported, drawn from an unusual set of suites, with no head-to-head table and no widely used SWE-bench-style result.
- No system card or safety evaluation found, despite strong cyber-agent scores.
- Free-tier reliability: retries and variable latency during the launch rush.
- Single-source details: the 124B/5.1B predecessor specs, the 88 tokens-per-second figure and the GDPval-AA Elo come from one outlet each.
How It Compares
A fair table has to rely on what is published for each model, and the gaps are the story.
| Model | Total / active | Context | Paid price per 1M tokens (in / out) | Weights |
|---|---|---|---|---|
| Ling 3.1 Flash | 560B / 25B | 262K served, 1M target | Unpublished (free until 13/10/2026) | Not released |
| GLM 5.3 Flash | 320B / 18B | 1M | About £0.11 / £0.38 ($0.15 / $0.50) | MIT, on Hugging Face |
| Qwen 3.8 Flash Next | 125B / 6B | 262K native, 1M with YaRN | About £0.12 / £0.35 ($0.16 / $0.47) | qwen-community-1.0, on Hugging Face |
| Ling 3.0 Flash | 124B / 5.1B | Not checked | Not checked | Not checked |
Rival figures are from Threat Frontier (prices, parameters, licences) and Build Fast with AI (Ling 3.0 Flash). Sterling figures are our approximate conversions. On the one benchmark shared by the two rivals, DeepSWE 1.1, GLM 5.3 Flash scores 63.4 and Qwen 3.8 Flash Next 58.7, while Ling 3.1 Flash has no published result. If you need a model with downloadable weights and a published price today, GLM 5.3 Flash and Qwen 3.8 Flash Next are the verifiable choices. For the largest Chinese open models, read our Qwen 4 and Kimi K3 vs Claude Fable 5 analyses.
Who Should Use It
- Agent tinkerers and open-source harness users: a free 560B-class model for the next fortnight is worth an afternoon in OpenCode, Kilo Code or the DeepSeek Harness.
- Developers evaluating sub-agent models: run your own task set, log tool-call failures and compare with your current executor before the price lands.
- Researchers: the benchmark selection (CyberGym, DRACO, SkillsBench) is unusual and worth reproducing independently.
- Not for: regulated or confidential data, production dependencies on a free ID, or anyone who needs open weights today.
A practical test plan: pick 20 real tasks from your own backlog, run them on Ling 3.1 Flash and on your current default model, record success, retries, tokens and wall-clock time, and finish before 13/10/2026. If it holds up, you will have the data to decide when the price arrives.
The Bottom Line
Ling 3.1 Flash is a big, ambitious release from InclusionAI: roughly 560B total and 25B active parameters, a 262K served context, a 1M design target and an agent-first benchmark card. The free window to 13/10/2026 makes it easy to try. But the evidence is thin and mostly second-hand: pricing, licence, weights, a system card and independent benchmarks are all missing or unverified.
Try it now, because it costs nothing. Judge it on your own tasks. And hold off on any commitment until InclusionAI publishes the details that this launch left out.
Sources
- Vercel changelog: Ling 3.1 Flash is now available on AI Gateway (30/09/2026; model IDs, free period, 262K context, setup).
- Vercel AI Gateway model page (262,144 context, 32,768 max output, provider Novita).
- Build Fast with AI: Ling 3.1 Flash review (release date, benchmarks, latency, Ling 3.0 Flash comparison; benchmark figure credit).
- BenchLM: Ling 3.1 Flash (benchmark list, provenance from @AntLingAGI posts, unranked status).
- LLM Reference: Ling-3.1-flash (InclusionAI background, licence and weights status, providers).
- Kilo Code: inclusionAI Ling 3.1 Flash (zero price, capabilities).
- Threat Frontier: Ling 3.1 Flash vs GLM 5.3 Flash vs Qwen 3.8 Flash Next (rival specs and prices).
- Hugging Face: inclusionAI/Ling-3.1-flash (listed, but could not be retrieved for this article).
- Julian Goldie SEO: Ling 3.1 Flash is INSANE (FREE!) (OpenCode demonstration).
Last updated: 05/10/2026. Sourced from Vercel, Build Fast with AI, BenchLM, LLM Reference, Kilo Code and Threat Frontier. Post-trial pricing, weights, licence and independent benchmarks were not available at the time of writing, and this article will be revised when they are.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.








