Reports of an AI model measured in the tens of trillions of parameters would once have sounded like marketing hyperbole. On 7 August 2026, the Financial Times gave the claim a byline: ByteDance, the Beijing-headquartered owner of TikTok and Doubao, is reportedly training a model that could reach 10 trillion parameters, a scale the paper's sources say is intended to approach Anthropic's Mythos, currently the most capable model family in the West.
Nothing about this story is confirmed by ByteDance itself. It rests on three anonymous sources, an early-stage training run, and a company that has not responded to requests for comment. That is enough to take seriously, not enough to take at face value. This article works through what is actually known, why the parameter-count framing can be misleading, and what would need to happen for a 10-trillion-parameter ByteDance model to matter in practice.
Note: this analysis is based on the Financial Times' 7 August 2026 report (via wire and syndication coverage), prior reporting on Kimi K3's training compute, and Anthropic's public model documentation. Where a figure is an estimate rather than an officially confirmed number, it is labelled as such throughout.
A rundown of the Financial Times report on ByteDance's reported 10-trillion-parameter model, alongside Meta's new coding agent and OpenAI's free-tier expansion.
Executive summary
- Reported 7 August 2026 by the Financial Times, citing three people familiar with the project; ByteDance had not responded to a request for comment at time of writing.
- Up to 10 trillion parameters, more than 3x Moonshot AI's Kimi K3 (2.8T), the largest openly released Chinese model so far.
- Still in early pre-training, a phase the FT's sources say will take three to six months on its own, before fine-tuning and evaluation even begin.
- Intended to rival Anthropic's Mythos, though Anthropic does not disclose official parameter counts, so the comparison is against industry estimates (Mythos 5 ≈ 8T, Fable 5 ≈ 5T), not confirmed figures.
- No benchmark data exists yet. This is a training-run report, not a launch, so there is nothing to test independently.
- Framed internally around originality: founder Zhang Yiming has reportedly told staff to avoid leaning on distillation from rival models for short-term gains.
What the Financial Times actually reported
The FT's account, since syndicated and summarised across outlets including Reuters-linked wire coverage, TechNode-adjacent trade press and multiple AI newsletters, says ByteDance is training a model with as many as 10 trillion parameters. The company's scale is not in question, ByteDance is one of the best-resourced technology companies in the world, but the specific figure, timeline and strategic intent all come from unnamed sources rather than an official ByteDance statement.
Three details stand out as more specific than the usual anonymous-sourcing boilerplate, which is part of why the report has been taken seriously rather than dismissed as speculation: the pre-training phase is described as requiring three to six months on its own, the final parameter count is explicitly described as not yet fixed, and the strategic comparison point named is Anthropic's Mythos rather than a vaguer "frontier lab" framing. That level of detail suggests sources with genuine visibility into the project, even though none of it is independently verifiable from outside ByteDance.
What the report does not contain is equally important: no architecture details (dense versus mixture-of-experts, layer count, context window), no training-data description, no compute budget, no confirmed release window, and no benchmark projections. This is a report about a training run in progress, not a product announcement.
Putting 10 trillion parameters in context
Parameter counts for Western frontier labs are almost never officially confirmed, Anthropic, OpenAI and Google DeepMind all treat the figure as effectively proprietary, so every comparison below mixes one reported number (ByteDance) with several industry estimates (everyone else). Treat the table as a rough map of scale, not a leaderboard.
| Model | Parameters | Status |
|---|---|---|
| ByteDance (unnamed, in training) | Up to 10T (reported) | Early pre-training, unconfirmed |
| Anthropic Mythos 5 | ≈8T (industry estimate) | Released, parameter count unconfirmed |
| Anthropic Fable 5 | ≈5T (industry estimate) | Released, parameter count unconfirmed |
| Moonshot AI Kimi K3 | 2.8T (confirmed) | Released, open weights |
| Alibaba Qwen 3.8 Max | 2.4T (confirmed) | Released |

Even taken at face value, the number places ByteDance's project above every publicly confirmed parameter count from any lab, East or West. That alone explains the attention the report has drawn, but it also means the comparison to Mythos is being made against an estimate Anthropic has never confirmed, so "bigger than Mythos" is a claim about two unverified or partially unverified numbers, not a settled fact.
Why parameter count is a weak signal of capability
Total parameter count made intuitive sense as a capability proxy in the dense-model era, when every parameter activated for every token, so a bigger model cost proportionally more to run and (roughly) knew more. Frontier models since 2024 have almost universally moved to sparse mixture-of-experts (MoE) architectures, where a routing mechanism activates only a small subset of the model's total parameters for any given token. Kimi K3's 2.8 trillion total parameters, for example, activate only a fraction of that per token in practice, which is precisely why a 2.8T MoE model can run at a fraction of the inference cost a 2.8T dense model would require.
This matters directly for how to read the ByteDance report. A 10-trillion-parameter MoE model with a small active-parameter fraction could plausibly be cheaper to run than Kimi K3, while a hypothetical 10-trillion-parameter dense model would be a genuinely enormous and expensive system to train and serve. The FT's report does not specify which, and until ByteDance discloses an architecture, headline comparisons of "10 trillion versus 8 trillion" are comparing two numbers whose practical meaning depends entirely on details neither company has published.
The more reliable lesson from 2025 and 2026's parameter-count race is that benchmark performance has repeatedly diverged from raw scale. Several smaller, more carefully trained models have outperformed larger predecessors on independent evaluations, and Meta's own Muse Spark 1.2 benchmark charts show a model trailing Claude Opus 5 despite competitive scale. Parameter count is a headline, not a leaderboard position.
ByteDance's position: Doubao, Seed and TikTok's data edge
ByteDance is not entering the frontier race from nowhere. Its Doubao assistant is one of the most-used AI products in China, and the company has shipped a steady cadence of Seed-branded models (Seed, Seedance for video, and others) into both consumer and developer channels over the past two years. What it has lacked, relative to Alibaba's Qwen, Moonshot's Kimi and DeepSeek's V-series, is a single model widely regarded as competitive with the absolute frontier rather than a strong regional alternative.
ByteDance's structural advantage is distribution and data. TikTok's global user base and Doubao's domestic reach give the company both a training-data pipeline and a deployment surface that few rivals can match, resources that matter as much for post-training and reinforcement learning as raw pre-training scale does. According to the FT's sources, founder Zhang Yiming has told staff to avoid leaning on distillation, using outputs from other labs' models to accelerate training, for short-term gains. Distillation has been a persistent point of controversy across the Chinese AI sector (and was alleged, though not proven, against several labs during 2025), and an explicit internal push against it reads as ByteDance positioning this specific model as an original, defensible asset rather than a fast-follower catch-up product.
A weekly AI news roundup covering the ByteDance report alongside Google's leadership reshuffle and OpenAI's Astra cyber-capability warning, with links to primary sources.
The compute problem: export controls and chip supply
The most concrete constraint on any Chinese frontier-scale training run remains US export controls on advanced AI accelerators. Moonshot AI reportedly used around 20,000 Nvidia chips to train Kimi K3, itself a sizeable cluster assembled despite restricted access to the newest Nvidia hardware generations. A model reportedly targeting a total parameter count several times larger than Kimi K3 would, all else equal, demand a training cluster to match, whether through more chips, longer training runs, smuggled or grey-market hardware, or (more likely, per most industry analysts) domestically produced accelerators from Huawei and others that remain behind Nvidia's frontier chips on raw performance but have closed the gap faster than most 2023-era forecasts expected.
This is also the strongest practical reason for scepticism about the 10-trillion figure holding through to a shipped product. Export controls do not make a 10T training run impossible, China has repeatedly demonstrated an ability to substitute chip efficiency and architectural cleverness for raw compute access, but they do make it materially harder and slower than the equivalent run would be for a well-funded US lab with unrestricted access to the newest Nvidia or custom silicon. The three-to-six-month pre-training estimate the FT reports is consistent with a compute-constrained, carefully resourced project rather than a rushed one.
Timeline: what "early pre-training" actually means
Pre-training is the first and most compute-intensive stage of building a large language model, the phase where the base model learns general language and world patterns from raw data before any instruction-tuning, reinforcement learning, or safety alignment happens. The FT's sources describe ByteDance's model as being in this early stage, with pre-training alone expected to take three to six months. Only after that would ByteDance move into fine-tuning, evaluation and safety testing, the phases that actually determine whether a model is usable, competitive or safe to release.
That timeline puts a plausible ByteDance model launch, if the project proceeds as currently planned, no earlier than late 2026 and more realistically into 2027, allowing for the post-pre-training work every other lab has needed months to complete. Given how many frontier training runs have been delayed, resized, or quietly abandoned industry-wide over the past two years (OpenAI's own Astra, still unreleased as of this month's cyber-capability warning, is one recent example), a specific launch date should be treated as speculative until ByteDance says otherwise.
How this fits the wider US-China AI race
The ByteDance report lands in a week already crowded with China-related AI news: Moonshot's Kimi K3 open-weights release amid a US-China trade dispute, continuing scrutiny of Nvidia chip access, and a steady cadence of Chinese labs (Alibaba, Zhipu AI, DeepSeek, Moonshot) each claiming ground on Western frontier models through open weights rather than closed APIs. A 10-trillion-parameter ByteDance model, confirmed or not, reinforces a narrative that has been building since early 2025: Chinese labs are no longer content to be fast followers on efficiency alone, and are explicitly targeting frontier-scale parity, not just competitive smaller models.
Whether that narrative survives contact with an actual product remains the open question. Kimi K3's 2.8 trillion parameters made headlines but did not, on independent benchmarks, make it the strongest available open model across every category; Qwen 3.8 Max shipped similarly bold scale claims alongside genuinely competitive but not category-leading results. A ByteDance model at 10 trillion parameters would be a scale milestone regardless of benchmark outcome, but scale milestones and capability leadership have not moved in lockstep for either Chinese or Western labs over the past 18 months.
Reasons for scepticism
- Anonymous sourcing, no company confirmation. Every specific figure in the report traces back to three unnamed people; ByteDance has not confirmed or denied any of it.
- The FT itself flags unverified details. The report says it could not independently verify every element of its sources' account.
- The final size is explicitly not fixed. Sources describe 10 trillion as a ceiling ByteDance is targeting during early pre-training, not a locked specification.
- No architecture disclosed. Whether this is a dense or sparse MoE model changes the practical meaning of the headline number substantially.
- Long history of scale claims underdelivering. Several 2025-era models with large publicised parameter counts did not translate scale into a clear benchmark lead.
- Compute constraints are real. Export controls make a training run at this scale materially harder to execute on the originally reported timeline than an equivalent US lab project.
None of this means the report is wrong. It means the appropriate level of confidence, based on what is publicly known today, is "a well-sourced report about an in-progress project," not "a confirmed frontier-model launch."
Bottom line
ByteDance training a model at up to 10 trillion parameters is a credible, well-sourced report, not a confirmed product. If it holds up, it would be the largest parameter count publicly attributed to any AI lab, and a genuine signal that ByteDance intends to compete at the absolute frontier rather than remain a strong-but-secondary Chinese lab. It would not, on its own, mean the model beats Mythos, Fable, or anything else, parameter count has repeatedly failed to predict benchmark rank over the past two years, and no benchmark, architecture, or release date exists yet to evaluate.
The more useful signal to watch for is not the parameter count itself but what ByteDance discloses when (if) the model ships: architecture, training compute, and independently verifiable benchmark results. Until then, this is a training-run report worth tracking, not a leaderboard entry.
Sources: the Financial Times' 7 August 2026 report on ByteDance's training project, as syndicated via multiple technology outlets, and prior public reporting on Kimi K3's training compute and parameter count.
Last updated: 8 August 2026, the day after the Financial Times report first surfaced. This article will be revised if ByteDance comments publicly or if further reporting changes the reported parameter count or timeline.
Get the free guide: Claude vs ChatGPT, Gemini & Grok
A 20-page playbook covering everything you need to choose and use the big four AI models in 2026, full cost and feature comparisons, what each is best (and worst) at, and how-tos for images, vectors, building a website, Claude Code and more.







