ByteDance is pre-training an AI model with as many as 10 trillion parameters, a scale that would put the TikTok parent ahead of every other Chinese lab and larger than outside estimates for Anthropic's most advanced Mythos 5. The model is three times the size of Moonshot's Kimi K3, the biggest Chinese model released to date, and pre-training is expected to run three to six months before fine-tuning and any potential release.
Parameter counts are not the whole story — data quality and training methods matter as much — but the number sets a hard ceiling on how much information a model can store. Industry estimates peg Mythos 5 at around 8 trillion parameters and Anthropic's Fable 5 at roughly 5 trillion, meaning ByteDance's target would exceed both if the final architecture holds. The exact size will be locked in later in the training cycle.
The push comes as Chinese labs have closed most of the visible benchmark gap with US frontier models. Recent releases from Moonshot and Alibaba trail only Fable 5 in certain evaluations, while Mythos 5 remains restricted to approved organizations following a temporary ban in June tied to security concerns. Multiple Chinese labs are now training models sized to match Fable 5, but ByteDance is pushing further than any of them.
Key facts
- 01ByteDance is pre-training a model with up to 10 trillion parameters, three times the size of Moonshot's Kimi K3.
- 02Industry estimates put Anthropic's Mythos 5 at 8 trillion parameters and Fable 5 at 5 trillion.
- 03Pre-training typically runs 3 to 6 months before fine-tuning and release.
- 04ByteDance's Doubao chatbot has 324 million monthly active users in China.
- 05The Seed research team, led by former Google DeepMind scientist Wu Yonghui, has roughly 2,000 members.
ByteDance has kept a lower profile than peers because most of its models are closed rather than open-weight. Its SeeDance video generation model ranks among the strongest globally, and Doubao, its consumer chatbot, is the most-used AI assistant in China with 324 million monthly active users. Over the past three years, ByteDance has out-invested every other Chinese tech giant on AI, building data centers, hiring researchers, and pouring resources into its Volcano Engine cloud unit and custom chip ambitions.
The Seed team behind the models is now about 2,000 people spread across China and overseas, covering core research, infrastructure, data labeling, and translation. It is led by Wu Yonghui, a former scientist at Google DeepMind. For more than a year, Seed has followed an independent development approach that avoids distilling other labs' models — the practice of training a smaller student model on outputs from a larger teacher — which some inside the company believe has slowed its pace against rivals that borrow more freely.
Founder Zhang Yiming has reinforced the strategy directly. Two weeks ago, in an internal meeting first reported by Latepost and The Information, Zhang told the Seed team to target world-leading model capabilities in the long run without worrying about falling behind in the near term. ByteDance's management believes only independent development can produce a model that outperforms competitors rather than merely tracks them.
“world-leading model capabilities”— Zhang Yiming, ByteDance founder
The 10 trillion parameter target is ambitious enough that execution risk dominates the story. Training runs at that scale routinely stall on data pipeline issues, hardware failures, or loss curves that flatten before the model reaches useful capability. A model that never leaves pre-training is not a model. ByteDance has not disclosed the compute budget, the chip mix, or the training data composition, and did not respond to a request for comment. Whether the final release matches Mythos 5 or Fable 5 on real-world tasks — coding, reasoning, agentic reliability — will depend on far more than raw parameter count.
ByteDance's bet is the most concrete signal yet that Chinese labs are no longer content to catch up. Anthropic and OpenAI have spent the past two years betting that test-time compute and reasoning matter more than sheer size, and Claude Opus 5 shipped at half the price of Fable 5 in part because Anthropic pulled back from ever-larger dense models. If ByteDance can turn 10 trillion parameters into a model that actually beats Mythos 5 on the benchmarks that matter, the frontier stops being an American story. If it can't, the case for scale as the winning strategy weakens further — and the labs that chose efficiency will look prescient.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




