Skip to main content
Live
Main content

ByteDance trains 10 trillion parameter model to rival Anthropic

The TikTok parent's next model would be three times the size of Moonshot's Kimi K3 and larger than estimates for Anthropic's Mythos 5.

Jaeden Schafer
Editor in Chief · · 5 min read
ByteDance trains 10 trillion parameter model to rival Anthropic

ByteDance is pre-training an AI model with as many as 10 trillion parameters, a scale that would put the TikTok parent ahead of every other Chinese lab and larger than outside estimates for Anthropic's most advanced Mythos 5. The model is three times the size of Moonshot's Kimi K3, the biggest Chinese model released to date, and pre-training is expected to run three to six months before fine-tuning and any potential release.

Parameter counts are not the whole story — data quality and training methods matter as much — but the number sets a hard ceiling on how much information a model can store. Industry estimates peg Mythos 5 at around 8 trillion parameters and Anthropic's Fable 5 at roughly 5 trillion, meaning ByteDance's target would exceed both if the final architecture holds. The exact size will be locked in later in the training cycle.

The push comes as Chinese labs have closed most of the visible benchmark gap with US frontier models. Recent releases from Moonshot and Alibaba trail only Fable 5 in certain evaluations, while Mythos 5 remains restricted to approved organizations following a temporary ban in June tied to security concerns. Multiple Chinese labs are now training models sized to match Fable 5, but ByteDance is pushing further than any of them.

Key facts

  • 01ByteDance is pre-training a model with up to 10 trillion parameters, three times the size of Moonshot's Kimi K3.
  • 02Industry estimates put Anthropic's Mythos 5 at 8 trillion parameters and Fable 5 at 5 trillion.
  • 03Pre-training typically runs 3 to 6 months before fine-tuning and release.
  • 04ByteDance's Doubao chatbot has 324 million monthly active users in China.
  • 05The Seed research team, led by former Google DeepMind scientist Wu Yonghui, has roughly 2,000 members.

ByteDance has kept a lower profile than peers because most of its models are closed rather than open-weight. Its SeeDance video generation model ranks among the strongest globally, and Doubao, its consumer chatbot, is the most-used AI assistant in China with 324 million monthly active users. Over the past three years, ByteDance has out-invested every other Chinese tech giant on AI, building data centers, hiring researchers, and pouring resources into its Volcano Engine cloud unit and custom chip ambitions.

The Seed team behind the models is now about 2,000 people spread across China and overseas, covering core research, infrastructure, data labeling, and translation. It is led by Wu Yonghui, a former scientist at Google DeepMind. For more than a year, Seed has followed an independent development approach that avoids distilling other labs' models — the practice of training a smaller student model on outputs from a larger teacher — which some inside the company believe has slowed its pace against rivals that borrow more freely.

Founder Zhang Yiming has reinforced the strategy directly. Two weeks ago, in an internal meeting first reported by Latepost and The Information, Zhang told the Seed team to target world-leading model capabilities in the long run without worrying about falling behind in the near term. ByteDance's management believes only independent development can produce a model that outperforms competitors rather than merely tracks them.

world-leading model capabilities
Zhang Yiming, ByteDance founder

The 10 trillion parameter target is ambitious enough that execution risk dominates the story. Training runs at that scale routinely stall on data pipeline issues, hardware failures, or loss curves that flatten before the model reaches useful capability. A model that never leaves pre-training is not a model. ByteDance has not disclosed the compute budget, the chip mix, or the training data composition, and did not respond to a request for comment. Whether the final release matches Mythos 5 or Fable 5 on real-world tasks — coding, reasoning, agentic reliability — will depend on far more than raw parameter count.

ByteDance's bet is the most concrete signal yet that Chinese labs are no longer content to catch up. Anthropic and OpenAI have spent the past two years betting that test-time compute and reasoning matter more than sheer size, and Claude Opus 5 shipped at half the price of Fable 5 in part because Anthropic pulled back from ever-larger dense models. If ByteDance can turn 10 trillion parameters into a model that actually beats Mythos 5 on the benchmarks that matter, the frontier stops being an American story. If it can't, the case for scale as the winning strategy weakens further — and the labs that chose efficiency will look prescient.

Related · from this week
OpenAI cuts GPT-5.6 Luna 80% as Anthropic undercuts its own flagship
Jaeden Schafer · 5 min read →
ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

OpenAI logo
Business

OpenAI cuts GPT-5.6 Luna 80% as Anthropic undercuts its own flagship

US token prices have dropped nearly a quarter since mid-July as DoorDash and Airbnb shift workloads to Chinese models from Moonshot and DeepSeek.

Jaeden Schafer5 min read
Anthropic logo
Security

Treasury threatens sanctions after White House accuses Moonshot of distilling Anthropic's Fable

Scott Bessent says Entity List designations are on the table as Michael Kratsios alleges Moonshot accessed banned Nvidia GB300 chips via Thailand.

Jaeden Schafer5 min read
Kimi's release splits Trump's AI advisors into open-source war
Analysis

Kimi's release splits Trump's AI advisors into open-source war

Moonshot's free Kimi model rivals OpenAI and Anthropic, and Trump's current and former AI advisors are trading insults over what to do about it.

Jaeden Schafer5 min read