Skip to content
Rush Commerce
AI & Automation3 min read

ByteDance's 10-trillion-parameter model isn't a plan

The FT says ByteDance is pre-training a 10T-parameter model. It has no benchmarks, no release date, and no price. Here's what to actually do with that.

The number moving around this weekend is 10 trillion. That is the reported parameter count of a model ByteDance has started pre-training — the largest ever attempted in China by roughly three to one. If you route any part of your stack through a cheap Chinese model provider, the useful question is not whether ByteDance's 10-trillion-parameter model wins a benchmark in 2027. It is what a company deciding to spend that much tells you about the token prices you are budgeting against today.

What actually happened

The Financial Times reported on August 7 that ByteDance is pre-training a model of up to 10 trillion parameters, citing three people familiar with the project. The Next Web's write-up puts the comparison plainly: that is more than three times the size of Moonshot's Kimi K3, at roughly 2.8 trillion parameters currently the largest Chinese model. Moonshot reportedly used around 20,000 Nvidia chips to train K3. ByteDance's run needs substantially more.

Two details matter more than the headline number.

First, the model is in early pre-training. That phase runs three to six months, and post-training, evaluation, and safety work follow it. Neither the final scale nor the architecture is settled. There is no model card, no benchmark, no license, no price, and no date. Ten trillion is the current upper bound of what is being considered, not a specification.

Second, founder Zhang Yiming has reportedly told the roughly 2,000-person Seed team not to lean on distillation — training on a stronger model's outputs — even if that costs ByteDance ground against domestic rivals in the short term. That is the expensive path on purpose.

We have written before about Qwen3.8-Max, a 2.4-trillion-parameter preview that shipped no evaluation evidence. Same shape, bigger number. Without the active-parameter count, the training budget, the data mix, and results, a parameter total tells you about capital expenditure, not capability.

Why a competitor's training run matters for your business

Here is the operator read. Your cheap-token line item exists because a handful of Chinese labs have been shipping capable open-weight models at aggressive prices. Distillation is part of why those models were cheap to produce. A large player publicly walking away from that shortcut is a signal that the next generation costs more to make — and things that cost more to make do not usually get cheaper to rent.

That does not mean prices spike next quarter. It means you should stop treating today's per-token rate as a floor you can build a margin on. We have watched enough model aliases get retired, subscriptions get paused, and peak-hour pricing appear to be confident about one thing: the cheap tier is a business decision, and business decisions reverse.

The practical move is boring and it works. Keep a routing layer between your product and any model. Pin exact model versions instead of latest. Know your cost per completed task, not your cost per million tokens. Then a vendor's 10-trillion-parameter ambition is somebody else's capital problem instead of your roadmap.

Key takeaways

  • The FT reports ByteDance is pre-training a model of up to 10T parameters — over 3x Kimi K3's ~2.8T
  • It is early pre-training: no benchmarks, no architecture lock, no release date, no price
  • Zhang Yiming has reportedly ruled out distillation, which makes the model more expensive to build
  • Parameter counts describe spend, not capability — active params, data, and evals decide that
  • Pin model versions and measure cost per completed task, not cost per million tokens

Your model choice should be a config value, not an architecture. We build systems with the routing layer in place from day one, so swapping providers is a deploy instead of a rewrite. See how we build vendor-agnostic AI or bring us your current model bill.

Sources: The Next Web, Financial Times.

  • #bytedance
  • #ai-models
  • #model-portability
  • #token-pricing
  • #vendor-risk
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.