ByteDance has reportedly begun pre-training an AI model with up to 10 trillion parameters on roughly 30,000 GPUs. The Financial Times reported the project on August 7, 2026. ByteDance's internal Seed team is targeting Anthropic's Mythos 5, described as the current US frontier.
By total parameter count, the proposed model would be more than three times the size of Moonshot AI's Kimi K3, identified as the largest publicly deployed Chinese model at 2.8 trillion parameters. But the more consequential detail is the reported choice of domestic Chinese hardware, with no NVIDIA chips in the training cluster. Completing a run this large would require ByteDance to solve software, networking, and reliability problems that adding more chips won't fix on its own.
What 10 trillion parameters means
ByteDance is reportedly using a Mixture of Experts, or MoE, architecture. Each input token passes through only a subset of specialized sub-networks called experts. The 10 trillion figure refers to the model's total parameters, not the number activated for each token.
A model of that size might activate 200 billion to 400 billion parameters per token. That range is an illustrative estimate, not a disclosed specification for Seed's model. It would still require substantial computation, but the total parameter count doesn't imply ten times the compute cost of a 1 trillion-parameter dense model, which activates its full network on each pass.
MoE allows a model to add capacity without increasing inference computation in direct proportion. It also creates training problems. The routing system has to distribute work across experts so that some don't become overloaded while others sit idle. Keeping that distribution efficient across 30,000 chips over months of training is a substantial systems challenge.
ByteDance has relevant experience. Its MegaScale research, published in early 2024, addressed distributed training at more than 10,000 GPUs. That provides a starting point, but tripling the chip count introduces more communication and coordination work. Results at 10,000 GPUs don't establish what throughput or reliability will look like at 30,000.
The move to Huawei hardware
The April 2025 H20 ban cut off what was described as the last NVIDIA product that could legally ship to China, pushing Chinese AI labs toward domestic alternatives. ByteDance reportedly placed a $5.6 billion order for Huawei Ascend chips in 2026, described as the largest known non-NVIDIA AI hardware deal.
The relevant products are the Ascend 910C for training and the Ascend 950PR for inference. The 910C is estimated to deliver roughly 60% of an H100's compute performance per chip. Huawei claims that the 950PR delivers 2.8 times the H20's FP4 throughput, a measure of computation using four-bit floating-point numbers. That vendor claim concerns a different chip, workload, and numerical format from the training comparison.
If the training cluster consists of 30,000 Ascend 910C chips, applying the 60% estimate gives roughly 18,000 H100s' worth of nominal compute. That would be a credible frontier-scale cluster, though not a dominant one by current standards. It is also only a rough comparison. Sustained training performance depends on how well the software and network keep the chips working.
China's GPU cloud has been consolidating around Huawei and domestic suppliers. ByteDance's reported order makes it an especially aggressive buyer in that market, where domestic hardware is increasingly treated as the viable route to expansion.
The software transition may be harder than compensating for lower per-chip performance. NVIDIA's CUDA ecosystem has roughly two decades of optimization behind it. That includes custom compute kernels, numerical libraries, and NCCL, the communication library used to coordinate work across GPUs. Integrations with PyTorch and JAX have also been tested extensively at scale.
Huawei's CANN, short for Compute Architecture of Neural Networks, and MindSpore provide a capable software stack, but they don't have equivalent coverage and optimization across the full range of operations. Custom training kernels developed for ByteDance's NVIDIA clusters need to be rewritten or translated for Ascend. Assumptions about numerical precision, network layout, and the latency of operations spanning multiple chips also need fresh validation.
Zhang Yiming reportedly instructed the Seed team to pursue world-leading capabilities through independent development rather than copying rivals. That approach has a clear engineering cost alongside the hardware bill.
Keeping 30,000 accelerators running
A training run lasting three to six months cannot depend on every chip staying healthy. At 30,000 accelerators, failures become a routine operating condition. Estimates of accelerator mean time between failures measured in months per chip imply frequent failures somewhere in the cluster, potentially multiple events each day or more.
Depending on the training architecture, one failed chip can stall the entire job or leave its state unusable. Checkpointing saves enough model and training state to resume after a failure. Saving checkpoints too often consumes storage bandwidth and training time; saving them too rarely risks losing hours of work. Fast restart and fault isolation determine how much of the cluster has to stop when something breaks.
ByteDance's MegaScale work addressed checkpoint timing, online repair without full-cluster restarts, and fault isolation at pod boundaries, which separate groups of machines. Those techniques were validated on NVIDIA hardware with better-understood failure modes and interconnect behavior. Their performance on Ascend remains an open question. Huawei's high-bandwidth memory, or HBM, and chip packaging may have different failure characteristics that aren't publicly documented.
The network presents another challenge. Distributed training repeatedly exchanges information across chips. AllReduce, for example, combines gradient values across workers so they can make consistent model updates. NVIDIA's NVLink and InfiniBand have well-understood performance profiles and years of tuning for these collective operations.
Huawei's High-Speed Interconnect is less publicly documented, with no cited peer-reviewed demonstration at the proposed 30,000-accelerator scale. The physical layout matters: chips must be spread across racks, pods, and data-center segments without forcing too much communication through slow or congested paths.
A poor topology could cost an estimated 20% to 30% of effective cluster throughput. Some bottlenecks may emerge only under sustained training conditions, potentially six weeks into a run. Those are estimates of the infrastructure risk, not measured results from ByteDance's cluster.
What a completed run would establish
The project tests an assumption behind US export controls: restricting access to NVIDIA hardware can create a capability gap large enough, and durable enough, to constrain Chinese AI development.
If Seed completes a frontier-scale MoE training run on domestic hardware, that would provide evidence of a workable alternative even if the model scores below Mythos 5 on every benchmark. Other Chinese labs, countries facing export restrictions, and hardware startups would have a concrete example to study. It would also support the view that export controls buy time rather than permanently prevent advanced model development.
Completion alone wouldn't establish cost competitiveness or software maturity. The amount of engineering work, downtime, and unused compute would still matter. A successful model produced through an unusually expensive run would demonstrate technical feasibility without necessarily offering an economical path for other labs.
Pricing is a separate possible consequence. ByteDance already operates Doubao, described as China's most widely used AI assistant, with hundreds of millions of users. If the new model approaches frontier performance, the company could deploy it at Chinese consumer price points, which have historically been below US equivalents on a per-token basis.
That could put downward pressure on frontier API prices beyond China. Alibaba's Qwen series and DeepSeek offer precedents for this competitive pressure. The effect of Seed's model would depend on successful training, the resulting capabilities, and how ByteDance chooses to make it available.
The useful evidence would be a systems paper
As of August 10, 2026, ByteDance hasn't committed to a public release date. The estimated three to six months of pre-training excludes alignment, safety evaluation, and other post-training work that can add months before launch. A tentative public-release forecast is late 2026 or early 2027, assuming the run completes cleanly. Infrastructure failures or longer post-training work could push that later.
A systems paper comparable to MegaScale would reveal more about the hardware transition than benchmark scores alone. Useful details would include the interconnect configurations that worked, checkpointing overhead, recovery times, and how the team balanced expert workloads across the cluster.
Those results would help distinguish nominal chip capacity from usable training performance. They would also show which parts of ByteDance's existing infrastructure transferred to Ascend and which required substantial new work.