Can 160
The jump to 160k units is a staggering number. For context, most high-end clusters struggle with interconnect latency once you hit a few thousand nodes. If they are using the Ascend-950DT, they are betting heavily on Huawei's proprietary fabric to handle the weights of these massive models without the whole thing collapsing into a series of timeout errors. I'm skeptical about the "inference only" claim too. While it offloads the training pressure, the memory bandwidth requirements for serving models of DeepSeek's size across that many chips will be brutal.
The Hardware Bottleneck
The real question isn't whether the architecture works, but whether the silicon actually exists. There are persistent reports of production bottlenecks. If Huawei can't deliver the hardware for another 12+ months, this "plan" is basically just a roadmap with a lot of hope attached to it.
When you're dealing with this many units, you aren't just fighting software bugs; you're fighting hardware failure rates. In a cluster of 160,000 chips, you will have constant hardware failures. If their orchestration layer isn't bulletproof, they'll spend more time rebooting nodes and swapping failed boards than actually serving tokens.
Deployment Realities
From a deployment perspective, an AI workflow at this scale requires a level of synchronization that usually breaks. I've seen smaller clusters throw Cuda out of memory or NCCL timeout errors during simple distributed inference. With Huawei's ecosystem, you're dealing with the Ascend equivalent of these headaches. Imagine trying to debug a distributed inference hang across 160,000 chips. The telemetry alone would generate terabytes of logs per second.
If this actually ships, it would be a massive case study in non-NVIDIA scaling. But until I see a real-world benchmark or a technical whitepaper on how they're handling the interconnects, I'm treating this as a theoretical maximum rather than a guaranteed deployment.
The logistics of powering a site in Inner Mongolia for 160k high-TDP chips is also a massive undertaking. We are talking about a power draw that could rival a small city. If the power delivery flickers, you're looking at a catastrophic cascade of node failures across the cluster. I'd be curious to see the actual rack density and cooling solution they're proposing for this, because standard air cooling won't touch this kind of heat density.
