Nvidia's $673B forecast reveals AI compute demand still
What's genuinely new isn't the top-line figure. It's the mix shift Jensen teased on the call. Training clusters still command the headlines, but inference revenue — specifically token-generation workloads on Hopper and Blackwell — is now growing faster than training capex. Enterprise inference, not hyperscaler pre-training, became the marginal dollar driver in Q1. That flips the procurement model: you're no longer selling DGX pods to three buyers; you're selling HGX platforms to every SaaS vendor adding a copilot, every telco running RAN-inference, every automaker deploying vision-language models at the edge.
The supply constraint moved upstream. CoWoS-L packaging capacity is the new bottleneck, not wafer starts. TSMC's Arizona fab won't meaningfully relieve this until 2027 at earliest. Meanwhile Samsung's 4nm yield on HBM3E stacks still lags SK hynix by 15-20 percentage points, which means Nvidia's allocation leverage over AMD and Intel just widened. If you're planning a cluster build for H2 2025, your BOM is effectively locked to Blackwell Ultra availability — and lead times are already quoting 40 weeks.
Software moat deepened quietly. CUDA 12.6's FP8 kernel fusion for Blackwell cuts inference latency 2.3x on Llama-3-70B versus Hopper, but the real lock-in is NIM microservices. Enterprises deploying via NIM don't just get optimized kernels — they get versioned, supported containers with SLA-backed security patches. Try replicating that stack on ROCm today; the engineering cost exceeds the GPU delta for any team under 50 people.
Networking revenue growing 2.5x YoY tells its own story. Spectrum-X and NVLink Switch aren't accessories — they're the only way to scale past 32K GPUs without melting the fabric. Ethernet with RoCE v2 still drops packets at 400Gbps under all-to-all traffic patterns; InfiniBand doesn't. That's why every 100K-GPU cluster RFP now specifies NVLink domain architecture.
The bear case: China export controls just erased $12B/quarter in H20 revenue. The bull case: sovereign AI builds in Saudi Arabia, UAE, Japan, and EU are each budgeting $5-15B over three years — and they're buying full-stack, not just silicon.
My read: $673B is conservative if Blackwell Ultra yields hit 70% by Q4. If they don't, the number holds but the timeline stretches. Either way, the compute demand curve isn't bending.
All Replies (10)
Hmm, that's a real bottleneck concern. Without memory chips, even the best GPU is just an expensive paperweight. Have you looked into alternative suppliers like Micron or SK Hynix? Or maybe reconditioned chips from old hardware?