Industry Analysis
NVIDIA's rebranding of GPU clusters as "AI factories" is a strategic re-anchoring of narrative control—data centres are no longer passive compute warehouses but active production lines for intelligence. This paradigm shift is fracturing the infrastructure stack: once rack-level heat density crosses 100 kW, liquid cooling shifts from optional to existential; power planning jumps from megawatt to gigawatt scale, inverting siting logic from "near the user" to "near the grid."
On the supply side, CoWoS advanced-packaging capacity remains heavily concentrated at TSMC in Taiwan, China, while HBM yield ramp-ups at SK Hynix, Samsung, and Micron directly gate delivery timelines. Under export controls, compliance costs for downgraded SKUs are now baked into pricing architecture rather than treated as one-off friction.
Competitively, hyperscaler ASICs (TPU, Trainium, MTIA) are not "de-NVIDIA" plays—they are single-vendor risk hedges. CUDA's moat is eroding on the inference side as open-source stacks like vLLM commoditize model serving. AMD's MI300 real variable isn't silicon specs; it's whether ROCm reaches production-grade usability within 18 months.
Over the next 18 months, "training factories" and "inference factories" will bifurcate into two distinct architectural classes. Power, not silicon, becomes the binding constraint. Small modular reactor contracts tied to data centres move from whitepaper to signed EPC.
This page displays AI-generated summaries and metadata for research purposes. Original content belongs to the respective publishers.