Industry Analysis
The 3.7x throughput figure is a red herring. The real signal: Vera Rubin compresses GPU economic life from 36 to 18 months, triggering a stranded-asset crisis across the neocloud sector. MLPerf v6.1's selection of Qwen3-VL and DeepSeek-R1 workloads is deliberate—long-context inference stresses HBM bandwidth and interconnect topology far more than peak FLOPS, shifting the competitive axis from transistor density to system-level architecture. For CoreWeave and Nebius, the question is no longer affordability; it's whether their GB300 collateral is being repriced in real time. The $3.7B convertible note is, in substance, a leveraged bet on hardware that may be obsolete within two depreciation cycles. Widening bond spreads are the market pricing that obsolescence risk. AMD's MI355X MLPerf entry is a strategic pivot: abandon the FLOPS arms race and anchor directly on cost-per-token-per-megawatt. Crusoe's 512-GPU aggregate throughput confirms the competitive unit has migrated from die to rack. Nvidia's moat was never silicon—it's NVLink fabric and the system-level lock-in it creates. Over the next 18 months, token deflation (40-60% cost reduction) will restructure LLM API pricing. CoWoS advanced packaging capacity in Taiwan, China will become the binding constraint, not chip design. The 'AI capex supercycle' narrative will fracture: hyperscalers amortizing over 36 months versus mid-tier operators forced into 18-month replacement cycles, whose financing costs will keep climbing.
This page displays AI-generated summaries and metadata for research purposes. Original content belongs to the respective publishers.