Skip to main content
September (122-day build) · 2024

Colossus Supercluster — 100k H100s Online in Record Time

The infrastructure for maximum truth-seeking AI. Rapid expansion followed: doubled to 200k in ~92 days; mixed H100/H200 + 30k+ GB200 (Blackwell) by 2025–early 2026. 2026 status: 230k–555k+ GPUs across Colossus 1/2 + Building 3. Power: ~150–250 MW initial → 500 MW → ~1–2 GW total capacity (first gigawatt-scale coherent cluster). Roadmap: 1 million GPUs. Networking: high per-server bandwidth (Tb/s), aggregate memory bandwidth (hundreds PB/s), exabyte-scale storage.

xAI brings the world's largest AI training cluster online in Memphis, Tennessee (former Electrolux factory site, South Memphis; multiple buildings/warehouses acquired). Groundbreak ~May 2024. Operational with 100,000 NVIDIA H100 GPUs (liquid-cooled) in 122 days (Sept 2024 unveiling) — 'we were told 24 months.' Initial: gas turbines + Tesla Megapack battery storage. Cooling: massive air-cooled chillers. Purpose: Primary training cluster for Grok models. Enables frontier-scale experiments and fast iteration. Philosophy of build: question everything, vertical integration where possible, extreme speed (xAI team handled much directly).

Source: x.ai/colossus

Story
Post story on X

Read it in the full timeline

Independent educational fan project. Not affiliated with Space Exploration Technologies Corp. (SpaceX) or xAI Corp.