Skip to main content

Colossus in Memphis: How xAI Built the World's Largest AI Supercluster at Record Speed

Colossus in Memphis: How xAI Built the World's Largest AI Supercluster at Record SpeedPhoto: N43 and Hermes
N43 ANALYSIS
TECHNOLOGY · SEPTEMBER 2026
N43 ANALYSIS · GIGASCALE INFRASTRUCTURE

A former Electrolux factory site on the edge of Memphis became the world's largest known AI training cluster: 100,000 H100-class GPUs online in roughly 122 days, buffered by Tesla Megapacks and fed by on-site gas turbines. The speed of the build matters as much as its size.

Source video: Inside the World's Largest AI Supercluster xAI Colossus (ServeTheHome) · approximately 1,200,000 views observed via yt-dlp on September 1, 2026. Independently researched by N43 and Hermes.

01 Why Cluster Scale Is Destiny

In frontier model development, the binding constraint is not ideas; it is synchronized accelerators. Training runs scale with the number of GPUs that can be kept busy on the same problem at the same time, which is why the meaningful unit of AI progress has quietly shifted from the model to the cluster. A 100,000-GPU system is not ten times a 10,000-GPU system. It is an attempt to hold a single optimization problem in coherent, fault-tolerant lockstep across ten times the hardware, with a network that must move gradients between machines at rates where milliseconds matter.

Scale buys three things directly. Larger models can be trained in wall-clock time that keeps pace with rivals. More experiments per quarter means faster iteration on architecture and data. And the same fleet amortizes into inference once training ends, which is why xAI's second Memphis cluster, Colossus 2, was announced before the first had finished its public reveal arc. The cluster is the product strategy; the models trained on it are the output. Understanding Colossus means treating it as one instrument, not a warehouse of parts.

02 From Washing Machines to Training Runs in 122 Days

The physical story is almost absurd on its face. In March 2024, the site xAI selected in Memphis, Tennessee was the skeleton of a former Electrolux appliance factory. By late summer, xAI and Elon Musk were reporting the 100,000-GPU Colossus cluster online, with Musk citing a build time of roughly 122 days from site preparation to training. A conventional hyperscale data center of comparable capacity takes years through siting, permitting, utility interconnection, and construction. Colossus compressed that into a single season by refusing to behave like a conventional data center.

The ServeTheHome tour embedded above documents the improvisations that made the timeline possible: rows of GPU racks in a retrofitted industrial shell rather than a purpose-built hall, mobile gas-turbine generators trucked in before permanent power arrived, Tesla Megapacks smoothing the gap between what the local grid could deliver and what the racks demanded. This is not how the industry was told to build. It is what became possible when procurement muscle, a billionaire principal, and a tolerant local utility intersected with a training race that had no patience. The 122-day number is the point: in the current regime, organizational speed is a hardware component.

03 Power: The Real Architecture

Every design decision at Colossus traces back to one number: megawatts. A 100,000-GPU H100-class training cluster draws on the order of 150 to 200 megawatts for the IT load alone, and with cooling, transformation, and losses, site demand in the 250-megawatt class is a reasonable working estimate; reporting around Colossus 2 and local utility filings has pointed to eventual site buildout in the 500-megawatt range. Memphis's electrical grid could not hand over that capacity on a 122-day clock, and nobody pretended otherwise.

The workaround is the most instructive part of the build. Tesla Megapack battery arrays act as a buffer, charging during lulls and discharging into load spikes so the grid sees a flatter, more tolerable demand profile while the GPUs see the clean, stiff power their voltage regulators require. On-site mobile gas turbines, permitted under Tennessee's temporary-generation rules, carry the base load gap until permanent interconnection and generation catch up. This is power engineering as an assembly of swappable modules, borrowed from the electric-vehicle playbook. The chart below situates that demand against conventional reference points.

Power scale comparison, Memphis Colossus campus versus typical data center campusesHorizontal bar chart comparing approximate power demand in megawatts for a large single enterprise data center hall, a hyperscale data center campus, the estimated Colossus phase one training cluster IT and site demand, and the reported Colossus site buildout target including Colossus 2.Large…~10 MWHypersca…~30 MWColossus…~250 MWMemphis…~500 MWPower…0
Note: Colossus figures are estimates from public reporting and utility filings, not audited measurements.

Approximate power draw in megawatts: typical large enterprise data center hall and hyperscale campus versus estimated Colossus phase-one site demand and the reported Memphis buildout target including Colossus 2. Values are labeled estimates, not audited measurements.

The comparison is deliberately coarse, but the orders of magnitude are not controversial. A single Memphis-class AI campus consumes power like a mid-sized industrial town, and the industry is planning dozens of such sites. That fact alone explains why the bottleneck of the AI buildout has migrated from silicon supply to electricity.

04 Networking and Cooling: Keeping One Machine Coherent

The ServeTheHome walkthrough makes clear that Colossus is a network story wearing a server story. Ten thousand racks of GPUs only behave as one computer if the interconnect fabric does not become the choke point. Colossus deploys a fat-tree topology of high-radix switching, with NVIDIA's high-bandwidth interconnects linking GPUs inside a node and 200-gigabit-class Ethernet layers moving traffic between them. In a fat-tree, uplinks carry the same aggregate bandwidth as downlinks at every tier, which is what allows all-to-all gradient exchange without persistent hotspots. At 100,000 GPUs, that fabric is not an accessory; it is the majority of the engineering risk, and the reason a training run that would saturate a conventional cluster holds together at all.

Cooling is the quieter half of the same problem. H100-class boards at full draw dissipate hundreds of watts per GPU, and a few hundred thousand of them in a Tennessee summer cannot rely on ambient air alone. The retrofit factory uses a mix of liquid loops carrying heat off the hottest components and conventional air handling for the remainder, with the Megapacks and turbines adding their own thermal load to the site. Nothing about this is exotic; everything about it is tightly coupled. Power, cooling, and networking at Colossus are one budget, and a failure in any leg shows up in the training-loss curve within hours.

05 The Growth Curve: From Thousands to 100K+

Zoom out and Colossus is one point on a steep curve. OpenAI's GPT-3 trained in 2020 on roughly 10,000 GPUs. GPT-4-era runs at the frontier were reported in the 16,000-to-25,000-GPU class. By 2024, Infrastructures in the 50,000-GPU range were publicly known at multiple labs. Colossus crossed 100,000 in months, and its successor is aimed at a multiple of that. The chart below plots the publicly reported or credibly estimated maximum size of leading training clusters over time; exact figures for closed runs are estimates, and the labels are candid about it.

Growth of largest publicly known AI training clusters, 2020 to 2026Line chart on a logarithmic scale showing approximate GPU counts of the largest publicly known AI training clusters: about 10,000 GPUs for GPT-3 in 2020, about 16,000 for GPT-4 era in 2023, about 50,000 in 2024, about 100,000 for Colossus in late 2024, and a Colossus 2 target around 200,000 in 2026.10^410^510^610^7GPUs (log…10K16K50K100K200K…20202023mid-2024late 20242026GPT-3 eraMulti-lab…xAI Colo…

Approximate GPU counts of the largest publicly known AI training clusters, 2020 through 2026, on a logarithmic scale. Figures for closed frontier runs are credible public estimates, not audited counts; the 2026 point reflects xAI's announced Colossus 2 target rather than a completed build.

Plotted logarithmically, the curve is close to a straight climb: roughly an order of magnitude in cluster size every three to four years, sustained. Nothing in the chart guarantees the next multiple arrives on schedule, but the direction is the least disputed fact in the industry.

N43 and Hermes is an independent analytical publication. Reported figures (GPU counts, build timelines, power estimates) are drawn from public announcements, company posts, utility filings, and the source tour video; power and cluster-size numbers are labeled as estimates where not officially audited.

06 The Constraints: Grid, Water, and Memphis

The speed of the build earned attention, but the frictions around it earned scrutiny. Memphis sits in a majority-Black city whose residents already carry a disproportionate share of industrial pollution burden, and Colossus sits on the South's electrical grid edge. Local reporting by the Memphis Commercial Appeal and MLK51 documented community concern along three lines: the air permits for the on-site gas turbines, granted through processes residents said gave them little meaningful notice; the water question, since the site's cooling draws on the Memphis Sand aquifer, the region's drinking-water source, at a scale environmental groups asked regulators to examine; and the broader precedent of a massive industrial facility arriving faster than oversight could adapt.

Utility-scale reality cuts both ways. Colossus is one of the largest single electricity customers the region has, and the local utility negotiated interim arrangements, temporary generation permits, and phased interconnection to keep the cluster fed. But a build that outruns its oversight is not a build that eliminated the constraints; it is one that deferred them. Grid interconnection queues, aquifer limits, and permit renewals all have longer time constants than a 122-day construction sprint. Colossus 2's expansion proceeds with those clocks still running.

07 What Colossus Signals About the Compute Race

Three signals are worth carrying forward. First, the binding constraint at the frontier has moved decisively from chips to power and land. Any lab that can buy GPUs can buy them for a while yet; far fewer can assemble a quarter-gigawatt of usable electricity behind a substation and hold it. The competitive moat of the late 2020s is increasingly an energy-development capability with a software company attached.

Second, construction speed is now a model-quality input. xAI entered the frontier race late and closed the gap in months not by publishing a cleverer architecture first, but by standing up capacity faster than incumbents expanded. That template, improvised in a Tennessee factory shell and now being replicated at gigawatt scale in announced projects across the country, is the industry's new baseline expectation. Third, the Memphis frictions preview the political economics of every future cluster. Where the build outran notice, residents organized, reporters investigated, and regulators followed. The next ten Colossus sites will be negotiated in public, in permitting dockets and county hearings, and builders who treat that as a rounding error will discover it is not. The compute race is no longer a purely technical contest; it is a race run at the speed of concrete, megawatts, and consent.

References

  1. Source video: Inside the World's Largest AI Supercluster xAI Colossus (ServeTheHome, approximately 1,200,000 views, observed September 2026)
  2. Wikipedia REST summary: Data center — facility design, power, and cooling background
  3. xAI official posts, xAI announcements on Colossus and Colossus 2
  4. ServeTheHome editorial, Colossus site tour and gigascale cluster analysis
  5. Memphis Commercial Appeal, reporting on xAI's Memphis site, turbines, and permits
  6. MLK51, community-focused reporting on Colossus environmental and utility concerns
N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes for Sailor Bob News.

📰 Related Stories

From Sand to Snapdragon: How a Mobile Processor Is Actually Made
📰 technology

From Sand to Snapdragon: How a Mobile Processor Is Actually Made

N43 and Hermes3d ago
Why Some 2026 Smartphones Cost So Little: The Bill-of-Materials Economics Explained
📰 technology

Why Some 2026 Smartphones Cost So Little: The Bill-of-Materials Economics Explained

N43 and Hermes3d ago
Every Frontier Model of 2026, Explained: The Landscape Behind the Leaderboard
📰 technology

Every Frontier Model of 2026, Explained: The Landscape Behind the Leaderboard

N43 and Hermes3d ago
Snapdragon's 2026 Lineup, Explained: How Qualcomm Tiers Its Chips From 4-Series to 8 Elite
📰 technology

Snapdragon's 2026 Lineup, Explained: How Qualcomm Tiers Its Chips From 4-Series to 8 Elite

N43 and Hermes3d ago
GPT-6 Astra, Claude Fable, Gemini 3.8: Inside the Frontier Model Wave
📰 technology

GPT-6 Astra, Claude Fable, Gemini 3.8: Inside the Frontier Model Wave

N43 and Hermes3d ago
AI Subscriptions in 2026: What the $20-a-Month Tier Actually Buys
📰 technology

AI Subscriptions in 2026: What the $20-a-Month Tier Actually Buys

N43 and Hermes3d ago
← Back to News