Articles & Research

Analysis

Grok, Colossus and the compute arms race

xAI built a 100,000-GPU cluster in 122 days, doubled it, and merged twice. The externalities arrived with the electricity.

Dr. Anthony Reyes FACS

Chair, ACS Specialist Group on Artificial Intelligence

August 2026 · 7 min read

Aisle of server racks inside a hyperscale data centre
Aisle of server racks inside a hyperscale data centre

The most instructive thing about xAI is not its models. It is the demonstration that a training cluster can be built in months rather than years, and what gets skipped to achieve it.

Models and mergers

Grok launched as a chatbot integrated with X in November 2023. xAI unveiled Grok 3 in beta on 19 February 2025, trained on the Colossus supercluster with a claimed tenfold increase in compute over the previous generation, and released Grok 4 on 9 July 2025, with a SuperGrok Heavy subscription tier offering a higher-compute Grok 4 Heavy variant. Subsequent point releases through 2026 are documented mainly by third-party release trackers rather than primary announcements, and specific dates should be verified before citation.

The corporate structure moved faster than the models. xAI acquired X in an all-stock transaction announced on 28 March 2025, valuing xAI at $80 billion and X at more than $33 billion. On 2 February 2026, major outlets reported that SpaceX and xAI had combined ahead of a planned public offering, at a reported valuation of $1.25 trillion, with orbital data centers cited as part of the rationale. Branding for the merged AI entity remains unsettled in public reporting.

“Speed to first token can be bought with permitting risk, commissioning shortcuts and local air quality. Those costs land on someone.”

Colossus, and the cost of speed

xAI states that Colossus in Memphis was built in 122 days to 100,000 NVIDIA H100 GPUs, then doubled to 200,000 in a further 92 days. Those figures are self-reported. Subsequent reporting described cooling failures and electrical problems attributed to the compressed schedule, requiring internal rework — the predictable result of collapsing a commissioning programme that normally runs in parallel with design.

Colossus 2 estimates from independent analysts diverge sharply, ranging from roughly 530,000 GPUs at around 950MW to over 550,000 GPUs at a claimed 2GW of site capacity, with capital costs estimated between $18 billion and $36 billion. None of these are company-confirmed, and members should quote them as analyst estimates with the spread attached.

The externality

The Memphis site generated sustained community opposition over the use of methane gas turbines for on-site power, with litigation involving local and national civil rights organisations. Reporting on federal involvement in that dispute during 2026 rests on limited sourcing and the Society does not treat it as established.

What is established is the general principle, and it is one ACS members will meet repeatedly. Time-to-first-token is an engineering objective that can be bought with permitting risk, commissioning shortcuts and local air quality. Those are not free; they are costs transferred to people who did not choose the trade. A member's professional duty to have regard for the public interest is engaged at the siting meeting, not after the turbines are running.

Artificial intelligenceGreen computing

Join the professional body behind this work

ACS members receive our research first, free CPD and ethics modules every year, and a route to professional registration assessed by their peers.

Become a member