All News
outageazurechatgptclaudegrokinfrastructure

Same morning, three AI flagships went dark. Whose infrastructure failed?

ChatGPT, Claude and Grok all broke at once on the same Thursday morning. OpenAI says routing, xAI says Memphis — trackers point at shared Azure East US.

Vlad MakarovVlad Makarovreviewed and published
2 min read
Same morning, three AI flagships went dark. Whose infrastructure failed?

ChatGPT, Claude and Grok all broke inside the same window on Thursday morning, September 3 — three rival frontier assistants down at once for roughly two hours, and nobody agrees why. OpenAI cites a "routing error" in its own systems. xAI blames its Memphis compute center. Anthropic names an "infrastructure issue" it has not explained. Third-party outage trackers point at a shared layer of Microsoft Azure East US — an explanation none of the vendors' statements yet support.

Three outages, three explanations

The vendors' accounts:

ServiceVendor's explanationSource
ChatGPT, CodexRouting error from ~7:43 am PT, fixed by ~8:17OpenAI via The Register
Claude"Infrastructure issue"; restored 9:16 am PT (16:16 UTC)Anthropic via The Register
Grok"Outage at our Memphis compute center"xAI on X

The scale gap is the most striking number. Per the shattered.io roundup of Downdetector snapshots, ChatGPT drew more than 37,000 reports at its peak; Claude peaked at 1,324 and Grok at roughly 1,365. Gemini, on Google's own cloud, stayed largely up with about 500 reports — most coverage reads that blip as users hopping over. Most services recovered within about two hours; Anthropic's partial outage ran closer to three.

The substrate nobody advertises

The Azure theory comes from trackers, not vendors. A Computing report says StatusGator logged ingress failures in Azure's East US region during the window and that all three services "rely heavily" on it. Yet The Register found nothing on Azure's status page, Cloudflare "rather emphatically" denied any problem, and no vendor blamed Microsoft. Treat Azure East US as an analysis, not a finding.

Memphis complicates it. xAI apologized to its "impacted compute partners" — and Anthropic has rented all of Colossus 1's capacity there since May, a deal The Register floated as "one potential reason" for Claude's woes. Hacker News converged on the same read: "xAI mentioned a Memphis outage, which is where Anthropic was leasing capacity on Colossus 1." An engineer who said they were OpenAI's incident commander replied the failure "was not related to the Astra launch" — the morning OpenAI began rolling out GPT-6 Astra to Daybreak defenders.

Concentration is the story either way

Whatever the root cause, the simultaneity is the lesson: competitors rest on overlapping metal. Anthropic alone rents Azure capacity, AWS, Google TPUs and all of xAI's Memphis build — when one substrate coughs, several "competitors" cough with it. It is also the case for models you can move: a local or open-weights deployment answers when the clouds do not. Watch for Azure post-incident reports, fuller vendor postmortems, and whether enterprises stop renting the same region under different brand names.

Related Articles

Scroll down

to load the next article