Nobody Is Saying Why OpenAI and Anthropic Had Outages Today
Original reporting by Wired

A rare confluence of technical issues temporarily disrupted leading AI frontier models from OpenAI, Anthropic, and xAI last Thursday morning, raising questions about shared infrastructure vulnerabilities. For several hours, users of OpenAI’s ChatGPT and Codex, Anthropic’s Claude, and xAI’s Grok experienced downtime or elevated errors, triggering immediate speculation that a common underlying service — such as a cloud provider or content delivery network — was to blame. This shared impact across key players in the rapidly expanding AI landscape prompted an industry-wide watch for a broader internet infrastructure problem.
The Shared Thread
However, as the dust settled, the narrative began to diverge, then reconnect in unexpected ways. OpenAI attributed its issues to a "routing error," a localized internal problem. Meanwhile, xAI, the AI venture led by Elon Musk, announced its Grok outages stemmed from an "outage at our Memphis compute center," notably adding an apology to "impacted compute partners." This detail quickly illuminated a potential link: Anthropic and xAI had publicly announced a compute partnership with xAI's parent company, SpaceX, just months prior. While Anthropic declined specific comment, their reported issues coinciding with xAI's and the acknowledgment of "compute partners" strongly suggested that a shared SpaceX compute facility was indeed the common point of failure for at least two of the affected models. The incident underscores the intricate and sometimes interconnected nature of the infrastructure powering today's most advanced AI systems.
Thursday's simultaneous outages across major frontier AI models presented a rare disruption, challenging the perception of always-on AI services. While initial speculation pointed to a shared external infrastructure provider, the detailed explanations from OpenAI, citing an internal routing error, and the admission from xAI regarding an issue at its Memphis compute center, complicated a simple narrative. Crucially, xAI's apology to "impacted compute partners" and the prior announcement of a compute partnership with Anthropic strongly suggest a direct link between at least two of the affected services. This incident underscores the intricate and often opaque web of dependencies underpinning modern AI, moving beyond individual platform issues to reveal deeper infrastructure ties.
Ensuring AI Resilience
This incident serves as a stark reminder of the foundational infrastructure challenges inherent in scaling advanced AI. The burgeoning reliance on powerful, capital-intensive compute clusters means that physical and logistical vulnerabilities, once primarily concerns for hardware firms, now directly impact the software layer and end-user experience of AI. For businesses and individuals increasingly integrating these models into critical workflows, the unexpected downtime raises immediate questions about service level agreements, disaster recovery strategies, and the true cost of centralized compute. Looking ahead, this episode will likely accelerate industry efforts to build more robust, redundant systems and diversify compute strategies, potentially fostering a more distributed and resilient AI ecosystem. As AI models transition from novel tools to critical utilities, ensuring their uninterrupted operation will not merely be a competitive advantage but a fundamental requirement for maintaining trust and fostering widespread adoption across all sectors.
Frequently asked questions
- Why did major AI chatbots like ChatGPT and Claude experience outages recently?
- Several leading AI models, including OpenAI's ChatGPT and Anthropic's Claude, experienced service disruptions last Thursday morning. OpenAI attributed its issues to a routing error, which made ChatGPT and Codex temporarily unavailable for some users. Anthropic reported a partial outage affecting Claude Mythos, Fable, and Opus models, stating they identified the cause and deployed a fix. While these incidents occurred simultaneously, both companies cited internal issues rather than a shared external cause for their respective downtimes.
- Were recent outages affecting multiple AI services connected or due to a single cause?
- While multiple major AI services, including OpenAI's ChatGPT, Anthropic's Claude, and xAI's Grok, experienced simultaneous outages last Thursday, the companies did not confirm a single shared cause. OpenAI cited a routing error, and Anthropic identified an internal issue. xAI, however, stated its Grok outage stemmed from a compute center issue and apologized to "impacted compute partners." Despite initial speculation about a common third-party provider, major internet infrastructure companies reported no related problems, suggesting independent issues or a more localized compute partnership impact.
- What caused the Grok AI chatbot to experience a service outage last Thursday?
- xAI's Grok chatbot experienced a service outage last Thursday morning across all platforms. The company, through its parent SpaceX, attributed the disruption to an outage at its Memphis compute center. xAI also issued an apology to its "impacted compute partners," suggesting broader implications related to its infrastructure. The issue was investigated and resolved within several hours, with service returning to normal traffic levels by late morning.