Codex is also down: OpenAI suffers simultaneous outages across all three of its core service lines, how should the downtime costs be calculated in the Agent era?
On the evening of July 25, OpenAI's API, ChatGPT, and Codex all encountered errors simultaneously, with performance degradation across 31 service components, and full service was restored 1 hour and 51 minutes later.
A single outage is not a major issue in itself, but the biggest problem is that OpenAI has not had a single fully normal day for 17 consecutive days.
A Close Call Outage
At 17:17 Beijing Time on July 25, OpenAI's official status page posted "Investigating": the error rate of multiple services was rising. At 18:02, the status was updated to "Monitoring", indicating that mitigation measures had taken effect; full service restoration was announced at 19:08.
The affected scope covers three product lines: 12 components of the API, 15 components of ChatGPT, and 4 components of Codex, totaling 31 service components with degraded performance. The incident start time recorded by third-party monitoring stations is 09:17 UTC, which is completely consistent with the status page.
Users have very intuitive experiences: failed requests, abnormal responses, and interrupted tasks. The impact on Codex is worth noting specifically — programming agents often get stuck for dozens of minutes while executing tasks, with services cutting off mid-operation. If a large project is running, there is a high risk of the entire project being left unfinished.
17 Consecutive Days of Operation With Degraded Performance
Looking at a one-month time frame, this incident cannot be regarded as an isolated case.
Records from third-party status monitoring platform Bifrost show that since July 9, OpenAI has not had a single day in a "fully normal" state: there were two Major Outages on July 12 and 16, and the remaining days alternated between Degraded Performance and Partial Outage.
Records from another monitoring station, incidenthub, are equally dense: on July 23 alone, OpenAI posted four separate incidents involving ChatGPT's error rate and latency; on the 24th, Codex Review encountered errors; and on the 25th, all three lines crashed simultaneously.
The official has remained silent on the cause, and there are two reasonable directions for deduction: the continuous rise of summer inference workload, coupled with the release pace of new models and new features, has kept the infrastructure operating at maximum capacity for a long time. These are all speculations, pending the official post-incident review.
In the Agent Era, the Calculus of Outages Has Changed
Two years ago, when ChatGPT went down, the main loss was the chat experience. The nature of this 2026 failure has shifted to a more serious level.
Behind the API are production systems: customer service robots, code pipelines, automated audits, and Agent workflows. A 111-minute service interruption means a halt to production lines. The advertising slogan below the monitoring page itself is a market signal — "OpenAI is down? Automatically route requests to healthy alternative models", showing that multi-model disaster recovery has become a business.
For enterprise solution selection, this 17-day record will bring one key indicator to the forefront: SLA. Model capability rankings change weekly, but reliability is scored on a daily basis. Capability gaps are calculated as percentages, while outage losses are calculated at 100%.
There are two reasonable deductions. First, multi-cloud and multi-model routing will evolve from a bonus feature to a standard configuration of enterprise AI architectures, and the risk exposure of relying on a single provider will be re-evaluated. Second, every major outage of an overseas flagship service is an opportunity for domestic models to absorb the overflow demand — on the premise that their own stability can withstand the same workload curve first.
OpenAI's engineering team will most likely release a post-incident review within a few days. But with the fact of 17 consecutive days of anomalies in place, the explanation the market is waiting for is probably far more complex than the simple phrase "error rate increased".