IMMUNO.OPS
Field Notes · Incident Analysis
On September 23, GitHub's API, organization creation, and Projects went down after database replicas detached, and the full effects took nearly nineteen hours to clear. Ten days earlier, a similar failure at the same company got a real root cause. This time, it got two sentences.
At 10:11 UTC on September 23, GitHub reported degraded performance for API Requests. Nine minutes later the scope widened: database replicas had detached, taking down organization creation and degrading the GitHub API and Projects. [1]
Core services recovered by 10:58 UTC, forty-seven minutes after the first report. But the incident wasn't actually over: Projects search results stayed stale for hours, and updates to issue labels kept lagging behind. GitHub closed the incident at 04:55 UTC the next day, eighteen hours and forty-four minutes after it began. [1]
Here's the explanation GitHub gave for the failure, in full:
"Database replicas have detached. We're working to restore the database replicas."GitHub Status, September 23, 2026
That's it. No mention of why the replicas detached, what triggered it, or what changed to stop it happening again. [1]
On September 13, GitHub had a database incident that, on paper, looks a lot like this one: replication trouble inside a shared cluster, cascading into failures across roughly 28 services. Issue creation failed for 96% of web requests at peak. New account signups failed more than 90% of the time. [2]
That time, GitHub didn't stop at describing the symptom. Its incident report explained the mechanism: a background data-cleanup job had started writing to a shared database cluster, and the safeguard meant to pace that job "watched only one health signal, how far the database replicas were lagging, and that signal stayed low the whole time. It did not account for the load building on the primary itself." A retry loop made it worse, repeatedly resending writes that were already failing and keeping the database saturated. [2]
GitHub also named what it was going to change after the September 13 incident: rate-limiting background jobs, monitoring primary load directly instead of only replica lag, adding request timeouts, bounding retries, and splitting the shared database cluster, with a two-week deadline attached. None of that appeared in the September 23 update.
GitHub hasn't published additional impact figures for the September 23 incident beyond its own status updates. What's documented is the timeline and, from ten days earlier, exactly how deep a GitHub incident report can go when it wants to.
Figures compiled from GitHub's public status page and incident reports, not additional impact disclosure from the company.
The gap here isn't technical capability. GitHub's own team proved, ten days earlier, that it can trace a database failure back to the exact safeguard that missed it. The gap is which incidents get that treatment and which get two sentences and a link to the status page.
Closing that gap isn't about writing a longer status update after the fact. It's about a chain of reasoning, done automatically, every time:
GitHub already knows how to do this. It published exactly that chain ten days before this incident. Immuno Ops is built on the assumption that every incident deserves it, not just the ones bad enough to force it: root cause and blast radius in the same view, every time.
We're not live yet. Beta opens later this year. If your team needs root cause and blast radius in the same place before the next "we're working to restore it" update, follow along.
Get notified at launch