Solicitar contato
Salesforce Still Hasn't Explained Why It Broke During Its Own Conference

IMMUNO.OPS

Field Notes · Incident Analysis

GitHub Explained a Database Failure Once. Ten Days Later, It Didn't.

On September 23, GitHub's API, organization creation, and Projects went down after database replicas detached, and the full effects took nearly nineteen hours to clear. Ten days earlier, a similar failure at the same company got a real root cause. This time, it got two sentences.

Sep 20265 min readInfrastructure · SRE · Observability

At 10:11 UTC on September 23, GitHub reported degraded performance for API Requests. Nine minutes later the scope widened: database replicas had detached, taking down organization creation and degrading the GitHub API and Projects. [1]

Core services recovered by 10:58 UTC, forty-seven minutes after the first report. But the incident wasn't actually over: Projects search results stayed stale for hours, and updates to issue labels kept lagging behind. GitHub closed the incident at 04:55 UTC the next day, eighteen hours and forty-four minutes after it began. [1]

Here's the explanation GitHub gave for the failure, in full:

"Database replicas have detached. We're working to restore the database replicas."GitHub Status, September 23, 2026

That's it. No mention of why the replicas detached, what triggered it, or what changed to stop it happening again. [1]

Ten days earlier, the same company did this differently

On September 13, GitHub had a database incident that, on paper, looks a lot like this one: replication trouble inside a shared cluster, cascading into failures across roughly 28 services. Issue creation failed for 96% of web requests at peak. New account signups failed more than 90% of the time. [2]

That time, GitHub didn't stop at describing the symptom. Its incident report explained the mechanism: a background data-cleanup job had started writing to a shared database cluster, and the safeguard meant to pace that job "watched only one health signal, how far the database replicas were lagging, and that signal stayed low the whole time. It did not account for the load building on the primary itself." A retry loop made it worse, repeatedly resending writes that were already failing and keeping the database saturated. [2]

GitHub also named what it was going to change after the September 13 incident: rate-limiting background jobs, monitoring primary load directly instead of only replica lag, adding request timeouts, bounding retries, and splitting the shared database cluster, with a two-week deadline attached. None of that appeared in the September 23 update.

One root cause,
fully explained.
Ten days later,
two sentences.

What the numbers show

GitHub hasn't published additional impact figures for the September 23 incident beyond its own status updates. What's documented is the timeline and, from ten days earlier, exactly how deep a GitHub incident report can go when it wants to.

10 daysbetween the two database-replication incidents at GitHub
18h44mfrom the September 23 incident's start to full resolution of its side effects
96%of web issue-creation requests that failed in the September 13 incident, the one that got a full explanation

Figures compiled from GitHub's public status page and incident reports, not additional impact disclosure from the company.

The gap here isn't technical capability. GitHub's own team proved, ten days earlier, that it can trace a database failure back to the exact safeguard that missed it. The gap is which incidents get that treatment and which get two sentences and a link to the status page.

What real root cause analysis requires

Closing that gap isn't about writing a longer status update after the fact. It's about a chain of reasoning, done automatically, every time:

  • Symptom → mechanism. Not "replicas detached," but which job, deploy, or connection limit caused them to detach in the first place.
  • Mechanism → blast radius. Which services and workflows actually depend on that replica set, mapped before the incident, not reconstructed after.
  • Blast radius → decision. If one database dependency can degrade org creation, the API, and Projects at once, that's a structural risk that belongs on a dashboard, not a rediscovered surprise every time replicas detach.

GitHub already knows how to do this. It published exactly that chain ten days before this incident. Immuno Ops is built on the assumption that every incident deserves it, not just the ones bad enough to force it: root cause and blast radius in the same view, every time.

We're not live yet. Beta opens later this year. If your team needs root cause and blast radius in the same place before the next "we're working to restore it" update, follow along.

Get notified at launch

Predict. Prevent. Prevail.

Kubernetes reliability intelligence.

Sources

  1. GitHub Status, "Incident across several services," githubstatus.com, Sep 23, 2026.
  2. GitHub Status, "Incident with several GitHub Services," githubstatus.com, Sep 13, 2026.