Quick thoughts on GitHub’s Sept 23 incident

On Mastodon, Ole Peder Brandtzæg pointed me to (yet) another GitHub incident that happened back on Sept 23. The write-up is very short, so I’ll just excerpt the whole thing here (emphasis mine):

Starting at 07:57 UTC on September 23, GitHub experienced elevated 500 and 404 responses across several application pages. This caused failures when installing GitHub Apps, creating organizations, and making some organization membership changes. Customers also experienced delayed label updates and stale search results in Projects.

The infrastructure failure was isolated to our Azure Central US region. The API errors were mitigated by 10:58 UTC on September 23. Projects’ processing continued to recover while an accumulated backlog was drained, and full service was restored at 04:55 UTC on September 24.

The incident was caused by a failed planned maintenance operation on a primary database. Automated recovery initiated an emergency database failover, after which several replicas in the affected region were unable to resume replication correctly. This reduced available database capacity and caused the API errors and downstream Projects processing delays.

We have mitigated the immediate failure mode. We are also improving maintenance safety checks, database failover handling, post-failover replica validation, and downstream processing resilience to reduce the likelihood and impact of similar incidents.

There really isn’t much to go on here, but let’s make the best of it.

Unlike previous GitHub incidents, this one doesn’t sound like a “hit a tipping point due to increased load” sort of failure mode. Instead, the trigger here was planned operations work. They don’t discuss the rationale for the planned maintenance, and so it may very well be that the work was being done to increase capacity, in which case it does rhyme with the other “increased load on GitHub” incidents. Or perhaps the planned maintenance was intended to deal with some other issue. There just isn’t enough information here.

The incident was caused by a failed planned maintenance operation on a primary database.

All we really know from this write-up is that something went wrong with the maintenance operation, because it describes the operation as “failed”. From the write-up, it sounds like they were running a database cluster, where they had a primary and replicas.

A primary-replica database cluster: all database nodes can service reads, and the primary services writes

The nice thing about this sort of setup is that, if the primary fails, then a replica can automatically take over as primary, which significantly reduces downtime, which is how this database cluster was configured.

Automated recovery initiated an emergency database failover, after which several replicas in the affected region were unable to resume replication correctly.

But, again, all we get from the write-up is something went wrong with this process. It sounds like a new replica was successfully promoted to primary, but the other replicas were somehow unable to connect to the new primary.

Another advantage of a primary-replica style database cluster is that you can divide up the read request traffic across the entire cluster, which gives you more capacity. Conversely, if you lose the replicas, then the primary has to service all of the requests. This increases the load on the primary.

When the replicas fail to replicate, the primary becomes responsible for servicing all of the read traffic

This reduced available database capacity and caused the API errors and downstream Projects processing delays.

Based on the write-up, it sounds like the primary was not able to handle all of the traffic that resulted from losing the replicas. Note that even though this incident was not triggered by an increase in load, it still ended up being a saturation-related failure mode because of the loss of capacity due to the effective loss of the replica nodes.

Looking over our omnipresent availability risks, it might hit the following three:

  • problem areas > saturation > databases (primary was saturated)
  • essential non-standard changes > mitigating an operational issue (planned maintenance went wrong)
  • essential increase in essential complexity > reliability subsystem (database failover went wrong)

Finally, even with the lack of detail here, we can still see hints of how multiple factors contributed to this incident:

  • the original need for the planned maintenance
  • whatever it was that went wrong during the maintenance itself
  • whatever it was that prevented the replicas from reading from the new primary
  • the primary node alone having insufficient capacity to handle all of the resulting traffic

Leave a comment