There was a light that never went out

Lex Neva (of SRE Weekly fame) pointed me to this recent GCP incident report, which describes networking issues they experienced in their us-central1 region. This incident is yet another reminder of why I don’t work in networking: I think that networking engineers are required to work within systems that are inherently unsafe by their very nature.

Now, I’m on the record for stating that human error is not a useful construct. However, I recognize that I’m in the minority here, and that most people believe that human error is a real and useful thing to talk about, so I’ll go with it here. I believe that for a system to be safe, it needs to be error-tolerant. An error-tolerant system is a system that does not fail hard even when a human makes an error. Achieving error-tolerance means having things like cross-checks, and the ability to recover quickly if an error does lead to a genuine problem. If you think about the way we typically roll out code to production, there are typically multiple checks involved (e.g., code review, unit tests, CI checks, tests done in a staging environment, canarying). In addition, we use techniques like progressive rollouts, blue-green deploys, and feature flags to reduce the harm if bad code makes it out to production.

I’m not a network engineer myself (I’ve never managed more than a single rack of servers, which involves a grand total of two switches), so I can’t speak from direct experience about the nature of the work. But every time I read an incident write-up that involves networking, it screams to me “this is not an error-tolerant system that network engineers have to work within”. Whenever you have to muck around with networking gear, very bad things can happen. From the GCP report, here’s the entire (sigh) “Root Cause” section (emphasis added):

The event was triggered by the inadvertent physical disconnection of network fiber-optic cables during routine hardware maintenance. A technician was performing a scheduled capacity upgrade on data center routers that support a fraction of capacity in the us-central1-b zone and a small fraction of capacity in the us-central1-f zone.

During the hands-on maintenance, optical fibers connecting the data center routers to the network fabric were unplugged and the optical transceivers were replaced with a transceiver that supports higher-density connectivity. This transceiver was not compatible with the transceivers still in place in the data center fabric, causing a loss of connectivity for each fiber.

The network is designed with redundant routers in separate data center rooms within the same building so that failure of one router does not interrupt service. The maintenance was intended to upgrade one router at a time over several days, with traffic diversions at each step and verification between steps that the network has returned to a fully-connected state.

However, a procedural error in the manually orchestrated upgrade process caused the complete list of transceiver replacements across all routers to be issued to the technician without instructions to sequence the work one router at a time. The maintenance workflow also did not include the expected human and software verification steps for detecting unintended disruption.

As a result, the technician sequentially unplugged all fiber paths across the affected devices within 13 minutes. The speed and nature of the error prevented warnings of incorrect action from reaching the engineer before connectivity was lost. Additionally, a standing procedure to halt if light is detected on any fiber-optic cable after unplug was not followed.

This resulted in compute capacity in the impacted zone being isolated from the network. Customers were unable to reach their virtual machines, and those virtual machines could not establish connections outside their zone.

The report attributes the incident to two errors:

  1. correctly following an incorrect procedure
  2. incorrectly following a correct procedure

Contrast the paucity of this analysis with the talk that Sean Klein gave at QCon SF 2025 about an Azure global networking outage that occurred back in January 2023. That also involved a networking engineer carrying out a change where the instructions contain a procedural error. However, Klein goes very deep into the history of how that came to be.

We’re also provided with no context about the nature of the work, so I have no idea how it came to be how the procedural error was introduced, or that the engineer didn’t follow the standing procedure. (Maybe they were distracted by some other aspect of the work that is part of the procedure?) All we have is the hindsight judgment that they didn’t do the right thing.

This line also jumped out at me:

 The maintenance workflow also did not include the expected human and software verification steps for detecting unintended disruption.

Our procedures are never going to be perfect. And that’s exactly why I wince every time I see “human didn’t follow the procedures” as a reason for an incident. Documented procedures are a tool, and the human consuming them has to use their judgment about how to apply the tool. If your system assumes humans perfectly following perfect procedures then, well, you’re in for some surprises.

Here are the follow-up action items (emphasis mine):

  • Reinforce training in us-central1 region and globally around safe maintenance techniques, including the mandatory requirement to verify whether light is on every unplugged fiber.
  • Finish the migration of network upgrade workflows to a fully automatically sequenced orchestration system. Most common regional network workflows were completed in 2025, and zonal workflows are in progress.
  • Complete the deployment of work stop alerting for actions resulting in the unintended disconnection of live fiber cables in Google data centers.
  • Complete the deployment of the automated system that will move regional traffic away from faulty zones, reducing the time from approximately 19 minutes in this instance to around 5 minutes.

Now, I’m confident that Google did a more extensive investigation than what they wrote up. But even so, attributing the outage to two human errors, and listing training as an action item (blame and train!) are big red flags to me that they don’t understand how work actually gets done in their own system.

I also note with interest how the the other actions involve deployment of new automation, which will likely introduce novel failure modes even as they eliminate the sorts of failure modes that occurred in this incident. I hope the Google SREs are ready for them.

In the meantime, I recommend they check out the book Behind Human Error.

Leave a comment