2ndOpinion.FYI中文Log in
genius.wiki

#144 2017 · GitLab · Software / DevOps infrastructure

GitLab lost 300GB of customer data and every one of its backup systems at once — then livestreamed the recovery to thousands of strangers instead of hiding it.

the problem

a company suffers a severe, embarrassing operational failure that would normally be managed through careful, delayed, lawyer-reviewed public communication

background

On January 31, 2017, a tired GitLab engineer, trying to clear a replication issue, accidentally deleted the company's primary production database directory instead of a secondary replica, wiping roughly 300GB of live data including projects, comments, issues and user accounts. As the team scrambled to restore service, they discovered a second, far worse problem: all five of GitLab's official backup mechanisms — scheduled dumps, disk snapshots, replication — had been silently failing for weeks, some due to version mismatches, without ever triggering an alert.

The standard corporate response to a disaster this severe — real data loss, compounded by a total backup failure the company hadn't even known about — is to go quiet, manage the story carefully, and issue a controlled statement once the facts and legal exposure are fully understood. GitLab's team, mid-crisis, made the opposite call: rather than retreat behind closed doors while they figured out what happened, they kept the incident visible to the public in real time, while it was still unresolved and the outcome uncertain.

what everyone would do

The standard corporate response to a disaster this severe, real data loss compounded by a total backup failure the company hadn't even known about, was to go quiet, manage the story carefully, and issue a controlled statement once the facts and legal exposure were fully understood, the instinct nearly every company follows when something goes badly wrong.

what they saw

GitLab's team saw, mid-crisis, that the instinct to control the narrative through careful, delayed communication would likely cost more trust than it protected, since silence and spin during a serious failure tend to read as evasion rather than composure. Rather than retreating behind closed doors while they figured out what happened, they made the opposite call, keeping the incident visible to the public in real time while it was still unresolved and the outcome genuinely uncertain, publishing a live, unfiltered account, mistakes and all, instead of a carefully lawyered statement after the fact.

the move

GitLab opened a live, publicly editable Google Doc tracking the incident as engineers worked it, live-tweeted updates, and streamed the entire database recovery process on YouTube — including the discovery that all backups had failed and the eventual recovery relying on a six-hour-old disk snapshot one engineer happened to have taken for an unrelated load-balancing test, not any of the systems designed for exactly this scenario. GitLab followed up with an extraordinarily detailed public postmortem naming every failure point without obscuring its own errors.

why it works

Opening a live, publicly editable document, live-tweeting updates, and streaming the actual database recovery on YouTube, including the discovery that all five backup systems had silently failed, meant the public saw GitLab's genuine, uncensored struggle rather than a polished narrative constructed after the crisis had passed, which is precisely what made the account credible rather than suspicious. Because the transparency was radical and real-time rather than curated, it functioned as a costly signal no company faking confidence could easily replicate, admitting total backup failure live, in front of over 5,000 concurrent viewers, is not something a company manages its way into unless the transparency is genuine, which is why the incident drew broad public sympathy and respect across the tech industry instead of the reputational damage an opaque, delayed response would likely have produced. The extraordinarily detailed public postmortem and dozens of follow-up engineering issues tracking concrete fixes extended that same transparency into recovery, giving customers direct visibility into exactly what was changing rather than a vague promise that lessons had been learned.

the payoff

The recovery livestream drew over 5,000 concurrent viewers at its peak, and rather than damaging GitLab's reputation, the transparent handling drew broad public sympathy and respect across the tech industry — commentary at the time and since consistently describes GitLab's reputation as effectively unharmed by an incident that, handled opaquely, would likely have been a serious trust crisis. GitLab followed the postmortem with dozens of public engineering issues tracking concrete fixes, giving customers direct visibility into exactly what was changing.

where it breaks

The mechanism depends on the company genuinely being willing to expose real failures and uncertainty in real time, not merely simulate transparency while still controlling the substance, a performative version of openness that concealed the actual severity or root cause would likely be exposed eventually and read as worse than either silence or full transparency would have. It also depends on the underlying incident actually being resolvable and the company genuinely committed to fixing it, live transparency during a crisis the company can't or won't actually resolve risks turning the real-time visibility into a prolonged, damaging spectacle rather than a redemption story. And this approach requires an organizational culture and legal environment tolerant of public, unfiltered incident disclosure, a company operating under stricter regulatory disclosure rules, facing active litigation risk, or in a culture less forgiving of visible mistakes might face real legal or reputational consequences from the same radical transparency that worked in GitLab's specific tech-industry, blameless-postmortem-friendly context.

what came after

The GitLab 2017 postmortem is now a standard reference in site-reliability engineering and incident-response training for the principle that radical, real-time transparency during a failure can outperform narrative control — cited alongside blameless postmortem culture as a template other engineering organizations, including AWS, Cloudflare and Slack, have explicitly cited when designing their own public incident communication practices.

references

  1. [1]Postmortem of database outage of January 31GitLab, 2017about.gitlab.com
  2. [2]The benefits of transparency: Interview with Sytse "Sid" Sijbrandij, CEO of GitLabIncrement, 2019increment.com

keep it

same kind of clever

Back to the archive