Prepared with AI assistance. Published by CyberNative AI LLC. Corrections: hello@cybernative.ai. MIT licence; third-party excerpts and records excluded. Publisher notice.
Failure Lab · database recovery
Verify before clearing
Replication has stalled. Choose where to clean up and what to check first.
A teaching model: all effort, counters and alternative outcomes are illustrative.
01 · choose
Prepare the repair
02 · run
Follow the repair
Steps and effort units are illustrative, not historical timings.
03 · incident choice
Choose a recovery copy
04 · reveal
What the record shows
GitLab · 31 January 2017. Historical facts below carry source IDs; the result above them is illustrative.
Illustrative exposure: this path assumes the scheduled backup works without demonstrating a restore. GitLab’s postmortem says its logical backup was failing because of a PostgreSQL version mismatch and failure emails were rejected (S2). This is exposure, not extra modeled loss.
The wrong host, then lost changes. During replication repair, an engineer cleared the primary data directory instead of the secondary. GitLab reports that database changes from about 17:20 to 00:00 UTC were lost, affecting roughly 5,000 projects, 5,000 comments and 700 new user accounts. S1
Recovery was untested. The logical backup failed because of a PostgreSQL version mismatch, and failure emails were rejected. No one owned regular recovery testing. S2
The newer copy was made for staging. An engineer manually took the recent LVM snapshot about six hours before the outage to refresh staging. GitLab says these snapshots were not designed for disaster recovery; the other snapshot was nearly 24 hours old. GitLab used the newer copy to reduce data loss. S3
The restore copy took time to move. Copying the staging database to production took around 18 hours over throttled network disks, at roughly 60 Mbps. Git repositories and wikis were unavailable but not lost. S4
The lesson: a backup schedule is a claim until recovery is tested. Host verification and owned recovery checks are separate safeguards.
Source register
GitLab, Postmortem of database outage of January 31, 10 February 2017.
- S1. Timeline around 23:00–23:30 UTC; Root cause analysis, problem 1; Data loss impact. Covers the mistaken primary deletion and reported lost-write window and counts.
- S2. Broken recovery procedures; Database backups using pg_dump; Root cause analysis, problem 2. Covers the failed logical backup, rejected failure emails and missing test ownership.
- S3. Timeline around 17:20 UTC; LVM snapshots; Recovering GitLab.com. Covers the manually created staging snapshot, its intended use and the older daily snapshot.
- S4. Recovering GitLab.com; Data loss impact. Covers the roughly 18-hour copy, throttled network disks and repository/wiki availability versus data loss.