Disaster Recovery¶
What to do when something has gone badly wrong. Work top to bottom: establish scope, stop making it worse, then recover.
First: do not make it worse¶
- Stop writes if data integrity is in doubt. Disabling service accounts is instant and reversible.
- Do not run
storage repair --apply. It cannot recover missing data and it removes evidence. - Do not delete anything, including a data directory that looks broken. Move it aside instead.
- Take a copy of the current state before attempting recovery.
Establish scope¶
record-store status --endpoint https://management.example.com
record-store storage inspect --endpoint https://management.example.com
The number that decides everything: metadata_without_data. Above zero means object
bytes are gone.
Server will not start¶
Ask the machine first. doctor reports every precondition at once, without starting
anything and without printing a secret:
| Check | Cause | Fix |
|---|---|---|
configuration |
Invalid or missing settings | Correct the listed settings |
data_directory |
Not writable, or the volume is mounted read-only | Ownership — uid 10001 in the container image |
data_directory_permissions |
World-writable data directory | chmod 0700 |
atomic_publication |
Temporary directory on a different filesystem | Move it onto the data directory's filesystem |
storage_format |
On-disk format from another release | Run the matching release, or restore into an empty directory |
restore_state |
A restore never finished | Run the restore again |
free_space |
The filesystem is nearly full | Free space before starting |
s3_listener / management_listener |
Another process holds the address | Stop it, or configure a different one |
credential_master_key |
Encryption on without a key | Set it — the original key |
object_encryption |
This directory holds encrypted payloads and no key is set | Set the key the payloads were written with |
It exits 0 when nothing failed and 7 when something did, so it also works in a pre-start hook.
Two failures cannot be seen ahead of time and appear only in the logs:
| Message | Cause | Fix |
|---|---|---|
| Data directory in use | Another process holds the lock | Stop it; check for a stale container |
| Schema newer than supported | Binary downgraded past a schema change | Use the newer binary, or restore a matching backup |
Data directory is corrupt¶
# 1. Preserve it
mv /var/lib/record-store /var/lib/record-store.broken
# 2. Check the backup before trusting it
record-store server verify-backup /backups/latest --level full
# 3. Restore into a fresh directory
mkdir -p /var/lib/record-store
record-store server restore /backups/latest --level full
# 4. Start and verify
record-store storage inspect --endpoint http://127.0.0.1:7601
record-store verify bucket uploads --endpoint http://127.0.0.1:7601
Metadata, payloads, and system records all come from the one backup, so there is no way to mix a Tuesday catalog with a Thursday payload directory. See Backup and Restore.
The master key is lost¶
This is unrecoverable, and it is worth being direct about what survives.
| Recoverable | |
|---|---|
| Object payloads, encryption off | Yes — plaintext on disk |
| Object payloads, encryption on | No |
| Service-account credentials | No |
| Webhook signing secrets | No |
| Share and embed capabilities | No |
| Bucket and object metadata | Yes |
With encryption off, you can stand up a new deployment with a new master key, recreate every service account and webhook, and restore the payloads.
With encryption on, the object bytes cannot be decrypted by anything.
Back the master key up separately, today, if you have not.
Objects are missing but metadata is present¶
metadata_without_data above zero.
- Identify them —
missing_payload_samplesin the inspection output gives examples. - Restore the affected objects from backup.
- Investigate the hardware. Silently vanished payloads usually mean a failing disk, a filesystem problem, or something outside Record Store writing to the data directory.
- Only once you have finished investigating, run
storage repair --applyto clean up.
Checksum mismatches¶
A mismatch means the bytes on disk changed after they were written. That is a hardware signal before it is a Record Store problem.
- Check SMART data and the kernel log for I/O errors.
- Check memory — bad RAM corrupts data on the way to disk.
- Restore the affected objects from backup.
- Fix the hardware before restoring, or you will do this again.
After any recovery¶
-
storage inspectreports no missing payloads -
verify bucketpasses on critical buckets - Applications can read and write
- A fresh backup is taken and tested
- The root cause is understood, not just the symptom
- The runbook is updated with what actually happened
Preparing in advance¶
The work that makes recovery possible is all done beforehand:
- The master key is backed up outside the data directory
- Backups run on a schedule and restores have been tested
- The runbook says where the master key and backups live
- Monitoring alerts before a disk fills, not after
- The data directory sits on redundant storage, so a single disk failure is not an outage
- Someone other than the person who built it can do all of this