Persistent Storage¶
Everything durable lives under one directory, storage.data_directory. In the
container image that is /var/lib/record-store.
Layout¶
<data_directory>/
├── .record-store.lock exclusive lock — one process per directory
├── metadata/
│ ├── catalog.redb buckets, objects, versions, multipart state
│ ├── credentials.redb service accounts, credentials, policies
│ ├── audit.redb the audit trail
│ ├── events.redb storage events and webhook state
│ ├── lifecycle.redb lifecycle scan cursors
│ └── sharing.redb share and embed capabilities
├── objects/ object payloads
├── system/ internal storage bookkeeping
└── tmp/ incomplete payloads
Two things follow from this layout:
metadata/andobjects/are one unit. Object payloads are meaningless without the catalog that names them. Never back up or restore one without the other.tmp/is disposable but must be on the same filesystem asobjects/. Committing an upload is a rename, and a rename across filesystems is a copy. Splitting them costs a full extra write per upload.
Requirements¶
| Filesystem | POSIX with working fsync and atomic rename — ext4, XFS, ZFS, APFS |
| Mode | Read-write, owned by the running user (uid 10001 in the image) |
| Exclusivity | One process per directory, enforced by .record-store.lock |
Not on NFS, SMB, or a network filesystem
Durability rests on fsync and atomic rename behaving as POSIX specifies. Network
filesystems commonly do not, which turns a reported-durable write into a lost one
after a crash. Use block storage.
The lock file is also unreliable there, so two processes can end up writing to the same directory and corrupting it.
Sizing¶
Budget for:
| Object payloads | Your data |
| Version history | Every non-current version of every object |
| Multipart parts | Parts of uploads not yet completed or aborted |
| Metadata | Grows with object count, not object size |
| Audit trail | Grows with request volume and is never pruned |
The gap between logical and physical bytes is version history plus multipart parts — watch both:
See Capacity Planning.
Docker volumes¶
A named volume is the simplest correct choice:
A bind mount works too, but the host directory must be writable by uid 10001:
Do not mount the data directory into two containers. The lock file will refuse the second, which is the correct outcome but not one you want to discover in production.
Backups¶
A filesystem snapshot of a running deployment can catch metadata mid-write. Use the built-in backup, which takes the data lock and records a checksum per file:
That covers metadata/. Back up objects/ with your normal file backup — payloads
are immutable once committed, so an incremental copy is safe.
The command refuses to write to a directory that already exists, so each run needs a fresh destination.
Restoring requires an empty metadata/ directory, verifies every checksum, and
refuses a backup from an incompatible format or a newer schema:
Both commands take the exclusive data lock, so the server must be stopped.
See Backup and Restore.
Separating temporary storage¶
[storage]
data_directory = "/var/lib/record-store"
temporary_directory = "/var/lib/record-store/tmp"
Only worth setting to move tmp/ somewhere with different characteristics — and only
on the same filesystem as objects/, for the rename reason above. Both
record-store server doctor and start-up compare the two filesystems and refuse a
deployment that has them apart, so this can no longer be got wrong silently.
What must never be in the data directory¶
The credential master key. It is what makes the sealed contents readable; storing it alongside them means one stolen backup is a complete compromise. Keep it in a secret manager, and back it up separately.