Why every database writes before it writes
A write-ahead log is the simplest idea in storage engineering and the one most systems get subtly wrong. Before a change touches a data page, the intent is appended to a sequential file, so a crash can be replayed instead of repaired by hand.
The contract
The rule has two halves: a page may only be flushed once its log record is durable, and a commit is only
acknowledged after the log is synced. Everything else, including fsync batching and
group commit, is an optimisation of those two sentences.
Durability is not a property of the disk. It is a property of the order in which you wait.
wal.append(record) // sequential, cheap
wal.sync() // the only expensive call
page.apply(record) // safe to do lazily
What to measure
- Sync latency at p50, p99 and p99.9, not just the mean.
- Bytes written per transaction, to spot log amplification.
- Recovery time after an unclean shutdown.

Social pill
#