Replication Slots
Watch a replication slot hold WAL open for a disconnected standby, measure exactly how much is retained, and clean it up safely
Without a replication slot, a standby that disconnects simply misses whatever WAL is generated while it is away — and if it comes back before the primary has recycled the segments it needs, it just catches up. But if the primary recycles a segment the standby still needed, the standby is permanently broken and needs a fresh base backup. A replication slot exists to prevent that: it tells the primary "do not remove any WAL past this point, no matter what, until I say otherwise" — tracked by restart_lsn in pg_replication_slots.\n\nThat guarantee is exactly what causes the disaster in the scenario above. A slot does not know or care whether its standby is actually connected — it holds WAL open regardless, for as long as the slot exists. This lab creates a slot, disconnects its standby on purpose, and measures the WAL retention grow in real time — then drops the slot and watches cleanup resume immediately.
Creating a Slot and Linking a Standby to It
SELECT pg_create_physical_replication_slot('standby1');
pg_basebackup -h 127.0.0.1 -p 5432 -U postgres -D /tmp/standby_data -R -X stream -c fast -S standby1
-S standby1 writes primary_slot_name = 'standby1' alongside the primary_conninfo that -R already generates — the standby will identify itself by this slot name on every connection, not just this first one.
Reading Slot State
SELECT slot_name, slot_type, active, restart_lsn FROM pg_replication_slots;
active reflects whether a standby is currently connected using this slot. restart_lsn is the oldest WAL position the slot is still holding open — while a standby is connected and caught up, this advances continuously; the moment it disconnects, restart_lsn freezes exactly where it was.
Measuring What Is Being Held Open
SELECT slot_name, active, restart_lsn,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal
FROM pg_replication_slots;
This number only grows while the slot's standby is disconnected and the primary keeps taking writes — exactly the "50GB and growing" scenario this lab reproduces on a much smaller, safer scale.
The Fix, and the Alternative
SELECT pg_drop_replication_slot('standby1');
Dropping the slot removes its retention guarantee immediately; the next checkpoint is then free to remove segments nothing else needs. The alternative, wal_keep_size, retains a fixed amount of WAL (in MB) regardless of whether any standby needs it — simpler, but with no way to guarantee a slow standby's exact requirement is covered, and no protection at all once that fixed amount is exceeded. A slot guarantees correctness for exactly as long as it exists; wal_keep_size guarantees a bounded disk cost but no correctness guarantee at all.
Physical Replication Slot
A named, persistent marker on the primary that guarantees WAL past a certain position is retained, specifically to protect one physical standby from ever losing WAL it has not replayed yet — even across that standby being completely offline. It survives a primary restart and must be explicitly dropped; it does not expire or clean itself up on its own.
restart_lsn
The oldest LSN a replication slot is still guaranteeing retention for. While its standby is connected and keeping up, this advances continuously alongside replay. The instant the standby disconnects, restart_lsn stops moving and stays frozen at exactly that point — which is also exactly the amount of WAL the primary can no longer safely discard.
🎰 Create the Slot
Create a physical replication slot before taking the base backup, so the standby can be linked to it from the start.
psql -U postgres -d beer_db -c "SELECT pg_create_physical_replication_slot('standby1');"student@lab:~$ psql -U postgres -d beer_db -c "SELECT pg_create_physical_replication_slot('standby1');" SET pg_create_physical_replication_slot -------------------------------------- (standby1,) (1 row)
🔎 Confirm It Starts Inactive
List replication slots and confirm this one shows inactive, with no restart_lsn yet.
psql -U postgres -d beer_db -c "SELECT slot_name, slot_type, active, restart_lsn FROM pg_replication_slots;"student@lab:~$ psql -U postgres -d beer_db -c "SELECT slot_name, slot_type, active, restart_lsn FROM pg_replication_slots;" SET slot_name | slot_type | active | restart_lsn -----------+-----------+--------+------------- standby1 | physical | f | (1 row)
📦 Link the Standby to the Slot
Take the base backup with -S so the standby identifies itself using this slot on every connection.
as-postgres pg_basebackup -h 127.0.0.1 -p 5432 -U postgres -D /tmp/standby_data -R -X stream -c fast -S standby1student@lab:~$ as-postgres pg_basebackup -h 127.0.0.1 -p 5432 -U postgres -D /tmp/standby_data -R -X stream -c fast -S standby1 2026-06-26 16:02:40.227 UTC [139] LOG: checkpoint starting: immediate force wait 2026-06-26 16:02:40.260 UTC [139] LOG: checkpoint complete: wrote 9 buffers (0.2%), wrote 3 SLRU buffers; 0 WAL file(s) added, 1 removed, 0 recycled; write=0.011 s, sync=0.003 s, total=0.033 s; sync files=6, longest=0.002 s, average=0.001 s; distance=7126 kB, estimate=7126 kB; lsn=0/8000078, redo lsn=0/8000024
🚀 Start the Standby, Watch the Slot Activate
Give the standby its port, start it, then confirm the slot is now active with a real restart_lsn.
as-postgres sh -c "echo 'port = 5433' >> /tmp/standby_data/postgresql.auto.conf"as-postgres pg_ctl start -D /tmp/standby_data -l /tmp/standby_log.txt -wpsql -U postgres -d beer_db -c "SELECT slot_name, active, restart_lsn FROM pg_replication_slots;"student@lab:~$ as-postgres sh -c "echo 'port = 5433' >> /tmp/standby_data/postgresql.auto.conf" student@lab:~$ as-postgres pg_ctl start -D /tmp/standby_data -l /tmp/standby_log.txt -w waiting for server to start.... done server started student@lab:~$ psql -U postgres -d beer_db -c "SELECT slot_name, active, restart_lsn FROM pg_replication_slots;" SET slot_name | active | restart_lsn -----------+--------+------------- standby1 | t | 0/90BDF28 (1 row)
🛑 Take the Standby Offline for "Maintenance"
Stop the standby and check the slot again — it does not disappear, and its restart_lsn does not clear.
as-postgres pg_ctl stop -D /tmp/standby_data -wpsql -U postgres -d beer_db -c "SELECT slot_name, active, restart_lsn FROM pg_replication_slots;"student@lab:~$ as-postgres pg_ctl stop -D /tmp/standby_data -w waiting for server to shut down.... done server stopped student@lab:~$ psql -U postgres -d beer_db -c "SELECT slot_name, active, restart_lsn FROM pg_replication_slots;" SET slot_name | active | restart_lsn -----------+--------+------------- standby1 | f | 0/911F3F0 (1 row)
📈 Generate Writes and Measure the Retention
Run a write workload and a WAL switch while the standby stays offline, then measure exactly how much WAL the slot is holding back from cleanup.
psql -U postgres -d beer_db -c "CREATE TABLE IF NOT EXISTS slot_demo (id serial PRIMARY KEY, payload text);" -c "INSERT INTO slot_demo (payload) SELECT repeat('a',100) FROM generate_series(1,50000) g;"psql -U postgres -d beer_db -c "SELECT pg_switch_wal();"psql -U postgres -d beer_db -c "SELECT slot_name, active, restart_lsn, pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal FROM pg_replication_slots;"student@lab:~$ psql -U postgres -d beer_db -c "CREATE TABLE IF NOT EXISTS slot_demo (id serial PRIMARY KEY, payload text);" -c "INSERT INTO slot_demo (payload) SELECT repeat('a',100) FROM generate_series(1,50000) g;" SET CREATE TABLE INSERT 0 50000 student@lab:~$ psql -U postgres -d beer_db -c "SELECT pg_switch_wal();" SET pg_switch_wal --------------- 0/9CFC9FC (1 row) student@lab:~$ psql -U postgres -d beer_db -c "SELECT slot_name, active, restart_lsn, pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal FROM pg_replication_slots;" SET slot_name | active | restart_lsn | retained_wal -----------+--------+-------------+-------------- standby1 | f | 0/911F3F0 | 15 MB (1 row)
🧹 Drop the Slot and Confirm Cleanup
Drop the slot and run a checkpoint to confirm WAL cleanup resumes immediately.
psql -U postgres -d beer_db -c "SELECT pg_drop_replication_slot('standby1');"psql -U postgres -d beer_db -c "CHECKPOINT;"student@lab:~$ psql -U postgres -d beer_db -c "SELECT pg_drop_replication_slot('standby1');" SET pg_drop_replication_slot --------------------------- (1 row) student@lab:~$ psql -U postgres -d beer_db -c "CHECKPOINT;" SET 2026-06-26 16:02:47.296 UTC [139] LOG: checkpoint starting: immediate force wait 2026-06-26 16:02:47.482 UTC [139] LOG: checkpoint complete: wrote 994 buffers (24.3%), wrote 3 SLRU buffers; 0 WAL file(s) added, 1 removed, 0 recycled; write=0.156 s, sync=0.005 s, total=0.187 s; sync files=39, longest=0.005 s, average=0.001 s; distance=11220 kB, estimate=11220 kB; lsn=0/83FF864, redo lsn=0/83FF810 CHECKPOINT (1 row)
Lab 2.5.3 complete. You have seen a replication slot cause real WAL growth on purpose, measured it exactly, and fixed it:\n\n\n Slot created and linked : ✅ pg_basebackup -S\n Active → inactive transition : ✅ restart_lsn frozen, not cleared\n Retention measured exactly : ✅ 15MB via pg_wal_lsn_diff()\n Cleanup after drop confirmed : ✅ checkpoint log line proves it\n
Enable JavaScript to run the live terminal and track your progress.