Retention, pruning & archive — bound the journal safely
rflow journals every run and step forever by default; history is how you
debug incidents and reconstruct runs. config.retention + the
rflow retention / rflow archive commands add a safe, operator-driven
lifecycle for old history:
rflow retention plan— see exactly what would be pruned, per class, before anything is deleted.rflow retention prune— prune it, in one cascade-correct transaction per class, only with an explicit--yes.rflow archive export— dump history to redacted JSONL for audit or cold storage.
Three guarantees:
- Absent config == today. With no
config.retentionblock, nothing is ever auto-pruned. - Never delete live state. Active runs, pending approvals, parked
wait_for:waits, parked delays, in-flight sends and running reorg responses are never prunable. A run is prunable only if it is terminal, old enough, and owns no live durable dependent. - Operator-driven. There is no background pruner and no delete at
boot. Rows disappear only when a human runs
rflow retention prune --yes.
Configuration
rflow_version: 1
name: treasury
config:
db_connection: postgresql://localhost/rflow
retention:
live_runs: 180d # succeeded/skipped runs older than this are prunable
replay_sessions: 14d # disposable `rflow replay` sessions
test_runs: 30d # disposable `rflow test` sessions
failed_runs: keep # failed/dead_letter runs — default keep (audit-first)
approvals: keep # decided approval rows — default keep
archive_before_prune: true
archive_path: ./archives
notifications:
channels:
ops:
console: {}
workflows:
daily-heartbeat:
trigger:
cron: { expression: "0 9 * * *" }
steps:
- id: report
notify: { channel: ops, message: "treasury heartbeat" }Every field is optional. Each window is a duration (180d, 14d, 12h)
or the literal keep; keep is the default for every window, so a partial
block only opts the classes you name into pruning.
| Field | Governs | Default |
|---|---|---|
live_runs | succeeded + skipped (benign terminal) runs | keep |
failed_runs | failed + dead_letter runs | keep (audit-preserving) |
replay_sessions | rflow replay sessions (via the replay-prune path) | keep |
test_runs | rflow test sessions | keep |
approvals | decided approval rows (approved/rejected/expired) | keep |
archive_before_prune | export a redacted archive before deleting | false |
archive_path | where archives are written (relative = project dir) | ./archives |
failed_runs and approvals default to keep on purpose: those are the rows
you most want to review after an incident.
rflow validate checks every window is keep or a positive duration and that
archive_path is non-empty, so a typo (live_runs: whenever) fails fast.
What is protected
A run is prunable only if all of these hold:
- its status is terminal for the class (
succeeded/skippedforlive_runs;failed/dead_letterforfailed_runs), and - its
finished_atis older than the window, and - it owns no live durable dependent: no pending approval, no parked
wait_for:wait, no step parked on a tx/delay/approval/event, and no running reorg response.
Plan (counting) and prune (deleting) share one SQL predicate, so they can never drift. Terminal runs are immutable, so the set is stable under a running executor.
approvals prunes only decided rows; a pending approval is never touched.
Replay/test sessions prune through the existing
rflow replay prune path, which refuses to delete a live
workflow that collides with the session namespace.
rflow retention plan
Reads only. Reports, per class, the prunable rows/sessions and the oldest/newest timestamp in that set:
$ rflow retention plan
treasury - retention plan (as of 2026-08-04 14:30:47 UTC)
class | window | prunable | oldest | newest
live_runs | 180d | 4210 runs | 2025-06-01 09:12 | 2026-02-05 23:59
failed_runs | keep | 0 runs | - | -
replay_sessions | 14d | 3 sessions | 2026-07-01 10:00 | 2026-07-18 14:22
test_runs | 30d | 0 sessions | - | -
approvals | keep | 0 approvals | - | -rflow retention prune
Prints the same plan, then acts. Deletion requires --yes. Without it,
prune prints the plan and a hint and deletes nothing. --dry-run is an explicit
plan-only run.
rflow retention prune # prints the plan, refuses (no --yes)
rflow retention prune --dry-run # prints the plan, exits (never deletes)
rflow retention prune --yes # actually prunesEach class is deleted in one transaction, children first (step rows, reorg
journal, dead-letters, budget reservations, …) so no foreign key is ever left
dangling. When archive_before_prune: true, the archive export runs first
and must succeed. If it fails, the whole prune aborts and nothing is
deleted.
rflow archive export
Dumps runs + their steps + sends + approvals + failures + reorg responses + replay sessions in a date range to a single JSONL file, secrets redacted by default:
rflow archive export --since 2026-01-01
rflow archive export --since 2026-01-01 --until 2026-03-01 --out ./cold-storage--since/--until are YYYY-MM-DD UTC dates (the --until day is inclusive;
omit it for "up to now"). Only --format jsonl exists today (CSV for selected
tables is a later add).
Format
One file, one JSON object per line, each tagged with its source table and
ordered deterministically ((timestamp, id) within each table):
{"table":"workflow_runs","row":{"id":"…","status":"succeeded","trigger_payload":{…}}}
{"table":"step_runs","row":{"id":"…","run_id":"…","output":{…}}}sends are the send_transaction rows of step_runs (their external_id /
tx_id / tx_hash), so there is no separate sends table.
Redaction
Every string leaf of every exported row is scrubbed of secret values before
it is written, the same discipline as the
config-version snapshot: every
redaction-masked field (declared secrets:, signer credentials, channel
tokens, the server bearer, http_call headers/hmac, command.env, …) plus
every ${VAR} env value the config references feeds the scrub set. Rendered
inputs, outputs, trigger payloads and tx summaries cannot leak a secret or
an Authorization header into the archive.
Backup & compliance workflow
A common cadence, run from cron or an operator runbook (rflow itself never schedules it):
rflow archive export --since <last-run-date>→ ship the JSONL to object storage / your warehouse loader.rflow retention plan→ review what will be pruned.rflow retention prune --yes→ prune (witharchive_before_prune: truethe export re-runs and gates the delete, so step 1 is belt-and-braces).
Old replay/test sessions can be reclaimed the same way, or ad-hoc with
rflow replay prune --older-than 7d.
Non-goals (v1)
- No data-warehouse export, no immutable compliance service, no cloud-storage integration. Archives are local JSONL you move where you like.
- No background auto-delete: retention is operator-driven through the CLI.
- CSV export lands later; JSONL is the v1 format.