Are you an LLM? Read llms.txt for a summary of the docs, or llms-full.txt for the full context.
Skip to content

Reliability & exactly-once

This page explains what rflow guarantees, how, and what it deliberately does not promise.

The guarantee is one run per retained trigger identity and idempotent relayer submission per step attempt. Completed journaled steps are not repeated during recovery. Arbitrary HTTP, notifications and local commands need their own idempotency: a crash after an external effect but before its result is recorded can cause that effect to run again. Chain finality and loss of the database journal are separate concerns.

The claim gate: one trigger occurrence, one run

Every trigger occurrence has a deterministic key: for events, chain_id:tx_hash:log_index (plus the phase for run_on: both triggers); for cron, the scheduled tick; for read/block triggers, the poll instant / block number. Starting a run is an

INSERT ... ON CONFLICT DO NOTHING

against a UNIQUE (workflow_name, trigger_key) constraint in Postgres. Whoever inserts the row owns the run; everyone else (a replayed block range, a reorged duplicate, an overlapping backfill, a second delivery from the indexer) is a no-op. The database is the arbiter: two runs can never exist for one occurrence.

For webhook triggers without idempotency_key, every accepted POST claims a fresh uuid key, so a caller's retry is a new run. Declare a dedupe identity with trigger.webhook.idempotency_key (e.g. ${{ headers['x-idempotency-key'] }} or ${{ body.event_id }}) and a duplicate POST within the TTL window creates no new run.

WebSocket stream triggers also need a stable idempotency_key to deduplicate redelivery within their TTL. A key does not guarantee that the upstream feed delivers every message: disconnected feeds or a full local queue can lose messages. Manual triggers and deliberate new runs have new identities. Pruning claim history, reusing expired webhook/stream keys or restoring an older database can allow an old occurrence to run again.

The journal: state before side effects

Every run and every step is journaled in Postgres (rflow.workflow_runs, rflow.step_runs). The invariant: state is persisted before the side effect. A step's row exists before its HTTP call fires or its transaction is queued, and its outcome (output, tx ids, error) is persisted when it settles.

So after a crash, recovery knows precisely where every run was:

  • steps that completed are never re-executed: their journaled output feeds later steps' expressions (including each foreach iteration, which journals under its own <id>[<i>] row)
  • a run parked in waiting_delay resumes its countdown against the persisted wake time
  • a run parked in waiting_tx resumes polling the same transaction
  • a run parked in waiting_event (wait_for) keeps its persisted wait rows and timeout deadline
  • a run parked in waiting_approval keeps its pending approval (the gate is a database row, not process memory)

Sends: idempotency keys, not hope

The dangerous window is a send: rflow asks the relayer to queue a transaction, and the process dies before the response lands. Did the tx get queued or not?

rflow journals a client-generated idempotency key before the request:

external_id = rflow:{run_id}:{step_id}:{attempt}

The embedded rrelayer stores it under a unique index. On recovery:

  1. look up the external_id: if the transaction exists, adopt it and keep waiting;
  2. if a successful lookup finds no transaction, mark that attempt as failed and create a fresh attempt with a new external_id;
  3. if the lookup is unavailable, fail closed rather than submit an ambiguous replacement. Normal failure handling keeps the run available for inspection.

The relayer's unique index deduplicates repeated submissions carrying the same step-attempt key. Recovery consults the original key before deciding whether a new attempt is needed. The key must exist client-side before the request because rrelayer's own tx id only comes back in the response, which the crash may have eaten.

After queueing, the stable tx_id drives status polling; gas bumping/rebroadcast (which change the tx hash, never the tx_id) are rrelayer's. Two per-send bounds cross that line: gas.hard_max_price rides the queue request and caps the bump path itself (an absolute ceiling), and valid_for's remaining window rides it as the transaction's own expiry, not just a pre-queue staleness check. With on_ceiling: expire, a send the market outruns stops bumping and expires (per valid_for when set, else the relayer's operator-wide window, 12h default), failing with the distinct gas_ceiling_exceeded reason instead of silently overpaying. hold keeps bidding at exactly the ceiling instead.

Compensation sends inside on_reorg: carry the same discipline under their own key shape, rflow-reorg:{run_id}:{fork_block}:{step_id}:{attempt}, journaled in rflow.reorg_step_runs before the relayer can see the transaction.

Approvals: the gate is a row

An approval-gated send parks the run with a persisted approval row. rflow approve/reject (or a timeout) settles that row; the executor picks the decision up and continues. A crash while parked loses nothing: the pending approval, its expiry, and the prepared transaction summary are all journaled, and an approved send still goes through recheck + re-simulation before broadcast. Not guaranteed: with on_timeout: proceed, an unanswered gate broadcasts (opt-in; rflow validate warns about it).

Sagas: wait_for guarantees

A wait_for step persists its event conditions and timeout deadline before parking, so sagas survive restarts. Every awaited (contract, event, network) gets indexed: pairs no declared trigger watches receive a wait-only subscription at boot, which routes deliveries only to the wait matcher (it can never create a run) and starts at the latest block. Resume depth is the per-condition confirmations: field: the default waits for confirmed-depth delivery (a head delivery only arms the wait), and confirmations: 0 opts into head-speed resumes with the reorg risk validate warns about. wait_for also parks per foreach item and inside finally: blocks, and on_timeout: goto crash-resume is exact: the jump decision is journaled with the settle, so a restart lands on the same named step.

Depth reduces reorg exposure; it does not prove finality. Waits armed by head deliveries check subsequent chain height without re-proving the log's inclusion. See the wait reorg limitations before using them to authorize a payment.

Recovery semantics

On every boot, before triggers start flowing:

  1. an advisory lock on (database, project) is taken: a second rflow process for the same project exits instead of double-running (or waits, if it is an HA standby)
  2. incomplete runs (queued / running / waiting_tx / waiting_delay / waiting_event / waiting_approval) are scanned
  3. waiting_tx steps reconcile against the relayer by external_id
  4. elapsed delays fire, in-flight steps resume from the journal
  5. incomplete on_reorg: responses (rflow.reorg_runs rows still running) resume through their own step journal: completed compensation steps never rerun, and interrupted compensation sends reconcile by their idempotency key

Then live triggers resume from their persisted cursors.

An HA takeover runs exactly this path: the promoted standby recovers the dead leader's in-flight work first, engines and triggers after, with the same exactly-once guarantees as any restart.

Replay and test isolation

rflow replay and rflow test run under namespaced workflow names (replay_<session>__<wf> / test_<uuid>__<wf>) with prefixed trigger keys, so a rehearsal can never collide with a live claim, a live executor never picks up rehearsal runs, and the workflow's own run history stays clean. Dry-run mode also stops every send before the relayer hand-off. Sessions are reclaimable with rflow replay prune, which only ever deletes namespaced rows, never a live workflow.

That boundary applies to the engine's workflow claims and native dry-run writes; it does not sandbox command:. Commands execute for real by default in a rehearsal. Reads, SQL queries and state/list expressions see the configured environment, not an automatic historical snapshot. See dry-run and historical-state limits.

Failed runs: the dead letter queue

A run that fails terminally (with on_failure: dead_letter, the default) keeps its full journal (every step's status, output, error, attempts and tx links) and lands in the DLQ. rflow runs list --failed shows them; rflow runs retry <id> re-fires one after you've fixed the cause. Repeated failures can trip a circuit breaker that pauses the workflow and pages someone.

Durable spend budgets survive restarts and races

Two workflow guards are Postgres-backed so they hold across a restart or an HA takeover:

  • rate_limit.durable: true persists the window in rflow.rate_limit_buckets as a fixed window claimed with one atomic INSERT … ON CONFLICT statement. A fresh process cannot reset a live window, and concurrent runs racing one bucket can never admit more than max. (The default rate_limit is in-process and resets on restart. See the caveat below.)
  • budgets: cap cumulative spend inside a rolling window. The cap is enforced by a reservation written before the send broadcasts, so racing runs cannot overshoot it.

A budget reservation follows the same money-path discipline as a send. Persist before the side effect, key by the send's idempotency external_id, let the journal drive recovery:

  1. Reserve — right before the relayer call, the send writes a durable rflow.budget_reservations row (one per asset) and atomically checks (live reservations + settled spend) + amount ≤ cap under a per-(budget, asset) lock. Over the cap: the send fails budget_exceeded, partial rows are released, and nothing broadcasts.
  2. Settle or release by outcome — a send that reaches the relayer flips its reservations to spent. A failure that proves the send never left the wallet (insufficient funds, a 429 rate-limit) releases them immediately. An ambiguous failure (rpc timeout, unknown transport error) stays reserved, because the transaction may have landed; the reconcile settles it exactly-once. A send queued but then failed on-chain (reverted / dropped / expired) moved no value, so its spent reservation is refunded to released too. A failed send never permanently consumes budget and its retry is not falsely refused.
  3. Reconcile — a crash between reserve and settle, or an unresolved ambiguous failure, leaves live reserved rows. The recovery pass reconciles them by external_id exactly like a send: provably landed becomes spent (reclaiming even a prematurely released row), never landed becomes released. Money is accounted at most once, and a queued value counts as committed spend (conservative for a cap).

Because the reservation is keyed by external_id, a re-run of the same attempt reserves the same row: a double reserve is one row, not double spend. See Budgets for the config surface and rflow budgets ls.

RPC failover: a dead provider is not a dead workflow

RPC quality dominates EVM workflow reliability. rflow's own RPC path (read: steps and read/block trigger polls, pre-flight simulation and gas estimation, wait_for confirmation/finality polls, receipt checks, the stall monitor) rides a health-aware failover stack when networks[].rpc lists more than one endpoint:

  • Traffic sticks to the primary while it is healthy. failover_after consecutive errors or a max_lag breach (a lagging-but-alive provider whose head goes stale while its 200s keep flowing) moves traffic to the first healthy fallback.
  • A background probe re-admits recovered endpoints with hysteresis: consecutive good probes, then probation until real traffic proves the endpoint, so a flapper's re-admission backs off instead of thrashing the active slot.
  • A wrong-chain endpoint is rejected outright: permanently evicted, never served, verified before its first request even when it was unreachable at boot.
  • Every request carries a 15s transport timeout, so a black-holed provider is a counted failure, never a wedged call.
  • When every endpoint is down the stack fails open to the full set rather than refusing to serve. A head unreadable from every endpoint is its own alert state (the stall monitor reports "RPC dead" instead of a silent "not stalled"); the alert opens only when endpoint health corroborates the total outage, so a one-off hiccup with a healthy fallback standing by never pages, and it closes on the first successful read.

This covers rflow's own reads and simulations; the embedded engines carry their own halves. The indexer receives the full url list and rotates its sticky active endpoint on failures/lag/wrong-chain, surfacing every switch as an event. Relayer broadcasts select health-aware among chain-verified, non-lagging endpoints, with gas estimation un-pinned from the first url. See the networks page for the exact behavior of each path.

What shutdown does — and does not — guarantee

On Ctrl-C / SIGTERM, rflow stops accepting new work, lets in-flight steps reach a journal-consistent point, and shuts the engines down (the relayer drains its queues). After a kill -9, recovery resumes claimed work from the durable journal and reconciles transaction submissions by their persisted idempotency keys. This depends on retaining the database and using the same project and relayer state; it is not a universal guarantee against duplicated external effects.

What is not guaranteed:

  • At-most-once for non-idempotent HTTP. A crash after an http_call fired but before its outcome was journaled means recovery re-executes that step. Sends are protected by idempotency keys; arbitrary HTTP cannot be. If your endpoint is not idempotent, send a stable idempotency key and have the receiver atomically deduplicate it with the effect. Include the step/item identity when a run makes multiple requests. HMAC authenticates a request; it does not deduplicate it. Same for notify: in that window: an alert may repeat, and provider failures can still prevent delivery.
  • At-most-once for commands. A local command: can repeat after an ambiguous crash or retry. Its own database writes, HTTP calls and direct transactions are outside rflow's relayer guarantees. Prefer commands that only calculate outputs; otherwise make their effects idempotent. Commands also execute by default during dry-run unless configured with dry_run: skip.
  • On-chain finality. Idempotent submission does not make a chain final. A reorg can still orphan the block your transaction landed in (see Reorgs) or the triggering event after a head-fired run already acted. confirmations: is the knob; on_reorg: is the response hook, whose response claim is deduplicated per (run, fork) and crash-resumable; its HTTP/notification steps have the same external-effect caveats as normal steps.
  • Missed cron ticks. By default, a cron tick that comes due while the process is down is skipped. Opt into catch_up: true to replay missed slots on boot, capped at the newest 100. Transient claim errors while the process is up are retried within the tick's window (the dedupe key is the scheduled time, so retrying can never double-fire).
  • Per-instance windows. rate_limit windows are in-process by default: a restart or takeover starts them fresh. Set rate_limit.durable: true for a Postgres-backed fixed window that a restart cannot reset. (Trigger throttle:, spend budgets and circuit-breaker state are Postgres-backed and survive restarts.)

Where state lives

Everything is in one Postgres (schemas rflow, relayer, rindexer). Back that up and you can rebuild a machine from scratch: the YAML is the definition, Postgres is the memory. Nothing else: no local files rflow can't regenerate (.rflow/ is generated runtime config, safe to delete).