Production operations
Use this runbook alongside self-hosting and the reliability guide. A successful repository test run cannot verify your provider permissions, actual workflow limits, production load or recovery objectives. Record those results for the exact release and project you deploy.
Release and preflight
- Use release
0.2.1or later for the transaction-admission, crash-recovery and webhook-secret fixes. Earlier binaries/images do not contain these fixes. Pin the container by digest, retain the previous digest and commit your project/ABI/script versions separately from secrets. - CI gates release publication on formatting, lint, Postgres-backed tests, the full local-chain E2E suite, docs, dependency review, deployment contracts and an actual image smoke test. That smoke test exercises TLS, both embedded engines, non-root runtime writes, command execution, bearer auth, standby takeover and a complete database dump/restore. It also injects an indexer startup failure and checks that foreign projects cannot change the database. Native candidate builds run alongside repository checks; both must pass before the draft release is prepared, and image smoke/publishing must pass before publication. Inspect all release-target results too.
- Run
rflow validate --preflight --path ./projectandrflow doctor --path ./projectwith the intended profile and production-equivalent providers. Doctor's cloud-signer checks are configuration checks; boot a staging copy with those providers to verify real permissions. Use separate staging credentials, database, wallets and project name. - Exercise representative workflows, retries, webhook deduplication, RPC failure and alert delivery in staging. Set appropriate confirmation depth, relayer ceilings, durable rate limits and idempotency keys for consequential external effects. Measure resource usage and latency at expected peak load.
- Set an owner for monitoring, dead letters and key rotation. Check that alerts reach that owner and establish an on-call response time.
The installer and generated rflowup verify SHA256SUMS before extraction or
replacement. Releases without checksums fail closed; use a reviewed source build
for historical releases. Check the release's build notice and
BUILD_METADATA.json for attestation availability. GitHub does not support
artifact attestations for this user-owned private repository, so private
releases supply unsigned build metadata and checksums, with no signed
provenance claim. Checksums detect corruption; unsigned metadata records the
commit and workflow run but does not authenticate them cryptographically.
Public-repository release builds require GitHub artifact attestations. When a release explicitly reports those attestations are available, verify an archive with GitHub CLI:
gh attestation verify ./rflow_linux-amd64.tar.gz --repo joshstevens19/rflowAttestation verification binds the downloaded archive to the repository's build identity. Linux ARM users must build from source or use an available matching container image; the installer refuses to substitute an AMD64 executable.
Database ownership and upgrades
Each database belongs to exactly one YAML project name, persisted in
rflow.project_identity. Independent projects need independent databases;
HA replicas use the same name and database. Startup, CLI database operations and
MCP check this identity before changing application state. Discovery/preflight
checks read the identity without claiming a fresh database. Restart the MCP server
after changing the project name or database connection; it rejects such edits
while retaining its original connection pool.
Stop all older binaries before the first upgrade: they do not enforce this guard.
Back up the entire database first. A fresh database is claimed automatically.
An existing database is adopted automatically only when all stored config versions
identify the requested project, or when it has no rflow workflow/relayer state.
Conflicting names, malformed snapshots, and unversioned rflow state are rejected.
Changing only the YAML name cannot bypass ownership, including while no leader runs.
If an older unversioned database has state, inspect its workflows, run history and relayer mappings to establish its owner before explicitly adopting it. For a confirmed project rename, stop every leader/standby and update every replica's YAML together. After taking a backup and verifying the intended database, an operator can set the identity explicitly:
BEGIN;
SELECT pg_advisory_xact_lock(8243395393945760611);
CREATE SCHEMA IF NOT EXISTS rflow;
CREATE TABLE IF NOT EXISTS rflow.project_identity (
singleton BOOLEAN PRIMARY KEY DEFAULT TRUE CHECK (singleton),
name TEXT NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
-- Replace this literal with the verified YAML name.
INSERT INTO rflow.project_identity (name) VALUES ('verified-project-name')
ON CONFLICT (singleton) DO UPDATE SET name = EXCLUDED.name;
COMMIT;This changes only the database identity; it does not isolate mixed project data or repair overwritten workflows. Recover contaminated databases from a verified backup and reconcile external effects before restarting. Retain the identity row in backups and keep it when pruning run/config history.
Upgrading send recovery
New transaction attempts save their resolved network and relayer identity before
submission, so crash recovery uses the original destination. Older versions,
including 0.2.0, did not save this identity. Before upgrading, pause new intake
and let in-flight sends finish if their network: template depends on mutable
state, foreach item, constants, secrets, or clock functions. Recovery of these
older attempts refuses to guess their original destination, even when a
transaction ID was journaled. Literal networks and templates based only on
persisted matrix, trigger, inputs, or events data can be recovered without
the new snapshot.
Inspect pending transactions and dead letters before restarting. Fixes that prevent new nonce gaps do not repair gaps already present in an older relayer queue; reconcile those transactions with the chain before enabling new sends. Retain the full database and transaction identities during investigation.
Fatal indexer errors, task panics and unexpected executor exits stop rflow start
with a nonzero exit code, so /health cannot remain healthy indefinitely after
an engine dies. The process also exits on leadership loss during engine startup.
A bounded indexer completing successfully is expected; the executor and other
triggers continue serving. Configure a restart policy and alert on repeated exits.
Database transport and backups
Set sslmode=require for remote Postgres. All three clients (rflow, indexer and
relayer) must trust the server certificate and its hostname. Install private CAs
in the image's operating-system trust store; never turn off verification to make
a connection succeed. See database setup.
Back up the entire database, including all rflow, indexer and relayer schemas.
The history archive/retention command is a redacted export, not a recovery backup.
Keep project files and secrets/provider recovery procedures separately. Choose and
record an RPO (acceptable data loss) and RTO (acceptable downtime); use managed
PITR and encrypted, access-controlled backup storage to meet them. Alert on failed
or stale backups and periodically restore into an isolated database.
A logical backup/restore rehearsal with PostgreSQL client tools:
# Set DATABASE_URL and RESTORE_DATABASE_URL securely in the environment first.
# RESTORE_DATABASE_URL must name an empty, isolated database.
umask 077
pg_dump --dbname="$DATABASE_URL" --format=custom --file=rflow-backup.dump
pg_restore --dbname="$RESTORE_DATABASE_URL" --no-owner --exit-on-error rflow-backup.dumpUse client tools compatible with your database major version. Verify journal, cursor, configuration-version and relayer-mapping rows after restore, then boot only against a staging chain/provider setup with the same isolated data model. Never connect a recovery rehearsal to funded production signers. Record restore duration, backup timestamp, row checks and ownership/permission checks.
A stale database restore can repeat real side effects. Transactions or HTTP requests completed after the recovery point may be absent from the restored journal. Before a real recovery, stop all old leaders and standbys, preserve the old database for investigation, reconcile pending/confirmed transactions and external idempotency records, and repair the journal before resuming automation. Do not treat a successful SQL restore as proof that it is safe to send funds.
Rollout and rollback
- Save the current image digest, project revision, schema state and backup/PITR recovery point. Rehearse the exact upgrade on a staging database copy first.
- Gracefully stop the old executor or use the Helm chart's
Recreaterollout. Keep all replicas on the same database and project name. Do not start two production executors using separate restored databases. - Boot the candidate, confirm only one leader, verify relayer addresses match
existing mappings, and check
/live,/health, authenticated/api/workflows, logs, pending runs and transaction reconciliation. - Verify a controlled canary workflow and monitor error rate, lag, dead letters, pending transactions and resource use before expanding traffic.
For rollback, stop the candidate and use the previous image/project only if its schema compatibility was verified in the rehearsal. Automatic schema application is not an automatic down-migration facility. Avoid restoring an old database just to undo a binary upgrade; follow the reconciliation procedure above if recovery requires restoring data. Keep a forward-repair plan when old-code compatibility cannot be established.
Health and exposure
/live reports HTTP-server liveness and remains 200 during a workflow liveness
breach. /health returns 503 when degraded and 200 when healthy. Use /live for
restart probes, /health for readiness/alerts. A waiting HA standby has no HTTP
server until takeover, so its process probe differs from its readiness probe.
Use bearer auth, TLS at the proxy and private database access. Keep /health
reachable by monitoring but restrict public exposure at the network/proxy layer.
After rotating Secret-backed environment values, restart the deployment; the
running process does not automatically reread Kubernetes Secrets.
Dependency review
Documentation dependencies are audited with npm audit. Rust CI runs
scripts/audit-dependencies.py: new vulnerabilities, unsound/yanked crates and
expired exceptions fail the gate. Reviewed exceptions are exact-version and
engine-revision scoped in security/dependency-exceptions.json, with reachability
evidence in security/dependency-review.md. They do not mean the affected upstream
packages are patched. Re-review them when changing engine or client integrations;
weekly CI also checks for new advisories and expired reviews.