Self-hosting
rflow is one binary plus one Postgres. There is nothing else to run.
Ports
rflow exposes one port: config.port (default 3940 in the scaffold), serving
health/status. The embedded engines are internal: rindexer's health server does not
bind, and rrelayer's API is off unless you opt in via
config.relayer_api. Expose 3940 to your
monitoring; expose nothing else.
Securing the port
By default the port is unauthenticated: fine on loopback or a private network, not for shared clusters or anything internet-adjacent. Two ways to harden it, usually combined:
Bearer auth (built in). Set
config.server to require
Authorization: Bearer <token> on /api/*. The viewer at / serves a public,
secret-free login shell; after you enter the token, it authenticates every API
request before displaying workflow data:
rflow_version: 1
name: my-project
config:
port: 3940
db_connection: ${DATABASE_URL}
server:
bind: 127.0.0.1:3940 # keep the port on loopback behind a proxy
auth:
token: "${{ secrets.RFLOW_API_TOKEN }}"
cors:
origins: ["https://ops.example.com"] # allow your dashboard's origin #
secrets:
RFLOW_API_TOKEN: ${RFLOW_API_TOKEN}/live and /health stay open for k8s probes (health_public defaults true); /metrics
follows auth unless metrics_public: true. Mint extra tokens with
rflow token create and revoke leaked ones without a
restart. /hooks/* webhook routes keep their own HMAC auth:.
Reverse proxy + TLS. rflow serves plain HTTP. Terminate TLS at a proxy
(nginx, Caddy, Traefik, or a cloud load balancer / k8s Ingress) and bind rflow
to loopback (bind: 127.0.0.1:3940) so only the proxy can reach it. The
chart's ingress values wire this up on Kubernetes.
DATABASE_URL
All state lives in Postgres, configured once via
config.db_connection, conventionally:
DATABASE_URL=postgresql://postgres:rflow@localhost:5448/postgresPoint it at any Postgres 14+ (managed is fine). rflow applies its schema on boot,
binds the database to the YAML project name and takes a per-project advisory lock (so a stray second replica exits instead of
double-running), and the embedded engines get the same connection injected. Backing
up this database preserves execution state; keep project files and secret/provider recovery procedures separately. See the production runbook.
Use a separate database for each independent project. Replicas of the same project share its database and name. A different name is rejected before engine startup or state changes. See database ownership and upgrades before upgrading an existing database or renaming a project.
For remote databases, append ?sslmode=require (or &sslmode=require if the URL
already has options). The pool and dedicated leadership connection negotiate TLS
and verify the certificate and hostname, as do both embedded engines. Install a
private database CA in the operating-system trust store used by all clients:
ARG RFLOW_IMAGE
FROM ${RFLOW_IMAGE}
USER root
COPY database-ca.crt /usr/local/share/ca-certificates/database.crt
RUN update-ca-certificates
USER 1000:1000Build with --build-arg RFLOW_IMAGE=ghcr.io/joshstevens19/rflow@sha256:... using
the reviewed image digest. Keep certificate verification enabled. An intentional
local connection to an authenticated database proxy can use the proxy's documented
transport settings; that does not secure a direct remote connection.
Docker compose
rflow new scaffolds a compose file for local Postgres. A full self-contained
local development deployment looks like (see the production runbook before exposing it):
volumes:
postgres_data:
driver: local
services:
postgresql:
image: postgres:16
shm_size: 1g
restart: always
volumes:
- postgres_data:/var/lib/postgresql/data
ports:
- 127.0.0.1:5448:5432
environment:
POSTGRES_PASSWORD: rflow
rflow:
image: ${RFLOW_IMAGE:?Set an image digest containing the current fixes}
restart: always
depends_on:
- postgresql
command: ["start", "--path", "/app/project"]
volumes:
- ./:/app/project
ports:
- 127.0.0.1:3940:3940
env_file:
- .env
environment:
DATABASE_URL: postgresql://postgres:rflow@postgresql:5432/postgresNote the DATABASE_URL override: inside the compose network Postgres is
postgresql:5432, not localhost:5448.
Helm
Use helm/rflow/values-production.yaml with an image digest containing these
fixes; the historical 0.1.0 release predates them. Create existing Secrets
rflow-env (at least DATABASE_URL) and rflow-api (rflow-api-token) through
your secret-management process. Keep credentials out of values files.
helm upgrade --install my-rflow ./helm/rflow \
-f helm/rflow/values-production.yaml \
-f project-values.yaml \
--set-file rflowConfig=./rflow.yaml \
--set image.digest="$RFLOW_IMAGE_DIGEST" \
--wait --timeout 10mYour project-values.yaml supplies projectFiles (relative ABI/script paths to
file contents) and env names for any additional Secret keys the project needs.
The chart copies the complete project into a writable volume, configures bearer
auth automatically, and binds the HTTP server to service.port.
Startup/liveness probes use /live, readiness uses /health. A fatal embedded
indexer or executor exit stops the host with exit code 1, allowing the orchestrator
to restart it. Successful bounded indexer backfills do not stop the service. Configuration
changes trigger a Recreate rollout to release the leader lock before starting
the replacement. Secret environment changes need an explicit rollout restart.
The production values enable non-root execution, a read-only root filesystem,
capability dropping and a same-namespace ingress NetworkPolicy. Adjust its
allow-list for monitoring or an ingress controller in another namespace.
Run one replica, or set ha.enabled: true for leader/standby.
Waiting standbys are not HTTP-ready; only the leader receives service traffic.
The project volume is ephemeral: durable state belongs in Postgres, and retained
command artifacts/archives need separate persistent storage. Large projects that
exceed ConfigMap limits need a custom volume/image deployment. The chart README
contains the complete project-file example and configuration details.
Deploying
Follow the provider guides for project files, database setup, secrets, health checks, upgrades, and cleanup:
Railway
Deploy on Railway with a project image and Postgres. The guide includes the stop-first update procedure needed for rflow's per-project lock.
AWS (EKS)
Deploy on AWS with EKS, the production Helm chart, and a private Postgres connection such as RDS.
GCP (GKE)
Deploy on GCP with GKE, the production Helm chart, and a TLS Postgres connection such as Cloud SQL.
High availability (leader/standby)
For failover, run extra replicas as standbys:
rflow start --standby # or in rflow.yaml:rflow_version: 1
name: my-project
config:
port: 3940
db_connection: ${DATABASE_URL}
ha:
standby: true # this instance waits instead of exiting on LockHeld
takeover_poll: 5s # how often it re-probes the lock (default 5s) How it works:
- exactly one leader executes; it holds the per-project Postgres advisory lock for the lifetime of its session
- standbys validate, connect and then park before the lock, logging
standby: waiting for leadershipwhile they poll - when the leader dies (clean shutdown,
kill -9, node loss, a network partition dropping its database session), Postgres frees the session-scoped lock and a standby wins it within one poll interval - the promoted standby runs the normal boot continuation: journal recovery first (resuming/adopting whatever the dead leader left in flight), then engines and triggers, the same exactly-once path as any restart
Fencing: the advisory lock IS the fence, and it only works because every replica uses the same project name and Postgres database. Separate databases means separate locks and two live executors: don't. The leader also watches its own lock session: if that session dies while the process is still alive (a Postgres restart, a load balancer dropping the idle connection), the leader exits immediately rather than keep executing unfenced, and the standby's takeover boot resumes everything from the journal.
rflow status shows the leadership state (leader running / no leader, via a lock
probe).
Graceful shutdown
Send SIGTERM (docker stop / kubernetes termination do this) and rflow stops accepting triggers, brings in-flight steps to a journal-consistent point, drains the relayer, and exits. A hard kill is also safe (see Reliability); recovery resumes from the journal on next boot.
Sizing
rflow is a single tokio process doing in-process function calls; CPU needs are
workload-dependent. The production Helm values start at 500m/512Mi requests
and 2 CPU/2Gi limits; these are starting values, not capacity guarantees. What actually matters:
- RPC quality — latency to your provider dominates end-to-end reaction time;
a fast HTTP provider and a tight
block_poll_frequencymatter most (ws:is parsed but not wired to the engines yet) - Postgres — every run/step is journaled; give it real storage, and back it up
- Backfill — historical
start_block: earliestruns are indexer-bound; moremax_block_range/CU budget helps