Are you an LLM? Read llms.txt for a summary of the docs, or llms-full.txt for the full context.
Skip to content

Self-hosting

rflow is one binary plus one Postgres. There is nothing else to run.

Ports

rflow exposes one port: config.port (default 3940 in the scaffold), serving health/status. The embedded engines are internal: rindexer's health server does not bind, and rrelayer's API is off unless you opt in via config.relayer_api. Expose 3940 to your monitoring; expose nothing else.

Securing the port

By default the port is unauthenticated: fine on loopback or a private network, not for shared clusters or anything internet-adjacent. Two ways to harden it, usually combined:

Bearer auth (built in). Set config.server to require Authorization: Bearer <token> on /api/*. The viewer at / serves a public, secret-free login shell; after you enter the token, it authenticates every API request before displaying workflow data:

rflow_version: 1
name: my-project
 
config:
  port: 3940
  db_connection: ${DATABASE_URL}
  server: 
    bind: 127.0.0.1:3940                       # keep the port on loopback behind a proxy 
    auth: 
      token: "${{ secrets.RFLOW_API_TOKEN }}"
    cors: 
      origins: ["https://ops.example.com"]     # allow your dashboard's origin #
 
secrets:
  RFLOW_API_TOKEN: ${RFLOW_API_TOKEN}

/live and /health stay open for k8s probes (health_public defaults true); /metrics follows auth unless metrics_public: true. Mint extra tokens with rflow token create and revoke leaked ones without a restart. /hooks/* webhook routes keep their own HMAC auth:.

Reverse proxy + TLS. rflow serves plain HTTP. Terminate TLS at a proxy (nginx, Caddy, Traefik, or a cloud load balancer / k8s Ingress) and bind rflow to loopback (bind: 127.0.0.1:3940) so only the proxy can reach it. The chart's ingress values wire this up on Kubernetes.

DATABASE_URL

All state lives in Postgres, configured once via config.db_connection, conventionally:

DATABASE_URL=postgresql://postgres:rflow@localhost:5448/postgres

Point it at any Postgres 14+ (managed is fine). rflow applies its schema on boot, binds the database to the YAML project name and takes a per-project advisory lock (so a stray second replica exits instead of double-running), and the embedded engines get the same connection injected. Backing up this database preserves execution state; keep project files and secret/provider recovery procedures separately. See the production runbook.

Use a separate database for each independent project. Replicas of the same project share its database and name. A different name is rejected before engine startup or state changes. See database ownership and upgrades before upgrading an existing database or renaming a project.

For remote databases, append ?sslmode=require (or &sslmode=require if the URL already has options). The pool and dedicated leadership connection negotiate TLS and verify the certificate and hostname, as do both embedded engines. Install a private database CA in the operating-system trust store used by all clients:

ARG RFLOW_IMAGE
FROM ${RFLOW_IMAGE}
USER root
COPY database-ca.crt /usr/local/share/ca-certificates/database.crt
RUN update-ca-certificates
USER 1000:1000

Build with --build-arg RFLOW_IMAGE=ghcr.io/joshstevens19/rflow@sha256:... using the reviewed image digest. Keep certificate verification enabled. An intentional local connection to an authenticated database proxy can use the proxy's documented transport settings; that does not secure a direct remote connection.

Docker compose

rflow new scaffolds a compose file for local Postgres. A full self-contained local development deployment looks like (see the production runbook before exposing it):

volumes:
  postgres_data:
    driver: local
 
services:
  postgresql:
    image: postgres:16
    shm_size: 1g
    restart: always
    volumes:
      - postgres_data:/var/lib/postgresql/data
    ports:
      - 127.0.0.1:5448:5432
    environment:
      POSTGRES_PASSWORD: rflow
 
  rflow:
    image: ${RFLOW_IMAGE:?Set an image digest containing the current fixes}
    restart: always
    depends_on:
      - postgresql
    command: ["start", "--path", "/app/project"]
    volumes:
      - ./:/app/project
    ports:
      - 127.0.0.1:3940:3940
    env_file:
      - .env
    environment: 
      DATABASE_URL: postgresql://postgres:rflow@postgresql:5432/postgres

Note the DATABASE_URL override: inside the compose network Postgres is postgresql:5432, not localhost:5448.

Helm

Use helm/rflow/values-production.yaml with an image digest containing these fixes; the historical 0.1.0 release predates them. Create existing Secrets rflow-env (at least DATABASE_URL) and rflow-api (rflow-api-token) through your secret-management process. Keep credentials out of values files.

helm upgrade --install my-rflow ./helm/rflow \
  -f helm/rflow/values-production.yaml \
  -f project-values.yaml \
  --set-file rflowConfig=./rflow.yaml \
  --set image.digest="$RFLOW_IMAGE_DIGEST" \
  --wait --timeout 10m

Your project-values.yaml supplies projectFiles (relative ABI/script paths to file contents) and env names for any additional Secret keys the project needs. The chart copies the complete project into a writable volume, configures bearer auth automatically, and binds the HTTP server to service.port.

Startup/liveness probes use /live, readiness uses /health. A fatal embedded indexer or executor exit stops the host with exit code 1, allowing the orchestrator to restart it. Successful bounded indexer backfills do not stop the service. Configuration changes trigger a Recreate rollout to release the leader lock before starting the replacement. Secret environment changes need an explicit rollout restart. The production values enable non-root execution, a read-only root filesystem, capability dropping and a same-namespace ingress NetworkPolicy. Adjust its allow-list for monitoring or an ingress controller in another namespace.

Run one replica, or set ha.enabled: true for leader/standby. Waiting standbys are not HTTP-ready; only the leader receives service traffic. The project volume is ephemeral: durable state belongs in Postgres, and retained command artifacts/archives need separate persistent storage. Large projects that exceed ConfigMap limits need a custom volume/image deployment. The chart README contains the complete project-file example and configuration details.

Deploying

Follow the provider guides for project files, database setup, secrets, health checks, upgrades, and cleanup:

Railway

Deploy on Railway with a project image and Postgres. The guide includes the stop-first update procedure needed for rflow's per-project lock.

AWS (EKS)

Deploy on AWS with EKS, the production Helm chart, and a private Postgres connection such as RDS.

GCP (GKE)

Deploy on GCP with GKE, the production Helm chart, and a TLS Postgres connection such as Cloud SQL.

High availability (leader/standby)

For failover, run extra replicas as standbys:

rflow start --standby       # or in rflow.yaml:
rflow_version: 1
name: my-project
 
config:
  port: 3940
  db_connection: ${DATABASE_URL}
  ha: 
    standby: true           # this instance waits instead of exiting on LockHeld 
    takeover_poll: 5s       # how often it re-probes the lock (default 5s) 

How it works:

  • exactly one leader executes; it holds the per-project Postgres advisory lock for the lifetime of its session
  • standbys validate, connect and then park before the lock, logging standby: waiting for leadership while they poll
  • when the leader dies (clean shutdown, kill -9, node loss, a network partition dropping its database session), Postgres frees the session-scoped lock and a standby wins it within one poll interval
  • the promoted standby runs the normal boot continuation: journal recovery first (resuming/adopting whatever the dead leader left in flight), then engines and triggers, the same exactly-once path as any restart

Fencing: the advisory lock IS the fence, and it only works because every replica uses the same project name and Postgres database. Separate databases means separate locks and two live executors: don't. The leader also watches its own lock session: if that session dies while the process is still alive (a Postgres restart, a load balancer dropping the idle connection), the leader exits immediately rather than keep executing unfenced, and the standby's takeover boot resumes everything from the journal.

rflow status shows the leadership state (leader running / no leader, via a lock probe).

Graceful shutdown

Send SIGTERM (docker stop / kubernetes termination do this) and rflow stops accepting triggers, brings in-flight steps to a journal-consistent point, drains the relayer, and exits. A hard kill is also safe (see Reliability); recovery resumes from the journal on next boot.

Sizing

rflow is a single tokio process doing in-process function calls; CPU needs are workload-dependent. The production Helm values start at 500m/512Mi requests and 2 CPU/2Gi limits; these are starting values, not capacity guarantees. What actually matters:

  • RPC quality — latency to your provider dominates end-to-end reaction time; a fast HTTP provider and a tight block_poll_frequency matter most (ws: is parsed but not wired to the engines yet)
  • Postgres — every run/step is journaled; give it real storage, and back it up
  • Backfill — historical start_block: earliest runs are indexer-bound; more max_block_range/CU budget helps