Fouria Operator Guide
On this page
1. Welcome
Fouria is a Cloudflare Worker that runs OpenClaw (the AI agent gateway) inside a Cloudflare Sandbox container. Each customer gets an isolated container; the Worker starts it, proxies chat traffic to it, snapshots its state to R2, and keeps it healthy.
Most of the time the system runs itself. This guide is for the moments when it does not: when a container is stuck, a backup will not restore, the AI gateway rejects requests, or the gateway will not start.
The three surfaces you operate
| Surface | Entry point | When to use |
|---|---|---|
| Health snapshot | GET /api/health (and ?passive=1) | Is the gateway up? What is its status? |
| Debug routes | /debug/* (only when DEBUG_ROUTES=true) | Deep inspection + emergency actions (stop, destroy, rescue, restart) |
| Diagnostics module | src/diagnostics.ts | The code that powers container rescue; gateway restart lives in src/gateway/process.ts + src/routes/debug.ts |
Mental model: the lifecycle
Request / cron / health probe
│
▼
ensureGateway() ── finds or starts the container
│
▼
start-openclaw.sh (onboard → moltlazy patch → plugins → gateway run)
│
▼
OpenClaw gateway (port 18789, Control UI + WebSocket)
│
▼
createSnapshot() → R2 (state preserved between restarts)
When any link in this chain breaks, the tools in this guide are your recovery path.
2. Getting Started
2.1 Prerequisites
You should be comfortable with:
- Reading a JSON health response (
curl … | jq). - Cloudflare concepts: Workers, Containers, R2, AI Gateway.
- That the gateway runs inside a container — most “server is down” issues are really “container is down or the gateway process inside it is down”.
2.2 Enabling diagnostics — how to turn on debug options
Debug tooling is off by default in production for safety. There are two layers:
Layer A — DEBUG_ROUTES (the /debug/* HTTP surface)
The /debug/* routes are gated by the DEBUG_ROUTES environment variable
(apps/fouria/src/index.ts:379, src/types.ts:34).
# In .dev.vars (local) or as a Worker var (production)
DEBUG_ROUTES=true
DEBUG_ROUTES !== 'true'→ every/debug/*call returns 404 immediately (index.ts:292).- Setting it is non-destructive: it only exposes the inspection/action endpoints; it does not change runtime behavior.
⚠️ Never leave
DEBUG_ROUTES=trueon a tenant-facing deployment. It exposes/debug/gateway-token,/debug/stop-gateway, and/debug/destroy-container. Use it on operator/ROOT deployments only.
Layer B — passive health probes (no extra config)
The public health endpoint GET /api/health is always on. Append ?passive=1 to inspect
a down instance without warming/cold-starting it (routes/public.ts:163):
curl "https://<worker>/api/health?passive=1" | jq .
Passive probes use a short 3-second bound (PASSIVE_PROBE_TIMEOUT_MS, config.ts:24) and
never trigger ensureGateway, so a dead container fails fast instead of spinning up for
minutes.
2.3 Your first diagnostic call
# 1. Is the gateway healthy? (fast, no warm-up)
curl "https://<worker>/api/health?passive=1" | jq '.status, .degraded, .containerError'
# 2. Full active snapshot (warms the container, schedules a backup)
curl "https://<worker>/api/health" | jq .
# 3. Deep dive (requires DEBUG_ROUTES=true)
curl "https://<worker>/debug/version" | jq .
curl "https://<worker>/debug/processes?logs=true" | jq .
Read the status field — it is one of running | starting | stopped
(health-probes.ts:67). Read degraded — when true, a permanent container error
(e.g. deleted image) requires a rescue (see §6.1).
3. Operation Guidelines — Debug Options & How to Activate Them
3.1 The debug endpoint catalog
All under /debug/*, active only when DEBUG_ROUTES=true. Each maps to logic in
src/diagnostics.ts or gateway/process.ts.
| Method & Path | What it does | Backed by |
|---|---|---|
GET /debug/version | OpenClaw + Node versions in the container | sandbox.startProcess |
GET /debug/processes?logs=true | List container processes (+ logs); returns 503 { status: "platform_updating" } while the sandbox runtime is being updated (deploy) | sandbox.listProcesses |
GET /debug/gateway-api | Probe the OpenClaw HTTP API | GATEWAY_API_TIMEOUT_MS |
GET /debug/cli | Run an OpenClaw CLI command | sandbox.exec |
GET /debug/logs | Container stdout logs | sandbox |
GET /debug/env | Sanitized environment (secrets masked) | maskEnvValues() |
GET /debug/container-config | OpenClaw config from inside container | sandbox.readFile |
GET /debug/gateway-logs | Gateway stdout (/tmp/gateway-stdout.log) | sandbox.exec |
GET /debug/gateway-token | The gateway auth token | resolveSecret |
GET /debug/gateway-lock | OpenClaw native lock state | gateway/lock.ts |
POST /debug/stop-gateway | Kill gateway without restarting | killGateway() |
POST /debug/restart-openclaw | Hard gateway restart (kill → clear startup cache → flag R2 restore → start-openclaw.sh reprovision) | killGateway() + resetStartupCache() + signalRestoreNeeded() |
POST /debug/reprovision | Force a container reprovision (destroy → clear startup/persistence caches → reset orphan streak); never touches the gateway | sandbox.destroy() + clearPersistenceCache() + resetStartupCache() |
POST /debug/destroy-container | Destroy the container entirely | sandbox.destroy() |
POST /debug/rescue-container | Diagnose + rescue a stuck container | rescueContainer() |
POST /debug/r2-put?key= | Write raw content to R2 (tests) | BACKUP_BUCKET.put |
Secret safety:
GET /debug/envruns every value throughmaskEnvValues()(diagnostics.ts:31). Any key matchingtoken,api_key,password,secret,credential,auth, orprivate_keyis replaced with••••••••. Never paste raw/debug/envoutput into tickets.
3.2 Reading the gateway lock (the most common “ghost” outage)
OpenClaw enforces single-instance ownership via three lock layers — a state lock, a
config lock, and an exclusive TCP bind on port 18789 (gateway/lock.ts:1,
docs.openclaw.ai/gateway/gateway-lock). A stale lock is the #1 cause of
"gateway already running (pid …)" / EADDRINUSE on restart.
curl "https://<worker>/debug/gateway-lock" | jq '.state'
state is one of (lock.ts:27):
| State | Meaning | Operator action |
|---|---|---|
free | No lock held — safe to start | None |
held-live | A live gateway owns the port | Adopt it; do not kill |
stale-reclaimable | Lock files present but owner PID gone | Auto-reclaimed on next start; restart is safe |
conflict | Lock error seen in the log | Wait/restart; avoid parallel starts |
unknown | Probe failed (container off) | Check container health first |
The hard restart path checkpoints state before killing the gateway, then waits for the next activation to restore that committed snapshot. OpenClaw’s native lock protects gateway ownership, but does not serialize filesystem mutations performed by onboarding, plugins, configuration, or backup operations.
3.3 Emergency actions — when to use each
| Action | Use when | Risk |
|---|---|---|
POST /debug/stop-gateway | Gateway wedged but config is fine; you want to manually start fresh | Low — does not destroy data |
POST /debug/restart-openclaw | Config valid but gateway process dead; you want a clean in-place restart | Low |
POST /debug/reprovision | Container lost/orphaned (e.g. destroyed from the dashboard) but the DO is alive; gateway RPCs are timing out | Medium — forces a cold start (R2 restore keeps state) |
POST /debug/destroy-container | You want the platform to provision a brand-new container (fresh state) | High — loses in-memory + unsnapshotted state |
POST /debug/rescue-container | Container is in Error mode / image deleted (see §6.1) | Medium — destroys stuck instance, keeps R2 backups |
3.4 The two diagnostic operations you will script most
Both live in src/diagnostics.ts and accept a dryRun option — always run dry first.
POST /debug/restart-openclaw — checkpointed hard gateway restart
Exposed as the diagnostic dashboard button (ROOT, DEBUG_ROUTES=true). It is the hard
restart path (the opposite of the Admin API’s soft reload):
- Find the live gateway process via
findExistingGatewayProcess(sandbox, { probeLive: true }). - Create and commit a snapshot of
/home/openclaw. killGateway(sandbox)—SIGTERMwithGATEWAY_KILL_GRACE_MS(5s) grace,SIGKILLonly if still active; uses the[o]penclawpkill bracket trick (no blanket"openclaw"pkill, no manual lock-file deletion — respects OpenClaw’s own single-instance lock).resetStartupCache()— clears the startup health cache.signalRestoreNeeded(BACKUP_BUCKET)— flags R2 restore so the next request reprovisions a fresh gateway viastart-openclaw.sh.
If the checkpoint fails, the gateway is not killed. There is no direct binary
launch; reprovisioning always goes through start-openclaw.sh.
The Admin API
POST /api/admin/gateway/restartis the soft reload: it sendsSIGUSR1to the liveopenclaw gateway runprocess (softReloadGateway), thenresetStartupCache(), and only falls back tosignalRestoreNeeded()when no live gateway is found.OPENCLAW_NO_RESPAWN=1(Dockerfile) makes the gateway reload in-process, so soft reload never spawns an orphan.
POST /debug/reprovision — force a container reprovision
The escape hatch for an orphaned/dead container whose Sandbox Durable Object is still alive — most commonly after the instance was destroyed from the Cloudflare Containers dashboard, which leaves the DO pointing at a placement that no longer exists.
clearPersistenceCache()/resetStartupCache()— drop the per-isolate “already restored / startup ok” facts about the old container.clearProvisioningFailures(sandbox)— reset the orphaned-instance failure streak.sandbox.destroy()— tear down the instance so the next RPC provisions a fresh one.
Unlike restart-openclaw, this never talks to the gateway, so it still works when the container
is completely gone and gateway RPCs are timing out.
curl -X POST "https://<worker>/debug/reprovision" | jq .rescueContainer(sandbox, env, options) — diagnostics.ts:296
Recovers a container stuck in Error mode (e.g. the fouria:unstable-<sha> image it
references was deleted by the stale-image cleanup cron). Sequence:
- Probe container health (
probeContainerHealth, shortPASSIVE_PROBE_TIMEOUT_MSbound). Healthy →{ action: 'none', ok: true }. - Repair image via the Cloudflare Containers API (
createContainersApiClient→rolloutImagewithstrategy: 'full_auto') if credentials are present and the current image ≠ expected (diagnostics.ts:339). - Destroy the stuck container (
sandbox.destroy()) so the platform provisions a fresh instance against the repaired image.
# Diagnose first (dry run — does nothing)
curl -X POST "https://<worker>/debug/rescue-container" \
-H 'content-type: application/json' -d '{"dryRun":true}' | jq .
# Actually rescue (requires CLOUDFLARE_API_TOKEN + CONTAINER_APPLICATION_ID)
curl -X POST "https://<worker>/debug/rescue-container" \
-H 'content-type: application/json' -d '{"dryRun":false}' | jq .
The result includes action ∈
none | destroy | repair-image | repair-image-and-destroy | error, a diagnosis
(health + errorClass), and the image report (current, expected, repaired,
rolloutId).
Error classification (what “permanent” vs “transient” means)
classifyContainerError() (diagnostics.ts:166) decides whether retrying helps:
- Permanent (no point retrying):
no such image,no matching app,no namespace configured,did not call start,out of memory,resource exhaustion,pid limit. → a rescue (image repair + destroy) is required. - Transient (recovers on its own):
connection refused,econnrefused,etimedout,timed out,not mapped,no container instance available. → just wait / warm up. - unknown → treat as transient unless health stays red.
This classification drives the degraded flag on the health snapshot (health-probes.ts:281):
a permanent error is what flips degraded:true and tells you a rescue is needed.
4. Working with Backups
State lives in the container at /home/openclaw and is preserved as squashfs snapshots
in R2 (persistence.ts). The gateway restarts from the latest snapshot on boot.
4.1 The persistence model
| Concept | Where | Notes |
|---|---|---|
| Backup store handle | backup-handle.json in R2 | Tracks up to maxVersions (default 3) |
| Backup objects | backups/<id>/data.sqsh[.enc], meta.json | .enc when encrypted at rest |
| Config versions | configs/<timestamp>.json | Kept 30 days (CONFIG_RETENTION_DAYS) |
| Restore marker | restore-needed key | Signals a restore on next boot |
| Lock | backup-lock (60s TTL) | Prevents concurrent scheduled backups |
4.2 Encryption at rest (F4E1)
When BACKUP_ENCRYPTION_KEY is set, every snapshot is post-encrypted with AES-256-GCM
before it lands in R2 (persistence.ts:405, createSnapshot). Config versions under
configs/ are encrypted with customMetadata.encrypted='true'.
# Generate a key (store it as a secret_text binding, NEVER in .dev.vars for prod)
openssl rand -base64 32
BACKUP_ENCRYPTION_KEYunset → backups are stored UNENCRYPTED (warned in logs).BACKUP_ENCRYPTION_REQUIRED=true→ backups fail closed if the key is missing (persistence.ts:419). Recommended in production.
Legacy plaintext backups remain restorable; decryption is transient and deleted in a
finallyblock (persistence.ts:223). Seedocs/plans/r2-encryption-at-rest.md.
4.3 How backups happen
- Scheduled —
scheduleBackupIfNeeded()runs (interval-gated, default every 5 min viaBACKUP_INTERVAL_MINUTES) each time/api/healthis called while not passive (public.ts:197). Guarded by a 60s R2 lock. - On hard restart —
POST /debug/restart-openclawfirst creates a checkpoint, then clears the startup cache and flags R2 restore. - Manual —
POST /api/admin/storage/synccallscreateSnapshot(). POST /debug/r2-putonly writes test objects; it does not create a snapshot.
Automatic backup trigger state machine
Since the Worker cron was removed, the only automatic snapshot trigger is a
non-passive GET /api/health — in practice the lerma HealthMonitor
(apps/lerma/lib/lerma/provisioning/health_monitor.ex, default every 120s).
Everything else in the state machine is a gate that must pass before an object
lands in R2.
stateDiagram-v2
direction TB
[*] --> Idle
state "Idle (no snapshot scheduled)" as Idle
state "Non-passive GET /api/health" as Probe
state "Interval elapsed?" as IntervalGate
state "backup-lock free?" as LockGate
state "Snapshotting (sandbox.createBackup)" as Snapshot
state "Encrypting at rest (F4E1)" as Encrypt
state "Committing catalog + pruning" as Commit
state "No backup written" as NoBackup
Idle --> Probe : lerma HealthMonitor (2 min; production/error clients only)
Idle --> Snapshot : hard restart checkpoint (POST /debug/restart-openclaw)
Idle --> Snapshot : manual sync (POST /api/admin/storage/sync)
Probe --> Idle : ?passive=1 skips backups
Probe --> IntervalGate : non-passive
IntervalGate --> Idle : < BACKUP_INTERVAL_MINUTES since last
IntervalGate --> LockGate : interval elapsed and BACKUP_BUCKET bound
LockGate --> Idle : backup-lock held (< 60s TTL)
LockGate --> Snapshot : lock acquired
Snapshot --> NoBackup : InvalidBackupConfigError / R2 error
Snapshot --> Encrypt : archive created
Encrypt --> NoBackup : BACKUP_ENCRYPTION_REQUIRED and key missing
Encrypt --> Commit : encrypted (.enc) or plaintext (warned)
Commit --> Idle : backup-handle.json written, oldest pruned
NoBackup --> Idle : lock released
The same state machine as an Eraser
flowchart. Render it by POSTing the DSL to the Eraser render API (see the
eraser-diagrams skill); it is kept inline so the diagram can be regenerated
without an API key.
direction down
Start [shape: oval]
Idle [shape: rectangle, label: "Idle - no snapshot scheduled"]
Probe [shape: rectangle, label: "Non-passive GET /api/health"]
HardRestart [shape: rectangle, label: "Hard restart checkpoint"]
ManualSync [shape: rectangle, label: "Manual sync"]
PassiveSkip [shape: oval, color: gray, label: "Skip (passive=1)"]
IntervalGate [shape: diamond, label: "Interval elapsed?"]
LockGate [shape: diamond, label: "backup-lock free?"]
Snapshot [shape: rectangle, label: "Snapshotting (sandbox.createBackup)"]
Encrypt [shape: rectangle, label: "Encrypting at rest (F4E1)"]
Commit [shape: rectangle, label: "Commit catalog + prune"]
NoBackup [shape: oval, color: red, label: "No backup written"]
Start > Idle
Idle > Probe : "HealthMonitor (2 min; production/error clients)"
Idle > HardRestart
Idle > ManualSync
Probe > PassiveSkip : "passive=1"
Probe > IntervalGate : "non-passive"
IntervalGate > Idle : "too soon"
IntervalGate > LockGate : "yes + BACKUP_BUCKET bound"
LockGate > Idle : "locked (<60s TTL)"
LockGate > Snapshot : "acquired"
HardRestart > Snapshot
ManualSync > Snapshot
Snapshot > NoBackup : "InvalidBackupConfigError / R2 error"
Snapshot > Encrypt : "archive created"
Encrypt > NoBackup : "encryption required, no key"
Encrypt > Commit : "encrypted or plaintext warned"
Commit > Idle : "catalog written, oldest pruned"
NoBackup > Idle : "lock released"| Automatic trigger | Entry point | Gates that must pass |
|---|---|---|
| Scheduled snapshot | non-passive GET /api/health (public.ts:200) | lerma HealthMonitor must poll the tenant; BACKUP_BUCKET bound; BACKUP_INTERVAL_MINUTES elapsed; 60s lock free; SDK presigned creds present (or DEV_MODE local bucket) |
| Hard-restart checkpoint | POST /debug/restart-openclaw | none (unconditional checkpoint before kill) |
| Manual snapshot | POST /api/admin/storage/sync | caller authorized (ADMIN/ROOT) |
Two operator traps that leave the bucket empty even though the worker is healthy.
First,
HealthMonitor.list_monitored/0only probes clients whosestatusisproductionorerror— a tenant left indraft(or any other status) is never probed, so no scheduled backup ever fires.Second, the Sandbox SDK (>= 0.12.9)
createBackup()on the remote path callsrequirePresignedURLSupport(), which throwsInvalidBackupConfigErrorunlessR2_ACCESS_KEY_ID,R2_SECRET_ACCESS_KEY,BACKUP_BUCKET_NAME, andCLOUDFLARE_ACCOUNT_IDare all set on the tenant worker. The failure is caught byscheduleBackupIfNeeded()and logged as[persistence] Scheduled backup failed:— the R2 bucket then contains onlybackup-handle.json(or nothing at all).Verify both before assuming backups are running.
To verify the R2 backup bucket is reachable without triggering a snapshot, use
POST /api/admin/storage/test-connectivity (ADMIN or ROOT). It HEADs the backup catalog
handle; a 404 (no backup written yet) still counts as connected. ROOT users can run it from
the Diagnostic Dashboard’s Test Backup Connectivity button even though the Backups page
is ADMIN-gated.
The number of retained versions is BACKUP_MAX_VERSIONS (default 3); oldest beyond that
are pruned (persistence.ts:306).
4.4 Restoring
- Automatic —
restoreIfNeeded()runs at container startup; it finds the newest valid backup and restores it. A missing/again-missingrestore-neededmarker is the signal. - Targeted —
restoreBackupById(sandbox, bucket, backupId, env)restores a specific snapshot (persistence.ts:236). Use the backup history + config version list to pick one.
Each entry in backup-handle.json records whether it was created with the SDK’s
localBucket mode (persistence.ts). That flag is forwarded to restoreBackup() so a
DEV_MODE (local-binding) backup restores through the local R2 binding instead of the
remote presigned-URL path — a backup created locally only exists in the local bucket, so a
remote restore 404s with Presigned URL download failed (exit code 22).
Entries created before the flag was persisted carry no localBucket; on restore it is
defaulted from the same DEV_MODE signal used at create time, so legacy local backups
still restore via the local binding. A restore that 404s at the presigned URL (archive
gone / unreachable) is treated like BACKUP_NOT_FOUND — the stale entry is dropped from
the catalog so one dead backup cannot block every subsequent boot.
Operator workflow for “I need to roll back to a known-good config”:
- List config versions:
GET /api/admin/storage/config(or inspectconfigs/in R2). - Confirm the snapshot you want exists in
backup-handle.json’sbackups[]. - If the container is up, stop the gateway (
POST /debug/stop-gateway), then call a targeted restore, then restart. If the container is down, the next boot auto-restores the newest.
⚠️ The current implementation uses the Sandbox backup API and the
BACKUP_BUCKETR2 binding. It does not mount R2 with s3fs. Never delete tenant backup objects outside the retention or decommission procedure.
4.5 Setting up R2 (one-time)
Scheduled backups run when the Worker has its BACKUP_BUCKET binding and the
interval/lock guards pass (scheduleBackupIfNeeded() in persistence.ts).
The bound bucket is provisioned per tenant, for example
fouria-backup-{slug}{-environment}{-platform}.
The Sandbox SDK remote path requires R2 S3 credentials. In
@cloudflare/sandbox >= 0.12.9,createBackup()(andrestoreBackup()) on the non-localBucketpath callsrequirePresignedURLSupport(), which throwsInvalidBackupConfigErrorunless all ofR2_ACCESS_KEY_ID,R2_SECRET_ACCESS_KEY,CLOUDFLARE_ACCOUNT_ID, andBACKUP_BUCKET_NAMEare present on the tenant Worker.hasBackupCredentials()only gates onBACKUP_BUCKET, so a misconfigured worker still runs the scheduler and then fails at snapshot time (logged as[persistence] Scheduled backup failed:) with an empty bucket. OnlyDEV_MODE=true(the local R2 binding path) avoids the credential requirement.
-
Bind the tenant bucket — the Lerma provisioner creates and binds the tenant R2 bucket as
BACKUP_BUCKETand setsBACKUP_BUCKET_NAME. -
Let lerma resolve it (it is now mandatory) — the provisioner resolves a credential per instance and fails provisioning (rolling back) when none can be produced; it no longer warns and deploys a tenant that silently cannot back up. Strategy is
R2_TOKEN_STRATEGY(auto|shared|per_instance):auto(default): usesR2_TOKEN_MAX_PER_ACCOUNT; a known limit of at most 250 selectsshared, a larger or unknown limit selectsper_instance. Cloudflare exposes no maximum-token endpoint, so the limit defaults to unknown → per-instance token minted at instance creation, scoped to the tenant bucket.shared: setR2_ACCESS_KEY_ID/R2_SECRET_ACCESS_KEYin the lerma environment (generate withmix lerma.r2_token --all-buckets).
Generated per-instance credentials are persisted on the instance record and reused across redeploys; a token minted for a failed brand-new provisioning is deleted on rollback.
CLOUDFLARE_ACCOUNT_IDis a plain var and is always injected. -
Ensure the client is being probed — the lerma
HealthMonitoris the only automatic trigger and only polls clients inproduction/error. A client indraftgets no scheduled snapshots regardless of credentials.# For a standalone/legacy Worker (not provisioned via lerma): npx wrangler secret put R2_ACCESS_KEY_ID npx wrangler secret put R2_SECRET_ACCESS_KEY
Sandbox snapshot behavior differs between deployed and local environments; local tests use the SDK’s
localBucketmode whenDEV_MODE=true(persistence.ts:398). The flag is persisted in the backup catalog so restores route to the same mode — a local backup restores via the local R2 binding, and a remote (presigned-URL) backup restores viaR2_ACCESS_KEY_ID/R2_SECRET_ACCESS_KEY/CLOUDFLARE_ACCOUNT_ID/BACKUP_BUCKET_NAME. Pre-persist entries default to the currentDEV_MODE. There is no s3fs mount.
5. Troubleshooting — Critical Issues
This section maps symptoms → root cause → the exact recovery from this guide.
5.1 Container stuck in Error / image deleted → RESCUE
Symptom: /api/health?passive=1 returns degraded:true,
containerError:"no such image" (or no matching app). The instance can never boot.
Why: the stale-image cleanup cron deleted the fouria:unstable-<sha> image the
application references (diagnostics.ts:42). The cleanup script now protects every
tag still referenced by a container application — including the versions of an
in-progress rollout — and aborts (fail closed) when that in-use set cannot be
resolved (.github/scripts/cleanup-stale-images.sh). Production deploys are also
restricted to stable release tags, so a production application can no longer be
pinned to the unstable-<sha> channel (Lerma.Cloudflare.ImageRegistry.stable?/1).
Fix: rescueContainer (§3.4):
curl -X POST "https://<worker>/debug/rescue-container" -d '{"dryRun":true}' | jq .
# confirm action would be repair-image-and-destroy, then:
curl -X POST "https://<worker>/debug/rescue-container" -d '{"dryRun":false}' | jq .
Credentials required: CLOUDFLARE_API_TOKEN (Containers:Edit),
CONTAINER_APPLICATION_ID, and either CONTAINER_IMAGE_TAG or an explicit
expectedImage (diagnostics.ts:304, .dev.vars.example:43). Without them, the image
cannot be repaired — only the container is destroyed (you then must re-deploy the
correct image tag).
NEVER set these on tenant workers — the rescue tool re-points the application image
and is operator/ROOT-only (.dev.vars.example:40).
5.2 Gateway “already running” / EADDRINUSE on restart
Symptom: restart fails with gateway already running (pid …) or failed to bind gateway socket.
Why: a previous gateway still holds the native lock / TCP bind when the new one
starts (gateway/process.ts).
Fix: for a clean state use killGateway() (the POST /debug/stop-gateway route), which
SIGTERMs with the GATEWAY_KILL_GRACE_MS (5s) grace and only SIGKILLs if still active,
then wait for the port to close before reprovisioning. POST /debug/restart-openclaw (hard
restart) does this then reprovisions via start-openclaw.sh. Inspect lock state with
GET /debug/gateway-lock.
5.3 Gateway process invisible but container live
Symptom: listProcesses() shows nothing, yet /health answers.
Why: the gateway runs as a detached child of start-openclaw.sh
(process.ts:427). The SDK does not track it, but it is healthy.
Fix: the live-gateway fallback (findExistingGatewayProcess with probeLive:true)
detects it via a bounded /health probe. Prefer restart/rescue APIs that use this path
rather than destroy-container (which would throw away a healthy instance).
5.4 Backup restore loops / “hadBackup but not restored”
Symptom: health startup.ok=false permanently with
Backup found but not restored.
Why: restoreIfNeeded found an entry but every backup expired/was missing
(persistence.ts:212). The code explicitly reports a clean start in this case so it
does not gate startup forever (persistence.ts:233).
Fix: confirm real backups exist (backup-handle.json → backups[]); if none, the
instance boots fresh (expected). If backups exist but restore fails, check
BACKUP_ENCRYPTION_KEY (encrypted backups cannot be read without it).
5.5 Platform & setup troubleshooting
| Symptom | Likely cause | Action |
|---|---|---|
npm run dev / deploy: Unauthorized | Cloudflare Containers not enabled for the account | Enable Containers in the Containers dashboard |
[sandbox-do] backup.restore … Presigned URL download failed (exit code 22) … 404 | restore hit the remote presigned-URL path for an archive that only exists in the local R2 binding (a DEV_MODE/localBucket backup), or BACKUP_BUCKET_NAME doesn’t match the BACKUP_BUCKET binding bucket | redeploy so the backup catalog persists localBucket (§4.4) — local backups restore via the local binding; for remote backups verify R2_ACCESS_KEY_ID/R2_SECRET_ACCESS_KEY/CLOUDFLARE_ACCOUNT_ID/BACKUP_BUCKET_NAME all point at the same bucket as the BACKUP_BUCKET binding (§4.5) |
| Gateway won’t start | missing/incorrect secrets | npx wrangler secret list + npx wrangler tail |
| Config changes not taking effect | stale Docker build cache | bump the # Build cache bust: comment in Dockerfile, then redeploy |
| Slow first request | cold start (1–2 min) | expected; later requests are faster |
| R2 not mounting | an R2 secret missing, or running wrangler dev | set all four R2 secrets (§4.5); R2 mount is production-only |
Access denied on /_admin or /api | CF_ACCESS_TEAM_DOMAIN / CF_ACCESS_AUD not set | configure Cloudflare Access (§8.1) |
| Devices not appearing in admin UI | device-list CLI takes 10–15 s (WS overhead) | wait and refresh |
| WebSocket issues in local dev | wrangler dev WS-proxy limitation | deploy for full WS functionality |
Windows: script exits 126 | CRLF line endings in shell scripts | add .gitattributes with * text=auto eol=lf |
5.6 Deploy-window “Could not list processes” / platform interrupted
Symptom: logs show [Gateway] Could not list processes: followed by a stack trace ending in
createPlatformInterruptedError / translatePlatformInterruption, and/or
/debug/processes returns a 503 with { "status": "platform_updating" }.
Why: the Sandbox SDK throws OperationInterruptedError (“Sandbox operation listProcesses was
interrupted while the platform was updating the sandbox runtime”) when the Durable Object isolate is
replaced — exactly what happens during a deploy / code update. It is a transient platform
lifecycle event, not an application failure, and it self-heals on the next request once the new
isolate is live.
Fix: nothing to fix. In findExistingGatewayProcess this is logged as a single concise
“platform was updating the sandbox runtime” line instead of a raw stack trace; /debug/processes
returns 503 platform_updating so debug tooling doesn’t misreport a crash. If it persists beyond
the deploy window, re-check wrangler tail and the container instance health in the Containers
dashboard.
5.7 Container terminated/destroyed → every request 500s or burns a cold-start budget
Symptom A — platform-terminated: the platform SIGTERMed the container (idle recycle,
maintenance, rescue) and after that all requests fail with [Gateway] Could not list processes:
plus sandbox.exec error … HTTP error! status: 500 — the container is dead and nothing ever
restarts it. The dashboard shows “Could not list processes” and the UI is fully bricked.
Symptom B — out-of-band destroyed (orphaned): the instance was destroyed from the Cloudflare
Containers dashboard (or a failed operator force-stop). The Sandbox DO is still alive but points
at a placement that no longer exists, so every request logs around no container instance available / maximum number of running container instances exceeded and each request burns the
SDK’s full provisioning retry budget (≈20s instance + 90s port) before giving up — CPU is spent
repeatedly trying to reach a container that can never be admitted, and the UI never recovers.
Why: the Worker holds a Sandbox stub for a DO whose container instance is gone. Retrying any
RPC against that instance fails forever, and the old code swallowed the error and returned null
instead of reprovisioning — so no fresh container was ever created
(gateway/sandbox-recovery.ts).
Fix (automatic): maybeRecoverSandbox classifies both failure classes and calls
recoverStaleSandbox (destroy the dead instance + reset the startup/persistence caches +
reset the orphan streak). The next RPC auto-provisions a fresh container and the normal cold-start
flow runs (R2 restore → start-openclaw.sh). Wired into findExistingGatewayProcess,
checkGatewayHealth, and the ensureGateway startProcess retry (gateway/process.ts).
- Stale (
HTTP error! status: 500) recovers immediately (isStaleContainerError). - Orphaned (
no container instance available) is retryable on its own, so it is only treated as an orphan once the failures spanORPHANED_INSTANCE_GRACE_MS(120s) — deliberately longer than the SDK’s ≈110s provisioning budget so a legitimately slow cold start is never cancelled mid-provision. The streak is cleared whenever a sandbox RPC succeeds (the container is demonstrably alive), so a healthy instance never trips the threshold later.
Recoveries are cooldown-guarded (STALE_RECOVERY_COOLDOWN_MS, 60s) to prevent destroy-storms
during request floods. Look for [SandboxRecovery] and repeated provisioning failures lines in
wrangler tail. Manual fallback: POST /debug/reprovision (works even when the gateway is gone),
or restart the container from the Containers dashboard.
NOT covered: permanent image errors (no such image) still require the rescue tool (§5.1),
and deploy-window OperationInterruptedError is expected and self-heals (§5.6).
6. Subsystem Deep-Dives
6.1 Working with Cloudflare Containers
- Image reference is built as
registry.cloudflare.com/<accountId>/<name>:<tag>(buildImageReference,diagnostics.ts:213).namedefaults tofouria(DEFAULT_CONTAINER_IMAGE_NAME). - Rescue needs the Containers API (
/accounts/{account}/containers/applications/...,createContainersApiClient,diagnostics.ts:235). It usesAuthorization: Bearer <CLOUDFLARE_API_TOKEN>and afull_autorollout. - Timeouts that keep a degraded/off container from hanging:
CONTAINER_INSTANCE_TIMEOUT_MS(20s),CONTAINER_PORT_TIMEOUT_MS(90s — kept generous for cold starts),CONTAINER_POLL_INTERVAL_MS(300ms) (config.ts:33). Override viaSANDBOX_INSTANCE_TIMEOUT_MS/SANDBOX_PORT_TIMEOUT_MS/SANDBOX_POLL_INTERVAL_MS(.dev.vars.example:78).
Container cost & lifecycle
The tenant container is a standard-1 instance (½ vCPU, 4 GiB memory, 8 GB disk).
Approximate 24/7 monthly cost (per Cloudflare Containers pricing):
| Resource | Provisioned | Approx. cost |
|---|---|---|
| Memory (4 GiB) | billed 24/7 | ~$26/mo |
| CPU (½ vCPU, ~10% used) | usage-based | ~$2/mo |
| Disk (8 GB) | billed 24/7 | ~$1.50/mo |
| Workers Paid plan | flat | $5/mo |
| Total | ~$34.50/mo |
Notes:
-
CPU is billed on active usage, not provisioned capacity. Memory/disk are billed on provisioned capacity for the full uptime.
-
Cold starts take 1–2 minutes; the first request after a sleep/wake is slow.
-
Reduce cost with
SANDBOX_SLEEP_AFTER— the container sleeps when idle instead of running 24/7. A container used ~4 h/day drops to roughly $5–6/mo compute + the $5 plan fee.npx wrangler secret put SANDBOX_SLEEP_AFTER # e.g. 10m, 30m, 1hnever(the default) keeps it alive indefinitely (gateway/process.ts:689).
6.2 R2 issues
| Symptom | Likely cause | Action |
|---|---|---|
| Backup stored UNENCRYPTED warning | BACKUP_ENCRYPTION_KEY unset | Set the key (secret_text binding) |
| Backup fails closed | BACKUP_ENCRYPTION_REQUIRED=true, key missing | Provide the key |
| Restore of encrypted backup fails | key mismatch / missing | Ensure BACKUP_ENCRYPTION_KEY matches the one used at write time |
| Scheduled backup never runs | missing R2_ACCESS_KEY_ID / R2_SECRET_ACCESS_KEY / BACKUP_BUCKET_NAME / CLOUDFLARE_ACCOUNT_ID | Sandbox SDK presigned-URL path needs all four; hasBackupCredentials (persistence.ts) gates the scheduler, requirePresignedURLSupport throws InvalidBackupConfigError at snapshot time |
| Old config versions pile up | retention is 30 days; normal | cleanupOldConfigVersions prunes automatically |
R2 persistence uses the Sandbox SDK backup API; there is no s3fs mount to inspect.
Verify backup state with GET /api/admin/storage/backups and never delete tenant
backup objects outside retention or decommission procedures.
6.3 AI Gateway issues
The only provisioned AI provider path is Cloudflare AI Gateway, configured when all
three are set (start-openclaw.sh:161, config.ts:97):
npx wrangler secret put CLOUDFLARE_AI_GATEWAY_API_KEY # your gateway API key
npx wrangler secret put CF_AI_GATEWAY_ACCOUNT_ID # Cloudflare account id
npx wrangler secret put CF_AI_GATEWAY_GATEWAY_ID # gateway id within the account
Cloudflare AI Gateway gives caching, rate limiting, analytics, and cost tracking, and the
cloudflare-unified-billing plugin extends the catalog to 150+ models (Google Gemini,
Anthropic Claude, OpenAI, xAI Grok, DeepSeek, Alibaba Qwen, MiniMax, Moonshot Kimi, and
more) via the AI Gateway REST API. Lerma provisions the AI-scoped
credential to tenant workers as both CLOUDFLARE_AI_GATEWAY_API_KEY and the REST alias
CLOUDFLARE_API_TOKEN.
Premade agents:
Production deployments ship no premade agents — only the main agent written by
openclaw onboard. The dev-only Fouria Builder and Fouria Q&A agent are injected
by moltlazy when DEV_MODE=true (they can self-install plugins and read per-tenant client
documentation materialized from R2).
| Agent | Environment | Primary model | Purpose |
|---|---|---|---|
fouria-builder | dev only | Claude Sonnet 4.6 | Read client proposal/phases, build a working POC |
fouria-qa | dev only | Claude Sonnet 4.6 | Run cctr corpus validation against the Builder’s POC |
| Symptom | Cause | Action |
|---|---|---|
| Onboard skips AI Gateway auth | one of the three vars missing | set all three; mark NEED_ONBOARD=true |
cloudflare-ai-gateway plugin missing | install timed out (network) | warm restore keeps it; cold start reinstalls (bounded, non-fatal) |
| Model requests fail with auth errors | wrong/expired gateway key | rotate CLOUDFLARE_AI_GATEWAY_API_KEY |
| No embedding / memory search errors | no embedding key (CF_AI_GATEWAY_OPENAI_COMPAT_KEY) | memory search auto-disabled (start-openclaw.sh:422) — expected |
The cloudflare-unified-billing plugin (40+ models via the REST API) is enabled
alongside; its enablement is written by the batched config patch, not a plugins enable CLI call, to avoid a hanging ClawHub registry refresh (start-openclaw.sh:300).
When migrating existing tenants from the old provisioner, re-provision the tenant and perform a cold restart after the AI credential is configured. Older R2 snapshots may contain a stale native-provider auth profile and the old random AI binding; a warm restore alone does not replace those persisted values.
6.4 Startup process issues
The startup sequence is start-openclaw.sh (STEPS array, start-openclaw.sh:69):
init → onboard → plugins → config-patch → config-validate → tunnel → gateway. Watch it
live via:
curl "https://<worker>/api/health" # status: starting while script runs
curl "https://<worker>/debug/gateway-logs" | tail -n 50| Symptom | Cause | Action |
|---|---|---|
Stuck at config-validate → startup aborts | invalid generated config | moltlazy validate locally; fix feature flags |
Stuck at onboard > 60s | AI Gateway plugin not preinstalled | the cloudflare-ai-gateway provider is preinstalled before onboard (so its plugin phase is a no-op); if it still times out, check ClawHub/npm reachability — onboard is capped at 60s and continues, gateway starts --allow-unconfigured |
| Gateway dies before port opens | gateway binary crash | GET /debug/gateway-logs for the crash; check EADDRINUSE |
STARTUP_TIMEOUT_MS (300s) exceeded | plugin installs slow | cold start is bounded; warm restores are faster |
| Loading page never advances | startup-progress.json not written | check /tmp/startup.log via debug/gateway-logs |
The loading page renders from /tmp/startup-progress.json + /tmp/startup.log
(routes/public.ts:131, read bounded by STARTUP_PROGRESS_TIMEOUT_MS).
8. Cloudflare Platform Setup
One-time operator setup tasks for the current FourIA Worker deployment.
8.1 Cloudflare Access (protects /_admin, /api, /debug)
-
Enable Access on the worker — Workers & Pages dashboard → your Worker → Settings → Domains & Routes →
workers.dev→ Enable Cloudflare Access. Copy the Application Audience (AUD) tag. -
Configure who can access — Zero Trust → Access → Applications → your worker app → add allowed identities (email, Google, GitHub, …).
-
Set the secrets:
npx wrangler secret put CF_ACCESS_TEAM_DOMAIN # e.g. myteam.cloudflareaccess.com npx wrangler secret put CF_ACCESS_PLATFORM_AUD # platform AUD (always required) npx wrangler secret put CF_ACCESS_CLIENT_AUD # optional per-tenant AUDAlternatively create a Self-hosted Access application and protect paths
/_admin/*,/api/*,/debug/*.
/debug/*must be Access-protected. Never expose it publicly — it can stop, destroy, and rescue containers.
8.2 Browser Automation (CDP)
A Chrome DevTools Protocol shim for browser automation. Set:
npx wrangler secret put CDP_SECRET # shared secret for CDP auth
npx wrangler secret put WORKER_URL # worker public URL
npm run deploy| Endpoint | Description |
|---|---|
GET /cdp/json/version | Browser version information |
GET /cdp/json/list | List available browser targets |
GET /cdp/json/new | Create a new browser target |
WS /cdp/devtools/browser/{id} | WebSocket connection for CDP commands |
All endpoints require ?secret=<CDP_SECRET>.
8.3 Obsidian Sync via Cloudflare Tunnel
Exposes the Obsidian LiveSync-compatible service (port 18790). Requires an active Cloudflare
zone for your domain.
cloudflared tunnel create obsidian-sync
cloudflared tunnel token <UUID> # then:
npx wrangler secret put CLOUDFLARE_TUNNEL_TOKEN
Add a proxied CNAME: name obsidian-sync, target <UUID>.cfargotunnel.com. Ingress
(~/.cloudflared/config.yml): hostname: obsidian-sync.yourdomain.com → service: http://localhost:18790. The gateway only starts the tunnel when
CLOUDFLARE_TUNNEL_TOKEN is set (start-openclaw.sh:497); otherwise obsidian-sync is
localhost-only.
8.4 Built-in skills
The container ships pre-installed skills in /root/clawd/skills/, including
cloudflare-browser (browser automation via the CDP shim; needs CDP_SECRET +
WORKER_URL). Scripts: screenshot.js, video.js, cdp-client.js.
9. Requirements & References
Requirements
- Workers Paid plan ($5/mo) — required for Cloudflare Sandbox containers.
- An AI provider key via Cloudflare AI Gateway (API key, account id, gateway id) — the only supported AI path. Unified Billing is available as an alternative (docs).
Free tiers exist for Cloudflare Access, Browser Rendering, AI Gateway, and R2 Storage.
See also
- OpenClaw · OpenClaw Docs
- Cloudflare Sandbox · Cloudflare Access
- Worker deep-dive · Root DEVELOPMENT guide
- Env vars reference
7. Quick Reference
Environment variables (operator-relevant)
| Variable | Purpose | Notes |
|---|---|---|
DEBUG_ROUTES | Enables /debug/* | Set true only on operator/ROOT deploys |
CLOUDFLARE_API_TOKEN | AI Gateway REST auth; Containers:Edit for rescue | AI-scoped tenant alias; rescue token on ROOT/ops |
CONTAINER_APPLICATION_ID | Target app for image rollout | Operator-only |
CONTAINER_IMAGE_TAG / CONTAINER_IMAGE_NAME | Expected image unstable-<sha> / fouria | Operator-only |
CLOUDFLARE_AI_GATEWAY_API_KEY / CF_AI_GATEWAY_ACCOUNT_ID / CF_AI_GATEWAY_GATEWAY_ID | AI Gateway auth | Required for AI provider |
BACKUP_ENCRYPTION_KEY / BACKUP_ENCRYPTION_REQUIRED | R2 encryption at rest | Recommended true in prod |
BACKUP_BUCKET_NAME / R2_ACCESS_KEY_ID / R2_SECRET_ACCESS_KEY | R2 persistence | All four (incl. CLOUDFLARE_ACCOUNT_ID) required for scheduled backups; SDK signs R2 presigned URLs |
BACKUP_INTERVAL_MINUTES / BACKUP_MAX_VERSIONS | Backup cadence / retention | Defaults 5 / 3 |
SANDBOX_INSTANCE_TIMEOUT_MS / SANDBOX_PORT_TIMEOUT_MS / SANDBOX_POLL_INTERVAL_MS | Container fail-fast bounds | See §6.1 |
SANDBOX_SLEEP_AFTER | Idle sleep to cut cost (never = always on) | e.g. 10m; see §6.1 cost note |
CLOUDFLARE_ACCOUNT_ID | Cloudflare account id (plain var) | Read by the Sandbox SDK for R2 presigned URL signing; legacy CF_ACCOUNT_ID is deprecated |
CF_ACCESS_TEAM_DOMAIN / CF_ACCESS_PLATFORM_AUD / CF_ACCESS_CLIENT_AUD | Cloudflare Access auth | Protects /_admin, /api, /debug |
CDP_SECRET / WORKER_URL | Browser automation (CDP) | See §8.2 |
CLOUDFLARE_TUNNEL_TOKEN | Obsidian sync ingress | See §8.3 |
STARTUP_LOG_VERBOSITY | steps/detailed/verbose boot log | Secret-free |
One-line recovery cheatsheet
# 1. Is it up? (no warm-up)
curl "https://<worker>/api/health?passive=1" | jq '{status,degraded,containerError}'
# 2. What's the lock state?
curl "https://<worker>/debug/gateway-lock" | jq .state
# 3. Degraded (no such image)? Rescue (dry first):
curl -X POST "https://<worker>/debug/rescue-container" -d '{"dryRun":true}' | jq .
curl -X POST "https://<worker>/debug/rescue-container" -d '{"dryRun":false}' | jq .
# 4. Gateway wedged but config OK? Stop, wait 8s, restart (via Admin API).
# 5. Lost state? Inspect backups, then restore by id or let next boot auto-restore.
This guide is generated from the source in apps/fouria/src/diagnostics.ts and its
dependencies (gateway/process.ts, gateway/health-probes.ts, gateway/lock.ts,
persistence.ts). When the behavior of
those modules changes, update this document in the same change.