Gateway troubleshooting
This page is the deep runbook. Start at /help/troubleshooting if you want the fast triage flow first.Command ladder
Run these first, in this order:openclaw gateway statusshowsRuntime: runningandRPC probe: ok.openclaw doctorreports no blocking config/service issues.openclaw channels status --probeshows connected/ready channels.
Anthropic 429 extra usage required for long context
Use this when logs/errors include:HTTP 429: rate_limit_error: Extra usage is required for long context requests.
- Selected Anthropic Opus/Sonnet model has
params.context1m: true. - Current Anthropic credential is not eligible for long-context usage.
- Requests fail only on long sessions/model runs that need the 1M beta path.
- Disable
context1mfor that model to fall back to the normal context window. - Use an Anthropic API key with billing, or enable Anthropic Extra Usage on the subscription account.
- Configure fallback models so runs continue when Anthropic long-context requests are rejected.
- /providers/anthropic
- /reference/token-use
- /help/faq#why-am-i-seeing-http-429-ratelimiterror-from-anthropic
No replies
If channels are up but nothing answers, check routing and policy before reconnecting anything.- Pairing pending for DM senders.
- Group mention gating (
requireMention,mentionPatterns). - Channel/group allowlist mismatches.
drop guild message (mention required→ group message ignored until mention.pairing request→ sender needs approval.blocked/allowlist→ sender/channel was filtered by policy.
Bundled plugin warning at startup
Use this when gateway logs includeConfig warnings: followed by
plugin not found: <id> for a plugin you expect to be bundled, especially when
plugins info or normal channel behavior suggests the plugin is actually
available.
plugins info <id>reportsOrigin: bundledandStatus: loaded.plugins.installs.<id>.sourceisbundled, or an older copied bundled record can be reconciled tobundled.- The stored
sourcePathstill points at the current bundled extension path. - The id appears only in stale config (
plugins.entries,allow,deny, orslots.memory) for a plugin you no longer use.
- Restart the gateway so startup reconciliation can refresh bundled install
records: Linux
systemctl --user restart openclaw-gateway; macOSwednesdayai gateway restart --deep. - Run
wednesdayai plugins update <id>to reconcile plugin install metadata before npm update selection runs. - Remove stale plugin references from
plugins.entries,plugins.allow,plugins.deny, orplugins.slots.memoryif the plugin is no longer wanted.
plugins.load.paths. Bundled plugins are
discovered from the installed WednesdayAI package, and a valid source: "bundled"
record is enough for validation and config-schema checks.
ACP startup identity reconcile fails at startup
Use this when gateway logs includestartup identity reconcile failed for <session-key>: code=ACP_BACKEND_UNAVAILABLE from the acp/manager subsystem,
or the summary line
acp startup identity reconcile (renderer=...): checked=N resolved=N failed=N
with a nonzero failed count.
At gateway startup the gateway re-resolves ACP sessions whose harness identity
is still pending. If a session’s backend is not healthy yet, the gateway waits
up to 30 seconds (polling every 250 ms) for it to report healthy, then retries
each deferred session once. Sessions whose backend is still unhealthy when the
wait expires count as failed and log this warning. They are not lost: each
session re-resolves on its next use.
acpx runtime backend registered (command: ..., expectedVersion: ..., pluginLocalInstall: ...)followed within about 30 seconds byacpx runtime backend ready.- No
acpx local binary unavailable or mismatched (...); running plugin-local installline. That line means the plugin-localacpxbinary does not match the pin and an npm install ran during startup. - Run the
commandpath from theacpx runtime backend registeredline with--version; it must print the same version as theexpectedVersionon that line, which is the exactacpxversion declared in the plugin’spackage.json.
- If
acpx runtime backend readynever appears, read the earlieracpx runtime setup failed: ...oracpx runtime backend probe failed after local installlines for the root cause, fix it (missing or brokenacpxbinary, adapter download, harness authentication), and restart the gateway. - If the plugin-local binary version is wrong, restart the gateway and let the plugin-local install re-pin it to the declared version, or reinstall the acpx plugin.
- Affected sessions need no manual repair: they re-resolve on their next use.
Dashboard control ui connectivity
When dashboard/control UI will not connect, validate URL, auth mode, and secure context assumptions.- Correct probe URL and dashboard URL.
- Auth mode/token mismatch between client and gateway.
- HTTP usage where device identity is required.
device identity required→ non-secure context or missing device auth.device nonce required/device nonce mismatch→ client is not completing the challenge-based device auth flow (connect.challenge+device.nonce).device signature invalid/device signature expired→ client signed the wrong payload (or stale timestamp) for the current handshake.unauthorized/ reconnect loop → token/password mismatch.gateway connect failed:→ wrong host/port/url target.
- waits for
connect.challenge - signs the challenge-bound payload
- sends
connect.params.device.noncewith the same challenge nonce
Gateway service not running
Use this when service is installed but process does not stay up.Runtime: stoppedwith exit hints.- Service config mismatch (
Config (cli)vsConfig (service)). - Port/listener conflicts.
Gateway start blocked: set gateway.mode=local→ local gateway mode is not enabled. Fix: setgateway.mode="local"in your config (or runopenclaw configure). If you are running OpenClaw via Podman using the dedicatedopenclawuser, the config lives at~openclaw/.openclaw/openclaw.json.refusing to bind gateway ... without auth→ non-loopback bind without token/password.another gateway instance is already listening/EADDRINUSE→ port conflict.
Gateway restart refused on invalid config
Symptom:openclaw gateway restart exits non-zero with Config invalid, names the config file and the exact issue path, and the gateway keeps running.
This is the pre-flight restart gate working as intended: the gateway never stops when the config it would boot into is already known-invalid. Before this gate, a bad config edit turned a restart into a crash loop under systemd or launchd supervision.
Recovery:
- Read the issue path in the refusal output, for example
plugins.entries.<id>.config: invalid config: must NOT have additional properties. - Repair with
openclaw doctor --fix, or fix the named key manually. - Verify with
openclaw config validate(printsConfig valid). - Retry
openclaw gateway restart.
gateway stopand shutdown signals (SIGTERM,SIGINT) are never gated;SIGUSR1watchdog restarts remain gated.- Watchdog restarts from config reload (
gateway.reload.mode="restart"or"hybrid") are gated the same way: an invalid edit is skipped, the gateway keeps running the previous valid config, and the refusal warning appears in the gateway log. The warning names the config file, the exact issue path, and the same repair commands (openclaw doctor --fix,openclaw config validate). - A first start of a stopped gateway is not gated; invalid config there surfaces as a normal startup error.
Channel connected messages not flowing
If channel state is connected but message flow is dead, focus on policy, permissions, and channel specific delivery rules.- DM policy (
pairing,allowlist,open,disabled). - Group allowlist and mention requirements.
- Missing channel API permissions/scopes.
mention required→ message ignored by group mention policy.pairing/ pending approval traces → sender is not approved.missing_scope,not_in_channel,Forbidden,401/403→ channel auth/permissions issue.
Telegram fleet polling or webhook conflicts
Use this when Telegram logs repeat messages likeNetwork request for 'getUpdates' failed!, account health shows conflicted,
or webhook mode receives requests but does not dispatch updates.
- Per-account
mode,state,owner,retryAt,poll.lastPollAt, andwebhook.lastRequestOutcomein the Channels UI. conflictedstate ortelegram.poll.conflictdiagnostics, which usually means two local accounts, two gateway processes, or a webhook and polling process are trying to own the same bot token.- Repeated
getUpdatesnetwork errors, which usually point to DNS, IPv6, TLS, proxy, firewall, or host egress problems toapi.telegram.org. - A
Polling stall detectedlog line, which means the built-in watchdog force-restarted a wedgedgetUpdatescycle; frequent stalls point to a stuck socket or proxy between the gateway andapi.telegram.org. - Webhook
rejected,malformed,timeout, or non-2xx outcomes, which point to a bad public URL, wrong secret token, reverse-proxy mismatch, or a slow webhook handler.
- Give each enabled
channels.telegram.accounts.<id>entry its ownbotTokenortokenFile;TELEGRAM_BOT_TOKENis only a fallback for the default account. - Ensure one bot token has exactly one owner: one configured account, one gateway process, and either polling or webhook mode.
- If polling, let WednesdayAI clear webhook ownership non-destructively at
startup; if a webhook owner reappears later, polling re-runs that cleanup
automatically on
409instead of conflict-looping. Treat any manualdrop_pending_updatesoperation as destructive unless you intentionally want to discard pending Telegram updates. - If direct Bot API egress is unstable, configure
channels.telegram.proxyor the Telegramnetworksettings described in the Telegram channel guide.
- /channels/telegram#multi-bot-fleet
- /channels/telegram#polling-webhook-and-ownership
- /diagnostics/otel-langfuse#telegram-fleet-diagnostics
Cron and heartbeat delivery
If cron or heartbeat did not run or did not deliver, verify scheduler state first, then delivery target.- Cron enabled and next wake present.
- Job run history status (
ok,skipped,error). - Heartbeat skip reasons (
quiet-hours,requests-in-flight,alerts-disabled).
cron: scheduler disabled; jobs will not run automatically→ cron disabled.cron: timer tick failed→ scheduler tick failed; check file/log/runtime errors.main job requires payload.kind="systemEvent"→ stored job type and payload are incompatible. Recreate or edit it as a Timeline message, Check first, then nudge, or isolated Assistant task job.wake-gate-empty→ a wake-gate cron check found no work, so the skipped run is expected.heartbeat skippedwithreason=quiet-hours→ outside active hours window.heartbeat: unknown accountId→ invalid account id for heartbeat delivery target.heartbeat skippedwithreason=dm-blocked→ heartbeat target resolved to a DM-style destination whileagents.defaults.heartbeat.directPolicy(or per-agent override) is set toblock.
Session run history is empty or stale
Main-session turns (chat, cron, heartbeat) are recorded as session-run rows in the configured session storage backend, powering run history and crash recovery. Recording is best effort: a failed store write never fails the turn itself, so a broken store stays invisible unless you look for the warnings. Ifsessions.runs.list returns no rows or stale timestamps:
session-run store write failedwarns from subsystemsessions/session-runs/recorder. The warn is rate limited to one per minute per operation and carries asuppressedcount of failures held back in that window, so a persistent failure shows as a steady drip instead of log flooding.- Schema drift after an upgrade. With
session.storage.backend: "postgres"or"sqlite"the gateway adds any missing session tables and columns automatically on boot, so a gateway restart repairs drift without manual SQL.
error field from one
warning line and check database connectivity, credentials, and permissions.
Related:
Node paired tool fails
If a node is paired but tools fail, isolate foreground, permission, and approval state.- Node online with expected capabilities.
- OS permission grants for camera/mic/location/screen.
- Exec approvals and allowlist state.
NODE_BACKGROUND_UNAVAILABLE→ node app must be in foreground.*_PERMISSION_REQUIRED/LOCATION_PERMISSION_REQUIRED→ missing OS permission.SYSTEM_RUN_DENIED: approval required→ exec approval pending.SYSTEM_RUN_DENIED: allowlist miss→ command blocked by allowlist.
Browser tool fails
Use this when browser tool actions fail even though the gateway itself is healthy.- Valid browser executable path.
- CDP profile reachability.
- Extension relay tab attachment for
profile="chrome".
Failed to start Chrome CDP on port→ browser process failed to launch.browser.executablePath not found→ configured path is invalid.Chrome extension relay is running, but no tab is connected→ extension relay not attached.Browser attachOnly is enabled ... not reachable→ attach-only profile has no reachable target.
If you upgraded and something suddenly broke
Most post-upgrade breakage is config drift or stricter defaults now being enforced.1) Auth and URL override behavior changed
- If
gateway.mode=remote, CLI calls may be targeting remote while your local service is fine. - Explicit
--urlcalls do not fall back to stored credentials.
gateway connect failed:→ wrong URL target.unauthorized→ endpoint reachable but wrong auth.
2) Bind and auth guardrails are stricter
- Non-loopback binds (
lan,tailnet,custom) need auth configured. - Old keys like
gateway.tokendo not replacegateway.auth.token.
refusing to bind gateway ... without auth→ bind+auth mismatch.RPC probe: failedwhile runtime is running → gateway alive but inaccessible with current auth/url.
3) Pairing and device identity state changed
- Pending device approvals for dashboard/nodes.
- Pending DM pairing approvals after policy or identity changes.
device identity required→ device auth not satisfied.pairing required→ sender/device must be approved.