Skip to main content

Gateway troubleshooting

This page is the deep runbook. Start at /help/troubleshooting if you want the fast triage flow first.

Command ladder

Run these first, in this order:
Expected healthy signals:
  • openclaw gateway status shows Runtime: running and RPC probe: ok.
  • openclaw doctor reports no blocking config/service issues.
  • openclaw channels status --probe shows connected/ready channels.

Anthropic 429 extra usage required for long context

Use this when logs/errors include: HTTP 429: rate_limit_error: Extra usage is required for long context requests.
Look for:
  • Selected Anthropic Opus/Sonnet model has params.context1m: true.
  • Current Anthropic credential is not eligible for long-context usage.
  • Requests fail only on long sessions/model runs that need the 1M beta path.
Fix options:
  1. Disable context1m for that model to fall back to the normal context window.
  2. Use an Anthropic API key with billing, or enable Anthropic Extra Usage on the subscription account.
  3. Configure fallback models so runs continue when Anthropic long-context requests are rejected.
Related:

No replies

If channels are up but nothing answers, check routing and policy before reconnecting anything.
Look for:
  • Pairing pending for DM senders.
  • Group mention gating (requireMention, mentionPatterns).
  • Channel/group allowlist mismatches.
Common signatures:
  • drop guild message (mention required → group message ignored until mention.
  • pairing request → sender needs approval.
  • blocked / allowlist → sender/channel was filtered by policy.
Related:

Bundled plugin warning at startup

Use this when gateway logs include Config warnings: followed by plugin not found: <id> for a plugin you expect to be bundled, especially when plugins info or normal channel behavior suggests the plugin is actually available.
Look for:
  • plugins info <id> reports Origin: bundled and Status: loaded.
  • plugins.installs.<id>.source is bundled, or an older copied bundled record can be reconciled to bundled.
  • The stored sourcePath still points at the current bundled extension path.
  • The id appears only in stale config (plugins.entries, allow, deny, or slots.memory) for a plugin you no longer use.
Fix options:
  1. Restart the gateway so startup reconciliation can refresh bundled install records: Linux systemctl --user restart openclaw-gateway; macOS wednesdayai gateway restart --deep.
  2. Run wednesdayai plugins update <id> to reconcile plugin install metadata before npm update selection runs.
  3. Remove stale plugin references from plugins.entries, plugins.allow, plugins.deny, or plugins.slots.memory if the plugin is no longer wanted.
Do not add bundled extension paths to plugins.load.paths. Bundled plugins are discovered from the installed WednesdayAI package, and a valid source: "bundled" record is enough for validation and config-schema checks.

ACP startup identity reconcile fails at startup

Use this when gateway logs include startup identity reconcile failed for <session-key>: code=ACP_BACKEND_UNAVAILABLE from the acp/manager subsystem, or the summary line acp startup identity reconcile (renderer=...): checked=N resolved=N failed=N with a nonzero failed count. At gateway startup the gateway re-resolves ACP sessions whose harness identity is still pending. If a session’s backend is not healthy yet, the gateway waits up to 30 seconds (polling every 250 ms) for it to report healthy, then retries each deferred session once. Sessions whose backend is still unhealthy when the wait expires count as failed and log this warning. They are not lost: each session re-resolves on its next use.
Look for:
  • acpx runtime backend registered (command: ..., expectedVersion: ..., pluginLocalInstall: ...) followed within about 30 seconds by acpx runtime backend ready.
  • No acpx local binary unavailable or mismatched (...); running plugin-local install line. That line means the plugin-local acpx binary does not match the pin and an npm install ran during startup.
  • Run the command path from the acpx runtime backend registered line with --version; it must print the same version as the expectedVersion on that line, which is the exact acpx version declared in the plugin’s package.json.
Fix options:
  1. If acpx runtime backend ready never appears, read the earlier acpx runtime setup failed: ... or acpx runtime backend probe failed after local install lines for the root cause, fix it (missing or broken acpx binary, adapter download, harness authentication), and restart the gateway.
  2. If the plugin-local binary version is wrong, restart the gateway and let the plugin-local install re-pin it to the declared version, or reinstall the acpx plugin.
  3. Affected sessions need no manual repair: they re-resolve on their next use.

Dashboard control ui connectivity

When dashboard/control UI will not connect, validate URL, auth mode, and secure context assumptions.
Look for:
  • Correct probe URL and dashboard URL.
  • Auth mode/token mismatch between client and gateway.
  • HTTP usage where device identity is required.
Common signatures:
  • device identity required → non-secure context or missing device auth.
  • device nonce required / device nonce mismatch → client is not completing the challenge-based device auth flow (connect.challenge + device.nonce).
  • device signature invalid / device signature expired → client signed the wrong payload (or stale timestamp) for the current handshake.
  • unauthorized / reconnect loop → token/password mismatch.
  • gateway connect failed: → wrong host/port/url target.
Device auth v2 migration check:
If logs show nonce/signature errors, update the connecting client and verify it:
  1. waits for connect.challenge
  2. signs the challenge-bound payload
  3. sends connect.params.device.nonce with the same challenge nonce
Related:

Gateway service not running

Use this when service is installed but process does not stay up.
Look for:
  • Runtime: stopped with exit hints.
  • Service config mismatch (Config (cli) vs Config (service)).
  • Port/listener conflicts.
Common signatures:
  • Gateway start blocked: set gateway.mode=local → local gateway mode is not enabled. Fix: set gateway.mode="local" in your config (or run openclaw configure). If you are running OpenClaw via Podman using the dedicated openclaw user, the config lives at ~openclaw/.openclaw/openclaw.json.
  • refusing to bind gateway ... without auth → non-loopback bind without token/password.
  • another gateway instance is already listening / EADDRINUSE → port conflict.
Related:

Gateway restart refused on invalid config

Symptom: openclaw gateway restart exits non-zero with Config invalid, names the config file and the exact issue path, and the gateway keeps running. This is the pre-flight restart gate working as intended: the gateway never stops when the config it would boot into is already known-invalid. Before this gate, a bad config edit turned a restart into a crash loop under systemd or launchd supervision. Recovery:
  1. Read the issue path in the refusal output, for example plugins.entries.<id>.config: invalid config: must NOT have additional properties.
  2. Repair with openclaw doctor --fix, or fix the named key manually.
  3. Verify with openclaw config validate (prints Config valid).
  4. Retry openclaw gateway restart.
Notes:
  • gateway stop and shutdown signals (SIGTERM, SIGINT) are never gated; SIGUSR1 watchdog restarts remain gated.
  • Watchdog restarts from config reload (gateway.reload.mode="restart" or "hybrid") are gated the same way: an invalid edit is skipped, the gateway keeps running the previous valid config, and the refusal warning appears in the gateway log. The warning names the config file, the exact issue path, and the same repair commands (openclaw doctor --fix, openclaw config validate).
  • A first start of a stopped gateway is not gated; invalid config there surfaces as a normal startup error.
Related:

Channel connected messages not flowing

If channel state is connected but message flow is dead, focus on policy, permissions, and channel specific delivery rules.
Look for:
  • DM policy (pairing, allowlist, open, disabled).
  • Group allowlist and mention requirements.
  • Missing channel API permissions/scopes.
Common signatures:
  • mention required → message ignored by group mention policy.
  • pairing / pending approval traces → sender is not approved.
  • missing_scope, not_in_channel, Forbidden, 401/403 → channel auth/permissions issue.
Related:

Telegram fleet polling or webhook conflicts

Use this when Telegram logs repeat messages like Network request for 'getUpdates' failed!, account health shows conflicted, or webhook mode receives requests but does not dispatch updates.
Look for:
  • Per-account mode, state, owner, retryAt, poll.lastPollAt, and webhook.lastRequestOutcome in the Channels UI.
  • conflicted state or telegram.poll.conflict diagnostics, which usually means two local accounts, two gateway processes, or a webhook and polling process are trying to own the same bot token.
  • Repeated getUpdates network errors, which usually point to DNS, IPv6, TLS, proxy, firewall, or host egress problems to api.telegram.org.
  • A Polling stall detected log line, which means the built-in watchdog force-restarted a wedged getUpdates cycle; frequent stalls point to a stuck socket or proxy between the gateway and api.telegram.org.
  • Webhook rejected, malformed, timeout, or non-2xx outcomes, which point to a bad public URL, wrong secret token, reverse-proxy mismatch, or a slow webhook handler.
Fix options:
  1. Give each enabled channels.telegram.accounts.<id> entry its own botToken or tokenFile; TELEGRAM_BOT_TOKEN is only a fallback for the default account.
  2. Ensure one bot token has exactly one owner: one configured account, one gateway process, and either polling or webhook mode.
  3. If polling, let WednesdayAI clear webhook ownership non-destructively at startup; if a webhook owner reappears later, polling re-runs that cleanup automatically on 409 instead of conflict-looping. Treat any manual drop_pending_updates operation as destructive unless you intentionally want to discard pending Telegram updates.
  4. If direct Bot API egress is unstable, configure channels.telegram.proxy or the Telegram network settings described in the Telegram channel guide.
Use the Channels UI for transport health and Sessions active runs for execution state. OTel and Langfuse are useful for fleet-wide trends, retry/flood-wait visibility, and correlating Telegram ingress with model and tool spans. Related:

Cron and heartbeat delivery

If cron or heartbeat did not run or did not deliver, verify scheduler state first, then delivery target.
Look for:
  • Cron enabled and next wake present.
  • Job run history status (ok, skipped, error).
  • Heartbeat skip reasons (quiet-hours, requests-in-flight, alerts-disabled).
Common signatures:
  • cron: scheduler disabled; jobs will not run automatically → cron disabled.
  • cron: timer tick failed → scheduler tick failed; check file/log/runtime errors.
  • main job requires payload.kind="systemEvent" → stored job type and payload are incompatible. Recreate or edit it as a Timeline message, Check first, then nudge, or isolated Assistant task job.
  • wake-gate-empty → a wake-gate cron check found no work, so the skipped run is expected.
  • heartbeat skipped with reason=quiet-hours → outside active hours window.
  • heartbeat: unknown accountId → invalid account id for heartbeat delivery target.
  • heartbeat skipped with reason=dm-blocked → heartbeat target resolved to a DM-style destination while agents.defaults.heartbeat.directPolicy (or per-agent override) is set to block.
Related:

Session run history is empty or stale

Main-session turns (chat, cron, heartbeat) are recorded as session-run rows in the configured session storage backend, powering run history and crash recovery. Recording is best effort: a failed store write never fails the turn itself, so a broken store stays invisible unless you look for the warnings. If sessions.runs.list returns no rows or stale timestamps:
Look for:
  • session-run store write failed warns from subsystem sessions/session-runs/recorder. The warn is rate limited to one per minute per operation and carries a suppressed count of failures held back in that window, so a persistent failure shows as a steady drip instead of log flooding.
  • Schema drift after an upgrade. With session.storage.backend: "postgres" or "sqlite" the gateway adds any missing session tables and columns automatically on boot, so a gateway restart repairs drift without manual SQL.
If warnings appeared before an upgrade and stopped after a restart, schema drift is one possible cause — the restarted gateway converged it. The same warning also fires for generic store failures, so if warnings persist after a healthy restart, capture the error field from one warning line and check database connectivity, credentials, and permissions. Related:

Node paired tool fails

If a node is paired but tools fail, isolate foreground, permission, and approval state.
Look for:
  • Node online with expected capabilities.
  • OS permission grants for camera/mic/location/screen.
  • Exec approvals and allowlist state.
Common signatures:
  • NODE_BACKGROUND_UNAVAILABLE → node app must be in foreground.
  • *_PERMISSION_REQUIRED / LOCATION_PERMISSION_REQUIRED → missing OS permission.
  • SYSTEM_RUN_DENIED: approval required → exec approval pending.
  • SYSTEM_RUN_DENIED: allowlist miss → command blocked by allowlist.
Related:

Browser tool fails

Use this when browser tool actions fail even though the gateway itself is healthy.
Look for:
  • Valid browser executable path.
  • CDP profile reachability.
  • Extension relay tab attachment for profile="chrome".
Common signatures:
  • Failed to start Chrome CDP on port → browser process failed to launch.
  • browser.executablePath not found → configured path is invalid.
  • Chrome extension relay is running, but no tab is connected → extension relay not attached.
  • Browser attachOnly is enabled ... not reachable → attach-only profile has no reachable target.
Related:

If you upgraded and something suddenly broke

Most post-upgrade breakage is config drift or stricter defaults now being enforced.

1) Auth and URL override behavior changed

What to check:
  • If gateway.mode=remote, CLI calls may be targeting remote while your local service is fine.
  • Explicit --url calls do not fall back to stored credentials.
Common signatures:
  • gateway connect failed: → wrong URL target.
  • unauthorized → endpoint reachable but wrong auth.

2) Bind and auth guardrails are stricter

What to check:
  • Non-loopback binds (lan, tailnet, custom) need auth configured.
  • Old keys like gateway.token do not replace gateway.auth.token.
Common signatures:
  • refusing to bind gateway ... without auth → bind+auth mismatch.
  • RPC probe: failed while runtime is running → gateway alive but inaccessible with current auth/url.

3) Pairing and device identity state changed

What to check:
  • Pending device approvals for dashboard/nodes.
  • Pending DM pairing approvals after policy or identity changes.
Common signatures:
  • device identity required → device auth not satisfied.
  • pairing required → sender/device must be approved.
If the service config and runtime still disagree after checks, reinstall service metadata from the same profile/state directory:
Related: