Troubleshooting Center
This page is generated from a structured incident catalog. Every issue follows “Symptom → Cause → Action → Verification”; do not skip verification.
Start from the symptom. Change one condition at a time and retain startup logs, Trace, and generation details.
Quick index
- Console cannot connect to the Runtime
- Console returns 401 or 403
- Endpoint remains offline
- Messages or notices are missing after refresh
- Configuration is rejected at startup
- Agent has no available model or provider
- Tool is awaiting approval or denied by policy
- A chat, channel, or repository does not route to a Workroom
- Container restarts or health checks fail
Console cannot connect to the Runtime
Symptom
Console remains offline or reconnecting and page data stops updating.
Cause
- The Runtime is stopped, or Console uses a different host, port, or base path.
- The reverse proxy does not forward SSE or buffers the event stream.
Action
- Run
npx zhin doctor, then inspect the HTTP listen address in startup logs. - Ensure the proxy disables buffering and keeps
/api/eventsconnections open.
Verification
curl -i http://127.0.0.1:8086/pub/healthshould succeed, and the Console header should return to Connected.
Console returns 401 or 403
Symptom
Requests remain rejected after login, or write actions report insufficient permissions.
Cause
- The Console token differs from the active generation's
http.token. - The session is read-only Demo mode, or its token lacks the required scope.
Action
- Reload
HTTP_TOKENfrom the deployment environment rather than browser history or stale config. - After changing the production token, publish a new generation and sign in again.
Verification
- Verify the same token with
curl -H "Authorization: Bearer $HTTP_TOKEN" http://127.0.0.1:8086/api/system/stats.
Endpoint remains offline
Symptom
The adapter is installed, but its Endpoint is offline and cannot receive or send messages.
Cause
- The instance configuration fails its plugin Schema or credentials are empty.
- The platform is unreachable, the webhook URL is wrong, or the account is rejected.
Action
- Inspect the latest error in Endpoint details, then compare it with the generated configuration fields.
- Revalidate credentials, callback URLs, and network egress against the platform guide.
Verification
- After reload, the Endpoint should become online; send a real direct-message probe to verify both directions.
Messages or notices are missing after refresh
Symptom
Live messages appear, but history is incomplete after refresh, reconnect, or SSE recovery.
Cause
- The selected Endpoint or Channel differs from the message's interaction space.
- The server event journal has a gap and the client must rebuild from authoritative HTTP APIs.
Action
- Reselect the target Endpoint and Channel and wait for recovery to finish.
- Confirm the database persistence directory is writable and the proxy does not cache history APIs.
Verification
- Send a uniquely worded message, refresh, and restart the Runtime; the HTTP history API should restore it.
Configuration is rejected at startup
Symptom
Startup reports Invalid Plugin config, an unknown top-level field, or environment expansion failure.
Cause
- A field name, type, or enum value differs from the installed version's Schema.
${VAR}is unset and expands to an empty value that fails validation.
Action
- Run
npx zhin doctorand fix the field at the reported path. - Check the current source and Schema in the generated configuration fields.
Verification
npx zhin doctorshould pass, followed bynpx zhin runtime startwithout validation errors.
Tool is awaiting approval or denied by policy
Symptom
An Agent turn stops at a tool step marked pending approval, denied, or cancelled.
Cause
- The working directory or tool is outside the active security policy.
- The approval was not handled, or a cancellation signal already ended the turn.
Action
- Inspect cwd, security policy, and approval details in Agent Studio; approve only understood side effects.
- After cancellation, start a new turn rather than replaying a tool call that may have produced side effects.
Verification
- Run a read-only probe; the tool and turn terminal states should agree and the approval record should be traceable.
A chat, channel, or repository does not route to a Workroom
Symptom
A message falls back to normal chat, or a task does not appear on the expected Workroom board.
Cause
- The Catalog lacks an exact interaction-space binding or still points to an old Agent.
- One Bot may serve multiple Workrooms, but the chat, channel, or repository identity is wrong or conflicting.
Action
- Check Bot, Endpoint, space ID, member roles, and Agent binding in Console Workroom configuration.
- After saving the Catalog, send a new message; historical messages are not reinterpreted.
Verification
- A new message should create a task/run in the target Workroom and show the matched space and Agent in details.
Container restarts or health checks fail
Symptom
Compose or Kubernetes reports unhealthy, CrashLoopBackOff, or persistence permission errors.
Cause
- The node user cannot write the
.zhinordatamount. - A required Secret, project config, or image tag was not deployed correctly.
Action
- Inspect the first error from
docker compose logs zhinorkubectl logs deploy/zhin. - Follow Production Deployment to verify ownership, Secrets, and immutable image tags.
Verification
docker compose psorkubectl rollout status deploy/zhinshould remain successful and the health endpoint should pass.