Skip to content

Incidents

First commands for any incident (with the production override from the Installation Guide; add -f deploy/compose.external-db.yml for an external database):

Terminal window
DC="docker compose -f deploy/compose.yml -f deploy/compose.production.yml"
$DC ps -a # which service is unhealthy or exited
$DC logs --since 30m api # structured logs; event ids below
curl -s https://<host>/api/readyz # database and schema readiness

GET /v1/diagnostics (administrators, also System → Diagnostics) is the support bundle: version, runtime, database and schema, license, metering, inference probe, core provider, last evidence sweep, health entries and every scheduled job with its next fire time. It contains no loan data and can be sent to support as is.

Symptom Meaning Do
EVIDENCE CHAIN BROKEN for loan ... (event 5120), or System → Diagnostics shows broken loans under the evidence sweep A stored evidence payload no longer hashes to its link: the row was changed outside the application (the application principal cannot update or delete evidence_events). Treat as a security event. Do not export or seal the loan. Preserve the database (back up now), compare with the last good backup, and involve the bank’s security team; GET /v1/loans/{ref}/evidence/verify names the first broken sequence number.
Pipeline stuck in processing; GET /v1/loans/{ref}/pipeline shows a stage retrying (StageWillRetry, event 5101) The inference container is down or slow; stages back off exponentially up to pipeline.max_attempts. $DC ps inference, $DC logs inference, $DC restart inference. Stages resume automatically; once fixed, use Reprocess on the loan if attempts were exhausted. On a small host, scanned packages can need a longer Inference__TimeoutSeconds or Ocr__TimeoutSeconds (Installation Guide, OCR tuning).
api never becomes healthy after an upgrade; /readyz reports schema unhealthy The schema_versions ledger lacks a script this api build requires. Check $DC logs migrator: exit 4 is drift (an applied script changed or is missing: restore the matching release’s scripts, never edit an applied one); exit 2 is a missing bootstrap variable. Re-run $DC up -d migrator once fixed (idempotent). Never edit schema_versions by hand.
Users cannot sign in; login_failed rows in GET /v1/audit?source=auth Policy (allowed domains, minimum length), an expired refresh token, or a rotated signing key. Settings → auth policy; a changed auth.signing_key (restart_required) needs an api restart and signs everyone out.
Sign-in links not arriving SMTP relay refused or misconfigured. Send a test from the SMTP settings (POST /v1/settings/test/smtp); check the relay’s logs.
Boarding record stays committing The api stopped mid-commit. The record is resumable: open Boarding and Commit again; the adapter resolves an already-boarded loan by inquiry and never posts twice.
Core connection test fails Endpoint, credentials or institution id wrong, or the file-export path not writable by the app user. Settings → core boarding; $DC exec api ls -ld /data/exports.
WatchFolderStuck (event 5114): a package folder stays in the watch folder The package was handled, but its folder is owned by another user, so the api cannot move it to processed/ or rejected/. It is skipped until the api restarts. $DC exec -u root api chown -R app:app /data/watch, then confirm the loan is in the queue and move or delete the leftover folder before restarting the api, so it is not picked up again. Make the process that drops packages write as a user the api can move.
Banner “No metering heartbeat delivered in the last 7 days” The heartbeat endpoint is unreachable (event 5131) or heartbeats are off. System → Metering & license: Send heartbeat now and read the last error; air-gapped installs generate the signed quarterly usage report instead.
Banner “License expired” and intake is read-only The license key has passed its expiry. Everything already in the system keeps working (review, boarding, evidence); install the renewed key: reopen the wizard (Settings → Other options) and enter it at step 9, then activate again.
429 responses Per-principal rate limit (RateLimiting__PerPrincipalPerMinute, default 300) or the login attempts limit (RateLimiting__LoginAttemptsPerMinute, default 10). Expected under abuse; raise a limit in the api environment only with security sign-off.

Escalate with: the diagnostics JSON, $DC logs --since 1h for the affected service, the loan reference(s), and the GET /v1/audit rows around the time of the incident. Loan documents never need to leave the bank for support.