Resilience and availability
What fails, what happens, and how long recovery takes. This is the statement a bank’s technology committee asks for; every claim in it is either enforced by the architecture or proven by the restore drill below.
Design point
Section titled “Design point”Bookend runs as a single compose stack on one host: one api node, one scheduler, one database. That is a deliberate fit for community-bank closing volumes (tens of loans a day, not thousands): the working set is small, every durable thing is in the database or the document volume, and the recovery unit is “restore the backup”, not “fail over a cluster”. The bank buys availability with its virtualization and database platforms, which it already operates; Bookend’s job is to be restart-safe and restore-safe so those platforms are sufficient.
What each component tolerates
Section titled “What each component tolerates”| Component | State it holds | On crash or restart |
|---|---|---|
api |
none: JWTs are stateless, refresh tokens and settings are rows, uploads are written to the store before the pipeline is scheduled | restart it; in-flight requests fail and the SPA retries; nobody is signed out (tokens live in the browser’s memory and re-issue on refresh) |
Pipeline and jobs (Quartz.NET in api) |
every stage attempt is a durable row in the qrtz_* tables (ADO.NET job store) |
the next start resumes exactly where it stopped; a missed nightly run (evidence sweep 02:00 UTC, heartbeat 03:15 UTC, retention 04:00 UTC) fires once at startup; jobs are marked non-concurrent, so a double start is harmless |
inference |
nothing durable: recognized page layouts are cached in memory per content hash and rebuilt after a restart | restart it; stages that could not reach it back off (retry_base_seconds × 2^(attempt − 1), capped at 10 minutes, up to pipeline.max_attempts) and resume when it returns; Diagnostics shows the inference probe failing meanwhile |
db |
everything relational: loans, values, findings, evidence chain, settings (secrets encrypted under the master key), scheduler state, audit | the bundled Postgres restarts with the stack (restart: unless-stopped); production installs can point DB__CONNECTION at the bank’s HA instance instead (see below) |
| Document volume | the PDF bytes every evidence link hashes; content-addressed by SHA-256, written before the database row that references them | files are immutable once written (deduplicated by hash), so a crash mid-upload leaves at worst an unreferenced file; the volume rides the host’s storage and is in every backup |
app-ui (nginx) |
none | restart it |
| Watch folder | transient: packages move to processed/ or rejected/ after ingest |
rescanned every poll; a package half-copied at crash time is picked up when its files go quiet (documents.watch_stable_seconds) |
| SMTP relay | none Bookend depends on | sign-in links fail visibly and are requested again; nothing else queues mail |
Two failure modes are worth naming because they are designed to be boring:
- Kill the api mid-pipeline and the run resumes from its last recorded stage after restart. Every upgrade exercises this (
down, thenupwith volumes kept). - Kill the api mid-boarding and two-phase boarding holds: the core adapter records the inquiry and import pair, and duplicate resolution by inquiry means a re-run neither double-boards nor loses the loan.
Database HA is the bank’s platform
Section titled “Database HA is the bank’s platform”The bundled Postgres container suits evaluation and small installs. A production install can set DB__CONNECTION to the bank’s managed instance: Postgres with streaming replication or Patroni, or SQL Server with an Availability Group (both dialects ship in db/). Bookend needs one application connection string and holds no state outside that database and the document volume, so the bank’s existing database HA, snapshots and DR replication apply without any Bookend-specific configuration. The same goes for the host: the compose stack is a stock VM workload, so VMware or Hyper-V restart policies and replication cover host loss.
RPO and RTO
Section titled “RPO and RTO”| Objective | Value | Why |
|---|---|---|
| RPO | the backup interval: nightly per backup-restore.md, so at most 24 hours; banks running the database on their own instance inherit its point-in-time recovery and can bring RPO to minutes |
pg_dump plus a volume tar is the floor, not the ceiling |
| RTO | minutes, not hours: the measured drill below restored and re-verified a small stack in under a minute; budget about 15 minutes at production scale (the dump restore dominates) plus the bank’s VM provisioning time if the host itself is lost | restore is a handful of commands from the runbook; no reconfiguration, because settings, keys and job state are all inside the backup |
A restore is only complete when the evidence chain verifies. That check is step 2 of the restore runbook, and the nightly evidence sweep repeats it from then on. A restored stack that cannot prove its chains is an incident, not a recovery.
Restore drill
Section titled “Restore drill”Rehearse a restore quarterly: restore the latest backup into a scratch stack (a throwaway database container and empty volumes; the live stack is untouched) and prove it. Last drill:
| Date | 2026-08-31 |
| Dataset | test stack after a full end-to-end test run: 2 sealed loans, 128- and 134-event chains, 6.7 MB document volume |
| Backup | pg_dump -Fc plus volume tar: 4 s |
| Restore | role plus pg_restore into a fresh Postgres: 6 s; volume untar, chown and api start: 18 s; about 30 s total |
| Proof | /readyz healthy · administrator sign-in (the restored master key decrypts the signing key) · evidence/verify returns intact: true on both sealed loans at full chain length · the evidence PDF renders from the restored bytes |
Drill procedure (scratch stack):
# fresh db container and fresh volumes; ports and names never collide with the live stackdocker compose -f drill-compose.yml up -d drill-db --waitdocker compose -f drill-compose.yml exec -T drill-db psql -U bookend -d postgres \ -c "CREATE ROLE bookend_app LOGIN PASSWORD '<app password>';" # the application principal from the grants scriptdocker compose -f drill-compose.yml exec -T drill-db pg_restore -U bookend -d bookend --no-owner < backup/bookend-<stamp>.dumpdocker run --rm -v drill_documents:/data/documents -v drill_exports:/data/exports -v "$PWD/backup":/backup alpine \ sh -c "tar xzf /backup/volumes-<stamp>.tgz -C / && chown -R 1654:1654 /data"docker compose -f drill-compose.yml up -d drill-api --wait # api image, DB__CONNECTION pointing at drill-db# prove: /readyz, sign in, GET /v1/loans/<ref>/evidence/verify returns intact:true, export the PDF; then `down -v`drill-compose.yml is a file you write for the drill, with two services: postgres:16-alpine (POSTGRES_DB and POSTGRES_USER set to bookend) and the installed bookend/api image with DB__CONNECTION pointed at the drill database, the backed-up BOOKEND_MASTER_KEY and BOOKEND_SKIP_SCRIPTS, and its own empty drill_documents and drill_exports volumes (declare them with those exact name: values so the commands above find them). Record the date and timings here after each drill.
What Bookend does not claim
Section titled “What Bookend does not claim”No automatic failover of the api or scheduler, no multi-node clustering, no zero-RPO replication of its own. If a bank’s volumes or availability requirements outgrow the single-node design, the honest path is scale-up (the stack is small) and the bank’s platform HA underneath, not a distributed-systems story bolted onto a closing-room tool.