Skip to content

Release notes

All notable changes to the Bookend Platform are recorded here. The file ships inside the API image and is served at GET /v1/releases. The format follows Keep a Changelog; versions follow SemVer.

  • Header-box values on scans: Tesseract returns a bordered header box cell by cell, so no OCR line carried the whole label row and the loan number could be missed on a scanned Note. The header-cell reader now assembles the label and value rows from every line in their band and cuts the cell from the union, without picking up the sentence printed under the box.
  • SMTP relays on implicit TLS (port 465) work: the wizard’s ssl mode opened a STARTTLS connection and hung for two minutes. The sender now maps ssl / starttls / none to the matching connection type and times out in 15 seconds. All three modes, and the wrong-mode case failing fast, are covered by integration tests against a TLS-enabled relay.
  • Evidence chain writers are serialized per loan inside the api, and the gate is held until the enclosing transaction completes, so two stages finishing together no longer collide on UNIQUE (loan_id, seq) (each collision was an ERROR in the bank’s database log; the retry remains as the backstop for a second api instance). Every unique-constraint violation now surfaces as a 409 Conflict instead of a 500 (for example, two operators saving a new profile or map version at once).
  • Scanned packages on small hosts: the inference service caches every recognized page, recognizes a document’s pages in parallel behind an admission gate sized for the host (Ocr__MaxConcurrentPages), and so resumes where an abandoned call stopped. The api’s per-call inference envelope is Inference__TimeoutSeconds (default 600 s, was 120 s). Before this, scanned splits could time out on a 2-vCPU host.
  • Tesseract runs with one OpenMP thread per process (Ocr__ThreadLimit 1, OMP_WAIT_POLICY=PASSIVE; was 2). Its workers busy-wait, so on cores shared with the rest of the stack two threads per process did not finish a page in two minutes where one thread takes four seconds. A per-page OCR timeout now answers 503 with a reason instead of an empty 500.
  • Point-and-name document mapping (Settings → Documents): upload one example, drag a box around each value and say which field it is. Bookend reads the box back on the spot with the value it would extract and the label it will follow, draws every value on the page when you check the template, and saves it as data. On a loan, administrators can fix where a value comes from by drawing on the source page; the box becomes the rule for that document type and the loan can be reprocessed. Under the hood: a region rule type anchored on the nearest label (survives scan offsets and re-flow; reduced certainty without its anchor), POST /v1/read-region on the inference service, and samples, read-region, prove and corrections endpoints under /v1/extraction-templates. The regex sentence rules and raw JSON remain under the technical details.
  • The metering heartbeat carries document_page_count (pages of every document received in the period) and license_key, and System → Metering lists pages next to closed loans. When the receiving endpoint answers with a license verdict, the install records it in its delivery history.
  • The api image includes a hash-password utility (dotnet Bookend.Api.dll hash-password, password on stdin) that prints an Argon2id hash, so a provisioning script can seed the first administrator.
  • OCR for scanned executed copies: pages without a text layer are rendered with PDFium and recognized by Tesseract 5 (LSTM) inside the inference container. Recognized words feed the same classification, extraction, execution-zone and provenance path as native text; fields report the ocr method with confidence scaled by recognition quality, and execution zones on scans are decided by ink above the underline. POST /v1/ocr on the inference service, the ocr pipeline stage records recognized pages and mean confidence, and /healthz reports the OCR engine.
  • BK-DEMO-004: the clean package delivered as 200-dpi JPEG scans, the OCR golden fixture, covered by the unit, contract and end-to-end test suites.
  • Production license validation: installs validate licenses against Bookend’s offline release key (Licensing__PublicKeyPath / Licensing__PublicKeyPem); the embedded development key is for demos only, and LicenseService.KeySource reports which key is active. The platform sends metering.api_key as X-Bookend-Metering-Key when configured, and an air-gapped install’s signed quarterly usage report is verifiable against its JWKS.
  • Signed release bundles: every release ships as one bundle with the container images, the DDL scripts, the runbooks and guides, the OpenAPI document, the SR 11-7 validation pack, SHA256SUMS (minisign-signed for official releases) and a VERIFY.md that explains how to check both.
  • Template studio: extraction extends to a bank’s own document families as data, never code. An admin uploads one sample, the configured AI model (any OpenAI-compatible endpoint, and a local model works; ai.endpoint / ai.model / ai.api_key) drafts a declarative template, the engine proves it by classifying and extracting that sample, and the reviewed overlay is saved to extraction.templates. Overlay rules run before the built-ins on every pipeline stage, and production extraction stays deterministic. /v1/extraction-templates (get, save, suggest) and the admin studio page. A document family the engine has no code for (an alternate boarding worksheet) classifies at 0.98 and extracts every field from a data overlay alone.
  • Extraction accuracy is measured, not asserted: a randomized corpus of closing packages with ground truth by construction (a fraction as scans at varied scanner settings) runs through the engine end to end, including the rules over the engine’s own output, and the report gives accuracy by field kind with pass thresholds. First run: 60 loans, 3,378 field reads, 100% exact on native input.
  • LAR profiles now parse CSV (header-addressed columns, delimiter sniffed, RFC 4180 quoting, $.Column[*] across data rows), XML (element paths, [*] / [n], trailing @attr) and DOCX (label lookups over tables and “Label:” paragraphs; the sample travels as base64) in addition to JSON. A PDF approval record is a package document: the extraction pipeline reads it, and the profile surface says so.
  • Retention enforcement: a nightly sweep (04:00 UTC) deletes the stored PDF bytes and derived renders of loans sealed longer than documents.retention_days ago. Database rows (metadata, SHA-256 hashes, extracted values, findings and the evidence chain) stay, so sealed packets still verify and still name every document by hash; only the bytes go. The store is content-addressed, so a file shared with a loan still inside policy is kept. The sweep is idempotent, and its outcome lands in the read-only documents.last_retention setting and in Diagnostics like any other job.
  • The dashboard is home: first in the navigation and the landing page after sign-in (the loan queue is one click away).
  • One Settings area (/settings): Rules, Documents (extraction templates), Approval record (LAR profiles), Core boarding (core field maps) and Other options (every remaining setting) are tabs in a single place instead of five navigation entries. The configuration pages read in plain language (“upload one example”, “paste a sample approval”, “what a real loan would board as”), with the JSON editors behind a “Show the technical details” toggle. Old addresses redirect.
  • Rules workbench (/rules, now Settings → Rules): the rule catalog is a first-class screen. Shipped rules can be switched off or on per bank policy; a disabled rule is skipped from the next reconciliation run onward, and the skip is recorded on the run and in the evidence chain, never silent. Administrators author bank rules in one sentence: pick a canonical field, pick the check (every source must agree, documents must match the approval, or must be present, optionally on one document type) and pick the weight (exception, warning or info), with no code and no JSON. The id is derived from the title (CUSTOM-…), and the rule is evaluated by the same deterministic engine as the shipped catalog, with the bank’s tolerances applied. GET /v1/rules now returns policy state; PUT /v1/rules/{id}/enabled and POST / PUT / DELETE /v1/rules/custom manage it (admin, audited). The generic evaluator carries the same 100% branch-coverage gate as every shipped rule.
  • Users & access page (/users): invite (72-hour set-password link) or create users, edit roles inline, deactivate and reactivate (sessions revoked), and the MFA lost-device reset, plus API keys: issue (plaintext shown once) with scopes, list, revoke. Previously all of this was API-only.
  • Work assignment in the app: the loan header gains Assign to me, Take over and Release. Assignment was already recorded in the evidence chain and drove the queue’s “My work” view, but had no UI. Settings gains the SMTP, core and inference test buttons (re-runnable after setup) and a reopen setup wizard action.
  • System → Settings page: every runtime setting editable in the app, by namespace: rule tolerances (rules.money_tolerance_cents, rate_tolerance, date_tolerance_days, name_normalization), reviewer reason codes and justification policy, sign-in policy (MFA, SSO), SMTP, documents, pipeline, metering, and the template-studio model (ai.*). Type-aware inputs, secrets masked (a blank save keeps the stored value), read-only catalog rows displayed as values, restart-required badges, and only edited keys submitted. Reconciliation rules themselves remain versioned code shipped with each release (ADR 0003); this page is where a bank tunes how they judge.
  • OIDC federation: “Continue with ” on the sign-in screen hands authentication to the bank’s OpenID Connect identity provider (Entra ID, Okta, ADFS, anything OpenID-certified), configured entirely in settings (auth.oidc_authority / client_id / client_secret / provider_name). The api is a confidential client: it seals the round trip’s state (nonce, redirect URI, 10-minute expiry) under the install’s master key, exchanges the code server-side, validates the id_token against the IdP’s JWKS (signature, issuer, audience, lifetime, nonce), and maps the proven email to an existing, active Bookend user (no just-in-time provisioning) before issuing the same token pair as password sign-in, so authorization, refresh rotation and audit are identical for every method. Contract-tested end to end against a stub IdP, including tampered state, wrong nonce and the unprovisioned-identity refusal.
  • MFA with authenticator apps: RFC 6238 TOTP on standard .NET cryptography, pinned to the RFC 4226/6238 published vectors. Two-step enrollment on the new account page (setup key and otpauth:// URI shown once, then a confirming code), an AES-GCM-protected secret, single-use codes with one step of clock tolerance, and enforcement per auth.login_totp_enabled: a confirmed user’s password sign-in demands the code (401 totp_required switches the login form), while unenrolled users are never locked out. Self-service disable needs a fresh code; DELETE /v1/users/{id}/totp is the admin lost-device reset (revokes sessions). Every step is an auth event, and the whole lifecycle is covered by unit vectors, an HTTP contract flow and an end-to-end journey that plays the authenticator app.
  • Resilience posture and restore drill: the resilience runbook states what each component tolerates (stateless api, persisted jobs that resume after restart, a content-addressed immutable document store, database high availability delegated to the bank’s platform), the RPO and RTO objectives, and a quarterly restore drill on a scratch stack. The first drill (2026-08-31) took 4 seconds to back up and about 30 seconds to restore at demo scale, and both sealed evidence chains re-verified intact: true from the restored bytes.
  • The inference image moved to the glibc-based .NET base image (ADR 0006) so it can rasterize pages with PDFium and carry Tesseract from the distribution packages.
  • POST /v1/lar-profiles rejected every request with “A mapping object is required”: the input validator rebuilds requests from declared fields only and the mapping element was never declared, so it was silently dropped. Found by the new CSV profile contract test.
  • First release: intake (upload, watch folder, LAR ingest), deterministic inference pipeline with page-level provenance, 17-rule reconciliation engine with 100% branch coverage, review workstation (accept, override, escalate, approval checklist), dashboard from recorded events.
  • Two-phase core boarding (stage, checker approval, commit) with file-export and Jack Henry jXchange adapters, versioned core field maps, wire staging with maker-checker and PDF or file-drop output.
  • Tamper-evident evidence chain with canonical-JSON hashing, nightly verification sweep, sealed loans, JSON and PDF evidence packets.
  • MCP server at /mcp (list_loans, get_loan, get_findings, get_boarding_preview, submit_package, get_evidence_summary, export_evidence, get_queue_stats) using the same bearer tokens and policies as REST; every call audited.
  • Metering heartbeat (daily, {install_id, version, period, closed_loan_count, health}), air-gapped signed quarterly usage report, license expiry makes intake read-only, diagnostics, release notes and audit endpoints.
  • The scheduled metering heartbeat waits until the install is activated (a fresh install reported an unlicensed-… id once); the operator’s manual run is unaffected.
  • Boarding no longer depends on crypto.randomUUID(), which browsers omit on plain-HTTP origins other than localhost; the idempotency key falls back to getRandomValues.
  • RS256 JWT access tokens with rotating refresh tokens, Argon2id password hashing, per-principal rate limiting, security headers, secrets encrypted at rest with the master key.