Skip to content

Validation pack (SR 11-7)

Under SR 11-7 / OCC 2011-12 guidance, proportionate to a community bank, the extraction and classification engine is an item in the bank’s model inventory. Bookend keeps the engine’s scope narrow and gives the bank what it needs to validate it independently.

  • Does: document classification, field location and extraction (including OCR on scanned pages), and signature, initials, date and notary zone detection.
  • Does not: decide whether terms agree (deterministic, versioned rules do that), approve anything (people do), or board and fund (people do, through the bank’s controls).

The engine is deterministic: the same document produces the same values on every run, and it runs on CPU inside the bank’s network.

Each release bundle has a validation-pack/ folder, covered by the bundle’s SHA256SUMS and signature like everything else in it:

Item Content
golden-packages/ The reference closing packages (a clean package, packages with money and execution exceptions, and the clean package as 200-dpi scans), each with its approval record, plus manifest.json: the expected document type, field values, execution zones and findings for every document. Synthetic data only.
rules/ One page per reconciliation rule: what it checks, which fields and documents it reads, its severity and the tolerances it applies. The same pages are served in the product and on this site under Rules.
accuracy-report.md When included: the extraction accuracy report from a corpus run, by field kind and by native versus scanned input, against pass thresholds.

The bundle also carries CHANGELOG.md (what changed in the release, including extraction and rule changes) and the OpenAPI document.

  • Reproduce the golden results. Load a golden package through Loans → New loan (or the API) and compare the extracted values, execution checks and findings with manifest.json. The install runbook uses the clean package as its acceptance check.
  • Measure accuracy on your own documents. Run representative packages, including your own forms mapped in the template studio, and review the values against the documents. Every value shows its source box on the page.
  • Monitor in production. Override rates by rule, reprocess counts and confidence distributions all derive from the evidence events and the dashboard.
  • Every extracted value records its method (native, ocr, model or dual), its confidence and its page and bounding box; the extraction event in the evidence chain records the engine version and whether the two extraction passes agreed. Disagreement caps a value’s confidence at 0.5, and low-confidence values route to a person.
  • Every reconciliation run records the ruleset version, the thresholds and any rules the bank has switched off.
  • Every decision records the person, the reason code and the justification.
  • The evidence packet is exportable and verifiable offline.

The LAR profile, the core field map and the bank’s taught document templates are configuration, not training: a taught field is a stored box and its anchor label, proved on the sample before it is saved and versioned in the settings audit. Nothing learned inside a deployment leaves it.