ADR 0006: Scanned pages are recognized by Tesseract 5 inside the inference service
Status: Accepted, 2026-08-30
Context
Section titled “Context”Executed closing packages often arrive as scans: a PDF whose pages are images with no text layer. The first release
had an ocr pipeline stage that answered “not available”, so a scan reached review with no fields, no execution
checks and no provenance. Production needs a scan to go through the same classification, extraction, execution-zone
and provenance path as a born-digital PDF, on CPU, inside the bank’s network, with no document leaving the container.
Open-source engines weighed, all CPU-capable and permissively licensed:
| Engine | Verdict |
|---|---|
| Tesseract 5 (LSTM): Apache 2.0, C++, packaged by every Linux distribution | Best fit. Printed English at 200 to 300 dpi is its strongest case; the tsv output gives every word with a box, a confidence and its block, paragraph and line position, which is exactly the shape the layout model needs; the engine and its language data are OS packages that receive security updates with the base image. |
| PaddleOCR / RapidOCR (PP-OCRv4): Apache 2.0, Python or ONNX | Stronger on photographs, rotated and low-quality text; heavier runtime (Paddle inference or ONNX Runtime plus model files), line-level rather than word-level boxes, and the .NET path is either a Python sidecar or hand-written pre- and post-processing. Kept as the natural upgrade if scan quality in the field demands it; the IOcrEngine boundary is where it would drop in. |
| EasyOCR, docTR | Python-only; a second runtime and a second container for no accuracy gain on printed forms. |
| Surya | GPL-3 weights with commercial restrictions; not usable in a commercial on-prem product. |
Two further choices:
- Bindings vs. process. The Tesseract .NET bindings (
Tesseract,TesseractOCR) P/Invokelibtesseractandlibleptonicaunder library names that differ from the distribution packages, which is a recurring source of container-only failures. Runningtesseractas a process per page costs a few milliseconds of spawn time against seconds of recognition, isolates the engine’s memory and any crash to one page, and letsOMP_THREAD_LIMITcap CPU use. - Rasterizing vs. extracting the embedded image. Scanned PDFs embed JPEG, CCITT G4, JBIG2 or Flate images at any placement. Rendering the page with PDFium (PDFtoImage, already the api’s rasterizer, see ADR 0004) yields one clean bitmap regardless of encoding, gives the pixels needed to detect underlines and ink for execution checks, and maps 1:1 onto the page’s coordinate space. PDFium has no musl build, so the inference image moves to a glibc base.
Decision
Section titled “Decision”Bookend.Inference.Servicegains anOcrmodule:PageRasterizer(PDFium, 300 dpi by default), thenTesseractOcrEngine(tesseract page.png stdout --dpi N --psm 3 -l eng tsv), thenTesseractTsvandOcrLayoutBuilder, which produces the samePageLayoutthe native path produces (words in points, baseline lines, reading-order blocks from Tesseract’s structure, underlines detected from the raster).LayoutProvideris the engine’s single source of layouts: native text where it exists, OCR where it does not, cached by content hash so the pipeline’s calls recognize a package once.- Each recognized page is cached on its own, and pages are recognized in parallel behind an admission gate
(
Ocr__MaxConcurrentPages, default processors divided byOcr__ThreadLimit, at most 8). A call the api abandons resumes at the first unrecognized page, and many documents at once queue instead of thrashing the host. - Each Tesseract process gets one OpenMP thread (
Ocr__ThreadLimit1,OMP_THREAD_LIMIT=1,OMP_WAIT_POLICY=PASSIVE). Tesseract’s OpenMP workers busy-wait, and on cores shared with the rest of the stack, two threads per process did not finish a page in two minutes where one thread takes about four seconds (measured on a 2-vCPU host before the defaults changed). - The engine contract gains
OcrAsync/POST /v1/ocrandPageInfo.OcrApplied/OcrConfidence;/healthzreports the OCR engine and version. On a recognized page, a layout-only value is reported with methodocrand every confidence is scaled by the page’s mean word confidence. Execution zones are decided by ink above the underline (SignatureLocator.InkFillThreshold) because a scan carries no typefaces. The bold-anchor requirement of the sentence rules applies to native pages only. - The api’s
ocrstage calls the engine and recordsdocument_pages.ocr_appliedand the mean confidence in the stage detail;extractandlocateproceed for scans whose pages were recognized and skip, with a clear message, for scans on an install whose inference service has no OCR engine. The api’s envelope for one inference call isInference__TimeoutSeconds(default 600), because a whole scanned document is recognized in one call. - The
BK-DEMO-004golden package is the clean package as 200-dpi JPEG scans; it is the OCR fixture for the unit, contract and end-to-end test suites. - The inference image is
mcr.microsoft.com/dotnet/aspnet:10.0withtesseract-ocr,tesseract-ocr-engandlibfontconfig1. Languages, DPI, page-segmentation mode, the per-page timeout and concurrency are configuration (Ocr__*), never code.
Consequences
Section titled “Consequences”- Scanned executed copies are validated like native ones; values carry provenance boxes on the scan itself.
- The inference image grows by about 130 MB (glibc base, PDFium, Tesseract and its English data). Acceptable on-prem.
- Recognition is a few seconds per page on CPU; the pipeline’s per-stage retry and the OCR timeout bound it. A per-page timeout answers 503 with a reason.
- OCR confidence is visible on every field (scaled confidence,
ocrmethod) and in the stage detail, so a poor scan routes values to a person through the existing review cap rather than being trusted silently. - A second engine is a new
IOcrEngine; nothing downstream ofOcrLayoutBuilderknows which engine ran.