Skip to content

ADR 0006: Scanned pages are recognized by Tesseract 5 inside the inference service

Status: Accepted, 2026-08-30

Executed closing packages often arrive as scans: a PDF whose pages are images with no text layer. The first release had an ocr pipeline stage that answered “not available”, so a scan reached review with no fields, no execution checks and no provenance. Production needs a scan to go through the same classification, extraction, execution-zone and provenance path as a born-digital PDF, on CPU, inside the bank’s network, with no document leaving the container.

Open-source engines weighed, all CPU-capable and permissively licensed:

Engine Verdict
Tesseract 5 (LSTM): Apache 2.0, C++, packaged by every Linux distribution Best fit. Printed English at 200 to 300 dpi is its strongest case; the tsv output gives every word with a box, a confidence and its block, paragraph and line position, which is exactly the shape the layout model needs; the engine and its language data are OS packages that receive security updates with the base image.
PaddleOCR / RapidOCR (PP-OCRv4): Apache 2.0, Python or ONNX Stronger on photographs, rotated and low-quality text; heavier runtime (Paddle inference or ONNX Runtime plus model files), line-level rather than word-level boxes, and the .NET path is either a Python sidecar or hand-written pre- and post-processing. Kept as the natural upgrade if scan quality in the field demands it; the IOcrEngine boundary is where it would drop in.
EasyOCR, docTR Python-only; a second runtime and a second container for no accuracy gain on printed forms.
Surya GPL-3 weights with commercial restrictions; not usable in a commercial on-prem product.

Two further choices:

  • Bindings vs. process. The Tesseract .NET bindings (Tesseract, TesseractOCR) P/Invoke libtesseract and libleptonica under library names that differ from the distribution packages, which is a recurring source of container-only failures. Running tesseract as a process per page costs a few milliseconds of spawn time against seconds of recognition, isolates the engine’s memory and any crash to one page, and lets OMP_THREAD_LIMIT cap CPU use.
  • Rasterizing vs. extracting the embedded image. Scanned PDFs embed JPEG, CCITT G4, JBIG2 or Flate images at any placement. Rendering the page with PDFium (PDFtoImage, already the api’s rasterizer, see ADR 0004) yields one clean bitmap regardless of encoding, gives the pixels needed to detect underlines and ink for execution checks, and maps 1:1 onto the page’s coordinate space. PDFium has no musl build, so the inference image moves to a glibc base.
  • Bookend.Inference.Service gains an Ocr module: PageRasterizer (PDFium, 300 dpi by default), then TesseractOcrEngine (tesseract page.png stdout --dpi N --psm 3 -l eng tsv), then TesseractTsv and OcrLayoutBuilder, which produces the same PageLayout the native path produces (words in points, baseline lines, reading-order blocks from Tesseract’s structure, underlines detected from the raster). LayoutProvider is the engine’s single source of layouts: native text where it exists, OCR where it does not, cached by content hash so the pipeline’s calls recognize a package once.
  • Each recognized page is cached on its own, and pages are recognized in parallel behind an admission gate (Ocr__MaxConcurrentPages, default processors divided by Ocr__ThreadLimit, at most 8). A call the api abandons resumes at the first unrecognized page, and many documents at once queue instead of thrashing the host.
  • Each Tesseract process gets one OpenMP thread (Ocr__ThreadLimit 1, OMP_THREAD_LIMIT=1, OMP_WAIT_POLICY=PASSIVE). Tesseract’s OpenMP workers busy-wait, and on cores shared with the rest of the stack, two threads per process did not finish a page in two minutes where one thread takes about four seconds (measured on a 2-vCPU host before the defaults changed).
  • The engine contract gains OcrAsync / POST /v1/ocr and PageInfo.OcrApplied / OcrConfidence; /healthz reports the OCR engine and version. On a recognized page, a layout-only value is reported with method ocr and every confidence is scaled by the page’s mean word confidence. Execution zones are decided by ink above the underline (SignatureLocator.InkFillThreshold) because a scan carries no typefaces. The bold-anchor requirement of the sentence rules applies to native pages only.
  • The api’s ocr stage calls the engine and records document_pages.ocr_applied and the mean confidence in the stage detail; extract and locate proceed for scans whose pages were recognized and skip, with a clear message, for scans on an install whose inference service has no OCR engine. The api’s envelope for one inference call is Inference__TimeoutSeconds (default 600), because a whole scanned document is recognized in one call.
  • The BK-DEMO-004 golden package is the clean package as 200-dpi JPEG scans; it is the OCR fixture for the unit, contract and end-to-end test suites.
  • The inference image is mcr.microsoft.com/dotnet/aspnet:10.0 with tesseract-ocr, tesseract-ocr-eng and libfontconfig1. Languages, DPI, page-segmentation mode, the per-page timeout and concurrency are configuration (Ocr__*), never code.
  • Scanned executed copies are validated like native ones; values carry provenance boxes on the scan itself.
  • The inference image grows by about 130 MB (glibc base, PDFium, Tesseract and its English data). Acceptable on-prem.
  • Recognition is a few seconds per page on CPU; the pipeline’s per-stage retry and the OCR timeout bound it. A per-page timeout answers 503 with a reason.
  • OCR confidence is visible on every field (scaled confidence, ocr method) and in the stage detail, so a poor scan routes values to a person through the existing review cap rather than being trusted silently.
  • A second engine is a new IOcrEngine; nothing downstream of OcrLayoutBuilder knows which engine ran.