Skip to content

ADR 0004: The api image is glibc-based because page rendering needs PDFium

Status: Accepted, 2026-08-29

The review workstation shows every extracted value on its source page with a bounding box. That needs a PDF rasterizer in the api container (GET /v1/documents/{id}/pages/{n}.png). The managed options were weighed:

  • PdfPig: pure managed; reads text and layout (the inference engine uses it) but does not rasterize.
  • PDFtoImage (SkiaSharp + PDFium): rasterizes reliably; PDFium native binaries are published for Windows, macOS and glibc Linux (bblanchon.PDFium.Linux). No musl build exists, and SkiaSharp needs libfontconfig.
  • Docnet.Core: also PDFium, same glibc-only constraint, less maintained.

Before this decision every image was built on mcr.microsoft.com/dotnet/*:10.0-alpine.

The api image builds on mcr.microsoft.com/dotnet/aspnet:10.0 (the default glibc-based .NET 10 tag) with libfontconfig1 and libgssapi-krb5-2 installed, still non-root (USER app), still with the same healthcheck and /data volume ownership. The migrator image stays on alpine; it has no native rendering dependency. (The inference image later moved to the same base for OCR, see ADR 0006.) PdfPageRenderer (Infrastructure) is the only consumer of PDFtoImage; renders are cached beside the original in the document store so a page is rasterized once.

  • The api image grows by roughly 100 MB versus alpine; acceptable for an on-prem appliance.
  • ICU is present in the glibc image by default, so InvariantGlobalization=false (required by SqlClient) needs no extra package there.
  • A future model-based inference runtime that also needs glibc natives can follow the same base image.
  • If a musl PDFium build appears the switch back is a two-line Dockerfile change; nothing in code depends on the distribution.