Local document models

Local modelson PaddleOCR-VLrunning on your hardware.

Open-weight models matched to your hardware, document family, and data residency — so extraction never has to leave your environment.

8

current local and open-weight options we evaluate for each document family

On-prem

deployment on your GPUs, a private VPC, or dedicated hardware we size with you

Hybrid

local models for volume; OpenAI or Anthropic only when the exception is worth it

Documents stay on-prem

pages never have to leave your network for extraction

Cost stops tracking tokens

hardware does the volume work; APIs handle exceptions only

Latest local options

Models we evaluate for on-prem document processing.

We match the model to the document family, volume, language mix, and hardware you already have — then keep a cloud API only for the exceptions.

Default high-volume parser0.9B

PaddleOCR-VL 1.6

PaddlePaddle

Compact vision-language OCR with 100+ languages, strong tables, seals, formulas, and layout. The usual first pick when thousands of pages need to become structured data on a single GPU.

Throughput and dense scans~3B MoE

DeepSeek-OCR v2

DeepSeek

Built for batch PDF-to-markdown at speed. Mixture-of-experts decoding keeps GPU-hour cost down when backlogs are large and pages are visually dense.

Extraction plus reasoning4B–30B+

Qwen3-VL

Alibaba

Reads the page and interprets it: messy invoices, charts, handwriting, and fields that need context. Use the 4B for edge boxes and larger variants when accuracy on hard layouts matters.

End-to-end key extraction4B

Qianfan-OCR

Baidu Qianfan

A single model for layout, parsing, and key-information extraction. Strong when you need vendor, amounts, dates, and line items without a separate pipeline.

Workstation multimodalE4B–31B

Gemma 4

Google DeepMind

Clean-license multimodal family that runs from a small edge variant up to a 31B dense model. A practical general reader for mixed office documents on one workstation GPU.

Long packs and binders109B / 17B active

Llama 4 Scout

Meta

Very large context for multi-page contracts, due-diligence rooms, and document sets that do not fit a typical window. Use when the job is the whole file, not a single page.

EU self-hosted OCRCompact container

Mistral OCR 4

Mistral

Enterprise OCR with structured blocks, bounding boxes, confidence scores, and 170 languages. Deploys in a single container when data residency and a European vendor matter.

Lightweight structured JSON258M

Granite-Docling

IBM

Tiny Docling-native model for structured document JSON. Fits constrained hardware and pairs well with a larger VLM only on the pages that need it.

How we choose

One document-processing stack. The model changes with the job.

Specialize the model on pages, not chat

When the job is turning thousands of similar documents into CSV or JSON, a compact OCR or document VLM beats a general cloud API on cost and residency. We start with PaddleOCR-VL, DeepSeek-OCR, or Granite-Docling and measure field accuracy on your samples.

  • Pick a parser sized for your GPU, not a chat model
  • Lock schemas for vendor, amounts, VAT, and line items
  • Batch overnight so hardware stays busy and cheap

Ready to run document processing locally?

Bring a sample pack. We will recommend the model mix, the hardware, and the handoff into CSV, JSON, or your API.

Book a discovery call