PaddleOCR-VL 1.6
PaddlePaddle
Compact vision-language OCR with 100+ languages, strong tables, seals, formulas, and layout. The usual first pick when thousands of pages need to become structured data on a single GPU.
Open-weight models matched to your hardware, document family, and data residency — so extraction never has to leave your environment.
current local and open-weight options we evaluate for each document family
deployment on your GPUs, a private VPC, or dedicated hardware we size with you
local models for volume; OpenAI or Anthropic only when the exception is worth it
Documents stay on-prem
pages never have to leave your network for extraction
Cost stops tracking tokens
hardware does the volume work; APIs handle exceptions only
Latest local options
We match the model to the document family, volume, language mix, and hardware you already have — then keep a cloud API only for the exceptions.
PaddlePaddle
Compact vision-language OCR with 100+ languages, strong tables, seals, formulas, and layout. The usual first pick when thousands of pages need to become structured data on a single GPU.
DeepSeek
Built for batch PDF-to-markdown at speed. Mixture-of-experts decoding keeps GPU-hour cost down when backlogs are large and pages are visually dense.
Alibaba
Reads the page and interprets it: messy invoices, charts, handwriting, and fields that need context. Use the 4B for edge boxes and larger variants when accuracy on hard layouts matters.
Baidu Qianfan
A single model for layout, parsing, and key-information extraction. Strong when you need vendor, amounts, dates, and line items without a separate pipeline.
Google DeepMind
Clean-license multimodal family that runs from a small edge variant up to a 31B dense model. A practical general reader for mixed office documents on one workstation GPU.
Meta
Very large context for multi-page contracts, due-diligence rooms, and document sets that do not fit a typical window. Use when the job is the whole file, not a single page.
Mistral
Enterprise OCR with structured blocks, bounding boxes, confidence scores, and 170 languages. Deploys in a single container when data residency and a European vendor matter.
IBM
Tiny Docling-native model for structured document JSON. Fits constrained hardware and pairs well with a larger VLM only on the pages that need it.
How we choose
When the job is turning thousands of similar documents into CSV or JSON, a compact OCR or document VLM beats a general cloud API on cost and residency. We start with PaddleOCR-VL, DeepSeek-OCR, or Granite-Docling and measure field accuracy on your samples.
From the blog
Bring a sample pack. We will recommend the model mix, the hardware, and the handoff into CSV, JSON, or your API.
Book a discovery call