A self hosted alternative to Nanonets, using Nanonets’ own open weights
TL;DR
Nanonets publish their own open weights: Nanonets-OCR2-3B is on HuggingFace as a fine tune of Qwen2.5-VL-3B, so the honest comparison is their hosted workflow against their own model running on your GPU.
Nanonets price work in blocks, at $0.02 for simple runs, $0.10 for standard AI and $0.30 for complex AI, and their own worked example puts a typical invoice at under $2 end to end.
A reserved L40S costs roughly $701 a month and handles 20,000 two page documents in about thirty hours, so the self hosted cost line barely moves as volume climbs.
Below roughly 2,000 documents a month Nanonets is the better buy, because the annual spend does not cover a build; above roughly 10,000 a month the per document meter dominates the budget.
Ækora is a consultancy and deployment engagement, not a product: no free tier, no trial, no validation interface to sign up for, and no published accuracy comparison against Nanonets.
Questions people ask
Is there a free alternative to Nanonets?
The weights are free, the deployment is not. Nanonets themselves publish Nanonets-OCR2-3B on HuggingFace, and PaddleOCR-VL-1.6, GLM-OCR and Docling are all open source. Downloading any of them costs nothing. What costs money is the GPU that runs them, the key and value extraction against your schema, the review queue, the audit trail and the person who keeps it all running. Treat open weights as removing a per document meter, not as removing cost. If you want something working today with no engineering, the free credits on the Nanonets Starter tier are a faster route.
Nanonets vs Docling: which should I use?
They are not competitors, which is why the question is confusing. Docling is a pipeline and toolkit from IBM Research Zurich, now hosted at the LF AI and Data Foundation under an MIT licence. It handles page layout, reading order, table structure and export to markdown, HTML or JSON, and it plugs in OCR backends including vision language models. Nanonets is a hosted workflow product with a validation interface, integrations and support. You would use Docling together with a model such as PaddleOCR-VL or Nanonets-OCR2-3B, and you would use it instead of Nanonets only if you are prepared to build the workflow yourself.
How much does Nanonets cost?
Their published pricing gives $50 in credits on the Starter tier with no card, then $100 a month for 100 credits. Runs are billed by block: $0.02 for simple operations, $0.10 for standard AI and $0.30 for complex AI. Their own worked example says a typical invoice workflow runs 4 to 6 blocks per document, so at $0.30 per complex block that is under $2 per invoice end to end. Growth is volume priced with up to 40 percent discount. Enterprise, which includes private cloud or on premise deployment and data residency, is quote only and the price is not published.
Which open weight OCR model scores highest right now?
On OmniDocBench v1.6, as reported in Roboflow’s open source OCR round up, PaddleOCR-VL-1.6 leads at 96.34 percent with 0.9B parameters under Apache-2.0, followed by MinerU2.5-Pro at 95.75 percent and GLM-OCR at 95.22 percent. Licences differ and matter: MinerU carries additional terms with usage and revenue thresholds, so it is not a clean Apache-2.0 grant. A leaderboard is a shortlist rather than an answer. Run fifty of your own difficult documents through the top three before you commit, because ranking on public benchmarks rarely survives contact with a supplier who sends photographs.
At what volume does self hosting beat Nanonets?
Using Nanonets’ own figure of roughly $1.50 per document, 2,000 documents a month is $36,000 a year. Subtract a continuously reserved L40S at about $8,410 a year and under $28,000 remains to pay for a build, a rollout and ongoing operations, which is not enough. At 10,000 documents a month the same arithmetic gives $180,000 a year against the same GPU line, and the gap widens every month volume grows. Below roughly 2,000 documents a month, buy Nanonets. Above roughly 10,000, a fixed cost build is worth pricing properly.
Does Ækora have a free trial or a self serve plan?
No. Ækora is a consultancy and deployment engagement, not a SaaS product. There is no sign up page, no free tier, no trial and no per seat plan. Work starts with scope, moves to a pilot on your own documents and ends with a rollout that you own and run inside your own infrastructure. That means weeks to a first production result rather than an afternoon. If you want a login and a card on file today, Nanonets sell exactly that and we would rather you bought it than waited for us.
Want this worked out on your documents?
We will price a real sample against OpenAI or Anthropic and tell you whether Bulk document processing or a local model on your hardware is the cheaper first step.
If you want a self hosted alternative to Nanonets, start with a fact Nanonets do not hide: they publish their own open weights. Nanonets-OCR2-3B sits on HuggingFace as a fine tune of Qwen2.5-VL-3B, free to download and run on your own GPU. So this page is not our model against their model. It is their hosted workflow at their published price, against their own model running inside your network with the workflow built around your process.
One thing to be clear about before you read further. Ækora is a consultancy and a deployment engagement, not a product. There is no free tier, no trial and no sign up page. You get scope, a pilot on your own documents, then a rollout you own and run. If you need invoices extracting this afternoon, sign up for Nanonets. That is a genuine recommendation and it appears more than once below.
What Nanonets actually sell
Nanonets sell a workflow product, and it is a good one. The model is the small part. What you pay for is the validation interface where a human fixes what the model got wrong, the approval routing, the integrations into your ERP or accounting system, the retraining loop and, on the upper tiers, support with an SLA behind it.
Their published pricing works in credits and blocks. Starter gives $50 in credits with no card, then $100 a month for 100 credits. Runs are priced by block complexity: $0.02 for simple operations, $0.10 for standard AI and $0.30 for complex AI. Growth is volume priced with up to 40 percent discount. Enterprise is custom and includes HIPAA and SOC 2 compliance, private cloud or on premise deployment, data residency in the US, EU or APAC, and dedicated support and SLAs.
The most useful number on their page is their own worked example: a typical invoice processing workflow runs 4 to 6 blocks per document, and at $0.30 per complex AI block that is under $2 per invoice end to end. Hold on to that figure. Everything below is measured against it.
Note
Nanonets already sell on premise deployment and EU data residency on their Enterprise tier. If your only problem is that documents must not leave your network or your jurisdiction, ask Nanonets for an Enterprise quote before you talk to anyone else. That price is not published, and we are not going to guess at it.
Nanonets publish the model
Definition · Nanonets-OCR2-3B
An open weight document model published by Nanonets on HuggingFace, built on Qwen2.5-VL-3B. It converts pages into structured markdown: LaTeX for equations, markdown and HTML for tables, Unicode symbols for checkboxes and radio buttons, tagged signature regions, handwriting across several languages, and flow charts as mermaid code.
It is genuinely used. The HuggingFace page showed roughly 10,800 downloads in the last month when we checked, with over 500 likes. One caveat that matters commercially: at the time of writing the model card carries no licence tag. A vendor's hosted terms of service does not travel with a weights download, so confirm the licence in writing before you put it into production. That applies to every open weight model on this page, not only this one.
The rest of the open weight shortlist
The Nanonets model is not automatically the right one for your pages. Measured on OmniDocBench v1.6 in Roboflow's current open source round up, the field looks like this.
Open weight document models. Parameter counts, licences and OmniDocBench v1.6 scores from Roboflow's open source OCR round up at blog.roboflow.com/best-open-source-ocr-models. Nanonets-OCR2-3B figures from its HuggingFace model card. Docling from its GitHub repository at github.com/docling-project/docling.
Model
Params
Licence
OmniDocBench v1.6
What it is for
PaddleOCR-VL-1.6
0.9B
Apache-2.0
96.34%
Highest scoring general document parsing at a small size
MinerU2.5-Pro
1.2B
Apache-2.0 with additional terms, including licensing requirements at stated usage or revenue thresholds
95.75%
Dense and academic documents
GLM-OCR
0.9B
MIT for the official model, Apache-2.0 for the PP-DocLayoutV3 component
95.22%
Small footprint with a clean permissive licence
dots.mocr
About 3B
MIT with additional terms
Not published in the round up
Layout aware page parsing
Nanonets-OCR2-3B
3B, a Qwen2.5-VL-3B fine tune
No licence stated on the model card
Not covered in the round up
Structured markdown with checkboxes, signatures and equations
Open weight document models. Parameter counts, licences and OmniDocBench v1.6 scores from Roboflow's open source OCR round up at blog.roboflow.com/best-open-source-ocr-models. Nanonets-OCR2-3B figures from its HuggingFace model card. Docling from its GitHub repository at github.com/docling-project/docling.
Two things to take from that table. First, no published head to head exists between the Nanonets hosted pipeline and any of these models on your documents, so nobody, us included, can honestly tell you which is more accurate on your invoices. Benchmarks give you a shortlist, not an answer. We go through how to build that shortlist in our open weight model round up. Second, Docling belongs in a different column entirely.
Docling is a pipeline, not a competing model
Nanonets versus Docling is a question people genuinely type, and the honest answer is that they sit at different layers. Docling, started by the AI for knowledge team at IBM Research Zurich and now hosted at the LF AI and Data Foundation under an MIT licence, is a toolkit: page layout, reading order, table structure, export to markdown, HTML or JSON, and pluggable OCR backends including vision language models. You would use Docling with one of the models above, not instead of one. It has become the default entry point for people building self hosted document pipelines, which is exactly why the comparison keeps getting made.
What self hosting actually costs
The compute is cheaper than most people expect, and the compute is not the expensive part.
Spheron's 2026 self hosting guide puts PaddleOCR-VL-1.6 on an L40S at $0.96 an hour, costing about $7.27 for 10,000 pages, against roughly $15.00 for AWS Textract at the same volume. Their throughput figures are sustained end to end estimates including PDF rasterisation, input and output and processing overhead, which they put at roughly half the peak decode rates. Repeat that caveat to yourself before you size anything: the throughput numbers on a model card are about double what you will actually get in production.
Work the arithmetic through. At $7.27 per 10,000 pages and $0.96 an hour, one L40S sustains roughly 1,300 pages an hour. A reserved L40S running continuously costs about $701 a month, or $8,410 a year. At 20,000 two page documents a month you are asking that card for about thirty hours of work, so the GPU line barely moves as volume climbs. That is the structural difference. Nanonets bill per document. A GPU bills per hour whether you feed it or not. Our GPU sizing note covers how to work out which card you actually need.
Nanonets modelled at five complex AI blocks per document ($1.50) using the block prices and worked example at nanonets.com/pricing, before any Growth tier discount of up to 40%. GPU line is a continuously reserved L40S at $0.96 an hour from spheron.network's 2026 self hosting guide. Build and operating costs are excluded and are not published here.
Documents a month
Nanonets a year at $1.50 each
Reserved L40S a year
Gap left for build and operations
500
$9,000
$8,410
$590
2,000
$36,000
$8,410
$27,590
5,000
$90,000
$8,410
$81,590
10,000
$180,000
$8,410
$171,590
20,000
$360,000
$8,410
$351,590
50,000
$900,000, or about $540,000 at the full 40% Growth discount
$8,410
$531,590 or more
Nanonets modelled at five complex AI blocks per document ($1.50) using the block prices and worked example at nanonets.com/pricing, before any Growth tier discount of up to 40%. GPU line is a continuously reserved L40S at $0.96 an hour from spheron.network's 2026 self hosting guide. Build and operating costs are excluded and are not published here.
Two honest caveats on that table. The Nanonets example is 4 to 6 blocks and not all of them need to be complex, so a well built workflow may land under $1.50 a document. And the right hand column is not savings. It is the budget available to cover a build, a GPU somebody has to keep running, monitoring, model updates and the hours your own people spend on the review queue.
The volume threshold, stated plainly
Below roughly 2,000 documents a month, buy Nanonets. At $1.50 a document that is $36,000 a year, and once you subtract a reserved GPU you have under $28,000 left to pay for a build, a rollout and ongoing operations. A serious on premise deployment does not come in under that and stay useful, and you would be trading a working validation interface for a project. That is a bad trade at that volume and we will say so on the call.
Between roughly 2,000 and 10,000 documents a month it depends on why you are asking. If the driver is cost, the numbers are tight, and the Nanonets Growth discount of up to 40 percent narrows them further. If the driver is that documents cannot leave your network, get the Nanonets Enterprise price first and compare like with like.
Above roughly 10,000 documents a month the per document meter becomes a dominant line in the budget and a fixed cost build starts to look structurally different. At 10,000 documents a month you are spending $180,000 a year, which buys an engagement and a GPU several times over, and the gap widens every month volume grows, because the self hosted side is mostly a card that sits idle. We set out that shape in more detail in our note on fixed cost versus per token invoice extraction.
What open weights do not give you
This is where most self hosting write ups go quiet, so here it is in full. Downloading Nanonets-OCR2-3B, or PaddleOCR-VL, or pointing Docling at a folder gives you pages converted to structured text. It does not give you:
Calibrated confidence scores per field. Without them you cannot tell which extractions need a human, which is the whole basis of a review queue.
Key and value extraction against your schema. The model returns a structured page, not your invoice fields, your supplier identifiers or your cost centres.
A human in the loop queue. That means a review interface, assignment, an audit trail and a path for corrections to feed back into the system.
Line item reconciliation against purchase orders or goods received notes.
A validation screen your accounts payable team will actually use.
Those five things are most of what Nanonets sell, and most of what a deployment engagement has to build. When we scope bulk document processing, that is what the number covers. The model is the cheap part.
Where Ækora loses
Time to first result. Nanonets works this afternoon. We work through scope, a pilot on your documents, then a rollout. That is weeks, not hours, and if your problem is urgent that difference matters more than the arithmetic above.
No validation interface as a product. We build review workflow into your systems rather than shipping a screen with a roadmap behind it. If you want a polished queue out of the box, Nanonets have one and we do not.
No SLA and no status page. You are buying an engagement with named people, not a platform with uptime commitments. Nanonets offer dedicated support and SLAs on Enterprise.
No vendor side SOC 2. Nanonets list HIPAA and SOC 2 compliance on their Enterprise tier. We deploy into your infrastructure, which inherits your controls and your certifications. For some buyers that is the right answer and for others it plainly is not.
No published accuracy comparison. We have not run a head to head against the Nanonets hosted pipeline, nobody else has published one either, and we are not going to claim a win we cannot show.
Low volume self serve. A few hundred documents a month, a card on file and a login is a shape we do not sell. Nanonets do.
We also cannot make you GDPR compliant. Keeping processing inside your own infrastructure removes a category of transfer and sub processor questions. The rest of the obligation stays with you.
How to decide this week
Get the Nanonets Enterprise quote if data residency is the driver. It is not published, so asking is the only way to know.
Count real monthly volume and pages per document, multiply by $1.50 and by twelve, then compare against the table above using your own numbers rather than ours.
Take fifty of your genuinely difficult documents, the smudged scans and the multi page tables and the supplier who sends photographs, and run them through two or three open weight models.
Size the hardware before you cost anything. A card idle most of the month and a card running flat out produce very different answers.
Decide what you are actually buying: a workflow product with a review screen and a support contract, or extraction inside your network with the workflow built around your process.
Where to go next
If the volume and the residency answers both point away from a hosted meter, our bulk document processing page covers what a pilot includes, what we need from you and how a rollout runs. If you are comparing against enterprise capture platforms rather than a per document meter, the Rossum comparison takes the same question from that side. If you already know the shape of the work, tell us the volume and the document types and we will tell you whether it is worth doing. If it is not, we will say so, and Nanonets will still be there.