13 August 202611 min readGDPR & EU AI Act

EU AI Act obligations for document processing: mostly minimal risk, with one Annex III trap

TL;DR

  • Reading invoices, receipts, contracts or claims into fields is not on the Annex III list, so the high-risk obligations that applied from 2 August 2026 mostly do not attach to document extraction.
  • Every deployer still owes the Art. 4 AI-literacy duty (since 2 February 2025) and the Art. 5 prohibited-practice check — and GDPR as usual.
  • The trap is “intended to be used”: if the extracted fields feed creditworthiness, insurance pricing, hiring or benefit eligibility, the system is high-risk and you are a deployer of it.
  • General-purpose model obligations under Art. 53 sit with the model provider, not with you. Open-weight models are exempt from some documentation duties, but not from the copyright policy or the training-content summary.
  • If you are pulled into high-risk, keep the automatically generated logs for at least six months — a local model on your own hardware satisfies that without a vendor dependency.

Questions people ask

Is document processing high-risk under the EU AI Act?
Mostly no. Art. 6(2) makes a system high-risk only if it falls under an Annex III use case, and invoice extraction, receipt capture, contract clause extraction and claims data capture appear nowhere on that list. It becomes high-risk when its output is used to make or materially inform an Annex III decision about a person — creditworthiness, life or health insurance pricing, recruitment, or access to public benefits.
When did the EU AI Act start applying to document processing?
In stages under Art. 113. The Act entered into force on 1 August 2024. Art. 4 AI literacy and the Art. 5 prohibitions have applied since 2 February 2025 to every deployer. General-purpose model obligations started on 2 August 2025 for model providers. The Annex III high-risk regime, Art. 50 transparency and Art. 26 deployer duties applied from 2 August 2026; Annex I product-embedded systems follow on 2 August 2027.
Am I a provider or a deployer under the EU AI Act if I use document AI?
A finance team using a document SaaS is a deployer under Art. 3(4); the vendor is the provider. If you download an open-weight model and assemble the system yourself, you are the deployer of that system and the model’s publisher is the general-purpose model provider. Under Art. 25 you become a provider of a high-risk system if you put your own name on it, substantially modify it, or change its purpose into an Annex III use.
Do open-source AI models have to comply with the EU AI Act?
Partly. Art. 53(2) exempts providers of models released under a free and open-source license, with weights and architecture public, from the technical-documentation duties in Art. 53(1)(a) and (b) — but not from the copyright policy or the training-content summary, nor if the model carries systemic risk. Many “open” licenses with acceptable-use or commercial restrictions may not qualify. As a deployer, pick a model whose provider has published the Art. 53 materials and record the version you run.
What does a deployer of a high-risk AI system have to do under Art. 26?
Use the system per the provider’s instructions; assign named, trained people with authority to override or stop it; ensure the input documents are relevant to the intended purpose; keep the automatically generated logs for at least six months; inform workers’ representatives before workplace deployment and tell people subject to a decision that AI was used. Deployers of creditworthiness and insurance-pricing systems, and public bodies, must also complete an Art. 27 fundamental rights impact assessment before first use.
Does Art. 50 transparency apply to document extraction?
Not to the extraction step. Art. 50 covers systems that interact directly with people, synthetic audio, image, video or text output, emotion recognition and biometric categorization, and deepfakes or AI-written text on public-interest matters. A back-office extractor turning a scanned invoice into a JSON row does none of these. It can matter if the same pipeline drafts customer-facing letters published as if a person wrote them, or if the intake form is a chatbot — those are separate systems.

Want this worked out on your documents?

We will price a real sample against OpenAI or Anthropic and tell you whether Bulk document processing or a local model on your hardware is the cheaper first step.