EU AI Act obligations for document processing: mostly minimal risk, with one Annex III trap
TL;DR
Reading invoices, receipts, contracts or claims into fields is not on the Annex III list, so the high-risk obligations that applied from 2 August 2026 mostly do not attach to document extraction.
Every deployer still owes the Art. 4 AI-literacy duty (since 2 February 2025) and the Art. 5 prohibited-practice check — and GDPR as usual.
The trap is “intended to be used”: if the extracted fields feed creditworthiness, insurance pricing, hiring or benefit eligibility, the system is high-risk and you are a deployer of it.
General-purpose model obligations under Art. 53 sit with the model provider, not with you. Open-weight models are exempt from some documentation duties, but not from the copyright policy or the training-content summary.
If you are pulled into high-risk, keep the automatically generated logs for at least six months — a local model on your own hardware satisfies that without a vendor dependency.
Questions people ask
Is document processing high-risk under the EU AI Act?
Mostly no. Art. 6(2) makes a system high-risk only if it falls under an Annex III use case, and invoice extraction, receipt capture, contract clause extraction and claims data capture appear nowhere on that list. It becomes high-risk when its output is used to make or materially inform an Annex III decision about a person — creditworthiness, life or health insurance pricing, recruitment, or access to public benefits.
When did the EU AI Act start applying to document processing?
In stages under Art. 113. The Act entered into force on 1 August 2024. Art. 4 AI literacy and the Art. 5 prohibitions have applied since 2 February 2025 to every deployer. General-purpose model obligations started on 2 August 2025 for model providers. The Annex III high-risk regime, Art. 50 transparency and Art. 26 deployer duties applied from 2 August 2026; Annex I product-embedded systems follow on 2 August 2027.
Am I a provider or a deployer under the EU AI Act if I use document AI?
A finance team using a document SaaS is a deployer under Art. 3(4); the vendor is the provider. If you download an open-weight model and assemble the system yourself, you are the deployer of that system and the model’s publisher is the general-purpose model provider. Under Art. 25 you become a provider of a high-risk system if you put your own name on it, substantially modify it, or change its purpose into an Annex III use.
Do open-source AI models have to comply with the EU AI Act?
Partly. Art. 53(2) exempts providers of models released under a free and open-source license, with weights and architecture public, from the technical-documentation duties in Art. 53(1)(a) and (b) — but not from the copyright policy or the training-content summary, nor if the model carries systemic risk. Many “open” licenses with acceptable-use or commercial restrictions may not qualify. As a deployer, pick a model whose provider has published the Art. 53 materials and record the version you run.
What does a deployer of a high-risk AI system have to do under Art. 26?
Use the system per the provider’s instructions; assign named, trained people with authority to override or stop it; ensure the input documents are relevant to the intended purpose; keep the automatically generated logs for at least six months; inform workers’ representatives before workplace deployment and tell people subject to a decision that AI was used. Deployers of creditworthiness and insurance-pricing systems, and public bodies, must also complete an Art. 27 fundamental rights impact assessment before first use.
Does Art. 50 transparency apply to document extraction?
Not to the extraction step. Art. 50 covers systems that interact directly with people, synthetic audio, image, video or text output, emotion recognition and biometric categorization, and deepfakes or AI-written text on public-interest matters. A back-office extractor turning a scanned invoice into a JSON row does none of these. It can matter if the same pipeline drafts customer-facing letters published as if a person wrote them, or if the intake form is a chatbot — those are separate systems.
Want this worked out on your documents?
We will price a real sample against OpenAI or Anthropic and tell you whether Bulk document processing or a local model on your hardware is the cheaper first step.
Most document-extraction pipelines are not “high-risk” under the EU AI Act. Reading an invoice, a receipt, a contract or a claim form into fields is not on the Annex III list, so the heavy obligations that started applying on 2 August 2026 — risk management, technical documentation, logging, human oversight, conformity assessment — mostly do not attach to them. What does apply to almost every European organization using such a system: the AI-literacy duty in Art. 4 (since 2 February 2025), the prohibited-practice check in Art. 5, and, if your extraction output feeds a decision that is on the Annex III list (creditworthiness, insurance pricing, hiring, benefit eligibility), the deployer duties in Art. 26. The model layer has its own rules: general-purpose model obligations (Art. 53) sit with the model provider, not with you, and open-weight models get a partial exemption.
Definition · Document processing under the EU AI Act
Under Regulation (EU) 2024/1689, a document-extraction system is classed as high-risk only if it is used in one of the Annex III domains or is a safety component of a product covered by Annex I. Back-office extraction of invoices, receipts, contracts and claims data is not listed; it becomes high-risk when its output is used to make, or materially inform, an Annex III decision about a person, such as creditworthiness, life or health insurance pricing, recruitment or access to public benefits.
The timeline, and where you are on it
The AI Act entered into force on 1 August 2024 and applies in stages under Art. 113. The dates that matter for a document pipeline:
Application dates of the EU AI Act (Regulation (EU) 2024/1689, Art. 113) as adopted. Dates are those in the regulation text; check for amendments before relying on them.
Date
What starts applying
Relevance to document extraction
1 Aug 2024
Entry into force
Clock starts; no obligations yet
2 Feb 2025
Art. 4 AI literacy; Art. 5 prohibited practices
Applies to every deployer, including of low-risk extraction
2 Aug 2025
Chapter V general-purpose AI model obligations; governance; penalties
Applies to the providers of the models you run, not to you as deployer
2 Aug 2026
General application: Annex III high-risk obligations, Art. 50 transparency, Art. 26 deployer duties
In force at time of writing; matters if your output feeds an Annex III decision
2 Aug 2027
Annex I product-embedded high-risk systems; legacy GPAI models must comply
Rarely relevant to document extraction
Application dates of the EU AI Act (Regulation (EU) 2024/1689, Art. 113) as adopted. Dates are those in the regulation text; check for amendments before relying on them.
So as of this article’s date, the high-risk regime is live. Two footnotes. Under Art. 111(2), a high-risk system placed on the market before 2 August 2026 is only caught when its design changes significantly — which does not help you if you are deploying now. And the Commission published a “digital omnibus” proposal in November 2025 that would tie some high-risk deadlines to the availability of harmonized standards; whether and how that has been adopted should be checked before you plan around a date. Nothing in that proposal changes the analysis of which tier you are in.
Is document extraction high-risk? Mostly no — with one trap
Art. 6(2) says a system is high-risk if it falls under one of the use cases in Annex III. Annex III lists eight areas: biometrics; critical infrastructure; education and vocational training; employment and workers’ management; access to essential private and public services (including creditworthiness, life and health insurance pricing, and public-benefit eligibility); law enforcement; migration and border control; and administration of justice. Invoice extraction, receipt capture, contract clause extraction and claims data capture appear nowhere on that list.
The trap is the phrase “intended to be used.” An extraction model is a component. If the fields it produces are used to evaluate a job applicant’s CV (Annex III, point 4), to assess creditworthiness from bank statements (point 5(b)), to price a health-insurance policy from a claims history (point 5(c)), or to decide public-benefit eligibility from uploaded documents (point 5(a)), the system that includes the extractor is high-risk, and the organization that uses it is a deployer of a high-risk system. The extractor did not become high-risk; the decision it feeds did. Art. 6(3) offers a derogation for systems that only perform a narrow procedural task or a preparatory task to an assessment — pure data capture that a human then evaluates can fall under it — but the assessment must be documented by the provider before the system is placed on the market or put into service (Art. 6(4)), and if you assembled the system yourself, that provider is you. The derogation never applies where the system profiles natural persons.
Likely risk tier for common document-processing use cases under the EU AI Act, and the obligations that follow. Tiering is our reading of Art. 6 and Annex III; confirm your own classification with counsel and document it under Art. 6(4) where you rely on the derogation.
Use case
Likely tier
What you owe as deployer
Accounts-payable invoice extraction to ERP
Minimal / not listed
Art. 4 literacy; Art. 5 check; GDPR as usual
Receipt and expense capture
Minimal / not listed
Same
Contract clause extraction for legal review
Minimal / not listed
Same; keep human review as the decision point
Claims data capture, human adjuster decides
Not listed; possible Art. 6(3) preparatory task
Document the Art. 6(3) assessment; Art. 4; GDPR Art. 9 if health data
Bank-statement extraction feeding an automated credit score
Art. 26; Art. 27 FRIA; registration under Art. 49(3) for public bodies
Likely risk tier for common document-processing use cases under the EU AI Act, and the obligations that follow. Tiering is our reading of Art. 6 and Annex III; confirm your own classification with counsel and document it under Art. 6(4) where you rely on the derogation.
Provider or deployer: which one you are
Art. 3(3) defines a provider as whoever develops an AI system or has it developed and places it on the market or puts it into service under their own name. Art. 3(4) defines a deployer as whoever uses an AI system under their own authority in a professional context. A finance team using a document SaaS is a deployer. Ækora, when we build and run a pipeline for you, is the provider of that pipeline and you remain the deployer. If you download Qwen3-VL and run it yourself, you are the deployer of the system you assembled, and — for the model layer — the model’s publisher is the GPAI provider.
Art. 25 is where the roles flip. A deployer becomes a provider of a high-risk system if it puts its own name on it, makes a substantial modification, or changes the intended purpose so that the system becomes high-risk. In practice: take a vendor’s invoice extractor, retrain it, and use it to score credit applicants, and you have become the provider with the full Art. 16 obligations. Keep the extractor as an extractor and you have not.
What a deployer of a low-risk extractor should still document
The regulation does not require a file for minimal-risk systems. Art. 4 and your own GDPR accountability do, and a short file is what makes the “not high-risk” position defensible when someone asks. Four items cover it: the intended purpose, in one sentence, with the decision it does not make; the Art. 6 classification reasoning and, where relied on, the Art. 6(3) derogation assessment; the AI-literacy measures taken for the people who operate and review the output; and the model provenance — which model, which version, which license, whether it is a GPAI model and who its provider is.
General-purpose models: obligations sit with the provider, open weights get a partial exemption
Chapter V (Art. 51–56) applies to providers of general-purpose AI models — a category that includes large vision-language models capable of reading documents, whether closed (GPT-4.1, Claude Sonnet 5) or open-weight (Qwen3-VL, Gemma 4, Llama 4 Scout, DeepSeek-OCR). Art. 53(1) asks providers to keep technical documentation, give downstream providers the information they need to integrate the model, put in place a copyright-compliance policy, and publish a summary of the training content. Art. 51 adds “systemic risk” duties for models trained above a compute threshold of 10²⁵ floating-point operations.
Open-source matters here in two ways. Art. 53(2) exempts providers of models released under a free and open-source license, with weights and architecture publicly available, from the documentation duties in 53(1)(a) and (b) — but not from the copyright policy and the training-content summary, and not at all if the model has systemic risk. Separately, Art. 2(12) says the regulation does not apply to AI systems released under free and open-source licenses unless they are placed on the market as high-risk or fall under Art. 5 or Art. 50. Two caveats. Many “open” model licenses carry acceptable-use restrictions or commercial thresholds that may not meet the regulation’s idea of free and open-source; and the exemption is for the system as released, not for what you build on it and use in your organization. As a deployer, your practical takeaway is simple: choose a model whose provider has published the Art. 53 materials, and record the version you run.
Art. 50 transparency: mostly not your problem
Art. 50 imposes transparency duties in four situations: systems that interact directly with people must say so; providers of systems generating synthetic audio, image, video or text must mark the output as machine-generated; deployers of emotion-recognition and biometric-categorization systems must inform the people affected; and deployers of deepfakes, or of AI-generated text published on matters of public interest, must disclose it. A back-office extractor that turns a scanned invoice into a JSON row does none of these things. Where it can touch Art. 50: if the same pipeline drafts a customer-facing letter from the extracted fields and you publish it as if a person wrote it, or if the intake form is a chatbot. Those are separate systems with their own obligations; the extraction step remains outside Art. 50.
If you are pulled into high-risk: record-keeping and human oversight
When an extraction output feeds an Annex III decision, the deployer duties in Art. 26 attach to the whole system. The ones that change how you run a pipeline:
Use per the instructions (Art. 26(1)) — the provider’s Art. 13 instructions define the envelope. Running an invoice extractor on bank statements for credit scoring is outside it.
Human oversight (Art. 26(2), Art. 14) — assign named, trained people with authority to override or stop the system. A confidence threshold that sends low-scoring pages to a reviewer is the mechanism; the named people are the compliance.
Input data relevance (Art. 26(4)) — to the extent you control which documents go in, you are responsible for them being relevant and representative of the intended purpose.
Logs (Art. 26(6), Art. 12) — keep the automatically generated logs for at least six months, longer if other law says so. This is one place where a local deployment helps: the logs are on your infrastructure and you decide retention, rather than depending on a vendor’s retention window.
Inform people (Art. 26(7), (11)) — workers’ representatives before deployment in the workplace, and natural persons subject to a decision that the system was used.
Fundamental rights impact assessment (Art. 27) — required for public bodies, private entities providing public services, and deployers of the creditworthiness and insurance-pricing systems in Annex III 5(b) and (c), before first use.
None of this is specific to whether the model is local or cloud. What is specific is who can produce the evidence. A local model on your own hardware with your own logging satisfies Art. 12 and 26(6) without a vendor dependency; a cloud API satisfies it only if the provider’s log retention and export match your six-month floor. For the data-protection side of the same pipeline — lawful basis, Art. 28 chains, transfers — see GDPR-compliant document AI; the two regimes stack rather than replace each other. For what running such a model on your own hardware involves, see what on-premise document AI is and the European buyer’s guide.
The compliance file: nine items, one folder
Whether you land in minimal risk or in Annex III, the same folder answers the question in an audit, a customer security review or a works-council meeting. Build it in this order.
System description. What goes in (document families), what comes out (fields), and the one decision it does not make. Half a page.
Classification memo. Your Art. 6 reasoning against each Annex III point, and the Art. 6(3) assessment if you rely on it. Date it before go-live.
Roles. Who is provider and who is deployer for each layer: model, pipeline, application. Note any Art. 25 trigger you are avoiding.
Model provenance. Model name, version, license, GPAI provider, and a link to the provider’s Art. 53 documentation and training-content summary.
Prohibited-practice check. A one-paragraph confirmation against Art. 5 — no emotion inference at work, no social scoring, no untargeted scraping.
AI-literacy record. Who operates and reviews the system, what training they had, when (Art. 4).
Oversight design. The confidence threshold, who reviews below it, and how they can override or stop the system.
Logging and retention. What is logged, where it lives, how long. Six months minimum if high-risk; state the number even if you are not.
Cross-reference to the GDPR file. The DPIA, the Art. 30 entry and the Art. 28 contracts for the same pipeline, so the two regimes point at one data path.
The most common failure we see is not misclassification. It is a pipeline that grew from “extract the invoice” into “extract the invoice and flag the supplier for credit hold” without anyone noticing the second half is an Annex III use. The cure is item 1: write down the decision the system does not make, and re-read it whenever a new field is added. If you want the cost and residency side of the same decision, start with why per-invoice API pricing stops making sense, then look at bulk document processing on an EU-only pipeline or local document models on your own hardware — both give you the logs, the model version and the data path that this file needs.