Mapping local open-weight document models to PiTuKri's eleven control areas
TL;DR
PiTuKri is the Finnish cloud services security assessment criteria set published by Traficom, currently at version 1.1, organised into eleven control areas containing 41 requirement cards.
Requirement EE-02 on legislation-derived risks decides most AI document procurements, because it asks which actors can be legally compelled to access the data rather than where the data physically sits.
Running open-weight document models on infrastructure you control does not remove PiTuKri obligations; it moves them from a cloud provider you must trust to an environment you can inspect, with physical and operational security becoming your burden.
Finland has no data localisation law, and neither does the EU; demand for local processing comes from classification rules and procurement, not from a localisation statute.
Published benchmark scores for open document models are measured on English and Chinese corpora, so Finnish-language accuracy has to be measured on your own documents rather than assumed.
Questions people ask
Can an AI document processing pipeline pass a PiTuKri assessment?
It can be assessed, and whether it passes depends on the data classification, the processing environment and the responsible parties, not on the AI component in isolation. PiTuKri assesses a service and its operating environment against eleven control areas. An AI pipeline is simply part of that environment. The practical difficulty is usually requirement EE-02, which demands a full description of physical locations, subcontracting chains, applicable jurisdiction and every actor that could be legally compelled to access the data. Many cloud-hosted AI services cannot supply that description in a form that survives scrutiny.
Does running models on-premise make PiTuKri compliance automatic?
No. PiTuKri is an assessment framework applied by a competent authority or assessment body, not a certification anyone holds in advance, and no supplier can confer it on you. Local deployment changes which criteria apply and who is responsible for them. It generally helps with the prerequisites, communications and encryption areas, and it hands you the physical security and operational security areas in full. Security management, personnel security and change management stay yours either way. If your organisation cannot carry physical and operational security, local deployment makes the assessment harder rather than easier.
Is Finnish public sector data legally required to stay in Finland?
There is no general data localisation mandate in Finnish or EU law. Regulation (EU) 2018/1807 on the free flow of non-personal data points the other way, and PiTuKri itself cites it while noting that it does not apply to location requirements imposed on national security and preparedness grounds. Real restrictions come from classification rules under the Tiedonhallintalaki and the government decree on document classification, and from procurement requirements. For security-classified material the practical conditions can be strict, but they arise from classification, not from a localisation statute.
How accurate are open-weight document models on Finnish text?
Nobody can honestly tell you without measuring on your documents. The headline benchmark for document parsing, OmniDocBench v1.6, documents its language categories as English, Simplified Chinese and mixed English-Chinese, with no Nordic language represented. Finnish is absent, and Finnish is morphologically rich and compounds heavily, which is exactly the kind of text where extraction quality diverges from English results. Take 200 to 500 pages of your own material in the formats you actually receive, and measure field-level accuracy against a human-labelled ground truth. That measurement is the only evidence worth putting into a procurement document.
What volume makes on-premise document AI worth it in Finland?
Below roughly 250,000 pages a year, buying GPUs and standing up a controlled environment usually costs more than paying a cloud API, and the engineering overhead is not recovered. Above that the arithmetic starts favouring fixed-cost local processing. Classification overrides the arithmetic entirely: if the documents are turvallisuusluokitellut asiakirjat, page count is irrelevant and the question becomes whether the environment can satisfy the criteria at all. At that point a service where a foreign provider retains technical access to cleartext is unlikely to qualify at any price.
Is PiTuKri being replaced?
Yes. On 29 January 2026 the Ministry of Finance and Traficom announced work on joint national cloud security criteria that will supersede PiTuKri and the cloud services sections of Julkri, with a public comment period that closed in April 2026 and completion scheduled for autumn 2026. Version 1.1 remains the published criteria set today and is what assessments currently reference. If you are scoping a procurement that lands in 2027 or later, ask Traficom about the transition before you commit to a specific control mapping.
Want this worked out on your documents?
We will price a real sample against OpenAI or Anthropic and tell you whether Local document models or a local model on your hardware is the cheaper first step.
Yes, an AI document processing pipeline can be assessed under PiTuKri, and the assessment gets simpler the less of it runs in someone else's cloud service. PiTuKri is Traficom's criteria set for evaluating the security of pilvipalvelut, and it is built on a shared responsibility model: every requirement lands on the cloud provider, on the customer, or on both. When you run open-weight document models on infrastructure your own organisation controls, a large block of requirements stops being a supplier's assurances you have to trust and becomes your own environment, which you can inspect.
That answer is narrower than it sounds. Moving the models onto your own hardware produces no assessment result and removes none of your obligations. It changes which criteria apply to whom. This page maps PiTuKri's control areas to what a local deployment changes.
Definition · PiTuKri
Pilvipalveluiden turvallisuuden arviointikriteeristö, the Criteria for Assessing the Information Security of Cloud Services, published by the National Cyber Security Centre at Traficom. The current published version is 1.1 (Traficomin julkaisuja 13/2020, dated March 2020). It exists to protect viranomaisten salassa pidettävä tieto when that information is processed in cloud services, and it addresses nationally classified information up to TL IV, touching on the general protection principles for international RESTRICTED.
What PiTuKri actually contains
PiTuKri v1.1 is organised into eleven control areas (osa-alueet) containing 41 requirement cards. Each card carries a two-letter prefix and a number, and states the requirement, its applicability, the data types it covers and its protection objective. The eleven areas are Esiehdot (EE), Turvallisuusjohtaminen (TJ), Henkilöstöturvallisuus (HT), Fyysinen turvallisuus (FT), Tietoliikenneturvallisuus (TT), Identiteetin ja pääsyn hallinta (IP), Tietojärjestelmäturvallisuus (JT), Salaus (SA), Käyttöturvallisuus (KT), Siirrettävyys ja yhteensopivuus (SI), and Muutostenhallinta ja järjestelmäkehitys (MH). It draws on BSI's C5, the CSA Cloud Controls Matrix, ISO/IEC 27001 and 27017, and Katakri 2015. The criteria and an Excel assessment tool are at kyberturvallisuuskeskus.fi, with an English edition alongside the Finnish.
Note
PiTuKri is being replaced. On 29 January 2026 the Ministry of Finance and Traficom announced joint national cloud criteria that will supersede PiTuKri and the cloud sections of Julkri, scheduled for completion in autumn 2026 (valtioneuvosto.fi). Everything below describes v1.1, which is what is published today. If you are scoping a procurement that lands in 2027, ask Traficom what the transition looks like before you commit to a control mapping.
EE-02 decides most AI procurements
The card that matters most for an AI document pipeline is EE-02, Lainsäädäntöjohdannaiset riskit, or legislation-derived risks. It sits in the prerequisites area, so it is assessed before the technical controls are worth discussing. EE-02 requires that the provider's descriptions let you evaluate, at minimum: the physical location of the data across its whole lifecycle including subcontracting chains; the location of the service's functions and components such as administration and backups; any other parties involved in producing the service; the applicable legislation and jurisdiction; and which actors may, by virtue of applicable legislation, have access to the data processed in the service. It then requires that those risks do not restrict the service's suitability for the use case.
Read that last item carefully, because it is the US CLOUD Act argument written by a Finnish authority. EE-02 does not ask where the disk sits. It asks who can be compelled to produce the data. An EU region is a location claim; EE-02 is a control claim. PiTuKri's commentary on the card notes that legislation-derived disclosure powers can reach both the physical location of confidential data and disclosure carried out from another country through management connections.
PiTuKri also closes the encryption escape hatch. Its section on the location of data and services states that generally applicable cryptographic protections do not bring significant additional protection against legislation-derived risks, and the footnote is specific: Bring Your Own Key schemes and HSMs placed in the provider's data centre limit, but do not typically prevent, provider access. Card SA-03 repeats it: as a starting point the provider always has access if the data exists in cleartext at any point in its lifecycle, for example when rendered as an image. Document AI renders documents as images to a model. That is the whole job.
The eleven control areas and what a local deployment changes
The table below maps each PiTuKri v1.1 control area to what changes when the document models run on hardware your organisation controls rather than in a cloud service. It is a planning aid, not an assessment. Requirement-level applicability depends on your use case, service model and data types, and PiTuKri devotes an annex to that question.
PiTuKri v1.1 control areas mapped to a local open-weight document AI deployment. Area names and codes are from the published criteria; column three is Ækora's view, not a Traficom position.
Osa-alue
Requirement codes
What changes when the models run on infrastructure you control
1. Esiehdot
EE-01, EE-02
The largest single change. With no external provider in the path, the answer to "who can be legally compelled to access this data" is your own organisation under Finnish law. EE-01 still has to be written.
2. Turvallisuusjohtaminen
TJ-01 to TJ-08
Unchanged and still yours: security principles, responsibilities, risk and incident management, continuity, classification and marking, compliance, supplier security. Local models write none of it.
3. Henkilöstöturvallisuus
HT-01 to HT-05
Applies to whoever administers the hardware. Vetting, confidentiality undertakings and need-to-know move to your own staff or a named integrator, not an unnamed global roster.
4. Fyysinen turvallisuus
FT-01 to FT-05
Becomes yours in full. A room you can walk into is auditable in a way a hyperscaler region is not. The cost: multi-layered physical protection is now your capital expense.
5. Tietoliikenneturvallisuus
TT-01, TT-02
Simplifies when inference runs inside an existing security zone: documents stop traversing a public network to reach a model. Network structure still has to be defended.
6. Identiteetin ja pääsyn hallinta
IP-01 to IP-03
Runs on your existing directory and admin practices, not a provider console. IP-03, hallintayhteydet, is where outsourced operating models most often come apart, including local ones.
7. Tietojärjestelmäturvallisuus
JT-01 to JT-05
Traceability, hardening, data separation, malware protection, transfer and deletion of protected assets. JT-03 separation is easier when the hardware is not shared with other tenants.
8. Salaus
SA-01 to SA-03
SA-03 allows lower-grade or unencrypted transfer inside approved physically protected areas where physical means give adequate protection. A local deployment can rely on that; a shared platform cannot.
9. Käyttöturvallisuus
KT-01 to KT-04
Capacity, backup and restore, and vulnerability management become your burden. This is the clearest running cost of local deployment, and the most underestimated.
10. Siirrettävyys ja yhteensopivuus
SI-01, SI-02
Open weights under a published licence are portable by construction. SI-02, destruction of data assets, is easier to evidence when you own the media.
11. Muutostenhallinta ja järjestelmäkehitys
MH-01, MH-02
Model, prompt and pipeline changes are system changes. Treat a new model version as a change requiring assessment, not a routine update.
PiTuKri v1.1 control areas mapped to a local open-weight document AI deployment. Area names and codes are from the published criteria; column three is Ækora's view, not a Traficom position.
Scope moves, it does not vanish
PiTuKri is explicit about the boundary. Its first diagram divides the processing environments: the cloud platform falls within PiTuKri assessment for administrative, physical and technical security; the customer system falls within Katakri and/or PiTuKri assessment as applicable; and the customer's other environments fall within Katakri assessment. The text says directly that Katakri 2015 can be used for the portions that are the customer's responsibility.
So the honest description is this: local deployment shrinks the part of your system that is a cloud service assessed against someone else's assurances, and grows the part that is your own environment assessed against your own evidence. Whether that is an improvement depends on whether you can carry the physical and operational areas. For a ministry with an accredited facility it usually is. For a small agency with no server room it is not.
Say it plainly, because vendors muddle it. Neither the EU nor Finland has a general data localisation mandate. Regulation (EU) 2018/1807 on the free flow of non-personal data points the other way, and PiTuKri cites it, noting it does not apply to location requirements imposed on national security and preparedness grounds. Demand for local processing in Finnish julkishallinto comes from classification rules, from the Tiedonhallintalaki (906/2019) and the decree on document classification (1101/2019), and from procurement. If a supplier tells you Finnish law requires data to stay in Finland, they have not read the criteria.
GDPR and the EU AI Act apply in Finland as elsewhere in the Union, and the Cybersecurity Act implementing NIS2 brought its obligations into force on 8 April 2025 (Traficom). None are localisation rules either. The GDPR side of a document pipeline is covered in GDPR-compliant document AI.
The models, and the Finnish-language problem
The technical half is less contested than it was two years ago. Open-weight document models now top the public parsing benchmarks. On OmniDocBench v1.6, PaddleOCR-VL-1.6 at 0.9B parameters scores 96.34% under Apache-2.0, MinerU2.5-Pro at 1.2B scores 95.75%, and GLM-OCR at 0.9B scores 95.22% under an MIT licence for the weights (Roboflow). Check the licences before you rely on them: MinerU's terms are Apache-2.0 with additional commercial conditions rather than clean Apache-2.0, and GLM-OCR's stack includes a component under a separate licence. Models of this size fit on a single professional GPU, which is what makes the physical security area affordable at all.
Why the benchmark number is not your number
OmniDocBench runs to 1,651 annotated PDF pages, and its documented language categories are English, Simplified Chinese and mixed English-Chinese (OmniDocBench). No Nordic language is represented, Finnish included. Finnish is morphologically rich, compounds heavily, and is poorly represented in document parsing evaluation. We found no credible published benchmark for Finnish-language document extraction at this model class, and we will not pretend a 96% English figure transfers.
The response is measurement, not assumption. Take 200 to 500 pages of your own documents, in the formats you receive them, and measure field-level accuracy against a human-labelled ground truth. If the documents are bilingual Finnish and Swedish, measure both. A comparison built on your own corpus is the only evidence worth putting in a procurement document. Our read on the model field is in the 2026 local OCR model comparison.
When this is not worth doing
Below roughly 250,000 pages a year, buying GPUs and standing up a controlled environment generally costs more than paying a cloud API, and the operational overhead is not recovered. At that volume a hosted service with a properly negotiated data processing agreement is usually the better answer, and you should be suspicious of anyone selling you hardware.
Classification overrides the arithmetic entirely. If the documents are turvallisuusluokitellut asiakirjat, cost is not the deciding factor and the page count is irrelevant. The question becomes whether the environment can satisfy the criteria at all, and a service where a foreign provider retains technical access to cleartext is unlikely to, whatever the price. PiTuKri's example conditions for higher assurance include classified data remaining physically inside Finland's borders throughout its lifecycle, and access to security-relevant system components being limited to Finnish citizens.
Classify first. Establish whether the documents are julkinen, salassa pidettävä, henkilötieto or turvallisuusluokiteltu, and at which level. Everything else follows from this.
Answer EE-02 for your current pipeline. Write down every physical location, every subcontractor, the jurisdiction and every actor that could be legally compelled to access the data. Most organisations find the problem here.
Decide who carries the physical and operational areas, FT and KT. If the answer is nobody, local deployment is not viable however good the models are.
Measure Finnish-language accuracy on your own documents before you commit to a model or a vendor.
Bring in Traficom early if you are heading towards viranomaisarviointi or hyväksyntä. Both processes are described in PiTuKri's third annex and neither is fast.
Where Ækora loses
Be direct, because the alternative wastes your time.
Ækora holds no PiTuKri assessment or accreditation. PiTuKri is applied to a service and its operating environment by a competent authority or an assessment body. It is not a badge a consultancy carries. Deploying locally changes which criteria apply and who is responsible for them. It produces no assessment result, and we confer no compliance on anyone.
No Finnish legal entity. AKORA SIA is a Latvian company. EU-domiciled resolves some questions, but it is not a Finnish supplier and some procurements treat that as disqualifying on its own.
No Finnish-language support commitment unless explicitly scoped and priced into the engagement. Working language is English.
No track record in Finnish julkishallinto and no history of taking a system through Traficom assessment. A Finnish integrator with that history may well be the better partner, and if your programme depends on the approval process, probably is.
No SLA and no instant start. This is a consulting and deployment engagement. Time to a first measured result on your own documents is weeks, and there is no product to try.
Where we are useful is narrower: working out whether local models can carry your document workload at the accuracy you need, sizing the hardware honestly, and building the pipeline that runs on it, described at local document models. If you have PiTuKri open and a pipeline to place, the first useful conversation is about classification and EE-02. Start it at contact.