Quick answer
Draft Annex 22 is the EU's proposed GMP guidance for data-trained AI and machine learning (AI/ML) models used in critical pharmaceutical manufacturing applications. It introduces AI-specific expectations for governance, validation, risk management, human oversight and ongoing monitoring, helping quality teams prepare for the safe and compliant use of AI in GMP-regulated environments.
Introduction
Ask a QA manager about their biggest challenge in 2026, and AI comes up fast. Cameras inspecting vials. Models flag deviations before anyone opens the batch record. This has moved off the slide deck and onto the shop floor.
A 2025 Define Ventures report, based on research that included executives from 16 of the top 20 pharmaceutical companies, stated that 85% of respondents from top-20 companies considered AI an “immediate priority.” Regulators had already noticed: the European Commission opened consultation on draft Annex 22 on 7 July 2025.
Brussels noticed. Draft Annex 22 is the first dedicated annex proposed for the EU GMP Guide on AI and machine learning in the manufacture of human medicinal products and active substances.
This article covers what Annex 22 is, why it exists, who it catches, and what you can start preparing now.
It is still a draft, so read it as a roadmap rather than settled law.
Key takeaways
What is Annex 22?
Draft Annex 22 is a proposed new annex to EudraLex Volume 4, the EU GMP guidance. The EMA GMDP Inspectors Working Group drafted it in cooperation with PIC/S, and the public consultation ran from 7 July to 7 October 2025.
There's no confirmed adoption date yet. EMA's current Inspectors Working Group work plan targets Q4 2026 for providing a final text to the European Commission. That is a target, not a confirmed publication or effective date; as of the date of publication of this article, no implementation period or effective date has been announced.
Draft Annex 22 was consulted alongside proposed revisions to Annex 11 and Chapter 4. Current Annex 11, dated 2011, continues to govern computerised systems generally until a revised text enters into operation. Draft Annex 22 adds model-specific expectations for data-trained AI/ML models embedded in critical applications with a direct impact on patient safety, product quality or data integrity.
Picture a vision system grading each tablet as accept or reject. That's exactly the kind of AI GMP application Annex 22 targets.
One point remains open. On 30 June and 1 July 2026, EMA held a workshop to gather expert input on regulatory pathways and guardrails for adaptive, probabilistic and generative AI. EMA says it is still considering consultation feedback. The current exclusion therefore remains the draft position; the workshop did not amend the July 2025 draft.
The scope: static, deterministic, and why generative AI is outside the current draft of Annex 22
Under the current draft, the model must first be a data-trained AI/ML model. Three further distinctions determine whether it can be used in a critical GMP application.
Static, not dynamic
The draft covers models whose parameters are fixed during operation. A retrained model can be introduced only as a controlled change, with an impact assessment and any necessary retesting. Continuously learning models are outside the current draft and should not be used in critical GMP applications.
That does not mean adaptive systems are impossible to validate in principle. It means the current draft does not yet provide a GMP pathway for them in critical applications.
Deterministic, not probabilistic
Using the draft's terminology, in-scope models must return identical outputs for identical inputs. Models that may return different outputs for the same input are outside the current draft.
Determinism helps repeatability, but reconstruction still depends on retaining the input, output, model version, configuration, and relevant audit-trail evidence.
Predictive, not generative
Prediction and classification are the examples used in the draft. Generative AI and LLMs are explicitly outside scope and should not be used in critical GMP applications. 'Agentic AI' is not a category defined in Annex 22. An agentic system that relies on an LLM, learns during use, or may produce different outputs for identical inputs would fall outside the current draft for critical applications.
The draft does not prohibit generative AI in non-critical GMP applications. It requires adequately qualified and trained personnel to remain responsible for ensuring that outputs are suitable for the intended use, and other GMP, data-integrity, security and confidentiality controls still apply.
Here is where it bites in practice. Suppose a vision system has inspected your vials reliably for a year, and a routine supplier update replaces the detection model underneath. The user interface may look unchanged, but the decision logic is no longer the validated version.
That is already a change to control under current Annex 11. Using draft Annex 22 as a gap-assessment baseline, the impact assessment should also determine whether model retesting is needed.
Which AI systems are currently in scope under draft Annex 22?
| AI system | Static | Deterministic | Predictive | Critical GMP | Draft status |
| Vision inspection | ✓ | ✓ | ✓ | ✓ | In scope |
| Tablet classification | ✓ | ✓ | ✓ | ✓ | In scope |
| Predictive maintenance | Usually | Usually | ✓ | Depends | Case-by-case |
| Chat GPT | ✗ | ✗ | ✗ | If critical | Outside scope |
| Microsoft Copilot | ✗ | ✗ | ✗ | If critical | Outside scope |
| Self-learning AI | ✗ | ✗ | ✓ | ✓ | Outside scope |
Why Annex 22 matters for AI in pharma
AI is already a strategic priority across much of large pharma. Define Ventures reported that 85% of respondents from top-20 pharmaceutical companies viewed AI as an immediate priority; its research included executives from 16 of the top 20 companies.
Until Annex 22 enters into operation, companies must apply the technology-neutral GMP requirements in Annex 11, supported where relevant by Annex 15 and ICH Q9(R1), together with established data-integrity guidance and industry good practice such as GAMP 5.
Annex 22 builds on familiar GMP principles, but it is not merely old wine in a new PDF. Its requirements for independent test data, staff independence, subgroup metrics, feature attribution, confidence thresholds and input-space drift are unusually specific for an EU GMP annex.
In April 2026, FDA gave "Inappropriate Use of Artificial Intelligence in Pharmaceutical Manufacturing" its own subsection in a warning letter after a manufacturer used AI agents to create specifications, procedures and master production or control records without adequate Quality Unit review. The case involved much wider CGMP failures, but the AI finding was explicit, not incidental.
Annex 22 vs EU AI Act
Annex 22 is not the only applicable framework. The EU AI Act applies horizontally to providers and deployers of AI systems in the EU. A “critical application” under Annex 22 is not automatically a “high-risk AI system” under the AI Act, so the two classifications and their resulting obligations should be assessed separately. AI-literacy obligations already apply, and certain transparency duties apply from 2 August 2026.
The message for QA is blunt. An AI output feeding a regulated decision is now something an inspector can ask you to defend, and you can't hand that accountability to a model.
New regulatory direction
In January 2026, EMA and FDA published joint Good AI Practice principles covering the medicines lifecycle, including manufacturing. They emphasise a clear context of use, risk-based governance, multidisciplinary expertise, data governance, performance assessment and lifecycle management. The principles are not GMP requirements or a substitute for Annex 11, but they point in the same general direction as draft Annex 22.
Annex 22 vs Annex 11 vs EU AI Act
| Annex 11 | Draft Annex 22 | EU AI Act | |
| Primary focus | Computerised systems | AI/ML models in critical GMP applications | AI regulation across sectors |
| Applies to |
All GMP computerised systems |
Data-trained AI in critical GMP | AI providers & deployers |
| Risk-based? | ✓ | ✓ | ✓ |
| AI-specific requirements | Limited | Extensive | Yes |
| Covers validation | ✓ | ✓ (AI-specific) | Partially |
| Covers supplier oversight | ✓ | ✓ | ✓ |
| Covers drift monitoring | - | ✓ | Sometimes |
| Legal status | Current GMP | Draft | Law |
The key AI governance principles behind draft Annex 22
Behind its technical clauses, the draft rests on principles any quality professional will recognise.
People and cooperation
Nobody selects, trains, validates and runs a model alone. It requires collaboration between process SMEs, QA, data scientists, IT and often external consultants, each with defined responsibilities and the appropriate level of access.
Documentation and supplier oversight
Whether the model is developed in-house or supplied, the regulated user should ensure that relevant documentation is available and reviewed. Buying an AI feature does not outsource the duty to understand and control its GMP use.
Risk first
Implementation should be proportionate to risk. A model that directly supports a batch disposition or in-process acceptance decision deserves more evidence and oversight than a lower-impact use. A genuinely non-critical use may fall outside Annex 22, but not outside all GMP controls.
Non-critical use is not prohibited by Annex 22, but that does not mean unrestricted use. The EU AI Act, GDPR where personal data are involved, confidentiality, cybersecurity and other applicable GMP controls may still apply.
Human oversight
As a general governance control, define the operator’s role, training and expected performance whenever a model informs a human decision. Annex 22’s specific monitoring and record-retention provisions apply where human oversight has been used to justify reducing the effort devoted to model testing.
Design the review step to counter automation bias. When results arrive quickly and at scale, rubber-stamping becomes the path of least resistance; a signature alone is not evidence of a meaningful review.
Some practitioners use 'decision integrity' as a non-regulatory shorthand for the ability to reconstruct the data, model output, context and human action behind a GMP decision. Useful concept, yes. Annex 22 term, not yet.
Governance is wider than model testing. It includes roles, policies, supplier controls, documentation, security, change management, monitoring, incident handling and periodic review.
What does Annex 22 mean for AI validation?
Draft Annex 22 applies to data-trained, static models with deterministic outputs, using the draft's terminology, in critical GMP applications. It supplements Annex 11 and uses a risk-based approach.
Annex 11 already requires lifecycle validation and periodic evaluation. Annex 22 adds AI-specific evidence before deployment and model-specific oversight during operation.
Before go-live
Before an AI model is used in a critical GMP application, the draft expects the following to be defined, approved and tested with process-SME involvement:
- Intended use, documented and approved. Describe the tasks the model will assist or automate, the process context, the full input sample space including common and rare variations, limitations, and possible erroneous or biased inputs. Approve this before acceptance testing.
- Metrics and acceptance criteria. Select case-appropriate metrics and set acceptance criteria before testing. Where the model replaces an existing process, the criteria should be at least as high as the performance of that process, which means the baseline must be known and comparable.
- Subgroup performance, not just the average. Where applicable, divide the input sample space into meaningful groups and assess performance for each group separately.
- Test data that are representative, statistically adequate and independent. The final test set must not be used during model development, training or model validation. Where it is separated before training, personnel involved in developing and training the model should never have had access to it. Access controls and audit trails should enforce that separation, no copies should exist outside the controlled repository, and the identity, timing and number of test uses should be recorded. Staff independence should also be maintained or, where that is impossible, supported through the proposed four-eyes arrangement.
- Test-data quality and processing. Use high-quality test data with verified labels; predefine and justify preprocessing, and fully document any cleaning or exclusions. Avoid AI-generated test data unless fully justified, and keep final physical test objects separate from those used for training or validation unless feature independence is proven.
- Controlled test execution. Use an approved test plan and script, demonstrate generalisation to new data, investigate deviations and failed acceptance criteria, and retain the actual test data and supporting records. The draft also expects testing to detect possible overfitting or underfitting.
- Explainability during testing: For critical applications, capture and record the features that contributed to a classification or decision. Where applicable, use feature-attribution methods or heat maps, then have the relevant experts review whether the model is relying on appropriate features.
- Confidence and thresholds. Where applicable, log the model’s confidence score for each prediction or classification during testing. Models used to predict or classify data should have an appropriate decision threshold. If confidence is very low, consider routing the result to an “undecided” outcome rather than issuing a potentially unreliable prediction or classification. The draft does not explicitly require operational confidence logging for every live output, so define that control according to intended use, risk and Annex 11 data-integrity needs.
Terminology note: Annex 22 uses “validation dataset (in AI)” for a dataset used during model development to optimise the trained model. It is distinct from GMP validation of the computerised system or application.
After go-live
Approval is not the finish line. The model, its host system and the process it supports must remain controlled and monitored in operation.
- Control every change. Any change to the model, its host system, the process or physical inputs should be documented and impact-assessed to determine whether retesting is needed. A decision not to retest must be justified.
- Configuration control. Keep a controlled record of the exact model and configuration that were approved, and use effective measures to detect unauthorised change.
- Performance and input-space monitoring. Monitor both the model's defined performance metrics and whether incoming data remain within the approved input sample space. The former can detect changes in the system or environment; the latter detects data drift.
- Human-review records. Where the model feeds a human decision and reduced model testing has been justified on that basis, define the operator's role, training and procedure, monitor performance and retain review records. Depending on criticality and the level of model testing, review or testing of every output may be necessary. It is not a blanket requirement for every AI system.
What this changes in practice
The shift is not from one-off to lifecycle validation; Annex 11 has required lifecycle control for years. The shift is from generic computerised-system evidence to additional model-specific evidence and monitoring.
That means more record-keeping, not less:
- the approved intended use and limitations
- the test datasets, test scripts and acceptance results
- proof the test set was independent and access-controlled
- feature-attribution evidence and expert review from acceptance testing
- confidence scores and thresholds during testing
- how it’s performing now
- human-review records where the approved process requires them
The payoff isn't lighter paperwork. It's fewer unknowns. At review or inspection, you can show that the model remains fit for its intended use, and anyone who asks can trace exactly what was done and why.
Recommended learning:
The problems that Annex 22 is designed to prevent
Behind all the rules sit a handful of real-world failures quality teams should watch for:
- Made-up information. Generative tools like chatbots can produce confident, polished text that’s simply wrong, like citing a guideline that doesn’t exist.
- Results you cannot reproduce. If identical inputs can produce different outputs, the behaviour is difficult to validate and reconstruct unless the exact input, output, model version, configuration and execution context are retained.
- Spurious decision logic. The model appears accurate but relies on irrelevant features, such as lighting, background or equipment artefacts, rather than the true quality attribute.
- Uncontrolled change and drift. A model, camera, process, product or input distribution changes without impact assessment, retesting or timely detection.
Hallucination, non-repeatability, confidential-data leakage and cybersecurity are also important AI risks.
The current draft addresses hallucination and non-repeatability mainly through its restrictions on generative and large-language models, dynamic models and outputs that may vary for identical inputs in critical applications.
Its testing, change-control and monitoring provisions address other model risks.
Data integrity, access control and system security remain governed through Annex 11 and the pharmaceutical quality system; where personal data are processed, GDPR applies separately. Annex 22 does not replace those controls.
How to prepare for Annex 22 today
While compliance cannot be claimed for an ineffective draft, do not wait for the final text to build your approach. A program based on what inspectors already ask about computerized systems will only require marginal adjustments rather than a full redesign. Delaying merely causes a six-month standstill and saves very little.
You can strengthen controls already required under current GMP and use the July 2025 draft as a gap-assessment baseline while monitoring changes to the final text.
Map your AI tools first. Current Annex 11 already requires an up-to-date inventory of relevant systems and their GMP functionality. Extend it to record the model's intended use, owner, supplier, version, whether it is data-trained, static or adaptive, reproducible or non-reproducible under identical inputs, generative or LLM-based, embedded in another system.
Rank each AI tool by its impact on GMP decisions. Rank every AI system that touches a GMP decision by how much damage a wrong output could cause. Score them on decision criticality, autonomy vs. human review, static/reproducible vs. adaptive, and whether they feed an official GMP record. Let that tier set the evidence bar: high tier means intended use, performance metrics, test data, oversight and monitoring; low tier means a documented rationale and basic controls. It's the GAMP 5 principle (effort proportionate to risk) applied to AI. That keeps scrutiny on the systems that need it and gives you a defensible answer when an inspector asks why.
Put an AI-use policy under document control. This is a governance recommendation, not an explicit Annex 22 clause. Define approved tools and use cases, prohibited data, required review, record-retention expectations, escalation routes and what may enter an official GMP record.
Make human oversight a real control. Assign appropriately trained and authorised reviewers, give them time and information to challenge the output, and define when review is per output, by exception or periodic.
Question and contract your vendors. Ask whether the model changes after deployment, how versions are identified, what test evidence and limitations exist, where data are processed, how incidents and updates are communicated, what audit trails are available, and whether you can test changes before release. Vague answers should trigger additional assessment.
Use an eQMS for what it is good at. It can control policies, training, supplier assessments, deviations, CAPA and change records. It may not be the right repository for model code, locked test datasets, feature-attribution evidence or live performance metrics unless those capabilities are integrated and validated. Annex 22 does not require an eQMS, and buying one does not outsource model governance.
What inspectors will expect
| Area | What inspectors may look for |
| Governance | Defined roles and responsibilities |
| Validation |
Intended use, metrics, test evidence |
| Documentation | Supplier documentation available |
| Risk management | Risk assessments proportional to impact |
| Monitoring | Ongoing performance and drift |
| Human oversight | Meaningful review process |
| Change control | Controlled updates and retraining |
Conclusion
Draft Annex 22 does not replace established GMP practice. It builds on intended use, documented evidence, change control, quality risk management and accountability, then adds AI-specific controls around test-data independence, subgroup performance, explainability, confidence and drift.
The July 2025 draft is deliberately restrictive about continuously learning or adaptive-in-use models, non-repeatable models and generative models in critical GMP applications. EMA has gathered expert input on whether guardrails and other mitigations could support broader use, but no revised text or policy position has yet been published.
Map the systems, define intended use and criticality, close supplier and data gaps, and plan how each GMP-relevant model will be tested and monitored, proportionate to its intended use, criticality and risk. Do that before the final text lands, not after.
For a wider look at how quality leaders are approaching AI and governance, Scilife's Global Quality Outlook report is a good next read.
Looking to build a stronger foundation for AI governance with Scilife's Smart Quality QMS?
Explore how a connected eQMS helps you stay compliant, inspection-ready and prepared for what's next.
FAQ
Is draft Annex 22 legally binding yet?
No. It is a consultation draft and is not yet in operation. The consultation ran from 7 July to 7 October 2025. EMA's current work plan targets Q4 2026 for providing a final text to the European Commission, but no publication or effective date has been confirmed.
Does draft Annex 22 ban generative or agentic AI in pharma?
Not across pharma. The current draft says generative AI and LLMs should not be used in critical GMP applications. It does not define 'agentic AI'. An agentic system that relies on an LLM, learns during use, or may produce different outputs for identical inputs would fall outside the current draft for critical applications. Non-critical GMP use is not prohibited, but qualified and trained personnel must ensure the output is suitable and other controls still apply.
How is draft Annex 22 different from Annex 11?
The current Annex 11 applies to computerised systems generally. Draft Annex 22 adds model-specific requirements for certain data-trained AI/ML models in critical applications. A separate draft revision of Annex 11 was consulted at the same time, but it is not yet the effective Annex 11.
Who is accountable for an AI system's output under draft Annex 22?
Under the combined GMP framework, the regulated user remains accountable for the system’s GMP use and supporting evidence, including where supplier documentation or testing is leveraged. Draft Annex 22 requires responsibilities to be defined and relevant documentation to be available and reviewed. It does not require the Quality Unit to approve every model output.
What’s the single most useful thing to do now?
Inventory and classify your current use of AI. You can't govern, validate or risk-assess what you've never mapped.
Does Annex 22 apply to AI tools you're already using, or only new ones?
The draft does not state a grandfathering rule or transition period. Existing in-scope systems should be gap-assessed, but the final implementation arrangements are not yet known.
What actually counts as "critical" under Annex 22?
The draft uses 'critical applications with direct impact on patient safety, product quality or data integrity.' Assess the intended use and decision pathway, not merely the department or software label. A non-critical use may fall outside Annex 22, while Annex 11 and other GMP controls can still apply.
Does Annex 22 apply to AI features buried inside software you already own, like your LIMS or MES?
Yes, if the embedded function is a data-trained AI/ML model used in an in-scope critical application. The supplier's marketing label is irrelevant; what matters is how the function works and the GMP decision it supports.
Will inspectors start asking about AI before Annex 22 is finalised?
They can already ask. Current Annex 11, the PQS, supplier oversight, validation, data-integrity and change-control requirements apply today. The 2026 FDA warning letter shows that US authorities can enforce existing CGMP requirements against inadequately controlled AI-assisted work. EU inspectors can separately rely on current Annex 11, the pharmaceutical quality system, validation, supplier-oversight, change-control and data-integrity requirements.
How does Annex 22 relate to ALCOA+?
It complements ALCOA+ rather than replacing it. Annex 11 and established data-integrity guidance govern trustworthy data and records; draft Annex 22 adds model-specific controls intended to make the supporting AI evidence credible.
Is Annex 22 the same as the EU AI Act?
No. Annex 22 provides GMP expectations for certain AI/ML models in critical manufacturing applications. The EU AI Act is a horizontal legal framework covering AI providers and deployers more broadly. “Critical” under Annex 22 and “high-risk” under the AI Act are separate classifications, and both may need to be assessed.




