AI in GxP that has already passed inspection.
Almost anyone can advise on AI governance in theory. Very few have put AI into live GxP quality and taken it all the way through a health-authority inspection. That is the work this practice is built on, and it is the difference between a slide about AI and a system an inspector accepts.
Decision integrity, the new layer on top of data integrity
For twenty-five years the discipline was data integrity: ALCOA++ on every record. AI adds a second layer. Every decision a model makes or shapes has to be traceable, explainable and defensible: what went in, what came out, who reviewed it, what they decided. Inspectors are already asking for exactly that record, and the FDA issued its first warning letter for uncontrolled AI in a GxP process in April 2026.
AI drafting deviations, assembling investigations and answering quality questions, so your people spend their judgment where it counts.
Every use case classified, validated to its risk tier, and monitored for drift, so AI never quietly becomes a finding.
AI tools running in live GxP quality, with the decision-integrity evidence an inspector can trace from input to disposition.
Hearing “validation, not AI” from your own quality function? That pushback is worth taking seriously, and it changes the sequencing rather than the destination.
Governed from the regulatory spine up
Governance precedes validation, and classification precedes both. The AI Governance Stack is the structure I built to put AI into regulated quality and bring it through inspection, six layers from the regulatory spine to live monitoring, with a four-tier risk model that decides how much rigour each use case carries.
Monitoring
Drift, performance and human oversight, watched after go-live because AI does not stay put.
Validation master plan
Model-specific and sized to the risk tier, not a one-size template.
Governance operating model
An AI Governance Board, a RACI, decision rights and reclassification triggers.
Four-tier risk classification
The most consequential hour in a model's life. Get the tier right and every control follows.
Reference architecture
Data, model, workflow, record and audit trail, mapped end to end.
Regulatory spine
Annex 11, the draft Annex 22, 21 CFR Part 11, FDA CSA and the FDA-EMA principles.
The AI Governance Stack
The six-layer framework: regulatory spine, reference architecture, classify and risk-tier, governance operating model, validation master plan, and live monitoring.
See the framework →AI-enabled QMS
Where AI goes in your quality system: the highest-volume, lowest-judgment steps, every tool validated, audit-ready and reviewed by a human who owns the output.
Explore the service →AI in GxP governance
Annex 22-aligned controls, the four-tier risk model, reclassification triggers, and the decision-integrity record inspectors ask to see.
Explore the service →How do you classify AI risk in GxP?
By consequence of error, never by model complexity. The four-tier model classifies each AI use case on what happens when the output is wrong, and every downstream decision follows from the tier: test depth, monitoring cadence, change-control rigour, even the seniority of the approver. A logistic regression that classifies deviations carries more validation weight than a deep neural network that draws a dashboard.
A human reads the output for awareness. No GxP record is created from it and no quality decision turns on it. A KPI dashboard or trend chart. Validation is light: intended use, data sources, a periodic review.
The AI recommends, and a qualified human explicitly approves before any GxP action. Deviation triage, chromatogram review flags. Structured validation: scripted OQ against pre-agreed criteria, PQ with real users, and the override always logged.
The AI decides the path the workflow takes, with a human override on the exception path. CAPA routing, change-control categorisation. Full IQ, OQ and PQ, subgroup acceptance criteria, continuous monitoring, mandatory explainability.
The AI acts with no mandatory human approval in the critical path. Truly autonomous systems for critical GMP decisions remain rare in 2026, and when in doubt you classify Tier 3 and revisit with evidence.
Four questions settle the tier in one structured hour, with the process owner, the data scientist, the QA lead and the validation lead in the same room. Does the output influence a GxP decision at all? If it were wrong on 10% of records for 6 months without detection, would patient safety, product quality or data integrity be at material risk? Does a qualified human review and approve every output before a GxP action? And can a human override the decision through an exception path that actually exists in the deployed workflow, rather than on the architecture diagram?
That 10% question is the one I lean on when a room is undecided, because it matches how AI actually fails. Models rarely collapse; they run 5 to 15% wrong over a sustained period, and no single record screams about it. The tier also has to stay honest after go-live: a broadened scope, a changed workflow, or an override rate that quietly falls toward zero all mean the classification needs a fresh look. A Tier 2 system whose reviewers stop overriding has become a Tier 3 system being signed by a rubber stamp.
The full treatment, with the classification record template, the reclassification triggers and two worked examples either side of the Tier 2/3 boundary, is Chapter 4 of the book.
How do you keep validated AI valid?
Traditional CSV ends at PQ sign-off, and the validated state holds until someone deliberately changes the system. AI breaks that assumption. A model can pass PQ on Tuesday, meet a changed input format on Wednesday, and be deciding on data it was never trained on by Friday, with no error logged anywhere. Monitoring is where the validation continues, and the draft Annex 22 treats it that way: continuous watch on performance, inputs, outputs and human oversight, with documented thresholds and a documented response when one is breached.
In practice the monitoring stack has five layers, each answering a different question. Performance metrics on a labelled sample of production records tell you whether the model is still right. Input-distribution checks catch the earliest signal, because the inputs shift before the performance does. Output-distribution checks catch the model responding to that shift. The override rate tells you what your human reviewers really think, and a sampled review of what drove each decision keeps the explainability honest.
Two of those signals deserve a special watch. An override rate climbing above roughly 30% says the model may no longer be fit. An override rate falling below 1% says the human review has probably become a formality, which quietly changes the system's risk tier. Both readings should trigger an investigation with a named owner and a close date, because a threshold breach that produces no recorded response is a control that does not exist.
The full metric stack, representative thresholds per tier, and the drift response ladder are Chapter 11 of the book; the monitoring layer of the AI Governance Stack explainer is free.
From an honest read to AI in control
Take the free AI-in-GxP Readiness Index, or we inventory and classify every AI touchpoint, including the ones inside vendor products.
A fixed-scope review of your highest-risk use cases against the four-tier model, ending in a board-ready view of risk, effort and options.
Stand up the governance board, write the model-specific validation plans, and design monitoring before go-live, not after.
Keep the tier honest as autonomy creeps, monitor for drift, and stay inspection-ready as a steady state.
Start with the readiness index Request a conversation
Weighing outside help? Read how to choose an AI validation partner, including the questions to ask anyone you shortlist, me included.
