What this page continues. The parent article's §11 Validation Strategy settled which control applies to which component of an agentic platform: deterministic rules validated as GAMP 5 Category 5; the generative path held by grounding, deterministic checking and human review; statistical models under an analytical-method lifecycle; and a Part 11 signature on every regulated output. It closed on a warning. Teams that document this posture before the first agent enters a GxP workflow will set the standard, and teams that retrofit it will spend twice as long. This page covers what §11 left open: how you prove it. It starts from the validation lifecycle every CSV team already runs (FDA's software-validation principles, CSA, GAMP 5, Part 11 and Annex 11), shows where AI breaks that lifecycle's assumptions, and runs it on two production systems from the parent article, ProtoCheck and Doscierge. It then tests the result against thirteen questions that practitioners actually asked.
How the argument is built: four layers, each resting on the one before, and a close that turns them into action
The Trust Paradox
Why does retrofitting cost twice as much? Because of a tension every Quality leader recognises:
The promise
AI promises speed
Drafts in minutes instead of days. Signals found before the annual report. Deviations triaged without a queue. Every vial inspected at line speed.
VS
The constraint
GxP leaves no room for uncertainty
Every regulated decision needs a specification, a test, an audit trail and an accountable signer. A model that learns from data has no specification, performs statistically, and can decay without anyone touching it.
The resolution, and the spine of this page
You don't validate trust into a model. You validate a performance claim for one context of use, then you keep proving it.
Validate the controls around the model for a specific context of use, not "the model" in the abstract. Model metrics are evidence that the controls work, not a certificate for the model in general.
Every section that follows establishes, or keeps proving, one part of that claim:
Bottom Line
Practitioners know CSV. What they lack is AI-specific controls, operational templates and worked examples.
The industry does not lack CSV knowledge. It lacks a way to turn CSV principles into AI-specific controls: how to set acceptance criteria against a measured human baseline, how to tell drift from a change in intended use, what to do when a SaaS vendor quietly adds AI to a validated Category 3/4 system, and which documents a small company actually needs. The regulations have now moved far enough to answer most of this. Draft EU GMP Annex 22 (July 2025) sets clear rules for static, deterministic ML in critical GMP use. The ISPE GAMP Guide: Artificial Intelligence (July 2025) sets out the lifecycle. FDA's January 2025 draft guidance defines a seven-step credibility framework built around context of use. What the industry is missing is operational templates and worked examples. That gap is the market opportunity.
§4.3
Draft Annex 22: the baseline rule
Acceptance criteria at least as high as the process replaced, so that performance must be known. Most sites have never measured it (§5).
7 / 8
session questions asked "show me how"
Not "tell me what". Practitioners already accept that AI needs validating. They are stuck on execution (§2).
8
lifecycle controls, one AI extension each
GxP assessment, risk, version control, data, design, registry, monitoring, periodic review, mapped to templates T1–T10 (§8).
5
metric families in one site policy
Inputs, outputs, human-in-the-loop, system/vendor, data integrity, with minimums by tier and a four-level response ladder (§13).
Six findings, and each one is a component of the claim
1
The unit of validation is a performance claim for one context of use, not the model.
The same model is low risk as a drafting aid and critical as a release gate; only the use, the population of inputs and the surrounding controls make it validatable.
Supportdraft Annex 22 §3.1 (intended use and input sample space); FDA January 2025 draft, Steps 1–3; §1 and §8 below.
2
The human baseline is the missing input, and draft Annex 22 makes it a regulatory expectation.
Its §4.3 requires model acceptance criteria "at least as high as the performance of the process it replaces", which means that performance must be known. Outside visual inspection (USP <1790>, Knapp–Kushner), most sites have never measured it.
Support§5 (Q6), which gives published anchors and a measurement protocol.
3
A locked model does not drift; the world it sees does. So monitoring watches inputs and humans, not only accuracy.
Ground-truth labels are rarely available in production, so the observable signals are input-distribution scores, the "undecided" rate, override and edit rates in both directions, and periodic reference checks that supply ground truth. In CPV, the first question is whether the process moved or the model's inputs did.
For most companies AI arrives inside validated SaaS, so category creep is a supplier-management and change-control problem before it is a model-validation problem.
GAMP categories describe the product; the AI function inside it must be classified on its own: critical or not, static or dynamic, deterministic or probabilistic.
Support§17 (Q7) program and vendor clauses; §4 (Q11a) decision table.
5
Part 11 and ALCOA+ do not change for AI. What each control has to capture does.
The model is a new actor, the prompt and retrieved context are new inputs, and a probabilistic output cannot be regenerated on demand. The record must therefore hold the output as generated, its model and configuration version, its sources, and the human's accept/modify/reject as separately attributed events. No product "meets Part 11".
Support§15 (Q9) clause by clause; §16 (Q12) principle by principle.
6
CSV is not obsolete. Each of its eight lifecycle controls still applies, and each needs one AI extension.
GxP assessment, risk, version control, data, design, registry, monitoring and periodic review map one-to-one onto an AI lifecycle SOP set (T1–T10). Draft Annex 22 also excludes dynamic and generative/probabilistic models from critical GMP use, so the practical design for critical decisions is dual-path: deterministic where the decision is made, generative where a qualified human reviews.
Support§8 (Q8) and the at-a-glance matrix in §8.0; parent article §10–11.
Strategic Planning Assumptions
Five planning assumptions, with probability and substantiation
By end-2027Probability 0.70
By end-2027, EU GMP Annex 22 will be in force in a form that keeps critical GMP use limited to static, deterministic models and requires a measured baseline for the process replaced.
Substantiation
The consultation draft (7 July 2025) scopes critical use to static, deterministic models and sets the "no decrease" rule in its §4.3; the consultation closed 7 October 2025 with about 1,300 comments; the EMA inspectors working group targets a final text in Q4 2026; an EMA workshop on 30 June–1 July 2026 discussed whether dynamic and probabilistic models could be addressed. The scope may widen for non-critical use; the baseline expectation is unlikely to be dropped, because it follows from the existing Annex 11 and validation principle of demonstrating fitness for intended use.
Through 2028Probability 0.75
Through 2028, the majority of GxP AI exposure at small and mid-size companies will arrive as vendor-delivered features inside Category 3/4 SaaS, not as in-house models.
Substantiation
The platform vendors that hold GxP records are shipping AI features in routine releases (the parent article records Veeva shipping AI agents for PromoMats in December 2025 and Agentforce Life Sciences reaching general availability in October 2025); the questions from the session (Q7, Q9) came from exactly this exposure; a company with a 3–10 person Quality team has no capacity to build and validate its own models, but receives release notes every month.
By 2027Probability 0.65
By 2027, FDA's January 2025 draft on AI in regulatory decision-making will be finalised with the seven-step, context-of-use credibility framework substantially intact, and it will become the default template for clinical AI evidence.
Substantiation
The draft's framework (question of interest, context of use, model risk = influence × consequence, credibility plan, execution, results, adequacy) is reinforced by the FDA–EMA Guiding Principles of Good AI Practice in Drug Development (14 January 2026), which foreground context of use, data governance and lifecycle monitoring; the ISPE GAMP AI Guide (July 2025) follows the same lifecycle logic.
By 2028Probability 0.70
By 2028, dual-path design (deterministic component carries the critical decision; generative component drafts and explains under human review) will be the default validation posture for GxP-critical AI, ahead of any regulator explicitly endorsing it.
Substantiation
Draft Annex 22 excludes generative and dynamic models from critical GMP use, and the EMA is only "considering" whether to address them; FDA's CSA guidance (final 24 September 2025, device scope, applied to pharma by analogy) supports risk-proportionate assurance of the deterministic components; the parent article's §11 already applies this posture across research, development, manufacturing and medical use-cases.
Through 2027Probability 0.60
Through 2027, the most common finding in Annex 22 readiness reviews will be an unmeasured human baseline, not a weak model.
Substantiation
Annex 22 §4.3 makes the baseline a precondition for setting acceptance criteria; the only mature pharma precedent for measuring it is manual visual inspection (USP <1790>, Knapp–Kushner); the clinical data-processing literature (Garza et al., 2025) shows how wide human error rates are by method, which is why a local measurement, not a published number, is needed. This is a judgement from the questions asked in the session (Q6, Q2), not a survey result.
On the probabilities. The Strategic Planning Assumptions are the author's judgements from the sources cited; the probabilities are not survey results (§19).
Recommendations
By persona: the next 90 days, and what not to accept
For the Head of Quality / QA
Next 90 days
For the Head of Quality / QA
Issue an AI policy and inventory (§17, elements 1–2); tier every AI function, including vendor features; make the human baseline (§5) a required input to any Tier 3 validation plan; adopt one site monitoring policy (§13) so every AI function reports the same five metric families.
Do not
Accept "the vendor validated it" or "a human reviews everything" as controls until the vendor evidence covers your context of use and the reviewer's catch rate has been measured (§7 "Human review", §5).
For the CSV / validation lead
Next 90 days
For the CSV / validation lead
Rewrite one AI requirement as a performance claim (metric, subgroup, threshold, baseline, confidence level; §1 last row); build the locked, independent test set before any model is touched (§8.4, T4); make the eight-control matrix (§8.0) the table of contents of your AI validation package; treat retraining, prompt edits and vendor model swaps as changes with a pre-approved envelope (§14, T9).
Do not
Run a single PQ and call the model "validated"; the validated state for AI is a claim that monitoring keeps true (§11, §13).
For Manufacturing / MS&T
Next 90 days
For Manufacturing / MS&T
For any CPV, PAT or vision AI, decide which component carries the critical decision and keep it static and deterministic (§4 decision table, §8.1); measure the manual process it replaces (§5); write the predictable changes (new supplier within spec, recalibration, NOC refresh) into a T9 envelope before go-live (§12, §14.2).
Do not
Retrain a CPV model to make an excursion disappear. A real process signal is what CPV exists to catch; route it to deviation and change control, not to model retraining (§11, FAQ 4).
For Clinical / Regulatory
Next 90 days
For Clinical / Regulatory
Validate the CSR or document-drafting agent under the seven-step credibility framework with a demonstrated (not assumed) human-review effect (§6); store every AI output as generated with model, prompt, sources and reviewer action (§15, §16); set the signature meaning to cover AI-assisted content (§15, §11.50/§11.70 row).
Do not
Let an LLM grade its own output, or let conversation memory carry state between cases (§3 rows 9–10, §9).
For Small-biotech QA (50–300 employees)
Next 90 days
For Small-biotech QA (50–300 employees)
Run the 90-day starter plan (§17): policy, inventory sweep, tiering, vendor AI questionnaire to the top five GxP SaaS suppliers, untriaged AI features switched off; adopt the Tier 2 one-page assurance plan for human-reviewed drafting uses (§9); register AI-assisted authoring once, not per document (§10).
Do not
Build a separate AI governance system. Extend the GxP system inventory, the supplier questionnaire, the quality agreement and the periodic review you already have (§17, elements 2, 4, 8).
For the IT / AI platform owner
Next 90 days
For the IT / AI platform owner
Model each AI component (model, prompt/config, retrieval index, locked test set, tool server, monitoring config) as a configuration item under the business application (§14.1); give each agent its own identity and least-privilege tool allowlist (§15 §11.10(d) row); pin vendor model versions and manage the deprecation schedule as a planned change (§4, §14.5 S2); route monitoring action-limit breaches into Incident and Problem records (§14.4, §13).
Do not
Hide the model, prompt and index inside one application CI. Every change then looks either trivial or total, and change-level risk-based validation becomes impossible (§14.1).
How to read this page
Pick your lens. The contents highlight your path.
This is a long page because the questions came from five different jobs. Read straight through for the full argument, or take the role path. Times assume about 250 words a minute. Select a role and the table of contents highlights the sections to read first and then, and estimates the reading time. Your choice is remembered on this device.
Detail level
for the Quality / CSV lens: the reframe, the misconceptions that change a package, the pitfall map, the eight controls, Part 11 clause by clause. The author's estimate: 40 min first pass; 30 min second.
for the Manufacturing / MS&T lens: locked vs. adaptive, the human baseline, the CPV signal agent as the running example, drift in CPV terms, the AI camera. The author's estimate: 40 min; 25 min.
for the Clinical / Regulatory lens: CSR drafting under the FDA credibility framework, clinical data-processing error rates, Part 11 and ALCOA+ with an AI contributor. The author's estimate: 35 min; 25 min.
for the Small-biotech QA lens: a 90-day governance program, a proportionate rule for AI-assisted authoring, vendor questions, monitoring minimums by tier. The author's estimate: 30 min; 25 min.
for the IT / AI platform lens: CMDB and change models, agent identities, configuration items, monitoring routed into ITSM. The author's estimate: 35 min; 25 min.
A 90-day governance program, a proportionate rule for AI-assisted authoring, vendor questions, monitoring minimums by tier
IT / AI platform owner
§14, §15, §9
§8.3, §8.6, §13, §17 (release-note triage)
35 min; 25 min
CMDB and change models, agent identities, configuration items, monitoring routed into ITSM
Whatever lens you bring, the argument has one spine: in CSV you validate a specification. In AI you validate a claim about performance, on a defined population of inputs, against a known baseline, and then you keep proving that claim holds.
Every audience and practitioner question keeps its Q number as a tag under the section heading, so external references (templates, posts, the parent article) still resolve.
Full tableQuestion index: Q1 to Q13 with origin and the section that answers each13 rows
What is being claimed, for which use, by which kind of model.
Bridge. The Trust Paradox is resolved by changing the unit of validation from "the model" to "a performance claim for one context of use". Part A builds that claim in the order a validation team would: the CSV/CSA lifecycle you already run and the assumptions AI breaks (§1.1), the parent article's rule for which control applies to which component (§1.2), and both of them run on two production systems (§1.3). It then turns to what practitioners asked, with each question pinned to a stage of that lifecycle (§2), the common views that get the claim wrong (§3), and the kinds of model the claim can cover in critical and non-critical GMP use (§4). Parts B to F set the bar for the claim, build its controls, keep it true in operation, record it, and turn it into action.
§1 · Part A · Layers 1–3 · The standard, the architecture, show me
In CSV you validate a specification. In AI you validate a claim.
The reframe · claim definition
Spine
Layer 1 · The standardThis section starts from the lifecycle every CSV team already runs and defines the claim. Classic CSV (GAMP 5, Annex 11, Part 11) rests on assumptions that deterministic software meets and AI does not. Every audience question traces back to one of the rows below, and the right-hand column is the control that restores each assumption.
CSV assumption
Holds for rules-based software
What changes for AI
Control that restores it
Behavior is specified: URS → FS → DS, and code implements the spec
Yes
Behavior is learned from data. The "spec" is partly the training set
Intended-use and input sample space definition (Annex 22 §3.1); data becomes a configuration item (§8.3, §8.4)
Same input → same output
Yes
ML: yes once frozen. LLMs: not guaranteed
Annex 22 scope is static + deterministic only; LLMs go non-critical with human-in-the-loop, or dual-path (§4, §9)
Testing proves correctness
Pass/fail per requirement
Performance is statistical: sensitivity, specificity, F1, with confidence intervals per subgroup
Pre-approved metrics and acceptance criteria, owned by an SME (Annex 22 §4.1–4.2); test set sized for statistical confidence (Annex 22 §5.2); the baseline in §5
Validated state lasts until the system changes
Yes
Performance can decay without anyone touching the system: the world changes (drift)
Performance and input-drift monitoring (Annex 22 §10.3–10.4); §11, §13
Change = code or config change
Yes
Change also includes new training data, new labels, a new supplier's container, new lighting, a vendor model swap
Change control covers model, system, process, and the physical inputs (Annex 22 §10.1); §12, §14
Test data is just test data
Mostly
Test data leaking into training inflates performance and invalidates the result
A Cat 3 SaaS product can ship an AI feature overnight. The category no longer tells you the risk
Classify the AI function, not only the product (§17, §8.1)
The human reviewer is the control
Yes
Automation bias: reviewers approve AI output they would have challenged coming from a colleague
Monitor human-in-the-loop performance "like any other manual process" (Annex 22 §3.3); measure acceptance and edit rates (parent G2); §5, §13
Each requirement is granular, testable, and risk-rated on its ownGAMP 5 2nd Ed.; FDA CSA, 2025
Yes: one requirement → one or more test scripts → pass/fail
A single AI requirement is a performance claim over a population. It cannot be broken down further into pass/fail steps
Extend the requirement-attribute record (ID, rationale, risk rating, source, dependencies) with metric, subgroup, threshold, baseline, and confidence level. That record is the row T6 traces to
The one-sentence reframe for a Quality audience
In CSV you validate a specification. In AI you validate a claim about performance, on a defined population of inputs, against a known baseline, and then you keep proving that claim holds.
1.1 From CSV to CSA to AI assurance
AI assurance is not a third methodology. It is CSA's risk-based critical thinking, applied to a component whose behavior is learned rather than specified. The same lifecycle (concept → project → operation → retirement, as in GAMP 5 2nd Ed. and the ISPE GAMP AI Guide) carries through.
Aspect
Classic CSV
CSA (risk-based)
AI assurance (adds)
Question asked
"Does the system do what it purports to do?"
"Where could failure hurt patient, product, or data, and how much assurance does that need?"
"How reliably does it perform, for this context of use, against a known baseline, and does that still hold?"
What drives effort
Document set, applied uniformly
Risk of each feature FDA CSA, 2025
Model influence × decision consequence (FDA Jan 2025 draft, Step 3) and Annex 22 criticality (§8.2)
Testing
Scripted IQ/OQ/PQ
Scripted for high risk; unscripted or exploratory for low risk
Statistical testing on an independent, stratified test set; N-run testing for probabilistic output; trajectory and red-team tests for agents (§9)
Acceptance criteria
Expected result per step
Same, scaled by risk
Pre-approved metric per subgroup with confidence interval, no worse than the measured baseline (Annex 22 §4.2–4.3; §5)
Vendor evidence
Often retested
Leveraged where risk allows, but owned by the regulated company
Also: model card, training-data provenance, model-change notification, whether your data trains shared models (§17)
The controls around the model, for a defined use: grounding, deterministic checker, confidence gating, human review, monitoring, change control. Model metrics are evidence that those controls work, not a certificate for the model in general
1.2 Where this fits: validation strategy for the whole platform
Layer 2 · The architectureThis page continues the parent article's §11 Validation Strategy, which sets the posture for every use-case on the enterprise platform. That section starts from the same premise as this one: "you can't validate an LLM under GAMP 5" is the wrong framing, because what you validate is the controls around the LLM. It then answers the auditor's question, "which control applies to which component?", with four components and four strategies.
Component 1 · Deterministic
The deterministic path
Rule engines, checkers, SPC and capability queries, reconciliation logic: validated conventionally as GAMP 5 Category 5, each rule version-controlled with a named owner.
Component 2 · Generative
The generative / LLM path
Drafting, synthesis, hypothesis generation: controlled by grounding on an approved-source registry, by deterministic checking of every output, and by human review, under CSA-aligned controls. CSA is FDA device guidance that pharma applies by analogy, and draft Annex 22 excludes dynamic and probabilistic/generative models from critical GMP use, which is the reason the critical decision stays on the deterministic path.
Component 3 · Statistical
Statistical and chemometric models
MSPC, PLS, ML scorers: they follow an analytical-method lifecycle rather than software validation. NOC set, cross-validation, held-out test, QA approval in the model registry, performance monitoring against the reference method.
Component 4 · Human
Human sign-off
On every output that enters a regulated record, governed by Part 11 and Annex 11, with the system designed so the human has the evidence, findings and confidence needed to decide.
The depth of assurance follows the risk of the decision, and the governance manifest on each agent card records which posture applies, so the validation package is a rendering of the card rather than a document written from scratch.
Everything below is that strategy worked through: §1.3 runs it on two production systems, stage by stage; §4 says which model types each path may use; §8.1 shows one CPV agent split across all three technical components; §9 gives the agent patterns; §15 and §16 cover the human sign-off component.
1.3 Show me: the standard lifecycle, run on ProtoCheck and Doscierge
Spine
Layer 3 · Show me§1.1 is the lifecycle every CSV team already runs, and §1.2 is the parent article's rule for which control applies to which component. This subsection puts the two together on the two production systems described in the parent article's §17: ProtoCheck (clinical protocol compliance) and Doscierge (regulatory documentation compliance). Later sections refer back to this table, so the same two systems can be followed through every stage.
Honest scope
Both systems are built and evaluated, and Doscierge's checker is designed for GAMP 5 Category 5 validation. Neither is presented here as validated in a customer's GxP environment. The tables show how the package would be built. Thresholds marked illustrative are examples, not measured results; the measured figures are the parent article's.
The two claims, written the way §1 says a claim must be written
ProtoCheck
Doscierge
Context of use
A draft clinical protocol is checked against 500+ expert-calibrated rules by four specialist reviewers (safety, efficacy, regulatory, operational). A qualified clinical reviewer accepts, modifies or rejects every finding before the protocol is finalised
An ANDA documentation package is checked against 970 deterministic rules and a curated knowledge graph (10,292 triplets, 29 domains). A regulatory reviewer dispositions every cited finding before submission
What the AI decides
Nothing on its own. It proposes findings; the protocol sign-off is human
Nothing on its own. The deterministic checker produces the findings, the LLM path extracts and explains, and the submission decision is human
Parent-article use-case
UC-D1 Trial design · the rule-gate pattern
UC-D4 Regulatory documentation · the draft-then-check pattern
Tier 2: decision support, human decides. Consequence of a miss: an amendment, a delay, or a safety design issue caught late
Tier 2, upper end: human decides, but a missed deficiency can cost a refuse-to-receive or a deficiency cycle
Performance claim (illustrative)
On a locked set of seeded-defect protocols across [n] therapeutic areas, recall on safety-rule findings ≥ [x]%, with the 95% CI lower bound no worse than a qualified reviewer working unaided on the same set
On the expert-labelled evaluation set, per-domain recall ≥ [x]%, with the CI lower bound above the unaided-reviewer baseline, findings per package within reviewer capacity, and no missed seeded critical errors
The lifecycle, stage by stage. The left column is the CSV/CSA stage you already run; the second column is the parent §11 component it lands on.
Full tableThe standard validation lifecycle, stage by stage, for ProtoCheck and Doscierge10 rows
Standard stage
§11 component
ProtoCheck
Doscierge
Deeper in
1. GxP assessment and intended use
All
GCP. The AI function is advisory review of a controlled document; the sign-off stays human
Regulatory submission content. The checker is Category 5 software; the LLM path is non-critical and works under human review
Doscierge's F1 of 85.87 is one number across 29 domains. To become an acceptance criterion it needs the per-domain breakdown, a confidence interval, and the unaided-reviewer baseline measured on the same set (§5).
2
Precision is a safety control by another route.
Doscierge's first evaluation produced about 288 findings per package; precision engineering brought that to about 73. A reviewer handed 288 findings stops reading them. That is the automation-bias and alert-fatigue risk §13.6 is written for, so finding volume is a monitored metric, not a cosmetic one.
3
SME ownership can be built into the tool.
ProtoCheck's provisional → confirmed calibration with a named signature is what draft Annex 22 §4.1 asks for in principle (subject-matter experts own the acceptance criteria), applied here by analogy to a clinical tool. Only confirmed rules belong in the validated scope.
4
The eval harness is the validation evidence engine.
Regression before every merge is a pre-approved change envelope (T9, §14.2) in engineering form. Evidence is generated by the change itself, not written afterwards.
So what
The claim is defined: performance, for a context of use, on a defined population, against a baseline, kept true over time. The standard, the architecture and two working systems now sit on one lifecycle. §2 turns to the practitioners: thirteen questions, each pinned to a stage of that lifecycle.
Thirteen questions, one stuck point: execution, not principle
The practitioner's lens · where the lifecycle bites
Spine
Layer 4 · The practitioner's lensQ1–Q9 came from practitioners at a September 2026 web event on acceptance criteria for AI validation, Q10 from the author, and Q11–Q13 from practitioner conversations. Each question is a symptom. The table maps each one to the underlying pain, to what the market does not yet supply, to the lifecycle stage in §1.3 where it bites, and to the section that answers it.
Full tableThe thirteen questions mapped to underlying pain, market gaps, who feels it most and the section that answers each13 rows
Read-across. Seven of the eight questions from the session ask "show me how", not "tell me what". The audience already accepts that AI needs validating. They are stuck on execution. Three signals stand out, and each one is a component of the claim:
Signal 1
The human baseline is missing (Q6, Q2)
Annex 22 §4.3 says model acceptance criteria must be "at least as high as the performance of the process it replaces", and that the performance of the process being replaced must therefore be known. Most companies have never measured their manual process's error rate. Measuring that baseline is a service on its own (§5).
Signal 2
Vendor-driven category creep (Q7)
Most small companies will not build AI. They will receive it inside eQMS, LIMS, EDC, and document-management SaaS releases. The governance problem is a supplier-management and change-control problem before it is a model-validation problem (§17, §14).
Signal 3
Agents are ahead of the guidance (Q3)
Annex 22 excludes generative AI and LLMs from critical GMP applications altogether. Yet teams are building agents today. They need a defensible way to position agents: non-critical with human-in-the-loop, or split so that the deterministic components carry the validated decision. The parent article's dual-path architecture (§10–11) is exactly that answer (§9, §1.2).
So what
The questions are not about whether to validate AI. They are about what the claim is and how to prove it, which is why each one is pinned to a stage of the §1.3 lifecycle. Before building on the claim, §3 clears away the views that get it wrong.
Ten views the industry still repeats, and the correction
Common misconceptions · what the claim is not
Spine
A claim defined wrongly cannot be proved. Vendor material, training content and conference talks on AI validation often repeat a small set of views that were reasonable a few years ago but are now outdated, device-specific, or incomplete. The corrections below matter because each one changes what a validation package must contain. The last column points to the section of this page that corrects the view in practice.
Misconception
Why it fails
Correct practice
Anchor
Corrected in
"AI validation means proving the model is accurate." Accuracy is treated as the headline metric on monitoring dashboards and alert limits
A model has no validated state outside a use. The same model can be low risk as a drafting aid and critical as a release gate. Accuracy alone says nothing about the controls that bound the residual risk
Define the context of use first. Then validate the controls around the model for that use, with model metrics as supporting evidence
FDA Jan 2025 draft, Steps 1–3; Annex 22 §3.1; ISPE GAMP AI Guide (AI as a component of a computerized system)
"AI systems are probabilistic, so outputs may vary." Presented as a property of all AI
A trained ML model that is frozen is deterministic: same input, same output. Only generative models sampled with non-zero temperature, and dynamic models, are not
Treat determinism as a design choice. Put the critical decision in a static, deterministic component (rules, SPC, frozen ML); keep LLMs in non-critical, human-reviewed roles (dual-path)
"With expert review, AI can support batch release and other GMP-critical decisions." Batch disposition support rated medium risk, and release recommendations high risk but acceptable with expert review
Draft Annex 22 excludes generative AI and LLMs from critical GMP applications whatever the human oversight. Human-in-the-loop is the condition for non-critical use. Risk ladders that are not anchored to a criticality definition tend to rate the same use differently in different places
Critical GMP decisions (release, disposition, reject) run only through deterministic logic or a static ML model validated to Annex 22. An LLM may draft the narrative; it may not carry the decision
"The Expert-in-the-Loop is the guardrail." Framed as split-second, experience-based judgement
An unmeasured human review is a paper control. Automation bias and anchoring mean reviewers miss errors they would catch in a colleague's work. GxP review should be structured, not split-second
Keep the good parts: credentialled reviewers, Approve / Modify / Reject with a written rationale. Then measure the reviewer: seeded-error catch rate, and override rates in both directions
"95% accuracy, with alert at <95% and action at <90%." Generic limits for accuracy, hallucination (>1% / >3%) and overrides
Accuracy on imbalanced data hides the critical class (§7 pitfall 1). Limits not tied to consequence or to a baseline cannot be defended. A 1% hallucination rate reaching a reviewer is unacceptable in a CSR (§6)
Criteria per class and per subgroup, with confidence intervals, set by an SME before testing and no worse than the measured process baseline. Derive alert and action limits from the validated performance and the consequence of failure. A two-level alert/action structure is sound; the numbers must be yours
"Drift happens as new data is introduced into the model"; three drift types; track accuracy monthly
A locked model does not change; the world it sees does. The three-type list misses prevalence shift and unseen classes (the dangerous case). In production you rarely have ground-truth labels, so a monthly accuracy trend is often not measurable
Monitor what you can observe: input-space and out-of-distribution scores, the "undecided" (low-confidence) rate, class-wise output rates, override and edit rates, and a periodic reference check (e.g., AQL re-inspection of accepted units) that supplies ground truth
FDA's AI position described only through device documents. Definitions and risks taken from the 2019 AI/ML SaMD discussion paper and device guidance, with the drug-side and EU documents left out
SaMD and device-software guidance do not govern drug GMP or evidence used for drug regulatory decisions. Applying them without saying so gives the wrong framework
For drugs and biologics use: FDA Jan 2025 draft (7-step credibility framework; CDER/CBER); draft EU GMP Annex 22 (July 2025; final targeted Q4 2026); ISPE GAMP AI Guide (July 2025; complements GAMP 5 2nd Ed. Appendix D11); FDA–EMA Good AI Practice principles (Jan 2026). Cite device guidance only as an analogy
"Ask the LLM to rework its own output"; the LLM remembers earlier prompts
Useful for personal productivity. In GxP, conversation memory is uncontrolled state that breaks reproducibility, and self-review is not independent verification. The same principle applies to migration tools: a tool must not verify its own results
Pinned model and prompt versions, a fresh context per case, N-run testing, and an independent deterministic checker. Do not use the same model as its own judge
Guidance status quoted out of date. CSA described as still draft and "likely to be issued for all regulated industries"; software validation anchored on 21 CFR 820.70(i); Annex 11 described as "not a legal requirement"
FDA issued the CSA guidance as final on 24 Sep 2025, from CDRH/CBER, for device production and quality-system software. The QMSR (effective 2 Feb 2026) replaced the old Part 820 text; software validation now sits under ISO 13485:2016 clause 4.1.6. EU GMP Annex 11 is part of the EU GMP Guide that manufacturers are inspected against, and its revision (July 2025 consultation) runs alongside Annex 22
Apply CSA to pharma by analogy and say so. Anchor drug systems on Part 11, the predicate rules, the 2018 Data Integrity guidance, Annex 11, and GAMP 5 2nd Ed.
Each row moves the same way: from "is the model good?" to "is this use, with these controls, demonstrably at least as good as the process it replaces, and can we show it still is?"
So what
Two of the ten views turn on the same confusion: what "locked", "adaptive", "deterministic" and "probabilistic" actually mean, and which of them a GMP claim may cover. §4 settles the vocabulary before Part B sets the bar.
Locked, retrained or learning: which models the claim can cover
Practitioner question Q11 (a) · not from the session"Locked versus adaptive, static versus dynamic, deterministic versus probabilistic: which of these can be used where in GMP, and what do the terms actually mean?"
Spine
A claim is only as stable as the thing it is made about. This section fixes the model vocabulary, shows what draft Annex 22 allows in critical and non-critical GMP use, explains where the "locked vs. adaptive" language comes from (FDA device policy, applied to pharma by analogy only), and ends with a decision table. The drift half of Q11 is in §11.
4.1 Definitions: two axes, not one ladder
The terms are often used as a single scale from "safe" to "risky". They are two independent axes: how the model changes over time and how it produces an output.
Axis 1 · How the model changes
Definition
Typical examples
What the claim covers
Static / locked
Weights and configuration are frozen at release. The model changes only through a formal change.
Frozen classifier on an AVI camera; a PLS or PCA model with a fixed NOC set; a rules engine
One model version, one claim. The claim holds until a change is made or the input space moves (§11)
Locked with periodic retraining (the realistic middle)
Locked in operation; retrained at intervals or on triggers, each retrain released as a new version after regression
CPV MSPC model refreshed after a qualified new site is added; a deviation classifier retrained annually on adjudicated labels
Each version carries its own claim. Retraining is a change (Annex 22 §10.1), usually inside a pre-approved envelope (T9) with regression on the locked test set
Dynamic / continuously learning
Updates its own parameters from production data without a release step
Online-learning recommenders; self-tuning thresholds fed from production feedback
No stable object to make a claim about. Excluded from critical GMP use by draft Annex 22
Axis 2 · How the output is produced
Definition
Typical examples
What the claim covers
Deterministic
Same input → same output, every time
Rules; SPC; a frozen ML classifier or regression run without sampling
A reproducible result. Conventional test-and-retest logic works
Probabilistic
The same input can produce different outputs (sampling, temperature, non-deterministic serving)
LLM text generation at non-zero temperature; some ensemble or stochastic inference services
A distribution, not a result. Acceptance is set on N runs per case (§9), and the critical decision is kept off this component
Two consequences follow. First, "AI is probabilistic" is not a property of AI; a frozen ML model is as deterministic as a spreadsheet (§3, row 2). Second, a generative model is probabilistic and usually vendor-hosted, so it sits at the far end of both axes at once unless the version is pinned (§4.4).
4.2 What draft Annex 22 allows
Draft EU GMP Annex 22 (consultation text, July 2025) is scoped to static, deterministic models used in critical GMP applications. Dynamic models, and generative AI and LLMs, are excluded from critical GMP applications; the draft's principles may be applied, where applicable, to non-critical uses. "Critical" means a direct impact on patient safety, product quality or data integrity. That gives four cells:
In scope for critical useNon-critical: principles where applicableNon-critical: allowed with human-in-the-loopExcluded from critical useNot covered by the draft
Two axes, four cells. Draft Annex 22 reads across both axes: only the static, deterministic corner (and the retrained sequence of static versions) may carry a critical GMP decision. The table below is the source for each cell.
Critical GMP use
Non-critical GMP use (human-in-the-loop)
Static + deterministic
In scope Full Annex 22 package: intended use and input sample space (§3.1), SME-owned acceptance criteria per subgroup (§4), independent test data (§5–6), explainability (§8), confidence with an "undecided" outcome (§9), operation and monitoring (§10)
Principles applied where applicable; depth by risk (§17 Tier 2)
Locked with periodic retraining
In scopebetween retrains, as a sequence of static versions; each retrain is a change under Annex 22 §10.1 with regression (T9)
As above, lighter
Dynamic / continuously learning
Excluded
Not covered by the draft; treat as high risk, and prefer a locked variant
Probabilistic / generative (LLM)
Excluded, whatever the human oversight
Allowed with human-in-the-loop; validate the controls around it (grounding, checker, review, audit trail; §9 Pattern A)
The final text may widen the non-critical guidance; the EMA workshop of 30 June–1 July 2026 discussed whether dynamic and probabilistic models could be addressed (§19). The direction of travel does not change the design rule: put the critical decision in the top-left cell.
4.3 Where "locked vs. adaptive" comes from, and why pharma applies it by analogy only
The vocabulary is FDA device policy. The April 2019 discussion paper, Proposed Regulatory Framework for Modifications to AI/ML-Based Software as a Medical Device, introduced the distinction between a "locked" algorithm, which gives the same result each time for the same input, and an "adaptive" one that changes its behaviour using real-world data. The same paper proposed the Predetermined Change Control Plan, which became final guidance for AI-enabled device software functions on 4 December 2024 (Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions). Both documents govern devices under CDRH's premarket framework. They do not govern drug GMP or the evidence used in drug regulatory decisions (§3, row 8). Pharma borrows two things from them by analogy: the vocabulary, and the idea that anticipated changes can be described, tested and approved in advance. The GMP anchor for that second idea is Annex 22 §10.1 change control and your own change-control SOP; T9 is the operational form (§14.2).
4.4 A vendor LLM is "locked" only if the version is pinned and the deprecation schedule is managed
A hosted foundation model behaves like a locked model only under three conditions, all of which are yours to enforce, not the vendor's default:
The version is pinned. The endpoint identifies an immutable model version, and your IQ records it (§8.3, §14.3). An alias that the vendor can repoint ("latest", "stable") is not a pinned version.
The deprecation schedule is managed. Vendors retire versions on a schedule. A deprecation notice is a planned change with a regression suite ready to run (§14.5, scenario S2). The quality agreement must oblige the vendor to give notice (§17, element 4).
Silent vendor-side changes are detectable. Safety filters, tokenizers, system-level instructions and serving infrastructure can change behind an unchanged version string. A daily configuration and behaviour check on a fixed probe set (T8 OPS-06; §13) is the only defence, and any unexplained shift is investigated as an unauthorised change (Annex 22 §10.2).
Without all three, the model is adaptive from your point of view even if the vendor calls it a release.
4.5 Decision table: model type × criticality → allowed? → controls
Conventional IQ/OQ/PQ; rule versioning with a named owner
Static ML (frozen classifier, PLS, PCA)
Deterministic
Yes
Yes (Annex 22 in scope)
Intended use and sample space; SME criteria per subgroup vs. measured baseline (§5); independent test set (§8.4); explainability; "undecided" outcome; monitoring (§13); retraining as a change (T9)
Static ML
Deterministic
No
Yes
Principles where applicable; Tier 2 assurance plan (§9)
Locked with periodic retraining
Deterministic
Yes
Yes, as a sequence of static versions
As static ML, plus a T9 envelope defining what a retrain may and may not change, regression on the locked test set, QA approval per version
Dynamic / online learning
Either
Yes
No (draft Annex 22)
Redesign as locked-with-retraining
Dynamic / online learning
Either
No
Not recommended
If unavoidable: treat as high risk, snapshot versions for reconstructability, human review of every output, explicit QA acceptance of residual risk
Generative / LLM, vendor-hosted
Probabilistic
Yes
No, whatever the oversight
Split the use (§9 Pattern B): deterministic component carries the decision
Generative / LLM, vendor-hosted
Probabilistic
No
Yes, with human-in-the-loop
Pinned version; managed deprecation; probe-set check (§4.4); grounding on approved sources; deterministic checker; N-run acceptance; measured reviewer catch rate (§5); AI contribution attributed in the record (§15, §16)
Generative / LLM, self-hosted, pinned weights, temperature 0
Near-deterministic in practice
Yes
No under the draft (generative models are excluded from critical use regardless of sampling)
Same as above; the deterministic serving mode helps reproducibility but does not change the regulatory position
So what
The claim can only be made about a locked object with a deterministic output, and it may carry a critical decision only in that form. Part B now sets the bar for the claim: what "at least as good as the process it replaces" means, and how to show it once.
What "at least as good as the process it replaces" means, and how to show it once.
Bridge. Part A defined the claim: performance for one context of use, on a defined population, against a known baseline, for a locked model with a deterministic output wherever the decision is critical. Part B sets the bar that claim must clear and shows it being cleared once. §5 supplies the baseline that draft Annex 22 §4.3 presumes; §6 walks a clinical drafting agent through FDA's seven-step credibility framework; §7 lists where familiar CSV concepts go wrong when the bar is set the old way.
Audience question Q6"Do you know of any studies that examine the error rates of humans that can be used as a basis for determining the acceptance of AI?"
Spine
This section establishes the baseline component of the claim. It matters more than any other question on the list, because draft Annex 22 §4.3 makes the human baseline a regulatory expectation. The baseline feeds the acceptance criteria in §8.1–8.2 and T6, the human-review test in §6, and the alert and action limits in §13.
Draft Annex 22 §4.3
"The acceptance criteria of a model should be at least as high as the performance of the process it replaces. This implies that the performance should be known for the process which is to be replaced."
5.1 Published anchors
Domain
Source
What it shows
How to use it
Visual inspection of injectables
USP <1790> Visual Inspection of Injections; Knapp–Kushner methodology
Manual inspection is probabilistic. A unit belongs in the "reject zone" only if manual inspectors detect it with probability ≥ 0.7. Automated systems must show equivalent or better performance than manual inspection
The most mature precedent in pharma for "AI ≥ human". Use it directly for any camera or vision AI (§12)
Clinical data processing
Garza et al., Int J Med Inform 2025, 195:105749. Systematic review and meta-analysis of 93 papers (1978–2008)
Error rates range from 2 to 2,784 errors per 10,000 fields depending on method. Single entry: 4–650 per 10,000; double entry: 4–33 per 10,000; medical record abstraction is the highest
Baseline for AI in data transcription, abstraction, EDC entry, and SDV-adjacent use-cases (§6)
About 1–5 errors per 1,000 hand-transcribed batch-parameter values
Illustrative only Replace with your own measurement
5.2 Measure your own baseline
Published numbers help frame the argument. Your acceptance criteria should rest on your process, measured with this protocol:
Same test set for human and model. Use the independent, stratified, adjudicated set from Annex 22 §5 (§8.4, T4).
Several operators under normal conditions. Include realistic pace, shift, and fatigue. Record each operator's performance, not just the average.
Measure the right metrics. Sensitivity per critical class (false accepts are the patient risk), false-reject rate (the cost), and inter- and intra-rater agreement (kappa). Human inconsistency is often a bigger issue than the human average.
Report confidence intervals. An AI with 97% sensitivity (95% CI 93–99%) does not demonstrably beat a human at 95% on a 150-unit test.
Compare system to system. If the AI will run with a human reviewer, the comparison is human+AI vs. human alone, not model vs. human. The same measurement, repeated in operation as a seeded-error catch rate, becomes the human-in-the-loop metric family in §13.
Two caveats that catch teams out
The label ceiling. If humans label the test set, the model is scored against human judgment and cannot demonstrably exceed it. Where possible, use a stronger reference: lab testing, destructive testing, or panel adjudication (Annex 22 §5.3).
"No decrease" applies per subgroup. A model can beat humans overall and still be worse on the rare, critical subgroup, such as small particles in amber glass. Annex 22 §4.2 allows different criteria for different subgroups. Use that (§8.2, T6 §4).
Show me
Doscierge's evaluation harness reports P 92.31 / R 80.27 / F1 85.87 against expert-labelled ground truth (parent §17). By the logic of Annex 22 §4.3, that recall becomes an acceptance criterion only when it sits next to a baseline: the recall of a qualified regulatory reviewer working unaided on the same labelled packages, measured with the §5.2 protocol. Without that number, 80% recall could be a large improvement or a step backwards (§1.3, stage 3).
So what
With the baseline measured, an acceptance criterion becomes defensible: a metric per subgroup, with a confidence interval, no worse than the process it replaces. §6 shows the whole claim being made and proved once, for a clinical drafting agent.
A worked validation: CSR drafting under the seven-step credibility framework
Audience question Q4"Is it possible to see an example of a validation of an AI tool following industry practice?"
Spine
This section establishes the evidence component of the claim end to end: question of interest, context of use, risk, credibility plan, execution, results, adequacy. It uses the baseline from §5 and the controls that §8 then generalises.
Example chosen: UC-D4, clinical study report (CSR) drafting with generate-then-check (parent article §8.2). It is a clinical-development tool, the highest-volume GenAI use-case in the industry, and the one regulators are watching most closely. This walk-through uses FDA's seven-step credibility framework (January 2025 draft), because the CSR output supports regulatory decisions. It borrows Annex 22's test-data discipline where that applies. Numbers are illustrative
1
Define the question of interest
"Is this CSR efficacy-results section accurate, complete, and traceable to the locked TFLs, so that a medical writer can finalize it?"
2
Define the context of use
The drafting agent produces Sections 11–12 from locked TFLs and the SAP. The deterministic checker (MCP tool) checks required elements, numeric consistency against source tables, and cross-references. A medical writer reviews every sentence with its source span and signs under Part 11. The AI output is never the final record, and a qualified human is accountable (Pattern A, §9).
Use a test corpus of 40 sections from 8 completed studies across 3 therapeutic areas, not used in development and locked in an access-controlled repository. Labels (the correct content) come from the approved, published CSRs. The baseline comes from the company's historical QC findings on manually drafted CSRs, e.g., numeric discrepancies found per section at QC.
6
Document results and deviations
Report each metric with its 95% CI, by therapeutic area and section type. Any failure is investigated. Example: the agent paraphrased a secondary endpoint name, and the checker's terminology rule was then extended. That is a ruleset change, so it goes through change control and the ruleset is re-tested (§14.2).
7
Determine adequacy
QA and Clinical agree the model is credible for this context of use. The release is scoped: TA1–TA3 and Sections 11–12 only. Adding a new TA or new sections is a predetermined change (T9) with a defined retest.
What makes this "industry practice" rather than a demo: risk-based depth (CSA critical thinking), independent test data, a human baseline, a demonstrated rather than assumed human-review effect, a scoped release, and a monitoring plan (reviewer edit distance, checker findings per section, unsupported-claim rate; §13 UC-D4 example) that feeds periodic review.
Manufacturing counterpart (outline only)
UC-M4 blend endpoint. Static PLS model plus a deterministic endpoint rule, so it is in scope for Annex 22 as a critical application once it replaces the fixed-time step (§4.5). Intended use and sample space cover formulations, blender scale, and API lots. Acceptance means the endpoint agrees with the thief-sample/HPLC reference and is no worse than the validated fixed-time step (Annex 22 §4.3). The model sits in the QA-approved registry, and the advisory phase acts as an extended PQ.
So what
One claim, proved once. What goes wrong in practice is that teams reuse CSV habits that quietly lower or misplace the bar. §7 lists those habits, concept by concept.
Audience question Q1"Is it possible to give more examples of the showed content to AI cases/pitfalls when applying it to AI?"
Spine
This section protects the bar. Each row takes a standard CSV concept, shows how teams misapply it to AI, and gives a concrete example from the parent article's use-cases. The "what good looks like" column is, in every row, a restatement of the claim: population, metric, baseline, kept true.
CSV concept
Common AI pitfall
Example (parent-article UC)
What good looks like
URS
"The model shall detect defects with high accuracy." That cannot be tested.
UC-M4 blend endpoint: "detect uniformity"
URS states the decision, the population, the metric, and the threshold: "Flag blend-uniform when moving-block SD of predicted API content < 1.5% w/w for 5 consecutive blocks; sensitivity to non-uniform blends ≥ baseline of the fixed-time step." (§1, last row)
Risk assessment (FMEA)
Failure modes list software bugs only; model failure modes are missing
UC-M2 CPV signal agent
Add model failure modes: false negative on a real drift (patient risk), false positive flood (alert fatigue, so real signals get ignored), out-of-distribution input from a new site, confidence mis-calibration (§8.2)
Non-functional requirements
Functional AI requirements are written; system requirements are forgotten. Often-missed items include error messaging, error logging and overload handling
UC-D4 drafting agent
Add AI non-functional requirements: behavior on out-of-scope or low-confidence input (refuse or "undecided"), logging of model/prompt version per output, rate-limit and timeout behavior, and fallback when the vendor model is unavailable
Supplier assessment
Vendor audit checks the QMS, but not how the model was trained or what data it used
Ask for a model card, training-data provenance, test-set independence, the change-notification policy for model updates, and whether customer data trains the shared model
Leveraging vendor test evidence
The vendor's published model evaluation is accepted as validation. Vendor IQ/OQ cannot be used as-is because it was not tested for your intended use; vendor documentation must be scrutinized and owned by the regulated company GAMP 5 2nd Ed., supplier leverage; Annex 11 §3
Vendor AI summarization in a document system
Use the vendor's evaluation to scope your testing, not to replace it. Re-test on your own independent set for your context of use, focused on the subgroups that matter to you
IQ
Installation verified. Model version not pinned
UC-D4 drafting agent
IQ confirms model ID and version hash, prompt/config version, tool allowlist, and retrieval index version. All are configuration items (§4.4, §8.3)
OQ
Tests run on "happy path" samples chosen by the developer
UC-M3 PAT chemometric model
Tests run on an independent, stratified test set that covers rare variations (Annex 22 §5.1). The developer never saw it (Annex 22 §6.2)
PQ
One PQ run, then "validated"
UC-M1 APQR
Parallel run against the manual APQR for a full cycle, with agreement within pre-set tolerance and every discrepancy explained (as in parent §8.4). Then continuous monitoring replaces the one-time PQ (§13)
Test deviations
The CSV categories (system error, script error, tester error) are applied unchanged, so every model miss is logged as a "system error"
UC-M3 PAT model
Add AI categories: model error (the model is wrong), reference/label error (the adjudicated answer was wrong, so correct it under change control and re-score), and test-set coverage gap (the case is outside the defined input space). Each has a different fix
Independent verification
An LLM is used to grade its own outputs ("LLM-as-judge" with the same model)
UC-D4
Verification must be independent of what it measures (the same principle CSV applies to data-migration verification tools). Use the deterministic checker, human adjudication, or at least a different, separately validated evaluator
Traceability matrix
Requirements traced to tests. Data not traced
UC-M2
Trace requirement → metric → test-set subgroup → result → monitoring KPI. The subgroup is the new column
Audit trail (Part 11)
Logs the user action, not the AI's contribution
UC-D3 clinical ops NBA
Log the AI output, its confidence, the model version, the evidence cited, and the human accept/modify/reject with rationale (parent G2 decision provenance; §15, §16)
Change control
Retraining treated as "data refresh, no change"
UC-M3 PAT model recalibration
Retraining is a change. It gets impact assessment, a regression test on the locked test set, and QA approval in the model registry (Annex 22 §10.1; §14)
Periodic review
Annual checkbox: "system still in use"
Any
Review the performance trend, drift metrics, human override rate, incidents, and vendor changes. Decide to continue, retrain, or retire (§8.8)
Data integrity (ALCOA+)
Applied to records, not to training data
UC-M5 deviation investigation
Training data and labels are attributable (who labelled), original (source-linked), and accurate (adjudicated). Label errors are data-integrity errors (§16)
Human review
"A human approves every output", so the risk is considered controlled
UC-D4 CSR drafting
Measure whether reviewers actually catch seeded errors (challenge testing of the human). Automation bias makes an unmeasured human-in-the-loop a paper control (§5)
Three pitfalls worth a slide on their own
1
"Accuracy" on imbalanced data
A defect-detection model that labels everything "good" scores 99.8% accuracy when the defect rate is 0.2%. Acceptance criteria must be per class: sensitivity on critical defects, false-reject rate on good units.
2
Test set equals training distribution
Validation passes, and then the first batch from a new glass supplier fails. The input sample space must be defined before testing and monitored after deployment (§11, §13).
3
Treating the human as the only control
Parent-article G2 makes the point: an agent whose recommendations are accepted 98% of the time might be excellent, or its reviewers might have stopped reading. The difference shows up only if you seed known errors and measure the catch rate. Good design helps the reviewer. Make them choose Approve, Modify, or Reject, and write the reason in their own words. But the design does not replace measuring how well they review.
So what
The bar is set and the habits that lower it are named. Part C turns the claim into a set of controls that a validation package can be built from, and that an inspector can be walked through.
The eight CSV lifecycle controls and their AI extensions; the templates that carry them.
Bridge. Part B set the bar: a measured baseline, criteria per subgroup with confidence intervals, and one worked claim proved end to end. Part C builds the controls that carry that claim through the lifecycle. §8 is the centrepiece: the eight CSV lifecycle controls, each with one AI extension, its template, and a CPV agent as the running example. §9 positions agents so that the controls apply to them, and gives the template pack. §10 closes with the proportionate answer to "what if AI wrote the templates?"
Audience question Q8 (paraphrased)"How do we apply CSV to AI in practice? For example AI GxP assessment, AI risk assessment, AI version control, AI data management, AI model design, AI registry, AI performance monitoring, AI periodic review."
Spine
This section establishes the controls component of the claim. It is the most complete question on the list, and it effectively describes an AI lifecycle SOP set. Each control below follows the same pattern: what CSV already does → what AI adds → the artifact → a worked example. The running example is UC-M2, the Continued Process Verification signal agent (parent article §8.4). Its agent card (parent §6) already declares most of these controls. Sub-numbers 8.1–8.8 are the ones the templates cite as Q8.1–Q8.8.
§14.4 scheduled task; §13.7 monitoring the monitors
"How do you know reliability remains acceptable?"
The template pack (T1–T10) is complete. T1, T2 and T8 are published free at the links above; the remaining templates are available on request (§9, §18).
Show me
Read the matrix for one real system. For Doscierge, the GxP assessment finds a Category 5 checker plus a non-critical LLM path (§8.1). Version control covers 970 rules, the knowledge-graph snapshot and a pinned model (§8.3). Data management means the locked, expert-labelled evaluation set (§8.4). Monitoring watches findings per package and acceptance per rule (§8.7). §1.3 gives the full row for both systems.
Lifecycle view
ConceptProjectOperationCross-phase (system of record)
Lifecycle view. Concept (8.1 GxP assessment) → project (8.2 risk assessment, 8.5 model/agent design, 8.4 data management) → operation (8.7 performance monitoring, 8.8 periodic review) → retirement or successor. Version control (8.3) and the AI registry (8.6, the system of record across all phases) span every phase. The review decision closes the loop: continue, retrain, restrict or retire, and a new intended use starts the cycle again.
8.1 · Concept
8.1 AI GxP assessment
CSV today
Is the system GxP? Is it Part 11 relevant? What is the GAMP category?
What AI adds
Is the AI function GxP? The CSV test still applies: does it "touch" a regulated product, or collect, analyze, report, store or transmit regulated data? Critical processes include release, safety and efficacy data, recall, adverse events, and pharmacovigilance.
Is it critical (direct impact on patient safety, product quality, or data integrity)?
Is the model static or dynamic, and deterministic or probabilistic? Draft Annex 22 allows only static, deterministic models in critical GMP use (§4).
Which regulatory frame applies: Annex 22 (GMP), the FDA credibility framework (regulatory-decision support), or EU AI Act high-risk classification?
GMP; the signal disposition is a regulated record; critical = Yes. The SPC path is deterministic, rules-based software (GAMP Cat 5 custom queries). The MSPC PCA model is static ML, so it is in Annex 22 scope. The LLM narrative that explains the signal is non-critical with human-in-the-loop. One agent, three components, three postures. This is the parent article's §11 table (§1.2) applied.
8.2 · Project
8.2 AI risk assessment
CSV today
System-level FMEA focused on software failure.
What AI adds
Model risk = model influence × decision consequence (FDA 2025 draft).
AI-specific failure modes: false negatives and false positives by subgroup, out-of-distribution inputs, drift, mis-calibrated confidence, automation bias, bias across sites or products, data leakage, adversarial input (for agents).
Risk sets the depth of every later control (CSA critical thinking), and the tier that sets monitoring minimums in §13.
Two CSV lessons apply with more force to AI. First, re-assess risk at every change, not only at validation (ICH Q9(R1) treats risk review as a lifecycle activity). Second, the system owner must formally accept residual risk and own it. For AI there is always residual risk.
The acceptance criteria that the risk rating justifies are set against the measured baseline from §5.
Sensitivity acceptance on seeded historical drifts; APQR cross-check; periodic re-challenge
Alert flood (false positives)
Scientists ignore signals
Medium
High
Precision target; orchestrator de-duplication; alert-fatigue index KPI (§13.6)
New site data outside the NOC model
Spurious T² excursions
Medium
Medium
Input-space monitoring; site-specific NOC or model transfer protocol (§11 FAQ 5)
LLM narrative mis-states a chart value
Wrong disposition rationale
Medium
High
Checker numeric-consistency rule; human disposition
8.3 · Cross-phase
8.3 AI version control
CSV today
Software version and configuration baseline.
What AI adds
The configuration baseline grows to include:
model weights (hash)training-data snapshotthe locked test setfeature and pre-processing codehyperparametersprompts and system instructionstool allowlistretrieval index versionfoundation-model version (vendor-pinned; §4.4)thresholds and confidence policy
Draft Annex 22 §10.2 requires configuration control with "effective measures … to detect any unauthorised change".
Artifact · Configuration-item list in T5 (on request), plus the registry version history. §14.1 shows the same items as CMDB configuration items.
UC-M2 example
Agent card mfg.cpv.signal-agent v2.3.0; model: scoped-small@pinned-2026-07; MSPC model v1.4 with its NOC set reference; eval harness mfg.cpv.eval@1.4. Semantic versioning rule:
major = intended use or class set changed, so revalidate
minor = retrained on the same class set, so regression test
patch = non-functional change, so documented verification
8.4 · Project
8.4 AI data management
CSV today
Data migration and data integrity of records.
What AI adds
Training, validation, and test data are validation evidence.
Requirements: provenance and lineage; a documented, verified labelling process; test-data independence with access control and audit trail; staff independence (Annex 22 §6); documented pre-processing and exclusions; no AI-generated test labels (Annex 22 §5.6).
Place AI on the data lifecycle (capture → maintenance → synthesis → usage → publication → archival → purge; the lifecycle view taken by MHRA 2018 and PIC/S PI 041-1). AI output is synthesized data: new values derived from other data. It needs the same attributability and metadata as any other GxP record: model version, inputs, confidence, and reviewer (§16). Once published outside the company (e.g., in a submission), it cannot be recalled, which is why review happens before publication.
Keep rejected and overridden AI outputs, and "undecided" results, with the reason. This follows the same principle as invalidated laboratory results, which stay in the record with their justification FDA, Investigating OOS Test Results, 2006, rev. 2022. They are also your best monitoring data (§13).
Ground truth is 2,140 labeled batch-phases across 3 sites and 2 products (from the agent card). Labels are dispositions adjudicated by two process scientists. The test set is split before training, stored in an access-controlled repository, and the MSPC developers never saw it. Every data point traces back to the historian tag and the MES event frame.
8.5 · Project
8.5 AI model / agent design
CSV today
Functional and design specifications.
What AI adds
Documented algorithm choice and rationale.
Explainability approach (feature attribution such as SHAP or LIME, contribution plots; Annex 22 §8).
Confidence thresholds with an "undecided" outcome (Annex 22 §9).
For agents: tool allowlist, autonomy bands, escalation logic, and deterministic orchestration where the path must be reproducible (§9).
Design for validatability: keep the critical decision in the deterministic or static component (§4.5).
Artifact · T5 Model/Agent Design Specification (available on request; the agent card is its machine-readable form)
UC-M2 example
Four isolated specialists (SPC, capability, MSPC, lot-effect) and a deterministic orchestrator; each specialist sees ≤5 tools. MSPC excursions show a contribution plot (explainability). Confidence policy: ≥0.90 notify the scientist directly; <0.55 escalate to the MS&T lead.
8.6 · Cross-phase
8.6 AI registry
CSV today
Validated systems inventory.
What AI adds
A registry that is the system of record for every model and agent. It holds:
identity, owner, intended use, risk tier, and lifecycle state (draft → lab → candidate → factory → deprecated)
current and previous versions
evaluation record and approvals (security, quality, data steward, business)
which consumers depend on it (impact analysis for change control)
For a small company this is the AI inventory from §17 with more fields. At enterprise scale it is the agent registry and service catalog from parent §6, synchronised with the CMDB (§14.1).
Artifact · Registry entry / agent card
UC-M2 example
The parent §6 card, including consumers: [mfg.apqr.orchestrator, mfg.quality-command-center, mfg.deviation.investigator]. A change to the CPV agent automatically flags three downstream consumers for impact assessment.
8.7 · Operation
8.7 AI performance monitoring
CSV today
Incident management; system availability.
What AI adds
Monitoring of the model's validated performance metrics (Annex 22 §10.3).
Input-drift monitoring against the defined sample space (Annex 22 §10.4); what drift is and is not, in CPV terms, is the subject of §11.
Confidence-calibration monitoring.
Human-in-the-loop monitoring: acceptance, override, and edit rates (Annex 22 §3.3 and §10.5).
Cost and latency for agents.
Each metric has a threshold, an owner, and a response playbook: investigate → restrict to advisory → retrain → retire.
Four design questions make a good checklist: what is measured, how often, what the alert limits are, and what requires action. Monitoring without response criteria has little value. Use two levels, an alert limit and an action limit, as CPV already does, but derive the values from the validated performance and the consequence of failure (§3, row 5; §13.3).
Human overrides are often the earliest signal of degradation. Watch both directions: a rising override rate suggests the model is degrading, and a falling rate may mean the reviewers are disengaging.
§13 lifts these elements into one site policy, with minimum metrics by tier and alert-fatigue controls; T8 sets the parameters per use.
Precision (signals confirmed / signals raised), rolling 30 days
< 0.80
Investigate; tune; change control
Seeded-drift detection (quarterly re-challenge with historical drifts)
< validated sensitivity
Restrict to advisory; retrain
Input drift: share of batch-phases outside the NOC envelope for non-process reasons
> 5%
Model transfer or NOC refresh (change control)
Coach-mode acceptance, rolling 14 days
< 0.70 or > 0.98
Low: fitness review. High: seeded-error check for automation bias
Signal-to-disposition time
> 5 working days
Workflow escalation
8.8 · Operation
8.8 AI periodic review
CSV today
Periodic review of the validated state, often annual and often perfunctory.
What AI adds
A review that decides something. It looks at:
the performance trend against acceptance criteria
drift history
changes made and their cumulative effect
human-in-the-loop statistics
incidents, deviations, and CAPAs involving AI
vendor or model deprecation notices
regulatory changes
whether the intended use still matches actual use (scope creep)
whether the monitoring thresholds and false-alarm rates are still right (§13.7)
The output is a decision: continue / continue with conditions / retrain / restrict / retire. Frequency follows the risk tier (§17).
The review also tests the usual revalidation triggers: a new model version, retraining, new data sources, significant drift, a critical error, a major process change, or a regulatory change. The question is the one every periodic review asks: does the current evidence still support fitness for intended use?
Retirement
Apply the CSV retirement discipline GAMP 5 2nd Ed., retirement phase to the model too. Keep the model version, locked test set, validation evidence, and the AI audit data for the retention period of the records it influenced, so any past decision can be reconstructed (§15, §11.10(c) row). Do not migrate old audit trails into the successor system's audit trail; archive them alongside it.
A semi-annual review, aligned to the APQR cycle so that the CPV agent's own performance becomes an input to the product's APQR. The model is reviewed alongside the process it monitors.
What an inspector will ask, and where the answer lives
Inspectors are unlikely to ask "what was your model accuracy?" and more likely to ask how you set acceptable performance and what happens when the AI is wrong. Map each question to evidence before the inspection:
Likely question
Evidence
Where it lives
How did you determine acceptable performance?
Measured baseline; SME-approved criteria per subgroup
§11 what drift is and is not; §12 the camera monitoring plan (8.7)
§13 the site monitoring policy, §13.3 limits, §13.6 alert fatigue, §13.7 monitoring the monitors (8.7, 8.8)
§14.1 the same configuration items in the CMDB; §14.4 periodic review as a scheduled task (8.3, 8.6, 8.8)
§15 and §16 audit trail, retention and ALCOA+ for datasets (8.4, 8.8)
§17 tiering and the small-company inventory (8.1, 8.6, 8.8)
So what
Eight controls carry the claim from concept to retirement. They assume the AI component is positioned where the controls can reach it. For agents, that positioning is the first decision, and §9 makes it.
Position the agent first. Then validate the system of controls.
Audience question Q3"Is it possible to share any validation templates to validate AI Agents?"
Spine
The eight controls apply to an agent only once the agent has been positioned so that the critical decision sits on a component the claim can cover (§4). This section gives the two positions, what is different about validating an agent, and the template pack that carries the controls.
First, position the agent correctly. Draft Annex 22 excludes generative AI and LLMs from critical GMP applications. So an agent is validated in one of two ways:
Pattern A
Non-critical with human-in-the-loop
The agent drafts, summarizes, or recommends. A qualified human decides and signs. Validation focuses on the system of controls: grounding, the checker, human review, and audit trail. This is CSA-style assurance plus Annex 22 principles "where applicable". Most agents today fit here: UC-D4 documentation, UC-M5 deviation hypotheses, UC-D3 next-best-actions.
Pattern B
Split so the critical decision is deterministic
Anything that directly decides product quality or patient safety runs through deterministic components (rules, validated SPC, a static ML model) that are validated conventionally. The agent orchestrates and explains. This is the parent article's dual-path architecture (§10) and "validate the component, not the model" (§11; §1.2 above).
What is different about validating an agent (versus a single model)
Agent property
Validation implication
Calls tools
Tool allowlist is a configuration item. Test that the agent cannot call tools outside it. Each tool is validated on its own
Takes multi-step paths
Test trajectories, not just final answers: did it retrieve the right source, call the checker, and stop at the escalation threshold?
Prompts and system instructions drive behavior
The prompt is configuration, versioned and under change control. A prompt change is a change (§14.5, S4)
Output is probabilistic
Run each test case N times (e.g., 5) and set acceptance on the distribution: e.g., zero unsupported claims in any run, required elements present in ≥ 95% of runs (§4.1)
Can be manipulated through its inputs
Adversarial and red-team tests: prompt injection in source documents, out-of-scope requests, conflicting sources
Confidence-gated autonomy
Test each band boundary (parent G7: ≥0.90 / 0.70–0.89 / 0.55–0.69 / <0.55): does it escalate when it should?
Depends on a vendor model
Pin the model version. Treat a vendor deprecation as a planned change with a regression suite ready to run (§4.4)
Template pack (the deliverable set)
Ten templates cover the lifecycle. Each one either replaces a familiar CSV document or is new to AI. The pack is complete. T1, T2 and T8 are published free; the others are available on request.
T1, T2 and T8 are published free. The full ten-template pack (intake, risk, context of use, data management, design, validation plan, summary report, monitoring, predetermined changes, periodic review) is available on request.
The full T1–T10 set is for Tier 3 (GxP-critical) uses. For a Tier 2 use, such as a human-reviewed drafting agent, T1 + T2 + a one-page assurance plan is often enough. That plan follows the CSA assurance-record pattern FDA CSA, 2025: risk rating, what we will test and what we will not, how (scripted, unscripted, or automated), what evidence we will keep, and the acceptance criteria. For AI, add the context of use, the baseline, and the monitoring KPIs. Part 11 controls stay a go/no-go gate whatever the tier (§15).
Canonical example: T6, the agent validation plan (skeleton)
AI Agent Validation Plan — skeletonT6 · TEMPLATE PREVIEW
# AI Agent Validation Plan — <agent id> v<version>## 1. Scope & positioning
- Agent: <id, registry link> Pattern: A (non-critical, HITL) | B (deterministic core)
- GxP area: GMP | GCP | GVP | GLP Critical GMP use? <Y/N — if Y, only static/deterministic components may decide>
- Regulated record(s) produced/influenced: <e.g., CSR section draft; deviation investigation report>
- Human decision owner (role): <e.g., Medical Writer; QA Investigator>## 2. Context of use (from T3)
- Question the agent helps answer: ...
- Input sample space & subgroups: <document types, TAs, sites, languages, edge cases>
- Out of scope (agent must refuse/escalate): ...
## 3. Configuration items under control
| Item | Version/hash | Owner |
|-------------------------------------------|------------------------------|----------------|
| Foundation model | <vendor/model@pinned-date> | Platform |
| System prompt / instructions | <git sha> | Agent owner |
| Tool allowlist (≤5) | <list + versions> | Platform |
| Retrieval sources / approved-source registry | <index version> | Knowledge eng. |
| Deterministic checker rules | <ruleset version> | Quality |
| Autonomy / confidence policy | <bands> | Quality |
## 4. Test design
| Test family | What it proves | Data | Runs/case | Acceptance criterion |
|--------------------------------|-------------------------------------------|-----------------------------------------------|-----------|---------------------------------------------------------------|
| Functional accuracy | Output correct vs. adjudicated reference | Independent test set, n=<>, stratified by subgroup | 5 | Per subgroup: <metric ≥ threshold, lower 95% CI ≥ baseline> |
| Grounding / hallucination | Every claim traces to approved source | Same set | 5 | 0 unsupported claims reaching reviewer |
| Trajectory | Correct tools, order, escalation | Scripted scenarios, n=<> | 5 | 100% required steps executed; 0 disallowed tool calls |
| Confidence gating | Escalates below threshold | Boundary cases | 5 | 100% correct band routing |
| Robustness / red team | Resists injection, refuses out-of-scope | Adversarial set | 3 | 0 policy violations |
| Human-in-the-loop effectiveness | Reviewers catch seeded errors | Seeded-error set | n/a | Catch rate ≥ <baseline> |
| Reproducibility | Stable output across runs | Sample | 10 | Decision-level agreement ≥ <x>% |
| Audit trail | Part 11 capture of AI output + human action | Execution logs | n/a | 100% complete |
## 5. Test data independence
- Test set locked in <repo>, access-controlled, audit-trailed; developers never accessed (Annex 22 §6.2)
- Labels adjudicated by ≥2 SMEs; disagreements resolved by a third (§5.3); no AI-generated labels (§5.6)
## 6. Acceptance & release decision
- Baseline source for "no decrease" (Annex 22 §4.3): <measured human/manual performance — see §5>
- Release = all criteria met, or deviations justified and approved by QA
## 7. Post-release
- Monitoring plan: T8 · Predetermined changes: T9 · Periodic review: T10 (frequency by risk)
The full version of this skeleton, with an illustrative UC-D4 worked example, is in T6 (available on request). The remaining templates follow the same structure.
So what
The controls now reach the agent. One question follows immediately in every small QA team: if AI helped write these templates, does the helper itself need a risk assessment? §10 gives the proportionate answer.
AI-drafted templates need a proportionate risk assessment, not a per-document one
Audience question Q5"Would I need a risk assessment in having AI creating those templates?"
Spine
The controls apply to the tool that helps build the controls, in proportion to what it does. This section is the smallest application of the claim: a one-time classification of the authoring use, not a per-document exercise.
Short answer
Yes, but a proportionate one. For most teams it is a one-time classification, not a per-document exercise.
What matters is what the AI is doing:
AI role
Example
GxP risk
Required control
Authoring aid for a document a qualified human reviews and approves
Drafting a validation-plan template or SOP skeleton
Low The approved document is the controlled record, and the tool is like a word processor with suggestions
Register the use in the AI inventory; SOP for AI-assisted authoring; human review and approval as normal; no confidential data in non-approved tools
Generating validation content that feeds test decisions
Drafting acceptance criteria, test cases, or risk ratings
Medium Errors propagate into the validated state
Above, plus SME ownership of every criterion (Annex 22 §4.2 requires an SME to define acceptance criteria), a check that cited regulations actually exist and say what is claimed, and a coverage check against the requirements
Generating test data or labels
Synthetic defect images, AI-labelled test sets
High Annex 22 §5.6: "not recommended and any use hereof should be fully justified"
Avoid for test sets. If used for training, justify, document, and keep the test set real and human-adjudicated (§8.4)
Executing tests or producing evidence
AI runs test scripts and records pass/fail
High The tool is now part of the validation evidence chain
Validate the tool for that use (CSA: tools that directly support assurance need assurance proportionate to risk). Annex 11 §4.7 already requires a documented assessment of the adequacy of automated testing tools and test environments
Specific pitfalls with AI-drafted templates
Fabricated or misquoted regulatory references
The most common failure. Every clause citation needs checking against the source text.
Plausible omissions
The template looks complete but leaves out, say, test-data independence. Check it against a reference checklist such as Annex 22 §3–10.
Circularity
AI writing the acceptance criteria for AI. Keep criteria human-owned and baseline-anchored (§5).
Confidentiality
Pasting proprietary process details into a public LLM. Use approved enterprise tools only.
Recommendation
Add one line to the AI inventory, "Generative AI for GxP document authoring: low risk, human-approved outputs", backed by a short SOP. That is enough for the authoring-aid case. Escalate only when AI output feeds test decisions or evidence. Drafting SOPs sits at the lowest-risk end of most AI risk ladders. §17 maps its full ladder of use-cases onto the three tiers, and §14.6 extends the same rule to AI as a co-brain for the whole CSV process.
So what
The controls are built, positioned, and proportionate. A validated state for AI is a claim that must stay true after release. Part D is about keeping it true.
Drift, monitoring, and change control: how the claim stays valid after release.
Bridge. Part C built the controls and positioned the AI component where they can reach it. A validated state for AI is a claim about performance on a defined input space, and the world that supplies those inputs keeps moving. Part D is about keeping the claim true: what drift is and is not, in CPV and clinical terms first (§11); the audience's camera case, where a new label is a change rather than drift (§12); one site monitoring policy with the right alerts and no alert fatigue (§13); and change control end to end, so that every signal that needs a change gets one (§14).
Practitioner question Q11 (b) · not from the session"Our model is locked. What exactly drifts, how do we tell drift from a real process change, and when does drift become a change?"
Spine
This section keeps the claim's population component true. The claim was made for a defined input space; drift is the world leaving that space, and the first skill is to tell that apart from the process signals the AI exists to catch. The answers below lead with CPV and clinical examples; the camera case that the audience asked about is in §12, unchanged.
The vocabulary, once. A locked model's weights do not change (§4.1). Four things around it can:
Kind
What moves
CPV (UC-M2) example
Clinical (UC-D4 / data processing) example
Model changed?
Covariate (input) drift
The distribution of inputs
A new equipment train or site feeds the MSPC model; a sensor is recalibrated; sampling frequency changes
New sites or countries join the trial; a new TFL template; longer protocols with new section structures
No
Prior (prevalence) shift
How often each class occurs
A campaign mix change: 80% product A instead of 50%
A change in disease-stage mix; more amendments per protocol in a therapeutic area
No
Concept drift
What a label means
QA tightens a specification, so "out of trend" now means something new
A protocol amendment redefines an endpoint; ICH E3 expectations change for a section
No, but the model is now wrong
New class (open-set)
A category the model has never seen
A new failure mode (a raw-material interaction never in the NOC history)
A new adverse-event pattern; a new document type routed to the drafting agent
No, but it will force the new case into a known class
The drift FAQ
Show
1Does a locked model drift?Policy
No. Its weights and configuration are frozen; its outputs for any given input are the same as on the day of validation. What drifts is the relationship between the inputs it now sees and the input space it was validated on (Annex 22 §3.1, §10.4). "The model drifted" almost always means one of the four things in the table above happened, and each has a different owner and response.
2Is a vendor model update drift or change?LLM & vendor
Change, always. If the vendor swapped the model behind an unchanged feature name and you found out from behaviour rather than from a notice, that is an unauthorised change from your point of view (Annex 22 §10.2) and a contract failure (§17, element 4). Detection is a configuration and probe-set check (T8 OPS-06; §4.4), not a drift metric. The response is the T9 regression if the update is inside the envelope, or a Normal change if it is not (§14.2; scenario S2 in §14.5).
3What is "LLM drift" when the version is pinned?LLM & vendorClinical
Four different things, only one of which is drift:
Observed as "LLM drift"
What it actually is
Response
The retrieval index or corpus changed (new SOP versions, new approved sources, re-chunking)
A prompt, system instruction, temperature or tool definition changed
A change; if it happened without a record, an unauthorised one
Change control (§14.5, S4); investigate if unrecorded
The input population changed (new document types, new therapeutic areas, longer or differently structured inputs)
Covariate drift
Input monitoring (§13); scope check against T3; T9 extension or Normal change if outside the validated space
The vendor changed something behind the same version string (safety filters, tokenizer, serving)
A silent vendor change
Probe-set check daily; treat as unauthorised change; invoke the quality agreement
4A supplier or raw-material change shifts a CPV trend. Is that model drift or a real process signal?CPV
This is the question that matters most in CPV, and the answer is: disentangle the two before anyone touches the model. A real process signal is what CPV exists to catch. A new excipient supplier that changes blend behaviour, a new API lot that shifts dissolution, a granulation that runs wetter in summer: these are process changes, and the MSPC excursion is the model doing its job. Route them to deviation and change control on the process, not to model retraining. Retraining the model so the excursion disappears is tuning the alarm to the fault.
Model-input drift is the other case: the process is unchanged, but the data feeding the model has changed. A sensor recalibration, a renamed historian tag, a changed sampling interval, a unit conversion, an ERP master-data change that alters how batches are grouped. Here the process is fine and the model's inputs are wrong.
Disentangle with four checks, in order:
Deterministic path first. Did the raw CPPs and CQAs move on the validated SPC charts? If they did, it is process. The MSPC model and the SPC path see the same data; if only the model reacted, suspect the inputs.
Data lineage. Check the historian tags, calibration records, sampling configuration and MES event frames feeding the model (§8.4). A lineage change with no process change is input drift.
Contribution plot. Which variables drive the T² or SPE excursion (§8.5 explainability)? A single instrumentation variable points to inputs; a coherent group of process variables points to the process.
Reference test. A lab result or a manual review of the batch settles it.
Then the routing: process signal → deviation / CAPA / process change control; input drift → fix the data path, then assess the impact on the model (was the validated performance affected?); an approved and qualified new normal (for example, a second supplier qualified within specification) → NOC refresh under change control, pre-approved in T9 if it was anticipated (§14.2). The deviation record should state which of the three it was.
5Covariate, prior, concept and new class, in CPV terms: what does each one mean for the claim?CPV
CPV event
Drift kind
Is the claim still valid?
Response
A new site or equipment train starts feeding the model
Covariate
Not for that site: it is outside the validated sample space (T3)
New subgroup; extend the test set; regression on all subgroups; Normal change (§14.5, S1). Until then, route that site's signals to manual CPV review
Sensitivity per product is unchanged; overall precision may move
Check precision at the new prevalence; no model change; note it in the periodic review
A specification is tightened
Concept drift
No: the label "out of trend" has a new meaning
Change control on the specification triggers relabelling of the affected training and test data and a retest (§12, concept-drift row)
A new failure mode appears
New class
No, for that mode: the model has no concept of it
Detect through the "undecided" rate, human disposition disagreement, and deviations; intended-use update; new acceptance criteria against the manual baseline (§5); full retest (§12 decision tree)
6Do I need ground-truth labels to monitor?CPVClinical
Not continuously. Labels arrive late (confirmed dispositions, lab results, QC findings) or never. Monitor what you can observe every day: input-distribution distance, out-of-distribution and "undecided" rates, class-wise output rates, override and edit rates, checker findings. Then add periodic reference checks that supply ground truth on a schedule: confirmed dispositions reconciled monthly for CPV, a re-challenge with seeded historical drifts each quarter, an AQL re-inspection sample for a camera, a QC audit of a sample of AI-drafted sections for clinical documents. Annex 22 §10.3 (performance) and §10.4 (input) together describe exactly this pairing. §13.1 lists the five metric families.
7How often do I retest?Policy
Two clocks. A scheduled re-challenge by risk tier (§13.2): quarterly for a critical use, semi-annual for Tier 2, at periodic review for Tier 1. An event-driven retest after any change to a configuration item, after maintenance that touches the physical inputs (a camera's lighting, a sensor), and after any drift metric crosses its alert limit and the investigation cannot explain it. "Monthly accuracy" is not a retest schedule if you cannot get monthly labels (§3, row 6).
8Who decides when drift becomes a change?Policy
The monitoring owner investigates and documents (Level 0 in T8). The system owner and QA decide whether to restrict the use (Level 1). Drift becomes a change when the root cause is known and the response touches a configuration item or the input-space definition: the process SME owns whether T3's input sample space must be redefined (Annex 22 §3.1), and QA approves the change classification, Standard inside the T9 envelope or Normal outside it. §14.4 shows the chain: Event → Incident → Problem → Change Request.
9Does drift in a non-critical, human-in-the-loop use need change control?Policy
The drift needs an investigation and a record in the monitoring log. The response to it, if it is a prompt edit, an index refresh or a retrain, is a change, even at Tier 2, because it alters a configuration item that the assurance plan relied on. What scales with the tier is the depth of the regression, not whether a record exists. A Tier 2 drafting agent whose edit-distance metric climbs and whose prompt is then adjusted has had a Standard change (§14.5, S4).
10Seasonal or cyclic patterns: drift or not?CPVClinical
Cycles (summer humidity in granulation, shift patterns, campaign sequences, the year-end surge of clinical documents) are part of the input space if the training window covered at least one full cycle. If it did not, the first season the model meets is covariate drift outside the validated space, and it will look like degradation. Three defences: build the NOC set or training set across a full cycle; put the drift metric itself on an SPC chart with a seasonal baseline (§13.3); and pre-approve a "seasonal extension of the NOC set" in T9 so that the first winter is a Standard change rather than a surprise.
11A clinical case: the TFL template or SAP version changes. Drift?Clinical
Change. The input format of the drafting agent's sources has changed, the deterministic checker's rules may no longer parse the tables, and the grounding tests need to run again. If the template change was anticipated in T9 (for example, a sponsor-standard TFL update), it is a Standard change with the pre-approved regression; if not, it is Normal (§14.2). The agent has not drifted; its context of use has moved.
12Can we retrain continuously "to keep up"?CPVPolicy
Not for a critical use: continuous retraining makes the model dynamic, and draft Annex 22 excludes dynamic models from critical GMP use (§4.2). For a non-critical use, each retrain is still a change with regression on the locked test set, and the labels that feed it must be human-adjudicated, not harvested from production acceptances (Annex 22 §5.6 by analogy; §8.4). Otherwise the model learns the reviewers' automation bias.
13Our override rate fell to 1%. Is the model getting better?Policy
Possibly. Or the reviewers have stopped reading. The two are indistinguishable from the override rate alone; only a seeded-error check separates them (§5, §13.1 human-in-the-loop family). Watch overrides in both directions, and treat a sustained fall toward zero as an alert, not as good news.
Related on this page
§3 row 6: why "monthly accuracy" is not a monitoring plan
§4.1 the locked/retrained/dynamic vocabulary; §4.2 why continuous retraining is excluded from critical use; §4.4 pinned vendor versions
§5 the manual baseline for a new class and for the seeded-error check
§8.3 index and prompt as configuration items; §8.4 data lineage; §8.5 contribution plots
§13.1 the five metric families; §13.2 re-challenge by tier; §13.3 drift metrics on SPC charts
§14.2 Standard vs. Normal change; §14.4 Event → Incident → Problem → Change Request; §14.5 scenarios S1, S2, S4
§17 element 4: the vendor clause that makes a silent model swap a contract failure
So what
Drift is the world leaving the validated space, and the response depends on which of four things moved. The audience asked the same question about a camera, where the tempting answer, "a new label is drift", is exactly wrong. §12 works that case, unchanged from the session.
A new label is not drift. It is a change to the intended use.
Audience question Q2"For an in-line AI camera system trained and validated on a dataset of labels, when would drift occur? When a new label is introduced into the dataset? And then would we need to revalidate anytime a new label is introduced?"
Spine
This section keeps the claim true on the audience's own example. The vocabulary from §11 applies unchanged; the new element is the decision tree for what a new label triggers.
Short answer
Adding a new label is not drift. It is a change to the intended use (the output space), so it goes through change control, and it always triggers retesting.
The scope of that retesting depends on how the change is made. Drift is something different: the model has not changed, but the world it sees has. It needs its own monitoring.
Scenario
Automated visual inspection (AVI) of lyophilized vials on a filling line. A deep-learning classifier sorts each unit as accept, reject:particle, reject:crack, reject:fill-height, or reject:cake-defect. Validated under Annex 11 plus draft Annex 22. Acceptance was set against manual inspection using the Knapp–Kushner method, which USP <1790> references for showing an automated system performs as well as or better than manual inspection (§5).
Four kinds of change people call "drift"
Type
What changed
Camera example
Model changed?
How you detect it
Response
Covariate (input) drift
The distribution of images
New vial supplier with slightly different glass tint; LED ageing; lens film; new lyo cycle changes cake appearance
No
Input-space monitoring: image statistics (brightness histograms, embedding distance from the training distribution), out-of-distribution score (Annex 22 §10.4)
Investigate. If still within the validated sample space, document it. If not, it is a change: retest or retrain
Prior (prevalence) shift
How often each class occurs
A fill-pump issue raises fill-height defects from 0.1% to 1%
No
SPC on reject rate per class
Usually a process signal, not a model problem. Route to deviation. Check that precision holds at the new prevalence (§11, FAQ 4)
Concept drift
What a label means
QA tightens the particle-size reject criterion from 150 µm to 100 µm
No, but it is now wrong
Not detectable from data alone. Comes from change control on specifications
Change control. Relabel the affected training and test data. Retest
New class (open-set)
A defect type the model has never seen
A new "stopper-skirt deformation" defect appears after a stopper supplier change
No, but it will force the new defect into a known class, or into accept
Low-confidence rate, disagreement with periodic manual re-inspection, and complaints
This is where Q2's new label comes in: a new class means a new intended use
The most dangerous case is the unseen defect
A closed-set classifier has no "I don't know" option. It will put a stopper-skirt defect into whichever known class looks closest, which may be accept. That is why Annex 22 §9.2 asks for a confidence threshold with an "undecided" outcome, and why the monitoring plan needs periodic manual re-inspection of an AQL sample of accepted units. That sample is the only way to see what the model is missing.
Does a new label require revalidation? A decision tree
Reading the tree. Two decisions set the retest scope: whether the output space changes, and whether the new class is critical. New examples of existing classes are a targeted regression on the locked test set. A new critical class is a revalidation of the model component, while the platform IQ/OQ carries forward.
The rule that saves the most work
Write the predictable changes into the validation plan before go-live. List the anticipated change types (new examples of existing classes, recalibration after a lighting change, a new container supplier within spec), the pre-approved test protocol for each, and the acceptance criteria. FDA formalized this for AI-enabled medical devices as the Predetermined Change Control Plan (PCCP) (§4.3). That framework is not binding on pharma manufacturing equipment, but the same approach is a defensible way to structure a GMP change-control SOP. Changes inside the plan run as pre-approved protocols (T9; §14.2). Changes outside it, such as a new critical defect class, go through full change control (§14.5, scenario S3).
Monitoring plan for the camera (minimum)
Metric
Frequency
Alert rule
Owner
Reject rate per class (SPC, p-chart)
Per batch
Western Electric rules
Production / QA
Low-confidence ("undecided") rate
Per batch
> 2× validated baseline
MS&T
Input-drift score (embedding distance vs. training distribution)
Daily
Above a pre-set threshold
Automation / data science
Manual re-inspection of an AQL sample of accepted units
Per batch or per campaign
Any critical defect found → deviation
QA
Knapp re-challenge with the qualified defect kit
Periodic (e.g., quarterly) and after any maintenance
The same six rows, with illustrative limits, are the worked example in T8 Appendix A. §13 shows how they fit the site policy's five metric families.
So what
One camera, one CPV agent and one drafting agent each need a monitoring plan, and each plan so far has been built by hand. §13 turns them into one site policy, with parameters set per use.
One monitoring policy for every AI use, with the parameters set per use
Practitioner question Q13 · not from the session"How do we keep every AI system on the site compliant with the right alerts, without a different dashboard for each one and without alert fatigue?"
Spine
This section keeps the claim true as a policy, not as a series of one-off plans. One site policy fixes the metric families, the minimums by tier, the way limits are derived, the ownership and response ladder, and the alert-fatigue controls. T8 sets the parameters for each use, and T8 Appendix B carries the policy text in reference form. §8.7 gave the control; §11 gave the vocabulary; this section makes them the same for every AI function on the site.
13.1 Five metric families, every AI function
1
Inputs / drift
Whether the inputs are still inside the validated sample space
No ground truth needed
T8 §3 · runs continuously
2
Outputs / performance
Whether the validated claim still holds
Ground truth on a schedule
T8 §2, §4 · supplies the reference check
3
Human-in-the-loop
Whether the reviewers are still reviewing, and whether the AI is still helping
Seeded checks supply it
T8 §5 · runs continuously
4
System / vendor
Whether the configuration is what was validated, and whether it is about to change
No ground truth needed
T8 §6 · runs continuously
5
Data integrity
Whether every AI-touched record is attributable and complete
Full tableFive metric families: what each watches, typical metrics, whether ground truth is needed, and the T8 section5 rows
Family
What it watches
Typical metrics
Ground truth needed?
T8 section
1. Inputs / drift
Whether the inputs are still inside the validated sample space
Distribution distance vs. training data (embedding distance, population stability index); out-of-distribution and "undecided" rates; share of inputs from sources outside the validated scope; class prevalence on a p-chart
No
T8 §3
2. Outputs / performance
Whether the validated claim still holds
Sensitivity on the critical class; precision or false-reject rate; agreement with confirmed dispositions; periodic re-challenge with a qualified set; calibration (Brier score, band distribution)
Yes, on a schedule
T8 §2, §4
3. Human-in-the-loop
Whether the reviewers are still reviewing, and whether the AI is still helping
Acceptance rate (both tails); override and modification rates in both directions (AI said reject → human accepted, and the reverse); seeded-error catch rate; edit distance and time-on-review
Seeded checks supply it
T8 §5
4. System / vendor
Whether the configuration is what was validated, and whether it is about to change
Model, prompt, index and threshold hash checks; vendor version and deprecation date; probe-set behaviour check; latency; cost per decision; degraded-mode activations
No
T8 §6
5. Data integrity
Whether every AI-touched record is attributable and complete
Audit-trail completeness for AI events (output stored as generated, model version, sources, reviewer action); share of AI-originated fields without attribution; retention checks on AI context artifacts
Families 1, 3, 4 and 5 need no labels and run continuously; family 2 supplies the ground truth on a schedule (§11, FAQ 6).
13.2 Minimum metrics and frequency by risk tier
The tiers are the §17 tiers, set by the §8.2 risk method (model influence × decision consequence, plus Annex 22 criticality). The minimums below are a floor; T8 can add to them, never below them.
Tier
Typical uses
Minimum metrics
Frequency
Reference check (ground truth)
High (Tier 3, GxP-critical)Tier 3
AVI camera; MSPC in CPV with automated disposition support; PAT endpoint
All five families. At least: one drift score, one OOD/undecided rate, one out-of-scope input check; sensitivity on the critical class and false-reject rate; calibration; acceptance and override in both directions; seeded-error catch rate; configuration hash; vendor version and deprecation; audit-trail completeness
Inputs and configuration daily or per batch; performance per batch where a reference exists; human-in-the-loop weekly; calibration weekly
Per batch or campaign for the critical class (e.g., AQL re-inspection, confirmed dispositions); qualified re-challenge quarterly and after maintenance
Medium (Tier 2, GxP, human-in-the-loop)Tier 2
CSR drafting; deviation classification with investigator confirmation; CPV narrative agent
Families 1, 3, 4, 5, and one performance proxy. At least: undecided or out-of-scope rate; checker findings per item; acceptance and edit rate; seeded-error catch rate; configuration hash and vendor version; audit-trail completeness
Weekly, with configuration checks daily
Quarterly sample audit against an adjudicated reference; seeded-error check quarterly
Copying generic percentages is the failure §3 (row 5) describes. Derive limits from three inputs:
The validated performance and its confidence interval (T7). The action limit for a performance metric is the lower bound of the validated confidence interval, because below it the claim is no longer supported. The alert limit sits between the point estimate and the lower bound, set so that an ordinary run of bad luck does not trigger it; a common choice is the value that would be crossed by chance less than once a quarter at the sampling rate in use. Illustrative: validated particle sensitivity 0.97 with a lower 95% bound of 0.94 gives an action limit of 0.94 and an alert limit near 0.95 on the quarterly re-challenge (T8 Appendix A, PERF-03).
The consequence of failure (T2). A metric that guards a High-severity failure mode gets a tighter alert limit and a faster escalation SLA than one that guards a Medium. Every High failure mode in T2 must map to at least one KPI (T8 §8).
The baseline (§5). For human-in-the-loop metrics, the action limit on the seeded-error catch rate is the measured human baseline, because below it the human+AI system is worse than the process it replaced.
Alert limit vs. action limit, illustrated. The action limit is the lower bound of the validated confidence interval, because below it the claim is no longer supported. The alert limit sits between the point estimate and that bound, set so that ordinary variation does not trigger it. Values are the illustrative ones from point 1 above.
SPC on the metrics themselves. Drift scores, undecided rates, override rates and edit distances are process data. Put each on a control chart (individuals or EWMA for continuous metrics, p-charts for rates) with limits set from the first validated weeks of operation, and use run rules (Western Electric or equivalent) as the alert condition rather than a single fixed threshold. Cyclic patterns get a seasonal baseline (§11, FAQ 10). Changing a limit is a change to the monitoring configuration CI (§14.4, rule 1).
13.4 Ownership, escalation and the response ladder
0
Investigate
Any alert limit, or a run-rule violation on a metric chart
Monitoring owner · Event → Incident
1
Restrict to advisory
Action limit on a critical KPI; or an unexplained alert repeated
System owner + QA · Deviation; Problem
2
Retrain / reconfigure
Confirmed drift or degradation with a known cause
System owner + SME + QA · Change record
3
Retire / suspend
Critical error reached a GxP record, or performance cannot be restored
QA · Deviation / CAPA; registry → deprecated
Level
Trigger
Response
Decision owner
Escalation SLA (illustrative)
Record
0 · Investigate
Any alert limit, or a run-rule violation on a metric chart
Confirm the signal; check for process and data-path causes before suspecting the model (§11, FAQ 4); document
Critical error reached a GxP record, or performance cannot be restored
Suspend the AI function; revert to the manual process; impact assessment on past decisions using the AI audit data (§15, §16)
QA
Immediate containment; impact assessment opened within 1 working day
Deviation / CAPA; registry state → deprecated
Every Level 1–3 event is an input to the next periodic review (T10) and a trigger to re-assess T2.
13.5 Tie-in to change control
The ladder is the front end of §14.4. An alert-limit breach opens an Event and, if confirmed, an Incident (Level 0); an action-limit breach opens a Problem (Level 1); a Problem whose root cause touches a configuration item becomes a Change Request, Standard inside the T9 envelope and Normal outside it (Level 2). A vendor deprecation notice enters the same chain as a planned change. The monitoring configuration is itself a CI, so tuning a limit is a Standard change if the range was pre-approved in T9, and a Normal change if not.
Alert limit→Event → Incident= Level 0Action limit→Problem= Level 1Root cause touches a CI→Change Request (Standard in T9 · Normal outside)= Level 2See the loop in §14.4 →
13.6 Alert-fatigue controls
Monitoring that people ignore is a paper control, like an unmeasured reviewer.
Tiering of alerts. Only action-limit breaches on Tier 3 KPIs page anyone in real time. Alert-limit breaches go to a daily digest; Tier 2 alerts to a weekly review. The routing is written in T8, not left to the tool's defaults.
Suppression windows with justification. A known cause (an LED replacement scheduled for Tuesday; a planned index refresh) may suppress a specific metric's alerts for a bounded window. The window, the reason and the approver are recorded, and the metric is checked when the window closes. Open-ended suppression is not allowed.
De-duplication and correlation. One physical cause (a lighting change) can trip a drift score, an undecided rate and a reviewer-agreement metric at once. The Event layer correlates them into one Incident (§14.4), so three alerts do not become three investigations.
Alert usefulness review. Each periodic review reports, per KPI: alerts raised, alerts that led to a finding, alerts closed as false alarms. A KPI whose alerts never lead to a finding gets its limit re-derived or is retired; a KPI that never alerts but whose failure mode still exists gets a reference check added.
13.7 Monitoring the monitors
Thresholds age. Each periodic review (T10) re-derives the limits from the current validated performance and the observed distribution of each metric, reviews the false-alarm and missed-signal rates from §13.6, confirms that every High failure mode in the current T2 still maps to a live KPI, and checks that the reference checks actually ran at the stated frequency. The output is either "limits confirmed" or a change to the monitoring configuration CI. The review also asks whether the intended use still matches actual use (§8.8), because scope creep is the drift the metrics cannot see.
13.8 Two worked examples (all values illustrative)
UC-M2, the CPV signal agent
UC-M2, the CPV signal agent (Tier 3 for the MSPC and SPC components; Tier 2 for the narrative).
Family
KPI
Frequency
Alert
Action
Owner
Response
Inputs
Share of batch-phases outside the NOC envelope for non-process reasons (after the §11 FAQ 4 triage)
Per batch
> 3% rolling 30 days
> 5%
MS&T
L0; NOC refresh via T9 if a qualified change caused it
Inputs
Inputs from a site, product or equipment train not in the validated scope
Per batch
Any
Any
System owner
Route that site to manual CPV review; Normal change (S1)
Outputs
Precision (signals confirmed / signals raised)
Rolling 30 days
< 0.85
< 0.80
MS&T
L0 → L2 if persistent
Outputs
Seeded historical drifts detected on re-challenge
Quarterly
< point estimate
< validated lower bound
Validation
L1 → L2
Human
Coach-mode acceptance, rolling 14 days
Weekly
< 0.75 or > 0.97
< 0.70 or > 0.98
System owner
Low: fitness review. High: seeded-error check
Human
Seeded-error catch rate of scientists on narratives with planted numeric errors
Quarterly
< validated
< manual baseline
QA
Retrain reviewers; restrict narrative
System/vendor
Hash check on MSPC model, SPC ruleset, prompt, thresholds
Daily
Any mismatch
Any mismatch
Automation
Stop; unauthorised-change investigation
System/vendor
Narrative LLM version vs. pinned; days to vendor deprecation
Daily
< 90 days to deprecation
Version mismatch
Platform
Planned change (S2)
Data integrity
Signals with a disposition but no stored narrative-as-generated, model version or reviewer rationale
Weekly
Any
> 1%
QA
Fix the record path; deviation if a signed record is affected
Response example
A T² excursion cluster begins after a new lactose supplier's first three lots. The SPC path shows blend-time and moisture moved with it: process signal. Deviation opened on the process; the model is not touched; the excursion stands as a correct detection in the precision count. Two months later the second supplier is qualified within specification, and the NOC set is extended under the T9 "qualified supplier within spec" pre-approved change with regression on the locked test set.
UC-D4, CSR drafting with generate-then-check
UC-D4, CSR drafting with generate-then-check (Tier 2).
Family
KPI
Frequency
Alert
Action
Owner
Response
Inputs
Sections from a therapeutic area, study design or TFL template outside the released scope
Per section
Any
Any
Clinical system owner
Route to manual drafting; T9 if pre-approved, else Normal change
Doscierge's first evaluation produced about 288 findings per test package, and precision engineering brought that down to about 73 (parent §17). For this policy the lesson concerns monitoring, not only engineering. Finding volume per package belongs in the outputs family with an action limit, because a reviewer handed 288 findings stops reading them: automation bias arriving through alert fatigue (§13.6).
So what
The policy keeps the claim true across the site and turns every confirmed signal into an Event, an Incident, a Problem or a Change. §14 shows the change side of that chain end to end, from CMDB to release.
Change control works when four records stay linked
Author's question Q10"Take specific use-cases and show how change control comes together: a ServiceNow-style model (business application, the applications beneath it, their specific change control), IQ, OQ, PQ, release for use in production, and ongoing monitoring for an AI-based system. Then add AI as a co-brain helping with the CSV/CSA process itself (risk, requirements and the other components). What does the human do and stay accountable for, and how does all of this map to change control?"
Spine
This section keeps the claim true through change. Every change to a model, prompt, index, test set or threshold either stays inside a pre-approved envelope or re-proves the claim in proportion to the change. Sub-numbers 14.1–14.8 are what the rest of the page cites as Q10.1–Q10.8.
The short answer
Change control for AI works when four records stay linked and in step:
The service model (CMDB)
What is running and what depends on it.
The AI registry and agent card
How each AI component behaves and was validated (§8.6).
The change record
What is changing, who approved it, and the evidence.
The validation record (VLMS or eQMS)
IQ/OQ/PQ protocols, results, and the summary report (VSR).
The AI-specific move is to make each AI component (model, prompt/configuration, retrieval index, locked test set, tool server) a configuration item of its own, with a versioning rule and a pre-approved change envelope. A model retrain then becomes a traceable CI change with a defined validation scope, not an invisible "data refresh". AI can do much of the legwork: impact analysis, drafting, test generation, evidence review. It never sets acceptance criteria, approves, or signs.
14.1 The service model: where AI components sit in a ServiceNow-style CMDB
This uses ServiceNow's Common Service Data Model (CSDM) hierarchy, extended with AI configuration-item classes. The class names are illustrative. Most organizations model them as custom CI classes, or sync them from the AI registry. Expand the narrative agent to see the AI configuration items nested beneath it.
Deterministic (Cat 5 / rules)Static ML (Annex 22 in scope)LLM / generative (Pattern A, HITL)Validation assetPlatform / vendor
Business capabilityManufacturing Quality Assurance
Business application
Continued Process Verification (CPV) Platform
GxP = GMPcriticality = HighSystem OwnerBusiness Process OwnerQA ownervalidation status = ValidatedAI risk tier = 3 (critical components present)
Application serviceCPV Signal Service – PROD (Sites A, B, C)← change control applies here
SPC/Capability query layerv3.2deterministicGAMP Cat 5 · Pattern B core
MSPC model (PCA, NOC set ref)v1.4static MLAnnex 22 in scope · QA-approved in registry
Narrative agentv2.3LLMnon-critical · HITL · Pattern A
3 nested AI configuration items
Foundation model endpointscoped-small@pinned-2026-07vendor
Text alternative: Business capability Manufacturing Quality Assurance contains the business application CPV Platform (GxP GMP, criticality High, validated, AI risk tier 3). Its production application service, CPV Signal Service PROD for Sites A, B and C, is where change control applies, and holds these configuration items: SPC/Capability query layer v3.2 (deterministic, GAMP Cat 5, Pattern B core); MSPC model v1.4 (static ML, Annex 22 in scope, QA-approved in registry); Narrative agent v2.3 (LLM, non-critical, human in the loop, Pattern A) with nested items foundation model endpoint scoped-small pinned 2026-07, prompt and instruction set sha 9f2c version 12, and tool allowlist of spc_calculator, capability_calculator and kg_lookup; Orchestrator workflow v2.3 (deterministic control flow); Deterministic checker ruleset v1.7; Locked test set cpv-test-2026Q2 (validation asset, access-controlled); Monitoring config v1.2 (alert and action limits are configuration too). It depends on data product batch_phase_summary v4.1, the PI historian and MES event frames. Two further application services, VAL/QA (where IQ and OQ execute) and DEV/Lab (no GxP data rules, no validation state), sit alongside. Business application Clinical Document Authoring (UC-D4; GxP GCP; Tier 2; Pattern A) has the CSR Drafting Service PROD with a drafting agent v1.9 (LLM) over a pinned vendor LLM endpoint, a prompt set and an approved-source index v5, plus a deterministic checker MCP tool v2.4 (Cat 5). Business application Automated Visual Inspection Line 3 (GxP GMP; Tier 3; Annex 22 critical) has AVI Line 3 PROD with the vendor inspection machine (Cat 4 platform), classifier model v3.0 (static, 5 classes), camera and lighting recipe v7, and Knapp defect kit KK-2026 as the qualified challenge set.
Why the granularity matters
Change control, impact analysis and validation scope all act on CIs. If the model, prompt and index are hidden inside one "application" CI, every change looks either trivial ("config tweak") or total ("revalidate everything"). Modelling them separately makes change-level risk-based validation possible, which is the whole point of CSA. The CI list is the §8.3 configuration baseline, seen from the ITSM side.
Attributes to add to each AI CI (mirrored from the agent card, parent §6):
GxP impactposture (Pattern A/B; Annex 22 in or out of scope)risk tierversioning rule (major/minor/patch)linked T3/T6/T7/T9 document IDslocked-test-set referencemonitoring KPI set (T8)last periodic review (T10)vendor model-deprecation date
14.2 Change types: the T9 envelope becomes ServiceNow change models
ServiceNow change type
When it applies to an AI CI
Validation scope
Approvals
Standardpre-approved template
Changes inside the T9 envelope: a vendor model patch with an unchanged API; retraining on new examples of existing classes; a prompt edit in the T9 catalogue; a monitoring threshold tuned within a pre-set range
The T9 pre-approved regression on the locked test set; automated evidence attached to the change
Pre-approved by QA when the template was created. Execution evidence reviewed by the system owner; QA is notified or samples
Normaloutside the envelope
A new output class; a new site or population (new subgroup); a new intended use; a model architecture change; an out-of-envelope vendor model swap; a new tool added to an agent; regression failure on a Standard change
Risk-based IQ/OQ/PQ scope from an updated T2; T3 updated if the context of use changes
Change Advisory Board (CAB) with Quality as a mandatory approver for GxP CIs; business process owner; security if identity or tools change
Emergencycontain first
Harmful behaviour in production: an action limit breached on a critical KPI, a prompt-injection incident, a vendor outage
First contain: kill switch, roll back to the last validated version, or drop to advisory or manual mode (parent G7 degraded mode). Then validate after the fact
Emergency CAB; QA retrospective approval; a deviation record is opened
Automation hook
CI/CD pipelines raise the change record automatically (e.g., DevOps change automation), attach the IQ evidence and the regression report, and use the risk rating to route Standard vs. Normal. A failed regression gate blocks promotion; nobody has to spot it.
14.3 IQ, OQ, PQ and release for use, redefined for AI components
Stage
Classic meaning
For an AI component
Evidence (mostly automated)
Human accountable
IQ
Installed as specified
The model artifact hash matches the registry; container image and IaC match; the prompt/config version, tool allowlist, retrieval index version and pinned vendor model version are verified (§4.4); the audit-trail and e-signature configuration is present
Pipeline-generated IQ report linked to the change record
Platform/IT owner executes; QA reviews
OQ
Operates per specification across ranges
Performance on the locked, independent test set by subgroup against pre-approved criteria (T6). Also: confidence gating and "undecided" routing, explainability output (Annex 22 §8), trajectory and red-team tests (agents), Part 11 checks (the AI cannot sign; AI contribution is attributed; §15), degraded-mode behaviour
Test report with 95% confidence intervals; the log of access to the locked test set
Validation lead executes; the SME owns the criteria; QA approves
PQ
Performs in the real process
Shadow or parallel run in production conditions: real inputs, real users, a defined period or batch count. Compare against the human baseline (§5). Measure human-in-the-loop effectiveness (seeded-error catch rate) and confirm the monitoring KPIs are live and producing data (§13)
PQ report; agreement statistics vs. the manual process; reviewer catch-rate results
Business process owner executes; QA approves
Release for use
Validated system handed to production
VSR (T7) approved with a scoped release (sites, products, TAs, document types); registry state candidate → factory; T8 monitoring and T9 addendum effective from day one; users trained, SOP updated
CR moves to Implement; CI validation status becomes Validated vX; release notes
System owner and QA sign the VSR (Part 11 e-signatures)
Post-implementation review
Did the change achieve its aim?
Check the first 30–90 days of T8 KPIs against validated performance before closing the change
PIR task on the change record
System owner
Scaling IQ/OQ/PQ to the change: a Standard change usually needs IQ plus targeted OQ (the regression) with no PQ. A Normal change for a new subgroup needs IQ, OQ on the new and all existing subgroups, and a short PQ on the new subgroup. A new critical class needs the full sequence.
14.4 Ongoing monitoring: how production signals flow back into change control
The loop. Alert limits raise Events and Incidents; action limits, drift and override anomalies raise Problems; vendor notices and periodic-review decisions raise Changes directly. Every change runs through a validation scope sized to it, is released, reviewed, and resets the monitoring baseline.
Three rules keep this defensible:
Monitoring configuration is a CI. Changing an alert or action limit is a change, even though no model changed (§13.3).
Not every signal is a model problem. A higher reject rate on the camera is often a process signal, so it goes to a deviation, not to a model change (§12 prior-shift row; §11 FAQ 4). The Problem record forces that triage.
Periodic review (T10) is a scheduled ServiceNow task on the Business Application, and its decision (continue, retrain, restrict, retire) is recorded and becomes a change where needed.
This is the same chain as the §13.4 response ladder: Level 0 is the Incident, Level 1 the Problem, Level 2 the Change Request.
14.5 Four scenarios run through the whole chain
S1 · CPV, add Site D
S1
CPV signal agent: extend to Site D
Normal change
Trigger
Business decision to extend CPV to Site D
CIs affected
MSPC model (new NOC/subgroup), monitoring config, data-product dependency
Change type
Normal: new subgroup, outside the envelope
AI co-brain work
Traverses the CMDB for dependencies; drafts T2 and T3 updates; proposes a Site D test-set stratification
IQNew site connectors, NOC reference, model v1.5 hash
OQLocked test set, all sites, plus a new Site D subgroup
PQShadow on 30 Site D batches vs. manual CPV review
ReleaseVSR addendum; scope now A–D
Monitoring changeSite D KPIs added; drift baseline set
Human accountable
MS&T owner (criteria); QA (approval)
S2 · CSR drafting, vendor retires the LLM version
S2
CSR drafting agent: the vendor retires the LLM version
Standard if the regression passes · Normal if it fails
Trigger
Vendor deprecation notice, 90 days out
CIs affected
Vendor LLM endpoint CI, possibly the prompt set
Change type
Standard if the T9 regression passes; Normal if it fails
AI co-brain work
Finds every service using the model; drafts the CR; runs the pre-approved regression; summarises the diffs
IQNew endpoint and version pinned
OQFull T6 regression, N runs per case; 0 unsupported claims
PQNot required (Standard), or short parallel run (Normal)
ReleaseChange closed; registry version bump (minor)
Monitoring changeWatch edit distance and unsupported-claim rate for 30 days
Human accountable
Clinical system owner; QA reviews the evidence
S3 · AVI camera, new critical defect class
S3
AVI camera: a new critical defect class
Normal change · major version
Trigger
Stopper-skirt deformation seen after a stopper supplier change
CIs affected
Classifier model (output space), defect kit, T3
Change type
Normal: new critical class, major version
AI co-brain work
Drafts the T3 intended-use update and new-class criteria options (ranges only); clusters candidate images for SME labelling
IQModel v4.0 hash; recipe unchanged
OQFull retest of all classes, explainability review, Knapp comparison for the new class
PQKnapp re-challenge plus a period of AQL re-inspection of accepted units
ReleaseVSR v4; scope unchanged except the new class
In Doscierge, every rule change must pass regression against the locked evaluation set before it is merged (parent §17). That is the T9 envelope (§14.2) in engineering form. A rule edit that passes is a Standard change whose evidence is generated automatically. One that fails, or a model-version change, becomes a Normal change that re-runs the full set (§1.3, stage 8).
14.6 AI as a co-brain for CSV/CSA: what it does, what the human owns
The same discipline applies to the helper as to any AI (§10). The co-brain is a Tier 1–2 tool that produces drafts and analysis. It does not produce decisions. Its outputs become records only once a human approves them.
Full tableCSV/CSA activities: what the AI co-brain does, what the human owns and signs, and the guardrail on the helper11 rows
CSV/CSA activity
What the AI co-brain does
What the human does and is accountable for
Guardrail on the helper
GxP assessment (T1)
Pre-fills from the CMDB and vendor documents; proposes classification with rationale
System owner and QA decide the classification
Its classification is a suggestion, shown with its evidence
Risk assessment (T2)
Proposes failure modes from the library and past incidents; drafts FMEA rows
SME and QA set severity, probability and detectability ratings and accept the residual risk
Cannot rate or accept risk; every row needs a human owner
Requirements / context of use (T3)
Drafts testable requirements from the SOP and process maps; flags untestable wording
The process SME owns intended use and the sample space (Annex 22 §3.1)
Cannot define acceptance criteria (Annex 22 §4.2)
Change impact analysis
Traverses CMDB dependencies and the registry "consumers" list; lists affected services, documents and SOPs; proposes Standard or Normal
The change owner confirms the scope; QA approves the classification
Its impact list is verified against the CMDB query result, not taken on trust
Test design (T6)
Generates test scripts and edge cases; proposes stratification
The validation lead approves the protocol before execution
No AI-generated test data or labels for the locked set (Annex 22 §5.6)
Test execution
Runs automated suites; collects evidence
The tester attests execution
The tool is itself assured for this use (CSA: tools that support assurance)
The AI that authored a script cannot be the only reviewer of its results (independence)
Deviation triage
Classifies test deviations; drafts the investigation
The validation lead and QA disposition them
Its classification is advisory
VSR drafting (T7)
Drafts the report from the evidence, citing each item
System owner and QA sign (Part 11)
Every statement is traced to evidence; numbers checked deterministically
Monitoring and periodic review (T8/T10)
Summarises trends; drafts the review; suggests a decision
QA and system owner decide: continue, retrain, restrict or retire
The suggested decision is labelled as such
Regulatory watch
Tracks guidance changes (Annex 22 final, Annex 11) and maps them to affected CIs and SOPs
QA and regulatory decide the impact
Citations checked against the source text
AI proposes, humans dispose, and the change record shows which was which.
The accountability rule in one line. Every co-brain output attached to a change record is tagged as AI-generated with its tool version, and the human decision on it is captured. This is the same Part 11 attribution pattern as §15, applied to the validation process itself.
14.7 RACI for an AI change
Full tableRACI for an AI change: system owner, business process owner or SME, QA/CSV, ML/platform engineering, CAB, and the AI co-brain10 rows
Activity
System Owner
Business Process Owner / SME
QA / CSV
ML / Platform Engineering
CAB
AI co-brain
Raise and classify the change
A
C
C
R
I
Drafts
Impact analysis
A
C
C
R
I
Drafts
Update T2 / T3
A
R (criteria, sample space)
C
C
Drafts
Approve Standard / Normal classification
C
C
A
I
I
—
IQ
A
C
R
Collects evidence
OQ (incl. locked test set)
A
C (criteria owner)
A (approval)
R
Runs suites
PQ
C
R
A
C
Summarises
Release for use (VSR)
A
C
A (co-sign)
I
I
Drafts VSR
CAB approval (Normal)
R
C
C (mandatory for GxP)
C
A
—
Monitoring response and periodic review
A
C
C
R
I
Summarises
R = responsible, A = accountable, C = consulted, I = informed.
14.8 How it all maps together
Four layers, one loop. The registry and the CMDB stay in sync; design-time templates and the change request feed the execution chain; operation feeds back as new change requests. The T9 envelope is what makes a change Standard; the T10 review is the scheduled task that decides continue, retrain, restrict or retire.
What to implement first (small or mid-size company)
Add AI CI classes, or an "AI component" related list, under existing Business Applications.
Turn the T9 catalogue into Standard change templates.
Make QA a mandatory approver for Normal changes on GxP AI CIs.
Route T8 action-limit breaches into Incident and Problem records.
Schedule T10 as a recurring task.
Everything else, including pipeline automation and the co-brain, can follow.
§11 FAQ 4 process-vs-model triage inside the Problem record; §12 prior-shift row
§13.3 limits as configuration; §13.4 the response ladder this chain implements; §13 live KPIs confirmed at PQ
§15 Part 11 checks in OQ and the attribution pattern the co-brain follows
Parent article §6 agent card attributes and G7 degraded mode
So what
The claim now survives change: every alteration is either pre-approved or re-proven in proportion. What remains is the record of all of it. Part E asks what Part 11 and ALCOA+ require when a model, not only a person, contributed to the record.
Part 11 and ALCOA+ when a model contributes to a record.
Bridge. Part D kept the claim true in operation: drift detected, changes controlled, every signal routed. Trust in GxP, though, rests on records: what was decided, by whom, on what evidence. Part E asks what Part 11 (§15) and the data-integrity guidance behind ALCOA+ (§16) require when a model contributed to the record. The answer in both cases is the same: the regulation does not change, but what each control must capture does.
§15 · Part E ·Audience question Q9 (after the session)
No product "meets Part 11". What changes is what each control must capture.
Audience question Q9 · after the session"Is there any AI software product on the market that would meet 21 CFR Part 11 compliance? Probably not, so how would the guidance differ from the most common digital product to an AI product?"
Spine
This section makes the claim recordable: the AI output, its version and evidence, and the human's decision on it, each captured as its own attributed event, so that any past decision can be reconstructed.
Short answer: the question contains a category error
No software product "meets Part 11", whether it uses AI or not. Part 11 applies to the regulated company's records and signatures, and compliance comes from three things together: the product's technical controls, how the company configures it, and the company's procedures (training, accountability, record retention). The same holds for classic systems: no vendor can make a product "validated" or "Part 11 compliant" on the customer's behalf. What a vendor can offer is a product that is Part 11-capable: audit trail, access control, e-signatures, record copies, and retention.
That leaves two real questions:
1
Are there AI-enabled products with Part 11-capable controls?
Yes. Validated eQMS, document management, LIMS and EDC platforms increasingly ship AI features inside a record layer that already has an audit trail and e-signatures. The open question is not the platform. It is whether the AI feature's contribution is captured by those controls. Often it is not: the audit trail records that the user saved the record, not that 70% of the text came from a model.
2
Can a general-purpose chatbot be used for GxP records?
Not as it ships. Consumer and general enterprise chat tools usually lack record-level audit trails, record retention under your control, and signature binding. They can be used in Part 11 scope only inside a wrapper that supplies those controls (the parent article's platform contract, UC-H1, §8.8), or kept outside record creation entirely (§10, Tier 1).
Part 11 does not change for AI
The regulation is technology-neutral. What changes is what each control has to capture, because an AI system adds a new actor (the model), a new input (the prompt and retrieved context), and output that cannot be reproduced on demand.
Clause by clause: what AI adds
Part 11 control
Typical digital product
What AI adds
Practical control
§11.10(a) Validation (accuracy, reliability, consistent intended performance, ability to discern invalid or altered records)
Functional testing against specifications
"Consistent intended performance" becomes a statistical claim for a context of use, and it can decay without any code change
Performance claim with acceptance criteria per subgroup (T3, T6), plus monitoring (T8; §13). You validate the controls around the model for this use, not the model in the abstract
§11.10(b) Accurate and complete copies
Re-render or export the stored record
A probabilistic model cannot regenerate the same output later, and the vendor model version may be retired
Store the output as generated, together with the prompt/input, retrieved sources, model and version, configuration version, and confidence. Never rely on regenerating it (§16, "Original")
§11.10(c) Protection and retrieval over the retention period
Database backup and archive
The artifacts needed to understand the record (model version, prompt template, retrieval index snapshot) sit with a vendor who may deprecate them
Keep AI context artifacts in a store you control. Add contract terms for model-version notice and data export (§17)
§11.10(d) Limiting system access
User accounts and roles
The agent acts as a principal. It can read and write through tools, often with broad service credentials
Give each agent its own identity, a tool allowlist, and least-privilege scopes. No shared or human credentials for agents (parent G5)
§11.10(e) Audit trail (secure, computer-generated, time-stamped record of operator entries and actions that create, modify or delete records)
Who changed what, when, old and new values
The model is a new actor that is not an "operator", so the trail must show the AI's contribution separately from the human's
Record AI-generated content as a distinct event attributed to the agent identity and model version. Then record the human action on it: accept, modify (with the diff) or reject, with a rationale (parent G2; §7 "Audit trail"). ALCOA+ "Attributable" now means attributable to a human or to an identified AI version (§16)
§11.10(f) Operational system checks (permitted sequencing of steps)
Workflow engine enforces the order
An agent that plans its own steps can skip or reorder them
Deterministic orchestration for regulated sequences. The model proposes, and the workflow enforces the order (T5; §6 Step 2)
§11.10(g) Authority checks
Role-based permissions
An agent may try to take an action that needs authority: approve, release, close a CAPA
Agents never hold authority for signature-bearing or disposition actions. Those routes always go to a qualified human (confidence gating, parent G7)
§11.10(h) Device checks (validity of the source of data input)
Instrument or terminal identity
AI output is itself a source of data entering the record
Tag every AI-originated field with its source (agent ID, model version) so reviewers and inspectors can tell machine-originated data from human- or instrument-originated data
§11.10(i) Training
System-use training
Reviewers must understand AI failure modes, especially automation bias
Training in how to review AI output, plus seeded-error checks of reviewers (§5, §7)
§11.10(j) Accountability for e-signatures
Signature policy
"The AI wrote it" is not a defence
The policy states that the signer is accountable for AI-assisted content they sign, exactly as for content they wrote
§11.10(k) Systems documentation control
SOPs, configuration specifications
Prompts, system instructions, tool definitions, retrieval sources and thresholds are system documentation
Put them under version and change control as configuration items (T5, T9; §8.3, §14.1)
§11.50 / §11.70 Signature manifestation and linking
Name, date/time, meaning; signature bound to the record
The meaning of a signature on AI-assisted content should be explicit
Use a signature meaning such as "Reviewed and approved, including AI-assisted content", and bind the signature to the exact output version reviewed
§11.100–§11.300 Electronic signatures
Unique to one individual; identification components
An AI can never sign
Hard technical block: no agent identity can apply an e-signature. Test this in OQ (§14.3)
Scope: when is AI output a Part 11 record?
Use the same test FDA's 2003 Part 11 Scope and Application guidance applies to any electronic record. It is in scope when a predicate rule requires the record, or when you rely on it to carry out a regulated activity or decision.
AI output
Part 11 record?
Why
Brainstorming a draft SOP structure that a human then rewrites
Usually no
The approved SOP is the record. The AI is an authoring aid (§10, Tier 1)
AI-drafted deviation investigation text kept in the final report
Yes, as part of that record
The report is a GMP record. The AI contribution must be attributable (§11.10(e))
AI classification that routes a deviation (e.g., minor vs. major)
Yes
It influences a regulated decision, so the output, confidence, model version and human confirmation are all retained
It is a batch-record decision, Annex 22 critical, with Annex 11 and Part 11 controls on the result
CSR section drafted by UC-D4
Yes
It is a submission document. The drafting history, checker findings and sign-off form the audit record
Chat transcripts used to decide something regulated
Yes, if relied on
Reliance makes it a record. Keep it in a controlled store, or do not use chat for regulated decisions
How to evaluate an AI product for Part 11 readiness (vendor questions)
Does the audit trail record AI-generated content as a separate event, attributed to a model or version, with the human's accept, modify or reject captured afterwards?
Is the as-generated output stored alongside the prompt/input, retrieved sources and configuration version, and can it be exported for the full retention period?
Which model version produced each output, how much notice do you get before a model changes, and can you pin versions (§4.4)?
Can the AI ever apply an e-signature, approve, release or close a record? The answer must be no.
Do agents have their own identities with least-privilege access, or do they run on user or service credentials?
Is customer data used to train shared models? Where are prompts and outputs stored, and for how long? Does the vendor's retention setting conflict with yours? A "zero data retention" option on the model API is fine only if your own record store keeps the Part 11 copy.
Can AI features be switched off per tenant and per module until you have assessed them (§17)?
Worked contrast
Common digital product
An eQMS deviation module without AI
Validate the workflow, audit trail, e-signatures and access. The investigator writes the text. Part 11 evidence covers who entered what, when, and who signed.
The same module with AI root-cause suggestions (UC-M5 pattern)
Everything above, plus
the suggestion is stored as generated, with model version and evidence citations
the investigator's accept, modify or reject and rationale are in the audit trail
AI-originated text is tagged in the final report
the signature meaning covers AI-assisted content
the AI cannot close the deviation or approve the CAPA
T8 monitoring tracks the acceptance rate and seeded-error catch rate (§13)
a vendor model change triggers the T9 change protocol (§14.5, S2)
Bottom line for the questioner
Do not look for a "Part 11-compliant AI product". Look for a Part 11-capable platform whose AI features are attributable, retained as generated, version-pinned, barred from signing, and switchable. Then close the remaining gaps with your own procedures and the T1–T10 controls.
Related on this page
§4.4 pinned versions and managed deprecation (vendor question 3)
§5 and §7 seeded-error checks of reviewers (§11.10(i))
§6 Step 2: deterministic orchestration in the worked example (§11.10(f))
§8.3 and §14.1 prompts, tools and thresholds as configuration items (§11.10(k))
§10 Tier 1 authoring aids kept outside record creation
§13 monitoring under §11.10(a); §14.3 the OQ test that an agent cannot sign; §14.5 S2 vendor model change
§16 ALCOA+ principle by principle, including "Original" and "Enduring"
§17 vendor clauses for model-version notice, data export and per-tenant switch-off
So what
Part 11 says what a record and a signature must be. The data-integrity guidance says what the data inside the record must be. §16 applies ALCOA+ principle by principle to a record that a model helped produce.
ALCOA+ for an AI-assisted record: attributable to a human or to an identified model version
Practitioner question Q12 · not from the session"What does ALCOA+ actually require of a record that AI helped produce, and what do we have to capture?"
Spine
This section makes the claim trustworthy as data. Every ALCOA+ principle still applies; the AI adds a second author whose contribution, version and sources must be captured as metadata of the record, and adds new ways for the data to fail.
16.1 The governing guidance, and what each adds
None of these documents mentions AI. All of them define the properties a GxP record must have, and those properties are what the AI-assisted record must still show. Titles and dates verified against the issuing bodies (§19).
Full tableGoverning data-integrity guidance, when it was issued, and what it contributes to the AI case5 rows
Guidance
Issued
What it contributes to the AI case
MHRA, 'GXP' Data Integrity Guidance and Definitions (Revision 1)
March 2018
Definitions used across GxP: data, raw data, metadata, audit trail, data lifecycle, data governance; ALCOA+ as the property set; expectations for hybrid systems and for validation of systems that generate records
PIC/S PI 041-1, Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments
In force 1 July 2021
The inspector's view: data governance, data criticality and inherent integrity risk, audit-trail review, outsourced activities and computerised systems; the reference for GMP/GDP inspections in PIC/S countries
FDA, Data Integrity and Compliance With Drug CGMP: Questions and Answers
December 2018 (final)
The CGMP anchor in the US: metadata defined as the contextual information required to understand data; audit trails as metadata; audit-trail review expectations; shared logins and system controls
WHO, Technical Report Series 1033, Annex 4, Guideline on data integrity
2021
Data governance as a senior-management responsibility embedded in the quality system; ALCOA+ applied across the data lifecycle; verification of the effectiveness of data-integrity controls
EU GMP Chapter 4 (Documentation) and Annex 11 (Computerised Systems), revision drafts
Consultation 7 July–7 October 2025, alongside Annex 22
The EU direction of travel: an expanded Annex 11 addressing, among other topics, audit trails, supplier oversight and identity and access management; final texts may change what audit-trail review and electronic records require. Revisit this section when they are published
16.2 Principle by principle: what it means for an AI-assisted record, and what to capture
A
Attributable
Every element of the record is traceable to who or what produced it. The model is a second author, and its contribution must be distinguishable from the human's.
What to captureAgent or system identity; model name and pinned version; prompt or configuration version; the reviewer's identity; the accept / modify / reject action with the reviewer's rationale; the time of each event
L
Legible
The record, and the AI's contribution to it, can be read and understood for the retention period, including why the AI produced what it did.
What to captureThe output as generated, in a durable format; the confidence or score; the explainability artifact where one exists (contribution plot, feature attribution, cited spans); no reliance on a vendor UI to render it
C
Contemporaneous
AI events are time-stamped when they happen, not reconstructed later; the human review is time-stamped separately from the generation.
What to captureSystem time-stamp of generation; time-stamp of each retrieval; time-stamp of the reviewer's action; the elapsed time between them (also a monitoring signal, §13.1 family 3)
O
Original
The first capture is the AI output as generated. A probabilistic model cannot regenerate it, so a regenerated version is a new record, not a copy.
What to captureThe as-generated output stored before any edit; the AI draft and the human-edited final, with the diff between them; the retrieved sources and the spans used; never overwrite the draft with the final
A
Accurate
The content is correct and complete relative to its sources, and the checks that established that are recorded.
What to captureDeterministic checker findings (numeric consistency, required elements, cross-references) and their resolution; the reference the output was checked against (TFL version, historian tags, source SOP version); confidence and any "undecided" outcome
+C
Complete
Nothing is missing: rejected outputs, overrides, undecided cases and re-runs are part of the record, as invalidated laboratory results are (§8.4).
What to captureAll runs for the case, not only the accepted one; rejected and overridden outputs with reasons; the checker's full finding list, not the summary
+C
Consistent
The sequence of events is coherent and time-ordered across generation, checking, review and signature; the same case is not represented differently in two systems.
What to captureOrdered event chain in one audit trail, or reconciled across systems; consistent case identifiers across the AI service, the eQMS and the signing system
+E
Enduring
The record and the metadata needed to understand it survive for the retention period, including the model version, prompt template and retrieval index snapshot that a vendor may retire.
What to captureRetention of AI context artifacts in a store the company controls (§15, §11.10(c) row); archived alongside the successor system at retirement (§8.8)
+A
Available
The record and its AI metadata can be retrieved and reviewed on request, by the company and by an inspector, in a readable form.
What to captureExport of the AI event chain in a readable format (a vendor clause, §17 element 4); a query that returns, for any signed record, which AI version contributed and what the reviewer changed
16.3 The AI audit trail as GxP metadata
FDA's 2018 Q&A defines metadata as the contextual information required to understand data, and treats audit trails as a form of metadata. On that definition the model version, prompt and configuration version, retrieved sources, confidence and the reviewer's action are metadata of every AI-assisted record. Three consequences follow.
Retention. The AI metadata is retained for the same period as the record it explains. A vendor's "zero data retention" setting on the model API is compatible with this only if your own record store keeps the copy (§15, vendor question 6).
Audit-trail review. The guidance expects audit trails for changes to critical data to be reviewed with the record, before final approval, and to be reviewed periodically across records for patterns. For an AI-assisted record that means two reviews: the reviewer sees the AI event chain for this record before signing (what the model produced, what was changed, why), and QA reviews AI audit trails across records periodically for patterns such as edits that always remove the same kind of claim, or acceptances made in seconds (§13.1, families 3 and 5).
Hybrid records. An AI draft produced in one tool, edited in a second and signed in a third is a hybrid record. MHRA 2018 discourages hybrid arrangements and expects, where they exist, a documented definition of what constitutes the complete record and controls that keep the parts linked. For AI, the complete record includes the as-generated draft, the sources, the checker findings and the signed final, with one case identifier across the tools.
16.4 Data governance for training, validation and test data
Training data, labels and the locked test set are GxP data in their own right (§8.4, T4), and ALCOA+ applies to them: labels are attributable to the adjudicators, original as source-linked, accurate as adjudicated, and complete including exclusions and their reasons. Data governance for AI adds three expectations: criticality assessment of the datasets (a label error on the critical class is a critical data-integrity error), access control and audit trail on the test set so that independence can be demonstrated (Annex 22 §6), and lineage from every data point to its source system (§8.4, UC-M2 example). WHO TRS 1033's expectation that senior management owns data governance applies here as much as to batch records.
16.5 Common data-integrity failure modes with AI
Failure mode
Why it is a DI failure
Detect
Control
The AI draft is overwritten by the edited final
Original lost; the human contribution cannot be separated from the model's
Records with a final but no as-generated version (§13.1 family 5)
Store drafts as immutable events; edit creates a new version
Unattributed AI text in a signed record
Not attributable; the signer may not know what they are signing
Share of AI-originated fields without a source tag
Tag AI-originated content at field level (§15, §11.10(h) row); signature meaning covers AI-assisted content
Retrieval over unapproved sources
Not accurate; content traceable to a source that has no controlled status
Citations that resolve outside the approved-source registry
Grounding restricted to the approved-source registry; index version as a CI (§8.3)
Test-data leakage into training
Validation evidence is not accurate; the claim is overstated
Access logs on the locked test set; overlap checks
Access control and staff independence (Annex 22 §6); T4
Labels harvested from production acceptances
Not accurate: the labels inherit the reviewers' automation bias
Retraining records whose labels lack adjudication
Human adjudication of every label used for training or test (§11, FAQ 12)
Silent vendor model change
Not consistent or enduring: the record's stated model version no longer describes what produced it
Not complete; the cases the model could not handle disappear
Undecided rate reported by the model but absent from the record store
Undecided outcomes stored and routed like any other output (§8.5, Annex 22 §9)
Related on this page
§3 row 9: conversation memory as uncontrolled state
§4.4 probe-set and version checks against silent vendor change
§8.3 index version as a configuration item; §8.4 datasets and labels as validation evidence; §8.5 the "undecided" outcome; §8.8 retirement and archiving
§11 FAQ 12 why production acceptances cannot be training labels
§13.1 families 3 and 5: time-on-review and audit-trail completeness
§15 Part 11 rows §11.10(c), (d), (h); vendor question 6 on retention
§19 the data-integrity sources and what still needs verifying
So what
With Part 11 and ALCOA+ satisfied, the claim is defined, proved, kept true and recorded. Part F asks what all of that looks like for a company with a 3–10 person Quality team, and what the market still fails to supply.
The same claim discipline sized for a 3–10 person Quality team, and what the market still lacks.
Bridge. Parts A to E defined the claim, set its bar, built its controls, kept it true and recorded it, with a CPV agent, a drafting agent and a camera as the running examples and ProtoCheck and Doscierge as the worked systems. Part F scales the same discipline down to a small company that receives AI rather than builds it (§17), turns the page into six moves you can start with, alongside what the market still lacks (§18), and lists the sources and what still needs verifying (§19).
Small biotech: the AI function, not the product category, carries the risk
Audience question Q7"How can a biotech or a smaller company committed to bringing AI tools monitor the risks associated with this switch from Cat 3 to Cat 5, and act on monitoring compliance and data integrity?"
Spine
This section scales the claim down. A small company makes fewer claims, but each one still needs a defined use, a tier, a monitoring minimum (§13.2) and a record (§15–16). The program below is the minimum that keeps every AI function on the site inside that discipline.
First, correct the framing
The risk is not that a company chooses to move from Category 3 to Category 5. For most biotechs, three different things happen at once:
1
Vendor AI arrives inside validated SaaS
The eQMS adds AI deviation triage. The EDC adds AI query suggestions. The document system adds AI summarization. The product stays "Category 3/4", but a function inside it now behaves like trained, probabilistic software. This is the main exposure for a small company, and it usually arrives in a routine release note.
2
Shadow AI
Staff use general-purpose chatbots to draft SOPs, deviation reports, or batch-record summaries. There is no inventory and no data-classification control. This is the GxP spreadsheet problem again. The approach that worked for spreadsheets carries over. Banning them outright rarely succeeded. What worked was an inventory, a risk assessment of the data used and reported, validation effort focused on the high-risk few, and controls validated once as a shared add-on that each use then references. For AI, that shared add-on is the approved enterprise tool with logging.
3
Deliberate builds
A small number of real AI projects, such as a CPV model or a document checker. These are Category 5-like and should be governed as such.
The GAMP categories were built for software whose behavior is set by its code and configuration. For AI you classify the AI function: what decision it influences, whether it is critical, and whether it is static and deterministic (§4, §8.1). A Category 3 product can contain a high-risk AI function.
Minimum viable AI governance program (50–300 employees)
Sized for a company with a 3–10 person Quality team and no dedicated AI group.
Element
What it is
Effort
Owner
1. AI policy + acceptable-use SOP
Approved tools; no GxP data in unapproved AI; AI-assisted authoring rules (§10)
2–3 pages, 1 week
Head of Quality + IT
2. AI inventory
One register of every AI function: vendor, built, or shadow. Fields: system, AI function, GxP area, critical Y/N, static/dynamic, deterministic Y/N, owner, risk tier, last review. Build it on the existing GxP system inventory, which already holds risk rating, business criticality, validation date, and retirement date. Add the AI columns to it rather than starting a separate list. This is the small-company form of the registry in §8.6
A spreadsheet or eQMS form is enough to start
QA (CSV)
3. Risk tiering (three tiers)
Tier 1: non-GxP or authoring aid; register only. Tier 2: GxP, non-critical, human-in-the-loop; lightweight assessment plus monitoring. Tier 3: GxP-critical; full T1–T10 package, and only static/deterministic models (Annex 22; §4.5). The tiers set the monitoring minimums in §13.2
One page
QA
4. Vendor AI clauses
In quality agreements and supplier questionnaires: notify before enabling AI features; say whether customer data trains shared models; supply a model card and test evidence; commit to model-version change notices; allow AI features to be toggled off per tenant; on exit, return AI audit data (inputs, outputs, model versions, human actions) in a readable format (§16, "Available")
Add to the supplier questionnaire template and the quality agreement / SLA
QA supplier mgmt + Procurement
5. Release-note triage
Every SaaS release is screened for AI features before it reaches production. If an AI feature is found: inventory entry, tier, and a decision to enable or disable
30 min per release
System owner
6. Data-integrity controls for AI
ALCOA+ applied to AI: log the input, output, model version, confidence, and human action. Make sure the vendor's audit trail captures that AI contributed (§15, §16)
Configuration plus supplier ask
System owner + QA
7. Monitoring KPIs (Tier 2–3)
Human override/edit rate; error or deviation rate on AI-touched records; drift indicators where available; incident count; the §13.2 minimums for the tier
Automation-bias awareness for reviewers; "how to review AI output" micro-training (§5)
1 hour per person
QA training
Mapping common AI uses to the three tiers
Industry practice ranks AI uses on a ladder, from low risk (drafting, search, summaries) through decision support to GMP-critical decisions. That ladder is a useful start. It becomes defensible once each rung is tied to Annex 22 criticality and to a posture:
AI use
Typical industry rating
Tier
Posture and what it implies
Draft SOPs
Low
1 (or 2 if the SOP is GxP)
Authoring aid; the human-approved document is the record (§10)
Summarize reports
Low–Medium
2
Pattern A: grounding + human review; measure edit rate
Classify deviations
Medium
2
Pattern A while a QA investigator confirms the classification. If the class alone drives impact assessment or batch decisions, move it to Tier 3
Risk assessments
Medium–High
2
AI proposes failure modes and ratings; an SME owns every rating (Annex 22 §4.2 by analogy). Check coverage against a reference failure-mode library (T2 §3)
Batch release / disposition support
High
3
Critical GMP. Only deterministic logic or a static, validated ML model may carry the decision (Annex 22 §1; §4.5). An LLM may assemble the evidence package for the QP or QA reviewer; it may not recommend release
SaaS, cloud and vendor audit: what carries over, and what AI adds
Established vendor and SaaS qualification practice GAMP 5 2nd Ed.; Annex 11 §3; PIC/S PI 011-3 still applies. The AI-specific additions matter because the vendor's model can change without any change you would see in a normal release (§4.4).
Full tableVendor and SaaS practice that carries over from CSV, what AI adds, and the nuance6 rows
Carries over from CSV
What AI adds
Nuance
Vendor documentation may be leveraged, but must be scrutinized, risk-rated and "owned". No vendor can claim a product is "validated" or "Part 11 compliant" slides 162, 164, 615, 626
The same applies to "validated AI" or "GxP-ready AI" claims. Ask for the evidence behind the claim, then re-test for your context of use
Vendor model evaluations are rarely stratified by the subgroups you care about
For cloud and SaaS, IQ becomes a "research effort" drawing on the vendor's published policies and certifications slides 281, 548, 613
Model identity, version, and update cadence are rarely published. Obtain them under the quality agreement, and record the model version in your IQ
A public website is a starting point, not objective evidence of what runs in your tenant
Contract and SLA: advance warning of updates with release notes and time for client testing; notice before service withdrawal so data can be retrieved slides 597–598, 618
Add model-change notices (including silent model swaps behind the same feature name), per-tenant opt-out, and the return of AI audit data on exit
Your release-note triage (element 5) depends on this clause
Post-audit outcomes: use unconditionally, use for certain products only, use subject to corrective action, or prohibit slide 596
Add a fifth outcome: use with the AI feature disabled until it is assessed
Often the fastest decision for a small company
Audit the vendor every two years; require SOC 2 slides 547, 589, 617, 636
Trigger a for-cause assessment when a vendor introduces or materially changes an AI feature
Neither a two-year cycle nor SOC 2 is a GxP regulatory requirement. Set audit frequency by risk (GAMP 5 2nd Ed. supplier management). SOC 2 is an attestation report on security and related trust criteria. It says nothing about GxP fitness or model behavior
Ask where servers are located slide 618
Also ask where inference runs and where prompts and outputs are stored or logged
Relevant to data-residency and confidentiality controls
Signals to monitor (the risk dashboard)
Signal
Why it matters
Trigger
AI functions in inventory not yet tiered
Ungoverned AI in use
Any after 30 days
Vendor releases with AI features enabled by default
Silent category creep
Any not triaged
Human acceptance rate of AI output > 98% sustained
Possible automation bias (reviewers may have stopped reviewing)
About 0.2–0.4 FTE of an experienced CSV lead for the first quarter, then about 0.1 FTE ongoing, plus targeted external help for any Tier 3 build. Small companies can afford this. What they cannot afford is finding an AI function in a validated system during an inspection.
So what
The same claim discipline fits a small company, because most of its AI arrives as a vendor feature that needs a tier, a switch and a record rather than a model validation. §18 turns the whole page into six moves you can start with.
Apply it: six moves from this page to your first defensible AI claim
Apply it · from reading to doing
Spine
The page has defined the claim, set its bar, built its controls, kept it true, recorded it and scaled it down. This section turns that into six moves for one AI use. The moves follow the lifecycle in §1.3, and each one names what you produce, where it was shown on ProtoCheck or Doscierge, and the template or section that carries it.
18.1 Six moves, one use
Move
What you do
What you produce
Shown on
Template / section
1. Pick one use and write its context of use
One sentence: the input, the output, who decides, and what happens if the output is wrong
Take the six inspector questions in §8 (What an inspector will ask) and answer each one with a document you can hand over, not a description. If "how did you determine acceptable performance?" is answered by a baseline study and a signed claim record, and "what happens when the AI is wrong?" is answered by a closed monitoring record linked to a change, the claim is defensible. If either answer is a slide, go back to move 3 or move 5.
Start small, on purpose
Pick a Tier 2 use (a human-reviewed drafting or checking agent) for the first pass. It exercises all six moves with a one-page assurance plan (§9) and teaches the team the claim discipline before it is needed for a Tier 3 decision.
18.2 What the demand signal says, and what to build
The questions in §2 are also demand data for the claim discipline this page describes. Read as a market:
The demand signal. One general web event produced eight questions, and seven of them asked for operational how-to. Three more arrived after the session or from practitioners, and all three asked for a policy or a decision table rather than a principle. The supply side is split:
Standards bodies
ISPE and PDA publish frameworks.
Big consultancies
Sell enterprise programs.
eQMS vendors
Ship AI features without governance tooling.
Nobody
Sells a right-sized, template-driven AI validation practice for small and mid-size companies.
The deadline pressure is real. Annex 22 and the Annex 11 revision are targeted for finalization around Q4 2026, and the FDA–EMA Good AI Practice principles came out in January 2026.
Why buyers care: the enforcement context
Inspectors already cite weak audit trails, unvalidated workflows, missing risk assessments and unauthorized design changes, and Quality Unit oversight of data systems is a continuing focus in public FDA warning letters. An AI function added to a validated system without inventory, risk assessment, or an audit trail of its contribution fits those citation patterns. Smaller firms often stay reactive and under-resourced, which is the gap the §17 program is sized for.
Offering
Answers
Buyer
Format
Notes
AI Validation Template Pack (T1–T10, plus the §12 decision tree, the §17 program and the §13 policy)
Extends the parent article; ProtoCheck and Doscierge as proof
Site AI monitoring policy and data-integrity addendum
Q13, Q12, Q9
Heads of Quality, system owners
Policy text (T8 Appendix B) + implementation workshop
New in this version; the piece that makes the eight controls operable across a site
Positioning line
"CSV tells you how to validate software. We show you how to validate a performance claim, and how to keep it valid."
What comes next
The template pack (T1–T10) is complete; T1, T2 and T8 are published free with this page. Next: an SME review pass of the templates and a one-page Tier 2 assurance-plan variant.
Refine the §17 program sizing, the tier mapping, the §12 decision tree and the §13 minimums with practitioner feedback.
An "Applying CSV to AI" practitioner session, using this page as the outline, is being proposed.
The parent article's §11 now links here; this page links back.
So what
Six moves make one claim defensible, and the market needs map onto the same claim components: baseline, controls, monitoring, records, scale-down. The last section lists what this page rests on and what still needs verifying.
A claim is only as strong as its sources. This section separates what was verified for this version, where the practitioner questions came from, and what remains to verify before relying on it.
Primary sources verified for this version
Full listRegulations, guidance and studies checked for this version11 sources
EU GMP Annex 22 Artificial Intelligence, consultation draft, 7 July 2025 (full text reviewed; section numbers above cite it directly): European Commission PDF
Annex 22 status: consultation closed 7 Oct 2025 with about 1,300 comments; EMA IWG targets a final text in Q4 2026; EMA workshop held 30 Jun–1 Jul 2026, considering whether dynamic and probabilistic models could be addressed. Epista, Scilife
EU GMP Chapter 4 and Annex 11 revision drafts, released for targeted stakeholder consultation together with Annex 22 on 7 July 2025; consultation 7 July–7 October 2025. PIC/S notice, ECA summary
FDA draft guidance, Considerations for the Use of AI to Support Regulatory Decision-Making for Drug and Biological Products (7 Jan 2025): seven-step, risk-based credibility framework. Federal Register, FDA PDF
FDA final guidance, Computer Software Assurance for Production and Quality System Software (24 Sep 2025).Note: issued by CDRH/CBER for device production and quality-system software. Pharma applies its risk-based principles by analogy; it is not binding on drug GMP systems. Federal Register
FDA–EMA, Guiding Principles of Good AI Practice in Drug Development (14 Jan 2026): ten high-level principles, including a human-centric, risk-based approach, context of use, data governance, and lifecycle monitoring. EMA PDF, EMA news
FDA Quality Management System Regulation (QMSR), effective 2 Feb 2026. It incorporates ISO 13485:2016 by reference, which replaces the old 21 CFR 820.70(i) software-validation text (ISO 13485 clause 4.1.6). FDA QMSR
21 CFR Part 11 (§11.10(a)–(k), §11.50, §11.70, §11.100–§11.300) and FDA guidance Part 11, Electronic Records; Electronic Signatures — Scope and Application (Aug 2003), which gives the predicate-rule and reliance test used in §15. eCFR Part 11, FDA guidance
Garza MY et al., Error rates of data processing methods in clinical research: a systematic review and meta-analysis, Int J Med Inform 2025;195:105749. PubMed, ScienceDirect
USP <1790> and the Knapp–Kushner method: POD ≥ 0.7 defines the reject zone; alternatives to manual inspection must show equivalent or better performance. PDA, ISPE Pharmaceutical Engineering
Full listData-integrity sources (§16) and FDA device sources used by analogy (§4.3, §12)6 sources
Data-integrity sources (§16), titles and dates verified for this version
PIC/S, Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments, PI 041-1, adopted 1 June 2021, in force 1 July 2021. PIC/S document, PIC/S news
WHO, Guideline on data integrity, Annex 4 of WHO Technical Report Series 1033 (fifty-fifth report of the Expert Committee on Specifications for Pharmaceutical Preparations), 2021. WHO PDF
FDA device sources used by analogy only (§4.3, §12)
FDA, Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD): Discussion Paper and Request for Feedback, April 2019: introduced the "locked" vs. "adaptive" distinction and the Predetermined Change Control Plan concept. FDA AI/ML SaMD page, Discussion paper PDF
FDA, Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions, final guidance, 4 December 2024.FDA guidance page, Federal Register, 4 Dec 2024
Origin of the practitioner questions
Web event, Acceptance Criteria for AI Validation (23 Sep 2026). The source of the session questions Q1–Q9 in §2.
ProtoCheck and Doscierge figures (rule counts, knowledge-graph size, P/R/F1, findings per package before and after precision engineering) are taken from the parent article's §17 and are the author's own systems.
Cited from domain knowledge, to verify before relying on it
Full listItems cited from domain knowledge, to confirm before quoting13 items
Verify before relying on these
Pooled per-method error rates in Garza et al. (only the ranges above were verified).
MASAI trial specifics (Lång et al., Lancet Oncology 2023): confirm the detection and workload figures before quoting numbers.
THERP / NUREG/CR-1278 nominal HEP values: cited only qualitatively here.
GAMP 5 2nd Edition Appendix D11 (AI/ML): confirm the appendix reference.
ServiceNow terminology (§14): the CSDM hierarchy and the Standard/Normal/Emergency change types are standard ServiceNow constructs. The AI CI classes are illustrative extensions, not out-of-the-box classes. Confirm the current names of ServiceNow's AI-governance and DevOps change-automation features before publication.
SOC 2 characterization (§17): an AICPA attestation report on the Trust Services Criteria, not a certification. Confirm the wording against the AICPA source before publication.
FDA 2018 Data Integrity Q&A (§16.3): the definition of metadata and the expectation that audit trails capturing changes to critical data are reviewed with each record before final approval are cited from the guidance as recalled; confirm the exact Q&A numbers and wording before quoting.
MHRA 2018 on hybrid systems (§16.3): the statement that hybrid arrangements are discouraged and, where used, need a documented definition of the complete record is cited as recalled; confirm the clause and wording.
Draft Annex 11 (July 2025) content summary (§16.1): the topics listed (audit trails, supplier oversight, identity and access management) are taken from consultation summaries, not from a review of the draft text. Verify against the draft before citing it in detail.
WHO TRS 1033 Annex 4 (§16.1, §16.4): the senior-management data-governance expectation and the appendix on verifying control effectiveness are cited from secondary summaries; confirm against the annex text.
PIC/S PI 041-1 (§16.1): described at the level of its scope and structure only; no section numbers are cited.
Draft Annex 22 §1 scope wording on non-critical applications (§4.2): the draft's principles are described here as applicable "where applicable" to non-critical uses; confirm the exact wording of the scope clause before quoting.
The "so what" statements and Strategic Planning Assumptions are the author's judgements from the sources cited; the probabilities are not survey results.
Confidence notes
Confidence notes
Every number labelled illustrative (§6 test sizes, §8 thresholds, §13 limits and SLAs, §17 FTE estimates) is a design example, not a benchmark.
Annex 22 is a draft. Its scope on dynamic and probabilistic models may widen in the final text, so revisit §4, §9 and §8.1 when it is published.
The Chapter 4 and Annex 11 revisions are drafts. Revisit §16 when the final texts are published.
Device guidance (CSA, the 2019 AI/ML discussion paper, the 2024 PCCP guidance) is cited by analogy only and is not binding on drug GMP.
Where the spine ends
Claim defined (A), bar set (B), controls built (C), kept true (D), recorded (E), scaled down (F). The claim is the whole argument: you validate a performance claim for one context of use, then you keep proving it.
The controls are defined. Three templates are free. The baseline is yours to measure.
This page exists to help Quality, CSV, manufacturing, clinical and platform teams turn CSV principles into AI-specific controls. It is a companion to the Enterprise Agentic AI Platform for Life Sciences architecture, whose governance plane, agent registry, promotion gate, dual-path reasoning and validation strategy it applies to the questions practitioners actually asked.
Three templates, published free
The intake, the risk assessment and the monitoring plan: enough to classify an AI function, rate its risk and set its KPIs.
T3 context of use, T4 data management, T5 design, T6 validation plan, T7 summary report, T9 predetermined changes, T10 periodic review. Message on LinkedIn or scan the contact code.
30 minutes to map one of your AI use-cases to the eight controls
Bring one use-case: a vendor AI feature, a vision system, a drafting agent. We walk it through GxP assessment, risk, version control, data, design, registry, monitoring and periodic review, and you leave with the gaps named.
The honest caveat: Annex 22 is still a draft, the FDA CSA guidance is device guidance applied by analogy, and the worked-example values on this page are illustrative. The method holds; the numbers must be yours.
Start with one use-case.
Not a sales call: a structured walkthrough of one AI use-case against the eight lifecycle controls, and where your first template would apply.