Part 2 · Continues §11 Validation Strategy · Life Sciences

Validation Strategy, AppliedAgentic AI in GxP: from the CSV/CSA lifecycle, through two production systems, to thirteen practitioner questions

OrchestraPrime · Nitin Bhatti, Founder & Platform ArchitectVersion 0.6 · 2026-09-23~135 min read straight through · role paths 25–40 minContinues §11 Validation Strategy of Enterprise Agentic AI Platform for Life Sciences — Solution Architecture (v2.1: §5.5 Governance Plane, §6 Agent Registry, §7.3 Promotion Gate, §10 Dual-Path Reasoning, §11 Validation Strategy)

What this page continues. The parent article's §11 Validation Strategy settled which control applies to which component of an agentic platform: deterministic rules validated as GAMP 5 Category 5; the generative path held by grounding, deterministic checking and human review; statistical models under an analytical-method lifecycle; and a Part 11 signature on every regulated output. It closed on a warning. Teams that document this posture before the first agent enters a GxP workflow will set the standard, and teams that retrofit it will spend twice as long. This page covers what §11 left open: how you prove it. It starts from the validation lifecycle every CSV team already runs (FDA's software-validation principles, CSA, GAMP 5, Part 11 and Annex 11), shows where AI breaks that lifecycle's assumptions, and runs it on two production systems from the parent article, ProtoCheck and Doscierge. It then tests the result against thirteen questions that practitioners actually asked.

How the argument is built: four layers, each resting on the one before, and a close that turns them into action

The Trust Paradox

Why does retrofitting cost twice as much? Because of a tension every Quality leader recognises:
The promise

AI promises speed

Drafts in minutes instead of days. Signals found before the annual report. Deviations triaged without a queue. Every vial inspected at line speed.

The constraint

GxP leaves no room for uncertainty

Every regulated decision needs a specification, a test, an audit trail and an accountable signer. A model that learns from data has no specification, performs statistically, and can decay without anyone touching it.

The resolution, and the spine of this page
You don't validate trust into a model. You validate a performance claim for one context of use, then you keep proving it.
Validate the controls around the model for a specific context of use, not "the model" in the abstract. Model metrics are evidence that the controls work, not a certificate for the model in general.

Every section that follows establishes, or keeps proving, one part of that claim:

Bottom Line

Practitioners know CSV. What they lack is AI-specific controls, operational templates and worked examples.

The industry does not lack CSV knowledge. It lacks a way to turn CSV principles into AI-specific controls: how to set acceptance criteria against a measured human baseline, how to tell drift from a change in intended use, what to do when a SaaS vendor quietly adds AI to a validated Category 3/4 system, and which documents a small company actually needs. The regulations have now moved far enough to answer most of this. Draft EU GMP Annex 22 (July 2025) sets clear rules for static, deterministic ML in critical GMP use. The ISPE GAMP Guide: Artificial Intelligence (July 2025) sets out the lifecycle. FDA's January 2025 draft guidance defines a seven-step credibility framework built around context of use. What the industry is missing is operational templates and worked examples. That gap is the market opportunity.

§4.3
Draft Annex 22: the baseline rule
Acceptance criteria at least as high as the process replaced, so that performance must be known. Most sites have never measured it (§5).
7 / 8
session questions asked "show me how"
Not "tell me what". Practitioners already accept that AI needs validating. They are stuck on execution (§2).
8
lifecycle controls, one AI extension each
GxP assessment, risk, version control, data, design, registry, monitoring, periodic review, mapped to templates T1–T10 (§8).
5
metric families in one site policy
Inputs, outputs, human-in-the-loop, system/vendor, data integrity, with minimums by tier and a four-level response ladder (§13).
3
templates published free
T1 intake, T2 risk, T8 monitoring. The rest of the ten-template pack is available on request.
Key Findings

Six findings, and each one is a component of the claim

1

The unit of validation is a performance claim for one context of use, not the model.

The same model is low risk as a drafting aid and critical as a release gate; only the use, the population of inputs and the surrounding controls make it validatable.

Supportdraft Annex 22 §3.1 (intended use and input sample space); FDA January 2025 draft, Steps 1–3; §1 and §8 below.
2

The human baseline is the missing input, and draft Annex 22 makes it a regulatory expectation.

Its §4.3 requires model acceptance criteria "at least as high as the performance of the process it replaces", which means that performance must be known. Outside visual inspection (USP <1790>, Knapp–Kushner), most sites have never measured it.

Support§5 (Q6), which gives published anchors and a measurement protocol.
3

A locked model does not drift; the world it sees does. So monitoring watches inputs and humans, not only accuracy.

Ground-truth labels are rarely available in production, so the observable signals are input-distribution scores, the "undecided" rate, override and edit rates in both directions, and periodic reference checks that supply ground truth. In CPV, the first question is whether the process moved or the model's inputs did.

SupportAnnex 22 §10.3–10.4; §11 (Q11 drift FAQ), §12 (Q2), §13 (Q13).
4

For most companies AI arrives inside validated SaaS, so category creep is a supplier-management and change-control problem before it is a model-validation problem.

GAMP categories describe the product; the AI function inside it must be classified on its own: critical or not, static or dynamic, deterministic or probabilistic.

Support§17 (Q7) program and vendor clauses; §4 (Q11a) decision table.
5

Part 11 and ALCOA+ do not change for AI. What each control has to capture does.

The model is a new actor, the prompt and retrieved context are new inputs, and a probabilistic output cannot be regenerated on demand. The record must therefore hold the output as generated, its model and configuration version, its sources, and the human's accept/modify/reject as separately attributed events. No product "meets Part 11".

Support§15 (Q9) clause by clause; §16 (Q12) principle by principle.
6

CSV is not obsolete. Each of its eight lifecycle controls still applies, and each needs one AI extension.

GxP assessment, risk, version control, data, design, registry, monitoring and periodic review map one-to-one onto an AI lifecycle SOP set (T1–T10). Draft Annex 22 also excludes dynamic and generative/probabilistic models from critical GMP use, so the practical design for critical decisions is dual-path: deterministic where the decision is made, generative where a qualified human reviews.

Support§8 (Q8) and the at-a-glance matrix in §8.0; parent article §10–11.
Strategic Planning Assumptions

Five planning assumptions, with probability and substantiation

By end-2027Probability 0.70

By end-2027, EU GMP Annex 22 will be in force in a form that keeps critical GMP use limited to static, deterministic models and requires a measured baseline for the process replaced.

Substantiation
The consultation draft (7 July 2025) scopes critical use to static, deterministic models and sets the "no decrease" rule in its §4.3; the consultation closed 7 October 2025 with about 1,300 comments; the EMA inspectors working group targets a final text in Q4 2026; an EMA workshop on 30 June–1 July 2026 discussed whether dynamic and probabilistic models could be addressed. The scope may widen for non-critical use; the baseline expectation is unlikely to be dropped, because it follows from the existing Annex 11 and validation principle of demonstrating fitness for intended use.
Through 2028Probability 0.75

Through 2028, the majority of GxP AI exposure at small and mid-size companies will arrive as vendor-delivered features inside Category 3/4 SaaS, not as in-house models.

Substantiation
The platform vendors that hold GxP records are shipping AI features in routine releases (the parent article records Veeva shipping AI agents for PromoMats in December 2025 and Agentforce Life Sciences reaching general availability in October 2025); the questions from the session (Q7, Q9) came from exactly this exposure; a company with a 3–10 person Quality team has no capacity to build and validate its own models, but receives release notes every month.
By 2027Probability 0.65

By 2027, FDA's January 2025 draft on AI in regulatory decision-making will be finalised with the seven-step, context-of-use credibility framework substantially intact, and it will become the default template for clinical AI evidence.

Substantiation
The draft's framework (question of interest, context of use, model risk = influence × consequence, credibility plan, execution, results, adequacy) is reinforced by the FDA–EMA Guiding Principles of Good AI Practice in Drug Development (14 January 2026), which foreground context of use, data governance and lifecycle monitoring; the ISPE GAMP AI Guide (July 2025) follows the same lifecycle logic.
By 2028Probability 0.70

By 2028, dual-path design (deterministic component carries the critical decision; generative component drafts and explains under human review) will be the default validation posture for GxP-critical AI, ahead of any regulator explicitly endorsing it.

Substantiation
Draft Annex 22 excludes generative and dynamic models from critical GMP use, and the EMA is only "considering" whether to address them; FDA's CSA guidance (final 24 September 2025, device scope, applied to pharma by analogy) supports risk-proportionate assurance of the deterministic components; the parent article's §11 already applies this posture across research, development, manufacturing and medical use-cases.
Through 2027Probability 0.60

Through 2027, the most common finding in Annex 22 readiness reviews will be an unmeasured human baseline, not a weak model.

Substantiation
Annex 22 §4.3 makes the baseline a precondition for setting acceptance criteria; the only mature pharma precedent for measuring it is manual visual inspection (USP <1790>, Knapp–Kushner); the clinical data-processing literature (Garza et al., 2025) shows how wide human error rates are by method, which is why a local measurement, not a published number, is needed. This is a judgement from the questions asked in the session (Q6, Q2), not a survey result.

On the probabilities. The Strategic Planning Assumptions are the author's judgements from the sources cited; the probabilities are not survey results (§19).

Recommendations

By persona: the next 90 days, and what not to accept

Next 90 days

For the Head of Quality / QA

  • Issue an AI policy and inventory (§17, elements 1–2); tier every AI function, including vendor features; make the human baseline (§5) a required input to any Tier 3 validation plan; adopt one site monitoring policy (§13) so every AI function reports the same five metric families.
Do not

Accept "the vendor validated it" or "a human reviews everything" as controls until the vendor evidence covers your context of use and the reviewer's catch rate has been measured (§7 "Human review", §5).

Next 90 days

For the CSV / validation lead

  • Rewrite one AI requirement as a performance claim (metric, subgroup, threshold, baseline, confidence level; §1 last row); build the locked, independent test set before any model is touched (§8.4, T4); make the eight-control matrix (§8.0) the table of contents of your AI validation package; treat retraining, prompt edits and vendor model swaps as changes with a pre-approved envelope (§14, T9).
Do not

Run a single PQ and call the model "validated"; the validated state for AI is a claim that monitoring keeps true (§11, §13).

Next 90 days

For Manufacturing / MS&T

  • For any CPV, PAT or vision AI, decide which component carries the critical decision and keep it static and deterministic (§4 decision table, §8.1); measure the manual process it replaces (§5); write the predictable changes (new supplier within spec, recalibration, NOC refresh) into a T9 envelope before go-live (§12, §14.2).
Do not

Retrain a CPV model to make an excursion disappear. A real process signal is what CPV exists to catch; route it to deviation and change control, not to model retraining (§11, FAQ 4).

Next 90 days

For Clinical / Regulatory

  • Validate the CSR or document-drafting agent under the seven-step credibility framework with a demonstrated (not assumed) human-review effect (§6); store every AI output as generated with model, prompt, sources and reviewer action (§15, §16); set the signature meaning to cover AI-assisted content (§15, §11.50/§11.70 row).
Do not

Let an LLM grade its own output, or let conversation memory carry state between cases (§3 rows 9–10, §9).

Next 90 days

For Small-biotech QA (50–300 employees)

  • Run the 90-day starter plan (§17): policy, inventory sweep, tiering, vendor AI questionnaire to the top five GxP SaaS suppliers, untriaged AI features switched off; adopt the Tier 2 one-page assurance plan for human-reviewed drafting uses (§9); register AI-assisted authoring once, not per document (§10).
Do not

Build a separate AI governance system. Extend the GxP system inventory, the supplier questionnaire, the quality agreement and the periodic review you already have (§17, elements 2, 4, 8).

Next 90 days

For the IT / AI platform owner

  • Model each AI component (model, prompt/config, retrieval index, locked test set, tool server, monitoring config) as a configuration item under the business application (§14.1); give each agent its own identity and least-privilege tool allowlist (§15 §11.10(d) row); pin vendor model versions and manage the deprecation schedule as a planned change (§4, §14.5 S2); route monitoring action-limit breaches into Incident and Problem records (§14.4, §13).
Do not

Hide the model, prompt and index inside one application CI. Every change then looks either trivial or total, and change-level risk-based validation becomes impossible (§14.1).

How to read this page

Pick your lens. The contents highlight your path.

This is a long page because the questions came from five different jobs. Read straight through for the full argument, or take the role path. Times assume about 250 words a minute. Select a role and the table of contents highlights the sections to read first and then, and estimates the reading time. Your choice is remembered on this device.

Detail level
for the Quality / CSV lens: the reframe, the misconceptions that change a package, the pitfall map, the eight controls, Part 11 clause by clause. The author's estimate: 40 min first pass; 30 min second.
for the Manufacturing / MS&T lens: locked vs. adaptive, the human baseline, the CPV signal agent as the running example, drift in CPV terms, the AI camera. The author's estimate: 40 min; 25 min.
for the Clinical / Regulatory lens: CSR drafting under the FDA credibility framework, clinical data-processing error rates, Part 11 and ALCOA+ with an AI contributor. The author's estimate: 35 min; 25 min.
for the Small-biotech QA lens: a 90-day governance program, a proportionate rule for AI-assisted authoring, vendor questions, monitoring minimums by tier. The author's estimate: 30 min; 25 min.
for the IT / AI platform lens: CMDB and change models, agent identities, configuration items, monitoring routed into ITSM. The author's estimate: 35 min; 25 min.

Whatever lens you bring, the argument has one spine: in CSV you validate a specification. In AI you validate a claim about performance, on a defined population of inputs, against a known baseline, and then you keep proving that claim holds.

Contents

Six parts, nineteen sections, one claim

Question index

Every audience and practitioner question keeps its Q number as a tag under the section heading, so external references (templates, posts, the parent article) still resolve.

Full tableQuestion index: Q1 to Q13 with origin and the section that answers each13 rows
QQuestion (short form)OriginSection
Q1More examples: CSV concepts and their AI pitfallsSession§7
Q2In-line AI camera: when does drift occur, does a new label force revalidation?Session§12
Q3Validation templates for AI agentsSession§9
Q4A worked validation of an AI tool, following industry practiceSession§6
Q5Do I need a risk assessment if AI creates the templates?Session§10
Q6Human error-rate studies as a basis for AI acceptanceSession§5
Q7Small biotech: monitoring the Category 3 → Category 5 shiftSession§17
Q8Applying CSV to AI: the eight lifecycle controlsSession§8 (Q8.1–Q8.8 = §8.1§8.8)
Q9Is there an AI product that meets Part 11, and how does the guidance differ?After the session§15
Q10Change control end to end: CMDB, IQ/OQ/PQ, release, monitoring, AI co-brainAuthor§14 (Q10.1–Q10.8 = §14.1§14.8)
Q11Locked vs. adaptive models (a), and the drift FAQ (b)Practitioner (not from the session)§4 (a), §11 (b)
Q12Data integrity for AI: what ALCOA+ requires of an AI-assisted recordPractitioner (not from the session)§16
Q13A monitoring policy for AI across the sitePractitioner (not from the session)§13
SpineABCDEF

Part A · Define the claim

Claim component Claim definition

What is being claimed, for which use, by which kind of model.

Bridge. The Trust Paradox is resolved by changing the unit of validation from "the model" to "a performance claim for one context of use". Part A builds that claim in the order a validation team would: the CSV/CSA lifecycle you already run and the assumptions AI breaks (§1.1), the parent article's rule for which control applies to which component (§1.2), and both of them run on two production systems (§1.3). It then turns to what practitioners asked, with each question pinned to a stage of that lifecycle (§2), the common views that get the claim wrong (§3), and the kinds of model the claim can cover in critical and non-critical GMP use (§4). Parts B to F set the bar for the claim, build its controls, keep it true in operation, record it, and turn it into action.

§1 · Part A · Layers 1–3 · The standard, the architecture, show me

In CSV you validate a specification. In AI you validate a claim.

The reframe · claim definition

Spine

Layer 1 · The standardThis section starts from the lifecycle every CSV team already runs and defines the claim. Classic CSV (GAMP 5, Annex 11, Part 11) rests on assumptions that deterministic software meets and AI does not. Every audience question traces back to one of the rows below, and the right-hand column is the control that restores each assumption.

CSV assumptionHolds for rules-based softwareWhat changes for AIControl that restores it
Behavior is specified: URS → FS → DS, and code implements the specYesBehavior is learned from data. The "spec" is partly the training setIntended-use and input sample space definition (Annex 22 §3.1); data becomes a configuration item (§8.3, §8.4)
Same input → same outputYesML: yes once frozen. LLMs: not guaranteedAnnex 22 scope is static + deterministic only; LLMs go non-critical with human-in-the-loop, or dual-path (§4, §9)
Testing proves correctnessPass/fail per requirementPerformance is statistical: sensitivity, specificity, F1, with confidence intervals per subgroupPre-approved metrics and acceptance criteria, owned by an SME (Annex 22 §4.1–4.2); test set sized for statistical confidence (Annex 22 §5.2); the baseline in §5
Validated state lasts until the system changesYesPerformance can decay without anyone touching the system: the world changes (drift)Performance and input-drift monitoring (Annex 22 §10.3–10.4); §11, §13
Change = code or config changeYesChange also includes new training data, new labels, a new supplier's container, new lighting, a vendor model swapChange control covers model, system, process, and the physical inputs (Annex 22 §10.1); §12, §14
Test data is just test dataMostlyTest data leaking into training inflates performance and invalidates the resultTest-data independence, access control, staff independence (Annex 22 §6); §8.4, §16
GAMP categories describe the systemCat 3/4/5A Cat 3 SaaS product can ship an AI feature overnight. The category no longer tells you the riskClassify the AI function, not only the product (§17, §8.1)
The human reviewer is the controlYesAutomation bias: reviewers approve AI output they would have challenged coming from a colleagueMonitor human-in-the-loop performance "like any other manual process" (Annex 22 §3.3); measure acceptance and edit rates (parent G2); §5, §13
Each requirement is granular, testable, and risk-rated on its own GAMP 5 2nd Ed.; FDA CSA, 2025Yes: one requirement → one or more test scripts → pass/failA single AI requirement is a performance claim over a population. It cannot be broken down further into pass/fail stepsExtend the requirement-attribute record (ID, rationale, risk rating, source, dependencies) with metric, subgroup, threshold, baseline, and confidence level. That record is the row T6 traces to
The one-sentence reframe for a Quality audience

In CSV you validate a specification. In AI you validate a claim about performance, on a defined population of inputs, against a known baseline, and then you keep proving that claim holds.

1.1 From CSV to CSA to AI assurance

AI assurance is not a third methodology. It is CSA's risk-based critical thinking, applied to a component whose behavior is learned rather than specified. The same lifecycle (concept → project → operation → retirement, as in GAMP 5 2nd Ed. and the ISPE GAMP AI Guide) carries through.

AspectClassic CSVCSA (risk-based)AI assurance (adds)
Question asked"Does the system do what it purports to do?""Where could failure hurt patient, product, or data, and how much assurance does that need?""How reliably does it perform, for this context of use, against a known baseline, and does that still hold?"
What drives effortDocument set, applied uniformlyRisk of each feature FDA CSA, 2025Model influence × decision consequence (FDA Jan 2025 draft, Step 3) and Annex 22 criticality (§8.2)
TestingScripted IQ/OQ/PQScripted for high risk; unscripted or exploratory for low riskStatistical testing on an independent, stratified test set; N-run testing for probabilistic output; trajectory and red-team tests for agents (§9)
Acceptance criteriaExpected result per stepSame, scaled by riskPre-approved metric per subgroup with confidence interval, no worse than the measured baseline (Annex 22 §4.2–4.3; §5)
Vendor evidenceOften retestedLeveraged where risk allows, but owned by the regulated companyAlso: model card, training-data provenance, model-change notification, whether your data trains shared models (§17)
After releaseChange control; annual periodic reviewRisk-scaled regressionContinuous performance, input-drift, and human-in-the-loop monitoring (Annex 22 §10.3–10.4; §8.7, §13)
What is validatedThe systemThe system, in proportion to riskThe controls around the model, for a defined use: grounding, deterministic checker, confidence gating, human review, monitoring, change control. Model metrics are evidence that those controls work, not a certificate for the model in general

1.2 Where this fits: validation strategy for the whole platform

Layer 2 · The architectureThis page continues the parent article's §11 Validation Strategy, which sets the posture for every use-case on the enterprise platform. That section starts from the same premise as this one: "you can't validate an LLM under GAMP 5" is the wrong framing, because what you validate is the controls around the LLM. It then answers the auditor's question, "which control applies to which component?", with four components and four strategies.

Component 1 · Deterministic

The deterministic path

Rule engines, checkers, SPC and capability queries, reconciliation logic: validated conventionally as GAMP 5 Category 5, each rule version-controlled with a named owner.

Component 2 · Generative

The generative / LLM path

Drafting, synthesis, hypothesis generation: controlled by grounding on an approved-source registry, by deterministic checking of every output, and by human review, under CSA-aligned controls. CSA is FDA device guidance that pharma applies by analogy, and draft Annex 22 excludes dynamic and probabilistic/generative models from critical GMP use, which is the reason the critical decision stays on the deterministic path.

Component 3 · Statistical

Statistical and chemometric models

MSPC, PLS, ML scorers: they follow an analytical-method lifecycle rather than software validation. NOC set, cross-validation, held-out test, QA approval in the model registry, performance monitoring against the reference method.

Component 4 · Human

Human sign-off

On every output that enters a regulated record, governed by Part 11 and Annex 11, with the system designed so the human has the evidence, findings and confidence needed to decide.

The depth of assurance follows the risk of the decision, and the governance manifest on each agent card records which posture applies, so the validation package is a rendering of the card rather than a document written from scratch.

Everything below is that strategy worked through: §1.3 runs it on two production systems, stage by stage; §4 says which model types each path may use; §8.1 shows one CPV agent split across all three technical components; §9 gives the agent patterns; §15 and §16 cover the human sign-off component.

1.3 Show me: the standard lifecycle, run on ProtoCheck and Doscierge

Spine

Layer 3 · Show me§1.1 is the lifecycle every CSV team already runs, and §1.2 is the parent article's rule for which control applies to which component. This subsection puts the two together on the two production systems described in the parent article's §17: ProtoCheck (clinical protocol compliance) and Doscierge (regulatory documentation compliance). Later sections refer back to this table, so the same two systems can be followed through every stage.

Honest scope

Both systems are built and evaluated, and Doscierge's checker is designed for GAMP 5 Category 5 validation. Neither is presented here as validated in a customer's GxP environment. The tables show how the package would be built. Thresholds marked illustrative are examples, not measured results; the measured figures are the parent article's.

The two claims, written the way §1 says a claim must be written

ProtoCheckDoscierge
Context of useA draft clinical protocol is checked against 500+ expert-calibrated rules by four specialist reviewers (safety, efficacy, regulatory, operational). A qualified clinical reviewer accepts, modifies or rejects every finding before the protocol is finalisedAn ANDA documentation package is checked against 970 deterministic rules and a curated knowledge graph (10,292 triplets, 29 domains). A regulatory reviewer dispositions every cited finding before submission
What the AI decidesNothing on its own. It proposes findings; the protocol sign-off is humanNothing on its own. The deterministic checker produces the findings, the LLM path extracts and explains, and the submission decision is human
Parent-article use-caseUC-D1 Trial design · the rule-gate patternUC-D4 Regulatory documentation · the draft-then-check pattern
Tier (§8.2)Tier 2: decision support, human decides. Consequence of a miss: an amendment, a delay, or a safety design issue caught lateTier 2, upper end: human decides, but a missed deficiency can cost a refuse-to-receive or a deficiency cycle
Performance claim (illustrative)On a locked set of seeded-defect protocols across [n] therapeutic areas, recall on safety-rule findings ≥ [x]%, with the 95% CI lower bound no worse than a qualified reviewer working unaided on the same setOn the expert-labelled evaluation set, per-domain recall ≥ [x]%, with the CI lower bound above the unaided-reviewer baseline, findings per package within reviewer capacity, and no missed seeded critical errors

The lifecycle, stage by stage. The left column is the CSV/CSA stage you already run; the second column is the parent §11 component it lands on.

Full tableThe standard validation lifecycle, stage by stage, for ProtoCheck and Doscierge10 rows
Standard stage§11 componentProtoCheckDosciergeDeeper in
1. GxP assessment and intended useAllGCP. The AI function is advisory review of a controlled document; the sign-off stays humanRegulatory submission content. The checker is Category 5 software; the LLM path is non-critical and works under human review§8.1 · T1
2. Risk assessmentAllWeight false negatives on safety rules highest. Failure modes: a missed finding, a rule applied out of scope, reviewer over-trustWeight false negatives on refuse-to-receive-relevant rules highest. False positives matter too, because noise tires the reviewer§8.2 · T2
3. Requirements become claimsDeterministic + generativeEach rule carries an ID, a CFR/ICH source, an owner and a calibration status. The agents' requirement is the performance claim aboveEach rule carries an ID, a citation and an owner. The claim is stated per domain, not as one headline F1§1 (last row) · §5
4. Configuration managementAllRule versions; the four agents' prompts and tool allowlists (≤5 tools each); the pinned model version970 rules; the knowledge-graph snapshot; extraction prompts; the pinned model version; the locked evaluation set§8.3 · §14.1
5. Test design (OQ/PQ)Deterministic: scripted. Generative: statisticalSeeded-defect protocols per therapeutic area; N-run consistency for each agent; the coordinator's merge tested for provenanceAn evaluation harness against expert-labelled ground truth (P 92.31 / R 80.27 / F1 85.87); every rule change passes regression before merge§6 · §8.4 · §9
6. Acceptance and releaseHuman sign-offOnly rules in confirmed status (named SME signature) are in the validated scope; provisional rules run in shadowRelease when the claim holds per domain and the reviewer effect is demonstrated, through the platform's promotion gate (parent §7.3)§8.5 · §8.8
7. MonitoringGenerative + humanFinding acceptance rate per rule; overrides in both directions; the share of findings reviewers editFindings per package (the volume the reviewer must work through); accept/reject per rule; input drift on new document types§11 · §13 · T8
8. Change controlAllA provisional → confirmed promotion needs a named SME signature; a model-version change re-runs the seeded-defect setA rule edit that passes regression is a Standard change; a model-version change or a knowledge-graph refresh re-runs the full evaluation§14 · T9
9. Periodic reviewAllRetire or recalibrate rules with persistently low acceptanceReview the per-domain trend against the claim; retire rules that no longer earn their place§8.8
10. RecordsHuman sign-offEach finding stores the rule ID, citation, confidence and model version; the reviewer's disposition is a separate, attributed eventThe same, plus the rule version and knowledge-graph snapshot that produced the finding§15 · §16

What the two systems teach

1

A headline metric is evidence, not yet a claim.

Doscierge's F1 of 85.87 is one number across 29 domains. To become an acceptance criterion it needs the per-domain breakdown, a confidence interval, and the unaided-reviewer baseline measured on the same set (§5).

2

Precision is a safety control by another route.

Doscierge's first evaluation produced about 288 findings per package; precision engineering brought that to about 73. A reviewer handed 288 findings stops reading them. That is the automation-bias and alert-fatigue risk §13.6 is written for, so finding volume is a monitored metric, not a cosmetic one.

3

SME ownership can be built into the tool.

ProtoCheck's provisional → confirmed calibration with a named signature is what draft Annex 22 §4.1 asks for in principle (subject-matter experts own the acceptance criteria), applied here by analogy to a clinical tool. Only confirmed rules belong in the validated scope.

4

The eval harness is the validation evidence engine.

Regression before every merge is a pre-approved change envelope (T9, §14.2) in engineering form. Evidence is generated by the change itself, not written afterwards.

So what

The claim is defined: performance, for a context of use, on a defined population, against a baseline, kept true over time. The standard, the architecture and two working systems now sit on one lifecycle. §2 turns to the practitioners: thirteen questions, each pinned to a stage of that lifecycle.

§2 · Part A · Layer 4 · The practitioner's lens

Thirteen questions, one stuck point: execution, not principle

The practitioner's lens · where the lifecycle bites

Spine

Layer 4 · The practitioner's lensQ1–Q9 came from practitioners at a September 2026 web event on acceptance criteria for AI validation, Q10 from the author, and Q11–Q13 from practitioner conversations. Each question is a symptom. The table maps each one to the underlying pain, to what the market does not yet supply, to the lifecycle stage in §1.3 where it bites, and to the section that answers it.

Full tableThe thirteen questions mapped to underlying pain, market gaps, who feels it most and the section that answers each13 rows
#Question (paraphrased)Underlying painWhat the market lacks todayWho feels it mostStage (§1.3)Answered in
Q1More examples of CSV concepts applied to AIPeople know CSV well but cannot translate it to AIA side-by-side "CSV concept → AI pitfall" playbookCSV/QA leads3–5 · claims, testing§7
Q2Camera system: when does drift occur, do new labels need revalidation?They cannot tell drift apart from change controlA decision tree for drift vs. change vs. new intended useSterile/OSD manufacturing, MS&T7–8 · monitoring, change§12, §11
Q3Validation templates for AI agentsNothing to start fromAgent-specific templates (tools, autonomy, trajectories)Everyone3–6 · claims to release§9
Q4A validation example that follows industry practiceNo public precedentAn end-to-end worked exampleQA, CSV consultants5–6 · testing, acceptance§6
Q5Is a risk assessment needed when AI writes the templates?"AI-for-compliance" feels circularA proportionate rule for AI-assisted authoringSmall QA teams2 · risk§10
Q6Studies of human error rates to base AI acceptance onAnnex 22 §4.3 requires a baseline; nobody has measured theirsPublished baselines plus a method for measuring your ownQA, statisticians3, 6 · claim, acceptance§5
Q7Small biotech: monitoring the Cat 3 → Cat 5 shiftVendors are adding AI to validated SaaS with no warningA minimum viable AI governance program sized for 50–300 peopleBiotech Heads of Quality1, 7 · assessment, monitoring§17
Q8GxP assessment, risk, versioning, data, design, registry, monitoring, periodic review for AIThey need the whole AI lifecycle, with AI-specific controls on top of CSVAn integrated AI lifecycle SOP setMid-to-large pharma CSV functions1–9 · all§8
Q9 (after the session)Is any AI product Part 11 compliant, and how does the guidance differ?They expect a product certification that doesn't exist, and their audit trails don't capture what the AI contributedA clause-by-clause Part 11 → AI control map and a vendor question setCSV, QA and IT system owners buying AI-enabled SaaS10 · records§15
Q10 (author)How does change control work end to end for AI, and where can AI help with CSV itself?AI components are hidden inside one application record, so every change is either "trivial" or "revalidate everything"CMDB-level AI configuration items, T9-driven Standard changes, IQ/OQ/PQ redefined for AI, and a co-brain RACICSV leads, ITSM/ServiceNow owners, QA8 · change control§14
Q11 (practitioner)Which models may be used where, and what is drift when the model is locked?"Locked" and "adaptive" are device terms applied loosely; drift is blamed for process signals and vice versaA model-type × criticality decision table and a drift FAQ in CPV and clinical termsMS&T, data science, QA1 (a) · 7 (b)§4, §11
Q12 (practitioner)What does ALCOA+ require of an AI-assisted record?The audit trail shows the human saved the record, not that a model wrote most of itA principle-by-principle capture list and the AI data-integrity failure modesQA, data-integrity leads, system owners10 · records§16
Q13 (practitioner)How do we keep every AI system compliant with the right alerts, without alert fatigue?Each AI use has its own ad-hoc dashboard, or none; limits are copied from vendor decksOne site policy: five metric families, minimums by tier, limits derived from validated performance, a response ladderHeads of Quality, system owners, platform teams7 · monitoring§13

Read-across. Seven of the eight questions from the session ask "show me how", not "tell me what". The audience already accepts that AI needs validating. They are stuck on execution. Three signals stand out, and each one is a component of the claim:

Signal 1

The human baseline is missing (Q6, Q2)

Annex 22 §4.3 says model acceptance criteria must be "at least as high as the performance of the process it replaces", and that the performance of the process being replaced must therefore be known. Most companies have never measured their manual process's error rate. Measuring that baseline is a service on its own (§5).

Signal 2

Vendor-driven category creep (Q7)

Most small companies will not build AI. They will receive it inside eQMS, LIMS, EDC, and document-management SaaS releases. The governance problem is a supplier-management and change-control problem before it is a model-validation problem (§17, §14).

Signal 3

Agents are ahead of the guidance (Q3)

Annex 22 excludes generative AI and LLMs from critical GMP applications altogether. Yet teams are building agents today. They need a defensible way to position agents: non-critical with human-in-the-loop, or split so that the deterministic components carry the validated decision. The parent article's dual-path architecture (§10–11) is exactly that answer (§9, §1.2).

So what

The questions are not about whether to validate AI. They are about what the claim is and how to prove it, which is why each one is pinned to a stage of the §1.3 lifecycle. Before building on the claim, §3 clears away the views that get it wrong.

§3 · Part A · Common misconceptions

Ten views the industry still repeats, and the correction

Common misconceptions · what the claim is not

Spine

A claim defined wrongly cannot be proved. Vendor material, training content and conference talks on AI validation often repeat a small set of views that were reasonable a few years ago but are now outdated, device-specific, or incomplete. The corrections below matter because each one changes what a validation package must contain. The last column points to the section of this page that corrects the view in practice.

MisconceptionWhy it failsCorrect practiceAnchorCorrected in
"AI validation means proving the model is accurate." Accuracy is treated as the headline metric on monitoring dashboards and alert limitsA model has no validated state outside a use. The same model can be low risk as a drafting aid and critical as a release gate. Accuracy alone says nothing about the controls that bound the residual riskDefine the context of use first. Then validate the controls around the model for that use, with model metrics as supporting evidenceFDA Jan 2025 draft, Steps 1–3; Annex 22 §3.1; ISPE GAMP AI Guide (AI as a component of a computerized system)§1, §8.1, §6
"AI systems are probabilistic, so outputs may vary." Presented as a property of all AIA trained ML model that is frozen is deterministic: same input, same output. Only generative models sampled with non-zero temperature, and dynamic models, are notTreat determinism as a design choice. Put the critical decision in a static, deterministic component (rules, SPC, frozen ML); keep LLMs in non-critical, human-reviewed roles (dual-path)Annex 22 §1 (scope: static, deterministic models in critical use); parent §10–11§4
"With expert review, AI can support batch release and other GMP-critical decisions." Batch disposition support rated medium risk, and release recommendations high risk but acceptable with expert reviewDraft Annex 22 excludes generative AI and LLMs from critical GMP applications whatever the human oversight. Human-in-the-loop is the condition for non-critical use. Risk ladders that are not anchored to a criticality definition tend to rate the same use differently in different placesCritical GMP decisions (release, disposition, reject) run only through deterministic logic or a static ML model validated to Annex 22. An LLM may draft the narrative; it may not carry the decisionAnnex 22 §1; §9 Pattern B; §17 tiering§4, §9, §17
"The Expert-in-the-Loop is the guardrail." Framed as split-second, experience-based judgementAn unmeasured human review is a paper control. Automation bias and anchoring mean reviewers miss errors they would catch in a colleague's work. GxP review should be structured, not split-secondKeep the good parts: credentialled reviewers, Approve / Modify / Reject with a written rationale. Then measure the reviewer: seeded-error catch rate, and override rates in both directionsAnnex 22 §3.3; §7 "Human review"; §5§5, §7, §13
"95% accuracy, with alert at <95% and action at <90%." Generic limits for accuracy, hallucination (>1% / >3%) and overridesAccuracy on imbalanced data hides the critical class (§7 pitfall 1). Limits not tied to consequence or to a baseline cannot be defended. A 1% hallucination rate reaching a reviewer is unacceptable in a CSR (§6)Criteria per class and per subgroup, with confidence intervals, set by an SME before testing and no worse than the measured process baseline. Derive alert and action limits from the validated performance and the consequence of failure. A two-level alert/action structure is sound; the numbers must be yoursAnnex 22 §4.1–4.3, §5.2; §5§5, §13
"Drift happens as new data is introduced into the model"; three drift types; track accuracy monthlyA locked model does not change; the world it sees does. The three-type list misses prevalence shift and unseen classes (the dangerous case). In production you rarely have ground-truth labels, so a monthly accuracy trend is often not measurableMonitor what you can observe: input-space and out-of-distribution scores, the "undecided" (low-confidence) rate, class-wise output rates, override and edit rates, and a periodic reference check (e.g., AQL re-inspection of accepted units) that supplies ground truthAnnex 22 §9.2, §10.3–10.4; §12 table; §8.7§11, §12, §13
"Adaptive, continuously learning algorithms can be more accurate than locked ones" (a view that traces to a 2021 device-policy report)Draft Annex 22 excludes dynamic (self-updating) models from critical GMP use. Autonomous updates are usually too risky where there is GxP impactRetraining is a change: impact assessment, regression on the locked test set, QA approval. Anticipated changes can be pre-approved (T9)Annex 22 §1, §10.1; FDA PCCP (devices, by analogy only)§4, §14
FDA's AI position described only through device documents. Definitions and risks taken from the 2019 AI/ML SaMD discussion paper and device guidance, with the drug-side and EU documents left outSaMD and device-software guidance do not govern drug GMP or evidence used for drug regulatory decisions. Applying them without saying so gives the wrong frameworkFor drugs and biologics use: FDA Jan 2025 draft (7-step credibility framework; CDER/CBER); draft EU GMP Annex 22 (July 2025; final targeted Q4 2026); ISPE GAMP AI Guide (July 2025; complements GAMP 5 2nd Ed. Appendix D11); FDA–EMA Good AI Practice principles (Jan 2026). Cite device guidance only as an analogy§19 sources§4, §6
"Ask the LLM to rework its own output"; the LLM remembers earlier promptsUseful for personal productivity. In GxP, conversation memory is uncontrolled state that breaks reproducibility, and self-review is not independent verification. The same principle applies to migration tools: a tool must not verify its own resultsPinned model and prompt versions, a fresh context per case, N-run testing, and an independent deterministic checker. Do not use the same model as its own judgeAnnex 22 §10.2 (configuration control); §9; parent §11§9, §7
Guidance status quoted out of date. CSA described as still draft and "likely to be issued for all regulated industries"; software validation anchored on 21 CFR 820.70(i); Annex 11 described as "not a legal requirement"FDA issued the CSA guidance as final on 24 Sep 2025, from CDRH/CBER, for device production and quality-system software. The QMSR (effective 2 Feb 2026) replaced the old Part 820 text; software validation now sits under ISO 13485:2016 clause 4.1.6. EU GMP Annex 11 is part of the EU GMP Guide that manufacturers are inspected against, and its revision (July 2025 consultation) runs alongside Annex 22Apply CSA to pharma by analogy and say so. Anchor drug systems on Part 11, the predicate rules, the 2018 Data Integrity guidance, Annex 11, and GAMP 5 2nd Ed.FDA CSA final (Sep 2025); FDA QMSR page; §19§1.2, §16, §19

Each row moves the same way: from "is the model good?" to "is this use, with these controls, demonstrably at least as good as the process it replaces, and can we show it still is?"

So what

Two of the ten views turn on the same confusion: what "locked", "adaptive", "deterministic" and "probabilistic" actually mean, and which of them a GMP claim may cover. §4 settles the vocabulary before Part B sets the bar.

§4 · Part A · Practitioner question Q11 (a)

Locked, retrained or learning: which models the claim can cover

Practitioner question Q11 (a) · not from the session"Locked versus adaptive, static versus dynamic, deterministic versus probabilistic: which of these can be used where in GMP, and what do the terms actually mean?"
Spine

A claim is only as stable as the thing it is made about. This section fixes the model vocabulary, shows what draft Annex 22 allows in critical and non-critical GMP use, explains where the "locked vs. adaptive" language comes from (FDA device policy, applied to pharma by analogy only), and ends with a decision table. The drift half of Q11 is in §11.

4.1 Definitions: two axes, not one ladder

The terms are often used as a single scale from "safe" to "risky". They are two independent axes: how the model changes over time and how it produces an output.

Axis 1 · How the model changesDefinitionTypical examplesWhat the claim covers
Static / lockedWeights and configuration are frozen at release. The model changes only through a formal change.Frozen classifier on an AVI camera; a PLS or PCA model with a fixed NOC set; a rules engineOne model version, one claim. The claim holds until a change is made or the input space moves (§11)
Locked with periodic retraining (the realistic middle)Locked in operation; retrained at intervals or on triggers, each retrain released as a new version after regressionCPV MSPC model refreshed after a qualified new site is added; a deviation classifier retrained annually on adjudicated labelsEach version carries its own claim. Retraining is a change (Annex 22 §10.1), usually inside a pre-approved envelope (T9) with regression on the locked test set
Dynamic / continuously learningUpdates its own parameters from production data without a release stepOnline-learning recommenders; self-tuning thresholds fed from production feedbackNo stable object to make a claim about. Excluded from critical GMP use by draft Annex 22
Axis 2 · How the output is producedDefinitionTypical examplesWhat the claim covers
DeterministicSame input → same output, every timeRules; SPC; a frozen ML classifier or regression run without samplingA reproducible result. Conventional test-and-retest logic works
ProbabilisticThe same input can produce different outputs (sampling, temperature, non-deterministic serving)LLM text generation at non-zero temperature; some ensemble or stochastic inference servicesA distribution, not a result. Acceptance is set on N runs per case (§9), and the critical decision is kept off this component

Two consequences follow. First, "AI is probabilistic" is not a property of AI; a frozen ML model is as deterministic as a spreadsheet (§3, row 2). Second, a generative model is probabilistic and usually vendor-hosted, so it sits at the far end of both axes at once unless the version is pinned (§4.4).

4.2 What draft Annex 22 allows

Draft EU GMP Annex 22 (consultation text, July 2025) is scoped to static, deterministic models used in critical GMP applications. Dynamic models, and generative AI and LLMs, are excluded from critical GMP applications; the draft's principles may be applied, where applicable, to non-critical uses. "Critical" means a direct impact on patient safety, product quality or data integrity. That gives four cells:

Model type on two axes, with what draft Annex 22 allows in critical and non-critical GMP use Two axes: how the model changes over time (static or locked at the top, locked with periodic retraining in the middle, dynamic or continuously learning at the bottom) and how the output is produced (deterministic on the left, probabilistic or generative on the right). Static plus deterministic: in scope for critical GMP use with the full Annex 22 package; in non-critical use the principles apply where applicable with depth by risk. Locked with periodic retraining, deterministic: in scope between retrains as a sequence of static versions, each retrain a change under Annex 22 section 10.1 with regression; non-critical as above but lighter. Probabilistic or generative, including LLMs: excluded from critical use whatever the human oversight; allowed in non-critical use with a human in the loop, validating the controls around it. Dynamic or continuously learning: excluded from critical use; not covered by the draft for non-critical use, treat as high risk and prefer a locked variant. The design rule: put the critical decision in the top-left cell. HOW THE OUTPUT IS PRODUCED → Deterministic Probabilistic / generative HOW THE MODEL CHANGES ↓ Static /locked Locked withperiodicretraining Dynamic /continuouslylearning Static + deterministic CRITICAL · IN SCOPEFull Annex 22 package: intended use & sample space, SME criteria per subgroup,independent test data, explainability, "undecided" outcome, monitoring NON-CRITICAL · principles where applicable; depth by risk(Tier 2) Locked with periodic retraining CRITICAL · IN SCOPE between retrainsA sequence of static versions; each retrain is a change underAnnex 22 §10.1 with regression (T9) NON-CRITICAL · as above, lighter Probabilistic / generative (LLM) CRITICAL · EXCLUDED, whatever the human oversight NON-CRITICAL · ALLOWED with human-in-the-loop Validate the controls around it: grounding, checker,review, audit trail (§9 Pattern A) A vendor LLM behaves as "locked" only if the version ispinned, the deprecation schedule is managed and silentvendor changes are detectable (§4.4) Self-hosted, pinned weights, temperature 0: still excluded from critical use (§4.5) Dynamic / continuously learning (either output type) CRITICAL · EXCLUDEDNo stable object to make a claim about. Redesign as locked-with-retraining NON-CRITICAL · not covered by the draft; treat as high risk, and prefer a locked variant If unavoidable: snapshot versions for reconstructability, human review of every output, explicit QA acceptance of residual risk Design rule: put the critical decision in the top-left cell
In scope for critical useNon-critical: principles where applicableNon-critical: allowed with human-in-the-loopExcluded from critical useNot covered by the draft
Two axes, four cells. Draft Annex 22 reads across both axes: only the static, deterministic corner (and the retrained sequence of static versions) may carry a critical GMP decision. The table below is the source for each cell.
Critical GMP useNon-critical GMP use (human-in-the-loop)
Static + deterministicIn scope Full Annex 22 package: intended use and input sample space (§3.1), SME-owned acceptance criteria per subgroup (§4), independent test data (§5–6), explainability (§8), confidence with an "undecided" outcome (§9), operation and monitoring (§10)Principles applied where applicable; depth by risk (§17 Tier 2)
Locked with periodic retrainingIn scope between retrains, as a sequence of static versions; each retrain is a change under Annex 22 §10.1 with regression (T9)As above, lighter
Dynamic / continuously learningExcludedNot covered by the draft; treat as high risk, and prefer a locked variant
Probabilistic / generative (LLM)Excluded, whatever the human oversightAllowed with human-in-the-loop; validate the controls around it (grounding, checker, review, audit trail; §9 Pattern A)

The final text may widen the non-critical guidance; the EMA workshop of 30 June–1 July 2026 discussed whether dynamic and probabilistic models could be addressed (§19). The direction of travel does not change the design rule: put the critical decision in the top-left cell.

4.3 Where "locked vs. adaptive" comes from, and why pharma applies it by analogy only

The vocabulary is FDA device policy. The April 2019 discussion paper, Proposed Regulatory Framework for Modifications to AI/ML-Based Software as a Medical Device, introduced the distinction between a "locked" algorithm, which gives the same result each time for the same input, and an "adaptive" one that changes its behaviour using real-world data. The same paper proposed the Predetermined Change Control Plan, which became final guidance for AI-enabled device software functions on 4 December 2024 (Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions). Both documents govern devices under CDRH's premarket framework. They do not govern drug GMP or the evidence used in drug regulatory decisions (§3, row 8). Pharma borrows two things from them by analogy: the vocabulary, and the idea that anticipated changes can be described, tested and approved in advance. The GMP anchor for that second idea is Annex 22 §10.1 change control and your own change-control SOP; T9 is the operational form (§14.2).

4.4 A vendor LLM is "locked" only if the version is pinned and the deprecation schedule is managed

A hosted foundation model behaves like a locked model only under three conditions, all of which are yours to enforce, not the vendor's default:

  1. The version is pinned. The endpoint identifies an immutable model version, and your IQ records it (§8.3, §14.3). An alias that the vendor can repoint ("latest", "stable") is not a pinned version.
  2. The deprecation schedule is managed. Vendors retire versions on a schedule. A deprecation notice is a planned change with a regression suite ready to run (§14.5, scenario S2). The quality agreement must oblige the vendor to give notice (§17, element 4).
  3. Silent vendor-side changes are detectable. Safety filters, tokenizers, system-level instructions and serving infrastructure can change behind an unchanged version string. A daily configuration and behaviour check on a fixed probe set (T8 OPS-06; §13) is the only defence, and any unexplained shift is investigated as an unauthorised change (Annex 22 §10.2).

Without all three, the model is adaptive from your point of view even if the vendor calls it a release.

4.5 Decision table: model type × criticality → allowed? → controls

Model typeOutputCritical GMP use?Allowed?Minimum controls (in addition to the eight in §8)
Rules / SPC / deterministic codeDeterministicYesYes (GAMP 5 Cat 5, Annex 11)Conventional IQ/OQ/PQ; rule versioning with a named owner
Static ML (frozen classifier, PLS, PCA)DeterministicYesYes (Annex 22 in scope)Intended use and sample space; SME criteria per subgroup vs. measured baseline (§5); independent test set (§8.4); explainability; "undecided" outcome; monitoring (§13); retraining as a change (T9)
Static MLDeterministicNoYesPrinciples where applicable; Tier 2 assurance plan (§9)
Locked with periodic retrainingDeterministicYesYes, as a sequence of static versionsAs static ML, plus a T9 envelope defining what a retrain may and may not change, regression on the locked test set, QA approval per version
Dynamic / online learningEitherYesNo (draft Annex 22)Redesign as locked-with-retraining
Dynamic / online learningEitherNoNot recommendedIf unavoidable: treat as high risk, snapshot versions for reconstructability, human review of every output, explicit QA acceptance of residual risk
Generative / LLM, vendor-hostedProbabilisticYesNo, whatever the oversightSplit the use (§9 Pattern B): deterministic component carries the decision
Generative / LLM, vendor-hostedProbabilisticNoYes, with human-in-the-loopPinned version; managed deprecation; probe-set check (§4.4); grounding on approved sources; deterministic checker; N-run acceptance; measured reviewer catch rate (§5); AI contribution attributed in the record (§15, §16)
Generative / LLM, self-hosted, pinned weights, temperature 0Near-deterministic in practiceYesNo under the draft (generative models are excluded from critical use regardless of sampling)Same as above; the deterministic serving mode helps reproducibility but does not change the regulatory position
So what

The claim can only be made about a locked object with a deterministic output, and it may carry a critical decision only in that form. Part B now sets the bar for the claim: what "at least as good as the process it replaces" means, and how to show it once.

SpineABCDEF

Part B · Set the bar

Claim component Baseline and evidence

What "at least as good as the process it replaces" means, and how to show it once.

Bridge. Part A defined the claim: performance for one context of use, on a defined population, against a known baseline, for a locked model with a deterministic output wherever the decision is critical. Part B sets the bar that claim must clear and shows it being cleared once. §5 supplies the baseline that draft Annex 22 §4.3 presumes; §6 walks a clinical drafting agent through FDA's seven-step credibility framework; §7 lists where familiar CSV concepts go wrong when the bar is set the old way.

§5 · Part B · Audience question Q6

The baseline nobody has measured

Audience question Q6"Do you know of any studies that examine the error rates of humans that can be used as a basis for determining the acceptance of AI?"
Spine

This section establishes the baseline component of the claim. It matters more than any other question on the list, because draft Annex 22 §4.3 makes the human baseline a regulatory expectation. The baseline feeds the acceptance criteria in §8.18.2 and T6, the human-review test in §6, and the alert and action limits in §13.

Draft Annex 22 §4.3

"The acceptance criteria of a model should be at least as high as the performance of the process it replaces. This implies that the performance should be known for the process which is to be replaced."

5.1 Published anchors

DomainSourceWhat it showsHow to use it
Visual inspection of injectablesUSP <1790> Visual Inspection of Injections; Knapp–Kushner methodologyManual inspection is probabilistic. A unit belongs in the "reject zone" only if manual inspectors detect it with probability ≥ 0.7. Automated systems must show equivalent or better performance than manual inspectionThe most mature precedent in pharma for "AI ≥ human". Use it directly for any camera or vision AI (§12)
Clinical data processingGarza et al., Int J Med Inform 2025, 195:105749. Systematic review and meta-analysis of 93 papers (1978–2008)Error rates range from 2 to 2,784 errors per 10,000 fields depending on method. Single entry: 4–650 per 10,000; double entry: 4–33 per 10,000; medical record abstraction is the highestBaseline for AI in data transcription, abstraction, EDC entry, and SDV-adjacent use-cases (§6)
Human reliability analysis (industrial)THERP, Swain & Guttmann, NUREG/CR-1278 (1983); HEART (Williams)Nominal human error probabilities for routine procedural tasks (reading, recording, checking), and the limited effectiveness of a second-person checkBaseline for batch-record review and transcription-type tasks when you have no local data. Always prefer local measurement
Screening with AI supportMASAI randomized trial (Lång et al., Lancet Oncology 2023): AI-supported mammography screeningAI-supported screening detected more cancers than standard double reading, with a much lower reading workloadShows the right comparison design: human + AI vs. human alone, randomized, on real workflow, rather than model vs. human on a curated set
Manual transcription in manufacturingParent article §8.4 (field estimate, APQR)About 1–5 errors per 1,000 hand-transcribed batch-parameter valuesIllustrative only Replace with your own measurement

5.2 Measure your own baseline

Published numbers help frame the argument. Your acceptance criteria should rest on your process, measured with this protocol:

  1. Same test set for human and model. Use the independent, stratified, adjudicated set from Annex 22 §5 (§8.4, T4).
  2. Several operators under normal conditions. Include realistic pace, shift, and fatigue. Record each operator's performance, not just the average.
  3. Measure the right metrics. Sensitivity per critical class (false accepts are the patient risk), false-reject rate (the cost), and inter- and intra-rater agreement (kappa). Human inconsistency is often a bigger issue than the human average.
  4. Report confidence intervals. An AI with 97% sensitivity (95% CI 93–99%) does not demonstrably beat a human at 95% on a 150-unit test.
  5. Compare system to system. If the AI will run with a human reviewer, the comparison is human+AI vs. human alone, not model vs. human. The same measurement, repeated in operation as a seeded-error catch rate, becomes the human-in-the-loop metric family in §13.
Two caveats that catch teams out
  • The label ceiling. If humans label the test set, the model is scored against human judgment and cannot demonstrably exceed it. Where possible, use a stronger reference: lab testing, destructive testing, or panel adjudication (Annex 22 §5.3).
  • "No decrease" applies per subgroup. A model can beat humans overall and still be worse on the rare, critical subgroup, such as small particles in amber glass. Annex 22 §4.2 allows different criteria for different subgroups. Use that (§8.2, T6 §4).
Show me

Doscierge's evaluation harness reports P 92.31 / R 80.27 / F1 85.87 against expert-labelled ground truth (parent §17). By the logic of Annex 22 §4.3, that recall becomes an acceptance criterion only when it sits next to a baseline: the recall of a qualified regulatory reviewer working unaided on the same labelled packages, measured with the §5.2 protocol. Without that number, 80% recall could be a large improvement or a step backwards (§1.3, stage 3).

So what

With the baseline measured, an acceptance criterion becomes defensible: a metric per subgroup, with a confidence interval, no worse than the process it replaces. §6 shows the whole claim being made and proved once, for a clinical drafting agent.

§6 · Part B · Audience question Q4

A worked validation: CSR drafting under the seven-step credibility framework

Audience question Q4"Is it possible to see an example of a validation of an AI tool following industry practice?"
Spine

This section establishes the evidence component of the claim end to end: question of interest, context of use, risk, credibility plan, execution, results, adequacy. It uses the baseline from §5 and the controls that §8 then generalises.

Example chosen: UC-D4, clinical study report (CSR) drafting with generate-then-check (parent article §8.2). It is a clinical-development tool, the highest-volume GenAI use-case in the industry, and the one regulators are watching most closely. This walk-through uses FDA's seven-step credibility framework (January 2025 draft), because the CSR output supports regulatory decisions. It borrows Annex 22's test-data discipline where that applies. Numbers are illustrative

1

Define the question of interest

"Is this CSR efficacy-results section accurate, complete, and traceable to the locked TFLs, so that a medical writer can finalize it?"

2

Define the context of use

The drafting agent produces Sections 11–12 from locked TFLs and the SAP. The deterministic checker (MCP tool) checks required elements, numeric consistency against source tables, and cross-references. A medical writer reviews every sentence with its source span and signs under Part 11. The AI output is never the final record, and a qualified human is accountable (Pattern A, §9).

Step 3: Assess model risk (model influence × decision consequence; §8.2, T2)

FactorAssessmentRationale
Model influenceMediumThe human reviews everything, but the draft frames the reviewer's thinking (automation bias)
Decision consequenceHighA wrong efficacy number in a CSR reaches the submission
Model riskMedium–HighThe checker and human review are the risk-reducing controls. Their effectiveness must be demonstrated, not assumed

Step 4: Credibility plan

ControlEvidence requiredAcceptance criterion (illustrative)
Numeric fidelityChecker cross-validates every number in the draft against the TFL100% of numeric mismatches detected on a seeded set of 200 errors
GroundingSpan-level citation on every sentence0 unsupported claims reaching the reviewer across 40 CSR sections × 5 runs
CompletenessRequired ICH E3 elements present≥ 98% of required elements drafted; 100% of missing ones flagged by the checker
Human review effectivenessWriters review drafts with seeded errorsCatch rate with AI + checker ≥ catch rate on the manual QC baseline (§5)
ReproducibilitySame inputs, 5 runsNo run introduces a numeric error the checker misses
ConfigurationModel, prompt, and ruleset versions pinnedIQ record (§4.4, §14.3)
5

Execute

Use a test corpus of 40 sections from 8 completed studies across 3 therapeutic areas, not used in development and locked in an access-controlled repository. Labels (the correct content) come from the approved, published CSRs. The baseline comes from the company's historical QC findings on manually drafted CSRs, e.g., numeric discrepancies found per section at QC.

6

Document results and deviations

Report each metric with its 95% CI, by therapeutic area and section type. Any failure is investigated. Example: the agent paraphrased a secondary endpoint name, and the checker's terminology rule was then extended. That is a ruleset change, so it goes through change control and the ruleset is re-tested (§14.2).

7

Determine adequacy

QA and Clinical agree the model is credible for this context of use. The release is scoped: TA1–TA3 and Sections 11–12 only. Adding a new TA or new sections is a predetermined change (T9) with a defined retest.

What makes this "industry practice" rather than a demo: risk-based depth (CSA critical thinking), independent test data, a human baseline, a demonstrated rather than assumed human-review effect, a scoped release, and a monitoring plan (reviewer edit distance, checker findings per section, unsupported-claim rate; §13 UC-D4 example) that feeds periodic review.

Manufacturing counterpart (outline only)

UC-M4 blend endpoint. Static PLS model plus a deterministic endpoint rule, so it is in scope for Annex 22 as a critical application once it replaces the fixed-time step (§4.5). Intended use and sample space cover formulations, blender scale, and API lots. Acceptance means the endpoint agrees with the thief-sample/HPLC reference and is no worse than the validated fixed-time step (Annex 22 §4.3). The model sits in the QA-approved registry, and the advisory phase acts as an extended PQ.

So what

One claim, proved once. What goes wrong in practice is that teams reuse CSV habits that quietly lower or misplace the bar. §7 lists those habits, concept by concept.

§7 · Part B · Audience question Q1

Where familiar CSV concepts go wrong on AI

Audience question Q1"Is it possible to give more examples of the showed content to AI cases/pitfalls when applying it to AI?"
Spine

This section protects the bar. Each row takes a standard CSV concept, shows how teams misapply it to AI, and gives a concrete example from the parent article's use-cases. The "what good looks like" column is, in every row, a restatement of the claim: population, metric, baseline, kept true.

CSV conceptCommon AI pitfallExample (parent-article UC)What good looks like
URS"The model shall detect defects with high accuracy." That cannot be tested.UC-M4 blend endpoint: "detect uniformity"URS states the decision, the population, the metric, and the threshold: "Flag blend-uniform when moving-block SD of predicted API content < 1.5% w/w for 5 consecutive blocks; sensitivity to non-uniform blends ≥ baseline of the fixed-time step." (§1, last row)
Risk assessment (FMEA)Failure modes list software bugs only; model failure modes are missingUC-M2 CPV signal agentAdd model failure modes: false negative on a real drift (patient risk), false positive flood (alert fatigue, so real signals get ignored), out-of-distribution input from a new site, confidence mis-calibration (§8.2)
Non-functional requirementsFunctional AI requirements are written; system requirements are forgotten. Often-missed items include error messaging, error logging and overload handlingUC-D4 drafting agentAdd AI non-functional requirements: behavior on out-of-scope or low-confidence input (refuse or "undecided"), logging of model/prompt version per output, rate-limit and timeout behavior, and fallback when the vendor model is unavailable
Supplier assessmentVendor audit checks the QMS, but not how the model was trained or what data it usedVendor AI in eQMS deviation triage (§17)Ask for a model card, training-data provenance, test-set independence, the change-notification policy for model updates, and whether customer data trains the shared model
Leveraging vendor test evidenceThe vendor's published model evaluation is accepted as validation. Vendor IQ/OQ cannot be used as-is because it was not tested for your intended use; vendor documentation must be scrutinized and owned by the regulated company GAMP 5 2nd Ed., supplier leverage; Annex 11 §3Vendor AI summarization in a document systemUse the vendor's evaluation to scope your testing, not to replace it. Re-test on your own independent set for your context of use, focused on the subgroups that matter to you
IQInstallation verified. Model version not pinnedUC-D4 drafting agentIQ confirms model ID and version hash, prompt/config version, tool allowlist, and retrieval index version. All are configuration items (§4.4, §8.3)
OQTests run on "happy path" samples chosen by the developerUC-M3 PAT chemometric modelTests run on an independent, stratified test set that covers rare variations (Annex 22 §5.1). The developer never saw it (Annex 22 §6.2)
PQOne PQ run, then "validated"UC-M1 APQRParallel run against the manual APQR for a full cycle, with agreement within pre-set tolerance and every discrepancy explained (as in parent §8.4). Then continuous monitoring replaces the one-time PQ (§13)
Test deviationsThe CSV categories (system error, script error, tester error) are applied unchanged, so every model miss is logged as a "system error"UC-M3 PAT modelAdd AI categories: model error (the model is wrong), reference/label error (the adjudicated answer was wrong, so correct it under change control and re-score), and test-set coverage gap (the case is outside the defined input space). Each has a different fix
Independent verificationAn LLM is used to grade its own outputs ("LLM-as-judge" with the same model)UC-D4Verification must be independent of what it measures (the same principle CSV applies to data-migration verification tools). Use the deterministic checker, human adjudication, or at least a different, separately validated evaluator
Traceability matrixRequirements traced to tests. Data not tracedUC-M2Trace requirement → metric → test-set subgroup → result → monitoring KPI. The subgroup is the new column
Audit trail (Part 11)Logs the user action, not the AI's contributionUC-D3 clinical ops NBALog the AI output, its confidence, the model version, the evidence cited, and the human accept/modify/reject with rationale (parent G2 decision provenance; §15, §16)
Change controlRetraining treated as "data refresh, no change"UC-M3 PAT model recalibrationRetraining is a change. It gets impact assessment, a regression test on the locked test set, and QA approval in the model registry (Annex 22 §10.1; §14)
Periodic reviewAnnual checkbox: "system still in use"AnyReview the performance trend, drift metrics, human override rate, incidents, and vendor changes. Decide to continue, retrain, or retire (§8.8)
Data integrity (ALCOA+)Applied to records, not to training dataUC-M5 deviation investigationTraining data and labels are attributable (who labelled), original (source-linked), and accurate (adjudicated). Label errors are data-integrity errors (§16)
Human review"A human approves every output", so the risk is considered controlledUC-D4 CSR draftingMeasure whether reviewers actually catch seeded errors (challenge testing of the human). Automation bias makes an unmeasured human-in-the-loop a paper control (§5)

Three pitfalls worth a slide on their own

1

"Accuracy" on imbalanced data

A defect-detection model that labels everything "good" scores 99.8% accuracy when the defect rate is 0.2%. Acceptance criteria must be per class: sensitivity on critical defects, false-reject rate on good units.

2

Test set equals training distribution

Validation passes, and then the first batch from a new glass supplier fails. The input sample space must be defined before testing and monitored after deployment (§11, §13).

3

Treating the human as the only control

Parent-article G2 makes the point: an agent whose recommendations are accepted 98% of the time might be excellent, or its reviewers might have stopped reading. The difference shows up only if you seed known errors and measure the catch rate. Good design helps the reviewer. Make them choose Approve, Modify, or Reject, and write the reason in their own words. But the design does not replace measuring how well they review.

So what

The bar is set and the habits that lower it are named. Part C turns the claim into a set of controls that a validation package can be built from, and that an inspector can be walked through.

SpineABCDEF

Part C · Build the controls

Claim component Controls

The eight CSV lifecycle controls and their AI extensions; the templates that carry them.

Bridge. Part B set the bar: a measured baseline, criteria per subgroup with confidence intervals, and one worked claim proved end to end. Part C builds the controls that carry that claim through the lifecycle. §8 is the centrepiece: the eight CSV lifecycle controls, each with one AI extension, its template, and a CPV agent as the running example. §9 positions agents so that the controls apply to them, and gives the template pack. §10 closes with the proportionate answer to "what if AI wrote the templates?"

§8 · Part C · Audience question Q8 · centrepiece

Eight controls, one AI extension each

Audience question Q8 (paraphrased)"How do we apply CSV to AI in practice? For example AI GxP assessment, AI risk assessment, AI version control, AI data management, AI model design, AI registry, AI performance monitoring, AI periodic review."
Spine

This section establishes the controls component of the claim. It is the most complete question on the list, and it effectively describes an AI lifecycle SOP set. Each control below follows the same pattern: what CSV already does → what AI adds → the artifact → a worked example. The running example is UC-M2, the Continued Process Verification signal agent (parent article §8.4). Its agent card (parent §6) already declares most of these controls. Sub-numbers 8.1–8.8 are the ones the templates cite as Q8.1–Q8.8.

8.0 The eight controls at a glance

#ControlCSV todayAI addsTemplateShown in useWhat the inspector asks
8.1GxP assessmentIs it GxP? Part 11? GAMP category?Is the AI function critical? Static or dynamic? Deterministic or probabilistic? Which frame applies?T1free§17 tiering; §4.5 decision table"Why is this AI function classified as it is, and who decided?"
8.2Risk assessmentSystem FMEA on software failureModel risk = influence × consequence; AI failure modes by subgroup; residual risk ownedT2free§6 Step 3; §5 baseline"Why is that threshold appropriate for that risk?"
8.3Version controlSoftware version and config baselineWeights, data snapshot, locked test set, prompts, tool allowlist, index, pinned vendor model, thresholdsT5 CI list on request§14.1 CMDB; §4.4 pinning"Which model version produced this record?"
8.4Data managementMigration; record integrityTraining/validation/test data as validation evidence; labelling; test-set independence; AI output as synthesized dataT4 on request§5.2 protocol; §16 governance"Show me the test set never touched by the developers."
8.5Model / agent designFS / DSAlgorithm rationale; explainability; confidence with "undecided"; tool allowlist; autonomy bands; validatable-by-designT5 on request§9 agent properties; §6 Step 2"What happens when the AI is wrong or unsure?"
8.6RegistryValidated systems inventorySystem of record for every model and agent: identity, owner, use, tier, lifecycle state, versions, consumersRegistry entry / agent card§14.1 sync to CMDB; §17 inventory"Who depends on this model, and how did you know to re-assess them?"
8.7Performance monitoringIncident management; availabilityValidated metric, input drift, calibration, human-in-the-loop, cost/latency; alert and action limits; response playbookT8free§13 site policy; §11 drift FAQ; §12 camera plan"How do you detect degradation, and what did you do the last time?"
8.8Periodic reviewAnnual, often perfunctoryA review that decides: continue / conditions / retrain / restrict / retire; reviews the revalidation triggers; retirement disciplineT10 on request§14.4 scheduled task; §13.7 monitoring the monitors"How do you know reliability remains acceptable?"

The template pack (T1–T10) is complete. T1, T2 and T8 are published free at the links above; the remaining templates are available on request (§9, §18).

Show me

Read the matrix for one real system. For Doscierge, the GxP assessment finds a Category 5 checker plus a non-critical LLM path (§8.1). Version control covers 970 rules, the knowledge-graph snapshot and a pinned model (§8.3). Data management means the locked, expert-labelled evaluation set (§8.4). Monitoring watches findings per package and acceptance per rule (§8.7). §1.3 gives the full row for both systems.

Lifecycle view

The eight AI lifecycle controls as a loop Eight controls arranged in a loop, read clockwise from the top: 8.1 GxP assessment (concept phase; is the AI function critical, static, deterministic?), 8.2 risk assessment (project phase; model influence times decision consequence), 8.3 version control (cross-phase; prompts, data and vendor model are configuration items), 8.4 data management (project; the test set is validation evidence), 8.5 model and agent design (project; confidence gate with an undecided outcome), 8.6 AI registry (cross-phase; every component owned and versioned with its consumers), 8.7 performance monitoring (operation; input drift, not just output errors), 8.8 periodic review (operation; a decision, not a checkbox). The periodic review loops back to the GxP assessment with a continue, retrain, restrict or retire decision. At the centre: validate the controls around the model for a specific context of use, not the model in the abstract. Each node links to its section below. LOOP BACK · continue / retrain / restrict / retire Validate the controls around the model for a specific context of use, not "the model" in the abstract. Click a control, or use Tab and Enter, to open it 8.1GxP assessmentcritical? static? deterministic? 8.2Risk assessmentinfluence × consequence 8.3Version controlprompts, data, model = CIs 8.4Data managementtest set = validation evidence 8.5Model / agent designconfidence gate + "undecided" 8.6AI registryowned, versioned, consumers 8.7Perf. monitoringinput drift, not just errors 8.8Periodic reviewa decision, not a checkbox CONCEPT → PROJECT → OPERATION → RETIREMENT · 8.3 version control and 8.6 registry span every phase
ConceptProjectOperationCross-phase (system of record)
Lifecycle view. Concept (8.1 GxP assessment) → project (8.2 risk assessment, 8.5 model/agent design, 8.4 data management) → operation (8.7 performance monitoring, 8.8 periodic review) → retirement or successor. Version control (8.3) and the AI registry (8.6, the system of record across all phases) span every phase. The review decision closes the loop: continue, retrain, restrict or retire, and a new intended use starts the cycle again.
8.1 · Concept

8.1 AI GxP assessment

CSV today

Is the system GxP? Is it Part 11 relevant? What is the GAMP category?

What AI adds
  • Is the AI function GxP? The CSV test still applies: does it "touch" a regulated product, or collect, analyze, report, store or transmit regulated data? Critical processes include release, safety and efficacy data, recall, adverse events, and pharmacovigilance.
  • Is it critical (direct impact on patient safety, product quality, or data integrity)?
  • Is the model static or dynamic, and deterministic or probabilistic? Draft Annex 22 allows only static, deterministic models in critical GMP use (§4).
  • Which regulatory frame applies: Annex 22 (GMP), the FDA credibility framework (regulatory-decision support), or EU AI Act high-risk classification?
Artifact · T1 AI Use Intake & GxP Assessment (published templatefree)
UC-M2 example

GMP; the signal disposition is a regulated record; critical = Yes. The SPC path is deterministic, rules-based software (GAMP Cat 5 custom queries). The MSPC PCA model is static ML, so it is in Annex 22 scope. The LLM narrative that explains the signal is non-critical with human-in-the-loop. One agent, three components, three postures. This is the parent article's §11 table (§1.2) applied.

8.2 · Project

8.2 AI risk assessment

CSV today

System-level FMEA focused on software failure.

What AI adds
  • Model risk = model influence × decision consequence (FDA 2025 draft).
  • AI-specific failure modes: false negatives and false positives by subgroup, out-of-distribution inputs, drift, mis-calibrated confidence, automation bias, bias across sites or products, data leakage, adversarial input (for agents).
  • Risk sets the depth of every later control (CSA critical thinking), and the tier that sets monitoring minimums in §13.
  • Two CSV lessons apply with more force to AI. First, re-assess risk at every change, not only at validation (ICH Q9(R1) treats risk review as a lifecycle activity). Second, the system owner must formally accept residual risk and own it. For AI there is always residual risk.
  • The acceptance criteria that the risk rating justifies are set against the measured baseline from §5.
Artifact · T2 AI Risk Assessment (published templatefree)
UC-M2 example
Failure modeEffectSeverityDetectabilityControl
Missed real drift (false negative)Out-of-trend batch released; late CAPAHighLowSensitivity acceptance on seeded historical drifts; APQR cross-check; periodic re-challenge
Alert flood (false positives)Scientists ignore signalsMediumHighPrecision target; orchestrator de-duplication; alert-fatigue index KPI (§13.6)
New site data outside the NOC modelSpurious T² excursionsMediumMediumInput-space monitoring; site-specific NOC or model transfer protocol (§11 FAQ 5)
LLM narrative mis-states a chart valueWrong disposition rationaleMediumHighChecker numeric-consistency rule; human disposition
8.3 · Cross-phase

8.3 AI version control

CSV today

Software version and configuration baseline.

What AI adds

The configuration baseline grows to include:

model weights (hash)training-data snapshotthe locked test setfeature and pre-processing codehyperparametersprompts and system instructionstool allowlistretrieval index versionfoundation-model version (vendor-pinned; §4.4)thresholds and confidence policy

Draft Annex 22 §10.2 requires configuration control with "effective measures … to detect any unauthorised change".

Artifact · Configuration-item list in T5 (on request), plus the registry version history. §14.1 shows the same items as CMDB configuration items.
UC-M2 example

Agent card mfg.cpv.signal-agent v2.3.0; model: scoped-small@pinned-2026-07; MSPC model v1.4 with its NOC set reference; eval harness mfg.cpv.eval@1.4. Semantic versioning rule:

  • major = intended use or class set changed, so revalidate
  • minor = retrained on the same class set, so regression test
  • patch = non-functional change, so documented verification
8.4 · Project

8.4 AI data management

CSV today

Data migration and data integrity of records.

What AI adds
  • Training, validation, and test data are validation evidence.
  • Requirements: provenance and lineage; a documented, verified labelling process; test-data independence with access control and audit trail; staff independence (Annex 22 §6); documented pre-processing and exclusions; no AI-generated test labels (Annex 22 §5.6).
  • ALCOA+ applies to datasets and labels (§16.4).
  • Place AI on the data lifecycle (capture → maintenance → synthesis → usage → publication → archival → purge; the lifecycle view taken by MHRA 2018 and PIC/S PI 041-1). AI output is synthesized data: new values derived from other data. It needs the same attributability and metadata as any other GxP record: model version, inputs, confidence, and reviewer (§16). Once published outside the company (e.g., in a submission), it cannot be recalled, which is why review happens before publication.
  • Keep rejected and overridden AI outputs, and "undecided" results, with the reason. This follows the same principle as invalidated laboratory results, which stay in the record with their justification FDA, Investigating OOS Test Results, 2006, rev. 2022. They are also your best monitoring data (§13).
Artifact · T4 Data Management Plan (available on request)
UC-M2 example

Ground truth is 2,140 labeled batch-phases across 3 sites and 2 products (from the agent card). Labels are dispositions adjudicated by two process scientists. The test set is split before training, stored in an access-controlled repository, and the MSPC developers never saw it. Every data point traces back to the historian tag and the MES event frame.

8.5 · Project

8.5 AI model / agent design

CSV today

Functional and design specifications.

What AI adds
  • Documented algorithm choice and rationale.
  • Explainability approach (feature attribution such as SHAP or LIME, contribution plots; Annex 22 §8).
  • Confidence thresholds with an "undecided" outcome (Annex 22 §9).
  • For agents: tool allowlist, autonomy bands, escalation logic, and deterministic orchestration where the path must be reproducible (§9).
  • Design for validatability: keep the critical decision in the deterministic or static component (§4.5).
Artifact · T5 Model/Agent Design Specification (available on request; the agent card is its machine-readable form)
UC-M2 example

Four isolated specialists (SPC, capability, MSPC, lot-effect) and a deterministic orchestrator; each specialist sees ≤5 tools. MSPC excursions show a contribution plot (explainability). Confidence policy: ≥0.90 notify the scientist directly; <0.55 escalate to the MS&T lead.

8.6 · Cross-phase

8.6 AI registry

CSV today

Validated systems inventory.

What AI adds

A registry that is the system of record for every model and agent. It holds:

  • identity, owner, intended use, risk tier, and lifecycle state (draft → lab → candidate → factory → deprecated)
  • current and previous versions
  • evaluation record and approvals (security, quality, data steward, business)
  • which consumers depend on it (impact analysis for change control)

For a small company this is the AI inventory from §17 with more fields. At enterprise scale it is the agent registry and service catalog from parent §6, synchronised with the CMDB (§14.1).

Artifact · Registry entry / agent card
UC-M2 example

The parent §6 card, including consumers: [mfg.apqr.orchestrator, mfg.quality-command-center, mfg.deviation.investigator]. A change to the CPV agent automatically flags three downstream consumers for impact assessment.

8.7 · Operation

8.7 AI performance monitoring

CSV today

Incident management; system availability.

What AI adds
  • Monitoring of the model's validated performance metrics (Annex 22 §10.3).
  • Input-drift monitoring against the defined sample space (Annex 22 §10.4); what drift is and is not, in CPV terms, is the subject of §11.
  • Confidence-calibration monitoring.
  • Human-in-the-loop monitoring: acceptance, override, and edit rates (Annex 22 §3.3 and §10.5).
  • Cost and latency for agents.
  • Each metric has a threshold, an owner, and a response playbook: investigate → restrict to advisory → retrain → retire.
  • Four design questions make a good checklist: what is measured, how often, what the alert limits are, and what requires action. Monitoring without response criteria has little value. Use two levels, an alert limit and an action limit, as CPV already does, but derive the values from the validated performance and the consequence of failure (§3, row 5; §13.3).
  • Human overrides are often the earliest signal of degradation. Watch both directions: a rising override rate suggests the model is degrading, and a falling rate may mean the reviewers are disengaging.
  • §13 lifts these elements into one site policy, with minimum metrics by tier and alert-fatigue controls; T8 sets the parameters per use.
Artifact · T8 Operational Monitoring Plan (published templatefree)
UC-M2 example (thresholds illustrative)
KPIThresholdResponse
Precision (signals confirmed / signals raised), rolling 30 days< 0.80Investigate; tune; change control
Seeded-drift detection (quarterly re-challenge with historical drifts)< validated sensitivityRestrict to advisory; retrain
Input drift: share of batch-phases outside the NOC envelope for non-process reasons> 5%Model transfer or NOC refresh (change control)
Coach-mode acceptance, rolling 14 days< 0.70 or > 0.98Low: fitness review. High: seeded-error check for automation bias
Signal-to-disposition time> 5 working daysWorkflow escalation
8.8 · Operation

8.8 AI periodic review

CSV today

Periodic review of the validated state, often annual and often perfunctory.

What AI adds

A review that decides something. It looks at:

  • the performance trend against acceptance criteria
  • drift history
  • changes made and their cumulative effect
  • human-in-the-loop statistics
  • incidents, deviations, and CAPAs involving AI
  • vendor or model deprecation notices
  • regulatory changes
  • whether the intended use still matches actual use (scope creep)
  • whether the monitoring thresholds and false-alarm rates are still right (§13.7)

The output is a decision: continue / continue with conditions / retrain / restrict / retire. Frequency follows the risk tier (§17).

The review also tests the usual revalidation triggers: a new model version, retraining, new data sources, significant drift, a critical error, a major process change, or a regulatory change. The question is the one every periodic review asks: does the current evidence still support fitness for intended use?

Retirement

Apply the CSV retirement discipline GAMP 5 2nd Ed., retirement phase to the model too. Keep the model version, locked test set, validation evidence, and the AI audit data for the retention period of the records it influenced, so any past decision can be reconstructed (§15, §11.10(c) row). Do not migrate old audit trails into the successor system's audit trail; archive them alongside it.

Artifact · T10 Periodic Review Record (available on request)
UC-M2 example

A semi-annual review, aligned to the APQR cycle so that the CPV agent's own performance becomes an input to the product's APQR. The model is reviewed alongside the process it monitors.

Summary: from CSV to the AI lifecycle

ControlCSV artifactAI artifactAnchor
GxP assessmentGxP/Part 11 assessmentT1 + critical / static / deterministic classificationAnnex 22 §1; EU AI Act
Risk assessmentSystem FMEAT2 with model risk = influence × consequenceFDA 2025 draft; ICH Q9(R1)
Version controlConfig baselineExpanded configuration items (weights, data, prompts, tools)Annex 22 §10.2
Data managementMigration / DIT4: provenance, labelling, test independenceAnnex 22 §5–6
Model designFS / DST5 / agent card: explainability, confidence, autonomyAnnex 22 §8–9
RegistrySystem inventoryAI registry with lifecycle states and consumersGAMP AI Guide; parent §6
Performance monitoringIncident mgmtT8: metric, drift, and human-in-the-loop KPIsAnnex 22 §10.3–10.5
Periodic reviewAnnual reviewT10 with a continue / retrain / retire decisionAnnex 11; GAMP 5

What an inspector will ask, and where the answer lives

Inspectors are unlikely to ask "what was your model accuracy?" and more likely to ask how you set acceptable performance and what happens when the AI is wrong. Map each question to evidence before the inspection:

Likely questionEvidenceWhere it lives
How did you determine acceptable performance?Measured baseline; SME-approved criteria per subgroupT6 §4; §5
Why is that threshold appropriate?Risk rating (influence × consequence); link from criterion to failure modeT2; T6
What happens when the AI is wrong?Confidence gating, deterministic checker, human review with measured catch rate, deviation/CAPA routeT5; T6; §7
How do you detect degradation?Monitoring KPIs with alert and action limits; reference re-checksT8; §13; §12 monitoring table
How do you know reliability remains acceptable?Periodic review decisions; trend recordsT10
Who decided, and on what basis?Part 11 audit trail of AI output, model version, and human Approve/Modify/Reject with rationaleRegistry + system audit trail (§15, §16)
So what

Eight controls carry the claim from concept to retirement. They assume the AI component is positioned where the controls can reach it. For agents, that positioning is the first decision, and §9 makes it.

§9 · Part C · Audience question Q3

Position the agent first. Then validate the system of controls.

Audience question Q3"Is it possible to share any validation templates to validate AI Agents?"
Spine

The eight controls apply to an agent only once the agent has been positioned so that the critical decision sits on a component the claim can cover (§4). This section gives the two positions, what is different about validating an agent, and the template pack that carries the controls.

First, position the agent correctly. Draft Annex 22 excludes generative AI and LLMs from critical GMP applications. So an agent is validated in one of two ways:

Pattern A

Non-critical with human-in-the-loop

The agent drafts, summarizes, or recommends. A qualified human decides and signs. Validation focuses on the system of controls: grounding, the checker, human review, and audit trail. This is CSA-style assurance plus Annex 22 principles "where applicable". Most agents today fit here: UC-D4 documentation, UC-M5 deviation hypotheses, UC-D3 next-best-actions.

Pattern B

Split so the critical decision is deterministic

Anything that directly decides product quality or patient safety runs through deterministic components (rules, validated SPC, a static ML model) that are validated conventionally. The agent orchestrates and explains. This is the parent article's dual-path architecture (§10) and "validate the component, not the model" (§11; §1.2 above).

What is different about validating an agent (versus a single model)

Agent propertyValidation implication
Calls toolsTool allowlist is a configuration item. Test that the agent cannot call tools outside it. Each tool is validated on its own
Takes multi-step pathsTest trajectories, not just final answers: did it retrieve the right source, call the checker, and stop at the escalation threshold?
Prompts and system instructions drive behaviorThe prompt is configuration, versioned and under change control. A prompt change is a change (§14.5, S4)
Output is probabilisticRun each test case N times (e.g., 5) and set acceptance on the distribution: e.g., zero unsupported claims in any run, required elements present in ≥ 95% of runs (§4.1)
Can be manipulated through its inputsAdversarial and red-team tests: prompt injection in source documents, out-of-scope requests, conflicting sources
Confidence-gated autonomyTest each band boundary (parent G7: ≥0.90 / 0.70–0.89 / 0.55–0.69 / <0.55): does it escalate when it should?
Depends on a vendor modelPin the model version. Treat a vendor deprecation as a planned change with a regression suite ready to run (§4.4)

Template pack (the deliverable set)

Ten templates cover the lifecycle. Each one either replaces a familiar CSV document or is new to AI. The pack is complete. T1, T2 and T8 are published free; the others are available on request.

T1
AI Use Intake & GxP Assessment
Is it GxP? Critical or non-critical? Which pattern (A or B)?
Replaces / extendsGxP assessment
Published free →
T2
AI Risk Assessment
Model- and agent-specific failure modes, risk = model influence × decision consequence (FDA 2025 draft)
Replaces / extendsFMEA / system risk assessment
Published free →
T3
Intended Use & Context of Use Specification
Decision supported, input sample space, subgroups, human role
Replaces / extendsURS
T4
Data Management Plan
Training/validation/test provenance, labelling process, independence controls
Replaces / extendsNew
T5
Model / Agent Design Specification
Architecture, tools, prompts, retrieval sources, autonomy policy, the agent card
Replaces / extendsFS / DS
T6
Validation Plan & Test Protocol
Metrics, acceptance criteria per subgroup, trajectory and red-team tests
Replaces / extendsValidation plan / OQ-PQ protocols
T7
Validation Summary Report
Results with confidence intervals, deviations, explainability review, release decision
Replaces / extendsVSR
T8
Operational Monitoring Plan
Performance and drift KPIs, thresholds, owners, response playbook; site policy reference in its Appendix B
Replaces / extendsNew (partly periodic review)
Published free →
T9
Change Control Addendum (predetermined changes)
Anticipated changes with pre-approved protocols
Replaces / extendsChange control SOP
T10
Periodic Review Record
Trend review, continue / retrain / retire decision
Replaces / extendsPeriodic review

T1, T2 and T8 are published free. The full ten-template pack (intake, risk, context of use, data management, design, validation plan, summary report, monitoring, predetermined changes, periodic review) is available on request.

Open T1 →Open T2 →Open T8 →Request the full pack →
Right-size the pack

The full T1–T10 set is for Tier 3 (GxP-critical) uses. For a Tier 2 use, such as a human-reviewed drafting agent, T1 + T2 + a one-page assurance plan is often enough. That plan follows the CSA assurance-record pattern FDA CSA, 2025: risk rating, what we will test and what we will not, how (scripted, unscripted, or automated), what evidence we will keep, and the acceptance criteria. For AI, add the context of use, the baseline, and the monitoring KPIs. Part 11 controls stay a go/no-go gate whatever the tier (§15).

Canonical example: T6, the agent validation plan (skeleton)

AI Agent Validation Plan — skeletonT6 · TEMPLATE PREVIEW
# AI Agent Validation Plan — <agent id> v<version>

## 1. Scope & positioning
- Agent: <id, registry link>          Pattern: A (non-critical, HITL) | B (deterministic core)
- GxP area: GMP | GCP | GVP | GLP     Critical GMP use? <Y/N — if Y, only static/deterministic components may decide>
- Regulated record(s) produced/influenced: <e.g., CSR section draft; deviation investigation report>
- Human decision owner (role): <e.g., Medical Writer; QA Investigator>

## 2. Context of use (from T3)
- Question the agent helps answer: ...
- Input sample space & subgroups: <document types, TAs, sites, languages, edge cases>
- Out of scope (agent must refuse/escalate): ...

## 3. Configuration items under control
| Item                                      | Version/hash                 | Owner          |
|-------------------------------------------|------------------------------|----------------|
| Foundation model                          | <vendor/model@pinned-date>   | Platform       |
| System prompt / instructions              | <git sha>                    | Agent owner    |
| Tool allowlist (≤5)                       | <list + versions>            | Platform       |
| Retrieval sources / approved-source registry | <index version>           | Knowledge eng. |
| Deterministic checker rules               | <ruleset version>            | Quality        |
| Autonomy / confidence policy              | <bands>                      | Quality        |

## 4. Test design
| Test family                    | What it proves                            | Data                                          | Runs/case | Acceptance criterion                                          |
|--------------------------------|-------------------------------------------|-----------------------------------------------|-----------|---------------------------------------------------------------|
| Functional accuracy            | Output correct vs. adjudicated reference  | Independent test set, n=<>, stratified by subgroup | 5 | Per subgroup: <metric ≥ threshold, lower 95% CI ≥ baseline> |
| Grounding / hallucination      | Every claim traces to approved source     | Same set                                      | 5         | 0 unsupported claims reaching reviewer                        |
| Trajectory                     | Correct tools, order, escalation          | Scripted scenarios, n=<>                | 5         | 100% required steps executed; 0 disallowed tool calls         |
| Confidence gating              | Escalates below threshold                 | Boundary cases                                | 5         | 100% correct band routing                                     |
| Robustness / red team          | Resists injection, refuses out-of-scope   | Adversarial set                               | 3         | 0 policy violations                                           |
| Human-in-the-loop effectiveness | Reviewers catch seeded errors            | Seeded-error set                              | n/a       | Catch rate ≥ <baseline>                                    |
| Reproducibility                | Stable output across runs                 | Sample                                        | 10        | Decision-level agreement ≥ <x>%                             |
| Audit trail                    | Part 11 capture of AI output + human action | Execution logs                              | n/a       | 100% complete                                                 |

## 5. Test data independence
- Test set locked in <repo>, access-controlled, audit-trailed; developers never accessed (Annex 22 §6.2)
- Labels adjudicated by ≥2 SMEs; disagreements resolved by a third (§5.3); no AI-generated labels (§5.6)

## 6. Acceptance & release decision
- Baseline source for "no decrease" (Annex 22 §4.3): <measured human/manual performance — see §5>
- Release = all criteria met, or deviations justified and approved by QA

## 7. Post-release
- Monitoring plan: T8   · Predetermined changes: T9   · Periodic review: T10 (frequency by risk)

The full version of this skeleton, with an illustrative UC-D4 worked example, is in T6 (available on request). The remaining templates follow the same structure.

So what

The controls now reach the agent. One question follows immediately in every small QA team: if AI helped write these templates, does the helper itself need a risk assessment? §10 gives the proportionate answer.

§10 · Part C · Audience question Q5

AI-drafted templates need a proportionate risk assessment, not a per-document one

Audience question Q5"Would I need a risk assessment in having AI creating those templates?"
Spine

The controls apply to the tool that helps build the controls, in proportion to what it does. This section is the smallest application of the claim: a one-time classification of the authoring use, not a per-document exercise.

Short answer

Yes, but a proportionate one. For most teams it is a one-time classification, not a per-document exercise.

What matters is what the AI is doing:

AI roleExampleGxP riskRequired control
Authoring aid for a document a qualified human reviews and approvesDrafting a validation-plan template or SOP skeletonLow The approved document is the controlled record, and the tool is like a word processor with suggestionsRegister the use in the AI inventory; SOP for AI-assisted authoring; human review and approval as normal; no confidential data in non-approved tools
Generating validation content that feeds test decisionsDrafting acceptance criteria, test cases, or risk ratingsMedium Errors propagate into the validated stateAbove, plus SME ownership of every criterion (Annex 22 §4.2 requires an SME to define acceptance criteria), a check that cited regulations actually exist and say what is claimed, and a coverage check against the requirements
Generating test data or labelsSynthetic defect images, AI-labelled test setsHigh Annex 22 §5.6: "not recommended and any use hereof should be fully justified"Avoid for test sets. If used for training, justify, document, and keep the test set real and human-adjudicated (§8.4)
Executing tests or producing evidenceAI runs test scripts and records pass/failHigh The tool is now part of the validation evidence chainValidate the tool for that use (CSA: tools that directly support assurance need assurance proportionate to risk). Annex 11 §4.7 already requires a documented assessment of the adequacy of automated testing tools and test environments

Specific pitfalls with AI-drafted templates

Fabricated or misquoted regulatory references

The most common failure. Every clause citation needs checking against the source text.

Plausible omissions

The template looks complete but leaves out, say, test-data independence. Check it against a reference checklist such as Annex 22 §3–10.

Circularity

AI writing the acceptance criteria for AI. Keep criteria human-owned and baseline-anchored (§5).

Confidentiality

Pasting proprietary process details into a public LLM. Use approved enterprise tools only.

Recommendation

Add one line to the AI inventory, "Generative AI for GxP document authoring: low risk, human-approved outputs", backed by a short SOP. That is enough for the authoring-aid case. Escalate only when AI output feeds test decisions or evidence. Drafting SOPs sits at the lowest-risk end of most AI risk ladders. §17 maps its full ladder of use-cases onto the three tiers, and §14.6 extends the same rule to AI as a co-brain for the whole CSV process.

So what

The controls are built, positioned, and proportionate. A validated state for AI is a claim that must stay true after release. Part D is about keeping it true.

SpineABCDEF

Part D · Keep it true in operation

Claim component Change and monitoring

Drift, monitoring, and change control: how the claim stays valid after release.

Bridge. Part C built the controls and positioned the AI component where they can reach it. A validated state for AI is a claim about performance on a defined input space, and the world that supplies those inputs keeps moving. Part D is about keeping the claim true: what drift is and is not, in CPV and clinical terms first (§11); the audience's camera case, where a new label is a change rather than drift (§12); one site monitoring policy with the right alerts and no alert fatigue (§13); and change control end to end, so that every signal that needs a change gets one (§14).

§11 · Part D · Practitioner question Q11 (b)

Locked models don't drift. The world does.

Practitioner question Q11 (b) · not from the session"Our model is locked. What exactly drifts, how do we tell drift from a real process change, and when does drift become a change?"
Spine

This section keeps the claim's population component true. The claim was made for a defined input space; drift is the world leaving that space, and the first skill is to tell that apart from the process signals the AI exists to catch. The answers below lead with CPV and clinical examples; the camera case that the audience asked about is in §12, unchanged.

The vocabulary, once. A locked model's weights do not change (§4.1). Four things around it can:

KindWhat movesCPV (UC-M2) exampleClinical (UC-D4 / data processing) exampleModel changed?
Covariate (input) driftThe distribution of inputsA new equipment train or site feeds the MSPC model; a sensor is recalibrated; sampling frequency changesNew sites or countries join the trial; a new TFL template; longer protocols with new section structuresNo
Prior (prevalence) shiftHow often each class occursA campaign mix change: 80% product A instead of 50%A change in disease-stage mix; more amendments per protocol in a therapeutic areaNo
Concept driftWhat a label meansQA tightens a specification, so "out of trend" now means something newA protocol amendment redefines an endpoint; ICH E3 expectations change for a sectionNo, but the model is now wrong
New class (open-set)A category the model has never seenA new failure mode (a raw-material interaction never in the NOC history)A new adverse-event pattern; a new document type routed to the drafting agentNo, but it will force the new case into a known class

The drift FAQ

Show
1Does a locked model drift?Policy

No. Its weights and configuration are frozen; its outputs for any given input are the same as on the day of validation. What drifts is the relationship between the inputs it now sees and the input space it was validated on (Annex 22 §3.1, §10.4). "The model drifted" almost always means one of the four things in the table above happened, and each has a different owner and response.

2Is a vendor model update drift or change?LLM & vendor

Change, always. If the vendor swapped the model behind an unchanged feature name and you found out from behaviour rather than from a notice, that is an unauthorised change from your point of view (Annex 22 §10.2) and a contract failure (§17, element 4). Detection is a configuration and probe-set check (T8 OPS-06; §4.4), not a drift metric. The response is the T9 regression if the update is inside the envelope, or a Normal change if it is not (§14.2; scenario S2 in §14.5).

3What is "LLM drift" when the version is pinned?LLM & vendorClinical

Four different things, only one of which is drift:

Observed as "LLM drift"What it actually isResponse
The retrieval index or corpus changed (new SOP versions, new approved sources, re-chunking)A change to a configuration item (§8.3)Change control; regression on the grounding tests
A prompt, system instruction, temperature or tool definition changedA change; if it happened without a record, an unauthorised oneChange control (§14.5, S4); investigate if unrecorded
The input population changed (new document types, new therapeutic areas, longer or differently structured inputs)Covariate driftInput monitoring (§13); scope check against T3; T9 extension or Normal change if outside the validated space
The vendor changed something behind the same version string (safety filters, tokenizer, serving)A silent vendor changeProbe-set check daily; treat as unauthorised change; invoke the quality agreement
4A supplier or raw-material change shifts a CPV trend. Is that model drift or a real process signal?CPV

This is the question that matters most in CPV, and the answer is: disentangle the two before anyone touches the model. A real process signal is what CPV exists to catch. A new excipient supplier that changes blend behaviour, a new API lot that shifts dissolution, a granulation that runs wetter in summer: these are process changes, and the MSPC excursion is the model doing its job. Route them to deviation and change control on the process, not to model retraining. Retraining the model so the excursion disappears is tuning the alarm to the fault.

Model-input drift is the other case: the process is unchanged, but the data feeding the model has changed. A sensor recalibration, a renamed historian tag, a changed sampling interval, a unit conversion, an ERP master-data change that alters how batches are grouped. Here the process is fine and the model's inputs are wrong.

Disentangle with four checks, in order:

  1. Deterministic path first. Did the raw CPPs and CQAs move on the validated SPC charts? If they did, it is process. The MSPC model and the SPC path see the same data; if only the model reacted, suspect the inputs.
  2. Data lineage. Check the historian tags, calibration records, sampling configuration and MES event frames feeding the model (§8.4). A lineage change with no process change is input drift.
  3. Contribution plot. Which variables drive the T² or SPE excursion (§8.5 explainability)? A single instrumentation variable points to inputs; a coherent group of process variables points to the process.
  4. Reference test. A lab result or a manual review of the batch settles it.

Then the routing: process signal → deviation / CAPA / process change control; input drift → fix the data path, then assess the impact on the model (was the validated performance affected?); an approved and qualified new normal (for example, a second supplier qualified within specification) → NOC refresh under change control, pre-approved in T9 if it was anticipated (§14.2). The deviation record should state which of the three it was.

5Covariate, prior, concept and new class, in CPV terms: what does each one mean for the claim?CPV
CPV eventDrift kindIs the claim still valid?Response
A new site or equipment train starts feeding the modelCovariateNot for that site: it is outside the validated sample space (T3)New subgroup; extend the test set; regression on all subgroups; Normal change (§14.5, S1). Until then, route that site's signals to manual CPV review
Campaign mix changes (product proportions, batch sizes)Prior shiftSensitivity per product is unchanged; overall precision may moveCheck precision at the new prevalence; no model change; note it in the periodic review
A specification is tightenedConcept driftNo: the label "out of trend" has a new meaningChange control on the specification triggers relabelling of the affected training and test data and a retest (§12, concept-drift row)
A new failure mode appearsNew classNo, for that mode: the model has no concept of itDetect through the "undecided" rate, human disposition disagreement, and deviations; intended-use update; new acceptance criteria against the manual baseline (§5); full retest (§12 decision tree)
6Do I need ground-truth labels to monitor?CPVClinical

Not continuously. Labels arrive late (confirmed dispositions, lab results, QC findings) or never. Monitor what you can observe every day: input-distribution distance, out-of-distribution and "undecided" rates, class-wise output rates, override and edit rates, checker findings. Then add periodic reference checks that supply ground truth on a schedule: confirmed dispositions reconciled monthly for CPV, a re-challenge with seeded historical drifts each quarter, an AQL re-inspection sample for a camera, a QC audit of a sample of AI-drafted sections for clinical documents. Annex 22 §10.3 (performance) and §10.4 (input) together describe exactly this pairing. §13.1 lists the five metric families.

7How often do I retest?Policy

Two clocks. A scheduled re-challenge by risk tier (§13.2): quarterly for a critical use, semi-annual for Tier 2, at periodic review for Tier 1. An event-driven retest after any change to a configuration item, after maintenance that touches the physical inputs (a camera's lighting, a sensor), and after any drift metric crosses its alert limit and the investigation cannot explain it. "Monthly accuracy" is not a retest schedule if you cannot get monthly labels (§3, row 6).

8Who decides when drift becomes a change?Policy

The monitoring owner investigates and documents (Level 0 in T8). The system owner and QA decide whether to restrict the use (Level 1). Drift becomes a change when the root cause is known and the response touches a configuration item or the input-space definition: the process SME owns whether T3's input sample space must be redefined (Annex 22 §3.1), and QA approves the change classification, Standard inside the T9 envelope or Normal outside it. §14.4 shows the chain: Event → Incident → Problem → Change Request.

9Does drift in a non-critical, human-in-the-loop use need change control?Policy

The drift needs an investigation and a record in the monitoring log. The response to it, if it is a prompt edit, an index refresh or a retrain, is a change, even at Tier 2, because it alters a configuration item that the assurance plan relied on. What scales with the tier is the depth of the regression, not whether a record exists. A Tier 2 drafting agent whose edit-distance metric climbs and whose prompt is then adjusted has had a Standard change (§14.5, S4).

10Seasonal or cyclic patterns: drift or not?CPVClinical

Cycles (summer humidity in granulation, shift patterns, campaign sequences, the year-end surge of clinical documents) are part of the input space if the training window covered at least one full cycle. If it did not, the first season the model meets is covariate drift outside the validated space, and it will look like degradation. Three defences: build the NOC set or training set across a full cycle; put the drift metric itself on an SPC chart with a seasonal baseline (§13.3); and pre-approve a "seasonal extension of the NOC set" in T9 so that the first winter is a Standard change rather than a surprise.

11A clinical case: the TFL template or SAP version changes. Drift?Clinical

Change. The input format of the drafting agent's sources has changed, the deterministic checker's rules may no longer parse the tables, and the grounding tests need to run again. If the template change was anticipated in T9 (for example, a sponsor-standard TFL update), it is a Standard change with the pre-approved regression; if not, it is Normal (§14.2). The agent has not drifted; its context of use has moved.

12Can we retrain continuously "to keep up"?CPVPolicy

Not for a critical use: continuous retraining makes the model dynamic, and draft Annex 22 excludes dynamic models from critical GMP use (§4.2). For a non-critical use, each retrain is still a change with regression on the locked test set, and the labels that feed it must be human-adjudicated, not harvested from production acceptances (Annex 22 §5.6 by analogy; §8.4). Otherwise the model learns the reviewers' automation bias.

13Our override rate fell to 1%. Is the model getting better?Policy

Possibly. Or the reviewers have stopped reading. The two are indistinguishable from the override rate alone; only a seeded-error check separates them (§5, §13.1 human-in-the-loop family). Watch overrides in both directions, and treat a sustained fall toward zero as an alert, not as good news.

So what

Drift is the world leaving the validated space, and the response depends on which of four things moved. The audience asked the same question about a camera, where the tempting answer, "a new label is drift", is exactly wrong. §12 works that case, unchanged from the session.

§12 · Part D · Audience question Q2

A new label is not drift. It is a change to the intended use.

Audience question Q2"For an in-line AI camera system trained and validated on a dataset of labels, when would drift occur? When a new label is introduced into the dataset? And then would we need to revalidate anytime a new label is introduced?"
Spine

This section keeps the claim true on the audience's own example. The vocabulary from §11 applies unchanged; the new element is the decision tree for what a new label triggers.

Short answer

Adding a new label is not drift. It is a change to the intended use (the output space), so it goes through change control, and it always triggers retesting.

The scope of that retesting depends on how the change is made. Drift is something different: the model has not changed, but the world it sees has. It needs its own monitoring.

Scenario

Automated visual inspection (AVI) of lyophilized vials on a filling line. A deep-learning classifier sorts each unit as accept, reject:particle, reject:crack, reject:fill-height, or reject:cake-defect. Validated under Annex 11 plus draft Annex 22. Acceptance was set against manual inspection using the Knapp–Kushner method, which USP <1790> references for showing an automated system performs as well as or better than manual inspection (§5).

Four kinds of change people call "drift"

TypeWhat changedCamera exampleModel changed?How you detect itResponse
Covariate (input) driftThe distribution of imagesNew vial supplier with slightly different glass tint; LED ageing; lens film; new lyo cycle changes cake appearanceNoInput-space monitoring: image statistics (brightness histograms, embedding distance from the training distribution), out-of-distribution score (Annex 22 §10.4)Investigate. If still within the validated sample space, document it. If not, it is a change: retest or retrain
Prior (prevalence) shiftHow often each class occursA fill-pump issue raises fill-height defects from 0.1% to 1%NoSPC on reject rate per classUsually a process signal, not a model problem. Route to deviation. Check that precision holds at the new prevalence (§11, FAQ 4)
Concept driftWhat a label meansQA tightens the particle-size reject criterion from 150 µm to 100 µmNo, but it is now wrongNot detectable from data alone. Comes from change control on specificationsChange control. Relabel the affected training and test data. Retest
New class (open-set)A defect type the model has never seenA new "stopper-skirt deformation" defect appears after a stopper supplier changeNo, but it will force the new defect into a known class, or into acceptLow-confidence rate, disagreement with periodic manual re-inspection, and complaintsThis is where Q2's new label comes in: a new class means a new intended use
The most dangerous case is the unseen defect

A closed-set classifier has no "I don't know" option. It will put a stopper-skirt defect into whichever known class looks closest, which may be accept. That is why Annex 22 §9.2 asks for a confidence threshold with an "undecided" outcome, and why the monitoring plan needs periodic manual re-inspection of an AQL sample of accepted units. That sample is the only way to see what the model is missing.

Does a new label require revalidation? A decision tree

Decision tree: does a new label require revalidation? Start: a new label or defect class is proposed. First decision: is the output space changing, meaning a new class the model must emit? If no, only new examples of existing classes are added: retrain with the same class set, run change control with impact assessment, a regression test on the locked independent test set where every existing subgroup must still meet acceptance criteria, then QA approval in the model registry as a new model version; the scope is a targeted retest, not full revalidation. If yes, a new class: second decision, is the new class critical, a patient-safety defect such as particle, crack or closure integrity? If yes: intended-use update under Annex 22 section 3.1, the SME re-approves the input sample space, new acceptance criteria for the new class are set against the manual baseline using Knapp to measure manual detection probability for the new defect, the independent test set is extended with adjudicated examples, all classes are fully retested because new classes shift decision boundaries, plus an explainability review; this is effectively a revalidation of the model component while the platform IQ and OQ carry forward. If no, the class is cosmetic or non-critical: the same path, but the test set can be smaller for the new class and interim handling can route the new defect to manual inspection until the retest completes. New label / defect class proposed Is the output space changing? (a new class the model must emit) NO New examples of existing classes only (e.g., more crack images) Retraining with the same class set Change control 1 · Impact assessment 2 · Regression test on the LOCKED independent test set: every existing subgroup must still meet its acceptance criteria (no decrease) 3 · QA approval in the model registry; new model version Scope: targeted retest, not full revalidation YES: new class Is the new class critical? (patient-safety defect: particle, crack, closure integrity) YES: critical Intended-use update (Annex 22 §3.1) · SME re-approves the input sample space · New acceptance criteria for the new class, set against the manual baseline (Knapp: measure manual detection probability for the new defect) · Extend the independent test set with adjudicated examples · Full retest of ALL classes (new classes shift decision boundaries for existing ones), plus explainability review (§8) = a revalidation of the model component; the platform IQ/OQ carries forward NO: cosmetic / non-critical Same path, but: · the test set can be smaller for the new class · interim handling can route the new defect to manual inspection until the retest completes
Reading the tree. Two decisions set the retest scope: whether the output space changes, and whether the new class is critical. New examples of existing classes are a targeted regression on the locked test set. A new critical class is a revalidation of the model component, while the platform IQ/OQ carries forward.
The rule that saves the most work

Write the predictable changes into the validation plan before go-live. List the anticipated change types (new examples of existing classes, recalibration after a lighting change, a new container supplier within spec), the pre-approved test protocol for each, and the acceptance criteria. FDA formalized this for AI-enabled medical devices as the Predetermined Change Control Plan (PCCP) (§4.3). That framework is not binding on pharma manufacturing equipment, but the same approach is a defensible way to structure a GMP change-control SOP. Changes inside the plan run as pre-approved protocols (T9; §14.2). Changes outside it, such as a new critical defect class, go through full change control (§14.5, scenario S3).

Monitoring plan for the camera (minimum)

MetricFrequencyAlert ruleOwner
Reject rate per class (SPC, p-chart)Per batchWestern Electric rulesProduction / QA
Low-confidence ("undecided") ratePer batch> 2× validated baselineMS&T
Input-drift score (embedding distance vs. training distribution)DailyAbove a pre-set thresholdAutomation / data science
Manual re-inspection of an AQL sample of accepted unitsPer batch or per campaignAny critical defect found → deviationQA
Knapp re-challenge with the qualified defect kitPeriodic (e.g., quarterly) and after any maintenanceDetection probability below the qualified valueQA / Validation
Human-in-the-loop reviewer performance (for manually reviewed rejects)MonthlyBelow operator qualification levelQA training

The same six rows, with illustrative limits, are the worked example in T8 Appendix A. §13 shows how they fit the site policy's five metric families.

So what

One camera, one CPV agent and one drafting agent each need a monitoring plan, and each plan so far has been built by hand. §13 turns them into one site policy, with parameters set per use.

§13 · Part D · Practitioner question Q13

One monitoring policy for every AI use, with the parameters set per use

Practitioner question Q13 · not from the session"How do we keep every AI system on the site compliant with the right alerts, without a different dashboard for each one and without alert fatigue?"
Spine

This section keeps the claim true as a policy, not as a series of one-off plans. One site policy fixes the metric families, the minimums by tier, the way limits are derived, the ownership and response ladder, and the alert-fatigue controls. T8 sets the parameters for each use, and T8 Appendix B carries the policy text in reference form. §8.7 gave the control; §11 gave the vocabulary; this section makes them the same for every AI function on the site.

13.1 Five metric families, every AI function

1
Inputs / drift
Whether the inputs are still inside the validated sample space
No ground truth needed
T8 §3 · runs continuously
2
Outputs / performance
Whether the validated claim still holds
Ground truth on a schedule
T8 §2, §4 · supplies the reference check
3
Human-in-the-loop
Whether the reviewers are still reviewing, and whether the AI is still helping
Seeded checks supply it
T8 §5 · runs continuously
4
System / vendor
Whether the configuration is what was validated, and whether it is about to change
No ground truth needed
T8 §6 · runs continuously
5
Data integrity
Whether every AI-touched record is attributable and complete
No ground truth needed
T8 §6 (OPS) and §16 · runs continuously
Full tableFive metric families: what each watches, typical metrics, whether ground truth is needed, and the T8 section5 rows
FamilyWhat it watchesTypical metricsGround truth needed?T8 section
1. Inputs / driftWhether the inputs are still inside the validated sample spaceDistribution distance vs. training data (embedding distance, population stability index); out-of-distribution and "undecided" rates; share of inputs from sources outside the validated scope; class prevalence on a p-chartNoT8 §3
2. Outputs / performanceWhether the validated claim still holdsSensitivity on the critical class; precision or false-reject rate; agreement with confirmed dispositions; periodic re-challenge with a qualified set; calibration (Brier score, band distribution)Yes, on a scheduleT8 §2, §4
3. Human-in-the-loopWhether the reviewers are still reviewing, and whether the AI is still helpingAcceptance rate (both tails); override and modification rates in both directions (AI said reject → human accepted, and the reverse); seeded-error catch rate; edit distance and time-on-reviewSeeded checks supply itT8 §5
4. System / vendorWhether the configuration is what was validated, and whether it is about to changeModel, prompt, index and threshold hash checks; vendor version and deprecation date; probe-set behaviour check; latency; cost per decision; degraded-mode activationsNoT8 §6
5. Data integrityWhether every AI-touched record is attributable and completeAudit-trail completeness for AI events (output stored as generated, model version, sources, reviewer action); share of AI-originated fields without attribution; retention checks on AI context artifactsNoT8 §6 (OPS) and §16

Families 1, 3, 4 and 5 need no labels and run continuously; family 2 supplies the ground truth on a schedule (§11, FAQ 6).

13.2 Minimum metrics and frequency by risk tier

The tiers are the §17 tiers, set by the §8.2 risk method (model influence × decision consequence, plus Annex 22 criticality). The minimums below are a floor; T8 can add to them, never below them.

TierTypical usesMinimum metricsFrequencyReference check (ground truth)
High (Tier 3, GxP-critical) Tier 3AVI camera; MSPC in CPV with automated disposition support; PAT endpointAll five families. At least: one drift score, one OOD/undecided rate, one out-of-scope input check; sensitivity on the critical class and false-reject rate; calibration; acceptance and override in both directions; seeded-error catch rate; configuration hash; vendor version and deprecation; audit-trail completenessInputs and configuration daily or per batch; performance per batch where a reference exists; human-in-the-loop weekly; calibration weeklyPer batch or campaign for the critical class (e.g., AQL re-inspection, confirmed dispositions); qualified re-challenge quarterly and after maintenance
Medium (Tier 2, GxP, human-in-the-loop) Tier 2CSR drafting; deviation classification with investigator confirmation; CPV narrative agentFamilies 1, 3, 4, 5, and one performance proxy. At least: undecided or out-of-scope rate; checker findings per item; acceptance and edit rate; seeded-error catch rate; configuration hash and vendor version; audit-trail completenessWeekly, with configuration checks dailyQuarterly sample audit against an adjudicated reference; seeded-error check quarterly
Low (Tier 1, non-GxP or authoring aid) Tier 1SOP skeleton drafting; internal summariesInventory entry current; acceptable-use SOP compliance; vendor version; shadow-use detectionAt periodic reviewNone required; spot check at periodic review

13.3 Setting alert and action limits

Copying generic percentages is the failure §3 (row 5) describes. Derive limits from three inputs:

  1. The validated performance and its confidence interval (T7). The action limit for a performance metric is the lower bound of the validated confidence interval, because below it the claim is no longer supported. The alert limit sits between the point estimate and the lower bound, set so that an ordinary run of bad luck does not trigger it; a common choice is the value that would be crossed by chance less than once a quarter at the sampling rate in use. Illustrative: validated particle sensitivity 0.97 with a lower 95% bound of 0.94 gives an action limit of 0.94 and an alert limit near 0.95 on the quarterly re-challenge (T8 Appendix A, PERF-03).
  2. The consequence of failure (T2). A metric that guards a High-severity failure mode gets a tighter alert limit and a faster escalation SLA than one that guards a Medium. Every High failure mode in T2 must map to at least one KPI (T8 §8).
  3. The baseline (§5). For human-in-the-loop metrics, the action limit on the seeded-error catch rate is the measured human baseline, because below it the human+AI system is worse than the process it replaced.
Illustrative control chart: validated performance, its confidence band, the alert limit and the action limit An illustrative control chart of particle sensitivity on the quarterly re-challenge. The validated point estimate is 0.97, shown as a dashed centre line; the 95 percent confidence band from 0.94 to 0.99 is shaded. The alert limit sits near 0.95, between the point estimate and the lower bound; the action limit equals the lower confidence bound, 0.94. Eight quarterly results drift downward: the sixth crosses the alert limit and opens a Level 0 investigation; the seventh crosses the action limit and triggers Level 1, restrict to advisory; the eighth, after the response, is back inside the band. Values are illustrative, not benchmarks. 95% CI of the validated performance · 0.94–0.99 SENSITIVITY ON THE CRITICAL CLASS (ILLUSTRATIVE) 1.00 0.97 0.95 0.94 0.90 Validated point estimate 0.97 Alert limit ≈ 0.95 · Level 0 investigate Action limit = 0.94 (lower CI bound) · Level 1 restrict to advisory alert → L0 action → L1 after response Q1Q2Q3Q4Q5Q6Q7Q8 Quarterly re-challenge with the qualified defect kit (T8 Appendix A, PERF-03) · illustrative values, not benchmarks
Alert limit vs. action limit, illustrated. The action limit is the lower bound of the validated confidence interval, because below it the claim is no longer supported. The alert limit sits between the point estimate and that bound, set so that ordinary variation does not trigger it. Values are the illustrative ones from point 1 above.

SPC on the metrics themselves. Drift scores, undecided rates, override rates and edit distances are process data. Put each on a control chart (individuals or EWMA for continuous metrics, p-charts for rates) with limits set from the first validated weeks of operation, and use run rules (Western Electric or equivalent) as the alert condition rather than a single fixed threshold. Cyclic patterns get a seasonal baseline (§11, FAQ 10). Changing a limit is a change to the monitoring configuration CI (§14.4, rule 1).

13.4 Ownership, escalation and the response ladder

0
Investigate
Any alert limit, or a run-rule violation on a metric chart
Monitoring owner · Event → Incident
1
Restrict to advisory
Action limit on a critical KPI; or an unexplained alert repeated
System owner + QA · Deviation; Problem
2
Retrain / reconfigure
Confirmed drift or degradation with a known cause
System owner + SME + QA · Change record
3
Retire / suspend
Critical error reached a GxP record, or performance cannot be restored
QA · Deviation / CAPA; registry → deprecated
LevelTriggerResponseDecision ownerEscalation SLA (illustrative)Record
0 · InvestigateAny alert limit, or a run-rule violation on a metric chartConfirm the signal; check for process and data-path causes before suspecting the model (§11, FAQ 4); documentMonitoring ownerOpened within 1 working day (Tier 3) / 5 (Tier 2)Monitoring log; ITSM Event → Incident (§14.4)
1 · Restrict to advisoryAction limit on a critical KPI; or an unexplained alert repeatedRestrict the AI function to advisory or to the unaffected scope; increase human review to 100% of affected outputsSystem owner + QADecision within 1 working day of the action-limit breach (Tier 3)Deviation; ITSM Problem
2 · Retrain / reconfigureConfirmed drift or degradation with a known causeChange via T9 if inside the envelope, else full change control; regression on the locked test set; PIR on the first 30–90 daysSystem owner + SME + QAChange raised within 5 working days of the root causeChange record (§14.2)
3 · Retire / suspendCritical error reached a GxP record, or performance cannot be restoredSuspend the AI function; revert to the manual process; impact assessment on past decisions using the AI audit data (§15, §16)QAImmediate containment; impact assessment opened within 1 working dayDeviation / CAPA; registry state → deprecated

Every Level 1–3 event is an input to the next periodic review (T10) and a trigger to re-assess T2.

13.5 Tie-in to change control

The ladder is the front end of §14.4. An alert-limit breach opens an Event and, if confirmed, an Incident (Level 0); an action-limit breach opens a Problem (Level 1); a Problem whose root cause touches a configuration item becomes a Change Request, Standard inside the T9 envelope and Normal outside it (Level 2). A vendor deprecation notice enters the same chain as a planned change. The monitoring configuration is itself a CI, so tuning a limit is a Standard change if the range was pre-approved in T9, and a Normal change if not.

Alert limitEvent → Incident= Level 0Action limitProblem= Level 1Root cause touches a CIChange Request (Standard in T9 · Normal outside)= Level 2See the loop in §14.4 →

13.6 Alert-fatigue controls

Monitoring that people ignore is a paper control, like an unmeasured reviewer.

  • Tiering of alerts. Only action-limit breaches on Tier 3 KPIs page anyone in real time. Alert-limit breaches go to a daily digest; Tier 2 alerts to a weekly review. The routing is written in T8, not left to the tool's defaults.
  • Suppression windows with justification. A known cause (an LED replacement scheduled for Tuesday; a planned index refresh) may suppress a specific metric's alerts for a bounded window. The window, the reason and the approver are recorded, and the metric is checked when the window closes. Open-ended suppression is not allowed.
  • De-duplication and correlation. One physical cause (a lighting change) can trip a drift score, an undecided rate and a reviewer-agreement metric at once. The Event layer correlates them into one Incident (§14.4), so three alerts do not become three investigations.
  • Alert usefulness review. Each periodic review reports, per KPI: alerts raised, alerts that led to a finding, alerts closed as false alarms. A KPI whose alerts never lead to a finding gets its limit re-derived or is retired; a KPI that never alerts but whose failure mode still exists gets a reference check added.

13.7 Monitoring the monitors

Thresholds age. Each periodic review (T10) re-derives the limits from the current validated performance and the observed distribution of each metric, reviews the false-alarm and missed-signal rates from §13.6, confirms that every High failure mode in the current T2 still maps to a live KPI, and checks that the reference checks actually ran at the stated frequency. The output is either "limits confirmed" or a change to the monitoring configuration CI. The review also asks whether the intended use still matches actual use (§8.8), because scope creep is the drift the metrics cannot see.

13.8 Two worked examples (all values illustrative)

UC-M2, the CPV signal agent (Tier 3 for the MSPC and SPC components; Tier 2 for the narrative).

FamilyKPIFrequencyAlertActionOwnerResponse
InputsShare of batch-phases outside the NOC envelope for non-process reasons (after the §11 FAQ 4 triage)Per batch> 3% rolling 30 days> 5%MS&TL0; NOC refresh via T9 if a qualified change caused it
InputsInputs from a site, product or equipment train not in the validated scopePer batchAnyAnySystem ownerRoute that site to manual CPV review; Normal change (S1)
OutputsPrecision (signals confirmed / signals raised)Rolling 30 days< 0.85< 0.80MS&TL0 → L2 if persistent
OutputsSeeded historical drifts detected on re-challengeQuarterly< point estimate< validated lower boundValidationL1 → L2
HumanCoach-mode acceptance, rolling 14 daysWeekly< 0.75 or > 0.97< 0.70 or > 0.98System ownerLow: fitness review. High: seeded-error check
HumanSeeded-error catch rate of scientists on narratives with planted numeric errorsQuarterly< validated< manual baselineQARetrain reviewers; restrict narrative
System/vendorHash check on MSPC model, SPC ruleset, prompt, thresholdsDailyAny mismatchAny mismatchAutomationStop; unauthorised-change investigation
System/vendorNarrative LLM version vs. pinned; days to vendor deprecationDaily< 90 days to deprecationVersion mismatchPlatformPlanned change (S2)
Data integritySignals with a disposition but no stored narrative-as-generated, model version or reviewer rationaleWeeklyAny> 1%QAFix the record path; deviation if a signed record is affected
Response example

A T² excursion cluster begins after a new lactose supplier's first three lots. The SPC path shows blend-time and moisture moved with it: process signal. Deviation opened on the process; the model is not touched; the excursion stands as a correct detection in the precision count. Two months later the second supplier is qualified within specification, and the NOC set is extended under the T9 "qualified supplier within spec" pre-approved change with regression on the locked test set.

UC-D4, CSR drafting with generate-then-check (Tier 2).

FamilyKPIFrequencyAlertActionOwnerResponse
InputsSections from a therapeutic area, study design or TFL template outside the released scopePer sectionAnyAnyClinical system ownerRoute to manual drafting; T9 if pre-approved, else Normal change
InputsChecker parse failures on source tablesWeekly> 2%> 5%PlatformInvestigate template change (§11, FAQ 11)
OutputsUnsupported claims caught by the checker per sectionWeeklyAbove the validated rateAny reaching a signed recordClinical QAL1: 100% second review
OutputsNumeric mismatches found by writers that the checker missedPer sectionAny2 in a quarterClinical QAL1 → ruleset change under change control
HumanWriter edit distance per section, and time-on-reviewWeeklySharp fall vs. validationMedical writing leadAutomation-bias check
HumanSeeded-error catch rate (planted numeric and endpoint errors)Quarterly< validated< manual QC baseline (§5)Clinical QARetrain; restrict to outline drafting
System/vendorPinned model version, prompt sha, approved-source index versionDailyAny mismatchAny mismatchPlatformStop; unauthorised change
Data integritySigned sections without the AI draft stored as generated, sources and reviewer actionWeeklyAnyAnyClinical QADeviation; fix before the next signature
Show me

Doscierge's first evaluation produced about 288 findings per test package, and precision engineering brought that down to about 73 (parent §17). For this policy the lesson concerns monitoring, not only engineering. Finding volume per package belongs in the outputs family with an action limit, because a reviewer handed 288 findings stops reading them: automation bias arriving through alert fatigue (§13.6).

So what

The policy keeps the claim true across the site and turns every confirmed signal into an Event, an Incident, a Problem or a Change. §14 shows the change side of that chain end to end, from CMDB to release.

§14 · Part D · Author's question Q10

Change control works when four records stay linked

Author's question Q10"Take specific use-cases and show how change control comes together: a ServiceNow-style model (business application, the applications beneath it, their specific change control), IQ, OQ, PQ, release for use in production, and ongoing monitoring for an AI-based system. Then add AI as a co-brain helping with the CSV/CSA process itself (risk, requirements and the other components). What does the human do and stay accountable for, and how does all of this map to change control?"
Spine

This section keeps the claim true through change. Every change to a model, prompt, index, test set or threshold either stays inside a pre-approved envelope or re-proves the claim in proportion to the change. Sub-numbers 14.1–14.8 are what the rest of the page cites as Q10.1–Q10.8.

The short answer

Change control for AI works when four records stay linked and in step:

The service model (CMDB)

What is running and what depends on it.

The AI registry and agent card

How each AI component behaves and was validated (§8.6).

The change record

What is changing, who approved it, and the evidence.

The validation record (VLMS or eQMS)

IQ/OQ/PQ protocols, results, and the summary report (VSR).

The AI-specific move is to make each AI component (model, prompt/configuration, retrieval index, locked test set, tool server) a configuration item of its own, with a versioning rule and a pre-approved change envelope. A model retrain then becomes a traceable CI change with a defined validation scope, not an invisible "data refresh". AI can do much of the legwork: impact analysis, drafting, test generation, evidence review. It never sets acceptance criteria, approves, or signs.

14.1 The service model: where AI components sit in a ServiceNow-style CMDB

This uses ServiceNow's Common Service Data Model (CSDM) hierarchy, extended with AI configuration-item classes. The class names are illustrative. Most organizations model them as custom CI classes, or sync them from the AI registry. Expand the narrative agent to see the AI configuration items nested beneath it.

Deterministic (Cat 5 / rules)Static ML (Annex 22 in scope)LLM / generative (Pattern A, HITL)Validation assetPlatform / vendor
Business capabilityManufacturing Quality Assurance
Business application
Continued Process Verification (CPV) Platform
GxP = GMPcriticality = HighSystem OwnerBusiness Process OwnerQA ownervalidation status = ValidatedAI risk tier = 3 (critical components present)
Application serviceCPV Signal Service – PROD (Sites A, B, C)← change control applies here
  • SPC/Capability query layerv3.2deterministicGAMP Cat 5 · Pattern B core
  • MSPC model (PCA, NOC set ref)v1.4static MLAnnex 22 in scope · QA-approved in registry
  • Narrative agentv2.3LLMnon-critical · HITL · Pattern A
    3 nested AI configuration items
    • Foundation model endpointscoped-small@pinned-2026-07vendor
    • Prompt / instruction setsha 9f2c… (v12)config
    • Tool allowlist[spc_calculator, capability_calculator, kg_lookup]config
  • Orchestrator workflowv2.3deterministiccontrol flow
  • Deterministic checker rulesetv1.7deterministicnumeric consistency, required elements
  • Locked test setcpv-test-2026Q2validation assetaccess-controlled
  • Monitoring config (T8 KPIs)v1.2configalert/action limits are configuration too
Depends on ► Data product batch_phase_summary v4.1 ► Historian (PI) ► MES event frames
Application serviceCPV Signal Service – VAL/QAIQ/OQ executed here
Application serviceCPV Signal Service – DEV/Labno GxP data rules; no validation state
Business capabilityClinical Development
Business application
Clinical Document Authoring (UC-D4)
GxP = GCPTier 2Pattern A
Application serviceCSR Drafting Service – PROD
  • Drafting agentv1.9LLM
    • Vendor LLM endpointpinnedvendor
    • Prompt setconfig
    • Approved-source indexv5config
  • Checker (MCP tool)v2.4deterministicCat 5
Business capabilitySterile Manufacturing
Business application
Automated Visual Inspection – Line 3 (§12)
GxP = GMPTier 3Annex 22 critical
Application serviceAVI Line 3 – PROD
  • Inspection machinevendor · Cat 4 platform
  • Classifier modelv3.0static ML5 classes
  • Camera/lighting recipev7config
  • Knapp defect kitKK-2026qualified challenge set

Text alternative: Business capability Manufacturing Quality Assurance contains the business application CPV Platform (GxP GMP, criticality High, validated, AI risk tier 3). Its production application service, CPV Signal Service PROD for Sites A, B and C, is where change control applies, and holds these configuration items: SPC/Capability query layer v3.2 (deterministic, GAMP Cat 5, Pattern B core); MSPC model v1.4 (static ML, Annex 22 in scope, QA-approved in registry); Narrative agent v2.3 (LLM, non-critical, human in the loop, Pattern A) with nested items foundation model endpoint scoped-small pinned 2026-07, prompt and instruction set sha 9f2c version 12, and tool allowlist of spc_calculator, capability_calculator and kg_lookup; Orchestrator workflow v2.3 (deterministic control flow); Deterministic checker ruleset v1.7; Locked test set cpv-test-2026Q2 (validation asset, access-controlled); Monitoring config v1.2 (alert and action limits are configuration too). It depends on data product batch_phase_summary v4.1, the PI historian and MES event frames. Two further application services, VAL/QA (where IQ and OQ execute) and DEV/Lab (no GxP data rules, no validation state), sit alongside. Business application Clinical Document Authoring (UC-D4; GxP GCP; Tier 2; Pattern A) has the CSR Drafting Service PROD with a drafting agent v1.9 (LLM) over a pinned vendor LLM endpoint, a prompt set and an approved-source index v5, plus a deterministic checker MCP tool v2.4 (Cat 5). Business application Automated Visual Inspection Line 3 (GxP GMP; Tier 3; Annex 22 critical) has AVI Line 3 PROD with the vendor inspection machine (Cat 4 platform), classifier model v3.0 (static, 5 classes), camera and lighting recipe v7, and Knapp defect kit KK-2026 as the qualified challenge set.

Why the granularity matters

Change control, impact analysis and validation scope all act on CIs. If the model, prompt and index are hidden inside one "application" CI, every change looks either trivial ("config tweak") or total ("revalidate everything"). Modelling them separately makes change-level risk-based validation possible, which is the whole point of CSA. The CI list is the §8.3 configuration baseline, seen from the ITSM side.

Attributes to add to each AI CI (mirrored from the agent card, parent §6):

GxP impactposture (Pattern A/B; Annex 22 in or out of scope)risk tierversioning rule (major/minor/patch)linked T3/T6/T7/T9 document IDslocked-test-set referencemonitoring KPI set (T8)last periodic review (T10)vendor model-deprecation date

14.2 Change types: the T9 envelope becomes ServiceNow change models

ServiceNow change typeWhen it applies to an AI CIValidation scopeApprovals
Standard pre-approved templateChanges inside the T9 envelope: a vendor model patch with an unchanged API; retraining on new examples of existing classes; a prompt edit in the T9 catalogue; a monitoring threshold tuned within a pre-set rangeThe T9 pre-approved regression on the locked test set; automated evidence attached to the changePre-approved by QA when the template was created. Execution evidence reviewed by the system owner; QA is notified or samples
Normal outside the envelopeA new output class; a new site or population (new subgroup); a new intended use; a model architecture change; an out-of-envelope vendor model swap; a new tool added to an agent; regression failure on a Standard changeRisk-based IQ/OQ/PQ scope from an updated T2; T3 updated if the context of use changesChange Advisory Board (CAB) with Quality as a mandatory approver for GxP CIs; business process owner; security if identity or tools change
Emergency contain firstHarmful behaviour in production: an action limit breached on a critical KPI, a prompt-injection incident, a vendor outageFirst contain: kill switch, roll back to the last validated version, or drop to advisory or manual mode (parent G7 degraded mode). Then validate after the factEmergency CAB; QA retrospective approval; a deviation record is opened
Automation hook

CI/CD pipelines raise the change record automatically (e.g., DevOps change automation), attach the IQ evidence and the regression report, and use the risk rating to route Standard vs. Normal. A failed regression gate blocks promotion; nobody has to spot it.

14.3 IQ, OQ, PQ and release for use, redefined for AI components

StageClassic meaningFor an AI componentEvidence (mostly automated)Human accountable
IQInstalled as specifiedThe model artifact hash matches the registry; container image and IaC match; the prompt/config version, tool allowlist, retrieval index version and pinned vendor model version are verified (§4.4); the audit-trail and e-signature configuration is presentPipeline-generated IQ report linked to the change recordPlatform/IT owner executes; QA reviews
OQOperates per specification across rangesPerformance on the locked, independent test set by subgroup against pre-approved criteria (T6). Also: confidence gating and "undecided" routing, explainability output (Annex 22 §8), trajectory and red-team tests (agents), Part 11 checks (the AI cannot sign; AI contribution is attributed; §15), degraded-mode behaviourTest report with 95% confidence intervals; the log of access to the locked test setValidation lead executes; the SME owns the criteria; QA approves
PQPerforms in the real processShadow or parallel run in production conditions: real inputs, real users, a defined period or batch count. Compare against the human baseline (§5). Measure human-in-the-loop effectiveness (seeded-error catch rate) and confirm the monitoring KPIs are live and producing data (§13)PQ report; agreement statistics vs. the manual process; reviewer catch-rate resultsBusiness process owner executes; QA approves
Release for useValidated system handed to productionVSR (T7) approved with a scoped release (sites, products, TAs, document types); registry state candidate → factory; T8 monitoring and T9 addendum effective from day one; users trained, SOP updatedCR moves to Implement; CI validation status becomes Validated vX; release notesSystem owner and QA sign the VSR (Part 11 e-signatures)
Post-implementation reviewDid the change achieve its aim?Check the first 30–90 days of T8 KPIs against validated performance before closing the changePIR task on the change recordSystem owner

Scaling IQ/OQ/PQ to the change: a Standard change usually needs IQ plus targeted OQ (the regression) with no PQ. A Normal change for a new subgroup needs IQ, OQ on the new and all existing subgroups, and a short PQ on the new subgroup. A new critical class needs the full sequence.

14.4 Ongoing monitoring: how production signals flow back into change control

Monitoring to event to incident to problem to change loop A closed loop. A T8 KPI breach at the alert limit raises a ServiceNow Event, which becomes an Incident (investigate; no configuration change). An action-limit breach, a drift score, an out-of-distribution rate or an override-rate anomaly escalates to a Problem record whose root cause is classified as drift, data, process or model. A vendor model-deprecation notice becomes a Problem or a planned Change. A periodic review decision from T10 becomes a Change Request directly. The Change Request is Standard if it is inside the T9 envelope, Normal otherwise, and Emergency if containment is needed. The change runs through the validation scope defined in 14.3, then release, then post-implementation review, after which the T8 baseline is updated and monitoring continues from the new baseline. T8 monitoringKPIs · drift · overrides · cost EVENTServiceNow INCIDENTinvestigate; no config change PROBLEMroot cause: drift? data? process? model? alert limit escalate Action-limit breach · drift scoreOOD rate · override-rate anomaly Vendor model-deprecation notice→ Problem, or a planned Change Periodic review decision (T10)continue · retrain · restrict · retire CHANGE REQUESTStandard if inside T9 · Normal otherwiseEmergency if containment is needed Validation scopeIQ / OQ / PQ sized to the change (14.3) Release → PIRpost-implementation review T8 baseline updatednew validated performance monitoring continues from the new baseline Every arrow into CHANGE REQUEST becomes a ServiceNow record. The monitoring configuration itself is a configuration item.
The loop. Alert limits raise Events and Incidents; action limits, drift and override anomalies raise Problems; vendor notices and periodic-review decisions raise Changes directly. Every change runs through a validation scope sized to it, is released, reviewed, and resets the monitoring baseline.

Three rules keep this defensible:

  1. Monitoring configuration is a CI. Changing an alert or action limit is a change, even though no model changed (§13.3).
  2. Not every signal is a model problem. A higher reject rate on the camera is often a process signal, so it goes to a deviation, not to a model change (§12 prior-shift row; §11 FAQ 4). The Problem record forces that triage.
  3. Periodic review (T10) is a scheduled ServiceNow task on the Business Application, and its decision (continue, retrain, restrict, retire) is recorded and becomes a change where needed.

This is the same chain as the §13.4 response ladder: Level 0 is the Incident, Level 1 the Problem, Level 2 the Change Request.

14.5 Four scenarios run through the whole chain

S1

CPV signal agent: extend to Site D

Normal change
Trigger
Business decision to extend CPV to Site D
CIs affected
MSPC model (new NOC/subgroup), monitoring config, data-product dependency
Change type
Normal: new subgroup, outside the envelope
AI co-brain work
Traverses the CMDB for dependencies; drafts T2 and T3 updates; proposes a Site D test-set stratification
IQNew site connectors, NOC reference, model v1.5 hash
OQLocked test set, all sites, plus a new Site D subgroup
PQShadow on 30 Site D batches vs. manual CPV review
ReleaseVSR addendum; scope now A–D
Monitoring changeSite D KPIs added; drift baseline set
Human accountable
MS&T owner (criteria); QA (approval)
S2

CSR drafting agent: the vendor retires the LLM version

Standard if the regression passes · Normal if it fails
Trigger
Vendor deprecation notice, 90 days out
CIs affected
Vendor LLM endpoint CI, possibly the prompt set
Change type
Standard if the T9 regression passes; Normal if it fails
AI co-brain work
Finds every service using the model; drafts the CR; runs the pre-approved regression; summarises the diffs
IQNew endpoint and version pinned
OQFull T6 regression, N runs per case; 0 unsupported claims
PQNot required (Standard), or short parallel run (Normal)
ReleaseChange closed; registry version bump (minor)
Monitoring changeWatch edit distance and unsupported-claim rate for 30 days
Human accountable
Clinical system owner; QA reviews the evidence
S3

AVI camera: a new critical defect class

Normal change · major version
Trigger
Stopper-skirt deformation seen after a stopper supplier change
CIs affected
Classifier model (output space), defect kit, T3
Change type
Normal: new critical class, major version
AI co-brain work
Drafts the T3 intended-use update and new-class criteria options (ranges only); clusters candidate images for SME labelling
IQModel v4.0 hash; recipe unchanged
OQFull retest of all classes, explainability review, Knapp comparison for the new class
PQKnapp re-challenge plus a period of AQL re-inspection of accepted units
ReleaseVSR v4; scope unchanged except the new class
Monitoring changeNew-class sensitivity KPI; undecided-rate baseline reset
Human accountable
QA and production (criteria, Knapp); QA approves
S4

CPV narrative agent: a prompt edit

Standard change
Trigger
Reviewers keep rewording one phrase (T8 edit-distance signal)
CIs affected
Prompt set CI
Change type
Standard: inside the T9 prompt catalogue
AI co-brain work
Proposes a prompt diff with before/after outputs on the locked set
IQNew prompt sha
OQRegression on the locked set; no metric decreases
PQNot required
ReleaseChange closed; patch version
Monitoring changeEdit-distance KPI watched
Human accountable
Agent owner; QA samples
S1: CPV, add Site DS2: CSR drafting, vendor retires the LLM versionS3: AVI camera, new critical defect classS4: CPV narrative prompt edit
TriggerBusiness decision to extend CPV to Site DVendor deprecation notice, 90 days outStopper-skirt deformation seen after a stopper supplier changeReviewers keep rewording one phrase (T8 edit-distance signal)
CIs affectedMSPC model (new NOC/subgroup), monitoring config, data-product dependencyVendor LLM endpoint CI, possibly the prompt setClassifier model (output space), defect kit, T3Prompt set CI
Change typeNormal: new subgroup, outside the envelopeStandard if the T9 regression passes; Normal if it failsNormal: new critical class, major versionStandard: inside the T9 prompt catalogue
AI co-brain workTraverses the CMDB for dependencies; drafts T2 and T3 updates; proposes a Site D test-set stratificationFinds every service using the model; drafts the CR; runs the pre-approved regression; summarises the diffsDrafts the T3 intended-use update and new-class criteria options (ranges only); clusters candidate images for SME labellingProposes a prompt diff with before/after outputs on the locked set
IQNew site connectors, NOC reference, model v1.5 hashNew endpoint and version pinnedModel v4.0 hash; recipe unchangedNew prompt sha
OQLocked test set, all sites, plus a new Site D subgroupFull T6 regression, N runs per case; 0 unsupported claimsFull retest of all classes, explainability review, Knapp comparison for the new classRegression on the locked set; no metric decreases
PQShadow on 30 Site D batches vs. manual CPV reviewNot required (Standard), or short parallel run (Normal)Knapp re-challenge plus a period of AQL re-inspection of accepted unitsNot required
ReleaseVSR addendum; scope now A–DChange closed; registry version bump (minor)VSR v4; scope unchanged except the new classChange closed; patch version
Monitoring changeSite D KPIs added; drift baseline setWatch edit distance and unsupported-claim rate for 30 daysNew-class sensitivity KPI; undecided-rate baseline resetEdit-distance KPI watched
Human accountableMS&T owner (criteria); QA (approval)Clinical system owner; QA reviews the evidenceQA and production (criteria, Knapp); QA approvesAgent owner; QA samples
Show me

In Doscierge, every rule change must pass regression against the locked evaluation set before it is merged (parent §17). That is the T9 envelope (§14.2) in engineering form. A rule edit that passes is a Standard change whose evidence is generated automatically. One that fails, or a model-version change, becomes a Normal change that re-runs the full set (§1.3, stage 8).

14.6 AI as a co-brain for CSV/CSA: what it does, what the human owns

The same discipline applies to the helper as to any AI (§10). The co-brain is a Tier 1–2 tool that produces drafts and analysis. It does not produce decisions. Its outputs become records only once a human approves them.

Full tableCSV/CSA activities: what the AI co-brain does, what the human owns and signs, and the guardrail on the helper11 rows
CSV/CSA activityWhat the AI co-brain doesWhat the human does and is accountable forGuardrail on the helper
GxP assessment (T1)Pre-fills from the CMDB and vendor documents; proposes classification with rationaleSystem owner and QA decide the classificationIts classification is a suggestion, shown with its evidence
Risk assessment (T2)Proposes failure modes from the library and past incidents; drafts FMEA rowsSME and QA set severity, probability and detectability ratings and accept the residual riskCannot rate or accept risk; every row needs a human owner
Requirements / context of use (T3)Drafts testable requirements from the SOP and process maps; flags untestable wordingThe process SME owns intended use and the sample space (Annex 22 §3.1)Cannot define acceptance criteria (Annex 22 §4.2)
Change impact analysisTraverses CMDB dependencies and the registry "consumers" list; lists affected services, documents and SOPs; proposes Standard or NormalThe change owner confirms the scope; QA approves the classificationIts impact list is verified against the CMDB query result, not taken on trust
Test design (T6)Generates test scripts and edge cases; proposes stratificationThe validation lead approves the protocol before executionNo AI-generated test data or labels for the locked set (Annex 22 §5.6)
Test executionRuns automated suites; collects evidenceThe tester attests executionThe tool is itself assured for this use (CSA: tools that support assurance)
Evidence reviewChecks completeness; flags anomalies; checks traceabilityThe QA reviewer judges pass or failThe AI that authored a script cannot be the only reviewer of its results (independence)
Deviation triageClassifies test deviations; drafts the investigationThe validation lead and QA disposition themIts classification is advisory
VSR drafting (T7)Drafts the report from the evidence, citing each itemSystem owner and QA sign (Part 11)Every statement is traced to evidence; numbers checked deterministically
Monitoring and periodic review (T8/T10)Summarises trends; drafts the review; suggests a decisionQA and system owner decide: continue, retrain, restrict or retireThe suggested decision is labelled as such
Regulatory watchTracks guidance changes (Annex 22 final, Annex 11) and maps them to affected CIs and SOPsQA and regulatory decide the impactCitations checked against the source text

AI proposes, humans dispose, and the change record shows which was which.

The accountability rule in one line. Every co-brain output attached to a change record is tagged as AI-generated with its tool version, and the human decision on it is captured. This is the same Part 11 attribution pattern as §15, applied to the validation process itself.

14.7 RACI for an AI change

Full tableRACI for an AI change: system owner, business process owner or SME, QA/CSV, ML/platform engineering, CAB, and the AI co-brain10 rows
ActivitySystem OwnerBusiness Process Owner / SMEQA / CSVML / Platform EngineeringCABAI co-brain
Raise and classify the changeACCRIDrafts
Impact analysisACCRIDrafts
Update T2 / T3AR (criteria, sample space)CCDrafts
Approve Standard / Normal classificationCCAII
IQACRCollects evidence
OQ (incl. locked test set)AC (criteria owner)A (approval)RRuns suites
PQCRACSummarises
Release for use (VSR)ACA (co-sign)IIDrafts VSR
CAB approval (Normal)RCC (mandatory for GxP)CA
Monitoring response and periodic reviewACCRISummarises

R = responsible, A = accountable, C = consulted, I = informed.

14.8 How it all maps together

How registry, CMDB, templates, change request, validation chain and monitoring map together Four layers. Systems of record: the registry and agent card sync with the CMDB configuration items of AI classes. Design-time: templates T1 to T5 (intake, risk, context of use, data, design) feed from the registry; the change request (Standard, Normal or Emergency) feeds from the CMDB. Execution chain: T6 protocol, then IQ, OQ, PQ, T7 validation summary report, release for use, and post-implementation review; the change request governs release for use. Operation: T8 monitoring raises an Event, then an Incident or Problem, which loops back as a new change request; the T9 envelope is the set of Standard change templates; the T10 review is a scheduled task on the business application whose decision becomes a change request. Across every layer, the AI co-brain drafts at every arrow and humans approve at every gate. SYSTEMS OF RECORD DESIGN-TIME EXECUTION CHAIN OPERATION Registry / agent cardidentity · intended use · risk tier · versions · consumers CMDB CI (AI classes)model · prompt · index · test set · tool server sync T1 · T2 · T3 · T4 · T5design-time: intake, risk, context of use, data, design Change RequestStandard · Normal · Emergency T6 protocol IQ OQ PQ T7 VSR Releasefor use PIR T8 monitoringKPIs · drift · overrides Event Incident / Problemloops back as a new CR T9 envelope= the Standard change templates T10 reviewscheduled task → decision → CR AI co-brain drafts at every arrow · humans approve at every gate
Four layers, one loop. The registry and the CMDB stay in sync; design-time templates and the change request feed the execution chain; operation feeds back as new change requests. The T9 envelope is what makes a change Standard; the T10 review is the scheduled task that decides continue, retrain, restrict or retire.

What to implement first (small or mid-size company)

  1. Add AI CI classes, or an "AI component" related list, under existing Business Applications.
  2. Turn the T9 catalogue into Standard change templates.
  3. Make QA a mandatory approver for Normal changes on GxP AI CIs.
  4. Route T8 action-limit breaches into Incident and Problem records.
  5. Schedule T10 as a recurring task.

Everything else, including pipeline automation and the co-brain, can follow.

So what

The claim now survives change: every alteration is either pre-approved or re-proven in proportion. What remains is the record of all of it. Part E asks what Part 11 and ALCOA+ require when a model, not only a person, contributed to the record.

SpineABCDEF

Part E · Records and trust

Claim component Records

Part 11 and ALCOA+ when a model contributes to a record.

Bridge. Part D kept the claim true in operation: drift detected, changes controlled, every signal routed. Trust in GxP, though, rests on records: what was decided, by whom, on what evidence. Part E asks what Part 11 (§15) and the data-integrity guidance behind ALCOA+ (§16) require when a model contributed to the record. The answer in both cases is the same: the regulation does not change, but what each control must capture does.

§15 · Part E · Audience question Q9 (after the session)

No product "meets Part 11". What changes is what each control must capture.

Audience question Q9 · after the session"Is there any AI software product on the market that would meet 21 CFR Part 11 compliance? Probably not, so how would the guidance differ from the most common digital product to an AI product?"
Spine

This section makes the claim recordable: the AI output, its version and evidence, and the human's decision on it, each captured as its own attributed event, so that any past decision can be reconstructed.

Short answer: the question contains a category error

No software product "meets Part 11", whether it uses AI or not. Part 11 applies to the regulated company's records and signatures, and compliance comes from three things together: the product's technical controls, how the company configures it, and the company's procedures (training, accountability, record retention). The same holds for classic systems: no vendor can make a product "validated" or "Part 11 compliant" on the customer's behalf. What a vendor can offer is a product that is Part 11-capable: audit trail, access control, e-signatures, record copies, and retention.

That leaves two real questions:

1

Are there AI-enabled products with Part 11-capable controls?

Yes. Validated eQMS, document management, LIMS and EDC platforms increasingly ship AI features inside a record layer that already has an audit trail and e-signatures. The open question is not the platform. It is whether the AI feature's contribution is captured by those controls. Often it is not: the audit trail records that the user saved the record, not that 70% of the text came from a model.

2

Can a general-purpose chatbot be used for GxP records?

Not as it ships. Consumer and general enterprise chat tools usually lack record-level audit trails, record retention under your control, and signature binding. They can be used in Part 11 scope only inside a wrapper that supplies those controls (the parent article's platform contract, UC-H1, §8.8), or kept outside record creation entirely (§10, Tier 1).

Part 11 does not change for AI

The regulation is technology-neutral. What changes is what each control has to capture, because an AI system adds a new actor (the model), a new input (the prompt and retrieved context), and output that cannot be reproduced on demand.

Clause by clause: what AI adds

Part 11 controlTypical digital productWhat AI addsPractical control
§11.10(a) Validation (accuracy, reliability, consistent intended performance, ability to discern invalid or altered records)Functional testing against specifications"Consistent intended performance" becomes a statistical claim for a context of use, and it can decay without any code changePerformance claim with acceptance criteria per subgroup (T3, T6), plus monitoring (T8; §13). You validate the controls around the model for this use, not the model in the abstract
§11.10(b) Accurate and complete copiesRe-render or export the stored recordA probabilistic model cannot regenerate the same output later, and the vendor model version may be retiredStore the output as generated, together with the prompt/input, retrieved sources, model and version, configuration version, and confidence. Never rely on regenerating it (§16, "Original")
§11.10(c) Protection and retrieval over the retention periodDatabase backup and archiveThe artifacts needed to understand the record (model version, prompt template, retrieval index snapshot) sit with a vendor who may deprecate themKeep AI context artifacts in a store you control. Add contract terms for model-version notice and data export (§17)
§11.10(d) Limiting system accessUser accounts and rolesThe agent acts as a principal. It can read and write through tools, often with broad service credentialsGive each agent its own identity, a tool allowlist, and least-privilege scopes. No shared or human credentials for agents (parent G5)
§11.10(e) Audit trail (secure, computer-generated, time-stamped record of operator entries and actions that create, modify or delete records)Who changed what, when, old and new valuesThe model is a new actor that is not an "operator", so the trail must show the AI's contribution separately from the human'sRecord AI-generated content as a distinct event attributed to the agent identity and model version. Then record the human action on it: accept, modify (with the diff) or reject, with a rationale (parent G2; §7 "Audit trail"). ALCOA+ "Attributable" now means attributable to a human or to an identified AI version (§16)
§11.10(f) Operational system checks (permitted sequencing of steps)Workflow engine enforces the orderAn agent that plans its own steps can skip or reorder themDeterministic orchestration for regulated sequences. The model proposes, and the workflow enforces the order (T5; §6 Step 2)
§11.10(g) Authority checksRole-based permissionsAn agent may try to take an action that needs authority: approve, release, close a CAPAAgents never hold authority for signature-bearing or disposition actions. Those routes always go to a qualified human (confidence gating, parent G7)
§11.10(h) Device checks (validity of the source of data input)Instrument or terminal identityAI output is itself a source of data entering the recordTag every AI-originated field with its source (agent ID, model version) so reviewers and inspectors can tell machine-originated data from human- or instrument-originated data
§11.10(i) TrainingSystem-use trainingReviewers must understand AI failure modes, especially automation biasTraining in how to review AI output, plus seeded-error checks of reviewers (§5, §7)
§11.10(j) Accountability for e-signaturesSignature policy"The AI wrote it" is not a defenceThe policy states that the signer is accountable for AI-assisted content they sign, exactly as for content they wrote
§11.10(k) Systems documentation controlSOPs, configuration specificationsPrompts, system instructions, tool definitions, retrieval sources and thresholds are system documentationPut them under version and change control as configuration items (T5, T9; §8.3, §14.1)
§11.50 / §11.70 Signature manifestation and linkingName, date/time, meaning; signature bound to the recordThe meaning of a signature on AI-assisted content should be explicitUse a signature meaning such as "Reviewed and approved, including AI-assisted content", and bind the signature to the exact output version reviewed
§11.100–§11.300 Electronic signaturesUnique to one individual; identification componentsAn AI can never signHard technical block: no agent identity can apply an e-signature. Test this in OQ (§14.3)

Scope: when is AI output a Part 11 record?

Use the same test FDA's 2003 Part 11 Scope and Application guidance applies to any electronic record. It is in scope when a predicate rule requires the record, or when you rely on it to carry out a regulated activity or decision.

AI outputPart 11 record?Why
Brainstorming a draft SOP structure that a human then rewritesUsually noThe approved SOP is the record. The AI is an authoring aid (§10, Tier 1)
AI-drafted deviation investigation text kept in the final reportYes, as part of that recordThe report is a GMP record. The AI contribution must be attributable (§11.10(e))
AI classification that routes a deviation (e.g., minor vs. major)YesIt influences a regulated decision, so the output, confidence, model version and human confirmation are all retained
AVI camera accept/reject per unit (§12)YesIt is a batch-record decision, Annex 22 critical, with Annex 11 and Part 11 controls on the result
CSR section drafted by UC-D4YesIt is a submission document. The drafting history, checker findings and sign-off form the audit record
Chat transcripts used to decide something regulatedYes, if relied onReliance makes it a record. Keep it in a controlled store, or do not use chat for regulated decisions

How to evaluate an AI product for Part 11 readiness (vendor questions)

  1. Does the audit trail record AI-generated content as a separate event, attributed to a model or version, with the human's accept, modify or reject captured afterwards?
  2. Is the as-generated output stored alongside the prompt/input, retrieved sources and configuration version, and can it be exported for the full retention period?
  3. Which model version produced each output, how much notice do you get before a model changes, and can you pin versions (§4.4)?
  4. Can the AI ever apply an e-signature, approve, release or close a record? The answer must be no.
  5. Do agents have their own identities with least-privilege access, or do they run on user or service credentials?
  6. Is customer data used to train shared models? Where are prompts and outputs stored, and for how long? Does the vendor's retention setting conflict with yours? A "zero data retention" option on the model API is fine only if your own record store keeps the Part 11 copy.
  7. Can AI features be switched off per tenant and per module until you have assessed them (§17)?

Worked contrast

Common digital product

An eQMS deviation module without AI

Validate the workflow, audit trail, e-signatures and access. The investigator writes the text. Part 11 evidence covers who entered what, when, and who signed.

The same module with AI root-cause suggestions (UC-M5 pattern)

Everything above, plus

  • the suggestion is stored as generated, with model version and evidence citations
  • the investigator's accept, modify or reject and rationale are in the audit trail
  • AI-originated text is tagged in the final report
  • the signature meaning covers AI-assisted content
  • the AI cannot close the deviation or approve the CAPA
  • T8 monitoring tracks the acceptance rate and seeded-error catch rate (§13)
  • a vendor model change triggers the T9 change protocol (§14.5, S2)
Bottom line for the questioner

Do not look for a "Part 11-compliant AI product". Look for a Part 11-capable platform whose AI features are attributable, retained as generated, version-pinned, barred from signing, and switchable. Then close the remaining gaps with your own procedures and the T1–T10 controls.

So what

Part 11 says what a record and a signature must be. The data-integrity guidance says what the data inside the record must be. §16 applies ALCOA+ principle by principle to a record that a model helped produce.

§16 · Part E · Practitioner question Q12

ALCOA+ for an AI-assisted record: attributable to a human or to an identified model version

Practitioner question Q12 · not from the session"What does ALCOA+ actually require of a record that AI helped produce, and what do we have to capture?"
Spine

This section makes the claim trustworthy as data. Every ALCOA+ principle still applies; the AI adds a second author whose contribution, version and sources must be captured as metadata of the record, and adds new ways for the data to fail.

16.1 The governing guidance, and what each adds

None of these documents mentions AI. All of them define the properties a GxP record must have, and those properties are what the AI-assisted record must still show. Titles and dates verified against the issuing bodies (§19).

Full tableGoverning data-integrity guidance, when it was issued, and what it contributes to the AI case5 rows
GuidanceIssuedWhat it contributes to the AI case
MHRA, 'GXP' Data Integrity Guidance and Definitions (Revision 1)March 2018Definitions used across GxP: data, raw data, metadata, audit trail, data lifecycle, data governance; ALCOA+ as the property set; expectations for hybrid systems and for validation of systems that generate records
PIC/S PI 041-1, Good Practices for Data Management and Integrity in Regulated GMP/GDP EnvironmentsIn force 1 July 2021The inspector's view: data governance, data criticality and inherent integrity risk, audit-trail review, outsourced activities and computerised systems; the reference for GMP/GDP inspections in PIC/S countries
FDA, Data Integrity and Compliance With Drug CGMP: Questions and AnswersDecember 2018 (final)The CGMP anchor in the US: metadata defined as the contextual information required to understand data; audit trails as metadata; audit-trail review expectations; shared logins and system controls
WHO, Technical Report Series 1033, Annex 4, Guideline on data integrity2021Data governance as a senior-management responsibility embedded in the quality system; ALCOA+ applied across the data lifecycle; verification of the effectiveness of data-integrity controls
EU GMP Chapter 4 (Documentation) and Annex 11 (Computerised Systems), revision draftsConsultation 7 July–7 October 2025, alongside Annex 22The EU direction of travel: an expanded Annex 11 addressing, among other topics, audit trails, supplier oversight and identity and access management; final texts may change what audit-trail review and electronic records require. Revisit this section when they are published

16.2 Principle by principle: what it means for an AI-assisted record, and what to capture

A

Attributable

Every element of the record is traceable to who or what produced it. The model is a second author, and its contribution must be distinguishable from the human's.
What to captureAgent or system identity; model name and pinned version; prompt or configuration version; the reviewer's identity; the accept / modify / reject action with the reviewer's rationale; the time of each event
L

Legible

The record, and the AI's contribution to it, can be read and understood for the retention period, including why the AI produced what it did.
What to captureThe output as generated, in a durable format; the confidence or score; the explainability artifact where one exists (contribution plot, feature attribution, cited spans); no reliance on a vendor UI to render it
C

Contemporaneous

AI events are time-stamped when they happen, not reconstructed later; the human review is time-stamped separately from the generation.
What to captureSystem time-stamp of generation; time-stamp of each retrieval; time-stamp of the reviewer's action; the elapsed time between them (also a monitoring signal, §13.1 family 3)
O

Original

The first capture is the AI output as generated. A probabilistic model cannot regenerate it, so a regenerated version is a new record, not a copy.
What to captureThe as-generated output stored before any edit; the AI draft and the human-edited final, with the diff between them; the retrieved sources and the spans used; never overwrite the draft with the final
A

Accurate

The content is correct and complete relative to its sources, and the checks that established that are recorded.
What to captureDeterministic checker findings (numeric consistency, required elements, cross-references) and their resolution; the reference the output was checked against (TFL version, historian tags, source SOP version); confidence and any "undecided" outcome
+C

Complete

Nothing is missing: rejected outputs, overrides, undecided cases and re-runs are part of the record, as invalidated laboratory results are (§8.4).
What to captureAll runs for the case, not only the accepted one; rejected and overridden outputs with reasons; the checker's full finding list, not the summary
+C

Consistent

The sequence of events is coherent and time-ordered across generation, checking, review and signature; the same case is not represented differently in two systems.
What to captureOrdered event chain in one audit trail, or reconciled across systems; consistent case identifiers across the AI service, the eQMS and the signing system
+E

Enduring

The record and the metadata needed to understand it survive for the retention period, including the model version, prompt template and retrieval index snapshot that a vendor may retire.
What to captureRetention of AI context artifacts in a store the company controls (§15, §11.10(c) row); archived alongside the successor system at retirement (§8.8)
+A

Available

The record and its AI metadata can be retrieved and reviewed on request, by the company and by an inspector, in a readable form.
What to captureExport of the AI event chain in a readable format (a vendor clause, §17 element 4); a query that returns, for any signed record, which AI version contributed and what the reviewer changed

16.3 The AI audit trail as GxP metadata

FDA's 2018 Q&A defines metadata as the contextual information required to understand data, and treats audit trails as a form of metadata. On that definition the model version, prompt and configuration version, retrieved sources, confidence and the reviewer's action are metadata of every AI-assisted record. Three consequences follow.

  • Retention. The AI metadata is retained for the same period as the record it explains. A vendor's "zero data retention" setting on the model API is compatible with this only if your own record store keeps the copy (§15, vendor question 6).
  • Audit-trail review. The guidance expects audit trails for changes to critical data to be reviewed with the record, before final approval, and to be reviewed periodically across records for patterns. For an AI-assisted record that means two reviews: the reviewer sees the AI event chain for this record before signing (what the model produced, what was changed, why), and QA reviews AI audit trails across records periodically for patterns such as edits that always remove the same kind of claim, or acceptances made in seconds (§13.1, families 3 and 5).
  • Hybrid records. An AI draft produced in one tool, edited in a second and signed in a third is a hybrid record. MHRA 2018 discourages hybrid arrangements and expects, where they exist, a documented definition of what constitutes the complete record and controls that keep the parts linked. For AI, the complete record includes the as-generated draft, the sources, the checker findings and the signed final, with one case identifier across the tools.

16.4 Data governance for training, validation and test data

Training data, labels and the locked test set are GxP data in their own right (§8.4, T4), and ALCOA+ applies to them: labels are attributable to the adjudicators, original as source-linked, accurate as adjudicated, and complete including exclusions and their reasons. Data governance for AI adds three expectations: criticality assessment of the datasets (a label error on the critical class is a critical data-integrity error), access control and audit trail on the test set so that independence can be demonstrated (Annex 22 §6), and lineage from every data point to its source system (§8.4, UC-M2 example). WHO TRS 1033's expectation that senior management owns data governance applies here as much as to batch records.

16.5 Common data-integrity failure modes with AI

Failure modeWhy it is a DI failureDetectControl
The AI draft is overwritten by the edited finalOriginal lost; the human contribution cannot be separated from the model'sRecords with a final but no as-generated version (§13.1 family 5)Store drafts as immutable events; edit creates a new version
Unattributed AI text in a signed recordNot attributable; the signer may not know what they are signingShare of AI-originated fields without a source tagTag AI-originated content at field level (§15, §11.10(h) row); signature meaning covers AI-assisted content
Retrieval over unapproved sourcesNot accurate; content traceable to a source that has no controlled statusCitations that resolve outside the approved-source registryGrounding restricted to the approved-source registry; index version as a CI (§8.3)
Test-data leakage into trainingValidation evidence is not accurate; the claim is overstatedAccess logs on the locked test set; overlap checksAccess control and staff independence (Annex 22 §6); T4
Labels harvested from production acceptancesNot accurate: the labels inherit the reviewers' automation biasRetraining records whose labels lack adjudicationHuman adjudication of every label used for training or test (§11, FAQ 12)
Silent vendor model changeNot consistent or enduring: the record's stated model version no longer describes what produced itProbe-set and version checks (§4.4)Pinned versions; contract notice; unauthorised-change investigation
Shared credentials for agentsNot attributable; the audit trail names a service account for every actorAgent actions logged under a human or shared identityOne identity per agent (§15, §11.10(d) row)
Conversation memory carrying state across casesNot consistent; the same input gives different outputs depending on historyReproducibility tests failing; outputs citing content from another caseFresh context per case (§3, row 9)
Suppressed or unlogged "undecided" outcomesNot complete; the cases the model could not handle disappearUndecided rate reported by the model but absent from the record storeUndecided outcomes stored and routed like any other output (§8.5, Annex 22 §9)
So what

With Part 11 and ALCOA+ satisfied, the claim is defined, proved, kept true and recorded. Part F asks what all of that looks like for a company with a 3–10 person Quality team, and what the market still fails to supply.

SpineABCDEF

Part F · Scale it

Claim component Scale-down

The same claim discipline sized for a 3–10 person Quality team, and what the market still lacks.

Bridge. Parts A to E defined the claim, set its bar, built its controls, kept it true and recorded it, with a CPV agent, a drafting agent and a camera as the running examples and ProtoCheck and Doscierge as the worked systems. Part F scales the same discipline down to a small company that receives AI rather than builds it (§17), turns the page into six moves you can start with, alongside what the market still lacks (§18), and lists the sources and what still needs verifying (§19).

§17 · Part F · Audience question Q7

Small biotech: the AI function, not the product category, carries the risk

Audience question Q7"How can a biotech or a smaller company committed to bringing AI tools monitor the risks associated with this switch from Cat 3 to Cat 5, and act on monitoring compliance and data integrity?"
Spine

This section scales the claim down. A small company makes fewer claims, but each one still needs a defined use, a tier, a monitoring minimum (§13.2) and a record (§1516). The program below is the minimum that keeps every AI function on the site inside that discipline.

First, correct the framing

The risk is not that a company chooses to move from Category 3 to Category 5. For most biotechs, three different things happen at once:

1

Vendor AI arrives inside validated SaaS

The eQMS adds AI deviation triage. The EDC adds AI query suggestions. The document system adds AI summarization. The product stays "Category 3/4", but a function inside it now behaves like trained, probabilistic software. This is the main exposure for a small company, and it usually arrives in a routine release note.

2

Shadow AI

Staff use general-purpose chatbots to draft SOPs, deviation reports, or batch-record summaries. There is no inventory and no data-classification control. This is the GxP spreadsheet problem again. The approach that worked for spreadsheets carries over. Banning them outright rarely succeeded. What worked was an inventory, a risk assessment of the data used and reported, validation effort focused on the high-risk few, and controls validated once as a shared add-on that each use then references. For AI, that shared add-on is the approved enterprise tool with logging.

3

Deliberate builds

A small number of real AI projects, such as a CPV model or a document checker. These are Category 5-like and should be governed as such.

The GAMP categories were built for software whose behavior is set by its code and configuration. For AI you classify the AI function: what decision it influences, whether it is critical, and whether it is static and deterministic (§4, §8.1). A Category 3 product can contain a high-risk AI function.

Minimum viable AI governance program (50–300 employees)

Sized for a company with a 3–10 person Quality team and no dedicated AI group.

ElementWhat it isEffortOwner
1. AI policy + acceptable-use SOPApproved tools; no GxP data in unapproved AI; AI-assisted authoring rules (§10)2–3 pages, 1 weekHead of Quality + IT
2. AI inventoryOne register of every AI function: vendor, built, or shadow. Fields: system, AI function, GxP area, critical Y/N, static/dynamic, deterministic Y/N, owner, risk tier, last review. Build it on the existing GxP system inventory, which already holds risk rating, business criticality, validation date, and retirement date. Add the AI columns to it rather than starting a separate list. This is the small-company form of the registry in §8.6A spreadsheet or eQMS form is enough to startQA (CSV)
3. Risk tiering (three tiers)Tier 1: non-GxP or authoring aid; register only. Tier 2: GxP, non-critical, human-in-the-loop; lightweight assessment plus monitoring. Tier 3: GxP-critical; full T1–T10 package, and only static/deterministic models (Annex 22; §4.5). The tiers set the monitoring minimums in §13.2One pageQA
4. Vendor AI clausesIn quality agreements and supplier questionnaires: notify before enabling AI features; say whether customer data trains shared models; supply a model card and test evidence; commit to model-version change notices; allow AI features to be toggled off per tenant; on exit, return AI audit data (inputs, outputs, model versions, human actions) in a readable format (§16, "Available")Add to the supplier questionnaire template and the quality agreement / SLAQA supplier mgmt + Procurement
5. Release-note triageEvery SaaS release is screened for AI features before it reaches production. If an AI feature is found: inventory entry, tier, and a decision to enable or disable30 min per releaseSystem owner
6. Data-integrity controls for AIALCOA+ applied to AI: log the input, output, model version, confidence, and human action. Make sure the vendor's audit trail captures that AI contributed (§15, §16)Configuration plus supplier askSystem owner + QA
7. Monitoring KPIs (Tier 2–3)Human override/edit rate; error or deviation rate on AI-touched records; drift indicators where available; incident count; the §13.2 minimums for the tierQuarterly reviewSystem owner
8. Periodic reviewAnnual (Tier 2), semi-annual (Tier 3): continue, restrict, retrain, or retire (§8.8)Part of existing periodic reviewQA
9. TrainingAutomation-bias awareness for reviewers; "how to review AI output" micro-training (§5)1 hour per personQA training

Mapping common AI uses to the three tiers

Industry practice ranks AI uses on a ladder, from low risk (drafting, search, summaries) through decision support to GMP-critical decisions. That ladder is a useful start. It becomes defensible once each rung is tied to Annex 22 criticality and to a posture:

AI useTypical industry ratingTierPosture and what it implies
Draft SOPsLow1 (or 2 if the SOP is GxP)Authoring aid; the human-approved document is the record (§10)
Summarize reportsLow–Medium2Pattern A: grounding + human review; measure edit rate
Classify deviationsMedium2Pattern A while a QA investigator confirms the classification. If the class alone drives impact assessment or batch decisions, move it to Tier 3
Risk assessmentsMedium–High2AI proposes failure modes and ratings; an SME owns every rating (Annex 22 §4.2 by analogy). Check coverage against a reference failure-mode library (T2 §3)
Batch release / disposition supportHigh3Critical GMP. Only deterministic logic or a static, validated ML model may carry the decision (Annex 22 §1; §4.5). An LLM may assemble the evidence package for the QP or QA reviewer; it may not recommend release

SaaS, cloud and vendor audit: what carries over, and what AI adds

Established vendor and SaaS qualification practice GAMP 5 2nd Ed.; Annex 11 §3; PIC/S PI 011-3 still applies. The AI-specific additions matter because the vendor's model can change without any change you would see in a normal release (§4.4).

Full tableVendor and SaaS practice that carries over from CSV, what AI adds, and the nuance6 rows
Carries over from CSVWhat AI addsNuance
Vendor documentation may be leveraged, but must be scrutinized, risk-rated and "owned". No vendor can claim a product is "validated" or "Part 11 compliant" slides 162, 164, 615, 626The same applies to "validated AI" or "GxP-ready AI" claims. Ask for the evidence behind the claim, then re-test for your context of useVendor model evaluations are rarely stratified by the subgroups you care about
For cloud and SaaS, IQ becomes a "research effort" drawing on the vendor's published policies and certifications slides 281, 548, 613Model identity, version, and update cadence are rarely published. Obtain them under the quality agreement, and record the model version in your IQA public website is a starting point, not objective evidence of what runs in your tenant
Contract and SLA: advance warning of updates with release notes and time for client testing; notice before service withdrawal so data can be retrieved slides 597–598, 618Add model-change notices (including silent model swaps behind the same feature name), per-tenant opt-out, and the return of AI audit data on exitYour release-note triage (element 5) depends on this clause
Post-audit outcomes: use unconditionally, use for certain products only, use subject to corrective action, or prohibit slide 596Add a fifth outcome: use with the AI feature disabled until it is assessedOften the fastest decision for a small company
Audit the vendor every two years; require SOC 2 slides 547, 589, 617, 636Trigger a for-cause assessment when a vendor introduces or materially changes an AI featureNeither a two-year cycle nor SOC 2 is a GxP regulatory requirement. Set audit frequency by risk (GAMP 5 2nd Ed. supplier management). SOC 2 is an attestation report on security and related trust criteria. It says nothing about GxP fitness or model behavior
Ask where servers are located slide 618Also ask where inference runs and where prompts and outputs are stored or loggedRelevant to data-residency and confidentiality controls

Signals to monitor (the risk dashboard)

SignalWhy it mattersTrigger
AI functions in inventory not yet tieredUngoverned AI in useAny after 30 days
Vendor releases with AI features enabled by defaultSilent category creepAny not triaged
Human acceptance rate of AI output > 98% sustainedPossible automation bias (reviewers may have stopped reviewing)Run a seeded-error check (§5, §13)
Human acceptance rate < 60%The tool isn't fit for purpose, and staff will route around itReview or retire
Deviations or CAPAs whose root cause involves AI outputDirect quality impactAny → CAPA and periodic-review input
Shadow-AI detections (proxy or DLP logs)Data-integrity and confidentiality exposureTrend
Regulatory changes (Annex 22 final, Annex 11 revision, FDA AI guidance final)Criteria may shiftReview the program

90-day starter plan

DaysActionOutput
0–30Issue the AI policy; run an inventory sweep (survey + SaaS admin consoles + release notes)Inventory v1, policy approved
31–60Tier every entry; send the vendor AI questionnaire to the top 5 GxP SaaS suppliers; switch off untriaged AI featuresTiered inventory; supplier responses
61–90Assess Tier 2–3 entries; set up the KPIs; train reviewers; run the first management reviewRisk assessments; KPI baseline; management review minutes
Cost reality (illustrative estimate)

About 0.2–0.4 FTE of an experienced CSV lead for the first quarter, then about 0.1 FTE ongoing, plus targeted external help for any Tier 3 build. Small companies can afford this. What they cannot afford is finding an AI function in a validated system during an inspection.

So what

The same claim discipline fits a small company, because most of its AI arrives as a vendor feature that needs a tier, a switch and a record rather than a model validation. §18 turns the whole page into six moves you can start with.

§18 · Part F · Apply it

Apply it: six moves from this page to your first defensible AI claim

Apply it · from reading to doing

Spine

The page has defined the claim, set its bar, built its controls, kept it true, recorded it and scaled it down. This section turns that into six moves for one AI use. The moves follow the lifecycle in §1.3, and each one names what you produce, where it was shown on ProtoCheck or Doscierge, and the template or section that carries it.

18.1 Six moves, one use

MoveWhat you doWhat you produceShown onTemplate / section
1. Pick one use and write its context of useOne sentence: the input, the output, who decides, and what happens if the output is wrongA context-of-use statementProtoCheck's first row in §1.3T1 · §8.1
2. Classify, tier and split the componentsGxP? Critical? Static or dynamic? Deterministic or probabilistic? Then influence × consequence, and which §11 component carries the decisionA tier and a component splitDoscierge: Category 5 checker, non-critical LLM pathT1, T2 · §4.5 · §8.2
3. Write the claim and measure the baselineMetric, subgroup, threshold, baseline, confidence level. Measure the process the AI replaces before you set the thresholdA claim record and a baseline studyDoscierge's per-domain claim against the unaided reviewer§1 (last row) · §5
4. Lock the test set and the configurationBuild the independent, expert-labelled set before the model is tuned; list every configuration itemA locked set and a configuration-item listDoscierge's evaluation set and knowledge-graph snapshot§8.3 · §8.4
5. Set monitoring limits before go-liveDerive alert and action limits from the validated performance; name who responds, and howA monitoring plan with a response ladderFindings per package and acceptance per ruleT8 · §13
6. Pre-approve the predictable changesWrite the changes you can foresee into an envelope, and link monitoring to change recordsA change envelope and linked recordsRegression before merge§14 · T9 (on request)
How you know it worked, at 90 days

Take the six inspector questions in §8 (What an inspector will ask) and answer each one with a document you can hand over, not a description. If "how did you determine acceptable performance?" is answered by a baseline study and a signed claim record, and "what happens when the AI is wrong?" is answered by a closed monitoring record linked to a change, the claim is defensible. If either answer is a slide, go back to move 3 or move 5.

Start small, on purpose

Pick a Tier 2 use (a human-reviewed drafting or checking agent) for the first pass. It exercises all six moves with a one-page assurance plan (§9) and teaches the team the claim discipline before it is needed for a Tier 3 decision.

18.2 What the demand signal says, and what to build

The questions in §2 are also demand data for the claim discipline this page describes. Read as a market:

The demand signal. One general web event produced eight questions, and seven of them asked for operational how-to. Three more arrived after the session or from practitioners, and all three asked for a policy or a decision table rather than a principle. The supply side is split:

Standards bodies

ISPE and PDA publish frameworks.

Big consultancies

Sell enterprise programs.

eQMS vendors

Ship AI features without governance tooling.

Nobody

Sells a right-sized, template-driven AI validation practice for small and mid-size companies.

The deadline pressure is real. Annex 22 and the Annex 11 revision are targeted for finalization around Q4 2026, and the FDA–EMA Good AI Practice principles came out in January 2026.

Why buyers care: the enforcement context

Inspectors already cite weak audit trails, unvalidated workflows, missing risk assessments and unauthorized design changes, and Quality Unit oversight of data systems is a continuing focus in public FDA warning letters. An AI function added to a validated system without inventory, risk assessment, or an audit trail of its contribution fits those citation patterns. Smaller firms often stay reactive and under-resourced, which is the gap the §17 program is sized for.

OfferingAnswersBuyerFormatNotes
AI Validation Template Pack (T1–T10, plus the §12 decision tree, the §17 program and the §13 policy)Q3, Q4, Q8, Q13CSV / QA leadsKit + walkthrough; T1, T2 and T8 published freeFastest to produce; the rest available on request
Human-baseline measurement studyQ6, Q2MS&T / QA at sites deploying vision or data AIFixed-scope engagementDirectly required by Annex 22 §4.3; almost nobody offers it
90-day AI governance starter for biotechQ7, Q5Biotech Head of QualityFixed-fee programPairs with practitioner training
Training course: "Applying CSV to AI"Q1, Q8, Q11QA / CSV practitionersOnline session or full dayThe session practitioners are asking for
Agent validation reference (dual-path, generate-then-check)Q3, Q4Pharma AI platform teamsArticle + reference implementationExtends the parent article; ProtoCheck and Doscierge as proof
Site AI monitoring policy and data-integrity addendumQ13, Q12, Q9Heads of Quality, system ownersPolicy text (T8 Appendix B) + implementation workshopNew in this version; the piece that makes the eight controls operable across a site
Positioning line

"CSV tells you how to validate software. We show you how to validate a performance claim, and how to keep it valid."

What comes next

  • The template pack (T1–T10) is complete; T1, T2 and T8 are published free with this page. Next: an SME review pass of the templates and a one-page Tier 2 assurance-plan variant.
  • Refine the §17 program sizing, the tier mapping, the §12 decision tree and the §13 minimums with practitioner feedback.
  • An "Applying CSV to AI" practitioner session, using this page as the outline, is being proposed.
  • The parent article's §11 now links here; this page links back.
So what

Six moves make one claim defensible, and the market needs map onto the same claim components: baseline, controls, monitoring, records, scale-down. The last section lists what this page rests on and what still needs verifying.

§19 · Part F · Sources

Sources and confidence notes

Spine

A claim is only as strong as its sources. This section separates what was verified for this version, where the practitioner questions came from, and what remains to verify before relying on it.

Primary sources verified for this version

Full listRegulations, guidance and studies checked for this version11 sources
  • EU GMP Annex 22 Artificial Intelligence, consultation draft, 7 July 2025 (full text reviewed; section numbers above cite it directly): European Commission PDF
  • Annex 22 status: consultation closed 7 Oct 2025 with about 1,300 comments; EMA IWG targets a final text in Q4 2026; EMA workshop held 30 Jun–1 Jul 2026, considering whether dynamic and probabilistic models could be addressed. Epista, Scilife
  • EU GMP Chapter 4 and Annex 11 revision drafts, released for targeted stakeholder consultation together with Annex 22 on 7 July 2025; consultation 7 July–7 October 2025. PIC/S notice, ECA summary
  • FDA draft guidance, Considerations for the Use of AI to Support Regulatory Decision-Making for Drug and Biological Products (7 Jan 2025): seven-step, risk-based credibility framework. Federal Register, FDA PDF
  • FDA final guidance, Computer Software Assurance for Production and Quality System Software (24 Sep 2025). Note: issued by CDRH/CBER for device production and quality-system software. Pharma applies its risk-based principles by analogy; it is not binding on drug GMP systems. Federal Register
  • ISPE GAMP Guide: Artificial Intelligence (July 2025, 290 pp). ISPE
  • FDA–EMA, Guiding Principles of Good AI Practice in Drug Development (14 Jan 2026): ten high-level principles, including a human-centric, risk-based approach, context of use, data governance, and lifecycle monitoring. EMA PDF, EMA news
  • FDA Quality Management System Regulation (QMSR), effective 2 Feb 2026. It incorporates ISO 13485:2016 by reference, which replaces the old 21 CFR 820.70(i) software-validation text (ISO 13485 clause 4.1.6). FDA QMSR
  • 21 CFR Part 11 (§11.10(a)–(k), §11.50, §11.70, §11.100–§11.300) and FDA guidance Part 11, Electronic Records; Electronic Signatures — Scope and Application (Aug 2003), which gives the predicate-rule and reliance test used in §15. eCFR Part 11, FDA guidance
  • Garza MY et al., Error rates of data processing methods in clinical research: a systematic review and meta-analysis, Int J Med Inform 2025;195:105749. PubMed, ScienceDirect
  • USP <1790> and the Knapp–Kushner method: POD ≥ 0.7 defines the reject zone; alternatives to manual inspection must show equivalent or better performance. PDA, ISPE Pharmaceutical Engineering
Full listData-integrity sources (§16) and FDA device sources used by analogy (§4.3, §12)6 sources

Data-integrity sources (§16), titles and dates verified for this version

  • MHRA, 'GXP' Data Integrity Guidance and Definitions, Revision 1, March 2018. MHRA PDF (gov.uk), MHRA Inspectorate blog, 9 Mar 2018
  • PIC/S, Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments, PI 041-1, adopted 1 June 2021, in force 1 July 2021. PIC/S document, PIC/S news
  • FDA, Data Integrity and Compliance With Drug CGMP: Questions and Answers, Guidance for Industry, final, December 2018. FDA guidance page, Federal Register, 13 Dec 2018
  • WHO, Guideline on data integrity, Annex 4 of WHO Technical Report Series 1033 (fifty-fifth report of the Expert Committee on Specifications for Pharmaceutical Preparations), 2021. WHO PDF

FDA device sources used by analogy only (§4.3, §12)

  • FDA, Proposed Regulatory Framework for Modifications to Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD): Discussion Paper and Request for Feedback, April 2019: introduced the "locked" vs. "adaptive" distinction and the Predetermined Change Control Plan concept. FDA AI/ML SaMD page, Discussion paper PDF
  • FDA, Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions, final guidance, 4 December 2024. FDA guidance page, Federal Register, 4 Dec 2024

Origin of the practitioner questions

  • Web event, Acceptance Criteria for AI Validation (23 Sep 2026). The source of the session questions Q1–Q9 in §2.
  • ProtoCheck and Doscierge figures (rule counts, knowledge-graph size, P/R/F1, findings per package before and after precision engineering) are taken from the parent article's §17 and are the author's own systems.

Cited from domain knowledge, to verify before relying on it

Full listItems cited from domain knowledge, to confirm before quoting13 items
Verify before relying on these
  • Pooled per-method error rates in Garza et al. (only the ranges above were verified).
  • MASAI trial specifics (Lång et al., Lancet Oncology 2023): confirm the detection and workload figures before quoting numbers.
  • THERP / NUREG/CR-1278 nominal HEP values: cited only qualitatively here.
  • GAMP 5 2nd Edition Appendix D11 (AI/ML): confirm the appendix reference.
  • ServiceNow terminology (§14): the CSDM hierarchy and the Standard/Normal/Emergency change types are standard ServiceNow constructs. The AI CI classes are illustrative extensions, not out-of-the-box classes. Confirm the current names of ServiceNow's AI-governance and DevOps change-automation features before publication.
  • SOC 2 characterization (§17): an AICPA attestation report on the Trust Services Criteria, not a certification. Confirm the wording against the AICPA source before publication.
  • FDA 2018 Data Integrity Q&A (§16.3): the definition of metadata and the expectation that audit trails capturing changes to critical data are reviewed with each record before final approval are cited from the guidance as recalled; confirm the exact Q&A numbers and wording before quoting.
  • MHRA 2018 on hybrid systems (§16.3): the statement that hybrid arrangements are discouraged and, where used, need a documented definition of the complete record is cited as recalled; confirm the clause and wording.
  • Draft Annex 11 (July 2025) content summary (§16.1): the topics listed (audit trails, supplier oversight, identity and access management) are taken from consultation summaries, not from a review of the draft text. Verify against the draft before citing it in detail.
  • WHO TRS 1033 Annex 4 (§16.1, §16.4): the senior-management data-governance expectation and the appendix on verifying control effectiveness are cited from secondary summaries; confirm against the annex text.
  • PIC/S PI 041-1 (§16.1): described at the level of its scope and structure only; no section numbers are cited.
  • Draft Annex 22 §1 scope wording on non-critical applications (§4.2): the draft's principles are described here as applicable "where applicable" to non-critical uses; confirm the exact wording of the scope clause before quoting.
  • The "so what" statements and Strategic Planning Assumptions are the author's judgements from the sources cited; the probabilities are not survey results.

Confidence notes

Confidence notes
  • Every number labelled illustrative (§6 test sizes, §8 thresholds, §13 limits and SLAs, §17 FTE estimates) is a design example, not a benchmark.
  • Annex 22 is a draft. Its scope on dynamic and probabilistic models may widen in the final text, so revisit §4, §9 and §8.1 when it is published.
  • The Chapter 4 and Annex 11 revisions are drafts. Revisit §16 when the final texts are published.
  • Device guidance (CSA, the 2019 AI/ML discussion paper, the 2024 PCCP guidance) is cited by analogy only and is not binding on drug GMP.
Where the spine ends

Claim defined (A), bar set (B), controls built (C), kept true (D), recorded (E), scaled down (F). The claim is the whole argument: you validate a performance claim for one context of use, then you keep proving it.

Contact · Free templates · Walkthrough

The controls are defined. Three templates are free. The baseline is yours to measure.

This page exists to help Quality, CSV, manufacturing, clinical and platform teams turn CSV principles into AI-specific controls. It is a companion to the Enterprise Agentic AI Platform for Life Sciences architecture, whose governance plane, agent registry, promotion gate, dual-path reasoning and validation strategy it applies to the questions practitioners actually asked.

Three templates, published free

The intake, the risk assessment and the monitoring plan: enough to classify an AI function, rate its risk and set its KPIs.

The full ten-template pack, on request

T3 context of use, T4 data management, T5 design, T6 validation plan, T7 summary report, T9 predetermined changes, T10 periodic review. Message on LinkedIn or scan the contact code.

30 minutes to map one of your AI use-cases to the eight controls

Bring one use-case: a vendor AI feature, a vision system, a drafting agent. We walk it through GxP assessment, risk, version control, data, design, registry, monitoring and periodic review, and you leave with the gaps named.

The honest caveat: Annex 22 is still a draft, the FDA CSA guidance is device guidance applied by analogy, and the worked-example values on this page are illustrative. The method holds; the numbers must be yours.

Start with one use-case.

Not a sales call: a structured walkthrough of one AI use-case against the eight lifecycle controls, and where your first template would apply.

Start the conversation on LinkedIn →
Nitin Bhatti
Nitin Bhatti
Founder & Platform Architect, OrchestraPrime
26 yrs enterprise technology · Fortune 50 pharma · TOGAF · PMP · MBA · Patent holder
linkedin.com/in/nitinbhatti · orchestraprime.ai
North Brunswick, NJ
Contact QR codeLinkedIn QR code
© 2026 OrchestraPrime LLC. All rights reserved. · Acceptance Criteria for AI in GxP — From CSV Principles to AI Lifecycle Practice, v0.6, 23 September 2026 · Worked-example values (test sizes, thresholds, limits, SLAs, FTE estimates, CMDB class names) are illustrative design examples, not benchmarks. Draft EU GMP Annex 22, the Chapter 4 and Annex 11 revision drafts and the FDA CSA guidance are cited with their status. The Strategic Planning Assumptions are the author's judgements; the probabilities are not survey results.