Right-Sized AutonomyA Draft GxP Model for Pharma Agents
One sentence. Pharma is putting agents into production faster than it is agreeing how far any one of them may act on its own. The question that matters is not "how autonomous can we get?" It is "what is the right amount of autonomy for this decision, at this criticality, given what this organization can prove today?"
The spine: right-sized autonomy. Maturity is the ability to justify autonomy, and to lower it. Regulation changes rows, not plans.
The ceiling at a glance: permitted autonomy falls as criticality rises
Default ceiling, overlay v0.1 (2026-09-23). A starting point your QRM owner adopts, tightens or relaxes with documented justification. No regulator has published it. Dots show the §8.2 rows as computed there (LLM unless marked; row 3 irreversible; row 8b deterministic).
ACE is a proposal. It is not an established standard, and no body named in this article has reviewed or endorsed it. Every mapping to an external framework is the author's interpretation and is labeled as such. The model text and tables are offered under CC BY 4.0 so that teams can adapt them, pilot them and send back what breaks.
The fragments exist. They do not compose. ACE composes them with one rule.
No public maturity model for pharma agents does three things at once: grade what an agent may do, tie that to GxP evidence, and work across functions. The fragments exist; they do not compose. ACE composes them with one rule: permitted autonomy = min(ceiling, enablement). The ceiling comes from the decision's criticality. The enablement comes from what the organization can prove. The target is the lowest level that delivers the value. Higher is not better: maturity is the ability to justify autonomy, and to lower it the day the evidence weakens. Regulation changes the ceiling rows, not the plan. This is v0.1, and it needs practitioners to break it.
Five findings, each with the section that supports it
The fragments do not compose.
BioPhorum's DPMM and ISPE Pharma 4.0 grade digital plants. ISPE's 2022 GxP AI maturity model grades how a model learns. CIOMS XIV names oversight modes for pharmacovigilance only. Cross-industry agent ladders (AWS, CSA, Capgemini, CMMI AIM) know nothing about GxP. Each is strong in one respect, and none answers the question a Quality head actually asks.
Action autonomy is ungraded in GxP.
The published pharma constructs grade model learning or digital capability. Agent risk sits in what the agent is allowed to do: read, draft, stage a change, commit to a GxP record, or submit to a health authority. The GAMP AI Guide's published table of contents contains no GenAI, LLM or agent headings; whether its body addresses them has not been checked for this draft. TOC only
Regulation already implies a falling ceiling.
Draft Annex 22 says generative and LLM models "should not be used in critical GMP applications" and requires a human in the loop for non-critical use. FDA's draft credibility framework rates risk by model influence and decision consequence. Together they imply a ceiling that falls as criticality rises, and falls further for non-deterministic models. Nobody has written it down as a table.
Most agentic value is outside GxP, and the risk concentrates at the hand-offs.
Commercial, corporate and research agents can run at far higher autonomy than a deviation agent. The inspection risk sits in the steps where a non-GxP agent hears an adverse event, touches a validated system or writes to a GxP record.
Capability is earned per dimension, not scored.
An organization with world-class data can still be unready for autonomous action if it cannot detect a silent model change. A single enterprise score hides exactly the dimension that binds.
Four assumptions, with probabilities stated so they can be tested
Probabilities are the author's judgment, stated so they can be tested and revised in v0.2.
The final text of EU GMP Annex 22 will keep an exclusion of generative/LLM models from critical GMP decisions, possibly with narrower wording.
Substantiation
Most production agents in top-20 pharma will operate at A1–A2 (inform and draft) in GxP work, with A3 and above concentrated in non-GxP functions.
Substantiation
At least one industry body (ISPE's GAMP Community of Practice, BioPhorum, the Pistoia Alliance or DIA) will publish an autonomy-level scheme for agents in regulated work.
Substantiation
Per-decision autonomy records (what the agent may do, who holds the human decision, and on what evidence) will be a routine audit request for HR and patient-program agents, driven by non-GxP rules rather than GxP ones.
Substantiation
What each role should do, in the order it matters
- Stop reporting one "agentic maturity" number. Report, per use case, the permitted level, the target level and what binds: the ceiling or a named capability dimension (§8).
- Put ACE inside your existing AI management system rather than beside it: ISO 42001 supplies the management system, and ACE supplies the per-decision ceiling and evidence for GxP and agents (§13.2).
- Own the ceiling. Adopt the default table in §6.5 as a starting point, then tighten or relax cells with documented justification under ICH Q9(R1). Version it.
- Ask for one filled-in Autonomy Assignment Record (§10.2) per agent in a GxP workflow before the next promotion decision.
- Make the hand-off steps your first inspection-readiness target: every place a non-GxP agent can receive an adverse event, touch a validated system or write a GxP record (§6.4).
- Classify commercial and medical agents as regulated non-GxP (C0-R) with sub-type caps: external promotion and decisions about individuals stay at A2 (§6.5).
- Require every agent that can hear an adverse event, product complaint or off-label request to detect and route it. That step inherits the GxP class, whatever the agent's home function (§6.4).
Pick your role. The contents highlight your path.
The article is written to be read straight through, but five kinds of readers will weight it differently. Select a role: the table of contents highlights the sections to read first and then, and the panel below shows what that role gains. Your choice is remembered on this device. Terms used across the series are defined in the box at the end of this section.
| If your role is… | Read first | Then | Approx. time | Why |
|---|---|---|---|---|
| Enterprise AI strategy (CDO, CAIO, head of AI platform) | Model on one page; §1, §4, §8, §13 | §3, §9 | ~15 min | The rule, the worked example and what it means for your next four quarters |
| Quality and QA leadership | Model on one page; §6, §8, §10.2 | §12, §14, §18 | ~18 min | The ceiling you would own, how it is justified, and what an inspector would be handed |
| CSV, validation and GAMP practitioners | §5, §6, §11 | §7, §12, Appendix B | ~20 min | Evidence per level, per-component ceilings, the GAMP lens and the crosswalk |
| Commercial, Medical and Privacy compliance | §6.3–6.5, §8 | §12, Appendix A (group D) | ~12 min | Regulated non-GxP classes, the boundary-crossing rule and the non-GxP overlay rows |
| Agentic Lab and Factory leads | §8, §9, §10 | §7, §11.2 | ~15 min | The calculator, the self-assessment and the building blocks behind the promotion gate |
| Role | Before | After |
|---|---|---|
| Enterprise AI strategy | One maturity score per company; autonomy treated as progress; platform investment sequenced by agent | Per-use-case permitted and target levels; the binding constraint named; investment sequenced by the dimension that unlocks the most use cases |
| Quality and QA | Ceiling decided case by case in review meetings; no versioned basis; new regulation triggers a policy rewrite | A default ceiling table owned by QRM, versioned as an overlay; a new rule changes a row and re-runs only the affected assessments |
| CSV and GAMP | Model-centric validation; tools, memory and orchestration outside the configuration baseline | Level-dependent evidence for tools, planning, memory and orchestration; per-component ceilings that make split designs consistent |
| Commercial, Medical and Privacy | "Not GxP" read as "no limit"; adverse events in rep conversations handled by training alone | Sub-type caps for promotion, individuals and finance; detection-triggered hand-offs designed as separate, lower-autonomy steps |
| Lab and Factory | Promotion gate criteria argued per agent | A B3 record per agent: permitted ≥ target, evidence attached, demotion triggers named |
Whatever lens you bring, the argument has one spine: right-sized autonomy. Maturity is the ability to justify autonomy, and to lower it. Regulation changes rows, not plans.
Terms used in this article (glossary; internal labels like G7, G8, T1–T10, UC-H1, B1–B9 and Pattern A/B are defined here)
| Term | Meaning here |
|---|---|
| A-level (A0–A5) | How much an agent does without a human deciding: Manual, Inform, Draft, Stage, Act, Direct (§5) |
| C-class (C0-B, C0-R, C1–C3) | The criticality of one decision: business, regulated non-GxP, then GxP supporting, indirect and direct (§6) |
| E-profile (E1–E5 × 8 dimensions) | What the organization can prove, scored per dimension (§7) |
| Ceiling | The highest A-level the decision's class and model type allow under the current overlay (§6.5) |
| Overlay | The versioned set of regulatory rows that sets ceiling cells and evidence thresholds (§12) |
| ODD (operating domain) | The declared scope an agent's level holds in, e.g., "deviation triage, non-sterile orals, Site X" |
| HITL / HOTL / HIC | Human in the loop (approves each output), on the loop (monitors and can stop), in command (sets goals and governs); CIOMS XIV vocabulary |
| Pattern A / Pattern B | Two ways to position an agent for validation. A: the agent drafts, a qualified human decides and signs. B: the critical decision runs on a validated deterministic component; the LLM explains |
| Confidence gating (the platform article's control G7) | Routing each output by calibrated confidence: autonomous, notify, suggest or escalate (§7.4) |
| Regional compliance as configuration (control G8) | The platform article's pattern of holding jurisdictional rules as versioned configuration, not code |
| T-templates (T1–T10) | The validation article's template pack, e.g., T1 intake and GxP assessment, T3 context of use, T6 validation plan, T8 monitoring plan, T9 change envelope |
| Lab → Factory gate | The platform article's promotion gate (§7.3 there) where a proof of concept becomes a governed production agent |
| Building blocks (B1–B9) | The nine records an organization builds along the Lab → Factory path, from the agent registry (B1) to the overlay register (B9) (§10.1) |
| Dimensions (D1–D8) | The eight E-profile dimensions: value, data, registry, assurance, oversight, security, monitoring, people (§7.2) |
| Rules R1–R5 | The calculator's rule tables: class, ceiling, reversibility, enablement, result (Appendix C) |
| UC-H1 (the platform contract) | The platform article's design for identity, tool access, logging and model-version notice as contract terms with a vendor (§13.3) |
Three axes, one rule, three questions per decision
Everything else in the article explains, justifies or applies this page.
Permitted autonomy = min( Ceiling[C-class, model type, reversibility, jurisdiction], Enablement[E-profile] )
Target autonomy = the lowest level that delivers the value case.
Promotion needs sustained evidence; demotion is immediate.
The ceiling
Default, overlay v0.1; generative/LLM component. Rows are criticality classes in ascending order; the cell is the highest permitted A-level. Full table with deterministic and dynamic columns, and the basis for every cell, in §6.5. The interactive version is the heatmap above.
| Class | A0 Manual | A1 Inform | A2 Draft | A3 Stage | A4 Act | A5 Direct |
|---|---|---|---|---|---|---|
| C0-B business | ● | ● | ● | ● | ● | ◐ sandbox only |
| C0-R regulated non-GxP | ● | ● | ● | ◐ internal only | ◐ internal, reversible | ○ |
| C1 GxP supporting | ● | ● | ● | ● | ○ | ○ |
| C2 GxP indirect | ● | ● | ● | ○ | ○ | ○ |
| C3 GxP direct | ● | ● | ○ | ○ | ○ | ○ |
● permitted · ◐ permitted only for the named sub-type · ○ above the ceiling
The enablement is read from an eight-dimension profile, not a single score (§7). A radar of the eight dimensions shows which one binds.
Three questions, asked per decision
Agents Are Shipping Faster Than the Rules That Bound Them
The one-page model answers a question the industry has not yet agreed how to ask. This section shows why it has become urgent.
Picture three people in the same company on the same Monday. Each is being asked the same question in different words: how far may this agent go on its own? None of them has a shared, written answer.
Has approved an internal "agentic factory" to move proofs of concept into production.
"How far may this agent go on its own?"
Has been asked to sign off the first agent that drafts deviation hypotheses.
"How far may this agent go on its own?"
Has just learned that the next-best-action agent in the field app hears what physicians tell sales representatives, including, occasionally, an adverse event.
"How far may this agent go on its own?"
The investment is real, and the value is not yet. Enterprise agentic platforms are now funded at board level; Merck, for example, announced a commitment of up to $1B with Google Cloud for an enterprise agentic platform (Google Cloud, Apr 2026). Several top-20 companies are building internal agentic factories to industrialize proofs of concept, and none has published the GxP gate criteria those factories apply. Meanwhile, the value numbers converge in single digits:
Self-reported "agentic adoption" varies widely with question framing: KPMG reports 85% with "significant or growing" agentic use (KPMG); Deloitte reports 35% with significant progress (Deloitte midyear 2026). Both are self-reported.
What is in production sits at the low end of autonomy. The best-known public examples are all vendor-reported: Novo Nordisk's clinical-study-report drafting, reported as ~10 weeks to ~10 minutes (Anthropic customer story); AstraZeneca's Development Assistant with 1,000+ users and 9 agents (ZenML/AWS case study); and Moderna's 750+ custom GPTs (OpenAI customer story). All three are read, draft and query work. No public example shows an agent autonomously closing a deviation, submitting an individual case safety report (ICSR) or releasing a batch.
That may be exactly right. The problem is that nobody can yet say why, for which decisions, or what evidence would justify going further. Without that answer, two failure modes appear together: agents held at "draft" everywhere because nobody can defend more, and agents quietly given write tools in functions that assume "not GxP" means "no limit".
Before proposing anything new, §2 shows what the industry already has, and where each piece stops.
Next → §2 · What ExistsWhat Exists: Six Families of Models, Each Stopping Short
§1 showed the question. The industry already holds most of the pieces of an answer; this section lays them side by side.
Pharma digital baselines are mature but pre-agentic. BioPhorum's Digital Plant Maturity Model (DPMM) has five named levels: Pre-digital plant, Digital silos, Connected plant, Predictive plant and Adaptive plant. About 20 experts from 11 biopharma companies built it in 2016, and version 3.0 (Nov 2023) added QC, QMS and process-development dimensions (BioPhorum DPMM 3.0↗ ref; ISA InTech 2017).1 The ISPE Pharma 4.0 Baseline Guide includes a maturity model and assessment tool whose level detail is paywalled (ISPE↗ ref).
The only pharma-native autonomy-to-validation construct predates agents. The ISPE D/A/CH working group's 2022 model crosses five Control Design stages (from "used in parallel to the normal GxP processes" to "running automatically and corrects itself") with six Autonomy stages (from fixed algorithms to "self-determines its task competency and strategy"). The combination maps to six validation levels; at Level VI the "validation concept [is] currently under development" (ISPE, Mar/Apr 2022↗ ref). Its autonomy axis measures how a model learns. An LLM agent is usually a locked model, so it sits low on that axis, yet it can have high action autonomy. The axis does not capture that.
The GAMP AI Guide is the lifecycle reference, and its maturity appendix is not public. The ISPE GAMP Guide: Artificial Intelligence (July 2025, 290 pp.) is the industry's AI lifecycle reference (ISPE↗ ref). Its published table of contents lists Appendix M10 "AI Maturity" (pp. 213–216) with headings for AI sub-system autonomy, AI sub-system adaptiveness and AI maturity levels, and Appendix M9 §20.2 "Human Autonomy and Control" (official TOC). TOC only This article draws only on that table of contents and on ISPE's own articles about the Guide; the Guide's body has not been reviewed for this draft. The M10 headings follow the same two-axis structure as the 2022 D/A/CH model, which is the best public proxy, but the correspondence is an inference from headings, not ISPE text.
| Model | Publisher · date | Levels | Pharma / GxP fit | Where it stops for agents |
|---|---|---|---|---|
| Digital Plant Maturity Model v3 | BioPhorum · 2016 / Nov 2023 | 5: Pre-digital → Adaptive | High (manufacturing) | No AI-autonomy or validation axis; plant scope |
| Pharma 4.0 Maturity Model | ISPE · Dec 2023 | 5 (names paywalled) | High (manufacturing) | No agent concept |
| AI Maturity Model for GxP | ISPE D/A/CH · Mar 2022 | 5 control × 6 autonomy → 6 validation levels | Highest: validation-mapped | Single system; learning, not action, autonomy; pre-LLM |
| GAMP Guide: AI, App. M10 | ISPE · Jul 2025 | "AI Maturity Levels" (content paywalled) | Highest (GxP lifecycle) | Level definitions not public; published TOC has no GenAI, LLM or agent headings |
| Gen AI maturity model for drug development | DIA Global Forum (CGI/Roche) · Jan 2025 | 4: interaction → self-learning | Medium (R&D, regulatory) | No oversight or validation link (DIA) |
| ML Maturity Ladder | Cradle.bio · 2025–26 | 0–4, ending in agents | Low (discovery) | No GxP (Cradle) |
| Biomanufacturing AI Maturity Index | Axio BioPharma · 2026 | Continuum | Medium-high | No autonomy levels (Axio) |
| Commercial agentic maturity curve | Axtria · 2026 | 4 stages (gated) | Medium (commercial) | No GxP; uses reversibility as a gate (Axtria) |
| CIOMS WG XIV oversight modes | CIOMS · Dec 2025 | HITL / HOTL / HIC | High (PV consensus) | Pharmacovigilance only; no enterprise capability axis |
| Pistoia Alliance agentic AI initiative | Pistoia · Sep 2025 → | None yet | High (consortium) | Plans an agent standard and validation metrics; no maturity or autonomy scheme found in a public search on 23 Sep 2026 (Pistoia) |
| Cross-industry agent ladders: AWS levels, scoping matrix and graduated autonomy; Capgemini L0–L5; CSA autonomy levels and governance model; Salesforce; IDC; Gartner | 2025–2026 | 4–6 per ladder | Low to medium | Best separation of capability, scope and trust (AWS); no criticality ceiling; no GxP |
| CMMI AI Maturity Model (AIM) | CMMI Institute / ISACA · Jul 2026 | CMMI levels; 31 practice areas | Low-medium | Enterprise capability and appraisal; no per-use-case autonomy or GxP (ISACA) |
Positions are the author's judgment from the §2 table, not a measured scale. ACE is drawn as a dashed outline: a draft aimed at the empty quadrant, not an established occupant of it.
Academic work supplies the missing concepts. Feng, McDonald and Zhang treat autonomy as "a deliberate design decision" separate from capability (Knight Institute, Jul 2025↗ ref). Safin and Balta separate autonomy from agency, whose top level, "committed writes to authoritative records", maps onto GxP records (arXiv 2605.12105↗ ref). A 2026 PHUSE paper gives each agent an intended use and an autonomy level and requires the orchestration platform to be validated (PHUSE 2026↗ ref). And a DIA/RegASK article said "scaling agentic AI demands a maturity roadmap", then deferred it (DIA, May 2026). The industry is asking for this artifact.
1 Several recent conference decks present a five-stage "digital organization maturity roadmap" (pre-digital → silo → connected → predictive → adaptive). That structure is DPMM's, restated at organization level, and should be credited to BioPhorum. ↩
Each fragment is strong somewhere. §3 names precisely where they leave a gap, and what regulation already says about the gap.
Next → §3 · Five GapsFive Gaps, and a Ceiling Regulation Already Implies
§2 showed six families of models. Read together, they leave five gaps, and the first two are the ones a Quality head feels.
Autonomy of action is ungraded in GxP.
Agent risk sits in what the agent is allowed to do: read, draft, stage a change, commit to a GxP record, or submit to a health authority. No published pharma model grades that. The GAMP community is starting to fill the gap one workflow at a time. The Jul–Aug 2026 Pharmaceutical Engineering article by Stoop and Walia reports that "several agentic AI tools" have been built for root-cause analysis, investigations and CAPA, and observes that translating the Guide's principles into operational safeguards "requires practical frameworks". It proposes four input-guardrail categories set by risk priority number (RPN) and a data-fitness score (PE, Jul–Aug 2026↗ ref).
Guardrail categories (sourced)
| Category | RPN | Minimum data fitness |
|---|---|---|
| 0 | ≥ 100 | Any: "AI recommendations are advisory only; SME review mandatory before action" |
| 1 | 50–99 | ≥ 80 |
| 2 | 20–49 | ≥ 70 |
| 3 | < 20 | ≥ 60 |
This is de facto oversight tiering from inside the GAMP community, scoped to one quality workflow. It is the strongest public evidence that practitioners want graded autonomy and do not yet have an enterprise-wide scheme.
No model encodes the regulatory ceiling.
Draft EU GMP Annex 22 (7 July 2025, still draft) says dynamic models, probabilistic-output models and "Generative AI and Large Language Models (LLM)… should not be used in critical GMP applications". In non-critical use, qualified personnel must stay responsible for outputs as a human in the loop, and §10.5 may require "review and/or test of every output" (EC, Annex 22 draft↗ ref). FDA's January 2025 draft guidance defines a seven-step credibility framework in which model risk is a combination of model influence and decision consequence (FDA draft guidance↗ ref). Model influence is an autonomy axis in all but name. Together these imply a ceiling that falls as criticality rises, and falls further for non-deterministic models. Nobody has written that ceiling down as a table, and no maturity model applies it.
"Higher is better" is the wrong message for pharma.
Vendor ladders put multi-agent or fully autonomous operation at the top. The design literature points the other way. SAE's own levels were described by critics as "nominal, rather than ordinal", meaningful only within an operating domain (Stayton and Stilgoe, IEEE TSSIT). NIST CSF 2.0 encourages higher tiers only "when risks or mandates are greater or when a cost-benefit analysis indicates" (NIST CSWP 29↗ ref). Capgemini reports trust in fully autonomous agents falling from 43% to 27% in a year, with about 4% of processes expected at its top level within three years (Capgemini, Jul 2025↗ ref). For a Quality audience, a model that recommends right-sized autonomy signals expertise; one that sells maximum autonomy is a red flag.
Nothing connects enterprise capability to per-use-case autonomy.
Organization models (IDC, CSA's governance model, CMMI AIM) score the enterprise. Agent ladders score the agent. No public instrument combines per-decision autonomy, write authority, validation state and organizational governance. The only dynamic scheme found is AWS's graduated autonomy: promotion after sustained performance, immediate demotion, and a safety floor "never averaged away" (AWS Architecture blog↗ ref). It has no regulated-industry specifics.
Value is not tied to maturity.
§1's value figures are all under 10%, and no survey maps value to autonomy level or operating model. Nobody can yet say whether higher autonomy pays in pharma. That is a reason to set targets by value case, not by ambition.
Five gaps, one common cause. Every fragment grades one thing on one ladder. §4 sets out a stance that keeps the three things apart and joins them with one rule.
Next → §4 · The Design StanceThe Design Stance: Right-Sized Autonomy
The gaps in §3 all come from collapsing different questions onto one ladder. ACE keeps them apart.
From here to §13, everything is the author's proposal unless it carries a citation. Tables that mix sourced and proposed content say which is which.
The axes are borrowed; say so plainly. ACE separates what the agent may do (Autonomy), where it may do it (Criticality) and what the organization can prove (Enablement). The separation follows SAE's level-versus-operating-domain split, IMDRF's significance-by-situation grid (IMDRF N12↗ ref) and CMMI's continuous representation. It also extends the "two axes, not one ladder" argument in the validation article (§4.1 there), which separates how a model changes from how it produces output; ACE adds a third axis for action.
Autonomy (A0–A5)
The six levels
Criticality (C0-B, C0-R, C1, C2, C3)
Enablement (E1–E5 × 8 dimensions)
The eight dimensions
The rule. Because the A-scale is ordinal in delegated authority, taking the minimum is meaningful:
Permitted A = min( Ceiling[C, model type, reversibility, jurisdiction], Enablement[E-profile] )
Target A = the lowest level that delivers the value case, set per use case and reviewed by Quality. Promotion needs sustained evidence; demotion is immediate.
Read it left to right. "Capped by ceiling" leads to redesign (for example, a Pattern B split) or acceptance, because the limit comes from the decision and the overlay, not from the organization. "Capped by capability" leads to investment in the named dimension, because the limit is the organization's to move.
The ceiling logic comes from Annex 22 and FDA's influence-by-consequence framing; enablement comes from CMMI- and DPMM-style capability; promotion and demotion come from AWS graduated autonomy. Pöppelbuß and Röglinger ask prescriptive maturity models to include an explicit decision calculus (ECIS 2011↗ ref); this is ACE's.
What is new is not the axes. It is three commitments:
A criticality ceiling owned by the QRM owner and versioned as an overlay
A boundary-crossing rule
risk concentrates where non-GxP agents touch GxP records or obligations, so those steps inherit the GxP class (§6.4).
"Higher is not better"
maturity means being able to show, for every agent, why its autonomy is exactly as high as the evidence and the decision allow, and to lower it the day the evidence weakens.
Stable core, pluggable overlay
Guidance on agents is arriving from many authorities at once. If an execution plan is written against one text, it must be rewritten each time a draft is finalized. ACE therefore has two layers:
| Layer | Contents | Changes when… | Who maintains |
|---|---|---|---|
| Stable core (build once) | The three axes; the rule; the four authority-neutral criticality attributes; the nine building blocks; the kinds of evidence per level; the GAMP lifecycle lens | Practitioner feedback shows the structure is wrong (v0.x revisions) | Model authors and, in time, a consortium |
| Regulatory overlay (versioned) | Ceiling cells; specific evidence thresholds ("review every output"); oversight wording; jurisdiction rows | An authority issues, finalizes or revises a text | Each organization's Quality and regulatory-intelligence owners, starting from the public register in Appendix A |
The platform article applies the same idea at the architecture level (control G8, "regional compliance as configuration"; see the glossary). ACE applies it to the maturity model.
With the stance fixed, the next three sections define each axis in turn, starting with the one that makes the model auditable: what an agent at each level actually does.
Next → §5 · Axis AAxis A: Six Autonomy Levels an Auditor Can Observe
§4 said A-levels are ordinal in delegated authority. Here is what each level looks like to someone watching the work, and what evidence holds it.
Each level is named for observable behavior, so an auditor can check it, and carries a one-word handle for use in conversation. The names avoid aspirational words like "transformational". Open a level's evidence to see the minimum that holds it (§5.1); each level needs everything below it, plus the additions shown. The evidence types are core; the specific thresholds are overlay and will change.
Evidence to hold A1
Evidence to hold A2
Evidence to hold A3
Evidence to hold A4
Evidence to hold A5
Mappings to AWS, CSA, Capgemini and ISPE control-design stages are in Appendix B. A use case can sit at different levels for different stages of the work. Following Parasuraman, Sheridan and Wickens' four functions (acquisition, analysis, decision, action), an agent can be A4 for analysis and A2 for decision. ACE records the level of the decision or action stage as the headline.
5.1 Evidence required to hold each level
This is where the model becomes auditable. Evidence accumulates: each level needs everything below it, plus the additions shown in the cards above. The same list as a table:
Evidence per level, as a table (5 rows, proposed)
| Level | Minimum evidence to operate at this level (proposed) |
|---|---|
| A1 Inform | Registered agent card; approved-source grounding; citation on every output; read-only tool allowlist tested |
| A2 Draft | + Intended-use and context-of-use statement; classification record; acceptance criteria against a measured human baseline; N-run acceptance for probabilistic output; accept/modify/reject attributed in the record (Part 11, ALCOA+) |
| A3 Stage | + Staging-area design; per-action approval audit trail; trajectory tests (right source, checker called, stops at threshold); prompt and tool configuration under change control |
| A4 Act | + Declared ODD with machine-readable bounds; confidence-banded escalation tested at each boundary; drift and override monitoring; sampling review plan; tested kill switch; performance at least equal to the process it replaces (Annex 22 draft "at least as high as"; Annex 11 draft "never increase overall risk") |
| A5 Direct | + Executive authorization; pre-approved change envelope; trust score sustained over rolling windows with immediate demotion; independent periodic review. Not proposed for any GxP use under current text |
An A-level is only meaningful against the decision it applies to. §6 defines that decision's criticality, and the ceiling it sets.
Next → §6 · Axis CAxis C: Criticality Is Decided per Decision
§5 graded what an agent may do. This section decides how much of that any one decision can bear, starting from a fact pharma often forgets: not everything in pharma is GxP.
6.1 Four authority-neutral attributes
Each decision is described by four attributes, and none of them names a regulator. A new rule therefore changes what a class is allowed to do, not what the class is.
| Attribute | Values (recorded per decision on the agent card) | Used for |
|---|---|---|
| I · Impact on patient safety, product quality or data integrity | None · Supporting (a human verifies independently before any GxP use) · Indirect (a GxP decision relies on the output) · Direct (the output is the GxP decision or record) | Sets the class |
| W · Write authority over GxP records (writes to non-GxP systems count as None here; reversibility covers them) | None · Draft store · Staged change · Commit to a primary GxP record · External regulatory submission | Can raise the class: commit or submission → C3; staged change to a GxP record → at least C2 |
| R · Reversibility | Reversible · Irreversible (cannot be recalled or corrected before it is relied on outside the process) | Ceiling modifier (§6.5) |
| M · Model type | Deterministic or static ML (including locked models retrained only under change control) · Generative / LLM · Dynamic (continuously learning) | Selects the ceiling column |
6.2 Five classes, in ascending order
The class is the highest that any attribute implies. Open a class to see its default ceiling cells and the basis for them (§6.5).
Ceiling cells and basis
author judgment Company risk appetite; OWASP least agency.
Ceiling cells and basis
By sub-type (same in all three model-type columns):
No agent both creates and approves. direct text for UK promotion (PMCPA: certification "cannot be fulfilled solely by the use of AI") and for solely automated decisions (GDPR Art. 22); by analogy for US promotion (21 CFR 202.1, Form 2253) and finance (COSO GenAI guidance); A-levels are author judgment.
Ceiling cells and basis
direct text (GMP): Annex 22 draft requires HITL for non-critical use; by analogy outside GMP; A-levels author judgment (see the note on C1 at A3 in §6.5).
Ceiling cells and basis
direct text (GMP): Annex 22 draft HITL and §10.5 "every output"; FDA draft (human review moderates model influence); by analogy for GVP and GCP; A-levels author judgment.
Ceiling cells and basis
direct text (GMP): Annex 22 draft §1 excludes GenAI/LLM, dynamic and probabilistic models from critical use. by analogy for GVP, GCP, GDP and submissions: GVP Module VI, CIOMS XIV, ICH E6(R3) sponsor accountability.
6.3 Three kinds of work in one company
A pharma company is not a GxP estate end to end. Much of its agentic value sits in regulated-but-not-GxP work, or in ordinary business work where the company's own risk appetite sets the limit. GxP status belongs to the decision, not to the function.
Levels in brackets are the author's proposal, computed from the default ceilings in §6.5. Inside the Supply-chain band: demand forecasting is C0-B; the GDP temperature-excursion disposition for a shipment is C3.
| Kind of work | Where it sits | What sets the ceiling | Example agents at realistic levels (LLM) author's proposal |
|---|---|---|---|
| GxP (C1–C3) | GLP, GCP, GMP, GDP and GVP work, and the submissions that draw on them | GxP overlay rows: Annex 22 and 11, FDA credibility and CSA, ICH E6(R3), CIOMS XIV, GAMP | SOP first draft (C1, A2–A3); deviation hypotheses (C2, A2); PV intake suggestion (C2, A2); batch-disposition support (C3, A1, Pattern B) |
| Regulated non-GxP (C0-R) | Commercial and marketing; medical affairs; patient and HCP personal data; finance; EU AI Act high-risk uses | Non-GxP overlay rows plus company compliance policy | Pre-call brief (A3); next-best-action suggestions (A3); promotional pre-check before MLR (A2); patient-support eligibility recommendation (A2); reconciliation below materiality (A4) |
| Business (C0-B) | Discovery exploration; HR; IT on non-GxP systems; procurement; corporate functions | Risk appetite, security policy (OWASP least agency, NIST), EU AI Act where applicable | IT incident remediation (A4); invoice matching (A4); in-silico design–make–test loop in a sandbox (A5) |
The same function contains both. Classification per decision is not pedantry:
6.4 The boundary-crossing rule, with an obligation trigger
An agent in a non-GxP function that touches a GxP record or event inherits that record's class for that step. Its other tools keep their own ceiling. The ceiling for the whole agent is never averaged across its steps.
The rule is enforced in two places:
Write authority
The registry records every GxP-writing tool. A commercial agent that can open an adverse-event intake record is classified at the GxP level for that tool call.
Obligation trigger
Pharmacovigilance, product-complaint and off-label-request obligations start when information is received, not when a record is written. So any agent that can receive such a signal must detect and route it, and the detection step inherits C2, even if the agent has no GxP write tool at all. An agent that hears an adverse event and does nothing is not low-risk; it is non-compliant.
Two design consequences follow. Each GxP hand-off (adverse event detected, excursion detected, validated-system change) is designed as a separate, lower-autonomy step. And the agent card (B1, §10) lists the obligation triggers the agent can receive, next to its tools. Worked-example rows 4 and 4a in §8.2 run this flow through the calculator.
6.5 The default ceiling, with its basis
The ceiling is a default starting point that your QRM owner adopts, tightens or relaxes with documented justification under ICH Q9(R1). No regulator has published it. It is inferred from the overlay rows in Appendix A, and it is the most change-prone table in the model.
Default ceiling, overlay v0.1 (as of 2026-09-23). Each cell is the highest permitted A-level before the reversibility modifier. The interactive version is the heatmap in the hero.
| Class | Generative / LLM | Deterministic or static ML | Dynamic (learning) | Basis |
|---|---|---|---|---|
| C0-B | A4; A5 in a declared sandbox with executive authorization | A4; A5 in sandbox | A4; A5 in sandbox | author judgment company risk appetite; OWASP least agency |
| C0-R | By sub-type (same in all three columns): external promotional release A2; decisions about individuals A2; finance: A4 below materiality or into suspense, A2 above, and no agent both creates and approves; internal suggestions to staff A3; other internal reversible work A4 | ← same | ← same | direct text for UK promotion (PMCPA: certification "cannot be fulfilled solely by the use of AI") and for solely automated decisions (GDPR Art. 22); by analogy for US promotion (21 CFR 202.1, Form 2253) and finance (COSO GenAI guidance); A-levels are author judgment |
| C1 | A3 | A4 | A2 | direct text (GMP): Annex 22 draft requires HITL for non-critical use; by analogy outside GMP; A-levels author judgment (see note) |
| C2 | A2 | A3 | A1 | direct text (GMP): Annex 22 draft HITL and §10.5 "every output"; FDA draft (human review moderates model influence); by analogy for GVP and GCP; A-levels author judgment |
| C3 | A1 (Pattern B: see §6.6) | A3; the human remains the release decision-maker | A1 | direct text (GMP): Annex 22 draft §1 excludes GenAI/LLM, dynamic and probabilistic models from critical use. by analogy for GVP, GCP, GDP and submissions: GVP Module VI, CIOMS XIV, ICH E6(R3) sponsor accountability |
Note on C1 at A3. At A3, steps run without approval only if they write nothing into a GxP record (retrieve, compare, draft into staging). Every item that enters the record is approved by a qualified person, which is consistent with Annex 22's human-in-the-loop condition. If your reading of the final §10.5 requires review of every output, set C1 LLM to A2 in your overlay. That is exactly the kind of row change the overlay exists for.
Reversibility modifier. An irreversible action (cannot be recalled or corrected before it is relied on outside the process) lowers the ceiling by one level, to a floor of A1. The modifier does not apply in C3, where the class ceiling already assumes reliance outside the process, or where a C0-R sub-type cap already governs the external act (external promotional release). It never compounds. The modifier follows Axtria's reversibility gating and Anthropic's finding that only 0.8% of observed agent tool calls were irreversible (Anthropic↗ ref).
Jurisdiction. When a use case falls under several jurisdictions, the lowest applicable overlay ceiling applies.
6.6 Ceilings apply per component, not per agent
A multi-agent system is not one decision-maker. ACE therefore applies ceilings per component:
Each component holds its own ceiling. The component that commits the GxP decision must be within the class ceiling for its model type.
This makes Pattern B consistent. In a batch-disposition design, the LLM component reads, summarizes and explains the evidence at A1; it may call the validated calculation read-only. The orchestration control flow is itself deterministic and validated, consistent with the PHUSE 2026 paper. The committing component is the deterministic disposition path at A3, and a qualified person approves the release. No component exceeds its ceiling, so the system is within the model. Rows 8, 8a and 8b in §8.2 show the same split through the calculator.
The ceiling says how far a decision may go. Whether an organization can actually hold that level is a separate question, and §7 answers it.
Next → §7 · Axis EAxis E: Capability Is Earned per Dimension
§6 set the ceiling from outside the organization. This section sets the other term of the rule, from inside it.
The profile is scored separately on each of eight dimensions. The shape matters more than any average; as Microsoft's adoption guidance notes, organizations "show characteristics of more than one level at the same time".
7.1 Five observable levels
7.2 Eight dimensions
Each dimension has one diagnostic question. Open a dimension for the evidence an auditor could inspect.
Evidence an auditor could inspect
Evidence an auditor could inspect
Evidence an auditor could inspect
Evidence an auditor could inspect
Evidence an auditor could inspect
Evidence an auditor could inspect
Evidence an auditor could inspect
Evidence an auditor could inspect
7.3 The enablement rule
Enablement is the highest A-level whose requirements, and the requirements of every level below it, are all met. A dimension not listed for a level has no minimum at that level.
| To enable… | Required minimum on the E-profile |
|---|---|
| A1 Inform | Dimension 3 ≥ E2 (the agent is registered) |
| A2 Draft | Dimensions 3, 4, 5 ≥ E2 |
| A3 Stage | Dimensions 4, 5, 7 ≥ E3 (and dimension 3 ≥ E2 from below) |
| A4 Act | Dimensions 5, 6, 7 ≥ E4; every other dimension ≥ E3; for C1–C3, dimension 4 ≥ E4 |
| A5 Direct | Dimensions 5, 6, 7 ≥ E5; every other dimension ≥ E4; C0-B only |
So an organization at E1 on dimension 3 enables A0: register first. An organization with world-class data can still be unready for A4 if it cannot detect drift (dimension 7).
Profile: E4 on dimensions 3, 5, 6 and 7; E3 on dimensions 1, 2, 4 and 8. Solid polygon: the profile. Dashed polygon: the requirement for the selected level; amber points mark the dimensions that fall short. Drag nothing here; the live version is in the calculator (§8).
7.4 How confidence gating fits inside the rule
The platform article's confidence-gating control (G7, see the glossary) routes each output by confidence: ≥ 0.90 autonomous, 0.70–0.89 notify, 0.55–0.69 suggest, < 0.55 escalate. ACE's permitted level decides which of those bands may be switched on:
| Band | Behavior | Enabled when permitted A is… |
|---|---|---|
| Escalate (< 0.55) | Routed to a human; output not used | Always on (A1 and above) |
| Suggest (0.55–0.69) | Offered as a flagged draft | A2 and above |
| Notify (0.70–0.89) | Staged; human notified and approves | A3 and above |
| Autonomous (≥ 0.90) | Executes within the ODD; sampled after the fact | A4 and above only |
When a band is not enabled, its outputs fall to the highest enabled band below it. Confidence must come from a calibrated evaluator or deterministic checker, not from the model's self-report, which is poorly calibrated. Confidence gating answers "should this output go that far?"; the ceiling answers "how far may this agent go at all?" The two are complementary.
Ceiling and enablement are now both defined. §8 runs them together on one organization's portfolio.
Next → §8 · Running the RuleRunning the Rule: The ACE Calculator and a Worked Example
§6 and §7 each produced one term of the rule. Here they meet, and the result shows why the same organization legitimately runs agents from A1 to A4 at once.
8.1 The calculator in five steps
The full rule tables are in Appendix C; the calculator below implements exactly those tables (R1–R5, see the glossary).
- Class. From I and W (§6.1–6.2), take the highest class implied. Apply the boundary-crossing rule and the obligation trigger per step.
- Ceiling. Look up the cell for class × model type (and the C0-R sub-type) in §6.5. If R = irreversible and the class is not C3 and no C0-R sub-type cap governs the act, subtract one level, to a floor of A1. Take the lowest across jurisdictions.
- Enablement. From the E-profile, find the highest level whose requirements, and those of every level below, are met (§7.3). Use the GxP condition on dimension 4 for C1–C3.
- Permitted = min(ceiling, enablement).
- Result. If target ≤ permitted: operate at target. If target > permitted: capped by ceiling when the ceiling is the lower term, capped by capability when enablement is the lower term, capped by both when they are equal.
Your organization E-profile
Keyboard: focus a point on the radar and use the arrow keys to change its level. The radar and the selects stay in sync.
Your decision per decision, not per function
Result
Set the decision and profile above; the result updates as you go.Your portfolio
| # | Use case | Class | R | Ceiling | Enablement | Permitted | Target | Result | Value metric to baseline |
|---|
8.2 Worked example
illustrative Illustrative organization and use cases; value metrics are examples of what to baseline, not measured results.
The organization's profile: E4 on dimensions 3, 5, 6 and 7; E3 on dimensions 1, 2, 4 and 8. Checking the rule: A3 needs dimensions 4, 5 and 7 ≥ E3 (met). A4 needs 5, 6 and 7 ≥ E4 (met) and every other dimension ≥ E3 (met), so non-GxP enablement is A4. For C1–C3, A4 also needs dimension 4 ≥ E4 (E3: not met), so GxP enablement is A3. A5 needs dimensions 5, 6 and 7 at E5 (not met).
All use cases use an LLM component unless marked.
| # | Use case | Class | R | Ceiling | Enablement | Permitted | Target | Result | Value metric to baseline |
|---|---|---|---|---|---|---|---|---|---|
| 1 | IT incident remediation, non-GxP infrastructure | C0-B | Reversible | A4 | A4 | A4 | A4 | Operate at A4 Confidence gating is the primary control, with sampling and a kill switch | Mean time to restore |
| 2 | In-silico design–make–test loop, research sandbox | C0-B (sandbox) | Reversible | A5 | A4 | A4 | A5 | Capped by capability Next move: dimensions 5, 6, 7 to E5 and 1, 2, 4, 8 to E4; run at A4 meanwhile | Design cycles per week |
| 3 | Purchase-order issue to suppliers | C0-B | Irreversible (binds a supplier) | A4 − 1 = A3 | A4 | A3 | A4 | Capped by ceiling (policy, via reversibility). Add a cancellation window before supplier acknowledgment to make the step reversible, or keep buyer approval | PO cycle time |
| 4 | Next-best-action suggestions to representatives | C0-R (internal suggestion) | Reversible | A3 | A4 | A3 | A3 | Operate at A3 | Call-plan preparation time |
| 4a | … its adverse-event detection step (obligation trigger) | C2 (inherited) | Reversible | A2 | A3 | A2 | A2 | Operate at A2 detect, draft the intake, route to PV for a human decision | Share of AE mentions routed within one business day |
| 5 | Intercompany reconciliation and posting, below materiality | C0-R (finance, below) | Reversible | A4 | A4 | A4 | A4 | Operate at A4 Above materiality the step is A2; creation and approval never in one agent | Days to close |
| 6 | Literature triage for SOP periodic review | C1 | Reversible | A3 | A3 | A3 | A2 | Operate at A2 A3 is permitted but not needed: higher is not better | Reviewer hours per SOP review |
| 7 | Deviation hypotheses | C2 | Reversible | A2 | A3 | A2 | A2 | Operate at A2 The ceiling, not capability, is the binding term | Days to probable root cause |
| 8 | Batch-release support, as first requested (one LLM agent staging the disposition) | C3 | n/a in C3 | A1 | A3 | A1 | A3 | Capped by ceiling (regulation). Redesign as Pattern B (rows 8a, 8b) | Release lead time |
| 8a | … Pattern B: LLM component explaining the evidence | C3 | n/a | A1 | A3 | A1 | A1 | Operate at A1 | (as row 8) |
| 8b | … Pattern B: deterministic disposition path (deterministic) | C3 | n/a | A3 | A3 | A3 | A3 | Operate at A3 a qualified person approves each release | (as row 8) |
The pattern is the point. The same organization legitimately runs agents from A1 to A4 at the same time. The highest autonomy sits where the decision is least regulated, not where the technology is most advanced. Two rows are worth a second look. Row 6 is the spine in miniature: the organization could go higher and chooses not to. Row 3 shows that "non-GxP" does not mean "no ceiling": irreversibility alone lowers it.
The organization's only GxP-binding capability gap is dimension 4 (assurance and validation). Raising it to E4 would lift GxP enablement to A4, yet it would change no permitted level in this portfolio, because every GxP row is held by its ceiling. That is useful to know before funding it.
The worked example assumed the profile was known. §9 shows how an organization produces that profile, and where to start if it has none.
Next → §9 · Self-AssessmentSelf-Assessment: Five Steps and Where to Start
§8 ran the rule on a known profile. This section is how a team produces its own, in one workshop, starting from its agent inventory.
9.1 Five steps
| Step | Action | Input | Output (building block, §10) |
|---|---|---|---|
| i. Profile | Score each of the eight dimensions E1–E5 from the diagnostic questions; attach evidence, not opinion | §7.2 evidence column | E-profile (radar) and the lowest dimensions |
| ii. Classify | Sort each use case into GxP, regulated non-GxP or business; record I, W, R, M and obligation triggers per decision; derive the class; apply the boundary-crossing rule per tool | Use-case description, ODD | B2 classification record |
| iii. Compute | Permitted = min(ceiling, enablement) | Overlay version, E-profile | B3 assignment record |
| iv. Gap | Compare permitted with target | Value case and baseline | Operate at target, capped by ceiling, capped by capability, or capped by both |
| v. Next steps | For capability caps, pick the dimension move that unlocks the most use cases; for ceiling caps, redesign (Pattern B, a reversible step) or accept | Gap list | A roadmap sequenced by dimension, not by agent |
9.2 Where am I, what next
ACE does not label a company "E3". Read the rows as "if your binding dimensions look like…". For a mixed profile, use the calculator: the useful output is the lowest required dimension for the next level. The "highest you can hold" column applies the default LLM ceilings of §6.5.
| If your binding dimensions look like… | Enablement | Highest you can hold (LLM, default ceilings) | Next three moves | Do not yet |
|---|---|---|---|---|
| Pilots nobody can list; no class on any agent (E1) | A0: register first | Nothing in production | Stand up the B1 registry; classify every agent (B2); measure one human baseline | Put any agent output into a GxP record |
| Every agent registered and classified; controls vary by team (E2) | A2 | A2 in C0-B to C2; A1 in C3 | Lifecycle SOPs; a promotion gate with a Quality vote (B6); B5 evidence packs | Staged or automatic writes (A3+) |
| Shared governance plane; gate operating; evidence packs standard (E3) | A3 | A3 in C0-B, C0-R internal and C1; A2 in C2; A1 in C3 | Monitoring and probe checks (B7); trajectory testing at scale; pinned models (B8) | A4 anywhere GxP-relevant |
| Monitoring live; drift and overrides visible; value measured (E4) | A4 | A4 in C0-B and internal reversible C0-R; A3 in C1; A2 in C2; A1 in C3 (deterministic components: A4 in C1, A3 in C2–C3) | Evidence-driven promotion and demotion; pre-approved change envelopes; overlay-version governance (B9) | A5 outside sandboxes |
| Autonomy moves up and down on evidence; overlay updates are routine (E5) | A5 in C0-B sandboxes; A4 elsewhere | The overlay ceiling in every class | Benchmark with peers; contribute overlay rows | Treat the ceiling as a target |
9.3 Minimum viable adoption: 30 and 90 days
The full model has many parts. Adoption does not need all of them at once.
A B1 registry entry and a B2 classification record for every agent that exists today
This alone finds the agents with GxP write tools or obligation triggers nobody knew about.
A B3 assignment record for every agent in a GxP workflow, and B6 gate criteria that require permitted ≥ target before promotion
The promotion gate now has a rule, not an argument.
The full overlay, the E-profile benchmark, the agent-property extensions
None of these is needed to start.
The self-assessment writes into records. §10 names the nine records an organization builds, and shows one filled in.
Next → §10 · Building BlocksBuilding Blocks: What an Organization Actually Builds
§9's steps each land in a building block. These are the things that persist when regulation changes: a new rule changes a block's content, never whether it exists.
10.1 Nine blocks along the Lab → Factory path
What a regulatory change touches · worked example
Regulatory change touches: new fields only.
Where to see a worked example: Platform §6
What a regulatory change touches · worked example
Regulatory change touches: rarely the class; the ceiling via the overlay.
Where to see a worked example: Validation §8.1 (T1)
What a regulatory change touches · worked example
Regulatory change touches: recomputed when the overlay changes.
Where to see a worked example: §10.2 below
What a regulatory change touches · worked example
Regulatory change touches: wording; review-every-output threshold.
Where to see a worked example: Platform §5.5
What a regulatory change touches · worked example
Regulatory change touches: added evidence items.
Where to see a worked example: Validation §9 (T3, T6)
What a regulatory change touches · worked example
Regulatory change touches: gate criteria reference new evidence.
Where to see a worked example: Platform §7.3
What a regulatory change touches · worked example
Regulatory change touches: thresholds.
Where to see a worked example: Validation §13 (T8)
What a regulatory change touches · worked example
Regulatory change touches: what counts as a change.
Where to see a worked example: Validation §14 (T9)
What a regulatory change touches · worked example
Regulatory change touches: this is where change lands.
Where to see a worked example: Appendix A
10.2 A sample B3 Autonomy Assignment Record
Inspectors do not grade maturity; they read records. This is what one would be handed. Illustrative: the organization, identifiers, names and dates are invented; the structure is the proposal.
Identity and scope
- Agent · version
- Deviation Hypothesis Assistant · v2.3 (LLM, vendor model pinned by version)
- Operating domain (ODD)
- Deviation investigations, non-sterile oral solid dose, Site X, EU and US markets
- Decision assessed
- Proposing probable root-cause hypotheses for a QA investigator
- Obligation triggers
- None it can receive (internal QMS data only)
Classification
- Attributes
- I = Indirect (the investigator relies on the hypotheses) · W = Draft store (investigation workspace; no write to the QMS record) · R = Reversible · M = Generative / LLM
- Class
- C2
- Overlay version · jurisdictions
- ACE overlay v0.1 (2026-09-23) · EU (Annex 22 draft, Annex 11 draft) and US (FDA CSA, draft credibility guidance)
Calculation
- Ceiling · basis
- A2 · Direct text (GMP): Annex 22 draft HITL, §10.5; A-level author judgment adopted by QRM owner, QRM file ref. QRM-AI-017
- Reversibility modifier
- Not applied (reversible)
- E-profile checked (date)
- 2026-09-10: D3 E4 · D4 E3 · D5 E4 · D7 E4 (A3 requirements met; A4 not met for GxP: D4 < E4)
- Enablement
- A3
- Permitted
- A2 (binding term: ceiling)
- Target
- A2: hypotheses drafted with cited evidence; the investigator accepts, edits or rejects each one
- Result
- Operate at A2
- Confidence bands enabled
- Escalate, Suggest (Notify and Autonomous disabled at A2)
Operation and approval
- Value case
- Days to probable root cause: baseline 11.0 (measured Q2 2026), target ≤ 7.0
- Evidence pack
- Context-of-use statement; acceptance vs. measured investigator baseline; 5-run acceptance; attributed accept/modify/reject in the record
- Demotion triggers
- Unsupported-claim rate above limit in monthly sample; investigator override rate outside band; unannounced vendor model change
- Approvers
- QA Head, Site X (QRM owner) · Deviation process owner · Agentic Factory lead
- Effective · next review
- 2026-10-01 · 2027-04-01, or on any overlay version change touching C2
The building blocks map onto a lifecycle that CSV teams already run. §11 is written for them.
Next → §11 · For CSV and GAMP PractitionersFor CSV and GAMP Practitioners
§10 named what gets built. This section places it inside the GAMP AI lifecycle, adds the four agent properties a model-centric lifecycle does not reach, and points to the full crosswalk.
11.1 The GAMP AI lifecycle as a directional lens
ACE is meant to be read through the GAMP AI lifecycle: an agentic layer on a structure CSV teams already use, not a competing lifecycle. The phase and gate names below come from the Guide's published table of contents only TOC only; the ACE columns are the proposal.
Phase and gate names from the GAMP AI Guide's published table of contents; ACE layers are the proposal. The B3 record is created at Concept and recomputed at the Concept → Project gate, which is the Lab → Factory gate.
| GAMP AI phase (TOC) | Phase gate (TOC) | ACE activity proposal | ACE evidence produced |
|---|---|---|---|
| Concept (incl. "Context of Use and Scope Definition", "Initial Risk Assessment") | PoC evaluation; transition to Project | Classify each decision (B2); declare ODD and target level; check enablement | Agent card (B1); context-of-use statement; measured baseline. Entry to Project = the Lab → Factory gate |
| Project (incl. "Model Design Space", "Testing of AI-Enabled Computerized Systems") | Reporting and release | Position the agent (Pattern A or B) before validating; confirm per-component ceilings | A2: acceptance vs. baseline, N-run. A3: trajectory tests, staging design. A4: boundary and escalation tests, kill switch |
| Operation (incl. "System Monitoring", "Operational Change and Configuration Management", "Periodic Review") | Periodic review | Promote on sustained evidence; demote immediately | Override and catch-rate data; drift and probe checks; prompt, model, tool and memory changes logged as changes |
| Retirement | Retirement reporting | Revoke identity and tool credentials | Archived model, prompt and memory versions for reconstructability |
Across all phases, the TOC lists appendices on human autonomy and control (M9) and AI maturity by sub-system autonomy and adaptiveness (M10). ACE is designed to sit beside M10's per-sub-system classification: adaptiveness feeds the model-type column of the ceiling, and the A-scale adds action autonomy and write authority. Whether M10's content already grades something close to action autonomy has not been checked for this draft; if it does, ACE's claim narrows to "extends M10 across functions and to write authority" (Appendix D).
11.2 Four agent properties, controlled as configuration
The Guide's published table of contents has no headings for these four properties. ACE treats each as a controlled configuration item with level-dependent evidence, following the agent-validation table in the validation article (§9 there).
| Agent property | Risk a model-centric lifecycle does not reach | ACE control (proposed) | Mandatory from |
|---|---|---|---|
| Tool permissions | Reach is set by tools, not the model; a write tool turns a draft into an action | Allowlist as a configuration item; write tools separated from read tools; negative tests proving the agent cannot call out-of-list tools; per-agent identity | A1 (read-only list) → A3 (staged write) → A4 (bounded commit) |
| Planning | A right answer can hide a wrong path: wrong source, skipped checker, missed escalation | Trajectory tests; step budget; checkpoints before consequential steps | A3 |
| Memory | Persistent memory changes behavior between releases, a form of adaptiveness | Memory scope on the card; no GxP record data in free memory; memory versioned or reset under change control; retrievable for inspection | A3 (any persistent memory) |
| Orchestration | Risk emerges from composition: scope elevation, error propagation | Orchestrator control flow deterministic and validated; per-component ceilings (§6.6); scope elevation logged; registry refuses disallowed compositions | A3 for any orchestrated GxP workflow |
11.3 Crosswalk in brief
ACE is framed through the GAMP AI lifecycle, extends the 2022 D/A/CH matrix to action autonomy, generalizes the 2026 guardrail categories beyond one workflow, complements DPMM and Pharma 4.0 as the agentic layer above the digital foundation, maps to CMMI AIM appraisal results as evidence for the E-profile, borrows AWS's scope, trust-tier and demotion logic, adopts CIOMS XIV's oversight vocabulary, operationalizes FDA's influence-by-consequence framing, and encodes draft Annex 22 as ceiling cells. The full crosswalk, row by row, is in Appendix B.
The ceiling cells that the crosswalk points to will move as texts are finalized. §12 shows how ACE absorbs that without rewriting any plan.
Next → §12 · The Regulatory OverlayThe Regulatory Overlay: Regulation Changes Rows, Not Plans
§6.5 called the ceiling "the most change-prone table in the model". This section is the change mechanism.
No regulator text specific to agentic AI sets a ceiling yet; the UK ICO's non-binding January 2026 report is the nearest thing. Ceilings are inherited from function-level rules, which is why ACE classifies per decision. The full register, 36 instruments in four groups (EU, US and international GxP; UK and Japan GxP; cross-cutting standards and security; regulated non-GxP and business regimes), is in Appendix A, overlay v0.1, as of 2026-09-23.
Change-impact procedure. When a rule is new or finalized, four steps follow, and the axes and building blocks stay untouched (the diagram in §4 shows the same four steps entering the overlay layer only):
Register it
Add or update one row: status, scope, owner.
Map it to attributes
Decide which of I, W, R, M it touches. It should rarely change a class definition.
Update the overlay
Change a ceiling cell, add an evidence item, or both. Bump the overlay version (v0.1 → v0.2).
Re-run only the affected assessments
Query the registry for use cases with the touched class × model type × jurisdiction; recompute; promote, hold or demote through the normal gate.
Scheduled and expected events, Q4 2026 to 2028
1 January 2027 touches C0-R and C0-B, not the GxP rows; the final Annex 22 date is targeted, not confirmed. The table beneath is the same list.
| When | Event | Row change | Assessments to re-run |
|---|---|---|---|
| Q4 2026 (targeted [S]) | Final EU GMP Annex 22 | If the GenAI exclusion survives: C3 cell unchanged; HITL evidence items added at A2–A3 for C1–C2 | LLM agents in C1–C2, EU and PIC/S scope |
| ~Dec 2026 (expected [S]) | FDA broadcast-advertising proposed rule | Promotion row: evidence items for broadcast content; C0-R external cap stays A2 | Agents drafting broadcast content |
| 1 Jan 2027 | California automated-decision rules; Colorado AI Act | State-law row: notice, opt-out, human review with authority to change | C0-R and C0-B agents making significant decisions about people |
| 15 Jan 2027 [S] | ICH E6(R3) Annex 2 effective | Data and system fitness evidence extended to decentralized and real-world data sources | Clinical agents ingesting such data |
| Jul 2027 (targeted [S]) | HIPAA Security Rule final | Asset inventory and risk analysis cover AI tools | Agents touching ePHI |
| 2 Dec 2027 | EU AI Act Annex III obligations | Art. 26 oversight, logging and worker notices mandatory | HR agents and other Annex III uses; confirm ceilings ≤ A2 |
| Late 2026–2027 [S] | NIST agent security overlays | Dimension 6 evidence and A3+ security items; no ceiling change | E-profile dimension 6; A3+ agents |
| Date unconfirmed | FDA credibility guidance finalized | Credibility-plan evidence formalized for C2–C3 uses supporting US regulatory decisions | Those use cases only |
| 2 Aug 2028 | EU AI Act Annex I obligations | Oversight evidence for AI in regulated products | Agents in device-regulated scope |
An overlay needs an owner, a budget and a place in the governance the company already runs. §13 turns the model into a four-quarter plan.
Next → §13 · Your Next Four QuartersWhat This Means for Your Next Four Quarters
§12 showed how the overlay changes. This section answers the executive's questions: who owns it, how it fits what we already run, what about the agents our vendors ship, and what it costs.
13.1 Who owns what
proposed RACI Adapt to your governance. R = responsible, A = accountable, C = consulted, I = informed.
| Decision | Quality / QRM owner | CAIO or Agentic Factory lead | Regulatory intelligence | Business / process owner | CSV lead |
|---|---|---|---|---|---|
| C-class of a decision (B2) | A | C | C | R | C |
| Ceiling table and overlay adoption | A | C | R | I | C |
| E-profile scoring | C | A/R | I | C | R |
| Target level and value case | C | C | I | A/R | I |
| Promotion or demotion (B6) | A (Quality vote) | R | I | R (readiness vote) | C |
| Overlay register maintenance (B9) | C | I | A/R | I | C |
13.2 Inside your ISO 42001 AI management system
ACE is not a second management system. ISO 42001 supplies the management system: policy, roles, risk and impact assessment, supplier controls and internal audit, with human oversight "in proportion to risk" but no numeric ceiling. ACE supplies the per-decision content for agents and GxP: the class, the ceiling, the evidence per level and the assignment record. In practice the B2 and B3 records become inputs to the 42001 impact assessment, the overlay register becomes a controlled document under the management system, and the E-profile becomes evidence for management review.
13.3 Supplier and embedded agents
For many companies, agents arrive inside platforms they do not build: CRM, quality and regulatory suites, IT service management, and agents run by CROs and CMOs. You cannot inspect their memory or orchestration, but you can still classify the decision. ACE's approach:
Classify by the W attribute of the vendor tool
and by the obligation triggers it can receive. A vendor agent that can write to your QMS or safety database is at least C2 for that tool, whoever built it.
Evidence it through supplier assessment and a platform contract
identity, tool access, logging and model-version notice as contract terms. The platform article's UC-H1 ("the platform contract", see the glossary) describes one design.
Cap it at the level you can evidence
If the vendor cannot show trajectory or boundary tests, the permitted level stops at A2, whatever the vendor's marketing says.
13.4 Value and effort
Value. Every B3 record carries a value metric with a measured baseline (§8.2). A target with no value case is not a target. Each promotion asks: "which metric moves, and is added autonomy the cheapest way to move it?"
Effort author's estimate, to be tested in pilots. A first E-profile and the classification of an existing inventory of 20–50 agents fits in one to two workshop days plus two to three weeks of evidence gathering. Overlay maintenance is a part-time regulatory-intelligence duty, reviewed quarterly and on each scheduled event in §12.
Registry, classification, first E-profile (§9.3).
B3 records and gate criteria for every GxP agent; overlay adopted by QRM.
Monitoring and demotion triggers live for A3+ agents; first overlay revision after the final Annex 22 text.
Re-profile; fund the dimension that unlocks the most use cases.
Ten Open Questions for Practitioners
The model is only useful if practitioners can break it. Each question is phrased so a QA head, CSV lead or digital leader can answer it in a sentence.
Safin and Balta treat agency as a separate axis; merging it keeps the model simple but may hide risk
Depends on the final Annex 22 text; the extension to GVP and GCP is by analogy
Multinationals need one policy; regulators diverge
Annex 22 §10.5 "every output" vs. reviewer fatigue and automation bias (AI Act Art. 14(4)(b))
The auditability claim rests on this
Benchmarks drove DPMM adoption
Neutral ownership is the strongest adoption lever
The claim that plans need no rewriting rests on it
Self-steering is the design goal; untested
Most agentic value sits outside GxP; the hand-off steps are where inspection risk concentrates
- Is three axes right, or should reversibility or write authority be a fourth explicit axis rather than an attribute and modifier?
- Is the C3 = A1 ceiling for LLM components right, now that ceilings apply per component?
- Should the ceiling differ for FDA-only vs. EU-marketed products?
- Does A3 Stage match how your Quality function wants to review agent actions, or is per-action approval unworkable at volume?
- Which evidence items in §5.1 would your inspectors actually ask for, and which are missing?
- Should the E-profile be self-assessed, assisted or appraised, and would you share anonymized profiles for a benchmark?
- Where should this live: ISPE GAMP CoP, BioPhorum, Pistoia, DIA, or a peer-reviewed journal, and would your organization pilot it?
- Is the stable core / overlay split in the right place? Which overlay rows are missing, especially for MHRA, PMDA and ICH?
- Can your team run the five-step self-assessment in one workshop, and does §9.2 point to the move you would actually make?
- Is the three-way split (C1–C3, C0-R, C0-B) right, and are the boundary-crossing rule and the obligation trigger workable in your registry and in inspections?
How to Contribute and What Happens Next
§14 asked the questions. This section says how answers reach v0.2, and what you get for giving them.
The ask is small and specific. Pick the two or three questions from §14 that touch your role, and answer each in a sentence or a paragraph. Or run §9.3's 30-day step on ten of your own agents and tell us where the classification broke. Either takes under an hour.
Break the draft. Two questions, under an hour.
Answers shape v0.2 (target Q1 2027). Contributors credited if they wish.
Where: a direct message to the author on LinkedIn, or the contact page. Comments are welcome; structured answers are easier to act on.
- Inner ring. 10–15 named reviewers (QA heads, CSV leads, GAMP CoP members, a PV head, an AI platform architect) are being asked the ten questions directly.
- Consortium ring. The draft is offered to ISPE, BioPhorum, Pistoia and DIA as input to their own work, with no expectation of endorsement.
- Public ring. Everyone else, through the channels above.
What you get back
v0.2 will publish the disagreements and what changed, credit contributors who agree to be named, and ship the overlay as a separately versioned register with a changelog. Target date for v0.2: Q1 2027, after the final Annex 22 text is expected.
License
The model text and tables are CC BY 4.0: adapt them, pilot them, publish your changes with attribution.
A product, a certification, or a claim of alignment with any body named here. Please do not publish a single enterprise "ACE score"; it invites the misreading the model exists to prevent.
A model that asks to be challenged has to show what its claims rest on. §16 does that.
Next → §16 · Evidence and ProvenanceEvidence and Provenance
§15 invited challenge. This section separates what is sourced from what is inferred, so challengers know where to aim.
sourced What is sourced
The landscape (§2), the regulatory texts (§3, §6.5 basis column) and the overlay status columns (Appendix A) are sourced and graded: [V] primary source read; [S] secondary source; [U] unverified. On 23 Sep 2026 the following were re-checked against primary sources: the Pharmaceutical Engineering guardrail categories (against the article), FDA's credibility framework (against the FDA draft PDF), and the Pistoia initiative's published outputs.
vendor-reported What is vendor-reported
The production examples in §1 (Novo Nordisk, AstraZeneca, Moderna) are published by vendors or their partners, not by the companies' own disclosures or independent audit.
author's judgment What is the author's judgment
Everything that makes ACE a model: level and class definitions, ceiling A-levels, the modifier and rules, the building blocks, the RACI, the effort estimate, the assumption probabilities and every crosswalk mapping. Where a table mixes sourced and inferred content, a column or note says which is which; Appendix D has the full list.
not read What was not read
The GAMP AI Guide's body (only its published TOC and ISPE's articles about it); several paywalled maturity models. Appendix D lists these and explains the grading.
This is / This is not
| This is | This is not |
|---|---|
| A draft model for deciding how far an agent may act, per decision | A standard, a certification or a regulator's position |
| A default ceiling a QRM owner adopts, tightens or relaxes with justification | A ceiling to copy verbatim into an SOP |
| A profile that names the dimension that binds | A single enterprise maturity score |
| An open invitation to break it (CC BY 4.0) | A product, or endorsed by any body named here |
What this validates / What remains to be proven
| What the evidence supports today | What remains to be proven |
|---|---|
| The fragments exist and do not compose (§2, sourced) | That practitioners find the three axes and five classes usable in a workshop (§14 Q9) |
| Draft Annex 22 and FDA's draft framework imply a ceiling that falls with criticality (§3, primary texts) | That the ceiling cells survive the final Annex 22 text and QRM review |
| The rule tables are internally consistent: every worked-example row is reproduced by Appendix C (and by this page's calculator self-test) | That the enablement thresholds predict safe operation; no pilot has run yet |
| Public production agents cluster at read and draft levels (vendor-reported) | Whether higher autonomy pays in pharma; no survey maps value to level |
ACE draws on two earlier articles in this series, which show how the evidence for each level can be produced in practice:
- Enterprise Agentic AI Platform for Life Sciences: the agent registry and card (§6), the Lab → Factory promotion gate (§7.3), confidence gating (G7), regional compliance as configuration (G8), the platform contract (UC-H1).
- Validation Strategy, Applied: Agentic AI in GxP: Pattern A and B (§9), the eight lifecycle controls (§8), the template pack (T1–T10), monitoring policy (§13) and change envelopes (§14).
These are illustrations of the building blocks, not the offer.
Conclusion and Contact
Every section has returned to one idea. Here it is, stated once more, with what to do next.
Consensus is forming in fragments: autonomy is a design decision, oversight is in, on or in command of the loop, and the ceiling falls with criticality and non-determinism. ACE puts those ideas into one instrument an organization can run on itself, and adds what they leave open: a versioned ceiling, the hand-off rule, and the knowledge that commercial, corporate and research work can and should run higher than GxP work.
That changes what maturity should mean.
A mature organization is not the one running the most autonomous agents. It is the one that can show, for every agent, why its autonomy is exactly as high as the evidence and the decision allow, and can lower it the day the evidence weakens.
Regulation will keep changing the rows. It should not have to change the plan. The spine, once more: right-sized autonomy. Maturity is the ability to justify autonomy, and to lower it. Regulation changes rows, not plans.
Three things you can do this month
Run §9.3's 30-day step on ten of your agents
Registry entry, classification, obligation triggers (§9.3).
Fill in one B3 record for the GxP agent closest to its promotion decision
The sample in §10.2 is the template; the calculator in §8 exports one.
Answer two of the ten questions in §14 and send them back
Under an hour; the route is in §15.
Start with two questions.
Not a sales call: a draft that asks to be broken, and a route for the answers. Message on LinkedIn, use the contact page, or scan the codes.

linkedin.com/in/nitinbhatti · orchestraprime.ai
North Brunswick, NJ


References
Primary sources first; secondary sources are cited inline where used. Grading as in Appendix D: [V] primary source read · [S] secondary · [U] unverified.
GxP guidance and industry standards
- ISPE. GAMP® Guide: Artificial Intelligence · July 2025 · Official table of contents (PDF) [V] TOC onlyThis article draws on the table of contents and ISPE's articles only.
- ISPE D/A/CH Affiliate AI working group. "AI Maturity Model for GxP Application: A Foundation for AI Validation." · Pharmaceutical Engineering, Mar/Apr 2022
- Stoop, R. and Walia, G. "Transforming Pharmaceutical Quality with GAMP® AI Guardrails." · Pharmaceutical Engineering, Jul/Aug 2026 [V]
- ISPE. Baseline Guide Vol. 8: Pharma 4.0, 1st ed. · Dec 2023
- BioPhorum. Digital Plant Maturity Model 3.0 · Nov 2023
- European Commission. EudraLex Vol. 4, Annex 22: Artificial Intelligence (draft for consultation) · 7 Jul 2025 · PDF [V] draft
- FDA. Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products (draft guidance) · Jan 2025 · PDF [V] draft
- FDA and EMA. Guiding Principles of Good AI Practice in Drug Development · 14 Jan 2026 · EMA PDF [V]
- CIOMS Working Group XIV. Artificial Intelligence in Pharmacovigilance · Dec 2025 · Working group page · Drug Safety summary article [S]
- ICH. Q9(R1) Quality Risk Management; E6(R3) Good Clinical Practice · ich.org [S]
- IMDRF. Software as a Medical Device: Possible Framework for Risk Categorization (N12) · 2014 · PDF
Cross-industry maturity and autonomy frameworks
- CMMI Institute / ISACA. CMMI AI Maturity Model · Jul 2026 · Announcement
- AWS. Levels of autonomy; Agentic AI Security Scoping Matrix; graduated autonomy. · 2025–2026 · Insights · Security · Architecture
- Cloud Security Alliance. Levels of Autonomy (Jan 2026); Agentic Governance Maturity Model (Mar 2026). · Blog · AGMM
- Capgemini Research Institute. Rise of agentic AI · Jul 2025 · PDF
- NIST. Cybersecurity Framework 2.0 (CSWP 29) · 2024 · PDF
- Pistoia Alliance. Agentic AI initiative announcement · Sep 2025 [V]
Research and methods
- Feng, K. J. K., McDonald, D. W. and Zhang, A. X. "Levels of Autonomy for AI Agents." · Knight First Amendment Institute, Jul 2025
- Safin and Balta. Autonomy vs. agency in AI agents · arXiv 2605.12105
- PHUSE US 2026, paper ML04. Validated orchestration of bounded agents · lexjansen.com
- Pöppelbuß, J. and Röglinger, M. "What makes a useful maturity model?" · ECIS 2011
- Becker, J., Knackstedt, R. and Pöppelbuß, J. "Developing Maturity Models for IT Management." · BISE, 2009
- Anthropic. "Measuring agent autonomy." · anthropic.com
Regulatory Overlay Register v0.1
As of 2026-09-23 for every group. Status cells are sourced and graded [V] / [S] / [U] (Appendix D). The implication, trigger and owner columns are the author's inference. Changelog at the end of the article.
Appendix AOverlay register: 36 instruments in four groupsA · GxP: EU, US and international (13) · B · GxP: UK and Japan (7) · C · Cross-cutting standards and security (4) · D · Regulated non-GxP and business (12). Filter by group, grade, owner or text.
| Group | Instrument | Status | Ceiling and evidence implication | Trigger to watch | Owner |
|---|---|---|---|---|---|
| A | EU GMP Annex 22 (AI), EMA with PIC/S | Draft 7 Jul 2025; consultation closed 7 Oct 2025 [V]; final targeted Q4 2026 [S] | LLM, dynamic and probabilistic: C3 ≤ A1; C1–C2 need qualified HITL. Evidence: intended use and input space; acceptance "at least as high as" the replaced process; "undecided" outcome; change and drift control | Final text; whether the GenAI exclusion survives; PIC/S, MHRA, PMDA adoption | GMP Quality / CSV |
| A | EU GMP Annex 11 rev. and Chapter 4 | Draft Jul 2025 [V]; final expected 2026 [S] | "Never increase overall risk" sets the A4 test; audit-trail review, e-signature, hybrid records | Final text | CSV / Data integrity |
| A | FDA draft guidance: AI to support regulatory decision-making | Draft Jan 2025 [V] | Model influence ≈ A-level; decision consequence ≈ C-class; 7-step credibility plan per context of use | Final guidance (date unconfirmed) | Regulatory Affairs / Quality |
| A | FDA Computer Software Assurance | Final Sep 2025 [V] | Lighter, risk-based assurance for C1; supplier evidence | Pharma-specific adoption statements | CSV |
| A | FDA–EMA Guiding Principles of Good AI Practice | Issued 14 Jan 2026; non-binding [V] | No numeric ceiling; context of use and lifecycle documentation | Conversion into formal guidance | Regulatory intelligence |
| A | EU AI Act Art. 14 (as amended by Reg. 2026/1744) | Annex III high-risk from 2 Dec 2027; Annex I from 2 Aug 2028 [V/S] | A4+ must support override and stop; oversight design; logging | 2 Dec 2027; Commission guidance on pharma high-risk uses | AI governance / Legal |
| A | EMA Reflection Paper; network data-science workplan 2026–28 | Paper final Sep 2024; workplan Feb 2026 [V] | No numeric ceiling yet; AI-in-clinical-development and AI-in-PV guidance planned | New EMA drafts | Regulatory intelligence |
| A | ICH E6(R3) and Annex 2 | Principles and Annex 1 effective in EU Jul 2025 [V]; Annex 2 effective 15 Jan 2027 [S] | Sponsor accountability cannot be delegated to an agent; data and system fitness extended to decentralized and real-world data | 15 Jan 2027 | Clinical QA |
| A | ICH Q9(R1) | Final (2023) [S] | The method for setting the ceiling: QRM formality scales with autonomy and GxP impact; QRM file per agent use case | ICH AI topic (none found) | QA (QRM owner) |
| A | ICH M15 (model-informed drug development) | Step 4, Jan 2026 [S] | Credibility framing for model-based evidence in submissions (C2–C3) | Regional implementation | Regulatory Affairs |
| A | PIC/S | Co-drafted Annex 22 [V] | Follows the Annex 22 row | Adoption of final Annex 22 | GMP Quality |
| A | CIOMS WG XIV (AI in PV) | Final Dec 2025 [S] | HITL / HOTL / HIC vocabulary; high oversight at launch, relaxed on proven performance | Adoption into GVP | QPPV office |
| A | ISPE GAMP AI Guide and 2026 guardrail article | Guide Jul 2025 [V, TOC only]; article Jul–Aug 2026 [V] | Directional; guardrail categories 0–3 in quality workflows | Second edition or agent guidance | CSV / GAMP liaison |
| B | MHRA Inspectorate blog: AI for GxP inspection responses | Published 29 Jun 2026; guidance-style [V] | AI drafts only; human technical review and accountable sign-off (by analogy, for any AI-drafted GxP record) | Extension into formal guidance | QA / RA |
| B | MHRA GxP data-integrity guidance; MHRA AI policy paper | DI 2018 [U]; policy Apr 2024 [S] | ALCOA+ for AI inputs and outputs; records attributable to accountable people | Update; formal UK Annex 22 position (none found) | QA / CSV |
| B | MHRA AI Airlock | Phase 2 report Jun 2026 [V] | Device scope only | Phase 3 cohort | RA (devices) |
| B | UK Clinical Trials Regulations | In force Apr 2026 [S] | No AI-specific clause found | MHRA AI-in-trials guidance | Clinical QA |
| B | PMDA AI Action Plan (internal) | Released Oct 2025 [V/S] | Signal only: final regulatory decisions remain with human staff [S] | PMDA guidance to industry | RA (Japan) |
| B | Japan AI Promotion Act and AI Basic Plan | In force 2025; plan revised Jul 2026 [S, single source] | Soft law; no ceiling | MHLW sector guidance | Legal / RA (Japan) |
| B | MHLW GMP/GQP/GVP on AI | Nothing found in English sources | Expect Annex 22 via PIC/S | PIC/S adoption in Japan | QA (Japan) |
| C | ISO/IEC 42001 (with 42005, 42006) | 2023–2025 [S] | No numeric ceiling; management-system layer above CSV (§13.2); impact assessments, supplier clauses, override logs | Certifications in large pharma (none found) | IT/Digital with QA |
| C | ISO/IEC 22989 (with 23894, TS 8200) | 2022–2024 [S/U] | Vocabulary for axis A (automation vs. autonomy) | Verbatim automation-level table (paywalled) | CSV / Architecture |
| C | NIST agent security overlays (SP 800-53); AI RMF | In development [S] | Raises dimension 6 evidence; no ceiling change | Draft publication | CISO |
| C | OWASP Top 10 for Agentic Applications | Published Dec 2025 [V] | Least agency: permissions no wider than the task | Annual revision | CISO |
| D | FDA OPDP/APLB promotion (21 CFR 202.1; Form 2253) | In force; heightened enforcement since Sep 2025 [S]; no AI-specific promotional guidance found | External promotion ≤ A2: draft, pre-check, route; human MLR approval and 2253 submission before dissemination; patient chatbots only on locked, pre-approved responses | Broadcast-advertising proposed rule (~Dec 2026); any letter citing AI content | Commercial compliance / MLR |
| D | ABPI Code / PMCPA AI Q&A (UK) | Revised Jun 2025 [V] | Certification "cannot be fulfilled solely by the use of AI": human signatory; ≤ A2 | Further PMCPA guidance | UK signatories |
| D | PhRMA, EFPIA codes; IFPMA AI principles | No AI text in codes; EFPIA monitoring [S] | No autonomous transfers of value | Code revisions | Ethics and Compliance |
| D | HIPAA and Security Rule proposal | Final targeted Jul 2027 [S] | Minimum necessary; business-associate agreements for AI vendors; no free-form PHI in general LLMs | Final Security Rule | Privacy / InfoSec |
| D | ACA Section 1557 §92.210 | Duties from May 2025 [U] | Bias assessment of patient-care decision tools where a covered entity | Enforcement posture | Privacy / Legal |
| D | US state AI laws | CA automated-decision rules and Colorado AI Act from 1 Jan 2027; Texas from 1 Jan 2026 [S] | CA: opt-out or human reviewers who can change the decision; CO, TX: disclosure | 1 Jan 2027; federal preemption activity | Privacy / HR / Legal |
| D | GDPR Art. 22 and 35 (SCHUFA; EDPB Op. 28/2024) | In force [U for case-law details] | Meaningful human decision-maker; a rubber-stamped recommendation can itself be the "decision"; decisions about individuals ≤ A2 unless an exception applies; DPIA first | GDPR omnibus negotiations | DPO |
| D | EU Data Act / EHDS | Phased to ~2029 [U] | Limits where agents may source data | Implementing acts | Data governance |
| D | SOX, COSO GenAI guidance, PCAOB | COSO guidance Feb 2026 [S] | Agent controlled as a user: A4 below materiality, A2 above; no agent both creates and approves | PCAOB findings on AI | Controller / Internal audit |
| D | EU AI Act, non-GxP uses | Art. 50 from Aug 2026; Annex III from 2 Dec 2027 [S] | Chatbots disclose AI; HR agents ≤ A2 with Art. 26 oversight | Art. 50 guidelines; 2 Dec 2027 | AI governance / HR |
| D | UK ICO Tech Futures: Agentic AI | Jan 2026, non-binding [V/S] | Direction of travel: data-flow maps, DPIAs covering tool use | ICO agentic guidance | DPO (UK) |
| D | FTC Section 5 | AI-accuracy policy statement proposed Jul 2026 [V] | Consumer-facing agent claims substantiated | Final statement | Legal |
Crosswalk to Frameworks Readers Already Use
author's judgment All relationship judgments are the author's.
Appendix BCrosswalk: how ACE v0.1 relates to twelve frameworks12 rows: GAMP AI Guide, A-level anchors, the 2022 D/A/CH matrix, the 2026 guardrail categories, DPMM and Pharma 4.0, CMMI AIM, AWS, CIOMS XIV, FDA's credibility framework, draft Annex 22 and 11, EU AI Act Art. 14 with OWASP and NIST, ISO/IEC 42001.
| Framework | How ACE v0.1 relates | Relationship |
|---|---|---|
| ISPE GAMP AI Guide (lifecycle; M9, M10 per TOC) | Phases and gates used unchanged as the lens (§11.1). M10's per-sub-system autonomy and adaptiveness, as named in the TOC, is kept as the per-component classification: adaptiveness feeds the model-type column; the A-scale adds action autonomy, write authority, tools, planning, memory and orchestration. Content of M10 not reviewed TOC only | Framed through; extends for agents |
| A-level anchors | A1 ≈ AWS Scope 1, CSA L1, ISPE CD1 "parallel to GxP process"; A2 ≈ AWS Scope 2, ISPE CD2 "actively approved", Annex 22 HITL; A3 ≈ AWS Scope 2–3, CSA L2; A4 ≈ AWS L3 / Scope 3, CSA L3, ISPE CD3; A5 ≈ AWS L4 / Scope 4, CSA L4–5, Capgemini L4–5. E-levels loosely follow DPMM (silos → adaptive) and CMMI (initial → optimizing) | Maps to |
| 2022 D/A/CH matrix | Control Design stages 1–5 map roughly to A1–A5; its validation levels remain the reference for learning systems | Extends |
| 2026 GAMP AI guardrail categories | RPN categories map to per-decision class handling inside a workflow (Category 0 "advisory only" ≈ an A1–A2 ceiling) | Consistent; generalizes |
| BioPhorum DPMM / ISPE Pharma 4.0 | E1–E5 follow the silos → adaptive arc; a DPMM score is an input to dimension 2 for manufacturing | Complements |
| CMMI AI Maturity Model | Appraisal results can evidence the E-profile; ACE adds the GxP ceiling and per-use-case autonomy | Maps to |
| AWS levels, scoping matrix, graduated autonomy | A-levels align with scopes 1–4; security dimensions inform dimension 6; promotion and demotion become the operating rule | Borrows; adds ceiling |
| CIOMS XIV | Human-role column of axis A, extended from PV to all functions | Adopts vocabulary |
| FDA draft credibility framework | Model influence ≈ A-level; decision consequence ≈ C-class | Operationalizes |
| Draft EU GMP Annex 22 and Annex 11 | C3 ceiling for LLMs, HITL at C1–C2, the "no decrease" test at A4 | Encodes as ceiling |
| EU AI Act Art. 14; OWASP least agency; NIST AI RMF | Override and stop at A4+; least-agency permissions; current-vs-target profiles | Controls vocabulary |
| ISO/IEC 42001 | Management system around ACE's per-decision content (§13.2) | Sits inside |
Calculator Rule Tables and Truth Table
These tables are the complete, table-driven specification of the ACE calculator. The calculator in §8 implements exactly these tables, and its on-load self-test checks every worked-example row against them.
Appendix CRules R1–R5, the truth table and the normative rule codeR1 class from attributes (5 rows) · R2 ceiling by class and model type (10 rows) · R3 reversibility modifier · R4 enablement (5 rows) · R5 result · truth table for uniform profiles (5 × 6 cells) · rule code as data tables.
R1. Class from attributes
Take the highest class that any row implies.
| Condition | Class |
|---|---|
| I = Direct, or W = commit to a primary GxP record, or W = external regulatory submission | C3 |
| I = Indirect, or W = staged change to a GxP record, or an obligation trigger can be received (detection step only) | C2 |
| I = Supporting | C1 |
| I = None and a named non-GxP regime applies | C0-R |
| I = None and no named regime applies | C0-B |
R2. Ceiling by class and model type
Default overlay v0.1.
| Class | LLM | Deterministic / static ML | Dynamic |
|---|---|---|---|
| C0-B | A4 (A5 if sandbox) | A4 (A5 if sandbox) | A4 (A5 if sandbox) |
| C0-R · external promotional release | A2 | A2 | A2 |
| C0-R · decision about an individual | A2 | A2 | A2 |
| C0-R · finance above materiality | A2 | A2 | A2 |
| C0-R · internal suggestion to staff | A3 | A3 | A3 |
| C0-R · finance below materiality or into suspense | A4 | A4 | A4 |
| C0-R · other internal reversible work | A4 | A4 | A4 |
| C1 | A3 | A4 | A2 |
| C2 | A2 | A3 | A1 |
| C3 | A1 | A3 | A1 |
R3. Reversibility modifier
If R = irreversible, class ≠ C3 and the C0-R sub-type is not "external promotional release": ceiling = max(A1, ceiling − 1). Otherwise unchanged. Applied once.
R4. Enablement
Highest level whose row, and every row above it, is satisfied.
| Level | D1 Value | D2 Data | D3 Registry | D4 Assurance | D5 Oversight | D6 Security | D7 Monitoring | D8 People | Scope |
|---|---|---|---|---|---|---|---|---|---|
| A1 | – | – | E2 | – | – | – | – | – | All |
| A2 | – | – | E2 | E2 | E2 | – | – | – | All |
| A3 | – | – | E2 | E3 | E3 | – | E3 | – | All |
| A4 | E3 | E3 | E3 | E3 (E4 if C1–C3) | E4 | E4 | E4 | E3 | All |
| A5 | E4 | E4 | E4 | E4 | E5 | E5 | E5 | E4 | C0-B only |
R5. Result
Permitted = min(ceiling after R3, enablement), taking the lowest ceiling across jurisdictions. If target ≤ permitted: operate at target. Else if ceiling < enablement: capped by ceiling. Else if enablement < ceiling: capped by capability. Else: capped by both.
Truth table: uniform profiles, LLM component, default ceilings, reversible, not sandbox
| Profile (all dimensions at) | Enablement (non-GxP / GxP) | C0-B | C0-R internal reversible | C0-R external or individual | C1 | C2 | C3 |
|---|---|---|---|---|---|---|---|
| E1 | A0 / A0 | A0 | A0 | A0 | A0 | A0 | A0 |
| E2 | A2 / A2 | A2 | A2 | A2 | A2 | A2 | A1 |
| E3 | A3 / A3 | A3 | A3 | A2 | A3 | A2 | A1 |
| E4 | A4 / A4 | A4 | A4 | A2 | A3 | A2 | A1 |
| E5 | A5 (C0-B only) / A4 | A4 (A5 in sandbox) | A4 | A2 | A3 | A2 | A1 |
The worked example in §8.2 is the test set for any implementation: an implementation of R1–R5 must reproduce every row. This page's calculator runs those eleven rows and the truth table as a self-test on load and reports the result under the calculator.
The rule tables as code (normative; what the calculator runs)
const LEVELS = ["A0","A1","A2","A3","A4","A5"]; // ordinal in delegated authority
const CLASSES = ["C0-B","C0-R","C1","C2","C3"]; // ascending
// R1: class from attributes (take the highest implied)
const CLASS_BY_I = { supporting:"C1", indirect:"C2", direct:"C3" };
const RAISE_BY_W = { staged:"C2", commit:"C3", submission:"C3" };
function classOf({I, W, regime, trigger, detectionStep}) {
let c = I==="none" ? (regime ? "C0-R" : "C0-B") : CLASS_BY_I[I];
const raise = x => { if (CLASSES.indexOf(x) > CLASSES.indexOf(c)) c = x; };
if (W==="staged" && I!=="none") raise(RAISE_BY_W.staged);
if (W==="commit" || W==="submission") raise(RAISE_BY_W[W]);
if (trigger && detectionStep) raise("C2"); // obligation trigger (§6.4)
return c;
}
// R2: default ceiling, overlay v0.1
const CEILING = {
"C0-B": {llm:"A4", det:"A4", dyn:"A4", sandbox:"A5"},
"C0-R": { "external-promotion":"A2", "individual-decision":"A2", "finance-above":"A2",
"internal-suggestion":"A3", "finance-below":"A4", "internal-reversible":"A4" },
"C1": {llm:"A3", det:"A4", dyn:"A2"},
"C2": {llm:"A2", det:"A3", dyn:"A1"},
"C3": {llm:"A1", det:"A3", dyn:"A1"}
};
// R3: reversibility modifier — once, floor A1; not in C3; not for external-promotion
function applyR(ceiling, cls, subtype, irreversible) {
if (!irreversible || cls==="C3" || subtype==="external-promotion") return ceiling;
return LEVELS[Math.max(1, LEVELS.indexOf(ceiling) - 1)];
}
// R4: enablement — minimum E-level per dimension (1..8), cumulative
const REQ = {
A1: {3:2},
A2: {3:2, 4:2, 5:2},
A3: {3:2, 4:3, 5:3, 7:3},
A4: {1:3, 2:3, 3:3, 4:3, 5:4, 6:4, 7:4, 8:3}, // GxP (C1–C3): 4 -> 4
A5: {1:4, 2:4, 3:4, 4:4, 5:5, 6:5, 7:5, 8:4} // C0-B only
};
function enablement(profile, cls) {
let best = "A0";
for (const L of ["A1","A2","A3","A4","A5"]) {
if (L==="A5" && cls!=="C0-B") break;
const req = {...REQ[L]};
if (L==="A4" && ["C1","C2","C3"].includes(cls)) req[4] = 4;
const ok = Object.entries(req).every(([d,min]) => profile[d] >= min);
if (!ok) break; // cumulative: stop at first failure
best = L;
}
return best;
}
// R5: result
function result({ceiling, enable, target}) {
const i = LEVELS.indexOf.bind(LEVELS);
const permitted = LEVELS[Math.min(i(ceiling), i(enable))];
if (i(target) <= i(permitted)) return {permitted, label:"Operate at target"};
if (i(ceiling) < i(enable)) return {permitted, label:"Capped by ceiling"};
if (i(enable) < i(ceiling)) return {permitted, label:"Capped by capability"};
return {permitted, label:"Capped by both"};
}
Methodology and Evidence Grading
How the model was built, how sources are graded, what was not accessed, and what is inferred rather than found.
Appendix DMethod, evidence tags, re-checks, what was not accessed, what is inferred, open verification itemsSix paragraphs. The full list of inferred content lives here; §16 summarizes it.
Purpose and method. ACE is both descriptive (where are we?) and prescriptive (how far may this agent go, and what next?). Its design follows Becker et al.'s procedure model for maturity-model development and Pöppelbuß and Röglinger's design principles, both cited in §18: problem definition (§1–§3), comparison with existing models (§2), iterative design (this v0.1), and evaluation through practitioner review and pilots with three to five organizations before v0.2. It is intended to pre-empt the standard critique that maturity models reflect one person's view.
Evidence tags. [V] primary source read; [S] secondary source (law firm, trade press, consultancy, or a search summary of a primary); [U] unverified. Inline citations without a tag point to the primary source.
Re-checked on 23 Sep 2026. Pharmaceutical Engineering guardrail categories and thresholds (against the article); FDA credibility framework (against the draft PDF: 7 steps; model influence and decision consequence); Pistoia initiative outputs (no maturity or autonomy scheme found).
Not accessed. The GAMP AI Guide body, including Appendices M9 and M10 and the model-category appendices; ISPE Pharma 4.0 level names; Gartner's agentic roadmap; Axtria's stage names; the ISO/IEC 22989 automation-level table. Every statement about the GAMP AI Guide's coverage is limited to its published table of contents and ISPE's articles. The correspondence between M10's headings and the 2022 D/A/CH axes, and the reading that M10 assesses each AI sub-system, are inferences from headings.
Inferred, not found. The three-way split, the C0-R / C0-B classes, the boundary-crossing rule and obligation trigger, the per-component rule, the attribute-based class definitions, every ceiling cell's A-level, the reversibility modifier, the enablement rule, the confidence-band mapping, the level and dimension names, the building blocks, the self-assessment and "where am I" table, the ACE columns of the GAMP lens, the agent-property controls, all crosswalk relationships, the overlay implication, trigger and owner columns, the RACI, the effort estimate, and the Strategic Planning Assumption probabilities. Survey figures on agentic adoption are self-reported and sensitive to question framing.
Open verification items for v0.2. Final Annex 22 text; confirmation of the Colorado amending legislation and the Japan AI Basic Plan revision against primary sources; the ISO/IEC 42001 Annex A numbering for human oversight; whether §92.210 is being enforced.
| Version | Date | Change |
|---|---|---|
| 0.1 | 2026-09-23 | First public draft for industry feedback. Model named ACE. Overlay register v0.1 (Appendix A). Calculator rules published (Appendix C). |