Skip to content
Innovation Insight × Reference Architecture · Life Sciences

Enterprise Agentic AI Platform for Life Sciences — Solution Architecture

OrchestraPrime · Nitin Bhatti, Founder & Platform ArchitectPublished September 2026~28 min read · role paths ~10–15 minVersion 2.1 (Medical & Safety tower expanded, September 17, 2026; supersedes v1, September 8, 2026)

One sentence. Every life-sciences company is running dozens of agentic AI initiatives across research, development, manufacturing, supply chain, commercial, and medical; almost all of them are being built as disconnected projects, and the cost of that fragmentation is measured not only in money but in decision latency — the weeks a protocol sits unamended, the months an adverse quality trend goes unseen, the days a batch waits for release.

This article describes the architecture for building those initiatives as one governed, validated, production-grade platform where every agent provides intelligent decision support — the human decides, the system explains — and where the fastest decision and the best decision become the same decision. It is published as thought leadership because we believe the patterns are right for the industry. But the real question — the one this article exists to help you think through — is not whether this architecture makes sense. It is whether your organization has the operating model, the semantic discipline, the validation posture, and the leadership to build it. That is a people question, not a technology question, and it is the question the rest of this article is designed to help you answer.

~$500K
per protocol amendment
Roughly half are avoidable with precedent intelligence up front. Months of delay each.
6–15 mo
APQR signal latency
A trend that should have triggered an investigation in March appears in next February's report.
15 days
the PV clock
Intake across five channels, manual coding, narrative drafting — a hard regulatory deadline.
137×
frontier vs scoped model cost
Claude 3.5 Sonnet vs. Haiku, bounded extraction/scoring tasks, Q3 2025, prompt caching enabled. A monolith routes everything to the frontier tier.
$2–5M
per standalone use-case
Twenty use-cases as twenty projects is a consulting program, not a platform.
Bottom Line

Not a technology project — a decision-velocity operating model

An enterprise agentic AI platform for life sciences is not a technology project — it is a decision-velocity operating model built on a shared semantic substrate, a composition-first agent framework, and a cross-cutting governance plane that makes the fastest decision and the best-governed decision the same decision. Organizations that build this as forty disconnected projects will spend 2–5× more, take 2–3× longer, and end up consolidating to a platform approach anyway.

Key Findings

Five findings that reorder the priorities

1

Decision latency is itself a compliance risk — and the industry under-weights it.

Slow APQR signals, late CPV detection, delayed protocol amendments, and missed PV clocks all carry the same regulatory exposure as wrong decisions — by a different route. The platform must be measured on decision velocity as a first-class metric alongside accuracy.

2

Composition over inheritance is an economic imperative, not a design preference.

A monolithic agent with 100 tools costs ~137× more per decision than composed bounded specialists, while also being unsecurable, unvalidatable, and impossible to promote from Lab to Factory. The unit of deployment is the ≤5-tool specialist; the unit of reuse is the service kit.

3

The semantic layer is the platform's durable moat.

Individual agents are replaceable; models deprecate quarterly. The federated knowledge graph — thin core ontology, domain-owned extensions, entity resolution, process intelligence — is the asset that compounds. It is also the single largest source of underestimation: the canonical attribute dictionary over regional LIMS instances is where most manufacturing AI programs stall.

4

GxP validation of hybrid AI systems is solvable today — with GAMP 5 Category 5 for deterministic paths, CSA-aligned controls for LLM paths, and analytical-method lifecycle for statistical/chemometric models.

The posture is defensible but not yet established precedent; early movers who articulate it to Quality before the first agent enters a GxP workflow will set the standard.

5

The hardest problem is organizational, not technical.

Hiring FDEs and knowledge engineers, earning Quality's trust, resisting the pressure to scale before the foundation is proven, and killing PoCs that don't move the metric — these are the risks an ED is accountable for. The platform reduces the technical barriers so the leadership problems become the binding constraint.

Strategic Planning Assumptions

Five planning assumptions, with probability and substantiation

By 2028Probability 0.70

By 2028, 60% of top-20 pharma companies that attempt enterprise agentic AI without a shared semantic substrate will consolidate to a platform approach — after spending 2–3× more than organizations that started with one.

Substantiation
Gartner's 2024 "Top Strategic Technology Trends" identifies AI-ready data foundations as a prerequisite for scaling GenAI; McKinsey's 2024 pharma AI survey found 70% of pilot programs stall at scale due to fragmented data and governance; Novartis's own Data42 and NERVE investments signal the same consolidation pattern from within the industry.
By 2027Probability 0.65

By 2027, the industry's default AI validation posture will converge on the hybrid model described here (GAMP 5 Cat 5 for deterministic paths + CSA-aligned controls for LLM paths), driven by FDA finalization of the 2025 draft AI guidance and EMA's adopted reflection paper (September 2024), and the ISPE GAMP Guide: AI (July 2025). Organizations that retrofitted validation will carry 12–18 months of rework.

Substantiation
FDA's January 2025 draft guidance "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products" (CDER/CBER) establishes risk-based credibility frameworks for AI in drug development; ISPE's GAMP 5 2nd Edition (2022) already accommodates agile and CSA approaches; ISPE GAMP Guide: Artificial Intelligence (July 2025, 290 pp) provides the industry's first comprehensive AI validation framework; PIC/S guidance PI 041-1 (2021) establishes data integrity expectations that map directly to the governance plane's G1–G3 controls; FDA–EMA "Guiding Principles of Good AI Practice in Drug Development" (January 2026) reinforce risk-proportionate validation aligned with intended context of use.
By 2029Probability 0.60

By 2029, "Chief Agentic Officer" or equivalent will be a named role at 40%+ of top-20 pharma companies, owning the Lab, Factory, and registry described here. The role will be measured on agent graduation rate, decision velocity, and reuse ratio — not on number of agents deployed.

Substantiation
The Chief AI Officer role appeared at 25%+ of Fortune 500 companies by 2024 (Deloitte AI Institute); pharma-specific variants (Head of AI/ML at Roche, VP AI Platform at Novartis) already exist; the agentic shift — from model deployment to agent lifecycle management — creates the operational mandate that the current CAIO title does not cover.
Through 2028Probability 0.75

Through 2028, organizations that measure agentic AI programs on "number of use-cases deployed" rather than decision-velocity improvement and domain-expert acceptance rate will report program dissatisfaction at 3× the rate of those using outcome metrics.

Substantiation
Gartner's 2024 "AI Value Realization" research found that organizations using outcome-based AI metrics reported 2.5× higher satisfaction than those using activity metrics; BCG's 2024 pharma AI survey identified "measuring the wrong things" as the #2 reason for AI program failure after data quality.
By 2028Probability 0.65

By 2028, Medical Affairs will be the function where agentic AI reaches governed production fastest in top-20 pharma — ahead of R&D and manufacturing — because its decisions (inquiry response, MSL preparation, insight synthesis, congress intelligence) are high-volume, text-native, and governed by approved-source discipline rather than GMP validation. But organizations that build those agents on a commercial CRM vendor's agent platform without an architecturally enforced medical/commercial firewall and a canonical interaction model will rebuild them at least once — on CRM migration, on a scientific-exchange finding, or both.

Substantiation
Peer-reviewed Medical Affairs AI proof points already exist inside top-20 companies — semantic clustering of unsolicited medical-information inquiries (PMC, 2023) and GenAI theme analysis of field medical insights (Pharmaceutical Medicine, 2024) — while the vendor layer is moving simultaneously: Agentforce Life Sciences reached general availability in October 2025 with pharma early adopters, Veeva shipped AI agents for PromoMats in December 2025, and Veeva CRM's end-of-support date (December 31, 2029) forces every Veeva CRM customer through a Vault CRM vs. Salesforce Life Sciences Cloud decision inside this window. The industry is therefore standing up Medical Affairs agents during a CRM transition — which is precisely when binding agents to CRM objects rather than a canonical model is most expensive.
Recommendations

By persona — what to do in Phase 1, and what not to measure

For the Chief Data & AI Officer / Head of AI Platform Strategy:
  • Publish the agent platform contract (identity, MCP tool access, observability, registry) before the second agent estate arrives — partner-built agents will deploy around the platform if the contract is not ready. Start in Phase 1.
  • Measure the platform on marginal cost per use-case (should fall 40–60% by use-case 5–6), reuse ratio (target ≥70%), and time-to-first-grounded-answer for new use-cases (target: days). Do not measure on number of agents deployed.
  • Fund the semantic layer as infrastructure, not as a project. The core ontology and KG-as-a-Service are consumed by every use-case; starving them starves everything.
For the Agentic Lab & Architecture Lead:
  • Enforce composition from day one — ≤5 tools per specialist, orchestrators that compose rather than inflate. The 137× cost difference compounds with every agent that violates the rule.
  • Build kill criteria into every PoC: if precision/recall/F1 does not improve over two iteration cycles, or coach-mode acceptance is flat, stop. Killing fast is a Lab success metric. Publish post-mortems as deprecated registry cards.
  • Staff the Lab with FDEs who speak the domain. The gap between "what the toxicologist needs" and "which specialist, which kit, which data product" is the gap the Lab exists to close.
For the Head of Agentic Factory:
  • Own the promotion gate with defined criteria, named reviewers, and an escalation SLA. The gate is the hardest organizational problem; without it, PoCs accumulate.
  • Build the support model (L1/L2/L3) before the first agent graduates — not after. L1 is domain super-users; L2 is platform SRE; L3 is pod re-formation. The model lifecycle (quarterly successor-model testing) is L2 work that most organizations forget.
  • Treat the Enterprise AI Operations Command Center as your network operations center. If you cannot see drift, cost, and acceptance-rate curves for every factory agent in one surface, you do not have a Factory.
For the Head of Semantic & Knowledge Engineering:
  • The federated ontology model — thin core, domain-owned extensions under a published modeling standard — is the only design that avoids both the central bottleneck and the disconnected-graph swamp. Staff the core thinly (3–5 ontologists); staff the domain extensions as part of each pod.
  • Entity resolution is a first-class service, not a nice-to-have. The same golden-record discipline commercial has applied to HCPs for years must extend to sites, investigators, batches, and materials. Site selection is a master-data problem before it is a prediction problem.
  • The canonical attribute dictionary over regional LIMS instances and CMO feeds is the single most underestimated work item in the platform. Budget for it explicitly; it is quality-owned work under change control.
For the Medical Affairs AI Platform Lead (product, architecture, or engineering):
  • Abstract the interaction model before the first agent ships. Bind agents and analytics to a canonical Medical Affairs data model — interaction, inquiry, insight, content, HCP, event, transfer of value — materialized from whichever CRM the persona works in (Veeva CRM today; Vault CRM or Salesforce Life Sciences Cloud tomorrow). The migration becomes a connector swap, not an agent rewrite. §8.7 describes the model.
  • Buy the vendor agents where they exist; build only where coverage stops. Vendor pre-review and quick-check agents for content, and the CRM vendor's engagement agents, run inside the platform contract (UC-H1) as first-pass triage. Build the enterprise's own claim-substantiation rules, the scientific-exchange boundary, MSL pre-call intelligence, insight theming, and congress extraction — the capabilities that carry the enterprise's medical judgment. Human MLR and medical review sign-off stays.
  • Make the firewall, the off-label boundary, and AE routing structural. Separate identity principal, separate registry scope, separate KG partition for medical; unsolicited-request and off-label rules per market (FDA vs. EFPIA) as versioned rules, not prompt instructions; adverse-event detection in inquiries and interactions as a deterministic, non-bypassable hook with human confirmation. A policy that lives in a prompt is not a control.
  • Measure the portfolio on decision velocity for the seven Medical Affairs decisions in §8.7 — inquiry response time, MSL prep time per interaction, abstract-to-alert latency, insight-to-strategy time, gap-to-funding-decision time, medical review cycle time, PV clock compliance — plus brief-acceptance rate in coach mode and citation completeness. Not on number of agents. Put AI telemetry (token cost, latency, drift, hallucination rate, prompt-injection attempts per agent, per product, per region) on one plane from day one; it is the TCO dashboard the portfolio will be judged on.
How to Read This Article

Pick your lens — the contents highlight your path

This is a long piece because the problem is a whole-enterprise problem. It is written to be read straight through, but four kinds of readers will want to weight sections differently. Select a role and the table of contents highlights your read first and then sections, estimates the reading time, and shows what your role gains. Your choice is remembered on this device.

for the platform strategy lens — platform economics, sequencing, competitive context, org design, and the measurement framework that tells you whether the bet is paying off.
Before — today's state
Three agent estates, no shared governance; each use-case a $2–5M standalone; no cross-functional decision-velocity metric; competitive landscape assessed vendor-by-vendor.
After — with the platform
One platform contract governs all agents; marginal cost drops 40–60% by use-case 5–6; one scorecard from target ID to batch release; coherence matrix shows reuse at a glance.
for the agentic lab lens — agent anatomy, composition rules, registry-driven reuse, and the kill criteria that keep the Lab honest.
Before — today's state
PoCs built as monoliths with 50–100 tools; no kill criteria; no registry; "done" means "demo works".
After — with the platform
≤5-tool bounded specialists; kill criteria enforced per iteration; every PoC leaves a registry card; "done" means the promotion gate passes.
for the agentic factory lens — promotion gates, support lens, validation posture, observability, and the command-center surfaces production agents feed.
Before — today's state
Validated agents without SLAs; no support model; model deprecation discovered in production; no observability across agents.
After — with the platform
L1/L2/L3 support tiers; quarterly successor-model testing; Enterprise AI Ops Command Center; confidence-gated degraded mode.
for the semantic & knowledge engineering lens — federated ontology governance, KG-as-a-Service, data-lake integration, run-time triplet materialization, process intelligence.
Before — today's state
Per-project ontologies that diverge; entity resolution only in commercial; KG-as-a-query vs. KG-as-a-service; data fabric aspirational.
After — with the platform
Thin core + federated domain extensions under a modeling standard; entity resolution as a first-class cross-domain service; canonical attribute dictionary over regional LIMS; process intelligence fused as confidence-scored edges.

Whatever lens you bring, the argument has one spine: decision velocity and decision quality are not a trade-off in pharma; they are the same engineering problem, and the platform is how you solve it once.

Section 1 · The Stakes

Why decision velocity is the new competitive advantage

Pharma has always known that a wrong decision is expensive. A missed safety signal, a mis-specified endpoint, a batch released against an unnoticed trend — these carry costs measured in patient harm, inspection findings, and write-offs. What the industry has been slower to internalize is that a slow decision carries the same order of cost, and that most of the industry's slowness is not caused by caution. It is caused by the time it takes to assemble evidence.

Consider where the calendar actually goes:

DecisionTypical latency todayWhat the latency costsWhy it is slowDemand type
Amend a clinical protocolWeeks to draft, months to implement~$500K per amendment; roughly half are avoidable with better design intelligence up frontPrecedent lives in a hundred prior protocols nobody has time to readType 1 waste
Confirm a process is in a state of control (APQR)Findings surface 6–15 months after the batches were madeA trend that should have triggered an investigation in March appears in the following February's reportHistorian export, Excel, manual batch reconciliation, 4–8 weeks per productType 1 waste
Detect process drift between batches (CPV)4–12 weeks from batch completion to chartThe first signal that something changed is a failed batchFive to ten of 30–50 CPPs actually trended; interactions invisibleType 2 / failure demand
Stop a blend when it is uniformFixed-time step: blend for N minutes, sample, wait for QC15–45 minutes per batch at $5K–$15K per batch-hour of capacityNo measurement-based endpoint; the clock decidesType 1 waste
Select trial sitesHistorically weeksSlow enrollment, under-enrolling sites discovered late, diversity commitments missedSite identity scattered across CTMS, EDC, CRO feeds with no golden recordType 1 waste
Close a deviation investigationDays to weeksBatch on hold; CAPA effectiveness unknownSimilar historical deviations live in PDFs and memoryFailure demand
Process an expedited adverse-event caseHard 15-day clockRegulatory non-complianceIntake across five channels, manual coding, narrative draftingType 1 waste
Answer a medical-information inquiry with approved contentHours to daysHCP trust; off-label riskMLR-approved content not retrievable with provenanceType 1 waste
Prepare an MSL for a KOL interaction1–3 hours per interactionFewer, shallower scientific exchanges; missed follow-ups; insights never capturedPrior interactions in the CRM, inquiries in another system, the KOL's abstracts in a third, approved SRLs in Vault — assembled by handType 1 waste
Turn a congress into medical intelligencePost-congress report in 4–6 weeksCompetitor data and KOL shifts reach TA leads after the strategy cycle has moved onThousands of abstracts read by a handful of people; no structured extraction, no KGType 1 waste
Value work — the decision itself (minutes)Type 1 waste — necessary only because of how the system is designedType 2 waste / failure demand — exists only because something upstream failed

Every row has the same shape: the evidence exists, it is scattered across systems and formats, and a skilled person spends most of their time finding and reconciling it rather than deciding. The decision itself — the part that requires a toxicologist, a medical writer, a QA lead, a process scientist — is often minutes. The assembly is weeks.

There is a sharper way to read that table, and anyone who has run a Lean or Six Sigma program will recognize it. Systems thinking — the Vanguard Method developed by John Seddon — asks first whether the work an organization does is value demand (what the customer actually needs) or failure demand (demand caused by the system's own prior failures), and classifies each activity as value work, Type 1 waste (non-value, but currently necessary given how the system is designed), or Type 2 waste (exists only because something upstream failed). Read that way, the table is almost entirely amber and red. Reading a hundred prior protocols is Type 1 waste — necessary only because precedent was never structured. Hand-reconciling batches for an APQR is Type 1. A failed batch that is the first signal of drift is Type 2, and the investigation it triggers is failure demand. The decision itself is the thin green thread of value work. Pharma's calendar is not consumed by deciding; it is consumed by waste the system generates for itself.

This is why agentic AI matters in pharma in a way that is different from how it matters in a consumer technology company. The value is not automation. Medicine will always be a human endeavor, and in regulated workflows the human must decide. The value is compressing the assembly so that the human decides sooner, with more of the evidence in front of them, and with the rationale captured so the next person can decide faster still. That is intelligent decision support, and it is the only framing under which agentic AI survives contact with GxP.

Three questions define whether an enterprise gets there. They are asked by three different executives, and a platform that cannot answer all three does not ship.

The Chief Data & AI Officer
— The Platform Question.

"We have three agent estates — internal, a CRM vendor's agents, an IT-operations partner's agents — and no shared governance. How do I prevent forty disconnected implementations?"

The Head of R&D AI
— The Validation Question.

"The PoC worked in the lab. Quality says they can't validate an LLM under GAMP 5. How do we get this to production?"

The VP Quality & Manufacturing
— The Trust Question.

"Show me which source supported which sentence in the investigation report. If I can't trace it, I can't sign it."

The rest of this article is an answer to those three questions. It starts by being honest about why the answer has to be different from what a technology company would build.

Section 2 · The Challenge

Why pharma agentic AI demands a different architecture

The cost of a slow decision is the cost of a wrong one

In consumer technology, an agent that is wrong five percent of the time is impressive, and an agent that is slow is merely annoying. In pharma, an agent that is wrong once in a regulatory submission, a safety report, or a batch-release decision creates an inspection finding — and an agent that is slow creates the same finding by a different route: the trend nobody saw, the amendment nobody avoided, the clock nobody met. Regulators have moved from asking "did you do the review" to "did the review find anything, and what did you do about it, and when." Latency is now a compliance property.

That is the reframing this section exists to make. The industry conversation about pharma AI has been dominated by the wrong-decision risk — hallucination, bias, black boxes — and that risk is real. But an architecture built only to prevent wrong decisions produces a system so gated, so manual, so bespoke per use-case, that it never compresses any calendar. The platform has to be designed for both failure modes at once: it has to make decisions faster and more governed simultaneously, or it makes neither. Four constraints make that harder in pharma than anywhere else.

Constraint 1 · Regulatory

GxP governance is non-negotiable, and it is per-decision, not per-system

21 CFR Part 11 audit trails. GAMP 5 validation. EU Annex 11 electronic records. ALCOA+ data integrity. Every agent that touches a regulated process needs immutable execution history, version-controlled rules, and human sign-off — not as a nice-to-have but as a condition of deployment. The subtlety is that the governance requirement attaches to the decision, not to the software. A drafting agent that produces a CSR section is subject to Part 11 at the moment a medical writer signs the section. A CPV agent that raises an SPC signal is subject to GMP at the moment a process scientist acts on it. The platform must therefore know, for every agent invocation, which regulatory frame applies to the decision it is supporting — and adjust the governance envelope accordingly. A single flat "compliance mode" is both too much for discovery science and not enough for batch release.

Regulatory constraint. Governance must be scoped to the decision, and the scoping must itself be auditable.

Constraint 2 · Fragmentation

Three agent estates, zero shared governance

Internal R&D agents. Commercial agents on a CRM vendor's agent platform. IT-operations agents built and operated by a systems integrator under a multi-year contract. Each estate arrives with its own identity model, tool access, observability, and evaluation practice. Without a shared platform contract, the enterprise ends up with three incompatible agent ecosystems that cannot be governed as one — and when an inspector asks "which agents can touch validated systems, and how do you know," the answer is three different answers from three different vendors.

Fragmentation risk. The platform's first product is a contract that partner-built agents deploy into, not around.

Constraint 3 · Scale

Dozens of use-cases × dozens of implementations is unsustainable

Target identification, generative chemistry, preclinical safety, trial design, site selection, operational precision, regulatory documentation, tech transfer, APQR, CPV, PAT, blend endpoint, deviation investigation, supplier risk, cold chain, HCP engagement, MLR review, medical inquiry, pharmacovigilance, semantic search. That is twenty before anyone in finance or HR has raised a hand. If each one builds its own ontology, its own agent framework, its own evaluation harness, and its own validation story, you have built a consulting program, not a platform — and at $2–5M per standalone implementation, the arithmetic ends the program by year three.

Scale blocker. The unit of reuse has to be a platform service, not a project template.

Constraint 4 · Velocity

Latency is itself a risk, and it compounds

This is the constraint the industry under-weights. Every month an adverse trend goes undetected is a month more batches ship against it. Every week a protocol design sits without precedent analysis is a week closer to an amendment. Every day a deviation investigation waits for someone to find the similar case from 2021 is a day of batch hold. These latencies compound across the pipeline: a slow site-selection decision delays enrollment, which delays the CSR, which delays the submission. A platform that governs beautifully but leaves the assembly time untouched has not reduced risk; it has formalized it.

Velocity risk. The platform must be measured on decision latency as a first-class metric, alongside accuracy and compliance.

Two constraints are properties of the domain; two are conditions the enterprise created for itself

A systems-thinking reader will notice that the four constraints are not of one kind. Constraints 1 and 4 are properties of the domain: GxP is the patient's purpose expressed as regulation, and latency compounds because the pipeline is sequential. No architecture removes them; the platform must serve them. Constraints 2 and 3 are system conditions — conditions the enterprise created for itself through procurement decisions, org design, and per-project funding, and which now generate most of the failure demand described in §1. Three agent estates exist because three contracts were signed separately. Dozens of implementations exist because use-cases are funded as projects. The distinction governs what the platform may honestly claim: against Constraints 1 and 4 it can only compress; against Constraints 2 and 3 it can remove the cause. An architecture that governs beautifully but leaves those two system conditions in place has automated the waste rather than removed it — and automated waste is still waste; it just runs cheaper.

These four constraints are why the architecture that follows looks the way it does. Governance is a cross-cutting plane, not a gate at the end. Agents are composed from bounded specialists rather than inflated into monoliths, because monoliths are neither fast nor validatable. The semantic layer is shared and federated, because that is the only way use-case N+1 is cheaper than use-case N. And every use-case is described in terms of the decision it accelerates, because that is the only value that matters.

Before the architecture, three principles. They are the acceptance criteria every layer has to meet.

Section 3 · Guiding Principles

Three acceptance criteria every component must meet

These are not aspirational. They are the acceptance criteria for every component in the platform. A use-case that violates any of them does not graduate from the Lab. They deliberately mirror the way leading R&D organizations have articulated their AI posture — because that articulation is right, and because a platform that speaks the enterprise's own language gets adopted.

Purpose-Driven

Address specific scientific and operational challenges — not AI for AI's sake. Every agent has a named business problem, a named user, and a named metric. If you cannot articulate the value of a single inference — what decision it accelerates, by how much, for whom — do not ship it.

This principle is why every use-case in §8 opens with the decision, not the model, and why the Lab's kill criteria in §7 are tied to a metric moving, not to a demo working.

Smart & Scalable

Reusable across therapeutic areas, sites, and geographies. One ontology, one agent framework, one validation strategy — configured per domain, not rebuilt. A new use-case is configuration, not a project.

This principle is why §5.3 insists on a thin core ontology with federated domain extensions, why §6 introduces a registry so agents can be discovered and composed rather than re-implemented, and why §9 measures the platform on the reuse ratio rather than on the number of agents deployed.

Human-Centered & Responsible

"Artificial intelligence needs to augment natural intelligence." Medicine will always be a human endeavor. Every agent provides intelligent decision support — rationale, provenance, an owner — and no agent exercises black-box autonomy in a regulated workflow. The goal is to raise the average to the level of the best: to shift the performance of many individuals closer to the performance of the best individual, by putting the best evidence and the best precedent in front of every decision-maker.

This principle is why §10's dual-path reasoning ends in a merged assessment a human signs, why §5.5's governance plane captures decision provenance and not just evidence provenance, and why §19 puts the promotion gate in the hands of domain experts rather than engineers.

Three principles, one consequence: the platform is a decision-support substrate, not an automation engine. That consequence is the thesis.

Section 4 · The Thesis

One substrate, every decision

What if every decision in your enterprise — from target selection to batch release to the next field call — ran on one semantic substrate, one agent framework, one governance plane, so that the fastest decision and the best-governed decision were the same decision?

That is the whole proposition, and it is worth being precise about why the two halves belong in one sentence. Speed without governance is what most enterprises have today in their ungoverned pilot estates: fast, impressive, unshippable. Governance without speed is what most enterprises get when Quality retrofits controls onto those pilots: validated, auditable, and so slow that nobody uses it. The industry has been treating these as a dial to be tuned — more speed here, more control there — when they are actually the same engineering problem viewed from two sides.

The reason they are the same problem is evidence assembly. The thing that makes a decision slow (finding and reconciling scattered evidence) is the same thing that makes it hard to govern (no provenance for what was found). Solve the assembly once — a semantic layer that makes enterprise evidence queryable with provenance, agents that reason over it with citations, a governance plane that records what was retrieved and what the human did with it — and speed and governance both fall out of the same mechanism. A recommendation that arrives with its evidence trail attached is both faster to act on and faster to audit.

That is what "raising the average to the level of the best" means in practice. The best toxicologist, the best medical writer, the best process scientist are the best partly because they remember the precedent. The platform makes the precedent available to everyone, with the source attached, at the moment of decision. The human still decides. The decision just arrives sooner and leaves a trail.

The four-layer architecture is how that substrate is built.

Section 5 · Platform Architecture

Four layers, one governance plane — each layer serves every domain

The platform is a four-layer stack — Experience, Agentic, Semantic & Knowledge, Data — with a cross-cutting governance plane that enforces GxP compliance, identity, audit trails, observability, and evaluation across all four. The layers are described top-down here because that is how a user experiences them; they are built bottom-up, which is how §13's roadmap sequences them. The single most important property of the stack is that each layer serves every functional domain. There is not a research stack and a manufacturing stack and a commercial stack. There is one stack, and the domains differ in which ontologies, which agents, which data products, and which governance envelope are configured for them.

Reading the diagram. Each column is a functional domain; each band is a layer. A cell answers "what does this layer give this domain?" — hover a column header to isolate a domain, hover a band to isolate a layer, click a cell to jump to the tower in §8. Requests and queries flow down the left edge; findings and provenance flow up the right; every layer reports audit, trace and policy into the Governance Plane on the right, whose eight categories are detailed in §5.5. Between Semantic and Data, the two-headed arrow is run-time triplet materialization (§5.4).

5.1 Experience Layer

What it is. Role-specific workspaces that consume from the agentic layer. Each persona sees findings, recommendations, and drafts through purpose-built interfaces — not a generic chatbot. The experience layer is where intelligent decision support becomes a decision: it is where the human accepts, modifies, or rejects, and where that choice is captured.

Design rule. No workspace talks to a model directly. Every surface consumes from registered agents (§6) through the platform contract, which is what makes a partner-built surface (a CRM vendor's field-force interface, an SI's IT-ops console) governable in the same way as an internal one.

Surfaces, by domain:

Surfaces, by domain — what the user sees and what the platform captures (15 surfaces)
SurfaceDomainWhat the user seesWhat the platform captures
Target evidence workspaceResearchTarget dossier with KG neighborhood, evidence tiers, cited liability briefGo/no-go rationale, accepted/rejected evidence
Med-chem review queueResearchRanked candidates, rule pass/fail lineage, property predictionsChemist selections as preference signal
Protocol design workspaceDevelopmentFindings, precedents, simulation side by sideEvery accepted/rejected recommendation
Study team inboxDevelopmentNext-best-actions with rationale, impact, ownerAccept/modify/reject, time-to-action
Regulatory writer toolDevelopment / MedicalDraft with per-sentence source, checker findingsEdit distance, sign-off under Part 11
Tech-transfer workbenchProduct DevelopmentProcess knowledge package, gap analysis vs. receiving siteTransfer decisions, open items
Quality investigation hubManufacturingSimilar deviations with provenance, root-cause hypotheses, CAPA historyInvestigator disposition, QA signature
CPV / APQR review consoleManufacturingSPC signals, capability, CPP-to-CQA relationships, cross-site panelsSignal disposition, limit approvals
PAT model consoleManufacturingModel version, prediction confidence band, endpoint advisoryOperator action, model-registry approvals
Supply risk boardSupply ChainMaterial, CMO, cold-chain, and geopolitical signalsMitigation decisions
Field-force assistantCommercialPre-call plan, MLR-only next-best-actionCall outcomes, AE sidecar routing
Medical inquiry deskMedicalFirewalled, approved-source responses in the inquirer's language, inquiry-cluster context, AE screen result, MSL triageResponse provenance, disposition, cluster assignment
MSL workbenchMedicalPre-call brief with every line cited (prior interactions, open inquiries, the HCP's recent abstracts, approved SRLs, what cannot be shared), territory engagement plan vs. the medical plan; works offlineBrief accept/edit, post-interaction structured insight capture — never auto-logged to CRM
Congress intelligence surfaceMedicalAbstract-level alerts by product, competitor, endpoint, and KOL within hours; KG neighborhood per abstract; proposed actionTA-lead disposition, actions taken, abstract-to-action trace
Medical strategy dashboardMedicalCorroborated insight themes with confidence and n_sources, inquiry-cluster trends, evidence gaps and candidate studies, congress signals — one surface per TAAdopt/monitor/dismiss on themes; fund/decline on study concepts; strategy changes cited to insights
Medical content review deskMedicalClaim-to-source findings, fair-balance and scientific-exchange rule results, cross-asset inconsistencies, vendor triage resultsMedical/legal/regulatory disposition, approval record (Part 11)
PV case workbenchMedical & SafetyStructured case, coding suggestions, cited narrative, clock statusReviewer sign-off, clock compliance
Command centersCross-domainMulti-agent feeds composed into decision hubs — §12Escalations, decision log
Agent marketplaceEnterpriseRegistry-driven discovery of agents and kits — §6Usage, approvals, chargeback

The transition from experience to agentic is the platform's most important boundary. Everything the user sees was produced by an agent; everything the user does is fed back to one.

Shared spaces — three levels of context

The experience layer is not a single surface. It is a hierarchy of shared spaces that establish what context is visible, to whom, and at what scope:

Domain space — the broadest shared context
What is sharedKG subgraph, registered agents, data products, command-center feeds, eval benchmarks, kit templates
What is privateNothing at this level — the domain is the broadest shared context
ExampleThe Manufacturing domain space: all sites see the same CPV signals, the same APQR roll-ups, the same deviation history, the same KG, the same registered agents
Sub-group / Division space
What is sharedTeam-specific filters, site-specific thresholds, working hypotheses, in-progress investigation state, draft analyses under review
What is privateRaw data from other sub-groups before aggregation; draft dispositions before sign-off
ExampleA site quality team's working space: they see their site's CPV charts, their open investigations, their draft APQR — not another site's drafts. But the domain space above them sees cross-site trends
Individual space
What is sharedPersonal queue, coaching feedback history, acceptance-rate curve, preferred formats, recent session context
What is privateFeedback corrections the individual has made (tone preferences, edit patterns), draft-in-progress before submission
ExampleA process scientist's workspace: their personal CPV review queue, their correction history that tunes future agent drafts, their individual acceptance rate that drives the coach-mode graduation curve
This three-level model mirrors how knowledge actually flows in a pharmaceutical organization: company-level observations are shared, team-level working state is collaborative within the team, and individual learning is private until it graduates. The graduation mechanism is the same trust-tier model the KG uses: an individual correction becomes a team pattern when corroborated by multiple users, and a team pattern becomes a domain standard when validated by the domain owner. The proven production pattern — where co-worker memory at the individual level (conversation history, edit corrections, preferred formats) promotes to company-level knowledge when 2+ users corroborate — is the same mechanism, applied at enterprise scale.

5.2 Agentic Layer — Lab & Factory

What it is. Two operating modes under one framework. The Lab develops PoCs with rapid iteration. The Factory industrializes validated agents into governed, monitored, production-grade deployments. A promotion gate enforces accuracy, safety, and validation thresholds between them. §7 describes the operating model in full; this section describes the anatomy of the agents themselves, because the anatomy is what makes Lab-to-Factory promotion possible at all.

Agent types. Five archetypes cover every use-case in §8:

ArchetypeJobTypical tools (≤5)Typical model tier
ExtractionTurn unstructured source into structured, provenance-tagged findingsDocument parser, KG writer, schema validatorScoped / small
ReasoningTraverse the KG, compare, score, synthesize with citationsKG query, rule engine, retrievalFrontier for synthesis; scoped for scoring
DraftingProduce sections, narratives, hypotheses from approved sources onlySource registry, template, citation checkerFrontier, tightly grounded
MonitoringWatch event streams, detect signals, suppress noiseStream reader, SPC/KRI calculators, KG lookupScoped / deterministic
OrchestratorCompose bounded specialists into a workflow; route by confidence; assemble the human-facing packageRegistry lookup, specialist invocation, decision logFrontier for planning, minimal tools

The six-layer anatomy. Every agent, regardless of archetype or domain, is built the same way: Model → System Prompt → Context Window → Tools → Hooks → Agent Rules. The model is pinned by version. The system prompt is versioned and owned. The context window is budgeted. Tools are the only way the agent touches the world, and there are no more than five of them visible to the model. Hooks are the deterministic pre/post steps (input validation, output checking, audit logging) that run whether the model cooperates or not. Agent rules are the declarative policy — what confidence threshold triggers escalation, which GxP class applies, who owns it. This anatomy is what gets registered in §6 and what the Factory validates in §11.

Modelpinned by version
System Promptversioned, owned
Context Windowbudgeted
Tools≤5 visible to the model
Hooksdeterministic pre/post
Agent Rulesthresholds · GxP class · owner

Composition over inheritance — the economic and architectural imperative

Here is what happens when you do not compose.

A team builds a "clinical operations agent." It starts well: three tools, one job. Then someone asks it to also check drug supply. Then to draft the monitoring visit letter. Then to look up the site's contract status. Then to pull the enrollment forecast. Eighteen months later the agent has a hundred tools, and four things have gone wrong at once.

The context window is spent before the user speaks. A hundred tool schemas consume tens of thousands of tokens of every single call — before a word of the actual task arrives. Prompt caching helps, until someone edits one tool description and invalidates the cache for the whole estate. The agent is paying full frontier-model price to re-read its own tool manual on every invocation.

Tool selection degrades. With five tools, a frontier model picks the right one almost every time. With a hundred, it picks a plausible one — and "plausible" is how an agent ends up querying the IRT system when it meant the CTMS. The failure is silent, and it is not the model's fault. It is the architect's.

Security boundaries blur. An agent that can read pharmacovigilance cases and write to the CRM and modify a monitoring plan is not one agent with a broad job; it is one identity principal with a broad blast radius. The least-privilege question — "what is the minimum this agent needs to touch?" — has no answer, because the answer is "everything."

Validation becomes impossible. GAMP 5 asks what the system does and how you know it does it. A hundred-tool agent's answer is "it depends on what the model chose." Category 5 validation of a system whose control flow is decided at run-time by a stochastic planner is not a validation strategy. It is a hope.

Now the number. In our measured deployments, the cost difference between a frontier model and a scoped model for the same bounded task is roughly 137× (Claude 3.5 Sonnet vs. Claude 3 Haiku on bounded extraction and scoring tasks, Q3 2025, prompt caching enabled). A monolith routes everything to the frontier tier, because a monolith cannot tell which sub-task is easy. A composed system routes extraction, normalization, scoring, and monitoring to scoped models — and reserves the frontier tier for synthesis, planning, and the handful of decisions that need it. On a portfolio of a dozen production agents handling tens of thousands of invocations a day, the difference between those two routing policies is not a line item. It is whether the program has a budget next year.

Monolith — "one agent, a hundred tools"

The context window before the user's first token

100 tool schemas · ~62%
system prompt
task
room
Tens of thousands of tokens re-read on every call, at frontier price — the agent pays to read its own tool manual before the task arrives.

Routes everything to the frontier tier, because a monolith cannot tell which sub-task is easy. One principal, every permission. Cat 5 scope is "everything".

Composed — bounded specialists, ≤5 tools each
137×measured cost difference, frontier vs scoped model, on the tokens that matter
prompt
task + evidence
room for reasoning
Hundreds of tokens of schema per specialist. Extraction, normalization, scoring and monitoring go to scoped models; the frontier tier is reserved for synthesis, planning, and the handful of decisions that need it.

On a dozen production agents handling tens of thousands of invocations a day, the difference between the two routing policies is not a line item. It is whether the program has a budget next year.

Monolith ("one agent, a hundred tools")Composed ("bounded specialists, ≤5 tools each")
Tool schema overhead per callTens of thousands of tokens, every callHundreds of tokens per specialist
Model tierFrontier for everything (cannot discriminate)Frontier for ~20% of calls; scoped for the rest
Relative cost per decision (normalized)~100~3–8
Tool-selection accuracyDegrades with tool countStable
Least-privilege storyNone — one principal, every permissionEach specialist has exactly its allowlist
GAMP 5 postureControl flow is stochastic; Cat 5 scope is "everything"Deterministic control flow at the orchestrator; specialists validated as bounded components
Failure isolationOne bad tool call can corrupt the whole runFailure is local to one specialist, and the orchestrator's hooks catch it
Promotion to FactoryWhole monolith must re-validate on every changeSpecialists promote independently; orchestrator re-validates only on topology change

The design rule is therefore the default: the orchestrator composes bounded specialists via natural language; it never inflates one agent with a hundred tools. Every specialist has ≤5 model-visible tools, its own eval record, its own identity, and its own place in the registry. This is not an architectural preference. It is the only configuration that is simultaneously affordable, secure, and validatable — and it is what the 137× figure is really measuring.

Proven pattern

ProtoCheck's four specialist reviewers (safety, efficacy, regulatory, operational) each hold ≤5 tools and apply expert-calibrated rules; a coordinator merges findings with provenance. Doscierge's checker is a deterministic MCP tool that an orchestrator calls, not an agent that decides — the deliberate GAMP 5 Category 5 ceiling.

Components in the layer: Extraction agents · Reasoning agents · Drafting agents · Monitoring agents · Orchestrator agents · Agent Registry & Service Catalog (§6) · MCP tool servers · Evaluation harness (P/R/F1 per use-case, per domain, per class) · Closed-loop evaluation (Model-loop, Dogfooding-loop, Production-loop — §7).

The agents reason over something. That something is the semantic layer.

5.3 Semantic & Knowledge Layer

What it is. The shared substrate every agent reasons over. A thin core ontology, centrally owned; domain ontologies federated to domain teams under a published modeling standard; the whole exposed as Knowledge-Graph-as-a-Service with query, provenance, and versioning APIs that agents call through MCP. No agent builds its own retrieval. None owns a private copy of the graph.

Why federated. A single central ontology becomes a bottleneck by month six — every domain is waiting on the ontology team to model its entities. A set of team-built graphs becomes a swamp by month twelve — nothing joins to anything. The federated model threads the needle: the core owns the entities everyone shares (product, compound, study, site, HCP/investigator, document, batch, material, specification), and each domain owns its extensions under a modeling standard that guarantees they join back to the core.

Core ontology — centrally owned, thin, versioned
Product · Compound · Study · Site · Investigator/HCP · Document · Batch · Material · Specification · Organization · Regulatory Requirement
Aligned to: IDMP, CDISC (SDTM/ADaM/PRM), MedDRA, ISA-88/95, FHIR
Research KG (domain-owned)
gene · protein · pathway · disease · assay · target-evidence tier
Aligned to: MONDO, HPO, GO, ChEBI, UniProt
Clinical KG
protocol · endpoint · population · design element · enrollment curve · monitoring finding · supply state · regulatory precedent
Aligned to: CDISC PRM/SDTM, ICH E6/E8/E9, CT.gov
CMC / Product Development KG
formulation · unit operation · CPP · CQA · design space · control strategy · analytical method · stability study · established condition (ICH Q12)
Aligned to: ICH Q8–Q14, USP, CTD Module 3
Manufacturing & Quality KG
batch · phase · equipment · historian tag · parameter · sample · test · result · specification · genealogy · deviation · CAPA · change · PAT model
Aligned to: ISA-88/95, ICH Q7/Q9/Q10, 21 CFR 211
Supply Chain KG
material lot · supplier · CMO · CoA/CoC · shipment · lane · temperature excursion · serialization event · demand signal
Aligned to: GS1, DSCSA, GDP
Commercial KG
HCP/HCO golden record · territory · product · indication · MLR-approved asset · interaction · consent
Aligned to: NPI, MDM golden record, MLR taxonomy
Medical & Safety KG
case · adverse event · seriousness · expectedness · causality · MedDRA PT · signal · medical inquiry · inquiry cluster · MSL interaction (type, topics, content shared) · field insight · insight theme · KOL tier · congress · abstract · publication · evidence gap · study concept · medical content · claim · transfer of value
Aligned to: MedDRA, E2B(R3), CIOMS, GVP, Sunshine Act / Open Payments, EFPIA Code

Explore every entity and predicate of the federated model interactively in the KG ontology explorer (§8.8).

Entity resolution is a first-class service

Compound ↔ study ↔ site ↔ HCP ↔ batch ↔ material — the same real-world thing appears under different identifiers in CTMS, EDC, LIMS, ERP, CRM, and the CRO's feed. The semantic layer owns the golden record and the crosswalk, using the master-data discipline that commercial organizations have applied to HCPs for years and R&D and manufacturing rarely have. Site selection (§8.2) is a master-data problem before it is a prediction problem; batch genealogy (§8.4) is a master-data problem before it is a trending problem.

Process intelligence is fused into the graph

Event-log mining over MES, QMS, CTMS, and ticketing systems — Celonis-style process discovery — surfaces what actually happens: where batches stall, which deviations recur, how long a protocol amendment really takes end to end. Those temporal and statistical findings are written into the KG as confidence-scored edges alongside the semantic and normative knowledge of what should happen. An agent asking "why is this batch late" can traverse from the batch to the phase where the delay concentrates to the deviation type that usually causes it to the CAPA that was supposed to fix it. A sample of the event log is enough for discovery; production monitoring needs the full stream.

Trust tiers and provenance on every edge

Every triplet carries a source, a trust tier, a confidence, a timestamp, and a version. Normative sources (regulations, approved specifications, signed rules) sit at the top tier; mined and inferred edges sit lower with a hard confidence cap. An agent's citation therefore carries not just where the fact came from but how much to trust it — which is what a reviewer needs to decide quickly.

NormativeStructuralOperationalCarry-fwdEnrichmentMinedcap 0.75

Regional KG partitioning — the graph has a geography

The federated model above is federated by domain. A global pharma graph must also be federated by jurisdiction, and the two partitionings are orthogonal: a Medical & Safety triple about a German patient and one about a Texan patient belong to the same domain and different residency zones. The design rule is that residency is a property of the triple, not of the graph, and the KG-as-a-Service endpoint is a router before it is a store.

GLOBAL CORE (replicated to every region; contains no personal data)
  Product · Compound · Study · Site · Material · Specification · Regulatory Requirement ·
  ontologies · signed rules · label versions by market · MedDRA/IDMP/CDISC alignment ·
  HCP professional attributes (specialty, affiliation, publications — public-record)
        │
        ├── REGIONAL PARTITION — EU        Azure West Europe (NL) [+ North Europe (IE) DR]
        │     patient-linked triples (case, subject, consent) · HCP interaction triples ·
        │     prescribing-linked triples · MSL note content · inquiry text
        │     Legal frame: GDPR Art. 9 (special categories), Art. 44–49 (transfers), Art. 35 (DPIA)
        │
        ├── REGIONAL PARTITION — US        Azure East US 2 (VA) [+ Central US DR]
        │     PHI-bearing triples (HIPAA 45 CFR 164), US HCP interactions, US prescribing
        │     Legal frame: HIPAA Privacy/Security Rules; state laws (CCPA/CPRA, WA MHMDA)
        │
        ├── REGIONAL PARTITION — CN        Azure China North (Beijing, operated by 21Vianet)
        │     all personal information and "important data" of Chinese origin
        │     Legal frame: PIPL Art. 38–40 (CAC security assessment before any export);
        │     HGRAC for human genetic resources — no export of genomic triples, period
        │
        ├── REGIONAL PARTITION — JP        Azure Japan East [+ Japan West DR]
        │     Legal frame: APPI (anonymized / pseudonymized info regime) · PMDA · J-GCP
        │     Exit: EU–Japan mutual adequacy decision; APEC CBPR
        │
        ├── REGIONAL PARTITION — BR        Azure Brazil South [+ Brazil Southeast DR]
        │     Legal frame: LGPD Art. 11 (sensitive health data) · Art. 33 (international transfers) · ANVISA · CONEP ethics
        │     Exit: ANPD standard contractual clauses (Resolution CD/ANPD 19/2024, finalized August 2024)
        │
        └── REGIONAL PARTITION — CH · ...  Azure Switzerland North · additional as operations require
              Legal frame: FADP (adequate under GDPR) — same mechanism, different legal-basis edges

Cross-border query governance

An agent never addresses a regional partition directly. It issues one federated query (SPARQL SERVICE blocks, or the equivalent property-graph federation) to the global KG-as-a-Service endpoint, which dispatches sub-queries to the regional endpoints the agent's card (§6) permits. Each regional endpoint evaluates the classification of the triples the sub-query would touch — not the identity of the caller alone — and applies one of three outcomes: return full results (non-PII triples), return aggregates only with a k-anonymity floor (PII triples, caller outside the region), or refuse and log (PII triples, caller not authorized for aggregates either). The federation layer joins what comes back.

The consequence that matters for an inspector: a signal-detection agent running in Basel can ask "how many serious hepatic events across all regions for product X in Q2" and receive five numbers, but it cannot receive five case narratives — and the audit trail (G3) shows both the question and the classification decision each region made.

The global patient node is pseudonymized, not absent

Schema-level absence of a patient entity may seem like the strongest privacy control — you cannot leak what the ontology cannot represent — but it is also the most analytically destructive. Cross-regional duplicate detection in pharmacovigilance (a legal obligation under GVP Module VI), global enrollment monitoring, cross-study exposure analysis, and any agentic workflow that reasons about patient journeys all require a global patient-level join key. CDISC SDTM already mandates USUBJID — a cross-study, pseudonymized subject identifier that FDA requires in every submission — so the sponsor already maintains a global pseudonymized patient identifier by regulatory design.

The architecture follows the pseudonymisation-domain model defined in EDPB Guidelines 01/2025 (paras 35–41): the control is not the absence of the entity but the boundary of the pseudonymisation domain — the guarantee that regional keys never enter the global tier.

Purpose-scoped tokens and two-layer keys

EDPB para 116 warns that a single consistent "person pseudonym" across all purposes carries a "comparatively high" risk of unauthorized attribution. The global KG holds two token families: T_pv = HMAC-SHA256(K_pv, canon(identifiers)) for pharmacovigilance duplicate detection (legal basis: GDPR Art. 6(1)(c) / Art. 9(2)(i); GVP Module VI), and T_res = HMAC-SHA256(K_res_study, canon(identifiers)) for Art. 89 research (enrollment monitoring, cross-study exposure). Cross-purpose linkage (T_pv ↔ T_res) is mediated only by the honest broker, logged, and requires DPO approval.

Two-layer key architecture. The regional partition computes a regional token using a regional key. An honest-broker service — operated in an attested trusted execution environment (TEE), organizationally separate from the analytics platform, reporting to Global Privacy / DPO — re-keys the regional token to the global purpose token. The global KG tier never holds any HMAC secret and therefore cannot compute a token from an identifier (the EDPB para 88 test for a valid pseudonymisation domain).

Re-identification requires threshold key recovery. Regional keys are split via 2-of-3 secret sharing (EDPB para 108): regional DPO, regional QA/PV head, and global privacy office. Re-identification is a human step-up workflow — never an agent action — requiring a documented legal basis and logged in the RoPA/DPIA.

Jurisdictional partition rules

EU/EEA/UK: tokens flow to the global tier; pseudonymisation serves as an Art. 46 supplementary measure per EDPB para 63. US: obtain an Expert Determination (45 CFR 164.514(b)(1)) covering the tokenized global dataset, because HMAC tokens are derived from PHI and fail Safe Harbor's 164.514(c) "not derived from" test — every major US tokenization vendor operates under Expert Determination. China: aggregate-only partition by default; a token export constitutes sensitive-PI transfer under PIPL Art. 38 and triggers a CAC security assessment (≥10,000 sensitive records/year under the March 2024 provisions) plus HGR clearance where genomic-related data is involved. Japan: trial data via APPI consent; NGMIL pseudonymized medical information cannot be provided to third parties — keep Japan RWD in-region. Brazil: LGPD Art. 13 §2 pseudonymisation; onward transfer restricted.

Output controls. Any global-tier agent output undergoes small-cell suppression: k≥11 for externally reportable results (aligned with CMS Cell Size Suppression Policy), k≥5 with complementary suppression as the internal floor. Automated agent-generated aggregates carry differential-privacy noise calibrated per NIST SP 800-226.

Reference architectures: N3C Linkage Honest Broker, PCORnet PPRL framework, German Medical Informatics Initiative federated Trusted Third Party (two-stage pseudonyms across 34 university hospitals), and EHDS Regulation (EU) 2025/327 (Health Data Access Body + Secure Processing Environment + statutory re-identification ban). All converge on the same shape: tokens at the shared tier, identity at the source, an honest broker in between, and a prohibition on user-initiated re-identification.

Agent scope = pseudonymisation-domain admission. Each agent's workload identity carries an explicit list of pseudonymisation domains it may enter (global:T_pv, global:T_res, regional:EU, etc.). Agents admitted to the global domain can read tokens but cannot invoke the honest broker. This aligns with the ICO's January 2026 agentic AI report (controllers remain responsible; restrict which databases an agent can reach) and the Entra Agent ID least-privilege pattern (GA May 2026: task-scoped RBAC, tool allowlists, JIT elevation, explicit provisions for agents accessing regulated data).

Entity resolution across regions

The golden record described above is itself partitioned. Cross-region matching runs on non-PII attributes only — NPI, ORCID, institutional affiliation, product, study identifier, site identifier — so that "Dr. Müller at Charité" resolves to the same global HCP node whether the interaction is logged in Berlin or Boston. PII-bearing attributes (date of birth, home address, prescribing, interaction content) are resolved locally inside the regional partition and never transit.

The merged golden record that the global core holds contains only the shareable fields plus regional pointers — opaque keys that resolve only when dereferenced by an agent running inside that region. This is the same crosswalk discipline the commercial MDM teams already practice, with the residency boundary made explicit in the schema rather than assumed in the deployment.

The golden record splits at the attribute, not the entity

In practice, a single HCP golden record spans at least five classification tiers. Professional attributes — NPI, institutional affiliation, publications, congress attendance, therapeutic focus — are shareable globally under legitimate-interest processing (GDPR Art. 6(1)(f)). Prescribing data is regional and often contractually restricted (the data-provider's license terms may prohibit export and secondary use outside the licensed geography; in some EU member states, prescriber-level data is prohibited outright). Interaction history — field calls, inquiry contacts, speaker engagements — is regional under purpose limitation (Art. 5(1)(b)). Transfer-of-value records range from fully public (US Open Payments / Sunshine Act) to country-by-country regional disclosure (EFPIA Code). KOL tiering scores are shareable as professional assessments, but the interaction data that informed the score stays regional.

The semantic layer must model these as separate attribute groups on one entity, each carrying its own residency tag — not one record with one classification. An agent that reads the tier can run globally; an agent that needs to explain the tier must run regionally.

Proven pattern

Doscierge: 10,292 triplets across 29 domains, 13-field triplet schema, KG-first pipeline. Enterprise Knowledge Vault reference architecture powering 8 production systems; the semantic layer that lifted retrieval accuracy from 75% to 99%+ (on a 212-question evaluation set against expert-labeled ground truth) at a Fortune 50 pharma. Mid-market manufacturing EKG: 1,786 triplets, 47 entity types, 50 predicates, 6 trust tiers with a T4 hard cap at 0.75.

Components in the layer: Core ontology · Domain KGs (seven, above) · Entity resolution & golden record · Process intelligence layer · Trust-tier & provenance model · KG-as-a-Service API (MCP) · Ontology governance agents (drift detection, orphan entities, cross-domain consistency).

The graph does not float. It is anchored to data, and how it is anchored determines whether it is fast, fresh, and inspectable.

5.4 Data Layer

What it is. Controlled, versioned data sources — the enterprise data lake and its data products, the transactional systems of record, external regulatory and evidence feeds — plus the approved-source registry that defines what an agent is allowed to cite. The platform connects to what the enterprise has. It does not replace it.

The enterprise has already invested here. Most large pharma organizations now have a curated data platform spanning ten to thirty years of clinical and preclinical data, a manufacturing data lake fed by historians across dozens of sites and many ERP instances, and a growing set of governed data products with named owners. The semantic layer must sit on top of that investment, not beside it. The question is how.

How the KG and the data lake fit together

The mistake is to treat the KG as a second copy of the lake. It is not. The KG is a semantic index over data products — it holds the reference and normative knowledge persistently (ontologies, regulations, signed rules, specifications, golden records) and materializes operational triplets at run-time from the conformed data products for the entity neighborhood an agent is actually reasoning about.

1 · TRANSACTIONAL SYSTEMS OF RECORD authority stays here · agents never write here MES LIMS (EU · US · APAC) ERP (multiple) QMS CTMS EDC IRT CRM PI / OSIsoft per site PAT vendor PCs CMO / supplier files (CoA, CoC · PDF · spreadsheet · EDI) Agents read data products and write decisions to the workflow tool the human uses (QMS, CTMS, CRM action) — never into a validated system of record. That boundary keeps the agent out of the transactional system's validation scope. land as-is 2 · ENTERPRISE DATA LAKE DATA FABRIC / VIRTUALIZATION · canonical attribute dictionary regional LIMS + site historian tags + CMO files → one logical test_results service quality-owned, under change control · joined to specifications by effective date, never by current version RAW ZONE landed as-is immutable CONFORMED ZONE data products · named owners batchesMS&T unit_operation_phasesMS&T parametersMS&T samples · testsQC test_resultsQC specificationsQA genealogySupply raw_material_resultsQC/Supply subjects · sitesClin Ops deviations · capaQA hcp_masterMDM The three products CPV & APQR cannot live without: historian CPPs (batch_phase_summary) · LIMS CQAs (test_results ← specs) · genealogy + CoA Join them → the CPP-to-CQA relationship — the scientific point of APQR and CPV — becomes a query. ANALYTIC-READY batch_phase_summarybatch_quality_profileenrollment_curves result_status (final / retest / superseded) carried through every query; specification joined by effective date on the batch date — the version that applied, not the version that is current. PROCESS INTELLIGENCE · event-log mining (MES · QMS · CTMS · ticketing) where batches stall · which deviations recur · how long an amendment really takes → mined findings a sample of the log suffices for discovery; monitoring needs the full stream · Celonis-type tools are a source, not a rival 3 · SEMANTIC LAYER — KG-as-a-Service query · provenance · versioning APIs via MCP PERSISTENT KG — the meaning • core + seven domain ontologies (§5.3)• regulations · guidance (Regulatory Requirement)• signed, versioned rules with calibration records• specifications with effective dates• golden records (HCP · site · material · batch)• mined process-intelligence edges (confidence-scored)• trust tier · confidence · timestamp · version on every edge which parameter is a CPP · what the validated range is · which specification version applied on the batch date RUN-TIME MATERIALIZED SUBGRAPH — the facts for THIS batch, projected from data-product rows on query batch B-2026-0417→ phase → parameter (CPP) → spc_signal→ sample → test → test_result → specification→ genealogy → material_lot → CoA → supplier • each triplet carries the row's provenance• each triplet carries the data-product version• discarded after the decision is logged — never a second copy freshness without duplication · validation scope stays with the data product materialize on query; never copy the lake confidence- scored edges AGENTIC LAYER — reads via MCP; cites with provenance Why not one big graph? Freshness, duplication, validation scope — the KG is a semantic index over governed data products, not a copy. Where data fabric earns its place: federated LIMS instances, site-specific historian tags, CMO spreadsheet feeds — one logical service, GROUP BY site instead of a heroic project. Transactional systems keep authority: MES for the batch, LIMS for the result, QMS for the deviation. Agents read data products and write decisions to the workflow tool the human uses.
Read it as a Quality head would. The LIMS of record is in zone 1. The APQR chart's numbers come from the conformed test_results and analytic-ready batch_phase_summary products in zone 2, materialized into the run-time subgraph in zone 3 on query. The KG knew which specification version applied because specifications are joined by effective date on the batch date — never by current version.
Diagram key — the four blocks of the KG-and-lake picture, in text
BlockContents
Transactional systems of recordMES · LIMS (regional instances) · ERP · QMS · CTMS · EDC · IRT · CRM · PI/OSIsoft historians · PAT vendor PCs
Enterprise data lakeRaw zone (landed as-is, immutable) · Conformed zone — data products with named owners: batches · unit_operation_phases · parameters · samples · tests · test_results · specifications · genealogy · raw_material_results (CoA/CoC) · subjects · sites · deviations · capa · hcp_master · Analytic-ready zone — batch_phase_summary, batch_quality_profile, enrollment_curves, ...
Data fabric / virtualization(canonical attribute dictionary over regional LIMS, site historians, CMO feeds)
SEMANTIC LAYER — KG-as-a-ServicePersistent: ontologies, rules, regulations, specifications, golden records, mined process-intelligence edges · Run-time materialized: batch → phase → parameter → result → spec → genealogy → material lot → CoA for THIS batch, projected from data-product rows with row-level provenance + data-product version

Three consequences of this design:

Freshness without duplication

The APQR agent asking about batch B-2026-0417 does not read a stale graph copy of the batch. It reads the batch_phase_summary and test_results rows for that batch, projected into triplets on demand, each carrying the row's provenance and the data product's version. The persistent graph supplies the meaning (which parameter is a CPP, what the validated range is, which specification version was effective on the batch date); the lake supplies the facts.

Data fabric earns its place where systems are federated

Regional LIMS instances, site-specific historian tag names, CMO results arriving as spreadsheets — a virtualization layer with a canonical attribute dictionary presents one logical test_results service over all of them. The dictionary is quality-owned work, under change control, joined to specifications by effective date, never by current version. It is the single most underestimated line item in a manufacturing AI program, and the one that makes cross-site comparison a GROUP BY site instead of a heroic project.

Transactional systems keep their authority

The MES remains the system of record for the batch; the LIMS for the result; the QMS for the deviation. Agents read from data products and write decisions back to the workflow tool the human uses — never directly into a validated transactional system. That boundary is what keeps the agent out of the transactional system's validation scope.

Batch genealogy, LIMS quality parameters, and historian process parameters are the three data products that manufacturing use-cases cannot live without. Genealogy resolves lot relationships (drug-substance lot to drug-product lots, raw-material lot to batch); the LIMS conformed services (samples → tests → test_results, given meaning by specifications) carry the CQAs; the historian-derived batch_phase_summary carries the CPPs. Join them and the CPP-to-CQA relationship — the scientific point of APQR and CPV — becomes a query. §8.4 is built on that join.

Classification-aware data fabric — residency is enforced where the data is served

The fabric described above unifies heterogeneous systems. It must also partition regulated ones, and the two jobs are done by the same layer so that there is one place to inspect. Every data product carries a residency classification alongside its owner and contract, and the fabric enforces it at the point of service — not at the agent, not at the network perimeter.

ClassificationWhat it coversWhere it livesHow the fabric serves it
RegionalPatient-level data (subjects, ICSRs, genomic results), HCP interaction records, prescribing and claims-derived data, MSL note content, inquiry text, employee dataThe regional zone of origin — Azure West Europe, East US 2, China North, Japan East, Brazil South — with no cross-border replica, including for DR (DR pairs stay inside the jurisdiction)Served only to agents whose card declares a matching regional scope (§6). A global agent's request triggers one of two paths: denial (default), or automated de-identification + human attestation — HIPAA Safe Harbor / Expert Determination for US data, GDPR Recital 26 anonymization standard for EU data, with a named data steward signing that the output is no longer personal data before it is released
RestrictedGenomic and genetic data (GDPR Art. 9(1) explicit enumeration), biometric identifiers, Chinese patient and HCP data subject to PIPL Art. 38–40 — any data category where REGIONAL classification plus de-identification is insufficient because the law or the ethics-committee approval is site-specific and non-transferableSame residency zone as REGIONAL — no cross-border replicaREGIONAL enforcement plus a mandatory second gate: ethics-committee approval with per-site scope, specific ICF consent language covering the derivative use, or a CAC security assessment with versioned reference. The distinction from REGIONAL is the exit path: REGIONAL data has a de-identification pathway to SHAREABLE-DERIVED; RESTRICTED data may have no pathway at all — Chinese genomic data under HGRAC is non-exportable, and German genetic data under GenDG requires per-subject ethics approval that does not transfer across sites
ShareableManufacturing quality (batches, results, specifications, genealogy), aggregated trial metrics (enrollment curves, site performance), HCP professional master data, regulatory intelligence, publications, ontologies, signed rulesReplicated globally; single logical serviceStandard platform contract; classification tag travels with the row
Shareable-derivedOutputs of a de-identification or aggregation step over REGIONAL inputs (insight themes, signal statistics, cohort counts)GlobalServed globally, but the row carries the lineage edge to its REGIONAL source and the attestation record; re-identification queries are blocked by the k-anonymity floor at the fabric
Mandated-flowICSRs (E2B(R3) format), SUSARs, annual safety reports — data that must cross borders under pharmacovigilance law regardless of privacy restrictionsOriginates regionally; transmitted to health authorities (EudraVigilance, FAERS, NMPA) via a separate safety-database gateway, never via the general-purpose data fabricDedicated pipeline through the validated safety database and its E2B(R3) submission gateway. The fabric does not touch it. The classification exists to prevent anyone from routing a case through the fabric "for convenience" — the moment an ICSR enters the fabric, the separation of duties that inspectors expect is broken

Three rules follow from the table:

Logical unification, physical separation

Regional LIMS, historian, and ERP instances are unified through the canonical attribute dictionary so that test_results is one logical service — but the unification is virtual. The fabric pushes the query down to each regional instance and joins the projections; it does not copy classified rows across a border to make the join convenient. For manufacturing data this is rarely a constraint (batch results are SHAREABLE). For the clinical and medical data products it is the constraint, and the reason the fabric — not a replication pipeline — is the integration pattern.

Consent is a checkpoint, not a column

For GDPR Art. 9 special-category data (health, genetic, biometric) and for any HCP data used beyond its collection purpose, the fabric calls the consent management platform before serving a row to any agent, and serves only rows whose consent scope covers the agent's declared purpose (from its card). A revoked consent is effective at the next query, not at the next batch refresh. The consent check is itself logged (G3) with the purpose evaluated, so that an Art. 15 subject-access request can be answered from the audit trail.

China is an asymmetric zone

PIPL localizes personal information and "important data" inside China; export requires a CAC security assessment (PIPL Art. 38, the 2024 cross-border data-flow provisions), and human genetic resources cannot leave under HGRAC rules at all. The fabric therefore treats the China North zone as REGIONAL-with-no-derivation-path by default: even SHAREABLE-DERIVED outputs require an explicit, versioned CAC assessment reference on the data product before the fabric will serve them outside the zone. Every other regional zone has a de-identification path to global; China's path is an approval, not an algorithm.

Contractual restrictions carry equal force. Data residency enforcement is not purely a privacy-law question. Prescribing data licensed from a third-party data provider (the industry's analogs of IQVIA and Symphony Health) typically prohibits export and secondary use outside the licensed geography under the data-use agreement — a contractual, not regulatory, restriction that the fabric must enforce identically. In some jurisdictions the restriction is harder than contractual: France prohibits prescriber-level prescribing databases for commercial promotion (CSP Art. L.4113-7), which means the fabric's classification for French prescribing is not REGIONAL for contractual reasons but RESTRICTED — the data exists only in aggregate form, and the agent that needs individual prescriber context for a French HCP must infer it from other signals (publication record, congress attendance, inquiry history) rather than from a dataset that does not exist. Similarly, CRO-provided site-performance data may carry confidentiality terms that restrict cross-sponsor use. The fabric's classification tag must therefore encode both legal basis and contractual basis, and the routing logic must check both. A data product tagged REGIONAL for GDPR reasons and one tagged REGIONAL for DUA reasons look the same to the agent — which is the point: the enforcement is uniform, the reason is metadata.

The approved-source registry

Nothing outside the registry can be cited by an agent. It is versioned and access-controlled, and it is scoped per use-case: the regulatory-documentation agent's registry contains locked TFLs, the approved protocol, and the IB; the field-force agent's registry contains MLR-approved assets only; the medical-inquiry agent's registry contains medical-affairs-approved responses. Same mechanism, different contents — which is the whole point of a platform.

Components in the layer: SDTM/ADaM clinical data · Regulatory feeds (FDA/EMA/ICH/WHO) · Real-world evidence · CTMS/EDC/IRT · CRM/MDM · QMS/LIMS/MES · Historians & PAT · ERP/supplier data · Enterprise data lake & data products · Data fabric / virtualization · Approved-source registry.

Every layer described so far is only trustworthy because of the plane that cuts across them.

5.5 Cross-Cutting Governance Plane

Governance in v1 of this architecture was a list of boxes. That is how governance usually gets drawn, and it is why governance usually gets ignored: a flat list gives an executive no way to ask "which of these do we actually have, and which ones matter for this use-case?" The plane is better understood as eight categories, each answering one question an inspector, a Quality head, or a CFO will eventually ask. Click a card for the full control list; each §8 use-case card carries a compact G1–G8 strip showing which categories apply at what intensity.

G1 · PROVENANCE
Evidence Provenance
Where did this come from, and how much should I trust it?
Approved-source registrySpan-level citationsTriplet provenanceTrust-tier caps
+ full control list
Approved-source registry · Span-level citations on every generated sentence · Triplet-level provenance (source, tier, confidence, timestamp, version) · Trust-tier caps on inferred edges.
Primary owner: Semantic & Knowledge Engineering
G2 · PROVENANCE
Decision Provenance
What did the human do with it, and why?
Accept / modify / rejectRationalePreference signalCoaching record (acceptance curve)
+ full control list
Accept / modify / reject capture on every recommendation · Rationale capture at sign-off · Preference signal from edits · The coaching record: the acceptance-rate curve that decides whether an agent graduates (§18).
Primary owner: Agentic Factory + domain owners
G3 · AUDIT
Traceability & Audit
Can you reconstruct this decision for an inspector in fifteen minutes?
21 CFR Part 11EU Annex 11E-signaturesExecution historyALCOA+Object lock
+ full control list
21 CFR Part 11 audit trail · EU Annex 11 electronic records · Electronic signatures · Per-execution history in stateful orchestration · ALCOA+ on every artifact · Immutable storage (object lock).
Primary owner: Quality / CSV
G4 · OBSERVABILITY
AI Observability
Is the system behaving today the way it behaved when we validated it?
TracingToken / cost accountingLatency & velocityDriftCalibration
+ full control list
Distributed tracing per agent invocation · Token and cost accounting per agent, per use-case, per domain · Latency and decision-velocity dashboards · Model drift and eval-score drift detection · Confidence calibration monitoring (reliability diagrams, Brier score).
Primary owner: Agentic Factory / Platform SRE
G5 · ACCESS
Identity & Access
Who — or what — is allowed to touch this?
Agent as principalRBACTool allowlistData classificationPartner onboarding
+ full control list
Agent as first-class identity principal · Role-based access control · Tool allowlist per agent identity · Data classification enforcement · Partner-agent onboarding under the platform contract.
Primary owner: Security / IAM
G6 · VALIDATION
Validation & Change Control
How do we know it works, and what happens when it changes?
GAMP 5 Cat 5CSA controlsModel pinningRule versioningRelease trains
+ full control list
GAMP 5 classification per component (Cat 5 for deterministic paths) · CSA-aligned controls for LLM paths · Model version pinning · Rule versioning with owner and calibration record · Release trains with regression gates · Change impact on validated state.
Primary owner: Quality / CSV + Factory
G7 · SAFETY
Autonomy & Safety Controls
When does the human take the wheel?
Confidence gatesHITL gatesDegraded modeCircuit-breakersKill switches
+ full control list
Confidence-gated autonomy (≥0.90 autonomous · 0.70–0.89 notify · 0.55–0.69 suggest · <0.55 escalate) · Human-in-the-loop gates at every regulated decision · Degraded-mode design (connector failure → confidence caps → human takes over, automatically) · Rate limits and cost circuit-breakers · Kill switches per agent.
≥0.90 autonomous0.70–0.89 notify0.55–0.69 suggest<0.55 escalate
Primary owner: Factory + domain owners
G8 · RESPONSIBLE AI
Responsible AI & Regulatory Alignment
Would we be comfortable explaining this to a regulator, a patient, or a journalist?
Bias monitoringPrivacyAgent cardsGuidance trackingPSMFAnnex III classify-upDPIA per agentResidency scope
+ full control list
Bias monitoring (site concentration, population representation) · Privacy controls (PHI, GDPR) · Model cards and agent cards in the registry · Regulatory-guidance tracking (FDA AI draft guidance, EMA reflection paper, EU AI Act classification) · AI use documented where required (PSMF for PV) · EU AI Act Annex III classify-up policy for agents touching trial participants, health-data profiling, safety-critical PV, or HCP profiling (Art. 22 GDPR) · DPIA-per-agent (GDPR Art. 35) as a Lab-to-Factory gate · Regional compliance as configuration — jurisdictional requirements as versioned KG edges, not code forks · Residency classification enforced at the fabric (§5.4) and on the agent card (§6).
Primary owner: AI Governance Council

Three categories deserve emphasis because they are the ones most architectures leave out.

Responsible AI (G8) is a set of gates, not a set of principles

Three of them are structural.

  • Annex III classify-up. The EU AI Act's Annex III list (Art. 6(2)) and its Commission guidelines will keep moving; an architecture that waits for the final word is one that re-validates every agent when it lands. The platform's policy is to treat as high-risk, by default, any agent in four classes: protocol-design and trial-design agents that shape which participants are enrolled and what they are exposed to; health-data profiling agents that score or segment patients or subjects (the clinical-operations intelligence class); safety-critical PV case-processing agents (UC-X1); and commercial agents that profile HCPs in a way that produces legal or similarly significant effects — which brings GDPR Art. 22 with it regardless of the AI Act's outcome. High-risk means the Art. 9 risk-management system, Art. 10 data-governance record, Art. 12 logging, Art. 14 human-oversight design, and Art. 26 deployer obligations are all satisfied by the platform contract — the confidence bands (G7), the traceability plane (G3), the agent card (§6) — so that classification is a flag on the card, not a project.
  • DPIA per agent. Every agent whose card declares a PII scope or a REGIONAL-* residency scope (§6) carries a completed Data Protection Impact Assessment before it can transition candidate → factory. The DPIA template is itself a platform artifact: agent pattern and kit; data categories and Art. 9 flags; legal basis per jurisdiction (consent, legitimate interest with balancing test, public-health exemption — as a table, one row per region the agent runs in); cross-border transfer mechanism per region pair (adequacy decision, SCCs with transfer-impact assessment, BCRs, CAC assessment reference); residual-risk assessment with the DPO's sign-off; and the re-assessment trigger — which changes to the card re-open the DPIA. The four-lane approval workflow (§6) gains a fifth reviewer, the DPO, for these cards.
  • Regional compliance as configuration. The regulatory ontology in the core KG holds jurisdictional requirements as versioned edgesrequirement ──applies_in──► jurisdiction, requirement ──effective_from──► date, requirement ──supersedes──► prior_version — with the same trust-tier and provenance discipline as any other triple. A protocol-compliance agent checking a study with sites in Germany, Ohio, and Shanghai does not run three code paths; it traverses the regulatory ontology for each site's jurisdiction and applies the rules it finds, each carrying its citation. When PIPL's implementing rules change, the ontology team publishes a new edge version; no agent is redeployed. This is the mechanism behind the §8.7 posture that rulesets are "configuration tagged by jurisdiction, never regional code forks."

Decision provenance (G2) is the coaching mechanism. Evidence provenance tells you what the agent saw. Decision provenance tells you what the expert did about it — and the aggregate of those decisions is the only honest measure of whether an agent is helpful. An agent whose recommendations are accepted 40% of the time is not ready, no matter how good its eval scores. An agent whose acceptance rate climbs monotonically over rolling two-week windows is learning the domain's boundaries. G2 is what turns human review from a cost into a signal.

AI observability (G4) is what makes the Factory a factory. Without per-agent, per-invocation tracing and cost accounting, the platform cannot tell which use-cases are paying for themselves, cannot detect drift before Quality does, and cannot answer the CFO's question. Observability is not a dashboard; it is the instrumentation that makes the other seven categories measurable.

The governance plane applies to every agent. The registry is where each agent declares which parts of it apply — including the one class of work where the model, not the query, has to travel.

5.6 Federated Learning — When the Model Must Travel Because the Data Cannot

Everything above the line keeps patient data inside its region by routing queries. There is one class of work that routing does not solve: training a model whose signal is spread across regional patient populations — a safety signal that is sub-threshold in every region and significant across them; a biomarker whose prevalence differs by ancestry; a recruitment-density model that needs every region's site history. For that class, the platform inverts the usual pattern: the data stays, the model moves.

                    AGGREGATION HUB — Basel (Azure West Europe)
                    receives:    DP-noised, encrypted model updates only
                    holds:       global model weights · per-node privacy ledger · round audit trail
                    never holds: a patient record, a gradient in the clear
                              ▲            ▲            ▲            ▲            ▲
             secure aggregation (SMPC or TEE): updates summed without any party decrypting an individual one
                              │            │            │            │            │
              ┌───────────────┴──┐ ┌───────┴─────┐ ┌────┴──────┐ ┌───┴────────┐ ┌─┴──────────┐
              │ EU node          │ │ US node     │ │ JP node   │ │ BR node    │ │ CN node    │
              │ West Europe      │ │ East US 2   │ │ Japan East│ │ Brazil S.  │ │ China North│
              │ trains on EU     │ │ HIPAA PHI   │ │ APPI      │ │ LGPD       │ │ PIPL/HGRAC │
              │ patient data     │ │             │ │           │ │            │ │ ASYMMETRIC │
              │ exports: noised  │ │ same        │ │ same      │ │ same       │ │ local-only │
              │ gradients, ε-    │ │             │ │           │ │            │ │ by default;│
              │ metered          │ │             │ │           │ │            │ │ opt-in per │
              └──────────────────┘ └─────────────┘ └───────────┘ └────────────┘ │ CAC assmt. │
                                                                                 └────────────┘

Four controls make it defensible, and each is auditable (G3):

Patient data never leaves the zone

Each regional training node runs inside its residency zone against the REGIONAL data products the fabric serves it (§5.4), under a card with REGIONAL-* scope (§6). What crosses the border is a model update — and only after differential-privacy noise is added at the node, before export. The hub is FEDERATED-scoped: it may receive aggregates, never records.

Privacy budget is metered per node

Every export spends epsilon; the hub's privacy ledger tracks cumulative ε per node per model lineage, and the node's DPIA (G8) sets the ceiling. When a node's budget is exhausted its contributions stop — automatically, not by policy reminder — until the DPO approves a new budget with a documented rationale. The ledger is what the answer to "how much have we learned about EU patients, in total" looks like when a supervisory authority asks.

Secure aggregation, so the hub cannot see an individual update

Updates are encrypted at the node and summed under secure multi-party computation or inside a trusted execution environment (Azure confidential computing), so that the hub decrypts only the sum. A single region's update — which can leak more than its noise budget suggests when the region is small — is never observable on its own.

China participates asymmetrically

PIPL localization and HGRAC make the China North node a fully local training node by default: it trains a local model on local data and contributes nothing to the hub. Participation in a federated round is an opt-in per model lineage that requires a CAC security assessment reference on the export, recorded on the node's card; the noised gradient is treated as a data export for that assessment, because the regulator will. The architecture does not assume China joins; it makes joining a configuration change rather than a redesign when the assessment is in hand.

Where the platform uses it

Use-caseRegional nodes train onHub learnsWhy routing alone was not enough
Federated signal detection (UC-X1)Regional ICSR databasesA disproportionality model that surfaces signals sub-threshold in every single regionAggregated counts per region lose the case-level covariates (age, co-medication, indication) that make a weak signal credible
Federated patient-density estimation (UC-D2)Regional claims-derived and EHR-derived cohort data, regional site historyA site-selection model that predicts enrollment from eligible-patient densityDensity is a patient-level attribute; the regional partitions cannot export it, and site history without it is what makes site selection a master-data problem rather than a prediction problem
Federated biomarker analysis (UC-R1, UC-X6)Regional genomic datasets — including China's, which cannot export under HGRACBiomarker-response associations across ancestriesGenomic data is the most localized data class the enterprise holds; a model trained on one region's ancestry mix is the bias the G8 monitoring exists to catch

The registry (§6) treats a federated model lineage as an agent: it has a card, a residency scope of FEDERATED, a DPIA, a privacy ledger, and a lifecycle. That is what makes the next section's discipline apply to it unchanged.

Section 6 · Agent Registry & Service Catalog

From "we built an agent" to "the enterprise has a capability"

A platform without a registry is a folder of prototypes. The registry is what turns "we built an agent" into "the enterprise has a capability" — discoverable, composable, governed, and billable. It is also the mechanism that makes Lab-to-Factory promotion (§7) a state transition rather than a re-implementation.

What is registered

Every agent, every MCP tool server, and every reusable service kit is registered with a signed agent card — the machine-readable contract that other agents, workspaces, and reviewers read.

Agent card — illustrative

InternalConfidentialPII← declared on every card at design time
# Agent card — illustrative
id: mfg.cpv.signal-agent
version: 2.3.0
archetype: monitoring
owner: MS&T Data Products (name, role, escalation)
domain: manufacturing
purpose: "Detect SPC rule violations and capability drift on batch_phase_summary; route to responsible scientist"
gxp_class: GMP / Part 11 (selective — signal disposition is a regulated record)
lifecycle_state: factory            # draft | lab | candidate | factory | deprecated

service_signature:
  inputs:
    - data_product: mfg.batch_phase_summary@>=4.1
    - data_product: mfg.test_results@>=2.0
    - kg_scope: manufacturing.cpv (read)
  outputs:
    - signal: cpv.rule_violation (schema v1.2, with provenance)
    - signal: cpv.capability_drift
  tools_visible_to_model: [spc_calculator, capability_calculator, kg_lookup]   # ≤5
  model: scoped-small@pinned-2026-07
  confidence_policy: {autonomous: ">=0.90", notify: "0.70-0.89", suggest: "0.55-0.69", escalate: "<0.55"}

eval_record:
  harness: mfg.cpv.eval@1.4
  ground_truth: 2,140 labeled batch-phases, 3 sites, 2 products
  precision: 0.93   recall: 0.88   f1: 0.905
  last_run: 2026-08-29
  coach_mode_acceptance: 0.81 (rolling 14d, n=212)

governance:
  evidence_provenance: G1 (triplet + row-level)
  decision_provenance: G2 (disposition captured)
  audit: G3 (Part 11 trail on disposition)
  observability: G4 (traced, cost-tagged)
  identity: G5 (principal mfg-cpv-signal; allowlist above)
  validation: G6 (GAMP 5 Cat 5 for SPC path; CSA for narrative)
  autonomy: G7 (policy above; degraded mode on historian connector loss)
  responsible_ai: G8 (agent card published)

approvals:
  security_review: approved 2026-07-14 (ticket, reviewer)
  quality_review:  approved 2026-08-02 (CSV lead)
  data_steward:    approved 2026-07-20 (MS&T data-product owner)
  business_owner:  approved 2026-08-05 (Head of MS&T)

usage_and_billing:
  invocations_30d: 18,420
  cost_30d: (tokens + compute, tagged to cost center MFG-QA)
  consumers: [mfg.apqr.orchestrator, mfg.quality-command-center, mfg.deviation.investigator]

Lifecycle state machine

draft lab candidate factory deprecated killed PoC → deprecated card with eval record + post-mortem signature change re-opens reviews PROMOTION GATE — all four lanes approved Any lifecycle transition and any change to a factory card triggers the four-lane review workflow. No consumers for 90 days → candidate for deprecated. The curated garden stays curated.
Securityidentity · allowlist · data classification
QualityGxP class · validation posture · audit trail
Data stewarddata-product contracts · approved-source scope
Business ownerpurpose · metric · adoption

The four lanes above are the review workflow described under "Approvals are a workflow, not a wiki page" below.

Service kits — the templatized unit of reuse

Individual agents are too fine-grained to be the unit of reuse. The unit of reuse is the service kit: a templatized bundle that a new use-case instantiates rather than builds.

Kit componentWhat it containsWhy it accelerates use-case N+1
Data contractNamed data products consumed, schema versions, freshness SLAs, approved-source registry scopeThe new team spends days mapping, not months integrating
Orchestrator templateStateful workflow skeleton (intake → specialists → merge → human gate → decision log), with hooks pre-wired for audit and observabilityControl flow is deterministic and already validated once
Specialist slotsTyped slots for extraction / reasoning / drafting / monitoring specialists, each with the ≤5-tool constraint enforcedComposition is the default; the monolith is structurally impossible
Eval harnessGround-truth loader, P/R/F1 per class, confidence calibration, acceptance-rate tracker, regression gate"If I can't give you precision and recall, I haven't finished" — from day one
Governance manifestWhich G1–G8 controls apply, pre-filled by GxP classThe Quality review starts from a known posture, not a blank page
Experience templateWorkspace shell with findings / provenance / accept-modify-reject built inDecision provenance is captured without the team remembering to

Examples of kits the platform ships:

Draft-then-check

Draft-then-check kit — drafting agent + deterministic checker (as an MCP tool) + KG-grounded reviewer + Part 11 sign-off. Instantiated for regulatory documentation (§8.2), PV narratives (§8.7), deviation reports (§8.4), tech-transfer packages (§8.3).

Rule-gate

Rule-gate kit — expert-calibrated rule engine with provisional → confirmed calibration loop, versioned rules with owners. Instantiated for protocol compliance, structural alerts in chemistry, seriousness/expectedness in PV, batch-record completeness.

Signal-and-NBA

Signal-and-NBA kit — isolated monitoring specialists + orchestrator that de-duplicates, prioritizes, composes next-best-actions with rationale/impact/owner, and suppresses noise from feedback. Instantiated for trial operations (§8.2), CPV (§8.4), supply risk (§8.5), and every command center (§12).

Evidence-dossier

Evidence-dossier kit — KG traversal + deterministic scoring tiers + cited synthesis → reviewable dossier. Instantiated for target validation (§8.1), site selection (§8.2), supplier qualification (§8.5).

Entity-resolution

Entity-resolution kit — cross-source matching, golden record, crosswalk maintenance. Instantiated for HCP/HCO, site/investigator, material/supplier, batch genealogy.

Agent scope classification — design-time access boundaries

Kits are not only discovered — they are used, and usage requires knowing what an agent is allowed to touch. Every agent card declares a scope classification at design time, analogous to how datasets are classified but promoted to the agent as a first-class property:

ScopeWhat the agent can accessRuntime enforcementExample
InternalAll platform data products within the agent's domain partition; broadest tool access; minimal frictionStandard platform contract (G5 identity, G3 audit)A Lab PoC agent exploring CPV trends across all sites
ConfidentialRestricted to named data products; tool access narrowed; outputs carry a classification tagData-product ACL check on every retrieval; output tagged for downstream handlingA deviation agent accessing supplier quality records under NDA
PIIPatient, HCP, or employee data — HIPAA, GDPR, or equivalent appliesDe-identification gate before any output leaves the agent; consent check on retrieval; no caching of raw PII in agent state; audit trail includes data-subject categoryA PV agent processing individual case safety reports; an HCP engagement agent accessing prescribing history

The classification is declared on the agent card, reviewed during the four-lane approval workflow (§6), and enforced at runtime by the governance plane (G5 identity + G1 approved-source + G3 audit). An agent classified Internal that attempts to access a PII data product is denied — the boundary is structural, not policy. This is the same principle that makes dataset classification work, extended to the entity that reasons over the data rather than the entity that stores it.

Scope escalation in multi-agent composition. The design-time classification is clear; what requires an explicit rule is how scope transitions when agents compose. When a Confidential-scoped orchestrator invokes an Internal-scoped specialist and passes it context that contains Confidential data, the specialist's effective scope elevates to the caller's scope for the duration of that invocation only. The elevation is not a grant — it is a delegated, time-boxed inheritance, and it is the audit event that matters.

The audit trail for every elevated invocation captures six fields: (1) caller identity — orchestrator agent ID, version, service principal, and the human session (if any) that triggered it; (2) callee identity — specialist agent ID and version, with its baseline scope recorded alongside the elevated scope; (3) timestamps at invocation and return (contemporaneous — ALCOA+); (4) structured purpose code from the DPIA's purpose register (G8), not free text; (5) context hash — a hash of the context payload, so an inspector can verify whether Confidential data actually crossed the boundary or only a reference to it did; (6) region binding — the elevated invocation must execute in the same region as the caller. A West Europe orchestrator cannot elevate an East US 2 specialist; the fabric rejects the call.

Three consequences follow. First, specialists must be deployable in every region where their orchestrators run, even if they only need Internal data — because elevation requires co-location. Second, elevated invocations are a first-class metric in the EU AI Act Art. 12 logging requirement and the Art. 72 post-market monitoring plan; a spike in elevations is a design smell — the orchestrator is passing more Confidential context than the specialist's task requires (GDPR Art. 5(1)(c) data minimization). Third, the DPIA for the orchestrator (G8) must enumerate every specialist it can elevate, with the purpose codes it may declare. Adding a new specialist to an orchestrator's call graph is a DPIA change-control event.

Regional scope enforcement — a second axis on the card

Beyond data classification (Internal / Confidential / PII), every agent card declares its deployment region and data residency scope. The two axes are independent: a PII agent may be REGIONAL-EU; an Internal agent may be GLOBAL. The residency scope is what the fabric (§5.4) and the regional KG endpoints (§5.3) check before serving a row.

Residency scopeWhere the agent runsWhat data it may accessCross-border output
GLOBALAny regionNon-PII data products only — reference data, ontologies, signed rules, SHAREABLE and SHAREABLE-DERIVED products, aggregated metricsUnrestricted
REGIONAL-EUAzure West Europe only (North Europe for DR)EU patient data, EU HCP interactions, EU clinical and safety dataOutputs containing personal data stay in the EU; de-identified outputs (GDPR Recital 26 standard, steward-attested) may be promoted to SHAREABLE-DERIVED
REGIONAL-USAzure East US 2 onlyUS patient data (HIPAA PHI), US HCP interactions, US prescribingSame mechanism — Safe Harbor or Expert Determination before promotion
REGIONAL-CNAzure China North onlyAll personal information and important data of Chinese origin (PIPL); no genomic export (HGRAC)No cross-border output of any kind without a CAC security-assessment reference on the output data product
REGIONAL-JPAzure Japan East onlyJapanese patient data (APPI), J-GCP clinical data, Japan HCP interactions, PMDA submissionsExit via EU–Japan mutual adequacy for EU-bound transfers; APEC CBPR for other destinations; de-identified outputs follow the same steward-attestation path as EU data
REGIONAL-BRAzure Brazil South onlyBrazilian patient data (LGPD Art. 11 sensitive health data), ANVISA-regulated clinical data, Brazil HCP interactionsExit via ANPD standard contractual clauses (finalized August 2024; grace period ended August 2025) or de-identification; CONEP ethics committee oversight on research-use data
FEDERATEDOrchestrator runs globally; data access happens regionallyThe orchestrator itself accesses only GLOBAL data; it dispatches sub-queries to REGIONAL-* specialists and receives aggregates that satisfy the k-anonymity floorThe orchestrator never holds an individual record; the audit trail shows the sub-query, the region, and the classification decision each regional specialist made

The FEDERATED scope is the architecturally interesting one, and the one a multinational cannot do without. A federated signal-detection orchestrator (UC-X1) runs in Basel and asks the same question of the EU, US, China, and Japan PV specialists; each specialist processes individual cases inside its region and returns disproportionality statistics — a reporting odds ratio, a case count above the floor, a MedDRA PT — not case data. The orchestrator sees a signal, not a patient. The same scope serves federated patient-density estimation for site selection (UC-D2) and federated theme aggregation from field medical insights (UC-X3).

Composition rules follow. A FEDERATED orchestrator may invoke REGIONAL-* specialists; a REGIONAL-EU agent may not invoke a REGIONAL-US one; a GLOBAL agent may invoke neither. The registry refuses the composition at discovery time, before any data moves — which is the point of putting the scope on the card rather than in a policy document.

Agents find agents

The registry exposes discovery and usage as an API. An orchestrator planning a workflow queries the registry for specialists by capability, domain, GxP class, scope classification, and eval score — and composes from what it finds. A deviation-investigation orchestrator does not implement CPV trend lookup; it discovers mfg.cpv.signal-agent, checks that its lifecycle state is factory, its scope classification is compatible, and its eval record is current, and invokes it through the platform contract. The discovery call and the invocation are both traced (G4) and both subject to the caller's allowlist (G5).

This is how the registry enforces composition: an orchestrator may only invoke registered specialists, and a specialist may only expose the tools on its card. There is no path by which an agent grows a hundred tools, because the card is the contract and the contract is reviewed.

Approvals are a workflow, not a wiki page

Lifecycle transitions — draft → lab → candidate → factory, and any change to a factory card — trigger a review workflow with four named reviewers: security (identity, allowlist, data classification), quality (GxP class, validation posture, audit trail), data steward (data-product contracts, approved-source scope), and business owner (purpose, metric, adoption). A card cannot reach factory with any approval missing, and a change to a factory card's service signature re-opens the relevant reviews. The workflow is where the human team reviews the agent's security and quality posture — not in a meeting, but against a card that says exactly what the agent can touch and how well it performs.

Usage tracking and chargeback

Every invocation is metered — tokens, compute, tool calls, latency — and tagged to the consuming use-case, domain, and cost center. This closes three loops at once:

Economics

The CFO can see cost per decision, per use-case, and whether the marginal cost of use-case N+1 is actually falling (§18).

Reuse measurement

The registry knows which kits and specialists are consumed by how many use-cases — the reuse ratio is a query, not an estimate.

Retirement

An agent with no consumers for ninety days is a candidate for deprecated. The registry is how the curated garden stays curated.

Proven pattern

MCP server published on the MCP registry (Linux Foundation AAIF, December 2025) — the interoperability contract partner agents and internal agents both speak. 18+ reusable infrastructure modules underlying 8 production systems: a new system stands up in days because the kit already exists.

With the registry in place, "promotion to production" becomes a defined state transition with defined reviewers. That is the operating model.

Section 7 · Lab → Factory

The operating model: people, process, technology — in that order

Every pharma executive knows the mantra: people, process, technology — in that order. Agentic AI programs fail when they invert it. The Lab-to-Factory model is, first, a people model: who comes together, for what purpose, with what authority. Then a process model: how a use-case moves from idea to governed production, and what the support commitments are on the far side. Only then a technology model — and the technology (§5, §6) is designed to make the people and process model cheap to run.

THE LAB · 2–4 week iteration cycles · evaluation logged, not validated THE FACTORY · release trains · run as a service for years BUSINESS TEAMsponsor · decision-makers FUNCTIONAL / DOMAIN TEAMSMEs who will use the agent TECH TEAMplatform · agent · KE · data · CSV VIRTUAL POD — purpose-scoped · time-boxed · meets in the artifacts, not in a room Ownsthe decision · the metricacceptance criteriathe promotion vote (funding) Must not be asked todesign the agenttune the promptevaluate the model Ownsco-design of behaviorrule calibration (provisional→confirmed)ground-truth labels · coach-mode reviewthe readiness vote Must not be asked tobecome AI engineers Ownsinstantiate the kit · build specialistswire data contracts · run eval harnessproduce the validation package Must not be asked todecide whether the agent is ready —that is the domain team's call FORWARD-DEPLOYED ENGINEERleads the pod · sits with the domaincarries context in both directions:toxicologist's need ↔ kit · specialist · data SHARED CENTRAL BENCHknowledge engineers · CSV · evalallocated to pods by platform leadershipso ontology decisions stay consistent LAB TOOLS kit instantiationKG sandboxdata-product catalogprompt playground(137× per task)synthetic + historicalground trutheval harness templatecoach-mode deployment KILL CRITERIAP/R/F1 flat 2 cyclesacceptance flatdemand test fails→ post-mortem card DEMAND TESTvalue / Type 1 / Type 2which condition theagent removes vs.absorbs faster;no industrializinga workaround PROMOTION GATE · all criteria required · chaired by Factory lead GATE CRITERIA (all required) 1 Eval record — P/R/F1 per class vs expert ground truth; calibrated 2 Coach-mode acceptance ≥70% sustained, rising (domain team verifies) 3 GxP classification; G1–G8 manifest complete; validation posture agreed 4 Data contracts stable — versioned, owned, SLA-backed 5 Composition compliance — ≤5 tools; deterministic control flow 6 Security review — identity, allowlist, classification 7 Named owner + support model — on-call, SLA 8 Demand baseline measured; which condition the agent changes VOTES → funding: business sponsor · readiness: domain team · validation: Quality/CSV Any one can hold. Escalation: steering committee, 2-week SLA. Lab lead does not vote on own PoC. FACTORY SUPPORT MODEL L1 · domain super-users2–3 per use-case · triagefeedback · coach re-entrysame business day L2 · Factory opsSRE + FDE rotationconnectors · drift · costhours · 24×7 for CC feeds L3 · pod re-formationrule recalibration · ontologymodel tier · re-validationscheduled into a release train Change control · Quality/CSVsignature / model / rule / contract change →impact → regression eval → re-approval Model lifecycle · quarterlyevery agent has a tested successormodel before the vendor's EOL date CLOSED-LOOP EVALUATION — three loops, continuously Model-loop (automated evals per release) · Dogfooding-loop (power users flag first) Production-loop (live samples → Reflect → Synthesize → change control) · a defect must pass all three Wildflower meadow → curated garden: every plant has a purpose, a place, a caretaker.
Three constituencies, one virtual pod, one gate. The pod forms around a use-case, runs through Lab and the gate, hands off to the Factory support model, and dissolves. The readiness vote belongs to the domain team, the validation vote to Quality, the funding vote to the business sponsor; any one can hold.

7.1 People — three teams, one virtual pod

Every use-case is delivered by a virtual pod that brings three constituencies together for a purpose and dissolves when the purpose is met. The pod does not need to be co-located, does not need to be a permanent org unit, and does not need to be large. It needs to be named.

ConstituencyWhoWhat they own in the podWhat they must not be asked to do
Business teamThe sponsor and the decision-makers whose calendar the use-case compresses — the Head of MS&T, the clinical-operations lead, the medical-writing headThe decision, the metric, the acceptance criteria, the promotion voteDesign the agent, tune the prompt, evaluate the model
Functional / domain teamThe subject-matter experts who will use the agent — process scientists, toxicologists, medical writers, PV physicians, QA leadsCo-design of the agent's behavior; rule calibration (provisional → confirmed); ground-truth labeling; coach-mode review; the promotion gate itselfBecome AI engineers
Tech teamPlatform engineers, agent engineers, knowledge engineers, data-product owners, CSV/validation engineersInstantiate the kit; build the specialists; wire the data contracts; run the eval harness; produce the validation packageDecide whether the agent is ready — that is the domain team's call

The pod is led by a Forward-Deployed Engineer (FDE) — an engineer who sits with the domain team, speaks the domain's language, and owns the translation between "what the toxicologist needs" and "which specialist, which kit, which data product." The FDE model is what makes the Lab fast: instead of requirements documents crossing an organizational boundary, one person carries the context in both directions. It is also what makes the Factory adoptable: the FDE who built the agent is the person the domain team trusts when it goes live.

Three constructs make the pod work at enterprise scale:

  • Pods are purpose-scoped and time-boxed. A pod forms around a use-case, runs through Lab and the promotion gate, hands off to the Factory support model, and dissolves. Members return to their home teams with the pattern in their heads. The next pod forms faster.
  • Pods share a bench. Knowledge engineers, CSV engineers, and eval specialists are scarce. They are pooled centrally and allocated to pods by the platform leadership, so that the ontology decision made in the CPV pod is consistent with the one made in the deviation pod.
  • Pods are virtual by default. The manufacturing SME is in one country, the FDE in another, the data-product owner in a third. The registry (§6), the kit, and the eval harness are the shared workspace; the pod meets in the artifacts, not in a room.

7.2 Process — what the Lab looks like from a use-case perspective

The Lab is not a research group. It is a rapid-prototyping environment whose job is to get from a named decision to a measured PoC in two to four weeks, and to kill PoCs that are not moving the metric.

What a pod finds when it walks into the Lab:

Lab toolWhat it doesTime saved
Kit instantiationPick the closest service kit (§6); the orchestrator skeleton, eval harness, governance manifest, and experience shell are generatedWeeks of scaffolding → hours
KG sandboxA branch of the semantic layer where the pod can extend the domain ontology, load candidate triplets, and test traversals without touching productionOntology iteration in days, with a merge path back to the federated model
Data-product catalogBrowse conformed data products, request access, bind the data contractNo bespoke extracts
Prompt & agent playgroundVersioned system prompts, tool binding, context-budget visualization, side-by-side model tier comparison (frontier vs scoped — the 137× question answered per task)Model-tier decisions made on evidence
Synthetic and historical ground truthCurated historical cases with expert labels; synthetic-data generators that mirror real distributions (not toy data) for domains where history is thinEval harness has something to measure against from week one
Eval harness templateP/R/F1 per class, confidence calibration, acceptance-rate tracker — pre-wired"Does it work?" has a number
Coach-mode deploymentShip to a small group of domain experts with thumbs-up/down on every output, in the real workspace shellAdoption signal before the Factory spends a dollar

Lab cadence and discipline

  • 2–4 week iteration cycles per PoC, against a real historical dataset and a real named decision.
  • Evaluation is logged, not validated. The Lab records everything (G3, G4 are on from day one) but does not carry the validation burden — that is what the gate is for.
  • Kill criteria are explicit. If precision/recall/F1 does not improve over two iteration cycles, or coach-mode acceptance is flat, the PoC stops. The pod writes a two-page post-mortem into the registry so the next pod does not repeat it. Killing fast is a Lab success metric, not a failure.
  • Every PoC produces a card. Even a killed PoC leaves a deprecated card with its eval record — institutional memory, not a lost notebook.
  • The demand test is a kill criterion too. Before the first iteration, the pod classifies the demand the use-case serves — how much of today's calendar is value work, Type 1 waste, Type 2 waste, and queue time — and names the system condition behind each. The PoC must then state which of those it removes and which it merely absorbs faster. An agent that assembles the same scattered evidence more quickly, leaving it scattered, copes with the condition; an agent that leaves the precedent structured in the KG for the next decision changes it. Both can be legitimate. But a PoC that can only cope, in a domain where the condition is cheap to change — a template field, an SOP step, a named data-product owner — does not proceed. The Lab does not industrialize a workaround.

7.3 Process — when and how a use-case transitions to Factory

The promotion gate is the hardest organizational problem in the model, and it deserves to be stated plainly: the Lab wants to ship, Quality says it is not validated, and the business sponsor wants the value now. Without defined criteria, a named owner, and an escalation path, PoCs accumulate and nothing reaches production. Think of it as transforming a meadow of wildflowers — beautiful but ungoverned — into a curated garden where every plant has a purpose, a place, and a caretaker.

Gate criteria (all required):

CriterionThresholdWho verifies
Eval recordP/R/F1 per class against expert-labeled ground truth, above the use-case's stated threshold; confidence calibratedTech team produces; domain team accepts
Coach-mode acceptance≥70% sustained over rolling two-week windows, with a rising curveDomain team — not engineering, not the sponsor
GxP classificationDecision's regulatory frame assigned; G1–G8 manifest complete; validation posture (§11) agreed with CSVQuality / CSV
Data contracts stableAll consumed data products at a versioned, owned, SLA-backed releaseData stewards
Composition complianceEvery specialist ≤5 tools; orchestrator control flow deterministic; cards completePlatform architecture
Security reviewIdentity, allowlist, data classification approvedSecurity
Named owner and support modelBusiness owner, technical owner, on-call rotation, SLA agreedBusiness + Factory
Demand baselineValue/failure-demand split measured (not estimated) before coach mode; the card names which system condition the agent changes and which it only absorbsDomain team + FDE

Gate ownership. The gate is chaired by the Factory lead, but the readiness vote belongs to the domain team, the validation vote belongs to Quality, and the funding vote belongs to the business sponsor. Any one of the three can hold; escalation is to the platform's steering committee with a two-week SLA. The Lab lead does not vote on their own PoC's promotion.

What changes at the gate:

DimensionLabFactory
MissionRapid PoC against real use-casesProduction-grade, governed, enterprise deployment
Cadence2–4 week iteration cyclesRelease trains with regression and validation gates
GovernanceLightweight — evaluation loggedFull — Part 11 audit trail, GAMP 5 / CSA, e-signatures
StackKit + sandbox KG + playgroundStateful orchestration, IaC, CI/CD, pinned models, production KG
MonitoringEval harnessEval harness + drift detection + cost accounting + performance dashboards + command-center feeds
ScaleOne team, one site, one studyMulti-study, multi-site, multi-TA, multi-geography
SupportThe podThe Factory support model (below)

7.4 Process — the support lens

The Factory is not "deployment." It is the commitment to run the agent as a service for years, through model deprecations, regulatory changes, and organizational turnover. The support model has to be as explicit as the gate.

TierWhoScopeTypical SLA
L1 — Domain super-usersTrained domain experts (2–3 per use-case, the champion network of §19)First-line triage: "is this a bad recommendation or a bad question?"; feedback capture; coach-mode re-entry if acceptance dropsSame business day
L2 — Factory operationsPlatform SRE + FDE on rotationConnector failures, degraded-mode events, cost anomalies, drift alerts, data-product contract breaksHours; 24×7 for command-center feeds
L3 — Pod re-formationOriginal pod members recalled for a scoped changeRule recalibration, ontology extension, model-tier change, re-validationScheduled into a release train
Change controlQuality / CSVAny change to a factory card's service signature, model version, rule set, or data contract → impact assessment → regression eval → re-approvalPer validated-state procedure
Model lifecyclePlatform architectureFrontier and scoped model deprecations tracked; every agent has a tested successor model before the vendor's end-of-life dateQuarterly review

Closed-loop evaluation is how production agents self-tune without leaving the validated state. Three loops run continuously: a Model-loop (automated evals against ground truth on every release), a Dogfooding-loop (internal power users flag failures before end-users see them), and a Production-loop (live decision samples diagnosed, failure classes reflected into prompt and rule adjustments, synthesized into the next version). A Reflect agent identifies failure classes; a Synthesize agent proposes patches — which go through change control like any other change. No single loop catches everything; a defect must pass all three undetected to reach a patient-facing decision.

7.5 Risk — the ED-level considerations

These are the risks a platform leader is accountable for, as distinct from the risks a project manager tracks.

RiskWhat it looks likeMitigation built into the model
Pilot purgatoryDozens of PoCs, none in production, sponsors losing faithKill criteria; gate with named owners and an escalation SLA; "PoCs promoted per quarter" as a platform KPI
Validation debtAgents in production with a retrofitted validation story; Quality discovers it during an inspectionGovernance manifest at Lab entry; validation posture agreed before the gate; G3/G4 on from day one
Shadow AITeams bypass the platform with their own vendor tools because the platform is slowLab time-to-first-grounded-answer measured in days; kits that are genuinely faster than going around
Three-estate fragmentationPartner agents ungoverned; three identity models; no shared observabilityPlatform contract published early (§13 Phase 1); partner onboarding as a registry workflow
Model-vendor concentrationOne frontier model in every agent; a pricing change or deprecation is an enterprise eventModel-agnostic kits; scoped models for 80% of calls; successor model tested per agent
Cost runawayToken spend grows 30–70% annually with no ownerPer-agent cost accounting; chargeback; composition rule; frontier-tier budget caps (G7)
Data-product decayThe conformed dataset the CPV agent relies on stops being maintained when the project team disbandsNamed data-product owners as a gate criterion; contract SLAs; breaks surface in L2
Regulatory changeFDA/EMA AI guidance finalizes with expectations the platform does not meetG8 regulatory tracking; validation posture documented and defensible; early engagement with Quality
TalentFDEs, knowledge engineers, and CSV engineers who understand hybrid AI are scarceCentral bench; training program (below); kits reduce the expertise needed per pod
Adoption failureTechnically excellent agent nobody usesCoach-mode acceptance as a gate criterion controlled by domain experts (§19)

7.6 Training — what to build into people, by role

AudienceWhat they need to be able to doFormat
Domain experts as agent supervisorsRead a confidence score; review a provenance trail; know when to override; give feedback that improves the agent2-day workshop + coach-mode participation. Not a six-month program.
Rule calibrators (SMEs)Take a provisional rule to confirmed: severity, scope, exceptions; sign itHalf-day + first calibration session with the FDE
Forward-deployed engineersSpeak the domain; instantiate kits; run the eval harness; write agent cards; lead a podApprenticeship on a live pod; platform certification
Knowledge engineersExtend a domain ontology under the modeling standard; write data contracts; map a LIMS to the canonical dictionaryModeling standard training; KG sandbox exercises
Quality / CSV engineersValidate a hybrid system: Cat 5 for deterministic paths, CSA-aligned controls for LLM paths; assess change impact on an agent cardJoint workshop with platform architecture; first validation package done together
Platform / Factory operationsRun L2: degraded mode, drift, cost anomalies, connector failuresRunbooks; game days
Executives and sponsorsRead the three metrics that matter per phase (§18); chair a gate; resist vanity metrics90-minute briefing; quarterly steering

Training is not a launch activity. It is a standing capability, because every pod that forms needs it and every gate depends on it.

The operating model is domain-agnostic by design. What differs by domain is the decisions, the ontologies, the data products, and the regulatory frame — which is what the next section lays out, tower by tower.

Section 8 · Cross-Functional Domain Use-Cases

Eight towers, thirty use-cases — and how few new things appear

The platform is only as credible as the decisions it accelerates. This section walks the enterprise tower by tower — Research, Development, Product Development, Manufacturing, Supply Chain, Commercial, Medical & Safety, and the Enterprise Horizontal that underlies them all — and for each tower asks the same questions: what decisions are slow today, what the domain knowledge graph has to know, which use-cases the platform serves, and how each use-case is a configuration of the shared substrate rather than a project.

The towers are not equal. Research is pre-GxP and lives in the Lab; Manufacturing carries the heaviest governance in the enterprise and lives in the Factory. The platform serves both because the substrate is the same and the governance envelope is configurable. The reader should notice, across all eight towers, how few new things appear after the first two or three use-cases: the same kits (§6), the same dual-path reasoning (§10), the same gate (§7). That repetition is the platform working.

A convention for the rest of this section: each use-case names its decision, its architectural pattern, its agent topology, its GxP frame, its key metric, and the proven pattern it inherits from — so that a reader can see the shared substrate underneath the domain surface. Each card's footer also shows its G1–G8 governance intensity.

Research · pre-GxPDevelopment · GCPProduct Dev · pre-GMP→GMPManufacturing · GMP/CSASupply · GMP/GDPCommercial · OPDPMedical · GVP/PSMFHorizontal · infra
8.1

Research Tower

Pre-GxP for target ID and generative chemistry · GLP / Part 11 begin at preclinical safety · Lab owns most of this tower; Factory owns the safety boundary
Decisions that are slow today

Which target to commit $50–100M of preclinical spend to. Which of ten thousand generated molecules to synthesize. Whether a compound's liability profile justifies advancing to IND-enabling studies. The industry's stated ambition is to halve the time from target to first clinical study while raising the probability of success — and that is precisely a decision-velocity-plus-quality problem.

Governance posture

Pre-GxP for target identification and generative chemistry (research tooling); GLP and Part 11 begin at preclinical safety, where the nonclinical overview enters the IND. The Lab owns most of this tower; the Factory owns the safety boundary.

What the platform does here — and does not

The scientists own the models: genetics, omics, structure prediction, generative chemistry, often via external partnerships. What those models do not come with is a governed evidence graph, a constraint layer, an auditable rationale, and reuse across therapeutic areas. The platform's contribution is the substrate, and the honest framing is that the platform makes the scientists' models auditable, reusable, and safe to act on — it does not replace them.

Research KG — what it must know
gene ──encodes──► protein ──participates_in──► pathway ──implicated_in──► disease
  │                  │                                                      │
  │ has_variant      │ has_structure (AlphaFold / experimental)             │ has_phenotype (HPO)
  ▼                  ▼                                                      ▼
variant ──associated_with──► phenotype ◄──observed_in── clinical_trial_outcome
                                                            ▲
compound ──binds──► target(protein) ──has_evidence──► evidence_item ──from_source──► publication | assay | dataset
  │                                       (tier: genetic / functional / clinical / literature; confidence; date)
  ├── has_structural_class ──► structural_alert / toxicophore
  ├── has_predicted_property ──► ADMET_profile
  ├── has_ip_status ──► patent_landscape
  └── has_preclinical_finding ──► finding (species, dose, endpoint, INHAND term) ──translated_to──► clinical_outcome
Aligned to: MONDO, HPO, GO, ChEBI, UniProt; core: Compound, Document
THE DOSSIER IS A SUBGRAPH WITH A SIGNATURE ON IT target Tprotein gene variant disease pathway outcome trial paper assay compound ──binds──► target every edge: source · tier · confidence · date geneticfunctionalclinicalliterature DETERMINISTIC · score tiers genetic support tier · replication counteffect-direction concordance · versioned rules AI PATH · cited synthesis target-discovery agent: novel associationswith evidence trails; validation agent:IP · competitive landscape · druggability→ liability brief with span-level citations MERGED DOSSIERfindings + brief +provenance + confidence biology leadgo / no-go+ rationale KIT · evidence-dossier → signature = audit record (G2)
A single interactive panel where the reader selects a target and watches the evidence neighborhood light up in the Research KG — genetic evidence in one color, functional in another, clinical in a third — with the deterministic score tiers rendered as a stacked bar beside the AI-synthesized brief, and the toxicologist's sign-off as the terminal node. The dossier is a subgraph with a signature on it.
UC-R1

Target Identification & Validation

Decision. Go / no-go on a target, defended by a biology lead against the question "why this one?"
Architectural pattern
Multi-omic datasets + genetics + imaging features + literature + internal assays
        
Research KG — typed edges, edge-level provenance, confidence, timestamp
        
  Deterministic path: evidence-scoring rules (genetic support tiers, replication
    count, effect-direction concordance) — versioned, auditable
  AI path: target-discovery agent traverses KG + literature for novel
    associations with evidence trails; validation agent cross-references
    IP, competitive landscape, druggability
        
Merged target dossier → biology lead reviews → go/no-go + rationale logged
Agent topology
Extraction: literature parser (PubMed, patents), multi-omic normalizer, assay structurer. Reasoning: KG traversal for target–disease links, druggability scorer, competitive IP analyzer. Orchestrator: evidence-dossier assembler with confidence scoring and provenance chain. Kit: evidence-dossier.
GxP frameNon-GxP (discovery)
Key metricTarget-to-candidate time; evidence-tier coverage per dossier
Proven patternEnterprise Knowledge Vault — KG with 10,292 triplets across 29 domains; RAG accuracy 75% → 99%+ over 14B+ data points
evidence-dossier
UC-R2

Generative Chemistry & Compound Design

Decision. Which generated candidates a medicinal chemist should spend synthesis budget on.
Architectural pattern
Generative engines (internal, partner, vendor) — pluggable via a common
  candidate-proposal contract (SMILES + conditioning + model version + seed)
        
Deterministic constraint gate: structural alerts, known toxicophores, IP
  blocklist, synthesizability threshold, property windows — every rejection
  logged with rule ID; each rule versioned + owner-signed
        
Multi-objective scoring specialists (potency, ADMET, selectivity) — each
  isolated with its own eval record
        
Ranked candidates + rationale + full lineage → med-chem review →
  selections captured as preference signal
Agent topology
Generative agent (external or internal, behind the contract). Evaluation specialists: ADMET predictor, synthesizability scorer, patent-landscape checker. Rule engine (deterministic path). Kit: rule-gate in front of a review queue. Model-agnostic by design — which matters when the enterprise has multiple discovery partnerships and more than one cloud.
GxP frameNon-GxP (pre-IND research tooling)
Key metricHit rate per synthesis cycle; cycle time
Proven patternDoscierge rule engine — 970 deterministic rules, GAMP 5 Category 5, every rejection with a rule ID
rule-gate
UC-R3

Preclinical Safety Assessment

Decision. Advance or stop, before IND-enabling studies — and survive the FDA question "why did you conclude this compound was safe to advance?"
Architectural pattern
Historical study corpus (SEND datasets, study reports, assay results,
  clinical outcomes of advanced compounds)
        
Extraction agents → structured findings with study/species/dose/endpoint provenance
        
Safety KG (compound ↔ structural class ↔ mechanism ↔ target ↔ finding ↔
  species ↔ clinical translation outcome)
        
  Deterministic path: expert-calibrated signal rules (e.g. hERG IC50 < X AND
    QT class effect in KG → flag cardiotox, severity S) — owner, calibration
    record, version
  AI path: KG-grounded LLM writes a cited "liability brief" — similar
    compounds, what happened, translational hit rate
        
Merged assessment → toxicologist reviews, signs → decision + rationale
  become the audit record
Agent topology
Extraction: SEND parser, study-report extractor, finding normalizer (MedDRA/INHAND). Reasoning: KG traversal for similar-compound outcomes, translational hit-rate calculator, liability-brief drafter. Rule engine: 500+ expert-calibrated safety rules with provisional → confirmed loop. Human gate: toxicologist sign-off under Part 11. Kits: rule-gate + draft-then-check.
GxP frameGLP / Part 11 — the nonclinical overview is IND content
Key metricP/R/F1 per liability class against compounds with known clinical outcomes; false-negative rate is the number that matters
Proven patternProtoCheck calibration — 500+ expert-calibrated rules, human-signed, patent-pending multi-agent system; Doscierge eval harness (P=92.31%, R=80.27%, F1=85.87% against expert-labeled ground truth)
rule-gatedraft-then-check
Transition. Preclinical safety is where the platform first crosses the regulatory boundary. Everything downstream lives on the other side of it.
8.2

Development Tower

GCP throughout · ICH E6(R3) and E8(R1) · Part 11 wherever an action touches a controlled document, a monitoring plan, or a supply decision · highest bar at regulatory documentation
Decisions that are slow today

Which protocol design to lock. Which sites to open. What to do about site 042, this week. Whether the CSR is ready for the reviewer. Trials designed with AI assistance finish ahead of schedule more often; the fundamental unit of work is the clinical trial, and this is the tower where the platform's home-turf patterns — ProtoCheck, the enterprise clinical-data integration, the documentation checker — apply most literally.

Governance posture

GCP throughout. ICH E6(R3) and E8(R1). Part 11 wherever an action touches a controlled document, a monitoring plan, or a supply decision. Highest bar at regulatory documentation.

Clinical KG — what it must know
study ──has_protocol──► protocol ──has_version──► protocol_version
  │                        ├── has_endpoint ──► endpoint ──precedent_in──► prior_study (internal | CT.gov)
  │                        ├── has_population ──► inclusion/exclusion_criterion ──enrollment_impact──► historical_rate
  │                        ├── has_design_element ──► design_element ──regulatory_precedent──► agency_letter | guidance
  │                        └── amended_for ──► amendment_reason (cost, delay)
  ├── conducted_at ──► site ──golden_record──► investigator (NPI/ORCID) ──has_performance──► enrollment_curve, data_quality, deviation_rate
  │                     ├── competing_trial ──► study
  │                     └── catchment ──► patient_density (RWE, filtered by I/E)
  ├── has_subject_status ──► subject_status (screened, enrolled, withdrawn)
  ├── has_supply_state ──► IRT_state (stock, expiry, resupply lead time)
  ├── has_monitoring_finding ──► KRI, deviation cluster, DQ anomaly
  └── produces_document ──► CSR | IB | CTD_module ──section──► required_content ──sourced_from──► approved_source (SDTM/ADaM, TFL)
Aligned to: CDISC PRM/SDTM/ADaM, ICH E6/E8/E9, CT.gov, FDORA diversity requirements
ONE LIFECYCLE, FOUR DECISIONS, THE SAME KITS RECURRING — hover a kit D1design Lock the protocol design-check (500+ rules)precedent-synthesiscomputational twin rule-gate ev-dossier D2sites Open the right sites entity resolution (golden record)feasibility + diversity scoringenrollment predictor + bias flag entity-res ev-dossier D3operations What to do this week supply · monitoring · recruitmentsignal specialists (isolated)orchestrator → NBA inbox signal-and-NBA D4documents Ready for the reviewer? drafting agent (span citations)checker as MCP tool (970 rules)Part 11 sign-off draft-then-check rule-gate CLINICAL KG — study · protocol · endpoint · site (golden record) · subject status · supply state · monitoring finding · document → approved source CLINICAL OPERATIONS COMMAND CENTER (§12) ← D3 signals · D2 site stall · X1 safety · D4 readiness
A trial-lifecycle strip — design → sites → operations → documentation — where each stage's decision is a node, the agents feeding it are stacked beneath, and the same kit icons recur across stages (rule-gate at design and documentation; signal-and-NBA at operations; entity-resolution at sites). Hovering a kit highlights every stage that consumes it — reuse, not four separate pictures.
UC-D1

Clinical Trial Design Intelligence

Decision. Lock the protocol — with endpoint precedent, enrollment risk, and regulatory issues surfaced before finalization. Amendments cost ~$500K+ each and delay trials by months; roughly half are avoidable.
Architectural pattern
Corpus: internal protocols + outcomes, CT.gov, publications, ICH E6/E8/E9,
  FDA/EMA letters, RWE
        
Trial-design ontology (indication ↔ endpoint ↔ population ↔ design element ↔
  regulatory precedent ↔ operational outcome) — aligned to CDISC PRM
        
  Design-check agent (ProtoCheck pattern): draft protocol → rule findings
    (compliance, precedent conflict, enrollment-risk criteria) with severity + citation
  Precedent-synthesis agent: "in this indication, 12 trials used endpoint X;
    4 amended for reason Y" — cited
  Twin/simulation agent: enrollment + dropout + site-burden model from
    historical SDTM/ops data; design alternatives compared
        
Design review workspace: findings, precedents, simulations side by side;
  every accepted/rejected recommendation logged
Agent topology
Design-check (500+ rules, 6+ TAs, FDA/EMA/ICH/WHO; human-in-the-loop mandatory per E6(R3)); precedent-synthesis (cross-trial analysis, cited); computational twin (enrollment simulation from SDTM/ADaM, screen-fail prediction, alternative comparison). Kits: rule-gate + evidence-dossier. Where a sponsor already has a protocol-intelligence tool, the design-check agent is the complement: the historical tool provides the intelligence; the rule engine provides the compliance gate and makes the recommendations auditable.
GxP frameGCP / E6(R3) / E8(R1) — the protocol is a controlled document
Key metricAmendments avoided; protocol cycle time
Proven patternProtoCheck — patent-pending, 500+ rules, 6+ therapeutic areas, four specialist reviewers, presented at BioNJ 2026
rule-gateevidence-dossier
UC-D2

Site Selection & Recruitment Intelligence

Decision. Which sites to open, ranked, with diversity-plan coverage as a first-class criterion — and which of the chosen sites will under-enroll, early enough to intervene.
Architectural pattern
Site/investigator master — entity-resolved across CTMS, EDC, CRO feeds, CT.gov,
  NPI/ORCID; golden record per site + per investigator (the MDM pattern)
        
Site KG (site ↔ investigator ↔ indication experience ↔ enrollment curves ↔
  data quality ↔ competing trials ↔ catchment patient density filtered by I/E)
        
  Deterministic scoring: regulatory readiness, equipment, performance
    thresholds, diversity coverage (FDORA) as a first-class criterion
  Predictive agent: enrollment forecast per site for THIS protocol's I/E;
    cited rationale per recommendation; bias flag on over-concentration in usual sites
        
Ranked site list + rationale + diversity coverage → feasibility team reviews →
  chosen sites feed operational monitoring (UC-D3)
Agent topology
Entity-resolution engine (golden record, crosswalk, 100K+ entity scale proven). Scoring specialists: feasibility rules, enrollment prediction, diversity calculator. Recommendation agent with bias detection and what-if analysis. Kits: entity-resolution + evidence-dossier. The under-appreciated insight: investigators are HCPs and sites are HCOs — the same golden-record discipline commercial has applied for years is what makes site intelligence durable across studies.
GxP frameGCP / FDORA diversity action plans
Key metricEnrollment velocity (time to last patient enrolled); screen-fail reduction; site-stall early-warning lead time
Proven patternFortune 50 pharma Customer MDM — 100K+ HCP/HCO, 20+ system integration; clinical data integration across 100+ studies (SDTM/ADaM)
entity-resolutionevidence-dossier
UC-D3

Operational Precision in Clinical Trials

Decision. What the study team should do this week about supply, monitoring, and recruitment — as next-best-actions with rationale, expected impact, and an owner. The failure mode is alert fatigue and unactionable recommendations. This is also the primary feed for the Clinical Operations Command Center (§12).
Architectural pattern
Event sources: IRT, EDC, CTMS, RBM/central monitoring, enrollment feeds
        
Trial-ops semantic layer (study ↔ site ↔ subject-status ↔ supply-state ↔
  monitoring-finding ↔ plan) — shared across ALL studies, not per study
        
Signal specialists (isolated, independently evaluated):
  Supply: stock-out / expiry / resupply-lead-time risk
  Monitoring: KRI drift, deviation clusters, data-quality anomalies
  Recruitment: enrollment-vs-plan, screen-fail trends, site stall
        
Orchestrator: de-duplicate, prioritize, compose NBA with rationale + expected
  impact + who-should-act; suppress noise from study-team feedback
        
Study team inbox: accept / modify / reject → logged (Part 11 where the action
  touches GCP) → feedback retrains prioritization
Agent topology
Three monitoring specialists + orchestrator on stateful orchestration with per-execution history. Kit: signal-and-NBA. Any recommendation can be reconstructed months later during an inspection — that is what stateful orchestration with execution history buys.
GxP frameGCP / Part 11 (selective — where actions touch monitoring plans or supply decisions)
Key metricSignal-to-action time (monthly → real-time); NBA acceptance rate; alert-fatigue index
Proven patternFortune 50 pharma real-time platform on 14B+ data points; 4-agent platform for 150+ regulated users; Doscierge 11-state Step Functions pipeline with 25 functions and full execution history
signal-and-NBA
UC-D4

Clinical & Regulatory Documentation

Decision. Is this CSR section / protocol section / IB / CTD module summary ready for the reviewer — with every sentence traceable to an approved source? Highest volume, most immediately measurable, and the use-case regulators are watching most closely.
The critical gap across the industry. Every major company now has GenAI tools that generate drafts at 80–85% acceleration. Almost nobody has a system that deterministically checks those drafts against the regulatory rulebook before a reviewer sees them. Generate-then-check is the architecture that makes this safe for production — and the checker is the harder, higher-value half.
Architectural pattern
Approved-source registry (SDTM/ADaM, locked TFLs, approved protocol, IB,
  prior submissions, style guide, CTD templates) — versioned, access-controlled;
  the ONLY citable source
        
Document ontology (CTD section ↔ required content ↔ source data types ↔
  regulatory requirement ↔ house-style rule)
        
Drafting agent: section-by-section; each sentence tagged with source-registry
  citations (span-level provenance); refuses unsupported claims
        
Checker path (Doscierge pattern): deterministic rules (structure, required
  elements, numeric consistency vs. source tables, cross-reference integrity)
  + KG-grounded LLM review (scientific, regulatory alignment)
        
Medical writer / regulatory reviewer: draft + findings + per-sentence source;
  edits captured as preference signal; sign-off under Part 11
Agent topology
Drafting agent. Checker (970 deterministic rules; structure, required-element, numeric-consistency; KG-grounded review) — exposed as a well-typed MCP tool, not an agent. Eval harness: hallucination rate per document type, numeric-consistency rate, reviewer edit distance, cycle time vs. baseline. Kit: draft-then-check.
Key architectural insight
The checker is a tool the agent calls, not an agent itself. Any agent workflow — any model vendor's, any internal orchestrator's — can call the compliance API as one step. For Category 5 validation, a model-driven tool-selection loop is a liability, not a feature; deterministic control flow is the deliberate ceiling.
GxP framePart 11 / Annex 11 (highest) — e-signature on sign-off; immutable audit trail on every edit. GAMP 5 Category 5 for the checker; CSA-aligned controls for the drafting LLM
Key metricDocument cycle time; reviewer edit distance; hallucination rate (target: zero unsupported claims reaching a reviewer)
Proven patternDoscierge — 970 rules, 10,292 KG triplets, checker + eval harness; Fortune 50 pharma GenAI function scaled to 150+ regulated users
draft-then-check
Transition. Development hands a validated process and a dossier to the next tower. Product Development is where the molecule becomes a manufacturable product — and where the knowledge that manufacturing will need is first written down.
8.3

Product Development Tower (CMC & Tech Transfer)

Largely pre-GMP but produces CTD Module 3 content and defines GMP-critical parameters · tech transfer is the GMP boundary · process validation Stage 1 lives here
Decisions that are slow today

Which formulation and process to lock. What the control strategy is — which parameters are critical, what the design space is. Whether the receiving site is ready to accept a transferred process. What to file as an established condition under ICH Q12 so post-approval changes do not require a supplement. This is the tower that writes the knowledge Manufacturing later uses, and the platform's leverage is making that knowledge structured from the start rather than reconstructed from PDFs a decade later.

Governance posture

CMC development is largely pre-GMP but produces CTD Module 3 content and defines GMP-critical parameters; tech transfer is the GMP boundary; process validation Stage 1 (Process Design) lives here.

CMC KG — what it must know
product ──has_formulation──► formulation ──has_component──► material (API, excipient) ──has_attribute──► material_attribute
  │                                                                                      └── from_supplier ──► supplier (→ Supply Chain KG)
  ├── has_process ──► process ──has_unit_operation──► unit_operation (ISA-88 recipe) ──has_parameter──► parameter
  │                                                        │                               ├── is_critical (CPP) ──justified_by──► DoE_study
  │                                                        │                               └── has_validated_range / design_space
  │                                                        └── has_equipment_class ──► equipment (→ Manufacturing KG)
  ├── has_quality_attribute ──► CQA ──measured_by──► analytical_method ──validated_by──► method_validation (ICH Q2/Q14)
  │                              └── has_specification ──► specification (version, effective date)
  ├── has_control_strategy ──► control_strategy ──links──► CPP ↔ CQA (with model / evidence)
  ├── has_stability_study ──► stability_study ──result──► stability_data (ICH Q1)
  ├── has_established_condition ──► established_condition (ICH Q12) ──change_category──► reporting_category
  └── transferred_to ──► site ──gap──► transfer_gap (equipment, method, material, training)
Aligned to: ICH Q8–Q14, USP, CTD Module 3, ISA-88; core: Product, Material, Specification, Batch
PRODUCT DEVELOPMENT IS WHERE THE MANUFACTURING KNOWLEDGE GRAPH IS BORN CMC / PRODUCT DEVELOPMENT KG MANUFACTURING & QUALITY KG formulation · DoE_study · design_spaceunit_operation (ISA-88 recipe)analytical_method · method_validationstability_study · stability_datacontrol_strategy (CPP ↔ CQA evidence)transfer_gap (equipment · method · material · training)written once, structured, with study provenance batch · unit_operation_phase · historian_tagsample · test · test_result (by effective date)genealogy · material_lot · CoAdeviation · CAPA · changeapqr · cpv_scope · spc_signal · pat_modelmined: stalls_at (phase, p50, confidence)tested against reality, batch after batch CPP · validated range CQA · analytical method specification (version, date) established condition (Q12) the same nodes — P1 locks them; M2 monitors them on day one without re-keying; P3 classifies changes against them equipment class → site equipment (P2 gap analysis) · supplier → Supply Chain KG · change (QMS) → established condition → reporting category
A "knowledge handoff" diagram: the CMC KG on the left, the Manufacturing KG on the right, and the CPP/CQA/specification/established-condition nodes drawn once, spanning both — because they are the same nodes. Product Development is where the manufacturing knowledge graph is born.
UC-P1

Process Knowledge & Control-Strategy Intelligence

Decision. Which parameters are critical, with what ranges, and why — defensible in the dossier and usable by CPV on day one of commercial manufacture.
Architectural pattern
DoE studies, characterization reports, scale-up data, analytical method files
        
Extraction agents → structured CPP/CQA candidates with study provenance
        
CMC KG: parameter ↔ attribute relationships with evidence strength; design space
        
  Deterministic path: criticality-assessment rules (ICH Q9 risk ranking,
    Q8 design-space criteria) — versioned, SME-signed
  AI path: control-strategy synthesis agent drafts the justification with citations
        
Process-development lead reviews → control strategy locked → CPP/CQA list
  becomes the CPV monitoring scope (UC-M2) without re-keying
Kitsrule-gate + draft-then-check
GxP framePre-GMP; CTD Module 3 content
Key metricTime from characterization complete to control strategy locked; proportion of CPPs with structured evidence links
Proven patternDoscierge CMC domains — KG covers stability, impurities (ICH Q3A thresholds), DMF, dissolution
rule-gatedraft-then-check
UC-P2

Tech Transfer & Site Readiness

Decision. Is the receiving site — internal or CMO — ready to run this process, and what are the gaps?
Architectural pattern
Sending-site process knowledge package (from CMC KG) + receiving-site capability
  (equipment, methods, materials, training records — from Manufacturing KG)
        
Gap-analysis agent: traverses process ↔ equipment class ↔ site equipment;
  method ↔ site method validation status; material ↔ approved supplier at site
        
Deterministic completeness rules (transfer protocol required elements per
  WHO TRS 961 Annex 7 / ICH Q10) + AI-drafted transfer protocol sections with citations
        
Transfer lead + receiving-site QA review → gaps become tracked items; protocol
  signed under Part 11 where GMP applies
Kitsevidence-dossier + draft-then-check
GxP frameGMP at the receiving site; Part 11 on the transfer protocol
Key metricTransfer cycle time; gaps discovered before vs. after first engineering batch
Proven patternEntity resolution across process ↔ equipment ↔ site; the same crosswalk discipline as MDM
evidence-dossierdraft-then-check
UC-P3

Post-Approval Change & Established Conditions (ICH Q12)

Decision. Does this proposed change touch an established condition, and what is the reporting category — before the change-control board meets?
Architectural pattern
Proposed change (from QMS) → change-classification agent traverses CMC KG:
  affected parameter → is it an established condition? → which reporting category?
  → which markets? → prior precedent for similar changes?
        
Deterministic path: reporting-category rules per market (versioned, RA-signed)
AI path: cited impact assessment draft
        
Regulatory-affairs reviewer disposition → decision logged; the KG edge is updated
Kitsrule-gate + evidence-dossier
GxP frameGMP change control; regulatory commitment tracking
Key metricChange-classification cycle time; classification disagreements with RA (target: falling)
rule-gateevidence-dossier
Transition. The control strategy, the specifications, and the genealogy of materials now exist as structured knowledge. Manufacturing is where they are tested against reality, batch after batch.
8.4

Manufacturing Tower

GMP / Part 11 / CSA — the highest bar · deterministic paths GAMP 5 Cat 5 · statistical and chemometric models under an analytical-method lifecycle with QA-approved model registry · LLM paths controlled by grounding, checking, and human sign-off
Decisions that are slow today

Is the process in a state of control — this month, not fifteen months from now? Is this batch's blend uniform — now, not after QC returns a thief-sample result? What caused this deviation, and did the last CAPA actually work? Can this batch be released? Manufacturing is the most heavily governed tower in the enterprise — GMP, Annex 11, ALCOA+, FDA's Computer Software Assurance guidance, draft EU GMP Annex 22 (consultation July–October 2025; final targeted Q4 2026) — and it is also the tower where decision latency is most directly a compliance exposure: a CPV program that generates a signal nobody acted on, followed by a batch failure, is a warning letter.

Governance posture

GMP / Part 11 / CSA — the highest bar. Deterministic paths validated as GAMP 5 Category 5; statistical and chemometric models managed under an analytical-method lifecycle with QA-approved model registry; LLM paths controlled by grounding, checking, and human sign-off.

Why this tower is its own tower, with four sub-domains

Manufacturing is not one use-case called "manufacturing." It is at least four distinct decision families that share one data foundation and one KG but differ in cadence, regulation, and user: the Annual Product Quality Review (retrospective, regulatory, QA-owned), Continued Process Verification (continuous, statistical, MS&T-owned), the PAT platform (real-time, spectral, model-lifecycle-owned), and Blend Endpoint Detection (the PAT platform's entry-point use-case, operator-facing). Deviation/CAPA investigation and batch-record review sit alongside them. Treating these as separate towers would lose the shared substrate; treating them as one use-case would lose the fact that they have different owners and different validation paths.

The data foundation the whole tower stands on

Three data products, joined: batch_phase_summary (CPPs from the historian, batch-aligned via MES event frames), the LIMS conformed services (samples → tests → test_results, given meaning by specifications; CQAs), and genealogy + raw_material_results (lot lineage; CoA/CoC from suppliers and CMOs). With those three, the CPP-to-CQA relationship is a query. Without them, every use-case below is a spreadsheet campaign. The KG holds the meaning persistently; the batch-level facts are materialized at run-time from the data products (§5.4).

Manufacturing & Quality KG — what it must know
product ──manufactured_as──► batch ──at_site──► site ──has_equipment──► equipment ──has_tag──► historian_tag (PI/OSIsoft, DeltaV)
  │                            │                                             └── has_pat_instrument ──► instrument (NIR, Raman, FTIR, FBRM)
  │                            ├── has_phase ──► unit_operation_phase (ISA-88; start/end from MES event frames)
  │                            │                   └── has_parameter_summary ──► parameter (CPP; mean/min/max/duration-out-of-range)
  │                            │                                                  └── validated_range ──from──► control_strategy (→ CMC KG)
  │                            ├── has_sample ──► sample (sample_point) ──has_test──► test ──has_result──► test_result (status: final/retest/superseded)
  │                            │                                                                              └── against ──► specification (version, effective date)
  │                            ├── genealogy ──► batch | material_lot (DS lot → DP lots; raw-material lot → batch)
  │                            │                   └── has_coa ──► raw_material_result (CoA / CoC from supplier or CMO)
  │                            ├── has_deviation ──► deviation ──root_cause──► cause ──addressed_by──► CAPA ──effectiveness──► outcome
  │                            ├── has_change ──► change (→ ICH Q12 established condition in CMC KG)
  │                            ├── has_spectral_record ──► spectrum (instrument, timestamp) ──scored_by──► pat_model (version, QA-approved)
  │                            └── mined_edge: stalls_at (phase, p50 duration, confidence) ← process intelligence
  ├── has_apqr ──► apqr (period, site, status) ──covers──► batch[] ──finding──► trend | OOS | OOT
  └── has_cpv_program ──► cpv_scope ──monitors──► parameter[] ──chart──► spc_signal (rule, batch, disposition)
Aligned to: ISA-88/95, ICH Q7/Q9/Q10, 21 CFR 211, EU GMP Ch. 1.10 / Annex 15; core: Product, Batch, Material, Specification, Site
ONE FOUNDATION · FIVE CADENCES · ONE GOVERNED SURFACE batch_phase_summaryhistorian-derived CPPs · MES event framesPI/OSIsoft per site → data fabric samples → tests → test_results ← specificationsLIMS conformed services · CQAsregional LIMS → canonical attribute dictionary genealogy + raw_material_resultslot lineage · CoA/CoC from suppliers & CMOsDS lot → DP lots; raw-material lot → batch MANUFACTURING & QUALITY KG — meaning persistent, facts materialized per batchbatch → phase → CPP → result → spec (by effective date) → genealogy → lot → CoA · CPP-to-CQA is a query M1 · APQRmonthly roll-up (not annual)QA-owned · 211.180(e) M2 · CPVnightly · every CPPMS&T-owned · SPC/MSPC M3 · PAT platformreal-time · spectramodel-lifecycle-owned M4 · Blend endpointper revolutionoperator-facing advisory M5 · Deviation/CAPAon eventQA signs · Part 11 MANUFACTURING QUALITY COMMAND CENTER (§12) — one NBA: "Line 2 EWMA shift correlates with resin lot L-7781; CoA method differs; DEV-2026-0412 on same lot; recommend hold"real-time PAT/CPV · daily quality huddle · monthly management review (ICH Q10) · every signal carries provenance you can walk + mined edge stalls_at (phase, p50, confidence) from process intelligence · + S1/S2 supply feeds (CoA exceptions, cold chain)
One diagram, three tiers. Bottom: the three data products (historian-derived CPPs, LIMS CQAs, genealogy/CoA) as three source lanes converging into the Manufacturing KG. Middle: the five use-cases as consumers of the same graph, with their cadences labeled (APQR: monthly roll-up; CPV: nightly; PAT: real-time; blend endpoint: per revolution; deviation: on event). Top: the Manufacturing Quality Command Center (§12) receiving all five feeds. One foundation, five cadences, one governed surface.
UC-M1

Annual Product Quality Review (APQR)

Decision. Is the process consistent, are the specifications still appropriate, and what needs to change — per product, per site, per year, with cross-site comparison. Mandated by 21 CFR 211.180(e) and EU GMP Chapter 1.10; a deficient APQR is one of the most common observations against otherwise well-run sites.
4–8 wksToday. 4–8 weeks of calendar time per product; 2.5–5 FTE of skilled labor for a five-product, three-site company; findings delivered 6–15 months after the batches were made; 12,000 batch-parameter values a year hand-transcribed at a 1–5 per 1,000 error rate; cross-site comparison effectively impossible.
Architectural pattern
batch_phase_summary + LIMS conformed services + genealogy + QMS extracts
  (deviations, changes, complaints) — validated data products
        
APQR orchestrator (scheduled — monthly, not annually): per product-site,
  assembles the review period's batch set; checks event_completeness_score
        
  Deterministic path: validated SPC/capability queries (Cpk/Ppk vs. validated
    range and spec), OOS/OOT listings, batch-count reconciliation vs. MES —
    GAMP 5 Cat 5 query layer
  AI path: trend-narrative agent drafts the interpretation section with every
    chart and number cited to its data-product row; deviation/CAPA/change
    cross-reference with provenance
        
Checker (draft-then-check): required elements per 211.180(e) / Ch. 1.10;
  numeric consistency between narrative and charts
        
QA reviewer sees data package + draft + findings → signs under Part 11 →
  the APQR becomes a periodic summary of an always-on system
Agent topology
Data-completeness monitor; SPC/capability calculator (deterministic); narrative drafter; checker (MCP tool); orchestrator. Kits: draft-then-check on a signal-and-NBA backbone. Parallel-run against the manual APQR for a full cycle before the manual process is retired — agreement within pre-defined tolerance on every CPP, every discrepancy explained, QA signs the PQ report.
GxP frameGMP / Part 11 / ALCOA+; the query layer is validated software
Key metricSignal latency (6–15 months → weeks); APQR data-gathering cycle time (weeks → days); reviewer time spent verifying vs. interpreting (60–70% verification today → the inverse)
Proven patternDoscierge checker pattern; ALCOA+ evidence chain (reconciliation records, immutable storage, access logs) so any chart is reproducible from raw historian data on demand during inspection
draft-then-checksignal-and-NBA
UC-M2

Continued Process Verification (CPV)

Decision. Has the validated process drifted — and which scientist needs to look, today. FDA Process Validation Stage 3; EU Annex 15 Ongoing Process Verification; ICH Q10 §3.2.1. A program, not an event, for the entire commercial life of the product.
4–12 wksToday. Quarterly or monthly Excel/Minitab exercise; 4–12 week lag from batch completion to chart; 5–10 of 30–50 relevant CPPs actually trended; each parameter monitored independently, so the pH-temperature interaction that shifts the glycosylation profile is invisible; every site a silo.
Architectural pattern
Batch-end event (MES) → nightly: phase-level statistics for EVERY CPP in the
  control strategy (marginal cost of one more parameter ≈ zero)
        
Monitoring specialists (isolated, each with its own eval record):
  SPC agent: Shewhart / CUSUM / EWMA; Western Electric / Nelson rule violations
  Capability agent: Cpk (within-batch) and Ppk (overall) vs. validated range;
    Ppk of each CQA vs. specification (from LIMS conformed services)
  MSPC agent (Phase 3): per-batch PCA on phase-level CPP vector; Hotelling T²
    and Q-residual vs. normal-operating-condition model; contribution plot on excursion
  Lot-effect agent: genealogy + raw_material_results — does a resin or media
    lot explain the shift?
        
Orchestrator: de-duplicate signals across agents; prioritize by risk (ICH Q9);
  compose NBA with rationale, the chart, the affected batches, and the
  responsible scientist; suppress known-benign patterns from feedback
        
Process scientist disposition (confirmed / benign / investigate) → logged;
  investigation opens in QMS; disposition feeds the APQR (UC-M1) and the
  Manufacturing Quality Command Center (§12)
Agent topology
Four monitoring specialists + orchestrator. Kit: signal-and-NBA. Cross-site CPV is GROUP BY site because sites share the canonical data model. The MSPC model is a statistical digital shadow of the process — a model of normal behavior against which reality is continuously compared.
GxP frameGMP; statistical engine validated (Cat 5); MSPC model under analytical-method lifecycle (NOC set, cross-validation, held-out detection test, QA approval in the model registry)
Key metricBatch-to-chart latency (4–12 weeks → overnight); CPP coverage (5–10 → all); signal-to-disposition time; deviations pre-empted by a CPV signal
Proven patternFortune 50 pharma real-time platform (14B+ data points); Doscierge stateful pipeline; the same signal-and-NBA kit as UC-D3
signal-and-NBA
UC-M3

Enterprise PAT Data Platform

Decision. Which batches are abnormal while they are running — and whether the chemometric model that says so is the approved version. NIR, Raman, FTIR, FBRM instruments across sites generate high-value data about what the material is at every moment; almost none of it leaves the vendor PC it was captured on, none is joined to the historian, and model provenance leaves with the scientist who built it.
Architectural pattern
Instrument PCs → scalar outputs (predictions, statistics, instrument health)
  flow through the historian-to-cloud channel like any tag; spectral vectors
  flow through a parallel high-throughput path into the same lake
        
Model registry: chemometric models (PLS, PCA, MSPC) as governed, versioned,
  QA-approved assets — not files on a laptop; every prediction carries model
  version + calibration set reference
        
Monitoring specialists: golden-batch MSPC (T²/Q per batch, in real time);
  instrument-health drift; model-performance drift (prediction vs. reference
  method)
        
Process-model agents (mechanistic, data-driven, hybrid) deployed to the floor
  with a confidence band; advisory first, control later, RTRT only if chosen
        
MS&T / QA console: model lifecycle, cross-site model transfer, method
  harmonisation; every batch's spectra retained and retrievable for a regulator
Agent topology
Ingestion monitors; MSPC scorer; model-drift monitor; model-lifecycle orchestrator (registry approvals as a workflow — the same pattern as §6). Kit: signal-and-NBA with a model registry extension. This platform is the prerequisite for continuous manufacturing (ICH Q13) and for any credible manufacturing digital twin.
GxP frameGMP; model lifecycle per analytical-method validation; GAMP 5 for the platform; Annex 11 for records
Key metricTime to first value (blend endpoint advisory, 6–9 months); time to multivariate monitoring on all site instruments (12–15 months); model provenance completeness (target: 100% of predictions traceable to an approved model version)
Proven patternRegistry-driven approval workflow (§6); trust-tiered provenance on every scored record
signal-and-NBAmodel registry
UC-M4

Blend Uniformity Endpoint Detection

Decision. Stop blending when the material is ready, not when the clock says so. An NIR probe on the blender reads the blend once per revolution; a validated chemometric model converts each spectrum to predicted API content; when scan-to-scan variation falls below threshold and stays there, the blend is uniform.
Why it is the PAT entry point. Lowest regulatory risk (the fixed-time step stays validated; the endpoint runs as an advisory signal first; the FDA PAT Guidance uses blend uniformity as its canonical example); fastest ROI (cycle-time savings begin the day the advisory goes live); most proven science (20+ years of precedent; PLS + moving-block standard deviation is the industry approach); highest platform reuse.
Architectural pattern
NIR spectrum per revolution → PAT platform (UC-M3) → PLS model (registry-approved
  version) → predicted API content → moving-block SD
        
Endpoint agent (deterministic): SD < threshold for N consecutive blocks →
  "blend uniform" advisory with confidence; every advisory logged with model
  version, spectra references, batch, timestamp
        
Operator sees advisory beside the validated fixed-time step (Phase 1–2);
  operator action logged. Phase 3: endpoint replaces fixed time under a
  validated change. Phase 4 (optional): RTRT — content-uniformity QC test eliminated.
        
Batch's endpoint evidence flows to CPV (UC-M2) and APQR (UC-M1) as a CPP
Agent topology
One deterministic endpoint specialist on the PAT platform; the human is the operator. Kit: signal-and-NBA (degenerate case: one signal, one action). The simplest agent in this article, and one of the most valuable.
GxP frameGMP; advisory mode requires no submission; RTRT requires one
Key metricBlend time saved per batch (15–45 min); manufacturing hours freed (50–150 h/yr per high-volume product); value of freed capacity ($250K–$2.25M/yr at $5K–$15K per batch-hour); payback 6–18 months
Proven patternConfidence-gated advisory → autonomous graduation (§18's coach-mode model, applied to a process step)
signal-and-NBA
UC-M5

Deviation & CAPA Investigation

Decision. What caused this deviation, has it happened before, did the last CAPA work — and can the batch be released?
Architectural pattern
Deviation opened in QMS → investigation agent retrieves similar historical
  deviations (KG traversal: product ↔ process ↔ equipment ↔ deviation ↔ root cause)
  with provenance; pulls the batch's CPV signals (UC-M2), PAT records (UC-M3),
  genealogy and CoA (UC-S1), and mined process-intelligence edges ("this phase
  stalls 3× more often on line 2")
        
  Deterministic path: batch-record completeness, spec-conformance, trending
    (ICH Q3A thresholds, stability limits); CAPA-effectiveness rules
  AI path: root-cause hypothesis draft with cited evidence; CAPA precedent summary
        
Checker: investigation-report required elements; numeric consistency
        
Investigator accepts / rejects hypotheses → QA signs under Part 11 →
  root cause and effectiveness become KG edges for the next investigation
Agent topology
Investigation (reasoning), batch-review (deterministic), narrative drafter, checker. Kits: evidence-dossier + draft-then-check.
GxP frameGMP / Part 11 / CSA (highest)
Key metricInvestigation cycle time (days saved per deviation); CAPA effectiveness rate; batch-hold hours
Proven patternDoscierge CMC domains (stability, impurities, DMF, dissolution) — GAMP 5 Category 5
evidence-dossierdraft-then-check
Transition. Every batch depends on materials that came from somewhere and will go somewhere. Supply Chain is where genealogy meets the world outside the site.
8.5

Supply Chain Tower

GMP for material release and supplier qualification · GDP for cold chain · DSCSA / serialization for traceability · Part 11 on release decisions
Decisions that are slow today

Is this incoming material lot acceptable — with the CMO's Certificate of Analysis reconciled against our specification, not just filed? Which shipment lane is at risk of a temperature excursion? Which supplier's qualification is about to lapse? Where will demand exceed supply next quarter, and which sites should absorb it? Supply chain decisions are latency-critical because a late decision becomes a shortage, and a shortage in pharma is a patient outcome.

Governance posture

GMP for material release and supplier qualification; GDP (Good Distribution Practice) for cold chain; DSCSA / serialization for traceability; Part 11 on release decisions.

Supply Chain KG — what it must know
material ──supplied_by──► supplier ──qualified_for──► material (status, audit date, expiry) ──at_site──► site
  │                          └── is_cmo ──► cmo ──produces──► batch (external) ──has_coa/coc──► raw_material_result (vs. specification)
  ├── has_lot ──► material_lot ──genealogy──► batch (→ Manufacturing KG)
  │                 └── shipped_via ──► shipment ──on_lane──► lane ──has_excursion──► temperature_event (GDP)
  ├── has_demand_signal ──► forecast (market, period) ──vs──► supply_plan ──at_risk──► shortage_risk
  ├── serialized_as ──► serialization_event (DSCSA: commission, ship, receive, verify)
  └── geopolitical_exposure ──► region_risk (single-source, sanctions, logistics)
Aligned to: GS1, DSCSA, GDP, ICH Q7 (supplier), ICH Q9; core: Material, Batch, Organization, Site
ONE TRACED LINE ACROSS THE SITE BOUNDARY SUPPLY CHAIN KG ⟷ MANUFACTURING KG — one continuous graph: material_lot ──genealogy──► batch ──has_coa──► raw_material_result ──against──► specificationcore: Material · Batch · Organization · Site — entity-resolved across ERP, LIMS, CMO feeds and serialization events site boundary supplier / CMOqualification status CoA / CoCPDF · spreadsheet · EDI shipment · lanetemperature telemetry receipt · releasePart 11 disposition genealogy → batchCPV lot-effect (M2) serializationcommission · ship · verify distributionDSCSA chain S1 · CoA intakereconcile vs spec by effective date S2 · supply riskexcursion · expiry · single-source S3 · traceabilityverification failures · chain gaps every sensor is a monitoring specialist on the signal-and-NBA kit; every disposition lands in genealogy and in the Manufacturing Quality Command Center
A material's journey as a single traced line — supplier → CoA → receipt → genealogy → batch → serialization → distribution lane — with the three agents that watch it drawn as sensors on the line, and the Manufacturing KG and Supply Chain KG shown as one continuous graph across the site boundary.
UC-S1

Incoming Material & CMO Quality Intake (CoA/CoC Reconciliation)

Decision. Accept this lot. Today, CMO results arrive as spreadsheets and PDFs; process engineers report spending the majority of their CPV time on extraction and aggregation, and CMO file handling is a disproportionate share of it.
Architectural pattern
CoA / CoC (PDF, spreadsheet, EDI) → extraction agent → structured raw_material_result
  with document provenance → standardized file-ingestion service with
  review-and-approve workflow → lands in the same conformed table as internal results
        
Deterministic path: reconciliation vs. specification (by effective date), test
  completeness, method equivalence flags; supplier-qualification status check
AI path: exception summary with citations when a value is missing, out of trend,
  or method differs
        
QC/QA disposition → release decision under Part 11 → lot appears in genealogy
  and in CPV lot-effect analysis (UC-M2) automatically
Kitsdraft-then-check (the checker is the reconciliation rule set)
GxP frameGMP / Part 11
Key metricLot-release cycle time; CMO batches appearing in CPV charts (target: 100%); transcription errors (target: zero)
Proven patternDoscierge extraction pipeline with KG-first grounding; ProtoCheck rule calibration
draft-then-check
UC-S2

Supply Risk & Cold-Chain Monitoring

Decision. Which shipments, suppliers, and lanes need action this week.
Architectural pattern
Signal specialists (temperature-excursion monitor on lane telemetry; supplier-qualification-expiry monitor; single-source and geopolitical exposure monitor; demand-vs-supply gap monitor) → orchestrator composing NBAs with rationale, impact, and owner → supply-risk board disposition → Supply Chain feed into the Manufacturing Quality Command Center (§12).
Kitsignal-and-NBA — the same kit as UC-D3 and UC-M2
GxP frameGDP; GMP for supplier status
Key metricExcursion-to-decision time; shortage events pre-empted; supplier audits scheduled before expiry (target: 100%)
signal-and-NBA
UC-S3

Serialization & Traceability Intelligence

Decision. Where is this saleable unit, is its provenance intact, and is a suspect-product investigation warranted?
Architectural pattern
Serialization events (DSCSA) as KG edges on the material lot; anomaly monitor for verification failures and chain gaps; investigation agent assembles the chain-of-custody dossier with provenance.
Kitevidence-dossier
GxP frameDSCSA / GDP
Key metricVerification-failure investigation time
evidence-dossier
Transition. The product leaves the site. From here, the decisions are about people — the HCPs who prescribe, the patients who take, and the medical and safety professionals who answer for both.
8.6

Commercial Tower

Promotional compliance (FDA OPDP, EFPIA, local codes) · MLR-approved content only · medical-vs-commercial firewall enforced by architecture, not policy · AE detection with mandatory reporting on every conversation
Decisions that are slow today

What should this rep say to this HCP on Thursday, using only MLR-approved content? Is this piece of content approved for this market, this indication, this audience? Which signals in the market — competitor launches, formulary changes, prescribing shifts — need a response? Commercial agentic AI is increasingly delivered on a CRM vendor's agent platform under a multi-year rollout; the platform's job is to make that estate governable and to put the semantic layer underneath it.

Governance posture

Promotional compliance (FDA OPDP, EFPIA and local codes); MLR-approved content only; medical-vs-commercial firewall enforced by architecture, not policy; adverse-event detection with mandatory reporting on every conversation.

Commercial KG — what it must know
hcp ──golden_record──► identifiers (NPI, CRM id, MDM id, ...) ──affiliated_with──► hco ──in_territory──► territory
  │                                                                                          └── has_formulary ──► formulary_status (product, payer)
  ├── has_specialty / has_prescribing_pattern (de-identified, market-level)
  ├── has_interaction ──► interaction (channel, date, content_used, outcome, consent_status)
  ├── has_consent ──► consent (channel, market, expiry)
  └── has_medical_inquiry ──► inquiry (→ Medical KG; firewalled)
product ──has_indication──► indication ──has_approved_asset──► mlr_asset (market, audience, version, expiry, claims[])
  │                                                                └── claim ──supported_by──► approved_source (label, CSR section)
  └── has_market_signal ──► signal (competitor, formulary, launch, pricing) ──from_source──► feed
Aligned to: NPI, MDM golden record, MLR taxonomy, local promotional codes; core: HCP, Product, Organization, Document
THE FIREWALL IS ARCHITECTURE, NOT POLICY separate identity principal · separate registry scope · separate KG partition COMMERCIAL ESTATE — CRM vendor platform, inside the contract C1 · field-force agentpre-call plan · NBA from MLR content ONLYeach claim cited to asset + source C2 · MLR review · C3 · market signalsclaim-to-source rules · fair balancecompetitor · formulary · launch feeds MLR ASSET REGISTRY — the ONLY content source heremlr_asset (market · audience · version · expiry) → claim → approved_source (label, CSR section) compliance gate: OPDP / local code rules · MLR expiry · consent check MEDICAL & SAFETY ESTATE X2 · medical-inquiry agentmedical-affairs approved sources onlyMSL triage where judgment is required X1 · pharmacovigilanceintake → coding → narrative → checker15-day clock enforced MEDICAL SOURCE REGISTRY — same mechanism, different contentsapproved_response ← medical_source · label by version and market off-label boundary rules per market · no promotional content · consent AE-DETECTION SIDECAR — mandatory, non-bypassable hook → PV, clock started SHARED ONLY: HCP / HCO GOLDEN RECORD (NPI · CRM id · MDM id) — entity-resolution kit
Two agent estates drawn on either side of a hard line — commercial on one side, medical on the other — sharing only the HCP golden record at the bottom and the AE sidecar routing across the top to PV. The MLR asset registry is the only content source on the commercial side.
UC-C1

HCP & Patient Engagement (Field Force and Self-Service)

Decision. Next-best-action for this HCP, from approved content, with provenance — and nothing else.
Architectural pattern
CRM vendor's agent platform as the experience layer
        
Customer & content semantic layer (HCP/HCO golden record ↔ product ↔
  indication ↔ MLR-approved asset registry with claim-level provenance)
        
Field-force agent: pre-call plan + NBA from MLR content ONLY, each claim cited
  to its approved asset and its approved source
Medical-inquiry agent: firewalled from commercial context by architecture
  (separate identity, separate registry scope, separate KG partition)
AE-detection sidecar: mandatory on every patient/HCP conversation; routes to
  PV (UC-X1) with the 24h / 15-day clock started
        
Compliance gate: OPDP / local code rules; MLR expiry; consent check
Agent topology
Field-force assistant; medical-inquiry agent (separate principal); AE sidecar (mandatory, non-bypassable hook). Kits: draft-then-check (the checker is the promotional-compliance rule set) + entity-resolution. The partner-built agents run inside the platform contract (§8.8, UC-H1).
GxP frameOPDP / EFPIA / GVP AE reporting
Key metricMLR violation rate (target: zero); AE capture rate; rep productivity; call quality
Proven patternFortune 50 pharma Commercial CRM + Customer MDM — 100K+ HCP/HCO, 20+ system integration, NBA in production
draft-then-checkentity-resolution
UC-C2

MLR Content Review & Claims Intelligence

Decision. Is this asset approvable — every claim supported, every reference current, every market rule met — before the MLR committee meets?
Architectural pattern
Draft asset → extraction agent structures claims → deterministic path: claim-to-source rules (each claim must trace to a label statement or approved study section; reference currency; fair-balance rules per market) → AI path: cited review memo → MLR reviewer disposition → approved asset registered with claim-level provenance, feeding UC-C1.
Kitdraft-then-check — the exact Doscierge pattern applied to promotional material
GxP frameOPDP / local codes
Key metricMLR cycle time; first-pass approval rate; post-approval claim challenges (target: falling)
draft-then-check
UC-C3

Commercial Intelligence & Market Signals

Decision. Which market movements need a response — with the evidence.
Architectural pattern
Signal specialists over competitor, formulary, launch, and prescribing feeds → orchestrator → NBA to brand teams → the Commercial Intelligence Command Center (§12).
Kitsignal-and-NBA
GovernanceData-use and privacy constraints (de-identified, market-level) enforced in the KG partition, not by policy
Key metricSignal-to-response time; signal precision (vs. noise)
signal-and-NBA
Transition. The AE sidecar hands off to the last tower. Medical & Safety is where the enterprise answers for the product after it is in the world.
8.7

Medical & Safety Tower

GVP / Part 11 / PSMF for safety · medical-affairs approved sources only · scientific-exchange and off-label boundaries as versioned rules per market (FDA vs. EFPIA) · Sunshine Act / Open Payments and EFPIA disclosure on every transfer of value · firewalled from commercial, architecturally · GDPR / HIPAA on HCP data · AI use in PV documented in the PSMF
Decisions that are slow today

Is this a serious, unexpected adverse event — and is the 15-day clock met? What should the response to this medical inquiry be, from approved sources, in the inquirer's language? What should this MSL know before Thursday's meeting with this KOL — prior interactions, open inquiries, the KOL's last three abstracts, the approved SRLs that answer what they are likely to ask, and what cannot be shared? Which of the six thousand abstracts at this congress change the evidence narrative for our product or a competitor's, and which therapeutic-area lead needs to know tonight? What is the field hearing, in aggregate, that should change the medical strategy — and is it a pattern or three anecdotes? Where are the evidence gaps that justify an investigator-initiated study or a real-world-evidence analysis? Is this medical slide deck approvable, and which claim is the one that will be challenged? Medical Affairs is where the platform's evidence-provenance discipline is tested by the hardest clocks (pharmacovigilance), the most explicit boundary rules (scientific exchange vs. promotion, per market), and — since the EMA/FDA AI reflection papers — the most explicit documentation requirement: AI use in pharmacovigilance must be documented in the PV System Master File.

Why this tower is often the right first wave

Medical Affairs work is text-native (inquiries, interaction notes, abstracts, publications, SRLs), high-volume, and governed by approved-source discipline rather than GMP validation. The regulatory frame is real — GVP, Part 11 on signed records, Sunshine Act / Open Payments and the EFPIA Disclosure Code on every transfer of value to an HCP, FDA and EFPIA scientific-exchange boundaries, GDPR on HCP data — but it is a frame the draft-then-check and rule-gate kits already satisfy. Several top-20 companies have published Medical Affairs AI proof points: semantic clustering of unsolicited medical-information inquiries; GenAI theme analysis of field medical insights; GenAI content drafting. What none of them has published is the platform underneath — the shared medical semantic layer, the canonical interaction model, the firewall, the telemetry plane — that turns three standalone tools into a portfolio where the fourth product costs less than the first. That is the gap this tower is designed to close, and it is why the Lab → Factory model (§7) applied to Medical Affairs is described at the end of this section.

Governance posture

GVP / Part 11 / PSMF for safety. Medical-affairs approved sources only. Scientific-exchange, unsolicited-request, and off-label boundaries enforced by versioned rules per market (FDA vs. EFPIA), not by prompt. Sunshine Act / Open Payments (42 CFR 403) and EFPIA Disclosure Code on every HCP transfer of value — advisory boards, consulting arrangements, investigator-sponsored study grants, congress sponsorships — each traced from source event to the reported aggregate spend figure with threshold alerting; the reportable amount is computed deterministically, never by an agent. Firewalled from commercial (architecturally, per §8.6). GDPR / HIPAA on HCP and patient data; multi-lingual, regionally parameterized pipelines — rulesets as configuration tagged by jurisdiction, never regional code forks.

Medical & Safety KG — what it must know
case ──received_via──► channel (literature, call center, social, site, partner, field) ──on_date──► clock_start
  ├── has_patient (de-identified) / has_reporter / has_product ──► product ──has_label──► label (version, market) ──expected_event──► meddra_pt
  ├── has_event ──► adverse_event ──coded_as──► meddra_pt ──seriousness──► criterion (death, hospitalization, ...)
  │                    ├── expectedness ──vs──► label (by version, by market)
  │                    └── causality ──assessed_by──► method (WHO-UMC, Naranjo) ──by──► pv_physician
  ├── has_narrative ──► narrative ──cited_to──► source_span[]
  ├── reported_to ──► authority (E2B(R3), deadline, status)
  └── contributes_to ──► signal ──across──► case[] ──from──► FAERS | EudraVigilance | internal
hcp (→ core golden record; medical partition: specialty, affiliations, KOL tier, consent, territory)
  ├── has_msl_interaction ──► msl_interaction (type: scientific exchange | advisory board | congress | virtual;
  │                              date; topics[]; content_shared[] ⊆ approved_content; follow_ups[])
  │                              └── yields ──► field_insight ──themed_as──► insight_theme (TA, product, topic;
  │                                                confidence CAPPED until n_sources ≥ N) ──supports──► medical_strategy
  ├── has_medical_inquiry ──► medical_inquiry ──clustered_in──► inquiry_cluster (semantic; volume trend)
  │                              └── answered_with ──► approved_response (SRL, language, market) ──from──► medical_source
  ├── authored ──► abstract | publication ──presented_at──► congress (session, date)
  │                    └── about ──► product | competitor | indication | endpoint (result, population, comparator)
  └── received_transfer_of_value ──► payment_event (nature, amount, date) ──reportable_under──► Sunshine Act | EFPIA Code
product ──has_indication──► indication ──has_evidence_gap──► evidence_gap (question, population, endpoint; sources[])
                                              └── addressed_by ──► study_concept (IIS | RWE | Phase 4) ──status──► proposed | funded | published
medical_content ──type──► SRL | slide deck | congress material | medical letter ──has_claim──► claim ──supported_by──► approved_source
                                              └── reviewed_by ──► medical_review (medical | legal | regulatory; status; Part 11)
label_change ──from──► label_version (by market) ──fans_out_to──► approved_asset[] (triggers Vault re-review queue; SRL refresh)
                                                 ──updates──► expectedness_reference (by market, for UC-X1 case processing)
Aligned to: MedDRA, E2B(R3), CIOMS, GVP, ICH E2A–E2F, Sunshine Act / Open Payments, EFPIA Code; core: Product, Document, HCP, Study
THE MEDICAL AFFAIRS LOOP — one HCP record, one interaction model, four surfaces, and a firewall no arrow crosses FIREWALL · commercial estate beyond this line (§8.6) X2 · MEDICAL INQUIRY DESKcluster → AE screen → cited responseoff-label boundary by rule, per market X4 · MSL WORKBENCHpre-call brief, every line cited · works offlinenever auto-logged to CRM X3·X6 · MEDICAL STRATEGY DASHBOARDthemes · evidence gaps · study conceptsadopt / monitor / dismiss · fund / decline X7 · MEDICAL CONTENT REVIEW DESKclaim-to-source · fair balance · cross-asset checkvendor triage in the contract · human sign-off X5 · CONGRESS INTELLIGENCEabstract → checker → KG → alert in hours HCP / KOLgolden record (MDM) every surface reads it; none owns a copy approved responses → what the MSL can share (X2 → X4) post-interaction insight capture→ corroborated themes (X4 → X3) evidence gaps → studies → publications → medical content (X6 → X7) approved content →inquiry responses (X7 → X2) signals KOL abstracts → brief CANONICAL MEDICAL AFFAIRS DATA MODEL — interaction · inquiry · insight · content · hcp · event · transfer_of_value · materialized at run-time Veeva CRM → Vault CRM | Salesforce LSC (connector swap) Vault MedComms · PromoMats (approved content) HCP/HCO MDM · Marketing Cloud · data platform
The Medical Affairs integration architecture. Medical Affairs sits on the seam of the enterprise's most fragmented system estate: a field CRM for MSL interactions and medical inquiries (Veeva CRM for most of the industry today, with an end-of-support date that forces a choice between Vault CRM and Salesforce Life Sciences Cloud before 2030); a content platform for medical materials (Vault MedComms, with PromoMats for anything that touches promotion); an enterprise HCP/HCO master; a data platform for interaction analytics; a marketing-cloud stack for omnichannel; and, increasingly, the CRM vendor's own agent layer. The architectural decision that makes every use-case below survivable is to abstract the interaction model: agents and analytics bind to a canonical Medical Affairs data model — interaction, inquiry, insight, content, HCP, event, transfer of value — materialized at run-time from whichever CRM the persona works in (§5.4's pattern). The CRM migration then becomes a connector swap, not an agent rewrite. System-of-record discipline follows: the HCP golden record stays in MDM; the approved-content record stays in Vault; the interaction and inquiry records stay in the CRM; engagement analytics live in the data platform; and agents read from data products, never from four systems, and write decisions back only to the workflow surface the human uses. The field CRM's offline-first constraint is inherited by anything the MSL touches — a brief that needs connectivity is a brief that does not get read.

The CRM transition is a data-governance event, not a migration project. The working pattern across multinationals navigating the Veeva CRM end-of-support window: bind agents and analytics to the canonical interaction model above — Interaction, Inquiry, Insight, Content, HCP, Event, Transfer of Value — materialized from whichever CRM the persona works in today. The canonical model is the data contract; the CRM connector is the adapter. When Veeva CRM gives way to Vault CRM or Salesforce Life Sciences Cloud (or both, in different markets, for three years), the migration is a connector swap, not an agent rewrite. Three implications for the platform: (1) every agent data contract references the canonical model, never CRM-native objects; (2) the dual-CRM state — some markets on Veeva, some on Salesforce LSC — is the default operating mode through at least 2029, so the canonical model must run both adapters simultaneously; (3) regional data residency must be re-verified per market during migration — a Salesforce LSC Hyperforce deployment is multi-tenant, and EU data residency requires explicit configuration (Hyperforce EU region + Data Residency add-on), not assumption.

WHERE THE 15 DAYS GO — the agents compress the segments before the human, not the human's TODAY intake 5 channels · 2d manual structuring · 2.5d MedDRA coding · 2d narrative drafting · 3d QC · 2d physician review · 2.5d submit DAY 15 — hard regulatory clock PLATFORM intake struct. coded narrated checked PV physician review + sign (Part 11) submit clock headroom — signal detection runs here; PSMF documents the AI role intake specialistsper channelcoding specialistPT + confidencenarrative agent · CIOMS · span citationschecker: E2B(R3) elements · field consistencyrules: seriousness · expectedness (label by version/market)causality support — signed by PV physicians the red segment is the only one that can stretch · kits: draft-then-check + rule-gate — the same substrate as UC-D4 with a different ontology and a harder clock
The case as a timeline with the clock drawn as a bar: receipt → structured → coded → narrated → checked → signed → submitted, with each agent's contribution shown as a segment and the human sign-off as the only segment that can stretch. The agents compress the segments before the human, not the human's.
UC-X1

Pharmacovigilance & Safety Case Processing

Decision. Seriousness, expectedness, causality — and a submitted, compliant case within the clock.
Architectural pattern
Intake channels: literature, call center, social media, clinical sites, partner
  reports, the commercial AE sidecar (UC-C1)
        
Intake specialists (per channel): extract structured case data with provenance;
  clock starts at receipt
        
  Deterministic path: seriousness, expectedness (vs. label version by market),
    causality-support rules — versioned, calibrated and signed by PV physicians
  Coding specialist: MedDRA PT suggestion with confidence
  AI path: narrative drafting agent — CIOMS format, span-level citations to source
        
Checker: E2B(R3) required elements; narrative-to-structured-field consistency
        
PV reviewer sign-off (Part 11) → submission → clock compliance recorded;
  AI role documented in the PSMF
        
Signal-detection specialist: cross-case pattern recognition; FAERS /
  EudraVigilance integration → signal dossier → safety physician review
Agent topology
Intake specialists; coding specialist; rule engine; narrative drafter; checker (MCP tool); signal-detection specialist. Kits: draft-then-check + rule-gate — the same substrate as regulatory documentation (UC-D4) with a different ontology and a harder clock.
GxP frameGVP / Part 11 / PSMF
Key metric15-day clock compliance; coding accuracy vs. PV-physician review; case processing time; signal lead time
Proven patternProtoCheck expert-calibrated rules with physician signature; Doscierge draft-then-check; eval harness per document type
draft-then-checkrule-gate
UC-X2

Medical Information & Inquiry Response

Decision. The response to this inquiry — from medical-affairs-approved sources only, with provenance, triaged to an MSL where judgment is required.
Architectural pattern
Inquiry (CRM Medical Inquiry object, call center, portal, email — any language) → intent, product, and market classification → semantic clustering against the rolling cluster set (the pattern published as semantic clustering of unsolicited medical-information inquiries), so that volume trends per cluster become an insight feed for UC-X3 → deterministic AE / product-complaint screen: any hit routes to PV (UC-X1) with the clock started and human confirmation required — a classifier and a rule, never a prompt → retrieval from the medical approved-source registry (SRLs, label by market and version, publications — the same registry mechanism as UC-D4, different contents) → response draft with span-level citations in the inquirer's language → deterministic checks (unsolicited-request and off-label boundary rules per market, FDA vs. EFPIA; no promotional content; consent) → medical-information specialist disposition or MSL triage → response logged with provenance back to the CRM inquiry record.
Kitsdraft-then-check + rule-gate
GovernanceArchitecturally firewalled from the commercial estate (separate principal, registry scope, KG partition)
Key metricResponse time; citation completeness (target: 100%); off-label boundary violations (target: zero); AE-in-inquiry capture rate; cluster-to-insight lead time
Proven patternDoscierge draft-then-check; Enterprise Knowledge Vault retrieval at 99%+ over a curated registry
draft-then-checkrule-gate
Regional data scope for Medical Affairs agents. Medical Affairs illustrates the regional challenge more sharply than any other tower because a single workflow combines HCP personal data (GDPR / CCPA), patient-adjacent data (inquiry content, adverse-event reports), and public scientific data (abstracts, publications). The residency scope on each agent's card (§6) and the classification-aware fabric (§5.4) resolve it per agent, not per tower:
Medical agentRegional data (stays in zone)Global data (shareable)Firewall and routing
MSL pre-call brief (UC-X4)Interaction history (regional CRM partition), inquiry history (regional MI), HCP prescribing (regional)Publications (PubMed), congress abstracts, KOL tier score (derived from public authorship and affiliation), approved SRLs by marketMedical-only principal; commercial interaction data excluded even within the same region. Card scope: REGIONAL-*, one instance per zone
Medical inquiry response (UC-X2)Regional SRL library (market-specific off-label rules), HCP identity, inquiry textApproved-source registry structure, indication / product reference, label version by marketAE-detection hook runs on every inquiry inside the zone; hits route to the regional PV agent with the clock started. Card scope: REGIONAL-*
Congress intelligence (UC-X5)Attendee HCP profiles where regionally enriched (interaction history, territory)Abstract text, competitor data, endpoint results — all publicOutput splits: the intelligence report is SHAREABLE; HCP-linked insights are REGIONAL. Card scope: GLOBAL for the abstract pipeline, REGIONAL-* for the attendee-enrichment step
Field insight capture (UC-X3)MSL note text — HCP PII plus scientific-exchange content — REGIONALTherapeutic-landscape KG, evidence gaps, insight themes (SHAREABLE-DERIVED)Raw insights regional; themed, de-identified insights promoted to the global insight repository after the steward attestation (§5.4). Card scope: FEDERATED orchestrator over REGIONAL-* extraction specialists
PV signal detection (UC-X1)Regional ICSR data, regional safety databases (EudraVigilance-mirrored EU cases, FAERS-mirrored US cases, NMPA-reported CN cases)Aggregate cross-regional signal statistics via federated queryMandatory HITL: a PV physician reviews every signal. Individual case transfer between regions — when a case genuinely must move — goes through E2B(R3) with the transfer mechanism recorded on the case, never through the agent. Card scope: FEDERATED
Evidence generation planning (UC-X6)Evidence-gap repository entries linked to specific regional patient populations; investigator-site performance data (regional CRM partition)Published literature, competitive landscape, congress abstracts (public), therapeutic-area strategy (internal)Gap identification is GLOBAL (public evidence); study-concept development requires REGIONAL patient-density and site data. Card scope: FEDERATED orchestrator dispatching to REGIONAL-* data-access specialists
Medical content review (UC-X7)SRL drafts with claim-level citations to regional label versions; regional off-label boundary rules (market-specific)Medical approved-source registry, core label text, regulatory requirements (SHAREABLE)Content creation is GLOBAL (label and regulatory reference); off-label boundary checking is REGIONAL-* per market because each market's approved indications and local regulations differ. Card scope: GLOBAL with REGIONAL-* rule-gate invocation per target market
The pattern is capture regionally, aggregate globally, act locally. The field-insight pipeline shows it end to end: the extraction specialist structures MSL notes inside the regional zone → the theme specialist produces de-identified themes with corroboration counts → themes are promoted to the global insight repository as SHAREABLE-DERIVED with lineage to their regional sources → TA leads in Basel see the global pattern on the medical strategy dashboard → MSLs in each territory act on local plans informed by the global theme, with the individual interactions that produced it never having left their region. The same three verbs describe signal detection (cases regional, signal global, label-change action local by market) and congress intelligence (attendee enrichment regional, abstract analysis global, KOL follow-up local).

CRM transition does not change the scope table. The canonical interaction model (§8.7 integration architecture) means the regional scope declared on each agent's card is CRM-agnostic. Whether the EU MSL works in Veeva CRM today or Vault CRM in 2028, the extraction specialist that structures interaction notes runs REGIONAL-EU against the same logical data product — msl_interaction — materialized from whichever CRM instance the region uses. The scope table above will survive the industry's CRM transition without a single card change, because the card declares a data residency scope, not a system binding. This is the practical consequence of abstracting the interaction model before the first agent ships.

UC-X3

Field Medical Insights & Medical Strategy

Decision. What the field is hearing, in aggregate, that should change the medical plan — and whether it is a pattern or three anecdotes.
Architectural pattern
Insight sources: MSL interaction notes (CRM), advisory-board minutes, congress
  debriefs (UC-X5), inquiry clusters (UC-X2), medical-information trends
        
Extraction specialist: structure each free-text insight → (HCP tier, product,
  indication, topic, sentiment, evidence cited, date) with source provenance;
  de-identify at the boundary; multi-lingual
        
Theme specialist: cluster insights into themes (the pattern published as GenAI
  theme analysis of field medical insights) — each theme carries n_sources,
  KOL-tier weighting, geography, and a confidence that is CAPPED until
  corroborated by ≥N independent interactions (the trust-tier rule of §5.3)
        
  Deterministic path: theme-significance rules (volume vs. baseline, tier
    weighting, recency, cross-region concordance) — versioned, signed by the
    medical excellence / insights lead
  AI path: synthesis agent drafts the periodic insights report per TA with
    every theme cited to its interactions; drafts the "what changed since
    last month" delta
        
Medical strategy dashboard (§5.1): themes, evidence gaps (→ UC-X6), inquiry
  clusters, congress signals in one surface; TA-lead disposition (adopt /
  monitor / dismiss) logged — the decision becomes a KG edge on the theme
Agent topology
Extraction specialist; theme specialist; significance rules; synthesis drafter; orchestrator publishing to the medical strategy dashboard through the command-center feed contract (§12). Kits: evidence-dossier + signal-and-NBA. The trap this design avoids is an insight tool that reports what the loudest MSL said. The confidence cap on uncorroborated themes and the tier weighting are what make the output a strategy input rather than a sentiment feed.
GxP frameNon-GxP; GDPR / HIPAA on HCP identity; insights are never re-used for promotional targeting (the firewall)
Key metricInsight-to-strategy time (quarterly → weekly); theme corroboration rate; share of medical-plan changes citing a structured insight
Proven patternFortune 50 pharma Medical Analytics platform — field and congress intelligence with NLP sentiment embedded in the operational platform rather than a standalone tool
evidence-dossiersignal-and-NBA
UC-X4

MSL Pre-Call Intelligence & KOL Engagement Planning

Decision. What this MSL needs to know before this interaction with this HCP — and, at territory level, which KOLs the medical plan says to engage, on which scientific objective, through which channel, next.
Architectural pattern
Trigger: scheduled interaction in the field CRM (call, virtual engagement,
  congress meeting) or the MSL's territory-planning session
        
Pre-call specialist (read-only against systems of record; ≤5 tools):
  KG lookup — HCP golden record (specialty, affiliations, KOL tier, consent)
  Interaction history — prior MSL interactions, topics, content shared, open follow-ups
  Inquiry history — this HCP's medical inquiries and their responses (UC-X2)
  Scientific activity — recent abstracts / publications / congress sessions
    (UC-X5); trial involvement (→ Clinical KG: investigator)
  Approved content — SRLs and medical materials matching the likely topics
        
Brief drafter: one page, every line cited — "what they asked last time, what
  they published since, what we can share, what we cannot"; the cannot-share
  list generated by the same off-label / unsolicited-request rules as UC-X2
        
Engagement-planning specialist (territory level): scientific-objective coverage
  per KOL vs. the medical plan; channel recommendation (face-to-face, virtual,
  congress, advisory board, medical email) from the MEDICAL consent and preference
  record — never from commercial engagement data; advisory board planning and
  execution (attendee selection, topic agenda, logistics) feeds the insight KG
  (UC-X3) and the transfer-of-value lineage (Sunshine Act / EFPIA)
        
MSL workbench (§5.1): brief + plan; MSL accepts / edits; NEVER auto-logged to
  the CRM — after the interaction, the MSL's structured insight capture feeds UC-X3
Agent topology
Pre-call specialist; brief drafter; engagement-planning specialist; orchestrator. Kits: evidence-dossier + draft-then-check (the checker is the scientific-exchange boundary rule set). HCP medical engagement scoring — built on the golden record's medical partition (interaction history, inquiry patterns, congress presence, publication activity, advisory-board participation) — is a shared service consumed by the engagement-planning specialist; architecturally separate from commercial engagement scoring, with medical next-best-actions derived from medical signals only. Considered and rejected: one general MSL assistant with thirty tools spanning CRM write, content, and analytics — tool-selection accuracy degrades and the compliance boundary blurs (§5.2). Omnichannel for Medical is this use-case's territory-level half: medical engagement runs on its own consent and preference record, its own approved-content registry, and its own channel analytics, sharing only the HCP golden record with commercial (§8.6's firewall).
GxP frameNon-GxP; GDPR / HIPAA on HCP data; Sunshine Act / EFPIA disclosure on any transfer of value the plan proposes (advisory boards, speaker programs); the field CRM's offline-first constraint — the brief must be usable without connectivity
Key metricMSL prep time per interaction (hours → minutes); brief acceptance rate in coach mode (§18); scientific-objective coverage per KOL; post-interaction insight capture rate
Proven patternVeeva CRM ↔ enterprise MDM interface architecture at a Fortune 50 pharma — Account synchronization, golden-record resolution, call activity and territory alignment; ProtoCheck's bounded-specialist composition
evidence-dossierdraft-then-checkentity-resolution
UC-X5

Medical Congress Intelligence Pipeline

Decision. Which of the thousands of abstracts, posters, and presentations at this congress change the evidence narrative for our product, our competitors, or our KOLs — and which TA lead needs to know tonight, not in the post-congress report six weeks later.
Architectural pattern
Sources: congress abstract books and late-breaker feeds (ASCO, ESMO, AHA, ...),
  presentation content where licensed, MSL congress debriefs, press signals
        
Extraction specialists (scoped model tier — thousands of abstracts, marginal
  cost matters, §5.2): abstract → (drug, indication, endpoint, result,
  population, comparator, authors → HCP golden record, institution) with provenance
        
Deterministic checker (the generate-then-check discipline of UC-D4): numeric
  extraction validated against the abstract text; endpoint vocabulary normalized;
  product / competitor entity-resolved before anything is written to the KG
        
Medical & Safety KG: congress ──session──► abstract ──about──► product | competitor |
  endpoint ──authored_by──► hcp (KOL tier); provenance-tagged; trust tier "operational"
        
Signal specialists (isolated, each with its own eval record):
  competitor-result watcher (vs. our label and open evidence gaps)
  KOL-activity watcher (who presented what, in whose session)
  safety-signal watcher (any adverse-event mention → PV screen, UC-X1)
  sentiment-on-our-data (from debriefs, where present)
        
Orchestrator → alerts to TA leads within hours, each with the abstract, the KG
  neighborhood, and a proposed action (brief the MSLs · update the SRL · flag
  to evidence generation); congress intelligence surface (§5.1); post-congress
  synthesis drafted with citations
Agent topology
Extraction specialists (scoped tier); checker (MCP tool); three-to-four signal specialists; orchestrator. Kits: signal-and-NBA + draft-then-check + entity-resolution — authors are HCPs are investigators, the same golden record as UC-D2 and UC-C1. Embargo rules for pre-publication abstracts and late-breaker data are enforced as a compliance control: the extraction agent respects embargo dates, and KG edges carry an embargo_status flag that gates downstream distribution until the embargo lifts.
GxP frameNon-GxP; copyright, licensing, and embargo on congress content enforced in the approved-source registry; any adverse-event mention screened to PV
Key metricAbstract-to-alert latency (weeks → hours); extraction precision on numeric endpoints; KOL-activity coverage; share of post-congress actions traceable to a cited abstract
Proven patternFortune 50 pharma Medical Congress intelligence — near-real-time sentiment and competitive intelligence from live congress events, leadership reacting within hours rather than weeks; today's version replaces the NLP engine with LLM extraction gated by a deterministic checker and a provenance-tagged KG
signal-and-NBAdraft-then-checkentity-resolution
UC-X6

Evidence Generation & Real-World Evidence Synthesis

Decision. Where are the evidence gaps for this product and indication — the questions HCPs keep asking that the label and the publications do not answer — and which of the proposed investigator-initiated studies, RWE analyses, or Phase 4 concepts closes the most important gap for the least money?
Architectural pattern
Gap detection: inquiry clusters (UC-X2) + insight themes (UC-X3) + congress
  signals (UC-X5) + payer / HTA questions ──► candidate evidence_gap
  (question, population, endpoint) with the sources that generated it
        
Evidence specialist: what already answers it — internal studies (→ Clinical KG),
  publications, RWE datasets (claims, EHR, registries; de-identified) —
  evidence tiers rendered as in UC-R1's target dossier
        
  Deterministic path: gap-priority rules (unmet-need weight, strategic-objective
    fit, HTA relevance, feasibility flags) — versioned, signed by the medical
    evidence lead; IIS-review completeness rules (protocol elements, budget,
    fair-market value, no promotional linkage)
  AI path: cited gap dossier; study-concept comparison ("three proposed IISs
    address gap G; concept 2 duplicates a registered trial; concept 3 uses an
    endpoint with prior regulatory precedent")
        
Evidence-generation committee reviews → fund / decline / merge logged with
  rationale; funded concepts become study nodes in the Clinical KG; the
  publication plan is drafted against the approved-source registry of study
  data (draft-then-check)
Agent topology
Gap-detection orchestrator (consumes the X2 / X3 / X5 feeds); evidence specialist; rule engine; dossier drafter; publication drafter + checker; systematic literature review (SLR) agent — screening with calibrated inclusion/exclusion thresholds, PICO extraction, evidence tables with source spans, reviewer adjudication on disagreements, and living-review change control (kit: classifier + extraction). Kits: evidence-dossier + rule-gate + draft-then-check.
GxP frameNon-GxP for gap analysis; GCP once a study is funded; publications under ICMJE / GPP guidelines with Part 11 sign-off on submitted manuscripts; SLR screening thresholds documented in the protocol; RWE under data-use agreements enforced by the PII scope class (§6)
Key metricGap-to-decision time; share of funded studies traceable to a structured gap; publication review cycles; duplicate-study avoidance; SLR screening time (weeks → days)
Proven patternFortune 50 pharma clinical-studies assessment and funding intelligence — post-commercialization tracking of investigator activity and indication-expansion signals to decide which studies to fund; clinical data integration across 100+ studies on SDTM/ADaM
evidence-dossierrule-gatedraft-then-check
UC-X7

Medical Content Review & Scientific-Exchange Compliance

Decision. Is this medical asset — an SRL, a slide deck, a congress booth panel, a medical letter — approvable: every claim supported by an approved source, fair balance met, no promotional drift, correct for this market — before the medical review committee meets?
Architectural pattern
Draft asset (Vault MedComms / PromoMats) → extraction agent structures claims and references → deterministic path: claim-to-source rules (each claim must trace to a label statement, an approved study section, or a peer-reviewed publication in the registry; reference currency; fair-balance and scientific-exchange rules per market; an off-label language detector as a rule, not a prompt) → AI path: cited review memo naming the three findings a reviewer is most likely to challenge → medical / legal / regulatory reviewer disposition → approved asset registered with claim-level provenance, feeding UC-X2 and UC-X4 as approved content.

Build vs. buy. Where the content platform ships its own review agents — vendor quick-check and content agents now exist for promotional material, and specialist MLR-automation vendors target large reductions in review labor — the platform consumes them under the platform contract (UC-H1) for first-pass triage and builds only where vendor coverage stops: the enterprise's own claim-substantiation rules, the medical (non-promotional) boundary, and the cross-asset consistency check ("this SRL and this slide deck make different statements about the same endpoint"). Human sign-off stays.
Kitdraft-then-check — the same Doscierge pattern as UC-C2, applied to medical rather than promotional material
GxP frameFDA / EFPIA scientific-exchange rules; Part 11 on the approval record
Key metricMedical review cycle time; first-pass approval rate; cross-asset inconsistencies caught before committee (rising at first, then falling as the registry cleans up)
draft-then-check
Medical AI Acceleration — the Lab → Factory model applied to this tower. A Medical Affairs AI portfolio typically arrives as three or four standalone proofs — an inquiry-clustering tool, an insights-theming tool, a content-drafting assistant — each with its own retrieval, its own prompts, its own vendor. The acceleration step is not another proof; it is the shared substrate: one medical semantic layer and canonical interaction model (above), one approved-source registry with medical contents, one firewall, one telemetry plane feeding TCO, one registry of bounded specialists. A sequence that has worked: (1) a demand study across the seven decisions above — most Medical Affairs calendars are dominated by Type 1 waste (finding the prior interaction, re-deriving the approved answer, reading the abstract book); (2) UC-X4 first, because the pre-call agent is the substrate-building first product — it forces the KG, the approved-source registry, MDM resolution, and the canonical interaction model into existence, all with the lowest validation burden (read-only, advisory, no regulated record written); MSL brief acceptance in coach mode is the adoption signal the whole portfolio needs; (3) UC-X2 second, because the MI response assistant is the highest-volume, highest-ROI product that inherits the registry, off-label rules, and AE classifier the first one built; (4) UC-X3 and UC-X5 third, because they turn the field and the congress into structured KG edges that UC-X6 then reasons over; (5) UC-X7 can run in parallel with (3)–(4) as a buy decision where the content platform's own agents are in place to triage — the enterprise builds only the claim-substantiation judgment layer; (6) the PV pipeline (UC-X1) is third-wave, co-owned with Safety, proving the same kits survive the hardest clock and the most-inspected process. The promotion gate is unchanged from §7; the validation vote here belongs to Medical Compliance and Legal rather than Quality, and the readiness vote belongs to the MSLs and medical-information specialists who will use it. The portfolio is measured on the seven decision latencies, brief-acceptance rate, citation completeness, and cost per decision — the same scorecard as every other tower (§18).
Transition. Seven towers, and one thing they have in common: none of them works without the horizontal underneath.
8.8

Enterprise Horizontal

The layer is infrastructure; GxP applies to what consumes it · GAMP 5 infrastructure qualification · Part 11 where relevant

Two use-cases serve every tower. Neither is a use-case in the traditional sense; both are the shared substrate that makes the towers cheap.

UC-H1

IT Operations & Partner-Built Agents (The Platform Contract)

Decision. How does the platform govern agents that a partner builds and operates — incident triage, change management, access provisioning, cloud cost — so that the enterprise does not end up with three incompatible agent estates?
Architectural pattern
Agent platform contract — published by the enterprise AI platform team:
  • Agent as a first-class identity principal (G5)
  • Tool access via MCP servers with policy enforcement (allowlist per identity)
  • Centralized observability and tracing (G4)
  • Registry card, eval record, and approval workflow required for production (§6)
  • Part 11 audit trail where change management touches validated systems (G3)
        
Partner-built agents (incident triage, change management, access provisioning,
  cloud cost) deploy INSIDE the contract — same identity model, same MCP tool
  access, same observability, same gate as internal agents
        
Platform-enforced controls: tool allowlist, rate limiting and cost governance,
  execution history in the central trace store, evaluation record before production
Key principle
Contract over bespoke. Partners build inside the platform patterns, not around them. The same contract governs the CRM vendor's commercial agents (UC-C1) and the IT partner's operations agents — which is what makes "three estates" into one governed estate.
GxP frameGAMP 5 infrastructure qualification; Part 11 where relevant
Key metricAgent-estate coherence — percentage of agents governed under the platform contract (target: 100%)
Proven patternMCP server published on the MCP registry (AAIF); 18+ IaC modules underlying 8 production systems
platform contract
UC-H2

Enterprise Knowledge & Semantic Search (KG-as-a-Service)

Decision. How fast can a new use-case get its first grounded, cited answer? The semantic layer's KPI is days from a new use-case request to its first grounded response — and it fails if it is either a central bottleneck or a set of disconnected team-built graphs.
Architectural pattern
Federated ontology governance (§5.3):
  • Thin core ontology — centrally owned
  • Domain ontologies — team-owned under a published modeling standard
        
KG-as-a-Service: query, provenance, versioning APIs via MCP; any agent calls it;
  none builds its own retrieval
        
Curation agents: auto-ingest from data products; entity extraction and linking;
  relationship discovery; freshness scoring and staleness flags
Governance agents: ontology drift detection; orphan entities; cross-domain
  consistency; trust-tier audits
Retrieval agents: provenance-first semantic search across all domain KGs;
  federated query
        
Aligned to public standards — CDISC, IDMP, MedDRA, MONDO, ChEBI, FHIR, ISA-88/95,
  GS1 — rather than invented
GxP frameThe layer is infrastructure; GxP applies to what consumes it
Key metricTime-to-first-grounded-answer (target: days); reuse ratio of core entities across domain KGs; provenance completeness
Proven patternEnterprise Knowledge Vault — powers all 8 production systems; RAG accuracy 75% → 99%+; 10,292 KG triplets
KG-as-a-Service

The federated ontology, explorable

Core entities in the center; seven domain rings around them; every entity and predicate named in the tower fragments above, plus the cross-domain joins. Click a node to light its neighborhood and see which use-cases consume it and which data products it materializes from. Click a domain label to isolate that subgraph. Pick a scenario to watch a traversal and the dual-path reasoning it produces. Esc resets.

100%
Normativeregulations · signed rules · 1.00Structuralontology · schema · 0.95+Operationalsystem-of-record rows · 0.90+Carry-forwardprior decisions · 0.80–0.90Enrichmentanalytics · 0.70–0.90Mined / Inferredprocess intelligence · HARD CAP 0.75
core entity domain entity evidence / provenance model / rule artifact ResearchClinicalCMC/PDManufacturingSupplyCommercialMedical edges: solid = normative/structural/operational · dashed = carry-forward · dotted = enrichment/mined

Interactive ontology explorer: core entities in the center, seven domain rings around them, a node click highlighting every domain that touches it — so that the reader sees, for example, that Batch is a core entity consumed by Product Development, Manufacturing, Supply Chain, and Medical (via product complaints), and that HCP is consumed by Development (investigator), Commercial, and Medical.

Eight towers. Twenty-plus use-cases. And, counted honestly, five kits, one governance plane, one gate, one graph. That count is the next section.
Section 9 · Enterprise Coherence

Factory over snowflakes — what gets shared

The platform's value is reuse. The previous section showed twenty-plus use-cases; this section shows how few distinct things they are built from — and makes the economic case that follows from that count.

9.1 The coherence matrix — what it shows and how to read it

Each filled cell means the use-case consumes this shared service — built once, owned once, validated once — rather than building its own.

The matrix answers one question per cell: does this use-case consume this shared platform service, or does it have to build its own? A filled cell means the use-case consumes the shared service — the service exists once, is owned once, is validated once, and the use-case is configured onto it. An empty cell means the service is not relevant to that use-case. There are no cells for "builds its own copy," because the platform rule is that there is no such option.

Columns are use-cases, grouped by tower and labeled with their §8 codes. Rows are shared platform services, grouped by the layer they live in. The final two columns show how many use-cases consume each service (the reuse count) and which team owns it — because a shared service without an owner decays. (Hover a column header for the use-case name; hover a row for the layer's role; the legend beneath lists all 23.)

Column key. R1 Target ID · R2 Gen Chem · R3 Preclinical Safety · D1 Trial Design · D2 Site Selection · D3 Ops Precision · D4 Reg Documentation · P1 Control Strategy · P2 Tech Transfer · P3 Q12 Change · M1 APQR · M2 CPV · M3 PAT Platform · M4 Blend Endpoint · M5 Deviation/CAPA · S1 CoA Intake · S2 Supply Risk · C1 HCP Engagement · C2 MLR Review · X1 Pharmacovigilance · X2 Medical Info · H1 Platform Contract · H2 KG-as-a-Service

The five Medical Affairs use-cases added in §8.7 (X3 insights, X4 MSL pre-call, X5 congress, X6 evidence generation, X7 medical content review) are omitted from the matrix for width. They consume only existing rows — core ontology and KG service, entity resolution, approved-source registry, the rule-gate / draft-then-check / signal-and-NBA / evidence-dossier kits, stateful orchestration, the eval harness, the registry, the platform contract, the human-in-the-loop gate, and the command-center feed contract — and add no new platform service. That is the point: a tower with seven use-cases and zero new rows.

How to read it in thirty seconds

Look down the Reuse column. Four services are consumed by essentially everything — the evaluation harness, the registry, the platform contract, and (with one exception) the core ontology and KG service. Those are the platform's foundation and are built first (§13, Phase 1). Then look across the Manufacturing columns (M1–M5): five use-cases in the most heavily governed tower in the enterprise, and not one row is new — every platform service they consume already exists by the time Development has scaled, including the model registry, which Development's simulation and prediction agents stood up first. What Manufacturing adds is its own data products and its own KG, not new platform services. That is the reuse dividend, made visible.

The same matrix as information flow

The dot matrix answers "does this use-case consume this service?" It does not show how services flow between layers. Hover a service band to see which towers consume it; hover a tower to see what it consumes; click a flow to list the use-cases at that intersection. Line thickness is the number of consuming use-cases.

Hover or click a flow to inspect it.

9.2 The reuse dividend — the economic case for one platform

Without sharing, each use-case reinvents its own KG, its own evaluation, its own governance — at an estimated $2–5M per standalone implementation. With the shared substrate, the marginal cost of use-case N+1 drops 40–60% versus use-case N. By use-case five or six, new agents are configuration, not projects. The platform's value is not in any single use-case; it is in the compounding reduction of time-to-value as the registry fills and the graph deepens.

There are two ways to lose the dividend, and the matrix guards against both:

  • Snowflakes. A use-case that builds its own retrieval, its own eval, its own audit trail — usually because the platform service was not ready when the team needed it. The mitigation is sequencing (§13) and the Lab's time-to-first-grounded-answer metric: the platform must be faster than going around it.
  • Monoliths. The reverse failure — one enormous agent that "does manufacturing." §5.2 made the case: performance degrades, token costs compound 30–70% annually, security boundaries blur, and validation becomes impossible. The platform pattern is composition: bounded specialists with narrow scopes and minimal tool sets, a coordinator that delegates, and a registry that makes the monolith structurally impossible. When connectors fail, degraded-mode design automatically caps confidence on affected data, actions fall below autonomy thresholds, and the human takes the driving seat — no manual intervention required to fail safe.

There is a demand-analysis reading of the dividend that is easier to defend to a CFO than a cost-per-project figure. A snowflake does not cost $2–5M once; it generates failure demand for years — the re-mapping when two ontologies disagree, the second validation package for the same rule engine, the reconciliation meeting when the CPV agent and the deviation agent describe the same batch differently. Those are Type 2 activities: they exist only because the platform service was not built once. The coherence matrix is therefore a map of failure demand avoided, service by service. The reuse count in the right-hand column is the number of times a system condition — per-project retrieval, per-project eval, per-project audit — was not recreated. Measured this way, the marginal-cost curve in §18 is not a productivity claim; it is the enterprise's failure-demand ratio falling as the registry fills.

Every one of these shared use-cases reasons the same way. That way is the next section.

Section 10 · Dual-Path Reasoning

Deterministic where you can, generative where you must, merged with provenance

Every use-case in the platform — discovery or GMP, research or field — follows the same reasoning architecture: deterministic where you can, generative where you must, merged with provenance. This is not a design preference. In a domain where every output may have to be explained to an inspector, it is the only architecture that can be both fast and defensible.

EVIDENCE — KG + data products + approved-source registry
↓            ↓

Deterministic path

  • Expert-calibrated rules — owner, version, calibration record
  • Provisional → confirmed: ships when an SME signs
  • Reproducible: same input → same output
  • Validatable: GAMP 5 Cat 5, conventional IQ/OQ/PQ
  • Inspectable: every rule readable, every firing traceable

Generative path

  • KG-grounded LLM — reasons over the semantic layer, not hallucinated context
  • Span-level citations: every claim traces to the approved registry
  • Refuses unsupported claims — architectural, not prompted
  • Controlled by grounding + checker + human — not by validating the model
↓            ↓
MERGED ASSESSMENT — one reviewable objectfindings + draft + provenance + confidence → human decides → decision + rationale become the audit record (G2)

Where each path lives, by tower. In Research, the deterministic path is evidence-tier scoring and structural alerts; the generative path is the target dossier and the liability brief. In Development, it is the 500-rule protocol check and the document checker versus precedent synthesis and drafting. In Manufacturing, it is validated SPC/capability queries and batch-record completeness versus trend narratives and root-cause hypotheses. In Commercial, it is promotional-compliance rules versus the pre-call plan. In Medical, it is seriousness/expectedness rules versus the case narrative. The proportion shifts — Manufacturing is deterministic-heavy, Research generative-heavy — but the shape never changes, which is why one validation strategy covers all of it.

TowerDeterministic pathGenerative pathBalance
ResearchEvidence-tier scoring; structural alertsTarget dossier; liability briefGenerative-heavy
Development500-rule protocol check; document checkerPrecedent synthesis; draftingBalanced
ManufacturingValidated SPC/capability queries; batch-record completenessTrend narratives; root-cause hypothesesDeterministic-heavy
CommercialPromotional-compliance rulesPre-call planBalanced
MedicalSeriousness / expectedness rulesCase narrativeDeterministic-heavy

Intelligent decision support. Human decides. System explains.

Intelligent decision support. Human decides. System explains. Every recommendation carries rationale, provenance, and an owner. The goal is to shift the performance of many individuals closer to the performance of the best. No black-box autonomy in regulated workflows — and, because the decision is captured (G2), the platform learns what "the best" looks like from the humans who are.

The same two paths are what make the system validatable.

Section 11 · Validation Strategy

"You can't validate an LLM under GAMP 5" is the wrong framing

You can validate the controls around the LLM. The question an auditor asks is "which control applies to which component?" — and the answer is three distinct strategies for three distinct system components, applied to every use-case in this article.

GAMP 5 Category 5

Deterministic path

Rule engines, structural alerts, completeness checkers, numeric-consistency validators, SPC/capability query layers, reconciliation logic.

Control. Same input → same output. Conventional IQ/OQ/PQ. Each rule version-controlled with a named owner and a calibration record. Change control on every rule change.

CSA-aligned controls (FDA Computer Software Assurance, final September 2025 — device guidance, applied to pharma by analogy)

Generative / LLM path

Drafting, synthesis, hypothesis generation, narrative.

Control. Controlled by grounding (approved-source registry — the only citable source), checking (the deterministic path post-processes every output), and human review. The model is qualified against its context of use; deterministic controls bound residual risk. Model version pinned; eval harness measures P/R/F1 against labeled ground truth before any deployment and on every release. Draft EU GMP Annex 22 (July 2025) is scoped to static, deterministic models and excludes dynamic (self-learning) and generative/probabilistic models from critical GMP use — the dual-path rationale this architecture implements: the critical decision stays on the deterministic path, and the LLM works where a qualified human reviews its output.

ICH Q2/Q14 posture · model registry

Statistical / chemometric models

MSPC, PLS, enrollment simulators, ML scorers.

Control. Analytical-method lifecycle, not software validation: NOC set, cross-validation, held-out detection test, QA approval in the model registry, performance monitoring against the reference method.

21 CFR Part 11 · EU Annex 11

Human sign-off

Every output that enters a regulated record.

Control. Electronic signature. The human is not a rubber stamp; the system is designed so the human has the context — evidence, findings, confidence — to decide. Edits captured as preference signal.

Risk-proportionate by design. CSA's critical-thinking approach means the depth of assurance follows the risk of the decision. The blend-endpoint advisory (UC-M4), which runs beside a still-validated fixed-time step, carries a lighter package than the same endpoint in RTRT mode. The APQR narrative agent (UC-M1), whose output a QA reviewer signs, carries a lighter package than the SPC query layer, whose numbers the narrative cites. The governance manifest on each agent card (§6) records which posture applies, so the validation package is a rendering of the card, not a document written from scratch.

A note on honest complexity

This validation posture is defensible but not established precedent. FDA's 2025 draft guidance on AI in regulatory decision-making does not yet address the hybrid model explicitly; EMA's adopted reflection paper (September 2024) leaves the same room. The CSA guidance (issued for devices, and applied to pharma manufacturing by analogy) provides the framework for risk-proportionate validation that avoids over-testing deterministic systems while establishing appropriate controls for AI components. Early movers who articulate this posture clearly to Quality and CSV — and document it before the first agent enters a GxP workflow — will set the standard. Those who retrofit it will spend twice as long.

Continued in Part 2. How to prove this posture is the subject of Part 2. It runs the standard CSV/CSA lifecycle stage by stage on ProtoCheck and Doscierge (§17), then works through the acceptance criteria against a measured human baseline, the eight lifecycle controls that carry CSV into AI, locked vs. adaptive models, drift, monitoring and data integrity, question by question, in Validation Strategy, Applied: Agentic AI in GxP.

Validated agents in production generate signals. Where those signals land is the next question — and the answer is not a dashboard.

Section 12 · Agentic Command Centers

Dashboards show what is happening. Decision hubs decide.

Industry analysts have been pushing "command centers" and "control towers" in pharma for several years, and the leading sponsors have built them: a control tower over hundreds of studies that cut non-screening sites to 12% in six months and cut data discrepancies from roughly 30,000 to 8,000; manufacturing insight centers with real-time visibility across sixty-plus sites and twenty ERP instances. These are real and valuable. They are also, almost universally, dashboards: they show what is happening. The agentic platform lets them become decision hubs: surfaces where multi-domain agent feeds arrive as next-best-actions with rationale, impact, and owner; where the human decides; and where the decision is logged.

12.1 What a command center is, architecturally

A command center is an Experience-layer composition over registered agent feeds. It has no logic of its own beyond composition, prioritization, and presentation; every signal it shows was produced by a registered monitoring or orchestrator agent under the platform contract, and every action taken from it is captured as decision provenance (G2).

The feed contract. An agent that wants to publish to a command center registers a feed on its card: signal schema, severity taxonomy, SLA, owner, drill-through pointers. The command center subscribes by domain and role. This is how a new use-case appears on the center's surface without anyone building a new dashboard — it is configuration on the registry.

Cross-domain correlation is the reason command centers are agentic

A dashboard shows the CPV signal on line 2 and the supplier's late CoA and the deviation opened yesterday as three tiles. An agentic command center's orchestrator traverses the KG — batch → genealogy → material lot → CoA → supplier — and presents one NBA: "Line 2 EWMA shift on dissolution correlates with resin lot L-7781; CoA method differs from spec; deviation DEV-2026-0412 opened on the same lot; recommend hold on remaining L-7781 inventory pending QC re-test; owner: site QA head." That is decision velocity: the assembly that used to take three people a day is done before the morning meeting.

CLINICAL OPERATIONSstudy leads · clin-ops heads · portfoliosite actions · supply reallocation · monitoringreal-time · daily stand-up · weekly portfolio MANUFACTURING QUALITYsite quality heads · MS&T · global qualitybatch disposition · investigation · supplier holdsreal-time PAT/CPV · daily huddle · monthly review COMMERCIAL INTELLIGENCEbrand leads · commercial ops · compliancecampaign responses · content refresh · fielddaily · weekly brand review MEDICAL AFFAIRSMA leadership · MSL directors · PV headsPV escalation · MSL deployment · medical strategyreal-time PV · daily inquiries · weekly strategy ENTERPRISE AI OPERATIONSFactory ops · platform leadership · AI governance councilrollback · coach re-entry · budget caps · gate schedulingthe NOC for the agent estate — observes the other four G4 observability feeds from every factory agent — drift · eval drift · acceptance curves · cost per decision · connector health · degraded-mode events COMMAND-CENTER FEED CONTRACT (registry): signal schema · severity taxonomy · SLA · owner · drill-through pointers — subscribe by domain and role D3ops signals D2site stall X1safety signals D4doc ready M2CPV M3PAT drift M4endpoint M5·M1CAPA·APQR S1·S2CoA·chain C3market C1field · AE C2MLR queue X2inquiry themes* X3insights X4MSL pre-call X5congress X6·X7evidence·MLR G4 platform-health agentsevery factory agent's traces Registry lifecycle · gate queuepromotions · deprecations · reviews Each box is a registered agent (§6) publishing a feed under the contract. Colors are towers. *de-identified and firewalled. WHERE A CONTROL TOWER ALREADY EXISTS, THE PLATFORM DOES NOT REPLACE IT The tower becomes a consumer of registered agent feeds — it gains NBAs with provenance and a decision log — and its existing analytics become data products the agents read. Where none exists, the command center is instantiated from the same experience template as any other workspace (§6), with feeds subscribed from the registry. Either way, the command center is a rendering of the platform, not a separate system to be integrated.

12.2 The five command centers — as the operator sees them

Five centers on every screen. Stat cards sense (system-of-record edges — every number is a graph query materialized for the person who needs to act). Alert cards decide (KG walks with min-path confidence). Drafted actions act (threshold-gated autonomy bands). Every approval or correction learns (decision provenance written back). Click a stat card to see the traversal that produced it; click an action to see which agent proposed it and on what evidence.

Clinical Operations Command Center. Domain feeds (from §8): D3 (supply, monitoring, recruitment signals) · D2 (site stall early warning, diversity coverage) · X1 (safety signals affecting ongoing studies) · D4 (document readiness for milestone submissions). Users: Study leads, clinical operations heads, portfolio leadership. Decisions made there: Site actions, supply reallocations, monitoring-plan changes (Part 11), enrollment interventions, portfolio-level trade-offs. Cadence: Real-time signals; daily stand-up; weekly portfolio review.

platform.pharma.internal/clinops/command-centerLIVE · 14 feeds
Clinical Ops
  • Command Center
  • Study Inbox
  • Sites
  • Supply
  • Documents
  • Decision Log
  • Agents
Portfolio stand-up · 42 active studies
Feeds: D3 ops signals · D2 site stall · X1 safety · D4 document readiness — Tue 08:00
1SenseT1–T2 system-of-record
2DecideT3–T4 walks · min-path confidence
Site 042 stalled 18 days — catchment density adequate; competing trial opened 3 km away on 8/21; screen-fail rate normal. Diversity plan coverage at this site is 12% of commitment.min-path 0.84 · site → competing_trial → study (CT.gov) · confidence band: NOTIFY
Stock-out risk, Study 1187, depot EU-2 — 14 days of stock; resupply lead time 19 days; 2 sites enrolling above plan.min-path 0.93 · IRT_state → lead_time → enrollment_curve · band: AUTONOMOUS (draft resupply)
KRI drift, 4 sites, protocol deviation cluster — all four share the same CRO monitoring team; deviation type: visit-window.min-path 0.78 · monitoring_finding → site → cro_team (mined edge) · band: NOTIFY
CSR 1187 §11.4 — 2 numeric-consistency findings vs locked TFL 14.2.1; drafting agent flagged unsupported claim, refused.deterministic checker · rule DOC-NUM-017 · Part 11 sign-off pending
3Actthreshold-gated NBAs
4Learndecision provenance → G2
08:14 Reallocation NBA (site 031) — accepted · rationale: "competing trial confirmed by CRA" · time-to-action 26 min
Mon Resupply draft depot US-1 — modified · quantity 180→220 · preference signal: buffer for holiday week
Mon Monitoring-visit NBA site 009 — rejected · "KRI drift explained by site staff turnover; not a data-quality issue" → suppression pattern learned
14d NBA acceptance 76% (rolling) · alert-fatigue index 0.21 ↓ · signal-to-action 4.1 h

Manufacturing Quality Command Center. Domain feeds (from §8): M2 (CPV signals) · M3 (PAT abnormal-batch, model drift) · M4 (endpoint advisories) · M5 (open investigations, CAPA effectiveness) · M1 (monthly APQR roll-up) · S1/S2 (CoA exceptions, cold-chain excursions, supplier status). Users: Site quality heads, MS&T, global quality leadership. Decisions made there: Batch disposition support, investigation prioritization, cross-site trend response, supplier holds, management review inputs (ICH Q10). Cadence: Real-time for PAT/CPV; daily quality huddle; monthly management review.

platform.pharma.internal/quality/command-centerLIVE · 22 feeds · 3 sites
Quality MS&T
  • Command Center
  • CPV Console
  • APQR
  • PAT Models
  • Investigations
  • Supply Intake
  • Decision Log
Daily quality huddle · Site 2 (Line 2 in focus)
Feeds: M2 CPV · M3 PAT · M4 endpoint · M5 investigations · M1 APQR roll-up · S1/S2 supply — Tue 07:30
1Sensehistorian · LIMS · genealogy
2Decidecross-domain correlation
Line 2 EWMA shift on dissolution correlates with resin lot L-7781 — CoA method differs from spec; deviation DEV-2026-0412 opened on the same lot yesterday; mined edge: this phase stalls 3× more often on Line 2.min-path 0.86 · batch → genealogy → material_lot → CoA → supplier · stalls_at mined edge capped 0.75 (evidence, not instruction)
PAT: batch 0419 T² excursion, contribution plot → moisture — instrument-health nominal; model PLS-BU-v3.2 approved; prediction vs reference drift 0.4% (within limit).min-path 0.90 · spectrum → pat_model → NOC model · band: NOTIFY
CAPA-2025-118 recurrence — same root-cause class (gasket wear) recurred twice in 6 months on Line 2 granulator; effectiveness check failing.deterministic rule CE-004 · QMS rows · band: NOTIFY
Cold-chain excursion, lane EU-2 → site 3 — 2 h 40 min above 8 °C; stability data supports 72 h excursion budget; 1 shipment affected.S2 monitor · lane telemetry · stability_data (ICH Q1) · band: SUGGEST
3ActNBAs with rationale · impact · owner
4Learnsignal disposition · QA sign-off
07:41 EWMA shift Line 2 — investigate · QA signed hold (Part 11) · disposition feeds APQR (M1) automatically
Mon Cpk signal, drying temp, Line 1 — benign · "seasonal HVAC; within validated range" → known-benign pattern suppressed
Mon CoA exception, lot L-7702 — accepted with edit · method equivalence documented → rule RM-007 calibration note
14d CPV acceptance 81% (n=212) · batch-to-chart latency overnight · signal-to-disposition 6.2 h · deviations pre-empted by CPV: 3

A VP Quality reads this screen without clicking: (a) the CPV signal alerting is the EWMA shift on dissolution, Line 2; (b) the recommended action is a hold on remaining L-7781 inventory pending re-test; (c) it was proposed by mfg.quality.orchestrator composing mfg.cpv.signal-agent v2.3.0; (d) the evidence is the validated SPC query, the genealogy link to six batches, the CoA method flag, and DEV-2026-0412 — with the mined edge carried as evidence, capped at 0.75.

Commercial Intelligence Command Center. Domain feeds (from §8): C3 (market signals) · C1 (field engagement outcomes, AE sidecar volume) · C2 (MLR queue and expiries) · X2 (medical inquiry themes, de-identified and firewalled). Users: Brand leads, commercial operations, compliance. Decisions made there: Campaign responses, content refresh priorities, field-force redirection, compliance interventions. Cadence: Daily; weekly brand review.

platform.pharma.internal/commercial/command-centerLIVE · 9 feeds · firewalled
Brand Intelligence
  • Command Center
  • Market Signals
  • Field Outcomes
  • MLR Queue
  • Compliance
  • Decision Log
Weekly brand review · Product X · US
Feeds: C3 market signals · C1 field outcomes + AE sidecar · C2 MLR queue · X2 inquiry themes (de-identified, firewalled) — Tue
1Sensemarket · field · content
2Decidesignals that need a response
Competitor launch, indication B, 12 overlapping territories — formulary status unchanged at 2 of 3 major payers; our approved comparative claim set is current.min-path 0.88 · signal → territory → formulary_status → mlr_asset · band: NOTIFY
Formulary tier change, payer P2 — moved to tier 3 in 4 states; affected HCPs: 1,140 (golden record); consent valid for 71%.min-path 0.90 · formulary_status → hco → hcp (golden record) → consent · band: NOTIFY
5 MLR assets expire in ≤30 days — 2 carry claims cited by 38% of last month's NBAs.deterministic · asset registry expiry · claim usage counts · band: NOTIFY
Inquiry theme rising: dosing in renal impairment — de-identified count only; content response is a Medical decision, not a commercial one.firewalled aggregate · routes to medical strategy (X3), not to the field
3ActMLR-approved content only
4Learnbrand-team decisions
Tue Field redirection NBA — accepted · rationale: "launch confirmed; claims current" · rep uptake measured from Thursday
Mon Refresh queue — modified · asset 3 deprioritized ("indication sunset in Q4") → registry usage weight updated
Fri Campaign NBA (payer P1) — rejected · "contract negotiation in progress" → signal suppressed for 30 days
14d MLR violations: 0 · AE capture by sidecar: 100% of flagged conversations · signal precision 0.87

Medical Affairs Intelligence Command Center. Domain feeds (from §8): X1 (PV signal detection, clock compliance) · X2 (medical inquiry triage, response quality, theme analysis) · X3 (field medical insights, confidence-scored themes) · X4 (MSL pre-call intelligence, brief acceptance rates) · X5 (congress intelligence, real-time session alerts) · X6 (evidence generation pipeline, RWE study status, SLR progress) · X7 (medical content review, MLR cycle time, claim substantiation). Users: Medical Affairs leadership, MSL directors, PV heads, medical information directors, evidence generation leads. Decisions made there: PV escalation and regulatory reporting, MSL deployment priorities, medical strategy inputs, congress coverage decisions, evidence gap prioritization, medical content refresh, medical/commercial firewall enforcement. Cadence: Real-time for PV clock; daily for inquiries; weekly medical strategy review; congress-driven alerts.

platform.pharma.internal/medaffairs/command-centerLIVE · 18 feeds · firewalled
Medical Affairs
  • Command Center
  • Inquiry Hub
  • MSL Insights
  • Congress Feed
  • Evidence Pipeline
  • PV Monitor
  • Decision Log
Weekly medical strategy · 4 therapeutic areas
Feeds: X1 PV signals · X2 inquiry themes · X3 insights · X4 MSL intelligence · X5 congress · X6 evidence · X7 content review — Tue 09:00
1Sensemedical landscape
2Decidemedical strategy signals
Congress abstract contradicts current dosing position, indication A — Phase III data from competitor presented at plenary; 3 KOLs in company advisory board are co-authors; medical response letter needed before media coverage cycle.min-path 0.91 · abstract → author → kol_advisory_board → company_position · band: ESCALATE
Inquiry theme rising: renal dosing adjustment, 3× baseline volume — 42 inquiries in 14 days vs. 14 baseline; approved response exists but references outdated PK data; label supplement pending.min-path 0.87 · inquiry_cluster → topic → approved_response → reference_currency · band: NOTIFY
MSL insight cluster: off-label inquiries in 3 territories — 8 MSLs reported similar HCP questions on unapproved combination; pattern corroborated by X2 inquiry data (firewalled join on de-identified topic only).min-path 0.79 · insight → topic → inquiry_theme (de-identified) · corroboration score 0.72 · band: NOTIFY
Evidence gap: no RWE for patient subgroup in indication B — X6 gap analysis shows competitor has published; advisory board recommended prioritization; study concept drafted.X6 gap analysis · competitive landscape · advisory board minutes (§5.3 KG) · band: SUGGEST
3Actapproved-source constrained
4Learnmedical decision provenance
09:22 Medical response (indication C, congress) — approved · medical director signed; distributed to MSL field within 4h of abstract publication
Mon Inquiry response update (dosing in hepatic impairment) — modified · medical reviewer added SmPC §4.2 reference · approved-source version incremented
Mon MSL deployment NBA (territory 7) — rejected · "HCP availability conflict with congress travel" → scheduling constraint learned
14d Inquiry SLA compliance 94% · MSL brief acceptance 78% (↑ from 61% at launch) · insight-to-strategy cycle 11 days · PV clock overdue: 0

Enterprise AI Operations Command Center. Domain feeds (from §8): G4 observability feeds from every factory agent — drift, eval-score drift, acceptance-rate curves, cost per decision, connector health, degraded-mode events · Registry lifecycle events · Gate queue. Users: Factory operations, platform leadership, AI governance council. Decisions made there: Agent rollback, coach-mode re-entry, budget caps, gate scheduling, incident response. Cadence: Real-time; weekly platform review; monthly steering.

platform.pharma.internal/factory/ai-opsLIVE · 31 factory agents · G4
Factory AI Ops
  • AI Ops Center
  • Agent Health
  • Registry
  • Gate Queue
  • Cost
  • Incidents
  • Steering
The network operations center for the agent estate
Observes the other four: every factory agent's traces, evals, acceptance curves, cost, connectors, degraded-mode events · registry lifecycle · gate queue
Feeds in fromClinical Operations CCManufacturing Quality CCCommercial Intelligence CCMedical Affairs Intelligence CC+ G4 observability from all 31 factory agents
1Senseagent health cards
2Decidedegraded mode · eval drift
mfg.cpv.signal-agent acceptance dropped 0.81 → 0.63 after the site-2 LIMS upgrade — test code TC-DISS renamed; conformed test_results contract v2.0 unaffected but dictionary mapping stale. Seen here before Quality sees it in an investigation.G4 + G2 · acceptance curve · data-product contract check · band: ESCALATE (below 0.70 gate threshold)
Eval drift: clin.doc.drafter F1 −3.4 vs release baseline — approved-source registry scope changed (new style guide v9); model pinned, prompt unchanged.nightly regression · model-loop · band: NOTIFY
Historian connector site 3 degraded 41 min — 3 agents auto-capped confidence to 0.55; 2 NBAs held for human; no manual intervention needed to fail safe.G7 degraded-mode design · automatic · resolved 06:52
Frontier-tier budget at 84% of monthly cap — one orchestrator routing 31% of calls to frontier (target ≤20%).G7 cost circuit-breaker · composition rule · band: NOTIFY
3Actrollback · coach re-entry · caps
4Learngate decisions · budget adjustments
Tue Gate: scm.supply-risk.orchestratorpromoted to factory · 4 lanes approved · acceptance 0.74 rising · demand baseline measured
Mon Gate: res.target.dossierheld · Quality: G6 manifest incomplete (model pinning) · escalation SLA 2 weeks
Fri PoC clin.site.chatbotkilled · acceptance flat 2 cycles · deprecated card + post-mortem published
Q Agent graduation rate 58% · killed PoCs 3 (Lab health) · agent-estate coherence 100% under contract · inspection drill: full trail in 11 min

The fifth is the one most organizations forget, and it is the one that keeps the other four honest. The Enterprise AI Operations Command Center is the network operations center for the agent estate: it is where the Factory sees that the CPV agent's acceptance rate dropped after a LIMS upgrade changed a test code, before Quality sees it in an investigation.

12.3 How command centers relate to existing control towers

Where a control tower already exists, the platform does not replace it. The tower becomes a consumer of registered agent feeds — it gains NBAs with provenance and a decision log — and its existing analytics become data products the agents read. Where none exists, the command center is instantiated from the same experience template as any other workspace (§6), with feeds subscribed from the registry. Either way, the command center is a rendering of the platform, not a separate system to be integrated.

Command centers are where the platform's value becomes visible to leadership daily. Getting there is a sequence.

Section 13 · Implementation Roadmap

18–30 months, four phases — build what everything consumes first

Total estimated timeline: 18–30 months depending on enterprise size, regulatory requirements, and the number of domain use-cases. Each phase delivers usable value; later phases compound on earlier ones. The sequencing follows the coherence matrix (§9): build what everything consumes first.

PhaseDurationWhat is builtWhat value it deliversGate to next phase
1 — Semantic & Platform Foundation3–5 months
  • Core ontology (IDMP, CDISC, MedDRA, ISA-88/95 alignment)
  • KG-as-a-Service with provenance and versioning APIs via MCP
  • Evaluation-harness framework
  • Agent platform contract (identity, MCP tool access, observability)
  • Agent registry v1 with the first two kits (draft-then-check, rule-gate)
  • Data-product catalog bound to the existing lake
  • Lab environment stood up
The Lab's first grounded, cited answer for a new use-case in days; partner-agent onboarding possible; the three-estates problem addressed before it hardensRegistry live; contract published; first PoC promoted through a defined gate
2 — First-Wave Production (Development tower)4–6 months
  • Regulatory documentation (D4) — highest ROI, most measurable
  • Trial design (D1) — direct amendment-avoidance value
  • Part 11 audit trail for regulated workflows
  • GAMP 5 / CSA validation for the first production agents
  • Approved-source registry
  • Clinical KG
  • Coach-mode metrics live
Document cycle time and amendments avoided — value the organization can see in the first quarter, validated against the hardest regulatory requirements (Part 11, GAMP 5)Two factory agents with sustained acceptance ≥70%; validation packages accepted by Quality
3 — Factory Scale & the Manufacturing Data Foundation6–9 months
  • Operational precision (D3) cross-study
  • Site selection (D2) on entity resolution
  • Commercial agents (C1) under the platform contract
  • IT-ops partner agents (H1) onboarded
  • Signal-and-NBA and evidence-dossier kits
  • Manufacturing data productsbatch_phase_summary, LIMS conformed services, genealogy/CoA — with data fabric over regional LIMS
  • Manufacturing & Quality KG
  • Clinical Operations Command Center
  • Enterprise AI Operations Command Center
Portfolio-wide study operations; all three agent estates governed as one; the data foundation that makes the manufacturing tower a configuration rather than a programCommand centers live with decision logs; reuse ratio ≥70% on new use-cases; manufacturing data products at validated release
4 — Manufacturing, Supply, Medical & Enterprise6–12 months
  • APQR (M1) parallel-run then cutover
  • CPV (M2) nightly, all CPPs, cross-site
  • PAT platform (M3) with blend endpoint (M4) as entry point
  • Deviation/CAPA (M5)
  • CoA intake (S1) and supply risk (S2)
  • Pharmacovigilance (X1) with clock enforcement
  • Medical information (X2)
  • Field Medical enablement — MSL pre-call intelligence (X4), insights (X3), congress intelligence (X5)
  • Evidence generation (X6) and medical content review (X7)
  • Product Development (P1–P3)
  • Manufacturing Quality and Commercial Intelligence Command Centers
  • Discovery use-cases (R1–R3) via the Lab pipeline
  • Full IQ/OQ/PQ for the regulated agent suite
The highest-GxP tower in production on infrastructure proven in Phases 2–3; enterprise-wide decision velocity measured on one scorecardPlatform TCO vs. standalone; agent graduation rate; inspection-readiness drill passed
Phase 1
3–5 months

Phase 1 — Semantic & Platform Foundation

What is builtCore ontology (IDMP, CDISC, MedDRA, ISA-88/95) · KG-as-a-Service via MCP · evaluation-harness framework · agent platform contract (identity, MCP tool access, observability) · registry v1 with draft-then-check and rule-gate kits · data-product catalog bound to the existing lake · Lab stood up · demand study per candidate use-caseValueFirst grounded, cited answer for a new use-case in days; partner-agent onboarding possible; the three-estates problem addressed before it hardens
Gate to nextRegistry live; contract published; first PoC promoted through a defined gate
Phase 2
4–6 months

Phase 2 — First-Wave Production (Development tower)

What is builtRegulatory documentation (D4) — highest ROI, most measurable · trial design (D1) — amendment avoidance · Part 11 audit trail · GAMP 5 / CSA validation for the first production agents · approved-source registry · Clinical KG · coach-mode metrics liveValueDocument cycle time and amendments avoided — value the organization can see in the first quarter, validated against the hardest regulatory requirements
Gate to nextTwo factory agents with sustained acceptance ≥70%; validation packages accepted by Quality
Phase 3
6–9 months

Phase 3 — Factory Scale & the Manufacturing Data Foundation

What is builtOps precision (D3) cross-study · site selection (D2) on entity resolution · commercial agents (C1) under the contract · IT-ops partner agents (H1) · signal-and-NBA and evidence-dossier kits · manufacturing data productsbatch_phase_summary, LIMS conformed services, genealogy/CoA — with data fabric over regional LIMS · Manufacturing & Quality KG · Clinical Ops and Enterprise AI Ops command centersValuePortfolio-wide study operations; all three agent estates governed as one; the data foundation that makes manufacturing a configuration, not a program
Gate to nextCommand centers live with decision logs; reuse ratio ≥70% on new use-cases; manufacturing data products at validated release
Phase 4
6–12 months

Phase 4 — Manufacturing, Supply, Medical & Enterprise

What is builtAPQR (M1) parallel-run then cutover · CPV (M2) nightly, all CPPs, cross-site · PAT platform (M3) with blend endpoint (M4) as entry point · deviation/CAPA (M5) · CoA intake (S1), supply risk (S2) · PV (X1) with clock enforcement · medical information (X2) · field medical enablement (X3–X5), evidence generation (X6), medical content review (X7) · Product Development (P1–P3) · Manufacturing Quality and Commercial Intelligence command centers · discovery (R1–R3) via the Lab · full IQ/OQ/PQ for the regulated suiteValueThe highest-GxP tower in production on infrastructure proven in Phases 2–3; enterprise-wide decision velocity on one scorecard
GatePlatform TCO vs standalone; agent graduation rate; inspection-readiness drill passed

Why a demand study precedes the first PoC

The first measurement in Phase 1 is not an eval score; it is a demand study. For each candidate use-case, the pod samples the current work — a month of protocol-design requests, a quarter of deviation investigations, one APQR cycle — and classifies every touchpoint: value demand or failure demand, and for the latter, which system condition caused it. The study takes days, costs almost nothing, and does three things no architecture can. It gives the Lab a baseline the value-realization claims in §18 are measured against rather than estimated. It ranks use-cases by eliminable failure demand rather than by sponsor enthusiasm. And it surfaces the cases where the cheapest intervention is not an agent at all — a data-product owner, a template field, an SOP step. Instrument demand before automating it; the platform's credibility rests on being able to show a before.

Why documentation and trial design first

D4 has the highest volume, the most measurable cycle-time impact, and the most immediate business case. D1 has the highest per-incident cost avoidance (~$500K per amendment). Starting with these two validates the platform against the hardest regulatory requirements while delivering value the organization can see in the first quarter. Everything later builds on proven infrastructure.

Why Medical Affairs can be the first-wave tower instead

Where the sponsoring organization is Medical Affairs rather than Development, the same Phase 2 logic holds with different use-cases: medical information (X2) has the volume and the measurability that D4 has, and MSL pre-call intelligence (X4) has the adoption signal that D1 lacks — an MSL either uses the brief on Thursday or does not. The approved-source registry, the draft-then-check and rule-gate kits, and the Part 11 audit trail are the same Phase 2 deliverables; the Clinical KG is replaced by the Medical & Safety KG and the canonical interaction model (§8.7). The gate criteria do not change. What changes is who holds the validation vote — Medical Compliance and Legal rather than Quality — and the fact that the tower is in the middle of a CRM transition, which is the strongest argument for building the interaction model in Phase 1 rather than binding the first agent to CRM objects.

Why manufacturing's data foundation is in Phase 3, not Phase 4

The LIMS conformed services and the canonical attribute dictionary are the single most underestimated line item in a manufacturing AI program (§5.4). Starting them in Phase 3 — while the Development tower is scaling — means the Manufacturing tower in Phase 4 is a configuration exercise rather than a data-engineering campaign. Organizations that have done this report that the first product-site takes the longest and every subsequent one is materially faster.

Why manufacturing is the first global-by-default agent

Manufacturing and quality use-cases — CPV, APQR, deviation investigation, tech-transfer knowledge packages — are the correct first Factory deployment for a global-by-default agent. The reason is residency: manufacturing data (batch records, test results, process parameters, deviations) is not patient data and carries no PII classification. A CPV signal-detection agent can access batch-phase summaries from every site without triggering a de-identification gate, a consent check, or a cross-border transfer assessment. That makes it the fastest path to proving the platform's cross-site value — and the correct example when asked to show an agent that scales to every site on Day 1.

Why blend endpoint is the PAT entry point

The first use-case pays for the foundation: advisory mode needs no submission, the science is twenty years proven, and capacity value begins the day the advisory goes live. The enterprise PAT platform is then already built when the second and third PAT use-cases arrive.

The roadmap is pharma-specific in its regulatory sequencing. The patterns are not. But the roadmap does not exist in a vacuum — the enterprise already has a technology estate, vendor relationships, and in many cases competing agent programs. Where the platform sits in that landscape is the next question.

Section 14 · Market Context

Where this platform sits — not "instead of," but "the layer that makes them coherent"

An enterprise agentic AI platform is not built on green field. It is built on top of, beside, and sometimes in tension with an existing technology estate — cloud platforms, pharma-vertical vendors, enterprise SaaS, and increasingly, those vendors' own agentic offerings. The platform's competitive positioning is not "instead of" any of these. It is "the semantic substrate and governance plane that makes all of them coherent."

14.1 Horizontal AI and data platforms

Cloud-native AI platforms & enterprise data platforms

Cloud-native AI platforms (AWS Bedrock, Azure AI, Google Vertex AI) provide model hosting, orchestration primitives, and increasingly agent frameworks. They are the compute layer this platform runs on — not a competitor to it. The platform is model-agnostic by design (§5.2); scoped models and frontier models may come from different providers. What the cloud vendors do not provide is the pharma semantic layer, the GxP governance plane, the domain-specific rule engines, or the federated ontology. An organization that builds agents directly on a cloud AI platform without a shared substrate will have fast agents and no coherence.

Enterprise data platforms (Snowflake, Databricks, Palantir Foundry) provide the data lake, the compute, and increasingly the data-product layer. The platform's Data Layer (§5.4) sits on top of these investments — it consumes their conformed data products and data fabric services, it does not replace them. Palantir Foundry's Ontology is the closest analog to the semantic layer described here; the difference is that Foundry's ontology is tied to Foundry's compute and is not designed for the federated, domain-owned model that pharma's cross-functional structure requires. The platform described here can sit on Foundry, on Databricks, on Snowflake, or on a bespoke AWS lake — the semantic layer is the abstraction.

14.2 Pharma-vertical vendors

Clinical, manufacturing/quality, and commercial platforms

Clinical platforms (Veeva Vault, Medidata Rave, IQVIA Orchestrated Clinical Trials) are the systems of record for clinical operations, regulatory documents, and safety. They will remain systems of record. The platform connects to them as data sources and writes decisions back to them as workflow actions — never around them. When these vendors ship their own agents (and they are), the platform contract (UC-H1) is how those agents are governed alongside internal agents. The alternative is three agent estates with three governance models — which is the problem §2 described.

Manufacturing and quality platforms (Veeva QMS, MasterControl, SAP QM, Honeywell/GE historians) are the transactional systems for deviations, CAPAs, batch records, and process data. The platform's manufacturing tower (§8.4) reads from data products derived from these systems and writes investigation decisions back to the QMS. The canonical attribute dictionary (§5.4) is the bridge between regional instances and site-specific configurations.

Commercial platforms (Veeva CRM, Salesforce Health Cloud, IQVIA) increasingly have their own agent strategies. Salesforce Agentforce and Veeva's agentic roadmap will deliver field-force and commercial agents. The platform's value is the semantic layer underneath (HCP/HCO golden record, MLR content provenance, the medical/commercial firewall) and the governance plane that ensures partner-built agents are governed the same way as internal ones.

Medical Affairs platforms are where the vendor agent layer is arriving fastest, and where the transition state is most acute. Three facts shape the architecture. First, Veeva CRM reaches end of support on December 31, 2029, and every Veeva CRM customer must choose between Vault CRM and Salesforce Life Sciences Cloud — many will dual-run, with medical and field-MSL workflows on one and commercial engagement on the other. Agentforce Life Sciences (GA October 2025, with top-20 pharma early adopters) is Salesforce's agent layer for that estate. Second, Veeva AI Agents for PromoMats (December 2025 — a Quick Check Agent for pre-review and a Content Agent for drafting) and specialist MLR-automation vendors such as Falcon MLR (Copli, acquired by Veeva in June 2026; targeting roughly 70% reduction in MLR review labor) mean content-review triage is becoming a buy, not a build; the enterprise's differentiated work moves up to claim substantiation, scientific-exchange boundaries, and cross-asset consistency (UC-X7). Third, Vault MedComms remains the medical content system of record and the source for any medical approved-source registry. The platform's position across all three is the one §8.7 describes: a canonical Medical Affairs interaction model that makes the CRM decision a connector swap, a medical registry and firewall the vendor agents run inside of, and a telemetry plane that shows what each vendor agent costs per decision alongside the enterprise's own.

14.3 Emerging agentic AI competitors

LLM-native agent platforms & process intelligence

LLM-native agent platforms (OpenAI's agent offerings, Anthropic's MCP ecosystem, Google's Gemini agents, Microsoft Copilot Studio) provide the model and the agent primitives. The platform described here is built on MCP (Model Context Protocol) — agents call tools through MCP servers, the registry exposes discovery as an API, the platform contract is an MCP contract. The distinction is that MCP is a protocol; the platform is the pharma-specific implementation — the ontologies, the rule engines, the validation strategy, the governance plane, the eval harness, the domain-specific kits.

Process intelligence platforms (Celonis, Signavio) provide event-log mining and process discovery. The platform's process-intelligence layer (§5.3) fuses mined temporal and statistical findings into the KG as confidence-scored edges. Celonis-type tools are a data source for the semantic layer, not a competitor to the platform. Where they compete is on the "command center" concept — Celonis's Execution Management System is architecturally similar to the agentic command centers in §12, but without the multi-agent composition, the dual-path reasoning, or the GxP governance plane.

14.4 What the platform uniquely provides

No single vendor — horizontal, vertical, or agentic — provides all five of these together:

CapabilityWho else provides itWhat the platform adds
Pharma-specific federated ontology with domain-owned extensionsPalantir Foundry (single-vendor ontology)Federated model; domain teams own extensions under a shared standard; aligned to IDMP/CDISC/MedDRA/ISA-88
GxP-grade governance plane (Part 11, GAMP 5, CSA, ALCOA+)Veeva QMS (for QMS workflows); none GxP-scopedGovernance scoped to the decision, not the system; eight categories with named owners; agent-as-identity-principal
Composition-first agent framework with registry and service kitsMCP + A2A (Linux Foundation AAIF; protocol + discovery); cloud agent platforms (AWS Bedrock AgentCore GA Oct 2025; Microsoft Agent 365 GA May 2026; Entra Agent ID)≤5-tool composition rule; registry with lifecycle states; five templatized kits; Lab-to-Factory promotion gate
Expert-calibrated deterministic rule engines alongside LLM pathsPoint solutions (Doscierge, specific vendors per domain)Cross-domain rule framework; dual-path reasoning; the checker-as-MCP-tool pattern that any workflow can call
Cross-functional decision-velocity measurementNone at this scopeOne scorecard from target ID to batch release to field call, measured on the same infrastructure (§18)

The platform's competitive position is not "better than Veeva" or "better than Palantir." It is: the integration layer, the governance contract, and the semantic substrate that makes the existing estate coherent — and makes use-case N+1 cheaper than use-case N.

Action for CDOs and platform leaders

Map your current vendor estate to the four-layer stack (§5). Identify which layer each vendor occupies. The gaps — typically the semantic layer, the governance plane, and the registry — are the platform's scope. The vendors are the platform's data sources and experience surfaces. Frame the business case as coherence cost avoided, not vendor replacement.

Section 15 · Organizational Design

The team shape by phase — how many, which roles, reporting to whom

The architecture (§5), the operating model (§7), and the roadmap (§13) define what gets built and how. This section answers the organizational question an ED building a greenfield division must answer: how many people, in which roles, at which phase — and what does the reporting structure look like?

The numbers below are calibrated for a top-20 pharma with 3-5 active therapeutic areas, 20-40 manufacturing sites, and an existing data-platform team. Scale down for mid-size; scale up for a company with more than 50 sites or a broader pipeline.

15.1 Headcount model by phase

PhaseDurationTotal headcount (new hires + allocated)CompositionKey hires
1 — Foundation3-5 months~15-20Platform core (6-8): architect, 2-3 platform engineers, 1-2 knowledge engineers, 1 eval/quality engineer, 1 DevOps/SRE. First Lab pod (4-6): 1 FDE, 1-2 agent engineers, 1 domain KE, 1-2 domain SMEs (allocated, not hired). CSV bench (2-3): 1 CSV lead, 1-2 CSV engineers (allocated from Quality or hired).Platform architect (the ED). Lead knowledge engineer. First FDE (the person who will lead the first pod and train the next FDE).
2 — First wave4-6 months~30-40Platform core grows to 10-12 (+ registry, observability, data-product binding). Second Lab pod (4-6). Factory nucleus (3-5): Factory lead, 1-2 SRE/ops, 1 release engineer. CSV grows to 4-5. Domain SME allocation doubles.Factory lead (reports to ED or is the ED for Agentic Factory). Second FDE. Data-product engineer (the person who owns the canonical attribute dictionary).
3 — Factory scale6-9 months~50-703-4 active Lab pods (12-18 pod members). Factory operations team (8-12): SRE, observability, support tiers, model lifecycle. Semantic & KE team (6-8): core ontologists, domain KEs allocated per pod, entity-resolution engineer. Data engineering (4-6): data-product owners for manufacturing data products, data-fabric engineer. CSV (5-7). Command-center development (2-3).Head of Semantic & KE (reports to ED or is the ED for that role). Command-center product lead. Manufacturing data-product owners (2-3 hires or allocations from MS&T).
4 — Enterprise6-12 months~80-120 (steady state)4-6 active Lab pods rotating. Factory at steady state (15-20: ops, support, release, model lifecycle). Semantic & KE (8-12). Data engineering (6-10). CSV (6-8). Platform core (12-15). Domain champion network (~30 SMEs, allocated not hired).Additional FDEs. Regional data-fabric engineers (for multi-site, multi-geography). PV and commercial domain KEs.

Where the effort actually goes

A common miscalculation in pod sizing: estimating engineering effort as if the work is mostly agent logic — prompt engineering, tool binding, evaluation. In practice, 40–60% of Year 1 engineering effort goes to data governance infrastructure: classification tagging on every data product, de-identification pipelines per region, consent verification services, dual-CRM connector maintenance during vendor transitions, and the canonical attribute dictionary that unifies regional LIMS instances. The agent itself — the part that reasons — is the smaller fraction. Teams that staff for agent engineering alone discover the governance gap in month three, when the agent works in the Lab sandbox and cannot access production data. Budget for governance infrastructure explicitly; it is the binding constraint, not the model.

15.2 Reporting structure

Two models work; a third does not.

Model A — Under the Chief Data & AI Officer.

The platform reports to the CDAO alongside the data-platform team and the existing analytics org. Advantage: natural alignment with the data layer and existing data-product owners. Risk: the platform is perceived as "IT" and domain adoption is slower.

Model B — As a peer to the CDAO, under the CTO or CEO of R&D.

A new division — Lab + Factory + Semantic & KE — with its own ED-level leads. Advantage: signal that this is a strategic investment, not an IT project; domain teams see it as a peer. Risk: requires executive sponsorship to avoid turf conflict with the existing data org.

Model C (anti-pattern) — Distributed across functions.

Each function (R&D, Manufacturing, Commercial) builds its own agents with a loose "center of excellence" for coordination. This is how three agent estates happen, and it is how the $2-5M-per-snowflake cost structure persists. The platform exists to prevent this.

The four ED roles that best match this architecture:

RoleOwnsReports toKey relationship
ED, Platform Strategy & InnovationPortfolio prioritization, roadmap, investment model, steering committee, platform contract, partner-agent onboardingCDAO or CTOOwns the coherence matrix; accountable for reuse ratio and marginal cost per use-case
ED, Agentic Lab & ArchitectureLab environment, kits, agent composition rules, FDE bench, PoC pipeline, kill decisionsPlatform Strategy ED or CDAOOwns time-to-first-grounded-answer; accountable for PoC throughput and kill rate
ED, Agentic FactoryFactory operations, promotion gate, support model (L1/L2/L3), release trains, model lifecycle, command centersPlatform Strategy ED or CDAOOwns agent graduation rate; accountable for production uptime, drift detection, and cost per decision
ED, Semantic & Knowledge EngineeringCore ontology, domain KG federation, entity resolution, KG-as-a-Service, data-product contracts, process intelligencePlatform Strategy ED or CDAOOwns the semantic layer's freshness, coverage, and provenance completeness; accountable for time-to-ontology-extension

15.3 The scarce roles

Three roles are the binding constraint on hiring, and the platform design is intentionally shaped to reduce the demand for them:

Forward-Deployed Engineers

Engineers who speak the domain. The FDE model (§7.1) means you need fewer of them than you would need project managers, because one FDE carries context in both directions. But you need some, and they take 6-12 months to develop. Invest in the first two hires; train the next cohort through pod apprenticeship.

Knowledge Engineers with pharma domain context

Ontologists who understand CDISC, ISA-88, MedDRA, and the modeling standard. The core team is 3-5; domain KEs are allocated per pod and trained through the modeling standard, not through a PhD.

CSV Engineers who understand hybrid AI

Quality professionals who can validate a system where the deterministic path is GAMP 5 Cat 5 and the LLM path is CSA-controlled. This skillset barely exists in the market. The mitigation is joint workshops (§7.6) and doing the first validation package together with the platform architecture team, so the template exists for the next ten.

Action for ED candidates and division builders

The first three hires define the culture of the division: the platform architect, the lead knowledge engineer, and the first FDE. Get those right and the pods will form correctly. Get them wrong and every subsequent hire inherits the wrong patterns.

The organizational design is specific to pharma. The architecture underneath it is not.

Section 16 · Beyond Life Sciences

Pattern portability — the same architecture, run in a second industry

The architecture described above — the enterprise knowledge graph, MCP connectors, confidence-gated autonomy, coach-to-production graduation, ALCOA+ audit trails, dual-path reasoning, the registry — is not pharma-specific. The same patterns have been built for mid-market manufacturing (quote-to-cash, procure-to-pay) on the same stack:

PatternPharma instantiation (this article)Mid-market manufacturing instantiation (production)
Enterprise knowledge graphFederated domain KGs on a thin core; trust-tiered provenance1,786 triplets, 47 entity types, 50 predicates, 6 trust tiers with a T4 hard cap at 0.75; same 13-field triplet schema, same data-quality gates
Graduated autonomyCoach mode → supervised → production, domain-expert-controlledCoach mode → production on rolling-window acceptance rates; thresholds ≥0.90 autonomous, 0.70–0.89 notify, 0.55–0.69 suggest, <0.55 escalate; test-driven design with regression suites at every gate
Decision velocityStateful orchestration with execution history; command centers11-state Step Functions pipeline; MCP connectors to ERP/CRM/email; real-time retrieval with no data warehouse; iterative build-measure-learn cycles to production
GovernancePart 11 / Annex 11 / GAMP 5 / CSAALCOA+ audit trail on every decision; same serverless substrate

Why this matters for pharma. Cross-industry portability is how you know an architecture is real and not a framework slide. The same EKG substrate, the same confidence-gated autonomy, the same audit trail — configured for a different domain with different entity types and different compliance requirements — proves the "one platform, many use-cases" thesis extends beyond a single vertical. A pharma platform team that adopts these patterns is adopting something that has been run end-to-end, twice, in two industries.

And within pharma, two production systems are the wedge.

Section 17 · Wedge Applications

Built, not theorized

Architectures are easy to draw. The question is always: does it work? These two production systems were built on the platform described above — same semantic layer, same agent framework, same governance plane, same dual-path reasoning. They are the first two use-cases. Everything that follows inherits what they proved.

ProtoCheck · maps to UC-D1

ProtoCheck — Clinical Trial Protocol Compliance

Maps to UC-D1 (Trial Design) and is the pattern behind every rule-gate in this article. A multi-agent architecture where four specialized reviewers — safety, efficacy, regulatory, operational — each apply expert-calibrated rules against a protocol document, and a coordinator merges findings with provenance.

500+Expert-calibrated rules across 6+ therapeutic areas; FDA, EMA, ICH E6/E8, WHO4Specialist agents, each ≤5 tools — composition over inheritance, in productionCFR / ICHCitation traceability on every findingPatent pendingMulti-agent compliance system; presented at BioNJ 2026

Platform patterns proven. Composition over inheritance (four bounded specialists); confidence-gated routing; human-in-the-loop mandatory; provisional → confirmed rule calibration with named SME signatures; ALCOA+ audit trail on every finding.

Doscierge · maps to UC-D4

Doscierge — Regulatory Documentation Compliance Engine

Maps to UC-D4 (Regulatory Documentation) and is the pattern behind every draft-then-check in this article. The insight that drove it: everyone accelerates draft generation (80–85%); nobody deterministically checks against the rulebook. Generate-then-check is the architecture. The checker is an MCP tool — deterministic, inspectable, GAMP 5 Category 5.

970Deterministic rules10,292KG triplets across 29 domains — stability, impurities (ICH Q3A), DMF, dissolution, and more11-stateStateful processing pipeline, 25 functions, full execution historyP 92.31 / R 80.27 / F1 85.87Evaluation harness against expert-labeled ground truthMCPChecker-as-tool pattern — the deliberate Category 5 ceiling

Platform patterns proven. Dual-path reasoning (deterministic rules + KG-grounded LLM); checker-as-MCP-tool; KG-first pipeline with run-time grounding; confidence-scored provenance on every finding; trust-tiered triplet schema.

The enterprise precedent

Before either wedge application, the same patterns ran at a Fortune 50 pharma: a real-time analytics platform over 14B+ data points; a 4-agent GenAI platform scaled from zero to 150+ regulated users; a Customer MDM of 100K+ HCP/HCO records integrated with CRM and 20+ systems; clinical data integration across 100+ studies on SDTM/ADaM; a semantic layer that lifted retrieval accuracy from 75% to 99%+; and a 21 CFR Part 11 / Annex 11 MLOps framework. The wedge applications are what those patterns look like when built from scratch, on a serverless substrate, with a registry and a governance plane designed in from day one.

Why wedge applications matter. A platform without applications is a strategy deck. ProtoCheck and Doscierge prove the thesis: two production systems, different domains, built on the same substrate. Use-case #3 inherits the semantic layer, the governance plane, the agent framework, the evaluation harness, and the audit trail. The marginal cost drops 40–60%. That is the platform economics working — not projected, observed.

Observed how? By measuring the right things at the right time.

Section 18 · Success Metrics

Measure the right things at the right time

Most agentic AI programs fail because they measure the wrong things at the wrong time. Month-one ROI on an LLM agent is meaningless. What matters is a staged measurement framework that tells you whether the approach needs tweaking before you scale it — and that puts decision velocity on the scorecard beside accuracy and compliance, because velocity is the value.

Phase 1 — Coach Mode (Months 0–3): Leading indicators

MetricWhat it tells youTarget
Agent acceptance rate% of agent suggestions the human actually uses≥70% sustained to graduate
Confidence calibrationIs the agent right when it says it is confident? (reliability diagram, Brier score)Calibration error falling
Time-to-first-valueDays from deployment to first expert-accepted outputDays, not weeks
Escalation qualityDoes the agent escalate the right things, or cry wolf?Escalation precision rising
Time-to-first-grounded-answer (Lab)Days from a new use-case request to a cited answer on the platformDays
Failure-demand ratio (baseline)Of the demand the use-case serves, what share is caused by the system's own prior failures — measured by classifying every touchpoint, not estimatedMeasured within 30 days; the number Phase 2 claims are tested against

Phase 2 — Supervised Autonomy (Months 3–9): Efficiency metrics

MetricWhat it tells youTarget
Decision cycle-time reductionProtocol design, deviation investigation, APQR data package, batch release review, case processing40–60%
Signal latencyBatch-to-CPV-chart; trend-to-APQR-finding; risk-to-study-team-actionWeeks/months → overnight/real-time
Error-catch rateFindings the agent surfaces that humans missed — the real value signalRising, per class
Reuse ratio% of shared platform components consumed by new use-cases vs. custom-built≥70% is platform health
Marginal cost per use-caseIs each additional use-case cheaper than the last?Falling 40–60% by use-case 5–6
Cost per decisionTokens + compute per accepted recommendation, by agent (from the registry)Falling as scoped models take share
Failure demand eliminatedShare of baseline failure demand whose cause the agent removed (evidence now structured, precedent now retrievable) vs. merely absorbed fasterRising by class; cause-removed ≥ absorbed

Phase 3 — Production Scale (Months 9–18): Outcome metrics

MetricWhat it tells youTarget
Inspection readinessCan you produce the full audit trail for any agent decision in under 15 minutes?Yes, drilled quarterly
Decision velocity, pipeline-wideClinical: months off enrollment; quality: days off batch release, months off signal latency; medical: clock complianceMeasured per tower on one scorecard
Agent graduation rate% of coach-mode agents promoted to supervised autonomy within 6 monthsRising; killed PoCs counted as Lab health
Agent-estate coherence% of agents — internal, CRM vendor, IT partner — under the platform contract100%
Platform TCOTotal cost vs. N standalone implementations at $2–5M eachDocumented dividend

Measuring coach mode: the critical leading indicator.

Coach mode is where you learn whether an agent is actually helpful or just generating plausible-looking output. Track the acceptance-rate curve over rolling two-week windows. A healthy agent shows a monotonically increasing curve as it learns domain boundaries. A flat or declining curve means the approach needs recalibration — retune the system prompt, add constraints, narrow the tool set, or revisit the model tier. The promotion gate from coach to supervised autonomy requires a sustained acceptance rate above threshold, verified by the domain expert team — not by the AI engineering team. This is decision provenance (G2) doing its job.

Why the failure-demand ratio keeps the other metrics honest.

Cycle-time reduction can be achieved by helping people cope faster with a broken system; it says nothing about whether the system is less broken. The failure-demand ratio does. If deviation investigations close 50% faster but the share of investigations caused by unseen drift is unchanged, the platform has industrialized the waste. If the ratio falls — because CPV now catches the drift before the batch fails — the platform has changed the system condition, and the cycle-time gain will hold even when the agent is switched off. The service-sector studies behind the method put failure demand at 25–75% of inbound demand in unexamined systems; pharma's evidence-assembly work exhibits every generator on that list.

Every metric above depends on people using the system. That is the last problem, and the hardest.

Section 19 · Enterprise Adoption & Change Management

The platform architecture is the easy part

The hard part is getting hundreds of scientists, medical writers, process engineers, quality leads, and regulatory specialists to trust an agent enough to integrate it into how they actually work. This requires a deliberate adoption strategy — not a training deck and a launch email. It is the people in people-process-technology, and §7's operating model only works if this section is taken as seriously as the architecture.

The human side
  • Champion network. Identify 2–3 domain experts per use-case who co-design the agent's behavior. They are the internal evangelists — not IT, not the platform team. When a process scientist sees a colleague act on a CPV signal from the agent, that outweighs any executive mandate. They are also the L1 support tier (§7.4).
  • Transparent limitations. Publish what the agent cannot do alongside what it can. Teams trust tools that are honest about their boundaries. "This agent does not replace your judgment on eligibility criteria" builds more adoption than "AI-powered protocol optimization." The agent card (§6) is the honest source.
  • Feedback loops. Every agent in coach mode has a thumbs-up/thumbs-down on every output. Aggregate weekly. Share the dashboard with users — they see their feedback changing the system. This is the single most effective adoption lever, and it is the same mechanism as decision provenance (G2).
  • Graduated trust. Start with "agent prepares a draft, human does everything else." Move to "agent prepares and routes, human reviews." Then "agent executes routine actions, human reviews exceptions." Never skip a step. The blend-endpoint use-case (§8.4) is this principle applied to a process step: advisory beside the validated method, then replacement, then — only if chosen — real-time release.
The organizational side
  • Governance alignment before exposure. Quality, Regulatory, IT Security, and Legal review the agent's operating envelope before users see it. Retroactive governance kills adoption — teams hear "legal pulled the plug" once and never come back. The registry approval workflow (§6) makes this a step, not a negotiation.
  • Operating model clarity. Lab → Factory means exploration is funded differently from production. Lab pods have permission to fail fast and kill criteria to fail against. Factory agents have SLAs and on-call. Mixing the two creates confusion about what "done" means.
  • Skill development. Train domain experts to be agent supervisors, not AI engineers (§7.6). They need to read a confidence score, review a provenance trail, and know when to override — not tune hyperparameters. Two days, not six months.
  • Executive cadence. Monthly steering with the three metrics that matter at each phase (§18). No vanity metrics, no "number of agents deployed." The question is always: are humans making better decisions faster with the system than without it?

The anti-pattern: "We built it, they didn't come"

The number-one failure mode in enterprise AI is a technically excellent system nobody uses. Every use-case in this platform includes a co-design phase with domain experts, a coach-mode period with measurable acceptance rates, and a promotion gate that domain experts control — not engineers, not executives. If the medical writers say the documentation agent is not ready, it stays in coach mode. If the process scientists say the CPV agent cries wolf, it goes back to the pod. Full stop. Adoption is earned through demonstrated reliability, not mandated through organizational authority.

Section 20 · Evidence & Provenance

Three evidence sources — distinguished, so the architecture is credible rather than aspirational

This article's claims rest on three distinct evidence sources. Distinguishing them is what makes the architecture credible rather than aspirational.

Built and operated at enterprise scale · Fortune 50 pharma, 2010–2025

What was built and operated at enterprise scale (Fortune 50 pharma, 2010-2025)

These are patterns validated in a 15-year career at a Fortune 50 pharmaceutical company with $45B+ revenue, 30,000+ employees, and active FDA/EMA regulatory engagement:

  • Real-time analytics platform over 14B+ data points — the data-layer and semantic-layer patterns in §5.3 and §5.4, operating across clinical, commercial, and manufacturing data at enterprise scale
  • 4-agent GenAI platform scaled from zero to 150+ regulated users — the agentic-layer patterns in §5.2, including composition, evaluation, and graduated deployment under GxP constraints
  • Customer MDM of 100K+ HCP/HCO records integrated with CRM and 20+ systems — the entity-resolution and golden-record discipline described in §5.3 and applied to site selection (UC-D2) and commercial (UC-C1)
  • Clinical data integration across 100+ studies on SDTM/ADaM — the clinical KG and data-product patterns in §8.2
  • Semantic layer that lifted retrieval accuracy from 75% to 99%+ — the KG-as-a-Service architecture in §5.3
  • 21 CFR Part 11 / EU Annex 11 MLOps framework — the governance-plane controls in §5.5 (G3, G6)

These are the reference experiences. They prove the patterns work at scale, in regulated environments, with real inspectors. They are not current products — they are the institutional knowledge the architecture is built on.

Re-implemented from scratch · OrchestraPrime, 2024–2026

What was re-implemented from scratch on a modern substrate (OrchestraPrime, 2024-2026)

These are production systems built on the serverless, registry-first, governance-by-design substrate described in this article:

  • ProtoCheck — 500+ expert-calibrated rules, multi-agent architecture (4 specialist reviewers, ≤5 tools each), human-in-the-loop mandatory, patent-pending, presented at BioNJ 2026. Maps to UC-D1. Proves: composition over inheritance, expert-calibrated rule engines, confidence-gated routing.
  • Doscierge — 970 deterministic rules, 10,292 KG triplets across 29 domains, 11-state stateful pipeline, 25 Lambda functions, measured eval harness (P=92.31%, R=80.27%, F1=85.87%), checker-as-MCP-tool, designed for GAMP 5 Category 5 validation, ALCOA+ audit trail. Maps to UC-D4. Proves: dual-path reasoning, KG-first pipeline, run-time grounding, trust-tiered provenance. Desirability score: 2/10 — technically validated, commercially unvalidated. This is an honest assessment, and it is why the platform thesis matters: a single-use-case product must find its own market; a platform component inherits value from the portfolio.
  • Mid-market manufacturing agentic platform — 1,786 KG triplets, 47 entity types, 50 predicates, 6 trust tiers, 11-state pipeline, MCP connectors to ERP/CRM/email, test-driven iterative development to production. Proves: cross-industry pattern portability.
  • MCP server on the MCP registry (Linux Foundation AAIF) — the interoperability contract. 18+ reusable IaC modules underlying 8 deployable system configurations.
Proposed · this article

What is proposed (this article)

The full cross-functional platform described in §5-§17 — the assembly of the above patterns into one governed, semantic-first operating system for AI-augmented decisions across the enterprise. The individual components have been proven; the assembly at enterprise scale has not. That is the gap between an architecture article and a shipped platform, and it is what Phases 1-4 of the roadmap (§13) are designed to close. The bet is that the components compose — and the evidence that they do is that two of them (ProtoCheck and Doscierge) already share the same semantic layer, the same governance plane, and the same evaluation harness, with a measured 40-60% marginal cost reduction on the second system.

What we got wrong, and what it taught us.

  • The first knowledge graph had 15,000+ triplets, and retrieval precision was 75%. The rebuild — curated, trust-tiered, 10,292 triplets — achieved 99%+ on a 212-question benchmark against expert-labeled ground truth. The lesson: for regulatory knowledge, curation beats scale. Every triplet should earn its place.
  • A 4-month spike tried to fine-tune a 7B/14B parameter model for regulatory compliance review. Best result: 0.597 on the detection criterion — all gates failed. The same KG and retrieval pipeline, connected to a frontier model, achieved 1.000. The lesson: don't spend months proving a small model can do what a frontier model does in one call. The methodology (structured retrieval, expert-calibrated rules) was right; the model capacity was the bottleneck.
  • Doscierge's first evaluation showed ~288 findings per test package. After iterative precision engineering cycles (enforcement gates, applicability predicates, correlator logic), it dropped to ~73 with higher precision. The lesson: the eval harness is not a report card — it is the engineering tool. A test-driven design approach — where every rule change must pass regression before merge — turns the eval harness into the primary development feedback loop. Without it, you are tuning by feel.
Section 21 · Contact

The architecture is defined. The patterns are proven at different scales. The question is yours.

This article exists to help you think through the architecture, the operating model, and the organizational design for an enterprise agentic AI platform — not to sell you one. The patterns described here were validated across three contexts:

  • At enterprise scale (Fortune 50 pharma, 15 years): Real-time analytics platform over 14B+ data points. A 4-agent GenAI platform scaled from zero to 150+ regulated users. Customer MDM of 100K+ HCP/HCO records integrated with CRM and 20+ systems. Clinical data integration across 100+ studies on SDTM/ADaM. A semantic layer that lifted retrieval accuracy from 75% to 99%+. A 21 CFR Part 11 / Annex 11 MLOps framework. These are the patterns — the ones that work at scale, in regulated environments, with real inspectors.
  • Re-implemented from scratch on a modern serverless substrate (OrchestraPrime): ProtoCheck (500+ expert-calibrated rules, multi-agent, patent-pending, presented at BioNJ 2026). Doscierge (970 deterministic rules, 10,292 KG triplets, 29 domains, 11-state pipeline, measured eval harness at P=92.31% / R=80.27% / F1=85.87%). A second full agentic platform in mid-market manufacturing on the same EKG substrate. MCP server on the MCP registry (Linux Foundation AAIF). 18+ reusable IaC modules. These prove the patterns compose on current infrastructure.
  • Proposed as an integrated enterprise architecture (this article): The full cross-functional platform described here, whose components have been proven individually but whose value is in their assembly as one governed, semantic-first operating system for AI-augmented decisions across the enterprise.

The honest assessment: the hardest parts of this program are not technical. They are organizational — hiring the right FDEs and knowledge engineers, earning Quality's trust before the first agent touches a GxP workflow, and resisting the pressure to scale before the foundation is proven. Those are leadership problems, and they are the ones this article is least able to solve on a page.

A 30-minute architecture conversation.

Not a sales call — a structured walkthrough of how this maps to your use-case portfolio, tower by tower, and where your registry and your first pod would start.

Start the conversation on LinkedIn →
Nitin Bhatti
Nitin Bhatti
Founder & Platform Architect, OrchestraPrime
26 yrs enterprise technology · Fortune 50 pharma · TOGAF · PMP · MBA · Patent holder
linkedin.com/in/nitinbhatti · orchestraprime.ai
North Brunswick, NJ
Contact QR codeLinkedIn QR code
© 2026 OrchestraPrime LLC. All rights reserved. Patent pending. · Enterprise Agentic AI Platform for Life Sciences — Solution Architecture, v2.1, September 2026 · Illustrative command-center data and scenario values are synthetic and shown for pattern demonstration.