Innovation Insight × Reference Architecture · Life Sciences
Enterprise Agentic AI Platform for Life Sciences— Solution Architecture
OrchestraPrime · Nitin Bhatti, Founder & Platform ArchitectPublished September 2026~28 min read · role paths ~10–15 minVersion 2.1 (Medical & Safety tower expanded, September 17, 2026; supersedes v1, September 8, 2026)
One sentence. Every life-sciences company is running dozens of agentic AI initiatives across research, development, manufacturing, supply chain, commercial, and medical; almost all of them are being built as disconnected projects, and the cost of that fragmentation is measured not only in money but in decision latency — the weeks a protocol sits unamended, the months an adverse quality trend goes unseen, the days a batch waits for release.
This article describes the architecture for building those initiatives as one governed, validated, production-grade platform where every agent provides intelligent decision support — the human decides, the system explains — and where the fastest decision and the best decision become the same decision. It is published as thought leadership because we believe the patterns are right for the industry. But the real question — the one this article exists to help you think through — is not whether this architecture makes sense. It is whether your organization has the operating model, the semantic discipline, the validation posture, and the leadership to build it. That is a people question, not a technology question, and it is the question the rest of this article is designed to help you answer.
~$500K
per protocol amendment
Roughly half are avoidable with precedent intelligence up front. Months of delay each.
6–15 mo
APQR signal latency
A trend that should have triggered an investigation in March appears in next February's report.
15 days
the PV clock
Intake across five channels, manual coding, narrative drafting — a hard regulatory deadline.
137×
frontier vs scoped model cost
Claude 3.5 Sonnet vs. Haiku, bounded extraction/scoring tasks, Q3 2025, prompt caching enabled. A monolith routes everything to the frontier tier.
$2–5M
per standalone use-case
Twenty use-cases as twenty projects is a consulting program, not a platform.
Bottom Line
Not a technology project — a decision-velocity operating model
An enterprise agentic AI platform for life sciences is not a technology project — it is a decision-velocity operating model built on a shared semantic substrate, a composition-first agent framework, and a cross-cutting governance plane that makes the fastest decision and the best-governed decision the same decision. Organizations that build this as forty disconnected projects will spend 2–5× more, take 2–3× longer, and end up consolidating to a platform approach anyway.
Key Findings
Five findings that reorder the priorities
1
Decision latency is itself a compliance risk — and the industry under-weights it.
Slow APQR signals, late CPV detection, delayed protocol amendments, and missed PV clocks all carry the same regulatory exposure as wrong decisions — by a different route. The platform must be measured on decision velocity as a first-class metric alongside accuracy.
2
Composition over inheritance is an economic imperative, not a design preference.
A monolithic agent with 100 tools costs ~137× more per decision than composed bounded specialists, while also being unsecurable, unvalidatable, and impossible to promote from Lab to Factory. The unit of deployment is the ≤5-tool specialist; the unit of reuse is the service kit.
3
The semantic layer is the platform's durable moat.
Individual agents are replaceable; models deprecate quarterly. The federated knowledge graph — thin core ontology, domain-owned extensions, entity resolution, process intelligence — is the asset that compounds. It is also the single largest source of underestimation: the canonical attribute dictionary over regional LIMS instances is where most manufacturing AI programs stall.
4
GxP validation of hybrid AI systems is solvable today — with GAMP 5 Category 5 for deterministic paths, CSA-aligned controls for LLM paths, and analytical-method lifecycle for statistical/chemometric models.
The posture is defensible but not yet established precedent; early movers who articulate it to Quality before the first agent enters a GxP workflow will set the standard.
5
The hardest problem is organizational, not technical.
Hiring FDEs and knowledge engineers, earning Quality's trust, resisting the pressure to scale before the foundation is proven, and killing PoCs that don't move the metric — these are the risks an ED is accountable for. The platform reduces the technical barriers so the leadership problems become the binding constraint.
Strategic Planning Assumptions
Five planning assumptions, with probability and substantiation
By 2028Probability 0.70
By 2028, 60% of top-20 pharma companies that attempt enterprise agentic AI without a shared semantic substrate will consolidate to a platform approach — after spending 2–3× more than organizations that started with one.
Substantiation
Gartner's 2024 "Top Strategic Technology Trends" identifies AI-ready data foundations as a prerequisite for scaling GenAI; McKinsey's 2024 pharma AI survey found 70% of pilot programs stall at scale due to fragmented data and governance; Novartis's own Data42 and NERVE investments signal the same consolidation pattern from within the industry.
By 2027Probability 0.65
By 2027, the industry's default AI validation posture will converge on the hybrid model described here (GAMP 5 Cat 5 for deterministic paths + CSA-aligned controls for LLM paths), driven by FDA finalization of the 2025 draft AI guidance and EMA's adopted reflection paper (September 2024), and the ISPE GAMP Guide: AI (July 2025). Organizations that retrofitted validation will carry 12–18 months of rework.
Substantiation
FDA's January 2025 draft guidance "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products" (CDER/CBER) establishes risk-based credibility frameworks for AI in drug development; ISPE's GAMP 5 2nd Edition (2022) already accommodates agile and CSA approaches; ISPE GAMP Guide: Artificial Intelligence (July 2025, 290 pp) provides the industry's first comprehensive AI validation framework; PIC/S guidance PI 041-1 (2021) establishes data integrity expectations that map directly to the governance plane's G1–G3 controls; FDA–EMA "Guiding Principles of Good AI Practice in Drug Development" (January 2026) reinforce risk-proportionate validation aligned with intended context of use.
By 2029Probability 0.60
By 2029, "Chief Agentic Officer" or equivalent will be a named role at 40%+ of top-20 pharma companies, owning the Lab, Factory, and registry described here. The role will be measured on agent graduation rate, decision velocity, and reuse ratio — not on number of agents deployed.
Substantiation
The Chief AI Officer role appeared at 25%+ of Fortune 500 companies by 2024 (Deloitte AI Institute); pharma-specific variants (Head of AI/ML at Roche, VP AI Platform at Novartis) already exist; the agentic shift — from model deployment to agent lifecycle management — creates the operational mandate that the current CAIO title does not cover.
Through 2028Probability 0.75
Through 2028, organizations that measure agentic AI programs on "number of use-cases deployed" rather than decision-velocity improvement and domain-expert acceptance rate will report program dissatisfaction at 3× the rate of those using outcome metrics.
Substantiation
Gartner's 2024 "AI Value Realization" research found that organizations using outcome-based AI metrics reported 2.5× higher satisfaction than those using activity metrics; BCG's 2024 pharma AI survey identified "measuring the wrong things" as the #2 reason for AI program failure after data quality.
By 2028Probability 0.65
By 2028, Medical Affairs will be the function where agentic AI reaches governed production fastest in top-20 pharma — ahead of R&D and manufacturing — because its decisions (inquiry response, MSL preparation, insight synthesis, congress intelligence) are high-volume, text-native, and governed by approved-source discipline rather than GMP validation. But organizations that build those agents on a commercial CRM vendor's agent platform without an architecturally enforced medical/commercial firewall and a canonical interaction model will rebuild them at least once — on CRM migration, on a scientific-exchange finding, or both.
Substantiation
Peer-reviewed Medical Affairs AI proof points already exist inside top-20 companies — semantic clustering of unsolicited medical-information inquiries (PMC, 2023) and GenAI theme analysis of field medical insights (Pharmaceutical Medicine, 2024) — while the vendor layer is moving simultaneously: Agentforce Life Sciences reached general availability in October 2025 with pharma early adopters, Veeva shipped AI agents for PromoMats in December 2025, and Veeva CRM's end-of-support date (December 31, 2029) forces every Veeva CRM customer through a Vault CRM vs. Salesforce Life Sciences Cloud decision inside this window. The industry is therefore standing up Medical Affairs agents during a CRM transition — which is precisely when binding agents to CRM objects rather than a canonical model is most expensive.
Recommendations
By persona — what to do in Phase 1, and what not to measure
For the Chief Data & AI Officer / Head of AI Platform Strategy:
Publish the agent platform contract (identity, MCP tool access, observability, registry) before the second agent estate arrives — partner-built agents will deploy around the platform if the contract is not ready. Start in Phase 1.
Measure the platform on marginal cost per use-case (should fall 40–60% by use-case 5–6), reuse ratio (target ≥70%), and time-to-first-grounded-answer for new use-cases (target: days). Do not measure on number of agents deployed.
Fund the semantic layer as infrastructure, not as a project. The core ontology and KG-as-a-Service are consumed by every use-case; starving them starves everything.
For the Agentic Lab & Architecture Lead:
Enforce composition from day one — ≤5 tools per specialist, orchestrators that compose rather than inflate. The 137× cost difference compounds with every agent that violates the rule.
Build kill criteria into every PoC: if precision/recall/F1 does not improve over two iteration cycles, or coach-mode acceptance is flat, stop. Killing fast is a Lab success metric. Publish post-mortems as deprecated registry cards.
Staff the Lab with FDEs who speak the domain. The gap between "what the toxicologist needs" and "which specialist, which kit, which data product" is the gap the Lab exists to close.
For the Head of Agentic Factory:
Own the promotion gate with defined criteria, named reviewers, and an escalation SLA. The gate is the hardest organizational problem; without it, PoCs accumulate.
Build the support model (L1/L2/L3) before the first agent graduates — not after. L1 is domain super-users; L2 is platform SRE; L3 is pod re-formation. The model lifecycle (quarterly successor-model testing) is L2 work that most organizations forget.
Treat the Enterprise AI Operations Command Center as your network operations center. If you cannot see drift, cost, and acceptance-rate curves for every factory agent in one surface, you do not have a Factory.
For the Head of Semantic & Knowledge Engineering:
The federated ontology model — thin core, domain-owned extensions under a published modeling standard — is the only design that avoids both the central bottleneck and the disconnected-graph swamp. Staff the core thinly (3–5 ontologists); staff the domain extensions as part of each pod.
Entity resolution is a first-class service, not a nice-to-have. The same golden-record discipline commercial has applied to HCPs for years must extend to sites, investigators, batches, and materials. Site selection is a master-data problem before it is a prediction problem.
The canonical attribute dictionary over regional LIMS instances and CMO feeds is the single most underestimated work item in the platform. Budget for it explicitly; it is quality-owned work under change control.
For the Medical Affairs AI Platform Lead (product, architecture, or engineering):
Abstract the interaction model before the first agent ships. Bind agents and analytics to a canonical Medical Affairs data model — interaction, inquiry, insight, content, HCP, event, transfer of value — materialized from whichever CRM the persona works in (Veeva CRM today; Vault CRM or Salesforce Life Sciences Cloud tomorrow). The migration becomes a connector swap, not an agent rewrite. §8.7 describes the model.
Buy the vendor agents where they exist; build only where coverage stops. Vendor pre-review and quick-check agents for content, and the CRM vendor's engagement agents, run inside the platform contract (UC-H1) as first-pass triage. Build the enterprise's own claim-substantiation rules, the scientific-exchange boundary, MSL pre-call intelligence, insight theming, and congress extraction — the capabilities that carry the enterprise's medical judgment. Human MLR and medical review sign-off stays.
Make the firewall, the off-label boundary, and AE routing structural. Separate identity principal, separate registry scope, separate KG partition for medical; unsolicited-request and off-label rules per market (FDA vs. EFPIA) as versioned rules, not prompt instructions; adverse-event detection in inquiries and interactions as a deterministic, non-bypassable hook with human confirmation. A policy that lives in a prompt is not a control.
Measure the portfolio on decision velocity for the seven Medical Affairs decisions in §8.7 — inquiry response time, MSL prep time per interaction, abstract-to-alert latency, insight-to-strategy time, gap-to-funding-decision time, medical review cycle time, PV clock compliance — plus brief-acceptance rate in coach mode and citation completeness. Not on number of agents. Put AI telemetry (token cost, latency, drift, hallucination rate, prompt-injection attempts per agent, per product, per region) on one plane from day one; it is the TCO dashboard the portfolio will be judged on.
How to Read This Article
Pick your lens — the contents highlight your path
This is a long piece because the problem is a whole-enterprise problem. It is written to be read straight through, but four kinds of readers will want to weight sections differently. Select a role and the table of contents highlights your read first and then sections, estimates the reading time, and shows what your role gains. Your choice is remembered on this device.
for the platform strategy lens — platform economics, sequencing, competitive context, org design, and the measurement framework that tells you whether the bet is paying off.
Before — today's state
Three agent estates, no shared governance; each use-case a $2–5M standalone; no cross-functional decision-velocity metric; competitive landscape assessed vendor-by-vendor.
After — with the platform
One platform contract governs all agents; marginal cost drops 40–60% by use-case 5–6; one scorecard from target ID to batch release; coherence matrix shows reuse at a glance.
for the agentic lab lens — agent anatomy, composition rules, registry-driven reuse, and the kill criteria that keep the Lab honest.
Before — today's state
PoCs built as monoliths with 50–100 tools; no kill criteria; no registry; "done" means "demo works".
After — with the platform
≤5-tool bounded specialists; kill criteria enforced per iteration; every PoC leaves a registry card; "done" means the promotion gate passes.
for the agentic factory lens — promotion gates, support lens, validation posture, observability, and the command-center surfaces production agents feed.
Before — today's state
Validated agents without SLAs; no support model; model deprecation discovered in production; no observability across agents.
After — with the platform
L1/L2/L3 support tiers; quarterly successor-model testing; Enterprise AI Ops Command Center; confidence-gated degraded mode.
for the semantic & knowledge engineering lens — federated ontology governance, KG-as-a-Service, data-lake integration, run-time triplet materialization, process intelligence.
Before — today's state
Per-project ontologies that diverge; entity resolution only in commercial; KG-as-a-query vs. KG-as-a-service; data fabric aspirational.
After — with the platform
Thin core + federated domain extensions under a modeling standard; entity resolution as a first-class cross-domain service; canonical attribute dictionary over regional LIMS; process intelligence fused as confidence-scored edges.
If your job is…
Read first
Then
Why
Enterprise AI platform strategy — portfolio, investment prioritization, the operating system for AI across the company
§1, §2, §4, §9, §13, §14, §15, §18
§5, §12
Platform economics, sequencing, competitive context, org design, and the measurement framework that tells you whether the bet is paying off
Three agent estates, no shared governance; each use-case a $2–5M standalone; no cross-functional decision-velocity metric; competitive landscape assessed vendor-by-vendor
One platform contract governs all agents; marginal cost drops 40–60% by use-case 5–6; one scorecard from target ID to batch release; coherence matrix shows reuse at a glance
Agentic Lab
PoCs built as monoliths with 50–100 tools; no kill criteria; no registry; "done" means "demo works"
≤5-tool bounded specialists; kill criteria enforced per iteration; every PoC leaves a registry card; "done" means the promotion gate passes
Agentic Factory
Validated agents without SLAs; no support model; model deprecation discovered in production; no observability across agents
L1/L2/L3 support tiers; quarterly successor-model testing; Enterprise AI Ops Command Center; confidence-gated degraded mode
Semantic & KE
Per-project ontologies that diverge; entity resolution only in commercial; KG-as-a-query vs. KG-as-a-service; data fabric aspirational
Thin core + federated domain extensions under a modeling standard; entity resolution as a first-class cross-domain service; canonical attribute dictionary over regional LIMS; process intelligence fused as confidence-scored edges
Whatever lens you bring, the argument has one spine: decision velocity and decision quality are not a trade-off in pharma; they are the same engineering problem, and the platform is how you solve it once.
Section 1 · The Stakes
Why decision velocity is the new competitive advantage
Pharma has always known that a wrong decision is expensive. A missed safety signal, a mis-specified endpoint, a batch released against an unnoticed trend — these carry costs measured in patient harm, inspection findings, and write-offs. What the industry has been slower to internalize is that a slow decision carries the same order of cost, and that most of the industry's slowness is not caused by caution. It is caused by the time it takes to assemble evidence.
Consider where the calendar actually goes:
Decision
Typical latency today
What the latency costs
Why it is slow
Demand type
Amend a clinical protocol
Weeks to draft, months to implement
~$500K per amendment; roughly half are avoidable with better design intelligence up front
Precedent lives in a hundred prior protocols nobody has time to read
Type 1 waste
Confirm a process is in a state of control (APQR)
Findings surface 6–15 months after the batches were made
A trend that should have triggered an investigation in March appears in the following February's report
Historian export, Excel, manual batch reconciliation, 4–8 weeks per product
Type 1 waste
Detect process drift between batches (CPV)
4–12 weeks from batch completion to chart
The first signal that something changed is a failed batch
Five to ten of 30–50 CPPs actually trended; interactions invisible
Type 2 / failure demand
Stop a blend when it is uniform
Fixed-time step: blend for N minutes, sample, wait for QC
15–45 minutes per batch at $5K–$15K per batch-hour of capacity
Site identity scattered across CTMS, EDC, CRO feeds with no golden record
Type 1 waste
Close a deviation investigation
Days to weeks
Batch on hold; CAPA effectiveness unknown
Similar historical deviations live in PDFs and memory
Failure demand
Process an expedited adverse-event case
Hard 15-day clock
Regulatory non-compliance
Intake across five channels, manual coding, narrative drafting
Type 1 waste
Answer a medical-information inquiry with approved content
Hours to days
HCP trust; off-label risk
MLR-approved content not retrievable with provenance
Type 1 waste
Prepare an MSL for a KOL interaction
1–3 hours per interaction
Fewer, shallower scientific exchanges; missed follow-ups; insights never captured
Prior interactions in the CRM, inquiries in another system, the KOL's abstracts in a third, approved SRLs in Vault — assembled by hand
Type 1 waste
Turn a congress into medical intelligence
Post-congress report in 4–6 weeks
Competitor data and KOL shifts reach TA leads after the strategy cycle has moved on
Thousands of abstracts read by a handful of people; no structured extraction, no KG
Type 1 waste
Value work — the decision itself (minutes)Type 1 waste — necessary only because of how the system is designedType 2 waste / failure demand — exists only because something upstream failed
Every row has the same shape: the evidence exists, it is scattered across systems and formats, and a skilled person spends most of their time finding and reconciling it rather than deciding. The decision itself — the part that requires a toxicologist, a medical writer, a QA lead, a process scientist — is often minutes. The assembly is weeks.
There is a sharper way to read that table, and anyone who has run a Lean or Six Sigma program will recognize it. Systems thinking — the Vanguard Method developed by John Seddon — asks first whether the work an organization does is value demand (what the customer actually needs) or failure demand (demand caused by the system's own prior failures), and classifies each activity as value work, Type 1 waste (non-value, but currently necessary given how the system is designed), or Type 2 waste (exists only because something upstream failed). Read that way, the table is almost entirely amber and red. Reading a hundred prior protocols is Type 1 waste — necessary only because precedent was never structured. Hand-reconciling batches for an APQR is Type 1. A failed batch that is the first signal of drift is Type 2, and the investigation it triggers is failure demand. The decision itself is the thin green thread of value work. Pharma's calendar is not consumed by deciding; it is consumed by waste the system generates for itself.
This is why agentic AI matters in pharma in a way that is different from how it matters in a consumer technology company. The value is not automation. Medicine will always be a human endeavor, and in regulated workflows the human must decide. The value is compressing the assembly so that the human decides sooner, with more of the evidence in front of them, and with the rationale captured so the next person can decide faster still. That is intelligent decision support, and it is the only framing under which agentic AI survives contact with GxP.
Three questions define whether an enterprise gets there. They are asked by three different executives, and a platform that cannot answer all three does not ship.
The Chief Data & AI Officer
— The Platform Question.
"We have three agent estates — internal, a CRM vendor's agents, an IT-operations partner's agents — and no shared governance. How do I prevent forty disconnected implementations?"
The Head of R&D AI
— The Validation Question.
"The PoC worked in the lab. Quality says they can't validate an LLM under GAMP 5. How do we get this to production?"
The VP Quality & Manufacturing
— The Trust Question.
"Show me which source supported which sentence in the investigation report. If I can't trace it, I can't sign it."
The rest of this article is an answer to those three questions. It starts by being honest about why the answer has to be different from what a technology company would build.
Section 2 · The Challenge
Why pharma agentic AI demands a different architecture
The cost of a slow decision is the cost of a wrong one
In consumer technology, an agent that is wrong five percent of the time is impressive, and an agent that is slow is merely annoying. In pharma, an agent that is wrong once in a regulatory submission, a safety report, or a batch-release decision creates an inspection finding — and an agent that is slow creates the same finding by a different route: the trend nobody saw, the amendment nobody avoided, the clock nobody met. Regulators have moved from asking "did you do the review" to "did the review find anything, and what did you do about it, and when." Latency is now a compliance property.
That is the reframing this section exists to make. The industry conversation about pharma AI has been dominated by the wrong-decision risk — hallucination, bias, black boxes — and that risk is real. But an architecture built only to prevent wrong decisions produces a system so gated, so manual, so bespoke per use-case, that it never compresses any calendar. The platform has to be designed for both failure modes at once: it has to make decisions faster and more governed simultaneously, or it makes neither. Four constraints make that harder in pharma than anywhere else.
Constraint 1 · Regulatory
GxP governance is non-negotiable, and it is per-decision, not per-system
21 CFR Part 11 audit trails. GAMP 5 validation. EU Annex 11 electronic records. ALCOA+ data integrity. Every agent that touches a regulated process needs immutable execution history, version-controlled rules, and human sign-off — not as a nice-to-have but as a condition of deployment. The subtlety is that the governance requirement attaches to the decision, not to the software. A drafting agent that produces a CSR section is subject to Part 11 at the moment a medical writer signs the section. A CPV agent that raises an SPC signal is subject to GMP at the moment a process scientist acts on it. The platform must therefore know, for every agent invocation, which regulatory frame applies to the decision it is supporting — and adjust the governance envelope accordingly. A single flat "compliance mode" is both too much for discovery science and not enough for batch release.
Regulatory constraint. Governance must be scoped to the decision, and the scoping must itself be auditable.
Constraint 2 · Fragmentation
Three agent estates, zero shared governance
Internal R&D agents. Commercial agents on a CRM vendor's agent platform. IT-operations agents built and operated by a systems integrator under a multi-year contract. Each estate arrives with its own identity model, tool access, observability, and evaluation practice. Without a shared platform contract, the enterprise ends up with three incompatible agent ecosystems that cannot be governed as one — and when an inspector asks "which agents can touch validated systems, and how do you know," the answer is three different answers from three different vendors.
Fragmentation risk. The platform's first product is a contract that partner-built agents deploy into, not around.
Constraint 3 · Scale
Dozens of use-cases × dozens of implementations is unsustainable
Target identification, generative chemistry, preclinical safety, trial design, site selection, operational precision, regulatory documentation, tech transfer, APQR, CPV, PAT, blend endpoint, deviation investigation, supplier risk, cold chain, HCP engagement, MLR review, medical inquiry, pharmacovigilance, semantic search. That is twenty before anyone in finance or HR has raised a hand. If each one builds its own ontology, its own agent framework, its own evaluation harness, and its own validation story, you have built a consulting program, not a platform — and at $2–5M per standalone implementation, the arithmetic ends the program by year three.
Scale blocker. The unit of reuse has to be a platform service, not a project template.
Constraint 4 · Velocity
Latency is itself a risk, and it compounds
This is the constraint the industry under-weights. Every month an adverse trend goes undetected is a month more batches ship against it. Every week a protocol design sits without precedent analysis is a week closer to an amendment. Every day a deviation investigation waits for someone to find the similar case from 2021 is a day of batch hold. These latencies compound across the pipeline: a slow site-selection decision delays enrollment, which delays the CSR, which delays the submission. A platform that governs beautifully but leaves the assembly time untouched has not reduced risk; it has formalized it.
Velocity risk. The platform must be measured on decision latency as a first-class metric, alongside accuracy and compliance.
Two constraints are properties of the domain; two are conditions the enterprise created for itself
A systems-thinking reader will notice that the four constraints are not of one kind. Constraints 1 and 4 are properties of the domain: GxP is the patient's purpose expressed as regulation, and latency compounds because the pipeline is sequential. No architecture removes them; the platform must serve them. Constraints 2 and 3 are system conditions — conditions the enterprise created for itself through procurement decisions, org design, and per-project funding, and which now generate most of the failure demand described in §1. Three agent estates exist because three contracts were signed separately. Dozens of implementations exist because use-cases are funded as projects. The distinction governs what the platform may honestly claim: against Constraints 1 and 4 it can only compress; against Constraints 2 and 3 it can remove the cause. An architecture that governs beautifully but leaves those two system conditions in place has automated the waste rather than removed it — and automated waste is still waste; it just runs cheaper.
These four constraints are why the architecture that follows looks the way it does. Governance is a cross-cutting plane, not a gate at the end. Agents are composed from bounded specialists rather than inflated into monoliths, because monoliths are neither fast nor validatable. The semantic layer is shared and federated, because that is the only way use-case N+1 is cheaper than use-case N. And every use-case is described in terms of the decision it accelerates, because that is the only value that matters.
Before the architecture, three principles. They are the acceptance criteria every layer has to meet.
Section 3 · Guiding Principles
Three acceptance criteria every component must meet
These are not aspirational. They are the acceptance criteria for every component in the platform. A use-case that violates any of them does not graduate from the Lab. They deliberately mirror the way leading R&D organizations have articulated their AI posture — because that articulation is right, and because a platform that speaks the enterprise's own language gets adopted.
Purpose-Driven
Address specific scientific and operational challenges — not AI for AI's sake. Every agent has a named business problem, a named user, and a named metric. If you cannot articulate the value of a single inference — what decision it accelerates, by how much, for whom — do not ship it.
This principle is why every use-case in §8 opens with the decision, not the model, and why the Lab's kill criteria in §7 are tied to a metric moving, not to a demo working.
Smart & Scalable
Reusable across therapeutic areas, sites, and geographies. One ontology, one agent framework, one validation strategy — configured per domain, not rebuilt. A new use-case is configuration, not a project.
This principle is why §5.3 insists on a thin core ontology with federated domain extensions, why §6 introduces a registry so agents can be discovered and composed rather than re-implemented, and why §9 measures the platform on the reuse ratio rather than on the number of agents deployed.
Human-Centered & Responsible
"Artificial intelligence needs to augment natural intelligence." Medicine will always be a human endeavor. Every agent provides intelligent decision support — rationale, provenance, an owner — and no agent exercises black-box autonomy in a regulated workflow. The goal is to raise the average to the level of the best: to shift the performance of many individuals closer to the performance of the best individual, by putting the best evidence and the best precedent in front of every decision-maker.
This principle is why §10's dual-path reasoning ends in a merged assessment a human signs, why §5.5's governance plane captures decision provenance and not just evidence provenance, and why §19 puts the promotion gate in the hands of domain experts rather than engineers.
Three principles, one consequence: the platform is a decision-support substrate, not an automation engine. That consequence is the thesis.
Section 4 · The Thesis
One substrate, every decision
What if every decision in your enterprise — from target selection to batch release to the next field call — ran on one semantic substrate, one agent framework, one governance plane, so that the fastest decision and the best-governed decision were the same decision?
That is the whole proposition, and it is worth being precise about why the two halves belong in one sentence. Speed without governance is what most enterprises have today in their ungoverned pilot estates: fast, impressive, unshippable. Governance without speed is what most enterprises get when Quality retrofits controls onto those pilots: validated, auditable, and so slow that nobody uses it. The industry has been treating these as a dial to be tuned — more speed here, more control there — when they are actually the same engineering problem viewed from two sides.
The reason they are the same problem is evidence assembly. The thing that makes a decision slow (finding and reconciling scattered evidence) is the same thing that makes it hard to govern (no provenance for what was found). Solve the assembly once — a semantic layer that makes enterprise evidence queryable with provenance, agents that reason over it with citations, a governance plane that records what was retrieved and what the human did with it — and speed and governance both fall out of the same mechanism. A recommendation that arrives with its evidence trail attached is both faster to act on and faster to audit.
That is what "raising the average to the level of the best" means in practice. The best toxicologist, the best medical writer, the best process scientist are the best partly because they remember the precedent. The platform makes the precedent available to everyone, with the source attached, at the moment of decision. The human still decides. The decision just arrives sooner and leaves a trail.
The four-layer architecture is how that substrate is built.
Section 5 · Platform Architecture
Four layers, one governance plane — each layer serves every domain
The platform is a four-layer stack — Experience, Agentic, Semantic & Knowledge, Data — with a cross-cutting governance plane that enforces GxP compliance, identity, audit trails, observability, and evaluation across all four. The layers are described top-down here because that is how a user experiences them; they are built bottom-up, which is how §13's roadmap sequences them. The single most important property of the stack is that each layer serves every functional domain. There is not a research stack and a manufacturing stack and a commercial stack. There is one stack, and the domains differ in which ontologies, which agents, which data products, and which governance envelope are configured for them.
Reading the diagram. Each column is a functional domain; each band is a layer. A cell answers "what does this layer give this domain?" — hover a column header to isolate a domain, hover a band to isolate a layer, click a cell to jump to the tower in §8. Requests and queries flow down the left edge; findings and provenance flow up the right; every layer reports audit, trace and policy into the Governance Plane on the right, whose eight categories are detailed in §5.5. Between Semantic and Data, the two-headed arrow is run-time triplet materialization (§5.4).
5.1 Experience Layer
What it is. Role-specific workspaces that consume from the agentic layer. Each persona sees findings, recommendations, and drafts through purpose-built interfaces — not a generic chatbot. The experience layer is where intelligent decision support becomes a decision: it is where the human accepts, modifies, or rejects, and where that choice is captured.
Design rule. No workspace talks to a model directly. Every surface consumes from registered agents (§6) through the platform contract, which is what makes a partner-built surface (a CRM vendor's field-force interface, an SI's IT-ops console) governable in the same way as an internal one.
Surfaces, by domain:
Surfaces, by domain — what the user sees and what the platform captures (15 surfaces)
Surface
Domain
What the user sees
What the platform captures
Target evidence workspace
Research
Target dossier with KG neighborhood, evidence tiers, cited liability brief
Pre-call brief with every line cited (prior interactions, open inquiries, the HCP's recent abstracts, approved SRLs, what cannot be shared), territory engagement plan vs. the medical plan; works offline
Brief accept/edit, post-interaction structured insight capture — never auto-logged to CRM
Congress intelligence surface
Medical
Abstract-level alerts by product, competitor, endpoint, and KOL within hours; KG neighborhood per abstract; proposed action
Corroborated insight themes with confidence and n_sources, inquiry-cluster trends, evidence gaps and candidate studies, congress signals — one surface per TA
Adopt/monitor/dismiss on themes; fund/decline on study concepts; strategy changes cited to insights
Medical/legal/regulatory disposition, approval record (Part 11)
PV case workbench
Medical & Safety
Structured case, coding suggestions, cited narrative, clock status
Reviewer sign-off, clock compliance
Command centers
Cross-domain
Multi-agent feeds composed into decision hubs — §12
Escalations, decision log
Agent marketplace
Enterprise
Registry-driven discovery of agents and kits — §6
Usage, approvals, chargeback
The transition from experience to agentic is the platform's most important boundary. Everything the user sees was produced by an agent; everything the user does is fed back to one.
Shared spaces — three levels of context
The experience layer is not a single surface. It is a hierarchy of shared spaces that establish what context is visible, to whom, and at what scope:
Domain space — the broadest shared context
What is sharedKG subgraph, registered agents, data products, command-center feeds, eval benchmarks, kit templates
What is privateNothing at this level — the domain is the broadest shared context
ExampleThe Manufacturing domain space: all sites see the same CPV signals, the same APQR roll-ups, the same deviation history, the same KG, the same registered agents
Sub-group / Division space
What is sharedTeam-specific filters, site-specific thresholds, working hypotheses, in-progress investigation state, draft analyses under review
What is privateRaw data from other sub-groups before aggregation; draft dispositions before sign-off
ExampleA site quality team's working space: they see their site's CPV charts, their open investigations, their draft APQR — not another site's drafts. But the domain space above them sees cross-site trends
Individual space
What is sharedPersonal queue, coaching feedback history, acceptance-rate curve, preferred formats, recent session context
What is privateFeedback corrections the individual has made (tone preferences, edit patterns), draft-in-progress before submission
ExampleA process scientist's workspace: their personal CPV review queue, their correction history that tunes future agent drafts, their individual acceptance rate that drives the coach-mode graduation curve
↑This three-level model mirrors how knowledge actually flows in a pharmaceutical organization: company-level observations are shared, team-level working state is collaborative within the team, and individual learning is private until it graduates. The graduation mechanism is the same trust-tier model the KG uses: an individual correction becomes a team pattern when corroborated by multiple users, and a team pattern becomes a domain standard when validated by the domain owner. The proven production pattern — where co-worker memory at the individual level (conversation history, edit corrections, preferred formats) promotes to company-level knowledge when 2+ users corroborate — is the same mechanism, applied at enterprise scale.
5.2 Agentic Layer — Lab & Factory
What it is. Two operating modes under one framework. The Lab develops PoCs with rapid iteration. The Factory industrializes validated agents into governed, monitored, production-grade deployments. A promotion gate enforces accuracy, safety, and validation thresholds between them. §7 describes the operating model in full; this section describes the anatomy of the agents themselves, because the anatomy is what makes Lab-to-Factory promotion possible at all.
Agent types. Five archetypes cover every use-case in §8:
Archetype
Job
Typical tools (≤5)
Typical model tier
Extraction
Turn unstructured source into structured, provenance-tagged findings
Document parser, KG writer, schema validator
Scoped / small
Reasoning
Traverse the KG, compare, score, synthesize with citations
KG query, rule engine, retrieval
Frontier for synthesis; scoped for scoring
Drafting
Produce sections, narratives, hypotheses from approved sources only
The six-layer anatomy. Every agent, regardless of archetype or domain, is built the same way: Model → System Prompt → Context Window → Tools → Hooks → Agent Rules. The model is pinned by version. The system prompt is versioned and owned. The context window is budgeted. Tools are the only way the agent touches the world, and there are no more than five of them visible to the model. Hooks are the deterministic pre/post steps (input validation, output checking, audit logging) that run whether the model cooperates or not. Agent rules are the declarative policy — what confidence threshold triggers escalation, which GxP class applies, who owns it. This anatomy is what gets registered in §6 and what the Factory validates in §11.
Modelpinned by version
→
System Promptversioned, owned
→
Context Windowbudgeted
→
Tools≤5 visible to the model
→
Hooksdeterministic pre/post
→
Agent Rulesthresholds · GxP class · owner
Composition over inheritance — the economic and architectural imperative
Here is what happens when you do not compose.
A team builds a "clinical operations agent." It starts well: three tools, one job. Then someone asks it to also check drug supply. Then to draft the monitoring visit letter. Then to look up the site's contract status. Then to pull the enrollment forecast. Eighteen months later the agent has a hundred tools, and four things have gone wrong at once.
The context window is spent before the user speaks. A hundred tool schemas consume tens of thousands of tokens of every single call — before a word of the actual task arrives. Prompt caching helps, until someone edits one tool description and invalidates the cache for the whole estate. The agent is paying full frontier-model price to re-read its own tool manual on every invocation.
Tool selection degrades. With five tools, a frontier model picks the right one almost every time. With a hundred, it picks a plausible one — and "plausible" is how an agent ends up querying the IRT system when it meant the CTMS. The failure is silent, and it is not the model's fault. It is the architect's.
Security boundaries blur. An agent that can read pharmacovigilance cases and write to the CRM and modify a monitoring plan is not one agent with a broad job; it is one identity principal with a broad blast radius. The least-privilege question — "what is the minimum this agent needs to touch?" — has no answer, because the answer is "everything."
Validation becomes impossible. GAMP 5 asks what the system does and how you know it does it. A hundred-tool agent's answer is "it depends on what the model chose." Category 5 validation of a system whose control flow is decided at run-time by a stochastic planner is not a validation strategy. It is a hope.
Now the number. In our measured deployments, the cost difference between a frontier model and a scoped model for the same bounded task is roughly 137× (Claude 3.5 Sonnet vs. Claude 3 Haiku on bounded extraction and scoring tasks, Q3 2025, prompt caching enabled). A monolith routes everything to the frontier tier, because a monolith cannot tell which sub-task is easy. A composed system routes extraction, normalization, scoring, and monitoring to scoped models — and reserves the frontier tier for synthesis, planning, and the handful of decisions that need it. On a portfolio of a dozen production agents handling tens of thousands of invocations a day, the difference between those two routing policies is not a line item. It is whether the program has a budget next year.
Monolith — "one agent, a hundred tools"
The context window before the user's first token
100 tool schemas · ~62%
system prompt
task
room
Tens of thousands of tokens re-read on every call, at frontier price — the agent pays to read its own tool manual before the task arrives.
Routes everything to the frontier tier, because a monolith cannot tell which sub-task is easy. One principal, every permission. Cat 5 scope is "everything".
Composed — bounded specialists, ≤5 tools each
137×measured cost difference, frontier vs scoped model, on the tokens that matter
prompt
task + evidence
room for reasoning
Hundreds of tokens of schema per specialist. Extraction, normalization, scoring and monitoring go to scoped models; the frontier tier is reserved for synthesis, planning, and the handful of decisions that need it.
On a dozen production agents handling tens of thousands of invocations a day, the difference between the two routing policies is not a line item. It is whether the program has a budget next year.
Monolith ("one agent, a hundred tools")
Composed ("bounded specialists, ≤5 tools each")
Tool schema overhead per call
Tens of thousands of tokens, every call
Hundreds of tokens per specialist
Model tier
Frontier for everything (cannot discriminate)
Frontier for ~20% of calls; scoped for the rest
Relative cost per decision (normalized)
~100
~3–8
Tool-selection accuracy
Degrades with tool count
Stable
Least-privilege story
None — one principal, every permission
Each specialist has exactly its allowlist
GAMP 5 posture
Control flow is stochastic; Cat 5 scope is "everything"
Deterministic control flow at the orchestrator; specialists validated as bounded components
Failure isolation
One bad tool call can corrupt the whole run
Failure is local to one specialist, and the orchestrator's hooks catch it
Promotion to Factory
Whole monolith must re-validate on every change
Specialists promote independently; orchestrator re-validates only on topology change
The design rule is therefore the default: the orchestrator composes bounded specialists via natural language; it never inflates one agent with a hundred tools. Every specialist has ≤5 model-visible tools, its own eval record, its own identity, and its own place in the registry. This is not an architectural preference. It is the only configuration that is simultaneously affordable, secure, and validatable — and it is what the 137× figure is really measuring.
Proven pattern
ProtoCheck's four specialist reviewers (safety, efficacy, regulatory, operational) each hold ≤5 tools and apply expert-calibrated rules; a coordinator merges findings with provenance. Doscierge's checker is a deterministic MCP tool that an orchestrator calls, not an agent that decides — the deliberate GAMP 5 Category 5 ceiling.
Components in the layer: Extraction agents · Reasoning agents · Drafting agents · Monitoring agents · Orchestrator agents · Agent Registry & Service Catalog (§6) · MCP tool servers · Evaluation harness (P/R/F1 per use-case, per domain, per class) · Closed-loop evaluation (Model-loop, Dogfooding-loop, Production-loop — §7).
The agents reason over something. That something is the semantic layer.
5.3 Semantic & Knowledge Layer
What it is. The shared substrate every agent reasons over. A thin core ontology, centrally owned; domain ontologies federated to domain teams under a published modeling standard; the whole exposed as Knowledge-Graph-as-a-Service with query, provenance, and versioning APIs that agents call through MCP. No agent builds its own retrieval. None owns a private copy of the graph.
Why federated. A single central ontology becomes a bottleneck by month six — every domain is waiting on the ontology team to model its entities. A set of team-built graphs becomes a swamp by month twelve — nothing joins to anything. The federated model threads the needle: the core owns the entities everyone shares (product, compound, study, site, HCP/investigator, document, batch, material, specification), and each domain owns its extensions under a modeling standard that guarantees they join back to the core.
Core ontology — centrally owned, thin, versioned
Product · Compound · Study · Site · Investigator/HCP · Document · Batch · Material · Specification · Organization · Regulatory Requirement
Aligned to: IDMP, CDISC (SDTM/ADaM/PRM), MedDRA, ISA-88/95, FHIR
protocol · endpoint · population · design element · enrollment curve · monitoring finding · supply state · regulatory precedent
Aligned to: CDISC PRM/SDTM, ICH E6/E8/E9, CT.gov
CMC / Product Development KG
formulation · unit operation · CPP · CQA · design space · control strategy · analytical method · stability study · established condition (ICH Q12)
Aligned to: ICH Q8–Q14, USP, CTD Module 3
Manufacturing & Quality KG
batch · phase · equipment · historian tag · parameter · sample · test · result · specification · genealogy · deviation · CAPA · change · PAT model
Aligned to: ISA-88/95, ICH Q7/Q9/Q10, 21 CFR 211
Supply Chain KG
material lot · supplier · CMO · CoA/CoC · shipment · lane · temperature excursion · serialization event · demand signal
Aligned to: GS1, DSCSA, GDP
Commercial KG
HCP/HCO golden record · territory · product · indication · MLR-approved asset · interaction · consent
Aligned to: NPI, MDM golden record, MLR taxonomy
Medical & Safety KG
case · adverse event · seriousness · expectedness · causality · MedDRA PT · signal · medical inquiry · inquiry cluster · MSL interaction (type, topics, content shared) · field insight · insight theme · KOL tier · congress · abstract · publication · evidence gap · study concept · medical content · claim · transfer of value
Aligned to: MedDRA, E2B(R3), CIOMS, GVP, Sunshine Act / Open Payments, EFPIA Code
Compound ↔ study ↔ site ↔ HCP ↔ batch ↔ material — the same real-world thing appears under different identifiers in CTMS, EDC, LIMS, ERP, CRM, and the CRO's feed. The semantic layer owns the golden record and the crosswalk, using the master-data discipline that commercial organizations have applied to HCPs for years and R&D and manufacturing rarely have. Site selection (§8.2) is a master-data problem before it is a prediction problem; batch genealogy (§8.4) is a master-data problem before it is a trending problem.
Process intelligence is fused into the graph
Event-log mining over MES, QMS, CTMS, and ticketing systems — Celonis-style process discovery — surfaces what actually happens: where batches stall, which deviations recur, how long a protocol amendment really takes end to end. Those temporal and statistical findings are written into the KG as confidence-scored edges alongside the semantic and normative knowledge of what should happen. An agent asking "why is this batch late" can traverse from the batch to the phase where the delay concentrates to the deviation type that usually causes it to the CAPA that was supposed to fix it. A sample of the event log is enough for discovery; production monitoring needs the full stream.
Trust tiers and provenance on every edge
Every triplet carries a source, a trust tier, a confidence, a timestamp, and a version. Normative sources (regulations, approved specifications, signed rules) sit at the top tier; mined and inferred edges sit lower with a hard confidence cap. An agent's citation therefore carries not just where the fact came from but how much to trust it — which is what a reviewer needs to decide quickly.
Regional KG partitioning — the graph has a geography
The federated model above is federated by domain. A global pharma graph must also be federated by jurisdiction, and the two partitionings are orthogonal: a Medical & Safety triple about a German patient and one about a Texan patient belong to the same domain and different residency zones. The design rule is that residency is a property of the triple, not of the graph, and the KG-as-a-Service endpoint is a router before it is a store.
GLOBAL CORE(replicated to every region; contains no personal data)
Product · Compound · Study · Site · Material · Specification · Regulatory Requirement ·
ontologies · signed rules · label versions by market · MedDRA/IDMP/CDISC alignment ·
HCP professional attributes (specialty, affiliation, publications — public-record)
│
├── REGIONAL PARTITION — EU Azure West Europe (NL) [+ North Europe (IE) DR]
│ patient-linked triples (case, subject, consent) · HCP interaction triples ·
│ prescribing-linked triples · MSL note content · inquiry text
│ Legal frame: GDPR Art. 9 (special categories), Art. 44–49 (transfers), Art. 35 (DPIA)
│
├── REGIONAL PARTITION — US Azure East US 2 (VA) [+ Central US DR]
│ PHI-bearing triples (HIPAA 45 CFR 164), US HCP interactions, US prescribing
│ Legal frame: HIPAA Privacy/Security Rules; state laws (CCPA/CPRA, WA MHMDA)
│
├── REGIONAL PARTITION — CN Azure China North (Beijing, operated by 21Vianet)
│ all personal information and "important data" of Chinese origin
│ Legal frame: PIPL Art. 38–40 (CAC security assessment before any export);
│ HGRAC for human genetic resources — no export of genomic triples, period
│
├── REGIONAL PARTITION — JP Azure Japan East [+ Japan West DR]
│ Legal frame: APPI (anonymized / pseudonymized info regime) · PMDA · J-GCP
│ Exit: EU–Japan mutual adequacy decision; APEC CBPR
│
├── REGIONAL PARTITION — BR Azure Brazil South [+ Brazil Southeast DR]
│ Legal frame: LGPD Art. 11 (sensitive health data) · Art. 33 (international transfers) · ANVISA · CONEP ethics
│ Exit: ANPD standard contractual clauses (Resolution CD/ANPD 19/2024, finalized August 2024)
│
└── REGIONAL PARTITION — CH · ... Azure Switzerland North · additional as operations require
Legal frame: FADP (adequate under GDPR) — same mechanism, different legal-basis edges
Cross-border query governance
An agent never addresses a regional partition directly. It issues one federated query (SPARQL SERVICE blocks, or the equivalent property-graph federation) to the global KG-as-a-Service endpoint, which dispatches sub-queries to the regional endpoints the agent's card (§6) permits. Each regional endpoint evaluates the classification of the triples the sub-query would touch — not the identity of the caller alone — and applies one of three outcomes: return full results (non-PII triples), return aggregates only with a k-anonymity floor (PII triples, caller outside the region), or refuse and log (PII triples, caller not authorized for aggregates either). The federation layer joins what comes back.
The consequence that matters for an inspector: a signal-detection agent running in Basel can ask "how many serious hepatic events across all regions for product X in Q2" and receive five numbers, but it cannot receive five case narratives — and the audit trail (G3) shows both the question and the classification decision each region made.
The global patient node is pseudonymized, not absent
Schema-level absence of a patient entity may seem like the strongest privacy control — you cannot leak what the ontology cannot represent — but it is also the most analytically destructive. Cross-regional duplicate detection in pharmacovigilance (a legal obligation under GVP Module VI), global enrollment monitoring, cross-study exposure analysis, and any agentic workflow that reasons about patient journeys all require a global patient-level join key. CDISC SDTM already mandates USUBJID — a cross-study, pseudonymized subject identifier that FDA requires in every submission — so the sponsor already maintains a global pseudonymized patient identifier by regulatory design.
The architecture follows the pseudonymisation-domain model defined in EDPB Guidelines 01/2025 (paras 35–41): the control is not the absence of the entity but the boundary of the pseudonymisation domain — the guarantee that regional keys never enter the global tier.
Purpose-scoped tokens and two-layer keys
EDPB para 116 warns that a single consistent "person pseudonym" across all purposes carries a "comparatively high" risk of unauthorized attribution. The global KG holds two token families: T_pv = HMAC-SHA256(K_pv, canon(identifiers)) for pharmacovigilance duplicate detection (legal basis: GDPR Art. 6(1)(c) / Art. 9(2)(i); GVP Module VI), and T_res = HMAC-SHA256(K_res_study, canon(identifiers)) for Art. 89 research (enrollment monitoring, cross-study exposure). Cross-purpose linkage (T_pv ↔ T_res) is mediated only by the honest broker, logged, and requires DPO approval.
Two-layer key architecture. The regional partition computes a regional token using a regional key. An honest-broker service — operated in an attested trusted execution environment (TEE), organizationally separate from the analytics platform, reporting to Global Privacy / DPO — re-keys the regional token to the global purpose token. The global KG tier never holds any HMAC secret and therefore cannot compute a token from an identifier (the EDPB para 88 test for a valid pseudonymisation domain).
Re-identification requires threshold key recovery. Regional keys are split via 2-of-3 secret sharing (EDPB para 108): regional DPO, regional QA/PV head, and global privacy office. Re-identification is a human step-up workflow — never an agent action — requiring a documented legal basis and logged in the RoPA/DPIA.
Jurisdictional partition rules
EU/EEA/UK: tokens flow to the global tier; pseudonymisation serves as an Art. 46 supplementary measure per EDPB para 63. US: obtain an Expert Determination (45 CFR 164.514(b)(1)) covering the tokenized global dataset, because HMAC tokens are derived from PHI and fail Safe Harbor's 164.514(c) "not derived from" test — every major US tokenization vendor operates under Expert Determination. China: aggregate-only partition by default; a token export constitutes sensitive-PI transfer under PIPL Art. 38 and triggers a CAC security assessment (≥10,000 sensitive records/year under the March 2024 provisions) plus HGR clearance where genomic-related data is involved. Japan: trial data via APPI consent; NGMIL pseudonymized medical information cannot be provided to third parties — keep Japan RWD in-region. Brazil: LGPD Art. 13 §2 pseudonymisation; onward transfer restricted.
Output controls. Any global-tier agent output undergoes small-cell suppression: k≥11 for externally reportable results (aligned with CMS Cell Size Suppression Policy), k≥5 with complementary suppression as the internal floor. Automated agent-generated aggregates carry differential-privacy noise calibrated per NIST SP 800-226.
Reference architectures: N3C Linkage Honest Broker, PCORnet PPRL framework, German Medical Informatics Initiative federated Trusted Third Party (two-stage pseudonyms across 34 university hospitals), and EHDS Regulation (EU) 2025/327 (Health Data Access Body + Secure Processing Environment + statutory re-identification ban). All converge on the same shape: tokens at the shared tier, identity at the source, an honest broker in between, and a prohibition on user-initiated re-identification.
Agent scope = pseudonymisation-domain admission. Each agent's workload identity carries an explicit list of pseudonymisation domains it may enter (global:T_pv, global:T_res, regional:EU, etc.). Agents admitted to the global domain can read tokens but cannot invoke the honest broker. This aligns with the ICO's January 2026 agentic AI report (controllers remain responsible; restrict which databases an agent can reach) and the Entra Agent ID least-privilege pattern (GA May 2026: task-scoped RBAC, tool allowlists, JIT elevation, explicit provisions for agents accessing regulated data).
Entity resolution across regions
The golden record described above is itself partitioned. Cross-region matching runs on non-PII attributes only — NPI, ORCID, institutional affiliation, product, study identifier, site identifier — so that "Dr. Müller at Charité" resolves to the same global HCP node whether the interaction is logged in Berlin or Boston. PII-bearing attributes (date of birth, home address, prescribing, interaction content) are resolved locally inside the regional partition and never transit.
The merged golden record that the global core holds contains only the shareable fields plus regional pointers — opaque keys that resolve only when dereferenced by an agent running inside that region. This is the same crosswalk discipline the commercial MDM teams already practice, with the residency boundary made explicit in the schema rather than assumed in the deployment.
The golden record splits at the attribute, not the entity
In practice, a single HCP golden record spans at least five classification tiers. Professional attributes — NPI, institutional affiliation, publications, congress attendance, therapeutic focus — are shareable globally under legitimate-interest processing (GDPR Art. 6(1)(f)). Prescribing data is regional and often contractually restricted (the data-provider's license terms may prohibit export and secondary use outside the licensed geography; in some EU member states, prescriber-level data is prohibited outright). Interaction history — field calls, inquiry contacts, speaker engagements — is regional under purpose limitation (Art. 5(1)(b)). Transfer-of-value records range from fully public (US Open Payments / Sunshine Act) to country-by-country regional disclosure (EFPIA Code). KOL tiering scores are shareable as professional assessments, but the interaction data that informed the score stays regional.
The semantic layer must model these as separate attribute groups on one entity, each carrying its own residency tag — not one record with one classification. An agent that reads the tier can run globally; an agent that needs to explain the tier must run regionally.
Proven pattern
Doscierge: 10,292 triplets across 29 domains, 13-field triplet schema, KG-first pipeline. Enterprise Knowledge Vault reference architecture powering 8 production systems; the semantic layer that lifted retrieval accuracy from 75% to 99%+ (on a 212-question evaluation set against expert-labeled ground truth) at a Fortune 50 pharma. Mid-market manufacturing EKG: 1,786 triplets, 47 entity types, 50 predicates, 6 trust tiers with a T4 hard cap at 0.75.
Components in the layer: Core ontology · Domain KGs (seven, above) · Entity resolution & golden record · Process intelligence layer · Trust-tier & provenance model · KG-as-a-Service API (MCP) · Ontology governance agents (drift detection, orphan entities, cross-domain consistency).
The graph does not float. It is anchored to data, and how it is anchored determines whether it is fast, fresh, and inspectable.
5.4 Data Layer
What it is. Controlled, versioned data sources — the enterprise data lake and its data products, the transactional systems of record, external regulatory and evidence feeds — plus the approved-source registry that defines what an agent is allowed to cite. The platform connects to what the enterprise has. It does not replace it.
The enterprise has already invested here. Most large pharma organizations now have a curated data platform spanning ten to thirty years of clinical and preclinical data, a manufacturing data lake fed by historians across dozens of sites and many ERP instances, and a growing set of governed data products with named owners. The semantic layer must sit on top of that investment, not beside it. The question is how.
How the KG and the data lake fit together
The mistake is to treat the KG as a second copy of the lake. It is not. The KG is a semantic index over data products — it holds the reference and normative knowledge persistently (ontologies, regulations, signed rules, specifications, golden records) and materializes operational triplets at run-time from the conformed data products for the entity neighborhood an agent is actually reasoning about.
Read it as a Quality head would. The LIMS of record is in zone 1. The APQR chart's numbers come from the conformed test_results and analytic-ready batch_phase_summary products in zone 2, materialized into the run-time subgraph in zone 3 on query. The KG knew which specification version applied because specifications are joined by effective date on the batch date — never by current version.
Diagram key — the four blocks of the KG-and-lake picture, in text
Block
Contents
Transactional systems of record
MES · LIMS (regional instances) · ERP · QMS · CTMS · EDC · IRT · CRM · PI/OSIsoft historians · PAT vendor PCs
Enterprise data lake
Raw zone (landed as-is, immutable) · Conformed zone — data products with named owners: batches · unit_operation_phases · parameters · samples · tests · test_results · specifications · genealogy · raw_material_results (CoA/CoC) · subjects · sites · deviations · capa · hcp_master · Analytic-ready zone — batch_phase_summary, batch_quality_profile, enrollment_curves, ...
Data fabric / virtualization
(canonical attribute dictionary over regional LIMS, site historians, CMO feeds)
SEMANTIC LAYER — KG-as-a-Service
Persistent: ontologies, rules, regulations, specifications, golden records, mined process-intelligence edges · Run-time materialized: batch → phase → parameter → result → spec → genealogy → material lot → CoA for THIS batch, projected from data-product rows with row-level provenance + data-product version
Three consequences of this design:
Freshness without duplication
The APQR agent asking about batch B-2026-0417 does not read a stale graph copy of the batch. It reads the batch_phase_summary and test_results rows for that batch, projected into triplets on demand, each carrying the row's provenance and the data product's version. The persistent graph supplies the meaning (which parameter is a CPP, what the validated range is, which specification version was effective on the batch date); the lake supplies the facts.
Data fabric earns its place where systems are federated
Regional LIMS instances, site-specific historian tag names, CMO results arriving as spreadsheets — a virtualization layer with a canonical attribute dictionary presents one logical test_results service over all of them. The dictionary is quality-owned work, under change control, joined to specifications by effective date, never by current version. It is the single most underestimated line item in a manufacturing AI program, and the one that makes cross-site comparison a GROUP BY site instead of a heroic project.
Transactional systems keep their authority
The MES remains the system of record for the batch; the LIMS for the result; the QMS for the deviation. Agents read from data products and write decisions back to the workflow tool the human uses — never directly into a validated transactional system. That boundary is what keeps the agent out of the transactional system's validation scope.
Batch genealogy, LIMS quality parameters, and historian process parameters are the three data products that manufacturing use-cases cannot live without. Genealogy resolves lot relationships (drug-substance lot to drug-product lots, raw-material lot to batch); the LIMS conformed services (samples → tests → test_results, given meaning by specifications) carry the CQAs; the historian-derived batch_phase_summary carries the CPPs. Join them and the CPP-to-CQA relationship — the scientific point of APQR and CPV — becomes a query. §8.4 is built on that join.
Classification-aware data fabric — residency is enforced where the data is served
The fabric described above unifies heterogeneous systems. It must also partition regulated ones, and the two jobs are done by the same layer so that there is one place to inspect. Every data product carries a residency classification alongside its owner and contract, and the fabric enforces it at the point of service — not at the agent, not at the network perimeter.
Classification
What it covers
Where it lives
How the fabric serves it
Regional
Patient-level data (subjects, ICSRs, genomic results), HCP interaction records, prescribing and claims-derived data, MSL note content, inquiry text, employee data
The regional zone of origin — Azure West Europe, East US 2, China North, Japan East, Brazil South — with no cross-border replica, including for DR (DR pairs stay inside the jurisdiction)
Served only to agents whose card declares a matching regional scope (§6). A global agent's request triggers one of two paths: denial (default), or automated de-identification + human attestation — HIPAA Safe Harbor / Expert Determination for US data, GDPR Recital 26 anonymization standard for EU data, with a named data steward signing that the output is no longer personal data before it is released
Restricted
Genomic and genetic data (GDPR Art. 9(1) explicit enumeration), biometric identifiers, Chinese patient and HCP data subject to PIPL Art. 38–40 — any data category where REGIONAL classification plus de-identification is insufficient because the law or the ethics-committee approval is site-specific and non-transferable
Same residency zone as REGIONAL — no cross-border replica
REGIONAL enforcement plus a mandatory second gate: ethics-committee approval with per-site scope, specific ICF consent language covering the derivative use, or a CAC security assessment with versioned reference. The distinction from REGIONAL is the exit path: REGIONAL data has a de-identification pathway to SHAREABLE-DERIVED; RESTRICTED data may have no pathway at all — Chinese genomic data under HGRAC is non-exportable, and German genetic data under GenDG requires per-subject ethics approval that does not transfer across sites
Shareable
Manufacturing quality (batches, results, specifications, genealogy), aggregated trial metrics (enrollment curves, site performance), HCP professional master data, regulatory intelligence, publications, ontologies, signed rules
Replicated globally; single logical service
Standard platform contract; classification tag travels with the row
Shareable-derived
Outputs of a de-identification or aggregation step over REGIONAL inputs (insight themes, signal statistics, cohort counts)
Global
Served globally, but the row carries the lineage edge to its REGIONAL source and the attestation record; re-identification queries are blocked by the k-anonymity floor at the fabric
Mandated-flow
ICSRs (E2B(R3) format), SUSARs, annual safety reports — data that must cross borders under pharmacovigilance law regardless of privacy restrictions
Originates regionally; transmitted to health authorities (EudraVigilance, FAERS, NMPA) via a separate safety-database gateway, never via the general-purpose data fabric
Dedicated pipeline through the validated safety database and its E2B(R3) submission gateway. The fabric does not touch it. The classification exists to prevent anyone from routing a case through the fabric "for convenience" — the moment an ICSR enters the fabric, the separation of duties that inspectors expect is broken
Three rules follow from the table:
Logical unification, physical separation
Regional LIMS, historian, and ERP instances are unified through the canonical attribute dictionary so that test_results is one logical service — but the unification is virtual. The fabric pushes the query down to each regional instance and joins the projections; it does not copy classified rows across a border to make the join convenient. For manufacturing data this is rarely a constraint (batch results are SHAREABLE). For the clinical and medical data products it is the constraint, and the reason the fabric — not a replication pipeline — is the integration pattern.
Consent is a checkpoint, not a column
For GDPR Art. 9 special-category data (health, genetic, biometric) and for any HCP data used beyond its collection purpose, the fabric calls the consent management platform before serving a row to any agent, and serves only rows whose consent scope covers the agent's declared purpose (from its card). A revoked consent is effective at the next query, not at the next batch refresh. The consent check is itself logged (G3) with the purpose evaluated, so that an Art. 15 subject-access request can be answered from the audit trail.
China is an asymmetric zone
PIPL localizes personal information and "important data" inside China; export requires a CAC security assessment (PIPL Art. 38, the 2024 cross-border data-flow provisions), and human genetic resources cannot leave under HGRAC rules at all. The fabric therefore treats the China North zone as REGIONAL-with-no-derivation-path by default: even SHAREABLE-DERIVED outputs require an explicit, versioned CAC assessment reference on the data product before the fabric will serve them outside the zone. Every other regional zone has a de-identification path to global; China's path is an approval, not an algorithm.
Contractual restrictions carry equal force. Data residency enforcement is not purely a privacy-law question. Prescribing data licensed from a third-party data provider (the industry's analogs of IQVIA and Symphony Health) typically prohibits export and secondary use outside the licensed geography under the data-use agreement — a contractual, not regulatory, restriction that the fabric must enforce identically. In some jurisdictions the restriction is harder than contractual: France prohibits prescriber-level prescribing databases for commercial promotion (CSP Art. L.4113-7), which means the fabric's classification for French prescribing is not REGIONAL for contractual reasons but RESTRICTED — the data exists only in aggregate form, and the agent that needs individual prescriber context for a French HCP must infer it from other signals (publication record, congress attendance, inquiry history) rather than from a dataset that does not exist. Similarly, CRO-provided site-performance data may carry confidentiality terms that restrict cross-sponsor use. The fabric's classification tag must therefore encode both legal basis and contractual basis, and the routing logic must check both. A data product tagged REGIONAL for GDPR reasons and one tagged REGIONAL for DUA reasons look the same to the agent — which is the point: the enforcement is uniform, the reason is metadata.
The approved-source registry
Nothing outside the registry can be cited by an agent. It is versioned and access-controlled, and it is scoped per use-case: the regulatory-documentation agent's registry contains locked TFLs, the approved protocol, and the IB; the field-force agent's registry contains MLR-approved assets only; the medical-inquiry agent's registry contains medical-affairs-approved responses. Same mechanism, different contents — which is the whole point of a platform.
Components in the layer: SDTM/ADaM clinical data · Regulatory feeds (FDA/EMA/ICH/WHO) · Real-world evidence · CTMS/EDC/IRT · CRM/MDM · QMS/LIMS/MES · Historians & PAT · ERP/supplier data · Enterprise data lake & data products · Data fabric / virtualization · Approved-source registry.
Every layer described so far is only trustworthy because of the plane that cuts across them.
5.5 Cross-Cutting Governance Plane
Governance in v1 of this architecture was a list of boxes. That is how governance usually gets drawn, and it is why governance usually gets ignored: a flat list gives an executive no way to ask "which of these do we actually have, and which ones matter for this use-case?" The plane is better understood as eight categories, each answering one question an inspector, a Quality head, or a CFO will eventually ask. Click a card for the full control list; each §8 use-case card carries a compact G1–G8 strip showing which categories apply at what intensity.
G1 · PROVENANCE
Evidence Provenance
Where did this come from, and how much should I trust it?
Approved-source registry · Span-level citations on every generated sentence · Triplet-level provenance (source, tier, confidence, timestamp, version) · Trust-tier caps on inferred edges.
Primary owner: Semantic & Knowledge Engineering
G2 · PROVENANCE
Decision Provenance
What did the human do with it, and why?
Accept / modify / rejectRationalePreference signalCoaching record (acceptance curve)
+ full control list
Accept / modify / reject capture on every recommendation · Rationale capture at sign-off · Preference signal from edits · The coaching record: the acceptance-rate curve that decides whether an agent graduates (§18).
Primary owner: Agentic Factory + domain owners
G3 · AUDIT
Traceability & Audit
Can you reconstruct this decision for an inspector in fifteen minutes?
21 CFR Part 11EU Annex 11E-signaturesExecution historyALCOA+Object lock
+ full control list
21 CFR Part 11 audit trail · EU Annex 11 electronic records · Electronic signatures · Per-execution history in stateful orchestration · ALCOA+ on every artifact · Immutable storage (object lock).
Primary owner: Quality / CSV
G4 · OBSERVABILITY
AI Observability
Is the system behaving today the way it behaved when we validated it?
Distributed tracing per agent invocation · Token and cost accounting per agent, per use-case, per domain · Latency and decision-velocity dashboards · Model drift and eval-score drift detection · Confidence calibration monitoring (reliability diagrams, Brier score).
Primary owner: Agentic Factory / Platform SRE
G5 · ACCESS
Identity & Access
Who — or what — is allowed to touch this?
Agent as principalRBACTool allowlistData classificationPartner onboarding
+ full control list
Agent as first-class identity principal · Role-based access control · Tool allowlist per agent identity · Data classification enforcement · Partner-agent onboarding under the platform contract.
Primary owner: Security / IAM
G6 · VALIDATION
Validation & Change Control
How do we know it works, and what happens when it changes?
GAMP 5 classification per component (Cat 5 for deterministic paths) · CSA-aligned controls for LLM paths · Model version pinning · Rule versioning with owner and calibration record · Release trains with regression gates · Change impact on validated state.
Would we be comfortable explaining this to a regulator, a patient, or a journalist?
Bias monitoringPrivacyAgent cardsGuidance trackingPSMFAnnex III classify-upDPIA per agentResidency scope
+ full control list
Bias monitoring (site concentration, population representation) · Privacy controls (PHI, GDPR) · Model cards and agent cards in the registry · Regulatory-guidance tracking (FDA AI draft guidance, EMA reflection paper, EU AI Act classification) · AI use documented where required (PSMF for PV) · EU AI Act Annex III classify-up policy for agents touching trial participants, health-data profiling, safety-critical PV, or HCP profiling (Art. 22 GDPR) · DPIA-per-agent (GDPR Art. 35) as a Lab-to-Factory gate · Regional compliance as configuration — jurisdictional requirements as versioned KG edges, not code forks · Residency classification enforced at the fabric (§5.4) and on the agent card (§6).
Primary owner: AI Governance Council
Three categories deserve emphasis because they are the ones most architectures leave out.
Responsible AI (G8) is a set of gates, not a set of principles
Three of them are structural.
Annex III classify-up. The EU AI Act's Annex III list (Art. 6(2)) and its Commission guidelines will keep moving; an architecture that waits for the final word is one that re-validates every agent when it lands. The platform's policy is to treat as high-risk, by default, any agent in four classes: protocol-design and trial-design agents that shape which participants are enrolled and what they are exposed to; health-data profiling agents that score or segment patients or subjects (the clinical-operations intelligence class); safety-critical PV case-processing agents (UC-X1); and commercial agents that profile HCPs in a way that produces legal or similarly significant effects — which brings GDPR Art. 22 with it regardless of the AI Act's outcome. High-risk means the Art. 9 risk-management system, Art. 10 data-governance record, Art. 12 logging, Art. 14 human-oversight design, and Art. 26 deployer obligations are all satisfied by the platform contract — the confidence bands (G7), the traceability plane (G3), the agent card (§6) — so that classification is a flag on the card, not a project.
DPIA per agent. Every agent whose card declares a PII scope or a REGIONAL-* residency scope (§6) carries a completed Data Protection Impact Assessment before it can transition candidate → factory. The DPIA template is itself a platform artifact: agent pattern and kit; data categories and Art. 9 flags; legal basis per jurisdiction (consent, legitimate interest with balancing test, public-health exemption — as a table, one row per region the agent runs in); cross-border transfer mechanism per region pair (adequacy decision, SCCs with transfer-impact assessment, BCRs, CAC assessment reference); residual-risk assessment with the DPO's sign-off; and the re-assessment trigger — which changes to the card re-open the DPIA. The four-lane approval workflow (§6) gains a fifth reviewer, the DPO, for these cards.
Regional compliance as configuration. The regulatory ontology in the core KG holds jurisdictional requirements as versioned edges — requirement ──applies_in──► jurisdiction, requirement ──effective_from──► date, requirement ──supersedes──► prior_version — with the same trust-tier and provenance discipline as any other triple. A protocol-compliance agent checking a study with sites in Germany, Ohio, and Shanghai does not run three code paths; it traverses the regulatory ontology for each site's jurisdiction and applies the rules it finds, each carrying its citation. When PIPL's implementing rules change, the ontology team publishes a new edge version; no agent is redeployed. This is the mechanism behind the §8.7 posture that rulesets are "configuration tagged by jurisdiction, never regional code forks."
Decision provenance (G2) is the coaching mechanism. Evidence provenance tells you what the agent saw. Decision provenance tells you what the expert did about it — and the aggregate of those decisions is the only honest measure of whether an agent is helpful. An agent whose recommendations are accepted 40% of the time is not ready, no matter how good its eval scores. An agent whose acceptance rate climbs monotonically over rolling two-week windows is learning the domain's boundaries. G2 is what turns human review from a cost into a signal.
AI observability (G4) is what makes the Factory a factory. Without per-agent, per-invocation tracing and cost accounting, the platform cannot tell which use-cases are paying for themselves, cannot detect drift before Quality does, and cannot answer the CFO's question. Observability is not a dashboard; it is the instrumentation that makes the other seven categories measurable.
The governance plane applies to every agent. The registry is where each agent declares which parts of it apply — including the one class of work where the model, not the query, has to travel.
5.6 Federated Learning — When the Model Must Travel Because the Data Cannot
Everything above the line keeps patient data inside its region by routing queries. There is one class of work that routing does not solve: training a model whose signal is spread across regional patient populations — a safety signal that is sub-threshold in every region and significant across them; a biomarker whose prevalence differs by ancestry; a recruitment-density model that needs every region's site history. For that class, the platform inverts the usual pattern: the data stays, the model moves.
AGGREGATION HUB — Basel (Azure West Europe)receives: DP-noised, encrypted model updates only
holds: global model weights · per-node privacy ledger · round audit trail
never holds: a patient record, a gradient in the clear
▲ ▲ ▲ ▲ ▲
secure aggregation (SMPC or TEE): updates summed without any party decrypting an individual one
│ │ │ │ │
┌───────────────┴──┐ ┌───────┴─────┐ ┌────┴──────┐ ┌───┴────────┐ ┌─┴──────────┐
│ EU node │ │ US node │ │ JP node │ │ BR node │ │ CN node │
│ West Europe │ │ East US 2 │ │ Japan East│ │ Brazil S. │ │ China North│
│ trains on EU │ │ HIPAA PHI │ │ APPI │ │ LGPD │ │ PIPL/HGRAC │
│ patient data │ │ │ │ │ │ │ │ ASYMMETRIC │
│ exports: noised │ │ same │ │ same │ │ same │ │ local-only │
│ gradients, ε- │ │ │ │ │ │ │ │ by default;│
│ metered │ │ │ │ │ │ │ │ opt-in per │
└──────────────────┘ └─────────────┘ └───────────┘ └────────────┘ │ CAC assmt. │
└────────────┘
Four controls make it defensible, and each is auditable (G3):
Patient data never leaves the zone
Each regional training node runs inside its residency zone against the REGIONAL data products the fabric serves it (§5.4), under a card with REGIONAL-* scope (§6). What crosses the border is a model update — and only after differential-privacy noise is added at the node, before export. The hub is FEDERATED-scoped: it may receive aggregates, never records.
Privacy budget is metered per node
Every export spends epsilon; the hub's privacy ledger tracks cumulative ε per node per model lineage, and the node's DPIA (G8) sets the ceiling. When a node's budget is exhausted its contributions stop — automatically, not by policy reminder — until the DPO approves a new budget with a documented rationale. The ledger is what the answer to "how much have we learned about EU patients, in total" looks like when a supervisory authority asks.
Secure aggregation, so the hub cannot see an individual update
Updates are encrypted at the node and summed under secure multi-party computation or inside a trusted execution environment (Azure confidential computing), so that the hub decrypts only the sum. A single region's update — which can leak more than its noise budget suggests when the region is small — is never observable on its own.
China participates asymmetrically
PIPL localization and HGRAC make the China North node a fully local training node by default: it trains a local model on local data and contributes nothing to the hub. Participation in a federated round is an opt-in per model lineage that requires a CAC security assessment reference on the export, recorded on the node's card; the noised gradient is treated as a data export for that assessment, because the regulator will. The architecture does not assume China joins; it makes joining a configuration change rather than a redesign when the assessment is in hand.
Where the platform uses it
Use-case
Regional nodes train on
Hub learns
Why routing alone was not enough
Federated signal detection (UC-X1)
Regional ICSR databases
A disproportionality model that surfaces signals sub-threshold in every single region
Aggregated counts per region lose the case-level covariates (age, co-medication, indication) that make a weak signal credible
Federated patient-density estimation (UC-D2)
Regional claims-derived and EHR-derived cohort data, regional site history
A site-selection model that predicts enrollment from eligible-patient density
Density is a patient-level attribute; the regional partitions cannot export it, and site history without it is what makes site selection a master-data problem rather than a prediction problem
Federated biomarker analysis (UC-R1, UC-X6)
Regional genomic datasets — including China's, which cannot export under HGRAC
Biomarker-response associations across ancestries
Genomic data is the most localized data class the enterprise holds; a model trained on one region's ancestry mix is the bias the G8 monitoring exists to catch
The registry (§6) treats a federated model lineage as an agent: it has a card, a residency scope of FEDERATED, a DPIA, a privacy ledger, and a lifecycle. That is what makes the next section's discipline apply to it unchanged.
Section 6 · Agent Registry & Service Catalog
From "we built an agent" to "the enterprise has a capability"
A platform without a registry is a folder of prototypes. The registry is what turns "we built an agent" into "the enterprise has a capability" — discoverable, composable, governed, and billable. It is also the mechanism that makes Lab-to-Factory promotion (§7) a state transition rather than a re-implementation.
What is registered
Every agent, every MCP tool server, and every reusable service kit is registered with a signed agent card — the machine-readable contract that other agents, workspaces, and reviewers read.
Agent card — illustrative
InternalConfidentialPII← declared on every card at design time
# Agent card — illustrativeid:mfg.cpv.signal-agentversion:2.3.0archetype:monitoringowner:MS&T Data Products (name, role, escalation)domain:manufacturingpurpose:"Detect SPC rule violations and capability drift on batch_phase_summary; route to responsible scientist"gxp_class:GMP / Part 11 (selective — signal disposition is a regulated record)lifecycle_state:factory# draft | lab | candidate | factory | deprecatedservice_signature:inputs:
- data_product:mfg.batch_phase_summary@>=4.1
- data_product:mfg.test_results@>=2.0
- kg_scope:manufacturing.cpv (read)outputs:
- signal:cpv.rule_violation (schema v1.2, with provenance)
- signal:cpv.capability_drifttools_visible_to_model:[spc_calculator, capability_calculator, kg_lookup]# ≤5model:scoped-small@pinned-2026-07confidence_policy:{autonomous: ">=0.90", notify: "0.70-0.89", suggest: "0.55-0.69", escalate: "<0.55"}eval_record:harness:mfg.cpv.eval@1.4ground_truth:2,140 labeled batch-phases, 3 sites, 2 productsprecision:0.93recall:0.88f1:0.905last_run:2026-08-29coach_mode_acceptance:0.81(rolling 14d, n=212)governance:evidence_provenance:G1 (triplet + row-level)decision_provenance:G2 (disposition captured)audit:G3 (Part 11 trail on disposition)observability:G4 (traced, cost-tagged)identity:G5 (principal mfg-cpv-signal; allowlist above)validation:G6 (GAMP 5 Cat 5 for SPC path; CSA for narrative)autonomy:G7 (policy above; degraded mode on historian connector loss)responsible_ai:G8 (agent card published)approvals:security_review:approved 2026-07-14 (ticket, reviewer)quality_review:approved 2026-08-02 (CSV lead)data_steward:approved 2026-07-20 (MS&T data-product owner)business_owner:approved 2026-08-05 (Head of MS&T)usage_and_billing:invocations_30d:18,420cost_30d:(tokens + compute, tagged to cost center MFG-QA)consumers:[mfg.apqr.orchestrator, mfg.quality-command-center, mfg.deviation.investigator]
Lifecycle state machine
Securityidentity · allowlist · data classification
QualityGxP class · validation posture · audit trail
Data stewarddata-product contracts · approved-source scope
Business ownerpurpose · metric · adoption
The four lanes above are the review workflow described under "Approvals are a workflow, not a wiki page" below.
Service kits — the templatized unit of reuse
Individual agents are too fine-grained to be the unit of reuse. The unit of reuse is the service kit: a templatized bundle that a new use-case instantiates rather than builds.
Kit component
What it contains
Why it accelerates use-case N+1
Data contract
Named data products consumed, schema versions, freshness SLAs, approved-source registry scope
The new team spends days mapping, not months integrating
Orchestrator template
Stateful workflow skeleton (intake → specialists → merge → human gate → decision log), with hooks pre-wired for audit and observability
Control flow is deterministic and already validated once
Specialist slots
Typed slots for extraction / reasoning / drafting / monitoring specialists, each with the ≤5-tool constraint enforced
Composition is the default; the monolith is structurally impossible
Rule-gate kit — expert-calibrated rule engine with provisional → confirmed calibration loop, versioned rules with owners. Instantiated for protocol compliance, structural alerts in chemistry, seriousness/expectedness in PV, batch-record completeness.
Signal-and-NBA
Signal-and-NBA kit — isolated monitoring specialists + orchestrator that de-duplicates, prioritizes, composes next-best-actions with rationale/impact/owner, and suppresses noise from feedback. Instantiated for trial operations (§8.2), CPV (§8.4), supply risk (§8.5), and every command center (§12).
Evidence-dossier
Evidence-dossier kit — KG traversal + deterministic scoring tiers + cited synthesis → reviewable dossier. Instantiated for target validation (§8.1), site selection (§8.2), supplier qualification (§8.5).
Entity-resolution
Entity-resolution kit — cross-source matching, golden record, crosswalk maintenance. Instantiated for HCP/HCO, site/investigator, material/supplier, batch genealogy.
Kits are not only discovered — they are used, and usage requires knowing what an agent is allowed to touch. Every agent card declares a scope classification at design time, analogous to how datasets are classified but promoted to the agent as a first-class property:
Scope
What the agent can access
Runtime enforcement
Example
Internal
All platform data products within the agent's domain partition; broadest tool access; minimal friction
Standard platform contract (G5 identity, G3 audit)
A Lab PoC agent exploring CPV trends across all sites
Confidential
Restricted to named data products; tool access narrowed; outputs carry a classification tag
Data-product ACL check on every retrieval; output tagged for downstream handling
A deviation agent accessing supplier quality records under NDA
PII
Patient, HCP, or employee data — HIPAA, GDPR, or equivalent applies
De-identification gate before any output leaves the agent; consent check on retrieval; no caching of raw PII in agent state; audit trail includes data-subject category
A PV agent processing individual case safety reports; an HCP engagement agent accessing prescribing history
The classification is declared on the agent card, reviewed during the four-lane approval workflow (§6), and enforced at runtime by the governance plane (G5 identity + G1 approved-source + G3 audit). An agent classified Internal that attempts to access a PII data product is denied — the boundary is structural, not policy. This is the same principle that makes dataset classification work, extended to the entity that reasons over the data rather than the entity that stores it.
Scope escalation in multi-agent composition. The design-time classification is clear; what requires an explicit rule is how scope transitions when agents compose. When a Confidential-scoped orchestrator invokes an Internal-scoped specialist and passes it context that contains Confidential data, the specialist's effective scope elevates to the caller's scope for the duration of that invocation only. The elevation is not a grant — it is a delegated, time-boxed inheritance, and it is the audit event that matters.
The audit trail for every elevated invocation captures six fields: (1) caller identity — orchestrator agent ID, version, service principal, and the human session (if any) that triggered it; (2) callee identity — specialist agent ID and version, with its baseline scope recorded alongside the elevated scope; (3) timestamps at invocation and return (contemporaneous — ALCOA+); (4) structured purpose code from the DPIA's purpose register (G8), not free text; (5) context hash — a hash of the context payload, so an inspector can verify whether Confidential data actually crossed the boundary or only a reference to it did; (6) region binding — the elevated invocation must execute in the same region as the caller. A West Europe orchestrator cannot elevate an East US 2 specialist; the fabric rejects the call.
Three consequences follow. First, specialists must be deployable in every region where their orchestrators run, even if they only need Internal data — because elevation requires co-location. Second, elevated invocations are a first-class metric in the EU AI Act Art. 12 logging requirement and the Art. 72 post-market monitoring plan; a spike in elevations is a design smell — the orchestrator is passing more Confidential context than the specialist's task requires (GDPR Art. 5(1)(c) data minimization). Third, the DPIA for the orchestrator (G8) must enumerate every specialist it can elevate, with the purpose codes it may declare. Adding a new specialist to an orchestrator's call graph is a DPIA change-control event.
Regional scope enforcement — a second axis on the card
Beyond data classification (Internal / Confidential / PII), every agent card declares its deployment region and data residency scope. The two axes are independent: a PII agent may be REGIONAL-EU; an Internal agent may be GLOBAL. The residency scope is what the fabric (§5.4) and the regional KG endpoints (§5.3) check before serving a row.
Residency scope
Where the agent runs
What data it may access
Cross-border output
GLOBAL
Any region
Non-PII data products only — reference data, ontologies, signed rules, SHAREABLE and SHAREABLE-DERIVED products, aggregated metrics
Unrestricted
REGIONAL-EU
Azure West Europe only (North Europe for DR)
EU patient data, EU HCP interactions, EU clinical and safety data
Outputs containing personal data stay in the EU; de-identified outputs (GDPR Recital 26 standard, steward-attested) may be promoted to SHAREABLE-DERIVED
REGIONAL-US
Azure East US 2 only
US patient data (HIPAA PHI), US HCP interactions, US prescribing
Same mechanism — Safe Harbor or Expert Determination before promotion
REGIONAL-CN
Azure China North only
All personal information and important data of Chinese origin (PIPL); no genomic export (HGRAC)
No cross-border output of any kind without a CAC security-assessment reference on the output data product
REGIONAL-JP
Azure Japan East only
Japanese patient data (APPI), J-GCP clinical data, Japan HCP interactions, PMDA submissions
Exit via EU–Japan mutual adequacy for EU-bound transfers; APEC CBPR for other destinations; de-identified outputs follow the same steward-attestation path as EU data
REGIONAL-BR
Azure Brazil South only
Brazilian patient data (LGPD Art. 11 sensitive health data), ANVISA-regulated clinical data, Brazil HCP interactions
Exit via ANPD standard contractual clauses (finalized August 2024; grace period ended August 2025) or de-identification; CONEP ethics committee oversight on research-use data
FEDERATED
Orchestrator runs globally; data access happens regionally
The orchestrator itself accesses only GLOBAL data; it dispatches sub-queries to REGIONAL-* specialists and receives aggregates that satisfy the k-anonymity floor
The orchestrator never holds an individual record; the audit trail shows the sub-query, the region, and the classification decision each regional specialist made
The FEDERATED scope is the architecturally interesting one, and the one a multinational cannot do without. A federated signal-detection orchestrator (UC-X1) runs in Basel and asks the same question of the EU, US, China, and Japan PV specialists; each specialist processes individual cases inside its region and returns disproportionality statistics — a reporting odds ratio, a case count above the floor, a MedDRA PT — not case data. The orchestrator sees a signal, not a patient. The same scope serves federated patient-density estimation for site selection (UC-D2) and federated theme aggregation from field medical insights (UC-X3).
Composition rules follow. A FEDERATED orchestrator may invoke REGIONAL-* specialists; a REGIONAL-EU agent may not invoke a REGIONAL-US one; a GLOBAL agent may invoke neither. The registry refuses the composition at discovery time, before any data moves — which is the point of putting the scope on the card rather than in a policy document.
Agents find agents
The registry exposes discovery and usage as an API. An orchestrator planning a workflow queries the registry for specialists by capability, domain, GxP class, scope classification, and eval score — and composes from what it finds. A deviation-investigation orchestrator does not implement CPV trend lookup; it discovers mfg.cpv.signal-agent, checks that its lifecycle state is factory, its scope classification is compatible, and its eval record is current, and invokes it through the platform contract. The discovery call and the invocation are both traced (G4) and both subject to the caller's allowlist (G5).
This is how the registry enforces composition: an orchestrator may only invoke registered specialists, and a specialist may only expose the tools on its card. There is no path by which an agent grows a hundred tools, because the card is the contract and the contract is reviewed.
Approvals are a workflow, not a wiki page
Lifecycle transitions — draft → lab → candidate → factory, and any change to a factory card — trigger a review workflow with four named reviewers: security (identity, allowlist, data classification), quality (GxP class, validation posture, audit trail), data steward (data-product contracts, approved-source scope), and business owner (purpose, metric, adoption). A card cannot reach factory with any approval missing, and a change to a factory card's service signature re-opens the relevant reviews. The workflow is where the human team reviews the agent's security and quality posture — not in a meeting, but against a card that says exactly what the agent can touch and how well it performs.
Usage tracking and chargeback
Every invocation is metered — tokens, compute, tool calls, latency — and tagged to the consuming use-case, domain, and cost center. This closes three loops at once:
Economics
The CFO can see cost per decision, per use-case, and whether the marginal cost of use-case N+1 is actually falling (§18).
Reuse measurement
The registry knows which kits and specialists are consumed by how many use-cases — the reuse ratio is a query, not an estimate.
Retirement
An agent with no consumers for ninety days is a candidate for deprecated. The registry is how the curated garden stays curated.
Proven pattern
MCP server published on the MCP registry (Linux Foundation AAIF, December 2025) — the interoperability contract partner agents and internal agents both speak. 18+ reusable infrastructure modules underlying 8 production systems: a new system stands up in days because the kit already exists.
With the registry in place, "promotion to production" becomes a defined state transition with defined reviewers. That is the operating model.
Section 7 · Lab → Factory
The operating model: people, process, technology — in that order
Every pharma executive knows the mantra: people, process, technology — in that order. Agentic AI programs fail when they invert it. The Lab-to-Factory model is, first, a people model: who comes together, for what purpose, with what authority. Then a process model: how a use-case moves from idea to governed production, and what the support commitments are on the far side. Only then a technology model — and the technology (§5, §6) is designed to make the people and process model cheap to run.
Three constituencies, one virtual pod, one gate. The pod forms around a use-case, runs through Lab and the gate, hands off to the Factory support model, and dissolves. The readiness vote belongs to the domain team, the validation vote to Quality, the funding vote to the business sponsor; any one can hold.
7.1 People — three teams, one virtual pod
Every use-case is delivered by a virtual pod that brings three constituencies together for a purpose and dissolves when the purpose is met. The pod does not need to be co-located, does not need to be a permanent org unit, and does not need to be large. It needs to be named.
Constituency
Who
What they own in the pod
What they must not be asked to do
Business team
The sponsor and the decision-makers whose calendar the use-case compresses — the Head of MS&T, the clinical-operations lead, the medical-writing head
The decision, the metric, the acceptance criteria, the promotion vote
Design the agent, tune the prompt, evaluate the model
Functional / domain team
The subject-matter experts who will use the agent — process scientists, toxicologists, medical writers, PV physicians, QA leads
Co-design of the agent's behavior; rule calibration (provisional → confirmed); ground-truth labeling; coach-mode review; the promotion gate itself
Instantiate the kit; build the specialists; wire the data contracts; run the eval harness; produce the validation package
Decide whether the agent is ready — that is the domain team's call
The pod is led by a Forward-Deployed Engineer (FDE) — an engineer who sits with the domain team, speaks the domain's language, and owns the translation between "what the toxicologist needs" and "which specialist, which kit, which data product." The FDE model is what makes the Lab fast: instead of requirements documents crossing an organizational boundary, one person carries the context in both directions. It is also what makes the Factory adoptable: the FDE who built the agent is the person the domain team trusts when it goes live.
Three constructs make the pod work at enterprise scale:
Pods are purpose-scoped and time-boxed. A pod forms around a use-case, runs through Lab and the promotion gate, hands off to the Factory support model, and dissolves. Members return to their home teams with the pattern in their heads. The next pod forms faster.
Pods share a bench. Knowledge engineers, CSV engineers, and eval specialists are scarce. They are pooled centrally and allocated to pods by the platform leadership, so that the ontology decision made in the CPV pod is consistent with the one made in the deviation pod.
Pods are virtual by default. The manufacturing SME is in one country, the FDE in another, the data-product owner in a third. The registry (§6), the kit, and the eval harness are the shared workspace; the pod meets in the artifacts, not in a room.
7.2 Process — what the Lab looks like from a use-case perspective
The Lab is not a research group. It is a rapid-prototyping environment whose job is to get from a named decision to a measured PoC in two to four weeks, and to kill PoCs that are not moving the metric.
What a pod finds when it walks into the Lab:
Lab tool
What it does
Time saved
Kit instantiation
Pick the closest service kit (§6); the orchestrator skeleton, eval harness, governance manifest, and experience shell are generated
Weeks of scaffolding → hours
KG sandbox
A branch of the semantic layer where the pod can extend the domain ontology, load candidate triplets, and test traversals without touching production
Ontology iteration in days, with a merge path back to the federated model
Data-product catalog
Browse conformed data products, request access, bind the data contract
No bespoke extracts
Prompt & agent playground
Versioned system prompts, tool binding, context-budget visualization, side-by-side model tier comparison (frontier vs scoped — the 137× question answered per task)
Model-tier decisions made on evidence
Synthetic and historical ground truth
Curated historical cases with expert labels; synthetic-data generators that mirror real distributions (not toy data) for domains where history is thin
Eval harness has something to measure against from week one
Eval harness template
P/R/F1 per class, confidence calibration, acceptance-rate tracker — pre-wired
"Does it work?" has a number
Coach-mode deployment
Ship to a small group of domain experts with thumbs-up/down on every output, in the real workspace shell
Adoption signal before the Factory spends a dollar
Lab cadence and discipline
2–4 week iteration cycles per PoC, against a real historical dataset and a real named decision.
Evaluation is logged, not validated. The Lab records everything (G3, G4 are on from day one) but does not carry the validation burden — that is what the gate is for.
Kill criteria are explicit. If precision/recall/F1 does not improve over two iteration cycles, or coach-mode acceptance is flat, the PoC stops. The pod writes a two-page post-mortem into the registry so the next pod does not repeat it. Killing fast is a Lab success metric, not a failure.
Every PoC produces a card. Even a killed PoC leaves a deprecated card with its eval record — institutional memory, not a lost notebook.
The demand test is a kill criterion too. Before the first iteration, the pod classifies the demand the use-case serves — how much of today's calendar is value work, Type 1 waste, Type 2 waste, and queue time — and names the system condition behind each. The PoC must then state which of those it removes and which it merely absorbs faster. An agent that assembles the same scattered evidence more quickly, leaving it scattered, copes with the condition; an agent that leaves the precedent structured in the KG for the next decision changes it. Both can be legitimate. But a PoC that can only cope, in a domain where the condition is cheap to change — a template field, an SOP step, a named data-product owner — does not proceed. The Lab does not industrialize a workaround.
7.3 Process — when and how a use-case transitions to Factory
The promotion gate is the hardest organizational problem in the model, and it deserves to be stated plainly: the Lab wants to ship, Quality says it is not validated, and the business sponsor wants the value now. Without defined criteria, a named owner, and an escalation path, PoCs accumulate and nothing reaches production. Think of it as transforming a meadow of wildflowers — beautiful but ungoverned — into a curated garden where every plant has a purpose, a place, and a caretaker.
Gate criteria (all required):
Criterion
Threshold
Who verifies
Eval record
P/R/F1 per class against expert-labeled ground truth, above the use-case's stated threshold; confidence calibrated
Tech team produces; domain team accepts
Coach-mode acceptance
≥70% sustained over rolling two-week windows, with a rising curve
All consumed data products at a versioned, owned, SLA-backed release
Data stewards
Composition compliance
Every specialist ≤5 tools; orchestrator control flow deterministic; cards complete
Platform architecture
Security review
Identity, allowlist, data classification approved
Security
Named owner and support model
Business owner, technical owner, on-call rotation, SLA agreed
Business + Factory
Demand baseline
Value/failure-demand split measured (not estimated) before coach mode; the card names which system condition the agent changes and which it only absorbs
Domain team + FDE
Gate ownership. The gate is chaired by the Factory lead, but the readiness vote belongs to the domain team, the validation vote belongs to Quality, and the funding vote belongs to the business sponsor. Any one of the three can hold; escalation is to the platform's steering committee with a two-week SLA. The Lab lead does not vote on their own PoC's promotion.
What changes at the gate:
Dimension
Lab
Factory
Mission
Rapid PoC against real use-cases
Production-grade, governed, enterprise deployment
Cadence
2–4 week iteration cycles
Release trains with regression and validation gates
Governance
Lightweight — evaluation logged
Full — Part 11 audit trail, GAMP 5 / CSA, e-signatures
Stack
Kit + sandbox KG + playground
Stateful orchestration, IaC, CI/CD, pinned models, production KG
The Factory is not "deployment." It is the commitment to run the agent as a service for years, through model deprecations, regulatory changes, and organizational turnover. The support model has to be as explicit as the gate.
Tier
Who
Scope
Typical SLA
L1 — Domain super-users
Trained domain experts (2–3 per use-case, the champion network of §19)
First-line triage: "is this a bad recommendation or a bad question?"; feedback capture; coach-mode re-entry if acceptance drops
Any change to a factory card's service signature, model version, rule set, or data contract → impact assessment → regression eval → re-approval
Per validated-state procedure
Model lifecycle
Platform architecture
Frontier and scoped model deprecations tracked; every agent has a tested successor model before the vendor's end-of-life date
Quarterly review
Closed-loop evaluation is how production agents self-tune without leaving the validated state. Three loops run continuously: a Model-loop (automated evals against ground truth on every release), a Dogfooding-loop (internal power users flag failures before end-users see them), and a Production-loop (live decision samples diagnosed, failure classes reflected into prompt and rule adjustments, synthesized into the next version). A Reflect agent identifies failure classes; a Synthesize agent proposes patches — which go through change control like any other change. No single loop catches everything; a defect must pass all three undetected to reach a patient-facing decision.
7.5 Risk — the ED-level considerations
These are the risks a platform leader is accountable for, as distinct from the risks a project manager tracks.
Risk
What it looks like
Mitigation built into the model
Pilot purgatory
Dozens of PoCs, none in production, sponsors losing faith
Kill criteria; gate with named owners and an escalation SLA; "PoCs promoted per quarter" as a platform KPI
Validation debt
Agents in production with a retrofitted validation story; Quality discovers it during an inspection
Governance manifest at Lab entry; validation posture agreed before the gate; G3/G4 on from day one
Shadow AI
Teams bypass the platform with their own vendor tools because the platform is slow
Lab time-to-first-grounded-answer measured in days; kits that are genuinely faster than going around
Three-estate fragmentation
Partner agents ungoverned; three identity models; no shared observability
Platform contract published early (§13 Phase 1); partner onboarding as a registry workflow
Model-vendor concentration
One frontier model in every agent; a pricing change or deprecation is an enterprise event
Model-agnostic kits; scoped models for 80% of calls; successor model tested per agent
The conformed dataset the CPV agent relies on stops being maintained when the project team disbands
Named data-product owners as a gate criterion; contract SLAs; breaks surface in L2
Regulatory change
FDA/EMA AI guidance finalizes with expectations the platform does not meet
G8 regulatory tracking; validation posture documented and defensible; early engagement with Quality
Talent
FDEs, knowledge engineers, and CSV engineers who understand hybrid AI are scarce
Central bench; training program (below); kits reduce the expertise needed per pod
Adoption failure
Technically excellent agent nobody uses
Coach-mode acceptance as a gate criterion controlled by domain experts (§19)
7.6 Training — what to build into people, by role
Audience
What they need to be able to do
Format
Domain experts as agent supervisors
Read a confidence score; review a provenance trail; know when to override; give feedback that improves the agent
2-day workshop + coach-mode participation. Not a six-month program.
Rule calibrators (SMEs)
Take a provisional rule to confirmed: severity, scope, exceptions; sign it
Half-day + first calibration session with the FDE
Forward-deployed engineers
Speak the domain; instantiate kits; run the eval harness; write agent cards; lead a pod
Apprenticeship on a live pod; platform certification
Knowledge engineers
Extend a domain ontology under the modeling standard; write data contracts; map a LIMS to the canonical dictionary
Modeling standard training; KG sandbox exercises
Quality / CSV engineers
Validate a hybrid system: Cat 5 for deterministic paths, CSA-aligned controls for LLM paths; assess change impact on an agent card
Joint workshop with platform architecture; first validation package done together
Platform / Factory operations
Run L2: degraded mode, drift, cost anomalies, connector failures
Runbooks; game days
Executives and sponsors
Read the three metrics that matter per phase (§18); chair a gate; resist vanity metrics
90-minute briefing; quarterly steering
Training is not a launch activity. It is a standing capability, because every pod that forms needs it and every gate depends on it.
The operating model is domain-agnostic by design. What differs by domain is the decisions, the ontologies, the data products, and the regulatory frame — which is what the next section lays out, tower by tower.
Section 8 · Cross-Functional Domain Use-Cases
Eight towers, thirty use-cases — and how few new things appear
The platform is only as credible as the decisions it accelerates. This section walks the enterprise tower by tower — Research, Development, Product Development, Manufacturing, Supply Chain, Commercial, Medical & Safety, and the Enterprise Horizontal that underlies them all — and for each tower asks the same questions: what decisions are slow today, what the domain knowledge graph has to know, which use-cases the platform serves, and how each use-case is a configuration of the shared substrate rather than a project.
The towers are not equal. Research is pre-GxP and lives in the Lab; Manufacturing carries the heaviest governance in the enterprise and lives in the Factory. The platform serves both because the substrate is the same and the governance envelope is configurable. The reader should notice, across all eight towers, how few new things appear after the first two or three use-cases: the same kits (§6), the same dual-path reasoning (§10), the same gate (§7). That repetition is the platform working.
A convention for the rest of this section: each use-case names its decision, its architectural pattern, its agent topology, its GxP frame, its key metric, and the proven pattern it inherits from — so that a reader can see the shared substrate underneath the domain surface. Each card's footer also shows its G1–G8 governance intensity.
Research · pre-GxPDevelopment · GCPProduct Dev · pre-GMP→GMPManufacturing · GMP/CSASupply · GMP/GDPCommercial · OPDPMedical · GVP/PSMFHorizontal · infra
8.1
Research Tower
Pre-GxP for target ID and generative chemistry · GLP / Part 11 begin at preclinical safety · Lab owns most of this tower; Factory owns the safety boundary
Decisions that are slow today
Which target to commit $50–100M of preclinical spend to. Which of ten thousand generated molecules to synthesize. Whether a compound's liability profile justifies advancing to IND-enabling studies. The industry's stated ambition is to halve the time from target to first clinical study while raising the probability of success — and that is precisely a decision-velocity-plus-quality problem.
Governance posture
Pre-GxP for target identification and generative chemistry (research tooling); GLP and Part 11 begin at preclinical safety, where the nonclinical overview enters the IND. The Lab owns most of this tower; the Factory owns the safety boundary.
What the platform does here — and does not
The scientists own the models: genetics, omics, structure prediction, generative chemistry, often via external partnerships. What those models do not come with is a governed evidence graph, a constraint layer, an auditable rationale, and reuse across therapeutic areas. The platform's contribution is the substrate, and the honest framing is that the platform makes the scientists' models auditable, reusable, and safe to act on — it does not replace them.
A single interactive panel where the reader selects a target and watches the evidence neighborhood light up in the Research KG — genetic evidence in one color, functional in another, clinical in a third — with the deterministic score tiers rendered as a stacked bar beside the AI-synthesized brief, and the toxicologist's sign-off as the terminal node. The dossier is a subgraph with a signature on it.
UC-R1
Target Identification & Validation
Decision. Go / no-go on a target, defended by a biology lead against the question "why this one?"
Architectural pattern
Multi-omic datasets + genetics + imaging features + literature + internal assays
↓Research KG — typed edges, edge-level provenance, confidence, timestamp
↓Deterministic path: evidence-scoring rules (genetic support tiers, replication
count, effect-direction concordance) — versioned, auditable
AI path: target-discovery agent traverses KG + literature for novel
associations with evidence trails; validation agent cross-references
IP, competitive landscape, druggability
↓
Merged target dossier → biology lead reviews → go/no-go + rationale logged
Agent topology
Extraction: literature parser (PubMed, patents), multi-omic normalizer, assay structurer. Reasoning: KG traversal for target–disease links, druggability scorer, competitive IP analyzer. Orchestrator: evidence-dossier assembler with confidence scoring and provenance chain. Kit: evidence-dossier.
GxP frameNon-GxP (discovery)
Key metricTarget-to-candidate time; evidence-tier coverage per dossier
Proven patternEnterprise Knowledge Vault — KG with 10,292 triplets across 29 domains; RAG accuracy 75% → 99%+ over 14B+ data points
evidence-dossier
UC-R2
Generative Chemistry & Compound Design
Decision. Which generated candidates a medicinal chemist should spend synthesis budget on.
Architectural pattern
Generative engines (internal, partner, vendor) — pluggable via a common
candidate-proposal contract (SMILES + conditioning + model version + seed)
↓Deterministic constraint gate: structural alerts, known toxicophores, IP
blocklist, synthesizability threshold, property windows — every rejection
logged with rule ID; each rule versioned + owner-signed
↓
Multi-objective scoring specialists (potency, ADMET, selectivity) — each
isolated with its own eval record
↓
Ranked candidates + rationale + full lineage → med-chem review →
selections captured as preference signal
Agent topology
Generative agent (external or internal, behind the contract). Evaluation specialists: ADMET predictor, synthesizability scorer, patent-landscape checker. Rule engine (deterministic path). Kit: rule-gate in front of a review queue. Model-agnostic by design — which matters when the enterprise has multiple discovery partnerships and more than one cloud.
GxP frameNon-GxP (pre-IND research tooling)
Key metricHit rate per synthesis cycle; cycle time
Proven patternDoscierge rule engine — 970 deterministic rules, GAMP 5 Category 5, every rejection with a rule ID
rule-gate
UC-R3
Preclinical Safety Assessment
Decision. Advance or stop, before IND-enabling studies — and survive the FDA question "why did you conclude this compound was safe to advance?"
Architectural pattern
Historical study corpus (SEND datasets, study reports, assay results,
clinical outcomes of advanced compounds)
↓
Extraction agents → structured findings with study/species/dose/endpoint provenance
↓Safety KG (compound ↔ structural class ↔ mechanism ↔ target ↔ finding ↔
species ↔ clinical translation outcome)
↓Deterministic path: expert-calibrated signal rules (e.g. hERG IC50 < X AND
QT class effect in KG → flag cardiotox, severity S) — owner, calibration
record, version
AI path: KG-grounded LLM writes a cited "liability brief" — similar
compounds, what happened, translational hit rate
↓
Merged assessment → toxicologist reviews, signs → decision + rationale
become the audit record
Agent topology
Extraction: SEND parser, study-report extractor, finding normalizer (MedDRA/INHAND). Reasoning: KG traversal for similar-compound outcomes, translational hit-rate calculator, liability-brief drafter. Rule engine: 500+ expert-calibrated safety rules with provisional → confirmed loop. Human gate: toxicologist sign-off under Part 11. Kits: rule-gate + draft-then-check.
GxP frameGLP / Part 11 — the nonclinical overview is IND content
Key metricP/R/F1 per liability class against compounds with known clinical outcomes; false-negative rate is the number that matters
Transition. Preclinical safety is where the platform first crosses the regulatory boundary. Everything downstream lives on the other side of it.
8.2
Development Tower
GCP throughout · ICH E6(R3) and E8(R1) · Part 11 wherever an action touches a controlled document, a monitoring plan, or a supply decision · highest bar at regulatory documentation
Decisions that are slow today
Which protocol design to lock. Which sites to open. What to do about site 042, this week. Whether the CSR is ready for the reviewer. Trials designed with AI assistance finish ahead of schedule more often; the fundamental unit of work is the clinical trial, and this is the tower where the platform's home-turf patterns — ProtoCheck, the enterprise clinical-data integration, the documentation checker — apply most literally.
Governance posture
GCP throughout. ICH E6(R3) and E8(R1). Part 11 wherever an action touches a controlled document, a monitoring plan, or a supply decision. Highest bar at regulatory documentation.
A trial-lifecycle strip — design → sites → operations → documentation — where each stage's decision is a node, the agents feeding it are stacked beneath, and the same kit icons recur across stages (rule-gate at design and documentation; signal-and-NBA at operations; entity-resolution at sites). Hovering a kit highlights every stage that consumes it — reuse, not four separate pictures.
UC-D1
Clinical Trial Design Intelligence
Decision. Lock the protocol — with endpoint precedent, enrollment risk, and regulatory issues surfaced before finalization. Amendments cost ~$500K+ each and delay trials by months; roughly half are avoidable.
Architectural pattern
Corpus: internal protocols + outcomes, CT.gov, publications, ICH E6/E8/E9,
FDA/EMA letters, RWE
↓Trial-design ontology (indication ↔ endpoint ↔ population ↔ design element ↔
regulatory precedent ↔ operational outcome) — aligned to CDISC PRM
↓Design-check agent (ProtoCheck pattern): draft protocol → rule findings
(compliance, precedent conflict, enrollment-risk criteria) with severity + citation
Precedent-synthesis agent: "in this indication, 12 trials used endpoint X;
4 amended for reason Y" — cited
Twin/simulation agent: enrollment + dropout + site-burden model from
historical SDTM/ops data; design alternatives compared
↓Design review workspace: findings, precedents, simulations side by side;
every accepted/rejected recommendation logged
Agent topology
Design-check (500+ rules, 6+ TAs, FDA/EMA/ICH/WHO; human-in-the-loop mandatory per E6(R3)); precedent-synthesis (cross-trial analysis, cited); computational twin (enrollment simulation from SDTM/ADaM, screen-fail prediction, alternative comparison). Kits: rule-gate + evidence-dossier. Where a sponsor already has a protocol-intelligence tool, the design-check agent is the complement: the historical tool provides the intelligence; the rule engine provides the compliance gate and makes the recommendations auditable.
GxP frameGCP / E6(R3) / E8(R1) — the protocol is a controlled document
Key metricAmendments avoided; protocol cycle time
Proven patternProtoCheck — patent-pending, 500+ rules, 6+ therapeutic areas, four specialist reviewers, presented at BioNJ 2026
rule-gateevidence-dossier
UC-D2
Site Selection & Recruitment Intelligence
Decision. Which sites to open, ranked, with diversity-plan coverage as a first-class criterion — and which of the chosen sites will under-enroll, early enough to intervene.
Architectural pattern
Site/investigator master — entity-resolved across CTMS, EDC, CRO feeds, CT.gov,
NPI/ORCID; golden record per site + per investigator (the MDM pattern)
↓Site KG (site ↔ investigator ↔ indication experience ↔ enrollment curves ↔
data quality ↔ competing trials ↔ catchment patient density filtered by I/E)
↓Deterministic scoring: regulatory readiness, equipment, performance
thresholds, diversity coverage (FDORA) as a first-class criterion
Predictive agent: enrollment forecast per site for THIS protocol's I/E;
cited rationale per recommendation; bias flag on over-concentration in usual sites
↓
Ranked site list + rationale + diversity coverage → feasibility team reviews →
chosen sites feed operational monitoring (UC-D3)
Agent topology
Entity-resolution engine (golden record, crosswalk, 100K+ entity scale proven). Scoring specialists: feasibility rules, enrollment prediction, diversity calculator. Recommendation agent with bias detection and what-if analysis. Kits: entity-resolution + evidence-dossier. The under-appreciated insight: investigators are HCPs and sites are HCOs — the same golden-record discipline commercial has applied for years is what makes site intelligence durable across studies.
GxP frameGCP / FDORA diversity action plans
Key metricEnrollment velocity (time to last patient enrolled); screen-fail reduction; site-stall early-warning lead time
Proven patternFortune 50 pharma Customer MDM — 100K+ HCP/HCO, 20+ system integration; clinical data integration across 100+ studies (SDTM/ADaM)
entity-resolutionevidence-dossier
UC-D3
Operational Precision in Clinical Trials
Decision. What the study team should do this week about supply, monitoring, and recruitment — as next-best-actions with rationale, expected impact, and an owner. The failure mode is alert fatigue and unactionable recommendations. This is also the primary feed for the Clinical Operations Command Center (§12).
Architectural pattern
Event sources: IRT, EDC, CTMS, RBM/central monitoring, enrollment feeds
↓Trial-ops semantic layer (study ↔ site ↔ subject-status ↔ supply-state ↔
monitoring-finding ↔ plan) — shared across ALL studies, not per study
↓Signal specialists (isolated, independently evaluated):
Supply: stock-out / expiry / resupply-lead-time risk
Monitoring: KRI drift, deviation clusters, data-quality anomalies
Recruitment: enrollment-vs-plan, screen-fail trends, site stall
↓Orchestrator: de-duplicate, prioritize, compose NBA with rationale + expected
impact + who-should-act; suppress noise from study-team feedback
↓Study team inbox: accept / modify / reject → logged (Part 11 where the action
touches GCP) → feedback retrains prioritization
Agent topology
Three monitoring specialists + orchestrator on stateful orchestration with per-execution history. Kit: signal-and-NBA. Any recommendation can be reconstructed months later during an inspection — that is what stateful orchestration with execution history buys.
GxP frameGCP / Part 11 (selective — where actions touch monitoring plans or supply decisions)
Key metricSignal-to-action time (monthly → real-time); NBA acceptance rate; alert-fatigue index
Proven patternFortune 50 pharma real-time platform on 14B+ data points; 4-agent platform for 150+ regulated users; Doscierge 11-state Step Functions pipeline with 25 functions and full execution history
signal-and-NBA
UC-D4
Clinical & Regulatory Documentation
Decision. Is this CSR section / protocol section / IB / CTD module summary ready for the reviewer — with every sentence traceable to an approved source? Highest volume, most immediately measurable, and the use-case regulators are watching most closely.
The critical gap across the industry. Every major company now has GenAI tools that generate drafts at 80–85% acceleration. Almost nobody has a system that deterministically checks those drafts against the regulatory rulebook before a reviewer sees them. Generate-then-check is the architecture that makes this safe for production — and the checker is the harder, higher-value half.
Drafting agent. Checker (970 deterministic rules; structure, required-element, numeric-consistency; KG-grounded review) — exposed as a well-typed MCP tool, not an agent. Eval harness: hallucination rate per document type, numeric-consistency rate, reviewer edit distance, cycle time vs. baseline. Kit: draft-then-check.
Key architectural insight
The checker is a tool the agent calls, not an agent itself. Any agent workflow — any model vendor's, any internal orchestrator's — can call the compliance API as one step. For Category 5 validation, a model-driven tool-selection loop is a liability, not a feature; deterministic control flow is the deliberate ceiling.
GxP framePart 11 / Annex 11 (highest) — e-signature on sign-off; immutable audit trail on every edit. GAMP 5 Category 5 for the checker; CSA-aligned controls for the drafting LLM
Key metricDocument cycle time; reviewer edit distance; hallucination rate (target: zero unsupported claims reaching a reviewer)
Proven patternDoscierge — 970 rules, 10,292 KG triplets, checker + eval harness; Fortune 50 pharma GenAI function scaled to 150+ regulated users
draft-then-check
Transition. Development hands a validated process and a dossier to the next tower. Product Development is where the molecule becomes a manufacturable product — and where the knowledge that manufacturing will need is first written down.
8.3
Product Development Tower (CMC & Tech Transfer)
Largely pre-GMP but produces CTD Module 3 content and defines GMP-critical parameters · tech transfer is the GMP boundary · process validation Stage 1 lives here
Decisions that are slow today
Which formulation and process to lock. What the control strategy is — which parameters are critical, what the design space is. Whether the receiving site is ready to accept a transferred process. What to file as an established condition under ICH Q12 so post-approval changes do not require a supplement. This is the tower that writes the knowledge Manufacturing later uses, and the platform's leverage is making that knowledge structured from the start rather than reconstructed from PDFs a decade later.
Governance posture
CMC development is largely pre-GMP but produces CTD Module 3 content and defines GMP-critical parameters; tech transfer is the GMP boundary; process validation Stage 1 (Process Design) lives here.
A "knowledge handoff" diagram: the CMC KG on the left, the Manufacturing KG on the right, and the CPP/CQA/specification/established-condition nodes drawn once, spanning both — because they are the same nodes. Product Development is where the manufacturing knowledge graph is born.
UC-P1
Process Knowledge & Control-Strategy Intelligence
Decision. Which parameters are critical, with what ranges, and why — defensible in the dossier and usable by CPV on day one of commercial manufacture.
Architectural pattern
DoE studies, characterization reports, scale-up data, analytical method files
↓
Extraction agents → structured CPP/CQA candidates with study provenance
↓CMC KG: parameter ↔ attribute relationships with evidence strength; design space
↓Deterministic path: criticality-assessment rules (ICH Q9 risk ranking,
Q8 design-space criteria) — versioned, SME-signed
AI path: control-strategy synthesis agent drafts the justification with citations
↓Process-development lead reviews → control strategy locked → CPP/CQA list
becomes the CPV monitoring scope (UC-M2) without re-keying
Kitsrule-gate + draft-then-check
GxP framePre-GMP; CTD Module 3 content
Key metricTime from characterization complete to control strategy locked; proportion of CPPs with structured evidence links
Decision. Is the receiving site — internal or CMO — ready to run this process, and what are the gaps?
Architectural pattern
Sending-site process knowledge package (from CMC KG) + receiving-site capability
(equipment, methods, materials, training records — from Manufacturing KG)
↓Gap-analysis agent: traverses process ↔ equipment class ↔ site equipment;
method ↔ site method validation status; material ↔ approved supplier at site
↓Deterministic completeness rules (transfer protocol required elements per
WHO TRS 961 Annex 7 / ICH Q10) + AI-drafted transfer protocol sections with citations
↓Transfer lead + receiving-site QA review → gaps become tracked items; protocol
signed under Part 11 where GMP applies
Kitsevidence-dossier + draft-then-check
GxP frameGMP at the receiving site; Part 11 on the transfer protocol
Key metricTransfer cycle time; gaps discovered before vs. after first engineering batch
Proven patternEntity resolution across process ↔ equipment ↔ site; the same crosswalk discipline as MDM
evidence-dossierdraft-then-check
UC-P3
Post-Approval Change & Established Conditions (ICH Q12)
Decision. Does this proposed change touch an established condition, and what is the reporting category — before the change-control board meets?
Architectural pattern
Proposed change (from QMS) → change-classification agent traverses CMC KG:
affected parameter → is it an established condition? → which reporting category?
→ which markets? → prior precedent for similar changes?
↓Deterministic path: reporting-category rules per market (versioned, RA-signed)
AI path: cited impact assessment draft
↓Regulatory-affairs reviewer disposition → decision logged; the KG edge is updated
Key metricChange-classification cycle time; classification disagreements with RA (target: falling)
rule-gateevidence-dossier
Transition. The control strategy, the specifications, and the genealogy of materials now exist as structured knowledge. Manufacturing is where they are tested against reality, batch after batch.
8.4
Manufacturing Tower
GMP / Part 11 / CSA — the highest bar · deterministic paths GAMP 5 Cat 5 · statistical and chemometric models under an analytical-method lifecycle with QA-approved model registry · LLM paths controlled by grounding, checking, and human sign-off
Decisions that are slow today
Is the process in a state of control — this month, not fifteen months from now? Is this batch's blend uniform — now, not after QC returns a thief-sample result? What caused this deviation, and did the last CAPA actually work? Can this batch be released? Manufacturing is the most heavily governed tower in the enterprise — GMP, Annex 11, ALCOA+, FDA's Computer Software Assurance guidance, draft EU GMP Annex 22 (consultation July–October 2025; final targeted Q4 2026) — and it is also the tower where decision latency is most directly a compliance exposure: a CPV program that generates a signal nobody acted on, followed by a batch failure, is a warning letter.
Governance posture
GMP / Part 11 / CSA — the highest bar. Deterministic paths validated as GAMP 5 Category 5; statistical and chemometric models managed under an analytical-method lifecycle with QA-approved model registry; LLM paths controlled by grounding, checking, and human sign-off.
Why this tower is its own tower, with four sub-domains
Manufacturing is not one use-case called "manufacturing." It is at least four distinct decision families that share one data foundation and one KG but differ in cadence, regulation, and user: the Annual Product Quality Review (retrospective, regulatory, QA-owned), Continued Process Verification (continuous, statistical, MS&T-owned), the PAT platform (real-time, spectral, model-lifecycle-owned), and Blend Endpoint Detection (the PAT platform's entry-point use-case, operator-facing). Deviation/CAPA investigation and batch-record review sit alongside them. Treating these as separate towers would lose the shared substrate; treating them as one use-case would lose the fact that they have different owners and different validation paths.
The data foundation the whole tower stands on
Three data products, joined: batch_phase_summary (CPPs from the historian, batch-aligned via MES event frames), the LIMS conformed services (samples → tests → test_results, given meaning by specifications; CQAs), and genealogy + raw_material_results (lot lineage; CoA/CoC from suppliers and CMOs). With those three, the CPP-to-CQA relationship is a query. Without them, every use-case below is a spreadsheet campaign. The KG holds the meaning persistently; the batch-level facts are materialized at run-time from the data products (§5.4).
Manufacturing & Quality KG — what it must know
product ──manufactured_as──► batch ──at_site──► site ──has_equipment──► equipment ──has_tag──► historian_tag (PI/OSIsoft, DeltaV)
│ │ └── has_pat_instrument ──► instrument (NIR, Raman, FTIR, FBRM)
│ ├── has_phase ──► unit_operation_phase (ISA-88; start/end from MES event frames)
│ │ └── has_parameter_summary ──► parameter (CPP; mean/min/max/duration-out-of-range)
│ │ └── validated_range ──from──► control_strategy (→ CMC KG)
│ ├── has_sample ──► sample (sample_point) ──has_test──► test ──has_result──► test_result (status: final/retest/superseded)
│ │ └── against ──► specification (version, effective date)
│ ├── genealogy ──► batch | material_lot (DS lot → DP lots; raw-material lot → batch)
│ │ └── has_coa ──► raw_material_result (CoA / CoC from supplier or CMO)
│ ├── has_deviation ──► deviation ──root_cause──► cause ──addressed_by──► CAPA ──effectiveness──► outcome
│ ├── has_change ──► change (→ ICH Q12 established condition in CMC KG)
│ ├── has_spectral_record ──► spectrum (instrument, timestamp) ──scored_by──► pat_model (version, QA-approved)
│ └── mined_edge: stalls_at (phase, p50 duration, confidence) ← process intelligence
├── has_apqr ──► apqr (period, site, status) ──covers──► batch[] ──finding──► trend | OOS | OOT
└── has_cpv_program ──► cpv_scope ──monitors──► parameter[] ──chart──► spc_signal (rule, batch, disposition)
Aligned to: ISA-88/95, ICH Q7/Q9/Q10, 21 CFR 211, EU GMP Ch. 1.10 / Annex 15; core: Product, Batch, Material, Specification, Site
One diagram, three tiers. Bottom: the three data products (historian-derived CPPs, LIMS CQAs, genealogy/CoA) as three source lanes converging into the Manufacturing KG. Middle: the five use-cases as consumers of the same graph, with their cadences labeled (APQR: monthly roll-up; CPV: nightly; PAT: real-time; blend endpoint: per revolution; deviation: on event). Top: the Manufacturing Quality Command Center (§12) receiving all five feeds. One foundation, five cadences, one governed surface.
UC-M1
Annual Product Quality Review (APQR)
Decision. Is the process consistent, are the specifications still appropriate, and what needs to change — per product, per site, per year, with cross-site comparison. Mandated by 21 CFR 211.180(e) and EU GMP Chapter 1.10; a deficient APQR is one of the most common observations against otherwise well-run sites.
4–8 wksToday. 4–8 weeks of calendar time per product; 2.5–5 FTE of skilled labor for a five-product, three-site company; findings delivered 6–15 months after the batches were made; 12,000 batch-parameter values a year hand-transcribed at a 1–5 per 1,000 error rate; cross-site comparison effectively impossible.
Architectural pattern
batch_phase_summary + LIMS conformed services + genealogy + QMS extracts
(deviations, changes, complaints) — validated data products
↓
APQR orchestrator (scheduled — monthly, not annually): per product-site,
assembles the review period's batch set; checks event_completeness_score
↓Deterministic path: validated SPC/capability queries (Cpk/Ppk vs. validated
range and spec), OOS/OOT listings, batch-count reconciliation vs. MES —
GAMP 5 Cat 5 query layer
AI path: trend-narrative agent drafts the interpretation section with every
chart and number cited to its data-product row; deviation/CAPA/change
cross-reference with provenance
↓Checker (draft-then-check): required elements per 211.180(e) / Ch. 1.10;
numeric consistency between narrative and charts
↓QA reviewer sees data package + draft + findings → signs under Part 11 →
the APQR becomes a periodic summary of an always-on system
Agent topology
Data-completeness monitor; SPC/capability calculator (deterministic); narrative drafter; checker (MCP tool); orchestrator. Kits: draft-then-check on a signal-and-NBA backbone. Parallel-run against the manual APQR for a full cycle before the manual process is retired — agreement within pre-defined tolerance on every CPP, every discrepancy explained, QA signs the PQ report.
GxP frameGMP / Part 11 / ALCOA+; the query layer is validated software
Key metricSignal latency (6–15 months → weeks); APQR data-gathering cycle time (weeks → days); reviewer time spent verifying vs. interpreting (60–70% verification today → the inverse)
Proven patternDoscierge checker pattern; ALCOA+ evidence chain (reconciliation records, immutable storage, access logs) so any chart is reproducible from raw historian data on demand during inspection
draft-then-checksignal-and-NBA
UC-M2
Continued Process Verification (CPV)
Decision. Has the validated process drifted — and which scientist needs to look, today. FDA Process Validation Stage 3; EU Annex 15 Ongoing Process Verification; ICH Q10 §3.2.1. A program, not an event, for the entire commercial life of the product.
4–12 wksToday. Quarterly or monthly Excel/Minitab exercise; 4–12 week lag from batch completion to chart; 5–10 of 30–50 relevant CPPs actually trended; each parameter monitored independently, so the pH-temperature interaction that shifts the glycosylation profile is invisible; every site a silo.
Architectural pattern
Batch-end event (MES) → nightly: phase-level statistics for EVERY CPP in the
control strategy (marginal cost of one more parameter ≈ zero)
↓Monitoring specialists (isolated, each with its own eval record):
SPC agent: Shewhart / CUSUM / EWMA; Western Electric / Nelson rule violations
Capability agent: Cpk (within-batch) and Ppk (overall) vs. validated range;
Ppk of each CQA vs. specification (from LIMS conformed services)
MSPC agent (Phase 3): per-batch PCA on phase-level CPP vector; Hotelling T²
and Q-residual vs. normal-operating-condition model; contribution plot on excursion
Lot-effect agent: genealogy + raw_material_results — does a resin or media
lot explain the shift?
↓Orchestrator: de-duplicate signals across agents; prioritize by risk (ICH Q9);
compose NBA with rationale, the chart, the affected batches, and the
responsible scientist; suppress known-benign patterns from feedback
↓Process scientist disposition (confirmed / benign / investigate) → logged;
investigation opens in QMS; disposition feeds the APQR (UC-M1) and the
Manufacturing Quality Command Center (§12)
Agent topology
Four monitoring specialists + orchestrator. Kit: signal-and-NBA. Cross-site CPV is GROUP BY site because sites share the canonical data model. The MSPC model is a statistical digital shadow of the process — a model of normal behavior against which reality is continuously compared.
GxP frameGMP; statistical engine validated (Cat 5); MSPC model under analytical-method lifecycle (NOC set, cross-validation, held-out detection test, QA approval in the model registry)
Key metricBatch-to-chart latency (4–12 weeks → overnight); CPP coverage (5–10 → all); signal-to-disposition time; deviations pre-empted by a CPV signal
Proven patternFortune 50 pharma real-time platform (14B+ data points); Doscierge stateful pipeline; the same signal-and-NBA kit as UC-D3
signal-and-NBA
UC-M3
Enterprise PAT Data Platform
Decision. Which batches are abnormal while they are running — and whether the chemometric model that says so is the approved version. NIR, Raman, FTIR, FBRM instruments across sites generate high-value data about what the material is at every moment; almost none of it leaves the vendor PC it was captured on, none is joined to the historian, and model provenance leaves with the scientist who built it.
Architectural pattern
Instrument PCs → scalar outputs (predictions, statistics, instrument health)
flow through the historian-to-cloud channel like any tag; spectral vectors
flow through a parallel high-throughput path into the same lake
↓Model registry: chemometric models (PLS, PCA, MSPC) as governed, versioned,
QA-approved assets — not files on a laptop; every prediction carries model
version + calibration set reference
↓
Monitoring specialists: golden-batch MSPC (T²/Q per batch, in real time);
instrument-health drift; model-performance drift (prediction vs. reference
method)
↓
Process-model agents (mechanistic, data-driven, hybrid) deployed to the floor
with a confidence band; advisory first, control later, RTRT only if chosen
↓MS&T / QA console: model lifecycle, cross-site model transfer, method
harmonisation; every batch's spectra retained and retrievable for a regulator
Agent topology
Ingestion monitors; MSPC scorer; model-drift monitor; model-lifecycle orchestrator (registry approvals as a workflow — the same pattern as §6). Kit: signal-and-NBA with a model registry extension. This platform is the prerequisite for continuous manufacturing (ICH Q13) and for any credible manufacturing digital twin.
GxP frameGMP; model lifecycle per analytical-method validation; GAMP 5 for the platform; Annex 11 for records
Key metricTime to first value (blend endpoint advisory, 6–9 months); time to multivariate monitoring on all site instruments (12–15 months); model provenance completeness (target: 100% of predictions traceable to an approved model version)
Proven patternRegistry-driven approval workflow (§6); trust-tiered provenance on every scored record
signal-and-NBAmodel registry
UC-M4
Blend Uniformity Endpoint Detection
Decision. Stop blending when the material is ready, not when the clock says so. An NIR probe on the blender reads the blend once per revolution; a validated chemometric model converts each spectrum to predicted API content; when scan-to-scan variation falls below threshold and stays there, the blend is uniform.
Why it is the PAT entry point. Lowest regulatory risk (the fixed-time step stays validated; the endpoint runs as an advisory signal first; the FDA PAT Guidance uses blend uniformity as its canonical example); fastest ROI (cycle-time savings begin the day the advisory goes live); most proven science (20+ years of precedent; PLS + moving-block standard deviation is the industry approach); highest platform reuse.
Architectural pattern
NIR spectrum per revolution → PAT platform (UC-M3) → PLS model (registry-approved
version) → predicted API content → moving-block SD
↓Endpoint agent (deterministic): SD < threshold for N consecutive blocks →
"blend uniform" advisory with confidence; every advisory logged with model
version, spectra references, batch, timestamp
↓Operator sees advisory beside the validated fixed-time step (Phase 1–2);
operator action logged. Phase 3: endpoint replaces fixed time under a
validated change. Phase 4 (optional): RTRT — content-uniformity QC test eliminated.
↓
Batch's endpoint evidence flows to CPV (UC-M2) and APQR (UC-M1) as a CPP
Agent topology
One deterministic endpoint specialist on the PAT platform; the human is the operator. Kit: signal-and-NBA (degenerate case: one signal, one action). The simplest agent in this article, and one of the most valuable.
GxP frameGMP; advisory mode requires no submission; RTRT requires one
Key metricBlend time saved per batch (15–45 min); manufacturing hours freed (50–150 h/yr per high-volume product); value of freed capacity ($250K–$2.25M/yr at $5K–$15K per batch-hour); payback 6–18 months
Proven patternConfidence-gated advisory → autonomous graduation (§18's coach-mode model, applied to a process step)
signal-and-NBA
UC-M5
Deviation & CAPA Investigation
Decision. What caused this deviation, has it happened before, did the last CAPA work — and can the batch be released?
Architectural pattern
Deviation opened in QMS → investigation agent retrieves similar historical
deviations (KG traversal: product ↔ process ↔ equipment ↔ deviation ↔ root cause)
with provenance; pulls the batch's CPV signals (UC-M2), PAT records (UC-M3),
genealogy and CoA (UC-S1), and mined process-intelligence edges ("this phase
stalls 3× more often on line 2")
↓Deterministic path: batch-record completeness, spec-conformance, trending
(ICH Q3A thresholds, stability limits); CAPA-effectiveness rules
AI path: root-cause hypothesis draft with cited evidence; CAPA precedent summary
↓Checker: investigation-report required elements; numeric consistency
↓Investigator accepts / rejects hypotheses → QA signs under Part 11 →
root cause and effectiveness become KG edges for the next investigation
Transition. Every batch depends on materials that came from somewhere and will go somewhere. Supply Chain is where genealogy meets the world outside the site.
8.5
Supply Chain Tower
GMP for material release and supplier qualification · GDP for cold chain · DSCSA / serialization for traceability · Part 11 on release decisions
Decisions that are slow today
Is this incoming material lot acceptable — with the CMO's Certificate of Analysis reconciled against our specification, not just filed? Which shipment lane is at risk of a temperature excursion? Which supplier's qualification is about to lapse? Where will demand exceed supply next quarter, and which sites should absorb it? Supply chain decisions are latency-critical because a late decision becomes a shortage, and a shortage in pharma is a patient outcome.
Governance posture
GMP for material release and supplier qualification; GDP (Good Distribution Practice) for cold chain; DSCSA / serialization for traceability; Part 11 on release decisions.
Supply Chain KG — what it must know
material ──supplied_by──► supplier ──qualified_for──► material (status, audit date, expiry) ──at_site──► site
│ └── is_cmo ──► cmo ──produces──► batch (external) ──has_coa/coc──► raw_material_result (vs. specification)
├── has_lot ──► material_lot ──genealogy──► batch (→ Manufacturing KG)
│ └── shipped_via ──► shipment ──on_lane──► lane ──has_excursion──► temperature_event (GDP)
├── has_demand_signal ──► forecast (market, period) ──vs──► supply_plan ──at_risk──► shortage_risk
├── serialized_as ──► serialization_event (DSCSA: commission, ship, receive, verify)
└── geopolitical_exposure ──► region_risk (single-source, sanctions, logistics)
Aligned to: GS1, DSCSA, GDP, ICH Q7 (supplier), ICH Q9; core: Material, Batch, Organization, Site
A material's journey as a single traced line — supplier → CoA → receipt → genealogy → batch → serialization → distribution lane — with the three agents that watch it drawn as sensors on the line, and the Manufacturing KG and Supply Chain KG shown as one continuous graph across the site boundary.
UC-S1
Incoming Material & CMO Quality Intake (CoA/CoC Reconciliation)
Decision. Accept this lot. Today, CMO results arrive as spreadsheets and PDFs; process engineers report spending the majority of their CPV time on extraction and aggregation, and CMO file handling is a disproportionate share of it.
Architectural pattern
CoA / CoC (PDF, spreadsheet, EDI) → extraction agent → structured raw_material_result
with document provenance → standardized file-ingestion service with
review-and-approve workflow → lands in the same conformed table as internal results
↓Deterministic path: reconciliation vs. specification (by effective date), test
completeness, method equivalence flags; supplier-qualification status check
AI path: exception summary with citations when a value is missing, out of trend,
or method differs
↓QC/QA disposition → release decision under Part 11 → lot appears in genealogy
and in CPV lot-effect analysis (UC-M2) automatically
Kitsdraft-then-check (the checker is the reconciliation rule set)
Proven patternDoscierge extraction pipeline with KG-first grounding; ProtoCheck rule calibration
draft-then-check
UC-S2
Supply Risk & Cold-Chain Monitoring
Decision. Which shipments, suppliers, and lanes need action this week.
Architectural pattern
Signal specialists (temperature-excursion monitor on lane telemetry; supplier-qualification-expiry monitor; single-source and geopolitical exposure monitor; demand-vs-supply gap monitor) → orchestrator composing NBAs with rationale, impact, and owner → supply-risk board disposition → Supply Chain feed into the Manufacturing Quality Command Center (§12).
Kitsignal-and-NBA — the same kit as UC-D3 and UC-M2
Decision. Where is this saleable unit, is its provenance intact, and is a suspect-product investigation warranted?
Architectural pattern
Serialization events (DSCSA) as KG edges on the material lot; anomaly monitor for verification failures and chain gaps; investigation agent assembles the chain-of-custody dossier with provenance.
Kitevidence-dossier
GxP frameDSCSA / GDP
Key metricVerification-failure investigation time
evidence-dossier
Transition. The product leaves the site. From here, the decisions are about people — the HCPs who prescribe, the patients who take, and the medical and safety professionals who answer for both.
8.6
Commercial Tower
Promotional compliance (FDA OPDP, EFPIA, local codes) · MLR-approved content only · medical-vs-commercial firewall enforced by architecture, not policy · AE detection with mandatory reporting on every conversation
Decisions that are slow today
What should this rep say to this HCP on Thursday, using only MLR-approved content? Is this piece of content approved for this market, this indication, this audience? Which signals in the market — competitor launches, formulary changes, prescribing shifts — need a response? Commercial agentic AI is increasingly delivered on a CRM vendor's agent platform under a multi-year rollout; the platform's job is to make that estate governable and to put the semantic layer underneath it.
Governance posture
Promotional compliance (FDA OPDP, EFPIA and local codes); MLR-approved content only; medical-vs-commercial firewall enforced by architecture, not policy; adverse-event detection with mandatory reporting on every conversation.
Two agent estates drawn on either side of a hard line — commercial on one side, medical on the other — sharing only the HCP golden record at the bottom and the AE sidecar routing across the top to PV. The MLR asset registry is the only content source on the commercial side.
UC-C1
HCP & Patient Engagement (Field Force and Self-Service)
Decision. Next-best-action for this HCP, from approved content, with provenance — and nothing else.
Architectural pattern
CRM vendor's agent platform as the experience layer
↓Customer & content semantic layer (HCP/HCO golden record ↔ product ↔
indication ↔ MLR-approved asset registry with claim-level provenance)
↓Field-force agent: pre-call plan + NBA from MLR content ONLY, each claim cited
to its approved asset and its approved source
Medical-inquiry agent: firewalled from commercial context by architecture
(separate identity, separate registry scope, separate KG partition)
AE-detection sidecar: mandatory on every patient/HCP conversation; routes to
PV (UC-X1) with the 24h / 15-day clock started
↓Compliance gate: OPDP / local code rules; MLR expiry; consent check
Agent topology
Field-force assistant; medical-inquiry agent (separate principal); AE sidecar (mandatory, non-bypassable hook). Kits: draft-then-check (the checker is the promotional-compliance rule set) + entity-resolution. The partner-built agents run inside the platform contract (§8.8, UC-H1).
Decision. Which market movements need a response — with the evidence.
Architectural pattern
Signal specialists over competitor, formulary, launch, and prescribing feeds → orchestrator → NBA to brand teams → the Commercial Intelligence Command Center (§12).
Kitsignal-and-NBA
GovernanceData-use and privacy constraints (de-identified, market-level) enforced in the KG partition, not by policy
Key metricSignal-to-response time; signal precision (vs. noise)
signal-and-NBA
Transition. The AE sidecar hands off to the last tower. Medical & Safety is where the enterprise answers for the product after it is in the world.
8.7
Medical & Safety Tower
GVP / Part 11 / PSMF for safety · medical-affairs approved sources only · scientific-exchange and off-label boundaries as versioned rules per market (FDA vs. EFPIA) · Sunshine Act / Open Payments and EFPIA disclosure on every transfer of value · firewalled from commercial, architecturally · GDPR / HIPAA on HCP data · AI use in PV documented in the PSMF
Decisions that are slow today
Is this a serious, unexpected adverse event — and is the 15-day clock met? What should the response to this medical inquiry be, from approved sources, in the inquirer's language? What should this MSL know before Thursday's meeting with this KOL — prior interactions, open inquiries, the KOL's last three abstracts, the approved SRLs that answer what they are likely to ask, and what cannot be shared? Which of the six thousand abstracts at this congress change the evidence narrative for our product or a competitor's, and which therapeutic-area lead needs to know tonight? What is the field hearing, in aggregate, that should change the medical strategy — and is it a pattern or three anecdotes? Where are the evidence gaps that justify an investigator-initiated study or a real-world-evidence analysis? Is this medical slide deck approvable, and which claim is the one that will be challenged? Medical Affairs is where the platform's evidence-provenance discipline is tested by the hardest clocks (pharmacovigilance), the most explicit boundary rules (scientific exchange vs. promotion, per market), and — since the EMA/FDA AI reflection papers — the most explicit documentation requirement: AI use in pharmacovigilance must be documented in the PV System Master File.
Why this tower is often the right first wave
Medical Affairs work is text-native (inquiries, interaction notes, abstracts, publications, SRLs), high-volume, and governed by approved-source discipline rather than GMP validation. The regulatory frame is real — GVP, Part 11 on signed records, Sunshine Act / Open Payments and the EFPIA Disclosure Code on every transfer of value to an HCP, FDA and EFPIA scientific-exchange boundaries, GDPR on HCP data — but it is a frame the draft-then-check and rule-gate kits already satisfy. Several top-20 companies have published Medical Affairs AI proof points: semantic clustering of unsolicited medical-information inquiries; GenAI theme analysis of field medical insights; GenAI content drafting. What none of them has published is the platform underneath — the shared medical semantic layer, the canonical interaction model, the firewall, the telemetry plane — that turns three standalone tools into a portfolio where the fourth product costs less than the first. That is the gap this tower is designed to close, and it is why the Lab → Factory model (§7) applied to Medical Affairs is described at the end of this section.
Governance posture
GVP / Part 11 / PSMF for safety. Medical-affairs approved sources only. Scientific-exchange, unsolicited-request, and off-label boundaries enforced by versioned rules per market (FDA vs. EFPIA), not by prompt. Sunshine Act / Open Payments (42 CFR 403) and EFPIA Disclosure Code on every HCP transfer of value — advisory boards, consulting arrangements, investigator-sponsored study grants, congress sponsorships — each traced from source event to the reported aggregate spend figure with threshold alerting; the reportable amount is computed deterministically, never by an agent. Firewalled from commercial (architecturally, per §8.6). GDPR / HIPAA on HCP and patient data; multi-lingual, regionally parameterized pipelines — rulesets as configuration tagged by jurisdiction, never regional code forks.
The Medical Affairs integration architecture. Medical Affairs sits on the seam of the enterprise's most fragmented system estate: a field CRM for MSL interactions and medical inquiries (Veeva CRM for most of the industry today, with an end-of-support date that forces a choice between Vault CRM and Salesforce Life Sciences Cloud before 2030); a content platform for medical materials (Vault MedComms, with PromoMats for anything that touches promotion); an enterprise HCP/HCO master; a data platform for interaction analytics; a marketing-cloud stack for omnichannel; and, increasingly, the CRM vendor's own agent layer. The architectural decision that makes every use-case below survivable is to abstract the interaction model: agents and analytics bind to a canonical Medical Affairs data model — interaction, inquiry, insight, content, HCP, event, transfer of value — materialized at run-time from whichever CRM the persona works in (§5.4's pattern). The CRM migration then becomes a connector swap, not an agent rewrite. System-of-record discipline follows: the HCP golden record stays in MDM; the approved-content record stays in Vault; the interaction and inquiry records stay in the CRM; engagement analytics live in the data platform; and agents read from data products, never from four systems, and write decisions back only to the workflow surface the human uses. The field CRM's offline-first constraint is inherited by anything the MSL touches — a brief that needs connectivity is a brief that does not get read.
The CRM transition is a data-governance event, not a migration project. The working pattern across multinationals navigating the Veeva CRM end-of-support window: bind agents and analytics to the canonical interaction model above — Interaction, Inquiry, Insight, Content, HCP, Event, Transfer of Value — materialized from whichever CRM the persona works in today. The canonical model is the data contract; the CRM connector is the adapter. When Veeva CRM gives way to Vault CRM or Salesforce Life Sciences Cloud (or both, in different markets, for three years), the migration is a connector swap, not an agent rewrite. Three implications for the platform: (1) every agent data contract references the canonical model, never CRM-native objects; (2) the dual-CRM state — some markets on Veeva, some on Salesforce LSC — is the default operating mode through at least 2029, so the canonical model must run both adapters simultaneously; (3) regional data residency must be re-verified per market during migration — a Salesforce LSC Hyperforce deployment is multi-tenant, and EU data residency requires explicit configuration (Hyperforce EU region + Data Residency add-on), not assumption.
The case as a timeline with the clock drawn as a bar: receipt → structured → coded → narrated → checked → signed → submitted, with each agent's contribution shown as a segment and the human sign-off as the only segment that can stretch. The agents compress the segments before the human, not the human's.
UC-X1
Pharmacovigilance & Safety Case Processing
Decision. Seriousness, expectedness, causality — and a submitted, compliant case within the clock.
Architectural pattern
Intake channels: literature, call center, social media, clinical sites, partner
reports, the commercial AE sidecar (UC-C1)
↓
Intake specialists (per channel): extract structured case data with provenance;
clock starts at receipt
↓Deterministic path: seriousness, expectedness (vs. label version by market),
causality-support rules — versioned, calibrated and signed by PV physicians
Coding specialist: MedDRA PT suggestion with confidence
AI path: narrative drafting agent — CIOMS format, span-level citations to source
↓Checker: E2B(R3) required elements; narrative-to-structured-field consistency
↓PV reviewer sign-off (Part 11) → submission → clock compliance recorded;
AI role documented in the PSMF
↓
Signal-detection specialist: cross-case pattern recognition; FAERS /
EudraVigilance integration → signal dossier → safety physician review
Agent topology
Intake specialists; coding specialist; rule engine; narrative drafter; checker (MCP tool); signal-detection specialist. Kits: draft-then-check + rule-gate — the same substrate as regulatory documentation (UC-D4) with a different ontology and a harder clock.
GxP frameGVP / Part 11 / PSMF
Key metric15-day clock compliance; coding accuracy vs. PV-physician review; case processing time; signal lead time
Proven patternProtoCheck expert-calibrated rules with physician signature; Doscierge draft-then-check; eval harness per document type
draft-then-checkrule-gate
UC-X2
Medical Information & Inquiry Response
Decision. The response to this inquiry — from medical-affairs-approved sources only, with provenance, triaged to an MSL where judgment is required.
Architectural pattern
Inquiry (CRM Medical Inquiry object, call center, portal, email — any language) → intent, product, and market classification → semantic clustering against the rolling cluster set (the pattern published as semantic clustering of unsolicited medical-information inquiries), so that volume trends per cluster become an insight feed for UC-X3 → deterministic AE / product-complaint screen: any hit routes to PV (UC-X1) with the clock started and human confirmation required — a classifier and a rule, never a prompt → retrieval from the medical approved-source registry (SRLs, label by market and version, publications — the same registry mechanism as UC-D4, different contents) → response draft with span-level citations in the inquirer's language → deterministic checks (unsolicited-request and off-label boundary rules per market, FDA vs. EFPIA; no promotional content; consent) → medical-information specialist disposition or MSL triage → response logged with provenance back to the CRM inquiry record.
Kitsdraft-then-check + rule-gate
GovernanceArchitecturally firewalled from the commercial estate (separate principal, registry scope, KG partition)
Key metricResponse time; citation completeness (target: 100%); off-label boundary violations (target: zero); AE-in-inquiry capture rate; cluster-to-insight lead time
Proven patternDoscierge draft-then-check; Enterprise Knowledge Vault retrieval at 99%+ over a curated registry
draft-then-checkrule-gate
Regional data scope for Medical Affairs agents. Medical Affairs illustrates the regional challenge more sharply than any other tower because a single workflow combines HCP personal data (GDPR / CCPA), patient-adjacent data (inquiry content, adverse-event reports), and public scientific data (abstracts, publications). The residency scope on each agent's card (§6) and the classification-aware fabric (§5.4) resolve it per agent, not per tower:
Medical agent
Regional data (stays in zone)
Global data (shareable)
Firewall and routing
MSL pre-call brief (UC-X4)
Interaction history (regional CRM partition), inquiry history (regional MI), HCP prescribing (regional)
Publications (PubMed), congress abstracts, KOL tier score (derived from public authorship and affiliation), approved SRLs by market
Medical-only principal; commercial interaction data excluded even within the same region. Card scope: REGIONAL-*, one instance per zone
Medical inquiry response (UC-X2)
Regional SRL library (market-specific off-label rules), HCP identity, inquiry text
Approved-source registry structure, indication / product reference, label version by market
AE-detection hook runs on every inquiry inside the zone; hits route to the regional PV agent with the clock started. Card scope: REGIONAL-*
Congress intelligence (UC-X5)
Attendee HCP profiles where regionally enriched (interaction history, territory)
Abstract text, competitor data, endpoint results — all public
Output splits: the intelligence report is SHAREABLE; HCP-linked insights are REGIONAL. Card scope: GLOBAL for the abstract pipeline, REGIONAL-* for the attendee-enrichment step
Field insight capture (UC-X3)
MSL note text — HCP PII plus scientific-exchange content — REGIONAL
Raw insights regional; themed, de-identified insights promoted to the global insight repository after the steward attestation (§5.4). Card scope: FEDERATED orchestrator over REGIONAL-* extraction specialists
PV signal detection (UC-X1)
Regional ICSR data, regional safety databases (EudraVigilance-mirrored EU cases, FAERS-mirrored US cases, NMPA-reported CN cases)
Aggregate cross-regional signal statistics via federated query
Mandatory HITL: a PV physician reviews every signal. Individual case transfer between regions — when a case genuinely must move — goes through E2B(R3) with the transfer mechanism recorded on the case, never through the agent. Card scope: FEDERATED
Evidence generation planning (UC-X6)
Evidence-gap repository entries linked to specific regional patient populations; investigator-site performance data (regional CRM partition)
Published literature, competitive landscape, congress abstracts (public), therapeutic-area strategy (internal)
Gap identification is GLOBAL (public evidence); study-concept development requires REGIONAL patient-density and site data. Card scope: FEDERATED orchestrator dispatching to REGIONAL-* data-access specialists
Medical content review (UC-X7)
SRL drafts with claim-level citations to regional label versions; regional off-label boundary rules (market-specific)
Medical approved-source registry, core label text, regulatory requirements (SHAREABLE)
Content creation is GLOBAL (label and regulatory reference); off-label boundary checking is REGIONAL-* per market because each market's approved indications and local regulations differ. Card scope: GLOBAL with REGIONAL-* rule-gate invocation per target market
The pattern is capture regionally, aggregate globally, act locally. The field-insight pipeline shows it end to end: the extraction specialist structures MSL notes inside the regional zone → the theme specialist produces de-identified themes with corroboration counts → themes are promoted to the global insight repository as SHAREABLE-DERIVED with lineage to their regional sources → TA leads in Basel see the global pattern on the medical strategy dashboard → MSLs in each territory act on local plans informed by the global theme, with the individual interactions that produced it never having left their region. The same three verbs describe signal detection (cases regional, signal global, label-change action local by market) and congress intelligence (attendee enrichment regional, abstract analysis global, KOL follow-up local).
CRM transition does not change the scope table. The canonical interaction model (§8.7 integration architecture) means the regional scope declared on each agent's card is CRM-agnostic. Whether the EU MSL works in Veeva CRM today or Vault CRM in 2028, the extraction specialist that structures interaction notes runs REGIONAL-EU against the same logical data product — msl_interaction — materialized from whichever CRM instance the region uses. The scope table above will survive the industry's CRM transition without a single card change, because the card declares a data residency scope, not a system binding. This is the practical consequence of abstracting the interaction model before the first agent ships.
UC-X3
Field Medical Insights & Medical Strategy
Decision. What the field is hearing, in aggregate, that should change the medical plan — and whether it is a pattern or three anecdotes.
Architectural pattern
Insight sources: MSL interaction notes (CRM), advisory-board minutes, congress
debriefs (UC-X5), inquiry clusters (UC-X2), medical-information trends
↓
Extraction specialist: structure each free-text insight → (HCP tier, product,
indication, topic, sentiment, evidence cited, date) with source provenance;
de-identify at the boundary; multi-lingual
↓
Theme specialist: cluster insights into themes (the pattern published as GenAI
theme analysis of field medical insights) — each theme carries n_sources,
KOL-tier weighting, geography, and a confidence that is CAPPED until
corroborated by ≥N independent interactions (the trust-tier rule of §5.3)
↓Deterministic path: theme-significance rules (volume vs. baseline, tier
weighting, recency, cross-region concordance) — versioned, signed by the
medical excellence / insights lead
AI path: synthesis agent drafts the periodic insights report per TA with
every theme cited to its interactions; drafts the "what changed since
last month" delta
↓Medical strategy dashboard (§5.1): themes, evidence gaps (→ UC-X6), inquiry
clusters, congress signals in one surface; TA-lead disposition (adopt /
monitor / dismiss) logged — the decision becomes a KG edge on the theme
Agent topology
Extraction specialist; theme specialist; significance rules; synthesis drafter; orchestrator publishing to the medical strategy dashboard through the command-center feed contract (§12). Kits: evidence-dossier + signal-and-NBA. The trap this design avoids is an insight tool that reports what the loudest MSL said. The confidence cap on uncorroborated themes and the tier weighting are what make the output a strategy input rather than a sentiment feed.
GxP frameNon-GxP; GDPR / HIPAA on HCP identity; insights are never re-used for promotional targeting (the firewall)
Key metricInsight-to-strategy time (quarterly → weekly); theme corroboration rate; share of medical-plan changes citing a structured insight
Proven patternFortune 50 pharma Medical Analytics platform — field and congress intelligence with NLP sentiment embedded in the operational platform rather than a standalone tool
evidence-dossiersignal-and-NBA
UC-X4
MSL Pre-Call Intelligence & KOL Engagement Planning
Decision. What this MSL needs to know before this interaction with this HCP — and, at territory level, which KOLs the medical plan says to engage, on which scientific objective, through which channel, next.
Architectural pattern
Trigger: scheduled interaction in the field CRM (call, virtual engagement,
congress meeting) or the MSL's territory-planning session
↓
Pre-call specialist (read-only against systems of record; ≤5 tools):
KG lookup — HCP golden record (specialty, affiliations, KOL tier, consent)
Interaction history — prior MSL interactions, topics, content shared, open follow-ups
Inquiry history — this HCP's medical inquiries and their responses (UC-X2)
Scientific activity — recent abstracts / publications / congress sessions
(UC-X5); trial involvement (→ Clinical KG: investigator)
Approved content — SRLs and medical materials matching the likely topics
↓Brief drafter: one page, every line cited — "what they asked last time, what
they published since, what we can share, what we cannot"; the cannot-share
list generated by the same off-label / unsolicited-request rules as UC-X2
↓
Engagement-planning specialist (territory level): scientific-objective coverage
per KOL vs. the medical plan; channel recommendation (face-to-face, virtual,
congress, advisory board, medical email) from the MEDICAL consent and preference
record — never from commercial engagement data; advisory board planning and
execution (attendee selection, topic agenda, logistics) feeds the insight KG
(UC-X3) and the transfer-of-value lineage (Sunshine Act / EFPIA)
↓MSL workbench (§5.1): brief + plan; MSL accepts / edits; NEVER auto-logged to
the CRM — after the interaction, the MSL's structured insight capture feeds UC-X3
Agent topology
Pre-call specialist; brief drafter; engagement-planning specialist; orchestrator. Kits: evidence-dossier + draft-then-check (the checker is the scientific-exchange boundary rule set). HCP medical engagement scoring — built on the golden record's medical partition (interaction history, inquiry patterns, congress presence, publication activity, advisory-board participation) — is a shared service consumed by the engagement-planning specialist; architecturally separate from commercial engagement scoring, with medical next-best-actions derived from medical signals only. Considered and rejected: one general MSL assistant with thirty tools spanning CRM write, content, and analytics — tool-selection accuracy degrades and the compliance boundary blurs (§5.2). Omnichannel for Medical is this use-case's territory-level half: medical engagement runs on its own consent and preference record, its own approved-content registry, and its own channel analytics, sharing only the HCP golden record with commercial (§8.6's firewall).
GxP frameNon-GxP; GDPR / HIPAA on HCP data; Sunshine Act / EFPIA disclosure on any transfer of value the plan proposes (advisory boards, speaker programs); the field CRM's offline-first constraint — the brief must be usable without connectivity
Key metricMSL prep time per interaction (hours → minutes); brief acceptance rate in coach mode (§18); scientific-objective coverage per KOL; post-interaction insight capture rate
Proven patternVeeva CRM ↔ enterprise MDM interface architecture at a Fortune 50 pharma — Account synchronization, golden-record resolution, call activity and territory alignment; ProtoCheck's bounded-specialist composition
evidence-dossierdraft-then-checkentity-resolution
UC-X5
Medical Congress Intelligence Pipeline
Decision. Which of the thousands of abstracts, posters, and presentations at this congress change the evidence narrative for our product, our competitors, or our KOLs — and which TA lead needs to know tonight, not in the post-congress report six weeks later.
Architectural pattern
Sources: congress abstract books and late-breaker feeds (ASCO, ESMO, AHA, ...),
presentation content where licensed, MSL congress debriefs, press signals
↓
Extraction specialists (scoped model tier — thousands of abstracts, marginal
cost matters, §5.2): abstract → (drug, indication, endpoint, result,
population, comparator, authors → HCP golden record, institution) with provenance
↓Deterministic checker (the generate-then-check discipline of UC-D4): numeric
extraction validated against the abstract text; endpoint vocabulary normalized;
product / competitor entity-resolved before anything is written to the KG
↓Medical & Safety KG: congress ──session──► abstract ──about──► product | competitor |
endpoint ──authored_by──► hcp (KOL tier); provenance-tagged; trust tier "operational"
↓
Signal specialists (isolated, each with its own eval record):
competitor-result watcher (vs. our label and open evidence gaps)
KOL-activity watcher (who presented what, in whose session)
safety-signal watcher (any adverse-event mention → PV screen, UC-X1)
sentiment-on-our-data (from debriefs, where present)
↓
Orchestrator → alerts to TA leads within hours, each with the abstract, the KG
neighborhood, and a proposed action (brief the MSLs · update the SRL · flag
to evidence generation); congress intelligence surface (§5.1); post-congress
synthesis drafted with citations
Agent topology
Extraction specialists (scoped tier); checker (MCP tool); three-to-four signal specialists; orchestrator. Kits: signal-and-NBA + draft-then-check + entity-resolution — authors are HCPs are investigators, the same golden record as UC-D2 and UC-C1. Embargo rules for pre-publication abstracts and late-breaker data are enforced as a compliance control: the extraction agent respects embargo dates, and KG edges carry an embargo_status flag that gates downstream distribution until the embargo lifts.
GxP frameNon-GxP; copyright, licensing, and embargo on congress content enforced in the approved-source registry; any adverse-event mention screened to PV
Key metricAbstract-to-alert latency (weeks → hours); extraction precision on numeric endpoints; KOL-activity coverage; share of post-congress actions traceable to a cited abstract
Proven patternFortune 50 pharma Medical Congress intelligence — near-real-time sentiment and competitive intelligence from live congress events, leadership reacting within hours rather than weeks; today's version replaces the NLP engine with LLM extraction gated by a deterministic checker and a provenance-tagged KG
Decision. Where are the evidence gaps for this product and indication — the questions HCPs keep asking that the label and the publications do not answer — and which of the proposed investigator-initiated studies, RWE analyses, or Phase 4 concepts closes the most important gap for the least money?
Architectural pattern
Gap detection: inquiry clusters (UC-X2) + insight themes (UC-X3) + congress
signals (UC-X5) + payer / HTA questions ──► candidate evidence_gap
(question, population, endpoint) with the sources that generated it
↓
Evidence specialist: what already answers it — internal studies (→ Clinical KG),
publications, RWE datasets (claims, EHR, registries; de-identified) —
evidence tiers rendered as in UC-R1's target dossier
↓Deterministic path: gap-priority rules (unmet-need weight, strategic-objective
fit, HTA relevance, feasibility flags) — versioned, signed by the medical
evidence lead; IIS-review completeness rules (protocol elements, budget,
fair-market value, no promotional linkage)
AI path: cited gap dossier; study-concept comparison ("three proposed IISs
address gap G; concept 2 duplicates a registered trial; concept 3 uses an
endpoint with prior regulatory precedent")
↓Evidence-generation committee reviews → fund / decline / merge logged with
rationale; funded concepts become study nodes in the Clinical KG; the
publication plan is drafted against the approved-source registry of study
data (draft-then-check)
Agent topology
Gap-detection orchestrator (consumes the X2 / X3 / X5 feeds); evidence specialist; rule engine; dossier drafter; publication drafter + checker; systematic literature review (SLR) agent — screening with calibrated inclusion/exclusion thresholds, PICO extraction, evidence tables with source spans, reviewer adjudication on disagreements, and living-review change control (kit: classifier + extraction). Kits: evidence-dossier + rule-gate + draft-then-check.
GxP frameNon-GxP for gap analysis; GCP once a study is funded; publications under ICMJE / GPP guidelines with Part 11 sign-off on submitted manuscripts; SLR screening thresholds documented in the protocol; RWE under data-use agreements enforced by the PII scope class (§6)
Key metricGap-to-decision time; share of funded studies traceable to a structured gap; publication review cycles; duplicate-study avoidance; SLR screening time (weeks → days)
Proven patternFortune 50 pharma clinical-studies assessment and funding intelligence — post-commercialization tracking of investigator activity and indication-expansion signals to decide which studies to fund; clinical data integration across 100+ studies on SDTM/ADaM
evidence-dossierrule-gatedraft-then-check
UC-X7
Medical Content Review & Scientific-Exchange Compliance
Decision. Is this medical asset — an SRL, a slide deck, a congress booth panel, a medical letter — approvable: every claim supported by an approved source, fair balance met, no promotional drift, correct for this market — before the medical review committee meets?
Architectural pattern
Draft asset (Vault MedComms / PromoMats) → extraction agent structures claims and references → deterministic path: claim-to-source rules (each claim must trace to a label statement, an approved study section, or a peer-reviewed publication in the registry; reference currency; fair-balance and scientific-exchange rules per market; an off-label language detector as a rule, not a prompt) → AI path: cited review memo naming the three findings a reviewer is most likely to challenge → medical / legal / regulatory reviewer disposition → approved asset registered with claim-level provenance, feeding UC-X2 and UC-X4 as approved content.
Build vs. buy. Where the content platform ships its own review agents — vendor quick-check and content agents now exist for promotional material, and specialist MLR-automation vendors target large reductions in review labor — the platform consumes them under the platform contract (UC-H1) for first-pass triage and builds only where vendor coverage stops: the enterprise's own claim-substantiation rules, the medical (non-promotional) boundary, and the cross-asset consistency check ("this SRL and this slide deck make different statements about the same endpoint"). Human sign-off stays.
Kitdraft-then-check — the same Doscierge pattern as UC-C2, applied to medical rather than promotional material
GxP frameFDA / EFPIA scientific-exchange rules; Part 11 on the approval record
Key metricMedical review cycle time; first-pass approval rate; cross-asset inconsistencies caught before committee (rising at first, then falling as the registry cleans up)
draft-then-check
Medical AI Acceleration — the Lab → Factory model applied to this tower. A Medical Affairs AI portfolio typically arrives as three or four standalone proofs — an inquiry-clustering tool, an insights-theming tool, a content-drafting assistant — each with its own retrieval, its own prompts, its own vendor. The acceleration step is not another proof; it is the shared substrate: one medical semantic layer and canonical interaction model (above), one approved-source registry with medical contents, one firewall, one telemetry plane feeding TCO, one registry of bounded specialists. A sequence that has worked: (1) a demand study across the seven decisions above — most Medical Affairs calendars are dominated by Type 1 waste (finding the prior interaction, re-deriving the approved answer, reading the abstract book); (2) UC-X4 first, because the pre-call agent is the substrate-building first product — it forces the KG, the approved-source registry, MDM resolution, and the canonical interaction model into existence, all with the lowest validation burden (read-only, advisory, no regulated record written); MSL brief acceptance in coach mode is the adoption signal the whole portfolio needs; (3) UC-X2 second, because the MI response assistant is the highest-volume, highest-ROI product that inherits the registry, off-label rules, and AE classifier the first one built; (4) UC-X3 and UC-X5 third, because they turn the field and the congress into structured KG edges that UC-X6 then reasons over; (5) UC-X7 can run in parallel with (3)–(4) as a buy decision where the content platform's own agents are in place to triage — the enterprise builds only the claim-substantiation judgment layer; (6) the PV pipeline (UC-X1) is third-wave, co-owned with Safety, proving the same kits survive the hardest clock and the most-inspected process. The promotion gate is unchanged from §7; the validation vote here belongs to Medical Compliance and Legal rather than Quality, and the readiness vote belongs to the MSLs and medical-information specialists who will use it. The portfolio is measured on the seven decision latencies, brief-acceptance rate, citation completeness, and cost per decision — the same scorecard as every other tower (§18).
Transition. Seven towers, and one thing they have in common: none of them works without the horizontal underneath.
8.8
Enterprise Horizontal
The layer is infrastructure; GxP applies to what consumes it · GAMP 5 infrastructure qualification · Part 11 where relevant
Two use-cases serve every tower. Neither is a use-case in the traditional sense; both are the shared substrate that makes the towers cheap.
UC-H1
IT Operations & Partner-Built Agents (The Platform Contract)
Decision. How does the platform govern agents that a partner builds and operates — incident triage, change management, access provisioning, cloud cost — so that the enterprise does not end up with three incompatible agent estates?
Architectural pattern
Agent platform contract — published by the enterprise AI platform team:
• Agent as a first-class identity principal (G5)
• Tool access via MCP servers with policy enforcement (allowlist per identity)
• Centralized observability and tracing (G4)
• Registry card, eval record, and approval workflow required for production (§6)
• Part 11 audit trail where change management touches validated systems (G3)
↓
Partner-built agents (incident triage, change management, access provisioning,
cloud cost) deploy INSIDE the contract — same identity model, same MCP tool
access, same observability, same gate as internal agents
↓Platform-enforced controls: tool allowlist, rate limiting and cost governance,
execution history in the central trace store, evaluation record before production
Key principle
Contract over bespoke. Partners build inside the platform patterns, not around them. The same contract governs the CRM vendor's commercial agents (UC-C1) and the IT partner's operations agents — which is what makes "three estates" into one governed estate.
GxP frameGAMP 5 infrastructure qualification; Part 11 where relevant
Key metricAgent-estate coherence — percentage of agents governed under the platform contract (target: 100%)
Proven patternMCP server published on the MCP registry (AAIF); 18+ IaC modules underlying 8 production systems
Decision. How fast can a new use-case get its first grounded, cited answer? The semantic layer's KPI is days from a new use-case request to its first grounded response — and it fails if it is either a central bottleneck or a set of disconnected team-built graphs.
Architectural pattern
Federated ontology governance (§5.3):
• Thin core ontology — centrally owned
• Domain ontologies — team-owned under a published modeling standard
↓
KG-as-a-Service: query, provenance, versioning APIs via MCP; any agent calls it;
none builds its own retrieval
↓
Curation agents: auto-ingest from data products; entity extraction and linking;
relationship discovery; freshness scoring and staleness flags
Governance agents: ontology drift detection; orphan entities; cross-domain
consistency; trust-tier audits
Retrieval agents: provenance-first semantic search across all domain KGs;
federated query
↓
Aligned to public standards — CDISC, IDMP, MedDRA, MONDO, ChEBI, FHIR, ISA-88/95,
GS1 — rather than invented
GxP frameThe layer is infrastructure; GxP applies to what consumes it
Key metricTime-to-first-grounded-answer (target: days); reuse ratio of core entities across domain KGs; provenance completeness
Proven patternEnterprise Knowledge Vault — powers all 8 production systems; RAG accuracy 75% → 99%+; 10,292 KG triplets
KG-as-a-Service
The federated ontology, explorable
Core entities in the center; seven domain rings around them; every entity and predicate named in the tower fragments above, plus the cross-domain joins. Click a node to light its neighborhood and see which use-cases consume it and which data products it materializes from. Click a domain label to isolate that subgraph. Pick a scenario to watch a traversal and the dual-path reasoning it produces. Esc resets.
100%
Normativeregulations · signed rules · 1.00Structuralontology · schema · 0.95+Operationalsystem-of-record rows · 0.90+Carry-forwardprior decisions · 0.80–0.90Enrichmentanalytics · 0.70–0.90Mined / Inferredprocess intelligence · HARD CAP 0.75
Interactive ontology explorer: core entities in the center, seven domain rings around them, a node click highlighting every domain that touches it — so that the reader sees, for example, that Batch is a core entity consumed by Product Development, Manufacturing, Supply Chain, and Medical (via product complaints), and that HCP is consumed by Development (investigator), Commercial, and Medical.
Eight towers. Twenty-plus use-cases. And, counted honestly, five kits, one governance plane, one gate, one graph. That count is the next section.
Section 9 · Enterprise Coherence
Factory over snowflakes — what gets shared
The platform's value is reuse. The previous section showed twenty-plus use-cases; this section shows how few distinct things they are built from — and makes the economic case that follows from that count.
9.1 The coherence matrix — what it shows and how to read it
Each filled cell means the use-case consumes this shared service — built once, owned once, validated once — rather than building its own.
The matrix answers one question per cell: does this use-case consume this shared platform service, or does it have to build its own? A filled cell means the use-case consumes the shared service — the service exists once, is owned once, is validated once, and the use-case is configured onto it. An empty cell means the service is not relevant to that use-case. There are no cells for "builds its own copy," because the platform rule is that there is no such option.
Columns are use-cases, grouped by tower and labeled with their §8 codes. Rows are shared platform services, grouped by the layer they live in. The final two columns show how many use-cases consume each service (the reuse count) and which team owns it — because a shared service without an owner decays. (Hover a column header for the use-case name; hover a row for the layer's role; the legend beneath lists all 23.)
Column key. R1 Target ID · R2 Gen Chem · R3 Preclinical Safety · D1 Trial Design · D2 Site Selection · D3 Ops Precision · D4 Reg Documentation · P1 Control Strategy · P2 Tech Transfer · P3 Q12 Change · M1 APQR · M2 CPV · M3 PAT Platform · M4 Blend Endpoint · M5 Deviation/CAPA · S1 CoA Intake · S2 Supply Risk · C1 HCP Engagement · C2 MLR Review · X1 Pharmacovigilance · X2 Medical Info · H1 Platform Contract · H2 KG-as-a-Service
The five Medical Affairs use-cases added in §8.7 (X3 insights, X4 MSL pre-call, X5 congress, X6 evidence generation, X7 medical content review) are omitted from the matrix for width. They consume only existing rows — core ontology and KG service, entity resolution, approved-source registry, the rule-gate / draft-then-check / signal-and-NBA / evidence-dossier kits, stateful orchestration, the eval harness, the registry, the platform contract, the human-in-the-loop gate, and the command-center feed contract — and add no new platform service. That is the point: a tower with seven use-cases and zero new rows.
How to read it in thirty seconds
Look down the Reuse column. Four services are consumed by essentially everything — the evaluation harness, the registry, the platform contract, and (with one exception) the core ontology and KG service. Those are the platform's foundation and are built first (§13, Phase 1). Then look across the Manufacturing columns (M1–M5): five use-cases in the most heavily governed tower in the enterprise, and not one row is new — every platform service they consume already exists by the time Development has scaled, including the model registry, which Development's simulation and prediction agents stood up first. What Manufacturing adds is its own data products and its own KG, not new platform services. That is the reuse dividend, made visible.
The same matrix as information flow
The dot matrix answers "does this use-case consume this service?" It does not show how services flow between layers. Hover a service band to see which towers consume it; hover a tower to see what it consumes; click a flow to list the use-cases at that intersection. Line thickness is the number of consuming use-cases.
Hover or click a flow to inspect it.
9.2 The reuse dividend — the economic case for one platform
Without sharing, each use-case reinvents its own KG, its own evaluation, its own governance — at an estimated $2–5M per standalone implementation. With the shared substrate, the marginal cost of use-case N+1 drops 40–60% versus use-case N. By use-case five or six, new agents are configuration, not projects. The platform's value is not in any single use-case; it is in the compounding reduction of time-to-value as the registry fills and the graph deepens.
There are two ways to lose the dividend, and the matrix guards against both:
Snowflakes. A use-case that builds its own retrieval, its own eval, its own audit trail — usually because the platform service was not ready when the team needed it. The mitigation is sequencing (§13) and the Lab's time-to-first-grounded-answer metric: the platform must be faster than going around it.
Monoliths. The reverse failure — one enormous agent that "does manufacturing." §5.2 made the case: performance degrades, token costs compound 30–70% annually, security boundaries blur, and validation becomes impossible. The platform pattern is composition: bounded specialists with narrow scopes and minimal tool sets, a coordinator that delegates, and a registry that makes the monolith structurally impossible. When connectors fail, degraded-mode design automatically caps confidence on affected data, actions fall below autonomy thresholds, and the human takes the driving seat — no manual intervention required to fail safe.
There is a demand-analysis reading of the dividend that is easier to defend to a CFO than a cost-per-project figure. A snowflake does not cost $2–5M once; it generates failure demand for years — the re-mapping when two ontologies disagree, the second validation package for the same rule engine, the reconciliation meeting when the CPV agent and the deviation agent describe the same batch differently. Those are Type 2 activities: they exist only because the platform service was not built once. The coherence matrix is therefore a map of failure demand avoided, service by service. The reuse count in the right-hand column is the number of times a system condition — per-project retrieval, per-project eval, per-project audit — was not recreated. Measured this way, the marginal-cost curve in §18 is not a productivity claim; it is the enterprise's failure-demand ratio falling as the registry fills.
Every one of these shared use-cases reasons the same way. That way is the next section.
Section 10 · Dual-Path Reasoning
Deterministic where you can, generative where you must, merged with provenance
Every use-case in the platform — discovery or GMP, research or field — follows the same reasoning architecture: deterministic where you can, generative where you must, merged with provenance. This is not a design preference. In a domain where every output may have to be explained to an inspector, it is the only architecture that can be both fast and defensible.
EVIDENCE — KG + data products + approved-source registry
↓ ↓
Deterministic path
Expert-calibrated rules — owner, version, calibration record
Provisional → confirmed: ships when an SME signs
Reproducible: same input → same output
Validatable: GAMP 5 Cat 5, conventional IQ/OQ/PQ
Inspectable: every rule readable, every firing traceable
Generative path
KG-grounded LLM — reasons over the semantic layer, not hallucinated context
Span-level citations: every claim traces to the approved registry
Refuses unsupported claims — architectural, not prompted
Controlled by grounding + checker + human — not by validating the model
↓ ↓
MERGED ASSESSMENT — one reviewable objectfindings + draft + provenance + confidence → human decides → decision + rationale become the audit record (G2)
Where each path lives, by tower. In Research, the deterministic path is evidence-tier scoring and structural alerts; the generative path is the target dossier and the liability brief. In Development, it is the 500-rule protocol check and the document checker versus precedent synthesis and drafting. In Manufacturing, it is validated SPC/capability queries and batch-record completeness versus trend narratives and root-cause hypotheses. In Commercial, it is promotional-compliance rules versus the pre-call plan. In Medical, it is seriousness/expectedness rules versus the case narrative. The proportion shifts — Manufacturing is deterministic-heavy, Research generative-heavy — but the shape never changes, which is why one validation strategy covers all of it.
Intelligent decision support. Human decides. System explains.
Intelligent decision support. Human decides. System explains. Every recommendation carries rationale, provenance, and an owner. The goal is to shift the performance of many individuals closer to the performance of the best. No black-box autonomy in regulated workflows — and, because the decision is captured (G2), the platform learns what "the best" looks like from the humans who are.
The same two paths are what make the system validatable.
Section 11 · Validation Strategy
"You can't validate an LLM under GAMP 5" is the wrong framing
You can validate the controls around the LLM. The question an auditor asks is "which control applies to which component?" — and the answer is three distinct strategies for three distinct system components, applied to every use-case in this article.
Control. Same input → same output. Conventional IQ/OQ/PQ. Each rule version-controlled with a named owner and a calibration record. Change control on every rule change.
CSA-aligned controls (FDA Computer Software Assurance, final September 2025 — device guidance, applied to pharma by analogy)
Control. Controlled by grounding (approved-source registry — the only citable source), checking (the deterministic path post-processes every output), and human review. The model is qualified against its context of use; deterministic controls bound residual risk. Model version pinned; eval harness measures P/R/F1 against labeled ground truth before any deployment and on every release. Draft EU GMP Annex 22 (July 2025) is scoped to static, deterministic models and excludes dynamic (self-learning) and generative/probabilistic models from critical GMP use — the dual-path rationale this architecture implements: the critical decision stays on the deterministic path, and the LLM works where a qualified human reviews its output.
ICH Q2/Q14 posture · model registry
Statistical / chemometric models
MSPC, PLS, enrollment simulators, ML scorers.
Control. Analytical-method lifecycle, not software validation: NOC set, cross-validation, held-out detection test, QA approval in the model registry, performance monitoring against the reference method.
21 CFR Part 11 · EU Annex 11
Human sign-off
Every output that enters a regulated record.
Control. Electronic signature. The human is not a rubber stamp; the system is designed so the human has the context — evidence, findings, confidence — to decide. Edits captured as preference signal.
Risk-proportionate by design. CSA's critical-thinking approach means the depth of assurance follows the risk of the decision. The blend-endpoint advisory (UC-M4), which runs beside a still-validated fixed-time step, carries a lighter package than the same endpoint in RTRT mode. The APQR narrative agent (UC-M1), whose output a QA reviewer signs, carries a lighter package than the SPC query layer, whose numbers the narrative cites. The governance manifest on each agent card (§6) records which posture applies, so the validation package is a rendering of the card, not a document written from scratch.
A note on honest complexity
This validation posture is defensible but not established precedent. FDA's 2025 draft guidance on AI in regulatory decision-making does not yet address the hybrid model explicitly; EMA's adopted reflection paper (September 2024) leaves the same room. The CSA guidance (issued for devices, and applied to pharma manufacturing by analogy) provides the framework for risk-proportionate validation that avoids over-testing deterministic systems while establishing appropriate controls for AI components. Early movers who articulate this posture clearly to Quality and CSV — and document it before the first agent enters a GxP workflow — will set the standard. Those who retrofit it will spend twice as long.
Continued in Part 2. How to prove this posture is the subject of Part 2. It runs the standard CSV/CSA lifecycle stage by stage on ProtoCheck and Doscierge (§17), then works through the acceptance criteria against a measured human baseline, the eight lifecycle controls that carry CSV into AI, locked vs. adaptive models, drift, monitoring and data integrity, question by question, in Validation Strategy, Applied: Agentic AI in GxP.
Validated agents in production generate signals. Where those signals land is the next question — and the answer is not a dashboard.
Section 12 · Agentic Command Centers
Dashboards show what is happening. Decision hubs decide.
Industry analysts have been pushing "command centers" and "control towers" in pharma for several years, and the leading sponsors have built them: a control tower over hundreds of studies that cut non-screening sites to 12% in six months and cut data discrepancies from roughly 30,000 to 8,000; manufacturing insight centers with real-time visibility across sixty-plus sites and twenty ERP instances. These are real and valuable. They are also, almost universally, dashboards: they show what is happening. The agentic platform lets them become decision hubs: surfaces where multi-domain agent feeds arrive as next-best-actions with rationale, impact, and owner; where the human decides; and where the decision is logged.
12.1 What a command center is, architecturally
A command center is an Experience-layer composition over registered agent feeds. It has no logic of its own beyond composition, prioritization, and presentation; every signal it shows was produced by a registered monitoring or orchestrator agent under the platform contract, and every action taken from it is captured as decision provenance (G2).
The feed contract. An agent that wants to publish to a command center registers a feed on its card: signal schema, severity taxonomy, SLA, owner, drill-through pointers. The command center subscribes by domain and role. This is how a new use-case appears on the center's surface without anyone building a new dashboard — it is configuration on the registry.
Cross-domain correlation is the reason command centers are agentic
A dashboard shows the CPV signal on line 2 and the supplier's late CoA and the deviation opened yesterday as three tiles. An agentic command center's orchestrator traverses the KG — batch → genealogy → material lot → CoA → supplier — and presents one NBA: "Line 2 EWMA shift on dissolution correlates with resin lot L-7781; CoA method differs from spec; deviation DEV-2026-0412 opened on the same lot; recommend hold on remaining L-7781 inventory pending QC re-test; owner: site QA head." That is decision velocity: the assembly that used to take three people a day is done before the morning meeting.
12.2 The five command centers — as the operator sees them
Five centers on every screen. Stat cards sense (system-of-record edges — every number is a graph query materialized for the person who needs to act). Alert cards decide (KG walks with min-path confidence). Drafted actions act (threshold-gated autonomy bands). Every approval or correction learns (decision provenance written back). Click a stat card to see the traversal that produced it; click an action to see which agent proposed it and on what evidence.
Site 042 stalled 18 days — catchment density adequate; competing trial opened 3 km away on 8/21; screen-fail rate normal. Diversity plan coverage at this site is 12% of commitment.min-path 0.84 · site → competing_trial → study (CT.gov) · confidence band: NOTIFY
Stock-out risk, Study 1187, depot EU-2 — 14 days of stock; resupply lead time 19 days; 2 sites enrolling above plan.min-path 0.93 · IRT_state → lead_time → enrollment_curve · band: AUTONOMOUS (draft resupply)
KRI drift, 4 sites, protocol deviation cluster — all four share the same CRO monitoring team; deviation type: visit-window.min-path 0.78 · monitoring_finding → site → cro_team (mined edge) · band: NOTIFY
Line 2 EWMA shift on dissolution correlates with resin lot L-7781 — CoA method differs from spec; deviation DEV-2026-0412 opened on the same lot yesterday; mined edge: this phase stalls 3× more often on Line 2.min-path 0.86 · batch → genealogy → material_lot → CoA → supplier · stalls_at mined edge capped 0.75 (evidence, not instruction)
CAPA-2025-118 recurrence — same root-cause class (gasket wear) recurred twice in 6 months on Line 2 granulator; effectiveness check failing.deterministic rule CE-004 · QMS rows · band: NOTIFY
Cold-chain excursion, lane EU-2 → site 3 — 2 h 40 min above 8 °C; stability data supports 72 h excursion budget; 1 shipment affected.S2 monitor · lane telemetry · stability_data (ICH Q1) · band: SUGGEST
3ActNBAs with rationale · impact · owner
4Learnsignal disposition · QA sign-off
07:41 EWMA shift Line 2 — investigate · QA signed hold (Part 11) · disposition feeds APQR (M1) automatically
Mon Cpk signal, drying temp, Line 1 — benign · "seasonal HVAC; within validated range" → known-benign pattern suppressed
Mon CoA exception, lot L-7702 — accepted with edit · method equivalence documented → rule RM-007 calibration note
14d CPV acceptance 81% (n=212) · batch-to-chart latency overnight · signal-to-disposition 6.2 h · deviations pre-empted by CPV: 3
A VP Quality reads this screen without clicking: (a) the CPV signal alerting is the EWMA shift on dissolution, Line 2; (b) the recommended action is a hold on remaining L-7781 inventory pending re-test; (c) it was proposed by mfg.quality.orchestrator composing mfg.cpv.signal-agent v2.3.0; (d) the evidence is the validated SPC query, the genealogy link to six batches, the CoA method flag, and DEV-2026-0412 — with the mined edge carried as evidence, capped at 0.75.
Competitor launch, indication B, 12 overlapping territories — formulary status unchanged at 2 of 3 major payers; our approved comparative claim set is current.min-path 0.88 · signal → territory → formulary_status → mlr_asset · band: NOTIFY
Formulary tier change, payer P2 — moved to tier 3 in 4 states; affected HCPs: 1,140 (golden record); consent valid for 71%.min-path 0.90 · formulary_status → hco → hcp (golden record) → consent · band: NOTIFY
5 MLR assets expire in ≤30 days — 2 carry claims cited by 38% of last month's NBAs.deterministic · asset registry expiry · claim usage counts · band: NOTIFY
Inquiry theme rising: dosing in renal impairment — de-identified count only; content response is a Medical decision, not a commercial one.firewalled aggregate · routes to medical strategy (X3), not to the field
3ActMLR-approved content only
4Learnbrand-team decisions
Tue Field redirection NBA — accepted · rationale: "launch confirmed; claims current" · rep uptake measured from Thursday
Congress abstract contradicts current dosing position, indication A — Phase III data from competitor presented at plenary; 3 KOLs in company advisory board are co-authors; medical response letter needed before media coverage cycle.min-path 0.91 · abstract → author → kol_advisory_board → company_position · band: ESCALATE
MSL insight cluster: off-label inquiries in 3 territories — 8 MSLs reported similar HCP questions on unapproved combination; pattern corroborated by X2 inquiry data (firewalled join on de-identified topic only).min-path 0.79 · insight → topic → inquiry_theme (de-identified) · corroboration score 0.72 · band: NOTIFY
Evidence gap: no RWE for patient subgroup in indication B — X6 gap analysis shows competitor has published; advisory board recommended prioritization; study concept drafted.X6 gap analysis · competitive landscape · advisory board minutes (§5.3 KG) · band: SUGGEST
3Actapproved-source constrained
4Learnmedical decision provenance
09:22 Medical response (indication C, congress) — approved · medical director signed; distributed to MSL field within 4h of abstract publication
Mon Inquiry response update (dosing in hepatic impairment) — modified · medical reviewer added SmPC §4.2 reference · approved-source version incremented
Mon MSL deployment NBA (territory 7) — rejected · "HCP availability conflict with congress travel" → scheduling constraint learned
14d Inquiry SLA compliance 94% · MSL brief acceptance 78% (↑ from 61% at launch) · insight-to-strategy cycle 11 days · PV clock overdue: 0
The network operations center for the agent estate
Observes the other four: every factory agent's traces, evals, acceptance curves, cost, connectors, degraded-mode events · registry lifecycle · gate queue
Feeds in fromClinical Operations CCManufacturing Quality CCCommercial Intelligence CCMedical Affairs Intelligence CC+ G4 observability from all 31 factory agents
1Senseagent health cards
2Decidedegraded mode · eval drift
mfg.cpv.signal-agent acceptance dropped 0.81 → 0.63 after the site-2 LIMS upgrade — test code TC-DISS renamed; conformed test_results contract v2.0 unaffected but dictionary mapping stale. Seen here before Quality sees it in an investigation.G4 + G2 · acceptance curve · data-product contract check · band: ESCALATE (below 0.70 gate threshold)
Eval drift: clin.doc.drafter F1 −3.4 vs release baseline — approved-source registry scope changed (new style guide v9); model pinned, prompt unchanged.nightly regression · model-loop · band: NOTIFY
Historian connector site 3 degraded 41 min — 3 agents auto-capped confidence to 0.55; 2 NBAs held for human; no manual intervention needed to fail safe.G7 degraded-mode design · automatic · resolved 06:52
Frontier-tier budget at 84% of monthly cap — one orchestrator routing 31% of calls to frontier (target ≤20%).G7 cost circuit-breaker · composition rule · band: NOTIFY
Q Agent graduation rate 58% · killed PoCs 3 (Lab health) · agent-estate coherence 100% under contract · inspection drill: full trail in 11 min
The fifth is the one most organizations forget, and it is the one that keeps the other four honest. The Enterprise AI Operations Command Center is the network operations center for the agent estate: it is where the Factory sees that the CPV agent's acceptance rate dropped after a LIMS upgrade changed a test code, before Quality sees it in an investigation.
12.3 How command centers relate to existing control towers
Where a control tower already exists, the platform does not replace it. The tower becomes a consumer of registered agent feeds — it gains NBAs with provenance and a decision log — and its existing analytics become data products the agents read. Where none exists, the command center is instantiated from the same experience template as any other workspace (§6), with feeds subscribed from the registry. Either way, the command center is a rendering of the platform, not a separate system to be integrated.
Command centers are where the platform's value becomes visible to leadership daily. Getting there is a sequence.
Section 13 · Implementation Roadmap
18–30 months, four phases — build what everything consumes first
Total estimated timeline: 18–30 months depending on enterprise size, regulatory requirements, and the number of domain use-cases. Each phase delivers usable value; later phases compound on earlier ones. The sequencing follows the coherence matrix (§9): build what everything consumes first.
Agent registry v1 with the first two kits (draft-then-check, rule-gate)
Data-product catalog bound to the existing lake
Lab environment stood up
The Lab's first grounded, cited answer for a new use-case in days; partner-agent onboarding possible; the three-estates problem addressed before it hardens
Registry live; contract published; first PoC promoted through a defined gate
2 — First-Wave Production (Development tower)
4–6 months
Regulatory documentation (D4) — highest ROI, most measurable
Trial design (D1) — direct amendment-avoidance value
Part 11 audit trail for regulated workflows
GAMP 5 / CSA validation for the first production agents
Approved-source registry
Clinical KG
Coach-mode metrics live
Document cycle time and amendments avoided — value the organization can see in the first quarter, validated against the hardest regulatory requirements (Part 11, GAMP 5)
Two factory agents with sustained acceptance ≥70%; validation packages accepted by Quality
3 — Factory Scale & the Manufacturing Data Foundation
6–9 months
Operational precision (D3) cross-study
Site selection (D2) on entity resolution
Commercial agents (C1) under the platform contract
IT-ops partner agents (H1) onboarded
Signal-and-NBA and evidence-dossier kits
Manufacturing data products — batch_phase_summary, LIMS conformed services, genealogy/CoA — with data fabric over regional LIMS
Manufacturing & Quality KG
Clinical Operations Command Center
Enterprise AI Operations Command Center
Portfolio-wide study operations; all three agent estates governed as one; the data foundation that makes the manufacturing tower a configuration rather than a program
Command centers live with decision logs; reuse ratio ≥70% on new use-cases; manufacturing data products at validated release
4 — Manufacturing, Supply, Medical & Enterprise
6–12 months
APQR (M1) parallel-run then cutover
CPV (M2) nightly, all CPPs, cross-site
PAT platform (M3) with blend endpoint (M4) as entry point
Deviation/CAPA (M5)
CoA intake (S1) and supply risk (S2)
Pharmacovigilance (X1) with clock enforcement
Medical information (X2)
Field Medical enablement — MSL pre-call intelligence (X4), insights (X3), congress intelligence (X5)
Evidence generation (X6) and medical content review (X7)
Product Development (P1–P3)
Manufacturing Quality and Commercial Intelligence Command Centers
Discovery use-cases (R1–R3) via the Lab pipeline
Full IQ/OQ/PQ for the regulated agent suite
The highest-GxP tower in production on infrastructure proven in Phases 2–3; enterprise-wide decision velocity measured on one scorecard
Platform TCO vs. standalone; agent graduation rate; inspection-readiness drill passed
Phase 1
3–5 months
Phase 1 — Semantic & Platform Foundation
What is builtCore ontology (IDMP, CDISC, MedDRA, ISA-88/95) · KG-as-a-Service via MCP · evaluation-harness framework · agent platform contract (identity, MCP tool access, observability) · registry v1 with draft-then-check and rule-gate kits · data-product catalog bound to the existing lake · Lab stood up · demand study per candidate use-caseValueFirst grounded, cited answer for a new use-case in days; partner-agent onboarding possible; the three-estates problem addressed before it hardens
Gate to nextRegistry live; contract published; first PoC promoted through a defined gate
Phase 2
4–6 months
Phase 2 — First-Wave Production (Development tower)
What is builtRegulatory documentation (D4) — highest ROI, most measurable · trial design (D1) — amendment avoidance · Part 11 audit trail · GAMP 5 / CSA validation for the first production agents · approved-source registry · Clinical KG · coach-mode metrics liveValueDocument cycle time and amendments avoided — value the organization can see in the first quarter, validated against the hardest regulatory requirements
Gate to nextTwo factory agents with sustained acceptance ≥70%; validation packages accepted by Quality
Phase 3
6–9 months
Phase 3 — Factory Scale & the Manufacturing Data Foundation
What is builtOps precision (D3) cross-study · site selection (D2) on entity resolution · commercial agents (C1) under the contract · IT-ops partner agents (H1) · signal-and-NBA and evidence-dossier kits · manufacturing data products — batch_phase_summary, LIMS conformed services, genealogy/CoA — with data fabric over regional LIMS · Manufacturing & Quality KG · Clinical Ops and Enterprise AI Ops command centersValuePortfolio-wide study operations; all three agent estates governed as one; the data foundation that makes manufacturing a configuration, not a program
Gate to nextCommand centers live with decision logs; reuse ratio ≥70% on new use-cases; manufacturing data products at validated release
Phase 4
6–12 months
Phase 4 — Manufacturing, Supply, Medical & Enterprise
What is builtAPQR (M1) parallel-run then cutover · CPV (M2) nightly, all CPPs, cross-site · PAT platform (M3) with blend endpoint (M4) as entry point · deviation/CAPA (M5) · CoA intake (S1), supply risk (S2) · PV (X1) with clock enforcement · medical information (X2) · field medical enablement (X3–X5), evidence generation (X6), medical content review (X7) · Product Development (P1–P3) · Manufacturing Quality and Commercial Intelligence command centers · discovery (R1–R3) via the Lab · full IQ/OQ/PQ for the regulated suiteValueThe highest-GxP tower in production on infrastructure proven in Phases 2–3; enterprise-wide decision velocity on one scorecard
GatePlatform TCO vs standalone; agent graduation rate; inspection-readiness drill passed
Why a demand study precedes the first PoC
The first measurement in Phase 1 is not an eval score; it is a demand study. For each candidate use-case, the pod samples the current work — a month of protocol-design requests, a quarter of deviation investigations, one APQR cycle — and classifies every touchpoint: value demand or failure demand, and for the latter, which system condition caused it. The study takes days, costs almost nothing, and does three things no architecture can. It gives the Lab a baseline the value-realization claims in §18 are measured against rather than estimated. It ranks use-cases by eliminable failure demand rather than by sponsor enthusiasm. And it surfaces the cases where the cheapest intervention is not an agent at all — a data-product owner, a template field, an SOP step. Instrument demand before automating it; the platform's credibility rests on being able to show a before.
Why documentation and trial design first
D4 has the highest volume, the most measurable cycle-time impact, and the most immediate business case. D1 has the highest per-incident cost avoidance (~$500K per amendment). Starting with these two validates the platform against the hardest regulatory requirements while delivering value the organization can see in the first quarter. Everything later builds on proven infrastructure.
Why Medical Affairs can be the first-wave tower instead
Where the sponsoring organization is Medical Affairs rather than Development, the same Phase 2 logic holds with different use-cases: medical information (X2) has the volume and the measurability that D4 has, and MSL pre-call intelligence (X4) has the adoption signal that D1 lacks — an MSL either uses the brief on Thursday or does not. The approved-source registry, the draft-then-check and rule-gate kits, and the Part 11 audit trail are the same Phase 2 deliverables; the Clinical KG is replaced by the Medical & Safety KG and the canonical interaction model (§8.7). The gate criteria do not change. What changes is who holds the validation vote — Medical Compliance and Legal rather than Quality — and the fact that the tower is in the middle of a CRM transition, which is the strongest argument for building the interaction model in Phase 1 rather than binding the first agent to CRM objects.
Why manufacturing's data foundation is in Phase 3, not Phase 4
The LIMS conformed services and the canonical attribute dictionary are the single most underestimated line item in a manufacturing AI program (§5.4). Starting them in Phase 3 — while the Development tower is scaling — means the Manufacturing tower in Phase 4 is a configuration exercise rather than a data-engineering campaign. Organizations that have done this report that the first product-site takes the longest and every subsequent one is materially faster.
Why manufacturing is the first global-by-default agent
Manufacturing and quality use-cases — CPV, APQR, deviation investigation, tech-transfer knowledge packages — are the correct first Factory deployment for a global-by-default agent. The reason is residency: manufacturing data (batch records, test results, process parameters, deviations) is not patient data and carries no PII classification. A CPV signal-detection agent can access batch-phase summaries from every site without triggering a de-identification gate, a consent check, or a cross-border transfer assessment. That makes it the fastest path to proving the platform's cross-site value — and the correct example when asked to show an agent that scales to every site on Day 1.
Why blend endpoint is the PAT entry point
The first use-case pays for the foundation: advisory mode needs no submission, the science is twenty years proven, and capacity value begins the day the advisory goes live. The enterprise PAT platform is then already built when the second and third PAT use-cases arrive.
The roadmap is pharma-specific in its regulatory sequencing. The patterns are not. But the roadmap does not exist in a vacuum — the enterprise already has a technology estate, vendor relationships, and in many cases competing agent programs. Where the platform sits in that landscape is the next question.
Section 14 · Market Context
Where this platform sits — not "instead of," but "the layer that makes them coherent"
An enterprise agentic AI platform is not built on green field. It is built on top of, beside, and sometimes in tension with an existing technology estate — cloud platforms, pharma-vertical vendors, enterprise SaaS, and increasingly, those vendors' own agentic offerings. The platform's competitive positioning is not "instead of" any of these. It is "the semantic substrate and governance plane that makes all of them coherent."
14.1 Horizontal AI and data platforms
Cloud-native AI platforms & enterprise data platforms
Cloud-native AI platforms (AWS Bedrock, Azure AI, Google Vertex AI) provide model hosting, orchestration primitives, and increasingly agent frameworks. They are the compute layer this platform runs on — not a competitor to it. The platform is model-agnostic by design (§5.2); scoped models and frontier models may come from different providers. What the cloud vendors do not provide is the pharma semantic layer, the GxP governance plane, the domain-specific rule engines, or the federated ontology. An organization that builds agents directly on a cloud AI platform without a shared substrate will have fast agents and no coherence.
Enterprise data platforms (Snowflake, Databricks, Palantir Foundry) provide the data lake, the compute, and increasingly the data-product layer. The platform's Data Layer (§5.4) sits on top of these investments — it consumes their conformed data products and data fabric services, it does not replace them. Palantir Foundry's Ontology is the closest analog to the semantic layer described here; the difference is that Foundry's ontology is tied to Foundry's compute and is not designed for the federated, domain-owned model that pharma's cross-functional structure requires. The platform described here can sit on Foundry, on Databricks, on Snowflake, or on a bespoke AWS lake — the semantic layer is the abstraction.
14.2 Pharma-vertical vendors
Clinical, manufacturing/quality, and commercial platforms
Clinical platforms (Veeva Vault, Medidata Rave, IQVIA Orchestrated Clinical Trials) are the systems of record for clinical operations, regulatory documents, and safety. They will remain systems of record. The platform connects to them as data sources and writes decisions back to them as workflow actions — never around them. When these vendors ship their own agents (and they are), the platform contract (UC-H1) is how those agents are governed alongside internal agents. The alternative is three agent estates with three governance models — which is the problem §2 described.
Manufacturing and quality platforms (Veeva QMS, MasterControl, SAP QM, Honeywell/GE historians) are the transactional systems for deviations, CAPAs, batch records, and process data. The platform's manufacturing tower (§8.4) reads from data products derived from these systems and writes investigation decisions back to the QMS. The canonical attribute dictionary (§5.4) is the bridge between regional instances and site-specific configurations.
Commercial platforms (Veeva CRM, Salesforce Health Cloud, IQVIA) increasingly have their own agent strategies. Salesforce Agentforce and Veeva's agentic roadmap will deliver field-force and commercial agents. The platform's value is the semantic layer underneath (HCP/HCO golden record, MLR content provenance, the medical/commercial firewall) and the governance plane that ensures partner-built agents are governed the same way as internal ones.
Medical Affairs platforms are where the vendor agent layer is arriving fastest, and where the transition state is most acute. Three facts shape the architecture. First, Veeva CRM reaches end of support on December 31, 2029, and every Veeva CRM customer must choose between Vault CRM and Salesforce Life Sciences Cloud — many will dual-run, with medical and field-MSL workflows on one and commercial engagement on the other. Agentforce Life Sciences (GA October 2025, with top-20 pharma early adopters) is Salesforce's agent layer for that estate. Second, Veeva AI Agents for PromoMats (December 2025 — a Quick Check Agent for pre-review and a Content Agent for drafting) and specialist MLR-automation vendors such as Falcon MLR (Copli, acquired by Veeva in June 2026; targeting roughly 70% reduction in MLR review labor) mean content-review triage is becoming a buy, not a build; the enterprise's differentiated work moves up to claim substantiation, scientific-exchange boundaries, and cross-asset consistency (UC-X7). Third, Vault MedComms remains the medical content system of record and the source for any medical approved-source registry. The platform's position across all three is the one §8.7 describes: a canonical Medical Affairs interaction model that makes the CRM decision a connector swap, a medical registry and firewall the vendor agents run inside of, and a telemetry plane that shows what each vendor agent costs per decision alongside the enterprise's own.
14.3 Emerging agentic AI competitors
LLM-native agent platforms & process intelligence
LLM-native agent platforms (OpenAI's agent offerings, Anthropic's MCP ecosystem, Google's Gemini agents, Microsoft Copilot Studio) provide the model and the agent primitives. The platform described here is built on MCP (Model Context Protocol) — agents call tools through MCP servers, the registry exposes discovery as an API, the platform contract is an MCP contract. The distinction is that MCP is a protocol; the platform is the pharma-specific implementation — the ontologies, the rule engines, the validation strategy, the governance plane, the eval harness, the domain-specific kits.
Process intelligence platforms (Celonis, Signavio) provide event-log mining and process discovery. The platform's process-intelligence layer (§5.3) fuses mined temporal and statistical findings into the KG as confidence-scored edges. Celonis-type tools are a data source for the semantic layer, not a competitor to the platform. Where they compete is on the "command center" concept — Celonis's Execution Management System is architecturally similar to the agentic command centers in §12, but without the multi-agent composition, the dual-path reasoning, or the GxP governance plane.
14.4 What the platform uniquely provides
No single vendor — horizontal, vertical, or agentic — provides all five of these together:
Capability
Who else provides it
What the platform adds
Pharma-specific federated ontology with domain-owned extensions
Palantir Foundry (single-vendor ontology)
Federated model; domain teams own extensions under a shared standard; aligned to IDMP/CDISC/MedDRA/ISA-88
Governance scoped to the decision, not the system; eight categories with named owners; agent-as-identity-principal
Composition-first agent framework with registry and service kits
MCP + A2A (Linux Foundation AAIF; protocol + discovery); cloud agent platforms (AWS Bedrock AgentCore GA Oct 2025; Microsoft Agent 365 GA May 2026; Entra Agent ID)
≤5-tool composition rule; registry with lifecycle states; five templatized kits; Lab-to-Factory promotion gate
Point solutions (Doscierge, specific vendors per domain)
Cross-domain rule framework; dual-path reasoning; the checker-as-MCP-tool pattern that any workflow can call
Cross-functional decision-velocity measurement
None at this scope
One scorecard from target ID to batch release to field call, measured on the same infrastructure (§18)
The platform's competitive position is not "better than Veeva" or "better than Palantir." It is: the integration layer, the governance contract, and the semantic substrate that makes the existing estate coherent — and makes use-case N+1 cheaper than use-case N.
Action for CDOs and platform leaders
Map your current vendor estate to the four-layer stack (§5). Identify which layer each vendor occupies. The gaps — typically the semantic layer, the governance plane, and the registry — are the platform's scope. The vendors are the platform's data sources and experience surfaces. Frame the business case as coherence cost avoided, not vendor replacement.
Section 15 · Organizational Design
The team shape by phase — how many, which roles, reporting to whom
The architecture (§5), the operating model (§7), and the roadmap (§13) define what gets built and how. This section answers the organizational question an ED building a greenfield division must answer: how many people, in which roles, at which phase — and what does the reporting structure look like?
The numbers below are calibrated for a top-20 pharma with 3-5 active therapeutic areas, 20-40 manufacturing sites, and an existing data-platform team. Scale down for mid-size; scale up for a company with more than 50 sites or a broader pipeline.
15.1 Headcount model by phase
Phase
Duration
Total headcount (new hires + allocated)
Composition
Key hires
1 — Foundation
3-5 months
~15-20
Platform core (6-8): architect, 2-3 platform engineers, 1-2 knowledge engineers, 1 eval/quality engineer, 1 DevOps/SRE. First Lab pod (4-6): 1 FDE, 1-2 agent engineers, 1 domain KE, 1-2 domain SMEs (allocated, not hired). CSV bench (2-3): 1 CSV lead, 1-2 CSV engineers (allocated from Quality or hired).
Platform architect (the ED). Lead knowledge engineer. First FDE (the person who will lead the first pod and train the next FDE).
2 — First wave
4-6 months
~30-40
Platform core grows to 10-12 (+ registry, observability, data-product binding). Second Lab pod (4-6). Factory nucleus (3-5): Factory lead, 1-2 SRE/ops, 1 release engineer. CSV grows to 4-5. Domain SME allocation doubles.
Factory lead (reports to ED or is the ED for Agentic Factory). Second FDE. Data-product engineer (the person who owns the canonical attribute dictionary).
3 — Factory scale
6-9 months
~50-70
3-4 active Lab pods (12-18 pod members). Factory operations team (8-12): SRE, observability, support tiers, model lifecycle. Semantic & KE team (6-8): core ontologists, domain KEs allocated per pod, entity-resolution engineer. Data engineering (4-6): data-product owners for manufacturing data products, data-fabric engineer. CSV (5-7). Command-center development (2-3).
Head of Semantic & KE (reports to ED or is the ED for that role). Command-center product lead. Manufacturing data-product owners (2-3 hires or allocations from MS&T).
4 — Enterprise
6-12 months
~80-120 (steady state)
4-6 active Lab pods rotating. Factory at steady state (15-20: ops, support, release, model lifecycle). Semantic & KE (8-12). Data engineering (6-10). CSV (6-8). Platform core (12-15). Domain champion network (~30 SMEs, allocated not hired).
A common miscalculation in pod sizing: estimating engineering effort as if the work is mostly agent logic — prompt engineering, tool binding, evaluation. In practice, 40–60% of Year 1 engineering effort goes to data governance infrastructure: classification tagging on every data product, de-identification pipelines per region, consent verification services, dual-CRM connector maintenance during vendor transitions, and the canonical attribute dictionary that unifies regional LIMS instances. The agent itself — the part that reasons — is the smaller fraction. Teams that staff for agent engineering alone discover the governance gap in month three, when the agent works in the Lab sandbox and cannot access production data. Budget for governance infrastructure explicitly; it is the binding constraint, not the model.
15.2 Reporting structure
Two models work; a third does not.
Model A — Under the Chief Data & AI Officer.
The platform reports to the CDAO alongside the data-platform team and the existing analytics org. Advantage: natural alignment with the data layer and existing data-product owners. Risk: the platform is perceived as "IT" and domain adoption is slower.
Model B — As a peer to the CDAO, under the CTO or CEO of R&D.
A new division — Lab + Factory + Semantic & KE — with its own ED-level leads. Advantage: signal that this is a strategic investment, not an IT project; domain teams see it as a peer. Risk: requires executive sponsorship to avoid turf conflict with the existing data org.
Model C (anti-pattern) — Distributed across functions.
Each function (R&D, Manufacturing, Commercial) builds its own agents with a loose "center of excellence" for coordination. This is how three agent estates happen, and it is how the $2-5M-per-snowflake cost structure persists. The platform exists to prevent this.
The four ED roles that best match this architecture:
Owns time-to-first-grounded-answer; accountable for PoC throughput and kill rate
ED, Agentic Factory
Factory operations, promotion gate, support model (L1/L2/L3), release trains, model lifecycle, command centers
Platform Strategy ED or CDAO
Owns agent graduation rate; accountable for production uptime, drift detection, and cost per decision
ED, Semantic & Knowledge Engineering
Core ontology, domain KG federation, entity resolution, KG-as-a-Service, data-product contracts, process intelligence
Platform Strategy ED or CDAO
Owns the semantic layer's freshness, coverage, and provenance completeness; accountable for time-to-ontology-extension
15.3 The scarce roles
Three roles are the binding constraint on hiring, and the platform design is intentionally shaped to reduce the demand for them:
Forward-Deployed Engineers
Engineers who speak the domain. The FDE model (§7.1) means you need fewer of them than you would need project managers, because one FDE carries context in both directions. But you need some, and they take 6-12 months to develop. Invest in the first two hires; train the next cohort through pod apprenticeship.
Knowledge Engineers with pharma domain context
Ontologists who understand CDISC, ISA-88, MedDRA, and the modeling standard. The core team is 3-5; domain KEs are allocated per pod and trained through the modeling standard, not through a PhD.
CSV Engineers who understand hybrid AI
Quality professionals who can validate a system where the deterministic path is GAMP 5 Cat 5 and the LLM path is CSA-controlled. This skillset barely exists in the market. The mitigation is joint workshops (§7.6) and doing the first validation package together with the platform architecture team, so the template exists for the next ten.
Action for ED candidates and division builders
The first three hires define the culture of the division: the platform architect, the lead knowledge engineer, and the first FDE. Get those right and the pods will form correctly. Get them wrong and every subsequent hire inherits the wrong patterns.
The organizational design is specific to pharma. The architecture underneath it is not.
Section 16 · Beyond Life Sciences
Pattern portability — the same architecture, run in a second industry
The architecture described above — the enterprise knowledge graph, MCP connectors, confidence-gated autonomy, coach-to-production graduation, ALCOA+ audit trails, dual-path reasoning, the registry — is not pharma-specific. The same patterns have been built for mid-market manufacturing (quote-to-cash, procure-to-pay) on the same stack:
Coach mode → production on rolling-window acceptance rates; thresholds ≥0.90 autonomous, 0.70–0.89 notify, 0.55–0.69 suggest, <0.55 escalate; test-driven design with regression suites at every gate
Decision velocity
Stateful orchestration with execution history; command centers
11-state Step Functions pipeline; MCP connectors to ERP/CRM/email; real-time retrieval with no data warehouse; iterative build-measure-learn cycles to production
Governance
Part 11 / Annex 11 / GAMP 5 / CSA
ALCOA+ audit trail on every decision; same serverless substrate
Why this matters for pharma. Cross-industry portability is how you know an architecture is real and not a framework slide. The same EKG substrate, the same confidence-gated autonomy, the same audit trail — configured for a different domain with different entity types and different compliance requirements — proves the "one platform, many use-cases" thesis extends beyond a single vertical. A pharma platform team that adopts these patterns is adopting something that has been run end-to-end, twice, in two industries.
And within pharma, two production systems are the wedge.
Section 17 · Wedge Applications
Built, not theorized
Architectures are easy to draw. The question is always: does it work? These two production systems were built on the platform described above — same semantic layer, same agent framework, same governance plane, same dual-path reasoning. They are the first two use-cases. Everything that follows inherits what they proved.
ProtoCheck · maps to UC-D1
ProtoCheck — Clinical Trial Protocol Compliance
Maps to UC-D1 (Trial Design) and is the pattern behind every rule-gate in this article. A multi-agent architecture where four specialized reviewers — safety, efficacy, regulatory, operational — each apply expert-calibrated rules against a protocol document, and a coordinator merges findings with provenance.
500+Expert-calibrated rules across 6+ therapeutic areas; FDA, EMA, ICH E6/E8, WHO4Specialist agents, each ≤5 tools — composition over inheritance, in productionCFR / ICHCitation traceability on every findingPatent pendingMulti-agent compliance system; presented at BioNJ 2026
Platform patterns proven. Composition over inheritance (four bounded specialists); confidence-gated routing; human-in-the-loop mandatory; provisional → confirmed rule calibration with named SME signatures; ALCOA+ audit trail on every finding.
Maps to UC-D4 (Regulatory Documentation) and is the pattern behind every draft-then-check in this article. The insight that drove it: everyone accelerates draft generation (80–85%); nobody deterministically checks against the rulebook. Generate-then-check is the architecture. The checker is an MCP tool — deterministic, inspectable, GAMP 5 Category 5.
970Deterministic rules10,292KG triplets across 29 domains — stability, impurities (ICH Q3A), DMF, dissolution, and more11-stateStateful processing pipeline, 25 functions, full execution historyP 92.31 / R 80.27 / F1 85.87Evaluation harness against expert-labeled ground truthMCPChecker-as-tool pattern — the deliberate Category 5 ceiling
Platform patterns proven. Dual-path reasoning (deterministic rules + KG-grounded LLM); checker-as-MCP-tool; KG-first pipeline with run-time grounding; confidence-scored provenance on every finding; trust-tiered triplet schema.
The enterprise precedent
Before either wedge application, the same patterns ran at a Fortune 50 pharma: a real-time analytics platform over 14B+ data points; a 4-agent GenAI platform scaled from zero to 150+ regulated users; a Customer MDM of 100K+ HCP/HCO records integrated with CRM and 20+ systems; clinical data integration across 100+ studies on SDTM/ADaM; a semantic layer that lifted retrieval accuracy from 75% to 99%+; and a 21 CFR Part 11 / Annex 11 MLOps framework. The wedge applications are what those patterns look like when built from scratch, on a serverless substrate, with a registry and a governance plane designed in from day one.
Why wedge applications matter. A platform without applications is a strategy deck. ProtoCheck and Doscierge prove the thesis: two production systems, different domains, built on the same substrate. Use-case #3 inherits the semantic layer, the governance plane, the agent framework, the evaluation harness, and the audit trail. The marginal cost drops 40–60%. That is the platform economics working — not projected, observed.
Observed how? By measuring the right things at the right time.
Section 18 · Success Metrics
Measure the right things at the right time
Most agentic AI programs fail because they measure the wrong things at the wrong time. Month-one ROI on an LLM agent is meaningless. What matters is a staged measurement framework that tells you whether the approach needs tweaking before you scale it — and that puts decision velocity on the scorecard beside accuracy and compliance, because velocity is the value.
Phase 1 — Coach Mode (Months 0–3): Leading indicators
Metric
What it tells you
Target
Agent acceptance rate
% of agent suggestions the human actually uses
≥70% sustained to graduate
Confidence calibration
Is the agent right when it says it is confident? (reliability diagram, Brier score)
Calibration error falling
Time-to-first-value
Days from deployment to first expert-accepted output
Days, not weeks
Escalation quality
Does the agent escalate the right things, or cry wolf?
Escalation precision rising
Time-to-first-grounded-answer (Lab)
Days from a new use-case request to a cited answer on the platform
Days
Failure-demand ratio (baseline)
Of the demand the use-case serves, what share is caused by the system's own prior failures — measured by classifying every touchpoint, not estimated
Measured within 30 days; the number Phase 2 claims are tested against
Findings the agent surfaces that humans missed — the real value signal
Rising, per class
Reuse ratio
% of shared platform components consumed by new use-cases vs. custom-built
≥70% is platform health
Marginal cost per use-case
Is each additional use-case cheaper than the last?
Falling 40–60% by use-case 5–6
Cost per decision
Tokens + compute per accepted recommendation, by agent (from the registry)
Falling as scoped models take share
Failure demand eliminated
Share of baseline failure demand whose cause the agent removed (evidence now structured, precedent now retrievable) vs. merely absorbed faster
Rising by class; cause-removed ≥ absorbed
Phase 3 — Production Scale (Months 9–18): Outcome metrics
Metric
What it tells you
Target
Inspection readiness
Can you produce the full audit trail for any agent decision in under 15 minutes?
Yes, drilled quarterly
Decision velocity, pipeline-wide
Clinical: months off enrollment; quality: days off batch release, months off signal latency; medical: clock compliance
Measured per tower on one scorecard
Agent graduation rate
% of coach-mode agents promoted to supervised autonomy within 6 months
Rising; killed PoCs counted as Lab health
Agent-estate coherence
% of agents — internal, CRM vendor, IT partner — under the platform contract
100%
Platform TCO
Total cost vs. N standalone implementations at $2–5M each
Documented dividend
Measuring coach mode: the critical leading indicator.
Coach mode is where you learn whether an agent is actually helpful or just generating plausible-looking output. Track the acceptance-rate curve over rolling two-week windows. A healthy agent shows a monotonically increasing curve as it learns domain boundaries. A flat or declining curve means the approach needs recalibration — retune the system prompt, add constraints, narrow the tool set, or revisit the model tier. The promotion gate from coach to supervised autonomy requires a sustained acceptance rate above threshold, verified by the domain expert team — not by the AI engineering team. This is decision provenance (G2) doing its job.
Why the failure-demand ratio keeps the other metrics honest.
Cycle-time reduction can be achieved by helping people cope faster with a broken system; it says nothing about whether the system is less broken. The failure-demand ratio does. If deviation investigations close 50% faster but the share of investigations caused by unseen drift is unchanged, the platform has industrialized the waste. If the ratio falls — because CPV now catches the drift before the batch fails — the platform has changed the system condition, and the cycle-time gain will hold even when the agent is switched off. The service-sector studies behind the method put failure demand at 25–75% of inbound demand in unexamined systems; pharma's evidence-assembly work exhibits every generator on that list.
Every metric above depends on people using the system. That is the last problem, and the hardest.
The hard part is getting hundreds of scientists, medical writers, process engineers, quality leads, and regulatory specialists to trust an agent enough to integrate it into how they actually work. This requires a deliberate adoption strategy — not a training deck and a launch email. It is the people in people-process-technology, and §7's operating model only works if this section is taken as seriously as the architecture.
The human side
Champion network. Identify 2–3 domain experts per use-case who co-design the agent's behavior. They are the internal evangelists — not IT, not the platform team. When a process scientist sees a colleague act on a CPV signal from the agent, that outweighs any executive mandate. They are also the L1 support tier (§7.4).
Transparent limitations. Publish what the agent cannot do alongside what it can. Teams trust tools that are honest about their boundaries. "This agent does not replace your judgment on eligibility criteria" builds more adoption than "AI-powered protocol optimization." The agent card (§6) is the honest source.
Feedback loops. Every agent in coach mode has a thumbs-up/thumbs-down on every output. Aggregate weekly. Share the dashboard with users — they see their feedback changing the system. This is the single most effective adoption lever, and it is the same mechanism as decision provenance (G2).
Graduated trust. Start with "agent prepares a draft, human does everything else." Move to "agent prepares and routes, human reviews." Then "agent executes routine actions, human reviews exceptions." Never skip a step. The blend-endpoint use-case (§8.4) is this principle applied to a process step: advisory beside the validated method, then replacement, then — only if chosen — real-time release.
The organizational side
Governance alignment before exposure. Quality, Regulatory, IT Security, and Legal review the agent's operating envelope before users see it. Retroactive governance kills adoption — teams hear "legal pulled the plug" once and never come back. The registry approval workflow (§6) makes this a step, not a negotiation.
Operating model clarity. Lab → Factory means exploration is funded differently from production. Lab pods have permission to fail fast and kill criteria to fail against. Factory agents have SLAs and on-call. Mixing the two creates confusion about what "done" means.
Skill development. Train domain experts to be agent supervisors, not AI engineers (§7.6). They need to read a confidence score, review a provenance trail, and know when to override — not tune hyperparameters. Two days, not six months.
Executive cadence. Monthly steering with the three metrics that matter at each phase (§18). No vanity metrics, no "number of agents deployed." The question is always: are humans making better decisions faster with the system than without it?
The anti-pattern: "We built it, they didn't come"
The number-one failure mode in enterprise AI is a technically excellent system nobody uses. Every use-case in this platform includes a co-design phase with domain experts, a coach-mode period with measurable acceptance rates, and a promotion gate that domain experts control — not engineers, not executives. If the medical writers say the documentation agent is not ready, it stays in coach mode. If the process scientists say the CPV agent cries wolf, it goes back to the pod. Full stop. Adoption is earned through demonstrated reliability, not mandated through organizational authority.
Section 20 · Evidence & Provenance
Three evidence sources — distinguished, so the architecture is credible rather than aspirational
This article's claims rest on three distinct evidence sources. Distinguishing them is what makes the architecture credible rather than aspirational.
Built and operated at enterprise scale · Fortune 50 pharma, 2010–2025
What was built and operated at enterprise scale (Fortune 50 pharma, 2010-2025)
These are patterns validated in a 15-year career at a Fortune 50 pharmaceutical company with $45B+ revenue, 30,000+ employees, and active FDA/EMA regulatory engagement:
Real-time analytics platform over 14B+ data points — the data-layer and semantic-layer patterns in §5.3 and §5.4, operating across clinical, commercial, and manufacturing data at enterprise scale
4-agent GenAI platform scaled from zero to 150+ regulated users — the agentic-layer patterns in §5.2, including composition, evaluation, and graduated deployment under GxP constraints
Customer MDM of 100K+ HCP/HCO records integrated with CRM and 20+ systems — the entity-resolution and golden-record discipline described in §5.3 and applied to site selection (UC-D2) and commercial (UC-C1)
Clinical data integration across 100+ studies on SDTM/ADaM — the clinical KG and data-product patterns in §8.2
Semantic layer that lifted retrieval accuracy from 75% to 99%+ — the KG-as-a-Service architecture in §5.3
21 CFR Part 11 / EU Annex 11 MLOps framework — the governance-plane controls in §5.5 (G3, G6)
These are the reference experiences. They prove the patterns work at scale, in regulated environments, with real inspectors. They are not current products — they are the institutional knowledge the architecture is built on.
Re-implemented from scratch · OrchestraPrime, 2024–2026
What was re-implemented from scratch on a modern substrate (OrchestraPrime, 2024-2026)
These are production systems built on the serverless, registry-first, governance-by-design substrate described in this article:
Doscierge — 970 deterministic rules, 10,292 KG triplets across 29 domains, 11-state stateful pipeline, 25 Lambda functions, measured eval harness (P=92.31%, R=80.27%, F1=85.87%), checker-as-MCP-tool, designed for GAMP 5 Category 5 validation, ALCOA+ audit trail. Maps to UC-D4. Proves: dual-path reasoning, KG-first pipeline, run-time grounding, trust-tiered provenance. Desirability score: 2/10 — technically validated, commercially unvalidated. This is an honest assessment, and it is why the platform thesis matters: a single-use-case product must find its own market; a platform component inherits value from the portfolio.
Mid-market manufacturing agentic platform — 1,786 KG triplets, 47 entity types, 50 predicates, 6 trust tiers, 11-state pipeline, MCP connectors to ERP/CRM/email, test-driven iterative development to production. Proves: cross-industry pattern portability.
MCP server on the MCP registry (Linux Foundation AAIF) — the interoperability contract. 18+ reusable IaC modules underlying 8 deployable system configurations.
Proposed · this article
What is proposed (this article)
The full cross-functional platform described in §5-§17 — the assembly of the above patterns into one governed, semantic-first operating system for AI-augmented decisions across the enterprise. The individual components have been proven; the assembly at enterprise scale has not. That is the gap between an architecture article and a shipped platform, and it is what Phases 1-4 of the roadmap (§13) are designed to close. The bet is that the components compose — and the evidence that they do is that two of them (ProtoCheck and Doscierge) already share the same semantic layer, the same governance plane, and the same evaluation harness, with a measured 40-60% marginal cost reduction on the second system.
What we got wrong, and what it taught us.
The first knowledge graph had 15,000+ triplets, and retrieval precision was 75%. The rebuild — curated, trust-tiered, 10,292 triplets — achieved 99%+ on a 212-question benchmark against expert-labeled ground truth. The lesson: for regulatory knowledge, curation beats scale. Every triplet should earn its place.
A 4-month spike tried to fine-tune a 7B/14B parameter model for regulatory compliance review. Best result: 0.597 on the detection criterion — all gates failed. The same KG and retrieval pipeline, connected to a frontier model, achieved 1.000. The lesson: don't spend months proving a small model can do what a frontier model does in one call. The methodology (structured retrieval, expert-calibrated rules) was right; the model capacity was the bottleneck.
Doscierge's first evaluation showed ~288 findings per test package. After iterative precision engineering cycles (enforcement gates, applicability predicates, correlator logic), it dropped to ~73 with higher precision. The lesson: the eval harness is not a report card — it is the engineering tool. A test-driven design approach — where every rule change must pass regression before merge — turns the eval harness into the primary development feedback loop. Without it, you are tuning by feel.
Section 21 · Contact
The architecture is defined. The patterns are proven at different scales. The question is yours.
This article exists to help you think through the architecture, the operating model, and the organizational design for an enterprise agentic AI platform — not to sell you one. The patterns described here were validated across three contexts:
At enterprise scale (Fortune 50 pharma, 15 years): Real-time analytics platform over 14B+ data points. A 4-agent GenAI platform scaled from zero to 150+ regulated users. Customer MDM of 100K+ HCP/HCO records integrated with CRM and 20+ systems. Clinical data integration across 100+ studies on SDTM/ADaM. A semantic layer that lifted retrieval accuracy from 75% to 99%+. A 21 CFR Part 11 / Annex 11 MLOps framework. These are the patterns — the ones that work at scale, in regulated environments, with real inspectors.
Re-implemented from scratch on a modern serverless substrate (OrchestraPrime): ProtoCheck (500+ expert-calibrated rules, multi-agent, patent-pending, presented at BioNJ 2026). Doscierge (970 deterministic rules, 10,292 KG triplets, 29 domains, 11-state pipeline, measured eval harness at P=92.31% / R=80.27% / F1=85.87%). A second full agentic platform in mid-market manufacturing on the same EKG substrate. MCP server on the MCP registry (Linux Foundation AAIF). 18+ reusable IaC modules. These prove the patterns compose on current infrastructure.
Proposed as an integrated enterprise architecture (this article): The full cross-functional platform described here, whose components have been proven individually but whose value is in their assembly as one governed, semantic-first operating system for AI-augmented decisions across the enterprise.
The honest assessment: the hardest parts of this program are not technical. They are organizational — hiring the right FDEs and knowledge engineers, earning Quality's trust before the first agent touches a GxP workflow, and resisting the pressure to scale before the foundation is proven. Those are leadership problems, and they are the ones this article is least able to solve on a page.
A 30-minute architecture conversation.
Not a sales call — a structured walkthrough of how this maps to your use-case portfolio, tower by tower, and where your registry and your first pod would start.