←
AI for Pharma & Life Sciences
Capable · M34 · lesson 34 of 35 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
System Prompts Grounded in ICH and FDA Guidance
📖
now learning

System Prompts Grounded in ICH and FDA Guidance

15 min

A platform team rolls out an enterprise medical-writing assistant across three therapeutic areas, and the system prompt at its core contains a single line of regulatory instruction: "Follow ICH guidance." It sounds responsible. It is, in practice, almost meaningless, because "ICH guidance" is not one document but a library of more than sixty harmonized guidelines spanning quality, safety, efficacy, and multidisciplinary tracks, each governing a different artifact with a different structure and a different set of expectations. A model told to follow ICH guidance, with no guidance named, does what models do with ambiguity: it produces something that has the surface texture of compliance, citing the right acronyms in the right tone, while conforming to no specific structure at all. The CSR it drafts is not aligned to ICH E3, because nothing told it E3 governs CSRs. The protocol section it produces ignores ICH M11, because M11 was never named. This lesson builds the alternative: a reusable system prompt pinned to the specific guidelines that govern the specific artifacts your function produces, and it shows why naming the guidance, rather than gesturing at it, is the difference between a system prompt that constrains the model and one that merely reassures the auditor.

What a System Prompt Actually Is and Why It Carries the Load

The system prompt is the block of instructions the deployment prepends to every conversation before your task ever reaches the model, and it does a disproportionate share of the work that determines whether output is safe. Where the per-task prompt from the previous lesson specifies one document, the system prompt encodes the standing behavior the model should exhibit across every document: the house style, the citation discipline, the refusal to invent, and, critically for this lesson, the mapping from artifact to governing guidance. Two deployments of the same underlying model can behave very differently because one has a system prompt that names the standards and one says "follow ICH," and the writer using the second tool will not understand why their drafts keep drifting structurally even when their per-task prompts are good.

Because the system prompt is set once and applied to thousands of drafts, the leverage is enormous and the failure is systemic rather than local. A weak per-task prompt produces one bad draft; a weak system prompt produces a bad base rate across an entire function's output for as long as it stays in production. This is why the system prompt is a governed artifact in a mature deployment, version-controlled and tested like any other configured component, rather than a paragraph someone typed during onboarding. The named author still owns each document, but the system prompt is the lever that sets how much verification each document will require, and a well-built one lowers the verification burden on every draft it touches.

The Failure Mode: "Follow ICH" Without Naming the Guidance

The central failure this lesson targets is precise and common: a system prompt that instructs the model to "follow ICH" or "comply with regulatory standards" without naming which guidance governs which document. The model cannot resolve this instruction, because compliance is not a generic property; it is conformance to a specific structure described in a specific guideline. ICH E3 prescribes the structure of a Clinical Study Report. ICH M4 prescribes the Common Technical Document organization. ICH M11 prescribes the harmonized clinical protocol template. These are different documents with different tables of contents, and an instruction that does not name them leaves the model to infer, from the genre of the request, which structure to produce, which is exactly the inference that goes wrong.

The danger is that the failure is invisible at the surface. A draft produced under "follow ICH" will mention ICH, will adopt a formal register, and will look, to a quick reader, like compliant content. The non-alignment lives in the structure: a CSR that omits a required E3 section, a protocol that does not follow the M11 ordering, a comparability narrative that ignores the ICH Q5E framework. Structural non-alignment is caught late, against the actual guideline, often during QC or, worse, by an assessor who reads for conformance as a matter of routine. The fix is not to tell the model to try harder; it is to name the guidance, so the model has a specific structure to conform to rather than a genre to imitate.

There is a second-order effect worth naming. When the system prompt does not bind artifacts to guidance, the citation discipline has nothing to anchor on either, because "cite the relevant guidance" is as empty as "follow ICH." A system prompt that names ICH E3 for CSR structure can also instruct the model to flag where its draft departs from the E3 section list, which turns the guideline into an active checklist rather than a decorative reference. Naming the guidance is the precondition for every downstream control that depends on it.

The Efficacy and Multidisciplinary Anchors: E3, M4, M11

The core of a clinical-writing system prompt is the binding of clinical artifacts to their governing efficacy and multidisciplinary guidelines. ICH E3 governs the structure and content of the Clinical Study Report, so the system prompt should state that any CSR section is drafted to ICH E3 and should name the E3 section the draft targets. ICH M4, with its M4E efficacy module, governs the Common Technical Document, so the system prompt should bind the Module 2.5 Clinical Overview and the Module 2.7 Clinical Summaries to M4 organization and M4E granularity, distinguishing the interpretive 2.5 from the detailed 2.7.3. These two guidelines cover most of the clinical writer's submission output, and naming them turns vague compliance into a specific structural target the model can reproduce.

ICH M11 is the newer and increasingly load-bearing anchor for protocol work. The M11 Clinical electronic Structured Harmonised Protocol, the CeSHarP template, prescribes a harmonized protocol structure now paired with controlled terminology, and a system prompt that names M11 lets the model draft eligibility criteria, endpoints, statistical considerations, and safety monitoring into the M11 section structure rather than into a sponsor's legacy template. For a clinical writer producing protocol sections, binding the draft to ICH M11 is the equivalent of binding a CSR section to E3: it gives the model the right table of contents to fill. The system prompt should name M11 explicitly and, where the deployment supports it, reference the controlled-terminology expectation so the model does not invent free-text where a coded value belongs.

The GCP, Safety, and Quality Anchors: E6(R3), E2B(R3), Q5E, Q9(R1)

Beyond the core clinical-writing artifacts, a function-grade system prompt names the guidelines that govern conduct, safety reporting, and CMC. ICH E6(R3), the revised Good Clinical Practice guideline, governs the conduct and documentation of trials and underlies artifacts such as the monitoring visit report and the deviation narrative; a system prompt serving clinical operations should name E6(R3) and the relevant sections, such as the monitoring expectations, so that drafted conduct documents align to the current GCP structure rather than to a superseded edition. Naming the revision is not pedantry: E6(R3) reorganized material relative to E6(R2), and a model anchored to the wrong revision produces a draft that is structurally dated.

ICH E2B(R3) governs the electronic transmission of Individual Case Safety Reports, and a pharmacovigilance system prompt should name it so that ICSR narrative drafting aligns to the E2B(R3) data structure and the expectations that surround it, which matters acutely as regions move to the R3 format. On the quality side, ICH Q5E governs the comparability of biotechnological products before and after a manufacturing change, and ICH Q9(R1) governs Quality Risk Management; a CMC-facing system prompt that names Q5E binds a comparability protocol to the right acceptance-criteria framework, and one that names Q9(R1) binds a risk assessment to the right QRM structure rather than to a generic FMEA the model pulls from its training data. Each named guideline converts a class of documents from imitation to conformance.

The discipline across all of these is the same: name the guideline, name the revision, and bind it to the specific artifact and section it governs. A system prompt that lists "ICH E6, E2B, Q5E, Q9" without revisions or bindings is only marginally better than "follow ICH," because the model still has to guess which governs what and which edition applies. The value is in the binding, the explicit statement that this artifact is drafted to that guideline at that revision, and the more of your function's output you bind, the lower the structural-error base rate across everything the tool produces.

The FDA-EMA Guiding Principles as the Cross-Cutting Layer

Above the artifact-specific guidelines sits a cross-cutting layer that a strong system prompt should also encode: the FDA-EMA Guiding Principles on the use of AI in regulatory work. These principles, including human-led decision-making, risk-based assessment, and transparency, are not document-structure rules; they are governance rules about how AI may participate in producing regulated content at all. A system prompt that internalizes them instructs the model to keep the human in the decision loop, to flag rather than decide on matters that require human judgment, and to make its contributions transparent and traceable, which is the behavioral counterpart to the structural conformance the ICH bindings provide.

Encoding the principles in the system prompt is what makes the tool's behavior defensible as a matter of governance rather than only as a matter of output quality. The transparency principle, for instance, supports the later cover-letter practice of disclosing AI involvement, and a system prompt that enforces traceable, source-grounded contributions produces drafts that are easier to disclose honestly. The risk-based principle justifies asking the model to raise its caution on higher-stakes content. Binding the ICH guidelines gives the model the right structure; encoding the FDA-EMA principles gives the deployment the right posture, and a mature system prompt does both, so that conformance and governance are built in rather than bolted on after an inspection finding.

Assembling the Reusable System Prompt

The reusable system prompt assembles these layers into a single governed block. It opens with a role and posture: the model is a regulatory and medical writer who keeps the human in the decision loop, flags missing data rather than filling it, and never states a result absent from the provided sources. It then declares the artifact-to-guidance bindings as an explicit table the model can consult: CSR sections to ICH E3; Module 2.5 and 2.7 to ICH M4 and M4E; protocol sections to ICH M11 with controlled terminology; conduct documents to ICH E6(R3); ICSR narratives to ICH E2B(R3); comparability protocols to ICH Q5E; risk assessments to ICH Q9(R1). It states the citation discipline, that every factual claim must be traceable to a provided source and flagged if it is not, and it encodes the FDA-EMA Guiding Principles as standing governance behavior.

The art is in keeping this block both complete and operable. A system prompt that names every ICH guideline in existence is as useless as one that names none, because the model cannot apply forty bindings to a single short request. The strong pattern is to bind the guidelines your function actually produces against, name their revisions, and instruct the model to ask which artifact it is drafting when the per-task prompt does not make the binding obvious, rather than guessing. This turns the system prompt into a router: the per-task prompt names the section, the system prompt supplies the governing structure for that section, and the two together give the model a specific target instead of a genre.

Finally, the system prompt is a versioned, tested artifact under change control, not a paragraph that drifts. Because it sets the base rate of structural correctness across thousands of drafts, a change to it is a change to the safety profile of the whole function, and it should be validated as such: documented, version-controlled, and tested against representative tasks before it ships, with its identity and version captured in the audit trail of every document it helps produce. The writer who can name which system-prompt version drafted a section, and show that it bound the artifact to the correct ICH revision, has a record that survives an inspection in a way that "we used the company AI" never will.

Testing the System Prompt Against Representative Tasks Before It Ships

A system prompt that binds artifacts to guidance is only as good as the evidence that it actually changes the model's behavior, and that evidence comes from testing, not from reading the prompt and finding it sensible. The discipline is to assemble a small, representative test set of real drafting tasks the function performs, a CSR Section 11 efficacy narrative, a Module 2.5.4 overview, an M11 protocol eligibility section, an E2B(R3) ICSR narrative, and to run each one through the candidate system prompt, then read the output against the actual governing guideline to confirm the binding took. A system prompt that names ICH E3 but produces a CSR draft missing a required E3 section has failed its own instruction, and the only way to discover that is to run the task and check the structure against the guideline's section list. Testing converts the system prompt from an aspiration into a measured control, and it is the step that distinguishes a governed deployment from one that merely has good intentions written into a hidden paragraph.

The test set should be designed to probe the failure modes the bindings are meant to prevent, which means it should include the cases where the model's training-data priors pull hardest against the correct structure. A biologic specification is a good probe for the Q6A-versus-Q6B binding, because the model's default is the more common small-molecule pattern, and a system prompt that correctly forces Q6B will produce the heterogeneity and process-impurity tests a Q6A draft would omit. An ICSR narrative is a good probe for the E2B(R3) binding and for the human-in-the-loop principle, because the model will otherwise produce a fluent causality assessment the regulation reserves for a human. Each probe task answers a specific question: does this binding survive contact with the model's strongest contrary instinct. A test set that only exercises the easy cases, where the model would have produced acceptable structure anyway, tells you nothing about whether the system prompt is doing any work, so the probes have to be chosen for difficulty, not for representativeness alone.

When a probe fails, the fix is almost always sharper binding rather than more words, and learning to read failures this way is the skill that keeps a system prompt operable as it matures. A CSR draft that drifts toward a journal-article structure under a system prompt that named ICH E3 usually means the binding was stated as a passive reference ("aligned to ICH E3") rather than as an active instruction ("produce the E3 section list and flag any required section you cannot populate from the sources"). A draft that applies the wrong revision usually means the revision was not named. The remediation is to make the binding directive and revision-specific, retest, and only then promote the change into the production system prompt under version control, with the new version and its test results recorded so that a future inspection can see not just that the system prompt named the right guidance but that it was tested and shown to enforce it. This testing loop is what makes the difference between a system prompt that reassures the auditor and one that actually constrains the model, and it is the practice that turns the artifact-to-guidance bindings from a list of acronyms into a demonstrated safety control.

Key Takeaways

  • "Follow ICH guidance" is almost meaningless because ICH is a library of more than sixty guidelines, each governing a different artifact. A model told to follow ICH with nothing named produces the surface texture of compliance while conforming to no specific structure, and the non-alignment hides in the structure where it is caught late, against the actual guideline.
  • The fix is to name the guidance and bind it to the specific artifact and revision it governs: CSR sections to ICH E3, Module 2.5 and 2.7 to ICH M4 and M4E, protocol sections to ICH M11 with controlled terminology, conduct documents to ICH E6(R3), ICSR narratives to ICH E2B(R3), comparability protocols to ICH Q5E, and risk assessments to ICH Q9(R1). The value is in the binding, not the list.
  • The system prompt carries a disproportionate share of the safety load because it is set once and applied to thousands of drafts. A weak per-task prompt produces one bad draft; a weak system prompt produces a bad structural base rate across an entire function for as long as it stays in production, which makes it a governed, versioned artifact rather than onboarding text.
  • The FDA-EMA Guiding Principles are the cross-cutting governance layer the system prompt encodes alongside the ICH bindings. Human-led decision-making, risk-based assessment, and transparency give the deployment the right posture, the behavioral counterpart to the structural conformance the ICH bindings provide, so conformance and governance are built in rather than bolted on after a finding.
  • A reusable system prompt is a router, not an encyclopedia. Bind the guidelines your function actually produces against, name their revisions, instruct the model to ask which artifact it is drafting rather than guess, and capture the system-prompt version in the audit trail so a writer can show a section was bound to the correct ICH revision.