AI-Assisted Clinical Study Protocol (CSP) Drafting Aligned to ICH M11
A clinical development team is standing up a Phase 2 study, and the protocol is the gating document. Nothing happens, no IRB submission, no site activation, no first patient consented, until there is a Clinical Study Protocol the regulators, the ethics committees, and the investigators can all read. In 2026 that protocol has a new and consequential shape. The ICH M11 guideline, the Clinical electronic Structured Harmonised Protocol or CeSHarP, reached Step 4 and moved into adoption, and it does something previous protocol templates never did: it fixes the table of contents, fixes the section numbering, and pins many fields to CDISC controlled terminology so the protocol becomes a structured, exchangeable, machine-readable document rather than free prose. This is exactly the kind of structure a language model thrives on, and exactly the kind of structure where an ungrounded model can do quiet damage. This lesson takes a target product profile and a prior Investigator's Brochure and drafts the high-stakes sections of an M11 protocol, the eligibility criteria, the objectives and endpoints, the statistical considerations, and the safety monitoring, with the source discipline that makes the draft an accelerant rather than a hazard. The protocol is the document every downstream artifact inherits from, so an error here propagates to the ICF, the CRF, the SAP, and the CSR. Drafting it well is the highest-leverage writing task in the chapter.
What ICH M11 Changes About Protocol Drafting
Before M11, every sponsor had its own protocol template, its own section order, its own headings, and a reviewer had to relearn the geography of each protocol they opened. M11 ends that. The template defines a fixed structure in which the level-one and level-two headings remain constant across all protocols, so the sequence becomes well known and predictable. Trial design is always Section 4. Inclusion and exclusion criteria live in Sections 5.4 and 5.5. Trial objectives, endpoints, and estimands sit in Section 6. Safety and its monitoring occupy Section 8. Statistical considerations, including sample size, are Section 9. A reviewer who knows the M11 map can navigate any compliant protocol instantly, and a model that knows the map can place content where it belongs. The structure is a gift to automation, because the hardest part of drafting, deciding what goes where, is now specified.
The second change is deeper. M11 is not only a document template; it is a data standard. The technical specification pins defined fields to controlled terminology governed by CDISC, so that a value like the sex of participants, the trial phase, the type of control, or an intercurrent-event strategy is not free text but a term drawn from a controlled list. This is what makes the protocol electronically exchangeable: a downstream system can read the structured fields without parsing prose. For the writer using a model, this matters enormously, because it means certain fields have a finite set of legal values, and a model that writes a plausible-sounding but non-conformant value has produced an error a validator will catch, or worse, a value that passes format but means the wrong thing. The discipline shifts from prose fluency to terminology conformance for the structured fields, and to source-grounded accuracy for the narrative ones.
The consequence for the workflow is that an M11 protocol has two kinds of content with two kinds of failure modes. The narrative sections, the rationale, the risk-benefit assessment, the detailed procedures, fail the way all AI writing fails: plausible invention that does not trace to a source. The structured fields fail a second way: a value outside the controlled terminology, or a controlled value that contradicts the narrative around it. A defensible M11 draft has to be checked on both axes, and the model that drafts it has to be steered on both. The lesson treats the protocol as a structured document first and a piece of prose second, because that is what M11 made it.
The Sources: Target Product Profile and Prior Investigator's Brochure
The source of truth for a protocol is different from the source of truth for a CSR. A CSR summarizes a completed trial with a locked database; a protocol describes a trial that has not happened yet, so its claims are not results but design decisions and safety facts inherited from what is already known. Two documents carry most of that inherited truth, and they are what you load into the context window before drafting. The first is the target product profile, the TPP, which states the intended indication, the target population, the desired efficacy and safety profile, the route and regimen, and the competitive context. The TPP is the design intent: it tells the model what the trial is trying to demonstrate, which constrains the objectives, the endpoints, and the population the protocol must define.
The second source is the prior Investigator's Brochure, which carries the accumulated nonclinical and clinical knowledge of the investigational product, and critically the Reference Safety Information that grounds the safety sections. The IB is where the known and potential risks live, where the pharmacology that justifies the starting dose lives, and where the expected adverse events that drive the safety monitoring plan are documented. A protocol's eligibility criteria, its dose rationale, and its safety monitoring cannot be invented; they must be derived from the IB. When you ask a model to draft eligibility criteria without the IB in the window, it will produce a generic, plausible list that reads like an oncology or a cardiology protocol from its training data, and that list may exclude exactly the wrong population or omit a contraindication the IB establishes. The IB is the safety ground truth, and a safety section drafted without it is the most dangerous output in the protocol.
There is a sequencing point that matters. The TPP constrains what the trial is for; the IB constrains what is safe and known. Load both, name them in a source manifest with version and date, and the model can draft sections that are anchored to design intent and to established safety. Load neither, and the model drafts a protocol for a generic drug that does not exist, fluent and wrong. The protocol is the one document in this chapter where the sources are not results to transcribe but knowledge to apply correctly, which makes the grounding discipline subtler and the verification more about judgment than about matching a number to a cell.
Drafting Eligibility Criteria From the IB and TPP
Eligibility criteria live in M11 Sections 5.4 and 5.5, inclusion and exclusion, and they are deceptively dangerous to automate because they read like boilerplate and are anything but. Each criterion encodes a safety judgment, a scientific judgment, or an operational constraint, and a wrong criterion either enrolls a patient who should have been excluded, a safety and ethics failure, or excludes a patient who should have qualified, a recruitment and generalizability failure. The model is fluent at producing criteria because criteria are highly patterned in the training corpus, which is precisely the trap: it will generate a confident, conventional list whether or not the list is right for this product.
The grounded approach prompts the model to derive each criterion from a named source and to mark which source justifies it. An exclusion for a specific hepatic impairment threshold must trace to the IB's pharmacology and known hepatotoxicity signal, not to the model's sense that oncology protocols usually exclude such patients. An inclusion defining the target population must trace to the TPP's intended indication. The prompt requires that each criterion carry a bracketed rationale source, that contraindications established in the IB appear as exclusions, and that the model flag any criterion it cannot ground rather than including it on convention. You then verify the list against the IB and TPP directly: every IB-established contraindication is reflected, every TPP population boundary is respected, and no criterion exists that the sources do not justify. A criterion the model invented because the corpus expects it is removed, because an unjustified exclusion narrows enrollment without a documented reason, and an unjustified inclusion may admit a patient the safety data warn against.
This is also where the structured-field discipline appears. Where M11 pins a field to controlled terminology, the model's value must conform. A population descriptor, an age-range field, or a condition coded to a controlled list has to use the legal term, and a free-text approximation that a validator will reject is a defect even when the prose around it is correct. You check the structured eligibility fields against the CDISC-governed terminology the template specifies, separately from checking the clinical justification of the criterion itself.
Objectives, Endpoints, and the Estimand Framework
M11 Section 6 carries the trial's objectives, endpoints, and, in a change that trips up unprepared drafters, the estimands. The estimand framework, introduced into protocol design by ICH E9(R1), forces the protocol to state precisely what treatment effect the trial will estimate, including how intercurrent events such as treatment discontinuation, rescue medication, or death are handled. M11 expects the estimand to be specified in structured form, with its attributes, the population, the treatment, the variable or endpoint, the intercurrent-event strategies, and the population-level summary, made explicit. This is exactly the kind of precise, attribute-by-attribute specification a model will happily approximate and quietly get wrong.
The danger here is subtle and specific. A model drafting Section 6 will produce a clean primary objective, a sensible primary endpoint, and an estimand that reads correctly, but the intercurrent-event strategy it chooses, treatment policy, hypothetical, composite, while-on-treatment, or principal stratum, is a design decision with major statistical and interpretive consequences, and the model has no basis for choosing it except pattern. An estimand whose intercurrent-event strategy does not match the TPP's question, or does not match what Section 9 will analyze, is an internal inconsistency that corrupts the entire trial. So the prompt requires the model to derive the objective and endpoint from the TPP, to state each estimand attribute explicitly, and to flag the intercurrent-event strategy as a decision requiring statistician sign-off rather than asserting it as settled. The model drafts the structure; the trial statistician owns the estimand, and the workflow makes that boundary explicit rather than letting the model's fluent default stand as if it were a decision.
Verification of Section 6 is cross-sectional. You confirm the primary endpoint in Section 6 is the one the sample size in Section 9 powers for, that the estimand's variable matches the endpoint, that the intercurrent-event strategy stated in Section 6 is the one the statistical analysis in Section 9 implements, and that the objective traces to the TPP's intended demonstration. An M11 protocol is a tightly coupled structure, and the estimand is the joint that connects the clinical question to the statistical method. A model can draft each section so it reads well in isolation while the joint is misaligned, and only a cross-sectional verification catches it.
Statistical Considerations and Safety Monitoring
Section 9, statistical considerations, includes the analysis populations, the analysis methods, the multiplicity handling, and the sample-size determination in Section 9.8. A model can draft a credible statistical section because the genre is patterned, and that is the hazard: a sample-size calculation that reads correctly but uses an effect size, a variance assumption, or an alpha allocation the statistician did not specify is a fabricated input wearing the costume of a derived result. The workflow treats the statistical section as model-assisted scaffolding that the trial statistician must populate and own. The prompt instructs the model to lay out the structure of Section 9 aligned to the SAP conventions and to the estimand in Section 6, to leave the numerical assumptions as explicitly flagged placeholders sourced from the statistician rather than inventing them, and to never assert a powered sample size without a stated, sourced assumption set. The statistician fills the assumptions; the model arranges them into compliant prose.
Safety monitoring, in Section 8, is where the IB grounding pays off most directly. The safety monitoring plan, the adverse-event collection, the stopping rules in Section 7.4, the dose-modification rules, the data monitoring committee charter triggers, must reflect the known and potential risks the IB establishes. A model that drafts a generic safety section omits the product-specific signals that are the entire reason this protocol needs particular monitoring. The prompt requires every monitored risk and every stopping rule to trace to an IB-documented signal or to an established class effect, and it requires the model to flag any safety provision it cannot ground. You then verify that every risk the IB calls out has a corresponding monitoring provision, and that the stopping rules and dose modifications are consistent with the IB's safety profile rather than with a generic template. A safety section that monitors for the wrong things, or fails to monitor for a known signal, is the kind of protocol defect that an IRB or a regulator flags and that, if it reaches the clinic, endangers participants.
Across both sections, the cross-references must resolve. The analysis populations defined in Section 9 must match the populations the eligibility criteria in Section 5 produce. The endpoints analyzed in Section 9 must match the endpoints defined in Section 6. The safety stopping rules in Section 7.4 must be consistent with the safety monitoring in Section 8 and the risks in the IB. M11's fixed structure makes these cross-references locatable, which is exactly why a model can be steered to maintain them and exactly why a human must verify them, because the model maintains the appearance of consistency more reliably than the fact of it.
Terminology Conformance and Structured Validation
Because M11 pins fields to CDISC controlled terminology, an M11 protocol can be validated in ways a free-prose protocol cannot, and this is an opportunity the workflow should exploit rather than treat as a burden. After the model drafts, the structured fields, the trial phase, the control type, the sex and age fields, the intercurrent-event strategy codes, the endpoint type classifications, are checked against the controlled terminology the technical specification governs. A value outside the controlled list is a conformance error a structured validator will catch before a human reviewer ever reads the prose, which means the writer can offload a category of checking to tooling and concentrate human attention on the judgments that tooling cannot make.
This split is the practical heart of M11-era drafting. Conformance of structured fields is a mechanical check, suited to validators and to deterministic tooling, and a model error there is cheap to find. Accuracy and groundedness of narrative content, the rationale, the eligibility justifications, the estimand strategy, the safety monitoring, is a judgment check, suited to a named human who reconciles each claim to the IB or the TPP, and a model error there is expensive to find and dangerous to miss. The writer who understands this division spends validator cycles on conformance and spends their own cycles on the clinical and statistical judgments, rather than reading every controlled-term field by eye while skimming the safety rationale that actually needs scrutiny.
The documentation closes the loop the same way every lesson in this chapter does. The AI use record pins the run, the model and version, the system prompt, the temperature, the timestamp, the TPP and IB versions loaded. The verification record shows that the narrative claims were reconciled to the IB and TPP, that the structured fields passed terminology conformance, that the cross-references resolve, and that the estimand and the statistical assumptions were signed off by the statistician who owns them. Tied to the protocol version under change control, these records make the M11 draft defensible under ICH E6(R3), which governs the conduct the protocol authorizes, and under the GCP audit that will eventually ask not whether AI drafted the protocol but whether a named human verified every design decision it contains.
What This Means for the Protocol Author
The protocol author who adopts this workflow gains the most leverage of any writer in the chapter, because the protocol is the document everything else inherits from, and an M11 protocol's fixed structure lets the model carry more of the scaffolding load than any other artifact. The model can place content correctly, maintain the section map, draft compliant prose, and arrange structured fields, in a fraction of the time a from-scratch draft takes. What it cannot do is choose the estimand's intercurrent-event strategy, set the statistical assumptions, justify an eligibility criterion against the safety data, or decide what to monitor, and the workflow's entire job is to keep those decisions with the humans who own them while letting the model accelerate everything around them.
The operating sequence is now familiar in shape and specific in content. Load the TPP and the prior IB so design intent and safety ground truth are in the window. Prompt for source-grounded sections where every eligibility criterion, every endpoint, every safety provision, and every statistical structure traces to a named source or is explicitly flagged as a human decision. Verify the narrative content against the IB and TPP for groundedness, verify the structured fields against CDISC controlled terminology for conformance, and verify the cross-references across Sections 5, 6, 7, 8, and 9 for internal consistency. Route the estimand and the statistical assumptions to the statistician for sign-off. Document the run and the verification, and sign only what the responsible humans have owned. The next lesson moves from authoring a new protocol to updating an existing Investigator's Brochure, the very document that grounded this protocol's safety, where new study reports and post-market safety data must be integrated with full traceability under ICH E6(R3).
Key Takeaways
- ICH M11 (CeSHarP) makes the protocol a structured, exchangeable document with a fixed section map and CDISC-governed controlled terminology. Trial design is Section 4, eligibility is 5.4 and 5.5, objectives and endpoints and estimands are Section 6, safety is Section 8, statistics and sample size are Section 9. The structure is a gift to automation and a new surface for conformance errors.
- The sources are the target product profile and the prior Investigator's Brochure, not a locked database. The TPP constrains what the trial is for; the IB carries the Reference Safety Information that grounds eligibility, dose rationale, and safety monitoring. A safety section drafted without the IB is the most dangerous output in the protocol.
- Eligibility criteria read like boilerplate and are not. Each criterion must trace to an IB-established contraindication or a TPP population boundary; a model will generate a confident conventional list whether or not it fits this product, so every criterion needs a sourced rationale and every unjustified one is removed.
- The estimand is the joint that couples the clinical question to the statistical method, and the model must not silently choose its intercurrent-event strategy. Section 6 states each estimand attribute explicitly, the intercurrent-event strategy is flagged for statistician sign-off, and verification confirms Section 6 and Section 9 agree on endpoint, variable, and analysis.
- Split the checking: validators for structured-field terminology conformance, named humans for narrative groundedness and design judgment. Conformance errors are cheap to catch with tooling; ungrounded safety rationale and estimand strategy are expensive to catch and dangerous to miss, so human attention goes where tooling cannot reach, all documented under ICH E6(R3).
Skill.re