←
AI for Pharma & Life Sciences
Capable · M32 · lesson 32 of 35 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Draft a Module 2.5 Clinical Overview Section from a Phase 3 CSR
📖
now learning

Draft a Module 2.5 Clinical Overview Section from a Phase 3 CSR

15 min

A Module 2.5 lead has a Phase 3 oncology Clinical Study Report open in one window and an empty Clinical Overview template in another. The CSR runs to six hundred pages with appendices. The deadline for the integrated 2.5 draft is Friday. The instinct, in 2026, is obvious: paste the executive summary and the synopsis into the enterprise large language model, ask for a draft Module 2.5.4 efficacy section, and start from something rather than nothing. That instinct is correct. The model will produce, in under a minute, a structurally sound efficacy section that reads like a senior writer wrote it. The danger is equally obvious to anyone who has read the previous chapter: the draft will contain factual claims and table citations that look exactly as confident as the true ones, and some of them will be wrong. This lesson is the workflow that turns that instinct into a defensible artifact. You will take a Phase 3 oncology CSR, produce a draft Module 2.5.4 efficacy section with built-in TLF citation validation where every claim maps to a named table number such as Table 14.2.1.4 for progression-free survival in the intent-to-treat population, verify the result against ICH E3 structure and the Module 2.5 expectations under ICH M4E, and document the verification step in a form that survives a Good Clinical Practice audit and an Office of New Drugs Information Request. The draft is the easy part. The traceability is the lesson.

What the 2.5.4 Is and Why It Is a Derivative Document

The Module 2.5 Clinical Overview sits inside the Common Technical Document as a critical, interpretive summary of the entire clinical program, and ICH M4E governs its structure. Section 2.5.4 is the Overview of Efficacy, and it is not a place for new analysis. It is a place where the sponsor states, concisely and with cross-references, what the efficacy data show and why they support the indication. Every number in the 2.5.4 originates somewhere else: in a Clinical Study Report written to ICH E3, and underneath that, in the Tables, Listings, and Figures package that the statistical programming team produced from the locked database. The 2.5.4 is a derivative document, which means its job is fidelity, not creativity. A claim in the 2.5.4 that does not reconcile to a CSR result, and through the CSR to a TLF table, is a defect regardless of how true it happens to be, because the reviewer cannot follow it home.

This is why the 2.5.4 is simultaneously the easiest and the most dangerous thing to hand a language model. It is easy because the structure is stable and well represented in the training corpus, so the model produces clean scaffolding without effort. It is dangerous because the derivative nature of the document is invisible to the model. The model does not know that the median progression-free survival it just wrote has to match the number in Table 14.2.1.4, because it does not know Table 14.2.1.4 exists or what it contains unless you put it in the context window. It writes the efficacy section the way the corpus taught it efficacy sections go, and efficacy sections go with hazard ratios and confidence intervals and table citations. The shape is right. Whether the shape is filled with your trial's true numbers is the entire question, and it is a question the architecture does not answer on its own.

Hold this frame for the rest of the lesson: you are not asking the model to summarize a trial. You are asking it to transcribe specific results from named sources into a specific structure, and then you are going to check that every transcription is faithful. That reframing, from summarization to disciplined transcription, is what makes the difference between an accelerant and a liability.

Staging the Source Set Before You Prompt

The single most consequential decision in this workflow happens before you write a word of prompt: what goes into the context window. The model reasons only over what is loaded, and a missing source is a silence, not a flag. If you load the executive summary and the synopsis but not the efficacy tables, you have asked the model to write an efficacy section whose source of truth is outside its world, and it will fill the gap with plausible invention rather than refuse. So the staging step is the verification step's precondition. You assemble the minimum source set that makes every claim checkable inside the window.

For a Module 2.5.4 efficacy section, that minimum set is specific. You load the CSR synopsis, which gives the model the design, the population definitions, and the headline results in prose. You load the CSR efficacy section, ICH E3 Section 11, which contains the primary and secondary analyses as narrative with their own table references. Critically, you load the relevant TLF tables themselves, the primary efficacy table such as Table 14.2.1.1 for the primary endpoint analysis, the progression-free survival table such as Table 14.2.1.4 for PFS by ITT, the overall survival table, and the key secondary and subgroup tables, as text or structured data the model can read. And you load the Statistical Analysis Plan extract that defines the analysis populations and the estimand, so the model has the definitions of intent-to-treat, per-protocol, and safety set rather than guessing which one a number belongs to. When the source set is staged this way, every factual claim the model writes has a corresponding source inside the window, which means every claim is checkable. When it is not staged this way, you have built a document you cannot defend.

There is a sequencing discipline here that pays off later. Name each source as you load it, and keep a manifest: document title, version, date, and the locator scheme it uses, so that when the draft cites Table 14.2.1.4 you can confirm not only that the table exists but that it is the version under change control. Context windows in 2026 enterprise deployments range from tens of thousands to over a million tokens, and a six-hundred-page CSR with full appendices may not fit even a large window, so a window stuffed past its limit silently drops the oldest content. Loading deliberately, rather than pasting everything and hoping, is how you avoid building a 2.5.4 on a CSR whose appendix listings were truncated without anyone noticing.

The Citation-Required Prompt Pattern for the 2.5.4

The prompt that produces a defensible 2.5.4 draft does not ask for an efficacy section. It asks for an efficacy section in which every factual sentence carries a source citation, and it forbids the model from writing any factual claim it cannot cite to a loaded document. This is the Show Your Sources pattern from the previous lesson, applied to the highest-stakes section in the dossier. The instruction reads, in substance: draft Module 2.5.4 aligned to ICH M4E and ICH E3, and for every statement of an efficacy result, append the source table or CSR section in brackets, for example progression-free survival result followed by a bracketed reference to Table 14.2.1.4, intent-to-treat population. If a claim cannot be supported by a loaded source, do not write the claim; instead write a bracketed flag stating that the source was not provided. Do not invent table numbers. State the analysis population for every result. The constraint that the model must flag missing data rather than fill it is the line that converts the model's worst habit into a usable signal.

Two features of this prompt matter disproportionately. First, requiring the analysis population on every result forces the model to commit to ITT versus per-protocol versus safety set explicitly, which makes a quiet population swap visible rather than buried in interchangeable phrasing. A hazard ratio that the model attaches to the wrong analysis set is a false claim even if the number itself is real, and the only way to catch it is to make the model state the set so you can check it. Second, requiring an inline bracketed citation on every factual sentence turns the draft into a self-documenting reconciliation worksheet. You do not have to guess which sentences are claims; the model has marked them, and each one points at the source you will check it against. The draft arrives pre-decomposed into the verification tasks you are about to perform.

You will also instruct the model on what not to do. It does not write a benefit-risk conclusion; that integration lives in Section 2.5.6 and is a human judgment the model is not permitted to make. It does not characterize a result as statistically significant unless the loaded SAP and table support the specific multiplicity-adjusted comparison, because significance language attached to a non-primary or non-adjusted endpoint is a common and costly drafting error. It does not soften or strengthen an effect with editorializing adverbs that the source does not justify. The prompt, in other words, encodes the house style and the regulatory constraints that a well-built system prompt would carry, and if your enterprise deployment already carries them in a hidden system prompt, your task prompt reinforces rather than replaces them.

The TLF Citation Validation Step, Claim by Claim

Now the actual work. The draft has arrived with bracketed citations on every factual sentence, and your job is to reconcile each one against the loaded TLF. This is not rereading. Rereading checks the prose against itself, and the prose is uniformly confident whether it is right or wrong. Reconciliation checks each claim against the table it cites, which is where the truth lives. You work down the draft sentence by sentence, and for each bracketed claim you perform four checks that mirror the claim decomposition from the previous lesson.

First, the structural check: does the cited table exist in the loaded TLF package, with that exact number? A citation to Table 14.2.1.4 is only valid if Table 14.2.1.4 is in the package and is the PFS-by-ITT table the SAP says it is. A format-valid citation to a nonexistent or mismatched table is the single most dangerous defect in the document, because it survives spell-check and a hasty reviewer and propagates into Module 2.7.3 and the integrated summary. Second, the value check: does the number in the sentence match the number in the cited cell, exactly, including the confidence interval bounds and the p-value? A transcribed hazard ratio of 0.71 that the model rendered as 0.68 is wrong even though both are plausible, and only the table tells you which. Third, the population check: is the analysis set the model stated the one the table reports? A PFS result is meaningless if it is silently attributed to per-protocol when the table is ITT. Fourth, the direction check: does the effect favor the arm the sentence says it favors? Arm swaps are rare and catastrophic, and a fluent sentence reverses a benefit into a harm without any change in tone.

Record the outcome of each reconciliation in a claim-reconciliation log, one row per claim, with the claim text, the claim type, the cited source and locator, the verification status, the reconciler's initials, and the date. Any row that cannot be resolved, an uncheckable citation, a value mismatch, a population the table does not support, blocks acceptance of the draft until it is corrected or removed. The log is not bureaucratic overhead. It is the artifact that answers the question a reviewer or an auditor will eventually ask, which is not whether you used AI but whether every claim in the Clinical Overview reconciles to a real source. The log is the document that says yes, here is the evidence, claim by claim.

Verifying Against ICH E3 and ICH M4E Structure

Content fidelity is one axis of verification; structural alignment is the other. The 2.5.4 has to sit correctly inside the Module 2.5 architecture that ICH M4E defines, and it has to faithfully represent the CSR that ICH E3 structures. These are separate checks from the claim reconciliation, and the model's clean scaffolding can lull you into skipping them. It should not. The structural conventions are where the model is safest, but safe is not certain, and the consequences of a misplaced or mischaracterized section are real.

The ICH M4E check asks whether the 2.5.4 stays in its lane. The efficacy overview summarizes and interprets at the program level; it does not reproduce the study-by-study detail that belongs in the Module 2.7.3 Summary of Clinical Efficacy, and it does not reach the benefit-risk integration of 2.5.6. A draft that drifts into pooled-analysis detail or that states an integrated benefit-risk conclusion has crossed a structural boundary, and you pull it back. The ICH E3 check asks whether the efficacy claims in the 2.5.4 faithfully represent the corresponding analyses in CSR Section 11, the efficacy evaluation, including the primary analysis, the sensitivity analyses, and the handling of the estimand's intercurrent events. A 2.5.4 that states a primary result without the sensitivity analysis that the CSR reports as qualifying it has summarized selectively, which is a fidelity failure even when every cited number is correct.

This structural verification is also where you confirm that the cross-references resolve in both directions. The 2.5.4 cites TLF tables, and it cites CSR sections, and a senior reviewer will follow those references. A cross-reference that points at the right CSR section but the wrong table, or the right table but a section number that shifted when the CSR was reorganized, is a broken link that erodes confidence. You verify that the document hangs together as a navigable structure, not just that its sentences are individually true, because a reviewer reads it as a structure and judges the sponsor by whether it holds.

Documenting the Verification Step So It Survives an Audit

The draft is reconciled and structurally verified, and now you make the work defensible. The documentation has two halves, and both are required. The first half is the AI use record: which model and version produced the draft, under what system prompt, at what temperature, on what date and time, with which exact sources loaded into the window, and the final human-edited artifact that resulted. Because temperature makes output vary run to run, this record pins the specific run that produced the text you kept, not a vague statement that the company AI was used. This record is what an Office of New Drugs Information Request about AI involvement, or a Good Clinical Practice audit of your drafting process, will ask to see, and it is what the FDA-EMA transparency principle implies you should be able to produce.

The second half is the claim-reconciliation log itself, tied to the document version under change control in your regulatory document system. The log demonstrates that every factual claim in the 2.5.4 was checked against a named source and that no unresolved discrepancy survived into the accepted draft. Together, the AI use record and the reconciliation log form a complete answer to the only question that matters under 21 CFR Part 11: can the integrity and traceability of this electronic record be demonstrated? The AI use record establishes how the draft was generated; the reconciliation log establishes that its content traces to truth; the named author's signature establishes who owns the gap between them. A document with all three is defensible. A polished 2.5.4 with none of them is a liability wearing the costume of finished work.

Note what this documentation is not. It is not a confession that AI was used, and it does not weaken the submission. The disclosure of AI involvement, where a sponsor chooses or is asked to make it, is strongest precisely when it is accompanied by evidence of a disciplined verification workflow. A reviewer who sees that the sponsor reconciled every claim to source and recorded the run is reassured, not alarmed. The sponsor is not asking the agency to trust the AI; the sponsor is demonstrating a validated process in which a named human verified every claim, which is exactly what the regulation has always required of medical writing, AI or not.

What This Means for the Writer on Friday

Return to the writer with the six-hundred-page CSR and the Friday deadline. The workflow this lesson describes does not slow that writer down; it changes where the time goes. The blank-page time, the hours spent assembling a first structural draft from scratch, collapses to minutes, because the model is genuinely excellent at scaffolding and at transcribing results it can see. The time moves to staging the source set deliberately and to the claim-by-claim reconciliation, which is the work that was always the real work and that a hand-drafted 2.5.4 also required, just less visibly. The writer who internalizes this stops experiencing the model as a magic summarizer and starts using it as a fast, fallible transcriptionist who must be checked against the source on every line.

The discipline reduces to a short sequence that you can run on any CSR-to-2.5 task. Stage the source set so every claim is checkable in the window, especially the TLF tables. Prompt for citation-required output that flags missing sources rather than filling them and that states the analysis population on every result. Reconcile every bracketed claim against its cited table for existence, value, population, and direction, logging each one. Verify the structure against ICH M4E and ICH E3 so the document sits correctly and its cross-references resolve. Document the run and the reconciliation, and sign only what you have verified. Do this, and the eleven-second draft becomes a thirty-minute defensible artifact instead of a thirty-second liability. The next lesson takes the same disciplined-transcription mindset upstream, to drafting a Clinical Study Protocol aligned to the ICH M11 structured template, where the source of truth is not a locked database but a target product profile and a prior Investigator's Brochure.

Key Takeaways

  • The Module 2.5.4 is a derivative document whose job is fidelity, not analysis. Every efficacy claim must reconcile to a CSR result under ICH E3 and, beneath it, to a named TLF table such as Table 14.2.1.4 for PFS by ITT. A claim that does not trace home is a defect regardless of whether it happens to be true.
  • Staging the source set is the precondition for verification, not a preliminary. Load the synopsis, the CSR efficacy section, the actual TLF tables, and the SAP extract that defines the analysis populations, so every claim the model writes has a checkable source inside the context window. A missing source is a silence, not a flag.
  • The citation-required prompt converts the model's worst habit into a worksheet. Require a bracketed source on every factual sentence, require the analysis population on every result, and forbid the model from writing any claim it cannot cite, instructing it to flag missing data rather than fill it. The draft arrives pre-decomposed into verification tasks.
  • Reconcile every claim on four axes: existence, value, population, and direction. Confirm the cited table exists with that exact number, the value matches the cell including interval and p-value, the analysis set matches the table, and the effect favors the stated arm. Record each in a claim-reconciliation log where any unresolved row blocks acceptance.
  • Defensibility requires both the AI use record and the reconciliation log, tied to a signature. The use record pins the specific run under 21 CFR Part 11; the reconciliation log shows content traces to source; the named author owns the gap. A polished 2.5.4 with neither is a liability, not a finished document.