←
AI for Pharma & Life Sciences
Capable · M25 · lesson 25 of 35 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI-Assisted PSUR / PBRER (ICH E2C(R2)) and DSUR (ICH E2F) Section Drafting
📖
now learning

AI-Assisted PSUR / PBRER (ICH E2C(R2)) and DSUR (ICH E2F) Section Drafting

15 min

It is week nine of the eleven-week clock before a PSUR is due to PMDA, and the safety writer has just exported the cumulative case tally and the interval line listings from the global safety database into the enterprise model, with one instruction: draft Section 16, Signal and Risk Evaluation, for this reporting interval, aligned to ICH E2C(R2). Forty seconds later a clean, sectioned draft appears, with a tidy paragraph for each of the three signals carried forward from the last cycle, a fourth paragraph for a new signal raised this interval, and a closing benefit-risk sentence that reads like the Qualified Person for Pharmacovigilance wrote it. That closing sentence is the problem. The model has quietly performed the one act that ICH E2C(R2) reserves for human judgment, the integration of benefit against risk, and it has done so in a register so confident that a tired reviewer will sign it. This lesson traces how a PSUR or PBRER under ICH E2C(R2), and the parallel DSUR under ICH E2F, can be drafted with AI section by section, where the database feed becomes the floor of truth, why Section 16 and Section 17 are different risk objects, and why the benefit-risk integration stays human all the way to the QPPV signature.

What the PSUR Actually Is, and Why Section 16 and 17 Carry the Weight

The Periodic Safety Update Report, named the Periodic Benefit-Risk Evaluation Report in the ICH E2C(R2) framework, is the marketed-product document in which a marketing authorization holder demonstrates, at defined intervals, that the accumulated safety experience still supports a favorable benefit-risk balance. It is not a case dump and it is not a literature review. It is an argued document with a fixed table of contents, running from the worldwide marketing authorization status in Section 1 through the estimated exposure in Section 5, the presentation of individual case histories and study data in the middle sections, and culminating in the two sections this lesson is built around. Section 16, the Signal and Risk Evaluation, is where every signal and every important identified and potential risk is summarized, assessed, and given a disposition. Section 17, the Benefit Evaluation, is where the holder restates the established efficacy and effectiveness in the approved indications.

The architecture matters because the two sections fail differently under AI. Section 16 is dense, structured, and repetitive across reporting cycles, which is exactly the shape a model drafts well, and exactly the shape in which a model can carry a stale disposition forward unchanged because the previous interval's text said so. Section 17 is shorter and more stable, drawn from the established efficacy in the labeling and the pivotal evidence, which makes it tempting to let the model assemble and tempting to forget that effectiveness claims must still trace to a source the holder can defend. Neither section is the benefit-risk conclusion itself. That conclusion lives in Section 18, the Integrated Benefit-Risk Analysis for Authorized Indications, and the discipline of this entire lesson is that Sections 16 and 17 are AI-assistable inputs while Section 18 is a human integration the model prepares for but never performs.

The DSUR, the Development Safety Update Report governed by ICH E2F, is the investigational-stage sibling. It covers products under clinical development rather than marketed products, its reporting clock runs from the development international birth date, and its purpose is to give regulators an annual account of whether the evolving safety profile of a drug still under study remains consistent with the protections built into the program. The DSUR has its own structured contents, and it reaches its own evaluation of the important risks identified during the reporting period together with a concluding assessment of whether the information obtained is consistent with the previous knowledge of the product's safety. The mapping this lesson uses is deliberate and approximate: the signal-and-risk evaluation work you do for PSUR Section 16 has a direct analogue in the DSUR's risk-evaluation section, and the benefit-oriented framing of PSUR Section 17 has its analogue in the DSUR's overall safety assessment, with the integration again reserved for human judgment.

The Database Feed Is the Floor of Truth, and Everything Above It Is Suspect

The single most important architectural decision in AI-assisted periodic-report drafting is where the numbers come from. A PSUR Section 16 is full of counts: the cumulative number of cases for a given preferred term, the interval count, the number that were serious, the number that were fatal, the number assessed as related. Every one of these is a query result against the global safety database, whether that database is Oracle Argus Safety, ArisGlobal LifeSphere, or another safety system of record, and every one of these must enter the draft as a transcribed value from a named, dated, reproducible query, not as a number the model produced because the paragraph needed one. The model, asked to draft a signal paragraph without the query output loaded, will generate a plausible cumulative count, and a plausible count of forty-one fatal cases is indistinguishable on the page from the true count of nineteen.

This is the literature-surveillance lesson's recall problem turned inside out. There the danger was the model silently dropping a case from a queue. Here the danger is the model silently inventing a tally that was never in a query. The control is the same in spirit and stricter in execution: the line listings, the summary tabulations, and the cumulative and interval counts are pulled from the safety database first, frozen as the data lock point output, and loaded into the model's context as the only permitted source of any number. The writer's rule is absolute. If a number in the draft cannot be traced to a cell in the loaded data-lock-point export, it is wrong until proven right, and the QPPV who signs the report owns that traceability under the same audit-trail expectations that govern any electronic record in a regulated safety system.

There is a second-order trap in the feed itself. A data lock point defines the cut-off date for data included in the report, and a model handed an export will happily summarize whatever rows are present without ever noticing that the export was run three days before the lock point and is missing the late-arriving serious cases. The model does not experience the gap between the export and the true data lock point as a problem, because it experiences nothing; it summarizes the rows it was given. So the human owns not only the transcription of each number but the provenance of the export itself, the confirmation that the query ran at the correct data lock point, against the correct product family, with the correct seriousness and listedness flags, before a single sentence is drafted on top of it.

Drafting Section 16 Signal by Signal, Not Report by Report

The productive unit of AI assistance in Section 16 is the individual signal or risk, not the whole section. A signal in this context is a piece of information that suggests a new potentially causal association, or a new aspect of a known association, that warrants further evaluation, and the holder maintains a running inventory of these across reporting cycles. The effective workflow loads, for one signal at a time, the relevant line listing, the cumulative and interval tabulations for the associated preferred terms, the previous interval's evaluation text for that signal, and any relevant aggregate analysis, and asks the model to draft that signal's paragraph in the E2C(R2) structure: the source and date the signal was raised, the cumulative and interval case counts, the clinical characterization, the assessment, and the disposition, which is whether the signal is refuted, remains under evaluation, or is closed as a new identified or potential risk.

Working signal by signal does two things. It keeps the loaded context tight enough that every number in the paragraph traces to a loaded source, and it forces the writer to confront the disposition for each signal as a distinct decision rather than letting the model carry a block of last interval's text forward wholesale. The carry-forward problem is the characteristic Section 16 failure. Because periodic reports are serial, the model trained on prior intervals, or simply handed the previous text, will reproduce a disposition of remains under evaluation for a signal that the new interval's data actually closes, or worse, reproduce a closed disposition for a signal that the new fatal cases should reopen. The disposition is a medical judgment about the current totality of evidence, and the model's strong prior toward textual consistency across intervals is precisely the force the human must resist.

The clinical characterization paragraph is where the model is most useful and most quietly wrong. Asked to characterize, say, a hepatotoxicity signal, the model writes a fluent summary of the typical presentation, the time to onset, the dechallenge and rechallenge pattern, and the confounders, and much of this is correct boilerplate for hepatotoxicity in general. Whether it is correct for your cases, in your interval, at your counts, is a separate question the model cannot answer from general knowledge, and the writer must reconcile every clinical claim against the actual line listing rather than against the model's well-trained sense of what a hepatotoxicity signal usually looks like. The shape of the characterization is safe; the specifics are the work.

Section 17 and the Trap of Letting the Model Argue Benefit

Section 17, the Benefit Evaluation, looks easy and is not. Its job is to summarize the important baseline efficacy and effectiveness information that supports the benefit side of the eventual integration, drawn from the approved indications, the pivotal evidence, and any new effectiveness data accrued in the interval. Because this material is stable across cycles and grounded in the labeling, it is the section writers are most tempted to let the model assemble with light supervision, and it is the section where an ungrounded model will most fluently overstate. A model asked to summarize the benefit of a marketed oncology product will reach for the strongest efficacy framing in its training corpus, which may be a hazard ratio from a pooled analysis or a press release rather than the specific, labeled, approved-indication claim the PSUR is entitled to make.

The discipline for Section 17 is that every effectiveness statement must trace to a defensible source: the approved labeling, the pivotal study results as filed, or a specifically cited new analysis from the interval, and nothing else. The model is permitted to organize and phrase the established benefit; it is not permitted to generate a benefit claim, to upgrade a labeled indication into a broader one, or to import efficacy language from outside the holder's defensible evidence base. This is the benefit-side mirror of the Section 16 numbers rule, and it matters because Section 17 is the input the integration will weigh against the risks in Section 18. An inflated Section 17 produces a falsely favorable integration, and the QPPV is signing a benefit-risk conclusion built on a benefit claim the holder cannot defend at inspection.

Why the Benefit-Risk Integration in Section 18 Stays Human

Everything in this lesson converges on a single reserved act. The integrated benefit-risk analysis weighs the established and new benefits in the approved indications against the identified and potential risks and the missing information, and reaches a conclusion about whether the balance remains favorable and whether any action is warranted. This is not a summarization task and it is not a structured-drafting task. It is a judgment under uncertainty, made by a named, accountable human who can be questioned by a regulator, and it is exactly the class of reasoning that frontier models perform fluently and unreliably. The model can write a paragraph that reads like a benefit-risk integration. It cannot make one, because making one requires weighing incommensurable quantities, a survival benefit against a rare fatal toxicity, under a value judgment about acceptable risk that belongs to the holder's pharmacovigilance system and ultimately to public health, not to a next-token sampler.

The reason this matters operationally, and not just philosophically, is that the model's integration will be plausible and will sometimes be wrong in the most dangerous direction. Handed a strong Section 17 and a Section 16 with a new fatal signal still under evaluation, the model tends to resolve the tension toward the favorable conclusion, because favorable benefit-risk conclusions dominate its training corpus of approved-product periodic reports. It produces a confident sentence that the benefit-risk balance remains positive, at the precise moment the human evaluation should be flagging that the new signal changes the calculus and may warrant a labeling change or a risk-minimization measure. The model's statistical pull toward the common conclusion is structurally biased against the rare moment when the conclusion should change, which is the only moment a periodic report truly matters.

So the workflow draws a hard line. The model drafts Sections 16 and 17 from loaded, traceable sources, signal by signal and claim by claim, and it may even produce a structured pre-read that lays out the benefits and risks side by side as inputs. It does not draft Section 18's conclusion, and where a deployment allows it to attempt one, that draft is treated as a prompt for human reasoning to react to, never as text to retain. The QPPV, or the DSUR's responsible safety physician, reads the assembled evidence, performs the integration as a human act, and signs. The named signatory owns the gap between a benefit-risk conclusion that looks correct and one that is correct, and in periodic safety reporting that gap is measured in patient harm.

Documenting the Run So It Survives a GVP Inspection

A PSUR or DSUR is a regulatory submission with a named author and a named approver, and an AI-assisted one inherits an obligation the eleven-second draft makes easy to forget: the work has to be reconstructable. Under EU Good Pharmacovigilance Practices and the equivalent expectations elsewhere, an inspector can ask how a given section was produced, and the answer cannot be that the company used its AI tool. The defensible record captures, per section drafted, the model and version, the system prompt, the exact data-lock-point export and query parameters loaded, the prompt, the temperature, the timestamp, and the final human-verified text, all linked to the controlled document version in the safety document management system. This is the same audit-trail discipline that governs any electronic record in a validated safety system, applied to the fact that part of the draft originated from a model.

The reason this is not bureaucratic theater is the carry-forward and the inflated-benefit failures described above. If a future inspection or a future signal review finds that a signal disposition was wrong, the question becomes whether the error was a human judgment made on correct data or an unreviewed model carry-forward on a stale export. Only the run record distinguishes those, and only one of them is defensible. The holder that can show the loaded export, the query parameters, the model output, and the human reviewer's confirmation or override for each signal disposition has a defensible process. The holder that can show only a finished PDF and a vague statement of AI assistance has a finding waiting to happen. The discipline is to make the periodic report's AI involvement as traceable as its numbers, because at inspection they will be questioned together.

Key Takeaways

  • The PSUR/PBRER under ICH E2C(R2) and the DSUR under ICH E2F are argued benefit-risk documents, not case dumps, and AI assists their structured sections, not their conclusions. Section 16 (Signal and Risk Evaluation) and Section 17 (Benefit Evaluation) are AI-draftable inputs; the Section 18 integrated benefit-risk analysis, and its DSUR analogue, is a reserved human act.
  • The global safety database export at the correct data lock point is the only permitted source of any number in the draft. If a count cannot be traced to a cell in the loaded data-lock-point export, it is wrong until proven right, and the human owns the provenance of the export as well as the transcription of each value.
  • Draft Section 16 signal by signal, not report by report, and resist the carry-forward. The characteristic failure is the model reproducing last interval's disposition unchanged when the new data should close, reopen, or reclassify the signal; the disposition is a medical judgment about the current totality of evidence.
  • Section 17 is where an ungrounded model overstates benefit most fluently. Every effectiveness statement must trace to the approved labeling, the pivotal evidence as filed, or a specifically cited interval analysis; an inflated benefit input produces a falsely favorable integration the QPPV cannot defend.
  • The benefit-risk integration stays human because the model is statistically biased toward the common favorable conclusion. It resolves the tension between a strong benefit and a new fatal signal toward positive precisely when the human should be flagging a change; document the run per section so an inspector can distinguish a human judgment on correct data from an unreviewed model carry-forward.