AI-Assisted Literature Surveillance and ICSR Triage
Every week a pharmacovigilance writer in a global safety department opens a literature surveillance queue with more than four hundred articles in it. The database queries against PubMed and EMBASE, run against the product's active substance and a list of synonyms, returned 412 hits this week. Somewhere in those 412 articles are perhaps ten that describe an individual case that meets the criteria for an ICSR, a real patient with a real suspected reaction to the product, buried among 380-plus papers that are pharmacokinetic studies with no case-level data, off-label discussions, review articles, duplicate indexing of the same paper, and mentions of the drug class that never touch a patient. The writer's job is to find the ten, document why the other four hundred were rejected, and do it again next week. This is the highest-volume triage task in pharmacovigilance, and it is where AI is most obviously useful and most subtly dangerous, because the governing metric here is not precision. It is recall. This lesson is about why that single word changes everything about how AI should and should not be used.
What Literature Surveillance Is and Why It Is Mandatory
Pharmacovigilance literature surveillance is the systematic, scheduled screening of the published scientific literature to identify reportable safety information about a medicinal product, principally individual case safety reports that appear in case reports, case series, and sometimes clinical study publications. It is not optional and it is not a courtesy. For products marketed in the European Union, the obligation is anchored in Article 57 of Regulation 726/2004 and operationalized through the EMA's medical literature monitoring service and the marketing authorization holder's own screening duty under EU Good Pharmacovigilance Practices Module VI. The holder must screen a defined set of literature on a defined cadence, identify cases that meet ICSR criteria, and report them within the same expedited and periodic timelines that govern any other case source. A case that was published in a journal the holder was obligated to screen, and that the holder missed, is a reporting failure regardless of where the case came from.
The EudraVigilance and Article 57 overlay adds a specific structure. The EMA monitors a defined list of active substances in selected medical literature and enters the resulting ICSRs into EudraVigilance directly, which means a holder must reconcile its own screening against the EMA's monitored-substance list to avoid both gaps and duplication. For substances on the EMA monitored list, the holder's obligation is shaped by what the agency already covers; for substances not on the list, the full screening burden sits with the holder. The writer running the weekly queue is therefore not just reading abstracts; they are operating inside a regulatory architecture that defines what must be screened, what the agency screens for them, and what timeline a found case inherits. The 412 hits are the raw input to a legally mandated process whose output is the set of cases the company is obligated to report.
The structure of the task is a funnel. Four hundred-plus hits enter at the top. A search-and-deduplication step removes exact and near-duplicate records. A relevance-screening step, reading title and abstract, removes the clearly irrelevant: the PK studies, the in vitro work, the reviews with no primary case data, the off-label mentions that describe no adverse event. A full-text assessment step examines the survivors to determine whether each describes a valid ICSR with the four minimum criteria. The ten or so that survive become cases; the four hundred that do not must each carry a documented reason for rejection. AI can accelerate every stage of this funnel, which is exactly why understanding the governing metric is the difference between accelerating it safely and breaking it quietly.
Recall Is the Governing Metric, and Why That Inverts Everything
In most classification problems, you balance precision and recall and accept a trade-off. Precision is the fraction of the articles the system flagged as relevant that actually are relevant; it measures how much noise you let through. Recall is the fraction of the truly relevant articles that the system successfully flagged; it measures how many real cases you caught. For literature surveillance, these two are not equal, and the asymmetry is the whole lesson. A precision failure means a writer wastes time reading an article that turns out to be irrelevant, an annoyance measured in minutes. A recall failure means a reportable case that existed in the screened literature was never flagged, never assessed, and never reported, a regulatory failure and a patient-safety failure measured in missed signals and inspection findings.
This asymmetry inverts the usual optimization. A tool tuned for high precision, one that only flags articles it is very confident are relevant, will be pleasant to use and will quietly drop the ambiguous cases at the margin, which are precisely the cases that most need human eyes. The borderline abstract, the case report with incomplete information, the publication where the adverse event is mentioned in passing, these are where recall is won or lost, and a high-precision tool sacrifices exactly them. The correct configuration for literature surveillance is the opposite of what feels efficient: the AI should be tuned to over-include, to flag everything that might be a case, accepting a higher false-positive rate in exchange for not missing a true case. The human then bears the cost of reading the false positives, and that cost is acceptable precisely because the alternative, a missed case, is not.
This is why AI in literature surveillance must never be configured as an autonomous filter that discards articles the human never sees. The dangerous deployment is the one that silently removes the bottom 380 and presents the writer with a clean list of 32 to review, because the writer cannot recover a case the tool dropped without telling them. The safe deployment uses AI to rank and cluster, to push the likely cases to the top and group the obvious noise, while keeping every article visible and every rejection a documented human decision. The model is a triage assistant that orders the work and explains its reasoning; it is not a gatekeeper that decides what the human is allowed to see. Recall as the governing metric means the human's field of view must include everything, and the AI's job is to make that complete field of view faster to work through, not smaller.
Documenting the Rejection Rationale, Article by Article
The part of literature surveillance that auditors actually inspect is not the ten cases that were reported. It is the four hundred that were not. An inspector reconstructing the screening process asks a simple, devastating question: show me why you rejected this article. If the answer is "the tool filtered it" or "it did not seem relevant," the screening process is undocumented and therefore indefensible, because the regulator cannot distinguish a deliberate, reasoned rejection from an accident or an oversight. Audit defensibility in literature surveillance lives entirely in the per-article rejection rationale, the documented reason each non-selected article was determined not to contain a reportable case.
This is where AI provides genuine, well-shaped value, because generating a consistent, structured rejection rationale for every article is exactly the kind of high-volume, pattern-following documentation the model is good at, and it is the kind of documentation humans do inconsistently when fatigued at article 340 of 412. The model can propose a rejection category for each non-selected article, no adverse event described, no identifiable patient, off-label use without an adverse reaction, duplicate of a record already assessed, review article with no primary case data, animal or in vitro study, and a one-line justification tied to the abstract. The human reviews and confirms the categorization, and the result is a complete, consistent, auditable screening log where every one of the 412 articles has a disposition and a reason. The model turns the most tedious and most audit-critical part of the task into a structured artifact, while the human retains the decision.
The discipline that keeps this defensible is that the rejection rationale must be a human-confirmed decision, not a model-asserted one, for the same reason recall is the governing metric: a model that auto-rejects an article with a generated rationale has made a screening decision the human never reviewed, and if that article was actually a case, the documented rationale is now a fabricated justification for missing it, which is worse than no documentation. The correct pattern is that the model proposes the disposition and the rationale, the human confirms or overrides it, and the log records that a human confirmed it. A rejection log full of plausible, model-generated rationales that no human actually reviewed is not audit defense; it is automated negligence with good formatting, and an inspector who discovers one wrongly-rejected case will read every other rationale as suspect.
Where the Model Helps and Where It Characteristically Misses
The model's strengths in literature surveillance map cleanly onto the funnel. It is excellent at deduplication, recognizing that the same study indexed in PubMed and EMBASE with slightly different metadata is one record, a tedious task it does tirelessly. It is strong at the first-pass relevance sort, reliably recognizing that a pharmacokinetic modeling paper or an in vitro receptor-binding study contains no individual case, which clears the obvious noise that consumes a human's attention. It is good at extracting the candidate case elements from a relevant abstract, the patient, the reaction, the suspect product, surfacing them for the human's valid-case assessment. And it is good, as established, at generating the structured rejection rationale at volume.
The characteristic misses are as important as the strengths, and they cluster at the margin where recall is decided. The model can miss a case that is described obliquely, where the adverse event is not stated in standard terminology but implied in clinical detail, because the model pattern-matches on the language of adverse-event reporting and a case described in unusual phrasing may not trigger it. It can mis-handle a non-English publication, where the case is real but the language or the translation obscures the signal. It can fail to recognize a case embedded in a larger paper, a single patient narrative inside a review or a clinical study report where the surrounding content is not case-level. It can be misled by negation, reading "no hepatotoxicity was observed" as a hepatotoxicity signal, or the inverse, missing a real signal phrased with clinical understatement. Each of these is a recall failure, a true case the model under-flagged, and each is why the human cannot be removed from the loop. The model's job is to make the human faster and more consistent across 412 articles; it is not to be trusted to catch the oblique case, which is exactly the case that matters most.
The practical consequence is a configuration principle: tune the model to surface and rank, set its threshold to over-include rather than under-include, and treat its relevance score as a prioritization aid rather than a verdict. An article the model scored as low-relevance is read by a human at the bottom of the ranked list, not discarded. The model's confidence orders the queue; it does not close it. This is the operational expression of recall as the governing metric, and it is the single configuration decision that most determines whether an AI-assisted surveillance process is safe or quietly broken.
Building the Surveillance Workflow as a Defensible Process
A defensible AI-assisted literature surveillance workflow is built around the funnel, the recall priority, and the audit trail simultaneously, and it assigns each stage a clear human-model division. The search stage is human-owned in its design: the query strategy, the database selection, the synonym and active-substance list, and the cadence are defined by the safety function and validated, because a case the query never retrieved is a case no downstream AI can recover. The model then runs deduplication and the first-pass relevance sort, ranking all retrieved articles by likelihood of containing a reportable case and clustering the obvious noise, while keeping every article in the human's field of view.
The human works the ranked list top to bottom, with the model's proposed disposition and rationale alongside each article. For the high-ranked candidates, the human performs the full valid-case assessment and routes confirmed cases into the ICSR process under their inherited timeline. For the rejected articles, the human confirms or overrides the model's proposed rejection category and rationale, producing the complete, human-confirmed screening log. Around this, the workflow records the query run and its parameters, the model and version, the relevance threshold (which must be set to favor recall), every disposition and who confirmed it, and the reconciliation against the EMA Article 57 monitored-substance list. The output is a screening process where every retrieved article has a human-confirmed disposition, every confirmed case entered the reporting stream on time, and an inspector can trace any article to a documented decision.
The strategic insight is that literature surveillance is the pharmacovigilance task where the governing metric and the deployment pattern are most tightly coupled, and getting the metric wrong corrupts the deployment in a way that is invisible until an inspection or a missed signal exposes it. A surveillance process optimized for precision feels efficient, presents clean numbers, and quietly loses the marginal cases that the entire obligation exists to catch. A surveillance process optimized for recall feels heavier, surfaces more false positives, and catches the case that would otherwise have been missed. The writer who understands that recall, not precision, is the metric will configure the tool to over-include, keep the human's field of view complete, and document every rejection as a human decision, which is the only version of AI-assisted surveillance that survives the question an inspector will eventually ask: show me why you rejected this one.
Why the Marginal Case Is the Whole Point
It is worth ending on the case at the margin, because it is the reason this lesson treats a single metric with such weight. The four hundred obvious rejections do not test the system; any tool and any tired human clears a pharmacokinetic modeling paper. The system is tested by the one abstract that describes, in non-standard language, a patient who developed a serious reaction the product's label does not list. That case is a potential new signal, the same kind of signal that the expedited report and the periodic safety report exist to surface, and literature surveillance is one of the channels through which it enters the safety system. If a precision-tuned tool drops that abstract below the threshold the human reviews, the signal does not merely arrive late. It does not arrive at all, and no one knows it was missed, because a missed case in literature surveillance leaves no trace unless someone re-screens the same literature and finds it.
This is the deepest reason recall governs. The cost of a missed case is not symmetric with the cost of a false positive, and it is not even fully knowable, because the organization cannot count the cases it never saw. The only defense against an unknowable cost is to refuse to let the tool create it: to keep the human's field of view complete, to tune for over-inclusion, and to make every rejection a reviewed decision. AI makes the writer faster and more consistent across the 412 articles, and that is a real and valuable acceleration. What AI must never do is decide, on its own and invisibly, that the marginal case was not worth the human's attention, because that single decision, made silently at scale every week, is how a literature surveillance process fails without anyone noticing until the signal it should have caught surfaces somewhere far more costly.
Key Takeaways
- Literature surveillance is a legally mandated funnel: 400-plus weekly PubMed and EMBASE hits down to the ten that meet ICSR criteria, with every rejection documented. The EU obligation anchors in Article 57 and EU GVP Module VI, with the EMA medical-literature-monitoring service and monitored-substance list defining what the agency screens and what burden remains with the marketing authorization holder.
- Recall, not precision, is the governing metric, and that inverts the usual optimization. A precision failure wastes a few minutes reading an irrelevant article; a recall failure means a reportable case in the screened literature was never flagged, assessed, or reported. The model must be tuned to over-include, accepting false positives so it never drops the marginal case.
- AI must rank and cluster, never autonomously filter what the human cannot see. The dangerous deployment silently discards the bottom hundreds and presents a clean short list, because the writer cannot recover a case the tool dropped. The model orders the complete field of view; it is not a gatekeeper that decides what the human is allowed to review.
- Audit defensibility lives in the per-article rejection rationale, which must be human-confirmed, not model-asserted. The model proposes a structured disposition and one-line justification for each non-selected article (no adverse event, no identifiable patient, duplicate, review with no primary data), and the human confirms it. A log of unreviewed, model-generated rationales is automated negligence with good formatting.
- The marginal case is the whole point, and a missed case leaves no trace. The model characteristically misses obliquely described, non-English, embedded, or negation-phrased cases, exactly where recall is decided. Because the organization cannot count the cases it never saw, the only defense against that unknowable cost is a complete human field of view, an over-inclusive threshold, and a reviewed decision on every rejection.
Skill.re