AI-Assisted Source Data Verification (SDV) Sampling and Prioritization
A central monitor opens the risk dashboard on a 600-patient Phase 3 cardiovascular trial and faces a question that used to have a brute-force answer: of the roughly forty-two thousand data points captured per visit across sixty sites, which ones does a human being need to verify against the source? The old answer was all of them, or close to it: 100 percent source data verification, a clinical research associate sitting beside a paper chart, reading every field of every case report form back to the medical record. That answer is gone. Under ICH E6(R3), the expectation is risk-based, which means the trial concentrates verification effort where an error would matter most to participant safety and data reliability, and accepts that low-risk fields do not all need a human eye. Into this gap walks the large language model and the central-monitoring signal layer, and the temptation is to let the system decide what gets verified. This lesson draws a hard line that the rest of the lesson defends: AI assists the prioritization of SDV, deciding which fields, which subjects, and which sites rise to the top of the queue. It does not perform the verification. The verification, the act of confirming that a recorded value matches the source, remains a human act with a human signature, and conflating the two is the error this lesson exists to prevent.
What Risk-Based SDV Replaced, and Why
For decades the default was 100 percent SDV, and it was expensive in a way that did not buy proportionate quality. A monitor verifying every field spends as much time confirming a correctly transcribed date of birth as confirming a primary-endpoint measurement, and the evidence accumulated over years that the vast majority of that effort caught trivial, low-impact discrepancies while the errors that actually threatened a trial, the eligibility violations, the unreported serious adverse events, the primary-endpoint inconsistencies, were not reliably caught by sheer volume of verification. ICH E6, first in the R2 addendum and now reinforced in R3, reframed the activity around risk: identify the data and processes critical to participant safety and the reliability of results, then design monitoring, including SDV, to concentrate on those. The R3 revision, with its compliance dates landing across 2026, pushes this further by treating quality as something built into the trial's design rather than inspected in at the end, and by expecting sponsors to define quality tolerance limits and to monitor centrally before sending anyone to a site.
The practical consequence is that SDV is now a sampling-and-targeting problem, not a coverage problem. The monitor must decide which fields are critical enough to verify at all, which subjects within a site carry elevated risk, and which sites warrant an on-site visit versus a remote or reduced-verification approach. Each of those decisions can be informed by data: the protocol's risk assessment names the critical variables, the central-monitoring layer surfaces statistical anomalies, and the site's history flags elevated baseline risk. This is genuinely a problem where pattern detection over large data helps, because a human cannot hold forty-two thousand data points per visit across sixty sites in their head and notice that one site's lab values cluster too tightly to be real. That is exactly the work the signal layer is built for, and exactly why the prioritization, not the verification, is the AI-amenable half.
The Three Targeting Questions: Which Fields, Which Subjects, Which Sites
Risk-based SDV decomposes into three nested targeting decisions, and the AI assists each differently. The first is which fields. A trial's data is not uniformly important: the primary and key secondary endpoints, eligibility criteria, informed-consent dates, serious-adverse-event reporting, and protocol-mandated safety assessments are critical, while many administrative and low-impact fields are not. The risk assessment that the protocol team builds, often expressed as a list of critical data elements, defines this tier, and AI can help by mapping the data dictionary against the risk assessment and proposing which fields fall into the high-verification tier, but the criticality judgment is a clinical and regulatory decision that the human owns. A model that demotes an eligibility field because it statistically looks stable is making a safety decision it has no authority to make.
The second question is which subjects. Within the critical fields, not every subject carries equal risk: a subject with a reported serious adverse event, an eligibility waiver, a large number of queries, or values that the central-monitoring layer flags as anomalous deserves closer verification than a subject whose record is clean and internally consistent. This is where the signal layer earns its keep, because it can score subjects by anomaly and concentrate the monitor's attention. The third question is which sites: a site with a deviation history, slow query resolution, staff turnover, or a statistical profile that diverges from its peers, such as implausibly low variability or digit-preference patterns suggesting fabricated data, rises in priority for an on-site visit. The signal-to-visit pattern is the spine of modern central monitoring: the system detects the signal, and the signal prioritizes where a human goes and what they verify when they get there.
The Line AI Must Not Cross: Prioritization Is Not Verification
This is the discipline the lesson is built around. Verification is the act of confirming that the value recorded in the electronic case report form matches the value in the source, the medical record, the lab report, the device printout, the original document where the observation was first captured. It is an evidentiary act, and it produces an attributable record under ALCOA-plus and 21 CFR Part 11 that says a named person compared the data to the source on a date and confirmed or queried it. An LLM cannot do this, not because it is not smart enough, but because it has no access to the source. The source is a paper chart in a clinic, a signed consent form in a binder, a printout in a regulatory file; the model sees only what is already in the electronic data, which means asking it to verify is asking it to confirm a value against itself, which is not verification at all. A model that reports a field as verified has, at best, confirmed internal consistency, and internal consistency is exactly what a transcription error or a fabrication preserves.
So the prioritization output is a queue, and the verification is a human act performed against that queue. The signal layer says this subject's potassium values at this site are statistically improbable; the central monitor escalates; the clinical research associate goes to the site, pulls the source lab reports, and confirms whether the recorded values match what the analyzer actually produced. The AI raised the question; the human answered it against the source. If anyone treats the model's anomaly score as a verification result, the trial has substituted a consistency check for an evidentiary one, and the gap will surface at an inspection where an investigator asks to see the source confirmation behind a field the system marked clean. The named monitor owns the verification, and the audit trail must show a human comparing data to source, not a model scoring data against patterns.
The Central-Monitoring Handoff: Signal Then CRA Visit Prioritization
The operational pattern that makes risk-based SDV work is a handoff between a central layer and a field role, and the named vendors in this space are worth knowing for vendor-neutral fluency. Platforms such as Saama, Medidata's Acorn-branded central-monitoring and risk-based-quality-management capabilities, and Lokavant ingest the trial's accumulating data and produce signals: key risk indicators trending out of bounds, quality tolerance limit excursions, site profiles that diverge from the study mean, anomalous data patterns at the subject and site level. These signals are the input to prioritization, not the conclusion. The central monitor reviews the signals, applies judgment about which are explicable and which are concerning, and converts the concerning ones into actions: a query, a request for clarification, a targeted remote review, or an on-site monitoring visit with a specific verification scope.
The clinical research associate then receives a prioritized visit, not a blanket instruction to verify everything. The visit might say: at this site, for these three flagged subjects, verify the primary-endpoint measurements, the eligibility criteria, and the serious-adverse-event source documents, because the signal layer flagged endpoint variability and the site has a deviation history. The CRA performs the actual SDV against the source and records the result. This is the handoff that ICH E6(R3) Section on monitoring contemplates: centralized monitoring detects, on-site monitoring confirms, and the two are coordinated so that scarce on-site effort lands where the risk is. The failure mode is treating the signal as the finding. A quality tolerance limit excursion is a prompt to investigate, not a confirmed quality problem, and a site that diverges statistically may have a benign explanation, a sicker population, a different local standard of care, that only a human inquiry will surface. The signal prioritizes the human; it does not replace the human.
Where the Prioritization Itself Can Go Wrong
Even confined to its proper role, AI prioritization has failure modes the monitor must guard against, and they mirror the bias and hallucination problems seen elsewhere. The first is that a model trained on historical monitoring data inherits historical attention: sites that were monitored heavily before generated more findings, which makes them look higher-risk, while sites that were under-monitored look clean precisely because no one looked. A naive risk model can therefore concentrate verification where it was always concentrated and leave genuinely risky but historically ignored sites under-verified, the same selection-bias dynamic that haunts site selection. The monitor has to ask whether a site looks low-risk because it is low-risk or because it has never been examined closely.
The second failure mode is over-trust in the anomaly score. A statistical flag is a hypothesis, and a confident-looking risk score can lull a monitor into treating flagged subjects as the only subjects worth attention, missing a quiet, consistent fabrication that produces no statistical anomaly because it was designed to look normal. The most dangerous data-integrity problems are sometimes the ones that do not trip a signal, and a prioritization system optimized to surface anomalies can create a blind spot for non-anomalous fraud. The third is scope creep in the other direction: a system that flags too much trains monitors to dismiss signals, the alert-fatigue dynamic, so that a real excursion gets waved away in a sea of false positives. The defensible posture treats the prioritization as one input to a human monitoring plan, calibrated against the protocol's risk assessment and reviewed by a monitor who retains the authority to verify something the system did not flag and to ask why the system flagged something that turns out benign.
SDV Versus SDR, and Where Prioritization Sits in the Mix
It is worth being precise that source data verification is only one tool in the monitoring kit, and AI prioritization touches several of them. Source data verification, SDV, is the field-to-source comparison this lesson centers on: does the recorded value match the source document. Source data review, SDR, is a broader on-site or remote examination of whether the source data and trial conduct are adequate, consistent, and compliant, looking at processes and documentation quality rather than verifying individual values one at a time. Central monitoring is the off-site statistical and review activity that uses the accumulating database to detect signals across sites without anyone traveling. Risk-based SDV under ICH E6(R3) blends these, using central monitoring to decide how much SDV and SDR each site needs, and the AI prioritization layer feeds that allocation decision rather than performing any of the three.
Holding these apart matters because the AI is useful at different intensities for each. For SDV, it targets which specific fields and subjects to verify, but a human must still perform the comparison against the source. For SDR, it can surface documentation gaps and consistency problems for a human to examine, but it cannot judge whether trial conduct was adequate. For central monitoring, it is closest to a primary engine, because statistical signal detection across a large database is exactly pattern work over data the model can actually see, with no external source required to raise a hypothesis. The monitor who understands which activity a given AI output belongs to will not make the category error of treating a central-monitoring anomaly score as if it were an SDV result, which is the most common way the line between prioritization and verification gets blurred in practice.
What the Monitor Documents, and Why It Survives Inspection
The artifact that makes AI-assisted SDV prioritization defensible is the monitoring plan and the record of how prioritization decisions were made and acted on. A risk-based monitoring plan under ICH E6(R3) names the critical data and processes, defines the quality tolerance limits, specifies the central-monitoring approach and the triggers that escalate to on-site verification, and describes how SDV is targeted rather than universal. When AI assists, the plan and its supporting records should make clear that the system produced signals and prioritization, that a named central monitor reviewed and dispositioned those signals, and that the actual verification was performed by a named CRA against identified source documents. An inspector reading this trail can see that a human made the risk judgment, that the AI informed but did not decide, and that every field marked verified was confirmed against a source by a person.
The record that fails inspection is the one that cannot distinguish a model's consistency score from a human's source confirmation, or that shows verification effort tracking the model's output without evidence that a human ever reviewed whether the prioritization was sound. If a site was reduced to minimal SDV because the model scored it low-risk, the plan must show the human rationale for accepting that, and ideally some confirmatory verification that the low-risk designation held. The standard the lesson teaches is that the prioritization is auditable as a decision, with the AI as a named, documented input and the human as the named decision-maker and verifier. The signal layer accelerates the monitor's attention; it does not absorb the monitor's accountability, and the trial that forgets this distinction is the trial that discovers, at a pre-approval inspection, that no one can point to the source behind a critical field the system marked clean.
What This Means for the Monitor on Monday
Use the signal layer to do what no human can: scan forty-two thousand data points per visit across sixty sites and surface the handful of subjects, sites, and fields where risk concentrates. Let it rank the queue, propose which sites need on-site visits, and flag the anomalies that deserve a human look. Then do the part that is irreducibly yours: review the signals with judgment, decide what to verify, send the CRA with a targeted scope, and confirm critical values against the actual source. Never let a model's anomaly score stand in for a source confirmation, never let historical attention define current risk without asking whether a clean site was simply unexamined, and never let the volume of signals train you to stop investigating them. The model decides where to look. You decide what is true, and you sign the record that says so. Risk-based SDV done well is not less rigorous than 100 percent verification; it is more rigorous, because it spends the scarce, irreplaceable resource of human verification where an error would actually hurt a participant or break the data.
Key Takeaways
- AI assists the prioritization of SDV, never the verification itself. Verification confirms a recorded value against a source the model cannot see; a model can at best confirm internal consistency, which is exactly what a transcription error or fabrication preserves, so a field the system scores clean still needs a human source confirmation under ALCOA-plus and 21 CFR Part 11.
- Risk-based SDV under ICH E6(R3) is a three-part targeting problem: which fields, which subjects, which sites. Critical data elements set the field tier (a human-owned safety judgment), the signal layer scores subjects by anomaly, and site history plus statistical divergence prioritizes on-site visits; the AI assists each, but the criticality call stays with the monitor.
- The central-monitoring handoff is signal then targeted CRA visit, not signal as finding. Saama, Medidata Acorn, and Lokavant surface KRI trends and quality-tolerance-limit excursions; the central monitor dispositions them and sends the CRA a scoped verification, because a statistical excursion is a prompt to investigate, not a confirmed quality problem.
- Prioritization has its own failure modes: historical attention bias, over-trust in anomaly scores, and alert fatigue. Under-monitored sites look clean because no one looked, quiet non-anomalous fabrication trips no signal, and too many flags train monitors to dismiss them; the monitor must retain authority to verify something unflagged and to question something flagged.
- The defensible artifact is an auditable monitoring plan where AI is a named input and the human is the named decision-maker and verifier. The trail must distinguish a model's consistency score from a human's source confirmation, document the rationale for any reduced-SDV designation, and let an inspector see that a person, not a model, confirmed every critical field against its source.
Skill.re