←
AI for Pharma & Life Sciences
Capable · M29 · lesson 29 of 35 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI-Assisted Site Selection and Feasibility Questionnaire Synthesis
📖
now learning

AI-Assisted Site Selection and Feasibility Questionnaire Synthesis

15 min

A study start-up lead has eighty completed feasibility questionnaires open across four browser tabs, a Phase 3 protocol with a 480-patient target and a non-negotiable last-patient-in date eleven months out, and a steering-committee meeting on Thursday that wants a recommended list of forty-five sites with a defensible enrollment forecast. Each questionnaire is twenty-two questions of free text and check-boxes: prior experience in the indication, principal-investigator workload, competing trials, patient-population access, site infrastructure, ethics-committee turnaround. Eighty sites, twenty-two questions, four countries, three languages of supporting documents. She pastes the lot into the enterprise large language model with one instruction: cluster these sites by likely enrollment and give me a ranked shortlist with projected FPCV and LPLV. Forty seconds later a clean ranking appears, top to bottom, with confident enrollment-rate estimates and a forecast that hits the timeline. It looks like a week of work compressed into a coffee break. It is also the single most bias-prone artifact in clinical operations, because a site-selection model trained on historical performance systematically rewards the sites that have always been chosen and quietly penalizes the newer, smaller, and more diverse sites that the trial needs in order to enroll a representative population. This lesson is about doing the synthesis well: what the model can legitimately compress, where the named SSU KPIs come from, and why the historical-bias caution is not a footnote but the center of the work.

What Feasibility Synthesis Actually Is, Mechanically

Site selection begins with a feasibility questionnaire sent to candidate sites, and the synthesis problem is that eighty completed questionnaires are eighty inconsistent documents. One investigator writes "we see roughly forty eligible patients a year"; another checks a box labeled "high"; a third attaches a referral-network diagram and a screening log from a competing trial. The large language model is genuinely good at the first move, which is normalization: reading eighty heterogeneous responses and projecting them onto a common structured schema so that "forty a year," "high," and the screening log all become comparable fields. This is extraction and classification work, the modalities the model does most reliably, because the task is to map varied surface language onto a fixed set of categories rather than to invent a fact. The synthesis lead who understands this draws the line correctly: the model harmonizes the responses into a table, and the human owns whether the table reflects reality.

The second move is clustering, grouping the eighty normalized sites into bands by enrollment likelihood, deviation history, and operational capability. Here the model is proposing structure, not measuring it, and the distinction matters. When the model places a site in the "high enrollment likelihood" cluster, it is producing the most statistically plausible label given the patterns in the questionnaire text and whatever historical data was loaded, not a verified prediction. A site can land in the top cluster because its investigator writes confident, fluent prose, while a site with a genuinely larger eligible population lands lower because its coordinator answered tersely. The fluency of the response is not the enrollment potential of the site, but the model cannot tell the two apart, and neither can a reader skimming the ranked output. The synthesis is useful precisely because it imposes structure on chaos; it is dangerous precisely because the structure looks like measurement when it is pattern.

The Three Clustering Dimensions and Where Each Comes From

The shortlist is built from three named dimensions, and each has a different source of truth that the synthesis lead must hold separately. Enrollment likelihood is forward-looking and comes from the questionnaire claims about eligible-patient volume, referral networks, competing trials, and the indication's local prevalence; it is the softest dimension because it is a prediction that sites have every incentive to inflate. Deviation history is backward-looking and comes from your own systems, the prior monitoring records, protocol-deviation logs, and audit findings for sites you have worked with before; it is the hardest dimension because it is recorded fact rather than self-report, but it is also the dimension most contaminated by historical bias, a point the caution section returns to. Operational capability is structural and comes from site infrastructure answers: staffing, equipment, certifications, ethics-committee turnaround, and whether the site can handle the protocol's specific procedural burden.

The synthesis lead's job is to keep these three dimensions from collapsing into a single seductive score. The model, left to its own devices, will happily produce one composite ranking that blends a site's self-reported enrollment optimism with its real deviation record and its infrastructure, and that blend hides exactly the trade-offs the steering committee needs to see. A site with outstanding infrastructure and a clean deviation history but a thin eligible population is a different bet from a site with a huge population and a messy deviation record, and a single number erases the difference. The disciplined output keeps the three dimensions visible as separate columns and lets the human weigh them against the protocol's actual constraints, because a protocol that needs rare-disease patients weights enrollment access very differently from a protocol drowning in eligible patients that needs flawless data quality.

Forecasting FPCV and LPLV Under Named SSU KPIs

The forecast the steering committee wants is anchored to two named milestones. FPCV, the first-patient-consented-visit, is the moment the first enrolled patient signs informed consent and is screened, and it is governed by the study start-up cycle: site identification, confidentiality agreements, feasibility, site qualification visit, contract and budget execution, ethics-committee and regulatory approval, site initiation visit, and green-light to enroll. LPLV, the last-patient-last-visit, marks the end of the treatment and follow-up period for the final enrolled patient and effectively closes the enrollment-and-conduct window. Between FPCV and LPLV sits the enrollment curve, and the forecast is the model's attempt to project that curve from per-site enrollment-rate assumptions, site activation timing, and screen-failure rate. The synthesis lead must know that the model is multiplying assumptions, not observing outcomes, and that the most fragile input is the per-site enrollment rate, the very number sites inflate.

The SSU KPIs that discipline this forecast are concrete and worth naming because a forecast that ignores them is fiction dressed as a plan. Cycle time from site identification to site activation, often tracked in weeks, sets how soon a site can contribute at all. Screen-failure rate, the proportion of consented patients who fail eligibility, deflates the gap between consented and randomized patients and is routinely underestimated in optimistic forecasts. Enrollment rate per site per month is the engine of the curve. Site activation rate, the fraction of selected sites that ever enroll a single patient, captures the painful reality that some selected sites never start, and the non-enrolling-site rate is one of the most expensive failures in study start-up. A forecast that hits the timeline only by assuming every site activates on schedule, enrolls at its self-reported rate, and screen-fails at an optimistic floor is exactly the forecast the model will produce by default, because each optimistic assumption is individually plausible and the model has no incentive to compound them pessimistically. The human applies the haircut.

The Historical-Bias Problem at the Center of Site Selection

This is the caution the lesson is built around, and it is not a compliance garnish; it is the dominant failure mode of AI-assisted site selection. A model that ranks sites by likely performance learns from historical data, and historical data records which sites were chosen before, which sites enrolled well before, and which sites had clean monitoring records before. Those historical patterns encode decades of selection bias: large academic centers in major cities were chosen repeatedly, accumulated experience and infrastructure and clean records because they were chosen, and therefore look like the safest bets in any model trained on that history. The newer community site, the site in an under-represented region, the site serving a more diverse patient population, has little or no track record, and absence of a track record reads to the model as risk. The model does not see "untested promise"; it sees "no evidence of performance" and ranks it down.

The consequence runs directly into a second regulatory and scientific imperative: trial populations are expected to reflect the patients who will ultimately use the drug, and under-enrollment of diverse and representative populations is a documented, named problem that FDA diversity-action-plan expectations and broad scientific consensus treat as a defect, not a preference. A site-selection model optimized purely for historical enrollment speed will recommend the same homogeneous roster of high-volume centers that produced the representation gap in the first place, and it will do so with confident numbers that make the recommendation feel objective. The synthesis lead who accepts the ranked output uncritically is not being neutral; they are laundering historical bias through a model that makes it look like data-driven rigor. The defensible practice is to treat the model's ranking as a hypothesis about operational capacity, to inspect explicitly whether the recommended roster meets the protocol's diversity and representativeness goals, and to deliberately weigh promising sites that the historical pattern penalizes. The model can surface candidates; it cannot be allowed to define the eligible pool, because its definition is the past.

The Questionnaire Claims the Model Cannot Verify

A feasibility questionnaire is a self-report, and self-reports are optimistic by construction, because a site that wants the trial has every reason to present its best case. When an investigator writes "we have approximately three hundred patients with this indication in active follow-up," the model has no way to verify the number; it transcribes the claim and may even smooth it into a confident enrollment projection downstream. The synthesis lead has to read the questionnaire the way an experienced study-start-up manager reads it, as a set of claims each carrying a credibility weight, not as data. The eligible-population claim, the competing-trials disclosure, the staffing claim, and the past-performance claim each deserve a different level of skepticism, and the most dangerous are the ones the site has an incentive to shade: eligible-patient volume up, competing trials down, prior screen-failure rates down.

The model amplifies this in a specific way. Because it normalizes eighty heterogeneous responses into clean comparable fields, it strips the texture that an experienced reader uses to discount a claim. A vague, hedged, over-promising answer and a precise, evidenced, conservative answer can both become the field "eligible patients per month: 25" after normalization, and the hedging that should have triggered skepticism is gone. The defensible workflow keeps the original response reachable behind every normalized field, so a recommendation can always be traced back to the actual words the site wrote, and a synthesis lead can ask why a site that wrote three vague sentences is sitting in the top cluster. Normalization is a convenience that must not become an erasure, and the audit trail for a site-selection decision should let a reviewer reconstruct which claim, in which questionnaire, drove which ranking.

The Vendor Landscape and the Human Handoff

Site-selection and feasibility AI is a real and maturing market, and naming the tools matters for vendor-neutral fluency. Platforms such as Medidata's clinical-trial intelligence layer, Saama, Lokavant, Reify Health, and IQVIA's clinical AI offerings ingest historical trial performance, claims and patient-flow data, and site-performance records to propose site lists and enrollment forecasts. These systems are more sophisticated than a raw LLM pasted with questionnaires, because they are grounded in structured performance databases rather than free-text self-reports, but they inherit the same historical-bias problem in a more entrenched form precisely because their training data is the industry's collective selection history. A vendor forecast that looks authoritative because it is built on millions of patient-flow records is still a forecast built on who was chosen before. The synthesis lead's skepticism does not relax because the tool is specialized; if anything, it sharpens, because the bias is harder to see.

The handoff that makes this defensible is the same shape as every other AI-assisted clinical-operations workflow: the model proposes, the human disposes, and the decision is documented under the named author's accountability. The synthesis lead takes the model's normalized table, clusters, and forecast as a first draft, applies the SSU-KPI haircut to the enrollment assumptions, explicitly tests the recommended roster against the protocol's diversity and representativeness goals, restores the questionnaire texture behind any borderline recommendation, and produces a shortlist that they can defend to the steering committee and, later, to a sponsor auditor asking how sites were chosen. The forty-second ranking is the start of that work. The defensible artifact is the shortlist plus the reasoning that explains why the model's ranking was accepted, adjusted, or overridden, site by site, because in study start-up the cost of a wrong site is measured in months of enrollment delay and, increasingly, in a representation gap that a regulator will name.

What This Means for the SSU Lead on Monday

The practical posture is to use the model aggressively for the work it does well and to refuse it the work it does badly. Let it normalize eighty questionnaires into a comparable schema, because that compression is real and reliable, and it turns a week of manual transcription into a reviewable table in minutes. Let it propose clusters and surface candidate sites you might not have considered, because widening the candidate pool is exactly where AI can counteract rather than entrench historical bias, if you point it that way. Do not let it deliver a single composite score, an unexamined enrollment forecast, or a roster you forward to the steering committee without applying the KPI haircut and the diversity test. The model has no stake in whether the trial enrolls a representative population on time; you do, and the named author of the site-selection recommendation owns the gap between a plausible ranking and a sound one. The forecast that hits the timeline on the first pass is the one to distrust most, because the model produced it by stacking optimistic assumptions, and the steering committee will hold you, not the model, to the date it printed.

Key Takeaways

  • The model normalizes and clusters feasibility responses well, but clustering is proposed structure, not measurement. A fluent questionnaire can lift a weak site into the top band while a tersely answered strong site sinks, because the model reads response fluency, not enrollment potential, and the ranked output hides the difference behind confident order.
  • Keep enrollment likelihood, deviation history, and operational capability as three separate dimensions, never one composite score. Each has a different source of truth (self-report, recorded fact, infrastructure) and a different credibility, and a single blended number erases exactly the trade-offs the steering committee needs to weigh against the protocol's actual constraints.
  • FPCV and LPLV forecasts are stacked assumptions, and the SSU KPIs are the discipline. Cycle time, screen-failure rate, per-site enrollment rate, and site-activation rate each deflate an optimistic curve; the default model forecast hits the timeline only by compounding individually plausible best cases, so the human applies the haircut.
  • Historical bias is the dominant failure mode, not a footnote. A model trained on who was chosen before rewards established high-volume centers and penalizes newer, smaller, and more diverse sites for having no track record, reproducing the representation gap that FDA diversity-action-plan expectations treat as a defect; the synthesis lead must test the roster against representativeness, not just speed.
  • Preserve questionnaire texture and document the reasoning, because self-reports are optimistic by construction. Normalization strips the hedging an experienced reader uses to discount inflated eligible-population and competing-trial claims; keep the original response reachable behind every field, and make the defensible artifact the shortlist plus the site-by-site rationale for accepting, adjusting, or overriding the model's ranking.