←
AI for Pharmacy
Strategic · M11 · lesson 11 of 19 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
PoC Design with a Safety Gate
📖
now learning

PoC Design with a Safety Gate

15 min

A specialty pharmacy ran a proof-of-concept (PoC), a small trial of a new AI tool before committing to it, and called it a success. The tool had cut prior-authorization (PA) assembly time from roughly twenty-five minutes to about five, the staff loved it, and the leadership was ready to roll it out across every site. Then a clinical pharmacist who had been quietly keeping her own notes asked to present at the wrap-up meeting. She had verified, by hand, every PA the tool produced during the pilot, and she had found that in about one of every twelve packages the tool had either fabricated a clinical criterion, asserted a diagnosis the chart did not support, or matched the request to an outdated payer rule. None of those errors had caused a denial yet, because she had caught them. But the pilot's success metric had been turnaround time, and turnaround time had nothing to say about the eight percent of packages that would have been wrong if no one had checked. The PoC had proven the tool was fast. It had not proven the tool was safe, and the two are not the same measurement. This lesson is about designing a proof-of-concept so that it answers the question that actually matters before you scale: not "is it faster?" but "is it safe enough to trust, and how do we know?"

What a PoC Is Actually For

A proof-of-concept exists to retire risk before you commit. It is the controlled, reversible, small-scale trial where you find out whether a tool does what you need it to do, on your data and in your workflow, while the cost of being wrong is still contained to a pilot rather than spread across your whole operation. The trap that the specialty pharmacy fell into is the most common one in pharmacy AI: designing the PoC to prove the easy thing, speed, because speed is visible, immediate, and exciting, while leaving the hard thing, safety, unmeasured. A PoC that only measures speed is not a smaller version of the real deployment; it is a different experiment that happens to look like one, and it will hand you a confident green light on a question you never actually asked.

The reason this matters so much in pharmacy is the patient-safety asymmetry that runs through this entire program. Speed is the easy win, and a wrong renal dose, a missed interaction, or a hallucinated coverage criterion is not an efficiency miss but a patient-safety event. A PoC designed around speed optimizes for the easy win and stays silent on the catastrophic risk, which means a successful speed-only PoC can actively mislead you into scaling a tool whose error rate you never measured. The whole purpose of a properly designed PoC is to refuse that trap by building the safety question into the experiment from the start, so that the tool has to prove not only that it is fast but that it is verifiably accurate enough to trust, and that the verification can keep up at scale.

A proof-of-concept that only measures speed proves the easy thing and hides the dangerous one. In pharmacy, the PoC has to prove the tool is safe, not just fast, or it has proven nothing worth scaling.

The Safety Gate: The Heart of the Design

The defining feature of a pharmacy PoC is the safety gate: an explicit, measured, pre-committed standard that the tool must meet on the safety dimensions before it is allowed to scale, no matter how good the efficiency numbers look. The gate is what turns a pilot from a demonstration into a decision. Without it, a PoC drifts toward whatever metric is easiest to celebrate, and in pharmacy the easiest metric to celebrate is exactly the one that does not protect patients. The gate forces the question to be answered on purpose.

Designing the gate means deciding, before the pilot begins, three things. First, what are the safety failure modes this tool could produce, named specifically for this use case: for a PA tool, that is fabricated criteria, unsupported clinical assertions, outdated payer rules, and silent misses of a relevant fact. Second, how will you measure each one, which in practice means human verification: during the pilot, a qualified person checks the tool's output against the source of truth and records what they find, because the only way to know the tool's error rate is to verify its output against the chart and the real criteria. Third, what rate of each failure mode is acceptable to scale, decided in advance so that the standard is not negotiated downward in the glow of a fast turnaround number. The specialty pharmacy's pilot failed not because the tool produced errors, every tool will, but because there was no gate: no pre-committed standard, so the eight percent error rate the pharmacist found had no threshold to fail against, and the speed number carried the decision unopposed.

It is worth being concrete about what the gate measures. It is not "did the tool work?" but "of the clinical claims and payer criteria the tool produced, what fraction were fully supported by the source, what fraction were unsupported or fabricated, what fraction were subtly wrong, and what relevant facts did it silently miss?" Those numbers, measured by real human verification on real cases, are the output of a safety gate. A tool that produces them at a rate the pharmacy decided in advance it could accept, with verification that fits the workflow, passes. A tool that is fast but cannot clear the gate does not scale, and the discipline of the design is that the gate decides, not the enthusiasm in the room.

A natural question is how to set the acceptable thresholds, and there is no universal number, because the right threshold depends on the consequence of the specific failure and on what the human verification step can reliably catch. For a use case where a human verifies every output before it has effect, the gate is asking a slightly different question than for a use case where the tool acts more autonomously: in the verified workflow, the threshold is about whether the error rate is low enough that a human reviewer, under real conditions, will actually catch the errors rather than be lulled by a stream of mostly-correct output into rubber-stamping the occasional wrong one. A tool that is wrong one time in fifty, where every error is caught, may be acceptable; a tool that is wrong one time in five will exhaust the verifier and produce automation bias, where the reviewer stops genuinely checking because checking almost never finds anything until the day it would have mattered. The gate, in other words, has to account for the human who sits behind it, because the safety of the whole system is the tool's error rate combined with the verifier's realistic ability to catch what the tool gets wrong. Setting the threshold is a clinical and operational judgment the pharmacy makes deliberately, in advance, and records, so that the standard is a decision rather than an accident of whoever was in the room when the speed number landed.

Verify Before You Scale, Not After

The sequencing here is the entire point, and it is easy to get backward. The seductive failure is to scale on the strength of the efficiency win and treat verification as something you will add later, once the tool is deployed and the time savings are already booked. That ordering is exactly wrong, because once a tool is scaled it is woven into the workflow, the staff depend on its speed, and the organizational momentum runs against pulling it back. Verification discovered to be inadequate after scaling is a problem you now have to solve while patients are already exposed to the tool's errors. The PoC exists precisely so that verification happens before the tool is load-bearing, while the trial is still small and reversible and the only cases at risk are the pilot cases a human is already checking.

This reflects the deeper truth the program returns to repeatedly: the job did not get smaller when AI arrived, it changed shape. The work shifted from producing the draft to verifying the draft against the source of truth, and a PoC that does not test whether that verification is possible, fast, and reliable has not tested the actual job. So the pilot has to run the verification in the same conditions the real workflow will, with the same people, the same time pressure, and the same volume scaled down, because verification that works when one careful pharmacist checks ten cases over a relaxed week may collapse when the same pharmacist has to check a hundred a day under normal pressure. The PoC measures not just whether the tool can be verified in principle but whether the verification holds up under the conditions that will actually exist, because a verification step that is real on paper and skipped in practice is automation bias waiting to produce the error nobody caught.

Design the Pilot to Surface Errors, Not Hide Them

A good PoC is designed to make the tool's errors visible, which sounds obvious until you notice how many pilots are quietly designed to do the opposite. A pilot run only on clean, straightforward cases will produce a low error rate and a happy result that tells you nothing about the messy cases that are where errors actually happen. A pilot where the people verifying are the same people championing the tool will tend to find fewer problems, because enthusiasm is a poor auditor. A pilot whose success is defined before the safety data is in will interpret ambiguous results charitably. Each of these is a way of designing a pilot to confirm what you hoped rather than to find what is true.

The countermeasures are concrete. Run the pilot on a representative range of real cases, deliberately including the hard ones: the messy faxes, the off-label referrals, the renal-dose adjustments, the rare interactions, the recently updated payer criteria, because the error rate on hard cases is the error rate that matters and the easy cases will flatter the tool. Assign verification to someone with the clinical competence to catch a subtle error and enough independence not to be invested in the tool succeeding, because the verification is only as good as the verifier's willingness to report a problem. Define the safety gate's thresholds before the pilot starts and write them down, so that a tool which comes in just over the acceptable fabrication rate has to fail rather than get rounded down because everyone already likes it. And track the silent misses specifically, the false negatives, by using cases where you know the correct answer and checking whether the tool found what it should have, because a tool's misses produce no visible error and a pilot that only reacts to visible errors will systematically undercount the most dangerous failure mode.

Measure the Right Things, Including the Uncomfortable Ones

A pharmacy PoC should produce a small dashboard of numbers, and the discipline is to insist that the safety numbers sit next to the efficiency numbers and carry equal weight in the decision. The efficiency side is the easy part and the part everyone wants: turnaround time, staff time saved, throughput. These are real and they matter, because the goldmine of pharmacy AI is the collapse of PA turnaround from roughly twenty-five minutes to about five, and a PoC should confirm the tool delivers it. But the efficiency numbers are only half the dashboard, and a PoC that reports only them has measured the half that was never in doubt.

The safety side is the half that decides whether scaling is responsible: the supported-claim rate, the unsupported or fabricated-claim rate, the rate of subtly wrong assertions, the silent-miss rate, the grounding currency of payer criteria, and, crucially, the verification burden, how long verification actually takes and whether it can keep pace at full volume. That last number is where many pilots quietly fail their own safety test: a tool that is fast to produce output but whose output takes so long to verify that the net time saving is small, or that the team starts skipping verification under pressure, has not actually delivered the efficiency win, it has relocated the risk. The honest PoC measures the whole workflow including verification, not just the tool's raw speed, because the speed that matters is the speed of producing a verified, trustworthy package, not the speed of producing an unverified one. Put the safety numbers and the efficiency numbers on one page, hold them to the pre-committed gate, and the PoC becomes what it is supposed to be: a decision instrument rather than a celebration.

The PoC as the Template for Everything After

The deepest value of designing a PoC with a safety gate is that it establishes the pattern the pharmacy will use for every AI tool it ever adopts, and it produces, almost as a byproduct, exactly the evidence a governed pharmacy needs. The discipline of naming the failure modes, measuring them through human verification, setting a pre-committed standard, and letting that standard decide is not a one-time procurement step; it is the operating model of a pharmacy that uses AI competently and under governance, which is precisely what the new URAC Health Care AI Accreditation (a national accreditation with a track for AI users) expects an organization to demonstrate. A PoC run this way generates an audit trail of what the tool produced, what was verified, what was found, and why the scale decision was made, and that record is the raw material of accreditation-grade documentation.

So the safety gate is not friction added to slow down an exciting tool; it is the mechanism by which a pharmacy earns the right to scale, and the record that proves it earned that right. The specialty pharmacy in the opening did eventually do this correctly: they paused the rollout, redesigned the pilot around a real gate, set a fabrication-rate threshold, measured the tool honestly on hard cases, and worked with the vendor until the grounded version cleared the bar. The tool they eventually scaled was the same tool, but the decision was different, because it was made on evidence about safety rather than enthusiasm about speed. That is the whole lesson. A PoC is not the moment you fall in love with a fast tool; it is the moment you make the tool prove, against a standard you set in advance, that it is safe enough to trust your patients to, and then you keep the proof. The pharmacy that internalizes this stops treating each new tool as a separate adventure and starts treating evaluation as a repeatable, governed process, which is both faster over time and far safer, because the gate is already built and the only question for each new tool is whether it can clear a bar the organization already knows how to measure.

Key Takeaways

  • A proof-of-concept (PoC) exists to retire risk before you commit; the common pharmacy trap is designing it to prove speed, the easy and visible win, while leaving safety, the dangerous and hidden risk, unmeasured.
  • The patient-safety asymmetry means a speed-only PoC can actively mislead you: it gives a green light on a question you never asked, because turnaround time has nothing to say about the fraction of packages that would have been wrong if no one checked.
  • The defining feature of a pharmacy PoC is the safety gate: a pre-committed, measured standard on the safety failure modes (fabricated criteria, unsupported assertions, outdated rules, silent misses) that the tool must clear before it scales, decided before the pilot so it cannot be negotiated down in the glow of a fast number.
  • The gate is measured by real human verification against the source of truth, producing the supported-claim rate, the unsupported or fabricated-claim rate, the subtly-wrong rate, and the silent-miss rate on cases where you already know the answer.
  • Verify before you scale, not after: once a tool is load-bearing the momentum runs against pulling it back, and verification found inadequate after scaling is a problem solved while patients are already exposed to the tool's errors.
  • Design the pilot to surface errors rather than hide them: run it on hard, representative cases including the messy ones; assign verification to someone competent and independent of the tool's success; write the thresholds down in advance; and track silent misses specifically.
  • Measure the whole workflow including the verification burden, because a tool whose output is fast to produce but slow to verify, or whose verification gets skipped under pressure, has relocated the risk rather than delivered the efficiency win.
  • A PoC run with a safety gate becomes the template for every AI tool the pharmacy adopts and produces the audit trail (what was produced, verified, found, and decided) that the URAC AI user track expects, making the gate the mechanism by which a pharmacy earns and proves the right to scale.