←
AI for Pharmacy
Strategic · M8 · lesson 8 of 19 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Evaluating Pharmacy AI Vendors
📖
now learning

Evaluating Pharmacy AI Vendors

15 min

A director of pharmacy sat through the best vendor demo she had ever seen. The tool ingested a referral, pulled the patient's labs and prior therapies, drafted a clean prior-authorization (PA) justification, matched it to the payer criteria, and produced a finished package in under a minute. Every example worked. The room was sold. Then she did the one thing that separated a buyer from a spectator: she asked the salesperson to run the same workflow on a referral she handed over on the spot, a messy real one from her own queue with a smudged fax, an off-label indication, and a renal note buried in the third page. The tool stumbled. It invented a lab value that was not in the document, asserted a diagnosis the record did not support, and cited a payer criterion that had been retired eight months earlier. Nothing about the demo had been dishonest. The demo had simply been built on clean, curated inputs, and her pharmacy does not run on clean, curated inputs. That gap, between the demo and the day, is where every pharmacy AI procurement decision is actually won or lost, and this lesson is about how to evaluate a vendor so that the gap is something you measured rather than something a patient discovers for you.

The Demo Is Not the Product

The first principle of vendor evaluation in pharmacy is that the demo is a performance, not a measurement. A demo is built to sell, which means it runs on inputs the vendor chose, on the happy path the vendor rehearsed, narrated by someone whose job is to make you feel that the hard parts are already solved. None of that is necessarily deceptive. It is simply that a demo answers the question "can this tool do the task under ideal conditions?" when the only question that matters for a pharmacy is "can this tool do the task under my conditions, on my worst inputs, without producing a confident error that reaches a patient?" Those are different questions, and the distance between them is exactly the distance between an efficiency win and a patient-safety event.

So the evaluator's job is to convert the demo from a performance into a test. That means refusing to let the vendor pick the inputs. Bring your own material: the smudged fax, the off-label referral, the patient with a renal note that changes the dose, the payer criterion you know was updated last quarter, the drug interaction that is real but rare. Hand these to the tool in the room and watch what it does. A vendor who is confident in the product will welcome this; a vendor who deflects, who insists the curated demo is representative, who promises that "the production version handles those," is telling you something important about where their confidence actually lives. The single most useful sentence a pharmacy evaluator can say in a vendor meeting is: "Let us run that on one of mine."

A demo proves the tool can succeed on inputs the vendor chose. Evaluation proves what it does on the inputs you cannot choose, which are the only inputs your patients will ever generate.

Evaluate Against the Failure Modes You Already Know

By the time a pharmacy leader reaches this evaluation, the earlier levels of this program have already named the failure modes that matter, and vendor evaluation is largely the discipline of testing a tool against each one rather than against a generic notion of "accuracy." The terminology lessons taught that every AI term is a question in disguise; vendor evaluation is where those questions get asked out loud, to a person who has to answer them about a specific product you are considering buying. You are not evaluating "AI." You are evaluating whether this tool, on your data, resists the specific ways pharmacy AI hurts patients.

The central failure mode is the hallucination: a fabricated dose, an invented interaction, a contraindication that is not real, or a coverage criterion the model conjured because it sounded plausible. In a clinical setting this is not a quality-of-life annoyance, it is the failure that hospitalizes someone. So the evaluation question is never "is it accurate?" in the abstract. It is "when this tool produces a clinical claim or a payer criterion, can I see the source it came from, and does that source actually say what the tool claims?" A tool whose output you cannot trace back to a verifiable source is a tool that asks you to trust fluent prose, and fluent prose is precisely what a hallucination wears. The second failure mode is the silent miss, the false negative: the real interaction the tool failed to surface, the renal flag it did not raise. Demos rarely reveal these because a miss produces no visible error in the room; it produces a clean-looking output that happens to be incomplete. To probe for misses you have to bring cases where you know the right answer and check whether the tool found what it should have, not just whether what it produced looks right.

The third is bias and inconsistency: a tool that performs well on common medications and conditions and quietly worse on the underrepresented ones, or that behaves differently on inputs that should be treated the same. The way you find this is not by asking the vendor whether their tool is biased, because no vendor will say yes; it is by testing across the actual range of your patients and your formulary, including the cases that are rare for the model but real for you. Across all three failure modes, the unifying evaluation principle is the same one the whole program turns on: the tool's job is to support the pharmacist's judgment, never to replace it, so the only acceptable design is one where a human can verify every load-bearing claim before it has effect, and the evaluation is really a test of whether that verification is possible, fast, and built into the workflow rather than bolted on after a sale.

The Verification Question for Each Claim Type

A useful way to structure a vendor evaluation is to walk through each kind of claim the tool will produce and ask the specific verification question that claim type demands, because "can I verify it?" is too vague to act on and a clinical claim, a payer criterion, and an extracted data point each fail differently. Take them in turn.

Extracted data points. When the tool pulls a lab value, a diagnosis code, a prior therapy, or a renal function from the record, the verification question is: does the extracted value match the source document exactly, and can the tool show me where in the document it came from? A good extraction tool links each pulled value back to the precise location in the chart so a pharmacist can confirm it in seconds; a poor one presents extracted values as a tidy list divorced from their source, which turns verification into a manual hunt and quietly invites the team to skip it. Ask to see the citation-to-source feature, and test it on a document where you know a value is ambiguous or buried.

Generated clinical text. When the tool drafts a justification, a counseling explanation, or a clinical summary, the verification question is: is every clinical assertion in this generated text traceable to something real in the record, or has the model added plausible-sounding content that no source supports? Generated text is where hallucination lives, because generation is the act of producing fluent content, and fluency is indistinguishable from accuracy to a tired reader. The evaluation move is to take a generated justification and check it claim by claim against the chart, counting how many assertions are supported, how many are unsupported, and how many are subtly wrong. A vendor whose tool you can audit this way, and whose numbers hold up on your inputs, is showing you something real.

Payer and formulary criteria. When the tool matches a request to a coverage rule, the verification question is: is this the current, correct criterion from the actual payer source, or a remembered approximation that may be outdated or invented? This is where grounding matters most. A tool that retrieves the live payer criteria and shows you the source is grounded; a tool that produces criteria from the model's training memory is guessing, and an outdated or fabricated criterion produces a denial, a delay, and in the case of a time-sensitive therapy, real patient harm. Ask directly: where do the criteria come from, how current is that source, and can the tool show me the source behind each matched criterion? The honest answer to "is it grounded?" is never just "yes"; it is "grounded on this specific source, refreshed on this cadence, and here is how you see it."

PHI Handling Is Part of the Evaluation, Not a Separate Track

A pharmacy leader new to vendor evaluation sometimes treats clinical performance and data protection as two separate reviews handled by two separate teams, the clinical people checking accuracy and the compliance people checking privacy. That separation is a mistake, because a tool that performs beautifully and mishandles protected health information (PHI) is not a usable tool; it is a breach with a good demo. PHI is individually identifiable patient health information, and the moment a pharmacy AI tool touches the chart, the pharmacy's obligations under the Health Insurance Portability and Accountability Act (HIPAA) are fully in play. Evaluation has to ask, alongside every accuracy question, where the patient data goes when it enters this tool.

The questions belong in the same evaluation conversation as the clinical ones: Is there a business associate agreement (BAA), the contract that legally binds a vendor handling PHI to protect it? Where is the data processed and stored? Is patient data used to train the vendor's models, and if so, under what restrictions, because a tool that improves itself by absorbing your patients' information is doing something the pharmacy must understand and control? Who at the vendor can access the data, and how is that access logged? These are covered in depth in this chapter's due-diligence lesson, but the point here is that they are not a downstream gate the clinical evaluation passes you to; they are part of whether the tool is acceptable at all. A pharmacy evaluating a vendor is asking one integrated question with two faces: does this tool produce verifiable, safe clinical output, and does it protect the patient's information while doing so? A no to either is a no to the tool.

Evaluate the Vendor, Not Only the Tool

The tool is what you see in the demo; the vendor is who you are actually buying from, and over a multi-year relationship the vendor's character matters as much as the current feature set. A pharmacy AI tool is not a static product you purchase once; it is a model that will be updated, retrained, and changed under you, often without your involvement, which means you are entering a relationship with whoever controls those changes. Evaluating the vendor means asking how they handle the things that will inevitably happen after the sale.

How does the vendor communicate model updates, and can a change to the model alter clinical behavior without notice? A retrained model can shift its outputs in ways that affect patient care, and a vendor who pushes silent updates is a vendor who can change your tool's clinical behavior between Friday and Monday without telling you. How does the vendor handle errors and incidents: when their tool produces a wrong clinical claim that a pharmacist catches, is there a path to report it, a response, a fix, or does the report vanish? How do they describe their own product's limits, because a vendor who claims their tool does not hallucinate, never errs, or removes the need for verification is either naive or selling, and either way is telling you they do not understand the patient-safety stakes you live with. The vendor you want is the one whose people can articulate exactly where their tool can fail and what they have built to contain it, because that honesty is the strongest evidence that the people building the tool understand the same asymmetry you do: that speed is the easy win and a confident clinical error is the unforgivable one.

How to Run the Evaluation in Practice

Pulling this together into something you can actually do, a pharmacy vendor evaluation is a structured test, not a meeting where you watch and decide on impression. Before the vendor arrives, assemble an evaluation set from your own pharmacy: a folder of real, varied, deliberately hard cases where you already know the right answer, including clean ones, messy faxes, off-label referrals, renal-dose cases, rare interactions, and at least one payer criterion you know was recently updated. This set is the most valuable thing you bring, because it converts every demo into a measurement and it is reusable across every vendor you evaluate, which lets you compare them on identical ground instead of on whose salesperson was more polished.

Then run each candidate tool against the set and score it on the dimensions that matter: traceability (can you see the source behind every claim?), extraction accuracy on your worst inputs, the rate of unsupported or fabricated clinical assertions, the rate of silent misses on cases where you know the answer, the currency and grounding of payer criteria, the speed and ease of human verification in the actual workflow, and the PHI posture (BAA, data location, model-training use, access controls). Score the vendor too: update communication, incident handling, and the honesty with which they describe their tool's limits. A tool that is fast but whose claims you cannot trace, or whose misses you cannot detect, or whose data posture is unclear, fails the evaluation no matter how impressive the demo was, because in pharmacy the cost of being wrong is not measured in time. The evaluation is not bureaucracy slowing down a purchase; it is the difference between buying a tool that makes your pharmacy faster and safer and buying a confident error generator with a good sales team. The director who handed over her own messy referral was not being difficult. She was doing the only thing that turns a demo into a decision: making the tool prove itself on the reality it will actually have to survive.

Key Takeaways

  • The demo is a performance built on inputs the vendor chose; evaluation is a measurement built on the messy, real inputs your patients actually generate, and the gap between them is where procurement decisions are won or lost.
  • Refuse to let the vendor pick the inputs: bring your own evaluation set of real, hard cases (smudged faxes, off-label referrals, renal-dose cases, rare interactions, a recently updated payer criterion) where you already know the right answer, and run every candidate on identical ground.
  • Evaluate against the failure modes the program already named: hallucination (fabricated doses, interactions, or criteria), the silent miss or false negative (the real signal the tool fails to surface), and bias or inconsistency across your full range of patients and formulary.
  • Ask the verification question that fits each claim type: for extracted data, does it match the source and link back to it; for generated clinical text, is every assertion traceable to the record; for payer criteria, is it the current grounded source rather than a remembered approximation.
  • A tool that you cannot trace to a verifiable source is asking you to trust fluent prose, which is exactly what a hallucination wears; "is it grounded?" is never answered by "yes" alone but by the specific source, its refresh cadence, and how you see it behind each claim.
  • Protected health information (PHI) handling is part of the same evaluation, not a separate track: a tool that performs beautifully but lacks a business associate agreement (BAA), has an unclear data location, or trains the vendor's models on your patients is a breach with a good demo.
  • Evaluate the vendor, not only the tool: how they communicate model updates that can silently change clinical behavior, how they handle reported errors, and how honestly they describe their tool's limits, because a vendor who claims their tool never hallucinates does not understand your patient-safety stakes.
  • Score every candidate on traceability, accuracy on your worst inputs, unsupported-claim and silent-miss rates, criteria grounding, verification speed in the real workflow, and PHI posture; a fast tool whose claims you cannot trace or whose misses you cannot detect fails the evaluation regardless of how good the demo looked.