←
CAP Certification
Aware · M38 · lesson 38 of 53 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Quality Decision Framework

10 min

Priya Raman runs content operations for Meridian Credit, a regional credit union. Her four-person team uses an AI assistant to draft member emails, help-center articles, and internal policy summaries. Last quarter she watched the same argument play out again and again. One writer shipped AI drafts almost untouched. Another rewrote every draft from scratch. A third burned an afternoon nudging the AI toward a paragraph she could have written herself in ten minutes. Same tool, three completely different instincts. What Priya's team lacked was not access to AI. It was a shared, defensible way to decide what to do with a draft once they had one.

The Three Decisions

Once you have evaluated a piece of AI output, there are only three things you can do with it, and naming them plainly is the first step toward choosing between them consistently. Accept means the output meets your standards, so you use it as it stands and move on. Refine means the output is close, so you give the AI feedback and iterate until it clears the bar. Restart means the output is not salvageable, so you begin again with a different prompt, a different system, or a different approach entirely. Everything that follows in this chapter exists to make that three-way choice reproducible rather than instinctive, so that two people looking at the same draft reach the same conclusion for the same stated reasons.

Risk Calibration

The single most common mistake Priya sees is treating "good enough" as a fixed bar. It is not. The bar moves with the stakes. A typo in an internal brainstorm costs nothing. The same typo in a rate-change notice mailed to 40,000 members can trigger complaints, corrections, and a compliance review. Risk calibration means setting your acceptance standard to match the consequences of being wrong before you look at the output, so the stakes drive the decision instead of your mood or your deadline. A brainstorm document, a hiring decision, a legal contract and a medical recommendation are not the same kind of work, and they should not share an acceptance standard.

Ask three questions to place any piece of work on a risk scale. Who sees it, and how many of them? Can a mistake be quietly corrected after the fact, or does fixing it require a formal retraction? Does an error create legal, financial, or safety exposure? The more people, the harder the correction, and the greater the exposure, the higher the risk tier, and the higher your acceptance bar should climb. Low-risk work covers creative pieces, brainstorms and internal drafts, where good enough genuinely is good enough. Medium-risk work covers customer-facing content, internal documentation and competitive analysis, where accuracy needs spot-checking and bias needs checking but perfection is not required. High-risk work covers legal documents, medical recommendations, hiring decisions and financial advice, where near-perfect alignment is the minimum, every quality dimension gets checked thoroughly, and a human expert reviews the final decision rather than the machine informing it unsupervised.

Priya turned this into a standing table her whole team uses. It is the first artifact of this chapter: a shared map from risk tier to the minimum quality she will accept and the review each tier requires.

Risk tierExample at MeridianMinimum alignment to acceptReview before it ships
LowInternal meeting notes, first-draft brainstorm, throwaway summary70%Author spot-check
MediumMember newsletter, help-center article, internal policy summary85%Full read by one reviewer
HighRegulatory disclosure, rate-change notice, anything with a legal or financial commitment95% or higherTwo reviewers plus subject-matter sign-off

The exact percentages are less important than the discipline of setting them in advance. When a writer knows before drafting that a help-center article must reach 85 percent and get a full read, the accept, refine or restart decision stops being a matter of taste and becomes a matter of measurement. It also stops being a negotiation with whoever is most senior in the room, because the standard was agreed when nobody was under deadline pressure.

Criteria for Accepting

Acceptance is a judgment across four dimensions, and it helps to run them in the same order every time. Accuracy asks how correct the output needs to be for this level of risk: critical facts must be right, while minor details can be imperfect if the overall direction holds. Completeness asks whether it covers the necessary ground, or whether using it as it stands would leave gaps that cause problems downstream. Relevance asks whether it actually addresses your need, staying focused rather than drifting, so that a reader would recognize it as an answer to the question they asked. Appropriateness asks whether the tone, format and style suit the use case, which reduces in practice to whether you would be comfortable putting this in front of your intended audience.

Those four combine into a fifth question, alignment with your standards, and that is where the risk tier does its work. For low-risk work, something in the region of 80 percent alignment is often perfectly acceptable and further polishing is waste. For high-risk work you want 95 percent or better before you even consider accepting. The same draft can therefore be a clear accept in one context and a clear refine in another, which is not inconsistency; it is the framework doing exactly what it was designed to do.

Criteria for Refining

Refine when the output is close but carries fixable gaps, roughly the band between 80 and 95 percent alignment where targeted feedback can do the work. The tell is that the bones are right and the problems are specific: correct structure but a wrong detail, accurate facts but a weak tone, a small omission, or an answer that is on-topic but shallow. If you can name exactly what is wrong in a sentence or two, refining is usually faster than restarting. Three further conditions should hold: the problem is fixable without starting over, of the kind that "make it more formal" or "add more examples" would address; the time cost of refining is genuinely lower than the cost of a fresh attempt; and you have real reason to believe iteration will improve it rather than spin your wheels.

The mirror image matters just as much. Do not refine when the output is fundamentally misaligned with what you need, somewhere below 70 percent, because feedback cannot rescue a draft aimed at the wrong target. Do not refine when you have already iterated four or five times without major improvement, since that pattern usually means the prompt, not the output, is the problem. Do not refine when the effort exceeds the value of a slightly better result, and do not refine when the deadline simply does not allow for iteration. Recognizing these earlier saves more time than any prompting technique.

Refining works best as a small number of targeted instructions rather than vague dissatisfaction. "Add the 60-day filing window and the provisional-credit timeline" gets a better result than "make this more complete." Give the AI the specific missing piece, and re-score the same way you scored the first draft so the decision stays measurable.

Worked example. Priya's team drafts a help-center article, "How to dispute a card transaction." They score it with ACRE, rating each dimension from 1 to 5, then average the four scores and read the average as a percentage on a 5-point scale, so a 4.0 average equals 80 percent. Here is the scorecard for the first draft. This scorecard is the second reusable artifact of the chapter; any writer can copy it for any piece of work.

DimensionScore (1-5)Note
Accuracy4Steps are correct but it omits the 60-day filing window
Completeness3Missing the provisional-credit timeline members care about
Relevance5Squarely answers the question a disputing member would ask
Appropriateness4Tone is fine, one sentence is too jargon-heavy

The average is 4.0, or 80 percent. A help-center article is Medium risk, so the bar is 85 percent. The draft misses that bar, but nowhere near by the margin that would signal a restart, and both faults have names. Decision: refine, with two specified fixes rather than a rewrite. The second pass adds the 60-day window and the provisional-credit timeline and softens the jargon. Re-scored, it reaches a 4.5 average, or 90 percent, clearing the 85 percent bar. Now the decision is accept, and Priya can point to the numbers if anyone asks why.

Criteria for Restarting

Restarting feels like a defeat and usually is not. It is the correct call when the output is fundamentally not useful and sits far below your standards; when the AI misunderstood what you were asking in a way that feedback will not repair; when the whole approach being taken is wrong, so no amount of correction will bend it back; or when you already have a clear idea of a better prompt or a better tool. The practical signal is that you cannot name the fix in a sentence. If your feedback would amount to describing the task again from the beginning, you are not refining, you are rewriting the prompt through an intermediary, and it is faster and cleaner to start over deliberately with what the failed attempt taught you.

Decision Tree Framework

Put together, the framework runs in five steps, and it becomes quick once it is habitual. Begin with a quick scan: does the output seem directionally correct? If not, you are probably restarting and can stop there. If it is, move to a full evaluation, applying ACRE and scoring accuracy, completeness, relevance and appropriateness individually rather than forming a single overall impression. Then calculate the alignment percentage by averaging the four dimensions, so you have a number rather than a feeling. Compare that number to the minimum for the risk level, which for Priya's team is 70 percent for low risk, 85 percent for medium and 95 percent for high. Then decide.

  • If alignment is above your minimum, accept, or refine if you have identified minor gaps worth closing.
  • If alignment is 10 to 20 percent below your minimum, refine.
  • If alignment is more than 20 percent below your minimum, restart.

The step people skip is the third one. Scoring the dimensions separately is what stops a single glaring flaw from dragging your assessment of an otherwise strong draft, and what stops a fluent, confident tone from carrying a draft that is quietly incomplete.

Cost-Benefit Analysis for Refinement

Sometimes the decision is not about quality at all. It is about effort, and four questions frame the trade-off. What does refining cost, in iterations and in time, and is the output actually converging or stuck? What is the benefit, meaning how much better will this realistically get and is that improvement worth the cost? What does accepting the imperfect version cost, in risk and in what could go wrong downstream? And what does starting over cost, once you count the time to write a better prompt and evaluate a fresh result?

Then make the trade-off consciously rather than by default. Sometimes accepting 85 percent quality is better than spending three hours to reach 90 percent, particularly when the extra five points are invisible to the reader. Sometimes it is not, because those five points are the difference between a notice that is accurate and one that is not. The framework does not make that call for you. It makes sure you know you are making it.

Context-Specific Guidance

Risk tier tells you how high to set the bar. Context tells you which way to lean when a draft sits near the line. Different kinds of work carry a sensible default posture, and naming that posture up front saves your team from re-litigating it every time a draft lands.

ContextDefault postureWhere the line falls
Customer-facing regulated noticeRestart-biased, heavy refineErrors are public and correctable only through a formal, visible retraction
Other customer-facing communicationRefine-biasedAccept when the message is clear and appropriate; refine if the tone is off; restart if it could damage the relationship or the brand
Internal draft or brainstormAccept-biasedCheap to fix downstream, and speed is the point
Creative and marketing copyRefine-biasedAccept when it is engaging and on-brand even if imperfect, because readers value voice over polish; restart if the whole direction is wrong
Technical documentationRefine-biasedAccept when it is accurate and complete enough for its audience; refine for clarity; restart if the technical details are wrong
Business analysis and strategyRefine-biasedAccept when the recommendations are sound and well-reasoned; refine if key considerations are missing; restart if the underlying assumptions are wrong
Code, config, or calculationsTest, do not eyeballAccept when it runs, reads clearly and handles edge cases; refine style issues; restart on bugs or architectural problems

Priya encoded these defaults into her team's style guide so a new writer inherits the collective judgment on day one instead of learning it through a painful mistake. When a member-facing rate notice comes through, everyone already knows it is restart-biased and high risk, and the AI draft is treated as a starting sketch rather than a shippable product. When it is an internal recap, the same team accepts an 80 percent draft and moves on. The posture is set by the work, not by who happens to be holding the keyboard that day.

Documentation for Decision

For decisions that matter, write down the reasoning at the time you make it. Record what you evaluated, meaning which quality dimensions you actually checked; what you found, meaning the strengths and the weaknesses; what you decided, accept, refine or restart; and why you decided it. A few short lines are enough, and they can live alongside the draft rather than in a separate system nobody opens.

That record does several jobs at once. It helps you make consistent decisions over time, because you can see what you accepted last month and hold yourself to the same line. It creates accountability, since a decision with a stated reason is one you can defend. It helps colleagues understand your standards without having to guess at them. And it leaves a record if questions come later, which for regulated or customer-facing work is the difference between explaining a decision and reconstructing it from memory.

Final Thoughts: Evaluation as a Practice

Critical evaluation of AI output is a skill, and like all skills it improves with practice. The frameworks in this lesson are tools for thinking systematically, not rituals to perform. Over time you will develop intuition and stop working through every step explicitly, which is the point. Even experienced evaluators benefit from returning to the systematic version occasionally, because intuition drifts and blind spots accumulate quietly. Running the full framework on a piece of work every so often is how you find out whether your instincts still match your standards.

The goal is not perfection. The goal is responsible use. Make decisions consciously, know why you accepted something or rejected it, build a track record of good calls, and hold systems accountable when they produce poor results. That is how you become trustworthy with AI, and it is what Priya's team gained: not better output on any single day, but the ability to explain, to each other and to anyone who asks, exactly why a draft shipped.

Anti-Patterns

Teams that adopt this framework and still argue about drafts are usually falling into one of these, and each has a specific correction inside the lesson.

  • One fixed bar for every piece of work. Treating "good enough" as a constant is the most common failure, and it produces both outcomes you do not want: rate-change notices held to the standard of a brainstorm, and internal recaps polished for hours. The bar belongs to the risk tier, not to the person.
  • Setting the standard after reading the output. A bar chosen once you have seen the draft will quietly settle wherever the draft already is, particularly under deadline. Naming the tier and the minimum before drafting is what makes the decision defensible afterwards.
  • Scoring by overall impression. A single holistic verdict lets one glaring flaw sink a strong draft, and lets a fluent, confident tone carry a draft that is quietly incomplete. Scoring accuracy, completeness, relevance and appropriateness separately is what surfaces the difference.
  • Refining when you cannot name the fix. If your feedback would amount to describing the task again from the beginning, you are rewriting the prompt through an intermediary. That is a restart, and doing it deliberately is faster than discovering it four passes later.
  • Iterating past the point of return. Sunk cost is persuasive. Four or five passes without real improvement, an effort cost above the value of the gain, or a deadline that cannot absorb another round all mean the honest answer is accept or restart.
  • Letting seniority decide instead of the standard. Where no bar was agreed in advance, the decision defaults to whoever is most senior or most insistent in the room, and the team learns that the framework is decorative.
  • Deciding without leaving a record. Undocumented decisions cannot be checked for consistency, cannot be explained later, and guarantee that the same argument recurs on the next borderline draft.

Practice Prompts

These are team exercises as much as individual ones, since the value of the framework is that two people reach the same conclusion for the same stated reasons.

  • Build your own tier table. List the three kinds of output your team produces most often, place each on the risk scale using the three questions about audience, correctability and exposure, and write the minimum alignment and the required review next to each.
  • Score something you already shipped. Take a recent piece of AI-assisted work, score it with ACRE after the fact, and check whether it would have cleared the bar you have just written. If it would not have, that is worth knowing before the next one.
  • Run a two-reviewer calibration. Have two colleagues score the same draft independently on all four dimensions and compare. Every dimension where you diverged is a place where your shared vocabulary is not yet shared.
  • Convert a vague rejection into named fixes. Find a draft someone sent back with a general complaint and rewrite the feedback as specific missing pieces, in the manner of naming a filing window and a timeline rather than asking for more completeness.
  • Time a refine against a restart. On one task, refine to the bar and record the time. On a comparable task, restart with a better prompt and record that. The comparison turns the cost-benefit questions from a guess into a number you own.
  • Write one decision record. For a single piece of work this week, note what you evaluated, what you found, what you decided and why. Keep it beside the draft, and see whether it changes the next conversation about that piece.

Reflection

Think about the last AI draft you sent back. Could you have named the dimension it failed on, or did the feedback amount to a feeling? If it was a feeling, the standard existed only in your head, which means nobody else on your team could have applied it and you could not have defended it if challenged. Now consider the inverse case: work you accepted because it read well and arrived on time. What risk tier did it actually belong to, and would you have accepted it had you set the bar before you read it rather than after? Then ask where your own instinct sits by default. Some people accept too readily and discover the problem downstream; others refine long past the point where the improvement is worth the hours. Knowing which you are is what tells you which part of this framework you personally need most.

Glossary

  • Accept, refine, restart: The three possible outcomes of evaluating an AI output, being use it as it stands, improve it with targeted feedback, or begin again with a different prompt, system or approach.
  • Risk calibration: Setting the acceptance standard to match the consequences of being wrong, decided before the output is read so that the stakes drive the decision rather than mood or deadline.
  • Risk tier: The band a piece of work falls into, determined by how many people see it, how hard a mistake is to correct, and whether an error creates legal, financial or safety exposure.
  • ACRE: The four evaluation dimensions applied to every draft: accuracy, completeness, relevance and appropriateness.
  • Alignment percentage: The average of the four dimension scores read as a percentage, giving a number that can be compared against the minimum for the risk tier.
  • Scorecard: The record of dimension scores and the note explaining each, which turns a judgment into something a colleague can check.
  • Default posture: The lean a category of work carries before scoring, such as restart-biased for regulated notices or accept-biased for internal drafts.
  • Decision record: A short written account of what was evaluated, what was found, what was decided and why, kept for consistency, accountability and later questions.

This lesson is the decision that follows evaluation, so it depends on the lessons that teach the evaluation itself. Critical Evaluation Framework establishes the habits of assessment this chapter turns into a decision, and Detecting Hallucinations & Errors and Identifying Bias cover two specific failures that the accuracy and appropriateness dimensions are meant to catch. Where the standard needs to hold across a whole team rather than in one head, Building Quality Rubrics for AI Outputs extends the scorecard here into a fuller instrument with behavior-anchored levels.

Closing

Priya's three writers were not disagreeing about quality. They were applying three private standards that had never been written down, which is why the argument recurred every week and never resolved. What the framework gave them was not better judgment but shared judgment, plus the ability to say out loud why a draft shipped. Start with the part that costs least: write down the risk tiers for the work your team actually produces, and agree the minimum for each while nobody is under deadline. Everything else in this chapter, the scoring, the refine band, the postures and the records, only works once that table exists.

Key Takeaways

  • Every evaluation ends in one of three decisions. Accept, refine, or restart. Naming the choice explicitly is what turns a vague sense of dissatisfaction into an action.
  • Calibrate the bar to the risk before you look at the output. Internal brainstorms and regulated notices should never share an acceptance standard, and high-risk work needs expert human review of the final decision.
  • Score the four ACRE dimensions separately. Accuracy, completeness, relevance and appropriateness, averaged into an alignment percentage, give you a number you can compare against a stated minimum.
  • Refine when you can name the fix in a sentence. Close-but-flawed drafts respond to targeted instructions; drafts you would have to describe from scratch are restarts wearing a disguise.
  • Know when to stop refining. Four or five passes without real improvement, an effort cost above the value of the gain, or a deadline that cannot absorb iteration all mean the answer is accept or restart.
  • Weigh effort as well as quality. Sometimes accepting 85 percent beats spending three hours to reach 90 percent, and sometimes it does not. Make that trade consciously.
  • Let context set the default posture. Regulated notices lean toward restart, internal drafts lean toward accept, and code gets tested rather than read.
  • Document decisions that matter. A short record of what you checked, what you found, what you decided and why creates consistency, accountability and a defensible trail.

Frequently Asked Questions

Do I have to score every piece of work? No, and a team that tries will abandon the framework within a month. Low-risk work gets an author spot-check, which is the whole point of tiering. Reserve the full ACRE scoring for the drafts where the decision is genuinely contested or the stakes are high enough that you may have to explain yourself later.

Are the percentages real measurements? They are a discipline rather than a measurement. Averaging four scores on a five-point scale gives you a number, and the number matters mainly because it was compared against a bar set in advance. The exact figures are less important than agreeing them before drafting, which is what converts a matter of taste into a matter of comparison.

What if two reviewers score the same draft differently? Treat the divergence as information about the criteria rather than about the reviewers. A dimension that two careful people read differently is a dimension whose definition is ambiguous in your context, and the fix is to agree what evidence a given score requires before you score the next draft.

How do I know a draft is close enough to refine rather than restart? The practical test is whether you can name what is wrong in a sentence or two. If you can, the bones are right and targeted instructions will usually fix it faster than beginning again. If naming the problem means describing the task from scratch, you are restarting whether you admit it or not.

What if the risk tier is genuinely unclear? Work through the three questions in order: who sees it and how many, whether a mistake can be quietly corrected or needs a formal retraction, and whether an error creates legal, financial or safety exposure. If it still sits between two tiers after that, treat it as the higher one, because the cost of over-reviewing an internal document is much smaller than the cost of under-reviewing a public one.

Does this replace human review for high-stakes work? No. High-risk work is defined partly by the fact that a human expert reviews the final decision, with the framework structuring that review rather than substituting for it. For legal, medical, hiring and financial material, the score tells you whether a draft is worth putting in front of a reviewer, not whether it can go out.