←
AI for Recruiters
Capable · M10 · lesson 10 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Feedback Loops: How to Report AI Errors and Improve System Performance

15 min

Wesley runs talent acquisition for a regional healthcare network with eleven recruiters and roughly 90 open requisitions at any given moment. His team leaned on AI for resume extraction, interview-note summaries, and candidate outreach, and for the first quarter it felt like a clean win. Then a hiring manager flagged that a summary credited a nurse candidate with five years of ICU experience the resume never mentioned. A week later, an outreach draft addressed a candidate by the wrong name. These were not random gremlins; they were signals, and Wesley's team had no way to capture, route, or learn from them. The errors evaporated into chat threads and memory. What changed his team's reliability was not a better AI tool. It was a feedback loop: a structured way to report what the AI got wrong, decide who needs to fix it, and turn each report into a measurable improvement.

Why AI Improves Through Feedback and Not by Accident

AI does not get better on its own inside your workflow. It improves through feedback. Every time you use it you learn something about how it behaves: where it fails, what it does well, which prompts hold up under pressure and which quietly drift. That learning is worth something only if it goes somewhere. Left in your head it improves your own judgment slightly and nothing else. Fed back into your prompts, your processes, and your vendor relationships, it compounds, and the same tool that produced a hallucinated ICU credential in January stops producing that class of error by March.

This is the difference between a team that uses AI and a team that manages it. The first team absorbs errors individually, each recruiter developing private workarounds nobody else benefits from. The second team treats each error as data about a system that can be adjusted. Wesley's teaching point to his recruiters was blunt: catching an error is only half the job. The other half is recording it in a way that lets someone change the thing that produced it. This lesson builds that second half, a feedback system with four steps that create continuous improvement rather than repeated firefighting.

What Actually Counts as an AI Error

Before you can report an error, the team has to agree on what one is. Vagueness here kills feedback loops, because half the team logs trivial wording preferences while the other half stays silent about hallucinations that affect hiring decisions. Wesley's team settled on five concrete error types, each with a one-line definition any recruiter could apply in seconds.

Wrong extractions and misinterpretations. The AI pulled a field from a resume incorrectly, giving a wrong job title, wrong dates, or a certification attributed to the wrong employer, or it read the input and drew the wrong meaning from it. These are common and easy to verify against the source document, which makes them the best training ground for a new reporter.

Hallucinations. The AI invented something that does not appear in the input at all, like the five years of ICU experience in Wesley's nurse summary. Hallucinations are the most dangerous extraction error because they look authoritative and a busy reviewer may not catch them. Nothing in the output signals that the detail was manufactured rather than read.

Omissions. The AI silently dropped information that mattered: a relevant gap in employment, a disqualifying licensing detail, a stated concern from an interviewer. Omissions sit between extraction errors and hallucinations and often reveal a prompt that under-specified what to include. They are the hardest category to notice, because a page with something missing looks exactly like a page with nothing wrong.

Biased output. The output contains discriminatory language, assumptions tied to a protected characteristic, or scoring that correlates with age, gender, national origin, or other protected classes. Under EEOC guidance, an AI tool that produces adverse impact in screening exposes the employer to liability regardless of intent, so bias errors get the highest scrutiny and the fastest escalation of any category here.

Formatting failures. The output is technically correct but unusable: it ignored the requested structure, exceeded a length limit, returned prose when a table was asked for, or broke a downstream paste into the applicant tracking system. Low severity individually, but a frequent formatting failure wastes minutes on every single use, and the cumulative cost shows up in the log even when no single instance seems worth reporting.

The test Wesley gave his team was simple: if the AI output, used as-is, would mislead a hiring decision, embarrass the company in front of a candidate, or create legal exposure, it is an error worth reporting. A stylistic preference is not. That one sentence resolved most of the arguments about what belonged in the log.

The Four-Step Feedback Loop: Capture, Categorize, Analyze, Act

A working feedback system has four steps, and skipping any one of them breaks the others. Capture without categorization produces an unreadable pile. Categorization without analysis produces tidy counts nobody acts on. Analysis without action trains your team to stop reporting. Wesley ran all four on a quarterly cycle, with capture happening continuously and the other three batched.

Step 1: Capture. Every time someone finds an AI error, they record four things while the context is fresh. What went wrong, stated specifically: "hallucinated years of experience," not "bad summary." Where it happened: screening, summarization, or outreach, with the tool and the prompt or template name, because without the prompt name you cannot fix the root cause. The evidence: what the error was, what the correct output would have been, and what the impact would have been if it had shipped. And the severity, using three tiers, where critical means the error affects a decision or creates legal exposure, major means a notable issue a reviewer would have to fix before using the output, and minor means it does not affect much. The capture tool itself can be anything low friction: a form, a spreadsheet, a shared document, or a tagged thread in a single team chat channel. Wesley used a shared spreadsheet so anyone could scan the full log in one view.

Step 2: Categorize. Group the errors by type using the five categories above, plus an "other" bucket so that nothing gets discarded for failing to fit. Categorization is what turns a list of incidents into a picture. Twelve separate frustrations look like bad luck; twelve wrong extractions all against the same clinical resume template look like a defect with an address.

Step 3: Analyze. Look for patterns across the categorized log rather than at individual entries. What is the most common error type? When does it happen, meaning which roles, which tools, which prompts, which kinds of source documents? What is the root cause behind the cluster? And what is the impact, in the sense of which errors are cheap annoyances and which would have changed a hiring outcome? The pairing of frequency and impact is what sets priority, because the most common error is not always the one worth fixing first.

Step 4: Act. For each pattern, take a specific action. Improve the prompt to prevent the error. Add a constraint to your process, such as a required verification step before a summary reaches a hiring manager. Change tools if the issue is a vendor limitation nothing on your side can fix. Train the team when the pattern is one of use rather than one of output. And escalate immediately when the issue touches legal or compliance ground. Each action needs an owner, because an action item with no name attached is a wish.

Documenting the Error: Input, Expected Versus Actual, Severity

A useful error report answers the same questions a good bug report does. Wesley's team standardized on six fields, which turn the capture step into something reproducible rather than a free-text complaint.

Context. Where it happened, which tool was used, and the prompt or template name. Input. What went in. For candidate data this is where privacy discipline matters: log a reference to the source document or a redacted excerpt, never paste a full resume with personal identifiers into a widely shared error tracker. Under GDPR and comparable regimes, the goal is to reproduce the error, not to create a second uncontrolled copy of candidate data in a system that was never assessed for it.

Expected output. What a correct response would have been: "the summary should reflect only experience stated on the resume." Actual output. What the AI actually produced, quoted exactly: "summary stated 5 years ICU experience; resume lists 2 years med-surg, no ICU." Severity. Critical for errors affecting a hiring decision or creating legal or compliance exposure, such as hallucinated experience, biased scoring, or leaked candidate data. Major for a notable error a reviewer would have to fix before using the output, such as a wrong title or a dropped qualification. Minor for cosmetic or low-impact issues like a formatting slip costing a few seconds. Impact and status. What the error would have caused if it had shipped, and whether the report is open, routed, or resolved.

The discipline is specificity. "Bad summary" is not a report. "Summary of candidate #4471 hallucinated 5 years ICU experience not present in the resume; critical; would have advanced an underqualified candidate to manager review" is a report someone can act on months later without needing to ask you what you meant.

Routing: Who Fixes What

Capturing an error is wasted effort if it lands nowhere. The most common failure Wesley saw before building his loop was a perfectly good error report sitting in a chat thread because nobody owned the next step. Routing solves that by assigning each error type to a destination before anyone needs one.

The prompt owner handles most errors. If a summarization prompt keeps hallucinating, or an extraction template keeps missing certifications, the fix is usually a prompt change: adding a constraint such as "include only experience explicitly stated in the source" or "return null for any field not present, never infer." On Wesley's team, each core prompt has a named owner responsible for revisions and version notes, so a change can be traced to the report that prompted it.

The vendor handles errors that no prompt change can fix: a model that consistently mangles a specific format, a tool feature that breaks, repeated behavior that looks like a model-level limitation. Reporting these matters because AI tools improve partly through aggregated user feedback, and a documented, reproducible report carries far more weight with a vendor than "it sometimes gets things wrong." If the issue proves unfixable at the vendor level, that is itself an input into tool selection.

The admin or program owner handles errors that cross into policy: a bias error, a candidate-data handling concern, anything with EEOC or GDPR implications. These do not just get a prompt patch. They get escalated, documented for the compliance record, and may trigger a pause on the affected workflow until reviewed. A simple routing rule keeps everything moving: minor and major errors tied to a prompt go to the prompt owner; persistent errors that survive prompt fixes go to the vendor; any critical bias or data error goes to the admin first, immediately, before anything else happens.

Where Feedback Comes From

Self-reported errors are the backbone of the system but not the whole of it. Wesley pulled from five sources, and each one surfaces a class of problem the others miss.

Your own use is the primary source. You run the tool, you notice the errors, you log them. Everything else in this lesson depends on this habit being reliable, because a system fed by nobody produces nothing. Team feedback is the second, and it needs a channel rather than an invitation. Wesley's recruiters had a single spreadsheet and a single tagged chat channel, and the constraint was deliberate: two reporting routes get used, five get ignored.

Candidates are a source most teams overlook. Occasionally a candidate will tell you directly that the AI got something wrong, whether that is "your email said I worked at a clinic I never worked at" or "your message called me by another name." Those reports are high signal because the candidate has no reason to invent them and every reason to be accurate about their own history. Listen, log them, and treat them with more weight rather than less, since for every candidate who mentions it there are others who noticed and said nothing.

Outcomes are feedback in a slower form. Track whether the candidates AI helped you screen in actually worked out, and whether candidates screened out turned out to fit somewhere else in your organization. Misses and false positives are data about your screening prompts even though nobody filed a report about them. Formal audits complete the set: periodically sample your AI-assisted decisions systematically and look for patterns of error or bias that no individual report would reveal on its own. An audit catches the distributed error, the one that is small in every instance and consequential in aggregate.

A Worked Example: Wesley's First Quarter of Error Logging

Over one quarter, Wesley's eleven recruiters logged 84 AI errors against roughly 1,400 AI-assisted tasks, a baseline error rate near 6 percent. When he categorized the log, the pattern was not evenly spread:

  • Wrong extractions: 31 errors, 37 percent of the total, almost all major severity, concentrated in the resume extraction prompt for clinical roles.
  • Formatting failures: 22 errors, 26 percent, nearly all minor, from outreach drafts ignoring the length limit.
  • Hallucinations: 14 errors, 17 percent, nine of them critical, from the summarization prompt inventing experience.
  • Misinterpretations and omissions: 13 errors, 15 percent, mixed severity.
  • Bias-flagged output: 4 errors, 5 percent, all critical, routed straight to the admin.

The analysis pointed at two root causes rather than eight separate problems. The clinical extraction prompt did not tell the model what to do when a field was absent, so it guessed, and that single gap produced both the wrong extractions and the hallucinated summaries downstream. The outreach prompt stated a tone but never a hard word limit, which explains the entire formatting cluster.

The actions followed directly from the analysis. The prompt owner rewrote the extraction template with one added constraint: "For any field not explicitly stated in the source document, return 'Not stated'. Never infer, estimate, or fill from typical career patterns." The summarization prompt got a parallel rule. The outreach prompt got an explicit "Keep under 90 words" line. Three sentences of prompt text, each traceable to a specific cluster in the log.

The following quarter, against a similar task volume, logged errors fell to 29, roughly a 2 percent error rate, down from 6 percent. Critical hallucinations dropped from nine to one. The four bias-flagged cases the admin reviewed led to retiring an experimental "culture fit" scoring prompt entirely, because its criteria could not be defended against an EEOC adverse-impact challenge. None of that improvement came from a smarter model. It came from reading the log and acting on what it showed.

Building a Culture Where People Actually Report

A feedback system only survives if reporting feels safe and worthwhile. The fastest way to kill one is to make people feel that logging an error reflects badly on them, or to collect reports that disappear into a void. Four principles keep the system alive.

Feedback is improvement, not blame. The framing is "we found an error, how do we fix it," never "someone made a mistake." Wesley made the unit of evaluation the prompt and the process rather than the recruiter who caught the problem, which is the only framing that survives contact with a bad quarter. A team that believes error reports are performance evidence will stop producing them within a month, and the errors will not stop, they will just stop being visible.

Make reporting easy. Do not build bureaucracy around it. A single spreadsheet row or a tagged message in one channel is the right amount of friction; a multi-field portal with required approvals is a system optimized against its own purpose. Every additional field you require costs you reports, and the reports you lose are disproportionately the ones filed by busy people in the middle of real work.

Close the loop, visibly. When someone reports an issue, tell them how it will be fixed, on what timeline, and later whether it worked. When Wesley's recruiter reported the outreach length problem, they heard back: here is the prompt change, it shipped this week, and formatting errors dropped by two-thirds. That follow-through is what keeps people reporting next month. A loop that takes input and never reports back trains the team to stop bothering, and once that happens rebuilding the habit is much harder than establishing it was.

Celebrate the catch. When someone catches an AI error that would have caused a real problem, that is a win and should be acknowledged as one. The recruiter who spots a hallucinated summary before it reaches a hiring manager is the system working exactly as designed, not evidence that AI is failing. Teams that treat catches as wins find more of them.

Anti-Patterns That Kill the Loop

Three failure modes account for most dead feedback systems. The first is having no feedback system at all. Errors happen, people work around them individually, and nothing is captured or analyzed. What goes wrong is that you keep making the same mistakes, because the knowledge required to prevent them lives in eleven separate heads and never assembles into a pattern. The defense is to build the system up front, before the errors accumulate, since retrofitting a log onto a quarter you have already forgotten produces nothing usable.

The second is feedback without action. You collect reports diligently and then do nothing with them. This is worse than not collecting, because the team invested effort and got nothing back, and the next time someone hesitates before filing they will decide it is not worth it. The system does not fail loudly; it just goes quiet. The defense is to close the loop every time: act on the pattern, then tell people what changed and what it did to the numbers.

The third is a blame culture. Reporting an error feels risky because it might reflect poorly on the reporter, on their judgment, or on the tool they championed. What goes wrong is that people hide errors instead of reporting them, and you lose exactly the reports that mattered most, since the highest-stakes errors are the ones most tempting to quietly fix and forget. The defense is framing feedback as improvement rather than fault, and applying that framing hardest when the error was expensive.

Practice

Work these in order. The first two produce the artifacts the last three exercise.

  • Design your feedback system. Decide what tool you will use, what information you will capture, who reports, and how often you analyze. Write it down as a one-page description a new team member could follow without asking you anything.
  • Create your capture template. Draft the fields for a simple error report: context including the prompt name, input reference, expected output, actual output, severity, and impact. Add your redaction rule for candidate data before anyone uses it.
  • Run a full cycle. Use AI for a real task over a week, capturing every error you hit. Categorize them by type, then analyze for patterns: which type dominates, when it happens, and what the likely root cause is.
  • Identify improvement actions. From your analysis, decide what changes: a prompt constraint, a process step, a training point, or a tool decision. Name an owner and a date for each.
  • Test the change. Run the improved prompt or process on comparable work and count the errors again. Did the rate move? If it did not, the diagnosis was wrong and the log will usually tell you why.
  • Set your routing rules. Write down which error types go to the prompt owner, which go to the vendor, and which go straight to the admin, then post the rules where reports are filed.

Reflection

  • What errors have you seen most often in your own AI use, and what is the pattern connecting them?
  • If you had a feedback system running for the last six months, what do you think it would have shown you about your prompts?
  • How would your team respond to being asked to report AI errors, and what does that answer tell you about the culture you are working with?
  • What would change about your process if you committed to one improvement action per month based on the log?
  • Which of your current AI outputs would nobody catch an error in, because nobody checks them against a source?
  • Who on your team would be the right owner for each of your core prompts, and do they know they own it?

Glossary

  • Feedback loop. The cycle of capturing errors, categorizing and analyzing them for patterns, and acting to prevent recurrence, then measuring whether the action worked.
  • Capture. Recording a specific error with its context, evidence, and impact at the moment it is found, while the details are still recoverable.
  • Categorize. Grouping logged errors by type, such as hallucination, misinterpretation, omission, or bias, so that patterns become visible across incidents.
  • Root cause. The underlying reason a cluster of errors occurred, typically a missing prompt constraint or an unspecified handling rule rather than a one-off model failure.
  • Improvement action. A specific change to a prompt, process, training, or tool made on the strength of the analysis, with a named owner attached.
  • Severity tier. The three-level rating that sets priority: critical for decision or compliance impact, major for errors requiring a fix before use, minor for cosmetic issues.
  • Routing. The rule assigning each error to whoever can actually fix it: the prompt owner, the vendor, or the admin with policy authority.

Closing

Every error is an opportunity to improve, but only if it is captured while it is still recoverable and acted on while it still matters. When you capture and act on feedback systematically, your AI use becomes measurably more reliable over time, and that reliability is what separates organizations with mature AI practices from those that keep rediscovering the same failures. Wesley's team did not get better because the model improved. They got better because eleven people started writing down what went wrong, someone read the list, and three sentences of prompt text changed as a result.

This week, build your feedback system. Pick the tool, define the fields, tell your team what counts as an error and what does not, and start capturing. At the end of the month, categorize what you collected, find the dominant pattern, and take one improvement action against it with your name and a date attached. Then tell the people who reported it what you did. Responsible organizations build feedback systems and act on them, and the acting is the part that is easy to skip and impossible to fake.

Key Takeaways

  • AI improves through feedback, not by accident. The learning you generate by using the tool is worth nothing until it feeds back into your prompts, your processes, and your vendor relationships.
  • Run the four-step loop: capture, categorize, analyze, act. Skipping any step breaks the others. Capture without categories produces a pile, categories without analysis produce counts, and analysis without action kills the system.
  • Define what counts as an error before you collect any. Wrong extractions, hallucinations, omissions, biased output, and formatting failures. The working test: if the output used as-is would mislead a hiring decision, embarrass the company, or create legal exposure, it is reportable. A style preference is not.
  • Document with six fields, and be specific. Context including the prompt name, input, expected output, actual output, severity, and impact. "Bad summary" is useless; "hallucinated 5 years ICU experience not in the resume, critical" is actionable.
  • Protect candidate data in the log itself. Record a redacted excerpt or a reference to the source document, never a full resume with personal identifiers, so that GDPR discipline is not undone by the error tracker.
  • Use three severity tiers to drive priority. Critical errors affect a decision or create compliance exposure, major errors need a fix before use, minor errors are cosmetic. Severity decides what escalates immediately versus what waits for the next prompt revision.
  • Route every error to a clear owner. Prompt owner for most fixes, vendor for model-level issues no prompt change resolves, admin first and immediately for any bias or candidate-data error with EEOC or GDPR implications.
  • Draw on all five feedback sources. Your own use, your team, candidates who tell you directly, outcomes that reveal misses and false positives, and periodic formal audits that catch what no individual report would.
  • Analyze for root cause, not incident count. Two missing prompt constraints explained the great majority of Wesley's 84 logged errors, which is why the fix was three sentences rather than eighty-four conversations.
  • Reported errors drive real, measurable improvement. Adding "return 'Not stated' for absent fields" and a hard word limit cut the error rate from roughly 6 percent to 2 percent and dropped critical hallucinations from nine to one in a single quarter.
  • Make reporting safe, easy, and closed-loop. Frame feedback as improving the prompt rather than blaming the person, keep logging to one low-friction step, celebrate the catch, and always report back what changed. A loop that never closes trains the team to stop reporting.

Frequently Asked Questions

How do I start a feedback system when nobody on my team is reporting anything today? Start with yourself and one artifact. Log your own errors for two weeks in a single spreadsheet with the six fields, then show the team the pattern you found and the prompt change it produced. A demonstrated fix is a far better recruitment pitch than a request to fill in a form, because it answers the question everyone is silently asking, which is whether reporting accomplishes anything. Then ask for reports on one narrow category first, such as hallucinations in summaries, rather than opening the floodgates to every kind of complaint. Narrow scope produces usable data and builds the habit; open scope produces a mix of severity levels nobody can prioritize.

How much detail is enough in a report? My recruiters will not fill in six fields. Six fields sounds heavier than it is, because four of them are a sentence or a copy-paste. The real minimum is that someone reading the entry three months later can reproduce the error and knows how bad it was, which in practice means the prompt name, the actual output quoted exactly, and the severity. If your team will not do six, start with those three and add the rest as the habit sets. What you cannot compromise on is specificity in the "what went wrong" line, because a log full of "bad output" entries cannot be categorized and therefore cannot be analyzed, which means the whole system collapses into a complaints box.

How often should we analyze the log rather than just adding to it? Monthly is a reasonable rhythm for most teams, with a fuller review each quarter. The constraint is that patterns need volume to become visible: analyzing six entries produces a false sense of a trend, while waiting a full year means acting on prompts you have since changed for other reasons. If a critical error appears, do not wait for the cycle. Critical severity is defined precisely by the fact that it needs action now, and the scheduled analysis is for finding the slow patterns rather than for handling the emergencies.

Is it worth reporting errors to the vendor, or does that go into a void? It is worth doing for the errors that survive your own prompt fixes, and the quality of the report determines whether it goes anywhere. A reproducible case with the exact input, the exact output, and the number of times you have seen it is a different artifact from "your tool makes things up sometimes." AI tools improve partly through aggregated user feedback, so a documented, reproducible report has a real chance of influencing the product. It also serves a second purpose regardless of the vendor's response: it creates a record that you identified the limitation and reported it, which matters when you are later deciding whether to renew and when you are explaining your diligence to anyone who asks.

What do I do when the error is a bias error? Route it to the admin or program owner first, immediately, before the prompt owner sees it and before anyone attempts a quiet fix. Bias errors are not prompt bugs with a compliance flavor; they carry legal exposure. Under EEOC guidance, an AI tool that produces adverse impact in screening exposes the employer to liability regardless of intent, so the sequence is escalate, document for the compliance record, and consider pausing the affected workflow until it has been reviewed. Wesley's four bias-flagged reports in one quarter led to retiring a scoring prompt outright, because its criteria could not be defended against an adverse-impact challenge. That is the right kind of outcome: a bias pattern is a reason to stop using something, not a reason to tune it until the complaints stop.

Our error rate went down. How do I know the improvement is real and not just people reporting less? That is exactly the right suspicion, and the way to test it is to look at composition rather than only at the count. In Wesley's second quarter the total fell and the critical hallucinations fell disproportionately, which is what a targeted fix should do, because the prompt change addressed the specific mechanism producing them. A drop caused by reporting fatigue tends to be flat across categories instead, since disengaged reporters stop logging everything rather than one thing. Two other checks help: run a small audit sampling outputs yourself, independent of what was reported, and watch whether people are still reporting minor issues, because the minor reports are the first to disappear when a team stops believing the loop closes.