←
AI for HR Certification
Strategic · M12 · lesson 12 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Incident Response: When AI Makes a Bad People Decision, Complete Playbook
📖
now learning

Incident Response: When AI Makes a Bad People Decision, Complete Playbook

15 min

Overview

Your AI recruiting tool screens out a qualified candidate from an underrepresented background. A manager escalates: "I reviewed this resume. They were clearly qualified. Why did the AI screen them out?"

You're now in an incident. What do you do? Who do you call? What's your timeline? Who communicates with the candidate? Do you reverse the decision? Do you contact regulators? Do you pause the tool?

Without a playbook, you're improvising under pressure. With a playbook, you follow a proven process: identify, assess, contain, communicate, remediate, learn.

This lesson teaches incident response for HR AI, the step-by-step playbook for when things go wrong. You'll learn what "incidents" actually are (not just dramatic stuff). You'll understand escalation levels and timelines. You'll build a response playbook with clear roles, communication templates, and remediation steps. And you'll learn to contain damage while using the incident to improve future practices.

AI incidents aren't rare. If you're deploying AI in HR, you will have incidents. The question isn't whether you'll face them. It's whether you'll handle them with confidence and speed. This lesson prepares you for that moment.

What Counts as an Incident? Severity Levels

Not everything is an incident. Not everything requires emergency action. But some things do. Define your severity levels upfront so you know how to respond.

Critical Incidents (Act Within 24 Hours):
- Bias discovered: AI systematically disadvantaging protected group (selection rates differ >15%)
- Data breach: Employee/candidate data exposed
- Discriminatory decision made and acted on: AI screened someone out, hiring decision was finalized based on that screening, they were rejected (legal exposure)
- Regulatory inquiry: DOL, EEOC, state regulator asking about your AI practices

High-Priority Incidents (Respond Within 48 Hours):
- Accuracy concerns: Tool performance dropped >10% from baseline
- User escalation: Multiple people reporting same issue
- Compliance concern: Tool may violate law (discovered GDPR compliance gap, for example)
- Vendor issue: Vendor went down, can't access tool, data inaccessible

Medium-Priority Incidents (Respond Within 1 Week):
- User confusion: Multiple people asking same question ("Why was I screened out?")
- Process breakdown: Tool isn't integrating smoothly with existing workflow
- Performance issues: Tool is slower than expected

Low-Priority Issues (Handle Normally):
- Individual user questions
- Feature requests
- Minor bugs

Incident Response Playbook: Six Steps

Step 1: Identify & Report (Immediate, Day 0)

Who reports? Anyone, hiring manager, user, candidate, employee.

How? Clear reporting channel:
- Email to [[email protected]] with subject: "[INCIDENT] [issue type]"
- Slack channel: #ai-incidents
- Phone call if critical

What to include:
- What happened? (Be specific)
- When did you notice it?
- Who's affected? (How many people?)
- What's the severity? (critical/high/medium/low)
- Any initial data/screenshots?

Example report:
"The resume screening AI screened out 15 candidates from a specific ethnic background yesterday. I reviewed a sample, all were qualified upon manual review. This looks like bias. CRITICAL."

HR AI Lead's action: Receive report. Acknowledge within 1 hour. Schedule governance trio meeting.

Step 2: Assess & Investigate (24-48 Hours)

Who assesses? Governance trio (HR lead + IT + Legal)

What do you do?


  • Confirm the issue. Is the claim accurate? Is this real or misunderstanding?
    - Pull data
    - Verify the pattern
    - Talk to relevant parties

  • Understand scope. How many people affected? How long has this been happening?
    - Review usage logs
    - Check affected candidates/employees
    - Determine time window (week? month?)

  • Initial root cause. Is this a tool problem, data problem, training problem, or user problem?
    - Model bias? (trained on biased data)
    - Data quality? (bad input data)
    - User misuse? (person using tool incorrectly)
    - Bug? (technical error)

  • Document findings. Write down what you found so you can reference it later.

Example investigation summary:
"Confirmed: AI screened out 15 candidates with [demographic indicator] over past week. Manual review of these candidates shows 12 were qualified. Issue existed for 1 week (since tool went live). Pattern: selection rate for Group A is 30%, for Group B is 55%. Difference of 25% (well above 4/5ths rule threshold of 20%). Appears to be model bias (model was trained on historical successful hires; historical data shows bias toward Group B). Estimated impact: 15 candidates affected, no hiring decisions made yet (these candidates haven't been formally rejected)."

Step 3: Contain & Mitigate (24-48 Hours)

If critical: Stop using the tool immediately. Revert to manual process.

If high-priority: Pause the problematic use case. Continue other uses if tool is multi-purpose.

If medium-priority: Continue using. Increase monitoring.

Example containment action:
"Paused AI resume screening effective immediately. Recruiting team reverts to manual screening for this role. All 15 candidates screened by AI will be manually reviewed before final decision. Tool resumes only after validation."

Step 4: Notify & Communicate (24-48 Hours)

Internal Notification (within 24 hours):
- Brief leadership (CHRO, CFO if business impact, General Counsel if legal risk)
- Inform affected teams (recruiting, compensation, etc.)
- Update incident log
- Determine: do we need external communication?

External Notification (timing depends on severity):

For bias issue:
- If candidates were rejected: Contact candidates. "We've identified an issue with our screening process. We're reviewing your application manually." Offer second chance.
- If no hiring decision made yet: No external notification needed (you caught it internally; no decision was finalized).
- If regulatory inquiry: Follow legal guidance. Respond to inquiry. Provide documentation of how you detected and responded.

For data breach:
- Regulatory notification within required timeline (GDPR: 30 days; some states: faster)
- Affected individuals notification (as required by law)
- Customers/stakeholders (if applicable)

For discriminatory decision:
- Contact affected individual immediately. "We made a decision using a flawed process. We want to make this right."
- Remediate (reverse decision if possible, offer appeal, make additional offer)

Communication Templates

Template 1: Internal Alert (To Leadership)

Subject: [INCIDENT] AI Bias Detected in Resume Screening

The resume screening AI made biased decisions. Details:

  • What: AI screened out candidates from [demographic group] at higher rate
    - When: [Dates when screening occurred]
    - Impact: 15 candidates affected, no hiring decisions finalized yet
    - Severity: HIGH
    - Root cause: Model was trained on historical data that reflected hiring bias
    - Status: Tool paused. Manual review process activated.
    - Next steps: Bias audit in progress. Will brief you Thursday with findings and remediation plan.

Template 2: Candidate Communication (If Decision Was Made Based on AI)

Subject: We're Reviewing Your Application

We identified an issue with our resume screening process. Your application may have been reviewed by a flawed screening system. We're manually reviewing your application now. We'll follow up within [timeline] with an update. Thank you for your patience.

Template 3: Regulatory Response (If Inquired by EEOC or Other Regulator)

[For EEOC or other regulator inquiry about AI hiring practices]

We use AI to screen resumes as part of our hiring process. Our screening process:

  • Uses [data sources] as input
    - Undergoes quarterly bias audits
    - Is reviewed by human recruiters before any hiring decision
    - Can be appealed by candidates who disagree with screening decision
    - Has safeguards [describe mitigations]

Results of our most recent bias audit: [attach report]. We found [no adverse impact / issue X which we remediated by Y].

Remediation: What to Do After

If Bias Is Confirmed:


  • Investigate root cause. Was the model trained on biased data? Does it use proxy variables?

  • Fix the issue. Options:
    - Retrain model with less biased data (oversampling underrepresented groups)
    - Accept some performance loss to be fairer
    - Add bias detection rules (flag decisions that differ by demographic group)
    - Require human review of borderline cases

  • Test the fix. Does the retrained model work? Is bias gone? Does accuracy hold?

  • Audit past decisions. Did you make bad decisions before you caught the bias? Review candidates screened out in past month. Are there candidates you should reconsider?

  • Notify affected candidates. If you screened them out based on biased screening, offer second chance. "We found a flaw in our process. We'd like to reconsider your application."

  • Document remediation. Write down what you did, what you found, how you fixed it.

If Data Breach Occurred:


  • Assess scope. What data was exposed? For how long? To whom?

  • Notify affected individuals. (Required by law in most cases)

  • Notify regulators. (If required, GDPR: yes; most US states: yes)

  • Implement security fixes. How did the breach happen? How do you prevent recurrence?

  • Legal settlement. If individuals were harmed, you may owe compensation.

If Accuracy Issue:


  • Get vendor to explain. What went wrong?

  • Test the fix. Does the fix actually improve accuracy?

  • Redeploy with monitoring. If vendor can fix, deploy and monitor closely.

  • Find alternative vendor. If vendor can't fix, start searching for replacement.

If Technical Failure (Tool Crashes):


  • Restore service. Get the tool back up ASAP.

  • Assess impact. How long was tool down? How many decisions affected? Can you make those decisions manually?

  • Implement redundancy. How do you prevent this in future? Backup system?

  • Document incident. Root cause, timeline, resolution.

Incident Response Timeline Example

Real-world example of a bias incident:

Day 0 (10am): Recruiter notices AI screened out qualified candidate from underrepresented background.

Day 0 (11am): Recruiter escalates to recruiting lead. Recruiting lead checks data. It's not isolated, 10+ candidates from same background screened out over past week.

Day 0 (2pm): Recruiting lead escalates to HR AI lead. "This looks like systematic bias."

Day 0-1 (By next morning): Governance trio confirms: this isn't isolated. 10+ candidates affected over past week. Pattern is clear: selection rate for Group A is 30%, for Group B is 55%. Looks like systematic bias.

Day 1 (10am): Decision: pause resume screening tool. Revert to manual screening for this role.

Day 1 (11am): Internal communication: Brief CHRO and Legal. Document findings.

Day 2-3: Investigate root cause: Bias audit of model. Discover: model was trained on 5-year-old hiring data that was biased toward Group B. New hire recommendations skewed Group B.

Day 4: Remediation plan: Retrain model with balanced training data. Implement bias monitoring. Add human review for borderline candidates.

Day 5: External communication: Contact 10 affected candidates. "We identified a flaw in our screening process. We're reviewing your applications manually. Are you still interested?" Ask if they'd like to reapply or have their resume reconsidered.

Day 7: Resume screening restarts with retrained model. Implement ongoing bias monitoring.

Day 30: Follow-up: Report to leadership on remediation and improvements.

Day 90: Quarterly audit: Bias audit on retrained model. Confirm no disparate impact.

Real-World Incident Examples and How They Were Handled

Example 1: The Screening Bias Discovery (Critical Incident)

A recruiting manager reviewing resumes noticed something odd: the AI screened in 45% of male candidates but only 28% of female candidates, all from the same talent pool. Both groups had similar qualification distributions.

Day 0 (2:30pm): Manager reports to HR AI lead with data
Day 0 (3:00pm): HR AI lead acknowledges, convenes governance trio
Day 0 (4:30pm): Governance trio confirms: pattern is real, spans past 3 months, 87 female candidates affected
Day 1 (9:00am): Decision made: pause screening immediately, manually review all 87 candidates
Day 1 (11:00am): Internal brief: CHRO, General Counsel, Finance (for business impact)
Day 2-3: Root cause analysis: model trained on 5-year historical data; 70% of historical hires were male; model learned this bias
Day 5: Retrain model with balanced training data; parallel test with 1,000 held-out candidates; new model shows 40% male, 42% female (no disparate impact)
Day 8: All 87 affected candidates manually reviewed; 31 found qualified and re-contacted
Day 30: Quarterly monitoring begins; subsequent audits show no disparate impact
Outcome: Tool redeployed successfully; 8 of the 31 candidates made it to interview stage; 2 hired. Incident handled transparently; candidate trust maintained.

Key success factors: Fast acknowledgment, clear governance structure, root cause investigation, transparent communication, and follow-through on remediation.

Example 2: The Data Breach (Critical Incident)

An IT audit discovered that candidate interview data (names, emails, phone numbers, interview notes containing sensitive information) was being stored in an unencrypted cloud folder that was inadvertently made public for 2 weeks. Approximately 2,000 candidates affected.

Day 0 (11:00am): IT reports to CHRO and General Counsel
Day 0 (12:00pm): Decision: confirm scope, assess regulatory obligation, plan notification
Day 0 (4:00pm): Scope confirmed: 2,000 candidates, 2-week exposure window, no evidence of misuse
Day 1: Legal determines: California, New York, and 4 other states require notification within 30 days
Day 2: Draft notification letter to affected candidates; brief leadership on regulatory and reputational risk
Day 3: Notification begins via email; set up hotline for candidate questions
Day 4: IT fixes the security issue; implements access controls; audits for recurrence
Day 7: Follow-up communication to candidates with offer of credit monitoring (within budget)
Day 30: Report filed with relevant state regulators as required
Day 60: Post-incident review: what happened, why, how to prevent recurrence
Outcome: Handled with transparency; minimal candidate backlash; regulatory compliance achieved; security improved.

Key success factors: Speed of detection, clear regulatory knowledge, transparent communication, remediation of the underlying issue, and follow-through on post-incident review.

Example 3: The Accuracy Drop (High-Priority Incident)

An HR leader running monthly audits on an AI performance evaluation tool notices something: accuracy on predicting "high performers" dropped from 87% to 71% in the latest month. Previous months were consistent.

Day 0: Discovery during routine audit
Day 0: Escalate to vendor (AI tool provider) asking what changed
Day 1: Vendor investigation begins; they identify: algorithm was updated 3 weeks ago; update not communicated to HR; update degraded performance
Day 2: Governance trio meets; decides: continue using (previous month still accurate; newer month can be manually reviewed), but require vendor to remediate
Day 3: Vendor submits a rollback plan
Day 5: Rollback deployed; accuracy returns to 87%
Day 7: Post-incident review: establish clear vendor communication requirements; require advance notice of algorithm changes
Outcome: Caught early due to robust monitoring; minimal impact; vendor accountability established.

Key success factors: Regular audits that catch problems early, clear escalation to vendor, decision framework that balances risk and business continuity, and follow-up governance improvements.

Incident Response Checklist

When incident is reported:

  • [ ] Acknowledge within 1 hour
    - [ ] Assess severity (critical/high/medium/low)
    - [ ] Convene governance trio within 4 hours if critical
    - [ ] Investigate and document findings within 24 hours
    - [ ] Make containment decision within 24 hours (pause/continue/modify)
    - [ ] Notify leadership within 24 hours if critical
    - [ ] Notify affected parties within 48 hours if necessary
    - [ ] Implement fix within 3-7 days
    - [ ] Validate fix (re-audit, re-test)
    - [ ] Resume tool after validation
    - [ ] Document entire incident for learning

Advanced: When to Escalate to External Parties

When to notify regulators (EEOC, DOL, state agencies):

  • Discriminatory decision was made and acted on (hiring, firing, pay decision based on biased AI)
    - Data breach involving sensitive employee/candidate data
    - Regulatory inquiry received
    - Pattern of disparate impact that suggests systemic discrimination

Example: If you screened out 50 women from 100 candidates (50% vs. 70% screening rate for men), that's disparate impact. If 30 of them made formal complaints, EEOC will likely investigate. Get ahead of it: proactive disclosure, document your remediation, show you took it seriously.

When to notify affected candidates/employees:

  • Decision based on biased AI was finalized and candidate/employee was rejected
    - Data breach affecting personal information
    - Adverse impact discovered and you want to reconsider affected individuals
    - System malfunctioned and affected their opportunity

When to involve external counsel:

  • Anticipated litigation or regulatory investigation
    - Data breach with potential legal liability
    - Public incident or media inquiries
    - Uncertainty about legal obligations

CALLOUT BOX: Incident Response Under Pressure

When incident happens:

Do: - Gather facts (don't make assumptions)
- Move fast (make decisions quickly)
- Communicate clearly (internal and external)
- Document everything (for legal protection and learning)
- Focus on containment first, analysis second
- Involve counsel early if legal risk exists

Don't:
- Panic (stay calm, follow the process)
- Blame (focus on fixing, not finger-pointing)
- Hide (transparency is better than secrecy)
- Delay (faster action = better outcome)
- Assume (verify before reacting)
- Post-mortem without learning (if incident happens, extract the lesson)

Deliverable: Your Incident Response Playbook (2 pages)

Create a document covering:

Page 1: Incident Classification & Response
- What counts as incident (critical/high/medium/low)
- Response steps (identify, assess, contain, communicate, remediate)
- Escalation triggers (when to notify leadership, regulators, candidates)
- Timeline expectations (how fast do we act?)

Page 2: Communication & Process
- Communication templates (internal alert, candidate communication, regulatory response)
- Remediation process (how do we fix each type of issue?)
- Incident log (how do we track and learn from incidents?)

What to Do Monday Morning


  • Define severity levels. What's critical vs. high vs. medium? Get leadership to agree.

  • Create incident reporting channel. Email address? Slack channel? Phone number? Make it clear how to report.

  • Establish response timeline. How fast do you respond to each severity level?

  • Schedule governance trio. Make sure they know to prioritize incident calls.

  • Write incident response playbook. Steps, templates, roles, timelines.

  • Share with team. Make sure everyone knows what to do if they discover an incident.

Key Takeaways


  • Incidents will happen. Have a plan so you respond quickly and effectively.

  • First 24 hours are critical. Identify, assess, contain. Speed matters.

  • Communication matters. Transparency builds trust. Hiding makes things worse.

  • Document everything. For compliance, learning, and legal defense.

  • Fix and learn. When incident happens, fix the immediate problem and fix the system issue so it doesn't happen again.

  • Bias audits and incident response are your insurance. They reduce likelihood and severity of incidents.

Post-Incident Learning and Improvement

One of the most important aspects of incident response is using the incident to improve your systems and processes. This turns a crisis into a learning opportunity.

Post-incident review (within 30 days):

  • What happened? Timeline of events, with key facts
    - Why did it happen? Root cause analysis. Not "the model had bias" but "the model was trained on biased historical data because we didn't have a process to audit training data for bias"
    - What did we do about it? Our response, timeline, decisions made
    - What could we have done differently? Would earlier detection have helped? Better monitoring? Different governance?
    - What are we changing? Specific process, policy, or governance changes to prevent recurrence

Example of a learning-driven improvement:

After the screening bias incident (Example 1), the company made three changes:
- Quarterly bias audits (not annual) for all employment AI
- Mandatory review of training data for all new AI models (before deployment)
- New governance rule: any model showing disparate impact >15% must have documented mitigation before redeployment

Cultural impact of learning from incidents:

When organizations handle incidents transparently and use them to improve, it sends a message: "We're serious about responsible AI. When something goes wrong, we fix it." This builds trust. Candidates and employees see that you care.

The alternative, hiding incidents, covering them up, not improving, gets discovered eventually. And when it does, the trust damage is far worse than the original incident.

FAQ

Q: If we find bias, are we liable?

A: Depends on timing. If caught before hiring decision: no liability. If hired based on biased screening: potential liability (discrimination claim). That's why you catch and remediate before decisions are made. Documentation of your audit and remediation is your defense.

Q: How long before we can resume using a tool after an incident?

A: Only after (1) identify root cause, (2) fix it, (3) test the fix with historical data, (4) get Legal sign-off. Don't rush back online. For a significant incident, expect 1-2 weeks minimum.

Q: Do we have to tell candidates about the bias we found?

A: Not legally required in most cases, but it's ethical. "We found a flaw in our process and corrected it" builds trust. Silence makes people suspicious. If candidates were affected by the biased decision (screened out unfairly), telling them and offering reconsideration is the right move.

Q: What if the incident is the vendor's fault?

A: Vendor is responsible, but you're accountable to your candidates and employees. Hold vendor accountable via contract (liability clauses, performance standards). But you still have to respond and remediate with your affected people. Your contract should specify vendor's obligation to notify you of issues, provide transparency, and implement fixes quickly.

Q: What if an employee or candidate is harmed and sues?

A: Your documentation of incident response is your defense. You can show: "We discovered the problem, assessed it, fixed it, and prevented further harm." That's a strong defensive position. Hiding the problem is indefensible.

Building Your Incident Response Culture

For incident response to work, you need a culture where problems are reported early, not hidden.

How to encourage reporting:

  • No blame for reporters. If someone reports a problem and they're punished, nobody will report next time. Make it safe to report.
    - Reward early detection. If someone catches a bias problem in testing before deployment, celebrate that. Early detection is vastly better than discovery after deployment.
    - Transparency about incidents. When incidents happen, communicate broadly (not gossip, but facts). "We found bias in our resume screening. Here's what we did. Here's what changed." This signals that reporting leads to action.
    - Regular simulations. Run "incident response drills" quarterly. Imagine scenarios. Walk through the playbook. This keeps the team sharp and identifies gaps in your process.

Developing incident response capability in your team:

  • HR AI lead or governance lead owns incident response (should be one person's primary responsibility)
    - Governance trio (HR, IT, Legal) trained on incident response (do a annual training)
    - Clear escalation paths documented and communicated
    - Incident response playbook reviewed and updated annually

What's Next

You've prepared for incidents. Now you need to manage vendor accountability when vendors cause problems. What do you negotiate into contracts? How do you hold vendors accountable? Next lesson: Third-Party AI Risk and Vendor Accountability.

Your incident response handles problems when they occur. Your vendor contracts prevent them from occurring. Your monitoring detects them early.