←
AI for Recruiters
Visionary · M9 · lesson 9 of 30 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Data Infrastructure: Collecting, Storing, and Analyzing Recruiting Data

15 min

Maya runs recruiting operations and analytics at Northwind, a 1,500-person software company that processes roughly 12,000 applications a year through its applicant tracking system. For two years, her team answered the same leadership questions the same painful way: someone exported a spreadsheet, hand-counted stages, reconciled it against a second export that disagreed, and produced a number three weeks after anyone needed it. The data existed. It was simply scattered across the tracking system, a calendar tool, three recruiters' private tabs, and an inbox. The turning point was not a new tool. It was deciding that recruiting data deserved the same treatment as finance data: one source of truth, defined fields, controlled access, a retention clock, and a small set of metrics everyone agreed to trust. This lesson walks through the infrastructure Maya built, the guardrails that kept it lawful and fair, and the reason fairness monitoring is impossible without any of it.

The ATS as Source of Truth

The first decision is the most consequential: the applicant tracking system is the system of record, and nothing competes with it. Every application, every stage move, every interview scorecard, and every offer lives there, and the answer to "how many candidates are in onsite for the backend role" is whatever the tracking system says, not whatever a recruiter remembers. The specific product matters far less than the discipline of designating one and refusing to let shadow spreadsheets accumulate authority.

Source of truth is a behavioral commitment, not a technical setting. The moment a hiring manager tracks candidates in a private sheet "because it is faster," the funnel numbers diverge and trust in every downstream metric erodes. Maya's rule was blunt: if a decision was made about a candidate, it gets recorded in the tracking system as a stage change with a reason, or it did not officially happen. That rule is what makes analysis possible at all, because analysis is only as good as the completeness of the underlying record. It is also what makes fairness monitoring possible, since a selection rate computed on a funnel with untracked decisions in it is a rate computed on a fiction.

What Data to Collect

Before you can analyze anything you have to decide what the record contains, and the answer is broader than most teams assume because fairness monitoring requires data that ordinary operational reporting never needs. Maya's specification covers four families of data, and each one exists for a reason she can articulate.

Data family Elements What it makes possible
Candidate data Name, contact details, education, skills, experience, and demographics such as gender, race, and age collected only with consent Identifying the person and, separately and voluntarily, grouping outcomes for fairness analysis
Recruiting data Source such as job board, referral, or recruiter; screening results as pass or fail and any score; interview results including interviewer rating and recommendation Locating which stage and which practice a disparity or a bottleneck comes from
Decision data Screening decision, interview decision, offer decision, acceptance decision, hire decision Computing selection rates at every decision point rather than only at the end
Outcome data Start date, ramp-to-productivity time, performance ratings at 30, 60, and 90 days and beyond, tenure, promotions Validating whether the hiring decisions were actually correct, by group

The demographic row carries the most caution and the most value. It is the only way to answer whether your process treats groups differently, and it is also the most sensitive category in the record, which is why it is collected with consent and, as the guardrails section explains, held apart from the people making hiring decisions. The outcome row is the one most organizations skip entirely, and skipping it means you can measure who you hired but never whether you hired well.

Establish Baselines Before You Deploy

The single most common infrastructure mistake is starting to collect the right data on the day the AI tool goes live. By then the comparison you most need is already impossible. Maya's rule is to establish data collection before AI deployment and to accumulate three to six months of pre-AI data first, so that the question "what was the selection rate by demographic group before AI?" has an answer.

Without baselines you cannot detect changes caused by AI, and that failure runs in both directions. If a screening tool introduces a disparity, you will see a number you have nothing to compare it against, and the debate will collapse into whether the number was always like that. If the tool improves things, you cannot demonstrate it, and the improvement becomes an anecdote rather than a finding. The baseline is also the only defensible way to answer a regulator or an executive who asks what changed when you introduced automation, because "we believe it got better" is not an answer and cannot be made into one after the fact.

Data Quality and Hygiene

A tracking system full of inconsistent data produces confident, wrong answers. The most common failures are mundane: rejection reasons entered as free text so that "not enough experience," "too junior," and "exp" are three different categories; sources logged as "LinkedIn," "Linkedin," and "LI" for the same channel; interview stages named differently across job templates so the funnel cannot be compared across roles. None of this is exotic, and all of it quietly destroys reporting.

Hygiene is enforced at the point of entry. Maya converted free-text fields into closed pick lists wherever a metric depended on them: a fixed set of rejection reasons, a controlled list of sources, standardized stage names applied through a shared job template. She added required fields so a stage move cannot be saved without a reason, and she scheduled a monthly review that flags records missing a source or stuck in a stage for more than 30 days. The goal is not perfection; it is a data set where the same real-world event is always recorded the same way, because that consistency is the precondition for every chart that follows.

Structured and Unstructured Data

Recruiting data comes in two shapes, and they serve different purposes. Structured data is the countable record: stage, source, dates, scorecard ratings on a defined scale, offer status. It lives in fields, it aggregates cleanly, and it is what dashboards are built on. Unstructured data is the prose: resume text, recruiter notes, interviewer feedback paragraphs, email threads. It is rich and essential for individual decisions, but it does not aggregate without first being categorized.

The practical implication is to capture the structured version of anything you intend to measure. If you want to report on why candidates are rejected, a free-text note is not enough; you need a structured rejection-reason field alongside it. The note explains the nuance for the one candidate; the field lets you count across thousands. Treating unstructured content as if it were analyzable, or trying to parse it after the fact, is where most homegrown reporting projects stall.

Collect data ethically or do not collect it. Candidates should understand what data you collect and why, which means informed consent about data use, protection of personally identifiable information, and compliance with GDPR, CCPA, and local privacy laws. Data collection that violates privacy erodes trust with the exact population you are trying to attract, and the reputational cost outlasts whatever analysis you gained.

Three mechanisms carry the weight. The first is transparent disclosure, stated plainly rather than buried: "We collect demographic information to monitor fairness in hiring. This information helps us ensure fair treatment of all candidates." The second is opt-in, meaning candidates choose whether to provide demographic information rather than having it inferred or required. The third is anonymization, removing personally identifiable information from the analysis so that the fairness work operates on groups rather than on named people.

Privacy compliance is not optional. GDPR requires explicit consent for processing personal data. CCPA gives candidates rights to access, delete, and know what you collect, which means your infrastructure has to be able to find every record about a person and act on it, a requirement that is trivial when there is one source of truth and nearly impossible when there are seven. Violating privacy laws creates legal liability and erodes trust with candidates. Notice that the compliance requirement and the analytics requirement point the same way.

Storage, Access, and Retention

Candidate data is sensitive, and it should be stored securely as a matter of course rather than as a project. Five measures make that concrete. Encryption at rest protects the stored copy and encryption in transit protects it moving over the network. Access control means only the recruiting team, HR, and data scientists who need candidate data can reach it, and none of it is public. Audit logging tracks who accessed what data and when. Regular audits have a security team review access patterns, identify anomalies, and test defenses. And a retention policy answers how long you keep data and deletes it when it is no longer needed.

Access is role-based and scoped to need. Recruiters see the candidates and reqs they work; hiring managers see their own pipelines; the analytics team queries de-identified aggregates, not named records. Nobody gets blanket access to the whole candidate database because their job title sounds senior. Access changes are logged, and access is removed promptly when someone changes roles or leaves. The audit log is what makes this enforceable rather than aspirational, because a policy nobody can verify is a policy nobody follows for long.

Retention is a clock, not a default of "keep everything forever." Under the GDPR's storage-limitation principle, personal data should not be kept longer than necessary for the purpose it was collected for, so Maya set explicit timers: rejected candidates purged or anonymized after a defined window, for example 12 months or the period local law permits for defending against discrimination claims, with a separate, consent-based talent pool for candidates who opted in to be contacted about future roles. Data minimization is the companion rule: collect only the fields the process genuinely needs, because every field you store is a field you must secure, justify, and eventually delete.

The Funnel Metrics Worth Tracking

A small set of metrics answers most questions leadership actually asks. Start with the funnel itself. Here is Northwind's year, computed illustratively from its 12,000 applications:

  • 12,000 applications to 1,800 recruiter screens: a 15 percent application-to-screen rate.
  • 1,800 screens to 600 interviews: a 33 percent screen-to-interview rate.
  • 600 interviews to 120 offers: a 20 percent interview-to-offer rate.
  • 120 offers to 90 hires: a 75 percent offer-acceptance rate.

End to end, 90 hires from 12,000 applications is an overall yield of 0.75 percent, and the inverse is the planning number recruiters care about: roughly 133 applications per hire at the top of the funnel, or about 6.7 interviews per hire deeper in it. Those ratios are what let Maya forecast: filling 30 net-new roles next quarter implies sponsoring about 4,000 applications and scheduling roughly 200 interviews, which is a staffing and scheduling problem her team can plan around months ahead.

Two more metrics complete the core set. Time-to-fill measures calendar days from job opening to accepted offer; if Northwind's median is 42 days and one engineering family runs at 71, that gap is a question worth investigating, not a number to bury. Source-of-hire breaks the 90 hires down by channel: say 36 from referrals at 40 percent, 27 from inbound applications at 30 percent, 18 from sourced outreach at 20 percent, and 9 from agencies at 10 percent. When referrals convert at several times the rate of cold applications but receive a fraction of the sourcing budget, the source-of-hire view is what reallocates spend. Conversion rate, time-to-fill, and source-of-hire are enough to run a serious operation; resist the urge to track fifty metrics nobody acts on.

Building the Analysis Capability

Collecting data and analyzing it are different capabilities, and the second one is where most recruiting functions stop short. Analysis is useless if it happens once a year, because the point of measuring fairness is to detect a problem while it is still small enough to fix cheaply. Maya's analysis layer has five components, and each one is a decision rather than a purchase.

The first is a data warehouse: a centralized location holding all recruiting data, organized for analysis rather than for transaction processing. For day-to-day operations the tracking system is sufficient, but serious cross-cutting analysis usually means syncing a copy of the structured data into a cloud warehouse such as BigQuery or Snowflake on a schedule, modeling the funnel once as governed tables, and connecting a business intelligence layer such as Looker or Tableau on top. Everyone then queries the same definitions instead of re-deriving "what counts as an interview" in each report. This is also where you decouple analysis from raw personal data, because the warehouse copy can be built to exclude direct identifiers.

The second is the metric set itself, defined once and written down: selection rates as the percentage of candidates who advance, the disparate impact ratio as the minority group's rate divided by the majority group's rate, time to decision, and offer acceptance rate. The third is reporting: automated reports generated weekly or monthly and shared with leadership and the fairness team, rather than pulled on request. The fourth is dashboards showing trends over time, comparing groups, and flagging anomalies so a human does not have to notice a drift by reading a table. The fifth is frequency, which decides whether the other four matter. Analysis must run weekly or monthly, not annually, because a fairness disparity discovered eleven months after it started has already affected everyone who applied in between.

Outcome Tracking After the Hire

Most recruiting data programs end at the hire decision, which means they can tell you what you did and never whether it was right. Once you hire someone, track outcomes: how long until full productivity, performance ratings, retention, promotions. That data validates whether your hiring decisions were correct and turns recruiting from a process you run into a process you can improve.

Five metrics carry it. Ramp-to-productivity asks how long until the hire is fully productive, and it should be similar across demographic groups. Performance ratings ask how hires actually perform, and specifically whether some groups systematically receive lower ratings. Retention asks whether some groups leave more often. Promotions ask whether some groups advance more frequently. Tenure asks how long people stay. Each of these is a number you probably already have somewhere in an HR system; the infrastructure work is joining it back to the recruiting record so the two can be read together.

The reason to do that joining is a specific question you cannot otherwise ask. "We hired more women with the tool. Are they performing as well as men? Are they staying?" If outcomes differ by demographic group, that reveals bias either in hiring or in post-hire treatment, and distinguishing between those two possibilities is itself informative: a group that performs well and leaves early points at something happening after the offer letter, which is not a recruiting problem but is certainly a recruiting signal. Either way you use the feedback to improve, and either way you cannot see it without outcome data connected to the hiring record.

Privacy and Fairness Guardrails

Storing candidate data for analysis is lawful and useful, but only inside guardrails. The hard line is special-category data. Under the GDPR, information such as race, ethnicity, health, religion, and sexual orientation is special-category data subject to heightened restrictions, and the safe operating posture is not to store it in the working recruiting record at all. Where diversity data is collected for fairness monitoring, it should be on an explicit, voluntary, separate basis, held apart from hiring decision-makers, and used only in aggregate. Anonymize or pseudonymize data for analysis so dashboards report on groups and trends, never on identifiable individuals, and so an analyst studying conversion rates is never looking at a named person's protected characteristics.

One legitimate and important analytical use is adverse-impact monitoring. A common screen, drawn from the EEOC's Uniform Guidelines, is the four-fifths rule: compare the selection rate of each group against the highest-selecting group, and a ratio below roughly 0.8, or four-fifths, flags a possible adverse impact that warrants investigation. If the highest-selecting group passes screening at 30 percent and another group passes at 18 percent, the ratio is 0.6, which is below the threshold and a signal to examine the process. The four-fifths rule is a screening heuristic, not proof of discrimination or a legal safe harbor, and a flagged result is the start of an inquiry into job-relatedness and process, not an automatic verdict. Run it on aggregated, de-identified data, document what you find, and route genuine concerns to legal and HR rather than acting on a single ratio in isolation.

Anti-Patterns

Deploying the tool before the baseline exists. This is rolling out an AI screening or sourcing tool and starting fairness measurement on the same day. It happens because the tool has a launch date and the data work does not, and because three months of collecting data with nothing to show for it looks like inaction. What goes wrong is permanent: you have post-AI numbers and no pre-AI numbers, so every conversation about whether the tool helped or harmed becomes an argument about a counterfactual nobody can produce. The counter is to sequence deliberately, collecting data first, running a three to six month pre-AI window second, and rolling out third, and to say out loud that the window is part of the deployment plan rather than a delay to it.

Letting shadow spreadsheets accumulate authority. This is a hiring manager tracking candidates in a private sheet because the tracking system feels slow. It happens one req at a time, always for a locally reasonable reason, and nobody ever decides to fragment the record. What goes wrong is that funnel numbers diverge, reconciliation becomes a recurring tax, and decisions recorded outside the system are invisible to every selection-rate calculation you run, which means your fairness monitoring is silently computed on a subset. The counter is Maya's blunt rule, backed by making the in-system path fast enough that the workaround stops being attractive.

Free text where a field belongs. This is collecting rejection reasons, sources, and stage names as prose, then discovering at reporting time that "not enough experience," "too junior," and "exp" are three categories and one channel is spelled three ways. It happens because free text is the path of least resistance at entry and the cost lands on someone else months later. What goes wrong is that you cannot count anything, and the reports you do produce are confidently wrong in ways nobody can see. The counter is closed pick lists on every field a metric depends on, required reasons on stage moves, and a monthly sweep for records missing a source or stuck in a stage.

Annual analysis. This is building dashboards and then reviewing them once a year, or pulling fairness numbers only when someone asks. It happens because the infrastructure feels like the hard part and the cadence feels like a scheduling detail. What goes wrong is that the detection lag becomes the damage: a disparity introduced in February is found in December, and everyone who applied in between was processed through it. On-request analysis also makes the review feel adversarial when it finally happens. The counter is automated weekly or monthly reporting delivered to leadership and the fairness team on a schedule, with dashboards that flag anomalies rather than requiring someone to notice a drift by reading a table.

Collecting sensitive data into the working record. This is putting demographic and other special-category fields directly into the candidate record where recruiters and hiring managers can see them, usually with good intentions about monitoring fairness. What goes wrong is that you have created exactly the exposure the fairness program exists to prevent: decision-makers can see protected characteristics at the moment of decision, and you are holding GDPR special-category data in the least defensible place possible. The counter is to collect diversity data explicitly, voluntarily, and separately, hold it apart from decision-makers, use it only in aggregate, and analyze on anonymized data.

Practice

  • Audit your current data infrastructure. Write down what candidate data you collect today across demographics, source, decisions, and outcomes; how it is stored and whether it is encrypted and audited; what analysis is currently possible including fairness metrics, dashboards, and frequency; and what is missing for comprehensive fairness monitoring. Turn the result into a gap analysis with prioritized investments.
  • Design your data collection specification. Identify each element you need for baseline and ongoing monitoring, the systems, forms, and processes that will capture it, how consent will be obtained, and how you will validate quality and catch errors at entry.
  • Write the privacy and security plan. Name the privacy regulations that apply to you, state how you will comply through consent, minimization, retention, and access control, list the security measures required including encryption, audit logging, and monitoring, and map which role needs access to which data.
  • Specify your fairness metrics and reporting. Decide which metrics you will track, how you will establish baselines including the pre-AI collection timeline, what your reporting frequency will be, and how results will be visualized and communicated to which stakeholders.
  • Design the outcome tracking system. Choose which post-hire outcomes you will track, where each will come from, how you will compare them by group to identify disparities, and how findings will feed back into hiring changes or post-hire interventions.
  • Run your funnel once, end to end, then run the four-fifths screen on it. Compute your application-to-screen, screen-to-interview, interview-to-offer, and offer-acceptance rates, derive applications per hire, and turn next quarter's hiring plan into a top-of-funnel volume target. Then, on anonymized data, compare each group's screening pass rate against the highest-selecting group's rate, mark any ratio below 0.8, and document where a flagged result would be routed.

Reflection

  • If you deployed an AI screening tool next month, what pre-deployment baseline could you actually produce, and how long would it take to build one you would defend?
  • Where in your organization does a hiring decision get made that never becomes a stage change in the system of record?
  • Which of your reporting fields are free text today, and which metric silently depends on one of them?
  • How long does your organization keep a rejected candidate's data, who decided that, and is it written down anywhere?
  • If a candidate exercised a right to know what you hold about them, how many systems would you have to search, and could you be confident you found everything?

Glossary

  • Source of truth. The single designated system whose record is authoritative. A behavioral commitment rather than a technical setting, and the precondition for every metric downstream.
  • Structured data. The countable record: stage, source, dates, ratings on a defined scale, offer status. It aggregates cleanly and is what dashboards are built on.
  • Unstructured data. Prose such as resumes, recruiter notes, and feedback paragraphs. Essential for individual decisions, useless for counting until categorized.
  • Data hygiene. Enforcing at entry that the same real-world event is always recorded the same way, through closed pick lists, required fields, and standardized stage names.
  • Baseline. The pre-deployment measurement, typically three to six months of data collected before an AI tool goes live, without which changes caused by the tool cannot be detected in either direction.
  • Selection rate. The percentage of candidates in a group who advance from one stage to the next.
  • Disparate impact ratio. One group's selection rate divided by the highest-selecting group's rate, screened against the four-fifths threshold of roughly 0.8.
  • Four-fifths rule. The EEOC Uniform Guidelines screen flagging a possible adverse impact when a group's selection rate falls below about 80 percent of the highest group's rate. A heuristic that starts an inquiry, not proof of discrimination or a safe harbor.
  • Data warehouse. A centralized copy of recruiting data organized for analysis rather than transactions, where the funnel is modeled once as governed tables everyone queries.
  • Encryption at rest and in transit. Protecting stored data and data moving across a network. Both are required; either alone leaves an open path.
  • Role-based access control. Scoping access to what a role needs, so recruiters see their reqs, managers see their pipelines, and analysts see de-identified aggregates. Audit logging, a record of who accessed what and when, is what makes it verifiable rather than aspirational.
  • Retention policy and data minimization. An explicit clock for how long each category of data is kept before purge or anonymization, reflecting the GDPR storage-limitation principle, paired with the rule of collecting only fields the process genuinely needs.
  • Anonymization and pseudonymization. Removing or replacing identifiers so analysis operates on groups and trends rather than on identifiable individuals.
  • Special-category data. Under the GDPR, data such as race, ethnicity, health, religion, and sexual orientation, subject to heightened restrictions and best kept out of the working recruiting record entirely.
  • Ramp-to-productivity. How long a hire takes to reach full productivity. An outcome metric that should be similar across groups and helps validate whether hiring decisions were correct.

Closing

Data infrastructure is the foundation of responsible AI in recruiting, and the reason is unglamorous: without it you are making decisions blindly. With it you have visibility into fairness, the capability to detect problems while they are small, and the ability to improve continuously rather than in response to a complaint. Every metric in this program, every selection rate and impact ratio and outcome comparison, is downstream of decisions Maya made about fields, consent, access, retention, and cadence.

The technical investment also has non-technical benefits that are easy to overlook. When your organization can articulate what data you collect, how you use it, how you protect it, and what fairness metrics you monitor, candidates and employees trust you more. Transparency about AI builds trust, trust builds reputation, and reputation is an advantage in a market where the people you most want to hire have choices. The infrastructure is what makes that articulation possible, because you cannot describe a system you have not built.

Key Takeaways

  • The applicant tracking system is the single source of truth. Designate one system as the record and enforce that every candidate decision is logged there. Shadow spreadsheets are the mechanism by which your funnel numbers, and your fairness calculations, stop being trustworthy.
  • Collect four families of data, not one. Candidate data including consented demographics, recruiting data such as source and screening and interview results, decision data at every stage, and post-hire outcome data. Most programs collect the first two and then cannot answer the questions that matter.
  • Establish baselines before AI deployment. Three to six months of pre-AI data is what lets you answer "what was the selection rate by demographic group before AI?" Without a baseline you cannot detect changes the tool caused in either direction, and the baseline can only be captured beforehand.
  • Data quality is enforced at entry, not cleaned up later. Convert measured fields into closed pick lists, require a reason on every stage move, and standardize stage names across templates. Consistent recording of the same event is the precondition for every chart you will build.
  • Capture the structured version of anything you intend to measure. A free-text rejection note explains one candidate; a structured rejection-reason field lets you count across thousands.
  • Consent is a mechanism, not a sentiment. Use transparent disclosure that states why demographic data is collected, make it opt-in, and anonymize for analysis. GDPR requires explicit consent for processing personal data and CCPA gives candidates rights to access, delete, and know what you hold.
  • Secure the data concretely, and put retention on a clock. Encryption at rest and in transit, role-based access scoped to need, audit logging of who accessed what and when, regular security review, and explicit purge or anonymization timers under the GDPR storage-limitation principle. Every stored field is one you must secure and eventually delete, which is why minimization is the companion rule.
  • Track conversion, time-to-fill, and source-of-hire, and stop there. A funnel of 12,000 applications to 90 hires yields roughly 133 applications per hire and turns hiring plans into staffing forecasts. Fifty unused metrics are worse than three acted-on ones.
  • Analysis capability is warehouse, metrics, reporting, dashboards, and frequency. Model the funnel once as governed tables, define the metric set in writing, automate weekly or monthly reports to leadership and the fairness team, and flag anomalies visually. Annual analysis means a disparity runs for a year before anyone sees it.
  • Track outcomes after the hire. Ramp-to-productivity, performance ratings, retention, promotions, and tenure, compared across groups, validate that hiring decisions were correct. Differing outcomes reveal bias in hiring or in post-hire treatment, and neither is visible without joining outcome data to the recruiting record.
  • Keep special-category data out of the working record. Collect diversity data explicitly, voluntarily, and separately from hiring decision-makers, use it only in aggregate, and run the four-fifths screen on anonymized data, treating a ratio under 0.8 as the start of an inquiry rather than a verdict.

Frequently Asked Questions

We are deploying a screening tool next month and have no baseline. What now? Say so explicitly rather than papering over it, and then do two things. First, see what can be honestly reconstructed from existing records: if stage data has been captured consistently, historical selection rates may be recoverable for a defined period and population, and that is legitimate as long as you state the coverage and its limits. Second, if the deployment can be sequenced, hold it long enough to collect a real pre-AI window, since three to six months of clean baseline is worth more than years of arguing about a counterfactual. Where neither is possible, start the clock now on a clean baseline for the next phase and document that the current rollout has no before-number.

How do we collect demographic data for fairness monitoring without creating the bias we are trying to detect? By separating collection from decision-making at every level. Ask for it through transparent disclosure explaining that it is used to monitor fairness, make it genuinely optional, hold it apart from the working candidate record that recruiters and hiring managers see, and use it only in aggregate. Under the GDPR much of this is special-category data with heightened restrictions, which is another reason the safe posture is to keep it out of the hiring record entirely. The analysis then runs on anonymized or pseudonymized data, so a dashboard shows group-level selection rates and never a named person's protected characteristics.

How long should we keep candidate data? Long enough to serve the purpose it was collected for and no longer, which is the GDPR storage-limitation principle rather than a fixed number. In practice that means an explicit timer per category with the reasoning written down: rejected candidates purged or anonymized after a defined window such as 12 months or the period local law permits for defending against discrimination claims, with a separate consent-based talent pool for people who opted in to future contact. The two failure modes are keeping everything forever, which maximizes both obligations and exposure, and deleting so aggressively that you cannot defend a claim or compute a longitudinal fairness trend. Decide deliberately and get the schedule reviewed by legal.

Why track post-hire outcomes at all? That is HR's data, not recruiting's. Because it is the only thing that tells you whether your hiring decisions were correct, and because a fairness program that stops at the offer letter can be badly misled. Tracking ramp-to-productivity, performance ratings, retention, promotions, and tenure by group is what lets you ask whether the people you hired are succeeding and staying. If outcomes differ systematically by group, that reveals bias either in hiring or in post-hire treatment, and the distinction matters: a group hired at parity that then leaves early is pointing at something happening after the offer, which is not a recruiting failure but is absolutely a recruiting signal. The infrastructure work is joining data you probably already hold back to the recruiting record so the two can be read together.