Training Program Design: Curriculum, Delivery, and Evaluation
Lena runs recruiting enablement at a 220-person fintech company with a talent acquisition team of 18 recruiters and sourcers. When the company rolled out an AI screening assistant, adoption stalled at roughly 30 percent within two months. The tool was not the problem. The team had never been taught what the model could and could not do, when to trust it, when to override it, or what the law required of them when they did. Lena's mandate was to fix that with a real training program, not a one-hour lunch-and-learn. What follows is the curriculum she designed, the way she delivered it, and how she proved it worked.
Start With a Needs Analysis, Not a Slide Deck
Before designing a single module, Lena spent a week understanding her starting point. A training program built on assumptions trains the wrong people on the wrong things, and the cost of that error is invisible until the program has already run and changed nothing. She surveyed all 18 team members and ran four short interviews across experience levels. The survey measured three things: current AI confidence on a one-to-five scale, the specific tasks people were already using AI for, and their biggest worries about it. The interviews existed to catch what a survey cannot, the hesitations people will say out loud but will not write down.
The results reshaped her plan. Average confidence was 2.4 out of 5, but it split sharply: four recent hires were comfortable experimenting, while six tenured recruiters were quietly avoiding the tool because they feared making a compliance mistake. The most common open-text concern was not "I don't understand AI" but "I don't know if I'm allowed to override it, and I don't know what to document when I do." That single insight told Lena her curriculum needed a heavy emphasis on judgment and compliance, not just button-clicking. A needs analysis is the difference between training people on what you assume they lack and training them on what they actually lack.
A complete needs analysis asks a third question alongside knowledge and concerns: what motivates this team. Knowledge tells you what to teach and concerns tell you what to address, but motivation tells you how to frame the program so people show up willing rather than compliant. Lena's tenured recruiters were not motivated by speed; they were motivated by not being the person who made a defensible process indefensible. Framing the program around confident, documentable decisions rather than around efficiency gains was what moved them, and it cost nothing to discover because the question was already in the survey. Ask what people want to be better at, not only what they are bad at.
Design the Curriculum in Sequenced Modules
Good curriculum moves from foundational to specific. You cannot teach someone to override an AI recommendation responsibly before they understand what the recommendation is and how it was produced, and you cannot teach documentation duties to someone who does not yet know which decisions carry them. Lena organized her program into six modules, sequenced so each one rests on the last, and she wrote the sequence down before she wrote a single slide so that the dependencies were visible and arguable.
Module 1, AI foundations: what the screening model actually does, what training data is, and why "the AI said so" is never a defensible reason for a decision. Module 2, the tool in practice: hands-on navigation of the specific assistant the team uses. Module 3, fairness and bias awareness: the four-fifths rule, how disparate impact arises, and why a model that looks neutral can still produce skewed outcomes. Module 4, the legal floor: NYC Local Law 144 obligations, EEOC and ADA considerations, and GDPR data-handling rules for candidate information. Module 5, judgment and escalation: when to trust a recommendation, when to override, and who to escalate to when something looks wrong. Module 6, documentation: what to record for every AI-assisted decision so the team can defend it later.
Notice the sequence. Compliance and fairness come after foundations and tool fluency, because rules only make sense once people understand the system the rules govern. A recruiter who does not understand that the model scores resumes against historical hiring patterns will not understand why bias monitoring matters, and will file the four-fifths rule under arbitrary policy rather than under mechanism. The same logic puts documentation last: it is the module that only makes sense once a learner has an opinion about which decisions are hard, because until then a decision log is paperwork rather than protection.
Write Learning Objectives You Can Actually Measure
A vague objective produces a vague training session and an untestable outcome. "Understand AI" cannot be assessed, because there is no observation that would settle whether a learner met it. "Explain how the screening tool generates a match score and name two situations that require a manual override" can be assessed, because a person either produces that explanation or does not. Lena wrote two or three observable objectives per module, each starting with an action verb a person can demonstrate in front of you.
For the fairness module, one objective read: "Apply the four-fifths rule to a sample selection-rate table and correctly identify whether adverse impact is present." For the legal module: "List the candidate-notification and bias-audit obligations Local Law 144 places on an employer using an automated employment decision tool." These objectives do double duty. They tell the instructor what to teach, and they hand the evaluator a ready-made test. If you cannot imagine how you would check whether someone met an objective, the objective is not written yet, and the honest response is to rewrite it rather than to promise yourself you will figure out assessment later.
Mix Delivery Methods to Match How People Learn
Different content demands different delivery. Pure lecture is efficient for transferring facts and useless for building judgment. Lena blended five methods and matched each to its purpose. Lecture carried the foundational concepts and the legal framework, where the goal is accurate information transfer. Hands-on practice carried tool navigation, because nobody learns software by watching slides. Small-group discussion carried the fairness material, because it surfaces the uncomfortable real cases people are afraid to raise alone and builds the psychological safety to admit uncertainty. Short videos covered the foundational module so new hires could self-pace before live sessions. Role-play carried the escalation module, because deciding to override an AI recommendation in front of a hiring manager is a high-judgment skill you can only build by rehearsing it.
The psychological-safety point deserves its own weight, because it is the reason the fairness module is a discussion rather than a lecture. Fairness training fails when it is delivered as a list of prohibitions to a room of people who each privately suspect they have already broken one. A small group where a recruiter can say "I have been letting the tool sort my pile and I only really look at the top ten" is a group that can be corrected. A lecture hall where the same recruiter stays silent produces a perfect attendance record and no change at all. Design the format for the confession you need to hear, not for the coverage you need to claim.
Here is the worked curriculum Lena delivered, with module hours across a four-week program totaling 12 instructional hours per recruiter.
| Week | Module | Delivery | Hours |
|---|---|---|---|
| 1 | AI foundations | Self-paced video plus live Q&A | 2 |
| 1 | The tool in practice | Hands-on lab with live candidate data in a sandbox | 2 |
| 2 | Fairness and bias awareness | Lecture plus small-group case discussion | 2.5 |
| 3 | The legal floor | Lecture on Local Law 144, EEOC, ADA and GDPR, with a short open-book quiz | 2 |
| 3 | Judgment and escalation | Role-play workshop with scripted override scenarios | 1.5 |
| 4 | Documentation | Workshop building each recruiter's decision-log template, plus a capstone | 2 |
The spread is deliberate. The two modules that carry the most legal and ethical risk, fairness and the legal floor, take 4.5 of the 12 hours between them, and the highest-judgment skill, escalation, gets live rehearsal rather than a slide. Reading the hours column as a statement of priorities is a useful discipline when you draft your own: whatever gets the most contact time is what your organization is actually saying matters, regardless of what the introduction claims. If your compliance content occupies twenty minutes at the end of a tool demo, you have designed a tool demo.
Tie Every Module to a Compliance Obligation
Training recruiters to use AI faster without teaching them to use it lawfully is a liability, not an enablement program. Lena treated compliance as a thread woven through the whole curriculum rather than a single dreaded module. Foundations explained why a model trained on past hiring can reproduce past patterns; the tool lab showed where in the interface a recruiter's override is recorded; escalation rehearsed the conversation a fairness concern actually requires. By the time learners reached the legal module, the obligations landed on prepared ground instead of arriving as trivia.
Bias awareness was anchored to the four-fifths rule, the long-standing screening test for adverse impact: if the selection rate for any protected group is less than four-fifths, or 80 percent, of the rate for the group with the highest selection rate, that is a signal of potential adverse impact worth investigating. Recruiters practiced reading selection-rate tables so they could recognize the pattern rather than treat fairness as an abstraction. The legal module covered NYC Local Law 144, which requires employers using an automated employment decision tool to conduct a bias audit and to notify candidates that such a tool is being used. It covered EEOC and ADA considerations, including the obligation to provide reasonable accommodation and to avoid tools that screen out candidates with disabilities. And it covered GDPR data-handling rules for any candidate data tied to the European Union, including the principle that candidates have rights over how their personal data is processed. The point was not to turn recruiters into lawyers. It was to make sure every person using the tool knew where the legal floor sat, so responsible use became the default rather than a happy accident.
Evaluate With the Kirkpatrick Four Levels
A program you cannot measure is a program you cannot defend or improve. Lena evaluated hers with the Kirkpatrick model, a widely used framework that measures training at four ascending levels. The levels ascend in value and descend in convenience, which is precisely why most programs report the first and quietly skip the rest.
Level 1, Reaction: did participants find the training relevant and engaging? Lena collected a short post-session survey after each module. Reaction is the weakest signal but the easiest to gather, and a session everyone hated rarely changes behavior, so it is worth collecting as a floor check rather than as evidence of success.
Level 2, Learning: did knowledge and skill actually increase? This is where her measurable objectives paid off, because a pre-and-post instrument is trivial to build once every objective is already written as an observable action. She ran a pre-assessment before the program and the same assessment after. The team's average score rose from 41 percent to 88 percent, and critically, every recruiter cleared the 80 percent threshold on the compliance items, which she had made a non-negotiable pass bar. Separating the compliance items into their own pass bar matters: a blended average lets a strong performer on tool mechanics hide a weak score on the rules, and the rules are the part with legal consequences.
Level 3, Behavior: did people change what they actually do on the job? This is the level most programs skip, and the one that matters most. Lena tracked two behaviors over the eight weeks after training: tool adoption and documentation completeness. Adoption rose from 30 percent to 82 percent of eligible screens, and the share of AI-assisted decisions with a complete decision log went from near zero to 91 percent. Behavior change, not test scores, is the real product of training. Note that both metrics were already being captured by systems the team used, which is what made Level 3 affordable; the practical trick is to choose behaviors your existing tooling can already see.
Level 4, Results: did the business outcome improve? Lena connected the program to two results her leadership cared about: average time-to-first-screen dropped meaningfully as recruiters trusted the tool for first-pass triage, and the team passed its first Local Law 144 readiness review with no documentation gaps. Level 4 is the hardest to attribute cleanly because many factors move business metrics at once, so she reported it as a contributing trend rather than a sole cause. Reporting it honestly is what preserves the credibility of the Level 2 and Level 3 numbers standing next to it.
Ask Which Delivery Method Actually Moved People
The evaluation question most programs never ask is comparative: do outcomes differ by training method? Lena's data could answer it because she had recorded which method carried which module and had per-module assessment results alongside per-behavior tracking. The modules delivered as hands-on lab and role-play produced the behavior changes she cared about; the modules delivered as lecture produced knowledge gains that showed up on the assessment. That is not an argument for abolishing lectures, since the legal floor genuinely is information that has to be transferred accurately. It is an argument for never using lecture where the objective is a judgment call.
Build the comparison into your design rather than reconstructing it afterwards. Tag every objective with the method that carries it, and when you review results, sort by method as well as by module. Over two or three cohorts a pattern emerges that tells you where to move hours: the content that people can already read for themselves, and the content that only lands when someone practices it under observation. Reallocating an hour from lecture to rehearsal is a cheaper intervention than adding an hour, and the method-level view is the only thing that tells you which hour to move.
Close the Loop and Keep the Program Alive
Evaluation data is only useful if it changes the next iteration. When Lena's Level 2 scores showed the GDPR section lagging the rest, she rebuilt it with a concrete candidate-data scenario instead of an abstract rule recap, and scores recovered in the next cohort. That is the whole point of measurement: it tells you which module to fix. The diagnostic works at the item level too, since a single question that most of a cohort misses usually points at one confusing slide rather than at a failure of the learners.
She also steered around the predictable failure modes. The first is rushing: shipping a tool before the team is trained, then spending months cleaning up the misuse that follows. Foundational work feels like a delay until you price in the cost of the mistakes it prevents. The second is treating training as a one-time event. Capability is built through repetition, refreshers, and a community of practice where recruiters bring real cases, not a single seminar that fades within weeks. The third is measuring only Level 1 reaction and declaring victory because people enjoyed the session. Enjoyment is not adoption, and adoption is not compliant adoption. The program that earns its keep is the one that can show, with numbers, that behavior on the job changed and that the change stayed inside the legal lines.
Anti-Patterns
The tool demo wearing a curriculum's clothes. This is a program whose modules are all named after screens: the dashboard, the candidate view, the summary panel. It happens because the vendor's enablement deck is free, complete, and available on Monday, while a curriculum built from your own needs analysis takes a week nobody budgeted. What goes wrong is that people learn where the buttons are and learn nothing about when to press them, so the first hard case, a strong candidate the model ranks low, is decided by whoever in the room sounds most confident. The counter is to write the judgment and documentation objectives first and let the tool walkthrough serve them, rather than the other way round.
Objectives that cannot be failed. "Understand AI," "be aware of bias," "appreciate the importance of compliance." It happens because these verbs are easy to agree on in a planning meeting and impossible to argue with, whereas a testable objective invites someone to ask whether the training will actually get people there. What goes wrong is that the program becomes unfalsifiable: no learner can be shown not to have met the objective, so no module can be shown to need rebuilding, and evaluation collapses into attendance. The counter is the imagination test, if you cannot picture the observation that would settle whether someone met the objective, rewrite it until you can.
One delivery method for everything. Usually this is lecture, because a deck scales and a role-play does not, and occasionally it is the opposite, an all-hands-on-lab program that never explains the model. It happens under scheduling pressure: a single ninety-minute slot for eighteen people is far easier to book than six differentiated sessions. What goes wrong is content-method mismatch, and it fails silently. People pass the quiz on when to override and still do not override, because the skill was never rehearsed in front of anyone. The counter is to assign a method to each objective at design time and to defend any objective about judgment that is not carried by practice or role-play.
Compliance quarantined at the end. The fairness and legal content is scheduled as the last module, or worse, as a follow-up session after go-live. It happens because compliance is the least popular content and organizers protect their attendance numbers by putting the enjoyable material first. What goes wrong is that a portion of the team is already using the tool on live candidates before they have been taught the four-fifths rule, the notice obligations, or what a decision log must contain, so the program creates exposure during the exact weeks it is supposed to be reducing it. The counter is a hard gate: no live use on candidate decisions until the fairness and legal modules are complete and passed.
Stopping at the smile sheet. The program reports a 4.6 average satisfaction score and a 100 percent completion rate and is declared a success. It happens because Level 1 data arrives in the room the session ends and Level 3 data requires waiting eight weeks and negotiating access to systems somebody else owns. What goes wrong is that the organization now believes a capability exists that does not, and it discovers otherwise through an incident rather than through a report. The counter is to pick the Level 3 behaviors during design, confirm at that point that existing systems can see them, and refuse to publish a result that contains only reaction data.
Training delivered long before the duty applies. The whole team is trained in March on a tool that reaches half of them in September. It happens when training is scheduled around trainer availability rather than around the moment of use. What goes wrong is straightforward decay: by the time the duty arrives, the escalation path and the documentation format are gone, and people improvise. The counter is to attach the high-consequence modules to the moment of relevance, at provisioning and at annual refresh, and to accept running the same session more than once as the cost of it working.
Practice
- Run the three-question needs analysis on your own team. Survey everyone on current AI confidence, the tasks they already use AI for, and their biggest worry, then add the motivation question: what would they like to be better at. Interview three or four people across tenure levels to catch what the survey misses. Compare the result with the curriculum you were about to build and note every module the data does not justify.
- Sequence your modules and defend the order. List the modules you intend to run, then for each one write the prerequisite it depends on. Any module whose prerequisite comes later in your sequence is misplaced. Pay particular attention to where fairness, legal and documentation sit relative to first live use.
- Rewrite five weak objectives. Take five objectives from an existing program, or five you have just drafted, and rewrite each to start with an action verb and name an observable performance. For each rewrite, sketch the test item you would use. Discard any objective for which you cannot write the test.
- Assign a delivery method to every objective. Build a two-column list of objective and method, then find every judgment-type objective assigned to lecture and move it to practice or role-play. Total the hours by method afterwards and check whether the split matches the risk profile of the content.
- Build the pre-and-post assessment before the first session. Draw the items directly from your objectives, separate the compliance items into their own block, and set an explicit pass bar for that block. Running it as a pre-assessment gives you a baseline; running the same instrument afterwards gives you Level 2 with no extra design work.
- Choose two Level 3 behaviors your systems can already see. Name the on-the-job behaviors that would prove the training landed, then confirm today that a system already records them and that you can get the report. If you cannot see a behavior, either instrument it before launch or pick a different one, because a Level 3 plan that depends on future access will not survive the quarter.
Reflection
- If your team's real gap is judgment rather than knowledge, would your current program detect that, or would it train tool mechanics either way?
- Which of your modules could a learner pass without being able to do anything differently on Monday?
- Where in your calendar does a recruiter first touch a live candidate with an AI tool, and which modules are complete by that date?
- What is the last piece of evaluation data that actually changed a module you run, and how long ago was that?
Glossary
- Needs analysis. The survey-and-interview pass that establishes current knowledge, active concerns, and motivation before any module is designed. It is what separates training people on their real gap from training them on your assumed one.
- Module sequencing. Ordering content so each module rests on the one before, foundational to specific. Rules taught before the mechanism they govern are memorized as arbitrary policy rather than understood.
- Learning objective. A statement of what a learner will be able to do, written with an action verb and an observable performance. If no observation would settle whether it was met, it is not yet an objective.
- Blended delivery. Matching method to content: lecture for accurate information transfer, hands-on lab for tool skill, small-group discussion for fairness, video for self-paced foundations, role-play for high-judgment skills.
- Psychological safety. The condition in which a learner will admit uncertainty or disclose current practice without fear of penalty. It is why fairness content is designed as small-group discussion rather than as a lecture.
- Role-play rehearsal. Practising a high-judgment act, such as overriding a recommendation in front of a hiring manager, under observation. Judgment skills are not reliably transferred by explanation alone.
- Kirkpatrick four levels. The evaluation framework measuring Reaction, Learning, Behavior and Results. Value ascends with the level and convenience descends, which is why most programs report only the first.
- Pre-and-post assessment. The same instrument run before and after the program, drawn from the objectives, producing the Level 2 learning measure with a defensible baseline.
- Compliance pass bar. A separate minimum score on the compliance items, held apart from the blended average so strong tool mechanics cannot mask weak knowledge of the rules.
- Level 3 behavior measure. An observed change in on-the-job conduct after training, such as adoption rate or decision-log completeness. Choose behaviors existing systems already record, or the measurement will not happen.
- Method-outcome comparison. Sorting evaluation results by delivery method rather than only by module, so you learn which hours to move from lecture into practice.
- Four-fifths rule. The adverse-impact screening test: a selection rate for any protected group below four-fifths, or 80 percent, of the highest group's rate is a signal of potential adverse impact worth investigating.
- Automated employment decision tool. The category of system that computationally screens or scores candidates, and the category to which Local Law 144 attaches bias-audit and candidate-notification obligations.
- Decision log. The record of an AI-assisted decision, including the recommendation, the human action taken, and the reasoning, produced so the decision can be explained later.
Related Lessons
- Training and Capability Building: Ensuring Staff Can Use AI Responsibly sets the wider capability agenda that this lesson turns into a specific curriculum, delivery plan, and evaluation design.
- Adult Learning Principles: How Recruiters Learn and Adopt New Tools explains why the delivery methods in this lesson work, and why relevance and immediate application matter more for experienced professionals than coverage does.
- Coaching and Support: Helping Individuals Build Confidence covers the between-session reinforcement that keeps a training spike from decaying before it becomes behavior.
- Measuring Adoption: Tracking Usage, Proficiency, and Impact develops the Level 3 and Level 4 measurement problem in more depth, including what to instrument and when.
- Creating Communities of Practice: Learning Networks is the structure that carries real cases after the formal program ends, and the answer to training as a one-time event.
- Hands-On Project: Design a Team Capability Building Program is where this curriculum design becomes one component of a full assessment, coaching, and measurement system.
Closing
Lena's program did not succeed because the content was novel. Everything in it is available in some form to anyone who looks. It succeeded because the design decisions were made in the right order: find out what the team actually lacks, sequence the modules so each rests on the last, write objectives specific enough to test, match every method to the content it carries, thread compliance through rather than bolting it on, and measure at the level where behavior lives instead of the level where satisfaction lives.
The parts that will be tempting to cut are the parts doing the work. The needs analysis looks like a week of delay. The role-play looks like an indulgence next to a deck that covers the same material in fifteen minutes. The Level 3 measurement looks like a reporting chore two months after everyone has moved on. Cut those three and you still have a training program; you just no longer have any way to know whether it did anything, and no reason to expect that it did.
Key Takeaways
- Start with a needs analysis, not a curriculum. Survey and interview your team on knowledge, concerns and motivation before you design anything. The gap you assume is rarely the gap that exists, and Lena's team needed judgment and compliance confidence far more than another tool demo.
- Sequence modules from foundational to specific. Teach what the AI does before you teach the rules that govern it. Compliance and fairness only make sense once people understand the system those rules are protecting against, and documentation only makes sense once a learner knows which decisions are hard.
- Write objectives you can test. "Understand AI" cannot be assessed; "apply the four-fifths rule to a selection-rate table" can. A measurable objective tells the instructor what to teach and hands the evaluator a ready-made test.
- Match delivery method to content. Lecture for facts and law, hands-on labs for tools, small-group discussion for fairness where psychological safety is the point, video for self-paced foundations, and role-play for high-judgment skills like overriding a recommendation.
- Weave compliance through the whole program. Bias awareness and the four-fifths rule, NYC Local Law 144 audit and notification duties, EEOC and ADA obligations including reasonable accommodation, and GDPR data handling are not a footnote. They are the difference between enablement and liability.
- Evaluate with all four Kirkpatrick levels. Reaction, Learning, Behavior, and Results. Most programs stop at whether people enjoyed it. The value lives at Level 3 behavior change and Level 4 business results, measured with pre and post data, and Level 4 should be reported as a contributing trend rather than a sole cause.
- Compare outcomes by delivery method, not only by module. Tag each objective with its method and sort your results both ways. That is the only view that tells you which hour to move from lecture into rehearsal.
- Close the loop. Use the evaluation data to fix the weak module, then run it again. Training is a continuous program with refreshers and a community of practice, not a one-time event that fades within a month.
Frequently Asked Questions
Twelve hours per recruiter is more than my organization will approve. What do I cut? Cut coverage, not the gate. The modules that can shrink are the ones whose content a competent adult can absorb from a document, typically the foundations material, which moves entirely to self-paced video with a short live Q&A for questions. What must not shrink is the fairness and legal block or the practice that turns judgment objectives into behavior, because those are the hours that carry legal consequence and the hours a reader cannot substitute for on their own. If the total has to come down, take it out of lecture time and protect the rehearsal and the compliance pass bar. A shorter program that still gates live candidate use on completed fairness training is defensible; a full-length program that lets people screen candidates before that module is not.
How do I evaluate at Level 3 when I do not own the systems that record behavior? Solve this during design rather than after delivery. Pick candidate behaviors, name the system that would show each one, and go ask for the report before you build the program, because the answer changes which behaviors you choose. Lena's two measures, adoption rate on eligible screens and decision-log completeness, worked partly because both were already visible in tooling her function used. Where nothing records the behavior you want, the fallback is a structured manager observation or a periodic sample review of AI-assisted decisions, which is more effort per data point but still gives you evidence about conduct rather than about opinion. What does not work is promising Level 3 and delivering Level 1 with a confident narrative attached.
Should the training differ for people who are already using AI heavily? Yes, but not in the direction most people assume. Confident users usually need less time on foundations and tool mechanics and more on the guardrails, because informal expertise is built by experimentation and experimentation does not teach anyone the four-fifths rule or what a notice obligation requires. The pattern in Lena's needs-analysis data is instructive: the split was not simply confident versus unconfident, it was recent hires experimenting freely alongside tenured recruiters avoiding the tool out of compliance fear. Those two groups need different framing and different reassurance, but both sit the same compliance block and clear the same pass bar, because the legal floor does not vary by how comfortable someone feels.
Skill.re