Redesigning Roles Around AI Without Breaking Trust
It is the annual review, and the manager has the form open on her screen. The first field is the job description, pulled automatically from the human resources system, and it reads: "Accurately code and enter supplier invoice data into the accounts payable system. Maintain a throughput standard of 180 documents per day." The person across from her has not hand-coded an invoice in seven months. Since the exception workflow went live, her day is verifying what the model proposes, arguing with the hard cases it cannot resolve, calling vendors, and telling the improvement group which proposals keep coming back wrong. She is better at her job than ever, and every artifact the company holds about that job describes work that no longer exists. The manager reads the field out loud, hears how absurd it sounds, and says the sentence that ends the meeting's usefulness: "Well, we all know that's out of date." They both laugh. Neither laughs a month later, when the promotion list comes out and the competency framework has no language for anything she does.
The Quiet Incoherence
You have spent this chapter building the 70 percent. BCG's 10-20-70 rule holds that roughly 10 percent of AI value comes from algorithms, 20 percent from technology and data, and 70 percent from people and process, and you have funded that 70 percent as a real workstream: a change plan with a budget, training that targets behavior, incentives that pay for the new work, a champion network carrying the load locally. One thing remains, and almost every program skips it.
The workflow redesign changed what the role does. Nobody updated the artifacts that say what the role is.
McKinsey's State of AI work keeps landing on the same finding: the roughly 6 percent of organizations reporting material impact are about three times more likely to have fundamentally redesigned workflows, and workflow redesign is among the strongest drivers of value. But a workflow is not only a sequence of steps in a swimlane diagram. It is those steps plus the people who perform them, and a person's relationship to their work runs through a stack of documents: job description, competency framework, performance metrics, pay band, promotion criteria, the hiring profile used for the next vacancy. Redesign the steps, leave that stack untouched, and you have done half a redesign. The half you skipped is the half employees can actually see.
What the skipped half produces is not a dramatic failure. It is a quiet incoherence. A clerk whose day is now 70 percent verification and exception judgment holds a job description about data entry. She is measured on a throughput key performance indicator (KPI) built for a task the model now performs, so her numbers look worse the more the program succeeds. She sits in a pay band benchmarked to the old work. She has learned three genuinely new skills and has no idea whether any of them counts, because no document in the building says so. Nothing here is exactly a lie, but out of date at this scale is functionally the same as untrue.
That incoherence is where trust dies. Not in the technology, which most people find fine to slightly useful, and not in the training, which most people sit through politely. Trust dies in the gap between what someone does all day and what their employer's official record says they do, because that record is what promotion committees read, what compensation reviews use, what headcount planning models, and what the next hiring manager recruits for. A person invisible in all four has correctly concluded that their new work does not count. MIT's autopsy of the 95 percent of enterprise generative AI pilots that produce no measurable profit-and-loss return names adoption without transformation as a core cause: the tool is used, the surrounding organization is unchanged. Roles are part of that organization.
If you change what a role does and not what the organization says the role is, you have not redesigned a job. You have created a job that exists only in practice, and practice-only jobs are invisible in every process that matters.
Role redesign closes the loop the workflow redesign opened. It has four moves and one artifact that carries them: the Role Change Brief, one per affected role. Its contents are the four moves in order: before and after task composition with time shares; responsibilities added and removed, with standards; skills now required; what this means for level, band, and path; what is decided and what is not; and the manager's conversation guide, because the last mile of role redesign is one person talking to another, and that goes badly by default.
Move One: Document the Task Shift Honestly, With Time Shares
Start with numbers from your instrumentation rather than your assumptions. By this stage you have the data: workflow telemetry, queue timings, case-handling logs, the sampling built for gate reviews. Turn it into a task composition: what percentage of a normal week goes to which category of work, before and after.
Do not run this as a survey of what people think they do. Self-reported time allocation is wrong in a predictable direction: people over-report work they consider skilled. Use the instrumentation, then show the draft to two people in the role and let them correct it. That step is not politeness: they will catch a category you missed, usually the one nobody logs, such as chasing a vendor for a purchase order number.
Here is the shape it takes for the invoice exceptions clerk in our running program, where the invoice-exception and vendor-inquiry workflows are now live. All figures are illustrative.
| Task category | Before (share of week) | After (share of week) |
|---|---|---|
| Read and categorize incoming exceptions | 55% | 0%, now proposed by the model |
| Data entry into the AP system | 20% | 0%, now written back automatically |
| Verification of AI proposals | 0% | 15% |
| Exception judgment on the hard cases | 15% | 40% |
| Vendor communication | 0%, handled by a separate desk | 15% |
| Queries and escalations | 10% | 20% |
| Improvement input to the program | 0% | 10% |
Now read what the table says, because this is the part most programs never say out loud. Routine content fell from 75 percent of the week to zero. Judgment content rose from 25 percent to 90 percent. That is not a downgrade dressed up as an opportunity. It is a measurable upskilling of the actual day, and the numbers make the case better than reassurance ever will. "An exciting opportunity to focus on higher-value work" is a sentence people have learned to distrust on contact. "Your judgment work went from 15 percent of your week to 40, and here is the queue data" is a sentence they can check.
Which brings the hard obligation. Where the shift is not favorable, show the numbers anyway. Some redesigns narrow a role. Some make it more monitored, because verification produces a per-item audit trail where the old batch work produced none. Publishing a table that pretends otherwise costs you everything the honest tables earned. Say it plainly, then attach a mitigation: a rotation into an adjacent queue, a share of the improvement work, an explicit path out of the role within a stated period, a redeployment conversation started early. An honest bad number with a mitigation is survivable. A dishonest good number is not, because the people in the role have the same instrumentation you do, in the form of their own Tuesday.
Move Two: Name and Value the New Responsibilities
The redesign created responsibilities nobody wrote down. In most programs they sit in accountability limbo: everyone relies on them, no document mentions them. List them explicitly and attach a standard to each, because a responsibility without a standard is a wish.
- Verification accountability. Someone signs off on the model's proposal, and that sign-off carries weight. It is the human gate you designed, with entry criteria and a review standard, not a courtesy click, and the person is accountable for what they approved, which means they need time, information, and the authority to reject. The gate specification's review standard is the job standard, and copying it verbatim into the role description is the cheapest coherence you will ever buy.
- Override judgment and reason-coding. The reviewer decides when the model is wrong and records why, using your reason codes. That coded reason is the program's highest-value learning input, and it is a skill: distinguishing "the model misread the document" from "the model read it correctly and the policy is ambiguous" is analytical work. Standard: override rate within the expected band, reason codes complete on every override.
- Escalation decisions. Knowing which cases must leave the desk, and how fast, is judgment under uncertainty with a cost attached in both directions. Standard: no unescalated case above the defined threshold.
- Drift noticing. The people closest to the queue see degradation weeks before a dashboard flags it, because they feel the texture of the wrongness change. Naming this as a responsibility, with a route to report it, converts an anecdote into a control. Standard: drift observations reported through the defined channel.
- Improvement contribution. Attending the review, bringing cases, proposing prompt or rule changes. Standard: contributes at a defined cadence.
- Coaching, for some roles. The experienced reviewer who teaches new joiners what a good verification looks like is doing training delivery. Say so in the role description rather than absorbing it as goodwill.
Now the question you cannot dodge and must not answer casually: does any of this change pay or level? Improvise an answer in a corridor and you have created a promise the organization did not make. Say "not my area" and you have told the person their new work is not worth a phone call.
The correct move is to route the question properly, with evidence. Take the task composition table, the responsibilities with standards, and the skills list to your human resources and compensation partners as a formal role review request. That is what those functions are for, and they can rarely act without exactly this evidence, which is why it so seldom happens: the evidence has never existed. Then be honest about the range of outcomes. The most common honest answer is: the band is unchanged for now, verification and judgment are added to the promotion criteria for the senior position in this job family and to the next hiring profile, and here is the path and the date of the next review. That answer disappoints nobody who was told it in advance. It infuriates everyone who was allowed to assume otherwise. Frame the review to HR as a retention question rather than a fairness plea: a seat that moved from 75 percent routine to 90 percent judgment is a scarcer skill to replace, whatever the band says this year.
Move Three: Update the Artifacts That Encode the Role
A role is not what happens. A role is what the organization's systems believe happens. Six artifacts encode that belief, and each has a specific failure mode when it goes stale.
| Artifact | What to change | Cost of leaving it stale |
|---|---|---|
| Job description | Rewrite accountabilities from the after-state task composition; copy the gate spec's review standard verbatim | The person's real work is invisible in every formal review of performance and potential |
| Competency framework | Add verification judgment, override reasoning, escalation decision-making, and improvement contribution as named, leveled competencies | Promotion panels have no language to credit the new skills, so they credit the old ones |
| Performance metrics | Retire the throughput measure of the automated task; measure verification quality, override accuracy, escalation timeliness, contribution | The metric rewards a task the model now performs, so success on paper requires resisting the program |
| Hiring profile | Recruit for judgment, document literacy, and comfort supervising machine output rather than keying speed | New hires arrive unable to do the job, and the profile broadcasts what the org still thinks the role is |
| Training curriculum tier | Move the role to the tier matching its new work, with verification and override reasoning as core modules | The role trains for the job it used to have |
| Career path map | Show where the redesigned role now leads, and what the next step requires | The role looks like a dead end, and the best people leave first |
Two rows deserve extra attention. Performance metrics formalizes the metric conflict you have already met: when the scorecard still counts the volume the AI now handles, you have built an incentive to work against your own program, and no communication survives a scorecard that punishes compliance. Fixing the metric in the role artifact makes the fix permanent rather than a local favor from one sympathetic manager.
The hiring profile is the most durable signal in the set and the most underrated, because everything else can be waved away as paperwork. What the organization recruits for is a public, budgeted statement of what the role has become, and everyone in the team reads the posting. When the next vacancy is advertised as "invoice exceptions specialist: reviews AI-proposed codings, exercises judgment on complex exceptions, manages vendor resolution," the incumbents learn more from that posting than from any town hall. When it is advertised as "data entry clerk, 180 documents per day," they learn something too: that the last eighteen months were not real.
Practice-only roles are invisible in every process that matters: promotion, compensation review, workforce planning, succession, headcount modeling. The person doing the hardest version of the job in the building can be made redundant by a spreadsheet that still calls them a keyer.
Move Four: The Conversation, Delivered by the Manager
The brief is not the change. The conversation is the change, and it must be delivered by the person's own manager, not by the program, not by HR, not by a recorded video from a sponsor. A message about someone's job carries the authority of whoever controls that job. A program-delivered message about your role is information. A manager-delivered message about your role is a commitment.
Do not script the words: scripted words are audible, and a person reading a script about your future is worse than silence. Teach the structure instead.
- Open with the specific task change and its numbers. Not "things are changing." "Your categorization work was 55 percent of your week and the tool handles it now. Your exception judgment was 15 percent and it is around 40. Here is the queue data." Specificity is the entire opening, because the person already knows the change happened and is waiting to find out whether you do.
- Name what the person is now accountable for. Read the added responsibilities and their standards. Being told precisely what you are on the hook for is respectful; being left to infer it is not. This is the moment the new work stops being extra and becomes the job.
- State what is decided and what is not, with dates. "The band is unchanged this year. Verification and override judgment go into the senior clerk promotion criteria and the next hiring profile. What we have not decided is what wave 3 means for headcount here, and I am not going to pretend otherwise. That decision sits in the November planning cycle and I will tell you the day I know." An unknown named with a date beats vague comfort, because vague comfort is what people were given before every bad surprise of their working lives, and they can smell it.
- Acknowledge the loss where there is one. In our people audit, craft pride came up in 15 of 43 interviews: veterans who were measurably fast and accurate at the categorization work and have just watched that skill made irrelevant by software. Saying "your speed and accuracy were real skill, it took years to build, and the tool taking it does not make it retroactively worthless" costs nothing and is the single highest-trust sentence available in this entire program. Then connect it forward: the new craft is the judgment work, and you already know what right looks like.
- End with their question, not your talking point. "What is the thing you most want an answer to?" Then answer it, or write it down and commit to a date. A conversation that ends with the manager's summary is a briefing.
Managers need three things from you: the Role Change Brief, a rehearsal (twenty minutes, you play the employee and ask the hardest question), and a written answer to that hardest question. It is almost always "is my job going away?" and every manager in the function must answer it the same way, agreed in advance. Two managers giving different answers in the same week does more damage than a month of silence.
The Cases That Need Distinct Handling
Three groups are not served by the standard brief, and pretending otherwise is how programs get a reputation for treating people as a category.
The veteran near retirement with no interest in retraining. Twenty-eight years in, four to go, and a polite refusal to become a verification specialist. This is not resistance, and treating it as a change-management problem is an insult. It is a rational allocation of a person's remaining working life. The honest response is an accommodation conversation with real options: a role weighted toward residual manual and vendor-facing work if that work exists, a move to an adjacent function outside the wave, a phased arrangement, or an early transition package. What you must not do is quietly place them on the training list, watch them fail the certification, and let that failure become the documented reason for a decision made on other grounds.
The high performer whose advantage was speed at the automated task. The most under-discussed transition in AI workforce change. Someone who was top of the team, whose identity at work was being fastest and most accurate, is now average, because the thing they were best at is the thing the machine does. They rarely complain. They go quiet, engagement drops, and eight months later they resign to a competitor with an older process. They need coaching, named as such: several weeks of structured support, an acknowledgment that the ground moved under them and it is not their failing, and deliberate work to find where their underlying strength applies now.
The person whose role genuinely shrinks. Sometimes the honest answer is that there is not a full job left. Say so early and run it properly: redeployment first, with real support and a real search, and if the truth is a reduction, a transition managed through HR with proper notice, outplacement, and the dignity of not being the last to know. Every instinct in a program pushes toward delay, because delay feels kind and is cheap this quarter. It is neither. Your integrity here is what makes your credibility usable everywhere else, because the workforce is watching how the hardest case is handled and calibrating how much to believe you in the easier ones.
The Brief in Full, and the Rules Around It
The artifact assembled for one role. Every number is hypothetical; the shape is what you copy.
Role Change Brief: Invoice Exceptions Clerk, Wave 1
1. Scope. 14 people in accounts payable (AP) operations, three sites, effective at wave 1 go-live. Author: program lead with the AP operations manager. Reviewed by: HR business partner, finance controller.
2. Task composition. Before: 55 percent read and categorize, 20 percent data entry, 15 percent exception judgment, 10 percent queries. After: 15 percent verification of AI proposals, 40 percent exception judgment on hard cases, 15 percent vendor communication, 20 percent queries and escalations, 10 percent improvement input. Source: queue instrumentation, weeks 1 to 6 of pilot, plus two incumbent reviews. Net: routine content 75 to 0 percent, judgment content 25 to 90 percent.
3. Responsibilities added, with standards. Verification accountability (gate specification section 4, review standard, verbatim). Override judgment with reason-coding (override rate inside the 8 to 18 percent band, reason code on 100 percent of overrides). Escalation decisions (no case above the value threshold unescalated beyond one business day). Improvement contribution (fortnightly review, two cases per cycle). Removed: manual categorization of routine exceptions; manual keying into the AP system.
4. Skills now required. Reading a model-proposed coding against source documents and naming the specific point of disagreement. Applying policy judgment where documents are ambiguous. Vendor conversation handling. Structured reason-coding. Recognizing pattern change in model output.
5. Level, band, and path. Band unchanged at wave 1. Verification and judgment competencies added to the promotion criteria for Senior Invoice Specialist this cycle, and to the hiring profile for the next vacancy. A formal role review with compensation is scheduled with HR at the wave 2 boundary, using this brief as the evidence pack. Training tier moves from 3 to 2.
6. Decided and not decided. Decided: no headcount reduction in waves 1 and 2; the role continues; the metric change takes effect at go-live. Not decided: what wave 3 automation of the vendor-inquiry queue means for headcount in this function. That question is genuinely open, sits with the finance and operations planning cycle, and has a November decision date. Every manager states this in the same words, and the date is real.
7. Manager conversation guide. The five-step structure, the rehearsal booking, and the written answer to "is my job going away?" agreed across all three site managers before any conversation happens.
Three Conversations, Three Legitimate Outcomes
The craft-pride veteran. Nineteen years, one of the fastest categorizers in the function, one of the 15 interviews where craft pride surfaced. Her manager opens with the numbers, then says her speed was real skill that took years to build. She is quiet a while, then asks the only question she cares about: who decides when the tool is wrong? The answer is: you do. Six months later she is the site's verification authority and a training champion. Her craft did not die. It moved up a level, and someone told her so in week one.
The speed-advantage high performer. Three years, consistently top of the throughput board, visibly flat since go-live. He says the thing people usually do not say: "I was the best at this and now I am average." His manager does not argue. He gets six weeks of structured coaching with a named coach, practice on ambiguous-policy cases, and honest feedback. He does not become the strongest judgment reviewer. He becomes the team's strongest vendor-resolution handler, because the decisive temperament that made him quick at coding makes him excellent with a supplier who has waited a month. The strength was real. The application changed.
The lateral move. One clerk works through the brief, thinks for a month, and decides the new job is not the job she wants. She asks about the procurement analyst opening. The program supports the move, HR runs it as an internal transfer, and the brief's skills list is excellent evidence for her application. Log this as a legitimate outcome, not attrition and not a failure. People are allowed to conclude that a redesigned role is not for them, and a program that treats every such move as a defeat will start hiding them.
The Three Program-Level Trust Mechanics
Three rules govern how briefs are used across a program. Each is cheap to follow, expensive to violate.
The sequencing rule: briefs land before go-live. A change explained in advance is a plan. The identical change explained afterward is a fait accompli, and people react to the two completely differently even when the words are the same. Two weeks after is an announcement that something was done to you, and no amount of good content recovers the difference. Put "role briefs delivered" in the go-live checklist as a gating item, not a communications task.
The consistency rule: the brief must match the systems. What the brief says the role is must match what the metrics reward, what the training teaches, and what the hiring profile seeks. If the brief celebrates judgment and the scorecard still counts documents processed, the workforce believes the scorecard, because the scorecard pays. A brief inconsistent with the systems around it is marketing, and people classify it correctly within a week.
Reversibility humility. Some of your role designs will be wrong: over-assigned verification that creates a bottleneck, under-assigned vendor work that orphans a queue, an override band that is nonsense in practice. Say so publicly and adjust. A program that revises a brief in month four with an honest note about what it got wrong buys more credibility than one quietly right first time, because it proves the briefs are living descriptions rather than pronouncements. Give each brief a ninety-day review date and honor it.
The Failure Story: The Silent Redesign
A retail bank rolls AI-assisted processing into a 60-person operations team handling payment investigations. The workflow redesign is good: the model drafts the case narrative and proposes a disposition, the investigator verifies and decides, cycle time falls by a third, quality holds. The numbers are illustrative; the pattern is not.
Nothing else changes. Job descriptions still describe the pre-AI task. The scorecard still counts cases closed per day, a volume the model now largely produces, so the metric measures the machine and credits the human. Eighteen months in, the bill arrives.
Two people are passed over for promotion. Both had become excellent at the new work, verification quality and disposition judgment, and none of it appears in the competency framework, so the panel has no vocabulary to credit it and defaults to the metrics it can read. The next two hires are recruited against the old profile, selected for keying speed, and arrive unable to do the job; one leaves inside a year, and the team quietly concludes that management does not know what the team does. Engagement scores drop eleven points, management attributes it to workload, funds a wellbeing initiative, and misses entirely. When the operations lead finally works it out, she summarizes it in one sentence that belongs on the wall of every AI program office: "We changed everyone's job and told no one, including HR."
Notice what did not fail. The technology worked. The workflow redesign was excellent, and on process metrics the project was a success anyone would put on a slide. The role redesign never happened, and the cost landed where the scorecard could not see it: on trust, and specifically on the program's next initiative, which arrived to a team that had learned exactly what happens when you cooperate with one of these.
That closes the chapter. The change plan is funded and owned, the training targets behavior, the incentives pay for the new work, the champions carry it locally, and the roles are documented, valued, and explained. The 70 percent is a running workstream with artifacts. Which raises the question the next chapter answers: with the program operating, how do you prove to a C-suite that it is producing value, controlling risk, and governing itself, in numbers a chief financial officer will accept?
What to Do Monday Morning
Five steps, one week, one role. Do not attempt the whole workforce.
- Pull the task-composition numbers from your instrumentation rather than guessing. Queue logs, handling times, override rates, sampling data. Build the before and after table for one role with a source line under it. If you cannot source a number, mark it as an estimate rather than laundering it into a fact.
- Write one Role Change Brief for your most-affected role. Two pages maximum, the seven parts in order, shown to two people who hold the role before anyone senior sees it. Their corrections are the quality control.
- Route the level and compensation question to HR with the evidence. Send the task composition, responsibilities with standards, and skills list as a formal role review request with a requested response date. Then stop answering the question in corridors, and tell managers to stop too.
- Update the hiring profile for the next vacancy in that role. The smallest task on the list and the loudest signal in the building. Do it before the vacancy exists, so the profile is ready rather than improvised under time pressure.
- Rehearse the hardest question with the manager who has to deliver it. Twenty minutes. You play the employee, ask "is my job going away?" and keep asking until the manager has an answer they can say in their own words without flinching. Then write it down and give the same words to every other manager in the function.
Key Takeaways
- Close the loop the workflow redesign opened: changing what a role does without changing what the organization says the role is creates a job that exists only in practice, invisible to promotion, compensation, and headcount planning.
- Document the task shift with time shares drawn from instrumentation rather than self-report, and publish the numbers even when unfavorable, attaching a mitigation instead of spin.
- Name the responsibilities the redesign created (verification accountability, override judgment and reason-coding, escalation, drift noticing, improvement contribution, coaching) and attach a standard to each, copying the gate specification's review standard as the job standard.
- Route the pay and level question to HR with the evidence pack rather than answering in a corridor, and be honest that "band unchanged, skills now promotable, here is the path and the date" is a legitimate answer.
- Update all six artifacts that encode the role (job description, competency framework, performance metrics, hiring profile, training tier, career path), treating the hiring profile as the most durable public signal of what the role has become.
- Let the manager deliver the conversation using structure rather than a script: specific numbers, named accountabilities, decided versus undecided with dates, acknowledgment of the loss, and the employee's own question at the end.
- Handle the three hard cases individually: honest accommodation for the veteran declining retraining, real coaching for the speed-advantage high performer, and a properly run transition where a role genuinely shrinks, because credibility everywhere depends on the hardest case.
- Enforce the three program rules: briefs land before go-live, the brief must match what the metrics reward and the hiring profile seeks, and wrong role designs get revised publicly at a ninety-day review.
Skill.re