From Scorecard to Pilot Charter
The steering committee says yes at 3:42 on a Thursday afternoon. There is a little laughter, a joke about robots, a sponsor who claps once and says "great work, let's move." You walk out of the room holding the most dangerous object in corporate life: an approval with nothing attached to it. Six months from now, that yes will mean whatever the loudest person in the room needs it to mean. The scope will be whatever it has quietly grown into. Success will be whatever the surviving slide says it always was. And if the pilot disappoints, there will be no way to end it, because nobody ever wrote down what dead looks like. This lesson, the finale of Level 2, is about the one document that prevents all of that: the Pilot Charter. Approval is a mood. A charter is a contract. Your job, in the week after the yes, is to convert the first into the second while everyone in the building is still honest.
The Most Dangerous Day of a Pilot's Life
Here is a pattern you can verify against any stalled pilot you have ever watched. The approval meeting goes well. Then nothing formal happens. Work starts on enthusiasm and hallway agreements. By week 2, someone important asks "while we're at it, could it also handle the free-text ones?" and, because no scope was ever written, saying yes costs nothing visible. By month 4, the early numbers are soft, and the definition of success shifts from cycle time to "learning value," a phrase that has never once appeared on a profit-and-loss statement. By month 9, the pilot cannot be killed, because killing it would require someone to state the criteria it failed, and there are none. It joins the 95 percent of enterprise generative AI pilots that MIT found deliver no measurable return, and it does so not because the readiness work was wrong but because the moment of maximum honesty was allowed to pass unsigned.
That moment matters more than most people realize, and the reason is not administrative. It is psychological. On the day of approval, nobody has sunk anything yet. The sponsor has spent no political capital defending results. The team has invested no months it needs to justify. The vendor invoice has not landed. On that day, and almost never again, every person involved can look at a sentence like "if week-6 verified accuracy is below 85 percent, the pilot reverts to manual" and sign it calmly, because it is still hypothetical. Ninety days later, the same sentence reads as an accusation. Sunk cost does not merely bias judgment; it re-writes memory, so that people sincerely believe success was always defined the way they now need it to be defined. The entire Pilot Charter is an instrument for making promises to your future, more biased self, and getting them witnessed.
The stakes of skipping it have been measured at industry scale. Gartner projects that over 40 percent of agentic AI projects will be canceled by the end of 2027, and the texture of those cancellations is the important part: they are overwhelmingly late, expensive, and acrimonious, executives pulling plugs out of exhaustion after budgets are spent and reputations are entangled. Compare that to the alternative this program has repeated since Level 1: a documented kill is a win. The difference between an acrimonious month-14 cancellation and a dignified week-6 kill is not the technology and not the team. It is whether the ending was pre-committed on paper before anyone had a reason to fight about it.
You have spent this chapter digging the well. You inventoried processes, scored candidates, wrote the one-page readiness report, and presented the findings without making enemies. The scorecard picked the target; the committee said yes. This lesson is stage five of the goldmine arc: the decision, with evidence, bound into a document that survives contact with month 4.
Approval is a mood; a charter is a contract. Sign the contract while everyone is still honest, because in ninety days nobody will be.
The Seven Clauses of the Pilot Charter
The Pilot Charter is a two-to-four page document, signed and dated, with seven clauses. Not seventeen. The point is not legal completeness; the point is that every clause closes one specific escape hatch through which approved pilots drift into the 95 percent. Take them slowly, because each one is a discipline, not a paragraph.
Clause 1: Scope, drawn in the negative too
Everyone writes an IN list. Almost nobody writes an OUT list, and the OUT list is the clause that actually works. Copyable language, using our running example: "IN scope: the three structured exception types (missing PO, price mismatch, quantity mismatch), representing 78 percent of monthly exception volume. OUT of scope: free-text exceptions, any field containing bank details, anything cross-border. Additions to scope require a written change request approved by the sponsor and trigger Clause 7."
Why the negative? Because scope creep never announces itself as scope creep. It arrives in week 6 as a friendly question from someone senior: "could it also look at the cross-border ones?" If your charter only has an IN list, that question is a judgment call, and judgment calls made under enthusiasm always land the same way. If your charter has an OUT list, the same question is instantly visible as a change request against a signed document. You have not forbidden anything; you have made drift cost a signature. The OUT list is a fence, and like all good fences it works not by being unclimbable but by making the climb deliberate and recorded.
Clause 2: Baseline, incorporated by reference
You built the baseline pack earlier in this level: volume, cycle time distribution, error and rework rates, cost per unit, every figure verified and versioned. The charter does not restate those numbers; it incorporates them the way a contract incorporates an exhibit: "The pre-pilot baseline is Baseline Pack v1.3, dated and attached as Exhibit A. All improvement claims will be computed against Exhibit A. No post-launch reconstruction of baseline figures will be accepted."
This clause is the before-photo formally entered into the record, and its timing is the whole trick: the baseline enters the contract BEFORE the tool arrives, while the numbers are still innocent. "Baseline or it didn't happen" has been this program's refrain since Level 1; Clause 2 is where the refrain becomes contract language. Remember the autopsy pattern from the MIT finding: most of the 95 percent could not prove value because nobody froze the before-state. A pilot whose baseline lives inside its charter can never suffer that fate, and a pilot whose baseline is "we'll pull the historicals later" already has.
Clause 3: Success criteria, pre-committed and measurable
Illustrative language for our running process: "The pilot is a success if, on in-scope volume: (a) verified extraction accuracy is at or above 85 percent at week 6, measured by the double-check protocol on a 100-document random sample; (b) median cycle time is under 2.0 days by week 10, against the baseline median of 3.2 days; and (c) the error-rework rate is no worse than the Exhibit A baseline."
Notice three pieces of craft. First, every criterion is a number with a date, so success is a measurement, not a mood. Second, and this is the piece most charters miss, the criteria specify VERIFIED numbers with the verification method named in the clause itself. The distinction matters because unverified success metrics are how zombie pilots pass their own exams: the tool reports its own accuracy, the vendor dashboard grades its own homework, and month 4 produces a triumphant chart nobody can trace to source. You spent this entire level building the verification habit; Clause 3 is where the habit gets contractual teeth. Third, criterion (c) is a floor, not a target: it guards against the silent trade where speed improves and quality quietly pays for it.
Clause 4: The kill condition, the charter's soul
If the charter had only one clause, it would be this one: "If verified extraction accuracy at week 6 is below 85 percent on the sample protocol, the pilot ends. In-scope volume reverts to the manual SOP, the remediation list from the data readiness report reopens, and a one-page closure note is filed. This clause executes on the measurement; no meeting is required to invoke it."
Study the design, because both halves are load-bearing. The trigger is automatic: it fires on a number crossing a threshold, not on anyone's courage. Kill decisions that require a meeting do not happen, because every attendee of that meeting has, by then, a reason to postpone it. The consequence is dignified: work reverts to a functioning SOP, the remediation backlog reopens, a note is filed. Nobody is blamed, nothing is hidden, and critically, the clause fires a design, not a person. The pilot did not fail; the pilot completed its experiment and returned a negative result, which is what experiments are for.
And here is why the signature timing is everything: the sponsor signed this clause when it was hypothetical, which is the only time anyone signs kill conditions honestly. Ask a sponsor to accept an 85 percent kill threshold in week 7, with the number sitting at 81 and their name on the project, and you will watch a sincere person discover twelve reasons the measurement was unfair. Ask the same sponsor on approval day, and they will sign in ten seconds, because it costs nothing yet. Gartner's projected 40 percent of agentic project cancellations will mostly be the first kind: late, expensive, acrimonious, everything a pre-committed kill is not. The kill clause is how your pilot, whatever happens, stays out of that story.
Clause 5: Roles, the RACI made real
RACI is the standard accountability chart: for each activity, who is Responsible for doing it, who is Accountable for the outcome, who is Consulted, who is Informed. Most RACI charts are wallpaper. The charter's version is short and sharp because it assigns the three roles that actually decide a pilot's integrity: "The team lead owns the verification gate (no AI output enters the workflow unchecked). The data steward owns remediation landing (the fixes from the readiness report arrive on schedule). The assessor owns measurement integrity (samples, protocols, and the week-6 number)."
Then one sentence that does more governance work than most steering committees: "The person who measures results cannot be a person whose bonus, budget, or reported performance depends on the result." Call it measurement independence, and name it as a principle in the document. It is the same reason financial auditors do not report to the sales director. In a pilot, the temptation is structural: the natural person to measure the tool is the enthusiastic owner who championed it, and that person, however honest, is the wrong instrument. Your role as assessor exists partly to be the independent gauge, and the charter is where that independence is granted in writing rather than negotiated later under pressure.
Clause 6: Cadence and evidence
The clause: "Weekly 30-minute review against the Clause 3 numbers, standing agenda, decisions logged. The following records are kept from day one: the verification log, the exception log for out-of-scope arrivals, and the weekly accuracy sample results. A drift check (re-running the sample protocol on current volume) runs every two weeks."
This is the least glamorous clause and the one that determines whether week 6 produces a number or an argument. Evidence has a property that money does not: it cannot be borrowed retroactively. If the logs were not kept from day one, the week-6 measurement is a reconstruction, and reconstructions are exactly the soft ground on which success gets renegotiated. Notice also what this clause really is: a promise about operating rhythm. Level 3 of this program devotes an entire chapter to pilot operations, the weekly instrumentation, drift monitoring, and gate discipline of a running pilot. Clause 6 is the charter pre-committing to the rhythm that L3 will teach you to run. The contract is written in L2; the operating manual arrives in L3.
Clause 7: The re-score trigger
The final clause keeps the charter alive without letting it rot: "The readiness assessment reopens, and this charter is re-ratified or amended, if any of the following occur: an approved scope change under Clause 1; a change in the source systems or data feeds described in the data readiness report; or a staffing change in any Clause 5 role. Amendments are written, versioned, and signed by the original signatories."
Your readiness scorecard was a photograph of the organization on a particular day. If the scope, the data, or the people change, the photograph is of a place that no longer exists, and every conclusion resting on it needs re-examination. Without this clause, charters become laminated relics, technically in force and practically ignored, and the pilot's real governance moves back into hallways. With it, the charter is a living document with a defined amendment procedure: change is allowed, drift is not, and the difference between them is a signature.
AI as Your Contract Counsel
Everything in this level has paired a human discipline with an AI acceleration, and chartering is no exception. AI plays three roles here, and none of them involves holding the pen at the signature line.
First, drafting. Feed the model your selection scoring record, the one-page readiness report, and the baseline pack summary, and ask it to draft the seven clauses in the structure above. This is the same pattern you used for SOP drafting: the AI produces an 80 percent draft in minutes from source documents you have already verified, and you spend your time on judgment rather than typing. Every number in the draft then gets traced back to its source artifact, because a charter with an invented figure is worse than no charter at all.
Second, adversarial stress-testing. This is the prompt pattern worth adding to your library, and it deserves a name: adversarial contract review. Paste a draft success criterion and prompt: "You are a motivated sponsor in month 4 whose pilot is underperforming. Find every ambiguity in this success criterion that you could exploit to claim success anyway. Then rewrite the criterion to close each loophole." Run it on a criterion like "cycle time improves significantly by the end of the pilot" and the model will cheerfully produce the exploits you would otherwise meet in a real meeting: significant by whose test, cycle time measured on which volume, the end of the pilot as defined by whom, improvement against which baseline. It is the cheapest red team you will ever hire, and it works because finding ambiguity in text is precisely the kind of task where the model's tirelessness beats your afternoon attention span.
Third, consistency checking. Before signature, have the AI cross-check every number in the charter against the baseline pack and the readiness report: does the 78 percent in Clause 1 match the pack's segmentation table, does the 3.2-day median in Clause 3 match Exhibit A, does the kill threshold match the number the committee actually approved. Then verify its check the way you verify everything, on a sample, yourself. The division of labor is the one this level has drilled into habit: AI accelerates the drafting and the checking; the human owns every figure and every signature. No exceptions on the document that governs everything else.
The Worked Example: The Invoice-Exception Charter in Full
Here is the complete charter for the process this level has followed from first inventory to committee approval. Every number is hypothetical, carried forward from the illustrative baseline pack, and the whole block is built to be copied and refilled with your own process.
| Pilot Charter: Invoice-Exception Extraction Pilot, v1.0, signed 14 days after approval |
|---|
| 1. Scope. IN: the three structured exception types (missing PO, price mismatch, quantity mismatch), 78 percent of the roughly 1,150 monthly exceptions. OUT: free-text exceptions, bank-detail fields, cross-border invoices. Scope changes require a written change request signed by the sponsor and trigger Clause 7. |
| 2. Baseline. Baseline Pack v1.3 attached as Exhibit A: 1,150 exceptions per month, median cycle time 3.2 days, p90 of 9 days, rework rate 11 percent, cost per exception 31 dollars. All improvement claims compute against Exhibit A only. |
| 3. Success criteria. On in-scope volume: verified extraction accuracy at or above 85 percent at week 6 (double-check protocol, 100-document random sample); median cycle time under 2.0 days by week 10; rework rate no worse than 11 percent. Verification method as specified; self-reported tool metrics do not satisfy this clause. |
| 4. Kill condition. If week-6 verified accuracy is below 85 percent, the pilot ends: volume reverts to the manual SOP v2.1, the remediation list from the Data Readiness Report reopens, and a closure note is filed within five business days. This clause executes on the measurement; no meeting is required. |
| 5. Roles. Team lead: owns the verification gate. Data steward: owns remediation landing. Assessor: owns measurement integrity and the week-6 sample. Measurement independence: no one measuring results has compensation or reported performance tied to the outcome. |
| 6. Cadence and evidence. Weekly 30-minute review against Clause 3; verification log, out-of-scope exception log, and weekly sample results kept from day one; drift check every two weeks. |
| 7. Re-score trigger. Assessment reopens on: approved scope change, change to source systems or data feeds, or staffing change in any Clause 5 role. Amendments written, versioned, re-signed. |
| Signatures, each with date: Sponsor (VP Finance Operations), Process Owner (AP Manager), Data Steward, Team Lead (Exceptions), Assessor (you). |
Two details of craft before we leave the example. The signature block is five names, not one, because each signature converts one future argument into a past agreement: the sponsor cannot later dispute the kill threshold, the process owner cannot later dispute the scope, the assessor cannot later be pressured off the measurement protocol without visibly breaking a signed commitment. And the charter is dated fourteen days after approval on purpose: the window for collecting signatures is while the yes is warm. Every week you wait, the approval cools, calendars fill, and hypothetical clauses start feeling real enough to negotiate. Charter drafting is a same-week activity, signature collection a same-fortnight one.
The Failure Story: The Unchartered Mirror
Now the same movie without the contract, assembled from the standard failure patterns and with illustrative figures. Another division of the same company approves a document-processing pilot the same quarter. Their team is enthusiastic and modern, and they explicitly decide against a charter: "we're running this agile, we'll figure out success as we learn." It sounds humble. It is actually the opposite: it is a claim that their future selves, under sunk cost and sponsor pressure, will define success more honestly than their present selves could. Nobody would say that sentence out loud, so it gets said as "agile" instead.
Month 3: scope has tripled, entirely by enthusiasm. Free-text documents came in because a director asked nicely; two adjacent teams joined because the demo was impressive. No single addition was a decision; there was no written scope for any of them to be measured against, so each was just a friendly yes. The baseline was never frozen, so when the team wants to show improvement, they discover the before-state must be reconstructed from memory and disputed exports, and every reconstruction is contested by whoever it inconveniences. Improvement has become unprovable in principle, the exact fate the MIT researchers found underneath the 95 percent.
Month 5: results disappoint, and the definition of success moves. Cycle time gives way to "quality of insights," which gives way, one steering deck later, to "organizational learning." Nobody decides this; each deck simply inherits the previous deck's softest slide and softens it further. This is what pre-commitment exists to prevent: not lying, which is rare, but the honest, incremental, collectively deniable renegotiation that happens whenever measurement is optional and memory is the referee.
Month 9: the pilot is nobody's fault and everybody's budget. It cannot be killed, because no kill condition exists and therefore any specific executioner would be choosing personal conflict on their own authority. It cannot be defended, because no success condition exists either. It has become the fully formed zombie pilot: consuming a license, a part-time team, and a monthly agenda slot, protected by the absence of any standard it could visibly fail. It runs to month 14, consuming roughly 280,000 dollars in licenses, integration hours, and staff time (illustrative), until a new executive kills it out of pure exhaustion during a budget review. And here is the detail that should stay with you: the autopsy cannot even say whether the idea was good. The process might have been a genuine goldmine. Nobody will ever know, because nothing was measured against anything. The readiness was never the problem. The contract was.
The Capstone: Nine Artifacts, One Credential, and the Bridge to Level 3
Sign the charter and something quietly significant becomes true: you are holding the complete AI-Readiness Assessment pack for one real process. Lay the artifacts out and count them, because you built every one across this level: the inventory row that surfaced the process; the verified process map with the invented step caught and corrected; the SOP drafted from interviews and verified by the people who live the work; the baseline pack, versioned and certified; the data readiness report with its severity-ranked gaps; the people readiness scorecard; the selection scoring record that made the choice auditable; the one-page readiness report leadership actually read; and now the signed Pilot Charter. Nine artifacts. Every figure in them carries a verification chain back to source. Produced by one assessor, with AI assistance, in roughly six working weeks (illustrative, and honest about the pace).
Feel the weight of that for a moment, because it is the L2 credential in physical form. A year or two ago, that stack of work was a consulting engagement: a team of three or four, a quarter of elapsed time, a six-figure invoice, and frankly, often less verification than yours carries. The economics moved because the drafting moved to the machine while the judgment stayed with you, which is the entire thesis of the level: BCG's 10-20-70 rule says the value was never mostly in the algorithms anyway, and McKinsey's finding that high performers are the ones who fundamentally redesign workflows tells you what the stack of paper is actually for. It is not documentation. It is the raw material of a redesign.
Which is exactly where Level 3 begins. The charter is a handoff document, and the person it hands off to is you, wearing the next title: AI Process Transformer. L3 takes the workflow the charter scoped and redesigns it rather than overlaying a tool on it (the very next lesson opens that argument: paving the cow path fails). It builds the human verification gates Clause 5 promised as real designed controls with entry criteria. It runs the pilot by the Clause 6 cadence, teaching the operating rhythm the charter pre-committed. And when week 6 and week 10 arrive, it executes the scale, iterate, or kill decision using the criteria your charter locked in while everyone was honest. The assessor becomes the transformer, and the contract you are about to sign is the first page of that job.
What to Do Monday Morning
If your assessment has an approval, warm or recent, this is the week the charter gets written. The sequence:
- Draft all seven clauses from your scoring record. Feed your selection scoring record, readiness report, and baseline pack summary to your AI assistant and ask for a first draft in the seven-clause structure. One hour, then trace every number back to its source artifact.
- Write the OUT list before the IN list. Sit with the process owner and name what the pilot will not touch: the data types, the fields, the geographies, the edge cases. If the OUT list is empty, you have not finished thinking.
- Run the adversarial ambiguity prompt on every success criterion. "Find the ambiguity in this criterion that a motivated sponsor could exploit in month 4, then close it." Rewrite until the red team comes back empty.
- Set the kill condition as a number, a date, and an automatic consequence. Check that it fires a design, not a person: reversion to the manual SOP, remediation reopened, closure note filed. If invoking it requires a meeting, it is not a kill condition yet.
- Secure measurement independence in the RACI. Name who measures, and confirm in writing that their bonus, budget, and reported performance are untouched by the result. If that person is currently the pilot's champion, fix it now, on paper.
- Collect the signatures while the approval is warm. Sponsor, process owner, data steward, team lead, assessor, each dated, within two weeks of the yes. Every cold week makes hypothetical clauses feel negotiable.
- File the complete nine-artifact capstone pack, versioned. Inventory row, verified map, SOP, baseline pack, data readiness report, people scorecard, scoring record, one-page report, signed charter. Version numbers on everything. This folder is your L2 capstone and your professional evidence.
Key Takeaways
- Treat approval as a mood and the charter as the contract that survives it: the drift into MIT's 95 percent happens in the unwritten gap between "the committee said yes" and "the pilot knows exactly what it is doing."
- Sign every clause while it is still hypothetical, because pre-commitment is the only honest signature: sunk costs will re-write everyone's memory of what success meant, including yours.
- Draw scope in the negative: the OUT list is the fence that turns week-6 scope pressure into a visible change request instead of an invisible drift.
- Incorporate the baseline by reference, versioned and attached, before the tool arrives; "baseline or it didn't happen" only protects you once it is contract language.
- Write success criteria on verified numbers with the verification method named in the clause, because unverified metrics are how zombie pilots pass their own exams.
- Make the kill condition automatic in trigger and dignified in consequence: it fires on a number, ends in a working SOP and a filed note, and blames a design rather than a person, which is why a documented kill is a win and Gartner's 40 percent of agentic cancellations mostly will not be.
- Enforce measurement independence in the RACI: the person who measures the pilot cannot be a person whose bonus depends on its result.
- File the nine-artifact capstone pack and carry the charter into Level 3 as a handoff document: the workflow redesign, human gates, operating cadence, and scale-or-kill decision it pre-commits are exactly what the AI Process Transformer level teaches you to execute.
Skill.re