Turning Messy Process Knowledge into Clean Inputs
The analyst was proud of the deliverable, and for two weeks everyone else was too. Asked to document the current invoice-exception process ahead of an AI readiness assessment, she had gathered everything she could find (an old standard operating procedure, a wiki page, a folder of help-desk tickets, transcripts of six meetings), pasted the whole pile into a model with a well-built four-part prompt, and received back twelve pages of clean, confident, beautifully structured process documentation. It read like the work of a senior consultant. Then, in a review meeting, a finance manager stopped on page four and frowned. The document stated, in fluent and authoritative prose, that exceptions above 5,000 dollars require director approval. That threshold had been raised to 25,000 dollars fourteen months earlier. The model had pulled it from a 2019 SOP, presented it as current truth, and wrapped it in prose so polished nobody thought to check. Unwinding the error, and the mistrust it seeded, took a week. The prompt was fine. The inputs were the problem, and this lesson is about the discipline that fixes them.
The Other Half of the Machine
In the previous lesson you built the four-part prompt: role, context, task, and constraints, with the cite-your-source rule as its seatbelt. That structure is genuinely half the machine. It tells the model who to be, what it may rely on, what to produce, and how to behave. But look again at the second part, context. The prompt structure tells the model to rely on the material you provide. It says nothing about whether that material deserves to be relied on. A perfectly constructed prompt wrapped around a contradictory, stale, unlabeled pile of documents is a beautiful engine bolted to a tank of contaminated fuel.
Here is the uncomfortable mechanical truth about how a language model treats the context you paste. It does not know which document is newer. It does not know that the wiki page was written by an intern and the SOP by the process owner. It does not know that the help-desk tickets describe what actually happens while the SOP describes what was supposed to happen in 2019. Unless you tell it otherwise, the model treats every sentence in its context window as equally true, equally current, and equally authoritative. When two sources disagree, it does not flag the disagreement; it quietly blends them, or picks whichever version is phrased more confidently, and moves on. When the pile has a gap, the model does not leave a gap. It fills the space with the most statistically plausible continuation, which in process documentation means an invented step that sounds exactly like a real one.
This is why the analyst's twelve pages were so dangerous. They were not visibly wrong. They were a smooth average of five conflicting accounts of one process, with the seams sanded off and the stale parts presented in the same confident register as the fresh ones. Garbage in, garbage out is the oldest rule in computing, but generative AI adds a cruel twist: garbage in, polished garbage out. The output does not look like garbage. It looks like the best process documentation your organization has ever produced, right up until someone who knows the process reads page four.
The model treats everything you feed it as equally true, equally current, and equally authoritative. Input hygiene is the discipline of making sure that assumption cannot hurt you.
The good news, and the reason this lesson sits in a program for process professionals rather than data scientists, is that the fix is not technical. Curating sources, checking dates, surfacing conflicts, labeling provenance: this is document control, this is records discipline, this is exactly the work you have done every time you prepared an audit binder or a due-diligence data room. Input hygiene is pure process work. You already have the muscles. This lesson gives them a target.
Where process knowledge actually lives
Before the discipline, the diagnosis. Ask any operations professional where the knowledge of a given process lives and, after a short laugh, you will get an inventory like this one.
- The official SOP, a standard operating procedure last revised in 2019, still marked "current" in the document management system because nobody owns the review cycle. Some of it is right. Some of it describes a system that was decommissioned two years ago.
- The wiki page a well-meaning team lead wrote in 2023 to "finally document how this really works." It contradicts the SOP in four places. It has no date visible on the page and no listed author.
- The ticket archive: 400 help-desk tickets in which the process's real exceptions, workarounds, and failure modes are recorded one frustrated paragraph at a time. This is the most honest source in the building and the least readable.
- The veteran's head. Maria has run the process for eleven years. The three most important rules exist nowhere on paper. One of them reverses what the SOP says, and everyone downstream quietly knows to follow Maria, not the document.
- Six meeting recordings in which the process was discussed, disputed, and partially redesigned, including one meeting where a decision was made and then reversed the following week in a hallway conversation that was never recorded at all.
None of this is a scandal. This is the normal condition of process knowledge in real organizations, and every assessment you will ever run starts from a pile like this one. The mistake is not having the pile. The mistake is feeding the pile to a model raw, because each of the pile's pathologies maps directly onto a model failure mode. Contradictions get averaged into a version of the process that nobody actually follows. Stale documents get treated as current because the model has no calendar. Gaps get filled with plausible invention. And the veteran's unwritten rules, the most operationally important knowledge of all, simply do not exist as far as the model is concerned, so the output is confidently silent about exactly the things that matter most.
Notice what this means for the shape of your job. The four-part prompt took you an hour to learn. Preparing the inputs for that prompt will routinely take you two hours per assessment question, and those two hours are not overhead to be minimized. They are the work. A readiness assessor who spends ten minutes on the prompt and two hours on the inputs has the ratio exactly right.
The Five Habits of Input Hygiene
Input hygiene is five habits, each small enough to practice on your next task and each mapped to a specific way that raw piles poison model output. Take them slowly; the worked example later in this lesson runs all five in sequence.
Habit 1: Triage before you feed
Not everything goes in. The single most common input mistake is completeness anxiety: the fear that leaving a document out means the model might miss something, so everything gets pasted. But a model's attention is not infinite, and more importantly, every irrelevant or marginal document you include is another opportunity for contamination. The 2019 SOP's obsolete approval threshold can only end up in the output if the 2019 SOP goes into the input.
Triage means selecting sources by relevance to the specific question you are asking, not to the process in general. If your question is "what is our current invoice-exception process?", the vendor onboarding SOP does not go in, even though it mentions invoices. The 2021 reorganization announcement does not go in, even though it renamed the team. A useful test: for each candidate document, complete the sentence "this source earns its place because it tells the model X about this question." If you cannot complete the sentence in operator terms, the document stays out. Mini-example: an assessor with 30 candidate documents for a question about complaint handling keeps 9. The other 21 are not deleted; they are listed in a "considered and excluded" note, one line each, so that a reviewer can see the triage happened rather than wondering what was missed.
Habit 2: Date and freshness check
The model has no calendar, so you must be its calendar. Every source that survives triage gets labeled with its date (creation or last meaningful revision, and if you cannot determine either, that itself is a finding) and a one-word freshness verdict: current, stale, or unknown. Current means someone who owns the process confirms it still describes reality. Stale means it is known or suspected to describe a past state. Unknown means nobody can vouch either way, which for decision purposes you treat as stale with an asterisk.
Stale does not always mean excluded. A stale SOP can be exactly what you need when the question is "how has this process drifted from its documented design?" The rule is that stale material goes into the prompt only with a warning label attached, in text the model can see: "The following SOP is dated March 2019 and is known to be outdated in at least the approval thresholds. Use it only as evidence of the documented design, not of current practice." Ten seconds of typing, and the single most common failure in the analyst's story becomes structurally impossible. The model can no longer present the 5,000 dollar threshold as current, because you told it, inside the context itself, that the document is old.
Habit 3: Surface conflicts instead of letting the model resolve them silently
When the SOP says exceptions go to the finance director and the ticket archive shows they actually go to a shared queue that a coordinator triages, you have a conflict. The raw-pile approach hands both sources to the model and lets it resolve the conflict silently, which it will do, smoothly, without telling you it did so, and without any principled basis for its choice. The hygienic approach does the opposite: you tell the model the conflict exists, and you instruct it to list conflicts, never to pick a winner.
The phrasing is worth memorizing: "Sources A and D disagree about the routing of exceptions. Do not resolve this disagreement. Present both versions side by side, attribute each to its source, and flag it as an open conflict requiring human confirmation." This habit is where input hygiene most directly serves your readiness work, because in an assessment, the conflicts are not noise to be cleaned away. They are the signal. A process whose official documentation contradicts its ticket trail is a process with a documentation debt, a training gap, or an unmanaged change, and every one of those is exactly what a readiness assessment exists to find. Let the model average the conflict away and you have not just gotten a wrong answer; you have destroyed the finding.
Habit 4: Provenance labels on every chunk
Every block of text you paste gets a prefix in square brackets: [SOURCE: Invoice-Exception SOP v3, dated 2019-03, owner: Finance Ops, status: STALE], then the text. Every chunk, every time, no exceptions. This feels fussy for the first day and then becomes automatic, and it is the habit that makes the previous lesson's cite-your-source rule actually function. You instructed the model to cite its sources for every claim; provenance labels are what give it sources to cite. Without labels, "cite your source" produces vague gestures at "the provided documents." With labels, it produces "per [SOURCE: Ticket archive, 2024-2025], exceptions are routed to the shared queue," and now every claim in the output is auditable in seconds. A reviewer can trace any sentence back to a specific document and a specific date. That traceability is the difference between AI output you can defend in a steering committee and AI output you have to defend with your personal credibility.
Habit 5: Redact before you paste
The last habit is the one that protects your career rather than your analysis. Before anything goes into a model, it gets a redaction pass: customer names, employee names beyond role titles, PII (personally identifiable information, meaning anything that identifies a specific person, like emails, phone numbers, account numbers), credentials and system passwords that somehow always end up pasted into tickets, and anything your organization's AI usage policy restricts. Help-desk tickets are the classic trap: they are your most honest source and also the one soaked in customer identifiers and the occasional "temp password is Summer2024!" that a support agent typed in a hurry.
The redaction pass on a typical Source Pack takes about ten minutes: find and replace names with role labels ("Customer A," "the AP clerk"), strip account numbers, delete credentials on sight. Ten minutes, against the alternative: being the named individual who pasted customer data into an external AI tool in violation of policy, discovered during the first audit. There is a full lesson on privacy, confidentiality, and AI policy later in this level, and it goes much deeper. For now, install the reflex: nothing goes into the context window before the redaction pass, ever, even when you are in a hurry, especially when you are in a hurry.
The Artifact: The Source Pack
The five habits become a repeatable deliverable when you bundle them into this lesson's named artifact: the Source Pack. A Source Pack is a curated, labeled bundle of inputs prepared for one specific assessment question. Not for a whole process, not for a whole department: one question, one pack. The discipline of scoping to a single question is what keeps triage honest and packs reusable.
A Source Pack has two components. The body is the set of source excerpts themselves, each carrying its provenance label and, where needed, its staleness warning, redacted and ready to paste. The cover is the Source Register, a one-page table that records, for every source in the pack:
- What it is: document type and a one-line description in operator terms.
- Date: creation or last meaningful revision, or "undated" as a recorded fact.
- Owner: the person or role accountable for the source, or "unowned," which is itself a readiness finding.
- Freshness verdict: current, stale, or unknown, with a word on how you know.
- Trust level: high, medium, or low, based on authorship and confirmation, not on how official the document looks.
- Known conflicts: which other sources it disagrees with, and about what.
The Source Register does three jobs at once. It forces you through habits one through four, because you cannot fill in the columns without doing the work. It travels with the AI output as evidence, so when someone asks "what was this analysis based on?", the answer is a page, not a shrug. And it quietly becomes one of the most useful assessment documents in its own right: a register showing that a core process is described by five sources, three of them stale, two of them in conflict, and one of them unowned has told a readiness story before the model generates a single word. More than once, an assessor building a Source Pack has found the assessment answer in the register and used the model only to confirm it.
Keep every pack you build. Source Packs compound: the register you built for the invoice-exception question is eighty percent of the register you will need for the invoice-approval question next month, and a shelf of packs is the beginning of exactly the AI-ready data practice the enterprise-scale statistics later in this lesson are about.
A Worked Example: Two Assessors, One Question
All numbers in this example are hypothetical and illustrative; the pattern is what you should carry away. Two assessors at the same mid-size company are separately asked the same assessment question: "what is our current invoice-exception process?"
The first assessor, the one from the opening scene, optimizes for speed. She exports everything the document system returns for "invoice exception," pastes roughly forty documents' worth of text into the model beneath a competent four-part prompt, and has twelve polished pages by lunch. Total preparation time: about twenty minutes. The output blends the 2019 SOP with the 2023 wiki, resolves their four contradictions silently (in three cases in favor of the older document, because it was phrased more formally), presents the obsolete 5,000 dollar approval threshold as current, and includes one processing step that appears in no source at all, invented to bridge a gap between two documents that described different eras of the process. The document circulates for two weeks and begins to be quoted in a pilot-scoping deck. Then the finance manager catches the threshold. Every claim in all twelve pages is now suspect, none of them cites a checkable source, and the assessor spends the better part of a week re-verifying the document line by line while the pilot-scoping work waits. Twenty minutes of preparation, roughly thirty hours of cleanup, and a durable dent in how much the finance team trusts anything labeled "AI-assisted."
The second assessor builds a Source Pack. The search returns 43 candidate documents. Triage against the specific question cuts that to 11 sources that each earn a one-line justification; the other 32 go into the considered-and-excluded list. Dating the survivors takes a round of emails and two short conversations, and produces 3 stale flags, including the 2019 SOP, which stays in the pack wearing an explicit warning label because the drift between documented design and current practice is relevant to the question. Cross-reading the survivors surfaces 2 conflicts, which are written into the Source Register and into the prompt itself with the do-not-resolve instruction: the SOP-versus-tickets routing disagreement, and a threshold discrepancy between the wiki and a controller's memo. A 30-minute interview with Maria, the eleven-year veteran, is written up into a dated, owned source of its own (this is how knowledge in someone's head becomes a document a model can use: you interview the head). The redaction pass strips customer names and one exposed system credential from the ticket excerpts. Total curation time: about 2 hours.
The model output from the pack is less smooth than the first assessor's. It is also worth incomparably more. Every claim carries a source tag. The two conflicts appear as flagged open questions with both versions presented side by side, and one of them, once escalated, turns out to be an unmanaged process change that nobody had communicated to the offshore team, a genuine finding with a real error-rate consequence. The stale SOP is quoted only as evidence of design drift. Where the pack is silent, the output says "not covered by provided sources" instead of inventing. The review meeting takes forty minutes, produces two action items, and ends with the finance manager asking whether the Source Register template could be used for the payment-run process too. Two hours of preparation against a week of unwinding: that is the trade, and you will make it dozens of times in this role.
| Symptom in the output | Raw pile | Source Pack |
|---|---|---|
| Contradictions between sources | Silently averaged into a version nobody follows | Listed as flagged conflicts with both versions attributed |
| Stale documents | Presented as current truth in confident prose | Quoted only with their date and a staleness warning |
| Gaps in the sources | Filled with plausible invented steps | Marked "not covered by provided sources" |
| Claims and citations | Uncheckable references to "the provided documents" | Every claim traceable to a labeled source and date |
| Sensitive data | Customer names and credentials pass straight through | Redacted before anything reaches the model |
| Review meeting | Line-by-line forensic re-verification | Discussion of two flagged findings |
| Reader's trust after one error | Collapses across the whole document | Contained: the register shows exactly what was relied on |
This Is Data Readiness, Practiced at Desk Scale
Zoom out from your desk for a moment, because the discipline you just learned has an enterprise-scale shadow, and the numbers attached to it are the stakes of this whole program. Gartner reported in February 2025 that 63 percent of organizations either lack AI-ready data practices or are unsure whether they have them, and predicted that through 2026, 60 percent of AI projects that lack AI-ready data will be abandoned. Read those two figures together and they describe, at the scale of entire project portfolios, exactly the failure you watched happen to the first assessor: organizations feeding models from piles they never triaged, dated, de-conflicted, labeled, or governed, and then abandoning the projects when the confident synthesis of garbage meets its finance manager.
The enterprise version of the problem involves data platforms, pipelines, and governance committees, and later levels of this program deal with it at that altitude. But the logic is identical, and this is the career point worth underlining. When you build a Source Pack, you are practicing data readiness at desk scale: source inventory, freshness assessment, conflict identification, provenance tracking, access and privacy control. These are the same disciplines, on one page instead of one platform. An assessor with a shelf of Source Registers does not just produce better AI outputs today; she is accumulating the ground-level evidence of exactly where the organization's process knowledge is stale, conflicted, and unowned, which is precisely the readiness picture the 63 percent are missing. The person in the building who can demonstrate, with artifacts, what AI-ready inputs look like in practice is not the person who owns the next stalled pilot. She is the person the steering committee starts inviting when the data-readiness question comes up, because she is the only one holding evidence instead of opinions.
What to Do Monday Morning
This lesson becomes a skill the first time you build a pack and feel the difference in the output. Here is the sequence.
- Pick one live assessment question, narrow and real: "what is our current X process?" for a process you are actually working on. One question, not a process area.
- Inventory the candidate sources: search the document system, the wiki, the ticket queue, and your own inbox, and list everything that plausibly bears on the question. Do not read deeply yet; just build the candidate list with a count.
- Run triage: keep only sources that earn a one-line "this tells the model X about this question" justification. Record the excluded ones in a considered-and-excluded list.
- Build the Source Register: for each survivor, fill in what it is, date, owner, freshness verdict, trust level, and known conflicts. Where knowledge lives in a person's head, book 30 minutes, interview them, and write it up as a dated, owned source.
- Run the redaction pass: ten minutes, names to roles, strip identifiers and credentials, check the pack against your organization's AI usage policy before anything is pasted anywhere.
- Run the diff experiment: take the same four-part prompt from the previous lesson and run it twice, once against the raw candidate pile and once against your Source Pack. Put the two outputs side by side and mark every place they differ: every silently resolved conflict, every stale claim presented as current, every invented step. That marked-up diff is the most persuasive two pages you can show a skeptical colleague, and it goes into your readiness portfolio next to the pack itself.
Key Takeaways
- Treat the four-part prompt as half the machine: the other half is what you feed it, because a model treats everything in its context as equally true, equally current, and equally authoritative unless you label it otherwise.
- Expect process knowledge to live in a contradictory pile (a stale SOP, a conflicting wiki page, hundreds of tickets, a veteran's head, meeting recordings), and never feed that pile to a model raw: it will average contradictions, present stale documents as current, and fill gaps with plausible invention.
- Practice the five input-hygiene habits: triage before you feed, date and freshness-check every source, surface conflicts explicitly, prefix every chunk with a [SOURCE: ...] provenance label, and redact before you paste.
- Instruct the model to list conflicts and never pick a winner, because in readiness work the disagreements between sources are findings, not noise, and silent resolution destroys them.
- Build a Source Pack for each assessment question: a curated, labeled, redacted bundle of sources fronted by a Source Register recording each source's type, date, owner, freshness verdict, trust level, and known conflicts.
- Budget roughly two hours of curation per question and defend that time as the work itself, against the illustrative alternative of twenty minutes of pasting followed by thirty hours of unwinding a polished, wrong synthesis.
- Spend the ten-minute redaction pass every single time: stripping PII, customer names, and credentials before pasting is the cheapest career protection available, and the full privacy discipline arrives later in this level.
- Recognize the Source Pack as data readiness at desk scale: Gartner's finding that 63 percent of organizations lack AI-ready data practices, and its prediction that 60 percent of AI projects without AI-ready data will be abandoned through 2026, are the enterprise-scale version of the raw pile, and your shelf of Source Registers is evidence you are on the other side of that statistic.
Skill.re