←
CAP Certification
Capable · M33 · lesson 33 of 54 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Documenting AI Systems

15 min

Emeka Whitfield spent three months building a machine-learning model that flagged procurement anomalies. When he handed it over to the compliance team, nobody could explain what it did or why. A contractor had written the core logic. The training data was on a laptop that had been wiped. Six weeks later the model was switched off, not because it failed, but because no one trusted something they could not describe.

That story is not unusual. AI systems get shut down, misused, or quietly ignored because the people running them cannot answer basic questions: what does this model do, what data trained it, what can go wrong, and who decides when it should stop. Documentation is how you answer those questions before someone asks them under pressure. Without it, a working system becomes a liability: hard to audit, hard to maintain, and dangerous once someone uses it beyond the scope it was validated for.

Why AI Documentation Is Different

Documenting an AI system is not the same as documenting software. With conventional software you document what the code does, deterministic steps producing predictable outputs. An AI model learns from data, and its behavior depends on patterns in that data. Two identical inputs can produce different outputs at different times, and edge cases are not bugs in the traditional sense; they are the boundaries of what the model learned. So your documentation has to cover things traditional software docs never needed: where the training data came from, what it did not include, how the model was evaluated, and what conditions cause it to perform poorly. You are not describing a system so much as a system that encodes assumptions.

Three patterns account for most documentation failures. Documentation lag is writing the documents after the system is built, from memory, by people already under pressure to start the next project; the result is inaccurate from the first day. Audience mismatch is developers writing for developers, so that business stakeholders cannot use the material and compliance reviewers cannot audit from it. Documentation decay is the rot that sets in as models are retrained, data sources change and use cases expand, while nobody updates the record because documentation is not treated as a component of the system. All three are structural problems, which means they need structural answers rather than exhortation.

The Documentation Stack

Complete documentation for a production AI system is not one document. It is a stack of seven layers, each written for a different reader and updated on a different rhythm. Deciding which layers your system actually needs is the first practical step, and for a low-stakes internal tool the honest answer is a subset rather than all seven.

  • System overview. One to two pages of plain language covering what the system does, why it was built, who uses it and what problem it solves. Revisited annually or when scope changes materially.
  • Architecture documentation. Data flows, components, integration points and infrastructure dependencies, written for engineers and updated whenever the architecture moves.
  • Data documentation. Sources, quality assessments, preprocessing pipelines, feature rationale and lineage. This layer is what makes reproducibility and bias auditing possible at all.
  • Model card. The standardized summary of intended use, out-of-scope use, performance overall and by segment, limitations and ethical considerations.
  • Operational runbook. Deployment, monitoring, alerting, rollback and incident procedures, written for whoever is on call when something breaks.
  • Governance record. The decision log, risk assessment, approval history and the description of how human oversight actually works in practice.
  • User guide. Plain-language guidance for end users on how to read the outputs, when to rely on them, and when to escalate to human judgment.

Model Cards

A model card is a structured document of roughly one to three pages summarizing everything a non-technical stakeholder needs to know about a model. The format was introduced by a major AI research lab in 2018 and has become the most widely adopted standard for AI transparency. A complete card covers model details such as name, version, type and maintenance contact; intended use, meaning the tasks and populations it was designed for; explicitly out-of-scope uses; training data, including sources, date range, composition and known gaps; evaluation results reported both overall and for the subgroups that matter; ethical considerations, including known fairness concerns and what was done about them; and the conditions under which the model performs poorly.

The subgroup reporting is the part most often skipped and the part auditors ask about first, because a single headline accuracy figure can conceal a model that works well for most cases and badly for a segment that was thin in the training data. Concrete caveats also do more work than careful hedging: a model card for a customer-churn predictor trained on data from 2021 to 2023 should say plainly that it was not trained on data from an economic downturn. That one sentence stops someone over-relying on it in a recession. A model card is not a boast sheet. It is a warning label and an owner's manual at the same time.

System Diagrams and Decision Logs

A model card describes the model. A system diagram describes everything around it: how data flows in, how outputs flow out, and what human steps sit in between. Draw it at two levels. The first is a high-level flow from raw data source through preprocessing, model, output, downstream action and human review, fitting on one page and making sense to a non-technical executive. The second is operational detail: which database, which endpoint, what format the output takes, where it lands, who gets notified. Keep diagrams in a version-controlled format, a text-based tool such as Mermaid or an editor like draw.io with the source exported. A diagram that exists only as an image in a slide deck will go stale within months.

Every AI project also involves dozens of choices invisible in the finished system. Why this model architecture rather than another? Why exclude records older than five years from training? Why set the confidence threshold at 0.72 instead of 0.8? A decision log captures the choice and the reasoning behind it, and need not be elaborate: a table with four columns covering date, decision, rationale and alternative considered, one row per significant choice. Decision logs serve three purposes. They stop teams re-debating settled questions, they let auditors understand why a system behaves as it does, and they give future maintainers context before they change something.

Data Dictionaries and Data Cards

A data dictionary defines every field that enters or leaves the AI system. For each field you record the name, a plain-English description, the data type, the allowed values or range, where it comes from, and what a blank or null value means. This matters because data fields accumulate hidden meaning. A field called status might carry values of 1, 2, 3 and 7, and only the original developer knows that 7 means archived before migration. When that developer leaves, the model starts treating archived records as active. You do not need a dictionary for every field in your database, only for the fields that cross the system boundary, which for most projects is 20 to 50 fields.

At the dataset level, the companion artifact is a data card: name, version and creation date; collection methodology and sources; size, format and update frequency; known quality issues; preprocessing applied; how sensitive information is handled; licence and usage restrictions; and any demographic or geographic gaps. Maintaining data cards next to model cards gives auditors the full picture of the evidence base, which is exactly what they will otherwise ask you to assemble under deadline.

The Living Governance Record

Governance documentation has to be a living record rather than a static artifact produced once for a sign-off meeting. A workable template has five parts. The system profile names the use case and the business, technical and data owners, along with the risk classification the system falls into under whichever framework applies to you. The risk assessment lists identified risks across accuracy, bias, security, privacy and dependency, with likelihood and impact, the mitigations in place, and, critically, the residual risks that were accepted and by whom.

The approval history records who approved deployment at each stage, what evidence they reviewed and what conditions they attached. The incident log captures failures, anomalous outputs and misuse with dates, responses and lessons learned. The review schedule sets mandatory review dates, every 12 months at minimum, plus triggering events such as retraining, scope expansion or an adverse incident. The residual-risk and approval sections matter because they convert an implicit organizational decision into an attributable one, which changes how carefully it gets made.

Keeping Documentation Current Without Creating a Burden

Documentation rots the moment someone changes the system without updating the record. The solution is not to write more documentation but to make updating it the natural end of every change. Three practices carry most of the weight. Attach documentation to your deployment checklist, so no model update goes live without a bumped model card version and a decision log entry; if it is not on the checklist it will not happen. Keep docs close to the code, in the same repository, so a developer pushing a change sees the documentation file next to their work. And set a recurring quarterly review in the system owner's calendar, no more than 30 minutes spent asking whether anything has changed that the documents do not reflect.

The rigorous version of this is documentation-as-code: treating documents as software artifacts that are version-controlled alongside the system, built from templates that require specific sections to be completed, validated by automated completeness checks that can fail a pipeline, and reviewed as part of code review itself. The operating rule is no deploy without docs, held as firmly as no deploy without tests. Two organizational habits make it stick. Assign a named documentation owner for each production system, so responsibility is a person rather than a team. And keep a visible documentation debt tracker alongside technical debt, so a gap found during an audit becomes a tracked item at management level rather than a private embarrassment. Culture follows from use: when people settle an argument by referencing the model card, documentation stops being overhead.

Black Boxes and Borrowed Systems

Two situations look like they make documentation impossible, and neither does. The first is the genuinely opaque model, typically a large neural network whose internal representations resist explanation. The response is to shift focus from internal mechanics to external behavior: performance across use case categories and demographic groups, known failure modes with examples, the conditions that trigger anomalous or low-confidence outputs, results from interpretability tools such as SHAP values, LIME explanations or attention visualizations with honest caveats about how far they can be trusted, and human evaluation studies characterizing output quality. That is more useful to almost every audience than claims of explainability that will not survive scrutiny.

The second is the vendor-provided system, where you cannot document internals because you have no access to them. Responsibility shifts rather than disappears. You document how the system is configured and integrated, the business decisions embedded in that configuration such as which fields are inputs and what thresholds trigger which actions, and your own validation testing along with the conditions under which you accept its output. You capture whatever documentation the vendor provides inside your governance record, and you keep a dependency log describing what would happen if the vendor discontinued, changed or degraded the service.

Regulatory Expectations

Documentation requirements are hardening across sectors, and the direction of travel is consistent even where the instruments differ. In financial services, the Federal Reserve's SR 11-7 model risk management guidance has required documentation for models used in credit, risk or pricing decisions for over a decade, and AI systems fall under it. In healthcare, the 21st Century Cures Act and subsequent FDA guidance require AI and machine-learning software functioning as a medical device to document design, training data and validation evidence. The EU AI Act, phased in from 2024 to 2027, requires technical documentation for high-risk systems covering system description, training data, performance metrics and conformity assessment records. In the public sector, Executive Order 14110 and equivalent national policies increasingly oblige agencies to document AI used in consequential decisions. Organizations outside all of these regimes still benefit from documenting to their standard, because it produces a defensible governance posture if anyone ever asks.

The Handoff Test

Here is a practical benchmark. Could a competent colleague who was not involved in building this system take it over in two weeks without talking to you? If the answer is no, the documentation is incomplete and you now know roughly where. A second test targets the user-facing layer: ask two or three representative users to complete a defined task using only the documentation, such as deciding whether the system suits a particular edge case. Where they struggle or draw the wrong inference, the documentation is failing regardless of how polished it looks to its authors.

Run both tests before you think you need to. Emeka ran his after the fact. His model card now carries a section headed What We Wish We Had Known, recording three assumptions baked into the training data that only surfaced during the compliance audit. That section alone has kept two subsequent teams from repeating the same mistakes, which is a fair description of what documentation is for.

Anti-Patterns

  • Writing the documentation after the build, from memory. The record is inaccurate on the day it is created, and everyone who relies on it inherits errors nobody knows are there.
  • Writing every layer for engineers. When the only documents are technical, the business owner cannot answer a stakeholder question and the reviewer cannot audit, so the material fails exactly when it is needed.
  • Treating the model card as a capability sheet. A card that lists strengths, omits out-of-scope uses and reports one headline accuracy number invites exactly the over-reliance it exists to prevent, because aggregate performance hides the segment where the model fails.
  • Storing diagrams as images in slide decks. With no version-controlled source the diagram cannot be updated cheaply, so it quietly starts describing a system that no longer exists.
  • Documenting field names without field meanings. A dictionary that omits what a null value or a legacy status code means will not prevent the misclassification it existed to prevent.
  • Assuming a vendor system needs no documentation of your own. You still own the configuration, the thresholds and the validation evidence, and documentation assigned to a team rather than a named person decays by default.

Practice Prompts

  • Pick one AI system in production and write its system overview in plain language on a single page. Show it to someone outside the team and ask what the system does.
  • Draft a model card for it and fill in the out-of-scope use section first. Notice how much of it you have to ask someone about.
  • Take your last significant design decision and write the decision log row for it: date, decision, rationale, alternative considered. Then check whether anyone else could reconstruct the rationale without you.
  • List the fields crossing the boundary of one AI system and mark every one whose null or legacy values you cannot explain. That is your data dictionary backlog.
  • Name the documentation owner for each AI system your team runs. Where you cannot name a person, you have found the system most likely to be undocumentable within a year.
  • Run the handoff test as a thought experiment with a specific colleague in mind, then write down the questions they would have to ask you. Those are the gaps to close first.

Reflection

Emeka's model was not switched off because it was wrong. It was switched off because nobody could establish that it was right, and in an organization facing a compliance question, unverifiable and wrong get treated the same way. The uncomfortable part is how ordinary the causes were: a contractor who moved on, a laptop that got wiped, design decisions that lived only in one person's memory. Think about a system your team depends on today. If the person who built it left this month, what would remain: something your organization can explain, defend and maintain, or a working artifact nobody has standing to vouch for?

Glossary

  • Model card: A short structured document summarizing a model's intended use, out-of-scope use, training data, evaluation results and limitations for a non-technical reader.
  • Data card: The dataset-level counterpart to a model card, recording collection methodology, size, quality issues, preprocessing, licensing and known gaps.
  • Data dictionary: A field-level reference defining the name, description, type, permitted values, source and null semantics of every field crossing the system boundary.
  • Decision log: A running record of significant design choices with the reasoning and alternatives considered, used to prevent re-debate and explain behavior to auditors.
  • Governance record: The living file combining system profile, risk assessment, approval history, incident log and review schedule for one AI system.
  • Documentation-as-code: Version-controlling documentation alongside the system, enforcing completeness through automated checks, and reviewing it as part of code review.
  • Documentation debt: The visible backlog of missing or outdated documentation, managed like technical debt rather than discovered during an audit.
  • Handoff test: Whether an uninvolved competent colleague could take over the system within two weeks using the documentation alone.

Closing

The goal is not a perfect documentation suite on day one. It is documentation accurate enough to hand off, accurate enough to audit, and maintained by a process light enough that it stays that way. The stack tells you which layers a system needs, the model card and data card carry what outsiders will ask about, the decision log and governance record preserve reasoning that would otherwise leave with a person, and the structural habits of checklists, proximity to the code, named owners and visible debt keep the whole thing from rotting. None of it is glamorous. It is the difference between a system your organization owns and one it merely operates until the day someone asks a question it cannot answer.

Key Takeaways

  • AI documentation must cover training data, limitations and assumptions, not just behavior. A model encodes what it learned, so the record must describe the evidence base and the boundaries, not only the function.
  • Documentation fails in three predictable ways. Lag, audience mismatch and decay are structural problems, so they need structural fixes rather than reminders to try harder.
  • Think in layers, and choose the ones your system needs. Overview, architecture, data, model card, runbook, governance record and user guide each serve a different reader on a different rhythm.
  • The model card is the highest-value single artifact. Intended use, out-of-scope use, subgroup performance and failure conditions stop a system being used where it was never validated.
  • Capture the reasoning, not just the result. Decision logs, data dictionaries and named acceptance of residual risk preserve context that otherwise walks out of the building with one person.
  • Make currency structural. Deployment checklists, docs in the code repository, a named owner, a quarterly review and a visible debt tracker beat any amount of individual discipline.
  • Opacity and vendor systems change what you document, not whether you document. Behavior characterization replaces internal mechanics, and configuration, validation and dependency records replace source-level detail. Test the result: the handoff test and a short user task test find gaps faster than another review by the authors.

Frequently Asked Questions

How much documentation is enough for a small internal tool? Enough that someone else could take it over and that you could explain it to a reviewer, which usually means a short system overview, a model card, a data dictionary for the fields crossing the boundary, and a decision log. The seven-layer stack describes a mature production system under regulatory attention; a low-stakes tool that borrows the whole structure produces documents nobody maintains, which is worse than a smaller set that stays accurate.

Who should actually write it? The people who made the decisions, with a named owner responsible for the record staying current. Delegating the writing wholesale to someone who was not in the room reliably produces documents that describe the architecture correctly and the reasoning not at all, and it is the reasoning that future maintainers and auditors need.

What if the model is a black box we genuinely cannot explain? Document the behavior instead of the mechanism. Performance by category and subgroup, known failure modes, the conditions that produce low-confidence outputs, and interpretability results carried with honest caveats serve almost every audience better than an explainability claim that cannot withstand a direct question.

How do we stop documentation going stale without adding process overhead? Tie the update to something that already happens. A deployment checklist item and a docs file sitting next to the code in the same repository cost almost nothing per change, whereas a separate documentation review meeting is the first thing dropped when the schedule tightens. Add a short recurring review to catch what the checklist misses, and keep the gaps visible rather than private.