Multimodal AI: Text, Image, Audio, Video
Gloria Nakamura had thirty-one minutes before the morning incident briefing, and a vendor demo that promised to change how her state emergency management agency would see the next wildfire. Gloria was a senior emergency-management analyst, fifteen years into a career that began with paper maps and a marine radio. The system on her screen claimed it could watch satellite imagery, listen to scanner traffic, read 911 transcripts, and watch traffic-camera video, all at once, and tell her in plain English where a fire was spreading and which neighborhoods to evacuate first. It was genuinely impressive. It was also, she suspected, dangerous in ways the sales engineer had not mentioned. Her job that morning was not to say yes or no. It was to figure out, modality by modality, where this thing could fail and what a human would have to check before anyone acted on what it said.
What Multimodal Actually Means
A modality is just a type of input: text, still images, audio, or video. Most AI systems you have met so far handle one. A language model reads and writes text. A transcription tool turns audio into text. An image classifier looks at a photo and labels it. Each does one thing.
A multimodal system handles more than one modality and, crucially, reasons across them together. Instead of running a separate tool for the satellite image, another for the radio audio, and then asking a human to connect the two, a multimodal model takes image, audio, and text into a single reasoning process and produces one combined answer. The technical term vendors use is a vision-language model, a system that turns an image into a numerical representation a language model can process alongside words. Newer systems extend that to audio and video.
Government works in every modality already. Agencies analyze documents that are both text and image, monitor through satellite imagery, review video recordings, and work with audio transcripts. A system that reasons across all of them at once is genuinely transformative, and the standard claim made for it is that this produces understanding richer than any single modality can provide. That claim is true when the fusion is correct. When it is wrong, the resulting error is harder to diagnose than a mistake in any single feed, because there is no second feed left to disagree with it.
How combining modalities differs from running tools side by side
Gloria drew the distinction on her whiteboard. Side-by-side single-modality tools give you four separate outputs and leave the human to reconcile them. If the transcription tool mishears an address, you can catch it because the satellite view disagrees. A multimodal model reconciles the feeds internally and hands you one confident conclusion. That is faster, and in a fast-moving incident speed saves lives. But the reconciliation now happens inside a system you cannot fully inspect. If the model decides the smoke on camera matches the wrong caller's location, it may present a single clean answer that is simply wrong, with no contradicting feed visible to warn you.
The difference matters for Gloria because emergencies are inherently multimodal. A wildfire is a thermal signature on a satellite image, a panicked caller on a 911 line, a plume of smoke on a traffic camera, and a string of dispatch notes, all describing one event. A single-modality tool sees one slice. A multimodal tool claims to see the whole picture, and that claim is exactly what makes it powerful and exactly what makes it risky.
Where Multimodal AI Earns Its Place
Before assessing risk, Gloria listed the uses that plausibly justify the technology in public-sector work. Naming the use case precisely is the first discipline, because "AI for emergency response" is too vague to govern.
- Imagery analysis, image plus text. Reading satellite or aerial photos to estimate flood extent, map building damage after a tornado, or detect change between a before image and an after image. The model labels structures and produces a textual damage summary a human analyst then verifies.
- Transcription and translation, audio to text. Turning 911 audio, public-comment sessions, or multilingual hotline calls into searchable text, sometimes translated. This supports both operations and accessibility.
- Accessibility captioning, audio and video to text. Generating captions for public meetings and emergency broadcasts. Section 508 of the Rehabilitation Act requires that information and communication technology used by federal agencies be accessible to people with disabilities, and accurate captions are a core part of that obligation. Auto-captions help, but they do not by themselves satisfy 508 if they are wrong.
- Document understanding, image plus text. Reading scanned forms, permits, or handwritten field reports, identifying which field contains what, and extracting structured data. Useful for damage-assessment paperwork that floods in after a disaster declaration.
- Video understanding, video plus image plus audio. Watching traffic or drone footage to flag events such as a road washout, a collapsed bridge, or a stalled evacuation route, combining what is seen with what is said on accompanying audio.
Each of these is real. None is safe to deploy without naming the specific failure mode and the human checkpoint, which is what Gloria built next.
Vision-Language Models
A vision-language model can do four things. It can look at an image and answer questions about it, such as what type of equipment is shown. It can look at an image and extract text from it, which is optical character recognition or OCR. It can look at an image and describe what it shows. And it can understand an image in the context of accompanying text, for example being handed an aerial photo and a street map and asked where the photo was taken.
Mechanically, a vision encoder trained on images converts the image into a numerical representation, and a language model trained on text processes that representation alongside words. The system learns to reason about images and text together rather than sequentially, which is why the output arrives as one integrated answer rather than a caption plus a separate analysis.
Government applications cluster in four areas. Document analysis extracts information from scanned forms, identifies fields, and classifies documents. Satellite imagery analysis identifies buildings and roads, detects change over time, and estimates damage after disasters. Quality assurance reviews inspection photographs to identify defects or non-compliance. Evidence analysis examines photographs of crime scenes or accident scenes and extracts details for a human investigator. The last of those carries obvious consequence and belongs in the highest tier of any oversight scheme you build.
Video and Audio Understanding
Video understanding systems watch footage and describe what is happening, answer questions about it, identify discrete events such as a person entering a building or a vehicle collision, track objects across frames, and combine visual understanding with audio transcription. The mechanism is a pipeline: the video is broken into frames, a vision model processes the frames, an audio model processes the sound, and a language model reasons across both to understand what is happening and what is being said.
The government applications divide sharply by how much they touch individuals. Traffic analysis monitors intersections, identifies accidents, and informs signal timing. Meeting transcription records public proceedings, transcribes the audio, summarizes content, and identifies action items. Public event monitoring watches large gatherings for safety and crowd control. Surveillance analysis monitors footage for security threats and flags what it judges to be suspicious behavior. Those four are not equivalent, and treating them as one category of deployment is how an agency ends up defending a decision it never consciously made.
Document Intelligence
Government agencies process enormous volumes of documents, and multimodal systems address the whole pipeline rather than one step of it. They extract text from scanned documents via OCR, interpret document layout and structure, classify documents by type, extract key information including tables, named fields, amounts, and dates, compare versions of a document to identify changes, summarize content, and answer questions about it.
The mechanism combines a vision model that analyzes the image of the page for layout, structure, and text location, an OCR component that extracts the characters, and a language model that interprets the content. The system reasons about visual structure and textual content together, which is what lets it know that a number sitting in a particular box on a form is an income figure rather than a case number.
Applications include benefits application processing, extracting applicant information and verifying completeness; permit processing, extracting application details and checking them against regulations; tax document processing, extracting and classifying and flagging anomalies; and contract analysis, extracting terms, comparing against templates, and identifying risk clauses. All four share a property worth naming: the extracted data becomes the input to a downstream decision, so an extraction error does not stay an extraction error for long.
Modality-Specific Risks
Different modalities fail in different ways. Treating "AI risk" as one undifferentiated thing is how agencies get surprised. Gloria worked through each.
Vision: misidentification and bias
A vision model can hallucinate, reporting something that is not in the image or missing something that is. On satellite imagery that might mean labeling a parking lot as floodwater because both reflect light similarly, or missing a damaged structure hidden under tree cover. Vision systems also carry bias: models trained on imbalanced data perform unevenly across groups, and facial-recognition and person-detection systems in particular have documented higher error rates for some demographic groups. If the training data was biased, the model may interpret images of certain groups differently, and that difference will not announce itself in the output.
For Gloria the conclusion was blunt. A vision model can support situational awareness, but it must never make an identification of a specific person that triggers a consequence for that person without human confirmation.
Audio: transcription errors in high-stakes settings
Transcription is not neutral. On a clear recording in standard English, a good model might reach 95 percent word accuracy. On a panicked 911 caller with background noise, an accent the model underperforms on, or a street name that sounds like another, accuracy can fall sharply. A single wrong digit in an address, or a misheard "not," can invert the meaning of an emergency report. Poor audio quality degrades everything downstream, because the language model reasons over the transcript rather than the sound. The higher the stakes, the more a human must verify the transcript against other evidence before it drives a dispatch decision.
Video: behavior interpretation, privacy, and authenticity
Video carries three extra hazards. The first is behavior interpretation. A model that flags "suspicious activity" is making a judgment that can be biased and is often a false positive, and in a public-safety context a false positive puts an innocent person under scrutiny. The second is privacy. Analyzing video of people can enable identification of individuals and tracking of their movements and habits across time, including in public spaces, and organizations routinely deploy such systems without working through the implications first.
The third is authenticity. Generative AI can produce convincing fake images, audio, and video, called deepfakes. During a disaster, a fabricated image of a collapsed dam or a faked audio clip of an official ordering an evacuation could spread panic faster than any correction. This is where content provenance matters. C2PA, the Coalition for Content Provenance and Authenticity, is an open technical standard for attaching tamper-evident content credentials to media, recording where it came from and how it was edited. Provenance does not prove that something is true. It lets an analyst distinguish a credentialed image from an emergency aircraft from an anonymous clip pulled off social media. Gloria's rule follows from that limit: media that drives an evacuation order must have verifiable provenance or independent corroboration.
Documents: OCR error and layout misreading
Document pipelines fail in three characteristic ways. OCR extracts text incorrectly, especially on poor-quality scans, misreading a digit or a name. The model misunderstands layout and pulls information from the wrong location on the page. And the system's output ends up inconsistent with what a human reading the same document plainly sees. Because OCR usually sits at the front of the pipeline, its errors do not stay local. They become the ground truth every downstream step reasons from, and teams frequently do not realize the extraction was wrong at all.
The combination risk
The most distinctive multimodal risk is cross-modal misalignment: the model fuses the feeds incorrectly. It hears one caller, sees a different fire, and confidently merges them into a single wrong conclusion. Sometimes it weights one modality too heavily; sometimes it aligns information from two modalities that do not actually correspond. A document image shows a table, OCR misreads some of its numbers, and the system reconciles the visual table with the mistaken text to reach a conclusion inconsistent with both. Because the output looks coherent, this error is the hardest to catch. The only reliable defense is a human checkpoint placed precisely where the fused conclusion would trigger an action.
Three Worked Scenarios
Disaster damage assessment using satellite imagery. A hurricane strikes a coastal region and FEMA needs to assess damage to homes, infrastructure, and land. Traditionally this means sending assessors on the ground, which is slow and dangerous. A multimodal system analyzes before and after satellite images, interprets them alongside street maps and property records, identifies damaged buildings and estimates severity, and flags areas for human assessment. Within hours the system identifies the worst-hit neighborhoods and flags affected schools, hospitals, and power plants, which helps prioritize resources and direct assessors. The challenge is that the system infers from overhead imagery: a building that looks unchanged may be devastated inside, and a debris pile may be temporary. Humans must verify before resource allocation decisions follow.
Benefits application processing. A state agency receives thousands of applications monthly, many incomplete or internally inconsistent, and staff spend their time checking them by hand. A multimodal system analyzes scanned forms, extracts applicant, household, and income information, verifies that required fields are complete, and flags inconsistencies such as one address on one form and a different address on another. Complete and consistent applications route onward, incomplete ones trigger a notice to the applicant, and inconsistent ones go to a human. The challenge is extraction accuracy across handwritten, typed, and poorly scanned forms, understanding which field corresponds to which piece of information, and flagging genuine problems while tolerating trivial variation.
Public meeting monitoring and summarization. A state legislature holds public hearings recorded on video and audio, and staff currently review the recordings by hand. A multimodal system transcribes the audio, interprets the visual channel to see who is speaking and what vote tallies are displayed, combines both into a summary, and identifies key decisions, action items, and contested votes. Staff review and refine, and the summary is published so citizens who could not attend can follow the proceedings. The challenge is that accuracy is not a convenience here: the audio may be unclear with multiple speakers and background noise, the system must understand what vote is being taken and what it means, and it must attribute statements to the right speaker. These summaries function as official records.
The Artifact: A Multimodal Use-Case Risk Register
Gloria's deliverable to her director was not a yes or a no. It was a risk register: one row per use case, naming the modality combination, the intended use, the failure mode that worries her most, the control that reduces it, and the specific human checkpoint that must occur before action. This is the document that lets an agency adopt multimodal AI deliberately rather than by accident. Every threshold that appears in the control column is a number the agency sets for itself and records in advance, not a standard anyone else imposed.
| Modality combination | Intended use | Primary failure mode | Control | Human checkpoint |
|---|---|---|---|---|
| Satellite image plus text | Estimate flood and fire extent for evacuation zoning | Misidentification, such as a parking lot read as water or damage hidden under canopy | Confidence threshold set in advance; require two independent image sources before a high-confidence label | GIS analyst confirms the boundary on a map before any evacuation zone is published |
| 911 audio to text | Transcribe and route emergency calls | Transcription error on an address, a name, or a negation in noisy audio | Flag low-confidence segments; retain the original audio; spell-back of address | Dispatcher verifies the address against the caller before dispatch |
| Meeting audio and video to captions | Section 508 captioning of public briefings | Inaccurate captions misstate official guidance | Human caption review for any live emergency broadcast | Communications staff approve captions before public posting |
| Traffic or drone video plus audio | Detect road washouts and blocked evacuation routes | False positive event; biased suspicious-behavior flag | Use for infrastructure events only, never person identification; log all flags | Operations officer confirms the event before closing a route |
| Scanned forms, image plus text | Extract structured data from damage-assessment paperwork | OCR error or layout misreading propagating into downstream decisions | Flag low-confidence extractions; retain the source image alongside the extracted value | Caseworker verifies names, amounts, and dates against the scan before the record is committed |
| Any incoming media, image, audio, or video | Use third-party media in situational awareness | Deepfake or doctored media triggers a false action | Require C2PA content credentials or independent corroboration | Analyst verifies provenance before media informs any public order |
With the register in hand, Gloria's recommendation was nuanced and defensible. Adopt the system for situational awareness and document understanding, where errors are caught downstream. Do not let it issue evacuation orders, identify individuals, or post public captions without the named human checkpoint.
Fitting Multimodal AI Into Existing Rules
Gloria did not need a new legal regime. Existing frameworks already reach these systems. The NIST AI Risk Management Framework, a voluntary and non-binding framework, treats each use case as something to map, measure, and manage, and its MAP and MEASURE functions ask precisely the two questions the register answers: what is this system used for, and how does it fail. Section 508 governs the captioning use case. If her agency were federal and any of these uses affected people's rights or safety, OMB Memorandum M-24-10 would require an impact assessment and meaningful human oversight before deployment, which is exactly what the human checkpoint column documents. Privacy obligations attach the moment video or audio captures identifiable people, and a privacy impact assessment belongs before deployment rather than after the first complaint.
The register is not a substitute for any of these frameworks. It is the working document that makes complying with them concrete, and when her agency later built its AI inventory, every row became a documented control with an owner attached.
Anti-Patterns
- Trusting multimodal output without verification. These systems reason across text and images, so their conclusions read as well-reasoned and teams accept them. A satellite system labels a building destroyed, the assessment drives a permit denial, and the building turns out to have been standing. Always verify output for high-stakes decisions: have human experts confirm imagery assessments before they drive decisions, spot-check document extractions, and have a person watch the video the system flagged before anyone acts on the flag. Establish escalation paths for unusual outputs, and for high-confidence outputs too, because confidence is not a reason to skip review.
- Ignoring misalignment between modalities. The system weights one modality too heavily or aligns information that does not correspond, reaching a conclusion inconsistent with both the image and the text taken separately. Test on cases where modalities deliberately conflict, verify that the system flags conflicts for human review rather than resolving them silently, have humans review outputs in the conflict zone, and build workflows that allow override when modalities are in tension.
- Privacy violations through video analysis. Behavior analysis, cross-frame tracking, and face identification together enable tracking of people's movements and habits, including in public spaces. Assess privacy impacts before deploying. Disable face recognition where it is not essential. Do not build systems that track specific individuals over time absent a compelling law enforcement need. Be transparent about what the system does, follow the legal requirements that apply to video analysis and surveillance, and put real oversight on it.
- Letting OCR errors propagate downstream. OCR sits at the front of the document pipeline, so a misread digit or misspelled name becomes the ground truth for everything after it, and teams often do not know the extraction was wrong. Verify OCR accuracy before extracted information drives decisions, have humans verify critical fields such as names, amounts, and dates, flag unusually low-confidence extractions for review, build workflows that let humans correct extractions before downstream processing, and monitor downstream outcomes for the signature of systematic extraction error.
- Treating fusion as inherently richer. Reasoning across modalities produces a better answer when the fusion is correct and a less detectable error when it is not. The single confident output is the selling point and the hazard in the same breath.
- Deploying against a category rather than a use case. "AI for emergency response" cannot be governed, controlled, or audited. "Read satellite imagery to estimate flood extent" can.
- Treating a provenance credential as proof of truth. Content credentials tell you where media came from and how it was edited. They do not tell you that what it depicts is accurate or current.
- Treating auto-captions as accessibility compliance. Captions that are wrong do not satisfy the obligation, and emergency broadcasts are exactly where errors matter most.
Practice Prompts
- Multimodal application assessment. Think of a workflow in your agency involving both text and visual information. Could a multimodal system improve it? What exactly would it need to do, and what would be the most critical failure mode?
- Modality conflict. Design a test case in which text and visual information conflict. How should the system handle it? Should it flag the conflict for human review, or should one modality take precedence by rule, and who decides?
- Accuracy verification. For a system processing documents or images, how would you verify accuracy? What error rate would be acceptable for that use case, and how would you sample and check outputs on an ongoing basis?
- Privacy impact assessment. If your agency deployed video analysis, what privacy impacts would it have? What safeguards would you require, who would hold oversight, and what would you disable by default?
- Error propagation. If your agency extracts information from images or documents, what happens when an extraction error reaches downstream systems? How would you detect it, and how would you correct records already committed?
- Build the register. Take one candidate use case and fill in all five columns: modality combination, intended use, primary failure mode, control, and human checkpoint. If you cannot name a checkpoint, you have found your finding.
Reflection
Take two minutes on a multimodal analysis your agency might realistically need: satellite imagery, document processing, or video review. What is the most damaging error the system could make, and who would be harmed by it? Then ask how that error would surface. Would anybody notice, or would the output simply be correct-looking enough to pass?
Now consider what the single fused answer costs you. In a side-by-side workflow, disagreement between feeds is itself a warning. A multimodal system removes that warning by design, in exchange for speed. That is often a trade worth making in an emergency, but it should be a decision someone made deliberately and wrote down, not a property the agency inherited from a demo.
Glossary
- Modality. A mode of perception or communication: text, image, audio, or video. Multimodal systems work across more than one.
- Vision-language model. A system that understands images and text together, answering questions about images, extracting text from them, and reasoning across the two.
- Vision encoder. The component that converts an image into a numerical representation a language model can process alongside words.
- OCR (optical character recognition). Technology that extracts text from images of documents, often the first step in a document pipeline and therefore a common source of propagated error.
- Video understanding. AI that analyzes footage to identify events, track objects across frames, and reason about what is happening visually and audibly.
- Alignment, multimodal. The degree to which understanding derived from different modalities is consistent and mutually reinforcing.
- Cross-modal misalignment. The failure in which a system fuses information from different modalities incorrectly, producing a coherent conclusion that matches none of the inputs.
- Hallucination, multimodal. Output that does not align with the actual content of any modality: seeing things in images that are not there, or reasoning that contradicts both the text and the visual evidence.
- Deepfake. Synthetic image, audio, or video generated to appear authentic.
- Content provenance. Recorded, tamper-evident information about where a piece of media came from and how it was edited. C2PA is an open standard for attaching such credentials.
- Human checkpoint. The named point in a workflow at which a person must confirm a fused conclusion before it triggers an action.
Related Lessons
- How Transformers and LLMs Work explains the language-model half of a vision-language system, including why fluent output is not evidence of correctness.
- Generative AI Deep Dive covers how image and video generation works, which is the other side of the deepfake problem.
- Supervised vs. Unsupervised vs. Reinforcement Learning supplies the learning paradigms underneath the vision and audio components.
- Privacy Impact Assessments for AI Systems is the required process for the video and audio uses described here.
- Minimum Risk Management Practices is where the human checkpoint column becomes a documented compliance obligation.
- NIST AI RMF: MAP, MEASURE, MANAGE gives the framework the risk register is designed to satisfy.
Closing
Agencies will increasingly work with systems that combine text, image, audio, and video, and the pressure to adopt them will come from exactly the situations where careful thought is hardest, namely fast-moving incidents with real consequences. The principle to hold on to is that multimodal systems can reason across modalities in ways no single-modality tool can, and that they are correspondingly more complex and harder to diagnose when they fail. Robust verification and human oversight are not optional add-ons to that capability. They are the condition under which the capability is usable at all.
Gloria did not walk into the briefing with a verdict. She walked in with a table, and the table changed the conversation from whether the system was impressive to which of its outputs anyone would be allowed to act on and who would be standing between the model and that action. Thirty-one minutes was enough because the questions were the right ones, and they are the same five questions for any modality combination an agency will ever be sold.
Key Takeaways
- Multimodal means reasoning across modalities, not running tools side by side. The system fuses text, image, audio, and video internally, which is faster and hides the reconciliation step from you.
- Richer understanding is conditional on correct fusion. When the fusion is right you get more than any single feed could give; when it is wrong you get one confident answer and no dissenting feed to warn you.
- Name the use case before you assess the risk. "AI for response" is ungovernable. "Read satellite imagery to estimate flood extent" is something you can write a control for.
- Each modality fails differently. Vision misidentifies and carries documented demographic bias; audio mistranscribes in noise and inverts meaning on a single missed word; video produces false positives on behavior and can be fabricated; document pipelines misread layout and propagate OCR errors downstream.
- Cross-modal misalignment is the signature multimodal risk. A confidently fused but wrong conclusion is the hardest error to catch, so place a human checkpoint exactly where the fused output triggers action.
- Never let a vision system identify a specific person into a consequence without human confirmation. Situational awareness is a supportable use. Identification that lands on an individual is not, absent a human in the path.
- Provenance is your defense against deepfakes, and it is not proof of truth. C2PA content credentials distinguish credentialed media from an anonymous clip. Media that drives a public order needs verifiable provenance or independent corroboration.
- Video analysis is a privacy decision before it is a technical one. Assess impacts first, disable face recognition where it is not essential, and do not track individuals over time without a compelling law enforcement need.
- Auto-captions do not automatically satisfy Section 508. Accessibility requires accurate captions, so emergency broadcasts need human caption review before posting.
- The risk register is the deliverable. Modality combination, intended use, failure mode, control, and human checkpoint in five columns turn an impressive demo into a deployment your agency can defend to an auditor.
Frequently Asked Questions
Is a multimodal system just several single-modality tools bundled together? No. Bundled tools produce separate outputs that a human reconciles, which means disagreement between them is visible to you. A multimodal system reconciles internally and returns one answer, so the reconciliation is faster and the error, when it happens, has no visible contradiction attached to it.
Where should the human checkpoint go? Exactly where the fused conclusion would trigger an action: before an evacuation zone is published, before a route is closed, before captions go public, before a record is committed, before media informs a public order. Placing it earlier wastes reviewer capacity; placing it later means the action already happened.
How accurate is automated transcription? On a clear recording in standard English a good model might reach 95 percent word accuracy. On noisy audio, a distressed caller, or an accent the model underperforms on, accuracy falls sharply, and a single missed negation or wrong digit can invert an emergency report. Treat the number as a best case, not a specification.
Do content credentials tell us a piece of media is genuine? They tell you where it came from and how it was edited, in a tamper-evident way. That distinguishes credentialed media from an anonymous clip, which is enormously useful during an incident. It does not tell you that what the media depicts is accurate, current, or relevant.
Can we use these systems for surveillance? Video analysis of people raises privacy obligations that attach the moment identifiable individuals are captured, and behavior flags carry documented bias and frequent false positives. Assess privacy impacts first, disable face recognition unless it is essential, avoid tracking individuals over time absent a compelling law enforcement need, and be transparent about what the system does.
Our OCR is very accurate. Do we still need human verification? For critical fields, yes. OCR sits at the front of the pipeline, so its errors become the ground truth for every downstream step and are rarely noticed at the point they occur. Verify names, amounts, and dates against the source image, flag low-confidence extractions, and monitor downstream outcomes for the signature of systematic error.
Do we need a new policy framework for multimodal AI? Generally not. The NIST AI Risk Management Framework, which is voluntary, maps and measures each use case; Section 508 governs captioning; OMB minimum practices require impact assessment and meaningful human oversight for rights-affecting federal uses; and privacy law attaches when identifiable people are captured. The risk register makes those obligations concrete rather than replacing them.
What if the modalities disagree with each other? That is the case to test for deliberately before deployment. The system should surface the conflict for human review rather than resolving it silently, and your workflow should let a person override the fused answer when the feeds are in tension.
Skill.re