Multi-Modal Prompting: Text, Images, and Documents
Nkechi Al-Rashidi manages procurement for a government infrastructure agency in the Gulf region. Her team reviews hundreds of vendor contracts each year: dense PDFs, sometimes scanned, sometimes in mixed Arabic and English. For years, extracting key terms for comparison meant hours of manual work per contract. Last year, she started uploading the PDFs directly to a multi-modal AI assistant. Now her team gets a first-pass extraction in under two minutes. "I kept thinking you could only type at these things," she said. "Nobody told me you could just hand it the document."
Multi-modal AI, meaning systems that accept and reason across more than one type of input, is one of the most underused capabilities in enterprise AI adoption. Most professionals have learned to type prompts. Fewer have learned to combine text with images, scanned documents, spreadsheets, or slides, which means they retype or summarize information the model could have read directly. This lesson closes that gap: what the capability is, the five use cases where it pays off fastest, how to structure prompts that include attachments, and where the verification risk sits.
What Multi-Modal Means
A multi-modal AI system is one that can process more than one type of input, or output, in a single interaction. The most common combination in today's enterprise tools is text plus images or documents. The important part is the phrase "in a single interaction": the model reasons across all the inputs simultaneously rather than handling them one at a time, which is what makes questions that span an attachment and your instructions answerable at all.
For practical purposes, the modal types you are most likely to use are these:
- Text: Typed prompts, pasted content, system instructions
- Images: Photos, screenshots, charts, diagrams, scanned forms
- Documents: PDFs, Word files, spreadsheets, which may contain text, tables, and embedded images simultaneously
- Audio and video: Emerging in some platforms, less common in current enterprise workflows
When you upload an image or a document, the model processes its visual and textual content together with your text prompt. That opens moves that text-only prompting cannot make. You can ask questions about what is in the image. You can request specific information from a document. You can ask the model to compare a screenshot against a written description. And you can combine multiple inputs and ask the model to synthesize across them, which is where most of the compounding value sits, because reconciling two sources by hand is exactly the kind of work that consumes a morning.
Five High-Value Multi-Modal Use Cases
1. Document Extraction and Comparison
This is Nkechi's use case: uploading structured documents and asking for specific information such as key dates, obligations, pricing terms, and exceptions. For dense documents like contracts, RFPs, or technical specifications, multi-modal AI can cut initial review time by 60% to 80% for experienced users. The extracted output still requires human verification, but the model finds the relevant passages and structures them faster than manual search, which changes the shape of the work from hunting to checking.
A prompt structure that works looks like this:
"Here is a vendor contract [upload PDF]. Extract the following fields into a table: contract duration, renewal terms, termination notice period, payment schedule, liability cap, and governing jurisdiction."
Two things in that prompt are doing the work. The field list is enumerated rather than described, so the model has an explicit checklist instead of a general instruction to find important terms. And the output format is named, which is what makes the results from a batch of contracts comparable to each other rather than a set of differently-shaped summaries you then have to normalize by hand.
2. Chart and Data Visualization Interpretation
Take a screenshot of a chart or dashboard and paste it into a prompt with a question about what you are seeing. The model can read axis labels, approximate values, identify trends, and surface anomalies. This is genuinely useful when you have a chart from a system you cannot export data from, which describes a great many internal dashboards and vendor portals, or when you want a second read on a visualization before you present it in a meeting.
This use case works best when you pair the image with a specific question rather than asking the model to "describe this chart," because an open request produces an inventory of what is visible rather than an answer. Specificity helps:
"What is the trend in Q3 versus Q2, and are there any months where the trend reverses?"
3. Form and Template Filling from Source Documents
Upload a source document, such as a past proposal, a case file, or a client intake form, alongside a blank template, and ask the model to populate the template with the relevant information. This is especially useful for organizations that produce many similar documents, including grant applications, project briefs, and onboarding checklists, where the structure is fixed but the content varies. The value here is not writing quality; it is that the tedious transcription step between two documents, which is where transposition errors usually enter, gets done in one pass that you then check.
4. Image-Based Quality Assessment
For teams that work with physical inspection, product quality, or visual compliance, multi-modal AI can serve as a first-pass assessment engine. Upload photos and describe what you are looking for: damage, compliance markers, specific features. The model will not replace expert inspection, and it should not be positioned as though it might, but it can triage a volume of images and flag the ones that need human attention. That reframing matters, because triage tolerates a certain rate of false positives while a verdict does not.
5. Mixed-Input Synthesis
Combine multiple input types in one prompt. Paste meeting notes as text, upload the slide deck from that meeting as images, and ask the model to produce a reconciled summary of decisions and action items. Or upload both a contract and an email thread and ask the model to identify any discrepancies between what was agreed in the contract and what was discussed subsequently. Discrepancy-finding across sources is the strongest version of this pattern, because it asks the model to do something reading is bad at, which is holding two documents in mind at once and noticing where they disagree.
Prompt Structure for Multi-Modal Inputs
Multi-modal prompts follow the same principles as text prompts but require a few additional habits. None of them are complicated, and all of them address a failure that only appears once an attachment is involved.
Reference the attachment explicitly. Do not assume the model knows why you uploaded a file. Say: "I've attached a PDF of our Q2 sales report. Please..." The model handles the attachment either way, but an explicit reference ensures it processes that attachment in the right context, particularly when there is more than one file in the conversation and the model has to work out which one your question refers to.
Specify the output format. For document extraction, ask for tables or structured lists. For image interpretation, ask for bullet points. Structure in the output request produces structure in the output, and the effect is more pronounced with attachments than with plain text prompts, because the source material is unstructured enough that the model has no default shape to fall back on.
Split complex multi-modal tasks. If you are asking the model to read a 40-page document and a 12-slide deck and synthesize across both, break the task into steps. First extract the key points from each source separately, then ask for the synthesis in a follow-up prompt using those extractions. This reduces the risk of the model glossing over important sections under context window pressure, and it has a second benefit: you can check the intermediate extractions, which means a synthesis error becomes traceable to a source rather than being a black box.
Specify what to ignore. Documents routinely contain headers, footers, boilerplate, and page numbers that are not relevant to your question and that compete for the model's attention. Tell it directly: "Ignore cover pages, headers and footers, and signature blocks." This is a small instruction with a disproportionate effect on scanned documents, where repeated page furniture can otherwise appear in the extraction as though it were content.
Quality and Verification Considerations
Multi-modal outputs introduce one additional verification risk beyond standard AI output review, and it is worth stating precisely because it is easy to under-rate. The model's reading of visual content can be wrong in ways that are not immediately obvious. A number read from a blurry table, a term extracted from a dense clause, a data point taken from a low-contrast chart: these can be misread without any accompanying signal of uncertainty. The output does not look less confident when the source was harder to read, which is exactly the opposite of what a human reviewer would produce under the same conditions.
The verification habit for multi-modal outputs is therefore to check at least the most consequential extracted facts against the source document directly, rather than sampling at random or trusting the whole extraction because part of it checked out. For high-stakes extractions, including contract values, regulatory requirements, and financial figures, treat the AI output as a finding-aid that accelerates your review rather than as a source of truth in itself. The distinction is practical, not philosophical: a finding-aid tells you which page to look at, and you still look.
Multi-modal AI is most powerful when it compresses the time you spend finding information, not when it replaces your judgment about that information.
Anti-Patterns
The failures below are all versions of treating an attachment as though it were pasted text, or treating an extraction as though it were a reading.
- Uploading without saying why. Attaching a file and asking a bare question leaves the model to infer the connection, which gets unreliable as soon as more than one file is in play. Name the attachment and its role.
- Asking the model to "describe this chart." Open-ended requests on images return an inventory of visible elements. A specific comparative question returns an answer you can act on.
- Requesting no output format. Extractions without a named format come back in whatever shape the source suggested, which makes comparison across documents a manual normalization job.
- Feeding everything in at once. A long document and a deck in a single prompt invites the model to skim under context window pressure, and leaves you no intermediate output to check when the synthesis is wrong.
- Letting page furniture into the extraction. Headers, footers, cover pages, and signature blocks are noise in almost every extraction task and should be excluded explicitly.
- Treating extraction as verification. The most damaging pattern. Visual misreads arrive with the same confident tone as correct reads, so consequential figures must be checked against the source, not merely against the plausibility of the output.
- Positioning image assessment as inspection. Multi-modal triage flags images for human attention. Presenting it as an expert verdict transfers accountability to a system that was never designed to carry it.
Practice Prompts
Run these against real material from your own work. Multi-modal technique is easy to read about and only becomes reliable once you have seen how your own documents behave.
- Field extraction. Take a real contract or specification and write an extraction prompt that enumerates the fields you need and names the output format, following the vendor contract pattern above. Then run it on a second document and check whether the two outputs are directly comparable.
- Chart interrogation. Screenshot a dashboard you cannot export data from. Ask it once as "describe this chart" and once as a specific comparative question about a trend and its reversals. Compare how usable the two answers are.
- Template population. Upload a past proposal alongside a blank template for the same document type and ask the model to populate it. Note every field it got wrong and what about the source made that field ambiguous.
- Split versus single-shot. Take a long document and a slide deck. Run the synthesis as one prompt, then run it again as separate extractions followed by a synthesis step. Compare what the single-shot version omitted.
- Exclusion instruction. Re-run an extraction on a scanned document with and without an explicit instruction to ignore cover pages, headers and footers, and signature blocks, and look at what changes.
- Verification drill. Take a completed extraction and check the most consequential figures against the source document directly. Record how many required correction; that rate is your local evidence for how much verification your context needs.
Reflection
Think about the documents and images you currently retype, summarize by hand, or re-read because you cannot export the underlying data. How much of that work exists only because you have been prompting these tools by typing at them? Then consider the other direction. Where in your work would a confident but wrong reading of a number, a clause, or a data point actually cause harm, and what would it take for you to notice? Nkechi's team moved from hours per contract to a first-pass extraction in under two minutes, and the part of the job that remained was verification, not reading. That is the trade this capability offers: it compresses finding, and it leaves judgment exactly where it was.
Glossary
- Multi-modal: An AI system that can process more than one type of input, or output, in a single interaction, reasoning across all inputs simultaneously.
- Modality: A category of input or output. The types most relevant in current enterprise work are text, images, documents, and, less commonly so far, audio and video.
- Document extraction: Pulling specific enumerated fields out of a structured document such as a contract, RFP, or technical specification, typically into a named output format.
- Mixed-input synthesis: Combining input types in one prompt, such as text notes plus slide images, and asking for a reconciled result across them.
- Discrepancy finding: Asking the model to identify where two sources disagree, for example a contract and a later email thread.
- Context window pressure: The condition in which the volume of supplied material makes it more likely that the model skims or omits sections, and the reason for splitting complex multi-source tasks into steps.
- Finding-aid: An output treated as a way of locating the relevant source passages quickly rather than as an authoritative statement of what those passages say.
- First-pass assessment: Automated triage of a volume of images or documents to flag items needing human attention, distinct from expert inspection or a final verdict.
Related Lessons
Multi-modal prompting builds on general prompting technique rather than replacing it. Anatomy of an Effective Prompt covers the structure that the attachment habits in this lesson extend, and Core Prompting Patterns supplies the patterns you will be applying to document and image inputs. Because extraction outputs need checking rather than trusting, Building Quality Rubrics for AI Outputs is the natural companion for defining what verified means in your context, and Building a Personal Prompt Testing Workflow helps you find out which of your multi-modal prompts hold up across documents before you rely on them.
Closing
The gap Nkechi described is common and almost entirely a matter of habit. People learn to type at these tools, and then keep typing, summarizing documents by hand for a system that could have read them. Closing that gap requires very little: know which modalities your tools accept, aim them at the five use cases where the payoff is largest, name your attachments and your output format, split the big jobs into checkable steps, and tell the model what to ignore. Then keep the verification discipline that the capability makes more rather than less important, because the one thing a confident extraction never tells you is that the source was hard to read.
Key Takeaways
- Multi-modal AI processes text, images, and documents together. You can combine these in a single prompt to get richer, faster analysis than text alone provides.
- The five highest-value use cases are document extraction and comparison, chart interpretation, form-filling from source documents, image-based triage, and mixed-input synthesis.
- Reference your attachment explicitly in the prompt, specify the output format, and split complex multi-source tasks into steps to maintain quality.
- Enumerate the fields you want rather than describing them generally, so that extractions from different documents come back directly comparable.
- Specify what to ignore, including cover pages, headers, footers, and signature blocks, to keep the model focused on content that matters in dense or scanned documents.
- Multi-modal extraction accelerates review. It finds and structures information faster than manual search, but consequential facts still need direct verification against the source.
- Visual content can be misread without uncertainty signals. Numbers, terms, and data points taken from images or PDFs require spot-checking, especially in high-stakes contexts.
Frequently Asked Questions
Does the model actually read a scanned PDF, or just the text layer? Multi-modal systems process the visual content of a document alongside any text, which is why scanned and image-heavy files can be handled at all. That is also precisely where misreading risk concentrates, so treat extractions from scanned sources as needing more verification rather than less.
How much should I upload at once? Less than you are tempted to. A long document and a deck submitted together invite skimming under context window pressure. Extract from each source separately, check those intermediate outputs, then ask for the synthesis in a follow-up prompt.
Can I trust the numbers it pulls out of a chart? Treat them as approximations to be confirmed. The model can read axis labels, approximate values, and identify trends, but a value taken from a low-contrast or blurry visualization can be wrong with no signal that it is. For anything consequential, go back to the source.
Is this suitable for regulated or high-stakes documents? It is suitable as a finding-aid that accelerates review, not as a source of truth. For contract values, regulatory requirements, and financial figures, the extraction tells you where to look and a person confirms what is there.
What about audio and video inputs? They are emerging in some platforms but remain less common in current enterprise workflows. The habits in this lesson, naming the input, specifying output format, splitting large tasks, and verifying consequential facts, transfer directly when those modalities become part of your toolset.
Skill.re