Skip to content
LLMs & Agents

What can you actually build with multimodal AI in 2026?

Document and form extraction, screenshot and UI understanding, image inspection, chart reading, audio transcription feeding a text model, and video summarisation. What each looks like, where multimodal models fail (small text, counting, spatial reasoning, confident wrong extraction), how to test on your own images, and the cost and privacy questions.

Nikhil De Silva · Founder, Square 1 AI6 min read

With multimodal AI you can build products that read documents and forms into structured data, understand screenshots and user interfaces, inspect photos for defects or damage, turn charts into tables, transcribe audio and feed it to a text model, and summarise video. These work well enough to ship today, provided you design around the failure modes: models misread small text with confidence, miscount objects, confuse spatial relationships and return extractions that look valid but are wrong. The products that succeed validate every output, measure accuracy on their own images and send the uncertain cases to a person.

This piece is about practical use cases rather than model rankings. Capabilities, prices and limits change often, so test current models on your own material before you commit.

What does "multimodal" mean in practice?

A multimodal model takes more than text as input, most commonly images such as photos, scans and screenshots, and in some cases audio or video, and answers in text. Some also generate or edit images. For a developer, an image becomes just another part of the prompt: you send the picture with an instruction such as "extract the supplier, date and total as JSON", and you get structured text back that your code can use.

That is different from classic computer vision, where you train a model for one fixed task such as detecting a particular object. Our background piece on AI image recognition covers how those vision models work. The two approaches overlap, and choosing between them is part of the job.

What can you actually build?

Use case What it looks like What to watch
Document and form extraction Invoices, receipts, applications and scanned forms turned into validated fields Misread digits, missed fields on dense or multi-page documents
Screenshots and UI understanding Triage of bug reports with screenshots, UI checks in testing, agents that read a screen Small interface text, which button belongs to which label
Image inspection and QA Checking product photos against listing rules, triaging site or damage photos for a person to review Rare defects the model has not seen, lighting and angle
Charts and diagrams Reading a chart into a table, describing a diagram for search or accessibility Precise values read off an axis, overlapping series
Audio to text Transcribing calls or meetings, then summarising, tagging or extracting actions with a text model Names, jargon and accents; errors in the transcript carry into the summary
Video summarisation Sampled frames plus the transcript, summarised or searched for moments Cost and latency, events between sampled frames

The strongest early projects share a pattern: a high-volume, boring task a person does by looking at something, where a wrong answer is caught cheaply.

Where does multimodal AI fail?

In ways that look like success, which is what makes them dangerous:

  • Small or low-quality text. Faced with tiny print, a blurry scan or a photo taken at an angle, a model will often produce a plausible reading rather than admit it cannot see. A total becomes a slightly different total.
  • Counting. Models are unreliable at counting many similar objects, such as items on a shelf or people in a crowd.
  • Spatial reasoning. Which label sits next to which box, which value belongs to which table cell, what is left or right of what: these relationships are weaker than the model's fluency suggests.
  • Confident wrong extraction. The output passes your schema, every field is filled, and one of them is wrong. Validation catches format errors, not reading errors.
  • Transcription errors that compound. A misheard name in an audio transcript becomes a confident fact in the summary built on it.

The defences are ordinary engineering. Validate against rules the document must obey, such as line items that sum to the total or dates that parse. Give the model an explicit way to say a field is unreadable, and treat that as a result rather than a failure. Crop and enlarge the region that matters instead of sending a whole page. Use a second method, such as a separate OCR pass, for fields where errors are costly, and send disagreements to a person. For counting, detection and anything that must run fast at high volume, a dedicated computer vision model is often the better tool.

How do you test a multimodal feature?

On your own images, never only on demos. Build a labelled set from real inputs, and include the bad ones on purpose: blurred, rotated, cropped, glare, handwriting, the unusual layout from the one supplier who does things differently. Record the correct answer for each field.

Then measure field by field, not document by document, because "most documents mostly right" hides the field that is always wrong. Compare models and prompts on the same set, rerun it on every change, and track the cases flagged as uncertain as well as the wrong ones. Our guide to evaluating LLM outputs covers golden sets and scoring.

What about cost, latency and privacy?

Cost and latency. Images, audio and especially video take more processing than a short text prompt, so they cost more and take longer. Resize images to the resolution the task needs, crop to the region that matters, send only the pages that contain what you want, and sample video frames rather than sending every one. Back-office work can run in batches; interactive features need streaming and a clear loading state. Measure the cost per document on your real inputs before you promise a price or a turnaround.

Privacy. Images carry more than they appear to. A photo can include faces, a licence plate or location data in the file's metadata; a screenshot can include other people's messages or customer records. Strip metadata, crop or redact what the task does not need, get consent where people are identifiable, set a retention period, and read the provider's data terms against your obligations. Medical, identity and children's images need particular care and often a separate review. Our piece on data privacy in AI systems covers the wider picture.

Where should you start this week?

  • Pick one task in your work where someone reads a document or image and types what they see.
  • Collect a small set of real examples, including the messy ones, and write down the correct fields for each.
  • Send them to a multimodal model with a schema and an explicit "unreadable" option, and score the results field by field.
  • Add one validation rule the document must satisfy and see how many errors it catches.
  • Note which fields fail most, and decide whether they need a person, a crop, or a different tool.

Where does Square 1 teach this?

Multimodal AI is an on-demand course recorded by an instructor and graded by Nova, the AI tutor: about seven and a half hours across five modules, for developers in Python who have made at least one LLM API call. The graded work is a document understanding feature with its extraction accuracy measured, a transcription pipeline with its word error rate reported, an image generation feature with a safety check, and a deployed multimodal product. Building with the Claude API includes a module on files, vision and documents with a graded document question-answering feature. The Computer Vision Bootcamp is twelve weeks, live on Zoom with one instructor, about 15 hours a week, from image datasets and classifiers through detection and a vision-language block with grounding accuracy measured to a real-time pipeline with monitoring; it asks for Python and the machine learning toolkit. All three are taking a waitlist today. The computer vision engineer role page describes the work when vision becomes the job.

Questions people ask

What can you build with multimodal AI?

Common production uses are extracting fields from documents and forms, understanding screenshots and user interfaces, inspecting photos for defects or damage, reading charts into tables, transcribing audio for a text model to summarise, and summarising video from sampled frames and transcripts.

Where do multimodal models fail?

They misread small or blurry text with confidence, count many similar objects unreliably, confuse spatial relationships such as which value belongs to which cell, and return extractions that pass a schema but contain a wrong value.

How do you test a multimodal feature?

Build a labelled set from your own real images, deliberately including blurred, rotated and unusual ones, and measure accuracy field by field rather than per document. Rerun the set on every model or prompt change.

When should you use a dedicated computer vision model instead?

For counting, object detection and fixed tasks that must run fast at high volume or on a device, a trained computer vision model is often more reliable and cheaper than a general multimodal model.

What privacy risks do images bring?

Images can contain faces, addresses, licence plates, other people's messages or location metadata. Strip metadata, crop or redact what the task does not need, get consent where people are identifiable, and set a retention period.

Free skill check · about 3 minutes

Where do you stand on Generative AI?

Five questions, and a skill breakdown the moment you finish: your strengths, the gaps to close, and what to learn next from real curriculum.

Start the Generative AI skill check

Free, with a student account — the check is the first entry in your record.

Learn this by building it

The programmes that teach what this piece covers, each ending in deployed work graded against a rubric you can read.