# What is multimodal AI?

> Multimodal AI understands more than text, like photos, PDFs or audio. What it means, a real receipt-photo example in n8n, and the mix-ups to avoid.

Source: https://learn.ynteractive.com/content/glossary/multimodal-ai · Updated 2026-10-01 · Free from Stepthrough (https://learn.ynteractive.com)

[AI basics](https://learn.ynteractive.com/content/glossary#topic-ai-basics) · [AI glossary](https://learn.ynteractive.com/content/glossary)

Updated October 1, 2026 · By [Robert Breen](https://learn.ynteractive.com/content/about)

Multimodal AI is AI that can take in, and sometimes produce, more than one kind of input, such as text, images, audio or video. A multimodal model can look at a photo of a receipt or a screenshot and answer questions about it, not just read typed words.

## Why it matters for a small business

Plenty of business information never arrives as clean text: receipt photos, pictures of a damaged part, a screenshot of an error, a scanned form, a voicemail. Multimodal models let you hand those straight to AI instead of retyping them first. That removes the most tedious step in many office jobs.

It also changes which model you pick. Not every model accepts images, and those that do vary in how reliably they read them. When a workflow needs to see something, model choice stops being only about cost and becomes about capability, and checking the AI's reading against the original becomes part of the routine.

## In a real lesson: AI Receipt Extractor: Receipts to Google Sheets with n8n

The [AI Receipt Extractor lesson](https://learn.ynteractive.com/content/ai-receipt-extractor-n8n-google-sheets) is built around a multimodal model. You start the n8n workflow with **On chat message**, because n8n's chat lets you attach files, then add an AI Agent. When you set up the **OpenAI Chat Model**, you open the **Model** dropdown and change it to **gpt-4o**. The voice-over explains the switch: the default is fine for text, but this agent has to read a picture, and gpt-4o reads images reliably and quickly.

To test it, you click **Open chat**, type "Heres a receipt.", click the **paperclip**, choose **receipt.png** (the photo, not the PDF) and send. The model looks at the image, finds the date, vendor and total, decides it is groceries, and calls the Google Sheets Tool to add a row to the **Invoices** sheet.

The lesson ends with a check you should always keep: click the receipt in your message to view it, confirm Corner Mart, April 24 and $31.57, then compare with the new row.

[Try this lesson free](https://learn.ynteractive.com/modules/receipt-extractor) or [read the step-by-step guide](https://learn.ynteractive.com/content/ai-receipt-extractor-n8n-google-sheets).

## Common confusions

### Multimodal AI vs OCR

[OCR](https://learn.ynteractive.com/content/glossary/ocr) turns text in an image into characters and stops there. A multimodal model can read that text and also understand the picture: which number is the total, what kind of store it is, what category fits.

### Multimodal vs generating images

Multimodal usually describes what a model can understand. [AI image generation](https://learn.ynteractive.com/content/glossary/ai-image-generation) is about creating pictures. Some products do both, but reading a photo and drawing one are separate abilities.

## Tips

- Check that the model you select accepts images before you build a workflow around photos.
- Clear, well-lit, straight-on photos give noticeably better readings.
- Always compare extracted numbers with the original image for the first batch of runs.

## Related terms

[OCR (optical character recognition)](https://learn.ynteractive.com/content/glossary/ocr) · [AI model](https://learn.ynteractive.com/content/glossary/ai-model) · [Data extraction](https://learn.ynteractive.com/content/glossary/data-extraction) · [Large language model (LLM)](https://learn.ynteractive.com/content/glossary/large-language-model) · [AI image generation](https://learn.ynteractive.com/content/glossary/ai-image-generation)

## Where you use it: free lessons

- [AI Receipt Extractor: Receipts to Google Sheets with n8n](https://learn.ynteractive.com/content/ai-receipt-extractor-n8n-google-sheets) (n8n, 12 min)
- [Summarize a Meeting Transcript with ChatGPT](https://learn.ynteractive.com/content/chatgpt-summarize-meeting-transcript) (ChatGPT, 9 min)

## Visual guide

Multimodal AI in a few slides, with the same guide written out as text.

- [5 Steps to Turn a Receipt Photo into a Google Sheets Row](https://learn.ynteractive.com/content/visual/5-steps-receipt-photo-to-google-sheets): A visual guide: an n8n AI agent reads a receipt photo, pulls out the date, vendor and total, picks a category and adds a row to Google Sheets.

[All visual guides](https://learn.ynteractive.com/content/visual)

## Frequently asked questions

**Can ChatGPT read photos and PDFs?**

Yes. In ChatGPT you can attach images and files with the plus button, and the models behind it can read them. In the meeting lessons, a transcript is attached as a Word file the same way.

**Which OpenAI model should I use for images in n8n?**

The receipt lesson uses gpt-4o because it reads images reliably. Check OpenAI's documentation for which current models accept image input before choosing another.

[All AI glossary terms, A to Z](https://learn.ynteractive.com/content/glossary) · [Free prompt templates](https://learn.ynteractive.com/content/prompts)
