Seeing: Ask AI About a Photo or Diagram
This is the capability that surprises people the most. You can hand an AI an image, ask a question about it in plain words, and get a useful answer. No special skills. No editing tools. Just a photo and a question.
This lesson covers what "seeing" means for a multimodal model, the kinds of image tasks that actually work well, and how to ask so you get a good answer.
What You'll Learn
- The four most useful things you can do with an image and a question
- How to combine an image with a text instruction
- The difference between AI reading an image and AI creating one
- Simple ways to ask so you get better results
Four things you can do with an image today
When people say a model can "see," they mean you can upload or photograph something and ask about it. As of mid-2026, these four tasks are reliable enough to use every day.
1. Ask about a photo. Snap a picture of anything and ask a question. What plant is this? Is this outfit formal enough? What does this warning light on my dashboard mean? The model looks at the image and answers in words.
2. Turn a diagram or chart into an explanation. Upload a diagram from a textbook, a flowchart, or a chart from a report and ask "explain this to me simply." Instead of decoding it alone, you get a plain-language walkthrough.
3. Pull text out of an image. Photograph a menu, a receipt, a whiteboard after a meeting, or a page of notes and ask the model to type it out, translate it, or summarize it. The model reads the text inside the picture.
4. Combine an image with an instruction. This is where it gets powerful. You give the model both a picture and a task in words, and it uses them together.
Combining an image with a text instruction
The single most useful pattern is image plus instruction. The image supplies the details. Your words supply the goal.
- ImageA photo of your fridge
- InstructionSuggest 3 dinners
- Model connects both
- Answer3 recipes from what it sees
A few real examples of the pattern:
- Photo of a math problem plus "explain each step, do not just give the answer."
- Screenshot of an error message plus "what is causing this and how do I fix it?"
- Picture of your bookshelf plus "recommend which of these to read first for a beginner."
- Photo of a nutrition label plus "is this a good choice if I am watching sugar?"
In every case the model is doing what you learned in the last lesson. It turns the image and your words into numbers in the same shared space, then lets your instruction shape how it reads the picture.
Reading an image is not the same as making one
This is a common point of confusion, so let us be clear.
- Reading an image means you give the model a picture and it understands it. That is what this lesson is about, and it is a core multimodal skill.
- Making an image means you give the model words and it generates a brand-new picture. That is a different capability called image generation.
Many tools can do both, but they are separate abilities. This course is about understanding the concept of a model that can read across senses. If your goal is to create images from a prompt, that is a whole craft of its own, and the free course AI Image Generation for Beginners is built for exactly that. If you want to understand how a computer recognizes objects in a picture under the hood, see Computer Vision Basics.
How to ask so you get a good answer
The image does half the work. Your words do the other half. A few simple habits make a big difference.
- Say what you want done, not just "what is this?" Compare "describe this chart" with "in two sentences, what is the main trend in this chart?" The second gets a sharper answer.
- Give context the image cannot show. "This is a receipt from a work trip. List only the food items and their prices." The model cannot know it was a work trip unless you say so.
- Ask one clear thing at a time. A pile of five questions about one photo often gets a muddled reply. Ask, read, then follow up.
- Use a clear image. Good light, straight angle, in focus. A blurry photo of tiny text is the number one reason the model gets it wrong. More on that in the limits lesson.
Key Takeaways
- You can ask about a photo, explain a diagram, pull text out of an image, and combine an image with an instruction.
- The most powerful pattern is image plus instruction: the picture gives details, your words give the goal.
- Reading an image (understanding it) is different from making an image (generating one). This course is about reading.
- To create images, take AI Image Generation for Beginners; this course covers the unifying concept.
- Ask for a specific action, add context the image cannot show, and use a clear, well-lit photo.

