Where You Already Use It: ChatGPT, Gemini, Claude
You do not need a special app to try multimodal AI. If you have opened ChatGPT, Google's Gemini, or Anthropic's Claude, you have already been using tools built on the ideas from this course. The little paperclip, camera, or microphone icon you have probably ignored is the door to everything we have covered.
This lesson is a concept-level tour. Exact buttons and features change often, so we focus on the pattern you will see across all of them rather than a click-by-click guide that would be stale in a month.
What You'll Learn
- The common multimodal pattern shared by the big chat tools
- What ChatGPT, Gemini, and Claude each let you do, at a concept level
- Why you should always check the current app for exact features
- How to spot the multimodal features in any tool you open
The pattern is the same everywhere
Under the branding, these tools work the same way. Each is a chat box where the model on the other side is multimodal. To use more than text, you look for a way to attach or capture something.
- Open the chat
- Find the attach or mic icon
- Add an image or speak
- Type or say your instruction
- Get a blended answer
Learn this pattern once and every tool feels familiar. Look for a paperclip or plus for files and photos, a camera for a live picture, and a microphone or waveform for voice.
The big three, at a concept level
As of mid-2026, all three of the most popular chat assistants are multimodal. Capabilities and plan limits shift often, so treat this as the general shape, not a spec sheet, and check the app for what is live today.
ChatGPT (OpenAI). You can upload images and screenshots and ask about them, and it has a real-time voice mode where you speak and it answers out loud. It can see, hear, and speak, which makes it a good all-round place to try every modality in one chat.
Gemini (Google). Gemini was built to be multimodal from the start, handling text, images, and audio together, with voice interaction as well. It also ties into Google's own apps, so multimodal help shows up in places you already work.
Claude (Anthropic). Claude can accept image uploads and answer questions about them, and it added a voice mode that lets you speak with it. It is a strong choice when you want to reason carefully over a document, a screenshot, or a diagram you have shared.
Multimodal chat assistants at a concept level, as of mid-2026 (check the app for current details)
| Criteria | ChatGPT | Gemini | Claude |
|---|---|---|---|
| Understands images | Yes | Yes | Yes |
| Voice conversation | Yes | Yes | Yes |
| Handy for | All-round, see and speak | Tied into Google apps | Careful reasoning on shared files |
ChatGPT
- Understands images
- Yes
- Voice conversation
- Yes
- Handy for
- All-round, see and speak
Gemini
- Understands images
- Yes
- Voice conversation
- Yes
- Handy for
- Tied into Google apps
Claude
- Understands images
- Yes
- Voice conversation
- Yes
- Handy for
- Careful reasoning on shared files
Why "check the app" is real advice, not a cop-out
AI tools change fast. Features roll out to some countries or plans before others, free and paid tiers differ, and names shift between updates. A course that promised "click the third button on the left" would be wrong within weeks, and following stale steps is frustrating.
So the durable skill is not memorizing today's menu. It is knowing what to look for. Once you understand that any of these tools is one multimodal model behind a chat box, you can open a fresh update, spot the attach and microphone icons, and get going without a tutorial.
How to find multimodal features in any tool
Next time you open an AI assistant, do this quick scan:
- Look near the message box for a paperclip, plus, or camera icon. That is for images and files.
- Look for a microphone or waveform icon. That is for voice.
- If you see them, the model is multimodal. Attach something, add an instruction, and try it.
- If your free plan blocks a feature, the app will usually say so. That is a plan limit, not a missing capability.
Key Takeaways
- ChatGPT, Gemini, and Claude are all multimodal as of mid-2026: they understand images and support voice conversation.
- They share one pattern: open the chat, find the attach or microphone icon, add your input, and give an instruction.
- Exact features, buttons, and plan limits change often, so always check the current app rather than a fixed set of steps.
- The durable skill is recognizing the multimodal icons in any tool, not memorizing today's menu.

