Hearing and Speaking: Voice and Audio
Typing is not always the fastest way to work with AI. Sometimes your hands are busy, or it is easier to just talk. Multimodal AI adds sound to the mix, so you can speak to a model and, in many tools, have it speak back.
This lesson covers what "hearing" and "speaking" mean for a multimodal model, the audio tasks that work well, and when talking beats typing.
What You'll Learn
- What it means for a model to take audio in and give audio out
- The main things you can do with voice today
- Why a real voice conversation feels different from typing
- The line between this concept course and hands-on voice courses
Audio in, audio out
Sound is just another modality. As you learned earlier, the model turns audio into numbers, the same way it turns text and images into numbers. Once your voice is numbers in that shared space, the model can work with it alongside everything else.
There are two directions:
- Audio in: you speak or upload a recording, and the model understands it. This is how voice input and transcription work.
- Audio out: the model replies with a spoken voice instead of, or alongside, text. This is how a model can read an answer aloud or hold a spoken conversation.
A truly multimodal tool can do both, which is what makes a back-and-forth voice chat possible.
What you can do with voice today
As of mid-2026, these audio tasks are common and dependable.
Talk instead of type. Instead of typing your question, you say it. Handy when you are cooking, walking, driving, or just thinking out loud. The model treats your spoken words the same as typed ones.
Have a spoken conversation. Many tools now offer a real-time voice mode where you speak, the model answers out loud, and you can interrupt and go back and forth like a phone call. This turns AI into something you can talk with, not just type at.
Turn a recording into text. Feed in a recorded meeting, a lecture, or a voice memo and ask for a transcript, a summary, or the key action items. This is transcription, and it saves hours of note-taking.
Get answers read aloud. The model can speak its reply, which helps when your eyes are busy or you simply prefer listening.
Two directions for audio in a multimodal tool
| Criteria | Audio in | Audio out |
|---|---|---|
| What happens | You speak or upload sound | The model replies with a voice |
| Everyday example | Dictate a message; transcribe a meeting | Hear an answer read aloud |
| Best when | Hands or eyes are busy | You prefer listening to reading |
Audio in
- What happens
- You speak or upload sound
- Everyday example
- Dictate a message; transcribe a meeting
- Best when
- Hands or eyes are busy
Audio out
- What happens
- The model replies with a voice
- Everyday example
- Hear an answer read aloud
- Best when
- You prefer listening to reading
Why a voice conversation feels different
Typing is turn-based and deliberate. A voice conversation is fast and natural. You can trail off, change your mind mid-sentence, or say "wait, go back." Modern voice modes are built to handle that flow, including letting you cut in before the model finishes.
That changes what AI is good for. Talking through a plan on a walk, practicing a language out loud, or rehearsing for an interview all feel natural by voice and awkward by typing. The model is the same underneath. The spoken interface just fits different moments in your day.
Remember the mental model from earlier: whether you type or talk, your input becomes numbers in the same shared space. Voice is not a different brain. It is a different door into the same one.
Where this course stops and a deeper one begins
This lesson is about the concept: sound is a modality, and a multimodal model can take it in and give it out. That is the piece that ties voice into the bigger picture.
It does not teach you to produce studio-quality speech, clone a voice, or build a polished transcription workflow. Those are real skills with their own tools and settings. The free course AI Voice & Audio covers text-to-speech, voice cloning, and transcription hands-on. If audio is your main interest, take that course next. Here, the goal is just to see how voice fits alongside text and images in one model.
Key Takeaways
- Sound is a modality too. The model turns audio into numbers in the same shared space as text and images.
- Audio in means you speak or upload sound; audio out means the model replies with a voice.
- You can dictate instead of type, hold a real spoken conversation, transcribe recordings, and hear answers read aloud.
- Voice suits hands-free and think-out-loud moments; the model underneath is the same as when you type.
- For hands-on text-to-speech, voice cloning, and transcription, take AI Voice & Audio.

