What is text-to-speech?
Updated · By Robert Breen
Text-to-speech (TTS) is software that reads written text out loud in a synthetic voice. Modern AI voices can sound close to a real person, with natural pauses and emphasis. It is the "mouth" of a voice agent and the engine behind AI voice-overs and read-aloud features.
Why it matters for a small business
For a small business, text-to-speech shows up in three places: a voice agent answering the phone, voice-overs for training and marketing videos, and accessibility (reading a page or a message to someone who would rather listen). It lets you produce spoken audio without booking a voice actor or recording every update yourself.
The catch is that text written for the eye often sounds wrong in the ear. Bullet lists, tables, abbreviations, long sentences and symbols like # or ** are fine on a screen and awkward or garbled when spoken. If an AI model writes the words and TTS reads them, the instructions have to ask for text that works out loud.
In a real lesson: Build an AI Follow-Up Agent for HVAC Service Calls
Stepthrough doesn't have a text-to-speech lesson yet. The closest skill you can practice is writing instructions that shape an AI's output for the channel it will be delivered in, which is the same problem TTS adds. In the AI Follow-Up Agent for HVAC Service Calls lesson, the agent for Cedar Ridge Heating & Air, a made-up HVAC company, has to pick a Channel (Text or Email) for each message.
The system message sets channel rules: "Keep text messages under 300 characters. Use email for installs." When you send the three test calls (Linda Morales, furnace tune-up; Tom Becker, replaced AC capacitor; Grace Chen, new heat pump install), the agent writes two short texts and one email, each sized for where it will be read.
For a spoken channel you would add rules the same way: short sentences, no lists or tables, write numbers the way you would say them, spell out abbreviations, and one question at a time. TTS reads exactly what it is given, so the shaping happens in the prompt.

Try this lesson free or read the step-by-step guide.
Common confusions
Text-to-speech vs voice cloning
Text-to-speech is the general technology. Voice cloning is one way to give it a voice, by copying a specific person's voice from recordings. You can use TTS with a stock voice and never clone anyone.
Text-to-speech vs speech-to-text
Text-to-speech speaks. Speech-to-text listens. In a voice agent, one hears the caller and the other answers.
Tips
- Listen to the output before you publish it. Names, prices and addresses are where AI voices most often stumble.
- Ask the model for "plain spoken text with no bullets, headings or symbols" when its answer will be read aloud. Markdown asterisks are a common giveaway.
- Keep a short pronunciation list for your business name, street names and products if your TTS tool supports it.
Related terms
Where to learn more
Prompt templates that use it
Frequently asked questions
- Can customers tell an AI voice from a real person?
- Sometimes not, which is why it is good practice to say up front when a caller or viewer is hearing an AI voice.
- Do I need a separate tool for text-to-speech?
- Usually yes. Some AI assistants can read answers aloud, but voice-overs and phone agents generally use a dedicated TTS service connected to your workflow.