How to · Audio

AI voiceover that doesn't sound AI

The difference between a believable AI voice and an obvious one is almost always in the script, not the model.

10 min read · Voiceover

AI voiceover gets blamed for the wrong sin. The complaint is usually 'it sounds robotic,' but the actual problem is that the script was written for the page, not for the breath. A great human VO artist would have rewritten the script in their head before the take - the AI doesn't get to do that.

If you do the rewrite up front, modern voice models - ElevenLabs, OpenAI, the newer Claude voices - produce takes that pass a casual listening test. Here's the working method.

Why most AI VO sounds AI

Three things give it away, in order of severity:

  1. No breath. Real speech has micro-breaths between clauses. The AI doesn't take any, so eight seconds in, you feel a robotic rhythm even if you can't name it.
  2. Flat Prosody. The model lands every clause at the same pitch contour. Real speakers vary the Pacing - they speed up on familiar phrases and slow down on emphasis.
  3. Punctuation as a stop sign. The model treats commas like full stops. If your script has nine commas in a sentence, the model is going to hammer all nine.
Each one is fixable in the script before you ever press generate.

Write for the breath

Read your script out loud, once. Mark every place you naturally take a breath with a period. Now rewrite the script using those periods as your sentence breaks.

It will feel like a regression - your beautiful semicolons and em-dashes are now full stops. That's the point. AI voice models honor full stops with a real pause; they ignore most other punctuation as a polite suggestion.

Before:

Most creators, when they start out, tend to overlook the importance of pacing - which, ironically, is the one variable they have direct control over.
After:
Most creators overlook pacing when they start out. That's ironic. It's the one variable they actually control.
Same idea, three sentences instead of one. The AI gives you three landings instead of one breathless run-on. Listen back and you'll hear it land like a human.

The one-thought-per-sentence rule

In writing, a sentence can carry two ideas - the second one in a subordinate clause. In speech, two ideas in one sentence almost always sound like the speaker is reading.

The fix is mechanical: every time you find an "and" or "but" or "which" mid-sentence, ask whether it should be a period. About 70% of the time the answer is yes. Your script gets longer on the page and shorter on the ear.

Punctuation cheat sheet for AI VO

  • Period - full pause. Use generously.
  • Comma - micro-pause. Use sparingly; the model ignores most of them.
  • Em-dash - works as a Pacing hold in ElevenLabs and OpenAI models, less reliable in others.
  • Ellipsis - produces a slow trailing inflection. Useful at the end of a sentence; avoid mid-sentence.
  • Question mark - lifts the final word. Use them; the models do this well.
  • Exclamation mark - usually overshoots. Replace with a period and rely on the line itself to read excited.

Match the voice to the script, not the script to the voice

Most creators pick a voice they like, then write the script. It's backwards. Pick the script's character first - sardonic narrator? warm explainer? gen-z confessional? - and audition voices against a sample paragraph that's already in that register.

ElevenLabs has roughly four voice archetypes that actually carry a full one-minute take: the Warm Authority (deep, slow), the Conversational Peer (mid-pitch, casual), the Bright Hostess (higher, faster, more upspeak), and the Documentary Narrator (slow, low-saturation). Mismatching script and archetype costs you more than any prompt-engineering trick.

Read direction beats voice direction

Here's the part nobody writes down: the AI doesn't read direction tags like (excited) or (whispered). They might work occasionally, but you can't ship a script that depends on them.

What works instead is read direction by word choice. If you want the line to read excited, write the line in a way that a human would read excited:

  • Replace abstract nouns with concrete verbs.
  • Shorten sentences.
  • Use sentence fragments. On purpose.
  • Repeat for emphasis. Repeat.
Read this line two ways:
"This is, in my opinion, a very effective technique for engagement retention."
"This works. Most things in this space don't."
Same general idea. The first one will sound like a corporate explainer no matter which voice you pick. The second one will read confident in any voice. The script did the work the direction tag couldn't.

SSML when you need surgical control

If you're doing serious VO work, the markup language SSML buys you another layer. You can specify pause length, emphasis, speech rate, even pitch. The major providers honor a subset of it.

<speak>
  Sixty seconds is a contract.
  <break time="450ms"/>
  Deliver the promise in the first three.
  <break time="200ms"/>
  Or the viewer leaves.
</speak>

A 450-millisecond break is the rhythm of a deliberate human pause. A 200-millisecond break is the rhythm of a confident continuation. Together, they're the punctuation the comma wishes it could be.

Budget reality

A 60-second script at 145 words is roughly 870 characters. At current ElevenLabs pricing, that's around \$0.25 per take. You'll do two or three takes on average to land the read, so plan \$0.50–\$0.75 per minute of finished video. The Voiceover Cost Estimator will compute the actual character count for your script and the corresponding cost across providers - useful when you're producing at scale and need to know whether the budget survives ten cuts a week.

When to record human VO instead

Three cases where I still go human:

  • Strong regional accent or dialect. The current AI voices flatten regional speech into a generic American or generic British. If the character matters, hire the character.
  • Performed emotion. Genuine sob, genuine laugh, the catch in the throat. Models simulate; they don't perform.
  • Brand-stakes narration. A company's signature voice, used across every video, should be a real person under contract. AI is great for ten cuts a week; it's risky for the one cut that defines the brand.
For everything else - the daily explainer, the news-briefing, the product walkthrough - modern AI Voiceover clears the bar, provided the script does its half of the work.

Bring it to VideoCue

Run the final script through the Voiceover Cost Estimator before you commit - it'll tell you the per-take cost and let you compare ElevenLabs, OpenAI, and the other supported providers side-by-side. The Filmwiki entry on Prosody is where to go next if you want to engineer the read more deeply, and Pacing covers the rhythm-level adjustments that separate a take that reads as written from a take that reads as performed.

Related Filmwiki terms

Your first cut
is on us.

+ New videoFree plan · no card · 1080p export