Text to Speech with a Different Voice for Each Character

How multi-voice text to speech actually works — speaker detection, casting for separation, per-line direction, and the three mistakes that make a cast scene unlistenable.

Almost every text-to-speech tool is built around one voice and one text box. That is the right design for narration, announcements and accessibility, and it is the wrong design for anything with more than one person in it.

If you have tried to make a scene with several characters using a single-voice tool, you know the workaround: run it once per character, download a pile of clips, and assemble them in an audio editor. It works. It also takes an afternoon per scene, and the moment you change a line you do it again.

This is how the multi-voice version works, and — more usefully — what makes the difference between a cast scene that is pleasant to listen to and one that is not.

What "multi-voice" actually means

Three distinct things, which are often conflated:

  1. Speaker detection. Working out from the text who is speaking each line.
  2. Casting. Assigning a specific voice to each speaker and keeping that assignment stable.
  3. Direction. Controlling emotion, pace and emphasis per line rather than per document.

A tool that only does the first is a parser. A tool that does the first two is a convenience. All three together is what changes how long a project takes.

Speaker detection: the formats that work

You do not have to convert your writing into a template. Three shapes are common and all of them are parseable.

Name-prefixed lines. MARIA: I told you not to come here. This is the most explicit and the least ambiguous. If you are writing specifically for audio, write like this.

Screenplay format. Character name above the line, parentheticals for direction, scene headings for structure. Exported text from most screenwriting software is already in this shape.

Prose with dialogue tags. "I told you not to come here," she said. This is how fiction is actually written, and it is the hardest of the three because attribution is dropped as soon as a rhythm is established.

For prose, expect to check the assignments. Long unattributed exchanges, three-way conversations and characters referred to by description rather than name are where detection fails. Scanning takes a minute; fixing it afterwards takes much longer.

Casting: separation beats accuracy

Here is the counterintuitive part, and it is the one thing most people get wrong.

Listeners identify speakers primarily by pitch band and speaking rate. Not by timbre, not by accent, and definitely not by personality. In audio there is no face to look at, so those two coarse features carry almost all the identification work.

This means your cast needs to be spread, and spread matters more than each individual voice being perfect for the role.

If two characters occupy the same pitch band and speak at similar speeds, they will blur into one person during any fast exchange — precisely the moments where following the scene matters most. It does not help that they have different personalities on the page. The listener cannot see the page.

The practical method:

  • Cast your two leads as far apart as the characterisation allows.
  • Place secondary characters in the gaps between them.
  • Differentiate by pace as well as pitch. A fast, eager voice and a slow, guarded one stay separate even when they are close in tone.
  • Put the narrator in a band no character occupies.

Then run the eyes-closed test: render a page of dialogue and listen without the script. If you cannot tell who is speaking, and you wrote it, nobody can.

Direction: mark less than you want to

The instinct with per-line emotion control is to use it on every line. This is the second big mistake, and the result is a scene where nothing stands out because everything is turned up.

Emotion in audio is read by contrast. A character who is level for eight lines and breaks on the ninth is devastating. A character who is emotional for all nine is exhausting after about forty seconds.

The workflow that produces good scenes:

  1. Set a default register per character — guarded, eager, flat, warm. Most of a scene runs on the default.
  2. Render the whole scene with defaults only and listen to it straight through.
  3. Mark the two or three lines where something actually changes.
  4. Render again.

In a five-minute scene you should be directing perhaps five lines by hand. If you are directing thirty, the scene is probably over-written rather than under-directed.

The third mistake: forgetting the gaps

Silence is the most underused tool in audio, and it is free.

The gap between lines is where a listener processes what was said and anticipates what comes next. Directors of radio drama spend a disproportionate amount of their time on gaps, because that is where the audience does their work.

In a line-by-line rendering model, your script layout is your timing. A line break is a pause. Splitting a line in two creates a beat in the middle of a thought. Writing a reaction as its own two-word line gives it a real moment rather than burying it in a parenthetical.

Three specific habits:

  • Put the beat before a punchline or a revelation on its own line.
  • Write interjections as lines, not as stage directions. Ugh. on its own line becomes a performance; (groans) becomes nothing, because renderers speak text, not directions.
  • Resist trimming silence in the edit. A tight cut destroys the thing you were building.

Keeping characters consistent

For anything longer than a single scene, consistency becomes the dominant problem, and it is not a technical one.

A listener will forgive a slightly flat read. They will notice instantly when a character sounds like a different person in episode four, because human hearing is tuned for exactly that.

Drift almost never comes from the model. It comes from a person re-choosing a similar-but-different voice weeks later. The fix is to store the decision: a character card holding the voice, the pace, the default emotion and the pronunciation fixes, under the character's name. Load the card, and there is nothing to re-decide.

While you are there, do a pronunciation pass at the start of a project rather than after nine chapters. Names, places and invented words are exactly what synthesis guesses wrong, and a correction that applies project-wide is worth twenty minutes on day one.

What to do about overlaps

Real conversations have people talking over each other, and it is one of the strongest signals that a scene is between humans rather than between settings.

No line-by-line renderer produces overlap, because each line is a separate performance. The answer is stems: export one audio file per character and slide them over each other in any editor. Overlapping the last two lines of an argument by half a second is a single drag and it is the most realistic thing you can do to a confrontation scene.

This is also why a mixed-down dialogue track is a limitation rather than a convenience. If you intend to add music, effects or spatial placement, you want the stems.

A short checklist

Before you render a scene:

  • Are the speaker assignments right, particularly in long exchanges?
  • Are any two characters in the same pitch band and speaking at the same speed?
  • Is the narrator clearly separated from every character?
  • Have you set defaults per character rather than directing every line?
  • Have the invented words been checked?

After you render:

  • Listen end to end at normal speed, without the script.
  • Can you follow who is speaking with your eyes closed?
  • Does anything feel long? It is longer than you think.
  • Have you left the silences alone?

None of this is about voice quality, which is the thing tools compete on and the thing that matters least once it is good enough. Casting, restraint and timing are what separate a scene people listen to from a scene people switch off.

Hear a sample scene

A two-character exchange with a narrator, rendered with three different voices. Placeholder audio while the public demo is produced.

More

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.