Synthesis

Expressive Text to Speech for Any Script

Ordinary text to speech converts a block of text with one voice. That is the right tool for narration and the wrong tool for a scene. This page covers the synthesis layer itself — what it can be told to do, and where it stops.

Voices to try it on

Noor

Noor

Newsreader clarity for exposition-heavy scenes

femaleadulten-USnewsclear
Quill

Quill

YouTube explainer voice that never sounds bored

maleyoung-adulten-USyoutubeclear
Dorian

Dorian

Audiobook baritone for literary fiction and slow reveals

maleadulten-GBaudiobookdeep
Solène

Solène

Continental accent, unhurried, very expensive sounding

femaleadulten-GBelegantaccent
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

neutraladulten-UScalmsoothing
Indie

Indie

TikTok voiceover: fast, flat, faintly amused

femaleyoung-adulten-USsocialfast
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

femaleadulten-USnarrationwarm
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

maleadulten-USnarrationdocumentary

Browse the whole library →

What it does

Seven emotional registers

Neutral, happy, sad, angry, fear, whisper and excited, set per line rather than per document. Not every voice supports every register — the library lists which ones each voice carries, because pretending otherwise produces a bad surprise at render time.

Pace control

A speed multiplier per character and per line. Pace does more work than people expect: it carries age, authority and emotional state more reliably than pitch does.

Pronunciation overrides

Correct a word once per project and it stays corrected. Essential for invented names, non-English names and any technical vocabulary.

Punctuation as direction

A full stop is a real pause, an ellipsis a longer one, a question mark lifts the phrase end. Line breaks create gaps. Your script layout is your timing control, which is why this page keeps telling you to split lines.

Provider-independent

The synthesis engine sits behind an interface. Today the default is a mock provider while the production engine is wired up; the same project will render through a different engine without any change to your casting or direction.

Every voice on every plan

The library is not a paywall. Plans differ on cast size, emotion control, clone slots and minutes — not on which voices you may use.

How to use it

  1. Paste the text

    One speaker or many. For a single voice, this behaves like any other text-to-speech tool.

  2. Choose the voice

    Filter by age, gender, language and style rather than scrolling. Audition on your own text, not on sample text.

  3. Set pace and register

    Start with neutral at a slightly slower pace than feels right. Almost every first attempt is too fast.

  4. Fix the pronunciations

    Find every name and invented word and correct it once, before you render anything long.

  5. Render and check on real hardware

    Play it on a phone speaker as well as headphones. Most audiences are on the phone.

Emotion control, and why less is more

The instinct with per-line emotion is to use it everywhere, and the result is uniform mush. Audiences read emotion by contrast, so a scene with two directed lines and twenty neutral ones lands far harder than a scene where everything is pushed.

The workflow that produces good results is: set a resting register per character, render with defaults only, listen straight through, then mark the two or three lines where something actually changes. In a five-minute scene, directing five lines by hand is about right.

Writing that a synthetic voice can perform

Four habits carry most of the difference. One intention per line — a sentence that swings from delight to horror is rendered as an average of both, so split it. Punctuation deliberately, because it is your timing. Interjections as their own lines, because a renderer speaks text and not stage directions. And short sentences for low voices, which lose intelligibility fast over long clauses.

None of these are workarounds for a limitation. They are the same notes a director gives a human performer, for the same reasons.

The limits, stated plainly

Physical vocalisation — screaming, sobbing, laughing while speaking, grunts — is the weakest area and will remain so for a while, because these are breathing events rather than speech events. Generate them for placeholder if you need to, and plan to replace them.

The second limit is interpretation. A synthetic read is a clean, consistent, well-paced performance of what you wrote. It is not a set of decisions about subtext. For most work that trade is worth making; it is worth making with clear eyes about which half of it is which.

Frequently asked questions

How is this different from other text-to-speech tools?

For single-voice work, not very. The difference appears the moment a script has more than one speaker: casting, per-character persistence and per-line direction are the product, and the synthesis is a component underneath them.

Which emotions are supported?

Neutral, happy, sad, angry, fear, whisper and excited. Each voice lists the registers it actually supports rather than claiming all seven.

Can I control pitch?

Pace, yes; pitch shifting, deliberately not beyond small amounts. Pitch-shifting a voice far from its natural range produces artefacts that are obvious on good headphones. Choose a voice in the right range instead.

Does it read numbers and dates correctly?

It applies the conventions of the voice's locale, so an American voice reads dates the American way. Mixed-convention text is where mistakes appear; check numbers first if your text mixes them.

What is the mock provider?

In this release the default engine returns a placeholder audio file so the whole pipeline — casting, direction, export — can be built and tested end to end. The provider interface, and the adapters for real engines, are in the codebase.

Can I use it for accessibility narration?

Yes, though for accessibility you want the clearest voice rather than the most characterful, and you should test at 1.5× or faster, since many screen-reader users listen well above real time.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.