
Noor
Newsreader clarity for exposition-heavy scenes
Synthesis
Ordinary text to speech converts a block of text with one voice. That is the right tool for narration and the wrong tool for a scene. This page covers the synthesis layer itself — what it can be told to do, and where it stops.

Noor
Newsreader clarity for exposition-heavy scenes

Quill
YouTube explainer voice that never sounds bored

Dorian
Audiobook baritone for literary fiction and slow reveals

Solène
Continental accent, unhurried, very expensive sounding

Halcyon
Meditation-adjacent narrator for slow, quiet scenes

Indie
TikTok voiceover: fast, flat, faintly amused

Wren
Bedtime-story narrator, slow and warm, never wakes anyone up

Hale
Documentary narrator with a hand on the listener's shoulder
Neutral, happy, sad, angry, fear, whisper and excited, set per line rather than per document. Not every voice supports every register — the library lists which ones each voice carries, because pretending otherwise produces a bad surprise at render time.
A speed multiplier per character and per line. Pace does more work than people expect: it carries age, authority and emotional state more reliably than pitch does.
Correct a word once per project and it stays corrected. Essential for invented names, non-English names and any technical vocabulary.
A full stop is a real pause, an ellipsis a longer one, a question mark lifts the phrase end. Line breaks create gaps. Your script layout is your timing control, which is why this page keeps telling you to split lines.
The synthesis engine sits behind an interface. Today the default is a mock provider while the production engine is wired up; the same project will render through a different engine without any change to your casting or direction.
The library is not a paywall. Plans differ on cast size, emotion control, clone slots and minutes — not on which voices you may use.
One speaker or many. For a single voice, this behaves like any other text-to-speech tool.
Filter by age, gender, language and style rather than scrolling. Audition on your own text, not on sample text.
Start with neutral at a slightly slower pace than feels right. Almost every first attempt is too fast.
Find every name and invented word and correct it once, before you render anything long.
Play it on a phone speaker as well as headphones. Most audiences are on the phone.
The instinct with per-line emotion is to use it everywhere, and the result is uniform mush. Audiences read emotion by contrast, so a scene with two directed lines and twenty neutral ones lands far harder than a scene where everything is pushed.
The workflow that produces good results is: set a resting register per character, render with defaults only, listen straight through, then mark the two or three lines where something actually changes. In a five-minute scene, directing five lines by hand is about right.
Four habits carry most of the difference. One intention per line — a sentence that swings from delight to horror is rendered as an average of both, so split it. Punctuation deliberately, because it is your timing. Interjections as their own lines, because a renderer speaks text and not stage directions. And short sentences for low voices, which lose intelligibility fast over long clauses.
None of these are workarounds for a limitation. They are the same notes a director gives a human performer, for the same reasons.
Physical vocalisation — screaming, sobbing, laughing while speaking, grunts — is the weakest area and will remain so for a while, because these are breathing events rather than speech events. Generate them for placeholder if you need to, and plan to replace them.
The second limit is interpretation. A synthetic read is a clean, consistent, well-paced performance of what you wrote. It is not a set of decisions about subtext. For most work that trade is worth making; it is worth making with clear eyes about which half of it is which.
For single-voice work, not very. The difference appears the moment a script has more than one speaker: casting, per-character persistence and per-line direction are the product, and the synthesis is a component underneath them.
Neutral, happy, sad, angry, fear, whisper and excited. Each voice lists the registers it actually supports rather than claiming all seven.
Pace, yes; pitch shifting, deliberately not beyond small amounts. Pitch-shifting a voice far from its natural range produces artefacts that are obvious on good headphones. Choose a voice in the right range instead.
It applies the conventions of the voice's locale, so an American voice reads dates the American way. Mixed-convention text is where mistakes appear; check numbers first if your text mixes them.
In this release the default engine returns a placeholder audio file so the whole pipeline — casting, direction, export — can be built and tested end to end. The provider interface, and the adapters for real engines, are in the codebase.
Yes, though for accessibility you want the clearest voice rather than the most characterful, and you should test at 1.5× or faster, since many screen-reader users listen well above real time.
Paste a script, let CastDub assign a voice to every character, and export a finished drama.
Create for freeNo credit card. Free plan renews every month.