
Nova
Ship AI that has read your file
Voice category
The interesting robot voices are not the buzzy ones. They are the ones that are almost human and wrong in one small way — perfect evenness, no breath, a pause in the wrong place. That is a performance choice, and it is one you can direct.

Nova
Ship AI that has read your file

Halcyon
Calm to the point of being unsettling

Noor
Announcement system clarity

Vex
Something wearing a reasonable tone
Samples are placeholders while the public demo audio is produced. Every voice is available on every plan — the library is not a paywall.
The first is the retro computer: heavily processed, syllable-timed, obviously mechanical. It is a nostalgia signal more than a character, and it is built entirely with effects rather than with performance. The second is the contemporary assistant: pleasant, even, mildly warm, indistinguishable from a good commercial text-to-speech voice — which, appropriately, is exactly what it is. The third is the android: fully human in timbre and subtly non-human in behaviour.
Only the third is really a casting decision, and it is the one worth spending time on, because it is the only one that can carry a dramatic scene. The other two are set dressing, and set dressing is cheap.
Machine speech in fiction is signalled almost entirely by word choice and structure. No contractions. No hedging. No sentence fragments. Complete grammatical sentences even under pressure. Precise numbers where a human would approximate — "in four minutes and ten seconds" rather than "in a few minutes".
The second tool is response mismatch. A human character says something emotional; the machine answers with information. That exchange establishes non-humanity faster than any amount of processing, and it survives translation, compression and bad speakers.
Every effect you bake in is a decision you cannot revisit. Ring modulation that sounded perfect on your monitors will be unintelligible on a phone. A doubled voice that worked in a quiet scene will fight the music in a loud one. And if you later need to change one word, the re-render will not match the processed original.
Generating clean audio and treating it in your editor costs a few minutes and keeps every one of those decisions open. Stem export exists largely for this reason.
There is something genuinely funny about using a synthetic voice to play a synthetic voice, and it is worth saying out loud: the reason a modern assistant voice is easy to produce here is that it is the same technology, unmodified. The uncanny android is harder, because it requires the model to be human enough to be believable and controlled enough to be wrong.
In practice this means the retro and assistant registers are close to free, and the android is where your direction time should go. Which is the opposite of where most people spend it.
Outside science fiction, the most productive use of an emotionally flat voice is comedy. A relentlessly pleasant automated system explaining something absurd, an in-world announcement delivering catastrophic news in the tone of a departure board, a navigation voice with opinions — all of these work because the register is completely fixed while the content escalates.
The writing rule is that the machine must never acknowledge the absurdity. The moment it comments on what is happening it becomes a character with a sense of humour, and the joke collapses into an ordinary funny voice. Keep the delivery neutral, keep the vocabulary procedural, and let everything else in the scene react instead.
A 1960s computer, a modern smart assistant and an uncanny android are three different sounds. The first is processed, the second is flat and pleasant, the third is human with something missing.
Start from a voice with naturally low emotional variation. Filtering a warm voice to sound robotic usually just sounds like a warm voice through a bad speaker.
Machine characters that say "cannot" instead of "can't" read as non-human immediately, with no processing at all. This is the highest-value edit on the page.
The unnerving quality of a machine voice comes from emotional flatness in situations that demand emotion. Resist the urge to direct.
Export clean and apply any filtering, doubling or ring modulation in your editor. A clean take can always be degraded; a degraded take cannot be recovered.
Not from the renderer, and deliberately so. That sound is a processing effect — ring modulation, vocoding, bit reduction — applied after the fact. Generate the performance clean and add the effect in your editor, where you can dial it back when it turns out to be unintelligible.
Because a processed voice is obviously a machine and the listener relaxes. A voice that is almost human but emotionally absent sits in the uncanny valley, and the discomfort comes from the listener not being able to categorise it.
Write extremely correct, complete sentences, keep the emotion neutral, and have it respond to emotional moments with procedural information. The mismatch between content and delivery does everything.
Not inside a single render. Generate the lines clean and apply increasing degradation to the later ones in your editor. Doing it in post also means you can tune the rate of decay, which is usually wrong the first time.
Very. Announcement systems, phone menus, in-world navigation voices and satirical corporate messaging all use the same register, and the comedy value of a relentlessly pleasant automated voice is close to inexhaustible.
Yes, though for accessibility you generally want the clearest voice rather than the flattest one. Filter for the clear style tag and check it at speed, since many screen-reader users listen well above real time.
Paste a script, let CastDub assign a voice to every character, and export a finished drama.
Create for freeNo credit card. Free plan renews every month.