Voice category

Robot Voice Generator

The interesting robot voices are not the buzzy ones. They are the ones that are almost human and wrong in one small way — perfect evenness, no breath, a pause in the wrong place. That is a performance choice, and it is one you can direct.

Four voices to start from

Samples are placeholders while the public demo audio is produced. Every voice is available on every plan — the library is not a paywall.

Three distinct robot registers

The first is the retro computer: heavily processed, syllable-timed, obviously mechanical. It is a nostalgia signal more than a character, and it is built entirely with effects rather than with performance. The second is the contemporary assistant: pleasant, even, mildly warm, indistinguishable from a good commercial text-to-speech voice — which, appropriately, is exactly what it is. The third is the android: fully human in timbre and subtly non-human in behaviour.

Only the third is really a casting decision, and it is the one worth spending time on, because it is the only one that can carry a dramatic scene. The other two are set dressing, and set dressing is cheap.

Writing the uncanny version

Machine speech in fiction is signalled almost entirely by word choice and structure. No contractions. No hedging. No sentence fragments. Complete grammatical sentences even under pressure. Precise numbers where a human would approximate — "in four minutes and ten seconds" rather than "in a few minutes".

The second tool is response mismatch. A human character says something emotional; the machine answers with information. That exchange establishes non-humanity faster than any amount of processing, and it survives translation, compression and bad speakers.

Why you should not process in the render

Every effect you bake in is a decision you cannot revisit. Ring modulation that sounded perfect on your monitors will be unintelligible on a phone. A doubled voice that worked in a quiet scene will fight the music in a loud one. And if you later need to change one word, the re-render will not match the processed original.

Generating clean audio and treating it in your editor costs a few minutes and keeps every one of those decisions open. Stem export exists largely for this reason.

A note on the obvious irony

There is something genuinely funny about using a synthetic voice to play a synthetic voice, and it is worth saying out loud: the reason a modern assistant voice is easy to produce here is that it is the same technology, unmodified. The uncanny android is harder, because it requires the model to be human enough to be believable and controlled enough to be wrong.

In practice this means the retro and assistant registers are close to free, and the android is where your direction time should go. Which is the opposite of where most people spend it.

The machine voice as a comic device

Outside science fiction, the most productive use of an emotionally flat voice is comedy. A relentlessly pleasant automated system explaining something absurd, an in-world announcement delivering catastrophic news in the tone of a departure board, a navigation voice with opinions — all of these work because the register is completely fixed while the content escalates.

The writing rule is that the machine must never acknowledge the absurdity. The moment it comments on what is happening it becomes a character with a sense of humour, and the joke collapses into an ordinary funny voice. Keep the delivery neutral, keep the vocabulary procedural, and let everything else in the scene react instead.

Script to audio in a few steps

  1. Decide which era of robot you want

    A 1960s computer, a modern smart assistant and an uncanny android are three different sounds. The first is processed, the second is flat and pleasant, the third is human with something missing.

  2. Cast flat rather than filtered

    Start from a voice with naturally low emotional variation. Filtering a warm voice to sound robotic usually just sounds like a warm voice through a bad speaker.

  3. Write without contractions

    Machine characters that say "cannot" instead of "can't" read as non-human immediately, with no processing at all. This is the highest-value edit on the page.

  4. Keep every line at neutral

    The unnerving quality of a machine voice comes from emotional flatness in situations that demand emotion. Resist the urge to direct.

  5. Add processing afterwards, if at all

    Export clean and apply any filtering, doubling or ring modulation in your editor. A clean take can always be degraded; a degraded take cannot be recovered.

Frequently asked questions

Can I get a classic buzzy robot voice?

Not from the renderer, and deliberately so. That sound is a processing effect — ring modulation, vocoding, bit reduction — applied after the fact. Generate the performance clean and add the effect in your editor, where you can dial it back when it turns out to be unintelligible.

Why does a flat voice sound creepier than a processed one?

Because a processed voice is obviously a machine and the listener relaxes. A voice that is almost human but emotionally absent sits in the uncanny valley, and the discomfort comes from the listener not being able to categorise it.

How do I make an AI character sound polite but wrong?

Write extremely correct, complete sentences, keep the emotion neutral, and have it respond to emotional moments with procedural information. The mismatch between content and delivery does everything.

Can the voice degrade over a scene?

Not inside a single render. Generate the lines clean and apply increasing degradation to the later ones in your editor. Doing it in post also means you can tune the rate of decay, which is usually wrong the first time.

Is a robot voice useful outside science fiction?

Very. Announcement systems, phone menus, in-world navigation voices and satirical corporate messaging all use the same register, and the comedy value of a relentlessly pleasant automated voice is close to inexhaustible.

Can I use the neutral voices for accessibility narration?

Yes, though for accessibility you generally want the clearest voice rather than the flattest one. Filter for the clear style tag and check it at speed, since many screen-reader users listen well above real time.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.