
Grit
A chest cavity roughly the size of a bus
Voice category
Monster voices work best when something human survives in them. Fully inhuman noise is a sound-design job; a monster that can hold a conversation is a character, and a character is what your scene actually needs.

Grit
A chest cavity roughly the size of a bus

Morrow
Under the floorboards, patient, knows your name

Pip
Small, chaotic and oddly friendly about it

Vex
Reasonable tone, unreasonable intentions
Every voice is available on every plan — the library is not a paywall.
The instinct with monsters is to reach for the lowest voice available and then push it lower. It rarely works, because the ear does not read size from frequency alone — it reads size from how long things take. Large bodies move slowly, breathe slowly and take their time between sentences, and a listener infers dimensions from that rhythm without ever thinking about it.
So the controls that matter are pace and line length. A creature that delivers four words, pauses, then delivers three more feels enormous even in a mid-range voice. The same voice running a full paragraph at conversational speed feels like a person in a costume. If you have line-level speed on Plus or Pro, take the creature one or two steps below the rest of the cast and leave it there for the whole scene.
The monsters that stay with an audience are the ones that can hold a conversation. Fully inhuman noise is impressive for about four seconds and then stops carrying information, which is why even the most alien creatures in published audio fiction are given a voice that can form words. A thing that can talk to you can also lie to you, bargain with you and be disappointed in you — none of which a roar can do.
This has a practical consequence for casting. Choose the voice that is closest to your creature while still being clearly articulate, and treat any further degradation as a post-production decision. A clean, growling read is raw material; a read that already sits at the edge of comprehension is a dead end, because every effect you add from there makes it worse.
It also means the writing has to be good. A monster with three lines needs those lines to do something other than threaten. Give it a preference, a complaint or a piece of information, and the scene will survive a voice that is only approximately what you imagined.
Almost every creature voice you admire is two or more takes stacked: the original, a copy pitched down, sometimes a third copy delayed by a few milliseconds so the consonants smear. None of that happens inside the render, and that is a deliberate boundary. Effects baked into generated audio cannot be undone, and the version you want in the final mix is almost never the version you wanted while writing.
The workflow is straightforward. Generate the creature's lines clean, export per-character stems on Pro so the creature arrives on its own full-length track, then do the stacking in your audio editor against the dialogue it is answering. Because the stems are the same length as the mix, everything lines up without any manual nudging.
One warning from experience: do not pitch the whole track uniformly. Pitch the moments — the entrance, the threat, the last line — and leave the conversational passages closer to the source. Uniform processing is what makes home-made creature audio sound like a filter rather than like a thing that is present in the room.
It will not scream, sob, gurgle, breathe wetly on command or produce any of the non-verbal noises that creature work loves. Those are body sounds rather than speech, and writing them as stage directions produces nothing: a line that is entirely in brackets is skipped, and anything else is read aloud as text. Source them separately and treat the generated dialogue as one layer among several.
It will not give you a truly non-human voice either. Everything starts from a human performance, which is a limitation if you wanted something genuinely alien and an advantage if you wanted something that could plausibly hold a conversation. The most economical way to signal inhumanity is rhythm — stress the wrong word, pause in the wrong place, answer a question slightly too fast.
And the same line generated twice will not be identical. Each take is a fresh performance with the same voice and slightly different delivery, which is genuinely useful for creatures: generate the roar-adjacent lines three or four times and keep the take where something in the timing is off in the right way.
Enormous means slow and low; small means fast and bright. Scale drives every other decision, so settle it before you audition anything — a creature that is both huge and quick reads as neither, and the listener spends the scene trying to picture it.
Cast a voice that can still articulate, then degrade it in post if you need to. A read that is already at the edge of intelligibility leaves you nothing to work with, whereas a clear growl can be pitched, filtered and doubled into almost anything.
Low, rough voices lose intelligibility quickly across long sentences, and listeners forgive a monster saying six words far more readily than one delivering a paragraph. Break the speech into separate lines; each line becomes its own performance with a real pause between them.
Two copies of the same line, one pitched down and offset by a few milliseconds, is the classic creature effect and it has to happen outside the render. Generate once, duplicate in your audio editor, and keep the original clean as your reference.
Layering needs the creature on its own track, away from the dialogue it is answering. Per-character stems come as a ZIP on Pro — the full mix, one full-length track per character, plus an SRT — which is exactly the shape creature work wants.
No. Roars, shrieks and snarls are sound design rather than speech, and the renderer performs written lines. Take the roar from a library or record it yourself, then generate the dialogue around it. Practically every published creature you can think of is that same combination.
Slow the pace and shorten the lines. Size reads as tempo far more than as pitch — a big animal moves slowly, breathes slowly and is never in a rush, and the ear picks that up long before it registers frequency. Line-level speed on Plus and Pro is the direct control; short sentences do most of it for free.
There is not. You choose a library character and direct the performance; you do not shift the pitch of the result. That is deliberate — a neutral adult voice dropped an octave sounds processed rather than large. Start from a character that is already close, and do any further shifting in your editor where you can hear it against the mix.
Mismatch one thing. A small voice saying something enormous, a friendly tone describing something appalling, or a creature that pauses in the wrong places. Wrongness is a pattern violation, and the listener only notices a violation if everything else is behaving normally.
Yes, and they should. Characters in one project are kept on different base voices until the pool for that gender is exhausted, so a pack of creatures will sound distinct up to the point where the project's ten distinct voices run out. Beyond that, separate them by pace and line length rather than hoping the voices will do it.
They do, and the small chaotic register is built for exactly that. Keep the pace slower than you think, avoid stacking two rough voices back to back, and remember that on small speakers a low growl mostly disappears — a friendly monster in children's audio is usually a bright voice with odd rhythm, not a deep one.
Paste a script, let CastDub assign a voice to every character, and export a finished drama.
Create for freeNo credit card. Free gives you a one-time pot of credits.