
Wren
Bedtime-story narrator, slow and warm, never wakes anyone up
Workflow
Children's audio has a specific grammar: a warm adult narrator carrying the prose, distinct character voices for the dialogue, and a pace considerably slower than adult narration. All three are decisions, and all three are easy to get wrong.
Child characters are synthetic models, not recordings of children. For anyone producing children's media this removes a genuine set of consent, scheduling and long-term-risk problems, which is why several publishers now prefer it for incidental characters.
Children's audio depends on a steady, warm adult voice that a young listener can follow across a whole story. Audition several on your actual text and pick the one that stays comfortable at length, not the one that sounds nicest for ten seconds.
The most common fault in amateur children's audio is speed. Set the pace slower than feels natural to you — young listeners need the extra space, and it is a setting rather than a performance skill.
Subtitles generated from your script are correct by construction, which matters for classroom use, for read-along editions, and for any platform that requires captions.

Wren
Bedtime-story narrator, slow and warm, never wakes anyone up

Milo
Nine-year-old with a loose tooth and a very large opinion

Hazel
Mum voice: patient, then suddenly not

Gran Thea
Grandmother telling a folk tale she half believes

Pip
Small chaotic gremlin voice for monsters that are friendly

Brix
Cartoon sidekick made of enthusiasm and bad ideas

Bram
Tavern keeper and accidental quest giver

Halcyon
Meditation-adjacent narrator for slow, quiet scenes
One idea per sentence. Children's prose is short not because children are simple but because listening is harder than reading.
The narrator carries most of the runtime and sets the tone. Everything else is chosen to fit around them.
Two or three characters, clearly separated. More than that and young listeners lose track, which is why classic children's stories have small casts.
Reduce it below your instinct and listen again. Almost every first attempt at children's audio is too fast.
Some excitement, some gentleness, nothing frightening at high intensity. A story at constant maximum energy is exhausting rather than exciting.
An SRT alongside the audio gives you a read-along edition and covers accessibility requirements for schools and platforms.
Published children's audio is remarkably consistent in form, and the consistency is not laziness. A warm adult narrator holds the frame. Character dialogue is voiced distinctly but within a narrow emotional range. Pace is slow, pauses are generous, and the whole thing is mixed with a narrow dynamic range so nothing jumps.
Each of those conventions solves a real problem. The steady narrator gives a young listener something to hold on to when they lose the thread. Distinct character voices remove the work of tracking who is speaking. Slow pace matches the speed at which children process spoken language, which is substantially slower than adults. And narrow dynamics mean a story can be listened to quietly at bedtime, which is when most of it is consumed.
If you change one thing after reading this page, slow down. Adults producing children's audio almost universally run too fast, because their own comprehension is instant and the silence feels like dead air. To a five-year-old it is not dead air; it is the time in which the sentence is being understood.
The practical method is to render a page, listen once at your instinctive pace, then reduce the pace and listen again. The second version will feel slow to you and will be approximately right. If you have access to an actual child, the test takes ninety seconds and settles the question permanently.
Children's stories have small casts for the same reason picture books have few characters per spread: attention is limited. Two or three speaking parts plus a narrator is the practical ceiling, and it is what most classic children's fiction uses.
Separate those parts as widely as you can. A high bright child, a low warm adult and a middle-register second character will remain distinguishable to a young listener even when they are half asleep, which is precisely the condition this material is consumed in.
Resist the urge to give every animal a comic voice. Two exaggerated voices in a story are delightful; six are noise, and they make the narration harder to follow rather than more fun.
A recurring request is to clone a parent's or grandparent's voice so a child can hear a familiar person read to them. The living-relative version of this is possible under our cloning policy with their recorded consent, and it is a genuinely lovely use. The deceased-relative version is not, because consent cannot be obtained, and no amount of family agreement substitutes for it.
Cloning a child's voice is never available, under any plan or agreement. A model of an identifiable child's voice is a tool for impersonating that child, and there is no creative requirement that justifies creating one.
Children's audio has a hard external constraint that no other format has: a substantial share of it is played at bedtime, by a tired adult, to a child who is supposed to fall asleep. That shapes everything. Episodes of eight to twelve minutes fit the ritual; twenty-five minute episodes do not, and parents quietly stop using them.
It also shapes the ending. A story that finishes on an exciting cliffhanger is a story that wakes a child up. The convention in bedtime audio is a deliberate wind-down in the last ninety seconds — slower pace, quieter delivery, a resolution rather than a hook — and it is the single most requested feature by the adults who actually press play.
Adults are unreliable judges of children's audio, consistently choosing material that is too fast, too dense and too clever. The correction is trivially available: play it to a child and watch.
You are looking for two things. Whether they can follow who is speaking without asking, which tells you the casting is separated enough. And where their attention goes, which is almost never where you expected. Ninety seconds of observation will teach you more than any amount of theory about pacing, and it will usually tell you to slow down again.
Free
$0 / month
Enough to finish a short scene and hear what your script sounds like cast.
Creator
$12 / month
For one person turning their own stories, fanfic or scripts into finished audio.
Studio
$29 / month
For serialised work: long-running dramas, audiobooks, dubbing pipelines.
Team
$69 / month
For a studio where a writer, a director and an editor touch the same project.
No. They are synthetic models, and no recording of a real child is involved in generating them. We also do not build voice clones of minors under any circumstances, including with parental consent — it is the one place our cloning policy has no exception.
Noticeably slower than adult audiobook narration, which itself runs around 150 words per minute. Start below that and check with an actual child if you can; adults consistently overestimate the right speed.
Dialogue yes, prose no. The convention of a warm adult narrator carrying the story with child voices only for speech is near-universal in published children's audio because young listeners find a steady adult voice easier to follow over time.
Yes, on any paid plan, and export the captions — many schools require them on audio-visual material. The slower pacing advice applies even more strongly to educational content.
It is one of the most common uses. Creator suits a weekly ten-minute episode comfortably. Keep the dynamic range narrow so a loud moment does not wake a child who has fallen asleep.
Children's stories can and should have tension, but handle it through pacing and pauses rather than through volume or harsh voices. A quiet, slow delivery of something ominous is both more effective and much less likely to genuinely frighten a small listener.
Paste a script, let CastDub assign a voice to every character, and export a finished drama.
Create for freeNo credit card. Free plan renews every month.