Workflow

Children's Story Audio

Children's audio has a specific grammar: a warm adult narrator carrying the prose, distinct character voices for the dialogue, and a pace considerably slower than adult narration. All three are decisions, and all three are easy to get wrong.

Why this works here

No child has to be recorded

Child characters are synthetic models, not recordings of children. For anyone producing children's media this removes a genuine set of consent, scheduling and long-term-risk problems, which is why several publishers now prefer it for incidental characters.

The narrator can be properly warm

Children's audio depends on a steady, warm adult voice that a young listener can follow across a whole story. Audition several on your actual text and pick the one that stays comfortable at length, not the one that sounds nicest for ten seconds.

Pace is under your control

The most common fault in amateur children's audio is speed. Set the pace slower than feels natural to you — young listeners need the extra space, and it is a setting rather than a performance skill.

Captions come out of the same project

Subtitles generated from your script are correct by construction, which matters for classroom use, for read-along editions, and for any platform that requires captions.

Voices that suit this kind of work

Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

femaleadulten-USnarrationwarm
Milo

Milo

Nine-year-old with a loose tooth and a very large opinion

malechilden-USkidscartoon
Hazel

Hazel

Mum voice: patient, then suddenly not

femaleadulten-USwarmdomestic
Gran Thea

Gran Thea

Grandmother telling a folk tale she half believes

femalesenioren-GBnarrationfolk
Pip

Pip

Small chaotic gremlin voice for monsters that are friendly

neutralchilden-USmonstercartoon
Brix

Brix

Cartoon sidekick made of enthusiasm and bad ideas

maleyoung-adulten-UScartooncomedy
Bram

Bram

Tavern keeper and accidental quest giver

malesenioren-GBdndwarm
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

neutraladulten-UScalmsoothing

Browse the whole library →

How to do it

  1. Write short sentences

    One idea per sentence. Children's prose is short not because children are simple but because listening is harder than reading.

  2. Cast the narrator first

    The narrator carries most of the runtime and sets the tone. Everything else is chosen to fit around them.

  3. Give each character a clear voice

    Two or three characters, clearly separated. More than that and young listeners lose track, which is why classic children's stories have small casts.

  4. Slow the pace deliberately

    Reduce it below your instinct and listen again. Almost every first attempt at children's audio is too fast.

  5. Keep the emotional range narrow

    Some excitement, some gentleness, nothing frightening at high intensity. A story at constant maximum energy is exhausting rather than exciting.

  6. Export with captions

    An SRT alongside the audio gives you a read-along edition and covers accessibility requirements for schools and platforms.

The grammar of children's audio

Published children's audio is remarkably consistent in form, and the consistency is not laziness. A warm adult narrator holds the frame. Character dialogue is voiced distinctly but within a narrow emotional range. Pace is slow, pauses are generous, and the whole thing is mixed with a narrow dynamic range so nothing jumps.

Each of those conventions solves a real problem. The steady narrator gives a young listener something to hold on to when they lose the thread. Distinct character voices remove the work of tracking who is speaking. Slow pace matches the speed at which children process spoken language, which is substantially slower than adults. And narrow dynamics mean a story can be listened to quietly at bedtime, which is when most of it is consumed.

Why the pace question is the whole game

If you change one thing after reading this page, slow down. Adults producing children's audio almost universally run too fast, because their own comprehension is instant and the silence feels like dead air. To a five-year-old it is not dead air; it is the time in which the sentence is being understood.

The practical method is to render a page, listen once at your instinctive pace, then reduce the pace and listen again. The second version will feel slow to you and will be approximately right. If you have access to an actual child, the test takes ninety seconds and settles the question permanently.

Casting a small, clear ensemble

Children's stories have small casts for the same reason picture books have few characters per spread: attention is limited. Two or three speaking parts plus a narrator is the practical ceiling, and it is what most classic children's fiction uses.

Separate those parts as widely as you can. A high bright child, a low warm adult and a middle-register second character will remain distinguishable to a young listener even when they are half asleep, which is precisely the condition this material is consumed in.

Resist the urge to give every animal a comic voice. Two exaggerated voices in a story are delightful; six are noise, and they make the narration harder to follow rather than more fun.

The consent question, stated plainly

A recurring request is to clone a parent's or grandparent's voice so a child can hear a familiar person read to them. The living-relative version of this is possible under our cloning policy with their recorded consent, and it is a genuinely lovely use. The deceased-relative version is not, because consent cannot be obtained, and no amount of family agreement substitutes for it.

Cloning a child's voice is never available, under any plan or agreement. A model of an identifiable child's voice is a tool for impersonating that child, and there is no creative requirement that justifies creating one.

Length, and the bedtime constraint

Children's audio has a hard external constraint that no other format has: a substantial share of it is played at bedtime, by a tired adult, to a child who is supposed to fall asleep. That shapes everything. Episodes of eight to twelve minutes fit the ritual; twenty-five minute episodes do not, and parents quietly stop using them.

It also shapes the ending. A story that finishes on an exciting cliffhanger is a story that wakes a child up. The convention in bedtime audio is a deliberate wind-down in the last ninety seconds — slower pace, quieter delivery, a resolution rather than a hook — and it is the single most requested feature by the adults who actually press play.

Testing with the actual audience

Adults are unreliable judges of children's audio, consistently choosing material that is too fast, too dense and too clever. The correction is trivially available: play it to a child and watch.

You are looking for two things. Whether they can follow who is speaking without asking, which tells you the casting is separated enough. And where their attention goes, which is almost never where you expected. Ninety seconds of observation will teach you more than any amount of theory about pacing, and it will usually tell you to slow down again.

What it costs

Free

$0 / month

Enough to finish a short scene and hear what your script sounds like cast.

  • 10,000 characters every month — the quota resets, it is not a one-time trial
  • 2 cast members per project
  • 5 minutes of finished audio
  • Exports carry a short CastDub watermark
Start free

Studio

$29 / month

For serialised work: long-running dramas, audiobooks, dubbing pipelines.

  • Unlimited cast members per project
  • 200 minutes of finished audio a month
  • Line-by-line emotion control
  • Multitrack and stem export (one WAV per character)
  • 3 voice clone slots
Choose plan

Team

$69 / month

For a studio where a writer, a director and an editor touch the same project.

  • 5 seats, shared projects and shared character cards
  • 500 minutes of finished audio a month
  • 10 voice clone slots
  • Priority rendering queue
Choose plan

Full comparison →

Frequently asked questions

Are the child voices real children?

No. They are synthetic models, and no recording of a real child is involved in generating them. We also do not build voice clones of minors under any circumstances, including with parental consent — it is the one place our cloning policy has no exception.

How slow should children's narration be?

Noticeably slower than adult audiobook narration, which itself runs around 150 words per minute. Start below that and check with an actual child if you can; adults consistently overestimate the right speed.

Should children's dialogue be read by child voices?

Dialogue yes, prose no. The convention of a warm adult narrator carrying the story with child voices only for speech is near-universal in published children's audio because young listeners find a steady adult voice easier to follow over time.

Can I use this for classroom material?

Yes, on any paid plan, and export the captions — many schools require them on audio-visual material. The slower pacing advice applies even more strongly to educational content.

Is it suitable for a bedtime story podcast?

It is one of the most common uses. Creator suits a weekly ten-minute episode comfortably. Keep the dynamic range narrow so a loud moment does not wake a child who has fallen asleep.

What about scary parts?

Children's stories can and should have tension, but handle it through pacing and pauses rather than through volume or harsh voices. A quiet, slow delivery of something ominous is both more effective and much less likely to genuinely frighten a small listener.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.