Workflow

AI Audio Drama Generator

An audio drama is a script performed by several people. Every general-purpose text-to-speech tool gives you one performer, which means the hard part — casting, directing, keeping characters distinct — is still yours. This is the page for the tool that does the hard part.

Why this works here

Casting happens automatically

Paste a scene and CastDub identifies who is speaking, in whatever format you already write — Name: prefixes, screenplay blocks, or prose with dialogue tags. Each speaker gets a proposed voice based on how you described them. You override what you disagree with, and the override sticks for the whole project.

Direction is per line, not per character

Drama lives on the moment a character's tone changes. Emotion is a property of the line, so a character can be level for eight lines and break on the ninth. That single mechanism is the difference between a read-through and a performance.

Characters stay themselves across episodes

Character cards carry a name, a persona, a voice, a default emotion and any pronunciation fixes. Episode twelve sounds like episode one because the card is the same object, not because somebody remembered what was chosen months ago.

Exports fit a real workflow

Take a mixed MP3 for a quick share, per-character stems for scoring and sound design, or an SRT if you are publishing with captions. Nothing forces you to finish inside one tool.

Voices that suit this kind of work

Kestrel

Kestrel

Low, weary swordswoman who has already lost this war once

femaleadulten-USdramaticgritty
Ashford

Ashford

Velvet-voiced villain who apologises before he ruins you

maleadulten-GBvillainsmooth
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

femaleadulten-USnarrationwarm
Cass

Cass

Seventeen and furious about it, in a good way

femaleteenen-USteenmodern
Torvald

Torvald

Knight commander who gives orders that men actually follow

maleadulten-GBheroicfantasy
Morrow

Morrow

Something under the floorboards that learned your name

neutraladulten-UShorrorwhisper
Juno

Juno

Podcast host who thinks out loud and never reads a script

femaleyoung-adulten-USpodcastconversational
Bram

Bram

Tavern keeper and accidental quest giver

malesenioren-GBdndwarm

Browse the whole library →

How to do it

  1. Bring your script

    Paste it or upload a text file. Any consistent format works; you do not have to reformat a screenplay into a template first.

  2. Review the cast

    Auto-casting proposes a voice per character. Audition alternatives back to back on the same line, which is the only comparison that actually settles a casting question.

  3. Set defaults per character

    Give each character a resting emotion and pace. Most of a scene runs on the default, so getting it right means you only direct the exceptions.

  4. Direct the turns

    Mark the specific lines where something changes — the confession, the threat, the joke that is not a joke. Leave everything else neutral so those lines have something to land against.

  5. Render and listen end to end

    Play the whole scene without stopping. Problems in an audio drama are almost always pacing problems, and pacing is invisible line by line.

  6. Export in the shape you need

    Mixed file, stems, or stems plus subtitles. If you plan to add music and effects, take the stems — a pre-mixed dialogue track is very hard to score under.

Why audio drama is having a moment

Audio fiction has grown steadily for a decade, driven by the same things that grew podcasting: commuting, chores, and the fact that audio is the only medium you can consume while doing something else. What changed recently is production cost. A drama needs several performers, and several performers means scheduling, studio time and a producer, which put the format out of reach for the people who most wanted to make it.

The result was a genre dominated either by well-funded productions or by solo narrators reading everything themselves. Neither is what audio drama is. The format's whole appeal is hearing a scene between people, and a single reader doing four voices is a compromise that listeners tolerate rather than enjoy.

Generated performance changes the economics rather than the craft. Casting, direction, pacing and writing all still matter — arguably more, because they are now the only things separating a good production from a bad one.

The four decisions that make a drama work

First, casting for separation. Listeners identify speakers primarily by pitch band and speaking rate, with no visual information at all. Two characters in the same band will blur no matter how different their personalities are on the page. Spread your cast deliberately and the scene becomes legible.

Second, restraint in direction. The temptation is to mark every line with an emotion, and the result is uniform mush. A scene with two directed lines and twenty neutral ones is far more affecting than a scene where everything is turned up.

Third, pacing. Audio has no equivalent of skimming. A scene that reads quickly on the page can run four minutes aloud, and listeners feel every second. The fix is almost always cutting, not speeding up.

Fourth, silence. Gaps between lines are where the listener does their work. Directors of radio drama spend most of their time on the gaps, and it is the single most underused tool available to someone producing alone.

A workflow that survives revisions

The reason people abandon audio drama projects is rarely the first render. It is the twentieth revision, when a line change means reassembling an episode by hand. Any workflow that does not make revision cheap will collapse under its own weight.

The structure that holds up is: one project per episode, one character card per recurring character, render per scene rather than per episode, and keep the script as the single source of truth. Change a line, re-render that scene, drop it back into the mix. Nothing else moves.

The corollary is that you should not do your mixing until the script is locked. It is very tempting to add music to a scene that sounds good, and very painful to redo it after a rewrite. Get the performance right first, then score once.

What this does not do

It does not write your script. There is no story generation here, deliberately — the interesting problem in audio drama is production, not ideas, and a generated script performed by generated voices is a product nobody has asked for twice.

It does not mix. No music beds, no reverb, no spatial placement. Those belong in an audio editor where you can hear everything together, and every attempt to fold them into a rendering tool has produced something worse than either.

It also does not handle overlapping dialogue in a single render. People talking over each other is a real dramatic device and it is a mixing operation: export the stems and slide them. That is one drag in any editor and it is the most realistic thing you can do to an argument scene.

What it costs

Free

$0 / month

Enough to finish a short scene and hear what your script sounds like cast.

  • 10,000 characters every month — the quota resets, it is not a one-time trial
  • 2 cast members per project
  • 5 minutes of finished audio
  • Exports carry a short CastDub watermark
Start free

Studio

$29 / month

For serialised work: long-running dramas, audiobooks, dubbing pipelines.

  • Unlimited cast members per project
  • 200 minutes of finished audio a month
  • Line-by-line emotion control
  • Multitrack and stem export (one WAV per character)
  • 3 voice clone slots
Choose plan

Team

$69 / month

For a studio where a writer, a director and an editor touch the same project.

  • 5 seats, shared projects and shared character cards
  • 500 minutes of finished audio a month
  • 10 voice clone slots
  • Priority rendering queue
Choose plan

Full comparison →

Frequently asked questions

What exactly is an audio drama?

A story told entirely in sound: dialogue, narration, effects and music, with no picture. It is the radio-play tradition, and it is having a substantial revival through podcast distribution. The defining feature is that characters are played by different voices, which is precisely what a single-voice tool cannot give you.

How is this different from ordinary text to speech?

Ordinary text to speech converts a block of text with one voice. This takes a script with several speakers, assigns a different voice to each, lets you direct each line separately, and mixes the result. The underlying speech synthesis is one component; the casting and direction layer is the product.

Do I need a specific script format?

No. Screenplay format, Name: lines, and ordinary prose with dialogue tags all work. If a scene is ambiguous — unnamed speakers, a lot of implied attribution — you will need to label those lines yourself, which takes a minute.

How many characters can a project have?

Two on the free plan, five on Creator, unlimited on Studio and Team. A typical audio drama scene has three to six speakers including the narrator, so Creator covers most single-creator work and Studio covers ensemble pieces.

Can I add music and sound effects?

Not inside CastDub — it renders performances, not mixes. Export the stems and score them in any audio editor. This is deliberate: dialogue and sound design are separate crafts and combining them in one tool tends to make both worse.

Can I sell what I produce?

Yes on any paid plan; commercial rights are included and there is no watermark. The free plan is for personal use and adds a short watermark. Your script remains yours in all cases.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.