Workflow

Script to Dialogue Audio

A script is a plan for a performance, and reading it silently tells you almost nothing about whether it works. Hearing it performed by distinct voices tells you within thirty seconds — and it now takes about that long to arrange.

Why this works here

A table read you can run alone

Table reads exist because writers cannot hear their own dialogue. Getting five people in a room to read a draft is the most valuable and least available feedback in screenwriting; rendering the same read takes minutes and can be repeated after every rewrite.

Screenplay format works without conversion

Character names in caps above their lines, parentheticals, scene headings — the standard layout is parsed directly. You do not have to strip formatting or rebuild the scene as a list.

Dialogue problems become audible

On the page, two characters with the same voice look different because their names are printed. In audio they are identical, which is exactly the diagnostic you want: if you cannot tell them apart by ear, neither can an audience.

It produces something usable, not just a test

The same render is a scratch track for animation, a pitch asset, a rehearsal reference for actors, or a finished audio version of a stage play. The output is not throwaway.

Voices that suit this kind of work

Kestrel

Kestrel

Low, weary swordswoman who has already lost this war once

femaleadulten-USdramaticgritty
Torvald

Torvald

Knight commander who gives orders that men actually follow

maleadulten-GBheroicfantasy
Cass

Cass

Seventeen and furious about it, in a good way

femaleteenen-USteenmodern
Juno

Juno

Podcast host who thinks out loud and never reads a script

femaleyoung-adulten-USpodcastconversational
Ashford

Ashford

Velvet-voiced villain who apologises before he ruins you

maleadulten-GBvillainsmooth
Noor

Noor

Newsreader clarity for exposition-heavy scenes

femaleadulten-USnewsclear
Sable

Sable

Noir detective narrating her own bad decisions

femaleadulten-USnoirsmoky
Hazel

Hazel

Mum voice: patient, then suddenly not

femaleadulten-USwarmdomestic

Browse the whole library →

How to do it

  1. Paste the scene in standard format

    Screenplay layout is parsed as written. Scene headings and action lines can be assigned to a narrator or excluded entirely.

  2. Decide what to do with action lines

    For a table read, have a narrator read them. For a scratch track, drop them. Both are one setting.

  3. Cast for separation, not for looks

    Two characters who read similarly on the page must sound different. This is the check the exercise exists for.

  4. Render the scene straight through

    Do not stop to fix lines. Hear the whole scene once at full length; pacing problems only show up continuously.

  5. Mark the beats

    Second pass: direct the three or four lines where the scene turns. Leave everything else neutral.

  6. Export for its purpose

    A mixed file for feedback, stems for an animatic, subtitles for a pitch video.

The feedback gap in scriptwriting

Writers in every dramatic medium face the same problem: the artefact they produce is not the artefact the audience receives. A script is instructions. Whether those instructions produce a good scene depends on performance, and performance is expensive and slow to arrange, so most scripts are revised many times before anyone hears them at all.

The traditional solution is the table read, and it works so well that professional productions build their schedules around it. The difficulty is access: a table read needs several people, a time everyone is free, and a draft you are willing to expose. Writers early in a project have none of those things, which is exactly when hearing the dialogue would help most.

What you learn from hearing a scene

Three things surface immediately. Length — scenes that read as brisk on the page routinely run twice as long aloud, and the difference is felt rather than measured. Voice distinctness — if two characters sound the same, they are the same, and no amount of character description on the page fixes it. And exposition — information that reads as natural is frequently unbearable when spoken, because a listener cannot skim.

A fourth, subtler one: rhythm. Good dialogue has varying line lengths, and a scene where every line is the same length develops a mechanical pulse that is invisible in text and obvious in audio. Hearing it once is worth more than any amount of advice about it.

Using generated reads without being misled

The failure mode is treating the render as a performance. It is not — it is a neutral reading at a consistent pace, which is precisely what makes it a good diagnostic and a poor judge of subtlety. A line that lands flat here may be excellent with an actor who knows what to do with it, and a line that sounds fine here may still be doing nothing.

The useful question is structural: can I follow this, can I tell who is speaking, is it too long, does the information arrive in a usable order. Those questions the render answers reliably. Questions about nuance it does not answer at all, and pretending otherwise leads to over-editing perfectly good material.

Beyond screenplays

The same workflow covers stage plays, radio scripts, game dialogue tables and interactive fiction. Game writing in particular benefits, because it is written as tables rather than scenes and is almost never heard until very late in production, by which point the structure is fixed.

Stage writing has its own use: a rendered read of a scene gives a director something to react to before casting, and it is common now for new-writing programmes to circulate audio versions of submitted scripts rather than asking readers to imagine them.

Using a rendered read in a writers' room

Rendered reads have found an unexpected use in collaborative writing: circulating a scene in audio rather than as pages. People who will not read eight pages of a colleague's draft will listen to four minutes of it while commuting, and the notes that come back are noticeably different — more about pacing and clarity, less about line-level phrasing.

That shift is useful at the structural stage of a project, when line polishing is premature anyway. Several television writers' rooms now circulate audio versions of outline scenes for exactly this reason.

Timing a scene properly

Screenwriters use a page-per-minute rule that is reliable for action and unreliable for dialogue, because delivery speed varies enormously with how a scene is played. A rendered read gives you a real number rather than an estimate, which matters most in formats with hard length limits: a ten-minute short, an advertisement, a festival submission.

Treat the number as a floor rather than an exact figure. Generated dialogue runs faster and more evenly than human performance, and it does not pause for business, reactions or the moment an actor takes before a difficult line. A scene that renders at four minutes will usually play at four and a half to five.

If you are cutting to length, the useful discovery is almost always which section feels long rather than which section is long. Those are different, and only listening finds the first one.

What it costs

Free

$0 / month

Enough to finish a short scene and hear what your script sounds like cast.

  • 10,000 characters every month — the quota resets, it is not a one-time trial
  • 2 cast members per project
  • 5 minutes of finished audio
  • Exports carry a short CastDub watermark
Start free

Studio

$29 / month

For serialised work: long-running dramas, audiobooks, dubbing pipelines.

  • Unlimited cast members per project
  • 200 minutes of finished audio a month
  • Line-by-line emotion control
  • Multitrack and stem export (one WAV per character)
  • 3 voice clone slots
Choose plan

Team

$69 / month

For a studio where a writer, a director and an editor touch the same project.

  • 5 seats, shared projects and shared character cards
  • 500 minutes of finished audio a month
  • 10 voice clone slots
  • Priority rendering queue
Choose plan

Full comparison →

Frequently asked questions

Is this a replacement for a real table read?

No, and it should not try to be. A room of actors gives you interpretation, questions and the moment where everyone laughs at a line you thought was serious. What this gives you is the structural half — pacing, clarity, whether characters are distinguishable — available at any hour, repeatable after every draft.

Does it read stage directions?

Only if you want it to. Assign action lines to a narrator for a full table read, or exclude them for a dialogue-only scratch track. Parentheticals are treated as direction rather than spoken.

Can I use this as a scratch track for animation?

Yes, and it is a common use. Animate to the generated track, then replace it with the recorded performance later — the timing will not match exactly, so plan for an adjustment pass rather than a drop-in swap.

What formats does it accept?

Plain text in standard screenplay layout, Name: line pairs, and prose with dialogue tags. Exported text from most screenwriting software falls into the first category without modification.

Can I share the render with collaborators?

Export the file and send it however you normally would. There is no collaborative editing in this release; Team plan seats share projects, but the review workflow is still export-and-send.

Will the timing match a real performance?

Roughly. Generated dialogue runs slightly faster and more evenly than human performance, and it does not pause for business or reactions. For rough timing it is reliable; for locking picture it is not.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.