Voice category

Audiobook Voice Generator

A single-voice audiobook is the standard because hiring a full cast is expensive, not because readers prefer it. CastDub lets one person produce the version with a real cast — narrator for the prose, a distinct voice for every character — from the manuscript they already have.

Four voices to start from

Samples are placeholders while the public demo audio is produced. Every voice is available on every plan — the library is not a paywall.

Why full-cast audiobooks are rare and popular

Listeners consistently rate full-cast productions above single-narrator ones, and publishers produce them rarely, because a cast multiplies studio time, scheduling and cost. The economics, not the preference, is what makes single-narrator the default. Any independent author who has priced a cast production knows exactly how quickly the numbers stop working.

Generating the performance changes the constraint. A cast costs the same as a narrator, so the decision becomes purely editorial: does this book benefit from distinct voices? For dialogue-heavy fiction the answer is almost always yes, and for ensemble fantasy it is emphatically yes, because keeping eight characters straight by ear is exactly what readers complain about.

The parts that still take real work

Three things absorb most of the effort. The pronunciation pass, which is unavoidable in any book with invented vocabulary or unusual names. Casting, which is worth doing slowly because a wrong narrator poisons the whole production. And listening — actually listening to every chapter at normal speed, which is the step people skip and then regret.

What has genuinely disappeared is the recording and re-recording loop. Fixing a typo in chapter nine used to mean a pickup session; now it means re-rendering one line. That single change is what makes independently produced audiobooks feasible for people who are not already audio professionals.

Chapter-level discipline

Treat chapters as the atomic unit for everything: rendering, checking, versioning and fixing. Keep a simple log of which chapter is at which stage, because on a forty-chapter book the state is genuinely hard to hold in your head and a single unchecked chapter is the one listeners will find.

Render settings should be identical across chapters, and this is where character cards earn their place: they carry the voice, pace and emotional default so chapter thirty-one matches chapter two without anyone remembering what was chosen in March.

Disclosure and reader expectations

The market position on synthetic narration is shifting quickly, and the only safe posture is transparency. Label it. Put it in the description, not just in the metadata. Readers who are fine with synthetic narration will buy anyway; readers who are not would have found out in chapter one.

There is also a craft argument for disclosure. It sets the right expectation, and a listener who knows what they are hearing judges it against the right standard — a clear, consistent, well-cast reading — rather than against a performance it was never trying to be.

Structure a listener can navigate

Print readers navigate visually: they see a chapter break, a section break, a change of point of view. Listeners have none of that, so structure has to be carried in the audio. Chapter headings on their own line with a real pause after. A longer gap at a scene break. A consistent way of signalling a point-of-view change, which in a cast production can simply be a change of narrator.

Front matter deserves particular thought. Long dedications, epigraphs and acknowledgements at the start of an audiobook are where listeners bail out, because they do not yet care. Most audio editions move them to the end, and it is worth doing the same with anything that is not the story.

Script to audio in a few steps

  1. Split the manuscript into chapters

    Work in the units you will publish in. A chapter is the right size to re-render after an edit and the right size to check in one listening session.

  2. Cast the narrator before anything else

    The narrator carries the majority of the runtime. Audition on your densest paragraph and listen for three full minutes before deciding.

  3. Let it find the dialogue

    Auto-casting detects quoted speech and dialogue tags in prose, so "she said" attributions become character assignments. Correct the ones it gets wrong and the correction sticks.

  4. Do a pronunciation pass

    Collect every name, place and invented word, check how each is spoken, and fix them once. Twenty minutes here saves re-rendering later chapters.

  5. Render and assemble

    Export each chapter as its own file, plus subtitles if you want a synchronised text edition. Assemble to your distributor's specification at the end.

Frequently asked questions

Can I sell an AI-narrated audiobook?

Commercial rights to the generated audio come with any paid plan. Whether a particular retailer accepts synthetic narration is a separate question with a moving answer — several major platforms now accept it with disclosure, others do not. Check your distributor's current policy before you produce forty hours.

Do I have to disclose that it is AI narrated?

Increasingly yes, and you should regardless. Most retailers that permit synthetic narration require it to be labelled, and listeners who discover it after purchase leave the reviews you would expect.

Can it do different voices for each character?

That is the entire point of the tool. A traditional audiobook has one reader performing everyone; here each character is genuinely a different voice, which makes long dialogue scenes far easier to follow.

How long does a novel take?

The rendering is fast; the work is casting, the pronunciation pass and listening. For a typical novel expect a day or two of real effort, most of it spent checking rather than generating.

What about footnotes, headings and front matter?

Give them their own lines with deliberate pauses. Non-fiction in particular falls apart when headings run into body text, because the listener loses the structure they would have seen on the page.

Which plan do I need?

Studio. Two hundred minutes a month is roughly three to four hours of finished audio, and you will want stem export so a single misread line can be replaced without re-rendering the chapter.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.