Voice category

Anime Voice Generator

Most anime voice generators hand you one voice and a text box. That is fine for a meme and useless for a scene. CastDub reads your script, notices there are four people talking, and gives each of them a separate performer — the bright lead, the clipped rival, the tired mentor, the gremlin. You direct; it renders.

Four voices to start from

Aria

Aria

Shonen lead who narrates her own fight scenes

teenenergeticheroic
Koji

Koji

Rival: clipped, competitive, secretly loyal

teensharprival
Lumen

Lumen

Cute on the surface, deadpan underneath

vtuberplayful
Torvald

Torvald

Mentor who gives orders men actually follow

adultdeepheroic

Samples are placeholders while the public demo audio is produced. Every voice is available on every plan — the library is not a paywall.

Why one-voice tools fall apart on anime scripts

Anime writing is unusually dense with speakers. A four-page scene can have a lead, a rival, two side characters, an off-screen announcer and a narrator, and half the comedy comes from how fast the voice changes. Feed that into a single-voice text-to-speech tool and every line arrives in the same register, so the reader has to do the casting in their own head — which is exactly the work you wanted the tool to do.

The workaround people invent is to run the tool once per character, download a pile of clips, and assemble them by hand in an audio editor. It works. It also takes an afternoon per scene, and the moment you change a line you do it again. CastDub exists because that loop is the actual bottleneck, not the voice quality.

What auto-casting actually does

Auto-casting reads the script and answers two questions: who is speaking, and what should they sound like. Speaker detection handles the formats people really write in — Name: prefixes, screenplay blocks with the character name centred, and prose with dialogue tags like "she said, quietly". Attribution is not magic; if your scene has three unnamed voices in the dark, you will have to label them yourself.

The second half is the recommendation. Every voice in the library carries a persona, an age band and a set of style tags, so a character described as a bored seventeen-year-old gets matched against teen voices with dry or sharp tags, not against the warm narrator. You can override every suggestion, and the override sticks for the rest of the project — swap a character once and they stay swapped across every scene in that project.

Directing emotion the way a scene needs it

The difference between a read-through and a performance is mostly timing and intensity. CastDub treats emotion as a per-line property: this line is neutral, this one is excited, this whispered aside is a whisper. That granularity matters more in anime than in almost any other genre, because the comedy convention is a hard cut between registers — a character shouting, then immediately muttering.

Line-level emotion control is a paid feature, which is deliberate. The free plan gives you the automatic pass, which is good enough to check that a scene works. Manual direction is what turns a working scene into one you would publish, so it sits on the Studio plan alongside multitrack export.

A practical note on writing: emotion tags do not rescue a line that is doing two things at once. If a sentence starts angry and ends sad, split it into two lines. The renderer treats each line as one performance, and so does a human actor.

Where anime creators use this

The three patterns we see most are fan audio dramas posted to YouTube and Discord, original-character scenes shared inside art communities, and indie visual novels that need placeholder voice acting before there is any budget for a cast. All three share the same constraint: one person, no studio, a script that keeps changing.

For fan work, the safe ground is original scripts with original characters, or transformative work in fandoms that tolerate it. The risky ground is cloning a named performer or reposting a licensed dub. The tool will let you build the first; our cloning policy is written specifically to stop the second.

Script to audio in a few steps

  1. Paste the scene

    Drop in dialogue in any shape you already write it — screenplay format, novel prose with quotation marks, or a plain list of Name: line pairs. CastDub does not need a template.

  2. Let it cast

    Auto-casting splits the scene into speakers and proposes a voice for each one, using the description you wrote for them. If your script says a character is fifteen and furious, you will not get a forty-year-old narrator.

  3. Swap anyone you disagree with

    Click a character, audition four or five alternatives back to back on the same line, and keep the one that sounds like the person in your head. The rest of the scene updates around the change.

  4. Direct the emotion

    Anime dialogue lives on sudden shifts — a joke that turns into a threat inside one breath. Set the emotion per line, not per character, so the shift actually lands.

  5. Export

    Render an MP3 or WAV of the full scene, or take the stems so you can drop the dialogue over your own music and effects in a video editor.

Frequently asked questions

Does this sound like real anime voice acting?

It sounds like a clean, directable read — closer to a table read than to a broadcast dub. The gap between synthetic and human is smallest in conversational lines and widest in screams, sobbing and long held notes. If your scene leans on those, record them yourself and let CastDub carry everything around them.

Can I make a character sound Japanese-accented in English?

Yes, by choosing a voice with that accent rather than by faking it with spelling. Writing dialogue phonetically to force an accent usually produces a caricature and worse pronunciation. Pick the voice, keep the spelling normal.

Can I dub an existing anime clip?

You can generate a dub track and line it up in your editor. CastDub exports SRT timings on paid plans, which is what most people use to sync. Be careful about rights: dubbing a licensed series and publishing it commercially is a licensing problem no tool can solve for you.

How many characters can one scene have?

The free plan allows two cast members per project — enough for a two-hander. Creator raises it to five and Studio removes the cap. Most fan scenes land between three and six speakers, so Creator covers a lot of people.

Can I clone a voice actor I like?

Only with their explicit, documented consent, and never for a public figure or a copyrighted performance. Cloning a voice you do not own is the fastest way to get a project taken down. Our cloning flow asks for a recorded consent statement before it will build a model.

Do I own what I make?

On paid plans you get a commercial licence to the generated audio. The free plan is for personal and non-commercial use and adds a short watermark to exports. The script is yours either way — we do not claim rights to your writing.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.