
Lumen
Cute on top, deadpan underneath
Voice category
A VTuber persona is a character, and characters need a consistent voice. CastDub is built for the scripted half of that work — lore videos, skits, intros, the clips where two characters talk — with a character card that keeps the voice identical months apart.

Lumen
Cute on top, deadpan underneath

Aria
Enthusiastic, slightly too loud

Tess
Dry, unimpressed, very online

Pip
The gremlin mascot everyone asks about
Samples are placeholders while the public demo audio is produced. Every voice is available on every plan — the library is not a paywall.
A VTuber channel is usually two production lines. Live streaming, which is unscripted and real-time, and everything else — intros, outros, lore videos, animated skits, clip compilations with added narration — which is scripted, edited and often the part that actually grows the channel. The second line is where a rendering tool belongs, and it is a surprisingly large share of the work.
It is also the half where consistency is hardest. Live, you are whoever you are that day. In scripted content, viewers compare a clip from January to one from October, and any drift in the persona's voice is immediately audible.
The most common mistake is auditioning voices before deciding who the character is. Without a persona, every voice sounds wrong in a way you cannot articulate, and the search becomes endless.
A workable persona brief is three lines long: an age band, an energy level, and a characteristic attitude. "Nineteen, high energy, sincerely enthusiastic about boring things" narrows a forty-voice library to about four candidates, and those four can be auditioned in ten minutes.
The channels that hold audiences over years tend to have more than one character: a mascot, a rival, a recurring bit character. That is difficult to produce when every voice costs a recording session and trivial when it does not, which is why multi-character lore content is one of the clearest wins here.
Give each of them a character card and keep them in one project per series. Then a two-year-old side character can reappear sounding exactly as they did, which is the kind of continuity audiences notice and reward.
Two norms matter and both are about honesty. Disclose that a voice is generated — the community's tolerance for this is much higher than its tolerance for finding out later. And never clone another creator's voice, even affectionately; in a scene built on persona identity it is read as identity theft, and correctly so.
Our cloning flow requires a recorded consent statement from the person whose voice is being cloned, which makes the legitimate case — cloning your own voice so you can script faster — straightforward and the illegitimate case impossible.
A persona that exists both live and in scripted content has a consistency problem that no tool solves entirely: your live voice and a generated voice will not match. The channels that handle this well do not try to hide the difference — they use the generated voice for a distinct kind of content, such as lore videos, animated segments or the mascot character, and let the live streams be themselves.
Trying to make scripted clips sound like the live persona invites a comparison you cannot win. Giving the scripted content its own identity sidesteps it entirely, and often produces something the channel did not have before.
The most durable channels in this space have a cast: a mascot, a rival, a recurring bit character with three catchphrases. Previously that meant a creator doing several voices themselves, which is a specific skill most people do not have.
A cast that costs nothing to add changes what kind of content is possible, and the pattern worth copying is a small number of very distinct characters who reappear rather than a large number of one-offs. Audiences form attachments to characters they see repeatedly, and attachments are what bring people back.
Age, energy level, what they are sarcastic about. A persona document makes casting fast; without one you will audition forty voices and like none of them.
Use lines your persona would actually say — a greeting, a complaint, a bit. Neutral sample text tells you nothing about whether the voice fits a character.
Voice, pace and default emotion travel together. Every future clip loads the card, which is how the persona stays identical across a year of uploads.
Lore videos and skits need a second and third voice. Casting them properly is what separates a channel with a world from a channel with a filter.
Clip editing involves a lot of retiming. Per-voice stems let you slide lines around without re-rendering anything.
No. This renders written scripts, it is not a real-time voice changer. For live streaming you want a real-time pitch and formant tool; CastDub is for the scripted content around the stream — intros, lore videos, skits, clip voiceover.
Character cards. Consistency fails when someone re-picks a voice from memory six months later, not because the model changes. Load the card and it is the same voice every time.
You can match the register — age, energy, brightness — which is what viewers actually respond to. Write down three adjectives for the design and filter the library by those style tags rather than auditioning at random.
That is between you and your audience, and the answer most communities land on is that disclosure makes it fine. Undisclosed is where it goes wrong, because VTuber audiences are unusually invested in the person behind the model.
No. It is prohibited here and it is one of the most actively policed forms of harassment in that community. Build an original persona instead.
Creator for solo clips, Studio once you are producing multi-character lore videos — the unlimited cast and stem export are the two features that matter for that format.
Paste a script, let CastDub assign a voice to every character, and export a finished drama.
Create for freeNo credit card. Free plan renews every month.