
Cass
Seventeen and furious about it, in a good way
Voice category
Teen characters carry most of young-adult fiction, anime and school-set drama, and they are the easiest to get wrong — cast too young and it reads as a cartoon, too old and the whole premise collapses. These voices sit in the band that actually sounds fifteen to nineteen.

Cass
Seventeen and furious about it, in a good way

Koji
Rival with a two-word maximum per sentence

Aria
The one who volunteers before thinking

Tess
Dry, flat, nothing impresses her
Samples are placeholders while the public demo audio is produced. Every voice is available on every plan — the library is not a paywall.
Adolescent voices are physically distinct from both child and adult voices — the range has settled but the resonance has not, and speech patterns are faster and more clipped than either neighbouring band. Text-to-speech libraries have historically skipped this band entirely, offering child voices and adult voices and nothing between, which is why so much amateur YA audio has teenagers who sound thirty.
Cast from a real teen band and a lot of writing problems disappear on their own. Dialogue that felt stilted in an adult voice often reads as perfectly natural in a seventeen-year-old's rhythm, because that is the rhythm it was written in.
Young-adult fiction runs on posture: the character who deflects with jokes, the one who says nothing, the one who is relentlessly sincere and embarrassing about it. These read as distinct people in prose because you can see their word choices. In audio they read as distinct only if the voices carry different default energy.
Set each character's resting emotion once, then direct the exceptions line by line. The emotional arc of most YA scenes is one long neutral stretch and two lines that break it, and marking exactly those two lines is more effective than pushing the whole scene towards drama.
The three that dominate are anime and manga fan audio, original-character scenes shared in art and writing communities, and vertical short drama, where school-set stories are one of the highest-volume genres worldwide. All three are produced fast, in volume, usually by one person, and all three benefit from being able to re-render a scene after a rewrite instead of scheduling a session.
Visual novels are a fourth, slightly different case: there the teen voices carry hundreds of short branching lines rather than continuous scenes, and character cards matter more than direction, because consistency across branches is what players notice.
Synthetic teen voices are convincing in conversation and less convincing at extremes. Shouting across a room, crying mid-sentence, laughing while talking — these are physical events, and the models approximate rather than reproduce them. Write around the extremes where you can, and where you cannot, consider recording those specific lines yourself and cutting them in.
This is not a permanent limitation, but it is the current one, and it is better to plan around it than to be surprised by it in the final mix.
Fifteen and nineteen are different instruments. Decide which end of adolescence your character sits at, then audition only voices from that band.
Teen scenes are group scenes — friends, a rival, an adult who does not understand. Cast them together so the ensemble balance is visible from the first render.
Most teen characters have a resting register: bored, eager, guarded. Set it as the character's default emotion so you only have to direct the exceptions.
Young-adult writing lives on the moment the armour drops. Mark those specific lines as sad or whispered and leave everything around them neutral — contrast does the work.
Your audience is listening on phone speakers or cheap earbuds. Check the mix there before you publish; high, bright voices behave differently on small drivers.
The library keeps teen voices as their own models rather than pitching adult voices up, which is the usual cause of that effect. The remaining tell is phrasing: if you write long, formally structured sentences, any voice will read as older than intended.
Write slang as ordinary words and it will be spoken naturally. Abbreviations are riskier — some are read as words and some spelled out, and which one you get depends on the abbreviation. Check them once and add pronunciation overrides where it matters.
English voices are the full library today. Japanese, Spanish, Brazilian Portuguese and Korean are the next languages we are adding, in that order, because that is where the audio drama and fan-fiction audiences are.
For first-person YA, yes, and it is a common choice — the narrator is the protagonist, so a teen voice is the correct one. Just budget more editing time: bright voices are less forgiving over an hour than a warm mid-range narrator.
Separate them by pace and attitude rather than by pitch. Two voices with the same energy will blur no matter how far apart their pitch is, whereas a fast eager voice and a slow guarded one stay separate even when they are close in tone.
Avoid sexualised content involving characters written as minors — it is against our terms regardless of how the audio is made, and accounts that generate it are closed.
Paste a script, let CastDub assign a voice to every character, and export a finished drama.
Create for freeNo credit card. Free plan renews every month.