
Cass
Seventeen and correct about it
Voice category
There are at least four kinds of angry, and they sound nothing alike: cold, hot, exhausted and frightened. Most voice tools only do loud. CastDub lets you cast and direct all four, which is what an argument scene actually needs.

Cass
Seventeen and correct about it

Torvald
Cold command, no volume required

Kestrel
Exhausted fury, the worst kind

Axel
Has never used an indoor voice
Samples are placeholders while the public demo audio is produced. Every voice is available on every plan — the library is not a paywall.
Cold anger is slower and quieter than the character's normal speech, with unusually precise consonants. It signals control, and it is the most threatening. Hot anger is faster, louder, with sentences that collide — it signals loss of control, which reads as danger of a different kind. Exhausted anger is flat and slow and comes from someone who has had this argument before. Frightened anger is fast and high and is the one most often miscast, because writers mark it angry when the character is actually scared.
Deciding which one you are writing before you direct it solves most argument-scene problems. The four are not interchangeable, and a listener identifies them instantly even if they could not name them.
An argument works the way a piece of music works: two lines that are distinguishable and rhythmically related. If both characters escalate at the same rate, in the same register, at the same speed, the scene turns into a wall of sound about forty seconds in and listeners disengage.
The standard fix is asymmetry. One character escalates, the other goes quieter. One speaks in long paragraphs, the other in four-word sentences. This is why interrogation scenes and confrontations between a hothead and a professional are such durable formats — the asymmetry is built into the premise.
Escalation needs a baseline. Open the scene with two or three genuinely neutral lines, even if they are about nothing, so the listener has a reference for what normal sounds like for these characters. Then move the temperature in steps rather than jumping.
In practice this means directing at the line level: neutral, neutral, tense, tense, angry, and then — the most effective move available — back down to neutral for the line after the peak. The drop after an outburst is what makes an argument feel like it happened between people rather than between settings.
Loud material in headphones is genuinely unpleasant if it is not controlled. Keep peaks well under the ceiling, and do not let a shouted line be ten decibels above the surrounding dialogue just because it was performed that way. Broadcast drama keeps the actual dynamic range much narrower than the performance suggests, and uses proximity and tone rather than level to signal volume.
If you export stems, this becomes straightforward: pull the shouted lines down and they still read as shouting, because listeners identify shouting from tone, not from level.
Arguments on the page and arguments in audio have different shapes. On the page, a long speech reads as forceful. In audio, a long speech in an argument reads as a monologue, and the other character disappears. The single most effective edit to a confrontation scene is to break the longest speech into three and give the other character two interruptions.
It also helps to write the subject change. Real arguments do not stay on topic; they escalate by widening, pulling in old grievances that have nothing to do with what started it. That widening is what makes a fictional argument feel real, and it gives the performance somewhere to go without simply getting louder.
Finally, write the ending before the peak. Arguments in fiction almost never resolve at maximum volume — they stop, awkwardly, with someone leaving or someone conceding something small. Directing that final line back down to neutral is what tells the audience the scene happened between people.
Cold anger is quieter and slower than normal speech. Hot anger is faster and louder. Writing one and directing the other is why most argument scenes feel wrong.
An argument between two identical registers is noise. Pair a fast escalating voice with a slow controlled one and the scene organises itself.
Mark early lines neutral, middle lines tense and one late line as full anger. A scene that starts at maximum has nowhere left to go.
Anger shortens sentences. Split any line longer than about fifteen words; the shorter rhythm does more than the emotion setting.
Real arguments have people talking over each other. Render the stems and overlap the last two lines by half a second — it is the single most realistic thing you can do to an argument.
It produces a raised, forceful delivery. A true shout — throat strain, volume that distorts — is a physical limit rather than a stylistic one, and if a scene depends on one full-blooded scream, that is the line to record yourself.
Because volume flattens dynamic range and the listener stops being able to read detail. A character speaking very quietly while furious keeps every detail audible and makes the audience lean in, which is the state you want them in during a confrontation.
Shorten it and slow it slightly. Loud, fast and long is the combination that turns into mush. Loud, short and deliberate stays clear at any volume.
Yes, though the strongest villain monologues are barely raised. Cast a controlled voice, keep the emotion near neutral, and let one specific line break the control.
Threatening audio directed at a real, identifiable person, and any cloned voice used to make someone appear to threaten another. Fictional threats between fictional characters are fine; the line is real people.
Studio gives you line-level emotion control and stem export, and arguments benefit from both more than most scenes. Creator will still produce a solid argument using the automatic emotional pass.
Paste a script, let CastDub assign a voice to every character, and export a finished drama.
Create for freeNo credit card. Free plan renews every month.