Voice category

Sad Voice Generator

Sadness in audio is almost always underplayed. The voice that breaks on every line stops being sad and becomes noise; the voice holding itself together is the one that lands. CastDub lets you set that restraint per line, so the one moment it slips actually counts.

Four voices to start from

Orrin

Orrin

Remembers the flood, will not discuss it

seniorrough

Samples are placeholders while the public demo audio is produced. Every voice is available on every plan — the library is not a paywall.

The restraint principle

There is a reason stage and radio directors spend most of their notes on emotional scenes telling performers to do less. An audience reads emotion by inference. If a character openly displays grief, the audience observes it. If a character is visibly trying not to display grief, the audience does the work of imagining what is underneath, and involvement is what makes a scene move someone.

This translates directly into how you should direct a synthetic performance. The instinct is to mark the scene as sad and let the renderer deliver sadness, and the result is uniformly mushy. The better approach is to keep the scene deliberately level and place the emotion in one or two specific places where the level breaks.

What sadness sounds like technically

Acoustically, sadness shows up as reduced pace, reduced pitch range, shorter phrases, and longer gaps between them. Note what is not on that list: volume, and pitch floor. A sad voice is not primarily a low voice. It is a voice that has stopped varying.

Because reduced variation is the marker, over-directing is self-defeating. Pushing intensity increases variation, which reads as agitation rather than grief. If you want a line to sound devastated, take things away from it.

Structuring an emotional scene

A structure that consistently works: an ordinary exchange about something practical, a line that touches the real subject, an immediate return to the practical, and then one line late in the scene where the practical conversation fails. This is how people actually behave, and it gives you natural places to put the two emotional marks a scene can support.

Scenes that open at full emotional intensity have nowhere to go, which is why they feel exhausting. Give yourself a floor to fall from.

Honest limitations

Current synthesis handles controlled, restrained sadness well and uncontrolled emotion poorly. Sobbing, breaking down mid-sentence, speaking while crying — these are where a listener will identify the audio as generated. That is not a reason to avoid emotional writing; it is a reason to write the restrained version, which is usually the better version anyway.

If a scene genuinely requires an unrestrained performance, record that line yourself or with a performer and cut it into the generated scene. Mixing human and generated audio in one project is completely normal and nobody hears the seam when the writing carries the moment.

Two mistakes that ruin a grief scene

The first is scoring it too early. Music under an emotional scene is enormously tempting and it does the audience's work for them — it tells them how to feel before they have decided. Radio drama directors traditionally hold music back until after the emotional beat has landed, so the music confirms the feeling rather than announcing it. If you are exporting stems, try one pass with no music at all and see whether the scene still works. If it does not, the problem is the writing.

The second is trimming the silences in the edit. Grief scenes are full of gaps, and a tight edit that removes them produces something that is technically neat and emotionally flat. Leave the pauses in, even when they feel too long while you are editing — they never feel too long to a listener who is following the story rather than watching a waveform.

Script to audio in a few steps

  1. Write the scene without the crying

    Draft it as if the character is determined not to show anything. What remains on the page is the actual scene; the emotion is what leaks through it.

  2. Cast a voice with weight, not one that sounds sad

    A voice that already sounds mournful gives you nowhere to go. Pick a steady voice and let the direction do the work.

  3. Keep almost every line neutral

    Set the character's default to neutral or restrained and mark only the two or three lines where the composure fails. Contrast is the entire mechanism.

  4. Slow the pace on the turn

    Grief shows up as slowing down and shortening sentences. Drop the pace on the lines that matter rather than adding intensity.

  5. Export and leave the silences

    Resist trimming the gaps. The pauses in a grief scene are doing more work than the lines, and a tight edit destroys them.

Frequently asked questions

Can the voice actually cry?

It can approximate a broken, unsteady delivery, but genuine crying — catching breath, sound collapsing mid-word — is a physical event and the weakest area for synthesis. Write around it. Almost every memorable grief scene in audio drama is a character not crying.

Why does my sad scene sound melodramatic?

Almost certainly because too many lines are marked as emotional. Try setting the whole scene to neutral, listening once, then marking a single line. Most people are shocked by how much stronger it is.

Is a lower, slower voice always sadder?

No, and this is the most common error. Sadness reads through restraint and shortened phrases more than through pitch. A bright voice going very quiet and very careful is devastating; a low voice speaking slowly just sounds tired.

Can I use this for content about real loss?

You can write whatever fiction you like. What you must not do is clone a deceased person's voice, including a relative's, without documented consent from them while living. This request comes up often and our answer does not change.

How long should a grief scene run?

Shorter than you want. Emotional scenes in audio have a shelf life of about ninety seconds before attention drops, regardless of quality. If yours runs four minutes, the fix is structural, not vocal.

Do I need emotion control for this?

Line-level emotion control is on Studio; the automatic pass on free and Creator will vary delivery reasonably but not precisely. For a scene where the emotional turn has to hit one exact line, Studio is the plan that gives you that.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free plan renews every month.