Synthesis

Free Online Text to Speech, Nothing to Install

Paste text, choose a voice, press play. CastDub's text to speech runs in the browser: forty AI voices in nine languages, an MP3 to download, and a free plan that does not ask for a card. It is built for scripts with several speakers, but it works just as well for one.

At a glance

What it is
A free AI voice generator and text-to-speech workbench built for scripts: it works out who is speaking, gives every character their own voice, lets you direct the emotion line by line and exports a finished mix — a full-cast audio drama rather than one narrator reading every part.
What you paste in
A script, pasted as text. Screenplay format, “Name: line” dialogue and ordinary prose with dialogue tags are all read as written, and so are the em-dash conventions used in Spanish, French, Portuguese and Russian. A note in brackets at the start of a line, or a line that is entirely inside brackets, becomes that line's acting direction: it is performed rather than spoken, and you can edit it in the line list. There is no file import, so paste the text instead of attaching a document.
What you get back
One mixed MP3 on every plan. SRT subtitles from Plus upward — written from your script rather than transcribed from the audio, so names are spelled right and the timings already line up. Per-character stems, one full-length track per character plus the mix, come as a ZIP on Pro. No WAV export.
Cast and voices
40 library characters, each with a portrait, a persona and a sample. Behind them sit 10 base voices — 4 female, 5 male, 1 androgynous — and the same 10 serve every language. Inside one project no two characters are given the same base voice until all 10 are taken, so a scene of up to 10 speakers separates cleanly; past that, two characters share a voice and are told apart by persona and by the direction each line carries. You choose the character, not the base voice.
Languages
9: English, 简体中文, 日本語, 한국어, Deutsch, Français, Español, Português, Русский. Every character performs in all of them — a character is a persona, not one recorded take — and a project's language is fixed when the project is created. Outside English you get the language's common accent; regional accents cannot be picked and are not promised. We have not had native speakers audition every language, so run your own lines through it before you commit.
Emotion and direction
Every line is performed by an AI voice actor working from a written direction, so the 7 emotions (neutral, happy, sad, angry, fear, whisper, excited) are acted rather than faked with settings. A resting emotion per character is available on every plan; overriding emotion line by line is Pro only, and 5-step speed control (0.75×–1.25×) comes with Plus and Pro. Regenerating a line gives a genuinely different take in the same voice, and editing one line re-synthesises only that line.
Price and open beta
A credit is one character of script, and credits are spent only when you download — generating and previewing cost nothing, within a ceiling of 30 lines per 24 hours on Free and up to 2,000 on Pro. Free is a one-time pot of 4,500 credits, about 5 minutes of finished audio, with 2 cast members per project. Paid plans are Basic $5, Plus $19, Pro $29 a month, 10% off yearly. Payments are not live yet: during the open beta every account gets the full Pro feature set at no cost, but a commercial licence still comes with a paid plan once plans open.
What it cannot do yet
Voice cloning is not available on any plan. There is no file import, no WAV export, no pronunciation dictionary, no pitch control, and no way to pick a specific base voice or a regional accent. No language has real child voices: child parts are played by adult performers, the way animation has always done it. Characters live inside one project and do not carry over to the next.

Voices to try it on

Noor

Noor

Newsreader clarity for exposition-heavy scenes

femaleadulten-USnewsclear
Use this voice
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

maleadulten-USnarrationdocumentary
Use this voice
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

femaleadulten-USnarrationwarm
Use this voice
Quill

Quill

YouTube explainer voice that never sounds bored

maleyoung-adulten-USyoutubeclear
Use this voice
Dorian

Dorian

Audiobook baritone for literary fiction and slow reveals

maleadulten-GBaudiobookdeep
Use this voice
Solène

Solène

Continental accent, unhurried, very expensive sounding

femaleadulten-GBelegantaccent
Use this voice
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

neutraladulten-UScalmsoothing
Use this voice
Juno

Juno

Podcast host who thinks out loud and never reads a script

femaleyoung-adulten-USpodcastconversational
Use this voice

Browse the whole library →

What it does

Works in the browser

Nothing to download or install. Open the studio in any modern browser, paste, listen. Projects are saved to your account, so a script you started on one machine is there on the next.

Forty voices, nine languages

Every voice speaks English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, and every voice is on every plan including Free. Pick by persona — a newsreader, a bedtime narrator, an audiobook baritone — rather than by locale code.

MP3 download

Export a mixed MP3 on any plan, or a ZIP with one MP3 per line and a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV option.

More than one speaker, when the text needs it

Write a name and a colon at the start of a line and that line goes to a second voice; everything else stays with the narrator. Ordinary online TTS gives you one voice per file. This one gives you a cast the moment the text turns into a conversation.

Direction on the line

A note in parentheses at the start of a line — (whispering), (fast) — is performed rather than read out. On Plus and Pro each line has its own pace; on Pro, its own emotion. Punctuation is timing on every plan: a full stop pauses, a line break pauses longer.

Free plan, no credit card

A one-time balance worth about five minutes of finished English audio — more in Chinese, Japanese and Korean — and generating or auditioning does not spend it; only downloading does. The demo on the homepage runs without an account at all.

How to use it

  1. Paste the text

    One paragraph or a whole script, as plain text. There is no file upload, so copy it out of the document. Names followed by a colon mark a second speaker; leave them out for a single voice.

  2. Pick a voice

    Filter the forty characters by gender, age and style, then audition on your own first sentence rather than on sample text. A voice that suits the sample and not your text is a common way to waste ten minutes.

  3. Listen line by line

    Each line is rendered on its own and takes a few seconds; generating the whole script runs several lines at once. Play the first few before you commit to the rest.

  4. Fix it in the text

    A mispronounced name gets respelled the way it sounds; a sentence that runs too fast gets split; a number that must be read a particular way is written out in words. There is no pronunciation setting, and you will not miss it.

  5. Download

    A mixed MP3, or one file per line with a CSV of durations. Downloading is what spends the balance; listening does not.

What "free online text to speech" usually leaves out

Most tools that rank for the phrase are free in one of three ways: free for a few hundred characters, free with the voices you would actually want locked behind a plan, or free until you try to download. It is worth being precise about which kind this is. All forty voices are available on the Free plan. Listening and regenerating cost nothing. The balance — roughly five minutes of finished English audio, more in Chinese, Japanese and Korean — is spent only when you download, and it is a one-time balance rather than a monthly one. The Free plan also caps generation at 30 lines a day and is licensed for personal, non-commercial use.

Paid plans change the numbers, not the voices: a monthly balance in place of the one-time one, more lines a day, commercial use, and on Plus and Pro the per-line pace control and subtitles. Nothing about the sound of the voices changes when you pay.

Getting a natural read out of any text to speech

Text written for the eye contains things that do not survive being spoken: long parenthetical asides, three nested clauses, a paragraph that relies on the reader glancing back. Ten minutes of adaptation — splitting sentences, cutting an aside, turning a list into a line per item — does more for the result than any voice setting. This is not a workaround for a limitation; it is the same edit a human narrator would ask for before a session.

Two habits matter most. Punctuate deliberately, because punctuation is the only timing control you have on the Free plan and the main one on any plan. And keep one idea per line: a sentence that swings from good news to bad is rendered as an average of both, and splitting it is the fix.

Numbers, dates and acronyms follow the conventions of the project's language and are usually right; text that mixes conventions is where mistakes appear. Audition the lines with figures in them first, and spell out anything whose reading matters.

When this is the right tool, and when it is not

If you want a web page or a PDF read aloud while you do something else, a reader extension is a better fit — it follows the text and does not ask you to paste. If you need a single voice reading an hour of prose with no direction, almost any online text to speech will do it, and some are cheaper per hour.

CastDub earns its place the moment the text has more than one voice in it, or the moment a line needs to be said a particular way. A dialogue, a presenter with a customer, a story with a narrator and characters, an explainer where the call to action should land differently from the paragraph before it: these are the cases where one flat voice per file stops being enough, and where casting and per-line direction — which is what this text to speech is built around — start to matter.

Frequently asked questions

Do I need to install anything?

No. It runs in the browser; there is no desktop app and no extension. Sign in with Google or an email address, paste, and listen. The demo on the homepage does not even need an account.

Is it actually free?

The Free plan is free, with limits stated on the pricing page: a one-time balance worth about five minutes of finished English audio, 30 generated lines a day, all forty voices, and personal non-commercial use. No credit card is taken at sign-up.

Which languages does it read?

English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian — every voice in every one of them. Outside English each voice uses the general standard of the language; regional accents cannot be chosen.

Can I download the audio?

Yes, as MP3: a single mixed file on every plan, or one file per line with a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV export.

Can I use the audio commercially?

On Basic, Plus and Pro, yes. The Free plan is for personal, non-commercial use — a licensing term, not a technical one; the audio is identical.

Can I upload a document?

No, only pasted text. Copy the text out of the file and paste it in; standard screenplay layout and "Name: line" pairs are recognised as speakers, everything else is read by the narrator.

Cast your first scene tonight

Paste a script, let CastDub assign a voice to every character, and export a finished drama.

Create for free

No credit card. Free gives you a one-time pot of credits.