Cette page n'est pas encore disponible en français ; la version anglaise est affichée.

Synthèse

Free Online Text to Speech, Nothing to Install

Paste text, choose a voice, press play. CastDub's text to speech runs in the browser: forty AI voices in nine languages, an MP3 to download, and a free plan that does not ask for a card. It is built for scripts with several speakers, but it works just as well for one.

En un coup d'œil

Ce que c'est
Un générateur de voix IA et un outil de synthèse vocale gratuits, pensés pour les scripts : il repère qui parle, donne à chaque personnage sa propre voix, vous laisse diriger l'émotion réplique par réplique et exporte un mixage fini — une fiction sonore entièrement distribuée, pas un narrateur unique qui lit tous les rôles.
Ce que vous collez
Un texte collé. Format scénario, « Nom : réplique » et prose ordinaire avec incises de dialogue sont lus tels quels, tout comme les dialogues au tiret de l'espagnol, du français, du portugais et du russe. Une ligne entièrement entre parenthèses est traitée comme une didascalie et n'est pas dite. Il n'y a pas d'import de fichier : collez le texte plutôt que de joindre un document.
Ce que vous récupérez
Un MP3 mixé dans toutes les formules. À partir de Plus, des sous-titres SRT écrits d'après votre texte et non transcrits depuis l'audio : les noms sont bien orthographiés et la synchro est juste d'emblée. Les pistes séparées par personnage, une piste pleine longueur par personnage plus le mixage, arrivent en ZIP dans Pro. Pas d'export WAV.
Personnages et voix de base
40 personnages en bibliothèque, chacun avec un portrait, un profil et un extrait. Derrière eux, 10 voix de base – 4 féminines, 5 masculines, 1 androgyne – les mêmes 10 dans toutes les langues. Dans un même projet, deux personnages ne reçoivent pas la même voix de base tant que les 10 ne sont pas prises : une scène jusqu'à 10 locuteurs se sépare donc proprement ; au-delà, deux personnages partagent une voix et se distinguent par le profil et par la direction portée par chaque réplique. Vous choisissez le personnage, pas la voix de base.
Langues
9 : English, 简体中文, 日本語, 한국어, Deutsch, Français, Español, Português, Русский. Chaque personnage joue dans toutes – un personnage est un profil, pas une prise enregistrée – et la langue d'un projet est fixée à sa création. Hors anglais, vous obtenez l'accent courant de la langue ; les accents régionaux ne se choisissent pas et ne sont pas promis. Nous n'avons pas fait écouter chaque langue par des locuteurs natifs : faites passer vos propres répliques avant de vous décider.
Émotion et direction réplique par réplique
Chaque réplique est jouée par un comédien IA à partir d'une indication de jeu écrite : les 7 émotions (neutre, joyeux, triste, en colère, peur, chuchoté, enthousiaste) sont donc jouées, pas imitées par réglages. Une émotion de base par personnage existe dans toutes les formules ; la remplacer réplique par réplique n'existe que dans Pro, et le réglage de vitesse à 5 crans (0.75×–1.25×) vient avec Plus et Pro. Régénérer une réplique donne une prise réellement différente, dans la même voix, et modifier une ligne ne resynthétise que cette ligne.
Prix et bêta ouverte
Un crédit vaut un caractère de texte, et les crédits ne partent qu'au téléchargement : générer et écouter ne coûte rien, dans la limite de 30 répliques par 24 heures en Free et jusqu'à 2,000 en Pro. La formule Free est une réserve unique de 4,500 crédits, environ 5 minutes de rendu fini, avec 2 personnages par projet. Les formules payantes sont à Basic $5, Plus $19, Pro $29 par mois, 10% de remise en annuel. Le paiement n'est pas encore en ligne : pendant la bêta ouverte, chaque compte dispose gratuitement de toutes les fonctions Pro, mais la licence commerciale reste liée à une formule payante, une fois les formules ouvertes.
Ce qui n'est pas encore possible
Le clonage de voix n'est proposé dans aucune formule. Il n'y a pas d'import de fichier, pas d'export WAV, pas de dictionnaire de prononciation, pas de réglage de hauteur, et aucun moyen d'exiger une voix de base précise ou un accent régional. Aucune langue n'a de vraie voix d'enfant : les rôles d'enfants sont joués par des comédiens adultes, comme l'animation l'a toujours fait. Les personnages n'existent qu'à l'intérieur d'un projet et ne passent pas au suivant.

Des voix pour l'essayer

Noor

Noor

Newsreader clarity for exposition-heavy scenes

féminineadulteanglais américaininfonet
Utiliser cette voix
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

masculineadulteanglais américainnarrationdocumentaire
Utiliser cette voix
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

féminineadulteanglais américainnarrationchaleureux
Utiliser cette voix
Quill

Quill

YouTube explainer voice that never sounds bored

masculinejeune adulteanglais américainYouTubenet
Utiliser cette voix
Dorian

Dorian

Audiobook baritone for literary fiction and slow reveals

masculineadulteanglais britanniquelivre audiograve
Utiliser cette voix
Solène

Solène

Continental accent, unhurried, very expensive sounding

féminineadulteanglais britanniqueélégantaccent
Utiliser cette voix
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

neutreadulteanglais américainposéapaisant
Utiliser cette voix
Juno

Juno

Podcast host who thinks out loud and never reads a script

fémininejeune adulteanglais américainpodcastconversationnel
Utiliser cette voix

Parcourir toute la bibliothèque →

Ce que ça fait

Works in the browser

Nothing to download or install. Open the studio in any modern browser, paste, listen. Projects are saved to your account, so a script you started on one machine is there on the next.

Forty voices, nine languages

Every voice speaks English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, and every voice is on every plan including Free. Pick by persona — a newsreader, a bedtime narrator, an audiobook baritone — rather than by locale code.

MP3 download

Export a mixed MP3 on any plan, or a ZIP with one MP3 per line and a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV option.

More than one speaker, when the text needs it

Write a name and a colon at the start of a line and that line goes to a second voice; everything else stays with the narrator. Ordinary online TTS gives you one voice per file. This one gives you a cast the moment the text turns into a conversation.

Direction on the line

A note in parentheses at the start of a line — (whispering), (fast) — is performed rather than read out. On Plus and Pro each line has its own pace; on Pro, its own emotion. Punctuation is timing on every plan: a full stop pauses, a line break pauses longer.

Free plan, no credit card

A one-time balance worth about five minutes of finished English audio — more in Chinese, Japanese and Korean — and generating or auditioning does not spend it; only downloading does. The demo on the homepage runs without an account at all.

Comment s'en servir

  1. Paste the text

    One paragraph or a whole script, as plain text. There is no file upload, so copy it out of the document. Names followed by a colon mark a second speaker; leave them out for a single voice.

  2. Pick a voice

    Filter the forty characters by gender, age and style, then audition on your own first sentence rather than on sample text. A voice that suits the sample and not your text is a common way to waste ten minutes.

  3. Listen line by line

    Each line is rendered on its own and takes a few seconds; generating the whole script runs several lines at once. Play the first few before you commit to the rest.

  4. Fix it in the text

    A mispronounced name gets respelled the way it sounds; a sentence that runs too fast gets split; a number that must be read a particular way is written out in words. There is no pronunciation setting, and you will not miss it.

  5. Download

    A mixed MP3, or one file per line with a CSV of durations. Downloading is what spends the balance; listening does not.

What "free online text to speech" usually leaves out

Most tools that rank for the phrase are free in one of three ways: free for a few hundred characters, free with the voices you would actually want locked behind a plan, or free until you try to download. It is worth being precise about which kind this is. All forty voices are available on the Free plan. Listening and regenerating cost nothing. The balance — roughly five minutes of finished English audio, more in Chinese, Japanese and Korean — is spent only when you download, and it is a one-time balance rather than a monthly one. The Free plan also caps generation at 30 lines a day and is licensed for personal, non-commercial use.

Paid plans change the numbers, not the voices: a monthly balance in place of the one-time one, more lines a day, commercial use, and on Plus and Pro the per-line pace control and subtitles. Nothing about the sound of the voices changes when you pay.

Getting a natural read out of any text to speech

Text written for the eye contains things that do not survive being spoken: long parenthetical asides, three nested clauses, a paragraph that relies on the reader glancing back. Ten minutes of adaptation — splitting sentences, cutting an aside, turning a list into a line per item — does more for the result than any voice setting. This is not a workaround for a limitation; it is the same edit a human narrator would ask for before a session.

Two habits matter most. Punctuate deliberately, because punctuation is the only timing control you have on the Free plan and the main one on any plan. And keep one idea per line: a sentence that swings from good news to bad is rendered as an average of both, and splitting it is the fix.

Numbers, dates and acronyms follow the conventions of the project's language and are usually right; text that mixes conventions is where mistakes appear. Audition the lines with figures in them first, and spell out anything whose reading matters.

When this is the right tool, and when it is not

If you want a web page or a PDF read aloud while you do something else, a reader extension is a better fit — it follows the text and does not ask you to paste. If you need a single voice reading an hour of prose with no direction, almost any online text to speech will do it, and some are cheaper per hour.

CastDub earns its place the moment the text has more than one voice in it, or the moment a line needs to be said a particular way. A dialogue, a presenter with a customer, a story with a narrator and characters, an explainer where the call to action should land differently from the paragraph before it: these are the cases where one flat voice per file stops being enough, and where casting and per-line direction — which is what this text to speech is built around — start to matter.

Questions fréquentes

Do I need to install anything?

No. It runs in the browser; there is no desktop app and no extension. Sign in with Google or an email address, paste, and listen. The demo on the homepage does not even need an account.

Is it actually free?

The Free plan is free, with limits stated on the pricing page: a one-time balance worth about five minutes of finished English audio, 30 generated lines a day, all forty voices, and personal non-commercial use. No credit card is taken at sign-up.

Which languages does it read?

English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian — every voice in every one of them. Outside English each voice uses the general standard of the language; regional accents cannot be chosen.

Can I download the audio?

Yes, as MP3: a single mixed file on every plan, or one file per line with a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV export.

Can I use the audio commercially?

On Basic, Plus and Pro, yes. The Free plan is for personal, non-commercial use — a licensing term, not a technical one; the audio is identical.

Can I upload a document?

No, only pasted text. Copy the text out of the file and paste it in; standard screenplay layout and "Name: line" pairs are recognised as speakers, everything else is read by the narrator.

Distribuez votre première scène ce soir

Collez un script, laissez CastDub donner sa voix à chaque personnage, et exportez une fiction sonore finie.

Créer gratuitement

Sans carte bancaire. Free vous donne une réserve unique de crédits.