Esta página todavía no está disponible en español; te mostramos la versión en inglés.

Síntesis

Free Online Text to Speech, Nothing to Install

Paste text, choose a voice, press play. CastDub's text to speech runs in the browser: forty AI voices in nine languages, an MP3 to download, and a free plan that does not ask for a card. It is built for scripts with several speakers, but it works just as well for one.

De un vistazo

Qué es
Un generador de voz IA y una herramienta de texto a voz gratis, pensados para guiones: averigua quién habla, da a cada personaje su propia voz, te deja dirigir la emoción línea por línea y exporta una mezcla terminada; una ficción sonora con reparto completo, no un único narrador leyendo todos los papeles.
Qué pegas
Un guion, pegado como texto. El formato de guion, «Nombre: línea» y la prosa corriente con acotaciones de diálogo se leen tal cual, igual que los diálogos con raya del español, el francés, el portugués y el ruso. Una línea que esté entera entre paréntesis se toma como acotación y no se dice. No hay importación de archivos: pega el texto en lugar de adjuntar un documento.
Qué obtienes
Un MP3 mezclado en todos los planes. Desde Plus, subtítulos SRT escritos a partir de tu guion y no transcritos del audio, así que los nombres van bien escritos y los tiempos ya cuadran. Las pistas por personaje, una pista de la misma duración por cada personaje más la mezcla, llegan en un ZIP en Pro. Sin exportación a WAV.
Personajes y voces base
40 personajes en la biblioteca, cada uno con retrato, perfil y muestra. Detrás hay 10 voces base – 4 femeninas, 5 masculinas, 1 andrógina – y son las mismas 10 en todos los idiomas. Dentro de un proyecto, dos personajes no reciben la misma voz base mientras queden libres: una escena de hasta 10 hablantes se separa limpiamente; a partir de ahí dos personajes comparten voz y se distinguen por el perfil y por la dirección que lleva cada línea. Eliges el personaje, no la voz base.
Idiomas
9: English, 简体中文, 日本語, 한국어, Deutsch, Français, Español, Português, Русский. Cada personaje actúa en todos ellos – un personaje es un perfil, no una toma grabada – y el idioma de un proyecto queda fijado al crearlo. Fuera del inglés obtienes el acento común de cada idioma; los acentos regionales no se pueden elegir ni se prometen. No hemos hecho que hablantes nativos escuchen cada idioma una por una, así que pasa tus propias líneas antes de decidirte.
Emoción y dirección línea por línea
Cada línea la interpreta un actor de voz de IA siguiendo una indicación de interpretación escrita, así que las 7 emociones (neutral, alegría, tristeza, ira, miedo, susurro, entusiasmo) están actuadas, no imitadas con ajustes. La emoción de base por personaje está en todos los planes; sustituirla línea por línea solo está en Pro, y el control de velocidad de 5 pasos (0.75×–1.25×) viene con Plus y Pro. Regenerar una línea da una toma realmente distinta con la misma voz, y cambiar una línea vuelve a sintetizar solo esa línea.
Precio y beta abierta
Un crédito es un carácter del guion y los créditos solo se gastan al descargar: generar y escuchar no cuesta nada, dentro de un tope de 30 líneas cada 24 horas en Free y hasta 2,000 en Pro. El plan Free es una bolsa única de 4,500 créditos, unos 5 minutos de audio terminado, con 2 personajes por proyecto. Los planes de pago son Basic $5, Plus $19, Pro $29 al mes, con un 10% de descuento en anual. Los pagos aún no están activos: durante la beta abierta cada cuenta tiene gratis todas las funciones de Pro, pero la licencia comercial sigue aparte y llega con un plan de pago cuando los planes abran.
Lo que todavía no puede hacer
La clonación de voz no está disponible en ningún plan. No hay importación de archivos, ni exportación a WAV, ni diccionario de pronunciación, ni control de tono, ni forma de pedir una voz base concreta o un acento regional. Ningún idioma tiene voces infantiles reales: los papeles de niño los interpretan actores adultos, como ha hecho siempre la animación. Los personajes existen solo dentro de un proyecto y no pasan al siguiente.

Voces para probarlo

Noor

Noor

Newsreader clarity for exposition-heavy scenes

femeninaadultainglés estadounidensenoticiasclara
Usar esta voz
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

masculinaadultainglés estadounidensenarracióndocumental
Usar esta voz
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

femeninaadultainglés estadounidensenarracióncálida
Usar esta voz
Quill

Quill

YouTube explainer voice that never sounds bored

masculinajoven adultainglés estadounidenseYouTubeclara
Usar esta voz
Dorian

Dorian

Audiobook baritone for literary fiction and slow reveals

masculinaadultainglés británicoaudiolibrograve
Usar esta voz
Solène

Solène

Continental accent, unhurried, very expensive sounding

femeninaadultainglés británicoeleganteacento
Usar esta voz
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

neutraadultainglés estadounidensecalmareconfortante
Usar esta voz
Juno

Juno

Podcast host who thinks out loud and never reads a script

femeninajoven adultainglés estadounidensepodcastconversacional
Usar esta voz

Explora la biblioteca completa →

Qué hace

Works in the browser

Nothing to download or install. Open the studio in any modern browser, paste, listen. Projects are saved to your account, so a script you started on one machine is there on the next.

Forty voices, nine languages

Every voice speaks English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, and every voice is on every plan including Free. Pick by persona — a newsreader, a bedtime narrator, an audiobook baritone — rather than by locale code.

MP3 download

Export a mixed MP3 on any plan, or a ZIP with one MP3 per line and a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV option.

More than one speaker, when the text needs it

Write a name and a colon at the start of a line and that line goes to a second voice; everything else stays with the narrator. Ordinary online TTS gives you one voice per file. This one gives you a cast the moment the text turns into a conversation.

Direction on the line

A note in parentheses at the start of a line — (whispering), (fast) — is performed rather than read out. On Plus and Pro each line has its own pace; on Pro, its own emotion. Punctuation is timing on every plan: a full stop pauses, a line break pauses longer.

Free plan, no credit card

A one-time balance worth about five minutes of finished English audio — more in Chinese, Japanese and Korean — and generating or auditioning does not spend it; only downloading does. The demo on the homepage runs without an account at all.

Cómo se usa

  1. Paste the text

    One paragraph or a whole script, as plain text. There is no file upload, so copy it out of the document. Names followed by a colon mark a second speaker; leave them out for a single voice.

  2. Pick a voice

    Filter the forty characters by gender, age and style, then audition on your own first sentence rather than on sample text. A voice that suits the sample and not your text is a common way to waste ten minutes.

  3. Listen line by line

    Each line is rendered on its own and takes a few seconds; generating the whole script runs several lines at once. Play the first few before you commit to the rest.

  4. Fix it in the text

    A mispronounced name gets respelled the way it sounds; a sentence that runs too fast gets split; a number that must be read a particular way is written out in words. There is no pronunciation setting, and you will not miss it.

  5. Download

    A mixed MP3, or one file per line with a CSV of durations. Downloading is what spends the balance; listening does not.

What "free online text to speech" usually leaves out

Most tools that rank for the phrase are free in one of three ways: free for a few hundred characters, free with the voices you would actually want locked behind a plan, or free until you try to download. It is worth being precise about which kind this is. All forty voices are available on the Free plan. Listening and regenerating cost nothing. The balance — roughly five minutes of finished English audio, more in Chinese, Japanese and Korean — is spent only when you download, and it is a one-time balance rather than a monthly one. The Free plan also caps generation at 30 lines a day and is licensed for personal, non-commercial use.

Paid plans change the numbers, not the voices: a monthly balance in place of the one-time one, more lines a day, commercial use, and on Plus and Pro the per-line pace control and subtitles. Nothing about the sound of the voices changes when you pay.

Getting a natural read out of any text to speech

Text written for the eye contains things that do not survive being spoken: long parenthetical asides, three nested clauses, a paragraph that relies on the reader glancing back. Ten minutes of adaptation — splitting sentences, cutting an aside, turning a list into a line per item — does more for the result than any voice setting. This is not a workaround for a limitation; it is the same edit a human narrator would ask for before a session.

Two habits matter most. Punctuate deliberately, because punctuation is the only timing control you have on the Free plan and the main one on any plan. And keep one idea per line: a sentence that swings from good news to bad is rendered as an average of both, and splitting it is the fix.

Numbers, dates and acronyms follow the conventions of the project's language and are usually right; text that mixes conventions is where mistakes appear. Audition the lines with figures in them first, and spell out anything whose reading matters.

When this is the right tool, and when it is not

If you want a web page or a PDF read aloud while you do something else, a reader extension is a better fit — it follows the text and does not ask you to paste. If you need a single voice reading an hour of prose with no direction, almost any online text to speech will do it, and some are cheaper per hour.

CastDub earns its place the moment the text has more than one voice in it, or the moment a line needs to be said a particular way. A dialogue, a presenter with a customer, a story with a narrator and characters, an explainer where the call to action should land differently from the paragraph before it: these are the cases where one flat voice per file stops being enough, and where casting and per-line direction — which is what this text to speech is built around — start to matter.

Preguntas frecuentes

Do I need to install anything?

No. It runs in the browser; there is no desktop app and no extension. Sign in with Google or an email address, paste, and listen. The demo on the homepage does not even need an account.

Is it actually free?

The Free plan is free, with limits stated on the pricing page: a one-time balance worth about five minutes of finished English audio, 30 generated lines a day, all forty voices, and personal non-commercial use. No credit card is taken at sign-up.

Which languages does it read?

English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian — every voice in every one of them. Outside English each voice uses the general standard of the language; regional accents cannot be chosen.

Can I download the audio?

Yes, as MP3: a single mixed file on every plan, or one file per line with a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV export.

Can I use the audio commercially?

On Basic, Plus and Pro, yes. The Free plan is for personal, non-commercial use — a licensing term, not a technical one; the audio is identical.

Can I upload a document?

No, only pasted text. Copy the text out of the file and paste it in; standard screenplay layout and "Name: line" pairs are recognised as speakers, everything else is read by the narrator.

Reparte tu primera escena esta noche

Pega un guion, deja que CastDub le dé una voz propia a cada personaje y exporta un audiodrama terminado.

Empieza gratis

Sin tarjeta. Free te da una bolsa única de créditos.