Esta página ainda não está disponível em português; exibindo a versão em inglês.

Síntese

Free Online Text to Speech, Nothing to Install

Paste text, choose a voice, press play. CastDub's text to speech runs in the browser: forty AI voices in nine languages, an MP3 to download, and a free plan that does not ask for a card. It is built for scripts with several speakers, but it works just as well for one.

Num relance

O que é
Um gerador de voz IA e ferramenta de texto para fala grátis, feitos para roteiros: descobre quem está falando, dá a cada personagem uma voz própria, deixa você dirigir a emoção fala a fala e exporta uma mixagem pronta — uma radionovela com elenco completo, não um narrador só lendo todos os papéis.
O que você cola
Um roteiro, colado como texto. Formato de roteiro, «Nome: fala» e prosa comum com marcações de diálogo são lidos como estão, assim como os diálogos com travessão do espanhol, do francês, do português e do russo. Uma linha inteira entre parênteses vira rubrica e não é falada. Não existe importação de arquivos: cole o texto em vez de anexar um documento.
O que você recebe
Um MP3 mixado em todos os planos. A partir do Plus, legendas SRT escritas a partir do seu roteiro e não transcritas do áudio, então os nomes saem certos e os tempos já batem. As faixas por personagem, uma faixa do mesmo tamanho para cada personagem mais a mixagem, vêm em ZIP no Pro. Sem exportação em WAV.
Personagens e vozes base
40 personagens na biblioteca, cada um com retrato, perfil e amostra. Atrás deles estão 10 vozes base – 4 femininas, 5 masculinas, 1 andrógina – e são as mesmas 10 em todos os idiomas. Dentro de um projeto, dois personagens não recebem a mesma voz base enquanto as 10 não estiverem ocupadas: uma cena com até 10 falantes se separa direitinho; além disso, dois personagens dividem uma voz e se distinguem pelo perfil e pela direção que cada fala carrega. Você escolhe o personagem, não a voz base.
Idiomas
9: English, 简体中文, 日本語, 한국어, Deutsch, Français, Español, Português, Русский. Todo personagem atua em todos eles – um personagem é um perfil, não uma gravação única – e o idioma de um projeto fica definido na criação. Fora do inglês você recebe o sotaque comum daquele idioma; sotaques regionais não dá para escolher e não são prometidos. Não pedimos a falantes nativos que ouvissem cada idioma um a um, então passe as suas próprias falas antes de decidir.
Emoção e direção fala a fala
Cada fala é interpretada por um ator de voz de IA a partir de uma indicação de interpretação escrita, então as 7 emoções (neutro, alegre, triste, raiva, medo, sussurro, empolgado) são atuadas, não imitadas com ajustes. A emoção padrão por personagem está em todos os planos; trocá-la fala a fala só existe no Pro, e o controle de velocidade em 5 níveis (0.75×–1.25×) vem com Plus e Pro. Gerar uma fala de novo dá uma tomada realmente diferente, na mesma voz, e mexer em uma linha só ressintetiza aquela linha.
Preço e beta aberto
Um crédito é um caractere do roteiro e os créditos só saem quando você baixa: gerar e ouvir não custa nada, dentro de um teto de 30 falas a cada 24 horas no Free e de até 2,000 no Pro. O plano Free é um bolo único de 4,500 créditos, mais ou menos 5 minutos de áudio pronto, com 2 personagens por projeto. Os planos pagos custam Basic $5, Plus $19, Pro $29 por mês, com 10% de desconto no anual. Os pagamentos ainda não estão no ar: durante o beta aberto toda conta tem de graça o conjunto completo de recursos do Pro, mas a licença comercial continua à parte e vem com um plano pago quando os planos abrirem.
O que ainda não dá para fazer
Clonagem de voz não está disponível em nenhum plano. Não há importação de arquivos, nem exportação em WAV, nem dicionário de pronúncia, nem controle de tom, nem como exigir uma voz base específica ou um sotaque regional. Nenhum idioma tem vozes infantis de verdade: papéis de criança são interpretados por atores adultos, como a animação sempre fez. Personagens existem só dentro de um projeto e não passam para o seguinte.

Vozes para testar

Noor

Noor

Newsreader clarity for exposition-heavy scenes

femininaadultoinglês americanojornalismoclara
Usar esta voz
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

masculinaadultoinglês americanonarraçãodocumentário
Usar esta voz
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

femininaadultoinglês americanonarraçãocalorosa
Usar esta voz
Quill

Quill

YouTube explainer voice that never sounds bored

masculinajovem adultoinglês americanoYouTubeclara
Usar esta voz
Dorian

Dorian

Audiobook baritone for literary fiction and slow reveals

masculinaadultoinglês britânicoaudiolivrograve
Usar esta voz
Solène

Solène

Continental accent, unhurried, very expensive sounding

femininaadultoinglês britânicoelegantesotaque
Usar esta voz
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

neutraadultoinglês americanocalmaacalentadora
Usar esta voz
Juno

Juno

Podcast host who thinks out loud and never reads a script

femininajovem adultoinglês americanopodcastconversa
Usar esta voz

Ver a biblioteca inteira →

O que ela faz

Works in the browser

Nothing to download or install. Open the studio in any modern browser, paste, listen. Projects are saved to your account, so a script you started on one machine is there on the next.

Forty voices, nine languages

Every voice speaks English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, and every voice is on every plan including Free. Pick by persona — a newsreader, a bedtime narrator, an audiobook baritone — rather than by locale code.

MP3 download

Export a mixed MP3 on any plan, or a ZIP with one MP3 per line and a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV option.

More than one speaker, when the text needs it

Write a name and a colon at the start of a line and that line goes to a second voice; everything else stays with the narrator. Ordinary online TTS gives you one voice per file. This one gives you a cast the moment the text turns into a conversation.

Direction on the line

A note in parentheses at the start of a line — (whispering), (fast) — is performed rather than read out. On Plus and Pro each line has its own pace; on Pro, its own emotion. Punctuation is timing on every plan: a full stop pauses, a line break pauses longer.

Free plan, no credit card

A one-time balance worth about five minutes of finished English audio — more in Chinese, Japanese and Korean — and generating or auditioning does not spend it; only downloading does. The demo on the homepage runs without an account at all.

Como usar

  1. Paste the text

    One paragraph or a whole script, as plain text. There is no file upload, so copy it out of the document. Names followed by a colon mark a second speaker; leave them out for a single voice.

  2. Pick a voice

    Filter the forty characters by gender, age and style, then audition on your own first sentence rather than on sample text. A voice that suits the sample and not your text is a common way to waste ten minutes.

  3. Listen line by line

    Each line is rendered on its own and takes a few seconds; generating the whole script runs several lines at once. Play the first few before you commit to the rest.

  4. Fix it in the text

    A mispronounced name gets respelled the way it sounds; a sentence that runs too fast gets split; a number that must be read a particular way is written out in words. There is no pronunciation setting, and you will not miss it.

  5. Download

    A mixed MP3, or one file per line with a CSV of durations. Downloading is what spends the balance; listening does not.

What "free online text to speech" usually leaves out

Most tools that rank for the phrase are free in one of three ways: free for a few hundred characters, free with the voices you would actually want locked behind a plan, or free until you try to download. It is worth being precise about which kind this is. All forty voices are available on the Free plan. Listening and regenerating cost nothing. The balance — roughly five minutes of finished English audio, more in Chinese, Japanese and Korean — is spent only when you download, and it is a one-time balance rather than a monthly one. The Free plan also caps generation at 30 lines a day and is licensed for personal, non-commercial use.

Paid plans change the numbers, not the voices: a monthly balance in place of the one-time one, more lines a day, commercial use, and on Plus and Pro the per-line pace control and subtitles. Nothing about the sound of the voices changes when you pay.

Getting a natural read out of any text to speech

Text written for the eye contains things that do not survive being spoken: long parenthetical asides, three nested clauses, a paragraph that relies on the reader glancing back. Ten minutes of adaptation — splitting sentences, cutting an aside, turning a list into a line per item — does more for the result than any voice setting. This is not a workaround for a limitation; it is the same edit a human narrator would ask for before a session.

Two habits matter most. Punctuate deliberately, because punctuation is the only timing control you have on the Free plan and the main one on any plan. And keep one idea per line: a sentence that swings from good news to bad is rendered as an average of both, and splitting it is the fix.

Numbers, dates and acronyms follow the conventions of the project's language and are usually right; text that mixes conventions is where mistakes appear. Audition the lines with figures in them first, and spell out anything whose reading matters.

When this is the right tool, and when it is not

If you want a web page or a PDF read aloud while you do something else, a reader extension is a better fit — it follows the text and does not ask you to paste. If you need a single voice reading an hour of prose with no direction, almost any online text to speech will do it, and some are cheaper per hour.

CastDub earns its place the moment the text has more than one voice in it, or the moment a line needs to be said a particular way. A dialogue, a presenter with a customer, a story with a narrator and characters, an explainer where the call to action should land differently from the paragraph before it: these are the cases where one flat voice per file stops being enough, and where casting and per-line direction — which is what this text to speech is built around — start to matter.

Perguntas frequentes

Do I need to install anything?

No. It runs in the browser; there is no desktop app and no extension. Sign in with Google or an email address, paste, and listen. The demo on the homepage does not even need an account.

Is it actually free?

The Free plan is free, with limits stated on the pricing page: a one-time balance worth about five minutes of finished English audio, 30 generated lines a day, all forty voices, and personal non-commercial use. No credit card is taken at sign-up.

Which languages does it read?

English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian — every voice in every one of them. Outside English each voice uses the general standard of the language; regional accents cannot be chosen.

Can I download the audio?

Yes, as MP3: a single mixed file on every plan, or one file per line with a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV export.

Can I use the audio commercially?

On Basic, Plus and Pro, yes. The Free plan is for personal, non-commercial use — a licensing term, not a technical one; the audio is identical.

Can I upload a document?

No, only pasted text. Copy the text out of the file and paste it in; standard screenplay layout and "Name: line" pairs are recognised as speakers, everything else is read by the narrator.

Escale a sua primeira cena hoje à noite

Cole um roteiro, deixe a CastDub dar uma voz a cada personagem e exporte um audiodrama pronto.

Criar de graça

Sem cartão de crédito. O Free dá um bolo único de créditos.