此页尚未翻译,暂以英文显示。

语音合成

Free Online Text to Speech, Nothing to Install

Paste text, choose a voice, press play. CastDub's text to speech runs in the browser: forty AI voices in nine languages, an MP3 to download, and a free plan that does not ask for a card. It is built for scripts with several speakers, but it works just as well for one.

一眼看全

这是什么
一个免费的 AI 配音与文字转语音工作台,专为剧本而做:自动认出谁在说话,给每个角色各自的声音,逐句调整情绪,最后导出一条成片——是全员配音的广播剧,不是一个旁白把所有人的台词念一遍。
输入什么
粘贴进来的剧本。剧本格式、「角色名:台词」、带说话标注的叙述文都照原样解析,西班牙语、法语、葡萄牙语、俄语用破折号起对白的写法也认得。台词行首的括号、以及整行都是括号的那几行,会当成这一句的表演提示演出来,不会念出来,也可以在台词表里直接改。没有文件导入,所以请把文字粘进来,而不是上传文档。
导出什么
所有套餐都能导出一条混好的 MP3。Plus 起可以导出 SRT 字幕——它是按你的剧本写的,不是从音频转写的,所以人名不会写错、时间轴本来就对得上。分轨导出(每个角色一条与成片等长的音轨,外加成片本身,打包成 ZIP)是 Pro 的功能。没有 WAV 导出。
角色与基础嗓子
库里有 40 个角色,每个都有立绘、人设和样音。它们背后是 10 把基础嗓子——女 4、男 5、中性 1——这 10 把在每种语言里都是同样的几把。同一个项目里,两个角色不会共用一把基础嗓子,直到 10 把全被占用:所以最多 10 个说话人的戏能干净地分开;再多就有两个角色共用一把嗓子,靠人设和每句的表演指示分辨。你挑的是角色,不是基础嗓子。
支持的语言
9 种:English、简体中文、日本語、한국어、Deutsch、Français、Español、Português、Русский。库里每个角色都能用这几种语言演——角色是人设,不是一段录好的音——项目的语言在新建时定下,之后不变。英语以外一律是该语言的通用口音,地区口音选不了,也不做保证。我们没有请各语言的母语者逐一审听过,所以先拿你自己的台词听一遍再定。
情绪与逐句表演
每一句都由 AI 配音演员照着一条书面的表演指示现场演,所以 7 种情绪(平静、开心、悲伤、愤怒、恐惧、耳语、兴奋)是真的演出来的,不是靠参数凑的。给角色设一个默认情绪所有套餐都能用;逐句覆盖情绪只有 Pro 有,5 档逐句语速(0.75×–1.25×)是 Plus 和 Pro。同一句重新生成是货真价实的另一遍,嗓子不变;改动一句只会重新合成那一句。
价格与公测
一个额度等于剧本里的一个字符,而且只在下载时扣——生成和试听不扣额度,但每 24 小时有合成行数上限:Free 档 30 行,Pro 档最多 2,000 行。Free 档给的是一次性的 4,500 额度,约合 16 分钟成片,每个项目 2 个角色。付费档每月 Basic $5、Plus $19、Pro $29,年付省 10%。支付还没上线:公测期间每个账号都直接拥有 Pro 的全部功能,不花钱;但商用授权不随公测放开,套餐正式开放后仍需付费档。
目前还做不到
声音克隆任何套餐都还没有。没有文件导入,没有 WAV 导出,没有发音词典,没有音高控制,也不能指定某一把基础嗓子或某个地区口音。任何语言都没有真童声:儿童角色由成人配音演员来演,动画配音一直就是这么做的。角色只存在于单个项目里,不会带到下一个项目。

拿这些音色试试

Noor

Noor

把说明密集的段落交给新闻播报级的清晰度

女声成年美式英语新闻清晰
用这个音色
Hale

Hale

一只手搭在听众肩上的纪录片旁白

男声成年美式英语旁白纪录片
用这个音色
Wren

Wren

睡前故事旁白,缓慢温暖,从不把人吵醒

女声成年美式英语旁白温暖
用这个音色
Quill

Quill

从不显得无聊的 YouTube 讲解音

男声青年美式英语视频讲解清晰
用这个音色
Dorian

Dorian

文学小说与慢热揭示用的有声书男中音

男声成年英式英语有声书低沉
用这个音色
Solène

Solène

欧陆口音,不紧不慢,听起来非常贵

女声成年英式英语优雅口音
用这个音色
Halcyon

Halcyon

近乎冥想引导的旁白,适合安静的慢戏

中性成年美式英语平静安抚
用这个音色
Juno

Juno

边想边说、从不念稿的播客主持

女声青年美式英语播客口语
用这个音色

浏览整个音色库 →

它能做什么

Works in the browser

Nothing to download or install. Open the studio in any modern browser, paste, listen. Projects are saved to your account, so a script you started on one machine is there on the next.

Forty voices, nine languages

Every voice speaks English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, and every voice is on every plan including Free. Pick by persona — a newsreader, a bedtime narrator, an audiobook baritone — rather than by locale code.

MP3 download

Export a mixed MP3 on any plan, or a ZIP with one MP3 per line and a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV option.

More than one speaker, when the text needs it

Write a name and a colon at the start of a line and that line goes to a second voice; everything else stays with the narrator. Ordinary online TTS gives you one voice per file. This one gives you a cast the moment the text turns into a conversation.

Direction on the line

A note in parentheses at the start of a line — (whispering), (fast) — is performed rather than read out. On Plus and Pro each line has its own pace; on Pro, its own emotion. Punctuation is timing on every plan: a full stop pauses, a line break pauses longer.

Free plan, no credit card

A one-time balance worth about five minutes of finished English audio — more in Chinese, Japanese and Korean — and generating or auditioning does not spend it; only downloading does. The demo on the homepage runs without an account at all.

怎么用

  1. Paste the text

    One paragraph or a whole script, as plain text. There is no file upload, so copy it out of the document. Names followed by a colon mark a second speaker; leave them out for a single voice.

  2. Pick a voice

    Filter the forty characters by gender, age and style, then audition on your own first sentence rather than on sample text. A voice that suits the sample and not your text is a common way to waste ten minutes.

  3. Listen line by line

    Each line is rendered on its own and takes a few seconds; generating the whole script runs several lines at once. Play the first few before you commit to the rest.

  4. Fix it in the text

    A mispronounced name gets respelled the way it sounds; a sentence that runs too fast gets split; a number that must be read a particular way is written out in words. There is no pronunciation setting, and you will not miss it.

  5. Download

    A mixed MP3, or one file per line with a CSV of durations. Downloading is what spends the balance; listening does not.

What "free online text to speech" usually leaves out

Most tools that rank for the phrase are free in one of three ways: free for a few hundred characters, free with the voices you would actually want locked behind a plan, or free until you try to download. It is worth being precise about which kind this is. All forty voices are available on the Free plan. Listening and regenerating cost nothing. The balance — roughly five minutes of finished English audio, more in Chinese, Japanese and Korean — is spent only when you download, and it is a one-time balance rather than a monthly one. The Free plan also caps generation at 30 lines a day and is licensed for personal, non-commercial use.

Paid plans change the numbers, not the voices: a monthly balance in place of the one-time one, more lines a day, commercial use, and on Plus and Pro the per-line pace control and subtitles. Nothing about the sound of the voices changes when you pay.

Getting a natural read out of any text to speech

Text written for the eye contains things that do not survive being spoken: long parenthetical asides, three nested clauses, a paragraph that relies on the reader glancing back. Ten minutes of adaptation — splitting sentences, cutting an aside, turning a list into a line per item — does more for the result than any voice setting. This is not a workaround for a limitation; it is the same edit a human narrator would ask for before a session.

Two habits matter most. Punctuate deliberately, because punctuation is the only timing control you have on the Free plan and the main one on any plan. And keep one idea per line: a sentence that swings from good news to bad is rendered as an average of both, and splitting it is the fix.

Numbers, dates and acronyms follow the conventions of the project's language and are usually right; text that mixes conventions is where mistakes appear. Audition the lines with figures in them first, and spell out anything whose reading matters.

When this is the right tool, and when it is not

If you want a web page or a PDF read aloud while you do something else, a reader extension is a better fit — it follows the text and does not ask you to paste. If you need a single voice reading an hour of prose with no direction, almost any online text to speech will do it, and some are cheaper per hour.

CastDub earns its place the moment the text has more than one voice in it, or the moment a line needs to be said a particular way. A dialogue, a presenter with a customer, a story with a narrator and characters, an explainer where the call to action should land differently from the paragraph before it: these are the cases where one flat voice per file stops being enough, and where casting and per-line direction — which is what this text to speech is built around — start to matter.

常见问题

Do I need to install anything?

No. It runs in the browser; there is no desktop app and no extension. Sign in with Google or an email address, paste, and listen. The demo on the homepage does not even need an account.

Is it actually free?

The Free plan is free, with limits stated on the pricing page: a one-time balance worth about five minutes of finished English audio, 30 generated lines a day, all forty voices, and personal non-commercial use. No credit card is taken at sign-up.

Which languages does it read?

English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian — every voice in every one of them. Outside English each voice uses the general standard of the language; regional accents cannot be chosen.

Can I download the audio?

Yes, as MP3: a single mixed file on every plan, or one file per line with a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV export.

Can I use the audio commercially?

On Basic, Plus and Pro, yes. The Free plan is for personal, non-commercial use — a licensing term, not a technical one; the audio is identical.

Can I upload a document?

No, only pasted text. Copy the text out of the file and paste it in; standard screenplay layout and "Name: line" pairs are recognised as speakers, everything else is read by the narrator.

今晚就给第一场戏配上音

粘贴剧本,让 CastDub 给每个角色配一个音色,直接导出一部完整的有声剧。

免费开始创作

无需绑卡。免费档给的是一次性额度。