このページはまだ日本語になっていません。英語版を表示しています。

音声合成

Free Online Text to Speech, Nothing to Install

Paste text, choose a voice, press play. CastDub's text to speech runs in the browser: forty AI voices in nine languages, an MP3 to download, and a free plan that does not ask for a card. It is built for scripts with several speakers, but it works just as well for one.

ひと目で分かる

これは何か
台本のために作られた、無料のAI音声生成・テキスト読み上げの作業台です。誰が話しているかを読み取り、キャラクターごとに別の声を割り当て、セリフ単位で感情を演出して、仕上がった音声を書き出します。一人のナレーターが全員分を読むのではなく、フルキャストのオーディオドラマになります。
入力するもの
貼り付けた台本です。シナリオ形式、「名前:セリフ」、地の文に会話タグが付いた文章はそのまま解析されます。スペイン語・フランス語・ポルトガル語・ロシア語のダッシュで始まる会話も読み取れます。セリフの行頭の括弧と、行全体が括弧のものは、その行の演出指示として演じられ、読み上げません。セリフ一覧でそのまま直せます。ファイル取り込みはないので、書類を添付するのではなく本文を貼り付けてください。
書き出せるもの
どのプランでもミックス済みの MP3 が 1 本。Plus 以上では SRT 字幕も書き出せます。音声から文字起こしするのではなく台本から作るので、固有名詞の表記が崩れず、タイミングも最初から合っています。キャラクターごとのステム(本編と同じ長さのトラックを人数分、ミックスと合わせて ZIP)は Pro の機能です。WAV の書き出しはありません。
キャラクターと基本ボイス
ライブラリのキャラクターは 40 人。それぞれに立ち絵、人物像、サンプルがあります。その裏にあるのは 10 種類の基本ボイス(女性 4、男性 5、中性 1)で、どの言語でも同じ 10 種類です。1 つのプロジェクト内では、10 種類が埋まるまで同じ基本ボイスが二人に割り当てられることはありません。話者が 10 人までなら声はきれいに分かれ、それを超えると二人が同じ声を共有し、人物像とセリフごとの演技指示で描き分けます。選ぶのはキャラクターであって、基本ボイスではありません。
対応言語
9 言語:English、简体中文、日本語、한국어、Deutsch、Français、Español、Português、Русский。ライブラリのキャラクターは全言語で演じます。キャラクターは人物像であって録音された一本の声ではないからです。プロジェクトの言語は作成時に決まり、あとから変わりません。英語以外はその言語の一般的なアクセントで、地域アクセントは選べませんし、保証もしていません。各言語のネイティブに一つずつ聴いてもらったわけではないので、まずご自分のセリフで試してから判断してください。
感情とセリフ単位の演技
セリフは一つずつ、文章で書かれた演技指示をもとに AI の声優がその場で演じます。7 種類の感情(ノーマル、喜び、悲しみ、怒り、恐れ、ささやき、興奮)はパラメータで似せたものではなく、実際に演じられたものです。キャラクターごとの既定の感情はどのプランでも設定でき、セリフ単位で感情を上書きできるのは Pro のみ、5 段階の速度調整(0.75×〜1.25×)は Plus と Pro です。同じセリフを生成し直すと、声は同じまま別のテイクになります。1 行を直したときに再合成されるのはその行だけです。
料金とオープンベータ
クレジットは台本の 1 文字ぶんで、消費するのはダウンロードのときだけです。生成と試聴では減りませんが、24 時間あたりの合成行数に上限があります(Free は 30 行、Pro は最大 2,000 行)。Free は一度きりの 4,500 クレジット、仕上がり約 13 分ぶんで、1 プロジェクトにキャラクターは 2 人までです。有料プランは月額 Basic $5、Plus $19、Pro $29、年払いで 10% 引きです。決済はまだ始まっていません。オープンベータの間はどのアカウントでも Pro の全機能を無料で使えますが、商用ライセンスは別で、プラン開始後は有料プランが必要です。
まだできないこと
ボイスクローンはどのプランでも使えません。ファイル取り込み、WAV 書き出し、発音辞書、ピッチ調整はなく、基本ボイスや地域アクセントを指名することもできません。どの言語にも本物の子どもの声はありません。子どもの役は大人の声優が演じます。アニメの吹き替えが昔からそうしてきたやり方です。キャラクターは一つのプロジェクトの中だけに存在し、次のプロジェクトには引き継がれません。

試してみる声

Noor

Noor

Newsreader clarity for exposition-heavy scenes

女性大人アメリカ英語ニュースクリア
この声で作る
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

男性大人アメリカ英語ナレーションドキュメンタリー
この声で作る
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

女性大人アメリカ英語ナレーション温かい
この声で作る
Quill

Quill

YouTube explainer voice that never sounds bored

男性青年アメリカ英語YouTubeクリア
この声で作る
Dorian

Dorian

Audiobook baritone for literary fiction and slow reveals

男性大人イギリス英語オーディオブック低音
この声で作る
Solène

Solène

Continental accent, unhurried, very expensive sounding

女性大人イギリス英語上品訛り
この声で作る
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

中性大人アメリカ英語落ち着き癒やし
この声で作る
Juno

Juno

Podcast host who thinks out loud and never reads a script

女性青年アメリカ英語ポッドキャスト会話調
この声で作る

ライブラリ全体を見る →

できること

Works in the browser

Nothing to download or install. Open the studio in any modern browser, paste, listen. Projects are saved to your account, so a script you started on one machine is there on the next.

Forty voices, nine languages

Every voice speaks English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, and every voice is on every plan including Free. Pick by persona — a newsreader, a bedtime narrator, an audiobook baritone — rather than by locale code.

MP3 download

Export a mixed MP3 on any plan, or a ZIP with one MP3 per line and a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV option.

More than one speaker, when the text needs it

Write a name and a colon at the start of a line and that line goes to a second voice; everything else stays with the narrator. Ordinary online TTS gives you one voice per file. This one gives you a cast the moment the text turns into a conversation.

Direction on the line

A note in parentheses at the start of a line — (whispering), (fast) — is performed rather than read out. On Plus and Pro each line has its own pace; on Pro, its own emotion. Punctuation is timing on every plan: a full stop pauses, a line break pauses longer.

Free plan, no credit card

A one-time balance worth about five minutes of finished English audio — more in Chinese, Japanese and Korean — and generating or auditioning does not spend it; only downloading does. The demo on the homepage runs without an account at all.

使い方

  1. Paste the text

    One paragraph or a whole script, as plain text. There is no file upload, so copy it out of the document. Names followed by a colon mark a second speaker; leave them out for a single voice.

  2. Pick a voice

    Filter the forty characters by gender, age and style, then audition on your own first sentence rather than on sample text. A voice that suits the sample and not your text is a common way to waste ten minutes.

  3. Listen line by line

    Each line is rendered on its own and takes a few seconds; generating the whole script runs several lines at once. Play the first few before you commit to the rest.

  4. Fix it in the text

    A mispronounced name gets respelled the way it sounds; a sentence that runs too fast gets split; a number that must be read a particular way is written out in words. There is no pronunciation setting, and you will not miss it.

  5. Download

    A mixed MP3, or one file per line with a CSV of durations. Downloading is what spends the balance; listening does not.

What "free online text to speech" usually leaves out

Most tools that rank for the phrase are free in one of three ways: free for a few hundred characters, free with the voices you would actually want locked behind a plan, or free until you try to download. It is worth being precise about which kind this is. All forty voices are available on the Free plan. Listening and regenerating cost nothing. The balance — roughly five minutes of finished English audio, more in Chinese, Japanese and Korean — is spent only when you download, and it is a one-time balance rather than a monthly one. The Free plan also caps generation at 30 lines a day and is licensed for personal, non-commercial use.

Paid plans change the numbers, not the voices: a monthly balance in place of the one-time one, more lines a day, commercial use, and on Plus and Pro the per-line pace control and subtitles. Nothing about the sound of the voices changes when you pay.

Getting a natural read out of any text to speech

Text written for the eye contains things that do not survive being spoken: long parenthetical asides, three nested clauses, a paragraph that relies on the reader glancing back. Ten minutes of adaptation — splitting sentences, cutting an aside, turning a list into a line per item — does more for the result than any voice setting. This is not a workaround for a limitation; it is the same edit a human narrator would ask for before a session.

Two habits matter most. Punctuate deliberately, because punctuation is the only timing control you have on the Free plan and the main one on any plan. And keep one idea per line: a sentence that swings from good news to bad is rendered as an average of both, and splitting it is the fix.

Numbers, dates and acronyms follow the conventions of the project's language and are usually right; text that mixes conventions is where mistakes appear. Audition the lines with figures in them first, and spell out anything whose reading matters.

When this is the right tool, and when it is not

If you want a web page or a PDF read aloud while you do something else, a reader extension is a better fit — it follows the text and does not ask you to paste. If you need a single voice reading an hour of prose with no direction, almost any online text to speech will do it, and some are cheaper per hour.

CastDub earns its place the moment the text has more than one voice in it, or the moment a line needs to be said a particular way. A dialogue, a presenter with a customer, a story with a narrator and characters, an explainer where the call to action should land differently from the paragraph before it: these are the cases where one flat voice per file stops being enough, and where casting and per-line direction — which is what this text to speech is built around — start to matter.

よくあるご質問

Do I need to install anything?

No. It runs in the browser; there is no desktop app and no extension. Sign in with Google or an email address, paste, and listen. The demo on the homepage does not even need an account.

Is it actually free?

The Free plan is free, with limits stated on the pricing page: a one-time balance worth about five minutes of finished English audio, 30 generated lines a day, all forty voices, and personal non-commercial use. No credit card is taken at sign-up.

Which languages does it read?

English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian — every voice in every one of them. Outside English each voice uses the general standard of the language; regional accents cannot be chosen.

Can I download the audio?

Yes, as MP3: a single mixed file on every plan, or one file per line with a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV export.

Can I use the audio commercially?

On Basic, Plus and Pro, yes. The Free plan is for personal, non-commercial use — a licensing term, not a technical one; the audio is identical.

Can I upload a document?

No, only pasted text. Copy the text out of the file and paste it in; standard screenplay layout and "Name: line" pairs are recognised as speakers, everything else is read by the narrator.

今夜、最初の1シーンをキャスティング

台本を貼り付ければ、CastDubが登場人物一人ひとりに声を割り当てます。あとは完成したボイスドラマを書き出すだけです。

無料で作ってみる

クレジットカードは不要です。無料枠は一度きりの持ち分です。