Этой страницы пока нет на русском — показана английская версия.

Синтез

Free Online Text to Speech, Nothing to Install

Paste text, choose a voice, press play. CastDub's text to speech runs in the browser: forty AI voices in nine languages, an MP3 to download, and a free plan that does not ask for a card. It is built for scripts with several speakers, but it works just as well for one.

Коротко о главном

Что это
Бесплатный генератор голоса на нейросети и озвучка текста, сделанные под сценарии: он определяет, кто говорит, даёт каждому персонажу свой голос, позволяет вести эмоцию от реплики к реплике и выгружает готовый микс — аудиоспектакль с полным составом, а не один диктор, читающий все роли.
Что вы вставляете
Сценарий, вставленный текстом. Сценарный формат, «Имя: реплика» и обычная проза с указаниями на говорящего читаются как написано, и диалоги через тире в испанском, французском, португальском и русском тоже распознаются. Строка целиком в скобках считается ремаркой и не произносится. Импорта файлов нет: вставьте текст, а не прикладывайте документ.
Что получаете
Один сведённый MP3 на любом тарифе. С Plus — субтитры SRT, написанные по вашему сценарию, а не расшифрованные со звука: имена написаны верно, а тайминги совпадают сразу. Отдельные дорожки по персонажам (по одной дорожке полной длины на персонажа плюс микс) приходят архивом ZIP на Pro. Экспорта в WAV нет.
Персонажи и базовые голоса
В библиотеке 40 персонажей, у каждого портрет, характер и образец. За ними стоят 10 базовых голосов — 4 женских, 5 мужских, 1 андрогинный — и это одни и те же 10 для всех языков. Внутри одного проекта двум персонажам не достаётся один базовый голос, пока не заняты все 10: сцена до 10 говорящих расходится чисто, а дальше двое делят голос и различаются характером и режиссурой каждой реплики. Вы выбираете персонажа, а не базовый голос.
Языки
9: English, 简体中文, 日本語, 한국어, Deutsch, Français, Español, Português, Русский. Каждый персонаж играет на всех них — персонаж это характер, а не один записанный дубль, — а язык проекта задаётся при создании и дальше не меняется. За пределами английского вы получаете распространённый акцент этого языка; региональные акценты выбрать нельзя, и мы их не обещаем. Мы не давали носителям прослушать каждый язык подряд, поэтому прогоните свои собственные реплики, прежде чем решать.
Эмоции и режиссура по репликам
Каждую реплику играет ИИ-актёр по выписанному указанию для игры, поэтому 7 эмоций (нейтрально, радость, грусть, гнев, страх, шёпот, воодушевление) именно сыграны, а не подделаны настройками. Базовая эмоция персонажа есть на любом тарифе; переопределять эмоцию по каждой реплике можно только на Pro, а регулировка скорости в 5 ступеней (0.75×–1.25×) идёт с Plus и Pro. Повторная генерация реплики даёт действительно другой дубль тем же голосом, а правка одной строки пересинтезирует только эту строку.
Цена и открытая бета
Один кредит — это один символ сценария, и кредиты списываются только при скачивании: генерация и прослушивание бесплатны, но есть потолок в 30 реплик за 24 часа на Free и до 2,000 на Pro. Тариф Free — разовый запас на 4,500 кредитов, это примерно 5 минут готового звука, а персонажей на проект — 2. Платные тарифы стоят Basic $5, Plus $19, Pro $29 в месяц, при оплате за год дешевле на 10%. Оплата ещё не запущена: во время открытой беты у каждого аккаунта бесплатно открыт весь набор возможностей Pro, но коммерческая лицензия сюда не входит и появляется вместе с платным тарифом, когда тарифы откроются.
Чего пока нельзя
Клонирования голоса нет ни на одном тарифе. Нет импорта файлов, нет экспорта в WAV, нет словаря произношения, нет регулировки высоты тона и нет способа потребовать конкретный базовый голос или региональный акцент. Ни в одном языке нет настоящих детских голосов: детские роли играют взрослые актёры, как это всегда делали в анимации. Персонажи существуют только внутри одного проекта и не переходят в следующий.

Голоса, на которых стоит попробовать

Noor

Noor

Newsreader clarity for exposition-heavy scenes

женскийвзрослыйамериканский английскийновостичёткий
Взять этот голос
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

мужскойвзрослыйамериканский английскийзакадровыйдокументальный
Взять этот голос
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

женскийвзрослыйамериканский английскийзакадровыйтёплый
Взять этот голос
Quill

Quill

YouTube explainer voice that never sounds bored

мужскоймолодойамериканский английскийYouTubeчёткий
Взять этот голос
Dorian

Dorian

Audiobook baritone for literary fiction and slow reveals

мужскойвзрослыйбританский английскийаудиокниганизкий
Взять этот голос
Solène

Solène

Continental accent, unhurried, very expensive sounding

женскийвзрослыйбританский английскийэлегантныйакцент
Взять этот голос
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

нейтральныйвзрослыйамериканский английскийспокойныйуспокаивающий
Взять этот голос
Juno

Juno

Podcast host who thinks out loud and never reads a script

женскиймолодойамериканский английскийподкастразговорный
Взять этот голос

Вся библиотека →

Что он умеет

Works in the browser

Nothing to download or install. Open the studio in any modern browser, paste, listen. Projects are saved to your account, so a script you started on one machine is there on the next.

Forty voices, nine languages

Every voice speaks English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, and every voice is on every plan including Free. Pick by persona — a newsreader, a bedtime narrator, an audiobook baritone — rather than by locale code.

MP3 download

Export a mixed MP3 on any plan, or a ZIP with one MP3 per line and a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV option.

More than one speaker, when the text needs it

Write a name and a colon at the start of a line and that line goes to a second voice; everything else stays with the narrator. Ordinary online TTS gives you one voice per file. This one gives you a cast the moment the text turns into a conversation.

Direction on the line

A note in parentheses at the start of a line — (whispering), (fast) — is performed rather than read out. On Plus and Pro each line has its own pace; on Pro, its own emotion. Punctuation is timing on every plan: a full stop pauses, a line break pauses longer.

Free plan, no credit card

A one-time balance worth about five minutes of finished English audio — more in Chinese, Japanese and Korean — and generating or auditioning does not spend it; only downloading does. The demo on the homepage runs without an account at all.

Как этим пользоваться

  1. Paste the text

    One paragraph or a whole script, as plain text. There is no file upload, so copy it out of the document. Names followed by a colon mark a second speaker; leave them out for a single voice.

  2. Pick a voice

    Filter the forty characters by gender, age and style, then audition on your own first sentence rather than on sample text. A voice that suits the sample and not your text is a common way to waste ten minutes.

  3. Listen line by line

    Each line is rendered on its own and takes a few seconds; generating the whole script runs several lines at once. Play the first few before you commit to the rest.

  4. Fix it in the text

    A mispronounced name gets respelled the way it sounds; a sentence that runs too fast gets split; a number that must be read a particular way is written out in words. There is no pronunciation setting, and you will not miss it.

  5. Download

    A mixed MP3, or one file per line with a CSV of durations. Downloading is what spends the balance; listening does not.

What "free online text to speech" usually leaves out

Most tools that rank for the phrase are free in one of three ways: free for a few hundred characters, free with the voices you would actually want locked behind a plan, or free until you try to download. It is worth being precise about which kind this is. All forty voices are available on the Free plan. Listening and regenerating cost nothing. The balance — roughly five minutes of finished English audio, more in Chinese, Japanese and Korean — is spent only when you download, and it is a one-time balance rather than a monthly one. The Free plan also caps generation at 30 lines a day and is licensed for personal, non-commercial use.

Paid plans change the numbers, not the voices: a monthly balance in place of the one-time one, more lines a day, commercial use, and on Plus and Pro the per-line pace control and subtitles. Nothing about the sound of the voices changes when you pay.

Getting a natural read out of any text to speech

Text written for the eye contains things that do not survive being spoken: long parenthetical asides, three nested clauses, a paragraph that relies on the reader glancing back. Ten minutes of adaptation — splitting sentences, cutting an aside, turning a list into a line per item — does more for the result than any voice setting. This is not a workaround for a limitation; it is the same edit a human narrator would ask for before a session.

Two habits matter most. Punctuate deliberately, because punctuation is the only timing control you have on the Free plan and the main one on any plan. And keep one idea per line: a sentence that swings from good news to bad is rendered as an average of both, and splitting it is the fix.

Numbers, dates and acronyms follow the conventions of the project's language and are usually right; text that mixes conventions is where mistakes appear. Audition the lines with figures in them first, and spell out anything whose reading matters.

When this is the right tool, and when it is not

If you want a web page or a PDF read aloud while you do something else, a reader extension is a better fit — it follows the text and does not ask you to paste. If you need a single voice reading an hour of prose with no direction, almost any online text to speech will do it, and some are cheaper per hour.

CastDub earns its place the moment the text has more than one voice in it, or the moment a line needs to be said a particular way. A dialogue, a presenter with a customer, a story with a narrator and characters, an explainer where the call to action should land differently from the paragraph before it: these are the cases where one flat voice per file stops being enough, and where casting and per-line direction — which is what this text to speech is built around — start to matter.

Частые вопросы

Do I need to install anything?

No. It runs in the browser; there is no desktop app and no extension. Sign in with Google or an email address, paste, and listen. The demo on the homepage does not even need an account.

Is it actually free?

The Free plan is free, with limits stated on the pricing page: a one-time balance worth about five minutes of finished English audio, 30 generated lines a day, all forty voices, and personal non-commercial use. No credit card is taken at sign-up.

Which languages does it read?

English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian — every voice in every one of them. Outside English each voice uses the general standard of the language; regional accents cannot be chosen.

Can I download the audio?

Yes, as MP3: a single mixed file on every plan, or one file per line with a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV export.

Can I use the audio commercially?

On Basic, Plus and Pro, yes. The Free plan is for personal, non-commercial use — a licensing term, not a technical one; the audio is identical.

Can I upload a document?

No, only pasted text. Copy the text out of the file and paste it in; standard screenplay layout and "Name: line" pairs are recognised as speakers, everything else is read by the narrator.

Озвучьте первую сцену уже сегодня

Вставьте сценарий, дайте CastDub подобрать голос каждому персонажу и экспортируйте готовый аудиоспектакль.

Создать бесплатно

Карта не нужна. Free даёт разовый запас кредитов.