이 페이지는 아직 한국어로 번역되지 않아서 영어판을 보여 드려요.

음성 합성

Free Online Text to Speech, Nothing to Install

Paste text, choose a voice, press play. CastDub's text to speech runs in the browser: forty AI voices in nine languages, an MP3 to download, and a free plan that does not ask for a card. It is built for scripts with several speakers, but it works just as well for one.

한눈에 보기

무엇인가
대본을 위해 만든 무료 AI 음성 생성기이자 텍스트 음성 변환 작업대입니다. 누가 말하는지 찾아내고, 인물마다 다른 목소리를 주고, 대사별로 감정을 연출하고, 완성된 믹스를 내보냅니다. 내레이터 한 명이 전부 읽는 게 아니라 전원 더빙 오디오 드라마가 됩니다.
무엇을 넣나
붙여 넣은 대본입니다. 시나리오 형식, 「이름: 대사」, 대화 표시가 붙은 산문 모두 쓰인 그대로 읽습니다. 스페인어·프랑스어·포르투갈어·러시아어에서 줄표로 대화를 여는 방식도 알아봅니다. 대사 맨 앞의 괄호와 줄 전체가 괄호인 줄은 그 대사의 연기 지시로 연기되고 읽지 않습니다. 대사 목록에서 바로 고칠 수 있습니다. 파일 가져오기는 없으니 문서를 첨부하지 말고 본문을 붙여 넣으세요.
무엇이 나오나
모든 플랜에서 믹스된 MP3 한 개. Plus부터는 SRT 자막도 나옵니다. 오디오를 받아쓴 것이 아니라 대본에서 만들기 때문에 이름 표기가 틀리지 않고 타이밍도 처음부터 맞습니다. 인물별 스템(본편과 같은 길이의 트랙을 인물 수만큼, 믹스와 함께 ZIP으로)은 Pro의 기능입니다. WAV 내보내기는 없습니다.
캐릭터와 기본 목소리
라이브러리 캐릭터는 40명이고 저마다 일러스트, 인물 설정, 샘플이 있습니다. 그 뒤에는 기본 목소리 10개(여성 4, 남성 5, 중성 1)가 있고, 어느 언어에서나 같은 10개입니다. 한 프로젝트 안에서는 10개가 다 찰 때까지 두 인물이 같은 기본 목소리를 쓰지 않습니다. 그래서 화자가 10명까지면 목소리가 깨끗하게 갈리고, 그보다 많아지면 두 인물이 한 목소리를 나눠 쓰며 인물 설정과 대사별 연기 지시로 구분됩니다. 고르는 것은 캐릭터이지 기본 목소리가 아닙니다.
지원 언어
9개: English, 简体中文, 日本語, 한국어, Deutsch, Français, Español, Português, Русский. 라이브러리의 모든 캐릭터가 이 언어들로 연기합니다. 캐릭터는 인물 설정이지 녹음된 목소리 하나가 아니기 때문입니다. 프로젝트의 언어는 만들 때 정해지고 이후 바뀌지 않습니다. 영어 외에는 그 언어의 일반적인 억양이며, 지역 억양은 고를 수 없고 보장하지도 않습니다. 언어마다 원어민이 하나하나 들어 본 것은 아니니, 직접 쓴 대사로 한 번 들어 보고 결정하세요.
감정과 대사별 연기
대사 하나하나를 AI 성우가 글로 적힌 연기 지시를 보고 그 자리에서 연기합니다. 그래서 7가지 감정(중립, 기쁨, 슬픔, 분노, 두려움, 속삭임, 들뜸)은 설정값으로 흉내 낸 것이 아니라 실제로 연기된 것입니다. 인물별 기본 감정은 모든 플랜에서 설정할 수 있고, 대사별로 감정을 덮어쓰는 것은 Pro만, 5단계 속도 조절(0.75×~1.25×)은 Plus와 Pro입니다. 같은 대사를 다시 생성하면 목소리는 그대로인 채 다른 테이크가 나오고, 한 줄을 고치면 그 줄만 다시 합성됩니다.
가격과 오픈 베타
크레딧 하나는 대본의 한 글자이고, 내려받을 때만 차감됩니다. 생성과 미리 듣기는 크레딧을 쓰지 않지만 24시간마다 합성 줄 수 상한이 있습니다(Free은 30줄, Pro은 최대 2,000줄). Free은 딱 한 번 주어지는 4,500 크레딧으로 완성본 약 11분어치이고, 프로젝트당 캐릭터는 2명까지입니다. 유료 플랜은 월 Basic $5, Plus $19, Pro $29, 연 결제는 10% 할인입니다. 결제는 아직 열리지 않았습니다. 오픈 베타 동안에는 모든 계정이 Pro의 전체 기능을 무료로 쓰지만, 상업적 이용 권리는 별개여서 플랜이 열린 뒤에는 유료 플랜이 필요합니다.
아직 안 되는 것
보이스 클로닝은 어떤 플랜에서도 쓸 수 없습니다. 파일 가져오기, WAV 내보내기, 발음 사전, 음높이 조절이 없고, 특정 기본 목소리나 지역 억양을 지정할 수도 없습니다. 어떤 언어에도 진짜 아이 목소리는 없습니다. 아이 배역은 성인 성우가 연기합니다. 애니메이션 더빙이 늘 해 온 방식입니다. 캐릭터는 한 프로젝트 안에만 있고 다음 프로젝트로 넘어가지 않습니다.

이 보이스들로 시험해 보세요

Noor

Noor

Newsreader clarity for exposition-heavy scenes

여성성인미국 영어뉴스또렷한
이 보이스로 만들기
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

남성성인미국 영어내레이션다큐멘터리
이 보이스로 만들기
Wren

Wren

Bedtime-story narrator, slow and warm, never wakes anyone up

여성성인미국 영어내레이션따뜻한
이 보이스로 만들기
Quill

Quill

YouTube explainer voice that never sounds bored

남성청년미국 영어유튜브또렷한
이 보이스로 만들기
Dorian

Dorian

Audiobook baritone for literary fiction and slow reveals

남성성인영국 영어오디오북저음
이 보이스로 만들기
Solène

Solène

Continental accent, unhurried, very expensive sounding

여성성인영국 영어우아한억양
이 보이스로 만들기
Halcyon

Halcyon

Meditation-adjacent narrator for slow, quiet scenes

중성성인미국 영어차분한포근한
이 보이스로 만들기
Juno

Juno

Podcast host who thinks out loud and never reads a script

여성청년미국 영어팟캐스트대화체
이 보이스로 만들기

라이브러리 전체 보기 →

무엇을 해 주나요

Works in the browser

Nothing to download or install. Open the studio in any modern browser, paste, listen. Projects are saved to your account, so a script you started on one machine is there on the next.

Forty voices, nine languages

Every voice speaks English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, and every voice is on every plan including Free. Pick by persona — a newsreader, a bedtime narrator, an audiobook baritone — rather than by locale code.

MP3 download

Export a mixed MP3 on any plan, or a ZIP with one MP3 per line and a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV option.

More than one speaker, when the text needs it

Write a name and a colon at the start of a line and that line goes to a second voice; everything else stays with the narrator. Ordinary online TTS gives you one voice per file. This one gives you a cast the moment the text turns into a conversation.

Direction on the line

A note in parentheses at the start of a line — (whispering), (fast) — is performed rather than read out. On Plus and Pro each line has its own pace; on Pro, its own emotion. Punctuation is timing on every plan: a full stop pauses, a line break pauses longer.

Free plan, no credit card

A one-time balance worth about five minutes of finished English audio — more in Chinese, Japanese and Korean — and generating or auditioning does not spend it; only downloading does. The demo on the homepage runs without an account at all.

쓰는 순서

  1. Paste the text

    One paragraph or a whole script, as plain text. There is no file upload, so copy it out of the document. Names followed by a colon mark a second speaker; leave them out for a single voice.

  2. Pick a voice

    Filter the forty characters by gender, age and style, then audition on your own first sentence rather than on sample text. A voice that suits the sample and not your text is a common way to waste ten minutes.

  3. Listen line by line

    Each line is rendered on its own and takes a few seconds; generating the whole script runs several lines at once. Play the first few before you commit to the rest.

  4. Fix it in the text

    A mispronounced name gets respelled the way it sounds; a sentence that runs too fast gets split; a number that must be read a particular way is written out in words. There is no pronunciation setting, and you will not miss it.

  5. Download

    A mixed MP3, or one file per line with a CSV of durations. Downloading is what spends the balance; listening does not.

What "free online text to speech" usually leaves out

Most tools that rank for the phrase are free in one of three ways: free for a few hundred characters, free with the voices you would actually want locked behind a plan, or free until you try to download. It is worth being precise about which kind this is. All forty voices are available on the Free plan. Listening and regenerating cost nothing. The balance — roughly five minutes of finished English audio, more in Chinese, Japanese and Korean — is spent only when you download, and it is a one-time balance rather than a monthly one. The Free plan also caps generation at 30 lines a day and is licensed for personal, non-commercial use.

Paid plans change the numbers, not the voices: a monthly balance in place of the one-time one, more lines a day, commercial use, and on Plus and Pro the per-line pace control and subtitles. Nothing about the sound of the voices changes when you pay.

Getting a natural read out of any text to speech

Text written for the eye contains things that do not survive being spoken: long parenthetical asides, three nested clauses, a paragraph that relies on the reader glancing back. Ten minutes of adaptation — splitting sentences, cutting an aside, turning a list into a line per item — does more for the result than any voice setting. This is not a workaround for a limitation; it is the same edit a human narrator would ask for before a session.

Two habits matter most. Punctuate deliberately, because punctuation is the only timing control you have on the Free plan and the main one on any plan. And keep one idea per line: a sentence that swings from good news to bad is rendered as an average of both, and splitting it is the fix.

Numbers, dates and acronyms follow the conventions of the project's language and are usually right; text that mixes conventions is where mistakes appear. Audition the lines with figures in them first, and spell out anything whose reading matters.

When this is the right tool, and when it is not

If you want a web page or a PDF read aloud while you do something else, a reader extension is a better fit — it follows the text and does not ask you to paste. If you need a single voice reading an hour of prose with no direction, almost any online text to speech will do it, and some are cheaper per hour.

CastDub earns its place the moment the text has more than one voice in it, or the moment a line needs to be said a particular way. A dialogue, a presenter with a customer, a story with a narrator and characters, an explainer where the call to action should land differently from the paragraph before it: these are the cases where one flat voice per file stops being enough, and where casting and per-line direction — which is what this text to speech is built around — start to matter.

자주 묻는 질문

Do I need to install anything?

No. It runs in the browser; there is no desktop app and no extension. Sign in with Google or an email address, paste, and listen. The demo on the homepage does not even need an account.

Is it actually free?

The Free plan is free, with limits stated on the pricing page: a one-time balance worth about five minutes of finished English audio, 30 generated lines a day, all forty voices, and personal non-commercial use. No credit card is taken at sign-up.

Which languages does it read?

English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian — every voice in every one of them. Outside English each voice uses the general standard of the language; regional accents cannot be chosen.

Can I download the audio?

Yes, as MP3: a single mixed file on every plan, or one file per line with a CSV of durations. Subtitles come out as SRT on Plus and Pro. There is no WAV export.

Can I use the audio commercially?

On Basic, Plus and Pro, yes. The Free plan is for personal, non-commercial use — a licensing term, not a technical one; the audio is identical.

Can I upload a document?

No, only pasted text. Copy the text out of the file and paste it in; standard screenplay layout and "Name: line" pairs are recognised as speakers, everything else is read by the narrator.

오늘 밤, 첫 장면을 캐스팅해 보세요

장면을 붙여 넣고, CastDub가 인물마다 보이스를 붙이게 하고, 완성된 오디오 드라마로 내보내세요.

무료로 만들어 보기

카드 등록 없이. 무료 요금제는 한 번만 주어지는 크레딧이에요.