이 페이지는 아직 한국어로 번역되지 않아서 영어판을 보여 드려요.

워크플로

Free AI Voice Over Generator

A voice over is the simplest thing text to speech does, and the one most tools get slightly wrong: one flat voice reading the whole script at one speed. This page is for the script that has a narrator and at least one other voice — a presenter and a customer, a host and a guest, an explainer with a character in it — rendered free, in the browser, as an MP3 you can lay under your video.

이 방식이 여기서 통하는 이유

One script, several voices, one file

Paste the whole script. Lines that start with a speaker's name go to a different character; everything else is read by the narrator. You get a single mixed MP3 with the pauses already in it, so the usual voice over routine — render each part separately, then line them up in an editor — disappears.

Direction, not just dictation

A line can be told how to be said. A note in parentheses at the start of a line — (warmly), (as if reading a warning label) — is performed rather than read aloud. On Pro, each line can also carry its own emotion; on Plus and Pro, its own pace. A voice over that changes gear at the right moment sounds recorded rather than generated.

Forty characters, nine languages, no paywall on voices

Every plan, including Free, can use any of the forty library characters, and each of them speaks English, Spanish, Portuguese, French, German, Russian, Chinese, Japanese and Korean. Voices are personas rather than a menu of accents: you pick the presenter you want, not a locale code.

Genuinely free to start

The Free plan comes with a one-time balance good for about five minutes of finished English audio, and generating or auditioning lines does not spend it — only exporting does. No credit card is asked for at sign-up, and the demo on the homepage runs without an account at all.

이런 작업에 어울리는 보이스

Quill

Quill

YouTube explainer voice that never sounds bored

남성청년미국 영어유튜브또렷한
이 보이스로 만들기
Indie

Indie

TikTok voiceover: fast, flat, faintly amused

여성청년미국 영어소셜빠른 말속도
이 보이스로 만들기
Juno

Juno

Podcast host who thinks out loud and never reads a script

여성청년미국 영어팟캐스트대화체
이 보이스로 만들기
Hale

Hale

Documentary narrator with a hand on the listener's shoulder

남성성인미국 영어내레이션다큐멘터리
이 보이스로 만들기
Noor

Noor

Newsreader clarity for exposition-heavy scenes

여성성인미국 영어뉴스또렷한
이 보이스로 만들기
Duke

Duke

Trailer voice. Three words per breath. All of them heavy.

남성성인미국 영어예고편저음
이 보이스로 만들기
Solène

Solène

Continental accent, unhurried, very expensive sounding

여성성인영국 영어우아한억양
이 보이스로 만들기
Beaudie

Beaudie

Australian field guide who finds everything hilarious

남성성인호주 영어억양야외
이 보이스로 만들기

라이브러리 전체 보기 →

이렇게 하면 돼요

  1. Paste the script

    Plain text, pasted. If the script has more than one voice, put the speaker's name and a colon at the start of those lines; unlabelled lines are read by the narrator. There is no file upload, so copy the text out of your document.

  2. Cast the narrator first

    Pick the character who carries most of the script, and audition on your own first sentence rather than on sample text. Quill and Hale are the safe explainer choices; Noor for anything that has to sound like news; Duke only if you actually want a trailer.

  3. Split long sentences

    Punctuation is your timing: a full stop is a pause, a line break a longer one. A sentence that carries two ideas is read as an average of both, so break it in two.

  4. Set pace before emotion

    On Plus and Pro, take the narrator a notch slower than feels right in the editor. Almost every first voice over is too fast for someone who is also watching a screen.

  5. Render straight through, then fix three lines

    Listen once at full length before editing anything. Direct only the lines where the video turns — the reveal, the call to action — and leave the rest neutral.

  6. Export and lay it under the picture

    Download the mixed MP3 on any plan, or the per-line ZIP with a CSV of durations if you need to slide individual sentences to picture. Subtitles come out as SRT on Plus and Pro at no extra cost.

What a voice over generator has to get right

Three things, in order. Pace: a voice over sits under pictures, and a listener who is also reading a screen needs the words to arrive slightly slower than in conversation. Most generated voice overs fail here first, not because the voice is wrong but because it is quick and even, and evenness reads as indifference. Clarity: product names, numbers and acronyms are the parts people actually need, and they are the parts a synthetic voice most often trips on. Register: an explainer, an advert and a compliance module want three different presenters, and a tool with one house voice cannot give you that.

CastDub approaches all three from the script rather than from a settings panel. Pace is set per line on Plus and Pro and, on every plan, by punctuation — a full stop is a real pause, an ellipsis a longer one, a line break a gap. Clarity is handled by auditioning the difficult words on their own and fixing them in the text. Register is a casting decision: forty characters with written personas, from a newsreader to a bedtime narrator to an arena announcer, and any of them can be the narrator.

The part that most voice over tools do not attempt is the second voice. Explainers keep sneaking dialogue in — a customer asking the question the video answers, a colleague objecting, a testimonial read in someone else's voice — and in a one-voice tool those lines are either read flat by the narrator or produced separately and glued on. Here they are lines with a name in front of them, cast to a second character, mixed into the same file with the right gap before and after.

Explainers, ads, product videos, e-learning: how the script changes

An explainer wants a presenter who sounds interested in the subject without performing interest. Quill was written for exactly this, and Hale is the calmer alternative for anything longer than three minutes. Keep sentences short, put the product name on its own line the first time it appears so you can audition it, and write the transitions — "here is the catch", "so what does that mean for you" — as separate lines, because those are the moments the voice should change gear.

Adverts and trailers are the one place where a pushed read is correct. Duke exists for the thirty-second spot; Axel for the sports-adjacent version of it. The trick is contrast: two lines pushed, the rest plain. A whole advert at trailer intensity is exhausting by the second sentence, and the call to action lands harder if the line before it is quiet.

E-learning and compliance modules are the largest voice over category by volume and the most sensitive to pace. Learners replay, so a slightly slow read is a feature. Noor's newsreader clarity suits definitions and procedures; Wren or Halcyon suit anything meant to lower the temperature. If the module includes a scenario — a manager and an employee, a customer and an agent — write it as a dialogue with names, and the scenario will be performed by two people instead of narrated by one.

Voice over, narration and dubbing are different jobs

Narration is a voice over that carries a story rather than a message; the same tools apply, but the casting leans towards the audiobook and bedtime voices, and the script leans towards longer sentences. Dubbing is something else again: it means replacing the speech already in a video with speech in another language, timed to the picture and ideally to the mouths. CastDub does not take a video as input and does not do that. What it does is generate the replacement track from a script you supply — which is most of the work of a simple dub, provided you translate the script yourself and are prepared to slide lines to picture afterwards.

If your actual need is to take an existing video and get it speaking another language automatically, a dedicated dubbing tool will serve you better. If your need is a clean, directed voice track from a script — in any of nine languages, with more than one voice when the script calls for it — that is what this page is for.

Timing a voice over to picture

There is no video import and no timeline here; the timing work happens in whatever editor you cut the video in. Two exports make that work easy. The mixed MP3 is the fast path: lay it on the timeline, and if the whole track runs a little long, trim the pauses between paragraphs rather than speeding up the voice. The per-line ZIP is the precise path: one MP3 per sentence and a CSV that lists each line's speaker, text and duration in milliseconds, so you can place every sentence exactly where the picture needs it.

Two habits save time. Write the script in the order the pictures will appear, with one visual beat per line, so the line list already matches your shot list. And decide the length before you record — a sixty-second video with a hundred and eighty words of script is going to be fast whatever voice reads it, and no amount of direction fixes a script that is simply too long for its slot.

Expect generated speech to run slightly faster and more evenly than a human session, and to leave no room for on-screen action unless you write it in. A line consisting of a few dots is the simplest way to reserve a beat for the picture; a line break between sentences is the second simplest.

비용은 이 정도

Free

US$0 / 월

 

내 대본에 목소리가 붙으면 어떤 소리가 나는지, 돈 한 푼 안 들이고 들어 보세요.

  • 생성과 미리 듣기는 크레딧이 안 빠져요 (24시간에 30줄까지)
  • 다운로드 크레딧 4,500개, 평생 한 번 (완성 오디오 11분쯤)
  • 캐릭터 보이스 40개 전부 — 유료 장벽은 라이브러리가 아니에요
  • 프로젝트당 등장인물 2명
  • 개인·비상업적 용도만

Basic

US$5 / 월

 

라이브러리를 전부 열고, 완성본을 가져가기 시작하는 단계.

  • 캐릭터 보이스 40개 전부
  • 매달 다운로드 크레딧 54,000개 (130분쯤)
  • 프로젝트당 등장인물 5명
  • 상업적 이용 라이선스 포함

Plus

US$19 / 월

 

속도를 직접 잡는 단계. 독백은 늦추고, 말다툼은 몰아치게.

  • 매달 다운로드 크레딧 180,000개 (450분쯤)
  • 대사별 말속도 조절
  • 프로젝트당 등장인물 무제한
  • SRT 자막 내보내기

전체 비교 →

자주 묻는 질문

Is the voice over generator really free?

Yes, within limits that are written down rather than discovered later. The Free plan has a one-time balance worth roughly five minutes of finished English audio (more in Chinese, Japanese and Korean, which cost fewer characters per minute), a daily cap of 30 generated lines, and it is for personal, non-commercial use. Paid plans reset monthly and allow commercial use. No credit card is asked for at sign-up.

Can I use the voice over in a YouTube video I monetise?

On a paid plan, yes — commercial use is part of Basic, Plus and Pro. On Free the licence is personal and non-commercial, so a monetised channel needs a paid plan. That is a term of service rather than a technical limit; the audio itself is the same.

What formats do I get?

A mixed MP3 on every plan, and a ZIP with one MP3 per line plus a CSV of durations on every plan. Separate stems per character on Pro, and an SRT subtitle file on Plus and Pro. There is no WAV export.

Can I choose a specific accent?

You choose a character, not an accent. In English the library includes British, American and Australian personas; outside English each voice speaks the general standard of that language — Latin American Spanish, Brazilian Portuguese, Mandarin — and regional accents are not selectable.

How do I fix a mispronounced brand name?

In the text. There is no pronunciation dictionary, so respell the word the way it sounds, hyphenate a compound, or write a number out in words if it matters how it is said. Audition the name on its own before rendering the whole script.

Will the same line sound identical if I regenerate it?

No — a regenerated line is a new take: same character, same voice, slightly different delivery. A rendered line keeps its audio until you edit that line, so a voice over you have approved does not drift. Regenerate when you want another read, not when you want the same one again.

오늘 밤, 첫 장면을 캐스팅해 보세요

장면을 붙여 넣고, CastDub가 인물마다 보이스를 붙이게 하고, 완성된 오디오 드라마로 내보내세요.

무료로 만들어 보기

카드 등록 없이. 무료 요금제는 한 번만 주어지는 크레딧이에요.