此页尚未翻译,暂以英文显示。
Comparison
CastDub as a Typecast alternative for scripted dialogue
Two tools built around characters rather than voice pickers, at nearly the same price. This page is about the parts that actually differ: what happens to a whole pasted script, what you get out at the end, and who has the feature the other does not.
Checked:
We compete with Typecast, so read this with that in mind and follow the links. Everything attributed to them was read from their own pricing page and home page on 2026-09-21.
We will start with the awkward part. Their paid ladder is $5, $19, $29 and $69 a month; ours is $5, $19 and $29. On both, credits are spent only when you download and generating is not metered the same way. That is not a coincidence in a market where everybody reads everybody's pricing page, and we would rather point at it than hope you do not notice.
What we could verify
| CastDub | Typecast | ElevenLabs | |
|---|---|---|---|
| Several characters from one pasted script? | Yes, and it is the whole product: paste a script, it splits the lines, works out who is speaking and gives each speaker one of 40 library characters. Up to 10 distinct base voices in one project. | Not stated on the pricing or home pages we read. (source) | “There is no limit to the number of speakers in a dialogue” (Text to Dialogue, Eleven v3). (source) |
| Free plan | Yes: 4,500 lifetime download credits (about 16 minutes of finished English audio), 30 synthesised lines per 24 hours, 2 cast members per project. Generating and previewing never costs credits on any plan. | “Voice generation & playback Unlimited” and “Lifetime download credits 3,000 credits (~5 mins)”, plus “HD (720p) video exports” and “1GB media storage”. (source) | $0 with “10k credits per month” and “3 Projects in Studio”. (source) |
| Attribution required on free downloads? | No. There is no attribution requirement; free use is personal and non-commercial as a licence term. | “Attribution is required for all content downloaded on the Free plan”. (source) | Not stated on their pricing page. (source) |
| Paid ladder | $5, $19 and $29 a month; credits are spent on download only. | Basic $5, Plus $19, Pro $29, Business $69 a month; “Credits are used when you download audio.” (source) | Starter $6, Creator $11, Pro $99, Scale $299, Business $990 a month. (source) |
| Direct emotion line by line? | Per line on Pro; on every plan each character carries a default emotion. 7 emotions, performed by the line rather than approximated with pitch and speed. | Pro: “Fine-tune voices with emotion controls”, “Smart Emotion's one-click AI voice adjustment”, “Advanced voice controls including intonation”; Plus: “Control voice speed”. (source) | Bracketed audio tags such as [sad] inside the text. (source) |
| Clone your own voice? | No. Voice cloning is not live here. If cloning your own voice is the point, we are the wrong tool today. | Basic: “Instant Cloning (1 slot)”; Plus: “Professional Cloning (1 slot in total)”; Pro: 2 slots; Business: 10 slots. (source) | Instant cloning from Starter; professional cloning from Creator. (source) |
| One audio file per character (stems)? | Pro only: a ZIP with the mix plus one full-length MP3 per character. | Not stated on their site. (source) | Not stated; Studio downloads are “MP3 or WAV”. (source) |
| Subtitles? | Plus and Pro: an SRT file, and exporting it costs no credits. | Not stated; the pricing page lists video exports up to HD (720p) on Free and higher on paid tiers. (source) | Automatic captions in Studio; an SRT file is not stated. (source) |
| Voices and languages | 40 library characters sitting on 10 base voices (4 female, 5 male, 1 neutral). The characters are personas with acting directions, not 40 separate recorded actors. 9: English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German, Russian. Every base voice speaks all of them. Regional accents cannot be chosen outside English, and we do not promise them. | “700+ AI voices” and “35+ languages” (the same page also shows a “600+ AI voices” figure). (source) | “74” languages listed against the paid tiers. (source) |
Checked 2026-09-21 against each company's own pages, linked below.
Where a row says “not stated”, we opened the listed pages and did not find the fact. Treat it as a question to ask them, not as a missing feature.
The CastDub column is interpolated from the product's own plan definitions when this page is built.
What actually differs, if the prices are the same?
The input. Typecast's pricing page does not state that one project takes a whole multi-speaker script and casts it; what it documents is a per-plan set of voice and emotion controls, voice cloning and video exports. CastDub's entire first step is the script: you paste a scene, it separates the speakers, creates a character for each and casts them from a library of 40. If your raw material is already a script with names in front of the lines, that difference is most of the time you will spend.
The output. We export a mixed MP3 on every plan, an SRT on Plus and Pro, and on Pro a ZIP with one full-length MP3 per character. Their pricing page lists “HD (720p) video exports” and higher-resolution video on paid tiers — a video-first pipeline, where ours is an audio-first one. If your deliverable is a video with a talking avatar, that is a point for them and a thing we simply do not do.
The part where they are plainly ahead: cloning. Their page lists “Instant Cloning (1 slot)” from Basic, “Professional Cloning (1 slot in total)” on Plus, two slots on Pro and ten on Business. We have no cloning at all.
Can it voice several characters in one script?
On our side, yes, and with a rule you should know before you plan a cast: there are 10 base voices (4 female, 5 male, 1 neutral), and within one project no two characters share one until the pool for that gender is used up. So four women in a scene get four different voices; eleven speakers do not get eleven. The 40 library characters are personas — age, manner, energy, and for English an accent direction — layered on those base voices.
On theirs, we could not verify it. Their home page advertises “700+ AI voices” (the same page also shows a “600+ AI voices and 35+ languages” line) and “adjustable emotion, speed, and dynamics”, but neither the home page nor the pricing page we read states that one project takes several characters from a pasted script. That is a gap in our checking, not a claim about their product — if it matters to you, ask them directly.
What we will not do is put a cross in that box. A comparison table where the competitor's unknowns quietly become failures is the kind of table nobody should trust, including ours.
Is there a free plan without attribution?
Theirs: “Voice generation & playback Unlimited”, “Lifetime download credits 3,000 credits (~5 mins)”, “Download trial voices”, “HD (720p) video exports”, “1GB media storage” — and, in their own words, “Attribution is required for all content downloaded on the Free plan”. Commercial licensing begins at Basic.
Ours: 4,500 lifetime download credits, about 16 minutes of finished English audio, 2 cast members per project, and a synthesis ceiling of 30 lines per 24 hours instead of a credit meter on previews. There is no attribution requirement on the file. Free use is still personal and non-commercial — that is a licence term, and the open beta, during which everything runs at Pro-level features and nothing is charged, does not lift it.
So the two free tiers differ mostly in shape: five minutes of credited audio with attribution against about 16 minutes without it, and unlimited playback on their side against 30 synthesised lines a day on ours.
Can I direct emotion per line?
Both, on the top plans, in different vocabularies. Their pricing page puts “Fine-tune voices with emotion controls” and “Emotion controls including Smart Emotion's one-click AI voice adjustment” plus “Advanced voice controls including intonation” on Pro, and “Control voice speed” on Plus.
Ours: a default emotion per character on every plan, then per-line override on Pro across 7 emotions, and a five-step speed control on Plus and Pro. The emotion is acted for that line rather than faked with a filter, which has one consequence worth planning around — pressing generate again on the same line gives a new take. Same voice, slightly different reading. When a line is nearly right, that is a feature; when you need an exact repeat of a previous render, it is not, and the previous render stays on disk untouched until you edit the line.
Both of us put the good emotion control on the most expensive tier, which is worth saying out loud rather than dressing up.
Which languages, and what about accents?
Their home page says “35+ languages”. We have 9 — English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German, Russian — and all 10 base voices speak every one of them, so there is no thin-language problem here where a language exists on paper with two usable voices.
Where we are careful: accents. Outside English you get the general accent of the language and cannot choose a region, and we do not promise one. We also have not had native reviewers sign off on every language, so the instruction we give everyone is the same — paste four of your own lines, listen, and decide from that rather than from a demo we chose.
If you are working in a language in their 35 and not in our 9, this comparison is over and they win it.
When you should pick Typecast instead
If you want your own voice in the piece, they have cloning at four plan levels and we have none. If your deliverable is video rather than audio, their pricing page sells video exports and ours does not exist. If you are doing single-speaker work — an explainer, an ad read, a course module — our casting step is overhead you would be paying for and never using.
And if you already work in their editor and it fits your head, that is worth real money. Switching tools costs a week of small frictions that no comparison page can price.
Pick us when the script is the artefact: several speakers, lines you will re-cut, and a mix that needs per-character stems and subtitles at the end. That is the whole of our case, and it is narrower than the price similarity might suggest.
How this page was written
We sell a competing tool. The only thing that makes a page like this worth reading is whether you can check it, so every Typecast fact here was read on 2026-09-21 from typecast.ai and the link is printed beside the claim and again at the bottom.
We have not scored anything, ranked anything, or put a cross where we simply did not know. Three rows say “not stated on their site”, and we left them that way rather than inferring an absence — a table whose unknowns all resolve in the publisher's favour is a sales sheet wearing a lab coat.
We also flagged the price similarity in the first paragraph instead of letting you discover it. Their ladder and ours are almost the same, and on both sides credits are consumed at download rather than at generation. Pretending that convergence is a coincidence would set the wrong tone for everything else on this page.
The CastDub column is generated from the plan definitions the product enforces, so it cannot quietly flatter us. If anything here is out of date, tell us and we will change it and move the Checked date.
Where each fact came from
Every row above that describes another product came from one of these pages. Open them yourself — vendors change their plans without telling anyone, including us.
- Typecast — Pricing — Free $0, Basic $5, Plus $19, Pro $29, Business $69; free plan has unlimited generation and playback, 3,000 lifetime download credits (~5 mins), HD 720p video export, 1GB storage, attribution required; credits used when you download audio; commercial licence from Basic; emotion controls on Pro, speed control on Plus; cloning capacity 1/1/2/10 by plan.
- Typecast — Home — “700+ AI voices”, “35+ languages”, “adjustable emotion, speed, and dynamics”, voice cloning; the page also shows a “600+ AI voices and 35+ languages” figure.
- ElevenLabs — Pricing — Plan prices and credits, commercial licence from Starter, cloning tiers, 74 languages.
- ElevenLabs — Text to Dialogue documentation — No limit on speakers in a dialogue; bracketed audio tags control delivery.
- ElevenLabs — Studio product guide — MP3 or WAV download by plan; automatic captions; per-speaker stems not stated.
常见问题
Typecast and CastDub cost almost the same. What am I actually choosing between?
Input and output. We take a whole pasted script, split it by speaker and cast it, then hand back stems and subtitles; their pricing page sells per-voice controls, cloning slots and video export. If your material is a script and your deliverable is audio, that is the difference. If it is a line of copy and a video, it is not.
Does CastDub require attribution on free exports?
No. No credit line is required on the audio. Free use is personal and non-commercial as a matter of the licence, and the open beta — during which every account runs at Pro-level features and nothing is charged — does not change that term.
Can I import a Typecast project?
No. There is no import of any kind: no .txt, .docx or .srt upload, no project file. You paste the script text and re-cast it, which takes a couple of minutes for a scene and longer for a series bible.
Does CastDub do video?
No. Audio only — a mixed MP3 on every plan, SRT on Plus and Pro, per-character MP3s on Pro. Their pricing page advertises video exports; if you need a video deliverable, that is a reason to pick them.
How many emotions are there, and on which plan?
7, selectable per line on Pro; every plan can set a default emotion per character. Plus and Pro also get a five-step speed control from 0.75× to 1.25×.
Which one has more voices?
Theirs, by their own numbers: their home page says “700+ AI voices” against our 40 library characters on 10 base voices. Whether that matters depends on whether you need many distinct voices in one scene — where our ceiling is 10 — or one voice that is exactly right, where a larger catalogue helps.