此页尚未翻译,暂以英文显示。

Comparison

CastDub as a PlayHT alternative for scripted dialogue

PlayHT's own documentation describes a model built for conversation, reached through an API. CastDub is a place to sit and cut a scene. This page is about which of those you need — and it is mostly a question about whether you write code.

Checked:

We compete with PlayHT, so follow the links. An unusual disclosure first: on 2026-09-21 play.ht refused our fetcher outright — the connection closed on their pricing page and on their home page, three attempts. So there are no PlayHT prices, free-tier terms or commercial-use terms on this page at all. Everything below comes from docs.play.ht, which did respond.

That limitation is worth stating loudly, because the most common question about any tool is what it costs, and on this page we cannot answer it for them.

What we could verify

 CastDubPlayHTElevenLabs
Several characters from one pasted script?Yes, and it is the whole product: paste a script, it splits the lines, works out who is speaking and gives each speaker one of 40 library characters. Up to 10 distinct base voices in one project.PlayDialog is “a more advanced model that can generate turn-based dialogues with multiple voices”, through the API. (source)“There is no limit to the number of speakers in a dialogue” (Text to Dialogue, Eleven v3). (source)
Editor, or API?A web editor. Paste, cast, direct, export. There is no public API for scripts.Documented as an API with voice engines, rate limits and latency guidance. (source)Both an API and the Studio editor. (source)
Emotion controlPer line on Pro; on every plan each character carries a default emotion. 7 emotions, performed by the line rather than approximated with pitch and speed.PlayDialog uses an “Adaptive Speech Contextualizer” to “control prosody, intonation, emotion and pacing”; per-line emotion labels are not stated. (source)Bracketed audio tags such as [sad] inside the text. (source)
Languages9: English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German, Russian. Every base voice speaks all of them. Regional accents cannot be chosen outside English, and we do not promise them.Play 3.0 Mini “Supports 36 languages”; PlayHT 2.0 Turbo is “English only”; PlayDialog multilingual is marked “Beta”; PlayDialog Turbo on Groq: “Only two language options are supported: english and arabic”. (source)“74” languages listed against the paid tiers. (source)
Clone your own voice?No. Voice cloning is not live here. If cloning your own voice is the point, we are the wrong tool today.“You can clone any voice instantly across languages with only 30 seconds of speech”. (source)Instant cloning from Starter; professional cloning from Creator. (source)
Price and free tier$5/month for Basic once payment is switched on. Nothing is being charged during the open beta, while every account runs at Pro-level features. Free tier: Yes: 4,500 lifetime download credits (about 16 minutes of finished English audio), 30 synthesised lines per 24 hours, 2 cast members per project. Generating and previewing never costs credits on any plan.Not verified. play.ht closed the connection to our fetcher on 2026-09-21, so we have no prices or free-tier terms from their own site.Free $0 with 10k credits a month; Starter $6; Creator $11; Pro $99; Scale $299; Business $990. (source)
One audio file per character (stems)?Pro only: a ZIP with the mix plus one full-length MP3 per character.Not applicable in the same sense: an API returns the audio you request per call, and assembly is yours. (source)Not stated; Studio downloads are “MP3 or WAV”. (source)
Subtitles?Plus and Pro: an SRT file, and exporting it costs no credits.Not stated on the documentation pages we read. (source)Automatic captions in Studio; an SRT file is not stated. (source)

Checked 2026-09-21. play.ht itself would not serve our fetcher, so every PlayHT row comes from docs.play.ht, which did.

Because of that, this page contains no PlayHT prices, free-tier limits or commercial terms. Read them on their own site before deciding anything.

The CastDub column is interpolated from the product's plan definitions at build time.

Can it voice several characters in one script?

Both, but at different altitudes. PlayHT's model documentation describes PlayDialog as “PlayHT's latest voice model built for fluid, emotive conversation” and “a more advanced model that can generate turn-based dialogues with multiple voices”. That is a genuine multi-speaker capability and it is the closest thing on this page to what we do — expressed as a model you call rather than a room you work in.

CastDub is the room. You paste a script, it splits the lines, creates a character per speaker and casts each from 40 library personas, and then you sit with the scene: re-generate one line, change one character's default emotion, reorder, export. Nothing in that loop requires an API key or a script of your own.

So the real question is not which produces dialogue — both do — but whether you want a building block or a workbench. If you are wiring voice into your own product, a model with an API is the correct choice and an editor is in your way.

How is emotion controlled?

On their side, largely automatically and with sophistication. The same documentation credits PlayDialog with “state-of-the-art” emotional capability and an “Adaptive Speech Contextualizer” used to “control prosody, intonation, emotion and pacing” — that is, the model reads the context of the conversation and decides the delivery. Per-line emotion labels are not stated on the pages we read.

On ours, explicitly: 7 emotions, a default per character on every plan and a per-line override on Pro, plus a five-step speed control on Plus and Pro. Each line is performed to its direction rather than filtered into shape.

There is a trade in here that is easy to miss. Contextual delivery is usually better across a whole conversation, because the model hears the turn before. Explicit per-line direction is better on the one line that has to break the pattern — the joke that lands flat on purpose, the confession that goes quiet. If most of your work is the first kind, their approach may simply produce better scenes with less effort.

Which languages?

Their documentation is specific per model, and the numbers vary a lot: Play 3.0 Mini “Supports 36 languages”; PlayHT 2.0 Turbo is listed as “English only”; PlayDialog's multilingual support is marked “Beta” in their comparison table; and for the Groq-hosted PlayDialog Turbo, “Only two language options are supported: english and arabic”.

That is worth reading twice if you are producing in anything but English, because the multi-speaker model and the broad language list are not the same model. We have no equivalent trap: all 10 base voices speak all 9 — English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German, Russian — and every feature works the same in each of them.

What we do not offer is regional accent choice outside English, or any promise of native-reviewed quality. Paste your own lines and listen; that is the only test that settles it.

Can I clone my own voice?

Theirs, yes: their documentation states you “can clone any voice instantly across languages with only 30 seconds of speech”. Thirty seconds is a low bar and cloning across languages is a real capability.

Ours, no. Voice cloning is not live at CastDub. When it arrives it will carry rules we have already published — recorded consent from the person, no public figures, never a minor — but the answer today is simply no.

As on every page in this family, that is a clean tiebreak: if your production depends on a specific real voice, use the tool that clones.

What do you get out at the end?

From us: a mixed MP3 on every plan, an SRT on Plus and Pro at no credit cost, and on Pro a ZIP with the mix plus one full-length MP3 per character for the mix session. No WAV — a real limitation if you are delivering to a spec.

From an API, what you get out is what you asked for: audio bytes per request, and the assembly — the gaps between lines, the timing, the subtitle file — is yours to build. Their documentation notes a limit worth planning around on the Turbo engine and rate limits that “depend on the API you are using and the plan”. None of that is a criticism; it is what an API is. But it means the comparison is really between our afternoon of clicking and your afternoon of plumbing.

Our credit model may matter here too: credits are spent only when you download, not when you generate or preview, with a per-day synthesis ceiling instead (30 lines on Free, 2,000 on Pro). Experimentation is not metered.

When you should pick PlayHT instead

When you are building software. If the voice is a feature of your own product — an agent, a game, a reading app — you need an API and a model, and a web editor is the wrong shape entirely. Their dialogue model is documented for exactly this.

When you need cloning, per their thirty-second instant clone. When you need one of the 36 languages on Play 3.0 Mini that we do not have, keeping in mind which model gets which languages. And when latency matters — their documentation has a page on reducing it, which tells you what kind of customer they design for.

Pick us when there is a person, not a program, sitting with a script: several speakers, lines to re-cut, and a need for stems and subtitles at the end without writing any code.

How this page was written

We are a competitor and this page is written by us, so the only thing that makes it worth your time is that it is checkable. Every PlayHT statement above is quoted from docs.play.ht, read on 2026-09-21, with the link beside it.

Three whole subjects are missing here — price, free-tier limits, commercial terms — because play.ht closed the connection on us. We could have taken those numbers from any of a dozen review sites. We did not, because a page that mixes verified quotes with second-hand numbers is less trustworthy than one that admits a hole, and because review-site prices for this product disagree with each other.

We have made no claim about whose voices sound better, and we have not put a cross anywhere we merely failed to find something. Where their capability is genuinely stronger — a documented dialogue model, thirty-second cloning, 36 languages on one of their engines — it is stated in their words, not paraphrased down.

Our column is generated from the plan code the product runs on, so it cannot drift into flattery.

Where each fact came from

Every row above that describes another product came from one of these pages. Open them yourself — vendors change their plans without telling anyone, including us.

  • PlayHT — Voice engines documentationPlayDialog “can generate turn-based dialogues with multiple voices” and uses an “Adaptive Speech Contextualizer” for “prosody, intonation, emotion and pacing”; Play 3.0 Mini “Supports 36 languages”; PlayHT 2.0 Turbo “English only”; PlayDialog multilingual marked Beta.
  • PlayHT — Dialog Turbo on Groq“Only two language options are supported: english and arabic.”
  • PlayHT — Documentation home“You can clone any voice instantly across languages with only 30 seconds of speech”; PlayDialog listed as a voice engine.
  • PlayHT — Rate limits“our APIs are rate-limited. The specific limits depend on the API you are using and the plan”.
  • PlayHT — PricingConnection closed to our fetcher on 2026-09-21; no prices or plan terms were taken from it.
  • ElevenLabs — PricingPlan prices and credits; cloning tiers; 74 languages.
  • ElevenLabs — Text to Dialogue documentationNo limit on speakers in a dialogue; bracketed audio tags control delivery.

常见问题

Why doesn't this page list PlayHT's prices?

Because their site would not serve our fetcher on 2026-09-21 — three attempts, connection closed each time. Every fact here comes from their documentation, which responded. Check their pricing page directly; we will add the figures once we can read them from their own page.

Is PlayDialog the same thing as CastDub's casting?

No, and the difference is the layer. PlayDialog is a model documented as generating “turn-based dialogues with multiple voices” through an API; CastDub is an editor that takes a whole pasted script, makes characters, casts them from 40 library personas and lets you re-cut individual lines without writing code.

Does CastDub have an API?

No public API for producing scripts. If you need to generate dialogue from inside your own software, an API-first provider is the right shape and we are not a candidate.

How many voices can one scene use?

10 distinct base voices (4 female, 5 male, 1 neutral). Beyond that, characters share a voice and are separated only by acting direction, so plan large ensemble scenes accordingly.

Which languages does CastDub support?

9: English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian, with every base voice speaking all of them. Regional accents outside English cannot be chosen and we do not promise them.

Can I get separate tracks for mixing?

On Pro: a ZIP with the mix and one full-length MP3 per character, plus an SRT. On Plus and Pro you get subtitles and a single mixed MP3. Everything is MP3 — there is no WAV export.

今晚就给第一场戏配上音

粘贴剧本,让 CastDub 给每个角色配一个音色,直接导出一部完整的有声剧。

免费开始创作

无需绑卡。免费档给的是一次性额度。