
Inés
Bilingual lead who switches register mid-sentence
Diese Seite gibt es noch nicht auf Deutsch, deshalb steht hier die englische Fassung.
Workflow
Hola, me llamo Jeannette. A beginner dialogue is six lines long, two people take turns, and the whole pedagogical point is that the learner can tell which of them is speaking. Read by one voice, it stops being a conversation; read by two, a unit that took ten minutes to write becomes listening material you can hand out the same day.
Turn-taking is most of what a beginner dialogue teaches. A learner who cannot hear the handover is parsing a paragraph rather than following a conversation, and no amount of punctuation on the page fixes that. Separate voices restore it, and the effect is strongest exactly where it matters most: at the level where every line is four words long.
A project is written in one language, so a bilingual pair means two projects with the same script translated. That is a minute of extra work and it gives you the comparison learners want: the same exchange, the same turn structure, heard twice. It is also the cleanest way to show a class that the pause and the intonation move when the language does.
Export a ZIP with one audio file per line plus a table listing the speaker and the text, on any plan. That is the raw material for gap-fill listening, repeat-after-me, dictation and flashcards — anything where a learner needs one utterance on demand rather than a track to scrub through.
Per-line speed control on Plus and Pro runs from noticeably slower than natural to slightly faster, so the same dialogue serves a first listening and a review. Slowing only the two lines carrying the new structure, and leaving the rest at normal pace, is more useful than slowing everything.

Inés
Bilingual lead who switches register mid-sentence

Juno
Podcast host who thinks out loud and never reads a script

Quill
YouTube explainer voice that never sounds bored

Noor
Newsreader clarity for exposition-heavy scenes

Hazel
Mum voice: patient, then suddenly not

Ezra
Visual-novel love interest with a two-second pause before honesty

Tess
Australian lead with a flat, dry delivery

Hale
Documentary narrator with a hand on the listener's shoulder
Speaker name, colon, line. Six to ten turns is the usual length for one teaching point, and short lines are easier for a learner to hold than long ones.
A project is created in one language and stays there. For a bilingual pair, make the second project with the translated script rather than mixing languages in one.
At beginner level the two voices should be as easy to tell apart as you can make them — different pitch bands and different speaking rates. Realism matters later; separation matters now.
Mark the question, the correction, the polite refusal. Leave the rest resting. A bracketed note before a line is performed as direction, so a note like [surprised] in your lesson draft comes through as one.
Play it once at full speed without the text. Anything you cannot catch at that level, a learner will not catch either — shorten the sentence or slow that line rather than repeating it louder.
The mixed file for playing in class, the per-line ZIP for drills and self-study, subtitles on Plus and Pro when you want a transcript learners can follow.
Every language course needs them and almost nobody has enough. A beginner dialogue is short, specific to one teaching point, and obsolete as soon as you change the vocabulary list — which makes it the worst possible candidate for a recording session and the best possible candidate for something you can regenerate in a minute.
The traditional workaround is the teacher reading both parts, doing a voice for the second speaker. Learners tolerate it and it costs the exercise more than teachers think: a beginner is still building the ability to segment continuous speech, and a single voice changing character is an extra signal to decode on top of the language itself.
The other workaround is a document-reading tool, which gives you one voice reading both parts with no attempt at character at all. For a vocabulary list that is fine. For a conversation it removes the structure the conversation was demonstrating.
Keep the turns short and roughly balanced. A dialogue where one speaker says thirty words and the other says two is a monologue with interruptions, and it teaches the learner nothing about taking a turn. Six to ten turns of four to twelve words is the shape most course books settle on for good reason.
Put the new structure in the middle, not the first line. The first exchange should be something the class already owns — a greeting, a name — so that everyone is inside the conversation before the difficulty arrives. The last line should close the exchange audibly, because learners need to know it has ended without being told.
Write the names carefully. A name the class cannot pronounce becomes the thing they remember about the lesson, and an unfamiliar proper noun is also the place where any speech engine is most likely to surprise you. Check the names first when you listen back.
Finally, write for listening rather than for reading. Subordinate clauses that are easy on the page are heavy in the ear, and a sentence that needs a comma to be understood usually needs to be two sentences instead.
With the per-line files and the table of speakers and text, a single six-line dialogue supports most of a unit's listening work. Play line four alone and ask what came before it. Play the audio with two lines removed and have learners reconstruct them. Give the learner one speaker's lines only and have them supply the other half aloud, which is the closest thing to speaking practice that self-study can produce.
For dictation, the per-line files remove the part learners hate, which is finding the sentence again. For repetition drills they remove the part teachers hate, which is pressing pause at exactly the right moment thirty times.
The same material also makes a usable assessment. Because you have the text of every line, a gap-fill built from the audio is a copy-and-paste job rather than a transcription job, and the answer key is correct by construction.
Comparing the same exchange across two languages is an old teaching move and an awkward one to resource, because you need recordings of both. Two projects with the same translated script give you exactly that, and the comparison works best when the casting rhymes: the same pair of registers, the same order of speakers, the same resting emotions.
What learners hear in the comparison is rarely the vocabulary. It is that the pause falls somewhere else, that the question rises differently, that politeness lands on a different word. Those are the things that make a learner sound foreign long after their grammar is fine, and they are almost impossible to teach from a page.
Two practical cautions. Characters exist inside a project, so the two versions are cast independently and will not use identical voices; choose deliberately in both rather than expecting a match. And translate for the same teaching point rather than word for word, because a literal translation of a natural greeting is usually neither natural nor a greeting.
It does not teach, assess or correct. It renders a script you wrote, with the words you wrote, which means the pedagogy is entirely yours and the material is exactly as good as your dialogue.
It will not produce a specific regional accent outside English, will not produce a learner's accent or deliberate errors, and has no pronunciation dictionary to force a reading. It also has no children's voices in any language, so a dialogue between school pupils is performed by adults playing young.
And a regenerated line is a fresh take rather than the same file again — the same voice, a slightly different delivery. For teaching material this is mostly convenient, because a line you did not like can simply be run again; just do not build an exercise that depends on two renders being bit-for-bit identical.
Free
US$0 / Monat
Hör, wie dein Skript besetzt klingt, bevor du irgendetwas bezahlst.
Basic
US$5 / Monat
Die ganze Stimmbibliothek öffnen und fertiges Audio mitnehmen.
Plus
US$19 / Monat
Führ das Tempo selbst: ein Geständnis langsamer, einen Streit schneller.
Pro
US$29 / Monat
Der Tarif fürs Hörspiel: jede Zeile mit eigener Emotion, jede Figur mit eigener Spur.
Nine: English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German and Russian. Every base voice speaks all of them, so there is no language with fewer options than another. Outside English what you get is that language's general accent — neutral Mandarin, Latin American Spanish, Brazilian Portuguese — and regional accents can neither be chosen nor promised. If your course teaches a specific regional variety, listen before you commit to it.
Yes, as two projects. A project's language is fixed when you create it, so the bilingual pair is the English project and the Spanish project with the same script translated. Cast them with the same pairing of registers in both — a lower voice for the first speaker, a brighter one for the second — and the comparison stays clean. Note that characters live inside a project, so the two projects are cast separately and the voices will not be identical.
Per-line speed runs from noticeably slower than natural to slightly faster, and it is available on Plus and Pro. It is genuine pacing rather than a stretched recording, so a slowed line still sounds like someone speaking carefully rather than like a tape running down. On plans without it, the text is the lever: shorter sentences, and a line break where you want the learner to have a beat.
Treat it as good and not authoritative. We have not had every language reviewed line by line by native-speaker teachers, so listen to your own material before you use it — that is the only test that matters for your specific vocabulary. There is no pronunciation dictionary and no way to override a reading; where a word comes out wrong the fix is the text itself, such as respelling a loanword or writing a Japanese term in kana. Proper nouns and numbers are the two places to check first.
A commercial licence comes with the paid plans and covers course material you sell. The free plan is for personal, non-commercial use, and the open beta opens the features rather than the licence terms. Your script and your translations stay yours throughout.
Up to a point, and the useful version of it is not an accent. You cannot ask for a foreign accent or a learner's errors, and you should not want to: beginner listening material works better when both speakers are clear. What you can do is direct the two roles differently — the teacher's lines calm and slightly slower, the learner's lines quicker and a bit uncertain — using the resting emotion for each character and a bracketed note on the lines where that changes.
Füge ein Skript ein, lass CastDub jeder Figur eine Stimme geben und exportier ein fertiges Hörspiel.
Kostenlos startenKeine Kreditkarte. Free gibt dir ein einmaliges Guthaben.