Production
Voice
Jalaran Voice is a local text-to-speech studio that turns a written script into narration on your own hardware, and hands the audio to the rest of the pipeline with a phoneme timeline riding inside the file.
Who Voice is for
For anyone narrating something they wrote — a video, a course, a walkthrough — who would rather not upload the script to a speech vendor and would rather not pay per character. It is also for people building characters, because Voice is where a script becomes a performance.
What Voice does
Voice runs Kokoro in a local container, so a script never leaves the machine and there is no per-character billing. Nine voices are available across American and British English, and the British ones are genuinely British — they run through British phonemisation rather than an American model wearing an accent label, which is the difference between "schedule" pronounced two different ways. A script can cast several voices in one clip with inline tags, and blank lines between paragraphs become real pauses in the audio rather than being collapsed away.
- Nine voices, with real British phonemisation for the British ones
- Multi-voice casting inside a single script
- Authored pauses and paragraph breathing
- WAV export carrying a phoneme timeline for exact lip-sync
How Voice works
Write or paste the script
Plain text. Blank lines between paragraphs become audible pauses, and you can author an explicit pause anywhere you want one.
Cast the voices
Tag spans with a voice to switch speaker mid-script. American and British voices coexist in a single clip.
Synthesise locally
Generation runs in a container on your own hardware at several times realtime on CPU. Nothing is sent to a speech API.
Export the WAV
The file carries a phoneme timeline in a custom chunk, so Puppet can lip-sync it exactly. Every other audio tool ignores the chunk safely.
What Voice does not do
Voice is English-only — the nine voices are American and British, and there is no Indonesian voice. It will not clone your voice or anyone else’s; the voices are the ones Kokoro ships, and that is a deliberate boundary rather than a missing feature. It refuses to truncate: a script longer than the limit is an error rather than audio that simply stops mid-sentence, because a clip that quietly ends early is worse than one that never generated. There is no background music, mixing or mastering.
Common questions
Does my script get sent to a speech API?
No. Synthesis runs in a container on your own hardware. There is no cloud text-to-speech vendor in the path and no per-character cost.
Can I clone my own voice?
No. Voice offers nine fixed Kokoro voices and does not do voice cloning. That is a deliberate limit, not a roadmap item.
What is the phoneme timeline for?
It records exactly which sound occurs when, and it rides inside the WAV itself. Puppet reads it to lip-sync precisely instead of estimating mouth shapes from the waveform, which was measured wrong on a large share of frames.
Explore the workspace
Jalaran is one workspace of 85 modules. Browse the rest of the arsenal: