Production

Ensemble

Jalaran Ensemble puts two characters into one rendered scene, each speaking only their own lines and sitting naturally idle through the other’s.

Who Ensemble is for

For anyone whose scripts have two people in them. Narration and monologue animate perfectly well one character at a time, but the moment a script becomes a conversation you are left cutting between two separate clips and calling it a scene. This removes that compromise.

What Ensemble does

Ensemble invents nothing. It connects four things that already worked separately and had never been wired together: a script cast with two speakers, the per-word record of which voice said what, the character animation, and the background keying. The real work is dividing the audio. For each of the two voices it builds a track running the full length of the scene where that voice’s own moments survive intact and every other moment becomes true digital silence — not quieter, but genuinely nothing, because the idle behaviour keys off measured energy and an almost-silent passage is one it never recognises as a pause. Both characters are then animated against tracks of identical length, which is what puts them on a single shared clock with no alignment step to get wrong afterwards. Finally both are separated from their own backdrops and blended into one frame together, side by side, over whichever background you chose.

  • Each character falls quiet through the other’s lines, from the recorded account
  • Both halves share one clock by construction, with no alignment pass afterwards
  • Blended into a genuinely shared frame rather than cut back and forth
  • You cast it yourself: who plays whom, and who stands where

How Ensemble works

  1. Bring a two-voice clip

    The audio produced from a script cast with two speakers. Who is in it gets read from the file itself rather than from the script, so editing the words afterwards cannot quietly desynchronise anything.

  2. Assign the two voices

    Choose which character plays which voice, and which side of the frame each one stands on. Neither is ever inferred for you.

  3. Pick a backdrop, or skip it

    A saved background if you have one. Without it the pair simply stand against a plain field instead.

  4. Collect the finished scene

    Both characters animated across the whole running time, composited together, carrying the original conversation as its soundtrack.

What Ensemble does not do

Drawn and illustrated characters only, and that restriction was measured rather than assumed. The photographic animator keeps a mouth gently moving whenever it has nothing to say, which passes unnoticed anywhere else but not here, where somebody is deliberately quiet for half the running time. Exactly two people, exactly two placements — left and right. Procedural idle motion only; a second performance driven from reference footage, kept in step with the first, is genuinely harder and is not attempted. Casting and dialogue belong upstream, so nothing here lets you rewrite a line or recast a part. And it costs what it looks like it costs: two full animations plus the blending, queued one behind the other.

Common questions

Why are photographic characters excluded?

Because the animator behind them was never taught to rest. Measured against a real two-voice recording, its mouth carries on moving through silence at roughly two thirds of its talking activity and never fully settles, while the drawn path drops to a third and genuinely closes. In an ordinary render nobody would notice. Here one person stands quiet for half the scene, and a faint constant mumble is the first thing you would see.

Why not simply cut between two clips?

You can, and sometimes that is the better edit. What cutting cannot give you is both faces present at once — the listener reacting while the speaker talks. That is the shot this exists to make, and it is why both characters are animated for the entire duration rather than only while they happen to be talking.

How does it know when each character should be quiet?

The speech was recorded word by word as it was generated, with the voice attached to each one. Those markings are what get read back, so the quiet passages are precisely where the other person was talking. Nothing is detected, estimated or inferred from the waveform.

What happens if one of the two fails?

The scene stops and tells you which one — by side and by voice — along with why that half failed. Two animations mean two chances to go wrong, and a message saying only that rendering failed would leave you guessing at which of them to look at.