One Timeline: How Voice, Captions and Sound Land on the Same Frame
Voice, captions, graphics and sound cut to one timeline, so every word lands on its frame, at any length, in every size and hook you want to test.
Most edits are built in layers, by hand. The picture goes down first. Someone lays a voice over it. Someone else types captions and nudges each one until it looks close. Sound effects go in last, placed by ear. Every layer is timed against the others by eye, so every change to one layer means checking all the rest.
We do it the other way round. One timeline holds every timing in the film. The picture, the captions, the graphics and the sound all read their timing from it. Change the timeline and everything moves together.
The timeline is a file
The film on our homepage is a good example, because it was built this way and you can watch it. It runs 63 seconds, which is 1,890 frames. Every one of those frame numbers lives in one file. The picture reads it. So does the sound: the film has 136 sound cues, and each one is generated from that file rather than placed by hand.
When we added ten seconds to the middle of the film this week, nobody moved a sound. The file changed, every later cue moved with it, and the new mix came back in sync. We checked by comparing the finished audio against the mix: the offset was zero.
That is the whole idea. If timing lives in one place, it cannot disagree with itself.
The voice sets the clock
On a narrated piece, the voice is the timeline. Before anything is recorded, each scene gets an estimated length from its script: the number of words at a normal speaking pace, plus a short lead in and a short tail. That gives a first cut to review.
Once the voice exists, the estimate is replaced by the real read. Speech-to-text gives a start and end time for every word. The scene lengths come from those times. A graphic that belongs to a phrase is keyed to the first word of that phrase, so it lands when the word is said, not when someone guessed it would be.
The voice can be your own recording, cleaned and levelled. It can be an AI voice. It can be a clone made from your own recordings, with your permission. An AI voice is generated one section at a time, so the breaths fall at the ends of thoughts the way a person takes them. Then it is timed to the word, the same as a recording.
Captions in your type
The captions come from the same word times, so they cannot drift from the voice. They are set in your typeface and your colours, not a stock caption style. The word being spoken can light up as it is said. You can have them burned into the picture, delivered as a caption file for the platform, or both.
Because the captions are data, changing one is a text edit. Fix a spelling, and every size and every version of the film picks up the fix on the next render.
Any length
None of this depends on the film being short. A long piece is a set of scenes, and each scene is timed from its own words. A thirty-second clip, a three-minute explainer, a full episode and a course with chapters all use the same method. The only difference is how many scenes there are.
Long pieces are where hand timing breaks down. Over forty minutes, small errors add up, and every revision means scrubbing through the whole thing to find what moved. When the timing is computed, a revision to one scene retimes that scene and everything after it, and the rest stays put.
If your team edits in-house, you get more than the finished film. We can hand over the scene clips, a timeline file your editor opens in Premiere, the voice on its own track and the caption file, so they can pick it up from there.
Every version, from the same frame
Once a film is approved, the versions are cheap. Three hooks, four sizes and three lengths is 36 versions of one cut: 16:9 for YouTube, 9:16 for Shorts, Reels and TikTok, 4:5 for the feed, 1:1 where square still works.
This matters most for testing. An A/B test only tells you something if the two versions differ in one thing. When a person re-cuts each version by hand, they drift in small ways: a caption lands a frame later, a shot runs a beat longer. When every version renders from the same timeline, hook A and hook B are identical except for the hook. If one wins, you know why.
Ads sit on top of this. When a piece has held attention, the paid versions come from the same session, in your own ad account.
What you get
- A voice timed to the word: your recording, an AI voice, or a clone of yours made with your permission.
- Captions in your type, from the same word times, burned in or as a file.
- Graphics and sound that land on the words they belong to.
- Any length, from a clip to a course with chapters.
- Every size and every hook you want to test, from one approved cut.
- If you want them, the scene clips, the edit timeline and the separate voice track.
This is what the Retainer makes every month. How it works has the steps. The first one is a free film about your company, made from your own site.
Your company film
Want it done for you?
Apply. If it fits, we make a film of about a minute about your company, from your own site, in your brand and your words. The call comes after, with the film in hand.