Videos
Slideshow video with narration
The reason most photo slideshows are unwatchable is that every picture gets the same three seconds and one voiceover runs across the lot. The interesting shot flies past, the filler lingers, and the narration is talking about image two while image five is on screen.
A frame is one picture plus its words
The unit here is not a slide, it is a frame: an image and the text that enters with it, held, then replaced by the next. A scene built from several such pairs shows one drag handle per frame in the timeline, so you set the hold for each picture individually. Frames can accumulate, building a pile or a grid on screen, or replace one another like a slideshow proper, where each frame's caption disappears as the next arrives.
One narration line per picture
Rather than a single track running over everything, a sequenced scene takes a separate narration entry per frame, index aligned to the beats. When any of those entries has audio it replaces the scene level narration entirely, so the voice you hear is the line written about the picture on screen. The director is told to keep each of those lines short, around five to eight words, because a frame is only on screen for a moment.
Timing that follows the voice
When you record or synthesise narration for a scene, the scene grows to fit the audio, adding the delay, the clip at its playback speed and a small tail. The original duration is remembered as a base, so shortening the line shrinks the scene back rather than leaving a gap. Turning on sequential mode does not change the total length either: the existing duration is spread evenly across the frames first, and you adjust from there.
Not a Ken Burns machine
It is worth being clear about what this does not do. There is no per photo pan and zoom you set by dragging a rectangle. Motion comes from the scene layout, from the pace setting, and from the entrance animation chosen per slot, of which there are nineteen. Pictures are not upscaled either, apart from a hero slot filling most of the frame, so a small photo stays at its native size rather than being blown up soft.
How it works, in three steps
Step 1
Upload the pictures in order
Display order is what the director reads, and images that clearly belong together are grouped and shown in one scene in the order given.
Step 2
Write a line per picture
In the brief, describe each photo in a few words. Those become the per frame narration, so the voice matches what is on screen.
Step 3
Drag the holds
Open the scene in the timeline and pull each frame handle. The minimum hold is a third of a second, which is enough for a fast montage.
The full walkthrough with screenshots is in the guide Add voiceover, captions and music to a video.
Limits worth knowing
- Eight images per generation, so a long photo set has to be split across more than one video.
- There is no per image pan and zoom control. Movement comes from the layout and the pace setting.
- A gallery layout tops out at eight pictures in one scene, and most hold three or four.
- Music loops over the whole video and ducks to 22 percent under the voice. It cannot be cut to the picture changes.
Questions people ask
Can each photo have a different length?
Yes, that is the point of frames. Each one has its own drag handle in the timeline and its own hold, with a minimum of 0.3 seconds. A hero shot can hold for three seconds while detail shots pass in one.
Do captions and narration have to match?
They do not have to, but they should. Captions are generated from the narration, so if you write the on-screen caption to say something different, the burned in text will follow the voice, not the caption.
Can I add music under it?
Yes. Thirteen tracks are in the library, or upload your own. Volume starts at 25 percent and caps at 60, and it ducks automatically wherever narration plays.
What resolution do the photos need to be?
Big enough for the slot they land in, since nothing is upscaled beyond a frame filling hero. Aim for at least the canvas dimension in the direction the slot is largest, so 1920 wide for a full width landscape hero.
Make your own video
The button opens the generator with this use case already described. Change the wording to match yours, generate, then edit anything you like.
Create a video with OneCraftRelated pages
AI voice options
There are two different voice pickers here and they work at different scales. When a video is generated, the director chooses from eight characters, four styles across two genders. In the builder afterwards, every voice Google offers is searchable per scene. Most people only need the first.
Add captions to a video
Most tools transcribe the audio to make captions, which means the captions inherit whatever the speech recogniser heard. Here they come from the other direction: the narration text you wrote is the source, the speech is synthesised from it, and the captions are generated from the same string.
Text to video with voiceover
The interesting problem in turning text into a video is not writing the scenes. It is timing: a spoken line and a visual cut are two different clocks, and every text to video tool has to decide which one wins. Here the voice is treated as one continuous track and the picture is fitted around it.
More finished work of this kind is on the video examples hub.