Videos
Text to video with voiceover
The interesting problem in turning text into a video is not writing the scenes. It is timing: a spoken line and a visual cut are two different clocks, and every text to video tool has to decide which one wins. Here the voice is treated as one continuous track and the picture is fitted around it.

One voice track, not one clip per scene
Narration is assembled into a single track laid over the scene sequence, and that track is the source of truth for timing, for the captions and for the length of the finished composition. It is not a set of clips each trapped inside its own scene. That distinction is what stops a video full of unnatural pauses where a line finished early and the cut had not arrived yet.
Dynamic stretches a little, calm stretches fully
The two pace settings resolve the clash differently. On dynamic, a scene grows by at most one and a half times to accommodate its line, and whatever is left of the line simply flows over the next cut, which is how ad editing actually works. On calm, the scene stretches to fit the whole line before cutting, which is what a tutorial needs because the viewer is following along on their own screen.
Per scene control once it exists
Every scene has its own narration text, its own language and voice, a delay in seconds before the voice starts, a playback speed of 0.85, 0.9, 1, 1.1 or 1.25 times, and a caption toggle. Regenerating narration grows the scene to fit the new audio and remembers the original duration as a base, so a shorter rewrite shrinks the scene back rather than leaving dead air at the end.
What the text does not control
Worth being clear. Your text becomes the script and the on-screen copy, but it does not choose the layouts, the colours, the transitions or the camera moves. Those come from the scene library and are deterministic. Which is why the same text run twice produces recognisably the same video rather than two different designs, and why editing the text afterwards does not rearrange the video around it.
How it works, in three steps
Step 1
Paste the text as a brief
Up to 8,000 characters. Write it as what you want said, not as a description of a video, and the director will divide it into scenes.
Step 2
Pick the pace deliberately
Dynamic for anything promotional, calm for anything instructional. It changes how narration and cuts negotiate, not just how it looks.
Step 3
Fix the lines that run long
Open the voice dialog on any scene that feels rushed, shorten the line or drop the speed to 0.9, and the scene retimes itself.
The full walkthrough with screenshots is in the guide Add voiceover, captions and music to a video.
Limits worth knowing
- Narration is synthesised text to speech. There is no emphasis markup, no pause tags and no per word control beyond delay and speed.
- Captions are derived from the narration text, so editing an on-screen caption does not change what the captions say.
- Each voiceover clip you generate in the builder spends one credit.
- The brief box takes 8,000 characters, which is a long script but not a document.
Questions people ask
Will it use my exact words?
Not automatically. In promotional mode it rewrites for pace, aiming at 5 to 8 words a scene. If you need your wording preserved, say so in the brief and then correct each scene's narration in the voice dialog, which is where the final text lives.
How many words fit in a video?
Fewer than you think. A promo scene carries roughly 5 to 8 words, so a six scene short video is about 40 words of speech. A tutorial step takes 10 to 25 words, which is why tutorials run so much longer for the same scene count.
Can I turn the voice off entirely?
Yes. The generator has a voiceover switch, and with it off you get a silent video with on-screen copy. Captions come from the narration, so with no narration there are no captions either.
What happens if my text is a list?
It tends to become a feature list, a checklist or a how it works scene rather than one scene per bullet. Those roles are built to carry three or four items at once, which is usually a better video than four near identical cuts.
Make your own video
The button opens the generator with this use case already described. Change the wording to match yours, generate, then edit anything you like.
Create a video with OneCraftRelated pages
AI voiceover languages
Language here does two separate jobs, and they can come apart. It tells the director which language to write every line in, and it selects the locale the synthesised voice comes from. The writing side accepts anything you type. The voice side has a fixed map of 31 entries.
Turn a one line idea into a video
A single sentence is a legitimate input, and it is worth knowing exactly how much is being invented on your behalf when you use one. The answer is: the arc, the scene roles, every word on screen, every spoken line, the voice, the tone and often the format as well.
Faceless video maker
Faceless is not a style, it is a constraint you accept for a reason. Nobody has to be available, nothing has to be filmed twice when a price changes, and the same video can be reissued in another language without anybody appearing to speak it. The trade is that every second has to be carried by type, motion and voice.
More finished work of this kind is on the video examples hub.