Videos

How to write a tutorial video script

A tutorial script is a list of steps with one sentence of why on each, spoken slower than an advert and shown one action at a time. The numbers are what make it work: seconds per step, words per step, and exactly one screen per step.

· Co-founder

7 min read · Published

The difference between a tutorial video that works and one that does not is almost never production quality. It is that the second one tries to teach two things in one shot, and the viewer, who is following along in another window, misses one of them and stops the video.

A tutorial is a list

Before any writing, get the task down as a numbered list of actions, each one a thing a person does: click something, type something, choose something. If a line on that list does not have a verb in it, it is context, not a step, and it belongs in the opening or in a diagram.

This sounds obvious and it is the part most scripts skip. People write a tutorial as prose, then discover halfway through recording that step four is really three steps, and the screenshot they took covers all of them at once. Microsoft’s style guide puts the same rule in one line: use a separate step for each instruction, and only combine short steps that happen in the same place in the interface.

The list also tells you whether the video is the right format at all. Two or three steps is a paragraph in your documentation. More than about ten and you are teaching a workflow, not a task, and it wants splitting. Guo, Kim and Rubin’s 2014 study of 6.9 million viewing sessions across 862 edX videos found median engagement time was at most six minutes whatever the video’s length, and their recommendation was blunt: plan the segments first, and keep each one under six minutes.

The five parts

A tutorial script has five kinds of scene and no others. The table at the foot of this post gives each one its budget in seconds and words.

The opening says what the viewer will be able to do at the end and where they are starting from. It is not a welcome, it is not the brand, and it is not a summary of the whole video. Six to twelve words.

Each step is one action on one screen. Five to nine seconds, ten to twenty five words.

The reason is an optional scene that explains why the thing works, and it is the only scene that gets a drawn diagram instead of a screenshot. Use it once, in the middle, when a step will otherwise feel arbitrary.

The recap lists the steps back in the order they were done. This is not padding. A person who was following along and fell behind uses it to find where they lost the thread.

The close gives one next action. Not three.

Words per step and why 25 is the ceiling

The ceiling is arithmetic. A step scene runs at most nine seconds. Twenty five words in nine seconds is about 167 words a minute, which is a brisk but comfortable speaking pace. Push to thirty five words and either the line is rushed or the step stretches past the point where a viewer is still looking at the screen.

The floor matters too. Fewer than ten words means you have written a caption, not an instruction, and the viewer is left to infer the part you left out. “Now open settings” is a caption. “Open Settings from the account menu, then choose Team members from the list on the left” is an instruction, and it is nineteen words.

For reference, the same edX study measured instructor speaking rates from 48 to 254 words a minute with a mean of 156, so 167 sits at the fast end of normal and 130 at the comfortable middle.

The reliable test is to read the line out loud with a timer. The speaking time calculator will do it for a whole script at once, which beats generating audio to find out that step six runs four seconds long.

Show one action per step

One step, one screen, one thing changing. The moment a step contains two clicks, the screenshot has to show both and the viewer has to guess the order.

This is also the rule that makes the visuals work. Each step carries a numbered badge, one instruction line and one marked point on the screenshot, and the marked point is where the cursor glides to before the camera eases in on it. Two marked points in one frame means two zooms in five seconds, which reads as fussiness.

If a step genuinely has no screen, it takes a drawn layout instead. A step that is “wait for the confirmation email” does not need a picture of an inbox.

The on-screen instruction line should say the same action as the narration in fewer words, never the same words. Mayer and Moreno’s review of multimedia learning found that when a picture is doing the explaining, narration alone beats narration plus matching on-screen text, with a median effect size of 0.69 across three studies. A caption that duplicates the voice makes the step harder to follow, not easier.

The “why” sentence

Every three or four steps, add half a sentence of reason to a step you already have. Not a new scene, just a clause.

“Choose Team members from the left, because permissions are set per team rather than per person.”

That clause is the difference between a viewer who can repeat the task and one who can adapt it. It also costs you about eight words, which is why it belongs inside an existing step rather than in a scene of its own.

Do not do it on every step. A script where every instruction is justified reads as defensive and doubles the runtime.

Pace and voice

Two things are set for you when the mode is tutorial rather than promotional, and both are correct.

The pace is forced to calm: every cut is a crossfade and the camera does one slow push in rather than drifting around. An advert’s hard cuts and moving camera make a walkthrough feel unstable, and a viewer trying to match your screen to theirs needs the frame to hold still.

The voice is forced to instructional, and it cannot be overridden by the model even when it asks for the energetic one. The full mapping of styles and voices, including which languages get which, is in the voiceover language guide.

Second person throughout. “You” and “your”, never “we” and never the passive. Name the actual button, in the actual capitalisation the interface uses, because the viewer is scanning the screen for that exact string.

The worked script

For a fictional invoicing tool called Harborline, five steps, with word counts.

Opening (9 words). “By the end of this you will have invited a teammate.”

Step 1 (17 words). “Open the account menu in the top right and choose Settings, then Team members from the left.”

Step 2 (21 words). “Click Invite member. Type their work email address, because the invitation is matched to the address, not to the name.”

Step 3 (19 words). “Choose a role. Editor can send invoices and cannot see the bank details, which is the right default.”

Step 4 (14 words). “Press Send invitation. The row appears immediately and stays greyed out until they accept.”

Step 5 (16 words). “Check the status column tomorrow. An invitation expires after seven days and Resend starts a fresh one.”

Recap (18 words). “Settings, then Team members. Invite member, work email, pick a role, send, and check the status column.”

Close (10 words). “Set their permissions next, in the same Team members panel.”

Every step is one screen and one action, no step names its own number, and the longest line is 21 words. Total narration is 124 words, which runs about 50 seconds, and the whole video lands near a minute and a quarter with the intro and recap around it.

If you have the screenshots already, turning a screenshot set into a video is the same script with the frames dropped in.

The five parts of a tutorial video, with the seconds and narration words each one is budgeted (read from the generator's scene rules on 1 September 2026)
PartRole nameSecondsNarration wordsWhat is on screen
Openingintro3 to 56 to 12What the viewer will be able to do, and where they are starting from
Each steptutorial5 to 910 to 25One screenshot, one step badge, one instruction line
The reasonhow_it_works5 to 910 to 25A drawn diagram, no screenshot
Recapchecklist5 to 910 to 25The steps listed back in the order they were done
Closeoutro3 to 56 to 12The one thing to do next

Questions people ask

How many steps should a tutorial video have?

Five to eight for a single task. At 5 to 9 seconds a step plus an opening and a close, eight steps is already over a minute, which is long for a screencast. If your list runs past ten, you are teaching two tasks and should cut it into two videos that link to each other. A shorter video gets rewatched.

Should I show my face?

For software, no. The screen is the thing being taught and a face in the corner competes with it for attention. Reserve a person on camera for the parts where trust matters more than instruction, such as an introduction from the team, and keep it out of the steps themselves. A voice does the human work perfectly well.

How fast should the narration be?

Slower than an advert. Ten to twenty five words in a step that runs 5 to 9 seconds lands between 120 and 170 words a minute, which is conversational rather than brisk. If a line feels rushed on playback, set that scene's narration speed to 0.9x or 0.85x rather than cutting words, because the words are the instruction.

Can I use screenshots instead of recording my screen?

Yes, and the capture tool takes still frames rather than motion video, so screenshots are the native input, not a compromise. One frame per step is what the format wants. Motion is added afterwards as a cursor glide to the point you mark, a click ripple and a slow zoom, which is more legible than a real recording of the same action.

Should every step have a title on screen?

Every step gets a numbered badge and one instruction line, and that line should say the same action as the narration in fewer words. Do not say the number out loud: the badge already carries it, and reading it aloud wastes half a second per step and makes the script harder to edit when you insert a step later.

Is a tutorial video a paid feature?

Tutorial mode sits behind the Starter plan and above. A free account can make promotional videos and gets a 402 with the type button locked when it selects tutorial. The script method here is worth following whoever you make the video with, since the timings come from how fast a person can read a screen, not from the tool.

How long should the whole video be?

Under two minutes for one task. The edX study of 6.9 million viewing sessions found people watch only two to three minutes of a tutorial video whatever its total length, and that they rewatch tutorials far more than lectures. That argues for short videos built to be skimmed and returned to, rather than one long one. In documentation, more and shorter wins.

Written by

Nuwan Madhusanka · Co-founder

Works across the builders and the export paths: how a form becomes a PDF, how a flyer canvas becomes a print file, and how a signed document carries its audit trail.

LinkedIn profile

Sources

Written and checked by the OneCraft team. Last checked .

Make your own video

Describe what you need and the generator writes and designs it, then you edit anything you like.

See what it can make

Read next

For the steps inside the builder, read the guideon this topic.