Videos
How to make an explainer video
To make an explainer video, write five scenes in order: the problem, the solution in one sentence, how it works in a few short steps, one piece of proof and a single ask. Write the narration first, give each scene one idea and one picture, and keep the whole thing well under the few minutes viewers actually stay for.
Indunil Asanka · Co-founder
7 min read · Published
An explainer video makes one idea understandable in about a minute. The reliable way to build one is five parts in a fixed order: the problem the viewer has, the solution in one sentence, how it works in three or four short steps, one piece of proof, and one ask. Write the narration for those parts first, give every scene one idea and one picture, and stop. Most explainers that fail do so because they try to explain three ideas, not because the pictures were wrong.
The example below the post is the Wren onboarding video, a fictional HR portal explained to a new starter in eight landscape scenes and 50 seconds. It uses the same bones: why you are here, the steps, a recap, who to ask, and a close.
Why five parts, in this order
Two pieces of learning research do most of the work here.
The first is about length. Philip Guo and colleagues studied 6.9 million video watching sessions across four edX courses and found that shorter videos were much more engaging, with median engagement time at most six minutes regardless of how long the video was. Viewers of tutorial videos engaged for only two to three minutes. An explainer that makes its point in under two minutes is working with viewers instead of against them. Wistia’s 2026 report points the same way for marketing video, putting average engagement for videos under a minute at 52 percent.
The second is about order. Richard Mayer and Roxana Moreno’s 2003 review, Nine Ways to Reduce Cognitive Load in Multimedia Learning, describes a pretraining effect: people understand an explanation better when they already know the names and behaviours of the parts involved. That is what the problem and solution scenes are for. By the time the mechanism starts, the viewer knows what the pieces are called and why they should care.
Scene one: the problem
State a problem the viewer already has, in one factual sentence. Not a caricature of a frustrated person, and not a statistic about the industry. The dashboard tour opens with its problem as a question, “Why did sign-ups drop last week?”, and every scene after it answers part of that question.
Name the thing the rest of the video depends on. In the Coppervale Water plan in the table under this post, the problem line is “Most water filters get changed when someone remembers, which is rarely.” The filter is the part; the forgetting is the problem. Both are now in the viewer’s head before any mechanism appears.
Keep the visual simple and drawn. A cartridge fading from white to grey says more than a stock photo of someone frowning at a tap.
Scene two: the solution in one sentence
Name the solution and what it does, in words a viewer could repeat. “Coppervale posts a new filter the week yours runs out.” One claim, no features list. Mayer and Moreno’s coherence effect is the reason: people learn better when interesting but extraneous material is left out. Every extra benefit in this scene competes with the one the viewer needs to carry into the next.
The Wren video does this in its intro: “Welcome. Your first day in Wren, in four steps.” It names the product, the job and the size of the job in nine words, so the viewer knows how much attention to budget.
Scene three: how it works, in short steps
This is the middle of the explainer and the part most likely to overload. Mayer and Moreno’s answer is segmenting: break the explanation into bite size segments with time between them rather than one continuous stream. In a video, a segment is a scene. One step, one scene, one picture.
The Wren video gives each of its four steps seven or eight seconds and a single action: “Complete your profile: photo, phone, emergency contact.” Then “Add your bank details under Pay; they are checked before the first run.” That second line carries its reason, which is what gets a step done today rather than next week.
Two more effects from the same review shape these scenes. Temporal contiguity: people learn better when the animation and the narration describing it happen at the same time, so the picture should change as the line is spoken, not after. Signalling: cues that show where to look reduce wasted effort, which is why the expense claim walkthrough spotlights the control each step is about and dims the rest.
Guo’s study adds a note on style: tablet drawing tutorials were more engaging than slides, and the authors recommend motion and continuous visual flow. Drawn steps that build on screen tend to hold attention better than static bullet slides.
Scene four: proof
Proof answers the doubt a viewer actually has after the steps. For Coppervale that doubt is effort, so the proof is “Swapping takes two turns: old one out, new one in.” For a product with measured results, it is one number with its sample, the way the Replyline launch video prints “across 32 beta support teams” under its 41 percent.
Internal explainers prove themselves differently. The Wren video’s proof that onboarding is finishable is its checklist: “By Friday: profile, bank details, policies, meet your buddy.” Four ticks turn a vague first week into something a starter can complete.
If the proof is a chart, say the number in the narration. The W3C’s accessibility guidance notes that charts, graphs and on screen text need description for people who cannot see them, and that description is easier when it is part of the script from the start.
Scene five: the ask
End with one action. For a customer explainer, it is the first step of using the thing: “Enter your filter model to see your first delivery date.” For an internal one, it can be as small as the Wren close, “See you Monday,” after a drawn three card scene tells the starter who to ask.
Do not add a second ask, a social link or a recap of the whole video. The explainer has done its job if the viewer knows what the thing is, why it matters and what to do next.
The table below plans Coppervale’s explainer as seven scenes, with the narration line, the visual and the research principle each scene leans on. Swap in your own lines and keep the column on the right; it is the part that stops the video drifting into a feature tour.
Common mistakes
Mechanism before motive. Starting with how it works leaves the viewer holding steps with no reason to remember them.
Full sentences on screen while the voice reads them. Mayer and Moreno’s redundancy effect found identical narration and on screen text at once hurt learning. Use short labels and let captions do the accessibility job.
Eight steps in one explainer. Teach three, link a tutorial for the rest.
Proof that answers a question nobody asked. Awards and logos rarely address the viewer’s real doubt.
Stock faces. Drawn diagrams explain ideas better, and an explainer does not need a presenter; see what a faceless video is.
A tutorial in disguise. If the viewer has to follow along click by click, write a tutorial instead; the structure for that is in how to write a tutorial video script.
Build it
The AI explainer video maker builds from a brief of up to 8,000 characters plus up to 8 images and 5 documents, and documents are summarised rather than transcribed, so put the key steps in the brief itself. Choose Promotional for an explainer that ends on a sales ask, or Tutorial, from the Starter plan, for one that teaches; tutorial steps run 5 to 9 seconds with 10 to 25 words of narration and never get the energetic voice. Scenes can carry 10 chart kinds and 6 diagram kinds, including a funnel, a cycle and Venn diagrams. Captions are on by default and the format is landscape or vertical. The steps are in make a promo video with an AI voiceover.
| Scene | Narration line | Visual | Principle it uses |
|---|---|---|---|
| 1. Problem | Most water filters get changed when someone remembers, which is rarely. | A drawn cartridge fading from white to grey across a calendar | Pretraining: name the part the rest depends on |
| 2. Solution | Coppervale posts a new filter the week yours runs out. | One date circled, an envelope beside it | Coherence: one claim, nothing extra |
| 3. How it works, step 1 | Tell it your filter model and how many people live with you. | Two labelled fields filling in | Segmenting: one step per scene |
| 4. How it works, step 2 | It works out the week that filter will be spent. | A row of weeks counting down to one highlighted week | Temporal contiguity: the picture moves as the line is spoken |
| 5. How it works, step 3 | The new one arrives in your letterbox with a return label for the old one. | Envelope in, arrow back out | Signalling: one highlighted object |
| 6. Proof | Swapping takes two turns: old one out, new one in. | A two step diagram of the cartridge | Prove the doubt the viewer actually has |
| 7. Ask | Enter your filter model to see your first delivery date. | The model field and one button | One action |
A finished example
This onboarding video example walks a new starter's first day in a fictional HR portal, and the script with per step seconds is below. Four steps cover profile, bank details, policies and the first week plan, the checklist scene turns Friday's obligations into four ticks, and a drawn three card scene answers the question every new starter actually has: who do I ask.
Read the onboarding video exampleQuestions people ask
How long should an explainer video be?
Shorter than you think. In Guo and colleagues' study of 6.9 million video watching sessions, median engagement time was at most six minutes whatever the video's length, and viewers of tutorial videos engaged for only two to three minutes. A five part explainer usually needs thirty to ninety seconds. If yours runs past two minutes, it is probably two explainers, or an explainer and a tutorial.
Should the words on screen repeat the narration?
Not in full. Mayer and Moreno's redundancy effect found people learned better when words came as narration rather than as identical narration and on screen text at once. Keep on screen text to a short label or number and let the voice carry the sentence. Captions are the exception: they serve people watching muted or who cannot hear, so keep them on.
Animation or live footage for an explainer?
Most explainers are about things with no picture, a process, a pricing model, a rule, so drawn scenes usually beat footage. Guo's study also found tablet drawing tutorials more engaging than slides, and recommended motion and continuous visual flow. Use real footage or screenshots where the thing being explained is physical or on a screen, and drawings for everything that is an idea.
How many steps can the how it works section hold?
Three or four is the comfortable range, each in its own scene or as one card each in a single scene. A long unbroken run of steps overloads viewers, which is the problem Mayer and Moreno's segmenting method was proposed to solve. If the process genuinely has eight steps, explain the three that matter in the explainer and teach all eight in a separate how to video.
What is the difference between an explainer and a tutorial?
An explainer answers what it is and why it matters, so a viewer understands the idea. A tutorial answers how to do it, so a viewer completes a task while watching. The explainer ends on an ask such as a trial or a next step; the tutorial ends when the task is done. Many products need both, linked together.
Do I write the script or the visuals first?
Write the narration first, one line per scene, then choose one visual that shows what each line says. Working from pictures first tends to produce a tour of nice images with a voice trying to connect them. The W3C's guidance adds a practical reason: description of charts and other visual information is easier when it is written into the script from the start.
Written by
Indunil Asanka · Co-founder
Builds the generation pipelines behind OneCraft: the slide, flyer and poster layout engines, the document grid and the render workers that turn a written brief into a finished file.
LinkedIn profileWritten and checked by the OneCraft team. Last checked .
Make your own video
Describe what you need and the generator writes and designs it, then you edit anything you like.
See what it can makeRead next
How many scenes in a promo video, counted across 20 videos
A short promo video runs six scenes: that is the median across the 5 promo videos and 10 creator style ads counted here, with scenes held for 4.3 and 4.8 seconds. Tutorials take eight scenes of about seven seconds, because a viewer is following steps rather than a pitch.
Video safe zones for Reels, TikTok and Shorts
Every vertical platform covers a different part of your frame with its own buttons and captions, and almost none of them will tell you where. This is what each one actually publishes, checked on 1 September 2026, and what to do about the parts nobody publishes.
How to write a tutorial video script
A tutorial script is a list of steps with one sentence of why on each, spoken slower than an advert and shown one action at a time. The numbers are what make it work: seconds per step, words per step, and exactly one screen per step.
For the steps inside the builder, read the guideon this topic.