Videos

How to make a video without filming

Most business videos are not filmed. They are assembled from images, screens, clips and a voice, and this post sets out the four sources, what each is honestly good for, and the one that still puts a person on screen.

· Co-founder

5 min read · Published

The thing that stops most small businesses making video is not the budget and not the editing. It is the assumption that somebody has to stand in front of a camera, and that if nobody will, there is no video to make. Look at the marketing videos you actually watch and most of them contain no original footage at all.

Filming is optional

A video is a sequence of frames with something over the top. Where the frames come from is a supply question, not a creative one, and there are four supplies. Every video worth making mixes at least two of them.

The reason to think about it this way is that the four have completely different failure modes. Generated images fail by being generic. Screenshots fail by being illegible. Your own clips fail by being too long. An AI creator fails by being uncanny. Knowing which one you are leaning on tells you what to watch for.

Source one: generated images

You write a prompt for an image slot and get a picture back, sized to fit the box it is going in. The generator snaps your slot’s shape to one of eight aspect ratios, from 21:9 down to 9:16, so a wide banner does not come back square and get cropped.

What this is genuinely good for: atmosphere, abstract concepts, backgrounds, and the shots that would need a location and a crew. A quiet office at dawn. A hand on a steering wheel. A city from above.

What it is not good for is your product. A generated picture of a coffee machine is a coffee machine, not yours, and anyone who knows the product will see it. Use generated images around the real ones, not instead of them.

The practical rule is that generated imagery should never be the shot the argument rests on. If the video’s claim is “our packaging is smaller”, the packaging has to be photographed.

Source two: screenshots and screen capture

For anything that lives on a screen, this is the best material you have, and it is worth being precise about what the capture tool does. It asks the browser to share a screen and takes a still PNG frame each time you press capture or hit the space bar. It does not record motion, and there is no recording to process afterwards.

That sounds like a limitation and is mostly the opposite. A step-by-step video wants one clean frame per step, not a real-time cursor wandering across a menu. Motion is added afterwards: mark the point that was clicked and the cursor glides there, a ripple fires, and the camera eases in at 1.5x or 2.1x over about a second.

It is desktop browsers only, because it depends on a screen sharing capability phones do not have. On a phone the generator tells you to upload screenshots instead, which works exactly as well.

If you are building a walkthrough this way, the tutorial script guide covers how many words each of those frames can carry.

Source three: clips and photos you already own

The most under-used supply. Almost every business has a folder of photographs taken for something else: the premises, the team, the product on a table, a delivery van. They are specific, they are real, and specificity is the whole difference between a video that could be anyone’s and one that is obviously yours.

Clips go in as mp4, webm or mov up to 200 MB, trimmed at both ends, and set either to cover the frame or to sit inside it. Images go in at up to 15 MB.

One honest limitation worth knowing before you plan around it: a clip’s own audio can be muted or kept, but the option to keep it and duck the music under it is not implemented. The music ducks under narration and not under clip audio. If a clip has speech in it, plan for that scene to have no voiceover over the top.

Source four: an AI creator

The one that does put a person on screen. You write one description of the person, a single portrait is generated from it, one synthetic voice speaks all their lines, and that portrait is animated for each talking beat.

The design constraint is deliberate: one persona, one voice, and at most three generated clips per video, so the face and the voice cannot drift between scenes. It sits behind the Pro plan.

Be clear-eyed about what it is. It is a stylised presenter, not a customer, and it must never be scripted as if it were reporting a real person’s experience, which since October 2024 is squarely within the United States rule on fabricated testimonials. What goes in its mouth is covered in the UGC script guide.

The comparison

The table at the foot of this post sets the four side by side, including what each one cannot do. Read the fourth column first. It is the one that decides which source a given scene needs.

Voice and captions carry the rest

Here is the part people skip. With no filmed footage, the voice is doing most of the work of holding the video together, and the captions are doing most of the work of being understood.

Feed video autoplays muted. LinkedIn says so in its own guidance and recommends captions for exactly that reason. TikTok’s creative advice assumes on-screen text and suggests keeping it to five to ten words a second. So the sequence that actually matters is: write the narration, caption it, and only then decide what pictures go behind it. Where on the frame those captions can safely sit is covered in the safe zone article.

That order also stops the most common failure of a no-filming video, which is a beautiful sequence of images that does not say anything.

What to make first

Take the material you already have and count it. If you have twenty screenshots, make the walkthrough. If you have thirty photographs of the product, make the promotional video. If you have neither, take twelve photographs on your phone this afternoon, because that will beat any generated set.

Then fill the gaps with the other three sources rather than starting from them. A video assembled around real material with generated images in the joins looks like a real company. A video assembled entirely from generated images looks like a template, and viewers have learned to recognise it.

Four ways to fill a video frame without a camera, with what each one cannot do and the limits the builder enforces (read from the source on 1 September 2026)
SourceWhat you supplyWhat it suitsWhat it cannot doLimits
Generated imagesA prompt per image slotBackgrounds, mood, concepts, anything you cannot photographShow your actual product or your actual premisesSnapped to one of 8 aspect ratios from 21:9 to 9:16; 2,000 character prompt
Screenshots and screen captureStill frames of your own screenSoftware demos, dashboards, any workflow that lives on a screenRecord motion; it captures stills, one per pressDesktop browsers only; each shot marked no zoom, 1.5x or 2.1x on one point
Clips and photos you ownFiles you already haveReal premises, real staff, the actual productDuck the music under the clip's own audio, which is not implementedmp4, webm or mov up to 200 MB; images up to 15 MB; trim, then cover or contain
An AI creatorOne written description of the personCreator-style ads that need a face and a first-person voiceBe two different people, or appear in more than three clipsOne persona, one portrait, one voice, at most 3 generated clips; Pro plan and up

Questions people ask

Can I record my screen as video?

Not as motion. The capture tool asks the browser to share a screen and then takes a still PNG frame each time you press Capture step or the space bar. Nothing is recorded and nothing is processed from a recording. For a walkthrough this is usually better, because one clear frame per step is easier to follow than a real-time cursor.

Do I need stock footage?

No, though it helps. A promotional video made entirely of text scenes reads as cheap, so the planner will ask for at least one footage scene when you have supplied no images of your own. If you have real photographs of the product or the premises, use those first. Generated images are for the parts you cannot photograph.

Is an AI creator obvious to viewers?

Often yes, and increasingly so as audiences see more of them. Treat it as a stylised presenter rather than someone viewers will take for real. The US rule on fake testimonials, in force since October 2024, covers a testimonial that misrepresents whether the person giving it exists or what their experience was. A synthetic presenter may explain a product; it must not describe a customer experience nobody had.

Can I use photos from my phone?

Yes, and they are usually the strongest material you have. Upload up to eight images with the brief, or add more later in the builder, at up to 15 MB each in PNG, JPEG, WebP or GIF. Every uploaded image is analysed for its shape and content so the planner can put a tall photo in a tall slot rather than cropping it.

What file formats can I upload?

Video clips as mp4, webm or mov up to 200 MB each. Images as PNG, JPEG, WebP, GIF or SVG up to 15 MB. Audio up to 25 MB. Documents attached to the brief can be PDF or plain text, and are summarised rather than transcribed, so a fifty page PDF contributes a summary and key points, not its full text.

Is there a watermark on the video?

Only if you add one. A wordmark or a logo can be placed in the bottom left, fading in after the opening scene and out before the closing one at 72% opacity. The logo is supplied as a URL rather than uploaded through a button, which catches people out. Nothing is added to the video unless you ask for it.

Which of the four should I start with?

Whichever one uses material you already have. Screenshots if you sell software, photographs if you sell a physical thing, generated images only for the gaps. Starting with generated images is the most common mistake, because it produces a video that looks like anyone's, and the specific real picture is what makes it look like yours.

Written by

Nuwan Madhusanka · Co-founder

Works across the builders and the export paths: how a form becomes a PDF, how a flyer canvas becomes a print file, and how a signed document carries its audit trail.

LinkedIn profile

Sources

Written and checked by the OneCraft team. Last checked .

Make your own video

Describe what you need and the generator writes and designs it, then you edit anything you like.

See what it can make

Read next

For the steps inside the builder, read the guideon this topic.