Videos

On screen text in video: size, duration and safe zones

On screen text works when it is large enough to read on a phone, stays up for at least 0.3 seconds per word, carries one idea per card and clears a 4.5 to 1 contrast ratio. Everything else, from font choice to animation, matters less than those four numbers.

· Co-founder

7 min read · Published

On screen text in a video should be big enough to read on a phone held at arm’s length, stay up for at least 0.3 seconds per word, say one thing per card, and sit at a contrast of at least 4.5 to 1 against whatever moves behind it. Get those four right and keep the text out of the parts of the frame that apps cover, and most text problems in video disappear.

None of these numbers were invented for social video. They come from the two places that have thought hardest about text over moving pictures: broadcast subtitling, where the BBC publishes detailed authoring rules, and web accessibility, where WCAG sets the contrast floor. They transfer well, because a viewer reading a headline over a moving background has exactly the problem a subtitle reader has.

Size: set a pixel floor, not a feeling

Text that looks generous on a laptop preview is often unreadable on a phone, because the preview is three times the physical size. The fix is to decide the size as a share of frame height and never go below it.

The BBC’s subtitle guidelines give that share. Subtitles are authored so each line fits a line height of 7 to 8 percent of the frame on 16:9 video, and 3.9 to 4.5 percent on 9:16 video. Work that out on the two canvases most videos use and something convenient happens: 7 to 8 percent of 1080 pixels is about 76 to 86 pixels, and 3.9 to 4.5 percent of 1920 pixels is about 75 to 86. The same pixel range works on both, so one floor covers a landscape and a vertical cut.

That floor is for text a viewer has to read. A headline should sit well above it. In the countdown offer video, scene 1 stacks Forty, Eight and Hours one word per beat, and the countdown card lets 48:00:00 loom over its single line of detail. The footer lines on the final scene, the terms, are the smallest text in the video, and they are the text that most needs to clear the floor.

Seconds on screen: count the words

The most common failure in short video is not small text but text that leaves before it is read. The BBC plans subtitles at 160 to 180 words per minute and sets a target minimum of about 0.3 seconds per word, so a four word subtitle should stay up for 1.2 seconds.

Apply it scene by scene. The countdown offer video runs four scenes in about 18 seconds:

Add time whenever the eye has other work to do. A price, a date, a code or a phone number is read more slowly than a slogan, and text that arrives during a transition loses the frames the transition takes.

Words per card: one idea, not the narration

Two findings from Mayer and Moreno’s 2003 review of multimedia learning decide how many words a card should carry. The redundancy effect: people understand a narrated presentation better when the same words are not shown as on screen text at the same time. The signaling effect: they understand it better when cues point out what matters, such as stressed key words and headings.

Together they give a simple rule. The narration carries the sentence; the card carries the word or number the viewer must keep. The gym membership video is a good model. The voice says “Joining fee, zero.” The card says FEE $0, with the code and the expiry line underneath. The screen does not repeat the voice; it gives the eye the fact to remember.

Tutorials bend the rule slightly because the text is often an instruction. The checklist scene in the expense claim how-to video lists the three things to attach while the voice summarises them, and it holds for 6 seconds so the list can be read twice.

Contrast: 4.5 to 1 is the floor

WCAG 2.2 success criterion 1.4.3 asks for a contrast ratio of at least 4.5 to 1 for normal text and 3 to 1 for large text, meaning 18 point or 14 point bold and up. It states that the requirement also covers images of text, and text burned into a video frame is exactly that.

Video makes contrast harder than a web page does, because the background moves. A white headline over a photograph can pass on the first frame and fail on the tenth when a bright window slides behind it. Check the worst frame of each scene, not the first. When a photo will sit behind text, put a panel, a scrim or a solid band under the words rather than trusting the image to stay dark.

Safe zones: keep text out of the interface

On vertical video, the apps draw their own names, captions, music tickers and buttons over your frame, and the top and bottom of the picture are the most crowded. The exact margins differ by platform and change often. The safe zone article collects what each platform publishes, and the short definition is in vertical video safe zones, so this post does not repeat the figures.

The rule for text is narrower than the rule for the frame. Anything that must be read, the offer, the code, the date, the call to action, goes in the middle band. The DCMP Captioning Key adds a second rule that applies to every format: text should not cover faces, mouths or graphics the viewer needs, and when it would, it moves.

The rules in one table

The table below the article puts every rule with a number next to its value and its source, so it can be pinned beside an edit. Two rows are worth reading twice. The line height rows give one pixel floor for both formats, and the width row explains why a vertical frame can hold a line that is 90 percent of its width while a landscape frame should stop at 68.

Common mistakes

Build it

The video builder ships 27 fonts, and the video fonts page shows them. Burned in captions are drawn as one dark pill centred at the bottom of the frame, capped at 72 percent of the frame width, with 26 pixel type on landscape and 30 pixel type on vertical, inset 44 pixels from the bottom on landscape and 120 on vertical, and switched on by default for every scene. Formats are landscape 1920 by 1080 and vertical 1080 by 1920. Switching between them rebuilds each scene’s layout for the new shape rather than cropping it, and hand edits are kept per format, so text can be resized in the vertical cut without disturbing the landscape one. The steps for checking each scene and exporting both versions are in the tutorial on exporting a video for YouTube, Reels and TikTok.

On screen text rules with a number attached, and where each number comes from (checked 13 September 2026)
RuleValueSource
Line height, landscape 16:97% to 8% of frame height, about 76 to 86 px on a 1080 px frameBBC subtitle guidelines, authoring font size
Line height, vertical 9:163.9% to 4.5% of frame height, about 75 to 86 px on a 1920 px frameBBC subtitle guidelines, authoring font size
Minimum time on screenAbout 0.3 seconds per word, so 1.2 seconds for four wordsBBC subtitle guidelines, target minimum timing
Reading speed to plan for160 to 180 words per minuteBBC subtitle guidelines, timing
Lines in one blockTwo on landscape or square, three on verticalBBC subtitle guidelines, number of lines
Width of a text block68% of a 16:9 frame, 90% of a 9:16 frameBBC subtitle guidelines, online maximum length
Contrast, normal textAt least 4.5 to 1 against what is behind itWCAG 2.2, success criterion 1.4.3
Contrast, large textAt least 3 to 1 for 18 point, or 14 point bold, and largerWCAG 2.2, success criterion 1.4.3
Narration and text togetherDo not show the narration word for word; show the key wordsMayer and Moreno 2003, redundancy and signaling
Faces and essential graphicsKeep text off them; move it up when it would cover themDCMP Captioning Key

A finished example

This countdown offer video example is a vertical flash sale promo told in four scenes, and the script with each scene's seconds is below. Forty eight hours lands as three stacked words, the offer scene stamps 30 percent off with the code TIDE30 on a ticket, the countdown card makes the deadline loom, and the close prints the exclusions in its footer instead of whispering them.

Read the countdown offer video example

Questions people ask

What font size should on screen text be in a video?

Think in share of frame height rather than points. The BBC authors subtitles at a line height of 7 to 8 percent of a landscape frame and 3.9 to 4.5 percent of a vertical one, which lands at roughly 75 to 86 pixels on both a 1080 and a 1920 pixel canvas. Treat that as the floor for text a viewer must read, and go larger for headlines.

How long should text stay on screen?

Allow about 0.3 seconds per word as a minimum, which is the BBC's target for subtitles, and add time when the viewer also has to look at a picture, a price or a code. A four word headline needs at least 1.2 seconds. A card with a dozen words of terms needs about four, and more if it arrives during a transition.

Should on screen text repeat the voiceover?

Not word for word. Mayer and Moreno's research on multimedia learning found people understand better when the same words are not presented twice at once, as narration and as printed text. Put the key word, the number or the offer on screen and let the voice carry the sentence. Captions are the exception, because they serve viewers who cannot hear the voice.

What contrast ratio does video text need?

WCAG 2.2 asks for at least 4.5 to 1 for normal text and 3 to 1 for large text, and it says the rule also applies to images of text, which is what text burned into a video frame is. Check the worst frame, not the first one, because a moving background can drop the ratio halfway through a scene.

How many words fit on one text card?

As few as carry one idea. A headline of two to five words reads in a glance, a supporting line of under ten words still reads in a second or two, and anything longer belongs in the narration. If a card needs two sentences to make sense, it is two cards, and the video is usually better for the extra cut.

Where should text go on a vertical video?

Away from the top strip, the bottom third and the right hand column, where the apps draw their own names, captions and buttons. The exact margins differ by platform and few platforms publish them. The linked safe zone article collects what each one does publish, and the practical rule is to keep anything that must be read inside the middle of the frame.

Written by

Indunil Asanka · Co-founder

Builds the generation pipelines behind OneCraft: the slide, flyer and poster layout engines, the document grid and the render workers that turn a written brief into a finished file.

LinkedIn profile

Sources

Written and checked by the OneCraft team. Last checked .

Make your own video

Describe what you need and the generator writes and designs it, then you edit anything you like.

See what it can make

Read next

For the steps inside the builder, read the guideon this topic.