Founder:pay once, never think about it again. Every feature we ship from here on is included.33 seats leftClaim a founder seat

Add captions to a video

Burns readable captions into the picture, timed per word.

The short answer

Adding captions to a video means burning timed text into the picture so it plays everywhere with the sound off. The two things that decide whether they work are timing and placement: captions timed per word hold attention in a way a whole line appearing at once does not, and captions placed clear of the platform’s own interface are the difference between readable and half covered.

First Short free. No account, no software, nothing to upload.

This is the working tool, not a description of one. The first Short is free and needs no account. Everything on this page describes what the product does today, checked August 2026.

Why it matters

The sound is off, and that changes the job

Short video is watched in places where sound is rude. A commute, an office, a bed with somebody asleep next to you, a queue. The clip has to work as a silent thing first and a heard thing second, and captions are what makes that possible. This is not a nicety bolted on for accessibility, although it is that too and that alone would be reason enough. It is the difference between a clip that plays and a clip that gets scrolled past.

It also changes what good captions look like. Subtitles for a film are designed to be unobtrusive, sitting quietly at the bottom, waiting to be read when needed. Captions for a Short are the opposite: they are a main element of the composition, they carry the pacing, and they are frequently the only thing moving on screen. Treating one like the other is why so many captioned Shorts feel flat.

Timing

Word by word, and the small lie that makes it feel right

A caption that appears a whole line at a time gives away the ending. The eye reads faster than the mouth speaks, so the viewer finishes the sentence before the speaker does and then waits, which is exactly the moment attention leaves. Lighting one word at a time keeps the reading speed and the speaking speed together, and the clip holds.

There is a small deliberate inaccuracy in how that is done here and it is worth explaining. A word lit at the exact millisecond the transcriber says it starts reads as late, because the viewer has already heard the beginning of it. So the highlight leads very slightly, by an amount small enough to still feel word accurate and large enough to stop feeling behind. Only colour and the box opacity ramp during that lead, never the size, because growing the word would shift every word after it on the line and the whole caption would jitter.

Individual words can also be emphasised, which is what turns a caption from a transcript into a piece of editing. There is an automatic pass that marks the words held longest and landing loudest, a few per clip and never two in a row, and you can add or remove them by hand.

Placement

The interface is drawn on top of your video

Every platform draws its own things over the picture. A caption placed at the very bottom of the frame, which is where a film would put it, ends up underneath a username, a description, a row of buttons, or a progress bar, depending on where it is playing. The safe area is narrower than the frame, and the practical answer is to keep captions comfortably above the bottom third and inside the sides.

The same applies at the top, where a header and a back button live. Captions here default to a position that clears both, and the position is adjustable, because the safe area is not identical on every platform and the clip you are making might only ever be going to one of them.

Contrast is the other half of readability, and the reason the caption styles here include a plate and an outline rather than only plain text. Plain white text over a bright, busy or moving background is unreadable at exactly the moments you most want it read. An outline keeps the letterforms legible against anything, and the stroke is painted behind the fill rather than over it, so thin faces do not get eaten from the inside.

Practical

Doing it without an editor open

None of this needs a timeline. Paste a link, let it transcribe, fix anything it misheard by retyping the line, pick a style, and export. The whole thing runs in a browser, which means it also runs on a phone, which is where a fair amount of this work actually gets done.

If you also want a separate subtitle file for a long upload elsewhere, that comes out of the same transcript as an SRT or VTT, so the burned in captions and the file cannot say different things.

Step by step

How to do it

01

Bring in the clip

Paste a public YouTube link. The first Short needs no account at all.

02

Transcribe it

Word level timings are produced automatically. If the video carries its own captions, those get used first, because a person usually wrote them.

03

Fix the words

Retype any misheard line and the timings follow. For a song, paste the lyrics and have them timed to the track instead of transcribed.

04

Style and place

Choose a face, a treatment and a position that clears the platform interface. Add emphasis to the words that land, or let the automatic pass mark them.

05

Export

A 1080 by 1920 MP4 with the captions burned in, plus an SRT or VTT file from the same words if you want one.

The hard facts

Specifications

The numbers rather than the adjectives, including the limits. If one of these is a dealbreaker it is better learned here than after paying.

TimingWord by word, with a small deliberate lead so it does not read late
EmphasisAutomatic on the strongest words, or marked by hand
TreatmentsPlain, outline with the stroke behind the fill, and a plate
PlacementDefaults clear of platform interface, and adjustable
ScriptsLatin and right to left including Arabic, with four Arabic faces
EmojiReal colour emoji, at most one per caption line
Filler wordsRemovable from the captions, not from the audio
Also exportsAn SRT or VTT file built from the same words
Output1080 by 1920 MP4, H.264, on every tier including free
Runs inAny modern browser. Nothing to install, works on a phone
Free tierOne Short a day, up to 30 seconds, with a small corner watermark
Account neededNot for your first Short. Captions and subtitle files need a free account
Questions

What people ask about this

How do I add captions to a video for free?
Paste a link here, let it transcribe, fix anything misheard and export. The free tier gives you one Short a day up to thirty seconds with a small corner watermark, and every caption style is included rather than held back for a paid plan. Captions need a free account because transcription costs real money.
Should captions be burned in or a separate file?
Burned in for a Short, because most people scrolling have the sound off and will not open a settings menu, and because some platforms strip caption tracks entirely. A separate file for a long upload, because the platform can index it, the viewer can turn it off and you can edit it later. Doing both is normal and both come from the same transcript here.
Why do word by word captions work better?
Because the eye reads faster than the mouth speaks. A whole line appearing at once gives away the end of the sentence, the viewer finishes early and waits, and that wait is where attention leaves. Lighting one word at a time keeps reading speed and speaking speed together.
Where should captions sit in the frame?
Comfortably above the bottom third and inside the sides, because every platform draws a username, a description, buttons or a progress bar over the picture. The very bottom of the frame, which is where film subtitles live, is exactly where an interface will cover them.
Can I change the font and colour?
Yes, and the choices are shown in their own faces rather than as a list of names to imagine. There are also three treatments: plain, an outline for readability over anything, and a plate. Over a bright or busy background use the outline or the plate, because plain white text disappears precisely when you want it read.
Do captions work in Arabic?
Yes, properly, with four Arabic faces and correct right to left rendering, letters joined and lines breaking in the right direction. This is worth checking on any tool you consider, because a product claiming a hundred languages often means a hundred Latin ones and Arabic arrives reversed.
Can I add emoji to the captions?
Yes, and it renders as real colour emoji in the exported video rather than as a monochrome outline or an empty box, which is the usual failure. There is a pass that marks at most one per caption line, on the strongest word in that line. The restraint is the point: an emoji on every other word stops reading as emphasis and starts reading as noise, which is what makes automatic emoji look cheap on other people’s videos.
Can it remove filler words?
From the captions, yes: ums, uhs and the like, plus "you know" and "I mean". The audio still says them, and that difference is worth knowing before you choose a tool on this feature, because Descript and quso.ai cut fillers out of the audio itself. Words such as "like" and "actually" are left alone here on purpose, since they are usually real words.
Can I caption a video that is not in English?
Yes. You can also translate the captions into a different language while the audio stays as it was, which is how you reach an audience that does not share your language without redubbing anything. A translated line is timed as a whole line, so the highlight is close rather than exact.

Try it on one video

The first Short is free, needs no account, and takes about two minutes. That answers more than any page about a tool can, including this one.

Open the Studio

Keep going