2026-10-0410 min readArticle

How to Add Captions to a Video for Free

Plenty of videos are watched with the sound off, in a noisy room, or by someone who simply follows text faster than speech. Captions fix all three problems, yet adding them still feels like a chore. This guide explains the two kinds of captions, how AI turns speech into timed text, the rules that make captions easy to read, and how to produce all of it with LuminaAudio’s free AI Caption Generator, which runs entirely in your browser.

Looking for the full list of guides? Visit the blog index →

Why captions are worth the effort

Think about how you watch video on your phone. On a train, in a queue, or next to someone who is asleep, the sound is off, and a video without text is a video you scroll past. Captions keep the message alive in those moments, which is why creators who publish short clips treat them as part of the edit rather than an extra.

There is an accessibility side as well. Captions are what make spoken content usable for people who are deaf or hard of hearing, and they also help viewers watching in a second language or dealing with a strong accent. The Web Content Accessibility Guidelines (WCAG) list captions for pre-recorded video with audio as a Level A requirement, so if you publish for a school, a business, or a public organisation, captions are not optional.

Finally, text is something computers can read and video is not. A caption file or transcript gives a platform exact words to index, which makes your video easier to find for the phrases people actually search. It will not guarantee rankings, but it gives search systems far more to work with than a title and a description.

Captions, subtitles and transcripts: what is the difference?

The three terms get mixed up constantly. Captions are timed text of what is said, and often include meaningful sounds such as [applause]. Subtitles assume the viewer can hear the audio but does not understand the language, so they translate the speech. A transcript is plain text with no timing at all, which makes it ideal for show notes, blog posts and quoting.

In everyday use, “subtitles” is the word most people type into a search box, even when they mean captions. The LuminaAudio AI Caption Generator covers the first and third ideas: it produces timed captions in English and a plain-text transcript. It does not translate between languages.

Sidecar files or burned-in captions: pick the right one

There are two ways to attach captions to a video, and choosing the wrong one is the most common mistake beginners make.

Sidecar files (SRT and VTT) are small text files that travel next to the video. Each entry holds a start time, an end time and some words. The viewer can switch them on or off, change the size on some players, and the platform can read them. This is the right choice for YouTube, Vimeo, course platforms and anything where accessibility and searchability matter. SRT is the older, near-universal format, while VTT is the web-native version and uses a full stop instead of a comma in its timestamps.

Burned-in captions (also called open captions) are drawn onto the pixels of the video itself. They cannot be turned off, but they also cannot go missing, and they can be styled with colour, outlines and motion. This is the right choice for TikTok, Instagram Reels and YouTube Shorts, where captions are part of the visual identity and many viewers never open a menu.

A sensible rule: publish a sidecar file wherever the platform accepts one, and burn captions in wherever it does not. For a long YouTube video that you also cut into Shorts, you will use both.

How AI turns speech into timed captions

Automatic captioning has two jobs. First it recognises the words, then it works out when each word starts and ends. The second part matters more than people expect: a perfectly spelled caption that appears half a second late feels broken, while a slightly misspelled one on time barely registers.

The LuminaAudio AI Caption Generator uses Whisper, an open speech recognition model, in its English-only form. The model is downloaded once to your browser (about 80 MB) and then runs on your own device through WebAssembly. Your audio is decoded locally, and no file is sent to a server for transcription. That is a real difference from most free caption sites, which upload your video to process it, then pay for that processing with watermarks, minute limits or upgrade prompts.

Because the work happens on your hardware, speed depends on your device and on how clean the recording is. To avoid a long wait on a long file, the tool transcribes in short windows and opens the editor as soon as the first stretch of speech is ready. The rest of the audio continues in the background, and you can already select, edit and style while it fills in.

The recogniser produces word-level timing, which is what enables effects like highlighting each word as it is spoken. It is accurate on clear speech, but it is still a machine. Names, brands, numbers and technical vocabulary are the usual trouble spots, so plan a short review pass every time.

What makes captions easy to read

Good caption styling is mostly restraint. Whatever tool you use, these rules hold up across platforms.

Keep each caption short. For short-form video, one to three words at a time feels energetic and keeps the eye in one spot. For talking-head or tutorial content, a full line of roughly 35 to 42 characters, with no more than two lines on screen, is a common convention borrowed from broadcast subtitling.

Give every caption enough time. A phrase should stay on screen for at least about a second, and longer if it has more words. If a caption flashes past, the viewer reads half of it and loses the thread.

Make contrast do the work. White or bright text with a dark outline or soft shadow stays readable over almost any background. Highlight colours are best used on one emphasised word, not on whole sentences.

Respect the safe zone. On vertical video the bottom of the frame is covered by usernames, descriptions and buttons, and the top is covered by tabs. Place captions in the lower-middle area, clear of faces and interface elements.

Choose one animation and stay with it. Motion draws attention, so using a different effect on every line turns captions into noise. One restrained style per video looks professional, and a stronger one is fine when the content itself is high-energy.

Meet the AI Caption Generator

LuminaAudio’s AI Caption Generator is a browser-based tool that takes an audio or video file and gives you editable, styled captions. There is nothing to install, no account to create and no watermark on the results. It accepts the common formats your browser can decode, including MP3, WAV, M4A, MP4, MOV and WebM.

It is designed around a simple promise: you should be able to go from a raw file to a finished caption file or a captioned video in one sitting, without moving between a transcription site, a text editor and a video editor.

Here is how a typical session goes. You drop in your file and the on-device model starts transcribing. When the first segment is ready the editor opens, showing your video, a zoomable timeline of caption blocks and a transcript you can edit. You choose a style, fix any words that were misheard, adjust timing where you want to, and export.

A tour of the features, and why each one exists

Rather than list features for their own sake, here is what each part of the editor is for.

Streaming transcription. Long recordings are the normal case for podcasts, lectures and interviews, so you are not asked to wait for 100 percent. You can start working on the first section while later sections are still being recognised.

A timeline you can actually edit. Caption blocks sit on a zoomable timeline. Drag the edges of a block to change when it appears and disappears. Double-click any caption to correct its spelling right where it sits, or use the Edit tab for a full list with timestamps. If you need a caption that was never spoken, such as a title card or a sound cue, select an empty stretch of the timeline and add a new caption there.

Select, delete and move. Drag across a range of captions to select them, then delete them, which leaves a clean gap, or move them to a different timeline. Moved captions keep their exact original duration. Undo and redo are available at every step, so experimenting is safe.

Multiple timelines. A second timeline can carry its own captions with its own colours, fonts and animation. That is useful when one line deserves special treatment, such as a call to action, a translated line you typed in yourself, or emphasis words that should look different from the main speech.

Styling and animation. More than 100 style presets, 85 fonts and 100 animations cover everything from a quiet lower-third to a loud karaoke sweep, with controls for colours, size, words per line, position and animation intensity. You can drag the captions anywhere on the frame, which matters for vertical video.

Exports. Download SRT or VTT for platforms that take sidecar files, TXT for a plain transcript, or a video with the captions burned in. The video export is limited to your original resolution rather than stretching a small clip into a bigger file, and you can choose a frame rate. It runs in real time, so a one-minute clip takes about a minute.

Step by step: caption a video from scratch

1. Prepare the audio. Clean speech is the single biggest factor in accuracy. If you can, trim long silences and reduce background music before you start. If the video is already recorded, you can still run it as it is.

2. Open the AI Caption Generator and drop in your file. The first time, the model downloads once (about 80 MB) and is then reused from your browser’s cache.

3. Wait for the editor to open. You will see your first captions within moments, even on a long file. Let the rest of the transcript continue in the background.

4. Review the words. Play through, pause on anything suspicious, and correct names, numbers and jargon by double-clicking the caption.

5. Pick a look. Start from a preset that fits your content, set how many words appear at once, then choose an animation and drag the captions to a safe position.

6. Fine-tune timing. Wherever a caption feels early or late, drag its edge on the timeline. Short, quick pops can stay tight, while longer phrases deserve more time on screen.

7. Export. Choose SRT or VTT for upload, TXT for notes, or the video with burned-in captions for social platforms. If you export video, keep the tab visible until it finishes.

Caption settings that work for common video types

TikTok, Reels and Shorts: one or two words per line, a bold preset, a high-contrast outline and a word-by-word animation. Put the captions in the lower-middle part of the frame and export a burned-in video.

Long YouTube videos: a plain, calm style or no burned-in captions at all. Upload the SRT or VTT in YouTube Studio so viewers can switch captions on, and so the platform gets accurate text to index.

Podcasts and interviews: use the TXT export for show notes and an SRT for the video version. The tool does not label speakers automatically, so add names such as “ANA:” at the start of lines when it helps listeners follow who is talking.

Tutorials and courses: two readable lines, 35 to 42 characters each, with a minimal style. Prioritise correct terminology, because a wrong word in a technical lesson is more confusing than a missing one.

Limits to know before you start

No tool is perfect, and it is better to know the edges in advance. The AI Caption Generator currently transcribes English only. It does not detect or translate other languages, and it does not identify different speakers. Fast speech, heavy background noise, strong accents and specialist vocabulary all reduce accuracy, so always review before you publish.

Video export records in real time inside the browser, so it needs the tab to stay open and visible, and it works best in desktop Chrome or Edge. Very long files need a device with enough memory and, ideally, a charger. Text exports such as SRT, VTT and TXT are instant.

These are the trade-offs of keeping everything private and free. If you need certified transcripts, many languages or a team review workflow, a paid human captioning service is the better fit. For fast, private English captions with strong styling control, an on-device editor is hard to beat.

A quick proofreading checklist

Before you export, spend two minutes on the things AI gets wrong most often: people’s names and brand names, numbers, dates and prices, homophones such as “their” and “there”, abbreviations and technical terms, and punctuation that changes meaning. Read the captions with the sound off once. If the video still makes sense, your captions are doing their job.

Where to go from here

Captions are one of the highest-return habits in video publishing: they widen your audience, keep sound-off viewers watching, and give platforms real text to index. Start with a short clip, run it through the AI Caption Generator, and compare the result with the rules above. After one or two videos, the workflow will take minutes.

If you also work with sound, LuminaAudio’s other free tools can help before you caption: trim dead air with the audio trimmer, convert a file with the audio converter, or shape a track with the slowed and reverb generator.

Try the Tool (Free)

Use LuminaAudio's free browser-based tools to process audio online. Experiment with speed, reverb, pitch, bass, sleep frequencies, and creative song transformations directly in your browser.

AI Caption Generator

Frequently Asked Questions

Should I upload an SRT file or burn captions into the video?

Upload an SRT or VTT file wherever the platform supports it, such as YouTube or Vimeo, because viewers can toggle it and the platform can index the text. Burn captions in for TikTok, Instagram Reels and Shorts, where styled, always-visible text is part of the format.

What is the difference between captions and subtitles?

Captions show what is said in the same language and may include sound cues. Subtitles translate speech for viewers who can hear it but do not understand the language. A transcript is untimed text of the same content.

Is the LuminaAudio AI Caption Generator really free?

Yes. There is no account, no watermark and no minute cap, because transcription runs on your own device instead of on a paid server. Your device’s memory and speed are the only practical limits.

Is my video uploaded when I generate captions?

No. The speech recognition model is downloaded to your browser once and runs locally, so your audio and video are processed on your device. You only need a connection for that first model download and for loading fonts.

How long should a caption stay on screen?

Aim for at least about one second for even the shortest phrase, and give longer lines more time so a viewer can finish reading before the next one appears. Quick one-word pops are the exception in fast short-form edits.

Do captions help a video get found in search?

They can help. Search systems cannot watch video, but they can read caption files and transcripts. Accurate text makes it easier for a video to match what people search for, though captions alone do not guarantee any ranking.

Can the caption generator handle languages other than English?

Not at the moment. It is built for English speech only and does not translate. For another language, use a multilingual transcription tool, or type the lines yourself in the editor and style them like any other caption.

How accurate are AI-generated captions?

Clear, single-speaker English with a decent microphone usually comes out well, while noise, accents, fast speech and jargon lower accuracy. Treat the result as a strong first draft and proofread names, numbers and technical terms before publishing.