Add subtitles to a video automatically

Vidmini listens to your video or audio file and writes the subtitles for you, using the open-source Whisper speech recognition model running on your own device. Review the text, then download an SRT or VTT file or burn the subtitles into the video.

No uploadNo watermarkNo signup

How it works

  1. 1

    Choose or drop a video or audio file (MP4, MOV, WebM, MKV, MP3, M4A, WAV, etc.), pick the spoken language or leave it on automatic detection, and choose a model.

  2. 2

    Click Auto Subtitles and wait for the transcription. The first time, the speech recognition model is downloaded once from vidmini.com.

  3. 3

    Read through the lines, fix or delete what is wrong, then download SRT or WebVTT, or burn the subtitles into the video.

Speech recognition that runs on your device

Most auto-caption services upload your file to a server and transcribe it there. Vidmini runs OpenAI's Whisper model, which is open source, directly in your browser through WebAssembly. Your audio is never sent anywhere: the only thing downloaded is the model itself, from vidmini.com, and your browser keeps it in its cache so the next videos start right away.

You can choose between two models. "More accurate" uses Whisper base, a download of about 102 MB including the runtime. "Faster, lighter" uses Whisper tiny, about 68 MB, which finishes sooner but makes more mistakes. On a typical laptop the base model transcribes at roughly the length of the video or a bit slower; phones take longer. The speed depends on your device, not on a queue.

Languages and accuracy

Leave the spoken language on automatic detection or pick it from the list, which includes English, French, Portuguese, Spanish, German, Italian, Dutch, Polish, Turkish, Russian, Ukrainian, Arabic, Hindi, Japanese, Korean and Chinese. Whisper knows about 100 languages, and automatic detection covers the ones not in the list. Picking the language yourself avoids a wrong guess on short clips. The tool transcribes what is said in the original language; it does not translate.

Recognition is good for clear speech: a voice-over, an interview, a talk to camera. It is weaker when music plays under the voice, with background noise, strong accents, several people talking at once, and names or brand words it has never heard. That is why Vidmini shows every line in an editor before you export anything.

Review and fix every line

After transcription, the subtitles appear as a list with their start time. Click a line to correct a word, a name or the punctuation, or delete a line that should not be there, such as a lyric picked up from background music. Lines are already split for reading: at most about two lines of around 42 characters per subtitle, the usual convention for TV and streaming.

Export SRT or WebVTT, or burn the subtitles in

SRT is the subtitle file almost every video editor, player and platform accepts, including YouTube, Premiere Pro, DaVinci Resolve and VLC. WebVTT (.vtt) is the format for web players and the HTML video element. Both are small text files you can still edit later, and viewers can turn them on or off.

If you want the text to be part of the image, burn the subtitles into the video: white text with a dark outline, in small, medium or large size, at the bottom or the top of the frame. This re-encodes the video, and the subtitles can no longer be switched off. For audio files, only SRT and WebVTT are available.

Accessibility and social videos

Subtitles make a video usable for deaf and hard-of-hearing viewers, for people watching in a second language and for anyone in a quiet place. On social networks, many videos autoplay muted in the feed: burned-in subtitles let people follow a Reel, a TikTok or a Short without turning the sound on.

For a cleaner result, cut what you don't need before transcribing. Remove Silence shortens long pauses, and Trim Video keeps only the part you want to publish. Files up to 3 hours long are accepted.

Frequently asked questions

Is my video uploaded to generate the subtitles?

No. The Whisper model is downloaded to your browser and runs on your device. The audio is decoded and transcribed locally and never leaves your computer or phone. Our Content Security Policy only lets the page talk to vidmini.com for its own code and the model files.

How accurate are the automatic subtitles?

With clear speech and little background sound, most lines come out right. Music, noise, accents, overlapping voices and proper names cause more errors, and the lighter model makes more mistakes than the more accurate one. Always read through the editor before exporting; fixing a few words takes a moment.

Can it translate subtitles into another language?

No. Vidmini transcribes the speech in the language it is spoken in. If you need a translation, export the SRT file and translate the text with the tool of your choice; the timings stay valid.

Which file should I choose: SRT, WebVTT or burned in?

Choose SRT for YouTube, video editors and most players. Choose WebVTT for a website video player. Burn the subtitles in when the platform has no subtitle upload or when you want them always visible, as on muted social feeds. Burning in re-encodes the video.

Why does the first use take longer?

The first time, your browser downloads the speech recognition model once: about 102 MB for "More accurate" or about 68 MB for "Faster, lighter". It is then cached by the browser, so later videos start without that download.

Is there a length limit?

Files up to 3 hours are accepted. Transcription time grows with the length and depends on your device: a laptop is much faster than a phone. For long recordings, trimming or removing silences first saves time.

Is it free, and is there a watermark?

It is free, with no account and no watermark. The subtitles you export contain only your text.