Subtitle Generator

Generate SRT or VTT subtitles from video or audio, with speaker labels. The AI runs entirely in your browser, so your footage never leaves your device.

lock

Your audio stays inside this browser tab, never reaches a server, and nothing is saved to your computer. The first visit prepares the AI tools in your browser cache (about 80 MB, one-time). After that, each transcript runs instantly.

infoRunning in compatibility mode. Transcription will work but may take longer. For the fastest experience, try Chrome or Edge on a recent computer.

upload_file

Pick an audio or video file to transcribe.

lightbulbTip

Record from your mic, capture another browser tab, or upload an audio or video file. Click any "Speaker 1" label to rename it, and click any line's text to fix the wording before exporting. Clearing a line deletes it.

What is the Subtitle Generator?

A subtitle generator listens to the speech in a video and turns it into timed caption files. This one runs the Whisper speech-recognition model directly in your browser using WebGPU or WebAssembly: it reads the audio track of your MP4, MOV, or WebM, transcribes it with word-level timing, tells speakers apart automatically, and exports standard SRT or VTT files that YouTube, Premiere, DaVinci Resolve, VLC, and every major player accept. Because the model runs on your device, the footage itself is never uploaded anywhere.

How to use the Subtitle Generator

  1. 1

    Upload a video or audio file

    Pick an MP4, MOV, WebM, MP3, or WAV. The first visit downloads the AI model into your browser cache (about 80 MB, one time); after that it starts instantly.

  2. 2

    Let the AI transcribe

    The audio track is transcribed with word-level timestamps, and speakers are detected and labeled automatically. A progress transcript streams in as it works.

  3. 3

    Edit lines and name speakers

    Click any line to fix the wording, clear a line to delete it, and click a speaker label to replace "Speaker 1" with a real name.

  4. 4

    Export SRT or VTT

    Download the finished captions as SRT for video editors and YouTube, VTT for web players, or plain TXT for a readable transcript.

Frequently Asked Questions

Which video formats can it subtitle?

Anything your browser can play: MP4, MOV, and WebM cover nearly all real-world files. The tool reads the file's audio track, so video resolution and codecs don't matter as long as the audio decodes.

Is my video uploaded to a server?

No. The speech-recognition model runs inside your browser via WebGPU or WebAssembly. You can open DevTools and watch the Network tab while it transcribes: after the one-time model download, no data leaves your machine.

What's the difference between SRT and VTT?

SRT is the older, universally supported format that video editors and YouTube expect. VTT is the web-native format used by HTML5 video players and supports speaker voice tags. This tool exports both, with speaker names included.

Can I edit the subtitles before exporting?

Yes. Click any line's text to correct it, clear a line entirely to delete it, and rename detected speakers to real names. Your edits flow straight into the exported SRT and VTT files.

How accurate is it, and which languages work?

It uses Whisper, the same model family behind most commercial transcription tools, and works best on clear English speech. Heavy accents, crosstalk, and loud background music reduce accuracy. That's what the built-in editor is for.

Is there a length or file-size limit?

No hard cap. Processing time scales with the recording length and your hardware. A recent laptop with WebGPU transcribes many times faster than real time, and very long recordings simply take proportionally longer.

Related Tools