Skip to content
toolsdocks

Generate subtitles from speech

Choose a Whisper model, add a video or audio file and generate timed captions. After a one-time model download from Hugging Face, the audio is processed in your browser and not uploaded.

On-device model
Loading tool…

How to use

  1. Add a video or audio file.
  2. Choose the model, the spoken language (or auto-detect) and the line length, then download the model (once) and generate.
  3. Review the cues, download SRT or VTT, open them in the subtitle editor, or burn them into the video.

Worked example

A 10-second, 117-character sentence is split into several cues, each at most two lines of 42 characters and 6 seconds on screen, timed in proportion to their text.

Supported formats and limits

InputMP4, MOV, WebM, MKV, MP3, WAV, M4A, Ogg, FLAC
OutputSRT, VTT, ASS, SBV, TTML, JSON, TXT
LimitsFirst use downloads the chosen Whisper model: about 44 MB (Tiny, English), 80 MB (Base) or 250 MB (Small) on the processor; WebGPU uses larger files. It is cached for later visits. Roughly real-time or slower without WebGPU; up to about 2 hours of audio.
EngineWhisper (transformers.js, ONNX Runtime Web) in a Web Worker, in 30-second sections cut at quiet points; segments are re-split into subtitle-sized cues

Limitations

  • Accuracy drops with music, overlapping speakers and heavy accents.
  • Model output can contain mistakes (names, numbers, technical words); review before publishing.
  • Cue timing comes from Whisper's segment timestamps, so a cue can start or end a little early or late.

Questions

How accurate are the subtitles?

Clear speech usually transcribes well, but accuracy drops with music, overlapping speakers and heavy accents, and names, numbers or technical words can be wrong. Cue times come from Whisper's segments and can be slightly early or late, so review before publishing.

Can it make English subtitles from another language?

Yes, with the multilingual models: set Subtitles in to English (translated). The Tiny English model only transcribes English.

How are long sentences split into cues?

Segments are re-split to your characters per line and lines per cue (for example two lines of 42 characters), at most 6 seconds each, timed in proportion to their text.

Guides

Privacy

On-device model. Processing runs in this browser. The open-source model files are downloaded once from the model host (Hugging Face) and cached; your content is not uploaded.

  • Hugging Face: Downloads of the open-source Whisper model files when you ask for them. No audio, video or text is sent.

See the privacy policy for how toolsdocks handles data.