Choose a local video or audio file
Pick a local MP4, MOV, WebM, MP3, WAV, or similar file.
Create subtitles on your device, edit every line, and style them in a live video preview. Free, with no upload, no sign-up, and no added watermark.
Pick a local MP4, MOV, WebM, MP3, WAV, or similar file.
When the speech model is available, it runs locally to create timed subtitle lines. Model download, browser support and device resources affect availability. You can also import SRT or VTT files.
Correct text and timing in the editor — speech recognition is good, but your ear is better.
Download SRT/VTT subtitle files or export MP4 with styled subtitles in a supported browser.
When the speech model is available, it runs locally to create timed subtitle lines. Model download, browser support and device resources affect availability. You can also import SRT or VTT files.
Audio decoding and speech recognition happen on your device. Your video, audio, and words never reach a server.
Fix wording, adjust start and end times, add or delete lines before you export anything.
Download SRT/VTT subtitle files or export MP4 with styled subtitles in a supported browser.
Open a local MP4, MOV, WebM or audio file and generate subtitles with browser-based speech recognition. Your media and transcript are not uploaded to DojoClip. The first transcription downloads an open-source speech model; the browser can cache it for later sessions. Local processing does not mean that the first visit works offline.
Export MP4 to burn your styled subtitles into the picture, so they stay visible wherever the video is played. Download SRT when you need an editable subtitle file for YouTube or another video editor, or VTT for web video players. SRT and VTT contain text and timing; the fonts, colors and animated styles shown in the preview are part of the MP4 export.
Search for a line, jump to its moment in the video, correct the words, and adjust start and end times. Merge short lines or add a missing caption. The transcript scrolls independently while the video stays in view. Choose a caption preset, preview fonts, and move the subtitle block to keep faces and important details clear. Animated word highlights use estimated timing and should be reviewed.
Clear speech gives you the best starting point; review names, punctuation, music-heavy passages and overlapping speakers. Automatic transcription accepts files up to 700 MB, with a 10-minute limit on low-memory devices and up to 30 minutes on other devices. Browser decoding, memory and encoding support still apply. MP4 export works in supported browsers such as current Chrome or Edge. You can import an existing SRT or VTT without downloading the speech model.
Shorten your recording with the free video trimmer, or use the audio extractor for an audio-only workflow. For a complete timeline, multiple clips, translation and assistant-guided editing, open Video Edit AI.
Yes. This tool is free, with no account or added DojoClip watermark. Processing happens on your device; browser support and available memory still limit the files you can use.
No. The audio is decoded and transcribed locally on your device. Nothing is sent to DojoClip — which matters for unreleased or confidential footage.
The model understands 90+ languages, including English, Spanish, French, German, Italian, Portuguese, Korean, Chinese, and Japanese. Use auto-detect or pick the spoken language for best results.
Clear speech transcribes very well. Heavy accents, background music, or crosstalk will need a few manual fixes — that is exactly what the built-in editor is for.
Yes. Choose a video, generate or import subtitles, select a style and choose Export MP4. The captions become part of the video image. MP4 export requires browser encoding support; use a current Chrome or Edge. You can also download SRT/VTT separately.
Speed depends on your device, browser, available acceleration and the selected speech model. The first run also needs a model download. Long files use more memory; there is no fixed speed guarantee.