Karaoke Captions is a caption generator that runs entirely on your own device. The speech recognition model is downloaded to your browser once and runs there, so your video and its audio are never uploaded to any server. Once the tool and model are loaded you can even switch off your internet connection and keep captioning.
Why “no upload” matters
- Client and unreleased work. Interviews, product launches, internal training and anything under NDA never leave your machine, so there's nothing to leak and no third party's retention policy to read.
- No waiting. A two-gigabyte file doesn't have to crawl up a hotel Wi-Fi connection before anything happens.
- No meter. Cloud caption tools pay for every minute they transcribe, which is why their free plans have minute caps and watermarks. Your device does this work, so there's no per-minute cost to pass on.
How it works
- Your browser reads the video file. Dropping a file gives the page access to it locally, the same way a photo editor opens an image. Nothing is sent anywhere.
- The audio is extracted on your device and converted to the 16 kHz mono format speech models expect.
- A speech model transcribes it, locally. The tool runs OpenAI's open-source Whisper model (the base size on computers, tiny on phones) through Transformers.js in a background worker. It uses your graphics card through WebGPU where the browser supports it, and your processor through WebAssembly otherwise.
- Every word gets its own timestamp, which is what makes word-by-word karaoke captions possible.
- The MP4 is encoded on your device too. Each frame is drawn with your captions and encoded with the browser's built-in video encoder (WebCodecs), then saved straight to your downloads.
The model files come from Hugging Face's public model hub the first time and are then cached by your browser, so later videos start straight away.
Check it yourself in two minutes
You shouldn't have to take a privacy claim on trust. Here are two ways to verify it with nothing but your browser:
1. Watch the network
- Open karaokecaptions.com in Chrome or Edge on a computer.
- Press F12 (or ⌥⌘I on a Mac) and open the Network tab.
- Drop in a video and let it caption.
- Sort the requests by size. You'll see the page's own files, the speech model downloading (only the first time) and a few small analytics pings. You won't see anything near the size of your video going out.
2. Pull the plug
- Caption one short video, so the tool and model are loaded and cached.
- Turn off Wi-Fi, or tick Offline in the Network tab.
- Drop in another video in the same tab. It transcribes and exports just the same, because nothing ever needed the internet.
What does leave your device: anonymous usage analytics (that a video was opened, its file type and size, that an export happened), never the file, its name or its words. Signing in is optional and only syncs saved caption styles. Details are in the privacy policy.
The trade-offs, honestly
- Accuracy. The models that fit in a browser are smaller than the largest cloud ones. Clear speech in English and other major languages comes out very well; heavy accents, crosstalk and loud music produce more mistakes. Every word is editable before export.
- Speed depends on your device. A recent laptop with WebGPU transcribes a short clip in seconds; an older phone takes longer. A computer is best for long videos.
- MP4 export needs Chrome or Edge. Safari and Firefox can transcribe and export SRT and VTT subtitle files, but burned-in MP4 export relies on WebCodecs features they don't fully provide yet.
- The first run downloads the model. Expect a one-time wait on a slow connection.
Private vs cloud caption tools
| Karaoke Captions | Typical cloud caption app | |
|---|---|---|
| Where your video goes | Nowhere. It stays on your device | Uploaded to the company's servers |
| Account | Not needed | Usually required |
| Free plan limits | None on the free styles | Minute caps, watermark or both |
| Works without internet | Yes, once loaded | No |
| Largest model size | Browser-sized | Server-sized |
For named comparisons see caption apps compared. Need only a subtitle file? Video to SRT uses the same on-device transcription.
Frequently asked questions
Is there a caption generator that doesn’t upload my video?
Yes. Karaoke Captions runs the speech model in your browser, so the video is processed on your device and never uploaded. You can confirm it in your browser’s Network tab.
Does it really work offline?
Yes, once the tool and its speech model have loaded (caption one video first). After that you can disconnect and keep transcribing and exporting in the same tab.
Which speech model does it use?
OpenAI’s open-source Whisper: the base model on computers and the tiny model on phones, run through Transformers.js with WebGPU or WebAssembly.
Is on-device transcription as accurate as cloud transcription?
For clear speech it is very good. The largest cloud models handle heavy accents, noise and crosstalk better. You can correct any word before exporting.
What languages does it support?
Whisper detects and transcribes many languages automatically. English is the most accurate.