On-device AI · No upload · Verify it yourself

A private caption generator that never uploads your video

The speech model runs inside your browser, on your own computer or phone. Your video and its audio never touch a server, and once it’s loaded it works with the internet off.

Caption a video privately → Free · no account · nothing uploaded

Karaoke Captions is a caption generator that runs entirely on your own device. The speech recognition model is downloaded to your browser once and runs there, so your video and its audio are never uploaded to any server. Once the tool and model are loaded you can even switch off your internet connection and keep captioning.

Why “no upload” matters

How it works

  1. Your browser reads the video file. Dropping a file gives the page access to it locally, the same way a photo editor opens an image. Nothing is sent anywhere.
  2. The audio is extracted on your device and converted to the 16 kHz mono format speech models expect.
  3. A speech model transcribes it, locally. The tool runs OpenAI's open-source Whisper model (the base size on computers, tiny on phones) through Transformers.js in a background worker. It uses your graphics card through WebGPU where the browser supports it, and your processor through WebAssembly otherwise.
  4. Every word gets its own timestamp, which is what makes word-by-word karaoke captions possible.
  5. The MP4 is encoded on your device too. Each frame is drawn with your captions and encoded with the browser's built-in video encoder (WebCodecs), then saved straight to your downloads.

The model files come from Hugging Face's public model hub the first time and are then cached by your browser, so later videos start straight away.

Check it yourself in two minutes

You shouldn't have to take a privacy claim on trust. Here are two ways to verify it with nothing but your browser:

1. Watch the network

  1. Open karaokecaptions.com in Chrome or Edge on a computer.
  2. Press F12 (or ⌥⌘I on a Mac) and open the Network tab.
  3. Drop in a video and let it caption.
  4. Sort the requests by size. You'll see the page's own files, the speech model downloading (only the first time) and a few small analytics pings. You won't see anything near the size of your video going out.

2. Pull the plug

  1. Caption one short video, so the tool and model are loaded and cached.
  2. Turn off Wi-Fi, or tick Offline in the Network tab.
  3. Drop in another video in the same tab. It transcribes and exports just the same, because nothing ever needed the internet.

What does leave your device: anonymous usage analytics (that a video was opened, its file type and size, that an export happened), never the file, its name or its words. Signing in is optional and only syncs saved caption styles. Details are in the privacy policy.

The trade-offs, honestly

Private vs cloud caption tools

Karaoke CaptionsTypical cloud caption app
Where your video goesNowhere. It stays on your deviceUploaded to the company's servers
AccountNot neededUsually required
Free plan limitsNone on the free stylesMinute caps, watermark or both
Works without internetYes, once loadedNo
Largest model sizeBrowser-sizedServer-sized

For named comparisons see caption apps compared. Need only a subtitle file? Video to SRT uses the same on-device transcription.

Frequently asked questions

Is there a caption generator that doesn’t upload my video?

Yes. Karaoke Captions runs the speech model in your browser, so the video is processed on your device and never uploaded. You can confirm it in your browser’s Network tab.

Does it really work offline?

Yes, once the tool and its speech model have loaded (caption one video first). After that you can disconnect and keep transcribing and exporting in the same tab.

Which speech model does it use?

OpenAI’s open-source Whisper: the base model on computers and the tiny model on phones, run through Transformers.js with WebGPU or WebAssembly.

Is on-device transcription as accurate as cloud transcription?

For clear speech it is very good. The largest cloud models handle heavy accents, noise and crosstalk better. You can correct any word before exporting.

What languages does it support?

Whisper detects and transcribes many languages automatically. English is the most accurate.

Keep your footage yours

Captions and subtitles, made on your own device.

Open the free tool →