Bilingual captions for X videos, with your own keys
Most videos on X have no captions. The extensions that add them lock AI caption generation behind a subscription, and the ones that let you bring your own key only apply it to translation; the speech recognition still goes through their servers. I wanted the other arrangement: both layers on my own keys, nothing proxied, and the original language kept on screen.
Acousmos Captions is that extension. MIT licensed, zero runtime dependencies, every provider integration is a plain fetch.
Keep the original line
If you follow a language at all, you usually understand part of what is said and lose the rest to speed. A translation-only caption throws away the part you had. Captions keeps the original line as the anchor and puts the translation under it. When the translation is off, the original is right there.
Your keys, both layers
Speech recognition runs on Deepgram (nova-3) or Soniox (stt-async-v5). Translation runs on any OpenAI-compatible endpoint (the default, gpt-4.1-mini), Google Gemini or Anthropic. All keys stay in chrome.storage.local, are never synced, and are sent only to the provider you picked. An X video is short; transcribing one costs a fraction of a cent on your own key.
The model never sees the timeline
This is the part I care about most. Timestamps come from the speech recognizer, word by word. The extension groups words into sentence-sized cues, merging across recognizer boundaries when a sentence was split and splitting only when a cue runs past 140 characters or 12 seconds. Every cue boundary is a recognizer timestamp.
The translation model then receives (id, text) pairs, a few lines of preceding context and your glossary, and never a timestamp. It is told to return exactly one item per id, in order; it cannot merge, split, reorder or re-time anything. A reply whose ids or count do not match is not trusted: the batch is split in half and retried, down to single lines, and only a line that still fails is left untranslated. Translations are patched back by id onto the original cues. Whatever the model does, the subtitles stay in sync with the audio.
How a video gets captioned
If the video ships its own subtitle track, that track is translated directly and no audio is fetched. Otherwise the audio is captured inside your logged-in browser session, the same requests the player itself makes: the audio-only HLS rendition when there is one, falling back to the smallest muxed variant or the progressive mp4. It is assembled and sent to the recognizer. Original captions appear the moment recognition finishes; translated lines fill in batch by batch. Results are cached locally, so re-watching is free, and the bilingual, original or translated track can be exported as SRT.
Installing
Until the Chrome Web Store listing is live: download a release zip, unzip it, open chrome://extensions, turn on Developer mode, and load the folder that contains manifest.json. Add one recognition key and one translation key in the settings page, open any X video, and press the CC pill on the player.
The source, the privacy notes and a Chinese README are in the repository.