Skip to content
helpyself

Transcribe Video

Local only — nothing is uploaded unless you save or share it.

Turn a video or audio recording into text, without uploading it anywhere

How to transcribe a video to text

  1. Drop in your video or audio file — it stays on your device
  2. Pick the spoken language, or let it work that out from the audio
  3. Press Transcribe and watch the text appear
  4. Copy it, or download it as TXT, Word, PDF, SRT or VTT

    Runs on your machine. Your recording is not uploaded.

    Share a link

    Saves this to your files and gives you a web address anyone can open. It stops working on its own.

    Link lasts
    Save this tool

    Keep Transcribe Video handy

    Bookmark it

    About this tool

    Drop in a video or an audio file and this writes down what is said in it. It works on the recordings people actually need transcribed — an interview, a lecture, a meeting, a podcast, a voice memo, a video you want captions or a searchable record for — and it handles more than twenty languages, including Danish, Dutch, German, Polish and Spanish. The unusual part is where it runs. The speech model is downloaded to your browser the first time you use it and everything after that happens on your own machine: your recording is never uploaded, never stored on a server, and never seen by anybody else. That matters more for this kind of file than for most — a recorded meeting or a patient interview is exactly the sort of thing you cannot paste into a website and hope. The first run downloads the model, which takes a moment. After that it is fast: on a laptop with a modern graphics chip a ten-minute video is transcribed in about a minute and a half, and even on an older machine with no GPU at all it finishes faster than the recording plays. You can watch the text appear as it goes, and stop early and keep what it found. When it is done you can copy the text, or save it as a plain TXT file, a Word document, a PDF, a Markdown file, a spreadsheet, or as SRT and VTT subtitles that a video player will read.

    Frequently asked questions

    Is my video uploaded anywhere?

    No. The speech model is downloaded to your browser and the transcription happens on your own machine — the recording is never sent to us or to anyone else. You can disconnect from the internet after the model has loaded and it will still work.

    Can it work out which language is spoken?

    Yes, and it asks the speech model itself — the same model that transcribes also identifies languages. On the Dutch recording we test with, it named Dutch with 96% confidence. The language it settled on appears in the dropdown as soon as it knows, so you can see it was right, or stop within seconds if it was not. If you already know the language, naming it yourself is still the surest route.

    How long does it take to transcribe a video?

    Less time than the recording lasts, on every machine we measured. A ten-minute video took about 100 seconds on a laptop with a graphics chip and about six and a half minutes on one without. The tool shows an estimate while it runs, and it gets more accurate as it goes because it measures your machine rather than guessing.

    What languages can it transcribe?

    Over twenty, including English, Danish, Dutch, German, French, Spanish, Italian, Polish, Finnish, Swedish, Norwegian, Portuguese, Russian, Ukrainian, Turkish, Arabic, Hindi, Japanese, Korean and Chinese. Choose one before you start, or let it work the language out from the audio — see below.

    Can it translate what it hears?

    Into English, yes — there is an option for it. The model can write down what it hears, or translate it into English, and it cannot do any other pair. Danish to German is not something it can do, so we do not offer it rather than offering it badly.

    What file formats does it accept?

    Anything your browser can play: MP4, MOV, WebM, MKV and AVI for video, and MP3, M4A, WAV, FLAC, OGG and Opus for audio. It reads the sound out of a video, so you do not need to extract the audio first.

    What can I save the transcript as?

    Plain text, Word (.docx), PDF, Markdown, CSV for a spreadsheet, JSON, and SRT or VTT subtitle files. The document formats can carry a timestamp against each paragraph, and the subtitle files carry the timings a video player needs.

    How accurate is it?

    Good, and honest about not being perfect. On a nine-minute Dutch interview it matched a desktop transcription tool nearly word for word; what it gets wrong is mostly proper nouns and names, which is what speech recognition is generally worst at. Read it through before you rely on it.

    Is there a length limit?

    Twenty minutes per file. That is a memory limit rather than a rule — decoded audio is held in your browser's memory, and past twenty minutes a tab is gambling with your machine. Split a long recording with our audio trimmer first.

    Why is the first run slower?

    It downloads the speech model. On a computer with a graphics chip that is about 920 MB, because only full precision transcribes correctly — we measured the smaller versions and one was ten times slower while the other quietly produced nonsense. Without a graphics chip it is about 240 MB. Your browser keeps it either way, so the second file starts immediately, and “Faster, rougher” uses a model around a sixth of the size.

    Where does the speech model come from?

    It is downloaded from Hugging Face, the public repository the model is published on. That is the one request this page makes to anybody else, and it goes the other way: files come down, nothing goes up. Your recording is never part of it — it is handed to the speech model inside your own browser and never touches the network. Everything else the tool needs, including the code that runs the model, is served from helpyself.com.

    Does it store anything on my computer?

    No cookies, and nothing from your recording or your transcript. The site keeps a few general preferences — your theme, whether you are signed in — and this tool adds one thing: the speech model itself, held in your browser's storage so the second file starts immediately. That is the 920 MB (or 240 MB without a graphics chip) described above, and it stays until you clear this site's data in your browser, which removes it completely.

    Why does my transcript have gaps in it?

    Because something in the recording was not speech. Music, applause, a held note or a room with a fan in it will send any speech model into a repeating loop — it writes the same phrase over and over, fluently and in the right language, and nothing about the output says it is invented. When that happens the passage is decoded again a few times with more randomness, and if it still loops it is left out and the summary says so. A gap you can see beats a paragraph of dialogue that was never spoken.

    Help improve this tool

    Found a bug, want a change, or need a tool we don't have?

    /transcribe-video

    ↑↓ to move · Enter to open · Esc to close