ScribeFit uses OpenAI's Whisper speech model. Press Download the engine and pick one. Standard is about 76 MB. High accuracy is about 240 MB and is part of Pro. The engine is saved in your browser's own storage, so you only do this once. After that, transcription works offline. Delete this engine removes it again.
Choose the language the recording is spoken in. About 70 are available. ScribeFit starts with your browser's language.
Drop an audio or video file — or click to choose one. MP3, M4A, WAV, FLAC, OGG, Opus, MP4, MOV and WebM all work, along with anything else your browser can decode.
The text appears as it is transcribed, with or without times. It runs roughly 3 to 10 times faster than real time, depending on your computer and the engine.
Press Copy, or download the text as a .txt file. With Pro you can also download subtitle files in .srt and .vtt format.
Everything happens in your browser. No upload, no account, no ads. The only thing ScribeFit ever downloads is the speech engine, and the request carries nothing about you or your recordings.
Yes — one recording at a time, up to 10 minutes long, is free, with the Standard engine and text output. Pro (one-time) unlocks recordings of any length, the High-accuracy engine, SRT and VTT subtitle files, and several recordings at once.
No. Transcription runs on your own device. The only download is the speech engine, once, when you press Download the engine. After that, ScribeFit works offline.
MP3, M4A, WAV, FLAC, OGG, Opus, MP4, MOV and WebM — anything your browser can decode. About 70 spoken languages; you pick the one in the recording.
Roughly 3 to 10 times faster than real time, depending on your computer and the engine you pick.