Choose your recording
Select a supported audio or video file. VideoShrink reads its audio track locally.
Private browser AI
Transcribe spoken video into editable, timestamped text without uploading the recording. Export a clean transcript or ready-to-use subtitle file.
Drop a file here, or select one from this device.
Three local steps
The media never enters an upload queue. Decoding and speech recognition happen on your device.
Select a supported audio or video file. VideoShrink reads its audio track locally.
Choose a spoken language or let the multilingual model detect it, then keep the tab open while it works.
Play from any timestamp, correct names or technical terms, and download TXT, SRT or VTT.
A different privacy model
Many transcription services upload media to a server. VideoShrink instead downloads a compact AI model and runs it against the decoded audio inside your browser.
A small local model protects privacy and avoids an account, but a larger cloud model may be faster or more accurate for long recordings, difficult audio, specialist vocabulary or speaker labeling.
Practical workflows
Export SRT or VTT cues for a video editor, course platform or accessible web player.
Use the TXT transcript as a searchable starting point, then rewrite and fact-check it.
Use timestamps to return to a quote, explanation or demonstration in the original recording.
Clear answers
No. VideoShrink reads the audio track and runs the speech-recognition model in your browser. The AI model is downloaded from Hugging Face on first use and cached by the browser, but your selected media is not sent with that model request.
The picker accepts common video and audio files including MP4, MOV, WebM, MKV, AVI, MP3, WAV, M4A, AAC, OGG and FLAC. Actual decoding depends on the audio codec supported by your browser, so MP4, WebM, MP3, M4A and WAV are the most dependable choices.
Accuracy varies with speech clarity, language, accents, names, background noise and the recording quality. This browser version uses a small multilingual Whisper model for practical local transcription, so review the result before publishing or quoting it.
Files must be smaller than 250 MB. To keep memory use practical, the current limit is 15 minutes on desktop browsers and 5 minutes on mobile devices.
Yes. Edit the timestamped segments and download them as SRT or WebVTT captions. You can also export the combined transcript as a UTF-8 TXT file.
Not in this local version. It creates timestamped speech segments but does not label or distinguish speakers. Speaker diarization requires a separate model and is not claimed by this tool.