Free Web Tool

Free Audio to Text

Transcribe any audio using Whisper AI — runs entirely in your browser. Private, free, no sign-up, no upload to servers.

🎵
Drop audio here or browse
MP3 · WAV · M4A · OGG · FLAC · WebM · max 25 MB
Language
🤖 Uses Whisper Tiny AI model (~40 MB, downloaded once and cached). All processing runs in your browser — your audio never leaves your device.

How to use

1
Upload an audio file (MP3, WAV, M4A…) or record directly from your mic.
2
Select a language or leave on Auto-detect. Then click Transcribe.
3
The first run downloads the Whisper AI model (~40 MB). After that it's instant — even offline.
4
Copy or download the transcript. Nothing is sent to any server.

Frequently Asked Questions

Is this free?
Yes — completely free, no sign-up, no limits.
Is my audio private?
Yes. Your audio never leaves your device. The Whisper AI model runs entirely inside your browser using WebAssembly — no server, no tracking.
What audio formats are supported?
Any format your browser can decode: MP3, WAV, M4A, OGG, FLAC, WebM. If it plays in your browser, it can be transcribed.
What languages does it support?
Whisper supports 99 languages including English, Chinese, Japanese, Korean, Spanish, French, German, Portuguese, Indonesian, Arabic, Russian and more. Use Auto-detect or pick a language for better accuracy.
Why does it take time on first use?
The Whisper AI model (~40 MB) is downloaded once from a CDN and cached by your browser. After the first use, transcription starts immediately — even offline.
Is there a file size limit?
25 MB per file. For longer recordings, consider splitting the audio first. Processing time depends on your device — modern laptops handle 5 minutes of audio in about 30 seconds.
Also try
🔊
Free Text to Speech
Convert text to natural speech using your device's voices. Free, no sign-up.
🎬
YouTube Caption Downloader
Download subtitles as SRT or TXT from any YouTube video. Free, no sign-up.

일반적인 사용 사례

Whisper는 현재 최강의 오픈소스 전사 모델입니다. 과거에는 Python, GPU, 커맨드라인이 필요해 일반 사용자에게는 접근 불가능했습니다. 브라우저 Whisper는 오디오 드롭 → 전사 → SRT 또는 TXT 다운로드로 바꿉니다 — 설치 없이, 계정 없이, 코드 없이.

프라이버시가 결정적입니다. 클라이언트 사이드 Whisper는 오디오가 절대 기기 밖으로 나가지 않는다는 뜻. 인터뷰 녹음, 의료 구술, 법적 증언, 기밀 회의 — 제3자 API에 올릴 수 없는 오디오도 안전하게 전사 가능.

사용 팁

모델 크기 절충: tiny/base는 가장 빠르지만 정확도 보통, 깨끗한 오디오용. small/medium은 실용적 스위트 스팟, 억양·잡음·전문 용어를 무난히 처리. large는 최고 정확도지만 느림 — 오디오 1분당 약 30초 연산 시간.

깨끗한 오디오가 큰 모델보다 중요합니다. tiny 모델도 깨끗한 녹음이면 large가 탁한 오디오를 처리하는 것보다 결과가 낫습니다. 녹음을 제어할 수 있다면 좋은 마이크, 배경 소음 억제, 근접 마이크 사용. 초장시간 파일(2+시간)은 브라우저 메모리에 걸릴 수 있음 — 오디오 편집기로 30분 청크로 나누고 각각 전사.

자주 묻는 질문

브라우저 Whisper가 데스크톱 버전만큼 정확한가요?

네 — 모델 가중치가 완전히 동일합니다. 차이는 계산 속도뿐이며 모델 품질이 아닌 기기에 달려있습니다.

지원하는 오디오 형식은?

MP3, WAV, M4A, MP4, WebM, OGG, FLAC — 브라우저가 Web Audio API로 디코딩할 수 있는 모든 것. 비디오 파일도 지원, 오디오 트랙 추출.

오프라인으로 작동하나요?

작동합니다 — 모델을 한 번 다운로드(브라우저에 캐시)하면 전사는 완전 오프라인. 비행기 안이나 에어갭 기기에서 유용.

지원 언어는?

Whisper는 99개 언어 지원. 알 수 없는 오디오에는 자동 감지; 알려진 언어는 명시하면 약간 더 빠르고 정확한 결과.

제 오디오 파일이 업로드되나요?

아니요. 오디오는 브라우저 내부에서 WebAssembly로 디코딩·전사됩니다. 업로드, 저장, 로그 없음.

관련 무료 도구

Built in the quiet hours, between everything else. Bookmark it for next time — and if it helped — Buy me a coffee
GistAI
GistAI App YT captions & AI summary Free on Google Play
If it saved you time — Buy me a coffee