Free Web Tool

Free Audio to Text

Transcribe any audio using Whisper AI — runs entirely in your browser. Private, free, no sign-up, no upload to servers.

🎵
Drop audio here or browse
MP3 · WAV · M4A · OGG · FLAC · WebM · max 25 MB
Language
🤖 Uses Whisper Tiny AI model (~40 MB, downloaded once and cached). All processing runs in your browser — your audio never leaves your device.

How to use

1
Upload an audio file (MP3, WAV, M4A…) or record directly from your mic.
2
Select a language or leave on Auto-detect. Then click Transcribe.
3
The first run downloads the Whisper AI model (~40 MB). After that it's instant — even offline.
4
Copy or download the transcript. Nothing is sent to any server.

Frequently Asked Questions

Is this free?
Yes — completely free, no sign-up, no limits.
Is my audio private?
Yes. Your audio never leaves your device. The Whisper AI model runs entirely inside your browser using WebAssembly — no server, no tracking.
What audio formats are supported?
Any format your browser can decode: MP3, WAV, M4A, OGG, FLAC, WebM. If it plays in your browser, it can be transcribed.
What languages does it support?
Whisper supports 99 languages including English, Chinese, Japanese, Korean, Spanish, French, German, Portuguese, Indonesian, Arabic, Russian and more. Use Auto-detect or pick a language for better accuracy.
Why does it take time on first use?
The Whisper AI model (~40 MB) is downloaded once from a CDN and cached by your browser. After the first use, transcription starts immediately — even offline.
Is there a file size limit?
25 MB per file. For longer recordings, consider splitting the audio first. Processing time depends on your device — modern laptops handle 5 minutes of audio in about 30 seconds.
Also try
🔊
Free Text to Speech
Convert text to natural speech using your device's voices. Free, no sign-up.
🎬
YouTube Caption Downloader
Download subtitles as SRT or TXT from any YouTube video. Free, no sign-up.

常見使用場景

Whisper 是目前最強開源轉錄模型。歷史上意味著 Python、GPU 和命令列——對一般使用者完全不可及。瀏覽器版 Whisper 讓流程變成:拖入音訊,轉錄,下載 SRT 或 TXT。零安裝,零帳號,零程式碼。

隱私是關鍵。客戶端 Whisper 意味著音訊永遠不離開裝置。採訪錄音、醫療口述、法律證詞、機密會議——任何不能上傳給第三方 API 的音訊,現在都可以安全轉錄。

使用技巧

模型大小的權衡:tiny/base 最快但精度一般,適合清晰音訊;small/medium 是實用甜點,處理口音、雜訊、術語都還不錯;large 精度最高但慢——每分鐘音訊約 30 秒計算時間。

乾淨音訊比大模型更重要。tiny 模型跑乾淨錄音的效果比 large 跑混濁音訊還好。有錄音控制時用好話筒、控制背景雜訊、近講。極長檔案(2+ 小時)在瀏覽器記憶體裡可能受限——用音訊軟體切成 30 分鐘塊分別轉錄。

常見問題

瀏覽器 Whisper 真的和桌面版一樣準嗎?

是的——模型權重完全相同。區別只是計算速度,取決於你的裝置而非模型品質。

支援哪些音訊格式?

MP3、WAV、M4A、MP4、WebM、OGG、FLAC——任何瀏覽器能透過 Web Audio API 解碼的格式。影片檔案也支援,會提取音訊軌道。

能離線運行嗎?

能——模型下載一次後(快取在瀏覽器裡),轉錄完全離線運行。適合飛機上或離線裝置。

支援哪些語言?

Whisper 支援 99 種語言。對未知音訊用自動偵測;已知語言可以明確設定獲得更快更準的結果。

我的音訊檔案會上傳嗎?

不會。音訊完全在瀏覽器裡透過 WebAssembly 解碼和轉錄,沒有上傳、儲存或日誌。

相關免費工具

Built in the quiet hours, between everything else. Bookmark it for next time — and if it helped — Buy me a coffee
GistAI
GistAI App YT captions & AI summary Free on Google Play
If it saved you time — Buy me a coffee