Free Web Tool

Free Audio to Text

Transcribe any audio using Whisper AI — runs entirely in your browser. Private, free, no sign-up, no upload to servers.

🎵
Drop audio here or browse
MP3 · WAV · M4A · OGG · FLAC · WebM · max 25 MB
Language
🤖 Uses Whisper Tiny AI model (~40 MB, downloaded once and cached). All processing runs in your browser — your audio never leaves your device.

How to use

1
Upload an audio file (MP3, WAV, M4A…) or record directly from your mic.
2
Select a language or leave on Auto-detect. Then click Transcribe.
3
The first run downloads the Whisper AI model (~40 MB). After that it's instant — even offline.
4
Copy or download the transcript. Nothing is sent to any server.

Frequently Asked Questions

Is this free?
Yes — completely free, no sign-up, no limits.
Is my audio private?
Yes. Your audio never leaves your device. The Whisper AI model runs entirely inside your browser using WebAssembly — no server, no tracking.
What audio formats are supported?
Any format your browser can decode: MP3, WAV, M4A, OGG, FLAC, WebM. If it plays in your browser, it can be transcribed.
What languages does it support?
Whisper supports 99 languages including English, Chinese, Japanese, Korean, Spanish, French, German, Portuguese, Indonesian, Arabic, Russian and more. Use Auto-detect or pick a language for better accuracy.
Why does it take time on first use?
The Whisper AI model (~40 MB) is downloaded once from a CDN and cached by your browser. After the first use, transcription starts immediately — even offline.
Is there a file size limit?
25 MB per file. For longer recordings, consider splitting the audio first. Processing time depends on your device — modern laptops handle 5 minutes of audio in about 30 seconds.
Also try
🔊
Free Text to Speech
Convert text to natural speech using your device's voices. Free, no sign-up.
🎬
YouTube Caption Downloader
Download subtitles as SRT or TXT from any YouTube video. Free, no sign-up.

常见使用场景

Whisper 是当前最强开源转录模型。历史上意味着 Python、GPU 和命令行——对普通用户完全不可及。浏览器版 Whisper 让流程变成:拖入音频,转录,下载 SRT 或 TXT。零安装,零账号,零代码。

隐私是关键。客户端 Whisper 意味着音频永远不离开设备。采访录音、医疗口述、法律证词、机密会议——任何不能上传给第三方 API 的音频,现在都可以安全转录。

使用技巧

模型大小的权衡:tiny/base 最快但精度一般,适合清晰音频;small/medium 是实用甜点,处理口音、噪声、术语都还不错;large 精度最高但慢——每分钟音频约 30 秒计算时间。

干净音频比大模型更重要。tiny 模型跑干净录音的效果比 large 跑浑浊音频还好。有录音控制时用好话筒、控制背景噪音、近讲。极长文件(2+ 小时)在浏览器内存里可能受限——用音频软件切成 30 分钟块分别转录。

常见问题

浏览器 Whisper 真的和桌面版一样准吗?

是的——模型权重完全相同。区别只是计算速度,取决于你的设备而非模型质量。

支持哪些音频格式?

MP3、WAV、M4A、MP4、WebM、OGG、FLAC——任何浏览器能通过 Web Audio API 解码的格式。视频文件也支持,会提取音频轨道。

能离线运行吗?

能——模型下载一次后(缓存在浏览器里),转录完全离线运行。适合飞机上或离线设备。

支持哪些语言?

Whisper 支持 99 种语言。对未知音频用自动检测;已知语言可以明确设置获得更快更准的结果。

我的音频文件会上传吗?

不会。音频完全在浏览器里通过 WebAssembly 解码和转录,没有上传、存储或日志。

相关免费工具

Built in the quiet hours, between everything else. Bookmark it for next time — and if it helped — Buy me a coffee
GistAI
GistAI App YT captions & AI summary Free on Google Play
If it saved you time — Buy me a coffee