Free Web Tool

Free Audio to Text

Transcribe any audio using Whisper AI — runs entirely in your browser. Private, free, no sign-up, no upload to servers.

🎵
Drop audio here or browse
MP3 · WAV · M4A · OGG · FLAC · WebM · max 25 MB
Language
🤖 Uses Whisper Tiny AI model (~40 MB, downloaded once and cached). All processing runs in your browser — your audio never leaves your device.

How to use

1
Upload an audio file (MP3, WAV, M4A…) or record directly from your mic.
2
Select a language or leave on Auto-detect. Then click Transcribe.
3
The first run downloads the Whisper AI model (~40 MB). After that it's instant — even offline.
4
Copy or download the transcript. Nothing is sent to any server.

Frequently Asked Questions

Is this free?
Yes — completely free, no sign-up, no limits.
Is my audio private?
Yes. Your audio never leaves your device. The Whisper AI model runs entirely inside your browser using WebAssembly — no server, no tracking.
What audio formats are supported?
Any format your browser can decode: MP3, WAV, M4A, OGG, FLAC, WebM. If it plays in your browser, it can be transcribed.
What languages does it support?
Whisper supports 99 languages including English, Chinese, Japanese, Korean, Spanish, French, German, Portuguese, Indonesian, Arabic, Russian and more. Use Auto-detect or pick a language for better accuracy.
Why does it take time on first use?
The Whisper AI model (~40 MB) is downloaded once from a CDN and cached by your browser. After the first use, transcription starts immediately — even offline.
Is there a file size limit?
25 MB per file. For longer recordings, consider splitting the audio first. Processing time depends on your device — modern laptops handle 5 minutes of audio in about 30 seconds.
Also try
🔊
Free Text to Speech
Convert text to natural speech using your device's voices. Free, no sign-up.
🎬
YouTube Caption Downloader
Download subtitles as SRT or TXT from any YouTube video. Free, no sign-up.

よくある使用シーン

Whisperは現在最強のオープンソース転写モデルです。従来はPython、GPU、コマンドラインが必要で一般ユーザーには手が届きませんでした。ブラウザ版Whisperなら音声をドロップ、転写、SRTまたはTXTをダウンロード——インストール不要、アカウント不要、コード不要。

プライバシーは決定的です。クライアントサイドWhisperなら音声はデバイスから出ません。インタビュー録音、医療口述、法的証言、機密会議——第三者APIに上げられない音声も安全に転写できます。

使い方のコツ

モデルサイズのトレードオフ:tiny/baseは最速だが精度は普通、クリーンな音声向け。small/mediumは実用的なスイートスポット、訛り・ノイズ・専門用語を程よく処理。largeは最高精度だが遅く、音声1分あたり約30秒の計算時間。

クリーンな音声は大きなモデルより重要です。tinyモデルでもクリーンな録音ならlargeで濁った音声を回すより結果が良いです。録音を制御できるならまともなマイク、背景ノイズ抑制、近接マイクで。超長時間ファイル(2+時間)はブラウザメモリに引っかかる可能性——音声ソフトで30分単位に分割してから転写を。

よくある質問

ブラウザWhisperはデスクトップ版と同じ精度ですか?

はい——モデル重みは完全に同じです。違いは計算速度のみで、モデル品質ではなくデバイスに依存します。

対応する音声フォーマットは?

MP3、WAV、M4A、MP4、WebM、OGG、FLAC——ブラウザがWeb Audio APIでデコードできるものすべて。動画ファイルも対応、音声トラックを抽出します。

オフラインで動きますか?

動きます——モデルを一度ダウンロード(ブラウザにキャッシュ)すれば、転写は完全オフラインで動作。飛行機内やエアギャップ環境で有用。

対応言語は?

Whisperは99言語対応。未知の音声には自動検出、既知言語は明示指定で若干高速・高精度の結果が得られます。

音声ファイルはアップロードされますか?

いいえ。音声はブラウザ内でWebAssemblyを介してデコード・転写されます。アップロード、保存、ログ記録は一切ありません。

関連する無料ツール

Built in the quiet hours, between everything else. Bookmark it for next time — and if it helped — Buy me a coffee
GistAI
GistAI App YT captions & AI summary Free on Google Play
If it saved you time — Buy me a coffee