How to Transcribe Audio to Text with AI for Free (2026 Guide)
Transcribe audio to text free with AI: compare Otter, TurboScribe, Word, iPhone and Pixel apps and local Whisper, with 2026 limits and step-by-step setup.

Table of contents
Difficulty
Beginner
Time required
30 minutes
Tools needed
- An audio or video file
- A free transcription tool (TurboScribe
- Otter or Whisper)
- A free AI chatbot
- Python 3 (optional for Whisper)
You've got a recording, maybe an interview, a lecture or a voice note you rambled into on a walk, and you want it as text without paying a monthly fee. Good news: in 2026 you can get a solid AI transcript for free. The catch is that every free option has a limit somewhere, whether that's minutes, file length, device or privacy.
Here's the short answer. For a one-off file under 30 minutes, TurboScribe's free plan is the easiest start. If you're on an iPhone 12 or later, Voice Memos and Notes transcribe recordings built in, and Pixel owners get the same from the Recorder app. If your audio is long, sensitive or you have lots of it, run OpenAI's Whisper on your own computer. It's open source, has no minute cap, and your audio never leaves your machine. Already paying for Microsoft 365? Word's Transcribe gives you 300 minutes of uploads a month with speaker labels.
The rest of this guide walks you through picking the right tool, getting a clean transcript, labeling speakers, tidying it up with a chatbot and exporting it in the format you need.
The free transcription tools worth knowing (September 2026)
| Tool | Free limit (as of September 2026) | Upload existing files? | Where it runs | Best for |
|---|---|---|---|---|
| TurboScribe (free) | 3 files a day, up to 30 min each | Yes | Cloud | Quick one-off files |
| Otter.ai Basic | 300 min a month, 30 min per conversation, 3 imports total (lifetime) | Only 3 ever | Cloud | Live notes and short meetings |
| Microsoft Word Transcribe | 300 min a month of uploads (needs a Microsoft 365 subscription) | Yes (.mp3, .m4a, .wav, .mp4) | Cloud | Existing Microsoft 365 users |
| Google Docs voice typing | No published limit | No, live mic only | Browser | Dictating as you speak |
| Apple Voice Memos and Notes | No published limit | Records in the app | iPhone 12 or later | iPhone voice notes and lectures |
| Google Pixel Recorder | No published limit | Records in the app | Mostly on device | Pixel owners |
| Samsung Voice Recorder (Transcript Assist) | Basic Galaxy AI features free | Records in the app | Galaxy devices, internet needed | Samsung owners |
| OpenAI Whisper (local) | Unlimited, limited only by your hardware | Yes, almost any format | Your computer | Long, private or bulk audio |
| MacWhisper (free tier) | Unlimited local transcription | Yes | Your Mac | Mac users who want an app, not a terminal |
A few notes on the fine print:
- Otter's free plan gives you 300 transcription minutes a month, a 30-minute cap per conversation and just 3 lifetime file imports, with exports limited to TXT (and MP3 audio). It's built for live recording and meetings, not for working through a folder of old files.
- Word Transcribe still exists and lives under Home, then Dictate, then Transcribe in Word for the web. Microsoft's support page says Microsoft 365 subscribers can transcribe up to 300 minutes of uploaded audio a month (30,000 with a Copilot license), in 80+ locales, and it needs Edge or Chrome. It isn't free on its own, but if your job or school already gives you Microsoft 365, it costs you nothing extra.
- Google Docs voice typing is free and works in the latest Chrome, Edge and Safari, but it only listens to your microphone in real time. You can't upload a file. Playing a recording through your speakers into the mic technically works, but the results are usually rough.
- TurboScribe lists 3 free transcriptions a day, each up to 30 minutes, with DOCX, PDF, TXT, SRT and VTT exports and speaker recognition, per its site. Its unlimited plan is listed at $10 a month billed annually.
Meeting bots that join Zoom or Teams calls and write notes for you are a different category. If that's what you're after, see our roundup of AI meeting note takers.
How to transcribe audio for free, step by step
Step 1: Prepare your audio
Transcription accuracy depends more on your recording than on which tool you pick. A few minutes of prep saves a lot of editing later.
- Get the best source you have. Use the original file, not a compressed copy sent through a messaging app.
- Cut the dead air. Long silences and music intros can confuse speech models. Whisper's own model card warns that it can "hallucinate," producing text that was never spoken, and silence is a common trigger.
- Use a common format. MP3, M4A and WAV work almost everywhere. Word Transcribe accepts .wav, .mp4, .m4a and .mp3.
- Split long files if needed. TurboScribe's and Otter's free plans cap files at 30 minutes, so a 90-minute lecture needs three chunks, or a local tool.
If you have ffmpeg installed (you'll need it for Whisper anyway), you can convert and split files from the terminal. The second command below cuts a file into 30-minute pieces:
Terminal: convert and split audio with ffmpeg # Convert any audio or video file to a 16 kHz mono WAV (ideal for whisper.cpp) ffmpeg -i lecture.m4a -ar 16000 -ac 1 -c:a pcm_s16le lecture.wav # Split a long file into 30-minute chunks without re-encoding ffmpeg -i lecture.mp3 -f segment -segment_time 1800 -c copy lecture_part%02d.mp3Step 2: Pick your tool by length, privacy and device
Match the tool to the job instead of defaulting to whatever you've heard of:
- Under 30 minutes, nothing sensitive: TurboScribe free. Upload, wait, download.
- You're recording right now on your phone: Voice Memos or Notes on iPhone 12 and later (Apple lists 10 supported languages, including English, Spanish, French, German, Japanese and Chinese), Recorder on a Pixel, or Voice Recorder with Transcript Assist on a Galaxy phone.
- Live meeting or class notes: Otter Basic, keeping the 30-minute cap in mind.
- Long recordings, lots of files or confidential audio: Whisper on your own computer. Therapy sessions, legal interviews, HR conversations and unpublished research don't belong on a free cloud service.
- You already have Microsoft 365: Word Transcribe, especially if you want the transcript in a Word document with speakers already labeled.
Step 3: Transcribe with a cloud tool or on your phone
For the easy routes, the process takes a couple of minutes:
- TurboScribe: sign up, drop in your file, choose the language, and turn on speaker recognition if there's more than one voice.
- Word for the web: open a document, go to Home, then Dictate, then Transcribe, and choose Upload audio. When it's done, you can add the whole transcript or selected sections to your document.
- iPhone Voice Memos: tap a recording, open the options, and choose View Transcript or Copy Transcript. In Notes, you can record audio inside a note, watch a live transcript and then add it to the note.
- Pixel Recorder: transcription runs as you record. If you picked the wrong language, open the recording, tap More, then Transcribe again (Pixel Help).
- Samsung: open a recording in Voice Recorder and use Transcript Assist, which also offers summaries and translation. Samsung says its basic Galaxy AI features are free.
Step 4: Or run Whisper on your own computer
Whisper is OpenAI's open source speech recognition model, released under the MIT license. It's what many paid transcription apps use under the hood, and you can run it yourself for free.
Is it good? OpenAI's research paper trained Whisper on 680,000 hours of multilingual audio. The best zero-shot model scored a word error rate (WER) of 2.5% on the LibriSpeech test-clean benchmark, and across other speech datasets it made 55.2% fewer errors on average than a supervised model with similar LibriSpeech scores. That second number matters more: it's why Whisper holds up on messy, real-world audio. The newer large-v3 model was trained on 1 million hours of weakly labeled plus 4 million hours of pseudo-labeled audio, and shows a 10% to 20% error reduction over large-v2 across a wide range of languages.
On languages, be realistic. The model card says training data covered 98 languages, but reports strong results in only about 10, and warns that accuracy drops for lower-resource languages and some accents. English, Spanish, French and German are in good shape. Smaller languages are worth a test clip first.
Here's the basic setup. You need Python and ffmpeg:
Terminal: install and run OpenAI Whisper # 1. Install ffmpeg (pick your system) brew install ffmpeg # macOS (Homebrew) sudo apt update && sudo apt install ffmpeg # Ubuntu or Debian choco install ffmpeg # Windows (Chocolatey) # 2. Install Whisper pip install -U openai-whisper # 3. Transcribe (the model downloads automatically the first time) whisper audio.mp3 --model small # Set the language and choose an output format (txt, srt, vtt, tsv, json or all) whisper interview.mp3 --model small --language English --output_format srtWhich model? The README lists six sizes. small (244M parameters, about 2 GB of VRAM) is a sensible default on an ordinary laptop. turbo (809M parameters, about 6 GB of VRAM) is an optimized version of large-v3 that the README lists as roughly 8 times faster than large, with a small accuracy trade-off. It's the one to pick if you have a decent GPU or an Apple Silicon Mac. Tiny and base are fast but noticeably less accurate. Without a GPU, expect long files to take a while.
Faster alternatives worth trying:
- faster-whisper: a reimplementation that's "up to 4 times faster than openai/whisper for the same accuracy while using less memory." Install it with
pip install faster-whisper. It doesn't need a separate ffmpeg install. - whisper.cpp: a C/C++ port with no Python required. It runs well on plain CPUs and Apple Silicon, and outputs TXT or SRT with the
-otxtand-osrtflags. - MacWhisper: a Mac app that wraps Whisper in a drag-and-drop window. The free version transcribes locally and exports TXT, SRT and VTT. Pro, a one-time €64 license, adds speaker recognition, batch jobs and DOCX, PDF and Markdown exports.
- faster-whisper: a reimplementation that's "up to 4 times faster than openai/whisper for the same accuracy while using less memory." Install it with
Step 5: Add speaker labels
A transcript of an interview without speaker names is hard to read. Your options depend on the tool:
- Automatic: Word Transcribe labels speakers as Speaker 1, Speaker 2 and so on, and lets you rename them. Pixel Recorder does the same on Pixel 6 and later, though Google says speaker labels work only in US English. TurboScribe and MacWhisper Pro also offer speaker recognition.
- Plain Whisper doesn't label speakers. Open source projects like WhisperX add it, but setup is more involved than a beginner guide should ask of you.
- The chatbot shortcut: for a two-person interview, a chatbot can often infer who's speaking from context (questions versus answers). Treat that as a draft and check it against the audio.
Step 6: Clean it up with an AI chatbot
Raw transcripts are full of "um," false starts, run-on sentences and misheard names. A free chatbot like ChatGPT, Claude or Gemini can tidy this up in seconds. Our ChatGPT vs Claude vs Gemini comparison covers the differences, but any of them works here.
The key is telling it what not to change. Otherwise it may "improve" quotes you need word for word.
Prompt: clean up a raw transcript Below is a raw AI transcript of a [interview / lecture / meeting] between [NAME 1] and [NAME 2]. Please: 1. Label each paragraph with the correct speaker name. If you're unsure who is speaking, write [SPEAKER?] instead of guessing. 2. Remove filler words (um, uh, you know, like) and false starts. 3. Fix punctuation and break it into readable paragraphs. 4. Correct obvious mishearings of these terms: [LIST NAMES, BRANDS, JARGON]. Flag anything else that looks wrong as [CHECK: ...]. 5. Keep timestamps if they exist. Do NOT summarize, reword or add anything. Keep the speakers' meaning and wording as close to the original as possible. TRANSCRIPT: [PASTE HERE]For long recordings, paste the transcript in sections of a few thousand words; free chatbot tiers can cut off or skip parts of very long inputs. Once it's clean, you can ask for a summary, action items or pull quotes in a separate message. Our guide to summarizing long PDFs with AI covers the same chunking approach for long documents, and how to write better prompts will help you tune the output.
Step 7: Export in the right format
Pick the format based on where the text is going:
- TXT: plain text for notes, search and pasting into a chatbot. Every tool here supports it, and it's Otter Basic's only text export.
- SRT or VTT: subtitle files with timestamps for YouTube, video editors and course platforms. Whisper, whisper.cpp, MacWhisper and TurboScribe can all produce these.
- DOCX: for editing, sharing and printing. Word Transcribe adds the transcript straight into a Word document, and TurboScribe and MacWhisper Pro export DOCX directly. Otherwise, paste your TXT into Word or Google Docs.
- Google Docs: Pixel Recorder can share a transcript straight to Google Docs or as a .txt file.
Before you publish or send anything, do one last listen on fast-forward while skimming the text. Names, numbers and technical terms are where every AI transcriber still slips up.
Which free option should you pick?
If we had to choose one path for most beginners, it'd be this: use your phone's built-in recorder for anything you capture yourself, TurboScribe for the occasional uploaded file, and Whisper (or MacWhisper) once you're dealing with long, frequent or private audio. That combination costs nothing and covers almost everything.
Otter's free plan is fine for short live sessions, but three lifetime imports make it a poor fit for existing files. Word Transcribe is excellent value if you already pay for Microsoft 365, and not worth subscribing for on its own. For more ways to cut busywork beyond transcription, see our best AI productivity tools for work.
Frequently asked questions
What is the best free AI to transcribe audio to text?
For most people, OpenAI's Whisper is the most capable free option because it's open source, runs on your own computer and has no minute limit. If you'd rather not install anything, TurboScribe's free plan (3 files a day, up to 30 minutes each, as of September 2026) is the easiest browser option.
Can Google Docs transcribe an audio file?
No. Google Docs voice typing only transcribes live speech from your microphone. It can't upload or read an audio file. For files, use TurboScribe, Word Transcribe (with Microsoft 365) or Whisper.
Is Microsoft Word's Transcribe feature free?
It's included with a Microsoft 365 subscription, not with free Word. Subscribers get up to 300 minutes of uploaded audio a month, and it works in Word for the web on Edge or Chrome.
How accurate is Whisper?
In OpenAI's paper, the best zero-shot Whisper model scored a 2.5% word error rate on the LibriSpeech clean benchmark and made far fewer errors than comparable models on messier real-world datasets. Accuracy drops with background noise, heavy accents and less common languages, so always proofread names and numbers.
Is it safe to upload private recordings to free transcription sites?
Cloud tools store your audio and text on their servers, so check their privacy policy before uploading anything confidential. For sensitive recordings, run Whisper, whisper.cpp or MacWhisper locally so the audio never leaves your computer.

Written by
Panoptix Editorial Team
Our editors test AI tools hands-on for weeks before we publish a word. We pay for our own subscriptions and never accept payment for rankings.
How we test AI tools →

