Skip to content

Rapidly and accurately convert WAV to text

Upload studio sessions, broadcast masters, or field recordings, and Scribe converts audio details into a precise, speaker-labeled transcript.

Studio recordings.pdf

From lossless WAV to text in minutes

Upload uncompressed audio and Scribe builds the transcript from the full signal, not a lossy copy. You get attributed, timestamped text ready for the edit.

1

Upload your WAV file

Drag in a file from your session folder, field recorder card, or cloud storage. Long, high-resolution WAV files upload without conversion or splitting.

2

Edit your transcript instantly

Correct names and technical terms, merge or split segments, and reassign speakers in the editor. Each word keeps its timestamp, so edits stay locked to the recording.

3

Export in any format you need

Download TXT, DOCX, or PDF for scripts and review copies, or SRT, VTT, and JSON when captions and logs need precise timing.

Not just transcription. Audio understanding

Uncompressed source audio deserves transcription that keeps up. Scribe reads the detail in your lossless files and returns structured, attributed text.

#1 Accuracy

Industry-leading transcription accuracy, delivering clean, editable text even in challenging audio conditions and across diverse accents and dialects.

Scribe beats all competing models in accuracy benchmarks

Edit the transcripts

Fix terminology, split takes, and reassign voices without leaving the transcript. Timing data follows every change.

Amidst the outer atmosphere of the planet Aurora, the sky shimmered with fractured light, as though the planet's veil were made of stained glass suspended in space.
Sensors pulsed with irregular patterns, the kind no algorithm could quite reconcile.
Amidst the outer atmosphere of the planet Aurora, the sky shimmered with fractured light, as though the planet's veil were made of stained glass suspended in space.

90+ Languages and accents

International broadcasts and multilingual field recordings transcribe accurately across 90+ languages, with the language detected automatically on upload.

Japanese
Hindi
Polish
Swedish
Mandarin
Vietnamese
French

Wide variety of formats

Upload WAV alongside MP3, FLAC, OGG, M4A, MP4, and MOV. Export to TXT, DOCX, PDF, JSON, SRT, VTT, or HTML.

Audio Event Tagging

Scribe records applause, coughing, and other non-speech events in the transcript, so the text reflects everything your microphones picked up.

Speaker Timestamps

Scribe labels up to 32 speakers and timestamps across the transcript, turning a panel, session, or control-room feed into an ordered, attributed record.

WAV Transcript Export Formats

Text file icon labeled "board_call.txt" on a textured background.

Transcribe WAV to TXT

Document icon with the filename "interview.docx" on a textured background.

Transcribe WAV to DOCX

A document icon labeled "meeting.pdf" on a textured background.

Transcribe WAV to PDF

Icon representing a JSON file named "playlist.json" on a textured background.

Transcribe WAV to JSON

File icon with HTML code and filename "video_ad.html" on a textured background.

Transcribe WAV to HTML

SRT file icon labeled "film.srt" on a textured gradient background.

Transcribe WAV to SRT

Audio file icon labeled "movie.avid" on a red-orange gradient background.

Transcribe WAV to AVID

Closed caption file icon labeled "series.vtt" on a textured background.

Transcribe WAV to VTT

Millions of words transcribed, and counting

  • I use ElevenLabs primarily for transcribing audio messages, and I find its accuracy to be a major highlight. This precision allows me to analyze students' reading fluency effectively, even when the speaker is a young student still learning to read, which is crucial for understanding each student's progress.
    G2 logo

    Pedro A.

    Head of technology

  • Perfect for transcribing interviews - and the voice quality is amazing when preparing for a speech.
    G2 logo

    Izabela M.

    Customer Experience Researcher

  • Remarkable inference speed of the Scribe v2 model by ElevenLabs, delivering near real-time latency on transcription requests, significantly faster than other models we've tried.
    G2 logo

    Vedaswaroop I.

    Founder

Turn audio to text today, starting at no cost

End-to-end audio Productions

Add human review to editing so your message always lands.

  • Synced captions and subtitles
  • Human edited translations
  • Predictable pricing
ElevenLabs Studio Capabilities

Audio to Text API and SDK

Integrate transcription directly into your product with a few lines of code.

  • Native SDKs for web and mobile
  • WebSocket and REST APIs
  • Community of 100k+ developers
Scribe API Graphic

Get started on the web

Turn audio to text using our ElevenCreative web platform.

  • 10k credits included, every month
  • 90+ languages and accents
  • Flexible pricing for larger volumes
Use TTS in the ElevenLabs Studio

Frequently asked questions

Create with the highest quality AI Audio