tomai
Log in
☁️ Cloud · AI · credits

Audio to Text

Transcribe speech from any video or audio file into plain text — powered by Qwen3-ASR

🪙 5 credits per transcription

Speech is transcribed by Qwen3-ASR running on our servers, then converted straight to plain text — no timestamps, no subtitle formatting, just the words. Need SRT/VTT captions instead? Use AI Subtitles.

ASR engine, languages and accuracy notes

Transcription runs Qwen3-ASR on the worker with these characteristics:

Audio up to 500 MB or 2 hours per job; typical turnaround is about one tenth of runtime. Output appears in-page with a copy button and downloads as UTF-8 TXT.

How it works

  1. 1

    Upload any video or audio file — format doesn't matter, the server decodes it.

  2. 2

    Pick the spoken language, or leave it on auto-detect.

  3. 3

    Copy the transcript or download it as a .txt file. Files auto-delete in 30 minutes.

  4. 4

    Copy the transcript from the page or download the .txt file; original audio is deleted after processing.

Transcripts with speaker-aware punctuation

Upload a recording, get punctuated, paragraphed text — not a wall of lowercase words — in any of 11 interface languages.

🎧

Qwen3-ASR Engine

The same speech recognition that powers our AI subtitles runs on the server — interviews, lectures and meetings come back as readable plain text.

📝

Any Format Accepted

Audio and video files alike, including containers your browser cannot decode — upload, and the server's transcoder handles the rest.

⏱️

Language Hint or Auto

Give the recognizer a language hint for better accuracy, or let it detect automatically — the transcript is clean text, ready to copy or search.

Frequently asked questions

How is this different from the AI Subtitles tool?

Same underlying Qwen3-ASR transcription — this tool just skips straight to plain text with no timestamps, cue numbers or subtitle-file editor. Use AI Subtitles instead if you need SRT/VTT/ASS caption files or burned-in subtitles.

How accurate is the transcript?

Accuracy depends on audio quality, accent and background noise, but Qwen3-ASR handles typical speech (meetings, interviews, podcasts, lectures) well. Very noisy audio or heavy cross-talk will still need manual proofreading.

What happens to my file?

The upload and the transcript are deleted automatically within 30 minutes. Jobs are only reachable with your session's signed token.

Does it identify who said what?

Not yet — diarization is on the roadmap. For interviews, the punctuation still makes turns visible in most cases, but names aren't attached to segments.

Related tools

What Is Audio to Text?

Audio to Text turns the speech in a recording into a clean, plain-text transcript. Powered by the same Qwen3-ASR speech recognition that runs our AI subtitles service, it delivers readable text without timestamps or formatting — ready to copy into notes, search through, or paste anywhere. It accepts audio and video files alike, in any format, including ones your browser cannot decode. Journalists turn interviews into notes, students convert lectures to study material, and professionals make meeting recordings searchable. A language hint, or automatic detection, helps the recognizer, and you can copy the result with one button or download it as a text file.

What this tool can do

  • 📝 Transcribe speech into plain, unformatted text
  • 🎬 Accept both audio and video files in any format
  • 🌐 Choose a language hint — auto, Chinese, English, Japanese, Korean, Spanish, French, German, Portuguese, Russian or Hindi
  • 📋 Copy the transcript with one click or download a text file
  • 🔍 Produce clean text that is easy to search and reuse

When you would use it

  • Turning an interview recording into written notes
  • Converting a lecture into study material you can skim
  • Making a meeting recording searchable afterwards
  • Getting text from a video voiceover for captions or an article

Your audio travels over an encrypted connection to our servers and is automatically deleted within 30 minutes. Accuracy depends on the recording: clear, single-speaker speech with low background noise transcribes well, while heavy noise, strong accents and overlapping speakers lower the quality — and the output is plain text only, with no speaker labels or timestamps. For best results, record in a quiet room and let one person speak at a time.

Related tools