图图音频

Audio to text with speakers and timestamps

TutuAudio uploads your recording to the cloud for speech recognition (this step requires an upload), labels speakers and adds timestamps. Long audio is processed in 10-minute chunks, so text appears as it goes. Export TXT, SRT, VTT, Markdown, JSON, CSV or LRC. No login needed; there is a daily free allowance metered by TutuAgent. It does not make video subtitles — use TutuCut for that.

Record

MP3, WAV, M4A, AAC, FLAC, OGG, WebM and the audio in MP4 videos

Transcription uploads the audio to the cloud; editing and effects still run in your browser

Drop in a recording and the text comes out section by section, with who said it and when.

How it works

  1. 1

    Drag a recording onto the page or click "Choose file"

  2. 2

    The audio is uploaded for recognition and text appears chunk by chunk

  3. 3

    Review it and export TXT, SRT or another format

FAQ

Is it free? Do I need an account?
No account needed. There is a daily free allowance metered by TutuAgent; if you use it up, come back the next day.
Is my recording uploaded?
Yes. Recognition requires uploading the audio to the cloud. Editing, deleting words and exporting afterwards happen in your browser without another upload.
Can it handle a one-hour meeting?
Yes. Long audio is recognized in 10-minute chunks, one after another, and earlier text shows up while the rest is still processing.
How accurate is it?
Clear speech works well. Heavy accents or dialects, people talking over each other and noisy backgrounds lower the accuracy, so give it a listen-through.
Can it add subtitles to a video?
This page works with audio only, but you can export SRT or VTT files. To caption a video directly, use TutuCut.

Keep going with this audio

Similar tools

Not here? For text-to-speech and AI voices use TutuSound; to edit video or pull audio from video use TutuCut. All audio tools