Transcription
Audio is uploaded for cloud recognition (this step uploads), with speaker labels and timestamps; long audio is recognized 10 minutes at a time with progressive results. Afterwards you can delete words in the transcript to cut the matching audio, and export TXT, SRT, VTT, Markdown, JSON, CSV or LRC.
AI assistant editing
Open the AI assistant in the editor and say “shorten all pauses to 0.5 seconds”, “delete every um and uh” or “halve the volume of the second clip”. It reads and writes your timeline (tracks, clips, volume, fades, speed, EQ, markers) and transcript; edits that change audio are previewed and applied only after you confirm, and every step can be undone. It does not run noise reduction, dereverb or loudness normalization, and it does not generate new voices.
What is not AI
Noise reduction (spectral gate), voice enhancer (denoise + dereverb + compression + limiter), hum removal, loudness normalization (ITU-R BS.1770), silence detection, BPM detection, vocal removal (center-channel cancellation) and all effects are traditional algorithms running in your browser without uploading.
Which models
TutuAudio builds no models; TutuAgent connects top-tier speech-recognition and language models and keeps updating them.
- Will the AI mess up my audio?
- Edits that change audio are previewed and applied only after you confirm, and every step can be undone.
- Do I need to sign in or pay?
- No. Both AI features include a daily free allowance metered by TutuAgent.
- Is my audio used to train models?
- TutuAudio does not train models. For transcription, audio is sent to the recognition service TutuAgent connects to.