Audio & voices
Aquilla has audio built in, both listening (text-to-speech) and recording with transcription (speech recognition). For audio, video, and subtitle files, see also Media view & timing for lining your recordings up against the original on the timeline.
Transcription runs entirely in your browser, and Aquilla also includes in-browser voice models that download once and cache locally, so after the first use you can listen and transcribe even offline, a real advantage in low-connectivity settings. Some higher-quality voices use a server instead; those need a connection.
Listening: text-to-speech (TTS)
Section titled “Listening: text-to-speech (TTS)”You can have any cell read aloud. Switch the editor to its Audio tab and each translated cell grows a voice control; the Voices tab in the sidebar holds your project’s cast.

- Your project keeps a cast of voices. Each voice is backed by a
text-to-speech engine, your choice of:
- OmniVoice (recommended): a hosted neural voice with no setup and no API key.
- Gemini: the highest-quality, promptable voices in many languages. It runs on Google’s servers, so it needs your own Google AI key (added in Project Settings).
- Kokoro (English) or MMS (many languages): free voice models that run in your browser, with no key and no connection once the model has downloaded.
- Assign a voice to a cell (or to a character whose lines share a voice), then play a single cell with one click, or queue several cells into a playback queue (the player bar along the bottom) to hear how a passage flows in the target language.


The first time you use an in-browser voice, Aquilla downloads its model and caches it; after that, playback works offline.
Every cell’s audio also lives in its detail panel’s Recording tab: the clip with a scrubber, where the audio came from (an AI-generated voice notes “AI generated voice. Drag a voice from the toolbar to regenerate, or:” with a Record over button), and, for recorded cells, the Takes strip where you pick the take to use, rename or delete takes, clean one up with Remove noise (adds a cleaned take), or Revert to the original recording.

Recording: speech recognition (ASR)
Section titled “Recording: speech recognition (ASR)”You can also record audio and have Aquilla transcribe it:
- Choose Record audio on a cell and a guided prompter opens, showing the text to read aloud. Press Space or click Start: a beep and a 3-2-1 countdown lead into recording (Aquilla asks for mic permission first, with a clear explanation of why).
- If your browser has blocked microphone access, the record button shows “Microphone access blocked — click for help”: click it for steps to re-allow the microphone in your browser’s site settings, then reload the page.
- An in-browser speech-recognition model transcribes what you said. This works on a take you’ve only just recorded, too, even offline or before it has synced to the server, so you can record and transcribe in one sitting without waiting for an upload.
- While it works, a small status badge next to the Transcribe button keeps you posted: a download percentage the first time the model fetches, a transcribing pill while it runs, and a word count when it lands. If something goes wrong you’ll see Transcription failed: click it for the details, with Retry and Dismiss right there. No more silently reverting buttons.
- Review the transcript and, if it’s good, accept it into the cell, which queues it as a normal edit.
- Prev / Next in the prompter step through the file cell by cell, so you can record a whole passage in one sitting, and you can transcribe several recordings at once when you have a batch.
Playing back a recording streams: long takes start playing almost immediately and you can jump around within them instantly, instead of waiting for the whole file to download first. One limit to know about: a single recording or clip can be at most 95 MB; anything larger is turned away with a clear message before any upload starts.

Cloning a voice
Section titled “Cloning a voice”Beyond the ready-made library, you can give a voice the timbre of a specific person. When you set up a voice for a character, add a reference recording: either record a short segment or upload an audio clip. Aquilla re-voices that voice’s text-to-speech output into the reference’s timbre. A voice with a reference attached is marked Cloned; leave the reference empty to use the base voice as-is.
Help assistant
AI answers from the Aquilla docs — may make mistakes.