Skip to content

Audio & voices

Aquilla has audio built in, both listening (text-to-speech) and recording with transcription (speech recognition). For audio, video, and subtitle files, see also Media view & timing for lining your recordings up against the original on the timeline.

Transcription runs entirely in your browser, and Aquilla also includes in-browser voice models that download once and cache locally, so after the first use you can listen and transcribe even offline, a real advantage in low-connectivity settings. Some higher-quality voices use a server instead; those need a connection.

You can have any cell read aloud. Switch the editor to its Audio tab and each translated cell grows a voice control; the Voices tab in the sidebar holds your project’s cast.

The workspace in the Audio tab: the sidebar shows the project's Voices, and each translated cell shows a voice selector reading Click a voice to generate; untranslated cells read Translate to voice this line
The Audio view. Translated cells are ready to voice; untranslated ones ask for text first.
  • Your project keeps a cast of voices. Each voice is backed by a text-to-speech engine, your choice of:
    • OmniVoice (recommended): a hosted neural voice with no setup and no API key.
    • Gemini: the highest-quality, promptable voices in many languages. It runs on Google’s servers, so it needs your own Google AI key (added in Project Settings).
    • Kokoro (English) or MMS (many languages): free voice models that run in your browser, with no key and no connection once the model has downloaded.
  • Assign a voice to a cell (or to a character whose lines share a voice), then play a single cell with one click, or queue several cells into a playback queue (the player bar along the bottom) to hear how a passage flows in the target language.
The new-voice dialog with TTS voice and Clone voice tabs, a name field, and four engine choices: OmniVoice (recommended), Gemini, Kokoro, and MMS
Adding a voice to the cast: pick an engine. OmniVoice needs no key or setup.
A cell after generating audio: a waveform with a play button sits above the voice selector; selecting the cell adds crop and volume controls, and the playback bar waits at the bottom of the workspace
Click a cell's voice to generate its audio: you get a waveform, play button, and trim controls.

The first time you use an in-browser voice, Aquilla downloads its model and caches it; after that, playback works offline.

Every cell’s audio also lives in its detail panel’s Recording tab: the clip with a scrubber, where the audio came from (an AI-generated voice notes “AI generated voice. Drag a voice from the toolbar to regenerate, or:” with a Record over button), and, for recorded cells, the Takes strip where you pick the take to use, rename or delete takes, clean one up with Remove noise (adds a cleaned take), or Revert to the original recording.

A cell's detail panel open on the Recording tab, showing the audio clip with a scrubber, a note that the audio is an AI generated voice, and a Record over button
The Recording tab: what this cell's audio is, where it came from, and the controls to redo it.

You can also record audio and have Aquilla transcribe it:

  • Choose Record audio on a cell and a guided prompter opens, showing the text to read aloud. Press Space or click Start: a beep and a 3-2-1 countdown lead into recording (Aquilla asks for mic permission first, with a clear explanation of why).
  • If your browser has blocked microphone access, the record button shows “Microphone access blocked — click for help”: click it for steps to re-allow the microphone in your browser’s site settings, then reload the page.
  • An in-browser speech-recognition model transcribes what you said. This works on a take you’ve only just recorded, too, even offline or before it has synced to the server, so you can record and transcribe in one sitting without waiting for an upload.
  • While it works, a small status badge next to the Transcribe button keeps you posted: a download percentage the first time the model fetches, a transcribing pill while it runs, and a word count when it lands. If something goes wrong you’ll see Transcription failed: click it for the details, with Retry and Dismiss right there. No more silently reverting buttons.
  • Review the transcript and, if it’s good, accept it into the cell, which queues it as a normal edit.
  • Prev / Next in the prompter step through the file cell by cell, so you can record a whole passage in one sitting, and you can transcribe several recordings at once when you have a batch.

Playing back a recording streams: long takes start playing almost immediately and you can jump around within them instantly, instead of waiting for the whole file to download first. One limit to know about: a single recording or clip can be at most 95 MB; anything larger is turned away with a clear message before any upload starts.

The record-audio prompter for a cell, showing the source text and the translation to read aloud, a microphone icon, instructions to press Space or click Start with a 3-2-1 countdown, and Prev and Next buttons
The recording prompter: it shows what to read, counts you in, and steps to the next cell when you're done.

Beyond the ready-made library, you can give a voice the timbre of a specific person. When you set up a voice for a character, add a reference recording: either record a short segment or upload an audio clip. Aquilla re-voices that voice’s text-to-speech output into the reference’s timbre. A voice with a reference attached is marked Cloned; leave the reference empty to use the base voice as-is.