Four modes in one app: clone a voice from a 5–10 second sample, Voice Design (TTS) by attributes, Multi-Voice Dialogue for podcasts and audio drama, and timestamped Speech-to-Text. 600+ languages, fully offline.
Same structure as the Voice Studio interface
Clone a voice from a 5–10s sample.
Create a new voice by attributes — no sample needed.
Multiple characters, one voice each. Podcasts, audio drama.
Transcribe speech, export .txt or .srt.
Translate .srt with an LLM — timing and line count stay locked.
Download TTS/ASR models, pick an LLM, check your hardware.
Run a local server so n8n, Make or Story Machine can call in.
Ships with <b>30 factory voices</b> (male/female, varied tones). Save your own, click <b>⭐</b> to pin favorites. Available in Voice Cloning, Voice Design and Multi-Voice Dialogue. Backup / restore via portable <code>.vcp</code> files to share across machines.
Live screenshots of the running app — not mockups.
The 5–10 second sample sits on the left, the text to read on the right. The <b>Voice Library</b> underneath holds 30 built-in voices plus the ones you save.

Write the script as <code><Name>: line</code>, cast a separate voice for every character, then push the whole script into the <b>render queue</b>.

Line count and timecodes are <b>immutable</b> — the AI may only change the words. A consistency table locks register, terminology and forms of address across the whole file.

Start the local server and let n8n, Make, a Python script — or <b>G-Labs Story Machine</b> — call in for audio. API key, sample body and a cURL example are right there.

From sample audio to WAV / MP3