09 / 19⏳ 3 min
Audio nodes
Audio nodes cover voice and sound: text to speech, voice cloning and design, and the audio-processing tools that go with video.
Text to speech
From the Add node menu open Audio → Speech to create a text-to-speech node:
- Type the text to read aloud, pick a voice and parameters, and generate a clip;
- The bar has rate, pitch and volume prosody sliders and pause / emphasis markup helpers (some are model extensions; a model that doesn't support them reads the markup as plain text);
- Text length is capped at about 50,000 characters, a soft on-screen counter.

Voice cloning and design
- Voice cloning: upload a reference clip to register a custom voice you can then synthesize with;
- Voice design: design a new voice from a description, without a reference clip;
- Registering a voice charges a registration fee and runs asynchronously in the background; a voice that is Processing has to become Ready before you can use it.
Related audio tools
Many audio capabilities actually live on the video node's toolbar (see Video nodes): split A/V, mix audio, vocal separation and the like all pull sound out or replace it, mostly for free.
TTS has prosody sliders and pause/emphasis markup · Cloning / designing a voice has a fee and becomes ready async · Split/mix and other audio tools are on the video node toolbar