Audio Models

Audio Models are the AI models available in the Audio Node's generation panel. The available model depends on the generation mode you select — Text to Speech, Text to Music, Text to Sound Effect, or Voice Design.


Available models by mode

ModeModelCustom voice
Text to SpeechFishAudioYes (Add Voice)
Text to SpeechElevenLabs V3No (official voices only)
Text to MusicMusic V1
Text to MusicLyria 3 Pro
Text to Sound EffectElevenLabs (sound effects)
Voice DesignElevenLabs (Voice Design)

Choosing a model

FishAudio runs locally in Zeemo and is the only Text to Speech model that supports custom voices. Use it when you want to create and reuse a specific voice across your projects — see Adding a custom voice on the Text to Speech page.

ElevenLabs V3 offers high-quality official voices for Text to Speech. It does not support custom voices — for that, use FishAudio.

Music V1 and Lyria 3 Pro both generate music tracks from a text description. Duration ranges from 5s to 300s in 30-second increments, and cost scales with length.

ElevenLabs (sound effects) generates short sound effects from a description, from 0.5s to 30s.

ElevenLabs (Voice Design) generates three voice samples of about 8 seconds each from a description of vocal qualities. Once you find a sample you like, save it as a custom voice for Text to Speech via FishAudio.


Parameters

ParameterApplies toDescription
VoiceText to SpeechChoose an official or custom voice
DurationText to Music, Text to Sound EffectOutput length (Music: 5–300s; Sound Effect: 0.5–30s)

Tip: The credit cost is shown on the Generate button before you confirm. For Music and Sound Effect modes, longer clips cost more — check the cost before generating long durations.

⚠️

Note: Custom voices created with FishAudio are available across all your canvases, not just the one where they were created.


Did this page help you?