Audio Models
Audio Models are the AI models available in the Audio Node's generation panel. The available model depends on the generation mode you select — Text to Speech, Text to Music, Text to Sound Effect, or Voice Design.
Available models by mode
| Mode | Model | Custom voice |
|---|---|---|
| Text to Speech | FishAudio | Yes (Add Voice) |
| Text to Speech | ElevenLabs V3 | No (official voices only) |
| Text to Music | Music V1 | — |
| Text to Music | Lyria 3 Pro | — |
| Text to Sound Effect | ElevenLabs (sound effects) | — |
| Voice Design | ElevenLabs (Voice Design) | — |
Choosing a model
FishAudio runs locally in Zeemo and is the only Text to Speech model that supports custom voices. Use it when you want to create and reuse a specific voice across your projects — see Adding a custom voice on the Text to Speech page.
ElevenLabs V3 offers high-quality official voices for Text to Speech. It does not support custom voices — for that, use FishAudio.
Music V1 and Lyria 3 Pro both generate music tracks from a text description. Duration ranges from 5s to 300s in 30-second increments, and cost scales with length.
ElevenLabs (sound effects) generates short sound effects from a description, from 0.5s to 30s.
ElevenLabs (Voice Design) generates three voice samples of about 8 seconds each from a description of vocal qualities. Once you find a sample you like, save it as a custom voice for Text to Speech via FishAudio.
Parameters
| Parameter | Applies to | Description |
|---|---|---|
| Voice | Text to Speech | Choose an official or custom voice |
| Duration | Text to Music, Text to Sound Effect | Output length (Music: 5–300s; Sound Effect: 0.5–30s) |
Tip: The credit cost is shown on the Generate button before you confirm. For Music and Sound Effect modes, longer clips cost more — check the cost before generating long durations.
Note: Custom voices created with FishAudio are available across all your canvases, not just the one where they were created.
Updated about 2 months ago

