Audio generation
Found this helpful? Share it:
Found this helpful? Share it:
Give the assistant text and it speaks it back in a natural voice. Every clip is saved to your Drive automatically, so you can download it, share it, or use it later 🔊
Audio generation Free and upworks on every signed-in account, the free tier included. You are charged per clip by length, so short text costs less.
The simplest way is to ask in a chat: paste or type the text and ask the assistant to read it aloud. It generates the clip with your current voice and settings and saves it to your Drive. You can also start a generation from the Generate audio action in the Drive browser toolbar.
idapt offers 11 text-to-speech models across providers like MiniMax, OpenAI, xAI, and Google, each with its own set of voices. The free default is grok-tts. Only the premium speech-2.8-hd model Pro and upis locked behind a subscription, so every other model is open to you.
Set the model and voice in Settings → Voice Mode. The voice list filters to the chosen model's provider, so you always see voices that actually work with it. Browse the full catalog on the Voice modelspage.
Beyond the voice itself, you can influence how a line is read:
Speed: from 0.5x for a deliberate pace up to 2.0x for quick previews.
Pitch: shift the voice up or down, from -12 to +12 semitones.
Emotion: steer the delivery toward happy, sad, angry, or neutral, or describe the feeling you want in your own words.
The effect varies by model and voice. Some express emotion and pitch shifts far more dramatically than others, so it is worth trying a couple before committing to a long clip.
Every clip is saved to a Generated Audio folder in your Drive, and you can find it in the Drive browser any time. Ask for a different folder or filename in your message to place it exactly where you want.
Each model caps how much text one clip can hold. The free default tops out around 4,000 characters, and other models go up to 10,000. For a longer script, split it into a few shorter clips.
This is about turning your own text into a saved audio file. If instead you want to hear an assistant reply spoken out loud, that is playback, not generation. See Listening to replies.
Related articles
Listening to replies
Play any reply out loud, choose the voice, and replay cached audio for free.
Image generation
Ask AI to create images from text descriptions, saved to your Drive automatically.
Chat settings and toggles
Customize each conversation from the composer: tool toggles, autonomy, compaction, and a spend limit.
Was this helpful?