Qwen3.0 Audio TTS Flash
by Qwen
Qwen's fast text-to-speech model for low-latency Chinese speech.
Qwen Audio 3.0 TTS Flash trades a little fidelity for responsive synthesis at a lower price, with the Loong voices. Serves through OpenRouter's speech API.
More details
Qwen3.0 Audio TTS Flash by Qwen is a fast-speed AI text-to-speech model. Qwen Audio 3.0 TTS Flash trades a little fidelity for responsive synthesis at a lower price, with the Loong voices. Serves through OpenRouter's speech API. Key capabilities include 274 voices across 12 languages, up to 4,096 characters per request.
Qwen3.0 Audio TTS Flash is priced at $0.02 per 1,000 characters of input text. You can generate speech with Qwen3.0 Audio TTS Flash on idapt.app alongside 200+ AI models in one workspace.
Voices
2 · 1 languagesAny of these voices can be paired with Qwen3.0 Audio TTS Flash. Pick a model for quality and speed, a voice for character and language.
Loong John
Chinese · Male
Longanhuan
Chinese · Female
Pricing
Cost per 1,000 characters
$0.02
Typical paragraph (~500 chars)
$0.009
Billed on input text length
Capabilities
Speed
Fast: Low-latency, real-time ready
Emotion Control
Natural delivery only
Max Input
4,096 characters per request
Voices
2 voices · 1 languages