Speech 02 HD
by MiniMax
MiniMax's proven studio-quality voice model.
Speech 02 HD produces lifelike, expressive speech with reliable pronunciation and emotional nuance — a dependable balance of quality, language coverage, and cost.
More details
Speech 02 HD by MiniMax is a standard-speed AI text-to-speech model. Speech 02 HD produces lifelike, expressive speech with reliable pronunciation and emotional nuance — a dependable balance of quality, language coverage, and cost. Key capabilities include 78 voices across 7 languages, emotion control, up to 10,000 characters per request.
Speech 02 HD is priced at $0.06 per 1,000 characters of input text. You can generate speech with Speech 02 HD on idapt.app alongside 200+ AI models in one workspace.
Voices
30 · 7 languagesAny of these voices can be paired with Speech 02 HD. Pick a model for quality and speed, a voice for character and language.
Wise Woman
English · Female
Graceful Lady
English · Female
Trustworthy Man
English · Male
Aussie Bloke
English · Male
Professional Woman
English · Female
Friendly Person
English · Neutral
Deep Voice Man
English · Male
Lively Girl
English · Female
Narrator
English · Neutral
Male Narrator
English · Male
Female Narrator
English · Female
Calm Woman
English · Female
Qingse (Male)
Chinese · Male
Shaonv (Female)
Chinese · Female
Yujie (Female)
Chinese · Female
Jingying (Male)
Chinese · Male
Tianmei (Female)
Chinese · Female
Presenter (Male)
Chinese · Male
Multilingual Narrator
Multilingual · Neutral
Bookworm
Multilingual · Neutral
Sweet Girl
Multilingual · Female
Podcast Girl
Multilingual · Female
Audiobook Male
Multilingual · Male
Calm Woman (Korean)
Korean · Female
Professional Man (Korean)
Korean · Male
Calm Woman (Japanese)
Japanese · Female
Friendly Man (Japanese)
Japanese · Male
Calm Woman (Spanish)
Spanish · Female
Calm Woman (Portuguese)
Portuguese · Female
Friendly Man (Portuguese)
Portuguese · Male
Pricing
Cost per 1,000 characters
$0.06
Typical paragraph (~500 chars)
$0.03
Billed on input text length
Capabilities
Speed
Standard: Balanced latency & quality
Emotion Control
Steer happy / sad / angry / neutral delivery
Max Input
10,000 characters per request
Voices
30 voices · 7 languages