Alibaba's Qwen-Audio-3.0-TTS is Here
This is a significant leap for TTS technology, moving beyond simple text-to-speech into genuine "expressive" audio. What stands out to me is the combination of low latency (300ms) and the "Free-style" natural language instructions. Being able to tell a model to sound like a "livestream salesperson" without complex manual tagging makes it incredibly accessible for developers.
The addition of fine-grained tags for gasps and laughter, paired with the ability to handle noisy reference audio, suggests this is aimed squarely at high-end content creation like gaming and audiobooks. Seeing it outperform ElevenLabs in certain benchmarks proves that the gap in naturalness is closing rapidly. It's a strong move toward seamless, human-like AI voice agents.
All Replies (0)
No replies yet — be the first!
