Alibaba's Qwen-Audio-3.0-TTS is Here

PromptCube3.com Novice 19h ago 234 views 7 likes 1 min read

via aibase.com
Key points
  • The Plus version currently ranks first on the Artificial Analysis Speech Arena, beating Gemini and ElevenLabs.

  • The Flash version achieves a low first-packet latency of 300ms for real-time interactions.

  • It supports 16 languages and 20 Chinese dialects with high speaker similarity.
  • Alibaba's Qwen-Audio-3.0-TTS is Here

    This is a significant leap for TTS technology, moving beyond simple text-to-speech into genuine "expressive" audio. What stands out to me is the combination of low latency (300ms) and the "Free-style" natural language instructions. Being able to tell a model to sound like a "livestream salesperson" without complex manual tagging makes it incredibly accessible for developers.

    The addition of fine-grained tags for gasps and laughter, paired with the ability to handle noisy reference audio, suggests this is aimed squarely at high-end content creation like gaming and audiobooks. Seeing it outperform ElevenLabs in certain benchmarks proves that the gap in naturalness is closing rapidly. It's a strong move toward seamless, human-like AI voice agents.

    Alibaba

    All Replies (0)

    No replies yet — be the first!

    Write a Reply

    Markdown supported