Alibaba’s Tongyi Lab has released Qwen-Audio-3.0-TTS, a production-oriented text-to-speech (TTS) system. The model ships in two variants from the same lineage. Flash targets real-time interaction.
Vocoder(GPU): vocoder 将生成出来的 codec token 映射回连续音频波形,和 audio encoder 恰好相反。 逻辑上,单次调用 vocoder 的计算通常很轻,但在高并发下,多条 AR loop 可能同时结束、一堵塞在 vocoder 的队列中。 不同 vocoder 的 streaming 行为也各不相同。
Please use our dedicated channels for questions and discussion. Help is much more valuable if it's shared publicly so that more people can benefit from it.