Alibaba’s Tongyi Lab has released Qwen-Audio-3.0-TTS, a production-oriented text-to-speech (TTS) system. The model ships in two variants from the same lineage. Flash targets real-time interaction.
Vocoder(GPU): vocoder 将生成出来的 codec token 映射回连续音频波形,和 audio encoder 恰好相反。 逻辑上,单次调用 vocoder 的计算通常很轻,但在高并发下,多条 AR loop 可能同时结束、一堵塞在 vocoder 的队列中。 不同 vocoder 的 streaming 行为也各不相同。
Please use our dedicated channels for questions and discussion. Help is much more valuable if it's shared publicly so that more people can benefit from it.
Some results have been hidden because they may be inaccessible to you
Show inaccessible results