HomeArticle

Alibaba has rolled out 5 new Qwen models in one go, slashing prices by up to 95%

智东西2026-09-24 08:17
Qwen releases a brand-new large speech model.

According to AI Technology Review's report on September 23, Qwen officially released the Qwen-Audio-3.1 series of speech large models today, launching 5 models at one time, upgraded the three core capabilities of Automatic Speech Recognition (ASR), Text-to-Speech (TTS) and real-time speech interaction (Realtime), and also launched the next-generation architecture-based audio understanding model ASR-Next and audio creation model TTS-Next.

Among them, the new-generation speech recognition large model Qwen-Audio-3.1-ASR has been tested on the public dialect test sets KeSpeech and WSYue. In 6 of the 11 subsets, it achieves better results than the two open-source models Doubao-ASR and Tencent Hy-ASR-3.0-preview, with an average Character Error Rate (CER) of only 4.55%.

▲ Public dialect test results (Source: Qwen Large Model)

In addition, the prices of all Qwen-Audio speech models have been reduced, among which TTS prices are cut by about 70%, Realtime prices by about 85%, and ASR prices by up to 95%.

In terms of pricing, Qwen-Audio-3.1-ASR-Flash costs 0.8 yuan for input and 2.7 yuan per million tokens for output; Qwen-Audio-3.1-TTS-Flash costs 1.5 yuan for input and 12 yuan per million tokens for output; Qwen-Audio-3.1-TTS-Next costs 6 yuan for input and 12 yuan per million tokens for output; Qwen-Audio-3.1-Realtime-Plus costs 5~40 yuan for input and 40~150 yuan per million tokens for output according to different tiers.

▲ Price list of Qwen-Audio-3.1 series speech models (Source: AI Technology Review illustration)

In this series, APIs for most models have been launched on the Qwen AI platform, while the API for ASR-Next is not yet available. At the 2026 Yunqi Conference which opened on September 22, Alibaba has already previewed several models of this release, and the debut simultaneous interpretation model Qwen3.8-LiveTranslate is not included in this release list.

In addition, Qwen Work also released two smart hardware products today, namely the first Agent hardware QwenNote A2 and the desktop robot Eva, and Qwen-Audio-3.1-Realtime will be gradually connected to these hardware devices.

QwenNote A2 is an AI assistant device that can be magnetically attached to the back of a mobile phone. It has built-in 4G network, can independently record and transcribe audio, perform translation without connecting to a mobile phone, and supports AI conversation, meeting summarization and proposal drafting.

The desktop robot Eva has Qwen Work built-in, which can conduct real-time voice interaction with users, complete tasks assigned by users, and support multi-device collaboration.

Experience links:

Qwen-Audio-3.1-ASR:

https://www.qianwenai.com/models/qwen-audio-3.1-asr-flash

Qwen-Audio-3.1-TTS:

https://www.qianwenai.com/models/qwen-audio-3.1-tts-flash

Qwen-Audio-3.1-TTS-Next:

https://www.qianwenai.com/models/qwen-audio-3.1-tts-next

Qwen-Audio-3.1-Realtime:

https://www.qianwenai.com/models/qwen-audio-3.1-realtime-plus

01.

Three core models upgraded:

16 dialects, cross-language synthesis, full-duplex

Qwen-Audio-3.1-ASR has five core capabilities: structured output, wide audio coverage, high recognition accuracy, fast response and clear audio processing.

Among them, Qwen-Audio-3.1-ASR adopts an end-to-end solution that jointly outputs speaker labels, timestamps and texts. It can handle turn-taking speeches, as well as preserve short interjections and overlapping voices, providing traceable structured records for meetings, interviews and dialogue analysis.

In terms of language coverage, one single model supports 30 languages and 16 Chinese dialects. In terms of vocabulary coverage, the model supports recognition of industry terms, professional entities and hierarchical hot words, and can refer to the historical context of long audio to maintain consistency of names, abbreviations and terms.

In terms of latency, under low-latency streaming recognition, the first character response of the model is about 160 milliseconds. Meanwhile, Qwen-Audio-3.1-ASR comes with a polishing function after transcription, which can automatically remove modal particles and repeated expressions and only reorganize semantics. It can automatically distinguish different speakers, realizing consistent character identity, time alignment and structured output for long audio.

In addition to public tests, in the ASR evaluation on Alibaba's self-built Chinese dialect test set, Qwen-Audio-3.1-ASR achieved an average CER of 10.38% in the original text transcription test of 16 Chinese dialects, and its recognition ability for Wenzhou dialect and Suzhou dialect is far ahead of other models.

▲ ASR evaluation results on Alibaba's self-built Chinese dialect test set (Source: Qwen Large Model)

The internally built Chinese dialect test set for AST evaluation is used to evaluate the ability to translate dialect speech into Mandarin and accurately retain the original meaning. Qwen-Audio-3.1-ASR performs best in 10 of the 11 dialect subsets, with an average semantic sentence accuracy of 82.10%.

▲ AST evaluation results on Alibaba's self-built Chinese dialect test set (Source: Qwen Large Model)

The upgrade focus of Qwen-Audio-3.1-TTS is the cross-language timbre synthesis capability. The same timbre can be naturally migrated across languages. In the official demo, Qwen-Audio-3.1-TTS can use one single timbre to speak Mandarin, Cantonese, Henan dialect, English and Japanese respectively.

In addition, the emotion, speech rate and expression mode of the synthesized speech of the model can be controlled by instructions. The same sentence can show different states such as happiness and sadness, which emphasizes more on real expression and makes the voice fit the specific content and usage scenarios.

The real-time speech interaction model Qwen-Audio-3.1-Realtime further strengthens the understanding of emotion, communication intention and context, supports real-time multi-language switching and tool calling, and enhances security and fact reliability.

In actual multi-person dialogue scenarios, full-duplex real-time interaction enables Qwen-Audio-3.1-Realtime to wait, insert and respond appropriately according to the dialogue process. Facing continuous multi-person communication, it can judge when to participate in the dialogue and when to continue waiting by combining the context.

Qwen-Audio-3.1-Realtime can adjust the response mode according to the user's tone and emotion, and continue to maintain the context and communication rhythm after switching languages in the dialogue. When the user puts forward demands such as query and search, Qwen-Audio-3.1-Realtime can also call external tools such as APIs and knowledge bases to complete the task, and then bring the results back to the dialogue.

According to the released technical report, the overall task success rate of Qwen-Audio-3.1-Realtime on the τ-Voice half-duplex speech-to-text adapted version increased from 78.4% to 82.0%; on Full-Duplex-Bench v1.5, the response rate to background speech decreased from 73.0% to 13.0%, which means the model is less interfered by irrelevant background speech.

▲ Qwen-Audio-3.1-Realtime technical report (Source: arXiv)

In addition, Qwen-Audio-3.1-Realtime is being integrated with Agents such as Qoder and Qwen Work, as well as smart hardware such as QwenNote, A2, Eva, Qwen AI glasses and Leqi AI glasses.

The official stated that in these scenarios, users do not need to keep their computers or mobile phones on all the time. They can directly initiate tasks via voice, and continue to add conditions, modify requirements or check the execution status in real-time conversations.

02.

Two new Next models:

One expands audio understanding, the other enables full audio creation

This time Alibaba also released two new audio models, Qwen-Audio-3.1-ASR-Next and Qwen-Audio-3.1-TTS-Next, which are used to supplement the audio understanding gaps beyond speech for ASR and TTS respectively.

Qwen-Audio-3.1-ASR-Next extends the sound processing range from text in speech to complete audio information. It can understand human emotions, environmental sounds and mechanical sounds, and complete sound description, event localization, audio Q&A and reasoning.

The official gave an example: for an audio that contains human dialogue, background music and environmental sounds at the same time, the model will output the dialogue text, describe what sounds are included in it, when the related events occur, and answer questions related to the entire content.

In the role-specific ASR evaluation, Qwen-Audio-3.1-ASR-Flash-Next and Qwen-Audio-3.1-ASR-Flash won five and three optimal results respectively, both outperforming the previous Fun-ASR cascade system.

Qwen-Audio-3.1-TTS-Next extends its capabilities from speech synthesis to audio creation. Based on inputs such as text, timestamps and reference audio, the model can uniformly generate human voices, sound effects and environmental sounds, and combine them into complete audio content.

In multi-round generation, the model can maintain consistent timbre for different characters, and present natural expressions such as pauses, turn-taking and emotional transitions.

Meanwhile, Qwen-Audio-3.1-TTS-Next supports fine-grained timestamp control, voice cloning and 48kHz output, which can arrange the dialogue, sound effects and content rhythm more precisely, covering scenarios such as podcasts, audio books, film and television, games and advertisements.

03.

Conclusion: Expanded capabilities, lower prices

Software and hardware provide traffic entrances

The Qwen-Audio-3.1 series of speech models continue to optimize in speech recognition, speech synthesis and real-time interaction capabilities, and the prices are further reduced. At the same time, ASR-Next and TTS-Next also extend their capability coverage to audio understanding and audio creation.

Alibaba positions speech as a natural entrance connecting people, content and AI. In the future, the popularization of new terminals such as AI glasses and desktop robots will also bring more demand for speech invocation.

This article is from the WeChat official account