Alibaba's Qwen team released Qwen-Audio 3.1 on September 23, 2026, a lineup of five speech models spanning recognition, synthesis, and real-time interaction, and paired the launch with steep price cuts across its audio APIs. Speech recognition pricing drops by up to 95 percent, text-to-speech by roughly 70 percent, and the real-time voice model by about 85 percent, according to the company and multiple outlets that covered the release. The models are the headline, but the pricing is the strategy. This piece separates the two, looks at what the new stack actually does, and asks what a cut of this size does to the economics of building with voice.
What actually shipped
Qwen-Audio 3.1 is not a single model. It is five, and the split matters. Three are upgrades to existing capabilities: automatic speech recognition (ASR), which turns speech into text; text-to-speech (TTS), which does the reverse; and a Realtime model built for low-latency, back-and-forth voice conversations. Two are new. TTS-Next is aimed at audio creation, and ASR-Next at deeper audio understanding.
The upgraded ASR now cleans up filler words and repetitions automatically and improves recognition across languages and dialects, per Alibaba's release notes as reported by The Decoder. ASR-Next goes further into analysis rather than transcription: it layers in multi-speaker identification with timestamps, emotion detection, and the ability to flag ambient sounds and machine noise. That is a shift from "what words were said" toward "who said them, how, and against what background."
TTS-Next is the more architecturally interesting of the two additions. It pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass, rather than synthesizing speech and layering audio separately. The practical pitch is one call that produces a finished audio scene instead of a clean but bare voice track.
The models are the headline, but the pricing is the strategy.
On the Qwen-Audio 3.1 release
The price cut is the actual news
Model launches arrive weekly. A price cut of up to 95 percent on a core API does not. To see why it matters, it helps to hold the three reductions side by side, because they are not uniform.
The deepest cut lands where differentiation is thinnest
Reported maximum price reduction on each Qwen-Audio 3.1 API, in percent. The cut is largest on speech recognition, the most commoditized of the three, and smaller on synthesis and real-time voice, where quality and latency still separate providers.
Figures are reported "up to" reductions; the actual saving depends on the specific model, tier, region and volume.
The pattern is worth reading. The deepest cut lands on ASR, the most commoditized of the three. Speech-to-text is a mature capability with many capable providers, open and closed, so price is close to the only lever left. TTS and the Realtime model, where quality and latency still separate providers, get smaller cuts. In other words, Alibaba discounted hardest where differentiation is thinnest and held more value back where its models can still compete on merit. That is a deliberate shape, not a flat sale.
There is a second signal in who benefits. A 95 percent cut does little for a hobbyist already inside a free tier. It matters enormously to anyone running voice at volume: call-center transcription, meeting capture, media subtitling, voice interfaces shipped inside consumer hardware. Coverage of the launch tied it explicitly to a push for AI creation and a hardware ecosystem, which lines up with the economics. Voice is the natural interface for earbuds, speakers, wearables, and appliances, and the cost of the model is what decides whether those devices can afford always-on speech.
The commoditization pattern, continued
This is not an isolated move. Chinese labs have spent 2026 competing aggressively on the price of inference across text, image, and now audio, and the direction of travel for voice APIs has been steadily downward. Each cut resets the floor for everyone else, because a developer comparing providers sees the same capability at a fraction of the cost and re-runs the math.

For the incumbents that built businesses on premium voice, the pressure is specific rather than existential. A cut like this does not erase the reasons a team chooses a particular TTS voice, a specific latency profile, or a compliance posture. What it does is shrink the price umbrella those reasons have to justify. When the gap between a premium voice API and a cut-price one widens from small to large, every buyer has to decide whether the premium still buys something they need. Some will say yes. The number who say yes falls as the gap grows.
What it changes for people building with voice
For a team shipping a product rather than tracking the model race, the useful takeaway is not "switch to Qwen." It is that the cost of voice is falling fast enough that it should not anchor a product decision. A feature that looked too expensive to run at scale six months ago may pencil out now, and may look different again in another six months as the next provider responds.
That volatility is the argument for staying flexible at the model layer. When ASR, TTS, and real-time voice are converging on commodity pricing from several providers at once, binding a product to one vendor's audio stack trades away the ability to follow the next price cut or quality jump. A model-agnostic approach, the kind Metir AI takes by routing across models from multiple providers, treats a launch like this one as an option to evaluate rather than a migration to execute. The provider that is cheapest or best this quarter is unlikely to be the one that is cheapest or best next quarter, and the teams that keep their choices open are the ones positioned to benefit either way.
The honest caveats
Two things are worth holding in view. First, headline price cuts are quoted as "up to," and the deepest number rarely applies to every tier or region. The real saving for a given workload depends on volume, the specific model, and where it runs, so the 95 percent figure is a ceiling, not a promise. Second, price is one axis. Latency, voice quality, language coverage, data-handling terms, and reliability all shape whether a voice API is right for a given use, and none of those is captured in a discount. The right read of Qwen-Audio 3.1 is that it lowers the cost of entry to serious voice work and raises the pressure on everyone pricing above it. What it does not do is settle, on price alone, which stack a given product should run.
Sources:
- Alibaba launches Qwen-Audio 3.1 with five new models and slashes AI audio prices by up to 95 percent | The Decoder
- Alibaba Ships Qwen-Audio 3.1 Stack, Cuts Voice APIs Up to 95% | AI Weekly
- Alibaba's Qwen-Audio 3.1 Slashes Voice API Prices by up to 95% | AlphaSignal
- Alibaba Slashes Voice AI Model Prices by Up to 95% in Push for AI Creation and Hardware Ecosystem | BigGo Finance
- Qwen-Audio-3.1: five speech models and API price cuts | AI/TLDR
Image credits
Header and in-body image: Alibaba Group headquarters, Hangzhou, by Thomas LOMBARD (Thecraft), via Wikimedia Commons, CC BY-SA 3.0. Illustrative of Alibaba as the company behind the Qwen models; not a depiction of the audio product.