Skip to content
Rush Commerce
AI & Automation4 min read

Qwen-Audio 3.1 cuts voice API prices up to 95%

Alibaba shipped five Qwen-Audio 3.1 models and cut ASR pricing up to 95%, TTS ~70% and realtime ~85%. Your voice stack just got a cheaper competitor.

The cheapest part of a voice agent used to be the part you did not build. That is changing fast. Alibaba's Qwen team shipped Qwen-Audio 3.1 — five models covering recognition, synthesis and realtime conversation — and cut prices across the lineup by up to 95%. If you are paying a US vendor per minute of speech, the Qwen-Audio 3.1 price cut is the number your next renewal negotiation should start from.

What actually happened

Per The Decoder, the September 23 release upgrades the three existing models — ASR, TTS and Realtime — and adds two new ones:

  • ASR-Next for audio understanding: multi-speaker labeling with timestamps, emotion detection, and recognition of ambient and machine sounds.
  • TTS-Next for audio creation: voice, sound effects and background audio generated in a single pass, aimed at audiobooks, podcasts, games and ads.

The price moves, consistent across independent reports: ASR down up to 95%, TTS down roughly 70%, realtime voice down roughly 85%. 36Kr published per-model list pricing in yuan — ASR-Flash at 0.8 in / 2.7 out per million tokens, TTS-Flash at 1.5 / 12, TTS-Next at 6 / 12, and Realtime-Plus ranging 5–40 in / 40–150 out by tier. Those are Alibaba Cloud list prices in China; convert and confirm against the region you would actually deploy in before you put them in a spreadsheet.

Capability moved too, not just price. The recognition model handles 30 languages and 16 Chinese dialects and strips filler words and repetitions automatically — the post-processing step most transcription pipelines currently pay a second vendor for.

Why cheap speech models matter for your business

Price cuts of this size reprice the whole category, not just one vendor. You do not need to deploy Qwen for this release to matter. You need to know the number exists when your incumbent's renewal lands. A 70% cut on synthesis from a credible competitor is leverage, and it costs you one email to use.

Falling ASR cost changes what is worth transcribing. At old prices, you transcribed the calls someone flagged. At a twentieth of that, you transcribe everything — every support call, every voicemail, every sales conversation — and the transcript becomes a searchable asset rather than a line item. Speaker labels with timestamps are what make that corpus queryable instead of a wall of text.

Do not swap vendors on price alone. Alibaba Cloud is the deployment path, and that carries real questions: where audio is processed, what your customer contracts say about data location, and what your industry's rules require. For a Phoenix services business recording customer calls, that is a conversation with your counsel before it is a conversation with your CTO. Answer it first; the discount is still there afterward.

Build the seam now. Put speech behind one internal interface — transcribe(audio) -> text, speak(text) -> audio — with the provider as configuration. We do this on every voice build, and the reason is exactly this week: the pricing under you moves by an order of magnitude and you want that to be a config change, not a refactor.

Key takeaways

  • Qwen-Audio 3.1 shipped September 23 with five models: upgraded ASR, TTS and Realtime plus new ASR-Next and TTS-Next
  • Reported cuts: ASR up to 95%, TTS roughly 70%, realtime voice roughly 85%
  • ASR-Next adds speaker labeling with timestamps, emotion detection and ambient sound recognition; TTS-Next generates voice, effects and background audio in one pass
  • Per-model yuan list pricing published by 36Kr — verify against your deployment region before budgeting
  • Recognition covers 30 languages and 16 Chinese dialects and removes filler words automatically
  • Use the number as renewal leverage with your current vendor even if you never deploy Qwen
  • Check data residency and contract terms before switching, and keep speech behind one internal interface so the provider stays a config value

Voice pricing moved 20x in a day — your architecture should not care. We build voice and transcription pipelines behind interfaces you own, so swapping providers is a config change and your call data stays where your contracts say it stays. See how we build vendor-agnostic systems, or send us your current voice bill.

Sources: The Decoder, 36Kr.

  • #qwen
  • #voice-ai
  • #ai-costs
  • #speech-recognition
  • #vendor-risk
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.