Alibaba’s Qwen team released Qwen3.8-Omni-Flash, the family’s first multimodal model built specifically for AI agents, and priced it at a fifth of Google’s comparable offering. On audio-video benchmarks, Qwen says the model comes close to matching Gemini 3.8 Flash — while charging $0.15 per million input tokens and $0.47 per million output, against Gemini 3.8 Flash’s introductory $0.75/$3.75. Google’s rate is set to double on 1 January 2027; Qwen’s is not. The release continues Alibaba’s run of open-weight challenges to the US frontier, this time aimed at the agent tier rather than the flagship league.
The price gap widens at the edges where agents actually operate. Qwen estimates audio input at under one US cent per hour, and 720p video with audio at one frame per second runs about $0.20 — before response costs. For a coding-agent workflow that watches screen recordings, listens to meeting audio, and ingests an hour of footage, the difference between fractions of a cent and real money is the difference between an experiment and a default setting.
What the model actually does
Qwen3.8-Omni-Flash processes audio and video together rather than as separate modalities bolted onto a text model. It draws conclusions across both streams and uses tools on its own to edit vlogs, translate short videos, or summarise feature-length films. The context window spans one million tokens, enough to hold a multi-hour recording or a large codebase with media attached.
The agent framing is where Alibaba has done the most work. Open-source Qwen-MM-Plugins add video editing, speaker recognition, PDF video notes, and reusable workflows to existing coding agents — Claude Code, Gemini CLI, and Qwen Code. Qwen-Live Harness enables real-time interaction through a camera and microphone, putting the model in the loop of a live video feed rather than a file upload. The model is available through Qwen Studio, Qwen Cloud, and the API.
The pricing pressure is structural, not promotional
This is the second time this quarter Chinese labs have forced a price reckoning on multimodal models. DeepSeek’s permanent 75 percent discount on its flagship made cheap-and-good the baseline expectation for text models, and Alibaba’s open-weight releases have repeatedly compressed the gap between Chinese and US frontier models to weeks. Extending that playbook to audio-video agents matters because multimodal capability is where the next wave of agentic products lives — and where Google has been charging its steepest premium.
The comparison comes with caveats. “Close to matching” is Qwen’s own characterisation, published alongside its own benchmark selections, and Google’s Flash tier is the introductory rate for a model that has been iterating on a faster cadence than any competitor — three Flash releases in six weeks by early September. Independent evaluations of the audio-video tasks, particularly on the messy end-to-end jobs like vlog editing, will be the real test.
Why it matters for what you use
For anyone building on AI agents, the practical change is that multimodal input stops being priced like a luxury. A workflow that transcribes and watches a full day of footage currently costs more than most teams will spend; at Qwen’s rates it rounds to zero. Combined with a million-token context and plugins that work inside agents developers already run, the model’s significance is less the benchmark deltas and more the default it sets: the second-largest AI economy on earth has decided multimodal agents are a commodity market, and it is pricing accordingly.
Google’s response will be worth watching. Its Flash line has taken the role of the volume product that undercuts everyone else, and it now finds itself undercut by roughly five to one on identical benchmark claims, with a scheduled price doubling arriving two weeks into 2027. The last time a Chinese lab opened a gap this wide, the response took months — and by then the workflow default had already moved.