Qwen3.8-Omni-Flash Undercuts Gemini on Audio and Video
Alibaba shipped an omni-modal model that takes text, images, audio in 113 languages and two-hour video into one million-token context, at $0.15 per million input tokens — five times under Gemini 3.8 Flash. Its agentic perception reads long video selectively, gaining accuracy while cutting tokens 45.7%. The weights are not open.
Alibaba released Qwen3.8-Omni-Flash on 18 September, an omni-modal model that takes text, images, audio and video into a single one-million-token context and returns text. It is built on Qwen3.8-Flash-Next, whose weights Qwen published in August, and it is priced low enough to change what kind of audio and video work is worth doing at all: $0.15 per million input tokens and $0.47 per million output, with cache hits at $0.016.
The input side is the interesting part. The model accepts audio in 113 languages and video up to two hours or 2GB, holding stable at 15 frames per second, and it handles stereo and four-channel spatial audio. QwenCloud lists the context as 991K tokens in and 131K out, with up to 262K of that available for reasoning. Thinking is on by default at reasoning_effort: xhigh and can be switched off. Function calling, web search, structured outputs, context caching and batch processing are all in the box, and the API speaks both the OpenAI and DashScope protocols.
Against Gemini 3.8 Flash, the model Qwen is plainly aiming at, the list rates are five times cheaper on input and eight times cheaper on output — and that gap widens on 1 January 2027, when Google's introductory pricing expires and its rates double. The comparison is not quite like for like: Google bills audio input above its text rate, while Qwen quotes a single rate across modalities, so for audio-heavy work the real spread is wider still. Qwen's own accounting says hourly audio ingestion now costs over 98% less than on Qwen3.5-Omni-Plus, audio-visual input over 93% less, and video roughly 89% less per hour.
The capability claim that deserves the most attention is not a benchmark score but an architectural one Qwen calls agentic perception. Rather than walking a long video linearly, the model decides which segments are worth attending to. On OmniVideoBench that shows up as accuracy rising from 63.4 to 67.8 while token consumption fell 45.7%, from 145,736 to 79,117. Getting more accurate while reading less is the rare kind of result that is cheap to verify and hard to fake, and it is the mechanism that makes the two-hour video ceiling more than a spec-sheet number.
The rest of the numbers are Qwen's own and have no independent validation yet: an average gain of more than 25% over Qwen3.5-Omni-Plus across 29 evaluations, +36.5 points on WildClawBench-MM, +22.3 on AgenticVBench, 69.6 on UniClawBench. Qwen describes the model's audio-visual performance as close to Gemini 3.8 Flash and its audio as above it. Treat all of that as a vendor's opening position. Third-party coverage is thin — BenchLM has scored it on 19 benchmarks out of 446 and declines to rank it, with 92.6 on LiveCodeBench v6 and 91.8 on MathVision at one end and 36.5 on HLE at the other.
The notable absence is the weights. Omni-Flash is API-only at launch, served from Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia, with no self-hosting option — a real break from the company whose open releases built its reputation, and the second such break in two months. When Alibaba opened Qwen3.8-Max in August it shipped not under Apache 2.0 but under a bespoke qwen3.8-max licence, the first Qwen flagship not to use the permissive terms. What Qwen did open here is tooling rather than a model: Qwen-MM-Plugins, Apache 2.0, a set of multimodal skills including omni-memory, omni-video2note and omni-chatcut that plug into Claude Code, Gemini CLI and Qwen's own CLI. That is a telling shape for a release — give away the thing that drives usage, meter the thing that does the work.
More on Gemini
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.