Models·3 min read
By BitsMindsSource: MarkTechPost

Qwen3.8-Omni-Flash Undercuts Gemini on Audio and Video

Alibaba shipped an omni-modal model that takes text, images, audio in 113 languages and two-hour video into one million-token context, at $0.15 per million input tokens — five times under Gemini 3.8 Flash. Its agentic perception reads long video selectively, gaining accuracy while cutting tokens 45.7%. The weights are not open.

Many kinds of input, one Qwen context On a deep violet field, a filmstrip, an audio waveform, a landscape photograph and a text page curve toward a single luminous sphere bearing the official purple Qwen star. A label below reads “1M context”, representing text, images, audio and video sharing one million-token context window. Aa 1M context BITSMINDS.COM
Share:

Alibaba released Qwen3.8-Omni-Flash on 18 September, an omni-modal model that takes text, images, audio and video into a single one-million-token context and returns text. It is built on Qwen3.8-Flash-Next, whose weights Qwen published in August, and it is priced low enough to change what kind of audio and video work is worth doing at all: $0.15 per million input tokens and $0.47 per million output, with cache hits at $0.016.

The input side is the interesting part. The model accepts audio in 113 languages and video up to two hours or 2GB, holding stable at 15 frames per second, and it handles stereo and four-channel spatial audio. QwenCloud lists the context as 991K tokens in and 131K out, with up to 262K of that available for reasoning. Thinking is on by default at reasoning_effort: xhigh and can be switched off. Function calling, web search, structured outputs, context caching and batch processing are all in the box, and the API speaks both the OpenAI and DashScope protocols.

List price per million tokensUS dollars · QwenCloud and Google list rates, September 2026Qwen3.8-Omni-FlashGemini 3.8 Flash012340.10.8Input0.53.8Output
Both are the published text-token rates. Google’s figures hold until 31 December 2026 and then double, to $1.50 and $7.50. Gemini bills audio input above its text rate; Qwen quotes one rate for every modality. Data: QwenCloud, Google.

Against Gemini 3.8 Flash, the model Qwen is plainly aiming at, the list rates are five times cheaper on input and eight times cheaper on output — and that gap widens on 1 January 2027, when Google's introductory pricing expires and its rates double. The comparison is not quite like for like: Google bills audio input above its text rate, while Qwen quotes a single rate across modalities, so for audio-heavy work the real spread is wider still. Qwen's own accounting says hourly audio ingestion now costs over 98% less than on Qwen3.5-Omni-Plus, audio-visual input over 93% less, and video roughly 89% less per hour.

The capability claim that deserves the most attention is not a benchmark score but an architectural one Qwen calls agentic perception. Rather than walking a long video linearly, the model decides which segments are worth attending to. On OmniVideoBench that shows up as accuracy rising from 63.4 to 67.8 while token consumption fell 45.7%, from 145,736 to 79,117. Getting more accurate while reading less is the rare kind of result that is cheap to verify and hard to fake, and it is the mechanism that makes the two-hour video ceiling more than a spec-sheet number.

The rest of the numbers are Qwen's own and have no independent validation yet: an average gain of more than 25% over Qwen3.5-Omni-Plus across 29 evaluations, +36.5 points on WildClawBench-MM, +22.3 on AgenticVBench, 69.6 on UniClawBench. Qwen describes the model's audio-visual performance as close to Gemini 3.8 Flash and its audio as above it. Treat all of that as a vendor's opening position. Third-party coverage is thin — BenchLM has scored it on 19 benchmarks out of 446 and declines to rank it, with 92.6 on LiveCodeBench v6 and 91.8 on MathVision at one end and 36.5 on HLE at the other.

The notable absence is the weights. Omni-Flash is API-only at launch, served from Beijing, Singapore, Hong Kong, Tokyo, Frankfurt and Virginia, with no self-hosting option — a real break from the company whose open releases built its reputation, and the second such break in two months. When Alibaba opened Qwen3.8-Max in August it shipped not under Apache 2.0 but under a bespoke qwen3.8-max licence, the first Qwen flagship not to use the permissive terms. What Qwen did open here is tooling rather than a model: Qwen-MM-Plugins, Apache 2.0, a set of multimodal skills including omni-memory, omni-video2note and omni-chatcut that plug into Claude Code, Gemini CLI and Qwen's own CLI. That is a telling shape for a release — give away the thing that drives usage, meter the thing that does the work.

More on Gemini

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Hill Climb: GPT-6 Astra versus Claude Fable 5.1 A teal desert buggy climbs a sandstone ridge on the left. An orange jeep climbs a green hill on the right, under an arc of gold coins. The two landscapes meet at a diagonal divide. 7 GPT-6 ASTRA CLAUDE FABLE 5.1 VS HILL CLIMB BITSMINDS.COM
Models

GPT-6 Astra vs Claude Fable 5.1: Hill Climb

Gemini 3.8 Live: thinking while the conversation continues An editorial illustration in a dark blue and violet room. A carefully drawn studio microphone and a tilted smartphone flank a translucent speech bubble carrying the multicoloured Gemini star. A continuous luminous audio waveform travels between them. Above the bubble, a separate arc connects small search, reasoning and completion symbols, representing background work continuing during a spoken conversation. The phone screen and visual paths are conceptual, not a reproduction of Google's actual interface or internal reasoning. Original vector illustration for BitsMinds, gemini-3-8-live-extended-thinking-voice. 16 September 2026. GEMINI 3.8 LIVE EXTENDED THINKING Gemini 3.8 LIVE The conversation continues Thinking. Still talking. BITSMINDS.COM
Models

Gemini 3.8 Live Tops Voice AI and Undercuts GPT

Atria Dawn Preview 744B · MIT · shipped with no announcement BITSMINDS.COM
Models

Atria Dawn: A 744B Agent Model, Shipped in Silence