Models·3 min read
By BitsMindsSource: Qwen (Alibaba Cloud)

Alibaba's Qwen3.7-Max Lands With a 1M-Token Context, AA Index 56.6, and a 35-Hour Agent Run

Qwen3.7-Max-Preview ships with a 1M-token context, extended-thinking mode, and benchmark gains across CritPt, Humanity's Last Exam, and Terminal-Bench Hard — clearing Gemini 3.5 Flash on the AA Intelligence Index.

Alibaba's Qwen3.7-Max Lands With a 1M-Token Context, AA Index 56.6, and a 35-Hour Agent Run
Share:

Alibaba's Qwen team formally unveiled Qwen3.7-Max-Preview at the Alibaba Cloud Summit on May 20, 2026, two days after it landed on the company's API platform and roughly a week after it stealth-debuted on the LM Arena leaderboard. The proprietary reasoning model doubles the prior generation's context window to 1 million tokens, ships an extended-thinking mode that can generate roughly 97 million reasoning tokens during evaluation, and is pitched explicitly as an agentic workhorse capable of sustaining "hundreds or even thousands" of tool calls in a single run.

On the Artificial Analysis Intelligence Index v4.0, Qwen3.7-Max scores 56.6 — a 4.8-point jump over Qwen3.6 Max Preview (51.8), and the highest mark ever posted by a Chinese model on that leaderboard. It sits fifth overall, ahead of Google's Gemini 3.5 Flash (55.3) and within striking distance of Gemini 3.1 Pro Preview (57.2) and Anthropic's Claude Opus 4.7 (57.3). OpenAI's GPT-5.5 still leads the pack at 60.2.

Qwen3.7-Max overall benchmark comparison chart

The gains are concentrated where they matter most for autonomous workflows. CritPt climbed 9.7 percentage points (3.7% to 13.4%), Humanity's Last Exam jumped 9.2 points (28.9% to 38.1%), and Terminal-Bench Hard — a brutal proxy for shell-driven agent reliability — rose 6.9 points to 50.8%. On coding-specific benchmarks the model claims Terminal Bench 2.0-Terminus at 69.7, SWE-Verified at 80.4, and SWE-Multilingual at 78.3, putting it inside the frontier band on every dimension that matters to a real software-engineering agent.

Qwen3.7-Max coding benchmark performance

Agentic capability is the headline pitch. On MCP-Mark Qwen3.7-Max scores 60.8 and on MCP-Atlas 76.4, and on SpreadSheetBench-v1 it lands an 87.0. Alibaba's own demos lean even harder: a 35-hour autonomous kernel-optimization run reportedly produced a 10x inference speedup, and the model is said to have sustained more than 1,000 sequential tool calls without falling out of its task. Those numbers come from internal testing and have not been independently verified, but the trajectory matches what frontier US labs have been claiming since GPT-5 and Claude Opus 4 shipped.

Reasoning is the other place where Qwen3.7-Max distinguishes itself from earlier open-weight flagships out of China. GPQA Diamond hits 92.4 — ahead of Claude Opus 4.6 Max's 91.3 — and on the Apex reasoning benchmark Qwen3.7-Max posts 44.5 against DeepSeek V4 Pro's 38.3. HMMT 2026 Feb climbs to 97.1. The trade-off, flagged by Artificial Analysis, is that the model now abstains more often on the AA-Omniscience knowledge battery (attempt rate fell from 67.3% to 48.0%) — a deliberate choice to say "I don't know" rather than confidently hallucinate.

Qwen3.7-Max reasoning and STEM benchmark performance

Access is rolling out via Alibaba Cloud Model Studio (DashScope) with OpenAI- and Anthropic-compatible endpoints, and the chat interface is live at chat.qwen.ai. Public pricing on OpenRouter currently lists $2.50 per million input tokens and $7.50 per million output tokens — a notable premium over Qwen3.6 Max Preview ($1.30 / $7.80) but still well below the Western frontier. The closed-weights "Preview" tag means terms and behavior may shift, and there is no open release planned. Even so, Qwen3.7-Max is the clearest signal yet that the gap between US and Chinese frontier labs is now measured in months and tenths of an Index point — not generations.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Claude Haiku 5.5 vs GPT-6 Luna: the BitsMinds Lab build-off A terracotta disc carved into a labyrinth, with the Claude mark inlaid at its centre, faces a brushed silver disc with the OpenAI mark inlaid over the shading of a crescent moon. A VS badge sits between them, the model names are painted below, and the line above reads Same briefs, same price. BITSMINDS LAB · SAME BRIEFS, SAME PRICE VS CLAUDE HAIKU 5.5 GPT-6 LUNA BITSMINDS.COM
Models

Claude Haiku 5.5 vs GPT-6 Luna: Lost in Thought

Mistral Large 4, le Chonk The pixel-block Mistral logo cast as a thick, heavy slab with deep extruded sides, its face banded yellow to orange to red, standing on a dark warm floor under a single light, above a label reading Large 4, 1.05T parameters, 52B active. MISTRAL LARGE 4 1.05T PARAMS · 52B ACTIVE BITSMINDS.COM
Models

Mistral Large 4: Europe’s 1-Trillion-Parameter Open Model

Claude Haiku 5.5: Anthropic's fastest model, at a tenth of Haiku 4.5's price A terracotta and ivory stopwatch, tipped to show its depth, has the official Claude asterisk inlaid in clay as the hub of its green sweep hand. A paper tag tied to the crown reads minus 90 percent against Haiku 4.5. Beside it, the title names Claude Haiku 5.5 and quotes its API prices for prompts up to 100,000 tokens: 10 US cents per million input tokens and 50 cents per million output tokens. The stopwatch is an editorial metaphor for Anthropic's claim that Haiku 5.5 is its fastest model at standard speed, not an Anthropic product; the 90 percent figure applies to prompts up to 100,000 tokens, and longer prompts cost five times as much. 51015202530354045505560 CLAUDE HAIKU 5.5 PER-TOKEN PRICE −90% VS HAIKU 4.5 PROMPTS UP TO 100K CLAUDE Haiku 5.5 $0.10 INPUT $0.50 OUTPUT PER MILLION TOKENS BITSMINDS.COM
Models

Claude Haiku 5.5 Is Out at a Tenth of Haiku 4.5's Price