Models·2 min read
By BitsMindsSource: Mistral AI

Mistral Medium 3.5 Lands With Cloud Coding Agents and 77.6% on SWE-Bench

Mistral fuses chat, reasoning, and coding into a single 128B dense model and pairs it with Vibe — async cloud-based coding agents that hand back work as pull requests instead of terminal output.

Mistral Medium 3.5 Lands With Cloud Coding Agents and 77.6% on SWE-Bench
Share:

Mistral AI on May 2 released Mistral Medium 3.5, a dense 128-billion-parameter model with a 256k context window that the French lab is positioning as its first true flagship — a single set of weights designed to handle instruction following, long-form reasoning, and coding without forcing developers to swap between specialist models. The release ships alongside Vibe Remote Agents, a new cloud execution layer that pulls coding work off the local machine and returns results as reviewable pull requests.

On benchmarks, Medium 3.5 hits 77.6% on SWE-Bench Verified, ahead of Mistral's own Devstral 2 and Alibaba's Qwen 3.5 397B A17B. It also posts a 91.4 on the τ³-Telecom agentic score, a test that grades reliability across multi-step tool calls. Reasoning effort is now configurable per API request, so a quick chat reply and a long-horizon agent run can use the exact same model with different compute budgets — a setup that mirrors recent moves from OpenAI and Anthropic to expose reasoning depth as a parameter rather than a separate SKU.

Vibe is the more strategically interesting half of the announcement. Coding agents launched from the Mistral Vibe CLI or directly from Le Chat now run asynchronously in sandboxed cloud environments, with full session teleportation from local to cloud and integrations across GitHub, Linear, Jira, Sentry, Slack, and Teams. Mistral is pitching the workflow as: kick off a refactor, dependency upgrade, or CI investigation, walk away, and review the resulting PR when the notification arrives. The same model also powers a new Work mode in Le Chat (preview) for cross-tool agentic tasks across email, calendar, and documents, with explicit approval gates before any sensitive action fires.

API pricing comes in at $1.50 per million input tokens and $7.50 per million output tokens, which puts Medium 3.5 well below GPT-5.5 and Claude Opus 4.7 on cost while undercutting most frontier-tier coding APIs. The model is available through Mistral's Pro, Team, and Enterprise plans, hosted on NVIDIA's build.nvidia.com, and shipped as an NVIDIA NIM container. Open weights are on Hugging Face under a modified MIT license, and Mistral says the model can be self-hosted on as few as four GPUs — an unusually accessible footprint for something competing in the agentic-coding tier.

The release reads as Mistral's clearest answer yet to the question of where it fits in a market dominated by OpenAI, Anthropic, and Google. Rather than chasing the largest possible base model, the lab is bundling a competitive coding model with cloud agent infrastructure and an open license — betting that the developers who care about self-hosting and per-request reasoning control are a market worth owning outright.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Many kinds of input, one Qwen context On a deep violet field, a filmstrip, an audio waveform, a landscape photograph and a text page curve toward a single luminous sphere bearing the official purple Qwen star. A label below reads “1M context”, representing text, images, audio and video sharing one million-token context window. Aa 1M context BITSMINDS.COM
Models

Qwen3.8-Omni-Flash Undercuts Gemini on Audio and Video

Hill Climb: GPT-6 Astra versus Claude Fable 5.1 A teal desert buggy climbs a sandstone ridge on the left. An orange jeep climbs a green hill on the right, under an arc of gold coins. The two landscapes meet at a diagonal divide. 7 GPT-6 ASTRA CLAUDE FABLE 5.1 VS HILL CLIMB BITSMINDS.COM
Models

GPT-6 Astra vs Claude Fable 5.1: Hill Climb

Gemini 3.8 Live: thinking while the conversation continues An editorial illustration in a dark blue and violet room. A carefully drawn studio microphone and a tilted smartphone flank a translucent speech bubble carrying the multicoloured Gemini star. A continuous luminous audio waveform travels between them. Above the bubble, a separate arc connects small search, reasoning and completion symbols, representing background work continuing during a spoken conversation. The phone screen and visual paths are conceptual, not a reproduction of Google's actual interface or internal reasoning. Original vector illustration for BitsMinds, gemini-3-8-live-extended-thinking-voice. 16 September 2026. GEMINI 3.8 LIVE EXTENDED THINKING Gemini 3.8 LIVE The conversation continues Thinking. Still talking. BITSMINDS.COM
Models

Gemini 3.8 Live Tops Voice AI and Undercuts GPT