AI News

Latest updates from the world of artificial intelligence

CLAUDE FABLE 5MYTHOS-CLASS, NOW PUBLICSWE-BENCH PRO80.3agentic coding · vs 69.2 OpusPRICE / 1M TOKENS$10 · $50input / outputCYBER ATTACK SUCCESS5.4%guardrails on · lower is saferAVAILABLE NOWAPI · CopilotPro · Max · Team · EnterpriseBITSMINDS.COMSource: Anthropic
Models

Anthropic Launches Claude Fable 5: Its Most Powerful Public Model Yet, With the Cyber Edges Sanded Off

Anthropic released Claude Fable 5 — the first publicly available model in its powerful Mythos class — scoring 80.3 on SWE-Bench Pro (vs 69.2 for Opus 4.8), priced at $10/$50 per million tokens, with classifier guardrails that block cyber, bio, and distillation requests and fall back to Opus 4.8.

Anthropic

Read more →
RUMOREDCLAUDEFABLE 5PUBLIC MYTHOS?claude-fable-5 · rumoredUNVERIFIEDMODEL CARD · LEAKEDclaude-fable-5ArchitectureClaude 5Long contextextendedMulti-turnimprovedCyber toolingExploit devUNLOCKINGwith guardrailsBITSMINDS.COMReporting: Alex Heath
Models

Claude Fable 5: The Internet Is Sure It Ships Today. Anthropic Hasn't Said a Word.

Prediction markets put it near 94%, a tech reporter says it lands June 9, and internal checkpoints named claude-fable-5 have surfaced — but Anthropic has confirmed nothing. Here is what the noise around Claude Fable 5, the rumored public release of its Mythos model, actually amounts to.

Reporting: Alex Heath; Polymarket

Read more →
Xiaomi Pushes a 1-Trillion-Parameter Model Past 1,000 Tokens a Second — on Eight Off-the-Shelf GPUs
Models

Xiaomi Pushes a 1-Trillion-Parameter Model Past 1,000 Tokens a Second — on Eight Off-the-Shelf GPUs

MiMo-V2.5-Pro-UltraSpeed broke the 1,000 tokens-per-second barrier on a one-trillion-parameter model using a single eight-GPU commodity node — roughly 15x faster than GPT-5.5 or Claude Opus, and without a line of custom silicon.

MarkTechPost

Read more →
OPEN MODELS · CODING · 1M CONTEXTMINIMAX M3 · JUNE 1MiniMax M3Open-weight · native multimodal · MSA attention~1/20 compute at 1M ctx · 9× faster prefill · 15× decodeBENCHMARK SCORECARD · MINIMAX-REPORTEDSWE-Bench Pro59%Terminal-Bench 2.166%SWE-fficiency34.8%BrowseComp83.5BROWSECOMPBeats Opus 4.783.5 vs 79.3Weights + technical report promised within 10 days · scores not yet independently verifiedBITSMINDS.COMSource: The Decoder · MiniMax
Models

MiniMax M3 Lands as an Open-Weight, Million-Token Coding Model That Claims to Edge Out GPT-5.5

The Chinese lab says its new open-weight model pairs frontier coding, a 1M-token context window and native multimodality on a sparse-attention architecture that cuts long-context compute 20x — but the weights and technical report are still days away, so every number is company-reported.

The Decoder

Read more →
REPORTEDLY LEAKEDCLAUDEOCEANUSANTHROPIC · MYTHOS LINEclaude-oceanus-v1-p · surfaced June 3 · unconfirmedBITSMINDS.COMSource: Cyber Security News
Models

Anthropic's Unreleased 'Claude Oceanus' Reportedly Leaked Through a Chinese Proxy Hours After Reaching Red Teamers

A model identifier, claude-oceanus-v1-p, surfaced in Anthropic's console on June 3 and went to vetted red-team testers — then, within hours, an unknown actor allegedly resold access through a Chinese API proxy. None of it is officially confirmed, but it offers an early look at the next model in the Mythos line.

Cyber Security News

Read more →
Aaheadline · in-image textSUBJECTLOGO#hex paletteIMAGE GENERATION · JUNE 3, 2026LAYOUT-NATIVEReve 2.0 + Ideogram 4Layout control is the new promptReve 2.0 — proprietary, 4K, layout editingIdeogram 4 — open-weight, 9.3B, in-image textBoth chase OpenAI GPT Image 2 and Google GeminiBITSMINDS.COMSource: Reve · Ideogram · Latent Space
Models

Reve 2.0 and Ideogram 4 Land the Same Day, Both Betting Image AI's Next Move Is Layout Control

On June 3, two image-AI startups shipped models built on the same wager: that the future of generation is precise layout control, not ever-longer prose prompts. Reve's proprietary 2.0 chases 4K and "images you can touch"; Ideogram's open-weight 4 nails in-image text and bounding-box placement. Both are gunning for OpenAI's GPT Image 2 and Google's Gemini.

Latent Space

Read more →
OPENAI · LIFE SCIENCESJUNE 2026 UPDATEGPT-RosalindRebuilt on GPT-5.5, tuned for biologyDrug discovery · Genomics · Medicinal chemistryOutperforms GPT-5.5 on the new LifeSciBench evalBITSMINDS.COMSource: OpenAI
Models

OpenAI Rebuilds GPT-Rosalind on GPT-5.5 and Widens Access — Its Science Model Now Beats the General One on Biology

OpenAI's life-sciences model just got a major upgrade. Rebuilt on GPT-5.5, it now outscores the general model on a new benchmark called LifeSciBench across drug discovery, genomics, and medicinal chemistry — and OpenAI is widening access beyond launch partners like Moderna, Amgen, and Novo Nordisk.

OpenAI

Read more →
BUILD 2026 · MICROSOFT TAKES AIM AT CLAUDE LOCKED ON MICROSOFT AI MAI-Thinking-1 FIRST IN-HOUSE REASONING MODEL 35B active · ~1T MoE · 256K ctx trained with zero distillation Claude ENTERPRISE CODING DEFAULT SWE-Bench Pro ≈ 53% — level with Opus 4.6 BITSMINDS.COM Microsoft's claims, pending independent benchmarks
Models

Microsoft Aims MAI-Thinking-1 Straight at Claude: a 35B Reasoning Model It Says Beats Sonnet 4.6 and Matches Opus 4.6 on Code

At Build 2026, Microsoft’s AI Superintelligence Team unveiled MAI-Thinking-1, its first in-house reasoning model — a sparse Mixture-of-Experts design with 35B active parameters (~1T total) and a 256K context window. Microsoft says human raters prefer it to Claude Sonnet 4.6 and that it matches Opus 4.6 on the SWE-Bench Pro coding benchmark, and it pointedly trained the model from scratch with zero distillation on commercially licensed data — no OpenAI involved.

Microsoft AI / TechTimes

Read more →
OPEN SOURCE · MIXTURE-OF-EXPERTS · APACHE 2.0JETBRAINS MELLUM2 · JUN 2Mellum212B total parameters · ~2.5B active per tokenMoE · 131K context · 2x faster inference6 variants · Base · Instruct · Thinking (RLVR)THE FOCAL-MODEL THESISFast, specialized parts orchestrated by frontier modelsROUTER · 8 OF 64 EXPERTS FIREOnly ~21% of parameters fire per tokenBITSMINDS.COMSource: JetBrains AI Blog · Hugging Face
Models

JetBrains Open-Sources Mellum2, a 12B Mixture-of-Experts Model Built to Be a Fast "Focal" Part, Not a Frontier Rival

JetBrains released Mellum2 under Apache 2.0 on June 2 — a 12B MoE model that activates just 2.5B parameters per token for 2x-faster inference, pitched as a fast specialized component for multi-model AI pipelines.

JetBrains AI Blog

Read more →
Microsoft Will Unveil Its Own GitHub Copilot Coding Model at Build — a Direct Shot at Claude Code
Models

Microsoft Will Unveil Its Own GitHub Copilot Coding Model at Build — a Direct Shot at Claude Code

According to The Information, Microsoft plans to debut a homegrown coding model for GitHub Copilot at its Build conference on June 2–3, part of a wider suite of in-house MAI models led by Mustafa Suleyman. The push to reduce reliance on OpenAI, Anthropic and Google comes as Copilot has lost ground to Anthropic’s Claude Code among developers.

The Information / TestingCatalog

Read more →
FRONTIER MODEL SHOWDOWN · WHO WINS? Three labs. Three strongest models. One fight. ANTHROPIC Claude Opus 4.8 AUTONOMY OPENAI GPT-5.5 “Spud” AGENTS GOOGLE Gemini 3.1 Ultra REASONING VS VS BITSMINDS.COM BitsMinds original analysis
Models

Claude Opus 4.8 vs GPT-5.5 vs Gemini 3.1 Ultra: Benchmarks

A BitsMinds analysis. We put each lab’s strongest model through the benchmark portfolio that actually separates frontier systems in 2026 — GPQA Diamond, ARC-AGI-2, AIME, SWE-bench, BFCL and more — then weigh price, context and speed. The short version: Gemini 3.1 Ultra is the sharpest reasoner and the best value, Claude Opus 4.8 owns real coding and agentic work, and GPT-5.5 is the strong all-rounder that everyone already has. Here is the full scorecard.

BitsMinds Analysis

Read more →
MODEL RELEASE · AGENTIC AI · ANTHROPICCLAUDE OPUS 4.8 · MAY 28Opus 4.8Anthropic’s most honest model yetSame price as 4.7 · $5 / $25 per M tokensFast mode · 2.5× speed · 3× cheaperEffort dial · plan-and-verify subagentsDYNAMIC WORKFLOW · PARALLEL SUBAGENTSPLANVERIFY ✓running 100s in parallel · one sessionBITSMINDS.COMSource: Anthropic · TechCrunch · Axios
Models

Anthropic Ships Claude Opus 4.8: Dynamic Workflows Spawn Hundreds of Parallel Subagents, and It’s the “Most Honest” Claude Yet

Anthropic released Claude Opus 4.8 on May 28, 41 days after 4.7 — adding Dynamic Workflows that plan and run hundreds of parallel subagents per session, a user-facing effort dial, a 2.5× faster fast mode at 3× lower cost, and its “most honest” self-review behavior yet, at 4.7’s price.

Anthropic / TechCrunch

Read more →
ANTHROPIC ® INTRODUCING OPUS 4.8 LEAK FORBIDDEN-STRINGS FILE · NPM SOURCE MAP · MARCH 31, 2026 BITSMINDS Source: WaveSpeed AI · decodethefuture.org
Models

Anthropic's Leaked Forbidden-Strings File Names Opus 4.8, Sonnet 4.8 and a New Tier Above Opus Called Capybara

Inside the 512K-line Claude Code source map that npm shipped on March 31 was a feature literally called Undercover Mode — and inside Undercover Mode was a hard-coded list of names it was supposed to scrub from public builds. That list spells out Opus 4.8, Sonnet 4.8, a never-released Sonnet 4.7 and animal codenames Capybara and Tengu — with Capybara described internally as a new tier sitting above Opus.

WaveSpeed AI / ChaoBro / AI Nexus Daily

Read more →
Alibaba's Qwen3.7-Max Lands With a 1M-Token Context, AA Index 56.6, and a 35-Hour Agent Run
Models

Alibaba's Qwen3.7-Max Lands With a 1M-Token Context, AA Index 56.6, and a 35-Hour Agent Run

Qwen3.7-Max-Preview ships with a 1M-token context, extended-thinking mode, and benchmark gains across CritPt, Humanity's Last Exam, and Terminal-Bench Hard — clearing Gemini 3.5 Flash on the AA Intelligence Index.

Qwen (Alibaba Cloud)

Read more →
Google's Gemini 'Omni' Video Model Leaks Days Before I/O 2026, Pointing to a Unified Multimodal System
Models

Google's Gemini 'Omni' Video Model Leaks Days Before I/O 2026, Pointing to a Unified Multimodal System

Test UI strings inside Gemini reveal a new Google video model that lets users remix and edit clips directly in chat, with a likely debut at I/O 2026 on May 19-20.

9to5Google

Read more →
Mira Murati's Thinking Machines Unveils TML-Interaction-Small, a Full-Duplex Voice Model That Beats GPT and Gemini
Models

Mira Murati's Thinking Machines Unveils TML-Interaction-Small, a Full-Duplex Voice Model That Beats GPT and Gemini

Mira Murati's Thinking Machines Lab released a research preview of TML-Interaction-Small, a 276B mixture-of-experts model that responds in under 0.4 seconds and listens while it speaks — outpacing GPT-realtime-2 and Gemini 3.1 Flash Live on the FD-bench latency test.

TechCrunch

Read more →
OpenAI Ships GPT-Realtime-2, Translate, and Whisper, Bringing GPT-5 Reasoning Into Voice Apps
Models

OpenAI Ships GPT-Realtime-2, Translate, and Whisper, Bringing GPT-5 Reasoning Into Voice Apps

OpenAI rolled out three new voice models on May 7 — a reasoning agent with a 128K context window, a 70-language live translator at $0.034 a minute, and a streaming Whisper that transcribes as you speak.

OpenAI

Read more →
Google Ships Gemini 3.1 Flash-Lite to GA at $0.25 per Million Input Tokens
Models

Google Ships Gemini 3.1 Flash-Lite to GA at $0.25 per Million Input Tokens

Gemini 3.1 Flash-Lite is now generally available on Vertex AI and Gemini Enterprise, pitched as the cheapest, fastest member of the Gemini 3 family for high-volume agentic, classification, and tool-use workloads.

Google Cloud Blog

Read more →
Mistral Medium 3.5 Lands With Cloud Coding Agents and 77.6% on SWE-Bench
Models

Mistral Medium 3.5 Lands With Cloud Coding Agents and 77.6% on SWE-Bench

Mistral fuses chat, reasoning, and coding into a single 128B dense model and pairs it with Vibe — async cloud-based coding agents that hand back work as pull requests instead of terminal output.

Mistral AI

Read more →
DeepSeek V4 Preview Closes Gap With Frontier Models at a Fraction of the Price
Models

DeepSeek V4 Preview Closes Gap With Frontier Models at a Fraction of the Price

Chinese lab DeepSeek released two open-weight V4 preview models with 1M-token context windows, matching closed-source rivals on reasoning while undercutting them on price.

TechCrunch

Read more →
NVIDIA Unveils Nemotron 3 Nano Omni: Open Multimodal Model with 9x Throughput
Models

NVIDIA Unveils Nemotron 3 Nano Omni: Open Multimodal Model with 9x Throughput

NVIDIA's new 30B-A3B mixture-of-experts model unifies vision, audio, and language into a single open system, delivering up to nine times the throughput of comparable open omni models for AI agents.

NVIDIA Blog

Read more →
OpenAI Launches GPT-5.5: Smarter, Faster, and Built for Agentic Work
Models

OpenAI Launches GPT-5.5: Smarter, Faster, and Built for Agentic Work

OpenAI's GPT-5.5, codenamed "Spud," brings 82.7% Terminal-Bench 2.0 performance and autonomous multi-step task execution — a step toward OpenAI's unified AI super app.

TechCrunch

Read more →
DeepSeek Releases V4: Open-Source Frontier Model at a Fraction of the Price
Models

DeepSeek Releases V4: Open-Source Frontier Model at a Fraction of the Price

DeepSeek's V4-Pro and V4-Flash arrive with 1.6 trillion parameters, a 1M-token context window, and Apache 2.0 licensing — matching near-frontier performance at 21x less cost than Claude Opus 4.7.

Simon Willison

Read more →
Google Launches Gemini 3.1 Pro: Tops ARC-AGI-2 Benchmark with 2M Token Context and Agentic Coding
Models

Google Launches Gemini 3.1 Pro: Tops ARC-AGI-2 Benchmark with 2M Token Context and Agentic Coding

Google's Gemini 3.1 Pro hits a verified 77.1% on the ARC-AGI-2 reasoning benchmark — more than double its predecessor — and arrives with a 2-million-token context window, multi-step agentic coding, and support across Gemini API, Vertex AI, and NotebookLM.

Google

Read more →