Research·2 min read
By BitsMindsSource: MarkTechPost / Together AI

Together AI Open-Sources OSCAR, a 2-Bit KV Cache That Cuts Long-Context Serving Costs in Half

OSCAR — Offline Spectral Covariance-Aware Rotation — squeezes the KV cache that bloats every long-context request down to two bits per element, with no measurable accuracy loss on Qwen3-32B or GLM-4.7 across reasoning, math and coding benchmarks. Together AI shipped it into SGLang on May 25, framing it as one of the more practical wins for actually-affordable 1M-token serving.

Together AI Open-Sources OSCAR, a 2-Bit KV Cache That Cuts Long-Context Serving Costs in Half
Share:

On May 25, Together AI open-sourced OSCAR (Offline Spectral Covariance-Aware Rotation), an attention-aware 2-bit quantization scheme for the KV cache — the per-token activation buffer that has quietly become the single most expensive thing about serving long-context LLMs. The release, written up by MarkTechPost, lands inside SGLang as a drop-in INT2 KV cache mode with full paged attention compatibility, meaning operators can switch on 2-bit storage without rewriting their serving stack.

The hard part of pushing a KV cache to two bits is not the math but the outliers. KV activations contain a handful of channels that carry extremely large values, with the rest of the channels well-behaved. Naive INT2 quantization — four representable levels per element — burns almost the entire range on those rare spikes, collapsing every normal value into one or two buckets and shredding attention quality. OSCAR’s answer is to rotate the activations into a basis where the outliers are spread across many channels before they ever hit the quantizer, using a spectral covariance analysis computed offline from a calibration set.

The runtime path is engineered to keep that math invisible. On the write side, each token is rotated, clipped to a calibration-derived percentile threshold, then quantized with per-token asymmetric INT2. On the read side, the INT2 kernel unpacks the packed bytes, dequantizes, applies the inverse rotation and hands results to the attention kernel — all in a single fused pass with no extra memory traffic. Because the rotation matrix and clipping thresholds are precomputed, there is no online overhead beyond the unpack-dequantize-rotate sequence that already lives inside any quantized attention kernel.

Together evaluated OSCAR against four reference models: Qwen3-4B-Thinking-2507, Qwen3-8B, Qwen3-32B and the 358B-parameter GLM-4.7-FP8 — the same model family covered in xAI’s and Alibaba Cloud’s recent long-context benchmarks. On AIME25, GPQA-Diamond, HumanEval, LiveCodeBench v6 and MATH500, INT2 OSCAR matched the FP16 baselines within a few tenths of a point, while shrinking the per-token KV cache footprint by roughly 8x compared to the FP16 baseline and 4x compared to typical INT8 caches.

The strategic read is simple: the long-context arms race has shifted from “who can claim a million tokens” to “who can actually afford to serve them.” Anthropic’s 1M-token context on Claude and Gemini’s 2M-token window are priced like luxury goods today in part because the KV cache scales linearly with sequence length and dominates the GPU memory bill. Plugging a 2-bit cache that survives reasoning benchmarks into the open serving stack is the kind of plumbing improvement that quietly resets the floor on what self-hosted long-context inference costs, and it is exactly the angle Together has been working since its Q1 funding round.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

OpenAI publishes 722 maths papers from an unreleased model An original monochrome still life on a pale grey desk. A tall stack of bound manuscripts is topped by a graphite cover with the OpenAI mark pressed into it and a small green seal with a tick. A loose page in front shows a line of number theory. Text on the left reads OPENAI · MATHEMATICS, 722 papers from one unreleased model, 372 families and 162 Lean-checked. The stack is an editorial metaphor; the tick marks Lean formalization, not independent peer review. BitsMinds original editorial vector artwork for openai-722-math-papers-internal-model. Figures verified on 7 October 2026 from github.com/openai/math (README, overview.tex, lean/formalization.yaml). Official OpenAI path from public/logos/openai.svg. L(s, χ) ≠ 0 for Re s > 7/8 OPENAI · MATHEMATICS 722 papers from one unreleased model 372 FAMILIES 162 LEAN-CHECKED BITSMINDS.COM
Research

OpenAI Posts 722 Math Papers From an Unreleased Model

Mythos cracks Rejetto HFS's random signing key A large ivory die on a cream field, the official Anthropic mark inlaid in clay on its front face. Its top face has split along a crack, and a brass key is rising out of it: the session signing key recovered from Math.random. Faint leaked random numbers drift in from the left; faint cookie fragments sit on the right. 0.73418 0.11902 0.58361 0.92047 0.30775 admin=1 sig:9f3c keygrip xs128+ BITSMINDS.COM
Research

A Bug Claude Mythos Found Was Exploited Within a Day

Meta Muse Spark's six math papers A fan of research manuscripts on a deep blue field. Five sheets behind carry gold check marks for the five open problems answered; the front sheet carries the official Meta mark, inlaid. Faint mathematical symbols float on either side. ∫ ∑ ψ |G| = 384 ∂ₜu λ ≥ 0 ℚₚ ≠ MUSE SPARK · THINKING 6 PAPERS · 5 OPEN PROBLEMS BITSMINDS.COM
Research

Meta Says Muse Spark Helped Crack Five Open Math Problems