OpenAI’s Astra for Law Scores 54% on Legal Research
Astra for Law wires a 230-million-URL index of US case law, statutes and regulations into GPT-6 Astra, and ships to selected firms through Trusted Access. OpenAI’s headline is a 40% improvement. The underlying score is 54.0% correct, against 38.7% for the same model with plain web search.
OpenAI introduced Astra for Law on 17 September: GPT-6 Astra packaged with a dedicated legal search index, instructions for legal analysis and writing, and a set of tools aimed at US research. It appears in the model picker as “GPT-6 Astra Law” and in the API as gpt-6-astra-law, and it goes first to selected firms through a Trusted Access programme in ChatGPT and Codex, with API access to follow.
The index is the substance of it. Astra for Law can search US case law, statutes, regulations, court rules and administrative decisions across a corpus OpenAI puts at more than 230 million URLs, with sources added daily. A collaboration with the Free Law Project, the nonprofit behind CourtListener, brings in a case-law collection covering over 99.9% of published US precedential decisions. OpenAI is careful to frame this as complementary to the licensed content firms already buy from providers such as Thomson Reuters rather than a replacement for it.
The number attached to the launch is a 40% improvement, and it is worth unpacking which 40%. OpenAI tested the full configuration against 200 US legal research questions from the private validation set of Vals AI’s Legal Research Bench, with both systems at their highest reasoning effort. Astra for Law passed the overall correctness check on 54.0% of questions; GPT-6 Astra with ordinary web search managed 38.7%. The 40% is the relative distance between those two figures.
On case-law questions specifically, OpenAI reports the configuration found 24% more reference cases than plain Astra with web search, and retrieved up to 54% more relevant passages from the correct opinions at matched reasoning effort. Those are retrieval gains, and retrieval is the part of legal research most amenable to an index. The absolute correctness level is the part a managing partner will read twice: on a benchmark built to measure whether the model finds the right authority and meets the evaluation’s criteria, the best configuration is wrong on nearly half the questions. That is a meaningful advance over the same model unassisted and still a long way from unsupervised.
OpenAI also published a head-to-head against Claude Fable 5.1, showing both models’ answers to a misrepresentation prompt and stating that Fable returned a holding that had been reversed on appeal in the litigation example, and reported finding no matching case in the transactional one. Two things are true about that: citing reversed authority is precisely the failure mode that has produced sanctions in real courtrooms, and this is a vendor-run comparison on vendor-chosen prompts with no published methodology. It is an argument for the index, not a measurement of Fable.
The governance layer is where OpenAI is trying to answer the objection that has kept frontier models out of privileged work. Eligible firms get Zero Data Retention on the API, and ChatGPT Enterprise usage is excluded from human review by default — a clause that lands three days after reporting that human contractors read and rate ChatGPT conversations. OpenAI says it is working with Latham & Watkins on information permissions, ethical walls, client instructions and firm oversight, the machinery a firm needs before a model touches a matter file.
Around the model sits an ecosystem play. There are 26 partner plugins connecting ChatGPT to tools firms already run — Relativity, Clio, iManage, Intapp, DeepJudge, and Thomson Reuters bringing HighQ matter context in while previewing a CoCounsel Legal connector — plus nine community plugins from LegalQuants, LECG and Skills.law carrying 47 adaptable skills. ChatGPT for Word went generally available the same day. And OpenAI’s forward-deployed engineers have been building bespoke systems inside firms: an agreement analyzer at Sullivan & Cromwell that folds the firm’s negotiating playbooks into deal review, a data-room diligence system at Ropes & Gray, and Cooley’s GO Public for IPO preparation, which propagates a change across a filing when the deal moves.
The commercial geometry is the interesting part. API customers including Harvey and Legora — the latter reported to be raising at a $10 billion valuation — will build on Astra for Law, so OpenAI is simultaneously their supplier and the owner of the legal index underneath their products. Its own framing is conciliatory: “Building on OpenAI should mean getting more from the ecosystem, not replacing it.” That holds while the model layer is the commodity and the workflow is the product. It gets harder to sustain each time OpenAI ships another piece of the workflow itself, and it sharpens the bet Kirkland & Ellis made by building rather than buying.
What has actually shipped is narrower than the announcement’s reach: a configuration, not a new model, available to a selected set of firms, evaluated by the company that built it. The useful test will not be the Vals number but whether Vals or another independent evaluator reproduces it on a held-out set — and whether 54% is the floor of a fast climb or the shape of what retrieval alone can buy.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.