Industry·4 min read
By BitsMindsSource: Hacktron AI

Claude Opus 5 Wrote the Exploit That Broke Into OpenAI

Three researchers at Hacktron chained a heap overflow in an image library with a single sign-on flaw, reached OpenAI employee Codex accounts and opened a pull request inside the internal monorepo. Opus 4.8 could not build the exploit. Opus 5, hours after launch, did it in three.

Hacktron used Claude Opus 5 to build the exploit chain that reached OpenAI's monorepo 3 PEOPLE · 72 HOURS · $6,500 BITSMINDS.COM
Share:

A three-person team at the San Francisco security startup Hacktron spent 72 hours in late July chaining two flaws until they were inside OpenAI’s GitHub organisation, and published the write-up this week. The detail that will travel is not the target. It is that the same exploit was attempted twice with two different models, and only the newer one could write it.

The way in was not OpenAI’s infrastructure but its community forum, which runs on Discourse. Discourse’s upload pipeline uses FastImage to sanity-check images, and FastImage does not understand the HEIC/HEIF format iPhones produce — so those files fell through to ImageMagick, which decodes them with libheif. The Debian 12 base image in Discourse’s Docker container shipped libheif 1.19.7, a build that, in Hacktron’s words, was missing some upstream security fixes that “were not back-ported to the libheif package.” An attacker-supplied photo was therefore an attacker-supplied heap buffer overflow.

Turning that overflow into reliable code execution is the hard part, and this is where the timeline gets pointed. Hacktron started on the image pipeline on 23 July. On 24 July they worked the bug with Claude Opus 4.8 — available to them through the access programme Anthropic runs for security researchers — and it repeatedly failed to produce an exploit that survived ASLR. Opus 5 shipped that evening. A fresh session with the same problem returned a working ARM64 exploit within three hours; the team ported it to x86-64 with a jemalloc configuration and confirmed local remote code execution through an image upload at 6am on 25 July. Four hours later an autonomous Claude agent reproduced the RCE against Discourse Cloud.

The second flaw is the one that made it OpenAI’s problem rather than Discourse’s. The forum authenticated through “Sign in with OpenAI” against auth.openai.com, and once the researchers held administrative access to the Discourse instance, a misconfiguration in that single sign-on flow let them take over employee ChatGPT and Codex accounts without any further authentication. One of those Codex accounts was connected to OpenAI’s GitHub organisation. To demonstrate the reach without touching anything sensitive, they opened pull request #1186742 in the internal openai/openai monorepo between 13:30 and 15:30 UTC, and then stopped testing.

Disclosure was fast on both sides. Hacktron filed through OpenAI’s Bugcrowd programme the same day; OpenAI confirmed its half was fixed at 22:49 UTC, fourteen hours after submission, and Discourse shipped its fix on 27 July. The bounty was $6,500, with OpenAI noting that the award covered the OpenAI-side finding rather than the work against Discourse, since the forum sat outside the scope of its programme. That scoping line is doing a lot of work: the forum was not OpenAI’s to test, and it was also the door.

Hacktron frames the OpenAI chain as one output of a wider libheif campaign it calls the HEIF Heist, which it says reached Slack, Meta, Zoom, Shopify and GitHub Enterprise among others. The economics it reports for that campaign are the uncomfortable part: two months, three researchers, and under $3,000 in model tokens. The OpenAI portion, by their account, took a few days of agent time and a few hours of human time. Work that would once have needed a well-resourced team and months of effort is being compressed into days, and the write-up argues plainly that “security assumptions must catch up with attacker capabilities.”

Two caveats are worth keeping attached to this. The researchers were white hats who reported and stopped, and the capability they used is the same one defenders are being handed — Anthropic and OpenAI both gate these models behind vetted-access programmes precisely so that this work happens in public write-ups rather than quietly. And the underlying bugs were ordinary: an unpatched library in a base image, and an SSO flow that trusted a forum more than it should have. No model invented either of them.

What the episode does supply is a clean, dated measurement of something the labs usually describe in the abstract. The same team, the same target, the same bug, one model generation apart: Opus 4.8 could not get past ASLR across several sessions, and Opus 5 cleared it in three hours. That is a sharper read on exploit-development capability than most published evaluations manage, and it arrived not from a red-team report but from a bug bounty filed the morning after a launch. The labs publish their cyber capability thresholds on their own schedule; the field is now generating its own datapoints, and it does not wait for the system card.

More on Claude

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Pew's global AI survey: uncertainty about the future of work An empty ochre office chair faces a computer in a quiet evening office. Beyond the window, buildings of unequal heights suggest an uneven economy. A large globe on the left is surrounded by 37 separate markers: 34 amber and three neutral. They represent surveyed publics, not individual respondents, countries mapped to locations, job counts or measured job losses. The caption states that in 34 of 37 publics, more people expect job losses than gains. This is an editorial illustration of public expectations reported in the Pew survey, not a prediction of what AI will actually do. Original vector illustration for BitsMinds, pew-global-ai-survey-2026-jobs-inequality. 17 September 2026. Survey finding, not observed labour-market outcomes. PEW / GLOBAL AI SURVEY 34 / 37 PUBLICS More expect job losses than gains AI BITSMINDS.COM
Industry

Most of the World Expects AI to Cost Jobs, Pew Finds

LawZero: a scientific instrument for understanding AI A warm-lit research laboratory. A mounted brass observation lens magnifies part of a branching AI network on a transparent specimen plate. The network continues unchanged outside the lens: this is examination, not a claim of proven safety. Canadian and German desk flags flank the instrument, with separate pledges labelled up to C$150 million and up to 100 million euros. A research notebook and a Scientist AI nameplate reference Yoshua Bengio's LawZero. The apparatus is an editorial metaphor for the proposed research programme. BitsMinds original vector illustration. Article: lawzero-bengio-canada-germany-300-million. 17 September 2026. Funding shown in the two original currencies. LawZero BENGIO'S SCIENTIST AI SCIENTIST AI CANADA · UP TO C$150M GERMANY · UP TO €100M BITSMINDS PUBLIC FUNDING / AI SAFETY RESEARCH
Industry

Bengio's LawZero Gets $300M From Canada and Germany

Project Lily: the human on the other side of the chat In a dark office lit by a warm desk lamp, a cream-coloured ChatGPT conversation unfurls from a deep green computer screen into a paper transcript. A human reviewer in a teal shirt leans over the page, one arm dropping out of view below the near edge of the desk. Black redaction bars conceal parts of the conversation, while other lines remain visible. A seven-point rating card with a check above six sits on the desk. This imagined editorial scene represents the human review and incomplete filtering described in the article; it is not an actual review interface or a claim that every conversation is reviewed. Illustration for BitsMinds, OpenAI Project Lily human ChatGPT prompt review. Original vector artwork, 15 September 2026. PROJECT LILY ChatGPT HUMAN FEEDBACK 1234567 BITSMINDS THE HUMAN IN THE LOOP
Industry

ChatGPT Chats Are Read and Rated by Human Contractors