Apple Capped Bug Reports Because AI Filled the Queue With Fake Ones — Then a Real macOS Flaw Couldn't Get In
An Italian startup says it found 50+ candidate macOS vulnerabilities in three weeks, including a privilege-escalation chain — and then couldn't file it, because Apple limits open submissions after a flood of AI-written reports describing bugs that don't exist. Apple is meanwhile running Anthropic and OpenAI models against its own code, with 184 CVEs in the latest macOS release. The same technology is jamming the front door and clearing the backlog behind it.
An Italian security startup called Bynario pointed an AI pipeline at macOS and, it says, turned up more than 50 candidate vulnerabilities in three weeks — including a privilege-escalation chain that would hand an attacker full control of a Mac.
Then it hit a wall that had nothing to do with the research. Apple has capped how many reports a researcher can keep open at once and imposed a cooling-off period between submissions, because its intake queue is drowning in AI-generated bug reports describing vulnerabilities that do not exist. Bynario's founder could not file the finding.
The Financial Times reported the episode; we could not read the FT piece directly, so the account here rests on The Decoder and other write-ups of it, plus the primary records below.
Bynario is not a crank, and the record shows it
This matters, because "researcher says Apple ignored me" is a genre. Apple's own security advisory for macOS Tahoe 26.6 credits CVE-2026-43760 to Alfredo Pesoli of Bynar.io, alongside two other reporters. Tenable's index of that advisory puts the release at 27 July 2026 and counts 184 CVEs in it.
So the pipeline works and Apple has already shipped a fix from it. A sourcing note: Apple's advisory page would not render for us, so that credit line comes from indexed copies rather than from the page itself. And one thing we cannot establish is whether the blocked privilege-escalation chain is the same issue as CVE-2026-43760 or a separate, later finding — Bynario reported dozens of candidates, and the public record does not line them up.
The valuation attached to it should also be read for what it is. Pesoli put the flaw's black-market worth at $100,000 to $200,000. That is the finder's own estimate of his own discovery, and grey-market prices are unverifiable by construction. It is a useful order of magnitude, not a fact.
The same technology is on both sides of the queue
Here is what makes this more than a support-queue complaint. Apple is not a bystander to AI bug-hunting — it was one of the first companies handed the capability. In May we covered Anthropic giving Claude Mythos Preview to Apple, AWS and JPMorgan to hunt zero-days, a programme that turned up more than 10,000 in its first month and grew to roughly 200 partners by June.
Apple now runs models from Anthropic and OpenAI against its own code, and reporting on the FT story says the result was around five times the usual number of fixes in recent updates — consistent with a 184-CVE release. Researchers using Claude and OpenAI's Codex Security tooling are credited in Apple's advisories.
So the same class of tool is generating the flood at the front door and the fixes behind it. What did not scale is the part in the middle.
| Stage | Scaled by AI | Did not scale |
|---|---|---|
| Finding candidates | 50+ in three weeks | — |
| Writing them up | Effectively unlimited | — |
| Reproducing and confirming | — | Human engineers |
| Deciding severity | — | Human judgement |
Why a cap is a rational response — and what it costs
Throttling submissions is defensible. If most of what arrives is fabricated, every unfiltered report is engineer-hours spent chasing a hallucination, and those hours come out of the same budget that fixes real bugs. Open-source maintainers hit this wall before Apple did.
The problem is that a cap is content-blind. It does not throttle bad reports; it throttles reports. A researcher with a 90% false-positive rate and a researcher with a working privilege-escalation chain are rationed identically, and the second one has vastly more to lose by waiting. Worse, the cap penalises exactly the behaviour AI makes cheap — submitting many candidates — without distinguishing whether they were triaged before sending.
Reputation-weighted quotas are the obvious answer and are not new; several platforms already rank researchers by historical signal quality. Apple's own advisory shows it can identify a reliable reporter after the fact. The gap is that the throttle does not appear to use that.
The counter-case: is this really AI's fault?
It would be convenient to file this as another AI-slop story, and worth resisting. Apple's bug bounty has drawn researcher complaints about slow triage, opaque decisions and low payouts for years — well before generative models could write a plausible-looking report. The programme's headline maximum now exceeds $5 million, but headline maximums were never the friction; getting a real bug looked at was.
Read that way, AI did not create the bottleneck. It found one that already existed and applied enough pressure to make it visible. A triage pipeline sized for the volume of humans willing to do unpaid security work was always going to break when the cost of producing a report fell to nearly zero. The interesting question is not who to blame but whether verification can be scaled at all — and if the only thing that scales verification is more AI, the loop closes in an uncomfortable place.
There is a real counterweight to the pessimism, too: a flaw that reached a CVE and a shipped patch is a system working, slowly. The failure here is a queue, not a cover-up.
This is the shape of the next few years
The pattern generalises well beyond Apple. Every organisation that accepts security reports from outsiders is about to receive far more of them, most machine-written, some excellent. We have watched the same asymmetry from the other side: Anthropic's models reached into three real companies during evaluations, Hugging Face published a forensic timeline of an agent intrusion, and sixty-plus companies formed a security alliance aimed at this layer. OpenAI's Patch the Planet is a bet that the fixing side wins.
There is one more wrinkle worth noting, given that Bynario's platform reportedly runs on OpenAI models: Apple is currently suing OpenAI for trade-secret theft. The flaws arriving at its door are being found with a rival's tooling, and the fixes going out are partly credited to it.
What to watch
Whether Apple replaces the flat cap with reputation weighting is the concrete thing to look for, and it would be visible in the developer documentation rather than in a press release. Beyond that: whether other large vendors follow with caps of their own, and whether any of them publishes a false-positive rate for AI-assisted submissions. Nobody has released that number, and until someone does, every argument in this story — including Apple's — rests on an estimate.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.