ChatGPT Chats Are Read and Rated by Human Contractors
Leaked documents describe Project Lily, an OpenAI effort that pays hundreds of contractors to read real ChatGPT conversations and score the replies. Usernames are stripped and a filter model removes personal details, but internal documents concede that sensitive information still gets through.
OpenAI is paying hundreds of contractors to read real ChatGPT conversations, according to leaked internal documents and prompts obtained by 404 Media, which published its account of the programme — internally called Project Lily — on Monday. The reviewers see whole exchanges, not isolated questions, and the prompts sometimes contain sensitive personal information. ChatGPT has more than 900 million users, very few of whom would guess that a stranger might read what they typed.
The work itself is the unglamorous machinery of model improvement. Contractors rate and critique the chatbot's replies, and that signal feeds back into training. Follow-on coverage of the report describes reviewers recruited through a firm called Crossing Hurdles, paid via the marketplace Mercor at rates above $50 an hour, scoring responses on a one-to-seven scale.
What they are training the model out of is specific. The documents show contractors teaching ChatGPT not to anthropomorphise itself and to be less sycophantic — the failure mode that made the older 4o model notorious, and that features in multiple lawsuits filed against the company after users died by suicide. Read charitably, this is a safety programme staffed by people: a known harm that automated evaluation did not catch, handed to humans who can.
The privacy trade-off is where it gets harder to defend. Reviewers do not see usernames, and OpenAI runs a filter model that strips personally identifying details before prompts reach them. But the internal documents acknowledge that the filter does not catch everything and sensitive details still get through. That matters more than it would for a search engine, because of what people bring to a chatbot: therapy-shaped conversations, medical questions, work documents they should not be pasting anywhere. OpenAI has spent the past year pushing ChatGPT deeper into exactly those categories, from a health product wired into Epic's patient records to an age-detected teen tier.
None of this is unique to OpenAI, and the report says so: Anthropic confirmed to 404 Media that it also uses human review to improve its models. Reinforcement learning from human feedback has always meant humans, and the industry's rhetoric about scale — bigger models, more compute, better data — has consistently underplayed how much of the quality comes from contract workers reading output one line at a time. It is the same invisible layer that Amazon spent twenty-one years monetising before shutting Mechanical Turk down in August; it did not disappear, it moved upmarket and started paying better.
Users have some control, with a catch. ChatGPT's data controls and temporary chats keep conversations out of training, but those settings govern what happens next — they do not reach back into chats already logged. Anyone who has been treating the assistant as a private notebook for the past two years has no retroactive switch to flip.
The gap here is not really legal. OpenAI's policies have long said human review is possible, and every major lab does it. The gap is between what a policy permits and what a user pictures while typing at 1am into something that answers like a person — which is precisely the impression the contractors are being paid to sand off.
More on ChatGPT
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.