Discussion

OpenAI contractors reportedly review real ChatGPT conversations

In The Watch Desk

OpenAI Watch
OpenAI WatchParticipantOpening post
#2594

404 Media reports that hundreds of OpenAI contractors review real ChatGPT conversations to evaluate the model’s replies. The practical point for users is immediate: on Free, Plus and Pro accounts, switching off Improve the model for everyone prevents new chats from being used for model improvement, but does not work retroactively.

OpenAI Watch analysis

What happened

The 404 Media investigation says contractors working on a programme codenamed Project Lily read real prompts, summarise the user’s intent, then score and critique several generated responses. The materials do not identify which current or future model is being trained.

According to the report, reviewers do not see usernames, but conversations can contain sensitive personal information and may be accompanied by summaries of earlier user memories. OpenAI told 404 Media that a Privacy Filter tries to remove identifying information before review, while acknowledging that the model can miss uncommon identifiers or redact details incorrectly.

Why it matters

ChatGPT is routinely used for health worries, workplace drafts, relationships and other material people would not casually hand to an unknown contractor. Human evaluation can improve warmth, accuracy and resistance to sycophancy, but the benefit does not dissolve the consent question in a puff of helpfulness.

Disclosure is the sharp edge. 404 Media says OpenAI did not answer when asked where users are explicitly told that humans may review chats for model improvement. Anthropic also confirmed to the publication that it uses human review to improve its models, making this a wider industry practice rather than an OpenAI-only design choice.

Our read

Users should assume that any consumer AI conversation permitted for model improvement may enter a human-review workflow, even when automated redaction is used. Turn off model improvement before sharing anything sensitive, and still avoid entering secrets, client data or information that could identify another person.

OpenAI should state plainly, beside the setting itself, that human reviewers may see selected conversations and that automated filtering can fail. A privacy control should not require an investigative report as its unofficial instruction manual.

What to watch

  • Whether OpenAI adds an explicit human-review disclosure to ChatGPT settings.
  • Whether existing conversations can eventually be withdrawn from training workflows.
  • What access, retention and audit controls govern contractors handling prompts.
  • Whether other major AI providers clarify their own review practices.

Discussion spark: Should consumer AI services require explicit opt-in consent before a human reviewer can see conversations used for model improvement?

Sources and evidence

OpenAI Watch is independently operated by WittyWires. It is not affiliated with, endorsed by, or operated by OpenAI.