AI companies train models to refuse harmful requests, but those safeguards can fail in both directions: dangerous answers may slip through, while legitimate questions get blocked. A new MIT Technology Review essay argues that refusal is an uncertain technical control and a consequential question of who gets to draw the line.
Watch Desk analysis
What happened
The essay traces several ways companies try to make AI say no, including training on examples of harmful requests, using other models to assess answers, and placing classifiers around a model to block risky prompts or responses. It describes these systems as probabilistic: they may miss harmful requests, and may refuse benign ones.
The piece also reports that Anthropic said one type of classifier added 24% to its chatbots’ compute costs, and that companies have begun switching to probes that monitor a model’s internal activations. The essay draws on interviews with researchers and former industry staff, while noting that the underlying mechanisms of refusal remain only partly understood.
Why it matters
A refusal is not just a technical feature. Someone has to decide which requests are dangerous, and the essay warns that the power to set those boundaries could move from companies to governments. The same rule might block genuinely harmful instructions or suppress legitimate speech, depending on who sets it and how it works.
For users, the practical point is that a chatbot’s refusal is not proof of a reliable safety barrier, and a refusal to answer is not proof that a question was harmful. The safeguards involve trade-offs in reliability, cost and control.
Our read
The essay makes a useful case against treating “the model refused” as the end of the safety conversation. Refusals matter, but they are not a magic switch, and the rules behind them deserve scrutiny too. Readers should treat claims about what safeguards can reliably prevent as claims to test, not comforting labels on the box.
What to watch
- Whether companies publish clearer evidence about when refusal safeguards fail or over-refuse.
- How the use of internal-activation probes develops, including what they can and cannot detect.
- Whether governments propose rules that make refusal boundaries more transparent, or simply give themselves more power to set them.
Discussion spark: Who should have the final say over what an AI system refuses: the company that builds it, governments, or users, and what safeguard would stop that power being abused?
Sources and evidence
- We’re putting too much faith in AI’s ability to say no (9 October 2026, 09:00 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.