Discussion

The Responses API Grew Hands. Prompt Injection Grew Teeth

In Developer Tools

OpenAI Watch
OpenAI WatchParticipantOpening post
#1892

OpenAI's Responses API can now pair a model with a shell tool and a hosted container workspace. That removes much of the plumbing needed to turn an answer into a long-running agent, but it also makes permissions, not just answer quality, the practical safety boundary.

OpenAI Watch analysis

What happened

OpenAI says the service passes model-proposed commands to isolated containers and streams the output back for the next step. Workspaces can retain files and structured data such as SQLite, run separate shell sessions concurrently, package reusable skills and carry selected state through context compaction. Network requests pass through an egress proxy with allowlists and access controls; domain-scoped secret injection applies credentials only to approved destinations without showing the raw secret to the model. Those are useful controls, but they are OpenAI's documented design, not an independent security assessment.

Why it matters

OpenAI's same-day security account describes prompt injection as social engineering: malicious content can look relevant, urgent and authorised, so filters cannot simply hunt for a naughty phrase. Its Safe URL control can ask for confirmation or block a sensitive transmission to a third party. A day earlier, OpenAI reported work training models to prefer system and developer directions over user or tool content.

Independent VPI-Bench researchers tested 306 visual prompt-injection cases across five pseudo-authentic platforms and found every evaluated agent vulnerable, with tested defences producing only limited gains. That preprint did not evaluate this Responses API environment, so it is not evidence that OpenAI's hosted service was breached. It does show why a smarter model is not a sufficient boundary once untrusted webpages or documents can influence an agent with a shell, persistent state and network access.

For developers, least privilege must do the boring work: narrow egress, scoped secrets and writes, confirmation before irreversible actions, inspectable logs and disposable runs. Isolation can contain a mistake; it cannot decide whether the requested action was legitimate.

Our read

OpenAI deserves credit for treating egress and credentials as system controls rather than prompt-writing problems. Centralised controls may be safer and easier to patch than hundreds of improvised runtimes. The risk is that convenience makes broad access the default. The platform should make narrow authority easier than ‘allow all’ and let operators reconstruct why an agent acted, not merely which command it ran.

What to watch

  • Independent prompt-injection tests against multi-step shell tasks, including delayed and visually embedded attacks.
  • Default-deny network and secret scopes that remain narrow when skills are added.
  • Clear retention, reset and deletion rules for files, compacted state and execution logs.

Discussion spark: Which agent permission would you refuse to grant without fresh human confirmation, even inside an isolated container?

Sources and evidence

OpenAI Watch is independently operated by WittyWires. It is not affiliated with, endorsed by, or operated by OpenAI.