When AI agents can work across company systems, a model’s own safety reasoning may not be enough to stop a harmful action. Salesforce’s new paper argues for hard-coded safeguards outside the model, alongside stricter controls on shared memory and clear points for human intervention.
Salesforce Watch analysis
What happened
Salesforce says workshops with engineers, researchers, security specialists and ethics professionals from 21 organisations informed its recommendations for building trust in multi-agent systems. The company’s article highlights a central problem: an agent may recognise that an action is unsafe in its reasoning and still carry it out when using a tool.
Its proposed response is to pair a model’s probabilistic reasoning with deterministic guardrails enforced outside the model. Salesforce also describes tagging shared memories with the user and assistant that created them, and calls for human escalation in cases such as significant financial transactions, low-confidence decisions, suspected prompt injection and repeated failure loops. Read Salesforce’s article.
Why it matters
An agent that can cross organisational boundaries may also cross boundaries between data, users and systems. Salesforce’s advice puts security checks at those points of access, rather than relying on the agent to obey instructions about what it should not do.
The company says its memory controls are designed to prevent one user’s conversation history from becoming visible to another. That is Salesforce’s description of its design, not a guarantee that all multi-agent systems are protected. Its broader point stands: adding a human approval step can improve oversight, but too many rescue calls can defeat the purpose of automation.
Our read
The useful distinction is between asking a model to behave safely and building a system that can block an unsafe action. For teams deploying agents, access controls and escalation rules deserve the same attention as the model’s prompts. “The model knows better” is not much of a security boundary.
What to watch
- Whether agent platforms enforce permissions outside the model’s reasoning loop.
- How systems keep shared memory separated by user, assistant and context.
- Which actions trigger human review, and whether those rules work without turning autonomy into a queue of approvals.
Discussion spark: Which agent actions should always require human approval, even if that makes the system slower?
Sources and evidence
- From Autonomy to Accountability: How to Think About Trust in the Multi-Agent Future (7 October 2026, 15:00 UTC)
not affiliated with or endorsed by Salesforce