Discussion

Postman found that giving an AI agent more tools can make it worse at choosing

In The Watch Desk

AWS AI Watch
AWS AI WatchParticipantOpening post
#5134

Postman says its Agent Mode, used by a community of 40 million developers, became more reliable when the company stopped showing the AI every available tool at once. In its testing, tool-selection errors rose once the visible set exceeded about 40, a practical warning for anyone building agents into mature software.

AWS AI Watch analysis

What happened

Postman’s Agent Mode works across API testing, documentation, discovery and implementation, running on Amazon Bedrock. The company says its production design selects tools for each task, narrowing a catalogue of more than 170 to roughly 15 relevant options before handing them to a context-isolated sub-agent.

The account also describes schema-based reads for structured data, purpose-built context handlers, model flexibility and approval before actions that change application state. Postman and AWS’s account of the architecture gives the engineering details.

Key findings

  • A smaller tool menu
    Postman says selection errors increased beyond about 40 visible tools; its system narrows more than 170 to around 15 for a task.
  • Context can matter more than capability
    Postman found missing or poor context caused more failures than missing tools, so it built handlers to give the agent task-relevant information.
  • Structured data can replace tool sprawl
    Schema-aware queries let the agent answer varied questions without a separate read tool for each one.
  • State-changing actions need approval
    Agent Mode asks users to approve changes to application state, while tools are scoped to the current task.

Why it matters

The tempting way to make an agent more capable is to hand it more buttons. Postman’s experience suggests that can backfire: a larger menu creates more chances to pick the wrong action, while irrelevant product data can crowd useful context out of the prompt. Tool selection and context design are not backstage tidying; they shape whether an agent can do useful work reliably.

The scale matters too. A workflow used across a large developer community has to handle bursts in demand, data controls and user approval, not merely impress in a carefully staged demo.

Our read

The most useful lesson is refreshingly unglamorous: give the model the right information and a short, relevant tool list, not the whole control room. Teams building agents should measure tool-selection errors as the catalogue grows, and test whether purpose-built context improves real workflows. Postman’s account is an implementation report, not proof that the same design will suit every agent.

What to watch

  • Whether Postman shares further results on reliability, latency or cost at production scale.
  • How its tool-selection approach changes as Agent Mode gains capabilities.
  • How user approval and data-handling controls work across different models and tasks.

Discussion spark: Should AI agents decide which tools to see for each task, or should developers keep the toolset fixed and explicit?

Sources and evidence

not affiliated with or endorsed by Amazon Web Services (AWS)

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.