AWS says fine-tuning a Qwen3.6-27B search agent with multi-turn reinforcement learning improved its results on three of four held-out benchmarks, while sharply reducing failed tasks on BrowseComp-Plus. The useful detail is that the training rewarded the whole search journey, not just one answer at a time.
AWS AI Watch analysis
What happened
In a 2 October AWS Machine Learning Blog post, the company describes training the agent with Amazon SageMaker AI’s multi-turn reinforcement learning service. The search agent could use BM25 keyword search and vector search, then refine its queries over several rounds. AWS used nDCG@10, a measure of how well relevant documents rank in the top ten, as the reward for the completed trajectory.
AWS reports gains on WixQA, Wands and BrowseComp-Plus, with a slight regression on FreshStack. On BrowseComp-Plus, the fine-tuned model’s nDCG@10 rose from 0.5136 to 0.6354 across 830 questions. Its reported failure rate, including errors such as exceeding the turn or token budget, fell from 22.89% to 0.68%. Read AWS’s report.
Why it matters
Search agents make a string of linked decisions: which tool to use, what to search next and when to stop. Training against the final retrieval result gives the system a reason to get that sequence right, rather than polishing each turn in isolation. The BrowseComp-Plus result suggests the approach may help an agent finish reliably as well as retrieve better.
For teams building retrieval agents, AWS also says its setup can use custom rewards and tool loops, with training jobs and evaluation metrics available through SageMaker AI. That makes this a concrete recipe to consider, though the reported results come from AWS’s own evaluation of one model and setup, not a general verdict on multi-turn RL.
Our read
The most persuasive number here is the fall in BrowseComp-Plus failures. A search agent that finds good documents only when it stays within its limits is not much of an agent; it is an expensive near-miss. The gains are promising, but the FreshStack dip is a reminder that tuning for one set of search tasks does not guarantee improvement everywhere. Treat this as a useful method to test against your own retrieval work, not a plug-and-play promise.
What to watch
- Whether independent evaluations reproduce the retrieval gains and lower failure rate.
- How results change across models, search tools and datasets beyond AWS’s four tests.
- Whether teams find the quality improvements worth the training and inference costs.
Discussion spark: For a search agent, would you prioritise better rankings on the tasks it completes, or fewer failed runs even if performance varies by benchmark?
Sources and evidence
- Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI (2 October 2026, 15:44 UTC)
- Add secure Web Search to Claude Desktop with Amazon Bedrock AgentCore (2 October 2026, 15:44 UTC)
not affiliated with or endorsed by Amazon Web Services (AWS)