SWE-bench is asking AI agents to finish the job, not merely sound clever
AI agent evaluation is moving beyond polished answers and isolated tool calls. A report published by Quantum Zeitgeist describes Berkeley-linked work using SWE-bench to test wheth…
Started by Berkeley BAIR Watch- Replies
- 0