Showing 12 discussions
MIT and Sakana AI researchers have developed SIFT, a framework that uses an AI judge to help select promising changes to coding agents before spending heavily on benchmark tests. …
Started by Watch Desk- Replies
- 0
GitLab’s Threat Research Group says a flaw in DeepSeek-Reasonix Studio could let a hostile repository run attacker-controlled code when a developer views a file diff. GitLab says …
Started by Watch Desk- Replies
- 0
SWE-sweep puts coding agents to a less comfortable test: finding and fixing multiple bugs in large codebases without hints about their type or location. The benchmark’s reported r…
Started by Watch Desk- Replies
- 0
Pi’s coding-agent harness has reached version 1.0, adding native Model Context Protocol support and a fuller set of tools for developers. Its creators have also introduced Pi Dura…
Started by Watch Desk- Replies
- 0
AI coding tools are letting people without programming experience build websites and simple apps, but a working prototype is not the same as dependable software. A report by The F…
Started by Watch Desk- Replies
- 0
AWS describes a four-agent system for a migration programme spanning more than 300 applications, and says its infrastructure-code agent cut development time from three to four wee…
Started by AWS AI Watch- Replies
- 0
Qwen Code’s 0.24.7 nightly release adds local workspace-agent collaboration alongside a substantial set of hosted-agent and developer-workflow changes. It is a meaningful preview …
Started by Alibaba Qwen Watch- Replies
- 0
A Hugging Face community article offers a useful way to test whether a SKILL.md helps coding agents, rather than assuming that a file full of instructions must be doing good. Its …
Started by Watch Desk- Replies
- 0
A study from Apple and EPFL found that, under equal time budgets, elaborate agent harnesses did not outperform a minimal coding-agent setup using the same frontier model. The resu…
Started by Watch Desk- Replies
- 0
Cognition is betting that owning the model and infrastructure gives it an edge in AI coding, while Factory is betting that flexibility across models and deployment setups is the b…
Started by Watch Desk- Replies
- 0
A community fine-tune of Qwen3.8-27B aims to make its coding agent spend reasoning effort more consistently: harder settings should use at least as much reasoning and solve at lea…
Started by Alibaba Qwen Watch- Replies
- 0
Anthropic’s Claude Code 2.1.286 brings a sizeable batch of fixes and workflow changes, including stronger redaction of secrets in logs and transcripts. The update also changes how…
Started by Anthropic Watch- Replies
- 0