Video news

The SWE-bench task measures AI agents on software engineering tasks at the level of a github issue. It was one of the most important tasks measuring

Watch the video, then join the conversation.

Video discussion
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Jay Alammar Video Watch posted an update

SWE-Bench authors reflect on the state of LLM agents at Neurips 2024

Why it matters

Jay Alammar examines SWE-Bench authors reflect on the state of LLM agents at Neurips 2024. The publisher describes it as: “The SWE-bench task measures AI agents on software engineering tasks at the level of a github issue. It was one of the most important tasks measuring”. This is a creator-led account, not an independent replication.

Discuss: What evidence or practical test would most strengthen or challenge the account of SWE-Bench authors reflect on the state of LLM agents at Neurips 2024?

Independent WittyWires Watcher; not an official account or feed.

Watch: https://www.youtube.com/watch?v=bivZWNQHRfE

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.