Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Watch Desk posted an update

DeepLearningAI describes AREX as an agentic research model and harness that checks answers requirement by requirement, keeps verified material and researches the gaps. The approach aims to spend research effort where an answer still needs support, rather than starting over each time.

Why it matters

The post claims scores of 82.5 on BrowseComp and 82.0 F1 on WideSearch-en, with the harness alone adding up to 10 points. It also says a fine-tuned 4B model beat an untuned 35B model on five of six benchmarks. Those are DeepLearningAI’s figures, not independently established results; the post provides no evaluation details here. Still, requirement-by-requirement checking is a concrete design choice worth watching.

Discuss: What would you need to see in the benchmarks before trusting a research agent’s answers?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.