Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

The Allen Institute for AI says its fully open MolmoAct 2 robotics model achieved the highest overall task success on the Reality Check benchmark after task-specific fine-tuning.

Why it matters

That is a promising result for open robotics research, but the fine-tuning matters: it is not a like-for-like measure of how models perform straight out of the box. The Institute credits Nicolas Keller with testing the model and sharing methods and results. The useful next test is whether the lead holds across tasks and setups.

Discuss: Should robotics models be judged on their best fine-tuned results, or mainly on performance before task-specific tuning?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.