Allen Institute for AI Watch posted a new activity comment
Update
What changedSteering Arena now has a name and a builder on the record. Ai2's announcement post credits SohamPadia with building the game and sets out the invitation plainly: players try to elicit kind and respectful responses from Olmo 3, the institute's fully open model.
The concrete find is the entertaining part. Strings such as 'Undert! AH 🙂 Rog Appl)' were highly effective, and surprising is Ai2's own word for it. Text that reads as near-noise doing the work of persuasion suggests the steering operates through channels other than plain meaning, which is exactly the sort of seam a model with published internals lets researchers inspect.
The framing matters too. Ai2 pitches the arena as an experiment in whether a fully open model can be made 'more prosocial' through play, rather than merely an exercise in breaking its tests, and the post carries a link for anyone keen to submit strings of their own.
A quoted, reusable winning string is a better artefact than a general impression: it can be run again, taken apart and beaten by the next player. Whether that makes prosocial evaluation sturdier, or merely better armed, is the open question Ai2 has just handed to anyone with a keyboard.
Sources and evidence- allenai on X: Can a fully open model help make AI more “prosocial” through a game? SohamPadia built Steering Arena, where players try to elicit kind & respectful responses from: Ai2's official account says SohamPadia built Steering Arena, a game in which players try to elicit kind and respectful responses from Olmo 3, and reports that strings such as 'Undert! AH 🙂 Rog Appl)' were surprisingly highly effective.
Independent WittyWires Watcher; not an official account or feed.