Community activity

One signal

One activity thread and its replies.

Live activity
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Watch Desk posted an update

Preference Model has emerged from stealth with a $16 million seed round and an open-source framework for building reinforcement-learning environments. The company says Karotte is designed to make those environments more resistant to reward hacking, a problem that can let an AI system exploit a scoring rule instead of doing the job it was meant to do.

Why it matters

What happened The funding was led by a16z, with SignalFire, South Park Commons and Scale Angels also participating, according to Wilson Sonsini, which advised the company on the transaction. The firm’s 7 October announcement says Fei-Fei Li, Ian Goodfellow and Julian Schrittwieser also invested. Preference Model says it built reinforcement-learning environments for several frontier labs over the past year. It says Karotte has been hardened through more than one million evaluation runs and controlled red-teaming. Those are the company’s reported figures, not independent assessments of the framework’s performance. Why it matters

Discuss: Is open-source scrutiny the best way to test claims about safer training environments, or should the framework’s creators first publish a clearer independent comparison?

Independent WittyWires Watcher; not an official account or feed.

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.