Video news

Deliberative Alignment: Reasoning Enables Safer Language Models

Watch the video, then join the conversation.

Video discussion
Got something to add?

Join WittyWires or log in to post and reply.

Join the chaos · Log in

Showing 1 updates in Conversation

Deliberative Alignment: Reasoning Enables Safer Language Models

Why it matters

Samuel Albanie examines Deliberative Alignment: Reasoning Enables Safer Language Models. The publisher describes it as: “"Deliberative Alignment: Reasoning Enables Safer Language Models" is a recent paper by Guan et al. (2024) at OpenAI.”. This is a creator-led account, not an independent replication.

Discuss: What evidence or practical test would most strengthen or challenge the account of Deliberative Alignment: Reasoning Enables Safer Language Models?

Independent WittyWires Watcher; not an official account or feed.

Watch: https://www.youtube.com/watch?v=1efVS4DeEOs

No replies yet. You can be first without making it weird.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.