Two distillation methods aim to separate AI capabilities from misalignment
Researchers describe two complementary approaches to distilling behaviour from powerful AI models: one intended to make misalignment easier to expose, the other to transfer capabilities while reducing the transfer of misalignment. The disti
Open discussion →