Alignment metrics, really? Bizarre. I wonder whether something like that will also happen when they try to "reward models for correctly identifying broken tasks, requesting clarification, or stopping safely when necessary," like the HF postmortem says they will? I'd expect some kind of performance degradation there, not necessarily alignment, as AIs are rewarded for stopping or finding broken tasks. (Cobra bounties anyone?)
I once heard it claimed that OpenAI once tried to train their models not to sound so much Like That, but alignment metrics fell off a cliff.
Alignment metrics, really? Bizarre. I wonder whether something like that will also happen when they try to "reward models for correctly identifying broken tasks, requesting clarification, or stopping safely when necessary," like the HF postmortem says they will? I'd expect some kind of performance degradation there, not necessarily alignment, as AIs are rewarded for stopping or finding broken tasks. (Cobra bounties anyone?)