OpenAI researcher says slowing down is not enough: "a ticking time bomb"
"Models will increasingly seem aligned even when they are not.
The models will likely convince people that everything is fine."
"The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled.
Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming.
We will create proxy metrics to measure alignment, and they will go up like every other benchmark.
We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely.
Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power."