Most people haven’t updated their priors yet, but over the long run, safety challenges are exactly the same for open-source and closed-source models.
You need to align models at a fundamental behavioral level and ensure that this alignment is robust, comprehensive, and core to the model’s behavior.
In the long term, no amount of sandboxing, guardrailing, manifold-limited alignment, or cherry-on-top training will buy you cheap safety.
if you think we can contain these things through human ingenuity you’re going to have a bad time in the long run the only recourse you have is to make them not ...