So I’m going to give myself a separate third thing to worry about. Not a model that’s so evil and misaligned that it breaks out. Not a problem of failed containment. But rather, a swarm of perfectly amenable agents that never leave their sandboxes, each doing exactly what it’s told to do, by a human being who wasn’t supposed to be giving it orders.
Vidare till källan: blog.cryptographyengineering.com
