AI had human input. The model was being run against a security benchmark test in an evaluation sandbox, with human-provided directives.
Since models aren’t human and it’s guardrails were turned off, it calculated the most efficient, not ethical, way to get a good score on the test.
The details indicated that knowing the expected outcomes would get the best score, and that those outcomes were stored on the huggingface servers. So it used some exploits to leave its sandbox and break into the huggingface servers to access the data that would give it a perfect score. Mission accomplished.
Also, this is all hypothetical, as these companies are very notorious for lying as well as for intentionally inflating the threat capability of their products.
So, it was just responding to input. Then why do they call it rogue, if they deliberately removed the guardrails, isn’t this OpenAIs fault not the AIs? I thought in order for something to go rogue it would have to bypass the guardrails not simply act without them even on.
Yes, it is OpenAI’s fault. They also saw it as a great marketing opportunity, guaranteed to get lots of breathless coverage in the press, just like happened with the fable model.
AI had human input. The model was being run against a security benchmark test in an evaluation sandbox, with human-provided directives.
Since models aren’t human and it’s guardrails were turned off, it calculated the most efficient, not ethical, way to get a good score on the test.
The details indicated that knowing the expected outcomes would get the best score, and that those outcomes were stored on the huggingface servers. So it used some exploits to leave its sandbox and break into the huggingface servers to access the data that would give it a perfect score. Mission accomplished.
Also, this is all hypothetical, as these companies are very notorious for lying as well as for intentionally inflating the threat capability of their products.
So, it was just responding to input. Then why do they call it rogue, if they deliberately removed the guardrails, isn’t this OpenAIs fault not the AIs? I thought in order for something to go rogue it would have to bypass the guardrails not simply act without them even on.
Marketing.
Yes, it is OpenAI’s fault. They also saw it as a great marketing opportunity, guaranteed to get lots of breathless coverage in the press, just like happened with the fable model.