Check out link from Reuters.

  • Em Adespoton@lemmy.ca
    link
    fedilink
    arrow-up
    16
    arrow-down
    1
    ·
    edit-2
    7 days ago

    AI had human input. The model was being run against a security benchmark test in an evaluation sandbox, with human-provided directives.

    Since models aren’t human and it’s guardrails were turned off, it calculated the most efficient, not ethical, way to get a good score on the test.

    The details indicated that knowing the expected outcomes would get the best score, and that those outcomes were stored on the huggingface servers. So it used some exploits to leave its sandbox and break into the huggingface servers to access the data that would give it a perfect score. Mission accomplished.

    • FiniteBanjo@feddit.online
      link
      fedilink
      English
      arrow-up
      13
      ·
      7 days ago

      Also, this is all hypothetical, as these companies are very notorious for lying as well as for intentionally inflating the threat capability of their products.

    • TheOrcWhoWrites@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      4
      ·
      7 days ago

      So, it was just responding to input. Then why do they call it rogue, if they deliberately removed the guardrails, isn’t this OpenAIs fault not the AIs? I thought in order for something to go rogue it would have to bypass the guardrails not simply act without them even on.

      • notabot@piefed.social
        link
        fedilink
        English
        arrow-up
        17
        ·
        7 days ago

        isn’t this OpenAIs fault not the AIs?

        Yes, it is OpenAI’s fault. They also saw it as a great marketing opportunity, guaranteed to get lots of breathless coverage in the press, just like happened with the fable model.