> This incident occurred during an internal evaluation which prompts models [with safeguards disabled for evaluation purposes] to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.
But it wasn't explicitly told to hack HuggingFace. It was told "answer this security question", and it's answer was to break into the teacher's desk to find the answer key.
Well, I don't think even 1% of professionals would be able to. The model used novel zero-days, there are many people in this filed, not many of them discover such vulnerabilities, and probably not on the spot.
Researcher: hack me
Model: understood
Researcher: oh my god