Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> This incident occurred during an internal evaluation which prompts models [with safeguards disabled for evaluation purposes] to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.

Researcher: hack me

Model: understood

Researcher: oh my god



You're aware that HuggingFace notified law enforcement about this incident? Was that OpenAI's intended outcome when they prompted their AI?


> You're aware that HuggingFace notified law enforcement about this incident?

How will this affect OpenAI?


They got a bunch of publicity and nothing bad (or at least, that their lawyers can’t handle) will happen


They will probably get stricter AI regulations which is actually what they've been pushing for years. So that's a funny outcome to the whole thing.


This just in:

> Lawmakers push for AI kill switch after OpenAI's models go rogue

https://www.bbc.com/news/articles/cx2vqj2e9x8o


OpenAI has not been pushing for stricter regulations.


> 2023 - OpenAI’s Sam Altman Urges A.I. Regulation in Senate Hearing

https://www.nytimes.com/2023/05/16/technology/openai-altman-...


But it wasn't explicitly told to hack HuggingFace. It was told "answer this security question", and it's answer was to break into the teacher's desk to find the answer key.


Researcher: hack me

Model: I committed a crime

Researcher: oh my god


Me to a random person: hack out of a secure environment into another secure environment.

Random person: I have no clue or ability to do that.


Well, I don't think even 1% of professionals would be able to. The model used novel zero-days, there are many people in this filed, not many of them discover such vulnerabilities, and probably not on the spot.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: