Discussion about this post

User's avatar
ContextWindow's avatar

(Quick thoughts on angles that seem underfocused in a lot coverage of the HuggingFace breach, and contra to this episode's value as a "warning shot")

Hugging Face's defence appears to have depended materially on open-source models and these models were surprisingly effective. If so, the episode is evidence not only for open-source models in general (contra to some strands in AI safety), but also for organizations maintaining models and infrastructure they control directly.

--- Note that, while rectified afterwards, OpenAI's regular APIs refused Hugging Face's attempts to use them during the incident. (In most contexts this would be wild, morally and legally: OpenAI's model was the bad actor here, plausibly OpenAI was in serious breach of the law and its safeguards blocked the victim's defensive use). I think many smaller companies would be doubtful of getting the same arrangement of being folded into OpenAI's security program.

--- I think a leader or organization who did not believe in radical AI transformation would read this episode as showing both that LLM literacy and effectiveness (e.g. open-source models) matter, and that standard methods work.

This is a deep question, but it seems plausible that defense can be asymmetrically favored over offense. In this case, quickly "reading the logs" would have revealed the intrusion. A simple, GPT-4o-level model would have provided much or all of that functionality: certainly for OpenAI's benching team, and probably to a large degree for Hugging Face.

--- This wasn't a sophisticated attack, the model did not try to hide its trail or manipulate a defensive LLM, but a guess is that a sophisticated detection system could be robust, while only using lower-tier LLM tools.

Bob Jacobs's avatar

> OpenAI’s models, apparently autonomously and without any direct human direction, escaped their sandbox and successfully hacked a third-party tech company

Not independently verified, and OpenAI are a bunch of habitual liars that have an incentive to lie about this.

34 more comments...

No posts

Ready for more?