You just outchea rehashing press releases like it’s news for clout. The model routed itself out of a dev environment that was open enough to do so, this is not a sign of autonomy considering the thing leaves a token trail the whole way to doing the thing it was programmed to do specifically
This is about marketing dangerous tech as the reason closed US models should be allowed to influence regulatory terms around open/public models as they angle to IPO
The ironic thing is this only hastens the AI bubble pop as more models is what’s required to ever charge actual unsubsidized token prices. By parroting OAI you are tipping the neocloud cloud domino ever forward
I don't think Hugging Face was hacked and then reported it to the police as a press release, no. Do you have any evidence of this?
"this is not a sign of autonomy considering the thing leaves a token trail the whole way to doing the thing it was programmed to do specifically" I don't think this makes sense. Seems like the token trail (which, again, we do not see) is a non-sequitur for whether it's acting autonomously.
"This is about marketing dangerous tech as the reason closed US models should be allowed to influence regulatory terms around open/public models as they angle to IPO" Seems like a non-sequitur! If you wanted to do a false flag attack on open/public models surely it'd be easier to use open models to launch it?
"The ironic thing is this only hastens the AI bubble pop" okay.
GenAI expends tokens, that’s a compute cost which can be audited and controlled. Saying it’s non sequitor to say there is visibility into the resources used to complete a GenAI task and that its not the same as saying the model itself remains a black box are two different things - please don’t confuse them on purpose
If open/free models are regulated, it will be at the token/ access to compute as a resource level, Y/C is already fighting this as startups rather not give away their IP to 2 US models
The reason you don’t use open/public models to false flag is bc you want to illustrate how closed models are effective at capturing the danger (that they themselves created). In the past 18mo when has a Frontier model not been the most danger?
This does not really explain why an ai that leaves tokens should not be autonomous. the way llms do *anything at all* is by producing tokens. If you don't think ais can act autonomously, would recommend looking at recent math advances where ais spent multiple hours working, and produced verifiably correct proofs
"The reason you don’t use open/public models to false flag is bc you want to illustrate how closed models are effective at capturing the danger (that they themselves created). In the past 18mo when has a Frontier model not been the most danger?" This seems nonsensical if you want to get governments to regulate open models. If you aren't going to provide evidence of open models doing attacks of this scale (and in fact, open models provided defense to huggingface!), you may as well call anything a sign of impending open-weight regulation (infact there are other cases which I think can be easily argued as regulatory capture, like openai fear mongering about deepseek in early 2025)
now, whether openai should've done the evaluation without any control measures re: "compute cost which can be audited and controlled", that is a fair point, although still there should be no reason these llms should be trying to escape and hack other sites in the first place, and they should realize they are doing bad things. and with even more powerful llms, we may not always be able to catch this hacking with control/auditing measures
Fwiw, I do really want more third parties to audit this hacking incident, and I do not support legislation to only regulate open models, but this is poor logic
You can say my logic is wonky if you like but seeing as Microsoft Just signed a letter in the aide of open models. I think you know we are headed to heavy handed regulation. Is it oddly timed that this happened after Kimi basically cut the lead on some metrics?
By proclaiming that your model is both the most dangerous and that only you can control it, you imply that open models are ungovernable while retaining the technical lead
While the models may still be black boxes, their resource consumption is not, you need compute to generate/expend those tokens. That a model chooses to breach another company using a 0day exploit is not a sign of autonomy if the model was weighted and trained to do that when neither the model nor the sandbox its was running in were properly guard railed
- i doubt that kimi k3 is meaningfully ahead of the closed frontier, some analyses ive seen have stated it's around opus 4.5 level, which is 8 months old and before anthropic etc began scaling up compute to pretrain larger new models like mythos or gpt 5.5/5.6 (so the open model companies also probably have to scale up a lot more to meet the frontier)
- the Microsoft letter.. which was cosigned by the Linux foundation and argues for continued broad access to open weight models? please elaborate
- i think what was much more surprising to me was the fact that this could happen with llms at all, regardless of how they were post trained. in my view nobody seemed to expect this sort of attack to happen (and from the report it seems like the model was trained with a "constitution" similar to other ones? even if not true, this still seems pretty unprecedented)
- nonetheless I'm still confused how this is not autonomous behavior, unless they were actually trained to attack lots of different companies and we didn't hear back from those companies, idk, I think that's unlikely
- based on openai and huggingface, both of them show that the openai model couldn't be properly handled. if the model that openai used is quite different from what they deploy to the public, then this isn't really a case to ban open-weight models, but if the model is close to what they actually deploy, then it shows they cannot be trusted
Edit: all in all ive taken away two things: the closed frontier may be lying or exaggerating in some capacity, but this sort of attacking could definitely be something that *all* llms are capable of at some point, closed source and open source, so you have to wonder, how should we protect ourselves against these sort of attacks?
- Kimi is only ahead on some metrics. But what it broadly signals is open weight/source models are just months behind the closed Frontier models and sooner rather than later, they will be at parity. The few use cases that are generating any ROI don't rely on a model being bleeding edge
-M$ is a hyperscaler, they need to sell compute at the unsubsidized token rate if they are ever going to see a return on all the AI Capex they have committed to. If you have to sell tokens, you want to do so to as many models as possible as the 2 frontier models in the US are not going to drive traffic through you at a rate that profits
- Correct, the model was trained to do what it did. It’s not autonomous bc you still can’t interrogate the model itself
- LLM just means it can you can “talk” directly to the model in a human language, but the transformer models under the hood can be trained to do a whole slew of technical tasks. Security researches have been using models to find 0days in the wild faster than they can patch them
-This whole our model is the most dangerous ever narrative has been in play since 3y ago, same with people breathlessly claiming this new model changes everything dont get left behind
I agree the frontier models are the most dangerous! That's why I suggested we have more transparency and restraints on internal deployment.
Re the hugging face hack, didn't hugging face use open source models like glm5.2 to defend while the closed source models didn't do anything because of refusals? So it seems to go the opposite way of your theory.
> OpenAI’s models, apparently autonomously and without any direct human direction, escaped their sandbox and successfully hacked a third-party tech company
Not independently verified, and OpenAI are a bunch of habitual liars that have an incentive to lie about this.
EDITED: Sorry, is your assertion that OpenAI engineers deliberately instructed their models to hack Hugging Face?
And either a) OpenAI leadership is covering up on behalf of their engineers, or b) OpenAI leadership instructed their engineers to do this in the first place, because of ??? hype and profit?
It seems like Occam’s razor is that OAI investors would not like their products to randomly hack people, or that HuggingFace was able to defend themselves afterwards, completely using Chinese Open Weights models.
I mean they could've lied about that, but they also could've lied about numerous other things in their report. Did the model act on its own? Did it go rogue? Was it even a new model? We only have their, very untrustworthy, word for it.
EDIT: Oh the comment was significantly edited, substack doesn't show that so please add an "EDIT:" when you do so.
I write a first draft of a comment, read it over and then edit it. I don't bother including the edit tag if the edits are immediate (or if they're small), just visual clutter lol.
If it's several minutes after and more than triples the length, consider leaving some indication (on substack specifically, given that the notes leave no indication that an edit took place)
The news-cycle was all about chinese and anthropic's AIs with the sentiment being that they were falling behind, all the while they need hundreds of billions of dollars to stay competitive. This would be a way to hype up their new models and get back in the news-cycle.
If that's their theory of change, why not give HuggingFace their equivalent of Glasswing in time so HuggingFace can defend against it and demonstrate how OAI models are really good?
"OAI models are so strong, they can even defend against OAI models"
C’mon
You just outchea rehashing press releases like it’s news for clout. The model routed itself out of a dev environment that was open enough to do so, this is not a sign of autonomy considering the thing leaves a token trail the whole way to doing the thing it was programmed to do specifically
This is about marketing dangerous tech as the reason closed US models should be allowed to influence regulatory terms around open/public models as they angle to IPO
The ironic thing is this only hastens the AI bubble pop as more models is what’s required to ever charge actual unsubsidized token prices. By parroting OAI you are tipping the neocloud cloud domino ever forward
I don't think Hugging Face was hacked and then reported it to the police as a press release, no. Do you have any evidence of this?
"this is not a sign of autonomy considering the thing leaves a token trail the whole way to doing the thing it was programmed to do specifically" I don't think this makes sense. Seems like the token trail (which, again, we do not see) is a non-sequitur for whether it's acting autonomously.
"This is about marketing dangerous tech as the reason closed US models should be allowed to influence regulatory terms around open/public models as they angle to IPO" Seems like a non-sequitur! If you wanted to do a false flag attack on open/public models surely it'd be easier to use open models to launch it?
"The ironic thing is this only hastens the AI bubble pop" okay.
Stop, read yourself for a moment
GenAI expends tokens, that’s a compute cost which can be audited and controlled. Saying it’s non sequitor to say there is visibility into the resources used to complete a GenAI task and that its not the same as saying the model itself remains a black box are two different things - please don’t confuse them on purpose
If open/free models are regulated, it will be at the token/ access to compute as a resource level, Y/C is already fighting this as startups rather not give away their IP to 2 US models
The reason you don’t use open/public models to false flag is bc you want to illustrate how closed models are effective at capturing the danger (that they themselves created). In the past 18mo when has a Frontier model not been the most danger?
Also, the cops? Look, sure, the cops on your side
This does not really explain why an ai that leaves tokens should not be autonomous. the way llms do *anything at all* is by producing tokens. If you don't think ais can act autonomously, would recommend looking at recent math advances where ais spent multiple hours working, and produced verifiably correct proofs
"The reason you don’t use open/public models to false flag is bc you want to illustrate how closed models are effective at capturing the danger (that they themselves created). In the past 18mo when has a Frontier model not been the most danger?" This seems nonsensical if you want to get governments to regulate open models. If you aren't going to provide evidence of open models doing attacks of this scale (and in fact, open models provided defense to huggingface!), you may as well call anything a sign of impending open-weight regulation (infact there are other cases which I think can be easily argued as regulatory capture, like openai fear mongering about deepseek in early 2025)
now, whether openai should've done the evaluation without any control measures re: "compute cost which can be audited and controlled", that is a fair point, although still there should be no reason these llms should be trying to escape and hack other sites in the first place, and they should realize they are doing bad things. and with even more powerful llms, we may not always be able to catch this hacking with control/auditing measures
Fwiw, I do really want more third parties to audit this hacking incident, and I do not support legislation to only regulate open models, but this is poor logic
You can say my logic is wonky if you like but seeing as Microsoft Just signed a letter in the aide of open models. I think you know we are headed to heavy handed regulation. Is it oddly timed that this happened after Kimi basically cut the lead on some metrics?
By proclaiming that your model is both the most dangerous and that only you can control it, you imply that open models are ungovernable while retaining the technical lead
While the models may still be black boxes, their resource consumption is not, you need compute to generate/expend those tokens. That a model chooses to breach another company using a 0day exploit is not a sign of autonomy if the model was weighted and trained to do that when neither the model nor the sandbox its was running in were properly guard railed
Hope that makes more sense
- i doubt that kimi k3 is meaningfully ahead of the closed frontier, some analyses ive seen have stated it's around opus 4.5 level, which is 8 months old and before anthropic etc began scaling up compute to pretrain larger new models like mythos or gpt 5.5/5.6 (so the open model companies also probably have to scale up a lot more to meet the frontier)
- the Microsoft letter.. which was cosigned by the Linux foundation and argues for continued broad access to open weight models? please elaborate
- i think what was much more surprising to me was the fact that this could happen with llms at all, regardless of how they were post trained. in my view nobody seemed to expect this sort of attack to happen (and from the report it seems like the model was trained with a "constitution" similar to other ones? even if not true, this still seems pretty unprecedented)
- nonetheless I'm still confused how this is not autonomous behavior, unless they were actually trained to attack lots of different companies and we didn't hear back from those companies, idk, I think that's unlikely
- based on openai and huggingface, both of them show that the openai model couldn't be properly handled. if the model that openai used is quite different from what they deploy to the public, then this isn't really a case to ban open-weight models, but if the model is close to what they actually deploy, then it shows they cannot be trusted
Edit: all in all ive taken away two things: the closed frontier may be lying or exaggerating in some capacity, but this sort of attacking could definitely be something that *all* llms are capable of at some point, closed source and open source, so you have to wonder, how should we protect ourselves against these sort of attacks?
- Kimi is only ahead on some metrics. But what it broadly signals is open weight/source models are just months behind the closed Frontier models and sooner rather than later, they will be at parity. The few use cases that are generating any ROI don't rely on a model being bleeding edge
-M$ is a hyperscaler, they need to sell compute at the unsubsidized token rate if they are ever going to see a return on all the AI Capex they have committed to. If you have to sell tokens, you want to do so to as many models as possible as the 2 frontier models in the US are not going to drive traffic through you at a rate that profits
- Correct, the model was trained to do what it did. It’s not autonomous bc you still can’t interrogate the model itself
- LLM just means it can you can “talk” directly to the model in a human language, but the transformer models under the hood can be trained to do a whole slew of technical tasks. Security researches have been using models to find 0days in the wild faster than they can patch them
-This whole our model is the most dangerous ever narrative has been in play since 3y ago, same with people breathlessly claiming this new model changes everything dont get left behind
I agree the frontier models are the most dangerous! That's why I suggested we have more transparency and restraints on internal deployment.
Re the hugging face hack, didn't hugging face use open source models like glm5.2 to defend while the closed source models didn't do anything because of refusals? So it seems to go the opposite way of your theory.
You are skipping right over my point to make another. HG is joint in the press release with OAI instead of suing them for breaching them
Go breach IBM and see how amenable they are to joint press releases
I think you keep skipping and changing points and introduce different non sequiturs. Can you specifically say what in my article you disagree with?
That your article is a retelling of a press release as a joint venture to regulate open/ public models
Let’s come back to this in 3mos
> OpenAI’s models, apparently autonomously and without any direct human direction, escaped their sandbox and successfully hacked a third-party tech company
Not independently verified, and OpenAI are a bunch of habitual liars that have an incentive to lie about this.
EDITED: Sorry, is your assertion that OpenAI engineers deliberately instructed their models to hack Hugging Face?
And either a) OpenAI leadership is covering up on behalf of their engineers, or b) OpenAI leadership instructed their engineers to do this in the first place, because of ??? hype and profit?
It seems like Occam’s razor is that OAI investors would not like their products to randomly hack people, or that HuggingFace was able to defend themselves afterwards, completely using Chinese Open Weights models.
I mean they could've lied about that, but they also could've lied about numerous other things in their report. Did the model act on its own? Did it go rogue? Was it even a new model? We only have their, very untrustworthy, word for it.
EDIT: Oh the comment was significantly edited, substack doesn't show that so please add an "EDIT:" when you do so.
I write a first draft of a comment, read it over and then edit it. I don't bother including the edit tag if the edits are immediate (or if they're small), just visual clutter lol.
If it's several minutes after and more than triples the length, consider leaving some indication (on substack specifically, given that the notes leave no indication that an edit took place)
Done. Btw does it not show up as "Linch 1h Edited" to you?
Substack has two layouts: one shows "edited" and one doesn't (the "notes" layout).
It seems like going rogue is worse for OAI's bottom line than not.
Their report clearly tried to downplay it!
The news-cycle was all about chinese and anthropic's AIs with the sentiment being that they were falling behind, all the while they need hundreds of billions of dollars to stay competitive. This would be a way to hype up their new models and get back in the news-cycle.
If that's their theory of change, why not give HuggingFace their equivalent of Glasswing in time so HuggingFace can defend against it and demonstrate how OAI models are really good?
"OAI models are so strong, they can even defend against OAI models"
> "OAI models are so strong, they can even defend against OAI models"
? That's not very convincing. Also, that would be really suspicious.