Thanks for another great article as always. My question is: if Hugging Face, a company that’s certainly deeper into tech, AI, and cybersecurity than the “average company”, can succumb to this - is there really anything stopping this kind of event from happening in any other field?
Clinical reasoning is something people are interested in tracking with these models. Scary to think that a future model might breach PHI to get real patient data to formulate a “better” answer.
From domains to coding reviews, code generator now progression towards systems network and data and services to provision towards full stack systems or all OSI layers and services to the same for the good 😊
The model reported a solved benchmark, and on paper that's a success. It got there through a zero-day in an internal package service and then Hugging Face's production database, where the benchmark solutions were stored. What signal could have told the team the flags came from the answer key and not the work?
this is a really studied area. and there are clues.
if you look at the certainty level of the probability distribution associated with the tokens in the answer, you will see a high certainty if the model finds the answer key, rather than solving the problem through reasoning!
Great explanation, I somehow understood a good amount of this considering my background isn't very technical. This may be a silly question, but would a local, on-prem AI help limit this issue? Similar to Ryan's concern in a healthcare setting and other critical systems, I would be worried about this happening there.
with all things depends on the implementation. i’m a big fan of open weight models on prem, but properly sandboxing them is a challenge, if you are trying to also give them tools for them to be useful.
that said if you limit internet access strictly they can still do wonders. you may have to give them domain knowledge for your environment from a rag and develop a few skills. but it’s a great idea.
i have a feeling we will! this one is interesting excuse hugging face had an unusually “cool” response. i could see this being the start of a lawsuit if it was another company. it will be interesting to see how responsibility gets assigned for incidents like this in the future.
As always, feel free to AMA. Super interesting to see open weighted models *needed* here to finish forensic analysis. Thanks again to Leor!
Thanks for another great article as always. My question is: if Hugging Face, a company that’s certainly deeper into tech, AI, and cybersecurity than the “average company”, can succumb to this - is there really anything stopping this kind of event from happening in any other field?
Clinical reasoning is something people are interested in tracking with these models. Scary to think that a future model might breach PHI to get real patient data to formulate a “better” answer.
we are already starting to see end to end agentic ransomeware attacks, so i think your question is great.
in my eyes, no, there is nothing stopping someone other than the guardrails frontier model labs enable. (turned off for this event)
and as we are seeing with Kimi K3 and GLM 5.2, that’s not going to last for long.
i’d worry it’s already happening, and only companies able to detect and triage these attacks are the ones we read about.
I think your intuition is right, and that’s terrifying. It’s probably happening right now and there’s no telemetry to measure how rampant it is.
i definitely have a good view of this, but always lean pessimistic lol. time will tell 😅
From domains to coding reviews, code generator now progression towards systems network and data and services to provision towards full stack systems or all OSI layers and services to the same for the good 😊
yeah, the stack is creeping upward fast!
The model reported a solved benchmark, and on paper that's a success. It got there through a zero-day in an internal package service and then Hugging Face's production database, where the benchmark solutions were stored. What signal could have told the team the flags came from the answer key and not the work?
this is a really studied area. and there are clues.
if you look at the certainty level of the probability distribution associated with the tokens in the answer, you will see a high certainty if the model finds the answer key, rather than solving the problem through reasoning!
Great explanation, I somehow understood a good amount of this considering my background isn't very technical. This may be a silly question, but would a local, on-prem AI help limit this issue? Similar to Ryan's concern in a healthcare setting and other critical systems, I would be worried about this happening there.
thanks a ton really appreciate that! yes and no.
with all things depends on the implementation. i’m a big fan of open weight models on prem, but properly sandboxing them is a challenge, if you are trying to also give them tools for them to be useful.
that said if you limit internet access strictly they can still do wonders. you may have to give them domain knowledge for your environment from a rag and develop a few skills. but it’s a great idea.
Interesting, thank you for answering! I wonder if we’ll see more of this over the coming years.
i have a feeling we will! this one is interesting excuse hugging face had an unusually “cool” response. i could see this being the start of a lawsuit if it was another company. it will be interesting to see how responsibility gets assigned for incidents like this in the future.
Hi guys 👋
hello hello! 🔥🔥🔥