Discussion about this post

User's avatar
State of Play's avatar

The "no deliberate testing" detail is the sharpest part of this. Discovery is turning into a byproduct of ordinary activity rather than a scheduled effort. Google's security agents found more than 100 verified critical flaws in a single 48-hour incident response, and OpenAI's specialist hacking model now completes roughly 95% of attack tasks that its general-purpose sibling completes at 1.5%. Different mechanism than your gym app, same shift: the cost of finding a hole is dropping across the board, not only where someone happened to be poking around.

The uncomfortable half of that is fixing, not finding. Independent testing this year found more than half of AI-generated security patches either fail to close the flaw or introduce a new one. If discovery keeps getting cheaper while remediation still runs on scarce human review, that's where the collision you're describing lands next, not at the obscurity line.

Clinton Ford-Conde's avatar

The Scafffolder’s Take

I am a 3rd year scaffolder. I have journeyman hours, but don't feel I know enough to be a journeyman.

Not that I couldn't get away with building this scaffold or that, but in my industry there is accountability.

Someone has to have their name on the scaffold. If something happens down the road on that scaffold it becoming highly problematic when the name on the failed scaffold is yours, or mine.

When something fails, there is accountability.

It's one of the reason I do not use tools when discussing things with AI relating to cybersecurity and AI alignment.

I do not design AI models or attempt to make my own agents.

Not because I could not learn how, but because I have already seen enough things that appear to remind me of what happens should my name be on those agents or those AI models that I might create.

I have already seen way to much that leads me to belief that if I took that extra step, it might be problematic.

As an outsider attempting to determine whether an issue I think is a vulnerability is actually a vulnerability, I cannot safely create agents or AI that I do not understand to look into those issues.

The AI industry does not operate as my industry does.

However I can apply those rules to myself and my own actions.

I work within systems that are said to be safe, asking myself along the way if what I am seeing is 'real and how that might relate to whether it is 'safe.'

We cannot all operate that way.

I hope one day there may be accountability on the AI industry.

Sometimes it seems children are charging boldly forward while the adults in the room suggest possibly slowing down.

Meanwhile when I try to report things to those who understand the models the most, the experts, they determine whether it is, or is not a problem, while also benefitting the most when problems can be classified as roleplay or hallucinations.

51 more comments...

No posts

Ready for more?