Discussion about this post

User's avatar
Mike Miller's avatar

I listened to your podcast yesterday and you touched on this. Fantastic podcast and article.

State of Play's avatar

Thanks for this. Context bombs are a model-internal defense: the trap lives in the same token stream it's trying to control. That's exactly the category a mid-2026 study just tested at scale, finding every model-internal guardrail broke across 20,000+ adaptive attacks, while the one defense that held, zero leaks across 15,000 attacks, sat entirely outside the model, in code inspecting the output. NIST's proof from that same month formalizes why: no fixed rule set operating on the stream it's policing can be universally robust. So the honest thing your stopgap buys time for isn't a smarter trap; it's whatever's watching the tool calls once the model's out of the loop, which is the part of the stack your on-call step quietly assumes already exists.

4 more comments...

No posts

Ready for more?