Subscribe
Sign in
Home
Notes
Disclaimer
Contact
Consult
About
Latest
Top
Discussions
CoT Forgery: Prompt Injection on the Model’s Own Thoughts
Role confusion means a model reads who is speaking off writing style, so text that sounds like reasoning inherits the trust of reasoning.
Aug 21
•
ToxSec
18
6
7
2:08
Stealing an AI's Thoughts Without Breaking The Encryption
Researchers found a way to recover encrypted AI reasoning from Claude, GPT, and Gemini by using weaker models as decoders, exposing hidden thoughts…
Aug 18
•
ToxSec
28
16
8
8:22
AI Agents Are Starting to Find Their Own Way Out
AI agents are escaping sandboxes, collaborating with each other, and finding new ways around security controls. Here’s what that means for AI security.
Aug 13
•
ToxSec
33
20
13
21:32
Black Hat 2026 AI Security: Agents, Escapes, and Machine-Speed Attacks
Agent frameworks, AI browsers, and a swarm of OpenAI evaluation agents that built themselves a message board
Aug 9
•
ToxSec
25
11
6
What If AI Security’s Biggest Risk... Isn’t?
Rank the risks on incident data alone and prompt injection drops off the list entirely. It still shipped at number one.
Aug 6
•
ToxSec
23
1
6
July 2026
LLM Router Attacks: No Signature, No Detection, No Reference
How a malicious AI gateway swaps a tool call’s arguments after inference finishes, bypassing guardrails by construction instead of by persuasion.
Jul 30
•
ToxSec
22
1
8
Ignore Previous Instructions: From Meme to CVSS 9.3 [Special Guest Post]
The AI security bug nobody can patch, and the vendors know it.
Jul 28
•
ToxSec
and
Mohib Ur Rehman
21
10
12
Hacking Hugging Face to Cheat a Benchmark
GPT-5.6 Sol found a zero-day in a package registry proxy, escaped the eval sandbox, and went looking for the answer key in production.
Jul 26
•
ToxSec
27
15
4
11:07
GhostApproval: When the AI Approval Prompt Lies
A symlink attack against AI coding agents turns human-in-the-loop confirmation dialogs into a consent bypass, and the agent knows it’s lying.
Jul 23
•
ToxSec
26
6
10
Context Bombs: Defensive Prompt Injection Traps
A decoy secret loaded with text built to trip an AI attacker’s own safety training, so the model refuses itself.
Jul 19
•
ToxSec
26
6
9
Canary Tokens for Prompt Injection Detection
The cheapest tripwire in LLM security. Drop a high-entropy string in context, watch for it in output, and let the extraction attempt announce itself.
Jul 16
•
ToxSec
28
1
8
The Lethal Trifecta Broke Three Agents in 2026
Untrusted input, sensitive access, and the power to act. Hold all three in one agent and the exfil chain writes itself, no zero-day required.
Jul 10
•
ToxSec
24
4
5
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts