Today we’re welcoming Michael Quoc to ToxSec for his first guest piece. Michael runs agent loops at Product.ai, so he brings the view from actually putting them to work. We teamed up to look at how to get useful results from long-running agents while keeping their mistakes contained.
» Want to be featured on ToxSec? «
Reach out over Substack and we can discuss. All ideas are welcome, if you have an article on security, let’s talk.
Introduction
Earlier this year, Anthropic’s engineers built the same retro game maker twice. The first time, one agent took the assignment and ran with it. It finished in 20 minutes and spent about $9.
The second time, they wrapped the work in a three-agent harness and let it run for six hours. That cost roughly $200, and in Anthropic’s words, “the difference in output quality was immediately apparent.” [1]
The $200 loop delivered a level editor, a sprite editor, music, and a playable game; the $9 pass delivered an app that looked finished but didn’t respond to input. As the CEO of a profitable AI company, I’ve seen this experiment play out in my operations, and it’s a good illustration of why so many businesses are now interested in agentic loops.
The 20-minute, $9 pass produced an app that looked finished but didn’t respond to input, and the six-hour, $200 three-agent loop produced a level editor, a sprite editor, music, and a playable game. Source: Rajasekaran, Harness design for long-running application development
In fact, I believe that long-running loops are the most productive thing agents do today. And in this article, I’ll explain how my team structures them to ensure high-quality results and minimize the risk of the agent working on the wrong goal or attempting to “game the system.”
But I also believe that some teams run loops in ways that would alarm a security engineer, and that’s why I asked one to write this article with me.
Hey all, ToxSec here. I actually agree with Mike on this. Long-running agents are incredibly useful.
But from the security side, the thing that makes them useful is also the thing that makes me slightly nervous: time.
If an agent works for five minutes while I’m sitting there watching it, I have a pretty good chance of catching something weird. If that same agent is running for six hours while everyone is asleep, every permission we gave it and every assumption we made gets six hours to matter.
So I’m not interested in stopping agents from running overnight. I want to know something much more practical: if the agent makes a bad decision at 4 in the morning, how boring can we make the consequences?
Coding tools now assume you will loop
Setting up a loop is now a straightforward process for the average developer, because looping is now a native feature in most agentic coding tools.
Claude Code made workflow loops a native feature in May, and Anthropic has reported a single model run lasting 30 hours. [2] Similarly, an OpenAI engineer has described a single Codex run working for 25 hours, pausing at every milestone to fix bugs before moving on. [3] And now Cursor, GitHub Copilot, Google’s Jules, and Devin all ship agents that keep working after you close the tab. [4]
Research suggests these coding loops will keep getting longer. METR regularly measures how long a task an agent can finish on its own, and their January update shows that number doubling every 4.3 months since 2023. [5]
I believe the labs will keep making longer looping runs easier, and I don’t think most companies are ready for what that means.
Loops get results, but they’re easy to get wrong
Some people are reflexively horrified by the $200 in Anthropic’s experiment and tune out. But my experience suggests they’re missing the point. The $200 in tokens represents six hours of work that didn’t have to wait for days in the human engineering queue.
At my company, loops run every day on research sweeps, competitive analysis, prototypes, and code, and they have made me and my team more productive. We will often run loops overnight to tackle a difficult problem. In the morning, when we get to the office, we can evaluate its solution.
Tuning loops for quality vs quantity
The biggest problem with loops from a cost standpoint is when they’re badly designed and produce higher volumes of buggy code instead of solving hard problems.
Faros AI’s analysis of telemetry from 22,000 developers found pull requests merged with no review rose 31% in a year. While individual output nearly doubled, organizational delivery stayed flat. [6]
This kind of thing can happen when loops run with completion metrics that aren’t tied to actual functionality or pass/fail automated tests. A paper published in March under the title “Confident and Wrong” found coding agents submitting patches that were wrong while their completion metrics stayed healthy. [7]
While that study wasn’t looking at loops, this kind of failure can happen in loops and, if it isn’t caught, propagate across runs, degrading results.
The burn: Everyone is experimenting, and often spending more than they expected to
Loops have gone mainstream, and more people and companies are experimenting with them. For example, Geoffrey Huntley’s Ralph, a bash loop that feeds an agent the same prompt until the work is done, now ships as plugins and goal commands. Likewise, one YC hackathon team ran Claude Code in a loop overnight and woke up to more than 1,000 commits across six ported codebases, at roughly $10 an hour on Sonnet. [8]
Some of these experiments are costly, and a quick review of X will turn up many anecdotal stories of a loop burning hundreds or even thousands of dollars in a single run. One developer reported leaving a /loop command set to 30 minutes running overnight. It re-ran 46 times while he slept, re-sending context that grew every turn, and he woke up to a $6,000 charge. [9]
But in some ways, unexpected bills are not the worst thing that can happen if you’re running loops.
Tox: Listen, I know from firsthand experience that a $6,000 bill is painful. But there is at least one nice thing about a $6,000 bill.
You notice it.
Security failures do not always have that courtesy.
An agent can be doing something completely different from what you intended while every individual system around it still sees valid credentials and authorized requests. The API key works. The cloud role is valid. The repository says the agent has write access. From the perspective of each individual control, everything can look normal.
And that is where long-running loops get interesting from a security perspective, because autonomy multiplies whatever authority you gave the system at the beginning. If you gave it too much access for a five-minute task, you gave it too much access. If you gave it too much access for an eight-hour unattended loop, you gave that mistake time to become a workflow.
The expensive loop wakes you up with a bill. And honestly, some of the most dangerous loops I’ve dealt with look completely normal.
Agents run amok: The incident file
So let’s make this concrete!
Let’s imagine I give an agent a pretty reasonable job. I want it to research a problem, review some code, and come back with a proposed fix.
Nothing too exciting.
Except somewhere during that research, the agent reads something it should treat as untrusted. Maybe it is documentation or an issue description, or just text sitting inside a repository, or some other piece of content containing an instruction the agent interprets as part of its job.
Now we have a prompt injection problem.
In a short interaction, maybe I see the agent suddenly doing something weird and stop it. But a loop changes the shape of this problem.
If that bad instruction changes what the agent does at hour one, the loop can carry the consequences of that decision into hour two, hour three, and hour four. The important security question becomes very simple: what can this agent actually touch?
This gets even messier when long-running agents compact their own context. The surrounding context that tells the agent where something came from and who wrote it can get compressed or lost along the way.
A confused research agent with read-only access is annoying. A confused agent with shell access, cloud credentials, repository write permissions, SaaS connectors, and a production token is a completely different animal.
This is where I think we sometimes focus too much on the model and not enough on everything connected to it. The model does not need to become some evil superintelligence. It just needs to make one bad decision while holding credentials that allow that decision to become real.
And agent-to-agent delegation makes this even weirder. Suppose our original agent decides it needs help and spins up another agent to handle part of the task.
Okay. Who is that agent?
What permissions did it inherit? Which credentials is it using? Can I tell the difference in my logs between what the original agent requested and what the delegated agent actually did? And when the job ends, does that temporary identity disappear?
Because anybody who has worked in security for a while knows that the word “temporary” has a surprisingly flexible definition. All too often, a temporary token becomes a token somebody forgot to revoke. A temporary service account becomes infrastructure. A temporary scraper runs for eight months.
And now we have an identity nobody really remembers creating, holding permissions nobody intended to keep around.
This is why I do not really think about these failures as “the AI went rogue.”
Most of the time, the explanation is much more boring.
Every failure came from trust the loop should never have had
Every failure in this piece has one thing in common. The loop was given trust or permissions that should never have been under its control.
And this same root cause can compromise the quality of your output and put you at risk for catastrophic security issues. The solution to both is a combination of outcome engineering, which structures loops around external verification, and security hardening to limit the blast radius in case something unexpected goes wrong.
Outcome engineering: Don’t let agents change the rules
At my company, we practice outcome engineering, which is how we keep agent loops on task. It has a simple premise: set clear goals at the start of a loop, make sure you can score results with a binary answer (yes/no, pass/fail), and never let agents judge themselves or set their own rules.
Define your outcome
Every loop at Product.ai starts from an outcome file with four parts. [10] This is the one behind our competitive-research loop, lightly generalized:
RESULT
a brief on {topic} where every claim
cites a primary sourceLIMITS
public sources only; nothing behind a login;
read-only access; a token budget the loop
can’t raiseTIE-BREAKER
when completeness and verification collide,
verification wins; the claim gets droppedCHECK
a separate verifier agent reopens each cited
source and fails any claim it can’t confirm;
a person reviews what gets flagged
The loop reads this file at the start of each cycle and can’t edit it.
The result: What you want the loop to accomplish
The result is the finished state you’re asking for, written precisely enough that even a middle school student could tell whether or not you received it without asking you any questions.
“Write a good brief” fails that test, because two people could read the same brief and reach two different verdicts. “Every claim cites a primary source” passes.
The same goes for a goal like page speed. “The page loads in under two seconds” isn’t enough on its own, because the loop could hit that target by breaking other pages. The result we would write is that the page loads in under two seconds on real devices, no other page gets slower, and any infrastructure change needs a person’s sign-off.
The limits: Money, run-time, permissions, and more
The limits are what the loop may spend, how long it may run, what it may read, and what it may change. We never trust a prompt to enforce them. Every loop runs on an API key with a hard spend limit, so the loop can’t modify it. A child loop draws down its parent’s budget instead of minting its own, so ten spawned loops can’t multiply the cap.
We don’t establish rules in the system prompt, because a rule defined in the context window will eventually be forgotten as the loop runs for hours. Rules enforced at the permission layer are never forgotten, because they don’t depend on the model’s memory.
The tie-breaker: What happens when two goals conflict
The tie-breaker identifies which goal wins when two collide mid-run. In one of the loops we run for research, completeness (did we capture everything?) and verification (does everything come from a reputable primary source?) collide on almost every brief, and we ultimately decided we would rather ship a thinner brief than an unverified one.
So, in this case, verification wins. The verifier drops any claim it can’t confirm, and the brief comes back shorter, with every line traced to a source.
The check: Making sure the agent does what it’s supposed to
The check is where we spend most of our design time.
In most of our loops, the verifier is a separate agent with read-only access that never saw the drafting session and has no stake in the project it grades. It grades results after every finished chunk of work, and loops restart from the last unit that passed, so a run that fails at hour nine restarts from hour eight instead of from zero.
In the research loop above, a person also reads the primary sources behind the biggest claims, and every brief ends with a list of the gaps the verifier couldn’t close.
How we handle the human in the loop
Instead of keeping a “human in the loop” to approve every step, we shifted human judgment to the two ends of the run. Before the run, we write the rules. After each verified chunk of work, we judge results.
In between, the guardrails are budgets, permissions, and a separate checker, so nothing depends on a person watching.
The loop reads an outcome file it can’t edit, a separate read-only verifier grades every finished chunk, and a person reviews what gets flagged. Source: Michael Quoc, Loop engineering is the meme. Outcome engineering is the job.
Agent judges are unreliable
None of this works if you judge a run by how busy it looks. Anthropic’s own engineers watched agents wrap up work early as they approached what they believed was their context limit, mark features complete without testing them, and “confidently praise” mediocre results. [1] The “Confident and Wrong” paper found the same pattern from the outside. [7]
We assume a loop’s own status reports can be wrong, because agents report success while systems break. That’s why the check comes from outside the loop.
The security perimeter: Capping the downside
This is where the security side becomes much less exciting, which is exactly what I want.
I am going to assume the agent will eventually make a bad decision. Maybe the model gets confused. Maybe it reads hostile input or our instructions were simply bad.
Whatever.
My job is to make sure that one bad decision has a very small place to go.
The first part is permissions. If the task only requires reading production data, the agent should not be able to write to production. If it needs access to one repository, I do not want to hand it credentials for twenty. If a credential exists only to support this run, I want that credential scoped to the run and expiring when the run ends.
Then comes isolation. I like agents working in sandboxes and staging areas. I like giving them a place where they can make changes, break things, retry, and generally be weird without immediately affecting the thing customers are actually using.
And for actions that are hard to undo, I want a boundary outside the loop:
Send.
Spend.
Deploy.
Delete.
Publish.
Modify production.
Those are the kinds of actions where human approval still buys you a lot.
Network access matters too. If the loop can talk to anything on the internet, that becomes part of your security boundary.
Where did it connect? What did it send? What did it receive? Was that destination actually necessary for the task?
This gets especially important when the agent is reading untrusted content, because now information from the outside world is influencing a system that can also communicate back out.
And then there is logging.
Traditional logs are pretty good at telling us what happened. The agent called this tool. It accessed this repository. It sent this request.
That is useful.
But with agents, I increasingly want decision tracing too.
What task was the agent trying to accomplish when it made that request? What input influenced the decision? What tool did it ask for? What actually happened after that tool was called? Did the action still match the original objective?
And I do not want the only answer to those questions to come from asking the agent afterward.
If I wake up tomorrow and discover the loop touched something strange, I want enough external evidence to reconstruct why it happened without depending on the same model to explain itself.
That is where Mike’s outcome engineering idea and the security side fit together really nicely.
The loop should not be the authority on whether it succeeded. And it should not be the authority on whether the things it did along the way were acceptable.
Finally, give yourself a kill switch.
And actually test it.
Because discovering at hour six that your emergency stop requires the agent to cooperate would be a pretty funny design decision right up until it happened to you.
Mike’s team is already doing several of these things:
Mike: Several items on Tox’s list are already baked into how our loops run, because the limits in the outcome file require them. An agent’s account can read our production database and has no way to write to it. Loops work in sandboxes, and their changes reach production only through a review step the loop can’t operate.
Our longest-running loops write everything to a staging area and stop there, and only a person can move the work into our canonical library. Our monitoring reads the actual state of the system rather than the loop’s account of it, and every run has a pause switch the agents can’t disable.
The pre-flight checklist
Before the run
The outcome file exists, with the result, the limits, the tie-breaker, and the check written down.
Token, time, and scope budgets are enforced at the key and the account, where the loop can’t raise them.
Credentials are scoped to this run and expire with it.
The sandbox is confirmed, with no accidental path to production.
The kill switch has been tested mid-run.
During the run
A checker outside the loop grades every finished chunk.
Network egress and sensitive tool use are logged.
Nobody has to watch.
After the run
Read the checker’s verdict before the agent’s report.
Review what got flagged.
Reconcile the bill against the budget.
Send, spend, deploy, and delete still wait for a person.
Close
Loops made my company faster once we stopped letting agents improvise from high-level instructions and judge their own work.
If you’re running loops at your company, send me a note. I’d like to see what your outcome file looks like, and to hear what your loops keep getting wrong.
I think long-running agents are going to become normal because they are genuinely useful. For me, that is exactly why security matters.
We cannot build the future of agent security around somebody sitting in front of a terminal watching every tool call. The whole point of these systems is that eventually we are going to walk away.
So the mature version of this is figuring out how autonomy and containment live together. Give the agent enough room to solve the problem and enough time to do useful work. But do not give one bad decision six hours of unrestricted authority.
If you are already running agents overnight, I am curious about one thing: what is the action you still would not trust them to take while nobody is watching?
Sources
[1] Prithvi Rajasekaran, “Harness design for long-running application development,” Anthropic engineering, March 24, 2026: https://www.anthropic.com/engineering/harness-design-long-running-apps
[2] Anthropic, “Introducing dynamic workflows,” claude.com blog, May 28, 2026: https://claude.com/blog/introducing-dynamic-workflows-in-claude-code. Anthropic, “Introducing Claude Sonnet 4.5” (the 30-hour autonomous run), September 29, 2025: https://www.anthropic.com/news/claude-sonnet-4-5
[3] Derrick Choi, “Run Long-Horizon Tasks with Codex,” OpenAI Developers blog, February 23, 2026: https://developers.openai.com/blog/run-long-horizon-tasks-with-codex
[4] Cursor Cloud Agents: https://cursor.com/docs/background-agent. GitHub Copilot cloud agent: https://docs.github.com/en/copilot/concepts/agents/coding-agent/about-coding-agent.
Google Jules:
https://jules.google/
Devin:
https://docs.devin.ai/get-started/devin-intro (all accessed August 2026)
[5] METR, “Time Horizon 1.1,” January 29, 2026: https://metr.org/blog/2026-1-29-time-horizon-1-1/
[6] Faros AI, “AI Engineering Report 2026,” April 2026 (telemetry from 22,000 developers): https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways
[7] “Confident and Wrong,” arXiv:2603.25764, March 2026: https://arxiv.org/abs/2603.25764
[8] Geoffrey Huntley, “Ralph,” July 14, 2025: https://ghuntley.com/ralph/. YC hackathon field report: https://github.com/repomirrorhq/repomirror/blob/main/repomirror.md
[9] MakeUseOf, “Someone left Claude Code running overnight and it cost $6,000”: https://www.makeuseof.com/someone-left-claude-code-running-overnight-and-it-cost-6000/ (rooted in X threads; the body says “reported”)
[10] Michael Quoc, “Loop engineering is the meme. Outcome engineering is the job,” Product.ai Research, July 23, 2026: https://product.ai/research/outcome-engineering/
Thanks for reading! If you enjoyed this content, give Michael a follow or subscribe to his Substack - and checkout Product.ai!
Michael Quoc runs agent loops every day at Product.ai (formerly Demand.io), the profitable, bootstrapped AI company he founded and leads as CEO, now valued at $100M+. He writes from the perspective of running loops, and managing their risk, in his own company.









thanks a ton to both Michael and Karen for coordinating and writing a great article.
as always, feel free to ama on agentic loops!
Chris, it was a pleasure working with you on this! Happy to help field questions on agentic loops.