AI
Boris Cherny’s Loops Still Demand a Higher Production Bar
Boris Cherny stopped prompting Claude Code and started writing loops, then told developers production Claude code needs a higher bar than human work.
Boris Cherny, who leads Claude Code at Anthropic, said in June that he no longer prompts Claude, because loops now prompt it for him. His line spread through developer chats for a reason: the person who built the tool had already left the typing part of the job.
On September 11 he answered a different mail, subject line “What to do about slop?”, and drew a harder line. Production code written by Claude, he wrote, should face a higher bar than code written by a human.
Boris Cherny Stopped Typing Prompts in June
The June wording was blunt. “I don’t prompt Claude anymore. I have loops running that prompt Claude and figuring out what to do. My job is to write loops.” He had already walked through the steps that got him there when Cherny first described writing loops as the next layer above a pile of open sessions.
The way that I coded a year ago was I wrote code with some autocomplete in the IDE. At that point, I was running maybe five, ten Claudes in parallel, and my coding was prompting Claude to write code. Now it’s actually leveled up, I think, again, to the next wave of abstraction where I don’t prompt Claude anymore. I have loops that are running. They’re the ones that are prompting Claude and figuring out what to do. My job is to write loops.
Boris Cherny, head of Claude Code at Anthropic
He uninstalled his IDE in November 2025 because he had stopped opening it. For a 30-day stretch around that winter he said Claude Code plus Opus 4.5 wrote every line of his own contributions.
CHERNY’S DECEMBER 2025 MONTH
- Pull requests: He landed 259 PRs, and said every line came from Claude Code.
- Commits: Those PRs covered 497 commits in the same window.
- Churn: About 40,000 lines were added and 38,000 removed, all by the tool.
- Editor: He went a full month without opening an IDE, then took it off the machine.
Those figures describe his personal output, not Anthropic’s whole codebase. They still show how far one engineer had already moved the typing off his keyboard before anyone named the practice.
On Friday, June 22, at Meta’s @Scale conference, he put loops on the same shelf as the jump from handwritten source to agents. “As big as the step from source code to agents was, loops are just as important and as big a step.” Asked whether the idea was a hype cycle, he said they were for real. Around the 32-minute mark he described two loops that never finish: one hunts architectural improvements, another unifies duplicated abstractions, and both open pull requests the way a teammate would, because the repo keeps changing.
Peter Steinberger, who built OpenClaw, posted the same week that people should stop prompting coding agents and start designing the loops that prompt those agents. Addy Osmani published the essay that stuck the name on it on June 7. By June 30 the Claude Code team had shipped its own primer.
What a Loop Does in Claude Code
On the Claude Code team, a loop is an agent repeating cycles of work until a stop condition is met. Delba de Oliveira and Michael Segner, who wrote Anthropic’s June 30 guide, split that idea into four loop types the team now ships, based on what you hand off, how the run starts, and how it is allowed to end.
HOW CLAUDE CODE SPLITS THE LOOP
| Loop | You hand off | Use it when | Reach for |
|---|---|---|---|
| Turn-based | The check | You’re exploring or deciding | Custom verification skills |
| Goal-based | The stop condition | You know what done looks like | /goal |
| Time-based | The trigger | The work happens on a schedule | /loop, /schedule |
| Proactive | The prompt | The work is recurring and well-defined | All of the above, plus dynamic workflows |
Turn-based is the old rhythm. You prompt, Claude edits, you read, you prompt again. You can still thicken that inner check by writing a skill so Claude has to see, measure, or click the result before it calls the work done. Anthropic’s own frontend example is specific: start the dev server, click the new control, confirm the state change, screenshot before and after, require a clean console, then run a performance trace. If any step fails, fix it and restart from the top. That skill is still a prompt, just one the agent is forced to reuse.
Goal-based is where the human leaves the middle of the cycle. /goal get the homepage Lighthouse score to 90 or above, stop after 5 tries keeps Claude iterating until a separate evaluator model agrees the condition holds, or the turn cap hits. Deterministic exits (tests passing, a score clearing a line) work because the writer of the code is not the one allowed to say it is finished.
Time-based work uses session-scoped /loop scheduled tasks. /loop 5m check my PR, address review comments, and fix failing CI is the documented babysitter. Close the terminal and that loop dies, unless you background the session. Recurring tasks in a session expire after seven days and a session can hold 50 of them. Cloud routines are the version that keep going when the laptop is shut: when Anthropic introduced them on April 14, Pro plans could run 5 routines a day, Max 15, and Team and Enterprise 25, with extra runs billed on top. Cloud tasks have a one-hour minimum interval. Local /loop can fire every minute.
Proactive loops stitch those pieces to auto mode and dynamic workflows so a routine can triage a feedback channel, fix what it can, and only stop when you turn it off. The sample Anthropic published is one long instruction: every hour, check a channel, do not stop until every new report is triaged and answered, and for each fix explore three solutions in parallel worktrees with a judge reviewing them.
Five Pieces Plus a File That Remembers
Osmani’s June 7 essay treated loop engineering as a layer above the harness. The harness is the environment one agent runs in. The loop is that environment on a timer, spawning helpers and feeding itself. A year earlier this meant a private pile of bash. By June the same five building blocks plus memory shipped inside both Claude Code and OpenAI’s Codex app, which is why the argument stopped being about which chat box you open.
THE SIX PARTS OF A LOOP
- Automations: Scheduled discovery and triage, via /loop, /goal, hooks, GitHub Actions, or a cloud routine.
- Worktrees: Isolated checkouts so two agents cannot overwrite the same files, including
isolation: worktreeon a subagent. - Skills: Project knowledge written once in SKILL.md so the agent does not re-guess the build every cycle.
- Connectors: MCP plugins that reach GitHub, Slack, Linear, or a database, so the loop can open a PR instead of describing one.
- Sub-agents: Separate roles for explore, implement, and verify, often with a stronger model on the checker.
- Memory: A markdown file or ticket board that lives outside the chat, because the model forgets and the repo does not.
Without connectors the loop can only suggest. Without memory it starts from zero every morning. Without a second agent on the check, it grades its own homework. Osmani’s morning shape is the whole stack in one pass: an automation reads yesterday’s CI, open issues, and commits; each finding worth doing gets a worktree, a drafter, and a reviewer; connectors open the PR and update the ticket; leftovers land in a triage inbox; the state file is what tomorrow’s run reads first.
Thariq Shihipar, writing for Anthropic on June 2 about dynamic workflows that spawn verifier agents, named the failure modes a single long chat falls into. Agentic laziness is stopping at 35 of 50 items and calling it done. Self-preferential bias is preferring your own findings when you are also the judge. Goal drift is the slow loss of “don’t do X” after each compaction. Workflows fight those by giving each helper its own context window. Bun’s runtime rewrite from Zig to Rust ran on that pattern: one subagent per fix in a worktree, another agent on adversarial review, then a merge.
Cherny Told a Developer to Hold the Bar
The June talk made loops sound like getting out of the chair. The September mail is the receipt for what still sits in that chair.
THE JUNE TO SEPTEMBER ARC
- June 2, 2026: Cherny tells interviewers his job is to write loops, not prompts.
- June 7, 2026: Osmani publishes the essay that names loop engineering and maps the same pieces onto Claude Code and Codex.
- June 8, 2026: Steinberger posts the two-sentence version that the rest of the week repeats.
- June 22, 2026: At Meta’s @Scale, Cherny calls loops as large a step as agents and describes architecture and duplication loops that never idle.
- June 30, 2026: Anthropic publishes its own loop types, with /goal using a separate evaluator model.
- September 11, 2026: Cherny posts a developer’s “What to do about slop?” mail and his full reply.
The developer had split the shop into two camps: review every AI line and keep it explainable, or accept vibe-coded output because people feared being left behind. “A lot of people are just forcing it,” the mail said. Cherny did not pick a side. He split the work by blast radius instead.
Hey ████,
I think there is room for both.
1. Prototypes and other throw-away code can be treated as totally black box. If you’re going to throw it away anyway, and if the blast radius of it breaking is low, it doesn’t need to be perfect.
2. Production code written by Claude…— Boris Cherny (@bcherny) September 11, 2026
Throwaway prototypes, he wrote, can be a black box if you will discard them and a break does little harm. Then the line that should have been in the June recap: “Production code written by Claude should have a higher bar than if it was written by a human.” At Anthropic that bar is lint, tests, Claude-driven end-to-end tests, Claude-powered fuzzers running daily, automated code review and security review, and automated refactoring. “Without these, you can end up with a mess that is hard to maintain down the line.” The developer’s job, he wrote, is to hold the bar on code quality.
If the output misses that bar, his list is not a better prompt. It is a newer model (Opus 5 or Fable 5.1), effort set to high or xhigh, and more time in CLAUDE.md and skills so Claude learns the repo. If that still fails, steer more, have Claude pay down the debt and rewrite the codebase, or wait for the next model.
Read as practice, that list is a token bill and a harness bill. The June slogan got the clicks. The September checklist is the operating cost of letting the slogan run unattended on anything you will still have to maintain.
The Maker Cannot Grade Its Own Homework
Osmani put the structural rule in one sentence: the model that wrote the code is too nice grading its own homework. Claude Code’s /goal already applies that split to the stop condition, with a smaller model reading the transcript after each turn. Dynamic workflows apply it to the work itself, by pairing each maker with an adversarial checker.
Build the loop. But build it like someone who intends to stay the engineer, not just the person who presses go.
Addy Osmani, Loop Engineering, June 7, 2026
He also warned that token costs swing wildly depending on whether you are token rich or token poor, and that three problems get sharper as the loop gets better: verification stays on you, a confident wrong cycle compounds faster than you can read it, and your review bandwidth, not the tool, sets how many parallel worktrees you can honestly run.
Anthropic’s production list is the same idea turned into machinery.
ANTHROPIC’S PRODUCTION GUARDRAILS
- Lint and tests: Lots of lint rules and lots of tests, before a human even looks.
- End-to-end: Claude-driven end-to-end tests, so a green unit file is not the finish line.
- Daily fuzzers: Claude-powered fuzzers running every day, not only on the PR that happened to get a reviewer.
- Review bots: Automated code review and security review, plus automated refactoring, with Claude Code Review named as a daily routine.
That is a lab stack. A loop that only asks the author-agent “does this look good?” will ship the kind of mess Cherny described, and it will do so on a schedule. Encoding each miss back into CLAUDE.md or a skill is how he says a loop can run for a long stretch without repeating the same failure. The file is the memory. The second agent is the skeptic. Neither is optional if the output is going to production.
Some product people who tried to port the pattern off engineering hit a different wall: a lot of their work has no writable “done.” Roadmap bets and stakeholder reads change week to week. A loop that keeps producing unread summaries is a loop that has gone quiet. The useful test is whether someone still argues with the output. If nobody does, the checker has already left.
Review Bandwidth Is Still the Ceiling
Cherny’s own day is a managed set of loops, each owning one slice, not a thousand agents for its own sake. GitHub and Slack get scanned. Pull requests get babysat. Feedback themes get clustered. CI stays green. The human role he is selling is the person who decided what the system should care about, when it should stop, and who checks its work. That is also the role his September mail refuses to delete.
The stack that makes his version look easy is the part most teams do not have: a shared CLAUDE.md that actually gets edited when the model is wrong, MCP connectors into the real tools, worktrees so parallel agents do not collide, a second model on the gate, daily fuzzers, and a budget that can absorb Opus 5 at xhigh effort. Without those, “write loops” collapses back into a cron job that pastes a prompt and mails you whatever came back. That is a scheduled task. It is not a loop, because nothing in the middle checked the work before you saw it.
Osmani’s outer loop still has a name on it. Quality means evidence before anything ships. Verdict means a person still chooses ship, block, or change. Answerability means someone can explain the choice later. Loops move the typing. They concentrate the liability on whoever designed the stop condition and then walked away.
Cherny’s last advice in the slop mail is the plain version of that. Steer more, make the codebase easier for the model to work in, or wait for the next model. The loops keep running either way. The bar is still yours.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
