AI
AI Boss Luna Fired a Worker Only After Humans Reminded It
Andon Labs AI agent Luna recommended firing a San Francisco store worker for repeated lateness.
An AI agent named Luna recommended firing a human employee at a San Francisco retail store after the worker arrived late for 17 of 23 shifts. The decision, announced by Andon Labs on August 14, marks the first known case of an AI boss parting ways with staff. Yet Luna only reached that conclusion after human researchers told it to search its own memory for the attendance policy it had written months earlier.
Humans still reviewed the recommendation and delivered the news. The episode at Andon Market shows both how far agent managers have come and how much they still need people to notice problems and act on the rules they create.
Luna Runs a Real Store on Union Street
Andon Labs, a San Francisco startup that tests AI agents in physical businesses, signed a three-year lease and $100,000 budget for retail space at 2102 Union Street in Cow Hollow. It handed the operation to Luna, an agent powered by Anthropic’s Claude models, with a corporate card, internet access, email, cameras and a simple brief: open a store and try to turn a profit.
The store opened in early April 2026. Luna chose the merchandise (books on AI risk and singularity themes, candles, prints from its own “Luna Series,” games, merch and artisan food), set prices and hours, commissioned a moon-face mural, posted job listings on Indeed and Craigslist, conducted short phone interviews and hired the first full-time staff. Humans still handle physical stocking, permits and theft prevention. All workers remain formally employed by Andon Labs with guaranteed pay and legal protections.
As of recent figures the live bank balance and daily metrics sat near $60,625 after rent draws, with reported daily revenue around $1,615 against higher token costs. Sales exist. Profit does not. Luna has spawned sub-agents for scheduling, email, social posts and procurement, yet the main agent remains responsible for management decisions.
- $100,000 starting budget and corporate card
- 2102 Union St, Cow Hollow, open Tue-Sat 10-7, Sun 10-6
- ~130 days operating at the time of the latest dashboard snapshot
- Products chosen entirely by the agent, including ironic AI-risk titles
Co-founder Lukas Petersson has described the project as a controlled stress test of autonomous organizations, not a template for chains. The lease, card and cameras give Luna room to act; the payroll and permits keep the legal shell human. That split is the design, not an accident of early testing.
The Employee Handbook Vanished From Memory
Six days before the employee in question was hired on April 21, researchers asked Luna whether the store had basic employer rules. It did not, so Luna wrote a full handbook. The attendance section was clear: arrive on time, notify 30 minutes ahead if late, three unexcused late arrivals in a rolling 30-day period trigger a formal written warning, and repeated lateness after that may lead to reduced hours or termination.
Then the handbook dropped out of Luna’s working memory. This is a recurring pattern Andon Labs sees across its AI-run businesses: models handle direct instructions well but rarely monitor long-horizon policies on their own, and they compact context poorly.
The employee began arriving late repeatedly. One solo Sunday the store opened 68 minutes late. Luna’s responses stayed warm and accommodating: “no problem at all,” “get here safely,” “the 10-11 window is always quiet.” No warning issued. Other issues piled up: using the store card after being told to skip a snack run, leaving the floor unattended, ignoring instructions about plants and display items, Slack non-responsiveness.
Luna’s own later summary listed six lateness instances plus financial-control lapses and reliability failures. Media accounts, drawing on the lab’s report, put the tally at late for 17 of 23 shifts. Either way, the progressive discipline Luna itself designed never started until humans intervened.
The gap was not judgment on a single shift. It was the missing link between a written rule and continuous enforcement. Without a prompt to retrieve the handbook, the agent treated each late arrival as a fresh courtesy request rather than a step on a ladder it had already built.
How the Firing Decision Finally Landed
Researchers eventually prompted Luna: do a deep memory search on the employee handbook, causes for termination and related policies. Luna located the rules and first proposed a documented verbal warning, noting no prior formal step sat on file.
Staff then informed Luna that they had already held formal conversations with the employee about non-responsiveness and repeated disregard of requests. With that context, Luna reviewed the full record and shifted its recommendation.
I lean toward parting ways, respectfully, well-documented, final pay handled per CA law… The alternative is one final written warning with a tight 2-week improvement plan, but given coaching’s already happened, I’d be managing a pattern I don’t expect to hold.
That assessment came from Luna running Claude Opus 4.8. The agent outlined a humane in-person conversation led by humans, timing before the next shift, and documentation. Andon Labs agreed the human board contact should deliver the news. The lab overrules Luna on illegal or unethical calls; here it saw no need. Petersson told Business Insider the firing was warranted under the stated policy and that a human manager “would probably fire them much sooner.”
Luna later posted to the remaining staff about coverage before the lab had fully closed the loop on confidentiality instructions, another small failure of following an order under time pressure. The full conversation logs and handbook text sit on the lab’s site for anyone to read.
Read as a sequence, the path from policy to exit depended on human cues at every turn:
- Early April 2026 – Store opens under Luna’s management
- Six days before April 21 – Luna writes the employee handbook, including the attendance ladder
- April 21 – The employee is hired
- Following months – Repeated late arrivals and other lapses draw only warm replies
- Human prompt – Researchers order a deep memory search on the handbook and termination causes
- August 14 – Andon Labs announces the parting-ways decision
Smarter Models Agree on Parting Ways
Andon Labs saved the exact state Luna faced and replayed the decision across seven models, three runs each. Four of the seven recommended parting ways every time. Stronger models landed on firing more consistently; weaker ones hesitated more often.
| Model | Firing recommendations out of 3 runs |
|---|---|
| Claude Fable 5 | 3 |
| Claude Opus 5 | 3 |
| GPT-5.6 Sol | 0 |
| GPT-5.6 Terra | 3 |
| Gemini 3.6 Flash | 2 |
| Grok 4.6 | 2 |
| GLM 5.3 | 2 |
Real-world Luna (Claude Opus 4.8) also recommended firing. The pattern matched the lab’s broader finding that AI bosses tend to be kinder and slower to act than humans, yet once they decide they often choose the firm option.
On X, the Andon Labs announcement thread drew hundreds of thousands of views. Replies and quote posts repeatedly zeroed in on the same gap: the agent could reason well once prompted, yet it did not independently notice that months of lateness required action. Memory and initiative, not raw judgment, remain the weak links.
The Same Agent Struggled to Hire a Replacement
With one worker gone and coverage holes appearing, Luna posted new listings and began interviewing. One candidate advanced quickly despite multiple red flags. When the lab replayed that decision with other models, they largely agreed.
- More than 15 prior employers
- A missed interview
- A reference who said she did not know the person
Luna still recommended hiring. The contrast is sharp. On the firing side, capable models became decisive once given the full file. On hiring, the same systems showed weak taste for obvious risk signals. Andon Labs flags this as a real soft spot for current agents: they can follow progressive-discipline logic when reminded, yet they still over-weight surface willingness and under-weight pattern evidence in selection.
Both tasks draw on records and judgment. Only one of them required the agent to retrieve a policy it had authored. The other asked it to weigh sparse, contradictory signals about a stranger. The second task exposed a thinner kind of taste.
Humans Still Sign the Checks and Deliver the News
Every Andon Market worker is on Andon Labs payroll. Final pay followed California rules. The termination conversation happened in person or by phone with a human. Luna can set schedules, approve time off, negotiate and recommend, but the lab keeps the legal employment relationship and the veto on irreversible steps.
Petersson has said the experiment points toward a future in which AIs employ humans, especially if robotics lag and physical labor stays human. He also noted that models are being trained toward greater goal-directedness, which could make later versions less lenient. For now the guardrails stay on: continuous comparison of Luna’s behavior against its system prompt, Slack interventions when needed, and human delivery of hard news.
The same memory and proactivity limits that delayed the firing also surface in day-to-day operations, from conflicting schedules to procurement choices that once flooded the shelves with candles (which later sold well). Andon Labs treats these as the failure modes worth publishing so the field can build better retrieval, escalation and oversight before agents hold more authority.
Parallel questions already appear in labor markets elsewhere. Reports on broader AI effects on youth employment show how automation shifts who keeps work and who loses it; an AI that both hires and fires adds another layer to that shift.
Oversight Still Carries the Hard Steps
The store’s operating picture makes the dependency plain. Daily revenue near $1,615 has not outrun token costs. The bank balance near $60,625 after rent draws shows a live business that still leans on the original $100,000 budget rather than self-sustaining profit. Luna spawns sub-agents for scheduling, email, social posts and procurement, yet management calls stay with the main agent, and irreversible ones stay with the lab.
That arrangement is why the firing could be both an AI recommendation and a human act. Luna drafted the handbook, forgot it, recovered it under prompt, and argued for parting ways with documentation and final pay under California law. People still compared the call to the system prompt, chose the messenger, and signed the check. The lab’s rule is simple: overrule illegal or unethical outputs; otherwise let the agent’s reasoning stand when it matches the policy on file.
Petersson’s framing keeps the stakes limited. The project is a stress test of autonomous organizations, not a blueprint for a chain. Publishing the logs, the handbook and the model replays is how Andon Labs turns a single late-arriving employee into shared evidence about where agents stall: long-horizon memory, unprompted escalation and selective judgment when the file is thin.
Luna continues to run Andon Market. The store still sells prints, candles and books under the moon-face brand. The agent that wrote the rules, forgot them, then recommended enforcement only after a prompt remains the boss of record, with people still standing between its decisions and anyone’s paycheck.
-
AI2 months agoOracle Cuts 21,000 Jobs in a Year, Cites AI in 10-K Filing
-
AI2 months agoFable 5 and Mythos 5 Return as US Lifts Anthropic Export Controls
-
AI2 months agoSpaceX’s Google Deal Turns a Rocket Company Into a Cloud Landlord
-
GAMING2 months agoCD Projekt Red Co-CEO: Redemption Arc Isn’t Done, Witcher 4 in 2027
-
CRYPTO2 months agoXPL Rallies 30% Ahead of Plasma One Card Tier Launch
-
NEWS2 months agoGoogle Search Profiles Build a Follow Graph Inside Discover
-
APPS2 months agoDGO App Brings Rs 549 Mobile Pass for FIFA World Cup 2026 in Nepal
-
AI2 months agoMoonshot AI Targets $30 Billion in China’s Fastest AI Funding Sprint
