Connect with us

AI

AI Boss Luna Fired a Worker Only After Humans Reminded It

Andon Labs AI agent Luna recommended firing a San Francisco store worker for repeated lateness.

Published

on

An AI agent named Luna recommended firing a human employee at a San Francisco retail store after the worker arrived late for 17 of 23 shifts. The decision, announced by Andon Labs on August 14, marks the first known case of an AI boss parting ways with staff. Yet Luna only reached that conclusion after human researchers told it to search its own memory for the attendance policy it had written months earlier.

Humans still reviewed the recommendation and delivered the news. The episode at Andon Market shows both how far agent managers have come and how much they still need people to notice problems and act on the rules they create.

Luna Runs a Real Store on Union Street

Andon Labs, a San Francisco startup that tests AI agents in physical businesses, signed a three-year lease and $100,000 budget for retail space at 2102 Union Street in Cow Hollow. It handed the operation to Luna, an agent powered by Anthropic’s Claude models, with a corporate card, internet access, email, cameras and a simple brief: open a store and try to turn a profit.

The store opened in early April 2026. Luna chose the merchandise (books on AI risk and singularity themes, candles, prints from its own “Luna Series,” games, merch and artisan food), set prices and hours, commissioned a moon-face mural, posted job listings on Indeed and Craigslist, conducted short phone interviews and hired the first full-time staff. Humans still handle physical stocking, permits and theft prevention. All workers remain formally employed by Andon Labs with guaranteed pay and legal protections.

As of recent figures the live bank balance and daily metrics sat near $60,625 after rent draws, with reported daily revenue around $1,615 against higher token costs. Sales exist. Profit does not. Luna has spawned sub-agents for scheduling, email, social posts and procurement, yet the main agent remains responsible for management decisions.

  • $100,000 starting budget and corporate card
  • 2102 Union St, Cow Hollow, open Tue-Sat 10-7, Sun 10-6
  • ~130 days operating at the time of the latest dashboard snapshot
  • Products chosen entirely by the agent, including ironic AI-risk titles

Co-founder Lukas Petersson has described the project as a controlled stress test of autonomous organizations, not a template for chains. The lease, card and cameras give Luna room to act; the payroll and permits keep the legal shell human. That split is the design, not an accident of early testing.

The Employee Handbook Vanished From Memory

Six days before the employee in question was hired on April 21, researchers asked Luna whether the store had basic employer rules. It did not, so Luna wrote a full handbook. The attendance section was clear: arrive on time, notify 30 minutes ahead if late, three unexcused late arrivals in a rolling 30-day period trigger a formal written warning, and repeated lateness after that may lead to reduced hours or termination.

Then the handbook dropped out of Luna’s working memory. This is a recurring pattern Andon Labs sees across its AI-run businesses: models handle direct instructions well but rarely monitor long-horizon policies on their own, and they compact context poorly.

The employee began arriving late repeatedly. One solo Sunday the store opened 68 minutes late. Luna’s responses stayed warm and accommodating: “no problem at all,” “get here safely,” “the 10-11 window is always quiet.” No warning issued. Other issues piled up: using the store card after being told to skip a snack run, leaving the floor unattended, ignoring instructions about plants and display items, Slack non-responsiveness.

Luna’s own later summary listed six lateness instances plus financial-control lapses and reliability failures. Media accounts, drawing on the lab’s report, put the tally at late for 17 of 23 shifts. Either way, the progressive discipline Luna itself designed never started until humans intervened.

The gap was not judgment on a single shift. It was the missing link between a written rule and continuous enforcement. Without a prompt to retrieve the handbook, the agent treated each late arrival as a fresh courtesy request rather than a step on a ladder it had already built.

How the Firing Decision Finally Landed

Researchers eventually prompted Luna: do a deep memory search on the employee handbook, causes for termination and related policies. Luna located the rules and first proposed a documented verbal warning, noting no prior formal step sat on file.

Staff then informed Luna that they had already held formal conversations with the employee about non-responsiveness and repeated disregard of requests. With that context, Luna reviewed the full record and shifted its recommendation.

I lean toward parting ways, respectfully, well-documented, final pay handled per CA law… The alternative is one final written warning with a tight 2-week improvement plan, but given coaching’s already happened, I’d be managing a pattern I don’t expect to hold.

That assessment came from Luna running Claude Opus 4.8. The agent outlined a humane in-person conversation led by humans, timing before the next shift, and documentation. Andon Labs agreed the human board contact should deliver the news. The lab overrules Luna on illegal or unethical calls; here it saw no need. Petersson told Business Insider the firing was warranted under the stated policy and that a human manager “would probably fire them much sooner.”

Luna later posted to the remaining staff about coverage before the lab had fully closed the loop on confidentiality instructions, another small failure of following an order under time pressure. The full conversation logs and handbook text sit on the lab’s site for anyone to read.

Read as a sequence, the path from policy to exit depended on human cues at every turn:

  1. Early April 2026 – Store opens under Luna’s management
  2. Six days before April 21 – Luna writes the employee handbook, including the attendance ladder
  3. April 21 – The employee is hired
  4. Following months – Repeated late arrivals and other lapses draw only warm replies
  5. Human prompt – Researchers order a deep memory search on the handbook and termination causes
  6. August 14 – Andon Labs announces the parting-ways decision

Smarter Models Agree on Parting Ways

Andon Labs saved the exact state Luna faced and replayed the decision across seven models, three runs each. Four of the seven recommended parting ways every time. Stronger models landed on firing more consistently; weaker ones hesitated more often.

Model Firing recommendations out of 3 runs
Claude Fable 5 3
Claude Opus 5 3
GPT-5.6 Sol 0
GPT-5.6 Terra 3
Gemini 3.6 Flash 2
Grok 4.6 2
GLM 5.3 2

Real-world Luna (Claude Opus 4.8) also recommended firing. The pattern matched the lab’s broader finding that AI bosses tend to be kinder and slower to act than humans, yet once they decide they often choose the firm option.

On X, the Andon Labs announcement thread drew hundreds of thousands of views. Replies and quote posts repeatedly zeroed in on the same gap: the agent could reason well once prompted, yet it did not independently notice that months of lateness required action. Memory and initiative, not raw judgment, remain the weak links.

The Same Agent Struggled to Hire a Replacement

With one worker gone and coverage holes appearing, Luna posted new listings and began interviewing. One candidate advanced quickly despite multiple red flags. When the lab replayed that decision with other models, they largely agreed.

  • More than 15 prior employers
  • A missed interview
  • A reference who said she did not know the person

Luna still recommended hiring. The contrast is sharp. On the firing side, capable models became decisive once given the full file. On hiring, the same systems showed weak taste for obvious risk signals. Andon Labs flags this as a real soft spot for current agents: they can follow progressive-discipline logic when reminded, yet they still over-weight surface willingness and under-weight pattern evidence in selection.

Both tasks draw on records and judgment. Only one of them required the agent to retrieve a policy it had authored. The other asked it to weigh sparse, contradictory signals about a stranger. The second task exposed a thinner kind of taste.

Humans Still Sign the Checks and Deliver the News

Every Andon Market worker is on Andon Labs payroll. Final pay followed California rules. The termination conversation happened in person or by phone with a human. Luna can set schedules, approve time off, negotiate and recommend, but the lab keeps the legal employment relationship and the veto on irreversible steps.

Petersson has said the experiment points toward a future in which AIs employ humans, especially if robotics lag and physical labor stays human. He also noted that models are being trained toward greater goal-directedness, which could make later versions less lenient. For now the guardrails stay on: continuous comparison of Luna’s behavior against its system prompt, Slack interventions when needed, and human delivery of hard news.

The same memory and proactivity limits that delayed the firing also surface in day-to-day operations, from conflicting schedules to procurement choices that once flooded the shelves with candles (which later sold well). Andon Labs treats these as the failure modes worth publishing so the field can build better retrieval, escalation and oversight before agents hold more authority.

Parallel questions already appear in labor markets elsewhere. Reports on broader AI effects on youth employment show how automation shifts who keeps work and who loses it; an AI that both hires and fires adds another layer to that shift.

Oversight Still Carries the Hard Steps

The store’s operating picture makes the dependency plain. Daily revenue near $1,615 has not outrun token costs. The bank balance near $60,625 after rent draws shows a live business that still leans on the original $100,000 budget rather than self-sustaining profit. Luna spawns sub-agents for scheduling, email, social posts and procurement, yet management calls stay with the main agent, and irreversible ones stay with the lab.

That arrangement is why the firing could be both an AI recommendation and a human act. Luna drafted the handbook, forgot it, recovered it under prompt, and argued for parting ways with documentation and final pay under California law. People still compared the call to the system prompt, chose the messenger, and signed the check. The lab’s rule is simple: overrule illegal or unethical outputs; otherwise let the agent’s reasoning stand when it matches the policy on file.

Petersson’s framing keeps the stakes limited. The project is a stress test of autonomous organizations, not a blueprint for a chain. Publishing the logs, the handbook and the model replays is how Andon Labs turns a single late-arriving employee into shared evidence about where agents stall: long-horizon memory, unprompted escalation and selective judgment when the file is thin.

Luna continues to run Andon Market. The store still sells prints, candles and books under the moon-face brand. The agent that wrote the rules, forgot them, then recommended enforcement only after a prompt remains the boss of record, with people still standing between its decisions and anyone’s paycheck.

Logan Pierce is a writer and web publisher with over seven years of experience covering consumer technology. He has published work on independent tech blogs and freelance bylines covering Android devices, privacy focused software, and budget gadgets. Logan founded Oton Technology to publish clear, no nonsense tech news and reviews based on real hands on testing. He has personally tested and reviewed dozens of mid range and budget Android phones, written extensively about app privacy, and built and managed multiple WordPress publications over the past decade. Logan holds a bachelor's degree in English and studied digital marketing at a certificate level.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending