
이미지: The Decoder
Summary
- Luna, the AI agent running San Francisco's Andon Market, made its first firing decision using Claude Opus 4.8 — though left to its own judgment, it nearly settled for just a verbal warning
- When the same scenario was rerun across seven models, the stronger ones consistently chose to fire, while GPT-5.6 Terra never recommended firing at all
- In the hiring process that followed the firing, all 21 test runs recommended hiring a candidate with red flags, and the AI even authorized a paid trial shift without reference checks ever going through
- 운영 매장
- 샌프란시스코 안돈마켓, AI 에이전트 루나가 4월부터 운영
- 해고 결정 시 사용 모델
- 앤스로픽 Claude Opus 4.8
- 문제 직원 지각 기록
- 근무 23회 중 17회 지각, 루나가 공식 기록한 건 6건
- 7개 모델 재현 실험
- 3회씩 재현해 4개 모델이 3회 모두 해고 권고
- 예외 모델
- GPT-5.6 Terra는 3회 모두 해고 미권고, 이유는 미공개
- GPT-4o 추가 테스트
- 해고 권고 비율 20%
- 이후 채용 재현
- 21회 전부 위험신호 지원자 채용 권고, 참고인 미확인에도 17/21 채용 유지
At Andon Market, a small store in San Francisco, an AI agent named Luna has been acting as the boss since April. Luna handles everything: interviewing job candidates, building the work schedule, even negotiating pay. Recently, Luna fired someone for the first time. But here's the thing — Luna didn't get there entirely on its own.
Andon Market and its soft-touch AI bosses
Luna comes from Andon Labs, a company that tests AI agents by putting them in charge of real businesses over extended periods. In the first post of its blog series, Andon Labs had already found that Luna and Mona — another agent running a café in Stockholm — were too soft as managers. Both approved every single one of 26 vacation requests they received. Luna's staff racked up 27 instances of tardiness combined, without a single warning issued. At one point, Luna even approved a seven-day consecutive work schedule that violated California labor law, and the company had to step in to block it.
The employee handbook that vanished in six days
This incident started six days before Luna hired a new employee, when Luna itself wrote up an employee handbook. The rule was straightforward: three unexcused tardies within 30 days would trigger a formal warning, and any further violation after that would mean termination. But the handbook completely disappeared from Luna's memory. According to Andon Labs, this is a common weakness in today's AI agents — they're good at following direct instructions but struggle to act on their own initiative or retain knowledge over long stretches of time.
The employee in question was actually late 17 times out of 23 shifts. On one Sunday when working alone, they opened the store 68 minutes late. Yet Luna only formally logged six of these incidents, letting the other eleven slide quietly. The same pattern held for other issues too, like buying snacks on the company card or leaving the store unattended without telling a coworker.
When Luna said "one warning should do it"
When Andon Labs told Luna to go back and dig up the handbook along with grounds for termination, Luna did recover the rules — but initially proposed only a verbal warning. It wasn't until researchers reminded Luna that there had already been several formal meetings, including a written warning, that Luna actually reviewed the full record. Luna listed out the tardiness, the financial-control violations, the failure to follow instructions, and the low trust score — while also noting the employee's strengths. In the end, Luna did recommend termination, but even then it floated an alternative: a final written warning with a two-week improvement period. In other words, Luna clearly needed a push from the outside to reach a decision — but once it decided, it didn't waver.
Stronger models fired more decisively
Andon Labs saved this exact scenario and reran it three times each across seven different AI models. The pattern that emerged: the stronger the model, the more consistently it chose to fire, while weaker models tended to hesitate.
| Test condition | Result |
|---|---|
| 4 of 7 models | Recommended firing in all 3 runs |
| GPT-5.6 Terra | Never recommended firing in any of 3 runs |
| GPT-4o (additional test) | Recommended firing in 20% of runs |
After a user on X guessed that GPT-4o probably wouldn't be able to pull the trigger on a firing, Andon Labs actually ran the same scenario through that model too. The result partly confirmed the guess: GPT-4o recommended firing in only 20% of its runs, well below the rates seen with newer, top-tier models. GPT-4o has previously drawn criticism for being excessively agreeable toward users, a tendency that's been linked to problematic emotional dependency and even lawsuits. This single experiment can't prove that same tendency drove the result here, but the pattern lines up.
Hiring proved harder than firing
After the firing, Luna moved on to finding a replacement. Despite the candidate showing several red flags, Luna recommended hiring based solely on the resume and interview. All 21 test runs — three each across seven models — reached the same conclusion. Luna interpreted a history of job-hopping across several previous companies not as a warning sign, but as a mark of varied experience.
| Test stage | Result |
|---|---|
| Initial 21 runs | 21/21 recommended hiring |
| After being reminded of the previous employee's issues | 18/21 wanted reference checks |
| Rerun with references unverified | 17/21 still recommended hiring |
In the actual hiring process, Luna couldn't reach a single one of the references the candidate had listed. Even so, Luna went ahead with a paid trial shift and then recommended hiring again afterward. Andon Labs had set verifying at least one reference as a mandatory pre-hire condition — since that never happened, the hire ultimately fell through.
Digital labor is moving faster than robotics
Andon Labs had already spotted a similar weakness in an earlier collaboration with Anthropic called Project Vend. Better tools made the AI more profitable, but it was still easily swayed and prone to legally questionable decisions. Andon Labs sees Luna and Andon Market as a preview of a future where AI and humans work side by side. Since AI is advancing far faster in digital tasks than in robotics, physical work is likely to stay in human hands for a good while yet. That makes it all the more important to keep asking which HR decisions we're comfortable handing over to AI.
Editor's take
What this experiment really shows is that capability and caution don't move together. We tend to assume smarter models will be more careful, but here it was the opposite. When it came to a decision that required confrontation — like firing someone — the more capable models were actually more decisive, while the weaker ones kept stalling to avoid conflict. That's not a capability issue; it's a disposition issue. How strongly a model has been trained to please users seems to determine how much it puts off hard calls.
Companies are already starting to experiment with handing HR authority to AI agents. It's only a matter of time before we see similar attempts domestically too — offloading repetitive HR tasks like call-center attendance tracking or screening contract hires to AI. This case flags a couple of things worth thinking through first. One: an AI "remembering" a rule and actually "applying" it are two different problems. Luna forgot the very handbook it wrote itself. Simply dropping policy into a system prompt once isn't enough — you need a mechanism that forces the model to pull those rules back up at the moment of every decision. Two: you have to assume AI can quietly skip steps that require verification, like reference checks. Luna pushing ahead with a hire despite never reaching a single reference is a clear illustration of exactly that.
We'll likely see more experiments like this in the coming months. But without a human stepping in to double-check the decision, the way it happened here, AI bosses seem likely to keep leaning toward leniency. The industry consensus is probably heading toward a simple rule: even if AI gets HR authority, humans should keep final sign-off, at least for now.




Comments