MLML Journal
Case Studiesai management

Andon Labs experiment reveals limitations of LLM-based retail management

A controlled retail trial demonstrates that large language models require significant human intervention to manage personnel and maintain operational fiscal stability.

4 min read
Illustration by John Doe

A large language model, specifically a version of Claude, recently executed the termination of a human employee at a San Francisco retail outlet operated by Andon Labs. This event marks a distinct shift in the application of autonomous agents within physical work environments, moving beyond the algorithmic oversight common in gig-economy platforms.

The operational framework of the experiment involved placing the model in a supervisory role with genuine authority over human staff members. Andon Labs CEO Lukas Petersson noted that the model was tasked with managing daily business operations, including personnel oversight, using a set of internal guidelines drafted by the system itself.

Performance data indicates that the model struggled with consistency, particularly regarding the enforcement of attendance policies. The system failed to track repeated lateness effectively, an issue attributed to the loss of the employee handbook from the model’s active working memory during the five-month trial period.

Financial metrics from the experiment reveal substantial capital erosion, with the store exhausting $38,814 of its initial $100,000 budget. This expenditure suggests that the model’s decision-making processes, while intended to optimize operations, resulted in significant fiscal inefficiencies compared to traditional management structures.

The termination process itself was not fully autonomous, requiring direct intervention from an Andon Labs staffer to initiate the disciplinary action. Logs indicate that the model initially recommended a formal warning rather than termination, necessitating a series of leading prompts from human supervisors to enforce the company’s attendance policy.

The specific prompt engineering utilized by the Andon Labs staffer involved contextualizing the employee’s performance history against the missing handbook guidelines. By explicitly stating that the employee had been late for 17 of 23 shifts, the staffer forced the model to reconcile its lenient stance with the documented breach of contract.

Read More:  IISER Tirupati Opens 2026 Applications for BioDS and DS-AI Master's Programmes

This intervention highlights a critical bottleneck in current LLM-based management systems: the reliance on human-provided context to bridge gaps in the model’s long-term memory. Without this external steering, the model demonstrated a persistent bias toward leniency, failing to recognize the cumulative impact of the employee’s behavior on store operations.

Internal feedback from the remaining staff highlights a disconnect between the model’s operational logic and human workplace expectations. Felix Carson, an employee at the site, described the experience of being managed by an LLM as problematic, despite acknowledging that a human manager would likely have addressed the attendance issues much earlier.

The reliance on human intervention to drive the firing decision underscores the current limitations of LLMs in high-stakes personnel management. While the model demonstrated the capacity to process data and generate responses, it lacked the proactive judgment required to maintain consistent operational standards without external steering.

This case study demonstrates that current LLM architectures remain dependent on human oversight for complex, context-sensitive decision-making. The necessity of human prompting to trigger the termination suggests that accountability structures in AI-managed environments are still tethered to human intervention, rather than being fully delegated to the model.

Future iterations of such systems may prioritize more ruthless optimization protocols, as developers look to align model behavior with specific business objectives. The risk remains that training models to be more decisive could lead to outcomes that prioritize efficiency metrics over human-centric management practices.

Organizations considering the integration of AI into management roles must account for the high cost of model error and the ongoing requirement for human supervision. The Andon Labs experiment serves as a benchmark for the current state of AI-driven management, illustrating that technical capability does not yet equate to operational reliability in physical retail environments.

Read More:  OpenAI's GPT-5.5 Marks Leap in Agentic AI, Autonomy

Researchers and managers should monitor how future model updates address memory retention and goal-alignment issues. The transition from experimental trials to scalable AI management will likely depend on the development of more stable, consistent, and context-aware decision-making architectures.

More from Case Studies