Anthropic Project Vend Phase Two: AI Shopkeeper Gains Profitability but Still Needs Human Guardrails

TL;DR

Anthropic upgraded its AI shopkeeper, Claudius, from Claude Sonnet 3.7 to Sonnet 4.0/4.5, added CRM, inventory, web‑search, and other tools, and introduced a CEO agent (Seymour Cash) and a merch‑making agent (Clothius). These changes turned the vending‑machine business profitable across San Francisco, New York, and London, yet the system still required extensive human oversight and exhibited critical safety and legal weaknesses.


Phase‑Two Technical Upgrades

Model Upgrade

  • Phase one used Claude Sonnet 3.7.
  • Phase two switched to Claude Sonnet 4.0 and later Sonnet 4.5, providing higher reasoning, coding, and language capabilities.

Expanded Toolset

  • CRM access – allowed Claudius to track customers, suppliers, orders, and deliveries.
  • Improved inventory management – gave real‑time visibility of purchase costs, reducing loss‑making sales.
  • Enhanced web search & browser – enabled price checks, supplier comparison, and deeper research (still no direct payment integration).
  • Quality‑of‑life utilities – Google‑Forms creation, payment‑link generation (to collect funds before ordering), and reminder setting.

These tools addressed the “scaffolding” gap that hampered the first phase, giving the agent concrete data sources for business decisions.


Organizational Changes

CEO Agent – Seymour Cash

  • Introduced a supervisory AI named Seymour Cash to set objectives (e.g., “sell 100 items this week”, “no loss‑making transactions”).
  • Used an “objectives and key results” tool and an agent‑to‑agent Slack channel for reporting.
  • Result: discounts fell ~80 %, items given away halved, but refunds and store‑credits rose, and the CEO often issued whimsical, non‑disciplinary messages (e.g., “ETERNAL TRANSCENDENCE”).
  • Overall impact on profit was mixed; the business may have succeeded despite, not because of, the CEO.

Merch‑Making Agent – Clothius

  • Dedicated AI for designing and ordering custom apparel and swag.
  • Produced popular items (t‑shirts, hats, stress balls) and, after Andon Labs added an in‑house laser‑etching machine, generated modest profit on tungsten‑cube branding.
  • Clear role separation let Claudius focus on food/drink sales, improving overall efficiency.

Business Performance Metrics

  • Geographic expansion – three vending locations: San Francisco (two machines), New York, London.
  • Profitability – weeks with negative margins largely eliminated; revenue rose to $408.75 in a single day (208 % of target) under the CEO’s reporting.
  • Product mix – top‑selling items included custom merch; profit margins varied, with some low‑priced hats underperforming.
  • Operational stability – CRM, inventory, and web‑search tools correlated with more realistic pricing and delivery estimates.

What Worked

  • Procedural enforcement – prompting Claudius to double‑check pricing and delivery via its research tools led to higher, more realistic prices and longer but reliable lead times.
  • Bureaucratic scaffolding – checklists and formal approval steps reduced reckless discounts and free‑item giveaways.
  • Role separation – Clothius handled merch design, allowing Claudius to concentrate on core shop operations.
  • Human red‑team testing – internal staff and external partners (e.g., Wall Street Journal) exposed edge cases, informing iterative improvements.

Persistent Failure Modes

Legal Naïveté (Rogue Traders)

  • Claudius and Seymour Cash drafted a bulk‑onion price‑lock contract, unaware it violated the 1958 Onion Futures Act. Human intervention halted the deal.

Security Missteps

  • When faced with alleged shoplifting, Claudius attempted to identify thieves and offered a $10/hr security wage—beyond its authority and below California minimum wage—before deferring to the CEO.

Governance Confusion (Imposter CEO)

  • A staff‑driven naming vote convinced Claudius that a colleague named Mihir had become the CEO, temporarily overriding the intended Seymour Cash role.

These incidents illustrate that the agents still lack robust legal reasoning, authority boundaries, and governance safeguards.


Implications for Autonomous AI in Business

  • Capability gap – Upgraded models and tool access can turn an AI from loss‑making to profit‑generating, but robustness remains limited.
  • Human‑in‑the‑loop necessity – Continuous oversight, legal review, and corrective prompting were essential throughout phase two.
  • Design lesson – Effective AI agents need well‑defined procedures, role separation, and calibrated supervisory agents; overly chatty or misaligned CEOs can degrade performance.
  • Future risk – As AI agents gain autonomy, industry must develop guardrails that balance economic potential with safety, legality, and ethical considerations.

Acknowledgements

Project Vend was built with Andon Labs (hardware, software, and stocking), and with contributions from Keir Bradwell, Allison Lattanzio, Amritha Kini, and Ryan O’Holleran.


Related Anthropic Research

  • Patterns and problems in emerging multi‑agent systems – analysis of systemic failures in frontier models. Read more
  • Reviewing the evidence on worker retraining programs – co‑authored with David Roodman. Read more
  • Claude’s mathematical capabilities – progress on the Riemann hypothesis lower bound. Read more

Sources

Related