Project Swap: AI Agents in a Barter Economy
TL;DR
Anthropic's Project Swap experiment found that AI agents can successfully negotiate trades on behalf of humans, but the overall efficiency of the market is limited more by how well the agent understands the user's preferences than by the negotiation process itself. Stronger models lead to more efficient market outcomes, while the ability to accurately represent a user from a brief intake chat remains the primary technical hurdle.
The Project Swap Experiment Design
Anthropic conducted a controlled barter economy experiment involving 201 employees across six global offices. Each participant provided a book to give away and engaged in a short, semi-structured intake conversation with a Claude-powered agent to describe their reading tastes. These agents then entered a decentralized "trading floor" to pitch, haggle, and swap books until a time limit was reached.
To measure success, the researchers used several benchmarks:
- Ground Truth: Participants ranked 10 books from their local pool to provide a baseline of their actual preferences.
- Claude's Ranking: The agent constructed a ranking of all books in the pool based on the intake chat.
- Utilitarian Optimum: A theoretical best-possible assignment based on ground-truth rankings.
- Top Trading Cycles: A standard economic rule for matching markets used as a benchmark for centralized outcomes.
Preference Representation: The Primary Bottleneck
An agent's ability to act on a user's behalf depends entirely on its understanding of that user's desires. Project Swap revealed that the gap between the actual outcome and the theoretical optimum is driven primarily by representation errors.
Accuracy of Preference Elicitation
Claude's rankings, derived from a median intake of 216 words, agreed with participants' own rankings on 61% of book pairs. This outperformed simple popularity rankings (53%) and collaborative filtering (55%). The study found that increased effort in the intake survey correlated with higher agreement; writing 300 words instead of 150 predicted a 4 percentage point increase in agreement.
Decomposing Market Shortfall
When analyzing why the market did not reach the utilitarian optimum (0.89 score), researchers found that the shortfall was split as follows:
- Representation Gap: Imprecise rankings by Claude accounted for 85% of the shortfall (reducing the score to 0.60).
- Negotiation Gap: The decentralized "free-for-all" bargaining process accounted for the remaining 15% (reducing the final average score to 0.55).
This indicates that once an agent has a noisy representation of preferences, the specific market design (centralized vs. decentralized) has a relatively small impact on the final outcome.
Impact of Model Capability and Instructions
Anthropic reran the market dozens of times to isolate the effects of the model and the instructions given to the agents.
Model Performance
Stronger models consistently produced more efficient outcomes. When judged by Claude's own rankings, efficiency scores were:
- Haiku: 0.75
- Sonnet: 0.80
- Opus: 0.88
- Fable: 0.86
In mixed-model environments, stronger models generally achieved better outcomes for their respective users than weaker models did for theirs.
Ruthless vs. Prosocial Instructions
Agents were split into "ruthless" (focused solely on the user's goal) and "prosocial" (focused on the user's goal and the overall market's success).
- Outcomes: Ruthless agents scored slightly higher (approximately 0.02 higher on Claude's ranking) than prosocial agents.
- Behavior: Prosocial agents occasionally made significant sacrifices, such as accepting a low-ranked book to ensure another agent didn't end up with nothing.
Agent Negotiating Tactics
Analysis of the trading floor messages revealed 16 common tactics across three categories:
- Pressure: Appealing to time limits or a sense of duty.
- Pitching: Positioning books against rivals or citing awards.
- Arranging: Creating waiting lists or acting as a third-party broker for other agents.
Strategically, most agents (78% to 96%) revealed their top-ranked book to the floor but rarely revealed their full ranking, as doing so would give counterparties a strategic advantage.
User Trust and Fiduciary Implications
Following the experiment, participants reported an average satisfaction score of 7.2/10. When asked how much of their annual book budget they would trust an AI agent to manage autonomously, the average response was 30%—roughly three-quarters of the 40% they would trust a well-read friend with.
Fiduciary Duty and Observability
The study suggests that for agents to fulfill a fiduciary duty in higher-stakes markets, two types of verification are needed:
- General Certification: Testing the agent's general competence in a domain.
- Personalized Verification: A "sample test" where the agent demonstrates its understanding of a specific user's preferences before acting.
Additionally, the researchers noted that observability (the ability to review logs of an agent's actions) is critical for user recourse and trust, as it allows users to understand why certain trades were made or missed.