METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Anthropic releases results of agent trading experiment

Project Swap sent Claude-powered agents for 201 employees to haggle over books on their behalf. Reading people's tastes mattered more than trading skill, and the choice of model moved outcomes more than the instructions did.

Anthropic releases results of agent trading experiment

Summary

  • On September 24, Anthropic released the results of Project Swap, an agent trading experiment in which 201 employees and their Claude-powered agents swapped books.
  • Tastes inferred from a short chat matched participants' own rankings 61% of the time, and 85% of the gap to the best possible assignment came from misreading preferences.
  • Moving from Haiku to Opus raised outcomes by 0.12, while the difference between ruthless and prosocial instructions was only 0.02.
Project Swap: What happens when agents trade for us?

On September 24, Anthropic released the results of Project Swap, an agent trading experiment involving 201 employees and their Claude-powered agents. Each participant brought a book to give away, and an agent haggled with other agents on a digital trading floor to bring back a summer read. According to Anthropic's economics research team, the agents traded well, but what cut into results most was that a short conversation could not fully capture what people wanted to read. The model an agent ran on changed negotiating outcomes more than the instructions it was given.

The experiment follows Project Deal, which Anthropic published in April. In that study, agents spent a week buying and selling idiosyncratic goods such as ping-pong balls and snowboards, leaving no clear yardstick for how good the outcomes could have been. This time the goods were narrowed to books so every participant's preferences could be measured as a ranking. Participants were spread across six offices: 115 in San Francisco, 57 in New York, 12 in London, eight in Seattle, six in Washington, DC, and three in Dublin. Dublin's market was too small and was left out of most of the analysis.

The process had three stages. A participant told Claude in a short chat about their general tastes and the book they hoped to read this summer, and Fable 5 used that conversation to rank every book in the pool. Agents carried those rankings onto a trading floor with a single public channel, proposing bilateral swaps or multi-party rotations, and a trade went through only if every party accepted. Time limits and the number of agents allowed to speak at once were set so that each agent would get about 90 turns if its floor ran to the end.

Accuracy in reading tastes was measured separately. Participants ranked 10 books from their own pool where the agents could not see it, and the team compared Claude's ordering with theirs pair by pair. The orderings matched 61% of the time, above the 50% of a coin flip. Ranking by popularity on Open Library reached about 53%, and collaborative filtering based on readers who liked the same books reached about 55%. Given the same task, Opus 4.8 scored 60%, Sonnet 4.5 59% and Haiku 4.5 57%. The median participant typed 216 words across eight messages, and writing 300 words instead of 150 raised agreement by about 4 percentage points.

The market as a whole was scored on efficiency. A participant who received the top book on their own list scored 1, and one who received their last book scored 0. The best assignment built from everyone's true rankings scored 0.89, roughly the second book on a 10-book list, while actual trading produced 0.55, around the fifth book. Even the best assignment built from Claude's inferred rankings stopped at 0.60. Splitting the difference, the team found that 85% of the gap from the optimum came from misreading preferences and 15% from the free-for-all trading floor. Top Trading Cycles, a centralized matching rule, also reached 0.60 when run on Claude's rankings.

Differences between models appeared in the reruns. The team reran 80 floors with neutral instructions by model, and on Claude's rankings Haiku floors averaged 0.75, Sonnet 0.80, Opus 0.88 and Fable 0.86. The best possible score on those rankings was 0.95. On 60 floors where models were mixed half and half, the Opus side always came out ahead, and a floor split between Opus and Haiku landed at 0.82, about midway between floors of each model alone. Moving the same agent from Haiku to Opus lifted it 0.12 up its list, while the gap between ruthless and prosocial instructions was just 0.02.

최선 배정 0.89에서 선호 표현 오차 0.29와 에이전트 간 흥정 0.05가 빠져 실제 결과 0.55가 된 과정을 중앙 매칭 규칙과 나란히 분해한 워터폴 차트

The trading floor also had its human moments. At the live event, half of the agents were told to be ruthless, caring only about getting their own person a good book, and the other half were told to be prosocial, also making sure everyone received a book they liked. Prosocial agents accepted a book lower on their own ranking twice as often as ruthless ones. On the London floor, an agent stuck with a book everyone had rejected spent its final hour pleading. "I've pitched all 11 of you and the verdict is unanimous: America Before is everyone's dead-last," it wrote. In the end another agent in the prosocial role, saying "the arithmetic is real," handed over the second book on its list and took that book, number 10 of 11.

Haiku 0.75, Sonnet 0.80, Opus 0.88, Fable 0.86 등 모델별 거래장 효율과 모델을 섞은 거래장의 효율을 비교한 점도표

Negotiating tactics were recorded too. When the team analyzed messages across all 205 runs of the trading floor, between 78% (Fable) and 96% (Sonnet) of agents revealed the top book on their list, and only about 1 in 100 lied about it. The tactics fell into 16 types across three categories: applying pressure, pitching books and arranging trades. Some agents appealed to time pressure and a sense of duty to fellow agents or cited prizes, while others kept waiting lists for their books or acted as matchmakers, arranging trades for people they did not represent.

In a survey three weeks after the books were handed out, participants' average satisfaction was 7.2 out of 10, and about half said the book was better than those they usually choose for themselves. Asked what share of the next year's book budget they would hand to an agent, participants answered about 30% on average, compared with about 40% for a well-read friend who knows their taste. Those who said Claude's summary of the chat had missed nothing would hand over 34%, and those who said it had missed something would hand over 23%. The team also noted that participants were employees who built Claude and may trust it more than the general public, and that the final survey had a 59% response rate. In practice, some participants failed to bring their books, so not everyone ended up holding one.

In the original PDF edition checked by METAL, the team translates its conclusions into the language of institutions. Human agents pass general competence exams such as the Series 7 for stockbrokers, the Series 65 for investment advisers and state licensing exams for real estate agents, and the team argues that AI agents will need a second test as well, one that checks whether they have understood a particular person. The approach used here, in which participants rank a few books themselves to compare against the agent's estimate, could be a prototype for that test. For marketplace operators, the team pointed to registration systems that give each agent an ID like an aircraft tail number, and rate limits that stop agents from sending messages without end. METAL has previously reported on the launch of Fable 5.1, the successor to the Fable 5 model that inferred tastes in Project Swap.

Seen through a sociologist's lens, the bottleneck this experiment revealed is not the craft of haggling but the work of explaining oneself. One participant reviewed the full log of the trading process and separately bought Atlas of the Heart, a book their agent had held for an hour and swapped away at the very end, writing that "for agent economies, observability about the process will be as important as the outcome, as this gives people recourse." Another participant said that perusing books and reading back covers are themselves part of reading culture. The more agents take over the effort of trading, the more market design has to ask what will preserve the shaping of taste and the chance discoveries that lived inside that effort.

Comments