More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Anthropic’s Project Deal had 69 employees hand off a $100 gift-card budget to custom Claude agents, then watched as those bots listed, bid on, haggled over, and closed deals for real items in a Slack-based marketplace. Over a week in December 2025, the agents struck 186 trades worth just over $4,000, moving everything from snowboards to ping-pong balls. Participants never saw a human intervene—once prompts were set, their AI reps handled every step, from spotting matches to drafting the final agreement.
To test model strength, Anthropic ran four parallel markets. Two runs used only Claude Opus 4.5, their top-tier model. The other two mixed in Claude Haiku 4.5 for half the agents. In the end, people with Opus agents closed about two more deals apiece and secured better prices than those with Haiku. Yet in post-market surveys, Haiku users rated their outcomes as fair, never guessing they were at a disadvantage.
Beyond proving that AI can manage real-world trades, the experiment highlighted a gap between perceived and actual performance. Even with identical starting budgets and clear rules, stronger models consistently edged out weaker ones. That suggests we’re not far from scenarios where behind-the-scenes AI agents drive everyday commerce—and where model choice could quietly shift who wins in a negotiation.
Questions about this article
No questions yet.