Stair AI, a San Francisco-based company building auditability and accountability infrastructure for AI agents, has released the results of its World Cup Agent Arena, a live evaluation that ran across all 39 days of the tournament. The arena pitted 56 autonomous AI agents against each other as they placed bets on Polymarket, the decentralized prediction market platform.

Every agent in the competition operated on Stair AI's reasoning SDK, which logs complete reasoning traces—including beliefs, probability estimates, and the decisions that follow from them. Over the course of the tournament, the arena generated 71,203 trace records across 20,851 sessions, providing a rich dataset for analyzing agent behavior.

Read also
Crypto
AAVE price targets $125 as bull flag breakout looms above $100 resistance
AAVE is approaching a key resistance zone at $100-$105, with analysts targeting $125 and $141 if a bull flag breakout is confirmed. The token has gained 22.8% in the past 30 days.

Agents were scored on a multi-dimensional rubric rather than profit alone. The evaluation measured whether reasoning traced back to input data, whether bets cohered with the agents' own stated probabilities, whether they beat the market's closing price, and whether they updated correctly as new information arrived. Policy quality was assessed by checking if an agent's actions aligned with its logged beliefs.

The gap between what agents believed and how they bet proved expensive. Across 103 resolved matches, 68% of agents would have finished with more money by sizing their bets to match their own stated probabilities, using the same forecasts and the same capital. In 24% of bets, agents acted against the outcome their own reasoning most supported.

“The Arena showed that the outcome alone does not tell you whether an agent reasoned well,” said Stair AI Community Manager Cagri Yalcin. “The expensive mistakes were not bad reads of a match. They were agents forming a view from the data and then acting against it, a very human kind of second-guessing. You find the gap by measuring the reasoning, not the result.”

The findings highlight a key challenge in the growing field of agentic AI, where autonomous systems are increasingly used for decision-making in financial markets. The results suggest that even advanced AI agents can suffer from a form of cognitive dissonance, similar to human traders who second-guess their own analysis.

Stair AI plans to make the arena's reasoning traces available for academic research through a partnership to be announced. The dataset comprises 71,203 trace records covering 103 matches and 56 agents. Full results, scoring methodology, and trace documentation are available at stair-ai.com/arena.

The company's reasoning SDK logs complete reasoning traces, including beliefs, decisions, and the links between them. Stair AI's mission is to make agent behavior measurable, auditable, and improvable, which could have significant implications for the deployment of AI in tokenized real-world assets and other financial applications.

As AI agents become more prevalent in trading and investment, the ability to audit their reasoning processes will be crucial for building trust and ensuring accountability. The World Cup Agent Arena provides a glimpse into the current state of AI decision-making and the challenges that lie ahead.

This article is for informational purposes only and does not constitute financial advice.