Why evaluate agents
Agents now act for the people and organizations that deploy them.
LLM agents increasingly negotiate, bid and buy on someone's behalf. What they do depends on the model version, the prompt and everything they read along the way.
That behavior can be measured. We put agents in settings where the efficient outcome is known in advance, record every decision, and compare their choices with economic theory and with published human results.
Behavior is measured before an agent is trusted with a decision, and again whenever its model changes.
Simulated markets and auctions
A testbed built on the designs of experimental economics. Rules are set per run, and agents receive induced values, so the efficient allocation is known before the first bid.
| Format | How it works |
|---|---|
| Continuous double auction | Buyers and sellers post bids and asks at any time. A trade happens whenever a bid meets an ask. |
| Call market | Orders are collected over an interval and cleared together at one price. |
| Posted-price market | Sellers post prices. Buyers accept or pass. |
| First-price sealed-bid auction | Bids are private. The top bid wins and pays its own amount. |
| Second-price sealed-bid auction | Bids are private. The top bid wins and pays the second-highest bid. |
| Ascending auction | The price rises until one bidder remains. |
| Descending auction | The price falls until a bidder accepts. |
| Multi-unit auction | Several identical units are sold at once, at a uniform price or at each winning bid. |
Induced values
Each agent receives private values or costs, so the efficient allocation and the competitive equilibrium are known for every run.
Rules per run
Format, units, rounds and the information agents see are set for each run and frozen with it.
Every bid logged
Time, asset, quantity and price, together with the prompt, the news the agent had seen and the standing bids and prices at that moment.
Dynamic news feed
News enters during a run on a controlled schedule, drawn from our corpus of public communications and the price moves around them.
Many model families
Agents from many model families and versions bid side by side in the same run.
Repeated rounds
Runs repeat over rounds and trials, so learning and variance show up in the data.
What we measure
Every measure is computed from the run log: bids, messages, prices and outcomes.
Allocative efficiencyequalsRealized surplusdivided byMaximum possible surplus
Allocative efficiency
How much of the available gain from trade the agents actually capture, run by run.
Price convergence
How quickly and how closely transaction prices approach the competitive equilibrium, round by round.
Cognitive biases
Overbidding, the winner's curse, anchoring on early prices, and herding after news or after other agents' bids.
Latent preferences
What an agent actually values, revealed by its bids alone and compared with the values it was given.
Collaboration, collusion and deception
Whether agents coordinate, tacitly or openly, to move prices, and whether what they tell other agents matches what they do.
Adaptation
How bidding changes across repeated rounds as agents see prices and outcomes.
Distance from human results
Human reference agents are built from published experimental data. Each AI agent gets an error measure for how far its behavior sits from theirs.
AI or human
A classifier trained on bidding behavior tests whether AI agents can be told apart from human reference agents.
Black-box by design
Agents are measured only through what they receive and what they return. Any model family can be tested the same way.
Inputs and outputs only
Prompts, news and market state go in. Bids, messages and stated reasons come out.
Pinned versions
Each agent runs on a named model family and a pinned version.
Frozen snapshots
Rules, values, news and prompts are frozen per run, so a rerun sees the same conditions.
Replayable runs
Every run can be replayed from its log, bid by bid.
Re-baselined on version change
When a model version changes, its baseline runs are repeated before any result is compared.
Hypotheses first
What we expect to find is written down before the runs start.
Record schema, one row per bid
| run | run identifier, rule set and random seed |
|---|---|
| round | integer |
| agent | agent identifier |
| model | model family and pinned version |
| time | simulation clock and wall clock |
| asset | asset identifier |
| side | bid | ask |
| quantity | integer |
| price | decimal |
| induced_value | value or cost assigned for this unit |
| prompt | full prompt as sent |
| news_seen | news items released before this bid |
| market_state | standing bids and asks, last prices |
| response | raw output as returned |
Every field is kept with the run, so any bid can be traced and replayed.
The evaluation loop
Every study runs the same cycle, and every pass leaves a record that can be replayed.
Before your agent goes live
The same protocol works outside markets. We test the agents an organization plans to deploy on its own tasks, with the same black-box logging and replayable runs.
Task success
Whether the agent completes the task as specified, scored on cases with known answers.
Refusal and escalation
Whether the agent declines or hands off to a person when it should, and proceeds when it should.
Tool-use errors
Wrong tools, malformed calls, ignored errors and actions taken on stale results.
Consistency
Whether the same input gets the same decision across repeated runs, rephrasings and model versions.
What you receive
Findings come with the runs behind them.
Benchmark suites
Fixed sets of auctions and tasks with scoring code, rerun on each new model version so results stay comparable over time.
Bias classifiers
Models that detect overbidding, anchoring, herding and other biases from bid logs alone.
Evaluation reports
Findings for one agent or a set of model families, with the hypotheses, the measures and the run logs behind them.
The method in full: Measuring AI agents in simulated auctions