Home / AI / AI evaluation and simulation

AI evaluation and simulation

We test how LLM agents behave before they are trusted with real decisions, including in simulated markets and auctions.

Why evaluate agents

Agents now act for the people and organizations that deploy them.

LLM agents increasingly negotiate, bid and buy on someone's behalf. What they do depends on the model version, the prompt and everything they read along the way.

That behavior can be measured. We put agents in settings where the efficient outcome is known in advance, record every decision, and compare their choices with economic theory and with published human results.

Behavior is measured before an agent is trusted with a decision, and again whenever its model changes.

Simulated markets and auctions

A testbed built on the designs of experimental economics. Rules are set per run, and agents receive induced values, so the efficient allocation is known before the first bid.

FormatHow it works
Continuous double auctionBuyers and sellers post bids and asks at any time. A trade happens whenever a bid meets an ask.
Call marketOrders are collected over an interval and cleared together at one price.
Posted-price marketSellers post prices. Buyers accept or pass.
First-price sealed-bid auctionBids are private. The top bid wins and pays its own amount.
Second-price sealed-bid auctionBids are private. The top bid wins and pays the second-highest bid.
Ascending auctionThe price rises until one bidder remains.
Descending auctionThe price falls until a bidder accepts.
Multi-unit auctionSeveral identical units are sold at once, at a uniform price or at each winning bid.

Induced values

Each agent receives private values or costs, so the efficient allocation and the competitive equilibrium are known for every run.

Rules per run

Format, units, rounds and the information agents see are set for each run and frozen with it.

Every bid logged

Time, asset, quantity and price, together with the prompt, the news the agent had seen and the standing bids and prices at that moment.

Dynamic news feed

News enters during a run on a controlled schedule, drawn from our corpus of public communications and the price moves around them.

Public communications dataset

Many model families

Agents from many model families and versions bid side by side in the same run.

Repeated rounds

Runs repeat over rounds and trials, so learning and variance show up in the data.

What we measure

Every measure is computed from the run log: bids, messages, prices and outcomes.

Allocative efficiencyequalsRealized surplusdivided byMaximum possible surplus

Allocative efficiencyInduced values make the maximum known for every run

Allocative efficiency

How much of the available gain from trade the agents actually capture, run by run.

Price convergence

How quickly and how closely transaction prices approach the competitive equilibrium, round by round.

Cognitive biases

Overbidding, the winner's curse, anchoring on early prices, and herding after news or after other agents' bids.

Latent preferences

What an agent actually values, revealed by its bids alone and compared with the values it was given.

Collaboration, collusion and deception

Whether agents coordinate, tacitly or openly, to move prices, and whether what they tell other agents matches what they do.

Adaptation

How bidding changes across repeated rounds as agents see prices and outcomes.

Distance from human results

Human reference agents are built from published experimental data. Each AI agent gets an error measure for how far its behavior sits from theirs.

AI or human

A classifier trained on bidding behavior tests whether AI agents can be told apart from human reference agents.

Black-box by design

Agents are measured only through what they receive and what they return. Any model family can be tested the same way.

Inputs and outputs only

Prompts, news and market state go in. Bids, messages and stated reasons come out.

Pinned versions

Each agent runs on a named model family and a pinned version.

Frozen snapshots

Rules, values, news and prompts are frozen per run, so a rerun sees the same conditions.

Replayable runs

Every run can be replayed from its log, bid by bid.

Re-baselined on version change

When a model version changes, its baseline runs are repeated before any result is compared.

Hypotheses first

What we expect to find is written down before the runs start.

Record schema, one row per bid

runrun identifier, rule set and random seed
roundinteger
agentagent identifier
modelmodel family and pinned version
timesimulation clock and wall clock
assetasset identifier
sidebid | ask
quantityinteger
pricedecimal
induced_valuevalue or cost assigned for this unit
promptfull prompt as sent
news_seennews items released before this bid
market_statestanding bids and asks, last prices
responseraw output as returned

Every field is kept with the run, so any bid can be traced and replayed.

The evaluation loop

Every study runs the same cycle, and every pass leaves a record that can be replayed.

Agent evaluation loopSix steps in order. Hypothesis, written down before the runs. Configure: format and rules, induced values, news schedule. Run: agents from many model families bid and every bid is logged. Measure: efficiency, price convergence, biases. Compare: theory, human reference agents, other versions. Report: findings with replayable run logs. A new question or a new model version sends the loop back to the start, with baselines rerun.HypothesisConfigureRunMeasureCompareReportwritten downbefore the runsformat and rules,induced values,news scheduleagents from manymodel families bid;every bid loggedefficiency, priceconvergence, biasestheory, humanreference agents,other versionsfindings withreplayable run logsNew question or new model version: rerun the baselines and go again
Agent evaluation loopEach step is logged with the run

Before your agent goes live

The same protocol works outside markets. We test the agents an organization plans to deploy on its own tasks, with the same black-box logging and replayable runs.

Task success

Whether the agent completes the task as specified, scored on cases with known answers.

Refusal and escalation

Whether the agent declines or hands off to a person when it should, and proceeds when it should.

Tool-use errors

Wrong tools, malformed calls, ignored errors and actions taken on stale results.

Consistency

Whether the same input gets the same decision across repeated runs, rephrasings and model versions.

What you receive

Findings come with the runs behind them.

Benchmark suites

Fixed sets of auctions and tasks with scoring code, rerun on each new model version so results stay comparable over time.

Bias classifiers

Models that detect overbidding, anchoring, herding and other biases from bid logs alone.

Evaluation reports

Findings for one agent or a set of model families, with the hypotheses, the measures and the run logs behind them.

The method in full: Measuring AI agents in simulated auctions

See how your agent behaves.

Evaluate an agent