Auctions give AI agents a test with a known answer. Economic theory predicts what a bidder should do in each format, and laboratory economics has measured what people actually do for more than sixty years. When an LLM agent bids under the same rules, its bids can be scored against both.
This note sets out Rillor's protocol for that test: how runs are set up, what is measured, how human reference agents are built and what every run records. It also draws on studies published through September 2026 that put LLM bidders in these formats, because their results shape what the protocol measures.
Control comes from induced values
In his 1962 study, Vernon Smith had subjects trade a good under double auction rules. Each seller held a private minimum price and each buyer a private maximum. Nobody saw the full supply and demand schedules, so nobody could compute the equilibrium price. Trading prices still came close to it, and their spread narrowed over five trading periods [1, 7].
The method behind that result is induced value, which Smith set out formally in 1976 [2, 7]. Each participant is told privately what a unit is worth to them. A buyer earns that value minus the price paid, and a seller earns the price minus its cost. Because the experimenter sets every value, the efficient allocation and the competitive equilibrium are known before trading starts [7].
The protocol gives agents induced values in the same way. Whatever preferences a model brings with it matter only through the payoffs the run defines, and every bid can be scored.
Formats and their benchmarks
Each format comes with a theoretical prediction and a published record of what people did in the laboratory.
| Format | Theory | What people did |
|---|---|---|
| Second-price sealed bid, private values | Bidding one's value is a dominant strategy (Vickrey, 1961) [3, 8] | Many bid above value. In Kagel and Levin's 1993 experiment, 67.2 percent of bids were above value and 27.0 percent within five cents of it [5] |
| Ascending clock | The same dominant strategy [8] | Subjects found it more easily than in the sealed version [8] |
| First-price sealed bid | No dominant strategy; equilibrium bids sit below value [8] | Persistent bidding above the risk-neutral equilibrium [5] |
| Common-value sealed bid | Bidders should discount for the winner's curse [8] | The winner's curse was common among inexperienced bidders, and larger groups bid more aggressively [6, 8] |
| Double auction | Prices settle at the competitive equilibrium [7] | Prices near equilibrium, with dispersion falling round by round [7] |
The protocol also covers descending (Dutch) auctions, multi-unit auctions and posted prices, with the rules fixed for each run.
What the protocol measures
Allocative efficiency is realized surplus as a share of the maximum possible surplus. Price convergence uses Smith's coefficient α: the standard deviation of transaction prices around the equilibrium price, as a percentage of that price, computed round by round [7, 12]. Bid ratios compare each bid with the benchmark bid for its format and value. In common-value formats the protocol counts how often winners lose money and by how much.
Efficiency alone is a weak test. In 1993 Dhananjay Gode and Shyam Sunder showed that programs submitting random bids and offers, barred only from trading at a loss, pushed double auction efficiency close to 100 percent [9]. A high efficiency score can come from the market rules alone. The protocol therefore reports efficiency next to measures of individual behavior: overbidding, anchoring on early prices, round-to-round adaptation and, in sessions with several agents, patterns consistent with collusion such as bid rotation and matched bids.
What recent studies found
Three studies published or revised in 2026 show why each of these measures earns its place.
Size and direction differ from people. Anand Shah and colleagues (revised September 2026) ran five LLMs across seven laboratory settings against human benchmarks rebuilt from published experiments. No model's deviation from theory fell within the human range in more than two of five private-value formats. In second-price auctions, humans mostly overbid, while most models that deviated underbid. The three large models without extended reasoning still preserved the human ordering of formats by difficulty: first-price harder than second-price, and ascending clocks easier than sealed bids. The reasoning model bid almost exactly at equilibrium [11].
Convergence is slower. Pawel Struski and colleagues (September 2026) replicated Smith's double auction with three populations of LLM traders. None reached full convergence. In Smith's data, α fell from 11.8 in round 1 to 3.5 in round 5. The closest LLM market went from 27.6 to 11.7, with efficiency rising from 0.79 to 0.91. In another population, agents improved bids a cent at a time without closing the spread, and efficiency ranged from 0.36 to 0.67 across rounds [12].
Wording moves outcomes. Sara Fish, Yannai Gonczarowski and Ran Shorrer put two LLM bidders in a repeated first-price auction under two prompts that differed by one clause. Under one prompt, bidders bid well below value and earned an average profit of 0.115 times their value. Under the other, they bid close to value and earned 0.006 times their value [13].
The protocol takes three rules from this work. Score size, direction and ordering separately. Treat the prompt as part of the treatment. Report efficiency only next to behavior.
Black-box by design
Agents are measured only through what they receive and what they return. Each run logs the prompt, the market state and any news item the agent saw, the raw output, the parsed bid, any parse failure and the timing.
Model versions are pinned and the environment is frozen, so every run can be replayed from its log. Each condition runs over repeated trials with recorded seeds, and hypotheses are written down before the runs. When a model version changes, the reference conditions run again on the new version before any comparison, and earlier results keep the version label they were produced under.
Pinning has limits. One model provider's documentation states that changes to its serving infrastructure can produce small differences in behavior even when the model ID and weights are unchanged [16]. Reference conditions therefore rerun on a schedule as well as on every version change.
Human reference agents
Published experiments mostly report summary statistics, such as overbidding shares and bid regressions, and seldom every bid. The protocol builds human reference agents as decision rules fitted to those statistics for each format, carries the fitting uncertainty forward, and compares orderings and directions where the reconstruction cannot support levels. Shah and colleagues took the same route, fitting a mixture of equilibrium bidders, overbidders and underbidders to each source [11].
Every LLM agent gets an error measure against the human reference under the same format and values. A classifier trained to tell LLM agents from human reference agents using bids alone adds one more measure. Agents from several model families run side by side under identical conditions.
Why it matters
When an agent bids, negotiates or buys on someone's behalf, its owner needs to know how it behaves when incentives are clear. An auction with induced values is one of the few settings where the right answer is known before the agent acts, so the gap between what it did and what it should have done is a number with a direction.
Read the full report
The full report, available as a PDF, sets out the run setup, the measures, the logging schema, the human reference construction and the limits of the protocol, with every source listed.