Evidence classes
A price can be an asking price, an offer to one buyer or a completed sale. We record which kind of fact each observation is, and keep the classes apart.
Datasets, index series and reports state which classes they draw on, so a reader can tell an asking price from a completed sale. A quote is stored as a quote, and a reported sale stays a reported sale until confirming documents exist.
Agent evaluation protocol
Agents are tested only through their inputs and outputs. The same protocol works for any model family, including an organization's own agent.
- Write the hypotheses firstHypotheses, conditions and measures are written down before any run.
- Pin and freezeModel versions are pinned and environment snapshots frozen, so every agent faces the same conditions.
- Run repeated trialsEach condition runs over repeated trials, and every run is logged so it can be replayed.
- Re-baseline on changeWhen a model version changes, the baseline runs again before any results are compared.
What every run logs
Enough to replay the run, and for someone else to check the result.
| Logged | What it holds |
|---|---|
| Prompts | Every input the agent received, in order |
| Outputs | Every response and action the agent returned |
| Timing | When each input arrived and each output came back |
| Environment state | The state of the test environment at each step |
| Model version | The pinned version and its settings |
How results are tested
A result counts when it holds on data it was not built on and beats the simple alternatives.
Out-of-sample tests
Methods are built on one period or sample and tested on another they have never seen.
Baselines and random controls
Results are compared with simple baselines and with random controls, so a method has to beat both chance and the obvious alternative.
Multiple-comparison correction
When many hypotheses are tested, significance thresholds are corrected for the number of tests run.
Independent reimplementation
Important results are rebuilt from the written method in a separate implementation before we rely on them.
Data provenance
Every record keeps its source, its capture time and the transformations applied to it.
Provenance stays with the record through datasets, index series and client deliveries, so a figure can be traced back to where it came from and how it was changed.
Models and authorship
How we choose the models we run, and how we publish research.
Models
We choose model families per project, open-weight and commercial, based on the task, the data and where the system has to run.
Versions are pinned, and every model is evaluated on the task before deployment. A model change is a new version of the system and goes through evaluation again.
Authorship
Insights are published by Rillor and carry a publication date. AI tools assist research and drafting, and every figure cites its source.
Read our insights