Home / Company

How we work

One set of rules covers our datasets, our AI systems and our compute work. Every measured figure carries a source and a date, and results can be traced and run again.

Evidence classes

A price can be an asking price, an offer to one buyer or a completed sale. We record which kind of fact each observation is, and keep the classes apart.

The evidence ladderSix evidence classes, from the bottom of the ladder to the top: specification, what the system is as its maker publishes it; listing, a public asking price at a point in time; quote, a price offered for a stated configuration and quantity; indicative price, a price level a source states for reference; reported sale, a sale as a buyer, seller or broker reports it; verified transaction, a completed sale with confirming documents.Verified transactionA completed sale with confirming documentsReported saleA sale as a buyer, seller or broker reports itIndicative priceA price level a source states for referenceQuoteA price offered for a stated configuration and quantityListingA public asking price at a point in timeSpecificationWhat the system is, as its maker publishes itCloser to what was paidDescribes the system
The evidence ladderEach observation carries exactly one class

Datasets, index series and reports state which classes they draw on, so a reader can tell an asking price from a completed sale. A quote is stored as a quote, and a reported sale stays a reported sale until confirming documents exist.

Evidence classes in GPU pricing

Agent evaluation protocol

Agents are tested only through their inputs and outputs. The same protocol works for any model family, including an organization's own agent.

  1. Write the hypotheses firstHypotheses, conditions and measures are written down before any run.
  2. Pin and freezeModel versions are pinned and environment snapshots frozen, so every agent faces the same conditions.
  3. Run repeated trialsEach condition runs over repeated trials, and every run is logged so it can be replayed.
  4. Re-baseline on changeWhen a model version changes, the baseline runs again before any results are compared.

What every run logs

Enough to replay the run, and for someone else to check the result.

LoggedWhat it holds
PromptsEvery input the agent received, in order
OutputsEvery response and action the agent returned
TimingWhen each input arrived and each output came back
Environment stateThe state of the test environment at each step
Model versionThe pinned version and its settings

How results are tested

A result counts when it holds on data it was not built on and beats the simple alternatives.

Out-of-sample tests

Methods are built on one period or sample and tested on another they have never seen.

Baselines and random controls

Results are compared with simple baselines and with random controls, so a method has to beat both chance and the obvious alternative.

Multiple-comparison correction

When many hypotheses are tested, significance thresholds are corrected for the number of tests run.

Independent reimplementation

Important results are rebuilt from the written method in a separate implementation before we rely on them.

Data provenance

Every record keeps its source, its capture time and the transformations applied to it.

Provenance on every recordData moves from its source, a provider, publisher or system, through capture, where the time is recorded in UTC, and through transformation, where each step is logged in order, into a record. The record carries three provenance fields: source, captured_at and transformations.SourceProvider, publisheror systemCaptureTime recordedin UTCTransformEach step loggedin orderRecordsourcecaptured_attransformationsEach step is written to the record
Provenance travels with the recordDatasets, index series and client deliveries

Provenance stays with the record through datasets, index series and client deliveries, so a figure can be traced back to where it came from and how it was changed.

Models and authorship

How we choose the models we run, and how we publish research.

Models

We choose model families per project, open-weight and commercial, based on the task, the data and where the system has to run.

Versions are pinned, and every model is evaluated on the task before deployment. A model change is a new version of the system and goes through evaluation again.

Authorship

Insights are published by Rillor and carry a publication date. AI tools assist research and drafting, and every figure cites its source.

Read our insights

Ask how we would test your system.

Talk to us