Home / Data

Data built for AI

Datasets that state their sources and their limits. Market data, GPU and neocloud pricing, legal and public records, and custom builds, each documented down to the record.

Six datasets, one standard

Each dataset states what it contains, where it comes from, how it stays current and what it can support.

DatasetWhat it holdsUsed forLink
Market dataDEX, U.S. stocks, futuresOrder-book snapshots and diffs, fills, liquidations and funding from on-chain perpetual futures venues. Trade and quote history for U.S. equities and futures.Forecasting, simulation calibration, price-discovery researchRecord
GPU cloud pricing and availabilityNeocloud and hyperscalerPrices and availability that GPU cloud providers publish through their pricing APIs and pricing pages, by GPU class, region, term and pricing type.Capacity planning, procurement, the Rillor Compute IndexRecord
GPU systems pricingSpecifications and pricesSystem specifications, normalized configurations and price observations, kept apart by evidence class: listing, quote, reported sale, verified transaction.Procurement comparison, valuation, financing diligenceRecord
Public communications and market responseEvent dataPublic statements aligned in time with price moves, with random-baseline controls.Event studies, news-driven simulation, methodology testingRecord
Legal datasetsDecisions and authoritiesLegal decisions and authorities with document structure and citations.Legal research, citation analysis, retrieval systemsRecord
Public procurement recordsU.S. contractingContracting notices, awards and documents from public U.S. sources, with document text extraction and change tracking.Supplier analysis, sourcing research, document AIRecord

Need data that does not exist yet? We scope, build and maintain custom datasets.

Custom datasets

DEX, U.S. stocks, futures

Market data

Order-book data from decentralized exchanges and trade and quote history for U.S. stocks and futures, aligned on one clock for research and AI.

Market data in detail
What it holds
On-chain perpetual futures order books rebuilt from snapshots and book diffs, with fills, liquidations and funding. Trades and quotes for U.S. equities and futures from 2018.
Sources
Public data published by on-chain perpetual futures venues. Trade and quote history from commercial data providers.
Record types
Book snapshot, book diff, fill, liquidation, funding rate, trade, quote.
Refresh
Historical archive. Extensions and recurring updates are set in the data agreement.
Used for
Forecasting research, calibrating simulated markets, price-discovery studies, backtesting.
Scope
Each rebuilt book reflects the snapshots and diffs its source published. Coverage windows are stated per instrument.
Access
Access is provided under a written data agreement. U.S. equities and futures history supports research and custom builds within the rights of each source.

Neocloud and hyperscaler

GPU cloud pricing and availability

What GPU cloud providers charge to rent a GPU, and what they report as available. Each observation records the price as published, its source and its UTC capture time.

GPU cloud pricing in detail
What it holds
Published prices and reported availability for 11 GPU classes, from H100 and A100 to GB300 NVL72 and MI355X, by provider segment, region, term and pricing type.
Sources
Official provider pricing APIs and published pricing pages, used within each source's terms.
Record types
Price observation, availability observation, raw snapshot.
Refresh
The methodology captures each provider source with a UTC timestamp and keeps every raw snapshot for audit.
Used for
Capacity planning, procurement, financing diligence, the Rillor Compute Index.
Scope
Prices and availability as each provider publishes them, normalized per GPU-hour.
Access
Access is provided under a written data agreement.

Specifications and prices

GPU systems pricing

What GPU server systems cost to own. Specifications and price observations for NVIDIA and AMD systems, each typed by the kind of evidence behind it.

GPU systems we supply
What it holds
System specifications, normalized configurations and price observations for HGX, NVL72 and Instinct systems.
Sources
Manufacturer specifications and system catalogs. Distributor, reseller and broker listings and quotes, where collection and use rights permit. Reported secondary sales with recorded provenance.
Record types
System specification, normalized configuration, price observation. Configurations are modeled by GPU family, GPU count, memory, CPU, interconnect, cooling, condition, location and seller type.
Refresh
Set per source. Every observation keeps its capture time and source.
Used for
Procurement comparison, valuation and depreciation, financing diligence, the RCI Hardware series.
Scope
Each observation keeps its evidence class, so a listing is read as a listing and a quote as a quote.
Access
Access is provided under a written data agreement.

Evidence classes

Evidence classWhat the record asserts
SpecificationSystem and component attributes as the manufacturer publishes them.
ListingA public offer, seen at a stated time.
QuoteA price supplied on request for a stated configuration and quantity.
Indicative priceA reported price level without a verifiable counterparty.
Reported saleA sale a source reports, without independent confirmation.
Verified transactionA completed sale with confirming evidence.

Event data

Public communications and market response

Public statements preserved as issued and aligned in time with price moves, with random-baseline controls for comparison.

How we use it in AI evaluation
What it holds
Selected public statements with their entities, classifications and timestamps, each paired with price observations in a declared window around it.
Sources
Public statements and their source metadata. Price series from public or licensed sources.
Record types
Statement as issued, entity reference, classification with its classifier version, time-aligned price observation.
Refresh
A bounded historical collection, extended to new windows and sources by agreement.
Used for
Event studies, news feeds for agent evaluation and simulation, methodology testing.
Scope
Records timing and association. Each window comes with random-baseline controls, so a move after a statement can be compared with moves at random times.
Access
Access is provided under a written data agreement.

U.S. contracting

Public procurement records

Contracting notices, awards and their documents from public U.S. sources, with the full text extracted and every revision kept.

What it holds
Notices, awards, suppliers, line items and attached documents, linked across sources.
Sources
Public U.S. procurement sources and the documents they publish.
Record types
Notice, award record, supplier entity, item record, document text, change record.
Refresh
Collected on a recurring schedule per source. Each change is kept as a new version.
Used for
Supplier analysis, sourcing research, document AI over long, mixed-format files.
Scope
Records as each public source publishes them, with document text extracted, including OCR for scanned files.
Access
Access is provided under a written data agreement.

Source data rarely arrives ready

Sources change, records disagree and identities drift. We turn that material into datasets a downstream system can inspect and use.

Fragmented sourcesThe records you need sit across many publishers, portals, filings and formats, with no shared index.
Inconsistent records and identitiesThe same entity appears under different names, identifiers, units and field conventions.
Changing formatsSources revise their structure and withdraw material, so one-time extracts go stale.
Rights and provenanceOrigin and usage rights get lost during collection, and records become unusable downstream.
Conflicting observationsTwo sources report different values for the same fact. We keep both, record which one is used and why.

How we build datasets

Every stage keeps the evidence the next stage needs: source context, identity decisions, transformations, versions and stated limits.

  1. CollectAcquire and preserve source material with its origin, capture time and the rights that apply to it.
  2. LabelIdentify records, entities, attributes and relationships, and assign each observation its evidence class.
  3. NormalizeReconcile structures, identities, units, terms and time, without erasing the source record.
  4. MaintainVersion the dataset, track changes, review quality, reconcile conflicts and publish its limits.

How we work in detail

Custom datasets

Need data that does not exist yet? Tell us what it has to cover. We scope it, build it and keep it current.

What you specify

Source universe
Which publishers, portals, filings, feeds or systems of your own are in scope.
Entities
The record and entity types that must resolve and join.
Fields
Required attributes and relationships, and the labels your systems expect.
History
How far back the records reach, as far as sources allow.
Geography
The jurisdictions and regions the dataset covers.
Refresh cadence
How current the data stays, and what a change should trigger.
Delivery
An authenticated API, a scheduled export or an agreed file format.

Each custom dataset comes with its schema, records with resolved identities, source provenance, lineage, version history, quality notes and stated limits.

Notices

Data rights. Source data is provided within the rights of each source.

Trademarks. NVIDIA, AMD and related product names are trademarks of their owners and are used only to identify products. No affiliation or endorsement is implied.

Tell us what data you need.

Request data