Home / Data
Data built for AI
Datasets that state their sources and their limits. Market data, GPU and neocloud pricing, legal and public records, and custom builds, each documented down to the record.
Six datasets, one standard
Each dataset states what it contains, where it comes from, how it stays current and what it can support.
| Dataset | What it holds | Used for | Link |
|---|---|---|---|
| Market data | Order-book snapshots and diffs, fills, liquidations and funding from on-chain perpetual futures venues. Trade and quote history for U.S. equities and futures. | Forecasting, simulation calibration, price-discovery research | Record |
| GPU cloud pricing and availability | Prices and availability that GPU cloud providers publish through their pricing APIs and pricing pages, by GPU class, region, term and pricing type. | Capacity planning, procurement, the Rillor Compute Index | Record |
| GPU systems pricing | System specifications, normalized configurations and price observations, kept apart by evidence class: listing, quote, reported sale, verified transaction. | Procurement comparison, valuation, financing diligence | Record |
| Public communications and market response | Public statements aligned in time with price moves, with random-baseline controls. | Event studies, news-driven simulation, methodology testing | Record |
| Legal datasets | Legal decisions and authorities with document structure and citations. | Legal research, citation analysis, retrieval systems | Record |
| Public procurement records | Contracting notices, awards and documents from public U.S. sources, with document text extraction and change tracking. | Supplier analysis, sourcing research, document AI | Record |
Need data that does not exist yet? We scope, build and maintain custom datasets.
Custom datasetsDEX, U.S. stocks, futures
Market data
Order-book data from decentralized exchanges and trade and quote history for U.S. stocks and futures, aligned on one clock for research and AI.
Market data in detail- What it holds
- On-chain perpetual futures order books rebuilt from snapshots and book diffs, with fills, liquidations and funding. Trades and quotes for U.S. equities and futures from 2018.
- Sources
- Public data published by on-chain perpetual futures venues. Trade and quote history from commercial data providers.
- Record types
- Book snapshot, book diff, fill, liquidation, funding rate, trade, quote.
- Refresh
- Historical archive. Extensions and recurring updates are set in the data agreement.
- Used for
- Forecasting research, calibrating simulated markets, price-discovery studies, backtesting.
- Scope
- Each rebuilt book reflects the snapshots and diffs its source published. Coverage windows are stated per instrument.
- Access
- Access is provided under a written data agreement. U.S. equities and futures history supports research and custom builds within the rights of each source.
Neocloud and hyperscaler
GPU cloud pricing and availability
What GPU cloud providers charge to rent a GPU, and what they report as available. Each observation records the price as published, its source and its UTC capture time.
GPU cloud pricing in detail- What it holds
- Published prices and reported availability for 11 GPU classes, from H100 and A100 to GB300 NVL72 and MI355X, by provider segment, region, term and pricing type.
- Sources
- Official provider pricing APIs and published pricing pages, used within each source's terms.
- Record types
- Price observation, availability observation, raw snapshot.
- Refresh
- The methodology captures each provider source with a UTC timestamp and keeps every raw snapshot for audit.
- Used for
- Capacity planning, procurement, financing diligence, the Rillor Compute Index.
- Scope
- Prices and availability as each provider publishes them, normalized per GPU-hour.
- Access
- Access is provided under a written data agreement.
Specifications and prices
GPU systems pricing
What GPU server systems cost to own. Specifications and price observations for NVIDIA and AMD systems, each typed by the kind of evidence behind it.
GPU systems we supply- What it holds
- System specifications, normalized configurations and price observations for HGX, NVL72 and Instinct systems.
- Sources
- Manufacturer specifications and system catalogs. Distributor, reseller and broker listings and quotes, where collection and use rights permit. Reported secondary sales with recorded provenance.
- Record types
- System specification, normalized configuration, price observation. Configurations are modeled by GPU family, GPU count, memory, CPU, interconnect, cooling, condition, location and seller type.
- Refresh
- Set per source. Every observation keeps its capture time and source.
- Used for
- Procurement comparison, valuation and depreciation, financing diligence, the RCI Hardware series.
- Scope
- Each observation keeps its evidence class, so a listing is read as a listing and a quote as a quote.
- Access
- Access is provided under a written data agreement.
Evidence classes
| Evidence class | What the record asserts |
|---|---|
| Specification | System and component attributes as the manufacturer publishes them. |
| Listing | A public offer, seen at a stated time. |
| Quote | A price supplied on request for a stated configuration and quantity. |
| Indicative price | A reported price level without a verifiable counterparty. |
| Reported sale | A sale a source reports, without independent confirmation. |
| Verified transaction | A completed sale with confirming evidence. |
Event data
Public communications and market response
Public statements preserved as issued and aligned in time with price moves, with random-baseline controls for comparison.
How we use it in AI evaluation- What it holds
- Selected public statements with their entities, classifications and timestamps, each paired with price observations in a declared window around it.
- Sources
- Public statements and their source metadata. Price series from public or licensed sources.
- Record types
- Statement as issued, entity reference, classification with its classifier version, time-aligned price observation.
- Refresh
- A bounded historical collection, extended to new windows and sources by agreement.
- Used for
- Event studies, news feeds for agent evaluation and simulation, methodology testing.
- Scope
- Records timing and association. Each window comes with random-baseline controls, so a move after a statement can be compared with moves at random times.
- Access
- Access is provided under a written data agreement.
Legal
Legal datasets
Legal decisions and authorities, structured for research and retrieval.
Grellum, the research system built on it- What it holds
- Decisions and authorities with document structure, issuing authority, citations and the relationships between them.
- Sources
- Decisions as published by the issuing authority, with the retrieval metadata that traces each record to its origin.
- Record types
- Published decision, source metadata, authority identity, document structure, citation.
- Refresh
- Checked for completeness against the source. Most recent full check 7 Oct 2026.
- Used for
- Legal research, citation analysis, retrieval systems, evaluation of legal AI.
- Scope
- Decisions as the issuing body publishes them. The issuing body's text remains the official version.
- Access
- Access is provided under a written data agreement.
U.S. contracting
Public procurement records
Contracting notices, awards and their documents from public U.S. sources, with the full text extracted and every revision kept.
- What it holds
- Notices, awards, suppliers, line items and attached documents, linked across sources.
- Sources
- Public U.S. procurement sources and the documents they publish.
- Record types
- Notice, award record, supplier entity, item record, document text, change record.
- Refresh
- Collected on a recurring schedule per source. Each change is kept as a new version.
- Used for
- Supplier analysis, sourcing research, document AI over long, mixed-format files.
- Scope
- Records as each public source publishes them, with document text extracted, including OCR for scanned files.
- Access
- Access is provided under a written data agreement.
Source data rarely arrives ready
Sources change, records disagree and identities drift. We turn that material into datasets a downstream system can inspect and use.
How we build datasets
Every stage keeps the evidence the next stage needs: source context, identity decisions, transformations, versions and stated limits.
- CollectAcquire and preserve source material with its origin, capture time and the rights that apply to it.
- LabelIdentify records, entities, attributes and relationships, and assign each observation its evidence class.
- NormalizeReconcile structures, identities, units, terms and time, without erasing the source record.
- MaintainVersion the dataset, track changes, review quality, reconcile conflicts and publish its limits.
Custom datasets
Need data that does not exist yet? Tell us what it has to cover. We scope it, build it and keep it current.
What you specify
- Source universe
- Which publishers, portals, filings, feeds or systems of your own are in scope.
- Entities
- The record and entity types that must resolve and join.
- Fields
- Required attributes and relationships, and the labels your systems expect.
- History
- How far back the records reach, as far as sources allow.
- Geography
- The jurisdictions and regions the dataset covers.
- Refresh cadence
- How current the data stays, and what a change should trigger.
- Delivery
- An authenticated API, a scheduled export or an agreed file format.
Each custom dataset comes with its schema, records with resolved identities, source provenance, lineage, version history, quality notes and stated limits.
Notices
Data rights. Source data is provided within the rights of each source.
Trademarks. NVIDIA, AMD and related product names are trademarks of their owners and are used only to identify products. No affiliation or endorsement is implied.