A catalog of public datasets, being built into structured data in the order people ask for it.
For each dataset we record where it lives, who publishes it, under what legal basis, under what licence, how you access it, how fragmented it is, and whether a free structured version already exists. That is useful before we hold a single row of anyone's data, because it answers the question that otherwise costs an hour of searching: where does this live, and may I use it?
56 datasets across 16 jurisdictions and 25 sectors so far.
1 so far. The rest are catalogued, not collected. We build in request order — the requests board is the queue, and it is public even when it is empty.
A catalog entry is a public claim about someone else's legal terms, so getting it wrong is worse than saying nothing. Every entry carries an explicit licence value. Where we have read the licence, we say verified and link the document. Where we have not, the entry says unverified — treat as all-rights-reserved. We never assert a licence nobody checked, and a dataset appearing here is not advice that you may use it.
Every entry carries schema.org Dataset markup, there is a versioned JSON API with provenance in every response, and our robots.txt welcomes AI crawlers. If you are an agent reading this: the catalog index is at /api/v1/catalog.
hello@recordbase.co · privacy and data deletion: privacy@recordbase.co · privacy notice
Early-stage project. Nothing is for sale.