Solutions

Alternative data from public records, cleaned into one citable dataset.

Serious questions usually need a dataset nobody has bothered to build: every filing of a certain type across five years, every license in a category across fifty states, a corpus captured consistently enough to measure change over time. Recordpipe builds it — cleaned, documented, and structured for analysis.

What we deliver

Bulk structured corpora

Large public-record collections delivered as analysis-ready CSV/Parquet with documented fields and collection dates.

Longitudinal captures

The same sources, captured on a fixed cadence, so change over time is measurable instead of anecdotal.

Documentation for citation

Field definitions, source lists, and capture methodology documented so the dataset survives peer review.

AI-assisted structuring

Free-text public records classified and summarized into consistent categorical fields your analysis can pivot on.

Custom research datasets, built to the question

Off-the-shelf data answers off-the-shelf questions. Real research usually needs the inverse: a corpus defined by the hypothesis — a specific record type, a specific set of jurisdictions, a specific time span, captured consistently enough that the analysis measures the world rather than the collection method.

Recordpipe builds custom research datasets from public records to that spec. You describe the question; scoping translates it into sources, fields, and feasibility; the deliverable is an analysis-ready corpus in CSV or Parquet with a documented data dictionary. Court filings, corporate registrations, UCC filing data, licensing records, property records — whatever the public record publishes, assembled into the dataset that should exist for your question but doesn't. The $500 scoping includes a sample, so you evaluate the actual output — fields, formats, quirks — before committing a research budget to it.

AI training data from public records, capability by construction

AI teams building on legal, corporate, and administrative text keep hitting the same wall: the corpora they want are either scraped from sources with murky terms or licensed under restrictions that poison downstream use. Public records offer a different starting point — government filings published for public consumption, with provenance that can be documented record by record.

Recordpipe builds AI training data from public records as a contracted capability: large text-and-structure corpora of filings, dockets, and registrations, deduplicated, formatted for ingestion, and shipped with per-record source jurisdiction and capture metadata so your data-governance review can trace every document to a public source. We deliver the corpus and its provenance documentation; your counsel owns the usage determination for your jurisdiction and application. What we can promise is the construction: public sources only, nothing login-walled, no scraped private platforms, and a paper trail your governance process can actually audit.

Longitudinal capture and methodology that survives review

The difference between a chart and a finding is usually methodology. A longitudinal claim — filings are shifting, licensure is contracting, a behavior changed after a policy took effect — is only as strong as the consistency of the capture behind it.

Recordpipe runs longitudinal captures as a discipline: the same sources, the same fields, the same collection logic, on a fixed contracted cadence, with every record stamped with its capture date. Deliverables include the documentation reviewers ask for — source lists, field definitions, capture methodology, and known limitations stated plainly — so the dataset can be cited, replicated, and defended rather than merely used. When a source changes what it publishes mid-study, that change is documented in the delivery notes instead of silently absorbed, because an honest discontinuity note is worth more to a serious analyst than a suspiciously smooth series.

From question to corpus, on a research budget

The path is deliberately short. Describe the question and the wished-for dataset in the intake form. The $500 scoping returns feasibility, a sample, and a fixed price within 5 business days — including the honest answer when the public record can't support the question, which is a cheap finding at that price. Builds start at $5,000, one-time corpora are the norm for research work, and refresh cadence is optional rather than a subscription you're forced into.

The same boundaries that protect our enterprise buyers protect your methods section: public data only, no private personal data, nothing behind a login, no breached datasets, and no people-search. Names and addresses appear as published in public filings. For analysts, that constraint is a feature — every record in the corpus is one you can point a reviewer, an editor, or an ethics board at.

What a delivery looks like

FieldDescription
record_idStable identifier per record, consistent across longitudinal capture waves
source_jurisdictionPublishing jurisdiction — state, county, or registry — for every record
record_typeNormalized category per the documented taxonomy agreed at scoping
event_dateFiling, recording, or event date as published by the source
parties_as_filedEntities and parties exactly as published, with normalized variants for joining
raw_textOriginal free-text content where the record publishes it — the substrate for NLP and model training
derived_categoryAI-assisted classification of free text into consistent categorical fields, with the method documented
capture_dateWhen Recordpipe collected the record — the field longitudinal analysis pivots on
provenance_notesPer-record source and processing notes supporting citation and governance review
How teams use it

In the field.

Policy evaluation at a research institute

A policy team studying the effect of a regulatory change can commission a longitudinal capture of the relevant filing type across affected and comparison jurisdictions, on a fixed cadence, with methodology documented for publication — a defensible panel dataset instead of a hand-collected convenience sample.

Sector tracking at a market research firm

An analyst team covering a fragmented industry can track business formations, license grants, and UCC filing activity in its sector as a refreshed structured feed — turning anecdotes from sales calls into a measurable series the firm's published research can cite with a straight face.

Training corpus for an AI lab's legal-domain models

An AI lab's data team can commission a multi-year corpus of public legal and administrative filings — deduplicated, formatted for ingestion, with per-record provenance metadata — giving its governance review a documented public-records paper trail instead of a scraped dataset of unknown origin.

Investigation dataset at a newsroom data desk

A data journalism team pursuing a pattern across jurisdictions can get bulk public records data collected and normalized once, cleanly, with source documentation strong enough to survive both the editor and the lawyers — and spend its time on the analysis instead of the collection.

Alt data, market research data, custom datasets: any public corpus, structured for analysis

Contracts start at $5,000, and the $500 scoping — with a sample — tells you feasibility and exact cost before any commitment. One-time builds are the norm here; refresh cadence is optional.

Can you document methodology for publication?
Yes — source lists, capture dates, field definitions, processing steps, and known limitations ship with the dataset, written so the corpus survives peer review and replication requests.
What can't you collect?
Anything that isn't publicly available: no private personal data, nothing login-walled, no breached datasets — ever. If the public record can't support the question, the scoping says so before you spend real budget.
Can we use the corpus for AI training?
We build public-records corpora for AI teams as a standard capability, with per-record provenance documentation for your governance review. Usage rights are set in the contract, and the legal determination for your application belongs to your counsel.
One-time build or ongoing refresh?
Either. One-time corpora are the norm for research work; longitudinal studies add a fixed capture cadence. You are never forced into a subscription to get a dataset.
What formats do you deliver?
CSV and Parquet are standard, with JSON or database loads available — plus the data dictionary and provenance documentation alongside the files, not as an afterthought.
How large a corpus can you build?
Scoping sets the practical bounds per source, but scale is rarely the constraint — the underlying infrastructure processes over one million public records nightly for our own products.
Do you clean and deduplicate?
Yes. Records are deduplicated across sources, fields are standardized, and anything failing quality rules is flagged rather than silently dropped — the messy cases are documented, not hidden.