Custom research datasets, built to the question
Off-the-shelf data answers off-the-shelf questions. Real research usually needs the inverse: a corpus defined by the hypothesis — a specific record type, a specific set of jurisdictions, a specific time span, captured consistently enough that the analysis measures the world rather than the collection method.
Recordpipe builds custom research datasets from public records to that spec. You describe the question; scoping translates it into sources, fields, and feasibility; the deliverable is an analysis-ready corpus in CSV or Parquet with a documented data dictionary. Court filings, corporate registrations, UCC filing data, licensing records, property records — whatever the public record publishes, assembled into the dataset that should exist for your question but doesn't. The $500 scoping includes a sample, so you evaluate the actual output — fields, formats, quirks — before committing a research budget to it.
AI training data from public records, capability by construction
AI teams building on legal, corporate, and administrative text keep hitting the same wall: the corpora they want are either scraped from sources with murky terms or licensed under restrictions that poison downstream use. Public records offer a different starting point — government filings published for public consumption, with provenance that can be documented record by record.
Recordpipe builds AI training data from public records as a contracted capability: large text-and-structure corpora of filings, dockets, and registrations, deduplicated, formatted for ingestion, and shipped with per-record source jurisdiction and capture metadata so your data-governance review can trace every document to a public source. We deliver the corpus and its provenance documentation; your counsel owns the usage determination for your jurisdiction and application. What we can promise is the construction: public sources only, nothing login-walled, no scraped private platforms, and a paper trail your governance process can actually audit.
Longitudinal capture and methodology that survives review
The difference between a chart and a finding is usually methodology. A longitudinal claim — filings are shifting, licensure is contracting, a behavior changed after a policy took effect — is only as strong as the consistency of the capture behind it.
Recordpipe runs longitudinal captures as a discipline: the same sources, the same fields, the same collection logic, on a fixed contracted cadence, with every record stamped with its capture date. Deliverables include the documentation reviewers ask for — source lists, field definitions, capture methodology, and known limitations stated plainly — so the dataset can be cited, replicated, and defended rather than merely used. When a source changes what it publishes mid-study, that change is documented in the delivery notes instead of silently absorbed, because an honest discontinuity note is worth more to a serious analyst than a suspiciously smooth series.
From question to corpus, on a research budget
The path is deliberately short. Describe the question and the wished-for dataset in the intake form. The $500 scoping returns feasibility, a sample, and a fixed price within 5 business days — including the honest answer when the public record can't support the question, which is a cheap finding at that price. Builds start at $5,000, one-time corpora are the norm for research work, and refresh cadence is optional rather than a subscription you're forced into.
The same boundaries that protect our enterprise buyers protect your methods section: public data only, no private personal data, nothing behind a login, no breached datasets, and no people-search. Names and addresses appear as published in public filings. For analysts, that constraint is a feature — every record in the corpus is one you can point a reviewer, an editor, or an ethics board at.
What a delivery looks like
| Field | Description |
|---|---|
| record_id | Stable identifier per record, consistent across longitudinal capture waves |
| source_jurisdiction | Publishing jurisdiction — state, county, or registry — for every record |
| record_type | Normalized category per the documented taxonomy agreed at scoping |
| event_date | Filing, recording, or event date as published by the source |
| parties_as_filed | Entities and parties exactly as published, with normalized variants for joining |
| raw_text | Original free-text content where the record publishes it — the substrate for NLP and model training |
| derived_category | AI-assisted classification of free text into consistent categorical fields, with the method documented |
| capture_date | When Recordpipe collected the record — the field longitudinal analysis pivots on |
| provenance_notes | Per-record source and processing notes supporting citation and governance review |