What ships with a custom data API
A custom data API from Recordpipe is a finished integration surface, not a database dump with a URL. The build includes:
- Endpoints designed around your product's queries — search, single-record fetch, history, and a change feed — shaped the way your developers would have designed them.
- API keys with per-key scoping, separating staging, production, and internal tooling from day one.
- Written documentation with example requests and responses your team can build against without a kickoff call.
- A versioned schema with a published change policy — additive changes flow in, breaking changes get a new version and a migration window.
- Webhook subscriptions for push delivery when records matching your criteria appear or change.
Staging keys are live during the build, so your integration and our pipeline are tested against each other before launch, not after.
Data as a service API: the pipeline is the product
The endpoint is the visible part. What makes this a data as a service API rather than a hosted spreadsheet is the managed collection layer behind it. On the cadence set in scoping, our automation gathers the underlying public records, normalizes them to your schema, deduplicates entities, and computes what changed — and those deltas are what feed the change endpoint and your webhooks.
Every record carries provenance: a capture date and a source category, so your product can show users how current a value is and your auditors can trace where it came from. The stance underneath doesn't move: publicly available sources only — nothing login-walled, nothing private, no breached data, under any contract.
Because collection and serving are one system, freshness is observable rather than promised: a status endpoint reports the last completed refresh per dataset, for your monitoring to alert on.
A managed scraping API for developers — not a scraping tool with an API
Most products sold as a scraping API take a URL and return raw page content. Your team still writes the parsers, still maintains them when pages change, and still owns data quality — the hard part never left your backlog. A managed scraping API for developers inverts that: you never send a URL, and you never see a page. You query records.
Extraction, parsing, validation, and repair all run on our side as a managed web scraping service, on the same production infrastructure that processes over one million public records nightly. Your integration surface is stable JSON with a contract we version. When a source restructures, we fix the pipeline behind the endpoint; your code doesn't know it happened.
Your developers' time goes into your product's features. The part that decays — source maintenance — sits with the vendor built to absorb it.
From scoping call to production keys
The path to a live endpoint is short and priced up front. Describe the dataset and how your product needs to query it; we reply within 2 business days with feasibility and an approach. The $500 scoping — credited in full against the contract — delivers a data sample and a proposed endpoint design within 5 business days, so your engineers review real JSON, not a slide.
From there: we finalize the schema together, stand up the pipeline and staging environment, your team integrates against staging keys, and launch happens when both sides sign off. The build is fixed-price — contracts run from $5,000 to $3 million depending on dataset scope and refresh cadence — and ongoing operation is a retainer quoted alongside it, with no metered-billing surprise.
When a custom API becomes a product
Some datasets stop being an internal dependency and start being a feature you sell. We build for that path from the start. Keys are multi-tenant by design, usage is metered per key, and the documentation is written cleanly enough to hand to your own customers. If your roadmap includes exposing the data inside your product — an in-app lookup, a verification step, an alerts feed — the API you commissioned is already shaped for it.
The commercial side scales the same way: licensing terms for downstream use are set in your contract, not renegotiated per customer you sign. Teams that later ship the endpoint as a product feature do it on the same build and the same pipeline, with capacity resized in the retainer rather than re-architected.
An example endpoint spec
| Endpoint | Method | Description |
|---|---|---|
| /v1/records/search | GET | Query the dataset by the parameters your product actually filters on — name, jurisdiction, status, date range. Paginated, with stable sort. |
| /v1/records/{id} | GET | Fetch a single record in full, including capture date and source category for every field group. |
| /v1/records/{id}/history | GET | Prior observed values for a record, so your product can show what changed and when. |
| /v1/changes | GET | Cursor-based change feed: everything added, updated, or no longer present at source since your last cursor. |
| /v1/webhooks | POST | Register a push subscription — your criteria, your endpoint, signed payloads. |
| /v1/webhooks/{id} | DELETE | Remove a subscription. Deliveries stop immediately; the audit log of past pushes remains queryable. |
| /v1/exports | POST | Start a bulk export job — full dataset or a filtered slice — in CSV, JSON, or Parquet. |
| /v1/exports/{id} | GET | Poll export status and retrieve the download link when the job completes. |
| /v1/status | GET | Pipeline health and last completed refresh per dataset, for your monitoring to alert on. |