# RegCorpus — ingestion license, in plain words

This is a plain-language summary for evaluation. The signed agreement is the controlling document.

## Why this corpus can be ingested

The material is public law and public records:

- **State-adopted regulation text is uncopyrightable.** Government edicts carry no copyright — *Georgia v. Public.Resource.Org*, 590 U.S. 255 (2020).
- **Department bulletins and enforcement orders are official public records** of state agencies.
- **We collect from the states directly.** No third-party copies, no reseller feeds, no vendor databases. Every document links to the official source it came from.

That last point is the one that matters for your legal review. Commercial legal publishers acquired their corpora under terms that restrict downstream reuse, and at least one has litigated over AI ingestion of its content — *Thomson Reuters v. Ross Intelligence* (D. Del. 2025), where the fair-use defense for training on Westlaw material was rejected. That is precisely the risk this corpus is built to avoid. We did not take our text from them.

## What the license grants

- **Ingestion rights.** You may load this corpus into your models, indexes, embeddings, retrieval systems, and products.
- **Provenance you can show your customers.** Every document carries its official URL, fetch date, and content hash, so when your product makes a claim, you can show where it came from.
- **Version history.** You can tell what a rule said on a given date, and what changed.

## What the license does not grant

- **NAIC model law text.** Copyrighted by the NAIC. We exclude it. What we carry is the text each state actually adopted, which is that state's law.
- **Material from sources that restrict reuse.** Where a source's terms cap redistribution, we carry facts and an official link, and no text. These records are labeled with `access_class` and `access_notes`. We do not route around anyone's terms.
- **A warranty that our copy is the law.** We are a faithful, dated, hashed copy of official sources, with a link to each one. For anything consequential, verify against the official source — we link it on every record so that is one click.

## How sources are cleared

Before a source is collected, its own terms are read in full and it is classified. Sources whose terms prohibit reproduction are not collected at all; the rule is enforced in the database, which refuses to store them. Sources that permit reuse only non-commercially are collected as facts and links, with no text stored. The per-state provenance tables on regcorpus.com show this posture publicly, source by source.

## Pricing shape

- **Corpus subscription** — bulk export plus daily deltas, for compliance teams.
- **Platform license with ingestion rights** — for companies building AI products on the data.
- **Redistribution / white-label** — for vendors serving the data inside their own product.

Pricing is published rather than quote-only, deliberately. For reference on the category: RegAlytics publishes $107,100/yr for API access and $157,500/yr for white-label redistribution. We are positioned well below that.

Contact: venu@aijourneymates.com
