Source provenance
Source provenance is the record of where a piece of regulatory information came from, when it was published and retrieved, and exactly what it said.
Source provenance is the record that ties a piece of information back to its origin: which authority published it, at what address, on what date, in which version, and what the text said when it was retrieved. The W3C's provenance standard defines provenance as "information about entities, activities, and people involved in producing a piece of data or thing, which can be used to form assessments about its quality, reliability or trustworthiness" (PROV-Overview, W3C Working Group Note, 30 April 2013).
In regulatory monitoring, provenance is what lets a reviewer check an item instead of trusting it. A usable record keeps at least the source URL, the issuing body, the publication date as the source states it, the retrieval date, and a verbatim excerpt of the passage that triggered the item. Each field prevents a specific failure. Regulators' pages often show a "last updated" or republication date rather than the original one, so an item without both dates can be mis-dated and treated as new. A summary without a verbatim excerpt cannot be checked against the text, which matters when a language model wrote the summary, because fluent output can still misstate its source. NIST's generative AI profile (NIST AI 600-1, July 2024) makes content provenance one of the four primary considerations its suggested actions address.
Provenance also has to survive time. Consolidated texts change, links rot and pages move, so the record should capture what was seen as well as where to look. In RegWatch, each finding keeps its source URL, dates and a verbatim excerpt, and the health of each source is tracked. To test any tool, pick five items at random and try to reach the exact passage behind each one. The LLM accuracy article turns that test into a control, and relevance scoring depends on it, because a score is only as sound as the text it was based on.
