Making public interest data discoverable
This lookup doesn’t exist yet. Securities, energy, research, finance, and nonprofits each have one. Seven precedents show how to build water’s — by standardizing the pointer, not the data.
Pick a case study to explore







Emerging technologies such as LLMs shift data policy priorities from standardization to discoverability, and then verifiability
- The barrier is discoverability, not data format. Tools read any rate sheet; they cannot reliably find each utility's current adopted sheet or confirm it is the version a board adopted.
- Build a registry of pointers, not a data standard. Map each utility to the canonical URL of its current adopted rate schedule, keyed to the PWSID (Public Water System Identification number), with the adoption recorded as data — adoption date, board action, effective date.
- Populate it through a filing that already exists. Add one field, the adopted rate-sheet URL, to California's Electronic Annual Report (eAR), which every public water system already files; this inherits statewide coverage at near-zero cost.
- Publish it openly; let an aggregator build discovery. The state or the data collaborative supplies coverage and authority; an aggregator such as whatwatercosts.org supplies the public search layer.
- Sequence the rollout: voluntary first, mandate second. Start with a small self-description file (an llms.txt for utilities) that cooperative agencies publish and the aggregator harvests, then pursue the eAR field for full coverage.
- Next: move from discovery to verifiability. Once the registry makes rate sheets discoverable, the next step confirms each pointer states the rates a board actually adopted and still in force — reconciling self-reported links against the board action of record, certifying identity through the signed crosswalk, and confirming the eAR field against the current schema before it carries weight.
A step by step guide to addressing water data discoverability
1. Anchor the registry on the PWSID
SEC EDGAR resolves a company name or ticker to a Central Index Key (CIK); the Internal Revenue Service joins each nonprofit's filings through its Employer Identification Number (EIN); the Legal Entity Identifier (LEI) does the same for financial firms. Water already holds the equivalent key in the PWSID, which persists through name changes and separates utilities with similar names. Build the registry on the PWSID and resolve names and aliases to it.
2. The registry points rather than hosts
EDGAR keeps a text document as the official filing and layers tagged data — XBRL, the eXtensible Business Reporting Language — on top, marked unofficial. The library harvesting protocol exposes a short record plus a pointer to where each document lives. The water asset is a table mapping PWSID to utility to the canonical rate-sheet URL. It need not hold or reformat the rate sheets, because the extractor reads them.
3. The registry records adoption status as a field
The verification problem, adopted rates against a draft study, is answered with metadata in every case that faced it. EDGAR stamps each submission with a filing type and an acceptance time and links amendments to their originals. Energy regulators derive the currently effective tariff from filed effective dates and supersession. The web's rel=canonical tag lets a publisher name the authoritative version among duplicates. The registry should record an adoption event on each pointer, the adoption date, the board action, and the effective date, so a field rather than a guess from the document text answers whether a rate sheet is in force.
4. An existing mandatory filing populates the registry at low cost
EDGAR, energy tariff filing, Form 990, and the LEI each rode an obligation that already existed: EDGAR moved a standing disclosure duty to one electronic channel, the Form 990 reform changed only the format and openness of a return nonprofits already filed, and LEI adoption followed reporting rules that required the identifier. California's counterpart is the eAR. Every public water system already files it, keyed to the PWSID, and its Section 8 collects rate-structure data but no link to the adopted schedule. Adding one field, the URL of the current adopted rate schedule with its adoption and effective dates, would turn an existing filing into the registry and cover every public water system in the state. This warrants confirmation against the current eAR schema before it reaches a client.
5. Where no authority covers every publisher, a low-friction convention spreads
No body holds authority over every U.S. water utility, just as none governed every research repository or every website. The library harvesting protocol (OAI-PMH, the Open Archives Initiative Protocol for Metadata Harvesting) and the web's self-description files spread because each asked little of the publisher. The water analogue is a small file at a fixed path on each utility's site, an llms.txt for utilities, that names the canonical rate-sheet URL, the adoption date, the rate structure, and local context. A utility webmaster can add it in one sitting, and an aggregator harvests and reconciles the rest.
6. Publish the registry openly and let an aggregator build discovery
The SEC publishes a free, CIK-keyed data feed and builds no consumer product; the Global Legal Entity Identifier Foundation (GLEIF) publishes its full dataset under a public-domain (CC0) license; the IRS released machine-readable returns, and ProPublica, not the IRS, built the search tool. An open PWSID-to-URL-to-adoption-metadata table, in CSV and JSON, lets whatwatercosts.org occupy that service-provider role while the state or the data collaborative supplies coverage and authority.
7. Adoption follows a dominant consumer or a mandate
The web conventions spread because Google rewarded publication with search visibility, the dynamic that also drove transit-schedule adoption in the earlier case set. The LEI spread under "no LEI, no trade" reporting rules. A voluntary water file will see uneven uptake until a widely used benchmarking or affordability tool makes accurate appearance depend on it. The identity layer follows GLEIF: a neutral body publishes an open, certified crosswalk from the PWSID to Department of Water Resources (DWR) and local identifiers, on the model of GLEIF's published mapping between the LEI and the bank Business Identifier Code (BIC), and signs it once the identifier and crosswalk exist.
Design implications
- Define the
llms.txt-for-utilities convention (canonical URL, adoption date, rate structure, context) and have whatwatercosts.org harvest it, recruiting utilities already in the project's contact set, such as Las Virgenes, Moulton Niguel, and Eastern Municipal, as initial adopters. - Advocate adding the adopted-rate-sheet-URL field to the State Water Board's eAR, which inherits statewide coverage from an existing mandate.
- Publish a PWSID-to-DWR-to-local crosswalk on the GLEIF model, certified and later signed, drawing the common grammar from the Moulton Niguel-led CaDC definitions whitepaper.
- Identify the consumer that rewards publication, and tie accurate appearance in a benchmarking or affordability tool to publishing the file.
For the nine investor-owned Class A water utilities that the California Public Utilities Commission (CPUC) regulates, serving about 16% of Californians, the registry can pull entries directly from the commission's advice-letter and rate-case records by utility number. The eAR field then covers the publicly owned systems no commission regulates, about 84% of roughly 50,000 community water systems nationally.
Limitations
- Voluntary participation produces incomplete coverage. Utilities without an incentive publish nothing, and the system must fall back to ordinary scraping for them.
- A self-asserted pointer can go stale or point to the wrong document. A published "current adopted" marker states the utility's claim, and a high-stakes use should reconcile it against the board action of record; self-reported URL fields require link validation.
- Persistence carries a standing operational cost. Keeping each pointer live is continuing work, which the scholarly identifier system funds through registrant fees and which a water registry must fund some other way.
- The smallest utilities report last and report least, as the smallest nonprofits do under the abbreviated Form 990-N. Coverage reaches completeness before it reaches uniform depth, so the design should phase in by size and keep a manual fallback.
- An identifier establishes reachability, not correctness. It locates the document; confidence in the content rests on the registrant's credibility and the adoption metadata.
SEC Electronic Data Gathering, Analysis, and Retrieval
A federal mandate routed every issuer's disclosures into one electronic repository, keyed to a stable identifier, with public access kept free.
Before EDGAR, a public company's disclosures existed as paper. A registrant printed its 10-K, prospectus, or proxy and mailed copies to the SEC, which kept them in public reference rooms. Anyone who wanted a current filing traveled to a reference room and stood at a photocopier, or paid a microfiche vendor such as Disclosure, Inc. The friction was not the document format; a 10-K was already a structured, regulated document. The friction was reachability. There was no single place where a reader could be confident of looking at a company's current, authoritative filing.
What turned EDGAR from a 1984 pilot into a discoverability protocol was a combination of five elements: a legal mandate, a single canonical repository, a stable identifier, an unambiguous notion of the authoritative filing, and free machine-reachable dissemination. The SEC adopted Regulation S-T in February 1993 and phased in mandatory electronic filing over roughly three years, completing on May 6, 1996. The graduated schedule and a hardship exemption resolved the mandate-versus-burden tension: coverage became universal while small filers were eased in.
Each filer receives a Central Index Key, the join key that lets a reader resolve a company name or ticker to a number and key all of that entity's filing history off it. The CIK persists across name changes. EDGAR stamps each submission with a filing type and acceptance datetime and threads amendments to originals, so the operative version is answered by metadata rather than guesswork. The SEC opened a free public web site in September 1995, and Congress later made free public access a statutory floor.
What this means for water rate data
EDGAR is the template for the core asset: a registry that resolves an identifier to a publisher's current filing, with currency carried as a stamped field rather than inferred from the document. The PWSID plays the CIK's role as the stable key. EDGAR also shows the discipline that made it useful — keep the canonical layer thin and the data feed free, and leave analysis to a third-party ecosystem, the role whatwatercosts.org would occupy.
The problem
Before EDGAR, discovery ran through intermediaries: the reference room, the microfiche vendor, the financial-data resellers, each adding cost, delay, and the risk that what a reader held was stale or a copy of a copy. Stanford's research library still warns that "before the mid-1990s, SEC filings can be challenging to find because electronic filings weren't required and EDGAR didn't exist." A small investor effectively could not get the same timely access as a Wall Street firm with a runner at the SEC. The SEC framed EDGAR's purpose against exactly this: to increase the efficiency and fairness of the securities market "by accelerating the receipt, acceptance, dissemination, and analysis of time-sensitive corporate information."
What failed first
The pre-EDGAR arrangements each fell short on reachability rather than content. The paper public-reference-room model gave authoritative access only to whoever physically showed up. The microfiche resale model (Disclosure, Inc. from the 1970s) broadened distribution but introduced a copy layer, added cost, and lagged the filing. Commercial financial-data products reformatted and indexed the data, but each was a proprietary silo with its own coverage gaps and its own price. EDGAR's own beginning was a long partial attempt: the SEC started the work in September 1984, the first pilot filing was a Form S-8 by the Southern Company that year, and for nearly a decade EDGAR was a voluntary pilot with thin coverage, useless as a canonical registry because a reader could never assume a given company was in it.
The mechanism that worked
In February 1993 the SEC adopted Regulation S-T (17 CFR Part 232) and began a graduated phase-in that completed on May 6, 1996, after which all public domestic companies were required to file on EDGAR except for hardship-exemption paper filings. By fall 1995, more than 92% of public companies were already filing electronically. The mandate later extended to foreign companies and governments (November 4, 2002) and to insider Forms 3/4/5 (June 30, 2003).
EDGAR became the single store of record. The September 1997 report to Congress recorded that the system had received and processed 1.3 million documents in roughly 491,000 submissions from over 28,000 filing entities, held in a single word-searchable database of nearly 53 gigabytes with 99.9% availability. Each filer receives a Central Index Key, "used on the SEC's computer systems to identify corporations and individual people who have filed disclosure with the SEC." The CIK is the join key, and the submissions record carries both current and former names. EDGAR is explicit about authority: "Only documents submitted to the EDGAR system in either plain text or HTML are official filings. PDF documents are unofficial copies of filings." The system stamps each submission with a filing type and acceptance datetime and threads amendments to originals.
The SEC opened its public web site in September 1995; by the March 1997 peak, 17 gigabytes were downloaded in a single day. When Congress pushed to privatize the system, the National Securities Markets Improvement Act of 1996 required that any privatization "maintain free public access to data filings in the EDGAR system." On top of that floor the SEC built machine interfaces: a full-text search over filings since 2001, daily indexes and RSS feeds, and the data.sec.gov REST APIs, which deliver JSON-formatted data, require no authentication or API keys, and are keyed by the 10-digit CIK, plus nightly bulk ZIP archives. In Release No. 33-9002, adopted January 30, 2009 (effective April 13, 2009), the SEC required financial statements in XBRL, phased in over three stages through June 15, 2011. XBRL did not replace the document: the plain-text or HTML filing stayed the official record and the XBRL exhibit was furnished, not filed.
Outcome
EDGAR became the assumed substrate of US capital markets, recorded at over 3,000 filings per day and more than 17 million filings as of mid-2025. The more important outcome is the third-party ecosystem: because dissemination was free, machine-readable, and keyed to a stable identifier, an industry of financial-data vendors, research platforms, and academic corpora built on top of it. The protocol the SEC owned was narrow, and it deliberately left the analysis and value-added products to others. The limits are equally clear: full-text keyword search reaches back only to 2001; the official filing is still a text or HTML document with XBRL furnished and historically uneven in tagging quality; coverage has holes by design; and the system answers "what was filed and when," not "is this disclosure true."
Sources
Utility Tariff Filing
Regulated utilities file their rates with a regulator, and the filed copy becomes the only legally operative one, which makes the regulator's system the canonical lookup.
Anyone who needs to know what an electric or gas utility currently charges faces a discoverability question before any data question. The price schedule exists somewhere, but finding the authoritative, current copy among a utility's own documents, press releases, and superseded versions is the hard part. Regulated energy utilities solved this structurally rather than technologically: the utility's prices are not merely published, they are filed with a regulator, and the regulator's copy becomes the single authoritative version. Under the filed rate doctrine, a utility may charge only the rate it has filed, and the act of filing, not any later approval, is what gives the rate legal force. As one court put it, "it is the filing of the tariffs, and not any affirmative approval or scrutiny by the agency, that triggers the filed rate doctrine."
FERC closed the discoverability gap with Order No. 714, "Electronic Tariff Filings" (Docket No. RM01-5-000), issued September 19, 2008 and published at 73 FR 57,515. The Final Rule required that all tariffs, tariff revisions, and rate-change applications for public utilities, natural gas pipelines, oil pipelines, and the federal power administrations be filed electronically. Mandatory electronic filing phased in beginning April 1, 2010 on a six-month staggered schedule, after which FERC no longer accepted paper tariff filings. The electronic format followed standards developed with the North American Energy Standards Board (NAESB), so filings from every utility shared a common structure.
eTariff filings are organized as discrete, addressable tariff records, each carrying an effective date and superseding a prior record, so the "currently effective" version of any record becomes a query rather than a manual reconstruction. FERC's eLibrary holds the filings and the public eTariff Viewer renders them. Filings are keyed to the utility and to docket numbers; California's CPUC mirrors this with a Utility Reference Number (U#) and numbered advice letters filed under General Order 96-B.
What this means for water rate data
Energy utilities are discoverable because law requires them to file their rates with a regulator, whose system becomes the authoritative lookup. Most water utilities have no such regulator: the nine investor-owned utilities the CPUC oversees serve about 16% of Californians, while roughly 84% of the country's 50,000 community water systems are publicly owned and set rates by a local board vote under Proposition 218, with no central filing. The transferable move is to manufacture the missing chokepoint cheaply, by adding a rate-sheet-URL field to the eAR, and to treat the adopting board resolution and its effective date as the water equivalent of the "currently effective tariff."
The problem
A utility serves a defined territory, publishes its own materials, and changes its prices on its own schedule. Multiply that by thousands of regulated utilities and the question "what does utility X charge today, and is this the version in force?" has no obvious place to be answered. A regulated tariff is the legally binding, published schedule of the rates, terms, and conditions a utility must charge. Because the filed copy is the only copy that matters legally, the regulator's filing system becomes the canonical lookup, and the current rate is wherever the regulator says the currently effective tariff lives.
What failed first
Before electronic filing, tariffs were paper. A utility maintained its rates as a set of individually numbered tariff sheets, and a rate change meant filing revised sheets that superseded specific prior sheets. The legal architecture was sound: California's Public Utilities Code §491 has long required that no rate change be made "except after 30 days' notice to the commission and to the public," with notice given "by filing with the commission and keeping open for public inspection new schedules." The findability was not. Reconstructing a utility's currently effective tariff meant assembling the latest non-superseded version of every sheet by hand. FERC framed the shift to electronic filing in terms of reduced storage and processing time, lower mailing and courier fees, and concurrent access for multiple parties.
The mechanism that worked
FERC Order No. 714 (124 FERC ¶ 61,270; Docket No. RM01-5-000), issued September 19, 2008 and published at 73 FR 57,515, required all tariffs, tariff revisions, and rate-change applications for public utilities, gas pipelines, oil pipelines, and the federal power administrations to be filed electronically, phasing in from April 1, 2010 on a six-month staggered schedule. Four design features turn that mandate into a discoverability mechanism. eTariff filings are organized as discrete, addressable tariff records carrying metadata including an effective date, the machine-readable inheritance of the paper tariff sheet. Because every record carries an effective date and supersedes a prior record, the system computes the "currently effective" version of any record as of any date while retaining superseded versions for the audit trail. FERC's eLibrary holds the filings and the public eTariff Viewer (etariff.ferc.gov) renders them; in FERC's description the viewer "allows the public access to view the status of Tariffs which have been submitted to FERC" and to "download currently effective tariff records of a particular tariff." Filings are keyed to the utility and to docket numbers. California's CPUC assigns each utility a U-number and moves rate changes through numbered advice letters filed under General Order 96-B, with tier-based effective-date rules; a utility whose gross intrastate revenues exceed $10 million "shall publish, and shall thereafter keep up-to-date, its currently effective California tariffs at a site on the Internet."
Outcome
The combination produced a single authoritative place to find any regulated utility's current rates, with the current version computable from effective dates and the supersession history preserved. For FERC-regulated entities the eTariff Viewer and eLibrary answer the question by identifier; at the state level the CPUC's advice-letter process plus the U# identifier and the internet-publication mandate provide the same answer for California's regulated energy utilities. The mechanism scales because it does not rely on each utility choosing to be findable: the filing obligation makes the regulator's index complete, the NAESB structure makes filings uniform, and the effective-date logic makes currency machine-determinable. The model presupposes a regulator with mandatory jurisdiction over every publisher in scope, which is mostly absent for publicly owned water, and it does not by itself standardize the contents of a rate sheet.
Sources
Open Archives Initiative Protocol
Data providers expose minimal metadata and a pointer; service providers harvest and aggregate, a voluntary federated model that asks little of each publisher.
By the late 1990s, scholarly output was scattering. Physicists posted preprints to arXiv, computer scientists used NCSTRL, economists used RePEc, and individual universities were standing up their own repositories. Each archive held real, current documents; none knew about each other. A reader had to know which archive held a paper, navigate to it, and search it on its own terms. The same shape of problem had appeared a generation earlier in libraries, where a book sat in thousands of separate institutions, each with its own card catalog. In both cases the obstacle was discovery across many independent holders, not the format of the underlying object.
Two complementary mechanisms carried discovery across independent holders. OCLC's WorldCat centralized the record: one library catalogs a work once to the MARC standard, and every other holder attaches a lightweight holding symbol to the existing record rather than re-describing the book. The cost of that model is its centralization; it works because participants agreed to join one cooperative. The Open Archives Initiative federated the index instead. Its harvesting protocol, OAI-PMH, grew from the Santa Fe meeting of October 21–22, 1999; version 2.0, the stable version still in use, was released on June 14, 2002 with Carl Lagoze and Herbert Van de Sompel as editors.
OAI-PMH is deliberately small: six verbs over plain HTTP, returning UTF-8 XML, with a mandatory floor of unqualified Dublin Core (15 flat elements). The data-provider / service-provider split is the heart of the design. A data provider exposes its own holdings, lightly described, with a pointer to where the full object lives; a service provider harvests across many providers, normalizes, and builds the discovery experience. Datestamps and selective harvesting answer the currency question: a service provider re-harvests only what changed.
What this means for water rate data
OAI-PMH is the voluntary path: thousands of independent repositories each expose a short record plus a pointer, and aggregators harvest them with no shared platform and no mandate. For water, the utility is the data provider — exposing its PWSID, adoption date, and rate-sheet URL — and whatwatercosts.org is the service provider that harvests and deduplicates by PWSID. The protocol's datestamps, which let a harvester detect what changed without re-reading everything, are the precedent for the adoption-date field that signals currency.
The problem
Each scholarly archive held valuable, current documents, but no single place told a reader what existed, who held the authoritative copy, and where that copy lived. The same gap had appeared in libraries: a book sat on shelves in thousands of separate institutions, each maintaining its own card catalog. The obstacle in both cases was discovery across many independent holders. This is the gap that blocks whatwatercosts.org today: a utility's current adopted rate sheet is a perfectly readable document, but nothing tells an aggregator which utility publishes it, at what canonical URL, and whether the version on that page is the current one.
What failed first
The instinct to solve scattered discovery with a single, rich, synchronous query protocol came first. Z39.50, an ANSI/NISO standard whose work began in the 1970s and which went through versions in 1988, 1992, 1995, and 2003, let a client send a live structured query to a remote library server. It worked and remains "the backbone of library resource sharing for over 30 years," but it is heavyweight: it assumes a well-resourced server, a rich and consistently applied MARC record, and substantial implementation effort on both ends. That was achievable for professionally staffed libraries inside the OCLC cooperative; it was not realistic for a graduate student running an e-print server. When the e-print community looked at the problem in 1999, the Santa Fe meeting "decided that a low-barrier solution was critical" and adopted metadata harvesting instead, on the reasoning that a model where each archive merely exposed a flat list of its metadata for periodic bulk pickup demanded almost nothing.
The mechanism that worked
OCLC began in 1967 as the Ohio College Library Center; Frederick G. Kilgour introduced shared cataloging in 1971 for 54 Ohio academic libraries. The core move was deduplication of description: one library catalogs a work once to MARC, and every other holder attaches a lightweight holding symbol rather than re-describing it. The result, WorldCat, crossed one billion holding symbols on August 11, 2005. Z39.50 is the protocol layer over this. The cost of the model is its centralization, a strong governance ask the repository world could not make.
OAI-PMH grew from the Santa Fe meeting of October 21–22, 1999, convened by Paul Ginsparg, Rick Luce, and Herbert Van de Sompel; version 1.0 went public in January 2001 and version 2.0 on June 14, 2002. The protocol is six verbs: Identify, ListMetadataFormats, ListSets, ListIdentifiers, ListRecords, and GetRecord, over plain HTTP GET or POST, with well-formed UTF-8 XML responses and a mandatory unqualified Dublin Core floor. The data-provider / service-provider split is the heart of the design: "Data Providers are repositories that expose structured metadata via OAI-PMH. Service Providers then make OAI-PMH service requests to harvest that metadata." An OAI-PMH record does not contain the paper; it contains a short Dublin Core description and an identifier, typically a URL, pointing to where the full object lives on the repository's own server. Every record carries a datestamp, and from/until arguments let a harvester ask only for records changed within a date range; deletions are explicit.
Outcome
The harvesting model built a real discovery layer over thousands of uncoordinated repositories. OAIster, started at the University of Michigan in 2002 on a Mellon grant, grew to tens of millions of records from more than 1,500 organizations, and OCLC absorbed it into WorldCat in 2009. BASE (Bielefeld Academic Search Engine) built a Lucene/Solr index over OAI-PMH harvests and now covers more than 12,000 content providers. CORE, built at the Open University by Petr Knoth from 2011, indexes over 400 million scholarly resources and hosts more than 25 million free-to-read full texts. None of these required the underlying repositories to adopt a common platform or move their content. The limits are instructive: the unqualified Dublin Core floor is shallow and "not adequate for describing journal articles," which is one reason Google Scholar built its own full-text crawling; an unexpectedly large number of repositories misconfigured their endpoints into "dark" ones that yield nothing; and the protocol carries metadata, not content. OAI-PMH became the cheap, voluntary substrate that the actual discovery services were built on.
Sources
The Web Self-Description Stack
Small files at predictable locations declare canonical URLs, freshness, and authority; adoption followed a dominant consumer rather than a mandate.
By the mid-1990s the web had a discovery problem that looked nothing like a data-format problem. Millions of independent sites published content that a crawler had to find, fetch without overloading the server, and recognize as current. No central registry listed what any site contained or when it last changed. A search engine confronting a new domain had three questions and no agreed way to ask them: which parts may I fetch, what are the canonical pages, and which changed recently enough to re-index. This is the same shape as the water-rate discovery problem. The web solved its version with a stack of small, voluntary self-description conventions at predictable locations, the direct ancestor of the llms.txt idea.
The stack is layered, each part a small file or tag at a predictable location solving one question. robots.txt (Martijn Koster, February 1994) established the pattern: a single plain-text file at a fixed path that any crawler checks without being told, later formalized as RFC 9309 in September 2022. Google introduced the Sitemaps protocol in June 2005, an XML file enumerating canonical URLs with an optional lastmod date; in November 2006 Google, Yahoo!, and Microsoft jointly announced support and stood up sitemaps.org. RSS and Atom (RFC 4287, December 2005) broadcast new and updated content, and WebSub (a W3C Recommendation on 23 January 2018) turned the freshness signal into a publisher push.
The remaining question, which of several URLs returning similar content is authoritative, was answered by the canonical link element, announced jointly by Google, Yahoo, and Microsoft on 12 February 2009 and formalized as RFC 6596 in April 2012. None of these spread by mandate. Each reached ubiquity because a dominant consumer made adoption pay, the same dynamic that drove transit agencies to publish GTFS once Google Maps consumed it.
What this means for water rate data
The web's self-description files are the working model for an llms.txt for utilities: name the canonical rate-sheet URL (the sitemap pattern), carry a last-adopted date as the freshness signal, mark which document is current versus draft (the rel=canonical pattern), and let a dated changelog stand in for change notification. The case's harder lesson is adoption — these conventions spread only once a dominant consumer rewarded them, so a water file needs an equivalent consumer, most plausibly a benchmarking or affordability tool utilities want to appear in correctly, rather than elegance.
The problem
A search engine confronting a new domain had three questions with no agreed way to ask them: which parts may I fetch, what are the canonical pages, and which changed recently enough to re-index. The blocker for water rate data is the same: knowing the canonical URL where a utility publishes its current adopted rates, and having confidence the document found there is the live version rather than a draft, an archived tariff, or a board-packet attachment from three cycles ago. Four tensions are native to any voluntary self-description scheme. A site-level file a crawler reads can only ever be advisory; RFC 9309 states "these rules are not a form of access authorization." Simplicity pulls against expressiveness. Self-description carries no inherent authority guarantee. And freshness signaling is a cheap way for a publisher to declare currency.
What failed first
The first instinct was centralized and curatorial: human-maintained directories such as the early Yahoo! catalog, and the idea that a complete list of all crawlers could be kept by hand. The robots.txt origin story records that in 1994 "the internet was small enough to maintain a complete list of all bots," an assumption that collapsed within a few years. Syndication's first attempt at a single format failed in a softer way: RSS was created at Netscape in March 1999, carried to RSS 2.0 by Dave Winer and UserLand in September 2002, and its specification was deliberately frozen, which left real needs unmet and produced a competing standard (Atom) rather than a unified one. Real-time freshness was the last gap; polling was the failure mode WebSub was built to retire.
The mechanism that worked
Martijn Koster proposed robots.txt in February 1994; by June 1994 it was a de facto standard observed by WebCrawler, Lycos, and AltaVista. Its lasting contribution is the pattern: a single file at a fixed path that any crawler knows to check. Google worked with Koster to bring it into the IETF, and RFC 9309 published in September 2022. Google introduced the Sitemaps protocol (version 0.84) in June 2005, an XML file enumerating canonical URLs (loc) with optional lastmod, changefreq, and priority. The decisive move was governance plus a consumer coalition: in November 2006 Google, Yahoo!, and Microsoft jointly announced support, bumped the schema to 0.90, and stood up sitemaps.org, with Google releasing the schema under a Creative Commons license to lower the barrier. RSS established the feed pattern; Atom rebuilt it under IETF stewardship and shipped as RFC 4287 in December 2005. PubSubHubbub, built by Brad Fitzpatrick and Brett Slatkin at Google in 2009, retired feed polling; renamed WebSub in October 2017, it became a W3C Recommendation on 23 January 2018. The canonical link element, announced jointly on 12 February 2009, lets a page declare the version that should be indexed, "a hint, not a directive." None of these spread by mandate; each reached ubiquity because a dominant consumer made adoption pay. Formal standardization arrived late, ratifying conventions that already worked.
Outcome
The stack became close to universal where a dominant consumer rewarded it and patchy where none did. robots.txt and sitemaps are effectively standard equipment for any site that wants to be found by search engines. RSS and Atom underpin podcasting and a large share of news distribution. WebSub powers real-time delivery for parts of the open social web. rel=canonical is routine SEO practice. The limits track the tensions: adoption is incomplete because it is voluntary; the signals are self-asserted, so they describe intent, not ground truth, and can be spoofed or gamed; and there is no authority guarantee, since a sitemap or rel=canonical tag tells you which URL the publisher prefers, not which document a regulator adopted.
Sources
Digital Object Identifier
A persistent identifier, a central resolver, and federated registration, so the identifier survives URL change and resolves to the current authoritative copy.
By the late 1990s scholarly publishing had moved online and citations began to break. A reference pointed at a publisher's URL, a location rather than a name; when a publisher redesigned its site or sold a journal, the location changed and the citation died. Norman Paskin, founding director of the International DOI Foundation, framed the failing directly: a stable identifier should let you "locate an entity, or provide services irrespective of changes in location or management responsibility of the entity." The damage is measurable: one 2022 study of more than 50,000 cited URLs found roughly 23% broken on average, rising toward 50% for older articles. Two distinct needs sit inside the problem: discovery (find where the authoritative copy lives now) and currency (confidence that the copy reached is authoritative and not a stale mirror or withdrawn draft).
The working solution stacks three layers. The Handle System, developed by Robert Kahn at CNRI with DARPA funding between 1992 and 1996 and first implemented in autumn 1994, is a general-purpose distributed name service: because the resolution target is stored in the handle record and can be changed by the registrant, the handle survives moves of the underlying object. A DOI is a handle with extra structure, metadata, and policy; its syntax is standardized as ANSI/NISO Z39.84-2005, and the system was initiated at the Frankfurt Book Fair in October 1997 and published as ISO 26324 on 23 April 2012. Crossref, founded by scholarly publishers in early 2000, is a registration agency: each member self-registers DOIs for its own content and deposits metadata.
Resolution is centralized through the Handle System and the doi.org proxy; registration is delegated to competing agencies (Crossref, DataCite, mEDRA, JaLC) and ultimately to publishers who self-register. Authority flows down from one standard; data entry and upkeep are pushed out to the parties who hold the documents.
What this means for water rate data
The DOI system fixes link rot by resolving a stable identifier to an ever-current URL that the document's owner maintains. A water registry can store an identifier per rate sheet rather than a bare link, so a utility's website redesign updates one pointer instead of breaking every downstream reference, and version metadata distinguishes the adopted schedule from superseded ones. Registration can be federated: a neutral body owns the resolver while a state board or the California Data Collaborative registers on behalf of the utilities that will not, so no central body has to crawl every website forever.
The problem
A bare URL names a location, and locations are mutable while the obligation to keep them stable sits with whoever happens to control the web server. The phenomenon has two faces publishing calls "link rot" (the URL stops resolving) and "reference rot" (the URL resolves but the content has changed). Both break the chain of evidence a citation is supposed to preserve. The solution was shaped by four tensions: persistence is a recurring operational obligation funded by charging registrants; resolution has to be globally consistent while registration scales to tens of thousands of publishers; an identifier guarantees the string resolves to a managed location but not that the document is correct; and the indirection only works if the registrant keeps the target URL current.
What failed first
The first approach was to cite the bare URL, which failed for the structural reason above. PURLs, introduced by OCLC in the mid-1990s, used an HTTP redirect service so a stable PURL could be repointed, but inherited the web's single-operator fragility and never acquired the cross-publisher governance and metadata layer scholarly citation needed. The URN effort at the IETF defined a syntax for location-independent names but left resolution largely unspecified, so URNs did not become clickable links. The deeper lesson, recorded in the DOI origin story, is that the fix could not be purely technical: the 1996 Association of American Publishers initiative "recognised the need to uniquely and unambiguously identify content entities, rather than refer to them by locations," and built "a social infrastructure of policies and formal agreements" alongside the technology.
The mechanism that worked
The Handle System, developed by Robert Kahn at CNRI with DARPA funding between 1992 and 1996, first implemented in autumn 1994, is a general-purpose distributed name service specified in IETF RFCs 3650, 3651, and 3652. A handle has a prefix naming a naming authority and a suffix giving the local name; a client passes the prefix to the Global Handle Registry, which returns the location of the responsible Local Handle Service. A DOI is a handle with extra structure: in Paskin's words, "the DOI System is one implementation of the Handle System; hence a DOI name is a Handle." The syntax (ANSI/NISO Z39.84-2005) is a prefix beginning with the directory indicator "10" plus a registrant code, and a suffix naming the object (10.1000/123456). The system was initiated at the Frankfurt Book Fair in October 1997 with the founding of the IDF, launched first applications in 2000, and was published as ISO 26324 on 23 April 2012. Crossref, founded in early 2000 as the Publishers International Linking Association, is a registration agency: each member self-registers DOIs and deposits metadata including the reference list. Content negotiation is the discovery primitive: "you don't need to know where a DOI is registered in order to retrieve its associated metadata." The roles separate cleanly: the IDF owns the "10" namespace, registration agencies get delegated authority, publishers self-register and keep their targets current, and the doi.org proxy plus Global Handle Registry provide one resolution path. The system distinguishes "the same object that moved" (handled silently by indirection) from "a new version" (handled in metadata with typed relationships such as IsNewVersionOf), and a retracted object's DOI is repointed to a tombstone page because "DOIs cannot be deleted."
Outcome
The mechanism became the backbone of scholarly discovery. The DOI Foundation factsheet reports on the order of 275 million DOI names assigned, over 155,000 distinct prefixes, and roughly one billion resolutions per month; Wikipedia puts cumulative DOIs at 391 million by 2025. Crossref alone connects on the order of 150 million metadata records from more than 160 countries. APA and MLA now instruct authors to cite the DOI in place of a URL, moving the persistence guarantee from the citing author to the infrastructure. The pattern generalized: DataCite extended DOIs to research datasets, ORCID gives every researcher a persistent identifier, and ROR (launched January 2019, jointly run by the California Digital Library, Crossref, and DataCite) does the same for institutions. The transferable core is the indirection plus the federation: a stable name that always resolves to the current authoritative location, maintained not by a central crawler but by the parties who hold the documents.
Sources
Legal Entity Identifier
A federated entity identifier, an open reference dataset, certified crosswalks to legacy schemes, and mandate-driven adoption.
When Lehman Brothers failed in September 2008, regulators and counterparties could not answer a basic question: how much exposure does my firm, or the market, have to this one entity? Lehman traded through a sprawl of legal entities, and every bank, custodian, and regulator recorded those entities under its own internal identifiers. The same legal entity appeared under inconsistent codes across systems, so exposures could not be summed. The Financial Stability Board's 2012 report to the G20 (8 June 2012) identified this directly as a gap exposed by the crisis. The diagnosis pointed to discoverability and crosswalk, not data scarcity: the exposure data existed inside hundreds of institutions, but no shared way existed to know that one dataset's record was the same legal entity as another's.
The solution combined a single technical standard, a federated three-tier governance structure, an openly published reference dataset, certified crosswalks, and a renewal regime. The LEI is a 20-character alphanumeric code defined by ISO 17442. Governance runs in three tiers: the Regulatory Oversight Committee (convened January 2013) governs in the public interest; the Global Legal Entity Identifier Foundation (GLEIF), a Swiss not-for-profit, held its inaugural board meeting in Zurich on 26 June 2014 and accredits issuers and publishes the database; and federated Local Operating Units register entities and verify their data against local sources. One global format, locally issued.
GLEIF publishes the entire dataset openly under a CC0 license: searchable free of charge, downloadable as the Golden Copy issued three times daily, and queryable through an API with fuzzy matching. As of 25 June 2026 the API reported 3,352,668 LEIs. A formal Certification of LEI Mapping process produces certified, openly published crosswalk files, including a BIC-to-LEI relationship file introduced with SWIFT on 8 February 2018 as the first open-source mapping.
What this means for water rate data
GLEIF is the model for the identity and crosswalk layer: a federated registry issues one identifier per entity, publishes the reference data openly, and certifies mappings to older identifier systems. For water, that is an open, certified — and eventually signed — crosswalk from the PWSID to DWR and local identifiers, so an extractor resolves "the same utility under different identifiers" against an authoritative source instead of guessing from names. Annual re-validation and a public challenge process keep the records current, and the verifiable LEI (vLEI, ISO 17442-3, 2024) shows that signing is an increment on an existing registry, not a separate system.
The problem
The 2008 crisis exposed the inability to reliably identify counterparties, which prevented quick aggregation of exposures to failing institutions. At the Cannes Summit in November 2011 the G20 called on the FSB to develop a global Legal Entity Identifier "to uniquely identify a party in any and all financial transactions," and at Los Cabos in June 2012 it endorsed the resulting recommendations. The design had to hold several tensions: a single worldwide scheme that no central body could register everywhere; a forcing function for adoption alongside a free public good on the consumption side; periodic re-verification against data decay; and open publication funded by paid registration.
What failed first
Before the LEI, entity identity lived in proprietary, domain-specific silos that did not crosswalk cleanly. D-U-N-S numbers identified businesses for commercial credit but were proprietary and licensed. BICs identified institutions for payment messaging, not the full population of legal counterparties. CUSIP and ISIN identified securities, not the legal entities that issued them. Beneath all of these sat each firm's internal counterparty IDs, which never reconciled across firms. There was no authoritative, signed mapping between a firm's internal "Lehman" record, its BIC, its D-U-N-S, and a regulator's record. Aggregating exposure meant manual, error-prone matching on names and addresses across incompatible systems, which the crisis showed could not be done fast enough when it mattered.
The mechanism that worked
The LEI is a 20-character alphanumeric code defined by ISO 17442. GLEIF states four foundational principles: a global standard, a single unique identifier per legal entity, supported by high-quality data, and a public good available free of charge. Each LEI connects to two layers of reference data: Level 1, "who is who," and Level 2, "who owns whom." Governance runs in three tiers set out by the FSB in 2012 and implemented by 2014: the Regulatory Oversight Committee (convened January 2013, now spanning more than 50 countries) at the top; GLEIF in the middle, established as a Swiss not-for-profit with its inaugural board meeting in Zurich on 26 June 2014, running the Central Operating Unit; and federated Local Operating Units at the base (28 pre-LOUs at founding) that register entities and verify their data against local sources. GLEIF publishes the entire dataset openly under CC0, searchable free of charge, downloadable as the Golden Copy three times daily, and queryable through the GLEIF API; as of 25 June 2026 the API reported 3,352,668 LEIs. A formal Certification of LEI Mapping process produces certified, openly published relationship files: BIC-to-LEI (with SWIFT, 8 February 2018, the first open-source mapping, CSV, monthly, free), ISIN-to-LEI (with ANNA, from April 2019), and others. To fight data decay, the ROC mandates annual re-validation against third-party sources; an entity that misses renewal is marked "lapsed"; and the Challenge Management facility lets anyone with proof file a data challenge.
Outcome
The Global LEI System moved from G20 concept (2011) to operating institution (GLEIF, 2014) to mass infrastructure, holding roughly 3.35 million LEIs as of 25 June 2026, published openly under CC0. Uptake was driven primarily by regulatory mandate. In the US, CFTC swap-data reporting rules (17 CFR Part 45) require an LEI conforming to ISO 17442; in the EU, EMIR requires LEIs for derivatives reports and MiFID II/MiFIR enforced a "No LEI, No Trade" rule from 3 January 2018. Tying market access to possession of the identifier converted a public-good registry into something firms had a direct reason to join. GLEIF has since standardized a cryptographically verifiable form, the verifiable LEI (vLEI), defined by ISO 17442-3 (2024) and built on the KERI protocol; it is standardized and in early production, and adoption volumes remain small relative to the conventional LEIs and should not be overstated.
Sources
Non-Profit IRS Reporting
A disclosure mandate, an identifier, and open release of the data; government supplies coverage and authority while a third party supplies findability.
The United States recognizes more than 1.8 million tax-exempt organizations under section 501(c). Each is legally independent and answers to no common registrar of records. A donor, journalist, academic, or regulator faces the same question: where is the authoritative financial disclosure for this organization, and how do I know I am looking at the current one? For most of the twentieth century, "public" meant a paper file an organization showed on request or a microfilm reel at an IRS reading room. The disclosure existed; discoverability did not. The instructive part is that the federal government solved the discoverability problem without building the discovery tool.
The working mechanism is a four-part stack. The mandate: organizations exempt under 501(c) must file an annual return in the Form 990 series, and section 6104 makes those returns public; the smallest file only the 990-N "e-Postcard," created by the Pension Protection Act of 2006, which also added automatic revocation after three consecutive years of non-filing. The canonical identifier: every organization is keyed by its Employer Identification Number, the join key that links the Business Master File record, the stream of 990 filings, and every third-party database.
The machine-readable turn came from litigation. Carl Malamud's Public.Resource.Org sought e-filed 990s in the format the IRS already held; in Public.Resource.Org v. IRS (N.D. Cal. 2015), Judge William H. Orrick ruled on January 29, 2015 that the IRS produce the records in machine-readable form. On June 15, 2016 the IRS began publishing e-filed 990s as XML, over one million filings, in an Amazon S3 bucket. The Taxpayer First Act, signed July 1, 2019, extended mandatory e-filing to all Form 990-series organizations and directed machine-readable public availability. ProPublica's Nonprofit Explorer then built the free, fast lookup the public uses.
What this means for water rate data
Form 990 shows the full lever: government mandates disclosure, assigns an identifier, opens the data, and third parties build the discovery tools the agency never builds. The water version rides the canonical pointer on a filing utilities already submit, the eAR, and lets whatwatercosts.org be the discovery layer ProPublica became for nonprofits. It also names the failure modes to plan for — currency lag between updates, thin filings from the smallest systems (the Form 990-N analogue), and self-reported links that need validation.
The problem
A nonprofit's annual return, the Form 990, was a public document by law, but for most of the twentieth century "public" meant a paper file shown on request or a microfilm reel at an IRS reading room. There was no single front door, no canonical pointer from an organization's name to its latest financial filing, and no way to ask the question of the whole sector at once. Four tensions shaped the case: a universal filing requirement makes a registry complete but a hospital system and a volunteer rescue group cannot carry the same paperwork; the IRS held many returns as structured data yet its public-release process converted them to images; the IRS is a tax collector, not a product company; and a 990 describes a fiscal year that has already closed, so "current" is structurally a year or more stale.
What failed first
Before 2016, every approach to making 990s findable hit the same wall: the IRS released returns as images. The legal duty was old and real. Internal Revenue Code section 6104 requires an exempt organization to make its returns available for public inspection and keep each available for three years, and an organization can satisfy the copy requirement by making the return "widely available" on the Internet. But this duty sat on each organization individually and produced no central index. GuideStar, founded in 1994 by Arthur "Buzz" Schmidt, tried to centralize this: from 1997 it posted information on all 501(c)(3) organizations in the IRS Business Master File, holding records on more than 600,000 nonprofits by that December, and from October 1999 it posted the 990 and 990-EZ returns themselves. But GuideStar obtained 990 images and digitized them largely by hand or with OCR, a private parsing burden that pushed structured data behind a paywall. The deeper failure was at the IRS: even when an organization e-filed, so the agency held the return as structured XML, the public-release process flattened it into an image.
The mechanism that worked
The stack is a disclosure mandate, a canonical identifier, a court-forced machine-readable release later codified by statute, and a third-party discovery layer. Organizations exempt under 501(c) must file an annual return in the Form 990 series; the tiering matters, with larger organizations filing the full 990, smaller ones the 990-EZ, private foundations the 990-PF, and the smallest (gross receipts normally $50,000 or less) only the Form 990-N "e-Postcard," created by Section 1223 of the Pension Protection Act of 2006, which also added automatic revocation after three consecutive years of non-filing. Every organization is keyed by its Employer Identification Number, the join key linking the Business Master File record, the 990 filings, and every third-party database. The machine-readable turn came from litigation: Public.Resource.Org filed a FOIA request for the e-filed 990s of nine organizations in the format the IRS already held, and in Public.Resource.Org v. IRS, 2015 WL 393736 (N.D. Cal. 2015), Judge William H. Orrick ruled on January 29, 2015 that the IRS produce the records in machine-readable form within 60 days, holding that financial distress "does not excuse an agency's duty to comply with the FOIA." On June 15, 2016 the IRS began publishing e-filed 990s as XML in the irs-form-990 Amazon S3 bucket, over one million filings from 2011 forward. The Taxpayer First Act, signed July 1, 2019, extended mandatory e-filing to all Form 990-series organizations and directed that the IRS make the information available "in a machine-readable format as soon as possible." ProPublica's Nonprofit Explorer ingests the raw filing data and exposes a free, fast lookup by organization name, person, or full text, with an open API, across over 1.8 million organizations; Candid (the February 5, 2019 GuideStar–Foundation Center merger) runs the commercial layer on the same disclosures.
Outcome
A donor, journalist, or researcher can now start from an organization's name, resolve it to an EIN, and reach its most recent machine-readable 990 in seconds, for free, with full-text and cross-organization search, a capability that did not exist before 2016. The sector became analyzable in bulk, and the open release commoditized what had been a proprietary moat, letting a free public-interest tool stand alongside a paid commercial one on the same data. The limits are real: currency is structurally lagged; the smallest organizations file only the 990-N, so the registry is complete in coverage but thin at the bottom; data quality varies because 990s are self-reported; and pre-universal-e-filing paper returns remain image-only. None of these undo the core achievement: the disclosure is mandated, the identifier is canonical, the data is open, and the discovery layer exists.
Sources
About
What this is
This document presents seven case studies in making canonical, current records discoverable across many independent publishers. It is the discoverability companion to the earlier five-case set on data standardization (GTFS, FHIR, UK Open Banking, Green Button, Australia BOM), and was prepared for the Protocolized / Water Protocol effort on California water-rate data. The earlier set examined how a sector converged on a shared data standard; this set examines how other domains made records reachable and verifiable before large language models existed. Each case was selected because its mechanism transfers to the water reachability and currency problem: a mandated repository (SEC EDGAR), a regulatory filing (utility tariff filing), metadata harvesting (OAI-PMH), web self-description, a persistent identifier (DOI, Handle, Crossref), a federated entity identifier (LEI, GLEIF), and open disclosure (IRS Form 990).
How the cases were researched
Each case was verified first against primary sources: sec.gov, ferc.gov, cpuc.ca.gov, gleif.org, openarchives.org, doi.org, irs.gov, the relevant Federal Register entries, the court rulings, and the ISO and RFC records. Source citations are listed at the end of each full case study. Several primary government pages (FERC, parts of sec.gov, and iso.org) returned HTTP 403 to automated fetching; in those instances the load-bearing facts were taken from search-index copies, cross-checked, and flagged in the relevant sources file. A few point-in-time statistics, including DOI counts, LEI lapse rates, and llms.txt adoption figures, vary across sources and are marked as approximate. The eAR Section 8 recommendation requires a check against the current eAR schema before it enters a client-facing deliverable. The cases were prepared in June 2026.
About the Protocol Institute
The Protocol Institute is an independent research organization studying protocols — the rules and coordination structures that shape interaction across diplomacy, software, medicine, governance, and beyond. Evolved from the Ethereum Foundation-funded Summer of Protocols program (2023–2025), it continues that work through research, publishing, and community building across organizational theory, infrastructure studies, and governance design.
Protocolized is its flagship publication; the AI Capability Maturity Model is one of its practitioner-facing frameworks, produced by the Protocols for Business leads.