Cross-Sector Evidence · Discoverability

Making public interest data discoverable

This lookup doesn’t exist yet. Securities, energy, research, finance, and nonprofits each have one. Seven precedents show how to build water’s — by standardizing the pointer, not the data.

Already have it:EDGARFERCCrossrefGLEIFIRS 990
The seven precedents

Pick a case study to explore

Summary

Emerging technologies such as LLMs shift data policy priorities from standardization to discoverability, and then verifiability

  • The barrier is discoverability, not data format. Tools read any rate sheet; they cannot reliably find each utility's current adopted sheet or confirm it is the version a board adopted.
  • Build a registry of pointers, not a data standard. Map each utility to the canonical URL of its current adopted rate schedule, keyed to the PWSID (Public Water System Identification number), with the adoption recorded as data — adoption date, board action, effective date.
  • Populate it through a filing that already exists. Add one field, the adopted rate-sheet URL, to California's Electronic Annual Report (eAR), which every public water system already files; this inherits statewide coverage at near-zero cost.
  • Publish it openly; let an aggregator build discovery. The state or the data collaborative supplies coverage and authority; an aggregator such as whatwatercosts.org supplies the public search layer.
  • Sequence the rollout: voluntary first, mandate second. Start with a small self-description file (an llms.txt for utilities) that cooperative agencies publish and the aggregator harvests, then pursue the eAR field for full coverage.
  • Next: move from discovery to verifiability. Once the registry makes rate sheets discoverable, the next step confirms each pointer states the rates a board actually adopted and still in force — reconciling self-reported links against the board action of record, certifying identity through the signed crosswalk, and confirming the eAR field against the current schema before it carries weight.
Findings

A step by step guide to addressing water data discoverability

1. Anchor the registry on the PWSID

SEC EDGAR resolves a company name or ticker to a Central Index Key (CIK); the Internal Revenue Service joins each nonprofit's filings through its Employer Identification Number (EIN); the Legal Entity Identifier (LEI) does the same for financial firms. Water already holds the equivalent key in the PWSID, which persists through name changes and separates utilities with similar names. Build the registry on the PWSID and resolve names and aliases to it.

2. The registry points rather than hosts

EDGAR keeps a text document as the official filing and layers tagged data — XBRL, the eXtensible Business Reporting Language — on top, marked unofficial. The library harvesting protocol exposes a short record plus a pointer to where each document lives. The water asset is a table mapping PWSID to utility to the canonical rate-sheet URL. It need not hold or reformat the rate sheets, because the extractor reads them.

3. The registry records adoption status as a field

The verification problem, adopted rates against a draft study, is answered with metadata in every case that faced it. EDGAR stamps each submission with a filing type and an acceptance time and links amendments to their originals. Energy regulators derive the currently effective tariff from filed effective dates and supersession. The web's rel=canonical tag lets a publisher name the authoritative version among duplicates. The registry should record an adoption event on each pointer, the adoption date, the board action, and the effective date, so a field rather than a guess from the document text answers whether a rate sheet is in force.

4. An existing mandatory filing populates the registry at low cost

EDGAR, energy tariff filing, Form 990, and the LEI each rode an obligation that already existed: EDGAR moved a standing disclosure duty to one electronic channel, the Form 990 reform changed only the format and openness of a return nonprofits already filed, and LEI adoption followed reporting rules that required the identifier. California's counterpart is the eAR. Every public water system already files it, keyed to the PWSID, and its Section 8 collects rate-structure data but no link to the adopted schedule. Adding one field, the URL of the current adopted rate schedule with its adoption and effective dates, would turn an existing filing into the registry and cover every public water system in the state. This warrants confirmation against the current eAR schema before it reaches a client.

5. Where no authority covers every publisher, a low-friction convention spreads

No body holds authority over every U.S. water utility, just as none governed every research repository or every website. The library harvesting protocol (OAI-PMH, the Open Archives Initiative Protocol for Metadata Harvesting) and the web's self-description files spread because each asked little of the publisher. The water analogue is a small file at a fixed path on each utility's site, an llms.txt for utilities, that names the canonical rate-sheet URL, the adoption date, the rate structure, and local context. A utility webmaster can add it in one sitting, and an aggregator harvests and reconciles the rest.

6. Publish the registry openly and let an aggregator build discovery

The SEC publishes a free, CIK-keyed data feed and builds no consumer product; the Global Legal Entity Identifier Foundation (GLEIF) publishes its full dataset under a public-domain (CC0) license; the IRS released machine-readable returns, and ProPublica, not the IRS, built the search tool. An open PWSID-to-URL-to-adoption-metadata table, in CSV and JSON, lets whatwatercosts.org occupy that service-provider role while the state or the data collaborative supplies coverage and authority.

7. Adoption follows a dominant consumer or a mandate

The web conventions spread because Google rewarded publication with search visibility, the dynamic that also drove transit-schedule adoption in the earlier case set. The LEI spread under "no LEI, no trade" reporting rules. A voluntary water file will see uneven uptake until a widely used benchmarking or affordability tool makes accurate appearance depend on it. The identity layer follows GLEIF: a neutral body publishes an open, certified crosswalk from the PWSID to Department of Water Resources (DWR) and local identifiers, on the model of GLEIF's published mapping between the LEI and the bank Business Identifier Code (BIC), and signs it once the identifier and crosswalk exist.

Design implications

  1. Define the llms.txt-for-utilities convention (canonical URL, adoption date, rate structure, context) and have whatwatercosts.org harvest it, recruiting utilities already in the project's contact set, such as Las Virgenes, Moulton Niguel, and Eastern Municipal, as initial adopters.
  2. Advocate adding the adopted-rate-sheet-URL field to the State Water Board's eAR, which inherits statewide coverage from an existing mandate.
  3. Publish a PWSID-to-DWR-to-local crosswalk on the GLEIF model, certified and later signed, drawing the common grammar from the Moulton Niguel-led CaDC definitions whitepaper.
  4. Identify the consumer that rewards publication, and tie accurate appearance in a benchmarking or affordability tool to publishing the file.

For the nine investor-owned Class A water utilities that the California Public Utilities Commission (CPUC) regulates, serving about 16% of Californians, the registry can pull entries directly from the commission's advice-letter and rate-case records by utility number. The eAR field then covers the publicly owned systems no commission regulates, about 84% of roughly 50,000 community water systems nationally.

Limitations

  • Voluntary participation produces incomplete coverage. Utilities without an incentive publish nothing, and the system must fall back to ordinary scraping for them.
  • A self-asserted pointer can go stale or point to the wrong document. A published "current adopted" marker states the utility's claim, and a high-stakes use should reconcile it against the board action of record; self-reported URL fields require link validation.
  • Persistence carries a standing operational cost. Keeping each pointer live is continuing work, which the scholarly identifier system funds through registrant fees and which a water registry must fund some other way.
  • The smallest utilities report last and report least, as the smallest nonprofits do under the abbreviated Form 990-N. Coverage reaches completeness before it reaches uniform depth, so the design should phase in by size and keep a manual fallback.
  • An identifier establishes reachability, not correctness. It locates the document; confidence in the content rests on the registrant's credibility and the adoption metadata.
Case 01 · US Securities Disclosure
Mandated Repository United States · 1993–present

SEC Electronic Data Gathering, Analysis, and Retrieval

A federal mandate routed every issuer's disclosures into one electronic repository, keyed to a stable identifier, with public access kept free.

Before EDGAR, a public company's disclosures existed as paper. A registrant printed its 10-K, prospectus, or proxy and mailed copies to the SEC, which kept them in public reference rooms. Anyone who wanted a current filing traveled to a reference room and stood at a photocopier, or paid a microfiche vendor such as Disclosure, Inc. The friction was not the document format; a 10-K was already a structured, regulated document. The friction was reachability. There was no single place where a reader could be confident of looking at a company's current, authoritative filing.

What turned EDGAR from a 1984 pilot into a discoverability protocol was a combination of five elements: a legal mandate, a single canonical repository, a stable identifier, an unambiguous notion of the authoritative filing, and free machine-reachable dissemination. The SEC adopted Regulation S-T in February 1993 and phased in mandatory electronic filing over roughly three years, completing on May 6, 1996. The graduated schedule and a hardship exemption resolved the mandate-versus-burden tension: coverage became universal while small filers were eased in.

Each filer receives a Central Index Key, the join key that lets a reader resolve a company name or ticker to a number and key all of that entity's filing history off it. The CIK persists across name changes. EDGAR stamps each submission with a filing type and acceptance datetime and threads amendments to originals, so the operative version is answered by metadata rather than guesswork. The SEC opened a free public web site in September 1995, and Congress later made free public access a statutory floor.

Mechanism
Federal mandate + single repository + free feed
Mandate
Regulation S-T, 1993; phase-in to May 6, 1996
Identifier
Central Index Key (CIK)
Structured layer
XBRL required from 2009 (Release 33-9002), furnished as unofficial

What this means for water rate data

EDGAR is the template for the core asset: a registry that resolves an identifier to a publisher's current filing, with currency carried as a stamped field rather than inferred from the document. The PWSID plays the CIK's role as the stable key. EDGAR also shows the discipline that made it useful — keep the canonical layer thin and the data feed free, and leave analysis to a third-party ecosystem, the role whatwatercosts.org would occupy.

Case 02 · US Energy Utilities
Regulatory Filing United States · 2008–present

Utility Tariff Filing

Regulated utilities file their rates with a regulator, and the filed copy becomes the only legally operative one, which makes the regulator's system the canonical lookup.

Anyone who needs to know what an electric or gas utility currently charges faces a discoverability question before any data question. The price schedule exists somewhere, but finding the authoritative, current copy among a utility's own documents, press releases, and superseded versions is the hard part. Regulated energy utilities solved this structurally rather than technologically: the utility's prices are not merely published, they are filed with a regulator, and the regulator's copy becomes the single authoritative version. Under the filed rate doctrine, a utility may charge only the rate it has filed, and the act of filing, not any later approval, is what gives the rate legal force. As one court put it, "it is the filing of the tariffs, and not any affirmative approval or scrutiny by the agency, that triggers the filed rate doctrine."

FERC closed the discoverability gap with Order No. 714, "Electronic Tariff Filings" (Docket No. RM01-5-000), issued September 19, 2008 and published at 73 FR 57,515. The Final Rule required that all tariffs, tariff revisions, and rate-change applications for public utilities, natural gas pipelines, oil pipelines, and the federal power administrations be filed electronically. Mandatory electronic filing phased in beginning April 1, 2010 on a six-month staggered schedule, after which FERC no longer accepted paper tariff filings. The electronic format followed standards developed with the North American Energy Standards Board (NAESB), so filings from every utility shared a common structure.

eTariff filings are organized as discrete, addressable tariff records, each carrying an effective date and superseding a prior record, so the "currently effective" version of any record becomes a query rather than a manual reconstruction. FERC's eLibrary holds the filings and the public eTariff Viewer renders them. Filings are keyed to the utility and to docket numbers; California's CPUC mirrors this with a Utility Reference Number (U#) and numbered advice letters filed under General Order 96-B.

Mechanism
Mandatory filing to a regulator; filed copy is operative
Order
FERC Order No. 714, 2008; e-filing from April 1, 2010
Identifiers
FERC docket; CPUC U-number; advice letters
Coverage limit
CPUC water IOUs serve ~16% of Californians; ~84% of systems are publicly owned

What this means for water rate data

Energy utilities are discoverable because law requires them to file their rates with a regulator, whose system becomes the authoritative lookup. Most water utilities have no such regulator: the nine investor-owned utilities the CPUC oversees serve about 16% of Californians, while roughly 84% of the country's 50,000 community water systems are publicly owned and set rates by a local board vote under Proposition 218, with no central filing. The transferable move is to manufacture the missing chokepoint cheaply, by adding a rate-sheet-URL field to the eAR, and to treat the adopting board resolution and its effective date as the water equivalent of the "currently effective tariff."

Case 03 · Scholarly Repositories
Metadata Harvesting International · 1999–present

Open Archives Initiative Protocol

Data providers expose minimal metadata and a pointer; service providers harvest and aggregate, a voluntary federated model that asks little of each publisher.

By the late 1990s, scholarly output was scattering. Physicists posted preprints to arXiv, computer scientists used NCSTRL, economists used RePEc, and individual universities were standing up their own repositories. Each archive held real, current documents; none knew about each other. A reader had to know which archive held a paper, navigate to it, and search it on its own terms. The same shape of problem had appeared a generation earlier in libraries, where a book sat in thousands of separate institutions, each with its own card catalog. In both cases the obstacle was discovery across many independent holders, not the format of the underlying object.

Two complementary mechanisms carried discovery across independent holders. OCLC's WorldCat centralized the record: one library catalogs a work once to the MARC standard, and every other holder attaches a lightweight holding symbol to the existing record rather than re-describing the book. The cost of that model is its centralization; it works because participants agreed to join one cooperative. The Open Archives Initiative federated the index instead. Its harvesting protocol, OAI-PMH, grew from the Santa Fe meeting of October 21–22, 1999; version 2.0, the stable version still in use, was released on June 14, 2002 with Carl Lagoze and Herbert Van de Sompel as editors.

OAI-PMH is deliberately small: six verbs over plain HTTP, returning UTF-8 XML, with a mandatory floor of unqualified Dublin Core (15 flat elements). The data-provider / service-provider split is the heart of the design. A data provider exposes its own holdings, lightly described, with a pointer to where the full object lives; a service provider harvests across many providers, normalizes, and builds the discovery experience. Datestamps and selective harvesting answer the currency question: a service provider re-harvests only what changed.

Mechanism
Voluntary metadata harvesting; data/service provider split
Protocol
OAI-PMH v2.0, 2002-06-14; six verbs over HTTP
Predecessor
Z39.50 (1988 onward); union catalog (OCLC, 1967/1971)
Metadata floor
Unqualified Dublin Core (15 elements), plus a pointer

What this means for water rate data

OAI-PMH is the voluntary path: thousands of independent repositories each expose a short record plus a pointer, and aggregators harvest them with no shared platform and no mandate. For water, the utility is the data provider — exposing its PWSID, adoption date, and rate-sheet URL — and whatwatercosts.org is the service provider that harvests and deduplicates by PWSID. The protocol's datestamps, which let a harvester detect what changed without re-reading everything, are the precedent for the adoption-date field that signals currency.

Case 04 · The Open Web
Self-Description International · 1994–present

The Web Self-Description Stack

Small files at predictable locations declare canonical URLs, freshness, and authority; adoption followed a dominant consumer rather than a mandate.

By the mid-1990s the web had a discovery problem that looked nothing like a data-format problem. Millions of independent sites published content that a crawler had to find, fetch without overloading the server, and recognize as current. No central registry listed what any site contained or when it last changed. A search engine confronting a new domain had three questions and no agreed way to ask them: which parts may I fetch, what are the canonical pages, and which changed recently enough to re-index. This is the same shape as the water-rate discovery problem. The web solved its version with a stack of small, voluntary self-description conventions at predictable locations, the direct ancestor of the llms.txt idea.

The stack is layered, each part a small file or tag at a predictable location solving one question. robots.txt (Martijn Koster, February 1994) established the pattern: a single plain-text file at a fixed path that any crawler checks without being told, later formalized as RFC 9309 in September 2022. Google introduced the Sitemaps protocol in June 2005, an XML file enumerating canonical URLs with an optional lastmod date; in November 2006 Google, Yahoo!, and Microsoft jointly announced support and stood up sitemaps.org. RSS and Atom (RFC 4287, December 2005) broadcast new and updated content, and WebSub (a W3C Recommendation on 23 January 2018) turned the freshness signal into a publisher push.

The remaining question, which of several URLs returning similar content is authoritative, was answered by the canonical link element, announced jointly by Google, Yahoo, and Microsoft on 12 February 2009 and formalized as RFC 6596 in April 2012. None of these spread by mandate. Each reached ubiquity because a dominant consumer made adoption pay, the same dynamic that drove transit agencies to publish GTFS once Google Maps consumed it.

Mechanism
Voluntary files at predictable paths; consumer-driven adoption
Core conventions
robots.txt (1994), Sitemaps (2005), RSS/Atom, rel=canonical (2009)
Standards
RFC 9309 (2022); RFC 4287 (2005); WebSub W3C Rec, 2018
Limit
Advisory only; signals self-asserted, not authority guarantees

What this means for water rate data

The web's self-description files are the working model for an llms.txt for utilities: name the canonical rate-sheet URL (the sitemap pattern), carry a last-adopted date as the freshness signal, mark which document is current versus draft (the rel=canonical pattern), and let a dated changelog stand in for change notification. The case's harder lesson is adoption — these conventions spread only once a dominant consumer rewarded them, so a water file needs an equivalent consumer, most plausibly a benchmarking or affordability tool utilities want to appear in correctly, rather than elegance.

Case 05 · Scholarly Publishing
Persistent Identifier International · 1997–present

Digital Object Identifier

A persistent identifier, a central resolver, and federated registration, so the identifier survives URL change and resolves to the current authoritative copy.

By the late 1990s scholarly publishing had moved online and citations began to break. A reference pointed at a publisher's URL, a location rather than a name; when a publisher redesigned its site or sold a journal, the location changed and the citation died. Norman Paskin, founding director of the International DOI Foundation, framed the failing directly: a stable identifier should let you "locate an entity, or provide services irrespective of changes in location or management responsibility of the entity." The damage is measurable: one 2022 study of more than 50,000 cited URLs found roughly 23% broken on average, rising toward 50% for older articles. Two distinct needs sit inside the problem: discovery (find where the authoritative copy lives now) and currency (confidence that the copy reached is authoritative and not a stale mirror or withdrawn draft).

The working solution stacks three layers. The Handle System, developed by Robert Kahn at CNRI with DARPA funding between 1992 and 1996 and first implemented in autumn 1994, is a general-purpose distributed name service: because the resolution target is stored in the handle record and can be changed by the registrant, the handle survives moves of the underlying object. A DOI is a handle with extra structure, metadata, and policy; its syntax is standardized as ANSI/NISO Z39.84-2005, and the system was initiated at the Frankfurt Book Fair in October 1997 and published as ISO 26324 on 23 April 2012. Crossref, founded by scholarly publishers in early 2000, is a registration agency: each member self-registers DOIs for its own content and deposits metadata.

Resolution is centralized through the Handle System and the doi.org proxy; registration is delegated to competing agencies (Crossref, DataCite, mEDRA, JaLC) and ultimately to publishers who self-register. Authority flows down from one standard; data entry and upkeep are pushed out to the parties who hold the documents.

Mechanism
Persistent identifier + central resolver + federated registration
Standards
Handle (RFC 3650–3652); DOI syntax Z39.84-2005; ISO 26324 (2012)
Resolver
doi.org proxy + Global Handle Registry
Versioning
DataCite IsNewVersionOf / IsPreviousVersionOf; tombstone pages

What this means for water rate data

The DOI system fixes link rot by resolving a stable identifier to an ever-current URL that the document's owner maintains. A water registry can store an identifier per rate sheet rather than a bare link, so a utility's website redesign updates one pointer instead of breaking every downstream reference, and version metadata distinguishes the adopted schedule from superseded ones. Registration can be federated: a neutral body owns the resolver while a state board or the California Data Collaborative registers on behalf of the utilities that will not, so no central body has to crawl every website forever.

Case 06 · Global Finance
Entity Identity International · 2012–present

Legal Entity Identifier

A federated entity identifier, an open reference dataset, certified crosswalks to legacy schemes, and mandate-driven adoption.

When Lehman Brothers failed in September 2008, regulators and counterparties could not answer a basic question: how much exposure does my firm, or the market, have to this one entity? Lehman traded through a sprawl of legal entities, and every bank, custodian, and regulator recorded those entities under its own internal identifiers. The same legal entity appeared under inconsistent codes across systems, so exposures could not be summed. The Financial Stability Board's 2012 report to the G20 (8 June 2012) identified this directly as a gap exposed by the crisis. The diagnosis pointed to discoverability and crosswalk, not data scarcity: the exposure data existed inside hundreds of institutions, but no shared way existed to know that one dataset's record was the same legal entity as another's.

The solution combined a single technical standard, a federated three-tier governance structure, an openly published reference dataset, certified crosswalks, and a renewal regime. The LEI is a 20-character alphanumeric code defined by ISO 17442. Governance runs in three tiers: the Regulatory Oversight Committee (convened January 2013) governs in the public interest; the Global Legal Entity Identifier Foundation (GLEIF), a Swiss not-for-profit, held its inaugural board meeting in Zurich on 26 June 2014 and accredits issuers and publishes the database; and federated Local Operating Units register entities and verify their data against local sources. One global format, locally issued.

GLEIF publishes the entire dataset openly under a CC0 license: searchable free of charge, downloadable as the Golden Copy issued three times daily, and queryable through an API with fuzzy matching. As of 25 June 2026 the API reported 3,352,668 LEIs. A formal Certification of LEI Mapping process produces certified, openly published crosswalk files, including a BIC-to-LEI relationship file introduced with SWIFT on 8 February 2018 as the first open-source mapping.

Mechanism
Federated identifier + open data + certified crosswalks + mandate
Standard
ISO 17442; 20-character code; GLEIF est. 2014
Open data
CC0 Golden Copy, three times daily; ~3.35M LEIs (June 2026)
Adoption lever
MiFID II "No LEI, No Trade" from 3 January 2018; CFTC 17 CFR Part 45

What this means for water rate data

GLEIF is the model for the identity and crosswalk layer: a federated registry issues one identifier per entity, publishes the reference data openly, and certifies mappings to older identifier systems. For water, that is an open, certified — and eventually signed — crosswalk from the PWSID to DWR and local identifiers, so an extractor resolves "the same utility under different identifiers" against an authoritative source instead of guessing from names. Annual re-validation and a public challenge process keep the records current, and the verifiable LEI (vLEI, ISO 17442-3, 2024) shows that signing is an increment on an existing registry, not a separate system.

Case 07 · US Nonprofits
Open Disclosure United States · 2016–present

Non-Profit IRS Reporting

A disclosure mandate, an identifier, and open release of the data; government supplies coverage and authority while a third party supplies findability.

The United States recognizes more than 1.8 million tax-exempt organizations under section 501(c). Each is legally independent and answers to no common registrar of records. A donor, journalist, academic, or regulator faces the same question: where is the authoritative financial disclosure for this organization, and how do I know I am looking at the current one? For most of the twentieth century, "public" meant a paper file an organization showed on request or a microfilm reel at an IRS reading room. The disclosure existed; discoverability did not. The instructive part is that the federal government solved the discoverability problem without building the discovery tool.

The working mechanism is a four-part stack. The mandate: organizations exempt under 501(c) must file an annual return in the Form 990 series, and section 6104 makes those returns public; the smallest file only the 990-N "e-Postcard," created by the Pension Protection Act of 2006, which also added automatic revocation after three consecutive years of non-filing. The canonical identifier: every organization is keyed by its Employer Identification Number, the join key that links the Business Master File record, the stream of 990 filings, and every third-party database.

The machine-readable turn came from litigation. Carl Malamud's Public.Resource.Org sought e-filed 990s in the format the IRS already held; in Public.Resource.Org v. IRS (N.D. Cal. 2015), Judge William H. Orrick ruled on January 29, 2015 that the IRS produce the records in machine-readable form. On June 15, 2016 the IRS began publishing e-filed 990s as XML, over one million filings, in an Amazon S3 bucket. The Taxpayer First Act, signed July 1, 2019, extended mandatory e-filing to all Form 990-series organizations and directed machine-readable public availability. ProPublica's Nonprofit Explorer then built the free, fast lookup the public uses.

Mechanism
Disclosure mandate + identifier + open data + third-party discovery
Identifier
Employer Identification Number (EIN)
Open release
Court order Jan 2015; bulk XML June 15, 2016; Taxpayer First Act 2019
Discovery layer
ProPublica Nonprofit Explorer; ~1.8M organizations

What this means for water rate data

Form 990 shows the full lever: government mandates disclosure, assigns an identifier, opens the data, and third parties build the discovery tools the agency never builds. The water version rides the canonical pointer on a filing utilities already submit, the eAR, and lets whatwatercosts.org be the discovery layer ProPublica became for nonprofits. It also names the failure modes to plan for — currency lag between updates, thin filings from the smallest systems (the Form 990-N analogue), and self-reported links that need validation.

About

About

What this is

This document presents seven case studies in making canonical, current records discoverable across many independent publishers. It is the discoverability companion to the earlier five-case set on data standardization (GTFS, FHIR, UK Open Banking, Green Button, Australia BOM), and was prepared for the Protocolized / Water Protocol effort on California water-rate data. The earlier set examined how a sector converged on a shared data standard; this set examines how other domains made records reachable and verifiable before large language models existed. Each case was selected because its mechanism transfers to the water reachability and currency problem: a mandated repository (SEC EDGAR), a regulatory filing (utility tariff filing), metadata harvesting (OAI-PMH), web self-description, a persistent identifier (DOI, Handle, Crossref), a federated entity identifier (LEI, GLEIF), and open disclosure (IRS Form 990).

How the cases were researched

Each case was verified first against primary sources: sec.gov, ferc.gov, cpuc.ca.gov, gleif.org, openarchives.org, doi.org, irs.gov, the relevant Federal Register entries, the court rulings, and the ISO and RFC records. Source citations are listed at the end of each full case study. Several primary government pages (FERC, parts of sec.gov, and iso.org) returned HTTP 403 to automated fetching; in those instances the load-bearing facts were taken from search-index copies, cross-checked, and flagged in the relevant sources file. A few point-in-time statistics, including DOI counts, LEI lapse rates, and llms.txt adoption figures, vary across sources and are marked as approximate. The eAR Section 8 recommendation requires a check against the current eAR schema before it enters a client-facing deliverable. The cases were prepared in June 2026.

About the Protocol Institute

The Protocol Institute is an independent research organization studying protocols — the rules and coordination structures that shape interaction across diplomacy, software, medicine, governance, and beyond. Evolved from the Ethereum Foundation-funded Summer of Protocols program (2023–2025), it continues that work through research, publishing, and community building across organizational theory, infrastructure studies, and governance design.

Protocolized is its flagship publication; the AI Capability Maturity Model is one of its practitioner-facing frameworks, produced by the Protocols for Business leads.

Contact

Reach us at team@protocol-institute.org.

More at protocol-institute.org.