Protocol Institute
Top ·Report summary 00Abstract 01Where rate data comes from 02Reading got cheap 03Case study · whatwatercosts 04Publish the address 05Beyond water rates AAppendix About

Protocol Institute · Protocol Vision · Water

It's time to discover, not file, water data

What if we didn't need to ask for another water data spreadsheet template, ever again?

For California water utility boards With the California Data Collaborative  ·  2026
What finding a rate should feel like
Long Beach Utilities DepartmentRates effective Oct 2025 $35.28 Verified
City of FresnoRates effective Jul 2026 $27.46 Verified
City of AnaheimRates effective Jul 2024 $43.65 Unverified
Real figures, read by machine from documents these utilities already publish. Monthly residential bill at 6,000 gallons. Two of the three can show you the document and the date a board adopted it. The third cannot, and its rates are two years old.
LLMs collapse the cost of data analysis
Now machine work
  • Locating a utility's rate schedule
  • Reading it in whatever format it was published
  • Normalizing it into a comparable structure
  • Computing a like-for-like benchmark bill
Still expert work
  • Cost-of-service analysis and rate design
  • Proposition 218 process and noticing
  • Preparing regulatory filings
  • Presenting to a board and to the public
Report summary
Getting water rate data is hard today, but it does not have to be.
1Challenge
Utilities spend too much staff time and consulting budget gathering and standardizing water rate data.

Every rate cycle a board is asked how its rates compare, and to whom. Answering means locating each peer's rate schedule, reading it, and converting it into a common form by hand, which is why the comparison arrives inside a paid rate study or a staff assignment running months.

>$100,000
Santa Clarita Valley Water Agency's 2024 rate study award. Peer comparison is one input.
2Example
Language models now read rate documents however they are published, making that work cheap.

whatwatercosts.org, assembled by one engineer from documents utilities already publish, matched hand-collected cost tables across 99 utilities. No utility filed anything, adopted a schema, or changed a format. The standardizing happens in the pipeline rather than at the agency.

<$0.05
Cost to read and standardize one utility's rates, within $1.98 of hand-collected cost tables. whatwatercosts.org
3Opportunity
What remains is finding the documents, so utilities need to publish a reliable link to the source.

Automated search located fewer than half the utilities on its own. A published link, with the date a board adopted the schedule and the date it took effect, takes one staff afternoon and requires no vendor, no procurement, and no change to any document.

Publish your rate data today →
Coverage rises from 47% to nearly all once the link is supplied. How to publish it · Email the CaDC
Figures as of August 2026 · Sources throughout

Water is a human right. It is also, in every California water agency's back office, a spreadsheet problem. The price of water is one of the most consequential numbers a household ever sees, a central pillar of operational planning, and too difficult to assemble correctly.

California water agencies need peer rate comparisons to set rates, inform customers, and defend increases, and the data comes at a steep cost. It has to be collected and managed by hand, which means that what should be a simple search query can cost six figures in consulting fees.1 California's water data movement has worked for a decade to streamline water reporting broadly and rate data specifically.

The programs that reached full coverage in other public sectors had either a regulator requiring every publisher to file or a large commercial buyer of the data. Water agencies have neither. The filing requirements that do exist collect hardcoded figures rather than the source policies and calculations that produce them, and no commercial player is large enough to force a change.

Large language models open a different path. Extraction and format conversion are now automatable even where the data is not standardized, and finding the document has become the constraint instead. This report examines that shift, presents whatwatercosts.org as evidence, and proposes three specific actions to make rate data discoverable and usable, none of which requires a change to what utilities publish or to state submission systems.

1Consulting costsRecent California rate study awards: Heber Public Utility District, $42,660, December 2025 (award report); Santa Clarita Valley Water Agency, $135,599, January 2024 (board packet). Both cover a full rate study rather than peer benchmarking alone.

Section 01

Where does rate data come from?

Every rate increase a California water utility proposes ends up in the same room. The board is seated, the Proposition 218 protest period has run,2 and somebody asks how the proposed rates compare to neighboring agencies and which agencies were chosen for the comparison. It is a fair question, it is asked in public, and answering it well is one of the more expensive things a utility does that produces no water.

Rate cycles run one to three years, and up to five, so the question comes around often enough to be a standing cost and rarely enough that nobody builds a permanent capability for it. The reliable way to answer it has been to prepare the data by hand, either through a consultant or by assigning a staff member to the work for several months. Neither approach produces anything reusable. The next cycle starts over.

The eAR collects a utility's tier structure but not the context behind it

Every public water system in California is required to file rate data; filing is mandatory and tied to funding eligibility. There are six requirements: the electronic Annual Report, the Urban Water Management Plan, the water loss audit, conservation reporting, and two further reports added under SB 606 and AB 1668.

Section 8A of the electronic Annual Report asks for a utility's most common residential rate structure.3 A water utility selects its structure type, its billing frequency, and its unit of measure. It then fills a grid recording the base rate, the upper bound of each consumption tier, and the cost per unit within it, once for single-family customers and again for multi-family.

The instructions ask for the rate that applies to most residential customers, at the most common meter size, and for an annual average where a rate varies through the year. The form does not include other customer classes, variation by meter size, or seasonal and drought adjustments. The reasoning behind an allocation is compressed into a free-text comment or dropped. The filing records what the rates are and does not record which document they came from, when a board adopted it, or when it took effect.

Meanwhile the source documents contain all of these components, because a utility writes them to run its own business rather than to answer a survey. Desk research suggests utilities are not strongly opposed to correcting the survey itself, and no party to it has been careless; the obstacle is the cost of modifying highly customized state data systems at systems-integrator rates.4

2Proposition 218California's constitutional requirement that property-related fee increases be noticed to parcel owners and be subject to a majority protest. Peer comparison is not required by the statute; it is what boards are asked for in practice.
3Section 8ACustomer Charges, in the electronic Annual Report filed by every public water system with the State Water Resources Control Board. The Board publishes guidance covering nine rate-structure types with worked examples. Guidance PDF · eAR portal
4Desk researchBased on published comment letters, agency materials, and the CaDC record rather than direct interviews. Scope to be stated before circulation.
What the state collects
eAR Section 8A, question A1.8 — Residential Rates & Charges Table, with the State Water Board's own worked example.
8. Customer Charges  ·  A1. Residential Water Rates and Charges
Base Rate (Fixed) + Usage Rate (Variable) monthly Hundred Cubic Feet Single & multi-family: Yes
Customer Class
& Billing Tiers
Base Rate Usage Rate Structure
Top Metric / UOM
Cost per Unit
of Measure
Single-family — Tier 12520.00
Tier 2do not fill45.00
Tier 3do not fill125.50
Tier 4do not fill188.00
Tier 5do not fill11.00
Multi-family — Tier 11045.00
Tier 2do not fill125.50
Tier 3do not fill188.00
Tier 4do not fill11.00
What stays in the rate schedule
The adopted document the utility wrote for its own operations, and which the filing does not point to.
Not carried by the filing
Rate schedule
The adopted document these figures were read from, stating the utility's charges in full. The filing records no URL, filename, or citation identifying it.
Adopting resolution
The board action that puts a schedule into force, carrying a number and a date. It is what distinguishes an adopted schedule from a draft or a rate study.
Effective date
The date a schedule begins to govern billing, together with the schedule it replaces. Without both, a document found today cannot be shown to be current.
Non-residential class
Commercial, industrial, institutional, and irrigation customers, each billed on a structure of its own that the residential table has no room to record.
Meter-size variation
The scaling of base charges by service connection size. The form asks for the rate at the system's most common meter size, so the remaining sizes are not recorded.
Seasonal variation
Rates that differ across the year, typically higher in summer. Where a rate is affected by season, the guidance directs the system to report an average.
Drought stage and pass-through
Conditional rates that take effect when a shortage stage is declared, along with surcharges recovering wholesale supply costs, both of which change what a customer is actually billed without appearing in the filing.
Allocation logic
The calculation assigning each household an efficient-use budget under a budget-based structure. The tiers are recorded; the derivation appears, at most, as a free-text comment.
Figure 1  The state's Section 8A grid, reproduced with the worked example the State Water Board publishes in its own reporting guidance. The filing captures the most common residential structure at the most common meter size, averaged if it varies through the year. It does not record which document produced those figures, when a board adopted it, or when it took effect, so a reader who wants the reasoning has to find the utility's rate schedule, and the filing does not say where that is.

Transit reached near-complete coverage through Google Maps and energy through FERC; water has neither lever

Several sectors have reached the coverage that water has not. Transit converged on GTFS, energy on tariff filing at the Federal Energy Regulatory Commission, and health records, UK banking, and Australian water resources each found their own route.5

Transit agencies published schedules in a common format because Google offered a free listing in Google Maps to any agency that did, and Google Maps was where their riders already were. The pull came from a commercial consumer large enough that joining was worth the work. The specification itself was deliberately kept as plain CSV so that any agency could edit it in a spreadsheet; it was criticized as technically old-fashioned and succeeded anyway.6 Energy tariffs became uniformly available because FERC required regulated utilities to file them, from April 2010, and the filed copy is the only legally operative one.7 The pressure came from a regulator with authority over every publisher.

California's water utilities sit outside both mechanisms. Publicly owned systems make up 54% of the roughly 49,700 community water systems nationally and serve 88% of the customers,8 and no commercial player derives enough value from their rate data to force the issue. A filing requirement does exist, and it collects a summary rather than pointing at the source. Continuing to imitate transit and energy, or other industries, is unlikely to produce a step change.

The 2018 rate standard didn't reach adoption because it required new technical capabilities

The Open Water Rate Specification, an award-winning initiative, was published in 2016 and won the state Water Data Challenge in 2018.9 It demonstrated a path towards a shared database of California water prices on the model of the energy sector's Utility Rate Database, and towards the analysis that becomes possible once rates are machine-readable: bill calculators, revenue and equity comparison across utilities, and affordability analysis joined to census data. Four analysis tools were built against it, and it supplied the 2017–2018 CA-NV AWWA rate survey.

Yet adoption lagged, because participating was too difficult. Submitting a rate structure meant working through a GitHub pull request, which is a technical capability most water agencies do not have and have no reason to acquire. The public record shows the pattern: forty rate structures were merged in March 2017 during a hackathon, the repository accumulated 500 specified utilities in total, and merged contributions stopped altogether in April 2018. Subsequent work on the submission interface did not move the number further. This case study suggests that adoption follows the cost of participating, not the quality of the specification.

7FERC Order 714Required electronic tariff filing by regulated energy utilities, phased in through 2010. Examined at length in the discoverability case set. Seven precedents
8Public system inventory26,918 of 49,672 community water systems are publicly owned, serving 88% of the population. EPA Safe Drinking Water Information System. SDWIS inventory
9OWRSOpen Water Rate Specification, California Data Collaborative, 2018. Moonshot award, state Water Data Challenge. Project page
5The other sectorsHealth records, UK Open Banking, Green Button in energy, and the Australian Bureau of Meteorology feed are covered in the five-case study set. npc.here.now/waterdataexploration
6Why GTFS workedThe specification was deliberately kept as plain CSV: "we wanted it to be as simple as possible so that agencies could easily edit the data, using any editor," and complexity "would have killed adoption." Five-case set

Section 02

LLMs and AI software have collapsed the cost of data analysis

Language models read source documents that were never written to be read by machines. A rate schedule issued as a PDF, an HTML table, or a scan of a board packet can now be parsed and its figures extracted at a cost measured in fractions of a cent, in whatever form the utility produced it.

That capability removes the reason most of the last two decades of water data work existed. Making rate data machine-readable meant asking every utility to publish it the same way, because the only way software could read four hundred documents was for all four hundred to be written alike. Language models have removed that requirement. What remains is a second and separate problem: finding the source document in the first place, and establishing that it is the one in force.

Identity, authority, recency and completeness are still needed

A model infers the rate structure from source documents, but is limited to what is written. As a result, it cannot extract facts about a document's relationship to the world. For example, it cannot know whether a rate schedule is outdated unless another document states it, or a utility itself publishes the information in another way.

What must be knownWhere that fact actually lives
IdentityIn a registrar's records. A rate schedule does not state which state system identifier the publishing agency holds, so matching a document to an agency is a lookup rather than a reading.
AuthorityIn the minutes of a board meeting. A document may assert that rates were adopted, and where it does the assertion can be stale or absent.
RecencyIn the existence of a later document that has not been seen. Nothing in the schedule in hand reports that it has been superseded.
CompletenessIn the list of who should have published. A process that reads only what it found cannot report what it missed, because absence produces silence rather than an error.

Improvements in model capability make reading steadily cheaper, but extraction will remain limited to what has been published and can be discovered.

The work that remains is correspondingly small. It requires an address for the source document, the date a board adopted it, and the date it took effect. That is a fraction of what a specification asks, and the record of the last twenty years indicates that the size of the ask determines whether this sector adopts anything.

Twenty years went to the format problem, and none to the location problem

Two independent conditions determine whether a utility's rates can be assembled into a dataset. The first is the format of what it publishes. The second is whether a third party can locate where it published. Twenty years of effort have gone into the first, with the results already described, because changing it asks each agency to restructure its documents, its systems, and its staff skills. Comparatively little effort has gone into the second, which asks for one line: a URL the utility already maintains for its own customers, published where a machine can read it.

Sorting the sector's options by who initiates the transfer, and by whose form the material takes when it arrives, produces four positions.10 California water rate reporting occupies one of them today.

10In data engineering termsSource and Reported correspond to schema-on-read and schema-on-write. The move from mandated submission to publication and discovery is ETL to ELT, with the transform step relocated from publisher to consumer.
SourceReported
Push · Source

Deposit

The agency transmits the operational artifact itself, unshaped.

  • Public Records Act responses
  • SEC EDGAR filings
Limited byThe recipient's processing cost, which scales with volume.
Target
Pull · Source

Publication and discovery

The agency posts what it already produces; the consumer locates it and imposes structure afterward.

Limited byFinding the document, and knowing it is in force.
Retained pushThe source-document address: URL, adoption date, effective date
Today
Push · Reported

Mandated submission

The agency computes a figure to a specification it did not write and files it on a fixed cycle.

  • eAR Section 8A, rate structure as free text
  • Federal tax returns
Limited byThe form itself. A new question means a change to a state system.
Pull · Reported

Commissioned collection

A consumer pays someone to gather a fixed field set from wherever it can be found.

  • A rate study's benchmarking task
  • Manual academic rate surveys
Limited byCost per record. Coverage stops where the budget stops.
Push · the publisher initiatesPull · the consumer initiates
Figure 2  Two axes of water rate data strategy. The horizontal axis records who initiates the transfer: under Push the publisher sends material to a named recipient, on the recipient's schedule and in the recipient's form, and under Pull the consumer fetches it from wherever the publisher already keeps it. The vertical axis records whose form the artifact takes: a Source artifact is the one the agency made to run its own business and keeps the full logic, and a Reported artifact is one figure computed to fill someone else's form. California water rate reporting sits in mandated submission, where coverage is complete because filing is mandatory and the form determines which questions can be asked. The recommendation moves the transformation work to publication and discovery without vacating push: one address remains a filed artifact, with a location as its payload rather than a figure.

The known failure of the publication and discovery position is the swamp, where the material is all present and nobody can find or trust any of it. The standard mitigation is a catalog or registry that carries the location and provenance of each source document.

Section 03 · Case study

whatwatercosts.org put 62% of bills within $5, but found only half the utilities

whatwatercosts.org is an open-source project holding residential rates for hundreds of utilities across the United States; its author later joined this team.11 It is a proof of concept rather than infrastructure, assembled from the PDFs and web pages utilities already publish, seeded from the federal system list and utility domains.

The whatwatercosts.org home page, showing a search field and a sortable table of utilities with their monthly bill at 6,000 gallons, price per thousand gallons, and population served.
Figure 3  The site as published, listing residential rates for 562 utilities serving 165 million people across 50 states, each with a standardized monthly bill at 6,000 gallons. Every figure was read from a document the utility already publishes. whatwatercosts.org

In this work a single engineer assembled, as a side project, a dataset the sector has treated as a six-figure procurement. The driver was the falling cost of reading rate sheet data, which means compute and engineering can now replace repeated manual collection. No utility changed a format, adopted a schema, or filed anything. No custom IT system, internal policy, or staff training changed. Conversion into a structured, comparable form happens in the pipeline rather than at the agency, with the output stored in a JSON refactoring of the Open Water Rate Specification, so the standardization work moved to the consumer side rather than disappearing.

Nearly two-thirds of the bills landed within $5 of a hand-collected dataset

The pipeline was validated against the North Carolina Environmental Finance Center's FY2026 cost tables across 99 utilities, comparing a standardized bill at 6,000 gallons per month.12

$1.98
Median absolute error against the manual dataset
62%
Within $5 of the benchmark bill
47%
First-pass automated coverage
$0.028
Cost per utility extracted
Figure 4  whatwatercosts.org validated against the North Carolina Environmental Finance Center FY2026 cost tables, 99 utilities. Accuracy on the utilities it reached is competitive with manual collection. Coverage is where it falls short.

Accuracy on the utilities the pipeline reached is already comparable to collecting the data by hand. Coverage and verifiability persisted as the gap. Utility data that could not be found on the web could not be included, and data that was found was not directly validated as current by each utility. The coverage gap is measurable, because when the source document's URL is supplied in advance, the cost per utility falls from $0.028 to $0.003 and coverage approaches complete.

Blind discovery
The pipeline finds the document on its own
Rates extracted47 of 100
With a published address
The URL is supplied in advance
Rates extracted96 of 100
Rates extracted Not reached — the source document was never located Each segment is one utility in a hundred; heavier rules mark tens.
Figure 5  Blind coverage of 47% is measured. The lower strip shows the coverage reported when URLs are known in advance, which the California benchmark described in Section 4 is designed to test at scale. Locating the source document is the bottleneck rather than the extraction.

Once the rates exist as data, the applications are close to free. Three already exist in the Open Water Rate Specification ecosystem: a household bill calculator, a cross-agency rate comparison, and a model of revenue and typical-bill impact for a proposed rate change.13 Each was straightforward to build once the underlying data was available, and each is currently limited to the handful of utilities whose rate structures were specified by hand.

The lookup that follows draws on the same data.14 It was assembled from the export in an afternoon, and it is the whole argument in one object: the rates are real, no utility did anything to put them there,15 and whether any given figure can be checked depends entirely on whether a source document was recorded.

14Snapshot, not a live callBuilt from the whatwatercosts export on 4 August 2026 and embedded in the page, so it keeps working regardless of the API. Coverage changes daily; the figures here do not.
15Which systems appearNinety-three California systems are in the export. The eighty-seven shown are those where recomputing the bill from the published tier structure reproduces the published benchmark exactly. Most of the six excluded are budget-based agencies, where the bill depends on a household allocation the single-tariff view cannot represent.
Figure 6  A working lookup over 87 California water systems, serving 19.4 million people, built from a snapshot of the whatwatercosts.org export taken 4 August 2026. Every figure was read by machine from the document each utility already publishes, and no utility changed anything to appear here. Move the slider to recompute a bill from the published tier structure. The provenance strip beneath each result is the part that matters: where a source was recorded, the record shows the document and the sentence the figure came from, and where none was recorded there is nothing to check the number against. Forty-four of the eighty-seven carry a source. That ratio is the problem this report is about.

The project was limited by coverage and maintenance responsibilities

Two limitations matter for how far the current dataset can be trusted.17 The seed list is filtered to systems serving 10,000 people or more, which excludes exactly the segment where the state's SAFER program records the concentration of failing and at-risk systems. And the federal list supplies database names rather than the names a utility is known by locally, so matching a record to the agency a reader recognizes is not always straightforward.

A third limitation is maintenance. A dataset that a sector comes to rely on should not depend on one volunteer's continued interest. Instead, utilities should publish their source documents at reliable addresses, so that analysis and applications, including ones that do not exist yet, can be built on utility-verified sources.

11whatwatercosts.orgOpen-source, independently built. Methodology at whatwatercosts.org/methodology; full data export at /api/export.json
12Environmental Finance CenterThe comparison dataset is collected manually and is a recognized reference for utility rate benchmarking. It reports the same location-versus-extraction pattern independently.
13OWRS toolingBillCalculator and RateComparison, California Data Collaborative.
16How "source recorded" is decidedThe lookup below marks each system as source recorded or not, on the single basis of whether the export carries a source document for it, which is directly observable. The export also carries a confidence rating of its own, whose values are not defined on the published methodology page and which varies independently of whether a source was recorded, so it is not used here.
17SAFERSafe and Affordable Funding for Equity and Resilience, the State Water Board program tracking failing and at-risk water systems.

Section 04

Become discoverable by publishing the source document's address

Something will answer the rate question whether or not utilities take part

Households are already directing rate questions to AI assistants, and those assistants answer from whatever they can find. Several already cite whatwatercosts.org rather than working the figures out independently. Something will answer the question of what water costs, and why, whether or not agencies participate in data discovery. The key question is whether the answer reaching water customers — locally, and at state and national level — originates in a source the utility stands behind.

The address therefore has to come from the utility rather than be inferred by whoever is collecting. If the collection side carries all of the work, then whoever runs the best pipeline becomes the effective authority on what every utility charges, which places a public fact under private control by default rather than by any decision. A published address reverses the relationship: the aggregator becomes an index, and the utility remains the author of its own numbers.

Three actions follow. Together they do one thing: they publish the address of the rate schedule a utility has already written, along with the date a board adopted it and the date it took effect. None of them changes the format of any document a utility publishes, and none requires a change to a state submission system.

The actionWho acts, and the board's part in itWhat it costsWhat it unlocks
One · Publish the current rate schedule at an address that does not move, with the date a board adopted it and the date it took effect, and list that address in the site's sitemap.xml. The utility. The board directs staff. No vendor, no procurement, no state involvement. One staff afternoon, then a few minutes each rate cycle. Rates become findable and checkable by anyone answering a question about them, including the utility's own customers. Roughly 60% of civic domains already serve a sitemap, so for most utilities this is a configuration change rather than a new capability.
Two · Record that same address, with its two dates, as one field in the electronic Annual Report. State Water Board. A board's route is to ask for it through the associations it already belongs to. A single field on a form that already exists. Statewide coverage, inherited from a mandate that already exists and is already tied to funding eligibility.
Three · Publish a certified crosswalk between the state system identifier and the other identifiers a utility appears under.18 The California Data Collaborative, as the sector's coordinating body. Coordination, and the benchmark work already underway. Records can be matched to the right agency reliably, so one place to look actually resolves to one utility.

Only the first is a board decision in the ordinary sense, and it is the one that produces most of the benefit. The second would settle the matter permanently and is a request to the state rather than something a board does directly. Whether the current Section 8 schema can carry such a field is being confirmed.

The ask is a URL, an adoption date, and an effective date

Nothing about the rate schedule itself changes. What is added is a statement of where it lives and what a board did with it.19

Example · a utility's published source-document address

# https://www.example-water.gov/llms.txt

# Example Municipal Water District
PWSID: CA1910xxx

## Rates
- Current residential rate schedule: https://www.example-water.gov/rates/schedule-2026.pdf
  Adopted: 2025-11-18 (Resolution 25-114)
  Effective: 2026-01-01
  Supersedes: /rates/schedule-2024.pdf
Figure 7  An address, a date the board adopted it, and a date it took effect. Nothing about the rate schedule itself changes.

Recording whether a board adopted the document an address points to, and when it took effect, is the same move that EDGAR and FERC already make. EDGAR stamps each submission with a filing type and an acceptance time and threads later amendments back to the original, and FERC derives which tariff is currently effective from filed effective dates and supersession. Both systems answer the currency question with a recorded field rather than by inference, which is the only way it can be answered reliably.

A pilot is testing whether the index can be assembled this cheaply

None of this needs to wait on the actions already set out, and a first attempt is under way.20 A crowdsourced collection is being run across the 455 California community water systems serving 10,000 people or more, which together cover roughly 37 million residents and the large majority of the state's metered connections. Two independent researchers work the same list, recording each utility's official domain, the direct URL of its current residential rate schedule, the effective date, and the apparent update cycle. Most utilities take one to two minutes. Disagreements between the two are resolved by the project team, and the rate at which they agree is itself reported as a measure of how findable this material is.

The budget for that collection is approximately $650, funded by four contributors at roughly $165 each.21

That figure buys a one-time collection of addresses by two researchers, and it is not the cost of a working registry. It excludes verifying what those researchers found, keeping the addresses current as rates change, and resolving disagreements about which document is authoritative. It also excludes whatever running the result as a dependable service would cost. It is evidence that the first step is cheap enough to attempt rather than a price for the finished thing. The collection is designed to test three things: whether extraction coverage rises above 90% once URLs are known, how accurately automated discovery finds those URLs without help, and what the floor of participation is, meaning how often a rate document can be found when a utility publishes nothing but its domain. The third question matters most for policy, because it establishes the smallest thing a utility could be asked to do.

The results and the dataset are to be published openly, including the cases where the approach does not work, since the value of the exercise depends on reporting those as well. If the California results hold, the project plans to extend extraction to the largest two thousand systems nationally over the following six months, calibrated against the California ground truth. The index of source-document addresses would then be published as open data, with a route for utilities, researchers, and practitioners to submit and correct entries.

Supporting the work

The California Data Collaborative coordinates this effort on behalf of its member agencies. Utilities interested in publishing a rate schedule address, or in the benchmark collection described above, can reach the Collaborative directly.

None of this depends on the technology getting better

The case so far rests on cost: a pilot in the hundreds of dollars against rate studies that run from forty thousand to six figures. That gap would hold even if model quality froze today, because the saving comes from not re-deriving structure by hand.

The same logic covers the planning-horizon question. A pipeline built around current models is a reasonable worry for a sector that plans in decades, but the recommendation is not a pipeline. It is an address, and two dates tying the record to the right agency, the facts that make a retrieved rate canonical by establishing identity, authority, recency and completeness. Because that artifact is data rather than a tool, it holds regardless of which model reads it, so there is no vendor lock-in and nothing to future-proof.

19llms.txtA plain text file at a predictable location describing a site's contents for automated readers, following the pattern of robots.txt and sitemap.xml. The Protocol Institute brand kit this report is built from publishes one.
18Identifier crosswalkThe model is GLEIF in financial reporting, which publishes certified mappings between entity identifiers as open reference data.
20Benchmark scope455 California community water systems serving ≥10,000 people, approximately 37 million residents. Two independent annotators, five-minute cap per utility, agreement rate reported as a metric. This is a working proposal dated July 2026, not a completed study; its figures describe what is planned rather than what has been produced.
21Budget detailTwo annotators at roughly 15 hours each, $18/hour, plus platform fee and $0.003 per utility for extraction. Approximately $650 in total, funded by four contributors. The figure covers collection only, and carries no allowance for verification, maintenance, dispute resolution, or operating a registry.

Section 05

Urban water plans and groundwater reports face the same test

Any public reporting domain can be sorted by two questions. Can the document be found, and how usable is it once found. Where documents are readable but scattered across independent publishers, findability is now the limiting factor and further work on submission standardization has limited upside against what automated extraction already does. Urban Water Management Plans, groundwater sustainability plan annual reports, and water loss audits all sit here, as does a great deal of municipal reporting that has nothing to do with water. Where a regulator already requires filing, findability comes free, and accuracy and verifiability are the questions that remain.

The same pattern recurs elsewhere. Automated readers absorb variation in form, which removes the reason most standardization efforts existed. What they do not absorb is disagreement about identity, location, and currency, and those are settled by publishing small agreed facts rather than by restructuring documents. The work that remains after the technology arrives is smaller than the work that came before it, and it is a different kind of work.

The practical consequence for a utility is that data which is liquid supports analysis that was previously not worth commissioning. A bill calculator and a rate comparison were straightforward to build once the underlying rates existed as data. The same will be true of questions nobody has thought to ask yet, which is the ordinary result of making something cheap that used to be expensive.

Appendix

Evidence and background

Both case sets are published in full. Between them they cover twelve systems in other sectors that faced a version of this problem, and the four sources this analysis leans on most.22

Case set · five sectors
How other sectors reached coverage
Five sectors that converged on a shared data standard, each examined for the mechanism that actually produced adoption rather than the standard's design.
GTFS · US transitFHIR · US health recordsUK Open Banking Green Button · US energyBureau of Meteorology · Australian water
Read the five-case set →
Case set · seven precedents
How canonical records were made discoverable before language models
Seven systems that solved the two problems this report is about — knowing where a publisher's current record lives, and knowing it is the one in force — none of which used machine learning.
SEC EDGARUtility tariff filingOAI-PMH union catalogs Web self-descriptionDOI & CrossrefLEI & GLEIFIRS Form 990
Read the seven-case set →
Background · protocol research
Where this way of thinking comes from
The Protocol Institute researches how coordination works among parties who share no common authority, which is the general form of the problem in this report. Its magazine and the Summer of Protocols research program publish the underlying work.
Protocol InstituteProtocolizedSummer of Protocols
Visit protocol-institute.org →
22SourcesAspen Institute 2017 on the Internet of Water; Stanford 2016; Berkeley 2018; Department of Water Resources 2018 on AB 1755 protocols.

About

About this report and the Protocol Institute

This analysis was conducted by the Protocol Vision team, a research group of the Protocol Institute that helps organizations see how protocols and AI shape their operations. The team examines persistent problems within organizations and industries, the coordination entanglements that produce them, and the simple interventions that move the performance frontier.

For more about the Protocol Institute, and about how protocols shape our lives, visit protocol-institute.org.

Protocol Research Lead
Rafa Fernandez
LinkedIn
Project Sponsor
Patrick Atwater
LinkedIn
whatwatercosts Creator
Maxwell Titsworth
LinkedIn