UC Berkeley said on October 2, 2026 that researchers there have used AI and data science to assemble what it describes as the first open-access database of local laws from across the United States, containing millions of unique ordinances from all 50 states. The university said a version of the database was published online in June and a paper describing the project was accepted to the Conference on Neural Information Processing Systems.
What the project built
According to UC Berkeley, the project collected every digitally accessible local law into a free, open-access database, spanning jurisdictions "from Alameda, California, to Zephyrhills, Florida." The university said online repositories have cataloged state and federal statutes for years, but that this is the first time such a database has existed specifically for local laws.
Lead researcher Diag Davenport, an assistant professor of technology policy, governance and society at Berkeley, said: "We've built a new kind of telescope — one pointed at governance itself." He added that it "could be a fundamental shift in the way people interact with the rules that govern their daily lives."
Berkeley said the dataset lets researchers search across local laws at a scale that was previously extremely difficult and identify patterns in how communities regulate housing, public space, business activity and everyday conduct.
How the dataset was assembled
The university said researchers first collected municipal and county laws from thousands of government and third-party websites. Davenport and postdoctoral scholar Denis Peskoff consulted lawyers while developing the collection process and designed it to meet the technical requirements imposed by the sites hosting the documents, Berkeley said. They worked with AI researcher Joe Barrow and Berkeley undergraduate Christopher Vu.
The larger technical challenge, according to the university, was converting nearly 10,000 documents — roughly 7 million pages, many stored as blurry, poorly structured or otherwise inaccessible PDFs — into usable data. The team used a vision-language optical character recognition model called LightOnOCR to extract text and structure.
Berkeley said processing the archive required a massive amount of computing power, and that the researchers then used models from OpenAI to tag and organize samples of the laws, distilling that work into buckets applied across the full corpus.
Why the researchers say it was hard
Davenport began considering the project six years ago while studying computer algorithms, systemic bias and the criminal legal system, the university said, and wanted to study bias in the law itself. He found no database of all local laws existed, with material scattered across proprietary company databases and local government websites covering more than 3,000 counties and roughly 6,000 other local jurisdictions.
"Of course, nothing's ever as simple as you want it to be," Davenport said. "Once you realize how fragmented it all is, it's easy to understand why no one's done the work." He also said: "I think it was actually impossible to do this work until six or 12 months ago."
What comes next
Berkeley said some people have already created searchable interfaces based on the team's work, and that the next stage is for computer scientists and other technically skilled users to train chatbots on specific subject areas — for example, Bay Area building codes to compare apartment requirements in Berkeley and El Cerrito. Davenport said any forthcoming tools based on the data must be free and publicly accessible.
The team will present its findings at NeurIPS in December, the university said. "The real promise here is making local government legible," Davenport said.
Key facts and where they come from
- The database contains millions of unique ordinances from all 50 states.
The resulting database contains millions of unique ordinances from all 50 states.
- The team processed nearly 10,000 documents totaling roughly 7 million pages.
turning nearly 10,000 often unwieldy documents — roughly 7 million pages, many stored as blurry, poorly structured or otherwise inaccessible PDFs — into data
- A vision-language OCR model called LightOnOCR was used to extract text and structure.
The team used a vision-language optical character recognition model called LightOnOCR to extract both text and structure from the documents.
- OpenAI models were used to tag and organize samples of the laws.
the researchers used models from OpenAI to tag and organize samples of the laws
- A version of the database was published online in June.
The team published a version of the database online in June
- A paper on the project was accepted to NeurIPS, to be presented in December.
a paper describing their project was accepted for the Conference on Neural Information Processing Systems
- The US has more than 3,000 counties and about 6,000 other local jurisdictions making laws, Berkeley said.
more than 3,000 counties and another 6,000 or so local jurisdictions that make their own laws
- Davenport said any tools built on the data must be free and publicly accessible.
Davenport said any forthcoming tools based on the data must be free and publicly accessible.
