The Internet Archive’s Information Stewardship Forum ran three days in March 2026 at the Archive’s headquarters at 300 Funston Avenue in San Francisco and drew 120 people whose job, in one form or another, is keeping US government information from disappearing. A report on what they found, published October 6, is titled “Who Is Government Information For?” The 120 attendees could not answer that question with a single map of what exists, what has been saved, and what is already gone, because no such map exists.
The Forum sits inside Democracy’s Library, the Archive’s project, launched in 2022, to collect government publications at every level of government and keep them accessible once agencies stop hosting them. The library held over 11 million items as of mid-2026, and the Archive has since become a designated Federal Depository Library, joining the more than 1,100 libraries that already receive and keep federal publications under that 1813 program, as the Archive described in June.
Why the count keeps changing
The Forum convened against a backdrop of genuinely contested numbers. The New England Journal of Medicine reported that the CDC removed 203 datasets, 13% of its online total, in three weeks between January 21 and February 11, 2025 — a drop from 1,519 datasets to 1,316. Separately, a National Security Archive brief on the Freedom of Information Act process noted that more than 2,000 datasets disappeared from data.gov in the weeks after the January 20, 2025 inauguration, leaving 306,770 datasets on the catalog, still more than 1,000 fewer than before the transition.
Those two figures measure different things at different moments and don’t reconcile cleanly, which is the point the Forum’s report keeps making. A newer tracker, dataindex.us, counts something narrower still: data products with documented, primary-source evidence of termination — an agency notice, a press release, a regulatory filing — excluding anything taken down and later restored. By that stricter definition, the tracker had logged 28 terminated data products as of July 2026, alongside hundreds of continuing datasets that lost specific variables, most often gender identity, sexual orientation, race and ethnicity fields. The same update cites Harvard’s Library Innovation Lab finding the data.gov catalog’s raw entry count swinging by thousands in a single month — up 6,000, then down 2,000 — from routine reorganization that has nothing to do with censorship. A raw catalog count, a named-dataset count, and a primary-source-verified termination count will each put a different number on the same event. None of the three groups producing them had, before March, sat in the same room.
How the Forum worked
The event ran under Chatham House Rule — plenaries, lightning talks, and “birds of a feather” breakout sessions — with attendees from libraries, archives, journalism, research, philanthropy and technology, spanning federal, state and local government information work. The report’s five findings read less like discoveries than confirmations of what each attendee already suspected about adjacent fields: nobody holds a complete inventory of what has been lost; preservation is not the same as access, because metadata, interfaces and the staff who know how to use a dataset can rot even when the raw bytes survive; the field runs on emergency triage that cannot scale into permanent infrastructure; records laws exist but enforcement does not, and some anti-scraping measures aimed at blocking AI training crawlers end up blocking the Archive’s own preservation bots too, since both arrive at a server as an anonymous HTTP request that a site operator’s firewall has no reliable way to tell apart. The report singles out local government records, climate data, health data, and disaggregated demographic data as the categories most likely to vanish without anyone outside a narrow specialist community noticing.
| Attendees agreeing (%) | |
|---|---|
| Strengthened cross-org connections | 93 |
| Found collaboration opportunities | 89 |
| Learned new tools or methods | 83 |
| Plan to stay in contact | 98 |
Source: Internet Archive
Those numbers measure a networking event’s success at networking, not whether any additional government record now has a backup copy it didn’t have in February. The survey was answered by attendees the Archive had already selected and who chose to show up; a forum about fragmented coordination finding that its own attendees felt better coordinated afterward is close to a tautology. The report does note that roughly 20 of the 120 respondents described a specific collaboration or initiative they intend to pursue — the only figure in the survey that points at a future action rather than a feeling about the past three days.
The call to action
The concrete output is a document, Preservation of Government Information: A Call to Action, drafted by James R. Jacobs of Stanford’s Free Government Information project with input from Data Rescue Project members who attended the Forum. It’s hosted both at the Data Rescue Project and at publicinformationtrust.org, and asks individuals and organizations to endorse building a “cooperative digital preservation infrastructure” and to push for what it calls an essential update to US information policy. Neither version of the page specifies a funding mechanism, a piece of legislation, or a target number of partner institutions, and neither publishes a running signatory count.
| From | To | How |
|---|---|---|
| Agency publishes a page or dataset | Routine Wayback Machine crawl | captured before any change |
| Agency publishes a page or dataset | Agency removes or alters the page | |
| Agency removes or alters the page | Volunteers flag the gap (Data Rescue Project) | only if someone notices |
| Routine Wayback Machine crawl | Democracy's Library | |
| Volunteers flag the gap (Data Rescue Project) | Democracy's Library | mirrored copy |
| Democracy's Library | Researchers, journalists, courts |
Based on Internet Archive
The diagram the Forum’s own findings imply has a weak link built in: the path from an edited or removed page to a rescued copy runs through a volunteer noticing the change at all, with no systematic alert in between. That’s the gap the report calls moving “from emergency response to durable infrastructure,” and it’s also the gap the three incompatible loss-counting methods above are symptoms of — nobody has instrumented the pipeline, so everyone is still counting by hand, after the fact, by whatever definition their own project happened to adopt.
US records statutes already obligate agencies to preserve what they produce; the Forum’s finding is that enforcement, not law, is what’s missing, and that gap is a policy choice rather than a technical limit. The Archive’s next event, Democracy’s Library: Tools for Participation, is scheduled for October 21, 2026, at its San Francisco headquarters and online. What it won’t settle on its own is the harder number: how many of the roughly 20 collaborations attendees said they intended to pursue turn into an actual archived record, checkable against a termination log that three different trackers still can’t agree how to define.
