Author:
Jacqueline J. DearbornPublished:
On September 10, 2026, Jacqueline J. Dearborn, Library Innovation Lab Fellow, gave a virtual talk about the Public Data Project’s International Data Infrastructure Research and Development effort. The transcript below has been lightly edited for clarity.
Hello everyone, before I get started, I want to give a shout-out to Molly Hardy and Jack Cushman at the Harvard Library Innovation Lab for giving me this opportunity to reflect on some of the most pressing challenges of our time.
With Jennifer Chapman as my counterpart, we’ve spearheaded the Public Data Infrastructure Research and Development project and through this work, we’ve gotten the chance to zoom way, way, way out and take a look at the global public data landscape.
Obviously this is a very huge, thorny, and complex topic. Much larger than any one talk could possibly resolve. So today, I share some preliminary insights that I hope will help us think together about public data infrastructure, long-term stewardship, and collective responsibility. In this spirit, I hope this talk is less of an ending and more of a beginning for all of us that spurs many future conversations to come.
Thriving. Struggling. Gone. Reborn?
Since I know some of you but not all of you, I thought it might be useful to tell you a little bit about myself. So, I began my library career right here at Harvard in 2011. I worked as a grad student at the Museum of Comparative Zoology’s Ernst Mayr Library to process and digitize collections for the Biodiversity Heritage Library. Since then, I’ve spent most of my time working on the infrastructures that hold public data.
As you can see by my career timeline in logos, the institutions and platforms have changed, but the mission to open up knowledge for all really hasn’t. I’ve moved through government, academia, libraries, museums, nonprofits, and research organizations. I’ve followed my purpose and, perhaps less romantically, I have also followed a lot of three-to-five-year funding cycles.
So while many of you may know me best as a member of the BHL Secretariat and Data Manager, I’ve also worked on and contributed to a wide range of other infrastructures, including Internet Archive, Digital Commonwealth, ADS, Boston Open Data (now known as Analyze Boston), Resource Watch, Dataverse, MassDocs, and a myriad of Wikimedia projects.
I suppose the charitable word to describe my career trajectory would be tenacity — or maybe just a stubborn refusal to abandon the work that I believe in. And all this moving around has given me an unusual vantage point. I’ve spearheaded new initiatives and I’ve been around projects that are thriving. I’ve watched others struggle, and heartbreakingly, I’ve watched some go extinct altogether. I’ve also seen some reborn and flourish again.
And when you’re working on the inside, you tend to measure success by fairly standardized metrics: Did we digitize the collection? How many pages? Did we build the API? Did we transcribe it? Did researchers use it? How many? All of these metrics definitely matter. But lately, I’ve become interested in different measures of success because data generation and stewardship is so incredibly labor-intensive.
Measuring success beyond funding cycles
Lately, I’ve begun to ask: Will the data still exist 20 years from now? What about 50? Or 100? And I don’t simply mean, will the files still be somewhere. I mean: Will anyone still understand what the data is, why it exists, and how to use it? Will there still be a community that cares for it? Will there still be enough funding and expertise to repair and maintain it?
These are very different metrics and they entail much more than preserving just the bits and bytes. And honestly, this is part of why I had to take the last year off. For me, this has been a critical moment to reflect. I spent a long time being very close to the machinery, especially in my last role as a data manager at the Smithsonian. And now I’m reflecting on the larger system that this machinery sits inside of.
And since I’m still figuring things out, I’m not going to give you some grand theory of the universe today. But I will give some personal observations I’ve made over the years. The things I’ve seen. The things I keep coming back to. As well as some uncomfortable truths that I think are becoming harder for us to ignore.
Straw, wood, or brick?
So before we get too serious, I want to begin this talk with a little story called “The Three Little Pigs.” This may not be where you expected things to go, but please bear with me. So, we’ve got 3 little pigs. One builds a house made of straw. One builds a house made of sticks. And one — the boring, sensible pig — builds a house made of brick. And then, of course, the big bad wolf comes along, huffing and puffing. And, well, we all know how this story ends.
I’ve been thinking about that story lately as a metaphor, and how it maps rather uncomfortably well onto the infrastructure that holds public data. Because we have built an awful lot of houses with data in them. Some are made of very strong materials. And some are built with straw.
And some infrastructure is beautiful, expensive, and sophisticated. But, nevertheless, it depends on one or two people holding the keys. And for a very long time, we as data home builders have tended to ask questions like: Is the data safe? Did we back it up? Is it online? Can people find it? Can the machines read it?
All of those are great questions. But I think we’re increasingly having to ask the harder ones: Is the house itself structurally sound? Have we built on shaky foundations? What happens if the wolf comes around? Are we ready? Because the big bad wolf isn’t really one thing, is it? Sometimes it’s a change in institutional priorities. Sometimes it’s a funding cut. Sometimes it’s an aging server well past its warranty. Sometimes it’s a person leaving after twenty-five years. And sometimes … it’s simply just time.
And that brings me to the loaded word I’ve been dancing around since the beginning of this talk: infrastructure.
The houses that public data live in
So far we’ve basically defined infrastructure as “some houses with data in them.” This word “infrastructure” may conjure thoughts of code, servers, platforms, networks, storage, software. But I think we’ve inherited a rather narrow technical definition of infrastructure. Because if the goal is to preserve and transmit knowledge across generations, then infrastructure is actually the entire system that allows data to be created, interpreted, trusted, preserved, and shared in the first place.
So let’s map it out. What are some of the core components of infrastructure? Well, there is obviously technology and data but we also have people, expertise, funding. And increasingly, I’ve been thinking about governance and the law as the wrapper around the entire house where human roles, responsibilities, decision rights, and accountability mechanisms all define how those other pieces actually work together.
Sure, there are twenty-seven ways to make this mental model more academically rigorous, but for now, I’m finding it really useful. Today I want to discuss the parts of the house that I know best which are the technical and deeply human dimensions. Because the various components in the house are not independent. They’re interdependent: Technology without expertise doesn’t get us very far. Expertise without funding disappears. Funding without governance becomes chaos. And so on.
So we shouldn’t think of infrastructure as a handful of separate pieces rather we should think of it as the house they build together. Like the electrical, plumbing, HVAC, and so on. All of these things have to work together if you want the house to function and keep its contents safe. And once you look at infrastructure this way, some of the things happening right now start to look less like isolated crises … and more like maybe we have a lot of straw houses and in them is our incredibly valuable public data.
When the house becomes brittle
Katherine Skinner, Director of Programs at Invest in Open Infrastructure, cuts straight to the issue for us:
Data is disappearing because the infrastructure in which it is nested was already brittle.
And in real time, we are watching that brittleness become increasingly visible. I’m not trying to catastrophize here, because this isn’t one single big-bad-wolf moment isolated to the U.S. Everywhere, infrastructure providers are facing some combination of AI bot traffic, leadership and funding shifts, aging technology, and the loss of key personnel.
Infrastructure survival can depend on a remarkably particular set of circumstances: a grant; a director; a willing host institution; a community champion; sometimes just a handful of people who really cared. And one of the most persistent “wolves” I have noticed has been chronic underfunding.
There is evidence for this. Invest in Open Infrastructure examined more than $550 million in grant funding across the public data ecosystem. Of the funding that flowed directly to infrastructure, nearly two-thirds supported research, development, and innovation. But only about 20 percent supported core operations and maintenance. So we are very good at funding shiny new things, but we are not so good at funding them to endure.
And when operational funding begins to evaporate, the effects compound. People leave; institutional knowledge leaves with them; the system stops evolving; users lose confidence; and the community shrinks. And from the outside, it might look as though the data is still there. And sometimes the files are technically still there. But the community-driven ecosystem that made new data generation possible begins to fall away.
When infrastructure stops evolving, in a sense, it stops living. And that is a different kind of loss. For those of us who have been close to the machinery, fighting to keep these systems alive, that loss can feel a lot like intense grief. Which brings me to an important distinction. When public data is endangered, rescuing the bits and bytes may be absolutely necessary. But rescue is only the first leg of a much longer intervention.
Rescue preserves the data. But stewardship preserves the conditions that allow that data to remain usable, meaningful, connected, and dynamically alive. If we copy the files somewhere safe but separate them from the communities, provenance, expertise, and infrastructure that gave them meaning, have we really secured them for the long horizon? Perhaps rescue is the emergency intervention. Stewardship is what has to come next.
Data, information, knowledge
So bear with me while I take you back to library school for a moment: A geo-coordinate can be recorded. A page image can be preserved. A database record can be copied. And we have the LOCKSS principles to safeguard this data. We can make multiple copies. We can checksum things. We can put them in geographically distributed storage. And we can keep them in highly durable formats.
But saving those data points doesn’t really mean we’ve preserved how to interpret and use them. And this difference really matters. Data can be recorded and replicated; information gives it context; knowledge tells us what it means and what to do with it — and much of that knowledge does not live in the files themselves. It lives in people. And that creates a major preservation problem. Because we are very good at preserving data. But people? You can’t LOCKSS a person … at least not yet.
You cannot LOCKSS a person
Since antiquity, human beings have always transmitted knowledge. We spoke it aloud; carved it into stone; wrote it on papyrus and parchment; printed it into books; and now we replicate it as digital data across the planet. The medium keeps changing, but throughout history, knowledge has survived through a continuous chain of human stewardship.
So despite all our new-fangled technology, I’m here to tell you something: the oral tradition is alive and well, folks. It never actually went away. Because ultimately, all of this knowledge is for human beings. We created this data. We are the core users of the public data ecosystem. And I think it’s time we start prioritizing the human use case — not just the AI use case.
AI can dramatically accelerate our ability to discover, connect, and use shared knowledge. But it does not eliminate the need for human stewardship. Its greatest promise is to expand humanity’s capacity to learn, understand, and act. And of everything in our infrastructure house — technology, funding, community, governance, law, expertise — the part I think we understand least well is the deeply embedded human one.
The people. Us.
Because we have a lot of good ideas about preserving data, but you can’t LOCKSS a person. You cannot make a second copy of somebody’s twenty-five years of experience and put it in another geographic location. You can’t checksum institutional memory. This kind of knowledge exists everywhere. We often call it tacit knowledge.
Sometimes that just means knowledge that has never been written down because, quite literally, its only storage medium is a human being. I’ve watched this happen. And I’ve participated in it too. I’m guilty: I’ve been the person who knew where something was; I’ve inherited systems from someone who later retired; and I’ve watched people leave organizations carrying extraordinary amounts of infrastructure out the door with them.
That’s changed how I think about resilience. Because sometimes the first failure isn’t a funding cut or a downed server. It’s a person. Somebody walks out the door, and the cascade begins. So when we talk about investing in infrastructure, we need to get much more comfortable saying something that sounds obvious, but apparently is not:
People are infrastructure. And if we don’t invest in the people who maintain, interpret, govern, and evolve our public data infrastructures, then I can’t really see how public data survives this decade. And that brings us back to the house. Because the same problem that happens when we put too much knowledge in one person’s head also happens when we put too much of our technology in one bespoke system.
We don’t need a gajillion Ferraris
So, the year is 2026 and we have built a remarkable number of beautiful, one-of-a-kind data platforms. Some of them are genuinely brilliant. They solve very specialized problems. They have really cool features and slick interactive interfaces that are optimized for a particular collection or community. And that’s not necessarily a bad thing.
But we’ve also developed a habit of building things that only one or two people in the world know how to maintain. And all that sparkle may be masking some very shaky foundations. And at some point recently I started thinking of data platforms as custom Ferraris. They’re beautiful, impressive, rare; but if something goes wrong, you really gotta hope the mechanic is answering their phone and won’t ask you for a hefty ransom to fix it.
We need a shared FOSS ecosystem
You know what I think? I think we need more Honda Civics. Not because Hondas are better cars, but because there is an enormous ecosystem around them. There are parts, mechanics, manuals, standards. There are many people who can look under the hood and have some idea what they’re looking at. That’s a really important property for public data infrastructure to have.
Because the question shouldn’t just be: can we build it? The question should also be: can more than one human maintain it? Can someone else easily contribute to its code base? Can another institution host it? Can a new person learn it? Can the feature that is genuinely unique sit on top of something shared? And I think this is where open source infrastructure becomes much more compelling than simply “free software.” The real value is not just that the code is available. The real value is that we can distribute the capacity to understand and maintain it.
We don’t need every institution to build its own bridge to some siloed data island. We need to build common infrastructure across datasets and collections using free and open source software, and then we do our distinctive, special-snowflake things on top. Because we need more houses built with brick. And in the virtual world, I think one of the key ingredients for a strong house is Free and Open Source Software (FOSS).
When convenience costs us sovereignty
Unfortunately, there is a rather enormous emerging issue with everything I just said. The FOSS ecosystem does not maintain itself. People maintain it. And many of the institutions that we have come to depend upon are losing the internal capacity to build, deploy, operate, and maintain open infrastructure.
As universities and research organizations outsource more of their technical capacity amidst constrained budgets, the immediate convenience is obvious: someone else manages the servers; someone else handles the updates; someone else keeps the lights blinking. And over time, institutions can stop employing the people who know how to build and operate infrastructure independently.
And this strategy might save us a short buck — but does it get us to the long horizon? One of my mentors, Professor Peter Cornwell, put this much more directly:
FOSS system platforms have advanced functionality; they share vigilance and overcome emerging security hazards via an expert community; they don’t attract license fees and operate indefinitely.
Many institutions have lost these skills: they’ve become support organizations for proprietary products. We must urgently set about restoring these capabilities.
And I think that last point is the one I want to underscore. We may eventually still have the FOSS code. But will we still have enough people who know how to deploy it? That is why there is some urgency here. If we want to preserve the option of building a global data commons on shared, open, interoperable foundations, we have to invest in the people capable of doing that — and we need to do it now.
Otherwise, we may discover that we still possess the blueprints … but have lost access to all of the builders.
What does resilience actually look like?
So I’ve just painted an elaborate picture using the Three Little Pigs and custom Ferraris to illustrate what human and technical fragility look like from the inside. Which brings us to the central question: What does resilience actually look like?
And for this, I want to borrow from biology — because I’ve spent most of my career in biodiversity and I can’t resist. Healthy ecosystems aren’t resilient because every component is invulnerable. They’re resilient because they are diverse, distributed, redundant, interconnected, adaptive, and regenerative. Instead of monocultures, we have biodiversity. Mother Nature really is the best designer, isn’t she?
And I think public knowledge infrastructure needs to start thinking in these terms. It’s not: How do we make one organization, brand, or institution the end-all, be-all house for the data? It’s: How do we build a system capable of surviving the failure of any one part?
That principle has to apply across the entire infrastructure house: Shared technology. Distributed expertise. Inclusive communities. Diversified funding. Transferable legal arrangements. Adaptive governance. Not one institution holding all the cards. Because if an organization has six months of runway left, that’s not resilience. That’s a countdown.
Resilience doesn’t mean making every component perfect. It means creating a system that is less dependent on perfection. We’re not trying to build one house that the wolf can never blow down. We’re trying to build an ecosystem where losing one house is a recoverable event — where the wider community has the resources, rights, knowledge, and capacity to carry the work forward.
We aren’t starting from scratch
Now, I realize I have spent quite a bit of time telling you what is going wrong. But I want to be very clear about something. We are not starting from scratch. In fact, many of the ideas we need already exist: We have LOCKSS and the wonderful idea that Lots of Copies Keep Stuff Safe. We have trusted repositories. We have open-source platforms. We have persistent identifiers. We have interoperable standards. We have a global community of experts operating all over the world.
And we already have enormous networks of libraries, archives, museums, universities, research organizations, governments, nonprofits, community groups, developers, collectives, and individual experts who care very deeply about keeping public data both sovereign and alive.
So the problem is not that humanity has no idea how to do this. Because we have many of the pieces. What we have not yet done is connect those pieces into a sufficiently resilient global system. LOCKSS taught us something incredibly important about data: Don’t depend on one copy. I think the challenge before us now is to take our learnings from LOCKSS one step further and apply them to not just the data but all of the dimensions of infrastructure:
Don’t depend on one institution. Don’t depend on one funder. Don’t depend on one platform. Don’t depend on one country. Don’t depend on one technical team. And certainly don’t depend on one person.
A global mycelium network
And there is something that makes a distributed system like this possible. Standards are the agreements systems make with one another. Law and governance are the agreements people and institutions make with one another. Together, these form a global mycelium network. Persistent identifiers. Shared protocols. Common data structures. Open-source stacks. And open governance processes that determine how we work best together.
Yes, these things feel painfully boring. Nobody is going to give a keynote about the metadata field they viciously argued about internally about for five years and finally standardized in 2017. But this is what makes interoperability and continuity possible. Shared agreements are what allow data and resources to move between systems. Because ultimately, standards and governance are relationships and social agreements expressed technically, organizationally, and legally.
And that’s why the next paradigm won’t be a collection of isolated successful data projects. It will be a unified community with shared foundations to steward the commons. And from my technical point of view, a shared FOSS ecosystem isn’t a nice-to-have. It’s the missing requirement for long-term public data stewardship. Which brings me to my next thought.
Towards a global data commonwealth
What would it look like if we stopped thinking about public data infrastructure as something that individual institutions steward and are ultimately responsible for? What if we thought about public data and infrastructure as a kind of shared inheritance? Not owned by one library or one institution or one country or even one generation. Something we steward collectively because the consequences of losing it are collective.
I don’t have a fully formed answer for what a global governance model looks like for that. But I’m hoping some of my friends here at Harvard Law Library and the people in the audience can lend a helping hand. And I actually think it’s useful to acknowledge that I don’t personally have all the answers. Because none of us do. And it is only through our own humility and our ability to cooperate and collaborate that any of this actually works.
I don’t think the answer lies in some bold visionary coming onto the scene to save the day. Or creating another organization, consortium, or intergovernmental interface. We already have all those things. Instead we need to answer the harder questions here: How do we connect what already exists? How do we change how resources flow? How do we protect and distribute expertise? How do we prevent single points of failure across all dimensions?
So as we near the end of this talk, I think we’ve evolved the narrow technical definition of “infrastructure.” To me, it’s beginning to sound a lot less like a bunch of data platforms and a whole lot more like a commons. Or perhaps a global commonwealth? Imagine a truly global network with different roles, different skills, and different interests — bound by the shared mission to steward public data and carry it forward for the collective benefit of all and future generations to come.
The future of the commons
A good friend of mine, Ben Vershbow, is the founder of As We May Think, which is a new lab dedicated to strengthening the digital commons. When I asked him what the future might look like, he had some interesting insights. He said:
A commons concentrated in a few prestigious and powerful institutions is a fragile commons. The future lies in a federation of diverse communities — attuned to their contexts, connected through shared standards and infrastructures — capable of stewarding knowledge together.
And I think the important word here is diverse. Not diversity as decoration or inclusion as a slogan. Diversity as resilience. A monoculture is fragile. A resilient commons needs many kinds of stewards: people, collectives, communities, organizations, and institutions — across geographies, disciplines, cultures, and various levels of power.
And if we genuinely want distributed stewardship, then we need to distribute resources more effectively and efficiently: money; authority; expertise; opportunity. Because we cannot decentralize stewardship if we are centralizing all the resources.
What now?
I’ve spoken to you today from the parts of the infrastructure house that I know best — the technical and intimately human side. But I’ve only lightly touched on the legal and international governance dimensions that will determine what a resilient global data commons might look like. That’s where my counterpart Jenn and the brilliant folks at the Library Innovation Lab, and many of you here, come in.
Jenn frames this moment as both a crisis and an opportunity to redesign and rethink the current paradigm. In her forthcoming paper, “Responding to the Global Access to Information Crisis through Sustainable Information Access Goals,” she writes:
In the face of mounting challenges to global information access, it is imperative that librarians, information professionals, and other stakeholders develop resilient, connected collections that counter current disruptions and withstand future attacks.
And I think that brings us back to the central point. None of us will solve this alone. The challenge is too big for any one institution, any one discipline, and probably too big for any one country. We may not know exactly what that future looks like yet. But I think we’ve reached the point where we need to start designing for it.
So who will steward it?
So, back to our original question: Who will steward humanity’s public data? I suppose I should admit something. I don’t actually know. But perhaps the answer starts with the people in this room. Perhaps many of us came here today because something about this question stirs something in your gut that transcends your career, your home institution, grant obligations, or personal loyalties to any particular brand, project, or platform.
And if we’re really serious about saving public data, then I think that asks that we all reflect a little bit on some very difficult questions:
- Have we built systems that depend too much on our own expertise, authority, institution, or vision?
- Are we making it easier for others to carry the work forward — or harder?
- And perhaps most uncomfortably: are there moments when stewardship means holding on, and others when it means letting go?
As we know, the public data landscape is shifting and we are experiencing a lot of data losses. Now more than ever, our community needs to come together, globally. To do this work, it will require that we think way beyond ourselves and commit multiple acts of humility, generosity, and service. Because ultimately stewardship means building for a future that does not depend on us.
Thank you all so very much for letting me have this soapbox moment. I’m incredibly grateful to be in a room filled with brilliant people willing to think about the future of public data — and I hope to collaborate with some of you to get busy building brick houses for all the data we’ve been rescuing as of late.