Nurturing a Dynamic & Coordinated Data Resilience Community
The long-term availability and usability of research data are increasingly at risk. Funding instability, infrastructure fragility, governance gaps, and political volatility have exposed systemic vulnerabilities across the data preservation and stewardship ecosystem. Research data resilience depends on more than preserving individual datasets; it requires sustained investment in the people, organizations, governance arrangements, and infrastructure that allow data to remain discoverable, usable, and continuously available. While some domains benefit from mature repositories and shared norms, many others rely on ad hoc, unevenly resourced efforts to safeguard critical datasets. Across the research ecosystem, preservation and continuity efforts often operate in silos, with limited coordination around contingency planning, provenance, governance models, and technical standards.
With generous support from the Portfolio to Protect Science, Phase I of this project focused on building a clearer understanding of these challenges and identifying shared priorities for action. ORCA conducted 18 key-informant interviews, complemented by desk research, with funders, repository leaders, technologists, archivists, librarians, researchers, community-based data stewards, and other key actors across the data resilience ecosystem. The resulting report, Data Resilience Interview Synthesis and Candidate Pilots, identifies seven recurring challenges and opportunities: the structural vulnerability created by heavy dependence on federal funding; the importance of sustaining people and infrastructure alongside individual datasets; the need to preserve continuity of data collection as well as archived data; the gap between making data available and making it reusable; the foundational importance of discoverability; the potential for state, local, and community infrastructure to advance both resilience and equity; and the growing need for effective governance as private-sector involvement increases.
These findings were translated into five candidate pilot areas: (1) a senior-junior mentorship network for data stewardship; (2) a pre-competitive research-industry data space; (3) curation for reuse, co-designed with a private-sector partner; (4) a sustainability funding-model playbook for public-good data; and (5) a distributed continuity stress test for critical public data. These pilots are candidate responses to recurring needs identified through the listening tour, rather than final selections or pre-established partnerships. Each builds on existing work and focuses on areas where convening, coordination, partnership development, or new funding pathways could strengthen the broader ecosystem.
In Phase II, supported by the Dana Foundation, ORCA will move from mapping vulnerabilities to piloting actionable interventions that strengthen resilience in practice. ORCA will convene practitioners, infrastructure and repository leaders, funders, researchers, and potential implementation partners to validate and refine the candidate pilots, then co-design, implement, test, and evaluate a focused set of interventions in live environments. The ultimate goal is to generate practical evidence about what works—and under what conditions—to accelerate adoption of resilient data practices and strengthen a more coordinated, sustainable, and equitable data resilience ecosystem.