Capturing data that exists only right after a disaster, in a form machine learning can use immediately
The National Science Foundation awarded roughly $1.63 million to the University of California-Berkeley to collect perishable field data after earthquakes, hurricanes, floods and wildfires through standardized workflows and publish it as open, machine-readable datasets. The aim is data that can enter machine learning models without extensive preprocessing.
Grant overview (primary data)
- Award amount$1,633,151
- RecipientUniversity of California-Berkeley (California)
- ProgramECI-Engineering for Civil Infr, Special Initiatives
- Period2026-10-01 〜 2031-09-30
- FunderU.S. National Science Foundation (NSF) / NSF
Key points
- Collecting field data that exists only right after earthquakes, hurricanes, floods and wildfires, and publishing it as open machine-readable datasets through standardized workflows ($1,633,151).
- The aim is to advance the GEER organization, with over two decades of experience, into the era of digital technology and data science.
- Operations adopt a three-tier structure: field deployment team, virtual data-mining team, and central communications unit.
- Datasets follow findability, accessibility, interoperability and reusability so they enter machine learning models without extensive preprocessing.
- The 120 NSF awards this site holds as of 2026-09-02 span 37 states with California leading at 24; the five-year period exceeds the median of three.
1Once it is cleaned up, it is gone
The site immediately after an earthquake or flood holds information no laboratory or simulation can reproduce: which retaining wall failed, which embankment settled, how the ground cracked. That evidence disappears within days once recovery work begins. Perishable is the word this project uses for such records — data whose window for collection is confined to just after the event. GEER is the framework that has organized that reconnaissance for more than two decades.
2Those who go, and those who do not
- 1Field deployment teamEnters the affected area and collects field performance data
- 2Virtual data-mining teamUses crowd-sourced information, social media streams and satellite remote sensing to support field planning and populate a public event clearinghouse
- 3Central communications unitConnects the two and organizes the flow of information
- 4Scientific advisory groupsCover extreme weather, earthquakes, ground and infrastructure instability, and emerging technologies
Splitting into three tiers follows from the constraint that only so many people can go on site. What can be learned from satellite imagery and public information is gathered without travelling, so the limited time on the ground goes to what can only be collected there. The organizational structure answers a question about allocating scarce disaster-response resources.
Broad participation in the virtual teams by students, engineers and public stakeholders is also part of the aim.
3What AI-ready means
The point of the project is not collecting but making what is collected usable by machine learning as it stands. All gathered datasets are to follow principles of findability, accessibility, interoperability and reusability, enabling direct integration into machine learning models without extensive preprocessing. Research data that is published but inconsistently formatted still costs enormous effort before anyone can use it. The design pays that cost up front.
4Within the distribution of states
The 120 NSF awards this site holds as of 2026-09-02 span 37 states, with California leading at 24, followed by Texas at 13, Massachusetts at 8 and Illinois at 6. The period runs 2026-10-01 to 2031-09-30 — five years, longer than the median of three at that date. An award involving data standardization and the running of an organization calls for a different investment of time than a single study. Amounts are the obligated amount as of the check date and may change.
Why it matters
Publishing data that is inconsistently formatted still leaves a large cost before use. Paying that cost at collection time is a design that carries to any effort trying to keep data as an asset rather than a byproduct.
FAQ
What is perishable data?
Why have a team that does not go on site?
Sources (primary)
Source: NSF Award Search (U.S. National Science Foundation, public domain). Amounts are the obligated amount. For privacy, we do not handle principal investigator names.
- NSF Award (original, official)
- NSF Award ID: 2629719