LT4CPR data

Data that can move with a crisis.

We are developing multilingual datasets grounded in crisis communication to help researchers and responders build more useful speech translation and situation reporting tools.

Releases in development
Coming soon

Two connected data resources.

Collection and documentation are underway. Release details, languages, licensing, and access instructions will be posted here as they are finalized.

01 · Speech Coming soon

Crisis speech translation corpus

Speech translation data for languages with limited digital resources, grounded in the communication needs that arise during crisis preparedness and response. This corpus is being developed in partnership with CLEAR Global.

Lead George Mason University Status In development
02 · Reporting Coming soon

Multilingual situation reporting corpus

Social media data designed to support research on turning multilingual crisis communication into useful, timely situation reports.

Lead University of Washington Status In development
How we approach data

Built to be useful beyond one study.

01Grounded in real crisis communication
02Centered on underserved languages
03Documented for responsible reuse
04Shaped with expertise across sectors
Stay connected

Follow future data releases.

Join the LT4CPR Slack community for project news, workshops, and opportunities to contribute.

Join the community