About LT4CPR

Building language technology for the moments it matters most.

LT4CPR is a collaborative research project building the language technology infrastructure needed for effective crisis preparedness and response.

The communities most vulnerable to disasters are often precisely those whose languages are least supported.
01 · The challenge

Powerful tools, unequal access.

Today’s machine translation, speech recognition, summarization, and text classification systems serve only a fraction of the world’s languages, leaving many communities without dependable tools during crises.

6,500+ languages are spoken or signed around the world.

02 · Our response

Two corpora, one shared infrastructure.

Together, these resources will help researchers, humanitarian organizations, and responders build and evaluate language tools where they are most needed.

Speech translation Datasets for languages with limited digital resources, led by George Mason University in partnership with CLEAR Global.
Situation reporting Social media datasets for report generation, led by the University of Washington.

Intellectual merit

The project produces datasets grounded in real crisis scenarios and languages with limited digital resources, while systematically identifying gaps in existing crisis response infrastructure and guiding future research.

Machine translation · Speech translation · Summarization · Classification

Broader impacts

LT4CPR centers languages spoken by communities most at risk but least served, working with NGOs that operate in crisis settings.

  • Faster triage and translation for responders
  • Information communities can access in their own language
  • Clearer communication across language barriers

This is an NSF Collaborative Research grant. Two institutions hold separate awards that together constitute one unified project: speech and NLP for languages with limited resources at GMU, and social media NLP and multilingual systems at UW.

George Mason University George Mason University
NSF 2346334PI: Antonios Anastasopoulos
View award
University of Washington University of Washington
NSF 2346335PI and coprincipal investigator: Fei Xia · William D. Lewis
View award
Grant period October 2024 to September 2027 NSF program CCRI · CISE Community Research Infrastructure
Official project language Read the full NSF abstract

Language technologies are promising and could have strong impact during disaster responses. They can help to triage text messages in a disaster to determine what aid to provide. Language technologies can translate vast amounts of data related to an ongoing pandemic. Responders can use these technologies to converse with victims during disaster responses. However, advances in language technologies to date are limited. They focus on a few dozen of the more than 6,500 languages spoken or signed in the world today. Current language technologies neglect millions of people. This especially impacts those who are most at risk for experiencing disasters. This project provides an infrastructure for language technology advancements for crisis response. The results will be useful for everyone, no matter the language they speak.

This project builds datasets of crisis communications using dedicated data collections and social media harvesting. These datasets will be applicable to curated crisis scenarios. They will use common language scenarios necessary to communicate with vulnerable populations. This approach helps people for whom language technologies are not typically developed. The project will bring together researchers from different disciplines. These include language technology researchers, experts in disaster relief, linguistics, and human computer interaction. The project will target representatives from the local speech communities to take part. To coordinate this effort, the project will organize yearly workshops and shared tasks with the communities.