Register
"*" indicates required fields
"*" indicates required fields
This page is a practical, applied data studio package for developing disease risk monitoring and mapping workflows, specifically for humanitarian emergency preparedness and response to events within international settings. It combines datasets relevant to humanitarian disease risk, up-to-date population and geospatial datasets (e.g., climate, elevation, topography) and imagery to give an emergency responder a one-stop location to find datasets most relevant to their use case(s), with a specific focus on vector-borne and water-borne diseases.
This data studio focuses on generalized disease transmission risk, seasonality, and outbreak detection, with guidance on data limitations and integration with complementary health and environmental systems, and prioritizes the protection of at-risk populations through the implementation of early alerts, warning systems, vulnerability mapping, and intervention targeting. Built for humanitarian responders, WASH and humanitarian clusters, outbreak disease control programs, researchers, and data teams.
These datasets are primarily linked to the systemic risk definition defined in the IPCC's 2021 report: "the potential for adverse consequences for human or ecological systems, arising from the interaction of climate-related hazards, exposure, and vulnerability." This allows us to assess disease risk using remotely-sensed datasets.
Note that this data studio is being created in conjunction with a disease risk decision making guide [link here].
Best Practices
Humanitarian situations are immensely complex, fast-moving situations where utilized datasets and the underlying context often changes depending on the situation.
Below, we have highlighted some best practices for how to utilize these datasets, grouped by the time period to be used:
| Time period | Actions to be taken |
| Either before an emergency occurs, or before the user goes to an emergency situation |
|
| During the emergency situation itself |
|
Additional Opportunities/Training Links
This is not an exhaustive list of training opportunities, but are potential mechanisms to obtain training in EO implementation in epidemiology. Please reach out to the team if you have additional suggestions!
The Disease Incidence and Resource Estimator (DIRE) from UC San Diego, UNICEF, and New Light Technologies. DIRE uses machine learning tools to provide dengue fever and malaria outbreak trends in Brazil and Peru, as well as predict the number of potential cases in the following month. They use epidemiological, socioeconomic, and environmental data to inform the model, many of which are open access.
The European Space Agency sponsored the Waterborne Infectious Diseases and Global Earth Observation in the Nearshore (WIDGEON) project to monitor the risk of waterborne diseases to humans using EO across India.
Researchers in the GeoHealth & Hydrology Lab at the University of Florida have developed the Vibrio Prediction Hub, which predicts the risk of cholera in parts of Africa, Asia, and Europe.
The Epidemic Prognosis Incorporating Disease and Environmental Monitoring for Integrated Assessment (EPIDEMIA) tool combines surveillance data and environmental data from EO to predict malaria outbreaks in Ethiopia. The team at University of Oklahoma and Ethiopian public health agencies, NGOs, and universities have made many of their forecasting tools publicly available.
The Lebanon WASH Sector’s 2025 Waterborne Disease Risk Map identified 85 regions that are “highly vulnerable” to waterborne diseases. This built on a cholera risk map created in 2022 and adds information on drought vulnerability to data on population, proximity to contaminated sources, WASH services and access, and case trends.
Graphic of a humanitarian snapshot provided by UN-OCHA, containing information about food insecurity, rainfall performance/disease outbreak, and overall displacement/conflict. It also contains figures about severely food-insecure people, acute malnutrition, internal displacement, refugees, cholera and mpox cases.
A case study documenting an early-stage use case, the Malaria Anticipation Project (MAP), developed by Medicins Sans Frontieres (MSF) and the Sweden Innovation Unit (SIU), and launched in 2021. Piloted in Lankien, South Sudan, the project explores whether historical health and climate data can be used to predict the malaria burden. The pilot used currently-tracked malaria cases seen in the facilities, and used predictions to identify when and where there could be a rise in malaria cases would help MSF and other actors to plan a more efficient response and potentially reduce malaria-related mortality and morbidity.
To prevent large-scale cholera outbreaks in the Democratic Republic of the Congo (DRC), the UN Central Emergency Response Fund (CERF) released four allocations over three years (pilot: 2022-2024; final stage: 2025) as part of an anticipatory action framework that contained cholera outbreaks, saved lives and maximized the impact of limited resources. The pilot was facilitated by OCHA in collaboration with the DRC’s National Cholera Elimination Plan, the United Nations Children’s Fund (UNICEF) and the World Health Organization (WHO), and enabled partners to respond in days, not weeks, to an uptick in cholera cases.
This is a dashboard containing information relevant to the WASH Cluster for Yemen, run by UNHCR. There are three sections to this dashboard: Humanitarian Needs and Response Plan (HNRP), 5W, and AWS/Cholera.
An overview of each of the displayed sections is below:
This is a dashboard from European Centres for Disease Control (ECDC) that includes information on a Daily Suitability Index (for Daily Vibrio Risk), Suitability Index (Max and Mean), and Forecast for the next 4 days. Information is discretized into Very Low/Low/Medium/High/Very High Category.
The Vibrio viewer is a real-time model that uses daily updated remote sensing data to examine worldwide environmental suitable conditions such as sea surface temperature and salinity for Vibrio spp. according to Baker-Austin et al. (Nature Climate Change 3, 73–77, 2013).
Infections caused by Vibrio species other than V. cholerae can also be serious notably for immunocompromised persons, but the overall occurrence is low despites an increase having been recently observed in Northern Europe. Note that only imported case of acute Vibrio cholerae infections are [notifiable cases] in [the EU]. Further work is on-going to improve this environmental suitability model in collaboration with NOAA CoastWatch, ECDC, CEFAS, University of Bath and the University of Santiago de Compostela (for more information in the model summary ). Early information about the environmental suitability will be of public health interest to assess the geographic extent of potential human exposure. Please note that this model has been calibrated to the Baltic Region in Northern Europe and might not apply to other worldwide settings prior to validation.
For the Baltic Sea, the model parameters are optimized for the following values: colour palette (boxfill/vibrio), number colour bands (10) , scale method (linear), legend range Min. value (0), and Max. value (28).
Dashboard created by UNHCR that contains a gridded modelling of a risk-based score, based on conflict events, conflict fatalities, high-temperature days, heavy precipitation days, drought accumulated, population estimate, and food security class.
This is a dashboard created by the UNHCR WASH Cluster State of Palestine, containing 6 separate pages: a dashboard showing number of cholera cases ("Community Assessed"), an overall assessment of water systems ("Overall Assessment"), costs and main sources of drinking water ("Water"), stacked barcharts showing hygiene needs by weekly/monthly basis ("Hygeine"), the results of a sanitation survey ("Sanitation"), and the results of a survey for Solid Waste Management ("Solid Waste Management").
This dashboard, created by the US Centers for Disease Control and Prevention's National Center for Emerging and Zoonotic Infectious Diseases (CDC/NCEZID), contains data on enteric bacterial, viral, and parasitic agents and other foodborne, waterborne, and fungal diseases. It serves as a repository of multiple surveillance systems (ncludes data from the System for Enteric Disease Response, Investigation, and Coordination (SEDRIC), the National Antimicrobial Resistance Monitoring System for Enteric Bacteria (NARMS), and the National Outbreak Reporting System (NORS)). Available dashboards include early signal detection, reported outbreaks, antimicrobial resistance monitoring, Vibriosis surveillance, Salmonella serotypes of concern, and Food Safety Funding, and all data goes down to the state level.
This interactive report, which is created by the Assistance Coordination Unit of Syria, is intended to show vaccine preventable diseases (suspected VPD curve, suspected measles, investigated measles, epi curve & complicated & symptoms; incidence rate (as a whole and by sub-district), and investigated measles cases). All data that is collected uses the Early Warning, Alert and Response Network (EWARN) surveillance system, managed by WHO EMRO.
The WASH Insecurity Analysis (WIA) is a sector-wide analytical framework designed to inform evidence-based decision-making across the humanitarian–development continuum. The WIA provides a common methodology for understanding WASH needs and risks by assessing WASH service levels, exposure to hazards, and WASH-related vulnerabilities. It is designed to be practical and scalable, leveraging commonly available secondary data to generate sub-national-level analyses. By identifying geographic areas and quantifying both the population at risk of WASH insecurity and those already experiencing it, the WIA enables governments, humanitarian actors, and development agencies to prioritize sectoral investments and improve coordination. It also supports localization efforts, offering insights that can guide preparedness and response before and after shocks.
Dashboard created by End Malaria Partnership, which contains separate pages on surveillance and intelligence, supply and commodities, country support, funding, gaps & donors, campaigns and community health workers (CHWs), and research and innovation.
Dashboard created by the WHO's Western Pacific Regional Office, that collates information on incidence rate, incidence, mortality, and testing of P. falciparum cases.
The Malaria Atlas Project are leaders in geospatial analysis, spatiotemporal statistical methods, machine learning, and computational disease models for malaria (both P. falciparum and P. vivax).The platform also provides malaria data at varying levels of detail to suit different needs. This suite of tools aims to support the needs of diverse audiences, including the general public, the media, policy analysts, engineers, as well as researchers.
Global elevation model (~30 m), created by the Shuttle Radar Topography Mission (SRTM) instrument, with improved processing over original SRTM, hydrologically corrected elevation model for flow accumulation, slope angle and terrain analysis, and landslide exposure modeling.
DEMs can be used to calculate elevation for either mosquito habitat suitability and cholera risk, or can be processed (conditioned) to allow for watershed analysis.
Multiple modalities are available for download:
Additional Links
Quick-start guides on SRTM usage can be found here:
NASA’s Integrated Multi-satellitE Retrievals for GPM (IMERG) product combines information from the GPM satellite constellation to estimate precipitation over the majority of the Earth's surface. Available in near-real time (provides high-resolution (~10 km) rainfall estimates obtained every 30 minutes), enabling timely, accurate flood forecasting, especially in areas that lack ground-based precipitation-measuring instruments, including oceans and remote areas.
IMERG fuses precipitation estimates collected during the TRMM satellite’s operation (1998 - 2015) with recent precipitation estimates collected by the GPM mission (2014 - present) creating a continuous precipitation dataset spanning over two decades.
Note: another precipitation dataset which combines IMERG datasets with ground sensors calibration is CHIRPS, produced by UCSB. More information can be found here (for the UCSB site) or here (for Google Earth Engine).
For CHIRPS, please note the following:
Multiple modalities for download includes:
The best practice is to follow these requirements:
Global daily measurements of nighttime lights imagery that reveals real-time infrastructure and power disruptions, both in magnitude and signature, providing rapid assessment of affected population centers and enabling targeted humanitarian responses.
Use any of the following websites, depending on data need:
Provides detailed roads, paths, buildings, POIs, and other infrastructure mapped by volunteers and local mappers, which can assist with routing and access analysis, station siting, evacuation planning, and base layers for dashboards.
You can access data directly from OpenStreetMap.org, or download data for export directly from OpenStreetMap using either the export instructions, the Overpass API, PlanetOSM, GeoFabrik (limited to certain geographies), or other sources.
Google’s Open Buildings dataset is a global building footprint dataset designed for humanitarian applications. The polygon data is comprised of a set of .csv files, which can be downloaded by individual cells on a global map. Similarly, point data can be downloaded as well. Google also provides a Colab notebook on how to download building footprint data for a specific country or geographic area.
From the data download website, you can either download the tiled version, use a provided Google Colab notebook to help subset to your region/area of interest, or download the data directly from Humanitarian Data Exchange (https://data.humdata.org/organization/google-open-buildings).
The polygon data (178 GB total) is composed of a set of CSV files, with one file per level 4 S2 cell that are up to 7.8 GB in size. Similarly, the points data (48 GB total) are up to 2.1 GB per file.
Note that the CSV datasets here encompass well-known text (WKT), which encompass the boundaries geometries.
Additional datasets that can also suffice here include the Microsoft Building Footprints dataset, which was derived between 2018-2021 using Maxar (Vantor) and Airbus imagery.
Near-global geographical mapping of surface water extent over land at a spatial resolution of 30 meters, allowing for water persistence mapping, with a temporal revisit frequency between 6-12 days. Using Sentinel-1 radar observations, DSWx-S1 maps open inland water bodies greater than 3 hectares and 200 meters in width, irrespective of cloud conditions and daylight illumination that often pose challenges to optical sensors.
Can be downloaded either from:
Further NASA access instructions are below (taken from above link):
Land surface temperature ~(1 km; day/night) global archive, collected by the MODIS sensor, useful for baseline heat detection, heat stress and drought/fire conditions and evapotranspiration proxies. For the purposes of disease risk, these are most appropriate to estimate drought conditions. Note that land surface temperature is NOT the same as air temperature.
From the NASA website for this data product:
The Land Surface Temperature (LST) and Emissivity daily data are retrieved at 1km pixels by the generalized split-window algorithm and at 6km grids by the day/night algorithm. In the split-window algorithm, emissivities in bands 31 and 32 are estimated from land cover types, atmospheric column water vapor and lower boundary air surface temperature are separated into tractable sub-ranges for optimal retrieval. In the day/night algorithm, daytime and nighttime LSTs and surface emissivities are retrieved from pairs of day and night MODIS observations in seven TIR bands. The product is comprised of LSTs, quality assessment, observation time, view angles, and emissivities.
ERA5-Land uses ERA5 atmospheric variables, such as air temperature and air humidity, as input to control the simulated land fields. This is called the atmospheric forcing. Without the constraint of the atmospheric forcing, the model-based estimates can rapidly deviate from reality. Therefore, while observations are not directly used in the production of ERA5-Land, they have an indirect influence through the atmospheric forcing used to run the simulation.
Note: Here we document the ERA5-Land dataset that, in its consolidated version, covers the period from January 1950 to 2-3 months before the present. In addition, the ERA5-Land-T version delivers non-checked close to Near-Real-Time (NRT) daily updates. ERA5-Land-T is synchronized with the close to NRT daily updates provided by the ERA5 climate reanalysis (ERA5T).
Global surface soil moisture (coarse, ~9–36 km depending on product) with frequent revisits, crucial for a variety of use cases, primarily floods, landslides, and droughts.
Downloadable or usable in the following links, such as:
Additional Links
NASA Global Land Data Assimilation System Version 2 (GLDAS-2) has three components: GLDAS-2.0, GLDAS-2.1, and GLDAS-2.2. GLDAS-2.0 is forced entirely with the Princeton meteorological forcing input data and provides a temporally consistent series from 1948 through 2014. GLDAS-2.1 is forced with a combination of model and observation data from 2000 to present. GLDAS-2.2 product suites use data assimilation (DA), whereas the GLDAS-2.0 and GLDAS-2.1 products are "open-loop" (i.e., no data assimilation). The choice of forcing data, as well as DA observation source, variable, and scheme, vary for different GLDAS-2.2 products.
The project has resulted in a massive archive of modeled and observed, global, surface meteorological data, parameter maps, and output which includes 1-degree and 0.25-degree resolution 1948-present simulations of the Noah, CLM, VIC, Mosaic, and Catchment land surface models.
We have linked ALL possible GLDAS resources here. However, note that there are a series of GLDAS resources that can be utilized for humidity, land use/land cover (specific to vegetation), and soil moisture.
Downloadable through NASA or through Google Earth Engine:
Additional Links:
Dynamic World is a near-realtime 10m resolution global land use land cover dataset, produced using deep learning, freely available and openly licensed. It is the result of a partnership between Google and the World Resources Institute, to produce a dynamic dataset of the physical material on the surface of the Earth. Dynamic World is intended to be used as a data product for users to add custom rules with which to assign final class values, producing derivative land cover maps.
Its algorithm calculates per-pixel probabilities across 9 land cover classes, using incoming Sentinel-2 satellite image. For every pixel in the image, the algorithm estimates the degree of tree cover, how built up a particular area is, or snow coverage if there’s been a recent snowstorm, for example.
This file is available from the program via two modalities:
Additional Links:
Only available at 500m resolution and reflecting time between 2000-2023, the Normalized Difference Vegetation Index is generated from the Near-IR and Red bands of each scene as (NIR - Red) / (NIR + Red), and ranges in value from -1.0 to 1.0. This product is generated from the MODIS/006/MCD43A4 surface reflectance composites.
Minimum spatial and temporal requirements:
Note that this dataset may only be captured until 2023.
Sentinel-2 (S2) is a wide-swath, high-resolution, multispectral imaging mission with a global 5-day revisit frequency. The S2 Multispectral Instrument (MSI) samples 13 spectral bands: visible and NIR at 10 meters, red edge and SWIR at 20 meters, and atmospheric bands at 60 meters spatial resolution. It provides data suitable for assessing state and change of vegetation, soil, and water cover.
Minimum spatial and temporal resolution:
WorldPop, based out of the University of Southampton, creates gridded population distribution datasets (various resolutions/years), with exposure mapping and prioritization for alerts, response logistics, and equitable resource allocation.
It is important to note that these are population estimates, spatially modeled off of existing datasets, and may exist in bottom-up (i.e., modeled from survey data points) or top-down (i.e., disaggregated using small-area estimation methods) methodologies.
Please refer to this document for more information: https://www.worldpop.org/choosing-the-right-worldpop-population-data-for-you/
Heavy uncertainty in rural/remote areas. Users must keep track of year and methodology in order to understand where their particular dataset can be used or not.
High-resolution global built-up area mapping dataset created by the European Commission, including built-up presence, settlement typologies, population layers (multi-year), degree of urban classification (as of late 2025), which is instrumental in pinpointing densely populated and vulnerable areas to enhance accuracy in evacuation planning.
Please refer to the GHSL website for more documentation on how to utilize their platform, as well as the degree of urbanization (DEGURBA) methodology utilized in their 2025 release.
COD-AB (administrative boundary) files are a type of Common Operational Dataset (COD) that is used within the emergency response framework. These are the preferred COD datasets to be utilized for any administrative boundaries, especially during an emergency management setting.
Administrative boundary CODs include:
COD-AB datasets can be linked by database or GIS to COD-PS datasets, when available using the P-codes as a key.
Important to note is that the COD-AB datasets may not be 'completely' edge-matched between multiple areal units. However, this also means that in use cases where administrative boundaries are indicated and edges are not a concern (e.g., island locations), COD-AB boundaries are the only geographic COD boundary dataset that exists.
Boundaries and gazetteers reflect the country’s specific administrative hierarchy. These datasets must be prepared before an emergency.
The access instructions depends on the dataset - please see general link above and add search terms if necessary.
Documentation about how to use these types of dataset from UN-OCHA here: https://knowledge.base.unocha.org/wiki/spaces/imtoolbox/pages/2557378679/Administrative+Boundaries+COD-AB
Please note that the link indicates the "official" administrative boundaries that have been consolidated by UN-OCHA. If additional boundaries are needed, you can look at the potential link with search terms here: https://data.humdata.org/search?cod_level=cod-enhanced&q=COD-AB.
COD-EM, or edge-matched administrative boundaries, are a type of Common Operational Dataset (COD) that is used within the emergency response ecosystem to assist in cartographic representations.
Administrative boundaries are available for some locations in an alternate edge-matched version (COD-EM) as well as the original definitive COD-AB version.
While COD-ABs remain the definitive, accepted, best-available administrative boundary sets for each country, COD-EM layers may be more suitable for cartographic visualizations and might also be useful for topological analysis of lower administrative level international connectivity.
COD-EMs are adapted versions of the definitive COD-AB that have had their external feature polygons extended or shrunk to fit the United Nations Geo Hub 1:1m boundary. They therefore fit the COD-EM of neighboring countries, when available. They may be suitable for cartographic visualizations and might also be useful for topological analysis of lower administrative level international connectivity.
COD-EMs are separate datasets on HDX containing the gazetteer from the COD-AB, a zipped folder of shapefiles, a zipped geodatabase, and the ITOS live geoservice.
Depends on dataset - please see link above for details.
Please note: if you intend to use datasets from Humanitarian Data Exchange, make sure that the asset you pull for your geography is listed as being COD-EM! Currently, the HDx website may not directly pull up edge-matched datasets when you indicate "COD-EM".
Usage guide from UN-OCHA can be found here: https://knowledge.base.unocha.org/wiki/spaces/imtoolbox/pages/3092348933/Edge-matched+administrative+boundary+CODs+COD-EM
Population Statistics CODs are the baseline population figures of a country's pre-crisis situation. To effectively allocate resources and efficiently channel assistance to those who need it most, information about the size, location, and demographic profile of the population is fundamental. This requires information about the likely age-/sex-profile of the population in a given geographic area.
Population statistics are required to inform programming in humanitarian response. Specifically, they are used:
The COD-PS is constructed based on the best available data principle – i.e. in a pragmatic way that takes the best available data and constructs an updated set of age- and sex-disaggregated population estimates for the current time period at the lowest practical level of geographic disaggregation possible.
As it is a humanitarian tool, the COD-PS is not required to be an official statistical output constructed according to the international standards of official statistics. It is rather intended to be updated annually according to the best-available humanitarian data standard, or as humanitarian needs and priorities change, and allows for the input data and estimation/projection methods to be of a lower standard than official statistical standards (e.g., decennial census).
This depends on the dataset. Broadly speaking, the best way of accessing this data includes:
A usage guide from UN-OCHA can be found here: https://knowledge.base.unocha.org/wiki/spaces/imtoolbox/pages/2493349951/Population+Statistics+COD-PS
Country-specific CODs (COD-CS) represent local hazards and operational requirements such as key infrastructure that could be impacted or used during relief operations such as schools, health facilities, and refugee camps, or topographical data such as rivers, land cover and elevation.
Unlike COD-AB, COD-PS, and COD-EM datasets, which should be prepared before an emergency, COD-CS datasets may be prepared before or after an emergency. The appropriate types of COD-CS depend on the nature and context of a country or emergency. See the following pages for further guidance.
Existing databases must satisfy various conditions to be a classified as a COD-CS.

This depends on the dataset(s) to be used. While there are guidelines for what constitutes a COD-CS dataset (see above), the usage of the COD-CS classification depends both on the classification set by the Incident Manager (IM) of the humanitarian setting (e.g., Health Cluster, WASH cluster) and the type of emergency that has been experienced.
CSV containing subnational p-codes, their corresponding administrative names, parent p-codes, and reference dates for the world (where available). Latin names are used where available.
Navigate to the website linked above and download the CSV of interest (need to figure out the differences between each dataset)
UN-OCHA guidance for how to use P-codes can be found at: https://knowledge.base.unocha.org/wiki/spaces/imtoolbox/pages/222265609/P-codes.
This is a shortening of the guidance document provided by the Global Taskforce of Cholera Control (GTFCC), specifically for medium- or high-cholera burden countries, available here.
An example application of this approach can be found in the literature in Kiama et al (2025).
Audience: Primarily WASH and health clusters located in medium-/high-burden countries that need to quickly identify priority geographies where multisectoral interventions must be deployed.
Deploy time: Depends on the mechanism. Once data is acquired, an Excel tool (available on the GTFCC website) can aid in data creation.
Prerequisites: Have available all national cholera program (NCP) operational geographic units (Excel spreadsheet and spreadsheets), population data (COD-PS preferred, although gridded population estimates can be utilized as needed); case and screening data within the last 5-15 years (preferred)
Data & Tools:
Collate all of the data for the priority index:
Make sure you address the following steps before calculating the priority index:
Address missing data
Determine appropriate cholera test positivity indicator
Calculate individual indicators, before the overall priority index
Calculate each of the four component subindicators individually for each NCP operational geographic unit: incidence, mortality, persistence, and cholera test positivity. These indicators are derived from epidemiologic and cholera testing data over the analysis period.
First, calculate the three main epidemiologic indicators: incidence, mortality, and persistence:
Incidence rate in an NCP operational geographic unit = the total number of cholera cases (including suspected cases and cases tested positive) reported in the unit over the analysis period / (the cumulative person-time (i.e., the sum of population of the geographic unit for each year over the analysis period)) * 100,000
This indicator is the number of cholera cases reported per 100,000 person-years over the analysis period.
Mortality rate in an NCP operational geographic unit = the total number of deaths attributed to cholera reported in the unit over the analysis period / (the cumulative person-time (i.e., the sum of the annual population over the period)) *100,000
This indicator is the number of deaths attributed to cholera reported per 100,000 person-years in the unit over the analysis period.
Cholera persistence in an NCP operational geographic unit = the number of weeks with at least one reported suspected cholera case over the analysis period / the total number of weeks over the analysis period
This indicator is the percentage of weeks with at least one reported suspected cholera case in the unit over the period of interest.
Determine whether to incorporate testing data in your PAMI index. To do so, you first calculate representativeness for testing, followed by the overall testing prevalence.
Weekly testing coverage (i.e., representativeness in testing) for cholera in an NCP operational geographic unit: the number of weeks with at least one reported suspected cholera case tested for cholera (regardless of the testing method and of the result) over the analysis period / the number of weeks with at least one reported suspected cholera case over the analysis period –
If this is relatively representative, you then can proceed with positivity rate
This indicator is the percentage of weeks with at least one suspected case tested for cholera among weeks with at least one suspected case reported in the unit over the analysis period.
Use this indicator to then determine if you calculate the positivity rate or the number of years of cases, following this flow-chart:

Calculate the 50th and 80th percentiles of each epidemiologic component subindicator’s distribution across all of the NCPs.
1. Once these two percentiles are calculated for each indicator, each of the NCP operational geographic units’ values for each indicator should be compared to these percentiles to determine the score category:
| Epidemiologic Indicator | Score | |||
| 0 points | 1 point | 2 points | 3 points | |
| Incidence | No case | > 0 and < median (50th percentile) | ≥ median (50th percentile) and < 80th percentile | ≥ 80th percentile |
| Mortality | No death | > 0 and < median (50th percentile) | ≥ median (50th percentile) and < 80th percentile | ≥ 80th percentile |
| Persistance | No case | > 0 and < median (50th percentile) | ≥ median (50th percentile) and < 80th percentile | ≥ 80th percentile |
Calculate the priority index for each NCP operational geographic unit by summing the scores of the component subindicators as follows:
Priority index = incidence score + mortality score + persistence score + cholera test positivity score (if applicable)
To aid in implementation, you can use the Excel tool created by the Global Taskforce on Cholera Control (GTFCC)
Join the spreadsheet of your priority index to the shapefile(s) of interest, using your GIS tool of choice (e.g., QGIS, ArcGIS). Then, map that shapefile of interest.
Conduct a stakeholder validation workshop for each NCP operational geographic unit. The exact setup of this workshop is outside of the scope of this how-to – please refer to the guidance document linked above for more information.
Outputs: Priority index, by each operational NCP geographic unit, and a map of all priority indices by each NCP geographic unit.
Validation & QA: Ensure that you check for missing data (review the linked guidance document for additional information)
Maintenance: Ensure that calculation is done every 5 years, whenever a new NCP dataset is done. As mentioned in the PAMI documentation, “as a general principle, PAMI analysis should be updated when a new version of an NCP is developed (typically every five years). Earlier updates may be considered if there are significant changes in [cholera epidemiology].”
Risks & Considerations: Ensure that missing data is sufficiently dealt with before creating the index.
This is a shortening of the guidance document provided by the Global Taskforce of Cholera Control (GTFCC), specifically for no- or low-burden countries, available PAMIs for cholera elimination for no-/low-burden countries (i.e., vulnerability index documentation)
An example application of this approach can be found in the literature in Kiama et al (2025).
Audience: Primarily WASH and health clusters located in no-/low-burden countries that need to quickly identify priority geographies where multisectoral interventions must be deployed.
Deploy time: Depends on the mechanism. Once data is acquired, an Excel tool can aid in data creation.
Prerequisites: Have available all national cholera program (NCP) operational geographic units (Excel spreadsheet and spreadsheets), population data (COD-PS preferred, although gridded population estimates can be utilized as needed); vulnerability assessment factors within the last 5-15 years (preferred)
Data & Tools:
Collate all of the data for the vulnerability index. Only utilize the indicators that are necessary to your geography of interest.
Make sure you address the following steps before calculating the vulnerability index:
Address missing data, specifically:
Calculate individual indicators, before the overall vulnerability index
Score each of your subindicators with a presence index (1/0), for each NCP, according to the following metrics:
| Generic vulnerability factor | Example of measurable indicator for PAMI identification | Scoring Criterion | |
| 0 | 1 | ||
| Confirmed cholera imported case(s) in the NCP operational geographic unit considered | NCP unit with at least one confirmed cholera outbreak reported over the past five year/analysis period | No | Yes |
| Cross-border areas adjacent to frequently cholera-affected areas or identified PAMIs in neighbouring country(ies) | NCP unit with at least one confirmed cholera case imported (from another country or another NCP operational geographic unit reported during the analysis period), or Cross-border NCP unit adjacent to cross-border areas frequently affected by cholera outbreaks or classified as PAMI in neighbouring countries | No | Yes |
| Location along major travel routes with transportation hubs | NCP unit located along transportation pathway(s) with transportation hub(s) | No | Yes |
| Major population gatherings | NCP unit hosting major population gathering(s) | No | Yes |
| High population density locations or overcrowded settings | NCP unit with high population density or overcrowded settings | No | Yes |
| High-risk populations | NCP unit with high-risk population | No | Yes |
| Hard-to-access populations | NCP unit with hard-to-access population | No | Yes |
| Population that received oral cholera vaccine (OCV) more than 3 years ago | NCP unit with a population vaccinated more than three years ago (two-doses OCV campaign with a coverage for both round >70%) | No | Yes |
| High risk for extreme climate and weather conditions | NCP unit exposed to extreme climate and weather conditions | No | Yes |
| Complex humanitarian emergency | NCP unit located in an area under a complex humanitarian emergency | No | Yes |
| Unimproved water | NCP unit with more than 30% of the population using unimproved water facility type (calculated as sum of % of population with unimproved service level and % of population using surface water) OR more than 15% of the population using surface water | Does not meet any of the two criteria | Meets one or more criteria |
| Unimproved sanitation | NCP unit with more than 50% of the population using unimproved sanitation facility type (calculated as % of population with unimproved sanitation service level and % of population practicing open defecation) OR more than 30% of the population practicing open defecation | Does not meet any of the two criteria | Meets one or more criteria |
| Limited access to hygeine | NCP unit with more than 50% of the population with no handwashing facility on premises | Does not meet the criteria | Meets the criteria |
Calculate the vulnerability index for each NCP operational geographic unit by summing the scores of the component subindicators. These can be done in an unweighted format (i.e., where all vulnerability factors are considered equally), or in a weighted format (i.e., where a particular vulnerability factor is (or factors are) emphasized compared to other factors, usually by multiplying it by a larger number (e.g., multiplying one indicator by 3 if it is thought to be 3x more 'predictive' of the indicator)).
Unweighted
Vulnerability index = ∑(0/1 subindicators)
Weighted
Vulnerability index = ∑(0/1 subindicators*weighting factor))
To aid in implementation, you can use the Excel tool created by the Global Taskforce on Cholera Control (GTFCC)
Join the spreadsheet of your vulnerability index to the shapefile(s) of interest, using your GIS tool of choice (e.g., QGIS, ArcGIS). Then, map that shapefile of interest.
Conduct a stakeholder validation workshop for each NCP operational geographic unit. The exact setup of this workshop is outside of the scope of this how-to – please refer to your guidance document(s) for information.
Outputs: Vulnerability index, by each operational NCP geographic unit, and a map of all vulnerability indices by each NCP geographic unit.
Validation & QA: Ensure that you check for missing data (review the guidance document as appropriate)
Maintenance: Ensure that calculation is done every 5 years, whenever a new NCP dataset is done. As mentioned in the PAMI documentation, “as a general principle, PAMI analysis should be updated when a new version of an NCP is developed (typically every five years). Earlier updates may be considered if there are significant changes in [vulnerability factors].”
Risks & Considerations: Ensure that missing data is sufficiently dealt with before creating the index.
Audience: GIS analysts with access to risk surfaces and geospatial raster and vector datasets (e.g., climatological, elevation, hydrological)
Deploy time: Download of data should take ~1 day, deployment should take 1 day
Prerequisites: Have basic familiarity with usage and implementation of rasters and vector datasets; have basic familiarity with GIS-related programs
Data & Tools:
Steps:
| Disease | Datasets |
| Cholera (and potentially other water-borne diseases) | · Land surface temperature (e.g., MERRA-2) – warm temperatures (28-34C) facilitates reproduction of V. cholerae. In addition, increased temperature -> higher evaporation -> lower water levels (water bodies)
· Precipitation/climate datasets (e.g., IMERG, CHIRPS)- heavy precipitation can allow for flood conditions · Water abundance/flooding (measured by water surface reflectance) - after heavy rainfall or floods, low‐lying lands that lack proper drainage or sewer systems experience, water stagnation, hampering the sanitation system, and accelerating cholera outbreaks · Compromised WASH systems – poor sanitation systems are linked to cholera outbreaks · Elevation (DEM) - elevation areas are more flood‐prone, which increases the chances of contact between people and water contaminated with V. cholerae. · Population Density (e.g., WorldPop) - high population density puts extra pressure on existing WASH systems · High-resolution imagery (optional, but suggested) – this can help us to identify areas of interest |
| Malaria (and other vector-borne diseases) | · Average temperature – higher temperatures are suitable for mosquito reproduction
· Rainfall (mm, total) – higher rainfall means the potential for stagnant water · Relative humidity – higher humidity conditions are ideal for mosquito propogation · Total number of malaria cases · Surface water / surface reflection – may reflect presence of stagnant water · Soil moisture – indicative of suitable mosquito habitats · Elevation (DEM) – lower elevations are more suitable for mosquito breeding habitats · Land use/land cover data – grasslands and forests may have access · Normalized flooding index (e.g., MODIS) |
Outputs: Map of overlaid risk estimates on top of relevant geospatial datasets
Validation & QA: Ensure that all datasets are available for the right spatial extents and temporal periods of interest, and that all raster datasets are at the same resolution.
Maintenance: All datasets must be temporally and spatially relevant. Make sure that all of your climate and hydrology datasets are of the time period(s) of interest. DEMs and LULCs are less likely to be updated in a timely manner, so find the most relevant dataset for your purposes.
Risks & Considerations: As the emergency situation progresses, ensure that all climate and hydrology datasets are updated. Downloading datasets on a daily basis and substituting in datasets in your maps is suggested.
No-Code Analysis: Threshold Calculation
Audience: Surveillance Epidemiologists
Deploy time: <1 day for data acquisition, ~1-3 hours for calculations (tops)
Prerequisites: Knowledge of Microsoft Excel or another spreadsheet application; ability to get access to historical case datasets.
Data & Tools:
Steps:
Outputs: Alert and/or outbreak thresholds
Validation & QA: Ensure that you have the correct datasets for the correct time periods and spatial areas of interest.
Maintenance: Update these on a weekly basis, or as data is released out into the public. Make sure to keep track of both past and future occurrences of the thresholds, to see if a location has started to exceed a known threshold.
Risks & Considerations: Ensure that the same time period (across multiple years) and the same spatial areas are represented in your baseline datasets.
Audience: Surveillance epidemiologists, data/decision makers, staff in WASH or humanitarian clusters that need to collapse environmental factors with epidemiological datasets.
Deploy time: <1 day for data collation, 1-3 hours for risk score calculation and mapping
Prerequisites: Knowledge of Excel or a spreadsheet program; knowledge of a GIS program (if you need to map the risk score)
Data & Tools:
Steps:
| Domain | Example Variables |
| Disease burden | incidence, prevalence |
| Exposure | vector density, pollution |
| Demographics | age, population density |
| Mobility | commuting, travel |
| Environment | rainfall, temperature |
| Socioeconomic | poverty, sanitation |
| Healthcare access | hospitals, ICU beds |
| Immunity | vaccination coverage |
| Admin_boundary | zonstat_var1 | zonstat_var2 |
| Admin_A | 1.30983 | 34756 |
| Admin_B | 3.49586 | 10294 |
| Admin_C | 2.39475 | 1827 |
| Admin_D | 5.02948 | 40987 |
For the above example, the two scaling approaches are shown for the first variable:
| Admin_boundary | zonstat_var1 | minmax_var1 | minmax_var1_FORMULA | zscore_var1 | zscore_var1_FORMULA |
| Admin_A | 1.30983 | 0 | =(B2-MIN($B$2:$B$5))/(MAX($B$2:$B$5)-MIN($B$2:$B$5)) | -1.27001494 | =(B2-AVERAGE($B$2:$B$5))/STDEV.P($B$2:$B$5) |
| Admin_B | 3.49586 | 0.58769777 | =(B3-MIN($B$2:$B$5))/(MAX($B$2:$B$5)-MIN($B$2:$B$5)) | 0.318570166 | =(B3-AVERAGE($B$2:$B$5))/STDEV.P($B$2:$B$5) |
| Admin_C | 2.39475 | 0.2916726 | =(B4-MIN($B$2:$B$5))/(MAX($B$2:$B$5)-MIN($B$2:$B$5)) | -0.48160501 | =(B4-AVERAGE($B$2:$B$5))/STDEV.P($B$2:$B$5) |
| Admin_D | 5.02948 | 1 | =(B5-MIN($B$2:$B$5))/(MAX($B$2:$B$5)-MIN($B$2:$B$5)) | 1.433049791 | =(B5-AVERAGE($B$2:$B$5))/STDEV.P($B$2:$B$5) |
EPI_i=∑k=1nwkZikEPI_i = \sum_{k=1}^{n} w_k Z_{ik}EPIi=k=1∑nwkZik
where:
iii = spatial unit (county, grid cell, district)
kkk = risk factor
wkw_kwk = weight
ZikZ_{ik}Zik = standardized variable
For our hypothetical example here:
| Admin_boundary | minmax_var1 | minmax_var2 | equal_weighting_index | equal_weighting_index_formula | unequal_weighting_index* | unequal_weighting_index_formula |
| Admin_A | 0 | 0.84088 | 0.840884 | =C2+H2 | 0.672707 | =(0.2*C2)+(0.8*H2) |
| Admin_B | 0.5877 | 0.21622 | 0.803913 | =C3+H3 | 0.290512 | =(0.2*C3)+(0.8*H3) |
| Admin_C | 0.29167 | 0 | 0.291673 | =C4+H4 | 0.058335 | =(0.2*C4)+(0.8*H4) |
| Admin_D | 1 | 1 | 2 | =C5+H5 | 1 | =(0.2*C5)+(0.8*H5) |
* For this worked example, (0.2 weight is given to var1, 0.8 weight given to var2)
Outputs: A disease risk score for every geographic area of interest (either in a spreadsheet format only (if in Excel, not joined to a GIS program), or a shapefile with indices attached..
Validation & QA: Ensure that you handle areas with missing data with either an imputation of the average or eliminate those areas if most variables are not collected in that area.
Maintenance: Update these risk scores on a daily basis, or as you have specific factors change (e.g., climate, humidity).
Risks & Considerations: Ensure that you consult with your stakeholders to better understand what variables are most relevant and least relevant for your use case.
Audience: Surveillance epidemiologists (if GIS-savvy), data managers, or individuals with heavy data skills
Deploy time: 1 day to acquire datasets, 1 day to produce results
Prerequisites: GIS tool of interest, case-specific data with dates incorporated
Data & Tools:
Data & Tools:
This mechanism creates a prevalence raster, that spatially models on x, y, and date. However, it does NOT capture the relative effect of other covariates.
| x | y | date |
| x1 | y1 | date1 |
| … | … | … |
| xn | yn | daten |
Regardless of the coding platform used, the following steps can be conducted for a basic kriging interpolation for a single time period of interest:
Outputs: If using ordinary kriging, you will have a series of rasters that show examples by week. If using the Space-Time cube functionalities, A stack of rasters ("space time cube")
Validation & QA: Ensure that your data is clean and that you have aggregated data to the format that you are interested in using it.
Maintenance: Conduct this every time a dataset is updated.
Risks & Considerations: This may be harder to interpret. Depending on the procedure used, there may be statistical parameters that need to be utilized with which the data analyst may not be immediately familiar.
Audience: Surveillance epidemiologists (if GIS-savvy), data managers, or individuals with heavy data skills
Deploy time: 1 day to acquire datasets, 1 day to produce results
Prerequisites: Use of a GIS program, including Raster Calculator Functionalities
Data & Tools:
The creation of a spatial risk surface is very similar in the algebra compared to the risk score mentioned above. However, instead of the risk score being aggregated down to different administrative units, you are using data purely in the raster format, which requires Map Algebra and Raster Calculator functionalities.
Using each of your raster datasets, you then will create a normalized version of each raster input.
The easiest version to create is the min-max scaled dataset, but you can also use z-scaling using the same approach.
First, identify your metric of interest, and right click on the raster of interest to get the appropriate values:
Perform this normalization procedure for each raster dataset of interest.
Once all of the raster datasets have been normalized, you will then use a new instance of Raster Calculator in your GIS program of interest. However, instead of performing calculations within a raster, you will be performing calculations between the rasters.
Make sure that you have all of your rasters in the same spatial resolution (grid cell size) and same spatial extent.
Outputs: A disease risk raster surface
Validation & QA: Ensure that all datasets are in the right spatial resolution and extents. Make sure that you are typing all formulas correctly into the Raster Calculator feature in your GIS program. Ideally, you also want to check your predictions afterwards using an AUC curve, but this would be more necessary for more extensive procedures (e.g., XGBoost, random forest).
Maintenance: To be performed as new datasets are incorporated in your datasets of interest. A best practice is to create all of the static normalized rasters at one time, and only update the dynamic normalized rasters as the data is updated.
Risks & Considerations: Ensure that you use the correct rasters. Note that this map algebra procedure is a basic risk raster, but is the most explainable type of risk raster that can be calculated. Other procedures, such as random forest or XGBoost, can be utilized to create more extensive rasters, but they are less predictable and potentially involve the user knowing their risks.
