AHEAD User Guide
Welcome
This reference guide describes the harmonization framework developed for the African Harmonized Early-Grade Assessments Dataset (AHEAD), which combines datasets from United States Agency for International Development (USAID)-funded Early Grade Reading Assessment (EGRA) and Early Grade Mathematics Assessment (EGMA) projects conducted across Africa. The data harmonization undertaken in this project is retrospective. Retrospective harmonization (also referred to as ex-post or output harmonization) involves harmonizing datasets after they have been collected. The goal is to develop a flexible, scalable, and analytically robust database that supports longitudinal, cross-sectional, and comparative analyses across diverse educational contexts.
Recognising that EGRA and EGMA assessments have been implemented over many years in different countries and programmes using different assessment instruments, languages, sampling designs, and variable naming conventions, this harmonization framework provides a systematic approach to identifying and standardising comparable measures while preserving the integrity and context of the original datasets.
This reference guide introduces the harmonized dataset and the long-term harmonization programme. It describes the structure and content of the pooled dataset; and provides guidance on variable naming conventions, assessment variables, identifiers, survey weights, missing data, and other key considerations for using the harmonized data.
0.1 Using this Guide
This user guide should be used alongside the complementary resources available through the DataFirst Open Data Repository and the AFLEARN website.
The harmonized dataset has been designed to preserve traceability to the original source data. Users who require additional variables, documentation, or project-specific information should consult the accompanying resources listed below.
- The DataFirst Open Data Repository hosts the pooled dataset and its structured metadata, including the DDI codebook (study- and variable-level documentation, labels, frequencies, project information).
- The linking-to-source file is provided alongside the pooled dataset; this file contains the actual mapping between harmonized and original identifiers. It uses the
school_idandstudent_idkeys, organized by country, project, and round, to let you join non-harmonized variables directly. - AHEAD Study Catalogue is the official registry of early-grade assessment surveys included in the AHEAD dataset. For each assessment survey, the catalogue provides country, year, project code, round/wave, languages, grades, study type, sub-task availability, survey-design setup (Stata
svysetand Rsvydesign), and guidance for linking harmonised records back to original source files.
Note: Understanding Projects and Constituent Studies
Throughout this guide, a distinction is made between projects and constituent studies. A project refers to a broader assessment programme, which may include one or more rounds of data collection. Each round is treated as a separate constituent study. Individual constituent studies are uniquely identified by the combination of country, project, and round, and this combination should be used when distinguishing between specific survey iterations. While harmonized variables share common names across projects, the values they contain remain specific to each constituent study and should be interpreted within their original survey context.