4 Context: demographics, survey and study designs
The pooled dataset contains both harmonized EGRA and EGMA assessment variables and a set of harmonized contextual variables describing pupils, schools, survey design, and study characteristics.
While the assessment variables are discussed later in this document, this section focuses on the contextual variable groups. These variables provide the information required to identify individual projects, account for differences in sampling and assessment contexts, and support appropriate comparative analyses.
The following subsections provide guidance on their interpretation and use, while detailed variable labels, coding schemes, and frequencies remain available in the accompanying DDI codebook.
4.1 Demographics and pupil context
The harmonized dataset contains demographic and pupil context variables describing learner characteristics, assessment context, and identifiers.
| Variable | Description |
|---|---|
gender |
Pupil gender |
grade |
Grade / class level |
age |
Pupil age (continuous) |
location |
Urban / rural |
lang_home |
Home language |
lang_assessment |
Language of assessment |
lang_instruction |
Language of instruction |
lang2_assessment |
Second language of assessment |
year, month |
Assessment year and month |
school_id, student_id |
School and pupil identifiers (unique within each country, project and round combination) |
consent |
Consent / assent flag |
Table 2: Demographic variables in the dataset
4.1.1 Important considerations
Pooling across studies: Use the harmonized country variable for country-level analyses. Individual constituent studies are uniquely identified by a combination of country, project, and round, which should be used when distinguishing between specific survey iterations.
Sparsity: Not every study collected every contextual variable. Before conducting pooled analyses, users should consult the interactive data browser on the AFLEARN website.
Grade: Although grade values have been harmonized to a common set of labels, the distribution of grades varies across projects or individual studies. Users should make comparisons with reference to each project’s stated target population and design. Note that the same grade label does not imply equivalent curricula or learning expectations across countries.
Age: Age is retained as a continuous variable wherever possible. Missing or non-response values follow the harmonized missing-value conventions described in the metadata record.
Location: Urban/rural coding is harmonized where collected. However, not all projects distinguish location.
Identifiers: The harmonized variables school_id and student_id are copied directly from the original datasets and are meaningful only within their individual studies. They should not be used to link observations across studies or survey rounds. The pooled dataset includes the harmonized identifier harm_student_id, described in Identifiers.
4.2 Survey design, weights, and geography
The dataset includes a set of variables describing the sampling design, geographic coverage, and survey structure of each project. These variables enable users to correctly account for survey design when conducting weighted analyses and to identify the geographic context of each project.
| Variable(s) | Description |
|---|---|
country |
Country |
year |
Year |
admin_div1–admin_div3 |
Administrative division codes are provided for up to three hierarchical levels. The administrative units represented at each level vary by country—for example, provinces and counties in Kenya, and divisions and subdivisions in the DRC. |
admin_div1_type–admin_div3_type |
Name of each administrative level in this study (e.g. province, county, division, region, district) — separates the code from what the level is called. |
sample_design_stages |
Number of sampling stages in the survey design. |
stage1–stage4 |
Sampling unit (SU / cluster) at each stage. |
strata1–strata4 |
Stratification grouping at each stage. |
fpc1–fpc4 |
Finite-population correction (FPC) factor at each stage. |
wt1–wt4 |
Stage weights |
wt_final |
Final weight — use this for design-based estimation. |
Table 3: Harmonized survey-design and geography variables
4.2.1 Important considerations
Harmonized names, project-specific contents: Survey design variables share common names across projects, but values are individual study-specific.
Within-survey rule: Design variables and weights must be used one study at a time. Never declare a survey design on the full harmonized file.
Multistage design variables: The number of applicable stages, strata, fpc, and stage-weight variables depends on the value of sample_design_stages. Variables for stages beyond the number used in a project’s survey design will not apply. For design-based analysis, use wt_final rather than the individual stage weights.
Geography (admin_div*): Administrative divisions are harmonized to admin_div1–admin_div3. Companion admin_div*_type fields document the local label (province, county, region, district, etc.) because administrative levels differ across countries.
No pooled representativity: Pooled descriptive statistics that ignore survey design are not nationally representative of any single country. Details of each study’s sampling design are available in the interactive browser.
Per-study survey set-up: The correct weighting and clustering specification for each study is listed in the EGRA/EGMA Project and Coverage Explorer. Do not apply one design to the entire pooled file.
Known gaps: Some studies lack weights or complete design documentation; see the Release Notes for the current release for known gaps.
4.3 Study Design
These concepts describe how a study was designed and fielded. Full project narratives and sampling descriptions, collection occasions and links to original files are in the interactive browser on the AFLEARN website.
| Concept | Meaning |
|---|---|
study_type |
Classifies the unit tracking structure of the study. Cross-sectional (cross_sectional) = independent samples drawn at each round. School panel (school_panel) = the same schools tracked across rounds. Pupil panel (pupil_panel) = the same pupils tracked across rounds. |
study_design |
Describes the methodological design of the study. Pre–post (pre_post) = measures taken before and after an intervention without a control group. quasi_experimental = non-randomised comparison group. RCT = randomised controlled trial. Descriptive = no intervention, monitoring or baseline only. Unknown = design not ascertainable from available documentation. |
round |
An integer indicating the sequential data collection occasion within a given project (1, 2, 3, ...), assigned in chronological order. |
| Treatment indicator | Identifies whether a record belongs to a treatment or control group. Control = in control group. Treatment = in treatment group. Treatment2 = in 2nd treatment group. Treatment3 = in 3rd treatment group. Not applicable = study has no intervention. |
| Level of randomization | The unit at which treatment and control groups were randomly assigned in experimental studies. Used to determine the appropriate level at which to cluster standard errors in treatment effect estimation. |
Table 4: Study design variables in the dataset