4 Context: demographics, survey and study designs

The pooled dataset contains both harmonized EGRA and EGMA assessment variables and a set of harmonized contextual variables describing pupils, schools, survey design, and study characteristics.

While the assessment variables are discussed later in this document, this section focuses on the contextual variable groups. These variables provide the information required to identify individual projects, account for differences in sampling and assessment contexts, and support appropriate comparative analyses.

The following subsections provide guidance on their interpretation and use, while detailed variable labels, coding schemes, and frequencies remain available in the accompanying DDI codebook.

4.1 Demographics and pupil context

The harmonized dataset contains demographic and pupil context variables describing learner characteristics, assessment context, and identifiers.

Variable Description
gender Pupil gender
grade Grade / class level
age Pupil age (continuous)
location Urban / rural
lang_home Home language
lang_assessment Language of assessment
lang_instruction Language of instruction
lang2_assessment Second language of assessment
year, month Assessment year and month
school_id, student_id School and pupil identifiers (unique within each country, project and round combination)
consent Consent / assent flag

Table 2: Demographic variables in the dataset

4.1.1 Important considerations

Pooling across studies: Use the harmonized country variable for country-level analyses. Individual constituent studies are uniquely identified by a combination of country, project, and round, which should be used when distinguishing between specific survey iterations.

Sparsity: Not every study collected every contextual variable. Before conducting pooled analyses, users should consult the interactive data browser on the AFLEARN website.

Grade: Although grade values have been harmonized to a common set of labels, the distribution of grades varies across projects or individual studies. Users should make comparisons with reference to each project’s stated target population and design. Note that the same grade label does not imply equivalent curricula or learning expectations across countries.

Age: Age is retained as a continuous variable wherever possible. Missing or non-response values follow the harmonized missing-value conventions described in the metadata record.

Location: Urban/rural coding is harmonized where collected. However, not all projects distinguish location.

Identifiers: The harmonized variables school_id and student_id are copied directly from the original datasets and are meaningful only within their individual studies. They should not be used to link observations across studies or survey rounds. The pooled dataset includes the harmonized identifier harm_student_id, described in Identifiers.

4.2 Survey design, weights, and geography

The dataset includes a set of variables describing the sampling design, geographic coverage, and survey structure of each project. These variables enable users to correctly account for survey design when conducting weighted analyses and to identify the geographic context of each project.

Variable(s) Description
country Country
year Year
admin_div1admin_div3 Administrative division codes are provided for up to three hierarchical levels. The administrative units represented at each level vary by country—for example, provinces and counties in Kenya, and divisions and subdivisions in the DRC.
admin_div1_typeadmin_div3_type Name of each administrative level in this study (e.g. province, county, division, region, district) — separates the code from what the level is called.
sample_design_stages Number of sampling stages in the survey design.
stage1stage4 Sampling unit (SU / cluster) at each stage.
strata1strata4 Stratification grouping at each stage.
fpc1fpc4 Finite-population correction (FPC) factor at each stage.
wt1wt4 Stage weights
wt_final Final weight — use this for design-based estimation.

Table 3: Harmonized survey-design and geography variables

4.2.1 Important considerations

Harmonized names, project-specific contents: Survey design variables share common names across projects, but values are individual study-specific.

Within-survey rule: Design variables and weights must be used one study at a time. Never declare a survey design on the full harmonized file.

Multistage design variables: The number of applicable stages, strata, fpc, and stage-weight variables depends on the value of sample_design_stages. Variables for stages beyond the number used in a project’s survey design will not apply. For design-based analysis, use wt_final rather than the individual stage weights.

Geography (admin_div*): Administrative divisions are harmonized to admin_div1admin_div3. Companion admin_div*_type fields document the local label (province, county, region, district, etc.) because administrative levels differ across countries.

No pooled representativity: Pooled descriptive statistics that ignore survey design are not nationally representative of any single country. Details of each study’s sampling design are available in the interactive browser.

Per-study survey set-up: The correct weighting and clustering specification for each study is listed in the EGRA/EGMA Project and Coverage Explorer. Do not apply one design to the entire pooled file.

Known gaps: Some studies lack weights or complete design documentation; see the Release Notes for the current release for known gaps.

4.3 Study Design

These concepts describe how a study was designed and fielded. Full project narratives and sampling descriptions, collection occasions and links to original files are in the interactive browser on the AFLEARN website.

Concept Meaning
study_type Classifies the unit tracking structure of the study. Cross-sectional (cross_sectional) = independent samples drawn at each round. School panel (school_panel) = the same schools tracked across rounds. Pupil panel (pupil_panel) = the same pupils tracked across rounds.
study_design Describes the methodological design of the study. Pre–post (pre_post) = measures taken before and after an intervention without a control group. quasi_experimental = non-randomised comparison group. RCT = randomised controlled trial. Descriptive = no intervention, monitoring or baseline only. Unknown = design not ascertainable from available documentation.
round An integer indicating the sequential data collection occasion within a given project (1, 2, 3, ...), assigned in chronological order.
Treatment indicator Identifies whether a record belongs to a treatment or control group. Control = in control group. Treatment = in treatment group. Treatment2 = in 2nd treatment group. Treatment3 = in 3rd treatment group. Not applicable = study has no intervention.
Level of randomization The unit at which treatment and control groups were randomly assigned in experimental studies. Used to determine the appropriate level at which to cluster standard errors in treatment effect estimation.

Table 4: Study design variables in the dataset