3 Harmonising MICS6 data

While the MICS6 questionnaires are largely standardised, there are important differences between countries. Variable names, education codes, languages, and omitted modules differ across surveys. Stacking fs.sav files does not give a comparable child-level file.

IPUMS MICS aligns much of the background data. It still keeps only one reading passage, so it is a poor base for foundational reading in multi-passage surveys (Getting IPUMS MICS files).

AFLEARN’s R script run-it.R unpacks the UNICEF zip and writes one child-level analysis file: shared names, selected household-list fields, and research-ready reading outcomes (passage length, scores, reading_status).

Data/AFLEARN Harmonised Data/mics6-fs-harmonised-v2.0.dta

plus the Excel map mics6-crosswalk.xlsx. Reading sits in that file. There is no second .dta.

If you do not yet have this file, go to Getting UNICEF MICS files. Put MICS_Datasets.zip beside Script/ in a project folder and run:

source("Script/run-it.R")

If you already have the file, you still merge only if:

  1. You need a UNICEF source variable the spec did not keep. Start from fs and add from hh or hl.
  2. You are analysing an IPUMS extract. Keep IPUMS as the analysis table and attach AFLEARN reading and FLS fields.

If the file already has what you need, skip the merge pages and go to the R or Stata tutorial.

3.1 The AFLEARN analysis file

This is the file run-it.R writes: one row per selected child in fs. Files, keys, and who is in fs sets out who that is. How foundational reading and numeracy are scored is in Reading and numeracy. If you do not yet have the file, start from Getting UNICEF MICS files.

prepare-mics-fs.R
        → Data/UNICEF/<survey>/
harmonize-mics-fs-v2.0.R     spec + crosswalk; already pulls selected hl fields
derive-mics-fl-vars-v1.0.R   passage length, scores, reading_status, max_reading_*
        → mics6-fs-harmonised-v2.0.dta

If you only unpacked the zip, source("Script/run-it.R") is enough. Re-run the two later scripts when you change the spec:

source("Script/harmonize-mics-fs-v2.0.R")
source("Script/derive-mics-fl-vars-v1.0.R")

3.1.1 Identifiers

AFLEARN Source / role
country_iso3, year Survey
HH1 / cluster Cluster (HH1)
HH2 / hhno Household within the cluster
LN / linech Selected child’s line number
sample IPUMS SAMPLE from the spec

Do not invent sample from country names. Use the value on the file and check it against your extract.

3.1.2 What is in the file

Area Examples
Survey design fsweight, fshweight
Geography urban/rural, region
Parents mother and father education
Schooling attendance, education level and grade
Child labour work, chores, hours
Parental involvement, functioning, discipline as fielded
Foundational learning numeracy items; reading scores and reading_status

Source names for every column are in mics6-crosswalk.xlsx (also in the guide repository). Look up the survey column before you pool a variable.

A shared name does not always mean a shared code. region, raw grades, and school_type stay country-specific. For education, use the harmonised levels (current_level_h, highest_level_h, mother_edu_h, father_edu_h) when you pool. current_grade is the year within that level: primary year 1 and secondary year 1 both carry the code 1.

The spec already attaches selected hl fields such as parent education and school type. Merge by hand only for variables that are not on this file (Adding variables from UNICEF files).

3.1.3 What derive added

harmonize-mics-fs-v2.0.R aligns names. derive-mics-fl-vars-v1.0.R then builds the reading fields researchers actually analyse:

Variable Meaning
passage_language, passage_length Language and word count of the main passage
words_att, words_incorrect Words attempted and missed
reading_score, reading_accuracy Words correct; share of passage length
read_comp_1–read_comp_5, read_comp_score Comprehension items and total
reading_skills Official foundational-reading classification
reading_status Where the child stopped, or the outcome reached
*B* / *C* and max_reading_* Later passages, and the best result across them

The official 90% + five-comprehension rule is in Reading and numeracy. A missing score is not one kind of missing: children leave the pathway at different points.

3.1.4 reading_status

Code Meaning
0 Not interviewed (interview_result > 1)
1 Age outside 7–14
2 Caregiver FLS consent not given
3 Child consent not given
4 Language mismatch
5 Child refuses the story
6 Failed practice
7 Attempted a passage (max_reading_score present)
8 Met 90% accuracy
9 Met foundational reading skills
10 Otherwise

Do not recode all non-attempts as children who cannot read

A child with no suitable booklet has not demonstrated the same thing as a child who attempted the passage and read no words correctly.

3.1.5 Checks

use "Data/AFLEARN Harmonised Data/mics6-fs-harmonised-v2.0.dta", clear

tab country_iso3 year
isid country_iso3 year HH1 HH2 LN

tab reading_status if age >= 7 & age <= 14
tab passage_language if country_iso3 == "MWI"
library(haven)
library(dplyr)

d <- read_dta("Data/AFLEARN Harmonised Data/mics6-fs-harmonised-v2.0.dta")

d %>% count(country_iso3, year)

d %>%
  count(country_iso3, year, HH1, HH2, LN) %>%
  filter(n > 1)

d %>%
  filter(age >= 7, age <= 14) %>%
  count(reading_status)

d %>%
  filter(country_iso3 == "MWI") %>%
  count(passage_language)

Malawi should show both assessment languages. If the key is not unique, stop before any merge.

3.2 Adding variables from UNICEF files

Decide what one row of the final dataset represents before you merge anything. For foundational learning, keep one row per selected child in fs.

The keys are in Files, keys, and who is in fs. The three common links are household (HH1 + HH2), the child’s own roster line (LN ↔︎ HL1), and another member such as the caretaker (FS4 ↔︎ HL1). UNICEF’s MICS6 Manual for Statistical Data Analysis describes the same identifiers.

The AFLEARN file already attaches selected hl fields (parent education, school type). Merge by hand only for variables that are not already on it.

Replace COUNTRY_SURVEY, variable_1, and variable_2 with the survey folder and the fields you need. If you work in Stata from UNICEF .sav files, import SPSS once and save temporary .dta versions.

3.2.1 Household variables onto fs

Many children can share a household; each HH1 + HH2 should appear once in hh.

use "Data/UNICEF/COUNTRY_SURVEY/fs.dta", clear

merge m:1 HH1 HH2 using ///
    "Data/UNICEF/COUNTRY_SURVEY/hh.dta", ///
    keepusing(variable_1 variable_2)

tab _merge
library(haven)
library(dplyr)

fs <- read_sav("Data/UNICEF/COUNTRY_SURVEY/fs.sav")
hh <- read_sav("Data/UNICEF/COUNTRY_SURVEY/hh.sav")

hh %>%
  count(HH1, HH2) %>%
  filter(n > 1)

hh_keep <- hh %>%
  select(HH1, HH2, variable_1, variable_2)

fs_merged <- fs %>%
  left_join(
    hh_keep,
    by = c("HH1", "HH2"),
    relationship = "many-to-one"
  )

3.2.2 The child’s own hl row

LN in fs is HL1 in hl. The selected-child link should be one-to-one.

import spss using ///
    "Data/UNICEF/COUNTRY_SURVEY/hl.sav", clear

keep HH1 HH2 HL1 variable_1 variable_2
rename HL1 LN
isid HH1 HH2 LN

tempfile hl_child
save `hl_child'

import spss using ///
    "Data/UNICEF/COUNTRY_SURVEY/fs.sav", clear

isid HH1 HH2 LN
merge 1:1 HH1 HH2 LN using `hl_child'
tab _merge
library(haven)
library(dplyr)

fs <- read_sav("Data/UNICEF/COUNTRY_SURVEY/fs.sav")
hl <- read_sav("Data/UNICEF/COUNTRY_SURVEY/hl.sav")

hl %>%
  count(HH1, HH2, HL1) %>%
  filter(n > 1)

hl_child <- hl %>%
  select(HH1, HH2, HL1, variable_1, variable_2)

fs_merged <- fs %>%
  left_join(
    hl_child,
    by = c("HH1", "HH2", "LN" = "HL1"),
    relationship = "one-to-one"
  )

3.2.3 Mother or caretaker

FS4 is another person’s line, not the child’s. Link HH1 + HH2 + FS4 to HL1.

import spss using ///
    "Data/UNICEF/COUNTRY_SURVEY/hl.sav", clear

keep HH1 HH2 HL1 variable_1 variable_2
rename HL1 FS4

tempfile hl_care
save `hl_care'

import spss using ///
    "Data/UNICEF/COUNTRY_SURVEY/fs.sav", clear

merge m:1 HH1 HH2 FS4 using `hl_care'
tab _merge
library(haven)
library(dplyr)

fs <- read_sav("Data/UNICEF/COUNTRY_SURVEY/fs.sav")
hl <- read_sav("Data/UNICEF/COUNTRY_SURVEY/hl.sav")

hl_care <- hl %>%
  select(HH1, HH2, HL1, variable_1, variable_2)

fs_merged <- fs %>%
  left_join(
    hl_care,
    by = c("HH1", "HH2", "FS4" = "HL1"),
    relationship = "many-to-one"
  )

3.2.4 Check the merge

A merge that completes without an error is not evidence that it is correct. Ask what one row represents in each file, whether the key is unique, and how many rows the analysis file should have afterwards.

tab _merge
* 1 = master only; 2 = using only; 3 = matched
* Do not drop if _merge != 3 until you know why rows did not match.
nrow(fs)
nrow(fs_merged)

fs_merged %>%
  count(HH1, HH2, LN) %>%
  filter(n > 1)

Adding an hh covariate does not change the weight. While the rows remain selected children, use fsweight (Files, keys, and who is in fs).

3.3 Attaching the AFLEARN file to IPUMS

Keep the IPUMS child extract for demographic, household, and schooling variables. Attach the AFLEARN file so reading includes every passage. The limit on IPUMS reading variables is in Getting IPUMS MICS files.

The v2 file already carries sample (IPUMS SAMPLE from the spec).

AFLEARN                 IPUMS
sample              <-> SAMPLE
cluster             <-> CLUSTER
hhno                <-> HHNO
linech              <-> LINECH

Do not guess IPUMS sample codes from country names

Use sample on the AFLEARN file and check it against SAMPLE in your extract. IPUMS identifiers distinguish survey year and, where needed, national and subnational samples.

Build the extract as children aged 5–17 and keep SAMPLE, CLUSTER, HHNO, and LINECH. See IPUMS MICS linking. Select the AFLEARN columns you need so names do not collide with IPUMS.

3.3.1 Check uniqueness

use "Data/ipums_mics_fs.dta", clear
duplicates report SAMPLE CLUSTER HHNO LINECH

use "Data/AFLEARN Harmonised Data/mics6-fs-harmonised-v2.0.dta", clear
duplicates report sample cluster hhno linech
library(haven)
library(dplyr)

ipums <- read_dta("Data/ipums_mics_fs.dta")
aflearn <- read_dta(
  "Data/AFLEARN Harmonised Data/mics6-fs-harmonised-v2.0.dta"
)

ipums %>%
  count(SAMPLE, CLUSTER, HHNO, LINECH) %>%
  filter(n > 1)

aflearn %>%
  count(sample, cluster, hhno, linech) %>%
  filter(n > 1)

If either check returns duplicates, stop before merging.

3.3.2 Join

Keep IPUMS as the analysis population. Not every fs child has a completed reading assessment; ages 5–6 and 15–17 are outside the FLS target, and eligible children can leave the pathway earlier. The question is whether the child record links, not whether every child has a score.

use "Data/AFLEARN Harmonised Data/mics6-fs-harmonised-v2.0.dta", clear

rename sample SAMPLE
rename cluster CLUSTER
rename hhno HHNO
rename linech LINECH

isid SAMPLE CLUSTER HHNO LINECH

tempfile aflearn
save `aflearn'

use "Data/ipums_mics_fs.dta", clear

merge 1:1 SAMPLE CLUSTER HHNO LINECH using `aflearn'
tab _merge
library(haven)
library(dplyr)

ipums <- read_dta("Data/ipums_mics_fs.dta")
aflearn <- read_dta(
  "Data/AFLEARN Harmonised Data/mics6-fs-harmonised-v2.0.dta"
)

analysis <- ipums %>%
  left_join(
    aflearn,
    by = c(
      "SAMPLE" = "sample",
      "CLUSTER" = "cluster",
      "HHNO" = "hhno",
      "LINECH" = "linech"
    ),
    relationship = "one-to-one"
  )

Do not drop unmatched records until you know why they did not match. Once AFLEARN reading is on the file, drop or rename the IPUMS reading variables so you do not use them by accident.

3.3.3 After the merge

Use the IPUMS age variable from your extract (often AGE). Malawi must show both assessment languages.

tab SAMPLE
tab reading_status if AGE >= 7 & AGE <= 14
tab SAMPLE passage_language
analysis %>%
  count(SAMPLE)

analysis %>%
  filter(AGE >= 7, AGE <= 14) %>%
  count(SAMPLE, reading_status)

analysis %>%
  count(SAMPLE, passage_language)

If the study uses only single-passage countries, the AFLEARN reading fields are still the simpler common definition: the same reading_status, passage length, and accuracy rule in every supported survey.