1 Getting MICS6 data

The Multiple Indicator Cluster Survey (MICS) is a national household survey implemented by countries under the programme developed by the United Nation’s Children’s Fund (UNICEF) to provide internationally comparable, statistically rigorous data on the situation of children and women. Among other topics included in the survey, MICS6 covers the Foundational Learning Skills (FLS) Module. This is a data collection tool that directly assesses whether children have acquired essential reading and mathematics skills at the early primary level. It focuses on children aged 7 to 14 and captures learning outcomes at grades 2 and 3.

To make the MICS6 microdata more accessible, this chapter of this guide shows how to download and unpack (the crucial!) the MICS6 for multi-country early foundational learning research. Generally, researchers can access the MICS6 microdata in two ways:

  1. UNICEF MICS: the original country survey files.
  2. IPUMS MICS: a custom extract with variables already harmonised across the surveys you select.
UNICEF MICS IPUMS MICS
What you receive Original country survey files A custom extract with harmonised variables across selected surveys
Best for Country-specific items; the FLS assessment; reproducing MICS indicators Multi-country work on demographic, household, and schooling variables
Preparation More: files must be unpacked and, for pooled work, harmonised Less: IPUMS supplies syntax that builds the extract
Software R or Stata: the .sav files open in either Stata to build the extract; then R or Stata for analysis
Single-country work Use this Adds little: IPUMS earns its keep across surveys

AFLEARN recommendation

If the analysis uses foundational reading, do not use the IPUMS reading variables for surveys that administered more than one reading passage. IPUMS keeps one passage and can drop a large share of children who completed the assessment. AFLEARN reconstructs reading outcomes from the complete UNICEF files.

The two routes complement each other: IPUMS for comparable background variables, UNICEF for every reading passage.

1.1 Getting UNICEF MICS files

UNICEF releases MICS microdata as SPSS (.sav) files. The files already carry variable and value labels, so you can open them in R or Stata without a separate codebook.

1.1.1 Download the country files

  1. Open https://mics.unicef.org/surveys.
  2. Log in if the site asks you to. Dataset downloads usually need an account.
  3. Filter to the surveys you want: round MICS6, then region or countries, data type MICS, and status Completed.
  4. Click Download MICS datasets. UNICEF packs every matching survey into one bulk zip, usually named MICS_Datasets.zip.

Create a project folder in your local computer, for example MICS project, and save MICS_Datasets.zip in it.

1.1.2 Unpack with AFLEARN’s scripts

The MICS_Datasets.zip bulk zip nests each country survey in further archives and wrapper folders, as shown in the illustration below. Unpacking a large number of country surveys by hand wastes time and invites path errors.

MICS_Datasets.zip
└── survey folder
    └── another archive
        └── wrapper folder
            ├── fs.sav
            ├── hh.sav
            └── ...

To avoid this, firstly download from AFLEARN’s repository Script.zip and unzip it into the same project folder, i.e., MICS project, so Script/ sits beside MICS_Datasets.zip. See blow:

MICS project/
├── MICS_Datasets.zip
└── Script/

Next, open R, set the MICS project folder as the working directory and source run-it.R as shown below:

setwd("C:/path/to/MICS project")
source("Script/run-it.R")

You have to ascertain that the working directory must be the folder that holds Script/. After a successful run, you will notice that there is now a Data/ folder with a UNICEF/ folder inside that carries all the unpacked and tidy country surveys you would have downloaded. Each survey then has a predictable path such as Data/UNICEF/BEN_2021_MICS6_v01_M/fs.sav.

MICS project/
├── MICS_Datasets.zip
├── Script/
└── Data/
    ├── UNICEF/
    │   ├── BEN_2021_MICS6_v01_M/
    │   │   ├── fs.sav
    │   │   ├── hh.sav
    │   │   ├── hl.sav
    │   │   └── ...
    │   └── ...
    └── AFLEARN Harmonised Data/
       

You will also notice a folder called AFLEARN Harmonised Data/. This folder carried the AFLEARN harmonised microdata file, which will be explained in depth in Chapter 3 on Harmonising MICS6 microdata.

Keep the original download

Keep the original MICS_Datasets.zip file. The clean country survey folders are for analysis, while the MICS_Datasets.zip file is the unchanged source.

After setting up, Understanding MICS6 data explains the survey files and the FLS module.

1.2 Getting IPUMS MICS files

IPUMS MICS lets you work with MICS microdata across countries without downloading every variable from every original survey microdata file. You select the unit of observation, country surveys and variables. IPUMS then supplies the original UNICEF source microdata together with syntax that harmonises those selections and builds a pooled microdata extract. For many multi-country research, this reduces data preparation. This includes the unpacking of the country surveys you previously saw in Getting UNICEF MICS data.

1.2.1 Download IPUMS MICS microdata

  1. Open IPUMS MICS https://mics.ipums.org/mics/
  2. Click SELECT DATA and choose a unit of observation. For the FLS module, use children aged 5–17.
  3. Select the MICS6 samples you need. Record the IPUMS sample identifiers, not only country names.
  4. Add variables to the Data Cart. Before you add one, check that it exists in your samples, that the universe matches, and that the codes are comparable. Keep identifiers, sample, and weight variables for linking and survey design.
  5. Create the extract. The download is source data plus Stata .do files. Run the main .do file to build the pooled dataset. An R workflow needs Stata once for this step; then read the .dta in R. See Opening an Extract.

One extract has one unit of analysis

Some household or parent items are already attached to the fs file while others need a link. See linking units of analysis.

IPUMS documents each variable’s universe, codes and questionnaire text. Comparisons go wrong when universes differ across surveys. See the IPUMS MICS FAQ.

1.2.2 Multi-passage reading

Do not use the IPUMS reading variables for a MICS6 survey that administered more than one reading passage.

In these surveys, IPUMS retains data for only one of the reading passages. Children assessed using another reading passage therefore appear to have no data on the IPUMS harmonised reading variables. This systematically excludes children according to the language or reading passage in which they were assessed and can make estimates of reading performance seriously misleading.

This issue is especially important in multilingual African countries. The FLS module allowed countries to adapt the reading assessment to local languages. The country-specific microdata files can therefore contain separate variables for different reading passages or languages, while the corresponding IPUMS harmonised reading variable represents only one of them.

Malawi MICS 2019–20 is the clear case. IPUMS records 1,720 children on FLWORDSATTEMPT variable (the English passage). The UNICEF microdata files show 3,970 who attempted Chichewa, of whom 130 also attempted English. Using IPUMS reading variables in this case drops most of the children who completed the reading assessment.

This limit applies to the harmonised reading variables in multi-passage surveys, not to the rest of IPUMS. To solve this issue, AFLEARN reconstructs the reading data from the original UNICEF microdata files, retaining all relevant passages and languages. All of this is shown in full detail in the Harmonising MICS6 microdata

Before you can reach that, Understanding MICS6 data explains the survey files and the FLS module in the next chapter.