4 Survey design, weights and standard errors
MICS6 is a complex stratified sample survey. An analysis needs to account for the survey design to produce correct population estimates and standard errors. For the fs microdata file, three pieces of information are crucial:
fsweight sampling weight for the selected child
PSU primary sampling unit
stratum sampling stratum
MICS6 does not draw children independently from a simple random sample. Households are selected through a stratified, clustered samplen design. Children living in the same sampled cluster can be more similar to one another than children selected independently from across the country. If the analysis ignores that structure, the estimated standard errors, confidence intervals and p-values can be incorrect. Weight, in general, solve a different problem. They simply account for unequal selection probabilities and survey adjustments. They do not, by themselves, tell the software (R or Stata) how observations are clustered or stratified. A correct design-based analysis therefore needs the weight, the primary sampling unit or cluster, and the stratum.
The basic rule
A weighted mean with an ordinary standard error is not a design-based estimate. Tell the software the weight, the cluster, and the stratum together.
Always check the exact design variable in the country dataset and survey documentation.
4.1 Which MICS weight
MICS attaches a weight to each questionnaire population.
| File / population | Weight |
|---|---|
| Household | hhweight |
| Women | wmweight |
| Men, where administered | mnweight |
| Children under five | chweight |
| Selected children aged 5–17 | fsweight |
The FLS module is administered within the questionnaire for children aged 5 - 17, for whihc one eligible child is randomly selected per household. A child in a household containing several eligible 5-17-year-olds has a lower probability of being selected for fs than an otherwise comparable child who is the only eligible child in their household. The fsweight accounts for the selected-child sampling process.
Perhaps more important is that the weight follows the analysis population, not the source of each variable. Suppose you start with fs and merge household wealth from hh. The unit of observation is still the selected child. Thus, always use fsweight, and not hhweight. Adding household wealth does not turn the analysis into a household analysis. Ask “what does one row represent?”. That question is more useful than asking which file a covariate originally came from.
It is common that the fs microdata file contain fshweight variable. This is labelled as the children 5-17 household sample weight. This is different from the fsweight. For ordinary child-level analyses of the randomly selected 5-17-year-old, including FLS outcomes, use fsweight. The latter weight variable incorporates the selection of one child from the eligible 5-17-year-olds in the household, whereas the household-level weight,fshweight, does not represent each selected child with that within-household selection adjustment.
4.2 Creating a survey design for one country survey
For one MICS6 survey the selected-child design is:
weight = fsweight
PSU = PSU
stratum = stratum
4.2.1 Check the design variables
Restrict to one survey first. Then count missing weights, PSUs, and strata, and how many PSUs sit in each stratum.
library(haven)
library(dplyr)
d = read_dta("Data/AFLEARN Harmonised Data/mics6-fs-harmonised-v2.0.dta") |>
filter(country_iso3 == "GHA") # subset to Ghana survey
d |>
summarise(
n = n(),
missing_weight = sum(is.na(fsweight)),
missing_psu = sum(is.na(PSU)),
missing_stratum = sum(is.na(stratum)),
n_psu = n_distinct(PSU),
n_strata = n_distinct(stratum)
)4.3 More than one survey, and the standard error
Appending surveys raises two questions: 1. How should the software see the design? 2. What population does a pooled number describe?
Making the primary sampling unit (PSU) codes unique answers the first. It does not answer the second.
4.3.1 Make PSU and stratum unique
The same numeric PSU and stratum appear in more than one survey. They are not the same sampling units. Build identifiers from country_iso3 + year. For an IPUMS extract, use SAMPLE. Tunisia 2018 and Tunisia 2023 are two surveys, thus country alone is not enough.
Always make both PSU and stratum survey-specific.
library(dplyr)
library(survey)
d <- d |>
mutate(
survey_id = interaction(country_iso3, year, drop = TRUE),
psu_pool = interaction(survey_id, PSU, drop = TRUE),
stratum_pool = interaction(survey_id, stratum, drop = TRUE)
)
des = d |>
as_survey_design(
ids = ~psu_pool,
strata = ~stratum_pool,
weights = ~fsweight,
data = d,
nest = TRUE
)If you are using HH1 in place of PSU, nest HH1 inside the same survey identifier.
4.3.2 What a pooled average means
For most comparisons in this guide, estimate by country, then plot or tabulate those estimates.
A single mean across all countries is a different estimand. Released fsweight values are normalised within each survey. Appended as they are, they do not produce a population-weighted regional figure. An equal-country average is another choice: estimate each country with its own design, then average those estimates with equal country weights, and say so.
Pooled regressions need the same decision: what the weights imply about each country’s contribution, whether country fixed effects belong, and whether the coefficient is a within-country or between-country relationship. There is no automatic MICS pooled weight.
4.3.3 Three ways to compute an average
| Approach | Point estimate | Standard error |
|---|---|---|
| Unweighted mean | Ignores unequal selection | Treats children as a simple random sample |
| Weighted mean, ordinary SE | Uses fsweight |
Still ignores clustering and strata |
| Survey-weighted mean, design-based SE | Uses fsweight |
Uses PSU and stratum |
Use the third!
4.3.4 Before you report
- One row is a selected
fschild. Merginghhorhldid not change the weight. - The weight is
fsweight, notfshweightorhhweight. - The design names the PSU (or
HH1), the stratum, and the weight together. - Ages 7–14 (or any other domain) were applied with
subpop()/subset(), not by dropping rows first. - Appended surveys have survey-specific PSU and stratum identifiers.
- A pooled number has an explicit estimand.
- The estimate is published with a confidence interval and an unweighted
N.