Exposure, Not Age: Why Multilingual Test Norms Need a Third Dimension
Age is a proxy for how much development a child has had. For multilingual children, it stops working, and English exposure explains more of the variance than age ever did.
The assumption hiding inside every set of test norms
Every norm-referenced test rests on an assumption so basic that it is almost never stated: that the children in the comparison group are the examinee’s true peers.
Peer, in practice, means same age. We stratify by gender, socioeconomic status, geography, race and ethnicity, but age is the variable doing the real work, because age is a proxy. It stands in for how much development a child has had. A ten-year-old has had ten years of exposure to the world, and so has every other ten-year-old, and therefore comparing them is fair.
For monolingual children, this is an excellent proxy. Age correlates nearly perfectly with how long a child has been immersed in the language the test is written in, because that immersion has been running continuously, twenty-four hours a day, since birth.
For multilingual children, it fails. Not subtly: comprehensively.
A ten-year-old whose life has been 60% English and 40% Spanish has not had ten years of English development. They have had roughly six. Comparing their performance to that of a ten-year-old monolingual is not comparing peers. It is comparing a child with six years of English to a child with ten and calling the difference ability.
“Age correlates perfectly with how old a person is and how much language they have been exposed to, but only if there has been one and only one language present.”
Sam, verbatim
Everything that follows is an attempt to take that sentence seriously.
Exposure explains more variance than age
The strong version of this claim, and the one worth defending: for multilingual examinees, English exposure accounts for more variance in test performance than age does.
This is not a marginal statistical observation. It reorders the entire logic of norming. If exposure is the dominant variable and we are controlling only for age, then we are controlling for the weaker of the two and treating the stronger one as noise, or worse, as signal about ability.
It also has an ironic implication for the vocabulary the field uses. Practitioners speak of “corrected” or “adjusted” scores when accounting for language background, as though the child’s obtained score were a flawed reading in need of repair.
“Their use gives the impression that the obtained score is not the actual or true estimate of the child, but is one that must be ‘fixed’ by considering other variables. It’s a power move. As I’ve pointed out, English exposure accounts for more variance than age, so in a sense, any correction we are making should be centered on age, not exposure!”
Sam, verbatim
This is not a quibble about terminology. “Corrected” carries a history: it implies deviation from a norm that is itself unquestioned, and it echoes earlier attempts at score adjustment that the field ultimately rejected. A naming convention that treats every exposure level as simply a different comparison group, rather than one privileged group plus corrections, is doing real conceptual work.
Quantity, not quality
The natural objection: exposure is not one thing. Two children with 30% English exposure may have had radically different experiences of it. One spoke English with educated adults; the other with playground friends. Surely quality matters.
Of course it does. The question is whether it matters enough to model.
The argument for setting quality aside rests on three points.
First, quantity is the dominant term. It is not possible for the average multilingual child to catch up to same-age monolingual peers in English development within the K–12 span, regardless of the intensity or quality of instruction. The reason is arithmetic: the monolingual child is accumulating English twenty-four hours a day, every day, measured in years. Quality differences operate at the margins of a gap that quantity created.
Second, quantity is zero-sum and quality is not. This is the property that makes the whole design work. A child cannot be exposed to two languages simultaneously in the same moment; acquisition shifts between them. If a child has spent 38% of their life hearing and processing Spanish, they have spent 62% with English. The two are complements and must sum to 100%. Quality has no such constraint (a child can have excellent exposure in both languages, or poor exposure in both), which means it cannot be inferred, only measured, and measuring it well requires an intake instrument no practitioner will tolerate.
Third, quality differences wash out in a properly constructed sample. Let quality vary randomly across the norm group. Children with rich exposure will perform slightly above the mean for their exposure level; children with impoverished exposure slightly below. The mean remains the appropriate expected performance for that quantity of exposure. This is exactly how norming handles every other unmeasured source of variance, and there is no principled reason to treat quality differently.
There is also precedent. The Ortiz PVAT asks a single question (at what age did the examinee first begin to learn English actively?), divides that by chronological age, and derives lifetime exposure from it. One question. No weighting, no settings inventory, no quality index. The resulting performance curves for low, moderate, and high exposure separate cleanly, unsmoothed.
“If a single question, like at what age were you first exposed to English, is sufficient with which to derive multilingual norms on a test of actual language ability, why then try and be more precise when we don’t know if there would necessarily be any added benefit?”
Sam, verbatim
Measuring exposure without exhausting the parent
A single question works. A somewhat better estimate is available for very little additional cost, and the design constraint is worth stating explicitly: the instrument has to be answerable by a parent in a few minutes, or by an examiner during intake without derailing the session. Examiners want to get to the testing. A twenty-item language-history survey will be skipped, guessed at, or filled in from the file.
An approach that respects that constraint uses developmental stages rather than settings:
- Early Childhood (birth through age 3): the home language environment, where basic interpersonal communicative skills begin.
- Preschool (age 4 through age 5): conversational language acquisition, often the first sustained exposure outside the home.
- Early School Years (age 6 through age 11): the transition from conversational language to cognitive academic language proficiency.
- Later School Years (age 12 through college): continued academic language development.
These divisions are not arbitrary. They track the well-documented sequence of first- and second-language acquisition, which gives the splits a theoretical foundation rather than a convenient one.
For each stage the respondent moves a single slider indicating the proportion of English exposure during that period. Because the slider is a proportion, it is inherently zero-sum: a child exposed to two languages from birth cannot be at 100% in both, and the instrument forces that recognition, which is itself clarifying for many families.
The computation is straightforward. Multiply each stage’s proportion by the number of months in that stage, sum, divide by the child’s chronological age in months. A child of 10 years 2 months (122 months) has completed the first two stages (48 months and 24 months) and is 50 months into the third. Stage proportions yielding 12, 12, and 45 months of English give 69 months total, or a Language Exposure Index of 57%.
Four questions at most, and fewer for a younger child who has not reached the later stages. No overlapping settings that can sum past the child’s actual age. No quality judgments. An estimate substantially more nuanced than a single age-of-first-exposure question, at almost no additional burden.
The three-dimensional norm plane
Here is the structural argument, and it is the part most likely to be misunderstood as a set of separate “special” norms.
It is tempting to describe a system like this as having two normative samples: a monolingual sample and a multilingual sample. That framing is wrong, and it invites exactly the criticism it should avoid: that multilingual children are being measured against a lesser or secondary standard.
The correct picture is a single sample occupying three dimensions: age × language exposure × ability.
Within that space, monolingual performance is not a separate sample. It is the edge of the plane: the line where exposure equals 100%. Immediately adjacent sits the line for 99% exposure, then 98%, continuing down to 1%. The monolingual group is a special case of the general model, not a different model.
“Instead of envisioning and presenting this as us having two normative samples, one MON and one MUL, we should always refer to it as just one sample that spans the entire linguistic range.”
Sam, verbatim
Two consequences follow, one conceptual and one practical.
Conceptually, it removes the asymmetry. Nobody is being corrected toward anybody. A child at 45% exposure is compared to other children of the same age at approximately 45% exposure: true peers, in the sense the whole enterprise of norming requires. A child at 100% is compared to other children at 100%. Neither comparison is privileged.
Practically, the sample is genuinely one sample and can be described as such. The total N is the total N.
Why the general population sample should not include English learners
This is the most counterintuitive recommendation in the design, and it is the one most likely to be challenged, so it needs the fullest defense.
Standard practice would include English learners in the general population norm group at their population proportion (roughly 20–22%) on representativeness grounds. It seems obviously correct. The sample should look like the country.
The argument against it is that ELL status is categorically unlike the other stratification variables. Gender, SES, and geography are demographic descriptors that we include because a sample unrepresentative of the population is a biased sample. Language exposure is not a descriptor. It is a direct determinant of performance on the measure itself, and it is being explicitly modeled as its own dimension. Including it a second time, unmodeled, inside the general population group does not improve representativeness. It reintroduces the confound the design exists to remove.
The consequence is concrete and severe. English learners have, by definition, less English development than their monolingual age peers. Mix them into a general population sample without controlling for exposure and the group mean drops. Because language background correlates strongly with race and ethnicity, any subsequent analysis of performance differences across racial and ethnic groups will detect a difference, and that difference will be language exposure wearing a demographic label. With sample sizes in the thousands, even trivial differences reach statistical significance.
That outcome would destroy the single most valuable fairness claim available. On the Ortiz PVAT, monolingual norms were constructed strictly (children with parents who spoke another language were deliberately excluded) precisely to eliminate language variance from that sample. The result was that on a direct measure of language, there was no statistically significant difference in performance as a function of race or ethnicity among any group. No other test in the field’s history has demonstrated that, and it happened specifically because language difference was not permitted to contaminate the monolingual comparison group. For the multilingual norms, race and ethnicity were not used as stratification variables at all; heritage language spoken was used instead, and no differences appeared among speakers of any heritage language.
“Monolinguals would always be compared to other monolinguals so there’s nothing wrong with that and multilinguals would always be compared to other multilinguals and there’s nothing wrong with that… ELL status is not a variable that creates any fairness by having it represented within a monolingual sample; it’s a variable that directly affects performance, unlike hair color, handedness, or favorite flavor of ice cream, and merits its own special handling.”
Sam, verbatim
One legitimate exception: instructional questions
There is a case for a combined general-population-plus-ELL comparison, and it is worth being precise about where it applies.
Diagnostic questions (does this child have a disability?) require true peer comparison without exception. Any other comparison is discriminatory and undermines the assumption of comparability that gives the score meaning.
Instructional questions are different. How far behind grade-level expectations is this child performing right now? is a legitimate question with a legitimate answer, and that answer requires comparison to the general population regardless of language background. School accountability runs on English achievement. A teacher planning intervention intensity needs to know the actual gap.
So: for achievement tests, and for direct measures of language and language-related abilities, offering both comparisons is appropriate and useful, provided the reporting makes unmistakably clear which question each one answers. The PVAT handles this by generating an “Instructional Level” narrative describing performance relative to the general population without reporting a score. That is deliberate, to keep a number that is discriminatory in a diagnostic context from circulating where it can be misused.
For cognitive tests that are not academic or language-related, there is no such exception. Exposure-controlled norms only.
What examiners should expect to see
One practical note, because it predicts a specific and misleading reaction.
When exposure-controlled norms are used, scores for multilingual children go up. Examiners accustomed to seeing 55, 60, and 65 for these students begin seeing 92, 93, and 100. The first instinct is that the test must be too easy.
It isn’t. Those scores are what emerge when a child is compared to genuine peers. The familiar depressed scores were never measurements of ability; they were measurements of a mismatch between the child and the comparison group. In eight years of the PVAT being in the field, no one has demonstrated that either the monolingual or multilingual norms are inaccurate.
When true peers are compared to true peers, the average score is 100. For monolinguals we have always taken that for granted. It should be equally unremarkable for everyone else.
Norming design and why ELLs should not enter the GP sample · Initial question on GP stratification · Quantity vs. quality of exposure · Weighting, PVAT precedent, and simplification · Developmental-stage survey design and LEI computation · Life stages and theoretical foundation · Single sample across the linguistic range · Norm sample as a three-dimensional plane · “Corrected” vs. exposure-based naming · Diagnostic vs. instructional norms
Enter your email to unlock the rest of this whitepaper
We’ll email you occasionally about our research, and store your address as described in our Privacy Policy. Unsubscribe anytime.
Working on norming or psychometrics for multilingual populations? We’d welcome your review.
Get involved →