Assessment
Last updated
Adaptive Testing and Validity: Keeping Measurement Sound
Computer-adaptive assessment can shorten tests and sharpen precision, but it must be designed with validity in mind. This article outlines how item selection, calibration, and scoring interact so that adaptive instruments remain interpretable and fair for high-stakes use.
Introduction
Computer-adaptive testing (CAT) selects each item partly from a candidate’s evolving ability estimate. It can concentrate measurement where it is most informative and avoid asking many questions that are far too easy or difficult. That efficiency is valuable, but it is not itself evidence of validity. The central question is still whether the resulting scores support the interpretations and decisions for which the assessment is used. The Standards for Educational and Psychological Testing treats validity as an evidence-based argument about intended score use, not as a permanent property of a test format.
Start with the Intended Interpretation
A CAT designed to estimate broad language proficiency has a different evidential burden from one designed to classify a learner as ready for a particular course. Developers should state the construct, target population, reporting scale, consequences, and acceptable uncertainty before choosing an algorithm. They then need evidence that the adaptive route still samples the required content and cognitive processes. Two candidates may see different items, but their tests must remain comparable representations of the same construct. A precise score from a narrow or distorted sample is not a valid score.
Item Selection and Calibration
Adaptive algorithms rely on a calibrated item bank. Items must be placed on a common scale with a suitable item response theory model, and the model must fit the population and response behaviour well enough for the intended use. Poorly estimated difficulty or discrimination parameters can influence which item appears next and then distort the final estimate. Calibration is therefore an operational process, not a one-off event: new items need pretesting, bank updates need linking, and parameter drift needs monitoring. Elements of Adaptive Testing covers item-pool design, parameter estimation, model fit and operational implementations in detail.
Content Coverage, Exposure and Security
Selecting only the statistically most informative item can overuse a small part of the bank or neglect content that matters to the blueprint. Operational CAT therefore needs constraints for domains, skills, stimulus types and other specifications, together with exposure controls that reduce repeated use of high-demand items. These controls create trade-offs: stronger security or tighter content balancing may reduce statistical information slightly, while a weak constraint system can make administrations less comparable. Simulations should test these trade-offs across realistic ability distributions before launch, followed by monitoring of live item exposure and blueprint completion.
Scoring and Interpretability
CAT scores are usually reported on the item bank’s scale, but the estimator, starting rule and stopping rule affect their behaviour. A test might stop after a fixed number of items, when standard error falls below a threshold, or when a classification decision becomes sufficiently stable. Stakeholders need to know what the scale means, how much uncertainty surrounds an individual result, and whether precision differs across the scale. Where scores feed bands or pass decisions, standard setting and classification studies should reflect the adaptive design. Near a cut score, an uncertainty interval or review rule can communicate more honestly than a bare point estimate.
Fairness in High-Stakes Use
High-stakes use requires evidence that relevant groups have equitable opportunities to demonstrate the construct. Differential item functioning should be investigated, but fairness review must also cover accessibility, device conditions, navigation, timing, accommodations and the interaction between adaptive selection and subgroup performance. Removing every flagged item is not automatically correct; statistical flags should trigger substantive review, impact analysis and documented decisions. Programmes should also compare completion rates, precision, classification error and item routes across relevant groups, while protecting privacy and avoiding simplistic conclusions from small samples.
Validation Continues After Launch
Before release, developers should run simulated and human-response studies covering item-bank sufficiency, extreme scores, non-standard response patterns, interruptions and accommodation paths. After release, they should monitor calibration drift, exposure, test information, route comparability, scoring stability, subgroup outcomes and unexpected behaviour. Material changes to the bank, algorithm, population or use case should trigger fresh evaluation rather than being treated as routine maintenance. A defensible programme keeps versioned evidence showing which bank, rules and scoring model produced each result.
EduZMS design recommendation
For an adaptive assessment platform, EduZMS recommends a versioned design record that connects the intended use, blueprint constraints, calibrated bank, selection and stopping rules, scoring model, accessibility decisions and monitoring thresholds. This is an EduZMS platform-design recommendation, not a requirement quoted from the cited publications.
Conclusion
Adaptive testing can make assessment more efficient, but defensible use depends on the whole system: construct definition, bank quality, constrained selection, transparent scoring, fairness evidence and continuing operational review. The right question is not simply whether the algorithm adapts successfully, but whether each reported score remains a sound basis for its intended interpretation and decision.
References
- Wainer, H., Dorans, N. J., Eignor, D., Flaugher, R., Green, B. F., Mislevy, R. J., Steinberg, L., & Thissen, D. (2000). Computerized Adaptive Testing: A Primer (2nd ed.). Lawrence Erlbaum Associates.
- van der Linden, W. J., & Glas, C. A. W. (Eds.). (2010). Elements of Adaptive Testing. Springer. https://doi.org/10.1007/978-0-387-85461-8
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. AERA.