Diagnosing Ceiling Effects and Unstable Nonlinearity in Short Ordinal Scales: A TIMSS 2023 Application
Abstract
1. Introduction
2. A Diagnostic Workflow for Short Ordinal Scales
3. Materials and Methods
3.1. Data
3.2. Software, Weights, Missing Data and Identification
3.3. Range: Endpoint Concentration and Conditional Information
3.4. Structure, Reliability and Method Variance
3.5. Variables and Models Used in the Applied Analyses
3.6. Functional Form
3.7. Simulation
4. Results
4.1. Range: The Scale Measured Poorly Among Confident Children
4.2. Structure: Localized Covariance Rather than a Separable Dimension
4.3. Method Variance
4.4. Functional Form: An Attenuated and Poorly Reproducible Conclusion
4.5. Simulation: When the Diagnosed Conditions Mattered
5. Discussion
5.1. Methodological Contribution
5.2. What the Application Shows About the TIMSS Digital Self-Efficacy Scale
5.3. Implications for Research on Children’s Digital Self-Efficacy and Cybervictimization
5.4. Limitations and Generalizability
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- ALMamari, K. (2026). Cross-national measurement invariance of TIMSS 2023 digital and environmental scales: An alignment optimization approach. Large-Scale Assessments in Education, 14, 31. [Google Scholar] [CrossRef] [Scilit]
- Baumgartner, H., & Steenkamp, J.-B. E. M. (2001). Response styles in marketing research: A cross-national investigation. Journal of Marketing Research, 38(2), 143–156. [Google Scholar] [CrossRef] [Scilit]
- Bond, T. N., & Lang, K. (2013). The evolution of the Black–White test score gap in grades K–3: The fragility of results. The Review of Economics and Statistics, 95(5), 1468–1479. [Google Scholar] [CrossRef] [Scilit]
- Chalmers, R. P. (2012). mirt: A multidimensional item response theory package for the R environment. Journal of Statistical Software, 48(6), 1–29. [Google Scholar] [CrossRef] [Scilit]
- Christensen, K. B., Makransky, G., & Horton, M. (2017). Critical values for Yen’s Q3: Identification of local dependence in the Rasch model using residual correlations. Applied Psychological Measurement, 41(3), 178–194. [Google Scholar] [CrossRef] [Scilit]
- Flora, D. B. (2020). Your coefficient alpha is probably wrong, but which coefficient omega is right? A tutorial on using R to obtain better reliability estimates. Advances in Methods and Practices in Psychological Science, 3(4), 484–501. [Google Scholar] [CrossRef] [Scilit]
- Flora, D. B., & Curran, P. J. (2004). An empirical evaluation of alternative methods of estimation for confirmatory factor analysis with ordinal data. Psychological Methods, 9(4), 466–491. [Google Scholar] [CrossRef] [Scilit]
- Ganzach, Y. (1997). Misleading interaction and curvilinear terms. Psychological Methods, 2(3), 235–247. [Google Scholar] [CrossRef]
- Green, S. B., & Yang, Y. (2009). Reliability of summed item scores using structural equation modeling: An alternative to coefficient alpha. Psychometrika, 74(1), 155–167. [Google Scholar] [CrossRef] [Scilit]
- Klein, A., & Moosbrugger, H. (2000). Maximum likelihood estimation of latent interaction effects with the LMS method. Psychometrika, 65(4), 457–474. [Google Scholar] [CrossRef] [Scilit]
- Livingstone, S., & Helsper, E. J. (2010). Balancing opportunities and risks in teenagers’ use of the internet: The role of online skills and internet self-efficacy. New Media & Society, 12(2), 309–329. [Google Scholar] [CrossRef] [Scilit]
- Livingstone, S., & Smith, P. K. (2014). Annual research review: Harms experienced by child users of online and mobile technologies: The nature, prevalence, and management of sexual and aggressive risks in the digital age. Journal of Child Psychology and Psychiatry, 55(6), 635–654. [Google Scholar] [CrossRef] [Scilit]
- Marsh, H. W., Guo, J., Dicke, T., Parker, P. D., & Craven, R. G. (2020). Confirmatory factor analysis (CFA), exploratory structural equation modeling (ESEM), and set-ESEM: Optimal balance between goodness of fit and parsimony. Multivariate Behavioral Research, 55(1), 102–119. [Google Scholar] [CrossRef] [Scilit]
- Maydeu-Olivares, A., & Coffman, D. L. (2006). Random intercept item factor analysis. Psychological Methods, 11(4), 344–362. [Google Scholar] [CrossRef] [Scilit]
- Paule, R. C., & Mandel, J. (1982). Consensus values and weighting factors. Journal of Research of the National Bureau of Standards, 87(5), 377–385. [Google Scholar] [CrossRef] [Scilit]
- Reise, S. P., Bonifay, W. E., & Haviland, M. G. (2013). Scoring and modeling psychological measures in the presence of multidimensionality. Journal of Personality Assessment, 95(2), 129–140. [Google Scholar] [CrossRef] [Scilit]
- Rhemtulla, M., Brosseau-Liard, P. É., & Savalei, V. (2012). When can categorical variables be treated as continuous? A comparison of robust continuous and categorical SEM estimation methods under suboptimal conditions. Psychological Methods, 17(3), 354–373. [Google Scholar] [CrossRef] [Scilit]
- Rutkowski, L., Gonzalez, E., Joncas, M., & von Davier, M. (2010). International large-scale assessment data: Issues in secondary analysis and reporting. Educational Researcher, 39(2), 142–151. [Google Scholar] [CrossRef] [Scilit]
- Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores. Psychometrika, 34(Suppl. 1), 1–97. [Google Scholar] [CrossRef] [Scilit]
- Sanchez, R. C. (2021). girth (Version 0.8.0) [Computer software]. Zenodo. [CrossRef]
- Simonsohn, U. (2018). Two lines: A valid alternative to the invalid testing of U-shaped relationships with quadratic regressions. Advances in Methods and Practices in Psychological Science, 1(4), 538–555. [Google Scholar] [CrossRef] [Scilit]
- Van de Gaer, E., Grisay, A., Schulz, W., & Gebhardt, E. (2012). The reference group effect: An explanation of the paradoxical relationship between academic achievement and self-confidence across countries. Journal of Cross-Cultural Psychology, 43(8), 1205–1228. [Google Scholar] [CrossRef] [Scilit]
- von Davier, M., Fishbein, B., & Kennedy, A. M. (Eds.). (2024a). TIMSS 2023 technical report: Methods and procedures. TIMSS & PIRLS International Study Center, Boston College. Available online: https://timss2023.org/methods (accessed on 18 August 2026).
- von Davier, M., Kennedy, A. M., Reynolds, K. A., Fishbein, B., Khorramdel, L., Aldrich, C. E. A., Bookbinder, A., Bezirhan, U., & Yin, L. (2024b). TIMSS 2023 international results in mathematics and science. TIMSS & PIRLS International Study Center, Boston College. [Google Scholar] [CrossRef] [Scilit]
- Weijters, B., Geuens, M., & Schillewaert, N. (2010). The stability of individual response styles. Psychological Methods, 15(1), 96–110. [Google Scholar] [CrossRef] [Scilit]
- Yen, W. M. (1993). Scaling performance assessments: Strategies for managing local item dependence. Journal of Educational Measurement, 30(3), 187–213. [Google Scholar] [CrossRef] [Scilit]



| Diagnostic Stage | Potential Failure | Recommended Analysis | Consequence If Undetected | Action If Detected |
|---|---|---|---|---|
| Range | Endpoint concentration | Endpoint mass; graded-response test-information and conditional-SE curves | Effects in the upper range may be compressed and imprecisely estimated | Confine inferences to the informative range, or adopt an IRT-based score before modeling |
| Structure | Local item dependence | Competing ordinal models (one-factor, correlated factors, correlated residuals, bifactor) with cross-validated comparison | Total-score dimensionality may be misidentified | Model the shared covariance; do not treat item subsets as separate competencies |
| Reliability | Specific-factor variance | Omega total, omega for each item subset, and omega hierarchical | High total reliability may conceal multidimensionality | If omega hierarchical is high, retain the total score; otherwise reconsider dimensionality before scoring |
| Method variance | Agreement-format covariance | External endorsement index and method-factor sensitivity model | Associations may partly reflect common response-format variance | Bound associations for shared-format variance; avoid response-style interpretations without direct evidence |
| Functional form | Unstable curvature | Alternative outcome specifications, splines, held-out validation, prediction intervals and practical magnitude | A quadratic term may not represent a stable substantive curve | Interpret a curve only if it survives an alternative outcome metric, a flexible form, out-of-sample validation, and a prediction interval |
| Property | Grade 4 (63 Systems) | Grade 8 (47 Systems) |
|---|---|---|
| % at ceiling (maximum on all 7 items) | 10.8 [8.6, 13.9] | 26.4 [14.4, 33.3] |
| Test information at θ = 0 (unidimensional GRM) | 4.04 [3.59, 4.64] | 4.35 [3.75, 5.28] |
| Test information at θ = +2 SD (unidimensional GRM) | 0.68 [0.57, 0.86] | 0.24 [0.14, 0.54] |
| Information loss, θ = 0 to +2 SD (unidimensional) | 83% | 95% |
| Information loss, θ = 0 to +2 SD (bifactor/testlet) | 80% | 89% |
| Conditional SE at θ = +2 SD (unidimensional|exact testlet) | 1.21|1.23 | 2.06|1.74 |
| Testlet (specific-factor) loading, items 1–3 | 1.01 | 1.18 |
| In-sample SRMR: one factor | 0.048 [0.042, 0.058] | 0.062 [0.057, 0.070] |
| In-sample SRMR: correlated two factors | 0.030 [0.028, 0.038] | 0.032 [0.027, 0.040] |
| In-sample SRMR: one factor + correlated residuals | 0.025 [0.022, 0.029] | 0.027 [0.022, 0.033] |
| Cross-validated SRMR: one factor | 0.057 [0.049, 0.065] | 0.068 [0.063, 0.077] |
| Cross-validated SRMR: correlated two factors | 0.043 [0.038, 0.048] | 0.044 [0.037, 0.050] |
| Cross-validated SRMR: one factor + correlated residuals | 0.038 [0.034, 0.046] | 0.040 [0.034, 0.047] |
| Systems where correlated residuals ≤ two factors (cross-validated) | 59/63 | 46/47 |
| Correlation between the two putative factors | 0.838 [0.817, 0.868] | 0.815 [0.781, 0.842] |
| Omega total | 0.828 [0.812, 0.844] | 0.885 [0.861, 0.910] |
| Omega, items 1–3 | 0.747 [0.711, 0.773] | 0.853 [0.795, 0.876] |
| Omega, items 4–7 | 0.747 [0.722, 0.771] | 0.826 [0.796, 0.861] |
| Omega hierarchical (general factor) | 0.787 [0.770, 0.805] | 0.833 [0.798, 0.866] |
| Cronbach alpha | 0.756 [0.736, 0.775] | 0.811 [0.779, 0.838] |
| Variance shared with external endorsement index | 4.9% [0.2, 13.4] | 6.1% [0.9, 18.3] |
| Latent variance reduction, endorsement-format factor (illustrative, one-system sensitivity estimate) | 32% | 44% |
| Diagnostic | Grade 4 | Grade 8 |
|---|---|---|
| Quadratic on standardized composite (pooled) | +0.027 [0.021, 0.033] | +0.051 [0.042, 0.061] |
| Pooled quadratic, national systems only (nested benchmarks removed) | +0.026 [0.019, 0.032] (k = 58) | +0.051 [0.041, 0.062] (k = 42) |
| Quadratic on composite, matched covariates (median) | +0.0281 | +0.0558 |
| Quadratic on binary outcome, matched covariates (median [IQR]) | 0.0093 [−0.0291, 0.0434] | 0.0259 [−0.0020, 0.0523] |
| Systems with positive binary quadratic | 40/63 | 33/45 |
| Sensitivity: binary quadratic additionally controlling traditional victimization | −0.0203 | −0.0328 |
| 95% prediction interval, fully adjusted | [−0.002, 0.039] | [−0.015, 0.055] |
| Between-system heterogeneity of the pooled quadratic (I2, τ2) | I2 = 77.1%, τ2 = 0.00042 | I2 = 88.4%, τ2 = 0.00083 |
| Held-out deviance: linear|quadratic|spline | 1.1802|1.1779|1.1776 | 1.3197|1.3191|1.3166 |
| Spline estimable (systems) | 63/63 | 45/45 |
| Best held-out specification (linear/quadratic/spline) | 28/11/24 | 12/10/23 |
| Linear fit at least as well as quadratic | 34/63 | 17/45 |
| Split-half reliability, primary model (Spearman–Brown [95% CI]) | 0.665 [0.487, 0.775] | 0.725 [0.478, 0.834] |
| Latent quadratic without|with endorsement-format factor | 0.028|0.036 | 0.021|0.026 |
| Quantity | Raw Summed Score | Latent (EAP) Score |
|---|---|---|
| Population projection coefficient, latent b2 = 0 | −0.0055 (mild), −0.0166 (severe) | −0.0012 (mild), −0.0004 (severe) |
| Rejection of no curvature, latent b2 = 0, mild | 5.35% (MCSE 0.36) | 4.75% (MCSE 0.34) |
| Rejection of no curvature, latent b2 = 0, severe | 16.73% (MCSE 0.59) | 5.47% (MCSE 0.36) |
| Grouped mean bias vs. projection estimand (max across displayed conditions) | ≤0.0013 | ≤0.0013 |
| Cell-level absolute bias (max)/coverage range [pooled across score types] | 0.00446/92.4–98.2% | — |
| Population projection coefficient, latent b2 = 0.10, severe | +0.0535 | +0.0747 |
| Power, latent b2 = 0.10, mild | 89.1% | 92.8% |
| Power, latent b2 = 0.10, severe | 70.8% | 88.5% |
| Estimated-parameter EAP: rejection, latent b2 = 0 (mild|severe) | — | 4.5%|6.0% (MCSE 0.73, 0.84) |
| Estimated-parameter EAP: power at b2 = 0.10 (mild|severe) | — | 91.6%|87.5% |
| Local dependence: rejection under the null (absent|present) [pooled across score types] | 8.4%|7.8% | — |
| Local dependence: power at b2 = 0.10 (absent|present) [pooled across score types] | 86.6%|83.9% | — |
| Workflow diagnostic: endpoint mass | 5.6% (mild) vs. 21.2% (severe) | — |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Sideridis, G.; Alghamdi, M. Diagnosing Ceiling Effects and Unstable Nonlinearity in Short Ordinal Scales: A TIMSS 2023 Application. Behav. Sci. 2026, 16, 1485. https://doi.org/10.3390/bs16091485
Sideridis G, Alghamdi M. Diagnosing Ceiling Effects and Unstable Nonlinearity in Short Ordinal Scales: A TIMSS 2023 Application. Behavioral Sciences. 2026; 16(9):1485. https://doi.org/10.3390/bs16091485
Chicago/Turabian StyleSideridis, Georgios, and Mohammed Alghamdi. 2026. "Diagnosing Ceiling Effects and Unstable Nonlinearity in Short Ordinal Scales: A TIMSS 2023 Application" Behavioral Sciences 16, no. 9: 1485. https://doi.org/10.3390/bs16091485
APA StyleSideridis, G., & Alghamdi, M. (2026). Diagnosing Ceiling Effects and Unstable Nonlinearity in Short Ordinal Scales: A TIMSS 2023 Application. Behavioral Sciences, 16(9), 1485. https://doi.org/10.3390/bs16091485
