Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (132)

Search Parameters:
Keywords = item response theory (IRT)

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
26 pages, 1319 KB  
Article
Psychometric Modeling of Academic Engagement and Dropout Propensity in Higher Education: An Item Response Theory Analysis
by Christina Modiati, George S. Androulakis and Stefanos Balaskas
Educ. Sci. 2026, 16(7), 1096; https://doi.org/10.3390/educsci16071096 - 8 Jul 2026
Viewed by 358
Abstract
Student dropout remains a persistent challenge for higher education systems, with consequences for student trajectories, institutional efficiency, and equity. This study investigates the relationship between student dropout propensity and academic engagement in a sample of 3099 students from all departments of the University [...] Read more.
Student dropout remains a persistent challenge for higher education systems, with consequences for student trajectories, institutional efficiency, and equity. This study investigates the relationship between student dropout propensity and academic engagement in a sample of 3099 students from all departments of the University of Patras. It asks whether academic engagement is associated with multidimensional dropout propensity and whether the instruments provide sufficient psychometric precision for identifying students at different levels of risk. Using the Utrecht Work Engagement Scale (UWES-9) and the APrISE-15 instrument, Item Response Theory (IRT) was applied to evaluate item functioning and scale adequacy. Higher engagement was associated with lower dropout propensity at both IRT-score level (r = −0.55) and observed-score level (r = −0.58). Dedication showed the strongest facet-level association with dropout propensity (r = −0.53), while the economic and personal dropout domains were most informative at elevated risk levels. These findings highlight the value of precision-aware assessment for identifying domain-specific risk profiles and informing targeted student-support strategies, including academic advising, wellbeing-oriented support, financial counseling, and social-integration initiatives. Full article
(This article belongs to the Special Issue Modern Psychometrics for Digital Assessment in Education)
Show Figures

Figure 1

15 pages, 805 KB  
Article
Teacher Educators’ Digital Proficiency and Sustainable Pedagogical Technology Use: An Integrated Model of Competence and Implementation
by Ester Aflalo and Moriya Vaknin
Sustainability 2026, 18(13), 6592; https://doi.org/10.3390/su18136592 - 29 Jun 2026
Viewed by 308
Abstract
This study proposes an integrated model examining the relationship between teacher educators’ digital proficiency and the frequency of their pedagogical use of digital tools. By promoting long-term capacity-building in digital competence, the model contributes to sustainable development goals in education, particularly in ensuring [...] Read more.
This study proposes an integrated model examining the relationship between teacher educators’ digital proficiency and the frequency of their pedagogical use of digital tools. By promoting long-term capacity-building in digital competence, the model contributes to sustainable development goals in education, particularly in ensuring inclusive, equitable, and high-quality learning environments. The study involved 156 faculty members from five teacher-training colleges in Israel. Digital proficiency was measured using a validated self-assessment questionnaire adapted from the SELFIE framework (Self-Reflection on Effective Learning by Fostering the Use of Innovative Educational Technologies). The questionnaire assessed perceived competence across three dimensions: (1) filtering and enhancing digital resources, (2) assessment, feedback, communication, and active learning, and (3) adaptive and creative learning. A second questionnaire examined how frequently educators used specific digital tools across four categories: collaboration, diversity and special needs, active and creative learning, and distance and hybrid learning. Data were analyzed using Item Response Theory (IRT) to generate proficiency scores and Structural Equation Modeling (SEM) to test associations. Results indicated moderate overall digital proficiency, with stronger competence in collaboration and communication and lower use of tools related to personalization and creativity. Significant positive associations were found between digital proficiency and all categories of tool use, especially creative and student-centered learning. Use also varied by gender, seniority, and professional role. The study underscores the importance of pedagogically informed professional development to support meaningful and inclusive digital integration. Full article
(This article belongs to the Special Issue Sustainable Educational Technologies and Improved Learning)
Show Figures

Figure 1

39 pages, 5650 KB  
Article
Integrating Three-Parameter Logistic IRT Models and Confirmatory Factor Analysis for Multidimensional Assessment of Academic Performance and Associated Factors in University Leveling Programs
by Erick P. Herrera-Granda, Paola V. Cabascango-Flores, Iván P. Sandoval-Palis, Tarquino Sánchez-Almeida, Ángel P. Villota-Cadena, María J. Aza-Espinosa, Ronie Martínez and Dayana E. Herrera-Granda
Appl. Sci. 2026, 16(12), 6248; https://doi.org/10.3390/app16126248 - 22 Jun 2026
Viewed by 211
Abstract
This study integrated Item Response Theory (IRT) models with ordinal survey instruments to establish a baseline psychometric framework and identify multidimensional factors associated with academic achievement among first-semester leveling students (N = 1558 pre-test; N = 1676 post-test) at the Escuela Politécnica Nacional, [...] Read more.
This study integrated Item Response Theory (IRT) models with ordinal survey instruments to establish a baseline psychometric framework and identify multidimensional factors associated with academic achievement among first-semester leveling students (N = 1558 pre-test; N = 1676 post-test) at the Escuela Politécnica Nacional, Ecuador. A dual-component methodology was employed in this study. Initially, an 80-item ordinal survey was utilized to assess eight latent constructs, yielding substantial validation metrics through Confirmatory Factor Analysis (CFA). Secondly, structured diagnostic assessments in core STEM and language subjects were calibrated using three-parameter logistic (3PL) IRT models via Expected A Posteriori (EAP) estimation. Results demonstrated high internal consistency (r = 0.93 between IRT and raw scores), with mean IRT-scaled ability θ¯ = 10.45 (SD = 3.51) on a 1–20 scale. Estimated item parameters yielded a mean discrimination of a¯ = 1.92 and a centered mean difficulty of b¯ = 0.05. The Orlando–Thissen SX2 goodness-of-fit test, applied at a significance threshold of p < 0.01, identified 19 items (23.75%) whose observed response patterns deviated significantly from model predictions, with the majority concentrated in the physics and chemistry content domains. Factor scores and performance outcomes were statistically contrasted against 24 categorical demographic variables, revealing differential performance patterns across student subgroups. This research provides validated psychometric instruments, reproducible IRT-LMS integration protocols, and empirical evidence supporting targeted interventions to strengthen university transition. Full article
Show Figures

Figure 1

24 pages, 11783 KB  
Article
Evaluating Inferential Statistics Filtering in High-Dimensional Item Feature Spaces for Predicting IRT Parameters
by Juyoung Jung, Yeonju Lee, Ae Kyong Jung, Seungwon Shin and Won-Chan Lee
Mathematics 2026, 14(10), 1662; https://doi.org/10.3390/math14101662 - 13 May 2026
Viewed by 379
Abstract
Predicting parameter estimates under item response theory (IRT) from expert-coded item features offers a scalable alternative to resource-intensive field testing. This study evaluates whether inferential feature selection can improve predictive accuracy for item difficulty and item discrimination using five filter methods: the Analysis [...] Read more.
Predicting parameter estimates under item response theory (IRT) from expert-coded item features offers a scalable alternative to resource-intensive field testing. This study evaluates whether inferential feature selection can improve predictive accuracy for item difficulty and item discrimination using five filter methods: the Analysis of Variance (ANOVA) F-test, Kendall’s Tau, the Kolmogorov–Smirnov test, the Anderson–Darling test, and the Energy Distance test. Models were trained using K-Nearest Neighbors (KNN) and Support Vector Regression (SVR) under random split and fixed-form cold-start partitioning strategies. Results show that the distributional properties of item features, rather than train–test splitting alone, drive predictive gains: distribution-based filter approaches, particularly the Kolmogorov–Smirnov test, consistently outperformed mean-based approaches by better capturing the full probability structure of the feature-parameter relationship. KNN benefited substantially from feature selection given its reliance on Euclidean distance, while SVR showed smaller gains due to its inherent regularization. Item discrimination generalized well to previously unseen test forms that share no calibration data with the training set, whereas item difficulty prediction was considerably more sensitive to distributional shifts when predicting entirely new, operationally administered forms. The main finding is that the distributional properties of item features are more important than the quantity of features for obtaining robust IRT parameter predictions. Full article
Show Figures

Figure A1

38 pages, 5892 KB  
Article
Psychometric Validation of the Scientific Epistemic Beliefs Questionnaire Among Mexican University Students Using Item Response Theory
by José Antonio Azuela, Laura Inés Ramírez-Hernández, Osvaldo Aquines-Gutiérrez, Wendy Xiomara Chavarría-Garza, Ayax Santos-Guevara and Humberto Martínez-Huerta
J. Intell. 2026, 14(5), 76; https://doi.org/10.3390/jintelligence14050076 - 2 May 2026
Viewed by 942
Abstract
This study examines the validity of the Spanish version of the Scientific Epistemic Beliefs (SEB) Questionnaire among university students in northeastern Mexico, considering multiple sources of evidence. The SEB measures four dimensions of epistemic beliefs: Source, Certainty, Development, and Justification. Data from pilot [...] Read more.
This study examines the validity of the Spanish version of the Scientific Epistemic Beliefs (SEB) Questionnaire among university students in northeastern Mexico, considering multiple sources of evidence. The SEB measures four dimensions of epistemic beliefs: Source, Certainty, Development, and Justification. Data from pilot (n = 150) and main (n = 791) samples were analyzed using Exploratory and Confirmatory Factor Analyses (EFA, CFA), Item Response Theory (IRT), and Differential Item Functioning (DIF). The results provided evidence consistent with a four-factor model, with adequate internal consistency (α = 0.85) and acceptable-to-good fit indices (CFI = 0.944, TLI = 0.936, RMSEA = 0.067, SRMR = 0.071) for a 22-item scale. IRT analyses indicated strong item discrimination, with Source and Certainty covering a broad range of the latent trait, while Development and Justification were more informative at lower to moderate levels. DIF analyses indicated negligible differences in item functioning by gender and academic semester, with minor DIF detected across faculties. Non-parametric analyses identified statistically significant but small differences, with females scoring slightly higher across all dimensions and variations also observed across academic semesters and faculties. Descriptive comparisons with published international data provide contextual evidence within a broader cross-cultural framework. Full article
(This article belongs to the Section Studies on Cognitive Processes)
Show Figures

Figure 1

20 pages, 398 KB  
Article
Robust-Mean–Geometric-Mean and Robust Haberman Linking with Invariant Item Discriminations Under Sparse Differential Item Functioning
by Alexander Robitzsch
Mathematics 2026, 14(9), 1549; https://doi.org/10.3390/math14091549 - 2 May 2026
Viewed by 329
Abstract
Comparison of two or multiple groups based on dichotomous items is a central task in item response theory (IRT) linking. This article considers the two-parameter logistic scaling model under sparse differential item functioning (DIF) in item intercepts and DIF-free item discriminations. Robust-mean-geometric-mean (RMGM) [...] Read more.
Comparison of two or multiple groups based on dichotomous items is a central task in item response theory (IRT) linking. This article considers the two-parameter logistic scaling model under sparse differential item functioning (DIF) in item intercepts and DIF-free item discriminations. Robust-mean-geometric-mean (RMGM) and robust Haberman (RHAB) linking are compared across several loss functions and under scaling models with noninvariant or invariant item discriminations. Two simulation studies show that invariant item discriminations improve the precision of estimated group means. In addition, the L0 loss function is generally preferable to the L1 and L0.5 loss functions when DIF proportions or sample sizes are large. Several empirical examples illustrate the proposed specifications. Full article
(This article belongs to the Special Issue Computational Statistics, Data Analysis and Applications)
21 pages, 1257 KB  
Article
Development and Validation of a Geometric Reasoning Test: Evidence from Preservice Teachers
by Khin Mimi Kyaw and Tibor Vidákovich
Educ. Sci. 2026, 16(5), 690; https://doi.org/10.3390/educsci16050690 - 27 Apr 2026
Viewed by 742
Abstract
This study developed and validated a curriculum-aligned instrument to assess preservice primary teachers’ geometric reasoning skills. Addressing the limited availability of domain-specific tools in teacher education research, the study examined preservice teachers’ conceptual strengths and weaknesses across key geometry domains relevant to primary [...] Read more.
This study developed and validated a curriculum-aligned instrument to assess preservice primary teachers’ geometric reasoning skills. Addressing the limited availability of domain-specific tools in teacher education research, the study examined preservice teachers’ conceptual strengths and weaknesses across key geometry domains relevant to primary mathematics teaching. A two-phase quantitative research design was employed. In Study 1, Confirmatory Factor Analysis (CFA) and Item Response Theory (IRT) were used to evaluate the psychometric properties of the instrument with a sample of 221 preservice teachers, providing evidence of construct validity and internal consistency. Geometric reasoning was conceptualised as a four-factor structure comprising Conceptualisation of Geometric Properties (GP), Geometric Transformation Reasoning (GT), Reasoning with Representations of Three-Dimensional Objects (RE), and Measurement Reasoning (MS). In Study 2, the validated Geometric Reasoning Test (GRT) was administered to a larger sample of 406 preservice primary teachers from three education colleges in Myanmar. Descriptive statistics and group comparisons were conducted using Welch’s t-tests and Welch’s ANOVA to examine differences by gender, year level, and institution. The findings indicate that preservice primary teachers’ geometric reasoning remains underdeveloped across training stages, highlighting the need for greater emphasis on geometry and spatial reasoning in teacher education. Full article
(This article belongs to the Section Curriculum and Instruction)
Show Figures

Figure 1

19 pages, 313 KB  
Review
Cognitive Diagnosis Computerized Adaptive Testing (CD-CAT) for Adolescent Internet Gaming Disorder: A Conceptual Assessment Framework
by Min Jia and Jing Liu
Behav. Sci. 2026, 16(4), 558; https://doi.org/10.3390/bs16040558 - 8 Apr 2026
Viewed by 634
Abstract
Internet Gaming Disorder (IGD) has become a major behavioral health concern among adolescents, yet current assessment tools remain limited. These tools often fail to capture the disorder’s complex symptom variations and lack clinical interpretability. This study, taking an interdisciplinary approach that combines clinical [...] Read more.
Internet Gaming Disorder (IGD) has become a major behavioral health concern among adolescents, yet current assessment tools remain limited. These tools often fail to capture the disorder’s complex symptom variations and lack clinical interpretability. This study, taking an interdisciplinary approach that combines clinical psychology and psychometrics, summarizes recent progress in understanding adolescent IGD and the development of its assessment methods. We compare the diagnostic criteria of the DSM-5 TR and ICD-11 and argue that the nine DSM-5 TR criteria are particularly suited for transformation into distinct diagnostic attributes due to their detailed and actionable nature. We then review the strengths and weaknesses of Classical Test Theory (CTT), Item Response Theory (IRT), and Cognitive Diagnostic Models (CDMs) in assessing IGD. The review emphasizes the limitations of total-score and single latent-trait approaches in capturing the disorder’s multidimensional symptoms. Based on these insights, we propose a conceptual assessment framework, Cognitive Diagnosis Computerized Adaptive Testing (CD-CAT), that integrates CDMs with computerized adaptive testing. Rather than presenting an empirically validated system, this framework offers a theoretically grounded proposal that specifies the key components, logical relationships, and methodological pathways necessary for advancing precision assessment of adolescent IGD. CD-CAT uses a system of attributes and a Q-matrix based on the DSM-5 TR criteria to efficiently classify IGD symptoms in adolescents, reducing the number of items required while enhancing clinical relevance. Lastly, we discuss the theoretical contributions of the proposed framework, acknowledge its limitations as a conceptual proposal, and outline directions for future empirical research. Full article
27 pages, 1388 KB  
Article
The Best of Two Worlds: IRT-Enhanced Automated Essay Interpretable Scoring
by Wei Xia, Jin Wu, Jiarui Yu and Chanjin Zheng
Behav. Sci. 2026, 16(4), 542; https://doi.org/10.3390/bs16040542 - 6 Apr 2026
Viewed by 1248
Abstract
The Automated Essay Scoring (AES) systems confront two fundamental challenges: opaque “black-box” decision-making that limits educator trust, and insufficient validation across linguistically diverse educational contexts. This study proposes IRT-AESF, an innovative framework that bridges educational measurement theory and artificial intelligence by integrating item [...] Read more.
The Automated Essay Scoring (AES) systems confront two fundamental challenges: opaque “black-box” decision-making that limits educator trust, and insufficient validation across linguistically diverse educational contexts. This study proposes IRT-AESF, an innovative framework that bridges educational measurement theory and artificial intelligence by integrating item response theory (IRT) with deep learning. The framework generates three theoretically grounded psychometric parameters: student ability, item difficulty, and item discrimination, which provide transparent and interpretable explanations for scoring decisions. We rigorously evaluated IRT-AESF through 5-fold cross-validation on three large-scale datasets comprising 41,328 authentic essays from English and Chinese educational settings, including both classroom assessments and high-stakes examinations. Results demonstrate statistically significant improvements over competitive baseline models, achieving an 8.4% relative increase in quadratic weighted kappa while maintaining robust cross-lingual performance. This research advances the development of transparent, trustworthy automated assessment systems that deliver not only scores but meaningful diagnostic insights for educational practice. Full article
Show Figures

Figure 1

22 pages, 569 KB  
Article
Student Involvement in Digital Tool Selection: A Pedagogical Approach to Critical Thinking-Oriented Learning
by Ester Aflalo
Educ. Sci. 2026, 16(4), 512; https://doi.org/10.3390/educsci16040512 - 25 Mar 2026
Viewed by 757
Abstract
Digital technologies are widely recognized for their potential to support active learning and foster higher-order cognitive skills, including critical thinking. However, limited research has examined the extent to which students are directly involved in selecting digital tools that shape their learning. This study [...] Read more.
Digital technologies are widely recognized for their potential to support active learning and foster higher-order cognitive skills, including critical thinking. However, limited research has examined the extent to which students are directly involved in selecting digital tools that shape their learning. This study investigates teachers’ ability to engage students in the selection and pedagogical use of digital technologies, with attention to practices supporting active, personalized learning and critical thinking. Data were collected from 156 educators across diverse disciplines in five teacher-training colleges in Israel using an online questionnaire assessing levels of digital tool use, from non-use to active student involvement. Item Response Theory (IRT) was applied to model teachers’ proficiency and examine differences across tools and background characteristics. Results indicate substantial variability in teachers’ ability to involve students, with particularly low involvement in tools related to problem-solving, differentiation, and personalized learning. Gender and institutional role were significant predictors, with female educators and those holding additional roles demonstrating higher proficiency. These findings highlight the importance of teachers’ techno-pedagogical competence in enabling student participation in digital decision-making and suggest that involving students in tool selection can support the development of critical thinking and learner agency in digitally mediated learning environments. Full article
Show Figures

Figure 1

24 pages, 632 KB  
Article
The Arabic Lubben Social Network Scale-6: Psychometric Validation, Measurement Invariance, and Social Support Profiles in Arabic-Speaking Older Adults
by Khaled Trabelsi, Waqar Husain, Hadeel Ghazzawi, Zahra Saif, Achraf Ammar and Haitham Jahrami
Eur. J. Investig. Health Psychol. Educ. 2026, 16(3), 40; https://doi.org/10.3390/ejihpe16030040 - 6 Mar 2026
Cited by 1 | Viewed by 1370
Abstract
This study aimed to translate, culturally adapt, and validate the Arabic version of the 6-Item Lubben Social Network Scale (LSNS-6). The LSNS-6 was translated, culturally adapted, and administered, alongside the Medical Outcomes Study Social Support Survey (MOS-SSS), to 327 Arabic-speaking adults aged 60 [...] Read more.
This study aimed to translate, culturally adapt, and validate the Arabic version of the 6-Item Lubben Social Network Scale (LSNS-6). The LSNS-6 was translated, culturally adapted, and administered, alongside the Medical Outcomes Study Social Support Survey (MOS-SSS), to 327 Arabic-speaking adults aged 60 years and older. Internal consistency was examined using Cronbach’s alpha and McDonald’s omega. Confirmatory factor analysis (CFA) tested the hypothesized two-factor structure (Family and Friends), and measurement invariance was evaluated across key sociodemographic and lifestyle variables. Convergent validity was assessed through correlations with MOS-SSS domains. Item response theory (IRT) analyses examined item discrimination and threshold parameters. Latent class analysis (LCA) explored whether the LSNS-6 could identify subgroups with distinct patterns of social connectedness and perceived support. The Arabic LSNS-6 demonstrated good internal consistency (α = 0.83; ω = 0.84) and supported the expected two-factor structure with satisfactory model fit (CFI = 0.963; TLI = 0.931; SRMR = 0.03). Convergent validity was evidenced by moderate correlations with overall perceived social support (r = 0.51). IRT analyses indicated strong discrimination for most items, and LCA identified four distinct latent classes. Overall, the Arabic LSNS-6 is a reliable and valid tool for assessing social isolation among older Arabic-speaking adults. Full article
Show Figures

Figure 1

20 pages, 1909 KB  
Article
Operationalising CTT and IRT in Spreadsheets: A Methodological Demonstration for Classroom Assessment
by António Faria and Guilhermina Lobato Miranda
Analytics 2026, 5(1), 12; https://doi.org/10.3390/analytics5010012 - 24 Feb 2026
Viewed by 1464
Abstract
The evaluation of student performance often relies on basic spreadsheet outputs that provide limited insight into item functioning. This study presents a methodological demonstration showing how widely available spreadsheet software can be transformed into a practical environment for psychometric analysis. Using a simulated [...] Read more.
The evaluation of student performance often relies on basic spreadsheet outputs that provide limited insight into item functioning. This study presents a methodological demonstration showing how widely available spreadsheet software can be transformed into a practical environment for psychometric analysis. Using a simulated dataset of 40 students responding to 20 dichotomous items, spreadsheet formulas were developed to compute descriptive statistics and Classical Test Theory (CTT) indices, including item difficulty, discrimination, and corrected item–total correlations. The demonstration was extended to Item Response Theory (IRT) through the implementation of 1PL, 2PL, and 3PL logistic models using forward-calculated item parameters. A smaller dataset of 10 students and 10 items was used to illustrate the interpretability of the indices and the generation of Item Characteristic Curves (ICCs). Results show that spreadsheets can support teachers in in-terpreting test data beyond total scores, enabling the identification of weak items, refinement of distractors, and construction of small-scale item banks aligned with competence-based curricula. The approach contributes to Sustainable Development Goal 4 (SDG 4) by promoting accessible, equitable, and high-quality assessment practices. Limitations include the instability of IRT parameter estimation in small samples and the need for teacher training. Future research should apply the approach to real classroom data, explore automation within spreadsheet environments, and examine the integration of artificial intelligence for adaptive assessment. Full article
Show Figures

Figure 1

24 pages, 344 KB  
Article
A Hybrid Neural-IRT Framework for Addressing Cold-Start Challenges in Computerized Adaptive Testing
by Almira Iskakova, Olga Salykova, Nauzhan Didarbekova, Irina Ivanova, Anara Akmoldina and Ainur Zhumadillayeva
Computers 2026, 15(2), 132; https://doi.org/10.3390/computers15020132 - 19 Feb 2026
Viewed by 1230
Abstract
Computerized adaptive testing (CAT) systems face major challenges at the beginning of test administration, when limited response data produces unstable ability estimates and poor item selection. This cold-start problem reduces measurement precision and testing efficiency, especially for students whose abilities diverge from population [...] Read more.
Computerized adaptive testing (CAT) systems face major challenges at the beginning of test administration, when limited response data produces unstable ability estimates and poor item selection. This cold-start problem reduces measurement precision and testing efficiency, especially for students whose abilities diverge from population norms. This study introduces a hybrid ability-estimation model that dynamically integrates neural network predictions with classical item response theory (IRT) estimation throughout the adaptive testing process. The neural component uses auxiliary student information-including demographics, prior performance, and early response patterns-to generate accurate initial ability estimates, while the IRT component preserves psychometric validity as response data accumulate. A dynamic fusion mechanism gradually shifts estimation weight from the neural model to the IRT model as more items are administered. Experimental validation on 2847 students across four subject domains shows that the hybrid approach reduces RMSE in ability estimation by 34.2% during the first five items compared with traditional CAT methods, while maintaining equivalent precision in later stages. The system also decreases the number of items required to reach target precision (SE < 0.3) by 28.7% on average, with the largest gains observed for students at ability extremes. Full article
(This article belongs to the Section Human–Computer Interactions)
Show Figures

Figure 1

14 pages, 304 KB  
Article
Revisiting the Geriatric Depression Scale: An IRT-Based 10-Item Screen Outperforms the GDS-15 in Diagnostic Accuracy and Efficiency
by Ji Won Han, Dae Jong Oh, Tae Hui Kim, Kyung Phil Kwak, Bong Jo Kim, Shin Gyeom Kim, Jeong Lan Kim, Seok Woo Moon, Joon Hyuk Park, Seung-Ho Ryu, Jong Chul Youn, Dong Young Lee, Dong Woo Lee, Seok Bum Lee, Jung Jae Lee, Jin Hyeong Jhoo and Ki Woong Kim
J. Clin. Med. 2026, 15(2), 473; https://doi.org/10.3390/jcm15020473 - 7 Jan 2026
Viewed by 975
Abstract
Background/Objective: Existing abbreviated Geriatric Depression Scales (GDSs), derived via Classical Test Theory (CTT), often sacrifice accuracy for brevity and retain non-specific items. We aimed to develop a minimum-item GDS maintaining diagnostic performance equivalent to the full 30-item scale (GDS30) using Item Response [...] Read more.
Background/Objective: Existing abbreviated Geriatric Depression Scales (GDSs), derived via Classical Test Theory (CTT), often sacrifice accuracy for brevity and retain non-specific items. We aimed to develop a minimum-item GDS maintaining diagnostic performance equivalent to the full 30-item scale (GDS30) using Item Response Theory (IRT). Methods: This cross-sectional study employed rigorous 5:5 split-sample cross-validation. Participants included 6525 older adults (aged ≥60 years) from community-based (Korean Longitudinal Study on Cognitive Aging and Dementia) and clinical settings (geropsychiatry clinic). Depression was diagnosed through standardized clinical interviews based on DSM-IV criteria. Two-parameter logistic IRT models estimated item discrimination and difficulty parameters. Sequential item reduction with DeLong tests identified the minimum number of items required to maintain GDS30-equivalent area under the curve (AUC). Results: The 10-item IRT-optimized scale (GDS10-IRT) achieved an AUC of 0.856 (95% CI: 0.809–0.895) in the validation set, showing no significant difference from GDS30 (AUC = 0.883; p = 0.396). Conversely, the 15-item GDS (GDS15) demonstrated significantly lower AUC than GDS30 (p < 0.001) despite having more items. GDS10-IRT achieved a 234% improvement in efficiency ratio (AUC/items) over GDS30. Notably, Item 16 (“feeling downhearted and blue”), identified as the most discriminating symptom (a = 2.53), is absent from the GDS15 but included in GDS10-IRT. Conclusions: IRT-based item selection achieves GDS30-equivalent diagnostic accuracy with only 10 items, outperforming the widely used GDS15. By recovering high-discrimination items excluded by CTT, the GDS10-IRT offers a more efficient, specific screening tool for late-life depression. Full article
(This article belongs to the Section Mental Health)
21 pages, 1761 KB  
Article
Developmental Change in Associations Between Mental Health and Academic Ability Across Grades in Adolescence: Evidence from IRT-Based Vertical Scaling
by Yuanqiu Ma, Youyou Duan, Yunxiao Qi, Ying Hu and Tour Liu
Behav. Sci. 2026, 16(1), 78; https://doi.org/10.3390/bs16010078 - 6 Jan 2026
Viewed by 779
Abstract
Adolescence is a critical period when rapid cognitive maturation coincides with heightened emotional vulnerability. This study examined the dynamic association between academic ability and mental health across early adolescence, focusing on vocabulary ability as a core indicator of academic ability. Using large-scale data [...] Read more.
Adolescence is a critical period when rapid cognitive maturation coincides with heightened emotional vulnerability. This study examined the dynamic association between academic ability and mental health across early adolescence, focusing on vocabulary ability as a core indicator of academic ability. Using large-scale data from Grades 1–12 (N = 13,412), a vertically scaled vocabulary ability scale was constructed based on Item Response Theory (IRT) and the Non-Equivalent Anchor Test (NEAT) design to achieve cross-grade comparability. Fixed-parameter calibration was then applied to an independent cross-sectional sample of middle school students (Grades 7–9, N = 401) in Tianjin, combined with the DASS-21 to assess internalizing symptoms (depression, anxiety, stress). Hierarchical multiple regression analyses revealed that higher vocabulary ability was significantly associated with lower levels of depression, anxiety, and stress, with the negative association strongest in Grade 8. The present study provides new empirical evidence for understanding the interactive mechanisms between academic and psychological development during adolescence. Methodologically, the study demonstrates the value of IRT-based vertical scaling in establishing developmentally interpretable metrics for educational and psychological assessment. Full article
Show Figures

Figure 1

Back to TopTop