Next Article in Journal
Adaptation and Validation of the Bern Illegitimate Tasks Scale (BITS) in the Context of a Portuguese Public University
Next Article in Special Issue
Autonomy-Supportive Classroom Climate: Its Multilevel Structure and Associations with Students’ Mathematics Achievement
Previous Article in Journal
Digital Learning Competence and Learning Performance Among Chinese Higher Vocational College Students: A Dual-Path Moderated Mediation Model
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Strategy Profiles in Solving Algebra Word Problems: A Person-Centered Bayesian Classification Approach for Chinese Students

School of Psychology, Shenzhen University, Nanshan District, Shenzhen 518060, China
*
Author to whom correspondence should be addressed.
Behav. Sci. 2026, 16(6), 953; https://doi.org/10.3390/bs16060953
Submission received: 24 April 2026 / Revised: 27 May 2026 / Accepted: 4 June 2026 / Published: 10 June 2026

Abstract

This study examined the reversal error phenomenon in Chinese students’ algebra word problem-solving, focusing on how sentence structures influence performance and problem representation strategies across educational levels. Using a person-centered Bayesian classification approach, this study analyzed individual differences in problem-solving strategies across different grade levels. Results confirmed the presence of reversal errors among Chinese students, with sentence structures significantly affecting error rates. Students demonstrated superior performance on congruent problems compared to incongruent problems, showing both fewer errors and faster response times across all grade levels. The study revealed that students processed congruent problems and incongruent problems using fundamentally different strategies. The analysis identified five distinct strategy profiles across grade levels, revealing grade-related differences in strategy use: younger students predominantly relied on direct translation, whereas older students more frequently employed analytic strategies. These findings advance our understanding of cognitive processes in algebra problem-solving and suggest targeted interventions for addressing the reversal error.

1. Introduction

Algebra represents a crucial transition from concrete to abstract mathematics. According to the National Council of Teachers of Mathematics (2000), the primary goal of algebra instruction is to enable students to use algebraic language as a tool for representing real-life situations and solving problems. This objective is primarily achieved through algebra word problems. However, research shows that students consistently struggle with these problems (Jupri & Drijvers, 2016; Soneira et al., 2021). A classic example of this difficulty is the “student-professor problem”, which demonstrates the reversal error (RE): “There are six times as many students as professors at this university. Use S for the number of students and P for the number of professors and write an equation.” In their seminal study, Clement et al. (1981) found that 37% of engineering freshman students failed to solve this problem correctly. The most common mistake was the reversal error, where students wrote “6S = P” instead of the correct equation “S = 6P”, thereby inverting the relationship between the variables. The reversal error has proven persistent across diverse groups, including adolescents (Martin & Bassok, 2005; MacGregor & Stacey, 1993), college students (Clement, 1982; Fisher et al., 2011), and even pre-service teachers and high school teachers (Pawley & Cooper, 1997; Soneira et al., 2021).
Cultural factors, language differences, and teaching methods significantly influence students’ mathematical performance (Cai, 2004; Jiang et al., 2017; Soneira et al., 2021). However, research on algebra word problems has been predominantly conducted in Western contexts, with only one notable exception: A study in Hong Kong, which was still conducted in English (Lopez-Real, 1995). This geographical and linguistic limitation is particularly noteworthy given that Chinese, as a logographic rather than phonetic language, may lead to different performance patterns among Chinese students when solving the student-professor problem compared to their Western counterparts. Additionally, previous research (González-Calero et al., 2015; Soneira et al., 2021) has generally analyzed the reversal error as a group phenomenon, potentially obscuring important individual differences in problem-solving approaches and cognitive processes. To address these gaps in the literature, the present study has two primary objectives: (1) Examine how varying sentence structures affect Chinese students’ performance in the student-professor problem. (2) Identify and analyze individual differences in problem-solving strategies using a person-centered approach. Through these objectives, this study aims to enhance our understanding of both cultural and individual factors that influence algebra word problem-solving abilities. The findings could inform the development of more effective teaching strategies and targeted interventions.

1.1. The Influence of Sentence Structures on the Reversal Error

Researchers have identified two primary strategies employed by students who make reversal errors: The static comparison strategy and the word order matching strategy (also known as direct translation strategy) (Clement, 1982; González-Calero et al., 2015; Soneira et al., 2021). The static comparison strategy occurs when students treat algebraic letters as labels rather than variables. For example, they might interpret “S” and “P” as representing “Students” and “Professors” instead of “the number of students” and “the number of professors.” Students using this strategy often misinterpret the equal sign as indicating comparison or association rather than mathematical equivalence. In the student-professor problem, they might interpret the equation “6S = P” as meaning “six students correspond to one professor.” However, research on these strategies has primarily focused on Western populations, where students regularly use initial letters to represent objects both in educational settings and daily life (e.g., C for cats, D for dogs). This practice may influence their tendency to treat variable letters as object labels in algebraic problem-solving (Knuth et al., 2005; McNeil et al., 2010; Soneira et al., 2018). The Chinese context offers a distinct perspective. As a logographic rather than phonetic language, Chinese has no inherent relationship between character spelling and pronunciation. Consequently, Chinese speakers rarely use characters as object labels. In Chinese mathematics education, generic letters (x and y) are the standard choice for representing variables. Our preliminary research with middle and high school students revealed no significant performance differences in the student-professor problem whether variables were represented using English letters, Pinyin (Chinese pronunciation) initials, or generic letters (x and y). These findings suggest that Chinese students are less likely to employ the static comparison strategy when solving algebraic word problems, including the student-professor problem. Given these observations, the present study will focus exclusively on investigating the direct translation strategy, which appears more relevant to Chinese students’ problem-solving approaches.
The direct translation strategy occurs when students write equations from left to right following the sentence order, leading to incorrect answers in the student-professor problem (Clement et al., 1981; Fisher et al., 2011). The direct translation hypothesis suggests that preventing this strategy could reduce reversal errors. Supporting this theory, González-Calero et al. (2020) found that Basque/Spanish bilingual pre-service teachers made fewer reversal errors when problems were presented in Basque rather than Spanish, as Basque’s linguistic structure naturally inhibits direct translation approaches. Sentence structure has emerged as a significant factor in reversal error occurrence. Soneira et al. (2018) demonstrated that pre-service teachers performed better with statements free from syntactic obstruction (e.g., “The number of students is six times greater than the number of professors”) compared to those with syntactic obstruction (e.g., “There are six students for every professor”). However, the interpretation of such statements can vary significantly across languages. For instance, in Chinese and Danish, the phrase “times greater than” can be semantically different from “times as many as”, potentially leading to a different solution (S = (6 + 1) P = 7P) (Jankvist & Niss, 2021). This linguistic variation means that statements considered unambiguous in one language may become unclear in another.
Notably, research with English-speaking students has produced contrasting results. Christianson et al. (2012) found no significant difference in reversal error rates between problems using helpful phrasing (“The number of students is six times the number of professors”) versus unhelpful phrasing (“There are six times as many students as professors”). These divergent findings across languages and cultures emphasize the importance of further investigating how sentence structure influences reversal errors in different linguistic contexts. In Chinese, two distinct expressions are commonly used for the student-professor problem: (1) “学生人数是教授的6倍” (literally “Students are professors six times”). This is designated as the congruent problem in this study because the direct translation strategy leads to the correct answer (S = 6P). This Chinese phrasing is more explicit and aligns naturally with the algebraic equation format. Unlike Western students, who show high rates of reversal errors (roughly 40–60%) on this kind of problems (Clement et al., 1981; Soneira et al., 2018), Chinese students may make fewer reversal errors with this phrasing due to its reduced syntactic ambiguity. (2) “每6个学生就有一个教授” (literally “There is one professor for every six students”). This is classified as the incongruent problem because the direct translation strategy results in an incorrect answer. This expression parallels English phrasing structures, and applying the direct translation strategy leads to erroneous solutions. While research has demonstrated that reversal errors persist across age groups (Cohen & Kanim, 2005; González-Calero et al., 2020; Pawley & Cooper, 1997), there remains a significant gap in our understanding of how these errors evolve over time. Few studies have conducted direct comparisons of different age groups within a single study, limiting our understanding of how reversal errors vary across grade levels. This lack of comparative data makes it difficult to assess how students’ problem-solving strategies and their susceptibility to reversal errors may change throughout their educational journey. To address this gap and enhance our understanding of reversal errors in the Chinese educational context, the first aim of this study was twofold: (1) Examine how different surface structures (congruent vs. incongruent phrasing) influenced Chinese students’ algebraic word problem-solving. (2) Investigate how these effects vary across educational levels in algebraic problem-solving strategies. Based on these identified linguistic differences, we propose two hypotheses: (1) Chinese students will demonstrate lower rates of reversal errors on congruent problems compared to their Western counterparts, due to the more explicit phrasing in Chinese. (2) On incongruent problems, Chinese students will show reversal errors rates similar to Western students, as the phrasing parallels English structure and encourages the use of the direct translation strategy (Clement et al., 1981; Kim et al., 2014).

1.2. Individual Differences in Strategy Use for Mathematical Problem-Solving

Most prior studies on students’ mathematical problem-solving have analyzed data at the group level (Li et al., 2023; Passolunghi et al., 2022). However, this aggregate methodology may obscure individual differences in the underlying cognitive processes involved. For example, Reinhold et al. (2020) did not find a significant effect of item congruency in fraction comparison tasks for 6th graders when analyzing the whole sample. This was because students with typical bias showed congruency effects in the opposite direction compared to students with reverse bias, causing the group level effect to effectively cancel out. Through cluster analysis, they identified three distinct subgroups: A typical bias cluster, a reverse bias cluster, and a no bias cluster. This exemplifies how person-centered statistical methods are needed to unmask individual differences in strategy use that can be masked by group-level analyses.
While an increasing number of researchers have employed person-centered statistical approaches to unmask individual differences in cognitive processes during mathematical problem-solving, most have utilized explorative techniques such as cluster analysis, latent class analysis, and latent profile analysis (Mo et al., 2026; Reinhold et al., 2020; Vanluydt et al., 2022). However, these explorative person-centered approaches are data-driven, meaning the results greatly rely on and are constrained by the specific sample. This severely impairs the stability and repeatability of the findings. Therefore, to robustly investigate individual differences in strategy use for algebraic word problems, the present study aimed to apply hypothesis-driven (non-explorative) person-centered statistical methods. The naive Bayesian classification approach (the naive Bayes classifier, the NBC) is one of the hypothesis-driven person-centered statistical methods, which is theory-driven and has been used in educational and psychological research in recent years (Culbertson, 2016; Jiang et al., 2026; Leuders & Loibl, 2020; Reinhold et al., 2023). In the naive Bayesian classification approach, all possible strategies that students may use in a certain context are considered as classes, and the features of each class are defined based on psychological theories and previous research findings. These features represent the likelihood of different response patterns within each class. For example, if the features are related to the accuracy of problem-solving, there would be corresponding features indicating the likelihood of achieving certain levels of accuracy on specific problems for each class of strategies. The key assumption in naive Bayesian classification is that the features within each class are independent of each other. With the pre-defined classes and features, the naive Bayesian classification approach analyzes each student’s responses and calculates the posterior probability of each class given the observed responses. Thus, if a student consistently demonstrates a particular response pattern that matches the features of a certain class, the naive Bayesian classification approach will assign a high probability to that class for the student. This process allows for the classification of students based on their observed behavior data.
Recently, some researchers have used the naive Bayesian classification approach to investigate individual differences in strategy use when solving fraction comparison tasks (Reinhold et al., 2023). They first identified the classes of comparison strategies based on the conceptual change account (Vamvakoussi & Vosniadou, 2004) and the dual process account (Dooren & Inglis, 2015). For each class, there is a feature that described the likelihood of the accuracy and response time to different types of fraction problems. Students were then classified by the naive Bayesian classification approach. The results not only corroborated existing strategy patterns identified in previous studies but also revealed additional composite strategy patterns. This hypothesis-driven approach helps to reduce misjudgments stemming from subjective factors and improves classification accuracy compared to exploratory clustering methods. Researchers have employed various methods to investigate the cognitive processes underlying the student-professor problem, with clinical interviews emerging as a valuable tool for revealing individual differences in problem-solving strategies (Adu-Gyamfi et al., 2015; Clement, 1982; Wollman, 1983). In a notable study, Adu-Gyamfi et al. (2015) conducted clinical interviews with college students to examine their cognitive processes during problem-solving. Their findings identified several distinct types of translation errors, including errors in attribute construction, implementation construction, equivalence construction, and unclear construction. However, these interview-based studies on strategy use tend to have limited sample sizes, reducing the representativeness and generalizability of the strategy classification results. Therefore, to comprehensively examine individual differences in cognitive processes underlying the student-professor problem, we proposed a more complete spectrum of hypothesized strategy use patterns based on potential underlying cognitive mechanisms. We then empirically tested the existence of these hypothesized strategy patterns through the naive Bayesian classification approach.

1.3. Cognitive Mechanisms Underlying Students’ Strategy Use

Our study examines student strategies in solving the student-professor problem through the lens of dual-process theory and the inhibitory control model. Dual-process theory (Evans, 2003; Evans & Stanovich, 2013) proposes two distinct cognitive processing systems: The heuristic system (S1) which is fast, effortless, and automatic and the analytic system (S2), which is slow, effortful, and demanding. While heuristic processes are typically adaptive and activate first, the analytic system must intervene when heuristic solutions conflict with correct outcomes. In algebra word problems, researchers have linked reversal errors to a left-right direct-translation heuristic (Fisher et al., 2011; González-Calero et al., 2020). Based on dual-process theory, we propose that the direct translation strategy emerges from the heuristic system, automatically activating when students encounter student-professor problems. This heuristic response likely stems from two factors: (1) The structural similarity between algebraic equations and sequential linear relational statements (Fisher et al., 2011). (2) Students’ repeated success using direct translation in solving mathematical word problems (Hegarty et al., 1995; Lubin et al., 2016; Passolunghi et al., 2022).
The effectiveness of these strategies varies by problem type: In congruent problems, direct translation (heuristic system) leads to correct equations and fast responses. In incongruent problems, direct translation yields incorrect solutions and fast responses, requiring students to employ the analytic strategy (analytic system) to achieve correct but slower responses.
While dual-process theory explains the coexistence of heuristic and analytical processes, it doesn’t fully account for how students successfully transition from heuristic to analytical approaches when solving incongruent problems. The inhibitory control model (Houdé & Borst, 2014, 2015) addresses this gap by emphasizing the critical role of inhibitory control mechanism. According to this model, misleading heuristics and correct strategies co-exist in the brain throughout development, but since heuristics are prepotent, students must engage inhibitory control to overcome them and reach correct solutions. Research has consistently demonstrated the importance of inhibitory control in this process (Jiang et al., 2019; Li et al., 2023). The inhibitory control model suggests two potential reasons for problem-solving failures (Jiang et al., 2020; Mevel et al., 2014): (1) Lack of conflict detection: Students fail to recognize the mismatch between the misleading strategy and the problem situation. (2) Insufficient inhibition: Students detect the conflict but cannot successfully inhibit the misleading strategy. This creates a hierarchical process where conflict detection serves as prerequisite for inhibitory control. Without conflict detection, there is no trigger for inhibitory mechanisms. Even with successful conflict detection, students may still apply the misleading heuristic if they cannot effectively inhibit it.
Based on this framework, we posit that successful resolution of the student-professor problem depends on two key mechanisms: (1) Conflict detection: Recognizing when direct translation strategies are inappropriate. (2) Inhibitory control: Successfully suppressing the misleading direct translation heuristic. The interaction of these mechanisms can create distinctive patterns in both problem-solving accuracy and response times. Inaccurate responses typically stem from two potential failures: Either a breakdown in conflict detection (such as failing to recognize contextual incompatibility) or a lapse in inhibitory control (continuing with the heuristic approach despite recognizing conflicts). To achieve accurate responses, both mechanisms must function successfully. Response times provide additional insight into these cognitive processes. Studies across various tasks (Babai et al., 2012; Frey et al., 2018; Jiang et al., 2020; Pennycook et al., 2012; Stupple et al., 2013) indicate that conflict detection requires more time than non-conflict scenarios, as it involves evaluating contextual ambiguity. Additionally, inhibitory processes—when suppressing a dominant but incorrect strategy—lead to longer response times compared to situations not requiring inhibition. Therefore, extended response times may indicate the engagement of both conflict resolution and inhibitory mechanisms.
In the present study, we propose four distinct profiles of students solving the student-professor problem. These profiles are: Automatic Translators, Conflict Recognizers, Strategic Suppressors, and Conceptual Integrators. Automatic Translators fail to detect conflicts in incongruent problems, applying direct translation indiscriminately. These students correctly solve congruent problems but make reversal errors on incongruent ones, displaying uniformly fast response times across both problem types. Conflict Recognizers recognize when direct translation conflicts with incongruent problems but fail to inhibit this inappropriate strategy despite detecting the conflict. Like Automatic Translators, they solve congruent problems correctly but make reversal errors on incongruent ones. However, they respond more slowly to incongruent problems due to the additional cognitive processing required for conflict detection. Strategic Suppressors demonstrate more sophisticated problem-solving abilities. They appropriately use direct translation for congruent problems while successfully inhibiting this strategy for incongruent problems in favor of an analytic approach. These students answer both problem types correctly but work more quickly on congruent problems. Conceptual Integrators represent the most advanced profile, handling both problem types with an efficient approach. They demonstrate consistently correct answers with similar response times across problem types, indicating genuine conceptual understanding (Reinhold et al., 2023). Beyond these four main profiles, some students may be classified as Random Guessers, responding arbitrarily to all problems regardless of type (Reinhold et al., 2023).

1.4. The Present Study

This study had two primary objectives: (a) To investigate how different sentence structures influenced Chinese students’ performance and problem representation strategies on the student-professor problem across different educational levels, particularly examining the congruency effects and testing the heuristic explanation; and (b) to profile individual differences in cognitive patterns of problem-solving using a person-centered Bayesian classification approach, thereby identifying distinct strategy profiles among students across different grade levels. To address these aims, we developed a two-alternative forced-choice (2AFC) computerized task comprising congruent and incongruent student-professor problems and recruited participants from middle school to college levels. Notably, the present study considered students’ behavioral data (accuracy and response time) simultaneously when defining the features for each class, in contrast to Reinhold et al.’s (2023) approach, in which the two response types were analyzed separately. This simultaneous approach yields more accurate and detailed classifications, as accuracy and response time are jointly generated by a single response rather than arising from two separate responses (Jiang et al., 2026). In a single-response event, accuracy and RT are not independent outputs of separate cognitive modules; rather, they are joint observable signatures of a unified underlying process—most commonly formalized as evidence accumulation. When a participant makes a single categorization decision, the same evidence accumulation dynamics that determine whether the response is correct (accuracy) also govern how long the process takes (RT). Treating them separately, as in the two-stage approach by Reinhold et al. (2023), imposes an artificial decomposition: it assumes that the cognitive information relevant to accuracy is fully captured in the first stage, leaving RT as a residual explained only by secondary factors. This misrepresents the cognitive reality in which parameters such as drift rate, boundary separation, and non-decision time simultaneously shape both accuracy and RT. We hypothesized the following: (1) Students would interpret congruent and incongruent problems as distinct problem types; (2) Sentence structures would affect reversal errors, such that students would perform better and respond faster on congruent problems compared to incongruent problems, regardless of grade level; (3) Individual differences in strategy use patterns would be found, and such individual differences would vary from grade levels.

2. Method

2.1. Participants

We used G*Power 3.1 (Faul et al., 2007) to calculate the sample size. The calculation yielded a total sample size of 76–192 (with 19–48 participants in each grade) with an effect size of 0.15–0.25 and a statistical power of 95%. To account for potential data loss, we recruited 49 seventh-grade students (27 boys, 22 girls; mean age: 13.39 ± 0.53 years), 52 eighth-grade students (27 boys, 25 girls; mean age: 13.57 ± 0.54 years) and 58 eleventh-grade students (42 boys, 16 girls; mean age: 17.22 ± 0.56 years) from a middle school in Zhaoqing, China. For young adults, we recruited 54 science and engineering (calculus-level) undergraduates or graduates (28 male, 26 female; mean age: 20.83 ± 1.89 years) from a university in Shenzhen, China. They received extra course credit for their participation. All participants reported normal or corrected-to-normal vision, had never participated in a similar experiment before, provided informed consent (or a guardian’s informed consent), and were tested in accordance with the national and international norms of using human participants. The present study was approved by the research ethics committee of the Shenzhen University.

2.2. Materials

In our study, there were 16 problems, with 8 congruent problems and 8 incongruent problems. For congruent problems, the direct translation strategy contributed to correct answers. However, correct solutions for incongruent problems conflicted with the direct translation strategy. The critical sentence of congruent problems was expressed as “The number of X times n is m times the number of Y”, while for incongruent problems it was “There are n X for every m Y”, where n and m were mutually prime integers ranging from 1 to 9 and X and Y were two categorically related objects (e.g., apples and pears). Note that when n was 1, congruent problems can be expressed as “The number of X is m times the number of Y.” Likewise, when m was 1, incongruent problems can be articulated as “There are n X for every Y.” Moreover, following previous studies (González-Calero et al., 2015; Kim et al., 2014; Soneira et al., 2018), we controlled for types of magnitudes (using discrete variables rather than continuous variables) and contextual clues (not suggesting which quantity was greater).
As shown in Table 1, in this study we developed a two-alternative, forced-choice (2AFC) computerized task. Specifically, each item provided two equations, one correctly represented the relational statement, the other was an inappropriate reverse representation. Participants were asked to select the correct equation for the student-professor problem. To validate the study instruments, we conducted pilot testing with twelve participants: Six eighth-grade students and six college students majoring in science and engineering, with equal gender representation in each group. Based on their feedback, we refined the items before administering them to the full study sample.

2.3. Procedure

Middle school students and high school students were tested in multimedia classrooms equipped with computers. As for college students, they performed the task individually in a lab. All participants were seated approximately 60 cm in front of the computers. The stimuli were presented with a screen resolution of 1280 × 768 pixels, using the E-Prime 3.0 software (Psychology Software Tools, Pittsburgh, PA, USA). To familiarize themselves with the tasks, participants were first asked to perform two practice trials without feedback, including one congruent problem and one incongruent problem. These trials were presented randomly and were not used in the formal experiment.
After completing practice trials, participants performed 16 experimental trials and the procedure was shown in Figure 1. Participants were requested to select the equation that correctly represented quantitative relationship as quickly and accurately as possible by pressing the “Q” key to choose the left equation or the “P” key to choose the right equation. Correct equations were equally distributed on the left and right sides across congruent and incongruent problems. To minimize practice effects and response patterns, we presented the experimental trials in a pseudorandom sequence, ensuring that no more than two congruent or incongruent problems appeared consecutively more than twice successively. Response times (RTs) and error rates (ERs) for all experimental trials were recorded.

2.4. Data Analysis

One participant’s data was lost during the experiment, leaving 212 students for the Bayesian classification analysis. A further two participants were excluded from the mean error rate and response time analyses because their RTs fell below 200 ms in more than 50% of trials, which may distort descriptive statistics for group-level performance (Simpson & Todd, 2017). These two participants were nonetheless retained in the Bayesian classification analysis, as this approach operates at the individual level and includes a hypothesized class of students who may adopt a random guessing strategy. At the trial level, RTs below 200 ms were removed, and RT outliers exceeding three standard deviations above or below each participant’s mean RT were also excluded (Simpson & Todd, 2017; Todd & Simpson, 2016).
Firstly, to examine how participants represented the two types of student-professor problems (congruent and incongruent), we conducted a confirmatory factor analysis (CFA) with Weighted Least Squares Mean and Variance adjusted estimation (WLSMV) using Mplus 7.0 (Muthén & Muthén, 2012). Next, we calculated the mean ERs and RTs separately for congruent and incongruent problems for each participant, and defined students’ responses for each problem by combining their accuracy and RT. Specifically, a student’s actual response to each item was defined as: Correct-faster, correct-slower, incorrect-faster, incorrect-slower, based on their accuracy and RT relative to the median RT for all problems. To determine the faster/slower categorization, we first normalized RTs by using the median RT as the reference point. Normalized RTs greater than 0 were classified as “slower”, and those less than 0 were classified as “faster”. Finally, we classified the students into distinct groups based on their response patterns (accuracy and response times) using a Bayesian classification approach.
The naive Bayesian classification approach assumes that all attributes (the evidence X) of the test samples (students) are mutually independent (Duda et al., 2012). Consequently, one student’s (e.g., the student k) posterior probabilities of all strategies can be obtained from the following formula:
P k C i | X = P k ( X | C i ) P k ( C i ) P k ( X ) = P k ( C i ) P k ( X ) j = 1 16 P k X j | C i   ( 1 j 16 ,   1 i 8 )
Note that P k ( C i | X ) expresses posterior probabilities of the student k using the strategy C i , P k ( X ) expresses the probability of the evidence X (i.e., response of the student k to each item), P k ( X | C i ) represented the conditional probability of the evidence X under the condition of the student k using the strategy Ci, P k ( C i ) expresses the probability of the student k using the strategy Ci. And then the student k is divided into the strategy based on the Bayes factors (BFs, probability ratios) from the following formula:
P k ( C i | X m ) P k ( C i + 1 | X m ) = m = 1 16 P k X m | C i P k X m | C i + 1 P k ( C i ) P k ( C i + 1 )   ( 1 m 16 ,   2 i + 1 9 )
According to Formula (1), we proposed a probability distribution for all strategies student k might use (i.e., Pk(Ci), we assigned 0.2 as an initial prior probability for every strategy since there were five patterns). This probability was then updated with each piece of evidence X (i.e., each student’s response to each item). Since the likelihoods for evidence X (i.e., P k ( X | C i ) ) had no exact probabilities, according to prior literature (Reinhold et al., 2023), we assigned 0.03 as a low probability, 0.5 as a medium (chance) probability, and 0.91 as a high probability for evidence X under the condition of student k using strategy Ci (see Table 2). For example, the probability of correctly (and faster) solving a congruent problem was 0.91 for student k using the direct translation strategy, while the likelihood was 0.03 for student k successfully (and faster) finishing an incongruent item using the direct translation strategy. The likelihood was 0.25 for the guessing strategy.
From Formula (1), we can see that if student k responded consistently to all (or most) items using one strategy, then cumulative evidence would yield a considerably increasing posterior probability for this strategy via Bayesian updating. For example, consider the order of student k’s responses on the first five student-professor problems with different responses (see Figure 2). For the first problem, the likelihood was high (0.91) for Automatic Translators, Conflict Recognizers, Strategic Suppressors, and Conceptual Integrators. Using Formula (1), the posterior probabilities for these four strategies were 0.23.1 The likelihood was medium (0.25) for Random Guessers, yielding a posterior probability of 0.06.2 Subsequently, the posterior probabilities from the first problem became the prior probabilities for the second problem, which were then updated based on the evidence from the second problem. After completing five items, the posterior probability of student k for Automatic Translators had accumulated to more than 99%. After finishing all 16 items, we obtained the final posterior probabilities.
Student k should be grouped based on the Bayes factors (probability ratios) from the Formula (2). We classified student k into the class of the strategy with the highest posterior probability if the Bayes factor of the two dominant strategies (the two strategies with the highest posterior probabilities) was greater than 3 (Jeffreys, 1961; Wagenmakers et al., 2018). For example, if student k had the highest posterior probability of 0.80 for Automatic Translators, and the second highest posterior probability of 0.10 for Conflict Recognizers, then the Bayes factor would be 0.80/0.10 = 8 and student k would be classified into the “Automatic Translators” class. If the Bayes factor of the two dominant strategies was less than 3 (Jeffreys, 1961; Wagenmakers et al., 2018), student k would be classified into the “Mixed Strategies Users” class.

3. Results

3.1. Result of a Confirmatory Factor Analysis

To evaluate the model fit, we applied the following criteria: the RMSEA should be less than 0.08, the CFI and TLI should be greater than 0.90, the χ2/df should be less than 3, and the WRMR should be less than 1 (Muthén & Muthén, 2012). Fit-comparison statistics further indicated that ΔCFI = 0.085 and ΔRMSEA = 0.137, both exceeding the recommended cutoffs for meaningful model improvement (Chen, 2007). As shown in Table 3, the two-factor model demonstrated superior fit relative to the one-factor model. All items loaded strongly on their respective factors, and the two factors were negatively correlated, supporting discriminant validity (see Appendix A Table A1). Together, these findings suggest that students perceived congruent and incongruent problems as qualitatively distinct problem types when solving the student-professor problems.

3.2. Descriptive Results

Descriptive results are shown in Table 4.

3.2.1. ERs

Shapiro–Wilk tests revealed that ERs violated the normal distribution for each group and each type of problem (ps < 0.01), thus we conducted a Wilcoxon signed-rank test. We found that participants performed better in congruent problems than incongruent problems at every grade level (see Figure 3). A Kruskal–Wallis test revealed that students’ grades were significantly related to their performance on congruent problems, χ2(3) = 8.16, p = 0.043, and incongruent problems, χ2(3) = 64.49, p < 0.001. For congruent problems, 7th graders performed worse than 11th graders and college students. However, there was no difference between 7th graders and 8th graders in the reversal error on those problems. No difference in performance was found among other grade levels. In contrast, for incongruent problems, both 7th graders and 8th graders committed more reversal errors than 11th graders and college students, while 11th graders and college students exhibited no difference in the reversal error on these problems. No significant differences in ERs were found between male and female students across all grader levels (see Appendix A Table A2).

3.2.2. RTs

Shapiro–Wilk tests confirmed that RTs were normally distributed for each group and problem type (ps > 0.05); thus we conducted a 2 (sentence structures: congruent problems, incongruent problems) × 4 (grade levels: 7th grade, 8th grade, 11th grade, college students) repeated measures ANOVA. We found a significant main effect of sentence structures, F (1, 209) = 27.33, p < 0.001, η2 = 0.17, with longer RTs for incongruent compared to congruent problems. The main effect of grades was also significant, F (3, 207) = 9.26, p < 0.001, η2 = 0.17. The Bonferroni post hoc test showed that 11th graders responded slower than the other grade levels, while 7th graders, 8th graders and college students did not differ between grade levels. The interaction between grades and sentence structures was not significant, F (3, 207) = 0.55, p = 0.65, η2 = 0.012. Further simple effect analysis suggested that 7th graders exhibited no difference between congruent and incongruent problems. However, 8th graders, 11th graders, and college students all had shorter RTs for congruent problems versus incongruent problems, p = 0.04, p < 0.001, p = 0.001, respectively (see Figure 3). No significant differences in RTs were found between male and female students across all grade levels (see Appendix A Table A3).

3.3. Individual Differences in Strategy Use

The Bayesian classification analysis revealed five distinct student profiles across all grade levels when solving the student-professor problem (see Figure 4), supporting our theoretical assumptions. However, the distribution of these profiles varied significantly by grade level. A chi-square test confirmed these grade-level differences in strategy use, χ2 (15, N = 212) = 79.92, p < 0.001. Among 7th and 8th graders, the highest posterior probabilities were observed for three profiles: Automatic Translators, Conflict Recognizers, and Random Guessers. In contrast, 11th graders and college students showed the highest posterior probabilities for Strategic Suppressors. When solving incongruent problems specifically, 7th and 8th graders relied more heavily on direct translation strategies (as Automatic Translators and Conflict Recognizers) and guessing strategies (as Random Guessers) compared to older students (see Figure 5). Conversely, 11th graders and college students more frequently employed the analytic strategy (as Strategic Suppressors) for incongruent problems. Notably, some students across all grade levels demonstrated highest posterior probabilities for multiple profiles, indicating the use of “mixed” strategies when approaching congruent or incongruent problems. These students were classified as Mixed Strategies Users. The complete distribution of students across all strategies and grades is presented in Table 5. Furthermore, we conducted a sensitivity analysis by systematically varying the three likelihood values across plausible ranges (e.g., 0.05/0.50/0.85). The results, presented in Appendix A Table A4, demonstrate that the main strategy profiles and classification outcomes remain qualitatively stable across these variations.
To further investigate the idiosyncratic strategy use of patterns of the Mixed Strategies Users group, we classified these individuals based on their dynamic shifts in strategy use across the 16-item task (see Table 6). Contrary to constituting a homogeneous group with consistent strategy use patterns, Mixed Strategies Users could be parsed into four qualitatively distinct transition profiles.
First, a subset of students exhibited the Automatic Translator to Conflict Recognizer profile (15.38% in Grade 7, 11.11% in Grade 8, and 27.27% in Grade 11). This pattern reflects a qualitative shift from the heuristic-based processing—characterized by correct-fast responses on congruent problems and incorrect-fast responses on incongruent problems—toward the conscious detection of response conflict, typically manifested as a deceleration in response times as cognitive control mechanisms engage, resulting in incorrect-slow responses on incongruent items.
Second, the Cautious Solver to Automatic Translator profile was the most prevalent among younger students (53.85% in Grade 7 and 88.89% in Grade 8), with markedly lower representation in Grade 11 (9.09%) and no occurrence among college students. Students in this profile initially employed controlled, deliberate processing on congruent problems but gradually shifted toward automatic application of the direct translation strategy, while persistently producing errors on incongruent items.
Third, the Strategic Suppressor to Conceptual Integrator profile predominated among older students (63.64% in Grade 11 and 69.23% among college students). This profile is characterized by an initial capacity to suppress the incorrect translation strategy on incongruent problems, which gradually gives way to efficient, schema-based processing that supports accurate performance on both congruent and incongruent problem types, reflected in correct-fast responses across both conditions.
Finally, a minority of students displayed the Genuinely Idiosyncratic Solver profile, observed exclusively among college students (30.77%) and one 7th-grade student (7.69%). This pattern is characterized by unsystematic fluctuations in strategy use across items, indicating the absence of a stable, rule-based response strategy and suggesting highly individualized approaches to problem solving that do not conform to any identifiable developmental trajectory.

4. Discussion

The present study had two primary aims: (1) to examine how sentence structure affected Chinese students’ performance and problem representation on the student-professor problem across educational levels; and (2) to identify individual differences in problem-solving strategies across grade level.
Our findings revealed that students across all grade levels perceived and processed congruent problems (e.g., “Students are professors six times”) and incongruent problems (e.g., “There are six students for every professor”) as fundamentally different. Consistently, students demonstrated both higher accuracy and faster response times on congruent problems compared to incongruent ones.
Most notably, we identified five distinct strategy profiles among students. Clear grade-related differences emerged: middle school students predominantly relied on direct translation strategies, whereas high school and college students demonstrated greater strategic flexibility. Specifically, these older students adapted their approach based on sentence structure, employing direct translation for congruent problems and shifting to analytic strategy for incongruent problems.
The confirmatory factor analysis revealed that Chinese students perceived congruent and incongruent problems as distinct types, despite both containing the same quantitative relationships. This finding indicates that students’ problem representations were influenced by sentence structures, suggesting they approached word problems based on surface structure (syntactic translation) rather than deep structure (quantitative relationship).
Despite recognizing these problems as distinct types, many students nonetheless attempted to apply direct translation universally, resulting in reversal errors on incongruent problems. This supports our hypothesis that sentence structure contributes to reversal errors in the student-professor problem. Notably, although Chinese syntax produced performance differences on congruent problems relative to Western peers (Fisher et al., 2011; González-Calero et al., 2020), Chinese students exhibited reversal errors on incongruent problems at rates comparable to those of their Western counterparts (Clement et al., 1981; Kim et al., 2014). We caution, however, against direct numerical comparison between our 2AFC error rates and those reported in production-task studies (Clement et al., 1981; Soneira et al., 2021). We acknowledge that our 2AFC recognition task likely yields a lower-bound estimate of reversal error prevalence and an overestimate of students’ true algebraic competence, because the recognition format reduces cognitive load in two ways: the presented alternatives provide additional cues that reduce working memory demands, and students are not required to generate a complete equation independently, thereby reducing executive processing costs. Contrary to this interpretation, Mo et al. (2026) provided validation evidence that Chinese eighth graders made fewer reversal errors on the production-task than participants in the present study who completed the 2AFC recognition task, for both congruent problems (11.44% vs. 15.74%) and incongruent problems (36.82% vs. 78.32%). This finding suggests that the 2AFC recognition format does not necessarily improve equation-solving performance.
Grade-level analysis revealed substantial differences in reversal error rates: 7th graders (78.83%) and 8th graders (78.32%) made considerably more errors than high school students (31.07%) and college students (27.39%). Given that algebraic equation concepts are introduced in Chinese 5th and 7th grades, these error patterns cannot be attributed solely to insufficient conceptual knowledge. Rather, we propose that these grade-related differences reflect variations in students’ inhibitory control abilities (Jiang & Li, 2017; Jiang et al., 2019; Lubin et al., 2016). While the direct translation strategy successfully solves congruent problems, incongruent problems require inhibiting this strategy—a process that imposes a cognitive cost reflected in longer response times. Supporting this interpretation, 11th graders and college students showed longer response times on incongruent versus congruent problems, suggesting successful inhibition of the direct translation strategy. In contrast, 7th graders showed no response time differences between problem types, indicating failure to inhibit the direct translation strategy. Because the present study did not include 9th and 10th graders, a complete developmental trajectory could not be established; future research should address this gap by incorporating a continuous grade-level sample.
The Bayesian classification analysis revealed distinct patterns of strategy use across grade levels. Among 7th and 8th graders, the predominant approach was direct translation for both congruent and incongruent problems, manifesting in two profiles: Automatic Translators and Conflict Recognizers. Automatic Translators failed to detect the differences between problem types, while Conflict Recognizers recognized the conflict between the direct translation and incongruent problems but failed to successfully inhibit this inappropriate strategy, consistent with the inhibitory control model. Notably, successful inhibition of the direct translation strategy on incongruent problems was extremely rare among middle school students with only one student each from 7th grade and 8th grade classified as Strategic Suppressors. In contrast, 11th graders and college students demonstrated markedly different strategy patterns. The Bayesian classification indicated that most of these older students successfully inhibited the direct translation strategy when solving incongruent problems. As Strategic Suppressors, they correctly solved both problem types by suppressing the misleading strategy when necessary.
The data revealed an intriguing pattern: Some students failed even on congruent problems (22.06% of 7th graders and 15.74% of 8th graders), despite presumably using the direct translation strategy. We propose this performance deficit stems from incomplete conceptual understanding of algebra. This interpretation is strengthened by the Bayesian classification results, which identified a substantial proportion of students who appeared to be guessing randomly on student-professor problems (Random Guessers: 34.69% of 7th graders and 19.61% of 8th graders).
Consistent with findings from Western samples, Chinese students similarly exhibited reversal errors on incongruent problems. Our analysis also identified a “mixed” group whose strategy use could not be classified into a single profile, suggesting that these students employed varied and flexible problem solving approaches—a pattern consistent with overlapping waves theory (Opfer & Siegler, 2007; Siegler, 1998). According to this theory, cognitive development involves increasing variability and flexibility in strategy use prior to consolidation into dominant approaches.
The Bayesian classification approach revealed important nuances that would have been obscured by aggregate analyses alone. By identifying distinct subgroups with various error patterns and strategy preferences, this fine-grained analysis provides a more complete understanding of the diverse reasoning profiles that underlie students’ performance on these problems.
This study, while offering valuable insights into the cognitive processes of algebraic word problem solving across grade levels, has several limitations. A fundamental challenge lies in inferring cognitive processes from behavioral data alone (accuracy and response times), as there remains an inherent gap between observable performance and underlying thought processes. Future research would benefit from a mixed-methods approach, combining quantitative measures like paper-pencil tests and eye tracking with qualitative methods such as clinical interviews, to build a more comprehensive understanding of students’ problem-solving cognition.
While our Bayesian classification approach successfully identified subgroups showing distinct response time patterns between congruent and incongruent problems, our interpretations of these patterns require further validation. We proposed that inhibitory control mechanisms underlie students’ ability to overcome the misleading direct translation strategy, with longer response times potentially reflecting this inhibition process. However, direct empirical evidence is needed to confirm whether slower or less accurate performance truly reflects an inhibition mechanism. Future studies could employ specialized methodologies, such as the Negative Priming paradigm, to specifically investigate the role of inhibitory control in overcoming the intuitive direct translation strategy when solving incongruent algebra word problems. Furthermore, although the naïve Bayesian classification analysis is hypothesis-driven, the distribution of the identified subgroups may be affected by a systematic sampling bias. For example, in the present study, the 11th-grade sample comprised substantially more boys than girls, and prior research has suggested that gender differences exist in inhibitory control and mathematical problem-solving. Such gender imbalance may influence the proportion of students assigned to different subgroups. Future studies should therefore aim for balanced gender representation across grade-level samples.
Finally, our findings have significant implications for algebra instruction. First, teachers should recognize the prevalence and persistence of reversal errors in students’ solutions to algebraic word problems. These errors often arise not solely from a lack of conceptual understanding but from the misapplication of intuitive translation strategies. When linguistic structures mislead them, students must inhibit these strategies to succeed. Solving such problems requires not only relevant knowledge and skills, but also the inhibitory control to override misleading heuristics or overlearned strategies. Teachers must be aware of this cognitive demand.
Second, the study underscores the importance of teachers gaining deeper insights into the diverse cognitive processes and strategy profiles students employ. For instance, slower problem-solving on incongruent problems may reflect inhibition demands, while direct guessing often signals gaps in foundational knowledge. A uniform instructional approach is unlikely to address these varied challenges effectively. Instead, curricula should focus on helping students recognize conflicts between linguistic and quantitative problem structures and strategically inhibit misleading heuristics when necessary. Such an approach could enhance algebraic reasoning instruction and address the persistent reversal errors. By understanding the range of strategies, error patterns, and cognitive demands students encounter, teachers can provide more targeted and effective support.

Author Contributions

Conceptualization, R.J. and X.L.; Methodology, Z.M.; Software, Z.M.; Formal analysis, Z.M.; Investigation, Z.M. and X.L.; Data curation, Z.M. and R.J.; Writing—original draft, Z.M.; Writing—review & editing, R.J. and X.L.; Visualization, R.J.; Supervision, X.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Ministry of Education of the People’s Republic of China (23YJC190011) and the Shenzhen Natural Science Fund (the Shenzhen Talent Program RCBS20231211090519027).

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki, and approved by the Research Ethics Committee of Shenzhen University (protocol code SZU PSY_2024_128 and date of approval: 19 September 2024).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The data presented in this study are openly available. We have made all our data publicly available on the Open Science Framework at https://osf.io/meetings/apa/ (accessed on 4 April 2026). To facilitate peer review, we provide a private read only link to view our data during the review process here https://osf.io/v3cja/ (accessed on 4 April 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1. Standardized factor loadings and inter-factor correlation for the two-factor CFA model.
Table A1. Standardized factor loadings and inter-factor correlation for the two-factor CFA model.
ItemFactor LoadingStandard ErrorCritical Ratiop-Value
Factor 1 (F1)
I10.9390.02439.117<0.001
I20.9240.02734.437<0.001
I30.950.02243.977<0.001
I40.9420.02340.238<0.001
I50.9260.02635.66<0.001
I60.9250.02734.575<0.001
I70.9540.02144.402<0.001
I80.910.0330.68<0.001
Factor 2 (F2)
I90.7970.07410.766<0.001
I100.7830.07210.86<0.001
I110.8690.05216.576<0.001
I120.7080.0858.344<0.001
I130.8970.06413.983<0.001
I140.8560.06912.406<0.001
I150.8290.0711.824<0.001
I160.7720.07510.334<0.001
Inter-factor correlation
F1 ↔ F2−0.2910.078−3.728<0.001
Table A2. Number of students classified into the hypothesized strategies in every grade (sensitivity analysis).
Table A2. Number of students classified into the hypothesized strategies in every grade (sensitivity analysis).
Strategy7th Graders
(n = 49)
8th Graders
(n = 51)
11th Graders
(n = 58)
College Students
(n = 54)
N%N%N%N%
Automatic Translators1122.452039.2235.171018.52
Conflict Recognizers918.371019.61813.7923.70
Strategic Suppressors12.0411.962543.102342.59
Conceptual Integrators0011.9658.6235.56
Random Guessers1734.691019.61610.3435.56
Mixed Strategies Users1122.45917.651118.971324.07
Table A3. Means and Standard Deviations of ERs for all students (M ± SD).
Table A3. Means and Standard Deviations of ERs for all students (M ± SD).
Dependent Variable7th Graders
(n = 48)
8th Graders
(n = 50)
11th Graders
(n = 58)
College Students
(n = 54)
Congruent
 Male17.15 ± 24.7414.27 ± 24.727.31 ± 14.998.70 ± 21.30
 Female27.86 ± 31.6417.33 ± 25.6417.81 ± 33.4315.59 ± 25.12
Incongruent
 Male74.41 ± 30.7378.54 ± 32.7025.38 ± 35.1924.89 ± 38.45
 Female82.58 ± 27.9378.08 ± 28.2846.00 ± 49.3829.89 ± 36.16
Table A4. Means and Standard Deviations of RTs for all students (M ± SD).
Table A4. Means and Standard Deviations of RTs for all students (M ± SD).
Dependent Variable7th Graders
(n = 48)
8th Graders
(n = 50)
11th Graders
(n = 58)
College Students
(n = 54)
Congruent
 Male6596.49 ± 2725.186261.55 ± 2156.648119.07 ± 2971.256411.97 ± 1811.27
 Female5391.18 ± 2338.996769.14 ± 2209.187234.09 ± 4459.696637.54 ± 2448.86
Incongruent
 Male6608.97 ± 3684.739056.01 ± 3398.5410058.44 ± 2827.308742.76 ± 2562.03
 Female7179.34 ± 4355.148700.32 ± 4666.4312059.62 ± 2877.387865.98 ± 2346.05

Notes

1
(0.2 × 0.91)/(0.2 × 0.91 + 0.2 × 0.91 + 0.2 × 0.91 + 0.2 × 0.91 + 0.2 × 0.25) = 0.23.
2
(0.2 × 0.25)/(0.2 × 0.91 + 0.2 × 0.91 + 0.2 × 0.91 + 0.2 × 0.91 + 0.2 × 0.25) = 0.06.

References

  1. Adu-Gyamfi, K., Bossé, M. J., & Chandler, K. (2015). Situating student errors: Linguistic-to-algebra translation errors. International Journal for Mathematics Teaching and Learning, 16(2), 1–29. [Google Scholar]
  2. Babai, R., Eidelman, R. R., & Stavy, R. (2012). Preactivation of inhibitory control mechanisms hinders intuitive reasoning. International Journal of Science and Mathematics Education, 10, 763–775. [Google Scholar] [CrossRef]
  3. Cai, J. (2004). Why do U.S. and Chinese students think differently in mathematical problem solving?: Impact of early algebra learning and teachers’ beliefs. The Journal of Mathematical Behavior, 23(2), 135–167. [Google Scholar] [CrossRef]
  4. Chen, F. (2007). Sensitivity of goodness of fit indexes to lack of measurement invariance. Structural Equation Modeling: A Multidisciplinary Journal, 14(3), 464–504. [Google Scholar] [CrossRef]
  5. Christianson, K., Mestre, J. P., & Luke, S. G. (2012). Practice makes (nearly) perfect: Solving ‘students-and-professors’-type algebra word problems. Applied Cognitive Psychology, 26(5), 810–822. [Google Scholar] [CrossRef]
  6. Clement, J. (1982). Algebra word problem solutions: Thought processes underlying a common misconception. Journal for Research in Mathematics Education, 13(1), 16–30. [Google Scholar] [CrossRef] [PubMed]
  7. Clement, J., Lochhead, J., & Monk, G. S. (1981). Translation difficulties in learning mathematics. The American Mathematical Monthly, 88(4), 286–290. [Google Scholar] [CrossRef]
  8. Cohen, E., & Kanim, S. E. (2005). Factors influencing the algebra “reversal error”. American Journal of Physics, 73(11), 1072–1078. [Google Scholar] [CrossRef]
  9. Culbertson, M. J. (2016). Bayesian networks in educational assessment: The state of the field. Applied Psychological Measurement, 40(1), 3–21. [Google Scholar] [CrossRef]
  10. Dooren, W. V., & Inglis, M. (2015). Inhibitory control in mathematical thinking, learning and problem solving: A survey. ZDM, 47(5), 713–721. [Google Scholar] [CrossRef]
  11. Duda, R. O., Hart, P. E., & Stork, D. G. (2012). Pattern classification. John Wiley & Sons. [Google Scholar]
  12. Evans, J. S. B. T. (2003). In two minds: Dual-process accounts of reasoning. Trends in Cognitive Sciences, 7(10), 454–459. [Google Scholar] [CrossRef]
  13. Evans, J. S. B. T., & Stanovich, K. E. (2013). Dual-process theories of higher cognition: Advancing the debate. Perspectives on Psychological Science, 8(3), 223–241. [Google Scholar] [CrossRef]
  14. Faul, F., Erdfelder, E., Lang, A.-G., & Buchner, A. (2007). G* Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods, 39(2), 175–191. [Google Scholar] [CrossRef]
  15. Fisher, K. J., Borchert, K., & Bassok, M. (2011). Following the standard form: Effects of equation format on algebraic modeling. Memory & Cognition, 39, 502–515. [Google Scholar] [CrossRef]
  16. Frey, D., Johnson, E. D., & De Neys, W. (2018). Individual differences in conflict detection during reasoning. Quarterly Journal of Experimental Psychology, 71(5), 1188–1208. [Google Scholar] [CrossRef]
  17. González-Calero, J. A., Arnau, D., & Laserna-Belenguer, B. (2015). Influence of additive and multiplicative structure and direction of comparison on the reversal error. Educational Studies in Mathematics, 89, 133–147. [Google Scholar] [CrossRef]
  18. González-Calero, J. A., Berciano, A., & Arnau, D. (2020). The role of language on the reversal error. A study with bilingual Basque-Spanish students. Mathematical Thinking and Learning, 22(3), 214–232. [Google Scholar] [CrossRef]
  19. Hegarty, M., Mayer, R. E., & Monk, C. A. (1995). Comprehension of arithmetic word problems: A comparison of successful and unsuccessful problem solvers. Journal of Educational Psychology, 87(1), 18–32. [Google Scholar] [CrossRef]
  20. Houdé, O., & Borst, G. (2014). Measuring inhibitory control in children and adults: Brain imaging and mental chronometry. Frontiers in Psychology, 5(616), 81504. [Google Scholar] [CrossRef]
  21. Houdé, O., & Borst, G. (2015). Evidence for an inhibitory-control theory of the reasoning brain. Frontiers in Human Neuroscience, 9, 148. [Google Scholar] [CrossRef]
  22. Jankvist, U. T., & Niss, M. (2021). The students-professors problem—The reversal error and beyond. Implementation and Replication Studies in Mathematics Education, 1(2), 190–226. [Google Scholar] [CrossRef]
  23. Jeffreys, H. (1961). Theory of probability (3rd ed.). Oxford University Press. [Google Scholar]
  24. Jiang, R., & Li, X. (2017). The overuse of proportional reasoning and its cognitive mechanism: A developmental negative priming study. Acta Psychologica Sinica 心理学报, 49(6), 745–758. [Google Scholar] [CrossRef]
  25. Jiang, R., Li, X., & Cai, M. (2026). From additive to multiplicative thinking: Individual cognitive patterns revealed through Bayesian classification. Journal of Educational Psychology. Advance online publication. [Google Scholar] [CrossRef]
  26. Jiang, R., Li, X., Fernández, C., & Fu, X. (2017). Students’ performance on missing-value word problems: A cross-national developmental study. European Journal of Psychology of Education, 32(4), 551–570. [Google Scholar] [CrossRef]
  27. Jiang, R., Li, X., Xu, P., & Chen, Y. (2019). Inhibiting intuitive rules in a geometry comparison task: Do age level and math achievement matter? Journal of Experimental Child Psychology, 186, 1–16. [Google Scholar] [CrossRef]
  28. Jiang, R., Li, X., Xu, P., & Lei, Y. (2020). Do teachers need to inhibit heuristic bias in mathematics problem-solving? Evidence from a negative-priming study. Current Psychology, 41(10), 6954–6965. [Google Scholar] [CrossRef]
  29. Jupri, A., & Drijvers, P. (2016). Student difficulties in mathematizing word problems in algebra. Eurasia Journal of Mathematics, Science and Technology Education, 12(9), 2481–2502. [Google Scholar] [CrossRef]
  30. Kim, S. H., Phang, D., An, T., Yi, J. S., Kenney, R., & Uhan, N. A. (2014). POETIC: Interactive solutions to alleviate the reversal error in student–professor type problems. International Journal of Human-Computer Studies, 72(1), 12–22. [Google Scholar] [CrossRef][Green Version]
  31. Knuth, E. J., Alibali, M. W., McNeil, N. M., Weinberg, A., & Stephens, A. C. (2005). Middle school students’ understanding of core algebraic concepts: Equality and variable. International Reviews on Mathematical Education, 37, 68–76. [Google Scholar] [CrossRef]
  32. Leuders, T., & Loibl, K. (2020). Processing probability information in nonnumerical settings–teachers’ Bayesian and non-Bayesian strategies during diagnostic judgment. Frontiers in Psychology, 11, 678. [Google Scholar] [CrossRef]
  33. Li, X., Xu, P., Jiang, R., & Chen, S. (2023). The role of inhibition in overcoming arithmetic natural number bias in the Chinese context: Evidence from behavioral and ERP experiments. Learning and Instruction, 86, 101752. [Google Scholar] [CrossRef]
  34. Lopez-Real, F. (1995). How important is the reversal error in algebra? In B. Atweh, & S. Flauvel (Eds.), Proceedings of the 18th annual conference of the mathematics education research group of Australasia (pp. 390–396). Mathematics Education Research Group of Australia. [Google Scholar]
  35. Lubin, A., Rossi, S., Lanoë, C., Vidal, J., Houdé, O., & Borst, G. (2016). Expertise, inhibitory control and arithmetic word problems: A negative priming study in mathematics experts. Learning and Instruction, 45, 40–48. [Google Scholar] [CrossRef]
  36. MacGregor, M., & Stacey, K. (1993). Cognitive models underlying students’ formulation of simple linear equations. Journal for Research in Mathematics Education, 24(3), 217–232. [Google Scholar] [CrossRef]
  37. Martin, S. A., & Bassok, M. (2005). Effects of semantic cues on mathematical modeling: Evidence from word-problem solving and equation construction tasks. Memory & Cognition, 33(3), 471–478. [Google Scholar] [CrossRef]
  38. McNeil, N. M., Weinberg, A., Hattikudur, S., Stephens, A. C., Asquith, P., Knuth, E. J., & Alibali, M. W. (2010). A is for apple: Mnemonic symbols hinder the interpretation of algebraic expressions. Journal of Educational Psychology, 102(3), 625–634. [Google Scholar] [CrossRef]
  39. Mevel, K., Poirel, N., Rossi, S., Cassotti, M., Simon, G., Houdé, O., & Neys, W. D. (2014). Bias detection: Response confidence evidence for conflict sensitivity in the ratio bias task. Journal of Cognitive Psychology, 27(2), 227–237. [Google Scholar] [CrossRef]
  40. Mo, Z., Jiang, R., & Li, X. (2026). Understanding individual differences in algebraic word problem solving: A latent class analysis of Chinese students’ performance across different problem formats. European Journal of Psychology of Education, 41(2), 40. [Google Scholar] [CrossRef]
  41. Muthén, L. K., & Muthén, B. O. (2012). Mplus user’s guide (version 7). Muthén & Muthén. [Google Scholar]
  42. National Council of Teachers of Mathematics. (2000). Principles and standards for school mathematics. NCTM. [Google Scholar]
  43. Opfer, J. E., & Siegler, R. S. (2007). Representational change and children’s numerical estimation. Cognitive Psychology, 55(3), 169–195. [Google Scholar] [CrossRef]
  44. Passolunghi, M. C., Blas, G. D. D., Carretti, B., Gomez-Veiga, I., Doz, E., & Garcia-Madruga, J. A. (2022). The role of working memory updating, inhibition, fluid intelligence, and reading comprehension in explaining differences between consistent and inconsistent arithmetic word-problem-solving performance. Journal of Experimental Child Psychology, 224, 105512. [Google Scholar] [CrossRef]
  45. Pawley, D., & Cooper, M. (1997). What can be done to overcome the multiplicative reversal error? In E. Pehkonen (Ed.), Proceedings of the 21st conference of the international group for the psychology of mathematics education (pp. 320–327). University of Helsinki. [Google Scholar]
  46. Pennycook, G., Fugelsang, J. A., & Koehler, D. J. (2012). Are we good at detecting conflict during reasoning? Cognition, 124(1), 101–106. [Google Scholar] [CrossRef] [PubMed]
  47. Reinhold, F., Leuders, T., & Loibl, K. (2023). Disentangling magnitude processing, natural number biases, and benchmarking in fraction comparison tasks: A person-centered Bayesian classification approach. Contemporary Educational Psychology, 75, 102224. [Google Scholar] [CrossRef]
  48. Reinhold, F., Obersteiner, A., Hoch, S., Hofer, S. I., & Reiss, K. (2020). The Interplay between the natural number bias and fraction magnitude processing in low-achieving students. Frontiers in Education, 5, 29. [Google Scholar] [CrossRef]
  49. Siegler, R. S. (1998). Emerging minds: The process of change in children’s thinking. Oxford University Press. [Google Scholar]
  50. Simpson, A. J., & Todd, A. R. (2017). Intergroup visual perspective-taking: Shared group membership impairs self-perspective inhibition but may facilitate perspective calculation. Cognition, 166, 371–381. [Google Scholar] [CrossRef] [PubMed]
  51. Soneira, C., Arnau, D., & González-Calero, J. A. (2018). An assessment of the sources of the reversal error through classic and new variables. Educational Studies in Mathematics, 99, 43–56. [Google Scholar] [CrossRef]
  52. Soneira, C., Bansilal, S., & Govender, R. (2021). Insights into the reversal error from a study with South African and Spanish prospective primary teachers. Pythagoras, 42(1), a613. [Google Scholar] [CrossRef]
  53. Stupple, E. J., Ball, L. J., & Ellis, D. (2013). Matching bias in syllogistic reasoning: Evidence for a dual-process account from response times and confidence ratings. Thinking & Reasoning, 19(1), 54–77. [Google Scholar] [CrossRef]
  54. Todd, A. R., & Simpson, A. J. (2016). Anxiety impairs spontaneous perspective calculation: Evidence from a level-1 visual perspective-taking task. Cognition, 156, 88–94. [Google Scholar] [CrossRef] [PubMed]
  55. Vamvakoussi, X., & Vosniadou, S. (2004). Understanding the structure of the set of rational numbers: A conceptual change approach. Learning and Instruction, 14(5), 453–467. [Google Scholar] [CrossRef]
  56. Vanluydt, E., Verschaffel, L., & Dooren, W. V. (2022). The role of relational preference in early proportional reasoning. Learning and Individual Differences, 93, 102108. [Google Scholar] [CrossRef]
  57. Wagenmakers, E. J., Marsman, M., Jamil, T., Ly, A., Verhagen, J., Love, J., Selker, R., Gronau, Q. F., Šmíra, M., Epskamp, S., Matzke, D., Rouder, J. N., & Morey, R. D. (2018). Bayesian inference for psychology. Part I: Theoretical advantages and practical ramifications. Psychonomic Bulletin & Review, 25(1), 35–57. [Google Scholar] [CrossRef]
  58. Wollman, W. (1983). Determining the sources of error in a translation from sentence to equation. Journal for Research in Mathematics Education, 14(3), 169–181. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Procedure of experimental trials.
Figure 1. Procedure of experimental trials.
Behavsci 16 00953 g001
Figure 2. One example of naive Bayesian classification approach on students’ response. Note. Congruent problems: the 1st item, 2nd item, and 4th item; incongruent problems: the 3rd item and 5th item.
Figure 2. One example of naive Bayesian classification approach on students’ response. Note. Congruent problems: the 1st item, 2nd item, and 4th item; incongruent problems: the 3rd item and 5th item.
Behavsci 16 00953 g002
Figure 3. Error rates (a) and response times (b) for students. Error bars mean the standard error of the mean. ns = non-significant; * p < 0.05; ** p < 0.01; *** p < 0.001.
Figure 3. Error rates (a) and response times (b) for students. Error bars mean the standard error of the mean. ns = non-significant; * p < 0.05; ** p < 0.01; *** p < 0.001.
Behavsci 16 00953 g003
Figure 4. Posterior probabilities of five distinct profiles of students in different grades.
Figure 4. Posterior probabilities of five distinct profiles of students in different grades.
Behavsci 16 00953 g004
Figure 5. Number of students (%) classified into the hypothesized strategies in every grade.
Figure 5. Number of students (%) classified into the hypothesized strategies in every grade.
Behavsci 16 00953 g005
Table 1. Examples of congruent problems and incongruent problems.
Table 1. Examples of congruent problems and incongruent problems.
Congruent ProblemsIncongruent Problems
In the classroom, the number of boys is four times the number of girls.
Use X for the number of boys and Y for the number of girls, then:
In the zoo, there are five tigers for every lion.
Use X for the number of tigers and Y for the number of lions, then:
4X = YX = 4YX = 5Y5X = Y
In the basket, the number of apples times five is seven times the number of pears.
Use X for the number of apples and Y for the number of pears, then:
In the basket, there are three oranges for every five peaches.
Use X for the number of oranges and Y for the number of peaches, then:
5X = 7Y7X = 5Y3X = 5Y5X = 3Y
Table 2. The likelihood for the evidence X under the condition of hypothesized strategies.
Table 2. The likelihood for the evidence X under the condition of hypothesized strategies.
Sentence StructuresCongruent ProblemsIncongruent Problems
AccuracyCorrectFalseCorrectFalse
Response TimeFasterSlowerFasterSlowerFasterSlowerFasterSlower
Automatic Translators0.910.030.030.030.030.030.910.03
Conflict Recognizers0.910.030.030.030.030.030.030.91
Strategic Suppressors0.910.030.030.030.030.910.030.03
Conceptual Integrators0.910.030.030.030.910.030.030.03
Random Guessers0.250.250.250.250.250.250.250.25
Table 3. Model fit indices for two hypothesized models.
Table 3. Model fit indices for two hypothesized models.
ModelRMSEACFITLIχ2dfWRMR
One-factor model0.1610.9130.900674.2101042.337
Two-factor model0.0240.9980.998116.0111030.760
Table 4. Means and Standard Deviations of ERs (%) and RTs (ms) for all students (M ± SD).
Table 4. Means and Standard Deviations of ERs (%) and RTs (ms) for all students (M ± SD).
Dependent Variable7th Graders
(n = 48)
8th Graders
(n = 50)
11th Graders
(n = 58)
College Students
(n = 54)
Error rates
Congruent22.06 ± 29.2215.74 ± 24.9610.21 ± 21.8712.15 ± 23.33
Incongruent78.83 ± 28.3278.32 ± 30.3531.07 ± 40.2427.39 ± 37.05
Response times
Congruent6057.95 ± 2604.186499.81 ± 2173.677886.18 ± 3404.476524.75 ± 2135.61
Incongruent6647.53 ± 4052.348858.40 ± 4079.7410,449.98 ± 2917.368277.53 ± 2463.86
Table 5. Number of students classified into the hypothesized strategies in every grade.
Table 5. Number of students classified into the hypothesized strategies in every grade.
Strategy7th Graders
(n = 49)
8th Graders
(n = 51)
11th Graders
(n = 58)
College Students
(n = 54)
N%N%N%N%
Automatic Translators1122.452039.2235.171018.52
Conflict Recognizers918.371019.61813.7923.70
Strategic Suppressors12.0411.962543.102342.59
Conceptual Integrators0011.9658.6235.56
Random Guessers1734.691019.61610.3435.56
Mixed Strategies Users1122.45917.651118.971324.07
Table 6. Number of students (%) of different strategy transition profiles across grade levels in the mixed-strategies-users group.
Table 6. Number of students (%) of different strategy transition profiles across grade levels in the mixed-strategies-users group.
Strategic Profiles7th Grade8th Grade11th GradeCollege
Automatic Translator to Conflict Recognizer2 (15.38)1 (11.11)3 (27.27)0 (0.00)
Cautious Solver to Automatic Translator7 (53.58)8 (88.89)1 (9.09)0 (0.00)
Strategic Suppressor to Conceptual Integrator1 (7.69)0 (0.00)7 (63.64)9 (69.23)
Genuinely Idiosyncratic Solver1 (7.69)0 (0.00)0 (0.00)4 (30.77)
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Mo, Z.; Jiang, R.; Li, X. Strategy Profiles in Solving Algebra Word Problems: A Person-Centered Bayesian Classification Approach for Chinese Students. Behav. Sci. 2026, 16, 953. https://doi.org/10.3390/bs16060953

AMA Style

Mo Z, Jiang R, Li X. Strategy Profiles in Solving Algebra Word Problems: A Person-Centered Bayesian Classification Approach for Chinese Students. Behavioral Sciences. 2026; 16(6):953. https://doi.org/10.3390/bs16060953

Chicago/Turabian Style

Mo, Zongzhao, Ronghuan Jiang, and Xiaodong Li. 2026. "Strategy Profiles in Solving Algebra Word Problems: A Person-Centered Bayesian Classification Approach for Chinese Students" Behavioral Sciences 16, no. 6: 953. https://doi.org/10.3390/bs16060953

APA Style

Mo, Z., Jiang, R., & Li, X. (2026). Strategy Profiles in Solving Algebra Word Problems: A Person-Centered Bayesian Classification Approach for Chinese Students. Behavioral Sciences, 16(6), 953. https://doi.org/10.3390/bs16060953

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop