Next Article in Journal
The Use of Internal State Terms by Individuals with Autism Spectrum Disorders: A Scoping Review
Previous Article in Journal
Orthographic Depth and Spelling Development in Immersion Education: A Predictive Framework of Spelling Errors in French
Previous Article in Special Issue
A Study of Grammatical Gradience in Relation to the Distributional Properties of Verbal Nouns in Scottish Gaelic
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Semantic Processing and Individual Variation: Experimental and Modeling Evidence from Quantifier Scope

1
Department of English, Purdue University, West Lafayette, IN 47907, USA
2
Department of Linguistics, Montclair State University, Montclair, NJ 07043, USA
3
Department of Psychology, Montclair State University, Montclair, NJ 07043, USA
4
Department of Linguistics, Purdue University, West Lafayette, IN 47907, USA
5
School of Languages and Cultures, Purdue University, West Lafayette, IN 47907, USA
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Languages 2026, 11(6), 126; https://doi.org/10.3390/languages11060126
Submission received: 9 April 2026 / Revised: 2 June 2026 / Accepted: 3 June 2026 / Published: 17 June 2026

Abstract

This study investigates the real-time processing of quantifier scope ambiguity using self-paced reading, comparing surface and inverse interpretations of a–every sentences. Reaction time data revealed greater processing difficulty for inverse scope, supporting an account grounded in semantic computation rather than purely heuristic-based parsing. Offline interpretation results further indicate that abstract structural operations are engaged during scope computation. Surprisal estimates derived from a pre-trained autoregressive language model GPT-2 small successfully predicted both the locus and direction of scope effects observed in the human data, with partial convergence in effect magnitude. When surprisal was not considered, language experience, but not working memory, more robustly accounted for variability in complex scope interpretations. Crucially, incorporating individual differences into surprisal-based analyses showed that both language experience and working memory capacity modulate surprisal effects in scope processing: higher proficiency and greater working memory are associated with increased sensitivity to surprisal, whereas lower levels show reduced or even reversed effects, suggesting weaker engagement in expectation-based processing. Together, these findings highlight the interplay among structural complexity, cognitive resources, language experience, and expectation-based mechanisms in shaping real-time semantic processing.

Graphical Abstract

1. Introduction

Ambiguity is pervasive in human language at both lexical and sentential levels. Understanding how people interpret ambiguous sentences reveals how multiple information sources are integrated and how this process interacts with other cognitive mechanisms. Most psycholinguistic research on sentence-level ambiguity has focused on syntactic ambiguity (e.g., The man saw the boy with a binocular; Altmann, 1998; Ferreira & Çokal, 2016; Pickering et al., 2000). The present study focuses on semantic ambiguity that arises from the interaction of logical operators (e.g., every, a), a relatively understudied phenomenon in psycholinguistic research of ambiguity, as illustrated in quantifier scope interpretation and processing. For example, the sentence A dog chased every cat permits two interpretations: (i) a single dog chased all the cats, or (ii) each cat was chased by a different dog.
There are three major motivations for the current study. First, the study aimed to investigate how the human parser resolves semantic ambiguity by integrating linguistic information during sentence processing, as indexed by the relative processing costs associated with computing inverse versus surface scope interpretations. Second, within expectation-based models of sentence processing (Hale, 2001; Levy, 2008), we employed model-generated surprisal value, an expectation-based metric, to examine its relationship with behavioral measures of processing difficulty, operationalized as reaction times (RTs). Specifically, we assessed the extent to which surprisal predicts semantic processing costs in quantifier scope resolution. Third, the study sought to account for individual differences in semantic processing attributable to working memory capacity and language experience, building on the view that variability is an inherent property of natural language grammar, particularly for phenomena characterized by input underdetermination, such as quantifier scope interpretation (Han et al., 2016; Zimmermann et al., 2025). In line with this perspective, recent work has begun to examine how individual differences modulate surprisal effects in language processing, including the relationship between surprisal’s predictive power and cognitive capacities (Haller et al., 2024) as well as language experience (Berzak & Levy, 2023). To this end, we also modeled the moderating effects of working memory capacity and language experience on the predictive power of surprisal in scope processing.

2. Literature Review

2.1. Scope Processing: Theories and Psycholinguistic Evidence

English sentences containing two logical operators, including an existential quantifier (a) and a universal quantifier (every), often give rise to scope ambiguity. For example, in (1), where the existential quantifier a appears in the subject position and the universal quantifier every in the object position, the sentence permits two interpretations. Under the surface scope reading (SSR) in (1a), the semantic scope mirrors the syntactic structure: the existential quantifier takes scope over the universal quantifier, with the scope relation directly determined by their surface positions. By contrast, the inverse scope reading (ISR) in (1b) involves the universal quantifier taking scope over the existential quantifier, such that the semantic scope relation does not align with the surface c-command relations, arguably due to a covert syntactic operation of Quantifier Raising (QR) at Logical Form (LF) (May, 1985).
(1)A dog chased every cat.
a.There was a single dog that chased each cat. (Surface scope)
b.For each cat, there was a different dog that chased it. (Inverse scope)
Although both interpretations are licensed by the grammar, they are not equally accessible at the computational level, with surface scope generally preferred over inverse scope. It is widely held that inverse scope incurs greater processing difficulty than surface scope; however, accounts differ as to the source of this cost. From a syntactic perspective, Anderson (2004) proposed the Processing Scope Economy (PSE) principle, according to which inverse scope representations are more complex because they require additional covert operations at Logical Form, thereby increasing processing cost. A closely related proposal, the Minimal Lowering Hypothesis (e.g., Fox, 2000), similarly assumes that surface scope constitutes the default interpretation, with inverse scope available only when required by contextual or pragmatic considerations and derived via non-default movement operations, rendering it less preferred.
In contrast to these syntactically grounded explanations, Kurtzman and MacDonald (1993) attributed scope preferences to principles of incremental sentence processing, arguing that the parser resolves ambiguity by integrating multiple probabilistic constraints as linguistic input unfolds (Trueswell et al., 1994). Central to this view is the Single Reference Principle, according to which comprehenders commit early to a single discourse referent upon encountering a singular indefinite in sentence-initial position (e.g., a dog). This commitment naturally supports a surface scope interpretation, which is compatible with a single referent, while disfavoring an inverse scope interpretation, which would require constructing multiple referents (e.g., multiple dogs).
Setting aside theoretical differences among accounts, empirical findings consistently point to a preference for surface scope among native speakers of English (Chu et al., 2014; Kurtzman & MacDonald, 1993; Scontras et al., 2017; Wu & Ionin, 2022). This preference has been documented primarily using variants of truth-value judgment tasks widely employed in semantic research (Fang & Francis, 2025). For example, using a sentence–picture matching task, Scontras et al. (2017) asked English native speakers to rate the match between sentences such as (1) and pictures depicting surface and inverse scope interpretations ((1a) and (1b), respectively) on a 7-point Likert scale. The results revealed significantly higher ratings for the surface scope interpretation than for the inverse scope interpretation, even though both interpretations were judged acceptable, with mean ratings exceeding 5.
In psycholinguistics, a much-studied question concerns how speakers resolve potential scope ambiguities during real-time comprehension, as investigated using time-sensitive behavioral methods such as self-paced reading and eye-tracking (Anderson, 2004; Brasoveanu & Dotlačil, 2015; Paterson et al., 2008; Zhou & Gao, 2009). These methods make it possible to probe initial scope interpretation more directly by measuring online processing difficulty (cf. Kurtzman & MacDonald, 1993). Crucially, they do so without explicitly drawing participants’ attention to alternative interpretations, instead comparing RTs across continuations that bias different interpretations of the same sentence.
A number of studies have examined the real-time processing of quantified sentences in simple transitive clauses in English (Dwivedi, 2013) and Chinese (Zhou & Gao, 2009). However, these studies have largely focused on configurations in which every precedes a, as in Every child climbed a tree. Using the sentence continuation paradigm introduced by Kurtzman and MacDonald (1993), they reported shorter RTs in regions disambiguating toward the surface scope interpretation relative to inverse scope, a pattern commonly taken to reflect increased processing costs associated with inverse scope.
Nevertheless, the reliance on every–a sentences to assess the availability or processing cost of inverse scope is problematic due to an entailment asymmetry. In such sentences, the inverse scope interpretation entails the surface scope interpretation; consequently, cases classified as involving inverse scope necessarily also satisfy the surface scope reading. This entailment relation renders every–a sentences poor diagnostics for isolating inverse scope, as any processing cost attributed to inverse scope may be attenuated or obscured by the concurrent availability of surface scope (Cowley et al., 2025; Scontras et al., 2017). Therefore, a–every sentences provide a more suitable test case for probing both the availability and the online computation of inverse scope interpretations.
Relatively few studies have examined the real-time processing of a–every configurations to test the claim that inverse scope is more costly. A notable exception is Anderson (2004), which found that native English speakers had greater difficulty accessing inverse scope than surface scope in a self-paced reading task. However, this finding is complicated by a potential morphological mismatch confound. Participants read a quantified sentence, repeated in (2), followed by a continuation that disambiguated toward either a surface scope (2a) or inverse scope interpretation (2b):
(2)A dog chased every cat.
a.… the dog was not very fast.
b.… the dogs were not very fast.
If scope had been determined by the time readers encountered the continuation, longer RTs for (2b) than (2a) would support a processing cost for inverse scope. However, this comparison is confounded by a number mismatch between “a dog” and “the dogs.” Thus, the longer RTs in (2b) may reflect morphological inconsistency under heuristic-driven good-enough processing (cf. Christianson, 2016; Ferreira, 2003; Fang et al., 2026a), rather than genuine scope computation.
To address this issue, Anderson introduced control conditions embedded in richer discourse contexts, which may have reduced inverse scope difficulty. Building on this concern, we introduce two control conditions (3) and (4) and compare them directly with the critical conditions. Stronger evidence for scope-driven processing would be obtained if longer RTs persist for (2b) than (2a), while no corresponding difference emerges between (3) and (4), thereby ruling out a heuristic-driven good-enough processing account.
(3)The same dog chased every cat, although the dog was not very fast.
(4)A different dog chased every cat, although the dogs were not very fast.

2.2. Surprisal in Sentence Processing

In cognitive science, one influential account of sentence comprehension and processing is the expectation-based model (Hale, 2001; Levy, 2008). Under this framework, the human parser incrementally constructs syntactic and semantic representations by actively generating probabilistic expectations about upcoming linguistic input, using these expectations to guide real-time parsing decisions. Within this approach, surprisal has been proposed as a formally defined computational metric of processing difficulty (Hale, 2001, 2016). Surprisal is defined as the negative log probability of a word given its preceding context and is assumed to be proportional to processing effort, such that words with higher surprisal incur greater cognitive cost.
With advances in machine learning and natural language processing, surprisal is now often estimated using neural Transformer-based language models. Surprisal derived from such models has been shown to more closely approximate human sentence processing behavior than surprisal estimated from n-gram models or context-free grammars (Aurnhammer & Frank, 2019; Wilcox et al., 2020). In the present study, we adopt the linking assumption between human sentence processing and next-word prediction in language models, following Futrell and Mahowald’s (2025) argument that, despite their fundamental differences, human language processors and language models are functionally similar in that both incrementally encode, decode, and predict linguistic information during comprehension. Psycholinguists have long investigated incremental processing through garden-path sentences, in which an initially preferred interpretation may later be revised. Expectation-based accounts have been shown to successfully capture such phenomena, and recent work has demonstrated that language models can reproduce key aspects of human incremental parsing behavior, including sensitivity to filler-gap dependencies (e.g., Kobzeva & Kush, 2024). In a similar vein, the sentences examined in the present study, such as (2), involve competing scope interpretations that emerge during incremental comprehension. Upon encountering “a dog”, comprehenders are likely to initially construct a representation involving a single referent, corresponding to the surface scope interpretation. Consequently, continuations containing the singular form “dog” are expected to receive higher probability and lower surprisal than continuations containing the plural form “dogs”, which is more compatible with an inverse scope interpretation. Encountering “dogs” therefore requires revision of the initially preferred interpretation, analogous to the reanalysis observed in garden-path sentence processing.
On this view, surprisal provides an empirically tractable link between distributional properties of linguistic input and the representations constructed during online comprehension. It thus holds promise for accounting for processing at multiple levels of linguistic representation, including syntactic and semantic structure. However, the extent to which model-generated word surprisal can capture specific incremental sentence processing phenomena remains an open empirical question—both in terms of whether it can qualitatively predict observed processing patterns and how well it can do so quantitatively (cf. Cong et al., 2023; Demberg & Keller, 2008; Oh et al., 2022; Smith & Levy, 2013; Wilcox et al., 2023).
Much work examining both the qualitative and quantitative fit between surprisal and human sentence processing has primarily focused on syntactic phenomena, particularly garden-path effects (e.g., Huang et al., 2024; Kobzeva & Kush, 2024; Van Schijndel & Linzen, 2021). These studies have shown that surprisal often succeeds in qualitatively capturing garden-path effects, in the sense that it correctly localizes processing difficulty to the regions where experimental effects are expected to arise during incremental sentence comprehension. However, more fine-grained quantitative evaluations reveal that (model-generated) surprisal tends to underestimate the magnitude of some garden-path effects observed in human data. Huang et al.’s (2024) large-scale self-paced reading benchmark (i.e., SAP) provides a systematic pipeline by evaluating psycholinguistic alignment across multiple syntactic constructions by comparing model-generated surprisal with human RTs using a linking function that maps surprisal and lexical predictors onto reading-time milliseconds. Their results show that while surprisal qualitatively tracks processing difficulty at disambiguating regions, it consistently underestimates effect magnitudes and fails to account for item-level variation, suggesting that probabilistic word-level expectations alone are insufficient to fully explain the cognitive cost of syntactic reanalysis.
Motivated by those findings, the current study extends prior work on syntactic processing to investigate the predictive power of word-level surprisal for semantic processing, focusing on quantifier scope interpretation, which involves linguistic operations at the level of Logical Form. Following the modeling pipeline described in Huang et al. (2024), we aim not only to qualitatively capture scope-related processing effects, but also to quantitatively assess the extent to which surprisal accounts for the observed RTs variation, that is, the proportion of overall processing cost attributable to probabilistic expectations over words.

2.3. Individual Variation in Grammatical Representation and Processing

Grammatical representation and processing are increasingly recognized as variable across native speakers rather than uniform (Polinsky, 2025; Zimmermann et al., 2025). This variability is especially pronounced in domains involving ambiguity and processing complexity (Roberts, 2012), where clear input is limited. Quantifier scope interpretation exemplifies this pattern, with substantial evidence showing systematic differences among native speakers in access to inverse scope (Han et al., 2016; Philipp & Zimmermann, 2025). Crucially, this variation is systematic rather than random, reflecting individual differences in factors such as working memory capacity (WMC) and language experience (LE).
In an offline picture-selection task, Fang and Wang (2025) found that WMC significantly affected English speakers’ scope interpretations, though not uniformly across configurations. Similarly, Philipp and Zimmermann (2025) attributed cross-linguistic differences in inverse scope availability, as well as interindividual variation, to differences in input exposure. Related evidence comes from Korean (Han et al., 2016) and Chinese (Fang et al., 2025), though these studies rely exclusively on offline measures. Beyond memory capacity, variation in language processing has also been linked to differences in linguistic experience (MacDonald & Christiansen, 2002). Considering both WMC and LE allows us to evaluate capacity-based and experience-based accounts of systematic variation in processing.
The present study extends this line of work by employing an online self-paced reading paradigm to examine how individual differences in WMC and LE manifest during real-time processing. More broadly, our focus on WMC and LE effects aligns with a gradient view of grammatical representation and interpretation (Francis, 2022), under which variability in scope interpretation and processing is expected rather than exceptional.
Research on individual differences extends beyond behavioral data to their interaction with surprisal-based predictions of reading times. Surprisal has typically been used as a group-level proxy for processing effort, with limited attention to its ability to predict behavior across individuals. Recent work has begun to address this gap (Berzak & Levy, 2023; Haller et al., 2024; Škrjanec & Demberg, 2026). For instance, Berzak and Levy (2023) show that higher L2 proficiency is associated with greater sensitivity to surprisal, while Škrjanec and Demberg (2026) demonstrate that domain-adapted language models improve RT predictions for expert readers. Building on this work, the present study examines whether modeling online processing benefits from surprisal estimates informed by individual differences in WMC and LE.

3. The Present Study

Through a combination of experimental and computational approaches, the present study addressed the following three research questions (RQs):
RQ1.
Does inverse scope incur greater processing costs than surface scope, as reflected in real-time ambiguity resolution during online sentence processing?
RQ2.
To what extent do surprisal differences between conditions predict (a) the locus of scope effects at a qualitative level and (b) the magnitude of those effects at a quantitative level in the behavioral data?
RQ3.
To what extent do individual differences in working memory capacity and language experience account for variability in the processing of scope-ambiguous sentences? Moreover, do these capacity- and experience-based differences modulate the influence of surprisal on scope-related processing effects?
Figure 1 outlines the methodological framework of the study. Participants complete a self-paced reading task on both experimental and filler items, yielding observed reaction times. In parallel, GPT-2 small is used to compute word-by-word surprisal, which is mapped to predicted RTs via a linking function trained on filler items. Comparing predicted and observed reading times allows us to estimate the magnitude and locus of quantifier scope effects. We further examine how individual differences modulate both behavioral responses and their alignment with model-based predictions.

4. Experiment 1: Self-Paced Reading with Human

4.1. Participants

Fifty-one native speakers of English (24 female; mean age = 34.4 years, SD = 8.1) participated in the experiment, which was approved by the Institutional Review Board (IRB) of the authors’ institution. Forty-five participants were recruited via Prolific (https://www.prolific.com/) and completed the experiment remotely on their own computers. The remaining six participants were recruited from a university in the northeastern United States and also completed the experiment online. All participants also completed the LexTALE vocabulary test and a backward digit span task as measures of English proficiency and verbal working memory, respectively. Table 1 below summarizes participants’ demographic and individual difference characteristics. All participants received monetary compensation.

4.2. Design and Materials

Forty-eight sets of experimental items were constructed, each comprising four conditions: two double-quantifier items and two baseline items, as illustrated in (5). Conditions (5a) and (5b) shared identical context sentences but differed in their continuation sentences, with (5a) disambiguating toward an SSR and (5b) toward an ISR. Conditions (5c) and (5d) served as baselines: in (5c), the continuation sentence was semantically compatible with a context sentence containing a semantically singular subject, whereas in (5d), it was compatible with a semantically plural subject.
ISR is assumed to be computationally more complex than SSR and thus to incur greater processing costs, reflected in longer RTs for (5b) than for (5a). To provide evidence for genuine semantic processing and rule out shallow processing accounts, no RT differences are expected between the baseline conditions (5c) and (5d).
Regions of interest (ROIs) were identical across the four conditions and comprised three regions: the noun dog/dogs as the critical region, followed by two spillover regions (was/were and not). We included two spillover regions because the first spillover region was relatively short in terms of character length; therefore, the second spillover region was added to optimally capture the potential experimental effects. All context sentences contained action verbs in the past tense.
(5)a.A dog chased every cat, although the dog was not very fast. (surface)
b.A dog chased every cat, although the dogs were not very fast. (inverse)
c.The same dog chased every cat, although the dog was not very fast. (same)
d.A different dog chased every cat, although the dogs were not very fast. (different)
The 48 sets of experimental items were distributed across four lists using a Latin Square design, such that each list contained only one item per condition from each set. Consequently, each participant saw 12 items per condition (i.e., 48 experimental items in total), together with 48 filler items. Filler sentences were approximately matched to the experimental items in length and varied in complexity (M = 13.90 words, SD = 2.73, range = 10–21). The fillers were adapted from Huang et al. (2024) and comprised four structural types: garden-path sentences, verb agreement constructions, relative clauses, and attachment ambiguity structures. Each structure included two conditions (garden-path vs. non–garden-path; agreement vs. disagreement; subject vs. object relatives; high vs. low attachment), with 12 items per structure (6 per condition). The filler items were distributed across two lists using a Latin Square design, resulting in 6 items per condition for each structure and a total of 48 filler items per list. To align with the four experimental lists, each filler list was paired with two experimental lists.

4.3. Procedure

The experiment was programmed and administered using Gorilla (Anwyl-Irvine et al., 2020). After providing informed consent, participants completed three tasks in sequence: an extended LexTALE task, a SPR task, and a backward digit span task. LexTALE is a lexical decision task originally developed by Lemhöfer and Broersma (2012) that uses knowledge of English lexical items as a proxy for English proficiency. Because LexTALE scores have been shown to correlate strongly with language experience across multiple aspects of language skill (Lemhöfer & Broersma, 2012), we used this measure to approximate individual differences in language experience that may underlie variability in scope processing and interpretation. Rather than using the original version of LexTALE, we adopted an extended version developed by Fiorentino et al. (2024), which adds 20 low-frequency real words and 10 additional nonwords to the original 40 words and 20 nonwords, yielding a total of 90 items (60 words and 30 nonwords). This extended version was intended to more sensitively capture variability in language experience among native speakers of English. In the LexTALE task, participants were asked to decide whether a visually presented letter string was a real English word by responding Yes or No. Participants were instructed to respond Yes if they believed the item to be a real English word, even if they did not know its meaning.
The LexTALE task was followed by a SPR task. In this task, participants read sentences using a word-by-word, moving-window paradigm (Just et al., 1982). After each sentence, participants answered a comprehension question presented on a separate screen with a binary response choice. For example, the comprehension question How many dogs chased cats? followed the sentences in (5a), with the response options one and several. Participants’ responses were taken to reflect their interpretation of the target sentence. For conditions (5a) and (5b), selecting one indicated an SSR, whereas selecting several indicated an ISR. For experimental items, both response options reflected grammatically licensed interpretations, so neither was treated as correct or incorrect. The left–right positions of the response options were counterbalanced across items to minimize response bias. Prior to the experimental trials, participants completed a block of five practice trials to familiarize themselves with the SPR procedure. No feedback was provided for any test items.
The next task was a backward digit span task used to measure working memory, adapted from the Working Memory Test Battery (Massonnié et al., 2022). On each trial, a sequence of digits was presented visually, one at a time, in the center of the screen. Digits were displayed at a rate of one per second (e.g., 2, 5, 3), and participants were instructed to recall the sequence in reverse order (e.g., 3, 5, 2) by clicking the digits in the corresponding on-screen boxes. Sequence length increased following successful recall, with higher spans indicating greater working memory capacity.
The experiment concluded with a language background questionnaire eliciting demographic information, including age, gender, and native language. The whole experiment took no more than one hour.

4.4. Analysis

Prior to analysis, outliers were excluded based on participants’ comprehension-question performance and implausible RTs. One participant was removed because their mean accuracy on a subset of unambiguous filler items was ≤75%. The remaining 50 participants were retained for analysis and demonstrated high overall comprehension accuracy (M = 94.3%, SD = 6.0%). At the trial level, RTs shorter than 100 ms or longer than 3000 ms were excluded. Visual inspection of the RT distribution indicated that the vast majority of observations fell within this range, and this trimming procedure resulted in the removal of 0.8% of the data. No additional trimming based on standard-deviation cutoffs was applied, following recommendations by Nicklin and Plonsky (2020), who caution that such procedures may eliminate legitimate observations. RTs were analyzed in their raw form and were not log-transformed, following Huang et al. (2024). This approach was adopted to preserve the interpretability of effect magnitudes in both human RT data and surprisal-derived RT measures. By retaining RTs in milliseconds, the size of experimental effects can be interpreted directly and compared more transparently across the two measures.
RTs were analyzed separately for each of the three ROIs using linear mixed-effects models (Bates et al., 2015). Binary response data from the comprehension questions were analyzed using logistic mixed-effects models (Jaeger, 2008). Both were fitted in R with the lme4 package (R version 4.5.1; R Core Team, 2025). Scope conditions and baseline conditions were analyzed in separate models. Following the recommendations of Schad et al. (2020), categorical predictors were contrast-coded and centered on the grand mean. Specifically, condition was coded as 0.5 (surface scope) and −0.5 (inverse scope for the scope analysis; same vs. different for the baseline analysis). These predictors were entered as fixed effects in all models.
Random-effects structures were initially specified to be maximal, as justified by the experimental design (Barr et al., 2013), including by-participant and by-item random intercepts, by-participant random slopes for condition, and by-item random slopes for between-participant predictors (i.e., WMC and LE), where applicable. When models failed to converge, the random-effects structure was simplified by first removing correlations among random effects and then iteratively removing the random effect accounting for the least variance until convergence was achieved. P-values were obtained using the Satterthwaite approximation to the degrees of freedom as implemented in the lmerTest package (Kuznetsova et al., 2017). Planned comparisons were conducted using the emmeans package (Lenth, 2025). All data and codes used for this study are available at the Open Science Framework (OSF): https://doi.org/10.17605/OSF.IO/GZ9CK.

4.5. Results

We first present the RT results for both scope and baseline conditions, followed by analyses of offline response patterns.

4.5.1. RTs

The results are shown in Figure 2 for the baseline and scope conditions, with the ROIs shaded in gray. The statistical modeling results are presented in Table 2. The left panel presents the results for the baseline conditions, whereas the right panel presents the results for the scope conditions. In the baseline conditions, the models revealed no reliable main effects of interpretation. Although a marginal effect emerged at the Spillover 1 region, it did not reach statistical significance, indicating that the two baseline conditions did not differ in RTs.
In contrast, in the scope conditions, significant main effects of scope interpretation were observed in both spillover regions (but not in the critical region). These effects were driven by longer RTs for ISR relative to SSR. The pattern aligns with our prediction that inverse scope incurs greater processing costs than surface scope. Importantly, the absence of effects in the baseline conditions rules out the possibility that the observed scope effects were merely due to shallow processing driven by morphological form mismatches, thereby supporting a genuine scope-related semantic processing account.

4.5.2. Offline Reponses

Patterns of responses to the comprehension questions are illustrated in Figure 3 for both scope and baseline conditions. Of theoretical interest, the primary focus of the analysis was on the scope conditions. In the scope conditions, “one” was the predominant response in the SSR condition, whereas this preference was attenuated in the ISR condition, with the two response options being more evenly distributed. Together, these patterns indicate a preference for SSR with the context sentence, particularly when the continuation sentence disambiguated toward an SSR. In the baseline conditions, response patterns numerically aligned with the intended singular and plural interpretations determined by the subject NPs (same for singular and different for plural). These observations were statistically confirmed by a series of logistic mixed-effects models: with the exception of the inverse interpretation bias condition (β = −0.38, SE = 0.40, p = .338), significant main effects of interpretation were observed for both the baseline and scope conditions (ps < .001).

5. Experiment 2: Modeling with LLM

We further examined whether the observed scope effects in the human SPR experiment, both in their temporal location and magnitude during sentence processing, can be predicted by surprisal differences across scope conditions. The goal is to seek qualitative and quantitative evidence for the predictive power of surprisal, as posited by expectation-based accounts of human sentence processing.

5.1. Language Model and Surprisal Derivation

Surprisal values were computed using GPT-2 (GPT-2 small; Radford et al., 2019). This model was selected for several reasons. First, GPT-2 is an autoregressive language model that predicts each word based on prior context, arguably mirroring human sentence processing, which is incremental and integrative in manner. Second, unlike larger proprietary models (e.g., GPT-3, GPT-4), GPT-2 publicly releases both model weights and token-level probability distributions, ensuring transparency and interpretability in the modeling results. Third, prior evidence indicates that GPT-2 small provides a better fit to human RT data than several larger GPT models (e.g., Oh et al., 2022; Oh & Schuler, 2023).
GPT-2 small was trained on the WebText corpus (approximately 40 GB of text) and contains 117 million parameters (Radford et al., 2019). Although, like other neural language models, it is not explicitly trained on particular syntactic or semantic representations, prior work has shown that GPT-2 can capture a wide range of syntactic and semantic regularities (e.g., Hu et al., 2024; Lampinen, 2024; Li et al., 2025; Wang et al., 2025), including phenomena related to scope ambiguity (Fang et al., 2026b; Kamath et al., 2024).
Surprisal values were operationalized as the probability of a word given its preceding context, where the linguistic representation is incrementally constructed based on all previously processed input. This approach is motivated by the assumption that LM-based surprisal captures aspects of incremental and predictive language processing that are conceptually analogous to those involved in human comprehension. From this perspective, language models provide a computational framework for investigating linguistic competence and performance, much as psycholinguistic experiments are used to infer the representations and processes underlying human language comprehension, including the interpretation and processing of quantifier scope examined in the present study. Surprisal values were computed using the Hugging Face Transformers library (Wolf et al., 2020) and executed on a GPU cluster. Token-level surprisal was calculated as the negative log probability of each token conditioned on the preceding context. To align surprisal estimates with word-level human data, values were aggregated at the word level by summing the surprisal of all subword tokens produced by the GPT-2 tokenizer.

5.2. Converting Surprisal into Predicted RTs

To generate quantitative predictions from surprisal values, we adopted the linking function approach from the Syntactic Ambiguity Processing (SAP) Benchmark (Huang et al., 2024). This approach operationalizes the linking hypothesis (Hale, 2001; Levy, 2008; Smith & Levy, 2013), which posits a linear relationship between surprisal and RTs: each additional bit of surprisal corresponds to a fixed increase in processing time (measured in milliseconds).
Following the SAP Benchmark methodology, we trained a linking function on filler sentences to estimate the conversion factors from surprisal to RTs, and then applied this trained model to generate predictions for experimental items.
The linking function was trained on filler sentences, which consisted of garden-path constructions, agreement constructions, attachment constructions, and relative clauses, all adapted from Grodner et al. (2003) as used in the SAP Benchmark. During the experiment, these filler sentences were interleaved with the experimental quantifier scope items. Because the linking function included lagged predictors for the three preceding words (positions n1 to n3), the first three words of each filler sentence lacked complete spillover information; the model was therefore trained on all remaining filler word positions with complete predictor values.
The linking function included the following predictors, following established practices in psycholinguistic modeling (Goodkind & Bicknell, 2018; Wilcox et al., 2020): (a) word-level surprisal from GPT-2, (b) word length in characters, (c) log-transformed word frequency (obtained from log word frequency (COCA; Davies, 2009)), and (d) word position within the sentence. To account for spillover effects, we included surprisal, word length, and word frequency values for the three preceding words (positions n1, n2, and n3). Additionally, we included interaction terms between word frequency and word length at each position, as these lexical properties are known to jointly influence RTs (Kliegl et al., 2004). All continuous predictors were standardized prior to conversion model fitting.
We fit a linear mixed-effects model predicting RTs from the predictors described above, with random intercepts for participants and items, using Formula (1):
R T   ~   β 0 + β 1 · P o s i t i o n + β 2 · S u r p r i s a l + β 3 · L o g F r e q + β 4 · L e n g t h + β 5 · L o g F r e q × L e n g t h + u p a r t i c i p a n t + w i t e m + ε
Surprisal, log frequency, and word length were included for the current word and three preceding word positions (n, n − 1, n − 2, n − 3), yielding spillover effects. The model included random intercepts for participants and items. The model was fit using the lme4 package (Bates et al., 2015) in R (R Core Team, 2025). The estimated coefficients from this model constitute the linking function, mapping surprisal (and other predictors) onto predicted RTs in milliseconds.
The trained linking function was then applied to the experimental quantifier scope sentences to generate predicted RTs for each word at each position. These predictions reflect the RTs that would be expected if processing difficulty were determined by the predictors captured in the linking function, most critically, GPT-2 surprisal values. By comparing these predictions to observed human RTs, we can evaluate the predictive power of surprisal values for the experimental effects observed in the human behavioral data of Experiment 1.
The linking function was trained on 17,065,464 observations from filler sentences (e.g., garden-path constructions). The model converged successfully, and the estimated coefficients are reported in Table 3.
The surprisal coefficient for the current word position (2.16 ms/bit) falls within the range typically observed in RT studies (Goodkind & Bicknell, 2018; Smith & Levy, 2013). Notably, the largest surprisal effect was observed at the n − 1 spillover position (3.01 ms/bit), consistent with previous findings that processing difficulty often manifests most strongly on the word immediately following the source of difficulty (Rayner & Duffy, 1986; Wilcox et al., 2020). The coefficients for n − 2 and n − 3 positions were smaller but still positive, indicating diminishing spillover effects across subsequent word positions.

5.3. Quantifying Surprisal Effects on RTs

To evaluate the linking function’s predictive accuracy, we generated predicted RTs for all experimental items (quantifier scope sentences) and compared them to observed human RTs. Table 4 presents the correlation between predicted and observed RTs by condition.
Across all four conditions, the linking function achieved correlations between predicted and observed RTs ranging from r = .48 to r = 0.55, corresponding to R2 values between 0.23 and 0.30. These values indicate that the linking function accounts for approximately 23–30% of the variance in RTs for semantic quantifier scope sentences. This level of predictive power is comparable to that observed in prior work applying surprisal-based models to RT prediction (Goodkind & Bicknell, 2018; Wilcox et al., 2020).
The core question of this analysis is whether GPT-2 surprisal values correctly predict the direction of RT differences between conditions in the ROIs where scope disambiguation occurs. Following Huang et al. (2024) and Kobzeva and Kush (2024), we adopt the criterion that the linking function successfully qualitatively captures the human behavior if the predicted effect is in the same direction as the effects observed in humans, and then we examine whether it captures the magnitude of the scope effects. Figure 4 below presents the word-by-word RT profiles for both observed (top) and predicted (bottom) RTs across relevant sentence regions.
Following the same approach used to analyze the observed human data, we fitted a series of linear mixed-effects models to examine how predicted RTs varied as a function of interpretations across both experimental and baseline conditions. The modeling results are reported in Table 5. Scope effects in the predicted RTs aligned with those observed in the behavioral data in terms of directionality with higher RTs for ISR than for SSR, emerging in the two spillover regions. The difference lies in the absence of a scope effect at the critical region in human data, which is present in the model predictions.
We further inspected the magnitudes of the experimental effects. The magnitude comparison between observed and predicted condition effects across regions of interest is presented in Figure 5. We compared the effect sizes only for the two spillover regions where the experimental effects were significant and comparable between predicted RT and behavioral data. At the first spillover position, the linking function captured the observed scope effect with high accuracy: the observed inverse-surface difference was 19.1 ms [3.0–35.3] (p = .021), while the predicted difference was 17.5 ms [15.8–19.3] (p < .001), corresponding to 91.7% magnitude capture. At the second spillover, the predicted effect remained directionally correct but smaller in magnitude: 9.3 ms [8.5–10.1] compared to the observed 14.7 ms [4.1–25.3] (p = .006), capturing 63.4% of the observed effect.

6. Individual Differences Effects

This section focuses on individual differences effects of WMC and LE on RT data, and on how these factors influence surprisal effects. As such, we first examine how WMC and LE modulate English speakers’ raw real-time processing of quantifier scope sentences. We then explored the interaction between surprisal and WMC/LE in accounting for both observed and predicted RTs. All continuous predictors were standardized prior to statistical modeling. WMC yielded both span-based and score-based measures. As these indices were highly correlated, we used the span-based measure in the reported models.
To inspect the effects of WMC and LE on observed human RT data, we constructed linear mixed-effects models for WMC and LE, respectively for their interaction with scope interpretation, because interaction of WMC and LE is of theoretical interest in this study. WMC did not yield any meaningful results across three ROIs. By contrast, robust LE effects were observed in both spillover regions, but not at the critical region. Specifically, a significant interaction between LE and interpretation bias emerged at the first spillover region (β = −22.93, SE = 8.59, p = .008) and the second spillover region (β = −19.49, SE = 5.37, p < .001). As illustrated in Figure 6 (first spillover region), continuation sentences were processed more rapidly as LE increased. This facilitative effect was particularly pronounced in the inverse scope condition, suggesting that the more costly interpretation benefited more strongly from greater language experience.
Second, we examined whether individual differences in WMC and LE modulate surprisal effects on human RTs. Specifically, we modeled the interaction among scope interpretation, WMC (or LE), and surprisal at the current word as well as the three preceding words. One model included WMC as the individual difference predictor, and a parallel model replaced WMC with LE. In both cases, raw RT served as the dependent variable. Log word frequency and word length were included as additive control predictors and did not enter into interaction terms. The model specification is provided below (Formula (2)). For reasons of space and interpretability, we report only the significant results.
RT ~ (surprisal(n) + surprisal(n − 1) + surprisal(n − 2) + surprisal(n − 3)) × condition (surface vs. inverse) × ID (WMC or LE) + log frequency + word length + (1 | participant) + (1 | item)
Figure 7 presents the full coefficient profiles for models examining the influence of WMC and LE on surprisal effects, with observed RTs as the dependent variable. Significant main effects emerged for LE at the critical region (β = 143.24, p = .036) and for WMC at the first spillover region (β = 258.66, p = .007). Notably, these effects were opposite to the expected direction: higher LE and WMC were associated with longer, rather than shorter, RTs. We do not pursue further interpretation of these main effects because the interaction results reported below offer a more informative account of how individual difference factors, including WMC and LE, modulate surprisal effects. These interactions are of primary theoretical interest for addressing RQ3.
At the critical word, a significant surprisal (n − 1) × LexTALE interaction emerged (β = 165.74, p = .010), indicating that higher proficiency readers exhibited a positive surprisal effect and lower proficiency readers showed a reversed slope (see Figure 8). At the first spillover position, surprisal (n − 2) interacted significantly with both LexTALE (β = 157.62, p = .019) (see Figure 9) and WMC (β = 192.56, p = .009) (see Figure 10), suggesting that higher proficiency and greater working memory capacity were associated with increased sensitivity to surprisal (n − 2), whereas lower proficiency and lower WMC participants showed attenuated or absent surprisal effects.
More importantly, the modulation of LE on surprisal effects differed across scope interpretations, as evidenced by a significant LexTALE × condition × surprisal (n − 3) interaction at the critical region (β = −33.62, p = .020; see Figure 11). This suggested that L2 readers with higher language proficiency showed a more pronounced surprisal (n − 3) effect in the surface scope condition, indicating heightened sensitivity to contextual fit at the point of disambiguation; however, in the inverse scope condition, surprisal effects were stronger among lower proficiency readers.
Finally, we examined individual differences in surprisal effects using surprisal-predicted RTs as the dependent variable. Figure 12 presents the full coefficient profiles from the three-way interaction models testing WMC and LE effects on predicted RTs.
Across positions, the predicted RT models showed robust main effects of surprisal. At the critical word, current word surprisal, log frequency, and word length were all significant predictors. At the first spillover, surprisal (n − 1) was the dominant predictor, with additional significant effects of surprisal (n − 2) and surprisal (n − 3). At the second spillover, current-word and surprisal (n − 2) remained significant. Focusing on the two-way interactions, a significant current-word surprisal × WMC interaction was observed at the critical word (β = −2.77, p = .012) (see Figure 13): although surprisal was generally associated with longer RTs, the slope was steeper for lower WMC participants and attenuated for higher WMC participants. Crucially, no three-way interactions were significant (all ps > .29), indicating that the linking function did not differentially modulate surprisal effects across scope conditions as a function of individual differences.

7. General Discussion

7.1. Asymmetrical Processing Costs for Surface and Inverse Scope Interpretations

The first research question concerned whether inverse scope incurs greater processing costs than surface scope and whether any observed asymmetry reflects genuine scope-related semantic processing. The results provide clear evidence that it does. We observed greater processing costs for inverse scope than for surface scope. These scope effects were primarily evident in spillover regions. In self-paced reading studies, it is common for the expected experimental effects not to emerge at the critical region itself, but rather in subsequent spillover regions (Mitchell, 1984).
A good-enough processing account could, in principle, explain longer RTs under inverse scope, attributing them to morphological mismatches (e.g., singular a dog vs. plural the dogs) rather than genuine scope computation. On this view, such mismatches may trigger shallow, heuristic-based interpretations (Christianson, 2016; Ferreira, 2003). However, the baseline conditions rule this out. If morphological mismatch drove the effect, (5d) should yield longer RTs than (5c), yet no difference was observed. This absence indicates that the inverse scope cost cannot be reduced to superficial form mismatch, but instead reflects genuine scope-related semantic processing.
Notably, the present study was not designed to directly test the good-enough processing account, but to consider and rule it out as an alternative explanation. Classic evidence for good-enough processing typically involves heuristic parsing driven by factors such as word order or plausibility (Christianson, 2016). In contrast, morphological (mis)match in NP number across sentences may not be a strong enough cue to reliably trigger such processing. Whether different heuristic cues vary in their ability to induce shallow processing—and thereby affect scope interpretation—remains an open question. Importantly, good-enough processing is not inevitable; its likelihood depends on speaker-level factors such as language experience and working memory resources (Karimi & Ferreira, 2016), which are beyond the scope of the present study.
Our comprehension question data provide further support for the claim that the higher cost of inverse scope reflects semantic computation involving both interpretations rather than shallow processing. The higher rate of “one” responses in the surface scope condition (Figure 3) is not surprising, either because surface scope is less costly or because of form-based priming between the singular NP and the response option “one.” More compelling, however, is the response pattern in the inverse scope condition. Despite potential form–meaning priming between the plural subject NP in the continuation and the response option “several,” participants selected “one” responses (57.2%) significantly more often than “several” responses (42.8%). Under a shallow or good-enough account, the plural form should have biased participants toward “several.” The fact that this did not occur suggests that participants were not relying on superficial morphological cues. Moreover, the lower selection rate of “several” cannot be attributed to a speed–accuracy trade-off because “several” responses were slower than “one” responses in both conditions. Together, these findings indicate that inverse scope processing reflects genuine semantic computation rather than shallow heuristic-based interpretation.
Our data confirm that inverse scope incurs greater processing costs than surface scope. The key question concerns the source of this asymmetry. Anderson (2004) attributes it to greater structural complexity, treating scope computation as a grammatical operation. In contrast, Kurtzman and MacDonald’s (1993) constraint-based model links the cost to processing principles such as the Single Reference Principle, whereby an initial commitment to a single referent must be revised to derive inverse scope. While both accounts predict higher costs for inverse scope, the grammatical account better fits our data. If the processing-based account were correct, early divergence should appear in the context sentence. However, as shown in Figure 2, early differences were minimal and even in the opposite direction, with slightly higher RTs for surface scope. The robust divergence instead emerged at disambiguation, consistent with a structure-driven explanation. Nonetheless, this study was not designed to adjudicate between theoretical accounts of scope processing, but to provide evidence for semantic processing and individual differences through scope interpretation.

7.2. Expectation-Based Explanations of Scope Interpretation Processing

Another important contribution of this paper was to examine the extent to which surprisal values derived from GPT-2 predict both the location and the magnitude of scope effects observed in human behavioral data. This inquiry is grounded in the assumption of a linear relationship between surprisal and reading times (Hale, 2001). It is crucial to evaluate model predictions not only at a qualitative level, namely, whether the model correctly identifies the locus of processing difficulty, but also at a quantitative level, namely whether it accurately captures the size of the observed effects.
Because word predictability has been shown to influence sentence processing more generally (Fang & Liu, 2026; Staub, 2025), examining the locus of scope-related effects allows us to assess whether scope interpretation can be accounted for within the expectation-based framework. If processing difficulty is proportional to surprisal across phenomena, such an account would be preferable on parsimony grounds. Furthermore, assessing the magnitude of scope-specific effects provides a more stringent test of the explanatory power of surprisal: only if model-derived surprisal values scale with observed RT differences can we conclude that they meaningfully capture the cognitive mechanisms underlying scope processing.
We found that surprisal was reasonably successful in predicting where scope effects would arise (two spillover regions) and, to some extent, in estimating the magnitude of those effects as reflected in the human behavioral data (see Figure 4). We followed the modeling procedures outlined in Huang et al. (2024; cf. Kobzeva & Kush, 2024; Van Schijndel & Linzen, 2021) to quantitatively link surprisal values to human RTs. Importantly, however, we extend their work beyond the prediction of syntactic processing phenomena (e.g., temporary ambiguity resolution) to the domain of semantic processing in the case of scope interpretation. Notably, the predictive relationship held to a meaningful extent even though the conversion factors were estimated from filler items involving syntactic structures (e.g., garden-path sentences) and subsequently applied to scope effects due to semantic computation. This cross-domain generalization suggests that surprisal-based estimates derived from one linguistic level may retain predictive power for processing phenomena at another. Such findings are consistent with the view that prediction may constitute a unified mechanism underlying sentence processing across different levels of linguistic representation (see Huang et al., 2024). More broadly, they underscore the potential of expectation-based metrics to model human language performance in an integrated manner.
Nevertheless, we should be generally cautious about concluding that scope effects can be fully reduced to surprisal or explained solely as a consequence of expectation-based processing. We cannot rule out the possibility that good-enough processing operated alongside surprisal-driven mechanisms for scope interpretation. This possibility is suggested by the clear difference between the different and same conditions in the predicted RTs, a contrast that was absent in the observed RT data (see Figure 4). The possibility that good-enough processing operates from the context sentence onward may explain why a meaningful condition difference emerged at the critical region in the predicted RTs, whereas no such difference was observed in the human data. However, we highlight that this explanation remains speculative and post hoc. Moreover, our study is not intended to be a mechanistic interpretability investigation, hence it was not designed to adjudicate among competing accounts of LM language processing, and therefore the present findings cannot be taken as evidence for any particular underlying mechanism.
More broadly, although the present study adopts the assumption that human sentence processing and language model processing share important functional similarities, the extent to which language models capture the complexity of human language processing remains an active area of investigation. Current research continues to examine the representational structures and computational constraints underlying language model predictions and how these compare with those involved in human sentence comprehension (Hanna & Mueller, 2025; Ryu & Lewis, 2025; Timkey et al., 2026). Future work could therefore investigate scope processing using mechanistic interpretability techniques (cf. Li et al., 2026) to probe potential parallels between the temporal dynamics of human semantic processing and the internal representations of language models. Such an approach would help clarify the extent to which computational metrics—particularly those derived from intermediate model layers, rather than final-layer outputs—can better account for human sentence processing (Hu et al., 2025; Kuribayashi et al., 2025).

7.3. Individual Differences in WMC, LE, and Surprisal Effects

We hypothesized that working memory capacity, as a crucial constraint on complex sentence processing, would modulate the processing of inverse scope interpretations due to their greater structural complexity. However, our results did not support this prediction. But this absence of WMC effects does not necessarily imply that WMC plays no role in scope processing; rather, it may suggest that the relationship between WMC and language processing is more complex than a simple capacity-based account (Roberts, 2012) would assume. For instance, Just and Carpenter (1992) proposed that language processing draws on a single pool of WMC resources independent of linguistic knowledge, predicting greater difficulty for individuals with lower WMC capacity when processing structurally complex constructions like inverse scope. Our online data, however, did not support this prediction.
A later refinement of this account was advanced by Caplan and Waters (1999), who argued for the existence of multiple resource pools, each dedicated to specific processing components. According to this two-stage model, sentence processing involves an initial, automatic interpretive stage, reflected in online measures, followed by a later, post-interpretive stage, typically manifested in offline interpretation. Crucially, only the latter stage is assumed to be susceptible to WMC interference (Caplan & Waters, 1999; James et al., 2018). Consistent with this framework, logistic mixed-effects models revealed significant interactions between scope interpretation bias and WM span (β = 1.52, SE = 0.71, p = .032), suggesting that greater WMC capacity facilitated inverse scope interpretations in offline interpretation. Together with the null effects of WMC on real-time processing, these findings are more compatible with the two-stage account of WMC effects.
Importantly, WMC effects are not uniformly observed across structures. Prior work (Fang & Wang, 2025) suggests that cognitive influences may be structure-dependent, with certain scope configurations engaging domain-general mechanisms more strongly than others. Cross-study comparisons must be interpreted cautiously due to methodological differences, but future research would benefit from integrating multiple cognitive measures and examining a broader range of scope constructions using both online and offline paradigms.
Turning to LE, our results indicate that greater naturalistic exposure enhances the efficiency of inverse scope processing. This finding dovetails with cross-linguistic evidence documenting substantial variability in quantifier scope interpretation (e.g., Brasoveanu & Dotlačil, 2015; Fang et al., 2025; Philipp & Zimmermann, 2025). One proposal is that aspects of scope representation are underdetermined by the input, leading to variability across individuals (Han et al., 2016). By directly measuring language experience, we provide empirical support for this view: richer exposure appears to strengthen access to inverse scope interpretations, reflecting cumulative experience with such mappings in naturalistic usage.
Our findings regarding the modulation of surprisal effects by WMC and LE contribute to the growing body of research examining individual differences in surprisal effects (Berzak & Levy, 2023; Oralova et al., 2026; Škrjanec & Demberg, 2026). To the extent that surprisal reflects expectation-driven processing, individual differences in surprisal sensitivity should arise naturally, given that predictive processing varies systematically across individuals (Verhagen et al., 2018).
Our results indicate that both language experience and cognitive resources, as indexed by LexTALE and working memory scores respectively, modulated the extent to which surprisal shaped scope processing, with these effects more clearly observable in the raw RTs than in the surprisal-predicted RTs. With respect to language experience, our findings extend prior work on L2 learners (Berzak & Levy, 2023; Oralova et al., 2026) by demonstrating similar modulation effects among native speakers: greater language experience was associated with increased sensitivity to surprisal. This heightened sensitivity reflects the facilitation of accumulated reading experience in driving predictive and integrative processing. More broadly, these findings align with experience-based and usage-based accounts of sentence processing (Wells et al., 2009; Verhagen et al., 2018), according to which distributional patterns derived from language use give rise to probabilistic expectations that guide real-time comprehension.
WMC also modulated surprisal effects with both observed and predicted RT data: robust surprisal sensitivity was observed among higher WMC participants. This pattern contrasts with findings reported by Haller et al. (2024), who documented the opposite relationship. One possible explanation for the strengthened surprisal effects associated with higher WMC is that greater cognitive capacity enables readers to maintain richer contextual and structural representations, generate sharper probabilistic expectations, and form more precise internal probability estimates. This, in turn, facilitates more efficient integration and updating when expectations are violated, thereby amplifying the impact of surprisal on RTs.
Another notable finding from the observed RT data is that participants with less language experience and working memory capacity exhibited a reversed surprisal pattern, such that surprisal was not positively associated with RT. This suggests reduced engagement in expectation-based processing among lower WMC and lower proficiency readers, who may instead rely more on shallow or heuristic interpretation strategies. This tendency appears particularly pronounced for less complex interpretations, namely surface scope readings, as in Figure 11.

8. Limitations and Suggestions for Future Studies

We acknowledge several limitations of the present study that warrant further investigation. First, although we are confident in concluding that semantic computation was operative during sentence processing, based on the contrastive performance between the baseline condition and the scope conditions, which allowed us to rule out heuristic parsing explanations, it remains unclear whether the comprehension questions targeting interpretation may have influenced online scope processing. Previous self-paced reading studies have shown that the type of comprehension questions following a sentence can affect processing patterns (Keating & Jegerski, 2015; Leeser et al., 2011). Future research should therefore incorporate an independent measure of scope interpretation and manipulate the type of comprehension questions following each sentence to examine how question type may influence scope interpretation during sentence processing.
Second, we remain cautious about drawing firm conclusions regarding the role of WMC in light of the null effects observed. One important limitation is that WMC was assessed solely via a backward digit span task. Prior research has suggested that digit span measures may exhibit lower reliability than more complex WMC tasks, such as written sentence recall (e.g., Liu & Murao, 2025), potentially limiting their predictive power. Future research should therefore incorporate a broader battery of WMC measures, including both digit span and reading span tasks, to obtain a more robust and comprehensive estimate of WMC within a given population.
In addition, other cognitive factors such as inhibitory control have been shown to affect scope interpretation in offline tasks (Fang & Wang, 2025). Future studies could include independent measures of inhibitory control to examine how it contributes to real-time processing of quantifier scope and to disentangle its relative contribution from that of WMC. Incorporating inhibitory control measures would also allow researchers to explore whether individual differences in inhibitory control modulate surprisal effects during sentence processing (cf. Haller et al., 2024).
Another question concerns how model properties, such as training data and model size, affect scope processing. One open issue is whether the distribution of a–every constructions across surface and inverse scope interpretations in naturalistic input mirrors their representation in model training data (Fang et al., 2026b). The degree to which model exposure aligns with human linguistic experience may help explain how well surprisal estimates from a given language model predict human scope processing.

9. Conclusions

The present study demonstrates greater processing difficulty for inverse scope interpretations than for surface scope interpretations, consistent with an account grounded in semantic computation rather than purely heuristic processing. Offline interpretation data further support the engagement of abstract structural operations during scope computation. Surprisal estimates derived from language models successfully predicted both the locus and direction of scope effects in the human behavioral data, with partial convergence in effect magnitude between modeling and empirical results. Our findings also underscore the importance of individual differences in semantic processing. When surprisal was not considered, language experience, but not working memory capacity, more robustly accounted for variability in scope interpretation, particularly for inverse scope. When surprisal was included, both language experience and working memory capacity modulated surprisal effects on RTs, albeit with region-specific differences in their influence. Overall, this study advances our understanding of semantic processing through the lens of quantifier scope interpretation, highlighting the interplay among structural complexity, experiential and cognitive resources, and expectation-based mechanisms in shaping real-time scope interpretation.

Author Contributions

Conceptualization, S.F., Y.L. and Y.C.; Methodology, S.F., Y.L. and Y.C.; Software, S.F., Y.L. and Y.C.; Validation, S.F., Y.L. and Y.C.; Formal analysis, S.F. and Y.L.; Investigation, S.F., Y.L. and Y.C.; Resources, S.F., Y.L. and Y.C.; Data curation, S.F. and Y.L.; Writing—original draft preparation, S.F. and Y.L.; Writing—review and editing, S.F., Y.L. and Y.C.; Visualization, S.F. and Y.L.; Supervision, S.F.; Project administration, S.F.; Funding acquisition, S.F. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by a Ross-Lynn Postdoctoral Fellowship awarded by Purdue University to Shaohua Fang.

Institutional Review Board Statement

The study was approved by the Institutional Review Board (IRB) of Purdue University (Protocol #IRB-2025-782; approved on 19 September 2025).

Informed Consent Statement

Informed consent was obtained from all participants involved in the study.

Data Availability Statement

All data and codes used for this study are publicly available at the Open Science Framework (OSF): https://doi.org/10.17605/OSF.IO/GZ9CK.

Acknowledgments

We thank the anonymous reviewers for their insightful and constructive feedback. We gratefully acknowledge support from the CLA Ross-Lynn Postdoctoral Fellowship, funded by the Office of the Vice President for Research and Partnerships at Purdue University and awarded to Shaohua Fang. We also acknowledge additional support from the Computation and Linguistic Meaning (CALM) Lab and the Experimental Linguistics Lab (ExLing) at Purdue University. We further thank the Gilbreth Cluster at Purdue University’s Rosen Center for Advanced Computing (RCAC) for providing the GPU computational resources used in this study. Finally, we thank the audiences at the Purdue Linguistics Symposium 2026 and the members of ExLing for their valuable feedback and helpful discussions.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Altmann, G. T. (1998). Ambiguity in sentence processing. Trends in Cognitive Sciences, 2(4), 146–152. [Google Scholar] [CrossRef] [PubMed]
  2. Anderson, C. (2004). The structure and real-time comprehension of quantifier scope ambiguity [Ph.D. thesis, Northwestern University]. [Google Scholar]
  3. Anwyl-Irvine, A. L., Massonnié, J., Flitton, A., Kirkham, N., & Evershed, J. K. (2020). Gorilla in our midst: An online behavioral experiment builder. Behavior Research Methods, 52(1), 388–407. [Google Scholar] [PubMed]
  4. Aurnhammer, C., & Frank, S. L. (2019). Comparing gated and simple recurrent neural network architectures as modelsof human sentence processing. Proceedings of the Annual Meeting of the Cognitive Science Society, 41, 112–118. [Google Scholar]
  5. Barr, D. J., Levy, R., Scheepers, C., & Tily, H. J. (2013). Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of Memory and Language, 68(3), 255–278. [Google Scholar] [CrossRef] [PubMed]
  6. Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. [Google Scholar] [CrossRef]
  7. Berzak, Y., & Levy, R. (2023). Eye movement traces of linguistic knowledge in native and non-native reading. Open Mind, 7, 179–196. [Google Scholar] [CrossRef] [PubMed]
  8. Brasoveanu, A., & Dotlačil, J. (2015). Strategies for scope taking. Natural Language Semantics, 23(1), 1–19. [Google Scholar]
  9. Caplan, D., & Waters, G. S. (1999). Verbal working memory and sentence comprehension. Behavioral and Brain Sciences, 22(1), 77–94. [Google Scholar] [CrossRef] [PubMed]
  10. Christianson, K. (2016). When language comprehension goes wrong for the right reasons: Good-enough, underspecified, or shallow language processing. Quarterly Journal of Experimental Psychology, 69(5), 817–828. [Google Scholar] [CrossRef] [PubMed]
  11. Chu, C. Y., Gabriele, A., & Minai, U. (2014). Acquisition of quantifier scope interpretation by Chinese- speaking learners of English. In Selected proceedings of the 5th conference on generative approaches to language acquisition North America (pp. 157–168). Cascadilla Proceedings Project. [Google Scholar]
  12. Cong, Y., Chersoni, E., Hsu, Y.-Y., & Blache, P. (2023). Investigating the effect of discourse connectives on transformer surprisal: Language models understand connectives, even so they are surprised. In Proceedings of the 6th BlackboxNLP workshop: Analyzing and interpreting neural networks for NLP (pp. 222–232). Association for Computational Linguistics. [Google Scholar]
  13. Cowley, S. H., Pearson, L., & Barner, D. (2025). Great expectations: Print exposure predicts resolution of quantifier scope ambiguity. Journal of Experimental Psychology: Learning, Memory, and Cognition, 51, 1837–1850. [Google Scholar] [CrossRef] [PubMed]
  14. Davies, M. (2009). The 385+ million word corpus of contemporary American English (1990–2008+): Design, architecture, and linguistic insights. International Journal of Corpus Linguistics, 14(2), 159–190. [Google Scholar] [CrossRef]
  15. Demberg, V., & Keller, F. (2008). Data from eye-tracking corpora as evidence for theories of syntactic processing complexity. Cognition, 109(2), 193–210. [Google Scholar] [CrossRef] [PubMed]
  16. Dwivedi, V. D. (2013). Interpreting quantifier scope ambiguity: Evidence of heuristic first, algorithmic second processing. PLoS ONE, 8(11), e81461. [Google Scholar] [CrossRef] [PubMed]
  17. Fang, S., & Francis, E. J. (2025). Truth-value judgment tasks in second language research. Language and Linguistics Compass, 19(5), e70019. [Google Scholar] [CrossRef]
  18. Fang, S., He, J., & Liu, X. (2026a). Scope interpretation in natural and artificial language processing. Acta Psychologica, 264, 106527. [Google Scholar] [CrossRef] [PubMed]
  19. Fang, S., Li, Y., & Cong, Y. (2026b). Semantic capacity in language learners and LLMs: A Case study of quantifier scope. In Proceedings of the fifteenth language resources and evaluation conference (LREC 2026) (Vol. 11, No. 16, pp. 9602–9617). European Language Resources Association (ELRA). [Google Scholar]
  20. Fang, S., & Liu, X. (2026). Predictive processing in L2 learners: Contributions of cognitive and affective factors. Psychological Research, 90(1), 17. [Google Scholar] [CrossRef] [PubMed]
  21. Fang, S., & Wang, S. (2025). L2 interpretation of quantifier scope: Influence of individual difference factors. Language and Cognition, 17, e84. [Google Scholar] [CrossRef]
  22. Fang, S., Wu, H., & Zhao, Y. (2025). Experimental investigation on quantifier scope in Chinese relative clauses. Linguistics Vanguard, 11(1), 247–262. [Google Scholar] [CrossRef]
  23. Ferreira, F. (2003). The misinterpretation of noncanonical sentences. Cognitive Psychology, 47(2), 164–203. [Google Scholar] [CrossRef] [PubMed]
  24. Ferreira, F., & Çokal, D. (2016). Sentence processing. In Neurobiology of language (pp. 265–274). Academic Press. [Google Scholar]
  25. Fiorentino, R., Özturhan, M., Collins, A., Yang, X., Swaim, B., Carreiras, M., Mancini, S., Hoffman, L., & Gabriele, A. (2024). Extending the English LexTALE to assess vocabulary knowledge in native speakers and advanced L2 learners. Available online: https://hsp2024.github.io/abstracts/submission_159.pdf (accessed on 24 April 2026).
  26. Fox, D. (2000). Economy and semantic interpretation (Vol. 35). MIT Press. [Google Scholar]
  27. Francis, E. (2022). Gradient acceptability and linguistic theory (Vol. 11). Oxford University Press. [Google Scholar]
  28. Futrell, R., & Mahowald, K. (2025). How linguistics learned to stop worrying and love the language models. arXiv, arXiv:2501.17047. [Google Scholar]
  29. Goodkind, A., & Bicknell, K. (2018). Predictive power of word surprisal for reading times is a linear function of language model quality. In Proceedings of the 8th workshop on cognitive modeling and computational linguistics (CMCL 2018) (pp. 10–18). Association for Computational Linguistics. [Google Scholar]
  30. Grodner, D., Gibson, E., Argaman, V., & Babyonyshev, M. (2003). Against repair-based reanalysis in sentence comprehension. Journal of Psycholinguistic Research, 32(2), 141–166. [Google Scholar] [CrossRef] [PubMed]
  31. Hale, J. (2001). A probabilistic Earley parser as a psycholinguistic model. In Second meeting of the North American chapter of the association for computational linguistics. Association for Computational Linguistics. [Google Scholar]
  32. Hale, J. (2016). Information-theoretical complexity metrics. Language and Linguistics Compass, 10(9), 397–412. [Google Scholar] [CrossRef]
  33. Haller, P., Bolliger, L., & Jäger, L. (2024). Language models emulate certain cognitive profiles: An investigation of how predictability measures interact with individual differences. In Findings of the association for computational linguistics: ACL 2024 (pp. 7878–7892). Association for Computational Linguistics. [Google Scholar]
  34. Han, C. H., Musolino, J., & Lidz, J. (2016). Endogenous sources of variation in language acquisition. Proceedings of the National Academy of Sciences of the United States of America, 113(4), 942–947. [Google Scholar] [CrossRef] [PubMed]
  35. Hanna, M., & Mueller, A. (2025). Incremental sentence processing mechanisms in autoregressive transformer language models. In Proceedings of the 2025 conference of the nations of the Americas chapter of the association for computational linguistics: Human language technologies (volume 1: Long papers) (pp. 3181–3203). Association for Computational Linguistics. [Google Scholar]
  36. Hu, J., Lepori, M. A., & Franke, M. (2025). Linking forward-pass dynamics in Transformers and real-time human processing. arXiv, arXiv:2504.14107. [Google Scholar]
  37. Hu, J., Mahowald, K., Lupyan, G., Ivanova, A., & Levy, R. (2024). Language models align with human judgments on key grammatical constructions. Proceedings of the National Academy of Sciences of the United States of America, 121(36), e2400917121. [Google Scholar] [CrossRef] [PubMed]
  38. Huang, K. J., Arehalli, S., Kugemoto, M., Muxica, C., Prasad, G., Dillon, B., & Linzen, T. (2024). Large-scale benchmark yields no evidence that language model surprisal explains syntactic disambiguation difficulty. Journal of Memory and Language, 137, 104510. [Google Scholar] [CrossRef]
  39. Jaeger, T. F. (2008). Categorical data analysis: Away from ANOVAs (transformation or not) and towards logit mixed models. Journal of Memory and Language, 59(4), 434–446. [Google Scholar] [CrossRef] [PubMed]
  40. James, A. N., Fraundorf, S. H., Lee, E. K., & Watson, D. G. (2018). Individual differences in syntactic processing: Is there evidence for reader-text interactions? Journal of Memory and Language, 102, 155–181. [Google Scholar] [CrossRef] [PubMed]
  41. Just, M. A., & Carpenter, P. A. (1992). A capacity theory of comprehension: Individual differences in working memory. Psychological Review, 99(1), 122–149. [Google Scholar] [CrossRef] [PubMed]
  42. Just, M. A., Carpenter, P. A., & Woolley, J. D. (1982). Paradigms and processes in reading comprehension. Journal of Experimental Psychology: General, 111(2), 228–238. [Google Scholar] [CrossRef] [PubMed]
  43. Kamath, G., Schuster, S., Vajjala, S., & Reddy, S. (2024). Scope ambiguities in large language models. Transactions of the Association for Computational Linguistics, 12, 738–754. [Google Scholar] [CrossRef]
  44. Karimi, H., & Ferreira, F. (2016). Good-enough linguistic representations and online cognitive equilibrium in language processing. Quarterly Journal of Experimental Psychology, 69(5), 1013–1040. [Google Scholar] [CrossRef] [PubMed]
  45. Keating, G. D., & Jegerski, J. (2015). Experimental designs in sentence processing research: A methodological review and user’s guide. Studies in Second Language Acquisition, 37(1), 1–32. [Google Scholar]
  46. Kliegl, R., Grabner, E., Rolfs, M., & Engbert, R. (2004). Length, frequency, and predictability effects of words on eye movements in reading. European Journal of Cognitive Psychology, 16(1–2), 262–284. [Google Scholar] [CrossRef]
  47. Kobzeva, A., & Kush, D. (2024). Grammar and expectation in active dependency resolution: Experimental and modeling evidence from Norwegian. Cognitive Science, 48(10), e13501. [Google Scholar] [CrossRef] [PubMed]
  48. Kuribayashi, T., Oseki, Y., Taieb, S. B., Inui, K., & Baldwin, T. (2025). Large language models are human-like internally. arXiv, arXiv:2502.01615. [Google Scholar]
  49. Kurtzman, H. S., & MacDonald, M. C. (1993). Resolution of quantifier scope ambiguities. Cognition, 48(3), 243–279. [Google Scholar] [CrossRef] [PubMed]
  50. Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. (2017). lmerTest package: Tests in linear mixed effects models. Journal of Statistical Software, 82, 1–26. [Google Scholar] [CrossRef]
  51. Lampinen, A. (2024). Can language models handle recursively nested grammatical structures? A case study on comparing models and humans. Computational Linguistics, 50(4), 1441–1476. [Google Scholar] [CrossRef]
  52. Leeser, M., Brandl, A., Weissglass, C., Trofimovich, P., & McDonough, K. (2011). Task effects in second language sentence processing research. In Applying priming methods to L2 learning, teaching, and research: Insights from psycholinguistics (pp. 179–198). John Benjamins Publishing Company. [Google Scholar]
  53. Lemhöfer, K., & Broersma, M. (2012). Introducing LexTALE: A quick and valid lexical test for advanced learners of English. Behavior Research Methods, 44(2), 325–343. [Google Scholar] [PubMed]
  54. Lenth, R. (2025). emmeans: Estimated marginal means, aka least-squares means (R package Version 1.11.2-8). R Foundation for Statistical Computing. [CrossRef]
  55. Levy, R. (2008). Expectation-based syntactic comprehension. Cognition, 106(3), 1126–1177. [Google Scholar] [CrossRef] [PubMed]
  56. Li, Y., Cong, Y., & Francis, E. J. (2025). Beyond binary animacy: A multi-method investigation of LMs’ sensitivity in English object relative clauses. In Proceedings of the workshop on cognitive modeling and computational linguistics (pp. 184–196). Association for Computational Linguistics. [Google Scholar]
  57. Li, Y., Cong, Y., & Francis, E. J. (2026). Mechanistic interpretability of animacy effects on structure choice in GPT-2. In Proceedings of the 30th conference on computational natural language learning. Association for Computational Linguistics. [Google Scholar]
  58. Liu, K., & Murao, R. (2025). Reliability and validity assessment of working memory measurements. Applied Psycholinguistics, 46, e3. [Google Scholar] [CrossRef]
  59. MacDonald, M. C., & Christiansen, M. H. (2002). Reassessing working memory: Comment on Just and Carpenter (1992) and Waters and Caplan (1996). Psychological Review, 109(1), 35–54. [Google Scholar] [CrossRef] [PubMed]
  60. Massonnié, J., Mareschal, D., & Kirkham, N. Z. (2022). Individual differences in dealing with classroom noise disturbances. Mind, Brain, and Education, 16(3), 252–262. [Google Scholar] [CrossRef]
  61. May, R. (1985). Logical form: Its structure and derivation (Vol. 12). MIT Press. [Google Scholar]
  62. Mitchell, D. C. (1984). An evaluation of subject-paced reading tasks and other methods for investigating immediate processes in reading. In D. E. Kieras, & M. A. Just (Eds.), New methods in reading comprehension research (pp. 69–89). Erlbaum. [Google Scholar]
  63. Nicklin, C., & Plonsky, L. (2020). Outliers in L2 research in applied linguistics: A synthesis and data re-analysis. Annual Review of Applied Linguistics, 40, 26–55. [Google Scholar] [CrossRef]
  64. Oh, B. D., Clark, C., & Schuler, W. (2022). Comparison of structural parsers and neural language models as surprisal estimators. Frontiers in Artificial Intelligence, 5, 777963. [Google Scholar] [CrossRef] [PubMed]
  65. Oh, B. D., & Schuler, W. (2023). Why does surprisal from larger transformer-based language models provide a poorer fit to human reading times? Transactions of the Association for Computational Linguistics, 11, 336–350. [Google Scholar] [CrossRef]
  66. Oralova, G., Chen, L., Clark, S., Teodorescu, D., Helfrich, M., Fyshe, A., Epp, C. D., & Perfetti, C. (2026). Surprisal in reading: Language models predict the N400 for L2 readers. Language, Cognition and Neuroscience, 41(1), 136–155. [Google Scholar]
  67. Paterson, K. B., Filik, R., & Liversedge, S. P. (2008). Competition during the processing of quantifier scope ambiguities: Evidence from eye movements during reading. Quarterly Journal of Experimental Psychology, 61(3), 459–473. [Google Scholar] [CrossRef] [PubMed]
  68. Philipp, M., & Zimmermann, M. (2025). An experimental comparison of the availability of inverse scope in English and German. Linguistic Inquiry, 56(2), 209–245. [Google Scholar] [CrossRef]
  69. Pickering, M. J., Traxler, M. J., & Crocker, M. W. (2000). Ambiguity resolution in sentence processing: Evidence against frequency-based accounts. Journal of Memory and Language, 43(3), 447–475. [Google Scholar] [CrossRef]
  70. Polinsky, M. (2025). Multiple grammars within linguistic populations: Distributions and theoretical implications. Linguistic Approaches to Bilingualism, 16(2), 101–128. [Google Scholar] [CrossRef]
  71. Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 9. [Google Scholar]
  72. Rayner, K., & Duffy, S. A. (1986). Lexical complexity and fixation times in reading: Effects of word frequency, verb complexity, and lexical ambiguity. Memory & Cognition, 14(3), 191–201. [Google Scholar] [CrossRef] [PubMed]
  73. R Core Team. (2025). R: A language and environment for statistical computing. R Foundation for Statistical Computing. Available online: https://www.R-project.org/ (accessed on 24 April 2026).
  74. Roberts, L. (2012). Individual differences in second language sentence processing. Language Learning, 62, 172–188. [Google Scholar] [CrossRef]
  75. Ryu, S. H., & Lewis, R. L. (2025). Memory for prediction: A Transformer-based theory of sentence processing. Journal of Memory and Language, 145, 104670. [Google Scholar] [CrossRef]
  76. Schad, D. J., Vasishth, S., Hohenstein, S., & Kliegl, R. (2020). How to capitalize on a priori contrasts in linear (mixed) models: A tutorial. Journal of Memory and Language, 110, 104038. [Google Scholar] [CrossRef]
  77. Scontras, G., Polinsky, M., Tsai, C. Y. E., & Mai, K. (2017). Cross-linguistic scope ambiguity: When two systems meet. Glossa: A Journal of General Linguistics, 2(1), 36. [Google Scholar] [CrossRef]
  78. Smith, N. J., & Levy, R. (2013). The effect of word predictability on reading time is logarithmic. Cognition, 128(3), 302–319. [Google Scholar] [CrossRef] [PubMed]
  79. Staub, A. (2025). Predictability in language comprehension: Prospects and problems for surprisal. Annual Review of Linguistics, 11(1), 17–34. [Google Scholar] [CrossRef]
  80. Škrjanec, I., & Demberg, V. (2026). Language models that match reader experience are better predictors of reading times. Journal of Memory and Language, 146, 104677. [Google Scholar] [CrossRef]
  81. Timkey, W., Dillon, B., & Linzen, T. (2026). Why are language models less surprised than humans? Testing the parse multiplicity mismatch hypothesis. arXiv, arXiv:2605.15440. [Google Scholar]
  82. Trueswell, J. C., Tanenhaus, M. K., & Garnsey, S. M. (1994). Semantic influences on parsing: Use of thematic role information in syntactic ambiguity resolution. Journal of Memory and Language, 33(3), 285–318. [Google Scholar] [CrossRef]
  83. Van Schijndel, M., & Linzen, T. (2021). Single-stage prediction models do not explain the magnitude of syntactic disambiguation difficulty. Cognitive Science, 45(6), e12988. [Google Scholar] [CrossRef] [PubMed]
  84. Verhagen, V., Mos, M., Backus, A., & Schilperoord, J. (2018). Predictive language processing revealing usage-based variation. Language and Cognition, 10(2), 329–373. [Google Scholar] [CrossRef]
  85. Wang, D. P., Sadrzadeh, M., Stanojević, M., Chow, W. Y., & Breheny, R. (2025). Extracting structure from an LLM-how to improve on surprisal-based models of human language processing. In Proceedings of the 31st international conference on computational linguistics (pp. 4938–4944). Association for Computational Linguistics. [Google Scholar]
  86. Wells, J. B., Christiansen, M. H., Race, D. S., Acheson, D. J., & MacDonald, M. C. (2009). Experience and sentence processing: Statistical learning and relative clause comprehension. Cognitive Psychology, 58(2), 250–271. [Google Scholar] [CrossRef] [PubMed]
  87. Wilcox, E. G., Gauthier, J., Hu, J., Qian, P., & Levy, R. (2020). On the predictive power of neural language models for human real-time comprehension behavior. arXiv, arXiv:2006.01912. [Google Scholar]
  88. Wilcox, E. G., Pimentel, T., Meister, C., Cotterell, R., & Levy, R. P. (2023). Testing the predictions of surprisal theory in 11 languages. Transactions of the Association for Computational Linguistics, 11, 1451–1470. [Google Scholar] [CrossRef]
  89. Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., & Davison, J. (2020). Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: System demonstrations (pp. 38–45). Association for Computational Linguistics. [Google Scholar]
  90. Wu, M. J., & Ionin, T. (2022). L1-mandarin L2-English learners’ acquisition of English double-quantifier scope. In Generative SLA in the age of minimalism (pp. 93–114). John Benjamins Publishing Company. [Google Scholar]
  91. Zhou, P., & Gao, L. (2009). Scope processing in Chinese. Journal of Psycholinguistic Research, 38(1), 11–24. [Google Scholar] [PubMed]
  92. Zimmermann, M., Wartenburger, I., & Georgi, D. (2025). Variability, its limits, and the performance–competence debate: Implications of linguistic variability for a theory of grammar. Annual Review of Linguistics, 12, 345–365. [Google Scholar]
Figure 1. Overall design and study pipeline.
Figure 1. Overall design and study pipeline.
Languages 11 00126 g001
Figure 2. RT profile for the baseline and scope conditions. Error bars represent the standard error of the mean. Significance levels: † indicating marginal significance (p < .10), * p < .05, ** p < .01.
Figure 2. RT profile for the baseline and scope conditions. Error bars represent the standard error of the mean. Significance levels: † indicating marginal significance (p < .10), * p < .05, ** p < .01.
Languages 11 00126 g002
Figure 3. Comprehension question responses across interpretations.
Figure 3. Comprehension question responses across interpretations.
Languages 11 00126 g003
Figure 4. Observed (top) vs. Predicted (bottom) RTs. Error bars represent 95% CI. Significance annotated for ROIs (*** p < .001, ** p < .01, * p < .05, † p < .10, n.s. = not significant). Shaded regions indicate the ROIs.
Figure 4. Observed (top) vs. Predicted (bottom) RTs. Error bars represent 95% CI. Significance annotated for ROIs (*** p < .001, ** p < .01, * p < .05, † p < .10, n.s. = not significant). Shaded regions indicate the ROIs.
Languages 11 00126 g004
Figure 5. Observed (human) and predicted condition effects at the critical word and two spillover positions, estimated as fixed-effect β coefficients from linear mixed-effects models.
Figure 5. Observed (human) and predicted condition effects at the critical word and two spillover positions, estimated as fixed-effect β coefficients from linear mixed-effects models.
Languages 11 00126 g005
Figure 6. Interaction of LexATLE and interpretation bias at spillover region 1.
Figure 6. Interaction of LexATLE and interpretation bias at spillover region 1.
Languages 11 00126 g006
Figure 7. Forest plots of three-way interaction models (surprisal × condition × ID factor) for human RTs. Points show β estimates (ms) with 95% CIs. Red filled points indicate significant effects (p < .05); gray open points indicate non-significant effects. Significance levels: * p < .05, ** p < .01.
Figure 7. Forest plots of three-way interaction models (surprisal × condition × ID factor) for human RTs. Points show β estimates (ms) with 95% CIs. Red filled points indicate significant effects (p < .05); gray open points indicate non-significant effects. Significance levels: * p < .05, ** p < .01.
Languages 11 00126 g007
Figure 8. Two-way interaction of LexTALE × surprisal (n − 1) at the critical region.
Figure 8. Two-way interaction of LexTALE × surprisal (n − 1) at the critical region.
Languages 11 00126 g008
Figure 9. Two-way interaction of LexTALE × surprisal (n − 2) at the spillover 1 region.
Figure 9. Two-way interaction of LexTALE × surprisal (n − 2) at the spillover 1 region.
Languages 11 00126 g009
Figure 10. Two-way interaction of WMC × surprisal (n − 2) at the spillover 1 region.
Figure 10. Two-way interaction of WMC × surprisal (n − 2) at the spillover 1 region.
Languages 11 00126 g010
Figure 11. Three-way interaction of LexTALE × condition × surprisal (n − 3) at the critical region.
Figure 11. Three-way interaction of LexTALE × condition × surprisal (n − 3) at the critical region.
Languages 11 00126 g011
Figure 12. Forest plots of three-way interaction models (surprisal × condition × ID factor) for predicted RTs. Significance levels: * p < .05, ** p < .01, *** p < .001.
Figure 12. Forest plots of three-way interaction models (surprisal × condition × ID factor) for predicted RTs. Significance levels: * p < .05, ** p < .01, *** p < .001.
Languages 11 00126 g012
Figure 13. Two-way interaction of WMC × surprisal (n) at the critical region.
Figure 13. Two-way interaction of WMC × surprisal (n) at the critical region.
Languages 11 00126 g013
Table 1. Participant demographics and individual difference measures (N = 51).
Table 1. Participant demographics and individual difference measures (N = 51).
MeasureMSDRange
Age (years)34.438.0918–45
LexTALE score87.527.6070.83–99.17
Backward digit span (length)6.271.861–10
Backward digit span (total score)25.008.520–43
Note. LexTALE scores are percentage-correct values (0–100). Backward digit span is reported both as the longest correctly recalled sequence (length) and as the total number of correctly recalled trials (score). Span length was used as the primary measure of working memory in subsequent analyses.
Table 2. Linear mixed-effects modeling results for the observed human data.
Table 2. Linear mixed-effects modeling results for the observed human data.
ConditionRegionObserved
β95% CIp
ScopeCritical−3.30[−18.2, 11.6].665
Spillover + 119.11[3.0, 35.3].021
Spillover + 214.73[4.1, 25.3].006
BaselineCritical14.08[−3.9, 32.1].126
Spillover + 113.72[−0.5, 28.0].060
Spillover + 25.31[−5.7, 16.3].345
Table 3. Fixed-effect estimates from the linear mixed-effects linking function trained on filler sentences.
Table 3. Fixed-effect estimates from the linear mixed-effects linking function trained on filler sentences.
PredictorEstimate (β)SEtp
word position0.37140.51130.726
surprisal_wn2.15581.16231.855
surprisal_wn−13.01280.84083.583*
surprisal_wn−21.28170.57572.226*
surprisal_wn−30.10290.51610.199
log frequency_wn−1.96680.8733−2.252*
log frequency_wn−11.05460.87721.202
log frequency_wn−22.04240.89722.276*
log frequency_wn−31.18860.83171.429
length_wn3.86060.94254.096*
length_wn−15.51050.91416.029*
length_wn−22.41880.94822.551*
length_wn−3−0.5970.9944−0.6
log freq × length_wn0.25660.20031.281
log freq × length_wn−1−0.0360.1968−0.183
log freq × length_wn−2−0.16940.2138−0.792
log freq × length_wn−3−0.4280.2236−1.914
Note. The table reports models with unscaled variables for ease of interpretation and comparability with previous studies. Significance levels: * p < .05.
Table 4. Correlation between predicted and observed RTs by condition.
Table 4. Correlation between predicted and observed RTs by condition.
ConditionObserved RTPredicted RTNrR2
Surface370.3379.64785.518.269
Inverse368.9384.24782.482.233
Same366.6381.35391.551.304
Different378.5385.65388.500.250
Table 5. Linear mixed-effects modeling results for the surprisal-predicted data.
Table 5. Linear mixed-effects modeling results for the surprisal-predicted data.
ConditionRegionPredicted
β95% CIp
ScopeCritical6.80[5.1, 8.5]<.001
Spillover + 117.53[15.8, 19.3]<.001
Spillover + 29.34[8.5, 10.1]<.001
BaselineCritical5.89[4.4, 7.4]<.001
Spillover + 116.12[14.5, 17.8]<.001
Spillover + 27.84[7.1, 8.6]<.001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Fang, S.; Li, Y.; Cong, Y. Semantic Processing and Individual Variation: Experimental and Modeling Evidence from Quantifier Scope. Languages 2026, 11, 126. https://doi.org/10.3390/languages11060126

AMA Style

Fang S, Li Y, Cong Y. Semantic Processing and Individual Variation: Experimental and Modeling Evidence from Quantifier Scope. Languages. 2026; 11(6):126. https://doi.org/10.3390/languages11060126

Chicago/Turabian Style

Fang, Shaohua, Yue Li, and Yan Cong. 2026. "Semantic Processing and Individual Variation: Experimental and Modeling Evidence from Quantifier Scope" Languages 11, no. 6: 126. https://doi.org/10.3390/languages11060126

APA Style

Fang, S., Li, Y., & Cong, Y. (2026). Semantic Processing and Individual Variation: Experimental and Modeling Evidence from Quantifier Scope. Languages, 11(6), 126. https://doi.org/10.3390/languages11060126

Article Metrics

Back to TopTop