1. Introduction
Digital technologies increasingly mediate not only what individuals attend to, but how attention is organized, stabilized, and fragmented at the collective level. Platform architectures structure visibility, relevance, and temporal coordination, thereby shaping the conditions under which public deliberation, collective inquiry, and shared understanding can emerge. Within philosophy of technology, this has generated sustained concern with attention capture, distraction, and the commodification of attentional capacities under platform capitalism [
1,
2]. Yet, despite the centrality of attention to digital life, collective attention itself remains insufficiently theorized.
Most philosophical and empirical discussions either treat collective attention as an aggregate of individual cognitive states or construe it as an achievement of norm-governed institutional practices. At the same time, empirical research documents declining attention spans, rapid topic churn, and polarization in online environments [
3,
4,
5]. What is missing is a clear account of what it means for attention to be collective in digitally mediated settings, especially where shared intentions, roles, or virtues are weak or absent.
A central limitation of existing debates is their reliance on a relatively thin conception of attention as a scarce psychological resource to be captured or depleted. While this framing has been productive for diagnosing harms, it obscures a more fundamental issue: the structural organization of relevance and persistence within socio-technical systems. High engagement does not entail stable collective orientation. Indeed, algorithmically amplified environments often exhibit intense activity alongside epistemic fragmentation and loss of shared relevance [
6,
7]. This article addresses that gap by reconceptualizing collective attention as an emergent socio-technical phenomenon grounded in inferential dynamics rather than in shared mental states or institutional norms. To clarify our contribution, it is useful to distinguish three levels of engagement with existing theories (
Table 1). First, we critique the assumption that collective attention requires shared intentions, explicit coordination, or institutional norms. Second, we adopt and extend the predictive processing framework [
8,
9] by operationalizing its concepts at the collective level. Third, our novel contribution lies in (a) reconceptualizing collective attention as an emergent property of precision-regulated inference, (b) introducing formal operational metrics (CFE and PAI) for its empirical identification, and (c) developing the concept of ‘precision hijacking’ to diagnose algorithmic distortion of collective attention.
The central claim is that collective attention can be empirically identified and analyzed through precision dynamics—patterns governing how persistence, relevance, and inferential commitment are regulated across interacting agents over time. The framework draws on predictive processing, which models cognition as the regulation of prediction error under uncertainty [
8,
9]. Within this framework, precision refers to the weighting assigned to prediction errors, determining which signals are treated as salient for belief updating and action. We define attention as the structuring of information intake to support action-guiding inference under uncertainty. This definition has three components: (a) attention is a structuring effort that organizes which environmental signals are received and processed; (b) this structuring is oriented toward improving inference about the world; and (c) it operates under uncertainty, where the reliability of signals must be continuously assessed. This definition extends but departs from both mentalist accounts [
10] and resource-based accounts [
11] by emphasizing the
inferential function of attention rather than its subjective experience or capacity constraints. Precision is thus not identical to attention, but functions as a control parameter shaping attentional allocation.
Extending this framework beyond the individual level allows collective attention to be conceptualized as a distributed inferential structure. The relevant question is no longer whether individuals attend to the same object, but whether a group exhibit aligned regulation of temporal engagement under uncertainty. We adopt a minimalist conception of collectivity: attention is collective when multiple agents are constrained by a shared inferential environment such that their attention becomes mutually responsive, even without explicit coordination. Digital platforms instantiate such environments by structuring visibility and feedback loops. Collective attention, in this view, is an emergent property of socio-technical systems rather than a psychological or intentional state alone.
While our framework draws on these cognitive science distinctions, it extends them to the collective level. The concept of attention has been extensively studied in cognitive science, where several key distinctions are well-established. Internal vs. external attention distinguishes between attending to internally generated information (e.g., memories, thoughts) versus external sensory input [
12]. Stimulus-driven (bottom–up) vs. goal-directed (top–down) attention distinguishes between attention captured by salient features of the environment versus attention guided by the observer’s goals and expectations [
13]. Selective attention refers to the ability to focus on relevant information while ignoring irrelevant distractors [
14]. Attention as resource allocation treats attention as a limited cognitive resource that must be distributed across competing tasks [
11].
Our framework builds on these distinctions while extending them to the collective level.
Table 2 maps our concepts onto these established distinctions.
This mapping clarifies that our framework does not replace cognitive theories of attention but extends them to the socio-technical level, where attention is no longer a property of individual minds but an emergent property of interacting agents within a shared inferential environment.
To operationalize collective precision dynamics, we introduce Prediction Error (PE) as a measure of temporal engagement density—the duration of collective attention sustained on a particular topic. Drawing on the well-established link between response time and cognitive effort in the cognitive science literature [
15,
16], we interpret PE as a proxy for the collective cognitive effort invested in an episode. We acknowledge that this measure conflates duration with potential difficulty; it is therefore best understood as a first-order indicator of how long a topic maintained collective attention rather than a direct measure of inferential challenge. This conservative interpretation allows us to empirically track patterns of collective engagement while recognizing the limitations of purely temporal measures.
On this basis, the paper distinguishes between two ideal-typical attention regimes. Organic attention arises when precision remains responsive to inferential demands and pragmatic goals, enabling collective orientation to stabilize even in the presence of disagreement. In such regimes, uncertainty motivates sustained inferential coordination, and disagreement structures attention productively rather than fragmenting it. Mechanistic attention arises when precision is externally redirected—most notably through engagement-optimizing algorithms—producing persistent prediction error without convergence. This phenomenon is described here as precision hijacking: attention is not merely captured, but inferential relevance is systematically distorted. This distinction reframes debates about digital distraction. The core problem is not excessive attention, but misallocated precision, whereby socio-technical systems decouple salience from pragmatic inferential success.
Philosophical theories of collective attention have often emphasized joint intention, shared practices, or attentional virtues as necessary conditions for attending together [
10,
17,
18]. While such accounts capture important features of deliberative and institutional contexts, they struggle to explain collective attention in digitally mediated environments characterized by anonymity, loose coupling, and algorithmic mediation. This article advances a non-institutional account of collective attention. Institutionalization is treated as a stabilizing factor that can increase durability, not as a constitutive requirement. Likewise, attentional virtues are not taken as explanatory primitives; they may describe stable collectives retrospectively, but they do not explain how collective attention emerges. The explanatory focus shifts instead to how socio-technical architectures structure inference, regulate persistence, and align relevance across distributed agents. This move is especially pertinent for philosophy of technology, as it foregrounds the mediating role of platforms and algorithms in shaping collective epistemic life—not merely through content control, but through the regulation of temporal coordination and relevance.
To move beyond conceptual analysis, the paper develops an empirical methodology for identifying collective attention in digital environments. Using temporal interaction data from Wikipedia talk pages, inferential episodes are reconstructed and analyzed through two system-level metrics: Collective Free Energy (CFE) and Precision Alignment Index (PAI). We interpret Prediction Error (PE)—the temporal duration of episodes—as a proxy for temporal engagement density, drawing on the established link between response time and cognitive effort in the cognitive science literature [
15,
16]. We acknowledge that this measure conflates duration with potential difficulty; it is therefore best understood as an indicator of how long a topic maintained collective attention. These measures capture, respectively, the degree of unresolved collective inference and the coherence of temporal engagement across participants. Wikipedia talk is treated as a baseline case of a minimally institutionalized digital ecology, acknowledging that Wikipedia’s governance structures—including editorial policies, reputation systems, and administrator oversight—shape attention dynamics in ways that distinguish it from fully unregulated spaces [
19,
20].
The paper makes three contributions. First, it reconceptualizes collective attention as an emergent inferential structure shaped by technological mediation. Second, it introduces formal operational metrics that render collective attention empirically tractable without reducing it to engagement volume or consensus. Third, it provides a framework for diagnosing algorithmic amplification as precision hijacking, offering new resources for critical analysis of digital platforms. The analysis shows bounded collective free energy and positive precision alignment, indicating a stable collective attention regime despite the absence of centralized enforcement or shared mental states. These findings support the central thesis: collective attention is best understood as a socio-technical phenomenon emerging from precision-weighted inferential coordination. We further hypothesize that algorithmically amplified platforms would exhibit distinct patterns—specifically, high CFE with low or chaotic PAI—though this hypothesis requires future empirical testing.
The remainder of the paper develops the theoretical framework in greater depth, presents the full methodology and results, and discusses implications for theories of attention and the governance of socio-technical systems.
2. Materials and Methods
2.1. Data Source and Case Selection
The empirical analysis draws on open access datasets from the Stanford Network Analysis Project (SNAP) repository, which provides large-scale temporal interaction networks suitable for studying socio-technical coordination [
21]. Among available datasets, the Wikipedia talk temporal network (wiki-talk-temporal.txt.gz) was selected as the primary case study.
Wikipedia talk pages are particularly well-suited for analyzing collective attention for three reasons. First, interactions are explicitly oriented toward epistemic negotiation—clarification, dispute resolution, and coordination around shared content—making inferential dynamics observable without requiring semantic interpretation. Second, Wikipedia talk operates with minimal algorithmic amplification, allowing collective attention to be studied in a setting approximating a minimally institutionalized baseline. Third, participation is loosely coupled and non-institutional: contributors are distributed, often anonymous, and not bound by formal roles.
We acknowledge that Wikipedia is not a fully unregulated space. Its governance structures—including editorial policies (e.g., Neutral Point of View, Verifiability), reputation systems, article rating mechanisms, and administrator oversight—constitute forms of institutional coordination that shape attention dynamics [
19,
20]. However, unlike algorithmically amplified platforms (e.g., Twitter, YouTube), Wikipedia’s attention dynamics are primarily shaped by these explicit governance mechanisms rather than by opaque recommendation algorithms. The term “minimally institutionalized” captures this middle ground between fully organic attention and heavy algorithmic intervention. This case selection allows us to examine collective attention dynamics under conditions of minimal algorithmic amplification, providing a foundation for future comparative research.
Each entry in the dataset consists of a triplet: (SRC) user identifier initiating a contribution, (DST) Wikipedia article identifier treated as an epistemic object, and (UNIXTS) Unix timestamp. All interactions are grouped by article identifier (DST), as articles are treated as epistemic loci around which collective attention is organized.
Data Structure. The raw dataset consists of timestamped interactions on Wikipedia talk pages. Each entry in the dataset is a triplet:
SRC: User identifier (anonymized) initiating a contribution.
DST: Wikipedia article identifier (the talk page where the contribution was made).
UNIXTS: Unix timestamp (seconds since 1 January 1970).
Table 3 shows which quantities are extracted directly from the raw data and which are derived through subsequent computation.
The raw data consists of unstructured temporal sequences of interactions, where each contribution is associated with a specific article (DST) and user (SRC). To analyze collective attention, we transform this raw data into structured inferential episodes through the following sequence:
Grouping: All interactions are grouped by article identifier (DST), as articles serve as epistemic loci around which collective attention is organized.
Sorting: Within each article, interactions are sorted chronologically by timestamp.
Segmentation: Consecutive contributions with inter-arrival times below a data-driven threshold (75th percentile) are grouped into inferential episodes.
Aggregation: Episode-level metrics (duration, contribution count) are computed, and contributor-level participation is aggregated across episodes.
This transformation preserves the temporal structure of collective engagement while making it amenable to quantitative analysis.
2.2. Data Preprocessing and Episode Segmentation
Within each article, interactions were sorted chronologically. Inter-arrival times between consecutive contributions were computed across the dataset. Rather than imposing an arbitrary time window, a data-driven threshold was inferred from the distribution of inter-arrival times. The 75th percentile of this distribution—approximately 31 h—was selected as the episode boundary. An inferential episode is defined as a maximal sequence of interactions on a given article where consecutive contributions occur within this threshold.
Inferential episodes were serialized into a structured JSON format mapping article identifiers to ordered lists of episodes, where each episode consists of [user_id, timestamp] pairs. This representation preserves temporal order and participant identity while remaining agnostic about semantic content.
2.3. Operational Definitions
Prediction Error (PE) is operationalized as a temporal measure of engagement density at the episode level:
where
is the duration of the inferential episode and
is the number of contributions within the episode. Drawing on the established link between response time and cognitive effort in the cognitive science literature [
19,
20], we interpret PE as a proxy for temporal engagement density—the collective cognitive effort sustained on a particular topic over time. Longer episode durations relative to interaction density indicate extended collective engagement. We acknowledge that this measure conflates duration with potential difficulty; it is therefore best understood as a first-order indicator of how long a topic maintained collective attention rather than a direct measure of inferential challenge. This conservative interpretation aligns with our theoretical framing of PE as a measure of temporal dynamics rather than semantic difficulty.
Precision (Π) is defined as the persistence of temporal engagement under elevated prediction error. At the contributor level, precision is operationalized as the proportion of participations in episodes whose PE exceeds the contributor’s own mean episode-level PE:
This operationalization follows from our theoretical distinction between two senses of “precision.” Conceptually, precision refers to the weighting parameter in active inference that determines the influence of prediction errors on belief updating—higher precision means the system trusts its own predictions and is less influenced by surprising signals [
8]. Operationally, we use Π as a measurable proxy for this concept. A contributor with high Π demonstrates persistent engagement with high-PE episodes, suggesting resistance to shifting attention away from temporally extended topics of collective concern.
Collective Free Energy (CFE) measures the aggregate burden of unresolved inference weighted by precision:
where
is the mean prediction error of episodes involving contributor
, and
is the number of contributors.
Precision Alignment Index (PAI) captures the structural coherence of temporal engagement across contributors as the variance of contributor-level precision values:
A central challenge in extending predictive processing concepts to collective systems is justifying the interpretation of aggregate quantities in terms of established concepts (
Table 4). Our justifications follow a consistent logic:
First, we identify the individual-level concept in cognitive science or active inference (e.g., prediction error as the mismatch between predicted and actual input).
Second, we identify the measurable proxy at the collective level that corresponds to the same functional role (e.g., episode duration as a proxy for prolonged collective engagement, which corresponds to the temporal cost of resolving prediction error).
Third, we ensure that the mathematical definition preserves the functional role of the concept while being computable from observable data (e.g., PE = duration/N captures temporal engagement density, which serves the same functional role as individual-level prediction error—indicating inferential demand).
This logic is consistent with the distributed cognition framework [
22], which holds that cognitive processes can be realized at the collective level through the coordination of multiple agents, without requiring a unified cognitive system. Our metrics capture emergent properties of the collective system that supervene on individual contributions without being reducible to them.
Our operationalization draws on established literature across multiple disciplines:
Cognitive Science: Response time as a measure of cognitive effort is well-established [
15,
16].
Predictive Processing: Prediction error, precision, and free energy are foundational concepts [
8,
9,
23].
Statistical Physics: Free energy as a measure of surprise or surprise-driven dynamics [
8].
Distributed Cognition: Cognitive processes can be realized at the collective level [
23].
Network Science: Variance as a measure of structural coherence [
24].
By grounding our operational definitions in this multidisciplinary literature, we ensure that our metrics are interpretable in terms of established concepts while being applicable to the specific context of collective attention in digital environments.
2.4. Analytical Workflow
The analytical procedure consists of the following steps:
Load and preprocess temporal interaction data;
Group interactions by epistemic object (article);
Segment interactions into inferential episodes using the 75th percentile inter-arrival time threshold;
Compute episode-level prediction error using Equation (1);
Aggregate episode participation per contributor;
Compute contributor-level precision using Equation (2);
Compute system-level CFE and PAI using Equations (3) and (4).
All analyses were implemented in Python 3.12.13. Custom scripts are available at [repository URL will be provided upon acceptance].
2.5. Data and Code Availability
The SNAP Wikipedia talk dataset is publicly available at
https://snap.stanford.edu/data/wiki-talk-temporal.html (accessed on 12 March 2026). No new data were collected. Analysis code and processed episode data are available from the corresponding author upon reasonable request. Restrictions apply only to the volume of processed episode data due to file size limitations.
2.6. Ethical Statement
This study used only publicly available, anonymized interaction data from Wikipedia. No human subjects were recruited, and no identifiable personal information was accessed. No ethical approval was required.
Parts of the article text were subjected to proof reading and rephrasing using Artificial Intelligence Software (ChatGPT-5.5) before being submitted. The reference list was managed and organized using Mendeley Reference Manager v2.125.0.
3. Results
3.1. Empirical Findings
Applying the precision-based operational framework to the Wikipedia talk temporal network yields a coherent empirical profile consistent with an organic collective attention regime, as defined in
Section 2. The system-level metrics indicate sustained inferential engagement without runaway instability, alongside structured alignment of inferential commitment across contributors.
Table 5 presents the core system-level metrics. The analysis comprised N = 379,978 articles and N = 191,372 contributors. At the article level, N = 277,529 articles had at least two contributors and were included in the article-level analyses.
These values must be interpreted relationally rather than absolutely. Taken together with the distributional analyses presented below, they indicate a collective attention regime characterized by sustained temporal engagement under uncertainty, rather than by algorithmically driven amplification or fragmentation.
3.2. Distribution of Prediction Error
Figure 1 (Prediction Error Distribution) shows a continuous, long-tailed distribution of prediction error across contributors. This pattern indicates that inferential difficulty varies widely across episodes and participants, but does not collapse into discrete modes or noise-driven spikes.
Table 6 provides descriptive statistics for prediction error at the contributor level. The wide standard deviation (9655.33) relative to the mean (10,899.53) indicates substantial heterogeneity in how different users experience episode duration, with the range extending from 0 to 74,821.75.
Such a distribution is incompatible with mechanistic attention regimes, where temporal engagement is repeatedly reset through rapid topic churn. Instead, the observed pattern suggests that unresolved uncertainty accumulates in a structured manner, reflecting sustained engagement with complex epistemic objects.
3.3. Precision as Inferential Persistence
Figure 2 (Precision Distribution) reveals a graded distribution of precision values, rather than polarization at extreme ends. Contributors exhibit varying degrees of inferential persistence, forming a stable spectrum from peripheral participants to highly persistent core contributors.
Table 7 provides descriptive statistics for precision (Π) at the contributor level. The median of 0.000 indicates that most contributors never persist in high-PE episodes, while the presence of ten users with perfect precision (Π = 1.000) suggests a subset of “core contributors” (see
Table 4). The interquartile range (Q1 = 0.000, Q3 = 0.385) indicates that 75% of contributors persist in ≤38.5% of high-PE episodes.
This structure supports two conclusions. First, collective attention does not require uniform commitment or shared virtue. Second, inferential roles are differentiated but stable, consistent with a division of epistemic labor rather than chaotic participation.
3.4. Relationship Between PE and Precision
Figure 3 (PE–Precision Scatter) provides the most direct test of the theoretical framework.
Table 8 presents the correlation analysis between PE and precision across articles.
The Pearson correlation revealed a statistically significant but very weak positive relationship (r = 0.059, p < 0.001). However, given the non-normal distribution of the data, the non-parametric Spearman correlation is more reliable and revealed a weak negative relationship (ρ = −0.172, p < 0.001). The discrepancy between these coefficients suggests that the relationship between temporal engagement density and precision is complex and likely moderated by other factors (e.g., article type, topic, contributor experience). The statistical significance of both correlations is attributable to the extremely large sample size (N = 277,529).
This pattern is partially consistent with organic attention. The weak correlation suggests that precision and temporal engagement are largely independent dimensions of collective attention, rather than being tightly coupled. This supports our framework’s distinction between collective free energy (the “surprise” generated by attention-demanding events) and precision alignment (the structural coherence of how contributors persist in response to such events).
3.5. Structure of Collective Free Energy
Figure 4 (CFE Contribution Curve) shows a gradual decay in individual contributions to collective free energy, rather than dominance by a small number of agents. This indicates that engagement with epistemic objects is distributed across the population rather than concentrated in isolated attention sinks.
Table 9 provides descriptive statistics for article-level CFE. The wide range of CFE values (0 to 1.44 billion) indicates that some articles generate substantially more collective free energy than others, likely reflecting differences in article importance, controversy, or the nature of epistemic negotiations. The interquartile range (Q1 = 30,190,528.39, Q3 = 79,547,628.59) shows that most articles fall within a bounded range.
The magnitude of CFE reflects the real temporal cost of collective inquiry. Crucially, it does not indicate disorder. Instead, bounded accumulation of CFE suggests regulated temporal persistence.
3.6. Precision Alignment and Collective Coherence
Figure 5 (Precision Alignment Boxplot) shows moderate dispersion in precision values, corresponding to the observed PAI value of 0.089. This level of alignment indicates structural coherence without homogenization.
Table 10 provides descriptive statistics for article-level PAI. The low mean PAI (0.013) and median (0.002) indicate that within most articles, contributors show relatively similar levels of precision. However, the range extends to 0.250, indicating that some articles exhibit substantial variation in how contributors persist in high-PE episodes. Contributors do not converge on identical patterns of engagement, but their temporal engagement patterns remain sufficiently aligned to support coordinated collective attention. This distinguishes organic collective attention from both consensus-driven models and fragmented engagement regimes.
These findings empirically substantiate the theoretical claim that collective attention can emerge as an emergent property of socio-technical inference, rather than as a mental, normative, or institutional achievement.
The high skew in precision (median = 0.000, with some users at 1.000) suggests that collective attention is driven by a subset of core contributors, consistent with research on Wikipedia governance [
19,
20]. The weak and complex correlation between CFE and PAI indicates that collective free energy and precision alignment are largely independent dimensions of collective attention, supporting our framework’s theoretical distinction. These findings provide the empirical foundation for the subsequent discussion of algorithmic amplification and digital governance.
4. Discussion
4.1. Reinterpreting Collective Attention
The results presented in
Section 5 support a central theoretical claim of this paper: collective attention in digital environments is best understood not as a shared mental focus or as a normatively coordinated institutional practice, but as an emergent property of precision-regulated inferential dynamics embedded in socio-technical systems. The Wikipedia talk ecology exhibits sustained collective attention without requiring shared intentions, explicit coordination, or virtue-based norms. Instead, attention stabilizes through the alignment of inferential persistence under uncertainty.
This finding challenges two dominant strands in the philosophy of attention. First, it goes beyond mentalist accounts that treat attention primarily as an individual-level phenomenon structured by internal priority relations [
10,
25]. Second, it problematizes institutional accounts that treat collective attention as dependent on normative role structures and shared attentional responsibilities [
18,
26].
The empirical profile observed here—bounded Collective Free Energy (CFE*) combined with moderate Precision Alignment (PAI*)—indicates that collective attention can emerge prior to and independently of institutionalization. Institutional norms may stabilize attention, but they are not a necessary condition for its formation.
4.2. Precision Regulation and Watzl’s Priority Structures
Watzl’s influential account of attention as the structuring of a priority space provides an important conceptual bridge between individual cognition and collective phenomena [
10]. On Watzl’s view, attention organizes mental activity by determining what is treated as more or less important for thought and action. While this account remains focused on individual agents, it leaves open the possibility that priority structures could be shaped externally.
The present results extend Watzl’s insight by showing that, in digital ecologies, priority structures are partially externalized into socio-technical architectures. Precision allocation functions as a system-level analogue of priority structuring: it determines which inferential signals exert influence over continued engagement. However, unlike Watzl’s internalist framework, precision regulation here is not governed by a unified agent or a stable mental architecture. Instead, it is distributed across contributors and mediated by platform design.
This externalization parallels the cognitive science distinction between internal and external attention [
12]: just as individuals can attend to internal thoughts or external stimuli, collectives can attend to internal (shared) representations or external (platform-mediated) signals. However, unlike individual attention, collective attention is not governed by a single priority structure but by the alignment of multiple priority structures across contributors.
Our framework also reframes the distinction between bottom–up and top–down attention. In cognitive science, bottom–up attention is driven by stimulus salience, while top–down attention is driven by goals and expectations [
13]. At the collective level, bottom–up attention corresponds to platform-driven engagement (e.g., algorithmic amplification of novel or emotionally charged content), while top–down attention corresponds to inferential need-driven engagement (e.g., sustained inquiry into unresolved epistemic issues). Organic attention regimes exhibit a balance between these, while mechanistic regimes are dominated by bottom-up, platform-driven engagement.
Correlation analysis between prediction error and refined precision (
Figure 3) reveals a complex relationship. The Pearson correlation showed a very weak positive relationship (r = 0.059,
p < 0.001), while the non-parametric Spearman correlation revealed a weak negative relationship (ρ = −0.172,
p < 0.001). This suggests that the relationship between temporal engagement density and persistence is non-linear and likely moderated by other factors, such as article type, topic, or contributor experience. The statistical significance of both correlations is attributable to the extremely large sample size (N = 277,529). The weak and complex correlation indicates that priority structures in Wikipedia talk are not simply responsive to temporal engagement density in a linear fashion. Instead, the relationship is more nuanced: while there is some coupling between engagement density and persistence, other factors play a significant role in determining how contributors allocate their attention. This nuanced pattern is precisely what distinguishes organic attention from mechanistic attention in the typology developed earlier.
4.3. Käslin’s Institutionalized Account and Its Limits
Isabel Käslin’s work on collective attention offers a contrasting perspective that emphasizes institutional coordination and attentional virtues. On her account, collective attention involves the normative coordination of relevance within shared practices, often supported by role differentiation, mutual accountability, and expectations of attentional responsibility [
18]. In later work, Käslin explicitly links collective attention to forms of acting together that presuppose shared commitments and practical identities [
26].
While this framework successfully captures paradigmatic cases of institutional attention—such as scientific collaboration, deliberative assemblies, or coordinated action—it struggles to accommodate digitally mediated collectives that lack stable norms or explicit role structures. The Wikipedia talk ecology analyzed here occupies precisely this intermediate space: it is neither fully institutionalized nor normatively unified, yet it exhibits stable collective attention.
The results therefore suggest a conceptual decoupling: institutionalization and attentional virtues may enhance the durability and normative quality of collective attention, but they are not constitutive of it. Collective attention can emerge from shared inferential constraints alone, without being grounded in virtue-based coordination. This does not refute Käslin’s account but rather situates it as a special case within a broader space of attention regimes. Institutionalized attention can now be understood as one end of a spectrum characterized by highly stabilized precision alignment, rather than as the default model of collective attention.
Our findings extend Käslin’s framework by identifying the enabling conditions for virtuous collective attention: before normative coordination can emerge, a system must first exhibit minimal attentional stability. The Wikipedia talk ecology exhibits this foundational stability—captured by bounded CFE and moderate PAI—even in the absence of the institutionalized procedures that Käslin emphasizes. Thus, Käslin’s framework provides the ideal-typical normative goal, while our computational model helps explain the foundational, pre-normative dynamics that make such coordination possible.
4.4. Algorithmic Amplification and Precision Hijacking
The contrast with algorithmically amplified platforms clarifies the distinctive features of the Wikipedia talk regime. On platforms optimized for engagement—such as Twitter/X, YouTube, or TikTok—precision is systematically redirected toward signals that maximize interaction rather than inferential resolution [
1,
2]. This process, conceptualized here as precision hijacking, decouples temporal persistence from epistemic difficulty. In such environments, rising prediction error does not motivate sustained inquiry. Instead, uncertainty becomes a resource for further amplification, producing cycles of attention without convergence. Mechanistic attention regimes are therefore characterized by high activity, high volatility, and escalating Collective Free Energy without corresponding alignment.
The empirical profile observed in Wikipedia talk—moderate CFE combined with structured PAI—stands in sharp contrast to this pattern. The weak and complex correlation between PE and precision suggests that precision remains partially tethered to temporal engagement, though the relationship is more nuanced than a simple linear coupling. The absence of engagement-driven ranking mechanisms likely plays a crucial role in preserving this coupling.
This comparison reframes familiar critiques of social media. The problem is not simply distraction or information overload, but the systematic reconfiguration of inferential priorities. Attention fragments not because individuals lack discipline or virtue, but because platforms intervene at the level of precision regulation. While our findings from Wikipedia provide a baseline characterization of organic attention dynamics, the hypothesized contrast with mechanistic regimes—where algorithmic amplification would produce high CFE with low or chaotic PAI—remains to be empirically tested. We view our current findings as establishing a baseline methodology and generating testable predictions for future comparative research with data from algorithmically amplified platforms.
4.5. Collective Attention Without Virtue: A Structural Account
One of the most significant implications of this study is that it provides an account of collective attention that does not rely on epistemic virtue as an explanatory primitive. While virtue-theoretic approaches rightly emphasize normative ideals of good attention [
27,
28], they struggle to explain how collective attention emerges in environments where such virtues are neither cultivated nor expected.
The precision-based framework developed here offers a complementary perspective. It explains attentional stability and fragmentation in terms of structural conditions rather than moral qualities. Contributors need not be virtuous, reflective, or normatively aligned for organic collective attention to arise. What matters is whether the socio-technical environment preserves the coupling between uncertainty and inferential persistence.
This shift has important methodological and ethical consequences. It allows attention regimes to be diagnosed empirically without attributing blame to users, and it reframes platform responsibility in terms of architectural design rather than content moderation alone.
4.6. Implications for Digital Governance and Platform Design
The findings have direct implications for the governance of digital platforms. If collective attention is regulated through precision dynamics, then governance interventions should target the allocation of inferential weight, not merely the removal of harmful content. Metrics such as CFE and PAI provide a way to assess whether platforms support or undermine collective sense-making. High CFE combined with low or chaotic PAI would indicate environments prone to epistemic exhaustion and polarization. Conversely, bounded CFE with structured PAI would signal environments conducive to sustained inquiry.
The empirical results from Wikipedia suggest that minimally institutionalized environments can sustain organic attention dynamics. This has implications for platform design: rather than assuming that institutionalization or algorithmic curation is necessary for collective attention, designers might consider how to preserve the conditions that allow organic attention to flourish.
Design interventions could include slowing mechanisms, transparency about ranking criteria, or affordances that preserve object-centered interaction. Importantly, such interventions need not impose normative consensus; they can aim simply to preserve inferential coupling—the complex, context-sensitive relationship between temporal engagement and persistence observed in Wikipedia talk. The precision-based framework advanced here thus offers a unifying perspective: it explains both organic and mechanistic attention regimes in terms of socio-technical regulation of inference. In doing so, it reorients philosophical inquiry from mental states and moral virtues toward the architectures that shape how we attend together in digital worlds.
4.7. Limitations and Future Directions
Several limitations of this study should be acknowledged. First, our analysis is restricted to a single platform—Wikipedia talk pages. While this provides a valuable baseline case of minimally institutionalized attention dynamics, it remains an open question whether our findings generalize to other digital environments. Future research should examine attention dynamics across platforms with varying degrees of institutionalization and algorithmic amplification. Second, our measure of PE operationalizes temporal engagement density rather than direct inferential difficulty. While we drew on literature linking response time to cognitive effort to justify this approach [
15,
16], future work incorporating content analysis of talk page discussions would provide valuable validation of the relationship between temporal engagement and genuine inferential work.
Third, the weak and complex correlation between PE and precision suggests that other factors—such as article importance, topic, or contributor experience—likely moderate this relationship. Future work should investigate these moderators to develop a more complete understanding of the dynamics of collective attention. Fourth, while our descriptive analysis reveals substantial heterogeneity in contributor persistence, future work should examine the longitudinal dynamics of precision: do contributors become more persistent over time, or is persistence a stable trait? Such analyses would further illuminate the mechanisms underlying collective attention.
Finally, our hypothesized contrast between organic and mechanistic attention regimes remains to be empirically tested. Future comparative research with data from algorithmically amplified platforms (e.g., Twitter, YouTube) is needed to validate the claim that mechanistic regimes exhibit high CFE with low or chaotic PAI.
5. Conclusions
This article set out to address a foundational gap in contemporary discussions of attention in digital environments: the tendency to assume the existence of collective attention without specifying the conditions under which it emerges, stabilizes, or fragments. By reconceptualizing attention as precision-regulated inference, the paper provides a framework capable of explaining collective attention as an emergent property of socio-technical systems rather than as a mental, normative, or institutional given.
The central theoretical contribution lies in shifting the analysis of attention away from foregrounding and mental priority structures toward the regulation of temporal engagement under uncertainty. We defined attention as the structuring of information intake to support action-guiding inference under uncertainty. This definition distinguishes our account from resource-based views that treat attention as a scarce capacity and from phenomenal views that emphasize subjective experience. On our account, the core function of attention is inferential: it determines which signals are treated as relevant for updating beliefs and guiding action. Collective attention, in turn, emerges when this structuring becomes mutually responsive across multiple agents within a shared inferential environment. On this basis, collective attention was characterized minimally, without presupposing shared intentions, virtues, or institutional coordination. This allowed for the identification of collective attention in digital ecologies where coordination is weak, distributed, or implicit.
Building on this framework, the paper introduced a precision-based typology distinguishing between organic and mechanistic regimes of collective attention. Organic attention arises when precision remains coupled to temporal engagement density, enabling sustained engagement under uncertainty. Mechanistic attention, by contrast, emerges when precision is redirected independently of inferential success, producing high activity without convergence. This typology clarifies why digital environments can exhibit intense engagement alongside epistemic fragmentation and why attention collapse is not simply a matter of distraction or cognitive overload.
Methodologically, the article translated these conceptual distinctions into operational metrics—Collective Free Energy (CFE) and Precision Alignment Index (PAI)—capable of diagnosing attention regimes using temporal interaction data. Applying this framework to Wikipedia talk interactions demonstrated that collective attention can remain stable without institutional enforcement or virtue-based coordination. The empirical analysis of Wikipedia talk pages (N = 379,978 articles, N = 191,372 contributors) revealed a collective attention regime characterized by substantial CFE (41,649,821.60), moderate PAI (0.089), and a highly skewed precision distribution (Median = 0.000, with ten users at Π = 1.000). The complex relationship between temporal engagement density and precision (Pearson r = 0.059, Spearman ρ = −0.172) suggests that these dimensions are partially coupled but moderated by other factors. The empirical results showed bounded collective free energy and structured precision alignment, consistent with an organic attention regime. Crucially, these findings undermine the assumption that institutionalization is a necessary condition for collective attention, while preserving its role in enhancing stability.
The broader philosophical implications are significant. First, the analysis reframes debates about digital attention from questions of individual responsibility and epistemic virtue to questions of architectural design and inferential regulation. Second, it situates institutional and virtue-based accounts of collective attention, such as Käslin’s, as special cases within a wider space of socio-technical attention regimes. Our findings extend Käslin’s framework by identifying the enabling conditions for virtuous collective attention: before normative coordination can emerge, a system must first exhibit minimal attentional stability—captured by bounded CFE and moderate PAI. Third, it provides a non-moralizing, empirically tractable basis for evaluating platform design and digital governance.
From a technological perspective, the framework suggests that the primary risk of algorithmically amplified platforms is not merely distraction but precision hijacking—the systematic redirection of inferential weight toward engagement-optimizing signals. While our findings from Wikipedia provide a baseline characterization of organic attention dynamics, the hypothesized contrast with mechanistic regimes—where algorithmic amplification would produce high CFE with low or chaotic PAI—remains to be empirically tested. Addressing this risk requires interventions that preserve the coupling between uncertainty and inquiry rather than imposing content-level controls alone. Metrics such as CFE and PAI offer one possible foundation for such interventions, enabling platform evaluation in terms of their effects on collective sense-making.
Several limitations remain. The present study focused on a baseline ecology with minimal algorithmic amplification. Additionally, our measure of PE operationalizes temporal engagement density rather than direct inferential difficulty, and the weak and complex correlation between PE and precision suggests that other factors—such as article importance, topic, or contributor experience—likely moderate this relationship. Future work should extend the framework to comparative analyses across platforms with different incentive structures and governance models. Further research is also needed to integrate semantic content and normative evaluation without collapsing the structural account into virtue theory.
In conclusion, this paper advances a new way of thinking about collective attention in digital societies. By treating attention as a socio-technical achievement governed by precision dynamics, it opens a path toward diagnosing, comparing, and redesigning the environments in which collective inquiry takes place. In an era marked by polarization, misinformation, and attention fragmentation, reclaiming collective attention is not merely a cognitive challenge—it is a philosophical and technological imperative.