1. Introduction
Although artificial intelligence (AI) is now central to organizational strategies, studies show that 70% of AI projects fail to deliver their intended outcomes [
1]. Many of these failures do not stem from technological deficiencies but from the misalignment between humans and technology, particularly organizations’ inability to integrate AI into human workflows, decision-making, and culture in a way that complements human expertise [
1,
2]. Recent research shows that AI’s organizational value depends on how effectively employees trust, understand, and adapt to AI tools rather than on the complexity of the algorithms themselves [
3]. At the same time, employees often struggle with unclear expectations, limited AI literacy, and misaligned mental models, which can hinder effective AI–human teaming [
4]. Industry evidence echoes these findings, noting that managers sometimes underestimate the cultural, psychological, social, and structural challenges that shape AI adoption in the workplace, thereby creating a gap between technical capability and human readiness [
5]. As a result, the value of AI integration increasingly depends on collaboration rather than automation alone.
In response, a growing body of reviews has mapped the evolution of AI–human collaboration and highlighted the need to understand it as a socio-technical rather than purely technological phenomenon [
2]. Systematic reviews show that the effects of AI at work are far more multidimensional than assumed, initially shaping employee attitudes, team dynamics, job design, and organizational structures simultaneously [
6]. Conceptual analyses also emphasize that successful AI integration depends on the quality of collaboration between humans and intelligent systems, with scholars arguing that AI should be viewed not as a tool but as a learning partner embedded within organizational routines [
1,
4]. Recent methodological and integrative reviews further highlight that AI–human collaboration unfolds through reciprocal adaptation, where humans refine AI inputs and AI augments human cognitive and analytical capabilities [
2]. Together, these reviews suggest that effective AI–human collaboration depends on organizational structures, workflows, and training that support shared agency.
Although AI–human collaboration is increasingly discussed in organizational research, its conceptual boundaries remain blurred. In this study, we define AI–human collaboration as the interactive process through which human agents and intelligent systems jointly perform work tasks, make decisions, and shape organizational outcomes. This definition reflects AI roles within organizations. Moreover, the extant literature distinguishes between two dominant paradigms. The first conceptualizes AI as a tool, where AI augments human capabilities but remains under human control. In this model, AI supports analytics, decision-making, or workflow efficiency while humans retain primary agency [
4,
7,
8]. The second paradigm frames AI as a teammate, where intelligent systems operate as semi-autonomous collaborators participating in task execution and decision processes. Rai et al. [
9] describe digitally mediated work environments in which tasks are determined, executed, and coordinated by both human and AI agents, reflecting this hybrid structure. Similarly, emerging research argues that advanced AI increasingly functions as a teammate rather than merely a technological instrument [
10]. Accordingly, our review captures studies spanning both tool-centric and teammate-centric perspectives.
Despite its transformative promise, the academic study of AI–human collaboration remains fragmented and conceptually diffused. Although recent review studies have provided important insights into the field, it still lacks an integrated overview and reveals considerable dispersion in theories, units of analysis, and methodological approaches [
6]. Many existing syntheses provide valuable but partial insights, often constrained by disciplinary silos or by a focus on isolated constructs rather than the broader intellectual structure of the field. For instance, Liu and Shen [
11] offered a descriptive account of publication activity but did not map how research streams relate or evolve over time. Likewise, Fragiadakis et al. [
2] advanced a methodological lens for evaluating collaboration quality, yet their focus on task-level dynamics leaves open questions about the organizational, behavioral, and socio-technical processes that underpin effective human–AI teaming. Other reviews, such as those on algorithmic trust [
12], interaction design [
13], and AI adoption in the workplace [
14], address specific facets but provide only fragmented snapshots of a field spanning human–computer interaction, organizational behavior, information systems, and human resource management (HRM). Taken together, these limitations underscore the need for a comprehensive bibliometric analysis capable of uncovering the field’s intellectual foundations, mapping its conceptual clusters, and identifying emerging frontiers that current narrative reviews cannot fully capture.
To address these limitations, this study employs bibliometric analysis to systematically map and synthesize the literature on AI–human collaboration in organizations. Specifically, it synthesizes peer-reviewed research published in recent years that examines how AI and employees interact, collaborate, and shape organizational practices. Bibliometric analysis addresses fragmentation by revealing how studies across organizational behavior, psychology, management, and information systems connect to one another, creating a clearer and more integrated view of the field. Bibliometric methodologies use publication and citation data to map the structure and evolution of scientific knowledge [
15]. Importantly, bibliometric analysis is not purely quantitative. Although it uses quantitative citation networks to identify relationships among documents [
16], it also incorporates qualitative interpretation through content analysis of key clusters and documents [
17]. In this sense, bibliometric studies combine statistical rigor with interpretive depth, enabling scholars to objectively identify influential works while also explaining the theoretical and methodological narratives they represent. Unlike traditional reviews that rely on selective reading and subjective synthesis, bibliometric analysis systematically reveals the field’s structure by using complete citation networks, allowing patterns and relationships that are invisible to manual reviews to emerge.
Using this methodological framework, the study integrates document co-citation and bibliographic coupling analyses to address two related research questions. First, what is the structure of the underlying intellectual foundation of the AI–human collaboration field? This question is examined through document co-citation, which measures how frequently two cited documents appear together in reference lists of AI primary documents to reveal the theoretical foundations and intellectual communities shaping the field [
16]. Second, considering the paths, strengths, and gaps in the structure and evolution of literature, what emerging themes and research fronts shape the current development of AI–human collaboration research, and how has this niche unfolded over time? This question is explored through bibliographic coupling, which examines how primary documents (the research articles directly returned from the focal keyword search and representing the core literature on the topic) cite the same secondary documents (the references cited within those primary documents). When multiple primary documents draw on overlapping references, it indicates that they are working on related ideas or forming a shared line of inquiry. Mapping these overlaps allows us to identify current themes, emerging research fronts, and the ways the field is beginning to organize itself. Together with the co-citation analysis, this approach provides a clear picture of both the established intellectual foundations and the evolving directions of AI–human collaboration in organizational contexts.
By taking this approach, the paper offers a contribution that traditional reviews cannot achieve. As scholars increasingly call for integrative, field-level perspectives to overcome conceptual fragmentation in emerging areas of work and technology [
1,
15], this study provides a systematic and theory-informed map of how AI–human collaboration research is organized and where it is moving. This analysis clarifies the field’s intellectual foundations, emerging research directions, and how research in this area develops over time. This clarity is essential as workplaces continue to grapple with the human, relational, and organizational implications of AI integration [
18,
19]. Overall, the study offers a coherent perspective for understanding why AI–human collaboration has become a critical frontier for organizational research and practice.
We begin by providing an overview of the bibliometric methodology, followed by a detailed description of our data collection process, inclusion criteria, and analytical procedures. We then present the methods and results for the two bibliometric studies: document co-citation analysis and bibliographic coupling. For each study, we identify and visualize the most influential documents and their organization into conceptual clusters, illustrating both the intellectual structure and the emerging research frontiers of the AI–human collaboration field. Finally, in the
Section 4, we integrate insights from both analyses to interpret the field’s evolution, identify dominant themes and theoretical foundations, and propose tangible directions for future research that can advance the understanding and practice of AI–human collaboration in organizational contexts.
2. Materials and Methods
In general, bibliometric methods are systematic and quantitative approaches that map the intellectual and conceptual structure of a research field by examining patterns of citations and co-occurrence [
15,
20]. They combine performance analysis, which assesses publication and citation activity, with science mapping, which visualizes how key studies, authors, and themes are interlinked [
21]. To address our specific research questions, we leverage two complementary methods, document co-citation and bibliographic coupling.
First, document co-citation identifies the intellectual foundations of AI–human collaboration research by examining how frequently two cited works appear together in the reference lists of publications within this field [
16]. In contrast, bibliographic coupling detects the current research front in AI–human collaboration by assessing how often two papers in this domain share overlapping references [
22]. Together, these techniques provide a comprehensive view of both the historical roots and the evolving directions of AI–human collaboration in organizational literature.
2.1. Data Sources and Scope
We retrieved the bibliometric data from the Scopus database (Elsevier, Amsterdam, The Netherlands), which provides extensive multidisciplinary coverage of peer-reviewed journals across management, psychology, organizational behavior, and information systems [
23]. We selected Scopus for its high citation accuracy, rich metadata, and export functionality suitable for bibliometric visualization [
15]. We imposed no language restrictions at the database level. The dataset includes documents published in multiple languages, as indexed in Scopus published between 2019 and 2026, including journal articles, conference proceedings, books, book chapters, and review papers. Non-English publications constituted a very small proportion (less than 2%) of the dataset and did not materially affect the overall network structure. To preserve conceptual and epistemic coherence, we restricted the review to scholarly publications indexed within management, organizational behavior, business, and related social science subject categories. Bibliometric best-practice guidelines emphasize the importance of clearly delimiting disciplinary scope to ensure interpretability and analytical validity of citation networks [
24]. Similarly, Raftopoulos and Hamari [
25] demonstrate that refining search boundaries to business and management domains is necessary when the objective is to capture socio-organizational interpretations of AI rather than technical system design. In line with this methodological guidance, technical and biomedical domains (e.g., computer science, engineering, medicine, chemistry, and physics) were excluded at the subject-category level. Bibliometric network techniques such as co-citation and bibliographic coupling assume a relatively coherent intellectual community; merging organizational scholarship with algorithmic or robotics-focused research would risk conflating distinct epistemic traditions and distorting cluster structure [
24]. Under this disciplinary configuration, no patent records were retrieved. Patent databases operate under different indexing logics and primarily document technological inventions rather than theoretical or empirical scholarly contributions embedded in academic citation networks.
Established bibliometric guidance further supports this boundary: Donthu et al. [
26], Mongeon and Paul-Hus [
27], and Pranckūtė [
28] all observe that academic citation networks and patent citation networks operate under fundamentally different inclusion logics, citation conventions, and disciplinary purposes, and combining them in management and organizational science research risks conflating distinct epistemic systems. We recognize the considerable scholarly value of patent corpora for tracing the technological lineage of AI as an invention. However, such an analysis addresses a different research question, namely how AI capabilities have evolved, rather than the question motivating the present study: how organizational research has conceptualized AI–human collaboration. A patent-based bibliometric mapping of AI invention is a complementary and worthwhile, but distinct, study. Within the present design, patents fall outside both the database used and the epistemic scope of the citation network being mapped.
2.2. Search Strategy and Inclusion Criteria
Following PRISMA guidelines [
29], we conducted a systematic search in November 2025 using field-tagged keywords that reflect the interdisciplinary nature of the AI-human collaboration domain. The search string included the following exact terms: “AI-human collaboration”, “AI integration in business”, “AI-human teaming”, “AI-human partner *”, “AI integration in organizations”, “AI integration in the workplace”, “AI in the workplace”, “AI socialization”, “AI social integration”, “AI sociotechnical”, “AI socio-technical”, “AI as social actor”, “AI social interaction”, and “AI integration”.
Consistent with the study’s objective of mapping AI–human collaboration within organizational and managerial contexts, the search strategy was intentionally constructed to reflect terminology predominantly used in management and information systems scholarship rather than the broader technical AI literature. Bibliometric methodology emphasizes that keyword selection must align with the conceptual focus of the research question to preserve interpretive coherence and avoid cross-domain distortion [
15,
26]. Preliminary scoping tests that included alternative formulations such as “human–AI collaboration”, “human-in-the-loop”, “hybrid intelligence”, and “AI-assisted decision making” substantially increased retrieval volume. However, inspection showed that most additional records originated from computer science, robotics, and human–computer interaction venues centered on system architecture and algorithmic control rather than organizational implementation. Accordingly, the final query retained “AI–human collaboration” and related expressions that more consistently capture workplace integration, socio-technical interaction, and managerial adaptation. This reflects a scope delimitation aligned with the managerial focus of the present study.
The Scopus search covered publications indexed between 2000 and 2026. However, the applied keyword combination yielded no eligible records prior to 2019. Consequently, the final dataset consisted entirely of publications published between 2019 and 2026.
The PRISMA-style flow counts recorded 2226 records identified and screened, of which 2178 were retained in the final dataset. Minor discrepancies between initial retrieval and exported records may reflect database indexing updates or metadata normalization procedures common in dynamic citation databases [
30]. The resulting dataset comprised 2178 primary documents. We exported the complete bibliographic metadata, including authors, titles, abstracts, keywords, and cited references, from Scopus in CSV format for subsequent co-citation and bibliographic coupling analyses. The dataset reflects the Scopus snapshot extracted in November 2025. Subsequent executions of the same query may yield additional records due to ongoing database updates in this rapidly evolving research domain.
2.3. Bibliometric Analysis Methods
To address the study’s research questions, the bibliometric analysis was organized into two complementary studies, each focusing on a distinct aspect of the AI–human collaboration literature. Study 1 examines the intellectual foundations of the field through document co-citation analysis, while Study 2 investigates the current research front and emerging themes using bibliographic coupling. The following subsections describe the procedures, thresholds, and analytical decisions applied in each study.
In bibliometric network analysis, different metrics capture distinct dimensions of document influence and relational positioning [
15,
26]. Citation count refers to the total number of times a document has been cited in the Scopus database and reflects overall citation impact. Co-citation strength reflects how frequently two cited documents appear together in the reference lists of primary studies [
15,
16]. Bibliographic coupling strength reflects the extent to which two primary documents share common references.
In this study, we ranked documents using total link strength, which represents the cumulative strength of all links connected to a given document within the constructed network [
17]. Total link strength reflects how strongly a document is connected within the citation network rather than how frequently it is cited overall.
Only documents meeting the predefined minimum citation threshold were included in the network visualization to ensure analytical clarity and readability. Inclusion in the displayed map therefore depends on threshold criteria and network connectivity rather than citation count alone. The designation “Top 10% Most Structurally Central Documents” refers to documents ranked within the highest 10% based on total link strength values within the respective network.
2.3.1. Study 1: Document Co-Citation—Methods and Analysis
Document co-citation focuses on how primary documents cite pairs of secondary documents together, revealing semantic similarity and intellectual connectedness among sources [
16,
31]. When two works are frequently co-cited in later publications, they are assumed to share related theoretical or conceptual content and to form part of the same invisible college or scholarly community [
32,
33]. Therefore, co-citation strength indicates both the degree of conceptual relatedness and the importance of a document within the intellectual structure of the field.
We performed the analysis using VOSviewer 1.6.20 (Centre for Science and Technology Studies, Leiden University, Leiden, The Netherlands) [
34]. The dataset consisted of 2178 primary documents retrieved from the Scopus search. Prior to analysis, the bibliographic data were cleaned and normalized following standard bibliometric preprocessing procedures. Specifically, we used a thesaurus file, where necessary, to merge synonymous terms and correct variations in author names, keywords, and cited references [
15]. Full counting was applied so that each co-citation link contributed equally to overall network strength, and association-strength normalization was used to account for differences in citation frequency across documents.
We conducted co-citation analysis on 15,078 secondary documents cited by the 2178 primary documents to uncover the intellectual structure underlying AI–human collaboration research. Following established bibliometric practice, a minimum co-citation threshold of five was applied to retain references that exhibit meaningful and recurrent citation pairing. As emphasized by Ferreira [
35], co-citation analysis aims to identify the “intellectual core” of a field by examining how frequently two works are cited together across publications. Documents that are only sporadically co-cited do not contribute to stable structural patterns and may introduce network fragmentation. Similarly, Muschetto and Siegel [
36] applied citation-based thresholds to retain only influential references and to improve the interpretability of the co-citation map, noting that thresholding reduces peripheral noise and enhances cluster clarity. Bahoo et al. [
37] further argue that filtering based on citation frequency ensures that only documents with demonstrated scholarly impact contribute to structural mapping.
In line with these methodological principles, the threshold of five was selected to balance inclusiveness and structural robustness: it excludes weakly connected references while preserving the field’s conceptual backbone. This procedure resulted in 304 secondary documents meeting the citation threshold. For visualization purposes and to enhance interpretive clarity, the top 100 documents ranked by co-citation strength were displayed in the final network map. To ensure that the identified clusters were not an artifact of a single cutoff decision, we conducted sensitivity analyses using adjacent threshold levels. The dominant clusters and their relative configurations remained substantively stable, indicating that the intellectual structure identified reflects coherent citation patterns within the domain rather than threshold-driven distortion.
We performed network construction and visualization using the VOS mapping layout and VOS clustering algorithm, a bibliometric technique introduced by [
38] that positions closely related documents near one another in two-dimensional space and assigns each document to a coherent cluster based on co-citation patterns. For visualization, total link strength was selected, with a minimum link count of 1 for each link connecting nodes. Node size represented total link strength weight, and color denoted cluster membership; a threshold of 10 documents per cluster was selected to ensure a good reflection of the cluster theme. Distance reflected co-citation proximity. The labeling strategy used VOSviewer’s default relevance-based term weighting. Clusters were interpreted through qualitative examination of their most frequently co-cited documents, focusing on shared theoretical frameworks and conceptual orientations. Approximately 10 percent of documents per cluster were reviewed in detail to validate labeling accuracy and thematic coherence.
The co-citation network thus provides a visual and statistical representation of the intellectual foundations of AI–human collaboration, highlighting the seminal works, dominant paradigms, and theoretical schools that underpin current research. All parameter files, threshold settings, and processed datasets are available upon request from the authors and are provided in the
Supplementary Materials (File S1).
2.3.2. Study 2: Bibliographic Coupling—Methods and Analysis
Bibliographic coupling provides complementary insights to document co-citation by offering a current and future-oriented perspective on the field’s development. Whereas document co-citation examines how secondary documents are cited together, thus reflecting established intellectual traditions, bibliographic coupling focuses on how primary (citing) documents share overlapping references, making it more suitable for identifying emerging research themes and contemporary scholarly alignments [
17]. In this approach, two documents are considered “coupled” when they cite one or more secondary sources in common; the greater the overlap in their bibliographies, the higher their “coupling strength”. Thus, bibliographic coupling captures the present state of research activity and signals the trajectories along which the AI–human collaboration field is currently evolving.
We based the analysis on the same dataset of 2178 primary documents retrieved from Scopus. To ensure interpretive clarity and computational manageability, we included only primary documents exceeding a minimum citation threshold of 20, resulting in a total of 219 coupled documents forming the bibliographic network. The use of citation thresholds in bibliometric mapping is methodologically established, as retaining sufficiently cited documents reduces peripheral noise and enhances structural robustness [
35,
37]. In bibliographic coupling specifically, minimum document thresholds help prevent excessive network fragmentation and improve thematic coherence [
35]. To evaluate the robustness of this decision, we conducted sensitivity analyses using adjacent citation thresholds. The dominant cluster configuration, thematic composition, and relative spatial positioning of core nodes remained substantively stable across tested values, indicating that the resulting network structure does not depend materially on a single cutoff specification but reflects coherent coupling patterns within the domain. The analysis employed full counting, which assigns equal weight to each shared reference [
39], a method appropriate for preserving the complete relational structure in interdisciplinary domains. Association-strength normalization was applied to account for variance in the number of references across documents [
34], thereby preventing highly reference-dense publications from disproportionately influencing link strength.
We used VOSviewer version 1.6.20 [
34] for visualization and network construction, applying the VOS mapping layout algorithm and VOS clustering technique to group documents based on similarity in their reference patterns. Node size represented total link strength (i.e., coupling strength); node color indicated cluster membership, and spatial distance reflected the degree of relatedness between documents. Clustering was computed on the full set of threshold-eligible documents (n = 219) prior to any graphical filtering, as clustering and visualization constitute analytically distinct stages in network construction [
26,
34]. The minimum number of documents per cluster was set to 10 to ensure meaningful thematic groupings while preventing fragmentation into marginal or statistically weak clusters, a common practice to enhance structural interpretability in science mapping [
26,
35]. The minimum link strength threshold was set to 1 to retain the complete relational structure of the network, allowing identification of even weakly connected yet conceptually relevant documents. Full counting was applied in the coupling analysis to ensure that all bibliographic links contributed equally to network construction [
39]. To maintain analytical clarity and enhance interpretability, the top 100 documents based on total link strength were selected for visualization. This reduction was applied solely at the display stage and did not influence cluster derivation, which was computed on the full network of 219 documents. Limiting displayed nodes in dense bibliometric networks is standard practice to improve graphical readability without altering the underlying structural configuration [
26,
34]. This criterion ensures that the most influential and central works are represented on the map, offering an accurate depiction of the intellectual and thematic organization of the AI–human collaboration literature.
4. Discussion
As artificial intelligence has moved from experimental novelty to an embedded part of organizational life, the central challenge is no longer merely adopting AI but understanding how humans and AI work together as partners to create value, meaning, and adaptive capability [
1,
74]. Recent industry evidence suggests that future productivity gains depend on developing skill partnerships between humans and AI. In these partnerships, AI amplifies human judgment, while humans contribute contextual understanding, ethical oversight, and strategic sensemaking rather than relying on automation alone [
75]. As AI has become embedded in organizational workflows, the literature increasingly emphasizes questions that extend beyond initial implementation, focusing on how organizations design effective AI–human collaboration that is trusted, ethically grounded, and performance-enhancing. Yet, research on AI–human interaction remains dispersed across organizational behavior, information systems, psychology, education, and human–computer interaction, often advancing in parallel rather than cumulatively. Prior reviews have surfaced important constructs (e.g., trust, adoption, ethics, outcomes), but the field still lacks a unified framework that shows how foundational traditions connect to today’s emerging research fronts. Within this context, the present study first examines what intellectual foundations have shaped how AI–human collaboration has been conceptualized in organizational research. The co-citation results reveal five foundational clusters that have shaped the field’s intellectual structure: (1) psychological and social foundations of AI, (2) organizational applications of AI in higher education, (3) ethical–cognitive foundations of generative AI, (4) AI literacy and educational transformation, and (5) behavioral foundations of AI adoption. Collectively, these foundations reflect an early and necessary phase of inquiry in which AI was largely conceptualized as a technology to be accepted, governed, and legitimized in contexts of uncertainty. In this phase, the literature prioritized constructs such as perceived usefulness and ease of use, intention to adopt, trust calibration, and fairness, providing guardrails to protect human agency and organizational legitimacy as AI systems entered workplace and institutional settings [
1,
6]. By focusing on constructs such as trust calibration and perceived fairness, this foundational literature established critical guardrails that protected employee agency and organizational legitimacy during early stages of AI diffusion. In this respect, the intellectual foundations identified through co-citation analysis directly support McKinsey’s argument that humans must feel safe, respected, and empowered to work with AI before meaningful collaboration can emerge [
76].
In addition, the co-citation structure reveals a marked institutional concentration within higher education contexts, particularly in studies examining teacher readiness, student perceptions, and the integration of generative AI systems such as ChatGPT. While this pattern may appear to narrow the organizational framing, it reflects the empirical distribution of the current knowledge base rather than a conceptual restriction. Recent reviews indicate that higher education has emerged as a primary institutional arena for examining AI adoption, pedagogical redesign, governance dynamics, and AI–human interaction processes, especially following the rapid diffusion of generative AI technologies [
53,
54]. The prominence of this sector within the intellectual base therefore represents an early institutional focus of AI–human collaboration research rather than a limitation of its broader organizational relevance.
At the same time, the co-citation structure also surfaces a limitation in the field’s foundations: adoption-oriented models tend to explain preconditions for use better than they explain the ongoing dynamics of collaboration once AI becomes embedded in routine work. When AI is treated primarily as a tool to be accepted rather than a collaborative work partner, the extant literature often pays greater attention to adoption than to how humans and AI co-adapt over time [
4]. The foundational map therefore clarifies why the field needs stronger frameworks for collaboration as a continuing sociotechnical process, not just an implementation outcome.
Recognizing these limitations, the present study then turns to the second research question of how research on AI–human collaboration is evolving, and what directions the field is moving toward as AI becomes embedded in organizational practice? Examining patterns of shared references among recent publications reveals that the bibliographic coupling results depict a distinctly emerging configuration of literature. Specifically, four active streams stand out: (1) AI governance, ethics, and humanization, (2) CRM adoption, capabilities, and organizational performance, (3) anthropomorphic AI and consumer emotional response, and (4) AI conversational agents and consumer experience dynamics. The coupling results show increasing attention to research examining interaction quality, relational cues, and governance mechanisms that shape ongoing collaboration rather than initial acceptance decisions. Furthermore, the coupling map suggests that research is moving toward a view of AI–human collaboration as an organizational capability that must be developed and maintained. In this emerging paradigm, outcomes depend less on the presence of AI and more on the organization’s ability to (a) redesign work for complementarity, (b) build human skills for oversight, interpretation, and judgment, and (c) embed ethical governance into everyday workflows. This trend aligns with sociotechnical perspectives that emphasize the joint optimization of technology, work design, and human systems rather than pursuing automation in isolation. Accordingly, this shift reflects increasing attention beyond adoption-centric perspectives toward understanding AI as interwoven with human practices and organizational processes [
2,
75].
In contrast to the intellectual foundations captured through document co-citation, which emphasize readiness and acceptance, the coupling results show that scholars are increasingly conceptualizing AI systems as collaborative partners engaged in interpretation, negotiation, judgment, and shared problem-solving. This pattern aligns conceptually with skill-partnership perspectives that emphasize redistributing work to leverage the complementary strengths of humans and intelligent systems rather than optimizing isolated tasks for automation [
59].
The relationship between the foundational and emerging clusters indicates the increasing scholarly attention paid to relational and practice-based dimensions, with important implications for both theory and practice. From a theoretical standpoint, a key contribution of this study is to show how the field’s intellectual foundations relate to its emerging fronts. The combined map suggests a layered structure linking foundational adoption theories with emerging collaboration-oriented research streams. Foundational work emphasizes acceptance, ethics, readiness, and behavioral intention; recent work gives greater emphasis on relational and practice-based conditions of collaboration, including humanization, conversational interaction, organizational capability development, and value creation in real contexts. This integration bridges a central gap in prior reviews by offering a field-level structure that connects fragmented streams and clarifies where cumulative theory-building is now possible. Moreover, the observed pattern suggests that future theory may benefit from extending beyond adoption-centric explanations toward frameworks that capture co-agency and co-adaptation: how humans and AI dynamically recalibrate roles, expectations, and trust over time; how collaboration quality emerges through interaction; and how organizational context shapes collaboration trajectories. In this view, collaboration is not a one-time condition that precedes use, but a developing capability shaped by job design, governance, training, and interaction design. Future research can advance theory by clarifying the mechanisms that link (a) interaction features (e.g., conversational design, anthropomorphic cues), (b) human cognitive and affective responses (e.g., trust, reliance, perceived agency), and (c) organizational conditions (e.g., readiness, leadership, digital culture) to sustained collaborative performance [
2,
4].
From a practical standpoint, these patterns suggest that organizations may need to move beyond one-off adoption strategies and invest in sustained capability-building efforts, including role-specific reskilling, leadership development in AI sensemaking and communication, workflow redesign for human–AI interdependence, and governance mechanisms that are embedded within everyday decision processes rather than imposed as external controls. Such practices directly address the limitations of earlier adopter-focused research and align with emerging human-centered and sociotechnical perspectives on AI-enabled work.
More specifically, the thematic structure of bibliometric clusters suggests that continuous, role-specific reskilling may be more aligned with the evolving research emphasis than generic AI literacy initiatives. Rather than focusing solely on technical familiarity, such initiatives may differentiate between technical AI literacy (understanding system capabilities and limits), interpretive literacy (evaluating output plausibility and bias), and coordination literacy (knowing escalation protocols when AI outputs conflict with contextual expertise). These skills enable employees to complement AI outputs and remain substantively engaged in decision processes [
3]. Leadership capability is equally critical, as leaders function as sense makers who shape how AI is understood, framed, and integrated within organizational narratives. Organizations can support AI-oriented leadership development through scenario-based decision simulations involving AI-supported recommendations, formal accountability mapping exercises clarifying when human override is required, and structured communication protocols that articulate the boundaries of AI authority within teams. By aligning AI use with organizational values and articulating clear expectations for human–AI collaboration, leaders influence identity, trust, and psychological safety in AI-enabled workplaces [
1]. In this context, fostering a collaboration culture that supports meaning and psychological safety is not peripheral but central to sustaining effective AI–human partnerships, particularly as AI systems become more autonomous and agentic [
74].
Research within the coupling clusters increasingly focuses on how work can be structured to support AI–human collaboration. Research on AI–human complementarity shows that better outcomes emerge when tasks are deliberately organized around differences between human and AI capabilities [
77]. Such outcomes are less likely when AI is treated only as a substitute for human labor. While bibliometric mapping does not prescribe specific redesign mechanisms, the concentration of complementarity-related research indicates that systematic workflow mapping and task reallocation represent promising directions for organizational experimentation. For example, organizations may conduct structured task decomposition analyses to distinguish activities suited for AI augmentation (e.g., data synthesis, anomaly detection) from those requiring contextual human judgment (e.g., ethical evaluation, negotiation), followed by pilot implementations and iterative refinement cycles to recalibrate task allocation. In parallel, although ethical governance has been extensively discussed in foundational research streams, further inquiry is needed to clarify how fairness, accountability, and human agency can be embedded within collaborative processes in ways that preserve flexibility, trust, and adaptability rather than imposing rigid oversight structures.
Another key implication relates to the evaluation of collaboration quality. The prominence of relational and interaction-oriented themes within the bibliometric configuration suggests increasing scholarly attention to dimensions beyond traditional performance metrics derived from human–machine interaction and automation research. Emerging frameworks emphasize interaction quality, mutual adaptation, and shared decision-making as potential indicators of effectiveness. However, there is still limited consensus regarding how such constructs should be operationalized, measured longitudinally, or compared across contexts [
2]. While bibliometric mapping does not assess metric validity directly, the clustering patterns indicate that evaluation of collaboration quality represents a growing area of conceptual development. Accordingly, future research may benefit from shifting attention from whether collaboration occurs to how it is designed, evaluated, and sustained within organizations.
5. Limitations
Despite its contributions, this study is subject to several limitations that stem primarily from methodological and design choices inherent in bibliometric research and should be addressed in future work. First, limitations related to data scope and coverage may affect the representativeness of the findings. The reliance on a single database (Scopus) may underrepresent relevant scholarship indexed elsewhere. Although no language restrictions were imposed at the search stage, the resulting dataset contained only a very small number of non-English publications. This distribution likely reflects the dominance of English-language outlets in this research domain rather than an explicit exclusion criterion; however, it may still limit the visibility of regionally grounded studies. As a result, some regional or disciplinary perspectives on AI–human collaboration may not be fully captured. In addition, while the search covered publications indexed between 2000 and 2026, no eligible records were identified prior to 2019. This pattern suggests that AI–human collaboration has emerged as a distinct research stream relatively recently; however, it also means that the temporal depth of the mapped network remains limited. Future studies employing longitudinal bibliometric designs may be better positioned to assess structural evolution over extended periods.
Second, the findings are influenced by methodological parameter choices that shape bibliometric network construction. Decisions regarding citation thresholds, normalization techniques, clustering resolution, and document selection are necessary for analytical clarity, yet they introduce a degree of parameter dependence that may affect network structure, cluster boundaries, and relative document prominence. While these choices follow established bibliometric guidelines, alternative parameter settings could yield slightly different configurations of the field. Importantly, bibliometric mapping identifies patterns of intellectual proximity rather than causal relationships or definitive theoretical progression. Consequently, interpretations regarding thematic evolution or structural alignment across clusters are better understood as analytical inferences derived from citation patterns rather than direct empirical demonstrations of field-wide transformation. In addition, bibliometric techniques are subject to time-lag effects, meaning that recently published but potentially influential studies may not yet have accumulated sufficient citations to appear prominently in the analysis.
Third, although data preprocessing procedures such as thesaurus files were applied to reduce redundancy and harmonize author names and references, residual issues related to self-citation or name variation may still influence relational strength within clusters. Moreover, the designation of highly influential or “top” documents is contingent upon the selected bibliometric indicators (e.g., citation frequency or total link strength), which may privilege established works over emerging contributions. These limitations are common in large-scale citation-based analyses and should be considered when interpreting cluster relationships and link strengths.
7. Conclusions
This study provides an integrative, field-level map of AI–human collaboration research and clarifies how the domain has evolved as AI has transitioned from peripheral automation to a structurally embedded element of organizational systems. By combining document co-citation and bibliographic coupling, we illuminate both (a) the intellectual foundations that have historically shaped how scholars conceptualize AI–human interaction and (b) the emerging research fronts that are now organizing contemporary inquiry. Across these analytical layers, the evidence suggests an evolution of the field: research increasingly moves beyond adoption- and tool-centric perspectives toward a sociotechnical collaboration lens, where value depends on how AI is embedded into work systems, how trust and human agency are sustained, and how organizations cultivate capabilities for effective AI–human teaming.
The co-citation structure shows that foundational research has focused on psychological and social mechanisms, behavioral acceptance frameworks, and ethical, cognitive concerns, work that established essential guardrails for legitimacy and responsible use. However, the coupling results demonstrate that current research increasingly emphasizes interaction and practice: conversational agents, humanization, capability-building, and governance mechanisms that support sustained collaboration over time. This convergence suggests that AI–human collaboration is better understood not as a one-time adoption outcome, but as an evolving sociotechnical capability developed through work design, organizational coordination mechanisms, and embedded governance.
This study contributes theoretically by organizing the fragmented literature into a coherent structure that supports cumulative theory building and by highlighting where next-generation models must go beyond acceptance to explain co-adaptation, role negotiation, and collaboration quality over time. Practically, it offers a roadmap for leaders: moving from “deploying AI” to designing sociotechnical collaboration systems, by redesigning workflows for complementarity, investing in role-specific skills (judgment, oversight, interpretation), and embedding ethical accountability within everyday decision processes. Ultimately, organizational outcomes may depend less on technical sophistication alone and more on the quality of AI–human collaboration mechanisms.