Next Article in Journal
Reusable Cognitive Digital Twins as a Foundational Paradigm for Intelligent Digital Ecosystems
Next Article in Special Issue
Modeling the Spreading of Fake News Through the Interactions Between Human Heuristics and Recommender Systems
Previous Article in Journal
An RL-Enhanced Multi-Agent Framework for Scalable and Intelligent Business Intelligence Systems
Previous Article in Special Issue
Explainable Reciprocal Recommender System for Affiliate–Seller Matching: A Two-Stage Deep Learning Approach
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Unlikely Pairs: A Decision-Support Recommendation Pipeline for Discovering Semantically Plausible Research Collaborations

by
Jorge Galán-Mena
1,2,
Martín López-Nores
1,*,
Daniel Pulla-Sánchez
3,
Luis Fernando Guerrero-Vásquez
3 and
Juan Pablo Salgado-Guerrero
2
1
AtlanTTic Research Center for Telecommunication Technologies, University of Vigo, 36310 Vigo, Spain
2
Economics and Business Management Faculty, Pontificia Universidad Católica del Ecuador, Quito 170143, Ecuador
3
GI-IATa, UNESCO Chair on Support Technologies for Educational Inclusion, Universidad Politécnica Salesiana, Cuenca 010105, Ecuador
*
Author to whom correspondence should be addressed.
Information 2026, 17(3), 254; https://doi.org/10.3390/info17030254
Submission received: 28 January 2026 / Revised: 26 February 2026 / Accepted: 28 February 2026 / Published: 3 March 2026

Abstract

Scientific collaboration is increasingly needed to address complex research challenges, yet identifying promising partners in the absence of prior co-authorship remains difficult. We present a decision-support pipeline for discovering researchers who have not previously worked together and whose collaboration is unlikely to emerge without deliberate intervention or institutional incentives. The approach leverages document-level semantic representations to estimate proximity between publications, aggregates these similarities at the author level, and surfaces collaboration opportunities that are not evident from the co-authorship graph. To support interpretation by decision makers, a separate LLM module proposes potential joint research directions, which are subsequently annotated with multi-label fields of study. We evaluate the pipeline through an institutional case study, analyzing 7531 publications from 2009 to 2024 using retrospective, temporally shifted windows. While only a small fraction of suggested pairs materialized spontaneously in subsequent periods, the collaborations that do emerge exhibit strong semantic alignment with the computed recommendations (high cosine similarity) and substantial thematic overlap. These results indicate that semantic proximity can act as an early indicator of latent complementarity between researchers without prior ties, supporting intentional institutional mediation and complementing topology-driven approaches that predict links under passive evolution.

Graphical Abstract

1. Introduction

As scientific output grows and research problems increasingly span disciplinary and organizational boundaries, collaboration enables the integration of complementary expertise, resources and perspectives, often leading to higher-impact and more innovative outcomes [1,2]. However, identifying suitable collaborators remains a nontrivial task in large and dynamically evolving research ecosystems. A particularly challenging case arises with potential collaborators who have no prior co-authorship history. Such pairs are structurally unlikely to collaborate under passive network evolution, since most new collaborations tend to emerge from existing social and topological proximity in co-authorship networks [3,4]. Nevertheless, some of these pairs may exhibit strong latent complementarities that are not yet expressed as collaborative ties. Identifying these opportunities is especially relevant for institutions aiming to foster interdisciplinary or novel collaborations through mediation mechanisms such as targeted calls, matchmaking events or seed funding. For example, universities seeking to improve their international positioning explicitly incentivize inter-institutional and international collaboration, and therefore require tools that help identify where such efforts have the highest potential. Likewise, national governments—for example, Ecuador’s—pursue increased visibility and impact of domestic research output for strategic, economic and societal reasons, and must decide how to prioritize limited resources for collaboration support. At the supranational level, programs promoted by the European Commission, such as the Twinning program (https://international-partnerships.ec.europa.eu/funding-and-technical-assistance/technical-assistance/twinning_en (accessed on 21 January 2026), are explicitly designed to connect research actors from emerging or less consolidated systems with leading institutions. In all these cases, the challenge is not to point out which collaborations are most likely to occur anyway, but to identify semantically plausible opportunities where targeted incentives or mediation could meaningfully influence the evolution of the collaboration network.
In this work, we focus on the discovery of such unlikely pairs, defined as researchers who have not previously worked together and whose collaboration is unlikely to emerge without deliberate intervention or institutional incentives. Here, institutional incentives refer to both monetary mechanisms (e.g., seed funding schemes, salary complements, mobility grants or research stays) and merit-based incentives (e.g., tenure and promotion criteria, performance evaluation systems or institutional recognition programs) designed to actively stimulate collaboration. Our objective is not to predict which collaborations are most likely to occur under passive network evolution, but to identify semantically plausible collaboration opportunities that would otherwise remain invisible to topology-driven methods. Thus, we address practical needs faced by research managers and policy makers who are expected to actively shape collaboration patterns rather than merely observe them.
Traditional methods for analyzing scientific collaboration have relied on bibliometric indicators, co-authorship analysis and social network measures [5,6,7,8,9,10]. These approaches are effective for characterizing existing collaborations, identifying influential actors and analyzing the structural properties of research networks. However, their practical deployment often requires extensive author and affiliation disambiguation, integration of heterogeneous metadata sources and modeling of large, temporally evolving graphs. Moreover, they primarily describe what has already occurred, rather than supporting the proactive discovery of new collaboration opportunities. More recent work has explored recommendation-based approaches, including topic modeling, machine learning techniques and hybrid systems that combine semantic and network information [11,12,13,14,15,16]. While these methods can highlight shared thematic interests or suggest potential partners, many of them remain focused on recommending research topics or reinforcing existing collaboration patterns. Topic-centric approaches, for example, can identify shared research areas but do not directly yield actionable author–author recommendations among researchers without prior ties [17,18,19].
A substantial line of research sees collaborator recommendation as a link prediction problem on co-authorship graphs, estimating which new edges are most likely to appear based on historical topology and social mechanisms such as common neighbors, triadic closure or preferential attachment [3,4,14]. These methods are well suited for forecasting likely future collaborations and often achieve strong predictive performance. However, because they rely on structural cues derived from past interactions, they tend to favor well-connected authors and consolidated communities, reinforcing existing collaboration patterns. As a result, they are not aligned with the goal of discovering collaborations that are structurally unlikely yet potentially valuable from a strategic or interdisciplinary perspective.
Our premise differs from topology-driven link prediction. We assume that semantic proximity between research outputs can act as an early indicator of latent complementarity, even when no prior collaboration exists and when structural predictors offer little support. Accordingly, we do not use prior co-authorship as a feature; instead, we use it as an explicit exclusion constraint. Rather than optimizing a predictive model to estimate the probability of future links, we adopt a decision-support perspective, aiming to prioritize candidates that may benefit from intentional mediation. To this end, we propose a modular pipeline that identifies unlikely pairs by estimating semantic proximity between publications, projecting these signals onto the author level, and producing ranked candidate collaborations that are independent of graph-topological reinforcement mechanisms. To support interpretation and deliberation by human decision-makers, a separate generative module proposes potential joint research directions and organizes them through field-of-study annotations. The system is designed to complement, rather than replace, predictive approaches by expanding the space of collaboration opportunities considered by institutions.
We evaluate the proposed approach through an institutional case study at Pontificia Universidad Católica del Ecuador, analyzing 7531 publications from 2009 to 2024 using retrospective shifted windows. Instead of reporting standard link-prediction metrics, we assess spontaneous materialization in subsequent periods and examine the semantic and thematic alignment between the computed recommendations and the collaborations that eventually emerge.
The main contributions of this work are threefold:
  • First, we provide a conceptual reframing of collaborator recommendation as a decision-support problem focused on unlikely pairs rather than predictive link estimation.
  • Second, we introduce a reproducible and modular semantic pipeline that operationalizes this framing at the institutional scale.
  • Third, we present an empirical evaluation showing that when collaborations do materialize, they exhibit strong semantic and thematic alignment with prior recommendations, supporting the relevance of semantic proximity as a prioritization criterion for intentional intervention.
The paper is organized as follows. Section 2 describes the proposed decision-support pipeline, including the problem formulation, data sources and the semantic mechanisms used to identify unlikely pairs. Section 3 presents the experimental setup and empirical evaluation. Section 4 discusses the results, their interpretation and their implications for collaboration discovery and institutional mediation. Finally, Section 5 summarizes the main conclusions and outlines directions for future work.

2. Methods

Our methodological approach is structured as a recommendation pipeline that begins with parameterized queries to OpenAlex [20], a free and open online catalog of the world’s scholarly research, including papers, researchers and institutions. The queries include an institutional identifier and a time range. Next, we calculate embeddings based on the textual metadata of each document (title and abstract); estimate proximities through cosine similarity by applying a threshold τ and a top-k filter with rules to exclude trivial matches; and project the document–document neighborhood onto the author–author plane to identify the unlikely pairs. An independent recommendation module, separate from the detection process, generates collaboration proposals with a Large Language Model (LLM) and organizes the output for decision-makers and researchers, as summarized in Figure 1.
Table 1 provides a concise processing summary of the data flow from OpenAlex records to the final LLM-generated proposals and their field-of-study annotations.
The target users are research managers and institutional decision-makers who act as human mediators, rather than an automated system optimizing rankings against a single predictive metric. In this context, the LLM module is not meant to improve link prediction, but to support interpretation and communication by generating concise, discussable research directions for the recommended pairs.

2.1. Problem Formulation and Intended Use

Given an institution-specific corpus of works up to a temporal cutoff t, our task is to produce a ranked set of author pairs ( a , b ) that have no recorded co-authorship before t and exhibit high semantic proximity induced by their publications. Unlike link prediction, we do not aim to estimate the probability of future collaboration under passive network evolution; instead, we prioritize decision-relevant candidates that could be activated through intentional mediation. In many settings, link prediction systems are deployed as algorithmic rankers designed to optimize predictive performance (e.g., AUC, precision/recall) and are consumed by automated recommendation workflows with limited need for case-by-case human interpretation. Accordingly, the pipeline is designed as a reproducible semantic discovery heuristic whose output is a shortlist enriched with human-interpretable research directions.
Formally, let C a b < t denote the indicator of whether authors a and b have co-authored at least one work before the cutoff t:
C a b < t = 1 x W < t : a , b A u t h ( x ) .
Our pipeline uses C a b < t only as an exclusion constraint and relies on a semantic score induced by publications, rather than on graph-topological predictors (e.g., common neighbors, triadic closure, preferential attachment or random-walk proximity) that are typical of link prediction settings.
From an operational standpoint, the method does not require training: it does not require labeled positive/negative examples, negative sampling or loss minimization. The only tunable elements are pipeline hyperparameters such as the similarity threshold τ and the top-k neighborhood size used during filtering, which the target users may tune to change the number of recommendations to assess. As a result, the output should be interpreted as a semantic-proximity shortlist, not as calibrated probabilities of future links. This is why, in Section 3, we evaluate the approach through an observational lower bound of spontaneous materialization and through ex post semantic and thematic alignment for the subset of pairs that materialize.

2.2. Data Collection

The corpus for our experiment was built from OpenAlex, which provides large-scale metadata on scholarly works and their associated entities (e.g., authors, institutions, venues, and citations) under persistent identifiers. OpenAlex is the underlying bibliographic source, but we evaluate on an institutional slice because the intended use case is institution-level decision support rather than global collaborator prediction.
Data collection begins with the institution identifier and time range, which are stored in a dedicated database collection. Concretely, the pipeline issues parameterized queries to the OpenAlex Works endpoint using the institution identifier and publication-year interval, selecting works whose authorships include that identifier in the institution’s lineage. Responses are retrieved in paginated batches and persisted in a document store with logical separation of works, per-year citation series, authorship–affiliation–country tuples, author records, and institution records. Each object is stored with OpenAlex persistent identifiers (work/author/institution) and an execution timestamp to ensure reproducibility and to enable contiguous temporal windowing in downstream stages of the pipeline.

2.3. Construction of Unlikely Pairs

To induce thematic proximity between works and thus between authors, the procedure operates at two levels (document–document and author–author), maintaining a consistent pipeline from semantic representation to author-level projection. Each document d i was represented by a SPECTER-2 vector [21] and 2 normalization was applied as per Equation (2) to ensure angular comparability.
v ^ i = e i e i 2
We use SPECTER-2 as an off-the-shelf document encoder designed for scientific text, providing a practical and reproducible semantic signal for content-based proximity without relying on co-authorship topology. For normalized vectors, the cosine similarity between two documents is the dot product (Equation (3)).
s ( d i , d j ) = v ^ i , v ^ j
Let V ^ R n × d denote the matrix whose i-th row is the normalized embedding v ^ i for document d i . The cosine similarity matrix is then obtained as V ^ V ^ . For computational efficiency, we evaluate this product in blocks: at each step, we take two row blocks A R n a × d and B R n b × d extracted from V ^ (possibly with A = B ), and then compute the corresponding similarity submatrix via A B (Equation (4)). Iterating over block pairs yields the full matrix S.
S ( A , B ) = A B , S i j ( A , B ) = s ( d i , d j )
Using the similarity matrix, only pairs with high semantic affinity were retained by applying a threshold τ (Equation (5)). In cases of high density, the matrix was reduced to the top-k neighbors per document. Before projecting onto the author plane, trivial matches were removed—specifically, identical documents (same ID/DOI) and document–document pairs that actually correspond to the same article co-signed by the same authors. This filter prevents identities or duplicates from distorting the semantic signal.
ε τ = { ( i , j ) : S i j τ }
The cleaned document–document neighborhood was projected onto the author–author graph by calculating, for each pair ( a , b ) , a proximity weight equal to the best match between their publications (Equation (6)).
w a b = max d i D a , d j D b s ( d i , d j )
Thus, the set of unlikely pairs P τ < t is defined by excluding prior co-authorship and selecting pairs whose publications exhibit high semantic proximity, so that recommendations do not depend on topological reinforcement mechanisms (Equation (7)). This definition yields a set of candidate author pairs with no collaboration history but sufficient affinity to suggest potential future cooperation.
P τ < t = { ( a , b ) : C a b < t = 0 w a b τ } .
For decision support, we rank candidates using a topology-agnostic score based solely on semantic proximity (Equation (8)). This design intentionally prioritizes semantically plausible opportunities that are not necessarily structurally favored by the co-authorship graph.
score ( a , b ) = w a b ( a , b ) P τ < t .

2.4. Generating New Research Directions for Each Pair

For each identified unlikely pair ( a , b ) , a standardized prompt was constructed summarizing the most representative works of both authors (title and abstract). This prompt was sent to an LLM (gpt-4o-mini, temperature 0.3) configured to return only a JSON containing a list of five potential joint research titles. This module is independent from the detection process, as it does not alter the semantic weighting or pair selection criteria; its goal is to enrich the recommendation with actionable content. All generated proposals are stored in the database with pair and time-window identifiers. Subsequently, these proposals are labeled using the S2FOS classifier [22] (23 fields of study) in multi-label mode.

2.5. Visualization Metrics for Proposal–Materialization Proximity

Uniform Manifold Approximation and Projection (UMAP) is used for the visualization of proposal and materialization centroids, offering a compact two-dimensional layout of the SPECTER-2 embeddings. It is worth noting that UMAP is a non-linear manifold learning method that prioritizes local neighborhood structure and may distort global distances; therefore, we do not use Euclidean geometry in the 2D projection as a quantitative proxy of semantic distance.
Let P a b be the set of generated proposal titles for a materialized pair ( a , b ) , and let G a b be the set of works co-authored by ( a , b ) observed after materialization. We embed both sets with the same SPECTER-2 encoder, denoting the (optionally 2 -normalized) embedding of an item z by v ^ ( z ) R d . We then define centroid embeddings as per Equation (9):
x ¯ a b = 1 | P a b | p P a b v ^ ( p ) , g ¯ a b = 1 | G a b | g G a b v ^ ( g ) .
All quantitative summaries in the visualization are computed in the original SPECTER-2 space using cosine distance (Equation (10)):
d cos ( a , b ) = 1 cos x ¯ a b , g ¯ a b .
We report a threshold-based proximity rate as Recall@t as per Equation (11):
Recall @ t = 1 n t ( a , b ) : M a b t + 1 = 1 1 d cos ( a , b ) t .
For visualization, we fit UMAP on the union of proposal and materialized embeddings for each window transition, using fixed hyperparameters ( n _ n e i g h b o r s = 10 , m i n _ d i s t = 0.3 , r a n d o m _ s t a t e = 7 ) to ensure layout reproducibility. The resulting 2D coordinates are used only to position centroids and to define hexagonal bins; the hexbin colors and ring radii are driven by d cos ( a , b ) from Equation (10), not by 2D Euclidean distances.

3. Results

To test our system, we analyzed the case of Pontificia Universidad Católica del Ecuador (PUCE). A total of 7531 publication records were retrieved from the period 2009–2024, maintaining persistent identifiers for works and authors to ensure traceability. The explicit collaboration network was characterized in four non-overlapping 4-year windows: 2009–2012, 2013–2016, 2017–2020 and 2021–2024, as illustrated in Figure 2. We adopt non-overlapping 4-year windows to preserve comparability across periods and to allow sufficient temporal span for collaboration dynamics to evolve beyond short-term fluctuations. This horizon is broadly consistent with typical multi-year governance and planning cycles in academic institutions, providing an interpretable frame for observing structural changes in collaboration patterns.
  • In the first window (2009–2012), the network contains 920 authors: 362 internal (PUCE, 39.35%) and 558 external (60.65%), with 309 scientific documents. Structurally, the average degree is 12.83, the density 0.014, the diameter 11 and the average path length 3.731.
  • In the last window (2021–2024), there has been a notable expansion: 13,495 authors in total, of which 4426 are internal (32.8%) and 9069 external (67.2%), with 3636 scientific documents. The structure is larger and more connected, with an average degree of 264, a density of 0.020 and a diameter of 27.
  • Between these two windows, the network shows a 14.67-fold increase in the number of authors and an 11.77-fold increase in total scientific documents published; local connectivity also increased 19.57 times compared to the initial network’s average degree. At the same time, the average path length increased to 1.57 times its initial value and the network diameter increased by a factor of 2.45—a pattern consistent with the incorporation of external communities and the expansion of the collaborative perimeter.
Figure 3 shows the network of potential links identified by our approach, setting the threshold τ = 0.90 (a sensitivity analysis over multiple τ values, and its implications for interpretation and visualization, is reported later in Section 4). For each window w t { 2009 - - 2012 , 2013 - - 2016 , 2017 - - 2020 } , red edges connect author pairs with no prior co-authorship whose document–document neighborhood contains at least one pair with cosine similarity s ( d i , d j ) τ and applying trivial-match filtering before projection. Under this criterion, 1555 potential links were obtained in 2009–2012, 6906 in 2013–2016 and 268,296 in 2017–2020.
For each identified unlikely pair, five joint work proposals are generated from a standardized prompt summarizing, per author, representative titles and abstracts from the “Potential” sections (see Figure 4). The prompt instructs the model to return a JSON object with the key items and a TI field per proposal, ensuring that the titles differ from the inputs. The example in Figure 4 shows the prompt content (left), the model’s JSON response (upper right) and the subsequent thematic annotation with S2FOS (lower right).
To support interpretability of the end-to-end flow (from semantic detection to proposal articulation and ex post verification), Table 2 presents two complete instances for new unlikely pairs detected in W t that later materialized in W t + 1 . In both cases, the proposals are generated using only the evidence available in W t (i.e., representative potential works), while the joint publication is verified in the subsequent window W t + 1 , making explicit the temporal separation between proposal generation and later materialization. For each case, we report the representative potential works used to instantiate the standardized prompt, the five proposal titles produced under a constrained JSON schema together with their S2FOS labels, and the collaboration observed in the next window with its S2FOS annotation.
To validate the proposed method, a retrospective validation with temporally shifted windows was applied. In each window W t (2009–2012, 2013–2016, 2017–2020), unlikely pairs were detected by using only the information from W t . The materialization of each pair ( a , b ) was verified in the next window W t + 1 using the indicator M a b t + 1 (Equation (12)) with a four-year shift: 2009–2012 → 2013–2016, 2013–2016 → 2017–2020 and 2017–2020 → 2021–2024.
M a b t + 1 = 1 x W t + 1 : a , b A u t h ( x ) .
Because the pipeline does not output calibrated link probabilities, we do not report standard link-prediction metrics such as the area under the receiver operating characteristic curve (AUC–ROC); instead, we report spontaneous materialization across shifted windows and assess ex post semantic and thematic alignment. Specifically, we introduce the Spontaneous Materialization Rate (SMR) as an observational lower bound, defined as the fraction of suggested unlikely pairs that materialize in W t + 1 in the absence of any matchmaking program or institutional intervention. Formally, for each window transition W t W t + 1 , we compute:
SMR t = n t | P τ < t |
where n t is the number of pairs ( a , b ) P τ < t such that M a b t + 1 = 1 , and | P τ < t | is the number of candidate unlikely pairs detected in W t under the similarity threshold τ .
The SMR—interpreted as an observational lower bound rather than a performance metric—is 0.514 % (8 out of 1555) for 2009–2012 → 2013–2016, 0.362 % (25 out of 6906) for 2013–2016 → 2017–2020 and 0.060 % (162 out of 268,296) for 2017–2020 → 2021–2024. Low values are expected given the task definition: the method targets pairs with no prior co-authorship, so without explicit mediation, only a small fraction should materialize in the subsequent period. For the subset that materialized, four properties are summarized:
1.
The semantic alignment between proposals generated in W t and confirmed works in W t + 1 , calculated as the mean cosine similarity between centroid embeddings (Equation (14)), where x ¯ a b is the centroid of the proposal-title embeddings for pair ( a , b ) and g ¯ a b is the centroid of the embeddings of the works observed after materialization (Equation (9)).
σ ¯ t = 1 n t ( a , b ) : M a b t + 1 = 1 cos x ¯ a b , g ¯ a b .
2.
The thematic overlap by S2FOS as a multi-label accuracy rate (Equation (15)).
p t F O S = 1 n t ( a , b ) : M a b t + 1 = 1 1 L a b p r o p L a b c o l l .
3.
The density of proposals per realized collaboration (Equation (16)).
D ¯ t = 1 n t ( a , b ) : M a b t + 1 = 1 K a b R a b
4.
A threshold-based proximity rate, reported as Recall@t (Equation (11)), computed over cosine distances d cos ( a , b ) in the original SPECTER-2 space.
As shown in Table 3, the absolute number of materialized pairs increases across the shifted windows (8 → 25 → 162), while the semantic alignment σ ¯ t (Equation (14)) remains high (0.9146, 0.9275 and 0.9159, respectively). The thematic overlap p t F O S (Equation (15)) is 100.00%, 100.00% and 77.16%. The proposal density D ¯ t (Equation (16)) is 2.98, 4.54 and 4.10, in the same order. Overall, these results are consistent with a recommender that prioritizes decision-oriented semantic proximity while maintaining thematic coherence among the collaborations that eventually materialize.
To visualize the proximity between proposal titles in W t and the materialized documents in W t + 1 , we fit UMAP to obtain a two-dimensional layout of the underlying SPECTER-2 embeddings. We stress that UMAP is used only as a visual scaffold: Euclidean geometry in the 2D projection is not interpreted as semantic distance. All quantitative summaries used in the visualization are computed in the original SPECTER-2 space using cosine distance, as defined in Section 2.5 (Equations (9) and (10)).
Figure 5 shows the proximity between generated proposals and materialized collaborations across shifted windows. The UMAP plane is used to obtain a two-dimensional visualization layout, only to position per-pair centroids and to define the hexagonal bins; Euclidean distances in the projection are not used for quantitative inference. The hexbin colors encode the mean cosine distance d cos ( a , b ) = 1 cos ( x ¯ a b , g ¯ a b ) —computed in the original SPECTER-2 space—for the pairs whose centroids fall in that region. In turn, the ring glyph radius encodes local distance magnitude, scaling with d cos (p95-scaled for visual stability).
Figure 6 shows the empirical distributions of d cos ( a , b ) across materialized pairs, together with Recall@t values defined over cosine-distance thresholds. This avoids interpreting 2D geometry as semantic distance, while still providing an interpretable visualization that is consistent with the high cosine-similarity alignment computed by Equation (14).
As shown in the figures, cosine distances are consistently small and the distributions concentrate most of their mass near low distances, with a long right tail (i.e., positive skew), indicating that the collaborations that materialize preserve semantic proximity with respect to the recommendations. The median cosine distance is 0.078 (2009–2012 → 2013–2016), 0.071 (2013–2016 → 2017–2020) and 0.08 (2017–2020 → 2021–2024). Using cosine distance thresholds in the original SPECTER-2 space, Recall@t reaches 62.50%, 92.00% and 67.90% at t = 0.10 , and 100% for all transitions at t = 0.20 . These results reinforce that the subset of materialized pairs remains semantically close to the generated proposals under an embedding-space distance, while the UMAP projection remains purely interpretive.

4. Discussion

We hypothesized that high document–document semantic proximity, even in the absence of prior co-authorship, indicates latent complementarity between pairs of authors and can therefore serve as a decision-support criterion for mediation-oriented prioritization. The empirical results are consistent with this: across the three temporal transitions analyzed, the absolute number of materializations increases (8 → 25 → 162), while semantic alignment between the generated proposals and the works that are actually published afterwards remains high (mean cosine similarity of 0.9146, 0.9275 and 0.9159) and thematic overlap measured via S2FOS remains substantial (100.00%, 100.00% and 77.16%), as shown in Table 3. This evidence indicates that the selection criterion based on cosine similarity with threshold τ identifies a shared thematic core that tends to persist when collaborations are eventually realized.
A brief sensitivity check for the cosine-similarity threshold τ is reported in Table 4. As expected, τ mainly controls the trade-off between candidate volume and semantic strictness: increasing τ reduces the pool of unlikely pairs | P τ < t | substantially (e.g., from 5,962 at τ = 0.86 to 1555 at τ = 0.90 and 125 at τ = 0.94 in 2009–2012 → 2013–2016). The corresponding SMR t ( τ ) typically increases as τ grows, largely because the denominator contracts, so this quantity should be interpreted jointly with the size of the pool. We used τ = 0.90 in Section 3 as a pragmatic operating point: it preserves non-zero materializations in all transitions while avoiding overly permissive pools at lower thresholds and overly sparse pools at higher thresholds. Importantly, | P τ < t | denotes the candidate pool after semantic filtering and prior to downstream prioritization; in operational use, decision makers would consume a ranked shortlist (e.g., top-N or percentile-based selection) derived from P τ < t according to available mediation resources.
The relatively low SMR values do not contradict the validity of the approach. By construction, unlikely pairs are defined by the absence of prior co-authorship, which places them outside the region of the co-authorship graph where new links are most likely to emerge spontaneously. In this context, the purpose of the system is not to maximize the probability of observing future edges under passive network evolution, but to prioritize semantically plausible opportunities that could become viable through active mediation, such as institutional introductions, targeted funding calls, seed grants, mobility programs, or incentive structures embedded in promotion and evaluation systems. In other words, the system identifies high-affinity collaboration seeds whose spontaneous materialization is expected to be rare. Thus, a low SMR should be interpreted as an inherent property of the task rather than as a failure of the approach. Conversely, a high SMR would suggest that the method is capturing collaborations that were already likely to emerge under passive network evolution, moving it closer to classical link prediction and weakening its exploratory intent. When materialization occurs, the observed semantic alignment and thematic overlap support the relevance of these candidates for strategic intervention. In this sense, the method shifts emphasis from predicting the most likely future edges to expanding the space of semantically plausible cross-community opportunities, which is better aligned with exploratory recommendation and institutional intervention goals.
The structural evolution described in the historical record of current and potential networks at PUCE provides a second interpretative angle. In contrast to topology-driven link prediction, which tends to reinforce existing collaboration patterns (e.g., via triadic closure and preferential attachment) and thus concentrate recommendations around already central authors and consolidated communities, the unlikely-pairs formulation has a different systemic intent. Growth in scale (authors and documents)—along with a substantial increase in mean degree, moderate densification and larger geodesic distances—indicates an expansion of the collaborative perimeter through the incorporation of external communities into the collaboration networks. In such networks, content-based suggested links may act as potential bridges between modules, potentially reducing effective paths between groups and diversifying thematic combinations. Moreover, because topology-driven predictors are anchored in observed connectivity patterns, their recommendations may concentrate on already dense regions of the graph and thus reproduce existing collaboration basins. By contrast, our exclusion of prior co-authorship and the absence of topological features aim to surface candidate ties that are structurally underrepresented, i.e., plausible cross-community bridges that can diversify mixing when coupled with intentional mediation. The networks of potential unlikely pairs also show that the recommendation space grows over time (from 1555 to 268,296 potential links), increasing the universe of options where mediation efforts can be focused, which is consistent with strategic goals such as interdisciplinarity and innovation.
The pipeline of research proposals and field-of-study annotation adds actionability to the recommendation. The field coherence observed in Table 3 is visually supported by Figure 5 and Figure 6, where UMAP is used as an exploratory layout to position centroids in two dimensions, while the quantitative summaries are grounded on cosine distances computed in the original SPECTER-2 space.
Finally, it is useful to synthesize the conceptual distinction. Topology-driven link prediction asks which collaborations are likely to occur anyway under passive network evolution, given historical structural cues. In contrast, unlikely-pair discovery asks which collaborations could occur if intentional mediation is applied, by surfacing semantically plausible opportunities that are not yet expressed in the co-authorship graph. In this sense, the two approaches are complementary rather than competing: link prediction supports forecasting and monitoring, whereas unlikely pairs support intervention-oriented opportunity discovery.

4.1. Limitations

The main limitations of this study are as follows:
  • The validation is retrospective with shifted windows; it does not capture the effect of interventions (e.g., an active matchmaking program) on materialization. Assessing intervention effects would require a prospective deployment with controlled exposure (e.g., randomized or quasi-experimental assignment of mediation resources) and follow-up over multi-year horizons, since collaboration formation and publication cycles are slow and noisy. In our setting, implementing such a study was not feasible within the project scope due to institutional coordination constraints (e.g., aligning multiple units, ensuring fair access to matchmaking resources and avoiding policy changes that confound outcomes), so we focus on a retrospective observational lower bound and leave causal evaluation of interventions to future work.
  • The coverage and normalization of metadata depend on OpenAlex, so missing abstracts, language biases or differential indexing may lead to underestimating real proximities. This comes in addition to the dependence on OpenAlex’s own disambiguation algorithms for assigning author profiles and affiliations.
  • The LLM component generates titles in a controlled format, but its practical usefulness depends on expert judgment. It should be viewed as an input for deliberation, not as a substitute. In addition, while the generative module is constrained and reproducible in format, future work should assess how variations in prompting or model choice may affect interpretive consistency.

4.2. Modularity and Replaceability

The presented pipeline is intentionally modular. Alternative scientific text encoders and proximity mechanisms could be substituted without changing the task formulation or the pipeline logic, although such substitutions would require recalibrating operational parameters and may affect density, noise tolerance and computational cost—thresholds are therefore not directly comparable across representation models.
Likewise, the role of the LLM module is intentionally decoupled from both pair detection and ranking. Its function is not to introduce additional predictive signals, but to act as a semantic articulation mechanism that translates abstract proximity relationships into concise, human-interpretable research directions. Because the task is narrowly scoped, constrained in format and run at low temperature, the pipeline does not rely on properties specific to a particular model. In this sense, the approach is largely model-agnostic within the class of contemporary instruction-following LLMs, and alternative generative or summarization techniques could be substituted without affecting the identification of unlikely pairs or the overall decision-support logic.

5. Conclusions and Future Work

The proposed framework responds to a practical need faced by research managers and policy makers who are expected to actively shape collaboration patterns rather than merely observe them. In many institutional and national contexts, collaboration is not left to emerge spontaneously, but is deliberately encouraged to achieve strategic objectives such as increasing international visibility, fostering interdisciplinarity or strengthening underrepresented research communities. In these settings, decision makers require tools that help them prioritize plausible collaboration opportunities among actors who do not yet share strong structural ties, but whose research profiles exhibit latent complementarity.
From a methodological standpoint, this work advances a reframing of collaborator recommendation as a decision-support problem rather than a purely predictive one. Instead of asking which collaborations are most likely to occur under passive network evolution, the unlikely-pairs formulation focuses on which collaborations could become viable if intentional mediation were applied. This distinction has direct implications for both system design and evaluation. By explicitly excluding prior co-authorship and avoiding graph-topological reinforcement mechanisms, the pipeline prioritizes semantically plausible opportunities that are structurally underrepresented in the historical collaboration network. As a result, the system is better aligned with exploratory and intervention-oriented objectives than with forecasting tasks. For information scientists and research information managers, this reframing highlights how semantic representations can be leveraged not only for retrieval or ranking, but also for institutional decision-making and policy design.
Empirically, the retrospective evaluation using shifted temporal windows shows that, while only a small fraction of recommended pairs materialize spontaneously in subsequent periods, the collaborations that do emerge exhibit high semantic alignment (high cosine similarity, equivalently low cosine distance) and substantial thematic coherence with the prior recommendations. These findings support the interpretation of semantic proximity as an early indicator of latent complementarity between researchers without prior ties. Importantly, the observed materialization rates should not be interpreted as a measure of predictive accuracy, but as a lower bound on spontaneous realization in the absence of any mediation. In contexts where institutional incentives or matchmaking mechanisms are introduced, these rates are expected to change, which further reinforces the role of the system as a prioritization tool rather than a predictor.
Operationally, the proposed pipeline is reproducible using open bibliographic graphs and is intentionally modular, allowing different embedding models, proximity mechanisms and generative components to be substituted without altering the underlying task formulation. This modularity facilitates transfer across institutions and supports integration into existing research management workflows, such as targeted funding calls, networking events or seed grant programs. The inclusion of a generative module further enhances actionability by translating abstract semantic proximity into concise, human-interpretable research directions that can support deliberation by both decision makers and researchers.
Future research can extend this work in several directions. First, controlled prospective studies could be conducted in which subsets of unlikely pairs receive explicit mediation, enabling direct comparison of materialization outcomes against baseline strategies. Second, the evaluation framework could be applied at broader scales, such as national or regional collaboration networks, to study systemic effects of intervention-oriented recommendation. In that vein, we are conducting an inter-institutional case study between the authors’ universities to foster collaborations aimed at co-developing PhD thesis projects, which will provide an opportunity to observe how recommendations translate into sustained research plans and supervisory structures. We are also initiating a case study that begins from the publications of all Spanish universities to encourage intra-national cooperation, motivated by the current fragmentation of research communities and the inefficiencies that arise when groups pursue overlapping agendas in parallel (see the reports of [23,24]). Finally, further exploration of alternative semantic representations and collaborative potential may help tailor the approach to different disciplinary contexts and institutional goals. In particular, for interdisciplinary matchmaking it may be necessary to go beyond semantic proximity and explicitly model complementarity (mutual added value) between researchers, since large cognitive distance can increase coordination costs and the benefits of diversity are often non-monotonic [25,26].

Author Contributions

Conceptualization, J.G.-M., M.L.-N. and J.P.S.-G.; Formal analysis, J.G.-M., M.L.-N. and D.P.-S.; Writing—original draft preparation, J.G.-M. and D.P.-S.; Writing—review and editing, J.G.-M., M.L.-N. and L.F.G.-V.; Visualization, J.G.-M. All authors have read and agreed to the published version of the manuscript.

Funding

The authors from the University of Vigo were supported by the Xunta de Galicia grant GPC-ED431B 2024/26 for the consolidation and structuring of competitive research units.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data supporting the findings of this study are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare that they have no conflicts of interest.

Abbreviations

AUC–ROCArea Under the Curve–Receiver Operating Characteristic
APIApplication Programming Interface
GPTGenerative Pre-trained Transformer
JSONJavaScript Object Notation
LLMLarge Language Model
PUCEPontificia Universidad Católica del Ecuador
S2FOSSemantic Scholar Fields of Study (multi-label classifier)
SMRSpontaneous Materialization Rate
UMAPUniform Manifold Approximation and Projection

References

  1. Haythornthwaite, C. Learning and knowledge networks in interdisciplinary collaborations. J. Am. Soc. Inf. Sci. Technol. 2006, 57, 1079–1092. [Google Scholar] [CrossRef]
  2. Pedersen, D.B. Collaborative Knowledge: The future of the academy in the knowledge-based economy. In On the Facilitation of the Academy; Brill: Leiden, The Netherlands, 2015; pp. 57–70. [Google Scholar]
  3. Liben-Nowell, D.; Kleinberg, J. The Link-Prediction Problem for Social Networks. J. Am. Soc. Inf. Sci. Technol. 2007, 58, 1019–1031. [Google Scholar] [CrossRef]
  4. Lü, L.; Zhou, T. Link prediction in complex networks: A survey. Phys. A Stat. Mech. Its Appl. 2011, 390, 1150–1170. [Google Scholar] [CrossRef]
  5. Lundberg, J.; Tomson, G.; Lundkvist, I.; Skår, J.; Brommels, M. Collaboration uncovered: Exploring the adequacy of measuring university-industry collaboration through co-authorship and funding. Scientometrics 2006, 69, 575–589. [Google Scholar] [CrossRef]
  6. Ullah, M.; Shahid, A.; Din, I.U.; Roman, M.; Assam, M.; Fayaz, M.; Ghadi, Y.; Aljuaid, H. Analyzing interdisciplinary research using Co-authorship networks. Complexity 2022, 2022, 2524491. [Google Scholar] [CrossRef]
  7. Abramo, G.; D’Angelo, C.; Solazzi, M. Assessing public–private research collaboration: Is it possible to compare university performance? Scientometrics 2010, 84, 173–197. [Google Scholar] [CrossRef]
  8. Bellanca, L. Measuring interdisciplinary research: Analysis of co-authorship for research staff at the University of York. Biosci. Horizons 2009, 2, 99–112. [Google Scholar] [CrossRef]
  9. Schlattmann, S. Capturing the collaboration intensity of research institutions using social network analysis. Procedia Comput. Sci. 2017, 106, 25–31. [Google Scholar] [CrossRef]
  10. Zhao, W.; Luo, J.; Fan, T.; Ren, Y.; Xia, Y. Analyzing and visualizing scientific research collaboration network with core node evaluation and community detection based on network embedding. Pattern Recognit. Lett. 2021, 144, 54–60. [Google Scholar] [CrossRef]
  11. Liang, W.; Zhou, X.; Huang, S.; Hu, C.; Xu, X.; Jin, Q. Modeling of cross-disciplinary collaboration for potential field discovery and recommendation based on scholarly big data. Future Gener. Comput. Syst. 2018, 87, 591–600. [Google Scholar] [CrossRef]
  12. Ye, G.; Xia, L. Analysis on cross-regional scientific research collaboration model. J. Libr. Sci. China 2019, 45, 79–95. [Google Scholar] [CrossRef]
  13. Hoang, D.T.; Tran, V.C.; Nguyen, T.T.; Nguyen, N.T.; Hwang, D. A consensus-based method to enhance a recommendation system for research collaboration. In Proceedings of the Asian Conference on Intelligent Information and Database Systems; Springer: Cham, Switzerland, 2017; pp. 170–180. [Google Scholar]
  14. Nguyen, T.T.; Nguyen, N.T.; Hoang, D.T.; Tran, V.C. Predicting Research Collaboration Trends Based on the Similarity of Publications and Relationship of Scientists. In Proceedings of the Asian Conference on Intelligent Information and Database Systems; Springer: Cham, Switzerland, 2020; pp. 15–24. [Google Scholar]
  15. Ye, G.; Wei, J.; Tan, Q.; Wu, C.; Song, X.; Li, S. Academic collaboration recommendation based on graph neural network and multi-attribute embedding. J. Inf. Sci. 2024. [Google Scholar] [CrossRef]
  16. Zhao, M.; Zhang, X.; Qin, H.; Ma, X.; Sun, H.; Sang, Y. Partner Recommendation Based on Scholar Embedding Method. In Proceedings of the 2023 7th International Conference on Communication and Information Systems (ICCIS); IEEE: Piscataway, NJ, USA, 2023; pp. 118–128. [Google Scholar]
  17. Guerra, J.; Quan, W.; Li, K.; Ahumada, L.; Winston, F.; Desai, B. Scosy: A biomedical collaboration recommendation system. In Proceedings of the 2018 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC); IEEE: Piscataway, NJ, USA, 2018; pp. 3987–3990. [Google Scholar]
  18. Yang, N.; Jo, J.; Jeon, M.; Kim, W.; Kang, J. Semantic and explainable research-related recommendation system based on semi-supervised methodology using BERT and LDA models. Expert Syst. Appl. 2022, 190, 116209. [Google Scholar] [CrossRef]
  19. Wu, M.; Zhang, Y.; Lu, J.; Lin, H.; Grosser, M. Recommending scientific collaborators: Bibliometric networks for medical research entities. In Proceedings of the Developments of Artificial Intelligence Technologies in Computation and Robotics: Proceedings of the 14th International FLINS Conference (FLINS 2020); World Scientific: Singapore, 2020; pp. 480–487. [Google Scholar]
  20. Priem, J.; Piwowar, H.; Orr, R. OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. arXiv 2022, arXiv:2205.01833. [Google Scholar]
  21. Singh, A.; D’Arcy, M.; Cohan, A.; Downey, D.; Feldman, S. SciRepEval: A Multi-Format Benchmark for Scientific Document Representations. In Proceedings of the Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Singapore, 2022. [Google Scholar]
  22. Kinney, R.; Anastasiades, C.; Authur, R.; Beltagy, I.; Bragg, J.; Buraczynski, A.; Cachola, I.; Candra, S.; Chandrasekhar, Y.; Cohan, A.; et al. The Semantic Scholar Open Data Platform. arXiv 2025, arXiv:2301.10140. [Google Scholar]
  23. European Commission. ERA Country Report 2023: Spain. European Research Area Platform. 2023. Available online: https://european-research-area.ec.europa.eu/country-report-spain (accessed on 26 February 2026).
  24. European Commission. ERA Country Report 2024: Spain. European Research Area Platform. 2024. Available online: https://european-research-area.ec.europa.eu/documents/country-report-spain (accessed on 26 February 2026).
  25. Feng, S.; Kirkley, A. Mixing Patterns in Interdisciplinary Co-Authorship Networks at Multiple Scales. Sci. Rep. 2020, 10, 7731. [Google Scholar] [CrossRef] [PubMed]
  26. Wang, G.; Gan, Y.; Yang, H. The Inverted U-Shaped Relationship between Knowledge Diversity of Researchers and Societal Impact. Sci. Rep. 2022, 12, 18585. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Recommendation pipeline for unlikely pairs and research proposal generation.
Figure 1. Recommendation pipeline for unlikely pairs and research proposal generation.
Information 17 00254 g001
Figure 2. Current collaboration networks in each period. Blue nodes represent authors from the institution; gray nodes represent external ones. Green links represent co-authorship ties.
Figure 2. Current collaboration networks in each period. Blue nodes represent authors from the institution; gray nodes represent external ones. Green links represent co-authorship ties.
Information 17 00254 g002
Figure 3. Unlikely pairs detected by our system in each window, with τ = 0.90 . Blue nodes represent authors from the institution; gray nodes represent external ones. Red links represent potential collaborations.
Figure 3. Unlikely pairs detected by our system in each window, with τ = 0.90 . Blue nodes represent authors from the institution; gray nodes represent external ones. Red links represent potential collaborations.
Information 17 00254 g003
Figure 4. Illustration of the research proposal pipeline, from model input to JSON output, with research proposals and field-of-study classification of research proposals. In this particular case, the proposals relate to Chagas disease research in Ecuador, including parasite genetic diversity, socioeconomic risk drivers, integrated vector-control strategies, environmental-change impacts on transmission, and congenital transmission in endemic regions. In the annotation, the predominant labels span Medicine and Environmental Science, with additional overlap into Biology and Sociology.
Figure 4. Illustration of the research proposal pipeline, from model input to JSON output, with research proposals and field-of-study classification of research proposals. In this particular case, the proposals relate to Chagas disease research in Ecuador, including parasite genetic diversity, socioeconomic risk drivers, integrated vector-control strategies, environmental-change impacts on transmission, and congenital transmission in endemic regions. In the annotation, the predominant labels span Medicine and Environmental Science, with additional overlap into Biology and Sociology.
Information 17 00254 g004
Figure 5. Hexbin map in the UMAP plane.
Figure 5. Hexbin map in the UMAP plane.
Information 17 00254 g005
Figure 6. Proximity between generated proposals and materialized collaborations across shifted windows.
Figure 6. Proximity between generated proposals and materialized collaborations across shifted windows.
Information 17 00254 g006
Table 1. Processing summary of the pipeline stages and their main artifacts.
Table 1. Processing summary of the pipeline stages and their main artifacts.
StepInput (Figure 1)ProcessingOutput
1Query parameters (institution ID, year range)Parameterized OpenAlex API queries under a temporal window W t .Retrieved identifiers and raw records.
2OpenAlex API (works and authors)Paginated retrieval and persistence with a timestamp for each run; window assignment for downstream stages.Corpus segmented into windows (titles/abstracts) and author metadata.
3Building unlikely pairs (SPECTER-2 + document distance + threshold filter + projection)SPECTER-2 encoding of titles/abstracts with 2 normalization. Cosine similarity s ( d i , d j ) and its distance counterpart 1 s . Threshold τ , optional top-k and trivial-match filtering. Document→author projection via w a b and exclusion of prior co-authorship C a b < t = 1 .Ranked unlikely pairs P τ < t with semantic scores w a b .
4Proposal generation and field annotation (OpenAI API + S2FOS)Prompting over representative works per pair; constrained JSON generation of proposal titles; multi-label field-of-study annotation via S2FOS (23 labels).Proposals and S2FOS labels stored per pair/window.
5Response (decision makers, researchers)Packaging of the shortlist for deliberation and communication.Shortlist of unlikely pairs and proposals.
Table 2. Two end-to-end examples of proposal generation for unlikely pairs that later materialized under the shifted-window validation step ( W t = 2013 2016 W t + 1 = 2017 –2020).
Table 2. Two end-to-end examples of proposal generation for unlikely pairs that later materialized under the shifted-window validation step ( W t = 2013 2016 W t + 1 = 2017 –2020).
Representative inputs in W t (potential)Generated proposal titles in W t (5 items) with S2FOSS2FOS overlapObserved collaboration in W t + 1 Coherence note
Pair: María F. Checa × Rafael E. Cárdenas
María F. Checa:
Title: Microclimate Variability Significantly Affects the Composition, Abundance and Phenology of Butterfly Communities in a Highly Threatened Neotropical Dry Forest.
DOI: https://doi.org/10.1653/024.097.0101
Rafael E. Cárdenas:
Title: Fine-scale climatic variation drives altitudinal niche partitioning of tabanid flies in a tropical montane cloud forest, Ecuadorian Chocó.
DOI: https://doi.org/10.1111/icad.12146
(1) Investigating Microclimate Effects on Tabanid Distribution in Tropical Forests. S2FOS: Environmental Science, Biology.
(2) Assessing Butterfly and Tabanid Responses to Climate Change in Ecuador. S2FOS: Environmental Science, Biology.
(3) Comparative Study of Insect Community Dynamics in Varying Microhabitats. S2FOS: Biology, Environmental Science.
(4) Microhabitat Preferences of Butterflies and Tabanids in Tropical Ecosystems. S2FOS: Environmental Science, Biology.
(5) Exploring Altitudinal Niche Partitioning of Tropical Insects. S2FOS: Biology, Environmental Science.
2/2
(100.00%)
Title: Forest stratification shapes allometry and flight morphology of tropical butterflies.
DOI: https://doi.org/10.1098/rspb.2020.1071
S2FOS: Environmental Science, Biology.
The prompt inputs already share a tight climatic/microclimatic framing of insect ecology (vertical/altitudinal gradients, niche partitioning, and fine-scale environmental constraints). The realized collaboration shifts the lens to flight morphology under forest stratification, but it stays inside the same explanatory layer where micro-environments shape insect traits and community structure.
Pair: Fabián Cueva × Iván Rueda Fierro
Fabián Cueva:
Title: Culture, change and learning in project-based organizations.
DOI: https://doi.org/10.29019/eyn.v7i1.255
Iván Rueda Fierro:
Title: Relationship between strategic plan and organizational learning as an element of knowledge management in higher education institutions.
DOI:
(1) Exploring Project Management Strategies for Organizational Learning in Higher Education. S2FOS: Education, Business.
(2) Cultural Adaptability and Strategic Planning in Project-Based Organizations. S2FOS: Business, Sociology.
(3) Enhancing Knowledge Management through Project-Based Learning Initiatives. S2FOS: Education, Business.
(4) Integrating Organizational Culture in Strategic Planning for Project Success. S2FOS: Business.
(5) Developing Learning Frameworks for Effective Project Management in Educational Institutions. S2FOS: Education.
2/3
(66.67%)
Title: Leadership and communication, fundamental competencies for a project director.
DOI: https://doi.org/10.29019/eyn.v9i1.445
S2FOS: Business, Education.
The generated proposals operationalize the same core management layer (organizational learning, planning, and knowledge management) in project-based and higher-education settings. The realized collaboration narrows to leadership and communication competencies for project directors, which is a direct, practice-facing slice of the same project-management/learning agenda rather than a thematic drift.
Table 3. Validation of materialized unlikely pairs.
Table 3. Validation of materialized unlikely pairs.
Window ( W t W t + 1 )Materialized (n)Mean Cosine
Similarity
Field of Study
Overlap (%)
Proposal
Density
2009–2012 → 2013–201680.9146100.002.98
2013–2016 → 2017–2020250.9275100.004.54
2017–2020 → 2021–20241620.915977.164.10
Table 4. Sensitivity of results to the cosine-similarity threshold τ .
Table 4. Sensitivity of results to the cosine-similarity threshold τ .
2009–2012 → 2013–20162013–2016 → 2017–20202017–2020 → 2021–2024
τ | P τ < t | n t SMR t ( τ ) | P τ < t | n t SMR t ( τ ) | P τ < t | n t SMR t ( τ )
0.865962100.17%30,689380.12%951,4152220.02%
0.88340990.26%16,716350.21%566,3622060.04%
0.90155580.51%6906250.36%268,2961620.06%
0.9250940.79%1683100.59%86,9401060.12%
0.9412510.80%32351.55%15,726240.15%
0.960000225620.09%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Galán-Mena, J.; López-Nores, M.; Pulla-Sánchez, D.; Guerrero-Vásquez, L.F.; Salgado-Guerrero, J.P. Unlikely Pairs: A Decision-Support Recommendation Pipeline for Discovering Semantically Plausible Research Collaborations. Information 2026, 17, 254. https://doi.org/10.3390/info17030254

AMA Style

Galán-Mena J, López-Nores M, Pulla-Sánchez D, Guerrero-Vásquez LF, Salgado-Guerrero JP. Unlikely Pairs: A Decision-Support Recommendation Pipeline for Discovering Semantically Plausible Research Collaborations. Information. 2026; 17(3):254. https://doi.org/10.3390/info17030254

Chicago/Turabian Style

Galán-Mena, Jorge, Martín López-Nores, Daniel Pulla-Sánchez, Luis Fernando Guerrero-Vásquez, and Juan Pablo Salgado-Guerrero. 2026. "Unlikely Pairs: A Decision-Support Recommendation Pipeline for Discovering Semantically Plausible Research Collaborations" Information 17, no. 3: 254. https://doi.org/10.3390/info17030254

APA Style

Galán-Mena, J., López-Nores, M., Pulla-Sánchez, D., Guerrero-Vásquez, L. F., & Salgado-Guerrero, J. P. (2026). Unlikely Pairs: A Decision-Support Recommendation Pipeline for Discovering Semantically Plausible Research Collaborations. Information, 17(3), 254. https://doi.org/10.3390/info17030254

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop