1. Introduction
Scientific collaboration networks constitute an important framework for understanding organization, knowledge diffusion, and systemic resilience in science [
1,
2,
3]. Empirical studies consistently report heavy-tailed degree distributions, small-world organization, and degree assortativity in such systems [
4,
5]. Despite the abundance of descriptive analyses, however, methodological consistency in identifying and comparatively characterizing mesoscopic structural regimes across subnetworks remains limited. In particular, the relationship between small-world organization and fractal self-similarity remains debated, both conceptually and in terms of the statistical criteria used for their identification [
6,
7]. Recent analyses of the Brazilian Agricultural Research Corporation (Embrapa) reinforce these structural patterns in an applied institutional context [
8]; however, a systematic, reproducible, and integrated multi-criteria framework for comparative subnetwork characterization is still lacking.
Community detection has been widely investigated in complex networks, including unipartite and bipartite representations, with approaches ranging from direct algorithms for bipartite graphs [
9] to modularity-based methods applied to projected networks. While these techniques effectively identify mesoscopic partitions, less attention has been devoted to the structural characterization of the resulting communities. In many empirical studies, communities are reported primarily as partitions, without evaluating whether they correspond to distinct topological organizations or structural regimes.
Complex network theory further suggests that large-scale coordination and robustness may emerge from repeated local interactions among agents. Iterative decisions involving cooperation, imitation, brokerage, and bounded rational search can generate recurrent structural patterns such as heavy-tailed weighted degree distributions, clustered organization, shortcut-induced small-world structure, or disassortative mixing, even in the absence of centralized planning. These recurrent structural patterns have been widely discussed in the literature as possible outcomes of decentralized interaction dynamics in social and rural institutional systems [
10,
11,
12]. However, relatively few studies translate these theoretical insights into explicit statistical criteria capable of distinguishing structural regimes within a single giant connected component.
From a network perspective, the emergence of a giant connected component can be interpreted as a phase-transition phenomenon in which link density exceeds a critical threshold. Importantly, such components are rarely homogeneous: multiple structural regimes may coexist within the same large-scale connected structure, reflecting thematic, institutional, spatial, or incentive-driven constraints. Evidence of this coexistence, including hub-dominated regions, clustered small-world areas, and sparsely connected fractal-like zones, has been reported in cooperative and organizational knowledge networks [
12,
13,
14]. Nevertheless, the absence of standardized analytical pipelines limits cross-study comparability and weakens reproducibility in comparative structural characterization.
This gap motivates the present study. The central research question addressed here is whether distinct and statistically identifiable structural regimes can coexist within a single giant connected component of a scientific collaboration network and how such regimes can be comparatively characterized using reproducible statistical criteria. To answer this question, we develop and demonstrate a statistically grounded and algorithmically explicit framework for comparative cluster-level structural characterization. Rather than treating community detection as an end in itself, we use it as an intermediate step in an integrated multi-criteria classification framework combining weighted degree-tail modeling, information-entropy measures of degree–degree correlations, small-world diagnostics, and fractal analysis based on the Song–Havlin–Makse box-covering renormalization framework. The complementary statistical and structural diagnostics are integrated through an explicit classification protocol that assigns network families while distinguishing supported and ambiguous classifications according to the overall consistency of the available evidence. The resulting categories should therefore be interpreted as comparative topological descriptions derived from statistical diagnostics, rather than definitive evidence of specific underlying generative mechanisms.
Applying this framework to the giant coauthorship component of Embrapa’s scientific production derived from the BDPA, we identify the coexistence of distinct network families, including BA-like/scale-free small-world, scale-free fractal (non-BA) and small-world (non-scale-free) organizations within the same connected collaboration system. Rather than assigning deterministic labels to individual communities, the proposed framework integrates complementary statistical and structural evidence to identify network families while explicitly representing the confidence associated with each classification. These findings indicate that large scientific collaboration networks may display substantial mesoscopic structural heterogeneity even after percolation, challenging interpretations based solely on global averages. By providing explicit decision rules and a reproducible integrated classification protocol, this study contributes a transferable methodological framework for the comparative topological characterization of heterogeneous collaboration networks.
2. Materials and Methods
The integrated multi-criteria classification framework adopted in this study follows a structured procedure designed to characterize the mesoscopic topological organization of empirical subnetworks in a statistically robust and reproducible manner. The framework consists of the following sequential stages:
First, a weighted coauthorship network is constructed from bibliographic metadata, and its giant connected component is extracted to ensure global connectivity and statistical stability.
Second, the weighted network is partitioned into communities using the Louvain community detection algorithm, allowing the analysis to focus on mesoscopic subnetworks that may exhibit distinct structural organizations.
Third, classical weighted structural metrics are computed for each community, including density, clustering coefficient, assortativity, average shortest-path length, and complementary topological descriptors. These metrics provide baseline information on connectivity, hierarchy, and local cohesion.
Fourth, small-world behavior is assessed by comparing the empirical clustering coefficient and average path length with Erdős–Rényi random graph baselines, yielding the small-worldness index , which quantifies the coexistence of strong local clustering and efficient global connectivity.
Fifth, the weighted degree-tail distribution of each community is analyzed through maximum-likelihood estimation. The lower cutoff is determined by Kolmogorov– Smirnov minimization, goodness-of-fit is evaluated by bootstrap resampling, and competing power-law, exponential, and log-normal models are compared using Akaike’s Information Criterion (AIC).
Sixth, fractal scaling is evaluated using the Song–Havlin–Makse box-covering renormalization framework. The relationship between the number of covering boxes and box diameter is modeled and compared using AIC-based model selection to determine whether the observed scaling is more consistent with a power-law model or an exponential decay.
Seventh, information-entropy measures of degree–degree correlations, including mutual information and conditional entropy, are computed to complement the structural characterization by quantifying the statistical dependence between connected node degrees.
Finally, all statistical and structural diagnostics are integrated through an explicit integrated classification protocol. Rather than relying on a single metric, the protocol combines complementary evidence from weighted degree-tail modeling, small-world diagnostics, fractal analysis, and information-theoretic measures to assign each community to a network family (BA-like/scale-free small-world, scale-free fractal (non-BA) and small-world (non-scale-free), or undefined) while explicitly distinguishing supported and ambiguous classifications according to the overall consistency of the available evidence.
This integrated framework ensures that the final classification emerges from the convergence of multiple complementary statistical and structural diagnostics, providing a reproducible and transferable methodology for the comparative characterization of heterogeneous mesoscopic regimes in complex collaboration networks.
2.1. Network and Giant Component Construction
The empirical dataset was obtained from the Brazilian Agricultural Research Database (BDPA), which indexes the scientific production of Embrapa researchers. Bibliographic metadata were retrieved directly from BDPA through its native JSON export interface and subsequently processed using custom Python routines in Python 3.12.3, using the NumPy 1.26.4, Pandas 2.1.4, NetworkX 2.8.8, and Matplotlib 3.6.3 libraries. The exported records include publication identifiers, publication year, author lists, institutional affiliations, and additional bibliographic metadata. Only peer-reviewed journal articles published between 1974 and 2024 were considered in this study.
Prior to network construction, duplicate publication records were removed, and author names were normalized to reduce textual inconsistencies arising from spelling variations, abbreviations, and formatting differences. Since BDPA does not provide persistent author identifiers (e.g., ORCID) for all records, a rule-based normalization and heuristic canonicalization procedure was applied before constructing the collaboration network. The procedure standardized trivial orthographic variations by removing leading Portuguese prepositions, eliminating unnecessary whitespace, and harmonizing common suffix abbreviations (e.g., jr” to junior”). Candidate author variants were subsequently identified using a conservative rule-based heuristic to reduce textual inconsistencies while avoiding unnecessary merges whenever potential ambiguities were detected. The resulting canonical author table was further subjected to quality-control procedures, including the detection of malformed or truncated names, consistency checks between intermediate network representations, and verification that isolated authors were preserved during network construction.
The coauthorship network was modeled as an undirected weighted graph, in which each node represents an author. For every publication containing multiple authors, all possible pairs of coauthors were connected, generating a complete clique associated with that publication. Whenever the same pair of authors collaborated in multiple publications, the corresponding edge weight was incremented according to the cumulative number of joint publications throughout the observation period. Thus, edge weights quantify the collaboration intensity between pairs of researchers.
After preprocessing, the largest connected component (giant component), containing 60,636 authors, was extracted using the connected-components algorithm implemented in the NetworkX library. Restricting the analysis to the giant component ensures that all subsequent structural analyses are performed on a connected network, avoiding artifacts associated with isolated nodes and disconnected subgraphs.
Communities were identified using the Louvain algorithm [
15], executed on the weighted network using collaboration frequencies as edge weights. The resulting partition contains 25 communities. The Louvain algorithm identifies communities by maximizing the weighted network modularity,
where
denotes the weight of the edge connecting nodes
i and
j,
is the weighted degree (node strength) of node
i,
is the total weight of the network,
and
denote the communities to which nodes
i and
j belong, and
is the Kronecker delta function, which equals 1 when both nodes belong to the same community and 0 otherwise. High values of
Q indicate a strong community structure, where the observed collaboration intensity within communities is significantly higher than expected under the weighted null model.
Throughout this study, the term degree refers to the weighted degree (node strength), defined as the sum of the weights of all edges incident on a node, unless otherwise stated. Consequently, all degree-distribution analyses were performed using the weighted degree. Community detection was also carried out on the weighted network, whereas structural descriptors based primarily on network topology, including the clustering coefficient, assortativity, and average shortest-path length, were computed using their standard graph-theoretic definitions.
2.2. Per-Cluster Structural Metrics
For each cluster we compute several structural quantities that capture its internal connectivity patterns:
Density quantifies how many edges exist relative to the maximum possible number of edges in a fully connected graph,
where
m is the number of edges and
n is the number of nodes in the cluster. Values of
close to 1 indicate highly connected subgraphs.
Average clustering coefficient [
4] measures the tendency of nodes to form triangles,
where
is the number of edges among the neighbors of node
i, and
is the degree of node
i. The term
represents the local clustering of node
i, while
C is the average over all nodes.
Degree assortativity [
5] captures degree correlations between connected nodes,
where
is the joint probability that an edge connects nodes of degrees
j and
k,
is the degree distribution of the nodes at the ends of edges, and
is the variance of
. Positive
r indicates that nodes tend to connect with others of similar degree (assortative mixing), while negative
r reflects disassortative networks.
Average path length represents the typical separation between pairs of nodes,
where
is the length of the shortest path between nodes
i and
j. For large clusters,
L is estimated by random sampling of node pairs.
To evaluate whether clusters exhibit small-world properties, we compare the empirical clustering coefficient and path length to those of equivalent Erdős–Rényi random graphs (
and
), and compute the small-worldness index [
16]:
where values of
indicate small-world behavior, combining high local clustering with short global distances.
2.3. Degree-Tail Modeling, Determination, and Model Selection
We assess whether cluster-level degree distributions exhibit power-law tails by combining maximum-likelihood estimation (MLE), Kolmogorov–Smirnov (KS) minimization to select the lower cutoff
, a bootstrap
p-value for goodness-of-fit, and Akaike’s Information Criterion (AIC) for model comparison, following [
17,
18,
19].
Data and notation: Let be the observed degrees in a given cluster. For a candidate lower cutoff , define the tail sample of size n.
2.3.1. Power-Law MLE in the Tail
For discrete degrees, the power-law probability mass above
is
where
is the Hurwitz zeta function ensuring normalization. The MLE for the exponent, using the standard
continuity correction, is
here,
;
denotes observed degrees in the tail.
2.3.2. Empirical Versus Theoretical CCDF and the KS Distance
Define the empirical complementary CDF (CCDF) on the tail as
and the fitted power-law CCDF (continuous approximation) as
The Kolmogorov–Smirnov statistic is the maximum vertical gap between these CCDFs:
We choose the cutoff that minimizes this discrepancy:
where
is the set of unique observed degrees.
2.3.3. Goodness-of-Fit via Bootstrap p-Value
To test the plausibility of the fitted tail, we generate
B synthetic tails from
, each of size
n, and compute their KS distances
. The bootstrap
p-value is
with
. Large
p (e.g.,
) indicates that the observed discrepancy is compatible with sampling noise under the fitted model.
2.3.4. Model Selection with AIC
We compare, on the same empirical tail, three candidate models: power-law (PL), exponential (EXP), and left-truncated log-normal (LN). Let
be the maximized likelihood and
the number of free parameters of model
. The Akaike Information Criterion is
and the preferred tail model is the one with the smallest AIC. Relative support is evaluated by
Following Burnham and Anderson [
19],
indicates that competing models receive comparable statistical support,
indicates considerably less support, and
provides strong evidence against the model relative to the best candidate.
In the present classification framework, however, clusters with are not considered to have a uniquely supported tail model. Instead, these cases are treated as statistically ambiguous and require complementary evidence from the remaining structural diagnostics before a final topological assignment is made.
For each cluster we report , KS , for the PL, EXP, and LN models, together with the selected best tail. The bootstrap KS test evaluates the absolute plausibility of the fitted distribution, whereas AIC quantifies the relative statistical support among competing models.
2.4. Fractality via Box Covering
To evaluate whether the collaboration clusters exhibit self-similar organization, we performed a box-covering analysis following the renormalization framework introduced by Song, Havlin, and Makse [
6] and further developed by Rozenfeld et al. [
7]. Let
denote the largest connected component of a cluster. For a given box diameter
, let
denote the number of covering boxes required to cover all vertices of the LCC, where each box contains only vertices whose pairwise shortest-path distance does not exceed
.
Fractal networks are expected to satisfy the scaling relation
where
denotes the box (fractal) dimension. In contrast, hub-dominated small-world networks exhibit an exponential decay,
reflecting the rapid collapse of distances during successive renormalization steps [
6,
7].
2.4.1. Operational Definition and Algorithmic Implementation
Computing the exact minimum box covering is an NP-hard problem [
6,
7]. We therefore implemented a deterministic greedy box-covering heuristic inspired by the original renormalization framework.
For each covering radius , the vertices of the largest connected component are first sorted in decreasing order of degree. The highest-degree uncovered vertex is then selected as the center of a new box. All vertices located within graph distance from this center are identified by a breadth-first search (BFS) and assigned to the same box. Covered vertices are removed from the remaining search space, and the procedure is repeated until every vertex has been assigned to a box.
The corresponding box diameter is defined as
ensuring that any two vertices belonging to the same box remain separated by a shortest-path distance not exceeding
.
The procedure is repeated for , corresponding to box diameters . For each radius, the resulting number of covering boxes is recorded.
To reduce finite-size effects associated with the trivial regime where the entire network is covered by only one or two boxes, only observations satisfying are retained. The largest contiguous scaling interval satisfying this condition is subsequently used for the regression analysis.
2.4.2. Model Fitting and Notation
Although both the degree-tail analysis (
Section 2.3) and the present box-covering analysis employ the Akaike Information Criterion (AIC), they address different statistical questions. In
Section 2.3, the AIC compares alternative probability distributions describing the weighted degree tail (power-law, exponential, and log-normal). Here, the AIC compares two competing scaling relationships between the box diameter (
) and the number of covering boxes (
): a power-law model, expected for fractal networks, and an exponential model, expected for non-fractal small-world networks.
The retained observations are fitted using two competing regression models,
The power-law model is fitted in log–log coordinates, whereas the exponential model is fitted in semi-log coordinates.
For the power-law regression, we compute the coefficient of determination,
which quantifies the quality of the fractal scaling.
Model comparison is performed using the Akaike Information Criterion,
where
is the residual sum of squares,
n is the number of fitted observations, and
is the number of fitted parameters.
The quantity
is then used to compare the competing scaling models. Positive values indicate better support for the power-law model (fractal organization), whereas negative values favor the exponential model (small-world organization).
Following the conventional interpretation of AIC differences, values satisfying indicate comparable support for the competing models, values between 4 and 7 indicate moderate evidence, and indicate strong evidence for the model with the lower AIC.
2.4.3. Decision Criteria
The box-covering analysis provides one component of the integrated multi-criteria classification framework proposed in this work. Fractal scaling is considered to be supported when the power-law scaling model is preferred over the exponential alternative (
) and the corresponding log–log regression exhibits high goodness of fit. Conversely, exponential scaling (
) supports a non-fractal small-world organization. Cases exhibiting comparable statistical support (
) or poor regression quality are regarded as inconclusive and are subsequently resolved by the integrated classification protocol described in
Section 2.5.
From the perspective of network renormalization, each covering box is replaced by a super-node, generating a coarse-grained representation of the original network. Fractal networks preserve their heterogeneous organization under successive renormalization steps, producing the power-law scaling described by Equation (
16). In contrast, hub-dominated small-world networks rapidly lose structural heterogeneity as hubs are merged, resulting in the exponential behavior described by Equation (
17) [
6,
7]. Consequently, the box-covering analysis provides an independent structural characterization that complements the degree-tail modeling and the small-world diagnostics within the integrated classification framework.
2.5. Integrated Classification Protocol
To reflect the overall consistency of the evidence, each network-type assignment was associated with one of two qualitative confidence levels:
-Supported: the statistical and structural diagnostics consistently converged toward the same topological regime.
-Ambiguous: competing statistical models exhibit comparable support (e.g., ), the bootstrap KS test does not provide sufficient evidence for a unique tail model, or different structural diagnostics lead to conflicting conclusions.
Here, the term “BA-like” is used exclusively as a phenomenological topological descriptor derived from the combined statistical and structural evidence rather than as proof of a specific network generative mechanism.
This integrative protocol ensures that the final classification is based on the consistent agreement among multiple independent statistical and structural diagnostics rather than on any single network metric. The corresponding decision workflow is summarized schematically in
Figure 1.
In the
Supplementary Materials File, we provide a validation of the framework presented in
Figure 1 against synthetic, benchmark networks with known structures. The pipeline was tested on canonical Barabási–Albert, Watts–Strogatz, and fractal/hierarchical networks; the classification accuracy, robustness, and sensitivity to parameter choices are reported.
2.6. Temporal Growth Analysis
To complement the structural characterization of the collaboration network, we estimated the temporal growth exponent
originally introduced by Barabási and Albert [
20]. Under the preferential-attachment framework, the cumulative degree of node
i evolves approximately as
where
is the cumulative weighted degree of author
i at time
t, and
denotes the year in which the author first appeared in the coauthorship network. Under the classical Barabási–Albert model, the expected value is
, whereas deviations from this prediction may reflect departures from idealized linear preferential attachment.
For each author, the cumulative weighted degree was computed annually from the year of the first publication recorded in the BDPA database until the end of the observation period (2024). The temporal variable used in the regression was therefore the elapsed time since the author’s entry into the collaboration network,
which avoids undefined values at the first observation and standardizes the temporal evolution of authors entering the network in different years.
The exponent
was estimated independently for each author by fitting a linear regression in logarithmic coordinates,
where
is the estimated temporal growth exponent and
is a normalization constant. Only regressions with statistically significant slopes (
) were retained for subsequent analyses. The resulting
estimates were then grouped according to the cluster membership obtained from the integrated structural classification protocol.
The temporal growth analysis is not part of the proposed classification framework and was not used to assign network types. Instead, the exponent is employed exclusively as a complementary descriptive indicator that provides an independent temporal perspective on the evolution of the structurally classified clusters. Consequently, the identification of BA-like/scale-free, scale-free fractal (non-BA), and small-world (non-scale-free) regimes relies entirely on the structural diagnostics described in the previous subsections.
Some limitations should be noted. Because authors enter the collaboration network at different times, the estimated exponents may be affected by finite observation windows and survivorship bias, particularly for authors with short publication histories. Furthermore, although only statistically significant regressions () were retained, no formal multiple-testing correction was applied. Therefore, the estimated values should be interpreted as descriptive indicators of temporal growth rather than as definitive evidence of preferential attachment, node fitness, or any specific network generative mechanism.
3. Results
Figure 2 shows the clusters identified using the Louvain algorithm.
Table 1 summarizes the main structural metrics computed for each community identified within the Embrapa coauthorship network. These parameters describe how densely each subnetwork is connected, the extent of clustering among neighboring nodes, the average distance between any two authors, and the degree of homophily or disassortative mixing. Together, these metrics allow us to compare structural regimes across clusters and evaluate how local connectivity patterns may affect robustness, information diffusion, and collaboration efficiency.
In general, denser and highly clustered subnetworks indicate more cohesive research groups, while lower density and longer path lengths are typical of large, heterogeneous communities that bridge distinct thematic or institutional domains [
1,
4]. Degree assortativity further reveals whether well-connected authors preferentially collaborate with similarly connected peers or with less-connected individuals, shaping the hierarchical or more egalitarian character of the collaboration structure [
21]. The data show strong heterogeneity among clusters. Larger communities (e.g., clusters 0–5) exhibit very low density (on the order of
) and moderate assortativity, consistent with scale-free small-world structures where a few highly connected authors act as bridges among multiple local groups. These clusters also show high average clustering coefficients (around 0.8) combined with moderate path lengths (
), indicating efficient local cohesion and global reach, a hallmark of small-world organization.
In contrast, smaller clusters (e.g., 12–15) are characterized by high density (>0.4) and clustering (>0.9), as well as short average path lengths (L< 2). These metrics reveal tightly knit research groups, often corresponding to specialized or institutionally localized collaborations. Clusters with negative assortativity (e.g., 2, 6, 17, 23) suggest hub-mediated communication, where central authors connect otherwise weakly linked nodes, an architecture that enhances reachability but can reduce resilience to targeted node removal. Conversely, positive assortativity values (e.g., clusters 7, 10, 13, 19) point to more balanced collaboration patterns, where partnerships among well-connected researchers reinforce the network’s structural stability and knowledge diffusion.
Overall, the observed combination of high clustering, moderate path lengths, and heterogeneous densities supports the interpretation of the Embrapa coauthorship system as a multi-scale, small-world network, simultaneously favoring local cohesion and global connectivity—an advantageous configuration for fostering interdisciplinary collaboration and long-term institutional resilience.
Table 2 presents the parameters resulting from the tail-fit analysis performed for each cluster in the Embrapa coauthorship network. The table includes, for each cluster, the cutoff value
, the number of observations in the degree tail (
), the power-law exponent
, the Kolmogorov–Smirnov distance (
D) and corresponding
p-value, as well as the AIC values for the three alternative models tested: power-law (PL), exponential (EXP), and log-normal (LN). The final column (“Best tail”) identifies the model providing the best statistical description of the empirical weighted degree distribution according to the minimum AIC.
The final network-type classification summarized in
Table 3 was obtained by applying the integrated multi-criteria protocol described in
Section 2.5. This protocol combines tail-model selection based on the weighted degree distribution, the bootstrap KS goodness-of-fit test, the small-worldness index (
), and the box-covering analysis to identify coherent mesoscopic structural regimes. The reported confidence level reflects the overall agreement among these independent statistical and structural diagnostics.
Table 3 summarizes the quantities entering the decision protocol together with the resulting network-type assignment and its associated confidence level.
Although the integrated classification protocol is based on the combined interpretation of multiple statistical and structural diagnostics, it is useful to report the fitted parameters associated with each candidate tail model.
Table 4 summarizes the estimated parameters obtained during the tail-fitting procedure for every cluster, including the power-law exponent (
), the parameters of the log-normal distribution (
and
), and the exponential decay rate (
). These quantities provide a complete description of the fitted statistical models and facilitate future comparisons with other empirical collaboration networks.
Reporting the fitted parameters also provides additional insight into the diversity of degree-tail behaviors observed across the collaboration network. While the integrated classification relies on multiple complementary diagnostics, the estimated model parameters quantify the characteristics of the candidate distributions themselves, allowing differences in scaling exponents, dispersion, and exponential decay rates to be directly compared among clusters. Such information may be useful in future studies investigating the evolution of collaboration patterns or benchmarking alternative network growth models.
The final network-type assignment is not determined by these parameters alone but results from the integrated multi-criteria protocol described in
Section 2 and summarized in
Table 3.
3.1. Supported BA-like/Scale-Free Small-World Clusters
According to the integrated classification protocol (
Table 3 and
Table 4), clusters 000, 002, 008, 009, 013, and 016 received supported BA-like/scale-free small-world classifications. These clusters consistently combine statistically supported heavy-tailed weighted degree distributions, pronounced small-world organization (
), and the absence of fractal scaling according to the box-covering analysis. Clusters 22, 23, and 24 were excluded from the classification because they contain too few nodes to provide reliable statistical estimates (
Table 1).
Figure 3 illustrates the topology of the supported BA-like/scale-free small-world clusters. The network layouts were generated using the OpenOrd algorithm [
21], which provided the clearest visualization of the internal modular organization and hub structure of each cluster. The adopted layout parameters were: liquid stage 25%, expansion stage 25%, cooldown stage 25%, crunch stage 10%, and simmer stage 15%; edge cut equal to 0.8; 19 computation threads; 2000 iterations; fixed time step of 0.2; and random seed—2527419387533434. This configuration emphasizes community boundaries while preserving the spatial separation of densely connected regions.
The estimated power-law exponents for clusters 000, 002, 008, 009, 013, and 016 are approximately 3.64, 2.88, 1.60, 1.75, 1.67, and 1.53, respectively. Together with the bootstrap KS test, the AIC-based tail-model selection, and the small-world diagnostics, these results support their classification as BA-like/scale-free small-world networks according to the integrated decision protocol.
The topologies shown in
Figure 3 reveal heterogeneous hub-dominated structures in which a relatively small number of highly connected authors coexist with many peripheral collaborators. Such an organization is compatible with heavy-tailed weighted degree distributions and decentralized collaboration patterns commonly observed in empirical scientific networks. Cluster 002 provides the strongest overall evidence for BA-like organization, combining a well-supported heavy-tailed distribution (
), pronounced hub dominance, and temporal growth behavior that is closest to the canonical Barabási–Albert scaling inferred from the
analysis.
It is important to emphasize that the coexistence of scale-free degree distributions and small-world organization is not unique to Barabási–Albert networks. A well-known example is the worldwide air transportation network, which simultaneously exhibits heavy-tailed connectivity, short average path lengths, high clustering, and strong modular organization [
22]. In such systems, highly connected hubs are not necessarily the most central vertices with respect to shortest-path betweenness because community boundaries constrain information flow. Consequently, preferential attachment and local clustering may coexist within a modular architecture, producing hybrid scale-free small-world regimes.
Accordingly, the term BA-like is employed here as a phenomenological topological descriptor rather than as evidence of a specific generative mechanism. The proposed framework identifies structural compatibility with the characteristic organization of Barabási–Albert-type networks based on the combined statistical and structural diagnostics, but it does not establish that the underlying collaboration dynamics necessarily follow the Barabási–Albert growth process. Alternative mechanisms, including finite-size effects, modular organization, or node fitness heterogeneity, may produce similar topological signatures.
Although the supported BA-like/scale-free small-world clusters share the same qualitative topological regime, they exhibit substantial variability in their scaling exponents, ranging from approximately to . Such variability reflects different degrees of hub dominance and degree heterogeneity while remaining consistent with the integrated classification protocol. Among the supported clusters, cluster 000 and cluster 002 exhibit the strongest overall agreement with the topological characteristics expected for BA-like organization, considering the combined statistical and structural diagnostics employed by the proposed framework.
The observed variability among the supported BA-like clusters suggests that preferential-attachment-like organization alone may not fully explain the structural diversity of the collaboration network. Similar deviations from idealized Barabási–Albert scaling have been widely reported in empirical complex networks and are frequently attributed to additional mechanisms such as finite-size effects, modular organization, or node fitness heterogeneity. These mechanisms may preserve the characteristic coexistence of heterogeneous connectivity and small-world organization while producing measurable differences in the degree distribution and temporal evolution of individual clusters.
This behavior has been extensively discussed in the network science literature. Bianconi and Barabási [
23] proposed a fitness model in which nodes possess intrinsic abilities to compete for new links, leading to heterogeneous growth rates and, in some cases, a condensation phenomenon where a small number of highly fit nodes dominate the connectivity structure. Likewise, empirical studies of scientific collaboration networks [
1,
24] have reported mixed scaling regimes in which heavy-tailed degree distributions coexist with truncated power-law or log-normal behavior as a consequence of finite-size effects, temporal constraints, and heterogeneous author productivity.
To further investigate these mechanisms, we analyze the temporal growth exponent
, originally introduced in the context of preferential attachment and node fitness [
20,
23]. Although the proposed classification framework is based exclusively on structural diagnostics, the
exponent provides an independent temporal perspective on the growth dynamics of the supported BA-like clusters and helps assess the extent to which their evolution is compatible with preferential-attachment-like behavior. The distribution of
across the classified clusters is presented in
Figure 4.
The distribution of the temporal growth exponent
for the supported BA-like/scale-free small-world clusters is shown in
Figure 4. Overall, the median
values are close to the canonical Barabási–Albert prediction of
, although noticeable variability is observed among individual clusters. This result indicates that the supported BA-like clusters exhibit temporal growth broadly compatible with preferential-attachment-like dynamics while preserving measurable differences in their rates of connectivity accumulation.
Among the supported clusters, cluster 002 exhibits the median value closest to the theoretical Barabási–Albert prediction, in agreement with the overall structural evidence discussed previously. The remaining clusters display moderate deviations from the theoretical reference while remaining within the range commonly observed for empirical collaboration networks. These differences reinforce the view that subnetworks belonging to the same mesoscopic structural regime may nevertheless evolve through quantitatively distinct temporal dynamics.
Such variability has been widely reported in empirical collaboration networks. Bianconi and Barabási [
23] proposed that node fitness may alter the probability of acquiring new links, producing growth patterns that depart from the idealized Barabási–Albert model. Likewise, empirical studies [
1,
24] have shown that finite-size effects, heterogeneous author productivity, temporal constraints, and modular organization may all contribute to deviations from the canonical preferential-attachment prediction while preserving heavy-tailed degree distributions and small-world organization.
The temporal analysis presented here should therefore be interpreted as complementary to the structural classification proposed in this work. While the network typology is determined exclusively from the integrated structural diagnostics described in
Section 2, the exponent
provides an independent temporal characterization of the supported BA-like clusters. Consequently, the observed deviations from the theoretical value
should not be interpreted as definitive evidence of preferential attachment, node fitness, or any specific network generative mechanism. Rather, they provide additional evidence that structurally similar collaboration communities may exhibit different temporal growth patterns, highlighting the intrinsic heterogeneity of large-scale scientific collaboration networks.
3.2. Ambiguous BA-like/Scale-Free Small-World Cases
The integrated classification protocol identified four clusters (003, 004, 014, and 015) whose statistical and structural diagnostics did not provide sufficiently consistent evidence for a definitive BA-like/scale-free small-world assignment. Although these clusters exhibited weighted degree distributions whose best-supported tail model was the power-law model, the combined evidence from the bootstrap KS test, model selection, and structural diagnostics was insufficient to satisfy all criteria required for a supported classification.
The corresponding network topologies are shown in
Figure 5. As in the previous section, the layouts were generated using the OpenOrd algorithm with identical visualization parameters to facilitate direct comparison with the supported BA-like clusters.
Clusters 003 and 004 illustrate situations in which the power-law model provides the best fit according to the AIC, yet the evidence remains statistically weak because the competing models exhibit comparable support (AIC ). Consequently, although their weighted degree distributions are compatible with heavy-tailed behavior, they cannot be distinguished robustly from alternative tail models based solely on the available statistical evidence. Their classification therefore remains ambiguous within the proposed decision protocol.
A similar situation is observed for clusters 014 and 015. While the power-law model also minimizes the AIC for these clusters, the statistical support is considerably weaker owing to both the very small
AIC values and the lack of consistent agreement among the complementary structural diagnostics. In addition, these clusters contain relatively small degree tails and unusually large estimated scaling exponents (
Table 4), indicating that the inferred power-law behavior is unstable and should be interpreted with caution. Under these conditions, small variations in the empirical tail may substantially alter the preferred statistical model.
For these reasons, the present framework intentionally avoids assigning definitive topological labels whenever the available evidence is inconclusive. Rather than forcing every cluster into a predefined category, the integrated protocol explicitly preserves ambiguous cases whenever the statistical or structural diagnostics fail to converge toward a unique interpretation. This conservative strategy reduces the risk of overinterpreting finite-size effects or statistical fluctuations as genuine topological regimes.
The existence of ambiguous clusters also illustrates an important characteristic of empirical collaboration networks. Real-world systems frequently exhibit transitional or mixed structural properties that cannot be adequately represented by idealized network models. Consequently, preserving an explicit ambiguous category constitutes an important feature of the proposed framework, allowing uncertain cases to be identified transparently instead of being artificially assigned to a specific topological regime.
3.3. Ambiguous Small-World (Non-Scale-Free) Cases
The integrated classification protocol identified clusters 001, 006, 011, 012, 017, 018, 019, 020, and 021 as cases that are structurally compatible with the small-world (non-scale-free) regime but whose statistical evidence is insufficient to support a definitive classification.
For all these clusters, the weighted degree distribution is better described by exponential or log-normal tails than by a statistically supported power-law model. However, some cases present limited statistical support for the preferred tail model or conflicting evidence among the complementary structural diagnostics. Consequently, according to the proposed decision protocol, these clusters are conservatively classified as ambiguous rather than being assigned a definitive small-world label.
Figure 6 illustrates the topology of these ambiguous small-world clusters. As in the previous sections, the layouts were generated using the Fruchterman–Reingold algorithm, which provides a clear visualization of the local connectivity patterns in relatively compact graphs.
Despite the absence of definitive statistical support, these clusters consistently exhibit the structural characteristics commonly associated with small-world organization. All of them display high clustering coefficients together with relatively short average path lengths, resulting in , which indicates efficient local cohesion combined with global reachability.
The ambiguous small-world clusters nevertheless exhibit considerable structural diversity. Cluster 001 represents a large and sparse collaboration community with relatively long average path lengths, whereas clusters such as 012, 019, 020, and 021 are considerably smaller and denser, exhibiting clustering coefficients above 0.9 and average shortest-path lengths below three (
Table 1). This diversity indicates that compact institutional groups and larger thematic collaboration communities may both display small-world organization despite exhibiting different connectivity patterns and degree-tail statistics.
Nevertheless, the integrated classification protocol intentionally avoids assigning definitive network types whenever the statistical evidence from degree-tail modeling remains inconclusive. Consequently, these clusters are interpreted as ambiguous small-world cases rather than confirmed representatives of a unique topological regime.
3.4. Supported Scale-Free Fractal (Non-BA) Clusters
According to the integrated classification protocol, clusters 005, 007, and 010 were consistently classified as supported scale-free fractal (non-BA) networks. These clusters simultaneously exhibit statistically supported heavy-tailed weighted degree distributions, fractal scaling supported by the box-covering analysis, and mutually consistent statistical and structural diagnostics according to the integrated classification protocol.
Figure 7 presents the topology of the supported fractal clusters. Identical visualization parameters were adopted for all layouts to facilitate structural comparison among the classified subnetworks. The layouts are presented solely to illustrate the qualitative structural organization of the clusters and were not used as classification criteria within the integrated classification protocol.
The three clusters exhibit the structural characteristics typically associated with fractal complex networks. Rather than being organized around one or a few dominant hubs that substantially shorten global distances, they display recursively nested modular structures connected through relatively sparse inter-module bridges. Such organization is consistent with the self-similar topology expected under the renormalization framework introduced by Song, Havlin, and Makse [
6] and further developed by Rozenfeld et al. [
7], in which the number of covering boxes decreases according to the power-law model with increasing box diameter, resulting in a finite fractal dimension. Similar approaches have been widely adopted in subsequent studies of fractal complex networks [
25,
26].
This structural organization differs substantially from the supported BA-like/scale-free small-world clusters presented previously. Although both topological regimes exhibit statistically supported heavy-tailed weighted degree distributions, the fractal clusters preserve modular separation across multiple scales instead of concentrating connectivity into a small number of globally dominant hubs. Consequently, communication and collaboration tend to occur through hierarchically organized modules connected by relatively few bridging authors, rather than through highly centralized hub-mediated interactions.
The estimated fractal dimensions (
,
, and
for clusters 005, 007, and 010, respectively;
Table 3) provide quantitative support for this interpretation. In all three cases, the box-covering analysis favored the power-law scaling over the exponential alternative, indicating that the observed modular organization remains statistically compatible with self-similar scaling across the analyzed range of box diameters. These results provide an independent structural diagnostic that complements the degree-tail modeling, bootstrap KS goodness-of-fit test, AIC model selection, and small-world analysis incorporated into the integrated classification framework.
From the perspective of scientific collaboration, this topology suggests communities organized around relatively autonomous research groups connected through a limited number of bridging collaborations. Such an organization combines strong local cohesion with controlled communication between modules, potentially increasing robustness against localized disruptions while preserving thematic specialization. Analogous self-similar modular architectures have been reported in empirical biological, technological, and collaboration networks, where fractal organization emerges from the coexistence of heterogeneous connectivity and hierarchical modularity rather than from purely hub-dominated growth processes [
7,
25,
26,
27].
Unlike the BA-like and small-world cases discussed previously, no ambiguous fractal clusters were identified by the proposed decision protocol. All subnetworks exhibiting supported fractal scaling also presented mutually consistent statistical and structural diagnostics across the complementary analyses employed in this work, resulting in supported classifications. This outcome indicates that, within the Embrapa coauthorship network, the fractal regime constitutes a well-defined mesoscopic organization that is consistently identified by the integrated decision protocol. However, it is important to note that the higher thresholds adopted in this protocol may have reduced the number of networks identified as fractal, as described in the
Supplementary Materials.
4. Discussion
Unless otherwise stated, the comparative analyses presented in this section summarize the structural properties of the network families identified by the integrated classification protocol. Each family includes both supported and ambiguous classifications, allowing the distributions to capture the overall structural variability associated with each topological regime. Consequently, the following boxplots are intended as descriptive summaries of the identified network families and do not constitute additional evidence supporting the classification itself.
4.1. Assortativity Patterns
Figure 8 summarizes the distribution of degree assortativity (
r) across the identified network families. The scale-free fractal family consistently exhibits positive assortativity, whereas the BA-like/scale-free small-world family displays lower and more variable values. The small-world (non-scale-free) family spans the widest assortativity range, including near-zero and slightly negative values. These observations suggest that the degree to which highly connected researchers preferentially collaborate with one another varies substantially among the identified network families and depends on the balance between hub formation, modular organization, and local collaboration patterns, in agreement with previous studies of scientific collaboration networks [
5].
4.2. Entropy Measures for Degree–Degree Mixing and Their Relation to Assortativity
Let
r denote degree assortativity (the Pearson correlation between the degrees at the two ends of edges) [
5,
28]. A typical measure in networks is the Shannon entropy of the degree distribution
However,
depends only on the marginal
and is invariant under any degree-preserving rewiring of the graph. Since
r can vary widely under degree-preserving rewirings, there is no generic monotone relationship between
(or its normalized variant) and
r. Thus,
is not appropriate for studying mixing patterns [
28,
29].
To capture the randomness or structure that generates assortativity, we work with the joint distribution of degrees at the two ends of edges. Let
for an undirected network (hence
) [
28]. Define the edge-end marginal
From
and
we build the following entropic quantities [
30]:
If the network is degree–degree independent (no mixing), then
, so
and
while
[
28,
30]. As the strength of degree–degree dependence increases (assortative or disassortative),
I increases and
decreases (neighbor degrees become more predictable); the
sign is given by
r [
30]. For comparability across clusters with different
q, it is useful to report a normalized measure such as
Although
is not the focus of the present analysis, it is defined here for completeness and may facilitate cross-study comparisons.
For each cluster we computed
(and its normalized version),
,
,
, and
, and merged these quantities with the network-family assignment obtained from the integrated classification protocol. Unless otherwise stated, the analyses presented below include both supported and ambiguous classifications within each network family. The panels in
Figure 9 summarize the key relationships.
Figure 9 presents two complementary two-dimensional projections of the correlation–information space. Panel (a) displays the relationship between the assortativity magnitude
and the degree–degree mutual information
I, whereas panel (b) shows assortativity
r as a function of the conditional entropy of degree correlations (
). In both panels, point colors identify the network family assigned by the integrated classification protocol. As discussed previously, each family includes both supported and ambiguous classifications and is presented here as a descriptive summary of the structural variability within each topological regime.
The two panels highlight complementary aspects of degree–degree organization. Because mutual information quantifies the statistical dependence between the degrees at the two ends of an edge irrespective of whether the correlation is assortative or disassortative, whereas the assortativity coefficient
r measures both the direction and strength of the linear degree correlation, no generic monotonic relationship between
I and
is expected [
28,
30]. This behavior is reflected in panel (a), where clusters exhibiting similar assortativity magnitudes may nevertheless present substantially different mutual-information values. Such variability indicates that networks with comparable Pearson correlations may possess distinct joint degree distributions and, consequently, different levels of statistical dependence between neighboring vertices. Mutual information therefore captures structural information that is not fully summarized by the assortativity coefficient alone [
30].
Panel (b) reveals a more complex relationship between assortativity and conditional entropy. While highly assortative clusters typically exhibit low conditional entropy values, clusters with similar assortativity coefficients may differ substantially in their conditional entropy. For instance, clusters with assortativity around 0.5 present conditional entropies ranging from approximately 0.68 (cluster 15) to 2.08 (cluster 11). This dispersion indicates that even when linear correlations are comparable, the predictability of neighboring degrees can vary considerably, reflecting differences in the underlying joint degree distributions that are not captured by the assortativity coefficient alone. Such behavior is consistent with the information-theoretic identity
which shows that the conditional entropy depends on both the mutual information and the marginal entropy of the edge-end degree distribution. As expected,
I is bounded by the entropy of the edge-end degree distribution
, with
[
30]. Consequently, networks characterized by narrower edge-end degree distributions tend to attain smaller mutual-information values, whereas normalized measures such as
facilitate comparisons among networks with different degrees of heterogeneity.
The information-theoretic descriptors provide additional context for interpreting the network families identified by the integrated classification protocol. Considerable overlap is observed among the families in both projections, indicating that neither mutual information nor conditional entropy should be regarded as independent discriminative features. The scale-free fractal clusters generally occupy an intermediate region of the information space, with mutual-information values ranging from approximately 1.56 to 2.33, consistent with the hierarchically modular organization expected for self-similar networks [
6,
27,
31,
32]. In contrast, BA-like/scale-free small-world and small-world (non-scale-free) clusters occupy a substantially broader portion of the correlation–information space, spanning both low- and high-information regimes. This dispersion indicates that similar topological classifications may still exhibit markedly different degree–degree mixing patterns, emphasizing that these information-theoretic quantities describe structural organization rather than the underlying network generation mechanism [
4,
5,
20,
28,
33,
34].
Overall, the two-dimensional projections reinforce that assortativity, mutual information, and conditional entropy provide complementary descriptions of degree–degree mixing. While assortativity summarizes the direction and linear strength of degree correlations, mutual information captures the overall statistical dependence encoded in the joint degree distribution, and conditional entropy quantifies the predictability of neighboring degrees. Rather than defining sharply separated regions for the identified topological families, these descriptors provide additional structural information that complements the statistical and topological diagnostics employed throughout this work. This interpretation is consistent with previous studies showing that different mesoscale organizations and scaling regimes often occupy partially overlapping regions of the structural phase space of complex networks, while remaining distinguishable through the combined use of multiple complementary descriptors [
7,
35].
4.3. Clustering and Path Length
Figure 10 and
Figure 11 summarize the distributions of the average clustering coefficient (
C) and average shortest-path length (
L) across the network families identified by the integrated classification protocol. Because each family includes both supported and ambiguous classifications, these comparisons are intended to illustrate overall structural tendencies rather than establish strict boundaries between topological regimes.
Average clustering coefficients remain high for all network families (
Figure 10), indicating that strong local cohesiveness is a common feature of the Embrapa coauthorship clusters. This observation is consistent with the prevalence of small-world characteristics in scientific collaboration networks, where dense local neighborhoods coexist with relatively short global distances [
4].
Differences become more apparent when considering the average path length (
Figure 11). Scale-free fractal families generally occupy the upper portion of the observed range, reflecting their more modular and hierarchically organized topology. In contrast, BA-like/scale-free small-world families tend to exhibit shorter path lengths, consistent with the presence of highly connected hubs that facilitate long-range communication across the network [
6,
7]. Small-world (non-scale-free) families generally present the shortest paths, although some overlap with the other families remains, reflecting the presence of ambiguous classifications within the integrated framework.
4.4. Betweenness Centrality Tails
Figure 12 presents the CCDF of betweenness centrality for the three network families identified by the integrated classification protocol. Since each family includes both supported and ambiguous classifications, the median curves and interquartile ranges summarize the overall distribution of brokerage patterns within each family rather than defining strict boundaries between topological regimes.
All three families exhibit heavy-tailed behavior, confirming the existence of a relatively small number of highly central nodes. Nevertheless, differences in the extreme tail remain apparent. The BA-like/scale-free small-world family (blue) tends to exhibit a somewhat sharper decay, whereas the scale-free fractal (green) and small-world (non-scale-free) (orange) families generally maintain higher survival probabilities at large betweenness values. These trends suggest that brokerage roles tend to be distributed more broadly among intermediary nodes in the latter families, whereas the BA-like/scale-free small-world family generally concentrates global connectivity in a smaller number of highly influential hubs, consistent with preferential attachment and hub-mediated connectivity [
5,
6,
7].
This pattern closely parallels observations reported for the worldwide air transportation network, which has also been shown to exhibit scale-free small-world structure while displaying nontrivial centrality organization due to its modular architecture [
22]. In that system, highly connected hubs are not necessarily the nodes with the highest betweenness centrality, because shortest paths are often confined within geographically or geopolitically defined communities. As a result, centrality becomes distributed across multiple regional connectors rather than concentrated exclusively in the largest hubs.
A similar mechanism may operate in the present coauthorship network. The BA-like/scale-free small-world family tends to exhibit a sharper decay in the betweenness CCDF, suggesting that global connectivity is generally concentrated in a relatively small number of hub authors, consistent with preferential-attachment-like dynamics. In contrast, the heavier tails observed for the scale-free fractal and small-world (non-scale-free) families indicate a more distributed intermediary structure, in which brokerage roles are shared among multiple authors. Such behavior is consistent with stronger modular organization and heterogeneous connectivity constraints, which redistribute shortest-path flows across several intermediary nodes rather than concentrating them in a single dominant hub. Taken together, these observations reinforce the interpretation that the hybrid scale-free small-world organization identified in this work may emerge from the interplay between preferential attachment and modular organization, rather than from a purely homogeneous Barabási–Albert generative process.
4.5. Distribution of Network Families
The integrated classification protocol partitions the 25 clusters extracted from the giant component into three principal network families–BA-like/scale-free small-world, small-world (non-scale-free) and scale-free fractal (non-BA)–together with a small number of undefined cases. The classification combines complementary evidence from weighted degree-tail modeling, small-world diagnostics (), and box-covering analysis, thereby integrating statistical and structural information rather than relying on a single topological descriptor. The spatial distribution of these network families within the giant component reveals a clear predominance of the BA-like/scale-free small-world family (63.4% of nodes when combining supported and ambiguous classifications), followed by the small-world (non-scale-free) family (27.5%), and the scale-free fractal (non-BA) family (9.0%). Undefined clusters account for only 0.08% of the giant component.
The BA-like/scale-free small-world family dominates the giant component, whereas small-world (non-scale-free) and scale-free fractal families occupy smaller but well-defined portions of the collaboration network. This predominance is consistent with previous studies reporting that mature coauthorship networks frequently develop scale-free small-world characteristics associated with cumulative growth and preferential-attachment-like dynamics [
2]. At the same time, the coexistence of multiple network families highlights the structural heterogeneity of the collaboration system, indicating that distinct mesoscopic organizations emerge within the same connected infrastructure.
This heterogeneity is further reflected in the weighted degree-tail models associated with each network family (
Table 5). Power-law tails are exclusively observed among the BA-like/scale-free small-world and scale-free fractal families, whereas the small-world (non-scale-free) family is consistently associated with exponential or log-normal tails. These results reinforce that the identified network families correspond to distinct structural organizations rather than arbitrary subdivisions of the giant component, suggesting different underlying mechanisms governing collaboration patterns across the network.
The distribution of network families reveals that the BA-like/scale-free small-world family accounts for the majority of the giant component, suggesting that cumulative-advantage-like dynamics remain an important organizing mechanism in the collaboration system. The smaller yet well-defined presence of scale-free fractal and small-world (non-scale-free) families, however, indicates that alternative structural organizations coexist within the same connected infrastructure, reflecting distinct collaboration patterns and connectivity constraints across different research communities.
5. Conclusions
This study introduces a multi-criteria methodological framework for identifying network families in large-scale collaboration networks. By integrating community detection, weighted degree-tail modeling, small-world diagnostics, fractal box-covering analysis, and entropy-based measures of degree–degree correlations, we propose a reproducible integrated classification protocol capable of distinguishing heterogeneous mesoscopic organizations within a single giant connected component.
The empirical application to Embrapa’s scientific coauthorship network (1974–2024) demonstrates the applicability of the proposed framework to a large real-world collaboration system. The identified network families, BA-like/scale-free small-world, scale-free fractal (non-BA), and small-world (non-scale-free), illustrate how complementary statistical and structural diagnostics can be combined into a unified decision protocol rather than interpreted independently. By jointly considering weighted degree-tail selection, small-world diagnostics, fractal scaling, and information-entropy measures, the proposed methodology provides a more comprehensive characterization of mesoscopic network organization than approaches based on individual descriptors.
At the system level, the predominance of the BA-like/scale-free small-world family suggests that cumulative-advantage-like dynamics remain an important organizing mechanism in large collaboration systems. At the same time, the coexistence of scale-free fractal and small-world (non-scale-free) families within the same giant component highlights the intrinsic structural heterogeneity that single-metric analyses often overlook. In this sense, the principal contribution of this work lies less in the specific proportions observed in the Embrapa network than in demonstrating that distinct network families can be systematically identified through an integrated methodological protocol.
The incorporation of entropy-based measures further enriches the characterization by revealing complementary aspects of degree–degree organization that are not captured by degree distributions or small-world metrics alone. Together with the structural diagnostics, these information-entropy measures provide an additional layer for interpreting how hierarchical organization, modularity, and degree correlations vary across the identified network families.
Beyond the empirical case examined here, the proposed framework offers a transferable analytical methodology for investigating collaboration, organizational, and production networks. Its principal strength lies in combining statistically grounded model selection with complementary structural diagnostics within a coherent decision framework, thereby improving reproducibility, interpretability, and cross-domain comparability. Consequently, the methodology is applicable not only to scientific collaboration networks but also to broader sociotechnical systems in which cooperation, information diffusion, resilience, and coordination emerge from centralized or decentralized interaction patterns.
6. Future Research Directions
The framework proposed in this work establishes a reproducible and interpretable methodology for the comparative characterization of mesoscopic structural regimes by integrating complementary statistical and structural diagnostics within a unified decision protocol. By combining multiple independent sources of evidence rather than relying on any single network descriptor, the proposed approach provides a robust foundation for identifying heterogeneous network families while remaining sufficiently flexible to accommodate future methodological refinements and applications. Building upon this foundation, several methodological and application-oriented research directions remain open.
A natural methodological extension consists of broadening both the validation benchmark and the statistical reference models adopted by the framework. Future studies may incorporate additional classes of synthetic networks, including configuration-model networks, stochastic block models, degree-corrected community models, random geometric networks, temporal network generators, and multiplex or multilayer structures. In particular, replacing or complementing the Erdős–Rényi baseline adopted in the present work with degree-preserving configuration-model nulls would provide a more stringent assessment of small-world organization in heterogeneous networks. Such analyses would enable the computation of ensemble-based confidence intervals and z-scores for clustering, path length, and small-worldness, thereby strengthening the statistical foundation of the proposed decision protocol while preserving its interpretability and generality.
Another important direction concerns the incorporation of temporal network analysis and percolation-based order parameters to investigate the emergence, persistence, and transition of network families over time. Extending the current static framework to temporal collaboration networks may reveal critical transitions in mesoscopic organization and improve our understanding of how cumulative growth, institutional changes, and collaboration dynamics influence structural evolution. Coupling this approach with Hamiltonian formulations or energy-landscape approaches may further provide a physics-inspired interpretation of the stability and transitions between different network families.
A particularly promising avenue involves enriching the structural analysis with semantic information extracted from publication metadata. Topic embeddings, domain clustering, and semantic representations assisted by large language models could be integrated with the structural diagnostics developed here to investigate whether specific research domains are preferentially associated with particular network families. Such integration would enable a combined structural-semantic characterization capable of identifying epistemic communities and collaboration patterns that are not apparent from network topology alone.
The transferability of the proposed framework to non-academic systems also represents an important research opportunity. Applying the methodology to production systems, agricultural value chains, agricultural cooperatives, innovation ecosystems, technological districts, and regional development networks may reveal whether the same network families identified in scientific collaboration also emerge in other sociotechnical systems. Comparative analyses across these domains could evaluate whether mesoscopic structural organization exhibits common patterns and whether the identified network families possess predictive value for resilience, innovation capacity, diffusion processes, governance, and organizational performance.
Another particularly relevant direction concerns the role of modular organization in shaping centrality patterns and producing deviations from idealized generative models such as the Barabási–Albert framework. Future studies may investigate, at both macroscopic and microscopic scales, how modularity influences the distribution of betweenness centrality and the emergence of brokerage roles. At the macroscopic level, this includes quantifying the relationship between global modularity and deviations from canonical BA scaling relationships. At the microscopic level, it involves determining whether community boundaries systematically prevent highly connected nodes from becoming dominant brokers. Such analyses may help explain why networks exhibiting similar weighted degree distributions can nevertheless display substantially different patterns of centrality and information flow, providing additional insight into the structural mechanisms underlying the distinct network families identified in this work.
Finally, the current framework may be expanded by incorporating additional structural descriptors, including rich-club organization, core–periphery structure, multilayer connectivity, weighted-edge heterogeneity, motif statistics, higher-order interactions, and persistent topological descriptors. Integrating these complementary diagnostics within the existing decision protocol may further improve classification resolution while preserving the interpretability, reproducibility, and modular design of the methodology.