2.1. Data Sources and Processing
This paper selects the list of award-winning projects of the Jiangsu Provincial Science and Technology Progress Award from 2006 to 2023, issued by the Jiangsu Provincial People’s Government, as the research object to explore the evolutionary process of innovation team collaboration networks led by leading enterprises. The reasons for choosing this sample are threefold. First, the Science and Technology Progress Award projects fully demonstrate the important characteristics of enterprise-led innovation teams, which are reflected in the fact that industry-leading enterprises take the lead in collaborating with universities, research institutions, or other enterprises to jointly advance these projects. Moreover, most of these projects originate from important tasks commissioned by national or provincial governments, focusing on breaking through critical bottlenecks in core technologies within industries [
31]. Second, the award-winning projects of the Science and Technology Progress Award are representative and authoritative. The Science and Technology Progress Award is an important award established by the Jiangsu Provincial People’s Government to recognize units and individuals that have made outstanding contributions in the field of science and technology. The selection process is rigorous and fair, ensuring the authority and credibility of the award. Furthermore, the Science and Technology Progress Award includes both applied scientific and technological achievements and basic research achievements, with award-winning projects covering multiple industries and fields, demonstrating high representativeness and broad coverage. Finally, the list of award-winning projects of the Jiangsu Provincial Science and Technology Progress Award is open and transparent, with reliable data sources, providing convenience for constructing the innovation collaboration network. Additionally, from 2006 to 2023, the selected sample spans a long period, effectively reflecting the historical changes and development trends of scientific and technological innovation collaboration in Jiangsu Province.
The award-winning projects of the Jiangsu Provincial Science and Technology Award are sourced from policy documents published on the official website of the Jiangsu Provincial People’s Government (
https://www.jiangsu.gov.cn/col/col84242/index.html (accessed on 8 July 2026)). A total of 4119 science and technology awards were conferred during the research period from 1 Lanuary 2006 to 31 December 2023. We preprocessed information including award grade, project title, completing entities and contributors based on the list of award winners, with detailed procedures as follows:
First, screen collaborative projects to construct the basic sample set. Only joint award projects involving two or more completing entities are retained, while independently completed projects undertaken by a single entity are eliminated directly, as they entail no cross-agent collaborative behaviors. This step identifies samples characterized by collaborative innovation linkages.
Second, code the types of completing entities. Drawing on the classification criteria proposed by Han Zenglin et al. [
41], four categories of innovation agents are distinguished via keyword matching of entity names: entities containing “university”, “college” or “school” are classified as universities; those with “research institute”, “research center”, “laboratory” or “hospital” are categorized as scientific research institutions; entities featuring “company”, “group”, “enterprise”, “mine” or “factory” are defined as firms; administrative entities with governance functions marked by keywords such as “government”, “department”, “bureau”, “commission”, “departmental office” are grouped into the government category (the Jiangsu Popular Science Writers Association is also incorporated into this category).
Third, define enterprise-led collaborative projects and verify the compliance of this operational definition. As stipulated in the Implementation Rules for the Measures of Jiangsu Provincial Science and Technology Awards, completing entities of provincial science and technology award projects are ranked in descending order of their contributions. The first completing entity provides capital, venues, overall technical schemes and cross-party resource coordination throughout the whole project, acting as the core leading agent in coordinating R&D and commercialization of research outcomes. In accordance with this official ranking rule, multi-entity collaborative projects where the first completing entity is an enterprise are defined as “enterprise-led” innovation projects. A total of 1156 such projects are screened out as analytical samples for the innovation collaboration network led by leading enterprises. This definition differentiates ordinary participating enterprises that merely occupy a high ranking from leading enterprises in charge of overall coordination, and validates the dominant position of enterprises in R&D collaboration from the perspective of official evaluation regulations.
Fourth, encode spatial information. Standard national administrative division codes are matched to identify the prefecture-level city affiliation of all completing entities, laying a foundation for analyzing the spatial evolution and regional heterogeneity of enterprise-led innovation collaboration networks.
2.2. Research Methods
Social Network Analysis (SNA) is a quantitative analytical tool for unpacking social ties and network structures. It integrates qualitative and quantitative methods to explore linkages between individuals and organizations as well as their underlying mechanisms, and interprets the regularities of complex interactive behaviors. At present, SNA is widely adopted in sociology, economics, innovation management, network science and other disciplines. This paper constructs an undirected innovation collaboration network based on award-winning projects of the Jiangsu Provincial Science and Technology Award from 2006 to 2023. The detailed operational procedures for network construction are listed as follows. First, disambiguation of institutional names: Standardize and unify the names of entities including enterprises, branches of state-owned enterprises, university-affiliated research institutes, university hospitals, various research centers and corporate subsidiaries, so as to eliminate identification biases caused by identical names for distinct entities or different names for the same entity. Second, classification of entity types: All nodes are categorized into four groups—universities, scientific research institutions, enterprises and government bodies—according to keyword matching of institutional names. Third, regional coding: Standard administrative division codes are matched to label the prefecture-level city where each institution is located. Fourth, rules for edge construction: If a single project involves multiple completing entities, collaborative ties are established pairwise among all participants within the project to form a complete graph structure. Fifth, edge weight assignment: Collaborative ties generated from different projects are assigned equal weights, and the number of edges is not adjusted according to project scale. Sixth, handling of repeated collaborations: If the same pair of entities win awards jointly on multiple occasions, multiple edges are retained to reflect the frequency of sustained collaboration between them.
After network construction, this paper calculates network characteristics from multi-level dimensions. Low-order indicators including network size, total number of edges, network density, network diameter, traditional clustering coefficient, average path length and average degree are selected to depict the overall macroscopic features of the network. Degree centrality, closeness centrality and betweenness centrality are adopted to identify the individual network position of each innovation entity. Furthermore, two high-order topological indicators, namely high-order clustering coefficient and high-order distance distribution, are introduced to deeply explore the dynamic evolutionary patterns of core network entities.
- (1)
Overall Characteristic Indicators
Network size refers to the number of all nodes in the network. Number of network edges refers to the number of all edges in the network. Network density is the ratio of the actual number of connections in the overall network to the maximum possible number of connections, reflecting the closeness of connections among nodes in the network. The higher the network density, the closer the connections between nodes. Its calculation formula is [
42]:
where M represents the number of actual connections, and N represents the number of network nodes. Obviously,
.
Network diameter is the maximum distance between any two nodes in the network, reflecting the maximum communication distance between any innovation agents in the network. Its calculation formula is:
where
is the distance between node
.
The clustering coefficient is the arithmetic mean of the clustering coefficients of all nodes in the network, characterizing the local connectivity and degree of clustering of the cooperation network. A higher clustering coefficient indicates the presence of more tightly connected clusters of nodes within the network. Its calculation formula is:
where
is the clustering coefficient of node
and
, where
is the number of neighbor nodes of node
, and
is the number of edges that actually exist among these
nodes.
Average path length is the average distance between any two nodes in the network, characterizing the average degree of separation among nodes in the network, i.e., how small the network is. A smaller average path length indicates that nodes in the network can reach each other more easily, reflecting the “small-world effect.” Its calculation formula is:
Average degree is the average of the degrees of all nodes in the network, characterizing the overall connectivity of the network. The larger the average degree, the better the connectivity. Its calculation formula is:
- (2)
Individual Characteristic Indicators
Individual characteristic indicators include degree centrality, closeness centrality, and betweenness centrality. Degree centrality measures the number of other nodes directly connected to a given node and is the most direct metric for characterizing node centrality in network analysis. In an innovation collaboration network, it reflects the central position of each innovation agent; a higher degree indicates more frequent interactions between the innovation agent and other agents, as well as more stable collaborative relationships. For a weighted undirected graph, its calculation formula is:
where
is the weight of the edge connected to node
.
Closeness centrality is measured by the reciprocal of the sum of the shortest paths between a node and all other nodes in the network, characterizing the extent to which a node is centrally located within the network. Nodes with high closeness centrality can reach other nodes in the network more quickly and are suitable for disseminating information within the network. Its calculation formula is:
Betweenness centrality is the number of shortest paths that pass through a given node in a network, reflecting the extent to which the node serves as a bridge along the shortest paths between other nodes. Nodes with high betweenness centrality frequently appear on shortest paths, controlling the flow of information and acting as “intermediaries” in the network. Its calculation formula is:
where
represents the number of shortest paths between node
and node
, and
represents the number of shortest paths among all shortest paths between node
and node
that pass through node
.
- (3)
Higher-Order Indicators
Higher-order statistical indicators in networks have diverse applications, as they can more accurately quantify and characterize the structural features of various complex systems, thereby providing deeper insights into their underlying evolutionary mechanisms [
43]. The higher-order topological indicators selected in this paper are the higher-order clustering coefficient and distance distribution.
The higher-order clustering coefficient involves more complex node associations and clustering patterns within a network, reflecting the closeness of connections among nodes at a higher level, and facilitates the analysis of both the overall organizational structure and local structural characteristics of the network. The x-order clustering coefficient of node
is defined as the proportion of node pairs at distance x among its neighbors after removing node
. Its calculation formula is [
44]:
where
denotes the number of neighbor pairs of
that are separated by a distance of
after deleting node
and its corresponding edges from the network;
represents the number of neighbor nodes of node
. The clustering coefficient distribution of node
. Since node
is eliminated from the network when calculating its clustering coefficient distribution, the maximum possible diameter of the newly generated network is
. Here,
stands for the x-order clustering coefficient value of node
, while the last entry
in
equals the proportion of neighbor pairs of
with no connecting path between them.
Conventional standard clustering coefficients only measure node agglomeration based on the count of triangular subgraphs, which results in a single-dimensional measurement and fails to fully capture diverse linkages among neighboring nodes. In contrast, high-order clustering coefficients extend measurement to higher-order subgraph structures. They enable refined characterization of multi-layered connections between the neighborhoods of innovation entities and accurately identify intricate topological features inside enterprise-led innovation collaboration networks. Analyzing network agglomeration characteristics and functional modular division rules via high-order clustering coefficients can clarify the internal logic behind structural resilience and stable operation of collaboration networks, and supply critical topological indicators for unpacking the dynamic evolutionary mechanisms of enterprise-led innovation collaboration networks. In empirical analysis of enterprise-led industry–university–research innovation collaboration networks, this indicator effectively excavates multi-level complex collaborative patterns across enterprises, universities and research institutes, clarifies diverse interaction logics and collaborative organizational frameworks within industrial chains and innovation alliances, and profoundly reveals the underlying evolutionary mechanisms of innovation resource aggregation and partnership iteration.
The calculation procedures of high-order clustering coefficients are illustrated with a small network consisting of seven nodes in
Figure 2a. After removing node
and its associated edges from this network, we obtain the network shown in
Figure 2b. We calculate the distances between
,
,
,
and
(the neighbors of
) in
Figure 2b, and the corresponding distance matrix is as follows:
The distance distribution of node
derived from the distance matrix is listed below.
| 1 | 2 | 3 | 4 | 5 | 6 |
| 4 | 3 | 2 | 1 | 0 | 0 |
| 0.4 | 0.3 | 0.2 | 0.1 | 0 | 0 |
Accordingly, the 1-order clustering coefficient is , the 2-order clustering coefficient is , and so on, with the 5-order clustering coefficient equal to .
High-order distances account for complex paths and relationships within networks, measuring inter-node distances through indirect paths involving multiple intermediate nodes. High-order distance distributions reflect the complexity and diversity of relationships between nodes. Given an order
, the k-order distance distribution vector of node
is defined as:
where
represents the k-order diameter of the network. The distribution probability is formulated as:
denotes an indicator function that takes the value of 1 if the condition holds and 0 otherwise. Unlike first-order distances that only focus on the optimal path with the fewest intermediate nodes, high-order distances incorporate multiple detoured and parallel indirect pathways containing numerous intermediate nodes in the network. By counting the share of nodes linked via high-order paths of varying lengths, high-order distance distributions fully capture the multiplicity, complexity and hierarchical disparities of indirect connections between nodes. While first-order distances merely describe the fastest connectivity route between two nodes, high-order distances simultaneously characterize multiple alternative indirect relationships between node pairs, and can quantify the richness, redundancy and multi-channel propagation features of relational ties.