1. Introduction
The patent cooperation network depicts the cooperative relationships formed by various patent applicants through joint patent filings. It constitutes a complex network encompassing inter-enterprise collaboration, collaboration between enterprises and universities, as well as collaboration between enterprises and research institutions. As an important manifestation of collaborative innovation among innovation agents, studies on the patent cooperation network can reveal the characteristics and pathways of collaborative innovation [
1,
2]. In recent years, to promote the innovation-driven effect of developed regions on their surrounding areas, the Chinese government has formulated strategies for regional collaborative innovation development. Through measures such as resource allocation, industrial restructuring, and tax incentives, the government aims to boost innovation cooperation among sub-regions within a major regional cluster. Typical examples of such collaborative development regions include the Beijing-Tianjin-Hebei region, the Yangtze River Delta region, and the Pearl River Delta region. Against this policy backdrop, numerous scholars have conducted research on intra-regional patent cooperation networks to evaluate the effectiveness and dynamics of regional collaborative innovation. Meanwhile, in the process of regional collaborative innovation, green technology innovation has attracted growing attention due to its significance for regional ecology, environmental protection, and sustainable development. Therefore, it is of great practical significance to study regional green patent cooperation networks, analyze their structural characteristics, and reveal their evolutionary trends through link prediction. On the one hand, predicting the future development trends of green technology innovation cooperation within a region enables the early identification of network structural deficiencies, the evaluation of the driving effect of collaborative innovation, and the optimization of policies for green technology collaborative innovation. On the other hand, link prediction can assist patent applicants in identifying potential partners and facilitating the realization of collaborative innovation.
Although extensive research on link prediction has been conducted based on knowledge cooperation networks, few studies have specifically focused on link prediction for regional green patent cooperation networks. Against this backdrop, this paper aims to construct a regional green patent cooperation network and perform link prediction to uncover its evolutionary trends. At the methodological level, the primary challenge lies in selecting appropriate prediction indicators and constructing an effective prediction model to enhance prediction accuracy. Existing link prediction studies have employed a wide array of indicators, such as those based on node similarity, path similarity, and content similarity, with each category encompassing multiple specific indicators. Scholars typically combine one or more types of indicators to improve prediction accuracy [
3,
4,
5]. Building on existing research, this paper proposes a novel hybrid model based on multidimensional features. Compared with previous related studies, the proposed model exhibits the following distinctive features: (1) For indicators within the same category, a multi-indicator coupling method based on the entropy weight method is introduced. (2) The heterogeneity of intermediate links and intermediate nodes in the network is simultaneously incorporated into path features. (3) Content similarity, node similarity, and path feature indicators are all integrated into the prediction model, and a method for calculating the optimal weights of these three indicator categories is proposed.
The remainder of this paper is organized as follows.
Section 2 reviews the related research.
Section 3 elaborates on the construction method of the prediction model.
Section 4 presents prediction experiments conducted on a green patent dataset of the Beijing-Tianjin-Hebei region and provides a comparative analysis of multiple prediction models.
Section 5 performs a predictive network analysis of link prediction results for the Beijing-Tianjin-Hebei green patent cooperation network and offers policy recommendations for the development of green technology collaborative innovation in the region.
2. Related Work
From the perspective of research objects, the regional green patent cooperation network falls within the category of patent cooperation networks. From the perspective of research methods, link prediction is categorized under the theoretical study of complex networks. Therefore, the review of related work will be organized around two aspects: patent cooperation networks and link prediction. Specifically, studies on patent cooperation networks cover three sub-themes: patent cooperation networks, green patent cooperation networks, and regional green patent cooperation networks. Studies on link prediction include three major types of prediction indicators: node-based, path-based, and semantic-based, as well as the relevant applications of combinations of these indicators in the prediction of knowledge cooperation networks.
- (1)
Research on patent cooperation networks
Utilizing patent applicant data to construct patent cooperation networks has long been a hot topic in the field of collaborative innovation. Tsay et al. [
1] constructed a global patent cooperation network in the field of artificial intelligence and analyzed the evolution characteristics of this network. Chen et al. [
2] collected patent cooperation application data of different types of organizations in China from 2007 to 2015 and constructed an inter-organizational patent cooperation network. Liu [
6] studied the relevant patent network of Tsinghua University in China and analyzed the important role of universities in collaborative innovation. Liu et al. [
7], based on the patent collaborative application data of the National Intellectual Property Administration, applied complex network theory and social network analysis methods to study the collaborative network of China’s wind energy industry. Mei et al. [
8] constructed a collaborative innovation network based on collaborative patent and paper data from the Beijing-Tianjin-Hebei region from 2010 to 2016 to explore the characteristics and efficiency of innovation cooperation among cities in this region.
Although patent data have been widely used in the study of the characteristics of collaborative innovation networks, there are few studies on green patent data and green innovation. Liu et al. [
9] utilized Chinese patent data and the patent classification numbers provided by the IPC Green List to construct a green patent cooperation network from 2007 to 2017. They analyzed the characteristics of this network from multiple perspectives, including time, region, and spillover effect. Zhou et al. [
10] collected green patent data of the Yangtze River Delta urban agglomeration from 2000 to 2016 and constructed a patent cooperation network, analyzing the evolution model of green collaborative innovation in the Yangtze River Delta region. Fan et al. [
11] analyzed the spatial correlation network of China’s green innovation using the gravity model and social network analysis model, and the results showed that there was a spatial effect in the green innovation correlation network of China. It was found that regions with strong green technological innovation capabilities did not have a significant impact on the green innovation of other provinces, while regions with medium innovation strength might play a crucial role in the green innovation linkage network. Wang et al. [
12] investigated the spatial correlation network structure of green innovation efficiency in the Yangtze River Delta region, revealing that the region’s green innovation efficiency in the Yangtze River Delta region was extremely unbalanced and the spatial network correlation density was low. Bai et al. [
13] proposed a framework for determining the impact of multiple relationship networks on green innovation. They employed text mining to construct multiple relationship networks, identified green patents from massive patent data through content analysis, and then utilized fuzzy set Qualitative Comparative Analysis (fsQCA) to examine the equivalent impact paths of green innovation. In the above studies on the structure of green patent cooperation networks, although relatively rich results have been achieved, they primarily focus on analyzing network characteristics, influencing factors, and evolution patterns.
- (2)
Research on link prediction
Link prediction in social networks refers to identifying missing connections within the network based on existing relationships or predicting whether unconnected nodes will establish connections in the future [
3]. For enterprises engaging in innovation cooperation, selecting the appropriate R&D partners is crucial in determining whether collaborative innovation can succeed and enhancing organizational performance [
14]. Therefore, identifying possible innovation partners through link prediction is also a hot topic in recent research on cooperative networks. The most commonly used indicators for innovation cooperation network link prediction can currently be divided into three categories: node-based, path-based, and semantic-based [
15,
16].
The prediction algorithm based on node features is based on the common neighbors of nodes and uses the structural similarity matrix of network nodes to represent the similarity between nodes, such as Common Neighbor Indicator (CN), Salton Indicator, Jaccard Indicator, Hub Promoted Index (HPI), LHN-I Indicator, Adam-Adar Indicator (AA), and Resource Allocation Indicator (RA) [
15,
17,
18,
19]. These algorithms primarily rely on local information and are logically straightforward to implement, offering high computational efficiency. Chen et al. [
2] introduced eight commonly used node similarity indicators to predict cooperative relationships based on the patent application data in 2016 and found that the CN index was particularly suitable for organizations’ decision-making on choosing unfamiliar partners in patent cooperation. Zhang et al. [
20] compared 10 common indicators based on node features and found that the AA index had the best prediction accuracy in the patent cooperation application network of the Guangdong-Hong Kong-Macao Greater Bay Area. Shi et al. [
21] used the entropy weight method and integrated the four most accurate prediction indicators based on node features to develop a new indicator, which proved more accurate in predicting cooperation between scientific and technological entities in the Beijing-Tianjin-Hebei region.
The prediction methods based on path features utilize network paths for prediction, such as Katz index [
22], LP index [
23,
24], and LHN-II [
17]. Among these methods, the Katz index makes predictions based on the global path information between two nodes, but it has a high time complexity in calculation [
22]. Zhou et al. [
23,
24] proposed the Local Path (LP) index, which only incorporates paths of length 2 and 3 into the prediction, balancing accuracy and time complexity, but it does not consider the issue of path heterogeneity. In recent years, prediction algorithms based on path heterogeneity features have continued to develop. Zhu et al. [
25] fully utilized the degree information of intermediate nodes in similarity calculation and proposed the Significant Path (SP) index. Empirical experiments on twelve different real-world networks show that this index outperforms mainstream link prediction baseline methods. Zhu et al. [
26] believe that the connectivity of intermediate nodes also affects path heterogeneity, and based on this, they proposed the semi-local index NSI (Neighbor Set Information). Some studies on innovation cooperation have also integrated path indicators with other indicators, Liu and Sun [
27], who fused the path-based feature indicators with the similarity of research interests, and Wang et al. [
28], who fused the CN, RA, Jaccard, AA based on node features with the Katz index based on path features to predict potential patent partners.
With the rapid development of innovation networks, some researchers have found that traditional prediction algorithms based on node and path features are no longer sufficient to meet the prediction requirements of cooperative networks, and have therefore turned to applying text semantic features for the prediction of cooperative networks. For example, Chuan et al. [
16] proposed a link prediction algorithm based on Latent Dirichlet Allocation (LDA) topic modeling from the perspective of content similarity, utilizing cosine similarity to calculate the similarity between topics and replacing the original prediction algorithm that relies on network structure features. Jeon et al. [
29] identified potential partners with the required technology by mining the relationships between words in patent claims. Wang et al. [
30] proposed a new algorithm that identifies R&D partners based on the similarity of subject–action–object semantics. Kang et al. [
31] further proposed a method for selecting sustainable industry–university–research cooperation partners based on an LDA topic model.
In research on link prediction in cooperative networks, there is a trend of integrating indicators based on node, path, and semantic features. The integration methods consider both the semantic features of the cooperative content and the structural features of the cooperative network. For example, Park et al. [
32] explored potential R&D partners through literature coupling and patent semantic analysis. Ding et al. [
33] developed a method for mining potential partnerships based on the similarity of authors’ paper contents and the structure of the cooperative network.
- (3)
Review of related work
Given the dual importance of regional collaborative innovation and green technology innovation, it is highly necessary to conduct link prediction for regional green patent cooperation networks. However, a review of existing studies indicates that research on regional green patent cooperation networks has mainly focused on network characteristic analysis, spatial effects, and innovation pathways, whereas link prediction research on such networks remains underexplored. This research gap has motivated the present study. In addition, the current research on prediction indicators has covered a wide range of perspectives, yet there remain notable limitations in the application of these indicators to Patent Network Prediction (PNP), as shown in
Table 1. First, integrating the network’s topological structure with semantic content similarity has emerged as a new trend in link prediction, but this approach has not yet been applied to the link prediction of regional green patent cooperation networks. Second, heterogeneous paths have not been fully considered, and few studies have correlated the specific topological structure of patent cooperation networks with the issue of path heterogeneity. Third, with regard to the combined prediction using multiple indicators, a unified indicator coupling method and an optimal weight calculation method are required. Therefore, how to integrate multidimensional prediction indicators into the prediction model and improve the accuracy of link prediction constitutes the primary problem to be addressed in conducting link prediction for regional green patent cooperation networks.
3. Research Method
To achieve link prediction for regional green patent cooperation networks, a comprehensive model based on multidimensional indicators is proposed. This model integrates predictive indicators from three dimensions: node features, path features, and content features. In terms of node features, the entropy weight method is employed to couple traditional node similarity indicators, enhancing the universality of the indicators. Regarding path features, the heterogeneous influences of intermediate links and intermediate nodes are integrated to fully emphasize the issue of heterogeneous paths. For content features, the LDA model is used to construct topic set vectors for nodes, and multiple similarity distances are integrated to calculate content similarity (CS). The methodological framework of the model construction is illustrated in
Figure 1. In the framework, the network is composed by nodes and edges, and the letters in the node circle means the number of the patent applicants.
3.1. Predictive Indicator Based on Node Similarity Metric (Coupling)
The network constructed based on patent cooperation data can be represented as G = (N, E, W), where N represents the set of nodes in the network, E represents the set of edges in the network, W represents the edge weight, which is the number of collaborations of the entities, and |N| represents the total number of network nodes.
Link prediction considering node feature similarity holds that the higher the similarity of nodes, the greater the possibility of cooperation [
15]. For example, the CN index refers to the number of common neighbors the two nodes share; the more likely the two nodes are to cooperate. Node similarity indicators have evolved over the years, and there are currently 10 commonly used indicators. The formulas are shown in
Table 2. In the formula,
represents the degree of node
x,
and represents the neighbor nodes of node
x.
The prediction accuracy of each of the above indicators varies in different networks. Therefore, in this paper, the 10 node similarity indicators are coupled using the entropy weight method, and the constructed coupled indicator is denoted as Coupling, as shown in Formula (1):
Among them, represents the weight value based on the entropy weight method, . represents the similarity score of nodes x and y under the coupling index, represents the similarity score of nodes x and y under the common neighbor (CN) structure, and is similar.
The entropy weight method draws on the concept of entropy in thermodynamics and uses information entropy to describe the amount of information of an event. In this paper, it represents the amount of information that each indicator accounts for in node features: the smaller the entropy value, the greater the dispersion degree of the indicator, and thus the greater its contribution to indicator ranking. For weight calculation using the entropy weight method, the first step is to perform sum standardization on each indicator, as shown in Formula (2).
where
j denotes the indicator number, which ranges from 1 to 10 in this paper;
i denotes the node pair number. Secondly, calculate the information entropy value of each indicator for all node pairs, as shown in Formula (3).
where
m represents the total number of node pairs to be predicted. Since the entropy value is inversely proportional to the contribution, a positive relationship transformation is performed on the entropy value, as shown in Formula (4).
3.2. Predictive Indicator Based on Heterogeneous Path Metric (HP)
The path characteristics in the cooperation network can also accurately predict the cooperative relationship. However, when constructing the path characteristics, it is necessary to distinguish the influence of different paths on the prediction results, so as to further improve the prediction accuracy of the path characteristics. In this study, in order to distinguish the heterogeneity of the paths between two entities, the heterogeneity of the intermediate edges and the heterogeneity of the intermediate nodes are simultaneously introduced into the path characteristics.
Intermediate edge heterogeneity: For a link with multiple intermediate nodes, the larger the degrees of the two end nodes, the more dispersed their cooperation intentions. Thus, the less likely these two nodes will cooperate [
26]. Here, the product of the degrees of the two end nodes is used to represent the link weight, that is, the heterogeneity of the intermediate edge. The link weight is negatively correlated with the cooperation probability. Therefore, for two nodes
x and
y in the network with a link between them, the influence of the link weight can be expressed as
. For the intermediate nodes between nodes
x and
y, all the intermediate nodes are represented by a set, denoted as
, and the total heterogeneity of the intermediate edge of this intermediate node set is denoted as
, which is expressed by Formula (5):
Intermediate node heterogeneity: Even with the heterogeneity of edge connections, when the product of the degrees of the nodes in the intermediate links is equal, it still cannot effectively differentiate the impact of paths on cooperation. Therefore, further consider the heterogeneity of intermediate nodes as a weighted weight to enhance the discrimination of path heterogeneity. Intermediate node heterogeneity refers to the higher the connectivity between intermediate nodes, the more stable the connection between end nodes, and the higher the possibility of cooperation [
25].
For node pairs
x and
y in the network that have at least one intermediate node, all the intermediate nodes are recorded as a whole Z, and the number of intermediate nodes is represented by |Z| = M. When M = 1, it is a second-order path, and its connectivity is represented as whether the intermediate node Z and the two neighboring nodes
x and
y can form a triangular ring. |
represents the actual number of triangular rings formed by all the surrounding neighbors of Z, and |
represents the number of all possible triangular rings that the surrounding neighbors of Z can form. Then, the second-order path heterogeneity score of the intermediate nodes Z of nodes
x and
y is recorded as Formula (6):
When M = 2, it becomes a third-order path, and its connectivity lies in whether the intermediate node set and the two neighboring nodes x and y can form a quadrilateral loop. At this time, the influence of the intermediate node set on similarity is similar to that when M = 1, that is, the actual number of quadrilateral loops formed is divided by the possible number of quadrilateral loops.
Considering the characteristic that the prediction accuracy decreases rapidly with the increase of path length, as well as the time complexity issue of the calculation process, this paper sets the number of intermediate nodes to M = 2. The heterogeneity path (HP) index
based on the intermediate edges and intermediate nodes is constructed, and is expressed by Formula (7):
Among them, represents the heterogeneity score of nodes x and y based on the heterogeneity of intermediate edges and intermediate nodes, represents the heterogeneity score of nodes x and y based solely on the heterogeneity of intermediate edges, represents the heterogeneity score of nodes x and y based solely on the heterogeneity of intermediate nodes, and α is the weight value of the third-order path heterogeneity index, with .
3.3. Predictive Indicator Based on Content Similarity Metric (CS)
In addition to the network structure information, the technical field of the patent applicant also has a significant impact on the cooperative relationship. Therefore, the content similarity between nodes needs to be introduced into the prediction model. The specific steps are as follows: The patent abstract is taken as the main content for text analysis. The LDA model is used to divide the patent abstract into themes. Then, vector similarity indicators such as the Jaccard coefficient, Euclidean distance, and Manhattan distance are utilized to calculate the content similarity of the patent, representing the degree of association of the applicant’s technical field.
Before applying the LDA model for topic classification, the optimal number of topics needs to be determined first. In this paper, the CV consistency score method is used to calculate the optimal number of topics K. Based on this, the LDA model is applied to calculate the topic distribution probability of each patent, and the topic with the highest probability value is selected as the topic of that patent. The calculation of probability is as shown in Formula (8).
Among them, P(w|d) represents the probability of word w in document d, P(w|t) represents the probability of word w in a specific topic t, and P(t|d)represents the probability of document d in a specific topic t.
After obtaining the topic of each patent, the topic is assigned to each applicant entity of the patent, indicating that the research direction of this applicant includes this topic. After assigning the topics one by one, a topic set vector can be formed for each applicant, representing the complete set of research directions of this applicant. Finally, the patent content similarity between applicants is calculated based on the topic set vectors. To improve the prediction effect, six common vector similarity indicators were selected, as shown in
Table 3. Based on the experimental results, the three indicators with the best prediction accuracy were selected, and then the entropy weight method was used for coupling.
The formula of similarity index for the coupled content is shown as follows:
Among them, represent the weight values based on the entropy weight method, and . represents the content similarity score of nodes x and y under the coupling index, while , , are the top three vector similarity indicators.
3.4. Comprehensive Index NPC Based on Network Topology Structure and Content Similarity
By integrating the coupling node similarity index Coupling, the path heterogeneity index HP, and the content similarity index CS, a new link prediction index is constructed and denoted as NPC (Node & Path & Content index). Its calculation is as shown in Formula (10):
Among them,
are the optimal weight values of each indicator obtained through mathematical programming. The mathematical model is as Formula (11):
Among them, represents the prediction accuracy of the NPC indicator, while , , and are the similarity score matrices generated for the corresponding indicators.
5. Analysis of Link Prediction Results
Based on the constructed novel prediction model, link prediction was performed on all unconnected node pairs in the green patent cooperation network of the Beijing-Tianjin-Hebei region. In typical new link prediction tasks, 0.1–1% of the total number of potential node pairs is usually selected to generate the predicted network [
15]. An analysis of the network spanning 2020–2023 shows that the proportion of actual edges to all possible edges is 0.2%; thus, this study adopted 0.2% as the threshold. The NPC indicators were calculated for all unconnected node pairs, and the top 0.2% of node pairs ranked by NPC scores were selected to generate the predicted network. In this way, the predicted network of green patent cooperation in the Beijing-Tianjin-Hebei region was obtained, as illustrated in
Figure 5. The prediction network contains 232 nodes and 219 edges. In the figure, nodes of different colors represent the main types of patent applicants; nodes of different sizes indicate the degree difference among applicants; and edges of different colors represent the patent cooperation relationships between different regions.
The analysis of the primary type data is shown in
Figure 6. It can be seen that in the prediction network, enterprises still dominate among the patent applicants, indicating that enterprises will remain the main force of innovation in the future. Such an innovative development trend not only conforms to the development law dominated by the market but also fully reflects the crucial role of enterprises in innovation. Among all the cooperation nodes, although research institutions account for about six percentage points more than universities, it can be clearly seen from the network diagram that the degree of university nodes is generally larger than that of research institution nodes, and they are also more likely to form structural holes. Therefore, enterprises and the government should pay more attention to the importance of universities in the future patent cooperation network, and make good use of the professional advantages and network intermediary role of universities. At the same time, attention should also be paid to the intermediary ability of research institutions, fully promoting the transformation of the vast knowledge resources of research institutions into productive forces and cultivating more influential research institutions.
By mapping the nodes in the prediction network onto the map of the Beijing-Tianjin-Hebei region, a regional distribution map of the future green patent cooperation network can be obtained, as shown in
Figure 7. It can be seen that the distribution of applicants in Beijing is the densest, followed by Tianjin, and the distribution in Hebei is the sparsest.
The analysis of the data in
Figure 7 is shown in
Figure 8. It can be observed that in the future green cooperation, the number of patent cooperation within the Beijing region accounts for more than half, while cross-regional cooperation between Tianjin and Hebei accounts for only 3.21%. This indicates that Beijing’s independent innovation ability in the future remains the strongest, while the cross-regional cooperation between Tianjin and Hebei is the weakest. The patent cooperation within the Hebei region (10.09%) is higher than that within the Tianjin region (8.72%), indicating that future innovation cooperation within Hebei is likely to be better than that within Tianjin, suggesting certain innovation potential. The data in
Figure 7 also indicate that a regional imbalance persists in the future innovation development of the Beijing-Tianjin-Hebei region, with cross-regional cooperation generally lower than internal regional cooperation. In terms of cross-regional cooperation, the patent cooperation between Beijing and Hebei (12.39%) is higher than that between Beijing and Tianjin (9.17%), indicating that Tianjin’s future cross-regional cooperation innovation ability still needs improvement. From the perspective of regional collaborative development, the government should promote more innovative cooperation between Beijing and Tianjin to achieve the radiation and leading role of innovation in Beijing. For instance, the government should provide special funds for Beijing-Tianjin collaborative innovation, build a Beijing-Tianjin innovation resource sharing platform, and establish a mechanism for two-way flow of talents between Beijing and Tianjin. In cross-regional cooperation, the cooperation between Tianjin and Hebei has the lowest proportion, indicating that both regions are relatively dependent on Beijing’s innovation resources, and there has not yet been a mutual driving effect in collaborative innovation between Tianjin and Hebei. To improve the situation, the government should make full use of the national-level platforms such as Tianjin Binhai New Area and Hebei Xiongan New Area, and jointly establish Tianjin-Hebei Industrial Cooperation Demonstration Parks. The parks should focus on fields with strong complementarity between the two sides such as new energy and high-end equipment manufacturing, guide upstream and downstream enterprises in the industrial chain to settle in, and promote innovative cooperation between enterprises in Tianjin and Hebei.
6. Conclusions
This paper, based on the green patent cooperation network in the Beijing-Tianjin-Hebei region from 2020 to 2023, established a new link prediction model. This model integrates the prediction indicators of node characteristics, path characteristics, and content characteristics, and improves each of these indicators. In terms of node similarity, the entropy weight method is used to couple 10 node similarity indicators to improve the prediction accuracy of multiple indicators in the patent cooperation network. In terms of path heterogeneity, the heterogeneity of intermediate edges and intermediate nodes in the network is simultaneously introduced into the path characteristics to construct a weighted prediction model for multi-level paths. In terms of content similarity, the LDA topic model is used to classify the patent abstracts, and the entropy weight method is coupled with the precision values of the higher accuracy Jaccard, Euclidean, and Manhattan indicators. Finally, considering the indicators of node similarity, path heterogeneity, and content similarity based on the model, a comprehensive prediction model NPC was constructed. The optimal weight values of the three-dimensional indicators were solved using GWO method which takes the AUC as the objective.
The results of the comparative experiments demonstrate that the NPC-based prediction model, which comprehensively incorporates the network’s node features, path features, and collaborative content, achieves significantly higher prediction accuracy than models relying on any single category of indicators. In the optimal weight combination of the NPC indicators, node similarity indicators contribute the most, followed by path indicators, while CS indicators make the smallest contribution. Despite their relatively low weight, CS indicators still remarkably improve the prediction accuracy, which underscores the necessity of indicator integration. An analysis of the predicted network was conducted from two dimensions: the distribution of agent types and regional distribution. The findings reveal that imbalances will persist in the future development of green technology collaborative innovation in the Beijing-Tianjin-Hebei region, and the driving effect of cross-regional collaborative innovation should be further strengthened.
Although the prediction performance of the model proposed in this paper is relatively satisfactory, there are still certain limitations and shortcomings. (1) The methodology adopted in this paper is indicator fusion, which integrates multiple mainstream prediction indicators to improve prediction accuracy. However, this method involves an excessive number of indicators, leading to extremely high computational complexity when applied to large-scale networks. Therefore, this method is not suitable for the prediction of large-scale networks. (2) Given that this method constructs a new prediction model based on indicator fusion, it is necessary to adjust the combination scheme of various indicators according to AUC during the fusion process, resulting in strong dependence on datasets. It is possible that indicator fusion may fail to improve prediction accuracy when applied to other datasets. (3) This paper only considers the link prediction of potential edges. However, practical patent cooperation networks are weighted networks, where existing cooperative connections still have the possibility of re-collaboration—in other words, the weights of edges also need to be predicted.
To address the aforementioned limitations, the future work of this study will focus on two aspects. First, more datasets of patent cooperation networks will be utilized to verify the effectiveness of the proposed method. Second, based on the weighted network data of patent cooperation, the weights of existing cooperative edges will be predicted, so as to more accurately evaluate the future evolutionary trends of the network.