1. Introduction
With the rapid development of emerging technologies such as big data, the Internet of Things, and the metaverse, data has transcended its traditional resource attributes and has become a core production factor driving social progress and economic development. It has emerged as a crucial foundation for driving educational innovation, promoting educational equity, and achieving educational modernization [
1,
2]. University data governance is an inevitable trend shaped by multiple factors, including the deepening of educational informatization, the surge in data volume, the increasing demand for scientific decision-making, tightening laws and regulations, and intensified international competition.
In 2015, the State Council issued the “Action Plan for Promoting Big Data Development”, emphasizing the creation of “education and cultural big data” and promoting the collection and sharing of foundational educational data. In 2018, the Ministry of Education (MOE) released the “Education Informatization 2.0 Action Plan”, proposing the use of education informatization to support and lead the modernization of education, making it a strategic choice for educational reform and development in the new era. In the same year, the “Regulations on Education Data Management for MOE Agencies and Directly Affiliated Institutions” were issued, clearly outlining the management of educational data collection, storage, sharing, public disclosure, and security. In 2021, the “Notice on Strengthening the Informatization of Educational Management in the New Era” by the MOE emphasized improving the management of educational data from four aspects: standardizing management, open sharing, quality assurance, and management effectiveness.
These developments indicate that the evolution of educational informatization continuously raises new demands for data governance in universities, and building a comprehensive data governance system has become an integral part of advancing the modernization of university governance systems and capabilities. The attention given by the government has driven universities to successively introduce data management regulations to guide their data governance practices. However, current university data governance policies face numerous challenges in the formulation process, such as the strategic orientation of policy goals, the completeness of institutional design, and the coordination of the policy framework. These issues not only impact the scientific and systematic nature of policy texts but also constrain the deepening of university data governance work. Therefore, conducting scientific, objective, and comprehensive evaluations of university data governance policies, analyzing the internal structure and constituent elements of policy texts, and identifying deficiencies in policy design to provide empirical evidence for policy formulation, adjustment, and improvement have become critical issues that need to be addressed in the field of university data governance.
2. Related Research
The concept of data governance has yet to reach a unified consensus within the academic community. Existing research commonly acknowledges three essential attributes of data governance: “cross-functional”, “strategic”, and “control-oriented”. These attributes manifest in six key dimensions: cross-functional collaboration, strategic asset positioning, decision-making authority and responsibility allocation, policy standardization, compliance monitoring, and framework attributes. Among these, the positioning of data as a strategic asset serves as the core foundation of the concept. For instance, Khatri and Brown [
3] explicitly define data governance as the “framework for the allocation of decision-making authority and responsibility regarding organizational data assets”, while DAMA International [
4] further emphasizes its function of “planning, monitoring, and executing data management through authority and control”. Additionally, cross-functional collaboration (involving business, IT, and data management departments), decision-making authority and responsibility allocation, policy standardization, and compliance monitoring together constitute the core elements of the definition, with over 80% of the literature referencing these dimensions [
5]. Notably, the emphasis in the definition of data governance varies according to the research context and industry characteristics. For example, research focusing on traditional data management [
6] places greater emphasis on data quality and lifecycle control, while studies in the context of big data environments [
7] introduce additional dimensions such as “data value quantification” and “privacy protection adaptation”. From an industry perspective, the healthcare sector [
8] highlights data security and compliance requirements, whereas corporate settings [
9] focus more on how data governance enhances business performance.
In the global wave of digital transformation, data has become a core strategic asset for universities, and the level of its governance directly impacts the enhancement of educational quality, breakthroughs in scientific research innovation, and improvements in management efficiency. Current research in the field of university data governance focuses on two major aspects: first, the theoretical construction of governance models, and second, the practical innovation of implementation paths. Both aim to improve the overall effectiveness of university data governance and mainly encompass maturity assessments, scenario-driven approaches, collaborative governance, and adaptation to smart education [
10,
11]. Regarding the former, there are primarily two research perspectives. One is represented by Zhang [
12] and Yasmine [
13], who advocate for the improvement of authoritative models, innovating theoretical frameworks based on authoritative frameworks such as DAMA and ADDIE in response to the governance needs of different educational contexts. The other is represented by Li et al. [
14], who focus on the deconstruction of the core elements of data governance to reconstruct governance frameworks. In the research on the implementation paths of university data governance, scholars such as Ramadhan [
10], Jim [
15], and Shen [
16] have proposed specific optimization strategies in areas like quality assurance, data standards, intelligent systems, management systems, and top-level design.
Policy evaluation, as a key component of public policy analysis, is crucial not only for optimizing the allocation of policy resources but also an important means to systematically review and assess the scientific and normative quality of policy texts. This process helps ensure the structural completeness of policies at the level of institutional design, providing a reference for subsequent policy improvement and revision. In the field of university data governance research, existing studies mainly focus on model framework construction or specific implementation path exploration, with limited quantitative evaluation of policy texts, particularly lacking in-depth analysis of data governance policies at the individual, micro-level within universities. Furthermore, there is a lack of systematic guidance on optimizing and improving data governance practices that align with the actual circumstances of universities. Data governance policies, as an indispensable component of university data governance, not only serve as a tangible representation of the theoretical framework but also play a significant role in guiding and advancing data governance practices within universities. This paper aims to fill the gap in the field of university data governance policy evaluation research by using the PMC index model to conduct a quantitative evaluation of university data governance policies. The PMC index model is comprehensive, highly quantifiable, flexible, visually demonstrative, and highly reproducible. It provides a holistic and objective assessment of the research object and serves as a powerful tool for decision-making. The model has been applied in various policy evaluations, such as those in digital economy, carbon reduction, mental health, and technology finance, and has been widely recognized by scholars [
17,
18,
19,
20].
This study examines the data governance policies of China’s “Double First-Class” universities (“Double First-Class” universities refer to higher education institutions included in China’s national initiative, the “World-Class Universities and First-Class Disciplines” program). By applying policy text mining and the PMC index model, an evaluation framework for university data governance policies was developed, consisting of 9 main indicators and 43 sub-indicators, followed by a quantitative analysis. Through the examination of micro-level data governance policies within individual universities, the study systematically analyzes the consistency, completeness, and strengths and weaknesses of these policies. This approach addresses a gap in current research on the evaluation of university data governance policies, providing a new perspective on the topic. The theoretical contributions of this study are as follows:
- (1)
Developing a PMC-based indicator system tailored for university data governance policies, enabling the structural quantification of policy texts;
- (2)
Expanding the application scope of the PMC model in the field of higher education governance;
- (3)
Providing empirical analysis based on data from 56 universities, revealing structural differences in policies and offering directions for optimizing university data governance policies;
- (4)
Enriching the academic discussion on digital governance, institutional policy design, and evidence-based evaluation in higher education.
3. Data Source, Collection, and Stakeholders
In this context, “big data” in this study is primarily understood as the technological and governance background that drives the rapid growth of university data volume, diversity, and circulation demands, thereby necessitating institutionalized data governance. The empirical material of this paper consists of policy texts, and the term “big data” does not refer to the collection of large-scale operational datasets from universities. Rather, computational text-mining techniques were applied to analyze policy corpora and quantify the comprehensiveness and structure of the policies. Data source and collection. The data analyzed in this study consist of publicly accessible data governance policy documents issued by China’s “Double First-Class” universities, including regulations, measures, guidelines, notices, and related institutional documents. These documents were collected by the authors through official university portals and publicly available information channels for academic research purposes only. The dataset is a text corpus of policy documents and does not involve the collection, processing, or identification of any personal data of faculty, staff, or students.
Purpose and potential users. The authors constructed the policy text corpus to conduct text mining (word segmentation, keyword extraction, and co-word network analysis) and to build a quantitative evaluation framework based on the PMC index model. Potential users of the findings include university data governance offices and information management departments (for policy benchmarking and improvement), education authorities and evaluators (for understanding policy maturity and common gaps), and researchers (for policy text analysis and methodological replication).
Stakeholders and interest in information sharing. The dissemination of data governance policy information primarily serves institutional transparency and governance improvement. Stakeholders who may be interested in public sharing and interpretation of such policy information include universities, education authorities, faculty and students (to clarify data usage boundaries, responsibilities, and rights protection), and the public (to understand institutional data governance rules). Importantly, the “sharing” discussed in this study refers to governance rules and policy transparency rather than the sharing of personal or sensitive operational data.
4. PMC Index Model for Higher Education Data Governance Policies
The PMC (Policy Modeling Consistency) index model was selected for this study because it is particularly suitable for the quantitative evaluation of policy texts. Unlike qualitative policy analysis methods that rely heavily on expert interpretation, the PMC model decomposes policy documents into structured variables and evaluates the internal consistency and completeness of policy design.
Compared with multi-criteria decision-making methods such as the Analytic Hierarchy Process (AHP), the PMC model adopts an equal-weight structure and binary assignment rules, which helps reduce the influence of subjective weighting and improves methodological transparency. In addition, the PMC framework allows researchers to transform complex policy texts into quantifiable variables, making it well suited for comparative analysis of multiple institutional policies.
This study combines the PMC model with text mining techniques to identify key governance elements in university data governance policies, enabling a systematic assessment of the internal consistency and completeness of policies across different universities.
At the same time, the limitations of the PMC model are recognized: the binary assignment may overlook detailed information in the policy texts; the equal weighting assumption for indicators may not reflect the actual importance differences between elements; and the text interpretation process carries a degree of subjectivity. To address these limitations, the following measures were taken: high-frequency word analysis and social network map analysis were used to ensure that the variable setup reflects the core themes of the policy; two researchers independently coded the data; and qualitative analysis was incorporated in the interpretation of results, using specific content from the policy texts to complement the quantitative information.
The PMC index model is a policy evaluation model that comprehensively considers various policy indicators.
Figure 1 illustrates the evaluation process of the PMC index model for university data governance policies: (1) Variable identification and indicator setting: Policy texts are preprocessed by removing irrelevant words, and high-frequency valid terms are coded. Based on the themes and content of the policies, variables at different levels and their corresponding quantities are determined to construct the evaluation indicator system. (2) Construction of a multi-input–output table: Each variable is quantitatively analyzed from multiple perspectives. Variables are independent and unordered, and they are constructed and assigned values according to policy characteristics and the evaluation system. (3) Calculation of the PMC index: Using specific formulas, values of secondary and primary variables are calculated step by step, ultimately yielding the PMC score. (4) Visualization via the PMC surface diagram: The results are matrixed and visualized, and the concavity of the surface intuitively reflects the relative strengths and weaknesses of the policies.
4.1. Text Selection
In 2015, the State Council issued the “Action Plan for Promoting the Development of Big Data”, marking the elevation of big data development and application to a national strategic level. Therefore, this study selects milestone time points in the history of big data development as the starting point for the statistical time period of the sample data. The time range for the sample data is from 2015 to 2024, during which the data governance policy texts published by China’s “Double First-Class” universities were collected as the research samples. The policy texts were obtained by sequentially visiting the official websites of these universities, based on the list of “Double First-Class” universities published by the Ministry of Education. To ensure the accuracy of the sample data, the following principles were followed during the sample collection and screening process: (1) Only policy texts directly related to data governance were selected, excluding those related to information management. Here, “data governance” refers to policies that primarily address the strategic oversight of data assets, including data security, data sharing, data quality, and data lifecycle management, following the definitions of Khatri and Brown [
3] and DAMA International [
4]. In contrast, “information management” encompasses a broader scope, such as information systems, IT infrastructure, and record management, which may include but are not limited to data governance. To maintain sample homogeneity, we excluded policies whose titles or core content focus on general information management (e.g., “University Information Management Regulations”) and retained only those that explicitly target data governance or data management as their main theme. This distinction was verified by reading the abstracts and key sections of each candidate policy. (2) Policy texts that have been repealed were excluded, and only the latest revised versions were selected as policy samples. (3) Policy texts with indeterminate publication dates, non-university main campus documents, and those unrelated to the theme were excluded. By December 2024, a total of 56 valid policy texts were collected, as shown in
Table 1.
4.2. Variable Identification and Indicator Setting
This study utilized the ROSTCM6 software to conduct text mining on the data governance policies collected from China’s “Double First-Class” universities. The process involved word segmentation and data cleaning by removing meaningless terms such as “according to” and “based on”. Subsequently, a social network map of the relevant high-frequency words was generated. As shown in
Figure 2, “data” and “management” emerge as the core keywords in the data governance policy texts of universities, with the highest node degree and the broadest radiating scope, positioning them at the center of the social network map. These keywords reflect the core elements and key tasks of university data governance. Additionally, other important keywords in the policy texts include “sharing”, “department”, “unit”, “usage”, “resources”, and “security”, which also exhibit substantial radiating scope and form cohesive subgroups. Among these, “sharing”, “security”, “usage”, and “services” represent the goals and tasks of data governance in universities, while “department”, “unit”, and “resources” signify the carriers and essential supports for data governance.
This study selects the top 50 most frequent terms for frequency analysis, as shown in
Table 2, which can provide reference for the setup of secondary variables. From
Figure 2 and
Table 2, it is evident that, in terms of the objectives of university data governance policies, high-frequency terms include “management”, which appears 2423 times, highlighting that enhancing data management capabilities is the primary goal of universities. Terms such as “norm” (486 occurrences), “standard” (384 occurrences), “regulation” (305 occurrences), and “principle” (414 occurrences) emphasize the basic approaches and methods of data management. The term “security”, appearing 1102 times, indicates the high importance universities place on data security, which is a core aspect of the policy goal of ensuring data safety. “Protection” (183 occurrences) further emphasizes the focus on personal data protection, indirectly reflecting the importance and urgency of ensuring data security. The term “quality” appears 317 times, and “integrity” (188 occurrences) highlights the focus on data quality management through data integrity to improve data quality. The term “sharing” appears 1541 times, as sharing is a key objective of management, and management is aimed at facilitating better sharing. The construction of “platforms” (506 occurrences) and “systems” (658 occurrences), as well as the safeguarding of “resources” (1104 occurrences), contribute to data sharing and utilization. The term “service” appears 449 times, reflecting the optimization of data service provision functions.
Regarding the content of university data governance policies, high-frequency terms include “collection”, “storage”, “sharing”, “security”, “quality”, and “maintenance”, indicating that Chinese universities primarily focus on data collection, storage, sharing, security, quality, and operational maintenance in their data governance efforts. High-frequency terms related to the main actors of data governance policies include “department” (1479 occurrences), “office” (211 occurrences), and “school” (1388 occurrences), indicating the active participation of universities in data management. The policy objects are more diverse, with relevant high-frequency terms such as “data”, “business”, “system”, “resources”, “management”, and “process”, showing that the policy targets not only various foundational data, business data, system data, and educational resource data within universities, but also the data management processes and essential carriers for data storage and sharing, such as data platforms. Furthermore, there are high-frequency terms that represent the policy basis, tools, and safeguards, including “nation”, “university”, “technology”, “norm”, “department”, and “safeguard”.
According to the PMC index model construction guidelines, a total of 9 primary variables and 43 secondary variables were identified. The identification process followed a dual-source approach combining empirical text mining and theoretical frameworks. Keyword frequency analysis was first performed using ROSTCM6 on the 56 policy texts (
Table 2 and
Figure 2). High-frequency terms such as “management”, “sharing”, “security”, “quality”, “collection”, “department”, “platform”, and “safeguard” revealed the core themes addressed by the policies. These terms were grouped into thematic clusters corresponding to potential policy dimensions. Next, these clusters were mapped to the classic PMC primary variables introduced by Estrada [
21] and subsequently expanded by scholars such as Zhao et al. [
22]. For example, terms related to governance goals (e.g., “management”, “security”, “sharing”) informed the policy objectives variable (X2); terms describing governance activities (e.g., “collection”, “storage”, “quality”) shaped the policy content variable (X3); terms indicating actors (e.g., “department”, “office”, “school”) contributed to policy entities (X4) and policy objects (X5); and terms like “technology”, “training”, “standards” fed into policy instruments (X7). Subsequently, the secondary variables were refined by linking specific high-frequency terms to distinct indicators. For instance, under policy content (X3), the presence of terms such as “collection”, “storage”, “sharing”, “security”, “quality”, and “maintenance” led to the creation of secondary variables X3: 2 through X3: 7. In cases where empirical terms were limited but theoretical significance was high (e.g., “service outsourcing” as a policy instrument), the set was supplemented based on insights from the data governance literature. Finally, the variable set was validated through cross-referencing with existing policy evaluation studies, ensuring that all relevant dimensions of data governance were comprehensively covered. This process yielded the final evaluation system shown in
Table 3, with each secondary variable clearly linked to both empirical frequency and scholarly precedent. The PMC evaluation indicator system for university data governance policies and its explanation are presented in
Table 3. Apart from the two classic variables, X1 (policy nature) and X9 (policy evaluation), the secondary variables X2 (policy objectives), X3 (policy content), X4 (policy entity), X5 (policy objects), X6 (policy basis), X7 (policy instruments), and X8 (policy safeguards) were independently organized based on the policy text content and high-frequency terms.
4.3. Establish a Multi-Input–Output Table
The multi-input–output table in the PMC index model allows for quantitative analysis of a single variable from multiple perspectives. It is crucial for the overall evaluation of policy structural completeness, calculating the PMC index, assigning weights to secondary variables, and determining the indicator weights of variables. Key characteristics of the table include independent variables with no predefined order, primary variables that can be subdivided into secondary variables with consistent weight distribution, and an unlimited number of secondary variables. According to the research findings of Zhang Yong’an and other scholars [
23], binary values [0, 1] are assigned to secondary variables, where the parameters for secondary variables are set to 1 if the policy text covers the relevant content, and 0 otherwise. All variables have the same weight with no predefined order.
4.4. PMC Index Calculation
The PMC index model evaluates the structural completeness and internal consistency of policy documents by decomposing them into multi-dimensional variables. Unlike outcome-based evaluations, the PMC index measures whether a policy text systematically covers essential governance elements.
- (1)
Variable Structure
Based on a literature review and characteristics of university data governance, nine first-level variables (X1–X9) were constructed, including policy objectives, governance structure, data sharing, security mechanisms, supervision and evaluation, etc. Each first-level variable contains several second-level variables reflecting specific policy elements.
- (2)
Assignment Rules
Each second-level variable is assigned a binary value:
1 = the policy text clearly contains the corresponding element.
0 = the element is absent.
For example, if a university policy explicitly mentions “data security responsibility division”, the corresponding variable is coded as 1; otherwise, it is coded as 0. The binary (0–1) assignment method follows the standard practice of the PMC index model proposed by Estrada [
21]. This approach enables qualitative policy text to be transformed into quantifiable variables by identifying whether specific governance elements are present in the policy document. Binary coding improves the clarity and replicability of policy evaluation by avoiding ambiguous intermediate scores.
- (3)
Calculation Steps
The first-level variable value is calculated as the arithmetic mean of its second-level variables:
where
Xi is the value of primary variable
i (with
i = 1, 2, …, 9),
Xij is the
j-th secondary variable under
Xi, and
ni is the number of secondary variables for that primary variable (as specified in
Table 3). The secondary variables are binary:
Xij = 1 if the policy text covers the relevant content, and 00 otherwise. The assignment of these binary values is guided by the keyword frequency analysis (
Table 2 and
Figure 2), which identifies the core themes of university data governance policies. For example, the presence of terms such as “collection”, “sharing”, or “security” in a policy leads to assigning 1 to the corresponding secondary variables under policy content (X3).
The overall PMC index is then computed as the sum of the nine primary variable scores:
Since each first-level variable ranges from 0 to 1, the theoretical PMC value ranges from 0 to 9.
As indicated by the prevailing policy research grading criteria and the quantitative evaluation outcomes of university data governance policies, the PMC index is categorized into four levels: excellent, good, moderate, and poor. It is evident that policies which have been assigned a PMC index of [8, 9] are to be considered excellent. This is due to the fact that such policies are deemed to contain reasonable content, to be effective to a strong degree, and to be well-suited to university data governance practices, exhibiting a high level of consistency. It is evident that policies which have been assigned a PMC index within the range of [7, 7.99] are typically regarded as being of a satisfactory nature. This observation indicates that the content of such policies is considered to be reasonably sound, with a high probability of achieving the anticipated policy outcomes. Policies that have an PMC index falling within the range of [5, 6.99] are designated as moderate, signifying that the policy content is deemed to be deficient and is not adequately aligned with the anticipated policy outcomes. Policies with a PMC index of [0, 4.99] are designated as poor, indicating that the policy content is inadequate, fails to achieve the expected outcomes, and exhibits a low level of consistency. These policies require significant improvement in data governance and should be revised immediately. The evaluation levels of the policy are illustrated in
Table 4.
4.5. Plotting of PMC Surface Diagram
The PMC surface diagram provides an intuitive and effective visualization tool for in-depth analysis of policies, clearly demonstrating the strengths and weaknesses of China’s higher education data governance policies across multiple dimensions. Based on the nine primary variables of the higher education data governance policy PMC index model, a 3 × 3 PMC matrix was constructed according to Equation (3), and a three-dimensional curve plot was generated. The degree of surface concavity directly reflects the quality of the policy. A smaller concavity indicates higher scores on the relevant variables, suggesting a higher policy level. Conversely, a larger concavity indicates deficiencies in the policy in corresponding areas. Note that the matrix
PMC is not a separate calculation of the PMC index but a rearrangement of the primary variable scores for visualization purposes.
5. Quantitative Analysis
5.1. Overall Evaluation of Higher Education Data Governance Policies
Based on the data governance policies of 56 “Double First-Class” universities, the secondary variables in
Table 4 were assigned values. The first-level variable values and PMC values were calculated using Equations (2) and (3), as shown in
Figure 3. The policies were then rated according to the evaluation levels in
Table 4, with the results presented in
Table 5.
To ensure methodological transparency, the coding of policy texts was conducted using a structured coding framework derived from the PMC indicator system. Each policy document was examined to determine whether the elements corresponding to the second-level indicators were present. The coding process was carried out independently by two researchers. Each coder reviewed the policy documents and assigned binary values to the indicators according to predefined coding rules. A value of 1 indicates that the policy explicitly includes the corresponding element, while 0 indicates that the element is absent. After the independent coding stage, the results were compared. Any discrepancies were re-examined by both coders, and the final coding decisions were determined through discussion until consensus was reached. This procedure helped improve the consistency and reliability of the coding results.
As demonstrated in
Table 5, within the ambit of China’s “Double First-Class” universities, 14 institutions, including Northwest A&F University and South China Agricultural University, attained PMC indices ranging from 8 to 9, thereby attaining an “excellent” evaluation level, accounting for a 25% share. Concurrently, 37 universities, such as Northeast Normal University and Wuhan University, garnered PMC indices between 7 and 7.99, thus receiving a “good” evaluation, with a 66.1% representation of the total. Furthermore, five universities, including Chongqing University and the University of International Business and Economics, attained PMC indices between 5 and 6.99, thus being designated as “moderate”, accounting for 8.9%. It is noteworthy that no policies were assigned a “poor” rating. Overall, as pioneers in the digital transformation of higher education, China’s “Double First-Class” universities actively align with the national data governance strategy, emphasizing the leading role of data governance policies in university informatization. The policies are comprehensive in nature, encompassing data types, collection, storage, sharing, security, and quality, thereby establishing a relatively complete governance system. The formulation of these policies is well-founded, with clear objectives and scientific principles. In terms of policy instruments and safeguards, a range of instruments are employed to effectively advance university data governance. These include informational support, technical assistance and regulatory oversight. In order to ensure the implementation of the aforementioned policies, a series of corresponding measures have been put in place. These include organizational restructuring, interdepartmental collaboration, and enhanced supervision and evaluation. Furthermore, the average PMC index for university data governance policies is 7.6, indicating an overall “good” level. The policy system demonstrates strong performance in both the rationality of tool allocation and the completeness of institutional design, offering comprehensive guidance for the standardized implementation of data governance.
5.2. Analysis of Data Governance Policies in Universities of Different Levels
Based on the PMC indices of data governance policies presented in
Table 5 and
Figure 3, this section presents PMC curve diagrams for universities of different levels and provides an analysis of their data governance policies.
5.2.1. Evaluation of Excellent Policies
There are significant differences in the data governance policy levels among different universities. Universities rated as excellent, such as Northwest A&F University, Sun Yat-sen University, China Conservatory of Music, and Tianjin University of Traditional Chinese Medicine, perform exceptionally well across multiple policy indicators. For example, in terms of policy content, many policies receive a full score of 1, reflecting the comprehensiveness and thoroughness of the policy content. On one hand, data classification is detailed and aligned with the specific needs of each institution. For instance, Northwest A&F University and Shanghai Ocean University precisely classify data based on their respective business activities, which promotes the integration and flow of data management with business operations. On the other hand, data lifecycle management is organized and systematic, covering all aspects from data collection, storage, sharing, security, quality, to operation and maintenance. Many universities follow principles such as “one datapoint, one source”, and take stringent measures at each stage to ensure data quality and security. For example, Northwest University of Technology and Sun Yat-sen University have established effective data security management systems and quality assurance mechanisms.
Moreover, the policy foundations are well-established, with universities referencing national and local laws and regulations related to data governance during the policy formulation process. Sun Yat-sen University, for instance, has formulated its data governance policy in accordance with the “Cybersecurity Law of the People’s Republic of China”, the “Interim Measures for the Management of Government Information Resources Sharing” issued by the State Council, the “Comprehensive Governance Action Plan for Cybersecurity in the Education Sector” issued by the Ministry of Education, and other relevant laws, regulations, and national policies. These policies comply with national legal requirements, ensuring the legality and standardization of the data governance work. This safeguards that all aspects of the data governance process are backed by law, reducing legal risks and safeguarding the legitimate rights and interests of the university, faculty, students, and other stakeholders. Additionally, the policies are closely aligned with the actual situation and development plans of the universities. For example, South China Agricultural University has integrated national policies such as the “Action Plan for Promoting Big Data Development” issued by the State Council and the “Decision on Vigorously Advancing Informationization Construction” with its own institutional planning, tailoring the policies to the specific needs and characteristics of the university, which enhances their relevance and adaptability.
These excellent policies are relatively mature in their formulation and provide comprehensive and effective guidance for data governance in universities. Their policy models serve as valuable references for other universities. Taking Sun Yat-sen University as an example, a PMC surface diagram is presented in
Figure 4. The PMC curve for excellent-level policies primarily concentrates in the range of 0.85–1.0. The policies align well with standards across multiple evaluation dimensions, although the dips in the X4 (policy subject) and X5 (policy object) dimensions are relatively large. The policy subject is somewhat singular and lacks diversified participation, while the policy object is broader but still has room for expansion.
5.2.2. Evaluation of Good Policies
In the evaluation system for university data governance policies, institutions such as Northeast Normal University, Wuhan University, Henan University, and China University of Mining and Technology (Beijing) are classified as having good-level policies. These universities generally achieve a high standard in areas such as policy nature and policy objectives. The data governance policies of these institutions are clearly defined and integrate data governance into the overall management of the university. They recognize the critical significance of data as an intangible asset and a strategic resource, laying a solid foundation for future efforts. The policy objectives are reasonably set, with a core focus on constructing a comprehensive data management system. The policies emphasize ensuring data quality, security, and sharing, with a commitment to supporting core activities such as teaching, research, and administration. For example, Henan University places a strong emphasis on strengthening the unified management and quality control of informatized data, establishing an effective system for data sharing, management, and protection, and promoting the effective application of data in teaching, research, and administration.
Additionally, some policy contents are solid and effective, with clear organizational structures and responsibilities. For instance, the Information Management Office at Wuhan University is responsible for overseeing and planning the university’s overall information data management, including data resource planning and standard setting. Several universities have established corresponding mechanisms for data quality management. For example, Northeast Normal University has developed a comprehensive data control system that spans from data collection to maintenance, requiring regular checks on data quality to ensure the authenticity and completeness of source data.
These universities’ data governance policies exhibit several strengths but also present areas for improvement. Some institutions’ data governance policies lack clear references to national laws, regulations, and policy frameworks. For instance, universities such as Beijing Normal University, Wuhan University, Hohai University, and Peking University fail to sufficiently incorporate national laws, regulations, and policies in their policy foundations. Data governance involves numerous legal frameworks, such as the Cybersecurity Law of the People’s Republic of China and the Data Security Law of the People’s Republic of China. The absence of explicit references to these laws reduces the legal authority of the policies.
Several universities, including Wuhan University, Dalian University of Technology, and Hohai University, show significant shortcomings in the area of talent training as a policy tool. Given the increasing specialization and technical nature of data governance, the lack of talent training means that data management staff may not be equipped with the latest and most effective data analysis methods and tools. Consequently, they might struggle to extract valuable information from vast amounts of data, which could negatively impact decision-making in teaching, research, and administration, thus limiting the overall effectiveness of data governance. For example, a PMC surface plot for Wuhan University, as shown in
Figure 5, indicates significant dips in X4 (policy subject) and X6 (policy foundation). Regarding the policy subject, its composition is relatively narrow, with a severe lack of multi-party participation. Moreover, the failure to explicitly reference relevant laws and regulations in the policy-making process weakens the policy’s authority.
Additionally, universities such as Peking University and Beijing Normal University have deficiencies in their reward and punishment mechanisms. Universities or information technology departments should recognize and reward departments and individuals who perform exceptionally in data governance. The absence of such incentive measures may result in low motivation for departments and individuals to actively participate in data governance, affecting the progress of the data governance initiatives. Furthermore, there is a lack of supervision and penalties for violations of data security regulations. A strict reward and punishment system can motivate data-producing and data-using units to focus on data quality and strengthen data security management. Holding individuals accountable for serious data security violations, such as data leakage or tampering, would effectively safeguard the integrity, confidentiality, and availability of data, ensuring that data governance policies achieve their objectives in securing data.
5.2.3. Evaluation of Moderate Policies
Moderate universities face certain limitations in their data governance policies, necessitating a thorough analysis and clear identification of areas for improvement to enhance the completeness and standardization of the policy framework. The policy content is not comprehensive. For example, the University of International Business and Economics scored relatively low in terms of policy content, with inadequate coverage in areas such as data collection and data operation and maintenance, which partially limits the systematicness of the policy system. From the perspective of policy design standards, data collection should adhere to principles such as authenticity, completeness, standardization, timeliness, and “one data source per data item”. The absence or vague expression of these elements in policy texts indicates gaps in the design of critical operational processes, which in turn affects the comprehensiveness and guidance capacity of the overall data governance policy framework. Additionally, the data operation and maintenance processes are not sufficiently standardized. The lack of clear operational standards and procedures leads to errors in operations such as data updates, corrections, and deletions, as data management personnel lack guidelines, which increases the likelihood of mistakes. Chongqing University has deficiencies in data storage. The information management department should develop a data security storage management plan, with multiple backups of critical data to ensure data security. Moreover, the university should establish a data management platform, providing both software and hardware facilities to ensure the platform’s data is reliable and secure, with regular backups and dynamic storage of shared global data, and support for historical data queries and analysis. Each unit should manage the data generated by its own system maintenance, establishing backup systems and emergency plans. The university’s archive management department should establish a mechanism for archiving historical data, collaboratively building a more complete data storage system. The application of policy instruments and the strength of policy support are insufficient. The talent training system is inadequate, with a lack of systematic training courses on data governance-related technologies and concepts, which leads to gaps in the professional knowledge and skills of data management personnel. Data standards and regulations are updated too slowly to keep up with the new demands of university business development and data management, resulting in poor data compatibility and interoperability, which hinders the circulation and sharing of data. The monitoring and evaluation mechanism is underdeveloped, lacking regular assessments and effective supervision of the implementation of data governance policies, making it difficult to identify and address problems that arise during policy implementation in a timely manner. For instance, Northwestern University and Southeast University have shown poor performance in policy instruments and policy safeguards, leading to inadequate means and implementation support for achieving policy goals.
Taking the University of International Business and Economics as an example, the PMC surface plot is shown in
Figure 6. The PMC surface of moderate policies is primarily concentrated in the range of 0.5–0.8, indicating significant shortcomings and areas for improvement in multiple aspects, such as policy content, policy basis, and policy safeguards. These elements have not reached an ideal state, resulting in an uneven and inefficient PMC surface. Therefore, a comprehensive review and optimization of the policy are necessary to enhance the completeness and coordination of the policy framework across all dimensions.
5.3. Analysis of First-Level Indicator Results
Figure 7 presents the radar diagram of policy mean values. With the exception of the relatively minor concavity in the policy indicators of the policy (policy entities), the mean values of each indicator are above 0.8. This indicates that the data governance policies of the universities are of high quality, particularly excelling in terms of policy nature and content. However, the policy entities still require further improvement.
Regarding policy nature, the average indicator value is 0.95, indicating that most universities perform well in terms of policy nature, covering various types such as descriptive and strategic policies. This suggests that universities focus on diversifying policy nature to meet different data governance needs. However, some universities lack clarity in terms of regulatory aspects in their policies, which limits the effective supervisory role of the policies in data governance and may result in insufficient enforcement during implementation.
Regarding policy objectives, the average indicator value is 0.86, suggesting that universities generally place significant emphasis on setting data governance objectives, covering aspects such as data security and data sharing. This indicates that universities have a certain level of planning and direction in data governance, recognizing the diversity and importance of data governance goals. However, some universities show low attention in certain goal dimensions, with inadequate goal-setting in areas like data sharing and improving data service levels, which hampers the support of data resources for certain institutional activities and decision-making.
Regarding policy content, the average indicator value is 0.94, indicating generally favorable results. Most universities cover a wide range of aspects in their data governance policies, suggesting that universities have a certain level of comprehensiveness in constructing policy content and are able to approach data governance from multiple angles. However, some universities face issues related to the absence of specific norms in data governance processes, such as the lack of regulations for data storage. For example, Renmin University of China lacks clear definitions of data types, resulting in a lack of clear norms and guidance in practical operations.
Regarding policy entities, the average indicator value is 0.5, indicating that most universities’ data governance policies are primarily driven at the institutional level, with insufficient participation from multiple stakeholders. This may result in the lack of diverse perspectives during policy formulation, making it difficult to meet the needs of all parties. Analysis of the policy texts shows that key actors, such as faculty, students, and academic departments, are largely absent, reflecting structural gaps in the completeness of actor participation. A system dominated by a single type of actor can limit the coverage of the policy framework in terms of resource integration and coordinated implementation, thereby affecting the overall systematicity and comprehensiveness of the data governance policy system.
Regarding policy objects, the average indicator value is 0.86, suggesting that universities’ data governance policies cover multiple entities, including data resources, secondary units, and faculty/students. However, there is relatively weak attention given to internal secondary units and faculty/students. Within the university data governance system, these entities are crucial components. Proper attention to secondary units ensures that data governance policies are accurately and effectively implemented in various departments. As producers, users, and managers of data, teachers and students are positioned as key participants in data governance within the policy texts. From the perspective of institutional design completeness, whether the policy clearly defines the roles and behavioral norms of these actors directly impacts the comprehensiveness of the policy framework in the actor dimension. The extent to which these institutional arrangements are reflected in the policy texts forms an important dimension for measuring the completeness of data governance policies and also affects the systematization and standardization of the policy framework when guiding the collaborative participation of diverse actors.
Regarding policy basis, the average indicator value is 0.87, indicating that most universities’ data governance policies are based on clear foundations, mainly grounded in national policies and the universities’ actual circumstances. This reflects the universities’ ability to follow higher-level policy requirements while considering their own practical situations, demonstrating a certain level of rationality and scientific approach. However, some universities lack strong alignment with national policies and institutional circumstances, failing to fully leverage the supporting role of policy foundations in policy formulation.
Regarding policy instruments, the average indicator value is 0.84, indicating some imbalance in the use of policy instruments. Some universities may overly rely on certain tools such as technological support, infrastructure, and strategic measures, while neglecting other tools, such as talent training and outsourcing. Talent training can directly promote data governance goals by addressing gaps in the professional knowledge and skills of data management personnel through training in data governance-related technologies and concepts. Additionally, outsourcing certain services to professional external providers can reduce universities’ investments in human, material, and financial resources, lowering operational costs and enabling a focus on core business and capabilities.
Regarding policy safeguards, the average indicator value is 0.89, indicating that the support measures are generally comprehensive, covering aspects such as organizational structure, departmental collaboration, supervision and evaluation, and reward–punishment mechanisms. This suggests that universities have established a certain system for policy safeguards, providing corresponding support for data governance efforts. However, some universities still have structural gaps in the planning of safeguard measures. This is mainly reflected in incomplete arrangements for personnel allocation and insufficiently developed incentive and constraint mechanisms. Specifically, the ambiguity or absence of provisions for data governance staffing in policy texts indicates a lack of institutional planning for implementing actors and professional coordination mechanisms. The absence or weakness of reward and punishment mechanisms further suggests shortcomings in the design of motivation and behavioral regulation. The incompleteness of these institutional elements, to some extent, limits the systematicness and practicability of the data governance policy framework.
Regarding policy evaluation, the average indicator value is 0.89, suggesting that universities’ data governance policies generally meet the requirements of clear accountability and detailed content. However, there is a need for improvement in terms of goals and policy foundations.
6. Conclusions and Recommendations
This study employs the PMC index model to construct an evaluation index system for data governance policies in higher education institutions. A quantitative evaluation was conducted on data governance policies released by 56 “Double First-Class” universities in China since 2015. The following conclusions were drawn: The overall data governance policies in “Double First-Class” universities are at a good level, with an average PMC index of 7.6. Among them, 25% of the policies are rated as excellent, 66.1% as good, and 8.9% as moderate, and none are classified as poor. This indicates that universities have made progress in building institutional frameworks for data governance policies. The policy system shows notable rationality in both tool allocation and institutional arrangements, providing a structured guide for the standardized implementation of data governance. However, there are significant differences in policy levels across universities. Excellent policies perform well in multiple evaluation dimensions, while good policies need improvement in some areas. Moderate policies exhibit problems such as insufficient reference to policy foundations and imbalanced policy tool structures, indicating room for optimization.
In terms of evaluation indicators, data governance policies in China’s “Double First-Class” universities perform well in X1 (policy nature), X3 (policy content), X6 (policy foundation), and X8 (policy assurance), suggesting that these aspects are comprehensively and consistently considered by the universities, with relatively scientific and effective content design. In contrast, the scores for X2 (policy objectives), X4 (policy subjects), X5 (policy objects), and X7 (policy instruments) are relatively low. The following areas can be considered for advancing data governance in Chinese universities:
Data governance policies in “Double First-Class” universities have an imbalanced allocation of attention to policy objectives. Existing policies overly focus on a specific domain, neglecting the multidimensional needs of data governance, which leads to a narrow policy focus. For example, policies tend to prioritize certain areas (such as educational data, scientific data, or information systems data), emphasizing data sharing while neglecting data security. To ensure the guiding nature of policy objectives, universities should balance their attention allocation when formulating data governance policies. This will help improve the institutional design of the policy framework and provide a more comprehensive data governance system to support the standardized operation of core university activities, such as teaching, research, and administration. First, conduct demand surveys with stakeholders to formulate diversified governance objectives based on governance needs; second, classify and prioritize governance objectives according to the university’s strategic focus and actual needs; third, strengthen inter-departmental communication and collaboration to form governance consensus, and dynamically adjust and optimize governance objectives; fourth, integrate data governance objectives into the overall goals of “Double First-Class” university development to drive the high-quality development of universities with efficient data governance.
- 2.
Constructing a Multilateral Co-governance Framework
There is a significant gap in the recognition of policy subjects in data governance policies at “Double First-Class” universities. The policy subjects are mainly limited to university offices and information departments, i.e., the departments responsible for policy formulation, release, and interpretation, neglecting the important roles played by faculty, students, subordinate units, and departmental data administrators in the data governance process. Strengthening the awareness of multilateral co-governance should become the optimization strategy for policy subjects in universities’ data governance. The policy formulation process should involve multiple stakeholders in decision-making. First, establish a multilateral co-governance concept. Shift from a managerial mindset to a participatory governance mindset, breaking down decision-making barriers, encouraging stakeholders to participate in discussions with an open attitude, integrating different perspectives, listening to various voices, and balancing multiple interests to develop more scientifically grounded and feasible policies. Second, create a multi-stakeholder participation mechanism. Establish a data governance policy formulation committee consisting of university leaders, faculty representatives, student representatives, administrative personnel, data administrators, and technical experts. Clearly define the tasks and responsibilities of participants at each stage, identify governance pain points and practical needs, and set up mechanisms for sorting and providing feedback on opinions. Finally, improve feedback and optimization mechanisms. Establish dynamic feedback channels, maintain continuous communication with stakeholders during the policy formulation process, collect timely feedback on policy drafts, and optimize policies based on this feedback. A policy evaluation group should be formed to comprehensively assess policy drafts, incorporating reasonable suggestions and making necessary adjustments to the policy content.
- 3.
Optimizing the Combination of policy instruments
There is significant room for optimization in the policy instruments of data governance in “Double First-Class” universities. A more balanced policy tool structure, with complementary functions, can better leverage synergies. On the one hand, increasing the proportion of supply-side policy instruments is essential. In addition to providing necessary information support, technical assistance, and infrastructure investments, it is crucial to enhance the cultivation of big data governance talents. As a key element of data governance, universities can provide professional human resources for data governance through talent training programs, such as offering data governance-related courses and organizing specialized training, thereby ensuring a steady supply of skilled professionals to meet the demand for specialized personnel. On the other hand, more attention should be given to demand-side policy instruments, such as encouraging service outsourcing, technical outsourcing, industry–university–research collaboration, and international exchanges, to stimulate data demand. Additionally, environmental policy instruments should be effectively utilized to provide multi-dimensional support, guidance, and facilitation to ensure the successful implementation of data governance in universities.