1. Introduction
The 2030 Agenda for Sustainable Development, adopted by the United Nations General Assembly in September 2015, positioned multistakeholder partnerships as essential governance instruments to bridge implementation gaps left by state-led multilateralism. SDG 17, specifically Target 17.16, calls upon governments, international organizations, civil society, and the private sector to forge collaborative arrangements that mobilize and share knowledge, expertise, technology, and financial resources in pursuit of the global goals [
1]. In the decade that followed, the UN Partnership Platform registered thousands of voluntary commitments. However, scholars have noted that the platform suffers from significant data inconsistencies and a lack of effective monitoring, with many registered initiatives failing to meet minimum partnership criteria [
2]. Indeed, an analysis of the platform reveals an extreme validation asymmetry: fewer than five percent of registered projects have achieved independent validation, designated as Checked-In status.
Throughout this study, we use validated to refer to projects assigned Checked-In status on the UN Partnership Platform, meaning an independent reviewer has completed the platform’s administrative verification process, and non-validated for projects with Checked-Out status, which remain registered but have not undergone this independent confirmation. Checked-Out status does not imply that a partnership is inactive, illegitimate, or fraudulent; it indicates only the absence of independent verification, which may reflect resource constraints in the review process as much as any property of the partnership itself. This validation asymmetry raises a fundamental governance question: to what extent does the language of partnership commitments reflect the platform’s own criteria for multistakeholder collaboration? The concept of SDG washing, whereby organizations project ambitions aligned with the SDGs without undertaking substantive action, has received growing scholarly attention. Heras-Saizarbitoria, et al. [
3] documented cherry-picking among 1370 organizations across 97 countries, finding superficial SDG engagement characteristic of symbolic reporting. van Zanten and van Tulder [
4] argued that many corporate strategies treat individual SDGs as isolated silos rather than as components of a systemic sustainability commitment. Costa, et al. [
5] extended this line of inquiry to a cross-sectoral analysis, distinguishing between SDG walking and SDG washing through discrepancy indices applied to Global Reporting Initiative data. However, these studies focus almost exclusively on corporate sustainability reporting. The analogous phenomenon within the UN partnership ecosystem, in which voluntary commitments may employ partnership language that diverges from independently verified partnership status, has not been examined computationally at scale.
The present study addresses this gap through an unsupervised machine learning analysis of the GENESIS database. Matsui, et al. [
6] demonstrated that BERT-based classifiers can accurately map organizational practices and challenges onto SDG goals, and applications of BERTopic have charted thematic landscapes in sustainable finance literature [
7]. Nevertheless, no prior study has deployed neural topic modeling on the UN Partnership Platform corpus to examine whether discourse patterns distinguish validated from non-validated MSPs.
We make three principal contributions. First, we apply BERTopic, integrating Sentence-BERT embeddings [
8], UMAP dimensionality reduction [
9], and HDBSCAN density-based clustering [
10,
11], to characterize the thematic architecture of UN SDG partnership commitments. Second, we introduce the Validated-Discourse Similarity Index (VDSI). This is a leave-one-out cosine similarity measure that operationalizes the similarity between a project’s textual embedding and the centroid of validated partnerships. Third, we conduct systematic SDG-level subgroup analyses with false discovery rate correction [
12], yielding actionable insights for partnership evaluation practice. The topic-modeling component of this analysis required fixing the target topic count to ensure reproducibility, since automatic topic-count selection proved unstable across runs; even under this fixed target, document-to-topic assignment retains moderate sensitivity to random initialization. The VDSI findings that anchor our central contribution, by contrast, do not depend on the specific topic partition and are stable under resampling. This distinction shapes how the findings should be read: as evidence of a real and policy-relevant pattern, not as a claim that the platform’s textual landscape has been definitively mapped. The remainder of this article is structured as follows.
Section 2 reviews the relevant literature.
Section 3 describes the data and methods.
Section 4 presents the results.
Section 5 discusses the findings.
Section 6 concludes.
2. Literature Review
Multistakeholder partnerships for sustainable development emerged as a mainstream governance instrument through two decades of iterative institutionalization. At the 2002 World Summit on Sustainable Development in Johannesburg, voluntary Type II partnerships were recognized as a complement to intergovernmental outcomes [
13]. Glass, Newig, and Ruf [
2] conducted a systematic survey of 192 MSPs registered on the platform, analyzing governance architecture, partner composition, and SDG coverage, and found that partnerships vary considerably in their institutional depth and functional contribution. While partnerships involving actors from multiple societal sectors are potentially more effective than those involving a single sector, MSPs still have untapped potential to leverage shared resources and capabilities to address complex interactions among the SDGs, particularly those prone to negative spillovers [
2].
The theoretical underpinnings of MSP scholarship draw on at least three intellectual traditions. In global governance theory, partnerships are analyzed as mechanisms that address real or perceived deficits created by traditional multilateral processes. Andonova [
14] demonstrated empirically that international organizations act as governance entrepreneurs, selectively catalyzing public–private partnerships in domains where structural conditions permit coalition formation. Pattberg and Widerberg [
15] identified nine conditions for MSP success, including a clear division of responsibilities, adequate funding, and robust monitoring mechanisms, noting that the evidence base for positive partnership performance remains thin compared with the aspirational discourse surrounding these instruments. Biermann, et al. [
16] outlined the urgent institutional reforms required for effective earth system governance to navigate planetary boundaries, with MSPs playing an increasingly critical role in implementation. In contrast, the earlier analysis by Biermann, et al. [
17] of fragmentation in global governance architectures provides the context within which partnership proliferation generates coordination challenges.
From an institutional theory perspective, partnerships are examined through the lens of organizational sociology, with scholars distinguishing between substantive engagement and symbolic conformity, in which the formal adoption of SDG language serves legitimating purposes without altering organizational behavior. Bäckstrand [
18] argued that the legitimacy, accountability, and effectiveness of multistakeholder partnerships must be assessed through independent criteria rather than self-referential declarations of intent. Westerman, et al. [
19] showed that when employees perceive a gap between espoused sustainability values and enacted practices, cynicism escalates, suggesting that aspiration–realization gaps have consequences that extend beyond external reputational management. In political economy, the voluntary, non-binding nature of UN partnership commitments has been critiqued as a source of accountability deficits, allowing powerful actors to shape global sustainability agendas without democratic oversight [
20]. Widerberg and Pattberg [
21] further argued that overcoming the accountability challenges inherent in transnational governance regimes requires structuring evaluation mechanisms around verifiable performance criteria rather than self-reported commitments.
The phenomenon of SDG washing has attracted increasing theoretical and empirical scrutiny. van Zanten and van Tulder [
4] proposed the nexus approach as a corrective, compelling companies to assess interactions across the full SDG system rather than selectively engaging with goals aligned with existing operations. Heras-Saizarbitoria, Urbieta, and Boiral [
3] documented SDG cherry-picking in 1370 sustainability reports, showing that organizations disproportionately reference SDGs already addressed in their core business activities. Costa, Tiburzi, Morales-Alonso, Calabrese, and Rosati [
5] operationalized this distinction through two complementary indices, demonstrating systematic discrepancies indicative of symbolic reporting in a cross-sectoral sample. None of these studies examined the UN Partnership Platform ecosystem, where the institutional context, comprising voluntary registration, minimal entry requirements, and inconsistent monitoring, may amplify aspiration–realization gaps.
Text mining and natural language processing have emerged as powerful tools for analyzing sustainability discourse at scale. The workhorse of classical topic modeling, Latent Dirichlet Allocation (LDA), introduced by Blei, et al. [
22], models documents as mixtures of latent topics, each characterized by a word-probability distribution. While LDA has been widely applied in sustainability research and SDG-related topic mapping, it treats words as exchangeable bags and cannot capture semantic similarity. The introduction of BERT (Bidirectional Encoder Representations from Transformers) [
23] and Sentence-BERT [
8] overcame this limitation by representing text as dense vectors encoding contextual meaning. BERTopic leverages these embeddings through a modular pipeline that applies UMAP for dimensionality reduction [
9] and HDBSCAN for density-based clustering [
10], extracting topic representations using a class-based term frequency–inverse document frequency procedure. Egger and Yu [
24] benchmarked BERTopic against LDA, non-negative matrix factorization, and Top2Vec on Twitter data, finding that BERTopic offers the greatest potential among embedding-based models for generating novel insights and achieving high interpretability in short-text social science contexts.
Applications of BERTopic to sustainability and governance topics have expanded rapidly. Raman, Ray, Das, and Nedungadi [
7] employed BERTopic to map the landscape of sustainable and green finance literature onto SDG clusters. Matsui, Suzuki, Ando, Kitai, Haga, Masuhara, and Kawakubo [
6] demonstrated that BERT-based semantic mapping can classify organizational sustainability practices to SDG goals with high accuracy. Lee, et al. [
25] applied BERTopic to compare academic and media framings of environmental, social, and governance themes. These studies establish the methodological viability of neural topic modeling for sustainability governance analysis but have not been directed at the UN Partnership Platform corpus.
The SDG-level heterogeneity of partnership engagement has received attention. van Zanten and van Tulder [
4] documented asymmetric corporate engagement across the SDGs, noting that companies tend to prioritize economic growth and industrialization while systematically under-representing goals focused on the biosphere. Similarly, Glass, Newig, and Ruf [
2] showed that climate action, quality education, and gender equality attract disproportionate MSP attention. Furthermore, Andonova, et al. [
26] demonstrated at the national policy level that transnational governance arrangements and domestic policies function as complements rather than substitutes, underscoring that structural institutional investment amplifies rather than replaces formal commitment. This pattern of symbolic alignment without corresponding substantive change echoes a broader family of ‘decoupling’ phenomena documented across organizational and environmental governance research: greenwashing, bluewashing (symbolic UN Global Compact membership without behavioral change), and more generally the gap between adopted policy and implemented practice that institutional theorists have long identified as a feature of organizations operating under external legitimacy pressure. What distinguishes the present setting is that the UN Partnership Platform is not a corporate disclosure regime but a multilateral governance infrastructure: the “audience” being addressed through partnership language is not shareholders or consumers but the intergovernmental process itself, and the accountability mechanism (independent validation) is built into the platform rather than externally imposed. This positions VDSI as a measure of the discourse-level dimension of decoupling within a governance-accountability system that already possesses, but underutilizes, a verification mechanism, rather than as a call for external regulation of a previously unaccountable domain. These threads point to a common construct: what Bäckstrand [
18] terms accountability through independent criteria, Widerberg and Pattberg [
21] frame as evaluation grounded in verifiable performance rather than self-reported commitments, and Westerman, Acikgoz, Nafees, and Westerman [
19] link to the erosion of legitimacy when espoused values diverge from enacted practice. Read together, these strands define governance accountability as the requirement that legitimacy claims be anchored in externally verifiable performance rather than self-description, a requirement the UN Partnership Platform formally provides through its validation mechanism but, as our findings show, substantially underutilizes. Our analysis extends these observations to the full GENESIS corpus, examining whether the structural advantage of validated MSPs varies systematically across SDGs.
While the preceding discussion situates governance accountability within multistakeholder partnership scholarship, this construct has been most extensively theorized in the corporate sustainability reporting literature, from which the UN Partnership Platform’s validation mechanism can usefully be read as a structural analogue. Cho, et al. [
27] argue that firms facing conflicting stakeholder and institutional pressures are structurally compelled toward organized hypocrisy, in which sustainability talk and organizational façades substitute for substantive practice change without this necessarily reflecting deliberate deception. This concept parallels the interpretive caution this study adopts toward the Inconsistent Non-Validated finding: symbolic alignment can arise as a structural response to institutional pressure rather than as evidence of intentional washing. Mandatory disclosure regimes were developed partly to counter this tendency; Christensen, et al. [
28] review the economic evidence on mandated sustainability disclosure, noting that frameworks such as the EU’s Non-Financial Reporting Directive shift firms from voluntary, self-selected disclosure toward standardized, externally structured reporting, though they characterize the causal evidence on whether such mandates improve substantive outcomes as still scarce. Mezzanotte [
29] shows that even under the EU’s subsequent Corporate Sustainability Reporting Directive and its double-materiality framework, the assessment of which impacts are material remains substantially discretionary, creating legal and interpretive uncertainty analogous to the ambiguity this study identifies in the UN Partnership Platform’s undocumented Checked-In criteria. Complementary reporting mechanisms have shown mixed capacity to close this gap: de Villiers, et al. [
30], synthesizing research from an AAAJ special issue on integrated reporting, note that its adoption has in some documented cases been driven as much by ongoing legitimacy struggles as by the substantive integration of financial and non-financial performance narratives it was designed to achieve. Most directly relevant to the present findings, Roszkowska-Menkes, et al. [
31] show empirically that neither adherence to Global Reporting Initiative guidelines nor third-party assurance meaningfully reduces selective disclosure of negative sustainability events, indicating that even codified, externally verifiable transparency mechanisms do not reliably convert a validation infrastructure into an assurance of substantive accountability. This body of work suggests that the UN Partnership Platform’s own validation mechanism, though similar in design to these corporate transparency instruments, is likely to be similarly limited in its capacity to ensure that discourse patterns reliably track independently verified partnership status.
6. Conclusions
The analysis is situated within a broader accountability frame articulated by Pogge and Sengupta [
43], who argued that SDG implementation must be assessed not merely against aspirational commitments but against the structural conditions that enable rights-consistent delivery. This study applied BERTopic neural topic modeling to 3807 project descriptions in the GENESIS WP4 Multistakeholder Partnerships Database [
32], introduced the Validated-Discourse Similarity Index as a novel measure of semantic alignment with validated-partnership discourse, and conducted systematic statistical comparisons between validated and non-validated partnerships.
Four principal conclusions emerge. First, the thematic architecture of UN SDG partnership discourse is highly concentrated, with approximately 61% of the corpus clustering within a Climate, Sanitation, and Access domain, alongside secondary marine/coastal and sustainable-textile domains. Second, a pervasive Inconsistent Non-Validated phenomenon characterizes the partnership ecosystem: 92.9% of registered projects employ language resembling that of independently validated partnerships, though raw semantic similarity discriminates validation status only modestly (AUC = 0.686), indicating a genuine but partial signal rather than indistinguishability between groups, a finding nonetheless robust across sensitivity analyses spanning a wide range of thresholds. This constitutes among the first corpus-scale empirical documentation of discourse-similarity patterns relative to validation status in the UN Partnership Platform and extends SDG-washing scholarship beyond corporate reporting into the arena of multilateral voluntary commitments. Third, partner count is the most consistent structural discriminator of validated partnerships (r = −0.389, medium effect; adjusted OR = 1.38 per standard-deviation increase after controlling for SDG scope, description length, duration, topic, and language), persisting across all six SDG subgroups examined. Fourth, an initial multi-seed and multi-parameter check revealed that BERTopic’s automatic topic-count selection was itself unstable (topic count ranging from 3 to 8 under nominally identical hyperparameters); tracing this to the automatic topic-merging step and fixing the target topic count resolved it, and a follow-up robustness check confirmed stable topic counts with moderate assignment-level agreement (mean ARI = 0.64), a methodological refinement we report transparently, together with the instability that motivated it, rather than treating automatic topic-count selection as reliable by default.
Platform administrators and accountability bodies may find partner count a useful preliminary signal to explore in validation outreach strategies, though the present analysis does not establish an operational cutoff: the observed mean partner count for non-validated projects is 4.64, while validated projects average 13.30, but sensitivity, specificity, and predictive value for any specific threshold remain unassessed. Whether partnerships reporting fewer than five partners warrant additional scrutiny is therefore best treated as a preliminary hypothesis requiring prospective, out-of-sample validation before any operational use, rather than as a ready-to-deploy screening rule. Any future algorithmic triage system built on such a threshold should be designed as an advisory tool rather than an automatic gate, with clear appeal mechanisms and human oversight, and with attention to the risk that genuine small-coalition partnerships, particularly those from under-resourced contexts in the Global South, may be disadvantaged. The VDSI is most appropriately deployed as a retrospective audit instrument or as a formative feedback tool provided to applicants at the point of registration rather than as an opaque filter, so that organizations can identify areas of low semantic alignment with validated-partnership discourse and strengthen their partnership architecture before seeking validation. In practice, VDSI is best understood as an early screening signal rather than a substitute for substantive assessment: projects whose discourse closely resembles that of already validated partnerships could be assigned lower review priority, allowing UN Partnership Platform administrators to concentrate limited verification capacity on projects whose language diverges most from the validated reference group. This positions VDSI within the broader responsible digital innovation agenda: transparency of the scoring logic, auditability of the model, and equitable access to feedback are essential governance requirements for any NLP-based platform accountability system. Future research should develop multilingual VDSI models [
36] to include non-English partnership documentation, an extension made more pressing by our finding that language-artifact documents differ systematically from the English-language corpus in validation rate and partnership-type composition (
Section 4.6), extract geographic and sectoral metadata via named-entity recognition to enable representativeness checks that the current dataset’s structure precludes, and apply dynamic topic modeling to track the temporal evolution of partnership discourse as the 2030 deadline approaches, extending the static year-level analysis reported here (
Section 4.6). Future applications of BERTopic to governance-accountability questions should fix the target topic count explicitly rather than relying on automatic topic-count selection and report multi-seed stability diagnostics under that fixed target, given the instability specifically traced to automatic selection in
Section 4.8. As the international community confronts a significant shortfall against SDG targets, the capacity to distinguish partnerships that mobilize genuine multi-sector coalitions from those that appropriate partnership language for legitimization has become a governance imperative [
5,
44]. Stibbe and Prescott [
33] offer a normative framework for building high-impact partnerships; our findings provide the empirical evidence base for evaluating departures from that standard at corpus scale.