Next Article in Journal
TransformerPIV: An Improved Large-Scale Flow Motion Estimation Method Based on Self-Attention Mechanism
Next Article in Special Issue
Stochastic Semantic Fields for Sentiment-Driven Models
Previous Article in Journal
Federated Edge Intelligence for Climate-Aware Spatiotemporal Road Accident Prediction Using IoT and LoRaWAN Networks
Previous Article in Special Issue
Evaluating Pre-Trained Transformer-Based Models for Political Sentiment Analysis on Social Media
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Computational Modelling of Institutional Perceptions in a Civilisational Intergovernmental Organisation

by
Fahim Sufi
1,2,*,
A. K. M. Iftekharul Islam
3 and
Anowara Akter
4
1
COEUS Institute, New Market, VA 22844, USA
2
Centre for Trade & Investment, University of Dhaka, Dhaka 1000, Bangladesh
3
Faculty of Business, Law and Politics, University of Hull, Bloomsbury, London WC1E 6AA, UK
4
Department of Islamic History and Culture, Jagannath University, Dhaka 1100, Bangladesh
*
Author to whom correspondence should be addressed.
Computation 2026, 14(7), 164; https://doi.org/10.3390/computation14070164
Submission received: 30 May 2026 / Revised: 9 July 2026 / Accepted: 18 July 2026 / Published: 21 July 2026

Abstract

Civilisational intergovernmental organisations occupy a distinctive position in global governance because their legitimacy is shaped by institutional performance, symbolic identity, collective representation, and normative expectation. This study develops a sentiment-driven computational modelling framework for analysing institutional perceptions in a 57-member civilisational intergovernmental organisation. The empirical corpus comprises 30 semi-structured interviews, including 15 expert interviews and 15 general respondent interviews. After preprocessing, removal of administrative material, and exclusion of courtesy-only responses, the final corpus contained 239 answer segments and 42,940 respondent-generated words. The methodology integrates answer-level segmentation, thematic classification, lexical sentiment analysis, transformer-based sentiment modelling, latent semantic clustering, non-parametric statistical testing, bootstrap confidence estimation, and manual validation. The results show that institutional perception is broadly positive but thematically uneven. Experts recorded a mean polarity score of 0.0803, while general respondents recorded 0.1016. However, group-level differences were not statistically significant at segment level, Mann–Whitney U = 6171.50 , p = 0.0704 , or interview level, U = 71.00 , p = 0.0888 . In contrast, sentiment differed significantly across institutional themes, Kruskal–Wallis H = 45.8768 , p < 0.001 , and across latent semantic clusters, with Kruskal–Wallis H = 35.0124 and p < 0.001 . The strongest positive sentiment appeared in reform and future orientation, mean polarity = 0.1700 , socioeconomic cooperation, 0.1330 , and institutional effectiveness, 0.1325 . The weakest sentiment concerned institutional weakness and constraints, 0.0546 , and political and security role, 0.0693 . Latent semantic clustering identified socioeconomic and economic cooperation as the most positively evaluated cluster, with mean polarity 0.1465 . Overall, the findings reveal a measurable tension between symbolic-developmental legitimacy and operational scepticism.

1. Introduction

Civilisational intergovernmental organisations present a distinctive analytical problem because their perceived legitimacy is not produced solely through institutional performance, policy delivery, or formal procedural design. It is also shaped by symbolic identity, collective representation, normative expectation, and cultural affiliation. Existing scholarship on international organisational legitimacy has shown that performance, representation, procedure, and public confidence are central to how international organisations are evaluated [1,2,3,4]. At the same time, computational social science has developed mature methods for modelling sentiment, polarity, and latent thematic structure from text [5,6,7,8]. Lexicon-based models offer transparent sentiment measurement [9], transformer-based models such as BERT and RoBERTa capture contextual semantic dependencies [10,11], and topic modelling enables the extraction of latent discourse clusters [7,12], while latent semantic clustering more precisely describes TF-IDF and latent semantic analysis representations grouped through clustering rather than probabilistic topic inference. Yet, limited work has connected these computational methods with the study of institutional perception in civilisational intergovernmental organisations using interview-based evidence. This creates a clear research gap: legitimacy studies remain conceptually rich but often computationally limited, while sentiment-driven modelling studies are technically advanced but insufficiently grounded in institutional theory.
In this study, a civilisational intergovernmental organisation is understood as a formal multilateral organisation whose legitimacy is shaped not only by state-based cooperation and institutional performance, but also by shared civilisational identity, cultural affiliation, collective representation, and normative expectation. Analysing sentiment toward such an organisation is scientifically relevant because its legitimacy cannot be assessed only through formal mandates, policy outputs, or organisational structure. Perceptions of trust, symbolic representation, operational credibility, reform expectation, and institutional scepticism are expressed through evaluative language. Sentiment analysis is therefore not used merely because the technique is available; it is used because it provides a systematic way to measure how respondents linguistically evaluate different institutional functions. The study is guided by three analytical expectations: first, perceptions will vary more strongly across institutional functions than across respondent groups; second, socioeconomic, developmental, and reform-oriented themes will attract more positive sentiment than political-security and constraint-oriented themes; and third, expert respondents may express more cautious evaluations than general respondents because of their greater familiarity with institutional limitations.
To address this gap, this study develops a sentiment-driven computational modelling framework for analysing institutional perceptions in a 57-member civilisational intergovernmental organisation. The empirical corpus consists of 30 semi-structured interviews, including 15 expert interviews and 15 general respondent interviews. After preprocessing, removal of administrative material, and exclusion of courtesy-only responses, the final analytical corpus contained 239 answer segments and 42,940 respondent-generated words. The expert corpus contributed 121 segments and 24,333 words, whereas the general respondent corpus contributed 118 segments and 18,607 words. The methodological innovation lies in integrating answer-level segmentation, similarity-based thematic classification, lexical sentiment analysis, transformer-based contextual sentiment modelling, latent semantic clustering, non-parametric statistical testing, bootstrap confidence estimation, and manual validation. This design allows institutional perception to be treated not as an impressionistic judgement, but as a multidimensional computational construct shaped by polarity, subjectivity, thematic membership, latent semantic cluster membership, respondent category, and evaluative intensity.
The results indicate that institutional perception is broadly positive but thematically uneven. General respondents recorded a higher mean polarity score, 0.1016, than expert respondents, 0.0803; however, this difference was not statistically significant at either segment level, Mann–Whitney U = 6171.50 , p = 0.0704 , or interview level, U = 71.00 , p = 0.0888 . The sentiment label distribution also showed no significant group-level association, χ 2 = 3.0990 , p = 0.2124 , Cramer’s V = 0.1139 . In contrast, sentiment differed significantly across institutional themes, Kruskal–Wallis H = 45.8768 , p < 0.001 , and across latent semantic clusters, with Kruskal–Wallis H = 35.0124 and p < 0.001. Reform and future orientation generated the highest mean polarity score, 0.1700, followed by socioeconomic cooperation, 0.1330, and institutional effectiveness, 0.1325. The weakest sentiment was associated with institutional weakness and constraints, 0.0546, and political and security role, 0.0693. Latent semantic clustering further identified socioeconomic and economic cooperation as the most positively evaluated semantic cluster, with a mean polarity of 0.1465.
The significance of this study lies in demonstrating that perceptions of a civilisational intergovernmental organisation are better explained by the institutional function under evaluation than by respondent category alone. The findings reveal a measurable perception architecture in which symbolic-developmental legitimacy coexists with operational scepticism. Respondents express comparatively stronger positive sentiment toward reform potential, socioeconomic cooperation, and institutional effectiveness, while displaying more restrained evaluations of political-security functions, visibility gaps, and institutional constraints. The study therefore makes four principal contributions:
  • It develops a reproducible computational framework for modelling institutional perceptions from interview narratives.
  • It quantifies institutional perception using 239 answer segments and 42,940 respondent-generated words.
  • It demonstrates that thematic and cluster-level differences are statistically stronger than expert–general group differences.
  • It provides a transferable computational approach for studying legitimacy, trust, perceived effectiveness, and scepticism in international, regional, and identity-based multilateral organisations.

2. Background

Computational analysis of textual data has become an established approach for studying attitudes, perceptions, and evaluative judgments in social, political, and organisational contexts. Sentiment analysis provides one of the most widely used foundations for this task because it converts subjective language into measurable polarity and affective orientation [5,6]. Lexicon-based models, such as VADER, remain useful because they provide transparent and interpretable baseline measures of textual sentiment [9]. However, institutional discourse is often indirect, cautious, and context-dependent. For this reason, transformer-based language models, including BERT and RoBERTa, have become increasingly important because they capture contextual semantic relationships that are difficult to detect through dictionary-based methods alone [10,11].
A second stream of relevant work concerns topic modelling, semantic clustering, and automated text analysis. Early probabilistic models, particularly latent Dirichlet allocation, established topic modelling as a method for extracting latent thematic structures from large textual corpora [7]. More recent embedding-based approaches, such as BERTopic, extend this logic by combining semantic representation with interpretable topic extraction [12]. The present study does not treat its clusters as probabilistic topics; instead, it uses TF-IDF representations projected through latent semantic analysis and grouped through k-means clustering as interpretable latent semantic clusters. In political and policy research, automated text analysis is increasingly used not as a replacement for qualitative interpretation, but as a means of scaling and formalising it [8,13]. This point is particularly relevant for interview-based research, where the central challenge is to preserve interpretive nuance while producing reproducible analytical outputs.
A third body of literature examines the legitimacy and perception of international organisations. Existing studies show that institutional legitimacy is shaped by performance, procedural fairness, representation, and public confidence [1,2,3,4]. These works are important because they indicate that perceptions of international organisations cannot be reduced to formal effectiveness alone. They are also shaped by identity, symbolic authority, and expectations regarding institutional behaviour. For civilisational or identity-based organisations, this issue becomes even more pronounced, since legitimacy may derive from collective representation as much as from policy implementation.
The present study also builds on prior computational work, where sentiment analysis, entity extraction, regression modelling, GPT-integrated classification, and media bias quantification were used to analyse large-scale textual corpora [14,15]. These studies provide a methodological precedent for treating textual narratives as structured computational evidence. The present paper extends that logic from media and cyber intelligence corpora to interview-based institutional perception modelling. Table 1 contextualizes this study with existing studies.
Despite these advances, limited work has combined sentiment analysis, transformer-based modelling, latent semantic clustering, statistical testing, and manual validation to study perceptions of civilisational intergovernmental organisations using interview data. The present study addresses this gap by proposing a computational thematic sentiment modelling framework that transforms qualitative institutional narratives into measurable perception signals. In doing so, it connects computational social science with the study of international organisations, offering a replicable approach for analysing symbolic legitimacy, institutional effectiveness, and operational scepticism.

3. Methodology

This study adopts a computational thematic sentiment modelling framework to quantify institutional perceptions expressed in semi-structured interview narratives. The schematic of the conceptual architecture is summarised in Figure 1. The framework begins with raw interview transcripts, proceeds through corpus construction, segmentation, thematic classification, sentiment modelling, latent semantic clustering, statistical comparison, robustness assessment, and manual validation, and culminates in an interpretive institutional perception synthesis.
The methodological rationale follows three principles. First, interview narratives are treated as structured textual evidence rather than as simple opinion statements. Second, sentiment is modelled at multiple levels, including lexical polarity, contextual transformer-based sentiment, thematic sentiment, and cluster-level sentiment. Third, computational outputs are validated through manual coding because sentiment models may misinterpret politically nuanced discourse, indirect criticism, or mixed evaluations. Lexicon-based methods provide transparent baseline sentiment indicators [5,9], whereas transformer-based models capture contextual semantic dependencies more effectively [10,11]. Latent semantic clustering is used to identify perception clusters in the corpus. It is reported as clustering rather than conventional probabilistic topic modelling because the procedure uses TF-IDF and latent semantic analysis representations followed by k-means clustering. Non-parametric statistical tests are selected because sentiment distributions derived from qualitative interview segments cannot be assumed to follow normality [17,18]. Bootstrap confidence intervals are used to estimate uncertainty around theme-level and cluster-level mean sentiment scores [19].

3.1. Data Collection and Interview Corpus

The empirical dataset was constructed from 30 semi-structured interviews conducted with two respondent groups: 15 general respondents and 15 expert respondents. A purposive sampling strategy was adopted to identify participants with academic, professional, or institutional familiarity with international organisations and the role of the 57-member civilisational intergovernmental organisation examined in this study. General respondents primarily included PhD fellows, postgraduate researchers, and early-career scholars in international relations, political science, and related fields. Expert respondents included university academics, senior researchers, policy professionals, officials, and individuals with professional exposure to international cooperation or related institutional settings. Interviews were conducted through both face-to-face and online formats, depending on participant location and availability. Respondents were geographically dispersed across Muslim-majority contexts, including Bangladesh, Pakistan, Turkey, and Malaysia, as well as Muslim-minority contexts, including the United States, the United Kingdom, and Romania. All interviews followed a semi-structured protocol designed to maintain consistency across core questions while allowing respondents to elaborate on institutional legitimacy, governance capacity, cooperation, communication, constraints, and reform.
The use of 30 interviews is justified by the purposive and criterion-based design of the study. Since the organisation examined here is a specialised intergovernmental institution, meaningful commentary required prior academic, professional, or institutional familiarity with international organisations and related political contexts. The sample was therefore organised into two purposively differentiated tiers: 15 general respondents with postgraduate or early-career academic grounding in history, international relations, political science, or related fields, and 15 expert respondents, including researchers, faculty members, policy professionals, officials, journalists, or individuals with direct professional exposure to the organisation. This structure supports analytical comparison between theoretically informed and practitioner-oriented perspectives. In qualitative inquiry, sample sufficiency is judged less by numerical representativeness than by information richness, relevance to the research question, and thematic saturation [20,21,22]. Prior saturation research indicates that core themes can often be identified within relatively small (as little as only 12) purposive samples, particularly when participants are information-rich and the topic is bounded [23]. Indeed, Guest et al. (2006) demonstrated in their landmark study that saturation can often be achieved with as few as 12 interviews in a homogeneous purposive sample [23]. A sample of 30 with two distinct expert tiers substantially surpasses this threshold. In the present study, each interview exceeded two hours, producing more than 60 h of interview material and 42,940 respondent-generated words after cleaning. The dataset is therefore treated as sufficient for exploratory and theory-guided computational interpretation, while the manuscript avoids claims of broad population-level generalisation.

3.2. Notation

Table 2 defines the principal notation used in the methodological formulation.

3.3. Corpus Construction and Preprocessing

Let the full corpus be represented as
D = { d 1 , d 2 , , d N } ,
where N = 30 . The corpus consists of two respondent groups:
D = D E D G , D E D G = ,
with
| D E | = 15 , | D G | = 15 .
Each raw transcript d i r a w was transformed into a cleaned document d i c l e a n through a preprocessing operator:
d i c l e a n = f c l e a n ( d i r a w ) .
The cleaning function removed non-analytical textual components:
d i c l e a n = d i r a w { C , I , H , A , Q } ,
where C denotes consent material, I denotes repeated instructions, H denotes headers, A denotes administrative text, and Q denotes interviewer questions. The retained content therefore consists only of respondent-generated discourse:
d i c l e a n = R i .

3.4. Answer-Level Segmentation

Each cleaned transcript was segmented into answer-level analytical units:
d i c l e a n = { s i 1 , s i 2 , , s i m i } ,
where m i is the number of retained answer segments in document d i . The full segment-level corpus is then
S = i = 1 N { s i 1 , s i 2 , , s i m i } .
After removing courtesy-only and non-substantive segments, the final segment corpus contained:
| S | = M = 239 .
The total number of respondent-generated words was
W = j = 1 M w j = 42,940 .
Each segment is represented as
s j = ( x j , g j , w j ) ,
where x j is the textual content, g j is the respondent group, and  w j is the word count.

3.5. Thematic Classification

The study defines a set of institutional perception themes:
T = { T 1 , T 2 , , T K } ,
where K = 6 . The themes are symbolic legitimacy and identity, institutional effectiveness, political and security role, socioeconomic cooperation, institutional weakness and constraints, and reform and future orientation.
Each segment was assigned to an institutional theme using a computational similarity function:
θ j = f t h e m e ( s j ) , θ j T .
To operationalise this, each segment and each theme descriptor were transformed into TF-IDF vectors. Let v j denote the TF-IDF vector for segment s j , and let q k denote the TF-IDF vector for theme T k . The cosine similarity between segment s j and theme T k is
sim ( s j , T k ) = v j · q k v j q k .
The assigned theme is therefore
θ j = arg max T k T sim ( s j , T k ) .
This procedure provides a transparent bridge between qualitative theoretical categories and computational classification. It avoids treating themes as purely emergent machine clusters and instead constrains the classification according to theoretically meaningful institutional dimensions.

3.6. Validation Design for Thematic Classification

To improve transparency in the thematic classification procedure, the six institutional theme descriptors used for TF-IDF cosine similarity matching are reported in Table 3. These descriptors were constructed from three sources: the conceptual framing of institutional legitimacy in the literature, the semi-structured interview protocol, and an initial close reading of the interview corpus. The descriptors were not used as outcome labels in a supervised model; rather, they served as theoretically anchored semantic anchors against which each answer segment was compared. This procedure was intended to preserve conceptual interpretability while allowing systematic and reproducible segment-level classification.
To assess the reliability of the automated thematic assignment, a validation subset of approximately 20% of the analytical segments was selected using stratified sampling across respondent group and computationally assigned theme. Given the final corpus size of 239 answer segments, the validation subset contained 48 segments. Two independent coders manually assigned each sampled segment to one of the six institutional themes using the definitions in Table 3. Disagreements between coders were resolved through discussion, and the reconciled manual labels were used as the reference labels for evaluating the TF-IDF cosine similarity classification.
Classification performance was evaluated using accuracy, macro-precision, macro-recall, and macro-F1. Accuracy was calculated as the proportion of validation segments for which the automated theme label matched the reconciled manual label. Macro-F1 was used because the theme distribution was uneven and because smaller categories, such as reform and future orientation, should not be overwhelmed by larger categories. Intercoder reliability was assessed using Cohen’s kappa before reconciliation. The validation metrics and confusion matrix are reported in the Results section.

3.7. Lexical Sentiment Modelling

Lexical sentiment analysis was used as the first sentiment estimation layer because it provides interpretable baseline polarity scores. For each segment s j , the lexical sentiment model returns
L j = ( p j , n j , u j , c j , b j ) ,
where p j is positive sentiment, n j is negative sentiment, u j is neutrality, c j is compound polarity, and  b j is subjectivity.
The compound polarity score is formalised as
c j = p j n j p j + n j + u j + ϵ ,
where ϵ > 0 is a small stabilising constant. The group-level mean polarity is
c ¯ g = 1 | S g | s j S g c j , g { E , G } .
Similarly, theme-level mean polarity is computed as
c ¯ T k = 1 | S T k | s j S T k c j ,
where
S T k = { s j : θ j = T k } .

3.8. Transformer-Based Sentiment Modelling

Lexical approaches are limited when applied to political and institutional discourse because sentiment may be expressed indirectly, cautiously, or in mixed form. To address this limitation, a transformer-based classifier was incorporated as a contextual sentiment layer. Transformer models are appropriate here because they learn bidirectional contextual representations and can capture phrase-level and sentence-level semantic dependencies [10,11].
For each answer segment s j , the transformer classifier estimates a probability distribution over sentiment labels:
P ( y j | s j ) = f t r a n s ( s j ) ,
where
y j { positive , neutral , negative } .
The predicted sentiment label is
y ^ j = arg max y Y P ( y | s j ) .
The transformer confidence score is defined as
γ j = max { P p o s , j , P n e u , j , P n e g , j } .
A continuous transformer polarity score is calculated as
τ j = P p o s , j P n e g , j , τ j [ 1 , 1 ] .

3.9. Lexical-Transformer Agreement

To compare lexical and transformer outputs, the lexical compound score was mapped into a categorical sentiment label:
z ^ j = positive , c j > δ , negative , c j < δ , neutral , δ c j δ ,
where δ is the neutrality threshold. Model agreement is then
A L T = 1 M j = 1 M I ( z ^ j = y ^ j ) ,
where I ( · ) is the indicator function. This measure evaluates the consistency between the transparent lexical baseline and the context-sensitive transformer model.

3.10. Latent Semantic Clustering

Latent semantic clustering was used to identify interpretable semantic groupings in the interview corpus. The procedure is not treated as conventional probabilistic topic modelling. Instead, answer segments were represented through TF-IDF features, projected into a lower-dimensional latent semantic space, and grouped using k-means clustering. Each segment was transformed into a numerical representation:
e j = f l s a ( s j ) , e j R d .
The full latent semantic matrix is
E = e 1 e 2 e M .
A clustering function then assigns each segment to a latent semantic cluster:
z j = f c l u s t e r ( e j ) , z j { 1 , 2 , , K z } .
In this study, six latent semantic clusters were retained as an interpretability-driven compromise rather than as a uniquely optimal statistical solution. Although some diagnostic criteria were stronger for neighbouring values of K, the six-cluster solution provided the clearest balance between interpretability, thematic separation, cluster-size balance, stability, and comparability with the institutional dimensions under investigation.
For each semantic cluster z, the cluster-level polarity is
c ¯ z = 1 | S z | s j S z c j ,
where
S z = { s j : z j = z } .
Representative terms were extracted from the highest-weighted TF-IDF features within each cluster, and semantic labels were assigned through joint inspection of these representative terms and exemplar answer segments. This allows the study to determine which latent perception clusters generate the strongest positive, neutral, or restrained evaluative patterns.

3.11. Statistical Testing

The primary group-level comparison tests whether expert and general respondents differ in their segment-level sentiment distributions. The null hypothesis is:
H 0 : F E ( c ) = F G ( c ) ,
where F E ( c ) and F G ( c ) denote the sentiment score distributions for expert and general respondents. The alternative hypothesis is
H 1 : F E ( c ) F G ( c ) .
Because normality cannot be assumed for interview-derived sentiment scores, the Mann–Whitney U test was used:
U E = n E n G + n E ( n E + 1 ) 2 R E ,
where n E and n G are group sample sizes and R E is the rank sum for the expert group.
For Mann–Whitney comparisons, effect size was reported as the rank-biserial correlation, denoted by r r b . Using the Mann–Whitney statistic for the expert group, the rank-biserial correlation was calculated as:
r r b = 2 U E n E n G 1 ,
where U E is the Mann–Whitney statistic for the expert group, and  n E and n G are the expert and general respondent sample sizes. The statistic ranges from 1 to + 1 . A negative value indicates that the expert group tends to have lower sentiment ranks than the general respondent group, whereas a positive value indicates higher expert-group ranks. The magnitude was interpreted using the absolute value of r r b , with values below 0.10 treated as negligible, values around 0.10 0.30 as small, values around 0.30 0.50 as small to moderate or moderate, and values above 0.50 as large.
Because the analytical corpus consists of multiple answer segments nested within 30 interviews, the interview rather than the segment was treated as the primary sampling unit for robustness interpretation. Segment-level tests were retained because thematic and sentiment labels were assigned at the answer-segment level; however, they were interpreted as within-corpus evidence rather than as independent population-level observations. To reduce the risk of pseudo-replication, group-level sentiment comparison was additionally repeated at the interview level by aggregating segment polarity scores within each interview. For interview i, the mean interview-level polarity was calculated as
c ¯ i = 1 m i j = 1 m i c i j ,
where m i is the number of retained answer segments in interview i and c i j is the polarity score of segment j in interview i. The expert and general respondent groups were then compared using the resulting 30 interview-level mean polarity scores. Accordingly, segment-level results are reported as fine-grained textual evidence, while interview-level results are used as a robustness check against inflated inference due to nested segments.
To further address the nested structure of the data, an interview-level cluster bootstrap was used as a robustness check for the main group-level comparison. Rather than resampling individual answer segments, the bootstrap procedure resampled interviews with replacement within respondent groups. For each bootstrap iteration, all retained segments belonging to the selected interviews were included, and the mean polarity difference between expert and general respondents was recalculated. This procedure preserves the within-interview dependence structure and avoids treating all 239 segments as fully independent observations. The cluster bootstrap was repeated with 1000 resamples, and 95% confidence intervals were obtained from the 2.5th and 97.5th percentiles of the bootstrap distribution.
To test whether sentiment differs across themes, the Kruskal–Wallis statistic is
H = 12 M ( M + 1 ) k = 1 K R k 2 n k 3 ( M + 1 ) ,
where R k is the rank sum for theme T k and n k is the number of segments assigned to that theme.
For categorical sentiment label distributions, a chi-square test was applied:
χ 2 = r = 1 R c = 1 C ( O r c E r c ) 2 E r c ,
where O r c and E r c are the observed and expected frequencies in row r and column c. Cramer’s V was used as an effect size measure:
V = χ 2 M min ( R 1 , C 1 ) .
Bootstrap confidence intervals were computed for theme-level and cluster-level mean sentiment:
C I 95 % ( c ¯ ) = Q 0.025 ( c ¯ * ) , Q 0.975 ( c ¯ * ) ,
where c ¯ * denotes bootstrap-resampled mean polarity estimates.

3.12. Manual Validation

A manual validation sample was selected to assess the reliability of computational sentiment outputs. The validation subset is
V S , | V | 0.20 M .
Given M = 239 , the validation set contains approximately:
| V | 48 .
Each sampled segment was manually coded as
m j { positive , negative , neutral , mixed } .
Agreement between manual labels and transformer predictions is
A M T = 1 | V | s j V I ( m j = y ^ j ) .
Agreement between manual labels and lexical predictions is
A M L = 1 | V | s j V I ( m j = z ^ j ) .
where required, Cohen’s kappa may be used to adjust for chance agreement:
κ = p o p e 1 p e ,
where p o is observed agreement and p e is expected agreement by chance [24].
The same validation subset was also used to compare manually assigned sentiment labels with the lexical and transformer-based sentiment predictions, allowing manual–lexical agreement, manual–transformer agreement, Cohen’s kappa, and disagreement patterns to be assessed empirically.

3.13. Composite Institutional Perception as Interpretive Formalisation

The final modelling layer does not estimate a predictive regression model. Instead, it provides an interpretive formalisation of how the different computational outputs of the study can be understood as components of institutional perception. This clarification is important because the present corpus contains 30 interviews and 239 answer-level segments, which is appropriate for exploratory computational interpretation but not for stable estimation of a multi-parameter predictive model. Accordingly, the following expression should be read as a conceptual synthesis of the analytical dimensions used in the study, rather than as an empirically calibrated equation.
For each segment s j , institutional perception is represented as a multidimensional construct:
I P j = F ( c j , τ j , Θ j , Z j , G j ) ,
where c j denotes lexical polarity, τ j denotes transformer-based polarity, Θ j denotes thematic membership, Z j denotes latent semantic cluster membership, and  G j denotes respondent group. The function F ( · ) is not estimated in this study. Rather, it summarises the analytical logic that institutional perception is shaped by lexical sentiment, contextual sentiment, thematic framing, latent semantic structure, and respondent category.
The corpus-level institutional perception pattern can therefore be interpreted as
I P = c ¯ g , c ¯ T k , c ¯ z , τ j , Θ j , Z j , G j ,
where c ¯ g represents group-level mean polarity, c ¯ T k represents theme-level mean polarity, and  c ¯ z represents cluster-level mean polarity. This formulation is intended to organise the empirical findings rather than to produce a single predictive score. It allows the study to interpret institutional perception as a structured configuration of symbolic legitimacy, socioeconomic and developmental confidence, political-security scepticism, institutional constraint, and reform-oriented expectation.
Future research using larger interview samples or survey-linked corpora may operationalise this formalisation as an estimated model by assigning weights to the lexical, contextual, thematic, semantic-cluster, and respondent-level components. In the present study, however, the formulation is used only as an interpretive synthesis framework for integrating the empirical results reported in the Results and Discussion sections.

3.14. Algorithmic Implementation

The full computational procedure is summarised in Algorithm 1. The algorithm presents the operational workflow used for corpus construction, thematic assignment, sentiment modelling, validation, robustness checking, and interpretive synthesis.
The computational complexity of Algorithm 1 was also analysed to clarify its scalability. Let M denote the number of answer segments, W the total number of respondent-generated words, K the number of predefined themes, K z the number of latent semantic clusters, d the latent semantic dimension, I the number of clustering iterations, B the number of bootstrap resamples, | V | the validation subset size, and T the maximum transformer token length. In this study, M = 239 , W = 42,940 , K = 6 , K z = 6 , d=100, B = 1000 , | V | = 48 , and T was capped at 512 tokens. Corpus cleaning, segmentation, and lexical sentiment analysis scale linearly with the corpus size, approximately O ( W ) . TF-IDF thematic assignment scales as O ( W + M K ) because each segment is compared with the predefined theme descriptors. Transformer-based sentiment inference represents the dominant contextual sentiment cost, approximately O ( M T 2 ) , due to the quadratic self-attention operation with respect to sequence length. Latent semantic clustering using TF-IDF, truncated SVD, and k-means scales approximately with the TF-IDF matrix size and the number of clustering iterations, while statistical testing and bootstrap estimation scale approximately as O ( M log M + B M ) . Manual validation scoring scales linearly with the validation subset, O ( | V | ) . Overall, because the corpus contains 239 answer segments and the transformer input length was capped at 512 tokens, the full pipeline remains computationally tractable on standard research computing hardware. For larger corpora, the main scalability constraint would arise from transformer inference, while latent semantic vectorisation and clustering can be scaled through sparse matrix operations and efficient batching.

3.15. Computational Implementation Details

To ensure reproducibility, all computational procedures were implemented in Python 3.11 using pandas, numpy, scipy, scikit-learn, textblob, transformers, torch, and matplotlib. The reproducibility package additionally reports the preprocessing script, parameter settings, lexical threshold sensitivity analysis, K-sensitivity diagnostics, per-class validation metrics, pairwise cluster tests, and cluster-bootstrap outputs. The interview corpus was analysed in English. Prior to modelling, transcripts were manually inspected and cleaned to remove consent text, repeated instructions, interviewer prompts, greetings, headers, and administrative material. Respondent-generated answers were then segmented into answer-level units. Text normalisation included removal of excessive whitespace, standardisation of line breaks, correction of obvious transcription spacing errors, and preservation of sentence punctuation. No stemming or lemmatisation was applied before sentiment modelling, because these operations can alter sentiment-bearing expressions. Stop-word removal was used only for TF-IDF-based thematic and latent semantic cluster representations, not for lexical or transformer sentiment analysis.
Algorithm 1 Computational Thematic Sentiment Modelling and Validation Procedure
Require: 
Raw interview corpus D r a w , respondent labels g i , theme descriptors Q , lexical model L, transformer model f t r a n s , latent semantic clustering model f c l u s t e r , threshold δ = 0.05 , validation ratio r = 0.20 , bootstrap iterations B = 1000
Ensure: 
Theme labels, sentiment scores, cluster labels, validation metrics, robustness statistics, and interpretive perception synthesis
  1:
Initialise cleaned segment corpus S
  2:
for each transcript d i r a w D r a w  do
  3:
    Remove consent text, instructions, greetings, headers, administrative text, and interviewer questions
  4:
    Segment respondent-generated text into answer-level units
  5:
    Remove courtesy-only and non-substantive segments
  6:
    Store each retained segment with respondent group g i and interview identifier i
  7:
end for
  8:
for each segment s j S  do
  9:
    Assign theme θ j using TF-IDF cosine similarity to theme descriptors
10:
    Compute lexical polarity c j , subjectivity b j , and lexical label z ^ j using threshold δ
11:
    Obtain transformer probabilities ( P p o s , j , P n e u , j , P n e g , j )
12:
    Compute transformer polarity τ j = P p o s , j P n e g , j and transformer label y ^ j
13:
    Generate latent semantic vector e j = f l s a ( s j )
14:
end for
15:
Apply f c l u s t e r to { e j } and assign semantic cluster label z j to each segment
16:
Compute group-level, theme-level, and cluster-level polarity summaries
17:
Compute bootstrap confidence intervals for theme-level and cluster-level mean polarity
18:
Conduct Mann–Whitney tests for expert–general comparison
19:
Conduct Kruskal–Wallis tests across themes and semantic clusters
20:
Conduct chi-square test for categorical sentiment label distributions
21:
Aggregate segment polarity scores within each interview to obtain interview-level means
22:
Recompute the expert–general comparison using interview-level mean polarity scores
23:
for  b = 1 to B do
24:
    Resample interviews with replacement within each respondent group
25:
    Retain all segments belonging to selected interviews
26:
    Compute bootstrap expert–general mean polarity difference Δ b
27:
end for
28:
Compute cluster-bootstrap confidence interval C I 95 % ( Δ ) from { Δ b } b = 1 B
29:
Select stratified validation subset V S , where | V | r | S |
30:
for each segment s j V  do
31:
    Compare manual theme labels with automated theme labels θ j
32:
    Compare manual sentiment labels with lexical labels z ^ j and transformer labels y ^ j
33:
end for
34:
Compute thematic accuracy, macro-precision, macro-recall, macro-F1, confusion matrix, and intercoder agreement
35:
Compute manual–lexical agreement, manual–transformer agreement, Cohen’s kappa, and disagreement patterns
36:
Integrate themes, semantic clusters, sentiment scores, validation evidence, and robustness checks into an interpretive institutional perception synthesis
37:
return Validation metrics, sentiment statistics, semantic clusters, robustness results, and interpretive perception synthesis
Lexical sentiment was computed using the TextBlob polarity and subjectivity estimator. For each answer segment s j , the lexical polarity score c j [ 1 , 1 ] and subjectivity score b j [ 0 , 1 ] were extracted. Segment-level lexical sentiment labels were assigned using a neutrality threshold of δ = 0.05 . Thus, a segment was classified as positive when c j > 0.05 , negative when c j < 0.05 , and neutral when 0.05 c j 0.05 . The stabilising constant used in the compound polarity expression was ϵ = 10 8 , included only to avoid division by zero in formal notation. Sensitivity checks were conducted using alternative neutrality thresholds of δ = 0.03 , δ = 0.05 , and  δ = 0.07 to verify whether the principal thematic conclusions remained stable.
The transformer-based sentiment layer was implemented using the pretrained checkpoint cardiffnlp/twitter-roberta-base-sentiment-latest from the Hugging Face transformers library. The model produces three sentiment probabilities corresponding to negative, neutral, and positive classes. The model was used in zero-shot transfer mode without fine-tuning, because the available corpus contained only 239 answer segments and was not sufficiently large for reliable supervised adaptation. Tokenisation was performed using the corresponding RoBERTa tokenizer, with truncation at 512 tokens and padding applied dynamically within each batch. For each segment, transformer polarity was calculated as τ j = P p o s , j P n e g , j , and model confidence was calculated as γ j = max ( P p o s , j , P n e u , j , P n e g , j ) .
Thematic classification was implemented using TF-IDF vectors and cosine similarity. Segment vectors were generated using TfidfVectorizer with unigram and bigram features, English stop-word removal, minimum document frequency of 2, and maximum document frequency of 0.90. Each of the six institutional themes was represented by a short descriptor containing theoretically relevant terms. A segment was assigned to the theme with the highest cosine similarity score. Latent semantic clustering was implemented using the same TF-IDF feature space projected through truncated singular value decomposition with 100 latent dimensions and normalisation before k-means clustering. K-means was run with K = 6, random state = 42, n init = 50, and the Lloyd algorithm. Representative cluster terms were extracted from the highest-weighted TF-IDF features within each cluster, and the robustness of the cluster number was assessed for K = 4, 5, 6, 7, and 8 using silhouette score, NPMI coherence, top-term diversity, adjusted Rand index stability, cluster-size balance, and interpretability. Table 4 highlights the implementation details for supporting research reproducibility. The selected K = 6 solution is therefore interpreted cautiously as a theoretically meaningful working solution, not as evidence that alternative cluster numbers are invalid.

4. Results

After preprocessing, removal of administrative material, and exclusion of courtesy-only responses, the final analytical corpus contained 239 answer segments from 30 interviews, corresponding to 42,940 respondent-generated words. The expert group contributed 121 segments and 24,333 words, whereas the general respondent group contributed 118 segments and 18,607 words. The expert responses were longer on average, with 201.10 words per segment compared with 157.69 words per segment for the general group, indicating a more elaborative discourse pattern among expert participants (as seen from Table 5).
The thematic classification validation results indicate that the TF-IDF cosine similarity procedure achieved acceptable agreement with manually reconciled labels. As shown in Table 6, the automated classification achieved an accuracy of 81.25%, macro-precision of 82.02%, macro-recall of 82.92%, and a macro-F1 of 81.96 percent. Intercoder agreement before reconciliation was substantial, with Cohen’s kappa of 0.78. The confusion matrix in Table 7 shows that most misclassifications occurred between conceptually adjacent categories, particularly symbolic legitimacy, institutional effectiveness, political-security role, and institutional weakness. The corresponding class specific precision, recall, and F1 scores are reported in Table 8. Most themes achieved F1 scores of at least 0.80, whereas institutional effectiveness recorded the lowest F1 score of 0.60, reflecting its greater semantic overlap with adjacent institutional categories. This indicates that the thematic classification procedure is sufficiently reliable for theme-level analysis, while also confirming that some semantic overlap is expected because institutional perception categories are related rather than fully independent.
Before interpreting the substantive sentiment results, manual validation was also conducted for the lexical and transformer-based sentiment outputs. The validation subset contained 48 answer segments, corresponding to approximately 20% of the final analytical corpus. Each segment was manually labelled as positive, negative, neutral, or mixed, and the reconciled manual labels were compared with the lexical and transformer sentiment labels. As shown in Table 9, the transformer model achieved stronger agreement with manual interpretation than the lexical model. The lexical model achieved manual agreement of 70.83% and Cohen’s kappa of 0.58, whereas the transformer model achieved manual agreement of 85.42% and Cohen’s kappa of 0.79. The disagreement patterns indicate that lexical errors were more frequent in mixed-valence statements, while transformer errors were more likely where criticism was subtle, diplomatic, or context-dependent.
The overall sentiment profile was cautiously positive. Expert respondents recorded a mean polarity score of 0.0803, while general respondents recorded a higher mean polarity score of 0.1016, as shown in Table 10. Median polarity followed the same pattern, with 0.0764 for experts and 0.1014 for general respondents. Mean subjectivity was almost identical across groups, 0.3422 for experts and 0.3456 for general respondents, suggesting that both groups expressed comparable levels of evaluative intensity. The distribution shown in Figure 2 confirms that both groups clustered in the mildly positive range, with only a small number of negative outliers.
The categorical sentiment distribution further supports this pattern as evident from Table 11. In the expert corpus, 85 of 121 segments were positive, representing 70.2%, while 31 segments were neutral, 25.6%, and 5 were negative, 4.1%. In the general corpus, 81 of 118 segments were positive, representing 68.6%, 36 were neutral, 30.5%, and only 1 was negative, 0.8%. Thus, the general group produced fewer explicitly negative segments, although the overall label distribution did not differ significantly between groups.
Sensitivity analysis confirmed that the categorical sentiment results were not dependent on the selected lexical neutrality threshold. Table 12 shows that the expert–general label association remained non-significant under all three examined thresholds, while the continuous Mann–Whitney comparison was unchanged because the threshold affects only categorical labels and not the underlying polarity scores.
As depicted in Table 13, Theme-level results revealed a clear asymmetry in institutional perception. The highest mean polarity was observed for reform and future orientation, 0.1700, although this theme contained only seven segments. Among the more substantively represented categories, socioeconomic cooperation produced the strongest positive sentiment, with 49 segments and a mean polarity of 0.1330. Institutional effectiveness followed closely, with 19 segments and a mean polarity of 0.1325. By contrast, institutional weakness and constraints produced the lowest mean polarity, 0.0546, across 43 segments. Political and security role also remained relatively restrained, with a mean polarity of 0.0693 across 53 segments. This pattern indicates that respondents evaluated developmental and reform-oriented functions more positively than coercive, political, or implementation-related functions.
Figure 3 shows that general respondents were more positive than experts in four of the six thematic categories. The largest group gap appeared in symbolic legitimacy and identity, where general respondents recorded a mean polarity of 0.099 compared with 0.057 among experts. Institutional weakness and constraints showed a similar directional difference, 0.075 for general respondents compared with 0.037 for experts. This suggests that expert discourse was more cautious when discussing identity, legitimacy, and structural limitations.
As seen from Table 14, latent semantic clustering identified six interpretable perception clusters. The most positive cluster was socioeconomic and economic cooperation, with a mean polarity of 0.1465 across 28 segments. Institutional capacity, member states, and constraints followed with a mean polarity of 0.1122 across 53 segments. Political role, rights, and member-state coordination produced a more restrained mean polarity of 0.0745, while formation, ummatic feelings, and solidarity recorded 0.0638. This pattern indicates that identity-based discourse was not uniformly celebratory, but was often embedded within reflective or historically qualified statements.
To test whether the six-cluster solution was an arbitrary modelling choice, sensitivity analysis compared K = 4 through K = 8 using silhouette score, NPMI coherence, top-term diversity, adjusted Rand index stability, cluster-size balance, and interpretability. Table 15 shows that no single value of K dominates on every criterion. The six-cluster solution was retained because it provided a defensible compromise between interpretability, separation, and comparability with the six institutional perception dimensions, while avoiding the more fragmented solutions obtained at K = 7 and K = 8.
The two-dimensional semantic cluster map in Figure 4 is used as an interpretive visualisation rather than as an additional statistical model. It shows a visibly structured semantic space in which socioeconomic and economic cooperation appears as a relatively coherent cluster, while political, rights-related, and identity-oriented discourse occupies more restrained semantic regions. The three-dimensional sentiment surface in Figure 5 is likewise illustrative and supports interpretation by showing that sentiment intensity is not uniformly distributed across the theme-cluster matrix. Peaks are concentrated around developmental and institutional performance regions, whereas lower surfaces correspond to weakness, political restraint, and visibility-related concerns.
As seen from Table 16, inferential testing confirms that the main differences lie across institutional dimensions rather than between respondent groups. The segment-level Mann–Whitney test comparing expert and general polarity scores produced U = 6171.50 and p = 0.0704 , which is not significant at the 0.05 level. At the interview level, the corresponding test produced U = 71.00 and p = 0.0888 . The chi-square test for sentiment label distribution by group was also not significant, χ 2 = 3.0990 , p = 0.2124 , with a weak association, Cramer’s V = 0.1139 . In contrast, polarity varied significantly across themes, Kruskal–Wallis H = 45.8768 , p < 0.001 , and across latent semantic clusters, Kruskal–Wallis H = 35.0124, p < 0.001.
To isolate the specific loci of these significant variations and establish an inferential foundation for the coexistence of symbolic-developmental optimism and operational skepticism, post hoc pairwise comparisons were conducted using Dunn’s test with a Holm-Bonferroni correction. While all 15 pairwise theme combinations were evaluated, Table 17 reports the primary theoretically salient contrasts that define these perceptual cleavages, along with their standardized effect sizes ( r = z / N ). These theme-level pairwise contrasts should be interpreted as exploratory post hoc evidence rather than as definitive interview-level inferential estimates, because additional interview-level pairwise bootstrapping across all theme contrasts was limited by small category sizes, particularly the reform and future orientation category.
The post hoc results reveal that sentiment towards Socioeconomic cooperation is significantly higher than both Institutional weakness and constraints ( z = 4.82 , p Holm < 0.001 ) and the Political and security role ( z = 3.91 , p Holm < 0.001 ), carrying moderate effect sizes. Similarly, reform and future orientation exhibits a significant positive separation from Institutional weakness ( z = 2.94 , p Holm = 0.013 ). These robust pairwise variations provide clear inferential evidence that respondents systematically decouple the organisation’s high normative/developmental utility from its operational limitations and structural constraints.
Because latent semantic clusters also showed significant polarity variation, pairwise cluster contrasts were additionally examined using Mann–Whitney tests with Holm correction. Table 18 reports the theoretically most relevant corrected contrasts, using rank-biserial correlation as the effect-size measure. These results show that the socioeconomic and economic cooperation cluster is significantly more positive than the political, identity, visibility, and formation-related clusters.
The inferential results were also checked at the interview level to address the nesting of answer segments within interviews. Although the segment-level Mann–Whitney test indicated no statistically significant expert–general difference, U = 6171.50 , p = 0.0704 , the same conclusion was retained when polarity scores were aggregated at the interview level. The interview-level Mann–Whitney test produced U = 71.00 , p = 0.0888 , indicating that the group-level difference remained non-significant when the interview, rather than the segment, was treated as the unit of comparison. This reduces concern that the expert–general comparison was driven by pseudo-replication across multiple answer segments from the same respondent. Nevertheless, the segment-level thematic and cluster analyses are interpreted as structured within-corpus evidence rather than as estimates supporting broad population-level generalisation.
An interview-level cluster bootstrap was also conducted to assess whether the main expert–general comparison was sensitive to the nesting of segments within interviews. In this procedure, interviews rather than individual segments were resampled, and all segments belonging to each selected interview were retained within each bootstrap sample. The resulting 95% cluster-bootstrap confidence interval for the expert–general mean polarity difference was [−0.0454, 0.0028]. Because this interval included zero, the result supports the conclusion that the expert–general sentiment difference should not be interpreted as statistically robust. This finding is consistent with both the segment-level Mann–Whitney test, (U = 6171.50), (p = 0.0704), and the interview-level Mann–Whitney test, (U = 71.00), (p = 0.0888). The same cluster-bootstrap logic was extended to the main latent semantic cluster claim by resampling interviews and recalculating mean polarity differences between the highest-polarity socioeconomic/economic cooperation cluster and each remaining cluster. As shown in Table 19, all intervals exclude zero, supporting the robustness of the main cluster-level interpretation. Theme-level bootstrap confidence intervals are reported in Table 13; additional pairwise theme bootstrapping was not emphasised because the reform category contained only seven segments and would yield unstable interview-level resamples.
Figure 6 shows that longer responses did not necessarily produce more positive sentiment. Instead, longer expert responses tended to spread across a broader polarity range, indicating more elaborated and qualified evaluation. Figure 7 presents the interview-level profile using deidentified numeric labels. Bubble size represents total word volume per interview. The plot shows that general interviews occupy several of the highest polarity and subjectivity positions, whereas expert interviews are more densely concentrated in the moderate polarity range. This supports the interpretation that expert respondents were more analytically cautious, while general respondents expressed somewhat stronger affective positivity.
Overall, the results show that perceptions of the organisation are not reducible to a simple positive or negative judgement. The dominant empirical pattern is a dual structure: respondents attach comparatively stronger positive sentiment to reform, socioeconomic cooperation, and institutional effectiveness, while expressing more restrained sentiment toward political-security roles, visibility gaps, and institutional constraints. The statistically significant differences across themes and semantic clusters, combined with non-significant differences between respondent groups, indicate that the object of evaluation matters more than the respondent category. The study therefore identifies a measurable perception architecture in which symbolic and developmental legitimacy coexist with operational scepticism.

5. Discussion

Before interpreting the empirical findings, it is useful to position the present study against the main bodies of literature on which it builds. Table 20 summarises how existing work on international organisational legitimacy, sentiment analysis, transformer-based language modelling, latent semantic clustering, and prior computational studies relates to the contribution of this paper. The comparison shows that the present study does not merely apply sentiment analysis to interview data; rather, it integrates institutional theory with a multi-layered computational framework for measuring perception through sentiment, theme, semantic cluster, and statistical validation.
To move from literature positioning to empirical interpretation, Figure 8 provides a synthesis of the institutional perception architecture identified in this study. The figure integrates the thematic classification, latent semantic clustering, and sentiment polarity results into a single three-layer representation. The left layer shows the six institutional themes, the middle layer shows the six latent semantic clusters, and the right layer condenses the analysis into two higher-order interpretive poles: symbolic-developmental legitimacy and operational scepticism. Node size represents segment volume, node colour represents mean polarity, and edge thickness represents the observed thematic-cluster association. This synthesis demonstrates that institutional perception is not organised around a simple positive-negative divide, but around a structured tension between reform-oriented, socioeconomic, and symbolic legitimacy on one side, and political-security, visibility, and institutional constraint-related scepticism on the other.
Building on Table 20 and Figure 8, the findings extend existing work on institutional legitimacy by showing that perception of a civilisational intergovernmental organisation can be quantified without reducing it to a single approval or disapproval score. Earlier studies have established that international organisational legitimacy depends on performance, representation, procedure, and confidence [1,2,3,4]. The present study adds a computational layer to that literature by demonstrating that institutional perception varies more sharply across themes and semantic clusters than across respondent categories. Expert and general respondents differed only modestly in overall sentiment, with mean polarity scores of 0.0803 and 0.1016, respectively. These group-level differences were not statistically significant at either segment level, Mann–Whitney U = 6171.50 , p = 0.0704 , or interview level, U = 71.00 , p = 0.0888 . By contrast, the variation across institutional themes was highly significant, Kruskal–Wallis H = 45.8768 , p < 0.001 , indicating that the object of evaluation matters more than respondent type.
The results reveal a dual perception structure. On one side, respondents expressed stronger positive sentiment toward reform and future orientation, mean polarity = 0.1700 , socioeconomic cooperation, 0.1330 , and institutional effectiveness, 0.1325 . On the other side, more restrained sentiment appeared in relation to institutional weakness and constraints, 0.0546 , and political and security role, 0.0693 . This suggests that the organisation is not perceived as simply successful or unsuccessful. Rather, it is evaluated through a differentiated structure in which symbolic and developmental legitimacy coexist with operational scepticism. The latent semantic clustering results strengthen this interpretation: socioeconomic and economic cooperation emerged as the most positively evaluated semantic cluster, with a mean polarity of 0.1465, whereas political, identity, visibility, and formation-related clusters generated more cautious sentiment. This interpretation is explicitly corroborated by our post hoc pairwise testing (Table 17 and Table 18), which demonstrates that the positive valuation of developmental and reformist efforts stands in statistically significant opposition to the systemic skepticism directed at institutional constraints and geopolitical execution. This pattern is consistent with the idea that civilisational organisations may retain legitimacy through representation, solidarity, and development-oriented aspirations, even when their political and implementation capacities are questioned.
Theoretically, this pattern can be interpreted through the distinction between different sources of institutional legitimacy. Developmental and socioeconomic functions are more likely to generate positive evaluations because they are associated with practical cooperation, shared benefit, low sovereignty cost, and future-oriented institutional promise. In these areas, respondents can recognise value even when implementation remains incomplete, because the organisation is perceived as providing a platform for cooperation, visibility, and collective aspiration. Political-security functions, by contrast, operate in a more contested domain [25]. They require coordination among sovereign member states, credible enforcement capacity, diplomatic unity, and the ability to respond to conflict, minority rights, and geopolitical crises. These are high-expectation and high-risk areas in which institutional limitations become more visible. The weaker sentiment toward political-security roles and institutional constraints should therefore not be read simply as rejection of the organisation. Rather, it reflects a legitimacy gap between what the organisation symbolically represents and what it can operationally deliver in politically sensitive domains [26]. This distinction helps explain why respondents can simultaneously value the organisation’s developmental, identity-based, and reform-oriented functions while remaining sceptical about its capacity for political coordination and security-related action.
Methodologically, the study contributes to computational social science by integrating multiple layers of analysis into a single reproducible framework. The pipeline combines answer-level segmentation, similarity-based thematic classification, lexical sentiment analysis, transformer-based contextual sentiment modelling, latent semantic clustering, non-parametric testing, bootstrap confidence estimation, and manual validation. This combination is important because qualitative interview data are often too context-dependent for simple automated scoring, while purely interpretive methods may not produce comparable indicators across groups and themes. The framework therefore provides a middle path: it preserves interpretive sensitivity while producing measurable outputs that can be compared statistically. The finding that sentiment differs significantly across themes and semantic clusters, but not across respondent groups, also illustrates the value of modelling perception at multiple analytical levels.
Several limitations should be acknowledged. The study is based on 30 semi-structured interviews and 239 answer-level segments, which is appropriate for exploratory and theory-guided computational analysis of interview narratives but does not support strong population-level generalisation. Although segmentation increases analytical granularity, the segments are nested within interviews and should not be interpreted as fully independent observations; therefore, both segment-level and interview-level comparisons were reported, with the main expert–general result remaining non-significant under interview-level aggregation. The purposive sampling strategy may also introduce respondent selection bias, since participants were selected for their academic, professional, or institutional familiarity with the organisation. In addition, the analysis depends partly on automated lexical and transformer-based sentiment tools, which may misinterpret diplomatic language, indirect criticism, cautious praise, culturally embedded meanings, or mixed institutional evaluations, even though manual validation was used to reduce this risk. The transformer checkpoint used in this study was originally designed for social-media sentiment classification rather than interview-based institutional and political discourse. Its outputs are therefore treated as a contextual robustness layer rather than as a domain-fine-tuned gold standard, and future work should compare or fine-tune general-domain and political-domain sentiment models on larger institution-specific corpora. Finally, the findings should be interpreted as perception patterns within the analysed corpus and organisation, rather than as definitive estimates of wider public or expert opinion or as automatically generalisable to all civilisational, regional, or intergovernmental organisations. Future research should extend the framework using larger interview samples, broader respondent categories, additional country contexts, multilingual corpora, and hierarchical or mixed-effects modelling.

6. Conclusions

This study developed and demonstrated a sentiment-driven computational framework for modelling institutional perceptions in a 57-member civilisational intergovernmental organisation. Using 30 semi-structured interviews, the study produced a final analytical corpus of 239 answer segments and 42,940 respondent-generated words. The evidence shows that institutional perception is broadly positive but unevenly distributed. General respondents were slightly more positive than experts, but the difference was not statistically significant. The stronger and more meaningful differences appeared across institutional themes and latent semantic clusters, where reform, socioeconomic cooperation, and institutional effectiveness attracted more positive sentiment, while institutional weakness, political role, and security-related functions generated more cautious evaluations.
The central conclusion is that institutional perception in civilisational intergovernmental organisations is best understood as a multidimensional construct rather than as a simple sentiment outcome. The organisation examined in this study is perceived through a measurable tension between symbolic developmental legitimacy and operational scepticism. This has implications for both computational methodology and international organisation research. Computationally, it shows that interview narratives can be transformed into valid and interpretable sentiment, theme, and semantic-cluster level indicators. Conceptually, it shows that legitimacy is not only a matter of institutional performance, but also of how different organisational functions are affectively and cognitively evaluated by stakeholders.
A limitation concerns sample size and inferential generalisability. The study is based on 30 semi-structured interviews and 239 answer-level segments. This design is appropriate for exploratory and theory-guided computational analysis of interview narratives, but it does not support strong population-level generalisation. The segment-level corpus increases analytical granularity, but the segments are nested within interviews and therefore should not be interpreted as fully independent observations. To reduce this concern, the study reports both segment-level and interview-level group comparisons, and the main expert–general result remains non-significant under interview-level aggregation. Even so, the findings should be interpreted as evidence of perception patterns within the analysed corpus rather than as definitive estimates of wider public or expert opinion. Future research should extend the design using larger interview samples, additional country contexts, and hierarchical or mixed-effects modelling to estimate respondent-level and segment-level variation more formally. Further extensions can proceed in three directions. First, larger multilingual corpora could be used to test whether similar perception structures appear across member states, regions, and language communities. Second, general-domain or political-domain transformer models can be compared and fine-tuned on institution-specific corpora to improve sensitivity to diplomatic, cultural, and politically cautious language. Third, the proposed framework can be applied comparatively to other international, regional, or identity-based multilateral organisations. Such extensions would allow researchers to examine whether the coexistence of symbolic legitimacy and operational scepticism is a broader feature of complex intergovernmental organisations.

Author Contributions

Conceptualisation, F.S., A.K.M.I.I. and A.A.; methodology, F.S.; software, F.S.; validation, A.A., F.S. and A.K.M.I.I.; formal analysis, F.S.; investigation, F.S.; resources, A.K.M.I.I.; data curation, F.S.; writing—original draft preparation, F.S.; writing—review and editing, F.S., A.K.M.I.I. and A.A.; visualization, F.S.; supervision, A.K.M.I.I.; project administration, A.K.M.I.I.; funding acquisition, A.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

This study was approved on the 20th July 2023 by the Faculty of Business, Law and Politics Research Ethics Committee at the University of Hull (HU67RX United Kingdom). The original research title approved by the Ethics Committee was “Regionalism: OIC- A case study”. (Ethics approval letter attached as a supplementary document).

Informed Consent Statement

Informed consent was obtained from all participants involved in the study.

Data Availability Statement

Due to platform terms, privacy considerations, and the potential sensitivity of conflict-related discourse, the full tweet text and user-level metadata are not publicly released. Aggregated statistics, derived construct counts, and non-identifying analytical outputs may be made available upon reasonable request. Any shared data will exclude direct user identifiers and will follow privacy-preserving procedures.

Acknowledgments

The authors express their sincere gratitude to the Faculty of Business, Law and Politics, University of Hull, for its oversight and ethical approval of this research (HU67RX, United Kingdom). Special appreciation is extended to the Department of Management, University of Dhaka, for facilitating access to respondents and for logistical support during the survey administration phase. The authors acknowledge the invaluable contributions of participating students whose informed and voluntary responses enabled the empirical foundation of this study.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
BERTBidirectional Encoder Representations from Transformers
CIConfidence Interval
IPInstitutional Perception
LDALatent Dirichlet Allocation
NLPNatural Language Processing
OICOrganisation of Islamic Cooperation
RoBERTaRobustly Optimized BERT Pretraining Approach
TF-IDFTerm Frequency–Inverse Document Frequency
VADERValence Aware Dictionary and sEntiment Reasoner

References

  1. Tallberg, J.; Zürn, M. The Legitimacy and Legitimation of International Organizations: Introduction and Framework. Rev. Int. Organ. 2019, 14, 581–606. [Google Scholar] [CrossRef] [Scilit]
  2. Dellmuth, L.M.; Tallberg, J. The Social Legitimacy of International Organisations: Interest Representation, Institutional Performance, and Confidence Extrapolation in the United Nations. Rev. Int. Stud. 2015, 41, 451–475. [Google Scholar] [CrossRef] [Scilit]
  3. Dellmuth, L.M.; Scholte, J.A.; Tallberg, J. Institutional Sources of Legitimacy for International Organisations: Beyond Procedure versus Performance. Rev. Int. Stud. 2019, 45, 627–646. [Google Scholar] [CrossRef] [Scilit]
  4. Steffek, J. Triangulating the Legitimacy of International Organizations: Beliefs, Discourses, and Actions. Int. Stud. Rev. 2023, 25, viad054. [Google Scholar] [CrossRef] [Scilit]
  5. Pang, B.; Lee, L. Opinion Mining and Sentiment Analysis. Found. Trends Inf. Retr. 2008, 2, 1–135. [Google Scholar] [CrossRef] [Scilit]
  6. Liu, B. Sentiment Analysis and Opinion Mining. Synth. Lect. Hum. Lang. Technol. 2012, 5, 1–167. [Google Scholar] [CrossRef] [Scilit]
  7. Blei, D.M.; Ng, A.Y.; Jordan, M.I. Latent Dirichlet Allocation. J. Mach. Learn. Res. 2003, 3, 993–1022. [Google Scholar]
  8. Grimmer, J.; Stewart, B.M. Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts. Political Anal. 2013, 21, 267–297. [Google Scholar] [CrossRef] [Scilit]
  9. Hutto, C.J.; Gilbert, E. VADER: A Parsimonious Rule-Based Model for Sentiment Analysis of Social Media Text. Proc. Int. AAAI Conf. Web Soc. Media 2014, 8, 216–225. [Google Scholar] [CrossRef] [Scilit]
  10. Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, Minneapolis, MN, USA, 3–5 June 2019; pp. 4171–4186. [Google Scholar]
  11. Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
  12. Grootendorst, M. BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure. arXiv 2022, arXiv:2203.05794. [Google Scholar]
  13. Isoaho, K.; Gritsenko, D.; Mäkelä, E. Topic Modeling and Text Analysis for Qualitative Policy Research. Policy Stud. J. 2021, 49, 300–324. [Google Scholar] [CrossRef] [Scilit]
  14. Sufi, F.K. Identifying the Drivers of Negative News with Sentiment, Entity and Regression Analysis. Int. J. Inf. Manag. Data Insights 2022, 2, 100074. [Google Scholar] [CrossRef] [Scilit]
  15. Sufi, F.K. A New Computational Method for Quantification and Analysis of Media Bias in Cybersecurity Reporting. IEEE Trans. Comput. Soc. Syst. 2025, 12, 4561–4570. [Google Scholar] [CrossRef] [Scilit]
  16. Barbieri, F.; Camacho-Collados, J.; Espinosa-Anke, L.; Neves, L. TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, Online, 16–20 November 2020; pp. 1644–1650. [Google Scholar] [CrossRef] [Scilit]
  17. Mann, H.B.; Whitney, D.R. On a Test of Whether One of Two Random Variables is Stochastically Larger than the Other. Ann. Math. Stat. 1947, 18, 50–60. [Google Scholar] [CrossRef] [Scilit]
  18. Kruskal, W.H.; Wallis, W.A. Use of Ranks in One-Criterion Variance Analysis. J. Am. Stat. Assoc. 1952, 47, 583–621. [Google Scholar] [CrossRef]
  19. Efron, B.; Tibshirani, R.J. An Introduction to the Bootstrap; Chapman and Hall: London, UK, 1993. [Google Scholar]
  20. Patton, M.Q. Qualitative Research and Evaluation Methods, 3rd ed.; Sage: Thousand Oaks, CA, USA, 2002. [Google Scholar]
  21. Creswell, J.W.; Poth, C.N. Qualitative Inquiry and Research Design: Choosing Among Five Approaches, 4th ed.; Sage: Thousand Oaks, CA, USA, 2018. [Google Scholar]
  22. Morse, J.M. The Significance of Saturation. Qual. Health Res. 1995, 5, 147–149. [Google Scholar] [CrossRef] [Scilit]
  23. Guest, G.; Bunce, A.; Johnson, L. How Many Interviews Are Enough? An Experiment with Data Saturation and Variability. Field Methods 2006, 18, 59–82. [Google Scholar] [CrossRef] [Scilit]
  24. Cohen, J. A Coefficient of Agreement for Nominal Scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef] [Scilit]
  25. Dellmuth, L.; Tallberg, J. Public Opinion and International Organizations. Rev. Int. Organ. 2026. [Google Scholar] [CrossRef] [Scilit]
  26. Tørstad, V. Can Transparency Strengthen the Legitimacy of International Institutions? Evidence from the UN Security Council. J. Peace Res. 2024, 61, 228–245. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Analytical architecture of the computational thematic sentiment modelling framework for institutional perception analysis, showing how interview narratives are transformed into validated thematic, sentiment, topic, statistical, and interpretive outputs.
Figure 1. Analytical architecture of the computational thematic sentiment modelling framework for institutional perception analysis, showing how interview narratives are transformed into validated thematic, sentiment, topic, statistical, and interpretive outputs.
Computation 14 00164 g001
Figure 2. Distribution of segment-level sentiment polarity by respondent group.
Figure 2. Distribution of segment-level sentiment polarity by respondent group.
Computation 14 00164 g002
Figure 3. Heatmap of mean sentiment polarity across institutional themes and respondent groups.
Figure 3. Heatmap of mean sentiment polarity across institutional themes and respondent groups.
Computation 14 00164 g003
Figure 4. Two-dimensional latent semantic cluster map of interview segments.
Figure 4. Two-dimensional latent semantic cluster map of interview segments.
Computation 14 00164 g004
Figure 5. Three-dimensional sentiment surface across themes and latent semantic clusters.
Figure 5. Three-dimensional sentiment surface across themes and latent semantic clusters.
Computation 14 00164 g005
Figure 6. Segment word count and sentiment polarity with subjectivity-weighted marker size.
Figure 6. Segment word count and sentiment polarity with subjectivity-weighted marker size.
Computation 14 00164 g006
Figure 7. Interview-level perception profile based on mean polarity, mean subjectivity, and total word volume.
Figure 7. Interview-level perception profile based on mean polarity, mean subjectivity, and total word volume.
Computation 14 00164 g007
Figure 8. Institutional perception architecture linking themes, latent semantic clusters, and the two interpretive poles of symbolic-developmental legitimacy and operational scepticism. Node size denotes segment volume, node colour denotes mean polarity, and edge thickness denotes thematic-cluster association.
Figure 8. Institutional perception architecture linking themes, latent semantic clusters, and the two interpretive poles of symbolic-developmental legitimacy and operational scepticism. Node size denotes segment volume, node colour denotes mean polarity, and edge thickness denotes thematic-cluster association.
Computation 14 00164 g008
Table 1. Summary of relevant research streams and their relevance to the present study.
Table 1. Summary of relevant research streams and their relevance to the present study.
Research StreamRepresentative StudiesRelevance to the Present Study
Sentiment analysisPang and Lee [5]; Liu [6]; Hutto and Gilbert [9]Provides the foundation for converting subjective interview narratives into measurable polarity, neutrality, and evaluative intensity.
Transformer-based language modellingDevlin et al. [10]; Liu et al. [11]; Barbieri et al. [16]Supports context-sensitive sentiment modelling where political and institutional evaluations may be indirect, cautious, or mixed.
Topic modelling, latent semantic clustering, and computational text analysisBlei et al. [7]; Grootendorst [12]; Grimmer and Stewart [8]; Isoaho et al. [13]Justifies the extraction of latent themes and semantic clusters from qualitative textual data while retaining interpretability.
Institutional legitimacy and perceptionTallberg and Zürn [1]; Dellmuth and Tallberg [2]; Dellmuth et al. [3]; Steffek [4]Establishes that perceptions of international organisations are shaped by performance, representation, confidence, and legitimacy.
Table 2. Notation used in the computational thematic sentiment modelling framework.
Table 2. Notation used in the computational thematic sentiment modelling framework.
SymbolDefinition
D Complete interview corpus
d i The i-th interview document
NNumber of interview documents, where N = 30
D E Expert interview corpus
D G General respondent interview corpus
S Set of answer-level segments after preprocessing
s j The j-th answer segment
MNumber of answer segments, where M = 239 after final cleaning
g j Respondent group label for segment s j , where g j { E , G }
w j Word count of segment s j
T Set of institutional themes
T k The k-th institutional theme
θ j Theme assignment of segment s j
p j Positive lexical sentiment score of segment s j
n j Negative lexical sentiment score of segment s j
u j Neutral lexical sentiment score of segment s j
c j Compound lexical polarity score of segment s j
b j Subjectivity score of segment s j
τ j Transformer-based polarity score of segment s j
γ j Transformer confidence score for segment s j
e j Latent semantic vector of segment s j
z j Latent semantic cluster assignment of segment s j
V Manual validation subset
m j Manual sentiment label assigned to segment s j
I P j Composite institutional perception score for segment s j
Table 3. Institutional theme descriptors used for TF-IDF cosine similarity classification.
Table 3. Institutional theme descriptors used for TF-IDF cosine similarity classification.
ThemeDescriptor Terms Used for Classification
Symbolic legitimacy and identitylegitimacy, identity, representation, solidarity, collective identity, cultural belonging, civilisational identity, symbolic authority, shared values, community, recognition, moral voice
Institutional effectivenesseffectiveness, performance, achievement, implementation, institutional role, capacity, coordination, governance, decision making, policy delivery, organisational function, practical contribution
Political and security rolepolitical role, diplomacy, conflict, security, mediation, international politics, rights advocacy, geopolitical influence, crisis response, peace, protection, minority rights, political coordination
Socioeconomic cooperationcooperation, development, economic cooperation, trade, education, science, technology, poverty, health, social progress, development partnership, investment, knowledge exchange
Institutional weakness and constraintsweakness, limitation, constraint, fragmentation, lack of unity, implementation gap, bureaucracy, ineffectiveness, internal division, lack of enforcement, political rivalry, institutional failure
Reform and future orientationreform, modernisation, future, transformation, improvement, strategic change, institutional renewal, digitalisation, youth, innovation, long-term vision, future relevance
Table 4. Computational implementation settings.
Table 4. Computational implementation settings.
ComponentImplementation Detail
Programming environmentPython 3.11 with pandas, numpy, scipy, scikit-learn, textblob, transformers, torch, and matplotlib.
Corpus languageEnglish interview transcripts.
PreprocessingRemoval of consent text, interviewer prompts, greetings, administrative text, repeated instructions, and headers; answer-level segmentation; whitespace and line-break normalisation.
Lexical sentiment modelTextBlob polarity and subjectivity estimator.
Lexical label thresholdPositive if c j > 0.05 , negative if c j < 0.05 , neutral otherwise. Sensitivity checked for δ = 0.03 , δ = 0.05 , and  δ = 0.07 .
Stabilising constant ϵ = 10 8 .
Transformer modelcardiffnlp/twitter-roberta-base-sentiment-latest, used without fine-tuning.
Transformer tokenisationRoBERTa tokenizer, maximum sequence length of 512 tokens, dynamic padding, truncation for longer segments.
Transformer polarity τ j = P p o s , j P n e g , j .
Thematic classificationTF-IDF unigram and bigram vectors with cosine similarity to predefined institutional theme descriptors.
TF-IDF settingsEnglish stop-word removal, min_df=2, max_df=0.90, unigram and bigram features.
Latent semantic representationTF-IDF unigram and bigram features projected by truncated SVD with 100 latent dimensions and normalisation.
Semantic clusteringk-means clustering with K = 6, random state = 42, n init = 50, and Lloyd optimisation.
Cluster-number robustnessK = 4, 5, 6, 7, and 8 compared using silhouette score, NPMI coherence, top-term diversity, adjusted Rand index stability, cluster-size balance, and interpretability.
Table 5. Corpus profile after preprocessing.
Table 5. Corpus profile after preprocessing.
GroupInterviewsSegmentsTotal WordsMean Words per SegmentMean Segments per Interview
Expert1512124,333201.108.07
General1511818,607157.697.87
Overall3023942,940179.677.97
Table 6. Validation metrics for thematic classification.
Table 6. Validation metrics for thematic classification.
MetricValueInterpretation
Validation sample size48 segmentsApproximately 20% of the final analytical corpus.
Intercoder agreement0.78Substantial agreement between the two independent manual coders before reconciliation.
Classification accuracy81.25%Proportion of automated theme labels matching reconciled manual labels.
Macro-precision82.02%Average precision across the six themes.
Macro-recall82.92%Average recall across the six themes.
Macro-F181.96 percentUnweighted mean of the six class-level F1 scores.
Table 7. Confusion matrix for TF-IDF cosine similarity thematic classification against reconciled manual labels.
Table 7. Confusion matrix for TF-IDF cosine similarity thematic classification against reconciled manual labels.
Manual Label/Automated LabelSLIEPSSCIWRF
SL1211000
IE030100
PS109010
SC110800
IW011060
RF000001
SL = Symbolic legitimacy and identity; IE = Institutional effectiveness; PS = Political and security role; SC = Socioeconomic cooperation; IW = Institutional weakness and constraints; RF = Reform and future orientation.
Table 8. Per-class thematic validation metrics.
Table 8. Per-class thematic validation metrics.
ThemeManual SupportAutomated CountPrecisionRecallF1
SL14140.85710.85710.8571
IE460.50000.75000.6000
PS11110.81820.81820.8182
SC1090.88890.80000.8421
IW870.85710.75000.8000
RF111.00001.00001.0000
Table 9. Manual validation results for lexical and transformer-based sentiment classification.
Table 9. Manual validation results for lexical and transformer-based sentiment classification.
Validation MeasureLexical ModelTransformer ModelInterpretation
Validation sample size48 segments48 segmentsApproximately 20% of the final analytical corpus.
Manual agreement70.83%85.42%Percentage of model labels matching reconciled manual sentiment labels.
Cohen’s kappa0.580.79Chance-adjusted agreement between model predictions and manual labels.
Most frequent disagreement patternMixed-valence segments misclassified as purely neutral or negative.Subtle diplomatic critiques misclassified as positive or neutral.Disagreements mainly involved mixed, indirect, or context-dependent institutional evaluations.
Main source of errorRigid token-matching failing to parse co-occurring positive and negative terms.Latent context dependence and implicit institutional scepticism.Lexical errors were more common where positive and negative words co-occurred; transformer errors were more common where institutional criticism was implied rather than explicit.
Table 10. Descriptive sentiment statistics by respondent group.
Table 10. Descriptive sentiment statistics by respondent group.
GroupMean PolarityMedian PolaritySD PolarityMean SubjectivityMedian SubjectivityMean Word Count
Expert0.08030.07640.08050.34220.3354201.10
General0.10160.10140.07710.34560.3344157.69
Table 11. Distribution of sentiment labels by respondent group.
Table 11. Distribution of sentiment labels by respondent group.
GroupNegativeNeutralPositive
Expert5 (4.1%)31 (25.6%)85 (70.2%)
General1 (0.8%)36 (30.5%)81 (68.6%)
Table 12. Lexical neutrality threshold sensitivity analysis.
Table 12. Lexical neutrality threshold sensitivity analysis.
δ Overall N/Neu/PExpert N/Neu/PGeneral N/Neu/P χ 2 pCramer’s V
0.037/43/1896/21/941/22/953.56290.16840.1221
0.056/67/1665/31/851/36/813.09900.21240.1139
0.074/100/1354/50/670/50/683.97040.13740.1289
Table 13. Theme-level sentiment summary with bootstrap confidence intervals.
Table 13. Theme-level sentiment summary with bootstrap confidence intervals.
ThemeSegmentsExpertGeneralMean PolarityMean Subjectivity95% CI
Reform and future orientation7520.17000.4072[0.1423, 0.1904]
Socioeconomic cooperation4925240.13300.3380[0.1155, 0.1494]
Institutional effectiveness198110.13250.3469[0.1004, 0.1605]
Symbolic legitimacy and identity6830380.08030.3664[0.0598, 0.1050]
Political and security role5330230.06930.3288[0.0566, 0.0818]
Institutional weakness and constraints4323200.05460.3221[0.0291, 0.0757]
Table 14. Cluster-level sentiment summary with representative terms.
Table 14. Cluster-level sentiment summary with representative terms.
Semantic ClusterSegmentsMean PolarityMean SubjectivityRepresentative Terms
Socioeconomic and economic cooperation280.14650.3525economic, OIC countries, development, socio economic, trade
Institutional capacity, member states, and constraints530.11220.3407member, member states, organisation, ability, hinder
Visibility, public knowledge, and media coverage280.07870.3439media, lack, role, public, coverage
Regional identity and organisational uniqueness320.07600.3763regional, identity, organisation, uniqueness, cooperation
Political role, rights, and member-state coordination720.07450.3132muslim, countries, member, political, rights
Formation, ummatic feelings, and solidarity260.06380.3864ummatic, feelings, formation, muslim, sense
Table 15. Robustness analysis for alternative numbers of latent semantic clusters.
Table 15. Robustness analysis for alternative numbers of latent semantic clusters.
KSilhouetteNPMIDiversityStability ARISize RangeInterpretability
40.06040.23810.87500.487529–128Interpretable but less stable
50.07080.31330.90000.620924–130Interpretable but imbalanced
60.08050.25890.83330.688926–72Balanced and substantively aligned
70.08780.28180.80000.718626–70More fragmented
80.09490.28680.82500.835426–40Stable but analytically finer-grained
Table 16. Summary of inferential statistical tests.
Table 16. Summary of inferential statistical tests.
TestStatisticp-ValueEffect SizeInterpretation
Mann–Whitney, segment-level group comparison6171.500.0704−0.1355Small group difference
Mann–Whitney, interview-level group comparison71.000.0888−0.3689Small to moderate group difference
Chi-square, label distribution by group3.09900.21240.1139Weak association
Kruskal–Wallis, polarity across themes45.8768< 0.001 Significant thematic variation
Kruskal–Wallis, polarity across semantic clusters35.0124< 0.001 Significant cluster variation
Note: Mann–Whitney effect sizes are rank-biserial correlations, r r b = ( 2 U E / ( n E n G ) ) 1 . Negative values indicate lower expert-group sentiment ranks relative to the general respondent group. The chi-square effect size is Cramer’s V.
Table 17. Selected Dunn’s post hoc pairwise comparisons for theme-level sentiment polarity with Holm-Bonferroni correction.
Table 17. Selected Dunn’s post hoc pairwise comparisons for theme-level sentiment polarity with Holm-Bonferroni correction.
Comparison Pair (Theme A vs. Theme B)Mean Diff.Dunn’s z p unadj p Holm Effect Size (r)
Socioeconomic coop. vs. Inst. weakness+0.07844.82< 0.001 < 0.001 0.312 (Moderate)
Socioeconomic coop. vs. Political/Security+0.06373.91< 0.001 < 0.001 0.253 (Small-Mod)
Reform/Future vs. Inst. weakness+0.11542.940.0030.0130.190 (Small)
Inst. effectiveness vs. Inst. weakness+0.07792.710.0070.0270.175 (Small)
Symbolic legitimacy vs. Inst. weakness+0.02571.840.0660.1980.119 (Negligible)
Table 18. Selected post hoc pairwise comparisons for latent semantic cluster sentiment polarity with Holm correction.
Table 18. Selected post hoc pairwise comparisons for latent semantic cluster sentiment polarity with Holm correction.
Comparison PairMean Diff.U p unadj p Holm r rb
Socioeconomic/economic coop. vs. Political/rights+0.07201660.00less than 0.001less than 0.0010.6468
Socioeconomic/economic coop. vs. Regional identity+0.0705738.00less than 0.001less than 0.0010.6473
Socioeconomic/economic coop. vs. Visibility/media gap+0.0678644.00less than 0.001less than 0.0010.6429
Socioeconomic/economic coop. vs. Formation/solidarity+0.0827566.00less than 0.0010.00580.5549
Institutional capacity vs. Political/rights+0.03782550.000.00140.01490.3365
Table 19. Interview-level cluster-bootstrap robustness for the highest-polarity latent semantic cluster.
Table 19. Interview-level cluster-bootstrap robustness for the highest-polarity latent semantic cluster.
ComparisonObserved Difference95% CIIncludes Zero
Socioeconomic/economic coop. minus formation/solidarity0.0827[0.0313, 0.1351]No
Socioeconomic/economic coop. minus institutional capacity0.0343[0.0104, 0.0581]No
Socioeconomic/economic coop. minus regional identity0.0705[0.0424, 0.0952]No
Socioeconomic/economic coop. minus visibility/media gap0.0678[0.0352, 0.0974]No
Socioeconomic/economic coop. minus political/rights0.0720[0.0451, 0.0988]No
Table 20. Positioning of the present study relative to relevant literature.
Table 20. Positioning of the present study relative to relevant literature.
Research DirectionMain Contribution of Prior WorkExtension Provided by the Present Study
Institutional legitimacy of international organisationsPrior studies show that legitimacy is shaped by institutional performance, representation, procedure, public confidence, and discursive legitimation [1,2,3,4].This study converts institutional perception into measurable sentiment, theme, and semantic-cluster level indicators using 239 answer segments and 42,940 respondent-generated words.
Sentiment analysis and opinion miningSentiment studies provide methods for measuring polarity, subjectivity, and evaluative orientation in textual data [5,6,9].This study applies sentiment modelling to interview-based institutional perception rather than to standard opinion corpora, news, or social media text.
Transformer-based language modellingTransformer models improve contextual interpretation of language by capturing semantic dependencies beyond dictionary-based word matching [10,11].This study incorporates transformer-based sentiment modelling as a contextual robustness layer for politically nuanced and institutionally cautious discourse.
Latent semantic clustering and computational text analysisTopic modelling and clustering methods extract latent discourse structures and support the formal analysis of political and policy texts [7,8,12,13].This study links semantic clusters with sentiment scores and institutional themes, showing significant variation across clusters, Kruskal–Wallis H = 35.0124 , p < 0.001 .
Prior computational studies by the authorPrevious work used sentiment analysis, entity extraction, regression, GPT-based classification, and media bias quantification for large-scale textual corpora [14,15].This study extends that computational logic from media and cyber intelligence data to interview-based institutional perception modelling.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sufi, F.; Islam, A.K.M.I.; Akter, A. Computational Modelling of Institutional Perceptions in a Civilisational Intergovernmental Organisation. Computation 2026, 14, 164. https://doi.org/10.3390/computation14070164

AMA Style

Sufi F, Islam AKMI, Akter A. Computational Modelling of Institutional Perceptions in a Civilisational Intergovernmental Organisation. Computation. 2026; 14(7):164. https://doi.org/10.3390/computation14070164

Chicago/Turabian Style

Sufi, Fahim, A. K. M. Iftekharul Islam, and Anowara Akter. 2026. "Computational Modelling of Institutional Perceptions in a Civilisational Intergovernmental Organisation" Computation 14, no. 7: 164. https://doi.org/10.3390/computation14070164

APA Style

Sufi, F., Islam, A. K. M. I., & Akter, A. (2026). Computational Modelling of Institutional Perceptions in a Civilisational Intergovernmental Organisation. Computation, 14(7), 164. https://doi.org/10.3390/computation14070164

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop