Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (48)

Search Parameters:
Keywords = document-level Sentiment Analysis

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
17 pages, 4317 KB  
Article
Bridging Aspect-Level and Document-Level Sentiment Analysis in Online Education Through Constrained Multi-Granularity Generative Modeling
by Shenyi Guo, Youchen Kao and Luchu Cao
Information 2026, 17(9), 868; https://doi.org/10.3390/info17090868 - 8 Sep 2026
Abstract
Automated sentiment analysis of online-education reviews is useful for understanding learner feedback. Classification-based methods usually capture only document-level polarity. They may miss aspect-level signals and may collapse to the majority class under the heavy imbalance typical of course reviews. When the task is [...] Read more.
Automated sentiment analysis of online-education reviews is useful for understanding learner feedback. Classification-based methods usually capture only document-level polarity. They may miss aspect-level signals and may collapse to the majority class under the heavy imbalance typical of course reviews. When the task is reformulated as generation, document-level and aspect-level outputs can be unified. However, out-of-vocabulary aspect labels, parsing failures, and weakly grounded links between granularities may also be introduced. Multi-perspective and Holistic Evaluation T5 (MHE-T5), a model built on the Text-to-Text Transfer Transformer (T5), is proposed as a constrained multi-granularity generative model. It emits aspect-level and document-level sentiment in one schema. The model combines grammar/finite-state machine (FSM)-constrained decoding, a document–aspect consistency coupling with a proved alignment property, and a cross-granularity contrastive objective. The decoding guarantee is limited to schema parse-validity and closed-vocabulary conformity; it does not guarantee semantic correctness of the selected aspect or polarity. Across four datasets, including a large rating-derived Coursera corpus, two human-annotated education aspect-based sentiment analysis (ABSA) datasets, and the standard Multi-Aspect Multi-Sentiment (MAMS) benchmark, generative models improve macro-averaged F1-score (Macro-F1) over Bidirectional Encoder Representations from Transformers (BERT) by 0.36 to 0.61 on the three datasets that carry discriminative baselines. MHE-T5 attains the highest document-level Macro-F1 among the evaluated benchmarks while providing formal schema-level guarantees on the closed-vocabulary settings. A controlled comparison with DeepSeek-V3 on identical examples, used as a large language model (LLM) baseline, shows that the fine-tuned 220M model is a competitive schema-constrained fine-grained aspect extractor under the fixed protocol. Full article
Show Figures

Figure 1

31 pages, 4442 KB  
Article
Explainable Transformer Models for Human Emotion Recognition: A Multi-Method Explainability Study in the Context of Mental Health
by Muhammad Azhar, Naureen Riaz, Waqar Azeem, Deshinta Arrova Dewi, Adeen Amjad and Muhammad Arman
Information 2026, 17(5), 496; https://doi.org/10.3390/info17050496 - 18 May 2026
Viewed by 716
Abstract
The ability to identify emotions based on written text is one of the core areas of Natural Language Processing (NLP) and has many applications in areas such as mental health monitoring, sentiment analysis, and dialogue systems. This study proposes an explainable emotion recognition [...] Read more.
The ability to identify emotions based on written text is one of the core areas of Natural Language Processing (NLP) and has many applications in areas such as mental health monitoring, sentiment analysis, and dialogue systems. This study proposes an explainable emotion recognition (EER) framework built on a fine-tuned RoBERTa-base model trained on the Emotions for NLP dataset with an accuracy of 92.4% and a weighted F1 score of 92.5%. To interpret the decision process of the EER model, we systematically applied four complementary explainable artificial intelligence (XAI) techniques to provide explanations and insights into how the model makes its predictions: SHAP for global token-level feature attribution, LIME for local instance-level explanations, multi-head attention visualization for structural interpretability, and integrated gradients via Captum for axiom-satisfying gradient-based attribution. Each of these four methods provides complementary multi-perspective views of EER model behavior, which can help increase model transparency, identify potential biases, and enable the responsible use of transformer-based models in critical environments (e.g., those requiring formal clinical documentation). Our experiments consistently show that the EER model identifies tokens as having the highest emotional expression level as the strongest predictive feature across methodological perspectives, with strong evidence of cross-methodological agreement regarding the semantic coherence of learned representations. Our findings have direct implications for the responsible implementation of AI-based emotion recognition systems in mental health support systems, where model user-interface transparency, bias mitigation, and clinical trust are necessary to ensure quality patient care. Full article
(This article belongs to the Special Issue Advances in Explainable Artificial Intelligence, 2nd Edition)
Show Figures

Figure 1

22 pages, 5221 KB  
Article
Hybrid Deep Neural Network with Natural Language Processing Techniques to Analyze Customer Satisfaction with Delivery Platform Manager Responses
by Salihah Alotaibi
Appl. Sci. 2026, 16(9), 4359; https://doi.org/10.3390/app16094359 - 29 Apr 2026
Cited by 2 | Viewed by 641
Abstract
Delivery services have drawn much attention and become of topmost significance in urban areas by presenting online food delivery selections for a diversity of dishes from a wide range of restaurants, decreasing both travel and waiting times. Customer data analysis acts as a [...] Read more.
Delivery services have drawn much attention and become of topmost significance in urban areas by presenting online food delivery selections for a diversity of dishes from a wide range of restaurants, decreasing both travel and waiting times. Customer data analysis acts as a cornerstone in corporate strategy, allowing enterprises to gather and interpret user feedback and helping them to make informed decisions that drive future business development. However, major knowledge gaps remain due to the scarcity of literature review studies on these delivery services, hindering a complete understanding of customer satisfaction in this sector. Furthermore, there has been little systematic research on managerial response tactics to online consumer complaints and negative reviews. Researchers have contributed by applying artificial intelligence, including deep learning and machine learning models, to analyze customer sentiment and understand customer brand perceptions. This study presents a Hybrid Deep Neural Network Model for Customer Satisfaction Analysis (HDNNM-CSA), with the aim of developing an efficient model which is capable of accurately classifying customer satisfaction levels in delivery apps based on textual responses provided by customer experience managers. To achieve this, the model initially pre-processes text data using text cleaning, emoji removal, normalization, tokenization, stop word removal, and stemming to clean and standardize the unstructured text data for further analysis. Following this, term frequency–inverse document frequency-based word embedding is utilized to transform the pre-processed text into meaningful feature representations. Lastly, an ensemble architecture involving bidirectional long short-term memory, temporal convolutional, and graph convolutional networks is deployed to classify customer satisfaction levels with managers’ responses. A series of experimental analyses are performed, and the results are examined for numerous features. A comparative analysis demonstrates the enhanced performance of the HDNNM-CSA technique with respect to existing approaches. Full article
Show Figures

Figure 1

23 pages, 3439 KB  
Article
Fear and Neutrality in Disaster Policy Communication: Emotion and Topic Structures from Text Analysis
by Soyoung Kim, Wooje Kim and Richard Clark Feiock
Adm. Sci. 2026, 16(5), 198; https://doi.org/10.3390/admsci16050198 - 23 Apr 2026
Viewed by 1103
Abstract
This study investigates emotional patterns in state government disaster guideline documents using keyword-level emotion analysis and TF–IDF based topic modeling, framing disaster policy communication as an emotional–cognitive dual structure, drawing from Situational Crisis Communication Theory. The findings demonstrate a strong negative relationship between [...] Read more.
This study investigates emotional patterns in state government disaster guideline documents using keyword-level emotion analysis and TF–IDF based topic modeling, framing disaster policy communication as an emotional–cognitive dual structure, drawing from Situational Crisis Communication Theory. The findings demonstrate a strong negative relationship between fear and neutrality, indicating a functional separation between risk awareness and administrative clarity. Nine topics were identified and organized into clusters centered on operational support, administrative structures, and policy frameworks, while content related to hazards and recovery emerged as a distinct semantic category based on cosine similarity analysis. In the integrated analysis of sentiment and topics, neutral language predominates, reflecting the cognitive dimension of government guidelines, with fear and sadness appearing as secondary but systematically patterned emotions. Fear concentrates in topics addressing hazardous conditions and risk-related content. Emotionally neutral language has traditionally been privileged in public administration, but the findings highlight disaster policy communication shaped by governance objectives that privilege specific emotional orientations aligned with coordination, participation, and risk management. State disaster guidelines function not only as technical instructions but also as structured communicative instruments that operate along a dual cognitive–emotional model, shaping public attention and response. Full article
Show Figures

Figure 1

22 pages, 1512 KB  
Article
A Data-Driven Multi-Granularity Attention Framework for Sentiment Recognition in News and User Reviews
by Wenjie Hong, Shaozu Ling, Siyuan Zhang, Yinke Huang, Yiyan Wang, Zhengyang Li, Xiangjun Dong and Yan Zhan
Appl. Sci. 2025, 15(21), 11424; https://doi.org/10.3390/app152111424 - 25 Oct 2025
Cited by 2 | Viewed by 2134
Abstract
Sentiment analysis plays a crucial role in domains such as financial news, user reviews, and public opinion monitoring, yet existing approaches face challenges when dealing with long and domain-specific texts due to semantic dilution, insufficient context modeling, and dispersed emotional signals. To address [...] Read more.
Sentiment analysis plays a crucial role in domains such as financial news, user reviews, and public opinion monitoring, yet existing approaches face challenges when dealing with long and domain-specific texts due to semantic dilution, insufficient context modeling, and dispersed emotional signals. To address these issues, a multi-granularity attention-based sentiment analysis model built on a transformer backbone is proposed. The framework integrates sentence-level and document-level hierarchical modeling, a different-dimensional embedding strategy, and a cross-granularity contrastive fusion mechanism, thereby achieving unified representation and dynamic alignment of local and global emotional features. Static word embeddings combined with dynamic contextual embeddings enhance both semantic stability and context sensitivity, while the cross-granularity fusion module alleviates sparsity and dispersion of emotional cues in long texts, improving robustness and discriminability. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of the proposed model. On the Financial Forum Reviews dataset, it achieves an accuracy of 0.932, precision of 0.928, recall of 0.925, F1-score of 0.926, and AUC of 0.951, surpassing state-of-the-art baselines such as BERT and RoBERTa. On the Financial Product User Reviews dataset, the model obtains an accuracy of 0.902, precision of 0.898, recall of 0.894, and AUC of 0.921, showing significant improvements for short-text sentiment tasks. On the Financial News dataset, it achieves an accuracy of 0.874, precision of 0.869, recall of 0.864, and AUC of 0.895, highlighting its strong adaptability to professional and domain-specific texts. Ablation studies further confirm that the multi-granularity transformer structure, the different-dimensional embedding strategy, and the cross-granularity fusion module each contribute critically to overall performance improvements. Full article
Show Figures

Figure 1

15 pages, 1374 KB  
Article
Stylometric Analysis of Sustainable Central Bank Communications: Revealing Authorial Signatures in Monetary Policy Statements
by Hakan Emekci and İbrahim Özkan
Sustainability 2025, 17(20), 8979; https://doi.org/10.3390/su17208979 - 10 Oct 2025
Cited by 1 | Viewed by 1157
Abstract
Sustainable economic development requires transparent and consistent institutional communication from monetary authorities to maintain long-term financial stability and public trust. This study investigates the latent authorial structure and stylistic heterogeneity of central bank communications by applying stylometric analysis and unsupervised machine learning to [...] Read more.
Sustainable economic development requires transparent and consistent institutional communication from monetary authorities to maintain long-term financial stability and public trust. This study investigates the latent authorial structure and stylistic heterogeneity of central bank communications by applying stylometric analysis and unsupervised machine learning to official announcements of the Central Bank of the Republic of Turkey (CBRT). Using a dataset of 557 press releases from 2006 to 2017, we extract a range of linguistic features at both sentence and document levels—including sentence length, punctuation density, word length, and type–token ratios. These features are reduced using Principal Component Analysis (PCA) and clustered via Hierarchical Clustering on Principal Components (HCPC), revealing three distinct authorial groups within the CBRT’s communications. The robustness of these clusters is validated using multidimensional scaling (MDS) on character-level and word-level n-gram distances. The analysis finds consistent stylistic differences between clusters, with implications for authorship attribution, tone variation, and communication strategy. Notably, sentiment analysis indicates that one authorial cluster tends to exhibit more negative tonal features, suggesting potential bias or divergence in internal communication style. These findings challenge the conventional assumption of institutional homogeneity and highlight the presence of distinct communicative voices within the central bank. Furthermore, the results suggest that stylistic variation—though often subtle—may convey unintended policy signals to markets, especially in contexts where linguistic shifts are closely scrutinized. This research contributes to the emerging intersection of natural language processing, monetary economics, and institutional transparency. It demonstrates the efficacy of stylometric techniques in revealing the hidden structure of policy discourse and suggests that linguistic analytics can offer valuable insights into the internal dynamics, credibility, and effectiveness of monetary authorities. These findings contribute to sustainable financial governance by demonstrating how AI-driven analysis can enhance institutional transparency, promote consistent policy communication, and support long-term economic stability—key pillars of sustainable development. Full article
(This article belongs to the Special Issue Public Policy and Economic Analysis in Sustainability Transitions)
Show Figures

Figure 1

42 pages, 3460 KB  
Review
A Survey of Multi-Label Text Classification Under Few-Shot Scenarios
by Wenlong Hu, Qiang Fan, Hao Yan, Xinyao Xu, Shan Huang and Ke Zhang
Appl. Sci. 2025, 15(16), 8872; https://doi.org/10.3390/app15168872 - 12 Aug 2025
Cited by 11 | Viewed by 8577
Abstract
Multi-label text classification is a fundamental and important task in natural language processing, with widespread applications in specialized domains such as sentiment analysis, legal document classification, and medical coding. However, real-world applications often face challenges such as high annotation costs, data scarcity, and [...] Read more.
Multi-label text classification is a fundamental and important task in natural language processing, with widespread applications in specialized domains such as sentiment analysis, legal document classification, and medical coding. However, real-world applications often face challenges such as high annotation costs, data scarcity, and long-tailed label distributions. These issues are particularly pronounced in professional fields like healthcare and law, significantly limiting the performance of classification models. This paper focuses on the topic of few-shot multi-label text classification and provides a systematic survey of current research progress and mainstream techniques. From multiple perspectives, including modeling under few-shot settings, research status, technical approaches, commonly used datasets, and evaluation metrics, this study comprehensively reviews the existing literature and advances. At the technical level, the methods are broadly categorized into data augmentation and model training. The latter includes paradigms such as transfer learning, prompt learning, metric learning, meta-learning, graph neural networks, and attention mechanisms. In addition, this survey explores the research and progress of specific tasks under few-shot multi-label scenarios, such as multi-label aspect category detection, multi-label intent detection, and hierarchical multi-label text classification. In terms of experimental resources, this review compiles commonly used datasets along with their characteristics and categorizes evaluation metrics that are widely adopted in few-shot multi-label classification settings. Finally, it discusses the key research challenges and outlines future directions, offering insights to guide further investigation in this field. Full article
Show Figures

Figure 1

19 pages, 1450 KB  
Article
Large Language Model-Based Topic-Level Sentiment Analysis for E-Grocery Consumer Reviews
by Julizar Isya Pandu Wangsa, Yudhistira Jinawi Agung, Safira Raissa Rahmi, Hendri Murfi, Nora Hariadi, Siti Nurrohmah, Yudi Satria and Choiru Za’in
Big Data Cogn. Comput. 2025, 9(8), 194; https://doi.org/10.3390/bdcc9080194 - 23 Jul 2025
Cited by 1 | Viewed by 3686
Abstract
Customer sentiment analysis plays a pivotal role in the digital economy by offering comprehensive insights that inform strategic business decisions, optimize digital marketing initiatives, and improve overall customer satisfaction. We propose a large language model-based topic-level sentiment analysis framework. We employ a BERT-based [...] Read more.
Customer sentiment analysis plays a pivotal role in the digital economy by offering comprehensive insights that inform strategic business decisions, optimize digital marketing initiatives, and improve overall customer satisfaction. We propose a large language model-based topic-level sentiment analysis framework. We employ a BERT-based model to generate contextualized vector representations of the documents, and then clustering algorithms are automatically applied to group documents into topics. Once the topics are formed, a GPT model is used to perform sentiment classification on the content related to each topic. The simulations show the effectiveness of this approach, where selecting appropriate clustering techniques yields more semantically coherent topics. Furthermore, topic-level sentiment polarization shows that 31.7% of all negative sentiment concentrates on the shopping experience, despite an overall positive sentiment trend. Full article
Show Figures

Figure 1

22 pages, 1345 KB  
Article
Integrating Financial Knowledge for Explainable Stock Market Sentiment Analysis via Query-Guided Attention
by Chuanyang Hong and Qingyun He
Appl. Sci. 2025, 15(12), 6893; https://doi.org/10.3390/app15126893 - 18 Jun 2025
Cited by 6 | Viewed by 2780
Abstract
Sentiment analysis is widely applied in the financial domain. However, financial documents, particularly those concerning the stock market, often contain complex and often ambiguous information, and their conclusions frequently deviate from actual market fluctuations. Thus, in comparison to sentiment polarity, financial analysts are [...] Read more.
Sentiment analysis is widely applied in the financial domain. However, financial documents, particularly those concerning the stock market, often contain complex and often ambiguous information, and their conclusions frequently deviate from actual market fluctuations. Thus, in comparison to sentiment polarity, financial analysts are primarily concerned with understanding the underlying rationale behind an article’s judgment. Therefore, providing an explainable foundation in a document classification model has become a critical focus in the financial sentiment analysis field. In this study, we propose a novel approach integrating financial domain knowledge within a hierarchical BERT-GRU model via a Query-Guided Dual Attention (QGDA) mechanism. Driven by domain-specific queries derived from securities knowledge, QGDA directs attention to text segments relevant to financial concepts, offering interpretable concept-level explanations for sentiment predictions and revealing the ’why’ behind a judgment. Crucially, this explainability is validated by designing diverse query categories. Utilizing attention weights to identify dominant query categories for each document, a case study demonstrates that predictions guided by these dominant categories exhibit statistically significant higher consistency with actual stock market fluctuations (p-value = 0.0368). This approach not only confirms the utility of the provided explanations but also identifies which conceptual drivers are more indicative of market movements. While prioritizing interpretability, the proposed model also achieves a 2.3% F1 score improvement over baselines, uniquely offering both competitive performance and structured, domain-specific explainability. This provides a valuable tool for analysts seeking deeper and more transparent insights into market-related texts. Full article
(This article belongs to the Special Issue Explainable Artificial Intelligence Technology and Its Applications)
Show Figures

Figure 1

37 pages, 3049 KB  
Article
English-Arabic Hybrid Semantic Text Chunking Based on Fine-Tuning BERT
by Mai Alammar, Khalil El Hindi and Hend Al-Khalifa
Computation 2025, 13(6), 151; https://doi.org/10.3390/computation13060151 - 16 Jun 2025
Cited by 5 | Viewed by 4818
Abstract
Semantic text chunking refers to segmenting text into coherently semantic chunks, i.e., into sets of statements that are semantically related. Semantic chunking is an essential pre-processing step in various NLP tasks e.g., document summarization, sentiment analysis and question answering. In this paper, we [...] Read more.
Semantic text chunking refers to segmenting text into coherently semantic chunks, i.e., into sets of statements that are semantically related. Semantic chunking is an essential pre-processing step in various NLP tasks e.g., document summarization, sentiment analysis and question answering. In this paper, we propose a hybrid chunking; two-steps semantic text chunking method that combines the effectiveness of unsupervised semantic text chunking based on the similarities between sentences embeddings and the pre-trained language models (PLMs) especially BERT by fine-tuning the BERT on semantic textual similarity task (STS) to provide a flexible and effective semantic text chunking. We evaluated the proposed method in English and Arabic. To the best of our knowledge, there is an absence of an Arabic dataset created to assess semantic text chunking at this level. Therefore, we created an AraWiki50k to evaluate our proposed text chunking method inspired by an existing English dataset. Our experiments showed that exploiting the fine-tuned pre-trained BERT on STS enhances results over unsupervised semantic chunking by an average of 7.4 in the PK metric and by an average of 11.19 in the WindowDiff metric on four English evaluation datasets, and 0.12 in the PK and 2.29 in the WindowDiff for the Arabic dataset. Full article
(This article belongs to the Section Computational Social Science)
Show Figures

Figure 1

22 pages, 4051 KB  
Article
Optimizing an LSTM Self-Attention Architecture for Portuguese Sentiment Analysis Using a Genetic Algorithm
by Daniel Parada, Alexandre Branco, Marcos Silva, Fábio Mendonça, Sheikh Mostafa and Fernando Morgado-Dias
Appl. Sci. 2025, 15(11), 6336; https://doi.org/10.3390/app15116336 - 5 Jun 2025
Viewed by 1467
Abstract
A sentiment analysis is a Natural Language Processing (NLP) task that identifies the opinion or emotional tone of documents such as customer reviews, either at the general or detailed level. Improving domain-specific models is important, as it provides smaller and better-suited models that [...] Read more.
A sentiment analysis is a Natural Language Processing (NLP) task that identifies the opinion or emotional tone of documents such as customer reviews, either at the general or detailed level. Improving domain-specific models is important, as it provides smaller and better-suited models that can be implemented by entities that own textual data. This paper presents a deep learning model trained on Portuguese restaurant reviews using recurrent and self-attention mechanisms, which have consistently delivered strong results in prior research studies. Designing an effective model involves numerous hyperparameters and architectural choices. To address this complexity, a discrete genetic algorithm was used to find an optimal configuration, selecting the layer types, placement of self-attention, dropout rate, and model dimensions and shape. A key outcome of this study was that the optimization process produced a model that is competitive with a Bidirectional Encoder Representation from Transformers (BERT) model retrained for Portuguese, which was used as the baseline. The proposed model achieved an area under the curve of 92.1% and F1-score of 75.4%, demonstrating that a small, optimized model can compete and even outperform larger state-of-the-art models. Moreover, this work helps address the scarcity of NLP resources for Portuguese, and highlights the potential of customized architectures over generic solutions. Full article
Show Figures

Figure 1

25 pages, 1964 KB  
Article
Hate Speech Detection and Online Public Opinion Regulation Using Support Vector Machine Algorithm: Application and Impact on Social Media
by Siyuan Li and Zhi Li
Information 2025, 16(5), 344; https://doi.org/10.3390/info16050344 - 24 Apr 2025
Cited by 2 | Viewed by 4095
Abstract
Detecting hate speech in social media is challenging due to its rarity, high-dimensional complexity, and implicit expression via sarcasm or spelling variations, rendering linear models ineffective. In this study, the SVM (Support Vector Machine) algorithm is used to map text features from low-dimensional [...] Read more.
Detecting hate speech in social media is challenging due to its rarity, high-dimensional complexity, and implicit expression via sarcasm or spelling variations, rendering linear models ineffective. In this study, the SVM (Support Vector Machine) algorithm is used to map text features from low-dimensional to high-dimensional space using kernel function techniques to meet complex nonlinear classification challenges. By maximizing the category interval to locate the optimal hyperplane and combining nuclear techniques to implicitly adjust the data distribution, the classification accuracy of hate speech detection is significantly improved. Data collection leverages social media APIs (Application Programming Interface) and customized crawlers with OAuth2.0 authentication and keyword filtering, ensuring relevance. Regular expressions validate data integrity, followed by preprocessing steps such as denoising, stop-word removal, and spelling correction. Word embeddings are generated using Word2Vec’s Skip-gram model, combined with TF-IDF (Term Frequency–Inverse Document Frequency) weighting to capture contextual semantics. A multi-level feature extraction framework integrates sentiment analysis via lexicon-based methods and BERT for advanced sentiment recognition. Experimental evaluations on two datasets demonstrate the SVM model’s effectiveness, achieving accuracies of 90.42% and 92.84%, recall rates of 88.06% and 90.79%, and average inference times of 3.71 ms and 2.96 ms. These results highlight the model’s ability to detect implicit hate speech accurately and efficiently, supporting real-time monitoring. This research contributes to creating a safer online environment by advancing hate speech detection methodologies. Full article
(This article belongs to the Special Issue Information Technology in Society)
Show Figures

Figure 1

17 pages, 1841 KB  
Article
Monitoring of Sustainable Development Trends: Text Mining in Regional Media
by Galina Chernyshova, Evgeniy Taran, Anna Firsova and Alla Vavilina
Sustainability 2025, 17(7), 3122; https://doi.org/10.3390/su17073122 - 1 Apr 2025
Cited by 3 | Viewed by 2588
Abstract
The monitoring of regional development sustainability is closely linked to the development of an indicator system that best meets stakeholders’ requirements, providing a solid foundation for strategic decision-making. In pursuit of progress in achieving the Sustainable Development Goals (SDG), efforts are continuously being [...] Read more.
The monitoring of regional development sustainability is closely linked to the development of an indicator system that best meets stakeholders’ requirements, providing a solid foundation for strategic decision-making. In pursuit of progress in achieving the Sustainable Development Goals (SDG), efforts are continuously being undertaken to refine and enhance the indicator framework. Implementing interdisciplinary approaches for a comprehensive assessment of sustainable development in regions allows for a swift expansion and augmentation of data on regional transformations. An important aspect of the study of sustainability at the regional level is the additional possibility of using unstructured news content through text mining methods. The issue of applying natural language processing techniques for Russian-language sources is significant, as a large number of relevant tools are developed for English. Additionally, the analysis of news content has several features that complicate the classification of sentiments of messages with mostly neutral wording. The proposed methodology for processing specific news content in assessing the sustainability of regional development was implemented. An application for data scraping was developed, data were collected taking into account the selected regions and periods, stop word dictionaries were configured, frequency analysis was implemented, and the sentiment analysis of the obtained slices was carried out. For the formed set of news documents related to sustainable development by keywords according to SDGs 1–17, for the regions of the Volga Federal District, a corpus of documents was obtained representing data for 2021, 2022, and 2023 for 14 regions. The analysis of key topics for different areas and periods was carried out using the cosine similarity measure. The developed approach to news analysis allows for increasing the efficiency of monitoring on various topics. This methodology has been tested for systemic and operational assessment in the dynamics of the sustainable development of regions. Text analysis methods within the framework of decision support at the regional level provide the opportunity to identify emerging trends. Full article
(This article belongs to the Section Development Goals towards Sustainability)
Show Figures

Figure 1

22 pages, 1670 KB  
Article
Word-of-Mouth Evaluation of Ancient Towns in Southern China Using Web Comments
by Yihan Zhang, Weizhuo Guo, Yanling Sheng and Shanshan Li
Tour. Hosp. 2025, 6(1), 25; https://doi.org/10.3390/tourhosp6010025 - 11 Feb 2025
Cited by 1 | Viewed by 3856
Abstract
With the rapid development of digital networks and communication technologies, traditional word-of-mouth (WOM) has transformed into electronic word-of-mouth (eWOM), which plays a pivotal role in improving the management and service quality of ancient town tourism. This study uses Python web scraping techniques to [...] Read more.
With the rapid development of digital networks and communication technologies, traditional word-of-mouth (WOM) has transformed into electronic word-of-mouth (eWOM), which plays a pivotal role in improving the management and service quality of ancient town tourism. This study uses Python web scraping techniques to gather eWOM data from the top ten ancient towns in southern China. Using IPA analysis, the analytic hierarchy process (AHP), Term Frequency–Inverse Document Frequency (TF-IDF), and cluster analysis, we developed a comprehensive eWOM evaluation framework. This framework was employed to perform word frequency analysis, sentiment analysis, topic modeling, and rating analysis, providing deeper insights into tourists’ perceptions. The results reveal several key findings: (1) Transportation infrastructure varies significantly across the towns. Heshun and Huangyao suffer from poor accessibility, while the remaining towns benefit from the developed transportation network of the Yangtze River Delta. (2) The volume of eWOM is strongly influenced by seasonal patterns and was notably impacted by the COVID-19 pandemic. (3) The majority of tourists express positive sentiments toward the ancient towns, with a focus on the available facilities. Their highest levels of satisfaction, however, are associated with the scenic landscapes. (4) A comprehensive eWOM analysis suggests that Wuzhen and Xidi–Hongcun are the most popular tourist destinations, while Zhujiajiao, Huangyao, Zhouzhuang, and Nanxun exhibit lower levels of both attention and visitor satisfaction. Full article
Show Figures

Figure 1

17 pages, 1865 KB  
Article
Improving Sentiment Analysis Performance on Imbalanced Moroccan Dialect Datasets Using Resample and Feature Extraction Techniques
by Zineb Nassr, Faouzia Benabbou, Nawal Sael and Touria Hamim
Information 2025, 16(1), 39; https://doi.org/10.3390/info16010039 - 10 Jan 2025
Cited by 6 | Viewed by 3677
Abstract
Sentiment analysis is a crucial component of text mining and natural language processing (NLP), involving the evaluation and classification of text data based on its emotional tone, typically categorized as positive, negative, or neutral. While significant research has focused on structured languages like [...] Read more.
Sentiment analysis is a crucial component of text mining and natural language processing (NLP), involving the evaluation and classification of text data based on its emotional tone, typically categorized as positive, negative, or neutral. While significant research has focused on structured languages like English, unstructured languages, such as the Moroccan Dialect (MD), face substantial resource limitations and linguistic challenges, making effective sentiment analysis difficult. This study addresses this gap by exploring the integration of data-balancing techniques with machine learning (ML) methods, specifically investigating the impact of resampling techniques and feature extraction methods, including Term Frequency–Inverse Document Frequency (TF-IDF), Bag of Words (BOW), and N-grams. Through rigorous experimentation, we evaluate the effectiveness of these approaches in enhancing sentiment analysis accuracy for the Moroccan dialect. Our findings demonstrate that strategic resampling, combined with the TF-IDF method, significantly improves classification accuracy and robustness. We also explore the interaction between resampling strategies and feature extraction methods, revealing varying levels of effectiveness across different combinations. Notably, the Support Vector Machine (SVM) classifier, when paired with TF-IDF representation, achieves superior performance, with an accuracy of 90.24% and a precision of 90.34%. These results highlight the importance of tailored resampling techniques, appropriate feature extraction methods, and machine learning optimization in advancing sentiment analysis for under-resourced and dialect-heavy languages like the Moroccan dialect, providing a practical framework for future research and development in NLP for unstructured languages. Full article
Show Figures

Graphical abstract

Back to TopTop