Next Article in Journal
Hydrological Threats to the Coasts of the Szczecin Lagoon in the Southern Baltic Sea
Next Article in Special Issue
Machine Learning-Informed Hydrological Response Time Modeling in Tropical Watersheds
Previous Article in Journal
UV-C LED Photolysis for Microbial Inactivation in Textile Industry Wastewater
Previous Article in Special Issue
Hybrid Projections of Nitrate Response to Hydroclimatic Variability in Selected U.S. Surface-Water and Groundwater Systems
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Machine Learning Approach to Hydrological Event Detection from News-Informed Social Media Alerts

1
International Research Centre for Artificial Intelligence Under the Auspices of UNESCO (IRCAI), 1000 Ljubljana, Slovenia
2
IHE Delft Institute for Water Education Under the Auspices of UNESCO, 2611 Delft, The Netherlands
3
AI Department, Jožef Stefan Institute, 1000 Ljubljana, Slovenia
4
UNESCO Chair on Water-Related Disaster Risk Reduction, 1000 Ljubljana, Slovenia
5
Faculty of Civil and Geodetic Engineering, University of Ljubljana, 1000 Ljubljana, Slovenia
6
Faculty of Mathematics and Physics, University of Ljubljana, 1000 Ljubljana, Slovenia
7
Aguas de Alicante, 03008 Alicante, Spain
8
Birla Institute of Technology and Science, Pilani 333031, India
*
Authors to whom correspondence should be addressed.
Water 2026, 18(15), 1820; https://doi.org/10.3390/w18151820
Submission received: 4 April 2026 / Revised: 9 June 2026 / Accepted: 10 July 2026 / Published: 27 July 2026

Abstract

Participatory citizenship plays a critical role in strengthening climate change resilience, particularly in the context of natural disasters such as floods and other hydrological extremes. Citizen-generated data shared through social media platforms offer valuable real-time insights that can complement traditional environmental monitoring systems. This study proposes a machine learning-based framework to analyze multilingual news data and global X (formerly known as Twitter) data that can complement street level sensor data for improved detection and understanding of extreme hydrological events: floods and landslides. The approach identifies and filters tweets related to hazards such as floods and contextualizes them with information extracted from news reports to enhance event characterization. In addition, sentiment and emotion analysis are applied to assess public reactions and perceived event intensity. By integrating physical event signals with societal responses, the method provides a broader perspective on disaster impacts and the effectiveness of emergency responses. The results highlight the potential of combining social media analytics and machine learning to support hydrological monitoring, enhance situational awareness, and contribute to more responsive disaster management strategies in the face of increasing climate-related risks.

Graphical Abstract

1. Introduction

Globally, the frequency and intensity of hydrometeorological extremes have increased, highlighting the need for innovative approaches to monitor, analyze, and manage these events [1]. Recent advances in Artificial Intelligence (AI), particularly in Natural Language Processing (NLP), together with the vast amount of data generated on the internet, provide new opportunities to improve the detection and understanding of such phenomena [2]. Online social media platforms, especially X (formerly known as Twitter), that we hereby refer to as its preceding name, Twitter, have emerged as valuable data sources for environmental monitoring and disaster analysis [3]. These platforms provide near real-time information on ongoing events and often include geolocation metadata, enabling timely and spatially relevant observations of emerging hazards [4]. The widespread use of Twitter has generated a substantial body of research, particularly in the analysis of tweet sentiment and opinion mining [5], offering important methodological foundations for the complementary perspective of the public opinion on the impacts of these events, and, e.g., the efficiency of the reactions of public authorities. The information captured on Twitter is complemented with the crawled open data of news published on these events across multilingual venues Beyond providing rapid information flows, media and social media platforms facilitate the integration of hydrology and socio-hydrology by capturing public perceptions and reactions to extreme events, including their impacts and consequences [6]. The georeferenced nature of the information shared on these platforms—whether explicitly mentioned in the post or inferred by the registered geolocation of the account—makes them a valuable resource for detecting, characterizing, and retrospectively analyzing hydrological events [7,8]. Although Twitter data related to localized urban water issues, such as leakage or infrastructure failures, may be limited, substantial volumes of information can be collected for larger-scale hydrological events [9]. Such information can act as an early warning signal for water resource managers monitoring social media channels, improving situational awareness and supporting faster response to unfolding events [10]. Despite this potential, several challenges remain in leveraging social media data for environmental monitoring [11,12]. These include the management of large and heterogeneous datasets, the extraction of meaningful signals from noisy data streams [13], and the effective use of rapidly evolving AI and machine learning techniques for knowledge discovery [14,15]. The study also highlights important limitations: social media data are inherently noisy, unevenly distributed across regions and languages, and subject to variations in volume and signal strength across events. These constraints underline the need for stronger quantitative validation, benchmarking against conventional monitoring systems, and closer integration with ground-truth hydrometeorological observations. Improving multilingual processing, explainability, and model robustness is essential for operational adoption.
Edge Intelligence in IoT-based Cyber-Physical Systems and the integration of multi-modal data provide a strong technical foundation for hybrid disaster detection systems, which can be significantly enhanced by digitally crowdsourced data from social media. Edge AI enables real-time, low-latency processing directly on distributed sensors, improving responsiveness and reliability during crises, especially when connectivity to centralized cloud systems is limited [16]. At the same time, multi-modal data fusion—combining environmental sensing, satellite imagery, infrastructure data, and geospatial information—has been shown to improve the accuracy and robustness of disaster prediction models [17]. State-of-the-art research highlights the emergence of social-physical sensing systems, where data from IoT sensors is integrated with social media signals to detect anomalies and events such as floods or earthquakes more effectively [18]. Social media and crowdsourced data provide a complementary, human-centric layer of real-time situational awareness, capturing localized impacts and enabling rapid information dissemination during emergencies [19]. Recent surveys further emphasize that combining AI, edge technologies, and social media analytics is key to building next-generation disaster management systems capable of prediction, detection, and response [20]. By fusing these heterogeneous data streams through machine learning, hybrid systems can cross-validate signals, reduce uncertainty, and generate more reliable early warnings, ultimately leading to more adaptive, scalable, and resilient disaster detection and response frameworks.
In recent years, sentiment analysis has increasingly been used to assess public responses to extreme events and crises [21,22]. Building on this line of research, this study explores the use of machine learning methods to analyze global Twitter data in order to improve the nowcasting and management of hydrological events following [23,24]. Sentiment analysis of social media leveraging transformer-based NLP models such as BERT and RoBERTa can infer the intensity and perceived impact of natural disasters by tracking real-time shifts in emotional tone, urgency, and distress signals in user-generated content [25,26,27]. Recent approaches further enhance this by integrating sentiment scores with geolocation, temporal dynamics, and multimodal data, improving the accuracy of disaster severity estimation and situational awareness in crisis informatics systems [28,29,30].
Beyond the integration of heterogeneous data sources, this study introduces a novel methodological perspective in which the dynamics of news coverage—captured through temporal evolution, co-occurring concepts, geographic references, and event-specific metadata—are treated as a structured knowledge layer that can enrich social-media-based detection mechanisms. Rather than using news solely as a validation source, the approach operationalizes news data as a driver for adaptive signal interpretation, enabling the calibration of Twitter-based alerts through context-aware features derived from how events are reported, evolve, and are framed across media outlets. This shifts social media monitoring from a purely reactive signal detection paradigm toward a hybrid, news knowledge-informed alerting framework that leverages both real-time citizen observations and the evolving narrative structure of reported disasters. The research addresses two key questions: (1) In which ways can the knowledge extracted from news sources improve social media–based alerts for hydrological events? and (2) Can sentiment analysis and emotion detection reveal meaningful indicators of event intensity that can support the calibration of such alerts? This paper contributes to the emerging field of socio-hydrological monitoring through social media analytics in several ways. First, it proposes a machine learning framework for the automated identification and filtering of tweets related to hydrological hazards at a global scale. Second, it integrates social media signals with contextual information extracted from news events to enhance the characterization and validation of detected occurrences. Third, the study applies sentiment and emotion analysis to assess public perception and the perceived intensity of hydrological events, providing insights into societal responses and the effectiveness of emergency actions. Finally, the work demonstrates how combining AI techniques with large-scale digital data streams can support improved situational awareness and decision-making in water-related disaster management.
The media and social media components of the proposed system enable the analysis of the relationship between news coverage and Twitter activity across specific topics, locations, and time periods. As illustrated in Figure 1 for the topic “storm,” the distribution of news articles and tweets exhibits similar temporal patterns, with closely aligned peaks that reflect the evolution and public attention surrounding the event.

2. Methodology

Social media and news media platforms contain vast amounts of data that can be mined for useful information, and are often used to gauge public sentiment on various issues. This can include everything from the timing and location of posts to the specific content of the posts themselves. These data can be analyzed to identify trends, patterns, and correlations that might not be apparent from a more traditional analysis. Moreover, social media platforms can also be used to crowdsource data. For example, users could be encouraged to post images or data about specific natural resources in their area. This can provide researchers with a wealth of data that would be difficult to gather through traditional methods. piece for the so-called forensic disaster analysis, which is a fundamental strategy for disaster management (https://www.undrr.org/disaster-forensics, accessed on 8 May 2025).

2.1. Nowcasting Extreme Weather from Enriched Twitter Events

Leveraging the advantages offered by social media data, the insights extracted from media sources are used to strengthen social-media-based signals, enabling more efficient nowcasting of hydrological events, clearer interpretation of causal relations between events and impacts, and improved sentiment analysis capturing public perceptions of, for example, flood management responses. In this study, nowcasting refers to the short-term prediction or early detection of water-related events immediately before or during their occurrence, based on the real-time ingestion and analysis of Twitter data [31]. Alerts are generated when anomalies are detected in the volume or content of tweets mentioning relevant concepts. These anomalies are identified through statistical and machine learning methods that detect deviations from previously defined baseline patterns.
The analysis of Twitter data was conducted through a research agreement established within the framework of the NAIADES project [32], where authors sourced a feed of tweets keyword-filtered for several water-related hazards identified as relevant by water sector stakeholders. The ingested dataset is pre-filtered for water-related topics. The dataset includes tweets related to floods (5791), droughts (676,300), heat waves (29,742), storms (29,757), and water-borne diseases (2066). Data collection relies on sets of related keywords in multiple languages, including English, Spanish, German, and Romanian, enabling broader geographic coverage and improved event detection (see Figure 2). Most tweets are geolocated at the country level, as many users do not share precise geolocation information. However, additional location cues can sometimes be inferred from the textual content of the tweets themselves. To improve event identification, tailored algorithms are used to detect mentions of specific water-related events in both English and local languages. These search queries can also be dynamically expanded using keywords identified in historical news data describing related events.
A key advantage of social media data is the ability to observe real-time public reporting on specific topics, often accompanied by location-related metadata. Twitter, in particular, has become an important resource for research, supported by its developer ecosystem and the extensive methodological work developed by the research community (https://developer.twitter.com/, accessed on 8 May 2025), including approaches for sentiment analysis in event detection, computed with the well established VADER methodology [33]. While Twitter data related to localized urban water issues—such as leakage or infrastructure failures—is often limited, large-scale hydrological events tend to generate substantial social media activity. Such signals can provide early indications of emerging hazards for water resource managers monitoring these channels.
Social media can be particularly useful in natural resource management, where public opinion can greatly influence policy decisions. For example, if there’s a public outcry on social media against a proposed wastewater project due to environmental concerns, it could influence the decision-making process. For example, the flood that occurred in Alicante on 21 August 2019 generated a large volume of positive sentiment in news coverage, largely associated with the rapid and effective response by the local water utility and municipal authorities. This example illustrates how sentiment indicators can help contextualize the public perception of crisis management and emergency actions. However, social media signals are often noisy and may contain incomplete or ambiguous information. To improve the reliability of Twitter-driven alerts for extreme water events, the proposed approach integrates knowledge extracted from media sources. News articles provide contextual narratives describing what occurred during an event, including its impacts, geographic scope, and temporal evolution. This contextual information is further complemented with historical environmental data, such as meteorological records, which help characterize the conditions under which specific events take place. Together, these sources provide additional explanatory power that helps filter noise in social media data and refine the interpretation of detected signals.

2.2. Knowledge Extraction from Multilingual News

An additional advantage of the proposed approach lies in the integration of knowledge extracted from related textual sources, such as news articles and scientific publications [34]. These sources provide contextual and descriptive information that can complement the real-time signals captured through social media (as shown in Figure 3). By analyzing historical and ongoing news coverage, relevant keywords, entities, and thematic patterns associated with specific hydrological events can be identified and incorporated into the social media monitoring framework. Similarly, scientific literature can provide domain-specific terminology and conceptual relationships that help refine event detection models, and contribute to misinformation strategies as in [35]. The integration of these knowledge sources enables the dynamic tuning of keyword sets and classification models, thereby improving the accuracy and reliability of automatic alerts triggered by social media signals [36]. The historical media dataset analyzed in this study covers more than ten years of worldwide news articles across approximately 60 languages. Exploration of this dataset reveals a substantial increase in reported events related to climate change and its impacts on water resources. In total, more than 590,000 news articles were identified, with prominent references to concepts such as drought (116,602 mentions), flood (108,987), temperature (85,411), and global warming (85,922). Two noticeable decreases in media coverage occur at the end of 2019 and again in 2021, reflecting the shift in global media attention toward the COVID-19 pandemic.
To operationalize this approach, text mining and machine learning techniques are applied to historical datasets in order to identify similar past incidents of extreme hydrological events. The analysis relies on event-related concepts extracted from media coverage, enabling the identification of relevant time periods and locations. Tweets posted within those temporal and spatial windows are then retrieved and filtered based on hashtags and keywords associated with the events. Through this process, the analysis captures how local populations discuss and react to hydrological hazards as they unfold. The integration of Twitter data, global news datasets, and environmental parameters allows the system to address several analytical objectives. First, it enables the detection of anomalies across different data streams, a classical anomaly detection task applied to identify unusual patterns in tweet activity, news reporting, or environmental indicators. Second, it supports the verification of reported incidents, combining ground-truth information such as observed heavy precipitation, news reports of damage, or increased social media activity related to flooding. Third, it allows the prediction of potential events, where the probability of a forthcoming event is estimated based on the occurrence of related preceding indicators. Finally, the system supports event investigation and explanation, using interactive dashboards that allow users to explore collected tweets and news articles, examine keyword distributions (e.g., tag clouds), analyze sentiment patterns, and interpret model outputs through explainable AI approaches.
In the investigation stage, the system examines the causes and impacts of events by combining signals from social media with contextual information extracted from news coverage. In this study, causality refers to the influence of an event on the emergence or evolution of related subsequent events [37]. To explore such causal relations, we rely on the analytical perspectives developed within the NAIADES Water Observatory framework [32]: (i) monitoring and exploration of news articles and social media streams; (ii) analysis of combinations of indicators over time to reveal underlying narratives and trends; and (iii) integration of scientific knowledge derived from published research. Together, these complementary perspectives enable a richer interpretation of hydrological events, linking environmental signals with societal responses and providing insights relevant to domains ranging from public health to infrastructure and water resource management.

3. Results

3.1. Data Extracted for the Improvement of Alert Triggers

To evaluate the capabilities of the proposed AI engine and its underlying knowledge infrastructure, four representative disaster case studies from 2021 were selected: (1) the July 2021 floods in Germany and Belgium, Europe; (2) the 2020 floods in Luwu, Indonesia, the 2020 landslide in Munnar, India; (3) the 2020 landslide in Munnar, India; and (4) the 2021 landslide in Atami, Japan. These cases were chosen to reflect diversity in geographic context, hazard type, scale, and data availability, thereby enabling a comprehensive assessment of the system’s ability to integrate and reason over heterogeneous, multilingual, and multi-source data.
Case Study 1: July 2021 Floods in Europe
The July 2021 floods in Europe constituted one of the most severe hydrometeorological disasters in recent European history, affecting multiple countries, including Germany and Belgium. The event was characterized by extreme precipitation, widespread infrastructure damage, and significant human and economic losses. For this case, a large-scale dataset was constructed comprising four main components. First, 20,498 news articles were collected from international and local media sources, providing extensive coverage of the disaster across different languages and perspectives. Second, these articles were organized into 4260 news events through clustering techniques, enabling the identification of distinct sub-events and temporal dynamics within the broader disaster narrative. Third, a substantial volume of 538,336 tweets related to water-related hazards (including floods and landslides) was collected for July 2021, of which 49,179 tweets explicitly referenced floods. This dataset captures real-time public discourse, situational awareness, and societal responses to the event (see Figure 4). This case study provides a high-density, multi-source data environment suitable for testing large-scale retrieval, cross-source synthesis, and multilingual reasoning.
Case Study 2: July 2020 Floods in Indonesia
The 2020 North Luwu flash flood in Indonesia, particularly affecting the North Luwu Regency in South Sulawesi, was driven by intense and prolonged rainfall that caused several rivers—including the Masamba, Rongkong, and Meli—to overflow. The flooding occurred during the night of 13 July 2020, rapidly inundating multiple districts with mud and debris, and prompting authorities to declare a state of emergency. The disaster gained national and international attention due to the scale of infrastructure damage, the number of displaced residents, and the challenging rescue operations in areas cut off by thick mud deposits and damaged roads. The dataset for this case includes 3876 news articles, reflecting both regional and global reporting on the floods and their aftermath. These articles were grouped into 102 news events, capturing key phases such as extreme rainfall conditions, river overflow, emergency response measures, and recovery efforts. In addition, a targeted set of 428 tweets was collected using the query “flood AND Indonesia,” providing a focused representation of social media discourse related to the event (see Figure 5). Compared to the European case, this dataset is more constrained in terms of social media volume but remains sufficiently rich to support analysis of cross-source relationships and event evolution. It also allows the evaluation of the system’s performance in contexts where data availability is more limited or filtered.
Case Study 3: Munnar August 2020 Landslide, India
The 2020 Pettimudi landslide in the Indian state of Kerala, particularly in the Pettimudi area near Munnar in Idukki district, was triggered by intense monsoon rainfall that led to a sudden and catastrophic debris flow. Occurring late at night on 6 August 2020, the landslide struck vulnerable labour settlements in a tea plantation, destroying rows of workers’ housing and causing a high number of casualties. The event drew national attention due to the loss of life, the remoteness of the affected area, and the challenges faced by rescue teams in accessing the site amid ongoing heavy rain and damaged infrastructure. The dataset for this case includes 2945 news articles, reflecting both national and international reporting on the disaster and its aftermath. These articles were grouped into 87 news events, capturing key phases such as the extreme rainfall conditions, the landslide occurrence, search and rescue operations, and subsequent recovery and policy discussions. In addition, a targeted set of 312 tweets was collected using the query “landslide AND Kerala,” providing a focused representation of social media discourse related to the event (see Figure 6).
Case Study 4: Atami July 2021 Landslide, Japan
The Atami landslide, which occurred on 3 July 2021 in Shizuoka Prefecture, Japan, was a rapid-onset debris flow triggered by intense rainfall. The event resulted in significant destruction in a localized urban area and highlighted the interaction between meteorological hazards and terrain instability. For this case, 1916 news articles were collected, documenting the causes, impacts, and response efforts associated with the landslide. These articles were clustered into 20 news events, representing different stages of the disaster, including the initial occurrence, search and rescue operations, and subsequent investigations. A small but targeted set of 19 tweets was also collected using the query “landslide AND (Atami OR Japan),” capturing limited but relevant social media signals (see Figure 7). This case represents a low-volume, high-specificity dataset, enabling the assessment of the AI agent’s ability to generate meaningful insights and responses in contexts where data is sparse but highly focused.
Together, the four case studies provide a diverse experimental framework for evaluating the proposed AI system. They cover different hazard types (floods and landslides), geographic regions (Europe and Asia), and data scales (from large-scale social media streams to highly targeted datasets). Moreover, the combination of news articles, event-level clustering, and social media data enables the analysis of both structured and unstructured information, as well as the temporal and semantic evolution of disaster-related knowledge.

3.2. Evaluation of the Methodology with Baseline Prototype

For the purpose of experimentally validating the assumption that historical news data can improve the performance and contextual awareness of social-media-based disaster detection systems, we developed a modern tweet-triggered flood and landslide alert framework capable of performing semantic feature extraction from real-time Twitter streams. To enrich the analysis beyond isolated tweet interpretation, the system was augmented with contextual information extracted from historically similar disaster events and geographically profiled locations, including recurring hydrological concepts, temporal reporting dynamics, infrastructure vulnerabilities, and regional risk characteristics derived from archived news coverage.
The proposed news-informed Twitter-triggered alert system builds upon the extraction of the most prominent flood- and landslide-related concepts from historical news coverage, together with the temporal dynamics of news reporting during past disaster events. Beyond identifying semantic patterns associated with hydrological hazards, the approach captures how information evolves across news media over time. This includes transitions from early warnings and eyewitness accounts to institutional responses and impact assessments. By learning these event-specific reporting trajectories, the system can contextualize real-time Twitter signals within historically observed disaster evolution patterns, improving the detection of weak precursor signals and reducing false positives. At the same time, the framework explicitly considers the influence of newsworthiness and media bias, acknowledging that disaster coverage is uneven across regions, populations, and event types, with rural or low-visibility events often receiving delayed or limited reporting. Integrating these considerations enables the construction of a more adaptive and context-aware alert system capable of combining social sensing from Twitter with the structured temporal and semantic knowledge embedded in historical news narratives for improved flood and landslide early warning.
We developed a zero-shot semantic feature extraction pipeline using Google Gemini 2.5 Flash [38]) to process the ingested tweet corpus (available at [39]). For each record, the model generated a structured feature vector across six continuous dimensions [0, 1]: immediacy, first-hand witness, location precision, severity index, distress signal, and hydro-relevance. Additionally, an event type field categorized signals as “flood,” “landslide,” or “null.” This approach leveraged the model’s multilingual reasoning to interpret tweets without further training.
To validate these vectors against expert judgment, we constructed a stratified annotation sample of 50 tweets per case study. Records were ranked by a composite score and partitioned into three zones to avoid an imbalance of non-event data: 30 high-scoring tweets (potential true positives) and 20 low-scoring tweets (potential true negatives). These tweets were manually assigned blindly binary expert labels yes/no (with labeled data available at [40]), providing the ground truth necessary to calculate precision, recall, F1-score, AUC and Cohen’s kappa for the framework’s evaluation.
The results presented in Table 1 highlight substantial differences in the performance of the proposed alert system across the evaluated case studies, particularly in terms of F1 score and inter-rater agreement. Higher performances were observed in the 2021 Belgium and Germany flood events, where the abundance of descriptive, geolocated, and continuously updated tweets contributed to a clearer identification of hydrological impacts and event progression on social media. Similarly high performances were also observed for the India and Japan case studies, suggesting that the semantic feature extraction framework was capable of effectively capturing disaster-related contextual information even in more localized and rapidly evolving landslide scenarios. In contrast, substantially lower precision and F1 scores were obtained for the Indonesia case study, reflecting the difficulty of distinguishing relevant disaster signals from noisy or weakly descriptive social media content. These results suggest that the quality, density, and contextual richness of tweets influence the system evaluation, since social media reporting during natural disasters remains highly across regions, event types, and communication practices. Unlike floods, which tend to evolve over longer temporal windows and generate sustained online discussion around inundation, transport disruption, and emergency response, landslides are highly localized and sudden phenomena that may occur with little warning and limited social media visibility. Consequently, eyewitness reporting is scarcer and frequently lacks precise contextual information, making it more difficult for automated systems to detect and semantically characterize the event in real time. Furthermore, differences in language usage, media attention, digital connectivity, and local social media practices across regions also influence the availability and quality of relevant Twitter signals, contributing to variations in system performance across case studies.
The current results of this machine learning framework and its validation should be interpreted in light of several limitations associated with the current experimental setting and the limited availability of historical and recent Twitter/X data, the relatively small manually annotated evaluation samples, and the number of case studies that could be systematically analyzed, the results nevertheless provide valuable evidence of the feasibility and practical relevance of the proposed methodology. Importantly, the selected case studies encompass a broad range of geographic regions, languages, hazard types, and data availability conditions, offering a challenging testbed that reflects the diversity encountered in real-world hydrological disasters. The framework demonstrated its ability to extract meaningful contextual knowledge from multilingual news sources and use that information to support the interpretation of social media signals, particularly in complex and data-sparse environments where traditional monitoring approaches often face limitations. Although additional large-scale benchmarking, comparisons with alternative methods, and longitudinal validation on future events would further strengthen the assessment, the study establishes a solid proof of concept for a novel news-informed social sensing approach. The continuous ingestion of multilingual news data provides the basis for an adaptive and progressively improving alert system, where each newly incorporated disaster contributes additional contextual knowledge for future calibration, retraining, and robustness evaluation across heterogeneous event conditions.As new disaster events are reported across multilingual news sources captured by the news engine crawlers, the framework incrementally enriches its understanding of recurring hydrological concepts, emerging terminology, regional vulnerabilities, and evolving reporting dynamics. This continuous ingestion process enables the adaptive refinement of alert thresholds, semantic associations, and event characterization patterns over time, improving the system’s capability to distinguish relevant signals from noisy or ambiguous social media activity. In practice, the accumulation of historical event knowledge contributes to progressively more robust and context-aware alerts, particularly in geographically diverse and multilingual environments where hazard-related expressions and reporting practices vary significantly. The approach therefore supports a form of continuously learning socio-hydrological intelligence, where each newly observed event contributes to improving the interpretation and calibration of future social-media-based flood and landslide alerts, ultimately strengthening early warning capacity and situational awareness for water-related disasters.

3.3. Estimating the Intensity of the Event from Sentiment Analysis and Emotion Detection

Sentiment analysis and emotion detection provide valuable complementary indicators for understanding the societal dimension of hydrological events. Beyond identifying the occurrence of hazards through social media signals, the analysis of public sentiment expressed in tweets and news coverage can help estimate the perceived intensity and impact of an event (see Figure 8 for the captured dynamics of that sentiment in the data). Sudden increases in negative emotions—such as fear, anxiety, or anger—may signal escalating impacts or insufficient emergency response, while positive sentiment may reflect effective mitigation measures or successful crisis management. Integrating these indicators into event monitoring systems therefore offers an additional layer of information that can support the calibration and interpretation of social-media-based alerts.
Beyond real-time monitoring, the analysis of social media also enables the forensic investigation of past events, allowing researchers to extract valuable information about the development of a phenomenon and its impacts on affected populations and their environment [41]. When combined with hydrometeorological observations, hydrological and hydraulic data, geomorphological analyses, and assessments of social, economic, and political factors, social media information can contribute to a more comprehensive understanding of disasters and support qualitative recommendations aimed at preventing future events. The added value of social media analysis is particularly significant in data-scarce environments, where traditional monitoring infrastructures may be limited [37].
Another important dimension of information available through social media relates to the preexisting social context, particularly the perception of risk among populations living in vulnerable areas. Public awareness and preparedness play a critical role in reducing disaster impacts. For example, insufficient risk perception among residents contributed significantly to the loss of human lives during the 2024 floods in the Valencia region [42]. Traditional approaches for assessing risk perception typically rely on surveys [43]; however, the long-term analysis of social media discourse can provide continuous and large-scale observations of public attitudes and concerns. Such information can complement survey-based methods and support the design of more effective risk communication and preparedness strategies.

4. Discussion

The results obtained in this study demonstrate the potential of integrating social media data with multilingual news information to improve the detection and interpretation of hydrological events, particularly floods and landslides. The four case studies analyzed—European and Indonesian floods, Indian and Japanese landslide—provide complementary insights into how data availability, geographic scale, and event characteristics influence the performance and usefulness of the proposed framework.

4.1. Achievements and Limitations of This Study

A first important observation concerns the complementarity between news data and Twitter signals. In large-scale events such as the July 2021 European floods, the high volume of both tweets and news articles enabled a rich and consistent representation of the event. The temporal alignment between peaks in news reporting and Twitter activity (as illustrated in Figure 1) suggests that social media can effectively capture real-time public awareness, while news sources provide structured and verified contextual information. This combination allows for improved filtering of irrelevant or noisy social media content, thereby enhancing the reliability of detected alerts. In contrast, in smaller-scale or more localized events such as the Atami landslide, the limited volume of tweets highlights the importance of news data as a primary source of contextual grounding. This demonstrates that the integration of heterogeneous data sources is particularly valuable in addressing data sparsity and uneven social media coverage across regions. The analysis also shows that news-informed keyword expansion and event contextualization play a key role in improving the detection of hydrological events from Twitter streams. By incorporating concepts, entities, and thematic patterns extracted from historical and real-time news articles, the system is better able to identify relevant tweets that might not explicitly use standard hazard-related keywords. This is particularly relevant in multilingual contexts, where different linguistic expressions and local terminology can otherwise limit detection performance. The results across the four case studies indicate that such enrichment contributes to a more comprehensive retrieval of event-related social media content, supporting earlier and more robust alert generation.
Another key factor in the study relates to the variability in data density and its implications for system performance. The European floods case, with hundreds of thousands of tweets, allows for strong statistical signals and reliable anomaly detection. In contrast, the Luwu (Indonesia) floods and especially the Atami (Japan) landslide illustrate scenarios where social media signals are weaker or more fragmented. These differences highlight a fundamental limitation of relying solely on social media for disaster monitoring, reinforcing the importance of integrating additional data sources such as news and environmental observations. The proposed framework addresses this limitation by combining multiple streams of information, enabling meaningful analysis even in low-data contexts. Nevertheless, a key limitation of the proposed approach relates to the representativeness and potential biases inherent in Twitter/X data. Social media users are not a statistically representative sample of the affected population, as platform usage varies significantly across regions, age groups, socioeconomic conditions, and levels of digital access. This challenge is not unique to social media analytics but is also well documented in citizen science initiatives, where voluntary participation often leads to sampling bias, uneven geographic coverage, and overrepresentation of more engaged or technologically enabled communities [44,45]. In addition, the content itself may be influenced by media amplification, coordinated messaging, or misinformation, introducing further distortions in the observed signals. To mitigate these effects, the proposed framework integrates Twitter data with multilingual news sources, which provide more structured and validated contextual information, enabling cross-source verification of detected events. Furthermore, anomaly detection is not based solely on absolute tweet volumes but on deviations from localized baselines, reducing sensitivity to systematic activity biases. Nevertheless, these measures do not fully eliminate representativeness limitations, and future work should incorporate more systematic bias quantification, including demographic-aware weighting, bot detection, and validation against ground-truth hydrometeorological observations and sensor data, to improve the robustness and generalizability of social-media-driven early warning systems.
The application of sentiment analysis and emotion detection provides an additional layer of insight into the societal dimension of hydrological events. The results suggest that variations in sentiment and emotional expressions in Twitter data can serve as proxies for the perceived intensity and impact of an event. For instance, increases in negative emotions such as fear, anxiety, or anger tend to coincide with periods of escalating hazard conditions or insufficient response capacity, while more positive sentiment may reflect effective emergency management or recovery phases. This aligns with the example discussed for the Alicante flood, where positive sentiment in news coverage corresponded to successful crisis response. When applied to Twitter data, such indicators can complement physical and environmental measurements by capturing how affected populations experience and react to disasters in real time. However, the interpretation of sentiment and emotion signals must be approached with caution. Social media expressions are influenced by multiple factors, including cultural context, media framing, and the presence of misinformation or automated accounts. Moreover, sentiment does not always directly correlate with objective measures of event severity. For example, highly destructive events in sparsely populated areas may generate limited social media activity, while less severe events in densely connected regions may produce disproportionately strong online reactions. Therefore, sentiment-based indicators should be considered as complementary signals, rather than standalone measures of event intensity.
By combining anomaly detection in Twitter data with contextual validation from news sources, the framework can support faster identification of emerging hydrological hazards. The inclusion of sentiment and emotion analysis further enhances situational awareness by providing insights into public perception, risk awareness, and the effectiveness of response measures. Such information can be valuable for emergency managers, water utilities, and policymakers in designing targeted communication strategies and improving disaster response. Though, the reliance on Twitter data introduces biases related to user demographics, geographic coverage, and platform usage patterns. Multilingual processing, while partially addressed, remains a challenge. In addition, the current analysis is primarily descriptive supported by an initial news-informed Twitter alert, and further work is needed to quantitatively evaluate the performance of the system and to validate the relationship between social media signals and physical event characteristics.

4.2. Recurring to Complementary Signals from Open Environmental Data and Edge AI Sensors

Deriving from the findings, an additional dimension emerging from the analysis is the extraction of structured indicators from news content, particularly through the quantification of reported impact factors such as affected population, infrastructure damage, economic losses, emergency response actions, and geographic extent. Unlike social media signals, which primarily reflect real-time perceptions and reactions, news articles often provide aggregated, verified, and progressively updated information on the consequences of an event. By systematically extracting and counting these impact-related indicators across news reports, it becomes possible to construct proxy measures of event severity and progression over time. These indicators can serve as a valuable complement to Twitter-based alerts by enabling context-aware calibration of detected anomalies, helping to distinguish between minor incidents and large-scale disasters. Furthermore, the integration of such impact metrics into the proposed framework contributes to the development of disaster resilience intelligence, where alert systems are not only triggered by the occurrence of an event but are also enriched with information about its potential consequences and required response capacity. This enhanced situational awareness can support more informed decision-making by emergency managers and water authorities, allowing for prioritization of interventions, allocation of resources, and improved communication with affected communities. In this sense, the combination of real-time social sensing with impact-oriented knowledge extracted from news media represents a significant step toward more holistic and actionable early warning systems for hydrological hazards.
Furthermore, TinyML and edge sensing technologies offer a complementary and highly scalable pathway to enhance hydrological event detection and monitoring [46]. Recent advances in low-power embedded machine learning enable the deployment of intelligent models directly on sensor nodes, such as weather stations, water level gauges, and environmental monitoring devices [47]. These TinyML-enabled sensors can process data locally, detecting anomalies in rainfall intensity, river levels, soil moisture, or drainage system performance in near real time while significantly reducing communication latency and bandwidth requirements. As highlighted by recent cyber-physical system approaches [48], integrating such edge intelligence within IoT-based weather networks allows continuous, fine-grained monitoring of flood-prone environments, particularly in data-scarce or infrastructure-limited regions. Moreover, adaptive sensing strategies can dynamically adjust sampling rates based on detected patterns, improving both energy efficiency and responsiveness. When combined with machine learning models for anomaly detection and forecasting, TinyML-enabled sensing infrastructures can act as an early warning layer, capturing physical signals of hydrological hazards before they are widely reported in social media or news sources.
Building on these sensing capabilities, the contribution of participatory citizenship through Edge AI/TinyML introduces a complementary and socially grounded layer to the monitoring framework. As demonstrated in the FLURMAC project, citizen-driven data—such as geo-referenced images, local observations, and community-reported flooding conditions—can be combined with high-resolution weather parameters including temperature, humidity, and rainfall, collected from distributed IoT-based weather stations. By leveraging federated learning and edge intelligence, these heterogeneous data sources can be processed locally on low-power devices, preserving data privacy while enabling real-time, context-aware decision making. This approach aligns with responsible AI practices by promoting data minimization, local autonomy, transparency, and inclusiveness, ensuring that communities are not only data providers but active participants in the co-creation of early warning systems. The fusion of physical environmental signals with citizen-generated inputs enhances the granularity and reliability of flood detection, particularly in data-scarce or rapidly changing urban environments, while also fostering trust, accountability, and societal engagement in climate resilience strategies.

4.3. Exploring Impact in Water Pollution and Contamination

The technological developments in the context of using social media and media as in [49] can be naturally refocused to monitor and analyze water contamination and extracting best practices utilizing text similarity. With the advantage of Big Data, trends and best practices can be derived from the insightful query over news and research on a certain water topic. The resourceful dataset of media publications (including blogs) and scientific publications are helpful in exploring causality through the identified best practices and resolutions from success stories in, e.g., water contamination by Escherichia Coli (around 40 thousand media articles and 7 thousand research papers). Figure 9 exemplifies the usage of text mining algorithms to cluster subtopics in the research of water pollution through the underlying MEDLINE open data to PubMed and review the existing knowledge on the flood topic over the published research articles and the patented technologies.
The pollution of receiving water bodies (related to city sewage and industrial waste discharge) is a much related focus of interest, as this issue is more frequent (especially in Europe) than drinking water pollution. It is a matter of concern for any utility managing the sewer network and/or the waste water treatment plants, and it is often detected by citizens (social networks) and published by the media. It can be identified by the search terms “sewage discharge”, “sewage disposal”, “industrial discharge”, “sewer overflow” and “untreated wastewater”. Wikipedia utilizes the term “Water pollution” (covered by 71 languages) with a section “Contaminants with an origin in sewage” that refers to this problem in general and can be used to query the worldwide news. On the other hand, the term “Waterborne diseases” (covered by 15 languages) extends this to the health conditions caused by pathogenic micro-organisms that are transmitted in water, covering also the scenario of long-term flooding scenarios caused by extreme weather events. Table 2 below shows how pollution deriving from discharges is a more common topic then the more general water-borne diseases, but the opposite trend is observed in the published science and accepted patents. On the other hand, the topic of floods has a much larger coverage on the news although floods generating diseases are 34% of that coverage, while the published science at that subtopic seems to be very small maybe due to the specificity of the problem. As noticed in Table 2, the characteristics of water pollution event make it harder to apply Twitter data analysis in most cases (pollution events usually have a much gradual evolution along time, and causes are often unknown, among other differences).
When analysing the dataset of news about water contamination, we use the following Wikipedia terms to recover the most common cases, where the languages cover that reflects the multilingual capacity of the system, is noted: Waterborne diseases (15 languages), covering conditions caused by pathogenic micro-organisms that are transmitted in water; Trihalomethane (18 languages), the subject of the first drinking water regulations; and the following microorganisms and organic compounds Protozoa (90 languages), Chloroform (60 languages), Coliform bacteria (19 languages) and Organochlorine (27 languages). Several other water contamination terms can be considered to generate queries, e.g., Bromoform, Bromodichloromethane, and Dibromochloromethane In what regards water bodies pollution we look for: (i) bacteria like Salmonella (90 languages), Shigella (36 languages), Campylobacter (31 languages), or Vibrio cholerae (41 languages); (ii) viruses like hepatitis A (93 languages), rotavirus (48 languages), coronavirus (120 languages), enteroviruses (30 languages); (iii) protozoa like Entamoeba histolytica (42 languages), Giardia lamblia (37 languages), Cryptosporidium parvum (14 languages); and (iv) helminths like Ascaris roundworm (23 languages), Ancylostoma hookworm (22 languages), and Trichuris whipworm (15 languages). This coverage reflects the capacity of information retrieval by text similarity, that allows to create real time alerts or to analyse the best practices from thousands of similar cases automatically identified.
In a similar fashion, one can use literature review for this purpose, taking into consideration the datasets of scientific articles and patents ingested, MAG and MEDLINE. Although the much bigger size of the MAG dataset of scientific articles covering a much wider range of domains, the MEDLINE dataset is much appropriate in this context even if limited to biomedical articles, but labelled by healthcare professionals with the MeSH Headings categories. As shown in Table 3 below, the coverage of some of the most prominent water bodies and contaminants has a significant presence in news, even in specific cases (e.g., entamoeba histolytica), as well as in the MEDLINE dataset using the MeSH Headings healthcare categories to perform the search. Nevertheless, not having the data labelled as in MEDLINE, the keyword search in the abstracts of the articles in MAG recovers a very small sample of articles related to the topics in analysis.
Building on these findings, the analysis can be extended beyond flood events to also consider combined sewer overflows (CSOs). These occur when rainfall, even if not intense enough to trigger flooding, exceeds the capacity of urban drainage systems at certain points [50]. In such cases, untreated mixtures of stormwater and wastewater are released into natural water bodies, such as rivers or coastal areas, contributing to environmental pollution. Because these events can be triggered by relatively moderate precipitation, they tend to occur more frequently than major floods and represent a persistent environmental challenge. In principle, preventing CSOs would require the complete separation of stormwater and wastewater systems; however, most urban infrastructures rely on combined sewer networks for practical and historical reasons. While dedicated sensing technologies can detect overflow events, their deployment at scale is often limited due to cost and infrastructure constraints. In this context, citizen observations become an important complementary source of information. People frequently witness and report such discharges, including through social media, providing valuable qualitative and sometimes quantitative insights into the extent and environmental impact of these events [51].

5. Conclusions

This study proposes a framework to monitor hydrological events, demonstrating the potential of news-based contextual knowledge integrated with large-scale Twitter data, improving the detection, characterization, and interpretation of natural hazards such as floods, droughts, and storms. The results highlight the value of combining citizen-generated data and media knowledge to enhance the monitoring and understanding of the dynamics of these phenomena and can provide cleaner and more coherent early signals of emerging events, continuously calibrated through dynamic, media-derived representations of real-world events. By utilizing reporting timelines, hydrological terminology, and historical event metadata as a calibration layer, we enrich standard Twitter-based alerts. The framework was evaluated against a dataset of 250 expert-annotated tweets across four case studies intentionally considered with diverse geographic, linguistic, and hazard-specific characteristics in order to evaluate the adaptability of the proposed framework across heterogeneous scenarios. Results demonstrate discriminative performance in regions with high social media density, indicating a threshold where supplemental news context currently does not offset the sparsity of local social data. Integrating sentiment analysis into the model revealed that shifts in public tone correlate closely with disaster escalation and the perceived adequacy of emergency responses, providing a more robust tool for situational awareness. They provide complementary insight into perceived severity, societal impacts, and public response, thereby extending purely physical observations with a social dimension.
The continuous integration of multilingual news data enables the system to iteratively refine its semantic representations, contextual associations, and event characterization patterns over time. As additional disaster events are incorporated into the historical knowledge base, the framework progressively improves its capability to identify weak precursor signals, reduce false positives, and adapt to regional and linguistic variations in hazard reporting and social media usage. This continuous enrichment process transforms the proposed approach from a static rule-based validation setting into a dynamically evolving learning framework, in which each newly observed hydrological event contributes additional contextual intelligence for future alert calibration and interpretation. Consequently, although the present validation primarily demonstrates the feasibility and initial effectiveness of the methodology, the framework has been designed to support progressively more robust performance through the longitudinal accumulation of disaster knowledge, larger annotated datasets, iterative benchmarking, and continuous retraining across diverse event conditions and multilingual environments.
In pair with the appropriate integration of complementary environmental data to improve the signals derived from social media, a promising direction for future research lies in the integration of real-time sensing data from distributed IoT and TinyML-enabled devices with social media–derived signals, such as those obtained from Twitter. While social media provides rapid, human-centric observations of unfolding events, sensor networks offer objective, continuous measurements of environmental conditions. The fusion of these complementary data streams can significantly enhance the reliability and timeliness of disaster detection systems. For instance, anomalies detected in water levels or rainfall intensity from edge sensors could be corroborated with spikes in social media activity and sentiment, reducing false positives and improving situational awareness. Conversely, early signals from social media could guide targeted analysis of sensor data in specific locations. This bidirectional integration enables the development of hybrid, multi-modal alert systems that combine physical measurements with societal responses, ultimately leading to more accurate, explainable, and actionable disaster alarms. Such approaches represent a key step toward next-generation disaster resilience intelligence systems that leverage both citizen-generated data and autonomous sensing infrastructures.
Future research will also explore the integration of Agentic Retrieval-Augmented Generation architectures to enhance the analytical capabilities of the proposed framework [52], combining large language models with autonomous AI agents capable of retrieving, integrating, and reasoning over heterogeneous information sources in real time. In the context of hydrological event monitoring, such systems could dynamically retrieve relevant knowledge from news archives, scientific literature, meteorological data, and social media streams to generate contextualized explanations of detected events. This approach integrating agentic AI approaches with socio-hydrological data streams, allows the system not only to detect anomalies but also to interpret them by connecting signals across multiple knowledge sources. By combining real-time citizen-generated data, contextual knowledge from media and scientific sources, and advanced machine learning techniques, future platforms could significantly improve early warning, situational awareness, and decision-making in the management of hydrological hazards under increasing climate variability.

Author Contributions

J.P.C.: Conceptualization, Methodology, Investigation, Data Curation, Writing—Original Draft, Project administration, Funding acquisition; G.C.P.: Conceptualization, Methodology, Investigation, Data Curation, Writing—Original Draft. O.T.: Software, Investigation, Writing—review & editing, Data Curation. M.M.: Conceptualization, Methodology, Writing—review & editing; I.N.: Software, Review & editing, Investigation; R.O.; Methodology, Data Curation, Software, Review & Editing; I.C.d.B.: Writing—review & editing, Investigation, Data Curation, Funding acquisition; N.G.: review & editing, Methodology, Investigation. All authors have read and agreed to the published version of the manuscript.

Funding

The authors thank the support of the European Commission on the H2020 NAIADES project (GA nr. 820985), HE RAIDO (GA nr. 101135800) and HE ELIAS (GA nr. 101120237). The work presented was also supported by the core program P2-0180, funded by the Slovenian Research Agency.

Data Availability Statement

Data used within this paper were either extracted from the NAIADES repositories (open data on water-related statistic indicators, MAG, MEDLINE and a sample of collected Twitter data) at IRCAI, available on request, the Event Registry repositories (data crawled from online news venues), or obtained through public dataset (e.g., soil moisture anomaly index). The labeled data for the engine evaluation is available at [40].

Acknowledgments

The authors acknowledge the collaboration of Eventregistry.org that provided the body, titles and metadata of the ingested news, and to Twitter Research that provided the filtered 1% feed of disaster-related tweets.

Conflicts of Interest

Author Ignacio Casals del Busto was employed by the Aguas de Alicante. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
GISGeographic Information Systems
IPCCIntergovernmental Panel on Climate Change
CSOsCombined Sewer Overflows
IHEInstitute for Water Education
IRCAIInternational Research Centre for Artificial Intelligence
MEDLINEMedical Literature Analysis and Retrieval System Online
MeSHMedical Subject Headings
NAIADESH2020 project on water management and AI

References

  1. IPCC. Climate Change 2014: Synthesis Report. Contribution of Working Groups I, II and III to the Fifth Assessment Report of the Intergovernmental Panel on Climate Change; Pachauri, R.K., Meyer, L.A., Eds.; IPCC: Geneva, Switzerland, 2014; 151p. [Google Scholar]
  2. Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
  3. Stefanidis, A.; Crooks, A.; Radzikowski, J. Harvesting ambient geospatial information from social media feeds. GeoJournal 2013, 78, 319–338. [Google Scholar]
  4. Middleton, S.E.; Middleton, L.; Modafferi, S. Real-time crisis mapping of natural disasters using social media. IEEE Intell. Syst. 2014, 29, 9–17. [Google Scholar] [CrossRef] [Scilit]
  5. Pak, A.; Paroubek, P. Twitter as a corpus for sentiment analysis and opinion mining. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC), Valletta, Malta, 17–23 May 2010. [Google Scholar]
  6. Di Baldassarre, G.; Viglione, A.; Carr, G.; Kuil, L.; Salinas, J.L.; Blöschl, G. Socio-hydrology: Conceptualising human–flood interactions. Hydrol. Earth Syst. Sci. 2013, 17, 3295–3303. [Google Scholar] [CrossRef] [Scilit]
  7. Crooks, A.; Croitoru, A.; Stefanidis, A.; Radzikowski, J. #Earthquake: Twitter as a distributed sensor system. Trans. GIS 2013, 17, 124–147. [Google Scholar]
  8. Bono, C.; Pernici, B.; Fernandez-Marquez, J.L.; Ravi Shankar, A.; Mülâyim, M.O.; Nemni, E. TriggerCit: Early flood alerting using Twitter and geolocation—A comparison with alternative sources. arXiv 2022, arXiv:2202.12014. [Google Scholar]
  9. Thompson, S.; MacVean, L.; Sivapalan, M. A stochastic water balance framework for lowland watersheds. Water Resour. Res. 2017, 53, 9564–9579. [Google Scholar] [CrossRef] [Scilit]
  10. Jongman, B.; Wagemaker, J.; Romero, B.; de Perez, E.C. Early flood detection for rapid humanitarian response: Harnessing near real-time satellite and Twitter signals. ISPRS Int. J. Geo-Inf. 2015, 4, 2246–2266. [Google Scholar] [CrossRef] [Scilit]
  11. Arvandi, A.; Hashemi, S.M.; Karimi, H. Twitter analysis in emergency management: Recent research trends and challenges. Soc. Netw. Anal. Min. 2024, 14, 154. [Google Scholar] [CrossRef] [Scilit]
  12. Soomro, S.-e.-H.; Boota, M.W.; Zwain, H.M.; Soomro, G.-Z. How effective is Twitter (X) social media data for urban flood management? J. Hydrol. Reg. Stud. 2024, 634, 131129. [Google Scholar] [CrossRef] [Scilit]
  13. Kamoji, S.; Kalla, M. Effective Flood prediction model based on Twitter Text and Image analysis using BMLP and SDAE-HHNN. J. Eng. Appl. Artif. Intell. 2023, 123, 106365. [Google Scholar] [CrossRef] [Scilit]
  14. Russell, S.J.; Norvig, P. Artificial Intelligence: A Modern Approach, 3rd ed.; Pearson Education Limited: London, UK, 2016. [Google Scholar]
  15. Hinton, G.; Deng, L.; Yu, D.; Dahl, G.E.; Mohamed, A.R.; Jaitly, N.; Senior, A.; Vanhoucke, V.; Nguyen, P.; Sainath, T.; et al. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Process. Mag. 2012, 29, 82–97. [Google Scholar] [CrossRef] [Scilit]
  16. Tenkale, M. Edge AI for Disaster Management in Smart Cities. Int. J. Sci. Res. Eng. Trends 2025, 11, 1395–1399. [Google Scholar]
  17. Krichen, M.; Abdalzaher, M.S.; Elwekeil, M.; Fouda, M.M. Managing natural disasters: An analysis of technological advancements, opportunities, and challenges. Internet Things Cyber-Phys. Syst. 2024, 4, 99–109. [Google Scholar] [CrossRef] [Scilit]
  18. Rashid, M.T.; Wei, N.; Wang, D. A survey on social-physical sensing: An emerging sensing paradigm that explores the collective intelligence of humans and machines. Collect. Intell. 2023, 2, 1–37. [Google Scholar] [CrossRef] [Scilit]
  19. UNDP; OCHA. Innovation in Disaster Management: Leveraging Technology to Save More Lives; UNDP: New Delhi, India, 2023. [Google Scholar]
  20. Aboualola, M.; Abualsaud, K.; Khattab, T.; Zorba, N.; Hassanein, H.S. Edge technologies for disaster management: A survey of social media and artificial intelligence integration. IEEE Access 2023, 11, 73782–73802. [Google Scholar] [CrossRef] [Scilit]
  21. Brynielsson, J.; Johansson, F.; Jonsson, C.; Westling, A. Emotion classification of social media posts for estimating people’s reactions to communicated alert messages during crises. Secur. Inform. 2013, 2, 7. [Google Scholar]
  22. Yang, P.; Dinh, L.; Stratton, A.; Diesner, J. Detection and categorization of needs during crises based on Twitter data. In Proceedings of the International AAAI Conference on Web and Social Media; AAAI Press: Buffalo, NY, USA, 2024. [Google Scholar]
  23. Liu, B.; Kelly, M.; Gong, L. Using tweets to mine cryptocurrency sentiment and its potential influence on price. In Proceedings of the 2015 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, Paris, France, 25–28 August 2015; pp. 1326–1331. [Google Scholar]
  24. Veigel, N.; Kreibich, H.; de Bruijn, J.A.; Aerts, J.C.J.H.; Cominola, A. Content analysis of multi-annual time series of flood-related Twitter (X) data. Nat. Hazards Earth Syst. Sci. 2025, 25, 879–891. [Google Scholar] [CrossRef] [Scilit]
  25. Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT; Association for Computational Linguistics: Minneapolis, MN, USA, 2019. [Google Scholar]
  26. Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
  27. Imran, M.; Castillo, C.; Diaz, F.; Vieweg, S. Processing Social Media Messages in Mass Emergency: A Survey. ACM Comput. Surv. 2015, 47, 1–38. [Google Scholar] [CrossRef] [Scilit]
  28. Alam, F.; Ofli, F.; Imran, M. CrisisMMD: Multimodal Twitter Datasets from Natural Disasters. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM); AAAI Press: Stanford, CA, USA, 2021. [Google Scholar]
  29. Kryvasheyeu, Y.; Chen, H.; Obradovich, N.; Moro, E.; Van Hentenryck, P.; Fowler, J.; Cebrian, M. Rapid assessment of disaster damage using social media activity. Sci. Adv. 2016, 2, e1500779. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Nguyen, D.; Al Mannai, K.A.; Joty, S.; Sajjad, H.; Imran, M.; Mitra, P. Robust Classification of Crisis-Related Data on Social Networks Using Convolutional Neural Networks. In Proceedings of ICWSM; AAAI Press: Montreal, QC, Canada, 2017. [Google Scholar]
  31. Novalija, I.; Papler, M.; Mladenić, D. Towards social media mining: TwitterObservatory. In Proceedings of the SIKDD2014, Ljubljana, Slovenia, 10 June 2014. [Google Scholar]
  32. Mikoš, M.; Bezak, N.; Costa, J.P.; Massri, M.B.; Novalija, I.; Jermol, M.; Grobelnik, M. Natural-Hazard-Related Web Observatory as a Sustainable Development Tool. In Progress in Landslide Research and Technology; Springer: Cham, Switzerland, 2022; Volume 1. [Google Scholar]
  33. Hutto, C.J.; Gilbert, E.E. VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. In Proceedings of the Eighth International Conference on Weblogs and Social Media (ICWSM-14), Ann Arbor, MI, USA, 1–4 June 2014. [Google Scholar]
  34. Leban, G.; Fortuna, B.; Brank, J.; Grobelnik, M. Event registry: Learning about world events from news. In Proceedings of the 23rd International Conference on World Wide Web, Seoul, Republic of Korea, 7–11 April 2014; pp. 107–110. [Google Scholar]
  35. Apostol, E.-S.; Truică, C.-O.; Paschke, A. ContCommRTD: A distributed content-based misinformation-aware community detection system for real-time disaster reporting. IEEE Trans. Knowl. Data Eng. 2024, 36, 5811–5822. [Google Scholar] [CrossRef] [Scilit]
  36. Hanny, D.; Schmidt, S.; Resch, B. Active learning for identifying disaster-related tweets: A comparison with keyword filtering and generic fine-tuning. In Intelligent Systems Conference; Springer Nature: Cham, Switzerland, 2024; pp. 126–142. [Google Scholar]
  37. Borga, M.; Comiti, F.; Ruin, I.; Marra, F. Forensic analysis of flash flood response. WIREs Water 2019, 6, e1338. [Google Scholar] [CrossRef] [Scilit]
  38. Google. Gemini 2.5 Flash. Google AI. Available online: https://gemini.google.com/ (accessed on 6 May 2026).
  39. Orel, R.; Pita Costa, J. News-Informed Social Media-Triggered Natural Disaster Alert System. IRCAI’s SDG Research GitHub Repository. 2026. Available online: https://github.com/IRCAI-SDGobservatory/disaster_alert (accessed on 8 May 2025).
  40. Orel, R.; Pita Costa, J. Labeled Twitter News Data on Floods and Lanslides in Europe, Indonesia, India and Japan. ZENODO. 2026. Available online: https://zenodo.org/records/20107120 (accessed on 8 May 2025).
  41. Ferreira, A.M.; Marchezini, V.; Mendes, T.S.G.; Trejo-Rangel, M.A.; Iwama, A.Y. A Systematic Review of Forensic Approaches to Disasters: Gaps and Challenges. Int. J. Disaster Risk Sci. 2023, 14, 722–735. [Google Scholar] [CrossRef] [Scilit]
  42. Martin-Moreno, J.M.; Garcia-Lopez, E.; Guerrero-Fernandez, M.; Alfonso-Sanchez, J.L.; Barach, P. Devastating “DANA” floods in Valencia: Insights on resilience, challenges, and strategies addressing future disasters. Public Health Rev. 2025, 46, 1608297. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Meneu, S.N.; Pitarch-Garrido, M.D.; Camarasa-Belmonte, A.M. Analysis of flood hazard perception in the municipality of Tavernes Blanques (Valencia, Spain): Its role in risk assessment. TERRA Rev. Desarro. Local 2021, 8, 68–97. [Google Scholar] [CrossRef] [Scilit]
  44. Kosmala, M.; Wiggins, A.; Swanson, A.; Simmons, B. Assessing data quality in citizen science. Front. Ecol. Environ. 2016, 14, 551–560. [Google Scholar] [CrossRef] [Scilit]
  45. Theobald, E.J.; Ettinger, A.K.; Burgess, H.K.; DeBey, L.B.; Schmidt, N.R.; Froehlich, H.E.; Wagner, C.; HilleRisLambers, J.; Tewksbury, J.; Harsch, M.A.; et al. Global change and local solutions: Tapping the unrealized potential of citizen science for biodiversity research. Biol. Conserv. 2015, 181, 236–244. [Google Scholar] [CrossRef] [Scilit]
  46. Samanta, R.; Saha, B.; Ghosh, S.K. A low-power low-cost system for disaster locations detection using esp32 cam and tinyml. In Proceedings of the 17th International Conference on COMmunication Systems and NETworks (COMSNETS), Bengaluru, India, 6–10 January 2025; IEEE: Piscataway, NJ, USA, 2025. [Google Scholar]
  47. Yarham, S.; Behjati, M.; Nordin, R. A Comprehensive Survey of TinyML Applications in Environmental Monitoring and Their Technical Challenges. In Proceedings of the International Conference on Computer, Internet of Things and Smart City (CIoTSC), Suzhou, China, 7–9 November 2025; IEEE: Piscataway, NJ, USA, 2025. [Google Scholar]
  48. Trigkas, A.; Piromalis, D.; Papageorgas, P. Edge intelligence in urban landscapes: Reviewing TinyML applications for connected and sustainable smart cities. Electronics 2025, 14, 2890. [Google Scholar] [CrossRef] [Scilit]
  49. Pita Costa, J.; Rei, L.; Bezak, N.; Mikoš, M.; Massri, M.B.; Novalija, I.; Leban, G. Towards improved knowledge about water-related extremes based on news media information captured using artificial intelligence. Int. J. Disaster Risk Reduct. 2024, 100, 104172. [Google Scholar] [CrossRef] [Scilit]
  50. Nilsen, V.; Lier, J.A.; Bjerkholt, J.T.; Lindholm, O.G. Analysing urban floods and combined sewer overflows in a changing climate. J. Water Clim. Change 2011, 2, 260–271. [Google Scholar] [CrossRef] [Scilit]
  51. Olcina, J.; Campos Rosique, A.; Casals del Busto, I.; Ayanz López-Cuervo, J.; Rodríguez Mateos, M.; Martínez Puentes, M. Resilience in the urban water cycle. In Rainfall Extremes and Adapting to Climate Change in the Mediterranean Area; Aquae Papers 8; Fundación Aquae: Alicante, Spain, 2018. [Google Scholar]
  52. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.T.; Rocktäschel, T.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2020. [Google Scholar]
Figure 1. Data collected on water disaster-related News vs. Tweets over time from April 2020 to October 2021 (disaster type: storm).
Figure 1. Data collected on water disaster-related News vs. Tweets over time from April 2020 to October 2021 (disaster type: storm).
Water 18 01820 g001
Figure 2. Detected location (country of the ingested tweets triggering alerts on flood and landslide events based on their computed sentiment.
Figure 2. Detected location (country of the ingested tweets triggering alerts on flood and landslide events based on their computed sentiment.
Water 18 01820 g002
Figure 3. The main topics extracted from news articles (on the left) and computed sentiment on landslides and climate change in 2020 with worldwide coverage shows that one of the biggest concerns is disaster relief and resilience occupying 2.6% of the news topics space (on the left), and the sentiment expressed through news outlets on the relation between these topics (on the right).
Figure 3. The main topics extracted from news articles (on the left) and computed sentiment on landslides and climate change in 2020 with worldwide coverage shows that one of the biggest concerns is disaster relief and resilience occupying 2.6% of the news topics space (on the left), and the sentiment expressed through news outlets on the relation between these topics (on the right).
Water 18 01820 g003
Figure 4. The reporting timeline of the news on the coverage of the 2021 floods in Europe (above), and the main concepts captured (below).
Figure 4. The reporting timeline of the news on the coverage of the 2021 floods in Europe (above), and the main concepts captured (below).
Water 18 01820 g004
Figure 5. The reporting timeline of the news on the coverage of the 2020 floods in Indonesia (above), and the main concepts captured (below).
Figure 5. The reporting timeline of the news on the coverage of the 2020 floods in Indonesia (above), and the main concepts captured (below).
Water 18 01820 g005
Figure 6. The reporting timeline of the news on the coverage of the 2020 landslide in India (above), and the main concepts captured (below).
Figure 6. The reporting timeline of the news on the coverage of the 2020 landslide in India (above), and the main concepts captured (below).
Water 18 01820 g006
Figure 7. The reporting timeline of the news on the coverage of the 2021 landslide in Japan (above), and the main concepts captured (below).
Figure 7. The reporting timeline of the news on the coverage of the 2021 landslide in Japan (above), and the main concepts captured (below).
Water 18 01820 g007
Figure 8. The sentiment analysis from news allows us to show the overall neutral and negative feedback on the flood event happening in Alicante on 21 August 2019. Above we have the distribution of collected tweets and news (on the right) with the average of positive, neutral and negative sentiment computed. Below we focus on the news captured in that time window, their computed average sentiment (on the left) and the time series of observed sentiment in the news on that thematic (on the right).
Figure 8. The sentiment analysis from news allows us to show the overall neutral and negative feedback on the flood event happening in Alicante on 21 August 2019. Above we have the distribution of collected tweets and news (on the right) with the average of positive, neutral and negative sentiment computed. Below we focus on the news captured in that time window, their computed average sentiment (on the left) and the time series of observed sentiment in the news on that thematic (on the right).
Water 18 01820 g008
Figure 9. Text mining-based clustering of water pollution subtopics using MEDLINE/PubMed data, integrating insights from research articles and patented flood-related technologies.
Figure 9. Text mining-based clustering of water pollution subtopics using MEDLINE/PubMed data, integrating insights from research articles and patented flood-related technologies.
Water 18 01820 g009
Table 1. Evaluation of the baseline tweet-triggered flood/landslide alert system using manually validated expert labels across the evaluated case studies.
Table 1. Evaluation of the baseline tweet-triggered flood/landslide alert system using manually validated expert labels across the evaluated case studies.
Case StudyAccuracyPrecisionRecallF1AUCκ
2021 Belgium84.0%0.93330.82350.87500.84930.6552
2021 Germany92.0%0.90000.96430.93100.91400.8361
2020 India82.0%0.73330.95650.83020.83010.6457
2020 Indonesia64.0%0.22730.83330.35710.72350.2077
2021 Japan82.0%0.70001.00000.82350.84480.6512
Table 2. The coverage of the water pollution, contamination and flood events across social media, news media and published science articles.
Table 2. The coverage of the water pollution, contamination and flood events across social media, news media and published science articles.
TopicNumber of Collected TweetsNumber of News ArticlesNumber of Scientific Articles
Water Bodies Pollutionnot collected387,5832425
Water-borne Diseases440,03328,8786187
Floods593,4696,877,52418,204
Floods generating diseasenot collected200,908191
Table 3. The coverage of the water pollution and water contamination in the news, and in the MAG and MEDLINE datasets.
Table 3. The coverage of the water pollution and water contamination in the news, and in the MAG and MEDLINE datasets.
Wikipedia TermMeSH Heading TermNr of Languages CoveredNr of NewsNr of PubMed ArticlesNr of Tweets
Trihalomet.Trihalomethane21242110320
ProtozoaProtozoan Infections832,964625911
ChloroformChloroform6719,87162761
SalmonellaSalmonella62188,71769,017135
RotavirusRotavirus5432,74911,24245
Entamoeba histolyticaEntamoeba histolytica4561965845
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Pita Costa, J.; Corzo Perez, G.; Topal, O.; Mikoš, M.; Novalija, I.; Orel, R.; Casals del Busto, I.; Goveas, N. A Machine Learning Approach to Hydrological Event Detection from News-Informed Social Media Alerts. Water 2026, 18, 1820. https://doi.org/10.3390/w18151820

AMA Style

Pita Costa J, Corzo Perez G, Topal O, Mikoš M, Novalija I, Orel R, Casals del Busto I, Goveas N. A Machine Learning Approach to Hydrological Event Detection from News-Informed Social Media Alerts. Water. 2026; 18(15):1820. https://doi.org/10.3390/w18151820

Chicago/Turabian Style

Pita Costa, Joao, Gerald Corzo Perez, Oleksandra Topal, Matjaž Mikoš, Inna Novalija, Rok Orel, Ignacio Casals del Busto, and Neena Goveas. 2026. "A Machine Learning Approach to Hydrological Event Detection from News-Informed Social Media Alerts" Water 18, no. 15: 1820. https://doi.org/10.3390/w18151820

APA Style

Pita Costa, J., Corzo Perez, G., Topal, O., Mikoš, M., Novalija, I., Orel, R., Casals del Busto, I., & Goveas, N. (2026). A Machine Learning Approach to Hydrological Event Detection from News-Informed Social Media Alerts. Water, 18(15), 1820. https://doi.org/10.3390/w18151820

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop