Previous Article in Journal
Aggregated Epidemic Localization and Spatiotemporal Diffusion Modeling Considering Road Network-Constrained Spatial Clustering and Tensor Field Analysis: A Case Study of COVID-19
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Urban Flood-Depth Estimation from Crowdsourced Image–Text Data Using Reference-Object Reasoning and Conditional Fusion

1
School of Environment and Spatial Informatics, China University of Mining and Technology, Xuzhou 221116, China
2
State Key Laboratory of Resources and Environmental Information System, Beijing 100101, China
3
School of Resources and Geosciences, China University of Mining and Technology, Xuzhou 221116, China
*
Author to whom correspondence should be addressed.
ISPRS Int. J. Geo-Inf. 2026, 15(9), 430; https://doi.org/10.3390/ijgi15090430 (registering DOI)
Submission received: 30 July 2026 / Revised: 14 September 2026 / Accepted: 17 September 2026 / Published: 20 September 2026

Abstract

Urban pluvial flooding is a rapidly evolving disaster that requires timely and fine-grained water-depth information for emergency response and impact assessment. However, street-scale water-depth observations are often difficult to obtain during short-duration flood events due to limited coverage and insufficient spatial detail of conventional monitoring systems. To alleviate this problem, this study proposes an image-led framework for urban street-scale flood-depth estimation from crowdsourced image–text data. The proposed method integrates reference-object-based visual evidence construction, textual depth extraction, and conditional image–text fusion to estimate water depth from crowdsourced image–text records. Specifically, image-based depth evidence is constructed by identifying reference objects, determining submerged parts, and establishing object-depth relationships through knowledge mapping. Textual information is further analysed through rule-based extraction and semantic verification to identify quantitative depth cues, including explicit values, ranges, and relative water-level descriptions. Considering the potential inconsistency between image and text observations, a conditional fusion strategy is designed to incorporate textual information only when quantifiability, directional, and image–text discrepancy conditions are satisfied. Experiments demonstrate that the proposed method can effectively extract quantitative flood-depth information from crowdsourced multimodal data and improve the utilization of heterogeneous disaster observations. The resulting depth estimates are intended as auxiliary evidence for post-disaster situational assessment rather than as substitutes for in situ water-level measurements.

1. Introduction

In recent years, global climate change has increased the frequency and intensity of extreme weather events, making urban pluvial flooding a major challenge for urban resilience and emergency management [1]. Due to its localised characteristics and rapid evolution, urban flooding can severely affect transportation systems, infrastructure operations, and emergency response activities. Among various flood-related information, water depth provides a direct indicator of inundation severity and supports flood impact assessment, road accessibility analysis, and emergency decision-making [2,3,4]. Therefore, obtaining timely and accurate street-scale water-depth information after flood events is essential for improving disaster response capabilities.
However, acquiring fine-grained water-depth observations during short-duration flooding events remains challenging. Conventional monitoring systems, such as hydrological stations and sensor networks, usually provide limited spatial coverage and may not capture rapidly changing local inundation conditions [2,5,6]. To address this limitation, researchers have explored alternative data sources, particularly image-based approaches, for flood-depth estimation [2,7,8]. Existing image-based methods mainly rely on supervised learning, geometric measurement, or reference-object reasoning [4,9]. Although these approaches have achieved promising results, their application in complex post-disaster environments remains challenging due to variations in observation conditions, the need for scenario-specific information, and uncertainties in interpreting visual flood evidence.
The rapid development of social media platforms has provided new opportunities for disaster monitoring and emergency response. Crowdsourced data, particularly image–text records shared by affected individuals, can provide timely and near-ground observations of disaster situations [10,11,12]. Previous studies have demonstrated the potential of social media data for disaster-related applications, including affected-area identification, location recognition, damage assessment, and situational awareness [13,14,15]. Compared with traditional monitoring systems, crowdsourced records can capture detailed local conditions that are difficult to obtain through fixed observation networks [5,16]. For urban flooding, images can provide direct visual evidence of water surfaces, roads, and submerged objects, while accompanying texts may contain complementary information, such as locations, timestamps, numerical depths, and relative water-level descriptions [17].
Nevertheless, extracting quantitative water-depth information from crowdsourced multimodal data remains challenging. Visual observations usually require identifying appropriate reference objects and establishing relationships between visible submerged parts and water depth. Furthermore, textual descriptions often contain implicit expressions, multiple values, temporal variations, and spatial ambiguity, which makes direct depth extraction difficult. Existing studies have shown that geolocated social media data contain semantic and spatial uncertainties [18,19,20], while multimodal information may involve differences in quality and potential conflicts [21]. Therefore, image and text should not be simply regarded as equivalent observations; instead, their complementary roles and reliability should be carefully considered when integrating crowdsourced multimodal information.
To address these challenges, this study develops an image-led framework for estimating street-scale urban flood depth from crowdsourced image–text records. Unlike approaches that directly generate depth values from images or combine image and text as equivalent observations, the proposed framework first converts visual observations into structured depth evidence from reference objects and then treats textual depth information as conditional auxiliary evidence. The image module identifies reference objects and submerged parts that are assigned to depth classes based on explicit knowledge relationships. The text module combines rule-based extraction with semantic review to assess whether textual depth cues are quantifiable. Text is incorporated into the final estimate only when predefined directional and image–text discrepancy conditions are satisfied. The methodology therefore focuses on controlling cross-modal inconsistency rather than assuming that multimodal fusion necessarily improves depth estimation.
The main contributions of this study are summarised as follows:
We propose an image-led framework for urban flood-depth estimation from crowdsourced image–text data. The framework treats images as the primary source of depth evidence and integrates textual information only as conditional auxiliary evidence, thereby avoiding direct equivalence between the two modalities.
We develop a structured depth-evidence construction strategy based on reference-object reasoning and text quantifiability assessment. The image module identifies reference objects and submerged parts and maps them to depth values through explicit knowledge relationships, while the text module combines rule-based extraction and semantic review to identify quantifiable depth cues from complex disaster descriptions.
We design a conditional image–text fusion strategy to control textual participation under cross-modal inconsistency. Textual evidence is incorporated only when quantifiability, directional, and image–text discrepancy conditions are satisfied, limiting the influence of temporally or spatially inconsistent descriptions on the final depth estimate.

2. Related Work

2.1. Crowdsourced Social Media Information for Urban Flood Sensing

User-generated social media content can serve as volunteered geographic information for urban flood monitoring and complement conventional sources such as fixed water-level sensors, remote sensing imagery, and hydrological stations [5,10,16]. These records are commonly produced near disaster events, provide near-ground observations of inundated roads, transport disruption, and public risk perception, and support dynamic, fine-grained analysis of urban flooding [11,22,23].
Previous studies have used natural language processing and machine-learning methods to extract spatial information, including inundated locations, road names, and landmarks, from social media text [24,25,26]. Water-depth quantification additionally requires explicit values and relative part descriptions to be distinguished together with their temporal and spatial contexts [27,28]. Consistent extraction of quantifiable water-depth information from unstructured disaster text remains challenging [29].
Flood data from social media also contain semantic and spatial uncertainty. Some records lack directly usable depth values, location tags may be coarse, and text may describe historical water levels, nearby locations, or subjective event severity [18,30,31]. Crowdsourced depth quantification therefore requires not only the identification of depth-related textual cues, but also assessment of their quantifiability and their spatiotemporal correspondence with the image scene.

2.2. Image-Based Urban Flood-Depth Estimation and Image–Text Evidence Fusion

Near-ground social media images and surveillance videos can record water surfaces, inundated objects, water boundaries, and scene context, thereby providing visual evidence for street-scale depth assessment [3,32,33]. Existing image-based methods mainly comprise whole-image classification or regression, geometric relationship estimation, and reference-object reasoning [4,9]. Supervised deep-learning models have been applied to urban flood-depth classification and continuous depth estimation, but usually require task-specific annotations, while the basis of their predictions and their stability across scenes remain constrained [34,35,36].
Geometry-based methods can estimate water depth using fixed cameras [37], optical refraction, or high-resolution digital elevation models [38,39]. These methods impose explicit physical constraints, but usually require stable viewpoints, camera calibration, or high-accuracy terrain data, which limits their direct application to social media images with unknown camera parameters, substantial viewpoint variation, and frequent occlusion [40,41].
Reference-object methods infer water level from identifiable parts of common urban objects, including pedestrians, cars, buses, bicycles, and roadside infrastructure. Human body parts provide references under different inundation conditions [42], while vehicle components such as tyres, doors, and windows have also been used to identify urban flood-water levels or depth classes [43,44,45]. These methods nevertheless remain affected by object visibility [46,47], water-boundary clarity, and whether the reference object represents the principal flooded area.
Large multimodal models have been applied to semantic interpretation and depth estimation in flood images, including the recognition of pedestrians, vehicles, water surfaces, and the condition of inundated objects [48,49]. Existing evaluations indicate that direct depth generation remains dependent on task prompts and output constraints, and that structured protocols or external knowledge are needed to reduce instability in object selection and depth quantification [50].
Previous work has already integrated social media image and text data for urban flood-depth estimation. Chu et al. [51] developed SDPO-MLLM, in which text depth extraction, image depth description, and water-level classification were formulated as multimodal learning tasks and optimised using SFT and segment-level DPO before the extracted results were quantified and integrated. In contrast, the present study does not rely on a unified fine-tuned MLLM for multimodal depth extraction. Image and text evidence are constructed separately, with the image retained as the primary source of local depth evidence and the text treated as conditional auxiliary evidence. General cross-modal studies further indicate that modalities may differ in information quality and that direct fusion can be affected by semantic ambiguity and cross-modal conflict [21,52]. The experiments therefore focus on whether and under what conditions textual evidence should modify an image-based estimate, rather than assuming that multimodal information should always be combined. Image-only, text-only, direct-average, fixed-weight, and conditional-fusion strategies are compared on the same fusion set to examine the effects of text quantifiability, directional constraints, and image–text discrepancy control. Structured disaster information derived from social media can be integrated with road networks, remote sensing imagery, and foundational geospatial data to support road-accessibility risk assessment, disaster-severity mapping, and emergency transport-loss evaluation [52,53,54]. Accordingly, the proposed workflow combines explicit reference-object knowledge mapping with conditional textual participation, allowing the contribution of each modality to be examined separately under cross-modal inconsistency rather than assuming a uniform benefit from multimodal fusion.

3. Materials and Methods

3.1. Construction and Analysis of the Crowdsourced Multimodal Dataset

Images related to urban flooding and their accompanying text published on Weibo [55] are used as the primary data source to construct a crowdsourced multimodal dataset for image-based depth quantification, text-based depth extraction, and image–text fusion evaluation. Data were collected mainly between 23 July and 1 August 2025, with Beijing as the principal study area. Retrieval was controlled using keywords including Beijing rainstorm, urban flooding, and road inundation. More than 17,558 public Weibo posts were collected, yielding 29,122 traceable single-image records, with additional image samples extracted from Weibo videos. The 29,122 traceable images constituted the initial candidate pool rather than the final experimental dataset. The initial image pool was first subjected to automatic flood/non-flood screening to remove irrelevant images. The remaining candidates were then manually reviewed to exclude residual irrelevant images and duplicate records. After this two-stage screening process, 1283 images were retained and manually annotated for the subsequent image-based experiments. The annotated image dataset was further divided into training, validation, and test subsets, with 193 images retained as the fixed test set for comparison among the image-based methods. After image screening, the accompanying texts paired with the retained image records were further screened for the text-based experiment. Texts unrelated to water-depth descriptions and duplicate text records were removed. Among the remaining records, 239 cleaned text samples containing annotatable depth information were retained and manually assigned text reference depths for the text-extraction experiment. Image reference depth was manually assigned from the relative relationship between visible reference-object parts and the water surface, with primary reference to the correspondence of human body parts and vehicle components under different inundation conditions [42,43]. Text reference depth was annotated solely from the accompanying text without using image information. All depth values were recorded in centimetres. Two annotators independently reviewed the samples, and disagreements were resolved through joint review. Because the image reference depths were inferred from visible scene evidence rather than field measurements, they should be interpreted as scene-based reference estimates rather than centimetre-level in situ observations. Image and text reference depths were annotated independently to avoid cross-modal information leakage. Relative-part descriptions from both modalities were converted using the same object-part and depth mapping to maintain a common depth representation.
The training and validation subsets were used for model training and parameter selection of the supervised ResNet-50 [56] and ViT-B/16 [57] models. The fixed test set of 193 images was used to compare the supervised models, direct vision–language model inference, and the proposed vision–language framework for the depth based on reference objects (RPD-VLM). Table 1 summarises the sample size, inclusion criteria, and corresponding section for each evaluation set.
The image-depth, text-extraction, image–text fusion, and ablation experiments used different task-specific evaluation sets. In particular, the 193-image test set was used to compare image-based methods, whereas the 202-sample fusion set consisted of paired image–text records for which an image-based prediction, a text-based prediction, and an image reference depth were all available. Therefore, these sample sets represent different experimental tasks rather than successive stages of a single filtering process. Table 1 summarises the sample size, inclusion criteria, and corresponding role of each experimental set.
The depth-estimation experiments were conducted at level of individual records, and precise sample-level georeferencing was not used as an input or evaluation criterion in the current experiments. The geographic scope was controlled by event- and location-related retrieval terms, whereas the positional accuracy of individual social media records was not independently evaluated.

3.2. Overall Framework

As shown in Figure 1, the overall study workflow comprises five stages: (1) crowdsourced multimodal dataset construction; (2) image-based depth-evidence construction using RPD-VLM; (3) textual depth-evidence extraction and quantifiability assessment using RPD-VLM-T; (4) image-led conditional fusion using RPD-VLM-Fusion; and (5) experimental evaluation and comparison. After dataset construction, the image and text pathways operate in parallel to produce image-based depth evidence D i m g and textual depth evidence ( D t x t , q t x t ), respectively. These outputs are combined only at the subsequent conditional-fusion stage. The experimental evaluation separately examines textual extraction, fusion strategies, and image-based depth estimation, including supervised and direct-VLM baselines. Precise geographic coordinates are not estimated in the current framework; location cues, when available from platform metadata or textual place names, require independent geocoding and positional validation before downstream spatial analysis.
The image module is the primary source of depth information. It uses a vision–language model (VLM) to identify the principal flooded area, candidate reference objects, and their spatial relationships with the water surface, and then generates image-based depth through reference-object knowledge mapping. The text module provides candidate auxiliary evidence. It first applies rules to parse explicit values, ranges, and common relative part descriptions, and then performs semantic review for complex text involving negation, multiple numbers, temporal change, spatial reporting, or colloquial expressions. Its outputs include text-based depth, text quantifiability, expression type, and evidence span. The fusion module calculates the image–text discrepancy on a common depth scale and determines whether text participates in the final estimate according to its quantifiability and a continuously decaying weight.

3.3. Construction of Image-Based Depth Evidence Through Reference-Object Knowledge Mapping

This image-based procedure constitutes the RPD-VLM framework introduced in Section 3.1. The image module converts crowdsourced flood images into structured evidence containing a reference object, submerged part, and water-depth value. The model first identifies the principal flooded area and candidate reference objects and then assesses the relative relationship between key object parts and the water surface. A reference-object knowledge base subsequently maps the object type and submerged part to a discrete depth class and its representative depth.
Image interpretation follows a two-stage procedure. The first stage locates the principal flooded area and identifies candidate reference objects. The second stage assesses the object–water relationship and submerged part within the candidate set. Qwen3-VL-8B-Instruct [58] is used to identify candidate reference objects and interpret object–water relationships, but does not directly generate the final depth value. A retrieval-augmented knowledge context is introduced between the two stages of image interpretation. After the first-stage detection, the identified reference-object types are normalised to predefined categories. The detected object categories are then used as retrieval keys to query a structured reference-object knowledge base. Only knowledge entries associated with the detected object types are retrieved, including candidate object parts, corresponding depth classes, and representative depth values. The retrieved entries are inserted into the second-stage prompt as contextual constraints for interpreting the object–water relationship and selecting the corresponding submerged-part key. The retrieved context does not directly provide the final depth estimate. Instead, after the VLM identifies the reference-object type and submerged part, the final image-based depth is assigned through the explicit object-part depth mapping described below. If no valid pair of reference objects can be established, no deterministic image-based depth is returned.
For depth quantification, this study follows the multi-reference-object classification scheme described in [51], mapping key parts of common urban objects, including people, non-motorised vehicles, and motor vehicles, to a common set of depth classes. The knowledge base is grounded in typical elevations of key human body parts and incorporates corresponding component-height relationships for bicycles, motorcycles, cars, and buses. It defines mapping from reference-object type and submerged part to depth class and representative depth, as illustrated in Figure 2.
Let the type and submerged part of candidate reference object i be denoted by o i and p i , respectively. Its depth class is:
l i = K l e v e l o i , p i
The corresponding representative depth is:
d i = K d e p t h l i
where K l e v e l maps the reference-object type and part to a depth class, K d e p t h maps the depth class to a representative depth, and d i is measured in centimetres.
When multiple valid reference objects occur in the same image, only candidates located within the principal flooded area, with a clear object–water relationship and visible key parts, are retained. Let C m a i n denote the set of valid candidates in the principal flooded area. The candidate with the largest mapped depth is selected:
c = a r g m a x c i C m a i n d i
The image-based depth is defined as:
D i m g = d c
This selection rule reduces the influence of shallow margins, local puddles, and areas outside the principal inundated zone. Before mapping, reference-object types and submerged parts are standardised to the categories defined in the knowledge base.
To examine how uncertainty associated with reference-object dimensions and the object-part depth mapping may propagate to the final numerical estimate, a scenario-based proportional sensitivity analysis was conducted. The perturbation magnitude of ±10% was selected following the sensitivity setting adopted in the previous reference-object-based flood-depth study [51]. This provides a moderate and symmetric scenario for evaluating the response of the estimates while maintaining methodological comparability with the previous study.
The ±10% perturbation was applied to both the representative depth levels and the final RPD-VLM estimates according to the following equation:
D δ = 1 + δ D , δ { 0.10 , 0 , + 0.10 }
where D denotes the depth value being perturbed. For the depth-level bound analysis, D represents a representative depth level defined in Figure 2. For the performance-propagation analysis, D represents the final RPD-VLM image-depth estimate for each sample. Reference-object recognition, submerged-part interpretation, manually assigned reference depths, and sample composition were kept unchanged. MAE, RMSE, MedAE, Acc@10, and Acc@20 were recalculated on the same common valid subset of 185 samples used for the main image-based comparison. Consequently, the 0% condition reproduces the baseline results reported in Section 4.3.1. The ±10% setting represents a controlled sensitivity scenario rather than an empirically estimated confidence interval for actual human or vehicle dimensions.

3.4. Extraction and Quantifiability Assessment of Depth Evidence from Social Media Text

The text module identifies quantifiable water-depth cues in the accompanying social media text and determines whether they can form an explicit text-based depth result, as shown in Figure 3. The output fields include text-based depth D t x t , text quantifiability q t x t , expression type, and evidence span.
Text-based depth evidence includes explicit values, interval expressions, and relative part descriptions. For explicit values, the rule layer identifies the numerical value and unit and converts them to centimetres:
D n u m = U x , u
where x is the numerical value in the text, u is the unit, and U is the unit-normalisation function.
For interval expressions with explicit lower and upper bounds, the midpoint is used:
D r a n g e = a + b 2
where a and b are the lower and upper bounds, respectively. Vague degree expressions without identifiable bounds are not forced into deterministic depth values.
For relative part descriptions, such as water reaching the ankle or covering half of a tyre, the reference-object type o t and part relationship p t are first identified and then converted to a representative depth using the same knowledge mapping as the image module. This shared mapping places relative-part evidence from the two modalities on the same depth scale before image–text discrepancy calculation:
D p a r t = M t e x t o t , p t
These results only form candidate text-based depths and do not indicate that the text is aligned with the time, location, or inundated area shown in the image.
Text containing multiple numbers, negation, temporal change, spatial reporting, range comparisons, or descriptions of a non-current scene is marked as complex and triggers semantic review. The semantic interpretation module uses the original text together with the rule-based candidate to determine whether the text refers to an explicit inundation state and whether the candidate should be retained, revised, or rejected.
The final text-based depth is defined as:
D t x t = D R , v a l i d   r u l e   r e s u l t   f o r   s i m p l e   t e x t , D S , v a l i d   s e m a n t i c r e v i e w   r e s u l t   f o r   c o m p l e x   t e x t , , n o   e x p l i c i t   d e p t h   v a l u e .
where D R is the rule-based result, D S is the semantic-review result, and indicates that the text module does not output a deterministic depth. This process yields D t x t and its quantifiability state q t x t . A value of q t x t = 1 indicates only that the text can be quantified; it does not establish consistency with the image scene. Text participation in the final estimate is determined subsequently by the directional constraint and image–text discrepancy control.

3.5. Image-Led Depth-Evidence Fusion Accounting for Textual Uncertainty

The fusion stage retains image-based depth D i m g as the primary estimate and treats text-based depth D t x t as candidate auxiliary evidence. The manually assigned reference depth is used only for experimental evaluation and does not enter the calculation of text participation or weight.
The discrepancy between image-based and text-based depth is defined as:
Δ d = D i m g D t x t
Text participation first requires q t x t = 1. Subject to this condition, text is allowed to provide a downward correction only when D t x t < D i m g ; otherwise, the image-based estimate is retained. This asymmetric rule is a conservative fusion constraint rather than a physical assumption that textual reports systematically underestimate water depth. Social media text may either under- or overstate the local inundation state because of temporal or spatial mismatch and subjective reporting. The constraint therefore prevents text from increasing the image-based estimate when cross-modal correspondence cannot be established.
When the quantifiability and directional conditions are satisfied, the text weight decays continuously with the image–text discrepancy. With a maximum text weight of 0.5, the weight is defined as:
w t x t = 0.5 , q t x t = 1 , D t x t < D i m g , Δ d 27 , 0.5 60 Δ d 60 27 , q t x t = 1 , D t x t < D i m g , 27 < Δ d 60 , 0 , else .
The final depth estimate is:
D f i n a l = 1 w t x t D i m g + w t x t D t x t = D i m g + w t x t D t x t D i m g
When the text is not quantifiable, the directional constraint is not satisfied, or the image–text discrepancy exceeds 60 cm, w t x t = 0 and the final result equals D i m g . When the discrepancy does not exceed 27 cm, text participates with the maximum weight. Between 27 and 60 cm, the text weight decreases linearly as the discrepancy increases.

3.6. Experimental Design and Evaluation Metrics

The experiments comprise three main tasks: textual depth-evidence extraction, image–text fusion, and image-based depth estimation. Textual depth extraction is first evaluated, followed by analysis of the conditions under which textual evidence participates in image–text fusion. The image-based depth method is then compared with supervised and direct-VLM baselines on the unified image test set. Additional analyses examine performance under different training-data proportions, depth ranges, and scene types, together with ablation experiments and representative cases. The framework does not assume a general preference for methods without task-specific supervised training. Supervised models are included as performance baselines, whereas RPD-VLM is investigated as a VLM-based inference pathway that does not require task-specific end-to-end depth-regression training.
The sample sets and corresponding inclusion criteria are summarised in Table 1. Because the text-extraction, image–text fusion, image-based depth estimation, and ablation experiments use different task-specific sample sets, methods are compared only within the same evaluation set. Text-extraction methods are evaluated against the text reference depth, whereas image-based depth estimation and image–text fusion are evaluated against the image reference depth. The fusion experiment uses samples for which an image-based prediction, a text-based prediction, and an image reference depth are all available. The image-depth comparison uses the unified image test set, with error metrics calculated on the common valid subset when direct comparison among methods is required. Coverage and extraction counts are calculated over the corresponding task-specific sample pool.
For continuous depth estimation, mean absolute error (MAE), root mean square error (RMSE), and median absolute error (MedAE) are used to quantify the difference between predicted and reference depths. Acc@10 and Acc@20 denote the proportions of samples with absolute errors no greater than 10 cm and 20 cm, respectively, whereas Out@30 and Out@50 denote the proportions with absolute errors greater than 30 cm and 50 cm. The text-extraction experiment additionally reports the number of successful extractions and extraction coverage. For paired comparisons between RPD-VLM and the baseline methods, statistical uncertainty was assessed using paired bootstrap resampling on the common valid subset. For each comparison, the 185 paired samples were resampled with replacement 10,000 times while preserving the correspondence between the predictions of the two methods and the same reference depth. Differences in MAE, RMSE, and Acc@20 were calculated for each resample. The 2.5th and 97.5th percentiles of the bootstrap distribution were used to form the 95% confidence interval (CI). A difference was considered statistically stable when the corresponding 95% CI did not include zero.

4. Results

4.1. Overview of Experimental Results

Table 2 summarises representative results for image–text fusion and image-based depth estimation. The complete textual depth-extraction results are reported separately in Section 4.2.1 to avoid duplication. Because these tasks use different sample sets, reference-depth definitions, and evaluation objectives, the reported metrics are interpreted only within each task and do not constitute a ranking across tasks. Detailed results are presented in Section 4.2 and Section 4.3.

4.2. Textual Depth-Evidence Extraction and Image–Text Fusion with Difference Control

This section evaluates the role of textual depth evidence in image–text fusion. Rule-only and RPD-VLM-T are first compared in terms of extraction coverage and depth error on quantifiable text. Image-only, text-only, unconditional fusion, fixed-weight fusion, and directionally constrained continuous-decay fusion are then compared on the same valid fusion set. The objective is not to establish that text generally outperforms images, but to examine whether text can participate as candidate corrective evidence when it is quantifiable, satisfies the directional condition, and remains within the controlled image–text difference range.

4.2.1. Textual Depth-Evidence Extraction

After the screening procedure described in Section 3.1, 239 cleaned accompanying texts with annotatable reference depths were retained as the input evaluation set for the text-extraction experiment. Both rule-only and RPD-VLM-T were applied to these same 239 samples. Of these 239 input samples, rule-only produced valid deterministic depth outputs for 143 samples, whereas RPD-VLM-T produced valid outputs for 221 samples. Thus, 143 and 221 represent method-specific successful extraction counts rather than different input sample sets. Errors are calculated against the textual reference depth Dref,txt. Rule-only is a deterministic baseline designed primarily for explicit numerical values and standard relative-part expressions. RPD-VLM-T first generates a rule-based candidate and triggers semantic review for complex text containing multiple numbers, negation, temporal change, spatial relay, or colloquial description, after which the candidate is retained, revised, or rejected. Table 3 summarises the overall extraction coverage and error results for the quantifiable text samples.
As shown in Table 3, RPD-VLM-T increases the number of successful extractions from 143 to 221 and the extraction coverage from 59.83% to 92.47%, indicating that semantic parsing recovers complex expressions not handled by the rules. The error metrics are calculated separately over the samples successfully extracted by each method. Because the two sets differ in size and textual difficulty, the table describes coverage and the quality of each method’s own outputs rather than providing a strict comparison of numerical accuracy. Table 4 reports the depth-extraction errors on the samples successfully extracted by both methods.
For the 143 samples successfully extracted by both methods, rule-only has lower MAE and RMSE and higher Acc@10 and Acc@20 than RPD-VLM-T. RPD-VLM-T therefore expands the coverage of textual depth evidence but does not improve the numerical accuracy of rule candidates on the common subset; some semantic reviews introduce additional error. The role of the text-extraction method is consequently interpreted as supplementing and reviewing complex expressions rather than providing a general accuracy gain over deterministic rules.

4.2.2. Comparison of Text-Participation Strategies and Fusion Performance

It should be noted that the text-only result in the fusion experiment differs conceptually from the text-extraction evaluation in Section 4.2.1. In Section 4.2.1, RPD-VLM-T is evaluated against the text reference depth D ref , txt , which measures whether the depth described in the text is correctly extracted. In the fusion experiment, text-only is evaluated against the image reference depth D ref , img , which measures whether the text-derived depth can directly represent the inundation depth depicted in the corresponding image. Therefore, the difference between the two results reflects not only different valid sample sets but also different evaluation targets and reference-depth definitions.
To examine the effect of different text-participation strategies, image-only, text-only, direct average, fixed weight, and directionally constrained continuous-decay fusion are compared on the same valid fusion set, as shown in Table 5.
Note: All error metrics use the image reference depth D ref , img as the evaluation benchmark. Total N denotes the number of samples in the common fusion evaluation set, with N = 202 for all strategies. Text-only uses the text-derived depth as an estimate of the corresponding image-scene depth and is therefore evaluated against D ref , img , rather than D ref , txt used in the text-extraction experiment. Participating N is the number of samples in which text enters the corresponding calculation; for RPD-VLM-Fusion, participation is defined by W t x t > 0. Changed N is the number of final predictions that differ from image-only.
Table 5 shows that unconditional text participation is associated with higher errors on this valid fusion set. Text-only has higher error than image-only, and neither direct average nor fixed weight improves on the image-only result. Social media textual depth should therefore not be treated directly as an equivalent observation of the local depth depicted in the current image. Under the image-led design, RPD-VLM-Fusion restricts the range of text participation. Relative to image-only, it shows lower MAE and RMSE, with little change in Acc@20. Text is therefore used only as candidate corrective evidence when the quantifiability, directional, and image–text difference conditions are satisfied. Without these controls, temporal, spatial, or object-level inconsistency may increase estimation error.

4.2.3. Directional Constraint and Image–Text Difference Control

A post hoc diagnostic analysis is conducted on the 202 valid fusion samples to examine the directional constraint and image–text difference control. As shown in Table 6, during fusion, text participation is still determined only from D t x t , D i m g , the text-quantifiability status, and the image–text difference Δ d .
Table 6 shows that text predictions are closer to the image reference depth more often when D t x t < D i m g than when D t x t D i m g . Therefore, D t x t < D i m g is used as a candidate direction for subsequent difference control, but it does not imply that text is necessarily more accurate than the image-based estimate. The directional constraint limits the entry of textual uncertainty into the final estimate under the image-led framework. This empirical asymmetry is specific to the present dataset and is not interpreted as a universal physical relationship between textual and image-derived water depths.
Among samples satisfying the directional constraint, the image–text discrepancy further determines the degree of text participation. Figure 4 shows that error reductions occur mainly when the discrepancy is small. When the directional constraint is not satisfied or Δ d > 60 cm, the result falls back to image-only. The directional constraint and discrepancy control therefore primarily restrict inconsistent text from entering the final estimate rather than increasing the number of text-participating samples. This observation is consistent with the image-dominated, conditionally text-assisted fusion design described in Section 3.5.

4.2.4. Image–Text Fusion Case Analysis

To illustrate text participation under different image–text discrepancy conditions, Figure 5 presents three representative samples corresponding to maximum-weight participation, continuously decayed participation, and discrepancy-based fallback. The cases are used only to explain the weight calculation and resulting changes. The manually assigned reference depth D r e f is used solely to compare errors after fusion and does not enter the calculation of the text weight or final depth.
In case (a), the text-based depth is lower than the image-based depth and the image–text discrepancy does not exceed 27 cm, satisfying the quantifiability, directional, and discrepancy-control conditions. With the maximum text weight, the final depth shifts from the image prediction towards the text prediction and becomes closer to the manually assigned reference depth. This case shows that text can serve as candidate corrective evidence when the two depth estimates are close and the directional condition is satisfied.
In case (b), the image–text discrepancy lies between 27 and 60 cm. As the discrepancy increases, the text weight decreases continuously from its maximum value. The final estimate remains dominated by the image prediction and is adjusted only to a limited extent. Compared with fixed-weight fusion, this mechanism limits the influence of text when the discrepancy is larger, although it does not ensure an equal reduction in error for every participating sample.
In case (c), the image–text discrepancy exceeds 60 cm, the text weight is set to 0, and the result falls back to the image-based depth. The text description differs substantially from the local inundation state shown in the image and may refer to another location, another object, or a non-current scene. Rejecting text participation in this case prevents a large cross-modal discrepancy from directly changing the image-based estimate.
These cases show that the method does not treat text as a consistent source of accuracy gain. Instead, text participation is limited according to quantifiability, the directional constraint, and image–text discrepancy. Text can provide candidate correction when the discrepancy is small, its influence decreases as the discrepancy increases, and the image-based depth output is retained beyond the rejection boundary.

4.3. Image-Based Depth-Quantification Results

In this section, RPD-VLM is not intended to replace fully supervised depth-regression models. Instead, it provides an alternative inference pathway that requires less task-specific end-to-end training and retains explicit intermediate evidence, including the reference-object type, submerged part, and mapped depth class. The following experiments therefore compare its accuracy with supervised and direct-VLM baselines while also examining its valid-prediction coverage, few-sample behaviour, and interpretable intermediate reasoning process.

4.3.1. Overall Quantitative Results

Valid-prediction coverage is first calculated for each method on the unified test set. Supervised regression models return a continuous depth for every test image, whereas direct VLM inference and RPD-VLM require interpretable depth cues in the image and therefore do not produce a valid estimate for a small number of samples. Table 7 reports valid-prediction coverage for all methods on the unified test set.
To ensure direct comparability of the error metrics, the main performance comparison is restricted to the common subset for which every method produces a valid continuous prediction. Table 8 presents the overall depth-estimation results on the common valid subset.
Table 8. Overall depth-estimation results on the common valid subset.
Table 8. Overall depth-estimation results on the common valid subset.
MethodCommon NMAE (cm)RMSE (cm)MedAE (cm)Acc@10Acc@20Out@30Out@50
ResNet-5018512.918.88.50.57300.81620.10270.0270
ViT-B/1618510.817.76.60.65410.83240.06490.0216
VLM-only18527.333.125.00.19460.45950.38380.0541
VLM + CoT18524.228.720.00.25410.54050.24860.0378
RPD-VLM18514.922.010.00.64860.80000.14590.0216
Note: Common N is the number of samples with a valid continuous depth prediction from every compared method.
Table 9 reports the paired bootstrap comparisons between RPD-VLM and the baseline methods.
Table 8 and Table 9 show that ViT-B/16 achieved the lowest error on the common valid subset, while ResNet-50 also performed stably. RPD-VLM had higher error than the supervised models but clearly outperformed VLM-only and VLM + CoT, indicating that structured reference-object reasoning reduces the instability of direct depth generation. Accordingly, RPD-VLM is evaluated in this study as a structured alternative to direct VLM depth generation rather than as a replacement for fully supervised depth-regression models.
RPD-VLM has lower error than VLM-only and VLM + CoT on the current common valid subset. This result indicates that structuring the image evidence can reduce the variability associated with direct continuous numerical output from a vision–language model.
Direct continuous outputs from the open-ended vision–language models are relatively dispersed in this experiment. RPD-VLM instead converts visual information into a reference-object type, submerged part, and discrete depth class before assigning a numerical depth, thereby retaining explicit intermediate interpretation steps. Figure 6 compares the predicted depths and manually assigned reference depths on the common valid subset against the 1:1 line. It is used to examine dispersion and systematic deviation; the numerical evaluation is provided in Table 8 and Table 9.
Figure 6 shows that the supervised models have more concentrated point distributions, direct VLM inference is more dispersed, and RPD-VLM lies between them. In the range containing most shallow-water samples, the RPD-VLM points are less dispersed than those from direct VLM inference. Departures from the reference line remain for intermediate- and high-depth samples, suggesting that reference-object visibility, object–water relations, and the spacing of discrete depth classes may limit image-based depth estimation.

4.3.2. Model Performance Trends Under Few-Sample Conditions with a Fixed Split

To examine the effect of task-annotation volume on supervised deep models, additional experiments train ResNet-50 and ViT-B/16 using different proportions of a fixed training split. Samples for each proportion are drawn from the same training partition, while the validation and test sets remain unchanged. RPD-VLM does not vary with the training proportion and serves as a fixed reference with low dependence on task-specific annotations.
Figure 7 shows changes in MAE and Acc@20 for ResNet-50, ViT-B/16, and RPD-VLM at different training proportions. The supervised models are retrained for each proportion and evaluated on the fixed full test set (N = 193), while the RPD-VLM result from the common valid subset (N = 185) is included as a reference.
Under this fixed split, the supervised models generally improve as the training proportion increases, indicating that end-to-end regression is sensitive to the amount of task-specific annotation. Because RPD-VLM does not require end-to-end depth-regression training on the current task, its result remains fixed across the different training proportions. It is therefore included as a reference for comparing the sensitivity of supervised models to task-specific training-data volume, rather than as evidence of superior few-sample performance.

4.3.3. Error Distributions Across Manually Assigned Depth Ranges

To examine applicability under different inundation depths, the samples are divided by manually assigned reference depth into 0–20 cm, 20–50 cm, and 50–100 cm groups, and the error differences among methods are calculated. The grouped analysis uses the 190 test samples for which RPD-VLM produces a valid prediction. Only three samples exceed 100 cm; these are omitted from the main table and considered only for qualitative interpretation. Table 10 summarises the error distributions across the manually assigned depth ranges.
Table 10 shows relatively low errors in the 0–20 cm shallow-water range. In these images, cues such as the road surface, water boundary, lower legs, and vehicle tyres are generally easier to interpret, and the image-based depth evidence produced by RPD-VLM is comparatively stable within this range.
Errors are higher for all methods in the 20–50 cm range than in the shallow-water group, with a larger increase for RPD-VLM. This pattern may be associated with transitional submerged positions around the knee, middle of a wheel, or kerb height, where a small part-assessment error can be converted into a larger continuous depth error.
Errors increase further in the 50–100 cm group. Visible reference objects are less frequent, key parts are more often occluded, and water boundaries are less clear, which may make object–water assessment and discrete-class mapping more difficult. Because the sample size in this range is limited, the result indicates a need for further evaluation of deep-water images rather than a stable general pattern.
The grouped results in Table 10 are based on the same 190 test samples with valid RPD-VLM predictions, and ResNet-50, ViT-B/16, and RPD-VLM are evaluated on this common set. The three samples above 100 cm are not reported separately in the main table. The grouped results are therefore used to interpret error sources and applicability boundaries across depth ranges, rather than to support a general conclusion for deep-water scenes.

4.3.4. Error Heterogeneity Across Scene Types

In addition to depth range, the type of urban spatial scene may be associated with the stability of image-based depth evidence. Scene types are assigned manually from the dominant spatial environment in each image. When several scene elements are present, the type is determined by the scene containing the main flooded area; samples that cannot be classified clearly are assigned to other/unclear. Table 11 presents the RPD-VLM error distributions across the manually assigned scene types.
Among the main scene categories with larger sample sizes, road scenes show a comparatively stable error distribution. These images commonly contain reference objects such as vehicle tyres, pedestrians’ lower legs, and kerbs, and the flooded area is more readily associated with the road plane, making object–water relations easier to interpret.
Parking-area scenes have higher errors. Possible reasons include the number of vehicles, complex occlusion, and the presence of multiple flooded areas with different depths in the same image. A single visible vehicle part may not represent the main flood depth, producing a discrepancy between the image-based evidence and the manually assigned reference depth.
Error sources are more heterogeneous in residential, bridge or culvert, and low-lying scenes. Residential scenes may be affected by complex backgrounds, inconsistent reference-object scales, and an unclear main flooded area. Bridge, culvert, and low-lying scenes more often contain deeper water, extensive water surfaces, low illumination, and perspective distortion, making the relation between the water surface and object parts more difficult to assess. The other/unclear category contains few samples and is included only for descriptive context. Figure 8 visualises the distribution of RPD-VLM absolute errors across scene types.

4.3.5. Ablation and Sensitivity Analyses for Image-Based Depth Quantification

Ablation experiments on the unified test set evaluate the contribution of the modules used for image-based depth quantification. The settings remove, in turn, the retrieval-augmented generation (RAG) knowledge context, main-flooded-area prioritisation, the first-stage candidate-reference-object constraint, and the explicit knowledge mapping of reference-object depth. Table 12 specifies the removed component and replacement mechanism for each setting so that the structural change associated with each result is explicit.
To avoid differences in valid sample counts affecting the comparison, ablation metrics are calculated only on the common subset for which the complete RPD-VLM and every ablation setting produce valid predictions. Table 13 reports the ablation results for image-based depth quantification on the common valid subset.
Table 14 reports the paired bootstrap tests comparing the complete RPD-VLM with the ablation settings.
Table 13 and Table 14 show that removing any module is associated with higher MAE and RMSE on the common valid subset. The increases are larger after removing the first-stage candidate-reference-object constraint or the knowledge mapping, indicating that candidate-object restriction and part-to-depth conversion are important components of image-based depth-evidence construction. Differences in MAE and RMSE consistently favour the complete RPD-VLM over the ablation settings. Acc@20 generally changes in the same direction, although the confidence intervals for the RAG-context and main-flooded-zone ablations include 0 and are therefore treated as supporting evidence only.
The RAG knowledge context and main-flooded-area prioritisation have smaller effects on overall MAE than the other two modules, but they are associated with the control of large errors in complex samples. The knowledge context constrains interpretation of reference objects and submerged parts, while main-flooded-area prioritisation helps the model focus on the principal inundated area and reduces interference from shallow margins or local puddles. The corresponding ablation results and bootstrap confidence intervals are visualised in Figure 9.
Overall, the ablation results indicate that the first-stage candidate-reference-object constraint, explicit reference-object depth mapping, RAG context and main-flooded-area prioritisation jointly reduce unconstrained object selection and direct numerical output.
To complement the ablation experiments, the sensitivity of the final RPD-VLM estimates to proportional mapping-scale uncertainty was evaluated. Table 15 reports the depth-level bounds and the resulting changes in performance under the ±10% scenarios.
For the representative depth levels, the maximum scenario-based deviation increased from 0.1 cm at the 1 cm level to 17.0 cm at the 170 cm level. On the 185 common valid samples, the mean, median, and maximum absolute changes in the final estimates were 3.2, 2.0, and 13.0 cm, respectively. The MAE ranged from 13.7 to 17.4 cm and the RMSE ranged from 19.4 to 25.2 cm, compared with baseline values of 14.9 and 22.0 cm.
The response was asymmetric. Although the −10% scenario reduced MAE and RMSE, Acc@20 also decreased from 0.8000 to 0.7784. The lower mean error under this scenario therefore does not indicate that the mapping should be reduced. Instead, the results show that proportional mapping uncertainty can affect both individual depth estimates and aggregate performance.

4.3.6. Representative Case Analysis

Figure 10 presents three representative cases: a low-error sample, a sample whose error increases after a key module is removed, and a complex deep-water sample with a large error under the complete method. The cases illustrate different error mechanisms rather than the overall statistical distribution.
In case (a), a clearly visible human ankle allows RPD-VLM to form an intermediate interpretation along the reference-object part and depth-class path.
Case (b) illustrates object-selection bias and overestimation that can occur when the first-stage detection constraint is removed.
Case (c) illustrates a difficult scene with multiple vehicles, occlusion, and an unclear water boundary.
All case comparisons use manually assigned reference depths rather than in situ measurements. The examples therefore illustrate relative differences under a common manual benchmark and should not be interpreted as centimetre-level physical validation.

5. Discussion

5.1. Conditions of Applicability

The method is not intended for all images of urban surface water. It targets crowdsourced flood image–text records containing interpretable reference objects and clear relations between those objects and the water surface. RPD-VLM converts an unstructured image into image-based depth evidence containing the reference-object type, submerged part, and knowledge-mapping relation, thereby retaining an intermediate interpretation path. It is not positioned as a replacement for a fully supervised depth-regression model. Instead, it provides a structured process for constructing candidate depth evidence without relying on end-to-end depth-regression training on the current task, particularly when disaster records need to be organised or the basis of an estimate needs to be reviewed.
With respect to image conditions, the method is more applicable to roads, streets, and parking-area entrances containing common reference objects such as pedestrians, vehicles, kerbs, or barriers. Spatial relations between these objects and the water surface are generally more interpretable. In blurred images, images captured from a high angle, scenes with unclear flood boundaries, occluded reference objects, or invisible key parts, the available image cues are insufficient and the output should be treated as a candidate estimate requiring further verification.
The text module has stricter conditions of applicability. Social media text may describe a historical water level, a nearby location, overall event severity, or a subjective impression and therefore cannot be equated directly with the local depth shown in the current image. Text participates in candidate correction only when it is quantifiable, satisfies the image-led directional condition, and remains within the controlled image–text difference range.
The method can support rapid screening of crowdsourced flood information, organisation of candidate depths before manual review, and mapping analyses that require an explicit evidence path. Samples involving deep water, severe occlusion, no stable reference object, or temporal or spatial inconsistency between image and text still require verification using field inspection, surveillance video, water gauges, or sensor observations.
The comparative results are consistent with previous studies showing that supervised image-based models can achieve strong flood-depth estimation performance when task-specific annotations are available [34,35,36], whereas direct numerical generation from large multimodal models remains sensitive to prompting and output constraints [50]. In the present experiments, ViT-B/16 achieved lower error than RPD-VLM, while the structured reference-object pathway yielded lower errors than direct VLM depth generation on the current common valid subset. This positions RPD-VLM as a structured alternative to direct VLM inference rather than as a replacement for supervised regression.

5.2. Error Sources and Uncertainty

The first source of error lies in the construction of image-based depth evidence. RPD-VLM depends on reference-object recognition, submerged-part assessment, and knowledge mapping, and uncertainty at any stage propagates to the final depth estimate. If the model selects an object outside the main flooded area or uses a reference object at a shallow margin to represent the principal inundation, the estimate may depart from the manually assigned reference depth. On the common valid subset, removal of the candidate-reference-object constraint is associated with a larger increase in error, indicating that object selection and object–water interpretation are important error sources.
A second source of uncertainty arises from the reference-object depth mapping. Variations in human body proportions, tyre diameters, vehicle ground clearance, and component heights may cause the physical depth corresponding to the same semantic part to differ from its representative value. The discrete depth classes also cannot fully represent within-class variation.
Under the ±10% scenario, the maximum deviation across the representative depth levels was 17.0 cm, while the MAE on the 185 common valid samples ranged from 13.7 to 17.4 cm. These results represent scenario-based mapping sensitivity rather than a confidence interval for actual object dimensions or a total uncertainty bound. Other uncertainties, including object selection, submerged-part interpretation, imaging conditions, and manually assigned reference depths, are not included.
A third source of error concerns the imaging conditions and spatial structure of crowdsourced scenes. Social media images generally lack standardised camera distance, camera height, and viewpoint information, and may be affected by rain, mist, low illumination, compression artefacts, or motion blur. Reflections from the water surface, unclear road–water boundaries, and vehicle occlusion reduce the interpretability of object–water relations. Among the larger scene categories, road scenes have a comparatively stable error distribution, whereas parking areas, bridges, culverts, and low-lying scenes more often contain multiple vehicles, occlusion, and unclear water boundaries.
Textual uncertainty is another source of error in image–text fusion. Even an explicit numerical value or relative-part expression may not correspond to the image acquisition time or the main flooded area in the image. Text may refer to a nearby intersection, a historical maximum, a local deepest point, or the poster’s subjective assessment of severity. Unconditional text participation is associated with higher error on the current fusion set. Text is therefore used as auxiliary evidence only when the quantifiability, directional, and image–text difference conditions are satisfied.
This finding also complements previous multimodal flood-depth studies that jointly exploit image and textual information, by showing that, in the present crowdsourced dataset, textual participation is beneficial only under restricted conditions and should not be assumed to provide a uniform gain.
In this framework, RPD-VLM-T is used to extract candidate depth evidence from accompanying text, whereas RPD-VLM-Fusion determines whether that evidence should modify the image-based estimate under the predefined quantifiability, directional, and image–text discrepancy conditions. This differs from multimodal approaches that jointly optimise image and text within a unified model, because textual evidence is evaluated separately before being allowed to participate in the final estimate.
The manually assigned reference depths also contain uncertainty. Crowdsourced data usually lack in situ water gauges or synchronous sensor measurements, and the reference depths are mainly determined from visible reference objects, scene context, and annotation rules. The reported errors are therefore more appropriate for comparing methods under a common manual reference benchmark than for representing absolute physical error relative to in situ depth. Future work can combine field observations, camera geometry, instance segmentation, and depth estimation to evaluate uncertainty in reference-object interpretation and scale conversion more directly.

5.3. Application Value and Future Work

The method does not use a single image as a substitute for an in situ water-gauge observation. Instead, it converts fragmented image–text records into comparable and traceable structured depth-evidence units. The image module retains the reference-object type, submerged part, and image-based depth evidence; the text module retains the textual depth evidence and quantifiability status; and the fusion module records whether text participates and the difference between the image and text predictions. These fields provide an evidence path for subsequent filtering, review, and aggregation.
Within a flood event, structured depth evidence with temporal and locational cues can be aggregated by road, community, or reported flood location to help identify repeatedly affected places and compare relative depths across records. Records with clear image evidence and text satisfying the participation conditions may support situational assessment. Records for which text does not participate, the reference object is unstable, or the image–text difference is large should be reviewed manually using the retained intermediate fields.
The spatiotemporal distribution of crowdsourced records is influenced by platform-user density, event attention, camera location, and posting behaviour and does not represent the actual spatial distribution of urban flooding. The location of a record may derive from a platform tag, a textual place name, or manual interpretation, each with uncertain positional accuracy and temporal correspondence. Precise georeferencing should therefore be regarded as a separate uncertainty source from depth estimation. The current experiments evaluate water depth at the record level and do not infer or validate exact geographic coordinates for every social media record. Consequently, the resulting depth estimates should not be interpreted directly as georeferenced flood-depth measurements or as a complete flood-depth map. For downstream spatial aggregation, location cues would require independent geocoding and positional validation before being associated with specific roads, communities, or inundated areas. Images from the same location may also represent different stages of inundation. The results are therefore more appropriately treated as candidate depth evidence and relative depth cues within a specified time window. Future spatial analyses should account for this sampling bias by incorporating population or platform-user density, applying spatial stratification or density-based weighting, and comparing crowdsourced observations with independent monitoring data where available.
RPD-VLM, RPD-VLM-T, and RPD-VLM-Fusion can serve as front-end evidence-construction modules for processing crowdsourced flood information and can be used together with field inspection, fixed cameras, water-level monitoring, rainfall data, or hydrodynamic models. Their role is to provide structured input for subsequent filtering and manual verification, not to replace field measurement or independently produce precise water-depth maps. Future work will focus on validating crowdsourced depth estimates against synchronous field observations, refining the reference-object depth mapping to account for variations in reference-object dimensions, improving georeferencing using textual locations and external geographic information, and evaluating the framework across different flood events and cities.

6. Conclusions

This study developed an image-led framework for urban street-scale flood-depth estimation from crowdsourced image and text data. The framework constructs structured depth evidence separately from images and accompanying text, with image-derived evidence retained as the primary source and text used only as conditional auxiliary evidence. Rather than directly combining the two modalities, the proposed fusion strategy controls textual participation using quantifiability, directional, and image–text discrepancy conditions, thereby limiting the influence of cross-modal inconsistency on the final depth estimate. The proposed method first extracts image-based depth evidence by recognizing reference objects, identifying submerged parts, and establishing object-part depth relationships through knowledge mapping. Compared with direct depth inference from vision–language models, this evidence-based strategy provides a more interpretable approach for transforming visual flood observations into depth estimates. Additionally, textual information from social media is analyzed through rule-based extraction and semantic verification to identify explicit values, range descriptions, and relative water-level expressions. The extracted textual information is incorporated as conditional auxiliary evidence under quantifiability, directional, and image–text discrepancy constraints, considering the potential inconsistency between image and text observations caused by differences in location, time, and flooding conditions.
Experimental results show that the proposed framework can construct structured water-depth evidence from crowdsourced image–text records and reduce the instability of direct VLM depth generation. Conditional text participation provides limited auxiliary correction under the specified constraints but should not be interpreted as a general accuracy gain from multimodal fusion. Because the reference depths are scene-based manual estimates rather than in situ measurements, the resulting depths are more appropriately used as auxiliary evidence for post-disaster situational assessment and subsequent manual verification. Future studies will further evaluate the framework using field measurements, sensor observations, cross-city flood events, and more detailed geometric constraints.

Author Contributions

Conceptualisation, Yaqin Sun and Yifan Zhang; methodology, Yong Zhang; software, Yong Zhang; validation, Yong Zhang and Xun Zhou; formal analysis, Yong Zhang; investigation, Yong Zhang; resources, Hui Yang and Kefei Zhang; data curation, Yong Zhang and Xun Zhou; writing—original draft preparation, Yong Zhang; writing—review and editing, Yaqin Sun, Yifan Zhang, Hui Yang, and Kefei Zhang; visualisation, Yong Zhang; supervision, Yaqin Sun and Yifan Zhang; project administration, Yaqin Sun, Hui Yang, and Kefei Zhang. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by the State Key Laboratory of Resources and Environmental Information System and funded by the Major Program of the National Natural Science Foundation of China, grant number 42394060.

Data Availability Statement

The processed data are maintained in the laboratory’s controlled research environment and are not publicly available due to privacy, copyright, platform terms of use, and potential re-identification risks. Anonymised derived data may be provided upon reasonable request to the corresponding author, subject to review of the proposed research use. Original social media posts and images cannot be redistributed.

Acknowledgments

We thank the editors and reviewers for their helpful comments that improved this paper.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Intergovernmental Panel on Climate Change (IPCC). Climate Change 2021—The Physical Science Basis: Working Group I Contribution to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change, 1st ed.; Cambridge University Press: Cambridge, UK, 2023; ISBN 978-1-009-15789-6. [Google Scholar]
  2. Song, J.; Shao, Z.; Zhan, Z.; Chen, L. State-of-the-Art Techniques for Real-Time Monitoring of Urban Flooding: A Review. Water 2024, 16, 2476. [Google Scholar] [CrossRef] [Scilit]
  3. Li, J.; Cai, R.; Tan, Y.; Zhou, H.; Sadick, A.-M.; Shou, W.; Wang, X. Automatic Detection of Actual Water Depth of Urban Floods from Social Media Images. Measurement 2023, 216, 112891. [Google Scholar] [CrossRef] [Scilit]
  4. Liu, B.; Li, Y.; Ma, M.; Mao, B. A Comprehensive Review of Machine Learning Approaches for Flood Depth Estimation. Int. J. Disaster Risk Sci. 2025, 16, 433–445. [Google Scholar] [CrossRef] [Scilit]
  5. Xiao, S.; Gu, H.; Shen, D.; Niu, Z.; Xiao, J.; Yu, F. A Systematic Review of Social Media-Enabled Flood Disaster Informatics: Method, Technology and Application. Int. J. Appl. Earth Obs. Geoinf. 2025, 144, 104935. [Google Scholar] [CrossRef] [Scilit]
  6. Huang, J.; Kang, J.; Wang, H.; Wang, Z.; Qiu, T. A Novel Approach to Measuring Urban Waterlogging Depth from Images Based on Mask Region-Based Convolutional Neural Network. Sustainability 2020, 12, 2149. [Google Scholar] [CrossRef] [Scilit]
  7. Ouyang, M.; Zeng, B.; Huang, G. A Deep Learning Method for Identifying Waterlogging Depth on Urban Roadways from Surveillance Camera Images. Remote Sens. Appl. Soc. Environ. 2026, 41, 101827. [Google Scholar] [CrossRef] [Scilit]
  8. Bentivoglio, R.; Isufi, E.; Jonkman, S.N.; Taormina, R. Deep Learning Methods for Flood Mapping: A Review of Existing Applications and Future Research Directions. Hydrol. Earth Syst. Sci. 2022, 26, 4345–4378. [Google Scholar] [CrossRef] [Scilit]
  9. Iqbal, U.; Perez, P.; Li, W.; Barthelemy, J. How Computer Vision Can Facilitate Flood Management: A Systematic Review. Int. J. Disaster Risk Reduct. 2021, 53, 102030. [Google Scholar] [CrossRef] [Scilit]
  10. Wang, J.; Zhang, N.; Liu, Y.; Liu, M.; Wang, X.; Li, Z. How Does Multi-Source Social Media Data Serve in Urban Flood Information Collection, Recognition, and Analysis? Water 2026, 18, 405. [Google Scholar] [CrossRef] [Scilit]
  11. Yan, Z.; Guo, X.; Zhao, Z.; Tang, L. Achieving Fine-Grained Urban Flood Perception and Spatio-Temporal Evolution Analysis Based on Social Media. Sustain. Cities Soc. 2024, 101, 105077. [Google Scholar] [CrossRef] [Scilit]
  12. Grassi, L.; Ciranni, M.; Baglietto, P.; Recchiuto, C.T.; Maresca, M.; Sgorbissa, A. Emergency Management through Information Crowdsourcing. Inf. Process. Manag. 2023, 60, 103386. [Google Scholar] [CrossRef] [Scilit]
  13. Wang, C.; Zhang, X.; Wu, J. Disaster Information Mining from a Social Perception Perspective: A Case Study of the “23·7” Extreme Rainfall Event in the Beijing–Tianjin–Hebei Region. Int. J. Disaster Risk Reduct. 2024, 115, 105056. [Google Scholar] [CrossRef] [Scilit]
  14. Acikara, T.; Xia, B.; Yigitcanlar, T.; Hon, C. Contribution of Social Media Analytics to Disaster Response Effectiveness: A Systematic Review of the Literature. Sustainability 2023, 15, 8860. [Google Scholar] [CrossRef] [Scilit]
  15. Guo, Q.; Jiao, S.; Yang, Y.; Yu, Y.; Pan, Y. Assessment of Urban Flood Disaster Responses and Causal Analysis at Different Temporal Scales Based on Social Media Data and Machine Learning Algorithms. Int. J. Disaster Risk Reduct. 2025, 117, 105170. [Google Scholar] [CrossRef] [Scilit]
  16. Feng, Y.; Sester, M. Extraction of Pluvial Flood Relevant Volunteered Geographic Information (VGI) by Deep Learning from User Generated Texts and Photos. ISPRS Int. J. Geo-Inf. 2018, 7, 39. [Google Scholar] [CrossRef] [Scilit]
  17. Kamoji, S.; Kalla, M. Effective Flood Prediction Model Based on Twitter Text and Image Analysis Using BMLP and SDAE-HHNN. Eng. Appl. Artif. Intell. 2023, 123, 106365. [Google Scholar] [CrossRef] [Scilit]
  18. Chow, T.E.; Yang, T.; Yan, Y. The Uncertainties and Potential of Locational-Based Social Media for Floodscape Mapping. Big Earth Data 2026, 1–28. [Google Scholar] [CrossRef] [Scilit]
  19. Yin, W.; Deuser, F.; Liu, Z.; Wei, J.; Luo, X.; Werner, M.; Li, H.; Xue, Y. Triple-Objective Cross-View Geolocalization of Disaster-Related VGI: The Case of Hurricane Ian. Int. J. Geogr. Inf. Sci. 2026, 40, 194–216. [Google Scholar] [CrossRef] [Scilit]
  20. Chen, Y.; Li, X.; Chen, X.; Hu, M. A Framework for Assessing the Credibility of Flood-Inundation Locations Derived from Social Media Using Multi-Source Data. Water Res. 2026, 292, 125224. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Yu, C.; Wang, Z. Cross-Modal Evidential Fusion Network for Social Media Classification. Comput. Speech Lang. 2025, 92, 101784. [Google Scholar] [CrossRef] [Scilit]
  22. Chen, Y.; Hu, M.; Chen, X.; Wang, F.; Liu, B.; Huo, Z. An Approach of Using Social Media Data to Detect the Real Time Spatio-Temporal Variations of Urban Waterlogging. J. Hydrol. 2023, 625, 130128. [Google Scholar] [CrossRef] [Scilit]
  23. Liu, Y.; Li, R.; Wang, S.; Wu, H.; Gui, Z. Deducing Flood Development Process Using Social Media: An Event-Based and Multi-Level Modeling Approach. ISPRS Int. J. Geo-Inf. 2022, 11, 306. [Google Scholar] [CrossRef] [Scilit]
  24. Ji, J.; Tan, Y.; Li, J.; Wang, X. An Automated End-to-End Pipeline for Identifying Fine-Grained Waterlogging Locations from Chinese Social Media. Geomat. Nat. Hazards Risk 2025, 16, 2551272. [Google Scholar] [CrossRef] [Scilit]
  25. Wang, S.; Li, R.; Wu, H.; Li, J.; Shen, Y. Fine-Grained Flood Disaster Information Extraction Incorporating Multiple Semantic Features. Int. J. Digit. Earth 2025, 18, 2448221. [Google Scholar] [CrossRef] [Scilit]
  26. Han, Y.; Liu, J.; Luo, A.; Wang, Y.; Bao, S. Fine-Tuning LLM-Assisted Chinese Disaster Geospatial Intelligence Extraction and Case Studies. ISPRS Int. J. Geo-Inf. 2025, 14, 79. [Google Scholar] [CrossRef] [Scilit]
  27. Yang, Y.; Wang, Q.; Li, W.; Wang, S.; Fang, S.; Sun, K.; Wang, S.; Wang, H.; Dai, X.; Li, Z.; et al. Automated Extraction of Spatiotemporal Disaster Knowledge for Urban Floods: A Multimodal Framework Based on LLMs and Agent. Int. J. Digit. Earth 2026, 19, 2640706. [Google Scholar] [CrossRef] [Scilit]
  28. Wang, M.; Zhang, J.; Cao, Y.; Li, S.; Chen, M. A Study on a Spatiotemporal Entity-Based Event Data Model. ISPRS Int. J. Geo-Inf. 2024, 13, 360. [Google Scholar] [CrossRef] [Scilit]
  29. Feng, Y.; Huang, X.; Sester, M. Extraction and Analysis of Natural Disaster-Related VGI from Social Media: Review, Opportunities and Challenges. Int. J. Geogr. Inf. Sci. 2022, 36, 1275–1316. [Google Scholar] [CrossRef] [Scilit]
  30. Zhang, X.; Zhang, X.; Zhang, Y.; Liu, Y.; Zhou, R.; Raxidin, A.; Li, M. A Multidimensional Study of the 2023 Beijing Extreme Rainfall: Theme, Location, and Sentiment Based on Social Media Data. ISPRS Int. J. Geo-Inf. 2025, 14, 136. [Google Scholar] [CrossRef] [Scilit]
  31. Li, B.; Zhang, Z.; Wang, X.; Ren, S. Assessing Stage-Dependent Impacts of Environmental Factors on Public-Perceived Flood Risk: Insights from a GeoXAI Framework. Environ. Impact Assess. Rev. 2026, 119, 108402. [Google Scholar] [CrossRef] [Scilit]
  32. Hao, X.; Lyu, H.; Wang, Z.; Fu, S.; Zhang, C. Estimating the Spatial-Temporal Distribution of Urban Street Ponding Levels from Surveillance Videos Based on Computer Vision. Water Resour. Manag. 2022, 36, 1799–1812. [Google Scholar] [CrossRef] [Scilit]
  33. Li, H.; Deuser, F.; Yin, W.; Luo, X.; Walther, P.; Mai, G.; Huang, W.; Werner, M. Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN. ISPRS J. Photogramm. Remote Sens. 2025, 220, 841–854. [Google Scholar] [CrossRef] [Scilit]
  34. Du, W.; Qian, M.; He, S.; Xu, L.; Zhang, X.; Huang, M.; Chen, N. An Improved ResNet Method for Urban Flooding Water Depth Estimation from Social Media Images. Measurement 2025, 242, 116114. [Google Scholar] [CrossRef] [Scilit]
  35. Wu, L.; Liu, Y.; Zhang, J.; Zhang, B.; Wang, Z.; Tong, J.; Li, M.; Zhang, A. Identification of Flood Depth Levels in Urban Waterlogging Disaster Caused by Rainstorm Using a CBAM-Improved ResNet50. Expert Syst. Appl. 2024, 255, 124382. [Google Scholar] [CrossRef] [Scilit]
  36. Du, W.; Liu, X.; Qian, M.; Xu, L.; Zhang, X.; Chen, N. Causal Attention-Based Water Depth Estimation for Complex Flooding Scenes in Social Media Images. Geomat. Nat. Hazards Risk 2026, 17, 2634964. [Google Scholar] [CrossRef] [Scilit]
  37. Noto, S.; Tauro, F.; Petroselli, A.; Apollonio, C.; Botter, G.; Grimaldi, S. Low-Cost Stage-Camera System for Continuous Water-Level Monitoring in Ephemeral Streams. Hydrol. Sci. J. 2022, 67, 1439–1448. [Google Scholar] [CrossRef] [Scilit]
  38. Qin, J.; Shen, P. Refraction-Based Waterlogging Depth Measurement Using Solely Traffic Cameras for Transparent Flood Monitoring. J. Hydrol. 2025, 655, 132917. [Google Scholar] [CrossRef] [Scilit]
  39. Zamanizadeh, M.; Cetin, M.; Shahabi, A.; Tahvildari, N. Depth Estimation in Urban Flooding Using Surveillance Cameras and High-Resolution LiDAR Data. Environ. Model. Softw. 2025, 192, 106572. [Google Scholar] [CrossRef] [Scilit]
  40. Moya, L.; Mas, E.; Koshimura, S. Sparse Representation-Based Inundation Depth Estimation Using SAR Data and Digital Elevation Model. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 9062–9072. [Google Scholar] [CrossRef] [Scilit]
  41. Luan, G.; Hou, J.; Wang, T.; Zhou, Q.; Xu, L.; Sun, J.; Wang, C. Method for Analyzing Urban Waterlogging Mechanisms Based on a 1D-2D Water Environment Dynamic Bidirectional Coupling Model. J. Environ. Manag. 2024, 360, 121024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Quagliarini, E.; Romano, G.; Bernardini, G. Investigating Pedestrian Behavioral Patterns under Different Floodwater Conditions: A Video Analysis on Real Flood Evacuations. Saf. Sci. 2023, 161, 106083. [Google Scholar] [CrossRef] [Scilit]
  43. Wan, J.; Qin, Y.; Shen, Y.; Yang, T.; Yan, X.; Zhang, S.; Yang, G.; Xue, F.; Wang, Q.J. Automatic Detection of Urban Flood Level with YOLOv8 Using Flooded Vehicle Dataset. J. Hydrol. 2024, 639, 131625. [Google Scholar] [CrossRef] [Scilit]
  44. Wan, J.; Shen, Y.; Xue, F.; Yan, X.; Qin, Y.; Yang, T.; Yang, G.; Wang, Q.J. DSC-YOLOv8n: An Advanced Automatic Detection Algorithm for Urban Flood Levels. J. Hydrol. 2024, 643, 132028. [Google Scholar] [CrossRef] [Scilit]
  45. Qiu, Y.; Zhou, X.; Wan, J.; Yang, T.; Zhang, L.; Zhong, Y.; Shen, L.; Ji, X. Automated Urban Flood Level Detection Based on Flooded Bus Dataset Using YOLOv8. Nat. Hazards Earth Syst. Sci. 2025, 25, 3525–3544. [Google Scholar] [CrossRef] [Scilit]
  46. M, D.; S, C.; C.M., B. Robust Human Detection System in Flood Related Images with Data Augmentation. Multimed. Tools Appl. 2023, 82, 10661–10679. [Google Scholar] [CrossRef] [Scilit]
  47. Gilroy, S.; Glavin, M.; Jones, E.; Mullins, D. An Objective Method for Pedestrian Occlusion Level Classification. Pattern Recognit. Lett. 2022, 164, 96–103. [Google Scholar] [CrossRef] [Scilit]
  48. Akinboyewa, T.; Ning, H.; Lessani, M.N.; Li, Z. Automated Floodwater Depth Estimation Using Large Multimodal Model for Rapid Flood Mapping. Comput. Urban. Sci. 2024, 4, 12. [Google Scholar] [CrossRef] [Scilit]
  49. Lin, L.; Zeng, Z.; Tang, C.; Xie, Y.; Liang, Q. Robust and Fast Sensing of Urban Flood Depth with Social Media Images Using Pre-Trained Large Models and Simple Edge Training. Hydrology 2025, 12, 307. [Google Scholar] [CrossRef] [Scilit]
  50. Lyu, H.; Zhou, S.; Wang, Z.; Fu, G.; Zhang, C. Assessing Large Multimodal Models for Urban Floodwater Depth Estimation. Water Resour. Res. 2025, 61, e2024WR039494. [Google Scholar] [CrossRef] [Scilit]
  51. Chu, T.; Chen, Y.; Zhu, R.; Zeng, F. Estimating Urban Flooding Depth by Integrating Multimodal Image-Text Data: A Segment-Level Direct Preference Optimization-Based Multimodal Large Language Model. ISPRS J. Photogramm. Remote Sens. 2025, 230, 895–917. [Google Scholar] [CrossRef] [Scilit]
  52. Zhu, H.; Meng, J.; Yao, J.; Xu, N. Feasibility of Emergency Flood Traffic Road Damage Assessment by Integrating Remote Sensing Images and Social Media Information. ISPRS Int. J. Geo-Inf. 2024, 13, 369. [Google Scholar] [CrossRef] [Scilit]
  53. Chen, Y.; Zhang, L.; Chen, X. A Framework for Using Event Evolutionary Graphs to Rapidly Assess the Vulnerability of Urban Flood Cascade Compound Disaster Event Networks. J. Hydrol. 2024, 642, 131783. [Google Scholar] [CrossRef] [Scilit]
  54. Lu, H.; Zhang, S.; Gao, Y.; Jin, H.; Zhao, P.; Gao, Y.; Li, Y.; Wang, W.; Zhang, Y. Using Social Media Data to Construct and Analyze Knowledge Graph for “7.20” Henan Rainstorm Flood Event. Int. J. Disaster Risk Reduct. 2025, 116, 105129. [Google Scholar] [CrossRef] [Scilit]
  55. Weibo. Available online: https://weibo.com/ (accessed on 31 August 2026).
  56. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Las Vegas, NV, USA, 2016; pp. 770–778. [Google Scholar]
  57. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
  58. Bai, S.; Cai, Y.; Chen, R.; Chen, K.; Chen, X.; Cheng, Z.; Deng, L.; Ding, W.; Gao, C.; Ge, C.; et al. Qwen3-VL Technical Report. arXiv 2025, arXiv:2511.21631. [Google Scholar]
Figure 1. Overall methodology and experimental evaluation workflow.
Figure 1. Overall methodology and experimental evaluation workflow.
Ijgi 15 00430 g001
Figure 2. Mapping between reference-object parts and discrete depth classes (Note: The figure illustrates the correspondence between the depth classes and common urban reference objects, including pedestrians, non-motorised vehicles, and motor vehicles).
Figure 2. Mapping between reference-object parts and discrete depth classes (Note: The figure illustrates the correspondence between the depth classes and common urban reference objects, including pedestrians, non-motorised vehicles, and motor vehicles).
Ijgi 15 00430 g002
Figure 3. Workflow for extracting and assessing quantifiable depth evidence from social media text.
Figure 3. Workflow for extracting and assessing quantifiable depth evidence from social media text.
Ijgi 15 00430 g003
Figure 4. Post hoc diagnostic of image–text discrepancy and change in fusion error. Note: A negative value indicates lower error after fusion. Direction fallback and Δd > 60 cm fallback indicate that text does not participate and the estimate returns to image-only.
Figure 4. Post hoc diagnostic of image–text discrepancy and change in fusion error. Note: A negative value indicates lower error after fusion. Direction fallback and Δd > 60 cm fallback indicate that text does not participate and the estimate returns to image-only.
Ijgi 15 00430 g004
Figure 5. Image–text fusion cases under different text-participation weights. Note: (a) When Δd ≤ 27 cm, text participates with the maximum weight. (b) When 27 < Δd ≤ 60 cm, the text weight decreases continuously with the image–text difference. (c) When Δd > 60 cm, the text weight is set to 0 and the final estimate falls back to the image depth. D r e f denotes an illustrative manually assigned reference depth.
Figure 5. Image–text fusion cases under different text-participation weights. Note: (a) When Δd ≤ 27 cm, text participates with the maximum weight. (b) When 27 < Δd ≤ 60 cm, the text weight decreases continuously with the image–text difference. (c) When Δd > 60 cm, the text weight is set to 0 and the final estimate falls back to the image depth. D r e f denotes an illustrative manually assigned reference depth.
Ijgi 15 00430 g005
Figure 6. Predicted depth versus manually assigned reference depth for the compared methods.
Figure 6. Predicted depth versus manually assigned reference depth for the compared methods.
Ijgi 15 00430 g006
Figure 7. Performance at different training proportions: (a) MAE; (b) Acc@20. Test sets: supervised models, N = 193; RPD-VLM reference, N = 185.
Figure 7. Performance at different training proportions: (a) MAE; (b) Acc@20. Test sets: supervised models, N = 193; RPD-VLM reference, N = 185.
Ijgi 15 00430 g007
Figure 8. Distribution of absolute errors for RPD-VLM across scene types. Note: The orange line within each box denotes the median, and the dashed horizontal lines indicate the 10 cm and 20 cm error thresholds.
Figure 8. Distribution of absolute errors for RPD-VLM across scene types. Note: The orange line within each box denotes the median, and the dashed horizontal lines indicate the 10 cm and 20 cm error thresholds.
Ijgi 15 00430 g008
Figure 9. Comparison of ablation results for image-based depth quantification: (a) MAE; (b) RMSE; (c) Acc@20. Error bars denote bootstrap 95% confidence intervals.
Figure 9. Comparison of ablation results for image-based depth quantification: (a) MAE; (b) RMSE; (c) Acc@20. Error bars denote bootstrap 95% confidence intervals.
Ijgi 15 00430 g009
Figure 10. Comparison of representative cases: (a) a successful sample with a clear reference object; (b) an error case after removal of a key module; (c) a difficult complex deep-water scene.
Figure 10. Comparison of representative cases: (a) a successful sample with a clear reference object; (b) an error case after removal of a key module; (c) a difficult complex deep-water scene.
Ijgi 15 00430 g010
Table 1. Experimental sample sets and data flow.
Table 1. Experimental sample sets and data flow.
Sample SetNInclusion Criterion
Annotated Image Dataset 1283Images Retained After Automatic Screening, Manual Secondary Review, Duplicate Removal, and Manual Annotation
Cleaned text samples for text extraction239Accompanying texts retained after removal of depth-irrelevant and duplicate records, with annotatable text reference depth
Unified image test set193Derived from the common train/validation/test split
Common valid subset for the main comparison185Valid predictions available for all compared methods
Valid RPD-VLM samples190Samples from the fixed 193-image test set for which RPD-VLM produced a valid depth estimate
Valid image–text fusion samples202Image prediction, text prediction, and image reference depth all available
Common valid subset for ablation167Valid outputs available for the full method and all ablation settings
Table 2. Representative results for image–text fusion and image-based depth estimation.
Table 2. Representative results for image–text fusion and image-based depth estimation.
ModuleRoleMethod/StrategyValid NMAERMSEAcc@20
Multimodal fusionImage-only baselineImage-only/RPD-VLM20213.218.60.8119
Text-only baselineText-only/RPD-VLM-T20235.144.50.3911
Simple fusion baselineDirect average20221.926.70.5743
Fixed-weight baselineFixed weight20215.819.60.7277
Fusion with image–text RPD-VLM-Fusion20212.517.20.8168
Image-depth quantificationSupervised image baselineResNet-5018512.918.80.8162
Supervised image baselineViT-B/1618510.817.70.8324
Direct VLM baselineVLM-only18527.333.10.4595
Direct VLM baselineVLM + CoT18524.228.70.5405
Image-based methodRPD-VLM18514.922.00.8000
Note: Valid N is the number of samples successfully extracted and included in error calculation. For the text-extraction task, Valid N denotes the number of successful deterministic depth extractions from the common input pool of 239 cleaned text samples. For the image–text fusion and image-depth experiments, N is the sample size of the corresponding evaluation set. Because the tasks differ in sample population, evaluation objective, and metric interpretation, results are compared only within each task. Complete metrics are reported in the corresponding sections.
Table 3. Overall extraction results for quantifiable text samples (error unit: cm).
Table 3. Overall extraction results for quantifiable text samples (error unit: cm).
MethodTotal NExtracted NCoverageMAERMSEMedAEAcc@10Acc@20
Rule-only2391430.59833.46.90.00.96500.9860
RPD-VLM-T2392210.92476.510.80.00.81900.9186
Table 4. Depth-extraction errors on the common successfully extracted samples.
Table 4. Depth-extraction errors on the common successfully extracted samples.
MethodCommon NMAE/cmRMSE/cmMedAE/cmAcc@10Acc@20
Rule-only1433.46.90.00.96500.9860
RPD-VLM-T1435.410.00.00.88110.9021
Table 5. Image–text depth estimates under different text-participation strategies.
Table 5. Image–text depth estimates under different text-participation strategies.
StrategyTotal
N
Participating NChanged NMAERMSEMedAEAcc@10Acc@20
Image-only2020013.218.610.00.61880.8119
Text-only20220217735.144.530.00.23270.3911
Direct average20220217721.926.720.00.28710.5743
Fixed weight20220217715.819.614.00.39600.7277
RPD-VLM-Fusion202595912.517.210.00.64360.8168
Table 6. Textual correction under directional and image–text difference conditions.
Table 6. Textual correction under directional and image–text difference conditions.
DirectionNBetter/Equal/WorseBetter RateImage Text Direct-Avg. Continuous
Dt < Di7030/3/370.428619.330.020.117.3
Dt ≥ Di1327/25/1000.053010.037.822.810.0
Note: Image, Text, Direct-avg, and Continuous denote image MAE, text MAE, direct-average MAE, and continuous-fusion MAE, respectively. All MAE values are in centimetres.
Table 7. Valid-prediction coverage on the unified test set.
Table 7. Valid-prediction coverage on the unified test set.
MethodTest NValid NInvalid NCoverage
ResNet-501931930100.00%
ViT-B/161931930100.00%
VLM-only193189497.93%
VLM + CoT193191298.96%
RPD-VLM193190398.45%
Common valid subset193185895.85%
Note: Valid N is the number of samples for which the method produces a valid continuous depth prediction, and Invalid N is the number for which no valid continuous prediction can be formed. Common valid subset is the set of samples with a valid prediction from every method.
Table 9. Paired bootstrap tests comparing RPD-VLM with the baseline methods.
Table 9. Paired bootstrap tests comparing RPD-VLM with the baseline methods.
ComparisonΔMAE (cm)95% CIΔRMSE (cm)95% CIΔAcc@2095% CI
ResNet-50+2.1[−0.6, 4.7]+3.2[−1.0, 7.5]−0.0162[−0.0865, 0.0541]
ViT-B/16+4.1[1.4, 6.9]+4.3[−0.9, 9.4]−0.0324[−0.1027, 0.0378]
VLM-only−12.4[−16.0, −8.8]−11.1[−15.9, −6.5]+0.3405[0.2432, 0.4378]
VLM + CoT−9.3[−12.5, −6.1]−6.7[−10.5, −2.9]+0.2595[0.1622, 0.3568]
Note: Δ is calculated as RPD-VLM minus the corresponding baseline. For MAE and RMSE, a negative value indicates lower error for RPD-VLM; for Acc@20, a positive value indicates higher performance. CI denotes confidence interval. Statistical comparison was performed using the paired bootstrap procedure described in Section 3.6.
Table 10. Comparison of image-based methods across manually assigned depth ranges.
Table 10. Comparison of image-based methods across manually assigned depth ranges.
MethodDepth RangeNMAE (cm)Acc@20
ResNet-500–20 cm1318.10.947
ResNet-5020–50 cm2415.80.708
ResNet-5050–100 cm3223.40.469
ViT-B/160–20 cm1316.90.916
ViT-B/1620–50 cm2413.50.667
ViT-B/1650–100 cm3217.70.688
RPD-VLM0–20 cm1319.00.931
RPD-VLM20–50 cm2424.10.667
RPD-VLM50–100 cm3230.30.406
Table 11. RPD-VLM error distributions across scene types.
Table 11. RPD-VLM error distributions across scene types.
Scene TypeNMAE (cm)RMSE (cm)MedAE (cm)Acc@10Acc@20Out@30
Road10312.218.510.00.7570.8930.078
Parking area2423.530.520.00.3750.5420.333
Residential area3615.723.810.00.5830.7500.167
Bridge, culvert, or low-lying area1716.523.810.00.5290.7060.235
Other/unclear108.514.63.50.8000.9000.100
Table 12. Definition of the ablation settings for image-based depth quantification.
Table 12. Definition of the ablation settings for image-based depth quantification.
Ablation SettingRemoved ComponentReplacement Mechanism
w/o RAG contextExternal knowledge contextReference-object part assessment using only the base prompt
w/o Main flooded-zone selectionMain-flooded-area prioritisation ruleImage-based depth evidence from the model’s default candidate object
w/o Stage-1 ConstraintFirst-stage candidate-reference-object constraintReference-object selection from a less constrained candidate set in the second stage
w/o Knowledge mappingExplicit reference-object depth mappingDirect VLM depth output followed by the same post-processing
Table 13. Ablation results for image-based depth quantification on the common valid subset.
Table 13. Ablation results for image-based depth quantification on the common valid subset.
MethodNMAE (cm)RMSE (cm)MedAE (cm)Acc@10Acc@20Out@30Out@50
RPD-VLM16715.922.810.00.61680.78440.15570.0240
w/o RAG context16721.531.410.00.52690.70060.22750.1018
w/o Main flooded-zone selection16719.029.310.00.61080.77250.19160.1138
w/o Stage-1 constraint16749.463.150.00.29940.40720.55090.4790
w/o Knowledge mapping16727.331.125.00.14970.38920.40720.0359
Note: The table is calculated only on the common subset for which the complete RPD-VLM and all ablation settings produce valid predictions. It is used to compare relative changes after individual modules are removed.
Table 14. Paired bootstrap tests comparing the complete RPD-VLM with the ablation settings.
Table 14. Paired bootstrap tests comparing the complete RPD-VLM with the ablation settings.
ComparisonΔMAE (cm)95% CIΔRMSE (cm)95% CIΔAcc@2095% CI
w/o RAG context−5.6[−9.6, −1.6]−8.5[−14.3, −2.6]+0.0838[0.0000, 0.1677]
w/o Main flooded-zone selection−3.1[−4.2, −2.1]−6.4[−8.0, −4.8]+0.0120[0.0000, 0.0299]
w/o Stage-1 constraint−33.5[−39.4, −27.8]−40.2[−45.9, −34.5]+0.3772[0.2934, 0.4611]
w/o Knowledge mapping−11.4[−15.1, −7.7]−8.2[−12.6, −3.9]+0.3952[0.2874, 0.4970]
Note: Δ is calculated as the complete RPD-VLM minus the corresponding ablation setting. For MAE and RMSE, a negative value indicates lower error for the complete RPD-VLM; for Acc@20, a positive value indicates higher performance. The 95% CIs are obtained from 10,000 paired bootstrap resamples. A difference is considered statistically stable when its 95% CI does not include 0.
Table 15. (a) Sensitivity of the final RPD-VLM estimates to ±10% proportional perturbations. Bounds of the representative depth levels defined in Figure 2. (b) Performance propagation on the 185 common valid samples.
Table 15. (a) Sensitivity of the final RPD-VLM estimates to ±10% proportional perturbations. Bounds of the representative depth levels defined in Figure 2. (b) Performance propagation on the 185 common valid samples.
(a)
Depth
Level (cm)
Range Under ±10%
Perturbation (cm)
Maximum
Deviation (cm)
00–00
10.9–1.10.1
109–111
2018–222
4036–444
6054–666
8072–888
10090–11010
130117–14313
150135–16515
170153–18717
(b)
PerturbationNMAE (cm)RMSE (cm)MedAE (cm)Acc@10Acc@20
−10%18513.719.49.00.60000.7784
0% (Baseline)18514.922.010.00.64860.8000
+10%18517.425.211.00.42700.7189
Note: Panel (a) reports the ±10% ranges for the depth levels defined in Figure 2. Panel (b) reports the performance obtained after applying the same perturbations to the final predictions of the 185 samples used in Table 8. The manually assigned reference depths were kept unchanged.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, Y.; Sun, Y.; Zhou, X.; Zhang, Y.; Yang, H.; Zhang, K. Urban Flood-Depth Estimation from Crowdsourced Image–Text Data Using Reference-Object Reasoning and Conditional Fusion. ISPRS Int. J. Geo-Inf. 2026, 15, 430. https://doi.org/10.3390/ijgi15090430

AMA Style

Zhang Y, Sun Y, Zhou X, Zhang Y, Yang H, Zhang K. Urban Flood-Depth Estimation from Crowdsourced Image–Text Data Using Reference-Object Reasoning and Conditional Fusion. ISPRS International Journal of Geo-Information. 2026; 15(9):430. https://doi.org/10.3390/ijgi15090430

Chicago/Turabian Style

Zhang, Yong, Yaqin Sun, Xun Zhou, Yifan Zhang, Hui Yang, and Kefei Zhang. 2026. "Urban Flood-Depth Estimation from Crowdsourced Image–Text Data Using Reference-Object Reasoning and Conditional Fusion" ISPRS International Journal of Geo-Information 15, no. 9: 430. https://doi.org/10.3390/ijgi15090430

APA Style

Zhang, Y., Sun, Y., Zhou, X., Zhang, Y., Yang, H., & Zhang, K. (2026). Urban Flood-Depth Estimation from Crowdsourced Image–Text Data Using Reference-Object Reasoning and Conditional Fusion. ISPRS International Journal of Geo-Information, 15(9), 430. https://doi.org/10.3390/ijgi15090430

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop