Skip to Content
Applied SciencesApplied Sciences
  • Article
  • Open Access

8 October 2026

26 Pages

MSGLF-Refine: Text-Guided Global–Local Fusion with Adaptive Gating for Fine-Grained Crisis Information Mining

,
,
,
and
1
School of Computer Science and Information Security, University of Emergency Management, Beijing 101601, China
2
Hebei Province University Smart Emergency Application Technology Research and Development Center, Beijing 101601, China
3
Planning and Development Department, Big Data Center of Ministry of Emergency Management, Beijing 100054, China
*
Authors to whom correspondence should be addressed.
Appl. Sci.2026, 16(19), 9951;https://doi.org/10.3390/app16199951 
(registering DOI)
This article belongs to the Section Computing and Artificial Intelligence

Abstract

Multimodal image–text posts on social media provide valuable evidence for disaster response, but fine-grained crisis classification remains difficult when modality reliability varies, categories overlap, and global representations obscure localized visual cues. This study proposes MSGLF-Refine, a text-guided global–local refinement framework based on OpenCLIP ViT-L/14-336. Global image–text fusion establishes a baseline representation; text-guided attention then retrieves relevant visual patches, and gated residual integration controls their contribution to the final prediction. On the official five-class CrisisMMD Task 02 split, a three-seed equal-logit ensemble achieves 92.44% accuracy, 92.41% weighted F1-score, and 90.28% macro F1-score. Published results provide benchmark context, while internally matched ablations show that local refinement maintains comparable average single-model performance to the global baseline and yields a larger observed ensemble gain. Leave-one-event-out experiments across seven disasters improve weighted F1-score over direct global concatenation for every held-out event. Text-length analysis examines performance across posts of different lengths. Counterfactual text interventions and visual evidence analysis offer complementary checks on the model’s evidence use. The results support global–local refinement on the historical CrisisMMD benchmark, while rare-class performance and generalization to current social media require further validation.

1. Introduction

1.1. Background and Motivation

During the first hours after a disaster, social media posts can reveal trapped individuals, road disruptions, infrastructure damage, volunteer activity, and requests for supplies. Studies of crisis communication and multimodal disaster assessment show that these streams can support situational awareness and operational triage [1,2,3,4]. However, actionable reports remain mixed with reposted news, emotional reactions, and irrelevant content. Emergency agencies must therefore do more than detect disaster-related posts: they must identify which messages require verification, routing, or intervention.
Fine-grained classification is difficult because the evidence in the two modalities is often incomplete or asymmetric. Text may express urgency without sufficient visual grounding, whereas an image may show damage without indicating whether the post concerns rescue, donation, infrastructure repair, or general awareness. Effective triage consequently requires global image–text alignment, localization of the visual evidence referenced by the text, and control over how strongly that local evidence changes the final decision.
The decisive evidence is also sparse and context dependent. Small objects or regions may distinguish semantically adjacent classes, and the reliability of each modality varies across posts. Class imbalance and event-specific vocabulary further complicate evaluation: a model can perform well on dominant classes or familiar events while missing rare, operationally important cases. We therefore report accuracy, macro F1-score, and weighted F1-score, evaluate multiple random seeds, and test transfer to held-out disaster events.
Interpretability is equally important in emergency-management use. A label is more useful when an analyst can relate it to specific evidence, such as an affected person, a damaged road, or relief supplies, and to language expressing an actionable condition or response activity. We therefore complement predictive metrics with counterfactual interventions on text and qualitative inspection of text-guided visual attention.
Across crisis corpora, four recurring properties motivate the proposed design: visual ambiguity among related categories; semantic conflict between image and text; contextual noise from salient disaster terms unsupported by the image; and non-humanitarian content that mentions an event without conveying actionable information [5,6,7,8].

1.2. Related Work

Research on social media disaster analysis has progressed from keyword retrieval and unimodal recognition to multimodal reasoning over visual context and linguistic intent. This subsection groups the developments most relevant to fine-grained crisis classification and distinguishes our evidence-retrieval objective from earlier global-fusion and decision-refinement approaches.
Early crisis-processing systems used keywords, bag-of-words features, topic models, and manually designed taxonomies [9,10]. They supported rapid filtering but were sensitive to vocabulary drift and event-specific expressions. Visual backbones such as VGG and DenseNet improved scene recognition [11,12], yet visually similar categories and weak image–text alignment remained difficult to resolve.
Transformer encoders adapted to general and social media language improved the contextual representation of short, noisy posts [13,14,15]. CrisisBERT further captures crisis-specific language [16], while HumAID and CrisisBench provide task-oriented corpora and benchmarks [7,8]. Text-only systems, however, cannot determine whether a linguistic claim is supported by the paired image.
CrisisMMD established a multi-event image–text benchmark for humanitarian information processing [1]. Subsequent studies examined semi-supervised learning, contrastive learning, and multimodal fusion on crisis posts [17,18,19,20]. Cross-modal categorization and graph-based context modeling further showed the value of joint evidence [21,22], consistent with broader systematic evidence on multimodal disaster-response systems [23]. Static or globally pooled fusion nevertheless risks suppressing small, category-defining regions.
Large-scale vision–language pretraining provides a stronger basis for alignment. ViT, CLIP, and ALIGN established scalable visual and image–text representation learning [24,25,26]. Region–word pretraining was advanced by VisualBERT and UNITER [27,28]. Later models reduced architectural overhead or broadened multimodal pretraining objectives [29,30,31,32]. These developments motivate the OpenCLIP backbone used here.
Attention-based interaction and co-attention provide general mechanisms for relating visual and linguistic evidence [33,34], while gated multimodal units regulate modality contributions [35]. Crisis-specific systems extend these ideas through bidirectional interaction, social graphs, or knowledge enhancement [22,36,37]. However, many approaches pool the image before fusion. Fine-grained categories may instead depend on a small road defect, an affected person, or a relief object that global pooling obscures.
Graph and social-context methods can exploit relations beyond an individual post, but they require metadata that may be unavailable or platform specific. MSGLF-Refine uses only the paired image and text, allowing its gains to be attributed to multimodal representation and evidence refinement.
Interpretability research distinguishes visual localization from behavioral faithfulness. Grad-CAM provides spatial evidence for image predictions [38], while attention-faithfulness research cautions that a plausible map is not necessarily a faithful explanation [39]. We therefore combine visual localization with token-removal and token-retention tests that measure whether ranked evidence changes or preserves the prediction under controlled intervention.
Interpretable multimodal evidence can also help analysts review individual predictions and apply operational triage rules consistently. When visual attention and influential textual cues converge on the same crisis-related entity or action, analysts can more readily assess whether a predicted category is supported by the underlying post. Conversely, disagreement between modalities may indicate ambiguity, misinformation, or insufficient evidence and should trigger additional human review. This evidence-centered perspective is particularly important for rare but high-priority categories, where classification errors may have disproportionate operational consequences. The framework improves aggregate performance and returns traceable multimodal cues that analysts can inspect when assessing model predictions for emergency decisions made under severe operational time pressure.

1.3. Research Gap and Contributions

CLIP-BCA-Gated [36] establishes bidirectional global image–text interaction with modality-level reliability gating. MKAN-Refine [40] incorporates text-conditioned patch attention and KAN-based nonlinear refinement, whereas MSGLF-Refine anchors its local-evidence pathway to a BCA-based global prediction and regulates the correction through a zero-initialized gated residual. MSGLF-Refine addresses a different remaining source of error: its text representation queries 576 visual patch tokens, and channel-wise gating with a zero-initialized residual scale controls how retrieved local evidence modifies the global BCA prediction. This extends the shared research line from global alignment and nonlinear decision refinement to explicit local-evidence retrieval and controlled global–local correction. The trade-off is additional attention and patch-feature processing, which is quantified against the internally matched BCA baseline in Section 3.5 a direct efficiency ranking against the published systems would require profiling them under the same hardware and preprocessing protocol.
The literature review identifies four limitations:
  • Global fusion struggles to distinguish humanitarian categories with similar visual appearances and overlapping textual semantics.
  • External knowledge and graph structures enrich contextual reasoning but cannot replace evidence refinement within the multimodal decision pathway.
  • Conventional modality gating adjusts image and text contributions without identifying which image regions support a particular post.
  • Unrestricted enhancement may disturb a reliable cross-modal baseline, requiring incremental correction that preserves the original prediction when additional evidence is uninformative.
MSGLF-Refine starts with representation-aligned baseline, retrieves text-conditioned patches, and applies gated residual correction when supplementary evidence improves the prediction. This design preserves broadly reliable cross-modal semantics while selectively incorporating localized visual cues that are consistent with the accompanying textual description. The adaptive gate further constrains noisy or irrelevant evidence from dominating the final representation, thereby supporting more robust discrimination among semantically adjacent humanitarian categories. Table 1 positions the method relative to representative developments in crisis classification and highlights its joint emphasis on local-evidence retrieval, modality-aware integration, and controlled decision correction.
Table 1. Evolution and technical differences among representative multimodal crisis-classification frameworks.
CLIP provides a shared semantic space through large-scale image–text contrastive learning [24]. CLIP-BCA-Gated [36] uses global bidirectional interaction and modality gating, whereas MKAN-Refine [40] uses KAN-based text-conditioned scoring of visual patches, symmetric cross-modal refinement, and reliability-aware fusion. MSGLF-Refine also draws on patch-level evidence, but uses multi-head text-guided retrieval from 576 patches and a zero-initialized gated residual to correct a BCA-based global prediction. Its distinct contribution is the controlled global–local correction pathway, not patch retrieval alone; no direct computational advantage over the two published systems is inferred without matched profiling.
To address this problem, we propose MSGLF-Refine, a multimodal framework for five-class crisis image–text classification. OpenCLIP and bidirectional cross-attention first establish a global baseline. The global text representation then queries visual patch tokens to retrieve category-relevant local evidence. Finally, channel-wise gating and a zero-initialized residual scale introduce only the correction supported by the combined global, local, and textual context.
The main contributions of this study are summarized as follows:
  • We develop an end-to-end framework that connects bidirectional global image–text interaction, text-guided patch retrieval, and global–local refinement for fine-grained crisis classification.
  • We introduce a text-guided local attention module and a channel-wise refinement gate that select task-relevant visual patches and regulate the contribution of global visual, local visual, and textual evidence.
  • We use a zero-initialized learnable residual scale so that training begins from the stable bidirectional cross-attention baseline and the refinement branch learns only the incremental correction required by each sample.
  • We evaluate the framework on the official CrisisMMD Task 02 split through multi-seed testing, controlled ablations, leave-one-event-out transfer, efficiency measurements, and counterfactual evidence analysis. The three-seed ensemble achieves 92.44% accuracy, 92.41% weighted F1-score, and 90.28% macro F1-score.
Figure 1 illustrates the local decision landscapes of the BCA baseline and MSGLF-Refine using four challenging examples from the official test set. In all four cases, the baseline predictions remain on the incorrect side of the decision boundary, whereas the complete model shifts them across the boundary toward the correct class. Together, these cases illustrate how targeted local evidence resolves ambiguity that remains after global multimodal fusion.
Figure 1. Local decision topology for representative challenging cases, illustrating how global–local refinement changes the class boundary relative to the BCA baseline.

2. Materials and Methods

MSGLF-Refine follows the forward-computation sequence described below. The model first builds a global cross-modal baseline, then retrieves text-conditioned visual evidence from patch tokens, and finally adds the resulting context through gated residual correction. The sequence keeps the stable semantic reference clearly separate from supplementary evidence used for refinement.

2.1. Task Formulation

Given a sample x = (I, T), where I denotes a disaster-related image and T denotes its accompanying social media text, the objective is to learn a parameterized mapping F θ that assigns x to one of five fine-grained humanitarian categories y in {0, 1, 2, 3, 4}. The model produces a five-dimensional logit vector and the corresponding posterior class probabilities, as defined in Equation (1).
y ^ = F θ I T ∈ R 5 ,     p c = e x p y ^ c ∑ k = 1 5 e x p y ^ k .
We follow the official CrisisMMD Task 02 training, validation, and test splits without event identifiers, social links, geographic metadata, or test-set reuse. Every sample is treated as one paired image–text observation, and the same five-class label space and evaluation subset are used for baselines, ablations, and the complete model. This fixed protocol ensures that differences arise from representation and fusion rather than auxiliary cues or favorable resampling.

2.2. Overall Architecture

As illustrated in Figure 2, MSGLF-Refine comprises six connected components: OpenCLIP image–text encoding, bidirectional cross-modal interaction, modality-gated fusion, text-guided local visual attention, gated global–local refinement, and residual classification. These components operate through two connected pathways. The baseline pathway predicts from global image and text semantics. The refinement pathway retrieves localized visual evidence for every image–text pair, while a gated residual controls its contribution to the baseline. The pathways share one OpenCLIP encoding stage, and downstream modules are optimized in the same forward graph.
Figure 2. MSGLF-Refine architecture and end-to-end data flow, including bidirectional cross-attention, modality-gated fusion, text-guided patch attention, adaptive global–local refinement, zero-initialized residual correction, and five-class prediction.
For each image–text pair, OpenCLIP first extracts a global description of the visual scene, a sequence of 576 spatial patch representations, and a sentence-level representation of the accompanying post. The BCA branch exchanges information between the global image and text features and then applies channel-wise modality gating. This process emphasizes complementary cross-modal information while retaining the original global descriptors. The resulting baseline representation can independently support classification, providing a stable semantic reference before localized evidence is introduced.
The refinement pathway begins by using the meaning of the post to retrieve relevant regions from the image. Instead of treating every patch as equally informative, the textual representation guides attention toward regions that are more consistent with the expressed humanitarian intent. The retrieved local feature may emphasize affected people, damaged infrastructure, relief supplies, or rescue activities while reducing the influence of unrelated background content. The patch representations are obtained with the existing OpenCLIP visual encoder, without introducing a separate image encoder.
The model then combines three complementary sources of information: the overall visual scene, the localized image evidence, and the semantic intent of the text. A nonlinear refinement network converts this combined context into a candidate correction, while a channel-wise gate determines which dimensions contain reliable supplementary evidence. The gated output does not replace the baseline representation directly. Instead, it is introduced through the adaptive residual mechanism defined in Equation (15). The residual scale is initialized to zero, so the complete model initially reduces to the BCA baseline and the newly introduced branch cannot immediately disturb the established prediction pathway. During training, the model learns the appropriate magnitude and direction of the correction according to the classification objective.
The architecture moves from global image–text context to text-guided local evidence and then applies a targeted gated decision correction. The controlled variants defined in Section 2.12 remove or replace one stage at a time, linking performance changes to specific mechanisms.

2.3. OpenCLIP Global and Patch-Level Representations

For the visual stream, the controlled experiments use OpenCLIP’s ViT-L/14-336 training/evaluation transforms to produce 336 × 336 inputs rather than directly stretching images into squares. Each transformed image is partitioned into a 24 × 24 grid of non-overlapping 14 × 14 patches. Following ViT and CLIP principles [24,25], the patches are embedded with positional information and processed by the visual Transformer. The projected class token yields the global 768-dimensional OpenCLIP representation v g in Equation (2), while a second pass through the same visual encoder yields 576 patch tokens P that can capture local evidence concerning people, roads, buildings, and relief supplies. Thus, v g and P provide related global and local views at different spatial granularities. Very small or peripheral details may nevertheless be weakened by resizing or cropping.
v g = Norm E v c l s I , P = Norm E v p a t c h I . v g ∈ R 768 ,   P ∈ R 576 × 768 .
For the textual stream, the paired post is tokenized using the OpenCLIP tokenizer and encoded by the text Transformer. The hidden state at the end-of-text position is projected and L2-normalized to obtain the sentence-level representation t g , as defined in Equation (3). This representation is used not only in the global BCA branch, but also as the semantic query for retrieving local visual evidence. In this way, the same linguistic intent connects global fusion with local-evidence extraction.
t g = Norm E t T ∈ R 768 .
OpenCLIP is fine-tuned rather than frozen, using a lower learning rate for the pretrained backbone than for the newly introduced layers. This preserves broad image–text semantics while adapting the representation space to the five crisis categories. Features are normalized before the attention and gating modules so that feature magnitude is not treated as modality reliability.

2.4. Bidirectional Cross-Attention and Modality Gating

Because image and text reliability varies across crisis-related posts, the baseline branch combines their global representations in two independently parameterized directions. Each direction applies eight-head scaled dot-product attention [33] to one global token from each modality, followed by a residual connection and layer normalization. This exchanges learned global features without selecting among text tokens or image patches.
In the visual-side interaction, the global visual vector v g serves as the query, while the global textual vector t g is the sole key and value. Since the key/value sequence has length one, its attention weight is one; the resulting text-derived transformation is combined with the visual residual to produce the representation in Equation (4).
v ∼ = L N v v g + M H A v ← t ( v g ,   t g ,   t g ) .
In the textual-side interaction, t g similarly queries the sole global visual key/value v g . The image-derived transformation is combined with the textual residual to produce the representation in Equation (5).
t ∼ = L N t t g + M H A t ← v ( t g , v g , v g ) .
The two directions therefore provide complementary global-feature transformations rather than query-dependent token or patch selection. Their residual outputs preserve the original modality representations, while the subsequent channel-wise gate learns how strongly each modality contributes to the fused feature. Spatially selective retrieval is performed separately by the text-guided local visual attention in Section 2.5.
The two interactive representations are concatenated and passed through a two-layer gating network with a sigmoid output, as formulated in Equation (6). The resulting 768-dimensional gate performs channel-wise modality selection through Equation (7): different semantic dimensions of the same post may rely more strongly on textual or visual evidence, rather than being constrained by a single sample-level modality weight.
g b = σ M L P b ( [ t ∼ ; v ∼ ] ) , g b ∈ 0 1 768 ,
m b = g b ⊙ t ∼ + 1 − g b ⊙ v ∼ .
The gated interaction feature is then projected together with the original global text and image representations to form the BCA baseline feature b, as shown in Equation (8). This information-preserving projection keeps direct discriminative cues from the pretrained representations available to the classifier, reducing the risk that the interaction layers will overwrite global semantics in noisy samples.
b = Dropout ReLU W b m b t g v g + b b . b ∈ R 768 .

2.5. Text-Guided Local Visual Attention

A global visual vector compresses the entire image into a single representation, although crisis categories are often determined by small visual regions. Road cracks, trapped individuals, damaged facilities, and relief supplies may occupy only a limited portion of an image. Unconditional average pooling would mix these signals with background content; therefore, MSGLF-Refine uses the text representation to determine which patches should be examined.
The local branch constructs a single semantic query from t g and uses the patch sequence P as both keys and values in an eight-head attention operation. The resulting attention weights select visual regions conditioned on the meaning of the post, as described in Equation (9).
A l = Softmax t g W l Q P W l K T d h ,
Here, d h denotes the dimensionality of each attention head.
Each attention head produces normalized weights over the 576 patches and computes a weighted sum of the corresponding value vectors. The outputs of all heads are concatenated, projected, and normalized to obtain the local-evidence representation l, as given in Equation (10).
l = L N l o c a l A l P W l V W l O ∈ R 768 .
This operation converts post semantics into visual region-selection conditions. Expressions such as volunteers, collapsed building, or flooding can therefore retrieve different local evidence from the same patch sequence. The representation l is not treated as an independent classifier; instead, it serves as a targeted supplement to the global visual scene. The ablation comparison between average-pooling variants and text-guided variants tests this design choice.

2.6. Gated Global–Local Refinement

Local evidence becomes reliable only when interpreted together with the global scene and the textual context. The model therefore concatenates the global visual representation v g , the local visual evidence l, and the global textual representation t g into the three-source context vector z in Equation (11), rather than replacing the global visual feature with the local feature.
z = v g l t g ∈ R 2304 .
This context jointly encodes the disaster background, the image regions referred to by the text, and the linguistic intent of the post.
The refinement network first maps the three-source context into the candidate correction vector r in Equation (12). The Gaussian Error Linear Unit (GELU) introduces smooth nonlinearity, while dropout reduces co-adaptation among the newly introduced parameters.
r = W r 2 Dropout G E L U ( W r 1 z + b r 1 ) + b r 2 .
A separate gating network with the same structural form produces the channel-wise refinement gate g r in Equation (13). The sigmoid activation constrains each channel to a bounded contribution range, enabling the model to regulate how much of the candidate correction should be retained.
g r = σ W g 2 Dropout G E L U ( W g 1 z + b g 1 ) + b g 2 .
The candidate correction vector is modulated by the channel-wise gate and then normalized to obtain the enhanced representation e, as given in Equation (14).
e = L N e n h g r ⊙ r ∈ R 768 .
This gate does not redistribute the overall importance of image and text modalities. Instead, it controls semantic channels inside the three-source refinement vector. The mechanism can preserve incremental cues related to affected individuals, damage, or rescue activities while suppressing noise triggered by background regions, slogans, or generic disaster words. Compared with direct concatenation or fixed enhancement, this sample-dependent channel selection is better suited to crisis data with overlapping category boundaries.

2.7. Adaptive Residual Correction and Classification

To prevent the randomly initialized refinement branch from disrupting the BCA baseline at the beginning of training, the model does not replace b directly with e. Instead, it introduces a learnable scalar alpha initialized to zero, so that the initial model exactly reduces to the BCA path. During optimization, alpha increases only when the refinement feature aligns with the task gradient and provides useful incremental evidence, as expressed in Equation (15).
f = b + α e , α 0 = 0 .
This design makes local correction a learnable decision rather than a fixed architectural assumption. Compared with a fixed alpha of 1, zero initialization prevents an untrained region-selection branch from dominating the prediction. Compared with disabling the branch entirely, it preserves the ability to correct global decision boundaries on difficult samples.
The fused representation f is passed to a two-layer multilayer perceptron (MLP) classifier with ReLU activation and dropout in the hidden layer, producing logits for the five target categories as specified in Equation (16).
h = Dropout R e L U ( W 1 f + b 1 ) , y ^ = W 2 h + b 2 .
The training objective is cross-entropy with label smoothing epsilon = 0.05 [41], as shown in Equation (17). Unlike focal loss and class-balanced loss [42,43], label smoothing retains the standard cross-entropy objective while reducing overconfidence on long-tail and boundary samples; it does not replace explicit post hoc calibration when calibrated probabilities are required [44].
L = − 1 N ∑ i ∑ c q i c log p i c . q i c = 1 − ε 1 y i = c + ε 5 .
Equations (2)–(17) define a single model from image–text encoding through residual classification. The ensemble procedure is specified separately in Section 2.9.
The following experimental protocol specifies the official data split, optimization and model-selection protocol, controlled variants, evaluation metrics, cross-event tests, efficiency measurements, and interpretability analyses used to assess MSGLF-Refine.

2.8. Dataset and Official Split

The experiments are conducted on CrisisMMD v2.0 Task 02, a five-class humanitarian classification task in which each sample contains a disaster-related image, paired social media text, and one category label. The five categories are affected individuals; rescue, volunteering, or donation effort; infrastructure and utility damage; other relevant information; and not humanitarian. We preserve the original five-class setting, including minority classes and the negative class, so that the model is evaluated under realistic semantic overlap, weak image–text alignment, and non-humanitarian noise.
To ensure direct comparability with previous studies, we strictly adopt the official split of 5119 training samples, 1097 validation samples, and 1098 test samples. The training set is used only for parameter updates, the validation set is used for checkpoint selection and early stopping, and the test set is accessed only after the protocol is fixed. Table 2 reports the class distribution. Because affected individuals accounts for only 17 test samples whereas the not-humanitarian class contains 568 samples, accuracy alone would obscure minority-class behavior; macro F1-score, weighted F1-score, and class-level metrics are therefore reported alongside accuracy.
Table 2. Class distribution in the official CrisisMMD Task 02 split [1].

2.9. Training Strategy and Implementation Details

All controlled comparisons use the same 336 × 336 OpenCLIP input, official split, OpenCLIP-supplied training and evaluation transforms, optimization framework, and checkpoint-selection rule. In the evaluated configurations, both train_resize_mode and resize_mode are set to clip_default, so the direct square-stretch alternative is not used. Except for the component being tested, training conditions remain fixed so that differences can be attributed to the evaluated mechanism.
For each random seed, a model is trained independently from the same pretrained OpenCLIP weights. Training uses AdamW [45] with differential learning rates, label smoothing [41], automatic mixed precision, an exponential moving average of parameters [46], and early stopping. After each epoch, validation accuracy alone is used to select the best EMA checkpoint, preventing test-set feedback from entering model selection. Applying one predefined criterion consistently across seeds also avoids retrospectively selecting checkpoints according to whichever test metric appears most favorable. Final test predictions are generated only after checkpoint selection has been completed. The experiments were implemented using Python 3.10.21, PyTorch 2.5.1, and open_clip_torch 2.32.0.
Figure 3 summarizes the optimization behavior. Across the three seeds, validation accuracy rises rapidly and then stabilizes above 90%, while training and validation losses approach similar low-loss regions without clear late-stage divergence. The best checkpoints occur at different epochs, supporting seed-specific validation selection rather than the use of a common final epoch.
Figure 3. Multi-seed optimization dynamics from actual epoch logs: (a) validation convergence with seed uncertainty; (b) optimization phase portrait and stable basin; (c) accuracy-macro F1-score learning manifold; and (d) checkpoint-selection landscape with best checkpoints, early stopping, and patience spans.
Images are resized to 336 × 336 pixels. The batch size is 8 with gradient accumulation over 2 steps. Training runs for at most 20 epochs, with one warm-up epoch followed by cosine learning-rate scheduling. The OpenCLIP backbone uses a learning rate of 5 × 10−6, and newly introduced modules use 5 × 10−5.
For final inference, we use the validation-driven three-seed ensemble defined in Equation (18). The ensemble consists of independently trained models with seeds 42, 2024, and 3407, each represented by its best validation checkpoint. Their logits are averaged with equal weights before the softmax operation. Equal-logit averaging improves repeatability without introducing an additional learned stacking model or validation-tuned ensemble weights.
y ^ e n s = 1 3 ∑ s ∈ { 42 , 2024 , 3407 } y ^ s , y * = a r g m a x c y ^ e n s , c

2.10. Baselines and Fair-Comparison Protocol

The comparison methods are organized into three groups. The first includes early CNN-based image–text fusion models, such as VGG-16 + CNN and DenseNet + BERT [11,12], to illustrate the progression in pretrained representations. The second includes recent crisis-specific or multimodal methods, including CrisisKAN [37], CLIP-BCA-Gated [36], and MKAN-Refine [40], as the closest methodological reference points; their published scores provide benchmark context because training protocols may differ. The third consists of internal variants evaluated with the same backbone, official split, optimization settings, and checkpoint-selection rule, allowing component contributions to be assessed under controlled conditions.
The internal comparison follows a progressive P0-A-C path. P0 directly concatenates global OpenCLIP image and text representations; A adds bidirectional cross-attention and modality gating; and C introduces text-guided local visual attention together with gated global–local residual refinement. This construction turns the ablation study into a controlled mechanism test rather than a collection of unrelated variants.

2.11. Evaluation Metrics and Statistical Protocol

We report accuracy, macro precision, macro recall, macro F1-score, weighted precision, weighted recall, and weighted F1-score. Macro F1-score weights all five classes equally and is sensitive to minority-class errors; weighted F1-score reflects performance under the observed class distribution. Both are necessary for assessing large-scale triage and rare actionable categories.
Repeatability is assessed with seeds 42, 2024, and 3407. Single-model results are reported as the mean ± sample standard deviation. The final system averages the logits of the three validation-selected checkpoints with equal weights. Paired significance tests are used when matched seed-wise or fold-wise comparisons are available.
To evaluate robustness beyond the official random split, we conduct leave-one-event-out experiments over seven disaster events. In each fold, one event is held out entirely as the test domain and the remaining events are used for training and validation. This setting measures event-level transfer under domain shift. We also report computational performance to characterize deployment cost, and we use counterfactual textual-evidence tests to examine whether model confidence depends on semantically meaningful crisis words rather than arbitrary token removal.

2.12. Controlled Variants and Mechanism Mapping

The controlled variants isolate distinct mechanisms under the same backbone and training protocol. P0 directly concatenates global image and text features; A adds bidirectional cross-attention and modality gating; B0 and B compare mean-pooled and text-guided local evidence without the refinement gate; C0 replaces text-guided evidence with mean pooling under the gated refinement structure; C1 retains text guidance and gating but fixes the residual scale to one; and C denotes the complete MSGLF-Refine model with text-guided local attention, adaptive gating, and a learnable residual scale. These controls support mechanism-specific comparisons rather than undifferentiated architectural changes.

3. Results

The evaluation covers overall and class-level performance, controlled ablations, cross-event generalization, computational efficiency, and textual and visual evidence for mechanisms under one consistent protocol.

3.1. Overall Comparison with Existing Methods

Table 3 and Figure 4 place the published Task 02 results alongside our evaluation. The three-seed MSGLF-Refine ensemble reaches 92.44% accuracy and 92.41% weighted F1-score, comparing favorably with the reported CLIP-BCA-Gated and MKAN-Refine figures. The architectural distinction is clearest in the internally matched variants: the BCA baseline provides global bidirectional fusion, whereas the complete model adds text-guided patch retrieval and gated residual correction. Since the published systems were not all re-profiled under identical conditions, we use Table 3 for benchmark context and Section 3.3 and Section 3.5 for controlled mechanism and efficiency comparisons within this study.
Table 3. Contextual comparison with published CrisisMMD Task 02 results (%).
Figure 4. Contextual comparison of published CrisisMMD Task 02 results and the MSGLF-Refine evaluation. (a) Reported F1-score ranking across selected representative baselines spanning early fusion, language-enhanced fusion, multimodal attention, vision–language pretraining, and adaptive gating and (b) class-wise F1 profile of the final MSGLF-Refine ensemble. Reporting conventions for the published entries are described in Table 3.

3.2. Main Results and Class-Level Performance

The final ensemble obtains 92.44% test accuracy, 92.41% weighted F1-score, and 90.28% macro F1-score. Table 4 shows that the performance is not confined to the dominant class: the model separates not-humanitarian posts while maintaining strong recognition of infrastructure damage, rescue activity, and other relevant information.
Table 4. Class-level test performance of the final MSGLF-Refine ensemble (%).
The not-humanitarian class achieves 93.07% precision, 94.54% recall, and a 93.80% F1-score. F1-scores are 93.02% for infrastructure and utility damage, 91.80% for rescue, volunteering, or donation, and 90.04% for other relevant information. Affected individuals achieves 100.00% precision but 70.59% recall on only 17 test instances, showing conservative prediction and substantial uncertainty for the smallest class.
The final ensemble correctly identifies 12 of the 17 affected-individual posts and misses five. Two missed posts are assigned to infrastructure and utility damage, and three to not humanitarian; none is sent to the rescue or other relevant categories. This pattern is compatible with the model’s 100.00% precision but 70.59% recall: when it predicts affected individuals it is precise on this test set, while competing damage and general-content interpretations absorb 29.41% of the true cases. Case inspection points to visually dominant infrastructure damage, ambiguous evidence of personal impact, and weak image–text grounding as the main observed boundaries (Appendix A, Table A3).
Across seeds 42, 2024, and 3407, test accuracy ranges from 91.35% to 92.08% (Figure 5). The confusion matrices show similar diagonal structure, while affected individuals remains the most variable class and rescue-related posts are sometimes confused with not-humanitarian content. The recurring pattern suggests that the main residual errors reflect class scarcity and semantic overlap rather than one favorable initialization.
Figure 5. Seed-wise error signatures. The (top row) compares row-normalized confusion matrices for the three independently trained checkpoints; the (bottom row) reports their class-wise precision, recall, and F1 profiles.

3.3. Controlled Ablation and Component Contributions

Table 5 presents the three-seed controlled ablation study. Figure 6 visualizes the corresponding stability, component-wise metrics, seed-level performance space, and controlled ensemble gains. P0 directly concatenates the global OpenCLIP image and text representations, while A introduces bidirectional cross-attention and modality gating to establish the BCA baseline. B0 and B examine mean-pooled and text-guided local fusion, respectively. C0 evaluates gated mean pooling, C1 uses a fixed residual connection, and C integrates text-guided local evidence, gated global–local refinement, and adaptive residual correction into the complete MSGLF-Refine model. All variants share the same backbone, data split, optimization settings, and checkpoint-selection protocol.
Table 5. Three-seed controlled ablation results (mean ± sample standard deviation, %).
Figure 6. Multi-seed controlled ablation: stability, component-wise metrics, seed space, and ensemble gain.
Bidirectional cross-modal interaction provides the clearest single-model gain, while adding local refinement preserves the strong BCA baseline. Relative to Variant A, Variant C changes the three-seed means by +0.03 percentage points (pp) in accuracy, −0.13 pp in macro F1-score, and +0.03 pp in weighted F1-score. Given the already strong BCA representation and the zero-initialized gated residual, these small changes indicate that the local pathway can be incorporated without materially disrupting the aggregate single-model performance.
Under equal-logit aggregation, Variant C shows a larger observed uplift relative to its own three-seed single-model mean. Variant C gains +0.820 percentage points (pp) in accuracy, +0.449 pp in macro F1-score, and +0.807 pp in weighted F1-score; the corresponding gains for Variant A are +0.486, +0.040, and +0.472 pp. The additional observed ensemble uplift for C is therefore +0.334, +0.409, and +0.335 pp, respectively. The disagreement diagnostics are compatible with this pattern: Variant C produces 61 seed-disagreement cases (5.56% of the test set) and classifies 62.30% of them correctly after aggregation, compared with 58 cases (5.28%) and 56.90% accuracy for Variant A. Together, these descriptive results support a plausible seed-dependent complementarity account of the larger observed ensemble gain, rather than uniform improvement in every seed; the differing disagreement subsets do not, by themselves, isolate a causal mechanism. The final Variant C ensemble reaches 92.44% accuracy, 90.28% macro F1-score, and 92.41% weighted F1-score, exceeding Variant A by +0.36, +0.28, and +0.36 pp. The bootstrap intervals include zero and the exact McNemar p value is 0.503; thus, the ensemble point estimates favor C, while additional matched runs would help assess the robustness of this pattern.
Text-length stratification provides a descriptive check on the query pathway. In the shortest empirical quartile (n = 284), Variant C reaches 90.49% accuracy versus 90.14% for Variant A; in the third quartile (n = 303), the corresponding accuracies are 92.08% and 91.09%. The other quartiles have equal accuracy, and no quartile shows an observed accuracy decrease for Variant C. These point estimates are compatible with retained performance among the shorter posts represented in this benchmark, but the subgroup bootstrap intervals include zero and token count is only a proxy for semantic informativeness. The fixed 11–20-token group contains only 17 samples, so performance for extremely sparse queries remains less certain (Appendix A, Table A4).

3.4. Generalization to Unseen Disaster Events

Seven leave-one-event-out experiments evaluate transfer to held-out historical disaster events (Table 6 and Figure 7). Relative to P0, the complete model improves the event mean for accuracy, macro F1-score, and weighted F1-score and raises weighted F1-score for every held-out event. Text-guided retrieval can focus on transferable entities such as affected people, damaged infrastructure, blocked roads, and relief supplies, while the residual gate limits harmful corrections for atypical image–text pairs. Event means and pooled accuracy are both reported because they respectively weight disasters equally and weight individual samples equally. The remaining between-event dispersion shows reduced, but not eliminated, sensitivity to event shift.
Table 6. Cross-event summary over seven leave-one-event-out evaluations (event mean ± standard deviation, %).
Figure 7. Cross-event generalization response of MSGLF-Refine under leave-one-event-out evaluation. (a) Shape-preserving performance profiles for P0, the BCA baseline (A), and the full model (C) across seven held-out disasters ordered by P0 macro F1-score. (b) Kernel-density distributions of event-wise accuracy and macro F1-score gains from P0 to MSGLF-Refine; rug marks denote individual events and dashed lines denote means.

3.5. Computational Performance

Table 7 quantifies the cost of global–local refinement. A single MSGLF-Refine model increases parameters and latency relative to the BCA baseline because it retains patch features and evaluates additional attention and gating layers, but it does not require a second image encoder. Its throughput remains suitable for batched triage and is substantially higher than that of the three-model ensemble. A single model suits continuous monitoring under the reported hardware setting; the ensemble can instead be used for offline analysis or selected high-consequence cases.
Table 7. Computational performance comparison.
Peak-memory measurements support the same deployment distinction. The ensemble stores three model instances and increases queueing and memory pressure during bursts; absolute latency will also vary with hardware, batch size, and preprocessing. A possible deployment would process the main stream with one model and flag low-confidence or high-risk posts according to calibrated confidence, prediction entropy, cross-modal disagreement, or category-specific thresholds.

3.6. Interpretability of Multimodal Evidence

Textual evidence is evaluated with leave-one-token-out attribution and counterfactual removal and retention at 10%, 20%, and 30% budgets (Figure 8). Feature-attribution methods motivate controlled perturbation analysis [47,48], while faithfulness-oriented research distinguishes plausible attention from evidence whose intervention changes model behavior [39,49,50]. Removing model-ranked tokens causes larger confidence changes and more label changes than removing random tokens, whereas retaining ranked tokens better preserves the original prediction. At the 20% budget, ranked removal reduces confidence by 0.1999 compared with 0.0312 for random removal. The consistent advantage across budgets indicates reliance on related phrases rather than a single disaster keyword.
Figure 8. Counterfactual textual-evidence fidelity. (a) Percentage-point advantages of model-ranked tokens over random controls for removal damage, retention preservation, and prediction flips at 10%, 20%, and 30% intervention budgets. (b) Radar summary of the same cross-budget evidence profile.
These scores measure model sensitivity rather than human-validated explanations: deletion can reduce fluency, and related words may compensate for one another. We therefore interpret the ranked-versus-random differences together with the localized visual evidence in Figure 9. Analysts compare influential textual cues with semantically relevant image regions, and manually review samples whenever the two modalities disagree materially.
Figure 9. Six-view visual evidence across five crisis categories. Rows 1–3 present original inputs, text-guided local attention, and semantic spotlight views for the model-attention sample set. Rows 4–6 present a second set of CrisisMMD inputs with synchronized saliency and spotlight views.
Prediction flips should also be considered with confidence damage. A label can remain unchanged despite an operationally important margin reduction, while a flip between adjacent classes does not by itself prove that the original evidence was spurious.
Figure 9 compares original images, text-guided attention, and semantic spotlight views across all five classes and varied scene conditions. Compact responses are expected when a small object resolves the category; broader responses are plausible for distributed evidence such as crowds, debris, or widespread infrastructure damage. Cases with concentrated textual importance but diffuse visual attention reveal possible modality imbalance and provide useful review cues.
Figure 10 provides a class-conditional view of textual evidence. Human-impact terms are most strongly associated with affected individuals, relief-action terms with rescue, volunteering, or donation, and damage-oriented terms with infrastructure and utility damage. Event names and broad disaster words are distributed across multiple categories, indicating that they mainly provide contextual information rather than sufficient evidence for a specific prediction.
Figure 10. Class-wise token attribution matrix for textual evidence. Scores are normalized within each class, and only non-negligible associations are annotated. Event-specific terms such as hurricane or earthquake names are treated as contextual cues rather than sufficient class evidence.
The class-conditional associations are operationally coherent: they distinguish who is affected, what is damaged, and whether a response activity is underway. Broad disaster terms and event names span multiple classes, whereas more specific human-impact, relief, and damage terms align with the corresponding categories. The counterfactual and visual analyses alone suggest that predictions do not rely exclusively on event-specific textual shortcuts, although these tests cannot establish causal reasoning.

4. Discussion

4.1. Why Global–Local Refinement Improves Fine-Grained Crisis Classification

The ablations support MSGLF-Refine as a decision-level refinement of a strong BCA baseline. The local branch preserves comparable average single-model performance while generating seed-dependent corrections on a small but consequential subset of samples. Although it does not reduce the across-seed dispersion of the aggregate metrics, it creates useful predictive diversity: the complete model attains both a larger uplift from equal-logit aggregation and higher accuracy on seed-disagreement cases than the BCA baseline. The principal contribution of local refinement is therefore selective complementarity—recovering local visual evidence that can be combined across seeds—rather than a uniform per-seed increase.
The affected-individual error analysis makes the practical value of class-level reporting clear. The ensemble avoids false positives for this category in the 1098-sample test set and correctly identifies 12 of its 17 true cases. The five missed cases cluster at meaningful class boundaries: two emphasize infrastructure damage and three are interpreted as non-humanitarian content. This precise yet selective behavior can help prioritize clear human-impact posts, while uncertain or competing-class cases should still be reviewed when recall is operationally important.

4.2. Potential Use in Human-Reviewed Crisis Triage

For emergency management, MSGLF-Refine can support human review of crisis posts through class scores, localized image evidence, influential text cues, and cross-modal disagreement. These outputs offer concrete cues for prioritizing verification and routing posts toward damage assessment, rescue coordination, or humanitarian support. The study evaluates predictive behavior and evidence patterns; it has not yet tested analyst interaction, decision time, or explanation usability. The appropriate application claim is therefore an evidence-assisted triage within a human-reviewed workflow.

4.3. Accuracy, Efficiency, and Deployment Modes

The single model retains the complete evidence pathway at 36.65 samples/s in the reported setting, whereas the three-seed ensemble provides the strongest accuracy and stability at approximately three times the backbone cost. A differentiated deployment could use one model for continuous monitoring and reserve the three-model ensemble for low-confidence or high-consequence samples. Because timing depends on hardware and batching, the reported measurements should be interpreted as relative costs rather than universal latency values.

4.4. Limitations and Future Work

The experiments establish strong benchmark performance under the official CrisisMMD Task 02 protocol, with several clear directions for extension. The affected-individual class contains only 17 test samples and five false negatives, so its 70.59% recall is sensitive to individual decisions; larger rare-class cohorts and recall-oriented calibration are needed before safety-critical use. OpenCLIP’s fixed 336 × 336 input preserves access to 576 patch representations but can still weaken very small or peripheral evidence, motivating multi-scale encoding or adaptive cropping. The text-guided query performs favorably in the observed shortest quartile, yet extremely short, noisy, or link-only posts are sparsely represented and merit targeted evaluation. Leave-one-event-out testing demonstrates transfer among historical disasters, not temporal or current-platform generalization; prospective time-separated and cross-platform studies should examine vocabulary, media style, and multilingual drift. Finally, no crisis analyst participated in the present evaluation. Human-review utility, decision time, trust calibration, and explanation usability require a dedicated user study; counterfactual and visual evidence analyses do not by themselves establish causal or human-validated explanations.

5. Conclusions

This study introduced MSGLF-Refine, combining bidirectional cross-attention, text-guided local visual retrieval, and adaptive gated residual correction for five-class crisis information mining. On the official CrisisMMD Task 02 split, its three-seed ensemble achieves 92.44% accuracy, 92.41% weighted F1-score, and 90.28% macro F1-score. Controlled comparisons show that local refinement maintains comparable single-model performance to the BCA baseline while producing a larger observed ensemble uplift. The complete model also achieves higher accuracy on its own seed-disagreement cases than the BCA ensemble does on its corresponding cases. In leave-one-event-out tests, it improves weighted F1-score over global concatenation for all seven held-out historical events. Class-level and text-length analyses identify rare-class and sparse-text cases that warrant further evaluation. Together, these results demonstrate the effectiveness of MSGLF-Refine for fine-grained crisis classification on CrisisMMD, while future work will assess its robustness on temporally separated and cross-platform data.

Author Contributions

Conceptualization, Y.Z. and S.L.; methodology, Y.Z.; software, Y.Z.; formal analysis, Y.Z.; investigation, Y.Z.; data curation, Y.Z.; writing—original draft preparation, Y.Z.; writing—review and editing, S.L., Z.P., Q.L. and G.L.; visualization, Y.Z.; supervision, S.L., Z.P., Q.L. and G.L.; project administration, S.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Key Research and Development Program of China, grant number 2024YFC3908000.

Institutional Review Board Statement

Not applicable. This study used an existing publicly available benchmark dataset and did not involve intervention with humans or animals.

Data Availability Statement

The CrisisMMD v2.0 data analyzed in this study are available through the dataset described in Reference [1], subject to the access conditions and platform terms specified by the dataset providers.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Appendix A.1. Statistical and Ensemble Complementarity Analysis

The analyses below use the official CrisisMMD Task 02 test split and the matched seeds 42, 2024, and 3407 reported in the main text. Variant A is the BCA baseline, and Variant C is the complete MSGLF-Refine model.
Table A1. Paired statistical comparison of Variants A and C at the seed and ensemble levels.
Although the inferential uncertainty prevents a claim of uniform per-seed improvement, all three ensemble point estimates favor Variant C. Table A2 shows a larger observed aggregation uplift for C relative to its own single-model mean, motivating a descriptive examination of prediction complementarity.
Table A2. Ensemble uplift and seed-disagreement diagnostics for Variants A and C.
Variant C shows a larger observed aggregation uplift for every reported metric. Its slightly broader disagreement set retains the same oracle coverage as Variant A, while the equal-logit ensemble classifies its respective disagreement cases 5.40 pp more accurately. Because the disagreement subsets are not identical, these diagnostics are consistent with useful seed-dependent complementarity but do not establish that the local module alone caused the larger ensemble gain.

Appendix A.2. Affected-Individual Error Analysis

The final three-seed equal-logit ensemble identifies 12 of 17 affected-individual test posts. The five false negatives are listed below; case interpretations are qualitative inspections of the paired image and text, not separate ground-truth annotations.
Table A3. False-negative destinations for affected-individual posts under final ensemble inference.
Four of the five misses are unanimous across the three seeds, while test index 636 is a boundary case in which mean-logit aggregation changes an otherwise correct two-seed majority. This distinction directs future work toward class-boundary calibration and review rules rather than treating every miss as random initialization noise.

Appendix A.3. OpenCLIP Text-Length Stratification

The following analysis groups the same 1098 test posts by valid OpenCLIP content-token count and compares the frozen Variant A and C equal-logit ensembles. Absolute-length bins are shown alongside empirical quartiles because only 17 posts fall into the 11–20-token bin.
Table A4. Accuracy by OpenCLIP content-token length for the BCA and MSGLF-Refine ensembles.
Across the empirical quartiles, the full model matches or exceeds the BCA ensemble in observed accuracy. This descriptive pattern is compatible with continued utility of text-guided retrieval in the sampled shorter-post range, while extremely sparse queries remain a useful target for further stress testing.

References

  1. Alam, F.; Ofli, F.; Imran, M. CrisisMMD: Multimodal Twitter datasets from natural disasters. In Proceedings of the International AAAI Conference on Web and Social Media; AAAI Press: Palo Alto, CA, USA, 2018; pp. 465–473. [Google Scholar] [CrossRef] [Scilit]
  2. Shetty, N.P.; Bijalwan, Y.; Chaudhari, P.; Shetty, J.; Muniyal, B. Disaster assessment from social media using multimodal deep learning. Multimed. Tools Appl. 2024, 84, 18829–18854. [Google Scholar] [CrossRef] [Scilit]
  3. Vieweg, S.; Hughes, A.L.; Starbird, K.; Palen, L. Microblogging during two natural hazards events: What Twitter may contribute to situational awareness. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems; ACM Press: New York, NY, USA, 2010; pp. 1079–1088. [Google Scholar]
  4. Imran, M.; Castillo, C.; Diaz, F.; Vieweg, S. Processing social media messages in mass emergency: A survey. ACM Comput. Surv. 2015, 47, 67. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Olteanu, A.; Vieweg, S.; Castillo, C. What to expect when the unexpected happens: Social media communications across crises. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing; ACM Press: New York, NY, USA, 2015; pp. 994–1009. [Google Scholar]
  6. Imran, M.; Mitra, P.; Castillo, C. Twitter as a lifeline: Human-annotated Twitter corpora for NLP of crisis-related messages. In Proceedings of the Tenth International Conference on Language Resources and Evaluation; ELRA: Paris, France, 2016; pp. 1638–1643. [Google Scholar]
  7. Alam, F.; Qazi, U.; Imran, M.; Ofli, F. HumAID: Human-annotated disaster incidents data from Twitter with deep learning benchmarks. In Proceedings of the International AAAI Conference on Web and Social Media; AAAI Press: Washington, DC, USA, 2021; pp. 933–942. [Google Scholar]
  8. Alam, F.; Sajjad, H.; Imran, M.; Ofli, F. CrisisBench: Benchmarking crisis-related social media datasets for humanitarian information processing. In Proceedings of the International AAAI Conference on Web and Social Media; AAAI Press: Palo Alto, CA, USA, 2021; pp. 923–932. [Google Scholar]
  9. Imran, M.; Castillo, C.; Lucas, J.; Meier, P.; Vieweg, S. AIDR: Artificial intelligence for disaster response. In Proceedings of the 23rd International Conference on World Wide Web Companion; ACM: New York, NY, USA, 2014; pp. 159–162. [Google Scholar]
  10. Olteanu, A.; Castillo, C.; Diaz, F.; Vieweg, S. CrisisLex: A lexicon for collecting and filtering microblogged communications in crises. In Proceedings of the International AAAI Conference on Web and Social Media; AAAI Press: Palo Alto, CA, USA, 2014; pp. 376–385. [Google Scholar]
  11. Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv 2015, arXiv:1409.1556. [Google Scholar]
  12. Huang, G.; Liu, Z.; Van Der Maaten, L.; Weinberger, K.Q. Densely connected convolutional networks. In Proceedings of the CVPR; IEEE: Piscataway, NJ, USA, 2017; pp. 4700–4708. [Google Scholar]
  13. Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the NAACL-HLT; ACL: Stroudsburg, PA, USA, 2019; pp. 4171–4186. [Google Scholar]
  14. Nguyen, D.Q.; Vu, T.; Nguyen, A.T. BERTweet: A pre-trained language model for English Tweets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations; ACL: Stroudsburg, PA, USA, 2020; pp. 9–14. [Google Scholar]
  15. Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. RoBERTa: A robustly optimized BERT pretraining approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
  16. Liu, J.; Singhal, T.; Blessing, L.T.M.; Wood, K.L.; Lim, K.H. CrisisBERT: A robust transformer for crisis classification and contextual crisis embedding. arXiv 2020, arXiv:2005.06627. [Google Scholar]
  17. Sirbu, I.; Sosea, T.; Caragea, C.; Caragea, D.; Rebedea, T. Multimodal semi-supervised learning for disaster tweet classification. In Proceedings of the COLING; ACL: Stroudsburg, PA, USA, 2022; pp. 2711–2723. [Google Scholar]
  18. Mandal, B.; Khanal, S.; Caragea, D. Contrastive learning for multimodal classification of crisis related tweets. In Proceedings of the Web Conference; ACM: New York, NY, USA, 2024; pp. 4555–4564. [Google Scholar]
  19. Ofli, F.; Alam, F.; Imran, M. Analysis of social media data using multimodal deep learning for disaster response. arXiv 2020, arXiv:2004.11838. [Google Scholar]
  20. Zou, Z.; Gan, H.; Huang, Q.; Cai, T.; Cao, K. Disaster image classification by fusing multimodal social media data. ISPRS Int. J. Geo Inf. 2021, 10, 636. [Google Scholar] [CrossRef] [Scilit]
  21. Abavisani, M.; Wu, L.; Hu, S.; Tetreault, J.; Jaimes, A. Multimodal categorization of crisis events in social media. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2020; pp. 14679–14689. [Google Scholar]
  22. Dar, S.S.; Rehman, M.Z.U.; Bais, K.; Haseeb, M.A.; Kumar, N. A social context-aware graph-based multimodal attentive learning framework for disaster content classification during emergencies. Expert Syst. Appl. 2025, 259, 125337. [Google Scholar] [CrossRef] [Scilit]
  23. Algiriyage, N.; Prasanna, R.; Stock, K.; Doyle, E.E.H. Multi-source multimodal data and deep learning for disaster response: A systematic review. SN Comput. Sci. 2022, 3, 92. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the ICML; PMLR: New York, NY, USA, 2021; pp. 8748–8763. [Google Scholar]
  25. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A. An image is worth 16 × 16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
  26. Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.H.; Li, Z.; Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the ICML; PMLR: New York, NY, USA, 2021; pp. 4904–4916. [Google Scholar]
  27. Li, L.H.; Yatskar, M.; Yin, D.; Hsieh, C.J.; Chang, K.W. VisualBERT: A simple and performant baseline for vision and language. arXiv 2019, arXiv:1908.03557. [Google Scholar]
  28. Chen, Y.C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; Liu, J. UNITER: Universal image-text representation learning. In Proceedings of the 16th European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 104–120. [Google Scholar]
  29. Kim, W.; Son, B.; Kim, I. ViLT: Vision-and-language transformer without convolution or region supervision. In Proceedings of the ICML; PMLR: New York, NY, USA, 2021; pp. 5583–5594. [Google Scholar]
  30. Singh, A.; Hu, R.; Goswami, V.; Couairon, G.; Galuba, W.; Rohrbach, M.; Kiela, D. FLAVA: A foundational language and vision alignment model. In Proceedings of the the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022; pp. 15638–15650. [Google Scholar]
  31. Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; Hoi, S.C.H. Align before fuse: Vision and language representation learning with momentum distillation. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; pp. 9694–9705. [Google Scholar]
  32. Li, J.; Li, D.; Savarese, S.; Hoi, S.C.H. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the ICML; PMLR: New York, NY, USA, 2023; pp. 19730–19742. [Google Scholar]
  33. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; p. 30. [Google Scholar]
  34. Lu, J.; Yang, J.; Batra, D.; Parikh, D. Hierarchical question-image co-attention for visual question answering. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2016; Volume 29, pp. 289–297. [Google Scholar]
  35. Arevalo, J.; Solorio, T.; Montes-y-Gómez, M.; González, F.A. Gated multimodal units for information fusion. arXiv 2017, arXiv:1702.01992. [Google Scholar]
  36. Li, S.; Liu, Q.; Pan, Z.; Wu, X. CLIP-BCA-Gated: A dynamic multimodal framework for real-time humanitarian crisis classification with bi-cross-attention and adaptive gating. Appl. Sci. 2025, 15, 8758. [Google Scholar] [CrossRef] [Scilit]
  37. Gupta, S.; Saini, N.; Kundu, S.; Das, D. CrisisKAN: Knowledge-infused and explainable multimodal attention network for crisis event classification. In Advances in Information Retrieval (ECIR 2024); Lecture Notes in Computer Science; Springer Nature: Cham, Switzerland, 2024; pp. 18–33. [Google Scholar] [CrossRef] [Scilit]
  38. Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Los Alamitos, CA, USA, 24–27 October 2017; pp. 618–626. [Google Scholar]
  39. Jain, S.; Wallace, B.C. Attention is not explanation. In Proceedings of the NAACL-HLT; Association for Computational Linguistics (ACL): Stroudsburg, PA, USA, 2019; pp. 3543–3556. [Google Scholar]
  40. Wang, Z.; Li, S.; Liu, Q.; Pan, Z.; Sun, X. MKAN-Refine: Fine-grained crisis information mining via reliability-aware nonlinear refinement and Kolmogorov-Arnold networks. IEEE Access 2026, 14, 66021–66039. [Google Scholar] [CrossRef] [Scilit]
  41. Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; Wojna, Z. Rethinking the Inception architecture for computer vision. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 2818–2826. [Google Scholar]
  42. Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Los Alamitos, CA, USA, 24–27 October 2017; pp. 2980–2988. [Google Scholar]
  43. Cui, Y.; Jia, M.; Lin, T.Y.; Song, Y.; Belongie, S. Class-balanced loss based on effective number of samples. In Proceedings of the 30th Annual IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 9268–9277. [Google Scholar]
  44. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On calibration of modern neural networks. In Proceedings of the ICML; PMLR: New York, NY, USA, 2017; pp. 1321–1330. [Google Scholar]
  45. Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. arXiv 2019, arXiv:1711.05101. [Google Scholar]
  46. Tarvainen, A.; Valpola, H. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  47. Ribeiro, M.T.; Singh, S.; Guestrin, C. Why should I trust you? Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2016), San Francisco, CA, USA, 13–17 August 2016; pp. 1135–1144. [Google Scholar]
  48. Lundberg, S.M.; Lee, S.I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; p. 30. [Google Scholar]
  49. DeYoung, J.; Jain, S.; Rajani, N.F.; Lehman, E.; Xiong, C.; Socher, R.; Wallace, B.C. ERASER: A benchmark to evaluate rationalized NLP models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), Virtual Event, 5–10 July 2020; pp. 4443–4458. [Google Scholar]
  50. Chefer, H.; Gur, S.; Wolf, L. Transformer interpretability beyond attention visualization. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual Event, 19–25 June 2021; pp. 782–791. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.