Next Article in Journal
Detection and Tracking of Medicanes Through DeMeTrA Self-Supervised Vision Transformer
Previous Article in Journal
Task-Guided CycleGAN for Domain-Adapted Coastal Boundary Segmentation in Remote Sensing Images
Previous Article in Special Issue
The Topological Detection of Spatially Proximate Emitters in Spaceborne-Radio-Environment Maps: An ImprovedPersistent-Homology Approach
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Marine Oil Film Segmentation Based on GCN-IGJO Method

1
Shenzhen Institute of Guangdong Ocean University, Shenzhen 518108, China
2
Naval Architecture and Shipping College, Guangdong Ocean University, Zhanjiang 524088, China
3
Guangdong Provincial Key Laboratory of Intelligent Equipment for South China Sea Marine Ranching, Guangdong Ocean University, Zhanjiang 524088, China
4
Technical Research Center for Ship Intelligence and Safety Engineering of Guangdong Province, Zhanjiang 524088, China
5
Saint-Petersburg Institute for Shipbuilding & Marine Technology, Guangdong Ocean University, Zhanjiang 524088, China
6
Saint Petersburg State Marine Technical University, 3 Lotsmanskaya Street, Saint Petersburg 190121, Russia
7
Navigation College, Dalian Maritime University, Dalian 116026, China
8
Faculty of Mechanical Engineering, Universiti Teknologi Malaysia, Johor Bahru 81310, Malaysia
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(16), 2660; https://doi.org/10.3390/rs18162660
Submission received: 28 May 2026 / Revised: 28 July 2026 / Accepted: 5 August 2026 / Published: 7 August 2026
(This article belongs to the Special Issue Advances in Remote Sensing Image Target Detection and Recognition)

Highlights

What are the main findings?
  • GCN-IGJO achieves satisfactory segmentation performance, with 96.7% precision, 82.8% recall, 89.2% F1/Dice, and 80.5% IoU, significantly outperforming GCN-GJO/Otsu, NO-GCN-IGJO, and GCN-LA across all metrics.
  • Both GCN and IGJO are individually critical: GCN enables effective ROI extraction via high-order spatial–contextual features (removing it drops IoU from 80.5% to 54.0%), while IGJO’s low-threshold bias strategies improve precision from 66.6% to 96.7% compared to the standard GJO/Otsu.
What are the implications of the main findings?
  • Verified the effectiveness of the GCN method in X-band radar image segmentation tasks, which is capable of adapting to characteristics of radar images such as multi-noise and weak targets.
  • The GCN + domain-informed metaheuristic framework is transferable, demonstrating that embedding physical priors (e.g., oil films reside in low gray-value regions) into an optimizer’s search design can substantially boost performance—a paradigm applicable to other weak-target segmentation problems in remote sensing.

Abstract

Marine oil spills pose significant threats to ecosystems and coastal economies, making accurate oil film detection from remote sensing data a critical task. This study proposes a two-stage method, GCN-IGJO, for segmenting oil films in challenging X-band radar images. The method first uses a Graph Convolutional Network (GCN) to learn high-order features from pixel data, enabling effective initial region extraction. Subsequently, an Improved Golden Jackal Optimization (IGJO) algorithm is introduced to perform precise threshold segmentation, incorporating specialized strategies to bias the search towards the low-intensity values characteristic of oil slicks. Experimental comparisons demonstrate that the proposed GCN-IGJO method achieves a superior balance between high precision and good recall, outperforming several baseline and alternative methods. The results validate the effectiveness of combining deep graph learning with an enhanced metaheuristic optimizer for the accurate segmentation of weak-target oil films.

1. Introduction

Global marine economic activities have been intensifying, with offshore oil exploitation and maritime transportation playing a particularly prominent role [1,2]. However, the risks contained in it are also constantly developing, which continue to threaten the marine ecological environment, coastal economy and human communities [3]. In recent years, the ecological and economic losses caused by various types of large-scale oil spill accidents around the world have been severe. The impact of the Deepwater Horizon oil spill on ecosystems in the northern Gulf of Mexico continues to this day [4,5] and an oil spill in Brazil in 2020 is considered to be the largest disaster in the tropical coastal area, which has affected the local ecosystem and human economic activities on a large scale [6]. Therefore, the accurate identification of oil spill film has important research value. It plays an indispensable role in improving the response speed of oil spill and preventing the expansion of oil spill damage [7]. At present, oil film recognition technology based on Synthetic Aperture Radar (SAR) remote sensing and X-band radar has become the mainstream method in oil film recognition, which is suitable for different application scenarios [8,9].
SAR utilizes the sensitivity of microwave echo signals to sea surface roughness, and realizes the detection through the smoothing effect of oil film on the sea surface capillary ripples. It has all-day and all-weather monitoring capabilities, especially for emergency monitoring under severe meteorological conditions [10,11]. Cui et al. proposed an Enhanced Unsupervised Domain Adaptation and iterative Pseudo-label Refinement (EUDA-PLR) method for cross-event oil spill SAR image segmentation, aiming to improve the universality of a single model [12]. Sun et al. proposed a Multi-scale Alignment and Fusion U-Shape Transformer network (MAF-UFormer), which enhances the feature representation of SAR images by integrating multi-scale fusion and proxy attention mechanisms and achieved good recognition results [13]. Sathya, S et al. proposed a multi-fusion deep learning model to accurately distinguish oil-like regions in SAR images [14].
Shore-based or ship-borne radar systems have unique advantages in oil film monitoring in offshore and port areas. Compared with SAR, radar monitoring based on X-band has higher spatial and temporal resolution and real-time performance, which can provide timely and accurate situation information for on-site emergency response [15]. Using the spatio-temporal deep neural network model, Zhao et al. extracted key spatial and temporal features from the synthesized X-band radar images and achieved effective estimation and prediction of wave height [16]. Serafino and Bianco conducted experiments and measurements on the actual ability of X-band radars to detect small garbage islands (SGI) floating on the sea surface, proving the potential of X-band radars in marine pollutant control [17]. Liu et al. proposed a method combining Otsu thresholding with aspect ratio analysis to suppress the noise in the X-band radar oil spill image, while preserving the original features of the radar image [18].
To address the challenges of weak boundary features and irregular structures in oil film regions within X-band radar imagery, this paper proposes a multi-modal identification method that integrates a Graph Convolutional Network (GCN) with an Improved Golden Jackal Optimization (IGJO) algorithm [19,20,21,22].
The main contributions of this work are summarized as follows:
  • A graph convolutional network (GCN)-based feature extraction module is proposed, which captures high-order semantic features of X-band radar images through double-layer graph convolution combined with residual connections, effectively integrating the spatial neighborhood information of pixels with nonlinear representations.
  • An improved golden jackal optimization algorithm (IGJO) is developed for oil film threshold segmentation. By redesigning a time-varying lower-bound reward function and incorporating a dynamic energy attenuation strategy, IGJO jointly optimizes inter-class variance and class balance constraints under the guidance of the golden jackal partner cooperation mechanism, significantly enhancing threshold search efficiency and robustness in oil film regions.
  • A hybrid framework integrating GCN-based feature learning with IGJO-driven threshold segmentation is constructed. Experimental results demonstrate that the proposed combination strategy outperforms conventional methods in reliability, providing an effective technical solution for marine oil film recognition in X-band radar imagery.

2. Experimental Data and Methods

2.1. Experimental Data and Preprocessing

The X-band marine radar data used in this study were acquired by the training ship “Yukun” of Dalian Maritime University during the Dalian Xingang port oil pipe-line explosion accident, in which over 1500 tons of crude oil spilled into the sea [23]. A Sperry Marine B.V. (New Malden, UK) X-band radar system was employed, featuring an 8 ft waveguide split antenna with horizontal polarization and full 360° azimuthal coverage. The radar operates at three pulse duration settings (50 ns/250 ns/750 ns) with corresponding pulse repetition frequencies of 3000 Hz/1800 Hz/785 Hz and selectable surveillance ranges of 0.5–12 nautical miles; the antenna rotational velocity is adjustable between 28 and 45 RPM. For this study, the detection range was set to 0.75 nautical miles (~1389 m), yielding an output image resolution of 1024 × 1024 pixels. The original radar image is displayed in the Plan Position Indicator (PPI) format, with the ship located at the image center and the radial direction corresponding to the detection distance. The specific radar original image is shown in Figure 1.
The acquired raw radar images undergo a five-stage preprocessing pipeline to suppress systematic noise, correct radiometric distortions, and enhance target visibility before subsequent analysis.
(1)
Polar-to-Cartesian coordinate conversion. The raw PPI image is inherently polar: the azimuth angle increases along the angular coordinate, and the radial distance increases outward from the center. This polar representation is unsuitable for constructing graph structures in pixel-level feature learning. The GCN module relies on spatial adjacency relationships defined over a regular Cartesian grid; each pixel’s position is represented by row and column coordinates, and the Euclidean distance in the joint spatial–gray-value feature space determines the graph edge weights. Operating directly on polar coordinates would produce geometrically distorted neighborhoods, where pixels with the same angular separation but different radial distances correspond to vastly different physical areas, undermining the consistency of local feature aggregation. To resolve this, the PPI image is converted to a Cartesian coordinate system via nearest-neighbor interpolation. In the converted image, the horizontal axis corresponds to the azimuth angle and the vertical axis corresponds to the detection range, providing a geometrically standardized 2D input in which each pixel has a well-defined spatial location that can be directly encoded into the GCN’s feature matrix and adjacency matrix [24].
(2)
Co-channel interference suppression. Marine radar images frequently contain co-channel interference (CCI): radially oriented bright streaks caused by electromagnetic pulses from nearby radars operating at the same frequency band. Since CCI spikes exhibit high local intensity gradients that substantially differ from both the sea clutter background and oil film regions, they can be effectively suppressed without compromising target information. A 3 × 3 mean filter is applied to attenuate these impulsive artifacts while preserving the overall image structure [25].
(3)
Speckle noise suppression. Speckle noise, inherent to coherent radar systems, manifests as granular fluctuations that degrade image quality and complicate pixel-level segmentation. To reduce speckle while retaining edge sharpness, a second 3 × 3 mean filter is applied. The mild kernel size is chosen deliberately: larger kernels would over-smooth the thin oil film boundaries that are critical for subsequent fine segmentation [26].
(4)
Grayscale normalization. Due to variations in radar receiver gain, antenna pattern, and range-dependent signal attenuation, raw pixel intensities may exhibit systematic biases and an uneven dynamic range across the image. A linear grayscale normalization is applied to map the pixel intensity values to a standardized range of [0, 1], which compensates for the intensity roll-off and ensures radiometric consistency across the full field of view for subsequent processing stages [27].
(5)
Local contrast enhancement. Oil films in X-band radar images typically appear as low-gray-value regions with weak contrast against the surrounding sea clutter, making direct thresholding unreliable. Contrast Limited Adaptive Histogram Equalization (CLAHE) is adopted to enhance local contrast while preventing noise amplification. The tile size is set to 8 × 8 pixels with a clip limit of 2.0, which effectively accentuates the boundaries between oil films and the background without over-enhancing the speckle residue [28].
Finally, to manage the computational complexity of the subsequent graph construction and GCN inference, the enhanced image is downscaled by a factor of s = 0.5, reducing the total pixel count to within approximately 50,000. An image of the preprocessing process is shown in Figure 2.

2.2. ROI Extraction Based on GCN

This section proposes an ROI initial extraction method based on a graph convolutional network. This method first uses the spatial and gray information of pixels to construct a sparse adjacency matrix to define the graph structure. Then, the high-order feature representation of pixels is learned by two-layer graph convolution operation. Finally, K-Means clustering is performed based on learning features, and ROI is obtained by morphological post-processing. The specific steps are shown in Figure 3.
The segmentation criteria in this study are established based on the following physical and computational principles. The oil films in X-band radar images exhibit consistently low gray values due to the damping of Bragg scattering by the oil layer, which serves as the fundamental physical prior. The spatial adjacency constraint, pixels within a 5 × 5 local neighborhood, is employed to model contextual relationships, ensuring that segmentation respects the continuity and boundary characteristics of oil slick regions. Additionally, high-order semantic features are extracted via graph convolution, transforming single-pixel gray values into 16-dimensional feature vectors that encode both appearance and spatial context, thereby providing a more discriminative representation for clustering. Taken together, these criteria form a multi-level segmentation framework that jointly leverages domain-specific priors, spatial structure, learned feature representations, and optimization constraints.
A critical challenge in radar-based oil spill segmentation is the presence of lookalike dark areas, such as low-wind zones, biogenic surface films, ship wakes, current shears, rain cells, and radar shadow regions. These areas appear as low-intensity regions and can be misclassified as oil films. The proposed method addresses this challenge through two complementary mechanisms. (1) The GCN-based feature learning module captures high-order spatial–contextual patterns that extend beyond intensity alone: oil slicks typically exhibit spatially coherent, smoothly varying boundaries, whereas many lookalikes (e.g., ship wakes or shadow regions) have distinct geometric signatures (linear, sharp-edged, or fragmented). By encoding the joint spatial–gray-value feature space, the GCN can learn to differentiate these structural differences. (2) The sparse adjacency graph with k = 1 nearest neighbor within a 5 × 5 window enforces local structural consistency; isolated dark pixels or small fragmented dark patches—common in speckle-induced lookalikes—are less likely to form stable graph neighborhoods and thus tend to be filtered out during graph convolution.

2.2.1. Sparse Adjacency Matrix

In order to save the memory space occupied by the algorithm and optimize the operation efficiency, before the operation of the graph convolution network, it is necessary to construct a sparse adjacency matrix for the preprocessed original image to provide a suitable graph structure for the graph convolution [29]. For each pixel in the figure, a three-dimensional feature vector X i is constructed respectively, and the set of feature vectors of all pixels constitutes the feature matrix R 3 :
X i = g i , x i , y i T R 3
where g i = I x i , y i 0 , 255 is the gray value of a single pixel before normalization, and x i and y i are the row and column coordinates of the pixel, respectively.
With each pixel i as the center, a local 5 × 5 pixel area is taken as the neighborhood, including 24 adjacent pixels except itself. Then, the Euclidean distance of the feature space between the central pixel X i and each adjacent pixel X j ( d i j ) is calculated:
d i j = x i x j 2 = g i g j 2 + x i x j 2 + y i y j 2
The Gaussian kernel function is used to convert the Euclidean distance into a similar weight w i j in the 5 × 5 neighborhood of each pixel:
w i j = e x p d i j 2 2 σ 2 , x j N i k
where k pixel with the smallest Euclidean distance is selected as its adjacent pixel, which is denoted as set N i k . k is set to 1 here: that is, only 1 nearest neighbor pixel of each pixel is retained. Finally, the adjacency matrix element that meets the nearest neighbor pixel condition is defined as A i j , and its element set is defined as sparse adjacency matrix A :
A i j = A j i = w i j ,   i f   x j N i k   o r   x i N j k 0 ,   o t h e r w i s e ,   A R N × N

2.2.2. Graph Convolutional Network Feature Learning Model

Let H l R N × d l be the node feature matrix of layer l, where d l is the feature dimension of layer l . Based on the sparse adjacency matrix A , the degree matrix D is constructed to normalize the adjacency matrix:
D = diag j = 1 N A i j + ε I
 = D 1 A
where ε = 10 8 is used to maintain numerical stability.
The graph convolution operation is performed, and α = 0.6 is set as the residual connection coefficient, which is selected based on the prior comprehensive consideration of retaining oil film characteristics and model stability. For H 0 = a , 0 , , 0 R N × 16 , only the first column is used as the normalized gray value, and the remaining columns are 0; then, after two layers of graph convolution operations with output dimensions of 16, the ReLU activation function is used to achieve high-order feature separation. The final output matrix H 2 R N × 16 is as follows:
H l + 1 = α H l + 1 α Â H l
ReLU x = m a x 0 , x
It is worth noting that the graph convolution operation in Equation (7) does not involve learnable weight matrices. Each propagation layer performs only neighborhood aggregation, degree normalization, and a residual connection followed by ReLU activation. This design follows the simplified graph convolution (SGC) paradigm established in simplifying graph convolutional networks [30]. The SGC demonstrates that the successive propagation of neighborhood information is the primary contributor to the effectiveness of graph convolutional networks in feature extraction tasks, rather than learned feature transformations. Operating GCN as a fixed parameter-free feature propagation process is well-suited to the single-image scenario, where the limited amount of pixel-level data makes supervised training of learnable weights both statistically challenging and computationally unnecessary. Consequently, the entire first stage is unsupervised, requiring no labeled data.

2.2.3. K-Means

The feature vectors of each pixel output by the graph convolutional network are used as input, and the K-Means algorithm is used for clustering. K-Means is a classical unsupervised learning method that partitions the dataset into K clusters by minimizing the within-cluster sum of squared distances. In this study, K-Means is chosen for the following reasons:
(1)
Efficiency: K-Means scales linearly with the number of pixels. For a 1024 × 1024 image, this translates to over one million nodes, making quadratic-complexity methods like spectral clustering and agglomerative clustering computationally prohibitive.
(2)
Compatibility with GCN output features: The GCN produces compact, well-separated 16-dimensional embeddings via residual graph convolution and ReLU activation, which naturally suit centroid-based clustering.
(3)
Deterministic and reproducible: With fixed initialization, K-Means yields consistent cluster assignments, which are important for fair experimental comparison.
(4)
The number of clusters is physically meaningful: K = 3 corresponds to the three physically distinct categories in the image—strong signal, weak signal, and background.
Thus, K-Means strikes an idealized balance between segmentation effectiveness, computational efficiency, and interpretability for this stage. The K-Means algorithm divides the feature vector into three clusters, assigns a cluster label to each vector, and outputs the segmentation result S i :
S i = m i n C 1 , , C k k = 1 K h i C k H i μ k 2
where K = 3 is the number of clusters, μ k is the center of the k th class, and H i is the graph convolution feature of pixel i . Then, the binarization and screening methods are used for morphological post-processing, and one of them is selected as the target region candidate. Filling the image holes and performing the area opening operation, the connected area with an area of less than 50 is removed, and the final region of interest is extracted, and the ROI image I R O I is the output.

2.3. Improved GJO

After the initial extraction of ROI, this stage addresses a three-class threshold optimization problem: given the ROI pixel set obtained, find the optimal threshold vector T that partitions the ROI pixels into three categories—strong-signal class, weak-signal class, and background class—such that the modified Otsu-based fitness function is defined. The lowest-threshold region is subsequently identified as the suspected oil film area.
To solve this optimization task, we propose an Improved Golden Jackal Optimization (IGJO) algorithm. The IGJO introduces four key improvements over the standard GJO metaheuristic:
(1)
Time-varying low-threshold reward function: A reward term that is positively correlated with the iteration count is added to the Otsu between-class variance objective, which progressively biases the search toward lower gray-level intervals where oil films are physically expected.
(2)
Low-threshold bias initialization strategy: A total of 70% of the population individuals are initialized in the low gray-value range, whereas the standard GJO uses uniform random initialization across the full range. This injects domain-specific prior knowledge into the metaheuristic search.
(3)
Dynamic energy-adaptive exploration–exploitation balance: The original GJO’s linear energy decay is replaced with a nonlinear attenuation mechanism, maintaining stronger exploration capability in early iterations and enhancing local exploitation in later iterations, thereby preventing premature convergence.
(4)
Late forced exploration mechanism: When the algorithm detects stagnation, a randomized injection mechanism with 20% probability is triggered to help the population escape local optima—a mechanism absent in the standard GJO.
Collectively, these improvements embed the physical characteristics of oil film images into the optimization framework of the metaheuristic optimizer, aligning the search behavior with domain-specific prior knowledge while preserving GJO’s population-based global search capability. The specific design of each improvement is detailed in the following subsections. The specific process is as follows (see Figure 4).

2.3.1. Statistics Definition

Define the non-background pixel set P , N non 255 = P as the number of non-background pixels. The improved GJO goal is to find the optimal threshold vector T = T 1 , T 2 T , where T 1 < T 2 : that is, two different thresholds for the three classifications of the extracted ROI. For any pixel value p P , the fitness function is designed. Firstly, all kinds of category statistics are defined, including the number of category pixels n k , the category weight w k , the category mean μ k and the overall mean μ T :
n k = p P Class p = k , k = 1 , 2 , 3
w k = n k N non 255
μ k = 1 n k p Class k p
μ T = k = 1 3 w k μ k

2.3.2. Fitness Function

Based on the Otsu criterion, the standard Otsu inter-class variance algorithm is used as the basis of the fitness function [31]. The set balance penalty term P balance is used as the inter-class variance coefficient to prevent Otsu from over-biasing during iteration, resulting in extreme segmentation and improving the practicality and robustness of the segmentation results:
σ B 2 T = k = 1 3 w k μ k μ T 2
P balance T = e x p 2 w 1 w 2
A time-varying low-threshold reward function is added to the fitness function, and a reward term positively related to the number of iterations is introduced into the standard Otsu between-class variance target. Definition R low T and a time-varying reward coefficient λ are as follows:
R low T = λ 1 T 1 + T 2 2 U 2
λ t = 0.1 + 0.4 · t t m a x , t 0 , t m a x
where t is the current number of iterations, and t m a x is the maximum number of iterations, and U = m a x P .
The final comprehensive fitness function formula is as follows:
F T = σ B 2 T · P balance T + R low T

2.3.3. Initialization Strategy

Using the low threshold bias initialization strategy, 70% of the population individuals are focused on the low gray range for searching at the beginning of the algorithm. The remaining 30% of the individuals maintain a full range of search. The parameters of the algorithm are set as follows: let the population size be N p and set the lower bound L and the upper bound U of the search space:
L = m i n p
U = m i n m a x p , μ p + 1.5 σ p
where μ p = mean p , σ p = std p .
For the a-th pixel, the threshold vector T a = T a 1 , T a 2 T is initialized with 70% probability in the lower range and 30% probability in the full range, and P I is the selection probability in the initialization stage. This is shown in Formula (21):
T a 1 ~ U L , L + 0.5 U L ,   T a 2 ~ U T a 1 , L + 0.5 U L ,   P I 1 = 0.7 T a 1 ~ U L , U ,   T a 2 ~ U T a 1 , U ,   P I 2 = 0.3
For each threshold value T a b of each pixel, the escape energy E 1 and the low threshold bias factor β a b are calculated as the basis of iterative strategy separation:
E 1 = E t × 2 r B 1 ,   r B ~ U 0 , 1
E t = 1.2 × 1 t t m a x 0.5
β a b = 0.2 × 1 T a b U

2.3.4. Iteration Strategy

The iterative process is divided into the exploration stage, development stage and later forced development stage, and different iterative strategies are adopted in different stages.
The exploration stage E 1 1 : Low threshold exploration is performed with probability 0.6 + β a b , otherwise standard exploration is performed; P E selects probability for the exploration stage, r E ~ U 0 , 1 :
T a b t + 1 = T b male E 1 × levy d × T b male T a b t λ t × r E × 3 ,   P E 1 = 0.6 + β a b T b female E 1 × levy d × T b female T a b t , P E 2 = 0.4 β a b
Levy d = 0.01 × μ l v l 1 / d , μ l ~ U 0 , σ l 2 , v l ~ U 0 , 1
σ l = Γ 1 + β l sin π β l 2 Γ 1 + β l 2 β l · 2 β l 1 2 1 / β l
where β l = 1.5 .
Development phase E 1 < 1 : Perform low threshold development with probability 0.5 + β a b , otherwise perform standard development. P D is the probability of the development phase, r D ~ U 0 , 1 :
T a b t + 1 = T b male E 1 × T b male T a b t λ t × r D × 2 ,   P D = 0.5 + β a b T b female E 1 × T b female T a b t , P D = 0.5 β a b
Late forced exploration stage ( t > 0.6 t m a x ): When the algorithm may fall into stagnation, the late forced exploration mechanism actively injects randomness and tries to break through the local optimum. It performs with a 20% probability for the global:
T a b temp = T a b male r C × 3 ,   r C ~ U 0 , 1 2 , P C = 0.2
If F T a b temp > F T a b male , T a b temp = T a b male .
The boundary constraint (30) and the sorting constraint (31) are applied to the threshold value to ensure that the threshold falls within the specified upper and lower bounds and that T a 1 < T a 2 :
T a b t + 1 = m a x L , m i n U , T a b t + 1
T a t + 1 = sort T a t + 1
According to the optimal threshold calculated by the improved GJO iteration, three ROIs are classified. The lowest threshold region is taken as the suspected oil film region, and the morphological post-processing and coordinate system transformation are carried out to facilitate the re-feedback to the polar coordinate system of the radar for suspected oil film identification.

2.4. Evaluation Index

The performance of the proposed GCN-IGJO algorithm is evaluated by establishing evaluation indicators, including accuracy, recall rate, precision, F1 score, Dice coefficient and IoU. The calculation of the index is based on four basic statistical units: True Positive (TP), False Positive (FP), True Negative (TN) and False Negative (FN). TP refers to the number of pixels correctly identified as oil film in the segmentation result, which indicates that the effect of the algorithm correctly locates the target area. FP refers to determining the actual background as the number of pixels of the oil film in the result, indicating the false recognition evaluation of the algorithm; TN refers to the number of background pixels correctly identified, indicating the ability of the algorithm to avoid the sea background; and FN is the number of pixels that identify the oil film as the background, which is related to the detection failure rate of the algorithm. The strict expert interpretation of the actual background and oil film pixel area can comprehensively and accurately compare the segmentation performance of the proposed algorithm on various indicators. In addition, we add total floating-point operations (FLOPs) to analyze the computational complexity.
Figure 5 presents the expert-annotated ground truth used for quantitative evaluation in this study. The annotation was performed by a remote sensing image interpretation specialist with extensive experience in marine radar oil spill identification. The ground truth delineates the oil-contaminated area based on the joint consideration of the low gray-level contrast between the oil film and the surrounding sea surface in the preprocessed radar image, the spatial continuity and morphological characteristics of the dark patch, and cross-referencing with the known spill location and extent documented during the Dalian Xingang port incident. This ground truth serves as the binary reference mask against which all quantitative metrics reported in Section 4 are computed. The annotation procedure was independently verified by a second interpreter to minimize subjective bias and ensure annotation reliability. The oil film area is marked in blue.

3. Results

The operating environment of this study is a hardware environment equipped with Intel Core i5-13420H (2.1 GHz, 4.6 GHz) and 16 GB DDR4 memory. The software environment used in the experiment is MATLAB R2024b. The relevant parameters of the IGJO step in the algorithm are set as follows. The maximum number of iterations t m a x is set to 60 and the population size is set to 25, which balances the convergence speed and search ability. The problem dimension d = 2 corresponds to two segmentation thresholds that need to be optimized. The lower limit of the search range L is the minimum gray value of the non-background pixel, and the upper limit U is compressed to within 1.5 times the standard deviation of the mean value to guide the algorithm to focus on the low gray area that is more likely to contain the oil film. The dynamic energy coefficient E controls the switching between global exploration and local development, and its value decays with iteration. The low threshold bias coefficient β a b increases linearly from 0.1 to 0.5, which makes the algorithm more inclined to find the low threshold in the later stage of iteration.
The oil film segmentation results of the GCN-IGJO algorithm are shown in Figure 6.
Figure 7 illustrates the intermediate neighborhood aggregation signal within the GCN feature propagation process. The color map ranges from blue to red, indicating low to high aggregated gray values; blue regions correspond to low-intensity neighborhoods (oil film candidates), while yellow-to-red regions represent high-intensity neighborhoods (sea clutter or strong scatterers). Compared to the original gray-level image, this aggregated signal suppresses isolated noise through the spatial constraints of the graph structure while preserving the overall spatial continuity of oil film regions. After fusion with the original signal via a residual connection, this intermediate representation serves as the first channel of the 16-dimensional feature vector output by the GCN, providing input for the subsequent K-Means clustering step. The final oil film extraction results are shown in Figure 8.

4. Discussion

4.1. Comparison of Performance Evaluation of Different Methods

Based on the evaluation index system established in Section 2.4, this section compares the proposed method with several partial differentiation methods. It mainly includes the effect comparison between IGJO and the original GJO method, the necessity comparison of the GCN method, and the effect comparison of two classical algorithms: Otsu and the local adaptive method in the threshold segmentation step.
Firstly, the GCN-IGJO method is qualitatively analyzed. As shown in the table, the accuracy of up to 96.7% means that the model is more reliable, with almost no false positives, and the segmentation results are accurate. At the same time, the recall rate of 82.8% indicates that the algorithm can effectively locate most of the target areas. This combination of high precision and good recall rate makes its comprehensive performance index (F1/Dice: 89.2%) and segmentation coincidence degree (IoU: 80.5%) reach a high level. In general, the recognition ability of the algorithm is comprehensive and reliable, and has practical value. It is worth mentioning that the traditional GJO algorithm is the same as the threshold result directly segmented by Otsu. Therefore, the comparison results of the two methods are no longer discussed separately in the following.
To further evaluate the practical viability of the proposed method, we estimate the theoretical FLOPs of each pipeline. As summarized in Table 1, the proposed GCN-IGJO requires approximately 1.98 GFLOPs, which is comparable to GCN-GJO and higher than the GCN-Otsu/GCN-LA pipelines. However, this increase in computation is justified by two factors. First, the improvement in comprehensive metrics relative to GCN-Otsu is substantial, representing a fundamentally better segmentation quality rather than a marginal gain. Second, the primary computational bottleneck is the K-Means clustering on the 16-dimensional GCN features, which is an inherent cost of the high-dimensional feature space that enables the precision gain. Future work may explore dimensionality reduction or approximate nearest-neighbor search to further reduce this overhead.

4.2. Comparison of GJO/Otsu Algorithms

The goal of the IGJO algorithm is to find two thresholds more accurately after extracting ROI. Three categories are divided: strong signal, weak signal and background, to achieve accurate identification of the oil film region. Compared with the original golden jackal optimization algorithm, the improved GJO algorithm proposed in this paper embeds the image feature depth of the oil film region into the optimization framework by optimizing the core mechanism of the original algorithm and improves the performance of the oil film image segmentation task. A time-varying low-threshold reward function is added to the fitness function, and a reward term positively related to the number of iterations is introduced. During the iteration process, the algorithm is guided to continuously move towards the solution space of the lower threshold to ensure that the segmentation result is consistent with the low gray characteristics of the oil film. The initialization strategy of low threshold bias is used to change the uniform random initialization method of the original algorithm in the solution space.
In the early stage of the algorithm, most of the population individuals are focused on the exploration of the low gray level interval, and the speed of the algorithm approaching the optimal solution is improved. The remaining individuals maintain the full range search and maintain the stability of the algorithm. The dynamic energy attenuation mechanism is added to the update mechanism, and the nonlinear attenuation is used to replace the linear attenuation, so as to maintain stronger exploration ability in the early stage of optimization and avoid premature results. In the later stage, local development is enhanced to balance the breadth and accuracy of search. In the iterative exploration stage, the symmetric Levy flight strategy with low threshold bias is combined to ensure the diversity of random jumps and the directional convergence of the target area in the global exploration. The oil film segmentation results and evaluation of the GJO method are as follows. Figure 9 shows the oil film mask of GCN-GJO/GCN-Otsu method. The orange boxes represent the main differences from the proposed method.
As can be observed from the segmentation results, both GJO and Otsu do not incorporate a lower-bound constraint on the threshold optimization objective [32]. Consequently, the iterative search tends to converge toward higher threshold values, resulting in coarser oil film segmentation and an elevated false alarm rate. The qualitative evaluation index shows that the performance of the improved GJO algorithm is better than that of the original algorithm: the accuracy rate is improved to a certain extent, indicating that the improved algorithm further refines the oil film recognition effect. The accuracy is improved from 66.6% to 98.6%, which shows that the improved algorithm greatly improves the credibility and can effectively distinguish the background and oil film. The recall rate is increased from 76.4% to 82.8%, which reflects the improvement of the algorithm’s ability to avoid underreporting and find out the real oil film area. The leading of F1 score and Dice coefficient shows that GCN-IGJO has achieved a good balance between accuracy and recall rate, and the comprehensive segmentation efficiency is excellent. The final IoU parameter shows that the oil film area segmented by the improved GJO algorithm is more consistent with the real area.
Interclass dispersion and convergence behavior, as shown in Figure 10, illustrates the fundamental behavioral difference between the two optimizers. The standard GJO, initialized uniformly across the full gray-value range, converges to thresholds separated by approximately 40–45 gray levels, with T2 settling at a high gray value near 115. This wide separation reflects the unconstrained maximization of between-class variance, which tends to push the upper threshold toward higher gray values to achieve maximal statistical separation. In contrast, IGJO’s low-threshold bias initialization directs 70% of the population to the lower gray-level interval, and the threshold evolution shows that both T1 and T2 converge rapidly to a tightly bounded low-value region. Quantitatively, the GJO achieves σ2 = 969.06 at T1 = 69.7 and T2 = 114.9, whereas IGJO’s Otsu component reaches σ2 = 861.77 at T1 = 58.0 and T2 = 82.1. The 11% reduction in σ2 is a deliberate design outcome. The low-threshold reward shifts the search toward the physically meaningful low-gray-value region where oil films reside, while the balance penalty term β maintains class weight equilibrium. The convergence curves demonstrate that both algorithms stabilize within the first 10 iterations, with IGJO exhibiting comparable convergence speed due to its domain-informed initialization. The segmentation metrics validate this trade-off: GCN-IGJO achieves 96.7% precision and 80.5% IoU, substantially exceeding GCN-GJO’s 66.6% precision and 55.3% IoU.

4.3. Effect Analysis of GCN

In this study, the graph convolution network regards each image pixel as a node in the graph, converts the original single gray value in the image into higher-order semantic features through graph convolution operation, and generates a 16-dimensional feature vector for each pixel. The vector synthesizes the gray information and spatial context relationship of each pixel value, and can better represent the category (such as oil film, sea surface, etc.) than the original pixel value. In order to measure the actual effect of the GCN, we will not use the GCN, but will directly use K-Means clustering to extract ROI. The IGJO segmentation method and the method of this study were compared; the experimental results are as follows. Figure 11 shows the oil film mask of K-IGJO method. The orange boxes represent the main differences from the proposed method.
It can be clearly seen from the segmentation results that, due to the loss of the feature extraction function of GCN, the results of direct K-Means clustering and segmentation confuse the oil film and some sea background images with similar features, and the oil film details cannot be accurately identified and the background cannot be separated correctly. It can also be seen from the comparison table that the algorithm processed by GCN is better than that without GCN in all aspects, indicating that GCN plays an indispensable role in the composition of our method.

4.4. Comparison with Local Adaptive Threshold Method

The local adaptive threshold segmentation method is a local threshold segmentation method. It calculates an independent threshold for each pixel or its neighborhood in the image, so as to deal with complex situations such as uneven illumination and background changes. Its image segmentation results are shown in Figure 12.
The GCN-LA method combined with the local adaptive threshold (accuracy 97.9%) went to the extreme, and its recall rate was only 11.2%, indicating that it was seriously missed and could hardly identify the target object completely, which was in line with the actual segmentation results. The GCN-IGJO method achieves a balance between accuracy and precision and obtains a considerable recall rate while maintaining high precision. This conclusion is consistent with the clear, coherent and concentrated visual performance of the segmented region output by GCN-IGJO, which proves that it is superior to the traditional threshold segmentation strategy in complex scenes.

4.5. Comparison of Deep Learning Methods

To further benchmark the proposed method against recent deep learning architectures, we additionally compare with three representative segmentation models: (i) U-Net with ResNet34 backbone, a widely adopted U-Net variant with an ImageNet-pretrained encoder [33,34]; (ii) Attention U-Net, which incorporates attention gates to focus on salient oil film regions during skip connections [35]; and (iii) SegFormer-B0, a lightweight hierarchical Transformer-based segmentation model with an all-MLP decoder [36]. These models were selected to cover both the U-Net variant family and the Transformer-based family, while keeping computational complexity manageable for the present single-image case study.
Due to the single-image constraint, training was conducted on 256 × 256 patches extracted from the preprocessed 512 × 2048 radar image with a stride of 128, yielding 16 patches. All patches were used for training with random horizontal/vertical flips for data augmentation. A weighted binary cross-entropy loss handled the severe class imbalance (oil film ≈ 2.07% of pixels). All models were trained for 150 epochs using the AdamW optimizer (initial learning rate 5 × 10−4, cosine annealing schedule). Full-image predictions were obtained via sliding-window inference (stride 128, overlapping predictions averaged). The detailed experimental results are shown in Table 2 and Figure 13.
As shown in Table 2, the proposed GCN-IGJO method achieves the highest F1 and IoU, outperforming all three deep learning baselines. Attention U-Net ranks second with an F1 of 86.0% and an IoU of 75.4%, demonstrating that attention gating is beneficial for oil film boundary preservation. U-Net (ResNet34) yields a precision of 92.6% but its recall drops to 72.6%, indicating that the plain skip connections miss a substantial portion of the oil slick. SegFormer-B0 attains the highest precision among the deep learning baselines but has the lowest recall, producing an F1 of only 66.8%. This conservative behavior suggests that Transformer-based models are more sensitive to training data scarcity under the single-image setting. Notably, GCN-IGJO’s precision surpasses all deep learning methods, confirming that the IGJO’s domain-informed low-threshold bias strategy is highly effective at suppressing false positives. Its recall is also substantially higher than that of the best deep learning competitor, Attention U-Net, demonstrating that combining GCN-based ROI extraction with prior-guided threshold optimization provides a more complete and reliable oil film segmentation than patch-based supervised learning when labeled data are extremely limited. These results reinforce the practical value of the proposed training-free optimization paradigm for remote sensing applications where annotated datasets are scarce.
It should be noted that the above comparison was conducted under highly restricted data conditions. Only one labeled image was available. This is quite common in remote sensing operations, as expert labeling is costly and oil spill incidents are rare, giving deep learning methods a natural disadvantage. The untrained nature of GCN-IGJO provides practical advantages in low-labeled scenarios.
Therefore, these results should be interpreted as follows: in the low-labeled scenario, the proposed GCN-IGJO method in this paper outperforms the deep learning baseline based on patch training in terms of segmentation performance. This finding supports the practical value of the no-training optimization paradigm in remote sensing applications with scarce labeled data. When labeled data is abundant, a complete comparison of deep learning methods based on multi-image datasets will be a valuable direction for future work and may narrow the performance gap; however, the current results demonstrate that in the case of limited data acquisition and labeling resources, GCN-IGJO provides a competitive alternative solution.

4.6. Effect Analysis of Parameters

4.6.1. K-Neighborhood Size

To evaluate the influence of the sparse adjacency graph neighborhood size on oil film segmentation performance, this section further conducts a systematic comparative experiment on three alternative neighborhood configurations: 3 × 3, 7 × 7, and 9 × 9. In all experiments, the GCN architecture, K-Means clustering, IGJO, the number of nearest neighbors, and all other parameters are kept fixed, and the same expert-annotated ground truth image is used as the reference. The proposed GCN-IGJO method with the default 5 × 5 neighborhood serves as the baseline for comparison. The experimental results are shown in Table 3. Figure 14 shows the output image of different neighborhood sizes.
The overall performance ranking across the four configurations is 5 × 5 > 3 × 3 > 9 × 9 > 7 × 7. After the 5 × 5 configuration adopted in this study, the 3 × 3 window achieves the best comprehensive performance; however, its recall is comparatively lower, which can be attributed to the smaller receptive field of the 3 × 3 neighborhood, limiting the range of feature propagation. The 7 × 7 window yields the highest recall, but its precision drops markedly to 54.8%, as the enlarged neighborhood tends to aggregate features from pixels belonging to different semantic categories, leading to the misclassification of background pixels as oil film. Considering both segmentation stability and accuracy, the 5 × 5 neighborhood is selected as the default configuration for this study. The 5 × 5 neighborhood provides sufficiently rich local context for the GCN while avoiding excessive cross-category feature contamination that arises from larger windows.

4.6.2. Threshold Upper Bound

The choice of upper bound U = μ p + 1.5 σ p is empirically validated by an ablation study across four settings: 1.0σ, 1.5σ, 2.0σ, and 3.0σ. For each setting, the complete segmentation pipeline was executed. The results are presented in Table 4. Figure 15 shows the output result of different upper bound settings.
The 1.5σ setting achieves the highest values across five of the six metrics. While 2.0σ yields a marginally higher recall of 86.8%, its precision drops to 64.0% and its IoU to 58.0%—a degradation of 32.7 and 22.5 percentage points, respectively, relative to 1.5σ. The 1.0σ and 3.0σ settings occupy intermediate positions but remain substantially inferior to 1.5σ in the comprehensive metrics.
Three conclusions are drawn. First, 1.5σ is the only setting that simultaneously breaks the 90% barrier on Accuracy and Precision, indicating that it identifies the optimal trade-off between the low-threshold bias and the sufficient search space. Second, larger bounds degrade overall performance because the optimizer wastes search resources in the high-gray-value region that contains no oil film pixels, effectively diluting the low-threshold bias. Third, tighter bounds restrict the solution space excessively and prevent the optimizer from finding the globally optimal threshold pair. Therefore, the compression to μ + 1.5σ is not merely a heuristic convenience but a rigorously validated design choice that maximizes the end-to-end segmentation quality.

4.7. Discussion on Computational Efficiency

The computational cost of the proposed GCN-IGJO pipeline is evaluated on the described hardware platform. In 50 runs, the average processing time for a single 1024 × 1024 image is approximately 51.9 s. The detailed breakdown is as follows: sparse adjacency matrix construction 45.4 s; two-layer GCN feature propagation 0.1 s; K-Means clustering and ROI binary extraction 1.5 s; IGJO iterative optimization 4.4 s; and morphological post-processing 0.4 s.
The adjacency matrix construction dominates the total runtime, accounting for over 87% of the computational cost. This is primarily due to the pixel-wise iterative loop over all 1048,576 nodes, each requiring pairwise Euclidean distance computation within the local 5 × 5 neighborhood and sparse matrix indexing. The GCN inference itself is negligible (0.1 s), confirming that the untrained, fixed-weight design is computationally lightweight. The IGJO, despite its iterative nature, completes in 4.4 s because the fitness evaluation is restricted to the non-background ROI pixels rather than the full image.
While the current runtime does not meet the real-time requirements of shipboard X-band radar systems operating at 24–45 RPM, it is acceptable for near-real-time monitoring and post-event forensic analysis. Adjacency matrix construction which is the primary bottleneck can be substantially accelerated through: (a) vectorized implementation or parallel computation; (b) pre-computed spatial distance look-up tables to avoid redundant Euclidean distance calculations; or (c) migrating the graph construction to a compiled language. These optimizations are identified as promising directions for future work.

4.8. Discussion on Wrong Segmentation

Although the proposed method achieves high overall accuracy on the final segmentation, a closer inspection of the prediction-versus-ground-truth overlay reveals two typical error patterns. The first pattern is false negatives concentrated along oil film boundaries, which is shown in Figure 16a. In these transitional zones, the backscatter contrast between the slick and the surrounding sea surface is weak, and the gray-value signature gradually merges with the background. Because the GCN-IGJO pipeline applies a low-threshold prior and a 5 × 5 spatial neighborhood to suppress speckle, it tends to be conservative at faint edges, causing thin boundary pixels to be omitted. The second pattern is isolated false positives, appearing as small red specks inside or adjacent to correctly detected oil regions. As shown in Figure 16b, these errors are mainly induced by strong sea-clutter fluctuations or by local low-backscatter background patches that have a gray-level distribution that is almost identical to oil films.
The effectiveness of the low-threshold bias strategy depends strongly on the sea state and the specific lookalike scenario. Under moderate sea conditions, wind-driven surface roughness ensures that most lookalike dark areas produce backscatter levels that are still distinguishable from genuine oil slicks after GCN feature enhancement, making the low-threshold bias both effective and robust. Under very calm seas, however, large regions of the sea surface exhibit near-zero backscatter due to insufficient wind stress; these regions can be nearly indistinguishable from oil films in both raw intensity and spatial texture, leading to an increase in false positives. Conversely, under high sea states, the breaking waves and foam generate strong clutter echoes that elevate the local average gray value, reducing contrast and potentially causing the low-threshold prior to miss thinner oil patches, which manifests as elevated false negatives. Additionally, man-made structures such as ship wakes and radar shadows produce elongated low-intensity features that sometimes survive the GCN filtering, contributing a small number of persistent false positives.

4.9. Multi-Scene Verification

To further evaluate the qualitative generalization capability of the proposed GCN-IGJO method across different marine oil spill scenes, an image under different oil spill scenarios is selected for the same pipeline segmentation. The original image after preprocessing is shown in Figure 17a. Compared with the image used in the previous article, this scene exhibits a different oil film morphology: the dark slick regions are more fragmented and distributed as elongated bands near the lower part of the image, with relatively weak backscatter contrast against the surrounding sea clutter.
The segmentation mask produced by the GCN-IGJO pipeline is presented in Figure 17b, and the overlay of the detected oil film regions on the original image is shown in Figure 17c. Qualitatively, the algorithm identifies the main low-backscatter slick bands and their fragmented extensions, producing spatially coherent segmentation results that are consistent with the visually interpreted oil film extent. The red contours in the overlay image indicate that the segmented regions align well with the dark areas in the original image, whereas most of the bright sea-clutter background is excluded.
No expert-annotated ground truth is available for this additional scene, so quantitative evaluation is not performed here. The purpose of this validation is to demonstrate the method’s adaptability to different oil film morphologies and imaging conditions rather than to re-evaluate the absolute accuracy metrics. From an experimental-needs perspective, the result supports the cross-scene transferability of the GCN-IGJO framework and confirms that its GCN-based feature extraction and IGJO low-threshold bias strategy remain effective when the oil slick presents a more fragmented and elongated spatial pattern.
It should be noted that the above verification is based solely on two X-band radar scenarios of the same oil spill event, and both were collected under similar sea conditions. A rigorous generalization ability assessment requires systematic evaluation under various sea conditions, oil types, and radar configurations. Obtaining such a diverse dataset of X-band radar oil spill data with real annotations remains a significant challenge in this field, as real oil spill events occur infrequently and monitoring activities are typically event-driven rather than systematic. Nevertheless, the consistent segmentation behavior observed in additional scenarios provides preliminary evidence for the transferability of the method, and the oil film morphology is significantly different from that in the main scenario. The construction of multi-condition datasets is listed as the top priority for future work.

4.10. Comparison of Feature Processing Methods for GCN Output

After the GCN module maps the input image to a 16-dimensional feature space, it needs to convert these high-dimensional features into a binary ROI mask. In this section, starting from the morphological characteristics of the binary mask image, we qualitatively compare the performance of three clustering methods, namely K-Means, Gaussian Mixture Model (GMM), and DBSCAN, in this step.
Figure 18 shows the binary ROI masks generated after clustering the output features of the GCN using three different methods. The white areas represent the candidate regions for oil films.
The ROI mask of K-Means (Figure 18a) appears as a continuous and dense block covering the oil film area in the lower half of the image, with a spatial distribution highly consistent with the expert annotation. The mask boundary is clear, the interior is complete, there are almost no scattered isolated noise points, and the overall continuity of the oil film strips is retained. The coverage range of the dark area is slightly larger than the true value, which is reasonable during the ROI extraction stage, as the subsequent IGJO threshold optimization can further filter out the truly oil film pixels within the ROI.
The ROI mask of GMM (Figure 18b) significantly reduces the mask area compared to that of K-Means (about 55% reduction) and shows obvious contraction at the mid-frame and edge areas. Specifically, the dark areas around the oil film region are excluded, but at the same time, the thin layer of the oil film region is also truncated, resulting in the disruption of the spatial continuity of the mask. This contraction effect is particularly evident at the bottom of the image: the fine branches of some oil film strips are fragmented or even completely disappear in the GMM mask.
The ROI mask of DBSCAN (Figure 18c) almost covers the entire image, losing the filtering function of ROI extraction. DBSCAN connects the high-density background areas with large areas of low-density dark regions, resulting in the mask being almost entirely marked as potential oil film areas from top to bottom. This mask does not effectively perform selective discretization on the output features of GCN and cannot be used in the subsequent threshold optimization process.

4.11. Comparison of Mainstream Optimizers

To further benchmark the proposed IGJO against mainstream population-based optimizers, we additionally compare it with Particle Swarm Optimization (PSO) and Genetic Algorithm (GA) under identical experimental conditions. All methods share the same population size, maximum iterations, and the standard Otsu inter-class variance as the fitness function, operating on the full gray-value search range of the ROI. The standard GJO is included as a baseline. The results are summarized in Table 5 and Figure 19.
Two key observations can be drawn from the table. First, the three optimized versions based on the standard Otsu method have basically the same segmentation indicators, which indicates that the specific selection of meta-heuristic optimization algorithms has negligible impact on the quality of the solution on the single-peak and well-characterized fitness function of Otsu’s inter-class variance. Second, IGJO significantly deviates from the above trend, finding distinct thresholds and achieving a significant improvement in segmentation performance. This confirms that the performance improvement mainly comes from the effective suppression of false positives, which is encoded in the search behavior of the IGJO fitness function with the low threshold bias term and the class balance penalty term. The above results support two conclusions: the superiority of GCN-IGJO over GCN-GJO should be attributed to the domain knowledge-driven fitness function design, rather than any inherent advantages of the GJO search mechanism over other optimization algorithms. Embedding physical priors into the optimization objective is more significant in improving the final segmentation quality than selecting or tuning the optimization algorithm itself; the equivalence of GJO/PSO/GA on the standard Otsu function further demonstrates that in this task, the selection of the fitness function completely dominates the impact of the optimizer selection on the segmentation quality.

5. Conclusions

Aiming at the key technical requirement of accurate identification of marine oil spill oil film based on X-band radar images, this study proposes and verifies an oil film segmentation method combining graph convolution network and improved golden jackal optimization algorithm (GCN-IGJO). In this study, a two-stage processing framework is constructed to effectively deal with the problem of blurred edges and low contrast with the background of oil film targets in radar images. In the first stage, the graph convolutional network (GCN) is used to perform high-order feature learning on the gray and spatial position information of image pixels, which lays a solid foundation for subsequent segmentation. The improved golden jackal optimization algorithm proposed in the second stage achieves the fine segmentation of the oil film region. Experimental results demonstrate that the GCN-IGJO method achieves superior performance across multiple evaluation metrics. These results confirm that the integration of graph deep learning with an improved meta-heuristic optimization algorithm provides an effective and innovative technical pathway for weak-target oil film segmentation under complex sea conditions.
The proposed method has the following main limitations:
(1)
The limitations of cross-scenario verification. The experimental evaluation is limited to a single oil spill event, with restrictions on the types of oil, the configuration of the radar system, and the range of sea conditions. Although Section 4.9 provides qualitative evidence of adaptability to different oil film forms, no systematic quantitative verification across various sea conditions, oil types, weathering stages, and radar hardware configurations has been conducted. This limitation stems from the scarcity of publicly available, labeled X-band radar oil spill datasets. We will prioritize this in subsequent data collection efforts.
(2)
The low-threshold prior is sea-state dependent. Under calm seas, extensive low-backscatter sea areas become nearly indistinguishable from genuine oil slicks in both intensity and spatial texture, leading to increased false positives. Under rough seas, breaking waves and foam can suppress thin oil film signals, elevating false negatives. Computational efficiency falls short of real-time requirements.
(3)
Computational efficiency. In the current implementation, the construction of the sparse adjacency matrix accounts for approximately 87% of the total running time, making this pipeline unable to meet the requirements for real-time deployment onboard the ship at the current radar speed. The main bottleneck is the pixel-level iterative loop for calculating pairwise Euclidean distances for the local neighborhood of each node. Solving this problem requires substantial engineering work, including vectorized tensor computations, spatial distance lookup tables, or migration of graph construction to a compiled language. This is beyond the scope of the method’s contribution in this paper. However, the design of the fixed-weight untrained GCN ensures that the feature propagation process itself can be negligible, confirming that the core of the algorithm is efficient, and the bottleneck is purely an implementation issue in the preprocessing stage. The actual real-time deployment will rely on these engineering optimizations, and we have identified this as a clear and well-defined path for future work.
(4)
The comparability with supervised deep learning is limited. To conduct a strict comparison between the GCN-IGJO and the deep learning segmentation model, a multi-image annotated dataset is required to train the latter under real data conditions. Due to the availability of only a single image in this study, the deep learning baseline is trained on a limited number of patches, which limits their representational capabilities. Although extensive data augmentation to some extent alleviates this problem, the reported performance gap may partly reflect data scarcity rather than fundamental algorithmic advantages. Therefore, we emphasize that our conclusions are limited to low annotation scenarios. Future work needs to establish a more comprehensive comparison between unsupervised optimization methods and supervised deep learning on larger annotated datasets.
Future work will address these limitations through multi-condition dataset collection covering varied sea states and oil types, adaptive graph construction to replace the current fixed neighborhood, and lightweight implementation strategies including vectorized distance computation or GPU parallelization to enable real-time shipboard deployment.

Author Contributions

J.X.: Writing—review and editing, Writing—original draft, Methodology, Supervision, Data curation, Conceptualization. X.L.: Writing—original draft, Methodology, Investigation, Formal analysis. Z.F.: Writing—review and editing, Methodology, Investigation, Data curation. M.S.: Writing—review and editing, Supervision, Conceptualization. M.Y.: Writing—review and editing, Methodology, Investigation. Z.G.: Writing—review and editing, Investigation, Data curation. B.C.: Writing—supervision, Data curation, Conceptualization. G.T.: Writing—Supervision, Data curation, Conceptualization. B.L.: Writing—Investigation, Formal analysis, Supervision. H.D.: Writing—review and editing, Methodology, Investigation. S.C.L.: Writing—review and editing, Methodology, Investigation. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Guangdong Basic and Applied Basic Research Foundation, grant number 2025A1515010886; National Natural Science Foundation of China, grant number 52271359; Postgraduate Education Innovation Project of Guangdong Ocean University, grant number 202421; and the Guangdong Provincial Key Laboratory of Intelligent Equipment for South China Sea Marine Ranching, grant number 2023B1212030003.

Data Availability Statement

The raw data supporting the conclusions of this article will be made available by the authors on request.

Acknowledgments

The authors would like to thank the anonymous reviewers for their valuable comments and suggestions.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Liu, J.; Guo, F.; Shi, Y.; Ding, R. Study on the Evolutionary Characteristics and Robustness of the Global Refined Oil Trade Network. Energy 2025, 331, 137032. [Google Scholar] [CrossRef] [Scilit]
  2. Kianfar, E. Current Situation and Future Outlook Petroleum Hydrocarbons in Marine Systems: A Review. Environ. Technol. Innov. 2025, 40, 104572. [Google Scholar] [CrossRef] [Scilit]
  3. Sharma, K.; Shah, G.; Singhal, K.; Soni, V. Comprehensive Insights into the Impact of Oil Pollution on the Environment. Reg. Stud. Mar. Sci. 2024, 74, 103516. [Google Scholar] [CrossRef] [Scilit]
  4. Beyer, J.; Trannum, H.C.; Bakke, T.; Hodson, P.V.; Collier, T.K. Environmental Effects of the Deepwater Horizon Oil Spill: A Review. Mar. Pollut. Bull. 2016, 110, 28–51. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Daly, K.L.; Jackson, G.; Remsen, A.; Kramer, K.; Dave, P.; Goldgof, D.B.; Hall, L. Marine Snow Dynamics in the NE Gulf of Mexico: Particle Abundance, Characteristics, and Impacts on Deepwater Horizon Oil Sedimentation. J. Geophys. Res. Oceans 2026, 131, e2025JC023316. [Google Scholar] [CrossRef] [Scilit]
  6. Soares, M.O.; Teixeira, C.E.P.; Bezerra, L.E.A.; Rossi, S.; Tavares, T.; Cavalcante, R.M. Brazil Oil Spill Response: Time for Coordination. Science 2020, 367, 155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Dong, S.; Feng, J.; Gu, Z.; Yin, K.; Long, Y. A Review of Artificial Intelligence and Remote Sensing for Marine Oil Spill Detection, Classification, and Thickness Estimation. Remote Sens. 2025, 17, 3681. [Google Scholar] [CrossRef] [Scilit]
  8. Marghany, M. Synthetic Aperture Radar Imaging Mechanism for Oil Spills; Gulf Professional Publishing: Cambridge, MA, USA, 2019; pp. 1–322. [Google Scholar] [CrossRef] [Scilit]
  9. Wang, L.; Liu, X.; Cheng, Y. Pre-Processing of Images from Shipborne X-Band Wave Measuring Radar. Int. J. Remote Sens. 2017, 38, 6268–6280. [Google Scholar] [CrossRef] [Scilit]
  10. Garcia-Pineda, O.; Staples, G.; Jones, C.E.; Hu, C.; Holt, B.; Kourafalou, V.; Graettinger, G.; DiPinto, L.; Ramirez, E.; Streett, D.; et al. Classification of Oil Spill by Thicknesses Using Multiple Remote Sensors. Remote Sens. Environ. 2020, 236, 111421. [Google Scholar] [CrossRef] [Scilit]
  11. Senthil Murugan, J.; Parthasarathy, V. AETC: Segmentation and Classification of the Oil Spills from SAR Imagery. Environ. Forensics 2017, 18, 258–271. [Google Scholar] [CrossRef] [Scilit]
  12. Cui, G.; Fan, J.; Zou, Y. Enhanced Unsupervised Domain Adaptation with Iterative Pseudo-Label Refinement for Inter-Event Oil Spill Segmentation in SAR Images. Int. J. Appl. Earth Obs. Geoinf. 2025, 139, 104479. [Google Scholar] [CrossRef] [Scilit]
  13. Sun, C.; Wang, D.; Xu, M.; Wei, S.; Liu, S.; Li, Z. MAF-UFormer: Oil Spill Detection in SAR Images Using Multi-Scale Alignment and Fusion U-Shaped Transformer Network. Reg. Stud. Mar. Sci. 2026, 94, 104773. [Google Scholar] [CrossRef] [Scilit]
  14. Sathya, S.; Murugan, J.S.; Selvam, K.; Bhuvana, S. Multimodal Fusion of Deep Convolutional Neural Networks with Optimization Algorithm for Oil Spill Detection and Lookalikes Classification Using SAR Images. J. Electr. Eng. Technol. 2026, 21, 2053–2072. [Google Scholar] [CrossRef] [Scilit]
  15. Zhao, J.; Tian, Y.; Wen, B.; Tian, Z. Unambiguous Wind Direction Field Extraction Using a Compact Shipborne High-Frequency Radar. IEEE Trans. Geosci. Remote Sens. 2020, 58, 7448–7458. [Google Scholar] [CrossRef] [Scilit]
  16. Zhao, X.; Liu, Z.; Hu, Q.; Sun, J.; Yang, X. Significant Wave Height Estimation and Prediction from Synthetic X-Band Radar Data by Spatio-Temporal Deep Neural Networks. Ocean Eng. 2025, 339, 122061. [Google Scholar] [CrossRef] [Scilit]
  17. Serafino, F.; Bianco, A. X-Band Radar Detection of Small Garbage Islands in Different Sea State Conditions. Remote Sens. 2024, 16, 2101. [Google Scholar] [CrossRef] [Scilit]
  18. Liu, P.; Shao, P.; Zhao, X.; Wang, X.; Chen, P.; Zhu, X.; Li, Y.; Xu, J.; Liu, B. Noise Reduction Based on the Characteristics of X-Band Marine Radar Images for Oil Spill Detection. Remote Sens. Lett. 2026, 17, 190–197. [Google Scholar] [CrossRef] [Scilit]
  19. Chopra, N.; Mohsin Ansari, M. Golden Jackal Optimization: A Novel Nature-Inspired Optimizer for Engineering Applications. Expert Syst. Appl. 2022, 198, 116924. [Google Scholar] [CrossRef] [Scilit]
  20. Khemani, B.; Patil, S.; Kotecha, K.; Tanwar, S. A Review of Graph Neural Networks: Concepts, Architectures, Techniques, Challenges, Datasets, Applications, and Future Directions. J. Big Data 2024, 11, 18. [Google Scholar] [CrossRef] [Scilit]
  21. Scarselli, F.; Gori, M.; Tsoi, A.C.; Hagenbuchner, M.; Monfardini, G. The Graph Neural Network Model. IEEE Trans. Neural Netw. 2009, 20, 61–80. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Ouyang, S.; Li, Y. Combining Deep Semantic Segmentation Network and Graph Convolutional Neural Network for Semantic Segmentation of Remote Sensing Imagery. Remote Sens. 2020, 13, 119. [Google Scholar] [CrossRef] [Scilit]
  23. Zhu, X.; Li, Y.; Feng, H.; Liu, B.; Xu, J. Oil Spill Detection Method Using X-Band Marine Radar Imagery. J. Appl. Remote Sens. 2015, 9, 095985. [Google Scholar] [CrossRef] [Scilit]
  24. Huang, W.; Liu, X.; Gill, E.W. Ocean Wind and Wave Measurements Using X-Band Marine Radar: A Comprehensive Review. Remote Sens. 2017, 9, 1261. [Google Scholar] [CrossRef] [Scilit]
  25. Li, N.; Liu, T.; Li, H. An Improved Adaptive Median Filtering Algorithm for Radar Image Co-Channel Interference Suppression. Sensors 2022, 22, 7573. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Zhang, L.; Zheng, L.; Wen, Y.; Zhang, F.; Bo, F.; Cen, Y. Effective SAR Image Despeckling Using Noise-Guided Transformer and Multi-Scale Feature Fusion. Remote Sens. 2025, 17, 3863. [Google Scholar] [CrossRef] [Scilit]
  27. Liu, J.; Zhou, C.; Chen, P.; Kang, C. An Efficient Contrast Enhancement Method for Remote Sensing Images. IEEE Geosci. Remote Sens. Lett. 2017, 14, 1715–1719. [Google Scholar] [CrossRef] [Scilit]
  28. Zuiderveld, K. Contrast Limited Adaptive Histogram Equalization. Graph. Gems 1994, 474–485. [Google Scholar] [CrossRef] [Scilit]
  29. Paul, S.; Chen, Y. Spectral and Matrix Factorization Methods for Consistent Community Detection in Multi-Layer Networks. Ann. Stat. 2020, 48, 230–250. [Google Scholar] [CrossRef] [Scilit]
  30. Wu, F.; Souza, A.; Zhang, T.; Fifty, C.; Yu, T.; Weinberger, K. Simplifying Graph Convolutional Networks. Int. Conf. Mach. Learn. 2019, 97, 6861–6871. [Google Scholar]
  31. Goh, T.Y.; Basah, S.N.; Yazid, H.; Aziz Safar, M.J.; Ahmad Saad, F.S. Performance Analysis of Image Thresholding: Otsu Technique. Measurement 2018, 114, 298–307. [Google Scholar] [CrossRef] [Scilit]
  32. Otsu, N. Threshold Selection Method from Gray-Level Histograms. IEEE Trans. Syst. Man Cybern. 1979, 9, 62–66. [Google Scholar] [CrossRef] [Scilit]
  33. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 December 2016; IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar] [CrossRef] [Scilit]
  34. Prajapati, K.; Bhavsar, M.; Mahajan, A. GSCAT-UNET: Enhanced U-Net Model with Spatial-Channel Attention Gate and Three-Level Attention for Oil Spill Detection Using SAR Data. Mar. Pollut. Bull. 2025, 212, 117583. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Oktay, O.; Schlemper, J.; Folgoc, L.L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N.Y.; Kainz, B.; et al. Attention U-Net: Learning Where to Look for the Pancreas. arXiv 2018, arXiv:1804.03999. [Google Scholar]
  36. Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
Figure 1. Original marine radar image.
Figure 1. Original marine radar image.
Remotesensing 18 02660 g001
Figure 2. Pre-processing image.
Figure 2. Pre-processing image.
Remotesensing 18 02660 g002
Figure 3. GCN flow chart.
Figure 3. GCN flow chart.
Remotesensing 18 02660 g003
Figure 4. IGJO algorithm flow chart.
Figure 4. IGJO algorithm flow chart.
Remotesensing 18 02660 g004
Figure 5. Visual interpretation image.
Figure 5. Visual interpretation image.
Remotesensing 18 02660 g005
Figure 6. Results image: (a) ROI; (b) incipient segmentation; (c) fine oil film segmentation.
Figure 6. Results image: (a) ROI; (b) incipient segmentation; (c) fine oil film segmentation.
Remotesensing 18 02660 g006
Figure 7. GCN process visualization diagram.
Figure 7. GCN process visualization diagram.
Remotesensing 18 02660 g007
Figure 8. The final result. The predicted oil film area is marked in red. (a) Cartesian system; (b) polar coordination system.
Figure 8. The final result. The predicted oil film area is marked in red. (a) Cartesian system; (b) polar coordination system.
Remotesensing 18 02660 g008
Figure 9. GCN-GJO/GCN-Otsu output image.
Figure 9. GCN-GJO/GCN-Otsu output image.
Remotesensing 18 02660 g009
Figure 10. Convergence process: (a) image of threshold convergence process. (b) Image of Otsu fitness convergence process.
Figure 10. Convergence process: (a) image of threshold convergence process. (b) Image of Otsu fitness convergence process.
Remotesensing 18 02660 g010aRemotesensing 18 02660 g010b
Figure 11. K-IGJO output image.
Figure 11. K-IGJO output image.
Remotesensing 18 02660 g011
Figure 12. Comparison with GCN-LA method.
Figure 12. Comparison with GCN-LA method.
Remotesensing 18 02660 g012
Figure 13. Output image of deep-learning methods. (a) U-Net (ResNet34 backbone); (b) Attention U-Net; (c) SegFormer-B0.
Figure 13. Output image of deep-learning methods. (a) U-Net (ResNet34 backbone); (b) Attention U-Net; (c) SegFormer-B0.
Remotesensing 18 02660 g013
Figure 14. Output image of different neighborhood sizes: (a) 3 × 3 neighborhood; (b) 7 × 7 neighborhood; (c) 9 × 9 neighborhood.
Figure 14. Output image of different neighborhood sizes: (a) 3 × 3 neighborhood; (b) 7 × 7 neighborhood; (c) 9 × 9 neighborhood.
Remotesensing 18 02660 g014
Figure 15. Output result of different upper bound settings. (a) 1.0σ; (b) 2.0σ; (c) 3.0σ.
Figure 15. Output result of different upper bound settings. (a) 1.0σ; (b) 2.0σ; (c) 3.0σ.
Remotesensing 18 02660 g015
Figure 16. Truth value contrast image. Green—accurate segmentation; blue—missing oil film; red—misdetection of oil film. (a) False positive area; (b) False negative area.
Figure 16. Truth value contrast image. Green—accurate segmentation; blue—missing oil film; red—misdetection of oil film. (a) False positive area; (b) False negative area.
Remotesensing 18 02660 g016
Figure 17. Output results of multi-scene verification. (a) Preprocessed image; (b) oil film mask image; (c) superimposed image.
Figure 17. Output results of multi-scene verification. (a) Preprocessed image; (b) oil film mask image; (c) superimposed image.
Remotesensing 18 02660 g017
Figure 18. ROI image of different clustering methods. (a) K-Means; (b) GMM; (c) DBSCAN.
Figure 18. ROI image of different clustering methods. (a) K-Means; (b) GMM; (c) DBSCAN.
Remotesensing 18 02660 g018
Figure 19. Output images of mainstream optimizers. (a) PSO; (b) GA.
Figure 19. Output images of mainstream optimizers. (a) PSO; (b) GA.
Remotesensing 18 02660 g019
Table 1. Comparison of algorithm performance.
Table 1. Comparison of algorithm performance.
GCN-IGJOGCN-GJONO-GCN-IGJOGCN-OtsuGCN-LA
Accuracy99.6%98.6%98.5%98.6%97.9%
Precision96.7%66.6%64.7%66.6%63.1%
Recall82.8%76.4%76.6%76.4%11.2%
F189.2%71.2%70.2%71.2%19%
Dice89.2%71.2%70.2%71.2%19%
IoU80.5%55.3%54.0%55.3%10.5%
FLOPs~1.98 G~1.98 G~0.26 G~0.83 G~0.85 G
Table 2. Comparison of deep-learning methods.
Table 2. Comparison of deep-learning methods.
U-Net (ResNet34 Backbone)Attention U-NetSegFormer-B0
Accuracy99.3%99.5%98.9%
Precision92.6%93.4%95.0%
Recall72.6%79.6%51.5%
F181.4%86.0%66.8%
Dice81.4%86.0%66.8%
IoU68.6%75.4%50.1%
Table 3. Comparison of the results of different neighborhood sizes.
Table 3. Comparison of the results of different neighborhood sizes.
3 × 35 × 5 (GCN-IGJO)7 × 79 × 9
Accuracy99.1%99.6%98.2%98.8%
Precision77.6%96.7%54.8%66.1%
Recall80.5%82.8%87.9%84.3%
F179.0%89.2%67.5%74.1%
Dice79.0%89.2%67.5%74.1%
IoU65.3%80.5%51.0%58.9%
Table 4. Comparison of the results of different upper bound settings.
Table 4. Comparison of the results of different upper bound settings.
1.0σ1.5σ (GCN-IGJO)2.0σ3.0σ
Accuracy92.7%99.6%91.4%93.46%
Precision71.2%96.7%64.0%75.0%
Recall83.2%82.8%86.8%81.8%
F176.7%89.2%73.4%78.3%
Dice76.7%89.2%73.4%78.3%
IoU62.2%80.5%58.0%64.3%
Table 5. Comparison of the results of mainstream optimizers.
Table 5. Comparison of the results of mainstream optimizers.
PSOGA
Accuracy98.8%98.8%
Precision67.5%65.9%
Recall83.9%84.6%
F174.8%74.1%
Dice74.8%74.1%
IoU59.8%58.8%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Xu, J.; Luo, X.; Fu, Z.; Sun, M.; Yan, M.; Guo, Z.; Chen, B.; Tu, G.; Liu, B.; Dong, H.; et al. Marine Oil Film Segmentation Based on GCN-IGJO Method. Remote Sens. 2026, 18, 2660. https://doi.org/10.3390/rs18162660

AMA Style

Xu J, Luo X, Fu Z, Sun M, Yan M, Guo Z, Chen B, Tu G, Liu B, Dong H, et al. Marine Oil Film Segmentation Based on GCN-IGJO Method. Remote Sensing. 2026; 18(16):2660. https://doi.org/10.3390/rs18162660

Chicago/Turabian Style

Xu, Jin, Xingchen Luo, Zhaobin Fu, Mengxin Sun, Minghao Yan, Zekun Guo, Binghui Chen, Gaorui Tu, Bingxin Liu, Haihui Dong, and et al. 2026. "Marine Oil Film Segmentation Based on GCN-IGJO Method" Remote Sensing 18, no. 16: 2660. https://doi.org/10.3390/rs18162660

APA Style

Xu, J., Luo, X., Fu, Z., Sun, M., Yan, M., Guo, Z., Chen, B., Tu, G., Liu, B., Dong, H., & Loon, S. C. (2026). Marine Oil Film Segmentation Based on GCN-IGJO Method. Remote Sensing, 18(16), 2660. https://doi.org/10.3390/rs18162660

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop