Next Article in Journal
A Static and Dynamic Combined Center of Mass Measurement Method Based on Multi-View Vision
Next Article in Special Issue
Lightweight Monocular Depth Estimation with Local Feature Enhancement Modules and Guided Data Augmentation
Previous Article in Journal
Plasmonic Field-Enhanced Raman Sensing Enables Rapid Trace Methanol Detection in Transformer Oil
Previous Article in Special Issue
Semi-Supervised Deep Image Stitching for Moving Elongated Objects
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Rail Light-Strip Abnormality Analysis from Color Inspection Images Using an Improved SegFormer and Geometric Rules

1
Infrastructure Inspection Research Institute, China Academy of Railway Sciences Corporation Limited, Beijing 100081, China
2
INMAI Railway Technology Co., Ltd., Beijing 100081, China
*
Author to whom correspondence should be addressed.
Sensors 2026, 26(16), 5292; https://doi.org/10.3390/s26165292
Submission received: 5 July 2026 / Revised: 6 August 2026 / Accepted: 12 August 2026 / Published: 21 August 2026

Highlights

What are the main findings?
  • Boundary-enhanced SegFormer improves rail-head and light-strip mask recovery in color inspection images.
  • Rail-head-constrained geometric rules identify eccentricity, width mutation and local integrity abnormalities.
What are the implications of the main findings?
  • Boundary quality directly affects light-strip centerline and width measurements.
  • Mask-derived geometric indicators provide interpretable evidence for rail maintenance analysis.

Abstract

Rail light-strip morphology reflects the wheel-rail contact condition. Reliable automatic analysis remains difficult. The strip is narrow and has weak boundaries, while specular reflection, rail-head texture and trackside background interfere with color inspection images. This study proposes a segmentation-guided geometric method for rail light-strip abnormality analysis. An improved SegFormer jointly segments the background, rail-head and light-strip regions. A boundary detail enhancement module refines weak rail-head and light-strip contours. Focal Loss emphasizes minority and hard boundary pixels. The rail-head mask provides the geometric reference for extracting the light-strip centerline, eccentricity, width sequence and connected-component morphology. The predicted masks are ordered using the corrected mileage record. Every 1000 original-resolution rows then form a consecutive 1 m detection unit. When a geometric rule is triggered, the method reports that unit’s 1 m mileage interval together with its eccentricity, width-change or local-integrity measurement. The model achieves 95.67% mean Intersection over Union (mIoU) on 3520 annotated images. It detects 845 of 876 positive units, with 96.46% recall, 89.23% precision and 92.70% F1-score. The resulting records identify abnormal 1 m mileage intervals and report the corresponding eccentricity, width-change, or local-integrity measurements for targeted manual review.

1. Introduction

The rail head is the bearing and guiding interface between the wheel and the track. Repeated wheel-rail contact forms a continuous bright running band on its surface, referred to as the rail light strip in this study. The strip center, width and local contour morphology reflect contact location, grinding condition and local rail-surface state [1,2,3,4]. Rail condition assessment combines track-geometry measurement, ultrasonic or eddy-current inspection and visual observation. These techniques provide complementary information on geometry, internal flaws and surface condition. A comprehensive rail inspection vehicle is an onboard platform that synchronizes these measurements with visible-light images and corrected mileage during line inspection. Its imaging subsystem provides continuous, non-contact color records at operating speed. This allows a suspicious interval to be traced back to the corresponding image sequence. Recent studies have used color cameras, laser profilometers and three-dimensional range imaging to locate running surfaces or detect rail-surface defects [5,6,7]. The visible running band considered here is not equivalent to a crack or another isolated defect. Its condition is expressed mainly by its lateral position, width and contour relative to the rail head. Reliable analysis therefore requires both regions to be extracted in the same geometric reference. The analysis must also remain tied to acquisition mileage because isolated frame numbers are not useful to inspectors.
Current light-strip inspection still relies mainly on manual image review and local intensity or threshold-based extraction [2,3,4]. Manual review is time-consuming for continuous line-scan data. Brightness-based extraction is sensitive to specular reflection, stains, rail-head texture and trackside background. General rail-surface defect detectors are useful for cracks, spalling and other localized defects [1,6,7,8,9,10]. Recent computer-vision studies have also improved railway-component detection under limited data, rail-surface defect detection through feature fusion and running-surface measurement under vehicle-based sensing [5,10,11]. However, their outputs are generally component boxes, defect categories, defect masks or laser-derived surface measurements. They do not directly convert color-image segmentation of the light strip and rail head into mileage-linked eccentricity, width-change and local-integrity indicators. A stable solution must preserve the weak boundaries of both regions and use the rail head as the reference for downstream measurements. Unlike image-level classification, the expected output must answer whether an abnormality occurs, where it is located and which geometric change causes the alarm. The proposed method therefore combines a boundary-enhanced three-class SegFormer with rail-head-constrained geometric rules. Its contribution is the coupling of boundary-sensitive segmentation and explicit geometric abnormality criteria, rather than segmentation alone. This design yields a 1 m mileage interval and the measured condition that triggered each abnormality record.
The remainder of this paper is organized as follows. Section 2 reviews vision-based rail inspection and semantic segmentation methods. Section 3 describes the improved SegFormer, geometric parameters and abnormality rules. Section 4 presents the dataset, implementation details and experimental results. Section 5 discusses practical implications and limitations, and Section 6 concludes the paper.

2. Related Work

Vision-based railway inspection has progressed from handcrafted features and template matching to deep networks capable of handling scale variation and cluttered trackside scenes [1,2,3,4]. Recent rail defect detectors and industrial surface-segmentation methods improve small-defect representation, class imbalance and pixel confidence [1,2,8,9]. However, these methods mainly identify cracks, spalling, scratches or other local defects. They do not explicitly describe the relative geometry between the wheel-rail running band and the rail head.
Semantic segmentation has evolved from convolutional encoder–decoder models and multiscale context aggregation [12,13,14,15] to transformer architectures with stronger long-range modeling [16,17,18,19]. Boundary-oriented methods commonly introduce an independent shape stream, model-agnostic boundary correction or boundary-sensitive evaluation [20,21,22]. Other context and edge-enhancement approaches improve semantic consistency or fine boundary recovery [23,24,25,26,27,28], while promptable segmentation models provide strong generic masks [29]. These developments confirm the importance of boundary quality when downstream measurements depend on narrow or elongated structures.
The proposed boundary detail enhancement module differs from a separate high-resolution shape stream and from post-processing that corrects an already generated mask. It is attached to the SegFormer decoder and fuses shallow boundary-sensitive features with high-level semantic features through channel alignment and an edge-aware spatial weight. This compact refinement is suited to rail light strips because their transverse width is small and their longitudinal structure is continuous. Small boundary displacements are amplified in subsequent centerline and width calculations. Thus, the study evaluates segmentation as both pixel classification and the geometric basis for rail-head-constrained abnormality analysis.

3. Materials and Methods

3.1. Overview of the Proposed Method

Figure 1 summarizes the segmentation-guided analysis pipeline. The color image is first divided into background, rail-head and light-strip regions. The rail-head mask defines the geometric reference, while the light-strip mask preserves the contact-trace morphology. The rule module then answers three operational questions. It identifies the mileage location, measures the strip shift relative to the rail-head centerline and describes width or local-contour changes. The outputs are therefore a location and measurable geometric indicators, rather than only a pixel mask.

3.2. Improved SegFormer for Rail-Head and Light-Strip Segmentation

SegFormer is used because its hierarchical Mix Transformer (MiT) encoder models long-range dependencies while retaining multiscale features [19]. The MiT-B2 variant is adopted for three-class segmentation of the background, rail head and light strip. A lightweight multilayer perceptron (MLP) decoder produces the initial prediction. The boundary detail enhancement module then refines the prediction before geometric analysis.
The input image is first encoded into multilevel feature maps. Low-level features preserve texture, brightness transition and boundary information, whereas high-level features provide more stable semantic context for distinguishing the rail head from the background. The decoder fuses these features and outputs a three-channel prediction map corresponding to the background, rail-head region and light-strip region. Compared with binary light-strip extraction, this three-class setting is more suitable for geometric analysis because the rail-head mask provides the reference frame for interpreting the light-strip position.
Weak or blurred boundaries are common in reflective rail images. The boundary detail enhancement module explicitly combines shallow texture information with the semantic feature produced by the SegFormer decoder. The two inputs are first aligned by 1 × 1 convolutions. Their concatenation is then processed by a boundary-guidance branch to generate a spatial weight map. This map enhances the semantic feature around rail-head and light-strip contours. Finally, the weighted semantic feature is concatenated with the aligned shallow feature and refined by a 3 × 3 convolution. The Gated Shape Convolutional Neural Network (Gated-SCNN) maintains a parallel shape stream [20]. SegFix corrects pixels after mask prediction [21]. By contrast, the proposed module performs feature refinement inside the decoder.
For an input image of H × W pixels, the MiT-B2 encoder outputs four feature maps F 1 , F 2 , F 3 , F 4 with channel numbers [64, 128, 320, 512] and nominal resolutions [ H 4 × W 4 , H 8 × W 8 , H 16 × W 16 , H 32 × W 32 ]. Each feature map is projected to 256 channels by the MLP decoder. F 1 is retained at the stage-1 resolution. F 2 , F 3 and F 4 are bilinearly resized with nominal factors of 2, 4 and 8, respectively. The resizing operation explicitly uses the F 1 spatial size, with align_corners set to false, to avoid mismatches caused by an odd image width. The aligned maps are concatenated in the order [ F 1 , U p F 2 , U p F 3 , U p F 4 ] to form a 1024-channel tensor. A 1 × 1 convolution followed by batch normalization and ReLU reduces it to the 256-channel decoder feature F d .
The shallow input of the boundary module is the stage-1 feature F 1 , and the high-level input is F d . A 1 × 1 convolution maps F 1 from 64 to 128 channels, and another 1 × 1 convolution maps F d from 256 to 128 channels. Both tensors therefore have 128 channels at H 4 × W 4 . They are concatenated in the order [ F d , F 1 ]. The concatenated tensor passes through a 3 × 3 convolution from 256 to 128 channels, followed by batch normalization and ReLU. A 1 × 1 convolution then maps 128 channels to 1 channel. Sigmoid converts the result into the boundary weight map. The aligned semantic feature is multiplied by one plus this weight map. It is then concatenated with the aligned F 1 feature. A final 3 × 3 convolution reduces the 256-channel tensor to a 64-channel refined feature. A 1 × 1 classifier produces three class logits, which are bilinearly resized to the input size.
With the 512 × 994 input used in the experiments, F 1 , F 2 , F 3 , F 4 have spatial sizes 128 × 249, 64 × 125, 32 × 63 and 16 × 32, respectively. All decoder features are resized to 128 × 249 before fusion. The boundary module outputs a 64 × 128 × 249 refined tensor and a 3 × 128 × 249 logit tensor. The final prediction is resized directly to 3 × 512 × 994. The boundary branch adds approximately 0.48 M trainable parameters, which is consistent with the 27.4 M to 27.9 M increase evaluated in Section 4.4.
P = S o f t m a x ( D θ ( E θ ( I ) ) ) , M ^ = a r g m a x P k
Here, I denotes the input color rail image. E θ and D θ denote the SegFormer encoder and decoder. P is the class probability map. k indexes the semantic class. M ^ is the predicted segmentation mask.
A = σ C 1 × 1 C 3 × 3 F s , F l F b = C 3 × 3 F s 1 + A , F l
Here, F s and F l are the 128-channel aligned semantic and shallow features, respectively. A is the single-channel boundary weight map. C k × k denotes a convolution with kernel size k × k . σ is the sigmoid function. F b is the 64-channel boundary-refined feature used by the final three-class classifier.
Focal Loss is further used during training to reduce the dominance of well-classified background pixels [30]. This setting corresponds to the loss component in Figure 2 and complements the boundary enhancement module in Figure 3. Most background pixels are classified with high confidence, whereas rail-head and light-strip pixels near weak or reflective boundaries are more difficult. The modulating factor in Focal Loss suppresses the loss contribution of these easy background pixels. Consequently, hard rail-head and light-strip pixels receive relatively greater optimization attention, which improves the stability of the boundaries used by the subsequent geometric measurements.
L f o c a l = 1 N n = 1 N k = 1 K α k ( 1 p n , k ) γ y n , k l o g ( p n , k )
Here, N is the number of pixels. K is the number of classes. y n , k is the one-hot label. p n , k is the predicted probability. α k was set to 0.25 for all three classes, and the focusing parameter γ was set to 2.0. The focusing term reduces the contribution of well-classified pixels rather than assigning different fixed weights to the three classes.
The improved structure retains the SegFormer encoder–decoder topology and adds only the boundary-refinement branch and Focal Loss. Section 4.4 evaluates its computational overhead. The baselines are DeepLab v3+, Segmenter, the Object-Contextual Representations Network (OCRNet) and the Bilateral Segmentation Network V2 (BiSeNetV2). They represent convolutional encoder–decoder segmentation, transformer segmentation, object-context modeling and real-time bilateral segmentation, respectively [15,17,24,25].

3.3. Geometric Rule-Based Abnormality Analysis

After obtaining the two masks, the prediction is restored from 512 × 994 pixels to the original resolution of 1024 × 1988 pixels by nearest-neighbor interpolation. The rail-head boundaries are then fitted, and their midpoint defines the rail centerline. The light-strip boundaries are sampled along the rail direction to obtain center position, width and local contour morphology. These quantities are associated with the corrected mileage record. Each output can therefore be traced to a consecutive 1 m detection unit, as illustrated in Figure 4.
The original longitudinal sampling resolution is 1 mm per image row. Therefore, the 1988 rows in one original image cover approximately 1.988 m of rail. Consecutive masks are ordered and concatenated according to the encoder-triggered and system-corrected mileage record. Every 1000 rows at the restored original resolution form one 1 m detection unit, and a unit may cross an image boundary. For row y in unit i , the left and right rail-head boundaries are denoted by x L r and x R r , and the light-strip boundaries by x L s and x R s . The row-wise light-strip width is b y = k x x R s y x L s y . Row-wise center positions and widths are aggregated by their median values within each 1 m unit to suppress isolated fluctuations caused by reflection and noise.
r i = m e d i a n x L r + x R r 2 , c i = m e d i a n x L s + x R s 2 B i = k x m e d i a n x R s x L s
Here, r i and c i are the rail-head and light-strip center positions in the i-th 1 m unit, respectively. B i is the corresponding light-strip width. k x is the horizontal image calibration coefficient in mm/pixel.
e i = k x c i r i , Δ B i = B i B i 1
Here e i is the signed lateral eccentricity of the light strip relative to the rail-head centerline, and Δ B i is the absolute width change between adjacent 1 m detection units. The eccentricity rule uses the absolute value e i in each unit rather than an adjacent-unit eccentricity difference.
D i = m a x Ω U i , Ω = 100 m e d i a n y Ω b y B i
m i = 1 D i > 10   m m L i 100   m m Δ B i 10   m m
A i = 1 e i > 6   m m + 1 Δ B i > 10   m m + m i
In Equation (6a–c), U i is the set of 1000 original-resolution rows in the i -th detection unit. Ω is a 100-row sliding window that corresponds to 100 mm in the longitudinal direction. D i is the maximum deviation between the local median width and the 1 m median width B i . L i is the longitudinal length covered by the merged abnormal windows. The binary indicator m i is activated only when a local width deviation exceeds 10 mm for at least 100 mm, and the adjacent-unit width-mutation rule is not triggered. Eccentricity abnormality is a positional condition and may coexist with local contour deformation. A i counts the triggered abnormality conditions, and a unit with A i > 0 is reported for manual review.
An eccentricity abnormality is reported when e i > 6   m m , and a width mutation is reported when Δ B i > 10   m m . A local integrity abnormality is reported when m i = 1 . It describes a short-range expansion, contraction or unilateral contour deformation within a 1 m unit, rather than an overall width change between adjacent 1 m units. The condition Δ B i 10   m m in Equation (6b) prevents the local-integrity category from duplicating width mutation. Eccentricity measures position, whereas local integrity measures contour morphology. Both indicators may therefore be activated in the same unit. The manually audited primary label is used for per-type statistics.
The rule thresholds are engineering mappings rather than parameters optimized on the validation set. DB34/T 3964-2021 specifies a 6 mm light-strip center offset [31]. For a local region longer than 100 mm, it also specifies an approximately 10 mm width difference relative to the overall mean [31]. The 6 mm value is used as the absolute eccentricity limit. The 10 mm value is used for both adjacent-unit width mutation and the local width-deviation test. The 100 mm local-window length follows the minimum longitudinal extent in the same local-abnormality requirement. The high-speed railway maintenance rule also identifies poor contact bands by widths below 20 mm or above 40 mm and by periodic width variation [32]. The 200 mm association threshold for neighboring local contour regions follows the criterion for adjacent surface damage in railway maintenance practice [33]. All rule calculations use masks restored to the original resolution. The longitudinal scale is 1 mm per row, and the calibrated horizontal scale is k x = 0.2   m m / p i x e l .
The geometric-rule hyperparameters are defined accordingly. They comprise a 6 mm absolute-eccentricity threshold, a 10 mm width-change threshold, a 100 mm local-window length and a 200 mm association distance. These values are applied after image measurements are converted to physical units. They remain identical for every segmentation model.
Longitudinal geometric sequences reduce alarms caused by isolated brightness changes. Reflection or stains may create high local contrast without a coherent centerline displacement, width change or contour deformation. A real abnormality is expected to persist spatially or produce a connected contour change. The rule module therefore evaluates local measurements together with connected-component morphology.
Each abnormal output record contains the 1 m start-end mileage interval, abnormality type, eccentricity, width change and local-integrity flag. The reported interval follows the corrected mileage sequence used to order and concatenate the masks. Therefore, a triggered unit can be traced to the corresponding inspection segment even when it crosses an image boundary. These are termed maintenance-related indicators because they can direct personnel to a specific interval for manual review. They do not constitute a maintenance grade or an automatic maintenance decision.

4. Experiments and Results

4.1. Dataset and Experimental Settings

The dataset contains 3520 color visible-light images acquired on multiple railway lines. Each original image has a resolution of 1024 × 1988 pixels. Pixel labels include background, rail head and light strip and are stored in the PASCAL Visual Object Classes (VOC) 2012 format [34,35]. The dataset was divided into 2464 training images and 1056 validation images at a ratio of 7:3. Splitting was performed by railway line and continuous acquisition section. Images from the same section and adjacent frames from the same sequence were assigned to only one subset, preventing section-level leakage.
The 1056-image validation set contained 876 manually audited ground-truth-positive 1 m units. These comprised 438 width-mutation, 292 eccentricity-abnormality and 146 local integrity abnormality units. The mutually exclusive primary labels are used for the per-type abnormality statistics. The counts represent detection units rather than images. Each 1024 × 1988 image covers approximately 1.988 m, and a 1 m unit may cross an image boundary. Negative validation units are not included in the 876 positive count but are retained when counting false positives. The abnormality labels are not additional semantic classes because normal and abnormal images share the same three-pixel categories. The complete dataset cannot be publicly released because it contains proprietary railway inspection data. Authorized access may be considered upon reasonable request.
Ground-truth masks were generated through human interactive annotation. Annotators first used interactive prompts to obtain the initial rail-head and light-strip regions. They then refined the contours manually at the pixel level, focusing on weak, reflective and locally deformed boundary areas. The refined masks were subsequently reviewed against the original color images. This review ensured consistent separation of the background, rail head and light strip.
The rail-surface imaging component was the GX3-LSM-02KGC-01A module (INMAI Railway Technology Co., Ltd., Beijing, China). It integrates a color line-scan camera, three-wavelength laser illumination, drive electronics, thermal management and a protective enclosure. The camera was a Linea LA-GC-02K05B color complementary metal-oxide-semiconductor (CMOS) line-scan camera (Teledyne DALSA, Waterloo, ON, Canada) with a 2048 × 2 pixel sensor layout and a pixel size of 7.04 × 7.04 μm. After region cropping, the effective line frequency was 45 kHz. Because the module uses line-scan imaging, line frequency, rather than area-camera frame rate, determines the longitudinal sampling capacity.
The camera was equipped with an LM25HC lens (Kowa Optronics Co., Ltd., Tokyo, Japan) with a focal length of 25 mm and an aperture of f/5.6. The nominal focusing distance from the module optical exit to the rail surface was 350 mm, and the specified working-distance range was 350–650 mm. Accordingly, 350 mm was used as the nominal installation-height reference. The configured camera field of view was 31.99°. The mounting angle in the vehicle-motion direction was adjustable from −8° to +8°, while the laser projection angle was 56°. The illumination combined 450, 520 and 650 nm laser bands to produce a white-light color image.
The camera accepted a 5–12 V external trigger through Line1 or Line2, and the laser source was synchronized by the camera trigger. This configuration allowed the rail head, bright light strip and surrounding fastener or ballast to be recorded in the same image. The fixed optical geometry also kept the rail-head and light-strip regions approximately aligned, providing a stable basis for annotation and geometric calibration.
The maximum inspection speed was 160 km/h. Image acquisition was triggered by the encoder, and the recorded position was corrected using the mileage information provided by the integrated inspection system. This mechanism maintained the correspondence between each line-scan image and its railway mileage under varying vehicle speeds. The acquisition system and imaging geometry are shown in Figure 5.
The experimental platform used an Intel Xeon 2.4 GHz central processing unit (CPU), an NVIDIA RTX 3090 graphics processing unit (GPU) and 256 GB of memory. The software environment comprised Ubuntu 20.04, Python 3.10, PyTorch 2.4 and Compute Unified Device Architecture (CUDA) 11.8. ImageNet-pretrained MiT-B2 parameters initialized the network [36]. During validation and inference, each 1024 × 1988 image was deterministically downsampled by 0.5 to 512 × 994 pixels.
Training samples were processed in the following order: image loading, annotation loading, RandomResize, RandomCrop, RandomFlip, PhotoMetricDistortion and PackSegInputs. RandomResize used the base scale (1024, 512), a ratio range of 0.5–2.0 and aspect-ratio preservation. RandomCrop used a crop size of 512 × 512 pixels and a maximum single-class ratio of 0.75. RandomFlip was applied with a probability of 0.5. PhotoMetricDistortion used the default brightness delta of 32, contrast range of 0.5–1.5, saturation range of 0.5–1.5 and hue delta of 18. No random rotation was used.
The model was trained for 40,000 iterations using stochastic gradient descent. The initial learning rate was 0.01, momentum was 0.9 and weight decay was 0.0005 [37]. Focal Loss used α = 0.25 and γ = 2.0 . A polynomial schedule with power 0.9 followed a 1500-iteration linear warm-up.
The physical mini-batch size was 8, with gradient accumulation over eight steps for an effective batch size of 64. The random seed was 42. Validation was performed every 1000 iterations. The checkpoint with the highest validation mIoU was selected, with light-strip IoU used as the tie-breaker. Preliminary training was used to finalize the learning rate, warm-up schedule, batch configuration and Focal Loss parameters. The criteria were loss convergence, validation stability and the memory capacity of the RTX 3090 platform. These settings control model optimization and are not used as thresholds in the geometric abnormality decision.
Evaluation is performed at three levels. Class-wise Intersection over Union (IoU) and mean Intersection over Union (mIoU) assess the three masks. Boundary F1-score (BF1) measures contour recovery within a 3-pixel tolerance [22]. Centerline mean absolute error (CL-MAE) and width mean absolute error (W-MAE) measure the geometric error propagated from segmentation. Recall, precision and F1-score assess the final abnormality labels obtained by applying the same rules to every segmentation model. The metrics are defined below.
I o U k = T P k T P k + F P k + F N k
m I o U = 1 K k = 1 K I o U k
B F 1 = 2 P b R b P b + R b
C L M A E = 1 N i = 1 N c i p r e d c i g t
W M A E = 1 N i = 1 N w i p r e d w i g t
R e c a l l = T P T P + F N , P r e c i s i o n = T P T P + F P , F 1 = 2 × P r e c i s i o n × R e c a l l P r e c i s i o n + R e c a l l
In Equations (7)–(12), the intersection and union pixel counts are computed for each of the three classes. Boundary precision and boundary recall use a 3-pixel matching tolerance. N is the number of valid longitudinal sampling rows. The predicted center and width are compared with ground truth. For abnormality analysis, true positives (TP), false positives (FP) and false negatives (FN) are counted from the geometric-rule outputs.

4.2. Segmentation Results

All ablation experiments were evaluated on the same validation set under identical training and inference settings. Compared with the baseline SegFormer, the boundary enhancement module increased mIoU by 1.08 percentage points and improved BF1 from 64.72% to 65.43%. Meanwhile, CL-MAE decreased from 1.80 to 1.15 px, and W-MAE decreased from 5.61 to 4.63 px. This indicates improved boundary localization and geometric measurement accuracy. Focal Loss mainly improved light-strip segmentation, increasing LS IoU from 93.88% to 94.76%. Using both components produced the best overall performance. The model reached 95.67% mIoU and 65.75% BF1. CL-MAE and W-MAE were lowest at 1.11 and 4.50 px, respectively. These results demonstrate the complementary effects of boundary refinement and hard-pixel learning. The ablation results are summarized in Table 1.
Figure 6 further illustrates the segmentation behavior. For normal samples, the rail-head mask remains continuous, and the light-strip mask preserves the elongated contact-band morphology. Stable masks are necessary because a fragmented reference or strip contour directly disturbs the centerline and width sequence.

4.3. Abnormality Analysis Results

BiSeNetV2 yields the highest precision of 90.11%, whereas the proposed model achieves 89.23%. The proposed model nevertheless provides the highest mIoU (95.67%), recall (96.46%) and F1-score (92.70%) among the compared methods. Relative to the original SegFormer, it improves mIoU by 2.81 percentage points, recall by 1.48 points and F1-score by 1.07 points. This indicates a better balance between missed and false abnormality detections. The comparative results are summarized in Table 2 and visualized in Figure 7.
Segmentation accuracy and abnormality-analysis performance are related but not identical. Similar mIoU values can lead to different recall or precision because the rule module depends on boundary stability and local mask shape. A small error on the light-strip boundary can change the measured eccentricity or width more than the same number of background pixels.
For each 1 m unit, the geometric module records the start-end mileage interval, triggered rule and corresponding eccentricity, width change and local integrity result. Consecutive masks are ordered using the encoder-triggered and system-corrected mileage record. Therefore, a triggered unit is localized to its associated 1 m mileage interval, including a unit that crosses an image boundary. A record with one or more triggered conditions is flagged for targeted manual review. Table 3 quantifies the abnormality decisions. The flagged records are used to prioritize the corresponding 1 m mileage intervals for manual review, without directly assigning maintenance grades or initiating automatic maintenance actions.
The validation set contained 876 ground-truth-positive 1 m units for geometric-rule evaluation. Each unit received one mutually exclusive primary label: width mutation, eccentricity abnormality or local integrity abnormality. A unit with multiple geometric manifestations was assigned the manually confirmed primary abnormality label for per-type counting. Negative validation units were retained when calculating false-positive counts. Table 3 reports the resulting TP, FP and FN counts.
Among the three categories, width mutation achieved the highest recall (98.17%) and precision (91.10%). Local integrity abnormality was more challenging, with 93.15% recall and 86.08% precision. Its local area changes and irregularly connected contours are more sensitive to reflection, stains and weak boundaries. Across all 876 positive units, 845 were correctly detected. The overall recall, precision and F1-score were 96.46%, 89.23% and 92.70%, respectively.
Figure 8 and Figure 9 examine representative abnormal morphologies and adverse imaging conditions. Each row compares the original image, ground truth, improved SegFormer prediction and light-strip error map. Red pixels are false positives, cyan pixels are false negatives, and gray pixels are correctly segmented light-strip regions.
Figure 10 traces the decision process for three representative validation cases. For L215, the prediction-derived width change is 11.0 mm and exceeds the 10 mm threshold, while the current-unit absolute eccentricity is 2.5 mm. For L732, the prediction-derived absolute eccentricity is 28.4 mm and exceeds the 6 mm limit, whereas the adjacent-unit width change is only 1.8 mm. For R1517, the maximum local width deviation is 33.0 mm, and its merged longitudinal extent is 205 mm, satisfying the local-integrity rule. Eccentricity abnormality describes the light-strip position relative to the rail-head centerline. It may coexist with a local contour abnormality, while the manually audited primary label is used for per-type statistics. All geometric measurements and rule decisions shown in the figure are calculated from the predicted masks. Ground-truth measurements are presented only as offline validation references. They do not participate in rule inference.

4.4. Computational Complexity and Deployment Efficiency

Deployment efficiency was evaluated at 512 × 994 pixels on the same RTX 3090 platform. With batch size 1, the MiT-B2 baseline has 27.4 M parameters. It requires 121.1 giga floating-point operations (GFLOPs). Its floating-point 32-bit (FP32) model is 109.6 MB, and its throughput is 38.7 frames/s. Peak GPU memory is 7.11 GB. The improved model has 27.9 M parameters, 122.5 GFLOPs, a 111.6 MB model, 38.3 frames/s throughput and 7.17 GB peak memory.
Geometric post-processing takes 2.58 ms per image. The complete improved pipeline therefore requires approximately 28.69 ms per image, corresponding to 34.9 frames/s. The boundary branch increases parameter count by 1.8%, computation by 1.2% and peak memory by 0.8%, while reducing throughput by about 1.0%. Focal Loss is used only during training and has no inference cost. The complete complexity results are summarized in Table 4.

5. Discussion

The method follows a perception-to-measurement process. SegFormer produces rail-head and light-strip masks, after which the rule module returns mileage, eccentricity, width change and local-integrity result. These outputs answer whether an abnormality is present, where it occurs and which geometric condition is triggered. This differs from a generic defect detector that returns only a class or bounding box.
The experiments also show why segmentation quality alone is insufficient. Higher mIoU does not automatically produce the highest precision after rule analysis. Boundary and centerline errors have a direct effect on geometric measurements. The BF1, CL-MAE and W-MAE results therefore complement mIoU and explain the improvement in abnormality recall and F1-score.
The rail-head mask is used as the geometric reference rather than treating the light strip as an isolated bright object. This reduces dependence on local intensity and makes each output traceable to a measured center shift, width change or local contour deformation. The indicators support targeted review but do not replace maintenance assessment.
Several limitations remain. The 6 mm absolute eccentricity and 10 mm width-mutation thresholds are engineering mappings whose transferability to different lines, speeds, illumination and camera installations requires external validation. The proprietary dataset limits independent reproduction. Reflection, stains and blurred boundaries can still cause localized false-positive and false-negative pixels, as shown in Figure 9. Table 3 shows that local integrity abnormality remains the most error-prone category. Reflection, stains and weak boundaries can disturb its local area and contour measurements.
Future work will evaluate the method on independent railway sections and the target onboard platform. It will also investigate normalized or adaptive thresholds, temporal consistency across consecutive images and the association between geometric indicators and verified maintenance records.

6. Conclusions

This study developed a rail light-strip abnormality analysis method that combines boundary-enhanced SegFormer segmentation with rail-head-constrained geometric rules. The method jointly extracts the rail head and light strip. Consecutive masks are ordered by corrected mileage and divided into 1 m units. For each triggered unit, the output record reports the start-end mileage interval, absolute eccentricity, width change and local-integrity result. This provides an explicit abnormality decision and identifies the geometric condition that triggered it.
The improved model achieves 95.67% mIoU, 65.75% BF1, 1.11 px CL-MAE and 4.50 px W-MAE. Among 876 ground-truth-positive 1 m units, 845 are correctly detected. Abnormality recall, precision and F1-score are 96.46%, 89.23% and 92.70%, respectively. Post-processing takes 2.58 ms per image, and the end-to-end throughput is 34.9 frames/s. These results support the use of the generated records for targeted manual review.
The reported interval is intended to guide targeted manual review rather than serve as an automatic maintenance grade. The longitudinal sampling and engineering rules used in this study define the localization resolution as 1 m. The available tests are limited to the inspected lines, acquisition configuration and RTX 3090 platform. Future work will evaluate independent railway sections and the target onboard system. It will also examine adaptive thresholds, temporal consistency across consecutive images and associations with verified maintenance records.

Author Contributions

Conceptualization, H.S.; Methodology, H.S. and Z.G.; Software, J.L.; Validation, N.W., L.W., S.W. and C.X.; Resources, Y.G. and N.W.; Data curation, Y.G., L.W. and S.W.; Writing—original draft, H.S.; Writing—review & editing, Z.G.; Supervision, Q.H.; Funding acquisition, Q.H. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Key Research and Development Program of China, grant number 2025YFF0519100; the China Academy of Railway Sciences Corporation Limited Research Fund, grant number 2023YJ041; and the INMAI Railway Technology Co., Ltd. Research Fund, grant number 2025IMXM09.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available from the corresponding author upon reasonable request and subject to authorization and approval for access to proprietary railway inspection data.

Acknowledgments

During the preparation of this manuscript, the authors used generative artificial intelligence tools for English-language editing, proofreading, and improving clarity. These tools were not used for research design, data generation, data analysis, or interpretation of the results. The authors reviewed and edited all AI-assisted content and take full responsibility for the content of this publication.

Conflicts of Interest

The authors were employed by either China Academy of Railway Sciences Corporation Limited or its wholly owned subsidiary, INMAI Railway Technology Co., Ltd. This study received funding from the National Key Research and Development Program of China, the China Academy of Railway Sciences Corporation Limited Research Fund, and the INMAI Railway Technology Co., Ltd. Research Fund.

References

  1. Du, Y.; Zhang, X.; Gao, Y.; Nan, Z. RSDNet: A New Multiscale Rail Surface Defect Detection Model. Sensors 2024, 24, 3579. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Faghih-Roohi, S.; Hajizadeh, S.; Nunez, A.; Babuska, R.; De Schutter, B. Deep Convolutional Neural Networks for Detection of Rail Surface Defects. In Proceedings of the International Joint Conference on Neural Networks; IEEE: New York, NY, USA, 2016; pp. 2584–2589. [Google Scholar] [CrossRef] [Scilit]
  3. Gan, J.; Li, Q.; Wang, J.; Yu, H. A Hierarchical Extractor-Based Visual Rail Surface Inspection System. IEEE Sens. J. 2017, 17, 7935–7944. [Google Scholar] [CrossRef] [Scilit]
  4. Gibert, X.; Patel, V.M.; Chellappa, R. Robust Fastener Detection for Autonomous Visual Railway Track Inspection. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2015; pp. 694–701. [Google Scholar] [CrossRef] [Scilit]
  5. Mauz, F.; Wigger, R.; Gota, A.-E.; Kuffa, M. Automatic Detection of the Running Surface of Railway Tracks Based on Laser Profilometer Data and Supervised Machine Learning. Sensors 2024, 24, 2638. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Xia, Y.; Han, S.W.; Kwon, H.J. Image Generation and Recognition for Railway Surface Defect Detection. Sensors 2023, 23, 4793. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Ming, G.; Zhou, B.; Luo, X.; Ling, R.; Zhou, M. Rail Surface Defect Detection Method Based on Deep Learning Method with 3D Range Image. In Advances in Frontier Research on Engineering Structures; Springer: Singapore, 2023; Volume 286, pp. 45–59. [Google Scholar] [CrossRef] [Scilit]
  8. Wang, X.; Xu, X.; Mei, X.; Guo, X. Localization and Pixel-Confidence Network for Surface Defect Segmentation. Sensors 2025, 25, 4548. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Leiñena, A.; Saiz, F.; Barandiaran, I. Latent Diffusion Models to Enhance the Performance of Visual Defect Segmentation Networks in Steel Surface Inspection. Sensors 2024, 24, 6016. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Zhong, Y.; Chen, G. Rail Surface Defect Detection Based on Dual-Path Feature Fusion. Electronics 2024, 13, 2564. [Google Scholar] [CrossRef] [Scilit]
  11. Gosiewska, A.; Baran, Z.; Baran, M.; Rutkowski, T. Seeking a Sufficient Data Volume for Railway Infrastructure Component Detection with Computer Vision Models. Sensors 2023, 23, 7776. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Long, J.; Shelhamer, E.; Darrell, T. Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2015; pp. 3431–3440. [Google Scholar] [CrossRef] [Scilit]
  13. Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention; IEEE: New York, NY, USA, 2015; pp. 234–241. [Google Scholar] [CrossRef] [Scilit]
  14. Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid Scene Parsing Network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 2881–2890. [Google Scholar] [CrossRef] [Scilit]
  15. Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Proceedings of the European Conference on Computer Vision; IEEE: New York, NY, USA, 2018; pp. 801–818. [Google Scholar] [CrossRef] [Scilit]
  16. Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P.H.S.; et al. Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 6877–6886. [Google Scholar] [CrossRef] [Scilit]
  17. Strudel, R.; Garcia, R.; Laptev, I.; Schmid, C. Segmenter: Transformer for Semantic Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 7262–7272. [Google Scholar] [CrossRef] [Scilit]
  18. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 10012–10022. [Google Scholar] [CrossRef] [Scilit]
  19. Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. In Advances in Neural Information Processing Systems; IEEE: New York, NY, USA, 2021; Volume 34, pp. 12077–12090. [Google Scholar]
  20. Takikawa, T.; Acuna, D.; Jampani, V.; Fidler, S. Gated-SCNN: Gated Shape CNNs for Semantic Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 5229–5238. [Google Scholar] [CrossRef] [Scilit]
  21. Yuan, Y.; Xie, J.; Chen, X.; Wang, J. SegFix: Model-Agnostic Boundary Refinement for Segmentation. In Computer Vision-ECCV 2020; Springer: Cham, Switzerland, 2020; pp. 489–506. [Google Scholar] [CrossRef] [Scilit]
  22. Cheng, B.; Girshick, R.; Dollar, P.; Berg, A.C.; Kirillov, A. Boundary IoU: Improving Object-Centric Image Segmentation Evaluation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 15334–15342. [Google Scholar] [CrossRef] [Scilit]
  23. Cheng, B.; Misra, I.; Schwing, A.G.; Kirillov, A.; Girdhar, R. Masked-Attention Mask Transformer for Universal Image Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 1290–1299. [Google Scholar] [CrossRef] [Scilit]
  24. Yuan, Y.; Chen, X.; Wang, J. Object-Contextual Representations for Semantic Segmentation. In Proceedings of the European Conference on Computer Vision; IEEE: New York, NY, USA, 2020; pp. 173–190. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Yu, C.; Gao, C.; Wang, J.; Yu, G.; Shen, C.; Sang, N. BiSeNet V2: Bilateral Network with Guided Aggregation for Real-Time Semantic Segmentation. Int. J. Comput. Vis. 2021, 129, 3051–3068. [Google Scholar] [CrossRef] [Scilit]
  26. Zhu, Y.; Xiao, N. Simple Scalable Multimodal Semantic Segmentation Model. Sensors 2024, 24, 699. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Shu, X.; Zhao, X. Multi-Resolution Learning and Semantic Edge Enhancement for Super-Resolution Semantic Segmentation of Urban Scene Images. Sensors 2024, 24, 4522. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Zhang, X.; Chen, X.; Gao, X. Semantic Guidance Fusion Network for Cross-Modal Semantic Segmentation. Sensors 2024, 24, 2473. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment Anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2023; pp. 4015–4026. [Google Scholar] [CrossRef] [Scilit]
  30. Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar] [CrossRef] [Scilit]
  31. DB34/T 3964-2021; Technical Specification for Rail Grinding and Maintenance of Urban Rail Transit. Anhui Provincial Administration for Market Regulation: Hefei, China, 2021.
  32. National Railway Administration. Rules for Maintenance of High-Speed Railway Lines, Guo Tie She Bei Jian Gui [2023] No. 15; National Railway Administration: Beijing, China, 2023.
  33. China State Railway Group Co., Ltd., Engineering and Electrical Department. Railway Line Maintenance and Repair; China Railway Publishing House: Beijing, China, 2021; Chapter 3, Section 2. [Google Scholar]
  34. Russell, B.C.; Torralba, A.; Murphy, K.P.; Freeman, W.T. LabelMe: A Database and Web-Based Tool for Image Annotation. Int. J. Comput. Vis. 2008, 77, 157–173. [Google Scholar] [CrossRef] [Scilit]
  35. Everingham, M.; Van Gool, L.; Williams, C.K.I.; Winn, J.; Zisserman, A. The Pascal Visual Object Classes Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef] [Scilit]
  36. Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Fei-Fei, L. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2009; pp. 248–255. [Google Scholar] [CrossRef] [Scilit]
  37. Sutskever, I.; Martens, J.; Dahl, G.; Hinton, G. On the Importance of Initialization and Momentum in Deep Learning. In Proceedings of the International Conference on Machine Learning; IEEE: New York, NY, USA, 2013; pp. 1139–1147. [Google Scholar]
Figure 1. Overview of the proposed rail light-strip abnormality analysis method.
Figure 1. Overview of the proposed rail light-strip abnormality analysis method.
Sensors 26 05292 g001
Figure 2. Architecture of the improved SegFormer for rail-head and light-strip segmentation, including the hierarchical encoder, lightweight MLP decoder, boundary detail enhancement module and Focal Loss.
Figure 2. Architecture of the improved SegFormer for rail-head and light-strip segmentation, including the hierarchical encoder, lightweight MLP decoder, boundary detail enhancement module and Focal Loss.
Sensors 26 05292 g002
Figure 3. Boundary detail enhancement module used in the improved SegFormer.
Figure 3. Boundary detail enhancement module used in the improved SegFormer.
Sensors 26 05292 g003
Figure 4. Mask-derived geometric parameters used for rail light-strip abnormality analysis.
Figure 4. Mask-derived geometric parameters used for rail light-strip abnormality analysis.
Sensors 26 05292 g004
Figure 5. Rail light-strip image acquisition setup. (a) Comprehensive inspection vehicle; (b) visible-light acquisition component; (c) imaging geometry between the acquisition component and rail head; (d) representative acquired color rail image.
Figure 5. Rail light-strip image acquisition setup. (a) Comprehensive inspection vehicle; (b) visible-light acquisition component; (c) imaging geometry between the acquisition component and rail head; (d) representative acquired color rail image.
Sensors 26 05292 g005
Figure 6. Segmentation results of the improved SegFormer on color rail inspection images. The left column shows original image patches and the right column shows corresponding masks.
Figure 6. Segmentation results of the improved SegFormer on color rail inspection images. The left column shows original image patches and the right column shows corresponding masks.
Sensors 26 05292 g006
Figure 7. Comparison of segmentation and abnormality-analysis performance.
Figure 7. Comparison of segmentation and abnormality-analysis performance.
Sensors 26 05292 g007
Figure 8. Typical examples of the three-rail light-strip abnormality categories. Rows (ac) show eccentricity abnormality, width mutation and local integrity abnormality, respectively. The columns show the original image, Ground truth, improved SegFormer prediction and light-strip error map. Red and cyan denote false-positive and false-negative light-strip pixels, respectively.
Figure 8. Typical examples of the three-rail light-strip abnormality categories. Rows (ac) show eccentricity abnormality, width mutation and local integrity abnormality, respectively. The columns show the original image, Ground truth, improved SegFormer prediction and light-strip error map. Red and cyan denote false-positive and false-negative light-strip pixels, respectively.
Sensors 26 05292 g008
Figure 9. Challenging rail light-strip samples and residual segmentation errors. Rows (ac) show uneven reflection with low illumination, weak contrast with blurred boundaries, and stain interference with blurred boundaries, respectively. The columns show the original image, ground truth, improved SegFormer prediction and light-strip error map. Red and cyan denote false-positive and false-negative light-strip pixels, respectively.
Figure 9. Challenging rail light-strip samples and residual segmentation errors. Rows (ac) show uneven reflection with low illumination, weak contrast with blurred boundaries, and stain interference with blurred boundaries, respectively. The columns show the original image, ground truth, improved SegFormer prediction and light-strip error map. Red and cyan denote false-positive and false-negative light-strip pixels, respectively.
Sensors 26 05292 g009
Figure 10. Representative geometric-rule abnormality decisions. Rows (ac) show width mutation, eccentricity abnormality and local integrity abnormality, respectively. The columns compare the original image, Ground truth, prediction with rail-head and light-strip centerlines, prediction-derived geometric measurements and the triggered rule. Ground truth is shown only as an offline validation reference and is not involved in rule inference. The white dashed line marks a boundary between adjacent 1 m detection units.
Figure 10. Representative geometric-rule abnormality decisions. Rows (ac) show width mutation, eccentricity abnormality and local integrity abnormality, respectively. The columns compare the original image, Ground truth, prediction with rail-head and light-strip centerlines, prediction-derived geometric measurements and the triggered rule. Ground truth is shown only as an offline validation reference and is not involved in rule inference. The white dashed line marks a boundary between adjacent 1 m detection units.
Sensors 26 05292 g010
Table 1. Ablation study with class-wise segmentation, boundary-quality and geometric-measurement metrics. BG: background; RH: rail head; LS: light strip; BF1: boundary F1 score computed with a 3-pixel tolerance; CL-MAE: centerline mean absolute error; W-MAE: width mean absolute error. Lower CL-MAE and W-MAE values indicate better performance.
Table 1. Ablation study with class-wise segmentation, boundary-quality and geometric-measurement metrics. BG: background; RH: rail head; LS: light strip; BF1: boundary F1 score computed with a 3-pixel tolerance; CL-MAE: centerline mean absolute error; W-MAE: width mean absolute error. Lower CL-MAE and W-MAE values indicate better performance.
VariantBG IoURH IoULS IoUmIoUBF1CL-MAEW-MAE
SegFormer95.00%89.70%93.88%92.86%64.72%1.80 px5.61 px
+Boundary enhancement96.10%91.60%94.12%93.94%65.43%1.15 px4.63 px
+Focal Loss96.35%91.25%94.76%94.12%64.96%1.38 px5.27 px
+Boundary enhancement + Focal Loss97.18%93.52%96.31%95.67%65.75%1.11 px4.50 px
Table 2. Abnormality analysis performance based on different segmentation models and the same geometric rules.
Table 2. Abnormality analysis performance based on different segmentation models and the same geometric rules.
ModelmIoURecallPrecisionF1-Score
DeepLab v3+90.52%87.67%85.91%86.78%
Segmenter91.31%91.10%85.81%88.38%
OCRNet91.67%93.15%88.31%90.67%
BiSeNetV292.12%93.61%90.11%91.83%
SegFormer92.86%94.98%88.51%91.63%
Improved SegFormer95.67%96.46%89.23%92.70%
Table 3. Per-type abnormality-analysis results on the validation set. Ground-truth-positive units are mutually exclusive primary abnormality labels. Bold values indicate the overall results.
Table 3. Per-type abnormality-analysis results on the validation set. Ground-truth-positive units are mutually exclusive primary abnormality labels. Bold values indicate the overall results.
Abnormality TypeGround-Truth Positive UnitsTPFPFNRecallPrecisionF1-Score
Width mutation43843042898.17%91.10%94.51%
Eccentricity abnormality292279381395.55%88.01%91.63%
Local integrity abnormality146136221093.15%86.08%89.47%
Overall8768451023196.46%89.23%92.70%
Table 4. Model complexity and deployment efficiency at 512 × 994 pixels. Frames per second (FPS) reports model inference throughput with batch size 1; post-processing time is measured separately.
Table 4. Model complexity and deployment efficiency at 512 × 994 pixels. Frames per second (FPS) reports model inference throughput with batch size 1; post-processing time is measured separately.
ModelParamsFLOPsModel SizeFPSPeak MemoryPost-Process
SegFormer-B227.4 M 121.1 G109.6 MB38.77.11 GB2.58 ms
Improved SegFormer27.9 M122.5 G111.6 MB38.37.17 GB2.58 ms
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Song, H.; Gou, Y.; Wang, N.; Wang, L.; Liu, J.; Wang, S.; Xia, C.; Han, Q.; Gu, Z. Rail Light-Strip Abnormality Analysis from Color Inspection Images Using an Improved SegFormer and Geometric Rules. Sensors 2026, 26, 5292. https://doi.org/10.3390/s26165292

AMA Style

Song H, Gou Y, Wang N, Wang L, Liu J, Wang S, Xia C, Han Q, Gu Z. Rail Light-Strip Abnormality Analysis from Color Inspection Images Using an Improved SegFormer and Geometric Rules. Sensors. 2026; 26(16):5292. https://doi.org/10.3390/s26165292

Chicago/Turabian Style

Song, Haoran, Yuntao Gou, Ning Wang, Le Wang, Junbo Liu, Shengchun Wang, Chengliang Xia, Qiang Han, and Zichen Gu. 2026. "Rail Light-Strip Abnormality Analysis from Color Inspection Images Using an Improved SegFormer and Geometric Rules" Sensors 26, no. 16: 5292. https://doi.org/10.3390/s26165292

APA Style

Song, H., Gou, Y., Wang, N., Wang, L., Liu, J., Wang, S., Xia, C., Han, Q., & Gu, Z. (2026). Rail Light-Strip Abnormality Analysis from Color Inspection Images Using an Improved SegFormer and Geometric Rules. Sensors, 26(16), 5292. https://doi.org/10.3390/s26165292

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop