A Direction-Aware Lightweight Network for Camera-Based Underground Mine Track Region Segmentation
Abstract
1. Introduction
- An underground mine track region segmentation task is formulated for front-camera imagery from rail-guided mine vehicles, with the foreground defined as the visible rail-bounded track surface in the local camera view.
- RailDLA is developed as a lightweight encoder–decoder for this task, combining a track context preconditioning block, an RDLA block, context-guided feature fusion, and a track axis proxy decoding branch for sparse, elongated foreground segmentation.
- Comparative accuracy and efficiency experiments are conducted on a self-constructed dataset of underground mine vehicle imagery; track region metrics and controlled runtime measurements are reported using the same GPU.
2. Related Work
2.1. Railway and Underground Track Perception
2.2. Track Region Segmentation and Dense Prediction
2.3. Real-Time Semantic Segmentation
2.4. Convolutional Attention, Transformers, and Linear Context
3. Method
3.1. Problem Formulation and Track Structural Prior
3.2. Network Overview

3.3. Track Context Preconditioning

3.4. Direction-Aware Rail Linear Attention

3.5. Structured Affinity and Stability

3.6. Track Axis Proxy Decoding
| Algorithm 1 Forward propagation of the track axis proxy decoder propagation of the track axis proxy decoder | |
| Require: fused decoder feature | |
| Require: branch width , strip kernel size , modulation factor | |
| Ensure: segmentation logits and proxy logit | |
| 1: | ▹ ConvModule |
| 2: | ▹ depthwise separable convolution |
| 3: , | ▹ depthwise strip convolutions |
| 4: | ▹ ConvModule |
| 5: | ▹ convolution |
| 6: | |
| 7: | |
| 8: | |
| 9: | ▹ classifier |
| 10: return | |
3.7. Geometry-Aware Objective
4. Experiments
4.1. Dataset and Task Setting
4.2. Evaluation Metrics
4.3. Implementation Details
4.4. Compared Methods
5. Results
5.1. Overall Quantitative Comparison
5.2. Focused Lightweight Accuracy Comparison
5.3. Ablation Study
5.4. CPU Runtime Check
5.5. Qualitative and Failure Case Analysis
5.6. Image-Level Five-Fold Stability Check
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Jiskani, I.M.; Zhou, W.; Hosseini, S.; Wang, Z. Mining 4.0 and climate neutrality: A unified and reliable decision system for safe, intelligent, and green & climate-smart mining. J. Clean. Prod. 2023, 410, 137313. [Google Scholar] [CrossRef]
- Zhang, H.; Li, B.; Karimi, M.; Saydam, S.; Hassan, M. Recent Advancements in IoT Implementation for Environmental, Safety, and Production Monitoring in Underground Mines. IEEE Internet Things J. 2023, 10, 14507–14526. [Google Scholar] [CrossRef]
- You, K.; Shao, H.; Chen, Z.; Yang, J.; Du, X.; Lin, Y.; Wang, Y. A physics-constrained multimodal LLM for fault diagnosis of pressurized water reactor coolant systems with imbalanced and under-sampled data. J. Ind. Inf. Integr. 2026, 52, 101143. [Google Scholar] [CrossRef]
- UH, D.; Prem Kumar, J. Enhanced Transfer Learning-Based CNN for Abnormal Human Activity Detection in Video Surveillance Using Spatial-Temporal Features. Cybern. Syst. 2025, 1–30. [Google Scholar] [CrossRef]
- Long, J.; Shelhamer, E.; Darrell, T. Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2015; pp. 3431–3440. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention; Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2018; pp. 801–818. [Google Scholar]
- Zhao, F.; He, Y.; Song, J.; Wang, J.; Xi, D.; Shao, X.; Wu, Q.; Liu, Y.; Chen, Y.; Zhang, G.; et al. Smart UAV-assisted blueberry maturity monitoring with Mamba-based computer vision. Precis. Agric. 2025, 26, 56. [Google Scholar] [CrossRef]
- Sun, J.; Li, D.; Kuai, Y.; Belotserkovsky, A.; Lukashevich, P. Towards Unified Transformer for UAV-Based Multi-Task Oriented Object Detection. IET Image Process. 2026, 20, e70343. [Google Scholar] [CrossRef]
- Zhao, F.; Xu, D.; Ren, Z.; Shao, X.; Wu, Q.; Liu, Y.; Wang, J.; Song, J.; Chen, Y.; Zhang, G.; et al. Mamba-based super-resolution and semi-supervised YOLOv10 for freshwater mussel detection using acoustic video camera: A case study at Lake Izunuma, Japan. Ecol. Inform. 2025, 90, 103324. [Google Scholar] [CrossRef]
- Chen, Z.; Yang, J.; Chen, L.; Feng, Z.; Jia, L. Efficient Railway Track Region Segmentation Algorithm Based on Lightweight Neural Network and Cross-fusion Decoder. Autom. Constr. 2023, 155, 105069. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Wang, X.; Girshick, R.; Gupta, A.; He, K. Non-Local Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 7794–7803. [Google Scholar]
- Yang, H.; Liu, Z.; Liu, W.; Wang, H.; Zhang, Y.; Wang, H. Graph-MDETR: A Graph-Guided Mamba-DETR Network for UAV Catenary Support Components Detection in Electrified Railways. IEEE Trans. Intell. Transp. Syst. 2026, 27, 6319–6332. [Google Scholar] [CrossRef]
- Duan, F.; Wang, H.; Yang, H.; Wei, C.; Zhang, C.; Song, Y.; Liu, Z. MRFM-IFCOS: An Anchor-Free Interactive Detector Based on Multireceptive Field Mamba for Detecting Catenary Support Components. IEEE Trans. Instrum. Meas. 2025, 74, 2553415. [Google Scholar] [CrossRef]
- Yang, H.; Hu, K.; Wang, H.; Hong, W.; Wang, X.; Wang, H.; Song, Y.; Liu, Z. BCLIP-ADer: A Bayesian Prompt Contrastive Language-Image Pretraining Method for Catenary Component Anomaly Detection in Electrified Railways. IEEE Trans. Transp. Electrif. 2026. [Google Scholar] [CrossRef]
- Yan, J.; Zhou, N.; Cheng, Y.; Zhang, F.; Wang, H.; Wang, M.; Jin, B.; Li, M.; Lu, Q.; Zhang, W. Application of Machine-Vision-Driven Physics-Informed Neural Networks in Pantograph–Catenary System State Detection. Mech. Syst. Signal Process. 2026, 257, 114577. [Google Scholar] [CrossRef]
- Yan, J.; Chen, B.; Zhang, F.; Cheng, Y.; Wang, H.; Wang, H.; Wang, M.; Li, T.; Zhang, W. Meta-Learning-Based Graph Convolutional Wavelet Network for Intelligent Dynamic Modeling of High-Speed Rail Subsystems. IEEE Trans. Veh. Technol. 2026, 1–16. [Google Scholar] [CrossRef]
- Gong, Y.; Lin, L.; Luo, Y.; Liu, H.; Gao, Y.; Zhao, J.; Song, Z.; Hu, X. Implicit Illumination-Aware Representation with Cross-Modal Prefusion Alignment for Universal Multispectral Pedestrian Detection. IEEE Trans. Neural Netw. Learn. Syst. 2026, 37, 3247–3261. [Google Scholar] [CrossRef] [PubMed]
- Chen, Z.; You, K.; Yang, J.; Chen, L.; Li, F.; Feng, Z.; Jia, L. A sparse-to-dense guided fusion framework for three-dimensional object detection in railway environments. Eng. Appl. Artif. Intell. 2026, 178, 115095. [Google Scholar] [CrossRef]
- Chen, Z.; Yang, J.; Chen, L.; Li, F.; Feng, Z.; Jia, L.; Li, P. RailVoxelDet: A Lightweight 3-D Object Detection Method for Railway Transportation Driven by Onboard LiDAR Data. IEEE Internet Things J. 2025, 12, 37175–37189. [Google Scholar] [CrossRef]
- Yang, G.; Jiang, Y.; Wang, S.; Chen, K. VinsFusion-Line: Binocular Vision Inertial Navigation Real-Time SLAM System Based on Line Features. Cybern. Syst. 2026, 57, 350–375. [Google Scholar] [CrossRef]
- Mounika, P.; Narayanan, B.; Balmuri, K.R. An Adaptive Trans-ResUnet++Based Segmentation and Hybrid CNN-Aided Classification for Detecting Breast Cancer from Mammogram Images. Cybern. Syst. 2025, 1–38. [Google Scholar] [CrossRef]
- Poudel, R.P.; Liwicki, S.; Cipolla, R. Fast-SCNN: Fast Semantic Segmentation Network. In Proceedings of the British Machine Vision Conference; BMVA Press: Durham, UK, 2019; p. 289. [Google Scholar]
- Chen, Z.; Yang, J.; Li, F.; Feng, Z.; Chen, L.; Jia, L.; Li, P. Foreign Object Detection Method for Railway Catenary Based on a Scarce Image Generation Model and Lightweight Perception Architecture. IEEE Trans. Circuits Syst. Video Technol. 2026, 36, 1377–1391. [Google Scholar] [CrossRef]
- Shajeena, J.; Govindasamy, B.; Gnanasundaram, M.; Joel, M.R. Mobile-Le Harmonic Fusion Network for Object Recognition and SiamMoT Based Multi-Object Tracking Using Video Surveillance. Cybern. Syst. 2026, 57, 866–896. [Google Scholar] [CrossRef]
- Chen, Z.; Guo, H.; Yang, J.; Jiao, H.; Feng, Z.; Chen, L.; Gao, T. Fast vehicle detection algorithm in traffic scene based on improved SSD. Measurement 2022, 201, 111655. [Google Scholar] [CrossRef]
- Yu, C.; Gao, C.; Wang, J.; Yu, G.; Shen, C.; Sang, N. BiSeNet V2: Bilateral Network with Guided Aggregation for Real-time Semantic Segmentation. Int. J. Comput. Vis. 2021, 129, 3051–3068. [Google Scholar] [CrossRef]
- Pan, H.; Hong, Y.; Sun, W.; Jia, Y. Deep Dual-Resolution Networks for Real-Time and Accurate Semantic Segmentation of Traffic Scenes. IEEE Trans. Intell. Transp. Syst. 2023, 24, 3448–3460. [Google Scholar] [CrossRef]
- Xu, J.; Xiong, Z.; Bhattacharyya, S.P. PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 19529–19539. [Google Scholar]
- Chen, Z.; Yang, J.; Feng, Z.; Zhu, H. RailFOD23: A dataset for foreign object detection on railroad transmission lines. Sci. Data 2024, 11, 72. [Google Scholar] [CrossRef] [PubMed]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 7132–7141. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2018; pp. 3–19. [Google Scholar]
- Yang, J.; Jiang, Y.; Jiang, D.; Chen, Z. Infrared–Visible Fusion via Cross-Modality Attention and Small-Object Enhancement for Pedestrian Detection. ISPRS Int. J. Geo-Inf. 2025, 14, 477. [Google Scholar] [CrossRef]
- Sreekala, K.; Maniraj, S.P.; Singh, A.; Singh, A.P.; Pyingkodi, M.; Inthiyaz, S. Enhancing Medical Diagnosis through Multimodal Image Fusion: A Novel Approach Using Modified Swin-Based Cross Attention Fusion. Cybern. Syst. 2026, 57, 765–807. [Google Scholar] [CrossRef]
- Peng, C.; Zhang, X.; Yu, G.; Luo, G.; Sun, J. Large Kernel Matters: Improve Semantic Segmentation by Global Convolutional Network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2017; pp. 4353–4361. [Google Scholar]
- Huang, Z.; Wang, X.; Huang, L.; Huang, C.; Wei, Y.; Liu, W. CCNet: Criss-Cross Attention for Semantic Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2019; pp. 603–612. [Google Scholar]
- Guo, M.H.; Lu, C.Z.; Hou, Q.; Liu, Z.; Cheng, M.M.; Hu, S.M. SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation. Adv. Neural Inf. Process. Syst. 2022, 35, 1140–1156. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations; Curran Associates, Inc.: Red Hook, NY, USA, 2021. [Google Scholar]
- Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P.H.S.; et al. Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2021; pp. 6881–6890. [Google Scholar]
- Strudel, R.; Garcia, R.; Laptev, I.; Schmid, C. Segmenter: Transformer for Semantic Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2021; pp. 7262–7272. [Google Scholar]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
- Katharopoulos, A.; Vyas, A.; Pappas, N.; Fleuret, F. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. In Proceedings of the International Conference on Machine Learning, PMLR; JMLR: New York, NY, USA, 2020; pp. 5156–5165. [Google Scholar]
- You, K.; Gu, Y.; Shao, H.; Wang, Y. A liquid-impulse neural network model based on heterogeneous fusion of multimodal information for interpretable rotating machinery fault diagnosis. Mech. Syst. Signal Process. 2026, 246, 113923. [Google Scholar] [CrossRef]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 4510–4520. [Google Scholar]
- Woo, S.; Debnath, S.; Hu, R.; Chen, X.; Liu, Z.; Kweon, I.S.; Xie, S. ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 16133–16142. [Google Scholar]
- Xu, M.; Zhang, Z.; Wei, F.; Hu, H.; Bai, X. Side Adapter Network for Open-Vocabulary Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 2945–2954. [Google Scholar]
- Zhu, C.; Suri, S.; Jose, C.; Oquab, M.; Szafraniec, M.; Wen, W.; Xiong, Y.; Labatut, P.; Bojanowski, P.; Krishnamoorthi, R.; et al. Efficient Universal Perception Encoder. arXiv 2026, arXiv:2603.22387. [Google Scholar]








| Input Size | Operator | t | c | n | s |
|---|---|---|---|---|---|
| Conv | – | 32 | 1 | 2 | |
| DSConv | – | 48 | 1 | 2 | |
| DSConv | – | 64 | 1 | 2 | |
| IRB | 6 | 64 | 3 | 2 | |
| IRB | 6 | 96 | 3 | 2 | |
| IRB | 6 | 128 | 3 | 1 | |
| Conv | – | 128 | 1 | 1 | |
| Track context preconditioning block | – | 128 | 1 | 1 | |
| RDLA block | – | 128 | 1 | 1 | |
| Bilinear upsampling | – | 128 | 1 | ||
| / | Context-guided feature fusion | – | 128 | 1 | 1 |
| Decoder DSConv | – | 128 | 2 | 1 | |
| Track axis proxy decoding branch | – | 64 | 1 | 1 | |
| Final segmentation head | – | 2 | 1 | 1 |
| (a) Common acquisition configuration | ||||||
| Item | Setting | |||||
| Platform and operation | Rail-guided mine vehicle with an NVIDIA Jetson Orin Nano; routine bidirectional recording | |||||
| Camera and modality | Front-facing Orbbec Gemini 335Le RGB-D camera (Orbbec, Shenzhen, China); RGB stream used; depth stream excluded | |||||
| RGB acquisition | or pixels; 30 frames/s; automatic exposure and white balance | |||||
| Mounting geometry | Front-center mounting, approximately 1.35 m above the rail top, with an 8° downward pitch | |||||
| Temporal sampling | One RGB frame retained every 100 recorded frames (approximately every 3.3 s) | |||||
| (b) Site coverage and retained-frame allocation | ||||||
| Acquisition volume | Image split | |||||
| Site | Routes | Videos | Frames | Train | Val. | Representative conditions |
| M1 | 3 | 24 | 2278 | 1822 | 456 | Inclined roadway; weak illumination |
| M2 | 2 | 20 | 1876 | 1501 | 375 | Dust; mixed pedestrian–vehicle traffic |
| M3 | 3 | 28 | 2689 | 2151 | 538 | Water reflection; multiple tracks |
| M4 | 2 | 24 | 2286 | 1829 | 457 | Curved track; illumination transitions |
| Total | 10 | 96 | 9129 | 7303 | 1826 | Four underground mines |
| Category | Setting |
|---|---|
| Training schedule | 100 epochs; validation every 10 epochs |
| Implementation framework | MMSegmentation 1.2.2 with MMEngine 0.10.7 and MMCV 2.2.0 |
| Training scope | Models corresponding to all accuracy entries in Table 4 were trained from scratch by the authors |
| Weight initialization | Default random initialization in MMSegmentation and PyTorch 2.4.0, without pretrained backbone weights |
| Random seed | 1,655,745,629 for the local training runs |
| Training hardware | NVIDIA GeForce RTX 3070 Laptop GPU with 8 GB memory |
| Checkpoint selection | Best validation mIoU over the 100-epoch schedule |
| Optimizer | Stochastic gradient descent (SGD) |
| Initial learning rate | 0.12 |
| Learning-rate schedule | Polynomial decay, power 0.9 |
| Momentum | 0.9 |
| Weight decay | |
| Batch size | 16 |
| Training crop | Random crop |
| Input scale augmentation | Random resize ratio from 0.5 to 2.0 |
| Horizontal flip | Probability 0.5 |
| Photometric augmentation | Photometric distortion |
| Complexity input size |
| Model | mIoU (%) | IoU (%) | Acc. (%) | Overall (%) | FPS ↑ | ||||
|---|---|---|---|---|---|---|---|---|---|
| Best ↑ | Final ↑ | Bg ↑ | Track ↑ | Bg ↑ | Track ↑ | aAcc ↑ | mAcc ↑ | ||
| SegNeXt-S [38] | 95.70 | 95.66 | 99.60 | 91.73 | 99.83 | 95.06 | 99.61 | 97.44 | – |
| SegNeXt-T [38] | 95.36 | 95.36 | 99.57 | 91.15 | 99.81 | 94.72 | 99.58 | 97.27 | 59.71 |
| DDRNet-23-slim [29] | 95.03 | 95.03 | 99.53 | 90.54 | 99.71 | 96.00 | 99.55 | 97.86 | 89.56 |
| PIDNet-M [30] | 94.98 | 94.98 | 99.52 | 90.44 | 99.66 | 97.00 | 99.54 | 98.33 | 38.77 |
| PIDNet-S [30] | 94.96 | 94.96 | 99.51 | 90.40 | 99.65 | 97.10 | 99.53 | 98.37 | 76.84 |
| RailDLA | 96.50 | 96.50 | 99.70 | 92.10 | 99.90 | 97.50 | 99.70 | 98.50 | 120.00 |
| ConvNeXtV2-Tiny [46] | 94.41 | 94.41 | 99.48 | 89.34 | 99.81 | 92.93 | 99.50 | 96.37 | – |
| FastSCNN [24] | 92.85 | 92.68 | 99.32 | 86.04 | 99.83 | 89.11 | 99.35 | 94.47 | 107.58 |
| SegFormer-B0 [42] | 89.76 | 86.36 | 98.76 | 73.97 | 99.90 | 75.54 | 98.80 | 87.72 | 16.33 |
| SAN-ViT-B16 [47] | 87.78 | 83.64 | 98.29 | 68.98 | 99.18 | 81.00 | 98.35 | 90.09 | 23.28 |
| EuPE-ViT-T16 [48] | 82.51 | 80.19 | 98.20 | 62.19 | 99.89 | 63.69 | 98.25 | 81.79 | – |
| Segmenter-ViT-T [41] | 77.69 | 64.82 | 96.88 | 32.75 | 99.95 | 33.08 | 96.93 | 66.52 | – |
| Model | Best mIoU ↑ | Final mIoU ↑ | Bg IoU ↑ | Track IoU ↑ | Track Acc. ↑ | mIoU ↑ | Track ↑ |
|---|---|---|---|---|---|---|---|
| Segmenter-ViT-T [41] | 77.69 | 64.82 | 96.88 | 32.75 | 33.08 | ||
| EuPE-ViT-T16 [48] | 82.51 | 80.19 | 98.20 | 62.19 | 63.69 | ||
| SAN-ViT-B16 [47] | 87.78 | 83.64 | 98.29 | 68.98 | 81.00 | ||
| SegFormer-B0 [42] | 89.76 | 86.36 | 98.76 | 73.97 | 75.54 | ||
| FastSCNN [25] | 92.85 | 92.68 | 99.32 | 86.04 | 89.11 | 0.00 | 0.00 |
| ConvNeXtV2-Tiny [46] | 94.41 | 94.41 | 99.48 | 89.34 | 92.93 | ||
| RailDLA | 96.50 | 96.50 | 99.70 | 92.10 | 97.50 | +3.65 | +6.06 |
| Variant | Track-Specific Component | Accuracy (%) | Complexity | |||||
|---|---|---|---|---|---|---|---|---|
| Local Geom. | Axis Decoder | Strip Attn. | mIoU ↑ | Track IoU ↑ | mIoU | Track | FLOPs (G) ↓ | |
| A: Plain lightweight variant | – | – | – | 92.85 | 86.04 | – | – | 0.927 |
| B: Track prior variant | ✓ | ✓ | – | 94.64 | 89.80 | 1.035 | ||
| C: RailDLA | ✓ | ✓ | ✓ | 96.50 | 92.10 | +1.86 | +2.30 | 1.044 |
| Variant | Geom. | TAP | Mod. | RDLA | Gate | mIoU ↑ | Track IoU ↑ | Track IoU |
|---|---|---|---|---|---|---|---|---|
| Full RailDLA | ✓ | ✓ | ✓ | ✓ | ✓ | 96.50 | 92.10 | – |
| w/o geom. loss | – | ✓ | ✓ | ✓ | ✓ | 95.96 | 91.32 | −0.78 |
| w/o TAP pathway | ✓ | – | – | ✓ | ✓ | 95.74 | 91.05 | −1.05 |
| w/o decoder mod. | ✓ | ✓ | – | ✓ | ✓ | 96.02 | 91.48 | −0.62 |
| w/o RDLA pathway | ✓ | ✓ | ✓ | – | – | 94.64 | 89.80 | −2.30 |
| Variant | H | V | Gate | Support | mIoU ↑ | Track IoU ↑ | Track IoU |
|---|---|---|---|---|---|---|---|
| Full H/V RDLA | ✓ | ✓ | ✓ | H + V | 96.50 | 92.10 | – |
| H-only | ✓ | – | ✓ | H | 95.36 | 90.64 | −1.46 |
| V-only | – | ✓ | ✓ | V | 95.18 | 90.32 | −1.78 |
| w/o gate | ✓ | ✓ | – | H + V | 95.74 | 91.12 | −0.98 |
| Model | Fold 1 | Fold 2 | Fold 3 | Fold 4 | Fold 5 | Mean ± Std |
|---|---|---|---|---|---|---|
| RailDLA | 96.42/91.94 | 96.57/92.24 | 96.48/92.05 | 96.36/91.82 | 96.55/92.18 | 96.48 ± 0.09/92.05 ± 0.17 |
| SegNeXt-S | 95.58/91.62 | 95.73/91.94 | 95.66/91.82 | 95.55/91.66 | 95.83/92.11 | 95.67 ± 0.11/91.83 ± 0.20 |
| RailDLA − SegNeXt-S | 0.84/0.32 | 0.84/0.30 | 0.82/0.23 | 0.81/0.16 | 0.72/0.07 | 0.81 ± 0.05/0.22 ± 0.10 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Published by MDPI on behalf of the International Society for Photogrammetry and Remote Sensing. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, H.; Ma, B.; Gong, J.; Jiang, D.; Yang, J.; Fan, K.; Chen, Z. A Direction-Aware Lightweight Network for Camera-Based Underground Mine Track Region Segmentation. ISPRS Int. J. Geo-Inf. 2026, 15, 351. https://doi.org/10.3390/ijgi15080351
Li H, Ma B, Gong J, Jiang D, Yang J, Fan K, Chen Z. A Direction-Aware Lightweight Network for Camera-Based Underground Mine Track Region Segmentation. ISPRS International Journal of Geo-Information. 2026; 15(8):351. https://doi.org/10.3390/ijgi15080351
Chicago/Turabian StyleLi, Haijun, Baolong Ma, Jianjun Gong, Dengyin Jiang, Jie Yang, Kuangang Fan, and Zhichao Chen. 2026. "A Direction-Aware Lightweight Network for Camera-Based Underground Mine Track Region Segmentation" ISPRS International Journal of Geo-Information 15, no. 8: 351. https://doi.org/10.3390/ijgi15080351
APA StyleLi, H., Ma, B., Gong, J., Jiang, D., Yang, J., Fan, K., & Chen, Z. (2026). A Direction-Aware Lightweight Network for Camera-Based Underground Mine Track Region Segmentation. ISPRS International Journal of Geo-Information, 15(8), 351. https://doi.org/10.3390/ijgi15080351

