Abstract
Automatic row guidance is important for efficient and low-loss corn harvesting. However, complex field conditions challenge reliable inter-row perception, while many deep learning models remain computationally demanding. To address this, a lightweight semantic segmentation network is proposed. A dataset covering challenging field conditions was constructed. Based on DeepLabV3+, GhostNetV2 was adopted to reduce computational cost, an SP-ASPP was designed to enhance the representation of elongated inter-row structures, and BiFormer was introduced to strengthen long-range contextual modeling. We jointly exploited directional multi-scale context and sparse long-range interactions to preserve continuous row-space structures under occlusion and background interference while maintaining a lightweight architecture. A composite loss combining Focal, Dice, and Boundary losses was further employed to improve region completeness and boundary localization. A navigation-line extraction algorithm was then developed to generate stable guidance paths. After structured pruning and TensorRT FP16 optimization, the model achieved an mIoU of 82.28% at 53.1 FPS on the edge platform. The extracted navigation line yielded a mean absolute lateral error of 4.2 cm and a mean absolute heading error of 2.17°. These results demonstrate that the proposed method provides accurate real-time navigation perception on resource-constrained hardware, supporting low-cost vision-based corn harvester guidance.
1. Introduction
Corn is one of the world’s major crops for food, feed, and industrial raw materials. Global corn production exceeds 1.2 billion tonnes, with the United States, China, Brazil, and Argentina as major producers [1]. Corn plays an important role in food security, livestock production, and bioenergy [2]. With the continued advancement of agricultural mechanization and intelligent technologies, automatic row guidance for corn harvesters has become an important research topic for improving harvesting efficiency, reducing operator workload, and minimizing crop losses [3,4]. During field operation, the harvester must travel steadily along the crop-row direction, as lateral and heading deviations directly affect the uniformity of crop feeding into the header, operational continuity, and harvesting losses [5]. Therefore, the accurate perception of the inter-row structure and the reliable extraction of traversable regions are essential for automatic row guidance of corn harvesters.
Corn harvesting environments are highly unstructured and dynamically varying, posing substantial challenges to automatic row-guidance systems [6]. During the middle and late growth stages, inter-row soil regions are frequently affected by leaf occlusion, weed intrusion, straw coverage, illumination variation, missing plants, structural discontinuities, and lodging. Early studies commonly employed contact-based sensors because of their simple structures and direct responses; however, their sparse structural information limits robustness when crop rows are discontinuous [7,8,9]. GNSS/RTK can provide high-precision global positioning, but cannot directly perceive local crop-row deformation or unexpected structural changes [10,11]. LiDAR can acquire accurate point-cloud information describing row geometry, but system cost and point-cloud processing complexity restrict low-cost deployment. [12,13,14]. Nevertheless, conventional image processing methods rely heavily on handcrafted features and rule-based segmentation, resulting in limited adaptability to illumination variations, occlusion, and weed interference.
The development of deep learning provides a new approach to perception in complex cornfield environments [15,16,17]. Convolutional neural networks (CNNs) can learn hierarchical features and exhibit greater robustness to illumination variations and background interference, while semantic segmentation enables pixel-level classification to obtain detailed crop or traversable regions for subsequent navigation-line extraction [18,19]. However, most existing studies have primarily focused either on improving segmentation accuracy with increasingly complex networks or on reducing computational cost through lightweight architectures [20]. Comparatively less attention has been paid to preserving the geometric continuity of elongated inter-row traversable regions when substantial portions of the row structure are obscured by leaves, weeds, or lodging. Lightweight CNNs tend to emphasize local features and may lose global structural cues, whereas models with stronger long-range modeling commonly increase computational burden. Consequently, a clear research gap remains in simultaneously achieving lightweight computation, long-range structural perception, and boundary-preserving segmentation for real-time corn-harvester guidance.
To address this research gap, this study proposes a lightweight structure-aware semantic segmentation network for the automatic row guidance of corn harvesters. DeepLabV3+ is adopted as the baseline framework because of its strong segmentation capability [21,22,23]. To improve computational efficiency, GhostNetV2 is introduced as the backbone, and an SP-ASPP module is designed to capture directional multi-scale features of elongated inter-row structures. In addition, BiFormer is incorporated to model sparse long-range dependencies under occlusion and structural discontinuities. A composite loss consisting of Dice, Focal, and Boundary losses is employed to improve region completeness and boundary localization. The traversable-area mask generated by the segmentation network is subsequently processed by a navigation-line extraction algorithm to obtain the guidance path. Figure 1 illustrates the overall workflow of the proposed automatic row-guidance system.
Figure 1.
Workflow of the automatic row guidance system.
First, a dataset of inter-row traversable regions in cornfields was constructed for model training and validation. The proposed model was then trained and comprehensively evaluated, followed by validation of its potential for edge-device deployment. The main contributions of this study are as follows:
- •
- A dataset of inter-row traversable regions covering the middle and late growth stages of corn was constructed.
- •
- A lightweight structure-aware semantic segmentation network based on the DeepLabV3+ framework was proposed for automatic row guidance of corn harvest.
- •
- A navigation-line extraction algorithm was developed to address row occlusion, missing plants, and structural variations in cornfields.
- •
- Edge-platform deployment and navigation-line extraction experiments were conducted, providing a reference for related research and practical applications.
The remainder of this paper is organized as follows. Section 2 describes dataset construction, network architecture, and the navigation-line extraction algorithm. Section 3 presents the evaluation metrics and deployment experiments on the onboard computing platform. Section 4 discusses the experimental results and the proposed design. Finally, Section 5 summarizes the main conclusions of this study.
2. Materials and Methods
2.1. Data Acquisition and Partitioning
Field images were collected for the automatic row guidance task of corn harvesters. A preliminary acquisition experiment was first conducted at camera heights of 0.5, 1.0, 1.5, 2.0, 2.5, and 3.0 m to evaluate crop-row visibility, inter-row coverage, and field of view. Based on these observations, 1.5, 2.0, and 2.5 m were selected as the primary acquisition heights. Data were collected using a handheld Intel RealSense D435i camera while walking at a speed of approximately 4 km/h. To simulate platform vibration and attitude variation during harvester operation, the camera was intentionally rotated within approximately ±10° in pitch and ±30° in yaw during acquisition. Images were recorded in corn fields in Jurong City, Jiangsu Province, China on 28 July 2025, and in Nanjing City, Jiangsu Province, China on 10 October 2025. The dataset covered representative field conditions, including normal row structures, inter-row ground, leaf occlusion, weed interference, missing plants, crop lodging, strong illumination and shadows, and row-end scenes. These variations improved the diversity of the collected dataset, as illustrated in Figure 2.
Figure 2.
Examples of the multi-scenario dataset.
The raw data were recorded as videos. One frame was extracted every 30 frames. After removing blurred, poorly exposed, and irrelevant samples, 1240 valid images were retained. To prevent information leakage between temporally adjacent frames, the dataset was partitioned at the video-sequence level, with frames from the same continuous sequence assigned exclusively to one subset. The dataset was divided into training, validation, and test sets at an approximate ratio of 7:1.5:1.5, containing 868, 186, and 186 images, respectively. Scene-balanced partitioning was adopted to maintain representative field conditions across the subsets.
2.2. Data Annotation and Augmentation
Pixel-level annotations were generated using Labelme 5.9.1 with a task-oriented binary labeling strategy. Inter-row traversable regions suitable for harvester guidance were labeled as foreground, whereas corn plants, leaves, weeds, shadows, and other non-traversable regions were labeled as background. This formulation focuses the model on the geometry and continuity of the traversable corridor. To ensure annotation quality, all masks were independently reviewed after initial labeling, and ambiguous boundaries or inconsistent samples were cross-checked and corrected by another annotator. Representative annotations are shown in Figure 3.
Figure 3.
Examples of image annotations.
Data augmentation was applied only to the training set, while the validation and test sets were left unchanged. The augmentation operations included random rotations within −30° to 30°, horizontal flipping, brightness and contrast perturbations, and noise injection. Geometric transformations were applied identically to the images and their segmentation masks to preserve pixel-wise correspondence. After augmentation, the training set was expanded to five times its original size. Examples are shown in Figure 4.
Figure 4.
Examples of data augmentation.
2.3. Framework Design
To balance segmentation accuracy, structural continuity, and real-time performance for corn harvester row guidance, the proposed network was developed on the DeepLabV3+ framework, as shown in Figure 5. The original Xception backbone provides strong feature representation but incurs relatively high computational cost [24]. Lightweight alternatives such as MobileNetV2 reduce model complexity but may sacrifice accuracy under occlusion and texture ambiguity [25]. Therefore, GhostNetV2 was adopted as the backbone to generate intrinsic features with standard convolutions and additional feature maps through inexpensive linear operations, thereby reducing computation while preserving representation capability [26].
Figure 5.
Architecture of DeepLabV3+ [19].
Because inter-row traversable regions are typically elongated, continuous, and strongly directional, conventional ASPP provides limited modeling of long-range dependencies along a dominant direction. Horizontal and vertical strip-pooling branches were therefore introduced to form SP-ASPP [27], as illustrated in Figure 6. This module enhances directional context aggregation and improves the representation of partially occluded, discontinuous, and blurred inter-row structures.
Figure 6.
Illustration of strip pooling [25].
To further enhance global context modeling, a BiFormer module was inserted after SP-ASPP, as shown in Figure 7 and Figure 8. Its bi-level routing attention selects highly relevant regions and performs sparse feature interaction among them, enabling long-range dependency modeling with controlled computational overhead [28]. This sequential design follows a coarse-to-global reasoning strategy: SP-ASPP first aggregates directional and multi-scale contextual features, after which BiFormer models long-range dependencies on the semantically enriched representation. Placing BiFormer after multi-scale aggregation allows global reasoning on compact high-level features while avoiding the cost of applying attention to high-resolution shallow features.
Figure 7.
Proposed BiFormer module.
Figure 8.
Principle of bi-level routing attention [26].
The resulting network integrates lightweight feature extraction, directional multi-scale modeling, and sparse global context enhancement. This provides a more solid model foundation for the stable and accurate segmentation of passable areas under complex occlusion, missing plants, and changing lighting conditions. The complete architecture is shown in Figure 9.
Figure 9.
Architecture of the proposed segmentation network.
2.4. Composite Loss Function
In corn-row guidance scenes, traversable and background regions may be imbalanced, while inter-row boundaries are frequently corrupted by leaves, weeds, and shadows. Cross-entropy loss alone therefore provides insufficient supervision for region completeness and boundary localization.
A composite loss combining Focal, Dice, and Boundary losses was adopted:
emphasizes hard pixels and suppresses the contribution of easily classified samples, alleviates foreground–background imbalance and improves region overlap, and explicitly constrains boundary localization.
is the weight coefficient of each term.
To avoid excessive boundary emphasis and fragmented predictions, the three terms were weighted as follows:
The weighting strategy was determined according to the complementary roles of the three loss terms and preliminary experiments. Focal and Dice losses were assigned dominant weights to preserve region-level classification and overlap consistency, whereas Boundary loss was given a smaller weight because excessive boundary supervision was observed to produce fragmented edge predictions and reduce regional continuity. The resulting objective jointly optimizes difficult-sample recognition, region consistency, and boundary accuracy, improving the completeness and continuity of traversable-area segmentation under challenging field conditions.
2.5. Navigation-Line Extraction
A bottom-up region-tracking and robust fitting method was developed to extract navigation lines from the predicted traversable-area masks. Small connected components were first removed, and morphological closing was applied to bridge local discontinuities. Because the distant image region is strongly affected by perspective compression, missing plants, and inter-plant gaps, mask regions above 0.7 H were excluded. This threshold was selected from the validation set and then fixed for all test images.
The remaining area was scanned upward using overlapping horizontal bands. For the 1280 × 720 images, the band height and vertical step were set to 14 and 5 pixels, respectively, providing approximately 64% overlap. These values were chosen to retain sufficient regional support while reducing missed narrow regions and excessive discretization. Candidate center points were selected according to trajectory continuity, region width, pixel density, and boundary proximity. When missing plants caused the mask to widen, the historical trajectory prediction was assigned greater weight to suppress center-point drift. Edge-connected interference regions inconsistent with the predicted trajectory were discarded or down-weighted. A robust straight line was first estimated using iteratively reweighted least squares with Huber weighting. The Huber threshold was defined as:
where 1.345 provides approximately 95% asymptotic efficiency for Gaussian residuals and 1.4826 is the normal-consistency correction factor for median absolute deviation (MAD). Samples satisfying retained unit weight, whereas larger residuals were assigned the weight:
Quadratic fitting was enabled only when at least 18 valid points were available and their longitudinal span exceeded 0.24 H. Its maximum lateral deviation from the robust straight-line reference was limited to , corresponding to 32 pixels at the adopted resolution. These curve-selection thresholds were treated as implementation-specific parameters, determined from the validation set and fixed throughout testing. Thus, curvature was constrained by an explicit geometric deviation threshold rather than an independent regularization coefficient.
This procedure suppresses the effects of distant missing-plant regions, edge-connected mask artifacts, and local outliers on navigation-line position and orientation. Lateral and heading deviations were then calculated at 0.25 H to characterize the current guidance state of the harvester.
3. Experiments and Results
3.1. Evaluation Metrics
Segmentation performance was evaluated using mean Intersection over Union (mIoU), mean Pixel Accuracy (mPA), and mean Precision (mPrecision):
where denotes the number of classes minus one, is the number of correctly classified pixels of class , is the total number of ground-truth pixels belonging to class , is the total number of pixels predicted as class . was used as the primary segmentation metric, while and reflected class-wise recall and prediction reliability.
Boundary F1 score () was used to evaluate boundary localization quality. Let and denote the predicted and ground-truth boundary pixels, respectively. A boundary pixel is considered correctly matched if a corresponding boundary exists within a tolerance distance . Boundary precision , boundary recall , and are calculated as:
where and denote the numbers of matched boundary pixels within distance . In this study, used a fixed tolerance of pixels to evaluate strict pixel-level boundary localization, whereas used:
To provide a resolution-normalized boundary measure for the 1280 × 720 images used in this study, the latter corresponds to approximately 11 pixels. Higher values indicate better agreement between predicted and ground-truth boundaries.
Model complexity and real-time performance were further assessed using the number of parameters, FLOPs, and inference speed in frames per second (FPS).
3.2. Training Platform and Settings
All models were trained on the platform listed in Table 1.
Table 1.
Training platform and software environment.
Input images were resized to 1280 × 720 pixels. Models were trained from scratch without pretrained weights. Adam was used as the optimizer with a maximum learning rate of 1 × 10−4 and a momentum parameter of 0.9, together with a cosine learning-rate schedule.
Data augmentation was applied only to the training set, while the validation set was used for convergence monitoring and model selection and the test set for final evaluation. mIoU was adopted as the principal criterion for segmentation performance.
3.3. Ablation Study
The original DeepLabV3+ with Xception was used as the standard baseline, while DeepLabV3+ with MobileNetV2 served as the lightweight baseline. The proposed components were then introduced sequentially to quantify their individual contributions, as summarized in Table 2.
Table 2.
Ablation-study configurations.
Table 3.
Ablation-study results.
Figure 10.
Qualitative segmentation results of experiments A–F.
The results show that backbone substantially affected both accuracy and efficiency, with GhostNetV2 providing the best trade-off. SP-ASPP only slightly reduced inference speed but better maintained the vertical continuity of inter-row regions. Adding BiFormer improved long-range feature interaction and suppressed background interference while maintaining real-time performance. The composite loss yielded an additional improvement in segmentation quality without increasing inference cost. Boundary evaluation further confirms this improvement: the final model achieved BF@3px and BF@0.75% diagonal scores of 21.31% and 63.59%, respectively, compared with 18.02% and 59.93% for the baseline. Overall, the ablation study confirms the effectiveness and complementarity of the proposed modifications.
3.4. Comparison with Other Models
To further evaluate the overall performance of the proposed method, we compared our model with two classic lightweight models, ResNet-FCN [29] and U-Net [30]. The results are presented in Table 4.
Table 4.
Comparison with other models.
Our model achieved the highest mIoU, mPA, mPrecision and F1@0.75%. Although its inference speed was lower than that of U-Net, it still satisfied real-time requirements and provided a better overall balance among segmentation accuracy and inference efficiency.
3.5. Nav-Line Error Analysis
To quantify the propagation of segmentation errors to downstream navigation, the same navigation-line extraction procedure was applied to both the predicted masks and the manually annotated masks. The navigation line derived from the manual annotation was used as the Ground Truth (GT), and the line extracted from the predicted was evaluated against it.
To evaluate the navigation error at the vertical position , let the horizontal axis of the navigation line be and the width of the image be . The lateral deviation is defined as:
The heading error is defined as:
To provide a physical interpretation of the lateral error, the pixel-level error was converted to the corresponding ground-plane distance using the calibrated camera intrinsics and the pinhole camera model. Assuming a locally planar ground surface and a level camera, the horizontal ground-plane scale at is given by:
where is the camera height above the ground in meters, and are the horizontal and vertical focal lengths in pixels, respectively, is the vertical coordinate of the principal point, and is the pitch angle of the camera relative to the horizontal plane (downwards is positive). The factor 100 converts the unit from meters to centimeters.
Accordingly, the lateral error in physical units is calculated as:
Table 5.
Navigation-line extraction errors.
Figure 11.
Examples of navigation-line extraction.
The predicted navigation lines were generally close to the reference lines in both position and orientation, indicating a limited propagation of segmentation errors to navigation estimates in most cases. The error distributions nevertheless exhibited a long tail, with a small number of difficult samples producing larger deviations. These extreme errors could be further reduced through navigation-line outlier detection and temporal smoothing across consecutive frames.
3.6. Model Pruning and Edge Deployment
To reduce computational cost, structured channel pruning was performed using the dependency-graph mechanism in Torch-Pruning 1.6.1, which physically removes channels and therefore enables hardware-level acceleration. R10, R20, and R30 denote models with nominal channel pruning ratios of 10%, 20%, and 30%, respectively; the actual reductions in parameters and FLOPs depend on the network structure and dependency constraints. The pruning results are reported in Table 6.
Table 6.
Structured pruning results.
Although pruning progressively reduced computational complexity, the corresponding speedup was modest because the network had already been designed with lightweight components. R10 provided the best trade-off between segmentation accuracy and computational efficiency.
Edge deployment was evaluated on an onboard computing platform consisting of an Intel N150 processor, 16 GB RAM, and an NVIDIA RTX PRO 2000, as shown in Figure 12. The platform had a peak power consumption of 80 W and was used for image preprocessing, traversable-area segmentation, and navigation-line extraction. The total hardware cost is similar to the cost range of edge systems based on Jetson Orin NX 16 GB. It provided sufficient computing space for current lightweight models and subsequent multitasking visual perception research.
Figure 12.
Onboard edge computing platform.
Real-time performance was evaluated using PyTorch FP16, CUDA Graph, and TensorRT FP16 inference, as summarized in Table 7. The reported latency represents the end-to-end perception time, including image memory copying, host-to-device (H2D) transfer and normalization, model inference, GPU-based argmax, and mask transfer back to host memory.
Table 7.
Edge inference performance.
TensorRT FP16 provided the largest acceleration, reducing latency to 18.83 ms and increasing the throughput of the R10 model to 53.10 FPS while maintaining near-identical pixel-level outputs. These results demonstrate that the proposed method can satisfy real-time processing requirements on the onboard edge platform while retaining high segmentation accuracy.
4. Discussion
Comparative experiments demonstrated that the proposed lightweight structure-aware segmentation network achieves a favorable balance between segmentation accuracy and computational efficiency for inter-row traversable-area perception. After R10 structured pruning and TensorRT FP16 optimization, the model achieved 53.1 FPS on the onboard edge platform while retaining an mIoU of 82.28%. Navigation-line evaluation yielded a mean absolute lateral error of 4.2 cm and a P95 error of 8.2 cm, together with a mean absolute heading error of 2.17° and a P95 error of 5.00°. Compared with previously reported vision-based row-perception methods, which often emphasize either segmentation accuracy or real-time efficiency, the proposed method maintains competitive segmentation performance while providing direct navigation-line outputs at real-time speed on resource-constrained hardware. In particular, the combination of directional contextual modeling and lightweight long-range interaction improves the continuity of traversable-region prediction under occlusion and discontinuous crop rows.
Several limitations remain. The dataset is limited in region, corn variety, growth stage, and environmental diversity, so generalization to different fields, seasons, lighting conditions, lodging, and lens contamination still needs validation. Typical failure cases were observed under severe lodging, where the original row structure became ambiguous, and under strong local glare or heavy leaf occlusion, where portions of the traversable region could be incorrectly fragmented or shifted. Future work should use external test sets and lightweight temporal fusion or state filtering to reduce frame-level fluctuations from occlusion and motion blur. Camera calibration, inverse perspective mapping, and coordinate transformation are required to convert pixel outputs into lateral and heading errors. Importantly, the present study included offline algorithm evaluation and inference tests on the onboard edge-computing platform but did not include closed-loop field experiments on an operating corn harvester. Therefore, the reported lateral and heading errors characterize the perception and navigation-line extraction stages rather than the final vehicle-tracking performance. Finally, the perception module should be integrated with controllers such as Pure Pursuit, Stanley, or MPC and evaluated in field trials for tracking accuracy, control stability, and operational reliability.
5. Conclusions
This study proposed a lightweight structure-aware semantic segmentation framework for the automatic row guidance of corn harvesters under challenging conditions such as leaf occlusion, weed interference, missing plants, and discontinuous crop rows. The main contributions are as follows:
- •
- An inter-row traversable-area dataset was constructed for corn-harvester row guidance, covering multiple camera heights and challenging conditions including weeds, leaf occlusion, missing plants, discontinuous rows, and complex illumination.
- •
- A lightweight structure-aware segmentation network was developed to balance segmentation accuracy and inference efficiency for inter-row perception, achieving an mIoU of 82.28% after R10 structured pruning.
- •
- A navigation-line extraction method was designed to suppress the effects of local segmentation noise, boundary fluctuations, and region discontinuities, yielding a mean absolute lateral error of 4.2 cm and a mean absolute heading error of 2.17°.
- •
- Structured pruning and onboard edge deployment were validated, with TensorRT FP16 inference reaching 53.1 FPS on the onboard edge platform, demonstrating real-time deployment capability.
Although the proposed method achieved promising performance, its robustness under more diverse field conditions remains to be established. The present study was limited to offline algorithm evaluation and onboard hardware testing without closed-loop field validation on an operating corn harvester. Future work will expand the dataset across regions, growth stages, and environmental conditions and will investigate camera calibration, vehicle-coordinate transformation, and closed-loop field experiments to further assess the practical applicability of the proposed system.
Author Contributions
Conceptualization, S.C. and B.Z.; methodology, S.C., F.K. and K.T.; software, S.C. and Z.M.; validation, S.C. and B.Z.; formal analysis, S.C. and B.Z.; investigation, S.C. and Z.M.; resources, S.C., F.K., K.T. and Y.S.; data curation, S.C., F.K. and Z.M.; writing—original draft preparation, S.C.; writing—review and editing, S.C. and Z.M.; visualization, S.C. and Z.M.; supervision, B.Z. and Z.M.; project administration, B.Z. and Z.M.; funding acquisition, B.Z. and Z.M. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by the Innovation Project of the Chinese Academy of Agricultural Sciences (31-NIAM-05) and the Youth Development of Innovation Project of the Chinese Academy of Agricultural Sciences (1102372100050080003).
Data Availability Statement
The data presented in this study will be made available on request.
Acknowledgments
The authors thank the editor and anonymous reviewers for providing helpful suggestions for improving the quality of this manuscript.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- World Agricultural Production. Available online: https://www.fas.usda.gov/sites/default/files/2025-05/production.pdf (accessed on 18 September 2026).
- OECD-FAO Agricultural Outlook 2025–2034. Available online: https://www.oecd.org/content/dam/oecd/en/publications/reports/2025/07/oecd-fao-agricultural-outlook-2025-2034_3eb15914/601276cd-en.pdf (accessed on 18 September 2026).
- Li, B.; Li, D.; Wei, Z.; Wang, J. Rethinking the crop row detection pipeline: An end-to-end method for crop row detection based on row-column attention. Comput. Electron. Agric. 2024, 225, 109264. [Google Scholar] [CrossRef] [Scilit]
- Yao, Z.; Zhao, C.; Zhang, T. Agricultural machinery automatic navigation technology. iScience 2024, 27, 108714. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huang, Y.; Jiang, Z.; Li, J.; Guo, G.; Zhang, J. Design on automatic row guidance system for corn harvester header based on improved PSO-PID. Mechatronics 2026, 116, 103490. [Google Scholar] [CrossRef] [Scilit]
- Cui, S.; Kong, F.; Tian, K.; Xie, Q.; Liu, Y.; Zhang, B.; Mu, Z. Review of automatic row alignment technology for intelligent agricultural machinery in the field. Smart Agric. Technol. 2026, 14, 102059. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Lv, S.; Pan, Y.; Wang, H.; Yang, R.; Zhao, Y.; Zhao, A.; Liu, Y. Auto-follow row control system for an autonomous maize harvester using adaptive segmented PID and row-deviation detection. Robot. Auton. Syst. 2026, 196, 105252. [Google Scholar] [CrossRef] [Scilit]
- Geng, A.; Hu, X.; Liu, J.; Mei, Z.; Zhang, Z.; Yu, W. Development and Testing of Automatic Row Alignment System for Corn Harvesters. Appl. Sci. 2022, 12, 6221. [Google Scholar] [CrossRef] [Scilit]
- Zhang, K.; Hu, Y.; Yang, L.; Zhang, D.; Cui, T.; Fan, L. Design and Experiment of Auto-follow Row System for Corn Harvester. Trans. Chin. Soc. Agric. Mach. 2020, 51, 103–114. [Google Scholar] [CrossRef]
- Hernández-Pajares, M.; Olivares-Pulido, G.; Graffigna, V.; García-Rigo, A.; Lyu, H.; Roma-Dollase, D.; de Lacy, M.C.; Fernández-Prades, C.; Arribas, J.; Majoral, M.; et al. Wide-Area GNSS Corrections for Precise Positioning and Navigation in Agriculture. Remote Sens. 2022, 14, 3845. [Google Scholar] [CrossRef] [Scilit]
- Yu, J.; Fang, H.; Zhang, X.; Wu, W.; He, Y. Tightly coupled GNSS/IMU/vision integrated system for positioning in agricultural scenarios. Comput. Electron. Agric. 2025, 239, 110478. [Google Scholar] [CrossRef] [Scilit]
- Zhang, B.; Xu, H.; Tian, K.; Huang, J.; Kong, F.; Mu, S.; Wu, T.; Mu, Z.; Wang, X.; Zhou, D. Research on Automatic Alignment for Corn Harvesting Based on Euclidean Clustering and K-Means Clustering. Agriculture 2024, 14, 2071. [Google Scholar] [CrossRef] [Scilit]
- Ban, C.; Wang, L.; Chi, R.; Su, T.; Ma, Y. A Camera-LiDAR-IMU fusion method for real-time extraction of navigation line between maize field rows. Comput. Electron. Agric. 2024, 223, 109114. [Google Scholar] [CrossRef] [Scilit]
- Biglia, A.; Zaman, S.; Gay, P.; Aimonino, D.R.; Comba, L. 3D point cloud density-based segmentation for vine rows detection and localisation. Comput. Electron. Agric. 2022, 199, 107166. [Google Scholar] [CrossRef] [Scilit]
- Hossain, M.S.; Rahman, M.; Rahman, A.; Kabir, M.M.; Mridha, M.F.; Huang, J.; Shin, J. Automatic Navigation and Self-Driving Technology in Agricultural Machinery: A State-of-the-Art Systematic Review. IEEE Access 2025, 13, 94370–94401. [Google Scholar] [CrossRef] [Scilit]
- Bai, Y.; Zhang, B.; Xu, N.; Zhou, J.; Shi, J.; Diao, Z. Vision-based navigation and guidance for agricultural autonomous vehicles and robots: A review. Comput. Electron. Agric. 2023, 205, 107584. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Wang, X.; Chen, X.; Zhang, Y.; Zhang, Y. Review on Key Technologies for Autonomous Navigation in Field Agricultural Machinery. Agriculture 2025, 15, 1297. [Google Scholar] [CrossRef] [Scilit]
- Liang, Z.; Zhou, J.; Chen, Y.; Zhang, Y.; Gemechu, T.T.; Li, L.; Zhou, H.; Aurangzaib, M. Semantic segmentation–based detection of exposed soil regions in paddy fields for a floating-type puddling and leveling operation. Comput. Electron. Agric. 2026, 244, 111494. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Jin, C.; Chen, M.; Cai, Z.; Liu, Z. Wheat Full-Width harvesting navigation line extraction method using improved Swin-Transformer. Comput. Electron. Agric. 2025, 239, 110881. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv 2021. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.-C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. arXiv 2018. [Google Scholar] [CrossRef] [Scilit]
- Yu, J.; Zhang, J.; Shu, A.; Chen, Y.; Chen, J.; Yang, Y.; Tang, W.; Zhang, Y. Study of convolutional neural network-based semantic segmentation methods on edge intelligence devices for field agricultural robot navigation line extraction. Comput. Electron. Agric. 2023, 209, 107811. [Google Scholar] [CrossRef] [Scilit]
- Nkwocha, C.L.; Wang, N. Deep learning-based semantic segmentation with novel navigation line extraction for autonomous agricultural robots. Discov. Artif. Intell. 2025, 5, 73. [Google Scholar] [CrossRef] [Scilit]
- Chollet, F. Xception: Deep Learning with Depthwise Separable Convolutions. arXiv 2017. [Google Scholar] [CrossRef] [Scilit]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.-C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. arXiv 2019. [Google Scholar] [CrossRef] [Scilit]
- Tang, Y.; Han, K.; Guo, J.; Xu, C.; Xu, C.; Wang, Y. GhostNetV2: Enhance Cheap Operation with Long-Range Attention. arXiv 2022. [Google Scholar] [CrossRef] [Scilit]
- Hou, Q.; Zhang, L.; Cheng, M.-M.; Feng, J. Strip Pooling: Rethinking Spatial Pooling for Scene Parsing. arXiv 2020. [Google Scholar] [CrossRef] [Scilit]
- Zhu, L.; Wang, X.; Ke, Z.; Zhang, W.; Lau, R. BiFormer: Vision Transformer with Bi-Level Routing Attention. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]
- Wu, W.; Huo, L.; Yang, G.; Liu, X.; Li, H. Research into the Application of ResNet in Soil: A Review. Agriculture 2025, 15, 661. [Google Scholar] [CrossRef] [Scilit]
- Gkologkinas, G.D.; Ntouros, K.; Protopapadakis, E.; Rallis, I. A Comparative Analysis of U-Net Architectures with Dimensionality Reduction for Agricultural Crop Classification Using Hyperspectral Data. Algorithms 2025, 18, 588. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.











