Figure 1.
Study area overview. Upper left: location of Jiangsu Province within China, with Jiangsu highlighted. Upper right: location of Suzhou City within Jiangsu Province, with Suzhou highlighted. Lower left: location of Changshu City within Suzhou City, with Changshu highlighted. Lower right: true-color composite of Changshu City showing the highly fragmented parcel structure typical of the region.
Figure 1.
Study area overview. Upper left: location of Jiangsu Province within China, with Jiangsu highlighted. Upper right: location of Suzhou City within Jiangsu Province, with Suzhou highlighted. Lower left: location of Changshu City within Suzhou City, with Changshu highlighted. Lower right: true-color composite of Changshu City showing the highly fragmented parcel structure typical of the region.
Figure 2.
Crop phenology calendar of the four cropping patterns in the study area, covering the observation period from October 2023 to October 2024. Growth stages are color-coded as indicated in the legend.
Figure 2.
Crop phenology calendar of the four cropping patterns in the study area, covering the observation period from October 2023 to October 2024. Growth stages are color-coded as indicated in the legend.
Figure 3.
Box plot of equivalent circular radius distribution for crop parcels across four cropping pattern classes, illustrating the predominance of small parcels in the study area.
Figure 3.
Box plot of equivalent circular radius distribution for crop parcels across four cropping pattern classes, illustrating the predominance of small parcels in the study area.
Figure 4.
High-resolution imagery of the study area overlaid with parcel vector boundaries (red outlines) generated in ArcMap 10.8 (Esri, Redlands, CA, USA), demonstrating the highly fragmented agricultural landscape.
Figure 4.
High-resolution imagery of the study area overlaid with parcel vector boundaries (red outlines) generated in ArcMap 10.8 (Esri, Redlands, CA, USA), demonstrating the highly fragmented agricultural landscape.
Figure 5.
Observation timestamp sequences of Sentinel-1 (35 time steps, upper) and Sentinel-2 (19 time steps, lower) from October 2023 to December 2024, illustrating the asynchronous and irregular sampling intervals inherent to multi-source satellite observations.
Figure 5.
Observation timestamp sequences of Sentinel-1 (35 time steps, upper) and Sentinel-2 (19 time steps, lower) from October 2023 to December 2024, illustrating the asynchronous and irregular sampling intervals inherent to multi-source satellite observations.
Figure 6.
Overall workflow of the proposed PAST framework, comprising three stages: (1) K-Shape-based label quality control, (2) parallel dual-branch classification using the temporal branch (DualTransformer-MLP) and the image branch (ConvFormer-CT), and (3) decision-level fusion for final parcel-level crop mapping.
Figure 6.
Overall workflow of the proposed PAST framework, comprising three stages: (1) K-Shape-based label quality control, (2) parallel dual-branch classification using the temporal branch (DualTransformer-MLP) and the image branch (ConvFormer-CT), and (3) decision-level fusion for final parcel-level crop mapping.
Figure 7.
Sensitivity analysis of the number of clusters k in K-Shape-based label quality control, evaluated across four crop classes using silhouette coefficient (blue, left axis; higher is better) and Davies–Bouldin index (red, right axis; lower is better). The selected k = 40 is marked by a vertical dashed line in each panel.
Figure 7.
Sensitivity analysis of the number of clusters k in K-Shape-based label quality control, evaluated across four crop classes using silhouette coefficient (blue, left axis; higher is better) and Davies–Bouldin index (red, right axis; lower is better). The selected k = 40 is marked by a vertical dashed line in each panel.
Figure 8.
Architecture of the temporal classification model (DualTransformer-MLP, DT). The model encodes asynchronous Sentinel-1 and Sentinel-2 time series via independent Transformer branches with sine–cosine positional encoding, followed by triple pooling (max/min/avg), MLP refinement, and cross-modal fusion for classification. Arrows indicate the data flow between modules, colors distinguish different functional modules, and ellipses denote omitted intermediate timestamps or features.
Figure 8.
Architecture of the temporal classification model (DualTransformer-MLP, DT). The model encodes asynchronous Sentinel-1 and Sentinel-2 time series via independent Transformer branches with sine–cosine positional encoding, followed by triple pooling (max/min/avg), MLP refinement, and cross-modal fusion for classification. Arrows indicate the data flow between modules, colors distinguish different functional modules, and ellipses denote omitted intermediate timestamps or features.
Figure 9.
Architecture of the image classification model (ConvFormer-CT). The model processes sub-meter resolution (0.8 m) high-resolution images from April and August through a shared ConvFormer_s18 backbone, with branch attention for adaptive temporal fusion and a planting pattern attention module for targeted enhancement of easily confused cropping pattern classes.
Figure 9.
Architecture of the image classification model (ConvFormer-CT). The model processes sub-meter resolution (0.8 m) high-resolution images from April and August through a shared ConvFormer_s18 backbone, with branch attention for adaptive temporal fusion and a planting pattern attention module for targeted enhancement of easily confused cropping pattern classes.
Figure 10.
Crop classification map of Changshu City generated by the proposed PAST framework, with three enlarged sub-regions illustrating parcel-level classification details in fragmented agricultural landscapes.
Figure 10.
Crop classification map of Changshu City generated by the proposed PAST framework, with three enlarged sub-regions illustrating parcel-level classification details in fragmented agricultural landscapes.
Figure 11.
Confusion matrices of the ablation experiments for three model configurations: full dual-branch model (left), time-series branch only (center), and image branch only (right).
Figure 11.
Confusion matrices of the ablation experiments for three model configurations: full dual-branch model (left), time-series branch only (center), and image branch only (right).
Table 1.
Features and calculation methods used in the study.
Table 1.
Features and calculation methods used in the study.
Satellite Platform | Feature | Formula | Resolution | Resampling |
|---|
| Sentinel-1 | VV | - | 10 m | - |
| VH | - | 10 m | - |
| Polarization Ratio | VV/VH | 10 m | - |
| Sentinel-2 | Blue | - | 10 m | - |
| Green | - | 10 m | - |
| Red | - | 10 m | - |
| Red Edge1 | - | 20 m → 10 m | Bilinear Interpolation |
| Red Edge 2 | - | 20 m → 10 m | Bilinear Interpolation |
| Red Edge 3 | - | 20 m → 10 m | Bilinear Interpolation |
| Near-Infrared | - | 10 m | - |
| Red Edge 4 | - | 20 m → 10 m | Bilinear Interpolation |
| SWIR (Shortwave Infrared) 1 | - | 20 m → 10 m | Bilinear Interpolation |
| SWIR 2 | - | 20 m → 10 m | Bilinear Interpolation |
| Jilin-1 02F | Blue | - | 0.8 m | - |
| Green | - | 0.8 m | - |
| Red | - | 0.8 m | - |
Table 2.
Sample quantity per cropping pattern.
Table 2.
Sample quantity per cropping pattern.
| Cropping Pattern | Original Count | After QC (Quality Control) | Final Dataset |
|---|
| Single Wheat | 3300 | 1100 | 1100 |
| Single Rice | 6000 | 1000 | 1000 |
| Rice–Wheat Rotation | 45,000 | 8600 | 1300 |
| Winter Oilseed Rape | 3300 | 1100 | 1100 |
| Total | 57,600 | 11,800 | 4500 |
Table 3.
Baseline model configurations.
Table 3.
Baseline model configurations.
| Model | Learning Rate | Batch Size | Optimizer | Epochs | Early-Stop Patience |
|---|
| LSTM | 5 × 10−3 | 68 | AdamW (wd 4 × 10−3) | 200 | 10 |
| TCN | 5 × 10−3 | 68 | AdamW (wd 4 × 10−3) | 200 | 10 |
| PatchTST | 5 × 10−3 | 68 | AdamW (wd 4 × 10−3) | 200 | 10 |
| Rocket | - | - | - | - | - |
| InceptionTime | 5 × 10−3 | 68 | AdamW (wd 4 × 10−3) | 200 | 10 |
| CNN | 1 × 10−5 | 68 | AdamW (wd 5 × 10−4) | 200 | 10 |
| ViT | 1 × 10−5 | 256 | AdamW (wd 5 × 10−4) | 200 | 10 |
| ResNet50 | 1 × 10−5 | 68 | AdamW (wd 5 × 10−4) | 200 | 10 |
Table 4.
Hyperparameter settings for both branches.
Table 4.
Hyperparameter settings for both branches.
| Hyperparameter | Time-Series Branch | Image Branch |
|---|
| Optimizer | AdamW | AdamW |
| Learning rate | 5 × 10−3 | 1 × 10−5 |
| Weight decay | 4 × 10−3 | 5 × 10−4 |
| Batch Size | 68 | 256 |
| Dropout | 0.3 | - |
| LR schedule | ReduceLROnPlateau | CosineAnnealingWarmRestarts |
| Max epochs | 200 | 200 |
| Early-stop patience | 10 | 15 |
| Input size | S1 seq = 35, S2 seq = 19 | 224 × 224 px |
| Data augmentation | - | Flip/Rot (±15°)/ColorJitter |
| Loss function | CrossEntropyLoss | Weighted CrossEntropyLoss |
Table 5.
Model efficiency of PAST, measured on a single NVIDIA GTX 1080 Ti GPU (NVIDIA Corporation, Santa Clara, CA, USA).
Table 5.
Model efficiency of PAST, measured on a single NVIDIA GTX 1080 Ti GPU (NVIDIA Corporation, Santa Clara, CA, USA).
| Component | Parameters | Trainable Params | Inference Latency (ms/sample) | Training Time |
|---|
| Temporal branch | 61,812 | 61,812 | 0.12 | 19.6 s |
| Image branch | 24,725,974 | 8,501,270 | 11.31 | 3.7 h |
| Full PAST | ≈24.79 M | ≈8.56 M | ≈11.43 | ≈3.7 h |
Table 6.
Time-series comparison of all parcels.
Table 6.
Time-series comparison of all parcels.
| Model | Precision | Recall | F1 |
|---|
| LSTM | 0.866 | 0.867 | 0.866 |
| TCN | 0.914 | 0.907 | 0.910 |
| Rocket | 0.903 | 0.899 | 0.901 |
| InceptionTime | 0.866 | 0.871 | 0.868 |
| PatchTST | 0.905 | 0.894 | 0.899 |
| Ours-Time Series Branch Only | 0.924 | 0.909 | 0.915 |
Table 7.
Comparison of time-series models for small parcels.
Table 7.
Comparison of time-series models for small parcels.
| Model | Precision | Recall | F1 |
|---|
| LSTM | 0.836 | 0.818 | 0.820 |
| TCN | 0.894 | 0.881 | 0.886 |
| Rocket | 0.872 | 0.870 | 0.870 |
| InceptionTime | 0.842 | 0.846 | 0.843 |
| PatchTST | 0.864 | 0.851 | 0.856 |
| Ours-Time Series Branch Only | 0.899 | 0.882 | 0.889 |
| Ours-Full Model | 0.921 | 0.898 | 0.906 |
Table 8.
Image comparison of all parcels.
Table 8.
Image comparison of all parcels.
| Model | Precision | Recall | F1 |
|---|
| CNN | 0.603 | 0.599 | 0.591 |
| ResNet50 | 0.507 | 0.470 | 0.450 |
| ViT | 0.686 | 0.680 | 0.678 |
| Ours-Image Branch Only | 0.724 | 0.718 | 0.719 |
Table 9.
Comparison of image models for small parcels.
Table 9.
Comparison of image models for small parcels.
| Model | Precision | Recall | F1 |
|---|
| CNN | 0.567 | 0.555 | 0.551 |
| ResNet50 | 0.531 | 0.450 | 0.441 |
| ViT | 0.658 | 0.630 | 0.636 |
| Ours-Image Branch Only | 0.681 | 0.666 | 0.670 |
| Ours-Full Model | 0.921 | 0.898 | 0.906 |
Table 10.
Comparison of different fusion strategies.
Table 10.
Comparison of different fusion strategies.
| Fusion Strategies | Precision | Recall | F1 |
|---|
| Three-modal Fusion | 0.833 | 0.810 | 0.814 |
| Decision-level Fusion | 0.936 | 0.919 | 0.926 |
Table 11.
Ablation experiment.
Table 11.
Ablation experiment.
| | Precision | Recall | F1 |
|---|
| Image Branch Only | 0.724 | 0.718 | 0.719 |
| Time-Series Branch Only | 0.924 | 0.909 | 0.915 |
| Full Dual-Branch Model | 0.936 | 0.919 | 0.926 |
Table 12.
Performance comparison with and without label quality control.
Table 12.
Performance comparison with and without label quality control.
| | Precision | Recall | F1 |
|---|
| w/o QC (Raw Labels) | 0.823 | 0.804 | 0.808 |
| w/QC (Ours) | 0.936 | 0.919 | 0.926 |