Target Identification via Multi-View Multi-Task Joint Sparse Representation
Abstract
1. Introduction
- (1)
- A multi-view multi-task method based on sparse representation, namely MVMT, is proposed for target identification. Compared with the previous related method, the proposed MVMT can not only use the learning based on sparse representation, but also introduces a variety of complementary perspective features to enhance the expression features of the targets;
- (2)
- Each perspective of each particle is regarded as a separate task, and the potential connections between different perspectives and different particles are considered in a multi-task learning framework;
- (3)
- To capture the outlier tasks that frequently occur in the particle sampling process, the coefficient matrix is decomposed into two cooperative parts to enhance the robustness of multi-task learning, and the posterior probability of outliers is set to zero to ensure that samples will not be taken in the resampling process.
2. Related Works
3. Problem Description and Modeling
4. Multiple Feature Extraction Methods for Target Identification
4.1. Scale-Invariant Feature Transformation (SIFT) Feature Extraction Method
4.2. Gradient Direction Histogram (HOG) Feature Extraction Method
4.3. Local Binary Pattern (LBP) Feature Extraction Method
4.4. Hu Invariant Matrix Feature Extraction Method
5. Target Identification Method Based on Multi-View Multi-Task Sparse Learning
5.1. Multi View Multi-Task Sparse Learning Target Identification Method
- (1)
- Initialize the target: Select the rectangular box of the target of interest on the first frame image;
- (2)
- Select the candidate template: Select the candidate template according to the Gaussian distribution model and process its gray image level.
- (3)
- Extract multi-view features: Extract the features of LBP, HOG, and SIFT view from the obtained candidate templates, and reduce the dimension of the extracted feature matrix;
- (4)
- Construct the redundant dictionary: Assemble the obtained single column vectors of each view of a given template into a new column vector, i.e., template vector, and cycle through other candidate templates. All the template vectors obtained are arranged in columns, and the identity matrix is added to form a redundant dictionary;
- (5)
- Multi-view multi-task sparse learning: Multi-view multi-task learning is performed on the templates in the redundant dictionary to solve the classification coefficient matrix C. For several particles selected by sampling around the target, select the candidate template with a small residual to replace the bad performance in the dictionary, and finally make the candidate template in the dictionary the optimal template set. For Equation (3), a gradient descent algorithm is proposed, so the target residual can be calculated as:
- (6)
- Update template: if the residual error between the training sample and the label is less than the given threshold, replace the corresponding template in the dictionary, as shown in Figure 10.
- (7)
- Determine target: Carry out probability conversion for each classification plane (each column) in matrix C and take the classification plane with the largest probability as the final target to be tracked. For the t-th frame image, the determination criteria are as follows:
- (8)
- Output box to calculate the target identification error.
5.2. Simplified MVMT Methods
6. Experiment and Result Analysis
6.1. Experimental Results and Analysis of Public Datasets
6.2. Experiment and Analysis of Custom Datasets
6.3. Qualitative Comparison and Analysis
6.4. Quantitative Comparison and Analysis
7. Conclusions
- (1)
- A target identification method based on machine learning is proposed. The traditional target identification method makes use of less information. In the case of many targets and complex backgrounds, tracking loss and error are common. To break through the limitations of traditional methods, we introduce a machine learning approach and provide a new idea for target identification by using sparse representation, dictionary learning method and particle filter framework;
- (2)
- By analyzing the features of targets in the surveillance video, the texture, edge, and geometry invariant features of targets are automatically extracted, and the multi-view feature learning method is introduced to realize the information fusion of multiple features. To obtain better results, a multi-task learning method is proposed based on the multi-view feature for joint learning, which makes full use of the information transfer function between different tasks, makes up for the deficiency of the traditional single learning task, improves the classification and recognition accuracy, and provides a theoretical basis for the accurate and stable identification.
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Tang, J.; Liu, F.; Zou, Y.; Zhang, W.; Wang, Y. An Improved Fuzzy Neural Network for Traffic Speed Prediction Considering Periodic Characteristic. IEEE Trans. Intell. Transp. Syst. 2017, 18, 2340–2350. [Google Scholar] [CrossRef] [Scilit]
- Yan, Y.; Zhang, S.; Tang, J.; Wang, X. Understanding characteristics in multivariate traffic flow time series from complex network structure. Phys. A Stat. Mech. Its Appl. 2017, 477, 149–160. [Google Scholar] [CrossRef] [Scilit]
- Chen, W.; Zhou, G.; Liu, Z.; Li, X.; Zheng, X.; Wang, L. NIGAN: A Framework for Mountain Road Extraction Integrating Remote Sensing Road-Scene Neighborhood Probability Enhancements and Improved Conditional Generative Adversarial Network. IEEE Trans. Geosci. Remote Sens. 2022, 60, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Chen, W.; Li, X.; Wang, L. Target Detection for Mine Remote Sensing Using Deep Learning. In Remote Sensing Intelligent Interpretation for Mine Geological Environment; Springer: Singapore, 2022; pp. 127–164. [Google Scholar]
- Tong, W.; Chen, W.; Han, W.; Li, X.; Wang, L. Channel-attention-based DenseNet network for remote sensing image scene classification. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2020, 13, 4121–4132. [Google Scholar] [CrossRef] [Scilit]
- Liu, M.; Hu, Q.; Wang, C.; Tian, T.; Chen, W. Daff-Net: Dual Attention Feature Fusion Network for Aircraft Detection in Remote Sensing Images. In Proceedings of the IEEE International Geoscience and Remote Sensing Symposium IGARSS, Brussels, Belgium, 11–16 July 2021; pp. 4196–4199. [Google Scholar]
- Ouyang, S.; Xu, J.; Chen, W.; Dong, Y.; Li, X.; Li, J. A Fine-Grained Genetic Landform Classification Network Based on Multimodal Feature-Extraction and Regional Geological Context. IEEE Trans. Geosci. Remote Sens. 2022, 60, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Craswell, N.; Mitra, B.; Yilmaz, E.; Campos, D.; Voorhees, E.M. Overview of the TREC 2019 deep learning track. arXiv 2020, arXiv:2003.07820. [Google Scholar]
- Zhang, S.; Tang, J.; Wang, H.; Wang, Y.; An, S. Revealing intra-urban travel patterns and service ranges from taxi trajectories. J. Transp. Geogr. 2017, 61, 72–86. [Google Scholar] [CrossRef] [Scilit]
- Kwon, J.; Lee, K.M. Visual tracking decomposition. In Proceedings of the IEEE Conference Computer Vision and Pattern Recognition (CVPR), San Francisco, CA, USA, 13–18 June 2010; pp. 1269–1276. [Google Scholar]
- Wang, Y.; Luo, X.; Ding, L.; Hu, S. Visual tracking via robust multi-task multi-feature joint sparse representation. Multimed. Tools Appl. 2018, 77, 31447–31467. [Google Scholar] [CrossRef] [Scilit]
- Lan, X.; Ye, M.; Zhang, S.; Zhou, H.; Yuen, P.C. Modality-correlation-aware sparse representation for RGB-infrared object tracking. Pattern Recognit. Lett. 2020, 130, 12–20. [Google Scholar] [CrossRef] [Scilit]
- Wan, M.; Gu, G.; Qian, W.; Ren, K.; Maldague, X.; Chen, Q. School of Electronic and Optical Engineering, Nanjing University of Science and Technology, Nanjing, China. Unmanned aerial vehicle video-based target tracking algorithm using sparse representation. IEEE Internet Things J. 2019, 6, 9689–9706. [Google Scholar] [CrossRef] [Scilit]
- Elfring, J.; Torta, E.; van de Molengraft, R. Particle filters: A hands-on tutorial. Sensors 2021, 21, 438. [Google Scholar] [CrossRef] [Scilit]
- Zhang, T.; Xu, C.; Yang, M.H. Learning multi-task correlation particle filters for visual tracking. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 41, 365–378. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nai, K.; Li, Z.; Gan, Y.; Wang, Q. Robust Visual Tracking via Multitask Sparse Correlation Filters Learning. IEEE Trans. Neural Netw. Learn. Syst. 2021, in press. [CrossRef] [Scilit] [PubMed]
- Gurkan, F.; Gunsel, B. Integration of regularized l1 tracking and instance segmentation for video object tracking. Neurocomputing 2021, 423, 284–300. [Google Scholar] [CrossRef] [Scilit]
- Javanmardi, M.; Qi, X. Robust structured multi-task multi-view sparse tracking. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), San Diego, CA, USA, 23–27 July 2018; pp. 1–6. [Google Scholar]
- Lu, R.; Liu, J.; Lian, S.; Zuo, X. Multi-view representation learning in multi-task scene. Neural Comput. Appl. 2020, 32, 10403–10422. [Google Scholar] [CrossRef] [Scilit]
- Yan, P.; Yao, S.; Zhu, Q.; Zhang, T.; Cui, W. Real-time detection and tracking of infrared small targets based on grid fast density peaks searching and improved KCF. Infrared Phys. Technol. 2022, 123, 104181. [Google Scholar] [CrossRef] [Scilit]
- Lukezic, A.; Vojir, T.; Zajc, L.C.; Matas, J.; Kristan, M. Discriminative correlation filter with channel and spatial reliability. In Proceedings of the IEEE Conference Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 4847–4856. [Google Scholar]
- Yuan, D.; Kang, W.; He, Z. Robust visual tracking with correlation filters and metric learning. Knowl. Based Syst. 2020, 195, 105697. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Li, R.; Lang, C.; Han, J. Task-wise attention guided part complementary learning for few-shot image classification. Sci. China Inf. Sci. 2021, 64, 120104. [Google Scholar] [CrossRef] [Scilit]
- Yang, Y.; Xing, W.; Zhang, S.; Gao, L.; Yu, Q.; Che, X.; Lu, W. Visual tracking with long-short term based correlation filter. IEEE Access 2020, 8, 20257–20269. [Google Scholar] [CrossRef] [Scilit]
- Kurani, A.; Doshi, P.; Vakharia, A.; Shah, M. A comprehensive comparative study of artificial neural network (ANN) and support vector machines (SVM) on stock forecasting. Ann. Data Sci. 2021, 1–26. [Google Scholar] [CrossRef] [Scilit]
- Abbass, M.Y.; Kwon, K.C.; Kim, N.; Abdelwahab, S.A.; El-Samie, F.E.A.; Khalaf, A.A.M. Visual tracking using convolutional features with sparse coding. Artif. Intell. Rev. 2021, 54, 3349–3360. [Google Scholar] [CrossRef] [Scilit]
- Bardow, P.; Davison, A.J.; Leutenegger, S. Simultaneous optical flow and intensity estimation from an event camera. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 884–892. [Google Scholar]
- Tripathi, R.P.; Ghosh, S.; Chandle, J.O. Tracking of object using optimal adaptive Kalman filter. In Proceedings of the IEEE International Conference on Engineering and Technology, Coimbatore, India, 17–18 March 2016; pp. 1128–1131. [Google Scholar]
- Zhang, S.; Yang, Y.; Zhang, M.; Mi, P. An Efficient Tracker via Multi-feature Adaptive Correlation Filter. J. Syst. Simul. 2022, 34, 1864. [Google Scholar]
- Weng, X.; Ivanovic, B.; Pavone, M. Mtp: Multi-hypothesis tracking and prediction for reduced error propagation. In Proceedings of the 33rd IEEE Intelligent Vehicles Symposium (IV), Aachen, Germany, 5–9 June 2022; pp. 1218–1225. [Google Scholar]
- Meshgi, K.; Oba, S.; Ishii, S. Efficient diverse ensemble for discriminative co-tracking. In Proceedings of the Conference Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 4814–4823. [Google Scholar]
- Wang, N.; Zhou, W.; Tian, Q.; Hong, R.; Wang, M.; Li, H. Multi-cue correlation filters for robust visual tracking. In Proceedings of the IEEE Conference Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 4844–4853. [Google Scholar]
- Jiao, L.; Wang, D.; Bai, Y.; Chen, P.; Liu, F. Deep learning in visual tracking: A review. IEEE Trans. Neural Netw. Learn. Syst. 2021. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dai, K.; Wang, D.; Lu, H.; Sun, C.; Li, J. Visual tracking via adaptive spatially-regularized correlation filters. In Proceedings of the IEEE Conference Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 4670–4679. [Google Scholar]
- He, A.; Luo, C.; Tian, X.; Zeng, W. A twofold siamese network for real-time object tracking. In Proceedings of the IEEE Conference Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 4834–4843. [Google Scholar]
- Gao, J.; Zhang, T.; Xu, C. Graph convolutional tracking. In Proceedings of the IEEE Conference Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 4649–4659. [Google Scholar]
- Dong, B.; Zhou, Y.; Hu, C.; Fu, K.; Chen, G. BCNet: Bidirectional collaboration network for edge-guided salient object detection. Neurocomputing 2021, 437, 58–71. [Google Scholar] [CrossRef] [Scilit]
- Wu, L.; Fang, S.; Ma, Y.; Fan, F.; Huang, J. Infrared small target detection based on gray intensity descent and local gradient watershed. Infrared Phys. Technol. 2022, 123, 104171. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Zhong, X. An object tracking method based on Mean Shift algorithm with HSV color space and texture features. Clust. Comput. 2019, 22, 6079–6090. [Google Scholar] [CrossRef] [Scilit]
- Humeau-Heurtier, A. Texture feature extraction methods: A. survey. IEEE Access 2019, 7, 8975–9000. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Chen, H.; Wu, H.; Huang, Y.; Yang, Y.; Zhang, W.; Xiong, P. Robust visual ship tracking with an ensemble framework via multi-view learning and wavelet filter. Sensors 2020, 20, 932. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mao, J.L. Adaptive multi-view learning and its application in image classification. J. Comput. Appl. 2013, 33, 1955–1959. [Google Scholar] [CrossRef] [Scilit]
- Deng, J.; Czarnecki, K. MLOD: A multi-view 3D object detection based on robust feature fusion method. In Proceedings of the 2019 IEEE Intelligent Transportation Systems Conference (ITSC), Auckland, New Zealand, 27–30 October 2019; pp. 279–284. [Google Scholar]
- Zhou, W.; Gao, S.; Zhang, L.; Lou, X. Histogram of oriented gradients feature extraction from raw bayer pattern images. IEEE Trans. Circ. Syst. II Express Briefs 2020, 67, 946–950. [Google Scholar] [CrossRef] [Scilit]
- Hassaballah, M.; Alshazly, H.A.; Ali, A.A. Ear recognition using local binary patterns: A comparative experimental study. Expert Syst. Appl. 2019, 118, 182–200. [Google Scholar] [CrossRef] [Scilit]
- Tareen, S.A.K.; Saleem, Z. A comparative analysis of sift, surf, kaze, akaze, orb, and brisk. In Proceedings of the IEEE International Conference on Computing, Mathematics and Engineering Technologies, Sukkur, Pakistan, 3–4 March 2018; pp. 1–10. [Google Scholar]
- Li, G.; Feng, Y. Moving object detection based on SIFT feature matching and K-means clustering. J. Comput. Appl. 2012, 32, 2824–2826. [Google Scholar]
- Hu, M.K. Visual pattern recognition by moment invariants. IRE Trans. Inf. Theory 1962, 8, 179–187. [Google Scholar]
- Chen, X.; Wang, S.; Shi, C.; Wu, H.; Zhao, J.; Fu, J. Robust ship tracking via multi-view learning and sparse representation. J. Navig. 2019, 72, 176–192. [Google Scholar] [CrossRef] [Scilit]
- Mei, X.; Hong, Z.; Prokhorov, D. Robust Multitask Multiview Tracking in Videos. IEEE Trans. Neural Netw. Learn. Syst. 2015, 26, 2874–2890. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hong, Z.; Mei, X.; Prokhorov, D.; Tao, D. Tracking via robust multi-task multi-view joint sparse representation. In Proceedings of the IEEE International Conference on Computer Vision, Sydney, NSW, Australia, 1–8 December 2013; pp. 649–656. [Google Scholar]















| Sequence | Ship-1 | Ship-2 | Ship-3 | Ship-4 | Ship-5 | Ship-6 |
|---|---|---|---|---|---|---|
| STMV | 160.4 | 280.4 | 96.1 | 179.2 | 109.1 | 110.8 |
| MTT | 28.57 | 7.12 | 25.52 | 38.06 | 20.52 | 10.42 |
| MTSV | 30.9 | 20.33 | 32.82 | 24.9 | 21.46 | 32.57 |
| L1T | 28.67 | 6.69 | 21.99 | 50.76 | 12.76 | 11.71 |
| MTMV | 9.21 | 4.56 | 7.00 | 11.04 | 3.01 | 5.99 |
| Sequence | Ship-1 | Ship-2 | Ship-3 | Ship-4 | Ship-5 | Ship-6 |
|---|---|---|---|---|---|---|
| STMV | 9.95 | 15.64 | 6.11 | 10.7 | 6.98 | 5.82 |
| MTT | 1.98 | 0.41 | 1.46 | 2.46 | 1.18 | 0.60 |
| MTSV | 2.01 | 1.22 | 1.89 | 1.52 | 1.31 | 1.88 |
| L1T | 1.99 | 0.39 | 1.23 | 3.08 | 0.75 | 0.66 |
| MTMV | 0.51 | 0.24 | 0.37 | 0.58 | 0.16 | 0.32 |
Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. |
© 2022 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Chen, J.; Zhang, Z.; Wen, X. Target Identification via Multi-View Multi-Task Joint Sparse Representation. Appl. Sci. 2022, 12, 10955. https://doi.org/10.3390/app122110955
Chen J, Zhang Z, Wen X. Target Identification via Multi-View Multi-Task Joint Sparse Representation. Applied Sciences. 2022; 12(21):10955. https://doi.org/10.3390/app122110955
Chicago/Turabian StyleChen, Jiawei, Zhenshi Zhang, and Xupeng Wen. 2022. "Target Identification via Multi-View Multi-Task Joint Sparse Representation" Applied Sciences 12, no. 21: 10955. https://doi.org/10.3390/app122110955
APA StyleChen, J., Zhang, Z., & Wen, X. (2022). Target Identification via Multi-View Multi-Task Joint Sparse Representation. Applied Sciences, 12(21), 10955. https://doi.org/10.3390/app122110955

