A Domain-Guided Feature-Fusion Framework for Ship Equipment Based on Multi-Type Features
Abstract
1. Introduction
2. Related Work
2.1. Mixed-Data Clustering Methods
2.1.1. Feature-Transformation-Based Methods
2.1.2. Distance-Metric-Based Methods
2.1.3. Dedicated Algorithms for Mixed Numerical and Categorical Data
2.1.4. Spectral, Density-Based, and Deep Clustering Alternatives
2.2. Heterogeneous Feature Fusion and Text Representation
2.3. Cluster Validity Evaluation and Interpretation
3. Data Characteristics Analysis of Ship Equipment
4. Clustering Method Based on Multi-Type Feature Fusion
4.1. Multi-Type Feature Fusion Method
4.1.1. Numerical Feature Processing
4.1.2. Categorical Feature Processing
4.1.3. Text Feature Processing
4.1.4. Feature-Weighted Fusion
4.2. Clustering Algorithm and Optimal Cluster Number Determination
4.2.1. Clustering Algorithm Selection
4.2.2. Determination of the Optimal Number of Clusters
4.3. Evaluation and Analysis of Clustering Results
- Internal Validity Evaluation
- 2.
- Engineering Consistency Evaluation
- 3.
- Clustering Stability Evaluation
5. Case Study
5.1. Experimental Settings
5.2. Sensitivity Analysis of Feature Weights
5.3. Determination of the Number of Clusters
5.4. Analysis of Clustering Results
5.4.1. Cluster-Size Distribution
5.4.2. Representative Cluster Profiles
5.4.3. Engineering Consistency Analysis
5.4.4. Clustering Stability Analysis
5.5. Analysis of Feature-Weighting Mechanism
5.5.1. Comparison of Internal Validity
5.5.2. Comparison of Engineering Consistency
5.6. Comparative Experimental Analysis
6. Conclusions
- (1)
- To address the coexistence of multiple attribute types in ship equipment data, a unified feature representation method integrating numerical, categorical, and textual information is developed. By applying type-specific preprocessing strategies and feature-weighted fusion, the proposed representation effectively captures functional, technical, and configuration characteristics of the equipment and improves similarity identification performance.
- (2)
- To overcome the difficulty of determining the cluster number using a single evaluation criterion, a multi-index cluster number determination strategy is proposed. By jointly considering internal validity indices, including SC, CH, and DB, together with cluster-size characteristics, the proposed strategy reduces the bias caused by individual evaluation metrics and achieves a balance between clustering quality and candidate reference-set availability.
- (3)
- A complete ship equipment clustering framework is established, including data preprocessing, multi-type feature fusion, clustering implementation, and result evaluation. Experimental validation using real ship equipment data demonstrates that the obtained equipment clusters exhibit strong interpretability in terms of name similarity, specification/model similarity, and engineering attribute consistency. Therefore, the proposed framework can generate engineering-meaningful candidate sets of similar equipment.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Wei, M.; Chow, T.W.S.; Chan, R.H.M. Clustering heterogeneous data with k-means by mutual information-based unsupervised feature transformation. Entropy 2015, 17, 1535–1548. [Google Scholar] [CrossRef] [Scilit]
- Liu, P.; Yuan, H.; Ning, Y.; Chakraborty, B.; Liu, N.; Peres, M.A. A modified and weighted Gower distance-based clustering analysis for mixed type data: A simulation and empirical analyses. BMC Med. Res. Methodol. 2024, 24, 305. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mousavi, E.; Sehhati, M. A generalized multi-aspect distance metric for mixed-type data clustering. Pattern Recognit. 2023, 138, 109353. [Google Scholar] [CrossRef] [Scilit]
- Gower, J.C. A general coefficient of similarity and some of its properties. Biometrics 1971, 27, 857–871. [Google Scholar] [CrossRef] [Scilit]
- Huang, Z. Extensions to the k-means algorithm for clustering large data sets with categorical values. Data Min. Knowl. Discov. 1998, 2, 283–304. [Google Scholar] [CrossRef] [Scilit]
- Yu, W.L.; Yu, J.J.; Fang, J.W. K-prototypes clustering algorithm for mixed attribute data. Comput. Syst. Appl. 2015, 24, 168–172. [Google Scholar]
- Zhang, Y.; Zhao, M.; Chen, Y.; Lu, Y.; Cheung, Y. Learning unified distance metric for heterogeneous attribute data clustering. Expert Syst. Appl. 2025, 273, 126738. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Zou, R.; Zhang, Y.; Zhang, Y.; Cheung, Y.; Li, K. Adaptive micro partition and hierarchical merging for accurate mixed data clustering. Complex Intell. Syst. 2024, 11, 84. [Google Scholar] [CrossRef] [Scilit]
- Silva, S.D.F.; Reis, D.C.J.; Reis, S.D.M. Enhancing clustering stability, compactness, and separation in multimodal data environments. Data Knowl. Eng. 2026, 162, 102536. [Google Scholar] [CrossRef] [Scilit]
- Ghashti, S.J.; Thompson, J.R.J. Mixed-type distance shrinkage and selection for clustering via kernel metric learning. J. Classif. 2024, 42, 311–334. [Google Scholar] [CrossRef] [Scilit]
- Wei, W.; Ding, X.X.; Guo, M.X.; Yang, Z.; Liu, H. A review of text similarity calculation methods. Comput. Eng. 2024, 50, 18–32. [Google Scholar]
- Wang, J.; Zhang, Z.; Yue, S. A validity index for clustering evaluation by grid structures. Mathematics 2025, 13, 1017. [Google Scholar] [CrossRef] [Scilit]
- Davies, D.L.; Bouldin, D.W. A cluster separation measure. IEEE Trans. Pattern Anal. Mach. Intell. 1979, 2, 224–227. [Google Scholar] [CrossRef] [Scilit]
- Amorim, D.C.R.; Makarenkov, V. Improving clustering quality evaluation in noisy Gaussian mixtures. Neurocomputing 2026, 680, 133330. [Google Scholar] [CrossRef] [Scilit]
- Xie, X.L.; Beni, G. A validity measure for fuzzy clustering. IEEE Trans. Pattern Anal. Mach. Intell. 1991, 13, 841–847. [Google Scholar] [CrossRef] [Scilit]
- Rousseeuw, P.J. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 1987, 20, 53–65. [Google Scholar] [CrossRef] [Scilit]
- Fan, S.; Ding, S.; Xue, Y. Self-adaptive kernel K-means algorithm based on the shuffled frog leaping algorithm. Soft Comput. 2016, 20, 4463–4478. [Google Scholar] [CrossRef] [Scilit]
- Ikotun, M.A.; Habyarimana, F.; Ezugwu, E.A. Cluster validity indices for automatic clustering: A comprehensive review. Heliyon 2025, 11, e41953. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Blasilli, G.; Kerrigan, D.; Bertini, E.; Santucci, G. Towards a visual perception-based analysis of clustering quality metrics. In Proceedings of the 2024 IEEE Visualization in Data Science (VDS), St. Pete Beach, FL, USA, 14 October 2024; pp. 15–24. [Google Scholar]
- Di Nuzzo, C. Advancing spectral clustering for categorical and mixed-type data: Insights and applications. Mathematics 2024, 12, 508. [Google Scholar] [CrossRef] [Scilit]
- Mbuga, F.; Tortora, C. Spectral clustering of mixed-type data. Stats 2022, 5, 1. [Google Scholar] [CrossRef] [Scilit]
- Cendana, M.; Kuo, R.-J. Categorical Data Clustering: A Bibliometric Analysis and Taxonomy. Mach. Learn. Knowl. Extr. 2024, 6, 1009–1054. [Google Scholar] [CrossRef] [Scilit]
- Aschenbruck, R.; Szepannek, G.; Wilhelm, A.F.X. Initialization strategies for clustering mixed-type data with the k-prototypes algorithm. Adv. Data Anal. Classif. 2025, 2025, 1–30. [Google Scholar] [CrossRef] [Scilit]
- Szepannek, G. Clustering large mixed-type data with ordinal variables. Adv. Data Anal. Classif. 2025, 19, 749–767. [Google Scholar] [CrossRef] [Scilit]
- Salton, G.; Buckley, C. Term-weighting approaches in automatic text retrieval. Inf. Process. Manag. 1988, 24, 513–523. [Google Scholar] [CrossRef] [Scilit]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the NAACL-HLT 2019, Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar]
- Reimers, N.; Gurevych, I. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the EMNLP-IJCNLP 2019, Hong Kong, China, 3–7 November 2019; pp. 3982–3992. [Google Scholar]
- Calinski, T.; Harabasz, J. A dendrite method for cluster analysis. Commun. Stat. 1974, 3, 1–27. [Google Scholar] [CrossRef] [Scilit]
- Miller, C.; Portlock, T.; Nyaga, D.M.; O’sUllivan, J.M. A review of model evaluation metrics for machine learning in genetics and genomics. Front. Bioinform. 2024, 4, 1457619. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hikmat, S.H.; Karwan, J.; Mohmed, R.S. A semantics-based clustering approach for online laboratories using K-means and HAC algorithms. Mathematics 2023, 11, 548. [Google Scholar] [CrossRef] [Scilit]


| Attribute | Engineering Interpretation | Type |
|---|---|---|
| Equipment name | Represents equipment function, application purpose, and component category | Text |
| Specification/model information | Represents model series, technical parameters, and compatibility relationships | Text |
| Technical category | Indicates the professional domain in which the equipment belongs | Categorical |
| Measurement unit | Represents measurement method and certain physical characteristics | Categorical |
| Number of installations per platform | Reflects equipment configuration scale | Numeric |
| Reference unit price | Reflects equipment economic value | Numeric |
| Perturbed Feature | SC Range | Maximum Relative Change in SC (%) | CH Range | Maximum Relative Change in CH (%) | DB Range | Maximum Relative Change in DB (%) |
|---|---|---|---|---|---|---|
| Equipment Name | 0.1649~0.1900 | 8.33 | 32.3387~42.4237 | 17.86 | 2.1569~2.2390 | 3.80 |
| Specification/Model Information | 0.1695~0.1855 | 5.76 | 32.4254~40.3181 | 12.01 | 1.9142~2.2906 | 11.25 |
| Technical Category | 0.1727~0.1835 | 3.96 | 34.3409~36.8666 | 4.60 | 1.9460~2.2046 | 9.78 |
| Measurement Unit | 0.1742~0.1841 | 3.16 | 34.7192~35.9957 | 3.55 | 2.1569~2.2454 | 4.10 |
| Number of Installations per Platform | 0.1764~0.1900 | 5.63 | 33.4970~37.5760 | 6.94 | 2.1397~2.2726 | 5.36 |
| Reference Unit Price | 0.1707~0.1890 | 5.11 | 30.6608~40.5565 | 14.82 | 1.8942~2.2066 | 12.18 |
| K | SC | CH | DB | Minimum Cluster Size | Median Cluster Size | Maximum Cluster Size | Number of Clusters < 5 Samples | Number of Clusters < 10 Samples |
|---|---|---|---|---|---|---|---|---|
| 10 | 0.079221 | 69.419174 | 2.811651 | 34 | 126.5 | 292 | 0 | 0 |
| 20 | 0.113148 | 46.860392 | 2.562038 | 14 | 53.0 | 170 | 0 | 0 |
| 30 | 0.151644 | 38.048495 | 2.322321 | 9 | 30.5 | 166 | 0 | 2 |
| 34 | 0.174688 | 37.247257 | 2.185793 | 7 | 31.5 | 150 | 0 | 1 |
| 36 | 0.179861 | 35.995673 | 2.156935 | 9 | 31.5 | 117 | 0 | 1 |
| 38 | 0.184378 | 34.946641 | 2.243590 | 8 | 28.0 | 115 | 0 | 1 |
| 42 | 0.194056 | 33.386295 | 2.144390 | 6 | 24.5 | 114 | 0 | 1 |
| 46 | 0.220381 | 32.484584 | 1.989814 | 4 | 22.5 | 107 | 1 | 3 |
| 60 | 0.239235 | 29.079498 | 1.806580 | 3 | 14.5 | 145 | 1 | 15 |
| 80 | 0.286044 | 26.529038 | 1.784084 | 2 | 13.0 | 62 | 2 | 25 |
| 100 | 0.331505 | 25.269601 | 1.573549 | 2 | 10.0 | 78 | 7 | 48 |
| Cluster Size | Representative Equipment Names | Representative Specification/Model Terms | Technical-Category Purity | Measurement-Unit Purity |
|---|---|---|---|---|
| 13 | Digital-to-analog conversion polarity switch board (6); digital-to-analog conversion control logic board (3); digital-to-analog conversion logic board (3); polarity switch board (1) | DA-PW (6); DA-LC (5); DA-LC (1); DA-PW (1) | 1.000 | 1.000 |
| 28 | Digital conversion board 1 (15); digital-to-analog conversion board 1 (6); buffer board-1 (2); digital synchronizer conversion board-2 (1); digital-synchronizer conversion board-2 (1) | SM1 (21); RTU-BUF-1 (2); RTU-DS-2 (2); HEA901 (1); PM-ADC5/12-7 (1) | 0.643 | 0.893 |
| 40 | 5 V/15 V power module (23); 15 V power module (6); 5 V/12 V power module (5); ±15 V power module (2); 5 V/15 V DC power module (2) | ZDM3 (21); PS-15 (5); ZDM2 (5); ZDM29 (3); PS-P15 (2) | 1.000 | 1.000 |
| 39 | Serial interface board (30); computer serial interface board (7); 429 interface board (1); 1.9 serial interface board (1) | SC-SR-5 (13); CJK7 (6); CJK1 (5); CP-SI (4); CMP-SIF (4) | 1.000 | 1.000 |
| 48 | Computer motherboard (18); computer processor board (15); computer central processor board (3); computer processing board (2); central processor board-1 (2) | JZB12 (10); JZB11 (8); CMP-CPU (5); JZB5 (3); ITP-CPU (3) | 0.896 | 1.000 |
| 25 | Heading data processing board (23); heading processing board (2) | HXC1 (14); HXC2 (10); CP-DPCH-1 (1) | 1.000 | 1.000 |
| 25 | CAN interface board (16); timing control switching board (7); 1.13 dual CAN interface board (1); CAN bus interface board (1) | JK3 (8); SXK1 (5); JK1 (5); CMP-CAN (3); SXK3 (2) | 1.000 | 1.000 |
| 34 | Temperature control detection board (8); system timing board (8); system power control board (7); operation control and display panel (6); system detection board (4) | ST-TC (8); ST-TS (8); ST-SC (7); ST-OC (6); ST-DT (4) | 0.588 | 1.000 |
| Grouping Strategy | Equipment Name Cosine Similarity | Specification/Model Cosine Similarity | Technical-Category Purity | Measurement-Unit Purity |
|---|---|---|---|---|
| Random grouping | 0.023208 | 0.022718 | 0.581554 | 0.647100 |
| Proposed clustering | 0.193760 | 0.154359 | 0.870632 | 0.801487 |
| Evaluation Metric | Mean | Standard Deviation | Minimum | Maximum | Coefficient of Variation (%) |
|---|---|---|---|---|---|
| SC | 0.180083 | 0.005458 | 0.171172 | 0.189018 | 3.031 |
| CH | 35.520807 | 0.340675 | 34.908789 | 36.215801 | 0.959 |
| DB | 2.120594 | 0.068253 | 1.965578 | 2.242227 | 3.219 |
| ARI (relative to seed = 42 partitions) | 0.502783 | 0.044023 | 0.425387 | 0.635839 | — |
| NMI (relative to seed = 42 partitions) | 0.785802 | 0.018138 | 0.755712 | 0.828540 | — |
| Pairwise ARI | 0.521421 | 0.051077 | 0.385326 | 0.676630 | — |
| Pairwise NMI | 0.797079 | 0.019674 | 0.744659 | 0.861424 | — |
| Clustering Scheme | SC | CH | DB |
|---|---|---|---|
| Weighted scheme | 0.179861 | 35.995673 | 2.156935 |
| Equal-weight scheme | 0.125394 | 31.491564 | 2.661584 |
| Scheme | Name Cosine Similarity | Specification/Model Cosine Similarity | Technical-Category Purity | Measurement-Unit Purity | Minimum Cluster Size | Median Cluster Size | Name Cosine Similarity | Specification/Model Cosine Similarity | Technical-Category Purity |
|---|---|---|---|---|---|---|---|---|---|
| Weighted scheme | 0.193760 | 0.154359 | 0.870632 | 0.801487 | 9 | 31.5 | 117 | 0 | 1 |
| Equal-weight scheme | 0.148335 | 0.129822 | 0.859480 | 0.905576 | 9 | 29.5 | 104 | 0 | 1 |
| Method | Equipment-Name Cosine Similarity | Specification/Model Cosine Similarity | Technical-Category Purity | Measurement-Unit Purity |
|---|---|---|---|---|
| One-Hot + K-Means | 0.114064 | 0.108339 | 0.854275 | 0.907063 |
| AW-KMeans | 0.146825 | 0.126743 | 0.891204 | 0.932617 |
| K-Prototypes | 0.100896 | 0.088911 | 0.837175 | 0.882528 |
| Gower + hierarchical | 0.046349 | 0.039971 | 0.987361 | 0.995539 |
| Proposed method | 0.193760 | 0.154359 | 0.870632 | 0.801487 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yin, R.; Zhai, Y.; Gu, Z.; Shao, S. A Domain-Guided Feature-Fusion Framework for Ship Equipment Based on Multi-Type Features. Algorithms 2026, 19, 798. https://doi.org/10.3390/a19090798
Yin R, Zhai Y, Gu Z, Shao S. A Domain-Guided Feature-Fusion Framework for Ship Equipment Based on Multi-Type Features. Algorithms. 2026; 19(9):798. https://doi.org/10.3390/a19090798
Chicago/Turabian StyleYin, Ruoyi, Yali Zhai, Zhengxuan Gu, and Songshi Shao. 2026. "A Domain-Guided Feature-Fusion Framework for Ship Equipment Based on Multi-Type Features" Algorithms 19, no. 9: 798. https://doi.org/10.3390/a19090798
APA StyleYin, R., Zhai, Y., Gu, Z., & Shao, S. (2026). A Domain-Guided Feature-Fusion Framework for Ship Equipment Based on Multi-Type Features. Algorithms, 19(9), 798. https://doi.org/10.3390/a19090798

