A Controlled Proof-of-Concept Study for Quantitative Estimation of Syrup Addition in Honey Using RGB Histogram Descriptors and Explainable Machine Learning
Featured Application
Abstract
1. Introduction
2. Materials and Methods
2.1. Study Design and Sample Preparation
2.2. Digital Image Acquisition
2.3. Image Selection, Masking, and Region of Interest
2.4. RGB Histogram Construction and Statistical Descriptors
2.5. Whole-Image and Random Patch Representations
2.6. Descriptor Usefulness and Orientation-Level Technical Repeatability
2.7. Regression Models
2.8. Group-Aware Validation and Performance Metrics
2.9. Restricted-Feature Reference Model Design
2.10. Explainable Artificial Intelligence and Descriptor-Family Analysis
2.11. Software and Reproducibility
3. Results
3.1. Image-Set Verification and Technical Repeatability
3.2. Evolution of RGB Distributions with Syrup Addition
3.3. Usefulness and Stability of Statistical Descriptors
3.4. Nested Comparison of Regression Methods
3.5. Feature-Family Ablation
3.6. Restricted-Feature Reference Analyses
3.7. Performance-Weighted Explainability Analysis
3.8. Whole-Image Versus Patch-Based Modeling
3.9. Final All-Data Model Refits
3.10. Scope and Applicability of the Final Models
4. Discussion
4.1. Principal Findings and Methodological Contribution
4.2. Descriptor Representation and Interpretation
4.3. Whole-Image Representation and Technical Robustness
4.4. Comparison with Imaging–AI Literature
4.5. Limitations, Transferability, and Future Development
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Walker, M.J.; Cowen, S.; Gray, K.; Hancock, P.; Burns, D.T. Honey authenticity: The opacity of analytical reports—Part 1, defining the problem. npj Sci. Food 2022, 6, 11. [Google Scholar] [CrossRef] [Scilit]
- Walker, M.J.; Cowen, S.; Gray, K.; Hancock, P.; Burns, D.T. Honey authenticity: The opacity of analytical reports—Part 2, forensic evaluative reporting as a potential solution. npj Sci. Food 2022, 6, 12. [Google Scholar] [CrossRef] [Scilit]
- Danieli, P.P.; Lazzari, F. Honey traceability and authenticity: Review of current methods most used to face this problem. J. Apic. Sci. 2022, 66, 101–119. [Google Scholar] [CrossRef] [Scilit]
- Valverde, S.; Ares, A.M.; Elmore, J.S.; Bernal, J. Recent trends in the analysis of honey constituents. Food Chem. 2022, 387, 132920. [Google Scholar] [CrossRef] [Scilit]
- Vázquez, L.; Armada, D.; Celeiro, M.; Dagnac, T.; Llompart, M. Authenticity of honey: Characterization, bioactivities and sensorial properties. Foods 2022, 11, 1301. [Google Scholar] [CrossRef] [Scilit]
- Zábrodská, B.; Vorlová, L. Adulteration of honey and available methods for detection—A review. Acta Vet. Brno 2014, 83, S85–S102. [Google Scholar] [CrossRef] [Scilit]
- Soares, S.; Amaral, J.S.; Oliveira, M.B.P.P.; Mafra, I. A comprehensive review on the main honey authentication issues: Production and origin. Compr. Rev. Food Sci. Food Saf. 2017, 16, 1072–1100. [Google Scholar] [CrossRef] [Scilit]
- Biswas, A.P.; Tasnim, M.; Süfer, Ö.; Das, S.C.; Sarker, S.; Zhang, M.; Islam, N. Honey adulteration detection: A comprehensive review of traditional and modern techniques. J. Food Meas. Charact. 2026, 20, 3929–3964. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.-H.; Gu, H.-W.; Liu, R.-J.; Qing, X.-D.; Nie, J.-F. A comprehensive review of the current trends and recent advancements on the authenticity of honey. Food Chem. X 2023, 19, 100850. [Google Scholar] [CrossRef] [Scilit]
- Minho, L.A.C.; Conceição, J.L.; Barboza, O.M.; Santos Junior, A.F.; dos Santos, W.N.L. Robust DEEP heterogeneous ensemble and META-learning for honey authentication. Food Chem. 2025, 482, 144001. [Google Scholar] [CrossRef] [Scilit]
- Gerginova, D.; Kurteva, V.; Simova, S. Optical rotation—A reliable parameter for authentication of honey? Molecules 2022, 27, 8916. [Google Scholar] [CrossRef] [Scilit]
- Gbashi, S.; Njobeh, P.B. Enhancing food integrity through artificial intelligence and machine learning: A comprehensive review. Appl. Sci. 2024, 14, 3421. [Google Scholar] [CrossRef] [Scilit]
- Schoder, D. Honey fraud as a moving analytical target: Omics-informed authentication within a multi-layer analytical framework. Foods 2026, 15, 712. [Google Scholar] [CrossRef] [Scilit]
- Fragkos, N.; Bouzembrak, Y.; Erasmus, S.W. The role of artificial intelligence in combating food fraud: A systematic literature review. Crit. Rev. Food Sci. Nutr. 2026, advance online publication. 1–19. [Google Scholar] [CrossRef] [Scilit]
- Phillips, T.; Abdulla, W. A new honey adulteration detection approach using hyperspectral imaging and machine learning. Eur. Food Res. Technol. 2023, 249, 259–272. [Google Scholar] [CrossRef] [Scilit]
- Calle, J.L.P.; Punta-Sánchez, I.; González-de-Peredo, A.V.; Ruiz-Rodríguez, A.; Ferreiro-González, M.; Palma, M. Rapid and automated method for detecting and quantifying adulterations in high-quality honey using Vis-NIRs in combination with machine learning. Foods 2023, 12, 2491. [Google Scholar] [CrossRef] [Scilit]
- Lanjewar, M.G.; Panchbhai, K.G.; Patle, L.B. Sugar detection in adulterated honey using hyperspectral imaging with stacking generalization method. Food Chem. 2024, 450, 139322. [Google Scholar] [CrossRef] [Scilit]
- Al Noman, M.A.; Nijhum, A.B.; Hossain, I.; Islam, M.S.; Sifat, I.M.; Aziz, M.G.; Rahman, A. Non-destructive adulterants detection in various honey types in Bangladesh using UV–VIS–NIR spectroscopy coupled with machine learning algorithms. LWT 2025, 228, 118125. [Google Scholar] [CrossRef] [Scilit]
- Ahmed, E. Detection of honey adulteration using machine learning. PLoS Digit. Health 2024, 3, e0000536. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shao, Y.; Shi, Y.; Xuan, G.; Li, Q.; Wang, F.; Shi, C.; Hu, Z. Hyperspectral imaging for non-destructive detection of honey adulteration. Vib. Spectrosc. 2022, 118, 103340. [Google Scholar] [CrossRef] [Scilit]
- Razavi, R.; Esmaeilzadeh Kenari, R. Ultraviolet–visible spectroscopy combined with machine learning as a rapid detection method to predict adulteration of honey. Heliyon 2023, 9, e20973. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hao, S.; Yuan, J.; Wu, Q.; Liu, X.; Cui, J.; Xuan, H. Rapid identification of corn sugar syrup adulteration in wolfberry honey based on fluorescence spectroscopy coupled with chemometrics. Foods 2023, 12, 2309. [Google Scholar] [CrossRef] [Scilit]
- Geană, E.-I.; Isopescu, R.; Ciucure, C.-T.; Gîjiu, C.L.; Joșceanu, A.M. Honey adulteration detection via ultraviolet-visible spectral investigation coupled with chemometric analysis. Foods 2024, 13, 3630. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- David, M.; Berghian-Grosan, C.; Magdas, D.A. Honey differentiation using infrared and Raman spectroscopy analysis and the employment of machine-learning-based authentication models. Foods 2025, 14, 1032. [Google Scholar] [CrossRef] [Scilit]
- Kou, Z.; Chen, G.; Li, S.; Yang, Z.; Ouyang, L.; Gong, Y. Identification of honey adulterated with syrup by Raman spectroscopy and chemometrics. Food Sci. 2024, 45, 254–260. [Google Scholar] [CrossRef]
- Ennahli, S.; Ajal, E.A.; Bajoub, A.; Hssaini, L. Rapid prediction of honey adulteration using machine learning-assisted FTIR spectroscopy and chemometrics. Food Control 2026, 187, 112146. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.; Xu, B.; Luo, H.; Ma, R.; Du, Z.; Zhang, X.; Liu, H.; Zhang, Y. Adulteration quantification of cheap honey in high-quality Manuka honey by two-dimensional correlation spectroscopy combined with deep learning. Food Control 2023, 154, 110010. [Google Scholar] [CrossRef] [Scilit]
- Boateng, A.A.; Sumaila, S.; Lartey, M.; Oppong, M.B.; Opuni, K.F.M.; Adutwum, L.A. Evaluation of chemometric classification and regression models for the detection of syrup adulteration in honey. LWT 2022, 163, 113498. [Google Scholar] [CrossRef] [Scilit]
- Rachineni, K.; Kakita, V.M.R.; Awasthi, N.P.; Shirke, V.S.; Hosur, R.V.; Shukla, S.C. Identifying type of sugar adulterants in honey: Combined application of NMR spectroscopy and supervised machine learning classification. Curr. Res. Food Sci. 2022, 5, 272–277. [Google Scholar] [CrossRef] [Scilit]
- Martinello, M.; Stella, R.; Baggio, A.; Biancotto, G.; Mutinelli, F. LC-HRMS-based non-targeted metabolomics for the assessment of honey adulteration with sugar syrups: A preliminary study. Metabolites 2022, 12, 985. [Google Scholar] [CrossRef] [Scilit]
- Hansen, J.; Kunert, C.; Raezke, K.-P.; Seifert, S. Detection of sugar syrups in honey using untargeted liquid chromatography–mass spectrometry and chemometrics. Metabolites 2024, 14, 633. [Google Scholar] [CrossRef] [Scilit]
- Nyarko, K.; Mensah, S.; Greenlief, C.M. Examining the use of polyphenols and sugars for authenticating honey on the U.S. market: A comprehensive review. Molecules 2024, 29, 4940. [Google Scholar] [CrossRef] [Scilit]
- Akyıldız, İ.E.; Uzunöner, D.; Raday, S.; Acar, S.; Erdem, Ö.; Damarlı, E. Identification of the rice syrup adulterated honey by introducing a candidate marker compound for brown rice syrups. LWT 2022, 154, 112618. [Google Scholar] [CrossRef] [Scilit]
- Punta-Sánchez, I.; Dymerski, T.; Calle, J.L.P.; Ruiz-Rodríguez, A.; Ferreiro-González, M.; Palma, M. Detecting honey adulteration: Advanced approach using UF-GC coupled with machine learning. Sensors 2024, 24, 7481. [Google Scholar] [CrossRef] [Scilit]
- Egido, C.; Saurina, J.; Sentellas, S.; Núñez, O. Honey fraud detection based on sugar syrup adulterations by HPLC-UV fingerprinting and chemometrics. Food Chem. 2024, 436, 137758. [Google Scholar] [CrossRef] [Scilit]
- Pourmoradian, A.; Barzegar, M.; Gharaghani, S.; Sahari, M.A. Honey adulteration detection using the HS-SPME-IMS technique combined with chemometric analysis. Food Chem. X 2025, 32, 103365. [Google Scholar] [CrossRef] [Scilit]
- Bodor, Z.; Benedek, C.; Urbin, Á.; Szabó, D.; Sipos, L. Colour of honey: Can we trust the Pfund scale?—An alternative graphical tool covering the whole visible spectra. LWT 2021, 149, 111859. [Google Scholar] [CrossRef] [Scilit]
- Zangirolami, M.S.; Valderrama, P.; Santos, O.O. Bibliometric study and potential applications in smartphone-based digital images: A perspective from 2013 to 2024. Food Chem. 2025, 482, 144106. [Google Scholar] [CrossRef] [Scilit]
- Meenu, M.; Kurade, C.; Neelapu, B.C.; Kalra, S.; Ramaswamy, H.S.; Yu, Y. A concise review on food quality assessment using digital image processing. Trends Food Sci. Technol. 2021, 118, 106–124. [Google Scholar] [CrossRef] [Scilit]
- Wu, D.; Sun, D.-W. Colour measurements by computer vision for food quality control—A review. Trends Food Sci. Technol. 2013, 29, 5–20. [Google Scholar] [CrossRef] [Scilit]
- Fan, Y.; Li, J.; Guo, Y.; Xie, L.; Zhang, G. Digital image colorimetry on smartphone for chemical analysis: A review. Measurement 2021, 171, 108829. [Google Scholar] [CrossRef] [Scilit]
- Capitán-Vallvey, L.F.; López-Ruiz, N.; Martínez-Olmos, A.; Erenas, M.M.; Palma, A.J. Recent developments in computer vision-based analytical chemistry: A tutorial review. Anal. Chim. Acta 2015, 899, 23–56. [Google Scholar] [CrossRef] [Scilit]
- Roda, A.; Michelini, E.; Zangheri, M.; Di Fusco, M.; Calabria, D.; Simoni, P. Smartphone-based biosensors: A critical review and perspectives. TrAC Trends Anal. Chem. 2016, 79, 317–325. [Google Scholar] [CrossRef] [Scilit]
- Teye, E.; Amuah, C.L.Y.; Lamptey, F.P.; Obeng, F.; Nyorkeh, R. Artificial intelligence for honey integrity in Ghana: A feasibility study on the use of smartphone images coupled with multivariate algorithms. Smart Agric. Technol. 2024, 8, 100453. [Google Scholar] [CrossRef] [Scilit]
- Kwiek, P.; Jakubowska, M. Color standardization of chemical solution images using template-based histogram matching in deep learning regression. Algorithms 2024, 17, 335. [Google Scholar] [CrossRef] [Scilit]
- Brar, D.S.; Aggarwal, A.K.; Nanda, V.; Kaur, S.; Saxena, S.; Gautam, S. Detection of sugar syrup adulteration in unifloral honey using deep learning framework: An effective quality analysis technique. Food Humanit. 2024, 2, 100190. [Google Scholar] [CrossRef] [Scilit]
- Ilias, B.; Abdelaziz, B.; Anas, E.-N.; Douzi, S.; Douzi, H. Leveraging RegNet and CBAM for precise detection of honey adulteration using thermal imaging. Sci. Rep. 2025, 15, 36555. [Google Scholar] [CrossRef] [Scilit]
- Shen, C.; Wang, R.; Nawazish, H.; Wang, B.; Cai, K.; Xu, B. Machine vision combined with deep learning-based approaches for food authentication: An integrative review and new insights. Compr. Rev. Food Sci. Food Saf. 2024, 23, e70054. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shi, S.; Zhang, K.; Tian, N.; Jin, Z.; Liu, K.; Huang, L.; Tian, X.; Cao, C.; Zhang, Y.; Jiang, Y. Spectroscopic techniques combined with chemometrics for rapid detection of food adulteration: Applications, perspectives, and challenges. Food Res. Int. 2025, 211, 116459. [Google Scholar] [CrossRef] [Scilit]
- Wójcik, S.; Ciepiela, F.; Jakubowska, M. Computer vision analysis of sample colors versus quadruple-disk iridium–platinum voltammetric e-tongue for recognition of natural honey adulteration. Measurement 2023, 209, 112514. [Google Scholar] [CrossRef] [Scilit]
- Arrighi, L.; de Moraes, I.A.; Zullich, M.; Simonato, M.; Barbin, D.F.; Barbon Junior, S. Explainable artificial intelligence techniques for interpretation of food models: A review. Artif. Intell. Rev. 2026, 59, 176. [Google Scholar] [CrossRef] [Scilit]
- Kapoor, S.; Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 2023, 4, 100804. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Roberts, D.R.; Bahn, V.; Ciuti, S.; Boyce, M.S.; Elith, J.; Guillera-Arroita, G.; Hauenstein, S.; Lahoz-Monfort, J.J.; Schröder, B.; Thuiller, W.; et al. Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 2017, 40, 913–929. [Google Scholar] [CrossRef] [Scilit]
- Varma, S.; Simon, R. Bias in error estimation when using cross-validation for model selection. BMC Bioinform. 2006, 7, 91. [Google Scholar] [CrossRef] [Scilit]
- Varoquaux, G. Cross-validation failure: Small sample sizes lead to large error bars. NeuroImage 2018, 180, 68–77. [Google Scholar] [CrossRef] [Scilit]
- Rosenblatt, M.; Tejavibulya, L.; Jiang, R.; Noble, S.; Scheinost, D. Data leakage inflates prediction performance in connectome-based machine learning models. Nat. Commun. 2024, 15, 1829. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dormann, C.F.; Elith, J.; Bacher, S.; Buchmann, C.; Carl, G.; Carré, G.; Marquéz, J.R.G.; Gruber, B.; Lafourcade, B.; Leitão, P.J.; et al. Collinearity: A review of methods to deal with it and a simulation study evaluating their performance. Ecography 2013, 36, 27–46. [Google Scholar] [CrossRef] [Scilit]
- Shannon, C.E. A mathematical theory of communication. Bell Syst. Tech. J. 1948, 27, 379–423, 623–656. [Google Scholar] [CrossRef] [Scilit]
- Joanes, D.N.; Gill, C.A. Comparing measures of sample skewness and kurtosis. J. R. Stat. Soc. Ser. D 1998, 47, 183–189. [Google Scholar] [CrossRef] [Scilit]
- Koo, T.K.; Li, M.Y. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J. Chiropr. Med. 2016, 15, 155–163. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wold, S.; Sjöström, M.; Eriksson, L. PLS-regression: A basic tool of chemometrics. Chemom. Intell. Lab. Syst. 2001, 58, 109–130. [Google Scholar] [CrossRef] [Scilit]
- Jolliffe, I.T.; Cadima, J. Principal component analysis: A review and recent developments. Philos. Trans. R. Soc. A 2016, 374, 20150202. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
- Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Statistical modeling: The two cultures. Stat. Sci. 2001, 16, 199–231. [Google Scholar] [CrossRef] [Scilit]
- Boulesteix, A.-L.; Binder, H.; Abrahamowicz, M.; Sauerbrei, W. On the necessity and design of studies comparing statistical methods. Biom. J. 2018, 60, 216–218. [Google Scholar] [CrossRef] [Scilit]
- Sauerbrei, W.; Abrahamowicz, M.; Altman, D.G.; le Cessie, S.; Carpenter, J. STRengthening analytical thinking for observational studies: The STRATOS initiative. Stat. Med. 2014, 33, 5413–5432. [Google Scholar] [CrossRef] [Scilit]
- Hyndman, R.J.; Koehler, A.B. Another look at measures of forecast accuracy. Int. J. Forecast. 2006, 22, 679–688. [Google Scholar] [CrossRef] [Scilit]
- Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30; Curran Associates, Inc.: Red Hook, NY, USA, 2017; pp. 4765–4774. [Google Scholar]
- Lundberg, S.M.; Erion, G.; Chen, H.; DeGrave, A.; Prutkin, J.M.; Nair, B.; Katz, R.; Himmelfarb, J.; Bansal, N.; Lee, S.-I. From local explanations to global understanding with explainable AI for trees. Nat. Mach. Intell. 2020, 2, 56–67. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fisher, A.; Rudin, C.; Dominici, F. All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously. J. Mach. Learn. Res. 2019, 20, 1–81. [Google Scholar]
- Molnar, C. Interpretable Machine Learning: A Guide for Making Black Box Models Explainable, 2nd ed.; Leanpub: Victoria, BC, Canada, 2022. [Google Scholar]
- Aas, K.; Jullum, M.; Løland, A. Explaining individual predictions when features are dependent: More accurate approximations to Shapley values. Artif. Intell. 2021, 298, 103502. [Google Scholar] [CrossRef] [Scilit]
- Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
- Virtanen, P.; Gommers, R.; Oliphant, T.E.; Haberland, M.; Reddy, T.; Cournapeau, D.; Burovski, E.; Peterson, P.; Weckesser, W.; Bright, J.; et al. SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nat. Methods 2020, 17, 261–272. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Harris, C.R.; Millman, K.J.; van der Walt, S.J.; Gommers, R.; Virtanen, P.; Cournapeau, D.; Wieser, E.; Taylor, J.; Berg, S.; Smith, N.J.; et al. Array programming with NumPy. Nature 2020, 585, 357–362. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- McKinney, W. Data structures for statistical computing in Python. In Proceedings of the 9th Python in Science Conference, Austin, TX, USA, 28 June–3 July 2010; pp. 56–61. [Google Scholar]
- Hunter, J.D. Matplotlib: A 2D graphics environment. Comput. Sci. Eng. 2007, 9, 90–95. [Google Scholar] [CrossRef] [Scilit]









| Method Group | Main Application in Honey Authentication | Advantages | Main Limitations | References |
|---|---|---|---|---|
| Hyperspectral and Vis–NIR imaging | Spectral–spatial detection and quantification of honey adulteration | Rapid; non-destructive; broad spectral information | High equipment cost; high-dimensional data; calibration transfer required | [15,16,17,18,19,20] |
| UV–Vis and fluorescence spectroscopy | Chromophore and fluorophore fingerprints for screening and quantification | Rapid; simple operation; low sample consumption | Limited molecular specificity; matrix and botanical-origin effects | [21,22,23] |
| FTIR and Raman spectroscopy | Vibrational fingerprinting and chemometric discrimination | Minimal sample preparation; rapid acquisition; molecularly informative bands | Overlapping bands; sensitivity to baseline and preprocessing; chemometric modelling required | [24,25,26,27,28] |
| NMR spectroscopy | Compositional profiling and adulterant-type classification | High chemical information content; reproducible fingerprints | High equipment cost; specialized expertise required; lower throughput | [29] |
| Chromatography and mass spectrometry | Targeted analysis of markers, sugars, and untargeted metabolomic profiling | High selectivity and sensitivity; compound-level information | Sample preparation; consumables; longer analysis; high cost | [30,31,32,33,34,35] |
| Ion mobility and other fingerprinting strategies | Volatile or untargeted profiling for rapid model-based authentication | Fast screening; multidimensional fingerprints; high throughput | Matrix and storage effects; calibration stability; instrument dependence | [11,36] |
| Family | Included Descriptors | Analytical Purpose | n |
|---|---|---|---|
| Location and percentiles | Per channel: mean, median (P50), mode, P1, P5, P10, P25, P75, P90, P95, and P99 | Location and shift of the intensity distribution | 33 |
| Dispersion and width | Per channel: variance, standard deviation, IQR, P95−P5 width, P99−P1 width, and FWHM | Within-image variability and histogram spread | 18 |
| Distribution shape | Per channel: skewness and excess kurtosis | Asymmetry and tail/peak shape | 6 |
| Information and peak | Per channel: entropy, energy, and dominant-peak height | Distribution concentration and peak prominence | 9 |
| Inter-channel | Ratios and differences for mean, median, mode, P10, P25, P75, and P90; six unstable B-denominator ratios excluded | Relative color balance between R, G, and B | 36 |
| Total candidate pool | Complete set entering the nested feature-selection workflow | Candidate predictors before fold-specific selection | 102 |
| Adulterant (%) | Mean | Median | SD | Skew. | Kurt. | Entropy | Energy | Width | Peak Position | ΔMean vs. 0% |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 208.19 | 208.30 | 2.50 | −0.280 | 0.644 | 3.358 | 0.114 | 8.10 | 208.50 | 0.00 |
| 5 | 209.28 | 209.00 | 2.53 | −0.701 | 4.195 | 3.334 | 0.117 | 7.80 | 209.00 | 1.09 |
| 10 | 209.50 | 209.60 | 2.42 | −0.209 | 0.353 | 3.312 | 0.118 | 7.80 | 209.60 | 1.30 |
| 15 | 209.95 | 210.00 | 2.35 | −0.311 | 1.672 | 3.259 | 0.122 | 7.90 | 210.00 | 1.75 |
| 20 | 211.38 | 211.30 | 2.45 | −0.515 | 2.994 | 3.306 | 0.118 | 7.60 | 211.60 | 3.19 |
| 25 | 211.81 | 212.00 | 2.36 | −0.279 | 1.318 | 3.268 | 0.121 | 7.80 | 212.00 | 3.61 |
| 30 | 212.49 | 212.80 | 2.39 | −0.433 | 2.212 | 3.282 | 0.120 | 7.20 | 212.90 | 4.29 |
| 35 | 212.83 | 213.00 | 2.32 | −0.447 | 2.565 | 3.236 | 0.124 | 7.50 | 213.00 | 4.64 |
| 40 | 215.08 | 215.10 | 2.47 | −0.871 | 5.942 | 3.295 | 0.120 | 7.80 | 215.20 | 6.89 |
| 45 | 214.34 | 214.30 | 2.51 | −1.453 | 11.147 | 3.264 | 0.123 | 7.20 | 214.60 | 6.15 |
| 50 | 215.74 | 216.00 | 2.61 | −2.001 | 16.166 | 3.278 | 0.123 | 7.40 | 216.00 | 7.55 |
| Adulterant (%) | Mean | Median | SD | Skew. | Kurt. | Entropy | Energy | Width | Peak Position | ΔMean vs. 0% |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 183.52 | 183.70 | 2.96 | −0.341 | 0.213 | 3.596 | 0.096 | 9.60 | 184.50 | 0.00 |
| 5 | 185.83 | 186.00 | 2.88 | −0.670 | 2.280 | 3.530 | 0.103 | 9.20 | 186.10 | 2.31 |
| 10 | 188.50 | 188.70 | 2.69 | −0.358 | 0.490 | 3.454 | 0.108 | 8.90 | 188.90 | 4.98 |
| 15 | 190.68 | 190.90 | 2.59 | −0.519 | 1.614 | 3.389 | 0.113 | 8.60 | 191.00 | 7.16 |
| 20 | 192.78 | 193.00 | 2.68 | −0.473 | 2.099 | 3.441 | 0.108 | 8.60 | 193.00 | 9.25 |
| 25 | 195.62 | 195.90 | 2.54 | −0.350 | 1.111 | 3.378 | 0.112 | 8.50 | 196.00 | 12.10 |
| 30 | 197.03 | 197.00 | 2.64 | −0.435 | 1.587 | 3.425 | 0.109 | 8.20 | 197.00 | 13.50 |
| 35 | 199.52 | 199.90 | 2.48 | −0.538 | 2.456 | 3.327 | 0.117 | 7.50 | 200.00 | 16.00 |
| 40 | 200.29 | 200.50 | 2.56 | −0.830 | 4.970 | 3.347 | 0.116 | 7.70 | 200.70 | 16.77 |
| 45 | 200.90 | 201.00 | 2.71 | −1.372 | 9.322 | 3.379 | 0.114 | 8.20 | 201.00 | 17.38 |
| 50 | 203.76 | 204.00 | 2.84 | −1.724 | 12.335 | 3.418 | 0.112 | 8.70 | 204.00 | 20.24 |
| Adulterant (%) | Mean | Median | SD | Skew. | Kurt. | Entropy | Energy | Width | Peak Position | ΔMean vs. 0% |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 1.37 | 1.00 | 1.51 | 1.012 | 0.385 | 2.282 | 0.253 | 4.00 | 0.00 | 0.00 |
| 5 | 1.29 | 1.00 | 1.47 | 1.066 | 0.518 | 2.223 | 0.266 | 4.00 | 0.00 | −0.08 |
| 10 | 1.34 | 1.00 | 1.50 | 1.029 | 0.410 | 2.257 | 0.259 | 4.00 | 0.00 | −0.03 |
| 15 | 1.76 | 1.10 | 1.67 | 0.758 | −0.129 | 2.521 | 0.201 | 5.00 | 0.00 | 0.39 |
| 20 | 3.69 | 3.80 | 2.36 | 0.361 | −0.322 | 3.201 | 0.119 | 8.00 | 3.30 | 2.32 |
| 25 | 11.73 | 12.00 | 2.81 | −0.161 | −0.004 | 3.535 | 0.100 | 9.20 | 12.00 | 10.36 |
| 30 | 15.24 | 15.10 | 2.98 | −0.113 | −0.072 | 3.622 | 0.094 | 10.00 | 15.30 | 13.87 |
| 35 | 24.02 | 24.00 | 2.83 | −0.260 | 0.080 | 3.538 | 0.100 | 9.40 | 24.40 | 22.65 |
| 40 | 21.22 | 21.20 | 2.98 | −0.211 | 0.085 | 3.614 | 0.095 | 10.00 | 21.70 | 19.85 |
| 45 | 23.44 | 23.70 | 2.97 | −0.206 | −0.005 | 3.611 | 0.095 | 9.70 | 23.90 | 22.07 |
| 50 | 35.85 | 36.00 | 2.97 | −0.363 | 0.342 | 3.602 | 0.097 | 9.20 | 36.40 | 34.47 |
| Feature | Family | Spearman Correlation ρ | Pearson Correlation r | Median CV (%) | ICC(1,1) | Between/Within |
|---|---|---|---|---|---|---|
| Mean G | location percentiles | 1.000 | 0.992 | 0.085 | 0.9988 | 29.2 |
| P95 G | location percentiles | 1.000 | 0.993 | 0.000 | 0.9988 | 29.2 |
| Mean R | location percentiles | 0.991 | 0.984 | 0.047 | 0.9944 | 13.4 |
| Mean B | location percentiles | 0.945 | 0.944 | 0.769 | 0.9999 | 96.5 |
| P95 B | location percentiles | 0.963 | 0.955 | 0.000 | 0.9996 | 53.4 |
| Mean(R−G) | interchannel | −0.973 | −0.972 | 0.320 | 0.9998 | 67.0 |
| Mean R/Mean G | interchannel | −0.973 | −0.971 | 0.032 | 0.9997 | 57.9 |
| Entropy B | information peak | 0.845 | 0.899 | 0.742 | 0.9985 | 25.5 |
| Entropy G | information peak | −0.718 | −0.743 | 1.252 | 0.7384 | 1.7 |
| SD R | dispersion width | 0.118 | 0.171 | 3.348 | 0.3660 | 0.8 |
| Model | R2nested-LOCO | RMSEnested-LOCO (pp) | 95% Bootstrap CI | MAEnested-LOCO (pp) | RMSEnested-LOCO,5–45 (pp) | SDorientation (pp) | Max Orientation Range (pp) | RPDnested-LOCO |
|---|---|---|---|---|---|---|---|---|
| PLSR | 0.980 | 2.22 | 1.60–2.76 | 1.96 | 2.25 | 1.01 | 5.06 | 7.47 |
| SVR | 0.970 | 2.74 | 1.54–3.80 | 2.16 | 2.93 | 0.70 | 4.64 | 6.05 |
| GB | 0.920 | 4.48 | 3.39–5.48 | 4.11 | 4.13 | 0.88 | 8.35 | 3.70 |
| XGBoost | 0.918 | 4.53 | 3.36–5.56 | 4.01 | 4.25 | 0.94 | 13.79 | 3.66 |
| RF | 0.909 | 4.77 | 3.71–5.73 | 4.41 | 4.61 | 0.81 | 9.25 | 3.48 |
| Feature Family | Candidate Features | Model | R2nested-LOCO | RMSEnested-LOCO (pp) | MAEnested-LOCO (pp) |
|---|---|---|---|---|---|
| Location and percentiles | 33 | PLSR | 0.984 | 2.03 | 1.67 |
| SVR | 0.972 | 2.64 | 1.91 | ||
| Dispersion and width | 18 | PLSR | 0.902 | 4.95 | 3.86 |
| SVR | 0.852 | 6.08 | 4.62 | ||
| Distribution shape | 6 | PLSR | 0.913 | 4.66 | 3.76 |
| SVR | 0.901 | 4.98 | 4.15 | ||
| Information and peak | 9 | PLSR | 0.907 | 4.83 | 4.31 |
| SVR | 0.860 | 5.93 | 4.42 | ||
| Inter-channel | 36 | PLSR | 0.917 | 4.55 | 3.86 |
| SVR | 0.568 | 10.39 | 6.48 |
| Model | Rank Within Method | Feature Representation | Retained Descriptors | RMSEselection-CV (pp) |
|---|---|---|---|---|
| OLS | 1 | G_mean | 1 | 2.230 * |
| SVR | 1 | Full descriptor pool | 102 of 102 | 1.340 |
| 2 | G-channel descriptors | 10 of 22 | 1.455 | |
| 3 | RGB channel means | 3 | 1.556 | |
| PLSR | 1 | G-channel descriptors | 20 of 22 | 1.604 |
| 2 | Full descriptor pool | 40 of 102 | 1.607 | |
| 3 | RGB channel means | 3 | 1.666 | |
| Random Forest | 1 | Full descriptor pool | 40 of 102 | 3.562 |
| 2 | G-channel descriptors | 5 of 22 | 4.207 | |
| 3 | RGB channel means | 3 | 4.302 | |
| Gradient Boosting | 1 | Full descriptor pool | 40 of 102 | 3.630 |
| 2 | G-channel descriptors | 22 of 22 | 3.902 | |
| 3 | RGB channel means | 3 | 4.005 | |
| XGBoost | 1 | G-channel descriptors | 22 of 22 | 3.850 |
| 2 | RGB channel means | 3 | 4.263 | |
| 3 | Full descriptor pool | 40 of 102 | 4.289 |
| Model | Whole R2nested-LOCO | Whole RMSEnested-LOCO (pp) | Whole MAEnested-LOCO (pp) | Patch R2nested-LOCO | Patch RMSEnested-LOCO (pp) | Patch MAEnested-LOCO (pp) | Patch SDorientation (pp) |
|---|---|---|---|---|---|---|---|
| PLSR | 0.980 | 2.22 | 1.96 | 0.939 | 3.90 | 3.48 | 1.24 |
| SVR | 0.970 | 2.74 | 2.16 | 0.963 | 3.06 | 2.26 | 1.65 |
| GB | 0.920 | 4.48 | 4.11 | 0.903 | 4.93 | 3.96 | 1.48 |
| RF | 0.909 | 4.77 | 4.41 | 0.915 | 4.60 | 3.70 | 1.67 |
| XGBoost | 0.918 | 4.53 | 4.01 | 0.858 | 5.96 | 5.07 | 1.24 |
| Model | Retained Descriptors | Final Selected Configuration | RMSE selection-CV (pp) |
|---|---|---|---|
| PLSR | 40 | 6 latent components; scale = False after fold-local standardization | 1.607 |
| SVR | 102 | linear kernel; C = 0.1; epsilon = 0.1 | 1.340 |
| RF | 40 | n_estimators = 300; max_depth = 8; min_samples_leaf = 3; max_features = 0.8 | 3.562 |
| GB | 40 | n_estimators = 200; learning_rate = 0.05; max_depth = 3; min_samples_leaf = 2; subsample = 0.7 | 3.630 |
| XGBoost | 40 | n_estimators = 300; learning_rate = 0.03; max_depth = 4; min_child_weight = 5; subsample = 0.7; colsample_bytree = 0.5; reg_lambda = 10; tree_method = hist; max_bin = 64 | 4.289 |
| Model | Final Specification | Intercept b0 | Five Largest Standardized-Coordinate Terms |
|---|---|---|---|
| PLSR | p = 40; n_components = 6 | 25.0 | +4.260473 zG_p95 −3.984216 zR_p99 +3.406883 zR_p05 −3.176364 zG_p01 +3.152930 zR_p25 |
| Linear SVR | p = 102; kernel = linear; C = 0.1; ε = 0.1 | 24.7 | +0.697515 zR_p75 +0.674334 zR_p25 +0.579621 zR_p95 +0.547492 zR_mean −0.544224 zB_kurtosis |
| Algorithm | Key Information | Reported RMSE (pp) | Our RMSE selection-CV (pp) | Our RMSEnested-LOCO (pp) | Ref. |
|---|---|---|---|---|---|
| Deep neural network regression | Smartphone RGB; 7 illuminations; 0–50%; external test | 0.46 | — | — | [50] |
| PLSR | Vis–NIR; 2 honey types; lower-cost honey; test set | 2.784 | 1.604 | 2.22 | [16] |
| PLSR | HSI, 400–1000 nm; fructose/sucrose; calibration/validation | 5.26 | 1.604 | 2.22 | [20] |
| SVR | Vis–NIR; Gaussian kernel; test set | 1.894 | 1.340 | 2.74 | [16] |
| RF | Vis–NIR; 500 trees; test set | 8.475 | 3.562 | 4.77 | [16] |
| Stacking ensemble | HSI; 32 concentration levels; test set and 10-fold CV | 0.493 (test); 1.27 (10-fold CV) | — | — | [17] |
| OLS | No comparable quantitative honey-imaging RMSE found | — | 2.230 | 2.48 | — |
| GB | No comparable quantitative honey-imaging RMSE found | — | 3.630 | 4.48 | — |
| XGBoost | No comparable quantitative honey-imaging RMSE found | — | 3.850 | 4.53 | — |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Dziubaniuk, M.; Kwiek, P.; Jakubowska, M. A Controlled Proof-of-Concept Study for Quantitative Estimation of Syrup Addition in Honey Using RGB Histogram Descriptors and Explainable Machine Learning. Appl. Sci. 2026, 16, 8452. https://doi.org/10.3390/app16178452
Dziubaniuk M, Kwiek P, Jakubowska M. A Controlled Proof-of-Concept Study for Quantitative Estimation of Syrup Addition in Honey Using RGB Histogram Descriptors and Explainable Machine Learning. Applied Sciences. 2026; 16(17):8452. https://doi.org/10.3390/app16178452
Chicago/Turabian StyleDziubaniuk, Małgorzata, Patrycja Kwiek, and Małgorzata Jakubowska. 2026. "A Controlled Proof-of-Concept Study for Quantitative Estimation of Syrup Addition in Honey Using RGB Histogram Descriptors and Explainable Machine Learning" Applied Sciences 16, no. 17: 8452. https://doi.org/10.3390/app16178452
APA StyleDziubaniuk, M., Kwiek, P., & Jakubowska, M. (2026). A Controlled Proof-of-Concept Study for Quantitative Estimation of Syrup Addition in Honey Using RGB Histogram Descriptors and Explainable Machine Learning. Applied Sciences, 16(17), 8452. https://doi.org/10.3390/app16178452

