1. Introduction
Color classification is a fundamental challenge in numerous domains, including art [
1], product design [
2], agriculture [
3], environmental monitoring [
4], and biomedical diagnostics [
5]. Across these industries, precise and consistent color identification underpins quality control, regulatory compliance, and functional effectiveness. For instance, in the textile and design sectors, accurate reproduction of standard color palettes is vital for customer satisfaction [
6,
7], while in biomedicine [
8] and agrifood industries [
9,
10,
11], color signatures often serve as proxies for health or ripeness assessment. Traditional techniques for color classification, such as visual inspection and handheld colorimeters, suffer from subjectivity, limited spectral resolution, and poor repeatability across varying lighting conditions. Spectrometry-based approaches [
12,
13], though significantly more accurate, are constrained by high costs, bulkiness, and the need for controlled laboratory environments. Consequently, there is a growing demand for miniaturized, portable, and cost-effective solutions capable of delivering high-fidelity spectral measurements in real-time, in real contexts. In parallel, recent advances in compact multichannel optical filtering structures have demonstrated highly selective wavelength separation across the visible to near-infrared range, further highlighting the potential of miniaturized optical platforms for sensing applications [
14]. Emerging innovations in discrete LED-based spectroscopy offer an appealing alternative. Unlike traditional spectrometers, LED-based systems utilize fixed-wavelength light sources to achieve spectral measurements with reduced complexity, making them ideal for wearable and embedded systems. Among such systems, the SENSIPATCH [
15,
16] as shown in
Figure 1, developed by Sensichips [
17], represents a cutting-edge solution for wearable spectral sensing. This 18 cm × 18 cm patch integrates a scattering-based spectrometer and the SENSIPLUS [
18] single-chip versatile sensor interface, enabling real-time spectral analysis across visible and near-infrared ranges. As shown in
Figure 2, the SENSIPATCH is worn directly on the arm, offering an ergonomic and portable form factor suitable for continuous use.
The system spectrometer module operates using four discrete LEDs (Blue: 468 nm, Green: 523 nm, Yellow: 593 nm, Red: 645 nm) and two NIR photodiodes (850 nm and 950 nm), supported by a central photodiode for light intensity capture. To compensate for the limited number of LEDs while still covering a broad portion of the visible and near-infrared spectrum (
Figure 3), the acquired current signals were analyzed using machine learning models. This approach allows the classification framework to exploit the combined information from all spectral channels for discrimination among the different PANTONE classes considered in this study.
This configuration provides broad spectral coverage while maintaining low power consumption and form-factor efficiency. Additionally, the device incorporates a 16-electrode bioimpedance matrix supporting spectroscopy from 10 Hz to 1.5 MHz, enabling the monitoring of body composition, hydration levels, respiration, and cardiovascular activity. Environmental parameters are captured using Al2O3-based gas sensors for sweat and vapor detection, as well as integrated relative humidity and temperature sensors. Electrochemical capabilities are provided via an embedded potentiostat/galvanostat for voltammetry, while an audio frequency response sensor (7 Hz–1 kHz) supports mechanical and vibrational signal capture.
The core of the proposed system is represented by SENSIPLUS. It is a fully programmable, single-chip integrated sensor interface that operates at low voltage and is fabricated using the UMC 0.18 μm CMOS process. Detailed information about its features can be found in [
19], while here we recall the main ones. A simplified block diagram of the SENSIPLUS interface is shown in
Figure 4. The system includes a programmable analog front end (AFE) for both stimulating the sensor and reading its response. A Digital Sub Unit oversees system operation and facilitates communication with an external host via the Communication Line, using standard protocols such as SPI, I2C, or I3C, or alternatively through a proprietary single-wire protocol called SENSIBUS. The SENSIPLUS is capable of acquiring signals from either external sensors connected to its analog terminals (P0–P3) or from its integrated sensors (S0–S15). Additional External Power Pads (Pp0–Pp3) can be used to stimulate devices that require currents in the order of tens of mA. As in the presented works, LEDs are connected to these pads. The acquisition of LEDs responses is performed relying on a single photodiode, whose response is converted in the digital domain by means of the embedded 20-bit analog-to-digital converter. Building upon prior applications of the SENSIPLUS microsensor in ECG monitoring [
20], optical spectroscopy [
21], and water pollutant detection [
22,
23], this work extends its utility to the domain of color science.
In this study, we specifically utilize the SENSIPATCH’s spectrometer module to collect color spectral data from standardized PANTONE [
24] color card samples. The aim is to exploit the platform’s multispectral sensing capabilities for accurate color classification, with data subsequently analyzed using machine learning techniques to develop a robust, wearable-compatible color recognition approach. The PANTONE color cards were used as physical reference samples, placed directly on the sensor surface to ensure consistent and standardized data collection across a broad spectrum of colors. As a globally recognized and standardized color system used across textiles, printing, plastics, and digital media, PANTONE [
25,
26] offers a controlled and reproducible palette, making it ideal for benchmarking color classification systems. Its discrete, industry-standard definitions ensure that spectral variations can be accurately linked to specific hues, enabling a rigorous and repeatable evaluation of the sensor’s performance. Moreover, the wide diversity of shades in the PANTONE library allows us to test the classifier’s sensitivity to subtle color differences, including pastel tones, deep pigments, and near-neutral grays. The primary objectives of this study are:
To evaluate the accuracy and robustness of the SENSIPATCH spectrometer module for classifying 100 PANTONE color samples under controlled lighting conditions, with the aim of enabling future extensions to unconstrained or variable lighting scenarios.
To develop a robust end-to-end machine learning solution capable of operating on wearable-acquired spectral data.
To investigate the impact of physical contact variation (firm vs. loose contact) on classification accuracy, simulating real-world wearable deployment scenarios.
To explore the feasibility of using compact, LED-based spectroscopy for scalable deployment in applications beyond color classification, such as biomedical diagnostics, quality control, and environmental monitoring.
Building on earlier preliminary investigations [
16] on PANTONE color classification with the SENSIPATCH system, this work considers a larger standardized dataset, includes two contact conditions, adopts a more detailed preprocessing procedure, and provides a broader comparative evaluation of machine learning models.
The rest of the paper is structured as follows:
Section 2 reviews recent literature on LED-based spectral sensing and color classification.
Section 3 describes the experimental protocol and data processing methods. Results and model evaluations are presented in
Section 4, followed by conclusions and future directions in
Section 5.
3. Methodology
Our proposed methodology consists of four key phases: establishing a controlled experimental environment, collecting spectral data, preprocessing the acquired data, and applying machine learning algorithms for accurate color classification.
Figure 5 presents the methodology pipeline used for this study.
3.1. Experimental Setup
To ensure precise and reproducible spectral data collection, we designed an experimental setup that minimizes ambient light interference and maintains consistent sensor contact. This setup consists of a custom-built black-box enclosure lined with black paper to create a controlled lighting environment, eliminating external light variations. Controlled lighting conditions are essential in spectral analysis, as variations in ambient illumination can introduce inconsistencies in sensor readings [
54]. The SENSIPATCH wearable system was placed inside this enclosure to simulate real-world close-contact configurations, where the sensors interact directly with the color sample. This setup aimed to replicate practical deployment scenarios where wearable optical sensors operate in skin-contact environments, ensuring high signal fidelity and reducing external interference.
3.1.1. Experimental Setup Under Varying Contact Conditions
To evaluate the impact of sensor contact pressure on spectral measurements, data collection was conducted under two different conditions:
Firm Contact Condition (FCCW)—With Weight: A small circular weight with a hollow center (resembling a ring) with a mass of approximately 100 g, corresponding to a force of about 0.981 N and a nominal contact pressure of approximately 9.81 kPa over an effective contact area of 1.0 cm2, was placed on the PANTONE color card to ensure uniform and consistent pressure between the sensor and the card. The hollow center design ensures that the weight does not obstruct the light emitted by the sensors, allowing unobstructed spectral readings. This condition simulates close-contact scenarios found in wearable sensors, where firm contact enhances spectral signal stability by minimizing air gaps.
Loose Contact Condition (LCCW)—Without Weight: The PANTONE color card was placed on the sensor without additional pressure. This scenario simulates potential variations in real-world use, where wearable sensors might experience slight detachment from the surface due to movement or external disturbances.
These two conditions were introduced to assess the impact of contact pressure on spectral measurement accuracy, as variations in sensor alignment and surface proximity can influence spectral reflectance readings.
Figure 6 represents the experimental setup used for this study.
Figure 7 presents the schematic representation of the experimental setup used for this study.
3.1.2. Measurement Procedure
A structured measurement sequence was implemented to ensure consistency in data acquisition. The following steps were performed for each experimental condition:
Positioning the SENSIPATCH: The SENSIPATCH was placed inside the enclosure, and the lid was closed to eliminate ambient light interference.
Baseline Measurement Collection: The sensors initially operated without any external material for the first 21 measurements, capturing baseline spectral readings under controlled conditions. These baseline measurements were subsequently used as an offset to stabilize the remaining measurements during preprocessing step.
Placement of PANTONE Color Card: After the 21st measurement, a PANTONE color card was carefully placed on the sensor surface under the predefined contact condition (either with or without weight). The lid was closed again to maintain a controlled environment.
Spectral Data Acquisition: Measurements resumed from the 22nd to the 121st reading to capture the spectral response of the color card under controlled conditions.
Repetition Across Multiple Color Cards: The process was repeated for 100 different PANTONE color cards, ensuring statistical robustness in the collected dataset.
This controlled approach ensures that spectral variations can be attributed solely to the interaction between the sensor and the PANTONE color card, rather than external environmental factors. The black-box enclosure and close-contact design enhance the reproducibility of measurements, reducing errors associated with stray light or sensor misalignment.
3.2. Data Collection
To build a robust dataset for PANTONE color classification using the SENSIPATCH wearable system, spectral data were collected from 100 distinct PANTONE color cards and a plain white paper sample. In addition, a baseline class was recorded with the sensor exposed to ambient conditions without any object placed on it, resulting in a total of 102 classes. The complete list of PANTONE color codes included in the dataset is presented in
Table 1. The data acquisition process was carefully designed to ensure measurement consistency and reliability across all samples. Each color card was subjected to five independent trials, each consisting of 121 individual sensor readings, yielding a comprehensive dataset of spectral responses for each class. To ensure accurate measurements and eliminate distance-related variability, each PANTONE card was placed in direct contact with the sensor surface, maintaining a consistent zero-gap configuration.
For each measurement sequence, the sensors were initially exposed to ambient conditions, without any PANTONE color card, in order to establish a baseline response. Following this baseline acquisition, a color card was positioned on the sensors, and data collection continued for 121 readings. This protocol was designed to capture potential sensor drift over time, thereby enhancing the accuracy and reliability of the dataset used for machine learning-based classification. In total, measurements were recorded for 100 distinct PANTONE color cards, a white paper reference, and a baseline condition representing the absence of any material on the sensors. Under two experimental conditions, with five repeated trials per card and 121 readings per trial, the final dataset comprised 123,420 samples (61,170 per condition; 102 × 5 × 121 × 2). The data acquisition rate was optimized to ensure efficiency without compromising accuracy. Each photodiode response was recorded approximately every 265 milliseconds, with all six sensors being sequentially acquired, leading to an overall sampling rate of approximately 1.6 s per complete sensor cycle.
The SENSIPATCH wearable system’s spectrometer operates with a lock-in amplifier. Optically, the six LEDs emit light in the VIS-NIR spectrum with a 120° emission angle. The emitted light strikes the colored surface of the PANTONE sample, which reflects it according to the spectral characteristics of the corresponding PANTONE code. The reflected light is then captured by the central photodiode, which converts it into a current signal acquired using the SLM-Studio software (version 1.2.5). The currents generated by the photodiode in response to the light reflected from the PANTONE surface under illumination by the six LEDs are the input for the machine learning models. In fact, as can be observed in
Figure 8 for the PANTONE 108 (Yellow) and PANTONE 1797 (Red), the currents associated with the two NIR signals (OUTPORT_AUX and OUTPORT_SHA) vary, and so they have a diagnostic value for the purposes of PANTONE color classification by the machine learning models used in the paper. So in this case, the AUX and SHA values have small variations between the two PANTONEs on the order of hundreds of micro-Amperes.
In particular, the active LED is excited by modulating its bias current with a sinusoidal waveform. The IN-PHASE component of the photodiode current is precisely detected, ensuring that artifacts from light sources other than the active LED are rejected. Each measurement corresponds to a specific LED configuration, with different wavelengths being stimulated to enhance the spectral differentiation of the PANTONE colors. The detailed LED configurations and corresponding wavelengths used in data acquisition are outlined in
Table 2. This carefully controlled data collection process ensures that the dataset is well-suited for the application of machine learning algorithms in PANTONE color classification, laying the foundation for accurate and scalable color identification using wearable sensor technology.
Figure 9 and
Figure 10 provide visual representations of the spectral data collected under FCCW and LCCW, respectively. Under the firm contact condition, signal traces for each of the six channels are smooth, stable, and well-separated, reflecting the benefits of consistent sensor-surface proximity. In contrast, the loose contact condition introduces higher variability, likely caused by inconsistent pressure and air gaps; while the fundamental spectral characteristics are still present, signal fluctuations and overlap between channels are more prominent under LCCW. These figures clearly demonstrate the influence of mechanical contact stability on spectral signal quality and underscore the importance of a robust experimental setup. A summary of the full-scale data acquisition parameters is provided in
Table 3.
3.3. Data Preprocessing
To ensure the quality, consistency, and suitability of the raw spectral data for machine learning classification, a structured preprocessing pipeline was implemented. The SENSIPATCH spectrometer module provides six raw input features per sample, corresponding to the in-phase current responses of the photodiode to sequential LED activations—four in the visible range and two in the near-infrared (NIR), as detailed earlier. These raw signals, while informative, are susceptible to baseline drift, environmental fluctuations, and amplitude variations across spectral channels. Therefore, preprocessing was essential to correct for sensor offsets, standardize the data scale, and enhance robustness for subsequent model training. The specific preprocessing steps applied baseline correction, bootstrapping for dataset expansion, and Z-score normalization, are described in the following subsections.
3.3.1. Baseline Correction
Before processing the measurements, baseline correction was applied to remove the sensor’s inherent response when no external material was present. The first 21 measurements, recorded with no PANTONE color card placed on the sensor, were used to establish the natural baseline response of the sensor. These readings captured any systematic variations due to environmental conditions and sensor behavior in the absence of external stimuli. As shown in
Figure 9 and
Figure 10, there is a sudden spike in the measurements when we put the PANTONE color card on the sensor. To correct for this baseline effect, the mean of the first 21 readings was computed, and this value was subtracted from each of the subsequent 100 spectral measurements collected with the PANTONE color card in place.
Mathematically, baseline correction was applied as:
where:
is the original sensor measurement for the ith data point,
is the mean of the first 21 baseline measurements (when no card was present).
is the baseline-corrected spectral response for the ith data point.
This step ensured that all color-specific spectral data were measured relative to a common baseline, improving reliability and consistency across all measurements.
The effects of this baseline correction procedure are graphically illustrated in
Figure 11 and
Figure 12, which shows baseline-adjusted signal curves for several representative color samples. The transformation from raw to corrected signals demonstrates enhanced alignment and reduced inter-channel variability. Notably, baseline-corrected data reveals clearer separation among channels and improved spectral stability, which is crucial for downstream machine learning tasks.
3.3.2. Bootstrapping for Dataset Expansion
To enhance dataset robustness and mitigate overfitting, bootstrapping was employed as a data resampling technique. Each file, containing 100 data points after baseline correction, was resampled to generate 2000 data points. Bootstrapping was performed by randomly drawing spectral readings with replacement from the corrected dataset, ensuring a diverse and expanded dataset for analysis.
This resampling process maintained the same statistical distribution as the original data while ensuring sufficient variability for training the machine learning models. By generating an expanded dataset, the system became less sensitive to minor variations in spectral readings caused by sensor noise, ambient fluctuations, or experimental conditions. The increased dataset size also improved the model’s generalization ability and robustness against overfitting. Bootstrapping was applied only to the training set within each cross-validation fold, after the measurement-level train/test split, while the test set was kept unchanged. Its purpose was not to generate new physical information, but to provide a more robust resampled training distribution from the limited number of post-baseline-correction samples available for each measurement.
3.3.3. Normalization Using Z-Score Standardization
After bootstrapping, Z-score normalization was applied to standardize spectral readings across different wavelengths and sensors. This transformation ensured that all spectral features had a mean of zero and unit variance, allowing for consistent feature scaling before applying machine learning models.
The Z-score normalization was computed using:
where:
is the spectral value after bootstrapping and baseline correction;
is the mean of the resampled dataset;
is the standard deviation of the resampled dataset.
For each cross-validation fold, the parameters required for Z-score normalization, namely the mean and standard deviation, were computed exclusively from the training data and then applied unchanged to the corresponding test data, thereby preventing data leakage. Since the six LED photodiode sensors operated at different wavelengths, this normalization helped mitigate variations in intensity levels across the different spectral bands, enabling the model to focus on spectral shape variations rather than absolute intensity differences.
3.4. Feature Importance Analysis of Multispectral Channels
In addition to normalization, we analyzed the contribution of each spectral channel using the built-in feature importance scores of the Random Forest classifier. Feature importance was computed for each cross-validation fold and averaged to obtain a stable ranking of the six wavelengths.
As shown in
Table 4, the visible channels (468–645 nm) contribute most strongly to classification performance, with the red (645 nm) and yellow (593 nm) wavelengths exhibiting the highest importance. This is consistent with the physical basis of color discrimination, as the PANTONE classes are defined primarily in the visible spectrum. The near-infrared channels (850 nm and 950 nm) show lower but non-negligible importance, indicating that they provide complementary information related to material- and pigment-dependent reflectance properties.
3.5. Color Classification with Machine Learning
After data preprocessing, the next step in our methodology is the application of machine learning algorithms for the classification of PANTONE colors based on the spectral data acquired using the SENSIPATCH wearable system. The goal is to develop a robust model that accurately distinguishes between different colors while ensuring generalization across varying experimental conditions. As described in the previous subsection, each input sample to the machine learning models consists of a six-dimensional feature vector. The input features are used directly for training after the preprocessing. The classification task is defined as a supervised multi-class problem with 102 output classes. These include 100 standardized PANTONE color cards, one plain white paper sample, and one baseline class representing measurements without any object on the sensor.
Selection of Machine Learning Models
To ensure optimal classification of spectral responses corresponding to different PANTONE colors, we employed a range of machine learning models, each chosen based on its suitability for handling high-dimensional spectral data and its classification efficiency:
Random Forest (RF): Selected for its robustness in handling high-dimensional sensor data and its interpretability via feature importance analysis. Its ensemble structure mitigates overfitting and enhances generalization capabilities [
55,
56].
Support Vector Machine (SVM): Chosen due to its ability to construct complex decision boundaries, which is particularly advantageous for sensor-derived datasets where non-linear relationships may exist among features [
57,
58].
Multilayer Perceptron (MLP): Selected as a classical neural network architecture capable of capturing non-linear relationships in numerical sensor data. Its fully connected structure enables the extraction of discriminative features from multivariate time series input, making it suitable for classification tasks involving wearable sensor measurements [
59].
Convolutional Neural Network (CNN): Employed due to its effectiveness in pattern recognition, especially when localized temporal or spatial features are relevant; while originally developed for image analysis, CNNs have shown strong performance on one-dimensional time-series data, such as those generated by sequential sensor readings, by leveraging convolutional filters to extract local structures [
60].
Long Short-Term Memory (LSTM): Utilized for its strength in modeling sequential dependencies, which is particularly beneficial when analyzing time-series sensor data where the order of readings conveys important contextual information. LSTM networks, a type of recurrent neural network (RNN), are designed to capture long-range dependencies and overcome vanishing gradient issues, making them suitable for tasks requiring temporal context modeling [
61].
Among the evaluated models, the MLP demonstrated the most consistent generalization performance across both experimental conditions. Its flexibility in tuning and suitability for deployment in embedded systems further support its selection as the primary model in this study, while the Support Vector Machine (SVM) achieved the highest accuracy under firm contact conditions, and Random Forest performed competitively on the combined dataset, these models exhibited more limited adaptability across conditions.
Table 5 summarizes the architectural configurations and training hyperparameters used for all models considered. These parameters were selected through empirical experimentation, where initial values were chosen based on standard practice and prior work in similar domains. To improve transparency, hyperparameter tuning was conducted iteratively using validation-based performance feedback, with the goal of maximizing classification accuracy while maintaining generalization across contact conditions. Specifically, practical ranges of learning rate, number of neurons, regularization strength, and training epochs were explored. Final configurations were selected based on convergence stability and consistent validation performance across folds; while no exhaustive hyperparameter search was conducted, the selected settings provided a reliable balance between model accuracy, computational efficiency, and robustness across varying contact scenarios.
4. Machine Learning Experiments and Results
To assess the performance of the proposed machine learning models, we conducted classification experiments under three experimental scenarios: FCCW, LCCW, and a combined condition (FCCW + LCCW). These scenarios reflect both ideal and real-world sensor placement configurations and allow evaluation of model robustness against variations in contact pressure. To ensure robust and unbiased performance evaluation, we adopted 5-fold cross-validation for both the FCCW and LCCW scenarios, as shown in
Figure 13, and 10-fold cross-validation for the Combined condition. These fold counts were directly aligned with the number of available measurements per class: five measurements in FCCW and LCCW individually, and ten in the Combined condition (i.e., five from each contact type). In each fold, one measurement per class was held out for testing while the remaining were used for training, ensuring measurement-level independence and preventing data leakage. This fold-wise approach helped mitigate overfitting and allowed each model to be evaluated on all measurement sequences in a rotation-based manner.
For each of the 102 classes (comprising 100 PANTONE color cards, one white paper reference, and one baseline), five independent measurements were recorded under each contact condition. Each measurement initially consisted of 121 sensor readings, of which 100 were retained after baseline correction. To increase dataset size and improve model robustness, bootstrapping was applied to each corrected measurement, resulting in 2000 samples per measurement. This produced a total of 1,020,000 samples per condition (102 classes × 5 measurements × 2000 samples). In both FCCW and LCCW experiments, four out of the five measurements per class were used for training, and the remaining one was used for testing. Therefore, each condition included 8000 training samples per class (4 measurements × 2000 samples) and 2000 test samples per class (1 measurement × 2000 samples), leading to a total of 816,000 training samples and 204,000 test samples per condition.
In the Combined condition, all ten measurements per class (five from FCCW and five from LCCW) were utilized as shown in
Figure 14. The training set comprised eight measurements per class, four from each contact condition, while the test set included the remaining two measurements (one from each condition). This resulted in 1,632,000 training samples (102 × 8 × 2000) and 408,000 test samples (102 × 2 × 2000). By incorporating both contact configurations in the training and evaluation phases, this setup enabled a more comprehensive and realistic assessment of model generalization under variable physical deployment scenarios.
Table 6 summarizes the measurement structure, number of samples, and the corresponding cross-validation folds used for each experimental condition.
4.1. Classification Performance Under FCCW
This subsection evaluates model performance under the Firm Contact Condition With Weight (FCCW), representing a stable and ideal sensor–card interface. In this setup, a circular weight with a hollow center was used to apply uniform pressure on the PANTONE color card, minimizing air gaps and enhancing signal stability.
Among all models, SVM achieved the highest average accuracy, 97.29%, with Random Forest 97.17% and MLP 97.04% following closely. CNN also performed competitively, whereas LSTM showed significantly lower accuracy, indicating limited effectiveness in this static setup.
Table 7 summarizes the five-fold cross-validation results.
To further analyze model behavior,
Table 8 presents the mean precision, recall, and F1 score for each classifier.
4.2. Classification Performance Under LCCW
This subsection evaluates model performance under the Loose Contact Condition Without Weight (LCCW), which simulates real-world wearable usage scenarios involving potential card misalignment and movement. In this setup, no weight was applied to the PANTONE color card, leading to reduced contact stability and increased signal variability.
Despite these challenges, most models maintained high classification accuracy, as shown in
Table 9. SVM achieved the highest mean accuracy of 97.14%, outperforming all other models under LCCW. MLP followed with a mean accuracy of 95.86%, while CNN and RF recorded 93.26% and 91.84%, respectively. LSTM again showed the weakest performance, with a mean accuracy of 76.33%, indicating limited robustness to contact variation.
To further assess model performance,
Table 10 reports the average precision, recall, and F1 scores. SVM continued to lead with an F1 score of 96.83%, reflecting strong generalization in the presence of contact inconsistency. MLP also performed well with an F1 score of 95.05%, reinforcing its robustness. CNN and RF followed, while LSTM remained the least effective model across all metrics.
These results suggest that while reduced contact stability does introduce classification variability, well-regularized models like SVM and MLP can maintain high performance. This reinforces the need for models that can tolerate minor sensor–card misalignment in real-world deployments.
4.3. Performance Under Combined Conditions (FCCW + LCCW)
To evaluate model robustness under realistic and variable sensor contact conditions, datasets from both FCCW and LCCW scenarios were merged to form a comprehensive dataset. This setup simulates practical deployment scenarios where contact pressure and alignment may vary. Classification performance was evaluated using 10-fold cross-validation, with results summarized in
Table 11.
Among the models, the MLP achieved the highest mean accuracy 96.08%, followed by Random Forest 94.96%. In contrast, SVM performance dropped notably to 88.40%, highlighting its sensitivity to contact variation. CNN and LSTM recorded lower average accuracies of 84.18% and 70.21%, respectively, with LSTM once again demonstrating the weakest performance across folds.
Table 12 presents the average precision, recall, and F1 score for each model under the combined condition. MLP maintained its leading position with the highest F1 score 95.37%, indicating strong generalization to varying contact dynamics. RF also performed well with an F1 score of 93.85%. Meanwhile, SVM’s decline in recall suggests reduced robustness under mixed-contact conditions. CNN showed moderate performance, and LSTM remained consistently the least effective.
These results confirm that training on combined datasets enables better generalization across heterogeneous contact conditions. MLP’s superior performance highlights its adaptability and makes it a strong candidate for deployment in real-world, variable-pressure environments. The drop in SVM’s performance further supports the need for models that can accommodate signal inconsistencies introduced by imperfect sensor placement.
4.4. Discussion and Comparative Insights
The experimental results reveal that machine learning models can effectively classify spectral responses from the SENSIPATCH wearable system across varying contact scenarios.
Figure 15 presents the performance of the machine learning models under FCCW, LCCW, and combined contact conditions. Among all models, the MLP consistently demonstrated strong and stable performance across all evaluation settings, particularly excelling in the combined setup with a mean accuracy of 96.08%. SVM also achieved competitive results, slightly outperforming MLP under the LCCW, while RF delivered strong baseline performance, particularly in stable environments. CNN exhibited reasonable generalization and handled noisy data better than LSTM, although its overall performance was marginally lower than that of MLP and SVM. LSTM networks, while suited to sequential data modeling, proved less effective in this context. This may be due to the relatively short sequence length and the absence of strong temporal dependencies in the spectral signal, which limits the advantage of recurrent architectures. Contact pressure was found to play a significant role in classification stability. The FCCW produced more consistent and well-separated spectral signatures, reflected in higher model performance across the board. However, the performance drop under LCCW was relatively modest, indicating that the system maintains resilience to moderate physical misalignment. This reinforces the importance of designing sensor systems and corresponding algorithms that tolerate real-world variability in deployment.
Furthermore, fold-to-fold variability observed in the LCCW and combined datasets underscores the value of training on diverse conditions. By incorporating both contact scenarios into a single dataset, models were better able to generalize across different sensor placements, which is critical for wearable applications. This combined training approach simulates realistic deployment conditions and prepares the model for variations in skin contact pressure and user motion. Among the evaluated models, the MLP was selected as the primary classifier for deployment due to its strong balance between classification accuracy, generalization capability, and architectural flexibility, while SVM demonstrated slightly higher accuracy under stable contact conditions, and RF performed competitively on the combined dataset, MLP consistently exhibited robust performance across both firm and loose contact scenarios. Moreover, MLPs are well-suited for iterative fine-tuning, scalable learning with larger datasets, and efficient integration into embedded systems through techniques such as pruning and quantization. This behavior can be attributed to the nature of the distortions introduced by contact variability in the sensing process. Changes in contact pressure do not result in uniform scaling of the signal but instead produce uneven variations across the six spectral channels, altering their relative relationships. In this context, the MLP is able to exploit these inter-channel patterns by learning combinations of spectral responses rather than relying on individual feature magnitudes. In contrast, models such as SVM operate on a fixed representation of the input space and are more sensitive to such non-uniform shifts, particularly when the relative structure of the features changes between training and testing conditions. This makes the MLP more suitable for handling the type of variability introduced by loose-contact measurements in the proposed system. These practical advantages, combined with their adaptability to real-world sensor variability, support the choice of MLP as the preferred model in this study.
To gain a deeper understanding of the classification behavior at the class level, we analyzed the model’s performance using both subset confusion matrices and a detailed class-wise accuracy table. Given the large number of classes (102), presenting the full confusion matrix was impractical. Instead, two targeted confusion matrices were generated: one highlighting classes with the highest accuracies (
Figure 16) and another focusing on classes with the lowest accuracies (
Figure 17).
The confusion matrix for the lower-performing classes revealed that certain colors, such as
PANTONE166,
PANTONE7720, and
PANTONE17-1456, exhibited greater misclassification rates. These misclassifications were predominantly observed between spectrally similar tones, notably between
PANTONE166 and
PANTONE152, or between
PANTONE7720 and
PANTONE635. Such overlaps are not unexpected, given the subtle differences in their reflectance spectra, particularly within muted red and dark green regions. Interestingly, some confusion was also noted between dark-toned samples and the
NO_PANTONE baseline class, suggesting that very low reflectance levels can sometimes blur the model’s decision boundaries. Nevertheless, even among these more challenging classes, the MLP model maintained respectable classification performance, with accuracies consistently above 75%. To further clarify this behavior,
Figure 18 shows representative spectral comparisons for selected challenging PANTONE pairs. In particular, the spectra of
PANTONE152 and
PANTONE166 exhibit very similar trends across the measured wavelength range, with only limited separation. Because the current SENSIPATCH spectrometer module relies on a small number of discrete optical channels, such subtle spectral differences are not always captured with sufficient resolution at the sensor level. Consequently, these classes may generate overlapping multispectral responses, which reduces the discriminative information available to the classifier and increases the likelihood of misclassification. The same figure also shows the spectral comparison between
PANTONE7720 and
PANTONE635, indicating that the observed difficulty is mainly associated with high spectral similarity between specific color pairs, while low overall reflectance may further reduce separability in some darker tones.
In contrast, the subset confusion matrix of high-performing classes showcased exemplary classification results. Twenty representative colors, including vivid tones such as
PANTONE1797,
PANTONE345, and
PANTONE535 were classified with 100% accuracy, without any observable misclassifications. These results demonstrate the model’s strong ability to distinguish between spectrally distinct colors, a feature that holds substantial promise for real-world wearable applications. It is worth noting that although only 20 classes are shown in the figure for clarity, a much larger number of classes actually achieved perfect classification. To complement the confusion matrices,
Table 13 provides a comprehensive view of the class-wise classification accuracies for all PANTONE color classes. The table reveals that an overwhelming majority of classes achieved exceptionally high accuracy, with 65 classes classified at 99% or above. Highly saturated and spectrally well-separated colors, such as
PANTONE100,
PANTONE1797,
PANTONE3278, and
WHITE_PAPER, were reliably distinguished by the MLP model. Conversely, a smaller group of colors particularly
PANTONE7720,
PANTONE166,
PANTONE152, and
PANTONE635 achieved lower, though still reasonable, accuracies between 77% and 85%. These findings align well with observations from the confusion matrix analysis, where muted or spectrally overlapping colors posed greater challenges.
Taken together, the insights from the subset confusion matrices and the class-wise accuracy table illustrate a consistent trend: when colors exhibit distinct and strong spectral features, classification is highly reliable. In cases where spectral differences are more nuanced, performance declines modestly but remains robust. Overall, the SENSIPATCH system, combined with a well-designed machine learning pipeline, demonstrates excellent potential for scalable, reliable, and real-world colorimetric sensing applications.
5. Conclusions and Future Works
This study presented a machine-learning-driven framework for accurate color classification using the SENSIPATCH, a compact, wearable spectral sensing system based on multi-wavelength LED–photodiode configurations. Through a structured experimental protocol involving firm and loose contact scenarios, we demonstrated the SENSIPATCH’s capability to capture stable and distinctive spectral signatures across 100 PANTONE color classes. Among the five machine learning models evaluated, the multilayer perceptron (MLP) consistently delivered the strongest generalization performance across both contact conditions, achieving a mean accuracy of 96.08% under combined scenarios. These results underscore the potential of integrating compact spectroscopic sensing with machine learning for practical, real-time color recognition in wearable applications. Graphical analyses of baseline-corrected and raw spectral data confirmed the importance of contact pressure and validated the effectiveness of the preprocessing pipeline. However, several limitations remain. The current system was evaluated in a controlled, enclosed environment designed to minimize ambient interference and ensure repeatability; while effective for benchmarking, this setting does not fully replicate real-world deployment conditions, where ambient lighting, motion artifacts, and sensor alignment variability can influence spectral readings. Preliminary evaluations from our previous work [
16] indicated stable performance under ambient lighting conditions, but further studies are needed to assess robustness under dynamic, uncontrolled environments. As part of future work, we plan to conduct additional experiments in open settings to evaluate performance across varying illumination, motion, and environmental conditions, as well as to investigate device-to-device consistency among multiple SENSIPATCH units. Another important consideration is surface geometry and compliance. The current study utilized flat, matte PANTONE color cards to ensure standardized and repeatable measurements. However, many real-world materials, such as skin, textiles, plastics, and food items, exhibit curvature, texture, gloss, or mechanical compliance. The system’s use of multiple directional LEDs and normalization techniques inherently provides some tolerance to such geometric and structural variations. Nonetheless, the current study is limited to flat, matte samples, and future work will include dedicated experiments on non-planar, textured, and compliant surfaces to extend the system’s applicability to more realistic materials and practical use cases.
Model-wise, while the MLP performed effectively in this study, further optimization is necessary to ensure compatibility with ultra-low-power microcontroller platforms suitable for embedded wearable applications. We also acknowledge that the current hyperparameter tuning relied on empirical experimentation. Future iterations will incorporate systematic optimization strategies, such as grid search, random search, or Bayesian optimization, to improve reproducibility and performance. Additionally, enhancing model interpretability will be an important step, particularly for applications in biomedical or safety-critical domains. Finally,
Table 14 compares the main technical characteristics of our system with the latest generation of commercial devices. It should be noted that SENSIPATCH is larger than the other products used for comparison, but this is due to the fact that some biomedical applications require large areas of investigation. Although scanning times are comparable, the present prototype is lighter, less expensive, and multi-purpose. Finally, thanks to the LEDs chosen, the range of wavelengths covered is wider.