Next Article in Journal
Development and Validation of an LC-MS/MS Method for the Quantitation of JNJ-64619178 (JNJ) in Mouse Plasma: Characterization of In Vitro and In Vivo Pharmacokinetic Properties
Next Article in Special Issue
Cold-Adapted Uric Acid-Degrading Lacticaseibacillus paracasei NEFU-6 Application in Kimchi “Paocai
Previous Article in Journal
Plant-Derived Modulators of Tumor Metabolism as Novel, Efficacious, and Low-Toxicity Therapeutic Agents for Cancer Treatment
Previous Article in Special Issue
GC-MS Analysis of Volatile Differences in Rice and Qingke Noodles Formulated with Functional Root Plant Flours
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Near-Infrared Spectroscopy-Based Discriminant Analysis for the Classification of Coffee Quality in Dry Parchment and Green Coffee

by
Claudia Rocio Gómez
*,
Aristófeles Ortiz
and
Valentina Osorio Pérez
National Center for Coffee Research, Cenicafé Km4 Chinchiná, Manizales 170009, Colombia
*
Author to whom correspondence should be addressed.
Molecules 2026, 31(9), 1395; https://doi.org/10.3390/molecules31091395
Submission received: 5 March 2026 / Revised: 6 April 2026 / Accepted: 20 April 2026 / Published: 23 April 2026
(This article belongs to the Special Issue 30th Anniversary of Molecules—Recent Advances in Food Chemistry)

Abstract

This study evaluates the potential of near-infrared spectroscopy (NIRS) combined with discriminant analysis to classify coffee quality based on sensory defects in dry parchment coffee (DPC) and green coffee. Spectral data were used to develop classification models, which were validated using both cross-validation and independent external datasets. Model performance was assessed using classification accuracy and Cohen’s kappa coefficient. The results demonstrate high classification accuracy for DPC (93.5%), with a Kappa coefficient indicating almost perfect agreement (κ = 0.90). In contrast, green coffee showed lower predictive performance (82.4%) and moderate agreement (κ = 0.55), reflecting the greater physicochemical complexity of this matrix. Importantly, the findings demonstrate that coffee quality can be reliably classified at the dry parchment stage, enabling early quality assessment without additional processing steps. This represents a significant advancement compared to previous studies, which have mainly focused on green or roasted coffee. Overall, these results highlight the potential of NIRS as a rapid, non-destructive, and objective tool for coffee quality assessment, with strong applicability in quality control and decision-making processes along the coffee production chain.

1. Introduction

Colombian coffee is widely recognized for its high quality, and this distinct sensory quality depends on multiple factors. The first factor is the composition of chemical compounds in coffee beans, which are transformed during roasting [1]. Other relevant factors include coffee variety, environmental conditions (such as climate and soil), and agronomic and postharvest practices, such as cultivation, crop management, and processing [2]. The sensory quality of coffee is evaluated through sensory analysis, which is carried out by trained judges according to specific attributes and standardized scoring protocols. Sensory analysis enables the identification of defective or off-flavor coffees, characterization of attributes, evaluation of intensity, and classification based on quality profiles [3,4]. Sensory defects are associated with inadequate practices, such as the collection of green beans, uncontrolled fermentation, drying interruptions or prolonged storage, which generate flavors such as immature, vinegary, musty, onion-like, earthy, vegetal, and papery (cardboard-like) [3,5,6,7].
Puerta [8] carried out a sensory quality analysis of samples from the Coffee Grower Cooperatives, over five months, which revealed the following sensory defects, in descending order: fermented (36%), woody (19%), immature (17%), and dirty-rough (14%), followed by aged and contaminated. Similarly, Osorio et al. [4] classified the main defects in Colombian coffee into four groups: fermented, rough, earthy and contaminated, which include the defects mentioned above [4,8]. Regarding the “aged” defect, Gallego et al. [5] performed a discriminant analysis and found that the influencing variables are the threshing yield factor (less than 71), stearic acid content (7.23–7.53%) and polyphenoloxidase (PFO) activity (0.000167–0.00019 U). Rendón et al. [9] also analyzed the “aged” defect and reported changes in the content of fatty acids (1.4–3.8 mg/g), lipid oxidation measured as TBARS (8.8–1.2 nmol MDA/g) and carbonyl groups (2.6–3.5 nmol/mg protein) [5].
Farah et al. [3] performed a correlation analysis between sensory quality and chemical composition. They reported that, for green coffee, trigonelline content decreased as the degree of sensory defects increased, with values ranging from 1.34 to 0.96 g/100 g (db). Another correlated compound class was total chlorogenic acids, which showed lower levels in sound green coffee (5.78 g/100 g db) and higher levels (7.02 g/100 g db) in defective samples. Pabón et al. [10] conducted a principal component analysis (PCA) and reported that earthy and fermented sensory defects are related to physical defects such as brocaded, black and vinegar beans. Regarding the chemical composition of green coffee, several authors have studied the relationships between chemical components and sensory attributes [11,12,13]. The lipid fraction of coffee, which includes total lipids and free fatty acids has been identified as an important determinant of sensory quality. Lipids and fatty acids act as carriers of aroma and flavor compounds and contribute to the body and texture of the beverage. In addition, specific fatty acids, such as linoleic acid and oleic acid, have been associated with sensory attributes such as aroma and acidity [14,15,16,17].
Alkaloids, particularly caffeine and trigonelline, represent an important class of chemicals compounds in coffee. These compounds are primarily associated with bitterness and, to a lesser extent, with other sensory attributes, such as aroma and acidity [18,19,20,21,22]. Sugars also play a key role in sensory quality, as they act as precursors of flavor, aroma, acidity and color. During roasting, sucrose is hydrolyzed into reducing sugars such as glucose and fructose, which contribute significantly to the taste and aroma of coffee [12,23,24]. Chlorogenic acids are considered precursors of the taste, acidity, astringency and bitterness in coffee. These compounds are naturally present in coffee beans, and the concentration of chlorogenic acids varies depending on coffee variety and roasting conditions [25,26,27].
Near-infrared spectroscopy (NIRS) is an analytical tool widely used for the evaluation, quality control and determination of chemical properties in agricultural products [28]. In coffee, NIRS has been used to differentiate varieties, roasting degrees, adulteration types, and both physical and sensory attributes as well as to quantify major chemical compounds; these applications have been developed for green coffee and roasted coffee [7,28,29,30,31,32,33,34,35,36]. The NIRS technique has evolved significantly in recent decades and is considered a precise and reproducible analytical method for both qualitative and quantitative analysis in the pharmaceutical, agri-food and chemical industries [37,38]. Its main advantages include rapid analysis, minimal or no sample preparation, non-destructive measurement, environmental friendliness, and the absence of chemical reagents [39,40].
To achieve the objectives of this research, sample data were analyzed using discriminant statistical analysis, evaluating the relationship between absorbance at each wavelength and sensory quality. The results were used to determine the linear combination of independent variables that maximized the separation between predefined classes [41,42]. In this type of calibration model, a confusion matrix is generated, allowing for the evaluation of classification performance. Correctly classified samples are referred to as true positives, whereas incorrectly classified samples are considered false positives. These results are used to calculate the percentage of success of the model [34,43,44].
This study presents calibration models developed to identify samples with and without sensory defects in both dry parchment coffee and green coffee. Additionally, the chemical composition of two evaluated coffee matrices (WSD and NSD) was analyzed.
Despite the increasing application of NIRS in coffee analysis, most studies have focused on single matrices, such as green or roasted coffee, and have typically addressed either chemical composition or quality classification independently. Moreover, limited studies have explored the integration of chemical and sensory information within a unified modeling framework. In addition, few studies have evaluated the performance of discriminant models across different coffee matrices under comparable conditions or explored alternative approaches such as RMS X residual-based model selection.
In this context, the present study aims to develop and evaluate NIRS-based discriminant models for the classification of coffee quality in both dry parchment coffee (DPC) and green coffee. The study also investigates the relationship between chemical composition and sensory defects, providing a more comprehensive and robust approach for coffee quality assessment under practical conditions.

2. Materials and Methods

The materials and methods selected for this research were designed to ensure an objective, reproducible, and consistent evaluation of coffee quality based on sensory defects and their correlation with the spectral characteristics. In this study, groups of coffee samples were initially established based on differences in quality and the most prevalent sensory defects at the country level. Subsequently, the samples were collected and subjected to sensory analysis by internationally trained and certified panels to ensure the reliability of the reference samples. The analysis was then carried out using near-infrared spectroscopy (NIRS), obtaining the spectral fingerprint of each sample. To develop the classification models, a principal component analysis (PCA) was first applied to explore the data structure and determine the potential separation between groups. Once the classification trends associated with quality were identified based on the spectral fingerprints, multiple discriminant models were evaluated, and the one with the best performance was selected. Finally, the performance of the selected model was evaluated, including cross-validation and independent external validation. Additionally, to complement the interpretation of the differences observed in the spectral fingerprints.

2.1. Coffee Samples

A total of 2054 dry parchment coffee (DPC) and 3834 green coffee samples were collected over a 23-month period, providing a comprehensive representation of Colombia’s diverse coffee-growing regions. The sampling frame encompassed 16 departments: Antioquia, Boyacá, Caldas, Cauca, César, Cundinamarca, Huila, Meta, Nariño, Norte de Santander, Quindío, Risaralda, Santander, Tolima, and Valle del Cauca. Samples were systematically characterized based on their sensory attributes, focusing on defect identification and quality profiling. Based on the sensory evaluation, samples were categorized into two primary groups:
With Sensory Defects (WSD): This group comprised four subgroups, over-fermented, rough, earthy, and contaminated, representing the most prevalent sensory defects in Colombian coffee.
No Sensory Defects (NSD): Samples in this group were classified into three distinct quality profiles (Table 1). Two coffee matrices, DPC and green coffee, were analyzed by NIRS.

2.2. Sensory Analysis

The sensory analysis of the samples was carried out by Q-Grader tasters certified by the Coffee Quality Institute (CQI), following a standardized laboratory protocol and the quality parameters established by the Specialty Coffee Association (SCA) [45]. The samples were carefully prepared prior to evaluation through threshing, selection of defect-free green beans, controlled roasting under standardized conditions, and subsequent cupping analysis. All sensory evaluations were conducted under controlled conditions to ensure the consistency and reproducibility of the results. The panel consisted of certified Q-Grader tasters with proven experience in coffee quality evaluation, ensuring high reliability of the sensory reference data. In addition, all tasting sessions adhered to standardized protocols for fragrance/aroma, flavor, aftertaste, acidity, body, balance, and defect identification. All participants voluntarily enrolled in the study after providing their written informed consent, in accordance with the ethical standards established by the institution.

2.3. Near-Infrared Spectroscopy (NIRS) Analysis

The samples were analyzed with NIRS XDS RCA (2012) equipment (FOSS, Hillerød, Denmark) in the wavelength range of 400–2498 nm; 110 g of DPC and green coffee were weighed into a rectangular reflectance cell 16 cm long, 5.0 cm wide and 5.0 cm high, with a quartz window. The samples were analyzed in duplicate, using the ISIscan program integrated with the analytical instruments. The number of samples used for model development and validation for each coffee matrix is presented in Table 2. This distribution reflects the natural composition of the dataset, including both samples with sensory defects (WSD) and without sensory defects (NSD), as well as the defined quality profiles. The relatively large number of samples per class strengthens model robustness and enables capturing the inherent variability of coffee under real production conditions.

2.4. Model Development

For spectral analysis and calibration model development, WinISI software (version 4, Infrasoft International, State College, PA, USA) was used. Spectral acquisition was performed using 32 scans per sample, and the resulting spectra were averaged to improve the signal-to-noise ratio. Background correction was carried out automatically by the instrument using a reference standard prior to spectral acquisition.
Prior to model development, principal component analysis (PCA) was performed to explore the structure of the spectral data and identify potential outliers. PCA is widely applied in NIR spectroscopy to reduce data dimensionality and detect anomalous samples based on spectral variability [28,40]. To improve spectral quality and reduce non informative variability, preprocessing techniques, including standard normal variate (SNV) and detrending (DT), were implemented. These methods correct scattering effects and baseline shifts in heterogeneous agricultural matrices such as coffee [46,47,48]. Outliers were initially identified using Mahalanobis distance (global H, GH > 3.0), a well-established criterion in NIRS to detect samples that deviate significantly from the population [49]. These samples were further evaluated using spectral residuals and leverage values to ensure robust detection and were subsequently excluded from the calibration dataset. Different mathematical treatments, including derivatives, wavelength selection, and smoothing, were evaluated. Derivative preprocessing enhances spectral resolution and helps remove baseline effects, whereas smoothing reduces high-frequency noise [49]. Additionally, different combinations of spectral pre-treatments and mathematical configurations available in the WinISI software were evaluated, including variations in derivatives, smoothing, and wavelength selection. Model performance was compared using several indicators provided by the software, such as PLS2 (%), correlation coefficient (R), maximum distance, Mahalanobis distance (GH), maximum spectral residual, and RMS X residual [50]. The final combination of treatments was selected based on its contribution to model stability, classification performance, and consistency between calibration and validation results [51]. After outlier removal, samples were randomly assigned to calibration and validation sets without enforcing a fixed proportion, in order to preserve the natural variability of the dataset. This strategy reflects real-world conditions, where class distributions are not necessarily balanced, and contributes to the robustness of the models.
Model selection was based on the consistency between calibration and validation results, prioritizing models that showed similar performance in both datasets. This approach reduces the risk of overfitting and ensures robust predictive capability under independent conditions [49]. Model performance was optimized by maximizing classification accuracy for each class (WSD and NSD) and for the defined quality profiles [50,52]. Performance was internally evaluated using cross-validation during model development and further assessed using an independent subset of samples reserved for validation [51]. Sensitivity and specificity were also calculated from the confusion matrix to evaluate the model’s ability to correctly identify samples with and without sensory defects. Sensitivity represents the proportion of correctly classified positive samples, whereas specificity reflects the proportion of correctly classified negative samples [53,54]. These complementary metrics provide a more comprehensive evaluation of classification performance, particularly in datasets with class imbalance, which is common in agricultural applications [55,56,57,58].

2.5. Model Performance Evaluation

Model performance was evaluated using overall accuracy and confusion matrix analysis. The confusion matrix provides a detailed representation of classification results by comparing reference (observed) and predicted classes, allowing the identification of correctly classified samples (true positives) and misclassified samples (false positives) [43,44].
In addition, Cohen’s Kappa coefficient (κ) was calculated to assess the agreement between predicted and reference classifications while accounting for agreement occurring by chance [59,60]. The coefficient was calculated as follows:
κ = (Po − Pe)/(1 − Pe)
where Po represents the observed agreement (overall accuracy), and Pe corresponds to the expected agreement by chance, calculated from the marginal totals of the confusion matrix. The interpretation of Kappa values followed the criteria proposed by Landis and Koch, where κ < 0 indicates poor agreement, 0.00–0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, and 0.81–1.00 almost perfect agreement.

2.6. External Validation of NIRS Models

External validation was performed using an independent set of samples that were not included in the calibration process. These samples were reserved prior to model development and used exclusively to evaluate model performance. The samples used in this study consisted of 69 defective and 55 non-defective dry parchment coffee samples, as well as 79 defective and 49 non-defective green coffee samples.

2.7. Prediction of Chemical Compounds by NIRS

Once the spectral information of the samples was obtained, the data were analyzed using calibration models to predict 10 chemical compounds in green coffee, using WinISI software (version 4, Foss Infrasoft International, USA). The performance of the quantitative models was evaluated using the relative prediction error, defined as:
Relative error = |1 − (X_NIR/X_ref)|
where X_NIR represents the mean value predicted by NIRS and X_ref corresponds to the mean value obtained using the reference laboratory method. This parameter is a dimensionless ratio that expresses the relative deviation between predicted and reference values. Values close to zero indicate higher model accuracy, whereas higher values indicate lower predictive performance. The results obtained for each compound are presented in Table 3 [33].

3. Results and Discussion

3.1. Quality of the Coffee Samples

The analyzed samples originated from 16 departments of Colombia: Antioquia, Boyacá, Caldas, Cauca, César, Cundinamarca, Huila, Meta, Nariño, Norte de Santander, Quindío, Risaralda, Santander, Tolima and Valle del Cauca. The samples were analyzed over 23 months. The DPC and green coffee samples were characterized by defects and quality. Descriptive statistical analysis revealed that for both coffee matrices, the most common defect class was earthy (Figure 1a and Figure 2a), and the most common sensory profile was Profile 2 (Figure 1b and Figure 2b). The most common defects identified such as over-fermentation and earthy aromas are primarily associated with processing issues post-harvest, as well as deficiencies in infrastructure and process control at the farm level. These factors can lead to inappropriate fermentation conditions, deficient drying practices, and contamination, which ultimately affect the sensory quality of the coffee.

3.2. Spectral Analysis

Figure 3 shows the mean absorbance spectra for dry parchment coffee (DPC) and green coffee samples. Green coffee (represented by purple and green lines) exhibited higher absorbance values across the entire wavelength range compared to DPC (represented by yellow and red lines). The main absorption peak for green coffee was observed at approximately 1940 nm, whereas for DPC it was located around 1920 nm. These differences are primarily associated with variations in moisture content between matrices. The spectral region around 1900–1950 nm corresponds to O–H stretching vibrations, which are strongly influenced by water content [39,40]. Green coffee typically retains higher moisture levels than DPC, which explains the increased absorbance in this region.
In addition to water-related absorptions, differences in chemical composition—particularly lipids and carbohydrates—also contribute to spectral variability through C–H and O–H vibrational overtones [28,40]. Structural differences, such as the presence of the parchment layer in DPC, may further influence light scattering effects, affecting both absorbance intensity and peak position. These spectral differences are relevant for model development, as they provide the underlying variability required for the discrimination between samples with and without sensory defects. Variations associated with moisture and lipid-related absorption bands have been previously linked to compositional and quality differences in coffee and other agricultural matrices [28,40].
To further explore the internal variability of the dataset, PCA was applied separately to samples with sensory defects. PCA of the spectral data for dry parchment coffee (DPC) and green coffee samples identified 24 anomalous samples out of 2098 for DPC (Figure 4a) and 32 out of 4622 for green coffee (Figure 4b), which were excluded from further analysis.
The score plots (Figure 4) show a clear separation between samples with (WSD) and without sensory defects (NSD) for both matrices. The first two principal components explained 95% of the variance for DPC and 80% for green coffee, indicating that most of the relevant spectral information is captured within a reduced dimensional space.
The variability described by PC1 is mainly associated with compositional differences influencing NIR absorption, particularly those related to water and organic constituents such as lipids and carbohydrates [28,40]. The lower variance explained in green coffee reflects a more heterogeneous matrix, likely due to differences in composition and structural characteristics among samples.
In contrast, PCA applied to samples with sensory defects (WSD) did not reveal a clear internal structure (Figure 5). The absence of clustering suggests that defective samples do not share a uniform spectral pattern, which is consistent with the diverse physicochemical origins of sensory defects. Processes such as lipid oxidation, fermentation, and degradation of carbohydrates may generate overlapping spectral responses in the NIR region [28,40]. This behavior indicates that unsupervised methods such as PCA are limited for differentiating variability within defective samples, although they remain useful for initial data exploration.
For samples without sensory defects (NSD), a different pattern was observed (Figure 6). In both DPC and green coffee, samples showed a tendency to separate into two groups: Group 1 (Profiles 1 and 2) and Group 2 (Profile 3). This separation reflects underlying differences in compositional and quality attributes within the NSD group. Profile 3 forms a more distinct cluster, suggesting a more consistent spectral signature compared to Profiles 1 and 2. Taken together, these results indicate that NIR spectroscopy can capture not only the presence or absence of sensory defects, but also variability within non-defective samples, supporting its application for classification into different quality profiles. Coffees with high sensory profiles are associated with higher sugar concentrations, better preservation of lipid composition, and alkaloids key precursors in the thermal reactions that occur during roasting. These compounds play a fundamental role in Maillard reactions, caramelization, and the generation of volatile aromatic compounds, which directly influence flavor complexity, aroma intensity, and mouthfeel. Consequently, the chemical composition of the green coffee matrix is a determining factor in its potential to express high sensory quality after roasting.

3.3. Development of Models

For each evaluated matrix and for the groupings obtained from PCA derived from the spectral analysis of the samples, discriminant classification models were developed and their performance indicators were compared. This approach enabled model selection based on quantitative and comparative criteria. Table 4 summarizes the results of the different model combinations generated using WinISI software, including the evaluated spectral pre-treatments and performance indicators. These results provide a comparative framework for the selection of the optimal model.
Among the evaluated strategies, the model based on the RMS X residual discriminant approach showed the best overall performance, achieving higher classification rates and a clearer separation between classes. This method classifies samples based on the root mean square (RMS) of spectral residuals, minimizing the distance between unknown samples and predefined class models. Its suitability for this study is associated with its ability to handle complex matrices such as coffee, where differences in composition and structural characteristics among samples influence spectral variability [49]. Overall, these results indicate improved discriminant capability and highlight the robustness of the selected model in capturing such differences.
Finally, it is important to contextualize the results obtained with the RMS-based approach against conventional classification methods widely employed in spectral data analysis, such as PLS-DA (Partial Least Squares Discriminant Analysis) and SVM-DA (Support Vector Machine Discriminant Analysis), which have demonstrated high performance in the classification of complex spectral datasets [61,62]. These supervised methods require careful calibration, validation, and parameter optimization to achieve reliable performance [63,64]. In contrast, the RMS-based method leverages the global variability of the spectral signal without the need for supervised training. This simplifies the analytical workflow, reduces methodological complexity, and facilitates interpretation of the results. While RMS-based classification may not always achieve the peak predictive accuracy of fully optimized PLS-DA or SVM-DA models, it provides a rapid and robust alternative, making it particularly suitable for applications that require fast and straightforward sample discrimination.
In this context, the RMS-based method is not intended to replace traditional supervised models but rather to complement them as an effective exploratory tool for the initial differentiation of samples with compositional and structural differences.

3.4. Discriminant Classification of Sensory Quality in Dry Parchment Coffee (DPC) and Green Coffee

The model with the best performance (Table 4) was further evaluated through confusion matrices to assess its classification ability. For dry parchment coffee (DPC), the model showed a high level of accuracy, with 997 true positives and 2 false positives for the NSD class (99.8%), and 404 true positives and 59 false positives for the WSD class (87.3%). The lower accuracy observed in samples with sensory defects suggests greater complexity in the spectral response of this group. In DPC, this variability may be mainly associated with physical factors related to the parchment structure. The presence of this layer, primarily composed of cellulose and hemicellulose, generates scattering effects in NIR radiation that can alter the spectral signal and reduce the model’s sensitivity to detect subtle differences between classes, as widely reported in heterogeneous solid matrices [49].
Cross-validation showed consistent behavior, with an overall accuracy above 90% and classification values similar to those obtained during calibration. The Kappa coefficient (κ = 0.90) indicated almost perfect agreement, confirming the robustness and stability of the model across different data subsets [59]. These results represent a significant advancement, as most NIR spectroscopy studies have focused on green or roasted coffee, whereas its application to dry parchment coffee has been limited. The ability to classify quality at this stage reduces the need for subsequent processes such as hulling, roasting, or sensory analysis, representing an operational advantage in terms of time, cost, and sample handling.
For green coffee, the confusion matrix showed an overall accuracy of 82.4%, with 84.5% for defective samples and 80.2% for non-defective samples, indicating lower discriminant performance compared to DPC. This decrease in performance can be attributed to the higher physicochemical heterogeneity of green coffee, particularly in terms of moisture content, chemical composition, and cellular structure, which directly affect the interaction of NIR radiation with the sample. As described by Osborne et al. [28] and Workman and Weyer [40], in complex biological matrices, the overlap of absorption bands associated with O–H, C–H, and N–H bonds—related to water, lipids, and carbohydrates—hampers the differentiation of specific spectral signals, reducing model sensitivity.
Cross-validation showed an overall accuracy of 78.4%, consistent with calibration results but lower than that observed for DPC, confirming reduced model stability for this matrix. This behavior has been previously reported in NIRS studies on green coffee. For example, Tolessa et al. [52] developed PLS models for sensory attributes such as acidity, sweetness, and aroma, reporting variability in predictive performance depending on the evaluated attribute. Similarly, Ribeiro et al. [53] showed that model accuracy depends on both sample complexity and the specific attribute analyzed. The Kappa coefficient (κ = 0.55) indicated moderate agreement, reinforcing the influence of intrinsic heterogeneity of green coffee on model performance. Overall, these results suggest that, although NIR spectroscopy is a useful tool for early-stage classification, its discriminant capacity in green coffee is limited by the compositional and structural complexity of the matrix. Recent studies provide further insight into the limitations observed in green coffee. Although near-infrared spectroscopy has demonstrated the ability to predict sensory attributes, predictive performance has been shown to depend strongly on the specific attribute and the intrinsic variability of the samples, particularly in less processed matrices [65]. Furthermore, large-scale FT-NIR studies have highlighted that spectral variability in green coffee is strongly influenced by compositional heterogeneity, requiring advanced preprocessing strategies to enhance signal discrimination. This complexity, associated with overlapping absorption bands and structural variability, has been identified as a key factor limiting model performance, which is consistent with the reduced discriminant capacity observed in the present study [66].

3.5. Classification by Quality Profiles

As previously mentioned, discriminant classification models were developed for both dry parchment coffee (DPC) and green coffee based on sensory quality profiles. Based on principal component analysis (PCA), two groups were defined: Group 1 (quality profiles 1 and 2) and Group 2 (quality profile 3) (Figure 6). For DPC, the model achieved an overall accuracy of 91.5%, with outstanding performance for Group 1 (99.2%) and lower accuracy for Group 2 (83.7%). Cross-validation confirmed this trend, with success rates of 90.7% and 79.4%, respectively, demonstrating the ability of NIRS to differentiate quality profiles in this matrix. The Kappa coefficient (κ = 0.87) indicated almost perfect agreement. The lower classification performance observed in Group 2 may be associated with the influence of the parchment structure, which, as reported for heterogeneous solid matrices [49], introduces scattering effects that hinder the detection of subtle spectral differences.
In green coffee, the model achieved an overall success rate of 94.6%, with a global error of approximately 5%. Unlike DPC, Group 2 showed near-perfect accuracy (99.7%), while Group 1 showed lower performance (89.6%). This trend was consistent in cross-validation (96.0% and 92.2%, respectively). The Kappa coefficient (κ = 0.85) also indicated almost perfect agreement. The relatively lower performance for Group 1 suggests that higher-quality samples may exhibit greater chemical and sensory variability, making spectral discrimination more challenging, consistent with the compositional complexity of green coffee described by Osborne et al. [28] and Workman and Weyer [40].
These findings confirm the potential of NIR spectroscopy for classifying coffee according to sensory quality profiles. In this context, Tolessa et al. [52] demonstrated that PLS-based models can predict sensory attributes with variable performance depending on the attribute evaluated, while Baqueta et al. [54] reported that prediction accuracy varies across individual sensory attributes. Together, these studies support the findings of the present work and reinforce the applicability of NIRS for objective coffee quality assessment. Recent literature reinforces the capability of near-infrared spectroscopy combined with multivariate analysis to model complex relationships between spectral signatures and sensory perception in coffee, supporting its application as a rapid and non-destructive tool for quality evaluation [65]. Recent studies have also demonstrated the potential of NIR spectroscopy combined with chemometric modeling for classification tasks related to post-harvest processing and coffee traceability. These approaches highlight the applicability of spectroscopic techniques not only for quality assessment but also for monitoring processing conditions along the value chain, reinforcing the broader applicability of NIRS-based models in the coffee sector [67]. These results highlight the potential of NIRS-based approaches not only for classification purposes but also for supporting decision-making processes in coffee quality control systems, particularly at early stages of the production chain.

3.6. External Validation

To evaluate the robustness of the model beyond calibration data, an external validation was performed using an independent set of samples not included in model development. For dry parchment coffee, 69 defective and 55 non-defective samples were used, achieving an overall accuracy of 91.1%. For green coffee, 79 defective and 49 non-defective samples were evaluated, obtaining an overall accuracy of 92.2%. The results were consistent with those obtained during calibration and cross-validation, confirming the stability, generalization capacity, and practical applicability of the proposed models. Furthermore, the models did not exhibit overfitting and effectively captured the relevant variability of the analyzed matrices. The robustness observed in the external validation is consistent with recent developments in NIR-based modeling approaches. The use of independent validation sets has been emphasized as a critical step to ensure model generalization and avoid overfitting, particularly in spectroscopic applications involving complex biological matrices [65]. In addition, studies based on large and diverse spectral datasets have demonstrated that model robustness and transferability improve when variability is adequately represented during model development, supporting the reliability of the models obtained in this study [66].

3.7. Chemical Composition of NSD and WSD Green Coffee

The chemical composition of WSD and NSD samples was estimated by the NIRS-based prediction models [35]. The results were first subjected to descriptive statistical analysis to identify trends, and then, multiple comparisons of the means were performed via the Tukey test. The results revealed significant differences for the ten tested chemical compounds. Table 5 shows the results of the analysis. Next, the characteristics of the coffee classes evaluated by groups of chemical compounds are described. The chemical composition of coffee beans is directly associated with their quality, as key components such as sugars, lipids, proteins, and alkaloids act as precursors for the development of flavor and aroma during roasting. Variations in these compounds influence the formation of desirable volatile and non-volatile compounds through Maillard reactions, caramelization, and other thermal processes, which directly affect sensory attributes such as sweetness, acidity, balance, and final sensory quality. Therefore, the intrinsic chemical profile of the green coffee bean is a fundamental determinant of its potential to achieve a high-quality sensory profile.

3.7.1. Alkaloids

The mean caffeine content in the NSD samples was slightly lower than that in the WSD samples, with average values of 1.10% and 1.09%, respectively. These values are within the range reported by other authors, such as Villegas et al. [36], who reported values ranging from 1.03% to 1.52% (db), and Hagos et al. [55], who reported values ranging from 0.96% to 1.23% (db). Regarding trigonelline, the NSD samples had a mean content of 0.80%, and the WSD samples had a content of 0.81% (wb). Similar results were reported by Farah et al. [3], with an average value of 0.96%, and Gallignani et al. [19], with a value of 1.0%.

3.7.2. Sucrose

WSD coffee presented lower average sucrose contents than NSD coffee, with values of 7.39% and 7.55% (wb), respectively. These results are consistent with previous studies carried out on NSD green coffee. Barbosa et al. [56] reported a similar sucrose content in NSD green coffee, with a value of 8.2% (db). Furthermore, Knopp et al. [57] analyzed the sugar content in arabica coffee of different origins and detected higher relative contents of sucrose, with average values between 7.6% and 8.2% (db).

3.7.3. Total Chlorogenic Acids

NSD coffee samples presented a higher content of total chlorogenic acids than WSD coffee samples; the average value obtained for the WSD group was 3.45%, and for NSD coffee, the average was 3.52% (wb). These results differ from those reported by Gallego et al. [5], who found no significant differences in the content of total chlorogenic acids between coffee with the “aged” defect and NSD coffee, with average values of 4.5% and 4.4%, respectively. However, the total CQA contents in NSD coffee are close to those reported by other researchers, such as Gómez et al. [33], who found an average CQA content of 4.2% (db), but lower than those reported by Villegas et al. [36], who found an average value of 5.23%.

3.7.4. Total Lipids

The total lipid content was greater in the WSD coffee samples, with an average value of 12.0%, than in the NSD coffee, which presented an average value of 11.78% (wb). These results coincide with those reported by Gallego et al. [5], who also found significant differences in total lipid content between coffee samples with the “aged” defect and NSD samples, with values of 11.87% and 11.0% respectively. Similarly, Rendón et al. [9] stored coffee for 15 months and determined that the coffee developed an “aged” defect and that the total lipid content increased compared with the initial value of 12%, which established that the increase was due to oxidation.

3.7.5. Free Fatty Acids

WSD coffee presented higher contents of arachidic and stearic fatty acids, with values of 4.04% and 8.70%, respectively, than NSD coffee, which showed values of 3.64% and 8.03% (wb). On the other hand, NSD coffee presented higher contents of oleic, palmitic and linoleic fatty acids, with average values of 9.64%, 43.42% and 35.99% (wb), respectively (Table 5). These results agree with those of Gallego et al. [5], who also identified significant differences in fatty acids in samples with the “aged” sensory defect. In addition, in this study, higher contents of linoleic and stearic fatty acids were detected, with values of 38.25% and 7.53% (db), respectively. Rendón et al. [9] also found a significant increase in fatty acids and related this change to the loss of sensory quality. The prediction accuracy of the NIRS models for each chemical component, expressed as relative prediction error, is presented in Table 3. These values correspond to the overall performance of the calibration models and are not specific to individual sample groups.
Although statistical differences were observed for several chemical compounds (Table 6), the magnitude of these differences was generally low. However, the higher variation observed for arachidic fatty acid suggests a potential link with the formation of sensory defects. Arachidic fatty acid is a long-chain saturated fatty acid associated with the lipid fraction of coffee. Changes in lipid composition can be related to oxidative processes and biochemical transformations occurring during postharvest handling and storage [5,9]. Lipid oxidation is known to generate volatile compounds that negatively affect sensory attributes such as aroma and flavor [9,14,15,16,17]. Additionally, microbial activity during uncontrolled fermentation processes may influence lipid metabolism, contributing to changes in fatty acid composition [5]. Therefore, the observed increase in arachidic fatty acid in WSD samples could be associated with degradation processes linked to the development of sensory defects. These results indicate that, although most chemical differences are small, specific compounds such as arachidic fatty acid may act as indicators of quality deterioration and contribute to the discrimination between coffee samples with and without sensory defects.

3.7.6. Chemical Composition of NSD Green Coffee

The NSD green coffee samples presented one of three quality levels (Table 1): Profile 1, “extraordinary” coffee; Profile 2, special; and Profile 3, standard. The quality profiles presented significant differences (Table 7). Profile 1 and Profile 2 were statistically the same but different from Profile 3 in terms of total lipid content and oleic and stearic fatty acid contents. Profile 3 showed a higher total lipid content (12.0%) than Profile 1 and Profile 2. Osorio et al. [12] reported that the total lipid content in green coffee at different stages of maturity ranged from 9.50% to 11.8%, and Echeverri-Giraldo et al. [68] reported values between 10.53% and 12.81%; thus, the values determined in this research for NSD coffee are within the reference range. With respect to the sucrose content, Profile 1 was significantly different from Profiles 2 and 3; Profile 1 showed a lower average sucrose content (7.44%). The chemical compounds that significantly differed among the three profiles were total chlorogenic acids and arachidic and palmitic fatty acids. The contents for total chlorogenic acids ranged from 3.39% to 3.67%, those for arachidic fatty acids ranged from 3.52% to 3.78%, and those for palmitic fatty acids ranged from 43.09% to 43.63%. The values obtained in this research are within the ranges reported by Osorio et al. [12], who found chlorogenic acid contents between 3.57% and 5.10%, arachidic fatty acid contents between 3.77% and 5.85%, and palmitic fatty acid contents between 38.85% and 45.27%. The compounds that did not show differences in content across the three quality profiles were caffeine, trigonelline and linoleic fatty acid. Osorio et al. [11] obtained caffeine contents of 1.09% to 1.16%, trigonelline contents of 0.87% to 0.88% and linoleic fatty acid contents of 33.03% to 34.9%, confirming the present findings.

3.8. Chemical Composition of WSD Green Coffee

The WSD samples with sensory defect were made up of the presented four main groups of defects identified in the sensory analysis: Group a, overfermented; Group b, rough; Group c, earthy; and Group d, contaminated, as shown in Table 1. The chemical information according to group of sensory defects was analyzed with the Tukey test; the results are presented in Table 8. The compounds that presented significant differences between the identified groups of sensory defects were total lipids and oleic fatty acid, with Group 1 (overfermented) having a higher content (average 12.45%) than the other groups. The group with the highest oleic fatty acid content was Group 4 (contaminated). Table 8 lists the minimum and maximum values, mean, standard deviation and group or result of the Tukey statistical analysis according to chemical compound and group of defects.

Overall Interpretation

The results obtained in this study confirm that near-infrared spectroscopy (NIRS), combined with discriminant analysis, provides a robust framework for classifying coffee samples according to sensory defects and quality profiles. The developed models achieved high classification performance for both dry parchment coffee and green coffee, enabling reliable discrimination between samples with and without sensory defects. Although the differences in chemical composition between classes were generally small, certain compounds, particularly arachidic fatty acid, showed a clearer association with quality deterioration. This suggests that specific components of the lipid fraction may contribute to the differentiation of defective samples.
A notable strength of this work lies in the use of samples collected from multiple coffee-producing regions in Colombia. This diversity introduces variability in environmental and processing conditions, which strengthens the robustness of the models and supports their applicability under real production scenarios. An additional contribution is the successful classification of coffee quality at the dry parchment stage. This finding demonstrates that it is possible to identify sensory quality without requiring threshing, roasting, or cupping, which represents a significant advantage in terms of operational efficiency. Furthermore, the integration of dry parchment coffee and green coffee within the same modeling strategy provides a broader analytical scope compared to previous studies focused on a single matrix. This expands the potential of NIRS as a tool applicable across different stages of coffee processing.

4. Conclusions

Unlike most studies reported in the literature, which have focused on a single matrix and regression models for specific sensory attributes, this study proposes a discriminant classification approach applicable to both dry parchment coffee and green coffee. This approach enables quality evaluation at early stages of processing, particularly in dry parchment coffee, representing a relevant contribution from an operational perspective. The integration of both matrices within a single modeling framework highlights the versatility of NIR spectroscopy and expands its application potential along the coffee production chain. From a chemical standpoint, the observed classification capability is associated with differences in sample composition and structure, which affect the interaction of NIR radiation with O–H, C–H, and N–H bonds related to water, lipids, and carbohydrates.
Overall, the results position NIR spectroscopy as a robust, rapid, and non-destructive tool for objective coffee quality evaluation, with strong potential for implementation in production environments. Furthermore, the developed models show high potential for application in quality control systems at key points in the supply chain, such as collection centers or export ports. In this context, their integration into NIRS equipment would enable rapid and objective verification of Colombian coffee quality prior to export, contributing to product standardization and traceability.

Author Contributions

Conceptualization, study design, and supervision: C.R.G.; methodology, data acquisition, formal analysis, and interpretation: A.O. and V.O.P.; writing—original draft preparation: C.R.G.; writing—review and editing: A.O. and V.O.P. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the National Coffee Research Center (Cenicafé) (Crossref Funder ID 100019597), Project number CAL101007.

Institutional Review Board Statement

The sensory tests were conducted in accordance with institutional ethical guidelines. Written informed consent was obtained from all participants prior to the sensory evaluation.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding authors.

Acknowledgments

The authors are grateful to Álvaro Gaitán (National Federation of Coffee Growers of Colombia); Santiago Jaramillo (National Coffee Research Center (Cenicafé)); Jenny Pabón, Cenicafé cupping panel, and Paola Calderón, Wilson Vargas, and Víctor Castañeda; Ruben Dario Medina in the Biometrics Discipline; and Fernando Osorio (Quality and Tasters of the Regional Centers of Almacafé).

Conflicts of Interest

The authors declare that they have no conflicts of interest.

References

  1. Osorio, V.; Pabón, J. Efecto de las temperaturas y tiempos de tueste en la calidad sensorial del café. Rev. Cenicafé 2022, 73, e73102. [Google Scholar] [CrossRef]
  2. Pohlan, H.; Janssens, M. Growth and Production of Coffee. In Soils, Plant Growth and Crop Production; UNESCO: London, UK, 2010; Volume 3, pp. 102–134. [Google Scholar]
  3. Farah, A.; Monteiro, M.C.; Calado, V.; Franca, A.S.; Trugo, L.C. Correlation between Cup Quality and Chemical Attributes of Brazilian Coffee. Food Chem. 2006, 98, 373–380. [Google Scholar] [CrossRef]
  4. Osorio, V. La calidad del Café. In Guía Más Agronomia, Más Productividad, Más Calidad; Cenicafé: Caldas, Colombia, 2021; pp. 219–234. ISBN 978-958-8490-49-6. [Google Scholar]
  5. Gallego, C.P.; Rodríguez, N. Identificación de algunas variables fisicoquímicas y microbiológicas asociadas con el defecto reposo en el café. Rev. Cenicafé 2021, 72, e72105. [Google Scholar] [CrossRef]
  6. Pabón, J.; Osorio, V. Efecto de la interrupción del secado mecánico en la calidad física y sensorial del café. Rev. Cenicafé 2022, 73, e73201. [Google Scholar] [CrossRef]
  7. Poltronieri, P.; Rossi, F. Challenges in Specialty Coffee Processing and Quality Assurance. Challenges 2016, 7, 19. [Google Scholar] [CrossRef]
  8. Puerta, G.I. Buenas prácticas para la prevención de los defectos de la calidad del café: Fermento reposado fenólico y mohoso. Av. Tec. Cenicafé 2015, 461, 1–12. [Google Scholar]
  9. Rendón, M.Y.; de Jesus Garcia Salva, T.; Bragagnolo, N. Impact of Chemical Changes on the Sensory Characteristics of Coffee Beans during Storage. Food Chem. 2014, 147, 279–286. [Google Scholar] [CrossRef] [PubMed]
  10. Pabón, J.; Osorio, V.; Imbachi, L.C. Calidad Física, Sensorial y Composición Química Del Café Cultivado En El Oriente Del Departamento de Caldas. Rev. Cenicafé 2021, 72, e72202. [Google Scholar] [CrossRef]
  11. Osorio, V.; Álvarez-Barreto, C.I.; Matallana, L.G.; Acuña, J.R.; Echeverri, L.F.; Imbachí, L.C. Effect of Prolonged Fermentations of Coffee Mucilage with Different Stages of Maturity on the Quality and Chemical Composition of the Bean. Fermentation 2022, 8, 519. [Google Scholar] [CrossRef]
  12. Osorio, V.; Medina, R.; Acuña, J.R.; Pabón, J.; Álvarez, C.I.; Matallana, L.G.; Fernández-Alduenda, M.R. Transformation of Organic Acids and Sugars in the Mucilage and Coffee Beans during Prolonged Fermentation. J. Food Compos. Anal. 2023, 123, 105551. [Google Scholar] [CrossRef]
  13. Osorio, V.; Pabón, J.; Gallego, C.P.; Echeverri-Giraldo, L.F. Efecto de las temperaturas y tiempos de tueste en la composición química del café. Rev. Cenicafé 2021, 72, e72103. [Google Scholar] [CrossRef]
  14. Cheng, B.; Furtado, A.; Smyth, H.E.; Henry, R.J. Influence of Genotype and Environment on Coffee Quality. Trends Food Sci. Technol. 2016, 57, 20–30. [Google Scholar] [CrossRef]
  15. Echeverri-Giraldo, L.F.; Ortiz, A.; Gallego, C.P.; Imbachí, L.C. Caracterización de la fracción lipídica del café verde en variedades mejoradas de Coffea arabica L. Rev. Cenicafé 2020, 71, 39–52. [Google Scholar] [CrossRef]
  16. Figueiredo, L.P.; Borem, F.M.; Ribeiro, F.C.; Giomo, G.S.; da Silva Taveira, J.H.; Malta, M.R. Fatty Acid Profiles and Parameters of Quality of Specialty Coffees Produced in Different Brazilian Regions. Afr. J. Agric. Res. 2015, 10, 3484–3493. [Google Scholar] [CrossRef]
  17. Villarreal, D.; Baena, L.M.; Posada, H.E. Análisis de lípidos y ácidos grasos en café verde de líneas avanzadas de Coffea arabica cultivadas en Colombia. Rev. Cenicafé 2012, 63, 19–40. [Google Scholar]
  18. López, J.R. Cuantificación de Cafeína, Trigonelina y Ácido Clorogénico por Cromatografía Líquida de Alta Resolución (HPLC). Bachelor’s Thesis, Universidad de Zamorano, Tegucigalpa, Honduras, 2008. [Google Scholar]
  19. Gallignani, M.; Torres, M.; Ayala, C.; Del, M.; Brunetto, R. Determination of caffeine in coffee by means Fourier transform infrared spectrometry. Rev. Téc. Fac. Ing. Univ. Zulia 2008, 31, 159–168. [Google Scholar]
  20. Ky, C.-L.; Louarn, J.; Dussert, S.; Guyot, B.; Hamon, S.; Noirot, M. Caffeine, Trigonelline, Chlorogenic Acids and Sucrose Diversity in Wild Coffea arabica L. and C. canephora P. Accessions. Food Chem. 2001, 75, 223–230. [Google Scholar] [CrossRef]
  21. Macheiner, L.; Schmidt, A.; Schreiner, M.; Mayer, H.K. Green Coffee Infusion as a Source of Caffeine and Chlorogenic Acid. J. Food Compos. Anal. 2019, 84, 103307. [Google Scholar] [CrossRef]
  22. Puerta, G.I. Composición química de una taza de café. Av. Téc. Cenicafé 2011, 414, 1–12. [Google Scholar]
  23. Bradbury, A.G.W. Chemistry I: Non-Volatile Compounds. In Coffee: Recent Developments; John Wiley & Sons, Ltd.: Hoboken, NJ, USA, 2001; pp. 1–17. ISBN 978-0-470-69049-9. [Google Scholar]
  24. Ginz, M.; Balzer, H.H.; Bradbury, A.G.W.; Maier, H.G. Formation of Aliphatic Acids by Carbohydrate Degradation during Roasting of Coffee. Eur. Food Res. Technol. 2000, 211, 404–410. [Google Scholar] [CrossRef]
  25. Marín, C.; Puerta, G.I. Contenido de ácidos clorogénicos en granos de Coffea arabica L. y C. canephora, según el desarrollo del fruto. Rev. Cenicafé 2008, 59, 7–28. [Google Scholar]
  26. Moreira, R.F.; Trugo, L.C.; de Maria, C.A.; Matos, A.G.; Santos, S.M.; Leite, J.M. Discrimination of Brazilian Arabica Green Coffee Samples by Chlorogenic Acid Composition. Arch. Latinoam. Nutr. 2001, 51, 95–99. [Google Scholar]
  27. Solís, L.; Herrera, C. Desarrollo de un método de análisis para la cuantificación de ácidos clorogénicos en café. Agron. Costarric. 2005, 29, 99–107. [Google Scholar] [CrossRef]
  28. Oliveira, L.; Franca, A. Applications of Near Infrared Spectroscopy (Nirs) in Food Quality Evaluation. In Food Quality: Control, Analysis and Consumer Concerns; Nova Science Publishers, Inc.: Hauppauge, NY, USA, 2011; pp. 131–180. ISBN 978-1-61122-917-2. [Google Scholar]
  29. Adnan, A.; Naumann, M.; Mörlein, D.; Pawelzik, E. Reliable Discrimination of Green Coffee Beans Species: A Comparison of UV-Vis-Based Determination of Caffeine and Chlorogenic Acid with Non-Targeted Near-Infrared Spectroscopy. Foods 2020, 9, 788. [Google Scholar] [CrossRef]
  30. Barbin, D.F.; de Souza Madureira Felicio, A.L.; Sun, D.-W.; Nixdorf, S.L.; Hirooka, E.Y. Application of Infrared Spectral Techniques on Quality and Compositional Attributes of Coffee: An Overview. Coffee—Sci. Technol. Impacts Hum. Health 2014, 61, 23–32. [Google Scholar] [CrossRef]
  31. Correia, R.M.; Tosato, F.; Domingos, E.; Rodrigues, R.R.T.; Aquino, L.F.M.; Filgueiras, P.R.; Lacerda, V.; Romão, W. Portable near Infrared Spectroscopy Applied to Quality Control of Brazilian Coffee. Talanta 2018, 176, 59–68. [Google Scholar] [CrossRef] [PubMed]
  32. Craig, A.P.; Franca, A.S.; Oliveira, L.S.; Irudayaraj, J.; Ileleji, K. Application of Elastic Net and Infrared Spectroscopy in the Discrimination between Defective and Non-Defective Roasted Coffees. Talanta 2014, 128, 393–400. [Google Scholar] [CrossRef] [PubMed]
  33. Gómez, C.R.; Ortiz, A.; Gallego, C.; Echeverri, L.F. Validación de curvas de calibración por NIRS para la predicción de compuestos químicos de café almendra. Rev. Cenicafé 2021, 72, e72204. [Google Scholar] [CrossRef]
  34. Gómez, C.R.; Ortiz, A.; Osorio, V. Predicción del origen regional del café de Colombia a partir de la técnica de espectroscopia de infrarrojo cercano—NIRS. Rev. Cenicafé 2022, 73, e73205. [Google Scholar] [CrossRef]
  35. Gómez, C.R.; Gallego, C.P.; Echeverri, L.F.; Pabón, J.; Ortiz, A.; Osorio, V. Determinación de compuestos químicos del café tostado por Espectroscopia de Infrarrojo Cercano (NIRS). Rev. Cenicafé 2023, 74, e74104. [Google Scholar] [CrossRef]
  36. Villegas, A.M.; Pérez, C.; Arana, V.A.; Sandoval, T.; Posada, H.E.; Garrido, A.; Guerrero, J.G.; Pérez, D.; García, J. Identificación de origen y calibración para tres compuestos químicos en café por espectroscopia de infrarojo cercano. Rev. Cenicafé 2014, 65, 7–16. [Google Scholar]
  37. Green, S.; Fanning, E.; Sim, J.; Eyres, G.T.; Frew, R.; Kebede, B. The Potential of NIR Spectroscopy and Chemometrics to Discriminate Roast Degrees and Predict Volatiles in Coffee. Molecules 2024, 29, 318. [Google Scholar] [CrossRef]
  38. Ramo, L.B.; Nobrega, R.O.; Fernandes, D.D.S.; Lyra, W.S.; Diniz, P.H.G.D.; Araujo, M.C.U. Determination of Moisture and Total Protein and Phosphorus Contents in Powdered Chicken Egg Samples Using Digital Images, NIR Spectra, Data Fusion, and Multivariate Calibration. J. Food Compos. Anal. 2024, 127, 105940. [Google Scholar] [CrossRef]
  39. Büning-Pfaue, H. Analysis of Water in Food by near Infrared Spectroscopy. Food Chem. 2003, 82, 107–115. [Google Scholar] [CrossRef]
  40. Zhu, M.; Long, Y.; Chen, Y.; Huang, Y.; Tang, L.; Gan, B.; Yu, Q.; Xie, J. Fast Determination of Lipid and Protein Content in Green Coffee Beans from Different Origins Using NIR Spectroscopy and Chemometrics. J. Food Compos. Anal. 2021, 102, 104055. [Google Scholar] [CrossRef]
  41. Bolaños, J.D. El método NIR combinado con el análisis quimiométrico PLS-da para determinar la adulteración del aceite de oliva con aceite de girasol. Pensam. Actual 2016, 16, 163–172. [Google Scholar] [CrossRef]
  42. Prades, C.; García Olmo, J.; Romero Prieto, T.; García De Ceca, J.L.; López Luque, R. Aplicación de la tecnología de espectroscopia de infrarrojo cercano (NIRS) a la clasificación por calidad del corcho en plancha. Congr. For. Esp. 2013, VI, 2–13. [Google Scholar]
  43. Ariza-López, F.J.; Rodríguez Avi, J.; Alba Fernández, M.V. Control estricto de matrices de confusión por medio de distribuciones multinomiales. Geofocus Rev. Int. Cienc. Tecnol. Inf. Geográfica 2018, 21, 215–226. [Google Scholar] [CrossRef]
  44. Sánchez, J.M. Análisis de calidad cartográfica mediante el estudio de la Matriz de Confusión. Pensam. Matemático 2016, 6, 9–26. [Google Scholar]
  45. Protocols & Best Practices. Available online: https://sca.coffee/research/coffee-standards (accessed on 2 December 2022).
  46. Barnes, R.J.; Dhanoa, M.S.; Lister, S.J. Correction to the Description of Standard Normal Variate (SNV) and De-Trend (DT) Transformations in Practical Spectroscopy with Applications in Food and Beverage Analysis—2nd Edition. J. Infrared Spectrosc. 1993, 1, 185–186. [Google Scholar] [CrossRef]
  47. García-Olmo, J. Clasificación y Autentificación de Canales de Cerdo Ibérico Mediante Espectroscopía en el Infrarrojo Cercano (NIRS). Ph.D. Thesis, Universidad de Córdoba, Montería, Córdoba, 2009. [Google Scholar]
  48. Isaksson, T.; Næs, T. The Effect of Multiplicative Scatter Correction (MSC) and Linearity Improvement in NIR Spectroscopy. Appl. Spectrosc. 1988, 42, 1273–1284. [Google Scholar] [CrossRef]
  49. Osborne, B.G.; Fearn, T.; Hindle, P.H. Practical NIR Spectroscopy with Applications in Food and Beverage Analysis; Longman Scientific & Technical: Essex, UK, 1993; ISBN 978-0-470-22128-0. [Google Scholar]
  50. Santos, J.R.; Sarraguça, M.C.; Rangel, A.O.S.S.; Lopes, J.A. Evaluation of Green Coffee Beans Quality Using near Infrared Spectroscopy: A Quantitative Approach. Food Chem. 2012, 135, 1828–1835. [Google Scholar] [CrossRef]
  51. Khuwijitjaru, P.; Boonyapisomparn, K.; Huck, C. Near-Infrared Spectroscopy with Linear Discriminant Analysis for Green ‘Robusta’ Coffee Bean Sorting. Int. Food Res. J. 2020, 27, 287–294. [Google Scholar]
  52. Tolessa, K.; Rademaker, M.; De Baets, B.; Boeckx, P. Prediction of Specialty Coffee Cup Quality Based on near Infrared Spectra of Green Coffee Beans. Talanta 2016, 150, 367–374. [Google Scholar] [CrossRef]
  53. Ribeiro, B.B.; Mendonça, L.M.V.L.; Assis, G.A.; Mendonça, J.M.A.; Malta, M.R.; Montanari, F.F. Evaluation of the Chemical and Sensory Characteristics of Coffea canephora Pierre and Coffea arabica L. Blends. Coffee Sci. 2014, 9, 178–186. [Google Scholar]
  54. Baqueta, M.R.; Coqueiro, A.; Valderrama, P. Brazilian Coffee Blends: A Simple and Fast Method by Near-Infrared Spectroscopy for the Determination of the Sensory Attributes Elicited in Professional Coffee Cupping. J. Food Sci. 2019, 84, 1247–1255. [Google Scholar] [CrossRef]
  55. Hagos, M.; Redi-Abshiro, M.; Chandravanshi, B.S.; Ele, E.; Mohammed, A.M.; Mamo, H. Correlation between Caffeine Contents of Green Coffee Beans and Altitudes of the Coffee Plants Grown in Southwest Ethiopia. Bull. Chem. Soc. Ethiop. 2018, 32, 13–25. [Google Scholar] [CrossRef]
  56. de Souza Gois Barbosa, M.; dos Santos Scholz, M.B.; Kitzberger, C.S.G.; de Toledo Benassi, M. Correlation between the Composition of Green Arabica Coffee Beans and the Sensory Quality of Coffee Brews. Food Chem. 2019, 292, 275–280. [Google Scholar] [CrossRef] [PubMed]
  57. Knopp, S.; Bytof, G.; Selmar, D. Influence of Processing on the Content of Sugars in Green Arabica Coffee Beans. Eur. Food Res. Technol. 2006, 223, 195–201. [Google Scholar] [CrossRef]
  58. Barge, M.S.; Garner, W.Y.; Ussary, J.P. Good Laboratory Practice Standards: Applications for Field and Laboratory Studies; American Chemical Society: Washington, DC, USA, 1992; ISBN 0-8412-2192-8. [Google Scholar]
  59. Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [PubMed]
  60. Cohen, J. A coefficient of agreement for nominal scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef]
  61. Barker, M.; Rayens, W. Partial least squares for discrimination. J. Chemom. 2003, 17, 166–173. [Google Scholar] [CrossRef]
  62. Ballabio, D.; Consonni, V. Classification tools in chemistry. Part 1: Linear models. PLS-DA. Anal. Methods 2013, 5, 3790–3798. [Google Scholar] [CrossRef]
  63. Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef]
  64. Wold, S.; Sjöström, M.; Eriksson, L. PLS-regression: A basic tool of chemometrics. Chemom. Intell. Lab. Syst. 2001, 58, 109–130. [Google Scholar] [CrossRef]
  65. Martins, R.N.; de Carvalho Pinto, F.D.A.; de Queiroz, D.M.; de Paula Corrêdo, L.; Cardoso, W.J.; Portes, M.F.; Valente, D.S.M.; Teófilo, R.F. Prediction of coffee sensory attributes using near-infrared spectroscopy. Coffee Sci. 2025, 20, e202369. [Google Scholar] [CrossRef]
  66. Bahamón-Monje, A.F.; Morales-Angulo, E.M.; Collazos-Escobar, G.A.; Gutiérrez-Guzmán, N. Large dataset of Fourier-transform near-infrared (FT-NIR) spectroscopy of green and roasted specialty coffee: Preprocessed spectra and sensory scores for machine learning-based quality monitoring. Data Brief 2025, 60, 111609. [Google Scholar] [CrossRef]
  67. Santos-Rivera, M.; Viswanathan, L.; Sheibani, F. Enhancing coffee quality and traceability: Chemometric modeling for post-harvest processing classification using near-infrared spectroscopy. Spectrosc. J. 2025, 3, 20. [Google Scholar] [CrossRef]
  68. Echeverri-Giraldo, L.F.; Pinzón Fandiño, M.I.; González Cadavid, L.M.; Rodriguez Marín, N.D.; Moreno Ríos, D.A.; Osorio Pérez, V. Determination of Lipids and Fatty Acids in Green Coffee Beans (Coffea arabica L.) Harvested in Different Agroclimatic Zones of the Department of Quindío, Colombia. Agronomy 2023, 13, 2560. [Google Scholar] [CrossRef]
Figure 1. (a) Proportions of defect classes and (b) quality profile types for dry parchment coffee (DPC).
Figure 1. (a) Proportions of defect classes and (b) quality profile types for dry parchment coffee (DPC).
Molecules 31 01395 g001
Figure 2. (a) Proportions of defect classes and (b) quality profiles in green coffee.
Figure 2. (a) Proportions of defect classes and (b) quality profiles in green coffee.
Molecules 31 01395 g002
Figure 3. Mean near-infrared absorbance spectra of dry parchment coffee (DPC) and green coffee samples (400–2498 nm). Sample groups are identified by color as indicated in the legend. The multicolored band (400–700 nm) represents the visible region, while the remaining regions correspond to near-infrared absorptions (e.g., O–H, C–H, N–H).
Figure 3. Mean near-infrared absorbance spectra of dry parchment coffee (DPC) and green coffee samples (400–2498 nm). Sample groups are identified by color as indicated in the legend. The multicolored band (400–700 nm) represents the visible region, while the remaining regions correspond to near-infrared absorptions (e.g., O–H, C–H, N–H).
Molecules 31 01395 g003
Figure 4. Principal component analysis (PCA) of (a) dry parchment coffee (DPC) and (b) green coffee samples, showing the distribution of samples in the principal component space. The axes correspond to selected principal components (e.g., PC1, PC3, and PC4 in (a), and PC3, PC6, and PC9 in (b)), chosen to enhance the visualization of sample separation. Green markers correspond to WSD samples, while blue markers represent NSD samples.
Figure 4. Principal component analysis (PCA) of (a) dry parchment coffee (DPC) and (b) green coffee samples, showing the distribution of samples in the principal component space. The axes correspond to selected principal components (e.g., PC1, PC3, and PC4 in (a), and PC3, PC6, and PC9 in (b)), chosen to enhance the visualization of sample separation. Green markers correspond to WSD samples, while blue markers represent NSD samples.
Molecules 31 01395 g004
Figure 5. Principal component analysis (PCA) results for samples with sensory defects: (a) dry parchment coffee (DPC) and (b) green coffee. The distribution of samples is shown in the principal component space. The axes correspond to selected principal components (PC1, PC2, and PC3 in (a), and PC2, PC3, and PC4 in (b)), chosen to enhance the visualization of sample patterns.
Figure 5. Principal component analysis (PCA) results for samples with sensory defects: (a) dry parchment coffee (DPC) and (b) green coffee. The distribution of samples is shown in the principal component space. The axes correspond to selected principal components (PC1, PC2, and PC3 in (a), and PC2, PC3, and PC4 in (b)), chosen to enhance the visualization of sample patterns.
Molecules 31 01395 g005
Figure 6. PCA results for samples without sensory defects: (a) DPC and (b) green coffee.
Figure 6. PCA results for samples without sensory defects: (a) DPC and (b) green coffee.
Molecules 31 01395 g006
Table 1. Description of the samples analyzed by NIRS.
Table 1. Description of the samples analyzed by NIRS.
MatrixFeatures:
With sensory defects—WSDOverfermented: Pulp, vinegar, fermented, musty, and onion
Rough: Immature, pungent and rough, and cereal-like
Earthy: Earthy, mold and aged
Contaminated: Phenol, smoke, contaminated and chemical
No sensory defects—NSDProfile 1Extraordinary>84 points
Profile 2Special82–84 points
Profile 3Standard79–82 points
Table 2. Number of samples per coffee matrix used for the prediction models.
Table 2. Number of samples per coffee matrix used for the prediction models.
Coffee MatrixModel TypeClass/GroupNumber
of Samples
Dry parchmentWSD and NSD classificationWSD *554
NSD *1500
NSD quality classificationProfile 1 and 21725
Profile 3727
GreenWSD and NSD classificationWSD829
NSD3005
NSD quality classificationProfile 1 and 2901
Profile 3447
* WSD: with sensory defects; NSD: no sensory defects.
Table 3. Relative prediction error by chemical compound.
Table 3. Relative prediction error by chemical compound.
Chemical CompoundRelative Prediction Error *
Total lipids0.003
Caffeine0.015
Trigonellin0.027
Sucrose0.007
Total chlorogenic acids (CQA)0.007
Palmitic fatty acid0.007
Linoleic fatty acid0.005
Oleic fatty acid0.001
Stearic fatty acid0.008
Arachidic fatty acid0.002
* Relative prediction error (dimensionless).
Table 4. Performance of discriminant classification models for different combinations of spectral pre-treatments and coffee matrices.
Table 4. Performance of discriminant classification models for different combinations of spectral pre-treatments and coffee matrices.
Coffee MatrixModel TypeMathematical Treatment *Pls2 (%)Correlation Coefficient (R) (%)Maximum Distance (%)Mahalanobis Distance (GH) (%)Maximum Spectral Residual (%)RMS X Residual (%)
DPCWSD vs. NSD classification3, 3, 3, 1053.139.456.456.993.5
Quality classification
(Group 1 and 2)
4, 3, 3, 15.758.941.959.25791.5
Green coffeeWSD vs. NSD classification1, 4, 4, 145.868.844.451.970.182.4
Quality classification
(Group 1 and 2)
3, 4, 4, 145.562.947.356.869.494.1
* Mathematical treatment is expressed as: derivative order, gap, first smoothing, and second smoothing, according to WinISI software conventions.
Table 5. Chemical composition by type of green coffee estimated by the NIRS technique.
Table 5. Chemical composition by type of green coffee estimated by the NIRS technique.
Chemical CompoundClassMinimum (%)Maximum (%)Mean (%)Standard Deviationp Value
CaffeineNSD0.831.351.10 B0.080.015
WSD0.651.541.09 A0.11
TrigonellinNSD0.660.950.80 B0.050.001
WSD0.651.020.81 A0.06
SucroseNSD6.568.707.55 B0.390.0001
WSD6.108.527.39 A0.43
Total chlorogenic acidsNSD2.844.363.52 B0.370.0001
WSD2.534.513.45 A0.36
Total lipidsNSD9.2714.1111.78 B0.770.0001
WSD8.8915.4312.00 A1.30
Arachidic fatty acidNSD1.905.463.64 A0.690.0001
WSD1.215.244.04 B0.74
Oleic fatty acidNSD8.4210.889.64 B0.410.0001
WSD7.8410.699.26 A0.49
Stearic fatty acidNSD6.289.988.03 A0.720.0001
WSD5.9910.618.70 B0.90
Palmitic fatty acidNSD40.4146.8343.42 B1.200.0001
WSD38.5448.3842.99 A1.76
Linoleic fatty acidNSD28.9639.9135.99 B1.270.0001
WSD25.6240.1935.23 A2.22
p Values correspond to one-way ANOVA. Different letters indicate significant differences according to Tukey’s post hoc test (p < 0.05).
Table 6. Chemical compounds, average contents by coffee class and corresponding percentages (%).
Table 6. Chemical compounds, average contents by coffee class and corresponding percentages (%).
CompoundWSDNSD10% Decimal Average (Difference)Criterion
Total lipids (%)12.0011.781.190.22
Caffeine (%)1.091.100.110.01
Trigonellin (%)0.810.800.080.01
Sucrose (%)7.397.550.750.16
Total chlorogenic acids (%)3.453.520.350.08
Arachidic fatty acid (%)4.043.640.380.40
Oleic fatty acid (%)9.269.640.940.38
Stearic fatty acid (%)8.708.030.840.67
Palmitic fatty acid (%)42.9943.424.320.43
Linoleic fatty acid (%)35.2335.993.560.76
Table 7. Chemical composition of coffee by quality profile.
Table 7. Chemical composition of coffee by quality profile.
Chemical CompoundQuality ProfileMinimum (%)Maximum (%)Mean (%)Standard Deviation
Total lipids19.2714.1111.74 A0.80
29.5014.0011.67 A0.73
39.6514.0512.04 B0.75
Caffeine10.931.311.11 A0.08
20.831.311.11 A0.08
30.881.351.08 A0.09
Trigonellin10.660.920.79 A0.05
20.660.940.80 A0.05
30.700.950.81 A0.05
Sucrose16.608.637.45 A0.40
26.768.707.56 B0.39
36.568.547.58 B0.38
Total chlorogenic acids12.874.183.39 A0.33
22.874.223.50 B0.36
32.844.363.67 C0.39
Arachidic fatty acid11.905.463.78 C0.61
22.025.233.64 B0.67
31.915.083.52 A0.77
Oleic fatty acid18.8310.849.74 A0.36
28.4810.889.68 A0.42
38.4210.619.46 B0.38
Stearic fatty acid16.429.988.16 B0.66
26.289.578.00 B0.68
36.379.777.99 A0.83
Palmitic fatty acid140.7346.3243.09 A1.07
240.4146.8343.43 B1.20
340.5746.2743.63 C1.23
Linoleic fatty acid128.9639.3436.08 A1.36
230.5639.6335.96 A1.23
330.6739.9135.99 A1.30
p Values correspond to one-way ANOVA. Different letters indicate significant differences according to Tukey’s post hoc test (p < 0.05).
Table 8. Chemical composition of WSD coffee by group of sensory defects.
Table 8. Chemical composition of WSD coffee by group of sensory defects.
StatisticalGroupMinimum (%)Maximum (%)Mean (%)Standard Deviation (n − 1)
Total lipidsa9.6715.4312.45 B1.10
b9.3015.0912.09 AB1.29
c8.9615.2611.89 A1.32
d8.8914.3311.89 A1.30
Caffeinea0.781.541.12 A0.10
b0.651.491.11 A0.18
c0.691.351.09 A0.10
d0.701.471.06 A0.12
Trigonellina0.670.980.81 A0.07
b0.690.970.82 A0.07
c0.650.980.81 A0.06
d0.741.020.83 A0.05
Sucrosea6.588.377.31 A0.36
b6.218.377.40 A0.52
c6.108.527.42 A0.42
d6.488.417.33 A0.48
Total chlorogenic acidsa2.534.273.39 A0.33
b2.994.213.47 A0.32
c2.904.513.46 A0.37
d2.984.343.43 A0.35
Arachidica1.215.134.06 A0.74
b2.745.164.22 A0.63
c2.105.244.00 A0.76
d1.905.124.05 A0.75
Oleic fatty acida7.849.899.08 A0.44
b8.3010.479.23 AB0.50
c8.0510.699.29 B0.49
d8.5610.409.39 B0.48
Stearic fatty acida5.9910.408.67 A0.86
b6.7810.368.86 A0.86
c6.5510.618.65 A0.90
d6.3810.508.85 A0.95
Palmitic fatty acida40.1048.0443.12 A1.96
b40.3447.4742.71 A1.62
c39.4548.1843.07 A1.77
d38.5447.4942.66 A1.52
Linoleic fatty acida26.6539.7535.39 A2.24
b28.9839.0734.98 A2.21
c25.6240.1935.18 A2.24
d29.7138.6635.52 A2.04
p values correspond to one-way ANOVA. Different letters indicate significant differences according to Tukey’s post hoc test (p < 0.05).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Gómez, C.R.; Ortiz, A.; Osorio Pérez, V. Near-Infrared Spectroscopy-Based Discriminant Analysis for the Classification of Coffee Quality in Dry Parchment and Green Coffee. Molecules 2026, 31, 1395. https://doi.org/10.3390/molecules31091395

AMA Style

Gómez CR, Ortiz A, Osorio Pérez V. Near-Infrared Spectroscopy-Based Discriminant Analysis for the Classification of Coffee Quality in Dry Parchment and Green Coffee. Molecules. 2026; 31(9):1395. https://doi.org/10.3390/molecules31091395

Chicago/Turabian Style

Gómez, Claudia Rocio, Aristófeles Ortiz, and Valentina Osorio Pérez. 2026. "Near-Infrared Spectroscopy-Based Discriminant Analysis for the Classification of Coffee Quality in Dry Parchment and Green Coffee" Molecules 31, no. 9: 1395. https://doi.org/10.3390/molecules31091395

APA Style

Gómez, C. R., Ortiz, A., & Osorio Pérez, V. (2026). Near-Infrared Spectroscopy-Based Discriminant Analysis for the Classification of Coffee Quality in Dry Parchment and Green Coffee. Molecules, 31(9), 1395. https://doi.org/10.3390/molecules31091395

Article Metrics

Back to TopTop