Next Article in Journal
Effects of Accelerated Fermentation on the Chemical Composition and Quality of Beer
Previous Article in Journal
Isolation and Structural Elucidation of Phytochemicals from Canarium luzonicum Leaves and Evaluation of Anti-Lung Cancer and Antileishmanial Activity
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Machine-Learning-Assisted Carbon Dots: From Algorithms to Applications and Beyond

1
School of Physics Science and Information Technology, Liaocheng University, Liaocheng 252000, China
2
Shandong Jinhengli Mechanical Manufacturing Co., Ltd., Taian 271200, China
3
Centre of Excellence for Nanotechnology, Department of Electronics and Communication Engineering, Koneru Lakshmaiah Education Foundation, Vaddeswaram 522302, Andhra Pradesh, India
*
Authors to whom correspondence should be addressed.
Molecules 2026, 31(10), 1696; https://doi.org/10.3390/molecules31101696
Submission received: 8 April 2026 / Revised: 13 May 2026 / Accepted: 14 May 2026 / Published: 17 May 2026

Abstract

Carbon dots (CDs) have emerged as frontier materials in multidisciplinary research owing to their unique optical properties and physicochemical characteristics. However, issues such as the reliance on trial-and-error experimentation for synthetic preparation and the difficulty in systematically revealing structure–activity relationships persist. In recent years, machine learning (ML) has provided a new paradigm for CD research through its powerful predictive and decision-making capabilities. This review first introduces the fundamental workflow of ML and the operational principles of several representative ML algorithms. It then summarizes the ML applications in CDs, including ML-optimized CD synthesis, ML-assisted detection in CD sensors, ML-based performance prediction, and ML-driven mechanism studies. Finally, the review outlines the future prospects for applications in this field, aiming to further advance the development of nanomaterials science.

Graphical Abstract

1. Introduction

Carbon dots (CDs), an emerging class of zero-dimensional carbon-based nanomaterials, exhibit broad application potential across various fields owing to their outstanding biocompatibility, tunable optical properties, low toxicity, high conductivity, and facile synthesis [1,2,3]. In terms of optical properties, CDs have the characteristics of high photostability, tunable luminescence and broad excitation spectrum, but their photoluminescence mechanism is still controversial and may be related to factors such as conjugated structure, surface state or molecular state [4,5]. In the field of biomedicine, CDs are widely used in bioimaging (such as in vitro cell labeling and in vivo tumor imaging), biosensing (such as the detection of metal ions, small molecules, and biopolymers), and treatment (such as photothermal therapy and drug delivery) [6,7,8]. Additionally, they exhibit excellent catalytic activity and electron mobility in electrocatalysis and energy devices [9]. Typically, at room temperature, CDs can be synthesized via top-down and bottom-up methods. The latter are more widely used because of their superior versatility and accessibility. Among these, environmentally benign methods, such as hydrothermal and microwave-assisted synthesis, are attractive because they use readily available, low-cost raw materials. At low temperatures, carbon dots with high quantum yields and excellent performance can be synthesized for use as fluorescent probes [10,11]. Biomass-derived CDs are regarded as an ideal alternative material for addressing issues such as the high toxicity and poor environmental compatibility of conventional semiconductor quantum dots [12,13]. However, CDs possess a wide variety of surface functional groups (such as amino, hydroxyl, and aldehyde groups), and their differing sizes, shapes, and anchoring methods result in complex structures, often making their synthesis difficult to control [14]. At the same time, optimizing the synthesis conditions of CDs often involves a complex trial-and-error process with limited predictability, leading to extended timelines and reduced synthesis efficiency [15]. Their practical applications still face challenges, including the absence of clear structure–activity relationships, structural inconsistencies, fluctuations in quantum yield, and potential cytotoxicity caused by degradation under light exposure [16,17,18].
As a core discipline of artificial intelligence, machine learning (ML) utilizes algorithmic models to enable computers to autonomously learn patterns and rules from data, thereby achieving prediction, classification, and optimization of complex systems. This approach has significant value in modern scientific research, particularly in contexts where data volume surpasses human analytical capacity or where underlying relationships are inherently complex. Therefore, it has great application potential in research fields such as biomedicine (analyzing genome and clinical data, studying disease mechanisms) [19], materials science (optimizing material properties, designing and discovering various nanomaterials) [20,21,22,23], fluid mechanics [24,25], industrial manufacturing (intelligent upgrades of predictive maintenance, quality control, and supply chain management) [26], and chemical reaction detection [27]. For instance, Zhang et al. [28] achieved the efficient substitution of the scarce element cobalt (Co) by using ML to interpret the influence of elements on alloy properties. Wen et al. [29] successfully screened refractory high-entropy alloys with both excellent strength and room-temperature ductility using a support vector regression model, nondominated sorting genetic algorithm, and K-means clustering algorithm (K-MCA), fully demonstrating the effectiveness of ML-assisted material design. In extreme environments, ML also plays a vital role. Under low-temperature conditions, ML not only predicts the key properties of sealing materials and deep-freeze damage factors in concrete but also identifies Cenchrus fungigraminus and elucidates the drying process of sludge [30,31,32,33]. Under high-temperature and high-pressure conditions, ML has enabled the prediction of material properties across diverse domains, including mechanics (elastic constants, phonon frequencies), thermodynamics (thermal conductivity, free energy), and electricity (superconductivity, electrical conductivity) [34,35,36,37,38].
Currently, ML is applied not only across diverse fields such as biomedicine, materials science, and chemical reaction detection, but it has also demonstrated distinct advantages and potential in CD research. For instance, Han et al. [39] employed the extreme gradient boosting (XGBoost) algorithm to construct human-interpretable decision trees, effectively guiding the synthesis of CDs with a QY of 39.3%. Pandit et al. [40] reported a CD-based sensor array that, when combined with ML techniques, can distinguish between eight proteins with 100% accuracy. Chen et al. [41] developed an ML strategy that successfully predicted the yield of CDs generated during the biochar production process. This review first outlines the workflow of ML and provides a detailed analysis of how typical ML algorithms operate. Subsequently, it comprehensively summarizes the applications of ML in optimizing CD synthesis, assisting CD sensors in detection, predicting performance, and investigating mechanisms (Figure 1). Finally, it discusses the challenges faced in applying ML in the CD field and explores future development prospects.

2. The Workflow of ML

ML is a system that learns from data through training, with its core function being the modelling of data and the optimization of model parameters via algorithms. A typical ML model follows a systematic workflow, which is crucial for studying its applications and principles across various domains. The general workflow of ML primarily encompasses data collection and preprocessing, model selection and training, and model evaluation and prediction.

2.1. Data Collection and Preprocessing

First, the ML process requires the collection and preparation of data as inputs to the algorithm. Generally, data sources include previously published papers, sample sets generated through experiments or simulations, and databases [42,43]. A dataset can be divided into a training set (used for model learning), a validation set (used for parameter tuning and model selection), and a test set (used to evaluate model performance) [44]. To reduce the randomness caused by data partitioning, cross-validation can be performed on the dataset. This involves dividing the dataset into several subsets, then taking turns using one subset as the test set and the rest as the training set, and calculating the average error after multiple iterations. Data quality plays a crucial role. Due to the potential presence of missing values, outliers, and redundant data in the raw data, preprocessing of the collected data is necessary. For example, in their study on the relationship between the physicochemical characteristics and antimicrobial activity of CDs, Bian et al. [45] encountered over 50% missing values for both hydrodynamic size and chemical composition. Consequently, they excluded features with higher proportions of missing values and selected those relevant to the learning process. To enhance data processing speed, enable data visualization, and preserve the most critical features of high-dimensional data, dimensionality reduction is also required. For instance, Talukder et al. [46] employed principal component analysis (PCA) for dimensionality reduction to enhance computational efficiency in intrusion detection for wireless sensor networks. At the same time, data standardization must be performed to scale feature data with different units and ranges to a common scale, ensuring the stability and efficiency of model training. If categorical variables are present in the data, the model cannot directly process this information. Feature encoding is required to convert categorical variables into numerical variables. Common encoding methods include one-hot encoding and label encoding. The former converts each category into a binary variable, which is suitable for unordered categories. The latter directly maps the category to an integer, which is suitable for ordered categories.

2.2. Model Selection and Training

The suitability of an ML model depends on the nature of the problem (classification, regression, clustering, reinforcement learning, and anomaly detection) and the characteristics of the data [47]. ML models can generally be categorized into two types: supervised and unsupervised. Supervised learning trains models using labeled data. In supervised learning, each training example requires both an input object and a corresponding output object, enabling the model to make predictions on new data. It can be categorized into two main types based on the nature of the target variable being predicted: classification and regression. Classification aims to predict a discrete, finite category label, whereas regression predicts a continuous numerical value. Common algorithms include linear regression, linear discriminant analysis (LDA), support vector machines (SVM), random forests (RF), gradient-boosted decision trees (GBDT), and artificial neural networks (ANN) [48,49,50]. For example, Mandal et al. [50] utilized the fluorescence response of carbon nanoparticles (CNPs) to predict heavy metal ions. This prediction was performed by mapping input variables to output variables, thus employing supervised learning. Unlike supervised learning, unsupervised learning is a statistical approach that uncovers underlying structures or distributions from unlabeled data. It relies on the inherent characteristics of the data itself for clustering, dimensionality reduction, and other operations. Cluster analysis groups data points into clusters, where data points within the same cluster are similar to each other. Dimensionality reduction is required to facilitate data analysis and visualization. This process maps high-dimensional data to a low-dimensional space while preserving as much original information as possible. Common algorithms include mean-based clustering, PCA, and hierarchical cluster analysis (HCA). Feature extraction using K-MCA generates K features from the R, G, and B channels, thereby improving the output accuracy of ML models [51]. Applying HCA to partition samples into distinct subsets facilitates the extraction of interpretable patterns, thereby enhancing the accuracy of predictive models [52]. Performing PCA on data reduces dimensionality [53], eliminating the influence of raw data and enabling regression analysis. However, ML sometimes relies excessively on manual feature engineering, making it challenging to handle complex nonlinear relationships. Deep learning, as a subset of ML, can automatically extract features through multi-layer neural networks (such as convolutional neural networks) and is primarily suited for unstructured data.

2.3. Model Evaluation and Prediction

By comprehensively evaluating the performance metrics of a model, one can more systematically assess its predictive accuracy and generalization capability. Generally, for regression problems, three performance metrics are commonly used to evaluate the performance of ML models: the coefficient of determination (R2), mean squared error (MSE), and Pearson correlation coefficient (r) [39]. Common metrics include the mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE) [43]. In classification problems, model performance is often evaluated using metrics such as accuracy, precision, recall, F 1 score, and the area under the receiver operating characteristic curve (AUC-ROC curve). To pursue optimal model generalization, cross-validation and hyperparameter optimization are sometimes employed to prevent overfitting or underfitting.
Precision is the proportion of samples with a true value of 1 among all samples predicted as 1. The equation is
P = T P T P + F P .
Recall is the proportion of samples predicted as 1 among all samples with a true value of 1. The equation is
r = T P T P + F N .
The F 1 score is a metric used to measure the accuracy of binary classification models and simultaneously accounts for both precision and recall. The equation is
F 1 = 2 P r P + r .
T P , F P , and F N refer to true positives, false positives, and false negatives, respectively. For example, Liu et al. [54] employed precision, recall, and the F 1 score as evaluation metrics to validate the performance of the You Only Look Once (YOLO) v3 model, ultimately calculating precision at 92.5%, recall at 100.0%, and the F 1 score at 96.1%. The model demonstrated high accuracy and excellent performance. Similarly, the AUC-ROC curve serves as a crucial tool for evaluating the performance of binary classification models. The ROC curve plots the relationship between the true positive rate and the false positive rate at different classification thresholds. The AUC represents the classifier’s ability to distinguish between positive and negative samples. Specific methods for measuring performance within the model are discussed in the following sections.

2.4. Common ML Algorithms

SVMs are powerful and flexible supervised learning algorithms primarily used for classification and regression tasks. The core idea of this algorithm is to find a decision boundary (hyperplane) with the maximum margin. However, this algorithm can only handle linearly separable data. Kernel functions have been introduced to address linearly inseparable data. For example, Liu et al. [55] employed kernel functions to address the nonlinear problem of identifying multiple lithofacies types. Through a nonlinear mapping, linearly inseparable data in the original low-dimensional space are transformed into a high-dimensional feature space, where the data become linearly separable. This approach flexibly resolves complex nonlinear problems. For example, Liu et al. [42] found that the SVM model performed optimally (test set R2 = 0.927) when using ML to investigate the properties of cement-based materials, enabling the capture of nonlinear trends in the data without complex optimization.
In unsupervised machine learning algorithms, the K-MCA is a fundamental, important, and widely applied learning algorithm. The general procedure involves randomly selecting K data points from a dataset containing N data points as initial cluster centers. The distance between each data point and all K cluster centers is calculated, and each data point is assigned to the nearest cluster center, forming K temporary clusters. Subsequently, the cluster center for each cluster is recalculated, which is the average of all data points within that cluster. The iterative process continues until convergence is reached, at which point the cluster centers no longer undergo significant changes. The quality of clustering can be assessed using the silhouette coefficient; the closer this coefficient is to 1, the better the clustering results. However, this algorithm requires K to be specified in advance; however, it is difficult to guarantee the number of meaningful values that exist in the data during actual processing. Because cluster-centered calculations are based on the average of all points within a cluster, they are highly susceptible to the influence of outliers [56]. Unlike K-MCA, HCA does not require a predefined K value or a specified initial center. It generates a dendrogram by splitting or merging samples layer by layer through top-down splitting and bottom-up aggregation, providing a visual representation of the relationships and classifications among samples [52]. For small sample sizes where exploratory data grouping is required, HCA yields more robust clustering results than K-MCA.
Ensemble learning is a highly significant learning method in ML and has been widely applied in industrial and medical fields in recent years. Decision tree ensemble learning refers to the use of multiple decision tree models of the same type, employing ensemble learning methods to enhance model performance and stability. Common algorithms include RF and GBDT. RF is a method that constructs multiple decision trees and averages their results to reduce overfitting. In contrast, GBDT is an ensemble approach that progressively builds decision trees to correct errors made by the preceding tree.
ML platforms and tools are indispensable for the efficient execution of modern machine learning and deep learning tasks. Representative examples based on the Python ecosystem include scikit-learn and TensorFlow. Scikit-learn is recognized as a premier library for traditional ML, distinguished by its comprehensive documentation and extensive suite of data processing utilities. It is commonly employed for data preprocessing, classification, regression, and dimensionality reduction, and supports the entire workflow—from feature engineering and model training to hyperparameter optimization and cross-validation—on conventional computing servers. For deep learning applications, TensorFlow serves as a primary framework [57]. At its core, TensorFlow represents computations using tensors and organizes operations within a dataflow graph architecture to facilitate model training. Notably, TensorFlow provides developers with the flexibility to design and train custom algorithms, making it particularly suitable for research and industrial-scale projects that demand precise control over model architecture and training procedures.

3. Application of ML in CDs

3.1. ML-Optimized Synthesis of CDs

CDs have garnered significant attention owing to their outstanding optical properties, straightforward synthesis routes, and diverse synthetic approaches [58]. However, the preparation of high-performance CDs not only depends on multiple factors, such as precursors, temperature, and reaction time [1], but even minor variations in the synthesis parameters can influence the properties of CDs [59]. Furthermore, screening CDs with unique optical properties typically relies on extensive trial-and-error experiments, a process that is both time-consuming and costly. Consequently, introducing ML as an auxiliary tool is regarded as an effective solution. By learning from and analyzing existing experimental data, ML methods can build predictive models, optimize parameters, and identify the optimal synthesis conditions, thereby enhancing the performance of CDs. At the same time, this approach can improve the efficiency of synthesizing specific CDs.

3.1.1. Single Performance of ML-Optimized CDs

CDs are widely used in the field of fluorescence sensing due to their excellent optical properties, including tunable emission wavelengths, high fluorescence quantum yield (QY), and good stability. Among these, QY is a key indicator of fluorescence emission efficiency, and understanding the relationships among synthetic parameters is crucial for optimizing QY. Han et al. [39] established a regression ML model for hydrothermally synthesized CDs to guide the synthesis of CDs with high QYs. Their design framework is illustrated in Figure 2A. They selected five key parameters (ethylene diamine (EDA) volume, precursor mass, reaction temperature, heating rate, and reaction time) as input features, with QY as the output target, thereby constructing the dataset. As shown in the correlation heatmap in Figure 2B, the selected features exhibit low correlation and high independence, confirming their validity. They subsequently evaluated the XGBoost regressor, multilayer perceptron (MLP), and SVM using three performance metrics: R2, MSE, and r. Based on the evaluation results, they selected the XGBoost model. Furthermore, using feature importance analysis, they determined that the key parameters in order of importance were the EDA volume, precursor mass, and reaction temperature. Finally, they screened the optimal synthesis conditions, conducted experimental verification, and obtained CDs with a high QY of 39.3%.
ML can also guide the synthesis of CDs with tunable third-order nonlinear optical susceptibility χ ( 3 ) and switching behavior. Nonlinear optics has important applications in laser technology, optical communications, and biomedical imaging [60,61,62]. Wang et al. [63] were the first to apply ML to guide the tunable third-order nonlinear optical susceptibility χ ( 3 ) and switching behavior. First, they select input features based on the synthesis conditions of CDs. Then, they construct experimental datasets under different experimental conditions and perform ML training. Furthermore, they employed four regression models (RF, XGBoost, SVM, and GBDT) and evaluated their performance, concluding that the GBDT model demonstrated optimal performance. Subsequently, they utilized the GBDT model for training and achieved significant progress in the data. Furthermore, they derived feature importance, revealing that reaction time accounted for 87% and reaction temperature for 8.9%. This indicates that the descriptor importance determining χ ( 3 ) is, in order of significance, reaction time, temperature, and so forth. Therefore, they optimized and tuned the features, successfully synthesizing CDs with χ ( 3 ) tunable from 0 to 1.79 × 10−8 esu, exhibiting excellent nonlinear optical properties and switching behavior.

3.1.2. Multi-Objective Optimization of ML-Based CDs

In ML-guided CD synthesis applications, most approaches optimize a single property by identifying the relationship between parameters and performance. In practical applications, it is often necessary to obtain CDs with multiple outstanding properties. Guo et al. [64] proposed an ML-guided multi-objective optimization strategy to indirectly guide CD synthesis by optimizing multiple properties (photoluminescence (PL) wavelength and PLQY). This strategy comprises four components: database construction, multi-objective optimization formulation, multi-objective optimization recommendation, and experimental validation. They employed an XGBoost ML model to construct a PL wavelength and PLQY prediction model based on eight synthetic parameters (temperature, time, catalyst, etc.). Notably, they adopted an iterative experimental optimization strategy, achieving full-color CDs with PLQY exceeding 60% across all colors after 20 iterations (40 experiments) (Figure 2C,D). The predictive accuracy of the ML model also improved with the number of iterations. As the number of iterations increased, the mean square error (MSE) for PLQY decreased from 0.45 to 0.1, whereas the MSE for PL wavelength remained consistently below 0.1, fully demonstrating the reliability of this approach (Figure 2E).
Figure 2. (A) Design framework for synthesizing CDs with high QY based on ML and hydrothermal experiments. (B) A heatmap of the Pearson correlation coefficients between selected characteristics of hydrothermally grown CDs. T: Temperature; t: time; Rr: Ramp rate; M: Mass; V: Volume; Y: Yield. Reprinted with permission from Ref. [39]. Copyright 2020, American Chemical Society. (C) Unified objective utility of MOO and design iterations. (D) Exploration of color under new synthetic experimental conditions. (E) MSE between predicted and actual target properties. Reprinted with permission from Ref. [64]. Copyright 2024, Nature Communications.
Figure 2. (A) Design framework for synthesizing CDs with high QY based on ML and hydrothermal experiments. (B) A heatmap of the Pearson correlation coefficients between selected characteristics of hydrothermally grown CDs. T: Temperature; t: time; Rr: Ramp rate; M: Mass; V: Volume; Y: Yield. Reprinted with permission from Ref. [39]. Copyright 2020, American Chemical Society. (C) Unified objective utility of MOO and design iterations. (D) Exploration of color under new synthetic experimental conditions. (E) MSE between predicted and actual target properties. Reprinted with permission from Ref. [64]. Copyright 2024, Nature Communications.
Molecules 31 01696 g002

3.1.3. ML-Guided Synthesis of Specific CDs

Red CDs exhibit high resistance to photo-bleaching and are biodegradable; however, their synthesis efficiency is low. To enhance the synthesis efficiency of red CDs, Luo et al. [58] first reported the use of ML to guide their synthesis. First, they collected synthetic data, then preprocessed it, which included labeling the data and splitting the dataset. Notably, this process employs multi-step feature engineering (XGBoost, one-hot encoding, and PCA). The XGBoost model was employed to extract data by combining input features across multiple dimensions. Categorical features were converted into numerical values using one-hot encoding, transforming each category into a binary vector. This process generated redundant information, prompting the use of PCA for dimensionality reduction by maximizing sample dispersion. Logistic regression was then employed to predict the synthesis conditions for red CDs. Ultimately, the constructed ML model achieved an AUC of 0.94 and an F 1 score of 0.94 in 10-fold cross-validation, identifying 93% of red CDs and 95% of non-red CDs, thereby enhancing the efficiency of synthesizing red CDs.

3.2. Application of ML in CD Sensor Detection

ML not only optimizes CD synthesis but also assists in distinguishing ions from organic compounds, thereby enhancing detection accuracy and speed. CDs function as fluorescent sensors by utilizing their outstanding fluorescent properties, such as high quantum yield and excellent photostability, to achieve highly sensitive and selective detection of target substances (ions and molecules) through changes in fluorescence signals (intensity and wavelength). Compared to traditional fluorescent probes, CD-based sensors offer advantages such as low toxicity, excellent biocompatibility, and low preparation costs. In the following section, we systematically elaborate on the role of ML in CD sensors.

3.2.1. Applications in Ion Detection

With the advancement of scientific inquiry, significant progress has been achieved in the application of CDs for metal ion detection. In recent years, extensive research efforts have been directed toward the development of CD-based fluorescent sensors for ion detection, as exemplified in Table 1.
By combining ML with fluorescence visualization, Zhang et al. [53] achieved the sequential quantitative detection of Al3+ and F in aqueous solutions. They employed synthetic sulfur-functionalized CDs as fluorescent probes and observed enhanced fluorescence intensity upon adding Al3+ at varying concentrations to the CDs solution, achieving a detection limit of 4.2 nmol/L. Fluorescence quenching occurred upon the addition of F. This phenomenon arises when quenchers interact with fluorescent molecules through processes such as charge transfer, electron transfer, or coordination bonding, leading to signal attenuation. The resulting detection limit was 47.6 nmol/L, demonstrating an “off-on-off” detection mode. They employed K-MCA, evaluating clustering tendency via Hopkins statistics: Hopkins statistics for Al3+ was 0.96, and for F it was 0.95. They then used PCA to reduce the dimensionality of three-dimensional variables to one-dimensional variables by comparing the magnitude of variance. The heatmap shows that there is a positive correlation with Al3+ and a negative correlation with F. Therefore, a linear regression model can be constructed based on the linear relationship between them. By predicting the concentrations in the test dataset, the final R2 values obtained were 0.927 for Al3+ and 0.987 for F, indicating that this model can be used for predicting and detecting pollutant concentrations.
ML can also assist in the detection of toxic heavy metal ions. Mandal et al. [50] developed an array sensor based on CNPs and integrated it with the predictive capabilities of artificial intelligence. This study achieved the first detection and differentiation of five toxic heavy metal ions (As (III), Cd (II), Hg (II), Cr (VI), and Pb (II)) listed by the World Health Organization and the U.S. Occupational Safety and Health Administration. The researchers analyzed sensor response images using multiclass classification algorithms, representing each pixel with red, green, and blue values. The extracted RGB dataset was then employed for pattern recognition. Testing revealed that deep-learning-based supervised algorithms enhanced the MLP to achieve optimal performance. The sensor array successfully distinguished heavy metal ions in labeled river water from those in sewage, demonstrating the concept of utilizing optical sensor data for accurate analytical predictions. First, the researchers synthesized nine types of CNPs with different surface functionalizations at a concentration of 50 ng μL−1 and characterized their structure, morphology, and photophysical properties. Subsequently, they constructed a fluorescent array sensor exhibiting distinct visual responses to toxic heavy metal ions, where interactions between different CNPs and heavy metal ions induced fluorescence changes. During this process, they considered the RGB values at 20 positions within the image and recorded the fluorescence responses of the nine distinct CNPs to each heavy metal ion. RGB values were extracted from digitally captured fluorescence images of each CNP–heavy-metal-ion combination to create a dataset, which was divided into training and testing sets for algorithmic learning and decision-making. Seven multiclass classification algorithms were then applied to analyze the sensor array data. The final results show that the enhanced MLP outperformed other methods, with the model converging after training. Concurrently, the average values for sensitivity, specificity, positive predictive value, and negative predictive value across all category labels were 92.45%, 79.67%, 93.75%, and 95.83%, respectively, demonstrating high accuracy even on a small dataset. Ultimately, the enhanced MLP was identified as optimal, validating the feasibility of the sensor array in real water samples. Notably, the pioneering use of generative adversarial networks (GANs) to augment the dataset enhanced the performance of deep learning algorithms in the automated heavy metal detection platform based on CNPs. The synthesized alizarin red S-based carbon nanoparticle fluorescent array sensor, combined with AI predictive analytics, effectively distinguished toxic heavy metal ions. The enhanced MLP algorithm demonstrated outstanding performance, whereas the sensor array exhibited robustness and practicality in real-world water sample detection. This proves that optical sensor data can be utilized for accurate analyte prediction without direct manual intervention.
Zhang et al. [51] developed an integrated ML approach to meet the demand for real-time, highly sensitive on-site detection of Cr (VI) in groundwater and drinking water. First, they synthesized N-doped blue light carbon dots (N-BCDs) with a QY of approximately 90%. Cr (VI) can be detected within 1 min by utilizing the internal filtering effect (IFE) quenching. Subsequently, they analyzed the fluorescence response of N-BCDs toward Cr (VI) at different concentrations using the Stern–Volmer plot of Cr (VI):
I I 0 = I + K S V [ M ] .
where I 0 and I are the luminescence emission intensities of the N-BCD analyte suspension, [ M ] is the molar concentration of the analyte (mM), and K S V is the Stern–Volmer constant. Calculations indicate a detection limit of 0.1574 μgL−1 for Cr (VI) concentrations ranging from 0 to 60 μgL−1. Furthermore, the fluorescence intensity at 425 nm progressively decreases with increasing Cr (VI) concentration. Through iterative adjustments of the cluster centers, compactness is achieved within each cluster. Subsequently, RGB and K-means feature extraction were applied in combination with ridge, XGBoost, SVR, and linear models to determine the concentrations. The final results demonstrated that the RGB and K-means feature extraction approach, integrated with ML models, achieved a fitting accuracy of 95.2%.
The combination of ML technology with nanoparticle fluorescence sensors enables not only the detection of individual chromium species but also the simultaneous quantification and identification of multiple chromium species. Khozani et al. [65] overcame the limitations of traditional chromium speciation methods by proposing a single-pore fluorescent sensor. This sensor utilizes orange-emissive thioglycolic acid-stabilized CdTe quantum dots (TGA-QDs) and CDs to detect chromium species. ML analysis of fluorescence spectral changes enables the quantitative detection and identification of multiple chromium species. The fluorescence spectrum of the QDs-CDs nanoprobe varies with concentration. Concurrently, adding different chromium species induces distinct spectral shifts, thereby enabling species differentiation based on quenching intensity. Partial least squares (PLS) regression modeling demonstrates strong linearity for Cr2+, Cr3+, and CrO42− within a specific concentration range. Thus, the QDs-CDs nanoprobe enables the simultaneous measurement of different chromium species. Subsequently, the authors employed an LDA model to distinguish individual chromium species from binary mixtures. The two-dimensional LDA score plot clearly delineated the differences between individual chromium species at varying concentrations. In practical water samples, both ML models were utilized for qualitative and quantitative analysis, thereby extending the application of the sensor to complex environmental samples.
Table 1. Examples of ions recently detected by CD sensors.
Table 1. Examples of ions recently detected by CD sensors.
Detected IonsDetection
Limit
Detection RangeSensor
Materials
References
Cr6+
Fe2+
Fe3+
Mn2+
Cu2+
Co2+
Ni2+
0.05 μM-QR-CDs
EDTA-Tb3+
[66]
Cr6+
Fe2+
Fe3+
Hg2+
-1–50 μMCPC-CDs[67]
Cd2+
Pb2+
Hg2+
0.15 μM
0.20 μM
0.09 μM
-AuNCs@NCDs[68]
Pb2+
Fe3+
12 nM
16 nM
1–100 μMVV-CDs[69]
Hg2+0.06 nM-CDs[23]
Hg2+6.2 nM-CQDs[70]
Fe3+0.135 μM0.3–3.3 μMN-CDs[71]
Fe3+0.039 μM0–150 μMCDs[39]
Fe3+0.91 μM1–200 μMCD@Eu-MOF[72]
Cu2+200 nM1–100 μMNS-CDs[73]
As3+16.8 nM0–200 nMCDs-MnO2[74]
Cr4+21.14 nM0.03–50 μMS, N-CDs[75]

3.2.2. Applications in Antibiotic Detection

Antibiotic residues pose serious threats to human health and the natural environment; however, their detection is hampered by high costs and cumbersome procedures. ML algorithms offer the advantages of being responsive, low-cost, simple, and fast, making them suitable for use in antibiotic detection. Relevant applications are shown in Table 2. Xu et al. [49] developed a dual-channel fluorescent sensor array based on two CDs (QR-CDs and CPC-CDs) and utilized ML to detect and distinguish four tetracyclines: Tetracycline (TC), oxytetracycline (OTC), doxycycline (DOX), and minocycline (MTC). Through the IFE, they observed that TCs induced distinct patterns of fluorescence intensity changes in CDs, thereby enabling the detection and differentiation of TCs. IFE refers to the phenomenon wherein the fluorescence intensity diminishes when fluorophores are present at high concentrations or coexist with other light-absorbing substances, owing to the absorption of excitation or emission light by these substances. Subsequently, they employed LDA and SVM to process multidimensional data. At different concentrations, each type of TC did not overlap with others. Ultimately, 100% discrimination of each TC type was achieved.
In array-based antibiotic detection, most arrays only perform detection and differentiation under specific concentrations and sample conditions, making it difficult to detect unknown samples and lacking a unified model. To address these challenges, Xu et al. [76] proposed a novel dual-emission fluorescence/colorimetric sensor array based on IFE, static quenching, and electrostatic interactions. This array utilizes highly fluorescent quantum yield CDs and CdTe quantum dots. Based on the addition of nine different antibiotics (ofloxacin (OFLX), amikacin sulfate (AMK), pefloxacin mesylate (PF), norfloxacin (NFC), kanamycin sulfate (KNM), DOX, metacycline (MTC), TC, and streptomycin (SM)), the fluorescence intensity (FI) and maximum emission wavelength (MEW) undergo changes. These distinct variations enable differentiation among the nine antibiotics. Within the tree-based pipeline optimization technology (TPTO) framework, stepwise prediction strategies were combined with ML to establish classification and concentration models, forming a unified sx model. TPTO is an automated ML tool and a tree-based model pipeline for predicting classification problems. It typically encompasses data cleaning, feature selection, feature processing, feature engineering, model selection, and hyperparameter optimization. By intelligently exploring thousands of potential pipelines, it identifies the optimal pipeline based on the data. Subsequently, the sx model was applied to identify nine antibiotics in deionized water, achieving 95% accuracy. The R2 value between the actual and predicted categories reached 0.9888, whereas the R2 between the actual and predicted concentrations for unknown samples was 0.9988. Both demonstrated linear relationships and excellent prediction performance. Furthermore, the sx model successfully identified binary and ternary mixtures with 100% accuracy.
To further assist antibiotic detection, Mandal et al. [77] proposed a deep learning recognition method based on nanoparticle array fluorescence imaging, which enabled multicategory antibiotic visualization without spectroscopic instruments. The detection principle relies on distinct fluorescence patterns emitted by all six antibiotics under a 350 nm laser wavelength. The researchers synthesized CNPs and optimized a nine-channel array (including eight metal ion channels) to capture antibiotic-induced fluorescence responses via a digital camera. Specifically, they selected pixel values for the cyan (C), magenta (M), yellow (Y), and key (K) channels; extracted CMYK values as feature data; and compared the performance of seven supervised algorithms: k-nearest neighbors (KNN), Gaussian Naive Bayes (GNB), global product classifier (GPC), RF, and artificial bee colony (ABC). The MLP algorithm achieved an accuracy of 0.82 and an F-measure of 0.81, outperforming the other algorithms. Subsequently, to enhance the MLP algorithm’s performance, they introduced GAN for data augmentation. This ultimately yielded the optimal algorithm, Aug-MLP (accuracy: 0.84, F-measure: 0.83), which demonstrated 100% accuracy in identifying antibiotics added to feed. The cause and mechanism of the fluorescence difference were identified using this array sensor.
Compared to traditional single-emission or dual-emission fluorescent probes, triple-emission fluorescent probes offer advantages, such as diverse color changes and reduced susceptibility to interference. Lu et al. [78] developed a portable device integrating a triple-emission fluorescent probe with ML-assisted smartphone technology, enabling quantitative and visual detection of TCs. The probe consists of blue-emitting carbon dots (BCDs) and red-emitting bovine serum albumin-protected copper nanoclusters (BSA-Cu NCs). The detection principle is based on the addition of different TCs, which induce shifts in the fluorescence peaks at 420 nm, 635 nm, and 520 nm, along with corresponding color changes, thereby enabling the detection of TCs. To achieve high-sensitivity, rapid, and convenient testing, they developed intelligent fluorescence analysis using YOLOv3 deep object detection. To extract and analyze color information, they compared six algorithms and selected different RGB channels for various TCs. Doxycycline hydrochloride (DC), Chlortetracycline Hydrochloride (CTC), and TC are suitable for the G channel, whereas OTC is suitable for the C/B channel. Finally, the least squares method was selected for fitting. To further demonstrate its application in real-world samples, they selected milk for testing TCs, achieving spiked recovery rates of 96.47–109.49%, highlighting the practicality of the sensor. Tan et al. [79] successfully synthesized a triple-emission ratio-based fluorescent probe and constructed a multimodal logic gate integrating fluorescence, colorimetric, and Ultraviolet detection channels. Specifically, they encapsulated CDs and Au NCs within ZIF-8 to synthesize a triple-emission ratio-based fluorescent probe (CD-Au NCs@ZIF-8). By integrating deep learning with smartphone tools, they achieved real-time, rapid identification of individual TCs. The detection principle involves the incorporation of TCs, which causes CD-Au NCs@ZIF-8 to exhibit altered fluorescence intensities at 440 nm and 640 nm and simultaneously generate a new fluorescence peak at 550 nm. Simultaneously, color changes occur. However, this only indicates the presence of TCs without enabling individual TC identification. Therefore, they developed a WeChat mini-program (96 Speckles) utilizing YOLOv5 and YOLOv8 algorithms for data processing and logic gate output. Ultimately, by incorporating different TCs, distinct variations in the ultraviolet–visible absorption spectrum peak of CD-Au NCs@ZIF-8 were achieved, thereby enabling the individual identification of the four TCs.
Table 2. Applications of ML in Antibiotic Detection.
Table 2. Applications of ML in Antibiotic Detection.
Antibiotics TestedDetection MechanismProbeProbe CompositionMLReferences
TC
OTC
DOX
MTC
IFEDual-channel fluorescent sensorQR-CDs,
CPC-CDs
SVM,
LDA
[49]
OFLX
AMK
PF
NFC
KNM
DOX
MTC
TC
SM
IFE, Static quenching, Electrostatic interactionsDual-emission fluorescence/colorimetric sensorCDs, CdTe quantum dotsTPTO,
ERF
[76]
AMP
CPFX
KAN
SMZ
TET
TMP
-Nine-channel arrayCNPsAug-MLP[77]
DC
CTC
OTC
TC
IFE, Sensitization mechanismsTriple-emission fluorescent probesBCDs, BSA-Cu NCsYOLOv3, Least squares method[78]
TTC
OTC
DC
CTC
Sensitization mechanismsTriple-emission fluorescent probesCD-Au NCsYOLOv5, YOLOv8[79]
Table notes: TC, Tetracycline; OTC, oxytetracycline; DOX, doxycycline; MTC, minocycline; OFLX, ofloxacin; AMK, amikacin sulfate; PF, pefloxacin mesylate; NFC, norfloxacin; KNM, kanamycin sulfate; MTC, metacycline; SM, streptomycin; ERF, Extreme Random Forest; AMP, Ampicillin; CPFX, Ciprofloxacin; KAN, Kanamycin; SMZ, Sulfamethoxazole; TET, Tetracycline; TMP, Trimethoprim; DC, Doxycycline hydrochloride; CTC, Chlortetracycline Hydrochloride; TTC, Tetracycline.

3.2.3. Application in the Detection of Certain Organic Compounds

ML technology can distinguish eight proteins with 100% accuracy. Pandit et al. [40] reported a biomolecular sensor based on a CD array for detecting proteins in buffers and human serum (Figure 3A). The detection principle relies on measurable fingerprint patterns generated by differential interactions with analytes, primarily driven by changes in the edge-state fluorescence of CDs, as illustrated in Figure 3B. At 460 nm, the eight proteins exhibited markedly different fluorescence changes in a buffer solution, enabling their identification based on these differences. To further determine the detection limit for proteins, they selected two representative proteins—hemoglobin (Hb) and β-galactosidase (β-Gal)—and observed distinct reaction patterns at varying concentrations in both PBS and human serum, exhibiting dose-dependent responses. ML algorithms outperformed LDA, with the final four ML algorithms achieving 100% accuracy. Ultimately, the detection limits for Hb and β-Gal in PBS were determined to be 25 nm, whereas those in human serum were found to be 50 nm.
ML can also assist in detecting bacteria. Soares et al. [80] developed an electrospun corn zein/curcumin carbon dot-based nanostructure electronic tongue for detecting Staphylococcus aureus in milk. This device performs comprehensive taste fingerprinting of liquid samples by combining sensor arrays with pattern recognition algorithms. During the detection process, they employed interactive document mapping (IDMAP) to reduce the dimensionality of capacitive spectra obtained from three sensor units (cCDOT, ZEI, and cCDOT/ZEI). This approach successfully distinguished pathogen samples at different concentrations from interfering samples. Subsequently, supervised learning algorithms (such as decision trees) were employed to construct a multidimensional calibration space (MCS), enabling automated classification and calibration. The study screened 57 frequency features using a decision tree model, generating an MCS for nine sample categories (including Staphylococcus aureus at varying concentrations and interferents) with an 86.1% classification accuracy. This model not only identified key distinguishing frequencies but also enhanced interpretability through rule visualization (ExMatrix). Based on this, they utilized the electronic tongue to distinguish between healthy cows, infected cows, and samples at different dilution levels, achieving a diagnostic accuracy of 80.1%. ML technology significantly enhanced the analytical capabilities of the carbon-dot sensor in complex environments, demonstrating its powerful potential for decision support throughout the process, from data dimensionality reduction to model calibration.
Deng et al. [81] employed a graphene quantum dots (GQDS) fluorescent sensor array. By utilizing a single GQDS sensing element combined with three different solvents (DMF: Dimethylformamide, DMSO: Dimethyl sulfoxide, EG: Ethylene glycol), they established a detection system capable of verifying the authenticity and quality of baijiu. They selected 21 organic compounds (acids, alcohols, aldehydes, and esters) present in baijiu at a concentration of 0.1 mM and analyzed them using PCA, HCA, and LDA. The first two principal components in the PCA algorithm account for a total of 90.3% of the information; however, the 21 organic compounds overlapped in the principal component plot, preventing the identification of all compounds (Figure 3C). Figure 3D,E demonstrate that HCA and LDA effectively distinguished all 21 organic compounds with 100% accuracy. Similarly, these methods can perfectly differentiate quaternary mixtures of methanol, acetaldehyde, butanoic acid ester, and ethyl acetate at various molar ratios.
Figure 3. (A) Sensing events in CD array-based sensing. (B) Mechanism of CD fluorescence changes in the presence of proteins. Reprinted with permission from Ref. [40] Copyright 2019, American Chemical Society. Identification of 21 organic molecules: (C) PCA; (D) HCA; (E) LDA. Reprinted with permission from Ref. [81] Copyright 2023, Analytical methods.
Figure 3. (A) Sensing events in CD array-based sensing. (B) Mechanism of CD fluorescence changes in the presence of proteins. Reprinted with permission from Ref. [40] Copyright 2019, American Chemical Society. Identification of 21 organic molecules: (C) PCA; (D) HCA; (E) LDA. Reprinted with permission from Ref. [81] Copyright 2023, Analytical methods.
Molecules 31 01696 g003

3.3. Application of ML in Predicting the Performance of CDs

The preparation of CDs faces issues such as long development cycles, high costs, and unpredictability. ML can address these issues by building models that extract patterns from synthesis parameters and characterization data, thereby enabling rapid and accurate prediction of unknown properties. However, compared to density functional theory (DFT), the predictive performance of ML largely depends on the quality and breadth of the training data and lacks a physical interpretation. DFT, on the other hand, has advantages in terms of mechanistic interpretation and validation, making it suitable for fields such as mechanistic interpretation of materials and design validation. Nevertheless, the two approaches are not mutually exclusive. For example, Schleder et al. [82] discussed the relationship between ML and DF, showing that combining them can optimize computational workflows, thereby broadening the scope of materials research and advancing the field.

3.3.1. Application in Predicting QY

In recent years, numerous studies have utilized ML to predict the yields of biochar and bio-oil from biomass, achieving significant progress. The high QY of CDs serves as a core metric for evaluating their luminescent performance. Chen et al. [41] pioneered the application of ML to predict the QY of CDs produced during biochar synthesis. First, they collected preparation parameters from 480 samples. First, based on the biochar production process, they identified 11 key parameters (cellulose, pyrolysis time, residence time, etc.). They collected preparation parameters for 480 samples under various combinations and split the dataset into a training set and a validation set in an 8:2 ratio. Subsequently, they evaluated six models, with the GBDT-R model demonstrating optimal QY prediction performance: R2 > 0.9, RMSE < 0.02 and MAPE < 3. Therefore, they employed the GBDT-R model for feature importance evaluation, identifying pyrolysis temperature, residence time, nitrogen content, and C/N ratio as the factors most significantly influencing QY. Furthermore, they simplified the parameters by selecting the top four features based on importance, with a relative error range of 0.00% to 4.60%. This model demonstrates good universality and provides a foundation for subsequent research on CD generation during biochar production.
Response surface methodology (RSM) is a statistical approach that integrates mathematical and statistical techniques based on polynomial equations and experimental data fitting. It predicts and optimizes multifactor systems by constructing response surface models [83]. Pudza et al. [84] employed a central composite design within RSM, selecting temperature, dosage, time, and solvent ratio as independent variables, with PLQY as the response value. When the optimal synthesis parameters were temperature 170 °C, dosage 0.1 g, time 100 min, and solvent volume 12 mL, the corresponding PLQY was 27.75%, with an RSM predicted value of 27.38% and R2 = 0.956. However, they acknowledged that while RSM effectively reduces experimental runs by examining interactions between factors, it has limitations in predicting nonlinear systems. ANN methods offer modeling capabilities for complex relationships. Therefore, they employed an ANN model from machine learning, pioneering the integration of the Levenberg–Marquardt backpropagation (LMBP) algorithm to construct a predictive model. This model incorporates temperature, dosage, time, and solvent ratio as inputs. After processing through hidden layers and optimizing the number of neurons, it outputs the PLQY, yielding a predicted value of 26.25% with R2 = 0.944. Figure 4A through Figure 4C demonstrate the reliability of this prediction method. The residual value between the RSM and ANN was ultimately found to be 0.0123, clearly indicating a high degree of consistency in the prediction results.

3.3.2. Application at the Predicted Wavelength

ML integrates multiple synthetic parameters and can also achieve a precise prediction of CD emission colors and wavelengths. Senanayake et al. [85] first selected seven input features (precursor molar ratio, reaction temperature, reaction time, etc.), compiled a dataset using carbon nanodot synthesis data from the literature, and divided it into training and testing sets. They then employed an ANN as the core algorithm in ML to construct three machines (M1–M3), where M1 served as a regression model and M2 and M3 as classification models. Without considering reaction temperature and time, they found that the training mean error from the M1 machine was 16.3 nm, with a maximum standard deviation of 17.1 nm. Through various tests, it was observed that the M1 machine could accurately predict blue and green CDs but failed to accurately predict red CDs. Notably, M2 and M3 differed by leveraging the strong correlation between the “color” feature and emission wavelength. They fed the predicted color from the classification models into the regression model and applied a scoring algorithm to derive the predicted emission wavelength. Models M2 and M3 reduced the training mean absolute error to 9.8 ± 11.1 nm and 9.6 ± 10.5 nm, respectively, demonstrating significant effectiveness. Finally, after model construction and parameter tuning, the predicted color accuracy reached 94%, with a minimum average wavelength error of 25.8 nm. This achieved a precise prediction of the emission color and wavelength, thereby reducing experimental uncertainty.
Additionally, Yan et al. [86] applied ML to the preparation of phosphorescent CDs, achieving highly accurate predictions of CD emission wavelengths and Stokes shifts. First, they constructed a dataset comprising 210 datasets, with input variables including reaction time, sulfuric acid volume, water, and ethanol, and output performance metrics encompassing fluorescence and phosphorescence emission wavelengths and Stokes shifts. The dataset was then split into a training set and a validation set in an 8:2 ratio. Subsequently, they trained the machine learning models and evaluated their performance using R2, MAE, and RMSE. The XGBoost model demonstrated the best performance through comparison. Subsequently, the predicted values were compared with the experimental values. The R2 values for fluorescence and phosphorescence emission wavelengths were 0.95 and 0.94, respectively, while the R2 values for Stokes shifts were 0.89 and 0.94, respectively. This confirmed the high accuracy and robustness of the XGBoost model. Finally, to further validate the model’s performance, the authors varied the parameters to synthesize nine novel CDs. The results show that the predicted values agreed well with the experimental data, establishing a solid foundation for phosphorescent CDs in fields such as information encryption and bioimaging.

3.4. Application of ML in Studying the Mechanism of CDs

ML can uncover hidden patterns and causal relationships within complex, high-dimensional experimental data, thereby accelerating the revelation of the “structure-synthesis-property” relationship in CDs—a feat that is difficult to achieve through traditional trial-and-error methods. The performance of CDs (such as fluorescence and catalytic activity) is influenced by nonlinear, synergistic interactions among multiple factors: precursors, synthesis methods, and reaction conditions (temperature, time, pH, etc.). Traditional data analysis methods are unable to extract deep insights from such data. ML algorithms (such as RF and ANN) excel at handling high-dimensional, nonlinear datasets. They can simultaneously consider dozens or even hundreds of variables to identify the most critical features, thereby establishing predictive models between complex “synthesis conditions” and “final performance.” Extracting key patterns from complex data deepens our understanding of the mechanisms and directly guides the rational design and screening of new materials.
Salahinejad et al. [87] investigated the mechanism of fluorescence quenching in cysteine-functionalized carbon quantum dots (Cys-CQDs). They synthesized Cys-CQDs and measured their fluorescence intensity quenched by 25 heavy metal ions. By extracting the physicochemical parameters of the metal ions, they established correlations between these parameters and quenching sensitivity using quantitative structure–property relationship (QSPR) models. Descriptor selection is fundamental to QSPR modeling. To extract more information, they employed stepwise regression, genetic algorithm (GA), and enhanced replacement method (ERM) techniques from machine learning. Simultaneously, they utilized multiple linear regression (MLR) and SVM to establish the QSPR model. The ratio of the luminescence peak area I 0 (without metal ions added to the Cys-CQD solution) to I Q (after addition) was used as the dependent variable in the QSPR model, and the SVM model outperformed the MLR model. ERM then extracts meaningful descriptions from the set of descriptors through iterative substitution and combinatorial optimization. Ultimately, ERM provides the covalent index ( C o v I n ), atomic environment number ( A E N ), and number of valence electrons ( N V E ) as the optimal subset of descriptors. Equation (5), derived from the MLR, indicates that the key parameters influencing heavy metal ion-induced fluorescence quenching of Cys-CQDs, in order of importance, are C o v I n , A E N , and N V E . This elucidates the structure–activity relationship governing CQD fluorescence quenching and provides theoretical guidance for designing highly selective metal fluorescent probes.
I 0 I Q = 2.643 ± 0.231 C o v I n + 0.275 ± 0.051 A E N 0.235 ± 0.027 N V E 5.115 ± 0.863
Furthermore, since the discovery of solid-state phosphorescence properties at room temperature in carbon-based quantum dot materials, such as carbon dots and graphene quantum dots, elucidating their phosphorescence mechanisms has remained a significant challenge in studying structure–property relationships. Li et al. [88] introduced ML to establish a quantitative structure–property relationship to explore the phosphorescence mechanism of GQDs. They introduced structural variability ( Ω ), porosity ( V ), and disorder ( S ). The equation is
S = 2 ln Ω + V .
Subsequently, they analyzed the structural disorder using ML and quantified the structural features. Simultaneously, they discovered that S is inversely correlated with the phosphorescence lifetime. Finally, they demonstrated that S can characterize the oscillator strength, thereby enabling the precise regulation of phosphorescence properties based on this finding.
CDs have emerged as alternative antimicrobial agents because of their unique physicochemical properties. Bian et al. [45] employed ML tools to investigate the relationship between the physicochemical characteristics of CDs and their antimicrobial capabilities, aiming to elucidate the underlying mechanisms. First, they constructed a dataset comprising 121 CD samples, which were processed to retain 17 features. The minimum inhibitory concentration (MIC) was used to determine antimicrobial activity. Next, feature importance was calculated. As shown in Figure 4D, the size of CDs, zeta potential, and bacterial species exerted the greatest influence on antimicrobial performance. This finding helps explain the relationship between physicochemical properties and antimicrobial efficacy. Additionally, they employed four ML classification algorithms—KNN, RF, XGBoost, and SVM—with MIC as the output metric, and evaluated the models based on the area under the ROC curve and accuracy (Figure 4E–H). The XGBoost model achieved an accuracy of 78.3% and an AUC of 0.90, indicating that it is the optimal model. Finally, they applied this model to predict ε-poly-L-lysine CDs (PL-CDs), thereby accelerating the design of highly effective antimicrobial CDs.

4. Conclusions and Perspective

ML has evolved from an emerging computational paradigm into a disruptive force driving progress across multiple disciplines. As a powerful data analysis technique, ML has deeply penetrated CD research, demonstrating immense potential. This study reviews recent ML applications in CDs, focusing on ML-optimized CD synthesis, sensor detection, predictive performance, and mechanistic studies. Generally, supervised learning is widely used in carbon dot research, but unsupervised learning also has unique value in specific contexts [14]. By establishing mapping models between synthesis parameters (such as precursor type, reaction temperature, time, and solvent) and target properties (QY and emission wavelength), ML effectively predicts optimal synthesis conditions, thereby accelerating the preparation of high-performance CDs. Concurrently, ML algorithms (such as support vector machines and convolutional neural networks) have been employed to process and analyze complex data generated by CD sensors (such as fluorescence spectra and images), thereby enabling high-precision, rapid identification and classification of multiple analytes and significantly enhancing the intelligence of detection. The reliability of ML depends on rigorous data and model optimization and therefore has inherent limitations. The current application of ML in CD research still faces several key challenges that limit its further development and broader application.
First, current research is generally constrained by small sample datasets (typically only a few hundred datasets), which severely limits the generalization capability and reliability of models. Future efforts should focus on expanding datasets at scale through high-throughput experiments and automated data collection. The field of CDs lacks a unified, open-source standard database. There is an urgent need to establish a standardized database containing precise synthesis conditions, detailed microstructural characterization (such as surface functional groups and particle size distribution), and multifunctional performance data. Promoting data format interoperability is also essential to facilitate data sharing. Differences in the size, shape, and anchoring methods of functional groups on CDs can easily lead to biases in model interpretation, feature redundancy that reduces accuracy, and confusion regarding the mechanisms of action of these functional groups. In practical applications, ML models cannot be applied directly; instead, feature engineering should be optimized by taking into account the surface structure of the CDs and the characteristics of their functional groups. Furthermore, current ML models are largely “black boxes,” lacking explanations for the intrinsic physicochemical mechanisms underlying the structure–property relationships of CDs. Future efforts should closely integrate theoretical calculations, incorporate structural parameters as model inputs to enhance interpretability, and develop multi-objective optimization algorithms to simultaneously regulate multiple performance metrics, such as emission wavelength, quantum yield, and lifetime. Finally, in practical applications, models should be further optimized to improve detection stability and reliability. Moreover, this approach can be extended to other applications, building a universal visualization sensing platform. For instance, in ion detection, the detection range can be expanded to cover more ions. Furthermore, applying this method to real-world environmental samples will enable the continuous optimization of detection techniques for increasingly complex environments. In the context of rapid AI advancements, we believe that machine learning—as a core branch—will infuse nanomaterials science with new intelligent vitality, achieving breakthrough applications in the field of carbon nanotubes.

Author Contributions

Conceptualization: D.S. (Dandan Sang) and F.J.; methodology: F.J. and H.W.; investigation: F.J., D.S. (Deyu Shen) and Q.W.; writing—original draft: F.J. and H.W.; data curation: D.S. (Deyu Shen), F.J. and Q.W.; writing—review and editing: D.S. (Dandan Sang), S.K. and Q.W.; supervision: D.S. (Dandan Sang) and Q.W.; visualization: Z.Z. and H.L.; validation: D.S. (Dandan Sang), Z.Z. and H.L.; resources and funding acquisition: D.S. (Dandan Sang) and Q.W. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China (grant number 62104090), the Natural Science Foundation of Shandong Province (grant numbers ZR2025MS68 and ZR2017QA013), and the Science and Technology Innovation “Double Ten Project” of Tai’an City (grant number 2025JSGG05).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study.

Conflicts of Interest

Authors Zhanfeng Zhang, Hang Li and Qinglin Wang were employed by the company Shandong Jinhengli Mechanical Manufacturing Co., Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Yu, J.; Yong, X.; Tang, Z.; Yang, B.; Lu, S. Theoretical Understanding of Structure-Property Relationships in Luminescence of Carbon Dots. J. Phys. Chem. Lett. 2021, 12, 7671–7687. [Google Scholar] [CrossRef] [PubMed]
  2. Xu, D.; Lin, Q.; Chang, H.T. Recent Advances and Sensing Applications of Carbon Dots. Small Methods 2019, 4, 1900387. [Google Scholar] [CrossRef]
  3. Liu, J.; Li, R.; Yang, B. Carbon Dots: A New Type of Carbon-Based Nanomaterial with Wide Applications. ACS Cent. Sci. 2020, 6, 2179–2195. [Google Scholar] [CrossRef]
  4. Ai, L.; Yang, Y.; Wang, B.; Chang, J.; Tang, Z.; Yang, B.; Lu, S. Insights into photoluminescence mechanisms of carbon dots: Advances and perspectives. Sci. Bull. 2021, 66, 839–856. [Google Scholar] [CrossRef]
  5. Macairan, J.R.; de Medeiros, T.V.; Gazzetto, M.; Yarur Villanueva, F.; Cannizzo, A.; Naccache, R. Elucidating the mechanism of dual-fluorescence in carbon dots. J. Colloid. Interface Sci. 2022, 606, 67–76. [Google Scholar] [CrossRef]
  6. Wang, B.; Song, H.; Qu, X.; Chang, J.; Yang, B.; Lu, S. Carbon dots as a new class of nanomedicines: Opportunities and challenges. Coord. Chem. Rev. 2021, 442, 214010. [Google Scholar] [CrossRef]
  7. Su, W.; Wu, H.; Xu, H.; Zhang, Y.; Li, Y.; Li, X.; Fan, L. Carbon dots: A booming material for biomedical applications. Mater. Chem. Front. 2020, 4, 821–836. [Google Scholar] [CrossRef]
  8. Nocito, G.; Calabrese, G.; Forte, S.; Petralia, S.; Puglisi, C.; Campolo, M.; Esposito, E.; Conoci, S. Carbon Dots as Promising Tools for Cancer Diagnosis and Therapy. Cancers 2021, 13, 1991. [Google Scholar] [CrossRef]
  9. Yu, J.; Song, H.; Li, X.; Tang, L.; Tang, Z.; Yang, B.; Lu, S. Computational Studies on Carbon Dots Electrocatalysis: A Review. Adv. Funct. Mater. 2021, 31, 2107196. [Google Scholar] [CrossRef]
  10. Long, C.; Qing, T.; Fu, Q.; Jiang, Z.; Xu, J.; Zhang, P.; Feng, B. Low-temperature rapid synthesis of high-stable carbon dots and its application in biochemical sensing. Dye. Pigment. 2020, 175, 108184. [Google Scholar] [CrossRef]
  11. Li, T.; Dong, Y.; Su, Y.; Li, Y.; Wang, J.; Hu, J.; Li, J. Facile preparation of low temperature carbon dots with long-wavelength emission and their sensing applications for crystal violet. Spectrochim. Acta A Mol. Biomol. Spectrosc. 2024, 310, 123863. [Google Scholar] [CrossRef]
  12. Wareing, T.C.; Gentile, P.; Phan, A.N. Biomass-Based Carbon Dots: Current Development and Future Perspectives. ACS Nano 2021, 15, 15471–15501. [Google Scholar] [CrossRef]
  13. Chahal, S.; Macairan, J.R.; Yousefi, N.; Tufenkji, N.; Naccache, R. Green synthesis of carbon dots and their applications. RSC Adv. 2021, 11, 25354–25363. [Google Scholar] [CrossRef] [PubMed]
  14. Duman, A.N.; Jalilov, A.S. Machine learning for carbon dot synthesis and applications. Mater. Adv. 2024, 5, 7097–7112. [Google Scholar] [CrossRef]
  15. Xu, Q.; Tang, Y.; Zhu, P.; Zhang, W.; Zhang, Y.; Solis, O.S.; Hu, T.S.; Wang, J. Machine learning guided microwave-assisted quantum dot synthesis and an indication of residual H2O2 in human teeth. Nanoscale 2022, 14, 13771–13778. [Google Scholar] [CrossRef]
  16. Liu, Y.Y.; Yu, N.Y.; Fang, W.D.; Tan, Q.G.; Ji, R.; Yang, L.Y.; Wei, S.; Zhang, X.W.; Miao, A.J. Photodegradation of carbon dots cause cytotoxicity. Nat. Commun. 2021, 12, 812. [Google Scholar] [CrossRef] [PubMed]
  17. Ethordevic, L.; Arcudi, F.; Cacioppo, M.; Prato, M. A multifunctional chemical toolbox to engineer carbon dots for biomedical and energy applications. Nat. Nanotechnol. 2022, 17, 112–130. [Google Scholar] [CrossRef]
  18. Behi, M.; Gholami, L.; Naficy, S.; Palomba, S.; Dehghani, F. Carbon dots: A novel platform for biomedical applications. Nanoscale Adv. 2022, 4, 353–376. [Google Scholar] [CrossRef]
  19. May, M. Eight ways machine learning is assisting medicine. Nat. Med. 2021, 27, 2–3. [Google Scholar] [CrossRef]
  20. Moosavi, S.M.; Jablonka, K.M.; Smit, B. The Role of Machine Learning in the Understanding and Design of Materials. J. Am. Chem. Soc. 2020, 142, 20273–20287. [Google Scholar] [CrossRef]
  21. Chen, J.; Luo, J.B.; Hu, M.Y.; Zhou, J.; Huang, C.Z.; Liu, H. Controlled Synthesis of Multicolor Carbon Dots Assisted by Machine Learning. Adv. Funct. Mater. 2022, 33, 2210095. [Google Scholar] [CrossRef]
  22. Tang, B.; Lu, Y.; Zhou, J.; Chouhan, T.; Wang, H.; Golani, P.; Xu, M.; Xu, Q.; Guan, C.; Liu, Z. Machine learning-guided synthesis of advanced inorganic materials. Mater. Today 2020, 41, 72–80. [Google Scholar] [CrossRef]
  23. Korah, B.K.; John, N.; John, B.K.; Mathew, S.; Bijimol, D.; Mathew, B. Carbon dots as a fluorescent ink and dual-mode probe for the efficient detection of doxycycline and Hg(II) ions. J. Mater. Res. 2022, 37, 3060–3070. [Google Scholar] [CrossRef]
  24. Brunton, S.L.; Noack, B.R.; Koumoutsakos, P. Machine Learning for Fluid Mechanics. Annu. Rev. Fluid. Mech. 2020, 52, 477–508. [Google Scholar] [CrossRef]
  25. Kochkov, D.; Smith, J.A.; Alieva, A.; Wang, Q.; Brenner, M.P.; Hoyer, S. Machine learning-accelerated computational fluid dynamics. Proc. Natl. Acad. Sci. USA 2021, 118, e2101784118. [Google Scholar] [CrossRef]
  26. Rai, R.; Tiwari, M.K.; Ivanov, D.; Dolgui, A. Machine learning in manufacturing and industry 4.0 applications. Int. J. Prod. Res. 2021, 59, 4773–4778. [Google Scholar] [CrossRef]
  27. Meuwly, M. Machine Learning for Chemical Reactions. Chem. Rev. 2021, 121, 10218–10239. [Google Scholar] [CrossRef]
  28. Zhang, H.; Fu, H.; Li, W.; Jiang, L.; Yong, W.; Sun, J.; Chen, L.Q.; Xie, J. Empowering the Sustainable Development of High-End Alloys via Interpretive Machine Learning. Adv. Mater. 2024, 36, e2404478. [Google Scholar] [CrossRef] [PubMed]
  29. Wen, C.; Zhang, Y.; Wang, C.; Huang, H.; Wu, Y.; Lookman, T.; Su, Y. Machine-Learning-Assisted Compositional Design of Refractory High-Entropy Alloys with Optimal Strength and Ductility. Engineering 2025, 46, 214–223. [Google Scholar] [CrossRef]
  30. Jia, H.; Tai, Z.; Lyu, R.; Ishikawa, K.; Sun, Y.; Cao, J.; Ju, D. Low-Temperature Sealing Material Database and Optimization Prediction Based on AI and Machine Learning. Polymers 2025, 17, 1233. [Google Scholar] [CrossRef]
  31. Zhao, Y.; Yang, B.; Zhang, K.; Guo, A.; Yu, Y.; Chen, L. Machine Learning Models for Predicting Freeze-Thaw Damage of Concrete Under Subzero Temperature Curing Conditions. Materials 2025, 18, 2856. [Google Scholar] [CrossRef] [PubMed]
  32. Xu, J.; Zheng, C.; Xie, Q.; Chen, Y.; Huang, Z.; Ye, D.; Kong, X.; Weng, H. Rapid and nondestructive identification of low-temperature stress severity in Juncao seedlings: Application of chlorophyll a fluorescence combined with visible-near infrared spectroscopy and machine learning. Plant Physiol. Biochem. 2026, 230, 110896. [Google Scholar] [CrossRef] [PubMed]
  33. Li, Y.; Yang, G.; Zhang, W.; He, D.; Yan, Y.; Jiang, J. Understanding the low-temperature drying process of sludge with machine learning in a sewage-source heat pump drying system. J. Environ. Manage 2025, 375, 124284. [Google Scholar] [CrossRef] [PubMed]
  34. Lee, H.; Gray, S.; Zhao, Y.; Castelluccio, G.M. Machine Learning Applied to Identify Corrosive Environmental Conditions. Front. Mater. 2022, 9, 830260. [Google Scholar] [CrossRef]
  35. Deng, J.; Stixrude, L. Thermal Conductivity of Silicate Liquid Determined by Machine Learning Potentials. Geophys. Res. Lett. 2021, 48, e2021GL093806. [Google Scholar] [CrossRef]
  36. Zhang, Z.; Csányi, G.; Alfè, D.; Zhang, Y.; Li, J.; Liu, J. Free Energies of Fe-O-Si Ternary Liquids at High Temperatures and Pressures: Implications for the Evolution of the Earth’s Core Composition. Geophys. Res. Lett. 2022, 49, e2021GL096749. [Google Scholar] [CrossRef]
  37. Zhang, Y.; Xu, X. Predicting the superconducting transition temperature of high-Temperature layered superconductors via machine learning. Phys. C 2022, 595, 1354031. [Google Scholar] [CrossRef]
  38. Murillo, M.S. Data-driven electrical conductivities of dense plasmas. Front. Phys. 2022, 10, 867990. [Google Scholar] [CrossRef]
  39. Han, Y.; Tang, B.; Wang, L.; Bao, H.; Lu, Y.; Guan, C.; Zhang, L.; Le, M.; Liu, Z.; Wu, M. Machine-Learning-Driven Synthesis of Carbon Dots with Enhanced Quantum Yields. ACS Nano 2020, 14, 14761–14768. [Google Scholar] [CrossRef]
  40. Pandit, S.; Banerjee, T.; Srivastava, I.; Nie, S.; Pan, D. Machine Learning-Assisted Array-Based Biomolecular Sensing Using Surface-Functionalized Carbon Dots. ACS Sens. 2019, 4, 2730–2737. [Google Scholar] [CrossRef]
  41. Chen, J.; Zhang, M.; Xu, Z.; Ma, R.; Shi, Q. Machine-learning analysis to predict the fluorescence quantum yield of carbon quantum dots in biochar. Sci. Total Environ. 2023, 896, 165136. [Google Scholar] [CrossRef]
  42. Liu, Y.; Mohammed, Z.M.A.; Ma, J.; Xia, R.; Fan, D.; Tang, J.; Yuan, Q. Machine Learning Driven Fluidity and Rheological Properties Prediction of Fresh Cement-Based Materials. Materials 2024, 17, 5400. [Google Scholar] [CrossRef]
  43. Wang, Z.; Wang, L.; Zhang, H.; Xu, H.; He, X. Materials descriptors of machine learning to boost development of lithium-ion batteries. Nano Converg. 2024, 11, 8. [Google Scholar] [CrossRef]
  44. Peng, J.; Muhammad, R.; Wang, S.L.; Zhong, H.Z. How Machine Learning Accelerates the Development of Quantum Dots? Chin. J. Chem. 2020, 39, 181–188. [Google Scholar] [CrossRef]
  45. Bian, Z.; Bao, T.; Sun, X.; Wang, N.; Mu, Q.; Jiang, T.; Yu, Z.; Ding, J.; Wang, T.; Zhou, Q. Machine Learning Tools to Assist the Synthesis of Antibacterial Carbon Dots. Int. J. Nanomed. 2024, 19, 5213–5226. [Google Scholar] [CrossRef] [PubMed]
  46. Talukder, M.A.; Khalid, M.; Sultana, N. A hybrid machine learning model for intrusion detection in wireless sensor networks leveraging data balancing and dimensionality reduction. Sci. Rep. 2025, 15, 4617. [Google Scholar] [CrossRef] [PubMed]
  47. Kumar, J.A.A.N.A. Machine Learning from Theory to Algorithms: An Overview. J. Phys. Conf. Ser. 2018, 1142, 012012. [Google Scholar] [CrossRef]
  48. Syah, R.; Ahmadian, N.; Elveny, M.; Alizadeh, S.M.; Hosseini, M.; Khan, A. Implementation of artificial intelligence and support vector machine learning to estimate the drilling fluid density in high-pressure high-temperature wells. Energy Rep. 2021, 7, 4106–4113. [Google Scholar] [CrossRef]
  49. Xu, Z.; Wang, Z.; Liu, M.; Yan, B.; Ren, X.; Gao, Z. Machine learning assisted dual-channel carbon quantum dots-based fluorescence sensor array for detection of tetracyclines. Spectrochim. Acta A Mol. Biomol. Spectrosc. 2020, 232, 118147. [Google Scholar] [CrossRef]
  50. Mandal, S.; Paul, D.; Saha, S.; Das, P. Deep learning assisted detection of toxic heavy metal ions based on visual fluorescence responses from a carbon nanoparticle array. Environ. Sci. Nano 2022, 9, 2596–2606. [Google Scholar] [CrossRef]
  51. Zhang, M.; He, H.; Huang, Y.; Huang, R.; Wu, Z.; Liu, X.; Deng, H. Machine learning integrated high quantum yield blue light carbon dots for real-time and on-site detection of Cr(VI) in groundwater and drinking water. Sci. Total Environ. 2023, 904, 166822. [Google Scholar] [CrossRef]
  52. Pravin Renold, A.; Arockia Abins, A.; Katiravan, J. IoT-Driven classroom air quality management with deep hierarchical cluster analysis. Sci. Rep. 2025, 15, 35595. [Google Scholar] [CrossRef]
  53. Zhang, Q.; Li, X.; Yu, L.; Wang, L.; Wen, Z.; Su, P.; Sun, Z.; Wang, S. Machine learning-assisted fluorescence visualization for sequential quantitative detection of aluminum and fluoride ions. J. Environ. Sci. 2025, 149, 68–78. [Google Scholar] [CrossRef] [PubMed]
  54. Liu, T.; Chen, S.; Ruan, K.; Zhang, S.; He, K.; Li, J.; Chen, M.; Yin, J.; Sun, M.; Wang, X.; et al. A handheld multifunctional smartphone platform integrated with 3D printing portable device: On-site evaluation for glutathione and azodicarbonamide with machine learning. J. Hazard. Mater. 2022, 426, 128091. [Google Scholar] [CrossRef] [PubMed]
  55. Liu, X.-Y.; Zhou, L.; Chen, X.-H.; Li, J.-Y. Lithofacies identification using support vector machine based on local deep multi-kernel learning. Pet. Sci. 2020, 17, 954–966. [Google Scholar] [CrossRef]
  56. Badillo, S.; Banfai, B.; Birzele, F.; Davydov, I.I.; Hutchinson, L.; Kam-Thong, T.; Siebourg-Polster, J.; Steiert, B.; Zhang, J.D. An Introduction to Machine Learning. Clin. Pharmacol. Ther. 2020, 107, 871–885. [Google Scholar] [CrossRef]
  57. Pang, B.; Nijkamp, E.; Wu, Y.N. Deep Learning with TensorFlow: A Review. J. Educ. Behav. Stat. 2019, 45, 227–248. [Google Scholar] [CrossRef]
  58. Luo, J.B.; Chen, J.; Liu, H.; Huang, C.Z.; Zhou, J. High-efficiency synthesis of red carbon dots using machine learning. Chem. Commun. 2022, 58, 9014–9017. [Google Scholar] [CrossRef]
  59. Muyassiroh, D.A.M.; Permatasari, F.A.; Iskandar, F. Machine learning-driven advanced development of carbon-based luminescent nanomaterials. J. Mater. Chem. C 2022, 10, 17431–17450. [Google Scholar] [CrossRef]
  60. Liu, Y.; Liu, B.; Song, Y.; Hu, M. Sub-30 fs Yb-fiber laser source based on a hybrid cascaded nonlinear compression approach. Chin. Opt. Lett. 2022, 20, 100006. [Google Scholar] [CrossRef]
  61. Hou, J.; Situ, G. Image encryption using spatial nonlinear optics. eLight 2022, 2, 3. [Google Scholar] [CrossRef]
  62. Wang, G.; Boppart, S.A.; Tu, H. Compact simultaneous label-free autofluorescence multi-harmonic microscopy for user-friendly photodamage-monitored imaging. J. Biomed. Opt. 2024, 29, 036501. [Google Scholar] [CrossRef]
  63. Wang, X.; Wang, H.; Zhou, W.; Zhang, T.; Huang, H.; Song, Y.; Li, Y.; Liu, Y.; Kang, Z. Carbon dots with tunable third-order nonlinear coefficient instructed by machine learning. J. Photoch Photobio A Chem. 2022, 426, 113729. [Google Scholar] [CrossRef]
  64. Guo, H.; Lu, Y.; Lei, Z.; Bao, H.; Zhang, M.; Wang, Z.; Guan, C.; Tang, B.; Liu, Z.; Wang, L. Machine learning-guided realization of full-color high-quantum-yield carbon quantum dots. Nat. Commun. 2024, 15, 4843. [Google Scholar] [CrossRef]
  65. Khozani, R.M.; Abbasi-Moayed, S.; Hormozi-Nezhad, M.R. Machine learning-assisted chromium speciation using a single-well ratiometric fluorescent nanoprobe. Chemosphere 2024, 357, 141966. [Google Scholar] [CrossRef]
  66. Xu, Z.; Chen, J.; Liu, Y.; Wang, X.; Shi, Q. Multi-emission fluorescent sensor array based on carbon dots and lanthanide for detection of heavy metal ions under stepwise prediction strategy. Chem. Eng. J. 2022, 441, 135690. [Google Scholar] [CrossRef]
  67. Liu, Y.; Chen, J.; Xu, Z.; Liu, H.; Yuan, T.; Wang, X.; Wei, J.; Shi, Q. Detection of multiple metal ions in water with a fluorescence sensor based on carbon quantum dots assisted by stepwise prediction and machine learning. Environ. Chem. Lett. 2022, 20, 3415–3420. [Google Scholar] [CrossRef]
  68. Wang, S.; Deng, G.; Yang, J.; Chen, H.; Long, W.; She, Y.; Fu, H. Carbon dot- and gold nanocluster-based three-channel fluorescence array sensor: Visual detection of multiple metal ions in complex samples. Sens. Actuat B Chem. 2022, 369, 132194. [Google Scholar] [CrossRef]
  69. Zulfajri, M.; Liu, K.-C.; Pu, Y.-H.; Rasool, A.; Dayalan, S.; Huang, G.G. Utilization of Carbon Dots Derived from Volvariella volvacea Mushroom for a Highly Sensitive Detection of Fe3+ and Pb2+ Ions in Aqueous Solutions. Chemosensors 2020, 8, 47. [Google Scholar] [CrossRef]
  70. Dhandapani, E.; Maadeswaran, P.; Mohan Raj, R.; Raj, V.; Kandiah, K.; Duraisamy, N. A potential forecast of carbon quantum dots (CQDs) as an ultrasensitive and selective fluorescence probe for Hg (II) ions sensing. Mat. Sci. Eng. B 2023, 287, 116098. [Google Scholar] [CrossRef]
  71. Kalanidhi, K.; Nagaraaj, P. Facile and Green synthesis of fluorescent N-doped carbon dots from betel leaves for sensitive detection of Picric acid and Iron ion. J. Photochem. Photobiol. A 2021, 418, 113369. [Google Scholar] [CrossRef]
  72. Zheng, Y.; Wang, X.; Guan, Z.; Deng, J.; Liu, X.; Li, H.; Zhao, P. Application of CD and Eu3+ Dual Emission MOF Colorimetric Fluorescent Probe Based on Neural Network in Fe3+ Detection. Part. Part. Syst. Char 2022, 39, 2200124. [Google Scholar] [CrossRef]
  73. Bisauriya, R.; Antonaroli, S.; Ardini, M.; Angelucci, F.; Ricci, A.; Pizzoferrato, R. Tuning the Sensing Properties of N and S Co-Doped Carbon Dots for Colorimetric Detection of Copper and Cobalt in Water. Sensors 2022, 22, 2487. [Google Scholar] [CrossRef]
  74. He, X.; Li, Y.; Yang, C.; Lu, L.; Nie, Y.; Tian, X. Carbon dots-MnO2 nanocomposites for As(III) detection in groundwater with high sensitivity and selectivity. Anal. Methods 2020, 12, 5572–5580. [Google Scholar] [CrossRef]
  75. Ji, Y.; Zou, X.; Wang, W.; Wang, T.; Zhang, S.; Gong, Z. Co-Doped S, N-Carbon dots and its fluorescent film sensors for rapid detection of Cr (VI) and Ascorbic acid. Microchem. J. 2021, 167, 106284. [Google Scholar] [CrossRef]
  76. Xu, Z.; Wang, K.; Zhang, M.; Wang, T.; Du, X.; Gao, Z.; Hu, S.; Ren, X.; Feng, H. Machine learning assisted dual-emission fluorescence/colorimetric sensor array detection of multiple antibiotics under stepwise prediction strategy. Sens. Actuat B Chem. 2022, 359, 131590. [Google Scholar] [CrossRef]
  77. Mandal, S.; Paul, D.; Saha, S.; Das, P. Multi-layer perceptron for detection of different class antibiotics from visual fluorescence response of a carbon nanoparticle-based multichannel array sensor. Sens. Actuators B Chem. 2022, 360, 131660. [Google Scholar] [CrossRef]
  78. Lu, Z.; Chen, S.; Chen, M.; Ma, H.; Wang, T.; Liu, T.; Yin, J.; Sun, M.; Wu, C.; Su, G.; et al. Trichromatic ratiometric fluorescent sensor based on machine learning and smartphone for visual and portable monitoring of tetracycline antibiotics. Chem. Eng. J. 2023, 454, 140492. [Google Scholar] [CrossRef]
  79. Tan, P.; Chen, Y.; Chang, H.; Liu, T.; Wang, J.; Lu, Z.; Sun, M.; Su, G.; Wang, Y.; Wang, H.D.; et al. Deep learning assisted logic gates for real-time identification of natural tetracycline antibiotics. Food Chem. 2024, 454, 139705. [Google Scholar] [CrossRef]
  80. Soares, A.C.; Soares, J.C.; Dos Santos, D.M.; Migliorini, F.L.; Popolin-Neto, M.; Dos Santos Cinelli Pinto, D.; Carvalho, W.A.; Brandao, H.M.; Paulovich, F.V.; Correa, D.S.; et al. Nanoarchitectonic E-Tongue of Electrospun Zein/Curcumin Carbon Dots for Detecting Staphylococcus aureusin Milk. ACS Omega 2023, 8, 13721–13732. [Google Scholar] [CrossRef]
  81. Deng, J.; Ma, Y.; Liu, X.; Xu, J.; Luo, H.; Luo, X.; Huo, D.; Hou, C. Identification of Chinese baijiu from the same brand based on a graphene quantum dots fluorescence sensing array. Anal. Methods 2023, 15, 5891–5900. [Google Scholar] [CrossRef]
  82. Schleder, G.R.; Padilha, A.C.M.; Acosta, C.M.; Costa, M.; Fazzio, A. From DFT to machine learning: Recent approaches to materials science–a review. J. Phys. Mater. 2019, 2, 032001. [Google Scholar] [CrossRef]
  83. Bezerra, M.A.; Santelli, R.E.; Oliveira, E.P.; Villar, L.S.; Escaleira, L.A. Response surface methodology (RSM) as a tool for optimization in analytical chemistry. Talanta 2008, 76, 965–977. [Google Scholar] [CrossRef]
  84. Yahaya Pudza, M.; Zainal Abidin, Z.; Abdul Rashid, S.; Md Yasin, F.; Noor, A.S.M.; Issa, M.A. Sustainable Synthesis Processes for Carbon Dots through Response Surface Methodology and Artificial Neural Network. Processes 2019, 7, 704. [Google Scholar] [CrossRef]
  85. Senanayake, R.D.; Yao, X.; Froehlich, C.E.; Cahill, M.S.; Sheldon, T.R.; McIntire, M.; Haynes, C.L.; Hernandez, R. Machine Learning-Assisted Carbon Dot Synthesis: Prediction of Emission Color and Wavelength. J. Chem. Inf. Model. 2022, 62, 5918–5928. [Google Scholar] [CrossRef]
  86. Yan, Z.; Li, J.; Zhou, S.; Yang, X. Machine learning-assisted controllable preparation of multicolor phosphorescent carbon dots. Carbon. 2025, 241, 120394. [Google Scholar] [CrossRef]
  87. Salahinejad, M.; Sadjadi, S.; Abdouss, M. Investigating fluorescence quenching of cysteine-functionalized carbon quantum dots by heavy metal ions: Experimental and QSPR studies. J. Mol. Liq. 2021, 334, 116067. [Google Scholar] [CrossRef]
  88. Li, Y.; Chen, L.; Yang, S.; Wei, G.; Ren, X.; Xu, A.; Wang, H.; He, P.; Dong, H.; Wang, G.; et al. Symmetry-Triggered Tunable Phosphorescence Lifetime of Graphene Quantum Dots in a Solid State. Adv. Mater. 2024, 36, e2313639. [Google Scholar] [CrossRef]
Figure 1. This review systematically outlines the application of ML in CDs research: ML optimization of CDs synthesis, ML-assisted detection using CDs sensors, ML for performance prediction, and ML in studying mechanisms.
Figure 1. This review systematically outlines the application of ML in CDs research: ML optimization of CDs synthesis, ML-assisted detection using CDs sensors, ML for performance prediction, and ML in studying mechanisms.
Molecules 31 01696 g001
Figure 4. (A) RSM predictions of experimental values. (B) ANN predictions of experimental values. (C) Comparison of RSM and ANN predictions. Reprinted with permission from Ref. [84]. Copyright 2019, Multidisciplinary Digital Publishing Institute. (D) Gini importance of features in the raw data set and features after extraction. The ROC curves and AUCs for KNN (E), SVM (F), XGBoost (G), and RF (H), respectively. Reprinted with permission from Ref. [45]. Copyright 2024, International Journal of Nanomedicine.
Figure 4. (A) RSM predictions of experimental values. (B) ANN predictions of experimental values. (C) Comparison of RSM and ANN predictions. Reprinted with permission from Ref. [84]. Copyright 2019, Multidisciplinary Digital Publishing Institute. (D) Gini importance of features in the raw data set and features after extraction. The ROC curves and AUCs for KNN (E), SVM (F), XGBoost (G), and RF (H), respectively. Reprinted with permission from Ref. [45]. Copyright 2024, International Journal of Nanomedicine.
Molecules 31 01696 g004
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Jia, F.; Wang, H.; Shen, D.; Sang, D.; Zhang, Z.; Li, H.; Kumar, S.; Wang, Q. Machine-Learning-Assisted Carbon Dots: From Algorithms to Applications and Beyond. Molecules 2026, 31, 1696. https://doi.org/10.3390/molecules31101696

AMA Style

Jia F, Wang H, Shen D, Sang D, Zhang Z, Li H, Kumar S, Wang Q. Machine-Learning-Assisted Carbon Dots: From Algorithms to Applications and Beyond. Molecules. 2026; 31(10):1696. https://doi.org/10.3390/molecules31101696

Chicago/Turabian Style

Jia, Fengjiao, Hengkai Wang, Deyu Shen, Dandan Sang, Zhanfeng Zhang, Hang Li, Santosh Kumar, and Qinglin Wang. 2026. "Machine-Learning-Assisted Carbon Dots: From Algorithms to Applications and Beyond" Molecules 31, no. 10: 1696. https://doi.org/10.3390/molecules31101696

APA Style

Jia, F., Wang, H., Shen, D., Sang, D., Zhang, Z., Li, H., Kumar, S., & Wang, Q. (2026). Machine-Learning-Assisted Carbon Dots: From Algorithms to Applications and Beyond. Molecules, 31(10), 1696. https://doi.org/10.3390/molecules31101696

Article Metrics

Back to TopTop