1. Introduction
Ovarian cancer is the fifth most common cause of death in women and the leading cause of gynecological-related mortality in the world [
1,
2,
3]. The incidence of ovarian cancer is growing at a rate of almost 250 thousand new cases every year and may further increase in the future as populations age [
2,
3,
4,
5]. Unfortunately, the diagnosis of most cases occurs when the cancer is already too advanced for surgery alone to be curative (stages 3 and 4) [
2,
3,
4]. At these stages, the cancer has already invaded widely and metastasized, thus accounting for the poor prognosis statistics associated with this disease [
1,
2,
3]. For these reasons, the detection of ovarian cancer should be at its early stage (stages 1 or 2) where surgery alone or in combination with other modalities can be curative [
3,
6]. Therefore, the focus of health care systems has been directed towards screening programs of at-risk female populations [
3,
7]. As this type of cancer shows no symptoms during initial stages, the early detection must rely on the least invasive and the most reliable screening tests to prevent potential damage to women’s reproductive systems due to further invasive examination [
2,
4,
5].
Currently, ovarian cancer screening programs are based on annual medical check-ups consisting of pelvic examination combined with a transvaginal ultrasound (TVUS) and serum CA-125 measurement by immunoassay [
3,
5]. TVUS imaging of the uterus, fallopian tubes, and ovaries for unusual masses detects benign and malignant masses with poor discrimination [
2,
5]. The CA-125 blood test has been used clinically for more than 30 years and is based on quantification of the CA-125 mucin-glycoprotein epitope in the blood, using a threshold cut-off value as a biomarker for ovarian cancer and other malignancies [
3,
5]. However, this biomarker fails to detect 50% of ovarian cancer at its early stage (FIGO/AJCC stage 1 and 2), 25% at stage 3, and 10% at stage 4, motivating researchers to search for new biomarkers [
3,
5]. Furthermore, the current costs of these screening programs are a huge burden for health care systems, leading many in the USA to stop supporting these programs due to the low accuracy for early stage detection [
2,
7]. Recently, some tests based on combining multiple novel blood biomarkers and circulating tumor-derived RNA signatures have shown higher sensitivities and have been proposed as potential future solutions [
2,
8,
9]. However, these approaches are also associated with a substantial increase in the costs per patient, which will make it difficult for many countries to implement screening programs for their general population. Based on these practical considerations, it is urgent to find novel and affordable solutions when conducting large screening programs for early detection of ovarian cancer.
Matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-ToF MS) is a very sensitive, affordable, and accurate technique for mass determination of biomolecules [
10,
11,
12,
13]. It enables precise detection of ionized peptides and proteins represented as peaks of mass-to-charge ratio (
m/
z). This technology has been considered to have high potential for clinical diagnostics and is already used in clinical microbiology [
10,
13]. One of the main reasons for this is the reduced cost per test in comparison with genetic-based testing and most currently used microbiological techniques for clinical identification of infection [
11,
12,
14]. Such 10-fold reductions in cost can enable screening programs to be rolled out, even in poorer countries. In the last years, MALDI-ToF MS has further evolved to become an ultra-fast high-throughput technology suitable for diagnostic screening purposes, with the capacity of each machine to generate approximately seven mass spectra per-patient samples per minute [
15]. Thus, this technology would be ideal for mass screening at-risk ovarian cancer populations within a sensible and reasonable processing time [
11,
12,
14]. Applying MALDI-ToF technology to ovarian cancer detection has previously been attempted using blood serum samples with an optimal sensitivity of 71% and specificity of 68% [
16]. However, although promising, this is still far from an ideal solution for population screening purposes.
More recently, MALDI-ToF application for human disease diagnostics has been enhanced by cutting-edge automated bioinformatics pipelines for spectral analysis [
17,
18,
19,
20]. This approach relies on the detection of mass spectral pattern recognition rather than specific biomarkers, which is the equivalent of simultaneously detecting multiple relevant biomarkers. Using mathematical models for disease scoring and machine learning approaches, these mass spectral patterns have been successfully applied to the detection of aneuploidies in spent IVF-embryo/blastocyst media and fetal Down’s syndrome from the urine of pregnant women, with high sensitivity and reasonable specificities [
20,
21,
22]. Further, these pipelines have been developed enabling automated and ultra-fast mass spectral data processing within seconds, creating software tools that enable screening for diseases in general populations by clinical laboratories [
18,
19]. Thus, it is highly feasible that this MALDI-ToF-based technology can be deployed for early-stage detection of ovarian cancer.
This study set out to test the hypothesis that this technological approach can provide a more suitable solution for general population screening programs. To address this question, we conducted a retrospective cohort study using blood serum samples from ovarian cancer patients in the development of MALDI-ToF MS-based diagnostic tools, for early- and late-stage cancer detection. Here, we report the diagnostic power and performance of these tools using distinct spectral profiles of MALDI-ToF MS generated from two experimental approaches, optimized using advanced bioinformatics and machine learning.
3. Results
Mass spectra of serum samples, from women with and without ovarian cancer, were generated using two protocols for MALDI-ToF MS for blood serum. For one, we used a CHCA matrix and for the other an SA matrix. Protocols were optimized for taking an average time of 80 minutes to generate a batch of 48 mass spectra (100 s per sample). We analyzed the mass spectral data generated from these protocols using our bioinformatics data pipeline. Most of the data generated through these experimental protocols were of high quality, 96% using CHCA matrix and 99% using SA matrix. The average mass spectra from patients with ovarian cancer at early and late stages were compared (
Figure 1). These spectra are characterized by having more pronounced peaks located between 200
m/
z and 700
m/
z using CHCA matrix (
Figure 1A) and between 2000
m/
z and 5500
m/
z using SA matrix (
Figure 1B). Multiple spectral changes between cancer stages and controls can also be observed across the entire mass spectrum, including mass regions of lower intensities (box I and II in
Figure 1). Proportionally, the averaged mass spectrum of early and late stages of ovarian cancer has substantial variation in comparison to women without cancer in particular mass regions across the entire mass spectrum (
Figure 2). Some regions show distinct variations among stages, which may potentially be used for classification purposes. However, there is a substantial degree of complexity of these variations across the spectrum and overlap between stages. In addition, one order of magnitude of intensity variation was found across the data set. This further complicates a direct application of mass spectral variations for classification purposes.
To simplify the mass spectral data for developing classification models, we focus on analyzing well-defined peaks, which were used as signatures for the pattern-based scoring function for the likelihood of having a particular stage of ovarian cancer. First, we optimized the extraction of well-defined peaks on the data by increasing the number of detected peaks while reducing the variability of total peaks detected (
Figure 3). We obtained an optimal increase in well-defined peak distribution ranging from 26 to 57 detected peaks in each sample for CHCA matrix mass spectral data (
Figure 3A). For SA matrix mass spectral data, the optimization reduced the number of peaks for a final distribution ranging from 12 to 113 detected peaks in each sample (
Figure 3B). Secondly, on this resulting data, we applied machine learning for training and validating predictive models based on mass spectral patterns, to optimize the binary classification of ovarian cancer detection on early stages (1 or 2) and late stages (3 and 4). We obtained six classification models, one for each type of matrix explored and specifically for detecting ovarian cancer at early stages (stage 1 or 2), stage 3, and stage 4. The performance evolution of these models is presented in
Figure 4. Our machine learning approach enables the evolution of predictive models with good to excellent classification power (98% < AUC > 70%) for ovarian cancer stage detection [
25,
26]. The maximum obtained sensitivities of these models were between 92% and 100% using CHCA and 100% for SA matrix. This indicates high potential for detecting ovarian cancer in early and late stages. However, models generated using SA matrix data show a much higher specificity compared with the ones obtained using CHCA matrix, having a much lower false positive rate. Among SA-generated models, we obtained final classification specificities of 92% for early stages, 96% for stage 3, and 82% for stage 4.
The selected spectral signatures learned during model evolution enabled us to obtain mass spectral patterns of the ovarian cancer stages that have a high classification potential. The patterns are composed of complex combinations of multiple characteristic peaks detected together, with multiple intensity variations (increase or decrease) in discrete mass spectral ranges (
Figure 5). Patterns differ among stages with only few signatures that have similar qualitative variation. Analysis of the individual contribution of each learned signature for the overall model specificity enables us to pinpoint the ones that impact specificity the most (
Figure 5). This information is useful for further tuning of the predictive models to balance the sensitivity and specificity of final deployed models. Thus, we further refined the models by neglecting the last learned pattern components, which increased the false positive rates, rendering a final model version with 97% specificity for SA-derived models and 86% specificity for CHCA-derived models (see
Figure 4).
Finally, we implemented the final refined version of the classification models into prototype pipelines, one for scoring cancer stages on mass spectral data generated using CHCA matrix and another for scoring SA matrix. Running these pipelines on the entire data set, we obtained an average data processing time to score all stages of 1.4 seconds per sample for CHCA data and 29.7 seconds for SA data. Using the selected cutoffs obtained during model evolution for refined versions, we have recapitulated the same model performance shown in
Figure 4 as expected (see
Supplementary Tables S1 and S2). The obtained score’s distribution was compared (
Figure 6) showing higher scoring differences among stages and the control for SA-matrix-derived models in comparison with CHCA-matrix-derived models. The results further show that models designed for detecting a particular stage often score higher than cutoffs on other stages in comparison with controls (see the example of SA-derived models in
Figure 7). This generates multiple hits for predicting ovarian cancer at different stages, indicating that our scoring models have low stage specificity. Yet, best scoring discrimination between a stage and the control is observed for the model that was optimized for that stage. Interestingly, the combination of algorithm predictions using a simple decision rule of at least one positive hit for cancer, regardless of stage, has improved the sensitivity of the individual algorithm to 94% for CHCA and 99% for SA (
Supplementary Tables S1 and S2, respectively). However, the rule base combination of algorithms also reduced the specificity to 92% for SA- and 71% for CHCA-derived models.
4. Discussion
Failure to detect ovarian cancer at an early stage has motivated many researchers to attempt improving current diagnostic tests or develop novel ones [
2,
3,
4]. In this study, we developed and optimized a MALDI-ToF mass spectral pattern-based bioinformatics pipeline, which enabled the detection of ovarian cancer in blood serum with high performance. In comparison with the reported sensitivities for the CA-125 test, used worldwide in screening programs, our approach provides a substantial gain in sensitivity and specificity for early- and late-stage disease detection [
3,
5]. These gains were almost of 50% in sensitivity for early-stage detection, suggesting that implementing our MALDI-ToF approach in screening programs would have a huge impact in preventing thousands of misdetected cases every year [
2,
3].
Other recently developed solutions to the problem are based on longitudinal bio-marker level monitoring, statistical models with additional bio-markers and microRNA quantification. In particular, measurement of plasma circBNC2 has enhanced the performance of CA-125 surveillance/monitoring boosting ovarian cancer detection [
5,
8,
9]. For early-stage disease, these approaches reported optimal sensitivities ranging from 82% to 92% with specificities between 86% and 95%. For later-stage ovarian cancers, longitudinal models and circBNC2-based tests report sensitivities of 90% to 100% with specificities of 84% to 95%. In comparison, our MALDI-ToF-based solution demonstrated here has higher sensitivity (95% to 99%) with equally higher specificities (97%), representing a significant gain. Our best solution outcompetes the performance of the recent approach of combining multiple longitudinal risk models which, although increasing sensitivity, report only equivalent specificity for early-stage ovarian cancer [
9].
This was possible due to the fine tuning of the MALDI-ToF MS data pattern recognition models towards improving the specificity, whilst keeping high sensitivity by removing particular signatures in the patterns that contributed to an increase in the number of false positives (
Figure 4 and
Figure 5). Thus, our MALDI-ToF-based approach provides an optimal solution and framework for maximizing the gain of true positives while minimizing the false positives. Balancing these is useful for designing optimal ovarian cancer screening programs to prevent unnecessary, invasive investigations in false positive cases or delayed diagnosis in the cases of false negatives [
2,
4,
5].
Interestingly, the models developed for a given stage have also scored positive for other stages and the combination of predicted outcomes using a decision rule further improved the cancer detection performance. This could be due to multiple intermediate cancer stages that share common mass spectral profiling or a continuous mass spectral profiling evolution that correlates with cancer development. Further research is still necessary for understanding how to correlate mass spectral profiling with cancer development to enable better stage classification or establishing a correlation with cancer aggressiveness.
The application of MALDI-ToF mass spectrometry for ovarian cancer detection on blood serum samples has previously been attempted, but the results have been poor and with experimental bias [
16,
27]. In a recent systemic review and meta-analysis of 18 studies evaluating the accuracy of MALDI-ToF MS for ovarian cancer, the reported overall sensitivity was 77% (95% CI: 73–80%) and specificity was 72% (95% CI: 70–74%) [
28]. In a more recent study, Swiatly et al. followed a traditional proteomic approach for multiple biomarker detection using high-performance liquid chromatography (HPLC) applying conventional statistical models [
6,
16,
29,
30]. This study reported only a 71% sensitivity with a 68% specificity. For classification purposes, we followed quite a different approach. Instead, we applied mass spectral pattern recognition and used novel machine learning algorithms to optimize predictive models [
17,
31,
32]. The substantial improvement in performance demonstrates a methodological breakthrough, which launches MALDI-ToF mass spectrometry as a highly sensitive technique for ovarian cancer diagnostics.
The reduced cost of running MALDI-ToF-based technology is a huge advantage for the implementation of our proposed solution for ovarian cancer screening programs. This advantage is mainly due to the reduced costs of reagents and the possibility of running multiple samples in one single run [
11,
14,
19]. In our approach we have not used HPLC for biomarker resolution prior to mass signature detection in a mass analyzer, as is traditionally employed in most mass spectral techniques in proteomics and metabolomics [
6,
29,
30]. Thus, this ‘dilute and shoot’ direct mass analysis enables further economization of operational costs, which are otherwise considerable from reagents and additional equipment. Indeed, the total estimated costs per test using our approach would be five to ten GBP, given a sample throughput of 50,000 a year over three years. It would approximately halve the cost when compared to the CA-125 screening test [
7,
33]. According to recent analysis, the major benefit to health care systems for each person’s case of early versus late stage detection of ovarian cancer is the saving of approximately 90 thousand GBP per year [
33]. Thus, the impact of implementing MALDI-ToF technology with our approach for screening would definitely improve current screening programs in terms of cost-effectiveness and long-term health economics. This is not the case for longitudinal follow-up monitoring models with additional biomarkers or circBNC2-based tests, since these techniques require the purchase of more expensive reagents [
11,
14].
However, there is still the need for governmental or private investment for implementing MALDI-ToF technology into clinical laboratories as each piece of equipment costs hundreds of thousands of GBP [
11,
14]. However, it can be argued that this investment is absolutely necessary, and the savings will pay off in the short-term, given that the hardware of this technology is used in other diagnostics such as microbiology and is being applied in the diagnosis of multiple other diseases [
14]. Future implementation of MALDI-ToF-based technology would enable government support for a broad number of screening programs, in particular in poorer countries [
4,
5,
14].
Another advantage of MALDI-ToF mass spectrometry for disease screening purposes is its capacity for generating multiple spectra in the scale of minutes [
17,
19]. We have further optimized and automatized the testing using bioinformatics, which rendered multiple results at a rate of a few seconds per sample. Excluding the time of sample preparation and calibration, each MALDI-ToF mass spectrometer has an estimated maximum capacity of generating approximately 620 results per day, with a potential increase if optimizing for speed of data acquisition. At this rate, only 102 MALDI-ToF machines working full-time would be required to screen the entire UK female population over 30 in one year [
34]. Thus, the approach could make it feasible to screen an entire relevant population group with an initial investment for laboratory setup or hiring specialized companies that already have MALDI-ToF technology.
In our work, we also optimized, tested, and compared the potential application of CHCA and SA matrices for ovarian cancer diagnostics. These have not been compared before for this purpose [
16]. Our results have pinpointed better diagnostic potential for SA-derived models in comparison to the CHCA. This was unexpected as CHCA matrix is known to generate more robust spectra in comparison to SA in urine, blood, serum, and culture media [
18,
20,
21,
22,
35]. Comparatively, mass spectral data acquired using SA matrix result in richer data; however, the acquisition simultaneously produces more background noise, often leading to the complex baseline that requires more sophisticated corrections [
18]. This has previously been solved using automated bioinformatics pipelines and is successfully applied now for ovarian cancer mass spectral analysis [
18].
There are several limitations to our study. First, our predictive models were not capable of providing an accurate classification between stages. This was evident during testing with multiple models on all data regardless of the diagnosed stage. Nevertheless, the low intra-stage specificity observed is irrelevant for screening purposes as other confirmatory methods such as TVUS offer support to further clinical diagnostics and therapy assessment [
2,
4,
5]. These can be applied as a second-stage evaluation following positive scoring by inexpensive serum MALDI-ToF MS screening [
2,
4]. Secondly, huge deviations in predictions can be obtained if experimental protocols, matrices, and mass spectrometer settings are not exactly followed, including sample preparation and storage. This has been observed before for other studies of data modeling using both CHCA and SA matrices [
18,
19,
20,
21,
36]. The reason for this is the high precision needed for our pattern recognition. Peak widening on a mass spectrum with experimental conditions and calibrations is detrimental [
18,
32,
37]. Third, the number of volunteers in different cancer stages and the control were limited for conducting more robust validations of predictive models [
26]. As the numbers for stages 1 and 2 were reduced, we had to aggregate these stages into one to have statistical meaning. In an ideal scenario, the numbers of volunteers within each group should be approximately a thousand, in order to split into large training and validation subsets [
26]. Consequently, the reported performance may slightly change if more data were to be tested.