Skip to Content
TechnologiesTechnologies
  • Review
  • Open Access

4 August 2026

36 Pages

Artificial Intelligence-Based Respiratory Sound Analysis: A Scoping Review of Digital Auscultation Technologies, Public Datasets, Signal Processing, and Deep Learning Methods

,
,
,
,
,
,
,
,
and
1
Metropolitan College (MET), Boston University, Boston, MA 02215, USA
2
Department of Electronics, Telecommunications and Space Technologies, Satbayev University, Almaty 050013, Kazakhstan
3
Department of Radio Engineering and Telecommunications, ALT University Named After Mukhametzhan Tynyshbayev, 97 Shevchenko Str., Almaty 050012, Kazakhstan
*
Authors to whom correspondence should be addressed.

Abstract

Respiratory sound analysis is becoming increasingly popular as a non-invasive method for detecting adventitious sounds and aiding in the diagnosis of respiratory diseases such as chronic obstructive pulmonary disease (COPD), asthma, and pneumonia. However, studies in this field differ significantly in datasets, recording devices, annotation techniques, preprocessing pipelines, model architectures, validation strategies, and evaluation metrics. This scoping review examines current research on artificial intelligence-based respiratory sound analysis, with a focus on datasets, acquisition and annotation practices, signal processing, feature representations, machine learning (ML) and deep learning (DL) methods, and evaluation protocols. The search identified 1056 database records, and 89 reports were included for the final evidence mapping. The reviewed datasets support event-, cycle-, recording-, and patient-level tasks and differ considerably in population, scale, acquisition hardware, annotation granularity, and label structure. The findings also show that acquisition and annotation are closely linked, creating potential device-, recording-site-, and label-related confounding. Methodologically, the literature can be summarized in four broad stages: handcrafted feature-based ML; deep spectrogram learning; representation learning and multimodality; and deployment-oriented, robustness-focused systems. Despite recent achievements, the field remains limited by small and imbalanced datasets, annotation uncertainty, device and population variability, inconsistent data splitting, and limited external validation. More reliable clinical use will require standardized acquisition and annotation, quality-controlled preprocessing, patient-independent evaluation, task-appropriate metrics, transparent reporting, and validation across independent devices, datasets, and clinical populations.

1. Introduction

Chronic respiratory diseases, including COPD and asthma, remain major causes of morbidity and mortality worldwide, while pneumonia and other lower respiratory infections also contribute substantially to the global respiratory disease burden [1,2]. Timely assessment and continued monitoring are therefore important for clinical management and the prevention of adverse outcomes. Auscultation is a widely used component of respiratory assessment, allowing clinicians to identify abnormal sounds such as wheezes, crackles, and diminished breath sounds [3]. However, the interpretation of lung sounds is observer-dependent and may vary across clinicians because of differences in experience and terminology [3]. These limitations have increased interest in Computerized Respiratory Sound Analysis (CRSA), which applies digital signal processing and automated methods to respiratory acoustics [4].
Existing reviews show that methodological practices remain inconsistent across automated respiratory sound studies. Garcia-Mendez et al. [5] reviewed 62 ML studies using public lung-sound databases and found that the ICBHI 2017 database was used in approximately two-thirds of the studies. However, most studies showed a high risk of bias or concerns related to patient selection, reference standards, and inconsistent methodological reporting [5]. Xia et al., likewise, reported substantial variation in respiratory audio databases, target conditions, ML pipelines, and experimental designs and identified unresolved challenges in remote respiratory screening [6]. Methodological differences are not limited to model design; they also extend to recording platforms and study populations. A systematic review of computerized respiratory sounds in pediatrics reported considerable variation in recording procedures, acoustic outcomes, and measurement approaches across studies [7]. Tabatabaei et al. showed that smartphone-based respiratory sound systems differ in their recording hardware, acoustic features, event-detection methods, classification algorithms, and intended applications, which complicates direct comparison and implementation [8]. A recent clinical review concluded that digital stethoscopes and AI may support more standardized and remote interpretation of lung sounds, but that clinical use still depends on high-quality datasets, effective noise management, and validation outside controlled research settings [3]. In parallel, multimodal approaches combining respiratory audio with visual or physiological information are emerging, although their clinical translation remains limited by heterogeneous data sources, alignment requirements, and insufficient external validation [9]. These reviews are useful for understanding specific parts of the field. Some mainly focus on public datasets, while others discuss pediatric respiratory sound analysis, smartphone-based recording, clinical auscultation, or multimodal approaches. However, they do not fully connect the task structure of the datasets with acquisition and annotation practices, preprocessing decisions, model selection, validation design, and deployment requirements.
This distinction is important because event-, cycle-, recording-, and patient-level analyses involve different prediction targets and cannot be evaluated using the same assumptions. Accordingly, this review examines the complete respiratory sound analysis pipeline, from data acquisition and annotation to model evaluation and deployment.
The main contributions of this review are as follows:
  • A task-based taxonomy of respiratory sound datasets that distinguishes event-, cycle-, recording-, and patient-level applications.
  • An analysis of the relationship between acquisition and annotation, including device variability, anatomical recording location, label ontology, and dataset-level confounding.
  • A comparative assessment of preprocessing methods, feature representations, and model families according to their suitability for different prediction units.
  • A four-stage conceptual synthesis of the field, covering handcrafted feature-based ML, deep spectrogram learning, representation learning, and multimodality, deployment-oriented and robustness-focused systems.
  • An evaluation of validation and deployment practices, including patient-independent data splitting, task-appropriate evaluation metrics, external validation, and quantitative reporting of lightweight models.
The remainder of the paper is organized as follows. Section 2 describes the review methodology, including the search, screening, data-charting, and synthesis procedures. Section 3 presents the results according to the five research questions. Section 4 discusses the main findings and methodological implications, and Section 5 presents the conclusions.

2. Methodology

2.1. Review Design

This study used a scoping review design to map the methodological landscape of AI-based respiratory sound analysis. This approach was appropriate because the literature covers diverse datasets, acquisition devices, annotation levels, preprocessing methods, feature representations, ML and DL models, prediction tasks, and evaluation protocols. Rather than identifying a single best-performing model, the review aimed to characterize the main research directions, compare methodological practices, and identify gaps in the available evidence.
The review was conducted using established scoping review methods and reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR) [10]. The completed PRISMA-ScR checklist is provided as Supplementary Material S1. The process included defining the research questions, identifying relevant records, applying eligibility criteria, screening and selecting reports, charting the included studies, and synthesizing the evidence using narrative and tabular methods. The review protocol was not prospectively registered. No formal critical appraisal or risk-of-bias assessment was performed, as the purpose of this scoping review was to map the literature, describe methodological patterns, and identify research gaps rather than to estimate pooled effects.

2.2. Research Questions

The review was structured around five research questions, each addressing a major component of the respiratory sound analysis pipeline. The questions and their purposes are presented in Table 1.
Table 1. Description and purpose of the RQs in this scoping review.

2.3. Search Strategy and Identification of Studies

The main database search focused on English-language journal articles, conference papers, review articles, and dataset or repository records published between 2021 and 2026. This time period was chosen to reflect current advances in AI-based respiratory sound analysis, specifically DL, transfer learning, transformer-based approaches, and deployment-oriented models. Earlier foundational studies published before 2021 were identified through backward reference checking and included only when they introduced key respiratory datasets.
The initial literature search was conducted on 10 May 2026, followed by a final search update and verification on 2 July 2026. Searches were conducted in IEEE Xplore, PubMed, ScienceDirect, Google Scholar, and the MDPI platform. The same core Boolean concepts were used across all sources, while the search interface, available fields, and filters varied by platform. The main search string was (“respiratory sound” OR “lung sound”) AND (“machine learning” OR “deep learning” OR “artificial intelligence” OR “AI”) AND (classification OR detection OR diagnosis). Google Scholar was used as a supplementary search source, and the first 100 relevance-ranked results were screened. Table 2 shows the search strategy and record identification procedure, including the databases searched, search strings, filters, and the number of records identified from each database.
Table 2. Search interfaces, applied filters, and numbers of records retrieved from each source.

2.4. Screening and Study Selection

The searches across the five sources identified 1056 records. Two additional foundational dataset reports published before 2021 were identified through backward reference checking, resulting in 1058 records before deduplication. Fifty-five duplicate records were removed in Zotero, followed by 46 additional duplicates identified in Rayyan, leaving 957 unique records. Three book sections and one preprint were then excluded because they did not meet the publication-type criteria. Consequently, 953 records proceeded to title-and-abstract screening. During title-and-abstract screening, supported by inspection of keywords and publication metadata, 363 records were excluded because they were outside the scope of the review. These comprised studies focused on cough, voice, speech, snoring, or other non-auscultation audio ( n = 45 ); imaging-only or other non-auscultation diagnostic methods ( n = 55 ); respiratory medicine without a lung-sound analysis component ( n = 70 ); general AI, ML, or audio-processing methods without a respiratory sound application ( n = 149 ); and records in which respiratory sounds were mentioned only tangentially and that did not report relevant acquisition, processing, dataset development, automated analysis, or model evaluation ( n = 44 ).
An additional 169 conference records were excluded during publication-overlap assessment when they described the same dataset, methods, experiments, and principal contribution as a more complete journal article identified in the search. In these cases, the journal version was retained for further assessment because it provided a more complete description of the methods, validation procedures, results, and study limitations. In total, 532 records were excluded during screening, leaving 421 reports for full-text retrieval and eligibility assessment. The study identification, screening, and selection process is summarized in Figure 1.
Figure 1. Flow diagram of study identification, screening, eligibility assessment, and inclusion for the scoping review.
Following title-and-abstract screening, 421 reports were retrieved and assessed against the eligibility criteria presented in Table 3. A total of 332 reports were excluded at this stage: 35 were short or preliminary conference papers with insufficient experimental detail, 123 used similar datasets without providing a distinct contribution, and 174 did not provide sufficient information to identify the preprocessing pipeline, model architecture, validation protocol, or train–test split. The remaining 89 reports were included in the evidence mapping.
Table 3. Eligibility criteria used for study selection.

2.5. Data Charting and Evidence Synthesis

Data from the included reports were charted using a structured evidence-mapping matrix. The extracted variables included publication type, dataset, population and sample size, acquisition device and recording location, annotation level, prediction task, preprocessing method, feature representation, model family, validation strategy, evaluation metrics, deployment characteristics, and the availability of public source code and trained model weights.
The charted information was organized according to the five research questions and summarized using narrative synthesis and comparative tables. Quantitative pooling was not performed because the included reports differed substantially in their datasets, prediction units, label structures, preprocessing pipelines, validation protocols, and evaluation metrics.

2.6. Open-Science Audit Methods

A total of 89 reports were included in the evidence mapping. Three additional references were cited outside the evidence map: Refs. [1,2] were used to support the epidemiological and clinical background, whereas Ref. [10] was cited as the methodological guideline for the scoping review. Most reports were identified through IEEE Xplore (49/89, 55.1%), followed by ScienceDirect (22/89, 24.7%), PubMed (8/89, 9.0%), and MDPI (7/89, 7.9%). Two reports (2.2%) were identified through backward reference checking, and one report (1.1%) through Google Scholar.
A study-level open-science audit was conducted for all 89 reports included in the evidence map. Review articles, dataset-only or challenge-description records, and studies that did not train or evaluate a computational model were classified as not applicable to the primary source code and trained-weights assessment. Dataset-focused reports were retained when they also trained and evaluated a computational classification model. This procedure resulted in 71 studies eligible for the primary audit. Among the 71 studies eligible for the code-and-weights audit, publicly accessible study-specific source code was identified for 9 studies (12.7%; Refs. [11,12,13,14,15,16,17,18,19]). Public study-trained model weights or checkpoints were identified for four studies (5.6%; Refs. [11,12,13,19]), all of which also provided source code. When the two resources were considered jointly, four studies (5.6%) provided both code and trained weights, five studies (7.0%) provided code without trained weights, and no study provided trained weights without source code. Neither publicly accessible study-specific code nor trained weights was identified for 62 studies (87.3%) at the time of verification (see Table 4).
Table 4. Availability of public study-specific source code and trained model weights among the eligible computational studies.

3. Results

3.1. Respiratory Sound Types

Respiratory sounds are acoustic signals produced by airflow through the respiratory system that can provide useful information regarding airway health and lung function. Respiratory sounds are generally categorized as normal and abnormal (or adventitious) sounds [20]. Normal lung sounds comprise tracheal, bronchial, bronchovesicular, and vesicular breath sounds, which are generated by physiological airflow through distinct parts of the airways and lungs [20]. Abnormal or adventitious lung sounds are frequently associated with disease alterations in the respiratory system. The most often analyzed abnormal sounds include crackles, wheezes, rhonchi, and stridor. Crackles are brief, discontinuous sounds, whereas wheezes are continuous musical sounds commonly associated with airway narrowing. Rhonchi are low-pitched continuous sounds, and stridor is a high-pitched sound associated with upper-airway obstruction [3,20]. In AI-based respiratory sound analysis, these sound categories are frequently used as target labels in classification tasks. Most studies focus on normal versus abnormal sound classification, crackle and wheeze detection, adventitious sound classification, or respiratory disease recognition using recorded lung sound patterns [20,21,22,23].

3.2. RQ1. What Respiratory Sound Datasets Are Used in AI-Based Respiratory Sound Analysis?

Dataset size, class balance, acquisition conditions, and annotation quality directly influence model development and evaluation. Unlike many traditional audio classification tasks, pulmonary sound analysis is extremely sensitive to clinical and technical parameters such as patient population, disease distribution, recording device, auscultation location, annotation level, and background noise [24,25,26,27,28]. As a result, the reported performance of ML and DL models cannot be interpreted without regard to the dataset on which they were trained and evaluated.
Public respiratory sound datasets play an important role in reproducible research and fair comparison of different methods. However, existing datasets differ substantially in terms of the number of subjects, number of recordings, label categories, annotation granularity, and acquisition protocols. Some datasets, for example, include respiratory cycle or event-level annotations for abnormal sound detection, while others focus on disease-level labels or severity evaluation [24,25,26,27,28]. Furthermore, class imbalance, small sample sizes, device variability, and unclear patient-independent evaluation protocols all have an impact on several datasets. These factors can affect the model results and, in some cases, lead to overly optimistic performance reporting [29]. The datasets are compared by prediction unit, annotation level, scale, and accessibility. Instead of distinguishing datasets as “clinical” or “benchmark”, the datasets are organized according to whether they support event-level detection, cycle-level classification, recording-level sound classification, patient-level diagnosis, or disease severity assessment.

3.2.1. Task-Based Dataset Taxonomy and Annotation Levels

Datasets were grouped according to their prediction unit and annotation level: event, respiratory cycle, recording, or patient. Event-level datasets provide temporal onset and offset annotations; cycle-level datasets assign labels to complete respiratory cycles; recording-level datasets label an entire audio file; and patient-level datasets associate one or more recordings with a diagnosis or severity category. The task structure, population, scale, and accessibility of the included respiratory sound datasets are summarized in Table 5.
Respiratory sound classification and respiratory sound event detection have distinct task structures. Classification algorithms typically operate on respiratory events or cycles with predefined temporal constraints. In pre-segmented classification, the interval of interest is supplied to the model in advance. Event detection additionally requires temporal localization of each event within a longer recording. The BioCAS 2024 challenge expanded respiratory sound analysis from pre-segmented samples to joint temporal localization and multiclass event classification [30].
Table 5. Task-based taxonomy, scale, and accessibility of respiratory sound datasets.

3.2.2. Event-Level and Cycle-Level Respiratory Sound Datasets

The ICBHI 2017 Respiratory Sound Database [24] is one of the most widely used respiratory sound datasets in the field of respiratory sound classification. This dataset was created for the ICBHI 2017 Scientific Challenge and is composed of data collected by two research groups in Portugal and Greece. The database contains 920 audio recordings from 126 participants, including both adults and children. The sounds were recorded from several auscultation sites in the trachea and chest using various digital stethoscopes and microphones. The recordings are stored in 16-bit audio format, with sampling rates of 4, 10, and 44.1 kHz and an average recording duration of 21.5 s. The database provides two annotation sets: 6,898 respiratory cycles labeled as normal, crackles, wheezes, or crackles plus wheezes, and 10,775 temporally localized crackle and wheeze events. For this reason, the ICBHI 2017 dataset is widely used as a primary reference dataset for comparing respiratory sound classification methods and evaluating the performance of models.
SPRSound was introduced by Zhang et al. as the first open-access pediatric respiratory sound database [25]. This database was jointly developed by Shanghai Jiao Tong University and its affiliated hospitals. The database contains 2683 respiratory sound recordings and 9089 annotated events gathered from 292 pediatric subjects. The annotation in the database is provided at two different levels: event-level and recording-level. At the event level, the labels include normal, rhonchi, wheeze, stridor, coarse crackle, fine crackle, and wheeze-and-crackle events, whereas the recording-level labels include normal, continuous adventitious sounds (CAS), discontinuous adventitious sounds (DAS), CAS and DAS, and poor quality. For the BioCAS 2024 challenge, the original SPRSound and Grand Challenge 2023 data were pooled for training and validation, while a separate Grand Challenge 2024 cohort of 1704 recordings and 7659 events from 324 participants was reserved as a blind test set. Unlike the 2023 classification benchmark, the 2024 challenge required temporal localization and classification of respiratory events in continuous recordings and included a separate audio-compression task [30,31].
HF_Lung_V1 [26] is an open-access lung sound database for breath-phase and adventitious-sound detection. It combines recordings from 261 patients in the TSECC database and 18 additional residents of a respiratory care ward or respiratory care center. TSECC recordings were acquired using a Littmann 3200 electronic stethoscope (3M, St. Paul, MN, USA), while the 18 RCW/RCC residents were recorded using both the Littmann 3200 and the HF-Type-1 multichannel system. The recorded lung sounds were stored in WAV format with a sampling rate of 4 kHz and a resolution of 16 bits. The duration of the recordings obtained with the Littmann 3200 device was initially 15.8 s. They were then shortened to 15 s, which was considered sufficient for clinical analysis. The dataset provides event-level temporal annotations, including the start and end points of inhalation, exhalation, wheeze, stridor, rhonchus, and crackles. For model development, wheeze, stridor, and rhonchus were grouped into the CAS category, whereas DAS labels related to crackles were not further divided into coarse and fine crackles. One significant annotation-related limitation is that each recording was annotated by a single labeler; thus, the labels were used as ground truth for model training and testing but did not represent perfect ground truth.
HF_Tracheal_V1 [32] is a specialized event-annotated tracheal sound database designed for automated respiratory monitoring. It contains 10,448 non-overlapping 15 s recordings obtained from 227 Taiwanese adults undergoing diagnostic or surgical procedures under monitored anesthesia care. The sounds were recorded from the left or right pretracheal region close to the thyroid cartilage using two smartphone-based devices, HF-Type-2 and HF-Type-3 (Taiwan), at a sample rate of 4 kHz and a resolution of 16 bits. The database contains temporal onset and offset annotations for 21,741 inhalation events, 15,858 exhalation events, and 6414 continuous adventitious sound events. CAS events were not further subclassified as wheeze, stridor, or rhonchus and also included snoring, whereas discontinuous adventitious sounds were not annotated. The data were divided into participant-independent training and test sets containing 186 and 41 individuals, respectively. Audio files can be accessed online, but annotation files require an application, making them only partially accessible.
The Peking University (PKU) Respiratory Dataset [33] was gathered from 40 local East Asian participants, including 20 healthy individuals and 20 hospitalized patients with lung diseases. Healthy participants’ lung sounds were recorded using an Exagiga Electric ETZ-1A (Zibo Exagiga Electric Co., Ltd., China) stethoscope at 44 kHz, while patient recordings were obtained using a 3M Littmann 3200 at 4 kHz. Two physicians labeled respiratory cycles as ’normal,’ ’crackle,’ or ’wheeze’ and excluded low-quality samples. All recordings were resampled to 10 kHz, filtered, normalized, segmented, and standardized to 6 s using cropping or padding. The dataset has 11,968 respiratory-cycle samples after time-shift augmentation; however, the original recordings and unaugmented cycles were not provided separately. The dataset was developed to supplement the ICBHI 2017 dataset. When combined with ICBHI, the additional PKU samples increased the CNN and ResNet50 models’ reported specificity and sensitivity. PKU respiratory dataset’s potential limitations include a small number of participants, the use of different recording devices for healthy and diseased groups, the inclusion of augmented samples in the reported total, unclear participant-level data splitting, a limited three-class annotation scheme, and no external validation.

3.2.3. Recording-Level and Patient-Level Respiratory Sound Datasets

In addition to publicly available datasets such as the ICBHI, some studies have provided private respiratory sound databases collected at specific medical centers to increase data diversity and clinical validity. For example, Liao et al. [34] created the LD-DF RSdb dataset from 145 patients (85 males and 60 females) at a single medical center using an electronic stethoscope and a mobile device. This dataset consisted of 9584 high-quality respiratory sound recordings: 6435 normal sounds, 2782 crackles, 208 wheezes, and 159 combined sounds, i.e., sounds that occur together as crackles and wheezes.
The CWLSD/KAUH dataset is relevant for patient-level diagnosis because it contains diagnostic categories, including asthma, pneumonia, COPD, bronchitis, heart failure, lung fibrosis, and pleural effusion. It is an open-access lung sound dataset recorded from the chest wall using an electronic stethoscope, as proposed by Fraiwan et al. [27,35]. Since the data were collected at King Abdullah University Hospital in Jordan, some studies also refer to this dataset as the KAUH dataset or the King Abdullah University Hospital dataset; therefore, CWLSD and KAUH should not be considered separate datasets. The dataset includes 112 participants, consisting of 35 healthy and 77 unhealthy subjects, with ages ranging from 21 to 90 years. It contains normal breathing sounds and seven pulmonary conditions: asthma, pneumonia, COPD, bronchitis, heart failure, lung fibrosis, and pleural effusion. The number of basic clinical recordings is 112, as one basic recording is provided for each subject. However, since each recording is exported using Bell, Diaphragm, and Extended filter modes, some studies report this dataset as 336 audio files, or lung sounds. In addition to diagnosis-level labels, the dataset provides sound-type categories, including normal, crepitations, wheezes, crackles, bronchial sounds, wheezes + crackles, and bronchial + crackles. The dataset therefore supports both recording-level sound classification and patient-level diagnosis.

3.2.4. Multimodal and Diagnosis-Oriented Datasets

RespiratoryDatabase@TR is an important multimodal public database in the field of respiratory sound analysis [28]. This dataset was collected at Antakya State Hospital and includes cardiopulmonary auscultation sounds, chest X-ray images, pulmonary function test (PFT) results, and St. George’s Respiratory Questionnaire for COPD Patients (SGRQ-C) questionnaire data. The sounds were recorded using two Littmann 3200 electronic stethoscopes, each with 12 lung-sound channels and four heart-sound channels. The database consists of 77 adult subjects, including healthy controls, asthma patients, and COPD severity groups from COPD0 to COPD4. Each recording was processed using bell, diaphragm, and extended filter modes, resulting in a total of 3696 auscultation sound files. The annotations were verified by two pulmonologists by jointly evaluating auscultation sounds, chest X-rays, and PFT results, and abnormal sound regions were labeled as ‘murmur,’ ‘crackle,’ and ‘wheezing.’ Therefore, RespiratoryDatabase@TR is a useful resource not only for lung sound analysis but also for multimodal respiratory disease assessment and COPD severity-related studies.
BRACETS [36] is an open-access bimodal respiratory collection, including 1097 respiratory sound recordings and 795 thoracic electrical impedance tomography (EIT) recordings from 78 adult volunteers recruited in Portugal and Greece. The study population comprises both healthy subjects and patients suffering from asthma, COPD, interstitial lung disease, and pulmonary infection. Respiratory sounds were collected from numerous chest auscultation sites using a 3M Littmann 3200 electronic stethoscope, while EIT data were captured using the Goe-MF II system (CareFusion, Höchberg, Germany). The repository contains subject-level diagnoses, clinical metadata, acquisition protocol information, and patient-independent cross-validation splits, but it lacks temporal annotations of individual adventitious respiratory sound events. Baseline experiments compared sound-only, EIT-only, and multimodal fusion models across sample- and subject-level classification tasks. The main advantages of BRACETS are its open availability, multimodal design, various auscultation locations, and reproducible patient-independent evaluation. BRACETS is limited by its small and imbalanced cohort and by the absence of event-level crackle and wheeze annotations. The baseline models used subject-isolated five-fold cross-validation but were not evaluated on an independent external cohort [36].

3.3. RQ2. How Are Respiratory Sounds Acquired and Annotated?

Respiratory sounds were acquired using commercial digital stethoscopes, custom smartphone-coupled or multichannel systems, and multimodal platforms. The recorded signal is influenced by the acoustic properties of the device, anatomical recording site, clinical environment, patient posture, and breathing pattern. These acquisition conditions also affect which respiratory events can be identified and how they are annotated. Acquisition and annotation should therefore be considered jointly when interpreting dataset labels and model performance [24,25,26,28,32,33,34,36].

3.3.1. Acquisition Hardware, Anatomical Location, and Recording Context

Across the reviewed studies, respiratory sounds were acquired using three broad configurations: commercial digital or electronic stethoscopes, custom smartphone-coupled or multichannel acoustic systems, and multimodal platforms combining auscultation with complementary physiological or clinical data. Recordings were obtained from tracheal or pretracheal regions and from multiple anterior, posterior, lateral, and back auscultation sites [24,25,26,28,32,33,34,36]. Several studies additionally developed low-cost, wireless, or portable systems for home monitoring, tele-auscultation, and screening in resource-limited settings [37,38].
Custom acquisition systems were also developed to support portable and remote respiratory monitoring. Geng et al. [37] developed a portable lung-sound acquisition system comprising a custom stethoscope head, an audio amplification circuit, an embedded microcontroller, and an ESP8266 Wi-Fi module (Espressif Systems, China) for wireless transmission of recordings to a computer. The system was evaluated using recordings collected from both home-managed and hospitalized patients. Shivaanivarsha et al. [38] developed a low-cost microphone-based digital stethoscope incorporating signal amplification and bandpass filtering before spectrogram-based CNN classification. Modena et al. [39] showed that the vibrating membrane and sensor design of electronic stethoscopes can influence lung sound acquisition quality.
The ICBHI 2017 dataset is a clear example of the acquisition heterogeneity problem. Its recordings were collected from the trachea and multiple anterior, posterior, and lateral chest locations at multiple research centers, using different electronic stethoscopes and microphone systems. This structure increases the true clinical heterogeneity of the dataset but may make it difficult to distinguish device, center, and anatomical location effects from pathological changes. Therefore, when evaluating ICBHI-based models, device- and center-related domain effects should be considered in addition to participant-independent splits [24].
The acquisition protocol in SPRSound is tailored to the features of the pediatric population. Because respiratory sounds in children may be weaker than in adults, and heart sounds may interfere more with anterior chest recordings, recordings were made from four posterior and lateral back locations. During recording, infants and young children may be in their parents’ arms, sitting, or in supine or prone positions [25].
HF_Lung_V1 combined sequential recordings acquired with a commercial Littmann 3200 stethoscope and synchronous multichannel recordings obtained using the custom HF-Type-1 system (Heroic Faith Medical Science Co., Ltd., Taiwan) at multiple chest-wall sites. Longer recordings were subsequently divided into 15 s files. Because part of the data originated from mechanically ventilated patients, ventilator and clinical-equipment noise may act as additional acoustic confounders [26].
HF_Tracheal_V1 recordings were obtained from the pretracheal region on the left or right side of the thyroid cartilage using HF-Type-2 and HF-Type-3 smartphone-based systems. Data were collected during monitored anesthesia care: participants were given oxygen via nasal cannula, most recordings were obtained during moderate sedation, and some participants underwent chin-lift or jaw-thrust maneuvers [32].
In a study comparing HF_Lung_V2 and HF_Tracheal_V1, models trained only on lung sounds performed poorly when applied to tracheal sounds, and models trained only on tracheal sounds performed poorly when applied to lung sounds. Mixed-set training and domain adaptation improved cross-domain performance. These findings indicate that anatomical recording location represents a distinct acoustic domain rather than a simple metadata variable [32].
In the PKU dataset, healthy participants were recorded with an Exagiga Electric ETZ-1A at a sampling rate of 44 kHz, while patients with pulmonary disease were recorded with a 3M Littmann 3200 at a sampling rate of 4 kHz. Although all files were subsequently resampled to 10 kHz, this procedure does not completely eliminate differences in sensor response, gain, and embedded filtering between devices. Therefore, the association between device type and clinical group poses a risk of device-label confounding [33].
The authors of the LD-DF RSdb database attempted to standardize the acquisition protocol. Sounds were recorded from six predefined chest locations using an electronic stethoscope and a smartphone app; each recording lasted 20 s, with patients breathing normally in a seated position, and the recording area left open to reduce noise from clothing friction. Limiting stethoscope movement and using daily quality assessment were aimed at reducing acquisition variability [34].
In multimodal datasets, acquisition extends beyond respiratory audio. RespiratoryDatabase@TR combines auscultation with chest radiography, pulmonary function testing, spirometric curves, and SGRQ-C data, whereas BRACETS combines respiratory sounds with thoracic EIT [28,36].
BreathSet represents a breath-monitoring domain that differs from conventional chest-wall auscultation datasets. It was collected from patients with COPD using a low-cost contact wearable device and includes normal, deep, heavy, and other breath categories. These labels are not directly equivalent to auscultatory-event labels such as crackle, wheeze, or rhonchus [36].
More recent work has used metadata explicitly to reduce domain dependence. Kim et al. [40] proposed adaptive metadata-guided supervised contrastive learning, in which environmental, demographic, and acquisition-related metadata were used to construct and reweight domain-aware training objectives. The method reduced domain dependency on both ICBHI 2017 and an independent respiratory sound dataset. These results suggest that device-agnostic respiratory sound analysis requires more than resampling or amplitude normalization.

3.3.2. Annotation Granularity, Label Ontology, and Quality Assurance

Annotation granularity varied across datasets. Event-level resources provide onset and offset times for respiratory phases or adventitious sounds. Cycle-level datasets assign labels to complete respiratory cycles, while recording- and participant-level datasets use whole-file sound categories or clinical diagnoses. These levels represent different prediction tasks and should not be compared directly.
In cycle- or segment-level classification, the model receives a pre-segmented respiratory interval. ICBHI 2017 provides two complementary annotation sets: respiratory-cycle labels comprising normal, crackle, wheeze, and crackle-wheeze categories; and precise temporal locations of crackle and wheeze events [24]. The PKU dataset identifies respiratory cycle boundaries and classifies them as normal, crackle, or wheeze [33].
An event-level annotation provides the exact start and end times of a sound event. HF_Lung_V1 provides temporal boundaries for inhalation, exhalation, CAS, and DAS, whereas HF_Tracheal_V1 provides inhalation, exhalation, and CAS events. SPRSound provides temporally bounded events labeled as normal, rhonchi, wheeze, stridor, coarse crackle, fine crackle, or wheeze-crackle combinations [25].
Recording-level annotations describe the content of the entire audio file. For example, in SPRSound, recordings are labeled as Normal, CAS, DAS, CAS&DAS, or Poor Quality. RespiratoryDatabase@TR and BRACETS use participant- and recording-level clinical diagnosis categories in addition to or instead of acoustic events [25,28,36].
In addition, the label ontology varies across datasets. The categories used by ICBHI 2017 are normal, crackle, wheeze, and crackle-wheeze cycle. Among adventitious sound categories, SPRSound distinguishes Rhonchi, Wheeze, Stridor, Coarse Crackle, Fine Crackle, and combined Wheeze-Crackle events. HF_Lung_V1 includes wheeze, stridor, and rhonchi in the CAS group, while DAS only includes crackles. HF_Tracheal_V1 includes snoring in CAS and does not annotate DAS [24,25,26,32].
SPRSound used a custom SoundAnn annotation tool and a multi-stage quality-assurance procedure involving 11 experienced pediatric physicians. Event and record annotations were subjected to independent evaluation and disagreement-resolution stages [25].
In LD-DF RSdb, the initial label was determined by the agreement of at least two annotators, and in case of disagreement, a panel of three experienced experts made the final decision. The entries were additionally rated on A, B, C, and D quality grades [34].
Table 6 summarizes the acquisition hardware, anatomical recording locations, annotation granularity, and principal label categories of the representative respiratory sound datasets.
Table 6. Acquisition and annotation characteristics of respiratory sound datasets.
In HF_Tracheal_V1, annotations made by a single labeler were reviewed by an independent inspector, and disagreements were resolved by consensus. In HF_Lung_V1, each recording was assigned to a single primary labeler, which may increase the likelihood of class and temporal-boundary uncertainty [26,32].
Annotation uncertainty is also relevant to the ICBHI 2017 respiratory cycle labels because a complete cycle may contain weak, overlapping, or temporally sparse adventitious events. Some recent methods reduce their dependence on precise local labels through multi-instance estimation of patch-level classes [41] or confidence-weighted segment aggregation [42]. These approaches may reduce the influence of ambiguous segments, but they do not explicitly correct potentially incorrect cycle-level annotations. Explicit noisy-label learning, expert reannotation, uncertainty-aware losses, and inter-annotator agreement were not consistently reported across the reviewed studies.
Quality-control procedures differed across datasets. SPRSound treats poor-quality recordings as a recording-level class, LD-DF RSdb uses A–D quality grades, and PKU excludes respiratory cycles judged to be of insufficient quality [25,33,34]. Because Grand Challenge 2024 requires temporal event localization in continuous recordings, its detection results should not be compared directly with accuracy values reported for pre-segmented classification tasks [30].

3.3.3. Acquisition–Annotation Coupling and Its Implications for Model Validity

Acquisition and annotation are closely linked. Recording hardware, anatomical location, clinical context, and acquisition protocol determine which acoustic phenomena are captured and how they can be labeled. Annotation should therefore be interpreted in relation to the measurement conditions under which the signal was recorded, rather than automatically treated as device-independent ground truth [24,25,26,32,33,36].
Lower-airway adventitious sounds and breath phases are studied using chest-wall recordings, whereas pretracheal recordings are sensitive to airflow, upper-airway blockage, stridor, snoring, and apnea-related patterns. The wearable-recorded BreathSet is designed to monitor regular, deep, and heavy breathing patterns. Thus, these domains are not completely interchangeable from a physiological and acoustic perspective [26,32,36].
Dataset-level confounding occurs when disease labels are systematically correlated with dataset-specific factors such as device type, recording site, sampling rate, clinical environment, or acquisition protocol. For example, in the PKU dataset, device type is associated with health status; in the ICBHI, multiple devices, centers, and recording sites are mixed in one dataset; HF_Tracheal_V1 is recorded in a procedural sedation context [24,32,33]. In such cases, the model may learn dataset identity rather than clinically relevant respiratory sound patterns. This may give high results in internal testing but show poor generalization to other datasets/devices.
When multiple recordings or short segments from the same participant are divided using an audio-level random split, recordings from that individual may appear in both the training and test sets. This creates data leakage and may produce overly optimistic performance estimates. Therefore, all recordings from the same participant should be assigned exclusively to one partition. Consistent with this requirement, BRACETS employed subject-isolated five-fold cross-validation, while the HF datasets adopted participant-independent training and testing procedures [26,32,36].
High classification accuracy on a pre-segmented respiratory cycle does not indicate that the model can detect events in continuous recordings. Likewise, accurate crackle or wheeze detection does not necessarily imply reliable disease diagnosis, because acoustic event recognition and patient-level clinical classification rely on different prediction targets and ground-truth definitions. For this reason, event-level, cycle-level, recording-level, and patient-level tasks should be evaluated as distinct problem settings, with separate reporting of the prediction unit, annotation level, and clinical ground truth [24,25,28,30,36].
Overall, respiratory sound datasets should not be compared solely by the number of participants, recordings, or reported accuracy. They should be evaluated as measurement-and-annotation systems in which hardware, anatomical location, clinical context, annotation granularity, label ontology, quality assurance, and validation design jointly determine the meaning and generalizability of model performance.

3.4. RQ3. What Preprocessing Techniques and Feature Representations Are Used for Respiratory Sound Analysis?

In the reviewed studies, the respiratory sound analysis pipeline included signal conditioning and denoising, segmentation, standardization, data augmentation, and feature representation. The specific methods used varied according to the recording conditions, prediction level, and temporal and frequency features of the observed respiratory event.

3.4.1. Signal Conditioning, Denoising, and Quality Control

Respiratory sound recordings frequently contain environmental noise, heart sounds, speech, motion artifacts, stethoscope friction, and device-related interference. Consequently, filtering and denoising are frequently applied before segmentation or feature extraction [32,36,43,44,45,46,47]. Rishabh and Kumar resampled the recordings to 4 kHz and applied a nominal 25–2000 Hz Butterworth band-pass filter, followed by amplitude normalization [43]. Shiri et al. [44] utilized a zero-phase Butterworth filter from 50 Hz to 2.5 kHz. Choi et al. [48] merged features from low-pass, high-pass, and bandpass signals into three channels, rather than limiting themselves to a single passband. Ohmshankar et al. [49] applied Savitzky–Golay filtering to smooth the lung sound waveform while preserving peak locations.
Other denoising strategies included spectral gating, wavelet decomposition, and variational mode decomposition. Dar et al. [50] employed a Hann window and spectral gating to generate a time-frequency mask based on the estimated noise spectrum. Gupta et al. used variational mode decomposition (VMD) to divide the signal into three modes, reduce low-frequency cardiac and ambient noise, and generate a gammatonegram from the retained respiratory component [51]. Shi et al. used Coiflet-2 wavelet reconstruction to separate the heart sound component by reducing frequencies below 100 Hz; 1 s was cut off at both ends to remove handling artifacts at the beginning and end of the recording [52]. Fava et al. [45] combined VMD and harmonic–percussive source separation with features derived from the RMS envelope, its frequency content, and autocorrelation. A binary classifier was then used to identify recordings containing sufficient respiratory information for subsequent analysis. Denoising seeks to recover usable respiratory information, whereas recording-level quality control determines whether an auscultation should be retained for analysis.
In one evaluation of pre-trained audio neural networks, the ICBHI score decreased from 81.14 for clean recordings to 77.08 under cardiac interference and 70.69 under hospital ambient noise [53]. In a separate physician-validation study, deep-learning-based audio enhancement improved classification under noisy conditions and increased diagnostic sensitivity by 11.61% during model-assisted assessment [54]. Because fine crackles are brief, low-amplitude transients, aggressive smoothing may attenuate clinically relevant signal components together with interference. For this reason, denoising performance should be assessed not only using signal-level measures but also through event-level crackle sensitivity.

3.4.2. Segmentation and Input Standardization

Segmentation determines the temporal unit presented to the model. The reviewed studies used fixed-duration windows, sliding windows, respiratory-cycle segmentation, event-centered segmentation, truncation, zero-padding, and repetition-based padding. These strategies were used to standardize variable-length recordings and generate inputs of consistent dimensions for model training.
Pham et al. combined heterogeneous ICBHI recordings into a single format, repeated short respiratory cycles, and then generated spectrogram patches [55]. Rishabh and Kumar used a smart padding strategy instead of conventional zero-padding: neighboring cycles were concatenated when appropriate, or the current cycle was repeated until an 8 s duration was reached [43]. Similarly, Shiri et al. extended short respiratory events to 7.15 s through cyclic repetition [44]. Wang et al. divided training recordings into 6 s windows with 50% overlap [56]. Zhantleuova et al. converted the recordings to 16 kHz mono signals and normalized them to zero mean and unit variance; because the recordings were obtained under controlled conditions, no additional denoising was applied [57].
Wu et al. [58] proposed a respiratory cycle-based segmentation strategy rather than treating fixed-length padding or trimming as the primary unit of analysis. Each complete respiratory cycle was retained as an information unit, thereby preserving the relationship between inhalation, exhalation, and adventitious sounds. The authors reported an approximately 10% increase in sensitivity compared with the alternative input-processing strategy.
The main preprocessing categories, representative methods, typical use, and potential limitations identified across the reviewed studies are summarized in Table 7.
Table 7. Preprocessing strategies used in respiratory sound analysis.

3.4.3. Data Augmentation

Data augmentation was widely used to address limited sample size and class imbalance. Common waveform-level transformations included pitch shifting, time stretching, temporal shifting, speed variation, noise addition, amplitude modification, and vocal tract length perturbation [44,56,59]. At the spectrogram level, the reviewed methods included cropping, horizontal shifting or flipping, time masking, and mixup [44,55,60,61].
Shiri et al. added Gaussian noise at signal-to-noise ratios of 15–30 dB to minority classes and applied cropping, horizontal flipping, and time masking to spectrograms [44]. Pham et al. used mixup to combine spectrogram patches [55], whereas Wang et al. introduced RandClipMix, which fuses randomly selected clips from recordings of the same class to generate additional training samples [56]. Roslan et al. evaluated both audio- and image-based augmentation for VGG16-based respiratory disease classification using spectrograms derived from the ICBHI 2017 dataset [60]. Although augmentation increased the reported overall accuracy, the improvement was not statistically significant.
A broader sensitivity analysis by Wang et al. [61] compared seven augmentation methods under common VGG-11 and ResNet-18 settings. Within that study, spectrogram flipping, mixup, and SpecMix produced some of the strongest results, although their relative effects varied across model and evaluation settings [61].

3.4.4. Time-Frequency Image Representations

Because respiratory sounds are non-stationary, their clinically relevant characteristics vary across both time and frequency. Therefore, time-frequency representations such as short-time Fourier transform (STFT) spectrograms, Mel spectrograms, log-Mel spectrograms, constant-Q transforms (CQT), wavelet scalograms, gammatonegrams, and cochleograms have been widely used.
Pham et al. compared log-Mel spectrograms, gammatone-based spectrograms, stacked MFCC representations, and CQT inputs. Their results showed that the optimal representation depended on the prediction level: an auditory-inspired representation performed better for cycle-level respiratory anomaly classification, whereas log-Mel spectrograms were more effective for recording-level disease prediction [55]. Representation performance therefore varied between cycle-level anomaly classification and recording-level disease prediction.
Auditory-inspired and multichannel representations were explored in several studies. Gupta et al. employed gammatonegrams to approximate cochlear frequency selectivity after variational mode decomposition-based signal processing [51]. Choi et al. constructed multichannel inputs by stacking log-Mel and MFCC matrices obtained from different frequency bands [48]. Mang et al. compared spectrogram, MFCC, CQT, and cochleogram representations and reported the best performance when cochleograms were combined with a Vision Transformer [62].
Shiri et al. [44] also found that representation performance depends on the model architecture. The authors compared STFT spectrograms, mel-spectrograms, and MFCC representations under similar experimental conditions. STFT inputs performed better with their custom CNN and VGG16 models, whereas MFCC inputs were more effective with InceptionV3 [44]. Wang et al. proposed an enhanced mel-spectrogram that emphasized frequency regions associated with adventitious respiratory sounds [56]. This comparison shows that the best time-frequency representation depends on the model architecture and the target task.
Phettom et al. used STFT representations for respiratory-cycle detection and abnormal lung sound classification using the ICBHI dataset [63]. After applying a fifth-order Butterworth band-pass filter, STFT images were generated and provided as input to a pretrained GoogLeNet model. The binary crackle-versus-wheeze task produced better results than the three-class classification of normal, crackle, and wheeze sounds. Ohmshankar et al. combined STFT, the Stockwell transform, and additional spectral features after Savitzky–Golay filtering [49]. The resulting feature set was processed using an enhanced long short-term memory (LSTM) model.
Wavelet-based methods were also used to provide multiresolution analysis. Laasya et al. divided 15 s recordings from the HF_Lung_V1 dataset into 3 s segments and generated continuous wavelet transform scalograms using an analytic Morlet wavelet [64]. These scalograms were used to distinguish healthy sounds from crackles, rhonchi, stridor, and wheeze. Their multiresolution structure is particularly suitable for representing short transient events such as crackles. Cansiz et al. used the tunable Q-factor wavelet transform, which provides adjustable time-frequency resolution according to the oscillatory characteristics of the signal [65].
The STFT window length directly controls the trade-off between temporal and frequency resolution. Short windows provide better localization of brief crackle events but reduce frequency resolution, whereas longer windows improve frequency resolution for sustained wheezes but may smear short discontinuous sounds. Therefore, the same STFT setting may not be optimal for both crackle detection and wheeze classification.
The principal time-frequency representations, their typical use, and their potential limitations are summarized in Table 8.
Table 8. Feature representations used in respiratory sound analysis.
Overall, representation performance varied with both the prediction task and the model architecture. STFT- and Mel-based inputs were used broadly across the reviewed studies, whereas auditory-inspired and wavelet-based representations provided alternative descriptions of spectral and multiresolution structure. Meaningful comparisons therefore require the representations to be evaluated using the same dataset, data split, prediction task, and model architecture.

3.5. RQ4. What Machine Learning and Deep Learning Models Are Used for Respiratory Sound Analysis?

The reviewed approaches described in RQ3 were grouped according to their feature-based learning, temporal modeling, transfer learning, and deployment strategies. The main groups included handcrafted-feature ML, CNN-based models, recurrent and temporal-convolutional architectures, transformers, pretrained audio models, and deployment-oriented systems. Traditional ML methods rely on manually designed descriptors; CNN-based models learn local spectro-temporal patterns; recurrent and temporal-convolutional models capture dependencies across respiratory frames; and transformer-based models use global attention to model longer acoustic context. Transfer learning, contrastive learning, and model compression have additionally been used to address limited labeled data and deployment constraints. Multi-task learning has also been explored to jointly classify lung sounds and lung diseases [66]. However, the reported performances are not directly comparable because the studies used different datasets, prediction units, class definitions, preprocessing pipelines, and validation protocols.

3.5.1. Handcrafted Feature-Based Models

Handcrafted feature-based models remain relevant when the available dataset is small, interpretability is important, or computational resources are limited. Tasar et al. [67] combined tunable Q-factor wavelet transform (TQWT)-based signal decomposition, Piccolo-pattern nonlinear descriptors, iterative neighborhood component analysis (INCA) feature selection, and conventional classifiers, with k-nearest neighbor (KNN) producing the strongest reported result. Koshta et al. [68] used discrete cosine transform (DCT)- and discrete Fourier transform (DFT)-based Fourier decomposition to obtain intrinsic frequency bands and classified their statistical descriptors using optimized support vector machine (SVM), KNN, and ensemble models. Fraiwan et al. [69] similarly combined Shannon, logarithmic-energy, and spectral entropy features with bagging and boosting classifiers.
These methods produce compact and relatively interpretable feature vectors and generally require fewer trainable parameters than deep networks. However, their performance depends strongly on manually selected transformations, statistical descriptors, and feature selection procedures. Their reported results should also be interpreted together with the validation design. For example, combining recordings from different sources may increase the apparent sample size but can introduce device-, center-, and population-related confounding when a clearly patient-independent evaluation protocol is not used.

3.5.2. CNN-Based Spectro-Temporal Models

CNNs were widely used because respiratory sounds are commonly represented as spectrograms, Mel-spectrograms, MFCC matrices, or wavelet-based images. CNN performance was closely linked to the front-end representation, frequency resolution, and effective receptive field. Alqudah et al. [70] explored DL approaches for lung-sound classification using audio-derived inputs.
Stas et al. [71] used a spectrogram-based CNN regression framework to estimate crackle counts rather than categorical sound labels. Segment-level predictions were aggregated to obtain a recording-level crackle count. The comparison of different convolution and pooling configurations suggested that several small kernels and stride-based pooling were more effective than a single large kernel. However, the model was trained primarily using simulated sounds and validated on a small real dataset, limiting conclusions regarding clinical generalization.
Dhavala et al. [72] used MFCC inputs and a CNN for subject-independent classification of healthy, chronic, and non-chronic pulmonary conditions. The subject-independent split was an important methodological strength because recordings from testing patients were excluded from training. Tong et al. [73] incorporated a feature-band attention module into a ResNet-based framework to emphasize physiologically relevant frequency regions for four-class ICBHI respiratory sound classification. Chanane et al. [74] compared STFT, CQT, Mel-spectrogram, and combined representations, whereas Wu et al. [21] processed STFT and wavelet representations in parallel using an improved Bi-ResNet architecture. Ma et al. [75] combined DenseNet-based classification with a sound preprocessing engine for respiratory disease diagnosis.
The reviewed studies show that CNNs are effective for detecting local spectro-temporal structures, but their output is highly dependent on the selected front-end representation. A CNN operating on a fixed-resolution STFT cannot recover temporal or frequency details that were lost during spectrogram construction. CNNs may also provide limited modeling of relationships across distant respiratory phases unless large receptive fields, dilated convolutions, recurrent layers, or attention mechanisms are added.
Roy and Satija [11] proposed the AsTFSONN architecture for asthmatic lung-sound classification, while their subsequent study [12] introduced a multi-head self-organized operational neural network for COPD detection. Although both studies reported high internal performance on the CWLSD dataset, results from a small single-dataset experiment should not be treated as equivalent to performance obtained using patient-independent or external validation.

3.5.3. Temporal and Hybrid Architectures

CNN-RNN and CNN-LSTM models were developed to combine local spectral-pattern extraction with temporal sequence modeling. The CNN component identifies local time-frequency structures, while LSTM or gated recurrent unit (GRU) layers model their order and evolution across the respiratory cycle.
Petmezas et al. [76] combined an STFT-based CNN with an LSTM and used focal loss to reduce the effect of class imbalance. The model achieved 52.78% sensitivity, 84.26% specificity, and 76.39% accuracy under inter-patient cross-validation. The difference between sensitivity and specificity illustrates that acceptable overall accuracy may coexist with limited detection of minority abnormal classes.
Hakki and Serbes [77] evaluated a CRNN with GRU layers for event-level wheeze detection. The reported results should be interpreted in relation to the number and duration of annotated wheeze events and the validation protocol. Papadakis et al. [78] compared CNN-only, recurrent-only, and CNN-LSTM configurations in AusculNET and found that the hybrid architecture provided a better balance between local feature extraction and temporal modeling. Its 8-bit quantization also reduced model size and inference time with only a small decrease in the ICBHI score.
Le et al. [29] used a multi-level temporal convolutional network after VGG19 feature extraction. Unlike recurrent layers, temporal convolutional networks (TCNs) use dilated temporal convolutions that can model long acoustic context while allowing parallel computation. Their comparison of recording-level and patient-level splits also demonstrated that recording-level evaluation can inflate reported performance when recordings from the same patient occur in both training and testing sets.
Recurrent models are appropriate when the sequential order of respiratory phases is important, but their computation is inherently sequential. TCNs provide larger temporal receptive fields and more parallel processing, although their effectiveness remains dependent on the selected dilation pattern and input-segment duration.

3.5.4. Transfer Learning and Transformer-Based Models

Transfer learning was primarily used to reduce dependence on large expert-labeled respiratory datasets. Common transfer-learning strategies included pretraining on a public respiratory dataset before adaptation to a smaller clinical cohort and fine-tuning an audio model pretrained on a large general-purpose sound dataset. Nguyen et al. [79] pretrained a multi-input CNN on ICBHI 2017 and fine-tuned it on a small multi-channel clinical dataset for crackle detection. Combining complete respiratory-cycle information with inspiration-phase information produced an F-score of 84.71% in the target domain. Roy and Satija [80] fine-tuned Yet Another Mobile Network (YAMNet), which had been pretrained on AudioSet, for COPD severity classification.
Transformer models use self-attention to model relationships between distant time-frequency regions. Wu et al. [58] evaluated an audio spectrogram transformer and proposed a dual-input configuration combining spectrogram and log-Mel representations. The attention mechanism enabled the model to integrate broader acoustic context than a conventional local CNN. In the reported AST experiment, sensitivity and specificity were 42.91% and 62.11%, respectively [58]. These results show that the use of global attention alone did not resolve the effects of limited data and class imbalance.
Transformers can capture long-range dependencies and support multi-input fusion, but they generally require more training data, memory, and computation than lightweight CNNs. Their benefit is therefore most plausible when the task requires broad temporal context and sufficient training or pretraining data are available.

3.5.5. Comparative Analysis of Model Families

Handcrafted models are computationally efficient and potentially interpretable but depend on manually selected descriptors. CNNs are well suited to local spectro-temporal patterns such as crackles and wheeze bands, although their performance is strongly influenced by the input representation and receptive-field size. CNN-RNN models add sequential modeling and are suitable for respiratory-cycle analysis, but recurrent processing increases latency and may overfit small datasets. TCNs capture longer temporal context with greater parallelism. Transformers model global relationships and complementary inputs but generally require more data and computation. Transfer learning and pretrained audio models improve label efficiency, although their performance can deteriorate when the pretraining and clinical domains differ.
Table 9 summarizes the principal model families used in respiratory sound analysis, together with their main modeling capabilities, typical use, and potential limitations.
Table 9. Comparative characteristics of major model families used in respiratory sound analysis.
Event-level crackle or wheeze detection requires fine temporal localization; cycle-level classification requires modeling the inspiration–expiration sequence; and record- or patient-level disease prediction requires aggregation across multiple events and respiratory cycles. Consequently, performance values obtained for these tasks should not be placed in a single ranking without considering the prediction unit, annotation level, patient-independent validation, and class distribution.

3.5.6. Deployment-Oriented and Lightweight Models

Lightweight respiratory sound models aim to reduce memory use, computational cost, and inference latency for digital stethoscopes, mobile phones, wearable sensors, and edge platforms. However, the term “lightweight” should be supported by quantitative measurements rather than by architecture naming alone.
Roy and Satija [13] proposed the Respiratory Disease Lightweight Inception Network (RDLINet) for mel-spectrogram-based respiratory disease classification. Although high internal accuracies were reported, the model combined datasets obtained using different devices, acquisition procedures, and annotation schemes. Therefore, its generalization should be assessed using patient-independent and external validation.
Hu et al. [81] implemented a ResNet-18-based respiratory classifier on a Xilinx Zynq ZCU102 field-programmable gate array (FPGA). Quantization-aware training, layer merging, memory optimization, and parallel processing reduced model size by 40% and computational latency by 70%, producing an inference latency of 16 ms with less than 2% degradation in the inference score. A smartphone-oriented implementation reduced a pruned and quantized CNN-LSTM model to approximately 0.38 MB with approximately 1% performance degradation [19]. Liu et al. [82] used a two-stage architecture that activated the computationally expensive fine-grained classifier only when an abnormal signal was detected, reducing the average computational workload and energy demand for normal recordings.
Park et al. [83] specifically framed lung sound classification as an on-device AI problem, emphasizing local inference for digital auscultation.
Table 10 summarizes the reported model size, inference latency or energy reduction, target hardware, and associated performance trade-offs of deployment-oriented respiratory sound models.
Table 10. Computational characteristics of deployment-oriented respiratory sound models.

3.5.7. Emerging Audio Representation Learning

Recent studies show that respiratory sound analysis is no longer limited to fully supervised classification. Newer approaches increasingly use contrastive learning, self-supervised learning, and pretrained audio foundation models to learn more general and reusable acoustic representations. Soni et al. [16] demonstrated that contrastive pretraining on heart and lung sounds can reduce dependence on large expert-labeled datasets. Moummad and Farrugia [17] extended this direction by incorporating age, sex, and other metadata into supervised contrastive pretext tasks. Their results indicate that metadata can provide useful supervisory information when class labels are scarce or imbalanced.
Castejón-Barrio and Gallardo-Antolín [59] trained a respiratory sound encoder using self-supervised contrastive learning on unlabeled recordings. The learned representations were then transferred to ICBHI 2017 and SPRSound classification tasks. The approach was particularly beneficial when only a limited proportion of labeled data was available, although its performance still depended on the acoustic similarity between the pretraining and target datasets.
Audio foundation models represent a related emerging direction. Niizumi et al. [18] compared publicly available foundation models across respiratory and heart sound tasks. General audio models performed competitively on relatively clean recordings but were less reliable on noisy tasks, indicating that large-scale pretraining does not automatically overcome the domain characteristics of clinical auscultation. Ehtesham et al. [84] applied Google’s Health Acoustic Representations model to pediatric respiratory sounds from SPRSound, using pretrained embeddings rather than learning all features from the limited target dataset.
LungListener extended multimodal representation learning by adapting an audio-language model to integrate lung-sound recordings with patient metadata for classification and analysis [85].
Self-supervised and foundation-model approaches improve label efficiency and transferability, but their benefits are constrained by domain mismatch, recording noise, device variability, and differences between general health audio and chest-auscultation signals. Domain-adaptation approaches have also been developed specifically for respiratory sounds. Kim et al. [15] proposed stethoscope-guided supervised contrastive learning to reduce acquisition-device-related distribution shift, whereas Huang et al. [86] introduced a contrastive embedding-based domain adaptation method for pediatric lung-sound recognition. Together with mixed-set training [32] and adaptive metadata-guided supervised contrastive learning [40], these studies indicate that device-agnostic performance requires explicit domain-aware objectives rather than resampling or amplitude normalization alone.

3.6. Cross-Study Synthesis: Four Overlapping Methodological Stages in Respiratory Sound AI

The reviewed literature can be organized into four overlapping methodological stages. These stages are not strictly chronological or mutually exclusive; rather, they provide a conceptual framework for comparing feature engineering, deep time-frequency modeling, representation learning, and deployment-oriented research.
In the first stage, researchers mainly used manually designed acoustic features together with classical ML classifiers. These approaches used features such as mel-frequency cepstral coefficients (MFCCs), wavelet coefficients, entropy measures, Fourier-based descriptors, nonlinear descriptors, and statistical features with classifiers such as SVM, KNN, Random Forest, bagging, boosting, and ensemble models [65,67,68,69]. Their main advantage is interpretability and low computational cost, but their performance depends strongly on manually selected features, preprocessing choices, and feature-selection procedures. The second stage moved toward DL models that analyze respiratory sounds as spectrogram-like time-frequency images. In this stage, respiratory sounds are transformed into STFT spectrograms, Mel-spectrograms, MFCC matrices, CQT representations, wavelet scalograms, gammatonegrams, and then processed using CNN, CNN–RNN, TCN, transformer, or transfer-learning architectures [21,29,44,49,55,58,62,63,76,78,79]. This stage improved benchmark performance and reduced dependence on handcrafted features, but it also introduced new concerns, including task mismatch, overfitting on small datasets, recording-level leakage, class imbalance, and limited external validation [23,29,72]. The third stage is now developing around methods that learn more general respiratory sound representations and combine audio with other types of information. Self-supervised learning, supervised contrastive learning, metadata-guided domain adaptation, pretrained audio foundation models, multimodal fusion, and audio-language models aim to improve label efficiency, domain robustness, and contextual interpretation [16,17,18,36,40,59,84,85,87,88]. These approaches are especially relevant because respiratory sound datasets are often small, imbalanced, noisy, and expensive to annotate. However, their clinical value depends on whether the learned representations generalize across devices, populations, recording sites, and independent clinical settings. The fourth stage addresses practical requirements for on-device and real-world auscultation. In this stage, model accuracy alone is insufficient. Respiratory sound AI systems must also address noise robustness, audio enhancement, signal-quality assessment, confidence-driven event-to-recording fusion, patient-independent evaluation, calibration, computational efficiency, privacy, and deployment on digital stethoscopes, smartphones, wearables, or edge hardware [19,42,46,47,53,54,81,82,89,90,91]. This stage reflects the transition from benchmark classification toward real-world digital auscultation and decision-support systems. The four overlapping methodological directions identified in the reviewed literature are summarized in Table 11.
Table 11. Four overlapping methodological directions in AI-based respiratory sound analysis.

3.7. RQ5. How Are AI-Based Respiratory Sound Analysis Models Evaluated?

3.7.1. Patient-Independent Versus Recording-Level Evaluation

Respiratory sound datasets commonly contain several recordings or segments from the same participant. Random splitting at the recording or segment level may therefore place data from one participant in both the training and test sets, introducing patient-level leakage and inflating performance estimates. Figure 2 illustrates the difference between these two splitting approaches. Le et al. [29] quantified this effect in an ML-TCN study on ICBHI 2017. Recording-level splitting produced higher scores than patient-level splitting across all four tasks, with a maximum difference of approximately 12 percentage points. Patient-independent splitting is therefore necessary when evaluating generalization to unseen participants.
Figure 2. Comparison of recording-level and patient-independent data splitting strategies: (a) recording-level splitting, in which segments from the same participant may be distributed across both the training and test sets, resulting in potential data leakage; (b) patient-independent splitting, in which all recordings and segments from each participant are assigned exclusively to either the training set or the test set.

3.7.2. Validation Protocols and Evaluation Metrics

Several included studies used subject- or patient-independent evaluation protocols. Such an evaluation strategy ensures that there is no patient-level overlap between training and testing data and allows for a fairer assessment of how well the model can generalize to previously unseen patients. For example, Dhavala et al. [72] used a subject-independent evaluation strategy, while Sreejith et al. [92] emphasized the need for patient-independent separation in ICBHI-based respiratory sound classification.
Evaluation metrics also differ depending on the prediction unit and task structure. Metrics, including accuracy, sensitivity, specificity, precision, F1 score, and area under the receiver operating characteristic curve (AUC), are frequently reported in studies on record- or patient-level classification tasks. However, event-level detection tasks necessitate additional metrics that capture the temporal localization of respiratory phases and adventitious sounds. For example, the breath phase and adventitious sound detection tasks on the HF_Lung_V1 [26] dataset were analyzed using a task-specific multi-level evaluation protocol. First, fivefold cross-validation was used on the training dataset to train and validate the models, and the final performance was evaluated using an independent testing dataset. In segment-level detection, predicted time segments were compared with ground-truth time segments, and the receiver operating characteristic (ROC) curve and AUC values were calculated. At the event level, predicted and reference events were matched using the Jaccard index, while differences in the detected event counts were summarized using mean absolute percentage error (MAPE). This protocol separates segment classification from event detection and evaluates both class discrimination and event correspondence.

4. Discussion

4.1. Respiratory Sound AI as a Measurement-to-Deployment Pipeline

Respiratory sound AI performance depends on the whole development pipeline, not only on the model architecture. Recording hardware, anatomical site, annotation level, preprocessing, input representation, data splitting, evaluation metrics, and deployment setting can all affect reported results [24,25,26,28,32,33,34,36]. For this reason, high internal accuracy does not necessarily mean that a model will generalize to continuous recordings or independent clinical populations [15,29,40,86]. Respiratory sound AI should therefore be evaluated as a measurement-to-deployment pipeline rather than as an isolated classification model.
The four-stage framework developed in this review should be interpreted as a set of overlapping methodological directions rather than a strict chronology. Handcrafted and deep spectrogram-based methods remain widely used, while representation learning, multimodal integration, and deployment-oriented optimization address different limitations of supervised benchmark classification. However, progress in architecture has not been matched consistently by external validation or clinical evaluation.

4.2. Dataset Heterogeneity, Annotation Uncertainty, and Task Mismatch

AI-based respiratory sound analysis still faces several methodological challenges. Public datasets such as ICBHI 2017, SPRSound, HF_Lung_V1, CWLSD/KAUH, RespiratoryDatabase@TR, and BRACETS have made it easier to reuse data and compare different methods. However, these datasets are not directly equivalent: they differ in patient population, recording conditions, label definitions, annotation levels, and prediction tasks [24,25,26,27,28,36]. Because datasets support different task structures, such as event-level detection, cycle-level classification, recording-level sound classification, patient-level diagnosis, disease severity assessment, and multimodal classification, results reported across these datasets should not be directly compared without considering the prediction unit and annotation level [24,25,26,28,30,36].
Another important limitation is the uncertainty of respiratory sound annotations. A respiratory cycle labeled as crackle or wheeze may contain only sparse, inconsistent, or temporally localized adventitious events, while event-level annotation requires precise onset and offset decisions [24,26,30]. Therefore, cycle-level classification, event detection, recording-level diagnosis, and patient-level disease prediction should be treated as distinct problem settings. Recent patch-level and confidence-weighted methods reduce reliance on treating every segment within a recording as equally informative [41,42]. However, they do not directly correct potentially incorrect cycle- or recording-level labels. Therefore, future studies should use more standardized acquisition protocols and describe annotation procedures more clearly. This would improve reproducibility, support fairer cross-dataset comparisons, and strengthen clinical applicability [24,32,33,36].

4.3. Preprocessing, Noise Robustness, and Clinical Validity

Preprocessing is not a neutral technical step. Filtering, denoising, segmentation, augmentation, and time-frequency conversion can change the acoustic properties of respiratory sound signals [43,44,45,50,51,52]. Because fine crackles are brief and low-amplitude, aggressive denoising may attenuate clinically relevant components together with interference. Denoising should therefore be evaluated using downstream event-level performance, not only signal-level quality measures.
Recent studies suggest that noise robustness should be considered a key requirement for clinical deployment, not just an optional technical improvement. Heart sounds, hospital ambient noise, speech, crying, stethoscope friction, clothing noise, and device artifacts can substantially reduce model performance [46,47,53]. Audio-enhancement and quality-control methods should be evaluated using event preservation, class-wise sensitivity, and downstream diagnostic performance rather than waveform smoothness or global accuracy alone.

4.4. Model Choice Should Follow the Prediction Unit

The reviewed studies do not show that one model architecture is clearly better than all others. Instead, the most suitable model depends on the prediction unit and the clinical goal. Handcrafted models are still useful when the task requires interpretable features and low computational cost [65,67,68,69]. CNNs are commonly applied to local spectro-temporal patterns [21,71,73], whereas CNN-RNN and TCN architectures were used to incorporate respiratory sequence information and longer temporal context [29,76,78]. Transformers and pretrained audio models may capture broader acoustic context, but they require sufficient data, careful regularization, and validation under realistic noise and domain-shift conditions [18,58,80,84].
In practice, the key question is not only whether a model achieves high accuracy, but whether it is appropriate for the task. Event detection requires temporal localization, cycle-level classification focuses on within-breath patterns, recording-level classification requires aggregation across segments, and patient-level diagnosis depends on information from multiple recordings and clinical context [24,25,26,28,30,42]. Thus, accuracy values should not be ranked across studies unless the prediction unit, annotation level, class distribution, split strategy, and external validation are comparable.

4.5. Evaluation, Deployment, and Reporting Transparency

Evaluation is still a major challenge in respiratory sound AI. Patient-independent splitting is essential because recordings or segments from the same participant may otherwise appear in both training and testing sets, causing data leakage and inflated performance [29,72]. For imbalanced datasets, overall accuracy and micro-averaged F1 can hide poor performance in rare classes. Therefore, studies should at least report class-wise sensitivity, specificity, and macro-averaged metrics. When possible, precision-recall analysis, confidence intervals, calibration, and sensitivity at fixed specificity thresholds should also be included [5,23,26,29].
Deployment-oriented research should report computational constraints together with accuracy. Model size, parameter count, FLOPs or MACs, memory footprint, inference latency, energy consumption, target hardware, and performance degradation after pruning or quantization are necessary to determine whether a system is suitable for smartphones, digital stethoscopes, wearables, or edge devices [13,19,22,78,81,82,89,90,91]. Finally, reproducibility should be improved by releasing code, trained weights, train-test partitions, configuration files, and preprocessing scripts whenever possible. Without such information, high reported performance remains difficult to verify or translate into clinical practice.

4.6. Limitations of the Scoping Review

This scoping review has several limitations. The search was limited to English-language publications and five information sources, and Google Scholar was used only as a supplementary source, with the first 100 relevance-ranked results screened. The main search focused on studies published between 2021 and 2026, while earlier studies were included only when they introduced foundational respiratory sound datasets identified through backward reference checking.
No formal risk-of-bias or quality appraisal was performed because the aim was to map the field rather than assess intervention effectiveness. The high heterogeneity of datasets, acquisition devices, annotation levels, prediction tasks, validation protocols, and evaluation metrics also prevented quantitative pooling or direct comparison of model performance. Finally, repository availability was checked at a specific time point, and the availability of code, trained weights, or related files may change over time. Funding sources of the individual included reports were not systematically charted.
Model-level reproducibility also remains limited. Among the 71 eligible computational studies, publicly accessible study-specific source code was identified for 9 studies (12.7%), whereas public study-trained model weights or checkpoints were identified for only 4 studies (5.6%). Neither resource was identified for 62 studies (87.3%) at the time of verification. More consistent release of executable pipelines, data partitions, preprocessing scripts, configuration files, software environments, and trained checkpoints is therefore needed to support independent validation and fair cross-study comparison.

5. Conclusions

This scoping review examined AI-based respiratory sound analysis across the complete methodological pipeline, including datasets, acquisition and annotation practices, preprocessing and feature representations, ML and DL models, evaluation protocols, deployment considerations, and reporting transparency.
The reviewed literature was organized into four overlapping methodological stages. These stages reflect the broader development of the field from handcrafted acoustic analysis to deep time-frequency modeling, representation learning and multimodal integration, and deployment-oriented, robustness-focused systems. Early studies [65,67,68,69] primarily relied on engineered acoustic descriptors, whereas subsequent research [21,29,58,76,78,79] increasingly adopted deep spectro-temporal modeling. More recent work [16,17,18,40,59,84,85,87] has expanded toward contrastive and self-supervised representation learning, pretrained audio foundation models, metadata integration, and multimodal analysis. In parallel, deployment-oriented studies [19,42,53,54,81,82] have placed greater emphasis on computational efficiency, noise robustness, external validation, and clinical applicability.
Although the field has made clear progress, several methodological challenges remain. Studies still differ widely in their datasets, recording devices, annotation schemes, preprocessing steps, validation protocols, and reported metrics [5,6,24,25,26,27,28,36]. Small datasets, class imbalance, device variability, uncertain labels, inconsistent patient-independent evaluation, and limited external validation continue to make clinical translation difficult [5,15,23,29,40].
Further progress will therefore require more than increasingly complex neural architectures. Clinical translation depends on standardized acquisition procedures, transparent annotation, patient-independent and external validation, reproducible reporting, and evaluation in real diagnostic workflows [45,53,54,61]. Multimodal systems combining respiratory sounds with patient metadata, spirometry, imaging, electrical impedance tomography, or vital signs may provide additional physiological and clinical context [28,36,85,87,88]. These developments are necessary to move the field beyond isolated benchmark classification toward reliable and deployable digital auscultation systems.

Supplementary Materials

The following supporting information is available: https://www.mdpi.com/article/10.3390/technologies14080486/s1, Supplementary Material S1: Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) Checklist according to Tricco et al. [10].

Author Contributions

Conceptualization, U.S. and N.S.; methodology, U.S. and P.A.; investigation, U.S., L.I., K.T., A.N., A.M., G.Y. and D.B.; resources, U.S., P.A., N.S. and M.Z.; data curation, U.S., L.I., K.T., A.N., A.M., G.Y. and D.B.; writing—original draft preparation, U.S.; writing—review and editing, U.S., P.A., L.I., K.T., A.N., A.M., N.S., M.Z., G.Y. and D.B.; visualization, U.S.; supervision, N.S. and M.Z.; project administration, N.S. and M.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan, grant number AP26197917, “Development of a machine learning-based recommendation system for the early diagnosis of respiratory diseases using lung sound analysis”.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

No new data were created in this study. Data sharing is not applicable to this article.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT (GPT-5.6 Sol, OpenAI; accessed on 3 August 2026) only to improve the organization and structural clarity of selected sections. The graphical abstract was also created with its assistance based on scientific content and instructions provided by the authors. All AI-assisted outputs were reviewed and edited by the authors, who take full responsibility for the final content.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
AUCArea Under the Curve
CASContinuous Adventitious Sounds
CNNConvolutional Neural Network
COPDChronic Obstructive Pulmonary Disease
CQTConstant-Q Transform
CRNNConvolutional Recurrent Neural Network
CRSAComputerized Respiratory Sound Analysis
CWLSDChest Wall Lung Sound Database
CWTContinuous Wavelet Transform
DASDiscontinuous Adventitious Sounds
DCTDiscrete Cosine Transform
DFTDiscrete Fourier Transform
DLDeep Learning
EITElectrical Impedance Tomography
FFTFast Fourier Transform
FPGAField-Programmable Gate Array
FLOPsFloating-Point Operations
GRUGated Recurrent Unit
ICBHIInternational Conference on Biomedical and Health Informatics
INCAIterative Neighborhood Component Analysis
KNNk-Nearest Neighbor
LSTMLong Short-Term Memory
MACsMultiply Accumulate Operations
MAPEMean Absolute Percentage Error
MFCCMel-Frequency Cepstral Coefficients
MLMachine Learning
ML-TCNMulti-Level Temporal Convolutional Network
PFTPulmonary Function Test
RDLINetRespiratory Disease Lightweight Inception Network
RMSRoot Mean Square
RNNRecurrent Neural Network
ROCReceiver Operating Characteristic
SGRQ-CSt. George’s Respiratory Questionnaire for COPD Patients
STFTShort-Time Fourier Transform
SVMSupport Vector Machine
TCNTemporal Convolutional Network
TQWTTunable Q-Factor Wavelet Transform
VMDVariational Mode Decomposition
YAMNetYet Another Mobile Network

References

  1. Oh, J.; Kim, S.; Yim, Y.; Kim, M.S.; GBD 2023 Global Chronic Respiratory Disease and Covid Collaborators; Hay, S.I.; Shin, J.I.; Yon, D.K. Global, Regional, and National Burden of Chronic Respiratory Diseases and Impact of the COVID-19 Pandemic, 1990–2023: A Global Burden of Disease Study. Nat. Med. 2026, 32, 197–223. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. GBD 2021 Lower Respiratory Infections and Antimicrobial Resistance Collaborators. Global, Regional, and National Incidence and Mortality Burden of Non-COVID-19 Lower Respiratory Infections and Aetiologies, 1990–2021: A Systematic Analysis from the Global Burden of Disease Study 2021. Lancet Infect. Dis. 2024, 24, 974–1002. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Moberg, A.; Emilsson, O.I.; Hansdottir, S.; Asmundsson, T.; Malinovschi, A.; Melbye, H.; Ludviksdottir, D. Lung auscultation-today and tomorrow: A narrative review. Expert Rev. Respir. Med. 2025, 19, 879–885. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Archana, M.; Dash, S.; HR, S.S.; V, S.D. Predicting the Severity of Pulmonary Disease from Respiratory Sounds using ML Algorithms. In Proceedings of the 2025 6th International Conference on Mobile Computing and Sustainable Informatics (ICMCSI), Goathgaun, Nepal, 7–8 January 2025; IEEE: New York, NY, USA, 2025; pp. 1750–1755. [Google Scholar] [CrossRef] [Scilit]
  5. Garcia-Mendez, J.P.; Lal, A.; Herasevich, S.; Tekin, A.; Pinevich, Y.; Lipatov, K.; Wang, H.-Y.; Qamar, S.; Ayala, I.N.; Khapov, I.; et al. Machine learning for automated classification of abnormal lung sounds obtained from public databases: A systematic review. Bioengineering 2023, 10, 1155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Xia, T.; Han, J.; Mascolo, C. Exploring machine learning for audio-based respiratory condition screening: A concise review of databases, methods, and open issues. Exp. Biol. Med. 2022, 247, 2053–2061. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Abreu, V.; Oliveira, A.; Duarte, J.A.; Marques, A. Computerized respiratory sounds in pediatrics: A systematic review. Respir. Med. X 2021, 3, 100027. [Google Scholar] [CrossRef] [Scilit]
  8. Tabatabaei, S.A.H.; Fischer, P.; Schneider, H.; Koehler, U.; Gross, V.; Sohrabi, K. Methods for adventitious respiratory sound analyzing applications based on smartphones: A survey. IEEE Rev. Biomed. Eng. 2021, 14, 98–115. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Zaben, S.O.; Zainon, W.M.N.W.; Sabry, A.H. Machine learning-based methods for detecting respiratory abnormalities using audio and visual analysis: A review. Results Eng. 2025, 26, 104744. [Google Scholar] [CrossRef] [Scilit]
  10. Tricco, A.C.; Lillie, E.; Zarin, W.; O’Brien, K.K.; Colquhoun, H.; Levac, D.; Moher, D.; Peters, M.D.J.; Horsley, T.; Weeks, L.; et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann. Intern. Med. 2018, 169, 467–473. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Roy, A.; Satija, U. AsTFSONN: A Unified Framework Based on Time-Frequency Domain Self-Operational Neural Network for Asthmatic Lung Sound Classification. In Proceedings of the 2023 IEEE International Symposium on Medical Measurements and Applications (MeMeA), Jeju, Republic of Korea, 14–16 June 2023; IEEE: New York, NY, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  12. Roy, A.; Satija, U. A Novel Multi-Head Self-Organized Operational Neural Network Architecture for Chronic Obstructive Pulmonary Disease Detection Using Lung Sounds. IEEE/ACM Trans. Audio Speech Lang. Process. 2024, 32, 2566–2575. [Google Scholar] [CrossRef] [Scilit]
  13. Roy, A.; Satija, U. RDLINet: A Novel Lightweight Inception Network for Respiratory Disease Classification Using Lung Sounds. IEEE Trans. Instrum. Meas. 2023, 72, 4008813. [Google Scholar] [CrossRef] [Scilit]
  14. Tran-Anh, D.; Vu, N.H.; Nguyen-Trong, K.; Pham, C. Multi-Task Learning Neural Networks for Breath Sound Detection and Classification in Pervasive Healthcare. Pervasive Mob. Comput. 2022, 86, 101685. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Kim, J.-W.; Bae, S.; Cho, W.-Y.; Lee, B.; Jung, H.-Y. Stethoscope-Guided Supervised Contrastive Learning for Cross-Domain Adaptation on Respiratory Sound Classification. In Proceedings of the 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Republic of Korea, 14–19 April 2024; IEEE: New York, NY, USA, 2024; pp. 1431–1435. [Google Scholar] [CrossRef] [Scilit]
  16. Soni, P.N.; Shi, S.; Sriram, P.R.; Ng, A.Y.; Rajpurkar, P. Contrastive learning of heart and lung sounds for label-efficient diagnosis. Patterns 2022, 3, 100400. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Moummad, I.; Farrugia, N. Pretraining Respiratory Sound Representations Using Metadata and Contrastive Learning. In Proceedings of the 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA, 22–25 October 2023; IEEE: New York, NY, USA, 2023; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  18. Niizumi, D.; Takeuchi, D.; Yasuda, M.; Nguyen, B.T.; Ohishi, Y.; Harada, N. Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis. In Proceedings of the 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Copenhagen, Denmark, 14–17 July 2025; IEEE: New York, NY, USA, 2025; pp. 1–4. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Oishee, T.T.; Anjom, J.; Mohammed, U.; Hossain, M.I.A. Leveraging deep edge intelligence for real-time respiratory disease detection. Clin. eHealth 2024, 7, 207–220. [Google Scholar] [CrossRef] [Scilit]
  20. Huang, D.M.; Huang, J.; Qiao, K.; Wang, X. Deep Learning-Based Lung Sound Analysis for Intelligent Stethoscope. Mil. Med. Res. 2023, 10, 44. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Wu, C.; Ye, N.; Jiang, J. Classification and Recognition of Lung Sounds Based on Improved Bi-ResNet Model. IEEE Access 2024, 12, 73079–73094. [Google Scholar] [CrossRef] [Scilit]
  22. Lu, X.; Fang, J.; Xiao, W.; Wu, J. Research Progress and Future Prospects in Intelligent Lung Sound Diagnosis: Models, Lightweight Design, and Hardware Platform Implementation. Biomed. Eng./Biomed. Tech. 2025, 70, 483–501. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Yu, S.; Yu, J.; Chen, L.; Zhu, B.; Liang, X.; Xie, Y.; Sun, Q. Advances and Challenges in Respiratory Sound Analysis: A Technique Review Based on the ICBHI2017 Database. Electronics 2025, 14, 2794. [Google Scholar] [CrossRef] [Scilit]
  24. Rocha, B.M.; Filos, D.; Mendes, L.; Serbes, G.; Ulukaya, S.; Kahya, Y.P.; Jakovljevic, N.; Turukalo, T.L.; Vogiatzis, I.M.; Perantoni, E.; et al. An open access database for the evaluation of respiratory sound classification algorithms. Physiol. Meas. 2019, 40, 035001. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Zhang, Q.; Zhang, J.; Yuan, J.; Huang, H.; Zhang, Y.; Zhang, B.; Lv, G.; Lin, S.; Wang, N.; Liu, X.; et al. SPRSound: Open-Source SJTU pediatric Respiratory Sound Database. IEEE Trans. Biomed. Circuits Syst. 2022, 16, 867–881. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Hsu, F.S.; Huang, S.R.; Huang, C.W.; Huang, C.J.; Cheng, Y.R.; Chen, C.C.; Hsiao, J.; Chen, C.W.; Chen, L.C.; Lai, Y.C.; et al. Benchmarking of eight recurrent neural network variants for breath phase and adventitious sound detection on a self-developed open-access lung sound database—HF_Lung_V1. PLoS ONE 2021, 16, e0254134. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Fraiwan, M.; Fraiwan, L.; Khassawneh, B.; Ibnian, A. A dataset of lung sounds recorded from the chest wall using an electronic stethoscope. Data Brief. 2021, 35, 106913. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Altan, G.; Kutlu, Y.; Garbi, Y.; Pekmezci, A.Ö.; Nural, S. Multimedia Respiratory Database (RespiratoryDatabase@TR): Auscultation Sounds and Chest X-Rays. Nat. Eng. Sci. 2017, 2, 59–72. [Google Scholar] [CrossRef] [Scilit]
  29. Le, K.-N.T.; Byun, G.; Raza, S.M.; Le, D.-T.; Choo, H. Respiratory Anomaly and Disease Detection Using Multi-Level Temporal Convolutional Networks. IEEE J. Biomed. Health Inform. 2025, 29, 4834–4846. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Zhang, Q.; Chen, C.; Yuan, S.; Zhang, J.; Yuan, J.; Huang, H.; Zhang, Y.; Pan, R.; Jiang, X.; Zhao, J.; et al. Meta: Data Compression and Event Detection Grand Challenge 2024 with SPRSound Dataset. IEEE Data Descr. 2024, 1, 122–130. [Google Scholar] [CrossRef] [Scilit]
  31. Zhang, Q.; Zhang, J.; Yuan, J.; Huang, H.; Zhang, Y.; Chen, C.; Lin, J.; Zhang, B.; Lv, G.; Lin, S.; et al. Respiratory Sound Track Grand Challenge 2023: Respiratory Sound Classification for SPRSound Dataset. IEEE DataPort, 20 February 2024. Available online: https://ieee-dataport.org/competitions/respiratory-sound-track-grand-challenge-2023-respiratory-sound-classification-sprsound (accessed on 3 June 2026).
  32. Hsu, F.-S.; Huang, S.-R.; Su, C.-F.; Huang, C.-W.; Cheng, Y.-R.; Chen, C.-C.; Wu, C.-Y.; Chen, C.-W.; Lai, Y.-C.; Cheng, T.-W.; et al. A dual-purpose deep learning model for auscultated lung and tracheal sound analysis based on mixed set training. Biomed. Signal Process. Control 2023, 86, 105222. [Google Scholar] [CrossRef] [Scilit]
  33. Zhou, G.; Liu, C.; Li, X.; Liang, S.; Wang, R.; Huang, X. An open auscultation dataset for machine learning-based respiratory diagnosis studies. JASA Express Lett. 2024, 4, 052001. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Liao, X.; Wu, Y.; Jiang, N.; Sun, J.; Xu, W.; Gao, S.; Wang, J.; Li, T.; Wang, K.; Li, Q. Automated detection of abnormal respiratory sound from electronic stethoscope and mobile phone using MobileNetV2. Biocybern. Biomed. Eng. 2023, 43, 763–775. [Google Scholar] [CrossRef] [Scilit]
  35. Fraiwan, M.; Fraiwan, L.; Khassawneh, B.; Ibnian, A. A dataset of lung sounds recorded from the chest wall using an electronic stethoscope [Dataset] (Version 3). Mendeley Data 2021. [Google Scholar] [CrossRef]
  36. Pessoa, D.; Rocha, B.M.; Strodthoff, C.; Gomes, M.; Rodrigues, G.; Petmezas, G.; Cheimariotis, G.-A.; Kilintzis, V.; Kaimakamis, E.; Maglaveras, N.; et al. BRACETS: Bimodal repository of auscultation coupled with electrical impedance thoracic signals. Comput. Methods Programs Biomed. 2023, 240, 107720. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Geng, Y.; Li, Z.; Sun, H. Research on Abnormal Lung Sound Recognition and Diagnosis Based on Improved CNN and Transformer. In Proceedings of the 2023 China Automation Congress (CAC), Chongqing, China, 17–19 November 2023; IEEE: New York, NY, USA, 2023; pp. 6730–6735. [Google Scholar] [CrossRef] [Scilit]
  38. Shivaanivarsha, N.; Sriram, A.; Saravaanan, S.; Rajesh, V. Respiratory Sound Analysis for Lung Disease Diagnosis. In Proceedings of the 2023 International Conference on Ambient Intelligence, Knowledge Informatics and Industrial Electronics (AIKIIE), Ballari, India, 2–3 November 2023; IEEE: New York, NY, USA, 2023; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  39. Modena, M.; Bertacchini, A.; Paltrinieri, F.; Dibiase, L.; Pancaldi, F. Modeling and Simulation of a Vibrating Membrane for the Acquisition of Lung Sounds. IEEE Sens. J. 2025, 25, 24421–24430. [Google Scholar] [CrossRef] [Scilit]
  40. Kim, J.-W.; Toikkanen, M.; Jalali, A.; Kim, M.; Han, H.-J.; Kim, H.; Shin, W.; Jung, H.-Y.; Kim, K. Adaptive metadata-guided supervised contrastive learning for domain adaptation on respiratory sound classification. IEEE J. Biomed. Health Inform. 2025, 29, 5381–5393. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Song, W.; Han, J. Patch-level contrastive embedding learning for respiratory sound classification. Biomed. Signal Process. Control 2023, 80, 104338. [Google Scholar] [CrossRef] [Scilit]
  42. Kala, A.; Elhilali, M. Multi-stage respiratory sound analysis: Confidence-driven wheeze and crackle detection. IEEE Trans. Biomed. Eng. 2026, 73, 1696–1704. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Rishabh; Kumar, D. Multi-Spectral Feature Extraction to Improve Lung Sound Classification Using CNN. In Proceedings of the 10th International Conference on Signal Processing and Integrated Networks (SPIN), Noida, India, 23–24 March 2023; IEEE: New York, NY, USA, 2023; pp. 186–191. [Google Scholar] [CrossRef] [Scilit]
  44. Shiri, G.; Bahrami, H.; Fallahi, A. Comparative Analysis of Time-Frequency Representations for Pediatric Respiratory Sound Classification Using Deep Learning. In Proceedings of the 32nd National and 10th International Iranian Conference on Biomedical Engineering (ICBME), Tabriz, Iran, 19–20 November 2025; IEEE: New York, NY, USA, 2025; pp. 417–424. [Google Scholar] [CrossRef] [Scilit]
  45. Fava, A.; Dianat, B.; Bertacchini, A.; Manfredi, A.; Sebastiani, M.; Modena, M.; Pancaldi, F. Pre-Processing Techniques to Enhance the Classification of Lung Sounds Based on Deep Learning. Biomed. Signal Process. Control 2024, 92, 106009. [Google Scholar] [CrossRef] [Scilit]
  46. Goutama, D.S.; Iskandar, A.A.; Rusyadi, R. Lung Sound Denoising with Adaptive Noise Cancellation of Heart Sounds on a Raspberry Pi-Powered Stethoscope. In Proceedings of the 2024 5th International Conference on Biomedical Engineering (IBIOMED), Bali, Indonesia, 23–25 October 2024; IEEE: New York, NY, USA, 2024; pp. 34–39. [Google Scholar] [CrossRef] [Scilit]
  47. Grooby, E.; He, J.; Kiewsky, J.; Fattahi, D.; Zhou, L.; King, A.; Ramanathan, A.; Malhotra, A.; Dumont, G.A.; Marzbanrad, F. Neonatal heart and lung sound quality assessment for robust heart and breathing rate estimation for telehealth applications. IEEE J. Biomed. Health Inform. 2021, 25, 4255–4266. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Choi, Y.; Choi, H.; Lee, H.; Lee, S.; Lee, H. Lightweight Skip Connections with Efficient Feature Stacking for Respiratory Sound Classification. IEEE Access 2022, 10, 53027–53042. [Google Scholar] [CrossRef] [Scilit]
  49. Ohmshankar, S.; Sudhagar, G. Lung Sound Classification via Improved Deep Architecture with Transform and Spectral Feature Set. In Proceedings of the 2024 IEEE International Conference on Information Technology, Electronics and Intelligent Communication Systems (ICITEICS), Bangalore, India, 28–29 June 2024; IEEE: New York, NY, USA, 2024; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
  50. Dar, J.A.; Srivastava, K.K.; Lone, S.A. Spectral Features and Optimal Hierarchical Attention Networks for Pulmonary Abnormality Detection from Respiratory Sound Signals. Biomed. Signal Process. Control 2022, 78, 103905. [Google Scholar] [CrossRef] [Scilit]
  51. Gupta, S.; Agrawal, M.; Deepak, D. Gammatonegram-Based Triple Classification of Lung Sounds Using Deep Convolutional Neural Network with Transfer Learning. Biomed. Signal Process. Control 2021, 70, 102947. [Google Scholar] [CrossRef] [Scilit]
  52. Shi, L.; Zhang, J.; Yang, B.; Gao, Y. Lung Sound Recognition Method Based on Multi-Resolution Interleaved Net and Time-Frequency Feature Enhancement. IEEE J. Biomed. Health Inform. 2023, 27, 4768–4779. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Roy, A.; Satija, U. Effect of auscultation hindering noises on detection of adventitious respiratory sounds using pretrained audio neural nets: A comprehensive study. IEEE Trans. Instrum. Meas. 2025, 74, 1–8. [Google Scholar] [CrossRef] [Scilit]
  54. Tzeng, J.-T.; Li, J.-L.; Chen, H.-Y.; Huang, C.-H.; Chen, C.-H.; Fan, C.-Y.; Huang, E.P.-C.; Lee, C.-C. Improving the robustness and clinical applicability of automatic respiratory sound classification using deep learning-based audio enhancement: Algorithm development and validation. JMIR AI 2025, 4, e67239. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Pham, L.; Phan, H.; Palaniappan, R.; Mertins, A.; McLoughlin, I. CNN-MoE-Based Framework for Classification of Respiratory Anomalies and Lung Disease Detection. IEEE J. Biomed. Health Inform. 2021, 25, 2938–2947. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Wang, F.; Yuan, X.; Liu, Y.; Lam, C.-T. LungNeXt: A Novel Lightweight Network Utilizing Enhanced Mel-Spectrogram for Lung Sound Classification. J. King Saud Univ. Comput. Inf. Sci. 2024, 36, 102200. [Google Scholar] [CrossRef] [Scilit]
  57. Zhantleuova, A.K.; Makashev, Y.K.; Duzbayev, N.T. Optimizing MFCC Parameters for Breathing Phase Detection. Sensors 2025, 25, 5002. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  58. Wu, C.; Huang, D.; Tao, X.; Qiao, K.; Lu, H.; Wang, W. Intelligent Stethoscope using Full Self-Attention Mechanism for Abnormal Respiratory Sound Recognition. In Proceedings of the 2023 IEEE EMBS International Conference on Biomedical and Health Informatics (BHI), Pittsburgh, PA, USA, 15–18 October 2023; IEEE: New York, NY, USA, 2023; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  59. Castejón-Barrio, A.; Gallardo-Antolín, A. Leveraging unlabeled data for lung sound classification through self-supervised contrastive learning. Biomed. Signal Process. Control 2026, 112, 108477. [Google Scholar] [CrossRef] [Scilit]
  60. Roslan, I.K.B.; Ehara, F. Detection of Respiratory Diseases from Auscultated Sounds Using VGG16 with Data Augmentation. In Proceedings of the 2024 2nd International Conference on Computer Graphics and Image Processing (CGIP), Kyoto, Japan, 12–14 January 2024; IEEE: New York, NY, USA, 2024; pp. 133–138. [Google Scholar] [CrossRef] [Scilit]
  61. Wang, Z.; Wang, Y.; Sun, Z. Sensitivity analysis of data augmentation methods on performance of deep learning model for lung sounds classification. Sci. Rep. 2025, 15, 39268. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Mang, L.D.; González Martínez, F.D.; Martinez Muñoz, D.; García Galán, S.; Cortina, R. Classification of Adventitious Sounds Combining Cochleogram and Vision Transformers. Sensors 2024, 24, 682. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  63. Phettom, R.; Theera-Umpon, N.; Auephanwiriyakul, S. Automatic Identification of Abnormal Lung Sounds Using Time-Frequency Analysis and Convolutional Neural Network. In Proceedings of the 2023 15th International Conference on Information Technology and Electrical Engineering (ICITEE), Chiang Mai, Thailand, 26–27 October 2023; IEEE: New York, NY, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  64. Laasya, R.A.; Muthulakshmi, M. Breathalytics: Deep Learning-Based Prediction of Abnormality in Respiratory Sounds using Continuous Wavelet Transform. In Proceedings of the 2024 3rd International Conference on Artificial Intelligence for Internet of Things (AIIoT), Vellore, India, 3–4 May 2024; IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  65. Cansiz, B.; Kilinc, C.U.; Serbes, G. Tunable Q-Factor Wavelet Transform-Based Lung Signal Decomposition and Statistical Feature Extraction for Effective Lung Disease Classification. Comput. Biol. Med. 2024, 178, 108698. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  66. Suma, K.V.; Koppad, D.; Kumar, P.; Kantikar, N.A.; Ramesh, S. Multi-Task Learning for Lung Sound and Lung Disease Classification. SN Comput. Sci. 2025, 6, 51. [Google Scholar] [CrossRef] [Scilit]
  67. Tasar, B.; Yaman, O.; Tuncer, T. Accurate respiratory sound classification model based on piccolo pattern. Appl. Acoust. 2022, 188, 108589. [Google Scholar] [CrossRef] [Scilit]
  68. Koshta, V.; Singh, B.K.; Behera, A.K.; G, R.T. Fourier Decomposition-Based Automated Classification of Healthy, COPD, and Asthma Using Single-Channel Lung Sounds. IEEE Trans. Med. Robot. Bionics 2024, 6, 1270–1284. [Google Scholar] [CrossRef] [Scilit]
  69. Fraiwan, L.; Hassanin, O.; Fraiwan, M.; Khassawneh, B.; Ibnian, A.M.; Alkhodari, M. Automatic Identification of Respiratory Diseases from Stethoscopic Lung Sound Signals Using Ensemble Classifiers. Biocybern. Biomed. Eng. 2021, 41, 1–14. [Google Scholar] [CrossRef] [Scilit]
  70. Alqudah, A.M.; Qazan, S.; Obeidat, Y.M. Deep learning models for detecting respiratory pathologies from raw lung auscultation sounds. Soft Comput. 2022, 26, 13405–13429. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Stas, T.; Lauwers, E.; Ides, K.; Verhulst, S.; Delputte, P.; Steckel, J. Convolutional Neural Network for the Detection of Respiratory Crackles. IEEE Access 2024, 12, 147301–147309. [Google Scholar] [CrossRef] [Scilit]
  72. Dhavala, A.; Ahmed, A.; Periyasamy, R.; Joshi, D. An MFCC Features-Driven Subject-Independent Convolution Neural Network for Detection of Chronic and Non-Chronic Pulmonary Diseases. In Proceedings of the 2022 3rd International Conference for Emerging Technology (INCET), Belgaum, India, 27–29 May 2022; IEEE: New York, NY, USA, 2022; pp. 1–9. [Google Scholar] [CrossRef] [Scilit]
  73. Tong, F.; Liu, L.; Xie, X.; Hong, Q.; Li, L. Respiratory Sound Classification: From Fluid-Solid Coupling Analysis to Feature-Band Attention. IEEE Access 2022, 10, 22018–22031. [Google Scholar] [CrossRef] [Scilit]
  74. Chanane, H.; Bahoura, M. Convolutional Neural Network-Based Model for Lung Sounds Classification. In Proceedings of the 2021 IEEE International Midwest Symposium on Circuits and Systems (MWSCAS), Lansing, MI, USA, 9–11 August 2021; IEEE: New York, NY, USA, 2021; pp. 555–558. [Google Scholar] [CrossRef] [Scilit]
  75. Ma, W.-B.; Deng, X.-Y.; Yang, Y.; Fang, W.-C. An Effective Lung Sound Classification System for Respiratory Disease Diagnosis Using DenseNet CNN Model with Sound Pre-processing Engine. In Proceedings of the 2022 IEEE Biomedical Circuits and Systems Conference (BioCAS), Taipei, Taiwan, 13–15 October 2022; IEEE: New York, NY, USA, 2022; pp. 218–222. [Google Scholar] [CrossRef] [Scilit]
  76. Petmezas, G.; Cheimariotis, G.-A.; Stefanopoulos, L.; Rocha, B.; Paiva, R.P.; Katsaggelos, A.K.; Maglaveras, N. Automated Lung Sound Classification Using a Hybrid CNN-LSTM Network and Focal Loss Function. Sensors 2022, 22, 1232. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  77. Hakki, L.; Serbes, G. Wheeze Events Detection Using Convolutional Recurrent Neural Network. In Proceedings of the 2023 Innovations in Intelligent Systems and Applications Conference (ASYU), Sivas, Turkiye, 11–13 October 2023; IEEE: New York, NY, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  78. Papadakis, C.; Rocha, L.M.G.; Catthoor, F.; Helleputte, N.V.; Biswas, D. AusculNET: A Deep Learning Framework for Adventitious Lung Sounds Classification. In Proceedings of the 2023 30th IEEE International Conference on Electronics, Circuits and Systems (ICECS), Istanbul, Turkey, 4–7 December 2023; IEEE: New York, NY, USA, 2023; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  79. Nguyen, T.; Pernkopf, F. Crackle Detection in Lung Sounds Using Transfer Learning and Multi-Input Convolutional Neural Networks. In Proceedings of the 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), Virtual Conference, 1–5 November 2021; IEEE: New York, NY, USA, 2021; pp. 80–83. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  80. Roy, A.; Satija, U. A Novel Melspectrogram Snippet Representation Learning Framework for Severity Detection of Chronic Obstructive Pulmonary Diseases. IEEE Trans. Instrum. Meas. 2023, 72, 4003311. [Google Scholar] [CrossRef] [Scilit]
  81. Hu, J.; Leow, C.S.; Tao, S.; Goh, W.L.; Gao, Y. Supervised contrastive learning framework and hardware implementation of learned ResNet for real-time respiratory sound classification. IEEE Trans. Biomed. Circuits Syst. 2025, 19, 185–195. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  82. Liu, B.; Wen, Z.; Zhu, H.; Lai, J.; Wu, J.; Ping, H.; Liu, W.; Yu, G.; Zhang, J.; Liu, Z.; et al. Energy-Efficient Intelligent Pulmonary Auscultation for Post COVID-19 Era Wearable Monitoring Enabled by Two-Stage Hybrid Neural Network. In Proceedings of the 2022 IEEE International Symposium on Circuits and Systems (ISCAS), Austin, TX, USA, 27 May–1 June 2022; IEEE: New York, NY, USA, 2022; pp. 2220–2224. [Google Scholar] [CrossRef] [Scilit]
  83. Park, J.; Jeong, C.; Choi, Y.; Hong, H.-K.; Jo, Y. Lung Sound Classification Model for On-Device AI. Appl. Sci. 2025, 15, 9361. [Google Scholar] [CrossRef] [Scilit]
  84. Ehtesham, A.; Kumar, S.; Singh, A.; Khoei, T.T. Pediatric Asthma Detection with Googleś HeAR Model: An AI-Driven Respiratory Sound Classifier. In Proceedings of the 2025 IEEE World AI IoT Congress (AIIoT), Seattle, WA, USA, 28–30 May 2025; IEEE: New York, NY, USA, 2025; pp. 103–109. [Google Scholar] [CrossRef] [Scilit]
  85. Liao, Z.; Luo, G.; Yan, H.; Wang, J.; Zhang, S.; Wu, J.; Yu, R.; Xu, L.; Chen, L. LungListener: Bootstrapping a Large-Scale Audio-Language Model for Lung Sound Classification and Analysis. In Proceedings of the 2025 IEEE Smart World Congress (SWC), Calgary, AB, Canada, 18–22 August 2025; IEEE: New York, NY, USA, 2025; pp. 987–992. [Google Scholar] [CrossRef] [Scilit]
  86. Huang, D.; Wang, L.; Lu, H.; Wang, W. A Contrastive Embedding-Based Domain Adaptation Method for Lung Sound Recognition in Children with Community-Acquired Pneumonia. In Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; IEEE: New York, NY, USA, 2023; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  87. Oshima, R.; Kamiya, T.; Kido, S. Improving Respiratory Sound Classification via Multimodal Learning with Patient Metadata and Robustness-Oriented Rotation Augmentation. In Proceedings of the 2025 25th International Conference on Control, Automation and Systems (ICCAS), Incheon, Republic of Korea, 4–7 November 2025; IEEE: New York, NY, USA, 2025; pp. 1262–1266. [Google Scholar] [CrossRef] [Scilit]
  88. Shayetreen, L.; Anani, A.; Tazin, T.M.; Marzan, U.; Afsar, S.R.; Noor, J. A Comprehensive Respiratory Evaluation: Incorporating Lung Sound and Disease Classification along with Spirometry Assessment. In Proceedings of the 2024 6th International Conference on Electrical Engineering and Information & Communication Technology (ICEEICT), Dhaka, Bangladesh, 2–4 May 2024; IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  89. Wanasinghe, T.; Bandara, S.; Madusanka, S.; Meedeniya, D.; Bandara, M.; Díez, I.D.L.T. Lung sound classification with multi-feature integration utilizing lightweight CNN model. IEEE Access 2024, 12, 21262–21276. [Google Scholar] [CrossRef] [Scilit]
  90. Majzoobi, F.; Khodabakhshi, M.B.; Jamasb, S.; Goudarzi, S. ConvLSNet: A lightweight architecture based on ConvLSTM model for the classification of pulmonary conditions using multichannel lung sound recordings. Artif. Intell. Med. 2024, 154, 102922. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  91. Abadade, Y.; Benamar, N.; Bagaa, M.; Chaoui, H. Empowering healthcare: TinyML for precise lung disease classification. Future Internet 2024, 16, 391. [Google Scholar] [CrossRef] [Scilit]
  92. Sreejith, R.; Ramasamy, R.K.; Mohd-Isa, W.-N.; Abdullah, J. Enhanced Lung Disease Classification Using CALMNet: A Hybrid CNN-LSTM-TimeDistributed Model for Respiratory Sound Analysis. IEEE Access 2025, 13, 135053–135073. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.