Next Article in Journal
From Profiles to Promising Paths: A Semantic Group Recommender for Novel Academic Topic Discovery
Previous Article in Journal
Cryptography-Based Security Authentication and Privacy Preservation of Cyber-Physical Power Systems: An Overview
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Phase-Coherent Four-Stage Pipeline for the Dereverberation of Quránic Recitation

1
College of Arts and Sciences, University of Nizwa, Nizwa 616, Oman
2
EuroMov Digital Health in Motion, Université de Montpellier, IMT Mines Alès, 30100 Alès, France
*
Author to whom correspondence should be addressed.
Information 2026, 17(7), 714; https://doi.org/10.3390/info17070714
Submission received: 23 June 2026 / Revised: 17 July 2026 / Accepted: 21 July 2026 / Published: 22 July 2026
(This article belongs to the Section Information Applications)

Abstract

The accuracy of spectro-temporal features for Makhaarij al-Huroof and Sifaat distinguishes between the ten canonical Qiraát recitation styles of the Holy Quran. However, real-world room reverberations blur formant contours and corrupt inter-word energies, thus making Qiraat discrimination difficult. The current dereverberation methods were designed to work under ordinary speech conditions and are not capable of preserving phonetic qualities for domain-specific purposes. This paper introduces a four-step, phase-consistent signal-processing approach prioritizing phonetic preservation over direct reverberation suppression. The four steps are: (1) adaptive noise-floor attenuation; (2) soft-voice activity detection using power-law boundary decay; (3) application-specific spectral contour adjustment from clean Quranic reference audio; and (4) Griffin–Lim algorithm-based phase correction. A total of 48 real-world room recordings were utilized for the evaluation of this approach based on Energy Ratio (ER), Spectral Contrast (SC), and Spectral Contour Stability (SCS)—measures specific to the Quran audio domain—alongside conventional speech-quality metrics. The proposed approach yielded the highest scores in three of seven metrics, namely SC (+40.11), SCS (+822.94), and PESQ (+1.251), alongside the second-highest Energy Ratio (+19.58 dB), while being superior to Spectral Subtraction, Wiener Filtering, and WPE Dereverberation approaches. Moreover, the perceptual enhancement was verified in a synthetic controlled experiment where the proposed approach scored an improved PESQ metric (+2.495; SNR −1.874 dB). The results illustrate the fact that an optimization for general-purpose metrics does not necessarily ensure phonetic preservation required for specific classification.

Graphical Abstract

1. Introduction

The Holy Quran was revealed in ten different Qira’aat, each signifying a particular method of reciting which is based on phonetic variations that have been passed down in academic chains [1,2]. Automatic detection of Qira’aat is a difficult task because the phonetic characteristics that differentiate between various Qira’aat are extremely vulnerable to acoustic distortion [1,2].
In contemporary Qira’at classification approaches, mel-frequency cepstral coefficients (MFCCs) [3] along with deep learning [1,4,5] have been utilized successfully for studio recordings, but these techniques fail miserably when applied to real-room audio recordings, where the RT60 (reverberation time) usually varies between 0.3 and 0.8 s [6]. Reverberations in real-room audio lead to spectral blurring that goes beyond phonemes, fills the spaces between words, blurs formant transitions, and reduces the high frequencies of fricative consonants like /s/, /sh/, and /t/ [7].
The three established dereverberation algorithms, Spectral Subtraction, ref. [7] Wiener Filtering ref. [8] and Weighted Prediction Error (WPE), ref. [9] were tested on general speech corpora. However, applying them to Quraan recordings, the algorithms are able to remove reverberation energy at the cost of losing the phonetic features required for Qira’at identification. In fact, WPE provides the lowest Spectral Contour Stability value among all evaluated approaches and is inferior even to the unprocessed reverberant speech.
The research proposes a solution to the described dilemma in a pipeline of four stages where maintaining phonetic features becomes the main goal of designing. As the ablation study in Section 4.6 demonstrates, no single stage accounts for the full improvement: adaptive noise-floor subtraction contributes most to the three magnitude-based Quranic metrics, while explicit phase recovery using the Griffin–Lim algorithm [9], applied after the magnitude-domain processing, eliminates metallic artifacts and restores perceptual naturalness.

Research Contributions

  • A four-stage, phase-coherent audio processing pipeline (see Figure 1) specifically designed for Quranic audio preprocessing.
  • A domain-specific spectral shaping gain curve, derived from the frequency energy profile of clean Quranic reference audio, that preserves vowel formants in the 1000–2000 Hz band and fricative energy above 4000 Hz.
  • The first use of Griffin–Lim phase reconstruction for Quranic voice enhancement shows that, after magnitude-domain processing, phase coherence is the main factor influencing perception naturalness.
  • Spectral Contour Stability (SCS), Spectral Contrast (SC), and Energy Ratio (ER) are three Quran-specific evaluation measures that reflect phoneme-level signal characteristics pertinent to Qira’at classification.
  • Results from a controlled synthetic experiment and 48 actual room recordings demonstrate a steady improvement over three well-established baseline techniques.

2. Related Work

2.1. Quranic Recitation Classification

Research into automated Quranic recitation analysis [10] has steadily advanced throughout the last decade from classical signal processing to end-to-end deep-learning approaches. Alkhateeb [1] identified the use of MFCC features for distinguishing between reciters and concluded that cepstral features [11] contain enough discriminant properties to be useful for supervised recitation identification. Nevertheless, it must be pointed out that this research was based on audio data recorded in a studio. Khan et al. [12]. Following suit, they expanded on this research line by analyzing ensemble classifiers, which helped in enhancing robustness for different reciters but kept the entire MFCC process susceptible to spectral smearing due to room acoustics. Ghori et al. [4] utilized deep neural networks to model Quranic recitation acoustically and succeeded in making significant improvements in word error rates in controlled environments.
Recent approaches have been more focused on end-to-end systems. Al-Harere and Al-Jallad [2] introduced a CNN-Bidirectional GRU encoder using connectionist temporal classification (CTC) training, achieving a new state-of-the-art on the Ar-DAD benchmark at 8.34% word error rate. Al-Issa et al. [13] demonstrated that using Whisper instead of DeepSpeech as the recognition engine results in a significant boost in performance for different genders, ages, and competencies. However, no reverberation removal preprocessing steps are involved in either system, and both will be negatively affected by the acoustic environment of an actual room. Recent surveys [5] indicate room acoustics as an unsolved problem in Quranic speech processing, which is a key motivation for this research.
This problem is exacerbated by the use of MFCC-based features. As MFCC analysis depends on the STFT magnitude spectrum, reverberation-caused degradation of either Spectral Contour Stability or Spectral Contrast will cause a decrease in the distinguishability of the feature vectors. Al-Ayyoub et al. [14] have shown experimentally that systems based on deep-learning methods, which rely on training on Quranic material, are sensitive to the quality of the input signals.

2.2. Classical Speech Dereverberation

Classical dereverberation methods operate on the STFT magnitude without explicit phase modeling. Spectral Subtraction [7] estimates a stationary noise floor from initial silence frames and subtracts it from the magnitude spectrum. Its main limitation is the generation of musical noise and non-stationary spectral artifacts that can be misidentified as phonemic content, particularly in the fricative frequency regions critical for Qira’at analysis.
The Wiener Filter [8] minimizes the mean-square error between the estimated and clean spectra, providing a theoretically optimal linear solution under stationary noise assumptions. However, it does not account for the time-varying structure of late reverberation.
The Weighted Prediction Error (WPE) method [9], extended in further work [15,16], models late reverberation as a linearly predictable component in the STFT domain and suppresses it with a multi-tap linear predictor of order K and delay Δ . WPE performs well on standard benchmarks such as the REVERB Challenge, but this study provides empirical evidence that it causes severe Spectral Contour Stability deterioration in Quranic speech (SCS = +148.29 vs. +822.94 for the proposed method) and collapses on Energy Ratio in 20 out of 48 real-room recordings (ER = 0.000 dB). This failure reflects a fundamental mismatch: WPE assumes stationary late reverberation, whereas real rooms exhibit time-varying impulse responses.

2.3. Deep-Learning and Diffusion-Based Dereverberation

Constraints inherent to classical approaches have motivated the exploration of data-driven methodologies. Wang et al. [17] performed a systematic analysis of DNNs for the task of speech dereverberation and showed that employing larger window sizes and transform-layer components leads to reliable performance gains. However, these networks were only developed using generic English datasets, disregarding any phonological properties unique to Quranic Arabic. Recent advancements in the field of diffusion-based models have enabled them to achieve state-of-the-art performance on speech enhancement tasks. Richter et al. [18] suggested a new approach named SGMSE, which is based on score-based diffusion and operates on the short-term Fourier transformation plane, demonstrating excellent quality results under various noise conditions. Lemercier et al. [19] designed a hybrid framework called StoRM, which is based on a combination of a predictive regression component and a diffusion-based module for post-refinement with fewer sampling iterations but with a comparable level of naturalness. However, there are three main challenges that hinder applying such methods directly for Qira’at preprocessing purposes. First, such systems need a large number of pairs for training [20]. Second, systems trained on general speech fail to generalize to Quranic Arabic since there is a difference in the articulation point and phonetic characteristics of the latter, which are normalized by formants lying outside the training data space. Third, inference using the diffusion approach is computationally expensive, rendering real-time applications impractical [21].
We acknowledge, however, that these limitations are not absolute. Fine-tuning a pretrained diffusion or SGMSE model on a small paired corpus of clean and reverberant Qur’anic recitation could, in principle, partially close the domain gap, adapting the learned prior toward the specific formant structure and articulation of Qur’anic Arabic without the full data requirement of training from scratch. We did not pursue this direction here for two reasons. First, no paired clean/reverberant Qur’anic corpus of sufficient size currently exists, and constructing one is a substantial undertaking in its own right. Second, the present work deliberately targets a training-free pipeline that can be deployed without any domain-specific data collection, which is the operational constraint under which Qur’anic preprocessing typically occurs. A controlled fine-tuning experiment, quantifying how far a modest Qur’anic paired set can recover the phonetic fidelity that general-corpus models lose, is a well-defined and worthwhile direction that we leave to future work.
In summary, as illustrated in Table 1, these findings highlight that no existing system, whether classical or based on deep learning, has been engineered to capture the domain-specific phonetic properties essential for Qira’at classification.

2.4. Griffin–Lim Phase Reconstruction

The Griffin–Lim Algorithm [9], first suggested in 1984, deals with the phase inconsistency that results when the magnitude of the STFT is changed without altering the phase. The algorithm switches back and forth between STFT and inverse STFT to minimize the spectral difference between the target and estimated magnitude. Later on, momentum-based acceleration was introduced by Perraudin et al. [22], resulting in an improved speed through reduced iteration counts. Griffin–Lim is popularly utilized in voice synthesis and spectrogram inversion [23] but has not been used for Quranic speech enhancement before. perceptual naturalness in Quranic dereverberation, raising informal listening ratings from 6.0 to 9.5, although the ablation study in Section 4.6 shows that this stage leaves the magnitude-domain Quran-Specific metrics largely unaffected, with adaptive noise-floor subtraction (Stage 2) driving those gains instead.

2.5. Research Gap

There is a persistent discrepancy within the literature. Classical models were developed and tested based on generic speech, and their metrics, SNR, PESQ, and STOI, do not capture the phonetic properties that define Qira’at styles. DL and diffusion approaches obtain higher scores on these metrics but suffer from domain mismatch and lack of training data. On the other hand, there have been architectural advances in the Qira’at style recognition literature without dealing with real-room acoustics issues.
There is no previous study that attempts to develop a training-free dereverberation approach that is optimized for preserving formant trajectory, fricative consonant energy, and inter-word silence patterns, the distinguishing factors of Qira’at style. This study seeks to fill this gap.
The originality of this work lies in three specific choices that distinguish it from all prior dereverberation. First, where classical and deep methods optimize a global distortion criterion (minimizing error against a reference, or maximizing a perceptual score), the proposed pipeline optimizes for the preservation of the acoustic cues that carry Qira’at identity: the formant trajectories that realize Makhaarij al-Huroof, the high-frequency energy of fricative Sifaat, and the inter-word silence that marks recitation boundaries. This is a different objective, not a different tuning of the same objective. Second, the pipeline is training-free by design, which sidesteps the domain mismatch that we show empirically defeats a state-of-the-art learned model (MetricGAN+) on Qur’anic Arabic. Third, phase coherence is preserved end-to-end by confining all magnitude modification to a single STFT block and restoring phase explicitly, rather than reconstructing the signal from a modified magnitude with inconsistent phase.
The link to downstream classification is direct and is quantified in Section 4.9. Because MFCC-based Qira’at classifiers derive their discriminative power from exactly the spectral properties the pipeline preserves, restoring those properties restores classification accuracy: on a labeled multi-style set, accuracy degraded by reverberation (96.1% to 94.1%) is fully recovered to 96.1% after processing, whereas methods that optimize general-purpose metrics do not preserve the cues on which that recovery depends. This is the concrete sense in which the proposed method improves Qira’at classification while competing dereverberation methods, despite comparable or better scores on standard metrics, do not.

3. Materials and Methods

The pipeline filters a mono reverberated audio file of the Quran through four successive steps, which work inside one STFT block. Such a mechanism provides phase consistency, meaning that the modification is made on the magnitude in Stages 1 to 3, while the phase estimation is made in Stage 4. Figure 2a shows a spectrogram of an example reverberant recording that has been enhanced to the one in Figure 2b. One can readily observe the continuous smeared energy in the enhanced spectrogram as against the clear black silence regions in the former. The pipeline works at a sampling frequency of 16,000 Hz, an FFT size of 1024, a hop length of 256, and a Hann windowing function.

3.1. Stage 1: Domain-Adapted Spectral Shaping

Room reverberation concentrates energy below 500 Hz, producing the characteristic orange-yellow smear in spectrograms that obscures inter-word silences. A frequency-dependent gain function g c ( f ) is applied to the STFT magnitude M ( f , t ) :
M s ( f , t ) = g c ( f ) × M ( f , t )
The gain curve g c ( f ) , in Figure 3, was derived from the frequency energy distribution of a clean studio recording of the Quran (Table 2). The 1000–2000 Hz band is preserved at 75% gain because it contains the vowel formants that distinguish Qira’at styles. The 4000–16,000 Hz band is preserved at 35% to maintain fricative consonant energy. These parameters were fixed for all 48 recordings and were not tuned per file.

3.2. Stage 2: Adaptive Noise-Floor Subtraction

After spectral shaping, residual noise is estimated from the quietest 15% of frames in the STFT block, corresponding to silence between words. The cleaned magnitude is computed as:
M c ( f , t ) = max M s ( f , t ) α · N ( f ) , β · M s ( f , t )
where N ( f ) is the mean of M s ( f , t ) calculated over the least noisy 15% of frames. The subtractive weighting factor α = 2.0 was chosen through experimentation: weights higher than 3.0 produce metallic artifacts, whereas weights lower than 1.5 allow too much noise. The spectral floor β = 0.001 ensures that no frequency bin is eliminated, maintaining the proper formant structure required for MFCC calculation.

3.3. Stage 3: Soft Voice Activity Detection with Power-Law Decay

While maintaining word boundaries, a gentle gate lessens the size of silent frames. The soft gate applies a smooth, energy-proportional gain as opposed to a hard gate, which creates click artifacts:
g soft ( t ) = g min + ( 1 g min ) × e norm ( t )
where e norm ( t ) is the normalized energy of the frame in the range [ 0 , 1 ] , and g min = 0.10 , which prevents whispered consonants from being completely eliminated. At the onset and offset of words, an exponential decay function defines the
g decay ( t + d ) = max g decay ( t + d ) , 1 d D p
where D = 45 frames (720 ms) and p = 2.0 . A gap-fill duration of 220 ms prevents fragmentation caused by intra-word micro-pauses, ensuring that Madd and Shaddah marks are not cut off. The final gate is the Hadamard product of g soft and g decay , smoothed with an 11-frame window. A comparison between these gate functions is illustrated in Figure 4.

3.4. Stage 4: Griffin–Lim Phase Reconstruction

Any alteration of the amplitude values during Stages 1 to 3, but without updating the corresponding phase values, results in a magnitude/phase inconsistency that gives rise to metallic artifacts and affects speech-related phase-dependent features. To overcome such an issue, the Griffin–Lim [9] algorithm proceeds as follows:
z n = M ( f ) × exp j · STFT ( iSTFT ( z n 1 ) )
with M ( f ) being the output magnitude of Stage 3 and z 0 initialized based on the phase of the original signal. The algorithm converged after ten iterations, representing a compromise between convergence and cost. The addition of Stage 4 improved the subjective ratings of the output signal from 6.0 out of 10 to 9.5, validating the hypothesis that perceived naturalness is determined by phase coherence. An example Griffin-Lim phase reconstruction and its convergence is graphically illustrated in Figure 5.

3.5. Algorithm Parameters

Table 3 summarises all pipeline parameters.

4. Results

4.1. Dataset

Two complementary datasets were used. The first consists of 48 real-room recordings of Quranic recitation (approximately 9.56 min total), collected from multiple reciters in natural indoor environments with varying reverberation conditions, ranging from small rooms to larger spaces such as prayer halls. Reverberation levels were not controlled, ensuring that the evaluation reflects realistic acoustic variability; all recordings were resampled to 16,000 Hz. The second dataset was a single studio-quality Quranic recording artificially reverberated using the Image Source Method [24] via the pyroomacoustics package [23], with room dimensions 6 × 5 × 3 m, RT60 = 0.4 s, source at [2.5, 3.5, 1.5] m, and microphone at [3.5, 2.0, 1.5] m. This controlled condition allows reference-based metrics (PESQ, STOI, SNR) to be computed against the known clean signal.

4.2. Baseline Methods

Four methods were compared under identical Python 3.12-based experimental conditions:
  • Spectral Subtraction [7]: noise estimated from the first 20 STFT frames; α = 6.0 ; window 512; hop 128.
  • Wiener Filter [8]: MMSE-optimal gain computed from an initial silence segment; same STFT parameters.
  • WPE Dereverberation [9]: nara_wpe package; K = 10 taps; Δ = 3 ; 5 iterations; window 512.
  • Reverberant Only: unprocessed signal serving as the degradation baseline.
  • MetricGAN+ [25]: GAN-based deep-learning speech enhancement network pretrained on VoiceBank-DEMAND dataset. No fine-tuning was conducted since, in real-world applications, domain-specific paired data for Qur’anic texts will not be available.

4.3. Evaluation Metrics

4.3.1. Standard Speech Quality Metrics

  • SNR (dB): signal-to-noise ratio relative to a clean reference (synthetic experiment only).
  • PESQ (ITU-T P.862) [26]: perceptual evaluation of speech quality; scale 1.0–4.5.
  • STOI [27]: short-time objective intelligibility; scale 0–1.
  • SI-SDR (dB): scale-invariant signal-to-distortion ratio.

4.3.2. Quran-Specific Metrics

Three metrics were designed to capture phonetic signal properties relevant to Qira’at classification:
  • Energy Ratio (ER, dB) is the ratio, separated by the 30th percentile threshold, between the average energy in active frames and the average energy in silent frames. Clear word boundaries are indicated by a high ER value, which also makes automatic segmentation easier.
  • The average Spectral Contrast over six frequency sub-bands (librosa) is represented by Spectral Contrast (SC). Higher values correlate with phoneme separability and show improved discrimination between harmonic and noise components.
  • Spectral Contour Stability (SCS): SCS quantifies how decisively the spectral envelope moves across phonemes relative to its per-frame spread. Reverberation flattens the temporal trajectory of the spectral centroid, as late reflections fill inter-phoneme intervals with quasi-stationary energy; SCS captures this as the temporal dispersion of the spectral centroid, normalized by the mean spectral bandwidth:
    SCS = σ c ( t ) μ b ( t ) × 10 3 ,
    where c ( t ) and b ( t ) are the spectral centroid and spectral bandwidth of frame t, both computed with librosa [28]. A signal whose envelope moves decisively between phonemes while remaining spectrally concentrated yields a high SCS; a reverberant signal, whose envelope is both static and diffuse, yields a low one. The metric is parameter-free, introducing no model order or detection threshold, so its value cannot be inflated by a favorable choice of settings. While SCS is introduced in this work as a composite descriptor, both of its constituents, the spectral centroid and the spectral bandwidth, are long-established features in audio signal analysis [28,29].

4.4. Controlled Experiment: Synthetic Reverberation

Table 4 and Figure 6 report results on the synthetic dataset. The proposed algorithm achieves the best SNR (−1.874 dB, an improvement of +1.365 dB over the reverberant baseline) and the best PESQ (+2.495, an improvement of +0.308). WPE achieves the highest Energy Ratio in this controlled condition (+38.193 dB) but at the cost of the worst PESQ (+1.116) and STOI (+0.174), consistent with previous findings that WPE generates phase artifacts that alter spectral signal properties [30].

4.5. Real-Room Evaluation: 48 Recordings

As can be observed from Table 5, as well as the corresponding radar chart in Figure 7, the proposed method achieves the best result in three of seven metrics: PESQ, Spectral Contrast, and Spectral Contour Stability. It leads on two of the three Qur’anic-specific measures (Spectral Contrast and SCS) and attains the second-highest energy ratio, where MetricGAN+ scores higher by aggressively suppressing low-energy regions at the cost of spectral structure.

4.6. Ablation Study

To quantify the individual contribution of each pipeline stage, an ablation study was conducted across all 48 recordings. Each stage was removed in isolation while the remaining three were retained, and the three Quran-Specific metrics were re-evaluated (Table 6).
The four steps work additively rather than duplicatively, each addressing a separate element of signal quality. Stage 2 (Adaptive Noise-Floor Subtraction) makes the greatest contribution to the spectral measures, providing a gain of + 289.4 for Spectral Contour Stability and a gain of + 9.82 for Spectral Contrast; without it, Spectral Contour Stability is reduced from 822.9 to 533.5, indicating its importance for reconstructing the spectra necessary for phoneme recognition. Stage 1 (Domain-Adapted Spectral Shaping) dominates the Energy Ratio enhancement ( + 5.97 dB); this is due to its purpose of removing low-frequency reverberant energy present between the words.
The contribution of Stage 4 (Phase Reconstruction) towards the three magnitude-related metrics ( Δ SCS = + 1.02 , Δ SC = + 0.19 ) is almost negligible. This comes naturally rather than being an anomaly because the magnitude-related metrics have been calculated using spectral magnitude, which has been determined till now through Stages 1–3, and Stage 4 works on phase only. The contribution of this stage is thus perceptual, i.e., it takes care of the metallic artifacts generated due to modification in magnitude, resulting in an increase in perceptual quality score from 6.0 to 9.5 on a scale of 10 and an improvement in PESQ score from 2.187 to 2.495 in the controlled experiment (Section 4.4).
Most importantly, none of the ablated configurations beat the whole pipeline on the set of metric values, and removing any stage is detrimental to at least one of the metrics compared to the complete system. It proves the necessity of all four stages: the magnitude shaping stages (1–3) determine the spectral characteristics used by the Quranic metrics, and the phase reconstruction stage (4) guarantees the perceptual naturalness not measured with magnitudes.

4.7. Energy Ratio Threshold Sensitivity

Reviewer feedback correctly noted that the Energy Ratio metric depends on a percentile threshold for separating speech from silence frames. To assess robustness, the silence/active percentile boundary was varied across five settings, from an aggressive 20th/80th split to a conservative 40th/60th split (Table 7).
As expected for any relative energy measure, the absolute energy ratio value scales with the chosen percentile: tighter thresholds (20th/80th) compare only the quietest and loudest frames, inflating the ratio, while wider thresholds (40th/60th) include more transitional frames and compress it. The operative property, however, is not the absolute value but the consistency of the advantage the proposed pipeline maintains over the reverberant baseline. Across every threshold, this advantage remains stable at + 6.6 to + 7.7 dB (mean + 7.04 dB), confirming that the measured improvement reflects genuine dereverberation rather than an artifact of threshold selection. Since the metric is applied comparatively rather than as an absolute quantity, this ranking-invariance is the relevant robustness criterion.

4.8. Feature Discriminability Experiment

In order to ensure that increases in Spectral Contrast and Energy Ratio resulted in better-separable feature spaces, an analysis of feature discriminability was performed. The MFCC features (78 dimensions: 13 MFCC coefficients with delta and delta-delta, mean and standard deviation) were computed for all 48 audio clips analyzed with each approach. Audio clips were split into two classes depending on the level of reverberation, calculated using the late-to-early energy ratio. The accuracy for a k-NN classifier ( k = 3 ) with leave-one-out cross-validation was found to be 77.1% with features taken from the output of the proposed algorithm, as opposed to 75.0% with no feature processing and 70.8% with WPE processing (Table 8). Intra-class tightness (mean distance to the centroid of the class) was reduced from 8.010 to 7.576, indicating that our approach not only results in tighter clusters but also increases their discriminative power. The corresponding feature-space separation is visualized in Figure 8.
Basically, WPE achieves the tightest intra-class compactness yet yields the lowest classification accuracy, indicating that it collapses feature diversity and degrades the discriminative information rather than preserving it. Meanwhile, the proposed algorithm does this nice tradeoff, finding the optimal balance between compactness and separability.

4.9. Qira’at Classification Validation

To assess whether the pipeline’s acoustic improvements carry over to downstream recognition, we evaluated Qira’at classification on a labeled set of 51 recordings spanning three reciters (Hafs: Abdul Basit, Al-Husary; Warsh: Al-Dosary), sourced from EveryAyah.com. Controlled reverberation (RT60 = 0.4 s, room 6 × 5 × 3 m) was applied via the Image Source Method (pyroomacoustics). MFCC features (78-dimensional) were classified using an SVM (RBF) under leave-one-out cross-validation, in three conditions: clean, reverberant, and pipeline-enhanced.
Under this controlled simulation, classification accuracy on the reverberant signal was restored to its clean-condition level after processing (Table 9). We note that the absolute recovery depends on the specific simulated room geometry; the robust and configuration-independent finding is that the pipeline markedly improves the three domain-specific acoustic metrics (Energy Ratio, Spectral Contrast, Spectral Contour Stability) that underlie phonetic discriminability, while preserving high classification accuracy. A broader validation across multiple room geometries and reciter-independent recitation styles is identified as future work.

4.10. Statistical Significance

Table 10 reports Wilcoxon signed-rank tests ( n = 48 ). Improvements in all three Quran-Specific metrics are statistically significant ( p < 0.001 ). PESQ also reaches significance ( p < 0.05 ) in this expanded evaluation, indicating that the perceptual advantage of the pipeline becomes clearer at larger evaluation scales.

5. Discussion

The three metrics outlined in Section 4.3.2 demonstrated interesting results, as discussed below.
Energy Ratio (+19.579 dB). The proposed method produces inter-word silence segments that are on average 19.579 dB quieter than active speech, a gain of +7.157 dB over Spectral Subtraction (+12.422 dB). WPE collapses on 20 of 48 recordings, with ER dropping to 0.000 dB, confirming that its linear prediction model fails under the time-varying reverberation of real rooms.
Spectral Contrast (+40.109). The proposed algorithm achieves the highest SC of all methods, indicating superior separation between harmonic and non-harmonic frequency components. Higher SC values correlate with better phoneme separability for downstream classifiers. WPE produces SC = +16.784, which is lower than the unprocessed reverberant signal (+29.748).
Spectral Contour Stability (+822.942). The proposed method attains an SCS 5.55× higher than WPE (+148.293) and 1.62× higher than the Wiener Filter (+508.441). The advantage is driven primarily by the numerator: dereverberation restores the temporal excursion of the spectral centroid, whose standard deviation rises from 742.8 (reverberant) to 1113.7, as suppressing late-reflection energy re-exposes the phoneme-to-phoneme movement of the spectral envelope. The denominator reinforces this, the mean spectral bandwidth contracting from 1565.4 to 1369.5 as the surviving energy concentrates. WPE moves the numerator the wrong way, collapsing centroid dispersion to 231.7 and flattening the spectral contour over time, which is why its SCS falls below that of the unprocessed signal. That this pattern holds across all 48 recordings, captured in uncontrolled rooms with different reciters, indicates that the domain-adapted gain curve generalizes rather than fitting a single acoustic condition.

5.1. Standard Metric Performance

Improvements in conventional metrics (SNR, PESQ, STOI) are generally smaller than those in the Quran-Specific metrics for two reasons. First, standard metrics cannot distinguish between modifications that preserve phonetic properties and those that destroy them. A high SNR achieved by suppressing high-frequency components carries no penalty in the SNR score yet sharply reduces SCS. Second, WPE achieves the best SNR (−0.400 dB) but simultaneously collapses on ER in 20 recordings, exposing an inconsistency between distance-to-reference metrics and task-relevant phonetic metrics.
Notably, the 48-file evaluation reversed the PESQ finding from the earlier 17-file experiment: the proposed algorithm now achieves the highest PESQ (+1.251), surpassing the original signal (+1.199) and all baselines, with statistical significance ( p < 0.05 ). This suggests that the perceptual advantage of the pipeline becomes more apparent at larger evaluation scales. MetricGAN+ Domain Mismatch. Although MetricGAN+ achieves state-of-the-art results on English speech corpora, its behavior on Qur’anic recitation reveals a clear domain mismatch. It improves the Energy Ratio (+22.851) but does so by aggressively suppressing low-energy regions, while simultaneously degrading Spectral Contrast (+27.000) below that of the unprocessed signal (+29.748) and leaving Spectral Contour Stability (+444.291) close to the reverberant baseline. In other words, the model trained on the English VoiceBank-DEMAND corpus optimizes a general enhancement objective at the expense of the spectral structure that carries Qira’at identity: the phonology of Qur’anic Arabic, characterized by Makhaarij al-Huroof and Sifaat, lies outside the model’s training distribution. The proposed training-free pipeline, by contrast, improves all three domain-specific metrics jointly, yielding SCS = +822.942, 1.85× that of MetricGAN+.
Comparison with State-of-the-Art Methods. Based on the Quarnic-specific metrics, the proposed pipeline was benchmarked against representative methods (see Figure 9) spanning three generations of dereverberation: classical spectral techniques (Spectral Subtraction, Wiener filtering), the linear-prediction–based WPE, and a recent deep-learning model, MetricGAN+, trained on the VoiceBank-DEMAND corpus. This spread lets us separate two questions: whether a method is effective in general, and whether it preserves the phonetic cues specific to Qur’anic recitation. The two do not coincide. WPE attains the best SNR on real recordings, and MetricGAN+ is state-of-the-art on English speech enhancement, yet both underperform on the domain-specific metrics that predict Qira’at discriminability. The proposed pipeline, by contrast, leads on three of the seven metrics, including two of the three Qur’anic-specific measures. The comparison therefore demonstrates effectiveness not by outperforming every method on every metric, which no single method does, but by being the only method that improves the acoustic properties on which downstream Qira’at classification depends, without requiring any domain-specific training data.
The relatively modest STOI (+0.107 vs. +0.117 for the original) reflects that STOI was designed for noisy rather than reverberant speech and penalizes the energy redistribution introduced by spectral shaping.

5.2. Sensitivity to Recording Conditions

These 48 real-room recordings have been recorded in varying indoor environments from small rooms to big prayer halls, by more than one reciter using unknown consumer-level equipment. This is intentional, as it tests the performance of the pipeline in realistic conditions, not idealized ones. This sets limits on our claims, which we acknowledge upfront.
We attempted to measure the reverberation time of each audio file with three techniques: a straightforward approach based on measuring the energy decay rate, local decay based on the peaks of the energy decay curve, and the Schroeder integration method from pyroomacoustics [23]. None of these yielded RT60 values that we are prepared to use. Blind RT60 measurement relies on a monotonic energy decay, such as that provided by an impulse response, while connected speech with silences on consumer hardware fails to provide one. We therefore have chosen not to give per-file RT60, and we take the lack of a valid figure as meaningful information: the audio is reverberant enough to foil blind RT60 estimation.
It is pertinent to the reviewer’s question about extremely reverberant conditions (RT60 > 0.8 s). As there is no parameter in the system tuned to a particular decay time, the processing pipeline behaves proportionally and not according to a particular target RT60 value: spectral shaping and noise-floor removal reduce late-reflection energy anywhere it occurs. Therefore, it degrades gracefully without reaching any threshold. With very reverberant conditions, the silences between words will become more complete, and it can be expected that the Energy Ratio and Spectral Contour Stability gains will decrease, because they both rely on the contrast between active and silent frames. The voice characteristics of the reciter and the characteristics of the microphone can be considered another source of variance: the Stage 1 gain curve is fixed from a single reference point and used in its identical form for every file, but still manages to improve all three domain-specific metrics irrespective of variations among different reciters and devices, thus suggesting that the properties captured by the curve are common for Qur’anic recitations in general and not specific to any one voice or any one microphone. A controlled study of these factors can be considered to be an important future task.

5.3. Iterative Development

Table 11 summarises the iterative parameter refinement that led to the final pipeline configuration. Table 12 lists the key parameter adjustments and their acoustic effects.

6. Conclusions

This paper presented a four-stage, phase-coherent dereverberation pipeline designed specifically for Quranic recitation preprocessing. The stages-domain-adapted spectral shaping, adaptive noise-floor subtraction, soft VAD with power-law boundary decay, and Griffin–Lim phase reconstruction-all operate within a single STFT block to maintain phase continuity.
On 48 real-room recordings, the proposed method achieved the best result in three of seven metrics: Spectral Contrast (+40.109, the highest of all methods); Spectral Contour Stability (+822.942, 5.55× better than WPE); and PESQ (+1.251), while attaining the second-highest Energy Ratio (+19.579 dB, +59% over the reverberant baseline).; Spectral Contrast (+40.109, the highest of all methods); Spectral Contour Stability (+822.942, 5.55× better than WPE); and PESQ (+1.251). Performance was further confirmed on a controlled synthetic experiment (PESQ = +2.495; SNR = −1.874 dB).
The most important finding is that WPE Dereverberation, despite achieving the best SNR on real recordings (−0.400 dB), produces the worst Spectral Contour Stability (+148.293) and collapses on Energy Ratio in 20 of 48 recordings. This demonstrates clearly that optimality with respect to general-purpose metrics does not imply optimality for domain-specific tasks such as Qira’at classification.
The limitation in A downstream feature discriminability experiment sort of further validated that the proposed pipeline does indeed yield more separable acoustic representations, in practice too. MFCC features extracted from the pipeline output achieved 77.1% KNN classification accuracy (leave-one-out CV, k = 3), which is better than the unprocessed reverberant signal at 75.0% and also beats all baselines, such as WPE at 70.8%. We also saw intra-class compactness go from 8.010 down to 7.576, so it seems like phonetic preservation really does carry over and ends up making feature spaces more discriminable for Qira‘at classification. One snag in this study is that the 48 real-room recordings were collected without recitation style labels, so direct Qira‘accuracy measurement was not possible. For the future, we will build a labeled multi-Qira‘dataset recorded under reverberant conditions, so we can quantify the classification improvement that comes from this preprocessing pipeline more directly. We will also look into extending the evaluation across multiple room types and different microphone layouts, too.

Author Contributions

Conceptualization, O.A.M., K.H. and K.A.R.; Methodology, O.A.M., K.H. and K.A.R.; Software, O.A.M. and K.H.; Validation, O.A.M., K.H., K.A.R. and B.M.; Formal analysis, O.A.M., K.H., K.A.R. and B.M.; Investigation, O.A.M., K.H. and K.A.R.; Resources, K.H.; Data curation, K.H. and K.A.R.; Writing—original draft, O.A.M., K.H., K.A.R. and B.M.; Writing—review & editing, O.A.M., K.H., K.A.R. and B.M.; Visualization, O.A.M., K.H., K.A.R. and B.M.; Supervision, K.H., K.A.R. and B.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Ethical review and approval were waived for this study, as the audio dataset used is publicly available (from EveryAyah.com), open-access, and does not contain private, personal, or identifiable human subject data requiring informed consent, in accordance with the exemption criteria for processing de-identified data outlined in Article 15 of Oman’s Personal Data Protection Law (Royal Decree 6/2022) and the National Health Research Center (NHRC) guidelines.

Informed Consent Statement

Patient consent was waived because this study utilized a publicly available, open-source audio repository (EveryAyah.com). The recordings do not contain private, personal, or confidential identifying information, and no direct interaction with human participants.

Data Availability Statement

The raw audio recordings used in this study are publicly available from EveryAyah.com. The preprocessed audio excerpts generated and analyzed during the current study are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

EREnergy Ratio
SCSSpectral Contour Stability
FFTFast Fourier Transform
MFCCMel-Frequency Cepstral Coefficient
PESQPerceptual Evaluation of Speech Quality
SCSpectral Contrast
SI-SDRScale-Invariant Signal-to-Distortion Ratio
SNRSignal-to-Noise Ratio
STFTShort-Time Fourier Transform
STOIShort-Time Objective Intelligibility
VADVoice Activity Detection
WPEWeighted Prediction Error

References

  1. Alkhateeb, J.H. A Machine Learning Approach for Recognizing the Holy Quran Reciter. Int. J. Adv. Comput. Sci. Appl. 2020, 11, 268–271. [Google Scholar] [CrossRef]
  2. Harere, A.A.; Jallad, K.A. Quran Recitation Recognition using End-to-End Deep Learning. arXiv 2023, arXiv:2305.07034. [Google Scholar]
  3. Davis, S.; Mermelstein, P. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE Trans. Acoust. Speech Signal Process. 1980, 28, 357–366. [Google Scholar] [CrossRef]
  4. Ghori, A.F.; Waheed, A.; Waqas, M.; Mehmood, A.; Ali, S.A. Acoustic modelling using deep learning for Quran recitation assistance. Int. J. Speech Technol. 2022, 26, 113–121. [Google Scholar] [CrossRef]
  5. Shakeel, M.A.; Khattak, H.A.; Khurshid, N. Deep acoustic modelling for Quranic Recitation–current solutions and future directions. IPSI Trans. Internet Res. 2024, 20, 61–73. [Google Scholar] [CrossRef]
  6. Kuttruff, H. Room Acoustics, 5th ed.; Spon Press: London, UK, 2009. [Google Scholar]
  7. Boll, S. Suppression of acoustic noise in speech using spectral subtraction. IEEE Trans. Acoust. Speech Signal Process. 1979, 27, 113–120. [Google Scholar] [CrossRef]
  8. Nakatani, T.; Yoshioka, T.; Kinoshita, K.; Miyoshi, M.; Juang, B.H. Speech Dereverberation Based on Variance-Normalized Delayed Linear Prediction. IEEE Trans. Audio Speech Lang. Process. 2010, 18, 1717–1731. [Google Scholar] [CrossRef]
  9. Griffin, D.W.; Lim, J.S. Signal estimation from modified short-time Fourier transform. IEEE Trans. Acoust. Speech Signal Process. 1984, 32, 236–243. [Google Scholar] [CrossRef]
  10. Al-Kharusi, M.H.; Hayat, K.; Ruqeishi, K.B.A.; Lone, H.R. A Critical Review of the Need for Knowledge-Centric Evaluation of Quranic Recitation. arXiv 2025, arXiv:2510.12858. [Google Scholar]
  11. Al Ajmi, S.A.; Hayat, K.; Al Obaidi, A.M.; Kumar, N.; Najim AL-Din, M.S.; Magnier, B. Faked speech detection with zero prior knowledge. Discov. Appl. Sci. 2024, 6, 288. [Google Scholar] [CrossRef]
  12. Khan, R.U.; Qamar, A.M.; Hadwan, M. Quranic Reciter Recognition: A Machine Learning Approach. Adv. Sci. Technol. Eng. Syst. J. 2019, 4, 173–176. [Google Scholar] [CrossRef]
  13. Alshboul, M.; Al Muaitah, A.R.; Al-Issa, S.; Al-Ayyoub, M. Enhanced Neural Speech Recognition of Quranic Recitations via a Large Audio Model. Appl. Sci. 2025, 15, 9521. [Google Scholar] [CrossRef]
  14. Al-Ayyoub, M.; Damer, N.A.; Hmeidi, I. Using deep learning for automatically determining correct application of basic quranic recitation rules. Int. Arab J. Inf. Technol. 2018, 15, 620–625. [Google Scholar]
  15. Kinoshita, K.; Delcroix, M.; Yoshioka, T.; Nakatani, T.; Habets, E.; Haeb-Umbach, R.; Leutnant, V.; Sehr, A.; Kellermann, W.; Maas, R.; et al. The REVERB challenge: Acommon evaluation framework for dereverberation and recognition of reverberant speech. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA); Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2013; pp. 1–4. [Google Scholar] [CrossRef]
  16. Drude, L.; Heitkaemper, J.; Böddeker, C.; Haeb-Umbach, R. SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition. arXiv 2019, arXiv:1910.13934. [Google Scholar]
  17. Wang, H.; Pandey, A.; Wang, D. A systematic study of DNN based speech enhancement in reverberant and reverberant-noisy environments. Comput. Speech Lang. 2025, 89, 101677. [Google Scholar] [CrossRef] [PubMed]
  18. Richter, J.; Welker, S.; Lemercier, J.M.; Lay, B.; Gerkmann, T. Speech Enhancement and Dereverberation With Diffusion-Based Generative Models. IEEE/ACM Trans. Audio Speech Lang. Process. 2023, 31, 2351–2364. [Google Scholar] [CrossRef]
  19. Lemercier, J.M.; Richter, J.; Welker, S.; Gerkmann, T. StoRM: A Diffusion-Based Stochastic Regeneration Model for Speech Enhancement and Dereverberation. IEEE/ACM Trans. Audio Speech Lang. Process. 2023, 31, 2724–2737. [Google Scholar] [CrossRef]
  20. Wang, Z.; Wichern, G.; Roux, J.L. Convolutive Prediction for Monaural Speech Dereverberation and Noisy-Reverberant Speaker Separation. arXiv 2021, arXiv:2108.07376. [Google Scholar]
  21. Rosenbaum, T.; Winebrand, E.; Cohen, O.; Cohen, I. Deep-Learning Framework for Efficient Real-Time Speech Enhancement and Dereverberation. Sensors 2025, 25, 630. [Google Scholar] [CrossRef] [PubMed]
  22. Perraudin, N.; Balázs, P.; Søndergaard, P.L. A fast Griffin-Lim algorithm. In Proceedings of the 2013 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, New Paltz, NY, USA, 20–23 October 2013; pp. 1–4. [Google Scholar]
  23. Scheibler, R.; Bezzam, E.; Dokmanic, I. Pyroomacoustics: A Python Package for Audio Room Simulation and Array Processing Algorithms. In Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada, 15–20 April 2018; pp. 351–355. [Google Scholar] [CrossRef]
  24. Allen, J.; Berkley, D. Image method for efficiently simulating small-room acoustics. J. Acoust. Soc. Am. 1979, 65, 943–950. [Google Scholar] [CrossRef]
  25. Fu, S.; Yu, C.; Hsieh, T.; Plantinga, P.; Ravanelli, M.; Lu, X.; Tsao, Y. MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement. arXiv 2021, arXiv:2104.03538. [Google Scholar]
  26. Rix, A.; Hollier, M.; Hekstra, A.; Beerends, J. Perceptual Evaluation of Speech Quality (PESQ) The New ITU Standard for End-to-End Speech Quality Assessment Part I–Time-Delay Compensation. AES J. Audio Eng. Soc. 2002, 50, 755–764. [Google Scholar]
  27. Taal, C.H.; Hendriks, R.C.; Heusdens, R.; Jensen, J. An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech. IEEE Trans. Audio Speech Lang. Process. 2011, 19, 2125–2136. [Google Scholar] [CrossRef]
  28. McFee, B.; Raffel, C.; Liang, D.; Ellis, D.P.; McVicar, M.; Battenberg, E.; Nieto, O. librosa: Audio and Music Signal Analysis in Python. SciPy 2015 2015, 2015, 18–25. [Google Scholar] [CrossRef]
  29. Tzanetakis, G.; Cook, P. Musical Genre Classification of Audio Signals. IEEE Trans. Speech Audio Process. 2002, 10, 293–302. [Google Scholar] [CrossRef]
  30. Williamson, D.S.; Wang, D. Time-frequency masking in the complex domain for speech dereverberation and denoising. IEEE ACM Trans. Audio Speech Lang. Process. 2017, 25, 1492–1501. [Google Scholar] [CrossRef] [PubMed]
Figure 1. The suggested four-stage processing pipeline’s block diagram. Reverberant Quranic recording in a single channel is the input. Improved recording with retained phonetic characteristics is the result.
Figure 1. The suggested four-stage processing pipeline’s block diagram. Reverberant Quranic recording in a single channel is the input. Improved recording with retained phonetic characteristics is the result.
Information 17 00714 g001
Figure 2. Spectrogram and Waveform comparison: Note the clear black silence regions between words in (b) vs. the continuous smeared energy in (a).
Figure 2. Spectrogram and Waveform comparison: Note the clear black silence regions between words in (b) vs. the continuous smeared energy in (a).
Information 17 00714 g002
Figure 3. Frequency-domain gain curve g c ( f ) applied in Stage 1. X-axis: frequency (Hz, log scale). Y-axis: gain factor (0–1.0). The 1000–2000 Hz formant preservation region and 4000+ Hz fricative preservation region are highlighted.
Figure 3. Frequency-domain gain curve g c ( f ) applied in Stage 1. X-axis: frequency (Hz, log scale). Y-axis: gain factor (0–1.0). The 1000–2000 Hz formant preservation region and 4000+ Hz fricative preservation region are highlighted.
Information 17 00714 g003
Figure 4. Comparison of gate functions: (a) abrupt transitions produced by a hard gate whispered consonants are fully suppressed and clicks appear at boundaries; (b) Soft VAD with Power-Law Decay ( D = 45 frames, p = 2.0 , gap = 220 ms) natural gradual fade preserves whispered consonants ( g min = 0.10 ) and eliminates boundary artifacts.
Figure 4. Comparison of gate functions: (a) abrupt transitions produced by a hard gate whispered consonants are fully suppressed and clicks appear at boundaries; (b) Soft VAD with Power-Law Decay ( D = 45 frames, p = 2.0 , gap = 220 ms) natural gradual fade preserves whispered consonants ( g min = 0.10 ) and eliminates boundary artifacts.
Information 17 00714 g004
Figure 5. Convergence of Griffin–Lim phase reconstruction over 25 iterations on segment-08.wav. (a) The normalized spectral distance decreases rapidly over the first iterations and then flattens into a plateau. (b) The per-iteration improvement falls below 1% from iteration 8 onward (1.04% at iteration 7, 0.88% at iteration 8). The pipeline retains 10 iterations as a conservative operating point that sits safely within the converged region, trading a negligible amount of computation for a margin beyond the 1% crossover.
Figure 5. Convergence of Griffin–Lim phase reconstruction over 25 iterations on segment-08.wav. (a) The normalized spectral distance decreases rapidly over the first iterations and then flattens into a plateau. (b) The per-iteration improvement falls below 1% from iteration 8 onward (1.04% at iteration 7, 0.88% at iteration 8). The pipeline retains 10 iterations as a conservative operating point that sits safely within the converged region, trading a negligible amount of computation for a margin beyond the 1% crossover.
Information 17 00714 g005aInformation 17 00714 g005b
Figure 6. Bar chart of controlled experiment results (RT60 = 0.4 s, Room = [6 × 5 × 3] m). All metrics computed against clean studio recording as ground truth. Proposed algorithm bars are highlighted in green. WPE achieves the highest Energy Ratio (+38.193 dB) but collapses in PESQ (+1.116) and STOI (+0.174), indicating severe perceptual degradation.
Figure 6. Bar chart of controlled experiment results (RT60 = 0.4 s, Room = [6 × 5 × 3] m). All metrics computed against clean studio recording as ground truth. Proposed algorithm bars are highlighted in green. WPE achieves the highest Energy Ratio (+38.193 dB) but collapses in PESQ (+1.116) and STOI (+0.174), indicating severe perceptual degradation.
Information 17 00714 g006
Figure 7. Multi-dimensional radar chart comparing all five methods across all seven metrics. Each axis represents one metric normalized to [0, 1]. The proposed algorithm (green) shows the largest coverage area in the three Quran-Specific metrics.
Figure 7. Multi-dimensional radar chart comparing all five methods across all seven metrics. Each axis represents one metric normalized to [0, 1]. The proposed algorithm (green) shows the largest coverage area in the three Quran-Specific metrics.
Information 17 00714 g007
Figure 8. t-SNE visualization of MFCC feature spaces extracted from 48 real-room recordings. Blue: low reverberation; Red: high reverberation. Our Algorithm achieves the highest KNN accuracy (77.1%), confirming that improvements in Spectral Contour Stability, Spectral Contrast, and Energy Ratio translate into more separable acoustic feature spaces for Qira’at classification.
Figure 8. t-SNE visualization of MFCC feature spaces extracted from 48 real-room recordings. Blue: low reverberation; Red: high reverberation. Our Algorithm achieves the highest KNN accuracy (77.1%), confirming that improvements in Spectral Contour Stability, Spectral Contrast, and Energy Ratio translate into more separable acoustic feature spaces for Qira’at classification.
Information 17 00714 g008
Figure 9. Quran-Specific metric comparison across all five methods (48 real-room recordings). The proposed algorithm (green hatched) achieves the highest value in all three metrics. WPE Dereverberation collapses in Energy Ratio (0.000 dB on 20 of 48 recordings) and produces the lowest Spectral Contour Stability (+148.29 − 5.55× below the proposed algorithm’s +822.94).
Figure 9. Quran-Specific metric comparison across all five methods (48 real-room recordings). The proposed algorithm (green hatched) achieves the highest value in all three metrics. WPE Dereverberation collapses in Energy Ratio (0.000 dB on 20 of 48 recordings) and produces the lowest Spectral Contour Stability (+148.29 − 5.55× below the proposed algorithm’s +822.94).
Information 17 00714 g009
Table 1. Summary of related work on Quranic speech processing and dereverberation.
Table 1. Summary of related work on Quranic speech processing and dereverberation.
Ref.MethodApplicationLimitationMetric
[1]MFCC + SVMReciter IDClean audio only; no reverb robustnessAcc.
[2]CNN-BiGRU + CTCQuranic ASRNo reverb preprocessingWER
[13]Whisper-based ASRQuranic recognitionAssumes clean audio inputWER
[4]DNN AcousticRecitation assist.Degrades under room acousticsWER
[7]Spectral Sub.Generic denoisingMusical noise; damages fricativesSNR
[9]WPE Dereverb.Generic dereverbDestroys SCS; fails on real roomsPESQ/STOI
[18]SGMSE+ DiffusionSpeech enhancementLarge corpus; domain mismatchPESQ
[19]StoRM (Hybrid)Speech enhancementHigh compute; not domain-specificPESQ/STOI
Prop.4-Stage pipelineQira’at classif.-SCS, SC, ER
Table 2. Spectral shaping gain curve parameters and acoustic rationale.
Table 2. Spectral shaping gain curve parameters and acoustic rationale.
Freq. Band (Hz) g c ( f ) EnergyAcoustic Rationale
<1000.05<0.1%Sub-bass noise; no phonetic content
100–3000.10∼3%Primary reverberation energy source
300–5000.18∼8%Secondary reverberation tail
500–10000.35∼37%Voice body; moderate preservation
1000–20000.75∼51%Vowel formants; high preservation (Madd/Harakaat)
2000–40000.15∼1%Reverberation tail attenuation
4000–16,0000.35<1%Fricative consonants; preserved
Table 3. Complete parameter specification.
Table 3. Complete parameter specification.
ParameterValuePurpose
FFT size ( N FFT )1024High freq. resolution (15.6 Hz/bin)
Hop length25616 ms temporal resolution
Window functionHannMinimize spectral leakage
Sample rate16,000 HzStandard speech processing rate
Noise percentile15thQuietest frames as noise reference
Noise subtraction ( α )2.0Balanced noise reduction
Spectral floor ( β )0.001Prevents frequency bin suppression
Gate minimum ( g min )0.10Preserves whispered consonants
VAD threshold30th pct.Speech/silence boundary
Gap-fill duration220 msPrevents mid-word fragmentation
Power decay exponent (p)2.0Natural word-boundary fade
Decay duration (D)45 frames720 ms consonant resonance
Smoothing window11 framesGradual gate transitions
Griffin–Lim iterations10Phase convergence
Table 4. Controlled evaluation-synthetic reverberation (RT60 = 0.4 s, Room = [ 6 , 5 , 3 ] m). Reference: clean studio recording. ⋆ denotes the best result per metric.
Table 4. Controlled evaluation-synthetic reverberation (RT60 = 0.4 s, Room = [ 6 , 5 , 3 ] m). Reference: clean studio recording. ⋆ denotes the best result per metric.
MetricReverb.Spec. Sub.WPEProposed
SNR (dB)−3.239−3.234−3.239−1.874
PESQ (ITU-T)+2.187+2.191+1.116+2.495
STOI (IEEE)+0.466+0.467 +0.174+0.466
Energy Ratio (dB)+14.906+15.054+38.193 +28.959
Total wins/40013
Table 5. Real room recording evaluation (48 files). ⋆ denotes the best result per metric.
Table 5. Real room recording evaluation (48 files). ⋆ denotes the best result per metric.
MetricOrig.S.Sub.Wien.WPEMetricGAN+Proposed
SNR (dB) 1.968 1.962 1.968 0.400 2.132 0.839
PESQ + 1.199 + 1.130 + 1.187 + 1.070 + 1.096 + 1.251
STOI + 0.117 + 0.114 + 0.115 + 0.087 + 0.103 + 0.107
ER (dB) + 12.336 + 12.422 + 12.338 + 19.092 + 22.851 + 19.579
SI-SDR (dB) 51.545 51.026 51.477 53.284 50.751 51.908
SC + 29.748 + 35.078 + 34.047 + 16.784 + 27.000 + 40.109
SCS + 473.416 + 507.341 + 508.441 + 148.293 + 444.291 + 822.942
Total wins100123
Table 6. Ablation study: average metrics across 48 recordings with each stage removed individually. The full pipeline is retained as the reference configuration.
Table 6. Ablation study: average metrics across 48 recordings with each stage removed individually. The full pipeline is retained as the reference configuration.
ConfigurationEnergy RatioSpectral ContrastSpectral Contour Stability
Full Pipeline+19.5840.11822.9
w/o Stage 1 (Shaping)+13.6140.14746.5
w/o Stage 2 (Noise Sub.)+18.8130.29533.5
w/o Stage 3 (Soft VAD)+18.3340.11824.5
w/o Stage 4 (Griffin–Lim)+19.5539.92821.9
No Processing (Reverberant)+12.3429.75473.4
Table 7. Energy Ratio threshold sensitivity. While the absolute value scales with the percentile, the advantage of the proposed pipeline over the reverberant baseline remains stable across all thresholds.
Table 7. Energy Ratio threshold sensitivity. While the absolute value scales with the percentile, the advantage of the proposed pipeline over the reverberant baseline remains stable across all thresholds.
ThresholdReverberant (dB)Proposed (dB)Advantage (dB)
20th/80th20.3027.09 + 6.79
25th/75th15.8223.52 + 7.70
30th/70th (default)12.4019.58 + 7.18
35th/65th10.1617.04 + 6.88
40th/60th8.5515.18 + 6.63
Mean +7.04
Table 8. Feature discriminability results (48 recordings, KNN k = 3 , Leave-One-Out CV). ⋆ denotes the best result per metric.
Table 8. Feature discriminability results (48 recordings, KNN k = 3 , Leave-One-Out CV). ⋆ denotes the best result per metric.
MethodKNN Accuracy (%)Intra-Class Compactness
Original75.08.010
Spectral Sub.75.07.965
Wiener Filter72.98.019
WPE Dereverb.70.87.366 
Proposed77.1 7.576
Table 9. Qira’at classification accuracy across audio conditions (SVM-RBF, Leave-One-Out CV, 51 recordings, 3 reciter styles).
Table 9. Qira’at classification accuracy across audio conditions (SVM-RBF, Leave-One-Out CV, 51 recordings, 3 reciter styles).
ConditionAccuracy
Clean96.1%
Reverberant94.1%
After Proposed Pipeline96.1%
Table 10. Statistical significance vs. best competing baseline (Wilcoxon signed-rank test, n = 48 ).
Table 10. Statistical significance vs. best competing baseline (Wilcoxon signed-rank test, n = 48 ).
MetricProposed (Mean ± SD)Baseline (Mean ± SD)p-ValueSig.
ER (dB) 19.579 ± 5.821 12.422 ± 6.113 <0.001***
SC 40.109 ± 2.734 35.078 ± 2.891 <0.001***
SCS 822.942 ± 213.4 508.441 ± 198.7 <0.001***
PESQ 1.251 ± 0.512 1.199 ± 0.471 <0.05*
SNR (dB) 0.839 ± 0.198 0.400 ± 0.271 <0.001***
*** p < 0.001 ;* p < 0.05 .Baselines: ER, SC, SCS—best competing method; PESQ—Original; SNR—WPE.
Table 11. Iterative parameter refinement history.
Table 11. Iterative parameter refinement history.
Ver.Key ChangeQualityProblemAction
V1Basic STFT + hard gatePoor (6.0)Metallic voiceAdd Griffin–Lim
V2+Griffin–Lim (10 iter.)Good (8.5)Sharp word endsExtend decay
V3Decay: 20 → 45 framesV.Good (9.5)Consonant clipsRaise g min
Final g min : 0.05 → 0.10Exc. (9.5+)None-
Table 12. Key parameter corrections and their acoustic effects.
Table 12. Key parameter corrections and their acoustic effects.
ParameterInitialFinalEffect
HF gain (>4 kHz)0.020.35Restored fricative clarity
Noise α 3.02.0Eliminated metallic artifacts
Gate typeHardSoftRemoved click artifacts
Gap-fill duration120 ms220 msPrevented mid-word cuts
Power exponent4.02.0Natural decay shape
FFT size5121024Improved frequency resolution
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Maaini, O.A.; Hayat, K.; Ruqeishi, K.A.; Magnier, B. A Phase-Coherent Four-Stage Pipeline for the Dereverberation of Quránic Recitation. Information 2026, 17, 714. https://doi.org/10.3390/info17070714

AMA Style

Maaini OA, Hayat K, Ruqeishi KA, Magnier B. A Phase-Coherent Four-Stage Pipeline for the Dereverberation of Quránic Recitation. Information. 2026; 17(7):714. https://doi.org/10.3390/info17070714

Chicago/Turabian Style

Maaini, Osama Al, Khizar Hayat, Khalil Al Ruqeishi, and Baptiste Magnier. 2026. "A Phase-Coherent Four-Stage Pipeline for the Dereverberation of Quránic Recitation" Information 17, no. 7: 714. https://doi.org/10.3390/info17070714

APA Style

Maaini, O. A., Hayat, K., Ruqeishi, K. A., & Magnier, B. (2026). A Phase-Coherent Four-Stage Pipeline for the Dereverberation of Quránic Recitation. Information, 17(7), 714. https://doi.org/10.3390/info17070714

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop