Author Contributions
Conceptualization, O.A.M., K.H. and K.A.R.; Methodology, O.A.M., K.H. and K.A.R.; Software, O.A.M. and K.H.; Validation, O.A.M., K.H., K.A.R. and B.M.; Formal analysis, O.A.M., K.H., K.A.R. and B.M.; Investigation, O.A.M., K.H. and K.A.R.; Resources, K.H.; Data curation, K.H. and K.A.R.; Writing—original draft, O.A.M., K.H., K.A.R. and B.M.; Writing—review & editing, O.A.M., K.H., K.A.R. and B.M.; Visualization, O.A.M., K.H., K.A.R. and B.M.; Supervision, K.H., K.A.R. and B.M. All authors have read and agreed to the published version of the manuscript.
Figure 1.
The suggested four-stage processing pipeline’s block diagram. Reverberant Quranic recording in a single channel is the input. Improved recording with retained phonetic characteristics is the result.
Figure 2.
Spectrogram and Waveform comparison: Note the clear black silence regions between words in (b) vs. the continuous smeared energy in (a).
Figure 3.
Frequency-domain gain curve applied in Stage 1. X-axis: frequency (Hz, log scale). Y-axis: gain factor (0–1.0). The 1000–2000 Hz formant preservation region and 4000+ Hz fricative preservation region are highlighted.
Figure 4.
Comparison of gate functions: (a) abrupt transitions produced by a hard gate whispered consonants are fully suppressed and clicks appear at boundaries; (b) Soft VAD with Power-Law Decay ( frames, , gap = 220 ms) natural gradual fade preserves whispered consonants () and eliminates boundary artifacts.
Figure 5.
Convergence of Griffin–Lim phase reconstruction over 25 iterations on segment-08.wav. (a) The normalized spectral distance decreases rapidly over the first iterations and then flattens into a plateau. (b) The per-iteration improvement falls below 1% from iteration 8 onward (1.04% at iteration 7, 0.88% at iteration 8). The pipeline retains 10 iterations as a conservative operating point that sits safely within the converged region, trading a negligible amount of computation for a margin beyond the 1% crossover.
Figure 6.
Bar chart of controlled experiment results (RT60 = 0.4 s, Room = [6 × 5 × 3] m). All metrics computed against clean studio recording as ground truth. Proposed algorithm bars are highlighted in green. WPE achieves the highest Energy Ratio (+38.193 dB) but collapses in PESQ (+1.116) and STOI (+0.174), indicating severe perceptual degradation.
Figure 7.
Multi-dimensional radar chart comparing all five methods across all seven metrics. Each axis represents one metric normalized to [0, 1]. The proposed algorithm (green) shows the largest coverage area in the three Quran-Specific metrics.
Figure 8.
t-SNE visualization of MFCC feature spaces extracted from 48 real-room recordings. Blue: low reverberation; Red: high reverberation. Our Algorithm achieves the highest KNN accuracy (77.1%), confirming that improvements in Spectral Contour Stability, Spectral Contrast, and Energy Ratio translate into more separable acoustic feature spaces for Qira’at classification.
Figure 9.
Quran-Specific metric comparison across all five methods (48 real-room recordings). The proposed algorithm (green hatched) achieves the highest value in all three metrics. WPE Dereverberation collapses in Energy Ratio (0.000 dB on 20 of 48 recordings) and produces the lowest Spectral Contour Stability (+148.29 − 5.55× below the proposed algorithm’s +822.94).
Table 1.
Summary of related work on Quranic speech processing and dereverberation.
| Ref. | Method | Application | Limitation | Metric |
|---|
| [1] | MFCC + SVM | Reciter ID | Clean audio only; no reverb robustness | Acc. |
| [2] | CNN-BiGRU + CTC | Quranic ASR | No reverb preprocessing | WER |
| [13] | Whisper-based ASR | Quranic recognition | Assumes clean audio input | WER |
| [4] | DNN Acoustic | Recitation assist. | Degrades under room acoustics | WER |
| [7] | Spectral Sub. | Generic denoising | Musical noise; damages fricatives | SNR |
| [9] | WPE Dereverb. | Generic dereverb | Destroys SCS; fails on real rooms | PESQ/STOI |
| [18] | SGMSE+ Diffusion | Speech enhancement | Large corpus; domain mismatch | PESQ |
| [19] | StoRM (Hybrid) | Speech enhancement | High compute; not domain-specific | PESQ/STOI |
| Prop. | 4-Stage pipeline | Qira’at classif. | - | SCS, SC, ER |
Table 2.
Spectral shaping gain curve parameters and acoustic rationale.
| Freq. Band (Hz) | | Energy | Acoustic Rationale |
|---|
| <100 | 0.05 | <0.1% | Sub-bass noise; no phonetic content |
| 100–300 | 0.10 | ∼3% | Primary reverberation energy source |
| 300–500 | 0.18 | ∼8% | Secondary reverberation tail |
| 500–1000 | 0.35 | ∼37% | Voice body; moderate preservation |
| 1000–2000 | 0.75 | ∼51% | Vowel formants; high preservation (Madd/Harakaat) |
| 2000–4000 | 0.15 | ∼1% | Reverberation tail attenuation |
| 4000–16,000 | 0.35 | <1% | Fricative consonants; preserved |
Table 3.
Complete parameter specification.
| Parameter | Value | Purpose |
|---|
| FFT size () | 1024 | High freq. resolution (15.6 Hz/bin) |
| Hop length | 256 | 16 ms temporal resolution |
| Window function | Hann | Minimize spectral leakage |
| Sample rate | 16,000 Hz | Standard speech processing rate |
| Noise percentile | 15th | Quietest frames as noise reference |
| Noise subtraction () | 2.0 | Balanced noise reduction |
| Spectral floor () | 0.001 | Prevents frequency bin suppression |
| Gate minimum () | 0.10 | Preserves whispered consonants |
| VAD threshold | 30th pct. | Speech/silence boundary |
| Gap-fill duration | 220 ms | Prevents mid-word fragmentation |
| Power decay exponent (p) | 2.0 | Natural word-boundary fade |
| Decay duration (D) | 45 frames | 720 ms consonant resonance |
| Smoothing window | 11 frames | Gradual gate transitions |
| Griffin–Lim iterations | 10 | Phase convergence |
Table 4.
Controlled evaluation-synthetic reverberation (RT60 = 0.4 s, Room = m). Reference: clean studio recording. ⋆ denotes the best result per metric.
| Metric | Reverb. | Spec. Sub. | WPE | Proposed |
|---|
| SNR (dB) | −3.239 | −3.234 | −3.239 | −1.874 ⋆ |
| PESQ (ITU-T) | +2.187 | +2.191 | +1.116 | +2.495 ⋆ |
| STOI (IEEE) | +0.466 | +0.467 ⋆ | +0.174 | +0.466 |
| Energy Ratio (dB) | +14.906 | +15.054 | +38.193 ⋆ | +28.959 |
| Total wins/4 | 0 | 0 | 1 | 3 |
Table 5.
Real room recording evaluation (48 files). ⋆ denotes the best result per metric.
| Metric | Orig. | S.Sub. | Wien. | WPE | MetricGAN+ | Proposed |
|---|
| SNR (dB) | | | | | | |
| PESQ | | | | | | |
| STOI | | | | | | |
| ER (dB) | | | | | | |
| SI-SDR (dB) | | | | | | |
| SC | | | | | | |
| SCS | | | | | | |
| Total wins | 1 | 0 | 0 | 1 | 2 | 3 |
Table 6.
Ablation study: average metrics across 48 recordings with each stage removed individually. The full pipeline is retained as the reference configuration.
| Configuration | Energy Ratio | Spectral Contrast | Spectral Contour Stability |
|---|
| Full Pipeline | +19.58 | 40.11 | 822.9 |
| w/o Stage 1 (Shaping) | +13.61 | 40.14 | 746.5 |
| w/o Stage 2 (Noise Sub.) | +18.81 | 30.29 | 533.5 |
| w/o Stage 3 (Soft VAD) | +18.33 | 40.11 | 824.5 |
| w/o Stage 4 (Griffin–Lim) | +19.55 | 39.92 | 821.9 |
| No Processing (Reverberant) | +12.34 | 29.75 | 473.4 |
Table 7.
Energy Ratio threshold sensitivity. While the absolute value scales with the percentile, the advantage of the proposed pipeline over the reverberant baseline remains stable across all thresholds.
| Threshold | Reverberant (dB) | Proposed (dB) | Advantage (dB) |
|---|
| 20th/80th | 20.30 | 27.09 | |
| 25th/75th | 15.82 | 23.52 | |
| 30th/70th (default) | 12.40 | 19.58 | |
| 35th/65th | 10.16 | 17.04 | |
| 40th/60th | 8.55 | 15.18 | |
| Mean | | | +7.04 |
Table 8.
Feature discriminability results (48 recordings, KNN , Leave-One-Out CV). ⋆ denotes the best result per metric.
| Method | KNN Accuracy (%) | Intra-Class Compactness |
|---|
| Original | 75.0 | 8.010 |
| Spectral Sub. | 75.0 | 7.965 |
| Wiener Filter | 72.9 | 8.019 |
| WPE Dereverb. | 70.8 | 7.366 ⋆ |
| Proposed | 77.1 ⋆ | 7.576 |
Table 9.
Qira’at classification accuracy across audio conditions (SVM-RBF, Leave-One-Out CV, 51 recordings, 3 reciter styles).
| Condition | Accuracy |
|---|
| Clean | 96.1% |
| Reverberant | 94.1% |
| After Proposed Pipeline | 96.1% |
Table 10.
Statistical significance vs. best competing baseline (Wilcoxon signed-rank test, ).
| Metric | Proposed (Mean ± SD) | Baseline (Mean ± SD) | p-Value | Sig. |
|---|
| ER (dB) | | | <0.001 | *** |
| SC | | | <0.001 | *** |
| SCS | | | <0.001 | *** |
| PESQ | | | <0.05 | * |
| SNR (dB) | | | <0.001 | *** |
Table 11.
Iterative parameter refinement history.
| Ver. | Key Change | Quality | Problem | Action |
|---|
| V1 | Basic STFT + hard gate | Poor (6.0) | Metallic voice | Add Griffin–Lim |
| V2 | +Griffin–Lim (10 iter.) | Good (8.5) | Sharp word ends | Extend decay |
| V3 | Decay: 20 → 45 frames | V.Good (9.5) | Consonant clips | Raise |
| Final | : 0.05 → 0.10 | Exc. (9.5+) | None | - |
Table 12.
Key parameter corrections and their acoustic effects.
| Parameter | Initial | Final | Effect |
|---|
| HF gain (>4 kHz) | 0.02 | 0.35 | Restored fricative clarity |
| Noise | 3.0 | 2.0 | Eliminated metallic artifacts |
| Gate type | Hard | Soft | Removed click artifacts |
| Gap-fill duration | 120 ms | 220 ms | Prevented mid-word cuts |
| Power exponent | 4.0 | 2.0 | Natural decay shape |
| FFT size | 512 | 1024 | Improved frequency resolution |