Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (276)

Search Parameters:
Keywords = speech embedding

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
22 pages, 1095 KB  
Article
Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh–English Overlapping Speech
by Rakhat Meiramov, Tomiris Rakhimzhanova, Adil Taibassarov, Zhanat Makhataeva and Huseyin Atakan Varol
Mach. Learn. Knowl. Extr. 2026, 8(8), 246; https://doi.org/10.3390/make8080246 - 14 Aug 2026
Viewed by 273
Abstract
Target-speaker automatic speech recognition (TS-ASR) enables transcription of a specific speaker in multi-talker environments, yet remains largely unexplored for multilingual, low-resource languages. Existing TS-ASR systems predominantly target monolingual English using diarization-based or speaker-embedding approaches, leaving a critical gap for languages such as Kazakh, [...] Read more.
Target-speaker automatic speech recognition (TS-ASR) enables transcription of a specific speaker in multi-talker environments, yet remains largely unexplored for multilingual, low-resource languages. Existing TS-ASR systems predominantly target monolingual English using diarization-based or speaker-embedding approaches, leaving a critical gap for languages such as Kazakh, where code-switching with Russian and English is commonplace. We propose Persona-ASR, a modular two-stage architecture. The first stage is an explicit target-presence gate that verifies whether the enrolled speaker appears in the mixture and emits a <no_target> token to suppress transcription when the speaker is absent, directly addressing the acoustic-hallucination failure mode of prior systems. The second stage performs enrollment-conditioned recognition: a 192-dimensional ECAPA-TDNN speaker embedding modulates a WavLM-Base-Plus encoder through feature-wise linear modulation (FiLM), while language-specific CTC heads enable joint Kazakh and English decoding without forcing Latin and Cyrillic symbols to compete in a single output space. To evaluate the system, we introduce KazMix3, a Kazakh overlap dataset for TS-ASR training, and PersonaMix, a controlled bilingual benchmark spanning same- and cross-language enrollment across varying interferer counts (1–3) and signal-to-noise ratios (3 to +3 dB). Persona-ASR outperforms a strong off-the-shelf cascade baseline by 13.3 WER points on English and 24.6 on Kazakh, and matches a published monolingual English baseline. On PersonaMix, speaker conditioning reduces relative word error rate by 40.7% on English and 59.3% on Kazakh mixtures over an unconditioned variant of the same model, and cross-language enrollment (unseen during training) remains effective, increasing average raw WER by only 4.1 points (English) and 2.2 points (Kazakh) relative to same-language enrollment. To our knowledge, Persona-ASR is the first TS-ASR system for the Kazakh language, establishing a foundation for multilingual personalized ASR in low-resource settings. Full article
Show Figures

Figure 1

18 pages, 3100 KB  
Article
Design and Efficacy of a Speaker Verification Method Combining CNN and Transformer for Secure Access Control
by Xiao Li, Xiao Hu, Kun Niu and Ling Tuo
Sensors 2026, 26(16), 5101; https://doi.org/10.3390/s26165101 - 12 Aug 2026
Viewed by 260
Abstract
In the rapidly evolving landscape of computer and mobile applications, the demand for secure access control has become increasingly pivotal. This paper introduces a speaker verification method aimed at remotely verifying an individual’s claimed identity, so as to achieve access authorization. The primary [...] Read more.
In the rapidly evolving landscape of computer and mobile applications, the demand for secure access control has become increasingly pivotal. This paper introduces a speaker verification method aimed at remotely verifying an individual’s claimed identity, so as to achieve access authorization. The primary objective is to develop a deep learning network capable of eliminating redundant and irrelevant information while learning robust deep speaker embedding descriptors that capture speaker-specific idiosyncrasies. This paper presents an AI framework combining convolutional neural network and transformer architectures, enhanced by (1) a stereoscopic attention mechanism that computes fine-grained attention weights across frequency, time, and channel dimensions with fewer parameters than existing CBAM or SE modules, and (2) a multi-feature aggregation mechanism that fuses supervised and unsupervised features at the utterance level to maximize complementary speaker information. These innovations introduce a new perspective for integrating acoustic and articulatory features, effectively addressing the challenges related to short-segment speech and cross-domain scenarios. Experimental results demonstrate that the AI framework achieves competitive performance under domain mismatch conditions. The methodology acts as a catalyst for advancing application authorization, highlighting the transformative potential of AI-driven innovations in the software engineering field. Full article
Show Figures

Figure 1

24 pages, 6076 KB  
Article
Spatial-Temporal Relation Enhancement for Speech Emotion Recognition from Acoustic Signals Using Fibonacci Encoding with Diverse Feature Fusion
by Shicong Huang, Jinghao Zhang, Jingchao Xu, Qi Zhang, Zihan Li, Songyang Wang, Zirui Qiu, Zijia Xiong, Jintian Liang and Changzeng Fu
Sensors 2026, 26(16), 5050; https://doi.org/10.3390/s26165050 - 9 Aug 2026
Viewed by 481
Abstract
Speech emotion recognition (SER) infers affective states from speech signals, but positional encoding for acoustic tokens remains underexplored in Transformer-based SER. Existing models often reuse encodings designed for text and do not explicitly account for the different sequential and two-dimensional structures of acoustic [...] Read more.
Speech emotion recognition (SER) infers affective states from speech signals, but positional encoding for acoustic tokens remains underexplored in Transformer-based SER. Existing models often reuse encodings designed for text and do not explicitly account for the different sequential and two-dimensional structures of acoustic representations. We propose Fibonacci Position Embedding (FPE) and Fibonacci Target Shutter (FTS). FTS constructs overlapping candidate-index sets over time–frequency token grids, and FPE samples a Fibonacci index and applies its modulo-wrapped, dimension-dependent phase rotation to query and key vectors. The modules are integrated into STRE-Former, which fuses Wav2Vec, log-mel spectrogram, and MFCC representations through asymmetric cross-representation attention with representation-specific positional encodings. We also introduce an implementation-consistent conditional-entropy formulation that quantifies uncertainty in recovering a token location from its sampled positional representation; this quantity characterizes positional ambiguity rather than downstream modeling capacity. Experiments over 64 positional-encoding combinations on IEMOCAP and MELD identify dataset-dependent highest-observed configurations, reaching 74.21% weighted accuracy on IEMOCAP-4, 74.54% on IEMOCAP-6, and 49.44% on MELD. These empirical observations suggest that the relative behavior of positional-encoding strategies may depend on the acoustic representation and evaluation dataset, rather than supporting a single universally optimal scheme. Full article
(This article belongs to the Special Issue Applications of Sensors in Emotion Recognition)
Show Figures

Figure 1

19 pages, 3881 KB  
Article
Multichannel Acoustic Beamforming for Speaker Localization and DOA-Based Tracking
by Jose Antonio Lopez-Olvera, Hector Manuel Perez-Meana, Elizabeth Garcia-Rios, Enrique Escamilla-Hernandez, Jose Manuel Carrichi-Chavez and Jesus Betancourt-Martinez
Electronics 2026, 15(15), 3410; https://doi.org/10.3390/electronics15153410 - 1 Aug 2026
Viewed by 345
Abstract
Real-time speaker localization and speech enhancement remain challenging for embedded acoustic systems operating in noisy and reverberant environments. This paper presents a multichannel beamforming framework based on Generalized Cross-Correlation with Phase Transform (GCC-PHAT) for Direction of Arrival (DOA) estimation and Delay-and-Sum (DAS) beamforming [...] Read more.
Real-time speaker localization and speech enhancement remain challenging for embedded acoustic systems operating in noisy and reverberant environments. This paper presents a multichannel beamforming framework based on Generalized Cross-Correlation with Phase Transform (GCC-PHAT) for Direction of Arrival (DOA) estimation and Delay-and-Sum (DAS) beamforming using a compact four-element circular MEMS microphone array. The proposed approach incorporates DOA smoothing and verification to improve localization stability before beamforming. Experimental evaluation in a controlled indoor environment demonstrated localization errors below 1°, while the beamforming stage achieved Signal-to-Noise Ratio (SNR) values predominantly between 25 dB and 40 dB, with peaks approaching 45 dB, and Root Mean Square Error (RMSE) values below 0.05 for most processing windows with an average processing time per frame of 3.473 ms and a memory consumption of 67.11 KB. Comparative results show that the proposed system provides competitive localization accuracy with a computationally simple processing pipeline, making it suitable for real-time embedded applications such as intelligent voice interfaces, videoconferencing systems, and service robotics. Full article
(This article belongs to the Section Circuit and Signal Processing)
Show Figures

Figure 1

25 pages, 1129 KB  
Article
Test-Time Adaptation for Personal Voice Activity Detection: VAD-Gated Test-Time Training and Speaker Embedding Adaptation
by Tai-You Chen, Chien-Chia Chiu, Jung-Shan Lin and Jeih-Weih Hung
Electronics 2026, 15(14), 3111; https://doi.org/10.3390/electronics15143111 - 15 Jul 2026
Viewed by 353
Abstract
Personal voice activity detection (PVAD) identifies whether each detected speech frame originates from a designated target speaker. Modern PVAD systems are typically trained offline and then deployed with frozen model parameters and a fixed, pre-enrolled speaker embedding, leaving them unable to adapt to [...] Read more.
Personal voice activity detection (PVAD) identifies whether each detected speech frame originates from a designated target speaker. Modern PVAD systems are typically trained offline and then deployed with frozen model parameters and a fixed, pre-enrolled speaker embedding, leaving them unable to adapt to distribution shifts at inference time such as unseen acoustic environments, changing speaking styles, or mismatches between enrollment and test conditions. Test-time training (TTT) and test-time adaptation have shown promise in language, vision, and several speech tasks, yet their behavior on PVAD has not been studied. In this work, we present an empirical study of two complementary test-time adaptation mechanisms built on top of the recently proposed FDE-Mamba backbone. The first is a VAD-gated TTT adapter, which instantiates the TTT-Linear formulation within the personalization pathway and augments it with a VAD-probability gate and exponential moving-average stabilization, adapting an internal weight matrix on the speaker-conditioned feature stream of each test utterance. The second is TEA (Test-time Embedding Adaptation), a scheme that keeps all model parameters frozen and instead adapts the target speaker d-vector itself via self-supervised objectives at inference time, directly targeting enrollment–test mismatch. We evaluate both mechanisms on the LibriSpeech PVAD benchmark across two backbones (LSTM-based FDE-RNN and Mamba-based FDE-Mamba), reporting category-wise average precision, mean average precision (mAP), accuracy, recall, precision, and real-time factor. We further isolate the effect of a post-hoc Gaussian smoothing step and report that, of all the components we examine, this task-agnostic smoothing accounts for the largest single accuracy gain on the FDE-Mamba backbone (accuracy 89.87%90.47%); the test-time adaptation mechanisms contribute a separate, smaller gain that is concentrated on speaker-discrimination metrics (mAP, precision) rather than on accuracy. Overall, the proposed test-time adaptation yields a consistent but modest improvement over the FDE-Mamba baseline (mAP 0.96050.9641, precision 0.8810.899), while slightly reducing recall and increasing inference cost when TEA is enabled. Through ablation studies, we quantify the independent and combined contribution of each component and characterize the recall–precision trade-off introduced by adaptation. These findings, together with a discussion of their limitations and cost–benefit profile, provide a measured baseline and design insights for future work on adaptive PVAD, particularly under stronger acoustic and enrollment mismatches than those captured by the LibriSpeech protocol. Full article
Show Figures

Figure 1

39 pages, 1627 KB  
Review
A Survey of LSTM Pedestrian Intention Prediction and Lightweight Methods for Intelligent Guide Sticks
by Yijia Cai, Fang Jing, Huafeng Qu, Yuxi Xie and Shafrida Sahrani
Future Internet 2026, 18(7), 362; https://doi.org/10.3390/fi18070362 - 15 Jul 2026
Viewed by 420
Abstract
The travel problem of visually impaired people is a worldwide issue that needs urgent attention. Although intelligent guide sticks provide obstacle detection and early warning through multi-sensor fusion and embedded algorithms, existing systems generally cannot model the temporal movement patterns of dynamic obstacles, [...] Read more.
The travel problem of visually impaired people is a worldwide issue that needs urgent attention. Although intelligent guide sticks provide obstacle detection and early warning through multi-sensor fusion and embedded algorithms, existing systems generally cannot model the temporal movement patterns of dynamic obstacles, such as pedestrians and vehicles, thereby hindering intention prediction and active obstacle avoidance. Long short-term memory (LSTM), with its gating mechanism, effectively captures long-term dependencies in trajectories and offers a promising solution. This review compares and analyzes LSTM against other mainstream temporal models under the resource constraints of intelligent guide sticks and finds that LSTM demonstrates a favorable combination in temporal modeling capability, lightweight maturity, and edge deployment feasibility. We categorize five lightweight techniques—architecture simplification, low-rank decomposition, structured pruning, quantization, and knowledge distillation—and examine their compression effectiveness, accuracy preservation, and hardware applicability across typical platforms. Furthermore, this review surveys application cases in speech guidance, trajectory prediction-based obstacle avoidance, positioning and navigation, edge computing, and Internet collaboration, exploring the diverse potential of LSTM in intelligent guide stick scenarios. The findings indicate that, after lightweight processing, LSTM models can meet the deployment requirements of resource-constrained edge devices, suggesting their potential feasibility on resource-constrained hardware platforms. However, existing applications still face challenges in balancing real-time performance and accuracy, meeting stringent resource constraints, and the absence of end-to-end validation on real intelligent guide stick prototypes. The reviewed evidence suggests that LSTM-based prediction represents a promising and practically valuable pathway for transitioning intelligent guide sticks from passive response to active prediction. Future research should prioritize real-world deployment validation, domain-specific data collection, and hardware-software co-design to realize its potential fully. Full article
(This article belongs to the Special Issue Distributed Intelligence for IoT and Smart Systems)
Show Figures

Graphical abstract

33 pages, 16638 KB  
Article
Task-State fMRI-Derived Whole-Brain Functional Topology-Constrained Spiking Neural Network with an Embedded Auditory Core Circuit for Speech Recognition
by Lei Guo and Yaxin Yang
Biomimetics 2026, 11(7), 481; https://doi.org/10.3390/biomimetics11070481 - 9 Jul 2026
Viewed by 327
Abstract
The topology of spiking neural networks (SNNs) plays an important role in determining their dynamic representation ability, recognition performance, and biological interpretability in speech recognition. However, most existing SNN reservoirs are constructed using random, regular, or manually designed connectivity patterns, which may not [...] Read more.
The topology of spiking neural networks (SNNs) plays an important role in determining their dynamic representation ability, recognition performance, and biological interpretability in speech recognition. However, most existing SNN reservoirs are constructed using random, regular, or manually designed connectivity patterns, which may not reflect the functional organization of the human brain during speech perception. In this study, we propose a task-state fMRI-constrained SNN framework for speech recognition. Human fMRI data acquired during naturalistic English audiobook listening are used offline to derive a task-state whole-brain functional topology, which serves as a biologically inspired structural prior for the recurrent connectivity of the SNN reservoir. Because the fMRI and downstream isolated-digit recognition tasks use different speech paradigms, this topology is interpreted as a general speech-listening prior rather than a digit-specific neural representation. The Schaefer-400 cortical parcellation is used to define 400 whole-brain functional nodes, all of which are retained to preserve distributed cortical interactions during speech listening. Within this topology, 7 SomMotB_Aud parcels are identified as auditory core nodes and analyzed as an embedded auditory circuit. Compared with resting-state fMRI, task-state fMRI shows enhanced functional connectivity among these auditory nodes, indicating task-related auditory-circuit activation. The resulting 400-node task-state topology is mapped onto the recurrent connectivity of the SNN reservoir. This mapping is regarded as a topology-constrained computational abstraction rather than a direct model of biological information transmission. During recognition, speech spike trains are the only external input, while fMRI data are used only for offline topology construction. Experimental comparisons with baseline SNNs show that the proposed topology improves recognition performance and biological interpretability. Resting-state topology comparison, auditory-core contribution analysis, threshold-sensitivity analysis, and statistical testing are further used to evaluate robustness. These findings suggest that speech-evoked whole-brain functional organization may provide an effective topology prior for biologically inspired speech recognition models. Full article
(This article belongs to the Section Biological Optimisation and Management)
Show Figures

Figure 1

26 pages, 3508 KB  
Article
Dual-Track Residual Framework for Residual Strength-Controlled Emotional Speech Synthesis
by Youdong Ding, Yafan Geng, Wenjing Yu and Feifan Cai
Appl. Sci. 2026, 16(13), 6613; https://doi.org/10.3390/app16136613 - 2 Jul 2026
Viewed by 527
Abstract
Recent text-to-speech (TTS) systems can synthesize natural and intelligible speech, but adding controllable emotional expression to a pretrained model while preserving target-speaker identity remains challenging. This setting is especially constrained when the acoustic backbone is kept frozen and emotional adaptation relies on additional [...] Read more.
Recent text-to-speech (TTS) systems can synthesize natural and intelligible speech, but adding controllable emotional expression to a pretrained model while preserving target-speaker identity remains challenging. This setting is especially constrained when the acoustic backbone is kept frozen and emotional adaptation relies on additional trainable modules. We study emotional adaptation for a frozen flow-matching TTS backbone and propose the dual-track residual framework (DTRF). The DTRF keeps the neutral-adapted Matcha-Base backbone unchanged, represents emotion as a neutral-anchor residual, and introduces two residual control paths: an asymmetric zero-initialized acoustic control branch for spectral vector field modulation and an emotional duration adapter (EDA) for duration-level prosody control. Rather than injecting emotion only into the acoustic path, the DTRF applies emotion control to both acoustic vector field prediction and phoneme duration prediction, jointly adjusting spectral realization and temporal prosody. A global neutral anchor converts absolute emotion embeddings into relative residuals so that the control signal describes the deviation from neutral speech toward the target emotion rather than an absolute style vector. During inference, a shared scalar factor α scales both residual paths, providing a practical residual strength interface for controllable emotion rendering. Moderate α values tend to increase emotional salience, whereas larger extrapolative values introduce trade-offs in naturalness, speaker similarity, and intelligibility. Experiments on the English subset of the Emotional Speech Dataset (ESD) show that the DTRF improves emotion-related metrics relative to the internal full-parameter updating and style token conditioning baselines, while maintaining a practical balance among speaker similarity, naturalness, and intelligibility. The emotion control modules contain approximately 27 M trainable parameters, corresponding to 29.61% of the full model parameters and a 70.39% reduction compared with full-parameter updating. These results suggest that jointly modeling acoustic and duration residuals can be an effective strategy for adding residual strength-controlled emotional rendering to a frozen flow-matching TTS model without full-backbone updating. Full article
(This article belongs to the Special Issue Deep Learning for Speech, Image and Language Processing)
Show Figures

Figure 1

51 pages, 1481 KB  
Article
A Hybrid Feature-Enhanced IndoBERT Framework with Controlled Semi-Supervised Learning for Low-Resource Indonesian Hate Speech Detection
by Shoffan Saifullah and Rafał Dreżewski
Appl. Sci. 2026, 16(13), 6478; https://doi.org/10.3390/app16136478 - 29 Jun 2026
Viewed by 573
Abstract
Low-resource hate speech detection remains a challenging task for Indonesian social media due to limited labeled annotations, highly informal linguistic expressions, and substantial lexical variability. Under such conditions, purely supervised transformer models often suffer from unstable semantic generalization, while conventional pseudo-labeling methods are [...] Read more.
Low-resource hate speech detection remains a challenging task for Indonesian social media due to limited labeled annotations, highly informal linguistic expressions, and substantial lexical variability. Under such conditions, purely supervised transformer models often suffer from unstable semantic generalization, while conventional pseudo-labeling methods are vulnerable to noisy unlabeled sample propagation. To address these limitations, this study proposes a hybrid feature-enhanced IndoBERT framework integrated with a controlled semi-supervised learning strategy. The proposed model combines contextual IndoBERT embeddings with abusive lexicon cues, handcrafted linguistic indicators, and TF-IDF–SVD statistical representations through a lightweight concatenation–projection feature fusion mechanism, while unlabeled data are incorporated via adaptive confidence thresholding and class-balanced pseudo-label selection to improve pseudo-label reliability. Extensive experiments were conducted under realistic low-resource supervision settings using only 5%, 10%, and 20% labeled data, and the proposed framework was systematically compared against representative baselines, including sparse lexical machine learning models, shallow neural architectures, multilingual transformers, IndoBERTweet, naive pseudo-labeling, and LLM-based prompting. The results show that model effectiveness is strongly supervision-dependent. Under the most extreme low-resource setting, compact statistical augmentation provides the most stable complementary signal, whereas under moderate low-resource supervision, the full hybrid representation combined with controlled semi-supervised learning yields the strongest and most consistent gains. The proposed Hybrid IndoBERT + controlled SSL framework outperforms all baselines at the 20% labeled setting, reaching an accuracy of 0.8654, Macro-F1 of 0.8633, and ROC-AUC of 0.9334. Additional analyses of pseudo-label reliability, calibration behavior, computational efficiency, and qualitative error patterns further show that the proposed framework improves low-resource robustness while maintaining comparable inference-time efficiency. These findings demonstrate that low-resource hate speech detection benefits most from the staged integration of contextual semantic modeling, interpretable linguistic cues, global lexical–statistical structure, and carefully regulated unlabeled data exploitation. Additional experiments using GPT-4o-mini and Llama-3.1-8B further demonstrate that the proposed framework remains competitive against general-purpose large language model prompting approaches under low-resource Indonesian hate speech detection scenarios. The proposed framework provides a practical and reproducible direction for hate speech detection in annotation-constrained social media environments. Full article
Show Figures

Figure 1

10 pages, 262 KB  
Proceeding Paper
Analytical Study of Key Techniques for Cross-Modal Feature Alignment and Decision-Level Fusion in Brain–Computer Interface-Virtual Reality Systems
by Dan Liu
Eng. Proc. 2026, 141(1), 19; https://doi.org/10.3390/engproc2026141019 - 29 Jun 2026
Viewed by 296
Abstract
Feature alignment and decision-level fusion in multimodal BCI–VR interaction were investigated using Transformer-based cross-modal embeddings, Lab Streaming Layer time synchronization, attention masks, and wavelet filtering for robust representation. A four-modal acquisition and synchronization platform covering electroencephalography, electromyography, eye-tracking, and speech was constructed, and [...] Read more.
Feature alignment and decision-level fusion in multimodal BCI–VR interaction were investigated using Transformer-based cross-modal embeddings, Lab Streaming Layer time synchronization, attention masks, and wavelet filtering for robust representation. A four-modal acquisition and synchronization platform covering electroencephalography, electromyography, eye-tracking, and speech was constructed, and fusion was achieved by introducing a stacking meta-learner together with a confidence-aware dynamic weighting mechanism. Prototype validation and comparative evaluations were conducted on virtual reality (VR) target-selection, trajectory-following, and object-manipulation tasks. The results showed that the proposed approach outperformed baselines such as weighted voting and independent single-modality classifiers in accuracy, cross-session and cross-subject generalization, and noise robustness, while achieving a measurable reduction in end-to-end response latency, indicating that an integrated semantic alignment–adaptive fusion pipeline enhanced stable outputs and robustness in multimodal interaction. The unified semantic alignment model tailored to BCI–VR can be used for establishing an integrated engineering workflow spanning synchronization, robust representation, and adaptive fusion, and for providing transferable evaluation metrics and application paradigms that offer methodological and technical references for scenarios such as rehabilitation training, virtual education, and intelligent control. Full article
20 pages, 12797 KB  
Article
Target Speaker Extraction with Cross-Correlation for Complex Spectra and Dual Post-Refinements
by Sangwook Han, Seonggyu Lee and Jong Won Shin
Appl. Sci. 2026, 16(13), 6420; https://doi.org/10.3390/app16136420 - 26 Jun 2026
Viewed by 446
Abstract
Target speaker extraction (TSE) aims to isolate speech spoken by a target speaker out of a mixture using speaker information in an enrollment utterance. Recently, several methods have been proposed that exploit the relationship between the enrollment utterance and the input mixture using [...] Read more.
Target speaker extraction (TSE) aims to isolate speech spoken by a target speaker out of a mixture using speaker information in an enrollment utterance. Recently, several methods have been proposed that exploit the relationship between the enrollment utterance and the input mixture using cross-attention, without extracting speaker embeddings from the enrollment. Previous approaches applied the cross-attention to the encoded representations or to the real and imaginary parts of the compressed spectrograms separately, which may not have a physical meaning. In this paper, we propose a two-stage TSE method with a physically interpretable modified cross-attention block and a dual post-refinement structure. In the first stage, the attention weights to fuse the enrollment and mixture are derived from the cross-correlation between the complex spectra for the two signals in a form analogous to the phase-sensitive mask. The fused features along with the mixture features were subsequently fed into a speech extraction network to obtain a coarsely extracted target speech. The second stage consists of two parallel branches, where one branch refines the first-stage output using the enrollment in a similar way to the first stage, and the other utilizes the mixture to complement possibly attenuated target speech. In addition, the low-dimensional speaker embeddings extracted from the enrollment and the first-stage output are incorporated into the second stage to exploit the speaker discriminability. Experimental results show that the proposed method consistently outperformed existing TSE methods on the Libri2Mix dataset under both clean and noisy conditions, in terms of speech quality, speech intelligibility, and signal distortion measures. Full article
Show Figures

Figure 1

35 pages, 4618 KB  
Article
Design of an Iterative Cross-Modal and Context-Aware Deep Analytical Framework for Hate Speech and Fake Post Detection on Social Media Sets
by Rakesh Bharati, Jyoti Bharti and Vasudev Dehalwar
Appl. Sci. 2026, 16(13), 6419; https://doi.org/10.3390/app16136419 - 26 Jun 2026
Viewed by 457
Abstract
There is an enormous rise in the amount of user-generated content on social media. That makes it easier for hateful and fake messages to spread, and threatens both societal stability and public trust in institutions. Most of current solutions have fundamental limitations due [...] Read more.
There is an enormous rise in the amount of user-generated content on social media. That makes it easier for hateful and fake messages to spread, and threatens both societal stability and public trust in institutions. Most of current solutions have fundamental limitations due to modal limitations (i.e., each solution only uses one type of data at a time), lack of user context integration, poor synchronization across different types of data, and poor resilience to manipulation by adversaries. As a result, most solutions are subject to compound loss in terms of their ability to generalize well, classify correctly, or remain reliable when deployed in real-world environments. To address all of the above challenges, we propose a comprehensive and modular analytical framework consisting of five interconnected components that integrate contextual representation learning, multimodal semantic alignment, graph-based propagation modeling, adaptive inference, and consistency validation for hate speech and fake post detection. First is our Context-Driven Social Vector Extraction methodology, which provides enriched contextual embeddings by extracting and combining text-based metadata, image-based metadata, temporal metadata, and behavioral metadata. We use those embeddings in our second module, Multimodal Label Fusion via Mutual Co-Attention (CMF-MCA). Our CMF-MCA module incorporates two transformers with co-attention mechanisms that can mutually annotate text and images. In our third methodology, Semantic Propagation Graph for Hate and Fake Correlation (SPG-HFC), we implement a relational graph attention mechanism that captures both the influence of semantics and how communities propagate information about hate and fake posts. The fourth module, Adaptive Modality Routing via Reinforcement (AMR-R), routes based on the modality of the input and whether the input is simple enough to be classified using machine learning or complex enough to require deep learning. Finally, our Counterfactual Consistency Validation Engine (CCVE) is used after prediction to validate that the model’s predictions are consistent with the output data by creating counterfactuals and validating them. Therefore, in addition to improving the overall accuracy of hate speech and fake post detections, our proposed framework also improves its scalability and inference reliability. Additionally, because our framework allows multimodal classifications that include both context and behavior, it enables the scalable and trustworthy development of content moderation systems. Full article
Show Figures

Figure 1

23 pages, 1105 KB  
Article
Leveraging Label-Attention Networks and POS Tagging for Generating Chinese Cloze Questions
by Yanyang Hou, Shufeng Xiong and Yang Li
Algorithms 2026, 19(6), 501; https://doi.org/10.3390/a19060501 - 22 Jun 2026
Viewed by 364
Abstract
Chinese cloze question generation for educational assessments requires identifying gap phrases that accurately reflect key knowledge points, posing significant challenges to automated systems. We observe that the syntactic boundaries revealed by part-of-speech (POS) tags closely align with the semantic boundaries of target gap [...] Read more.
Chinese cloze question generation for educational assessments requires identifying gap phrases that accurately reflect key knowledge points, posing significant challenges to automated systems. We observe that the syntactic boundaries revealed by part-of-speech (POS) tags closely align with the semantic boundaries of target gap phrases. Motivated by this observation, we propose a multi-task learning framework in which gap phrase identification serves as the primary task and POS tagging as a complementary auxiliary task. The two tasks share a common BERT-BiLSTM encoder, enabling mutual reinforcement of both syntactic and semantic representations through joint training. To further capture the interaction between label semantics and contextual word representations, we introduce a label-attention mechanism that models dependencies between the global word sequence and candidate label embeddings. Additionally, we construct a refined POS tag subset by excluding categories whose boundaries show no alignment with gap phrase boundaries, thereby strengthening the correspondence between the two tasks. Evaluated on a real-world dataset of 20.5K questions spanning five academic disciplines, our method achieves an F1 score of 65.85%, with a Recall of 67.79%, representing improvements of 2.12% and 4.35% over the prior state-of-the-art, respectively. These results demonstrate that exploiting the alignment between syntactic and semantic structures through joint learning is effective for generating educationally meaningful fill-in-the-blank questions. Full article
(This article belongs to the Special Issue Deep Learning Methods and Applications)
Show Figures

Figure 1

25 pages, 478 KB  
Article
A CEFR-Graded Lexicon and Morphology-Aware Benchmarks for Kazakh Lexical Complexity Prediction
by Gulnur Yerkebulan, Akerke Akanova, Zhantore Galymzhan and Nazira Ospanova
Technologies 2026, 14(6), 346; https://doi.org/10.3390/technologies14060346 - 9 Jun 2026
Viewed by 443
Abstract
Graded lexical resources aligned with the Common European Framework of Reference for Languages (CEFR) and lexical complexity prediction remain limited for low-resource Turkic languages, and the extent to which existing predictive models generalize to agglutinative morphology is unresolved. We introduce the first CEFR-graded [...] Read more.
Graded lexical resources aligned with the Common European Framework of Reference for Languages (CEFR) and lexical complexity prediction remain limited for low-resource Turkic languages, and the extent to which existing predictive models generalize to agglutinative morphology is unresolved. We introduce the first CEFR-graded lexicon for Kazakh, containing 4561 lemma–part-of-speech (POS) entries across A1–C1, and use it to test whether explicit morphology improves lexical complexity prediction. We compare handcrafted morphological features, XLM-RoBERTa contextual embeddings, and fusion models that combine both signal types on held-out CEFR classification. Our best model, a gated fusion of contextual embeddings with morphological features, achieves a macro-averaged F1 score of 0.360 and a mean absolute error of 1.125 on the held-out test set. Morphology provides useful information beyond character-level cues, contextual representations are strong on their own, and combining them yields the best supervised performance for this task. The paper therefore contributes a new CEFR resource for Turkic languages and evidence that morphology-aware modeling is useful for Kazakh lexical difficulty prediction. The results support Sustainable Development Goal 4 (Quality Education) by enabling objective assessment of learning-material complexity and adaptive Kazakh language learning. The derived lexicon and code are publicly available. Full article
Show Figures

Figure 1

24 pages, 1730 KB  
Article
An Unsupervised Subspace Weighting Co-Clustering Framework for Hate Speech Detection Patterns in Social Media
by Maya Sultan ALGhafri, Imran Khan and Abdelhamid Abdesselam
AI 2026, 7(6), 204; https://doi.org/10.3390/ai7060204 - 4 Jun 2026
Viewed by 587
Abstract
The exponential growth of social media has revolutionized global communication, enabling instant idea exchange and transforming information sharing into a worldwide phenomenon while simultaneously accelerating the spread of abusive and hateful content that threatens online harmony and poses a serious risk to online [...] Read more.
The exponential growth of social media has revolutionized global communication, enabling instant idea exchange and transforming information sharing into a worldwide phenomenon while simultaneously accelerating the spread of abusive and hateful content that threatens online harmony and poses a serious risk to online community integrity and public trust. Although supervised deep learning approaches achieve impressive accuracy for hate speech detection, they remain fundamentally reliant on extensive annotated corpora, and their lack of interpretability makes them insufficient for transparent and scalable real-world hate speech detection. This study presents a category-oriented unsupervised architecture for English hate-speech detection and classification that substantially reduces reliance on large labeled datasets by requiring only minimal supervision (10% of labels for post hoc cluster interpretation), ensuring transparency and a high degree of semantic interpretability. We introduce an unsupervised Subspace Weighting Co-Clustering framework that uses HateBERT-driven contextual embeddings, enabling simultaneous interpretable feature weighting and semantic understanding for robust hate-speech detection. The obtained embeddings are further structured using the Subspace Weighting Co-Clustering approach, which enables the unsupervised discovery of latent subspaces and the organization of tweets into semantically coherent hate categories. The comprehensive evaluation shows that the framework achieves superior accuracy over existing methods, providing a more robust and effective mechanism for digital platforms to identify and mitigate hate speech and promote safer online interactions. Full article
Show Figures

Figure 1

Back to TopTop