Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (75)

Search Parameters:
Keywords = word structure retrieval

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
32 pages, 2960 KB  
Article
When AI Gets It Wrong: Hallucinations and Trust Recalibration in E-Commerce Using a Sequential Mixed-Methods Approach
by Sayyed Khawar Abbas, Hafiz Muhammad Junaid and Aseel Smerat
J. Theor. Appl. Electron. Commer. Res. 2026, 21(8), 277; https://doi.org/10.3390/jtaer21080277 - 17 Aug 2026
Viewed by 627
Abstract
Generative AI shopping assistants are becoming a primary touchpoint for online consumers, yet they often produce convincing but inaccurate content, a phenomenon known as AI hallucination that poses an under-studied risk to consumer trust in e-commerce. This study explains that phenomenon by building [...] Read more.
Generative AI shopping assistants are becoming a primary touchpoint for online consumers, yet they often produce convincing but inaccurate content, a phenomenon known as AI hallucination that poses an under-studied risk to consumer trust in e-commerce. This study explains that phenomenon by building and testing a moderated mediation model grounded in Expectation Violation Theory, Epistemic Vigilance Theory, and Algorithmic Trust Repair Theory. A sequential, exploratory mixed-methods design was used: a qualitative phase identified the dimensions and configurational pathways of consumer trust withdrawal using the Gioia methodology and fuzzy-set Qualitative Comparative Analysis, and a subsequent large-scale quantitative phase tested and refined the resulting model across a multi-country European sample using partial least squares structural equation modeling and Necessary Condition Analysis. The results show that exposure to hallucinations triggers expectation violation, activating epistemic vigilance and reducing perceived AI competence; this sequence drives trust recalibration, reflected in lower continued-use and purchase intentions and greater negative word-of-mouth. AI literacy, prior trust, and transparency cues significantly moderate these relationships, and structural trust repair mechanisms, namely retrieval-augmented generation and uncertainty disclosure, prove more effective than purely communicative repair strategies. Theoretically, this study advances a dynamic account of trust recalibration in AI-mediated commerce; practically, it offers concrete guidance for platform design and regulatory policy under the EU AI Act. Full article
(This article belongs to the Special Issue AI-Enabled Marketing and Information Dynamics)
Show Figures

Figure 1

27 pages, 3976 KB  
Article
An LLM-Based Framework for the Automatic Generation of SysML Models
by Baoran An, Tao Lei and Guangtai Tian
Sensors 2026, 26(16), 5133; https://doi.org/10.3390/s26165133 - 13 Aug 2026
Viewed by 459
Abstract
Model-based systems engineering (MBSE) takes Systems Modeling Language (SysML) as the industrial standard modeling language, yet cloud Large Language Model (LLM)-based SysML generation faces limited domain data, model hallucinations, high hardware cost and confidential data leakage risks. This paper builds a 914-sample SysML [...] Read more.
Model-based systems engineering (MBSE) takes Systems Modeling Language (SysML) as the industrial standard modeling language, yet cloud Large Language Model (LLM)-based SysML generation faces limited domain data, model hallucinations, high hardware cost and confidential data leakage risks. This paper builds a 914-sample SysML PlantUML corpus and proposes a fully offline lightweight framework based on Qwen2.5-Coder-7B-Instruct, integrating 4-bit NF4 Quantized Low-Rank Adaptation (QLoRA) fine-tuning, vector-free Jaccard same-diagram reference retrieval and a three-round syntax correction loop. PlantUML executes syntax parsing while Graphviz only renders layouts. Tested on 131 samples covering five structural and behavioral SysML v1 diagram types, the plain-prompt baseline achieves word-set semantic F1 of 52.94%, and the retrieval-enhanced variant lifts the zero-retry syntax pass rate from 92.37% to 99.28%, with F1 slightly dropping to 50.64%. Running fully local without cloud data transmission, this pipeline offers a privacy-safe lightweight solution for SysML PlantUML modeling and does not support SysML-exclusive requirement or parametric diagrams. Full article
Show Figures

Figure 1

23 pages, 3142 KB  
Systematic Review
Sustainable Production System for the Cultivation and Processing of Coffea arabica L. From Oaxaca, Mexico: A Systematic Review
by Jesica Ariadna Jiménez-Mendoza, Magdaleno Caballero-Caballero, Fernando Chiñas-Castillo, Luis Humberto Robledo-Taboada, Luis Eduardo García-Mayoral, Rafael Alavez-Ramírez, José Luis Montes-Bernabe and María Eugenia Silva-Rivera
Sustainability 2026, 18(16), 8164; https://doi.org/10.3390/su18168164 - 10 Aug 2026
Viewed by 334
Abstract
The sustainability of Coffea arabica L. production in Oaxaca, Mexico, is increasingly threatened by climate change, biodiversity loss, and pest pressure, along with other factors that undermine the responsiveness of small-scale producers in the state, such as trade restrictions and coffee-related regulatory frameworks. [...] Read more.
The sustainability of Coffea arabica L. production in Oaxaca, Mexico, is increasingly threatened by climate change, biodiversity loss, and pest pressure, along with other factors that undermine the responsiveness of small-scale producers in the state, such as trade restrictions and coffee-related regulatory frameworks. These challenges are particularly acute in mountainous regions like Oaxaca, where coffee cultivation plays a central role in rural livelihoods, cultural identity, and territorial development. Studies demonstrate that localized models are vital for rural improvement, reducing the carbon footprint, and maintaining economic viability. Despite extensive research on agronomic, environmental, and market factors, sustainability strategies for coffee production often remain fragmented and insufficiently integrated. This systematic review used the Scopus Review database, searching by title, abstract, and keywords such as “sustainable coffee production” from 2015 to 2026. A total of 1085 documents were retrieved and processed using VOSviewer software to generate a bibliographic map in which frequently used words are grouped by color to show their relationships. Using Oaxaca as a regional case study, the article synthesizes the key factors influencing coffee productivity and quality, examines the main socio-environmental challenges, and proposes a conceptual framework that integrates agroecosystem management with socioeconomic processes under external climate and market pressures. The proposed framework highlights the central role of agroforestry systems and ecosystem services in improving climate resilience, conserving biodiversity, and supporting quality-oriented value chains. These processes generate feedback loops that influence farmers’ livelihoods, food security, and territorial sustainability. While grounded in the context of Oaxaca, the socio-ecological review and conceptual framework presented here are applicable to other Arabica-producing regions facing similar challenges, providing a structured basis for future research, policy design, and integrated sustainability strategies. Full article
(This article belongs to the Section Sustainable Urban and Rural Development)
Show Figures

Graphical abstract

25 pages, 4561 KB  
Article
Stream-Aware Vocabulary Demands in Singapore Secondary Tamil Textbooks: A Morphology- and Multiword-Unit-Sensitive Corpus Analysis, with a Textbook-Faithful GenAI Item Benchmark
by Kingston Palthamburaj and Mercedes Premalatha Ramesh
Educ. Sci. 2026, 16(8), 1270; https://doi.org/10.3390/educsci16081270 - 10 Aug 2026
Viewed by 285
Abstract
Vocabulary learning research increasingly calls for curriculum-aligned evidence and methods that reflect how vocabulary is selected, sequenced, and assessed in real educational systems. However, stream-differentiated curricular pathways may entail distinct lexical demands that are rarely quantified for morphologically rich languages such as Tamil, [...] Read more.
Vocabulary learning research increasingly calls for curriculum-aligned evidence and methods that reflect how vocabulary is selected, sequenced, and assessed in real educational systems. However, stream-differentiated curricular pathways may entail distinct lexical demands that are rarely quantified for morphologically rich languages such as Tamil, where surface-form variation and digitization artifacts complicate textbook profiling. This study reports a stream-aware corpus analysis of Singapore Secondary 3 Tamil materials (Express, Higher Tamil, Normal (Academic), Normal (Technical)) using a transparent entry-extraction workflow from list-structured PDF resources. We introduce a morphology- and multiword-unit (MWU)–sensitive profiling approach that (i) operationalizes the unit of analysis as list entries (single-word or multiword), (ii) applies conservative Unicode normalization to reduce Tamil PDF segmentation artifacts, and (iii) avoids reliance on opaque lemmatization, treating results as auditable derived indicators rather than definitive lexical-family counts. The corpus comprises 36 pages and yields 1911 extracted entries (1784 unique in the union of streams; 1871 unique when counted within streams), with MWUs comprising 47.7–57.4% of extracted entries and typical MWUs spanning 2–4 words. Between-stream overlap of unique entries is low (pairwise Jaccard 0.010–0.043), and within-stream recycling across pages is limited (0.9–2.1% of within-stream unique entries repeated on ≥2 pages), suggesting substantial stream-specific instructional emphasis and limited built-in recycling in vocabulary-list components. As an innovation component aligned with emerging GenAI practices, we specify and pilot a textbook-faithful item-generation benchmark in which vocabulary practice items are generated under retrieval constraints and evaluated via transparent, classroom-oriented quality screens (faithfulness/grounding, morphological correctness, appropriateness, distractor plausibility, and bilingual gloss quality). The paper concludes with practical outputs for stream-aligned scope-and-sequence planning and a reproducible pipeline that can be extended as additional materials become available. Full article
Show Figures

Figure 1

18 pages, 4630 KB  
Article
Real-Time Sign Language Interpretation via Customized Sign Language Gloves and Motion Retrieval
by Chien-Hua Chen, Chih-Yuan Yao and Shih-Hsuan Hung
Sensors 2026, 26(15), 4884; https://doi.org/10.3390/s26154884 - 3 Aug 2026
Viewed by 380
Abstract
A sign language interpretation system aims to translate sign gestures into spoken or written language in real time, enabling signers and non-signers to communicate in their familiar linguistic forms. However, vision-based approaches suffer from hand occlusion, lighting variability, and complex backgrounds, while Deep [...] Read more.
A sign language interpretation system aims to translate sign gestures into spoken or written language in real time, enabling signers and non-signers to communicate in their familiar linguistic forms. However, vision-based approaches suffer from hand occlusion, lighting variability, and complex backgrounds, while Deep Neural Network (DNN)-based methods incur heavy computational costs that hinder real-time use on resource-constrained platforms. In this paper, we propose sign language gloves and a lightweight motion retrieval method for real-time sign language interpretation that runs on mobile devices and embedded systems. The sign language gloves integrate flex sensors, an inertial measurement unit (IMU), and pressure sensors to accurately capture gesture features, including finger bending angles, hand orientation, movement trajectories, and fingertip contacts with body parts, enabling recognition of touch-based gestures. For the motion retrieval method, we build a comprehensive gesture dataset with the gloves and perform feature analysis on each sign language gesture to avoid redundant information in the dataset. During interpretation, our system employs a feature-labeling mechanism to ensure gesture distinguishability and a gesture retrieval algorithm to evaluate movement continuity and similarity. This allows the system to identify corresponding feature labels and consolidate them into complete sign language vocabulary entries. The proposed motion retrieval method is characterized by low computational complexity and a well-defined data structure. This makes it suitable for integration into embedded systems, offering real-time performance and high portability for practical deployment. In our experiments, the proposed system achieved an average recognition accuracy of 92% on a gesture dataset covering 300 sign language words. Full article
(This article belongs to the Section Biomedical Sensors)
Show Figures

Figure 1

16 pages, 15271 KB  
Article
Digitization and Preservation of Cultural Heritage: Translating the Coptic Scripts on Artifacts
by Argyro Kontogianni, Antreas Kantaros, Theodore Ganetsos, Panagiotis Kousoulis, Melina G. Mouzala and Evangelos Papakitsos
Appl. Sci. 2026, 16(14), 7147; https://doi.org/10.3390/app16147147 - 16 Jul 2026
Viewed by 380
Abstract
Coptic represents the final stage of the Ancient Egyptian language and remains an important component of Christian and Mediterranean cultural heritage. Although several digital resources exist for Coptic textual corpora, Greek-oriented computational tools for the interpretation of Coptic inscriptions on artifacts remain limited. [...] Read more.
Coptic represents the final stage of the Ancient Egyptian language and remains an important component of Christian and Mediterranean cultural heritage. Although several digital resources exist for Coptic textual corpora, Greek-oriented computational tools for the interpretation of Coptic inscriptions on artifacts remain limited. This study presents the design and early implementation of a semi-automated software tool for the computer-assisted translation of Coptic inscriptions into Greek, with optional English support. The tool combines a Coptic–Greek digital dictionary, an interactive character-selection interface, and two dictionary-search strategies: a length/alphabetically structured linear search and a weighted linear search based on expected word frequency. The application is intended to support scholars working with inscriptions on fragile or fragmented cultural heritage objects, where full automation is not realistic and human supervision remains essential. The paper describes the linguistic and material challenges of Coptic inscriptions, the structure of the lexical database, the interface design, and the planned use of Coptic corpora for improving retrieval efficiency. The proposed approach contributes to cultural heritage digitization by offering a practical, expandable, and user-oriented framework for supporting the study, interpretation, and preservation of Coptic inscriptions. Full article
(This article belongs to the Special Issue Application of Digital Technology in Cultural Heritage)
Show Figures

Figure 1

44 pages, 996 KB  
Article
Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition
by Diederik Aerts, Jonito Aerts Arguëlles, Lester Beltran, Suzette Geriente, Roberto Leporini, Massimiliano Sassoli de Bianchi and Sandro Sozzo
Entropy 2026, 28(6), 622; https://doi.org/10.3390/e28060622 - 1 Jun 2026
Cited by 1 | Viewed by 1388
Abstract
We present the results of cognitive tests on conceptual combinations, performed using specific Large Language Models (LLMs) as test subjects. In the first test, performed with ChatGPT (GPT-5.5 Thinking) and Google Gemini Advanced (Gemini 1.5 Pro), we show that Bell’s inequalities are significantly [...] Read more.
We present the results of cognitive tests on conceptual combinations, performed using specific Large Language Models (LLMs) as test subjects. In the first test, performed with ChatGPT (GPT-5.5 Thinking) and Google Gemini Advanced (Gemini 1.5 Pro), we show that Bell’s inequalities are significantly violated, which indicates the presence of a ‘non-classical probability model’ with probabilities that do not satisfy Kolmogorov’s axioms. In the second test, also performed using ChatGPT and Gemini, we identify the presence of ‘Bose–Einstein statistics’, rather than the intuitively expected ‘Maxwell–Boltzmann statistics’, in the distribution of the words contained in large-size texts. Interestingly, these findings mirror the results previously obtained in both cognitive tests with human participants and information retrieval tests on large corpora. Taken together, they point to the ‘systematic emergence of non-classical quantum-like structures in conceptual-linguistic domains’, regardless of whether the cognitive agent is human or artificial. Although LLMs are classified as neural networks for historical reasons, we believe that a more essential form of knowledge organization takes place in the distributive semantic structure of vector spaces built on top of the neural network. It is this meaning-bearing structure that lends itself to a phenomenon of evolutionary convergence between human cognition and language, slowly established through biological evolution, and LLM cognition and language, emerging much more rapidly as a result of self-learning and training. We analyze various aspects and examples that contain evidence supporting the above hypothesis. We also advance a unifying framework that explains the pervasive quantum organization of meaning that we identify. Full article
(This article belongs to the Section Multidisciplinary Applications)
Show Figures

Figure 1

24 pages, 3010 KB  
Article
Retrieval-Augmented Generation-Based Earth Surface System Association Network Optimization and Data Recommendation
by Jiangbing Sun, Yan Zhang, Longxing Tian, Jiali Li, Miao Tian, Jie Chen, Liufeng Tao and Qinjun Qiu
ISPRS Int. J. Geo-Inf. 2026, 15(5), 199; https://doi.org/10.3390/ijgi15050199 - 2 May 2026
Viewed by 1011
Abstract
The scientific data of the Earth surface system is characterized by multi-source heterogeneity and dynamic correlation, so constructing an efficient data association network and enabling intelligent knowledge services is a hot topic. Nevertheless, confronted with the existing challenges of onerous data acquisition, inadequate [...] Read more.
The scientific data of the Earth surface system is characterized by multi-source heterogeneity and dynamic correlation, so constructing an efficient data association network and enabling intelligent knowledge services is a hot topic. Nevertheless, confronted with the existing challenges of onerous data acquisition, inadequate precision of data recommendation, excessive time and labor consumption, as well as insufficient semantic reasoning in intelligent question-and-answer (Q&A) systems, we propose an intelligent framework that integrates dynamic optimization and retrieval-augmented generation (RAG) technology to address the problems of strong subjectivity in the setting of edge weight thresholds in association networks and insufficient semantic inference in intelligent Q&A. First, a multidimensional association network is constructed based on metadata features, redundant edge pruning is achieved through dynamic threshold analysis, and key nodes are identified by combining complex network centrality theory to optimize network structure and storage efficiency. Secondly, the RAG-based intelligent Q&A model is designed to transform the association triples into a paragraph-based knowledge base, generate a domain Q&A dataset using a large language model GPT-4o, and fine-tune the word embedding model to improve the semantic representation accuracy. Experiments show that the number of network edges is reduced by about 70% after optimization, and the node importance analysis accurately identifies key data nodes; the fine-tuned model improves each index by 6% on average in the retrieval task, and the Q&A system significantly outperforms the traditional method in terms of indexes such as relevance and completeness. This study provides innovative solutions for the intelligent service of scientific data in Earth surface systems and promotes the deep integration of association networks and generative AI. Full article
(This article belongs to the Special Issue LLM4GIS: Large Language Models for GIS)
Show Figures

Figure 1

23 pages, 6022 KB  
Review
Research Trends on Invasive Marine Species in the Mediterranean: A Bibliometric and Topic Modeling Analysis
by Dimitris Klaoudatos, Stefanos Gkourtsoulis, Dimitris Pafras and Alexandros Theocharis
Oceans 2026, 7(3), 37; https://doi.org/10.3390/oceans7030037 - 24 Apr 2026
Viewed by 1615
Abstract
The Mediterranean Sea is both a global biodiversity hotspot and the world’s most heavily invaded marine region, where non-indigenous species arrivals are accelerating under intensifying shipping, Suez Canal traffic, aquaculture, and climate warming. Yet, despite rapidly growing research activity, a comprehensive synthesis of [...] Read more.
The Mediterranean Sea is both a global biodiversity hotspot and the world’s most heavily invaded marine region, where non-indigenous species arrivals are accelerating under intensifying shipping, Suez Canal traffic, aquaculture, and climate warming. Yet, despite rapidly growing research activity, a comprehensive synthesis of the scientific literature on Mediterranean marine invasions has been lacking. This study provides the first Mediterranean-wide combined bibliometric and topic-modeling analysis of invasive marine species research, using 3521 unique documents retrieved from Scopus and Web of Science. We quantify temporal growth in publications and citations, map the conceptual structure of the field through co-citation, co-word, and topic modeling, and reveal pronounced regional and thematic biases. Latent Dirichlet Allocation resolves 13 coherent topics, dominated by first records of non-native species, invasive macroalgae, alien species diversity, and ecological impacts, with strong signals for Lessepsian migration and climate-driven range shifts, particularly in the Eastern Mediterranean. Spatial and thematic analyses reveal pronounced regional biases, with invasion hotspots in the Aegean and Levantine seas contrasted by comparatively sparse coverage of western and central sub-basins, and notable gaps in predictive modeling and socioeconomic assessments. The results underscore the need to rebalance effort toward under-studied regions and themes, while leveraging existing collaboration networks and methodological advances to support MSFD (Marine Strategy Framework Directive) implementation, International Maritime Organization (IMO) instruments, and broader ecosystem-based management. The reproducible framework presented here offers a baseline for periodically tracking research evolution and guiding adaptive, transboundary governance of Mediterranean marine bio-invasions. Full article
Show Figures

Figure 1

45 pages, 6682 KB  
Article
A Multidimensional MIR Analysis of Acoustic, Linguistic and Cultural Gaps Between Maskandi and Western Music Genres
by Absolom Muzambi, Tebatso Gorgina Moape and Bester Chimbo
Appl. Sci. 2026, 16(8), 3802; https://doi.org/10.3390/app16083802 - 14 Apr 2026
Viewed by 1142
Abstract
Contemporary Music Information Retrieval (MIR) and Natural Language Processing (NLP) systems are increasingly applied to diverse musical traditions, yet they are largely grounded in Western musical and linguistic assumptions. This study examines whether commonly used MIR features and multilingual NLP models adequately represent [...] Read more.
Contemporary Music Information Retrieval (MIR) and Natural Language Processing (NLP) systems are increasingly applied to diverse musical traditions, yet they are largely grounded in Western musical and linguistic assumptions. This study examines whether commonly used MIR features and multilingual NLP models adequately represent the acoustic, linguistic, and cultural structures of Maskandi music in comparison to Western music and identifies where representational gaps and biases arise. A multidimensional framework was employed, comprising acoustic and structural MIR analysis, linguistic and semantic lyrical analysis, and bias analysis. A curated dataset of 60 recordings and corresponding lyrics was analysed using rhythm and beat features, pitch contour measures, structural self-similarity, timbre embeddings, semantic similarity, lexical diversity, metaphor density, topic modelling, multilingual embeddings, and dataset-level audits. The results reveal systematic representational failures: beat tracking showed lower median IOI coefficient of variation for Maskandi (0.028) versus Western music (0.040, p = 0.0199) yet exhibited greater algorithmic instability, tempo averaged 131.16 BPM versus 111.69 BPM (p = 0.000262), pitch glide proportions were significantly higher in Maskandi (0.34 vs. 0.16), on-beat energy ratios differed substantially (2.26 vs. 1.19, p < 0.0000007), semantic similarity revealed high intra-genre coherence for Maskandi (0.73) versus Western (0.25), metaphor density approached zero in Maskandi versus up to 7 per 100 words in Western lyrics, topic modeling produced two compact clusters for Maskandi versus 6 dispersed clusters for Western, timbre embeddings achieved a 0.405 silhouette score, dataset audits revealed 0% Maskandi representation across seven major MIR corpora with African traditions comprising <3%. The study concludes that statistical separability does not imply representational adequacy and highlights the need for culturally grounded MIR and NLP representations to support diverse musical traditions. Full article
(This article belongs to the Special Issue Large Language Models and Knowledge Computing)
Show Figures

Figure 1

24 pages, 4909 KB  
Article
UniTriM: Unified Text–Image–Video Retrieval via Multi-Granular Alignment and Feature Disentanglement
by Yangchen Wang, Yan Hua, Yingyun Yang and Wenhui Zhang
Electronics 2026, 15(7), 1424; https://doi.org/10.3390/electronics15071424 - 30 Mar 2026
Viewed by 651
Abstract
With the proliferation of multimodal content on social media, creators increasingly require tools that can retrieve both images and videos relevant to a single textual query. However, existing cross-modal retrieval methods are typically confined to binary (text–image or text–video) settings and struggle with [...] Read more.
With the proliferation of multimodal content on social media, creators increasingly require tools that can retrieve both images and videos relevant to a single textual query. However, existing cross-modal retrieval methods are typically confined to binary (text–image or text–video) settings and struggle with fine-grained semantic alignment and spatiotemporal information imbalance. To address this issue, we propose UniTriM, a unified framework for text–image–video joint retrieval. First, UniTriM supports concurrent retrieval of semantically relevant images and videos given one textual input. To overcome the scarcity of text–image–video triplet data, we introduce a self-attention-based keyframe selection strategy that converts existing text–video datasets into triplet format. Second, we design a multi-granularity similarity alignment module that captures hierarchical semantics by modeling patch–frame–video and word–triple–sentence structures and jointly optimizes intra- and cross-granularity alignments to enhance fine-grained cross-modal correspondence. Third, to alleviate the inherent spatiotemporal information imbalance between static images and video-aligned text descriptions, we introduce a feature disentanglement module that disentangles spatial-related features from text and aligns them explicitly with image representations. Experiments conducted on three benchmark datasets MSR-VTT, MSVD, and DiDeMo demonstrate that UniTriM achieves state-of-the-art performance on joint retrieval tasks. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

19 pages, 992 KB  
Article
Hybrid Music Similarity with Hypergraph and Siamese Network
by Sera Kim, Youngjun Kim, Jaewon Lee and Dalwon Jang
Big Data Cogn. Comput. 2026, 10(3), 96; https://doi.org/10.3390/bdcc10030096 - 21 Mar 2026
Viewed by 880
Abstract
This paper proposes a novel method for measuring music similarity. Existing music similarity measurements have often been used for music appreciation, but this paper proposes a method for measuring the similarity between music samples which are used for music production. Conventional music recommendation [...] Read more.
This paper proposes a novel method for measuring music similarity. Existing music similarity measurements have often been used for music appreciation, but this paper proposes a method for measuring the similarity between music samples which are used for music production. Conventional music recommendation approaches often rely on either metadata-based similarity or audio-based feature similarity in isolation, which limits their effectiveness in sample-based recommendation scenarios where both compositional context and acoustic characteristics are important. To address this limitation, the proposed framework combines a hypergraph-based information similarity module with a feature-based similarity module learned using Siamese networks and triplet loss. In the information-based module, metadata attributes such as beats per minute (BPM), genre, chord, key, and instrument are modeled as vertices in a hypergraph, and Random Walk–Word2Vec embeddings are learned to capture structural relationships between music samples and their attributes. In parallel, the feature-based module employs vertex-specific Siamese networks trained on instrument and key classification tasks to learn perceptual similarity directly from audio signals. The two modules are trained independently and jointly utilized at the recommendation stage to provide attribute-specific similarity results for a given query sample. Results show that the proposed system achieves high Precision@k across multiple attributes and forms stable similarity structures in the embedding space, even without relying on user interaction data. These results reflect embedding consistency evaluated over the entire dataset where training and retrieval are performed on the same sample pool, rather than generalization to unseen samples. These results demonstrate that the proposed hybrid framework effectively captures both structural and perceptual similarity among music samples and is well suited for sample-based music recommendation in music production environments. Full article
Show Figures

Figure 1

37 pages, 4447 KB  
Article
A Citation-Based Main Path Analysis of Tinnitus Research (1984–2025): Knowledge Evolution, Thematic Clusters, and Emerging Research Directions
by Tang-Min Hsieh, Kai-Ying Chen and Hsin-Yu Hsieh
Appl. Sci. 2026, 16(5), 2474; https://doi.org/10.3390/app16052474 - 4 Mar 2026
Viewed by 1522
Abstract
Over the past four decades, tinnitus research has grown into a highly interdisciplinary field spanning auditory science, neuroscience, psychology, and clinical medicine. Yet how knowledge across subfields has been inherited, diversified, and integrated over time still lacks traceable structural evidence. To address this [...] Read more.
Over the past four decades, tinnitus research has grown into a highly interdisciplinary field spanning auditory science, neuroscience, psychology, and clinical medicine. Yet how knowledge across subfields has been inherited, diversified, and integrated over time still lacks traceable structural evidence. To address this gap and move beyond frequency-oriented reviews and bibliometric studies that mainly report “hot topics” and prolific contributors, the present study reconstructs the intellectual evolution of tinnitus research (1984–2025) using citation-network-based main path analysis (MPA). From the Web of Science Core Collection, 6584 records were initially retrieved, of which 6354 formed a mutually linked core citation network (96.5%), indicating high coverage and analyzability. SPLC (Search Path Link Count)–weighted MPA was applied to extract global and key-route main paths capturing dominant knowledge trajectories and major branches. Cluster and co-word analyses were then integrated to delineate seven evolutionary stages and five major thematic clusters. This framework identifies bridging works and turning points and reveals how emerging lines—neuromodulation, implant-related treatments, and digital/telehealth CBT—branch from and later converge with established neurobiological and psychological pathways rather than appearing in isolation. Overall, the field has progressed from early psychoacoustics and spontaneous otoacoustic emissions through cochlear-injury plasticity, central gain, and limbic–auditory network models, and most recently toward mechanism-oriented diagnostics, individualized assessment, and targeted interventions. Full article
Show Figures

Figure 1

29 pages, 10558 KB  
Article
AI-Powered Interpretation of Traditional Village Landscape Language: An Analysis of Xinye Village in Zhejiang, China
by Yanying Liang, Tao Chen and Zizhen Hong
Sustainability 2026, 18(5), 2183; https://doi.org/10.3390/su18052183 - 24 Feb 2026
Viewed by 1100
Abstract
Amidst rapid urbanization and modernization, numerous traditional villages in China face severe challenges, including landscape homogenization and the erosion of their distinctive characteristics. Addressing this issue requires a method capable of systematically identifying, analyzing, and reconstructing both the landscape and its underlying cultural [...] Read more.
Amidst rapid urbanization and modernization, numerous traditional villages in China face severe challenges, including landscape homogenization and the erosion of their distinctive characteristics. Addressing this issue requires a method capable of systematically identifying, analyzing, and reconstructing both the landscape and its underlying cultural features. This study proposes a digital analytical approach that integrates multimodal artificial intelligence with landscape language theory to address the homogenization of cultural landscapes in traditional Chinese villages. Taking Xinye Village in Zhejiang Province as a case study, the research systematically decodes its landscape spatial narratives and underlying cultural genes. This framework systematically deconstructs village landscapes across four levels: “vocabulary, context, grammar, and semantics”. The village image database is first automatically recognized and statistically analyzed by computer vision technology, which extracts 31 core landscape vocabulary items from three main categories and nine subcategories. Second, Retrieval-augmented Generation technology is employed to synthesize from the constructed domain-specific corpus, a natural context structured around Yuhua Mountain and Daofeng Mountain, as well as a cultural context based on ancestral hall order, connected through folk activities, and idealized by farming and reading passed down through generations. Building on this framework, a multimodal model was used to examine the spatial composition and combinatorial laws of landscape features. Six essential dimensions—spatial layout, visual order, element combination, functional relationships, circulation layout, and scale correlations—revealed the spatial grammar of shuikou landscape. Lastly, the semantic values conveyed by the landscape vocabulary were thoroughly analyzed across three dimensions—form, function, and culture—by integrating a knowledge base. This work creates a landscape language atlas of Xinye Village by combining these studies and using a linguistic model of “character-word-sentence-paragraph”. By methodically deciphering the clan’s cultural code of “farming and reading passed down through generations”, this clearly reconstructs the spatial narrative logic from micro-elements to macro-patterns. This research not only advances the study of landscape language in traditional villages from qualitative description toward a systematic, digital, and interpretable paradigm but also provides an operational theoretical and methodological foundation for the in-depth interpretation, conservation, and transmission of traditional village cultural landscapes. Full article
Show Figures

Figure 1

9 pages, 1037 KB  
Proceeding Paper
Hybrid Dictionary–Retrieval-Augmented Generation–Large Language Model for Low-Resource Translation
by Reen-Cheng Wang, Cheng-Kai Yang, Tun-Chieh Yang and Yi-Xuan Tseng
Eng. Proc. 2025, 120(1), 52; https://doi.org/10.3390/engproc2025120052 - 5 Feb 2026
Cited by 1 | Viewed by 1568
Abstract
The rapid decline of linguistic diversity, driven by globalization and technological standardization, presents significant challenges for the preservation of endangered languages, many of which lack sufficient parallel corpora for effective machine translation. Conventional neural translation models perform poorly in such contexts, often failing [...] Read more.
The rapid decline of linguistic diversity, driven by globalization and technological standardization, presents significant challenges for the preservation of endangered languages, many of which lack sufficient parallel corpora for effective machine translation. Conventional neural translation models perform poorly in such contexts, often failing to capture semantic precision, grammatical complexity, and culturally specific nuances. This study addresses these limitations by proposing a hybrid translation framework that combines dictionary-based pre-translation, retrieval-augmented generation, and large language model post-editing. The system is designed to improve translation quality for extremely low-resource languages, with a particular focus on the endangered Paiwan language in Taiwan. In the proposed approach, a handcrafted bilingual dictionary is the first to establish deterministic lexical alignments to generate a symbolically precise intermediate representation. When gaps occur due to missing vocabulary or sparse training data, a retrieval module enriches contextual understanding by dynamically sourcing semantically relevant examples from a vector database. These enriched words are then processed by an instruction-tuned large language model that reorders syntactic structures, inflects verbs appropriately, and resolves lexical ambiguities to produce fluent and culturally coherent translations. The evaluation is conducted on a 250-sentence Paiwan–Mandarin dataset, and the results demonstrate substantial performance gains across key metrics, with cosine similarity increasing from 0.210–0.236 to 0.810–0.846, BLEU scores rising from 1.7–4.4 to 40.8–51.9, and ROUGE-L F1 scores improving from 0.135–0.177 to 0.548–0.632. These results corroborate the effectiveness of the proposed hybrid pipeline in mitigating semantic drift, preserving core meaning, and enhancing linguistic alignment in low-resource settings. Beyond technical performance, the framework contributes to broader efforts in language revitalization and cultural preservation by supporting the transmission of Indigenous knowledge through accurate, contextually grounded, and accessible translations. This research demonstrates that integrating symbolic linguistic resources with retrieval-augmented large language models offers a scalable and efficient solution for endangered language translation and provides a foundation for sustainable digital heritage preservation in multilingual societies. Full article
(This article belongs to the Proceedings of 8th International Conference on Knowledge Innovation and Invention)
Show Figures

Figure 1

Back to TopTop