Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (110)

Search Parameters:
Keywords = corpus-based translation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
23 pages, 1261 KB  
Systematic Review
AI-Driven Food Fraud Detection Systems: A Critical Systematic Review of the Detection–Prevention Gap
by Orlando Meneses Quelal, David Pilamunga Hurtado and Marco Burbano Pulles
Foods 2026, 15(18), 3185; https://doi.org/10.3390/foods15183185 - 9 Sep 2026
Abstract
The economic impact of food fraud is difficult to quantify precisely, because fraud is structurally designed to evade detection; available estimates are indirect projections rather than direct forensic accounting and are commonly cited in the range of USD 10–15 billion annually. The integration [...] Read more.
The economic impact of food fraud is difficult to quantify precisely, because fraud is structurally designed to evade detection; available estimates are indirect projections rather than direct forensic accounting and are commonly cited in the range of USD 10–15 billion annually. The integration of artificial intelligence (AI) with analytical instrumentation has generated a rapidly expanding body of research aimed at detecting adulteration, mislabeling, and substitution across food matrices. This systematic review examines the extent to which AI-assisted instrumental technologies contribute to food fraud prevention (as distinct from laboratory detection) and characterizes the structural factors that constrain real-world translation. A systematic search of the peer-reviewed literature published between 2021 and 2026 yielded 83 eligible records (80 primary studies and 3 review articles) after applying predefined inclusion criteria. Data were extracted into a structured seven-sheet workbook covering study characteristics, instrumental technologies, AI architectures, performance metrics, industrial-validation status, implementation evidence, and methodological quality. The corpus shows consistently high reported analytical accuracy under controlled laboratory conditions (median of extractable classification accuracies ≈ 99–100%; ≥95% in 86% of studies with an extractable value). At the same time, 68 of 83 studies (82%) reported no external validation, no study (0/83) achieved inter-laboratory validation, no study documented routine-monitoring application, and only one study reported testing in a genuine industrial environment. The most frequently featured platforms were NIR spectroscopy and electronic-nose arrays (each featuring in 30/83 studies, frequently in data-fusion combinations), followed by gas-chromatography-based systems (16/83) and hyperspectral imaging (13/83). Classical machine learning predominated (57/83 studies coded as classical ML, with a further 11 hybrid ML/DL designs and 12 deep-learning-only designs). A direct statistical comparison found no significant difference in reported accuracy between classical-ML and deep-learning studies (median 100% vs. 98.2%; Mann–Whitney U test, p = 0.16). A pre-specified test of the hypothesis that high reported accuracy is itself a marker of overfitting was not supported by the corpus: reported accuracy was not negatively associated with external-validation status (Fisher’s exact p = 0.51) or with methodological-quality score (Spearman ρ = 0.15, p = 0.23). Methodological quality was predominantly moderate (49/83 scored 3/5; 22 scored 2/5; 11 scored 4/5; one study scored 5/5), and 19/83 (23%) carried a high risk of bias. The review’s central observation—a measurable gap between demonstrated laboratory detection and evidenced real-world prevention—is well supported by the deployment, inter-laboratory, and routine-monitoring data. We deliberately separate this strongly evidenced conclusion from weaker inferences (e.g., the overfitting hypothesis) that the corpus cannot currently establish, and we outline a validation-driven, deployment-oriented research agenda. Full article
(This article belongs to the Section Food Engineering and Technology)
Show Figures

Figure 1

19 pages, 530 KB  
Article
Rethinking Xuanzang’s “New Translation” Style: A Corpus-Based Analysis of Three Chinese Translations of the Vimalakīrti Sutra
by Yanfei Zhao
Religions 2026, 17(9), 1039; https://doi.org/10.3390/rel17091039 - 4 Sep 2026
Viewed by 224
Abstract
In the history of Chinese Buddhist translation, Xuanzang has conventionally been associated with a “new translation” style, indicating a departure from earlier “old” translations. This study reconsiders the nature of this “newness” through a corpus-based comparison of Xuanzang’s translation of the Vimalakīrti Sutra [...] Read more.
In the history of Chinese Buddhist translation, Xuanzang has conventionally been associated with a “new translation” style, indicating a departure from earlier “old” translations. This study reconsiders the nature of this “newness” through a corpus-based comparison of Xuanzang’s translation of the Vimalakīrti Sutra with two earlier versions by Kumārajīva and Zhi Qian. The analysis reveals greater textual continuity between Xuanzang and Kumārajīva, with a higher LCS retention rate (61.68%) than that between Zhi Qian and Kumārajīva (48.71%). A symmetric LCS check preserves the direction of this result but substantially narrows the contrast. That said, textual continuity should not be equated with direct borrowing. Xuanzang often reworked inherited formulations through expansion and lexical revision, producing a translation marked by greater semantic and conceptual precision. This study argues that Xuanzang’s translation strategies were jointly shaped by target cultural norms, ideological and textual contexts, and his own translation poetics. The “newness” of his translation thus lay less in a complete break with earlier versions than in the selective preservation, expansion, and reconfiguration of inherited textual resources. Full article
Show Figures

Figure 1

26 pages, 1453 KB  
Review
Silk-Derived Antibacterial Hydrogels: Material Identity, Mechanistic Evidence, and Translation
by Hongmei Wang, Bingbing Xia, Yanlin Zhang and Xiaojuan Mi
Gels 2026, 12(9), 809; https://doi.org/10.3390/gels12090809 - 3 Sep 2026
Viewed by 254
Abstract
Silk fibroin (SF)- and silk sericin (SS)-based antibacterial hydrogels are increasingly engineered as local antimicrobial platforms, yet cross-study interpretation is limited by inconsistent material reporting and by conflation of bacterial inhibition with tissue repair. We performed a structured evidence-mapping and critical synthesis of [...] Read more.
Silk fibroin (SF)- and silk sericin (SS)-based antibacterial hydrogels are increasingly engineered as local antimicrobial platforms, yet cross-study interpretation is limited by inconsistent material reporting and by conflation of bacterial inhibition with tissue repair. We performed a structured evidence-mapping and critical synthesis of a frozen 2020–July 2026 corpus of 94 references. The original 46-record core map was re-audited at the original-article level: 43 full-text-verified, non-retracted primary studies were retained for detailed evidence grading, 2 records available only at abstract/database level were retained descriptively but not graded, and 1 subsequently retracted study was excluded from quantitative synthesis. Among the 43 graded studies, metal-ion/nanozyme/catalytic systems were most common (12/43, 27.9%), followed by release-mediated (11/43, 25.6%), multimodal (9/43, 20.9%), contact-active/anti-adhesive (6/43, 14.0%), and light-responsive systems (5/43, 11.6%). Sixteen studies (37.2%) used deliberately infected animal models, whereas only 4 (9.3%) reached a biofilm or adherent-bacteria-level endpoint in the graded map. Biological claim ceilings (C0–C5) are assessed independently from translation gates spanning material identity, reproducibility, mechanism, host safety, sterilization/storage, resistance, long-term fate, and deployment. Across mechanisms, SF and SS most often function as structural, interfacial, or transport-regulating matrices; direct silk-dependent bactericidal causality remains uncommon. The central translational deficit is failure to quantitatively link silk molecular identity and network architecture to antimicrobial exposure, bacterial killing, host selectivity, and long-term material fate. Full article
Show Figures

Graphical abstract

40 pages, 7162 KB  
Article
SSPA: Enhancing Pseudo-Corpus Quality on Tibetan Machine Translation via Semantic-Syntax Prealignment
by Yidong Sun, Dongxu Liu, Jiale Zhang and Youcheng Wang
Computers 2026, 15(9), 576; https://doi.org/10.3390/computers15090576 - 2 Sep 2026
Viewed by 246
Abstract
Tibetan-to-English machine translation (MT) models frequently falter under extreme domain data scarcity, often producing translations that violate the distinctive agglutinative rules of Tibetan and suffer from domain-specific stylistic mismatches. To overcome these limitations, we propose Semantic-Syntax Prealignment (SSPA), an innovative corpus generation framework. [...] Read more.
Tibetan-to-English machine translation (MT) models frequently falter under extreme domain data scarcity, often producing translations that violate the distinctive agglutinative rules of Tibetan and suffer from domain-specific stylistic mismatches. To overcome these limitations, we propose Semantic-Syntax Prealignment (SSPA), an innovative corpus generation framework. SSPA constructs high-quality pseudo-parallel pairs by explicitly minimizing the deviation between the syntactic-semantic profiles of generated samples and professional reference texts. Specifically, source-target structural representations are standardized through length-unified truncation and terminology normalization, followed by a dual-domain alignment process that maximizes syntactic cosine similarity under rigorous structural constraints. We further augment these aligned frames via a cross-length dynamic filling mechanism, which is integrated with an Expectation-over-Transformation (EOT)-based style regularization mechanism specifically adapted for stylistic perturbations, to simulate authentic linguistic variations. Extensive evaluations on our newly constructed Tibetan Medicine-Tibetan English (TM-TE) dataset demonstrate that SSPA significantly outperforms existing competitive baselines. Notably, SSPA achieves a BLEU-4 score of 36.2 and improves long-sentence BLEU-4 by 16.8 points, with a parser-verified grammatical compliance rate of 96.2%. The framework exhibits remarkable cross-domain adaptability and stylistic consistency, offering a robust, versatile solution for low-resource Tibetan professional domain MT. Full article
Show Figures

Figure 1

25 pages, 2192 KB  
Systematic Review
From Sustainability Recognition to Documented Outcomes: A Systematic Review and Study-Level Evidence Translation Analysis Across Productive Sectors
by Jorge Andrés Sarmiento Rojas, Fabián David Güiza Pinzón and Oscar Julian Alarcón Argüello
Sustainability 2026, 18(17), 8921; https://doi.org/10.3390/su18178921 - 31 Aug 2026
Viewed by 200
Abstract
Sustainability is widely recognized in project-based and productive-sector research, yet published studies do not always make visible how this recognition enters control routines, resource-allocation rules, and documented outcomes. This systematic review and study-level evidence translation analysis examined 133 unique full-text studies drawn from [...] Read more.
Sustainability is widely recognized in project-based and productive-sector research, yet published studies do not always make visible how this recognition enters control routines, resource-allocation rules, and documented outcomes. This systematic review and study-level evidence translation analysis examined 133 unique full-text studies drawn from Scopus and Web of Science evidence matrices covering publications from 2008 to 2026. A bilingual rule-based coding layer organized candidate evidence across ten sustainability-control domains and three stages: recognition (R), operationalization (O), and documented outcome evidence (E). Two independent university professor reviewers, one with a sustainability focus, and one with a statistics focus and each with more than five years of professional experience, then coded the complete corpus under the pre-specified protocol; disagreements were resolved through third-author adjudication. Observed agreement was 0.921 for R, 0.831 for O, and 0.883 for E. In the final matrix, R, O, and E were present in 53.1%, 34.7%, and 45.4% of study-domain observations, respectively, and full R-O-E completion occurred in 29.7%. The pattern was non-nested: documented outcomes were sometimes reported without an explicit account of the operational pathway. Generalized estimating equations confirmed lower odds for O (OR = 0.423, 95% CI 0.378–0.473) and E (OR = 0.703, 95% CI 0.619–0.799) relative to R. A four-class Bernoulli latent-class model minimized BIC and identified stakeholder-governance-transition, lifecycle-system integration, digital-data evidence, and measurement–performance profiles. The resulting Sustainability Evidence Translation Control Architecture locates documentary discontinuities between purpose, controls, decisions, and outcomes while avoiding unsupported claims about symbolic intent, implementation failure, or causal organizational transformation. Full article
Show Figures

Figure 1

23 pages, 3164 KB  
Systematic Review
Nanoparticle-Based Therapies for Myocardial Injury and Heart Failure: A Systematic Review and Translational Appraisal of Preclinical Evidence
by Ayesha Jabeen, Ilaria Barison, Honoria Ocagli, Bruna Fata, Diego Perazzolo, Cristina Basso, Roberto Luisetto, Silvia Pozzo, Fabrizio Mancin, Enrico Grisan, Dario Gregori, Annalisa Angelini, Marny Fedrigo and Chiara Castellani
Biomolecules 2026, 16(9), 1245; https://doi.org/10.3390/biom16091245 - 27 Aug 2026
Viewed by 212
Abstract
Background: Heart failure remains a leading cause of morbidity and mortality, and current therapies rarely repair established myocardial damage. Nanoparticle-based interventions have been investigated across heterogeneous models of myocardial injury, remodeling, cardiomyopathy, and heart failure, but the distribution and translational maturity of this [...] Read more.
Background: Heart failure remains a leading cause of morbidity and mortality, and current therapies rarely repair established myocardial damage. Nanoparticle-based interventions have been investigated across heterogeneous models of myocardial injury, remodeling, cardiomyopathy, and heart failure, but the distribution and translational maturity of this evidence remain unclear. Methods: A systematic search of PubMed, Embase, Scopus, and Web of Science was conducted from database inception to June 2024. Eligible reports were mapped according to disease model, experimental system, carrier-level nanoparticle platform, payload, route, comparator, outcomes, biodistribution, safety assessment, and translational characteristics. Reports of non-therapeutic nanoparticle exposure were retained in a separate contextual safety/toxicology stratum and were not included in the therapeutic evidence-density map. Risk of bias was evaluated using design-appropriate tools. Results: Of 2640 records screened, 157 independent studies met the criteria: 140 in the main therapeutic/platform evidence map and 17 in a separate contextual safety/toxicology stratum. Within the main corpus, polymeric systems were the largest platform class (n = 50), followed by inorganic/mineral (n = 35), biological/biomimetic (n = 24), lipid-based (n = 23), carbon-based (n = 5), and hybrid/multicomponent systems (n = 3). Evidence was concentrated in acute myocardial injury (n = 76), while direct same-agent comparisons, long-term safety assessment, repeated dosing, quantitative biodistribution, and clinically aligned heart-failure models remained limited. Conclusions: The field demonstrates substantial formulation diversity and biological activity, but translation is constrained by fragmented characterization, sparse comparative evidence, and incomplete assessment of biological fate and safety. Full article
(This article belongs to the Section Bio-Engineered Materials)
Show Figures

Figure 1

20 pages, 1564 KB  
Review
Wearable Technology in Winter Sports: A Cross-Domain Synthesis and a Conceptual Framework for the Cold-Context Translational Gap
by Zbigniew Waśkiewicz
Appl. Sci. 2026, 16(17), 8471; https://doi.org/10.3390/app16178471 - 25 Aug 2026
Viewed by 300
Abstract
Winter-sport wearable technology spans motion and force sensing, physiological monitoring, thermal intervention, flexible bioelectronics, equipment-integrated systems, and safety technologies. This structured critical review synthesizes an evidence base of 80 unique scholarly records identified through a systematic Boolean search executed on 15 August 2026 [...] Read more.
Winter-sport wearable technology spans motion and force sensing, physiological monitoring, thermal intervention, flexible bioelectronics, equipment-integrated systems, and safety technologies. This structured critical review synthesizes an evidence base of 80 unique scholarly records identified through a systematic Boolean search executed on 15 August 2026 in four standard academic databases (Scopus, Web of Science Core Collection, PubMed, and IEEE Xplore), which retrieved 1466 records (808 unique after cross-database deduplication), supplemented by backward/forward citation chasing for eligible records not indexed in these databases. The corpus comprises 32 direct winter-sport records, 15 cold-context translational records, 14 contextual validation records, and 19 secondary/background records. For empirical records containing sufficient information, validation maturity was additionally coded on a seven-stage ordinal scale; 54/80 records could be staged without inference, whereas 26/80 were retained as ‘not staged’. The corpus shows that translational maturity is strongly domain dependent. Motion and kinematic sensing frequently reaches real winter-sport training or field settings, whereas antifreezing hydrogels and flexible bioelectronics have advanced substantially in conductivity, adhesion, self-healing, conformability, and low-temperature operation but remain concentrated at material, integrated-device, and human-demonstration stages. The five recurring constraints—thermodynamic, interface, ecological, connectivity, and equity—are therefore reframed as non-equivalent, context-dependent dimensions rather than universal burdens. The revised architecture also distinguishes digitally mediated sense–decide–actuate loops from material-native stimulus–response and hybrid pathways, while continuous remote monitoring is treated as one option within an energy–communication trade-space. The resulting framework links evidence type, validation depth, system interface, and deployment context without equating commercial availability with scientific validation. Full article
(This article belongs to the Special Issue Advances in Biomechanics and Sports Medicine)
Show Figures

Figure 1

30 pages, 4199 KB  
Systematic Review
Credible Sovereignty: Operationalizing AI Governance Across Infrastructure, Data, and Models: A Systematic Review
by Raghu Raman and Prema Nedungadi
AI 2026, 7(9), 327; https://doi.org/10.3390/ai7090327 - 24 Aug 2026
Viewed by 359
Abstract
Claims of AI sovereignty are increasingly invoked but operational control remains uneven. Claims to control are made through national models, sovereign clouds, data localization mandates, and procurement rules; however, whether such claims translate into demonstrable control over how AI systems are run, inspected, [...] Read more.
Claims of AI sovereignty are increasingly invoked but operational control remains uneven. Claims to control are made through national models, sovereign clouds, data localization mandates, and procurement rules; however, whether such claims translate into demonstrable control over how AI systems are run, inspected, and contested remains poorly understood. This paper introduces credible sovereignty, the gap between declared and demonstrable control in deployment, as a conceptual lens for analyzing AI governance to examine how this gap is opened and closed across infrastructure, data, and model supply chains. Using a PRISMA-guided social-science corpus and machine learning-based BERTopic modeling, validated through topic diversity and topic separation diagnostics and triangulated through close reading, the analysis identifies four governance logics through which sovereignty is contested: data infrastructure and legitimacy frameworks; techno-bloc diplomacy and infrastructure politics; European regulatory sovereignty; and community-driven sovereignty in the Global South. Across these logics, sovereignty is enacted less through national capabilities than through proxy mechanisms—certification regimes, procurement clauses, cloud governance, and deployment architectures—each carrying trade-offs between autonomy, dependence, and accountability. Rereading the corpus through an Antecedents–Decisions–Outcomes lens yields a testable research agenda: antecedents that push actors toward sovereignty seeking; design and governance choices that translate ambition into implementation; and outcomes—resilience, inclusion, accountability—against which sovereign AI programs should be assessed. This paper reframes sovereignty as a layered operational capability rather than a discursive claim and links computational synthesis to a normative construct that applies across jurisdictions and scales. Full article
(This article belongs to the Section AI Systems: Theory and Applications)
Show Figures

Figure 1

33 pages, 3739 KB  
Article
SEM-PDPL: Semantic Exposure Graphs for Privacy-Law-Informed Risk Assessment of Public Social-Media Data
by Heba Ismail
Information 2026, 17(8), 803; https://doi.org/10.3390/info17080803 - 20 Aug 2026
Viewed by 301
Abstract
Public social-media content often contains self-disclosed personal attributes that appear low-risk in isolation but become privacy-relevant when linked across posts, platform accounts, or user-level traces. Existing research has advanced privacy-sensitive content detection, de-anonymization analysis, social-media research ethics, and privacy-compliance workflows; however, limited work [...] Read more.
Public social-media content often contains self-disclosed personal attributes that appear low-risk in isolation but become privacy-relevant when linked across posts, platform accounts, or user-level traces. Existing research has advanced privacy-sensitive content detection, de-anonymization analysis, social-media research ethics, and privacy-compliance workflows; however, limited work operationalizes how personal-data disclosures combine structurally and how these structures can be translated into auditable governance actions. This paper proposes SEM-PDPL, a computational, privacy-law-informed risk-assessment framework for modeling public social-media exposure as semantic exposure graphs and mapping graph patterns to controls aligned with the United Arab Emirates Personal Data Protection Law (PDPL) and compatible with GDPR principles. SEM-PDPL combines governance scoping; a PDPL-informed disclosure taxonomy; hybrid extraction using rule-based methods; named-entity recognition; fine-tuned BERT; and schema-constrained large language model annotation, followed by graph construction at post, platform, corpus, and persona levels. The framework is evaluated on a synthetic multi-platform corpus of 1095 posts generated for 150 personas across 290 platform accounts. Results show that, within this controlled synthetic corpus, fine-tuned BERT provides the strongest extraction performance among six evaluated methods, achieving a macro-F1 of 0.975. Graph analysis shows that exposure density increases with aggregation, rising from 0.275 at post level to 1.000 at corpus level, and from 0.859 at platform level to 0.967 at persona level. Across all graph resolutions, quasi-identifiers emerge as the dominant weighted-degree and betweenness node, indicating that ordinary location, employer, school, and demographic cues often function as bridges connecting sensitive categories such as health and biometric data to identifying information. These findings indicate that, within this controlled corpus, privacy risk in public social-media data is not only attribute-based but also structurally graph-shaped. SEM-PDPL contributes an explainable and reproducible framework for identifying exposure hubs, sensitive bridges, and aggregation risks before applying masking, minimization, exclusion, retention, or review controls. The framework does not automate legal compliance; rather, it provides evidence-based decision support for privacy-aware social-media analytics. Full article
(This article belongs to the Special Issue Semantic Networks for Social Media and Policy Insights)
Show Figures

Figure 1

27 pages, 6219 KB  
Article
Context-Sensitive N-Gram Word Partitioning for Improving the Quality of Turkish Word Embeddings
by Hayri Volkan Agun
Appl. Sci. 2026, 16(16), 8184; https://doi.org/10.3390/app16168184 - 17 Aug 2026
Viewed by 274
Abstract
Current advances in neural network models have improved state-of-the-art performance in natural language processing tasks such as named-entity recognition, sentiment analysis, and machine translation. In particular, neural language models are applied to encode information in word embeddings. These approaches are generally trained on [...] Read more.
Current advances in neural network models have improved state-of-the-art performance in natural language processing tasks such as named-entity recognition, sentiment analysis, and machine translation. In particular, neural language models are applied to encode information in word embeddings. These approaches are generally trained on large corpora using semi-supervised learning. Word embeddings encode the syntactic and semantic properties of words as dense vectors. In agglutinative languages such as Turkish, Finnish, and Hungarian, word-embedding construction is challenging because extensive suffixation and polysemy can cause information loss. To overcome these limitations, character n-grams are often preferred for embedding representations. Nevertheless, character n-grams do not guarantee the capture of information in long word sequences. In this study, a method that partitions word sequences according to frequent patterns within a given context is proposed for training a neural language model. In this respect, likelihood- and ranking-based inference are combined with n-gram and syllable partitioning for word-embedding generation from a text corpus. The proposed approach provides a language-agnostic, context-sensitive segmentation mechanism that can complement language processing methods such as lemmatization, morphological analysis, and stemming. For embedding generation, the SkipGram and FastText models are used, and the effects of word partitioning are evaluated using analogy, named-entity recognition, POS tagging, sentiment analysis, and morphological disambiguation datasets for Turkish. The results indicate task-dependent and generally limited improvements over traditional token-based word-embedding extraction. In particular, skip n-gram partitioning produces a substantial improvement over partitioning based on frequent-ngrams, sentencepiece-bpe, sentence-unigram and morfessor. No consistent relationship was observed across tasks between performance and either graph density or the average number of distinct n-grams per sentence. Full article
(This article belongs to the Special Issue Natural Language Processing: Modeling, Optimization and Application)
Show Figures

Figure 1

28 pages, 852 KB  
Article
Multi-Metric Evaluation of Translation-Based Cross-Lingual Sentiment Consistency Using Large Language Models and Neural Machine Translation
by Esra Duruoglu Cetin and Cagri Sahin
Appl. Sci. 2026, 16(16), 7878; https://doi.org/10.3390/app16167878 - 7 Aug 2026
Viewed by 440
Abstract
In today’s globalized and digitally connected world, individuals increasingly share emotions, opinions, and experiences across multiple languages, making accurate translation essential for cross-lingual sentiment analysis. Although machine translation (MT) is widely used in multilingual applications, the relationships among translation quality, semantic similarity, and [...] Read more.
In today’s globalized and digitally connected world, individuals increasingly share emotions, opinions, and experiences across multiple languages, making accurate translation essential for cross-lingual sentiment analysis. Although machine translation (MT) is widely used in multilingual applications, the relationships among translation quality, semantic similarity, and sentiment consistency remain insufficiently understood. This study investigates the performance of six LLM-based systems (GPT-4o-mini, Gemini 2.5 Flash-Lite, Qwen 2.5, Llama 3.1, Mistral 7B, and NiuTrans LMT) and four NMT-based systems (Google Translate, Microsoft Translator, NLLB-200, and LibreTranslate-v1.5) in maintaining classifier-mediated sentiment consistency across twelve translation directions involving English, Spanish, French, and Chinese. Experiments were conducted on the Multilingual Amazon Reviews Corpus (MARC), comprising 84,000 randomly sampled user reviews. A multidimensional evaluation framework was used, combining sentiment-consistency metrics (accuracy, weighted F1, MCC, and SSR), translation-quality estimation (COMET-QE), and semantic-similarity assessment (LaBSE). Statistical significance was examined using the Friedman and Nemenyi post hoc tests. The results show that GPT-4o-mini, Gemini 2.5 Flash-Lite, and Google Translate consistently ranked among the strongest systems across multiple evaluation dimensions. Performance differences were particularly pronounced in translation directions involving Chinese, highlighting the influence of language-specific structural characteristics. Furthermore, semantic similarity and translation quality exhibited only moderate relationships with sentiment consistency, indicating that high semantic similarity does not necessarily guarantee strong sentiment consistency. Overall, the findings demonstrate the importance of multidimensional and statistically grounded evaluation frameworks for assessing cross-lingual sentiment consistency and provide practical insights into the strengths and limitations of contemporary MT systems. Full article
(This article belongs to the Special Issue New Trends in Natural Language Processing, 2nd Edition)
Show Figures

Figure 1

23 pages, 378 KB  
Article
Overtourism Communication and Public Visitor Guidance in Kyoto, Japan: A Soft Urban Governance Perspective
by Hermann Kimo Boukamba
Tour. Hosp. 2026, 7(8), 232; https://doi.org/10.3390/tourhosp7080232 - 6 Aug 2026
Viewed by 935
Abstract
Overtourism management increasingly uses public communication to frame how visitors are expected to move, behave, and understand their responsibilities in pressured destination spaces. Yet, in heritage cities, such communication is often examined through isolated signs, etiquette campaigns, visitor advisories, or congestion notices rather [...] Read more.
Overtourism management increasingly uses public communication to frame how visitors are expected to move, behave, and understand their responsibilities in pressured destination spaces. Yet, in heritage cities, such communication is often examined through isolated signs, etiquette campaigns, visitor advisories, or congestion notices rather than as part of a wider visitor-management communication system. This paper examines Kyoto as a case of overtourism communication in an urban heritage setting. Using document-based content analysis supported by exploratory statistical tests, it maps 64 public-facing communication artifacts published between 2016 and 2025 and codes them by message orientation, artifact type, source environment, and spatial focus. The findings show that exclusionary messaging does not dominate Kyoto’s public-facing overtourism communication within the analyzed corpus. Dispersal and behavioral guidance is the most common orientation, while restriction is targeted and secondary, and invitational place guidance remains comparatively limited. Message orientation is significantly associated with artifact type, with restriction concentrated in signs and posters and invitational place guidance appearing mainly in guides and brochures. Semi-institutional actors account for most artifacts, indicating the central role of destination-management organizations and other visitor-facing tourism bodies in translating management concerns into public guidance. Spatially, guidance-oriented communication appears broadly across the city, while restriction has its clearest place-based concentration in Eastern Kyoto, especially Gion-related areas. A smaller set of messages invites visitors to explore wider Kyoto or less conventional places beyond standard guidebook routes. Drawing on the lens of soft urban governance, the paper argues that public-facing overtourism communication can be understood as one visible layer through which visitor-management expectations are framed, spatially differentiated, and made publicly legible. As a document-based analysis, the study does not measure visitor reception, compliance, or behavioral effects, which remain priorities for future research. Full article
40 pages, 60827 KB  
Article
IDS-Based Accessibility Validation in BIM Using ISO 21542:2021 Door Criteria: A Case Study of the Urla Summer Villa
by Murat Aydın
Buildings 2026, 16(15), 3057; https://doi.org/10.3390/buildings16153057 - 2 Aug 2026
Viewed by 404
Abstract
Building Information Modeling (BIM) has become a central paradigm in the architecture, engineering, and construction industry, enabling integrated management of design and construction data. Accessibility standards are essential to ensure inclusive and usable built environments, yet manual verification of ISO 21542:2021 criteria in [...] Read more.
Building Information Modeling (BIM) has become a central paradigm in the architecture, engineering, and construction industry, enabling integrated management of design and construction data. Accessibility standards are essential to ensure inclusive and usable built environments, yet manual verification of ISO 21542:2021 criteria in BIM models is time-consuming and error-prone. Following the formal standardization of the Information Delivery Specification (IDS) in 2024, this study investigates its potential for systematic accessibility validation. Using the Design Science Research methodology, ISO 21542 door requirements were translated into IDS format and applied to the Urla Summer Villa case. The workflow integrated IDS Maker for XML rule definition, ArchiCAD for BIM, and BIMvision with the IDS Checker plugin for automated validation. The evaluation revealed an overall compliance rate of 50%, with non-compliances in door width, height, threshold, and handle parameters. Results indicate that IDS can detect accessibility inconsistencies early, provide transparent reporting, and strengthen quality assurance in BIM-based design processes. This study contributes to expanding the limited corpus of accessibility-focused IDS applications and highlights IDS as a practical tool for advancing inclusive design practices within digital construction workflows. The study is limited to a single case and door elements only, within a specific software ecosystem. Future research should extend IDS-based accessibility validation to ramps, stairs, elevators, and diverse BIM platforms to strengthen generalizability. Full article
(This article belongs to the Special Issue Emerging Technologies and Workflows for BIM and Digital Construction)
Show Figures

Figure 1

34 pages, 758 KB  
Article
MythoBiLLM: BiLSTM-Guided Parameter-Efficient Fine-Tuning of Large Language Models for Coherent Summarization and Generation of Indian Mythological Texts
by Shweta Bansal, Sumendra Yogarayan and Siti Fatimah Abdul Razak
Information 2026, 17(8), 726; https://doi.org/10.3390/info17080726 - 27 Jul 2026
Viewed by 282
Abstract
Indian mythological narratives contain long event chains, recurring characters, moral conflicts, interactions between human and divine agents, and source-specific narrative styles. General-purpose large language models can generate fluent text while losing character continuity, thematic relations, or source-supported events. This study presents MythoBiLLM, a [...] Read more.
Indian mythological narratives contain long event chains, recurring characters, moral conflicts, interactions between human and divine agents, and source-specific narrative styles. General-purpose large language models can generate fluent text while losing character continuity, thematic relations, or source-supported events. This study presents MythoBiLLM, a parameter-efficient framework for summarization and continuation generation from Indian mythological texts. The framework combines a frozen Llama 3.2 3B-Instruct backbone, LoRA-based adaptation, and a gated BiLSTM narrative-memory adapter. A corpus of public-domain English translations from the Ramayana, Mahabharata, Bhagavad-Gita, Vishnupuranam, Harivamsha, Hindu Tales, and Indian Myth and Legend contains 3,684,838 word-level tokens and 6057 segmented passages. Evaluation covers language modeling, summarization, continuation generation, entity consistency, theme retention, component ablation, robustness, human assessment, and statistical testing. Relative to LLM+LoRA, the complete framework reduces average perplexity from 23.4 to 19.8. In controlled comparisons, the BiLSTM adapter achieves an MCS of 0.713 on both tasks, compared with 0.699 for the parameter-matched MLP adapter, 0.704 for independently trained long-context LoRA, and 0.708 for retrieval augmentation. Full MythoBiLLM reaches MCS values of 0.762 for summarization and 0.744 for continuation generation. After entity consistency and style alignment are excluded from MCS, the complete configuration retains the highest scores of 0.751 and 0.731. These findings support the complete framework on the evaluated corpus, while the controlled comparisons indicate a modest complementary contribution from the BiLSTM and do not identify it as the sole source of the performance gains. Full article
Show Figures

Graphical abstract

25 pages, 1630 KB  
Article
A Hybrid NLLB and Large Language Model Pipeline for Diachronic Intralingual Translation of 16th-Century Slovene Literary Heritage
by Žan Tomaž Šprajc, Rok Sekirnik, Vlasta Kučiš and David Jesenko
Appl. Sci. 2026, 16(14), 7317; https://doi.org/10.3390/app16147317 - 21 Jul 2026
Viewed by 441
Abstract
Modernising historical literature into contemporary language is a form of diachronic intralingual translation that supports access to written cultural heritage. For low-resource languages such as Slovene, this task is hindered by orthographic, lexical and syntactic shifts, as well as the scarcity of parallel [...] Read more.
Modernising historical literature into contemporary language is a form of diachronic intralingual translation that supports access to written cultural heritage. For low-resource languages such as Slovene, this task is hindered by orthographic, lexical and syntactic shifts, as well as the scarcity of parallel data. We present a two-stage pipeline that combines a fine-tuned No Language Left Behind (NLLB) model with Claude Opus 4.8 post-editing for the modernisation of 16th-century Slovene literature, retaining archaic words and phrases while normalising the alphabet, orthography and grammar. Using Jurij Dalmatin’s 1584 Bible and its 2017 modernised edition, we constructed an aligned parallel corpus of 14,876 sentence and clause-level pairs through Bohorič-to-Gaj normalisation and LaBSE-based embedding alignment. The hybrid pipeline achieved the best overall scores, reaching BLEU 45.78, CHRF 67.68 and METEOR 71.72, with TER 43.25 and CER 34.00. Its gains over the standalone LLM baseline were large and statistically significant across all the metrics, while its improvement over the fine-tuned NLLB model was smaller and significant mainly for overlap-based measures. We applied the pipeline further to Tulščak’s Kerszhanske leipe molitve (1579), producing the first preliminary modernisation of the earliest known Slovene prayer book assessed qualitatively and tested the generalisation on additional 16th-century texts, including an out-of-domain legal text. The results demonstrate that combining task-specific neural machine translation with controlled LLM post-editing offers a practical strategy for modernising low-resource historical texts and contributes a reusable methodology for digital cultural heritage preservation. Full article
(This article belongs to the Special Issue Artificial Intelligence Technologies in Cultural Heritage)
Show Figures

Figure 1

Back to TopTop