Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (62)

Search Parameters:
Keywords = document-topic embeddings

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
38 pages, 2627 KB  
Article
Reliable Criterion Retrieval for Sensor-Instrumented Road Infrastructure: Diagnosing and Correcting a Title-Framing Bias in Dense Retrieval over Korean Design Documents
by Byeong-Cheol Kim and Byung-Jik Son
Sensors 2026, 26(17), 5683; https://doi.org/10.3390/s26175683 - 7 Sep 2026
Abstract
Retrieval-augmented generation increasingly serves as the knowledge backbone for sensor-informed infrastructure decisions, where a field engineer must retrieve the clause stating a design criterion, not a document about the topic. In twenty years (2005–2024) of Korea Expressway Corporation design-practice guidelines (HWP/HWPX), we identify [...] Read more.
Retrieval-augmented generation increasingly serves as the knowledge backbone for sensor-informed infrastructure decisions, where a field engineer must retrieve the clause stating a design criterion, not a document about the topic. In twenty years (2005–2024) of Korea Expressway Corporation design-practice guidelines (HWP/HWPX), we identify a framing bias in dense retrieval as follows: embedding models over-weight the topical aboutness of a section title relative to the criterion in its body. The corpus invites this failure as follows: 63.2% of sections carry plan-framed titles, yet 86.4% of them carry criterion-type language in the body. A controlled counterfactual (n = 80) holding the body fixed and rewriting only the title isolates the effect (Δcos = 0.030, dz = 1.63, p = 1.1 × 10−14), reproduces it on a second embedding family, and decomposes it into term-frequency, early-position, and title-framing components; the framing residual (dz = 0.53) survives primacy controls. A length- and frequency-matched neutral-token control splits that residual further into a token-composition component that replicates on both embedding families and a plan-framing component that reaches significance on one (dz = 0.46). The bias buries framing-prone criterion documents by tens to hundreds of ranks; standard remedies are partial. We propose Criterion-Aware Retrieval (CAR), which hypothesizes the sought criterion at query time; on the 21 low-overlap queries that the bias hits hardest it outperforms both BM25 and the weighted-RRF hybrid after Holm correction (MRR 0.271 vs. 0.048 and 0.120). A cross-encoder reranker ranks better still (0.376) at no language-model cost but cannot exceed the recall of the pool it reorders (0.714 against CAR’s 0.857): the two address different failure modes, and widening the pool is what the framing bias calls for. A parsing-pathway comparison shows that the structured pathways measured are near-lossless while PDF loses table content, justifying our HWPX-derived ground truth. Full article
(This article belongs to the Special Issue Smart Infrastructure for Sensor-Driven Systems)
Show Figures

Figure 1

44 pages, 4771 KB  
Article
Evaluating LLM-Based Retrieval-Augmented Generation for Soil Science Question Answering
by Karla Topić, Marina Bagić Babac and Vedran Mornar
Information 2026, 17(9), 859; https://doi.org/10.3390/info17090859 - 4 Sep 2026
Viewed by 183
Abstract
Retrieval-augmented generation (RAG) systems for scientific literature require evidence-based choices of document segmentation, representation, retrieval, and generation components, particularly when the source collection varies in topical specificity and document structure. This study addresses the lack of an end-to-end, component-level comparison of these choices [...] Read more.
Retrieval-augmented generation (RAG) systems for scientific literature require evidence-based choices of document segmentation, representation, retrieval, and generation components, particularly when the source collection varies in topical specificity and document structure. This study addresses the lack of an end-to-end, component-level comparison of these choices for soil science question answering. A three-stage evaluation was conducted across general, domain-specific, and geospatial soil science corpora. The corpus combines foundational soil science books, peer-reviewed research articles, European soil monitoring material, and geospatial mapping publications, thereby covering both broad disciplinary concepts and specialized scientific evidence. The study compares four chunking strategies, three embedding models, five retrieval methods, and five large language models. In Experiment 1, semantic chunking with text-embedding-3-large achieved the highest aggregate retrieval scores (recall@1 = 0.824; MRR = 0.819), whereas text-embedding-3-small delivered practically comparable performance at lower cost. In Experiment 2, hybrid reciprocal rank fusion achieved recall@5 values of 0.957, 0.960, and 0.647 for the general, domain-specific, and geospatial corpora, respectively; the cross-encoder reranker showed weaker rank quality on scientific content. In Experiment 3, model responses attained BERTScore values of 0.909–0.927 and faithfulness of at least 0.993; these automated measures indicate low contradiction with retrieved context but do not establish answer completeness or human-perceived correctness. The study provides a reproducible component-level evaluation design, characterizes the effect of corpus specificity on RAG retrieval, and identifies a practical configuration for soil science literature retrieval. Among the models retained for direct aggregate comparison, Llama 3.1 8B offered the most favorable observed balance of answer quality, latency, cost, and model openness. Full article
Show Figures

Graphical abstract

29 pages, 722 KB  
Article
TMCAS: Efficient Large Language Model-Assisted Topic Modeling for Civil Aviation Safety Reports
by Xiangge Li, Haofeng Wang, Xiuting Zhou, Yan Ren, Zhi Tian, Weidong Liang and Gengsong Wang
Electronics 2026, 15(15), 3462; https://doi.org/10.3390/electronics15153462 - 5 Aug 2026
Viewed by 276
Abstract
Voluntary safety reports provide valuable information for identifying potential risks and improving safety management in civil aviation. However, these reports are often large in volume, unstructured in format, and rich in domain-specific terminology, making manual analysis costly, inefficient, and difficult to scale. To [...] Read more.
Voluntary safety reports provide valuable information for identifying potential risks and improving safety management in civil aviation. However, these reports are often large in volume, unstructured in format, and rich in domain-specific terminology, making manual analysis costly, inefficient, and difficult to scale. To address these challenges, this paper proposes TMCAS, an efficient large language model-assisted topic modeling framework for civil aviation safety reports. The proposed framework combines domain-adapted text embeddings, density-based clustering, representative sampling, noise repair, and large language model-based topic generation. Specifically, a contrastive learning-based fine-tuning strategy is introduced to enhance the semantic representation of aviation safety texts. An HDBSCAN-based clustering and sampling mechanism is then designed to select representative reports and reduce the computational cost of large language model inference, while a noise-repair strategy is used to improve topic coverage. Finally, large language models are employed to generate interpretable sentence-level topic labels and descriptions. Experiments demonstrate that TMCAS achieves superior clustering and interpretability while substantially reducing inference cost compared with document-wise LLM baselines. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

22 pages, 2060 KB  
Systematic Review
Mapping the Methodological Bifurcation of Quantitative Portfolio Optimization: A PRISMA-Compliant Systematic Review with BERTopic–SPECTER Analysis (2003–2025)
by Gharmili Meryem, Boudri Imane and Alj Abdelkamel
J. Risk Financ. Manag. 2026, 19(8), 582; https://doi.org/10.3390/jrfm19080582 - 3 Aug 2026
Viewed by 355
Abstract
Quantitative portfolio optimization has accelerated sharply since 2018, with deep learning and reinforcement learning agents now competing with the mean–variance framework that defined six decades of research. Existing narrative reviews struggle to track this expansion. We screen 832 documents from Scopus and Web [...] Read more.
Quantitative portfolio optimization has accelerated sharply since 2018, with deep learning and reinforcement learning agents now competing with the mean–variance framework that defined six decades of research. Existing narrative reviews struggle to track this expansion. We screen 832 documents from Scopus and Web of Science under PRISMA 2020 and retain 589 unique articles spanning 2003–2025. Applying BERTopic with SPECTER scientific embeddings, UMAP and HDBSCAN, we identify five coherent topics with a mean coherence of 0.864: classical mean–variance (T0; n = 270), deep reinforcement learning (T1; n = 116), machine learning return forecasting (T2; n = 87), covariance estimation and robust optimization (T3; n = 52)—and metaheuristics (T4; n = 56). A rank-weighted similarity analysis, designed to neutralise the c-TF-IDF collinearity artefact, shows that deep reinforcement learning is the most isolated paradigm. The two methodological families bifurcate over time: AI/deep learning approaches grow from 3.6% of annual output before 2018 to 40.2% afterwards, while classical methods retain volume but lose share. We synthesise the empirical practices of each family along five dimensions critical to applied finance and identify three under-explored integration frontiers. Full article
Show Figures

Figure 1

28 pages, 6917 KB  
Article
Mediating Pathways to Sustainable Investment: A TOE Framework for AI-Driven Green Fintech Adoption in Banking
by Reem A. Abdalla, Lamya Abbas Hidaytalla and Gulnar Sadat Mulla
J. Risk Financ. Manag. 2026, 19(7), 496; https://doi.org/10.3390/jrfm19070496 - 3 Jul 2026
Viewed by 869
Abstract
Purpose: Despite growing research on green fintech and sustainable finance individually, no systematic theoretical framework explains how AI-driven green fintech solutions can be adopted in banking for sustainable investment purposes. This paper addresses this demonstrated gap by developing the first bibliometrically grounded, TOE-based [...] Read more.
Purpose: Despite growing research on green fintech and sustainable finance individually, no systematic theoretical framework explains how AI-driven green fintech solutions can be adopted in banking for sustainable investment purposes. This paper addresses this demonstrated gap by developing the first bibliometrically grounded, TOE-based conceptual framework for AI-driven green fintech adoption in banking. Design/Methodology/Approach: A two-phase approach is employed. First, a bibliometric analysis of 79 Scopus-indexed documents (2020–2026) using bibliometrix in R provides quantitative evidence of the research gap through keyword co-occurrence networks, thematic mapping, and trend topic analysis. Second, building on this evidence, a conceptual framework integrating the Technology–Organization–Environment (TOE) framework with three mediating constructs, technological readiness, sustainability culture, and regulatory support is developed and five theoretical propositions are derived. Findings: The bibliometric analysis reveals an annual growth rate of 78.4% in the field and confirms that the TOE framework has never occupied the motor themes quadrant of the green fintech literature. The proposed framework theorizes three mediated pathways through which technological, organizational, and environmental conditions translate into improved sustainable investment outcomes including enhanced ESG transparency, increased green investment allocation, and SDG alignment. Practical Implications: The framework provides bank executives with three actionable intervention points: technological infrastructure investment, sustainability culture embedding, and regulatory engagement and offers policymakers evidence-based guidance for designing supportive green fintech adoption frameworks. Originality/Value: This study presents a conceptual framework that is, to the authors’ knowledge, the first to combine TOE theory, AI-driven green fintech, a banking context, an explicit three-mediator architecture (technological readiness, sustainability culture, regulatory support), and sustainable investment outcomes as the dependent variable, grounded in reproducible bibliometric evidence. Existing studies address subsets of these dimensions; none integrates all six simultaneously. Full article
(This article belongs to the Section Financial Technology and Innovation)
Show Figures

Figure 1

24 pages, 14896 KB  
Article
Analyzing Post-Disaster Public Reactions in Turkish Social Media Through Topic Modeling and Hybrid Sentiment Classification
by Ayşe Meydanoğlu, Serpil Aslan, Emirhan Denizyol, Mesut Toğaçar, Abdurrezzak Ekidi, Yunus Emre Temiz, Tuncay Karateke, Ramazan Erten, Beyzade Nadir Çetin, Enes Saylan and Hatice Çakmak
Electronics 2026, 15(13), 2911; https://doi.org/10.3390/electronics15132911 - 2 Jul 2026
Viewed by 419
Abstract
Social media has emerged as a crucial environment for examining public sentiment during disasters, providing immediate insights into collective emotions and urgent expectations. This research examines the emotional reactions expressed on Turkish posts shared on the X platform (formerly Twitter) following the 6 [...] Read more.
Social media has emerged as a crucial environment for examining public sentiment during disasters, providing immediate insights into collective emotions and urgent expectations. This research examines the emotional reactions expressed on Turkish posts shared on the X platform (formerly Twitter) following the 6 February 2023 earthquake by employing an integrated method that combines topic modeling and topic-based sentiment analysis. Data were collected between 10 February 2023 and 28 February 2023. A large dataset consisting of 305,000 tweets was compiled, and 296,836 tweets remained for analysis after preprocessing and filtering procedures. Latent Dirichlet Allocation (LDA), enhanced with term frequency-inverse document frequency weighting and bigram extraction techniques, was applied to identify prominent themes, including rescue operations, appeals for assistance, communication about missing persons, and disaster management. The sentiment polarity within each topic was determined using a hybrid deep learning model incorporating Bidirectional Encoder Representations from Transformers (BERT) embeddings Convolutional Neural Networks (CNN), Bidirectional Long Short-Term Memory (BiLSTM) layers, and FastText representations. This model reached a classification accuracy of 94%, with F1-scores of 0.91 and 0.95, recall values of 0.90 and 0.96, and precision values of 0.92 and 0.95, achieving higher performance than the evaluated baseline models. The findings indicate that supportive, solidarity-oriented, and resilience-related communication patterns were among the most frequently observed positive sentiment expressions, whereas negative sentiments appeared more frequently in discussions regarding delays in aid delivery and perceived shortcomings in institutional response. This study presents a scalable and flexible framework for analyzing sentiment in Turkish-language crisis communication, providing insights that may support disaster response monitoring and decision-making processes as well as the development of systems for tracking public reactions in real time. Full article
(This article belongs to the Section Computer Science & Engineering)
Show Figures

Figure 1

44 pages, 820 KB  
Article
An Information-Geometric Justification for Composite Coherence in Event-Based Narrative Extraction
by Brian Keith-Norambuena
Entropy 2026, 28(7), 732; https://doi.org/10.3390/e28070732 - 28 Jun 2026
Viewed by 322
Abstract
Graph-based narrative extraction relies on a coherence function to score transitions between events, but the coherence metrics in current use are defined operationally and lack an information-theoretic foundation. We study the composite metric C=A·T, where A is the [...] Read more.
Graph-based narrative extraction relies on a coherence function to score transitions between events, but the coherence metrics in current use are defined operationally and lack an information-theoretic foundation. We study the composite metric C=A·T, where A is the angular similarity of document embeddings and T=1dJS is the topic proximity through the Jensen–Shannon distance of soft cluster memberships, and we provide an information-geometric reading of this metric together with an axiomatic characterization of the geometric-mean combinator. On the product manifold Sd1×Δ+K1, the negative log-coherence decomposes additively into an angular and a topic cost. Because the Riemannian metric tensor induced by the Jensen–Shannon distance on the simplex is proportional to the Fisher information matrix, the topic component is locally consistent with the Fisher–Rao metric singled out by Chentsov’s theorem. Within a parametric family of combinators (the compensability spectrum), the geometric mean is the unique combinator consistent with four natural axioms (a boundary/veto condition, symmetry, log-additivity, normalization), and the construction also motivates a proper product metric d× that we use as a reference distance. Experiments on four corpora spanning news and academic domains (40 to 6000 documents), three general-purpose embedding families (GPT-4/ada-002, MPNet, MiniLM-L6) plus citation-aware SPECTER2, and three alternative topic models (LDA, soft k-means, GMM) are consistent with the framework: the Fisher identity holds with R0.99, the geometric mean tracks d× closely (ρ=0.999), and a downstream LLM-as-judge consistency check shows that the geometric mean is not empirically dominated by any alternative combinator or single-channel baseline. Sweeping the compensability spectrum, the bottleneck-coherence gap between extracted storylines and random sequences splits into a symmetric component—maximized at the geometric mean on the four corpora above and a fifth, human-navigation corpus—and a displacement term; a cross-modal case study on a human-curated image narrative reproduces the same effect in a second modality. Together, these results provide an information-geometric justification for the composite coherence metric and articulate the conditions under which the geometric mean is the natural choice. Full article
(This article belongs to the Special Issue Information Theory in Artificial Intelligence)
Show Figures

Figure 1

23 pages, 9226 KB  
Article
A Method for Comment Text Feature Mining via Integrated Keyword Extraction, Clustering, and Sentiment Analysis
by Jinbao Song, Jiahui Cai, Yijun Wang, Kai Wang, Shiwen Cui and Nuo Xu
Appl. Syst. Innov. 2026, 9(6), 124; https://doi.org/10.3390/asi9060124 - 11 Jun 2026
Viewed by 645
Abstract
In recent years, short video platforms have rapidly developed into important media for cultural dissemination. The interactions of netizens in short video comment sections not only reflect their focus on cultural content but also contain rich emotional attitudes. However, given the vast and [...] Read more.
In recent years, short video platforms have rapidly developed into important media for cultural dissemination. The interactions of netizens in short video comment sections not only reflect their focus on cultural content but also contain rich emotional attitudes. However, given the vast and fragmented nature of comment data, accurately extracting keywords, identifying cultural themes, and analyzing sentiment tendencies pose significant challenges in understanding netizens’ cultural perceptions. To address these challenges, this study proposes a text analysis framework that integrates keyword extraction, clustering analysis, and sentiment analysis to explore the core topics and emotional characteristics of cultural dissemination in short video comment sections. Firstly, to address the challenge of balancing statistical information and semantic understanding in short-text keyword extraction, this paper proposes the TF-IDF-KeyBERT Integrated Algorithm (TKIA) keyword extraction algorithm, which integrates Term Frequency–Inverse Document Frequency (TF-IDF) and Key Bidirectional Encoder Representations from Transformers (BERT). Experiments on the CSL dataset demonstrate improvement in the F1@5 metric, showing its potential to enhance keyword extraction performance for short texts. Secondly, to address the difficulty of simultaneously considering semantic representation capability and clustering flexibility in short-text clustering analysis, this paper designs the Self-Supervised Contrastive Enhanced Clustering (SCEC) algorithm by integrating self-supervised contrastive learning with a soft clustering strategy. Compared to baseline methods, SCEC improves clustering accuracy (ACC) by 17.5% on AGNews and 6.8% on THUCNews, suggesting a more effective way to reveal the underlying structure of cultural topics. Finally, to address the challenge of effectively leveraging both text structural information and global semantic features in short-text sentiment analysis, this paper develops the BERT-GCN Cross-Attention (BGC) Model, integrating BERT embeddings and Graph Convolutional Network (GCN)-based structural features via a Cross-Attention mechanism. On the My_weibo_senti_100k dataset, the BGC model achieves a 2.45% increase in Macro-F1 and a 2.41% improvement in accuracy over strong baselines, offering its ability for high-precision modeling of user sentiment. This study offers effective data support and technical pathways for applications such as cultural content understanding, personalized recommendation, and user emotion guidance. Full article
(This article belongs to the Special Issue Smart and Human-Centered Rehabilitation Technologies and Systems)
Show Figures

Figure 1

21 pages, 2399 KB  
Article
Research on Framework for and Strategies of Green Energy Consumption Based on Unsupervised Machine Learning
by Jun Lyu, Yu Shu and Shuo Wang
Energies 2026, 19(11), 2733; https://doi.org/10.3390/en19112733 - 5 Jun 2026
Cited by 1 | Viewed by 410
Abstract
Documentary videos on green energy consumption are widely distributed via platforms such as YouTube, yet the verbal framing strategies embedded in their subtitle transcripts remain systematically understudied. This study applies the Analysis of Topic Model Networks (ATMN)—an unsupervised machine learning approach combining LDA [...] Read more.
Documentary videos on green energy consumption are widely distributed via platforms such as YouTube, yet the verbal framing strategies embedded in their subtitle transcripts remain systematically understudied. This study applies the Analysis of Topic Model Networks (ATMN)—an unsupervised machine learning approach combining LDA topic modeling, semantic network analysis, and hierarchical clustering—to subtitle transcripts extracted from 60 YouTube green energy consumption documentaries. Three distinct framing communities are identified: (1) the Technological Supply Frame, which foregrounds zero-carbon resources, renewable generation, smart grid systems, and AI-enabled energy management as the technical foundation of decarbonization; (2) the Socioeconomic Transition Frame, the most thematically expansive, which positions the energy transition simultaneously as an economic opportunity, a behavioral imperative, and a systemic industrial transformation spanning green investment, end-use substitution, industrial decarbonization, and green mobility; and (3) the Ecological Governance Frame, which integrates ecological co-benefits with international climate commitments to construct the transition as a globally mandated planetary responsibility. Together, these frames reveal a richer and more multi-dimensional verbal framing landscape than previously documented in the green energy communication literature, extending beyond techno-optimism or environmentalism to encompass financial, governance, and behavioral dimensions within a single integrated corpus. The identified framing strategies offer actionable guidance for policymakers, energy enterprises, and media producers seeking to accelerate green energy consumption transition through targeted, evidence-based video communication. Full article
Show Figures

Figure 1

23 pages, 466 KB  
Article
The Knowledge-Coherence Framework for Narrative Extraction: An Empirical Study on Scientific Literature
by Brian Keith-Norambuena and Carolina Flores-Bustos
Analytics 2026, 5(2), 18; https://doi.org/10.3390/analytics5020018 - 4 May 2026
Viewed by 865
Abstract
Narrative extraction builds coherent ordered sequences of documents that trace how concepts develop over time, and is a growing area of information retrieval. In this work we focus on scientific literature, using a corpus of 3549 IEEE visualization research papers (1990–2022). A natural [...] Read more.
Narrative extraction builds coherent ordered sequences of documents that trace how concepts develop over time, and is a growing area of information retrieval. In this work we focus on scientific literature, using a corpus of 3549 IEEE visualization research papers (1990–2022). A natural hypothesis is that augmenting embedding-based pathfinding with explicit domain knowledge should improve narrative quality. We present the Knowledge-Coherence Framework (KCF), which integrates structured metadata from OpenAlex into narrative extraction (building on the Narrative Trails algorithm), and conduct a systematic empirical investigation along three axes: (1) the effect of embedding model choice (MiniLM vs. SPECTER), (2) the effect of knowledge augmentation (with and without, plus sensitivity to the knowledge weight α), and (3) the reliability of LLM-based evaluation (cross-agreement among 13 large language models). Throughout, mathematical coherence denotes the geometric mean of angular and topic similarity between consecutive documents along a path—an automatic, model-computed quantity inherited from Narrative Maps and Narrative Trails—while narrative quality refers to the LLM-judged construct. Using up to 600 evaluation pairs, we find that embedding model choice has a large effect on mathematical coherence (SPECTER: 0.94 vs. MiniLM: 0.81) and that, contrary to expectations, knowledge augmentation does not improve LLM-judged narrative quality—it slightly decreases it for both embeddings. Notably, the two notions dissociate: SPECTER produces the most mathematically coherent paths, yet MiniLM paths receive the highest LLM narrative-quality scores (5.87 vs. 5.36 out of 10). Alpha sensitivity analysis over five values (α{0.0,0.3,0.5,0.7,1.0}, 500 pairs) confirms that LLM scores remain essentially flat while mathematical coherence steadily declines with increasing knowledge weight. Cross-model evaluation with 13 LLM judges shows high inter-model agreement (median Pearson r=0.71), supporting evaluation reliability. The main practical takeaways are that (i) embedding model choice, not knowledge augmentation, is the more consequential design decision, and (ii) mathematical coherence and LLM-judged narrative quality are distinct optimization targets that practitioners should not conflate. Full article
Show Figures

Figure 1

40 pages, 3738 KB  
Article
Knowledge Evolution in the Mobile Industry via Embedding-Based Topic Growth and Typology Analysis
by Sungjin Jeon, Woojun Jung and Keuntae Cho
Systems 2026, 14(4), 415; https://doi.org/10.3390/systems14040415 - 9 Apr 2026
Viewed by 949
Abstract
The mobile industry has experienced long-run changes in its knowledge structure, including identifiable transition points observable through embedding-based semantic analysis. Using abstracts from 86,674 mobile industry publications published between 2005 and 2024, we embed documents with SPECTER2, build year-specific embedding distributions, and derive [...] Read more.
The mobile industry has experienced long-run changes in its knowledge structure, including identifiable transition points observable through embedding-based semantic analysis. Using abstracts from 86,674 mobile industry publications published between 2005 and 2024, we embed documents with SPECTER2, build year-specific embedding distributions, and derive knowledge regimes by combining change-point detection with inter-year distribution distances. We then extract regime-specific topics via clustering and reconstruct topic lineages by aligning topic similarities to classify inheritance, differentiation, convergence, and disappearance. The analysis delineates three regimes spanning 2005 to 2012, 2013 to 2019, and 2020 to 2024, with pronounced transitions around 2012 to 2013 and 2019 to 2020. Regime 1 centers on foundational technologies such as wireless communication, power, sensors, and reliability. Regime 2 expands toward platforms, apps, and data analytics alongside cross-domain convergence. Regime 3 is characterized by strengthened 5G operations and data-driven services, together with the independent rise in policy, governance, and regulation topics. Transitions reflect recombination built on inherited knowledge rather than abrupt replacement, and post-transition topics display distinct growth typologies by network position and growth pattern. By integrating embedding-based changepoint detection with topic lineage reconstruction, we provide a reproducible account of regime transitions and quantitative evidence to inform the timing of corporate R&D, standard and platform strategies, and policy and regulatory design. Full article
Show Figures

Figure 1

26 pages, 1536 KB  
Article
GraphGPT-Patent: Time-Aware Graph Foundation Modeling on Semantic Similarity Document Graphs for Grant-Time Economic Impact Prediction
by Tianhui Fang, Junru Si, Chi Ye and Hailong Shi
Appl. Sci. 2026, 16(6), 2737; https://doi.org/10.3390/app16062737 - 12 Mar 2026
Viewed by 821
Abstract
Predicting the future impact of technical economic documents at release time is challenging due to delayed supervision signals, long-tailed label distributions, and time- and domain-dependent shifts in language and topics. Moreover, similarity graphs derived from text embeddings can be noisy due to boilerplate [...] Read more.
Predicting the future impact of technical economic documents at release time is challenging due to delayed supervision signals, long-tailed label distributions, and time- and domain-dependent shifts in language and topics. Moreover, similarity graphs derived from text embeddings can be noisy due to boilerplate and evolve under temporal drift, making robustness and leakage-free evaluation essential. We formulate grant-time patent impact prediction as a node classification and within-domain ranking problem on a large-scale semantic similarity document graph built from patent text embeddings, avoiding any future citation leakage. The document graph is constructed via ANN Top-K retrieval and similarity thresholding, enabling scalable and reproducible sparsification on hundreds of thousands of nodes. We propose GraphGPT-Patent, which adapts a reversible graph-to-sequence foundation backbone to local subgraphs extracted from the similarity network. The model incorporates time- and domain-conditioned edge reliability to suppress drift-induced and template-driven pseudo-similarity, and optimizes a joint objective coupling high-impact classification with ranking consistency within comparable groups. Experiments on USPTO granted patents (2000–2022) across three high-volume CPC domains and three evaluation horizons show consistent gains over text-only and GNN baselines, achieving up to 0.94 recall for the positive class and improved macro-average recall across nine settings. Temporal shift analyses further quantify the effect of training-data freshness, while explanation subgraphs provide auditable structural evidence of model decisions. The proposed framework offers an effective graph-based learning pipeline for scalable impact prediction and downstream triage under strict information constraints. Full article
Show Figures

Figure 1

14 pages, 3450 KB  
Article
From the Lab to the Land: Challenges of Upscaling Biobased Materials for Architecture
by Mercedes Garcia-Holguera
Appl. Sci. 2026, 16(4), 1990; https://doi.org/10.3390/app16041990 - 17 Feb 2026
Cited by 2 | Viewed by 786
Abstract
The field of biology offers great inspiration for sustainable design solutions through the exploration and implementation of biobased materials in architecture. Research on this topic is increasingly viewed as a key pathway to addressing climate change, partly because biobased materials have lower embedded [...] Read more.
The field of biology offers great inspiration for sustainable design solutions through the exploration and implementation of biobased materials in architecture. Research on this topic is increasingly viewed as a key pathway to addressing climate change, partly because biobased materials have lower embedded energy, can be integrated into circular economy strategies, can be produced locally, and in some cases, biobased materials have been shown to have similar or improved mechanical and hygrothermal properties compared to standard construction materials. However, significant challenges need to be addressed to facilitate a smooth and consistent transition toward a biobased construction industry. Some of these barriers relate to growth processes, cultural perceptions, standardization, and mass production of materials. Another barrier is transitioning from micro-scale structures developed in laboratory settings to metre-scale structures used in architectural applications. Upscaling biobased materials requires adjustments in growth techniques, workspaces, material manipulation tools, and post-processing to ensure the materials meet the requirements for use in the built environment. This document examines bacterial cellulose in this context, illustrating the process followed to upscale the production of the material and adapt it from a controlled lab environment to a larger architectural scale. The study presents and assesses the steps taken to adapt lab growing conditions, harvesting and drying techniques, and coating choices, among other critical procedures. The barriers and opportunities encountered through this process contribute to the ongoing discussion on shifting from traditional to biobased materials in the built environment. Moreover, this research underscores the transformative role that biobased materials like bacterial cellulose can play in advancing sustainable architectural practices and highlights the importance of interdisciplinary efforts to bridge laboratory research and large-scale built design. Full article
Show Figures

Figure 1

14 pages, 488 KB  
Article
The Evolution of Nanoparticle Regulation: A Meta-Analysis of Research Trends and Historical Parallels (2015–2025)
by Sung-Kwang Shin, Niti Sharma, Seong Soo A. An and Meyoung-Kon (Jerry) Kim
Nanomaterials 2026, 16(2), 134; https://doi.org/10.3390/nano16020134 - 19 Jan 2026
Cited by 4 | Viewed by 2237
Abstract
Objective: We analyzed nanoparticle regulation research to examine the evolution of regulatory frameworks, identify major thematic structures, and evaluate current challenges in the governance of rapidly advancing nanotechnologies. By drawing parallels with the historical development of radiation regulation, the study aimed to [...] Read more.
Objective: We analyzed nanoparticle regulation research to examine the evolution of regulatory frameworks, identify major thematic structures, and evaluate current challenges in the governance of rapidly advancing nanotechnologies. By drawing parallels with the historical development of radiation regulation, the study aimed to contextualize emerging regulatory strategies and derive lessons for future governance. Methods: A total of 9095 PubMed-indexed articles published between January 2015 and October 2025 were analyzed using text mining, keyword frequency analysis, and topic modeling. Preprocessed titles and abstracts were transformed into a TF-IDF (Term Frequency–Inverse Document Frequency) document–term matrix, and NMF (Non-negative Matrix Factorization) was applied to extract semantically coherent topics. Candidate topic numbers (K = 1–12) were evaluated using UMass coherence scores and qualitative interpretability criteria to determine the optimal topic structure. Results: Six major research topics were identified, spanning energy and sensor applications, metal oxide toxicity, antibacterial silver nanoparticles, cancer nano-therapy, and nanoparticle-enabled drug and mRNA delivery. Publication output increased markedly after 2019 with interdisciplinary journals driving much of the growth. Regulatory considerations were increasingly embedded within experimental and biomedical research, particularly in safety assessment and environmental impact analyses. Conclusions: Nanoparticle regulation matured into a dynamic multidisciplinary field. Regulatory efforts should prioritize adaptive, data-informed, and internationally harmonized frameworks that support innovation while ensuring human and environmental safety. These findings provide a data-driven overview of how regulatory thinking was evolved alongside scientific development and highlight areas where future governance efforts were most urgently needed. Full article
(This article belongs to the Section Environmental Nanoscience and Nanotechnology)
Show Figures

Figure 1

33 pages, 465 KB  
Article
A Multi-Stage NLP Framework for Knowledge Discovery from Crop Disease Research Literature
by Jantima Polpinij, Manasawee Kaenampornpan, Christopher S. G. Khoo, Wei-Ning Cheng and Bancha Luaphol
Mathematics 2026, 14(2), 299; https://doi.org/10.3390/math14020299 - 14 Jan 2026
Cited by 1 | Viewed by 1173
Abstract
Extracting and organizing knowledge from the agricultural crop disease research literature are challenging tasks because of the heterogeneous terminologies, complicated symptom descriptions, and unstructured nature of scientific documents. In this study, we developed a multi-stage natural language processing (NLP) pipeline to automate knowledge [...] Read more.
Extracting and organizing knowledge from the agricultural crop disease research literature are challenging tasks because of the heterogeneous terminologies, complicated symptom descriptions, and unstructured nature of scientific documents. In this study, we developed a multi-stage natural language processing (NLP) pipeline to automate knowledge extraction, organization, and integration from the agricultural research literature into a domain-consistent crop disease knowledge graph. The model combines transformer-based sentence embeddings with variational deep clustering to extract topics, which are further refined via facet-aware relevance scoring for sentence selection to be included in the summary. Lexicon-guided named entity recognition helps in the precise identification and normalization of terms for crops, diseases, symptoms, etc. Relation extraction based on a combination of lexical, semantic, and contextual features leads to the meaningful generation of triplets for the knowledge graph. The experimental results show that the method yielded consistently good results at each stage of the knowledge extraction process. Among the combinations of embedding and deep clustering methods, SciBERT + VaDE achieved the best clustering results. The extraction of representative sentences for disease symptoms, control/treatment, and prevention obtained high F1-scores of around 0.8. The resulting knowledge graph has high node coverage and high relation completeness, as well as high precision and recall in triplet generation. The multi-stage NLP pipeline effectively converts unstructured agricultural research texts into a coherent and semantically rich knowledge graph, providing a basis for further research in crop disease analysis, knowledge retrieval, and data-driven decision support in agricultural informatics. Full article
Show Figures

Figure 1

Back to TopTop