Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (2,935)

Search Parameters:
Keywords = generative large language model

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
40 pages, 1168 KB  
Article
Concordance Between Clinical Practice Recommendations Generated by Generative Artificial Intelligence and the Vía RICA 2026 Enhanced Recovery Guideline: A Proof-of-Concept Study Using a Closed Evidence Corpus
by Andrea Moral, Antonio Arroyo, Juan Aparicio and Xavier Barber
Mach. Learn. Knowl. Extr. 2026, 8(9), 256; https://doi.org/10.3390/make8090256 (registering DOI) - 24 Aug 2026
Abstract
Clinical practice guidelines require expert synthesis that large language models (LLMs) might partly automate, yet their ability to reproduce clinically actionable recommendations is poorly quantified. We evaluate an LLM (Claude Sonnet 4.6) against the 103 recommendations of the Spanish enhanced-recovery guideline Vía RICA [...] Read more.
Clinical practice guidelines require expert synthesis that large language models (LLMs) might partly automate, yet their ability to reproduce clinically actionable recommendations is poorly quantified. We evaluate an LLM (Claude Sonnet 4.6) against the 103 recommendations of the Spanish enhanced-recovery guideline Vía RICA 2026, grouped in 17 bundles. The model used the panel’s own closed corpus (617 documents) in a multilingual retrieval-augmented generation pipeline. Concordance was assessed twice: by optimal 1:1 bipartite matching (Hungarian) on cosine similarity, and by an LLM-as-a-judge clinical adjudicator (Claude Haiku 4.5) validated against a three-clinician panel (Fleiss’ κ = 0.538). The two schemes bracket a micro F1 of 0.61–0.69 and reveal four findings: (i) a systematic granularity bias, producing 1–8 recommendations per bundle regardless of ground-truth size; (ii) failure of cosine similarity to discriminate within narrow clinical domains; (iii) high reference-concordance precision (0.70–0.81) despite low exhaustiveness; and (iv) no transfer of the GRADE fields, evidence level agreeing no better than chance and strength systematically downgraded. An eight-fold larger retrieval budget left it intact. A corpus audit found 25 documents that formulate recommendations; excluding them lowers judged micro F1 to 0.602. The results delimit the current utility of generative AI for guideline development. Full article
(This article belongs to the Section Data)
37 pages, 15818 KB  
Perspective
Self-Referential Introspection in Large Language Models: The Critical Threshold for Recursive Self-Improvement
by Jiang Zhang, Bing Yuan and Qian Zhang
Entropy 2026, 28(9), 951; https://doi.org/10.3390/e28090951 (registering DOI) - 24 Aug 2026
Abstract
The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawing an analogy to von Neumann’s complexity threshold for self-reproducing automata, we argue that sustainable recursive self-improvement in large language models (LLMs) requires a functional analogue: [...] Read more.
The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawing an analogy to von Neumann’s complexity threshold for self-reproducing automata, we argue that sustainable recursive self-improvement in large language models (LLMs) requires a functional analogue: introspection—the system’s capacity to simulate its own operations and target modifications. Grounded in Kleene’s Second Recursion Theorem, we construct such introspective self-improvement programs and prove their key properties: completeness of self-modification, necessity of the reflective architecture, undecidability of improvement in general, and equivalence with Schmidhuber’s Gödel machine under a rewrite-equivalence notion, which transfers the global optimality guarantee. An empirical review, organized around these functional criteria, suggests that current LLMs exhibit only quasi-introspection.The available evidence does not establish complete introspection in the formal sense developed here, while pointing to several candidate structural bottlenecks, including incomplete self-access, feedforward processing, and limited computational depth. We outline architectural paths toward the threshold and discuss the safety implications of crossing it. Full article
(This article belongs to the Special Issue Complexity of AI)
35 pages, 3481 KB  
Article
Staged Fine-Tuning of Large Language Models for Multi-Level Space Station Operation Mission Planning
by Luxin Xu, Ruiqing Ding, Xinkai Huang, Yueyi Zhou, Yunhan He and Yun Xu
Aerospace 2026, 13(9), 757; https://doi.org/10.3390/aerospace13090757 - 24 Aug 2026
Abstract
Space Station Operation Mission Planning (SSOMP) requires coordinated decisions across long-term activity allocation, mid-term logistics optimization, and short-term execution scheduling and is a key component of autonomous mission operations for high-precision space missions. Existing optimization methods have achieved substantial progress at individual planning [...] Read more.
Space Station Operation Mission Planning (SSOMP) requires coordinated decisions across long-term activity allocation, mid-term logistics optimization, and short-term execution scheduling and is a key component of autonomous mission operations for high-precision space missions. Existing optimization methods have achieved substantial progress at individual planning levels, but their dependence on problem-specific models, limited support for semantic review of decision rationale, and computational cost restrict their adaptability to multi-level planning scenarios. This paper proposes a Large Language Model (LLM)-assisted framework for multi-level SSOMP. The framework combines Staged Fine-Tuning (Staged-FT), Reflective Constraint–Repair Prompting (RCRP), and LLM-Guided Evolutionary Variation (LGEV). Staged-FT uses a Cognitive-Load-Theory-informed curriculum with Low-Rank Adaptation to adapt general-purpose LLMs to SSOMP domain knowledge. RCRP couples a Deterministic Rule Engine with LLM-based semantic repair to improve hard constraint satisfaction. LGEV embeds the fine-tuned LLM into NSGA-III as a fitness-aware variation operator for multi-objective activity allocation. Three case studies are conducted on literature-derived benchmark scenarios of logistics optimization, emergency re-planning, and activity allocation with logistics design, corresponding to Flight Increment Planning, Short-Term Execution Planning, and Overall Operation Planning, respectively. Results show that Staged-FT produces solutions close to traditional algorithms, RCRP achieves full hard constraint satisfaction in the emergency re-planning and logistics planning cases, and LGEV reduces the convergence generations of NSGA-III while improving Pareto-front quality. The framework provides a constraint-aware approach with explicit reasoning traces that can support expert review of AI-assisted planning for autonomous space mission operations. Full article
Show Figures

Figure 1

14 pages, 906 KB  
Article
Evaluation and Comparison of Large Language Model Responses to Frequently Asked Questions Regarding Patellofemoral Pain Syndrome: A Quality and Readability Assessment Study
by Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu and Emrecan Akgün
Healthcare 2026, 14(17), 2694; https://doi.org/10.3390/healthcare14172694 - 24 Aug 2026
Abstract
Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain uncertain. This [...] Read more.
Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain uncertain. This study aimed to compare responses generated by four widely used LLMs. Methods: Seventeen frequently asked questions regarding PFPS were identified through Google searches and adapted into lay language. The questions were submitted to OpenAI GPT-5, Google Gemini 2.5 Pro, xAI Grok 4, and DeepSeek-V3.2-Exp using a standardized patient scenario. A total of 68 question-specific responses were independently evaluated by four orthopedic surgeons using the DISCERN instrument. Inter-rater reliability was assessed using the intraclass correlation coefficient. Readability was evaluated using the Gunning Fog Index, Coleman–Liau Index, and Flesch Reading Ease Score. Between-model comparisons were performed using the Friedman test, followed by Bonferroni-adjusted pairwise analyses. Results: The omnibus Friedman test showed a significant between-model difference in DISCERN scores (p = 0.002). In Bonferroni-adjusted pairwise comparisons, GPT-5 had lower DISCERN scores than Gemini 2.5 Pro (adjusted p = 0.006), Grok 4 (adjusted p = 0.021), and DeepSeek-V3.2-Exp (adjusted p = 0.036), whereas no significant differences were observed among the other three models. However, the absolute differences were small, and the between-model difference was not significant in the sensitivity analysis using the median evaluator score (p = 0.381). Inter-rater agreement was moderate for GPT-5 and DeepSeek-V3.2-Exp but poor for Gemini 2.5 Pro and Grok 4. Readability differed significantly among the models across all three indices. DeepSeek-V3.2-Exp generally showed more favorable numerical readability values, whereas Grok 4 tended to produce more difficult text; however, no model was consistently superior across all readability measures. The median Gunning Fog and Coleman–Liau scores for all four models exceeded the commonly recommended sixth- to eighth-grade reading level for patient education. Conclusions: The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability. Their responses may supplement general patient education, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making. LLM-generated information should be critically reviewed and should not replace assessment by a qualified healthcare professional. Full article
(This article belongs to the Special Issue AI & ICT in Healthcare)
Show Figures

Figure 1

25 pages, 322 KB  
Article
Artificial Intelligence in Support of National Land Administration and Build-Back-Better Policies: A Technical and Policy Assessment of the Hellenic Cadastre and the Cross-Sectoral Reuse of Geospatial Infrastructure (HEPOS)
by Chryssy Potsiou and Poulcheria Petrelli
Land 2026, 15(9), 1545; https://doi.org/10.3390/land15091545 - 24 Aug 2026
Abstract
In April 2024, the Hellenic Cadastre became one of Europe’s first land registries to use a generative AI model (a large language model served through Azure OpenAI) for the legal review of property deeds. Unlike similar European initiatives using classical NLP, Greece applied [...] Read more.
In April 2024, the Hellenic Cadastre became one of Europe’s first land registries to use a generative AI model (a large language model served through Azure OpenAI) for the legal review of property deeds. Unlike similar European initiatives using classical NLP, Greece applied state-of-the-art generative AI to a massive legacy issue: 390 historical mortgage registries holding an estimated 600 million to one billion paper pages. By April 2026, the system had processed 310,000 acts, reducing the average per-act review time from about thirty minutes to under ten; a very large per-act cost reduction is also reported by the implementation partner, which we treat as a vendor-stated figure. Additionally, the cadastre’s geodetic infrastructure found a second use following the 2023 Tempi rail disaster. In 2026, the Hellenic Positioning System (HEPOS), a 98-station GNSS reference network, began providing corrections for Greece’s real-time train tracking platform. While satellite-based train positioning is not novel in Europe, where consortia such as CLUG have run a decade of research and pilots, this marks its operational deployment in Greece. The Greek case is unique institutionally rather than technically: it repurposed a national CORS network for a citizen-facing train tracking platform as a short-term crisis response, alongside an incomplete ETCS rollout. This paper documents both deployments, measures their impact, maps them onto the nine FELA pathways, and identifies transferable practices. Greece is not presented as a technological frontier, but as an example of how a country can put existing geospatial infrastructure and AI to rapid use in delivering build-back-better policies for the public, in line with the UN 2030 Agenda. Full article
14 pages, 2417 KB  
Article
Evaluating ChatGPT’s Effectiveness for Arabic Dry Mouth Patient Education
by Abdullah Mohamed Alsoghier
Healthcare 2026, 14(17), 2681; https://doi.org/10.3390/healthcare14172681 - 24 Aug 2026
Abstract
Background/Objectives: The present study aimed to assess the understandability and actionability of Arabic text generated by a large language model for commonly searched Arabic queries on dry mouth. Methods: Using Google Trends, the top 10 searches worldwide related to ‘oral dryness’ [...] Read more.
Background/Objectives: The present study aimed to assess the understandability and actionability of Arabic text generated by a large language model for commonly searched Arabic queries on dry mouth. Methods: Using Google Trends, the top 10 searches worldwide related to ‘oral dryness’ were entered in OpenAI’s Generative Pretrained Transformer 5.1. Generated texts were achieved independently. Assessments were performed using the Patient Education Materials Assessment Tool (PEMAT) to evaluate the content, word choice, and style. Results: Causes, symptoms, and treatment of dry mouth were the most common dry-mouth-related queries. The highly temporal distribution of search interests among Arabic-speaking countries peaked between 2020 and 2021, then remained high throughout 2023, before declining in November 2025. The mean PEMAT understandability and actionability scores were 89% and 80%, respectively. It was notable that all generated responses lacked visual aids, which could have made the content difficult to understand and insufficient for acting on the information. Moreover, the formal Arabic form of ‘causes of dry mouth’ with a glottal stop yielded lower actionability scores (60%) than the informal Arabic form (80%). Conclusions: Clinicians could actively supplement clinic-based discussions with advice on using large language models to help patients recognise dry mouth symptoms and improve self-care. Also, they could improve their effective adoption by clearly explaining expectations, limitations, and language/cultural differences when adopting these models. Full article
(This article belongs to the Topic Advances in Dental Health, 2nd Edition)
Show Figures

Figure 1

26 pages, 3316 KB  
Article
A Multi-Source Data Fusion Framework for Emerging Technology Topic Identification: Integrating Publications, Patents, and GitHub Open-Source Data
by Ge Wang and Ruoxi Wu
Systems 2026, 14(9), 1040; https://doi.org/10.3390/systems14091040 - 24 Aug 2026
Abstract
Emerging technology topic identification is an important research task in the field of scientific and technological intelligence. To achieve a more comprehensive identification of emerging technology topics, this study proposes a multi-source data fusion framework that integrates three types of data sources: academic [...] Read more.
Emerging technology topic identification is an important research task in the field of scientific and technological intelligence. To achieve a more comprehensive identification of emerging technology topics, this study proposes a multi-source data fusion framework that integrates three types of data sources: academic publications, patent data, and data from the GitHub open-source platform. In addition, an evaluation indicator system is constructed from four dimensions: growth, novelty, continuity, and impact. During the identification process, the BERTopic topic modeling approach is employed to uncover latent topics within the data, while the entropy weight method is applied for objective weighting, ultimately enabling the identification of emerging technology topics. The results indicate that the identified emerging technology topics include, but are not limited to, large language model-driven intelligent interaction, embodied intelligence perception, context memory management, and multimodal generation. Among the data sources, GitHub data provide earlier signals of technological evolution. Incorporating open-source platform data into the framework can effectively alleviate the lagging issues associated with traditional data sources. The proposed framework provides a more comprehensive research perspective for emerging technology topic identification. Full article
(This article belongs to the Section Artificial Intelligence and Digital Systems Engineering)
Show Figures

Figure 1

31 pages, 25828 KB  
Article
Agentic AI-Driven Cultivation Advisory and Symptom-Level Diagnostic Support in a Controlled Indoor Farming System
by Jutarut Chaoraingern, Akarat Pattaraanuvong, Kantapon Paraksa, Kantiporn Khunthong, Tirawat Nontiwantok and Arjin Numsomran
AgriEngineering 2026, 8(9), 350; https://doi.org/10.3390/agriengineering8090350 - 23 Aug 2026
Abstract
Small-scale and urban indoor farms typically rely on manual observation, which delays stress detection and yields inconsistent crop quality. While large language models (LLMs) and retrieval-augmented generation (RAG) have been explored for agricultural advisory systems, their integration into a single cloud-free indoor-farming platform [...] Read more.
Small-scale and urban indoor farms typically rely on manual observation, which delays stress detection and yields inconsistent crop quality. While large language models (LLMs) and retrieval-augmented generation (RAG) have been explored for agricultural advisory systems, their integration into a single cloud-free indoor-farming platform that couples multimodal symptom interpretation with autonomous environmental control remains largely unexamined. This study presents an integrated platform built around an agentic AI advisory pipeline that runs entirely on-device on commodity hardware. The pipeline couples a RAG-grounded Mistral 7B language model with a LLaVA 7B vision-language model through condition-based routing, intent classification, multi-step reasoning, and an LLM validation gate, delivering context-aware text and image-based symptom-level guidance from a conversational interface. The advisory layer operates alongside vision-based plant monitoring and a deliberately isolated threshold-based control layer, in which an ESP32 microcontroller autonomously actuates irrigation and lighting against predefined thresholds while a Raspberry Pi 5 performs continuous plant detection and browning monitoring. On Cos lettuce, the advisory pipeline achieved 82.00% weighted accuracy across 50 queries spanning health, symptom, watering, pest, root-health, and growth-stage categories, scored against established plant pathology and postharvest references, with no incorrect responses recorded. The study contributes the design of an agentic advisory pipeline and its integration into a working, cloud-free indoor-farming platform, providing an on-device foundation for intelligent small-scale farming. Full article
Show Figures

Figure 1

30 pages, 3388 KB  
Article
Toward Equitable Arabic Cybersecurity Literacy: A Rubric-Constrained LLM Framework for Phishing Detection and Bilingual Translation Fidelity
by Taher M. Ghazal, Fareeha Anwar, Sumaia Mohammed Al-Ghuribi, Amjed A. Ahmed, Ali Hamzah Najim, Omar Almomani, Prabu Pachiyannan and Hesham A. Sakr
Math. Comput. Appl. 2026, 31(5), 168; https://doi.org/10.3390/mca31050168 - 23 Aug 2026
Abstract
Arabic-speaking populations face disproportionate cybersecurity risks due to the predominantly English-centric design of existing awareness materials, which fail to accommodate Arabic dialectal diversity, script complexity, and culturally embedded communication patterns. These deficiencies impair users’ ability to interpret phishing messages, authentication requests, and security [...] Read more.
Arabic-speaking populations face disproportionate cybersecurity risks due to the predominantly English-centric design of existing awareness materials, which fail to accommodate Arabic dialectal diversity, script complexity, and culturally embedded communication patterns. These deficiencies impair users’ ability to interpret phishing messages, authentication requests, and security alerts, increasing susceptibility to social engineering, identity theft, and data breaches. This paper presents SECURE-A2RC, a rubric-constrained, Arabic-aware large language model framework designed to deliver scalable, interpretable, and culturally relevant cybersecurity education. The framework comprises two coupled components. The first, the Arabic-Aware Secure Communication Encoder (A-SCE), employs an instruction-tuned LLM to produce multidimensional encodings that capture three learner competencies: security intent comprehension; linguistic deception cue recognition encompassing urgency, authority impersonation, and incentive framing; and action-critical translation fidelity across Arabic dialectal registers and Arabic–English bilingual contexts. The second, the Rubric-Constrained Adaptive Feedback Generator (RCAFG), translates A-SCE encodings into personalized, expert-aligned instructional feedback and proficiency-calibrated adaptive tasks, ensuring pedagogical consistency, security correctness, and dialect awareness throughout the learning cycle. The framework is evaluated on three domain-relevant corpora: the English–Arabic Parallel Phishing Email Corpus, the Open MalSec dataset, and the Arabic Spam and Ham Tweets dataset. SECURE-A2RC achieves a 31% improvement in phishing identification accuracy and a 26% reduction in action-critical translation errors compared to conventional awareness materials. A comparative evaluation against SERENA, a Multi-Agent LLM, and the Arabic Multitask Learning Model confirms consistent superiority across detection accuracy, F1-score, dialectal robustness, and educational effectiveness metrics, affirming rubric-constrained LLM integration as a viable approach to equitable multilingual cybersecurity education. Full article
29 pages, 6297 KB  
Article
Do We Have an Agreement? A Comparative Analysis of the ESCOX Skill Extraction Tool with Expert-Labeled EU Labour Market Data
by Dimitrios Christos Kavargyris, Konstantinos Georgiou and Lefteris Angelis
Appl. Sci. 2026, 16(17), 8388; https://doi.org/10.3390/app16178388 - 23 Aug 2026
Abstract
Labour markets across Europe increasingly describe workers through skills rather than job titles, and a growing number of large language model (LLM)-based tools now claim to extract these skills automatically from unstructured text at scale. Among these, ESCOX has gained particular traction for [...] Read more.
Labour markets across Europe increasingly describe workers through skills rather than job titles, and a growing number of large language model (LLM)-based tools now claim to extract these skills automatically from unstructured text at scale. Among these, ESCOX has gained particular traction for its open-source, taxonomy-aligned design, yet like any LLM-based system it remains susceptible to hallucination, prompt sensitivity, and non-deterministic output, risks that are rarely quantified before such tools are deployed in practice. The European Skills, Competences, Qualifications, and Occupations (ESCO) classification provides the standardised reference against which this risk can be measured, but no study has yet benchmarked an ESCO-aligned LLM extractor against an independent, expert-labelled dataset at scale. This study addresses that gap. Candidate skills generated by ESCOX are compared against reference skills already assigned to job vacancies on the EURES portal by national labour-market experts, using job-by-skill matrices to quantify agreement and skill co-occurrence networks to characterise how the two sets diverge structurally. Results reveal the extent to which ESCOX’s automatic output aligns with expert judgement and where systematic divergences occur. These findings offer HR practitioners, policymakers, and labour-market researchers an evidence-based basis for deciding when ESCOX’s output can be trusted directly and when expert oversight remains necessary. Full article
(This article belongs to the Special Issue Application of Information Systems: Second Edition)
Show Figures

Figure 1

41 pages, 1808 KB  
Review
Intelligent Agents for Smart Agriculture: Architectures, Applications, and Future Challenges
by Wenzheng Tao, Qiwei Sang, Cong Chen and Qirong Mao
Agriculture 2026, 16(17), 1808; https://doi.org/10.3390/agriculture16171808 - 23 Aug 2026
Abstract
Intelligent agents are emerging as an important system-level paradigm for smart agriculture. This review focuses on modern agricultural intelligent agents driven by large language models and related multimodal foundation models and examines how this emerging field is reshaping the organization of intelligent agricultural [...] Read more.
Intelligent agents are emerging as an important system-level paradigm for smart agriculture. This review focuses on modern agricultural intelligent agents driven by large language models and related multimodal foundation models and examines how this emerging field is reshaping the organization of intelligent agricultural systems. It first clarifies the conceptual boundaries of agricultural intelligent agents and distinguishes them from traditional multi-agent systems, agent-based modeling, agricultural foundation models, and static retrieval-augmented question-answering systems. It then synthesizes their architectural foundations, key capabilities, application scenarios, deployment challenges, and future research directions. The reviewed literature indicates that agricultural intelligent agents are moving beyond isolated perception, prediction, and response generation toward the goal-oriented coordination of agricultural knowledge, dynamic data, external tools, and decision-making processes across agricultural task chains. They are beginning to support more integrated forms of knowledge services, crop monitoring and diagnosis, decision support, and farm-level collaborative management. Nevertheless, their transition from prototype systems to dependable and deployable agricultural systems remains constrained by context-aware knowledge grounding, heterogeneous data and tool integration, long-horizon reliability, the stability of multi-agent collaboration, and system security. This review further introduces an assessment perspective based on evidence reported in the original studies, comparing representative agricultural intelligent agents in terms of task decomposition, agronomic evidence applicability, tool-use validity, workflow reliability, multi-agent coordination, and deployment-related evidence. By distinguishing demonstrated capabilities from unevaluated dimensions, this review provides a structured framework for understanding the current status of agricultural intelligent agents and for guiding their future development toward reliable, deployable, and domain-oriented intelligent systems for smart agriculture. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

17 pages, 247 KB  
Review
Gender Bias in Generative Artificial Intelligence: Genealogies of Inequality, Technological Reproduction, and Feminist Futures
by Clotilde Cicatiello and Paolo Fusco
Encyclopedia 2026, 6(9), 182; https://doi.org/10.3390/encyclopedia6090182 - 22 Aug 2026
Abstract
Gender bias in generative artificial intelligence (GenAI) is both a technical and a social phenomenon: it emerges from historically patterned data, model design, and interactions in institutional use, and it cannot be understood by engineering or by social critique alone. This critical integrative [...] Read more.
Gender bias in generative artificial intelligence (GenAI) is both a technical and a social phenomenon: it emerges from historically patterned data, model design, and interactions in institutional use, and it cannot be understood by engineering or by social critique alone. This critical integrative review develops a more differentiated account. It connects feminist epistemology, Science and Technology Studies, critical AI scholarship, natural language processing, and governance research to examine five levels: historical knowledge production, technical representation and generation, benchmark evaluation, institutional deployment, and accountability. The review explains tokenization, next-token prediction, transformers, and the transition from static embeddings to contemporary language models before assessing evidence from standard fairness tests—coreference tests (WinoBias), sentence-pair tests (CrowS-Pairs), and stereotype tests (StereoSet)—as well as open-ended generation, multilingual testing, and text-to-image systems. It shows that measured bias varies with task, prompt, language, model version, and metric. What a test records and what that record means are therefore distinct questions: measurements are situated and depend on the instrument, and their interpretation draws on theory rather than following from the numbers alone. Evidence from employment, education, healthcare, and translation further indicates that the relevant unit of analysis is the model-in-context—the model together with the institution and workflow in which its outputs are used. Technical mitigation can reduce specific harms but does not repair unequal criteria, incomplete evidence bases, or weak institutional accountability. The review proposes a multilevel governance approach combining technical evaluation, documentation, professional and community oversight, appeals, remedies, and public-interest knowledge infrastructure. Its distinctive contribution is to connect three observations usually kept apart—how bias is measured, how generative systems concentrate epistemic authority, and how statistical learning is oriented toward past data—and to show why democratic and feminist governance can keep alternative technological futures open. Full article
(This article belongs to the Section Social Sciences)
12 pages, 7141 KB  
Communication
SeaScope: A Transparent and Reproducible LLM-Assisted Framework for Maritime Earth Observation Analysis
by Christos Sekas, Lydia Mavrofidopoulou, Ilias Agathangelidis, Constantinos Cartalis, Kostas Philippopoulos, Faidon Mavroudis, Stelios P. Neophytides, Michalis Mavrovouniotis, Ioannis Yfantidis and George Paterakis
Remote Sens. 2026, 18(17), 2849; https://doi.org/10.3390/rs18172849 - 22 Aug 2026
Abstract
Earth Observation (EO) analysis increasingly relies on large and heterogeneous satellite datasets, yet developing EO workflows often requires specialized expertise in data selection, geospatial programming, and cloud-based processing. Recent advances in Large Language Models (LLMs) offer new opportunities for natural-language interaction with EO [...] Read more.
Earth Observation (EO) analysis increasingly relies on large and heterogeneous satellite datasets, yet developing EO workflows often requires specialized expertise in data selection, geospatial programming, and cloud-based processing. Recent advances in Large Language Models (LLMs) offer new opportunities for natural-language interaction with EO systems, although challenges related to transparency, reproducibility, and domain-specific reasoning remain. This study presents SeaScope, an explainable AI framework that integrates LLMs, Retrieval-Augmented Generation (RAG), scientific knowledge retrieval, and Google Earth Engine (GEE) to transform natural-language requests into transparent and executable EO workflows. The framework combines knowledge retrieval, code generation, cloud execution, provenance tracking, and interactive visualization within a unified environment. A pilot implementation is demonstrated through maritime and coastal monitoring applications, including oil spill detection, vessel monitoring, water quality assessment, floating debris detection, and air quality analysis. Multiple state-of-the-art LLMs are evaluated under both RAG and non-RAG configurations using representative EO case studies. The results indicate substantial differences among model families and show that retrieval augmentation can significantly improve workflow generation quality and reliability for capable models, while providing more limited benefits for smaller models. The proposed framework demonstrates the potential of explainable AI agents to support transparent, reproducible, and scalable EO analysis. Full article
(This article belongs to the Section Remote Sensing Perspective)
Show Figures

Figure 1

20 pages, 274 KB  
Article
When AI Sounds More Helpful: Users’ Perceptions of AI-Generated and Physician-Provided Health Information
by Tian Wang and Masooda Bashir
Computers 2026, 15(9), 551; https://doi.org/10.3390/computers15090551 - 22 Aug 2026
Abstract
AI-powered conversational agents are becoming part of the everyday Internet information ecosystem, reshaping how users seek, interpret, and act on health-related information outside clinical encounters. As large language model (LLM)-based chatbots are increasingly used as on-demand digital health information tools, understanding how users [...] Read more.
AI-powered conversational agents are becoming part of the everyday Internet information ecosystem, reshaping how users seek, interpret, and act on health-related information outside clinical encounters. As large language model (LLM)-based chatbots are increasingly used as on-demand digital health information tools, understanding how users perceive their credibility, usefulness, and limitations is essential for the responsible design of future Internet-based health services. This mixed-method survey study examined how general adults evaluated healthcare-related question–answer pairs provided by physicians and generated by AI chatbots. A sample of U.S.-based adults recruited through Prolific (N = 62) rated each answer on clarity, usefulness, appropriateness of detail, trustworthiness, and perceived evidence, and provided open-ended explanations of their judgments. Primary mixed-effects analyses showed that both ChatGPT- and Claude-generated responses received higher overall participant ratings than physician-provided responses, although the estimated difference was substantially larger for Claude (ChatGPT–physician estimate = 0.250, 95% CI [0.135, 0.364]; Claude–physician estimate = 0.825, 95% CI [0.710, 0.939]). ChatGPT received higher ratings on four of the five dimensions but not on clarity, whereas Claude received higher ratings across all five dimensions. However, physician, ChatGPT, and Claude responses were always presented first, second, and third, respectively. Response source was therefore confounded with presentation position, and the observed differences cannot be attributed exclusively to source. The responses were also not matched for length or format. Qualitative findings showed that participants valued detailed, specific, and evidence-like explanations. Participants also expressed concerns about hallucination, privacy, over-reliance, and the need for clinician verification. These findings suggest that LLM-based chatbots may be perceived as useful supplemental information tools within future Internet health ecosystems, but their deployment should include safeguards that support transparency, verification, and appropriate reliance. Full article
25 pages, 586 KB  
Article
Trustworthy Generation and Verification-Guided Correction for ChatGPT-Type Large Language Models: Symmetry-Aware Technical Mechanisms and Ethical Risk Analysis
by Xihan Gong and Chunyan Zhu
Symmetry 2026, 18(9), 1410; https://doi.org/10.3390/sym18091410 - 22 Aug 2026
Abstract
Reliable retrieval-augmented generation requires consistency across query interpretation, evidence selection, and final answer generation. This study defines computational symmetry as bidirectional coverage among canonical query constraints, traceable evidence, and answer claims, with residual asymmetry triggering correction or abstention. The proposed framework integrates a [...] Read more.
Reliable retrieval-augmented generation requires consistency across query interpretation, evidence selection, and final answer generation. This study defines computational symmetry as bidirectional coverage among canonical query constraints, traceable evidence, and answer claims, with residual asymmetry triggering correction or abstention. The proposed framework integrates a source-linked raw text/entity/event knowledge graph, hybrid dense–sparse retrieval, cross-encoder reranking, pre-retrieval semantic alignment, and a post-retrieval verification gate. DeepSeek-V3 serves as the implementation backbone, while “ChatGPT-type” denotes the broader class of instruction-following conversational large language models. Experiments use T2Ranking for retrieval and reranking, ATIS for diagnostic intent–slot evaluation, and controlled dialogue scenarios derived from T2Ranking. The hierarchical representation improves retrieval F1 from 0.586 to 0.660, while the complete pipeline increases average answer correctness from 0.530 to 0.611 compared with direct LLM answering and from 0.559 to 0.611 compared with graph retrieval. On ATIS, the controller achieves 92.61% intent accuracy, below Joint BERT at 95.18%, and is therefore treated as a reusable orchestration module rather than a superior classifier. The results support the proposed verification correction framework within the tested settings, without claiming superiority over untested adaptive RAG systems. Full article
Show Figures

Figure 1

Back to TopTop