Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (780)

Search Parameters:
Keywords = gemini

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
32 pages, 1270 KB  
Article
An Exploratory Study of the Transferability of Smart Mobility Solutions Across Seven European Mobility Living Labs: Towards an Interoperability and Governance Framework
by Shaghayegh Rahnama, David Escuin, David Cipres and Lorena Polo
Smart Cities 2026, 9(9), 145; https://doi.org/10.3390/smartcities9090145 (registering DOI) - 5 Sep 2026
Abstract
Smart city mobility systems face a persistent gap between digital innovation and large-scale deployment. While Mobility-as-a-Service platforms and Mobility Living Labs have generated promising local solutions, transferring these solutions across cities remains constrained by data fragmentation, proprietary platforms, and divergent governance frameworks. This [...] Read more.
Smart city mobility systems face a persistent gap between digital innovation and large-scale deployment. While Mobility-as-a-Service platforms and Mobility Living Labs have generated promising local solutions, transferring these solutions across cities remains constrained by data fragmentation, proprietary platforms, and divergent governance frameworks. This study proposes and operationalizes a policy-oriented framework for assessing the transferability of smart mobility solutions across urban contexts, drawing on the GEMINI project spanning seven European cities: Amsterdam, Copenhagen, Helsinki, Munich, Paris-Saclay, Porto, and Turin. A structured dataset covering application features, interoperability characteristics, technical specifications, and implementation challenges was developed and analysed using comparative indicators including Ease of integration, interoperability level, and replication potential. The results demonstrate that transferability depends on the interaction of technical, governance, and institutional conditions rather than technical compatibility alone. Applications built on open standards and modular architectures show significantly higher replication potential, while legacy systems, fragmented governance structures, and restrictive data-sharing arrangements remain the most persistent barriers. To operationalize the framework, a web-based Knowledge Hub was developed and deployed as a live decision-support platform, currently serving all seven Mobility Living Labs and openly accessible to cities across Europe. The framework offers a replicable model for smart city mobility governance. Full article
24 pages, 11303 KB  
Article
Can Urban Planners Rely on Large Language Models? An Exploratory Evaluation of Four Frontier Models on Two Contested Greek Masterplans
by Aristotelis Vartholomaios and Dionysia Georgia Ch. Perperidou
Smart Cities 2026, 9(9), 143; https://doi.org/10.3390/smartcities9090143 - 3 Sep 2026
Viewed by 295
Abstract
Urban planners are increasingly using Large Language Models (LLMs) to draft project appraisals, raising the question of whether such systems can act as independent evaluators. Here we test this question on two contested urban regeneration projects in Greece: Hellinikon in Athens and TIF-HELEXPO [...] Read more.
Urban planners are increasingly using Large Language Models (LLMs) to draft project appraisals, raising the question of whether such systems can act as independent evaluators. Here we test this question on two contested urban regeneration projects in Greece: Hellinikon in Athens and TIF-HELEXPO in Thessaloniki. A validation rubric was compiled for each case from court rulings, professional body statements, peer-reviewed scholarship and civic-movement documentation. Four frontier LLMs (ChatGPT 5.4, Claude Opus 4.6, Gemini 3.0 Pro, Grok 4) evaluated each case under four prompts: minimal, comprehensive-impartial, comprehensive-sceptical and comprehensive-advocating. Three of the four models covered 77–79% of rubric positions under comprehensive prompts. Coverage depended more on the neutral factsheet than on model choice. When the framing flipped from sceptical to advocating, sceptical-position coverage fell by 9 percentage points and advocating-position coverage rose by 21 percentage points. Positions articulated by local professional and academic bodies were consistently missed. Multi-model ensembles reached 94% and 84% sceptical-position coverage on Hellinikon and TIF-HELEXPO. This evidence is inconclusive but suggests that LLM outputs still require expert verification of locally documented positions, though multi-prompt and multi-model elicitation with a neutral factsheet can support drafting planning reviews. Full article
Show Figures

Figure 1

13 pages, 808 KB  
Proceeding Paper
Comparative Evaluation of AI-Assisted Systematic Literature Reviews: NotebookLM, ChatGPT-4o, and Gemini 1.5 Pro in Wearable Technology, Smart Textiles, and Sustainable Fashion Research
by Selma Bulut, Beyza Buzol Mülayim and Bozhana Stoycheva
Eng. Proc. 2026, 154(1), 29; https://doi.org/10.3390/engproc2026154029 - 3 Sep 2026
Viewed by 117
Abstract
This study compares three AI platforms, Google NotebookLM (Google Labs, May 2026 release/access), ChatGPT-4o, and Gemini 1.5 Pro, for systematic literature review (SLR) tasks in sustainable textile and fashion research. A corpus of 40 peer-reviewed articles (2019–2025) was queried using 12 standardized prompts [...] Read more.
This study compares three AI platforms, Google NotebookLM (Google Labs, May 2026 release/access), ChatGPT-4o, and Gemini 1.5 Pro, for systematic literature review (SLR) tasks in sustainable textile and fashion research. A corpus of 40 peer-reviewed articles (2019–2025) was queried using 12 standardized prompts across four dimensions: source traceability, thematic synthesis, research gap identification, and output utility. NotebookLM achieved the highest weighted score (8.63/10) due to its document-grounded, structurally constrained architecture, which produced no out-of-corpus citations under the conditions of this study. ChatGPT-4o (8.50/10) excelled in narrative synthesis; Gemini 1.5 Pro (7.78/10) produced superior structured tables. Five corpus-grounded and two inferred gaps, including XAI, Federated Learning, and Edge AI, were identified. A four-step hybrid AI workflow is proposed. Full article
Show Figures

Figure 1

14 pages, 3765 KB  
Article
Benchmarking Large Language Models on Long-Tail Plant Taxonomic Knowledge with PTTB-600
by Jian He, Jiamin Xiao, Hong Qu and Lei Xie
Diversity 2026, 18(9), 541; https://doi.org/10.3390/d18090541 - 3 Sep 2026
Viewed by 124
Abstract
Plant taxonomic knowledge contains a long tail of infrequently encountered names, diagnostic characters, and nomenclatural decisions, yet model reliability across this distribution remains unclear. We developed the Chinese-language PTTB-600, comprising 200 general, 300 ordinary specialized, and 100 long-tail fill-in questions, and evaluated 31 [...] Read more.
Plant taxonomic knowledge contains a long tail of infrequently encountered names, diagnostic characters, and nomenclatural decisions, yet model reliability across this distribution remains unclear. We developed the Chinese-language PTTB-600, comprising 200 general, 300 ordinary specialized, and 100 long-tail fill-in questions, and evaluated 31 large language models (LLMs) or run modes under closed-book conditions without retrieval augmentation. The first author drafted the question bank and answer key; three coauthors with doctorates in plant taxonomy reviewed them independently. All models scored at least 197/200 on general questions, and 21 achieved full marks. The six highest-scoring models answered 291–297/300 ordinary specialized questions (97.0–99.0%) but achieved 63.0–90.0% accuracy on long-tail fill-in questions. Gemini 3.1 Pro Preview ranked first at 587/600; ranks two through six formed a closely spaced cluster with no significant adjacent differences after Holm correction. Across 11 within-family comparisons, thinking-mode runs yielded 20–65 additional correct answers, chiefly on specialized and fill-in tasks. Factual errors were uncommon in routine undergraduate content and concentrated in the generation of rare genus names, fine diagnostic distinctions, and alternative nomenclatural treatments. Top-performing LLMs can provide reliable support for routine teaching under instructor oversight, whereas long-tail identifications and nomenclatural decisions require verification against authoritative sources. Full article
(This article belongs to the Section Plant Diversity)
Show Figures

Figure 1

27 pages, 37947 KB  
Article
Generative Artificial Intelligence as a Collaborative Designer for Elementary Computer Science Education
by Fatema Nasrin, Mohammad Mohi Uddin and Feiya Luo
Educ. Sci. 2026, 16(9), 1426; https://doi.org/10.3390/educsci16091426 - 2 Sep 2026
Viewed by 236
Abstract
Artificial Intelligence (AI) is becoming ubiquitous in our everyday lives. In the United States, AI has increasingly been recognized as an important part of computer science (CS) standards for elementary school students. Prior research work has shown that AI-related content can be taught [...] Read more.
Artificial Intelligence (AI) is becoming ubiquitous in our everyday lives. In the United States, AI has increasingly been recognized as an important part of computer science (CS) standards for elementary school students. Prior research work has shown that AI-related content can be taught at upper elementary grade levels. However, classroom-ready, standards-aligned materials for CS in the context of AI remain limited at the elementary level. This paper showcases how educators may use generative AI to design standards-aligned materials for teaching CS and AI in elementary schools. Specifically, this paper presents a design case in which generative AI (Google Gemini) serves as a collaborative designer to co-author four comic-based stories for grades 4–6. The four comic stories, titled “The AI Journey with Milo and Zada,” “Zappy the AI Robot,” “The AI Adventures,” and “Think Like an Engineer,” are mapped to specific AI standards in the Alabama Digital Literacy and Computer Science (DLCS) Course of Study. This work provides a replicable example of using generative AI to co-design standards-aligned instructional materials with responsible use of AI. This work is also useful for instructional designers, elementary school teachers, and policymakers in elementary computer science education. Full article
(This article belongs to the Special Issue K-12 Computer Science Education in the Era of AI)
Show Figures

Figure 1

25 pages, 1005 KB  
Article
The Hidden Thirst of AI: A Framework for Estimating Direct, Indirect, and Scarcity-Adjusted Freshwater Consumption per LLM Query
by Bhanu Sharma, Amit Tiwari, Rashanjot Kaur, Kathleen Marshall Park and Eugene Pinsky
Green 2026, 1(2), 8; https://doi.org/10.3390/green1020008 - 1 Sep 2026
Viewed by 123
Abstract
Large-scale artificial intelligence systems increasingly disclose energy and carbon metrics, but their freshwater costs remain less consistently measured. This paper introduces the Water Cost of Intelligence (WCI), a per-query metric combining direct water consumed for on-site data-center cooling with indirect water consumed during [...] Read more.
Large-scale artificial intelligence systems increasingly disclose energy and carbon metrics, but their freshwater costs remain less consistently measured. This paper introduces the Water Cost of Intelligence (WCI), a per-query metric combining direct water consumed for on-site data-center cooling with indirect water consumed during electricity generation, weighted by local scarcity using Aqueduct 4.0 Baseline Water Stress (BWS) scores. We first reconstruct Google’s disclosed Gemini direct-water figure from Google’s own reported parameters, an internal consistency check on the implementation rather than an independent validation. Expanding the accounting boundary to include electricity-generation water raises the estimate for a median large language model (LLM) prompt by 179% under a uniform national water-intensity value. Parameterizing that intensity by the regional generation mix instead changes the estimate substantially and reverses the regional ordering, depending on whether hydroelectric reservoir evaporation is allocated to generation: the same grid is the least water-intensive of those studied under one convention and the most water-intensive under the other. A region cannot be characterized as water-efficient in terms of electricity without first establishing that convention. Direct-only reporting can be internally accurate yet boundary-incomplete, and regional scarcity can change the interpretation of identical physical water use by an order of magnitude. Full article
Show Figures

Figure 1

26 pages, 3442 KB  
Article
Diagnostic Accuracy of Multimodal Large Language Models for Four-Class Benchmark of Oral Autoimmune Blistering Diseases: A Multicenter Paired Study
by Asmaa Abou-Bakr, Salma M Saad, Nevine H. Kheir El Din, Abdullah Bin Nabhan, Amal Bajonaid, Asma Saleh Almeslet and Fatma E. A. Hassanein
Diagnostics 2026, 16(17), 2761; https://doi.org/10.3390/diagnostics16172761 - 28 Aug 2026
Viewed by 134
Abstract
Background/Objectives: To compare the diagnostic performance of Claude Opus 4.7 and Gemini Pro 3 for the differential diagnosis of oral autoimmune blistering diseases (AIBDs) and evaluate their diagnostic reasoning, confidence, and calibration. Materials and Methods: This retrospective multicenter paired diagnostic accuracy study included [...] Read more.
Background/Objectives: To compare the diagnostic performance of Claude Opus 4.7 and Gemini Pro 3 for the differential diagnosis of oral autoimmune blistering diseases (AIBDs) and evaluate their diagnostic reasoning, confidence, and calibration. Materials and Methods: This retrospective multicenter paired diagnostic accuracy study included 200 clinicopathologically confirmed AIBD cases (50 each of pemphigus vulgaris, mucous membrane pemphigoid, bullous pemphigoid, and linear IgA bullous dermatosis). Each case was independently assessed by both models using identical standardized clinical information and clinical photographic inputs. The task required forced-choice classification among the four predefined diseases. Histopathological and direct immunofluorescence findings were used exclusively to establish the clinicopathological reference diagnosis and were not provided to the AI models. The reference diagnosis was established by clinicopathological correlation. The primary outcome was diagnostic accuracy. Secondary outcomes included disease-specific diagnostic performance, Cohen’s κ, ROC analysis, calibration, confidence, reasoning quality, management recommendations, and error patterns. Pre-consensus inter-rater reliability of the two human assessors was also evaluated using Cohen’s κ for binary outcomes and weighted Cohen’s κ for the ordinal reasoning-quality score. Results: Claude achieved significantly higher diagnostic accuracy than Gemini (92.0% vs. 86.0%, p = 0.012), stronger agreement with the reference standard (κ = 0.893 vs. 0.813), and superior discrimination (macro-AUC 0.998 vs. 0.965). Claude demonstrated higher key diagnostic-feature identification (92.0% vs. 86.0%; p = 0.012) and higher clinical-reasoning scores (61.0% vs. 40.0% of responses rated good; Wilcoxon p < 0.001; r = 0.47), whereas management recommendations did not differ significantly (100.0% vs. 98.0%; p = 0.125). Calibration results were metric-dependent: Claude had a lower one-vs-rest Brier score (0.0468 vs. 0.0558), whereas Gemini had a lower expected calibration error (0.083 vs. 0.251). For both models, the predominant error was misclassification of linear IgA bullous dermatosis as mucous membrane pemphigoid. Conclusions: Both multimodal LLMs showed high performance in this controlled four-class benchmark, with Claude Opus 4.7 outperforming Gemini Pro 3 in overall accuracy and reasoning quality. However, these findings do not establish autonomous diagnostic capability, clinical effectiveness, or safety. The LABD–MMP misclassification and metric-dependent calibration highlight important limitations. The models should therefore be regarded as investigational adjunctive decision-support tools requiring clinician oversight and diagnostic verification. Prospective external and human-in-the-loop validation is required before clinical implementation. Full article
(This article belongs to the Special Issue Application of Artificial Intelligence to Oral Diseases)
Show Figures

Figure 1

22 pages, 1390 KB  
Article
Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models
by Ibrahim Güler, Armin Kraus, Gerrit Grieb and Henrik Stelling
Bioengineering 2026, 13(9), 1000; https://doi.org/10.3390/bioengineering13091000 - 28 Aug 2026
Viewed by 290
Abstract
Background: Burn extent guides triage, transfer and fluid resuscitation, yet its clinical estimation is imprecise and observer-dependent. Multimodal large language models (MLLMs) process clinical photographs without task-specific training, but their error has rarely been separated into systematic and random components or their performance [...] Read more.
Background: Burn extent guides triage, transfer and fluid resuscitation, yet its clinical estimation is imprecise and observer-dependent. Multimodal large language models (MLLMs) process clinical photographs without task-specific training, but their error has rarely been separated into systematic and random components or their performance across skin tones characterized. Methods: Three state-of-the-art MLLMs (Gemini 3.1 Pro, GPT-5.6 Sol, and Fable 5) each assessed 153 burn photographs five times under an identical prompt. The tasks were as follows: burned proportion of the imaged field, against an expert-guided pixel-wise segmentation (tolerance ± 10 percentage points, pp); burned percentage of total body surface area (TBSA), against physician consensus (±2 pp); and binary Fitzpatrick skin tone (FST; light I–III versus dark IV–VI). The first of these was the primary endpoint. Results: The primary endpoint was in the range of 32.5–70.2%, with TBSA at 69.5–77.5%. All models compressed the estimation range (slopes 0.58–0.76, intercepts +10.0 to +22.2 pp); one multiplicative constant per model brought errors differing more than twofold into a 1.4 pp range. Across repeated queries, the median within-image range was 5.0–25.0 pp; averaging the five answers reduced error by only 0.24–2.22 pp. FST accuracy was 83.8–91.9% against a majority-class baseline of 81.0%. Conclusions: Averaging repeated answers removes only the smaller, random component; the larger, systematic one persists and requires calibration against reference data before clinical use can be considered. Full article
Show Figures

Graphical abstract

16 pages, 32393 KB  
Communication
Under the Sea: Detection of Explosives Underwater Using Raman Spectroscopy
by Dominika Łucja Sobczuk, Karol Zalewski, Mateusz Szala, Grzegorz Siedlewicz and Vasile Sorin Balan
Molecules 2026, 31(17), 2994; https://doi.org/10.3390/molecules31172994 - 27 Aug 2026
Viewed by 223
Abstract
Secondary explosives such as 1,3,5-trinitro-1,3,5-triazacyclohexane (RDX) and 2,4,6-trinitrotoluene (TNT) were commonly used in the production of many types of munitions during World War II and still continue to be used, for example, 1,3,5,7-tetranitro-1,3,5,7-tetrazocane (HMX) and pentaerythritol tetranitrate (PETN). After the war, numerous remaining [...] Read more.
Secondary explosives such as 1,3,5-trinitro-1,3,5-triazacyclohexane (RDX) and 2,4,6-trinitrotoluene (TNT) were commonly used in the production of many types of munitions during World War II and still continue to be used, for example, 1,3,5,7-tetranitro-1,3,5,7-tetrazocane (HMX) and pentaerythritol tetranitrate (PETN). After the war, numerous remaining munitions were disposed of by dumping them into the Baltic Sea. The same was true for the Black Sea, but due to a lack of documentation about these dumpsites, the awareness of hazards and knowledge about the scale of the problem have significantly increased in the last 30 years. After decades spent in salty water, the shells of said projectiles corroded, and currently, they pose a risk of environmental contamination and explosion. Because of this, new underwater detection methods that are reliable, quick, remote and, at the same time, insensitive to the external interference of sea water are needed. The detection of explosives used in projectiles can be performed using field Raman spectrometers. Raman spectroscopy is an analytical method that is non-invasive: it allows one to perform analysis without taking samples out of the projectile. Field detectors are water-resistant due to the frequent need for decontamination. Therefore, in this study, a ThermoScientific Gemini analyzer (Waltham, MA, USA) was used. The purpose of the study was to determine how the Raman spectra of pressed explosives are influenced by two main factors: the distance between the Raman probe and the sample and the type of medium that stands between the Raman probe and the sample. For the purpose of this study, four explosives: TNT, RDX, PETN, and HMX were tested through various layers of distilled water, tap water, silica dispersion, Black Sea water and, finally, Baltic Sea water. Full article
(This article belongs to the Special Issue Structure and Properties of Energetic Materials)
Show Figures

Figure 1

27 pages, 3230 KB  
Article
Lateral and Axial Camera Motions Reveal Two Fixed-Error Components in Active Infrared Stereo
by Yunfeng Ji, Yongze Lu, Shanshan Wang and Wei Li
Mathematics 2026, 14(17), 3074; https://doi.org/10.3390/math14173074 - 26 Aug 2026
Viewed by 148
Abstract
Multi-frame depth fusion reduces only errors that vary across observations; a static active infrared stereo camera can retain a repeatable spatial residual field. We characterized a Gemini 336L over 1.5–3.5 m and performed a Gemini 335 check. Residuals after plane fitting comprised averageable [...] Read more.
Multi-frame depth fusion reduces only errors that vary across observations; a static active infrared stereo camera can retain a repeatable spatial residual field. We characterized a Gemini 336L over 1.5–3.5 m and performed a Gemini 335 check. Residuals after plane fitting comprised averageable temporal and fixed terms. After registration to checkerboard coordinates, sampled fixed fields admitted an exact finite-view mean-centered partition into view-invariant and view-varying constituents. Lateral displacement decorrelated the varying constituent at 1.5–2.5 m; fixed-pose optical substitution supported its speckle association. The invariant constituent persisted across same-range lateral views and showed fractional-disparity periodicity and axial phase-dependent correlation, supporting a disparity-locking association. At 3.5 m, eight lateral views retained 57% of fixed-field mean square, whereas four axial views reduced residual root-mean-square (RMS) from 13.87 to 7.48 mm (ratio 0.540). Median total and fixed-only single-frame disparity-equivalent residuals were 0.068 and 0.056 px. For static acquisition, total error is persistent bias squared plus temporal variance, not a conditional covariance bound. Results motivate combining lateral speckle and axial disparity-phase diversity. Axial fusion was verified for registered planar residual fields at one range endpoint, not across ranges or in full-scene three-dimensional (3D) reconstruction. Full article
Show Figures

Graphical abstract

17 pages, 834 KB  
Article
Ranking Soil Quality Indicators Using the SMART Criteria, AHP Method and Chatbots
by Alexandre Marco da Silva and Jakub Kostecki
Sustainability 2026, 18(17), 8713; https://doi.org/10.3390/su18178713 - 25 Aug 2026
Viewed by 276
Abstract
Soil is a complex environmental component characterized by numerous physical, chemical, and biological attributes that determine its capacity to provide ecosystem services. While composite soil health indices are traditionally derived from static literature reviews or costly expert panels, this study evaluates the potential [...] Read more.
Soil is a complex environmental component characterized by numerous physical, chemical, and biological attributes that determine its capacity to provide ecosystem services. While composite soil health indices are traditionally derived from static literature reviews or costly expert panels, this study evaluates the potential of general-purpose large language models as a rapid, low-cost aid for AI-driven environmental decision support. We systematically evaluate the conceptual reliability and mathematical consistency of four distinct Large Language Models (LLMs): ChatGPT, Gemini, Claude, and Copilot, in executing an Analytic Hierarchy Process (AHP) matrix integrated with the SMART criteria. The evaluation prioritized seven indicators: soil organic matter (SOM), earthworm presence, aggregate stability, electrical conductivity, available nutrients, pH, and water infiltration capacity. All platforms generated mathematically consistent AHP matrices, with consistency ratios below accepted thresholds. Across ten independent runs per platform, soil organic matter received the highest mean global weight (~0.22) and earthworms the lowest (~0.08); inter-platform agreement was statistically significant (Kendall’s W = 0.75, p = 0.007), yet single-run variability reached ~50%, showing that repeated prompting is required. This approach offers a practical framework for rapid, preliminary indicator screening under resource-constrained conditions, while revealing model-specific divergences and possible training-data biases that warrant further investigation. However, our findings suggest AI-generated AHP outputs should be treated as preliminary decision-support results and should be verified through cross-platform comparison, repeated prompting, expert review, and site-specific validation. Full article
(This article belongs to the Special Issue Land Degradation, Soil Conservation and Reclamation)
Show Figures

Figure 1

31 pages, 8052 KB  
Article
Large Language Models in Peer Review: Decision Alignment, Review-Text Characteristics, and Human–AI Aggregation at ICLR 2025
by Zhihe Yang, Xiaoyu Zhou, Hongsa Wang, Yuxin Jiang and Xinjie Zhang
Publications 2026, 14(3), 55; https://doi.org/10.3390/publications14030055 - 25 Aug 2026
Viewed by 377
Abstract
Large language models (LLMs) are increasingly employed in scholarly peer review, yet their suitability as autonomous evaluators remains uncertain. Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet [...] Read more.
Large language models (LLMs) are increasingly employed in scholarly peer review, yet their suitability as autonomous evaluators remains uncertain. Using the ICLR 2025 review process, this study compares 2401 human reviews with 7203 reviews produced in separate, context-isolated API runs using Claude Sonnet 4.5, GPT-5.2 Thinking, and Gemini 3 Pro Preview across decision agreement, review-text characteristics, inter-model consistency, and human–AI aggregation. Raw LLM scores showed systematic leniency and score compression. A 0.1-point grid search identified thresholds of 6.2, 6.3, and 6.7 for Claude, GPT, and Gemini, respectively; repeated stratified cross-validation reproduced these thresholds. When applied without retuning to a stratified balanced sample of 300 ICLR 2024 papers, decision-agreement accuracy was 0.927, 0.913, and 0.930. Independent human coding of research type and primary field showed substantial pre-adjudication agreement (Cohen’s kappa = 0.774 and 0.714), and the recalculated analyses did not support H3. Review-text indicators showed similar structural completeness across sources but uneven critical-section length; these descriptive measures do not establish review quality. Human-containing aggregation rules showed higher agreement with conference decisions than corresponding AI-only rules, without establishing independent review quality or causal complementarity. A textual-overlap check found very low exact eight-gram containment, and manual inspection of the highest-similarity 1% found shared manuscript content or domain terminology rather than reviewer-specific evaluative language; possible prior exposure nevertheless could not be excluded. Full article
(This article belongs to the Special Issue Large Language Models Across the Lifecycle of Scholarly Publishing)
Show Figures

Figure 1

17 pages, 3818 KB  
Article
Flotation Behavior of Lepidolite Using Conventional Amine and Gemini Collectors
by Sergio Vladimir Mejia, Andrés Ramirez-Madrid and Leopoldo Gutierrez
Minerals 2026, 16(9), 864; https://doi.org/10.3390/min16090864 - 25 Aug 2026
Viewed by 226
Abstract
Lepidolite flotation was investigated using Armeen C, a conventional primary amine, and a cationic Gemini surfactant (hexanediyl-α,ω-bis(dimethyldodecylammonium bromide), 12-6-12) (HBDB), as collectors. Microflotation tests, electrophoretic mobility measurements, adsorption isotherms, and dynamic foamability experiments were conducted to compare their performance and interaction mechanisms. HBDB [...] Read more.
Lepidolite flotation was investigated using Armeen C, a conventional primary amine, and a cationic Gemini surfactant (hexanediyl-α,ω-bis(dimethyldodecylammonium bromide), 12-6-12) (HBDB), as collectors. Microflotation tests, electrophoretic mobility measurements, adsorption isotherms, and dynamic foamability experiments were conducted to compare their performance and interaction mechanisms. HBDB produced high lepidolite recoveries at relatively low concentrations of 25–75 mg/L, particularly between pH 5 and 9, whereas Armeen C required higher dosages of 100–300 mg/L to achieve comparable recoveries over a broader pH range. Electrophoretic mobility and adsorption results confirmed the interaction of both collectors with the lepidolite surface, although Armeen C showed stronger charge reversal and higher adsorption density. In contrast, HBDB exhibited substantially stronger foamability, with DFI values ~7.2–7.8 higher than those of Armeen C. These results indicate that HBDB can enhance lepidolite flotation through efficient surface hydrophobization and strong froth stabilization. Full article
(This article belongs to the Collection Flotation Theory and Technology)
Show Figures

Graphical abstract

33 pages, 2364 KB  
Article
The Paradigm Shift in Education: Stakeholder Perceptions of Generative AI in Teaching–Learning Dynamics
by Stoica Silviu-Ionel and Vasciuc Sandulescu Cristina Gabriela
Sustainability 2026, 18(17), 8678; https://doi.org/10.3390/su18178678 - 24 Aug 2026
Viewed by 374
Abstract
The study explores the paradigm shift in education brought about by the introduction of generative artificial intelligence (AI) tools, focusing on educational stakeholders’ self-reported perceptions rather than observed changes in teaching or learning outcomes. We consider stakeholders’ views on AI-based technologies within the [...] Read more.
The study explores the paradigm shift in education brought about by the introduction of generative artificial intelligence (AI) tools, focusing on educational stakeholders’ self-reported perceptions rather than observed changes in teaching or learning outcomes. We consider stakeholders’ views on AI-based technologies within the teaching–learning process. The current study uses a cross-sectional empirical survey design with a sample of N = 917 respondents, including teachers, students, administrators, and management. It examines the use of advanced AI technologies such as ChatGPT, Gemini, DeepSeek, and Grok, and stakeholders’ perceived connection between digital skills and classroom performance, student motivation, and critical thinking. We also discuss the ethical dilemmas and structural challenges that accompany this digital change. Inferential statistics, such as One-Way ANOVA and the Pearson Chi-Square test, show statistically significant differences in perceptions and regulatory expectations across organizational responsibilities. The findings contribute to understanding how advanced digitalization is perceived to reshape traditional academic roles, offering practical insights for creating effective, responsible, and sustainable teaching practices. Full article
(This article belongs to the Section Sustainable Education and Approaches)
Show Figures

Figure 1

14 pages, 906 KB  
Article
Evaluation and Comparison of Large Language Model Responses to Frequently Asked Questions Regarding Patellofemoral Pain Syndrome: A Quality and Readability Assessment Study
by Oktay Polat, Berk Koncalıoğlu, Mert Gündoğdu and Emrecan Akgün
Healthcare 2026, 14(17), 2694; https://doi.org/10.3390/healthcare14172694 - 24 Aug 2026
Viewed by 178
Abstract
Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain uncertain. This [...] Read more.
Background: Patellofemoral pain syndrome (PFPS) is a common cause of anterior knee pain, and patients increasingly use large language models (LLMs) to obtain general medical information. However, the quality, reliability, and readability of LLM-generated responses to patient-oriented questions regarding PFPS remain uncertain. This study aimed to compare responses generated by four widely used LLMs. Methods: Seventeen frequently asked questions regarding PFPS were identified through Google searches and adapted into lay language. The questions were submitted to OpenAI GPT-5, Google Gemini 2.5 Pro, xAI Grok 4, and DeepSeek-V3.2-Exp using a standardized patient scenario. A total of 68 question-specific responses were independently evaluated by four orthopedic surgeons using the DISCERN instrument. Inter-rater reliability was assessed using the intraclass correlation coefficient. Readability was evaluated using the Gunning Fog Index, Coleman–Liau Index, and Flesch Reading Ease Score. Between-model comparisons were performed using the Friedman test, followed by Bonferroni-adjusted pairwise analyses. Results: The omnibus Friedman test showed a significant between-model difference in DISCERN scores (p = 0.002). In Bonferroni-adjusted pairwise comparisons, GPT-5 had lower DISCERN scores than Gemini 2.5 Pro (adjusted p = 0.006), Grok 4 (adjusted p = 0.021), and DeepSeek-V3.2-Exp (adjusted p = 0.036), whereas no significant differences were observed among the other three models. However, the absolute differences were small, and the between-model difference was not significant in the sensitivity analysis using the median evaluator score (p = 0.381). Inter-rater agreement was moderate for GPT-5 and DeepSeek-V3.2-Exp but poor for Gemini 2.5 Pro and Grok 4. Readability differed significantly among the models across all three indices. DeepSeek-V3.2-Exp generally showed more favorable numerical readability values, whereas Grok 4 tended to produce more difficult text; however, no model was consistently superior across all readability measures. The median Gunning Fog and Coleman–Liau scores for all four models exceeded the commonly recommended sixth- to eighth-grade reading level for patient education. Conclusions: The evaluated LLMs showed small and method-dependent differences in DISCERN-based information quality and variable differences in readability. Their responses may supplement general patient education, but the findings should not be interpreted as evidence of factual accuracy, clinical safety, or suitability for individualized decision-making. LLM-generated information should be critically reviewed and should not replace assessment by a qualified healthcare professional. Full article
(This article belongs to the Special Issue AI & ICT in Healthcare)
Show Figures

Figure 1

Back to TopTop