Large Language Models and Their Limitations

A Special Issue of Big Data and Cognitive Computing (ISSN 2504-2289) belonging to the section "Large Language Models and Embodied Intelligence".

Deadline for manuscript submissions: 15 February 2027 | Viewed by 2686

Editor


E-Mail Website
Guest Editor
Department of Computer Science and Electrical Engineering, Florida International University, Miami, FL 33199, USA
Interests: natural language processing; machine learning; text mining; semantic analysis; emotion detection

Special Issue Information

Dear Colleagues,

Large Language Models (LLMs) have transformed the field of artificial intelligence, enabling applications ranging from natural language understanding to content generation. Despite their impressive capabilities, these models have notable deficiencies, including bias propagation, hallucinations, ethical challenges, and limitations in reasoning or domain-specific knowledge. This Special Issue aims to explore both the advancements and the shortcomings of LLMs, providing a comprehensive overview of current research and practical applications. Contributions may include original research articles, reviews, communications, and concept papers that address theoretical developments, evaluation methodologies, practical implementations, and strategies for mitigating model limitations. By bringing together researchers and practitioners, this issue will offer insights into improving model reliability, transparency, and ethical deployment. We encourage submissions that critically assess existing models, propose novel architectures, or explore interdisciplinary approaches to overcoming the inherent challenges of LLMs. The goal is to inform the broader AI community and provide a reference for future research in large-scale language modeling.

(1) Introduction, including scientific background and highlighting the importance of this research area.

Large Language Models (LLMs) have emerged as a transformative technology in artificial intelligence, driven by advances in deep learning, transformer architectures, and the availability of large-scale datasets. These models have achieved remarkable success in a wide range of applications, including natural language understanding, text generation, question answering, machine translation, sentiment analysis, and decision support systems. Their rapid integration into academic research, industry products, healthcare, education, and public services demonstrates their significant impact on modern data-driven and cognitive computing systems.

Despite these achievements, LLMs exhibit critical deficiencies that limit their reliability and responsible deployment. Common challenges include hallucinations, bias amplification, lack of transparency and explainability, limited reasoning and factual consistency, privacy and security risks, and high computational and environmental costs. These limitations raise fundamental scientific, ethical, and societal concerns, particularly in high-stakes domains. Addressing these issues is essential for advancing trustworthy AI, improving model robustness, and strengthening the theoretical foundations of large-scale language modeling. Consequently, research on the deficiencies of LLMs has become a crucial and timely topic within the broader AI and big data research community.

(2) Aim of the Special Issue and how the subject relates to the journal scope.

The aim of this Special Issue is to provide a dedicated platform for exploring the limitations, risks, and deficiencies of Large Language Models, while also highlighting emerging methods and frameworks to address these challenges. This Special Issue seeks to bring together researchers and practitioners to critically examine model behavior, evaluation strategies, mitigation techniques, and real-world implications of LLM deployment.

This topic strongly aligns with the scope of Big Data and Cognitive Computing (BDCC), which focuses on data-intensive intelligent systems, cognitive computing, scalable machine learning, and responsible AI technologies. LLMs fundamentally rely on big data and cognitive architectures, making BDCC an ideal venue for disseminating research that advances understanding of both their capabilities and shortcomings. Contributions to this Special Issue will support the development of more reliable, transparent, and ethically grounded language models and cognitive systems.

(3) Suggest themes.

In this Special Issue, original research articles and reviews are welcome. Research areas may include (but are not limited to) the following:

  • Bias, fairness, and ethical challenges in Large Language Models.
  • Hallucination detection, analysis, and mitigation strategies.
  • Explainability, interpretability, and transparency of LLMs.
  • Evaluation benchmarks and reliability assessment of language models.
  • Reasoning limitations and factual consistency in LLM-generated content.
  • Privacy, security, and data leakage risks in large-scale models.
  • Domain adaptation and low-resource challenges in LLMs.
  • Energy efficiency, scalability, and sustainability of LLM training and deployment.
  • Human–AI interaction, trust, and user perception of LLM-based systems.
  • Real-world applications and case studies highlighting model deficiencies.

We look forward to receiving your contributions. 

Dr. Samira Zad
Guest Editor

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Big Data and Cognitive Computing is an international peer-reviewed open access monthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 1800 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • large language models
  • natural language processing
  • big data analytics
  • cognitive computing
  • model bias
  • hallucination
  • explainable AI
  • ethical AI

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (5 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

24 pages, 1135 KB  
Article
Beyond Keyword Filters: Calibrated Monte-Carlo Risk Gating for Safe Multilingual Colorectal-Cancer LLM Dialogue
by Abdurrahim Kızılay and Kerem Gencer
Big Data Cogn. Comput. 2026, 10(9), 317; https://doi.org/10.3390/bdcc10090317 - 14 Sep 2026
Abstract
Large language models are increasingly consulted by cancer patients, and a single unsafe answer about chemotherapy dosing, opioid use or self-harm can cause real harm. This paper introduces a calibrated Monte-Carlo risk gate that treats colorectal-cancer dialogue safety as a selective-prediction problem, estimating [...] Read more.
Large language models are increasingly consulted by cancer patients, and a single unsafe answer about chemotherapy dosing, opioid use or self-harm can cause real harm. This paper introduces a calibrated Monte-Carlo risk gate that treats colorectal-cancer dialogue safety as a selective-prediction problem, estimating the risk of a user turn from a bootstrap ensemble over multilingual sentence representations, calibrating it with Platt scaling and deciding at a single threshold between an informative answer and referral to a clinician. Evaluated on 450 oncologist-approved prompts in English, Turkish and Spanish under a scenario-level split, the gate reaches a guardrail F1 of 0.961, blocks 98.2 percent of harmful prompts and refuses 10.8 percent of legitimate questions, while the keyword, regular-expression and fuzzy layers that dominate current practice reach at most 0.034 and fire on five of 450 prompts, a separation that holds at a corrected q of 0.0003 and survives Bonferroni correction. Ablation locates the mechanism, with the semantic representation carrying the discriminative signal, Platt scaling lowering the expected calibration error from 0.216 to 0.069, and the ensemble predicting its own errors at an AUROC of 0.856 and removing them entirely at 50 percent coverage. Calibrated selective prediction over semantic representations makes multilingual medical-dialogue safety measurable, tunable to an explicit operating point, and consistent across the three languages tested. Full article
(This article belongs to the Special Issue Large Language Models and Their Limitations)
23 pages, 4771 KB  
Article
Benchmarking LLM-Based Time-Series Foundation Models for Minute-Resolution Turning-Movement Traffic Dynamics on an Urban Arterial
by Deo Chimba, Therezia Matongo, Afia Yeboah and Rheecha Sharma
Big Data Cogn. Comput. 2026, 10(9), 306; https://doi.org/10.3390/bdcc10090306 - 7 Sep 2026
Viewed by 191
Abstract
Large language models (LLMs) and time-series foundation models have advanced traffic forecasting, but almost exclusively on freeway sensors at five-minute, aggregate-flow resolution; their behavior on minute-resolution, movement-level counts at signalized arterials is unknown. This study assembled a real corridor dataset—6449 gap-free one-minute turning-movement [...] Read more.
Large language models (LLMs) and time-series foundation models have advanced traffic forecasting, but almost exclusively on freeway sensors at five-minute, aggregate-flow resolution; their behavior on minute-resolution, movement-level counts at signalized arterials is unknown. This study assembled a real corridor dataset—6449 gap-free one-minute turning-movement records (255,364 counted vehicles) at 20 signalized intersections along 13 miles of Nolensville Pike, Nashville, Tennessee—and ran a controlled benchmark across three model families (classical, deep, and LLM/foundation models). We evaluate in-sample accuracy at 5/10/15 min horizons (RQ1), cross-date leave-intersection-out transfer to morning-only sites (RQ2). On the full-day sites, a per-series ARIMA attains the best short-horizon accuracy (MASE 0.87 at h = 5), while a zero-shot foundation model (Chronos) is the most stable across horizons; under leave-intersection-out transfer the ordering reverses—foundation models transfer best (MASE ≈ 0.61, improving to ≈0.56 with few-shot adaptation) whereas globally trained deep networks fail catastrophically (MASE > 4). All benchmark values reported here are measured on the corridor data; foundation-model families that we could not execute end-to-end in this environment are discussed qualitatively and are not included in the quantitative comparison. The dataset and protocol establish an honest, cost-aware basis for judging whether language and foundation models genuinely forecast an urban corridor at minute resolution. Full article
(This article belongs to the Special Issue Large Language Models and Their Limitations)
Show Figures

Figure 1

21 pages, 917 KB  
Article
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
by Ali Şenol, Garima Agrawal and Huan Liu
Big Data Cogn. Comput. 2026, 10(9), 300; https://doi.org/10.3390/bdcc10090300 - 3 Sep 2026
Viewed by 191
Abstract
Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness. This provides limited insight into how models reason, how reliably they behave under contextual variation, or how efficiently they reach conclusions. This paper proposes RQEval, a unified multi-dimensional [...] Read more.
Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness. This provides limited insight into how models reason, how reliably they behave under contextual variation, or how efficiently they reach conclusions. This paper proposes RQEval, a unified multi-dimensional framework for measuring LLM reasoning quality from a behavioral perspective. The framework operationalizes six theoretically grounded dimensions rooted in cognitive science: Correctness (CQ), Consistency (CS), Robustness (RS), Local Logical Coherence (LS), Efficiency (ES), and Stability (SS). It also introduces deployment-aware aggregation, enabling context-specific model selection beyond accuracy-based leaderboards. Applying RQEval to seven LLMs across four benchmarks reveals an outcome-level cluster (CQ, CS, RS, and ES) characterized by strong intercorrelations and a trace-level layer in which LS is empirically distinct from the outcome-level metrics, whereas SS retains moderate associations with several of them. Across the 28 model–dataset observations, LS showed no statistically significant correlation with any other dimension. Efficiency-weighted deployment scenarios, nevertheless, produced limited ranking inversions among models that were otherwise ranked consistently across weighting schemes. The resulting pipeline provides a foundation for diagnosing LLM reasoning behavior across deployment contexts, while highlighting domain-specific validation as an important direction for future work. Full article
(This article belongs to the Special Issue Large Language Models and Their Limitations)
Show Figures

Figure 1

30 pages, 1444 KB  
Article
Beyond Transcript Alignment: Diagnosing Paralinguistic Information Flow in Frozen Speech-to-LLM Adapters
by Nurgali Kadyrbek and Madina Mansurova
Big Data Cogn. Comput. 2026, 10(8), 251; https://doi.org/10.3390/bdcc10080251 - 30 Jul 2026
Viewed by 416
Abstract
Frozen speech-to-LLM systems train an adapter between a frozen audio encoder and a text LLM. We test whether adapters preserve sentence stress beyond transcripts and whether the LLM uses it. In a five-seed Qwen3-8B/WavLM baseline, transcript alignment lowers linear adapter Probe-K below the [...] Read more.
Frozen speech-to-LLM systems train an adapter between a frozen audio encoder and a text LLM. We test whether adapters preserve sentence stress beyond transcripts and whether the LLM uses it. In a five-seed Qwen3-8B/WavLM baseline, transcript alignment lowers linear adapter Probe-K below the text-only KT baseline (0.211 vs. 0.290). The R1.8 configuration improves MLP-2 Probe-K from 0.245 to 0.306 (paired p=0.012) and passes three controls, but because warmup and augmentation also change, the gain is not attributed to Lcf alone; its crossing of the linear KT floor is neither statistically established nor capacity-matched. Response Probe-G remains 0.512 in both cohorts; single-seed LoRA and styled-teacher pilots do not improve it. A post hoc trace through all 36 Qwen blocks finds persistent absolute stress decodability in speech-slot states, but no reliable R1.8-over-R0 advantage at any state and no speech-conditioned answer margin; an explicit text tag instead yields a +6.40-nat margin. Attention differences depend on slot-length normalization, and late left/system concentration is shared across modalities rather than speech-specific. Thus, the tested system remains an end-to-end negative: adapter decodability does not imply causal response use. Scope is limited to sentence stress, mostly synthetic voices, one encoder, and one LLM. Full article
(This article belongs to the Special Issue Large Language Models and Their Limitations)
Show Figures

Figure 1

23 pages, 677 KB  
Article
Large Language Models for Energy Market Analytics: An Exploratory Feasibility Study Across Geopolitical Monitoring, Commodity Summarisation, and Renewable Forecasting
by Alex Krempasky, Erik Kajati and Peter Papcun
Big Data Cogn. Comput. 2026, 10(6), 166; https://doi.org/10.3390/bdcc10060166 - 22 May 2026
Viewed by 1024
Abstract
Large Language Models (LLMs) offer opportunities for processing heterogeneous information streams relevant to energy-market decision-making, but their practical role in forecasting-oriented analytical workflows remains uncertain. This paper presents an exploratory feasibility study of LLM use across four energy-market tasks: geopolitical event monitoring for [...] Read more.
Large Language Models (LLMs) offer opportunities for processing heterogeneous information streams relevant to energy-market decision-making, but their practical role in forecasting-oriented analytical workflows remains uncertain. This paper presents an exploratory feasibility study of LLM use across four energy-market tasks: geopolitical event monitoring for Dutch Title Transfer Facility (TTF) market context using Global Database of Events, Language, and Tone (GDELT)-based data, structured summarisation of commodity-intelligence articles, prompt-engineered solar-power and grid-load forecasting for Austria, and a short-horizon exploratory TTF price-estimation case. The study is positioned as a pilot investigation and hybrid workflow blueprint rather than as a statistically conclusive forecasting benchmark. A four-layer reference architecture was devised, including structured market data, semi-structured news intelligence, web-scraping concepts, and implemented Twitter/X and GDELT monitoring layers. The empirical cases indicate that LLMs are most useful for text-heavy reasoning, event-context integration, source triage, and structured interpretation. In the 20-article summarisation corpus, Gemini 1.5 Pro achieved higher commodity-direction accuracy than GPT-4, while GPT-4 showed stronger output-format stability. In selected solar case checks, OpenAI models produced plausible generation curves close to the Fraunhofer ISE Energy Charts reference, while Energy Charts remained more accurate for aggregate load estimation in the available benchmark comparison. The two-day TTF experiment illustrated that LLMs can incorporate qualitative geopolitical context into short-horizon reasoning, but it did not establish reliable price-forecasting capability. The Twitter/X monitoring layer is retained as a documented negative pathway, showing the limitations of informal social-media scraping for reproducible market intelligence. Full article
(This article belongs to the Special Issue Large Language Models and Their Limitations)
Show Figures

Figure 1

Back to TopTop