Next Article in Journal
A Novel Lightweight Framework for Real-Time Pavement Crack Segmentation Based on Knowledge Distillation
Previous Article in Journal
Coordinated Cyber–Physical Attack Strategies in Power Systems Considering Defense Resource Allocation and Emergency Dispatch Responses
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
This is an early access version, the complete PDF, HTML, and XML versions will be available soon.
Review

Large Language Models: From Internal Architecture and Distributed Training Optimisation to Adaptation Strategies

Institute of Robotics and Cybernetics, Faculty of Electrical Engineering and Information Technology, Slovak University of Technology in Bratislava, Ilkovičova 3, 841 04 Bratislava, Slovakia
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(15), 7849; https://doi.org/10.3390/app16157849
Submission received: 30 June 2026 / Revised: 2 August 2026 / Accepted: 4 August 2026 / Published: 6 August 2026

Abstract

This paper presents a unified technical survey of Large Language Models (LLMs), connecting three layers of the modelling pipeline that existing surveys address in isolation: internal architecture, distributed training optimisation, and downstream adaptation. Its organising principle is the dependency between these layers—how a choice at one constrains what remains feasible at the next. The survey examines fundamental mechanisms (tokenisation, scaled dot-product attention, activation functions, and normalisation, including RMSNorm and pre- versus post-normalisation placement), then the engineering of training at scale: data, tensor, and pipeline parallelism, hybrid schemes, mixed-precision training with BF16 and FP8, ZeRO-Offload memory management, activation checkpointing, and compute-optimal scaling laws together with the conditions under which they fail. The adaptation section covers supervised and instruction fine-tuning, a comparison of parameter-efficient methods (LoRA, QLoRA, adapters, prefix and prompt tuning), Reinforcement Learning from Human Feedback with its reward-hacking failure mode, alternatives including DPO, KTO and Constitutional AI, Retrieval-Augmented Generation beyond the basic pipeline, and decoding strategies. Practical configuration guidance is given for 7B, 70B and trillion-parameter regimes. Dedicated treatments of Mixture-of-Experts architectures, long-context modelling, and hardware-aware co-design close the survey, with open challenges classified by origin and severity.
Keywords: large language models; distributed training; parallelism; fine-tuning; retrieval-augmented generation large language models; distributed training; parallelism; fine-tuning; retrieval-augmented generation

Share and Cite

MDPI and ACS Style

Lukáč, M.; Duchoň, F.; Ivan, J.; Dekan, M.; Zelenay, E.; Kocúr, M.; Trepáčová, I.; Jurov, T.; Pšenka, R. Large Language Models: From Internal Architecture and Distributed Training Optimisation to Adaptation Strategies. Appl. Sci. 2026, 16, 7849. https://doi.org/10.3390/app16157849

AMA Style

Lukáč M, Duchoň F, Ivan J, Dekan M, Zelenay E, Kocúr M, Trepáčová I, Jurov T, Pšenka R. Large Language Models: From Internal Architecture and Distributed Training Optimisation to Adaptation Strategies. Applied Sciences. 2026; 16(15):7849. https://doi.org/10.3390/app16157849

Chicago/Turabian Style

Lukáč, Martin, František Duchoň, Jakub Ivan, Martin Dekan, Eduard Zelenay, Maroš Kocúr, Izabela Trepáčová, Tomáš Jurov, and Radoslav Pšenka. 2026. "Large Language Models: From Internal Architecture and Distributed Training Optimisation to Adaptation Strategies" Applied Sciences 16, no. 15: 7849. https://doi.org/10.3390/app16157849

APA Style

Lukáč, M., Duchoň, F., Ivan, J., Dekan, M., Zelenay, E., Kocúr, M., Trepáčová, I., Jurov, T., & Pšenka, R. (2026). Large Language Models: From Internal Architecture and Distributed Training Optimisation to Adaptation Strategies. Applied Sciences, 16(15), 7849. https://doi.org/10.3390/app16157849

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop