Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (28)

Search Parameters:
Keywords = in-memory processing framework

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
42 pages, 3136 KB  
Article
New DTMOS-Based Charge- and Flux-Controlled Memtranstor Emulators
by Predrag Petrović
Appl. Sci. 2026, 16(15), 7551; https://doi.org/10.3390/app16157551 - 29 Jul 2026
Viewed by 153
Abstract
Memtranstors are emerging higher-order memory elements that establish a state-dependent constitutive relationship between electric charge and magnetic flux, making them attractive for adaptive analog electronics, neuromorphic computing, nonlinear dynamical systems, and memory-enabled signal processing applications. However, existing memtranstor emulators predominantly rely on operational [...] Read more.
Memtranstors are emerging higher-order memory elements that establish a state-dependent constitutive relationship between electric charge and magnetic flux, making them attractive for adaptive analog electronics, neuromorphic computing, nonlinear dynamical systems, and memory-enabled signal processing applications. However, existing memtranstor emulators predominantly rely on operational amplifiers, analog multipliers, current conveyors, or behavioral models, leading to increased circuit complexity and limited suitability for monolithic CMOS integration. This paper presents a unified transistor-level dynamic-threshold MOS (DTMOS) framework for realizing both charge-controlled and flux-controlled memtranstor emulators. The proposed architectures synthesize direct and inverse memtranstances through capacitive state integration, state-dependent DTMOS conductance modulation, and current-domain affine processing, thereby eliminating the need for composite active building blocks. Closed-form analytical expressions are derived for both constitutive relations and explicitly related to transistor-level parameters, bias conditions, and state-storage elements. The theoretical framework is further supported by comprehensive analyses of channel-length modulation, finite output resistance, device mismatch, DTMOS body-effect deviations, leakage mechanisms, pseudo-resistor non-idealities, parasitic capacitances, and small-signal stability. Cadence Virtuoso simulations performed in a 180 nm triple-well CMOS technology validate the analytical predictions and demonstrate the characteristic butterfly shaped pinched hysteresis loops of both emulators. The proposed circuits operate from a single 0.8 V supply while dissipating approximately 22 μW and 36 μW for the charge-controlled and flux-controlled realizations, respectively, and exhibit electronic tunability, together with robustness against process and temperature variations. Representative implementations of reconfigurable frequency-selective circuits and a memtranstor-based envelope detector further demonstrate the practical applicability of the proposed architectures. To the best of the author’s knowledge, this work presents the first unified transistor-level DTMOS constitutive synthesis framework for realizing both direct and inverse memtranstive behavior, providing a scalable foundation for future adaptive mixed-signal integrated circuits, programmable analog memory systems, neuromorphic hardware, and in-memory computing platforms. Full article
Show Figures

Figure 1

11 pages, 7634 KB  
Article
CMOS-Compatible AlScN Memristor on Silicon Exhibiting Short-Term Memory for Reservoir Computing
by Woohyun Park, Hyojeong Chae, Maria Rasheed and Sungjun Kim
Biomimetics 2026, 11(8), 519; https://doi.org/10.3390/biomimetics11080519 - 23 Jul 2026
Viewed by 317
Abstract
We report a CMOS-compatible ferroelectric memristor based on a TiN/AlScN/n+ Si metal ferroelectric semiconductor (MFS) structure, fabricated entirely via low-temperature sputtering processes. The ultrathin AlScN film exhibits robust ferroelectricity with a high remanent polarization (2Pr ≈ 80.91 μC/cm2) and [...] Read more.
We report a CMOS-compatible ferroelectric memristor based on a TiN/AlScN/n+ Si metal ferroelectric semiconductor (MFS) structure, fabricated entirely via low-temperature sputtering processes. The ultrathin AlScN film exhibits robust ferroelectricity with a high remanent polarization (2Pr ≈ 80.91 μC/cm2) and excellent endurance over 105 cycles, while maintaining uniform switching across cells. Notably, the use of a heavily doped silicon bottom electrode enables full compatibility with conventional back-end-of-line (BEOL) CMOS processes and facilitates integration with silicon-based circuits. Beyond stable memory performance, the device demonstrates volatile short-term memory (STM) behavior originating from depolarization field-induced polarization relaxation, which is essential for neuromorphic dynamics. Leveraging this STM feature, the device was implemented as a physical reservoir in a reservoir computing (RC) framework, achieving 97.64% classification accuracy on the MNIST dataset using temporally coded inputs. These results highlight the potential of AlScN-based ferroelectric memristors as dynamic CMOS-compatible building blocks for in-memory and neuromorphic computing. Full article
(This article belongs to the Section Bioinspired Sensorics, Information Processing and Control)
Show Figures

Figure 1

56 pages, 6689 KB  
Review
AI-on-Chip Systems: A Cross-Layer Review of Architectures, Interconnects, Design Automation, and Embedded Intelligence
by Mohamed M. Morsy
Electronics 2026, 15(12), 2645; https://doi.org/10.3390/electronics15122645 - 15 Jun 2026
Viewed by 2349
Abstract
The rapid growth of artificial intelligence (AI) workloads is reshaping semiconductor design across architecture, interconnect, memory hierarchy, packaging, timing, and design automation. Rather than converging on a single hardware solution, the field is expanding into a heterogeneous ecosystem that includes data-center graphics processing [...] Read more.
The rapid growth of artificial intelligence (AI) workloads is reshaping semiconductor design across architecture, interconnect, memory hierarchy, packaging, timing, and design automation. Rather than converging on a single hardware solution, the field is expanding into a heterogeneous ecosystem that includes data-center graphics processing units (GPUs), edge neural processing units (NPUs), and application-specific integrated circuits (ASICs), field-programmable gate array (FPGA)-based and hybrid AI system-on-chip (SoC) platforms, chiplet-enabled systems, and emerging beyond-conventional-silicon approaches such as photonic, neuromorphic, and analog in-memory processors. This paper presents a comprehensive review of AI-on-chip systems from a cross-layer perspective. It examines AI chip architectures and hardware platforms, network-on-chip (NoC) designs for AI communication patterns, and algorithm–hardware co-design methods for model acceleration, including compression, quantization, and sparsity-aware optimization. It also reviews clocking, synchronization, and clock-domain-crossing (CDC) challenges in large heterogeneous systems and chiplets, as well as manufacturing, advanced packaging, and reliability issues, including two-and-a-half-dimensional (2.5D) and three-dimensional (3D) integration, thermal and mechanical constraints, assembly quality, and long-term yield considerations. In parallel, the paper surveys the growing role of AI in chip design itself, covering machine-learning-assisted analysis, Bayesian and reinforcement-learning-based optimization, and the emerging use of large language models (LLMs) and AI agents for register-transfer level (RTL) generation, design-space exploration, and autonomous electronic design automation (EDA) workflows. Finally, it discusses beyond-silicon AI chip directions and the broader economic and industry context shaping cloud, on-premises, and edge deployment. By integrating these topics into a unified framework, this review highlights the key technological drivers, system-level tradeoffs, and future research directions that will define next-generation scalable, reliable, and energy-efficient AI-on-chip systems. Full article
(This article belongs to the Topic AI Agents: Progress, Architecture, and Applications)
Show Figures

Figure 1

45 pages, 4664 KB  
Review
Bridging Architectures, Mapping, and Learning for DNN Acceleration with Processing-in-Memory and In-Memory Computing Systems
by Syeda Munazza Marium and Song Chen
Microelectronics 2026, 2(2), 10; https://doi.org/10.3390/microelectronics2020010 - 10 Jun 2026
Viewed by 739
Abstract
Processing-in-memory and in-memory computing (PIM/IMC) are increasingly explored to mitigate the von Neumann data-movement bottleneck that limits deep neural network (DNN) performance and energy efficiency. Progress, however, remains fragmented across device substrates, architectural prototypes, mapping and scheduling methods, compiler toolchains, and benchmarking practices, [...] Read more.
Processing-in-memory and in-memory computing (PIM/IMC) are increasingly explored to mitigate the von Neumann data-movement bottleneck that limits deep neural network (DNN) performance and energy efficiency. Progress, however, remains fragmented across device substrates, architectural prototypes, mapping and scheduling methods, compiler toolchains, and benchmarking practices, making results hard to compare and slowing deployment. This survey synthesizes developments from 2019–2025 along four coupled axes: (i) memory substrates and architectural design, (ii) mapping, partitioning, and scheduling, including learning- and graph-based strategies, (iii) compilers and end-to-end deployment flows, and (iv) benchmarking datasets, metrics, and reporting norms. Drawing on over twenty representative platforms spanning static random-access memory (SRAM) and dynamic random-access memory (DRAM), emerging non-volatile, capacitive, and photonic substrates, we clarify the trade-offs separating analog/charge-domain IMC from digital SRAM/DRAM-centric PIM, including reported peaks up to 600 TOPS/W and 1.5 TOPS/mm2. We organize mapping frameworks into a unified reference taxonomy, identify recurrent evaluation pitfalls that undermine reproducibility, and highlight persistent gaps in training support, robustness under non-idealities, and coverage of large-scale GNN workloads. Finally, we outline a five-phase roadmap from benchmark standardization to industrial validation toward compiler-integrated, GNN-informed PIM/IMC systems validated on production-scale workloads. Full article
Show Figures

Figure 1

22 pages, 1133 KB  
Article
Elastic IoT Ontologies for Industry 4.0: Methodological Approach and Hybrid Architecture
by Larysa S. Globa and Serhii M. Ushakov
Future Internet 2026, 18(5), 264; https://doi.org/10.3390/fi18050264 - 17 May 2026
Viewed by 413
Abstract
Industry 4.0 requires IoT ontologies that are interoperable, scalable, and adaptive in non-stationary industrial environments. This study combines methodological ontology optimization with a hybrid elastic framework for dynamic semantic updates and feedback-driven refinement. The methodological component systematizes literature and industrial practices to identify [...] Read more.
Industry 4.0 requires IoT ontologies that are interoperable, scalable, and adaptive in non-stationary industrial environments. This study combines methodological ontology optimization with a hybrid elastic framework for dynamic semantic updates and feedback-driven refinement. The methodological component systematizes literature and industrial practices to identify structural gaps and derive practical requirements. The engineering component integrates truth-table-based data structuring, vector–matrix automata for real-time classification and clustering, and in-memory event processing for low-latency operation. Experimental evaluation across no-drift, abrupt-drift, gradual-drift, and cyclic-drift scenarios shows a trade-off between semantic proximity and operational robustness: the rule-based approach reaches lower semantic distance in drift regimes, while the hybrid approach delivers higher stability and fewer false alarms in cyclic dynamics. All tested configurations preserve sub-millisecond processing latency, supporting edge/fog deployment. The results indicate that combining methodological analysis with elastic architecture is a practical pathway from static to adaptive IoT ontologies and a relevant step toward human-centric Industry 5.0 systems. Full article
(This article belongs to the Special Issue Cyber-Physical Systems in Industrial Communication Systems)
Show Figures

Figure 1

33 pages, 745 KB  
Article
XAI-Driven Malware Detection from Memory Artifacts: An Alert-Driven AI Framework with TabNet and Ensemble Classification
by Aristeidis Mystakidis, Grigorios Kalogiannnis, Nikolaos Vakakis, Nikolaos Altanis, Konstantina Milousi, Iason Somarakis, Gabriela Mihalachi, Mariana S. Mazi, Dimitris Sotos, Antonis Voulgaridis, Christos Tjortjis, Konstantinos Votis and Dimitrios Tzovaras
AI 2026, 7(2), 66; https://doi.org/10.3390/ai7020066 - 10 Feb 2026
Cited by 1 | Viewed by 2592
Abstract
Modern malware presents significant challenges to traditional detection methods, often leveraging fileless techniques, in-memory execution, and process injection to evade antivirus and signature-based systems. To address these challenges, alert-driven memory forensics has emerged as a critical capability for uncovering stealthy, persistent, and zero-day [...] Read more.
Modern malware presents significant challenges to traditional detection methods, often leveraging fileless techniques, in-memory execution, and process injection to evade antivirus and signature-based systems. To address these challenges, alert-driven memory forensics has emerged as a critical capability for uncovering stealthy, persistent, and zero-day threats. This study presents a two-stage host-based malware detection framework, that integrates memory forensics, explainable machine learning, and ensemble classification, designed as a post-alert asynchronous SOC workflow balancing forensic depth and operational efficiency. Utilizing the MemMal-D2024 dataset—comprising rich memory forensic artifacts from Windows systems infected with malware samples whose creation metadata spans 2006–2021—the system performs malware detection, using features extracted from volatile memory. In the first stage, an Attentive and Interpretable Learning for structured Tabular data (TabNet) model is used for binary classification (benign vs. malware), leveraging its sequential attention mechanism and built-in explainability. In the second stage, a Voting Classifier ensemble, composed of Light Gradient Boosting Machine (LGBM), eXtreme Gradient Boosting (XGB), and Histogram Gradient Boosting (HGB) models, is used to identify the specific malware family (Trojan, Ransomware, Spyware). To reduce memory dump extraction and analysis time without compromising detection performance, only a curated subset of 24 memory features—operationally selected to reduce acquisition/extraction time and validated via redundancy inspection, model explainability (SHAP/TabNet), and training data correlation analysis —was used during training and runtime, identifying the best trade-off between memory analysis and detection accuracy. The pipeline, which is triggered from host-based Wazuh Security Information and Event Management (SIEM) alerts, achieved 99.97% accuracy in binary detection and 70.17% multiclass accuracy, resulting in an overall performance of 87.02%, including both global and local explainability, ensuring operational transparency and forensic interpretability. This approach provides an efficient and interpretable detection solution used in combination with conventional security tools as an extra layer of defense suitable for modern threat landscapes. Full article
Show Figures

Figure 1

23 pages, 360 KB  
Article
In-Memory Shellcode Runner Detection in Internet of Things (IoT) Networks: A Lightweight Behavioral and Semantic Analysis Framework
by Jean Rosemond Dora, Ladislav Hluchý and Michal Staňo
Sensors 2025, 25(17), 5425; https://doi.org/10.3390/s25175425 - 2 Sep 2025
Cited by 3 | Viewed by 2324
Abstract
The widespread expansion of Internet of Things devices has ushered in an era of unprecedented connectivity. However, it has simultaneously exposed these resource-constrained systems to novel and advanced cyber threats. Among the most impressive and complex attacks are those leveraging in-memory shellcode runners [...] Read more.
The widespread expansion of Internet of Things devices has ushered in an era of unprecedented connectivity. However, it has simultaneously exposed these resource-constrained systems to novel and advanced cyber threats. Among the most impressive and complex attacks are those leveraging in-memory shellcode runners (malware), which perform malicious payloads directly in memory, circumventing conventional disk-based detection security mechanisms. This paper presents a comprehensive framework, both academic and technical, for detecting in-memory shellcode runners, particularly tailored to the unique characteristics of these networks. We analyze and review the limitations of existing security parameters in this area, highlight the different challenges posed by those constraints, and propose a multi-layered approach that combines entropy-based anomaly scoring, lightweight behavioral monitoring, and novel Graph Neural Network methods for System Call Semantic Graph Analysis. Our proposal focuses on runtime analysis of process memory, system call patterns (e.g., Syscall ID, Process ID, Hooking, Win32 application programming interface), and network behavior to identify the subtle indicators of compromise that portray in-memory attacks, even in the absence of conventional file-system artifacts. Through meticulous empirical evaluation against simulated and real-world Internet of Things attacks (red team engagements, penetration testing), we demonstrate the efficiency and a few challenges of our approach, providing a crucial step towards enhancing the security posture of these critical environments. Full article
(This article belongs to the Special Issue Internet of Things Cybersecurity)
Show Figures

Figure 1

25 pages, 1615 KB  
Article
Efficient Parallel Processing of Big Data on Supercomputers for Industrial IoT Environments
by Isam Mashhour Al Jawarneh, Lorenzo Rosa, Riccardo Venanzi, Luca Foschini and Paolo Bellavista
Electronics 2025, 14(13), 2626; https://doi.org/10.3390/electronics14132626 - 29 Jun 2025
Cited by 5 | Viewed by 2705
Abstract
The integration of distributed big data analytics into modern industrial environments has become increasingly critical, particularly with the rise of data-intensive applications and the need for real-time processing at the edge. While High-Performance Computing (HPC) systems offer robust petabyte-scale capabilities for efficient big [...] Read more.
The integration of distributed big data analytics into modern industrial environments has become increasingly critical, particularly with the rise of data-intensive applications and the need for real-time processing at the edge. While High-Performance Computing (HPC) systems offer robust petabyte-scale capabilities for efficient big data analytics, the performance of big data frameworks, especially on ARM-based HPC systems, remains underexplored. This paper presents an extensive experimental study on deploying Apache Spark 3.0.2, the de facto standard in-memory processing system, on an ARM-based HPC system. This study conducts a comprehensive performance evaluation of Apache Spark through representative big data workloads, including K-means clustering, to assess the effects of latency variations, such as those induced by network delays, memory bottlenecks, or computational overheads, on application performance in industrial IoT and edge computing environments. Our findings contribute to an understanding of how big data frameworks like Apache Spark can be effectively deployed and optimized on ARM-based HPC systems, particularly when leveraging vectorized instruction sets such as SVE, contributing to the broader goal of enhancing the integration of cloud–edge computing paradigms in modern industrial environments. We also discuss potential improvements and strategies for leveraging ARM-based architectures to support scalable, efficient, and real-time data processing in Industry 4.0 and beyond. Full article
Show Figures

Figure 1

43 pages, 2159 KB  
Systematic Review
A Systematic Review and Classification of HPC-Related Emerging Computing Technologies
by Ehsan Arianyan, Niloofar Gholipour, Davood Maleki, Neda Ghorbani, Abdolah Sepahvand and Pejman Goudarzi
Electronics 2025, 14(12), 2476; https://doi.org/10.3390/electronics14122476 - 18 Jun 2025
Cited by 2 | Viewed by 3950
Abstract
In recent decades, access to powerful computational resources has brought about a major transformation in science, with supercomputers drawing significant attention from academia, industry, and governments. Among these resources, high-performance computing (HPC) has emerged as one of the most critical processing infrastructures, providing [...] Read more.
In recent decades, access to powerful computational resources has brought about a major transformation in science, with supercomputers drawing significant attention from academia, industry, and governments. Among these resources, high-performance computing (HPC) has emerged as one of the most critical processing infrastructures, providing a suitable platform for evaluating and implementing novel technologies. In this context, the development of emerging computing technologies has opened up new horizons in information processing and the delivery of computing services. In this regard, this paper systematically reviews and classifies emerging HPC-related computing technologies, including quantum computing, nanocomputing, in-memory architectures, neuromorphic systems, serverless paradigms, adiabatic technology, and biological solutions. Within the scope of this research, 142 studies which were mostly published between 2018 and 2025 are analyzed, and relevant hardware solutions, domain-specific programming languages, frameworks, development tools, and simulation platforms are examined. The primary objective of this study is to identify the software and hardware dimensions of these technologies and analyze their roles in improving the performance, scalability, and efficiency of HPC systems. To this end, in addition to a literature review, statistical analysis methods are employed to assess the practical applicability and impact of these technologies across various domains, including scientific simulation, artificial intelligence, big data analytics, and cloud computing. The findings of this study indicate that emerging HPC-related computing technologies can serve as complements or alternatives to classical computing architectures, driving substantial transformations in the design, implementation, and operation of high-performance computing infrastructures. This article concludes by identifying existing challenges and future research directions in this rapidly evolving field. Full article
Show Figures

Figure 1

24 pages, 2877 KB  
Article
Memory-Efficient Batching for Time Series Transformer Training: A Systematic Evaluation
by Phanwadee Sinthong, Nam Nguyen, Vijay Ekambaram, Arindam Jati, Jayant Kalagnanam and Peeravit Koad
Algorithms 2025, 18(6), 350; https://doi.org/10.3390/a18060350 - 5 Jun 2025
Viewed by 4290
Abstract
Transformer-based time series models are being increasingly employed for time series data analysis. However, their training remains memory intensive, especially with high-dimensional data and extended look-back windows, while model-level memory optimizations are well studied, the batch formation process remains an underexplored factor to [...] Read more.
Transformer-based time series models are being increasingly employed for time series data analysis. However, their training remains memory intensive, especially with high-dimensional data and extended look-back windows, while model-level memory optimizations are well studied, the batch formation process remains an underexplored factor to performance inefficiency. This paper introduces a memory-efficient batching framework based on view-based sliding windows operating directly on GPU-resident tensors. This approach eliminates redundant data materialization caused by tensor stacking and reduces data transfer volumes without modifying model architectures. We present two variants of our solution: (1) per-batch optimization for datasets exceeding GPU memory, and (2) dataset-wise optimization for in-memory workloads. We evaluate our proposed batching framework systematically using peak GPU memory consumption and epoch runtime as efficiency metrics across varying batch sizes, sequence lengths, feature dimensions, and model architectures. Results show consistent memory savings, averaging 90% and runtime improvements of up to 33% across multiple transformer-based models (Informer, Autoformer, Transformer, and PatchTST) and a linear baseline (DLinear) without compromising model accuracy. We extensively validate our method using synthetic and standard real-world benchmarks, demonstrating accuracy preservation and practical scalability in distributed GPU environments. The proposed method highlights batch formation process as a critical component for improving training efficiency. Full article
(This article belongs to the Section Parallel and Distributed Algorithms)
Show Figures

Figure 1

22 pages, 3570 KB  
Article
High-Performance Computing and Parallel Algorithms for Urban Water Demand Forecasting
by Georgios Myllis, Alkiviadis Tsimpiris, Stamatios Aggelopoulos and Vasiliki G. Vrana
Algorithms 2025, 18(4), 182; https://doi.org/10.3390/a18040182 - 22 Mar 2025
Cited by 7 | Viewed by 3394
Abstract
This paper explores the application of parallel algorithms and high-performance computing (HPC) in the processing and forecasting of large-scale water demand data. Building upon prior work, which identified the need for more robust and scalable forecasting models, this study integrates parallel computing frameworks [...] Read more.
This paper explores the application of parallel algorithms and high-performance computing (HPC) in the processing and forecasting of large-scale water demand data. Building upon prior work, which identified the need for more robust and scalable forecasting models, this study integrates parallel computing frameworks such as Apache Spark for distributed data processing, Message Passing Interface (MPI) for fine-grained parallel execution, and CUDA-enabled GPUs for deep learning acceleration. These advancements significantly improve model training and deployment speed, enabling near-real-time data processing. Apache Spark’s in-memory computing and distributed data handling optimize data preprocessing and model execution, while MPI provides enhanced control over custom parallel algorithms, ensuring high performance in complex simulations. By leveraging these techniques, urban water utilities can implement scalable, efficient, and reliable forecasting solutions critical for sustainable water resource management in increasingly complex environments. Additionally, expanding these models to larger datasets and diverse regional contexts will be essential for validating their robustness and applicability in different urban settings. Addressing these challenges will help bridge the gap between theoretical advancements and practical implementation, ensuring that HPC-driven forecasting models provide actionable insights for real-world water management decision-making. Full article
Show Figures

Figure 1

25 pages, 1936 KB  
Article
A Scalable Framework for Sensor Data Ingestion and Real-Time Processing in Cloud Manufacturing
by Massimo Pacella, Antonio Papa, Gabriele Papadia and Emiliano Fedeli
Algorithms 2025, 18(1), 22; https://doi.org/10.3390/a18010022 - 4 Jan 2025
Cited by 19 | Viewed by 7141
Abstract
Cloud Manufacturing enables the integration of geographically distributed manufacturing resources through advanced Cloud Computing and IoT technologies. This paradigm promotes the development of scalable and adaptable production systems. However, existing frameworks face challenges related to scalability, resource orchestration, and data security, particularly in [...] Read more.
Cloud Manufacturing enables the integration of geographically distributed manufacturing resources through advanced Cloud Computing and IoT technologies. This paradigm promotes the development of scalable and adaptable production systems. However, existing frameworks face challenges related to scalability, resource orchestration, and data security, particularly in rapidly evolving decentralized manufacturing settings. This study presents a novel nine-layer architecture designed specifically to address these issues. Central to this framework is the use of Apache Kafka for robust, high-throughput data ingestion, and Apache Spark Streaming to enhance real-time data processing. This framework is underpinned by a microservice-based architecture that ensures a high scalability and reduced latency. Experimental validation using sensor data from the UCI Machine Learning Repository demonstrated substantial improvements in processing efficiency and throughput compared with conventional frameworks. Key components, such as RabbitMQ, contribute to low-latency performance, whereas Kafka ensures data durability and supports real-time application. Additionally, the in-memory data processing of Spark Streaming enables rapid and dynamic data analysis, yielding actionable insights. The experimental results highlight the potential of the framework to enhance operational efficiency, resource utilization, and data security, offering a resilient solution suited to the demands of modern industrial applications. This study underscores the contribution of the framework to advancing Cloud Manufacturing by providing detailed insights into its performance, scalability, and applicability to contemporary manufacturing ecosystems. Full article
Show Figures

Figure 1

20 pages, 740 KB  
Article
A Variation-Aware Binary Neural Network Framework for Process Resilient In-Memory Computations
by Minh-Son Le, Thi-Nhan Pham, Thanh-Dat Nguyen and Ik-Joon Chang
Electronics 2024, 13(19), 3847; https://doi.org/10.3390/electronics13193847 - 28 Sep 2024
Cited by 2 | Viewed by 2620
Abstract
Binary neural networks (BNNs) that use 1-bit weights and activations have garnered interest as extreme quantization provides low power dissipation. By implementing BNNs as computation-in-memory (CIM), which computes multiplication and accumulations on memory arrays in an analog fashion, namely, analog CIM, we can [...] Read more.
Binary neural networks (BNNs) that use 1-bit weights and activations have garnered interest as extreme quantization provides low power dissipation. By implementing BNNs as computation-in-memory (CIM), which computes multiplication and accumulations on memory arrays in an analog fashion, namely, analog CIM, we can further improve the energy efficiency to process neural networks. However, analog CIMs are susceptible to process variation, which refers to the variability in manufacturing that causes fluctuations in the electrical properties of transistors, resulting in significant degradation in BNN accuracy. Our Monte Carlo simulations demonstrate that in an SRAM-based analog CIM implementing the VGG-9 BNN model, the classification accuracy on the CIFAR-10 image dataset is degraded to below 50% under process variations in a 28 nm FD-SOI technology. To overcome this problem, we present a variation-aware BNN framework. The proposed framework is developed for SRAM-based BNN CIMs since SRAM is most widely used as on-chip memory; however, it is easily extensible to BNN CIMs based on other memories. Our extensive experimental results demonstrate that under process variation of 28 nm FD-SOI, with an SRAM array size of 128×128, our framework significantly enhances classification accuracies on both the MNIST hand-written digit dataset and the CIFAR-10 image dataset. Specifically, for the CONVNET BNN model on MNIST, accuracy improves from 60.24% to 92.33%, while for the VGG-9 BNN model on CIFAR-10, accuracy increases from 45.23% to 78.22%. Full article
(This article belongs to the Special Issue Research on Key Technologies for Hardware Acceleration)
Show Figures

Figure 1

44 pages, 8494 KB  
Review
Survey of Deep Learning Accelerators for Edge and Emerging Computing
by Shahanur Alam, Chris Yakopcic, Qing Wu, Mark Barnell, Simon Khan and Tarek M. Taha
Electronics 2024, 13(15), 2988; https://doi.org/10.3390/electronics13152988 - 29 Jul 2024
Cited by 37 | Viewed by 23580
Abstract
The unprecedented progress in artificial intelligence (AI), particularly in deep learning algorithms with ubiquitous internet connected smart devices, has created a high demand for AI computing on the edge devices. This review studied commercially available edge processors, and the processors that are still [...] Read more.
The unprecedented progress in artificial intelligence (AI), particularly in deep learning algorithms with ubiquitous internet connected smart devices, has created a high demand for AI computing on the edge devices. This review studied commercially available edge processors, and the processors that are still in industrial research stages. We categorized state-of-the-art edge processors based on the underlying architecture, such as dataflow, neuromorphic, and processing in-memory (PIM) architecture. The processors are analyzed based on their performance, chip area, energy efficiency, and application domains. The supported programming frameworks, model compression, data precision, and the CMOS fabrication process technology are discussed. Currently, most commercial edge processors utilize dataflow architectures. However, emerging non-von Neumann computing architectures have attracted the attention of the industry in recent years. Neuromorphic processors are highly efficient for performing computation with fewer synaptic operations, and several neuromorphic processors offer online training for secured and personalized AI applications. This review found that the PIM processors show significant energy efficiency and consume less power compared to dataflow and neuromorphic processors. A future direction of the industry could be to implement state-of-the-art deep learning algorithms in emerging non-von Neumann computing paradigms for low-power computing on edge devices. Full article
(This article belongs to the Special Issue AI for Edge Computing)
Show Figures

Figure 1

20 pages, 1280 KB  
Article
A Novel Algorithm for Multi-Criteria Ontology Merging through Iterative Update of RDF Graph
by Mohammed Suleiman Mohammed Rudwan and Jean Vincent Fonou-Dombeu
Big Data Cogn. Comput. 2024, 8(3), 19; https://doi.org/10.3390/bdcc8030019 - 21 Feb 2024
Cited by 2 | Viewed by 3512
Abstract
Ontology merging is an important task in ontology engineering to date. However, despite the efforts devoted to ontology merging, the incorporation of relevant features of ontologies such as axioms, individuals and annotations in the output ontologies remains challenging. Consequently, existing ontology-merging solutions produce [...] Read more.
Ontology merging is an important task in ontology engineering to date. However, despite the efforts devoted to ontology merging, the incorporation of relevant features of ontologies such as axioms, individuals and annotations in the output ontologies remains challenging. Consequently, existing ontology-merging solutions produce new ontologies that do not include all the relevant semantic features from the candidate ontologies. To address these limitations, this paper proposes a novel algorithm for multi-criteria ontology merging that automatically builds a new ontology from candidate ontologies by iteratively updating an RDF graph in the memory. The proposed algorithm leverages state-of-the-art Natural Language Processing tools as well as a Machine Learning-based framework to assess the similarities and merge various criteria into the resulting output ontology. The key contribution of the proposed algorithm lies in its ability to merge relevant features from the candidate ontologies to build a more accurate, integrated and cohesive output ontology. The proposed algorithm is tested with five ontologies of different computing domains and evaluated in terms of its asymptotic behavior, quality and computational performance. The experimental results indicate that the proposed algorithm produces output ontologies that meet the integrity, accuracy and cohesion quality criteria better than related studies. This performance demonstrates the effectiveness and superior capabilities of the proposed algorithm. Furthermore, the proposed algorithm enables iterative in-memory update and building of the RDF graph of the resulting output ontology, which enhances the processing speed and improves the computational efficiency, making it an ideal solution for big data applications. Full article
Show Figures

Figure 1

Back to TopTop