Next Article in Journal
Library-Oriented Code Generation Method Combining Execution-Feedback Re-Retrieval and Repair Mechanisms
Previous Article in Journal
Research on Task Assignment Method Based on Multi-Strategy Improved Whale Migration Hybrid Genetic Algorithm
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Modular VHDL Framework for Deterministic FPGA Acceleration of MR Fingerprinting Reconstruction

by
Mattia Ricchi
1,2,
Fabrizio Alfonsi
2,
Marco Barbieri
3,
Alessandra Retico
4,
Alessandro Gabrielli
2,5,*,
Leonardo Brizi
5,† and
Claudia Testa
2,5,†
1
Department of Computer Sciences, University of Pisa, Largo Bruno Pontecorvo 3, 56127 Pisa, Italy
2
Division of Bologna, National Institute for Nuclear Physics (INFN), Viale Berti Pichat 6/2, 40127 Bologna, Italy
3
Department of Biomedical Engineering, University of Basel, Hegenheimermattweg 167B/C, 4123 Allschwil, Switzerland
4
Division of Pisa, National Institute for Nuclear Physics (INFN), Largo Bruno Pontecorvo 3, 56127 Pisa, Italy
5
Department of Physics and Astronomy, University of Bologna, Viale Berti Pichat 6/2, 40127 Bologna, Italy
*
Author to whom correspondence should be addressed.
†
These authors contributed equally to this work.
Appl. Sci. 2026, 16(19), 9935; https://doi.org/10.3390/app16199935 (registering DOI)
Submission received: 26 August 2026 / Revised: 1 October 2026 / Accepted: 3 October 2026 / Published: 8 October 2026

Abstract

Neural networks (NNs) for real-time medical image processing in resource-constrained environments benefit from hardware platforms that ensure deterministic timing, low latency, and fine-grained control over computational resources. This work introduces a fully custom and modular hardware framework for implementing fully connected neural networks (FCNs) on FPGAs, targeting online quantitative MR fingerprinting (MRF) reconstruction without relying on high-level synthesis tools. The proposed approach is based on a reusable base node that integrates numerical operations within a compact, parameterizable hardware unit. A single block of 16 such nodes is instantiated and reused in time multiplexing over a fixed schedule, so that the number of clock cycles of an inference is a property of the design rather than of the data. The framework is validated through a three-way comparison: the proposed custom accelerator, three hls4ml-generated designs, and the reference software implementation. The complete network was executed on an AMD/Xilinx Alveo U250 board over a 5000-sample dataset covering a broad range of T 1 and T 2 values: the accelerator agrees with the reference 8-bit quantized model on every one of the integer outputs, so that the residual error of the deployed system, a mean absolute percentage error (MAPE) of 2.37 % on T 1 and 11.07 % on T 2 with respect to the ground truth, is entirely that of the quantized network and not of the hardware. Timing closure at 83.3 MHz and post-implementation analysis indicate a 10 3 speed-up over CPU execution, and an estimated energy reduction of 10 4 , with limited FPGA resource usage and approximately 8.6 W power consumption. None of the three hls4ml configurations tested reached a working bitstream on the target device, while the proposed architecture completed place-and-route and was executed on hardware, supporting its use as a building block for online quantitative MRI reconstruction.

1. Introduction

Artificial neural networks (NNs) constitute one of the most powerful and flexible tools of modern technology [1]. Inspired by the structure and functioning of the human brain, NNs are composed of computational units called neurons, organized in layers [2], and connected to each other to learn representations of complex data. Thanks to their capability to also model nonlinear relations and to generate examples and synthetic data, NNs have been applied with great success in a large variety of fields [3], such as computer vision [4], natural language processing [5,6], medical diagnosis and medicine [7,8,9].
One of the greatest advantages of NNs is their architectural flexibility [1,3]: models such as fully connected NN (FCN), convolutional NN (CNN), or recurrent NN (RNN) can be adapted to a wide diversity of problems simply by changing the structure or parameters. Moreover, thanks to the availability of tools and advanced software libraries such as TensorFlow [10] and PyTorch [11], the development and the training of different NN models have become increasingly accessible.
On the other hand, NNs present a few disadvantages which may limit their use. First, they have high computational demand, requiring high computing resources, limiting their application in embedded and low-cost devices. NNs also typically rely on Graphics Processing Units (GPUs) or Tensor Processing Units (TPUs), which provide high computational throughput but can be costly and power-hungry. Moreover, clusters of GPUs can be expensive and are not always available, especially in real-time or portable applications. In addition, NNs based on CPU or GPU have a variable latency, which could compromise their use in applications such as medical diagnostics or robotic control.
In recent years, the adoption of NNs in embedded applications has registered a substantial growth, also due to the need for real-time inferences in contexts where latency, energy consumption, and portability are critical factors [12,13]. In recent years, supervised deep learning methods have been introduced to directly map temporal fingerprints to relaxation times maps by either fully connected networks [7,8] or convolutional neural networks [14] or other models. In particular, field-programmable gate arrays (FPGAs) [15] have established themselves as a particularly suitable hardware platform for the implementation of machine learning models due to their reconfigurable architecture, massive parallelism capacity, and flexibility in managing logic resources [16,17].
The implementation of an NN on FPGA can be done at different levels. The most common solution is based on High-Level Synthesis (HLS) tools, which allow software models, like TensorFlow [10], Keras, and PyTorch [11], to be converted into FPGA-synthesizable code, often in languages such as C/C++ or SystemC. The alternative route, adopted here, is the direct description of every functional block at the register-transfer level (RTL) in a hardware description language, which requires a substantially larger development effort but leaves scheduling, memory organization, and arithmetic precision under explicit designer control. The trade-off between these two routes, and the point at which it becomes binding on a real device, is the subject of this work.

1.1. Related Work

Among HLS-based toolflows, Vitis AI [18] and hls4ml [19] have gained particular attention in the scientific community, the former mainly for CNNs [20] and more complex models, the latter mostly in particle-physics applications [21] requiring rapid inference. hls4ml generates FPGA firmware starting from pre-trained models, offering a simplified workflow accessible even without in-depth hardware design skills [19]. Several studies have nevertheless reported that HLS flows introduce a non-negligible architectural overhead in terms of logic-resource usage and latency, together with limited support for non-standard frameworks and devices, insufficient hardware-performance profiling, and reduced flexibility in adapting to rapidly evolving ML architectures [22,23,24,25].Their high level of abstraction limits fine-grained control over hardware resources, memory management, and timing, which may affect performance and numerical reproducibility in latency-critical applications such as medical image reconstruction [8,22,26], and HLS-generated accelerators may be difficult to reuse across different network architectures and precision requirements.
These limitations are the reason why, in high-throughput acquisition systems requiring deterministic sub-microsecond response, such as the on-the-fly trigger systems developed for high-energy physics experiments [26,27], fully custom low-level firmware is still adopted despite the increased development effort. The present work applies the same design philosophy to neural-network acceleration for quantitative MRI based on MRF.

1.2. Contribution

This work proposes a fully customized and modular VHDL framework for FCN acceleration on FPGAs, aimed at enabling systematic and reusable low-level NN design rather than ad hoc, application-specific implementations. Every computational component is implemented at the register-transfer level without relying on vendor IP cores or high-level synthesis tools. The architecture is built around a reusable “base node” that integrates quantization-aware arithmetic, activation, and requantization operations into a compact, parameterizable unit. This design ensures fine-grained control over memory organization and dataflow, enabling deterministic latency and predictable resource utilization, while explicit integer-only operations guarantee numerical consistency with quantization-aware training, making the proposed approach particularly suitable for real-time applications requiring strict timing. The novelty of this work lies in a modular VHDL design methodology based on a reusable hardware node, which can be systematically instantiated to build FCNs of different sizes while ensuring predictable resource utilization and deterministic latency and not in the implementation of a specific FCN accelerator on FPGA per se, since FPGA-based neural network accelerators are already widely used. The methodology is validated through the implementation of an MRF reconstruction network and compared against hls4ml-generated designs under realistic implementation constraints. The main contributions of this paper are: (1) the definition of a reusable and parameterizable VHDL base-node abstraction integrating quantization-aware arithmetic, activation, and requantization operations within a single hardware unit; (2) a deterministic FCN construction method based on controlled replication and scheduling of 16-node processing blocks, allowing for predictable latency and resource usage; (3) the validation of the proposed architecture through the implementation of an MRF reconstruction network and a comparison against hls4ml designs, including synthesis, routing, and timing-closure constraints on the target FPGA board.
The properties targeted by the proposed architecture are also those required by trustworthy cyber-physical systems (CPSs) in a safety-relevant medical imaging workflow, and the design is therefore presented here as a candidate building block for such systems. In quantitative MRI, reconstruction modules may operate as part of a real-time acquisition–processing pipeline, where bounded latency and reproducible numerical behavior are essential. By implementing the FCN directly in VHDL, the proposed framework, integrating an on-chip parameter storage, avoids runtime variability typical of CPU/GPU execution and enables low and fixed-latency inference after timing closure. This control over timing, arithmetic precision, resource allocation, and hardware isolation supports deterministic operation, data integrity, and reliable real-time performance, which are key requirements for trustworthy and secure CPS in medical applications.

2. Materials and Methods

Before describing the individual components, it is helpful to provide an overview of the overall workflow, which is structured into five stages and shown in Figure 1. (i) The original FCN of Barbieri et al. [8] is adapted to the resource budget of the target device and retrained, then converted to a full-integer representation through quantization-aware training (QAT), producing the quantized reference model (Section 2.1). (ii) A reusable base node implementing the integer inference kernel of that model in hardware is written in VHDL (Section 2.2). (iii) A block of 16 base nodes, an on-chip parameter store, and a control module built on a fixed schedule table are assembled into the complete FCN (Section 2.3 and Section 2.4). (iv) The design is verified at four distinct levels, namely functional VHDL simulation, on-device execution of a block of 16 nodes, on-device execution of the complete network over the whole test set, and post-implementation timing and power analysis after place-and-route (Section 2.5). (v) The same trained model is finally implemented with hls4ml in three configurations and compared against the proposed design (Section 2.6).
The current status of firmware development for FPGA allows several approaches for the NN acceleration through hardware. A low-level HDL design approach, in which every firmware block is designed in VHDL without high-level synthesis support, has been selected. The VHDL code was written and managed using the AMD/Xilinx Vivado Design Suite, which was also used for synthesis, implementation, and place-and-route of the proposed architecture. The chosen FPGA board is the Alveo U250 which features 1.7M LUTs, 3.4M of FFs, 12k Digital Signal Processors (DSPs) and 2.6k Block RAM (BRAM). Prior to FPGA deployment, the FCN was implemented and evaluated on a workstation equipped with an Intel® Xeon® W7-2575X CPU, 64 GB of DDR5 ECC memory, and an NVIDIA® T400 GPU. The execution times obtained on this platform are used as a reference for comparison with the FPGA-based implementation.

2.1. Network Adaptation and Quantization

The structure of the original model, derived from Barbieri et al. [8], was modified to the one shown in Figure 2: 20 input neurons, hidden layers of 256, 128, 64, 32, and 16 neurons, and 2 linear outputs for T1 and T2 (ReLU in all other layers). The NN takes 20 inputs given by the first principal components of the k-SVD–truncated, denoised signal [7]. The modification consists of removing the two widest hidden layers of the original network, eliminating its largest weight matrices: the 1024 × 512 matrix alone has 524 , 288 parameters, yielding a >14× parameter reduction. Because removing >90% of parameters can degrade performance, the reduced topology was validated before hardware design. Both networks were trained from scratch on the same dataset of 2.5 × 10 8 simulated IR-FISP MRF signals spanning wide relaxation times, SNRs, and phases, with identical training settings, and reached comparable error.
Finally, since the target architecture operates exclusively on integers, the retrained model was converted to a full-integer representation through quantization-aware training (QAT) [29,30], carried out with the TensorFlow Model Optimization Toolkit and exported with the TensorFlow Lite converter. An 8-bit quantization was adopted for both parameters and activations, with 32-bit accumulators.
Table 1 separates the contribution of the two steps on the same 5000-sample test set. The architecture simplification costs 0.68 percentage points of global MAPE, whereas the transition to 8-bit arithmetic costs 1.20 points, almost twice as much. The accuracy of the deployed model is therefore limited by the numerical representation rather than by the reduced capacity of the network. This model, simplified, retrained, and quantized to 8-bit integer precision, is the one that the hardware architecture implements, and the one against which the numerical correctness of the FPGA design is verified.

2.2. Node Functioning and VHDL Implementation

The most straightforward method to represent the functioning of the human brain is to consider it as a complex network of neurons that receive inputs from other neurons and subsequently transmit information via membrane and action potentials. Based on these principles, the concept of “perceptron” [31], or artificial neuron, was introduced, which is a mathematical function designed to take a vector of numerical inputs and output their weighted linear combination. A bias is then added to the linear combination and a non-linear activation function is applied to obtain the final neuron activation. Thus, the general function of a single artificial neuron, also called a node, is as follows.
y = σ ∑ i = 1 N i n p u t s x i · w i + b
where x i represents the single input, weighted by the parameter w i ; b is instead the bias and σ is the activation function, which is an important function used to introduce non-linearity into the model.
Since Equation (1) represents the generic operations of each node within a network, it was implemented generically as an independent VHDL package. Thus, it can be used for all the node operations within the desired NN by simply adjusting the input data and configuring the parameters of the node accordingly.
FPGAs natively support integer arithmetic; floating-point operations are possible but cost more logic, latency, and power. Therefore, to meet resource and latency budgets, the node was implemented in VHDL using only integer-only operations. A feasible method of achieving an integer NN is through QAT [29,30] on the pre-trained network. Hence, the operations that have been coded in VHDL are intended to simulate the functionality of a QAT network, and are as follows:
y i n t = σ ∑ i = 1 N i n p u t s ( x i q − z p i n ) · w i q + b q
where the symbol q indicates that the variable was previously quantized and z p i n is the zero input point for effective quantization. In this work, an 8-bits quantization was adopted, and in VHDL code, as also represented in Figure 3, this translates as follows:
Applsci 16 09935 i001
Figure 3. Detailed schematic of the base node implementation used in the VHDL architecture. The quantized inputs x i q , weights ω i q , and bias b q are processed through the integer-only computational pipeline, which performs zero-point subtraction, multiplication, accumulation, and ReLU activation. The quantization-aware operations mirror those used during QAT [29], ensuring numerical consistency between software and hardware. The final activation y i n t is requantized to 8-bit precision before being propagated to the next stage.
Figure 3. Detailed schematic of the base node implementation used in the VHDL architecture. The quantized inputs x i q , weights ω i q , and bias b q are processed through the integer-only computational pipeline, which performs zero-point subtraction, multiplication, accumulation, and ReLU activation. The quantization-aware operations mirror those used during QAT [29], ensuring numerical consistency between software and hardware. The final activation y i n t is requantized to 8-bit precision before being propagated to the next stage.
Applsci 16 09935 g003
The result of the weighted sum in Equation (2) can easily reach 32-bit word representation. To be given as input to the following levels of the NN, the output number y i n t must be quantized to 8 bits. The quantization of the output in Equation (2) is realized in 6 steps and can be summarized in the following function:
y i n t q = round y i n t · M 0 + 2 n − 1 ≫ n + z p o u t , y i n t q ∈ − 128 , 127
Initially, the process multiplies by an integer scaling factor M 0 . Following that, rounding up to the nearest integer is achieved by adding 2 n − 1 , where n represents a shift factor. At this stage, the computation results are maintained in 64 bits to prevent overflow, and they are subsequently shifted right by the 2 n division. Next, the output zero point z p o u t gets added to the shifted value. In the final step, this number is restricted to fit within the int 8 range [ − 128 ; 127 ] and then converted to the appropriate signal type in VHDL.
In this work, the Rectified Linear Unit (ReLU) activation function σ = max ( 0 , x ) was selected for its simplicity within the base node package and is implemented as a straightforward if condition:
Applsci 16 09935 i002
The modularity of the VHDL block allows one to use the node as a generic entity in different NN contexts, simply modifying the input parameters such as weights, biases, zero-points, and scale factors.
To test the proper functioning of the implemented package for base node calculations, 16 nodes were completely mapped onto the FPGA board. Both hardware and software implementations used the same inputs, weights, and biases, and Python (3.11.10) and FPGA results were then compared.

2.3. Fully Connected Network Implementation

The adapted and quantized network described above provides the first practical motivation for designing an efficient hardware realization of dense layers on FPGA.
Having defined the simplest computational unit, i.e., the node, a complete NN can now be constructed consisting of nodes arranged in several layers. The very first is defined as the input layer, the final one is the output layer, and those in between are called hidden layers. In this case, each layer receives as input a vector that contains the outputs of the nodes in the previous layer. If every neuron in a hidden layer is connected to all the neurons in the previous layer, as shown in Figure 4, then such an arrangement is called an FCN.
To optimize the FPGA implementation and find a trade-off between parallelism and logical resources consumption, in the proposed FPGA architecture, the FCN has been implemented by instantiating a single block of 16 base nodes in hardware and iterating on it, in time multiplexing, for the whole network: the 498 neurons are partitioned into 32 groups of at most 16, so that the hardware cost is that of 16 neurons while the network behaves as if all 498 were present.
The specific value of 16 was selected on the basis of three quantitative considerations. First, the arithmetic cost of the block is set by the product of its width and of the largest fan-in it must sustain: with 16 nodes and a widest layer of 256 inputs, the multiply–accumulate array requires 16 × 256 DSP units, plus four per node for the 32-bit requantization multiplier, that is, the 4160 units, corresponding to 33.85 % of the device. Doubling the block width to 32 nodes would require 8320 DSPs (67.7% of the available resources), substantially increasing resource utilisation and reducing the headroom available for placement and routing. The 16-node configuration was therefore selected as a more conservative design point, providing sufficient parallelism while retaining a greater resource margin for timing closure. Reducing it to 8 would double the length of the schedule and hence the inference latency. Second, 16 divides exactly all the layer widths of the network (256, 128, 64, 32, 16), so that the partition never crosses a layer boundary and each layer is assigned a whole number of groups. Third, a block of this width keeps the fan-out of the parameter-selection logic within a range that place-and-route can absorb without congestion.
The control module coordinates the FCN execution and is responsible for:
  • Input loading: activations from the previous layer are gathered into a buffer array, ensuring synchronous data transfer at each clock cycle;
  • Weight and bias management: parameters are retrieved from the dedicated internal memory block, where they are preloaded and organized in hierarchical matrices tailored to each layer’s dimensions;
  • Pipeline control: the computation is managed by two counters, which regulate address generation, layer transitions, and the temporal sequencing of input/output storage;
  • Layer parameterization: quantization-related parameters such as scale factors, zero-points, and shift factors are updated dynamically at each layer, ensuring that the integer-only arithmetic remains consistent with the QAT;
  • Output collection: the computed activations are stored in another buffer array, either to be propagated to the next layer or to serve as the final network output.
Compared to fully automated HLS tools, this low-level VHDL code offers much greater control. Every step of the computation—from weight memory placement, through scheduling multiplications and additions, to rounding strategy in requantization—is explicitly defined. Unlike programmable devices such as GPUs and CPUs, FPGAs are internally synchronized with a clock signal that defines what happens step by step. For this reason, the system guarantees deterministic latency, precise use of resources, and the ability to fine-tune numerical precision adapted to application-level constraints. This kind of control is typically not present or limited in HLS-generated designs, where the optimization is left to rely on the heuristics of the compiler and the libraries.
The computation of a layer proceeds deterministically. At the beginning of each clock cycle, input activations are loaded into a buffer array. Each node calculates the weighted summation of the inputs, adds the corresponding bias, performs requantization steps, and applies ReLU activation. Outputs of the 16-node block are produced in parallel and saved into the corresponding positions in the output vector. Once the outputs of a layer have been completed, the system updates the quantization parameters and begins computation for the following layer without host intervention.
This design shows the advantages of a custom-designed architecture: it is deterministic, scalable, and modular. Within the limits imposed by the available FPGA resources, memory bandwidth, routing complexity, and timing-closure constraints, larger FCN architectures can be constructed by duplicating the 16-node blocks or by chaining multiple instances together, while maintaining the same deterministic execution model and explicit control over arithmetic precision. Besides, interoperability with on-board FPGA memory ensures that weights and biases are correctly distributed to the computational units, reducing Peripheral Component Interconnect Express (PCIe) transactions and improving throughput.

2.4. FPGA Internal Memory

Another key aspect in the FPGA implementation of NNs is related to memory management. Weights and biases need to be provided to the computation units at each clock cycle during inference, and using external PCIe transfers would lead to high latency and bandwidth usage. To overcome this issue, a custom memory block was created to store all the parameters locally inside the FPGA and to make them available efficiently during computation.
The module has two sub-modules dedicated to weight and bias storage, respectively. They are both instantiated in read-only mode for inference, because weights and biases are pre-loaded before execution. Data are represented as std_logic_vector first and then cast into signed integers to be compatible with arithmetic operations defined in the base_node module.
The design of the memory is guided by the following principles:
  • Hierarchical organization of weights: parameters are grouped according to the size of each layer. For example, the first layer requires 256 neurons with 20 inputs each, while subsequent layers progressively reduce in size down to 2 neurons with 16 inputs. This organization guarantees compact storage and predictable addressing.
  • Efficient bias handling: biases are preloaded into dedicated registers and synchronized with the computation pipeline, ensuring that each node receives its corresponding parameter in the correct cycle.
  • Parallel access: multiple weights can be made available simultaneously through the use of generative constructs, feeding the 16-node computation blocks without additional multiplexing overhead.
  • Explicit data conversion: the casting from std_logic_vector to signed guarantees consistency with the quantization scheme adopted, avoiding mismatches between memory representation and arithmetic logic.
  • Reduced host transfers: by preloading all parameters within the FPGA, communication via PCIe is minimized, improving both latency and energy efficiency.
The other part of the memory management strategy is the method implemented in the Top_module to operate sequentially. At every rising edge of the clock, the design iterates through the 16 nodes and updates both their biases and weights according to the current cycle counter: the weights are progressively loaded in blocks from the global weights matrix into local registers, with different time intervals corresponding to different portions of the layer. At the first rising edge of the clock, the registers are reset back to their initial values, and then later cycles fill in 20, 256, 128, 64, 32, then 16 entries, filling unused positions with zeros if necessary. This hierarchical loading plan follows the decreasing size of the layers and ensures that only the appropriate subset of the weights is active during any one cycle.
Biases are handled the same way: they are initialized at the beginning of each calculation, and while the counter is incremented, they are loaded in from the preloaded bias matrix. This design guarantees that every node sees the correct weights and bias in parallel with the processing pipeline.
The memory module, coupled with the Top module of the FCN, supports the network to be executed in autonomous mode completely within the FPGA, realizing the modularity of design of this work.

2.5. NN Performance Evaluation

To evaluate the performance of the implemented NN, a dataset consisting of 5000 samples was used as input for both the software-based and hardware-based network implementations. Each of these samples corresponds to a simulated MRF signal generated for a single pixel, representative of an MRF image. In other words, every data point in the set models the temporal signal evolution that would be measured from one pixel under specific tissue and acquisition conditions. Since the implemented model is an FCN, the reconstruction process is performed in a pixel-wise manner, meaning that each pixel is processed independently and no spatial information from neighboring pixels is used. Consequently, each input sample is representative of a generic image pixel, and the complete reconstruction of an image can be obtained by applying the same network to all pixels individually. This makes the chosen dataset representative of the entire T 1 and T 2 domain, despite being composed of independent samples.
The simulated MRF signals were generated following the methodology described in [8]. These works describe how to obtain realistic MRF time courses by incorporating appropriate tissue parameters (such as T 1 and T 2 relaxation times) and acquisition settings (such as flip angle patterns and repetition times). In particular, the simulated T 1 and T 2 values were sampled uniformly and independently over a wide physiological range ( T 1 : [0.0438–3.9990 s] and T 2 : [0.0017–2.9990 s]), covering relaxation times representative of different tissue types and anatomical regions, with the same forward model, the same IR-FISP pulse sequence, and the same augmentation and preprocessing chain used for the training data. This choice was made to increase the robustness of the trained FCN and avoid specialization to brain-only applications. The 5000 samples were generated independently of, and are disjoint from, the 2.5 × 10 8 signals used to train the network, and were never seen by the model during training or validation. Using this common set of 5000 simulated MRF signals—referred to as the 5000-sample dataset in the rest of the paper—for the ground truth, the floating-point model, the quantized software model, and the FPGA implementation enables a direct comparison with the tested hls4ml configurations.
The validation was carried out at four distinct levels:
(i) Functional simulation. A behavioral VHDL simulation of the complete FCN reproduces the logical behavior of the firmware exactly, including the processing sequence and the logical timing of operations, but it does not reproduce the physical propagation delays through the FPGA fabric. It is therefore used to establish functional correctness, and the cycle count of an inference and the latency it yields are expressed at the nominal simulation clock of 200 MHz .
(ii) Module-level on-board verification. A block of 16 base nodes was mapped onto the physical device and its outputs compared, for identical inputs, weights, and biases, against the software reference. This step verifies the arithmetic of the elementary unit but it is not sufficient to establish the correctness of the entire network.
(iii) Full-network on-board execution. The complete FCN was configured on the Alveo U250 and executed on the physical device over the entire 5000-sample test set. The input samples were streamed to the accelerator and the integer outputs were read back and compared, before dequantization, with those produced by the reference 8-bit quantized model executed in software by the TensorFlow Lite integer kernel. This is the level at which the expression “deployed on the target FPGA” is used in this work: the design was synthesized, placed, routed, loaded onto the device, and run on the complete test set.
(iv) Post-implementation analysis. After place-and-route, static timing analysis was used to determine the maximum achievable clock frequency and to verify timing closure, and vectorless power analysis to estimate power consumption. All latency and power values reported for the implemented design derive from these reports and are estimates produced by the design tools, not metered measurements.
Latency, resource utilization, numerical error, and power consumption were selected as evaluation metrics because they directly determine the feasibility of FPGA accelerators in real-time and embedded contexts, where timing determinism, hardware footprint, numerical fidelity and energy efficiency are equally critical. The Python/Keras execution time used as a baseline was measured on the workstation over the complete 5000-sample dataset rather than on a single inference, because software execution time varies from run to run due to scheduling, cache behavior, and runtime overhead.
The design was developed with the AMD/Xilinx Vivado Design Suite 2022.2, and the software reference with TensorFlow/Keras 2.15.0 and the TensorFlow Model Optimization Toolkit 0.8.0. Quantization is symmetric per tensor for the weights and asymmetric per tensor for the activations, with 8-bit parameters and activations and 32-bit accumulators, as described by Equation (3).

2.6. Three-Way Evaluation Pipeline

In order to provide a comparison with existing FPGA-oriented toolflows, the same FCN model was also implemented using the hls4ml framework [19]. The conversion started from the trained Keras model, leaving the default hls4ml configuration parameters unchanged. Network parameters and activations are automatically mapped to fixed-point representations with a precision of 8 bits, with wider accumulators for intermediate sums. In this way, the generated design reflects a standard and widely adopted baseline for NN implementation on FPGA, without introducing additional manual optimizations that could bias the comparison.
Three different hls4ml implementations were tested. In the first case (model 2.1), the entire network is implemented fully in parallel, where all nodes are instantiated simultaneously. The second (model 2.2) applies a strong serialization, so that only a very small fraction of the nodes is physically implemented and iterated over in time. The third (model 2.3) adopts an intermediate configuration chosen to mirror the proposed architecture, which instantiates 16 nodes of the full network and processes the remaining nodes sequentially. As illustrated in Figure 5, the proposed VHDL implementation is evaluated against the hls4ml implementations and the software-based model in terms of latency and resource usage.

3. Results

3.1. Functional Simulation and Numerical Accuracy

A dedicated simulation of the FCN was used to verify the correctness of the VHDL implementation. Figure 6 reports the waveform of a complete inference cycle. The shown signals include the input activations fed to the 16-node block, and the final quantized outputs produced at the end of the computation pipeline (T1_tb and T2_tb). The output values are shown in an integer 8-bit representation. To obtain the final T 1 and T 2 values a scaling to the float 32 representation is necessary. Intermediate control signals (e.g., layer counters and internal buffering lines) are omitted for clarity, as they do not affect the numerical validation.
From the simulation waveforms, it is also possible to clearly and easily measure the processing time of each signal. Each inference follows the same fixed sequence of operations, corresponding to a fixed schedule of 95 clock cycles. At the nominal simulation frequency of 200 MHz , the event-to-event interval measured in the waveform is 472 ns . This interval covers the whole on-chip pipeline, from the loading of the input activations into the input buffer to the writing of the two output codes into the output buffer, and it includes the on-chip buffer transactions needed to make the next input sample available.
The numerical correctness of the design was then established on the physical device. The complete network was configured on the Alveo U250 and executed over the entire 5000-sample test set, and its integer outputs were compared, before dequantization, with those produced by the reference 8-bit quantized model: the two agree on every one of the 5000 output T1 and T2 couples. It follows that the error metrics of the FPGA implementation and those of the QAT model are identical, and that the residual error of the deployed system with respect to the ground truth, summarized in Table 2, is entirely the error of the quantized network. Consistently with the decomposition of Table 1, that error is dominated by the 8-bit numerical representation, which costs 1.20 percentage points of global MAPE, rather than by the topology simplification, which costs 0.68 ; the hardware contributes nothing.
The MPE remains within a tenth of a percentage point of zero for T1 ( 0.12 % ), but reaches − 3.12 % for T2, against 0.02 % for the same network in floating point, indicating that quantization introduces a systematic overestimation to the T2 estimate. This bias is concentrated where the output representation is coarsest: resolved by decile of the true value, it is − 22.6 % over the lowest tenth of the T2 range and falls below one percentage point above 800 ms . Since the output scale is 15.8 ms and the output zero-point coincides with the bottom of the representable range, a true T2 of a few tens of milliseconds can be represented by only two distinct numbers and is necessarily pushed upwards. If the lowest decile is excluded, the MAPE on T2 falls from 11.07 % to 7.61 % .

3.2. FPGA Implementation

After functional validation, the FCN was synthesized, placed, and routed on the target board. Table 3 shows the resource utilization extracted from the post-implementation report, which indicates that the design occupies only a modest fraction of the available resources, without any timing violations.
After place-and-route and successful deployment on the FPGA device, post-implementation timing analysis reported a maximum achievable clock frequency of 83.3 MHz , lower than the 200 MHz used in simulation. This value represents the practical operating limit of the current firmware, imposed by the critical paths of the arithmetic and control logic, and it is the frequency at which the accelerator was actually run. Figure 7 shows the spatial distribution of the instantiated modules after place-and-route.
Since the number of clock cycles per inference is fixed by construction, the latency can be derived from the implemented clock frequency. The fixed schedule of 95 cycles corresponds to 475 ns at the nominal simulation frequency of 200 MHz , consistent with the 472 ns event-to-event interval measured in the functional simulation waveform. At the post-implementation operating frequency of 83.3 MHz , the same 95-cycle schedule corresponds to a practical per-inference latency of 1.14 μs. This value was also confirmed during the actual execution of the inference workload on the Alveo U250, providing experimental confirmation of the expected latency. This represents the practically achieved per-inference latency, and the corresponding processing time for the full 5000-sample dataset is 5.7 × 10 − 3 s , compared with 7.69 s for the software baseline, resulting in an acceleration factor of 1.35 × 10 3 .
Post-implementation power analysis estimated a power consumption of 8.6 W at the operating frequency of 83.3 MHz . Combining this with the rated power of the reference workstation ( 250 W ), the processing of the 5000-sample dataset corresponds to an estimated energy consumption of 1.9 × 10 3 J for the workstation and 4.9 × 10 − 2 J for the FPGA, i.e., an estimated energy-efficiency improvement of approximately 3.9 × 10 4 , as shown in Table 4. Considering that a typical MRF brain slice, such as those shown in Figure 2 contains ∼ 180 × 180 pixels, it results in a processing time of 3.69 × 10 − 2 s on the FPGA against 49.8 s on the workstation, and an estimated energy consumption of 3.17 × 10 − 1 J against 1.25 × 10 4 J .

3.3. hls4ml Comparison

The proposed VHDL architecture is compared with three tested hls4ml configurations, starting from the same quantized Keras model and using default configuration parameters. Table 5 shows the synthesis estimates obtained for three representative reuse-factor configurations, together with the corresponding resource utilization, maximum achievable clock frequency, and inference latency, together with the synthesis results of the proposed architecture and the original software-based model for comparison purposes. The development of the three hls4ml configurations became necessary as the default settings failed in the implementation; as shown in Table 5 model 2.1 requires more than the available logic resources, exceeding the device capacity with 114% of LUTs and 100% of DSPs, despite achieving a low estimated latency of 93.6 ns .
As a second attempt, a strong serialization factor of 1024 is applied (model 2.2). While this configuration drastically reduces DSP usage, it still consumes 43 % of LUTs and limits the maximum frequency to 114 MHz (Table 5). More importantly, although implementation completed successfully, timing closure could not be achieved. The final design reported a Worst Negative Slack (WNS) of − 8.62 ns and a Total Negative Slack (TNS) of −73.35 μs, with 38,571 failing timing endpoints, indicating that the synthesized design architecture cannot reliably operate at the target clock frequency.
Finally, the third configuration (model 2.3) was tested. As reported in Table 5, even in this case the hls4ml design occupies 42 % of LUTs and 11 % of DSPs, and reports a post-synthesis maximum clock frequency of 147 MHz , with an estimated latency of 1.03 μs. During full place-and-route, the hls4ml configuration with reuse factor 32 failed during implementation because the router reported a global routing congestion level of 7, preventing successful routing and timing closure despite the moderate logic utilization.

4. Discussion

The primary contribution of this work is not the implementation of a single FCN on FPGA, but the definition of a reusable low-level hardware design methodology that allows FCNs to be systematically constructed from a parameterizable node abstraction. This introduces a new approach to the use of NNs for quantitative MRI image reconstruction, based on hardware-accelerated inference of an FCN network for MRF. The success of this approach could open perspectives for the development of embedded systems for the real-time reconstruction of quantitative maps, to be applied directly in the context of MRI acquisitions in clinical settings. FPGAs have shown to be particularly suitable hardware platforms for the implementation of deep learning algorithms due to their flexibility in configuration that can be tailored to specific computational patterns, enabling optimized parallelism, low-latency processing, and efficient use of on-chip resources while maintaining relatively low power consumption. Here, we propose a modular approach for implementing FCNs on FPGAs based on a node function that already integrates all required operations, so that one only needs to instantiate this node an appropriate number of times to construct the full network, thereby greatly simplifying the design and implementation of FCNs.

4.1. Numerical Accuracy Evaluation

The comparison carried out on the device over the entire 5000-sample test set showed agreement on all 5000 integer output T1 and T2 couples, before dequantization. Every int8 value computed by the accelerator is the same value that the reference quantized model produces for the same input. The residual error of the deployed system is therefore entirely the error of the quantized network, and the hardware adds nothing to it.
In line with previous deep-learning MRF studies, reconstruction performance was evaluated using MAPE, MPE, and RMSE with respect to the ground truth. Since the accelerator reproduces the quantized model exactly, the question of accuracy moves entirely from the hardware to the numerical representation, and Table 1 locates it precisely: removing the two widest hidden layers costs 0.68 percentage points of global MAPE, whereas the transition to 8-bit arithmetic costs 1.20 , almost twice as much. The dominant source of the residual error is therefore the numerical representation, not the architecture simplification and not the FPGA implementation. The same conclusion follows from the MPE, which remains within a tenth of a percentage point of zero for T 1 but reaches − 3.12 % for T 2 : quantization introduces a systematic overestimation, concentrated in the lowest decile of the T 2 range where the 15.8 ms output step is a large fraction of the value being represented. This bias is deterministic, and is therefore consistent with the reproducible nature of the architecture. More importantly, it identifies a well-defined solution: widening the output representation beyond eight bits, an intervention confined to the requantization of the last layer that would leave the base node, the memory organization, and the scheduling unchanged and would use a negligible fraction of the available resources.
It is worth being explicit about what the FPGA implementation is and is not intended to do: it does not improve the intrinsic predictive accuracy of the neural network, but deploys it in a deterministic, verifiable, and bounded-latency hardware form. From a quantitative MRI perspective, relative errors in the 8–15% range for T 2 mapping are commonly reported in vivo [32,33,34], so that the 11.07 % observed on the synthetic test set falls within the range in which quantitative-imaging applications already operate. This is a limitation of the numerical representation of the model, not of its hardware realization, and it is precisely the component that widening the output stage would remove.

4.2. Post-Implementation Timing and Power Analysis

After successful implementation and deployment of the FCN on the target FPGA board, post-implementation resource usage analysis (Table 3 and Figure 7) shows the efficiency and scalability of the proposed FCN architecture. The design occupies only 12.29% of the available LUTs and 6.51% of FFs, indicating that the control logic and data path are implemented with a limited footprint. The usage of LUTRAM is negligible, confirming that the architecture does not rely heavily on distributed memory resources. A more significant portion of resources is represented by the DSP blocks, with 33.85% utilization. This behavior is expected, as the FCN computation is dominated by multiply–accumulate operations. The DSP usage grows in a predictable and regular manner with the number of instantiated nodes, making performance and scaling easy to control during architectural exploration.
Post-implementation timing analysis indicated a maximum operating frequency of 83.3 MHz . This resulting frequency is lower than the targeted 200 MHz . However, an acceleration factor of 1.35 × 10 3 is achieved compared to the software. An additional advantage of the selected operating point is power efficiency. Increasing the clock frequency on FPGA platforms leads to an approximately linear increase in dynamic power consumption. According to the FPGA post-implementation power analysis report, operation at 83.3 MHz is associated with an estimated power consumption of 8.6 W . Thus, the FPGA implementation suggests an energy-efficiency advantage of about four orders of magnitude for the 5000-sample dataset. In fact, considering a typical MRF brain slice composed of approximately 180 × 180 pixels, the estimated reconstruction time is about 49.8 s on the workstation and 3.69 × 10 − 2 s on the FPGA. As shown in Table 4, this corresponds to an energy consumption of approximately 1.25 × 10 4 J and 0.32 J , respectively, further highlighting the significant advantage of the FPGA-based implementation in both latency and energy efficiency, demonstrating the potential of the proposed approach for energy-efficient deployment in large-scale or real-time medical imaging applications. Both sides of this comparison are estimates and not metered measurements: the FPGA value derives from design-tool power analysis, which uses circuit-level switching activity to predict dissipation and does not account for board-level power-delivery losses or host idle power, and the workstation estimate derives from the rated power of the platform rather than from a wall-power measurement taken during execution. The relative comparison in Table 4 is reported here as an order of magnitude.
Unlike software implementations on CPU or GPU, where execution time varies depending on memory access patterns, scheduling, and parallel workload, the FPGA design operates with a fully deterministic pipeline, and each inference is completed with constant and predictable time. It is worth being precise about the origin of this property: the latency is not observed to be constant, it is constructed to be constant, since the schedule is encoded in a constant table fixed before synthesis and the entire parameter set of 49 , 170 parameters resides on-chip, so that no transfer over the PCIe interface takes place during inference. The number of clock cycles required by an inference is therefore a property of the design and not of the data. The acceleration factor of 1.35 × 10 3 actually achieved after implementation, together with this fixed and predictable latency, is what makes the solution relevant for real-time applications.
From a trustworthy cyber-physical systems perspective, the deterministic behavior of the proposed FPGA implementation is as important as its acceleration factor. In a real-time MRI acquisition/reconstruction workflow, bounded latency and reproducible numerical behavior are essential for reliable operation, while verifiable hardware execution and controlled on-chip data and parameter storage can contribute to system integrity and security. These are properties that would facilitate integration into safety-relevant medical imaging systems. No system-level verification, no interfacing with acquisition hardware, no safety or security analysis, and no clinical evaluation have been performed in this work, as they fall beyond the scope of the present study. The discussion of clinical and safety-critical scenarios in this paper should be read as a statement of potential application rather than of demonstrated capability.

4.3. Scaling Behavior

Only one network size and one degree of parallelism were implemented in this work, so that the scalability of the architecture cannot be claimed as an experimental result. It can, however, be stated as an explicit model, because the composition of the resource usage is known by construction. The DSP occupation, which is the resource that will be exhausted first, decomposes as
N DSP = W · F max + 4 W = 16 × 256 + 4 × 16 = 4160 ,
where W = 16 is the width of the node block and F max = 256 the largest fan-in it must sustain, the four additional units per node being the 32-bit requantization multiplier. The number of clock cycles of an inference is likewise fixed before synthesis and is given by
N cyc = ∑ ℓ = 1 L n ℓ W + d ,
with n ℓ the number of neurons of layer ℓ and d the constant drain of the node pipeline at each layer transition, so that the inference latency is N cyc / f clk .
Two consequences follow, and they are the reason why the two directions of growth can be quantified before implementation. Adding layers, or widening a layer beyond W, lengthens the schedule of Equation (5) and leaves the arithmetic occupation of Equation (4) unchanged; it costs latency, not DSPs. Widening the node block, instead, costs DSPs in proportion to W and reduces the latency correspondingly. The depth of the network is therefore absorbed by time multiplexing and does not enter the resource budget at all, whereas F max and W determine it entirely. What Equations (4) and (5) do not predict is timing closure, which depends on placement and routing and can become binding well before the resource budget does; the maximum achievable frequency of a scaled design must therefore still be established by implementation. Within these limits, and provided that the parameter set continues to fit on-chip, the framework can be extended to larger FCN configurations in a controlled way.

4.4. hls4ml Model Analysis

As a final point of comparison, the proposed VHDL architecture was evaluated against an implementation generated with the hls4ml framework, starting from the same quantized Keras model and using default configuration parameters. Although hls4ml is designed to simplify the deployment of NNs on FPGA, the practical realization of the considered FCN turned out to be non-trivial and required substantial effort to obtain meaningful synthesis results.
As a first attempt, the entire network is implemented fully in parallel (model 2.1), where all nodes are instantiated simultaneously. This configuration was obtained by using all the default hls4ml settings, using it as a pure automated tool. As shown in Table 5, with these configurations, the design requires more than the available logic resources, showing how a brute-force parallel mapping of the FCN is not practical for networks of this size, and results in the impossibility of deploying such a model on the FPGA board. To reduce resource usage, a second hls4ml model (model 2.2) was built, using a strong serialization factor of 1024. In this way, only a very small fraction of the nodes is hardware-developed and used iteratively to cover all the operations of the FCN. Whilst drastically reducing the necessary resources, it still requires 43 % of the available LUTs. The reuse-1024 implementation successfully completed place-and-route. However, timing closure was not achieved, with a WNS of − 8.62 ns affecting more than 38,000 timing endpoints. This indicates that, although aggressive serialization greatly reduces DSP utilization, the resulting architecture introduces long critical paths that prevent operation at the intended clock frequency.
Finally, model 2.3 was tested to provide an additional comparison between the proposed modular VHDL approach and hls4ml. As shown in Table 5, also in this case the implementation on the FPGA board failed to complete. Consequently, the reported latency of 1.03 μ s and maximum frequency of 147 MHz for model 2.3 cannot be considered achievable without substantial redesign or additional optimization. Although the overall LUT utilization remained at only 42%, implementation still failed because the router reported a global congestion level of 7, indicating that routing complexity rather than raw resource availability became the limiting factor. In contrast, the proposed custom VHDL architecture successfully completed place-and-route and achieved timing closure at 83.3 MHz . The placement shown in Figure 7 confirms a controlled resource utilization with available device headroom, supporting the scalability and practical deployability of the design.
To sum up, none of the three hls4ml configurations produced a working bitstream, and consequently no hls4ml baseline could be executed on the device. The comparison reported here therefore documents that, for this network, on this device, and under the default and reuse-factor configurations tested, the automated flow did not reach a deployable implementation while the proposed custom design did. It does not establish that the proposed architecture is superior to hls4ml in general: a different device, a smaller network, or a more extensive manual tuning of the hls4ml directives, of the precision settings, or of the partitioning strategy could well yield a working and competitive design, and no such exploration was attempted here.

5. Conclusions

This work presents a fully custom FPGA implementation of a quantized FCN for MRF map reconstruction, built around a modular and reusable VHDL node module architecture. Rather than targeting a single ad hoc accelerator, the proposed approach defines a systematic methodology for constructing FCNs with deterministic, pipelined inference and fixed latency, while providing substantial acceleration over CPU-based execution.
The design was synthesized, placed, routed, configured on the target Alveo U250 board, and executed on the device over the entire 5000-sample test set, where it agreed with the reference 8-bit quantized model on every one of the 5000 integer output T1 and T2 couples. The accuracy of the accelerator, a MAPE of 2.37 % on T 1 and 11.07 % on T 2 with respect to the ground truth, is therefore not merely comparable to that of the software model but identical to it, and the residual error is attributable to the 8-bit representation rather than to the hardware. After place-and-route, the design closed timing at 83.3 MHz , giving a fixed per-inference latency of 1.14 μs and an acceleration factor of 1.35 × 10 3 relative to the software baseline, with 12.29 % LUT and 33.85 % DSP occupation. The associated energy-efficiency figure of 3.9 × 10 4 should be interpreted as an indicative estimate based on design-tool power analysis and on the rated power of the reference workstation, and not as a direct wall-power measurement.
The comparison with the hls4ml framework was carried out under the default and reuse-factor configurations of the tool, and none of the three produced a deployable design on this device for this network, whereas the proposed architecture did. This supports the usefulness of a low-level architectural approach when strict timing and implementation constraints apply, but it is a statement about this case study and not a general claim of superiority: automated toolflows remain the appropriate choice for smaller networks, for rapid design-space exploration, and whenever development effort is the dominant constraint.

5.1. Limitations

Four limitations should be made explicit. First, the validation was carried out entirely on simulated MRF signals produced by the same forward model used to generate the training data, so the accuracy values reported here do not constitute a validation on in vivo data, where field inhomogeneity, partial-volume effects, and motion become further sources of error. Second, the energy comparison is based on estimates on both sides, since the FPGA figure derives from post-implementation power analysis and the workstation figure from rated power; a metered measurement would be required before it could be reported as a result rather than as an order of magnitude. Third, only one network topology and one degree of parallelism were implemented, on a single high-end device: the scaling relations of Section 4.3 indicate how resource usage and latency should behave elsewhere, but this was not verified, and in particular timing closure on smaller devices, where the logic budget would become binding considerably sooner, remains to be established. Fourth, the architecture is currently restricted to fully connected networks with ReLU or linear activations whose parameter set fits entirely on-chip; convolutional and encoder–decoder models would require a different memory organization and scheduling strategy.

5.2. Future Work

Four developments follow directly from these limitations. The most immediate is widening the output representation beyond eight bits, an intervention confined to the requantization of the last layer that would leave the base node, the memory organization, and the scheduling unchanged, would use a negligible fraction of the available logic, and would address the dominant component of the residual error. The second is validation on in vivo acquisitions, which is a prerequisite for any clinical claim. The third is implementation on smaller and less expensive devices, which is the condition for the portable and point-of-care deployments discussed above. The fourth is the automatic generation of the parameter store and of the schedule table from a trained model, which would turn the methodology presented here into a complete design flow: both tables are data derived from the topology and from the quantization parameters, and generating them from a single description would also guarantee their mutual consistency.
Beyond the specific network implemented in this work, the proposed design represents a first step toward the development of a modular VHDL library for neural network acceleration on FPGA. By encapsulating arithmetic, quantization, activation, and control into reusable building blocks, the framework facilitates systematic low-level NN design and lowers the barrier to adopting FPGA-based neural accelerators, bridging the gap between algorithm development and efficient hardware deployment for real-time medical imaging and other latency-critical applications.

Author Contributions

Conceptualization, A.G., C.T., L.B. and M.R.; methodology, M.R. and F.A.; software, M.R. and F.A.; validation, M.R. and F.A.; formal analysis, M.R.; investigation, M.R. and F.A.; resources, A.G. and F.A.; data curation, M.R., M.B. and L.B.; writing—original draft preparation, M.R.; writing—review and editing, M.R., F.A., M.B., A.R., L.B., C.T. and A.G.; figures, L.B., A.G., F.A. and M.R.; supervision, C.T. and A.G.; project administration, C.T., A.G., L.B. and C.T. contributed equally to this work as senior authors. All authors have read and agreed to the published version of the manuscript.

Funding

This work was partially supported by the National Institute for Nuclear Physics (INFN), Bologna Section, which provided the computing infrastructure, FPGA boards, and software licenses required for the development and testing of the proposed architectures. Additionally, human resources and laboratory facilities were made available by the Department of Physics and Astronomy at the University of Bologna. No further external funding was received.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data used in this study consist of simulated Magnetic Resonance Fingerprinting (MRF) signals generated over a broad range of physiologically relevant T 1 and T 2 values, as described in the manuscript. The dataset was created for benchmarking and validating the software and FPGA implementations of the proposed neural network. Since all samples can be reproduced from the simulation procedure reported in the paper, no dedicated public repository is associated with this work. Additional information and representative data samples are available from the authors upon reasonable request.

Acknowledgments

The authors gratefully acknowledge the Istituto Nazionale di Fisica Nucleare (INFN), Sezione di Bologna, for providing access to the computational infrastructure, FPGA hardware, and software licenses that made this research possible. The authors also thank Camilla Marella for her contribution throughout the development of this work.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
BRAM  Block RAM
CLBConfigurable Logic Block
CNNConvolutional Neural Network
CPUCentral Processing Unit
DSPDigital Signal Processor
FCNFully Connected Network
FFFlip-Flop
FPGAField-Programmable Gate Array
GPUGraphics Processing Unit
HLSHigh-Level Synthesis
LUTLook-Up Table
LUTRAMLook-Up Table RAM
MAPEMean Absolute Percentage Error
MMCMMixed-Mode Clock Manager
MRIMagnetic Resonance Imaging
MRFMagnetic Resonance Fingerprinting
NNNeural Network
PCIePeripheral Component Interconnect Express
QATQuantization-Aware Training
ReLURectified Linear Unit
RNNRecurrent Neural Network
TNSTotal Negative Slack
TPUTensor Processing Unit
VHDLVHSIC Hardware Description Language
WNSWorst Negative Slack

References

  1. Alom, M.Z.; Taha, T.M.; Yakopcic, C.; Westberg, S.; Sidike, P.; Nasrin, M.S.; Hasan, M.; Van Essen, B.C.; Awwal, A.A.S.; Asari, V.K. A State-of-the-Art Survey on Deep Learning Theory and Architectures. Electronics 2019, 8, 292. [Google Scholar] [CrossRef] [Scilit]
  2. LeCun, Y.; Bengio, Y.; Hinton, G. Deep Learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [Scilit]
  3. Dong, S.; Wang, P.; Abbas, K. A survey on deep learning and its applications. Comput. Sci. Rev. 2021, 40, 100379. [Google Scholar] [CrossRef] [Scilit]
  4. Feng, X.; Jiang, Y.; Yang, X.; Du, M.; Li, X. Computer vision algorithms and hardware implementations: A survey. Integration 2019, 69, 309–320. [Google Scholar] [CrossRef] [Scilit]
  5. Ahmed, N.; Saha, A.K.; Al Noman, M.A.; Jim, J.R.; Mridha, M.; Kabir, M.M. Deep learning-based natural language processing in human–agent interaction: Applications, advancements and challenges. Nat. Lang. Process. J. 2024, 9, 100112. [Google Scholar] [CrossRef] [Scilit]
  6. Supriyono; Wibawa, A.P.; Suyono; Kurniawan, F. Advancements in natural language processing: Implications, challenges, and future directions. Telemat. Inform. Rep. 2024, 16, 100173. [Google Scholar] [CrossRef] [Scilit]
  7. Barbieri, M.; Lee, P.K.; Brizi, L.; Giampieri, E.; Solera, F.; Castellani, G.; Hargreaves, B.A.; Testa, C.; Lodi, R.; Remondini, D. Circumventing the curse of dimensionality in magnetic resonance fingerprinting through a deep learning approach. NMR Biomed. 2022, 35, e4670. [Google Scholar] [CrossRef] [Scilit]
  8. Barbieri, M.; Brizi, L.; Giampieri, E.; Solera, F.; Manners, D.N.; Castellani, G.; Testa, C.; Remondini, D. A deep learning approach for magnetic resonance fingerprinting: Scaling capabilities and good training practices investigated by simulations. Phys. Med. 2021, 89, 80–92. [Google Scholar] [CrossRef] [Scilit]
  9. Fiscone, C.; Curti, N.; Ceccarelli, M.; Remondini, D.; Testa, C.; Lodi, R.; Tonon, C.; Manners, D.N.; Castellani, G. Generalizing the Enhanced-Deep-Super-Resolution Neural Network to Brain MR Images: A Retrospective Study on the Cam-CAN Dataset. eNeuro 2024, 11, ENEURO.0458-22.2023. [Google Scholar] [CrossRef] [Scilit]
  10. Abadi, M.; Agarwal, A.; Barham, P.; Brevdo, E.; Chen, Z.; Citro, C.; Corrado, G.S.; Davis, A.; Dean, J.; Devin, M.; et al. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. 2015. Available online: https://www.tensorflow.org (accessed on 17 January 2024).
  11. Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. arXiv 2019, arXiv:1912.01703. [Google Scholar] [CrossRef] [Scilit]
  12. Lata, K.; Singh, P.; Saini, S.; Cenkeramaddi, L.R. Privacy-preserving brain tumor detection using FPGA-accelerated deep learning on Kria KV260 for smart healthcare. Comput. Methods Programs Biomed. Update 2025, 8, 100205. [Google Scholar] [CrossRef] [Scilit]
  13. Xiong, S.; Wu, G.; Fan, X.; Feng, X.; Huang, Z.; Cao, W.; Zhou, X.; Ding, S.; Yu, J.; Wang, L.; et al. MRI-based brain tumor segmentation using FPGA-accelerated neural network. BMC Bioinform. 2021, 22, 421. [Google Scholar] [CrossRef] [Scilit]
  14. Soyak, R.; Navruz, E.; Ersoy, E.O.; Cruz, G.; Prieto, C.; King, A.P.; Oksuz, I. Channel Attention Networks for Robust MR Fingerprint Matching. IEEE Trans. Biomed. Eng. 2022, 69, 1398–1405. [Google Scholar] [CrossRef] [Scilit]
  15. Amano, H. (Ed.) Principles and Structures of FPGAs, 1st ed.; Springer: Singapore, 2018; p. 231. [Google Scholar] [CrossRef] [Scilit]
  16. Yan, F.; Koch, A.; Sinnen, O. A survey on FPGA-based accelerator for ML models. arXiv 2024, arXiv:2412.15666. [Google Scholar] [CrossRef] [Scilit]
  17. Mittal, S. A Survey of FPGA-based Accelerators for Convolutional Neural Networks. Neural Comput. Appl. 2020, 32, 1109–1139. [Google Scholar] [CrossRef] [Scilit]
  18. Xilinx/AMD. Vitis AI: Integrated Development Environment for AI Inference. Available online: https://github.com/Xilinx/Vitis-AI (accessed on 23 December 2025).
  19. Fahim, F.; Hawks, B.; Herwig, C.; Hirschauer, J.; Jindariani, S.; Tran, N.; Carloni, L.P.; Guglielmo, G.D.; Harris, P.; Krupa, J.; et al. hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices. arXiv 2021, arXiv:2103.05579. [Google Scholar] [CrossRef] [Scilit]
  20. Boudjadar, J.; Islam, S.U.; Buyya, R. Dynamic FPGA reconfiguration for scalable embedded artificial intelligence (AI): A co-design methodology for convolutional neural networks (CNN) acceleration. Future Gener. Comput. Syst. 2025, 169, 107777. [Google Scholar] [CrossRef] [Scilit]
  21. Duarte, J.; Han, S.; Harris, P.; Jindariani, S.; Kreinar, E.; Kreis, B.; Ngadiuba, J.; Pierini, M.; Rivera, R.; Tran, N.; et al. Fast inference of deep neural networks in FPGAs for particle physics. J. Instrum. 2018, 13, P07027. [Google Scholar] [CrossRef] [Scilit]
  22. Jariri, N.; Chaman, M.; Majdoubi, R.; Hadjoudja, A. Evaluating HLS4ML For Converting a CNN Model Into High Level Synthesis. In Proceedings of the 2025 5th International Conference on Advanced Research in Computing (ICARC), Belihuloya, Sri Lanka, 19–20 February 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  23. Sarg, M.; Khalil, A.H.; Mostafa, H. Efficient HLS Implementation for Convolutional Neural Networks Accelerator on an SoC. In Proceedings of the 2021 International Conference on Microelectronics (ICM), New Cairo City, Egypt, 19–22 December 2021; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  24. Ramhorst, B.; Lončar, V.; Constantinides, G.A. FPGA Resource-aware Structured Pruning for Real-Time Neural Networks. In Proceedings of the 2023 International Conference on Field Programmable Technology (ICFPT); IEEE: Piscataway, NJ, USA, 2023; pp. 282–283. [Google Scholar] [CrossRef] [Scilit]
  25. Athur, D.K.; Pawar, R.; Arora, A. Out-of-the-Box Performance of FPGAs for ML Workloads Using Vitis AI. In Proceedings of the Applied Reconfigurable Computing. Architectures, Tools, and Applications; Giorgi, R., Stojilović, M., Stroobandt, D., Brox Jiménez, P., Barriga Barros, Á., Eds.; Springer: Cham, Switzerland, 2025; pp. 123–139. [Google Scholar]
  26. Kokkinis, A.; Siozios, K. Fast Resource Estimation of FPGA-Based MLP Accelerators for TinyML Applications. Electronics 2025, 14, 247. [Google Scholar] [CrossRef] [Scilit]
  27. Gabrielli, A. Readout demonstrator for a large-scale pixel-detector conforming to the ATLAS Phase-II Upgrade. Nucl. Instrum. Methods Phys. Res. Sect. A Accel. Spectrom. Detect. Assoc. Equip. 2020, 978, 164409. [Google Scholar] [CrossRef] [Scilit]
  28. Landman, B.A.; Huang, A.J.; Gifford, A.; Vikram, D.S.; Lim, I.A.L.; Farrell, J.A.; Bogovic, J.A.; Hua, J.; Chen, M.; Jarso, S.; et al. Multi-Parametric neuroimaging reproducibility: A 3-T resource study. NeuroImage 2011, 54, 2854–2866. [Google Scholar] [CrossRef] [Scilit]
  29. TensorFlow Model Optimization Team. Quantization Aware Training Guide. 2025. Available online: https://www.tensorflow.org/model_optimization/guide/quantization/training (accessed on 23 December 2025).
  30. Nahshan, Y.; Chmiel, B.; Baskin, C.; Zheltonozhskii, E.; Banner, R.; Bronstein, A.M.; Mendelson, A. Loss aware post-training quantization. Mach. Learn. 2021, 110, 3245–3262. [Google Scholar] [CrossRef] [Scilit]
  31. Rosenblatt, F. The perceptron: A probabilistic model for information storage and organization in the brain. Psychol. Rev. 1958, 65, 386–408. [Google Scholar] [CrossRef] [Scilit]
  32. Lo, W.C.; Bittencourt, L.K.; Panda, A.; Jiang, Y.; Tokuda, J.; Seethamraju, R.; Tempany-Afdhal, C.; Obmann, V.; Wright, K.; Griswold, M.; et al. Multicenter Repeatability and Reproducibility of MR Fingerprinting in Phantoms and in Prostatic Tissue. Magn. Reson. Med. 2022, 88, 1818–1827. [Google Scholar] [CrossRef] [Scilit]
  33. Jordanova, K.V.; Ogier, S.E.; Russek, S.E.; Stoffer, C.M.; Stupic, K.F.; Buonincontri, G.; Nittka, M.; Keenan, K.E. Adult brain T1 and T2 values measured at 3 T using magnetic resonance fingerprinting with phantom validation. J. Appl. Clin. Med. Phys. 2025, 26, e70289. [Google Scholar] [CrossRef] [Scilit]
  34. Reitz, S.C.; Hof, S.M.; Fleischer, V.; Brodski, A.; Gröger, A.; Gracien, R.M.; Droby, A.; Steinmetz, H.; Ziemann, U.; Zipp, F.; et al. Multi-parametric quantitative MRI of normal appearing white matter in multiple sclerosis, and the effect of disease activity on T2. Brain Imaging Behav. 2017, 11, 744–753. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Flow diagram of the five-stage development and evaluation workflow: adaptation and quantization of the reference FCN, design of the reusable base node in VHDL, implementation of the complete network by time-multiplexing a single 16-node block over all 498 neurons, verification at four distinct levels, and comparison against three hls4ml configurations and the software reference. Darker boxes mark the outcome of each stage; section numbers in the left-hand labels refer to the subsections in which each stage is detailed.
Figure 1. Flow diagram of the five-stage development and evaluation workflow: adaptation and quantization of the reference FCN, design of the reusable base node in VHDL, implementation of the complete network by time-multiplexing a single 16-node block over all 498 neurons, verification at four distinct levels, and comparison against three hls4ml configurations and the software reference. Darker boxes mark the outcome of each stage; section numbers in the left-hand labels refer to the subsections in which each stage is detailed.
Applsci 16 09935 g001
Figure 2. Example instantiation of the FCN that motivated the development of the proposed modular FPGA architecture. The first principal components obtained after k-SVD truncation of the MRF signal are provided as input to a 256–128–64–32–16 network, which outputs the quantitative relaxation parameters T1 and T2. The corresponding reconstructed parameter maps are shown on the right. Brain MRF maps were obtained by processing real acquisitions downloaded from the Multi-Modal MRI Reproducibility Resource repository [28].
Figure 2. Example instantiation of the FCN that motivated the development of the proposed modular FPGA architecture. The first principal components obtained after k-SVD truncation of the MRF signal are provided as input to a 256–128–64–32–16 network, which outputs the quantitative relaxation parameters T1 and T2. The corresponding reconstructed parameter maps are shown on the right. Brain MRF maps were obtained by processing real acquisitions downloaded from the Multi-Modal MRI Reproducibility Resource repository [28].
Applsci 16 09935 g002
Figure 4. Block-level architecture of the FCN implemented on FPGA. The network is composed of an input layer, five hidden layers, and an output layer, where each node corresponds to a reusable Neuron Block. The inset shows the internal architecture of a single Neuron Block, which constitutes the fundamental building unit of the network.
Figure 4. Block-level architecture of the FCN implemented on FPGA. The network is composed of an input layer, five hidden layers, and an output layer, where each node corresponds to a reusable Neuron Block. The inset shows the internal architecture of a single Neuron Block, which constitutes the fundamental building unit of the network.
Applsci 16 09935 g004
Figure 5. Schematic overview of the three-way evaluation pipeline used in this work. The same input dataset, composed of 5000 simulated MRF samples, is processed by (1) the proposed custom-designed FPGA-based hardware accelerator, (2) three hls4ml-generated FPGA designs, corresponding to different degrees of parallelism, and (3) the reference software implementation running in Python/Keras. The outputs of all three branches are compared in terms of latency and resource usage.
Figure 5. Schematic overview of the three-way evaluation pipeline used in this work. The same input dataset, composed of 5000 simulated MRF samples, is processed by (1) the proposed custom-designed FPGA-based hardware accelerator, (2) three hls4ml-generated FPGA designs, corresponding to different degrees of parallelism, and (3) the reference software implementation running in Python/Keras. The outputs of all three branches are compared in terms of latency and resource usage.
Applsci 16 09935 g005
Figure 6. Waveform of the VHDL simulation for the FCN. The figure shows the input activations provided to the 16-node block, and the final quantized output activations produced after one complete inference cycle.
Figure 6. Waveform of the VHDL simulation for the FCN. The figure shows the input activations provided to the 16-node block, and the final quantized output activations produced after one complete inference cycle.
Applsci 16 09935 g006
Figure 7. Post-implementation placement of the complete custom FCN architecture on the target FPGA device, including on-chip memory blocks for weights and intermediate activations. The figure shows the spatial distribution of logic resources after place-and-route. The design, shown in green, occupies a limited and well-confined portion of the device, leaving sufficient headroom for further replication or integration with additional processing modules.
Figure 7. Post-implementation placement of the complete custom FCN architecture on the target FPGA device, including on-chip memory blocks for weights and intermediate activations. The figure shows the spatial distribution of logic resources after place-and-route. The design, shown in green, occupies a limited and well-confined portion of the device, leaving sufficient headroom for further replication or integration with additional processing modules.
Applsci 16 09935 g007
Table 1. Error metrics of the three model configurations on the 5000-sample test set, computed with respect to the ground-truth relaxation times. The first two rows isolate the effect of the architecture simplification, the last two the effect of the 8-bit quantization.
Table 1. Error metrics of the three model configurations on the 5000-sample test set, computed with respect to the ground-truth relaxation times. The first two rows isolate the effect of the architecture simplification, the last two the effect of the 8-bit quantization.
T1T2
Configuration MAPE (%) MPE (%) RMSE (ms) MAPE (%) MPE (%) RMSE (ms)
Original, floating point 2.03 − 0.16 73 7.65 0.25 144
Adapted, floating point 2.15 − 0.66 75 8.89 0.02 145
Adapted, 8-bit quantized 2.37 0.12 78 11.07 − 3.12 149
Table 2. Error metrics of the three model configurations and of the FPGA implementation, computed over the same 5000-sample test set with respect to the ground-truth relaxation times. The last two rows are identical because they describe the same quantized model, executed respectively by the reference integer kernel in software and by the accelerator on the device: the two agree on every one of the 5000 output T1 and T2 couples of the test set. The first three rows are reproduced from Table 1.
Table 2. Error metrics of the three model configurations and of the FPGA implementation, computed over the same 5000-sample test set with respect to the ground-truth relaxation times. The last two rows are identical because they describe the same quantized model, executed respectively by the reference integer kernel in software and by the accelerator on the device: the two agree on every one of the 5000 output T1 and T2 couples of the test set. The first three rows are reproduced from Table 1.
T1T2
Configuration MAPE (%) MPE (%) RMSE (ms) MAPE (%) MPE (%) RMSE (ms)
Original, floating point 2.03 − 0.16 73 7.65 0.25 144
Adapted, floating point 2.15 − 0.66 75 8.89 0.02 145
Adapted, 8-bit quantized (software) 2.37 0.12 78 11.07 − 3.12 149
FPGA, 8-bit quantized * 2.37 0.12 78 11.07 − 3.12 149
* Identical to the preceding row: the two implementations agree on every one of the 5000 T1 and T2 output couples of the test set.
Table 3. Resource utilization of the proposed FCN design on the target FPGA device (AMD/Xilinx Alveo U250) after implementation (place-and-route).
Table 3. Resource utilization of the proposed FCN design on the target FPGA device (AMD/Xilinx Alveo U250) after implementation (place-and-route).
ResourceEstimationAvailableUtilization %
LUT212,4321,728,000 12.29
LUTRAM20791,040 0.01
FF225,0353,456,000 6.51
DSP416012,288 33.85
IO2676 0.30
MMCM116 6.25
Table 4. Estimated comparison of reconstruction time and energy consumption between the workstation and the FPGA implementation. Energy values are estimates on both sides, neither is a metered measurement. FPGA energy is derived from post-implementation power analysis, while workstation energy is estimated from typical rated power consumption.
Table 4. Estimated comparison of reconstruction time and energy consumption between the workstation and the FPGA implementation. Energy values are estimates on both sides, neither is a metered measurement. FPGA energy is derived from post-implementation power analysis, while workstation energy is estimated from typical rated power consumption.
5000-Sample DatasetSingle Brain Slice ( 180 × 180 Pixels)
Platform Time Energy Time Energy
Workstation 7.69 s 1.9 × 10 3 J 49.8 s 1.25 × 10 4 J
FPGA 5.7 × 10 − 3 s 4.9 × 10 − 2 J 3.69 × 10 − 2 s 3.17 × 10 − 1 J
Improvement 1.35 × 10 3 3.9 × 10 4 1.35 × 10 3 3.9 × 10 4
Table 5. Comparison of resource usage and performance for the FCN implemented with three different approaches: (Model 1) the proposed custom VHDL architecture on FPGA, (Model 2) FPGA designs generated with hls4ml using three different reuse factors, and (Model 3) the reference software-based implementation on CPU. For the hls4ml designs all values are post-synthesis estimates, since none of the three configurations produced a working bitstream. For the proposed design, the post-layout frequency and latency are the values obtained after place-and-route and actually used on the device.
Table 5. Comparison of resource usage and performance for the FCN implemented with three different approaches: (Model 1) the proposed custom VHDL architecture on FPGA, (Model 2) FPGA designs generated with hls4ml using three different reuse factors, and (Model 3) the reference software-based implementation on CPU. For the hls4ml designs all values are post-synthesis estimates, since none of the three configurations produced a working bitstream. For the proposed design, the post-layout frequency and latency are the values obtained after place-and-route and actually used on the device.
ModelReuse Factor% LUTs% DSPsMax Frequency (MHz)Freq Post-Layout (MHz)Latency (ns)
Custom FPGA132 12.29 33.85 200 83.3 1.14 · 10 3
hls4ml2.11114100192✗✗
2.21024430114✗✗
2.3324211147✗✗
Python/Keras31//////// ( 1.5 ± 0.6 ) · 10 6
Red color for the resources is used to highlight that the design saturates the available logics. The red crosses are used to highlight that the design did not close the timing and could not be implemented on the FPGA board.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ricchi, M.; Alfonsi, F.; Barbieri, M.; Retico, A.; Gabrielli, A.; Brizi, L.; Testa, C. A Modular VHDL Framework for Deterministic FPGA Acceleration of MR Fingerprinting Reconstruction. Appl. Sci. 2026, 16, 9935. https://doi.org/10.3390/app16199935

AMA Style

Ricchi M, Alfonsi F, Barbieri M, Retico A, Gabrielli A, Brizi L, Testa C. A Modular VHDL Framework for Deterministic FPGA Acceleration of MR Fingerprinting Reconstruction. Applied Sciences. 2026; 16(19):9935. https://doi.org/10.3390/app16199935

Chicago/Turabian Style

Ricchi, Mattia, Fabrizio Alfonsi, Marco Barbieri, Alessandra Retico, Alessandro Gabrielli, Leonardo Brizi, and Claudia Testa. 2026. "A Modular VHDL Framework for Deterministic FPGA Acceleration of MR Fingerprinting Reconstruction" Applied Sciences 16, no. 19: 9935. https://doi.org/10.3390/app16199935

APA Style

Ricchi, M., Alfonsi, F., Barbieri, M., Retico, A., Gabrielli, A., Brizi, L., & Testa, C. (2026). A Modular VHDL Framework for Deterministic FPGA Acceleration of MR Fingerprinting Reconstruction. Applied Sciences, 16(19), 9935. https://doi.org/10.3390/app16199935

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop