Next Article in Journal
Unified EEG Feature Extraction for Cross-Subject Driver State Recognition and a Leakage-Free Safe-Stop Trigger Mechanism
Previous Article in Journal
trafficBCDB: A Traffic-Aware Blockchain Database for Verifiable Transportation Data Sharing
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

C2fDeploy: Function-Preserving Graph Rewriting to Eliminate Runtime Split Overhead on FPGA Deep Learning Processing Units

1
School of Electrical and Information Engineering, Zhengzhou University, Zhengzhou 450001, China
2
School of Aerospace and Intelligent Equipment, Xihua University, Chengdu 610039, China
3
School of Integrated Circuits and Electronics, Beijing Institute of Technology, Beijing 100081, China
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(19), 4431; https://doi.org/10.3390/electronics15194431
Submission received: 16 August 2026 / Revised: 23 September 2026 / Accepted: 24 September 2026 / Published: 26 September 2026

Abstract

Efficient deployment of neural-network detectors on field-programmable gate array (FPGA) accelerators depends not only on model complexity but also on compiler-visible graph structure. On deep learning processing unit (DPU) platforms, unsupported operators can fragment execution between accelerator and host execution domains. We present C2fDeploy, a training-free rewrite for the YOLOv8 C2f block that moves channel splitting from the activation graph to an offline partition of the trained projection and batch-normalization parameters. The transformed block preserves the 32-bit floating-point (FP32) function without retraining or additional parameters. On the GRAZPEDWRI-DX fracture-detection task, the original and rewritten graphs produced identical FP32 test metrics and comparable 8-bit integer (INT8) accuracy. Direct tensor-level FP32 comparison at the outputs of all eight rewritten C2f blocks yielded an aggregate mean absolute error of 9.73 × 10 − 8 and a relative L 2 error of 1.97 × 10 − 7 , providing numerical verification beyond detection-level metrics. Compilation for the Kria KV260 consolidated nine DPU subgraphs into one and removed the C2f-related host-side slicing operations. In a same-checkpoint whole-XModel benchmark, C2fDeploy improved whole-XModel graph execution throughput by 95.2 × and reduced energy per execution on the 5 V system-on-module (SOM) rail by 98.2%. Direct runtime profiling further showed that DPU compute-unit busy-time utilization increased from 0.36% to 90.75%, while system-wide CPU utilization decreased by 89.5%. Aggregate APM-observed external-memory bandwidth increased from 40.20 to 3250.56 MB/s as accelerator execution became more continuous, whereas normalized APM-observed traffic decreased by 14.8% per graph execution. These results show that compiler-aware graph rewriting can remove deployment bottlenecks without changing the trained detector.

1. Introduction

Real-time automatic analysis of pediatric wrist radiographs on resource-constrained edge platforms requires both reliable detection accuracy and efficient inference. Pediatric wrist fractures can be difficult to localize when subtle fracture lines overlap with normal anatomical structures, motivating computer-assisted detection methods [1,2,3]. The GRAZPEDWRI-DX dataset provides expert-annotated pediatric wrist radiographs and has become a reproducible benchmark for fracture localization [4].
Recent work on this dataset has primarily focused on detection accuracy through data augmentation, alternative YOLO variants, attention mechanisms, and multi-scale feature fusion [5,6,7,8,9,10]. However, translating an accurate detector to an edge platform based on a field-programmable gate array (FPGA) introduces a different challenge: the compiled operator graph can determine whether computation remains on the accelerator or is fragmented across accelerator and host execution.
Deployment efficiency is not determined by parameter count or floating-point operations alone. Measured latency can depend strongly on operator structure, memory access, and compiler lowering [11]. Deep learning compilers therefore optimize both graphs and operators, including fusion, layout selection, loop structure, and mapping to hardware primitives [12,13]. Similar optimization considerations have also appeared in other artificial intelligence applications. For example, hyperspectral remote sensing image classification has investigated spatial–spectral feature extraction and representation strategies to improve classification performance under complex imaging conditions [14]. Likewise, flight arrival time prediction has explored modular deep learning frameworks and structured prediction pipelines to improve predictive modeling performance [15]. These studies illustrate the broader importance of model structure, feature representation, and processing-pipeline design in practical AI systems. The present work focuses on a different deployment-oriented issue: how compiler-visible graph structure affects execution efficiency on an FPGA-based DPU platform.
Unsupported operators can divide a network across accelerator and central processing unit (CPU) execution domains, introducing tensor transfers and numerical-format conversions at their boundaries. FPGA deployments must therefore be evaluated with compiler partitioning and board measurements, not only with theoretical complexity [16,17,18,19,20,21].
This study examines the channel split in the YOLOv8 Cross-Stage Partial block with two convolutions (C2f) [22]. In the Vitis AI 2.5 compilation path used here, the split was lowered to CPU-side strided_slice operations. Eight C2f blocks in the retained deploy-s-512 model produced sixteen such operations and divided the compiled network into nine deep learning processing unit (DPU) subgraphs and nineteen CPU subgraphs.
C2fDeploy removes this runtime split by partitioning the already trained projection filters and batch-normalization state along the output-channel dimension. It follows the general idea of simplifying an inference graph after training [23,24,25], but it does not introduce a richer training-time block and later fuse it. Both deployment branches are obtained directly from the trained parameters, so no new weights, approximation, or retraining are required. This distinguishes C2fDeploy from architecture-level substitution or approximation of unsupported operators [26].
The central claim of this study is not that C2fDeploy reduces model size or improves fracture-detection accuracy. Instead, it presents the same learned 32-bit floating-point (FP32) function to the deployment compiler in a different operator form. The evaluation therefore follows a causal chain: exact parameter mapping, preservation of detector behavior, removal of CPU-side split subgraphs, and the resulting changes in board-level throughput, energy efficiency, and end-to-end deployment behavior.
We evaluate the method in four stages: FP32 function preservation, 8-bit integer (INT8) quantization behavior, operator-level XModel partitioning, and direct board performance. In addition, an end-to-end pipeline analysis is performed to distinguish accelerator-level acceleration from practical application throughput. The main contributions are as follows:
  • C2fDeploy, a training-free rewrite that replaces the runtime C2f split with an offline partition of the trained convolution and batch-normalization parameters.
  • An exact-arithmetic derivation and a same-checkpoint FP32/INT8 evaluation of function preservation and quantization behavior, including tensor-level error verification at all eight C2f outputs.
  • Xilinx Intermediate Representation (XIR)-based analysis of operator assignments and subgraph partitioning, together with same-checkpoint KV260 measurements of accelerator throughput, energy efficiency,
    DPU/CPU utilization, APM-observed external-memory activity, and complete deployment latency.

2. Materials and Methods

The experiments distinguish detector behavior, compiler partitioning, and board performance. Original C2f and C2fDeploy are compared with the same deploy-s-512 checkpoint. Comparisons between deploy-s-512 and deploy-n-512 describe final operating points and are not treated as an ablation of the rewrite.

2.1. Dataset, Image Processing, and Model Configuration

Experiments used the public GRAZPEDWRI-DX pediatric wrist trauma dataset [4]. The 20,327 images were partitioned at the patient level into 14,204 training images, 4094 validation images, and 2029 test images. Images from the same patient were assigned to only one split, preventing patient-level overlap among the training, validation, and test sets. The resulting patient-level partition is summarized in Table 1. Fracture annotations were mapped to one detection class, denoted as fracture. No new patient data were collected. The radiographs were treated as grayscale intensity images. To match the three-channel input expected by the YOLOv8 implementation and the compiled DPU model, each single-channel image was replicated across three identical channels without pseudo-color conversion. The same image decoding, channel conversion, resizing, and normalization procedure was used for training, validation, test evaluation, post-training quantization (PTQ) calibration, and board inference.
Two model scales were retained. The higher-capacity deploy-s-512 configuration used the depth and width settings of the YOLOv8-s configuration, whereas the lower-capacity deploy-n-512 model used the corresponding YOLOv8-n settings defined in the Ultralytics implementation [22]. The final detector configurations used ReLU in place of the default SiLU activation to match the deployment configuration used in this study. This activation choice was fixed before the controlled Original C2f versus C2fDeploy comparison and was retained throughout training, FP32 evaluation, PTQ, compilation, and board execution. It was therefore part of the common model configuration rather than part of the C2fDeploy transformation. Both models used one fracture class, and no attention module was present in the evaluated FP32, INT8, or compiled models.
The retained DPU-ready training run used 640 × 640 inputs, a maximum of 100 epochs, a batch size of 16, and an early-stopping patience of 15 epochs. The model was initialized from the previously trained best_cloud.pt checkpoint. Optimizer selection was explicitly configured as AdamW rather than being delegated to the Ultralytics automatic optimizer mode. The initial learning rate was 0.001, the final learning-rate fraction was 0.01, and the weight decay was 5 × 10 − 4 . Cosine learning-rate scheduling was enabled.
The recorded training configuration further specified a momentum argument of 0.937, a three-epoch warm-up, a warm-up momentum of 0.8, and a warm-up bias learning rate of 0.1. The detector loss gains were 7.5 for box regression, 0.5 for classification, and 1.5 for distribution focal loss. The random seed was fixed to 0, deterministic execution was enabled, and automatic mixed precision was disabled.
The recorded augmentation configuration used HSV gains of 0.015, 0.7, and 0.4 for hue, saturation, and value, respectively; translation of 0.1; scaling of 0.5; horizontal flipping with a probability of 0.5; and mosaic augmentation with a probability of 0.2. Rotation, shear, perspective transformation, and vertical flipping were disabled. Mosaic augmentation was disabled during the final 10 epochs, while mixup, cutmix, and copy-paste augmentation were also disabled. The validation split was used for training-time monitoring and checkpoint selection.
Export, PTQ, test evaluation, compilation, and board experiments used fixed 512 × 512 inputs with batch size 1 during deployment.
The controlled transformation experiment used the same deploy-s-512 checkpoint for both graph variants. Checkpoint and calibration-list digests are included in the archived reproducibility materials. Because the checkpoint, activation function, image-processing pipeline, evaluation set, and deployment input were held fixed, differences between Original C2f and C2fDeploy can be attributed to the graph rewrite rather than to training or preprocessing variation.
The training, FP32 validation, Vitis AI PTQ/compilation, and KV260 board runtime stages used separate software environments. In particular, the Python 3.13.5/Ultralytics environment used for model training and FP32 validation was not used as the Vitis AI compilation or embedded board-runtime environment. PTQ and XModel compilation were performed using the Vitis AI 2.5 toolchain, whereas board execution used the Vitis AI 2.5.0 runtime/library stack with XRT 2.13.479 and Python 3.10.12. Separating these stage-specific software stacks avoids implying compatibility requirements among packages that were never co-installed or executed in the same environment.
The exact PyTorch build of the original GPU training host was not retained in the archived records and is therefore not assigned an unverified version in Table 2. The separately retained FP32 validation log identifies Python 3.13.5, PyTorch 2.9.0+cpu, and Ultralytics 8.3.240 for that evaluation stage.

2.2. C2fDeploy Rewrite

Vitis AI converts a quantized model into the Xilinx Intermediate Representation (XIR) and partitions the graph according to whether each operation can be executed by the selected DPU [27]. Let
X ∈ R N × C in × H × W
be the input to a C2f block. Its first 1 × 1 projection produces 2 C h output channels:
Z = W ∗ X + b , W ∈ R 2 C h × C in × 1 × 1 ,
followed by batch normalization (BN) and a point-wise activation,
Y = σ BN ( Z ) , Y ∈ R N × 2 C h × H × W .
The activation is divided along the channel dimension:
[ U 0 , U 1 ] = Split C ( Y ) , U 0 , U 1 ∈ R N × C h × H × W .
Let B j ( · ) denote the j-th bottleneck. Starting from V 0 = U 1 ,
V j = B j ( V j − 1 ) , j = 1 , … , m ,
and
F C 2 f ( X ) = CV 2 Concat [ U 0 , U 1 , V 1 , … , V m ] .
In the evaluated compilation path, Equation (4) was lowered to CPU-side strided_slice. C2fDeploy moves the split from the runtime activation graph to an offline partition of the trained parameters, as illustrated in Figure 1.

2.3. Parameter Mapping and Function Preservation

The original projection weight and optional bias are partitioned along the output-channel dimension:
W = W ( 0 ) W ( 1 ) , b = b ( 0 ) b ( 1 ) ,
where
W ( 0 ) , W ( 1 ) ∈ R C h × C in × 1 × 1 .
For inference-mode BN,
BN c ( z c ) = γ c z c − μ c ν c + ϵ + β c ,
where γ c and β c are learned parameters and μ c and ν c are stored running statistics. These channel-wise vectors are partitioned using the same output-channel indices:
γ = [ γ ( 0 ) ; γ ( 1 ) ] , β = [ β ( 0 ) ; β ( 1 ) ] , μ = [ μ ( 0 ) ; μ ( 1 ) ] , ν = [ ν ( 0 ) ; ν ( 1 ) ] .
The transformed projection branches are
U ^ 0 = σ BN ( 0 ) ( W ( 0 ) ∗ X + b ( 0 ) ) , U ^ 1 = σ BN ( 1 ) ( W ( 1 ) ∗ X + b ( 1 ) ) .
Starting from V ^ 0 = U ^ 1 , the original bottleneck sequence and final projection are retained:
V ^ j = B j ( V ^ j − 1 ) , j = 1 , … , m ,
F C 2 fDeploy ( X ) = CV 2 Concat [ U ^ 0 , U ^ 1 , V ^ 1 , … , V ^ m ] .
Proposition 1.
If the convolution parameters, channel-wise BN parameters, and stored BN statistics are partitioned according to Equations (7) and (10), both branches retain the source BN epsilon and inference configuration, and σ acts independently on each tensor element, then
F C 2 fDeploy ( X ) = F C 2 f ( X )
in exact arithmetic.
Proof. 
Let P i , i ∈ { 0 , 1 } , select the corresponding C h output channels. Convolution output channels are determined by independent output filters, so
P i ( W ∗ X + b ) = W ( i ) ∗ X + b ( i ) .
Inference-mode BN acts independently on each channel, and  σ acts independently on each tensor element. Therefore,
P i σ BN ( W ∗ X + b ) = σ BN ( i ) ( P i ( W ∗ X + b ) )
= σ BN ( i ) ( W ( i ) ∗ X + b ( i ) ) = U ^ i .
Thus U ^ 0 = U 0 and U ^ 1 = U 1 . Since each B j and CV 2 is unchanged, induction gives V ^ j = V j for all j, and Equation (14) follows.    □
C2fDeploy does not change the projection parameter count or its theoretical multiply–accumulate count. Ignoring an optional convolution bias,
P Original = 2 C h C in = C h C in + C h C in = P C 2 fDeploy ,
and, for a feature map of size H × W ,
M Original = H W ( 2 C h ) C in = 2 H W C h C in = M C 2 fDeploy .
Any compiler-level difference therefore arises from graph representation rather than pruning or reduced convolutional arithmetic.

2.4. Quantization and XModel Compilation

Each trained C2f module was replaced by a C2fDeploy module with the same input and output channels, hidden width, bottleneck count, shortcut setting, group setting, and expansion ratio. The first C h filters of the trained 1 × 1 projection were assigned to branch 0, and the remaining C h filters were assigned to branch 1. Convolution biases, when present, and the channel-wise BN weight, bias, running mean, and running variance were partitioned using the same output-channel indices. Both branches retained the source BN epsilon and inference configuration. Because num_batches_tracked is a scalar buffer rather than a channel-wise quantity, it was copied to both branches. The bottleneck sequence and final C2f projection were transferred without modification. All eight C2f modules in deploy-s-512 were converted, and no parameter was estimated, optimized, or fine-tuned during the transformation.
FP32 evaluation used the complete held-out test split and the same image decoding, resizing, normalization, confidence threshold, non-maximum suppression (NMS), bounding-box conversion, and metric implementation for both graph variants.
To verify function preservation below the detection-metric level, an additional FP32 tensor comparison was performed using the same checkpoint, 512 × 512 input size, and deterministic image preprocessing. Forward hooks recorded the outputs of each of the eight corresponding C2f and C2fDeploy blocks. For every paired tensor, the maximum absolute error, mean absolute error (MAE), root-mean-square error (RMSE), and relative L 2 error were accumulated over the 2029-image held-out test set. The final raw network output before non-maximum suppression was compared using the same error measures.
PTQ and XModel compilation were performed in a software environment separate from the Python/Ultralytics environment used for model training and FP32 validation. This separation reflects the stage-specific requirements of the Vitis AI 2.5 deployment toolchain and avoids implying that all reported software components were required to coexist in one Python environment.
PTQ was performed with Vitis AI 2.5 [27]. Both graphs used a 1 × 3 × 512 × 512 input and the same ordered set of 200 training images for calibration. No validation or test image was used for calibration. The quantized models were evaluated on the complete test split. The shared checkpoint and calibration set isolate the graph rewrite, but do not imply bit-level INT8 equivalence [28,29,30].
The quantized models were compiled for the B4096 DPU target. XModel metadata were inspected with xdputil, and an XIR Python script recorded the top-level subgraphs and the operation types assigned to each CPU subgraph. To avoid ambiguity in XIR terminology, throughout this paper we use subgraph to denote an XIR structural unit counted by the inspection script. Each reported top-level subgraph is classified according to its device assignment as USER, DPU, or CPU; therefore, the reported total number of top-level subgraphs is the sum of these three categories. The Original C2f XModel contains 29 top-level subgraphs, comprising 1 USER subgraph, 9 DPU subgraphs, and 19 CPU subgraphs, whereas C2fDeploy contains 5 top-level subgraphs, comprising 1 USER subgraph, 1 DPU subgraph, and 3 CPU subgraphs.
A CPU subgraph and a CPU-assigned operation refer to different levels of the XIR structure. A CPU subgraph is a structural unit assigned to CPU execution, whereas a CPU-assigned operation is an individual operation contained within such a subgraph. Consequently, the number of CPU-assigned operations can exceed the number of CPU subgraphs. In the Original C2f XModel, nineteen CPU subgraphs contain 51 CPU-assigned operations, whereas the three CPU subgraphs in C2fDeploy contain three CPU-assigned operations. Throughout the remainder of the manuscript, subgraph is used when referring to a counted XIR structural unit, whereas execution domain is reserved for generic descriptions of accelerator-side and host-side execution. The term region is therefore avoided for XIR structural counts.
Full commands, model digests, and compiler logs are retained with the archived materials.

2.5. Board Measurement Protocols

A direct board comparison was performed with Original C2f and C2fDeploy XModels generated from the same deploy-s-512 checkpoint and calibration list. Both models were executed on the same KV260 and B4096 target [31]. Whole-XModel throughput was measured with xdputil benchmark using complete-graph execution and one worker thread. One smoke run per model was excluded, followed by five formal runs of approximately 60 s in alternating order.

2.6. Runtime Utilization and External-Memory Profiling Protocol

To directly characterize accelerator utilization, host-side CPU activity, and external-memory behavior, two additional profiling experiments were performed using the same Original C2f and C2fDeploy XModels, KV260 board, B4096 DPU configuration, batch size, and single-worker whole-XModel execution protocol used in the formal throughput experiment. Five formal runs were collected for each graph in alternating order. Reported uncertainties correspond to the sample standard deviation across the five independent formal runs rather than temporal variation among samples within a single run.
DPU utilization was measured as the time-weighted busy fraction of the XRT-exposed DPU compute unit. The ZOCL/XRT-reported compute-unit control state was sampled through the sysfs interface at a target interval of 1 ms. Under the XRT-managed kernel control convention, bit 2 of the control/status register corresponds to ap_idle; therefore, samples with ap_idle=0 were classified as non-idle compute-unit states [32]. Let s i ∈ { 0 , 1 } denote whether sample i reports a non-idle compute unit and let Δ t i = t i + 1 − t i . The time-weighted DPU compute-unit busy-time utilization was calculated as
U DPU = ∑ i s i Δ t i ∑ i Δ t i × 100 % .
Only samples collected between the xdputil benchmark 0% and 100% execution markers were included. The achieved mean sampling interval was approximately 1 ms. This metric represents compute-unit busy time and should not be interpreted as internal processing-element or MAC-lane occupancy.
CPU utilization and external-memory activity were collected in a separate synchronized profiling experiment. Per-core CPU utilization was obtained from successive Linux /proc/stat samples at approximately 1 Hz for the four Arm Cortex-A53 cores. The system-wide CPU utilization at each sampling instant was defined as the arithmetic mean of the four per-core utilization values,
U CPU ( t ) = 1 4 ∑ k = 0 3 U k ( t ) ,
and the run-level CPU utilization was the temporal mean over the retained steady-state samples.
External-memory activity was sampled through the AXI Performance Monitor (APM) interface used by the Vitis AI profiling infrastructure [27]. Read and write counters from the five monitored APM ports were sampled at 0.1 s intervals and converted to bandwidth values in MB/s using the same APM profiling interface employed by the Vitis AI tracing infrastructure. Aggregate APM-observed bandwidth was calculated as
B APM = ∑ p = 0 4 B read , p + B write , p .
For the CPU/APM experiment, the formal benchmark interval was identified from the same 0% and 100% execution markers, and 1 s was removed from each end of the interval to exclude startup and shutdown transients. The remaining samples were averaged to obtain one CPU-utilization value and one read, write, and aggregate APM-bandwidth value for each formal run.
Because sustained bandwidth depends on the number of graph executions completed per unit time, an additional normalized traffic metric was calculated as
D APM = B ¯ APM T graph ,
where B ¯ APM is the mean aggregate APM-observed bandwidth in MB/s and T graph is the simultaneously measured whole-XModel throughput in graph executions/s. The resulting unit is MB/execution. This quantity is reported as normalized APM-observed external-memory traffic and is not interpreted as a complete accounting of all system-level DDR traffic.
This benchmark includes CPU fallback operations and format conversions inside the XModel, but excludes image-file input/output, application resizing, detection decoding, NMS, and visualization. It is therefore reported as graph executions per second rather than application frames per second. Reciprocal latency was calculated as 1000 / throughput and is an aggregate equivalent, not a measured p50 or p95 latency.
During each run, the Linux hwmon interface associated with the onboard INA260 power monitor (Texas Instruments, Dallas, TX, USA) was sampled. On the KV260, this telemetry path measures the 5 V V CC _ SOM rail supplying the K26 system-on-module (SOM) [33,34]. For each run,
η graph = T graph P ¯ SOM , E graph = P ¯ SOM T graph ,
where T graph is whole-XModel throughput. Uncertainty is reported as the sample standard deviation across the five formal runs. The reported run-to-run variability does not include systematic uncertainty associated with the onboard INA260 monitor or the telemetry interface. Potential sources include sensor gain and offset error, finite telemetry resolution, temperature dependence, and asynchronous hwmon sampling. Furthermore, because the monitor measures the 5 V V CC _ SOM rail, the reported power represents the complete K26 SOM rather than isolated DPU power.
As an independent external cross-check of the onboard power measurements, an additional input-side measurement was performed at the 12 V supply boundary of the KV260 starter kit. A DC20 inline DC voltage/current/power display (specified measurement ranges: 6–200 V DC and 0–20 A) was inserted in the external power path supplying the board. This supplementary measurement has a different boundary from the onboard INA260: the INA260 measures the 5 V V CC _ SOM rail supplying the K26 SOM, whereas the external instrument measures power at the starter-kit 12 V input. The absolute power values from the two measurement paths are therefore not expected to be equal.
The Original C2f and C2fDeploy XModels were evaluated using the same single-worker whole-XModel benchmark protocol. Five valid runs were retained for each graph. The runs were performed in alternating order, and each benchmark lasted approximately 60 s. External input power was recorded at approximately 10, 20, 30, 40, and 50 s during each run. The five readings within each run were averaged to obtain one external input-power value for that run. One interrupted Original C2f benchmark, which did not complete the 60 s execution interval, was excluded and replaced by a complete repeat.
For each valid run, external input energy per graph execution was calculated as
E ext , r = P ¯ ext , r T r ,
where P ¯ ext , r is the mean externally observed 12-V input power for run r and T r is the corresponding whole-XModel throughput. Reported uncertainties are the sample standard deviation across the five valid runs. Because no manufacturer accuracy specification was available for the external display beyond its stated voltage and current ranges, this measurement is treated as an independent power-trend and energy-efficiency cross-check rather than as a calibrated replacement for the onboard INA260 measurement.
The final deploy-s-512 and deploy-n-512 models were also measured with a custom Vitis AI Runtime (VART)/XIR application. Figure 2 summarizes the measurement boundary.
For timed DPU-stage execution, let L i denote the latency of execution i. Mean latency and throughput were calculated as
L ¯ DPU = 1 N ∑ i = 1 N L i , N = 2500 ,
FPS DPU = 1000 L ¯ DPU ,
where latency is expressed in milliseconds. During the same timed loop, the INA260 interface was sampled at approximately 100 Hz. This sampling rate characterizes average active power over the timed interval but may not capture short-duration power transients. Let P j denote power sample j. Mean active SOM-rail power was
P ¯ SOM = 1 M ∑ j = 1 M P j ,
where M is the number of power samples collected during the active interval. The derived throughput-efficiency and energy metrics were
η DPU = FPS DPU P ¯ SOM ,
E DPU = P ¯ SOM FPS DPU .
To evaluate complete application-level performance, a separate matched end-to-end comparison was performed between the Original C2f and C2fDeploy deploy-s-512 XModels generated from the same checkpoint. Both XModels were executed on the same Kria KV260 with a fixed 512 × 512 input size, batch size 1, the same radiograph, and the same GraphRunner-based application pipeline.
For each XModel, five complete-pipeline warm-up executions were excluded, followed by 50 timed repetitions. Each timed repetition included image loading, host preprocessing, input tensor handoff, complete XModel execution, output dequantization/copy, YOLO decoding, and non-maximum suppression. Visualization and result-file writing were excluded from the timed interval. Stage latencies were recorded separately, while total end-to-end latency was measured directly over the complete included pipeline.
Because the same radiograph was repeatedly accessed during this controlled comparison, the image-loading measurements may be influenced by the Linux page cache and should not be interpreted as a storage-device bandwidth benchmark. The purpose of this experiment was to compare Original C2f and C2fDeploy under identical application-level conditions rather than to characterize storage system performance.

2.7. Reproducibility and Reporting Boundaries

Archived materials include the training configuration, conversion script, model and calibration digests, quantization and compiler settings, XModels, XIR reports, benchmark logs, application-profile records, synchronized power measurements, per-run CPU-utilization records, raw APM samples, DPU compute-unit state samples, external input-power observations, corresponding whole-XModel benchmark logs, and the associated five-run profiling summaries. These materials are available from the corresponding author on reasonable request during peer review and will be publicly archived before publication.
Compiler reports establish subgraph partitioning and CPU-assigned operation counts but do not themselves measure runtime utilization or external-memory traffic. Accordingly, CPU utilization, DPU compute-unit busy time, and APM-observed external-memory activity were measured separately using runtime telemetry. DPU utilization denotes the time-weighted fraction for which the XRT-exposed DPU compute unit reported ap_idle=0; it does not represent internal PE- or MAC-lane occupancy. Similarly, APM bandwidth characterizes traffic observed on the monitored accelerator memory interfaces and is not interpreted as a complete measurement of all DDR traffic generated by the processing system.
The xdputil whole-XModel benchmark, custom VART DPU-stage measurements, GraphRunner complete end-to-end measurements, CPU/APM profiling experiment, and DPU-state profiling experiment use different measurement boundaries and are therefore reported separately. Results obtained under one timing boundary are not combined with measurements from another boundary to construct synthetic latency values. The utilization and memory measurements are used to explain runtime behavior, whereas the primary throughput and energy values remain those obtained from the independent formal whole-XModel benchmark.

3. Results

3.1. FP32 Preservation and INT8 Quantization

Table 3 compares Original C2f and C2fDeploy on the complete test set using the same deploy-s-512 checkpoint. The FP32 graphs produced identical precision, recall, mAP50, and mAP50–95. After PTQ, the small differences were mixed in direction and did not indicate an accuracy gain or a meaningful loss.
The tensor-level comparison confirmed that the implemented rewrite remained numerically equivalent within FP32 finite-precision effects. Across the 2029-image held-out test set and all eight C2f outputs, the maximum absolute difference was 1.43 × 10 − 5 , with an aggregate MAE of 9.73 × 10 − 8 , RMSE of 2.40 × 10 − 7 , and relative L 2 error of 1.97 × 10 − 7 . At the final raw network output, the maximum absolute difference was 5.65 × 10 − 4 , while the MAE, RMSE, and relative L 2 error were 6.42 × 10 − 6 , 1.44 × 10 − 5 , and 7.31 × 10 − 8 , respectively. The larger maximum raw-output difference reflects downstream accumulation of very small floating-point differences, whereas the relative error remained negligible. The block-wise and aggregate tensor-level comparison results are summarized in Table 4.
The identical FP32 values across all four aggregate metrics are consistent with the exact-arithmetic derivation and indicate that the implemented parameter transfer preserved detector behavior at the reported evaluation level. These values were obtained from independent evaluations of the two graph variants; their equality was observed rather than imposed. After PTQ, C2fDeploy differed from Original C2f by + 0.003726 in precision, − 0.001691 in recall, + 0.000522 in mAP50, and  + 0.000629 in mAP50–95. The changes are small and mixed in direction; they are therefore interpreted as quantization variation rather than an accuracy improvement. Relative to the common FP32 mAP50–95 value, the decreases were 0.005551 for Original C2f and 0.004922 for C2fDeploy.

3.2. Compiled-Graph Partitioning

The Original C2f XModel contained twenty-nine top-level subgraphs, comprising one USER subgraph, nine DPU subgraphs, and nineteen CPU subgraphs. The nineteen CPU subgraphs contained sixteen strided_slice, sixteen float2fix, and nineteen fix2float operations. Figure 3 illustrates the change, and Table 5 reports the corresponding XIR counts.

Example of C2f Block Partition Behavior

To further clarify how runtime channel splitting causes graph fragmentation, we provide a representative dataflow walk-through of a single C2f block. As illustrated in Figure 3, the original C2f block first performs a projection convolution on the DPU and generates an intermediate activation tensor. Using the tensor dimension order reported by XIR, the intermediate activation before channel separation is represented in NHWC layout as Y ∈ R 1 × 512 × 512 × 128 . Since runtime channel splitting is not directly supported by the DPU execution path in the adopted Vitis AI compilation flow, the compiler maps this operation to a CPU-side strided_slice operation.
Consequently, the activation tensor must cross the DPU–CPU boundary. The execution sequence can be summarized as
Y → fix 2 float Y CPU → strided _ slice ( U 0 , U 1 ) → float 2 fix ( U 0 q , U 1 q ) ,
where the channel split is executed in the CPU execution domain rather than as part of a continuous DPU execution path. The intermediate branches then return to the DPU for subsequent convolutional processing. When multiple C2f blocks are stacked in YOLOv8, these repeated domain transitions accumulate and produce a fragmented XIR structure containing multiple DPU and CPU subgraphs.
In contrast, C2fDeploy moves the channel partition from runtime activation processing to offline parameter transformation. The two projection branches are generated directly from the partitioned convolution parameters and remain inside the accelerator-supported execution path:
X → Conv ( 0 ) , Conv ( 1 ) ( U ^ 0 , U ^ 1 ) .
Here, ( U ^ 0 , U ^ 1 ) denotes the intermediate branch tensors that replace the runtime channel split; it does not denote the output of the complete C2fDeploy block. The downstream bottleneck sequence and final projection remain unchanged, as defined in Equation (13). Therefore, C2fDeploy removes the need for runtime strided_slice and the associated tensor-format conversions between the DPU and CPU execution domains. This split-replacement mechanism explains why the fragmented XIR subgraph structure observed in the Original C2f model is consolidated after deployment.
C2fDeploy removed all sixteen C2f-related slicing operations and all sixteen associated float2fix boundaries. The number of top-level subgraphs decreased from 29 to 5, the DPU subgraphs were consolidated from 9 to 1, and CPU-assigned operations decreased from 51 to 3. The three remaining CPU subgraphs each contain one output-side fix2float operation and are unrelated to the original C2f split. Because model parameters and theoretical convolutional arithmetic were unchanged, the difference is attributable to operator representation and compiler partitioning rather than pruning or reduced model arithmetic.

3.3. Same-Checkpoint Whole-XModel Performance

Table 6 reports the direct Original-versus-Deploy board comparison. C2fDeploy increased whole-XModel throughput by 95.2-fold and reduced the reciprocal latency equivalent by 98.95%. Active SOM-rail power was higher, but energy per graph execution was 98.20% lower because execution completed much sooner.
C2fDeploy increased whole-XModel throughput from 0.5158 ± 0.0001 to 49.13 ± 0.04 graph executions/s, corresponding to a 95.2 × improvement. The reciprocal latency equivalent decreased from 1938.7 ± 0.2 to 20.36 ± 0.02 ms/execution. Because both XModels were generated from the same checkpoint and evaluated with the same board, DPU target, thread count, and benchmark mode, the observed difference isolates the effect of the compiled-graph rewrite. This result is consistent with the XIR analysis, which showed the removal of repeated CPU-side slicing, numerical-format conversions, and accelerator–host transitions.
Mean active V CC _ SOM power increased from 5.087 ± 0.003 to 8.735 ± 0.003 W. This increase is consistent with the direct runtime-utilization measurements reported in Section 3.4, where DPU compute-unit busy-time utilization increased from 0.3562 ± 0.0092 % for Original C2f to 90.7465 ± 0.1356 % for C2fDeploy. Despite the higher active power, the shorter execution time reduced SOM-rail energy per graph execution from 9.862 ± 0.006 to 0.1778 ± 0.0002 J, a 98.2% reduction. These values characterize energy delivered through the 5 V V CC _ SOM rail during the whole-XModel benchmark; they do not represent isolated DPU dynamic energy or total 12 V starter-kit input energy.
To independently cross-check the direction and energy-efficiency implication of the onboard INA260 measurements, power was also observed at the external 12 V starter-kit input during a separate five-run whole-XModel experiment. Table 7 summarizes the results. Because this measurement includes the complete starter-kit input boundary rather than only the 5 V V CC _ SOM rail, its absolute power values are higher and should not be directly equated with the INA260 values.
The independent input-side measurement reproduced the same qualitative power behavior observed with the onboard INA260. C2fDeploy increased external starter-kit input power from 8.64 ± 0.05 to 12.50 ± 0.05 W as accelerator execution became substantially more continuous. However, the corresponding approximately 95.05 × throughput increase reduced external-input energy per graph execution from 16.76 ± 0.11 to 0.255 ± 0.001 J, corresponding to a 98.48% reduction.
The external measurement and the INA260 measurement have different electrical boundaries and are therefore not expected to yield identical absolute power values. Nevertheless, both independently show higher active power for C2fDeploy together with a large reduction in energy per completed graph execution. The external result therefore provides a supplementary system-input-level cross-check of the power and energy trends obtained from the onboard SOM-rail telemetry.

3.4. Runtime Utilization and External-Memory Profiling

Direct runtime profiling revealed a substantial redistribution of execution activity between the DPU and host CPU after the graph rewrite. Table 8 summarizes the DPU compute-unit busy-time utilization, system-wide CPU utilization, and APM-observed external-memory behavior. Values are reported as the mean ± sample standard deviation across five independent formal runs.
The Original C2f graph exhibited a DPU compute-unit busy-time utilization of only 0.3562 ± 0.0092 % . In contrast, C2fDeploy increased the measured busy-time utilization to 90.7465 ± 0.1356 % , corresponding to an absolute increase of 90.3903 percentage points. During the DPU-state profiling experiment, the throughput remained at approximately 0.5160 graph executions/s for Original C2f and 48.9673 ± 0.0257 graph executions/s for C2fDeploy, closely matching the independent formal throughput measurements. Thus, the 1 ms status sampling did not materially perturb the execution behavior.
The CPU measurements showed the complementary effect. System-wide CPU utilization decreased from 25.3365 ± 0.0187 % for Original C2f to 2.6536 ± 0.0318 % for C2fDeploy, an 89.53% reduction. Because the reported system-wide value is normalized across the four Arm Cortex-A53 cores, the Original C2f utilization corresponds to approximately one core-equivalent of aggregate CPU capacity. Per-core traces typically showed the workload concentrated on one core, with occasional migration across cores. This behavior is consistent with the CPU-assigned strided_slice and numerical-format-conversion operations identified by the XIR analysis. After the rewrite, these CPU-side operations were largely removed and only low host-side activity remained during accelerator execution.
The external-memory results show a different but complementary trend. Aggregate APM-observed bandwidth increased from 40.2036 ± 0.0963 MB/s for Original C2f to 3250.5641 ± 2.0062 MB/s for C2fDeploy. The increase does not indicate a larger memory cost per completed graph execution. Instead, the Original graph spends most wall-clock time in fragmented CPU-side execution and DPU–CPU transitions, leaving the accelerator memory path intermittently active. C2fDeploy keeps the DPU-supported path active substantially more continuously, thereby increasing sustained memory-interface activity per unit time.
After normalization by simultaneously measured whole-XModel throughput, aggregate APM-observed traffic decreased from 77.8675 ± 0.1788 MB/execution to 66.3395 ± 0.0169  MB/execution, corresponding to a 14.80% reduction. Therefore, C2fDeploy simultaneously increases sustained accelerator-side memory activity while reducing the APM-observed external-memory traffic associated with each completed graph execution.
Taken together, the utilization and bandwidth measurements provide direct runtime evidence for the execution mechanism inferred from the XIR structure. The Original graph combines very low sustained DPU activity with substantial host CPU involvement, whereas C2fDeploy shifts execution toward sustained DPU operation, substantially reduces CPU utilization, and decreases normalized external-memory traffic per completed graph execution.

3.5. Final Operating Points and Application Profile

The final C2fDeploy models provide two operating points (Table 9). Deploy-s-512 retains the higher detection accuracy, whereas deploy-n-512 provides a lower-latency deployment configuration. Both models benefit from the same graph rewriting strategy and therefore maintain the compiler-friendly execution structure introduced by C2fDeploy.
Deploy-s-512 achieved 46.542 FPS under the DPU-stage measurement boundary, while deploy-n-512 further reduced the execution latency to 7.857 ms and increased throughput to 127.280 FPS. These values characterize the timed accelerator-stage execution loop and therefore exclude image-file access, complete host-side preprocessing, detection decoding, NMS, and visualization. They represent accelerator capability after removing compiler-induced graph fragmentation rather than complete application throughput.
To provide a controlled application-level comparison, a matched complete end-to-end GraphRunner profile was performed for Original C2f and C2fDeploy using the same deploy-s-512 checkpoint, KV260 platform, input radiograph, 512 × 512 input size, batch size 1, preprocessing pipeline, and post-processing implementation. The timed interval included image loading, host preprocessing, input tensor handoff, complete XModel execution, output dequantization/copy, YOLO decoding, and NMS.
Table 10 summarizes the mean stage-wise latency across 50 timed complete-pipeline repetitions after five excluded warm-up executions. The Original C2f pipeline required 2045.923 ms per image, corresponding to 0.489 FPS, whereas C2fDeploy required 86.410 ms, corresponding to 11.573 FPS. The resulting complete end-to-end speedup was approximately 23.7 × .
Within the same GraphRunner experiment, whole-XModel execution decreased from 1970.381 ms for Original C2f to 20.636 ms for C2fDeploy, corresponding to a 95.5 × XModel-stage speedup. This independently corroborates the 95.2 × whole-XModel throughput improvement obtained from the separate five-run xdputil benchmark, while using a different runtime measurement path.
The reported total end-to-end latency was measured directly over the complete included pipeline. Because the individual stage means in Table 10 are rounded independently, their displayed sum may differ from the directly measured total by 0.001 ms.
Post-processing, defined as YOLO decoding plus NMS, required 4.628 ms for Original C2f and 3.899 ms for C2fDeploy. These values are derived from the corresponding decoding and NMS rows and are therefore not added separately to the total end-to-end latency.
Across the 50 timed complete-pipeline repetitions, total end-to-end latency for Original C2f had a standard deviation of 10.627 ms, a median of 2049.017 ms, and a 95th percentile (p95) of 2057.323 ms. The corresponding C2fDeploy values were 0.739 ms, 86.181 ms, and 87.786 ms, respectively.
The matched comparison shows that the complete application-level speedup is smaller than the graph-execution speedup because image loading, preprocessing, tensor handoff, output conversion, detection decoding, and NMS remain outside the C2fDeploy graph rewrite. Nevertheless, after these preprocessing and post-processing stages are included, C2fDeploy retains a substantial 23.7 × end-to-end speedup over Original C2f.

4. Discussion

4.1. Effect of Operator Representation on Compilation

C2fDeploy neither prunes the network nor reduces its theoretical convolutional arithmetic. Under the same checkpoint, calibration set, compiler, and DPU target, the controlled change is the representation of the first C2f projection and split. Replacing the runtime split with two parameter-derived projection branches removed the repeated CPU subgraphs associated with slicing and format conversion and consolidated DPU-supported computation into one DPU subgraph.
The direct runtime measurements provide additional evidence for the mechanism suggested by the XIR structure. Original C2f achieved a DPU compute-unit busy-time utilization of only 0.3562 ± 0.0092 % , while its system-wide CPU utilization was 25.3365 ± 0.0187 % . C2fDeploy reversed this execution distribution: DPU busy-time utilization increased to 90.7465 ± 0.1356 % , while CPU utilization decreased to 2.6536 ± 0.0318 % . These complementary measurements indicate that the performance difference is not merely associated with a different number of compiled subgraphs; it corresponds to a fundamental change from fragmented host-assisted execution to sustained accelerator execution.
The external-memory measurements are consistent with the same interpretation. C2fDeploy increased aggregate APM-observed bandwidth from 40.20 to 3250.56 MB/s because the DPU and its memory path remained active for a much larger fraction of wall-clock time. However, normalized APM-observed traffic decreased from 77.87 to 66.34  MB/execution. Thus, the higher sustained bandwidth reflects greater accelerator activity rather than an increase in external-memory traffic per completed graph execution.
This result illustrates why parameter count and floating-point operation count are incomplete deployment indicators. Two graphs can represent the same FP32 function and require the same theoretical convolutional arithmetic, yet they produce different compiler partitions, host workloads, accelerator busy times, memory-interface behavior, and board-level performance.

4.2. Comparison with Source-Level C2f Replacement and Alternative Deployment Strategies

Although C2fDeploy achieves substantial acceleration by removing compiler-induced graph fragmentation, similar deployment problems can also be addressed through alternative strategies. These approaches differ from C2fDeploy in terms of modification scope, engineering effort, and preservation of the original model function.
A straightforward alternative is to modify the original YOLOv8 C2f implementation at the source-code level by replacing the runtime channel split with two independent convolutional branches. If these branches are initialized using the same exact partition of the trained convolution and batch-normalization parameters, such a source-level implementation can also preserve the original function without retraining. By contrast, an independently initialized source-level redesign would require training and re-validation. The practical distinction of C2fDeploy is that it performs this exact transformation systematically after training, starting from a standard trained C2f checkpoint, without requiring modification of the training-time model definition or training workflow. This allows the original training pipeline and checkpoint format to be retained while generating a compiler-friendly deployment graph.
Another possible solution is to implement unsupported operations using hardware-specific extensions, such as Vitis AI custom layers or fused-layer mechanisms. These methods provide maximum flexibility because dedicated kernels or compiler extensions can explicitly define the execution behavior of new operators. However, they usually require additional hardware–software co-design effort, including custom kernel development, compiler integration, runtime maintenance, and platform-specific optimization. C2fDeploy avoids these requirements by rewriting the graph using existing compiler-supported convolutional operators while preserving the original computation.
C2fDeploy is also different from approximation-based optimization methods, such as the DPU-aware attention approximation approach proposed by Karki et al. [26]. These methods improve accelerator efficiency by reducing the computational cost of attention-related operations through approximation. Although such strategies can provide significant hardware benefits, they may introduce accuracy–efficiency trade-offs because the original computation is modified. In contrast, C2fDeploy focuses on deployment-oriented graph rewriting and preserves the original learned function through mathematically equivalent parameter partitioning.
Table 11 summarizes the differences among these strategies. Compared with source-level modification and custom hardware extensions, C2fDeploy requires lower engineering effort while maintaining exact functional equivalence. Compared with approximation-based optimization, C2fDeploy achieves acceleration without modifying the learned representation or introducing an approximation error.

4.3. Generalizability and Applicability Boundaries of C2fDeploy

Although C2fDeploy was validated using the Vitis AI 2.5 compilation flow and the DPUCZDX8G architecture on the Kria KV260 platform, the underlying principle is not limited to a single FPGA device. The key requirement of C2fDeploy is that the compiler generates inefficient host-side execution for channel splitting operations while the subsequent branch computations remain independent and can be represented by separated parameterized paths. Therefore, the effectiveness of C2fDeploy depends on the interaction between compiler behavior, target accelerator architecture, and graph representation rather than on the YOLOv8 model alone.
For newer Vitis AI versions or different DPU configurations, such as B512, B1024, or B4096-based platforms, the deployment benefit may vary. If future compiler versions provide optimized hardware implementations for channel splitting operators, the performance improvement obtained from C2fDeploy may be reduced. However, the proposed parameter transformation remains applicable as a graph-level deployment optimization because it does not depend on a specific DPU instruction set or hardware primitive.
Beyond YOLOv8 C2f blocks, the partition-copy strategy can potentially be extended to other split-based architectures, including CSP-style networks, other YOLO variants, and residual bottleneck structures containing channel partition operations. Such migration requires that the split branches are functionally independent and that the original parameters can be equivalently mapped to separated computational paths. Architectures with strong cross-branch interactions, dynamic routing, or data-dependent partitioning may require additional modifications and cannot be directly transformed.
Therefore, C2fDeploy should be considered a deployment-oriented graph rewriting principle rather than a model-specific optimization. Its applicability is determined by three conditions: (1) the existence of compiler-unfriendly split operations, (2) the availability of mathematically equivalent branch decomposition, and (3) the preservation of model semantics after offline parameter transformation.

4.4. Hardware and Application-Level Implications

Although C2fDeploy achieves a 95.2 × improvement in the whole-XModel xdputil benchmark, this improvement should be interpreted within the corresponding measurement boundary. The whole-XModel benchmark evaluates compiled graph execution, including DPU computation, CPU-assigned graph operations, and the associated tensor-format conversions, but it excludes application-level operations such as image-file access, radiograph decoding, preprocessing, detection decoding, NMS, and visualization. Therefore, the measured 95.2 × acceleration reflects the removal of compiler-induced execution fragmentation rather than an equivalent acceleration of every stage of the complete medical imaging workflow.
The direct utilization measurements provide a hardware-level explanation for this graph-execution speedup. During formal profiling, Original C2f kept the DPU compute unit non-idle for only 0.3562 ± 0.0092 % of the benchmark interval while system-wide CPU utilization reached 25.3365 ± 0.0187 % . C2fDeploy increased DPU compute-unit busy-time utilization to 90.7465 ± 0.1356 % and reduced CPU utilization to 2.6536 ± 0.0318 % . The graph rewrite therefore changes not only the number of XIR subgraphs but also the measured runtime balance between host and accelerator execution.
The matched GraphRunner comparison in Table 10 provides a direct measure of how this graph-level acceleration propagates to the complete application pipeline. Under the same deploy-s-512 checkpoint, KV260 platform, input radiograph, and pre/post-processing implementation, total end-to-end latency decreased from 2045.923 ms for Original C2f to 86.410 ms for C2fDeploy. This corresponds to an approximately 23.7 × complete end-to-end speedup.
Within the same GraphRunner experiment, whole-XModel execution decreased from 1970.381 ms to 20.636 ms, corresponding to a 95.5 × XModel-stage speedup. This independently corroborates the 95.2 × improvement observed in the formal five-run xdputil benchmark. Agreement between these two independent runtime measurements supports the conclusion that the dominant acceleration originates from removal of compiler-induced graph fragmentation.
The matched profile also demonstrates the resulting bottleneck shift. In the Original C2f pipeline, whole-XModel execution accounted for approximately 96.3% of the measured total end-to-end latency. After the rewrite, XModel execution accounted for approximately 23.9% of the C2fDeploy end-to-end latency, while approximately 76.1% occurred outside XModel execution. Consequently, the benefit of the approximately 95.5 × XModel-stage speedup is diluted to approximately 23.7 × at the complete application level because the surrounding host-side stages are not accelerated by the graph rewrite.
The remaining host-side latency arises from several stages of the practical medical imaging pipeline. Radiographs must be loaded from storage, decoded and preprocessed into the tensor representation required by the detector, transferred to the runtime input buffer, and subsequently processed through output conversion, YOLO decoding, and NMS. In the controlled GraphRunner experiment, the same radiograph was repeatedly accessed; therefore, image-loading latency may be influenced by the Linux page cache and should not be interpreted as a storage-device bandwidth measurement.
Several system-level optimizations could reduce this remaining bottleneck. First, asynchronous pipeline execution can overlap image acquisition or loading, CPU preprocessing, and accelerator inference by using multiple or double-buffered input buffers. Second, prefetching or preloading can reduce repeated storage-access latency when a sequence of radiographs is processed. Third, persistent buffer allocation and buffer reuse can reduce repeated host-side memory allocation and copy overhead. Finally, where supported by the runtime and memory architecture, zero-copy or reduced-copy data paths could decrease unnecessary movement between host-side preprocessing buffers and accelerator-accessible memory. These optimizations are complementary to C2fDeploy because they target the surrounding application pipeline rather than the neural-network graph itself.
From a practical PACS perspective, C2fDeploy therefore remains valuable even though the complete application does not inherit the full 95.2 × – 95.5 × graph-execution acceleration. The directly measured complete-pipeline improvement was approximately 23.7 × . C2fDeploy removes the dominant compiler-induced accelerator bottleneck, after which image access, preprocessing, output conversion, and post-processing become a substantially larger fraction of the remaining latency. C2fDeploy should therefore be viewed as an accelerator-side foundation for efficient medical edge inference rather than as a standalone solution to every source of end-to-end latency.

4.5. Function Preservation and Quantization

The function-preservation claim is supported jointly by the exact parameter mapping, the identical FP32 detection metrics, and the tensor-level comparison at all eight rewritten C2f outputs. The aggregate C2f-output MAE was 9.73 × 10 − 8 and the relative L 2 error was 1.97 × 10 − 7 , while the final raw network output retained a relative L 2 error of only 7.31 × 10 − 8 . These small nonzero differences are consistent with finite-precision FP32 execution rather than a change in the represented function. Selecting channel groups after the original projection is equivalent, in exact arithmetic, to applying the corresponding partitioned projection and BN state before the split. The downstream bottleneck sequence and final projection are unchanged.
The INT8 results are not bit-identical because the two branches can receive different activation scales and finite-precision kernels. The small, mixed-direction metric differences therefore do not indicate an accuracy gain, but they show that the rewrite did not introduce a meaningful loss under the shared calibration protocol.

4.6. Relation to Existing Re-Parameterization and DPU-Aware Rewrites

C2fDeploy is related to structural re-parameterization methods such as RepVGG, Diverse Branch Block, and OREPA, but its objective and transformation stage are different. These methods generally introduce richer structures during training and convert them into simpler inference-time forms. C2fDeploy instead starts from an already trained C2f block and partitions its existing projection and BN state to remove a compiler-visible runtime split. No additional training-time branch is introduced, and no branch is learned specifically for deployment.
The method also differs from DPU-aware architecture substitution and operator approximation. Both C2fDeploy branches are obtained directly from the trained parameters; no unsupported operator is approximated, and no newly initialized projection is trained. The contribution is therefore not a new detector architecture, but an exact change in how the learned detector is represented to the deployment compiler. Linking this same-checkpoint rewrite to operator-level XIR evidence and direct board measurements distinguishes the present study from comparisons based only on different model variants or theoretical complexity.

4.7. Limitations

Several limitations should be considered when interpreting the reported results.
First, this study evaluates C2fDeploy using the Vitis AI 2.5 compilation flow, the DPUCZDX8G architecture, and the Kria KV260 platform. Although the proposed rewrite principle is not fundamentally restricted to this environment, the measured acceleration depends on how a specific compiler version maps channel partition operations and unsupported operators. Therefore, future evaluation with newer Vitis AI releases and different DPU configurations is required to fully characterize the portability of the method.
Second, the reported DPU utilization represents compute-unit busy time rather than internal arithmetic occupancy. Specifically, utilization was calculated from the time-weighted fraction for which the XRT-exposed DPU compute unit reported ap_idle=0. This directly measures whether the DPU compute unit is idle or non-idle over the benchmark interval, but it does not resolve the instantaneous occupancy of individual processing elements, MAC lanes, or internal execution pipelines. Consequently, the reported 90.7465 ± 0.1356 % C2fDeploy value should be interpreted as DPU compute-unit busy-time utilization rather than PE- or MAC-level utilization.
Third, the reported memory-bandwidth values characterize the accelerator-side interfaces observed by the platform AXI Performance Monitor. They therefore quantify APM-observed external-memory activity rather than all traffic on the complete system DDR subsystem. The normalized MB/execution quantity is derived from mean APM-observed bandwidth divided by simultaneously measured graph throughput and should likewise be interpreted as normalized monitored traffic, not as a complete byte-accurate accounting of every system memory transaction.
Fourth, CPU utilization was reported as a system-wide value normalized across the four Arm Cortex-A53 cores. This definition facilitates comparison between the two graph variants, but it can obscure the distribution of work among individual cores. Per-core records showed that the Original C2f workload typically consumed approximately one core-equivalent of aggregate CPU capacity, although the active workload could migrate between cores.
Fifth, the external 12 V input-power measurement was introduced as an independent cross-check rather than as a traceable calibration of the onboard INA260. The external DC display provided voltage, current, and power readings and had specified operating ranges of 6–200 V DC and 0–20 A, but a manufacturer accuracy specification was not available. In addition, the external readings were sampled manually at five steady-state time points per run rather than recorded continuously. Consequently, the external results are used to verify the direction and approximate magnitude of the observed power and energy-efficiency changes, while the onboard INA260 remains the primary high-rate power measurement used for the main SOM-rail results. The two measurement paths also have different electrical boundaries: the external instrument measures the 12 V starter-kit input, whereas the onboard INA260 measures the 5 V V CC _ SOM rail. Their absolute power values should therefore not be interpreted as directly interchangeable.
Sixth, the current evaluation focuses on YOLOv8 C2f structures. Although the partition-copy principle can potentially be extended to other split-based architectures, successful migration requires that the split branches remain functionally independent and that the trained parameters can be equivalently decomposed. Networks containing dynamic routing or strong cross-branch dependencies may require additional modifications.
Finally, C2fDeploy targets deployment optimization after model training. Therefore, it complements rather than replaces broader system-level optimization approaches, including data-pipeline acceleration, memory optimization, and runtime scheduling strategies required for complete medical AI deployment.

5. Conclusions

C2fDeploy removes the YOLOv8 C2f runtime split by partitioning the trained projection and batch-normalization parameters, preserving the detector’s FP32 function without retraining or introducing additional parameters. On the Kria KV260 platform using the Vitis AI 2.5 compilation flow, the proposed rewrite eliminated the C2f-related CPU slicing operations, consolidated DPU computation into one subgraph, and increased same-checkpoint whole-XModel throughput by 95.2 × in the formal xdputil benchmark while reducing SOM-rail energy per execution by 98.2%.
An independent 12 V starter-kit input-power cross-check reproduced the same energy-efficiency trend, yielding an approximately 98.5% reduction in external-input energy per graph execution.
Direct runtime profiling further showed that the graph rewrite substantially changed hardware utilization. DPU compute-unit busy-time utilization increased from 0.3562 ± 0.0092 % for Original C2f to 90.7465 ± 0.1356 % for C2fDeploy, while system-wide CPU utilization decreased from 25.3365 ± 0.0187 % to 2.6536 ± 0.0318 % . Aggregate APM-observed external-memory bandwidth increased from 40.20 ± 0.10 to 3250.56 ± 2.01 MB/s as accelerator execution became substantially more continuous. After normalization by whole-XModel throughput, however, APM-observed traffic decreased from 77.87 ± 0.18 to 66.34 ± 0.02 MB/execution, indicating that the higher sustained bandwidth was accompanied by lower monitored memory traffic per completed graph execution.
An independent matched GraphRunner comparison produced a 95.5 × XModel-stage speedup, closely corroborating the formal whole-XModel benchmark. When image loading, host preprocessing, tensor handoff, complete XModel execution, output dequantization/copy, YOLO decoding, and NMS were all included, complete end-to-end latency decreased from 2045.923 ms for Original C2f to 86.410 ms for C2fDeploy, corresponding to an approximately 23.7 × application-level speedup. The resulting difference between graph-level and end-to-end acceleration demonstrates that, once compiler-induced graph fragmentation is removed, surrounding host-side stages constitute a substantially larger fraction of practical deployment latency.
Although validated on YOLOv8 and the KV260 platform, C2fDeploy represents a general deployment-oriented graph rewriting principle for split-based neural architectures. Its applicability depends on whether compiler-unfriendly partition operations can be transformed into equivalent parameterized execution paths while preserving model semantics. Therefore, C2fDeploy complements rather than replaces model architecture optimization and system-level pipeline optimization, and it provides a practical approach for improving accelerator-friendly representations in FPGA-based edge AI deployment.

Author Contributions

Conceptualization, S.H., X.J. and H.W.; methodology, S.H., X.J. and X.L.; software, S.H.; validation, S.H. and W.H.; formal analysis, S.H.; investigation, S.H.; resources, X.J., W.H. and H.W.; data curation, S.H.; writing—original draft preparation, S.H.; writing—review and editing, X.J., H.W., W.H. and X.L.; visualization, S.H.; supervision, X.J. and H.W.; project administration, X.J. and H.W.; funding acquisition, H.W. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Talent Introduction Project of Xihua University, grant number ZX20250131.

Institutional Review Board Statement

Not applicable. This study used a publicly available, de-identified dataset and did not collect new patient data.

Informed Consent Statement

Not applicable.

Data Availability Statement

The GRAZPEDWRI-DX dataset is publicly available through Figshare at https://doi.org/10.6084/m9.figshare.14825193. Derived artifacts supporting this study are available from the corresponding author upon reasonable request. A public archival package will be deposited before publication.

Conflicts of Interest

The authors declare no conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
APMAXI performance monitor
BNBatch normalization
C2fCross-stage partial block with two convolutions
CPUCentral processing unit
CUCompute unit
DPUDeep learning processing unit
FP3232-bit floating point
FPGAField-programmable gate array
INT88-bit integer
MACMultiply–accumulate
NMSNon-maximum suppression
PEProcessing element
PLProgrammable logic
PSProcessing system
PTQPost-training quantization
SOMSystem-on-module

References

  1. Kuo, R.Y.L.; Harrison, C.; Curran, T.-A.; Jones, B.; Freethy, A.; Cussons, D.; Stewart, M.; Collins, G.S.; Furniss, D. Artificial Intelligence in Fracture Detection: A Systematic Review and Meta-Analysis. Radiology 2022, 304, 50–62. [Google Scholar] [CrossRef] [Scilit]
  2. Raisuddin, A.M.; Vaattovaara, E.; Nevalainen, M.; Nikki, M.; Järvenpää, E.; Makkonen, K.; Pinola, P.; Palsio, T.; Niemensivu, A.; Tervonen, O.; et al. Critical Evaluation of Deep Neural Networks for Wrist Fracture Detection. Sci. Rep. 2021, 11, 6006. [Google Scholar] [CrossRef] [Scilit]
  3. Hardalaç, F.; Uysal, F.; Peker, O.; Çiçeklidağ, M.; Tolunay, T.; Tokgöz, N.; Kutbay, U.; Demirciler, B.; Mert, F. Fracture Detection in Wrist X-ray Images Using Deep Learning-Based Object Detection Models. Sensors 2022, 22, 1285. [Google Scholar] [CrossRef] [Scilit]
  4. Nagy, E.; Janisch, M.; Hržić, F.; Sorantin, E.; Tschauner, S. A Pediatric Wrist Trauma X-ray Dataset (GRAZPEDWRI-DX) for Machine Learning. Sci. Data 2022, 9, 222. [Google Scholar] [CrossRef] [Scilit]
  5. Ju, R.-Y.; Cai, W. Fracture Detection in Pediatric Wrist Trauma X-ray Images Using YOLOv8 Algorithm. Sci. Rep. 2023, 13, 20077. [Google Scholar] [CrossRef] [Scilit]
  6. Ahmed, A.; Imran, A.S.; Manaf, A.; Kastrati, Z.; Daudpota, S.M. Enhancing Wrist Abnormality Detection with YOLO: Analysis of State-of-the-Art Single-Stage Detection Models. Biomed. Signal Process. Control 2024, 93, 106144. [Google Scholar] [CrossRef] [Scilit]
  7. Chien, C.-T.; Ju, R.-Y.; Chou, K.-Y.; Chiang, J.-S. YOLOv9 for Fracture Detection in Pediatric Wrist Trauma X-ray Images. Electron. Lett. 2024, 60, e13248. [Google Scholar] [CrossRef] [Scilit]
  8. Chien, C.-T.; Ju, R.-Y.; Chou, K.-Y.; Xieerke, E.; Chiang, J.-S. YOLOv8-AM: YOLOv8 Based on Effective Attention Mechanisms for Pediatric Wrist Fracture Detection. IEEE Access 2025, 13, 52461–52477. [Google Scholar] [CrossRef] [Scilit]
  9. Du, S.; Wei, Y. ASC-YOLO: Multi-Scale Feature Fusion and Adaptive Decoupled Head for Fracture Detection in Medical Imaging. Appl. Sci. 2025, 15, 9031. [Google Scholar] [CrossRef] [Scilit]
  10. Zhang, W.; Ji, S. Rehabilitation Driven Optimized YOLOv11 Model for Medical X-Ray Fracture Detection. Sensors 2025, 25, 5793. [Google Scholar] [CrossRef] [Scilit]
  11. Vasu, P.K.A.; Gabriel, J.; Zhu, J.; Tuzel, O.; Ranjan, A. MobileOne: An Improved One Millisecond Mobile Backbone. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 7907–7917. [Google Scholar]
  12. Chen, T.; Moreau, T.; Jiang, Z.; Zheng, L.; Yan, E.; Shen, H.; Cowan, M.; Wang, L.; Hu, Y.; Ceze, L.; et al. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation, Carlsbad, CA, USA, 8–10 October 2018; pp. 578–594. [Google Scholar]
  13. Xu, Z.; Xu, J.; Peng, H.; Wang, W.; Wang, X.; Wan, H.; Dai, H.; Xu, Y.; Cheng, H.; Wang, K.; et al. ALT: Breaking the Wall between Data Layout and Loop Optimizations for Deep Learning Compilation. In Proceedings of the Eighteenth European Conference on Computer Systems, Rome, Italy, 8–12 May 2023; pp. 199–214. [Google Scholar]
  14. Chen, H.; Li, Y.; Zheng, B.; Chen, T. Hyperspectral Remote Sensing Image Classification Based on Domain-Level Complementarity of Spatial-Spectral Component. Neural Netw. 2026, 198, 108610. [Google Scholar] [CrossRef] [Scilit]
  15. Deng, W.; Li, K.; Zhao, H. A Flight Arrival Time Prediction Method Based on Cluster Clustering-Based Modular With Deep Neural Network. IEEE Trans. Intell. Transp. Syst. 2024, 25, 6238–6247. [Google Scholar] [CrossRef] [Scilit]
  16. Zeng, K.; Ma, Q.; Wu, J.W.; Chen, Z.; Shen, T.; Yan, C. FPGA-Based Accelerator for Object Detection: A Comprehensive Survey. J. Supercomput. 2022, 78, 14096–14136. [Google Scholar] [CrossRef] [Scilit]
  17. Galliera, R.; Suri, N. Object Detection at the Edge: Off-the-Shelf Deep Learning Capable Devices and Accelerators. Procedia Comput. Sci. 2022, 205, 239–248. [Google Scholar] [CrossRef] [Scilit]
  18. Zhai, J.; Li, B.; Lv, S.; Zhou, Q. FPGA-Based Vehicle Detection and Tracking Accelerator. Sensors 2023, 23, 2208. [Google Scholar] [CrossRef] [Scilit]
  19. Zaharia, C.; Popescu, V.; Sandu, F. Hardware–Software Partitioning for Real-Time Object Detection Using Dynamic Parameter Optimization. Sensors 2023, 23, 4894. [Google Scholar] [CrossRef] [Scilit]
  20. Shi, K.; Wang, M.; Tan, X.; Li, Q.; Lei, T. Efficient Dynamic Reconfigurable CNN Accelerator for Edge Intelligence Computing on FPGA. Information 2023, 14, 194. [Google Scholar] [CrossRef] [Scilit]
  21. Wang, Y.; Liao, Y.; Yang, J.; Wang, H.; Zhao, Y.; Zhang, C.; Xiao, B.; Xu, F.; Gao, Y.; Xu, M.; et al. An FPGA-Based Online Reconfigurable CNN Edge Computing Device for Object Detection. Microelectron. J. 2023, 137, 105805. [Google Scholar] [CrossRef] [Scilit]
  22. Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLOv8, Version 8.3.240, Used in the Present Study; Ultralytics Inc.: Frederick, MD, USA, 2025. Available online: https://github.com/ultralytics/ultralytics (accessed on 27 July 2026).
  23. Ding, X.; Zhang, X.; Ma, N.; Han, J.; Ding, G.; Sun, J. RepVGG: Making VGG-Style ConvNets Great Again. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 13733–13742. [Google Scholar]
  24. Ding, X.; Zhang, X.; Han, J.; Ding, G. Diverse Branch Block: Building a Convolution as an Inception-Like Unit. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 10886–10895. [Google Scholar]
  25. Hu, M.; Feng, J.; Hua, J.; Lai, B.; Huang, J.; Gong, X.; Hua, X.-S. Online Convolutional Re-Parameterization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 568–577. [Google Scholar]
  26. Karki, S.; Ahmed, Q.A.; Jungeblut, T. No Attention, No Problem: DPU-Aware Attention Approximation in Modern YOLO on FPGA. arXiv 2026, arXiv:2607.13106. [Google Scholar]
  27. AMD. Vitis AI User Guide, Version 2.5; UG1414; AMD: Santa Clara, CA, USA, 2022; Available online: https://docs.amd.com/r/2.5-English/ug1414-vitis-ai (accessed on 27 July 2026).
  28. Nagel, M.; Fournarakis, M.; Amjad, R.A.; Bondarenko, Y.; van Baalen, M.; Blankevoort, T. A White Paper on Neural Network Quantization. arXiv 2021, arXiv:2106.08295. [Google Scholar]
  29. Hubara, I.; Nahshan, Y.; Hanani, Y.; Banner, R.; Soudry, D. Accurate Post Training Quantization with Small Calibration Sets. Int. Conf. Mach. Learn. 2021, 139, 4466–4475. [Google Scholar]
  30. Li, Y.; Gong, R.; Tan, X.; Yang, Y.; Hu, P.; Zhang, Q.; Yu, F.; Wang, W.; Gu, S. BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction. In Proceedings of the 9th International Conference on Learning Representations, Virtual Event, 3–7 May 2021. [Google Scholar]
  31. AMD. DPUCZDX8G for Zynq UltraScale+ MPSoCs Product Guide, Version 4.0; PG338; AMD: Santa Clara, CA, USA, 2022; Available online: https://docs.amd.com/r/4.0-English/pg338-dpu (accessed on 27 July 2026).
  32. AMD. Vitis Unified Software Platform Documentation: Application Acceleration Development, Version 2022.1; UG1393; AMD: Santa Clara, CA, USA, 2022; Available online: https://docs.amd.com/r/2022.1-English/ug1393-vitis-application-acceleration/Control-Requirements-for-XRT-Managed-Kernels (accessed on 22 September 2026).
  33. Texas Instruments. INA260 Precision Digital Current and Power Monitor with Low-Drift, Precision Integrated Shunt; Data Sheet SBOS656C; Texas Instruments: Dallas, TX, USA, 2016; Available online: https://www.ti.com/product/INA260 (accessed on 27 July 2026).
  34. AMD. Kria KV260 Vision AI Starter Kit Data Sheet; DS986, Revision 1.3; AMD: Santa Clara, CA, USA, 2025; Available online: https://docs.amd.com/r/en-US/ds986-kv260-starter-kit/Product-Details (accessed on 27 July 2026).
Figure 1. Original C2f and the proposed C2fDeploy transformation. (a) Original C2f generates 2 C h channels with one 1 × 1 projection and performs a runtime channel split, which was lowered to CPU-side strided_slice operations in the evaluated Vitis AI 2.5 path. (b) C2fDeploy partitions the trained projection and channel-wise BN state into two C h -channel branches, removing the runtime split while preserving the downstream topology, parameter count, and theoretical convolutional arithmetic. Solid black arrows indicate tensor flow; red dashed boxes highlight the runtime channel split and its CPU-side execution path, green boxes denote the two partitioned projection branches, and the black dashed annotation indicates that no runtime split remains after offline parameter partitioning.
Figure 1. Original C2f and the proposed C2fDeploy transformation. (a) Original C2f generates 2 C h channels with one 1 × 1 projection and performs a runtime channel split, which was lowered to CPU-side strided_slice operations in the evaluated Vitis AI 2.5 path. (b) C2fDeploy partitions the trained projection and channel-wise BN state into two C h -channel branches, removing the runtime split while preserving the downstream topology, parameter count, and theoretical convolutional arithmetic. Solid black arrows indicate tensor flow; red dashed boxes highlight the runtime channel split and its CPU-side execution path, green boxes denote the two partitioned projection branches, and the black dashed annotation indicates that no runtime split remains after offline parameter partitioning.
Electronics 15 04431 g001
Figure 2. KV260 deployment workflow and measurement protocol. (a) Application workflow from radiograph input through Arm-based preprocessing, VART/XIR execution of the INT8 DPU subgraph, Arm-based output processing, and production of the detection result. (b) Synchronized DPU-stage timing and power measurement. After 50 excluded warm-up executions, 2500 DPU-stage executions were timed while the onboard INA260 interface sampled the 5 V V CC _ SOM rail over the same active interval. The derived throughput, throughput per watt, and energy per execution characterize the timed DPU-stage loop rather than total starter-kit input power or complete application latency. Arrows indicate the data or measurement flow. In panel (a), blue boxes denote Arm/VART-based processing stages and the orange box denotes DPU computation. In panel (b), blue and purple elements denote the latency and power measurement tracks, respectively; the green interval denotes the timed DPU-stage execution window, and the red dashed box denotes excluded warm-up runs.
Figure 2. KV260 deployment workflow and measurement protocol. (a) Application workflow from radiograph input through Arm-based preprocessing, VART/XIR execution of the INT8 DPU subgraph, Arm-based output processing, and production of the detection result. (b) Synchronized DPU-stage timing and power measurement. After 50 excluded warm-up executions, 2500 DPU-stage executions were timed while the onboard INA260 interface sampled the 5 V V CC _ SOM rail over the same active interval. The derived throughput, throughput per watt, and energy per execution characterize the timed DPU-stage loop rather than total starter-kit input power or complete application latency. Arrows indicate the data or measurement flow. In panel (a), blue boxes denote Arm/VART-based processing stages and the orange box denotes DPU computation. In panel (b), blue and purple elements denote the latency and power measurement tracks, respectively; the green interval denotes the timed DPU-stage execution window, and the red dashed box denotes excluded warm-up runs.
Electronics 15 04431 g002
Figure 3. Compiled XModel partitioning under the same deploy-s-512 checkpoint, quantization protocol, compiler configuration, and DPU target. (a) Original C2f introduces CPU-side strided_slice and associated fix2float/float2fix conversions, fragmenting the accelerator-supported path into nine DPU subgraphs. (b) C2fDeploy removes the runtime split and consolidates the DPU-supported computation into one DPU subgraph. The three remaining CPU subgraphs contain only output-side fix2float conversions. Subgraph and operation counts were obtained from XIR inspection.
Figure 3. Compiled XModel partitioning under the same deploy-s-512 checkpoint, quantization protocol, compiler configuration, and DPU target. (a) Original C2f introduces CPU-side strided_slice and associated fix2float/float2fix conversions, fragmenting the accelerator-supported path into nine DPU subgraphs. (b) C2fDeploy removes the runtime split and consolidates the DPU-supported computation into one DPU subgraph. The three remaining CPU subgraphs contain only output-side fix2float conversions. Subgraph and operation counts were obtained from XIR inspection.
Electronics 15 04431 g003
Table 1. Patient-level partition of the GRAZPEDWRI-DX images used in this study.
Table 1. Patient-level partition of the GRAZPEDWRI-DX images used in this study.
SplitImagesPercentageRole
Training14,20469.88%Model fitting and PTQ calibration pool
Validation409420.14%Model monitoring and checkpoint selection
Test20299.98%Final FP32 and INT8 evaluation
Table 2. Stage-specific software environments used in this study. Software versions are reported according to the experimental stage in which they were used and should not be interpreted as components of a single shared Python environment.
Table 2. Stage-specific software environments used in this study. Software versions are reported according to the experimental stage in which they were used and should not be interpreted as components of a single shared Python environment.
StageSoftware/ToolchainKey Configuration
TrainingPython 3.13.5; project-customized Ultralytics 8.3.240AdamW explicitly selected; 640 × 640 input; maximum 100 epochs; batch size 16; patience 15; l r 0 = 0.001 ; l r f = 0.01 ; weight decay 5 × 10 − 4 ; cosine learning-rate scheduling enabled; AMP disabled; seed 0; deterministic execution enabled
FP32 validationPython 3.13.5; PyTorch 2.9.0+cpu; Ultralytics 8.3.240 512 × 512 input; batch size 1
PTQ/compilationVitis AI 2.5200 fixed training images for calibration; input shape 1 × 3 × 512 × 512 ; target DPUCZDX8G_ISA1_B4096
Board runtimePython 3.10.12; Vitis AI Runtime/Library 2.5.0; XRT 2.13.479AMD Kria KV260 Vision AI Starter Kit (AMD, Santa Clara, CA, USA); target DPUCZDX8G_ISA1_B4096
Table 3. Controlled FP32 and INT8 comparison on the held-out test set of 2029 images. Both graph variants used the same deploy-s-512 checkpoint and the same ordered calibration set.
Table 3. Controlled FP32 and INT8 comparison on the held-out test set of 2029 images. Both graph variants used the same deploy-s-512 checkpoint and the same ordered calibration set.
Graph and StagePrecisionRecallmAP50mAP50–95
Original C2f, FP320.941600.875790.949410.63525
C2fDeploy, FP320.941600.875790.949410.63525
Original C2f, INT80.933400.876770.945980.62970
C2fDeploy, INT80.937120.875080.946500.63033
Table 4. Tensor-level FP32 comparison between Original C2f and C2fDeploy. Errors were accumulated over corresponding activation tensors for all eight rewritten C2f blocks. The final row reports the aggregate across all C2f outputs; the raw network output before NMS is reported separately below the table.
Table 4. Tensor-level FP32 comparison between Original C2f and C2fDeploy. Errors were accumulated over corresponding activation tensors for all eight rewritten C2f blocks. The final row reports the aggregate across all C2f outputs; the raw network output before NMS is reported separately below the table.
BlockMax. Abs. ErrorMAERMSERelative L 2 Error
model.2 7.629395 × 10 − 6 6.802229 × 10 − 8 1.821058 × 10 − 7 1.120294 × 10 − 7
model.4 9.536743 × 10 − 6 1.651143 × 10 − 7 3.341610 × 10 − 7 2.942311 × 10 − 7
model.6 7.390976 × 10 − 6 1.280678 × 10 − 7 2.900479 × 10 − 7 3.307216 × 10 − 7
model.8 5.245209 × 10 − 6 8.857794 × 10 − 8 2.164007 × 10 − 7 1.965278 × 10 − 7
model.12 6.318092 × 10 − 6 6.611016 × 10 − 8 1.879275 × 10 − 7 1.918238 × 10 − 7
model.15 1.335144 × 10 − 5 9.390036 × 10 − 8 2.209009 × 10 − 7 3.037120 × 10 − 7
model.18 1.430511 × 10 − 5 8.902064 × 10 − 8 2.538226 × 10 − 7 3.350905 × 10 − 7
model.21 5.364418 × 10 − 6 1.006492 × 10 − 7 2.381456 × 10 − 7 1.983245 × 10 − 7
All C2f outputs 1.430511 × 10 − 5 9.732755 × 10 − 8 2.396549 × 10 − 7 1.965703 × 10 − 7
Table 5. Compiled XModel structure of Original C2f and C2fDeploy.
Table 5. Compiled XModel structure of Original C2f and C2fDeploy.
Compilation MetricOriginal C2fC2fDeploy
Total top-level subgraphs295
USER subgraphs11
DPU subgraphs91
CPU subgraphs193
CPU strided_slice operations160
CPU float2fix operations160
CPU fix2float operations193
Total CPU-assigned operations513
Table 6. Same-checkpoint, single-thread whole-XModel comparison on the Kria KV260. Values are the mean ± sample standard deviation across five formal runs. Displayed values are rounded for readability; unrounded results are retained in the archived benchmark logs.
Table 6. Same-checkpoint, single-thread whole-XModel comparison on the Kria KV260. Values are the mean ± sample standard deviation across five formal runs. Displayed values are rounded for readability; unrounded results are retained in the archived benchmark logs.
MetricOriginal C2fC2fDeployRelative Change
Whole-XModel throughput (graph executions/s) 0.5158 ± 0.0001 49.13 ± 0.04 95.2 × higher
Reciprocal latency equivalent (ms/execution) 1938.7 ± 0.2 20.36 ± 0.02 98.95% lower
Mean V CC _ SOM power (W) 5.087 ± 0.003 8.735 ± 0.003 71.7% higher
Throughput per watt (executions/s/W) 0.1014 ± 0.0001 5.624 ± 0.005 55.5 × higher
V CC _ SOM energy per graph execution (J/execution) 9.862 ± 0.006 0.1778 ± 0.0002 98.2% lower
Table 7. Independent external 12 V input-power cross-check for the same-checkpoint Original C2f and C2fDeploy XModels. Power values are based on five steady-state observations per approximately 60 s run and are reported as the mean ± sample standard deviation across five valid runs.
Table 7. Independent external 12 V input-power cross-check for the same-checkpoint Original C2f and C2fDeploy XModels. Power values are based on five steady-state observations per approximately 60 s run and are reported as the mean ± sample standard deviation across five valid runs.
MetricOriginal C2fC2fDeployRelative Change
Whole-XModel throughput (executions/s) 0.515571 ± 0.000051 49.00434 ± 0.01991 95.05 × higher
External 12 V input power (W) 8.64 ± 0.05 12.50 ± 0.05 44.72% higher
External-input energy per graph execution (J/execution) 16.76 ± 0.11 0.255 ± 0.001 98.48% lower
Table 8. Runtime utilization and APM-observed external-memory measurements for the same-checkpoint Original C2f and C2fDeploy XModels. DPU utilization denotes the time-weighted compute-unit busy fraction obtained from the XRT-exposed ap_idle state and does not represent internal PE- or MAC-lane occupancy.
Table 8. Runtime utilization and APM-observed external-memory measurements for the same-checkpoint Original C2f and C2fDeploy XModels. DPU utilization denotes the time-weighted compute-unit busy fraction obtained from the XRT-exposed ap_idle state and does not represent internal PE- or MAC-lane occupancy.
MetricOriginal C2fC2fDeployRelative Change
DPU CU busy-time utilization (%) 0.3562 ± 0.0092 90.7465 ± 0.1356 + 90.3903 percentage points
System CPU utilization (%) 25.3365 ± 0.0187 2.6536 ± 0.0318 89.53% lower
APM read bandwidth (MB/s) 28.3452 ± 0.0476 2289.9858 ± 1.5290 80.79 × higher
APM write bandwidth (MB/s) 11.8584 ± 0.0911 960.5783 ± 0.4875 81.00 × higher
Aggregate APM bandwidth (MB/s) 40.2036 ± 0.0963 3250.5641 ± 2.0062 80.85 × higher
Normalized aggregate APM traffic (MB/execution) 77.8675 ± 0.1788 66.3395 ± 0.0169 14.80% lower
Table 9. Accuracy and custom VART DPU-stage operating points of the final C2fDeploy models. Power and energy refer to the 5 V V CC _ SOM rail during the timed DPU execution loop.
Table 9. Accuracy and custom VART DPU-stage operating points of the final C2fDeploy models. Power and energy refer to the 5 V V CC _ SOM rail during the timed DPU execution loop.
ModelFP32 mAP50–95INT8 mAP50–95Latency (ms)FPSPower (W)FPS/WEnergy (J)
deploy-s-5120.6350.63021.48646.5428.6015.4120.18480
deploy-n-5120.6270.6187.857127.2808.05015.8110.06325
Table 10. Complete end-to-end GraphRunner latency comparison between Original C2f and C2fDeploy using the same deploy-s-512 checkpoint and application pipeline. Stage values are mean latencies across 50 timed repetitions after five excluded complete-pipeline warm-up executions. Visualization and result-file writing were excluded.
Table 10. Complete end-to-end GraphRunner latency comparison between Original C2f and C2fDeploy using the same deploy-s-512 checkpoint and application pipeline. Stage values are mean latencies across 50 timed repetitions after five excluded complete-pipeline warm-up executions. Visualization and result-file writing were excluded.
Pipeline StageOriginal C2fC2fDeploy
Image loading35.978 ms35.629 ms
Host preprocessing31.519 ms22.864 ms
Input tensor handoff0.569 ms0.557 ms
Whole-XModel execution1970.381 ms20.636 ms
Output dequantization/copy2.847 ms2.824 ms
YOLO decoding4.055 ms3.311 ms
NMS0.573 ms0.588 ms
Total end-to-end latency2045.923 ms86.410 ms
Detection throughput0.489 FPS11.573 FPS
Table 11. Qualitative comparison of C2fDeploy and alternative deployment strategies.
Table 11. Qualitative comparison of C2fDeploy and alternative deployment strategies.
StrategyMain MechanismAdditional TrainingFunction PreservationImplementation Complexity
Source-level C2f replacementModify YOLOv8 C2f and replace split with independent convolution branchesNot required with exact parameter mapping; otherwise requiredExact with equivalent parameter mappingMedium; requires modified model definition and deployment integration
Vitis AI custom/fused layerHardware-specific kernels or compiler extensionsNoYes if correctly implementedHigh; requires hardware–software co-design
Attention approximation methods (e.g., Karki et al. [26])Reduce attention computation through approximation strategiesUsually noApproximateMedium–high; requires accuracy–efficiency evaluation
C2fDeployOffline partition of trained convolution/BN parameters with equivalent rewritingNoExactLow; uses existing compiler-supported operators
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ji, X.; Hu, S.; Wang, H.; Hao, W.; Li, X. C2fDeploy: Function-Preserving Graph Rewriting to Eliminate Runtime Split Overhead on FPGA Deep Learning Processing Units. Electronics 2026, 15, 4431. https://doi.org/10.3390/electronics15194431

AMA Style

Ji X, Hu S, Wang H, Hao W, Li X. C2fDeploy: Function-Preserving Graph Rewriting to Eliminate Runtime Split Overhead on FPGA Deep Learning Processing Units. Electronics. 2026; 15(19):4431. https://doi.org/10.3390/electronics15194431

Chicago/Turabian Style

Ji, Xiang, Shuaifei Hu, Haofei Wang, Wanming Hao, and Xiangnan Li. 2026. "C2fDeploy: Function-Preserving Graph Rewriting to Eliminate Runtime Split Overhead on FPGA Deep Learning Processing Units" Electronics 15, no. 19: 4431. https://doi.org/10.3390/electronics15194431

APA Style

Ji, X., Hu, S., Wang, H., Hao, W., & Li, X. (2026). C2fDeploy: Function-Preserving Graph Rewriting to Eliminate Runtime Split Overhead on FPGA Deep Learning Processing Units. Electronics, 15(19), 4431. https://doi.org/10.3390/electronics15194431

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop