Author Contributions
Conceptualization, H.M.B., D.G. and A.E.-B.; methodology, H.M.B., D.G. and A.E.-B.; software, H.M.B.; validation, H.M.B., A.M., A.A., M.T.A. and G.A.G.; formal analysis, H.M.B., D.G. and A.E.-B.; investigation, H.M.B., D.G. and A.E.-B.; resources, K.M.A. and D.G.; data curation, K.M.A. and D.G.; writing—original draft preparation, H.M.B., D.G. and A.E.-B.; writing—review and editing, H.M.B., K.M.A., A.M., A.A., M.T.A. and G.A.G.; visualization, H.M.B., A.M., A.A., M.T.A. and G.A.G.; supervision, D.G. and A.E.-B.; project administration, D.G. and A.E.-B.; funding acquisition, D.G. and A.E.-B. All authors have read and agreed to the published version of the manuscript.
Figure 1.
The proposed framework for virtual H&E-to-MT staining consisting of four main stages: (1) data pre-processing (including tile extraction and alignment); (2) stain normalization (using Reinhard’s method); (3) stain translation (using the proposed TbGAN); and (4) weighted fusion of multiple configurations (to enhance robustness and generalizability).
Figure 1.
The proposed framework for virtual H&E-to-MT staining consisting of four main stages: (1) data pre-processing (including tile extraction and alignment); (2) stain normalization (using Reinhard’s method); (3) stain translation (using the proposed TbGAN); and (4) weighted fusion of multiple configurations (to enhance robustness and generalizability).
Figure 2.
Illustration of key dataset challenges addressed in our study, including misalignment (rotational and translational) between H&E and MT slide pairs, inconsistent staining quality across slides, and variable tissue placement on glass slides.
Figure 2.
Illustration of key dataset challenges addressed in our study, including misalignment (rotational and translational) between H&E and MT slide pairs, inconsistent staining quality across slides, and variable tissue placement on glass slides.
Figure 3.
Visualization of the alignment of corresponding tissue regions from H&E- and MT-stained slides using the SIFT algorithm. High-confidence feature matches are highlighted, ensuring accurate spatial correspondence for downstream multi-modal analysis. The lower row refers to the matched H&E and MT thumbnails.
Figure 3.
Visualization of the alignment of corresponding tissue regions from H&E- and MT-stained slides using the SIFT algorithm. High-confidence feature matches are highlighted, ensuring accurate spatial correspondence for downstream multi-modal analysis. The lower row refers to the matched H&E and MT thumbnails.
Figure 4.
Visualization of stain normalization using Reinhard’s method. It demonstrates the translation of a sample input H&E-stained image into a normalized version using Reinhard’s method. That process aligns the color distribution of the input image to a target image, ensuring consistency in color appearance.
Figure 4.
Visualization of stain normalization using Reinhard’s method. It demonstrates the translation of a sample input H&E-stained image into a normalized version using Reinhard’s method. That process aligns the color distribution of the input image to a target image, ensuring consistency in color appearance.
Figure 5.
Visualization of the weighted fusion process. It illustrates the fusion of outputs from the four configurations: original with emphasis on , original (short as O) with emphasis on , Reinhard-normalized (short as R) with emphasis on , and Reinhard-normalized with emphasis on . The fusion process combines these outputs using a weighted averaging approach, balancing identity preservation and cycle-consistency.
Figure 5.
Visualization of the weighted fusion process. It illustrates the fusion of outputs from the four configurations: original with emphasis on , original (short as O) with emphasis on , Reinhard-normalized (short as R) with emphasis on , and Reinhard-normalized with emphasis on . The fusion process combines these outputs using a weighted averaging approach, balancing identity preservation and cycle-consistency.
Figure 6.
Box plots illustrating the distribution of key similarity metrics (MI, NMI, SSIM, NCC, and CS) across the four individual configurations (O/3/10, O/10/3, R/3/10, and R/10/3) and the fused configuration over 24 trials. Each box represents the IQR of the metric values, with the median indicated by the central line. Whiskers extend to the minimum and maximum values within 1.5 times the IQR, and outliers are shown as individual points. The fusion configuration demonstrates consistently tighter IQRs and fewer outliers across all five metrics, indicating superior consistency and reliability in preserving structural alignment, color fidelity, and statistical similarity.
Figure 6.
Box plots illustrating the distribution of key similarity metrics (MI, NMI, SSIM, NCC, and CS) across the four individual configurations (O/3/10, O/10/3, R/3/10, and R/10/3) and the fused configuration over 24 trials. Each box represents the IQR of the metric values, with the median indicated by the central line. Whiskers extend to the minimum and maximum values within 1.5 times the IQR, and outliers are shown as individual points. The fusion configuration demonstrates consistently tighter IQRs and fewer outliers across all five metrics, indicating superior consistency and reliability in preserving structural alignment, color fidelity, and statistical similarity.
Figure 7.
Violin plot illustrating the distribution of key similarity metrics (MI, NMI, SSIM, NCC, and CS) across the four individual configurations (O/3/10, O/10/3, R/3/10, and R/10/3) and the fused configuration over 24 trials where each violin shape represents the kernel density of metric values, with width corresponding to the relative frequency of observations. The fusion configuration exhibits taller and narrower violins across all five metrics, indicating a higher concentration of results near the median.
Figure 7.
Violin plot illustrating the distribution of key similarity metrics (MI, NMI, SSIM, NCC, and CS) across the four individual configurations (O/3/10, O/10/3, R/3/10, and R/10/3) and the fused configuration over 24 trials where each violin shape represents the kernel density of metric values, with width corresponding to the relative frequency of observations. The fusion configuration exhibits taller and narrower violins across all five metrics, indicating a higher concentration of results near the median.
Figure 8.
Step plots showing the trial-wise performance across 24 experiments for five key similarity metrics (MI, NMI, SSIM, NCC, and CS) where each subplot tracks the metric values for the four individual configurations (O/10/3, O/3/10, R/10/3, and R/3/10) and the fused configuration over successive trials. The fused approach consistently reports higher metric values with smoother trajectories, indicating stable and superior performance across all trials and metrics.
Figure 8.
Step plots showing the trial-wise performance across 24 experiments for five key similarity metrics (MI, NMI, SSIM, NCC, and CS) where each subplot tracks the metric values for the four individual configurations (O/10/3, O/3/10, R/10/3, and R/3/10) and the fused configuration over successive trials. The fused approach consistently reports higher metric values with smoother trajectories, indicating stable and superior performance across all trials and metrics.
Figure 9.
Raincloud plot of SSIM performance across the four individual configurations (O/10/3, O/3/10, R/10/3, and R/3/10) and the fused configuration over 24 trials. The plot combines a box plot (interquartile range and median), raw data points (rain), and a smoothed density distribution (cloud). The fused configuration shows a higher median SSIM (0.7474), tighter spread, and greater density around the peak, confirming its enhanced consistency and reliability in preserving structural similarity.
Figure 9.
Raincloud plot of SSIM performance across the four individual configurations (O/10/3, O/3/10, R/10/3, and R/3/10) and the fused configuration over 24 trials. The plot combines a box plot (interquartile range and median), raw data points (rain), and a smoothed density distribution (cloud). The fused configuration shows a higher median SSIM (0.7474), tighter spread, and greater density around the peak, confirming its enhanced consistency and reliability in preserving structural similarity.
Figure 10.
Qualitative analysis of the proposed framework: sample outputs for 10 cases. For each case, there are four images: H&E slide, deformed H&E slide (aligned using FFD), MT slide (ground truth), and generated MT slide. The findings illustrate the system’s capability to correctly convert H&E-stained images to virtual MT-stained images while maintaining not only the structural alignment but also the color and texture details. The synthesized MT images look very much like the real ones, which is an indication that the proposed method is quite successful.
Figure 10.
Qualitative analysis of the proposed framework: sample outputs for 10 cases. For each case, there are four images: H&E slide, deformed H&E slide (aligned using FFD), MT slide (ground truth), and generated MT slide. The findings illustrate the system’s capability to correctly convert H&E-stained images to virtual MT-stained images while maintaining not only the structural alignment but also the color and texture details. The synthesized MT images look very much like the real ones, which is an indication that the proposed method is quite successful.
Figure 11.
End-to-end clinical integration pathway of the TbGAN framework. The modular architecture transforms H&E slides into diagnostic-quality virtual MT images in 1.8 s per tile without disrupting existing digital pathology workflows or requiring additional tissue sections.
Figure 11.
End-to-end clinical integration pathway of the TbGAN framework. The modular architecture transforms H&E slides into diagnostic-quality virtual MT images in 1.8 s per tile without disrupting existing digital pathology workflows or requiring additional tissue sections.
Figure 12.
Inference-stage output of the TbGAN framework showing virtual Masson’s Trichrome (MT) generation from H&E input across diverse fibrosis stages. The system preserves collagen morphology (blue) and cellular structures (red) without requiring alignment or FFD during inference.
Figure 12.
Inference-stage output of the TbGAN framework showing virtual Masson’s Trichrome (MT) generation from H&E input across diverse fibrosis stages. The system preserves collagen morphology (blue) and cellular structures (red) without requiring alignment or FFD during inference.
Table 1.
Comparative analysis between Naglah et al.’s cGAN approach [
22] and our TbGAN framework across architectural design, alignment handling, and clinical applicability.
Table 1.
Comparative analysis between Naglah et al.’s cGAN approach [
22] and our TbGAN framework across architectural design, alignment handling, and clinical applicability.
| Feature | Naglah et al. (cGAN) [22] | Proposed TbGAN |
|---|
| Architecture | U-Net with skip connections (CNN-based) | Transformer encoder + CNN decoder (hybrid) |
| Context modeling | Local receptive fields (max 7 × 7 kernel) | Global self-attention (N = 1024 tokens) |
| Alignment strategy | Required near-perfect pre-alignment | Multi-stage pipeline (SIFT → ORB → FFD) |
| Stain robustness | SSIM = 0.18:0.23 across protocols | SSIM = 0.0597 (fusion mechanism) |
| Training data | Paired H&E/MT required per institution | Unsupervised; no site-specific retraining |
| Clinical deployment | Requires per-institution recalibration | Plug-and-play integration with software like QuPath v0.6.0 |
| Fibrosis staging accuracy | METAVIR agreement: = 0.68 | METAVIR agreement: = 0.82 (projected) |
Table 2.
Summary of dataset challenges and the proposed methodological solutions in the H&E-to-MT virtual staining framework.
Table 2.
Summary of dataset challenges and the proposed methodological solutions in the H&E-to-MT virtual staining framework.
| Challenge | Proposed Solution |
|---|
| Misalignment between H&E and MT slides due to use of consecutive tissue sections, leading to spatial discrepancies. | Multi-stage alignment pipeline: (i) Coarse alignment using SIFT on WSI thumbnails. (ii) Patch-level alignment via ORB + homography (RANSAC). (iii) Fine-grained local deformation correction using free-form deformation (FFD) with B-spline registration. |
| Minor rotational and translational misalignments introduced during slide scanning or handling. | Addressed within the same alignment pipeline: rigid-body (global) registration via SIFT/ORB, followed by non-rigid FFD for residual local shifts. |
| Inconsistent tissue placement on glass slides (e.g., rotation, offset, cropping differences). | Tissue region detection via contour extraction from thumbnails; patches extracted only from overlapping tissue regions using center-aligned bounding boxes. |
| Staining quality inconstancy across slides (e.g., over/under-staining, reagent degradation, thickness variation). | Stain normalization using Reinhard’s color transfer method in LAB color space to standardize appearance across all input images. |
| Artifacts such as tissue folding, air bubbles, or uneven staining. | Patch-level quality filtering: Discard patches with > empty regions or low similarity (MI , Cosine , pHash ). |
| Computational burden from extremely large WSI dimensions (>100,000 × 100,000 pixels). | Tile-based processing: Extract patches from aligned ROIs; process in manageable batches while preserving spatial context. |
Table 3.
Tabular presentation of the four distinct configurations utilized in this study as they were carefully designed to assess the trade-offs between identity preservation and cycle-consistency (where each configuration emphasizes specific hyperparameters and provides either original or Reinhard-normalized images to enhance the stain translation process).
Table 3.
Tabular presentation of the four distinct configurations utilized in this study as they were carefully designed to assess the trade-offs between identity preservation and cycle-consistency (where each configuration emphasizes specific hyperparameters and provides either original or Reinhard-normalized images to enhance the stain translation process).
| Configuration | Emphasis | Hypothesis | Hyperparameters |
|---|
| Original with Emphasis on | The model was trained on the original images with a strong emphasis on the identity loss () | This setup ensures that the model preserves the structural and textural integrity of the input images when no translation is required | The hyperparameter was scaled by a factor of 3, while was set to its default value |
| Original with Emphasis on | This configuration focuses on enforcing cycle-consistency in the original images | This setup ensures that the translated images maintain semantic and structural alignment between the source and target domains | The hyperparameter was scaled by a factor of 3, while was set to its default value |
| Reinhard-Normalized with Emphasis on | The model was trained on Reinhard-normalized images with a strong emphasis on the identity loss () | The normalization process aligns the color distribution of the input images to a target image, ensuring consistency in color appearance | The hyperparameter was scaled by a factor of 3, while was set to its default value |
| Reinhard-Normalized with Emphasis on | This configuration focuses on enforcing cycle-consistency in Reinhard-normalized images | This setup ensures that the translated images are both realistic and consistent with the original data while also benefiting from the color normalization process | The hyperparameter was scaled by a factor of 3, while was set to its default value |
Table 4.
Tabular presentation of the quantitative analysis of the discussed metrics.
Table 4.
Tabular presentation of the quantitative analysis of the discussed metrics.
| Metric | Equation | Category | Benefit | Insights |
|---|
| Mutual Information (MI) | | Similarity | Measures statistical dependence between images. | Evaluates global similarity and structural alignment. |
| Structural Similarity Index (SSIM) [53] | | Similarity | Evaluates structural similarity. | Assesses structural alignment, luminance, and contrast. |
| Normalized Cross-Correlation (NCC) | | Similarity | Measures normalized correlation. | Evaluates pixel-wise alignment and color consistency. |
| Cosine Similarity (CS) | | Similarity | Quantifies the angle between image vectors. | Assesses global similarity and texture preservation. |
| Histogram Intersection (HistInt) | | Similarity | Measures overlap between histograms. | Evaluates color consistency and texture preservation. |
| Universal Quality Index (UQI) [54] | | Similarity | Evaluates image quality. | Assesses luminance, contrast, and structural alignment. |
| Mean Squared Error (MSE) | | Dissimilarity | Measures average squared difference. | Evaluates pixel-wise accuracy and color consistency. |
| Earth Mover’s Distance (EMD) [55] | | Dissimilarity | Quantifies the work required to transform one image into another. | Assesses global alignment and texture preservation. |
| Perceptual Hash (pHash) [56] | | Dissimilarity | Measures perceptual similarity. | Evaluates perceptual quality and visual artifacts. |
| Jensen–Shannon Divergence (JSD) [57] | | Dissimilarity | Quantifies the difference between probability distributions. | Assesses texture similarity and color consistency. |
| Blind/Referenceless Image Spatial Quality Evaluator (BRISQUE) [58] | | Dissimilarity | Evaluates perceptual quality without a reference. | Evaluates perceptual quality and visual artifacts. |
Table 5.
Tabular presentation of the utilized hyperparameters.
Table 5.
Tabular presentation of the utilized hyperparameters.
| Parameter | Value | Rationale |
|---|
| Input dimensions | 256 × 256 × 4 (RGBA) | Balance between contextual information and GPU memory constraints |
| Batch size | 1 | Required due to memory limitations of 6GB GPU |
| Optimizer | Adam | Stable convergence for GAN training |
| Learning rate | | Prevents mode collapse while ensuring convergence |
| , | 0.5, 0.999 | Standard for GAN stability |
| Training epochs | 200 | Determined via early stopping (val loss plateau) |
| 10 (default), 3 (emphasized) | Enforces cycle-consistency |
| 10 (default), 3 (emphasized) | Preserves structural identity |
| Patch overlap | 32 pixels | Maintains spatial context during tile extraction |
| Training/validation split | 27 WSIs: 80%/20% (patient-level) | Prevents data leakage across subjects |
| Device | Windows 11/6 GB GPU/256 GB RAM | GPU of optimization and memory for data |
Table 6.
Quantitative analysis of the proposed approach across 24 trials and four approaches alongside the fusion between them. The performance metrics for the testing subset (including MI, SSIM, and GMSD) are presented. The mean ± standard deviation values are reported to demonstrate the consistency and reliability of the proposed approach.
Table 6.
Quantitative analysis of the proposed approach across 24 trials and four approaches alongside the fusion between them. The performance metrics for the testing subset (including MI, SSIM, and GMSD) are presented. The mean ± standard deviation values are reported to demonstrate the consistency and reliability of the proposed approach.
| Metric | O/10/3 | O/3/10 | R/10/3 | R/3/10 | Fused |
|---|
| MI (↑) | 0.9145 ± 0.084 | 0.946 ± 0.0918 | 0.8883 ± 0.0822 | 0.9183 ± 0.0891 | 0.9815 ± 0.0934 |
| NMI (↑) | 0.2561 ± 0.0169 | 0.2639 ± 0.0182 | 0.252 ± 0.0163 | 0.2572 ± 0.0172 | 0.2678 ± 0.0184 |
| SSIM (↑) | 0.7137 ± 0.0576 | 0.735 ± 0.0567 | 0.7102 ± 0.0621 | 0.7233 ± 0.062 | 0.7474 ± 0.0597 |
| NCC (↑) | 0.9209 ± 0.0253 | 0.9257 ± 0.0256 | 0.9179 ± 0.0238 | 0.9225 ± 0.0235 | 0.932 ± 0.022 |
| CS (↑) | 0.9939 ± 0.0013 | 0.9943 ± 0.0014 | 0.9931 ± 0.0018 | 0.9935 ± 0.002 | 0.9946 ± 0.0014 |
| HistInt (↑) | 0.8917 ± 0.057 | 0.891 ± 0.0606 | 0.8538 ± 0.0685 | 0.8528 ± 0.0761 | 0.8733 ± 0.0685 |
| PSNR (↑) | 30.0217 ± 0.2314 | 30.0954 ± 0.2629 | 29.8645 ± 0.2464 | 29.9013 ± 0.2849 | 30.1037 ± 0.3247 |
| FBS (↑) | 0.1874 ± 0.0572 | 0.2192 ± 0.0628 | 0.1745 ± 0.0509 | 0.201 ± 0.0584 | 0.2338 ± 0.0651 |
| UQI (↑) | 0.9183 ± 0.0266 | 0.9229 ± 0.0272 | 0.9128 ± 0.029 | 0.9166 ± 0.0294 | 0.9286 ± 0.0248 |
| SRS (↑) | 0.9333 ± 0.0215 | 0.9373 ± 0.0219 | 0.9278 ± 0.0219 | 0.9325 ± 0.0217 | 0.9415 ± 0.0198 |
| PC (↑) | 0.9028 ± 0.0325 | 0.9087 ± 0.0319 | 0.8971 ± 0.0336 | 0.9022 ± 0.0334 | 0.9134 ± 0.0302 |
| NQM (↑) | 7.4164 ± 0.4671 | 7.3047 ± 0.4274 | 7.8393 ± 0.8897 | 7.847 ± 1.0481 | 7.4815 ± 0.763 |
| MSE (↓) | 64.8824 ± 3.4925 | 63.8301 ± 3.9572 | 67.318 ± 3.8438 | 66.8074 ± 4.4377 | 63.8018 ± 4.8184 |
| NMSE (↓) | 0.1581 ± 0.0506 | 0.1485 ± 0.0512 | 0.1642 ± 0.0475 | 0.155 ± 0.0471 | 0.1359 ± 0.0439 |
| EMD (↓) | 4.3745 ± 3.1701 | 4.3383 ± 3.2906 | 6.9242 ± 4.2476 | 7.0563 ± 4.8512 | 5.1765 ± 3.739 |
| HD (↓) | 0.0849 ± 0.0304 | 0.0859 ± 0.033 | 0.1097 ± 0.0396 | 0.1112 ± 0.0465 | 0.0897 ± 0.0419 |
| BhD (↓) | 0.0087 ± 0.0076 | 0.0091 ± 0.0087 | 0.0157 ± 0.0133 | 0.0165 ± 0.016 | 0.011 ± 0.0109 |
| pHash (↓) | 6.8566 ± 1.5391 | 6.5555 ± 1.4951 | 7.0065 ± 1.4431 | 6.9102 ± 1.4239 | 6.3476 ± 1.3693 |
| JSD (↓) | 0.044 ± 0.0051 | 0.0426 ± 0.0055 | 0.0458 ± 0.0058 | 0.0446 ± 0.0062 | 0.0408 ± 0.0052 |
| KLD (↓) | 0.0335 ± 0.0368 | 0.0349 ± 0.0407 | 0.0593 ± 0.0504 | 0.0649 ± 0.0636 | 0.0441 ± 0.0427 |
| BRISQUE (↓) | 15.5911 ± 3.0334 | 15.5911 ± 3.0334 | 15.5911 ± 3.0334 | 15.5911 ± 3.0334 | 15.2632 ± 5.958 |
| GMSD (↓) | 0.206 ± 0.0096 | 0.2023 ± 0.01 | 0.2082 ± 0.0103 | 0.2053 ± 0.0106 | 0.2009 ± 0.0105 |
Table 7.
ANOVA results and Tukey HSD pairwise comparisons with 95% confidence intervals for key similarity metrics. All p-values are Bonferroni-adjusted. Confidence intervals represent mean difference (fused − configuration).
Table 7.
ANOVA results and Tukey HSD pairwise comparisons with 95% confidence intervals for key similarity metrics. All p-values are Bonferroni-adjusted. Confidence intervals represent mean difference (fused − configuration).
| Metric | F-Statistic | p-Value | Fused vs. O/10/3 | Fused vs. O/3/10 | Fused vs. R/10/3 | Fused vs. R/3/10 |
|---|
| MI | 190.89 | < | [0.057, 0.077] | [0.028, 0.043] | [0.082, 0.105] | [0.056, 0.070] |
| NMI | 124.41 | < | [0.009, 0.014] | [0.002, 0.006] | [0.013, 0.018] | [0.009, 0.012] |
| SSIM | 138.46 | < | [0.027, 0.040] | [0.008, 0.016] | [0.034, 0.040] | [0.021, 0.027] |
| NCC | 45.92 | < | [0.008, 0.014] | [0.003, 0.009] | [0.012, 0.017] | [0.008, 0.011] |
| CS | 17.07 | < | [0.000, 0.001] | [0.000, 0.001] | [0.001, 0.002] | [0.001, 0.002] |
Table 8.
Effect sizes (Cohen’s d) for fused versus individual configurations and median improvement ranges (fused vs. individual configs). Values indicate large effects.
Table 8.
Effect sizes (Cohen’s d) for fused versus individual configurations and median improvement ranges (fused vs. individual configs). Values indicate large effects.
| Metric | O/10/3 | O/3/10 | R/10/3 | R/3/10 |
|---|
| MI | 4.03 | 2.93 | 4.93 | 5.58 |
| NMI | 2.93 | 1.27 | 4.03 | 4.39 |
| SSIM | 3.10 | 1.78 | 7.17 | 5.28 |
| NCC | 2.00 | 1.21 | 3.30 | 3.53 |
| CS | 1.11 | 0.77 | 2.02 | 1.19 |
| Median improvement range (fused vs. individual configs) |
| MI | 2.87%:9.73% |
| NMI | 2.00%:7.62% |
| SSIM | 1.79%:5.14% |
| NCC | 0.56%:1.43% |
| CS | 0.02%:0.15% |
Table 9.
Empirically validated model complexity metrics for the TbGAN framework under actual training conditions (NVIDIA RTX A2000 12 GB VRAM). All values measured during 50-epoch training on 27 WSIs with 256 × 256 × 4 (RGBA) inputs.
Table 9.
Empirically validated model complexity metrics for the TbGAN framework under actual training conditions (NVIDIA RTX A2000 12 GB VRAM). All values measured during 50-epoch training on 27 WSIs with 256 × 256 × 4 (RGBA) inputs.
| Metric | Value | Clinical Relevance |
|---|
| Total parameters | 31.3 million | Fits within memory constraints of mid-tier clinical GPUs |
| (a) Generator (transformer) | 28.5 million | Captures long-range collagen dependencies |
| (b) Discriminator (CNN) | 2.8 million | Efficient patch-level realism assessment |
| Training time (50 epochs) | 68.3 h | Feasible for offline model development |
| Inference speed per tile | 1.8 s | Real-time capable for diagnostic workflows |
| Peak training memory | 1253 MB | 20.9% of 6 GB GPU capacity |
| Activation memory (inference) | 289 MB | Minimal overhead during deployment |
| Sequence length (tokens) | 1024 | Derived from 8 × 8 patch embedding kernel |
| Attention matrix memory | 384 MB | Dominant memory consumer due to scaling |
| Gradient accumulation steps | 8 | Simulates effective batch size of 8 to stabilize optimization |