Design and Implementation of a Prefetcher in a Key Performance Subsystems of RISC-V Processors
Abstract
1. Introduction
2. Backgroud
2.1. Analysis of Stride Prefetcher
2.2. Analysis of Best-Offset Prefetcher
2.3. Analysis of SMS Prefetcher
3. Proposed Hardware Prefetcher Implementation
3.1. Proposed L1 Dcache Hardware Prefetcher Implementation
3.1.1. Proposed Stride Prefetcher Implementation
3.1.2. Proposed Software Template Prefetcher Implementation
3.1.3. Proposed L1 Dcache Hardware Prefetcher Control Register Configuration
- templata_en: Used to select either the software template prefetch address or the hardware stride prefetch address as the output. The reset value is 0, selecting the hardware stride prefetch address by default.
- pref_en: Used to control the enable signal of the l1_dcache hardware prefetcher. The reset value of this field is 1, meaning the hardware prefetcher is enabled by default. It can be disabled by configuring this bit to 0 via a CSR instruction.
- Cnt_limit: Used to control the threshold for the confidence counter. Each entry in the l1_dcache hardware prefetcher has a confidence counter. When the confidence counter value exceeds the set threshold, data is prefetched into the L1 cache; otherwise, it is only prefetched into the L2 cache. The confidence threshold is set as cnt_limit + 3. Thus, if cnt_limit is not explicitly set, the default threshold is 3.
- H_step: Used to set the number of prefetch strides. When h_step is not set, each prefetch adds one stride, i.e., prefetch_addr = vaddr + stride. When h_step is set, each prefetch adds (h_step + 1) strides, i.e., prefetch_addr = vaddr + (h_step + 1) * stride.
3.2. Proposed L2 Cache Hardware Prefetcher Implementation
3.2.1. Optimization of SMS Prefetcher Implementation
3.2.2. Optimization of Best-Offset Prefetcher Implementation
3.2.3. L2 Cache Prefetcher Control Register Configuration
- Sms_enable: Controls the enable signal for the SMS prefetcher. The reset value of this field is 1, meaning the SMS prefetcher is enabled by default. It can be disabled by setting this bit to 0 via a CSR instruction.
- Delay_q_cycle: Controls the delay time (in cycles) for the Best-Offset prefetcher’s Delay Queue. Tuning this delay helps find the optimal balance between timeliness and coverage for best prefetch performance. The reset value is 60, meaning after reset, missed addresses wait 60 cycles in the Delay Queue before entering the RR Table. This value can be adjusted via CSR instructions for better prefetch performance.
- Max_score: Represents the maximum score threshold. During a learning session, if the current highest score is greater than or equal to this threshold, it indicates a Best Offset has been learned, and the learning session will end, outputting this Best Offset. The reset value is 31, adjustable via CSR instruction.
- Min_score: Represents the minimum score threshold. Upon completion of a BO learning session, if the highest score among the 46 offsets is below this threshold, it is concluded that no suitable prefetch offset for the current program flow was found. In this case, the Best Offset is set to 0, instructing the L2 Cache not to perform prefetching until the next BO learning session completes successfully. The reset value is 10, adjustable via CSR instruction.
- Max_round: Represents the number of learning rounds per session. The Best-Offset prefetcher has a list of 46 offsets. When a miss_addr arrives, it sequentially selects an offset, calculates miss_addr - offset, and checks if the result exists in the RR Table. If found, the score for that offset is incremented. Processing from the first to the last offset constitutes one round. Max_round defines how many rounds comprise one learning session. The reset value is 100, meaning each session learns for 100 rounds. Adjustable via CSR instruction, a valid range is 0–127.
- Fixed_best_offset and Fixed_bo_en: Fixed_bo_en indicates whether the prefetcher should use a fixed Best Offset. When fixed_bo_en is set, the prefetcher uses the value in Fixed_best_offset as the fixed Best Offset. The reset value of Fixed_bo_en is 0 (do not use fixed BO). The reset value of Fixed_best_offset is 1. Can be set via CSR instruction to control whether a fixed value is used for prefetching.
- Bo_clear: BO prefetcher clear. Setting this bit to 1 causes all learning states of the BO prefetcher to be cleared. The reset value is 0. Setting this bit achieves the effect of clearing the BO prefetcher.
- Bo_enable: BO prefetcher enable. The reset value of this bit is 1, meaning the Best-Offset prefetcher is enabled by default. It can be disabled by configuring this bit to 0 via a CSR instruction.
- Pcue: This is a unique field for the sprefetch_cfg register, used to select the BO prefetcher configuration in User (U) mode. When in U mode and this field is set to 0, the prefetcher uses the configuration from sprefetch_cfg. When set to 1, it uses the configuration from uprefetch_cfg.
4. Results of the Hardware Implementation
4.1. Performance Analysis
4.2. Physical Implementation
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| RISC-V | Reduced Instruction Set Computing—Five |
| FPGA | Field-Programmable Gate Array |
| CPU | Central Processing Unit |
| L1 | Level 1 |
| L2 | Level 2 |
| RV64 | RISC-V 64-bit |
| RVA23 | RISC-V Architecture Roadmap 2023 |
| CMOS | Complementary Metal–Oxide–Semiconductor |
| CLB | Configurable Logic Block |
| LUT | Look-Up Tables |
| URAM | Ultra RAM |
| DSP | Digital Signal Processors |
References
- Patterson, D. 50 Years of computer architecture: From the mainframe CPU to the domain-specific tpu and the open RISC-V instruction set. In Proceedings of the 2018 IEEE International Solid-State Circuits Conference—(ISSCC), San Francisco, CA, USA, 11–15 February 2018; pp. 27–31. [Google Scholar] [CrossRef]
- Yang, S.; Shao, L.; Huang, J.; Zou, W. Design and Implementation of Low-Power IoT RISC-V Processor with Hybrid Encryption Accelerator. Electronics 2023, 12, 4222. [Google Scholar] [CrossRef]
- Patterson, D.; Waterman, A. The RISC-V Reader: An Open Architecture Atlas, 1st ed.; Strawberry Canyon: Berkeley, CA, USA, 2017. [Google Scholar]
- Falsafi, B.; Wenisch, T.F. A Primer on Hardware Prefetching; Springer Nature: Berlin/Heidelberg, Germany, 2022. [Google Scholar]
- Tse, J.; Smith, A. CPU cache prefetching: Timing evaluation of hardware implementations. IEEE Trans. Comput. 1998, 47, 509–526. [Google Scholar] [CrossRef][Green Version]
- Chen, J.; Loi, I.; Flamand, E.; Tagliavini, G.; Benini, L.; Rossi, D. Scalable Hierarchical Instruction Cache for Ultralow-Power Processors Clusters. IEEE Trans. Very Large Scale Integr. (VLSI) Syst. 2023, 31, 456–469. [Google Scholar] [CrossRef]
- Wu, Y.; Serrano, M.; Krishnaiyer, R.; Li, W.; Fang, J. Value-Profile Guided Stride Prefetching for Irregular Code. In Proceedings of the Compiler Construction, Grenoble, France, 8–12 April 2002; Horspool, R.N., Ed.; Springer: Berlin/Heidelberg, Germany, 2002; pp. 307–324. [Google Scholar]
- Michaud, P. A Best-Offset Prefetcher. In Proceedings of the 2nd Data Prefetching Championship, Portland, OR, USA, 13 June 2015. [Google Scholar]
- Somogyi, S.; Wenisch, T.F.; Ailamaki, A.; Falsafi, B.; Moshovos, A. Spatial Memory Streaming. SIGARCH Comput. Archit. News 2006, 34, 252–263. [Google Scholar] [CrossRef]
- Michaud, P. Best-offset hardware prefetching. In Proceedings of the 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA), Barcelona, Spain, 12–16 March 2016; pp. 469–480. [Google Scholar] [CrossRef]
- Waterman, A.; Asanovic, K.; SiFive, Inc.; CS Division; EECS Department; University of California, Berkeley. The RISC-V Instruction Set Manual Volume I: Unprivileged Isa; Document Version; RISC-V: Zurich, Switzerland, 2019; Volume 20191213, pp. 1–4. [Google Scholar]
- Waterman, A.; Asanovic, K.; Hauser, J. The RISC-V Instruction Set Manual Volume II: Privileged Architecture; RISC-V Foundation: Zurich, Switzerland, 2019; pp. 1–4. [Google Scholar]
- Chen, T.F.; Baer, J.L. Effective hardware-based data prefetching for high-performance processors. IEEE Trans. Comput. 1995, 44, 609–623. [Google Scholar] [CrossRef]
- Wang, T.; Yu, L.; Zhuang, W. Implementation and optimization of cache data prefetcher based on SPARC processors. In Proceedings of the Fourth International Conference on Algorithms, Microchips, and Network Applications (AMNA 2025), Yangzhou, China, 7–9 March 2025; Taheri, J., Chen, L., Eds.; International Society for Optics and Photonics, SPIE: Bellingham, WA, USA, 2025; Volume 13576, p. 135760L. [Google Scholar] [CrossRef]
- Sutherland, M.; Kannan, A.; Jerger, N.E. Not quite my tempo: Matching prefetches to memory access times. In Proceedings of the Data Prefetching Championship Workshop, Portland, OR, USA, 13 June 2015. [Google Scholar]
- Mohapatra, S.; Panda, B. Drishyam: An Image is Worth a Data Prefetcher. In Proceedings of the 2023 32nd International Conference on Parallel Architectures and Compilation Techniques (PACT), Vienna, Austria, 21–25 October 2023; pp. 51–61. [Google Scholar] [CrossRef]
- Jamet, A.V.; Vavouliotis, G.; Jiménez, D.A.; Alvarez, L.; Casas, M. A Two Level Neural Approach Combining Off-Chip Prediction with Adaptive Prefetch Filtering. In Proceedings of the 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA), Edinburgh, UK, 2–6 March 2024; pp. 528–542. [Google Scholar] [CrossRef]
- McCalpin, J.D. Memory Bandwidth and Machine Balance in Current High Performance Computers. In IEEE Computer Society Technical Committee on Computer Architecture (TCCA) Newsletter; IEEE: Piscataway, NJ, USA, 1995; Volume 2. [Google Scholar]
- Nair, A.A.; John, L.K. Simulation points for SPEC CPU 2006. In Proceedings of the 2008 IEEE International Conference on Computer Design, Lake Tahoe, CA, USA, 12–15 October 2008; pp. 397–403. [Google Scholar] [CrossRef]
- Intel Corporation. Intel® Core™ i3-550 Processor (4M Cache, 3.20 GHz). Available online: https://www.intel.cn/content/www/cn/zh/products/sku/48505/intel-core-i3550-processor-4m-cache-3-20-ghz/specifications.html (accessed on 13 December 2025).
- Intel Corporation. Intel® Celeron® Processor J1900 (2M Cache, up to 2.42 GHz). Available online: https://www.intel.com.tw/content/www/tw/zh/products/sku/78867/intel-celeron-processor-j1900-2m-cache-up-to-2-42-ghz/specifications.html (accessed on 13 December 2025).
- Alibaba DAMO Academy. XuanTie C908. Available online: https://www.xrvm.cn/product/xuantie/C908 (accessed on 13 December 2025).
- Alibaba DAMO Academy. OpenC910 Datasheet. Available online: https://github.com/T-head-Semi/openc910 (accessed on 13 December 2025).
- RISC-V International. RVA23 Profile. Available online: https://lists.riscv.org/g/tech-golden-model/attachment/265/0/rva23-profiles-internal-review-20240321%20.pdf (accessed on 13 December 2025).











| (a) Performance and L1 Data Cache (DC) Metrics | ||||||
| Program | Prefetch | Bandwidth (MB/S) | Perf. Boost | dc_access | dc_access _miss | dc miss rate |
| stream _copy | off | 8081 | 42.08% | 1124991 | 839928 | 74.66% |
| on | 11,481 | 1124983 | 588008 | 52.27% | ||
| stream _scale | off | 5934 | 61.22% | 2125010 | 800728 | 37.68% |
| on | 9566 | 2125014 | 550407 | 25.90% | ||
| stream _add | off | 7520 | 50.29% | 2124985 | 1724361 | 81.15% |
| on | 11,301 | 2124981 | 961775 | 45.26% | ||
| stream _triad | off | 7119 | 32.06% | 3124976 | 1638880 | 52.44% |
| on | 9401 | 3124985 | 996264 | 31.88% | ||
| (b) L2 Cache Metrics Details | ||||||
| Program | Prefetch | l2_request | l2_request _miss | l2 miss rate | ||
| stream _copy | off | 250033 | 164288 | 65.71% | ||
| on | 390762 | 177908 | 45.53% | |||
| stream _scale | off | 250024 | 160212 | 64.08% | ||
| on | 414970 | 178314 | 42.97% | |||
| stream _add | off | 375024 | 332491 | 88.66% | ||
| on | 590056 | 349182 | 59.18% | |||
| stream _triad | off | 375021 | 333398 | 88.90% | ||
| on | 587287 | 353557 | 60.20% | |||
| CPU Name | Year | Micro-Arch. | Clock (GHz) | Stream_ Copy (MB/s) | Stream_ Scale (MB/s) | Stream_ Add (MB/s) | Stream_ Triad (MB/s) |
|---|---|---|---|---|---|---|---|
| Intel i3-550 [20] | 2010 | Nehalem | 3.2 | 8916 | 8703 | 9299 | 9335 |
| Intel Atom J1900 [21] | 2014 | Bay Trail | 2.4 | 7082 | 7051 | 7404 | 7597 |
| XuanTie C908 [22] | 2022 | RISC-V 64 | 2.0 | 5953 | 3477 | 5568 | 4534 |
| XuanTie C910 [23] | 2024 | RISC-V 64 | 2.0 | 8106 | 8115 | 6027 | 6063 |
| HCC75 (this work) | 2025 | RISC-V 64 | 2.5 | 11,481 | 9566 | 11,301 | 9401 |
| Benchmark | Speed-Up |
|---|---|
| 400. perlbench | 1.33% |
| 401. bzip2 | 13.87% |
| 403. gcc | 17.99% |
| 429. mcf | 1.30% |
| 445. gobmk | 2.07% |
| 456. hmmer | 117.46% |
| 458. sjeng | −3.12% |
| 462. libquantum | 32.05% |
| 464. h264ref | 1.15% |
| 471. omnetpp | 187.73% |
| 473. astar | 17.96% |
| 483. xalancbmk | 26.31% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
He, G.; Zhao, Y.; Xiang, Y.; Li, L. Design and Implementation of a Prefetcher in a Key Performance Subsystems of RISC-V Processors. Electronics 2026, 15, 319. https://doi.org/10.3390/electronics15020319
He G, Zhao Y, Xiang Y, Li L. Design and Implementation of a Prefetcher in a Key Performance Subsystems of RISC-V Processors. Electronics. 2026; 15(2):319. https://doi.org/10.3390/electronics15020319
Chicago/Turabian StyleHe, Guoqiang, Yanbo Zhao, Yang Xiang, and Li Li. 2026. "Design and Implementation of a Prefetcher in a Key Performance Subsystems of RISC-V Processors" Electronics 15, no. 2: 319. https://doi.org/10.3390/electronics15020319
APA StyleHe, G., Zhao, Y., Xiang, Y., & Li, L. (2026). Design and Implementation of a Prefetcher in a Key Performance Subsystems of RISC-V Processors. Electronics, 15(2), 319. https://doi.org/10.3390/electronics15020319

