Hybrid Event–Frame Sensing for Human-Perceptual Imaging and Machine Vision
Highlights
- This review presents a sensor-oriented taxonomy of hybrid event–frame sensing architectures and systems, including dual-camera event–frame systems, optically aligned event–frame systems, pixel-level shared hybrid image sensors, stacked CIS–DVS hybrid image sensors, homogeneous-pixel sensing systems, and event-only reconstruction systems.
- It analyzes key sensor specifications—latency, spatial resolution, color fidelity, power consumption, and form factor—and links them to configuration and design choices for hybrid event–frame sensing.
- Stacked CIS–DVS hybrid image sensors are identified as one of the most competitive architectures across the reviewed specifications, offering a strong balance among compact form factor, synchronized event–frame sensing, spatial resolution, power efficiency, and system integration.
- Color fidelity, demosaicing, event-pixel ratio, calibration, and edge-AI deployment remain key challenges for future human-perceptual imaging and machine-vision applications.
Abstract
1. Introduction
2. Design and Taxonomy of Hybrid Event–Frame Sensing Architectures
2.1. Operating Principle of Frame-Based RGB Image Sensors
2.2. Operating Principle of Event-Based Sensors
2.3. Dual-Camera Systems
2.4. Optically Aligned Event–Frame Systems
2.5. Pixel-Level Shared Hybrid Sensors
2.6. Stacked or Co-Integrated Single-Chip Sensors
2.7. Homogeneous-Pixel Computational Hybrid Sensing
–1, F(ti) − F(ti−1) ≤ –TOFF
0, otherwise
2.8. Event-Only Computational Video Reconstruction
3. Sensor Specifications
3.1. Latency
3.2. Spatial Resolution
3.3. Color Fidelity
3.4. Power Consumption
3.5. Form Factor
4. Sensor Configuration and Challenges
4.1. Sensor Configuration
4.2. Open Challenges
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| ADC | Analog-to-Digital Converter |
| AER | Address-Event Representation |
| AI | Artificial Intelligence |
| AIoT | Artificial Intelligence of Things |
| APS | Active Pixel Sensor |
| AR | Augmented Reality |
| VR | Virtual Reality |
| BSI | Backside Illumination |
| CDS | Correlated Double Sampling |
| CFA | Color Filter Array |
| CIS | CMOS Image Sensor |
| CNN | Convolutional Neural Network |
| CMOS | Complementary Metal-Oxide Semiconductor |
| CV | Computer Vision |
| DNN | Deep Neural Network |
| DR | Dynamic Range |
| DVS | Dynamic Vision Sensor |
| EVS | Event Vision Sensor |
| FAR | False Acceptance Rate |
| FD | Floating Diffusion |
| FLOPs | Floating-Point Operations |
| FPN | Fixed-Pattern Noise |
| HDR | High Dynamic Range |
| IMU | Inertial Measurement Unit |
| ISP | Image Signal Processing |
| LPIPS | Learned Perceptual Image Patch Similarity |
| mAP | mean Average Precision |
| MIPI | Mobile Industry Processor Interface |
| MTF | Modulation Transfer Function |
| PPD | Pinned Photodiode |
| PSNR | Peak Signal-to-Noise Ratio |
| QCFA | Quad Color Filter Array |
| RGB | Red, Green, and Blue |
| ROI | Region of Interest |
| RTK GPS | Real-Time Kinematic Global Positioning System |
| SLAM | Simultaneous Localization and Mapping |
| SNR | Signal-to-Noise Ratio |
| SSIM | Structural Similarity Index Measure |
| UAV | Unmanned Aerial Vehicle |
References
- Fossum, E.R. CMOS image sensors: Electronic camera-on-a-chip. IEEE Trans. Electron Devices 1997, 44, 1689–1698. [Google Scholar] [CrossRef] [Scilit]
- Fossum, E.R.; Hondongwa, D.B. A review of the pinned photodiode for CCD and CMOS image sensors. IEEE J. Electron Devices Soc. 2014, 2, 33–43. [Google Scholar] [CrossRef] [Scilit]
- Bigas, M.; Cabruja, E.; Forest, J.; Salvi, J. Review of CMOS image sensors. Microelectron. J. 2006, 37, 433–451. [Google Scholar] [CrossRef] [Scilit]
- Hasinoff, S.W.; Sharlet, D.; Geiss, R.; Adams, A.; Barron, J.T.; Kainz, F.; Chen, J.; Levoy, M. Burst photography for high dynamic range and low-light imaging on mobile cameras. ACM Trans. Graph. 2016, 35, 192. [Google Scholar] [CrossRef] [Scilit]
- Lichtsteiner, P.; Posch, C.; Delbruck, T. A 128 × 128 120 dB 15 μs latency asynchronous temporal contrast vision sensor. IEEE J. Solid-State Circuits 2008, 43, 566–576. [Google Scholar]
- Posch, C.; Matolin, D.; Wohlgenannt, R. A QVGA 143 dB dynamic range frame-free PWM image sensor with lossless pixel-level video compression and time-domain CDS. IEEE J. Solid-State Circuits 2011, 46, 259–275. [Google Scholar] [CrossRef] [Scilit]
- Brandli, C.; Berner, R.; Yang, M.; Liu, S.-C.; Delbruck, T. A 240 × 180 130 dB 3 μs latency global shutter spatiotemporal vision sensor. IEEE J. Solid-State Circuits 2014, 49, 2333–2341. [Google Scholar] [CrossRef] [Scilit]
- Gallego, G.; Delbruck, T.; Orchard, G.; Bartolozzi, C.; Taba, B.; Censi, A.; Leutenegger, S.; Davison, A.J.; Conradt, J.; Daniilidis, K.; et al. Event-based vision: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 154–180. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Delbruck, T.; Linares-Barranco, B.; Culurciello, E.; Posch, C. Activity-Driven, Event-Based Vision Sensors. In Proceedings of the IEEE International Symposium on Circuits and Systems, Paris, France, 30 May–2 June 2010. [Google Scholar]
- Mueggler, E.; Huber, B.; Scaramuzza, D. Event-based, 6-DOF pose tracking for high-speed maneuvers. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, Chicago, IL, USA, 14–18 September 2014. [Google Scholar]
- Vidal, A.R.; Rebecq, H.; Horstschaefer, T.; Scaramuzza, D. UltimateSLAM? Combining events, images, and IMU for robust visual SLAM in HDR and high-speed scenarios. IEEE Robot. Autom. Lett. 2018, 3, 994–1001. [Google Scholar] [CrossRef] [Scilit]
- Mueggler, E.; Rebecq, H.; Gallego, G.; Delbruck, T.; Scaramuzza, D. The event-camera dataset and simulator: Event-based data for pose estimation, visual odometry, and SLAM. Int. J. Robot. Res. 2017, 36, 142–149. [Google Scholar] [CrossRef] [Scilit]
- Amir, A.; Taba, B.; Berg, D.; Melano, T.; McKinstry, J.; di Nolfo, C.; Nayak, T.; Andreopoulos, A.; Garreau, G.; Mendoza, M.; et al. A low power, fully event-based gesture recognition system. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017. [Google Scholar]
- Maqueda, A.I.; Loquercio, A.; Gallego, G.; García, N.; Scaramuzza, D. Event-based vision meets deep learning on steering prediction for self-driving cars. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018. [Google Scholar]
- Nozaki, Y.; Delbruck, T. Temperature and parasitic photocurrent effects in dynamic vision sensors. IEEE Trans. Electron Devices 2017, 64, 3239–3245. [Google Scholar] [CrossRef] [Scilit]
- Graca, R.; Delbruck, T. Unraveling the paradox of intensity-dependent DVS pixel noise. arXiv 2021, arXiv:2109.08640. [Google Scholar]
- Lin, S.; Ma, Y.; Guo, Z.; Wen, B. DVS-Voltmeter: Stochastic process-based event simulator for dynamic vision sensors. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022. [Google Scholar]
- Rebecq, H.; Ranftl, R.; Koltun, V.; Scaramuzza, D. High speed and high dynamic range video with an event camera. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 1964–1980. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pan, L.; Hartley, R.; Scheerlinck, C.; Liu, M.; Yu, X.; Dai, Y. High frame rate video reconstruction based on an event camera. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 2519–2533. [Google Scholar] [PubMed]
- Tulyakov, S.; Gehrig, D.; Georgoulis, S.; Erbach, J.; Gehrig, M.; Li, Y.; Scaramuzza, D. Time Lens: Event-based video frame interpolation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021. [Google Scholar]
- de Tournemire, P.; Nitti, D.; Perot, E.; Migliore, D.; Sironi, A. A large scale event-based detection dataset for automotive. arXiv 2020, arXiv:2001.08499. [Google Scholar]
- Gehrig, M.; Aarents, W.; Gehrig, D.; Scaramuzza, D. DSEC: A stereo event camera dataset for driving scenarios. IEEE Robot. Autom. Lett. 2021, 6, 4947–4954. [Google Scholar] [CrossRef] [Scilit]
- Kodama, K.; Sato, Y.; Yorikado, Y.; Berner, R.; Mizoguchi, K.; Miyazaki, T.; Tsukamoto, M.; Matoba, Y.; Shinozaki, H.; Niwa, A.; et al. 1.22 μm 35.6 Mpixel RGB hybrid event-based vision sensor with 4.88 μm-pitch event pixels and up to 10K event frame rate by adaptive control on event sparsity. In Proceedings of the IEEE International Solid-State Circuits Conference, San Francisco, CA, USA, 19–23 February 2023. [Google Scholar]
- Lu, Y.; Messikommer, N.; Xu, X.; Chen, L.; Chen, Y.; Zubić, N.; Scaramuzza, D.; Xiong, H. Hybrid event-frame sensors: Modeling, calibration, and simulation. arXiv 2025, arXiv:2511.18037. [Google Scholar]
- Delbruck, T.; Lang, M. Robotic goalie with 3 ms reaction time at 4% CPU load using event-based dynamic vision sensor. Front. Neurosci. 2013, 7, 223. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rebecq, H.; Gehrig, D.; Scaramuzza, D. ESIM: An open event camera simulator. In Proceedings of the Conference on Robot Learning, Zürich, Switzerland, 29–31 October 2018. [Google Scholar]
- Hu, Y.; Liu, S.-C.; Delbruck, T. v2e: From video frames to realistic DVS events. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Nashville, TN, USA, 19–25 June 2021. [Google Scholar]
- Davies, M.; Srinivasa, N.; Lin, T.-H.; Chinya, G.; Cao, Y.; Choday, S.H.; Dimou, G.; Joshi, P.; Imam, N.; Jain, S.; et al. Loihi: A neuromorphic manycore processor with on-chip learning. IEEE Micro 2018, 38, 82–99. [Google Scholar] [CrossRef] [Scilit]
- Merolla, P.A.; Arthur, J.V.; Alvarez-Icaza, R.; Cassidy, A.S.; Sawada, J.; Akopyan, F.; Jackson, B.L.; Imam, N.; Guo, C.; Nakamura, Y.; et al. A million spiking-neuron integrated circuit with a scalable communication network and interface. Science 2014, 345, 668–673. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- El Gamal, A.; Eltoukhy, H. CMOS image sensors. IEEE Circuits Devices Mag. 2005, 21, 6–20. [Google Scholar] [CrossRef] [Scilit]
- Gunturk, B.K.; Altunbasak, Y.; Mersereau, R.M. Color plane interpolation using alternating projections. IEEE Trans. Image Process. 2002, 11, 997–1013. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, X.; Gunturk, B.; Zhang, L. Image demosaicing: A systematic survey. In Proceedings of the SPIE Visual Communications and Image Processing, San Jose, CA, USA, 27–31 January 2008. [Google Scholar]
- Menon, D.; Calvagno, G. Color image demosaicking: An overview. Signal Process. Image Commun. 2011, 26, 518–533. [Google Scholar] [CrossRef] [Scilit]
- Heide, F.; Steinberger, M.; Tsai, Y.-T.; Rouf, M.; Pajak, D.; Reddy, D.; Gallo, O.; Liu, J.; Heidrich, W.; Egiazarian, K.; et al. FlexISP: A flexible camera image processing framework. ACM Trans. Graph. 2014, 33, 231. [Google Scholar]
- Kim, Y.; Jung, Y.; Sul, H.; Koh, K. A 1/1.12-inch 1.4 μm-pitch 50Mpixel 65/28nm stacked CMOS image sensor using multiple sampling. In Proceedings of the IEEE International Symposium on Circuits and Systems, Monterey, CA, USA, 21–25 May 2023. [Google Scholar]
- Park, P.K.J.; Kim, J.; Ko, J. Motion blur-free high-speed hybrid image sensing. Sensors 2025, 25, 7496. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Son, B.; Suh, Y.; Kim, S.; Jung, H.; Kim, J.S.; Shin, C.; Park, K.; Lee, K.; Park, J.; Woo, J.; et al. A 640 × 480 dynamic vision sensor with a 9 μm pixel and 300 Meps address-event representation. In Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 5–9 February 2017. [Google Scholar]
- Posch, C.; Serrano-Gotarredona, T.; Linares-Barranco, B.; Delbruck, T. Retinomorphic event-based vision sensors: Bioinspired cameras with spiking output. Proc. IEEE 2014, 102, 1470–1484. [Google Scholar] [CrossRef] [Scilit]
- Boahen, K.A. Point-to-point connectivity between neuromorphic chips using address events. IEEE Trans. Circuits Syst. II Analog Digit. Signal Process. 2000, 47, 416–434. [Google Scholar] [CrossRef] [Scilit]
- Suh, Y.; Choi, S.; Ito, M.; Kim, J.; Lee, Y.; Seo, J.; Jung, H.; Yeo, D.H.; Namgung, S.; Bong, K.; et al. A 1280 × 960 dynamic vision sensor with a 4.95-μm pixel pitch and motion artifact minimization. In Proceedings of the IEEE International Symposium on Circuits and Systems, Seville, Spain, 10–21 October 2020. [Google Scholar]
- Gehrig, D.; Loquercio, A.; Derpanis, K.G.; Scaramuzza, D. End-to-end learning of representations for asynchronous event-based data. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
- Zhu, A.Z.; Thakur, D.; Özaslan, T.; Pfrommer, B.; Kumar, V.; Daniilidis, K. The multivehicle stereo event camera dataset: An event camera dataset for 3D perception. IEEE Robot. Autom. Lett. 2018, 3, 2032–2039. [Google Scholar] [CrossRef] [Scilit]
- Binas, J.; Neil, D.; Liu, S.-C.; Delbruck, T. DDD17: End-to-end DAVIS driving dataset. arXiv 2017, arXiv:1711.01458. [Google Scholar]
- Tomy, A.; Paigwar, A.; Mann, K.S.; Renzaglia, A.; Laugier, C. Fusing event-based and RGB camera for robust object detection in adverse conditions. In Proceedings of the IEEE International Conference on Robotics and Automation, Philadelphia, PA, USA, 23–27 May 2022. [Google Scholar]
- Rebecq, H.; Horstschaefer, T.; Scaramuzza, D. Real-time visual-inertial odometry for event cameras using keyframe-based nonlinear optimization. In Proceedings of the British Machine Vision Conference, London, UK, 4–7 September 2017. [Google Scholar]
- Hidalgo-Carrió, J.; Gallego, G.; Scaramuzza, D. Event-aided direct sparse odometry. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022. [Google Scholar]
- Wang, K.; Liu, S.; Shi, H.; Shi, L.; Chen, H. Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026. [Google Scholar]
- Cao, H.; Zhang, Z.; Xia, Y.; Li, X.; Xia, J.; Chen, G.; Knoll, A. Embracing events and frames with hierarchical feature refinement network for object detection. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Scheerlinck, C.; Rebecq, H.; Gehrig, D.; Barnes, N.; Mahony, R.; Scaramuzza, D. Fast image reconstruction with an event camera. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision, Snowmass Village, CO, USA, 1–5 March 2020. [Google Scholar]
- Pan, L.; Scheerlinck, C.; Yu, X.; Hartley, R.; Liu, M.; Dai, Y. Bringing a blurry frame alive at high frame-rate with an event camera. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
- Gehrig, D.; Gehrig, M.; Hidalgo-Carrió, J.; Scaramuzza, D. Video to events: Recycling video datasets for event cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020. [Google Scholar]
- Sun, L.; Sakaridis, C.; Liang, J.; Jiang, Q.; Yang, K.; Sun, P.; Ye, Y.; Wang, K.; Van Gool, L. Event-based fusion for motion deblurring with cross-modal attention. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022. [Google Scholar]
- Wang, Z.; Hamann, F.; Chaney, K.; Jiang, W.; Gallego, G.; Daniilidis, K. Event-based continuous color video decompression from single frames. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Nashville, TN, USA, 11–15 June 2025. [Google Scholar]
- Kim, T.; Jeong, J.; Cho, H.; Jeong, Y.; Yoon, K.-J. Towards real-world event-guided low-light video enhancement and deblurring. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Guo, M.; Chen, S.; Gao, Z.; Yang, W.; Bartkovjak, P.; Qin, Q.; Hu, X.; Zhou, D.; Uchiyama, M.; Fukuoka, S.; et al. A Three-Wafer-Stacked Hybrid 15-MPixel CIS + 1-MPixel EVS With 4.6-GEvent/s Readout, In-Pixel TDC, and On-Chip ISP and ESP Function. IEEE J. Solid-State Circuits 2023, 58, 2955–2964. [Google Scholar] [CrossRef] [Scilit]
- Guo, M.; Chen, S.; Gao, Z.; Yang, W.; Bartkovjak, P.; Qin, Q.; Hu, X.; Zhou, D.; Uchiyama, M.; Fukuoka, S.; et al. A 3-Wafer-Stacked Hybrid 15MPixel CIS + 1MPixel EVS with 4.6GEvent/s Readout, In-Pixel TDC and On-Chip ISP and ESP Function. In Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 19–23 February 2023. [Google Scholar]
- Finateu, T.; Niwa, A.; Matolin, D.; Tsuchimoto, K.; Mascheroni, A.; Reynaud, E.; Mostafalu, P.; Brady, F.; Chotard, L.; LeGoff, F.; et al. A 1280 × 720 Back-Illuminated Stacked Temporal Contrast Event-Based Vision Sensor With 4.86 μm Pixels, 1.066 GEPS Readout, Programmable Event-Rate Controller and Compressive Data-Formatting Pipeline. In Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA, 16–20 February 2020. [Google Scholar]
- Kondo, T.; Takemoto, Y.; Kobayashi, K.; Tsukimura, M.; Takazawa, N.; Kato, H.; Suzuki, S.; Aoki, J.; Saito, H.; Gomi, Y.; et al. A 3D Stacked CMOS Image Sensor With 16Mpixel Global-Shutter Mode and 2Mpixel 10,000 fps Mode Using 4 Million Interconnections. In Proceedings of the IEEE Symposium on VLSI Circuits, Kyoto, Japan, 16–18 June 2015. [Google Scholar]
- Hwang, J.H.; Kim, J.H.; Kim, S.; Lee, J.; Park, J. A Numerical Method of Aligning the Optical Stacks for All Pixels. Sensors 2023, 23, 702. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, J.; Kim, S.; Kim, J.; Kim, J.; Paik, J. Crosstalk Correction for Color Filter Array Image Sensors Based on Lp-Regularized Multi-Channel Deconvolution. Sensors 2022, 22, 4285. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xhakoni, A.; Theuwissen, A.J.P.; Charbon, E. A 0.5 MP, 3D-Stacked, Voltage-Domain Global Shutter CMOS Image Sensor. Sensors 2023, 23, 9448. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vivet, P.; Thonnart, Y.; Pillonnet, G.; Moritz, C.A.; Bernard, C.; Biswas, S.; Clermidy, F.; Durupt, J.; Lefevre, L.; Lemaire, R.; et al. Advanced 3D Technologies and Architectures for 3D Smart Image Sensors. In Proceedings of the Design, Automation & Test in Europe Conference & Exhibition (DATE), Florence, Italy, 25–29 March 2019. [Google Scholar]
- Inoue, M.; Hara, T.; Tani, T.; Kondoh, Y.; Hirose, S.; Namiki, A. Motion-Blur-Free High-Speed Video Shooting Using a High-Speed Mirror Drive System. Sensors 2017, 17, 2483. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rengarajan, V.; Zhao, S.; Zhen, R.; Glotzbach, J.; Sheikh, H.; Sankaranarayanan, A.C. Photosequencing of Motion Blur Using Short and Long Exposures. arXiv 2019, arXiv:1912.06102. [Google Scholar]
- Nguyen, C.M.; Martel, J.N.P.; Wetzstein, G. Learning Spatially Varying Pixel Exposures for Motion Deblurring. arXiv 2022, arXiv:2204.07267. [Google Scholar]
- Yang, D.; Koskinen, S.; Kämäräinen, J.-K. Active Short-Long Exposure Deblurring. In Proceedings of the International Conference on Pattern Recognition, Montreal, QC, Canada, 21–25 August 2022. [Google Scholar]
- Youn, S.J.; Kim, S.; Choi, J.; Lee, S.; Kim, J.; Kim, H.; Park, J. Design of Low-Noise CMOS Image Sensor Using a Hybrid-Correlated Multiple Sampling Technique with Adaptive Dual-Gain Analog-to-Digital Converter. Sensors 2023, 23, 9551. [Google Scholar] [PubMed]
- Agarwal, A.; Hansrani, J.; Bagwell, S.; Rytov, O.; Shah, V.; Ong, K.L.; Blerkom, D.V.; Bergey, J.; Kumar, N.; Lu, T.; et al. A 316MP, 120FPS, High Dynamic Range CMOS Image Sensor for Next Generation Immersive Displays. Sensors 2023, 23, 8383. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, J.; Yao, Y.; Han, G.; Li, X.; Yang, C.; Xu, Z. Compact All-CMOS Spatiotemporal Compressive Sensing Video Camera with Pixel-Wise Coded Exposure. Opt. Express 2016, 24, 9013–9024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, H.; Park, D.; Jeong, W.; Kim, K.; Je, H.; Ryu, D.; Chun, S.Y. Efficient Unified Demosaicing for Bayer and Non-Bayer Patterned Image Sensors. arXiv 2023, arXiv:2307.10667. [Google Scholar]
- Reinbacher, C.; Graber, G.; Pock, T. Real-Time Intensity-Image Reconstruction for Event Cameras Using Manifold Regularisation. In Proceedings of the British Machine Vision Conference, York, UK, 19–22 September 2016. [Google Scholar]
- Munda, G.; Reinbacher, C.; Pock, T. Real-Time Intensity-Image Reconstruction for Event Cameras Using Manifold Regularisation. Int. J. Comput. Vis. 2018, 126, 1381–1393. [Google Scholar] [CrossRef] [Scilit]
- Rebecq, H.; Ranftl, R.; Koltun, V.; Scaramuzza, D. Events-to-Video: Bringing Modern Computer Vision to Event Cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
- Weng, W.; Zhang, Y.; Xiong, Z. Event-Based Video Reconstruction Using Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021. [Google Scholar]
- Qu, Q.; Shen, Y.; Chen, X.; Chung, Y.Y.; Liu, T. E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning. Proc. AAAI Conf. Artif. Intell. 2024, 38, 4632–4640. [Google Scholar] [CrossRef] [Scilit]
- Cadena, P.R.G.; Qian, Y.; Wang, C.; Yang, M. Sparse-E2VID: A Sparse Convolutional Model for Event-Based Video Reconstruction Trained with Real Event Noise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, BC, Canada, 18–22 June 2023. [Google Scholar]
- Ercan, B.; Eker, O.; Saglam, C.; Erdem, A.; Erdem, E. HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks. arXiv 2023, arXiv:2305.06382. [Google Scholar]
- Wang, Z.; Lu, Y.; Wang, L. Revisit Event Generation Model: Self-Supervised Learning of Event-to-Video Reconstruction with Implicit Neural Representations. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Xu, C.; Zhou, H.; Chen, L.; Chen, H.; Hu, Z.Z.; Lu, Z.; Zhou, Y.; Chung, V.; Qu, Q.; Cai, W. A Survey of 3D Reconstruction with Event Cameras. Comput. Vis. Media 2026, 1–37. [Google Scholar] [CrossRef] [Scilit]
- Muthusamy, R.; Ayyad, A.; Halwani, M.; Swart, D.; Gan, D.; Seneviratne, L.; Zweiri, Y. Neuromorphic Eye-in-Hand Visual Servoing. IEEE Access 2021, 9, 55853–55870. [Google Scholar] [CrossRef] [Scilit]
- Loch, A.; Haessig, G.; Vincze, M. Event-Based High-Speed Low-Latency Fiducial Marker Tracking. arXiv 2021, arXiv:2110.05819. [Google Scholar]
- Monforte, M.; Gava, L.; Iacono, M.; Glover, A.; Bartolozzi, C. Fast Trajectory End-Point Prediction with Event Cameras for Reactive Robot Control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, BC, Canada, 18–22 June 2023. [Google Scholar]
- Angelopoulos, A.N.; Martel, J.N.P.; Kohli, A.P.S.; Conradt, J.; Wetzstein, G. Event-Based Near-Eye Gaze Tracking Beyond 10,000 Hz. IEEE Trans. Vis. Comput. Graph. 2021, 27, 2577–2586. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chamorro, W.; Andrade-Cetto, J.; Solà, J. High Speed Event Camera Tracking. arXiv 2020, arXiv:2010.02771. [Google Scholar]
- Everding, L.; Conradt, J. Low-Latency Line Tracking Using Event-Based Dynamic Vision Sensors. Front. Neurorobot. 2018, 12, 4. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, D.; Parra, A.; Latif, Y.; Chen, B.; Chin, T.-J.; Reid, I. Asynchronous Optimisation for Event-Based Visual Odometry. arXiv 2022, arXiv:2203.01037. [Google Scholar]
- Hamara, A.; Kilpatrick, B.; Baratta, A.; Kofink, B.; Freeman, A.C. Low-Latency Scalable Streaming for Event-Based Vision. arXiv 2024, arXiv:2412.07889. [Google Scholar]
- Köhler, S.; Lovisotto, G.; Birnbach, S.; Baker, R.; Martinovic, I. They See Me Rollin’: Inherent Vulnerability of the Rolling Shutter in CMOS Image Sensors. arXiv 2021, arXiv:2101.10011. [Google Scholar]
- Wang, Z.; Ji, X.; Huang, J.-B.; Satoh, S.; Zhou, X.; Zheng, Y. Neural Global Shutter: Learn to Restore Video from a Rolling Shutter Camera with Global Reset Feature. arXiv 2022, arXiv:2204.00974. [Google Scholar]
- Gallego, G.; Rebecq, H.; Scaramuzza, D. A Unifying Contrast Maximization Framework for Event Cameras, with Applications to Motion, Depth, and Optical Flow Estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018. [Google Scholar]
- Taverni, G.; Moeys, D.P.; Li, C.; San Segundo Bello, D.; Delbruck, T. Front and Back Illuminated Dynamic and Active Pixel Vision Sensors Comparison. IEEE Trans. Circuits Syst. II Express Briefs 2018, 65, 677–681. [Google Scholar] [CrossRef] [Scilit]
- Ramesh, B.; Yang, H.; Orchard, G.; Le Thi, N.A.; Zhang, S.; Xiang, C. DART: Distribution Aware Retinal Transform for Event-Based Cameras. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 2767–2780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Beck, M.; Maier, G.; Hinz, G.; Beyerer, J. An Extended Modular Processing Pipeline for Event-Based Vision in Automatic Visual Inspection. Sensors 2021, 21, 6143. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, T.; Catrysse, P.B.; El Gamal, A.; Wandell, B.A. How Small Should Pixel Size Be? In Proceedings of the SPIE 3965, Sensors and Camera Systems for Scientific, Industrial, and Digital Photography Applications, San Jose, CA, USA, 22–28 January 2000. [Google Scholar]
- Farrell, J.; Xiao, F.; Kavusi, S. Resolution and Light Sensitivity Tradeoff with Pixel Size. In Proceedings of the SPIE 6069, Digital Photography II, San Jose, CA, USA, 15–19 January 2006. [Google Scholar]
- Chen, X.; George, N.; Agranov, G.; Liu, C.; Gravelle, B. Sensor Modulation Transfer Function Measurement Using Band-Limited Laser Speckle. Opt. Express 2008, 16, 20047–20059. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Williams, D.; Burns, P.D. Diagnostics for Digital Capture Using MTF. In Proceedings of the IS&T Archiving Conference, San Antonio, TX, USA, 20–23 April 2001. [Google Scholar]
- Mostafavi, I.S.M.; Choi, J.; Yoon, K.-J. Learning to Super Resolve Intensity Images from Events. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020. [Google Scholar]
- Wang, L.; Kim, T.-K.; Yoon, K.-J. EventSR: From Asynchronous Events to Image Reconstruction, Restoration, and Super-Resolution via End-to-End Adversarial Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 14–19 June 2020. [Google Scholar]
- Li, H.; Li, G.; Liu, H.; Shi, L. Super-Resolution of Spatiotemporal Event-Stream Image Captured by the Asynchronous Temporal Contrast Vision Sensor. arXiv 2018, arXiv:1802.02398. [Google Scholar]
- Xiao, Z.; Kai, D.; Zhang, Y.; Zha, Z.; Sun, X.; Xiong, Z. Event-Adapted Video Super-Resolution. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Agranov, G.; Berezin, V.; Tsai, R.H. Crosstalk and Microlens Study in a Color CMOS Image Sensor. IEEE Trans. Electron Devices 2003, 50, 4–11. [Google Scholar] [CrossRef] [Scilit]
- Huo, Y.; Fesenmaier, C.C.; Catrysse, P.B. Microlens Performance Limits in Sub-2 μm Pixel CMOS Image Sensors. Opt. Express 2010, 18, 5861–5872. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Blockstein, L.; Yadid-Pecht, O. Crosstalk Quantification, Analysis, and Trends in CMOS Image Sensors. Appl. Opt. 2010, 49, 4483–4488. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khabir, M.; Alaibakhsh, H.; Karami, M.A. Electrical Crosstalk Analysis in a Pinned Photodiode CMOS Image Sensor Array. Appl. Opt. 2021, 60, 9640–9650. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Takahashi, S.; Huang, Y.-M.; Sze, J.-J.; Wu, T.-T.; Guo, F.-S.; Hsu, W.-C.; Tseng, T.-H.; Liao, K.; Kuo, C.-C.; Chen, T.-H.; et al. A 45 nm Stacked CMOS Image Sensor Process Technology for Submicron Pixel. Sensors 2017, 17, 2816. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gnanasambandam, A.; Elgendy, O.; Ma, J.; Chan, S.H. Megapixel Photon-Counting Color Imaging Using Quanta Image Sensor. Opt. Express 2019, 27, 17298–17310. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ma, J.; Anzagira, L.; Fossum, E.R. A 1 Megapixel Quanta Image Sensor Jot Device with Sub-0.3e− Read Noise and Photon Counting Capability. IEEE Electron Device Lett. 2017, 38, 1141–1144. [Google Scholar]
- Mahato, S.B.; De Ridder, J.; Meynants, G.; Raskin, G.; Van Winckel, H. Measuring Intra-Pixel Sensitivity Variations of a CMOS Image Sensor. arXiv 2018, arXiv:1805.01843. [Google Scholar]
- Marinelli, R.; Della Corte, F.G.; De Nicola, F.; Rendina, I. Optical Performances of Lensless Sub-2 Micron Pixel for CMOS Image Sensors. Prog. Electromagn. Res. B 2011, 31, 265–281. [Google Scholar] [CrossRef] [Scilit]
- Arfin, R.; Niegemann, J.; McGuire, D.; Bakr, M.H. Adjoint-Assisted Shape Optimization of Microlenses for CMOS Image Sensors. Sensors 2024, 24, 7693. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- ISO 17321-1:2006; Graphic Technology and Photography—Colour Characterisation of Digital Still Cameras (DSCs)—Part 1: Stimuli, Metrology and Test Procedures. ISO: Geneva, Switzerland, 2006.
- CIE 015:2018; Colorimetry, 4th Ed. CIE Central Bureau: Vienna, Austria, 2018.
- Luo, M.R.; Cui, G.; Rigg, B. The Development of the CIE 2000 Colour-Difference Formula: CIEDE2000. Color Res. Appl. 2001, 26, 340–350. [Google Scholar] [CrossRef] [Scilit]
- Vázquez-Corral, J.; Connah, D.; Bertalmío, M. Perceptual Color Characterization of Cameras. Sensors 2014, 14, 23205–23229. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Karaimer, H.C.; Brown, M.S. Improving Color Reproduction Accuracy on Cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018. [Google Scholar]
- Jiang, J.; Liu, D.; Gu, J.; Süsstrunk, S. What Is the Space of Spectral Sensitivity Functions for Digital Color Cameras? In Proceedings of the IEEE Workshop on Applications of Computer Vision, Clearwater Beach, FL, USA, 15–17 January 2013. [Google Scholar]
- Han, S.; Matsushita, Y.; Sato, I.; Okabe, T.; Sato, Y. Camera Spectral Sensitivity Estimation from a Single Image under Unknown Illumination by Using Fluorescence. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 16–21 June 2012. [Google Scholar]
- Parmar, M.; Reeves, S.J. Selection of Optimal Spectral Sensitivity Functions for Color Filter Arrays. IEEE Trans. Image Process. 2010, 19, 3190–3203. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hirakawa, K.; Wolfe, P.J. Spatio-Spectral Color Filter Array Design for Optimal Image Recovery. IEEE Trans. Image Process. 2008, 17, 1876–1890. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tian, Q.; Lansel, S.; Farrell, J.; Wandell, B.A. Automating the Design of Image Processing Pipelines for Novel Color Filter Arrays. In Proceedings of the SPIE 8660, Digital Photography IX, Burlingame, CA, USA, 3–7 February 2013. [Google Scholar]
- Yu, X.; Tanaka, M.; Monno, Y.; Okutomi, M. Joint Design of Camera Spectral Sensitivity and Color Correction Matrix with Noise Consideration. Electron. Imaging 2024, 36, 286-1–286-6. [Google Scholar] [CrossRef] [Scilit]
- Scheerlinck, C.; Rebecq, H.; Stoffregen, T.; Barnes, N.; Mahony, R.; Scaramuzza, D. CED: Color Event Camera Dataset. arXiv 2019, arXiv:1904.10772. [Google Scholar]
- Cohen, K.; Hershko, O.; Levy, H.; Mendlovic, D.; Raviv, D. Illumination-Based Color Reconstruction for the Dynamic Vision Sensor. Sensors 2023, 23, 8327. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pini, S.; Borghi, G.; Vezzani, R. Learn to See by Events: Color Frame Synthesis from Event and RGB Cameras. arXiv 2018, arXiv:1812.02041. [Google Scholar]
- Rudnev, V.; Elgharib, M.; Theobalt, C.; Golyanik, V. EventNeRF: Neural Radiance Fields from a Single Colour Event Camera. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023. [Google Scholar]
- Finlayson, G.D.; Qiu, G.; Qiu, S. Designing Color Filters That Make Cameras More Colorimetric. IEEE Trans. Image Process. 2020, 29, 8534–8544. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, J.; Lin, S.; Li, Y.; Kang, S.B. High-Quality Color Image Reconstruction from RGBW Color Filter Array. In Proceedings of the IEEE International Conference on Image Processing, Melbourne, VIC, Australia, 15–18 September 2013. [Google Scholar]
- Jun, J. A Comprehensive Methodology for Optimizing Read-Out Timing and Reference DAC Offset in High Frame Rate Image Sensing Systems. Sensors 2023, 23, 7048. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Choo, H.S.; Youn, D.-H.; Choi, H.; Kim, G.Y.; Kim, S.Y. The Design of a Low-Noise CMOS Image Sensor Using a Hybrid Single-Slope Analog-to-Digital Converter. Sensors 2024, 24, 8131. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, D.; Song, M.; Choe, B.; Kim, S.Y. A Multi-Resolution Mode CMOS Image Sensor with a Novel Two-Step Single-Slope ADC for Intelligent Surveillance Systems. Sensors 2017, 17, 1497. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ramesh, M.; Rossi, D.; Lecca, M.; Gottardi, M.; Farella, E.; Benini, L. An Event-Driven Ultra-Low-Power Smart Visual Sensor. IEEE Sens. J. 2016, 16, 5344–5353. [Google Scholar] [CrossRef] [Scilit]
- Ramesh, B.; Ussa, A.; Della Vedova, L.; Yang, H.; Orchard, G. PCA-RECT: An Energy-Efficient Object Detection Approach for Event Cameras. In Proceedings of the Asian Conference on Computer Vision Workshops, Perth, Australia, 2–6 December 2018; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2019; Volume 11367. [Google Scholar]
- Ussa, A.; Della Vedova, L.; Padala, V.R.; Singla, D.; Acharya, J.; Lei, C.Z.; Orchard, G.; Basu, A.; Ramesh, B. A Low-Power End-to-End Hybrid Neuromorphic Framework for Surveillance Applications. arXiv 2019, arXiv:1910.09806. [Google Scholar]
- Paissan, F.; Gottardi, M.; Farella, E. Enabling Energy Efficient Machine Learning on an Ultra-Low-Power Vision Sensor for IoT. arXiv 2021, arXiv:2102.01340. [Google Scholar]
- Oike, Y.; El Gamal, A. CMOS Image Sensor with Per-Column ΣΔ ADC and Programmable Compressed Sensing. IEEE J. Solid-State Circuits 2013, 48, 318–328. [Google Scholar] [CrossRef] [Scilit]
- Couniot, N.; Streel, G.; de Botman, F.; Lusala, K.A.; Flandre, D.; Bol, D. A 65 nm 0.5 V DPS CMOS Image Sensor with 17 pJ/Frame.Pixel and 42 dB Dynamic Range for Ultra-Low-Power SoCs. IEEE J. Solid-State Circuits 2015, 50, 2419–2430. [Google Scholar] [CrossRef] [Scilit]
- Choi, J.; Park, S.; Cho, J.; Yoon, E. A 3.4-μW Object-Adaptive CMOS Image Sensor with Embedded Feature Extraction Algorithm for Motion-Triggered Object-of-Interest Imaging. IEEE J. Solid-State Circuits 2014, 49, 289–300. [Google Scholar] [CrossRef] [Scilit]
- Kim, B.; Matthias, T.; Kreindl, G.; Dragoi, V.; Wimplinger, M.; Lindner, P. Advances in Wafer Level Processing and Integration for CIS Module Manufacturing. Int. Symp. Microelectron. 2010, 2010, 000378–000384. [Google Scholar] [CrossRef] [Scilit]
- Han, H.; Kriman, M.; Boomgarden, M. Wafer Level Camera Technology—From Wafer Level Packaging to Wafer Level Integration. In Proceedings of the 2010 11th International Conference on Electronic Packaging Technology & High Density Packaging, Xi’an, China, 16–19 August 2010. [Google Scholar]
- Wuu, S.-G.; Chen, H.-L.; Chien, H.-C.; Enquist, P.; Guidash, R.M.; McCarten, J. A Review of 3-Dimensional Wafer Level Stacked Backside Illuminated CMOS Image Sensor Process Technologies. IEEE Trans. Electron Devices 2022, 69, 2766–2778. [Google Scholar] [CrossRef] [Scilit]
- Asif, M.S.; Ayremlou, A.; Sankaranarayanan, A.; Veeraraghavan, A.; Baraniuk, R.G. FlatCam: Thin, Bare-Sensor Cameras Using Coded Aperture and Computation. IEEE Trans. Comput. Imaging 2017, 3, 384–397. [Google Scholar] [CrossRef] [Scilit]
- Antipa, N.; Kuo, G.; Heckel, R.; Mildenhall, B.; Bostan, E.; Ng, R.; Waller, L. DiffuserCam: Lensless Single-Exposure 3D Imaging. Optica 2018, 5, 1–9. [Google Scholar] [CrossRef] [Scilit]
- Gissibl, T.; Thiele, S.; Herkommer, A.; Giessen, H. Two-Photon Direct Laser Writing of Ultracompact Multi-Lens Objectives. Nat. Photonics 2016, 10, 554–560. [Google Scholar] [CrossRef] [Scilit]
- Lin, R.; Tsai, D.P. Ultracompact wide-FOV near-infrared camera with a wafer-level manufactured meta-aspheric lens. eLight 2026, 6, 19. [Google Scholar] [CrossRef] [Scilit]
- Stork, D.G.; Gill, P.R. Lensless Ultra-Miniature CMOS Computational Imagers and Sensors. In Proceedings of the SENSORCOMM 2013, The Seventh International Conference on Sensor Technologies and Applications, Barcelona, Spain, 25–31 August 2013. [Google Scholar]
- Zoberbier, M.; Hansen, S.; Hennemeyer, M.; Tonnies, D.; Zoberbier, R.; Brehm, M.; Kraft, A.; Eisner, M.; Völkel, R. Wafer Level Cameras—Novel Fabrication and Packaging Technologies. In Proceedings of the International Image Sensor Workshop, Bergen, Norway, 25–28 June 2009. [Google Scholar]
- Dierickx, B.; Meynants, G.; Scheffer, D. Near 100% Fill Factor CMOS Active Pixels. In Proceedings of the IEEE CCD & Advanced Image Sensors Workshop, Bruges, Belgium, 5–7 June 1997. [Google Scholar]
- Wang, Z.; Pan, L.; Ng, Y.; Zhuang, Z.; Mahony, R. Stereo Hybrid Event-Frame (SHEF) Cameras for 3D Perception. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, Prague, Czech Republic, 27 September–1 October 2021. [Google Scholar]
- Ryoo, W.; Nam, G.; Hyun, J.-S.; Kim, S. Event Fusion Photometric Stereo Network. Neural Netw. 2023, 168, 521–530. [Google Scholar]
- Bonazzi, P.; Vogt, C.; Jost, M.; Qin, H.; Khacef, L.; Paredes-Vallés, F.; Magno, M. RGB-Event Fusion with Self-Attention for Collision Prediction. arXiv 2025, arXiv:2505.04258. [Google Scholar]
- Zhou, Z.; Wu, Z.; Boutteau, R.; Yang, F.; Demonceaux, C.; Ginhac, D. RGB-Event Fusion for Moving Object Detection in Autonomous Driving. In Proceedings of the IEEE International Conference on Robotics and Automation, London, UK, 29 May–2 June 2023. [Google Scholar]
- Wang, X.; Wu, Z.; Rong, Y.; Zhu, L.; Jiang, B.; Tang, J.; Tian, Y. SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Event-frame Based Recognition. arXiv 2023, arXiv:2308.04369. [Google Scholar]
- Li, Z.; He, H. EIFNet: Leveraging Event-Image Fusion for Robust Semantic Segmentation. arXiv 2025, arXiv:2507.21971. [Google Scholar]
- Xie, B.; Deng, Y.; Shao, Z.; Li, Y. EISNet: A Multi-Modal Fusion Network for Semantic Segmentation With Events and Images. Trans. Multimed. 2024, 26, 8639–8650. [Google Scholar] [CrossRef] [Scilit]
- Dong, J.; Zhuang, H.; Yang, H.; Pan, L. RGB-Event Fusion for Robust Lane Detection. In Proceedings of the British Machine Vision Conference, Sheffield, UK, 24–27 November 2025. [Google Scholar]
- Fan, L.; Yang, J.; Zhang, J.; Lian, X.; Shen, H.; Hu, D. Efficient Spiking Neural Network for RGB–Event Fusion-Based Object Detection. Electronics 2025, 14, 1105. [Google Scholar] [CrossRef] [Scilit]
- Shi, Y.; Li, M.; Chen, N.; An, W. Sparse-Gated RGB-Event Fusion for Small Object Detection in the Wild. Remote Sens. 2025, 17, 3112. [Google Scholar] [CrossRef] [Scilit]
- Zhu, L.; Zheng, Y.; Zhang, Y.; Wang, X.; Wang, L.; Huang, H. Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Wu, Y.; Fan, Z.; Chu, X.; Ren, J.S.; Li, X.; Yue, Z.; Li, C.; Zhou, S.; Feng, R.; Dai, Y.; et al. MIPI 2024 Challenge on Demosaic for HybridEVS Camera: Methods and Results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA, 17–18 June 2024. [Google Scholar]
- Xu, S.; Sun, Z.; Zhu, J.; Zhu, Y.; Fu, X.; Zha, Z.-J. DemosaicFormer: Coarse-to-Fine Demosaicing Network for HybridEVS Camera. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA, 17–18 June 2024. [Google Scholar]
- Lu, Y.; Xu, Y.; Ma, W.; Guo, W.; Xiong, H. Event Camera Demosaicing via Swin Transformer and Pixel-Focus Loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA, 17–18 June 2024. [Google Scholar]
- Zhou, S.; Zeng, H.; Lu, Y.; Chen, Y.; Liu, J.; Su, J. Lightweight Quad Bayer HybridEVS Demosaicing via State Space Augmented Cross-Attention. arXiv 2025, arXiv:2508.06058. [Google Scholar]
- Zhou, S.; Zeng, H.; Lu, Y.; Shao, T.; Tang, K.; Chen, Y.; Liu, J.; Su, J. Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS Demosaicing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025. [Google Scholar]
- Lu, Y.; Qian, Y.; Rao, Z.; Xiao, J.; Chen, L.; Xiong, H. RGB-Event ISP: The Dataset and Benchmark. In Proceedings of the International Conference on Learning Representations, Singapore, 24–28 April 2025. [Google Scholar]
- Zheng, W.; Fu, H.; Wang, X.; Kang, H.; Wang, C.; Liu, J.; Xu, Z.; Zhang, H.; Ma, H. EvRAW: Event-Guided Structural and Color Modeling for RAW-to-sRGB Image Reconstruction. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025. [Google Scholar]
- Sun, Q.; Yang, Q.; Li, C.; Zhou, S.; Feng, R.; Dai, Y.; Sun, W.; Zhu, Q.; Loy, C.C.; Gu, J.; et al. MIPI 2023 Challenge on RGBW Remosaic: Methods and Results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, BC, Canada, 18–22 June 2023. [Google Scholar]
- Sun, Q.; Yang, Q.; Li, C.; Zhou, S.; Feng, R.; Dai, Y.; Sun, W.; Zhu, Q.; Loy, C.C.; Gu, J.; et al. MIPI 2023 Challenge on RGBW Fusion: Methods and Results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, BC, Canada, 18–22 June 2023. [Google Scholar]
- Zeng, H.; Feng, K.; Cao, J.; Huang, S.; Zhao, Y.; Luong, H.; Aelterman, J.; Philips, W. Inheriting Bayer’s Legacy: Joint Remosaicing and Denoising for Quad Bayer Image Sensor. arXiv 2023, arXiv:2303.13571. [Google Scholar]
- Zheng, B.; Yuan, X.; Slabaugh, G.; Leonardis, A. Quad Bayer Joint Demosaicing and Denoising Based on Dual-Branch Deep Neural Network. Proc. AAAI Conf. Artif. Intell. 2024, 38, 7420–7428. [Google Scholar]
- Qian, G.; Wang, Y.; Gu, J.; Dong, C.; Heidrich, W.; Ghanem, B.; Ren, J. Rethinking Learning-Based Demosaicing, Denoising, and Super-Resolution Pipeline. In Proceedings of the IEEE International Conference on Computational Photography, Pasadena, CA, USA, 1–3 August 2022. [Google Scholar]
- Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.-H. Restormer: Efficient Transformer for High-Resolution Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022. [Google Scholar]
- Chen, L.; Chu, X.; Zhang, X.; Sun, J. Simple Baselines for Image Restoration. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022. [Google Scholar]
- Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; Timofte, R. SwinIR: Image Restoration Using Swin Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, Montreal, QC, Canada, 11–17 October 2021. [Google Scholar]
- Chen, X.; Wang, X.; Zhou, J.; Qiao, Y.; Dong, C. Activating More Pixels in Image Super-Resolution Transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023. [Google Scholar]
- Tu, Z.; Talebi, H.; Zhang, H.; Yang, F.; Milanfar, P.; Bovik, A.; Li, Y. MAXIM: Multi-Axis MLP for Image Processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022. [Google Scholar]
- Wang, Z.; Cun, X.; Bao, J.; Zhou, W.; Liu, J.; Li, H. Uformer: A General U-Shaped Transformer for Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022. [Google Scholar]
- Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.-H.; Shao, L. Multi-Stage Progressive Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 19–25 June 2021. [Google Scholar]
- Wu, Y.; Fan, Z.; Shinozaki, H.; Zhang, F.; Li, X.; Baudron, A.; Feng, W.; Zhao, S.; Han, J.; Li, C.; et al. MIPI 2025 Challenge on Deblurring for Hybrid EVS Camera: Methods and Results. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, Honolulu, HI, USA, 19–20 October 2025. [Google Scholar]
- Schwartz, E.; Giryes, R.; Bronstein, A.M. DeepISP: Toward Learning an End-to-End Image Processing Pipeline. IEEE Trans. Image Process. 2019, 28, 912–923. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liang, Z.; Cai, J.; Cao, Z.; Zhang, L. Cameranet: A Two-Stage Framework for Effective Camera ISP Learning. IEEE Trans. Image Process. 2021, 30, 2248–2262. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.-H.; Shao, L. CycleISP: Real Image Restoration via Improved Data Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020. [Google Scholar]
- Ignatov, A.; Timofte, R.; Van Vu, T.; Minh Luu, T.; Pham, T.X.; Van Nguyen, C.; Kim, Y.; Choi, J.-S.; Kim, M.; Huang, J.; et al. Learned Smartphone ISP on Mobile NPUs with Deep Learning, Mobile AI 2021 Challenge: Report. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Virtual, 19–25 June 2021. [Google Scholar]
- Hseih, B.-C.; Siddiqui, H.; Luo, J.; Georgiev, T.; Atanassov, K.; Goma, S.; Cheng, H.-Y.; Sze, J.J.; Lin, R.J.; Chou, K.Y.; et al. New Color Filter Patterns and Demosaic for Sub-Micron Pixel Arrays. In Proceedings of the International Image Sensor Workshop, Vaals, The Netherlands, 8–11 June 2015. [Google Scholar]
- Kim, Y.; Kim, Y. High-Sensitivity Pixels with a Quad-WRGB Color Filter and Spatial Deep-Trench Isolation. Sensors 2019, 19, 4653. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zeng, H.; Cai, J.; Li, L.; Cao, Z.; Zhang, L. Learning Image-Adaptive 3D Lookup Tables for High Performance Photo Enhancement in Real-Time. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 2058–2073. [Google Scholar] [PubMed]
- Perevozchikov, G.; Mehta, N.; Afifi, M.; Timofte, R. Rawformer: Unpaired Raw-to-Raw Translation for Learnable Camera ISPs. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Zeng, H.; Luong, H.; Philips, W. Wavelength-Embedding-Guided Filter-Array Transformer for Spectral Demosaicing. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Jiang, Y.; Liu, L.; Wu, D.; Huang, F.; Fu, Q.; An, T.; Niu, Y.; Zheng, C. Color Image Demosaicking: A Systematic Survey of Algorithms, Performance Evaluation, and Open Challenges. IEEE Access 2025, 13, 193049–193070. [Google Scholar] [CrossRef] [Scilit]
- Zheng, X.; Liu, Y.; Lu, Y.; Hua, T.; Pan, T.; Zhang, W.; Tao, D.; Wang, L. Deep Learning for Event-Based Vision: A Comprehensive Survey and Benchmarks. arXiv 2023, arXiv:2302.08890. [Google Scholar]
- Chakravarthi, B.; Verma, A.A.; Daniilidis, K.; Fermuller, C.; Yang, Y. Recent Event Camera Innovations: A Survey. arXiv 2024, arXiv:2408.13627. [Google Scholar]
- Wang, Z.; Chen, Q.; Liu, M.; Perrone, D.; Pei, Y.R.; Zou, Z.; Tan, S.; Han, T.; Lu, G.; Xu, Z.; et al. Event-Based Eye Tracking. AIS 2024 Challenge Survey. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA, 17–18 June 2024. [Google Scholar]
- Chen, Q.; Gao, C.; Liu, M.; Perrone, D.; Pei, Y.R.; Wang, Z.; Zou, Z.; Tan, S.; Han, T.; Lu, G.; et al. Event-Based Eye Tracking. 2025 Event-Based Vision Workshop. arXiv 2025, arXiv:2504.18249. [Google Scholar]
- Wang, Q.; Zhang, Y.; Yuan, C.; Li, J.; Cheng, X.; Tang, L.; Li, X.; Zhou, Y. DailyDVS-200: A Comprehensive Benchmark Dataset for Event-Based Action Recognition. arXiv 2024, arXiv:2407.05106. [Google Scholar]
- Xia, R.; Cai, J.; Leng, L.; Wang, L.; Liu, C.; Cheng, R.; Tang, Y.; Zhou, P. Temporal-Guided Visual Foundation Models for Event-Based Vision. arXiv 2025, arXiv:2511.06238. [Google Scholar]
- Creß, C.; Zimmer, W.; Purschke, N.; Doan, B.N.; Kirchner, S.; Lakshminarasimhan, V.; Strand, L.; Knoll, A.C. TUMTraf Event: Calibration and Fusion Resulting in a Dataset for Roadside Event-Based and RGB Cameras. arXiv 2024, arXiv:2401.08474. [Google Scholar]
- Ghosh, S.; Gallego, G. Event-Based Stereo Depth Estimation: A Survey. arXiv 2024, arXiv:2409.17680. [Google Scholar]
- Zimmer, W.; Wardana, G.; Sritharan, S.; Zhou, X.; Song, R.; Knoll, A. TUMTraf V2X Cooperative Perception Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024. [Google Scholar]
- Liao, W.; Zhang, X.; Yu, L.; Lin, S.; Yang, W.; Qiao, N. Synthetic Aperture Imaging with Events and Frames. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022. [Google Scholar]
- Tulyakov, S.; Fleuret, F.; Krawczuk, I.; Gehrig, D.; Gehrig, M.; Scaramuzza, D. Time Lens++: Event-Based Frame Interpolation with Parametric Non-Linear Flow and Multi-Scale Fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022. [Google Scholar]
- Lin, S.; Zhang, Y.; Yu, L.; Zhou, B.; Luo, X.; Pan, J. Autofocus for Event Cameras. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022. [Google Scholar]
- Zhang, J.; Yang, X.; Fu, Y.; Wei, X.; Yin, B.; Dong, B. Object Tracking by Jointly Exploiting Frame and Event Domain. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021. [Google Scholar]
- Shang, W.; Ren, D.; Zou, D.; Ren, J.S.; Luo, P.; Zuo, W. Bringing Events into Video Deblurring with Non-Consecutively Blurry Frames. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021. [Google Scholar]
- Jiang, Z.; Xia, P.; Huang, K.; Stechele, W.; Chen, G.; Bing, Z.; Knoll, A. Mixed Frame-/Event-Driven Fast Pedestrian Detection. In Proceedings of the IEEE International Conference on Robotics and Automation, Montreal, QC, Canada, 20–24 May 2019. [Google Scholar]
- Mitrokhin, A.; Ye, C.; Fermüller, C.; Aloimonos, Y.; Delbruck, T. EV-IMO: Motion Segmentation Dataset and Learning Pipeline for Event Cameras. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, Macau, China, 3–8 November 2019. [Google Scholar]
- Stoffregen, T.; Scheerlinck, C.; Scaramuzza, D.; Barnes, N.; Mahony, R.; Kleeman, L.; Yu, X.; Rebecq, H. Reducing the Sim-to-Real Gap for Event Cameras. In Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23–28 August 2020. [Google Scholar]
- Park, P.K.J. A Novel Quantised Image Sensing for Machine Vision. Electron. Lett. 2026, 62, e70515. [Google Scholar] [CrossRef] [Scilit]
- Park, P.K.J.; Kim, J.; Ko, J.; Chang, Y. High-Speed Image Restoration Based on a Dynamic Vision Sensor. Sensors 2026, 26, 781. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Park, P.K.J.; Kim, J.; Ko, J.; Chang, Y. Event-Based Machine Vision for Edge AI Computing. Sensors 2026, 26, 935. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Park, P.K.J.; Kim, J.; Ko, J.; Chang, Y. Low-Latency Machine Vision Based on a Neuromorphic Vision Sensor. Electronics 2026, 15, 2828. [Google Scholar] [CrossRef] [Scilit]
- Jiang, D.; Wang, H.; Li, T.; Gouda, M.A.; Zhou, B. Real-Time Tracker of Chicken for Poultry Based on Attention Mechanism-Enhanced YOLO-Chicken Algorithm. Comput. Electron. Agric. 2025, 237, 110640. [Google Scholar] [CrossRef] [Scilit]
- Zhang, B.; Li, Z.; Ma, Q.; Zhang, J.; Xiang, Z.; Jiang, D. Silhouette-Based Cross-View Motion Gait Recognition via a Multi-Scale Temporal Difference Unit. Electronics 2026, 15, 2512. [Google Scholar] [CrossRef] [Scilit]



















| Sensor Type | Advantages | Limitations | Representative Applications |
|---|---|---|---|
| Frame-based RGB image sensor | High spatial resolution; dense texture and color information; human-interpretable images; mature ISP and computer-vision pipelines; compatibility with large RGB datasets. | Motion blur under long exposure; limited temporal resolution at low frame rate; redundant data output in static scenes; high bandwidth and power at high resolution or high frame rate; limited dynamic range for single-exposure imaging. | Mobile imaging, surveillance, human-viewable video, object recognition, scene understanding, automotive cameras, industrial inspection. |
| Event-based sensor | Very low latency; high temporal resolution; high dynamic range; sparse and activity-driven output; reduced redundant data; suitable for high-speed motion and low-power sensing. | No direct dense RGB or absolute intensity output; weak response to static scenes; event noise and background activity; dependence on bias and contrast thresholds; irregular asynchronous data require specialized algorithms. | High-speed motion analysis, optical flow, gesture recognition, robotics, visual odometry, SLAM, automotive perception, neuromorphic edge sensing. |
| Hybrid image sensor | Combines dense spatial/color or intensity information with fast temporal event information; improves motion robustness, latency, and energy efficiency; supports both human-perceptual imaging and machine vision. | Increased sensor and system complexity; spatial/temporal calibration issues in multi-sensor systems; increased per-pixel circuit complexity in pixel-level shared sensors; color-fidelity and demosaicing challenges in stacked sensors; more complex fusion algorithms and benchmarks. | Motion blur-free imaging, video deblurring, frame interpolation, SLAM, object/person detection, gesture recognition, AR/VR, robotics, automotive perception, and always-on AIoT sensing. |
| Architecture Type | Integration Level | Advantages | Limitations | Suitable Applications |
|---|---|---|---|---|
| Dual camera | Module-level hybrid vision system | High flexibility; easy prototyping; independent sensor selection; suitable for dataset collection and algorithm evaluation. | Parallax; synchronization error; extrinsic calibration complexity; larger form factor; higher module-level power. | Robotics, autonomous driving datasets, object detection, SLAM, visual odometry, multimodal perception research. |
| Optically aligned | Module-level hybrid vision system | Reduced parallax; improved event–frame registration; suitable for dense fusion, deblurring, interpolation, and controlled experiments. | Light loss due to beam splitting; optical complexity; larger optical module; residual calibration still required. | Event-guided deblurring, frame interpolation, HDR/low-light imaging, image restoration, dataset acquisition. |
| Pixel-level shared | Physically integrated hybrid image sensor | Intrinsic spatial alignment between frame and event outputs; no external event–frame geometric calibration; compact single-sensor implementation; synchronized APS intensity frames and DVS events. | Increased per-pixel circuit complexity; reduced fill factor or larger pixel pitch; limited color capability in many DAVIS implementations; dual-mode readout complexity; less flexibility than dual-camera systems. | Visual odometry, SLAM, high-speed tracking, robotics, event-guided reconstruction, visual–inertial perception, and event–frame algorithm benchmarking. |
| Stacked | Physically integrated hybrid image sensor | High functional density; compact single-chip form factor; parallel CIS and DVS data paths; high event throughput; on-chip ISP/ESP possible. | High fabrication complexity; Cu–Cu bonding yield; inter-layer noise; thermal coupling; calibration and testing difficulty. | High-end mobile imaging, automotive perception, robotics, AR/VR, high-resolution hybrid sensing. |
| Homogeneous-pixel | Computationally hybrid vision system | Preserves regular CIS pixel layout; avoids event-pixel-induced static artifacts; compatible with small CIS pixel pitch and mature ISP pipelines. | Pseudo-events are discrete-time; latency limited by high-speed readout; not equivalent to true asynchronous DVS; higher internal readout bandwidth may be required. | Motion blur-free imaging, high-speed video reconstruction, mobile human-perceptual imaging, frame-domain event-like sensing. |
| Event-only algorithmic reconstruction | Computationally hybrid vision system | Event-only hardware; high temporal resolution; high dynamic range; frame-like grayscale output can be generated without a physical frame sensor. | Ill-posed intensity recovery; no direct color information; static scene ambiguity; dependence on learned priors and computation; possible hallucination or smoothing artifacts. | Event-camera visualization, high-speed grayscale video generation, robotics, event-based perception, downstream image-based algorithms. |
| RGB-to-DVS Ratio | Event-Pixel Occupancy | Evidence Level | Design Interpretation |
|---|---|---|---|
| 31:1 | 3.125% | Preliminary hypothesis | Low event-density, image-quality-oriented design point; not experimentally validated |
| 15:1 | 6.25% | Preliminary hypothesis | Intermediate trade-off hypothesis; not established as a practical baseline |
| 7:1 | 12.5% | Published MIPI 2024 evidence | Feasible under the specific challenge pattern, dataset, algorithms, and PSNR/SSIM metrics |
| Specifications | A [55,56] | B [23] | C [36] |
|---|---|---|---|
| Architecture | Stacked | Stacked | Homogeneous-pixel |
| CIS resolution | 4096 × 3680 (15 Mp) | 35.6 Mp | 4032 × 3024 (12 Mp) |
| Pixel ratio | 15:1 | 3:1 | N/A |
| CIS pixel pitch (μm) | 2.2 | 1.22 | 1.8 |
| DVS pixel pitch (μm) | 8.8 | 4.88 | 1.8 |
| Power Consumption (mW) | 64 (DVS-only) | 525 | 845 |
| Remarks | High-throughput DVS readout up to 4.6 GEvents/s with in-pixel TDC and on-chip ISP/ESP functions. | High-resolution RGB CIS combined with event pixels; adaptive event-sparsity control supports up to 10,000 event frames/s. | Pseudo-DVS is generated from high-speed CIS frame differencing; no static bad pixels are introduced by dedicated DVS pixels. |
| Architecture | Applications | Latency or Rate | Performance | Dynamic Range/Power | Calibration/ Computational Cost |
|---|---|---|---|---|---|
| Dual-camera hybrid vision system [196] | Roadside RGB–event object detection using the TUMTraf Event dataset | NR | Event–frame fusion improved detection performance by up to 9% mAP during daytime and 13% mAP at night relative to RGB-only detection | NR | Targetless extrinsic calibration for multiple moving objects; numerical residual calibration error NR |
| Dual-camera CIS–DVS system [208] | Post-capture motion-deblurred image restoration | NR | PSNR: 18.52 → 38.72 dB; SSIM: 0.683 → 0.911; normalized MTF50 ratio: 0.39 → 0.99 | NR | Spatial, temporal, resolution, and disparity compensation; computational cost NR |
| Optically aligned event–frame system [200] | Beam-splitter-based video frame interpolation using Time Lens++ | NR | Reconstruction improved by up to 0.2 dB PSNR and 15% LPIPS; dataset contained more than 100 scenes | NR | Beam-splitter-based spatial alignment; numerical residual calibration error NR |
| Pixel-level shared hybrid sensor [11] | Event–frame–IMU visual SLAM using a DAVIS-type sensor | NR | Accuracy improved by 130% over an event-only pipeline and 85% over a frame-only visual–inertial pipeline | NR | Intrinsically co-located frame and event measurements; IMU synchronization required |
| Stacked CIS–DVS sensor [23] | High-resolution RGB and event acquisition | Up to 10,000 event frames/s for the 35.6-Mpixel sensor; up to 4.6 GEvents/s for the three-wafer-stacked sensor | 35.6-Mpixel RGB with 4.88 μm event pixels; three-wafer-stacked implementation with 15-Mpixel CIS and 1-Mpixel EVS | 67.8 dB/ 525 mW | NR |
| Stacked/inserted-pixel HybridEVS [160,161,162,163,164] | Quad-Bayer HybridEVS demosaicing | NR | Best MIPI 2024 result: 44.8464 dB PSNR and 0.9854 SSIM for the approximately 7:1 RGB-to-DVS pattern | NR | NR |
| Pixel-aligned RGB–event ISP system [165] | Image enhancement | Processing time: 10 ms | Overall average PSNR of 32.47 dB | NR | InvertISP excels in computational efficiency with 1.41 GFLOPs |
| Homogeneous-pixel pseudo-DVS system [36] | Motion-blur-free hybrid image sensing | 1440 fps pseudo-DVS output; approximately 1 ms event-like latency | Normalized MTF50 ratio improved from approximately 0.4 to 0.8 after motion compensation | Approximately 845 mW for the demonstrated high-speed implementation | NR |
| Event-only computational reconstruction [18,73] | Events-to-video reconstruction using E2VID | Reconstructed-frame interval configurable; complete runtime depends on implementation | Reported to outperform preceding reconstruction methods by more than 20% in image quality | NR | NR |
| Event-only machine vision [209] | Human detection, pose estimation, and hand-posture recognition for edge AI | Human detection: 15 ms; pose estimation: 6 ms; hand posture: 14.31 ms | Human-detection computation: 5.8 G → 81 M FLOPs; pose mAP: 0.95 → 0.94; hand-posture recall: 99.19%, FAR: 0.0926% | NR | Pose-model size: 127 → 19 MB; detection speed-up greater than 11× |
| Event-only low-latency machine vision [210] | Person detection, gesture recognition, and SLAM on mobile processors | Person detection: 92 ms; gesture recognition: 20 ms; SLAM: 15.9 ms | Task-specific accuracy and robustness metrics reported separately in the source | NR | Exynos 7570, Exynos 5422, and Snapdragon 845; task-specific event representations |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Park, P.K.J.; Kim, J.; Ko, J. Hybrid Event–Frame Sensing for Human-Perceptual Imaging and Machine Vision. Sensors 2026, 26, 5127. https://doi.org/10.3390/s26165127
Park PKJ, Kim J, Ko J. Hybrid Event–Frame Sensing for Human-Perceptual Imaging and Machine Vision. Sensors. 2026; 26(16):5127. https://doi.org/10.3390/s26165127
Chicago/Turabian StylePark, Paul K. J., Junseok Kim, and Juhyun Ko. 2026. "Hybrid Event–Frame Sensing for Human-Perceptual Imaging and Machine Vision" Sensors 26, no. 16: 5127. https://doi.org/10.3390/s26165127
APA StylePark, P. K. J., Kim, J., & Ko, J. (2026). Hybrid Event–Frame Sensing for Human-Perceptual Imaging and Machine Vision. Sensors, 26(16), 5127. https://doi.org/10.3390/s26165127

