Next Article in Journal
Model-Contingent Polarity Bias in Large Language Model Annotation: Implications for Semantic Multimedia Personalization
Next Article in Special Issue
LACE-Net: A Swin Transformer with Local Frequency-Domain Energy and Adaptive Contrast Enhancement for Fine-Grained Land Cover Classification
Previous Article in Journal
Verification of the Methods of Digital Monitoring of Information Space Based on Coding Theory Tools
Previous Article in Special Issue
SESQ: Spatially Aware Encoding and Semantically Guided Querying for 3D Grounding
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

DaN: A Comprehensive Semi-Real Dataset for Extreme Low-Light Image Enhancement

1
China Nuclear Industry 23 Construction Co., Ltd., Beijing 100080, China
2
Cyberspace Security Institute, Beijing University of Posts and Telecommunications, Beijing 100876, China
*
Author to whom correspondence should be addressed.
Computers 2026, 15(5), 261; https://doi.org/10.3390/computers15050261
Submission received: 15 March 2026 / Revised: 10 April 2026 / Accepted: 13 April 2026 / Published: 22 April 2026
(This article belongs to the Special Issue Advanced Image Processing and Computer Vision (2nd Edition))

Abstract

Extreme low-light image enhancement (ELLIE) targets the restoration of visual quality under ultra-dim environments (<0.1 lux). Conventional image signal processing (ISP) pipelines often fail in such scenarios due to the limitations of heuristic, hand-crafted algorithms. While deep learning has advanced the field via end-to-end mapping, existing models suffer from constrained generalization and suboptimal perceptual fidelity, primarily stemming from the scarcity of large-scale, high-diversity datasets. To bridge this gap, we present the Day and Night (DaN) dataset, a semi-synthetic benchmark synthesized through a rigorous physics-based noise model. This approach effectively captures authentic noise characteristics while enabling the scalable generation of paired samples across multifaceted illumination conditions and scenes. Furthermore, we propose No Longer Vigil (NLV), a fully differentiable AI-ISP framework. By replacing traditional rigid blocks with adaptive non-linear networks, NLV facilitates scene-dependent transformations without requiring manual priors. Comprehensive evaluations demonstrate that our method significantly outshines state-of-the-art approaches, yielding a 4.15 dB gain in PSNR and a 0.026 improvement in SSIM.

1. Introduction

Extreme low-light image enhancement (ELLIE) has become essential for full-color imaging in environments with illumination below 0.1 lux, such as night surveillance, autonomous systems, and medical diagnostics. The main challenge lies in the suboptimal performance of ineffective hand-crafted algorithms. For example, although traditional image signal processing (ISP) methods use infrared cameras in dark scenes, the lack of color data prevents precise color representation [1,2,3,4,5], and while increasing ISO sensitivity or extending exposure time can improve photon capture, these hand-crafted algorithms ultimately result in increased noise and motion blur, making them ineffective for dynamic scenes [6,7,8,9,10,11,12,13].
Recent deep learning-based approaches have circumvented the difficulties of hand-crafted algorithm design through end-to-end learning. These approaches effectively suppress noise by learning to map noisy images to clear ones [14,15,16]. Moreover, there has been a notable shift in network training data from RGB three-channel images to single-channel raw Bayer array images, which contain more comprehensive information [17,18,19,20,21,22,23,24,25].
The existing methods mentioned above fundamentally depend on the quality of the training datasets [26,27,28,29,30,31,32]. While foundational benchmarks such as SID [18] and ELD [19] propelled early advancements in raw-domain restoration, their empirical utility is constrained by severe environmental homogeneity. Specifically, the static nature and restricted scene variability within these collections lead to poor algorithmic generalization when facing dynamic, semi-real low-light conditions. Furthermore, existing corpora predominantly neglect fine-grained structural restoration (e.g., text and complex textures) under sub-0.1 lux illuminance, leaving a significant gap in robust model evaluation. Specifically, the current LOL series datasets (illumination > 10 lux) are inadequate for extremely low-light conditions (<0.1 lux). Furthermore, these datasets are missing elements such as text or objects, leading to blurry images, as illustrated in Figure 1. Moreover, the existing metric scores do not accurately represent visual quality: The test datasets employ shooting strategies similar to the training datasets, including ISO, exposure time, and camera. This results in high metrics that do not align with visual image quality.
The limitations of the existing datasets highlight the need for a comprehensive dataset dedicated to extreme low-light image enhancement. In response, we propose Day and Night (DaN), a large-scale semi-synthetic dataset comprising 4200 raw format images spanning diverse scenarios, City, Suburban, Indoor, and Wild, captured at multiple times of day. We partitioned DaN into two components, among which 100 pairs were utilized for testing, and 4000 images combined with the noise model were employed for training. DaN integrates a physics-based noise model to simulate the multimodal Gaussian noise distribution in the real world, significantly reducing the time cost of data collection and expanding the data volume to infinity. Meanwhile, the independent distribution of Gaussian noise guarantees independence among the data. In addition, DaN contains data instances that exhibit character-like elements or analogous structures, which are formally defined as fine-grained features following the methodology established in [33,34]. A comparative analysis among the proposed DaN and previous datasets [18,19,24,35] is presented in Table 1.
Furthermore, existing methods [18,19,25,36,37,38] stack multiple UNet-based deep neural networks (DNN) for denoising while disregarding other ISP components, such as white balance and color correction. This oversight results in unreal color distortion when deployed in the real world. Thus, we introduce No Longer Vigil (NLV), a pure AI-ISP pipeline. The proposed NLV comprises three components: Denoising Networks (DNs), White Balance (WB), and Color Correction (CC). These components are realized by neural networks instead of manually fixed parameters. For example, white balance is a fixed parameter matrix M W B R H × 1 × 1 in traditional ISP, whereas in our NLV, it is predicted by a fully connected layer. The integration of ISP theory and deep learning eliminates the priors and compensates for the inexplicability of neural networks.
We conducted comprehensive experiments on the proposed DaN dataset and the NLV pipeline. We first trained and tested our NLV method on the existing SID dataset. The results demonstrate state-of-the-art performance, with our NLV outperforming existing LED methods by +4.15 dB in PSNR and +0.026 in SSIM. Subsequently, six existing methods [18,19,25,36,37,38] and our NLV method were trained on our DaN dataset and tested on the existing SID dataset. The results show that the same method performs better when trained on our DaN dataset. Finally, we conducted tests on our DaN dataset. The results demonstrated that methods with comparable performance on the existing SID dataset exhibit a considerable gap on our DaN dataset.
Our contributions are summarized below:
  • We have developed a semi-synthetic dataset termed Day and Night (DaN). To our knowledge, DaN represents the first large-scale dataset specifically focused on extreme low-light image enhancement (ELLIE). Our DaN surpasses existing datasets in comprehensiveness. Its diverse and fine-grained features render it more challenging.
  • We introduce No Longer Vigil (NLV), a pure AI image signal processing (AI-ISP) pipeline that reconstructs the traditional ISP pipeline by deep neural networks (DNNs) instead of merely replacing some components. NLV automatically processes signals, eliminating the priors.
  • Extensive experiments show that our NLV pipeline outperforms existing methods in terms of the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM). Our DaN can serve as a benchmark for extremely low-light image processing.

2. Related Works

2.1. ELLIE Datasets

Existing datasets like SID [18] and LOL [35] initiated raw sensor data collection, but were restricted by their small size and lack of variety. In contrast, the LOLV2 and ELD datasets [24,35] use generative models to increase data volume. However, they involve implicit relations, making it seem as if the model has already “seen” the test sets. Our DaN dataset stands out because (1) our DaN is 10 times larger than the previous one, encompassing broader types, periods, and fine-grained features, and (2) it ensures independence without implicit mappings, making them more challenging and suitable for benchmarking.

2.2. ELLIE Approaches

Existing approaches can be classified into two categories: traditional ISP pipeline and learning-based methods.
  • Traditional ISP pipeline: The traditional ISP pipeline [6,7,8,9,10,11,12,13] encompasses multiple components, such as denoising, demosaic, color correction, and white balance. These components are artificially designed and equipped with fixed parameters. The well-designed ISP pipeline is fabricated into chips and is prevalently present in major camera manufacturers. This kind of pipeline is hardware-dependent and demonstrates poor performance in extremely low light due to the constraints of CMOS.
  • Deep learning-based methods: Deep learning-based methods usually stack one to three DNNs for learning implicit noisy-normal image mapping. Various models have since been put forward, notably [18,25,36,37,39,40,41]. These models merely focus on denoising yet overlook other ISP components, thereby causing blurriness, color distortion, and the loss of fine-grained features. Our NLV method, in conjunction with ISP theory, has redesigned a pure AI-ISP pipeline suitable for ELLIE. The pipeline parameters are all adaptive, bypassing the inconvenience parameter tuning.

2.3. Evaluation Metrics

The current evaluation metrics include peak signal-to-noise ratio (PSNR) [42], measuring the logarithmic ratio of peak signal power to noise level, and the structural similarity index (SSIM) [43], evaluating perceptual fidelity through structural analysis. Additionally, researchers perform subjective evaluations of visual details.

3. Materials: The Proposed DaN Dataset

This section introduces our proposed DaN (Day and Night) dataset. First, we outline the methodology for data collection and annotation. Then, we present the proposed noise model designed to generate noisy images. Finally, we provide a quantitative assessment of our DaN dataset.

3.1. Dataset Collection

We collected 4731 raw Bayer images for the training set, which were categorized into indoor and outdoor subsets. The indoor images encompassed as many chromatic objects as possible. The outdoor images were collected from cities, suburbs, and wild landscapes, enhancing generalization. The brand of cameras is Sony A7S2, and all images were stored in the “.ARW” raw format. Furthermore, we deliberately captured images containing plentiful fine-grained features. The data collection was accomplished by five people.
We analyzed the quality and quantity of the images and classified them into four categories: Indoor (Object) and Outdoor (City, Suburb, Wild). Then, we conducted a thorough review to eliminate highly repetitive images. We excluded those containing sensitive content, such as pornography and political material, and any content infringing on privacy rights. After meticulous screening, we ultimately chose 4000 raw Bayer images as the training set of our DaN.
We also gathered 100 pairs of raw Bayer images for the test set. Each pair comprises a noisy image captured in a highly dim light environment (<0.1 lux) and a clean image taken during the daytime (>100 lux) that has been manually aligned. These images were taken by five different mobile phones and are unrelated to the images in the training set.

3.2. Annotation Pipeline

We employ a noise model to generate noise and mark it on the Ground Truth as a label. Theoretically, the noise model expands the training set to an infinite size. The whole process includes three steps, noise modeling, noise parameter selection, and manual annotation, as shown in Figure 2.

3.2.1. Noise Modeling

The noise of camera imaging in low illumination environments is mainly shot noise, dark noise, and various additive noise [44]. Shot noise follows the Poisson distribution, is related to the input signal, and can be expressed as [45]
N shot Poisson [ T · α · Φ ] ,
where T represents the exposure time, measured in second (S). The α denotes the photon conversion rate, and Φ indicates the luminous flux. Dark noise adheres to Poisson distribution [46]:
N dark Poisson [ T · Ψ ] ,
where Ψ is the temperature, measured in Kelvin (K). Additive noise encompasses read noise and noise generated during the digital-to-analog conversion process, and it conforms to a Gaussian distribution. Additive noise is
N add N [ 0 , σ 2 ] ,
σ 2 represents the variance of the additive noise. Therefore, The signal of the camera is represented as follows:
N = g ( I + N S h o t ) + g · N d a r k + N a d d ,
where I is the original image without noise, and g is an amplification factor of the camera, pertains to the global gain, and is proportional to ISO. N follows Gaussian distribution.

3.2.2. Noise Parameters Selection

According to formulation (4), g, the variances of N a d d , N s h o t , and N d a r k need to be calibrated, and the mean and variance of the signal are
E ( N ) = T · ( a · Φ + Ψ ) · g ,
σ ( N ) 2 = T · ( a · Φ + Ψ ) · g 2 + N add 2 ,
If we take a dark frame image, the signal is 0, and thus, the mean of the image can be expressed as follows:
E ( I ) = T · ( a · 0 + Ψ ) · g = T · Ψ · g = g · N d a r k ,
this expression represents the g · N d a r k , which will be used to replace the dark noise in the subsequent processing. According to Formulas (5) and (6),
σ ( I ) 2 = E ( I ) · g + N add 2 ,
The mean and variance of image pixels exhibit a linear relationship, with the slope being g and the intercept representing the additive noise variance σ a d d 2 . So we captured many gray scale charts, calculated the mean and variance of each pixel, and plotted the mean-variance coordinate graph to complete the calibration, as shown in Figure 2. N S h o t can be inversely deduced once g, σ a d d 2 , and N d a r k have been calibrated. The noise parameters are distinct under different configurations. Therefore, we have calibrated a total of 24 sets of noise parameters altogether. Parameters are randomly chosen when generating noise according to (4).

3.2.3. Manual Annotation

The noise generated by the physics-based noise model is distributed independently, and theoretically, no implicit mapping exists. In order to match the real world, we conduct an iterative selection of the noise that is most consistent with the real environment and align it spatially with the real data through cross-correlation matching [47], thereby minimizing the spatial offset and overexposure.

3.3. Dataset Statistics

Our DaN dataset comprises 4200 Ground Truths with a resolution ranging from 1440 × 2560 to 4320 × 7680 pixels, extracted from cameras in raw format. Among them, 4000 images are used for training, and in combination with the noise model, there are theoretically an infinite number of training data pairs. 100 pairs of real images are used for evaluation. The comprehensive details of the dataset are shown in Figure 3.

Ethical Statement

We hold the copyright for the photographs that we collect and adhere strictly to the requirements established in the data collection and annotation process. To further enhance our ethical compliance, we will take the following measures: (1) Implement rigorous data anonymization techniques to safeguard personal information. (2) Ensure transparency regarding data sources and collection methods in our evaluations. (3) Commit to ongoing reviews and remain prepared to delete or modify any data collected from sources that may be considered ethically inappropriate or that have not provided the requisite authorization.

4. Methods

In extremely low-light conditions, the traditional image signal processing (ISP) pipeline often produces poor images. Current learning-based models [18,25,36,37,39,40,41] that utilize multiple layers of deep neural networks (DNN) for image denoising perform only one stage within the ISP process. The inadequacy of the ISP components leads to blurred images, color distortion, and lack of fine-grained features. To tackle these problems, we introduce No Longer Vigil (NLV), an entirely AI-based ISP framework. NLV integrates and reconfigures the homogeneous linear elements found in conventional ISPs. It consists of three main components: a Denoising Network (DN), White Balance (WB), and Color Correction (CC). The primary innovations of our NLV pipeline over existing methods are: (1) Our NLV has comprehensively reorganized and redesigned the structure and sequence of traditional ISP components. (2) All parameters within NLV are self-adjusting, contrary to the static, hand-crafted parameters of traditional ISPs. The structure of the proposed NLV pipeline is illustrated in Figure 4.

4.1. Denoising Network (DN)

In the conventional ISP pipeline, the denoising process is executed on RGB images containing three channels and is typically performed towards the latter stages of the pipeline. Conversely, in the proposed NLV framework, we prioritize denoising and combine it with black level correction (BLC), lens shading correction (LSC), and wavelet transformation. The initial raw image R i n R H × W × 1 is processed to obtain the denoised raw Bayer images R o u t R H × W × 1 .
R o u t = C o n v u p ( C o n v d o w n ( R i n , φ d n 1 ) , φ d n 2 ) ,
where C o n v u p and C o n v d o w n denote the upsampling convolutional network and downsampling convolutional network, respectively. φ d n represents the parameters. Both C o n v u p and C o n v d o w n have 4 convolutional layers, each with a kernel of ( 3 × 3 ) , a stride of 1, and a padding of 0.
Unlike standard spatial filters, the wavelet transform facilitates a multi-scale time-frequency analysis. By projecting the features into the frequency domain, it separates complex noise distributions from the underlying image content. The mathematical formulation utilized in our network can be expressed as
X ( x , y ) = x n = ( y n x ) ψ n x ,
where x defines the scaling factor that dictates the resolution, y acts as the translation parameter for spatial shifting, and ψ represents the selected mother wavelet function.

4.2. White Balance (WB)

At varying temperatures, the pixel values for the same color differ [48]. To correct for this, white balance is utilized. Standard signal processing systems typically compensate for chromatic aberrations by applying rigid, pre-calibrated temperature configurations, such as a standard 5500 K daylight profile, which lack adaptability in complex dark scenes. In contrast, our NLV pipeline uses a fully connected layer to estimate the color temperature parameter. The expression for WB is
φ W B = F C W B ( R o u t , φ F C ) ,
R W B = R o u t · φ W B ,
where φ F C represents the parameter of the fully connected layer, and R W B is the output. The WB contains a convolutional layer with a 3 × 3 convolution kernel, a stride of 1, and a padding of 0.

4.3. Color Correction (CC)

Standard signal processors typically decouple chromatic adaptations into distinct gamut mapping and gamma adjustment phases, which can compound errors under extreme noise. To circumvent this, our NLV architecture intrinsically fuses these non-linear transformations into a cohesive, data-driven mapping step. The generalized color restoration is parameterized by a dedicated convolutional mechanism, defined as
R C C = C o n v C C ( R W B , φ C C ) ,
where R C C denotes the output of our NLV pipeline, φ C C represents the learnable parameters of the color correction, C o n v C C represents the color correction network. The R C C is subsequently downsampled to a three-channel RGB image I R H 2 × W 2 × 3 by the demosaic algorithm, and can be saved in image formats such as “PNG” or “JPEG”. CC consists of two convolutional layers with 3×3 convolution kernels, stride 1, and padding 0.

5. Results

5.1. Training Details

All baseline methods were trained on the existing SID dataset and our DaN dataset, respectively. For fairness, we unified the training protocols. Model optimization was driven by the L1 objective function, iteratively minimized via the Adam algorithm. We applied a continuous learning rate scheduling strategy, initiating at a base of 10 4 and decaying to a lower bound of 10 6 across 100 training cycles. To facilitate robust feature extraction, spatial dimensions of the input batches were standardized to 512 × 512 pixel patches through a random cropping strategy. Hardware acceleration for all network training and evaluation phases was provided by a single 48 GB NVIDIA A40 Tensor Core GPU.

5.2. Comparisons with State-of-the-Art Methods

We utilized the existing SID dataset as the foundational dataset. For comparison purposes, we employed six methods [18,19,25,36,37,38] sourced from conferences and journals of 2023–2024, including ICCV, CVPR, and TPAMI. Testing on the SID dataset reveals that our NLV pipeline outperforms the current SOTA methods. Moreover, evaluations on our DaN dataset demonstrate that DaN boosts the model’s performance and provides a more comprehensive representation of the model’s capabilities, establishing it as a viable benchmark. The findings are displayed in Table 2.

5.3. Case Study

To thoroughly assess the functional aspects of our proposed DaN dataset and NLV pipeline, we conducted an in-depth visual study using six representative samples. We compared these samples against three leading extreme low-light enhancement techniques: ELD [19], RetinexFormer [25], and LED [38]. Although these baseline methods show similar results on standard quantitative metrics (PSNR: 36.00 ± 1.00 dB, SSIM: 0.870 ± 0.030), our qualitative analysis highlighted significant perceptual differences in key regions. Detailed visual analysis is shown in Figure 5.
Visual inspection reveals that our NLV pipeline, when trained on the proposed DaN dataset, outperforms current methods that are trained on existing datasets. It particularly excels in color restoration, as illustrated in the sample comparison in the fourth row, and in the restoration of fine-grained features, such as the characters depicted in the first row. Existing methods are plagued by issues like blurriness, excessive exposure, and an inability to restore characters effectively. Although our NLV pipeline mitigates these challenges, it does not solve them perfectly, indicating that the DaN dataset presents significant challenges. Thus, we conclude that our NLV pipeline exhibits superior performance in the restoration of fine-grained features and color, which can be attributed to the incorporation of the Denoising Network (DN), White Balance (WB), and Color Correction (CC). Furthermore, the DaN datasets proves to be a more challenging benchmark for ELLIE.

5.4. Ablation Study

  • Our NLV pipeline components: The boxes of different colors are zoom areas, used for comparing the details’ restoration. In the ablation experiment, the effects of the two components of the NLV were investigated: white balance and color correction. The test was conducted on SID and DaN dataset, and the experimental results are presented in Table 3. The results demonstrate that two components significantly influence the ISP pipeline.
  • Wavelet Transform: To validate the necessity of the wavelet module within our pipeline, we benchmarked it against alternative frequency-domain techniques, namely the Fourier, Hilbert, short-time Fourier, Wigner, and Radon transforms. The quantitative results indicate that the wavelet approach preserves structural fidelity most effectively, as shown in Table 4. We attribute this advantage to its multi-resolution characteristics, which inherently decouple high-frequency degradation from critical image structures, rendering it highly effective for restoring textures and edges in severely degraded low-light environments.

5.5. Generalization Evaluation

To further evaluate the model’s robustness, we conducted experiments in an industrial setting—specifically, robotic gas metal arc welding (GMAW) under extremely low-light conditions, where illumination is limited to the molten pool region. Post-weld quality assessment relies heavily on visual inspection of the weld seam; however, images captured in such low-light environments suffer from severe degradation in contrast, noise, and detail fidelity, thereby compromising defect detection accuracy. To assess semi-real deployability and generalization capability, we implemented the proposed method on an embedded NVIDIA Jetson Orin platform. Quantitative and qualitative evaluations demonstrate that our model significantly enhances image visual quality, restoring structural details, improving contrast, and suppressing noise. As shown in the Figure 6, the enhanced images enable more reliable downstream analysis: defect detection accuracy improves by 12.92% compared to results obtained using the original low-quality images.

6. Discussion

Are DaN datasets genuinely significant?The experimental findings reveal that the DaN dataset adequately addresses the limitations inherent in existing datasets, resulting in a marked improvement in training effectiveness, as illustrated in Table 2. We propose that the key advantage of DaN lies in its exceptional generalization capabilities. In contrast, current SID datasets are restricted by their limited scale, a factor often leading to overfitting when applied to train complex models with numerous parameters. Such an overfitting can substantially compromise training outcomes. However, the extensive data provided by DaN substantially reduce the risk of overfitting. Furthermore, the multicategory design of DaN allows it to attain outstanding performance, even with unfamiliar data. A cross-visualization analysis was performed to examine the link between the DaN dataset and the NLV pipeline, as shown in Figure 7. The analysis indicates that both the DaN dataset and the NLV pipeline enhance the performance of the model, with a greater dependence of visual outcomes on the DaN dataset.
How does the fully processing NLV provide advantages over a single processing? The most notable aspect of NLV is its redesigned AI-ISP framework, which significantly diverges from traditional methods that focus primarily on developing denoising networks. These traditional methods, with their exclusive focus on noise reduction, overlooked the incorporation of other crucial ISP components, thus greatly constraining the overall efficacy of the model. This constraint emerges from the intricate nature of the ISP pipeline, which includes various non-linear operations that are not straightforwardly translatable to deep neural networks. Furthermore, the shared parameters native to convolutional layers add complexity to implementing essential tasks, such as white balance and color correction. Consequently, the comprehensive NLV process, which amalgamates multiple processing stages, exhibits a distinct advantage over single-process networks that concentrate only on denoising deep neural networks.
How do the components of NLV interact and collaborate? In conventional ISP systems, the individual components usually function alone without mutual connections, such as denoising and color correction. However, the components within NLV show interdependence, working harmoniously to create an integrated system, as shown in Table 3. Removing any component from this system will cause the entire model to collapse. Implicit associations between our NLV components are not present in traditional ISP.

7. Conclusions

This paper introduces Day and Night (DaN), a comprehensive dataset designed for extreme low-light image enhancement (ELLIE). When combined with our noise model, DaN provides a limitless pool of training data. Furthermore, we have developed an entirely Artificial Intelligence-based image signal processing (AI-ISP) pipeline known as No Longer Vigil (NLV). This novel pipeline has redefined the AI-ISP framework, suggesting a new path for future research. The NLV pipeline improves image quality in extremely low-light environments and is expected to be beneficial for applications like night surveillance, cave exploration, and maritime navigation. Nonetheless, our investigation of the NLV components is potentially not exhaustive; for instance, the denoising network could be complemented with conventional filtering techniques. Our future plan includes implementing the proposed NLV model in real camera systems and testing these devices across various scenarios.

Author Contributions

Conceptualization, Q.S. and S.L.; methodology, Q.S.; software, H.L.; validation, L.S. and K.L. (Kun Lu); formal analysis, K.L. (Kangtai Liu); investigation, Q.S.; resources, H.L.; data curation, H.L.; writing—original draft preparation, S.L.; writing—review and editing, Q.S.; visualization, H.L.; supervision, S.L.; project administration, Q.S. and Y.F.; funding acquisition, Q.S. and Y.F. All authors have read and agreed to the published version of the manuscript.

Funding

The authors declare that financial support was received for the research, authorship, and/or publication of this article. This research was funded by China National Nuclear Corporation (Project Number: CNI23-W-KY-2024-01), and the funded content was “Image Enhancement in Extremely Low Light Welding Environments Based on Computer Vision”.

Data Availability Statement

Dataset is available on request from the authors.

Conflicts of Interest

Authors QiuYang Sun, Hong Li, YingChao Feng, LiuQing Sun, Kun Lu and KangTai Liu were employed by the company China Nuclear Industry 23 Construction Co., Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ELLIEExtreme Low-light Image Enhancement
ISPImage Signal Processing
DNNDeep Neural Network
WBWhite Balance
CCColor Correction
PSNRPeak Signal-to-Noise Ratio
SSIMStructural Similarity Index

References

  1. Jiang, L.J.; Ng, E.Y.K.; Yeo, A.C.B.; Wu, S.; Pan, F.; Yau, W.Y.; Chen, J.H.; Yang, Y. A perspective on medical infrared imaging. J. Med. Eng. Technol. 2005, 29, 257–267. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Kastberger, G.; Stachl, R. Infrared imaging technology and biological applications. Behav. Res. Methods Instrum. Comput. 2003, 35, 429–439. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Diakides, N.A.; Bronzino, J.D. Advances in medical infrared imaging. In Medical Infrared Imaging; CRC Press: Boca Raton, FL, USA, 2007; pp. 19–32. [Google Scholar]
  4. Altınoğlu, E.İ.; Adair, J.H. Near infrared imaging with nanoparticles. Wiley Interdiscip. Rev. Nanomed. Nanobiotechnol. 2010, 2, 461–477. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Türker-Kaya, S.; Huck, C.W. A review of mid-infrared and near-infrared imaging: Principles, concepts and applications in plant tissue analysis. Molecules 2017, 22, 168. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Toet, A. Color the night: Applying daytime colors to nighttime imagery. In Proceedings of the Enhanced and Synthetic Vision 2003; SPIE: Bellingham, WA, USA, 2003; Volume 5081, pp. 168–178. [Google Scholar]
  7. Kriesel, J.; Gat, N. True-color night vision cameras. In Proceedings of the Optics and Photonics in Global Homeland Security III; SPIE: Bellingham, WA, USA, 2007; Volume 6540, pp. 65–74. [Google Scholar]
  8. Liu, S.; Feng, C.; Wang, X.; Wang, H.; Zhu, R.; Li, Y.; Lei, L. Deep-flexisp: A three-stage framework for night photography rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022; pp. 1211–1220. [Google Scholar]
  9. Zini, S.; Rota, C.; Buzzelli, M.; Bianco, S.; Schettini, R. Back to the future: A night photography rendering ISP without deep learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2023; pp. 1465–1473. [Google Scholar]
  10. Yang, J.B., Sr.; Lu, Y.; Wang, L.; Zhao, K.; Yang, C.; Liu, Y.C.; Chai, X.H. Research on starlight level broad spectrum full color imaging technology. In Proceedings of the AOPC 2019: Optical Sensing and Imaging Technology, 2019; SPIE: Bellingham, WA, USA, 2019; Volume 11338, pp. 470–482. [Google Scholar]
  11. Toet, A. Applying daytime colors to multiband nightvision imagery. In Proceedings of the Sixth International Conference on Information Fusion (FUSION); SPIE: Bellingham, WA, USA, 2003. [Google Scholar]
  12. Toet, A.; de Jong, M.J.; Hogervorst, M.A.; Hooge, I.T.C. Perceptual evaluation of colorized nighttime imagery. In Proceedings of the Human Vision and Electronic Imaging XIX, 2014; SPIE: Bellingham, WA, USA, 2014; Volume 9014, pp. 276–289. [Google Scholar]
  13. Zuo, W.; Zhang, L.; Song, C.; Zhang, D. Texture enhanced image denoising via gradient histogram preservation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2013; pp. 1203–1210. [Google Scholar]
  14. Zhang, K.; Zuo, W.; Zhang, L. FFDNet: Toward a fast and flexible solution for CNN-based image denoising. IEEE Trans. Image Process. 2018, 27, 4608–4622. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Guo, S.; Yan, Z.; Zhang, K.; Zuo, W.; Zhang, L. Toward convolutional blind denoising of real photographs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2019; pp. 1712–1722. [Google Scholar]
  16. Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.H.; Shao, L. Learning enriched features for fast image restoration and enhancement. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 1934–1948. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Wei, K.; Fu, Y.; Zheng, Y.; Yang, J. Physics-based noise modeling for extreme low-light photography. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 8520–8537. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Chen, C.; Chen, Q.; Xu, J.; Koltun, V. Learning to see in the dark. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2018; pp. 3291–3300. [Google Scholar]
  19. Wei, K.; Fu, Y.; Yang, J.; Huang, H. A physics-based noise formation model for extreme low-light raw denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2020; pp. 2758–2767. [Google Scholar]
  20. Zhang, Y.; Qin, H.; Wang, X.; Li, H. Rethinking noise synthesis and modeling in raw denoising. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2021; pp. 4593–4601. [Google Scholar]
  21. Moran, N.; Schmidt, D.; Zhong, Y.; Coady, P. Noisier2noise: Learning to denoise from unpaired noisy data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2020; pp. 12064–12072. [Google Scholar]
  22. Calvarons, A.F. Improved Noise2Noise denoising with limited data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2021; pp. 796–805. [Google Scholar]
  23. Zhang, C.; Han, W.; Zhou, Y.; Shen, J.; Xu, C.Z.; Liu, W. Leveraging Frame Affinity for sRGB-to-RAW Video De-rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2024; pp. 25659–25668. [Google Scholar]
  24. Yang, W.; Wang, W.; Huang, H.; Wang, S.; Liu, J. Sparse Gradient Regularized Deep Retinex Network for Robust Low-Light Image Enhancement. IEEE Trans. Image Process. 2021, 30, 2072–2086. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Cai, Y.; Bian, H.; Lin, J.; Wang, H.; Timofte, R.; Zhang, Y. Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2023; pp. 12504–12513. [Google Scholar]
  26. Jin, X.; Han, L.H.; Li, Z.; Guo, C.L.; Chai, Z.; Li, C. Dnf: Decouple and feedback network for seeing in the dark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2023; pp. 18135–18144. [Google Scholar]
  27. Abdelhamed, A.; Brubaker, M.A.; Brown, M.S. Noise flow: Noise modeling with conditional normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2019; pp. 3165–3173. [Google Scholar]
  28. Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.H.; Shao, L. Cycleisp: Real image restoration via improved data synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2020; pp. 2696–2705. [Google Scholar]
  29. Jang, G.; Lee, W.; Son, S.; Lee, K.M. C2n: Practical generative noise modeling for real-world denoising. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2021; pp. 2350–2359. [Google Scholar]
  30. Wang, Y.; Huang, H.; Xu, Q.; Liu, J.; Liu, Y.; Wang, J. Practical deep raw image denoising on mobile devices. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2020; pp. 1–16. [Google Scholar]
  31. Maleky, A.; Kousha, S.; Brown, M.S.; Brubaker, M.A. Noise2noiseflow: Realistic camera noise modeling without clean images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022; pp. 17632–17641. [Google Scholar]
  32. Kousha, S.; Maleky, A.; Brown, M.S.; Brubaker, M.A. Modeling srgb camera noise with normalizing flows. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022; pp. 17463–17471. [Google Scholar]
  33. Wah, C.; Branson, S.; Welinder, P.; Perona, P.; Belongie, S. The Caltech-Ucsd Birds-200-2011 Dataset; Technical Report; California Institute of Technology: Pasadena, CA, USA, 2011. [Google Scholar]
  34. Khosla, A.; Jayadevaprakash, N.; Yao, B.; Li, F.F. Novel dataset for fine-grained image categorization: Stanford dogs. In Proceedings of the CVPR Workshop on Fine-Grained Visual Categorization (FGVC); IEEE: Colorado Springs, CO, USA, 2011. [Google Scholar]
  35. Wei, C.; Wang, W.; Yang, W.; Liu, J. Deep Retinex Decomposition for Low-Light Enhancement. arXiv 2018, arXiv:1808.04560. [Google Scholar] [CrossRef] [Scilit]
  36. Feng, H.; Wang, L.; Wang, Y.; Fan, H.; Huang, H. Learnability enhancement for low-light raw image denoising: A data perspective. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 370–387. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Wang, Y.; Yu, Y.; Yang, W.; Guo, L.; Chau, L.P.; Kot, A.C.; Wen, B. Exposurediffusion: Learning to expose for low-light image enhancement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2023; pp. 12438–12448. [Google Scholar]
  38. Jin, X.; Xiao, J.W.; Han, L.H.; Guo, C.; Zhang, R.; Liu, X.; Li, C. Lighting every darkness in two pairs: A calibration-free pipeline for raw denoising. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2023; pp. 13275–13284. [Google Scholar]
  39. Lamba, M.; Mitra, K. Restoring extremely dark images in real time. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2021; pp. 3487–3497. [Google Scholar]
  40. Monakhova, K.; Richter, S.R.; Waller, L.; Koltun, V. Dancing under the stars: Video denoising in starlight. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022; pp. 16241–16251. [Google Scholar]
  41. Zheng, N.; Zhou, M.; Dong, Y.; Rui, X.; Huang, J.; Li, C.; Zhao, F. Empowering low-light image enhancer through customized learnable priors. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2023; pp. 12559–12569. [Google Scholar]
  42. Hore, A.; Ziou, D. Image quality metrics: PSNR vs. SSIM. In Proceedings of the 20th International Conference on Pattern Recognition (ICPR); ACM: New York, NY, USA, 2010; pp. 2366–2369. [Google Scholar]
  43. Sara, U.; Akter, M.; Uddin, M.S. Image quality assessment through FSIM, SSIM, MSE and PSNR—A comparative study. J. Comput. Commun. 2019, 7, 8–18. [Google Scholar] [CrossRef]
  44. Boie, R.A.; Cox, I.J. An analysis of camera noise. IEEE Trans. Pattern Anal. Mach. Intell. 1992, 14, 671–674. [Google Scholar] [CrossRef] [Scilit]
  45. Blanter, Y.M.; Büttiker, M. Shot noise in mesoscopic conductors. Phys. Rep. 2000, 336, 1–166. [Google Scholar] [CrossRef] [Scilit]
  46. Schöberl, M.; Senel, C.; Fößel, S.; Bloss, H.; Kaup, A. Non-linear dark current fixed pattern noise compensation for variable frame rate moving picture cameras. In Proceedings of the 17th European Signal Processing Conference (EUSIPCO); IEEE: Piscataway, NJ, USA, 2009; pp. 268–272. [Google Scholar]
  47. Anaya, J.; Barbu, A. Renoir—A dataset for real low-light image noise reduction. J. Vis. Commun. Image Represent. 2018, 51, 144–154. [Google Scholar] [CrossRef] [Scilit]
  48. Weng, C.C.; Chen, H.; Fuh, C.S. A novel automatic white balance method for digital still cameras. In Proceedings of the 2005 IEEE International Symposium on Circuits and Systems (ISCAS); IEEE: Piscataway, NJ, USA, 2005; pp. 3801–3804. [Google Scholar]
Figure 1. Comparisons of various methods for extremely low-light images. We focus on two types of objects including caps of pens and the title of a paper. Here the title is “Practical Poissonian-Gaussian Noise Modeling and Fitting for Single-image Raw-Data”. (a) Unprocessed image. (b) The output of a traditional ISP. (c) The output of a machine learning method trained on a small-scale dataset. (d) The output of our NLV trained on the proposed DaN dataset. (e) Reference normal light image. We observe that our NLV with DaN yields much clearer images under the extremely low-light setting. The boxes of different colors are zoom areas, used for comparing the details’ restoration.
Figure 1. Comparisons of various methods for extremely low-light images. We focus on two types of objects including caps of pens and the title of a paper. Here the title is “Practical Poissonian-Gaussian Noise Modeling and Fitting for Single-image Raw-Data”. (a) Unprocessed image. (b) The output of a traditional ISP. (c) The output of a machine learning method trained on a small-scale dataset. (d) The output of our NLV trained on the proposed DaN dataset. (e) Reference normal light image. We observe that our NLV with DaN yields much clearer images under the extremely low-light setting. The boxes of different colors are zoom areas, used for comparing the details’ restoration.
Computers 15 00261 g001
Figure 2. The data annotation process. Firstly, the parameters are calibrated and then randomly selected into the noise model. The generated noise is added to the Ground Truth.
Figure 2. The data annotation process. Firstly, the parameters are calibrated and then randomly selected into the noise model. The generated noise is added to the Ground Truth.
Computers 15 00261 g002
Figure 3. (a) Statistical analysis of the proposed DaN Dataset. (b) Presentation of our DaN dataset. (c) Fine-grained feature data of our DaN.
Figure 3. (a) Statistical analysis of the proposed DaN Dataset. (b) Presentation of our DaN dataset. (c) Fine-grained feature data of our DaN.
Computers 15 00261 g003
Figure 4. (a) The noise model for generating noisy images. (bd) Illustration of the proposed NLV pipeline. The NLV pipeline encompasses Denoising Network (DN), White Balance (WB), and Color Correction (CC). DN is employed for noise reduction, WB is employed to counterbalance the chromatic aberration, and CC restores the true image color and fine-grained features. The arrows represent the direction of feature flow. For the specific layer structure, please refer to the legend at the bottom of the figure.
Figure 4. (a) The noise model for generating noisy images. (bd) Illustration of the proposed NLV pipeline. The NLV pipeline encompasses Denoising Network (DN), White Balance (WB), and Color Correction (CC). DN is employed for noise reduction, WB is employed to counterbalance the chromatic aberration, and CC restores the true image color and fine-grained features. The arrows represent the direction of feature flow. For the specific layer structure, please refer to the legend at the bottom of the figure.
Computers 15 00261 g004
Figure 5. Visual comparisons between our NLV and other cutting-edge methods. We employed the same ISP as ELD to amplify and post-process the input images. The PSNR and SSIM values of the images are presented below the images. The boxes of different colors are zoom areas, used for comparing the details’ restoration.
Figure 5. Visual comparisons between our NLV and other cutting-edge methods. We employed the same ISP as ELD to amplify and post-process the input images. The PSNR and SSIM values of the images are presented below the images. The boxes of different colors are zoom areas, used for comparing the details’ restoration.
Computers 15 00261 g005
Figure 6. Two pictures of the equipment for pool welding: one is the original and the other is the enhanced version.
Figure 6. Two pictures of the equipment for pool welding: one is the original and the other is the enhanced version.
Computers 15 00261 g006
Figure 7. Cross-visualization comparison.
Figure 7. Cross-visualization comparison.
Computers 15 00261 g007
Table 1. Comparisons between our DaN dataset and existing ones. DaN comprises four primary categories throughout the entire day, from morning till night. The fine-grained features signify the existence of data dedicated to fine-grained learning. A “” indicates that the dataset contains images of this category, while an “” indicates that it does not.
Table 1. Comparisons between our DaN dataset and existing ones. DaN comprises four primary categories throughout the entire day, from morning till night. The fine-grained features signify the existence of data dedicated to fine-grained learning. A “” indicates that the dataset contains images of this category, while an “” indicates that it does not.
DatasetScenariosMorningDaytimeNightFine-GrainedIlluminanceGround TruthTotal
ELD [19]OneNone>0.5 lux10 +
LOL [35]OneNone>10 lux5001000
LOLV2-real [24]TwoFew>10 lux7891578
LOLV2-sync [24]TwoNone>10 lux10002000
SID [18]TwoFew>0.1 lux4245094
DaN (Ours)FourPlentiful<0.1 lux4200 +
Table 2. Cross experiments on existing datasets and the proposed DaN dataset.
Table 2. Cross experiments on existing datasets and the proposed DaN dataset.
MethodTrain DatasetLOLV1LOLV2-RealLOLV2-SyncSIDDaN
PSNRSSIMLPIPSPSNRSSIMLPIPSPSNRSSIMLPIPSPSNRSSIMLPIPSPSNRSSIMLPIPS
SID [18]SID27.350.7360.46228.240.7210.47829.040.7760.39830.760.8100.35120.980.5120.672
DaN (ours)28.150.7420.44127.610.7420.44529.260.7350.41232.150.8510.30521.060.5250.658
ELD [19]SID35.420.8250.15835.120.8620.14236.110.8190.17636.300.8720.14524.160.5870.523
DaN (ours)37.210.8610.11237.650.8710.10838.010.8910.09839.440.9010.08624.230.5710.509
ExpDiff [37]SID35.210.8160.16734.520.7620.20237.260.8510.12635.000.8080.18626.620.6040.455
DaN (ours)36.040.8210.14935.260.8010.18540.020.9100.07238.020.8470.12127.020.6270.423
Retinexformer [25]SID35.120.8250.16232.850.7710.23835.260.8410.15234.440.8260.19622.250.5260.602
DaN (ours)35.230.8420.14833.740.8290.19837.680.8690.11536.120.8420.14522.850.5390.578
PWN [36]SID35.950.8270.14434.920.8260.15636.820.8790.10837.870.8340.13225.120.6380.412
DaN (ours)36.880.8210.13836.920.8820.10237.120.8620.12540.060.9220.07826.800.6310.382
LED [38]SID36.810.9020.08937.110.8970.09635.910.8430.14839.340.9310.06226.940.6490.364
DaN (ours)36.720.8940.09538.280.9110.07837.200.8700.11840.260.9300.05727.490.6680.345
NLV (ours)SID37.110.8950.08337.580.9100.07136.820.8530.12442.870.9430.04528.360.7090.312
DaN (ours)38.060.9020.07439.620.9200.06437.200.9120.10543.490.9570.03829.520.7320.286
Table 3. The ablation experiments of the NLV components.
Table 3. The ablation experiments of the NLV components.
MethodSIDDaN
PSNRSSIMPSNRSSIM
w/o WB22.420.26120.840.304
w/o CC23.870.44922.490.430
NLV (Ours)43.490.95729.520.732
Table 4. The ablation experiments of wavelet transform.
Table 4. The ablation experiments of wavelet transform.
MethodSIDDaN
PSNRSSIMPSNRSSIM
w/o transform40.630.91226.630.701
Fourier41.030.90327.030.703
Hilbert41.120.89127.120.719
Short-time Fourier40.210.91227.210.712
Wigner42.510.94228.510.715
Radon42.150.93628.150.726
Wavelet (Ours)43.490.95729.520.732
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sun, Q.; Liu, S.; Li, H.; Feng, Y.; Sun, L.; Lu, K.; Liu, K. DaN: A Comprehensive Semi-Real Dataset for Extreme Low-Light Image Enhancement. Computers 2026, 15, 261. https://doi.org/10.3390/computers15050261

AMA Style

Sun Q, Liu S, Li H, Feng Y, Sun L, Lu K, Liu K. DaN: A Comprehensive Semi-Real Dataset for Extreme Low-Light Image Enhancement. Computers. 2026; 15(5):261. https://doi.org/10.3390/computers15050261

Chicago/Turabian Style

Sun, Qiuyang, Shaonan Liu, Hong Li, Yingchao Feng, Liuqing Sun, Kun Lu, and Kangtai Liu. 2026. "DaN: A Comprehensive Semi-Real Dataset for Extreme Low-Light Image Enhancement" Computers 15, no. 5: 261. https://doi.org/10.3390/computers15050261

APA Style

Sun, Q., Liu, S., Li, H., Feng, Y., Sun, L., Lu, K., & Liu, K. (2026). DaN: A Comprehensive Semi-Real Dataset for Extreme Low-Light Image Enhancement. Computers, 15(5), 261. https://doi.org/10.3390/computers15050261

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop