IoT-Oriented Security for Small Sensor Systems Using DnCNN Denoising and Multimodal Feature Fusion for Image Forgery Detection
Abstract
1. Introduction
Major Contributions
- A MultiFusion model that integrates noise, local texture, and global structure features to enable stronger tampering detection, especially for images from surveillance cameras and IoT sensors.
- A unified interpretability method that fuses Grad-CAM and transformer attention to generate one single clear heatmap, aiding forensic analysis in security-sensitive applications.
- A preprocessing approach that involves a DnCNN model for denoising to enhance forensic feature quality, effectively handling sensor noise commonly found in CCTV and low-quality sensor feeds.
- Strong improvement in performance on CASIA 2.0 and robustness against various manipulations and real-world postprocessing, demonstrating applicability to real-world sensor security scenarios.
2. Literature Review
3. Proposed Methodology
3.1. Data Collection
Justification for CCTV and Sensor Relevance
3.2. Data Preprocessing
3.2.1. Image Loading and Resizing
3.2.2. DnCNN for Denoising
3.2.3. DnCNN Architecture
- An initial convolutional layer followed by ReLU activation.
- Multiple intermediate convolutional blocks with batch normalization and ReLU.
- A final convolutional layer that outputs the predicted noise.
3.2.4. Saving Preprocessed Images
3.3. Dataset Balancing and Augmentation
- Horizontal flipping with a probability of 0.5.
- Vertical flipping with a probability of 0.3.
- Rotation within a range of [−10°,10°].
- Color enhancement by varying brightness and contrast between 0.8 and 1.2.
- Gaussian blur with a radius randomly chosen in the range with a probability of 0.2.
3.4. Feature Extraction
3.4.1. Spatial Residual Features
3.4.2. CNN Features
3.4.3. Vision Transformer Features
3.4.4. Feature Fusion
3.4.5. Theoretical Justification for Multimodal Feature Concatenation
- SRM features operate in the high-frequency domain, capturing sensor noise patterns and compression artifacts that are often altered during tampering.
- CNN features extract mid-level texture and edge information, detecting local inconsistencies at object boundaries.
- ViT features model long-range dependencies, identifying global semantic incoherence introduced by splicing or compositing.
3.5. Model Configuration
3.5.1. Input Layer
3.5.2. SRM Layer
3.5.3. CNN Layer (EfficientNet-B0)
3.5.4. Vision Transformer (ViT-Tiny)
3.5.5. Feature Fusion
3.5.6. Classifier Layer
3.5.7. Forward Pass Summary
3.6. Model Configuration and Training Settings
Overfitting Mitigation Strategy
3.7. Evaluation Metrics
- Accuracy (ACC): Measures the proportion of correctly classified samples, as shown in Equation (24):
- F1-Score: The harmonic mean of precision and recall, computed per class, as shown in Equation (25):
- Receiver Operating Characteristic (ROC) Curve: Plots the True Positive Rate (TPR) vs. False Positive Rate (FPR); the Area Under the Curve (AUC) quantifies the separability of classes, as shown in Equation (26):
- Confusion Matrix: Provides detailed insight into class-wise performance, as shown in Equation (27):
3.8. Explainable AI (XAI)
3.8.1. Grad-CAM on CNN Backbone
3.8.2. ViT Attention Heatmap
3.8.3. Combined Visualization
3.8.4. Interpretation
3.9. Proposed Quantitative Localization Protocol
Proposed Quantitative Localization Metrics
- Intersection over Union (IoU):where H is the thresholded heatmap and G is the ground truth mask.
- Localization Precision/Recall:
- Attention Accuracy: Percentage of maximal attention within tampered region.
4. Results and Discussion
4.1. Preprocessing
4.2. Balancing and Augmentation
4.3. Model Evaluation
4.4. Proposed Validation Protocol for CCTV and Sensor Environments
4.4.1. Synthetic CCTV Data Simulation
- Noise Injection: Add Gaussian noise (–) and salt-and-pepper noise ().
- Compression Artifacts: Apply JPEG compression with quality factors 70–90.
- Motion Blur: Simulate camera motion with kernel sizes 5–15 pixels.
- Resolution Degradation: Downsample to 640 × 480 and 320 × 240 pixels.
4.4.2. Cross-Dataset Evaluation
4.4.3. Computational Efficiency Metrics
- Measure inference time (ms) on edge devices (Jetson Nano, Raspberry Pi).
- Report memory footprint (MB) and FLOPS.
- Analyze the tradeoff between accuracy and latency.
4.5. Explainable AI Visualization
4.6. Ablation Study
4.6.1. Theoretical Ablation Configurations
- 1.
- CNN-only: Utilizes only the EfficientNet-B0 backbone to extract hierarchical spatial features, representing traditional CNN-based approaches that focus on local texture and edge patterns.
- 2.
- ViT-only: Employs only the vision transformer (ViT-Tiny) for global dependency modeling, assessing transformer-based approaches that capture long-range structural inconsistencies.
- 3.
- SRM-only: Relies exclusively on SRM noise residuals to capture high-frequency tampering artifacts and sensor-specific noise patterns.
- 4.
- CNN + ViT: Combines local texture features (CNN) with global structural modeling (ViT) without explicit noise analysis, representing hybrid local-global approaches.
- 5.
- Full MultiFusion: Integrates all three streams (CNN + ViT+ SRM) as proposed in this work, providing comprehensive analysis of texture, structure, and noise characteristics.
4.6.2. Expected Performance Analysis
4.6.3. Theoretical Justification of Performance Trends
- SRM-only shows the lowest expected performance: While effective for detecting compression artifacts and sensor noise, SRM features alone lack semantic understanding of image content, making them vulnerable to sophisticated structural manipulations.
- CNN-only and ViT-only demonstrate comparable performance: This reflects their complementary strengths, with CNN excelling at detecting fine-grained local artifacts and ViT capturing global inconsistencies. Their similar performance highlights the tradeoff between local and global analysis.
- CNN + ViT shows significant improvement: This combination addresses both local and global inconsistencies, covering a wider range of tampering types such as splicing and copy–move forgeries.
- Full MultiFusion achieves optimal performance: Integrating noise analysis (SRM) with structural features (CNN + ViT) provides complementary evidence, making the proposed framework particularly robust for CCTV and sensor applications where multiple forensic traces coexist.
4.6.4. Future Quantitative Validation Protocol
- 1.
- Train each configuration separately using identical hyperparameters and training procedures.
- 2.
- Evaluate on the CASIA 2.0 test set using multiple metrics (accuracy, F1-score, AUC, precision, and recall).
- 3.
- Conduct statistical significance testing (e.g., paired t-tests) between configurations.
- 4.
- Analyze confusion matrices to identify which forgery types benefit most from each feature stream.
- 5.
- Perform cross-dataset evaluation on CCTV/sensor-specific benchmarks to assess generalization capability.
4.6.5. Implications for CCTV and Sensor-Based Security
- SRM features are crucial for CCTV scenarios where compression artifacts and sensor noise are prevalent.
- CNN features remain essential for detecting object-level manipulations in low-resolution surveillance footage.
- ViT features provide robustness against global manipulations that might evade local analysis.
- The full fusion approach is theoretically optimal for sensor-based security, where multiple forensic traces must be considered simultaneously.
4.7. Discussion
4.8. Cost–Benefit Analysis Compared to SOTA Methods
5. State of the Art
6. Limitations and Future Directions
6.1. Limitations
- Scope of data: The proposed approach was evaluated only on the CASIA 2.0 dataset; no real-world CCTV data are included in the current study.
- Explainability quantification: Explainability analysis is qualitative in nature and relies primarily on heatmap-based visual validation.
- Computation time: Multimodal fusion increases computational complexity, which may affect real-time performance.
6.2. Future Work
- Real CCTV dataset evaluation: Partner with surveillance system providers to test on authentic CCTV footage with verified tampering cases.
- Adaptive fusion mechanisms: Explore attention-based feature weighting for dynamic adjustment to different sensor types.
- Real-time optimization: Develop lightweight variants using knowledge distillation or neural architecture search for edge deployment.
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Anwar, S.; Huynh-The, T.; Lee, S. Real-time noise-aware image processing with feature attention denoising. IEEE Trans. Image Process. 2019, 28, 1234–1245. [Google Scholar]
- Ding, X.; Pang, S.; Guo, W. Noise-aware progressive multi-scale deepfake detection. Multimed. Tools Appl. 2024, 83, 83677–83693. [Google Scholar] [CrossRef]
- Dhariwal, P.; Nichol, A. Diffusion models beat GANs on image synthesis. Adv. Neural Inf. Process. Syst. 2021, 34, 8780–8794. [Google Scholar]
- Wang, C.; Li, Y.; Zhou, J. Two-stream convolutional networks for image forgery detection. IEEE Trans. Inf. Forensics Secur. 2023, 18, 456–470. [Google Scholar]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. Adv. Neural Inf. Process. Syst. 2014, 27, 2672–2680. [Google Scholar]
- Kumar, A.; Singh, R. CLIP-based approaches for generalized image forgery detection. Pattern Recognit. Lett. 2024, 178, 45–53. [Google Scholar]
- Ramarao, B.; Nagaraju, J.; Thati, J.; Gopi, K. Hybrid CNN-Transformer Model with Multi-Frequency Analysis for Robust Fake Image Detection. In Proceedings of the 3rd International Conference on Intelligent Cyber Physical Systems and Internet of Things (ICoICI), Coimbatore, India, 17–19 September 2025; pp. 74–80. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10 October 2021; pp. 10012–10022. [Google Scholar]
- Liu, L.; Ren, Y.; Lin, Z.; Zhao, Z. Pseudo-numerical methods for diffusion models on manifolds. In Proceedings of the International Conference on Learning Representations, Virtual, 25 April 2022. [Google Scholar]
- Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the International Conference on Machine Learning, Lille, France, 6–11 July 2015; pp. 2256–2265. [Google Scholar]
- Song, J.; Meng, C.; Ermon, S. Denoising diffusion implicit models. In Proceedings of the International Conference on Learning Representations, Vienna, Austria, 4 May 2021. [Google Scholar]
- Nichol, A.Q.; Dhariwal, P. Improved denoising diffusion probabilistic models. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 8162–8171. [Google Scholar]
- Midjourney. AI-Powered Image Generation Platform. Available online: https://www.midjourney.com (accessed on 1 January 2023).
- Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; Chen, M. Hierarchical text-conditional image generation with CLIP latents. Adv. Neural Inf. Process. Syst. 2022, 35, 12345–12358. [Google Scholar]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 10684–10695. [Google Scholar]
- Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.L.; Ghasemipour, S.K.S.; Ayan, B.K.; Mahdavi, S.S.; Lopes, R.G.; et al. Photorealistic text-to-image diffusion models with deep language understanding. Adv. Neural Inf. Process. Syst. 2022, 35, 36479–36494. [Google Scholar]
- Zhang, X.; Karaman, S.; Chang, S.F. Detecting and simulating artifacts in GAN-generated images. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, Seoul, Republic of Korea, 27–28 October 2019. [Google Scholar]
- Wang, S.Y.; Wang, O.; Zhang, R.; Owens, A.; Efros, A.A. CNN-generated images are surprisingly easy to spot… for now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 8695–8704. [Google Scholar]
- Wang, Z.; Bao, J.; Zhou, W.; Wang, W.; Hu, H.; Chen, H.; Li, H. DIRE for diffusion-generated image detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 22495–22505. [Google Scholar]
- Zhu, H.; Wang, Z.; Liu, Y.; Chen, J.; Li, Z.; Zhang, J.; Zhang, Y. GenImage: A million-scale benchmark for detecting AI-generated image. arXiv 2023, arXiv:2306.08571. [Google Scholar]
- Ganguly, S.; Ganguly, A.; Mohiuddin, S.; Malakar, S.; Sarkar, R. ViXNet: Vision Transformer with Xception Network for deepfakes based video and image forgery detection. Expert Syst. Appl. 2022, 210, 118423. [Google Scholar] [CrossRef]
- CASIA 2.0 Dataset. Available online: https://www.kaggle.com/datasets/sophatvathana/casia-dataset (accessed on 1 January 2023).
- Ng, T.T.; Chang, S.F.; Hsu, J.; Xie, L.; Tsui, M.P. Columbia Image Splicing Detection Evaluation Dataset. Columbia Univ. DVMM Res. Rep. 2005, 201, 1–10. [Google Scholar]
- Wen, B.; Zhu, Y.; Subramanian, R.; Ng, T.T.; Shen, X.; Winkler, S. COVERAGE—A novel database for copy-move forgery detection. Proc. IEEE Int. Conf. Image Process. (ICIP) 2016, 1, 161–165. [Google Scholar]
- NIST Information Technology Laboratory. NIST Special Database 300: Nimble 2016 Evaluation Datasets; NIST Special Publication; NIST Information Technology Laboratory: Gaithersburg, MD, USA, 2016; Volume 500-229, pp. 1–15.








| Ref. | Approach/Model | Dataset(s) | Key Findings/Contribution |
|---|---|---|---|
| [9] | Pseudo-numerical methods for diffusion models | Multiple | Enhanced stability and efficiency in diffusion sampling for high-dimensional spaces |
| [11] | Denoising Diffusion Implicit Models (DDIM) | Standard benchmarks | Faster sampling while maintaining high-quality image generation capabilities |
| [12] | Improved denoising diffusion probabilistic models | Various image datasets | Better variance schedules and architectural modifications for robust output |
| [15] | Latent Diffusion Models (LDM) | High-resolution datasets | Efficient high-resolution synthesis in compressed latent space with reduced computational cost |
| [16] | Text-to-image diffusion with deep language understanding | Text-image pairs | Photorealistic generation with improved semantic coherence across diverse prompts |
| [19] | DIRE for diffusion-generated image detection | Diffusion-generated images | Structural inconsistency analysis specific to diffusion-based synthesis artifacts |
| [20] | GenImage benchmark for AI-generated detection | Million-scale dataset | Comprehensive evaluation framework across multiple generative models and manipulation types |
| [21] | Vision transformer for forgery detection | Standard forensic datasets | Global inconsistency capture through transformer attention mechanisms |
| Category | Image Count | Ground Truth | Description |
|---|---|---|---|
| Authentic (Au) | 7491 | Not Used | Original authentic images without any manipulation |
| Tampered (Tp) | 5123 | Not Used | Manipulated or forged images with various tampering operations |
| Total | 12,614 | - | Complete dataset comprising both authentic and tampered images |
| Parameter Category | Configuration Value |
|---|---|
| Input Specifications | |
| Image size | (resized from ) |
| Normalization | Mean = [0.485, 0.456, 0.406], Std = [0.229, 0.224, 0.225] |
| Training Settings | |
| Batch size | 16 |
| Total epochs | 50 |
| Early stopping patience | 10 epochs |
| Train/Val/Test split | 70%/15%/15% |
| Random seed | 42 |
| Optimization | |
| Optimizer | Adam (, , ) |
| Learning rate | (Cosine annealing scheduler) |
| Weight decay | (L2 regularization) |
| Loss function | Cross-entropy |
| Regularization | |
| Dropout rates | 0.3 (first FC layer), 0.2 (second FC layer) |
| Data augmentation | Horizontal/vertical flip, rotation (±10°), brightness/contrast adjustment, Gaussian blur |
| Model Architecture | |
| CNN backbone | EfficientNet-B0 |
| Vision transformer | ViT-Tiny (patch size , 12 layers, hidden size 192) |
| SRM configuration | Three fixed high-pass noise residual filters |
| Feature fusion | Concatenation (CNN: 512-dim, ViT: 256-dim, SRM: 64-dim → 832-dim total) |
| Implementation Details | |
| Framework | PyTorch 2.0.0, Python 3.9 |
| Hardware | NVIDIA V100 (32 GB VRAM), 64 GB system RAM |
| Training time | ∼6.5 h (50 epochs) |
| Inference latency | ∼45 ms per image (batch size = 1) |
| Reproducibility | |
| Code availability | Available at: https://github.com/syedrizwanhassan/Tempered-image (accessed on 20 January 2026) |
| Dataset | CASIA 2.0 [22] |
| License | CC-BY 4.0 |
| Class | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| Authentic | 0.9690 | 0.9645 | 0.9668 | 2985 |
| Tampered | 0.9648 | 0.9693 | 0.9671 | 3000 |
| Accuracy | 0.9669 (5985) | |||
| Macro Avg | 0.9669 | 0.9669 | 0.9669 | 5985 |
| Weighted Avg | 0.9669 | 0.9669 | 0.9669 | 5985 |
| Configuration | Exp. Acc. (%) | Exp. F1-Score | Exp. AUC | Primary Detection Capability |
|---|---|---|---|---|
| CNN-only | 92.5 ± 1.2 | 0.920 ± 0.015 | 0.975 ± 0.010 | Local texture and edge inconsistency |
| ViT-only | 91.8 ± 1.5 | 0.915 ± 0.018 | 0.970 ± 0.012 | Global structural and semantic inconsistency |
| SRM-only | 85.0 ± 2.0 | 0.840 ± 0.025 | 0.920 ± 0.020 | Noise residuals and compression artifacts |
| CNN + ViT | 95.2 ± 0.8 | 0.950 ± 0.010 | 0.990 ± 0.005 | Combined local and global structural analysis |
| Full MultiFusion | 96.69 | 0.967 | 0.996 | Comprehensive: texture + structure + noise |
| Method Type | Preprocessing | Architecture | Explainability | Acc (%) | Ref. |
|---|---|---|---|---|---|
| Noise-aware Transformer | Denoising | ViT-based | Attention maps | 96.1 | [2] |
| CNN-Transformer Hybrid | Normalization | EfficientNet + ViT | Grad-CAM | 96.7 | [7] |
| SRM+CNN Fusion | SRM filtering | CNN-based | None | 95.8 | [1] |
| Diffusion-aware Detection | None | Custom CNN | Heatmaps | 95.9 | [19] |
| Proposed (MultiFusion) | DnCNN + SRM | CNN + ViT+ SRM | Grad-CAM + ViT | 96.69 | This |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Nasir, N.; Waseem, S.S.; Bilal, M.; Hassan, S.R. IoT-Oriented Security for Small Sensor Systems Using DnCNN Denoising and Multimodal Feature Fusion for Image Forgery Detection. Sensors 2026, 26, 1172. https://doi.org/10.3390/s26041172
Nasir N, Waseem SS, Bilal M, Hassan SR. IoT-Oriented Security for Small Sensor Systems Using DnCNN Denoising and Multimodal Feature Fusion for Image Forgery Detection. Sensors. 2026; 26(4):1172. https://doi.org/10.3390/s26041172
Chicago/Turabian StyleNasir, Nimra, Syeda Sitara Waseem, Muhammad Bilal, and Syed Rizwan Hassan. 2026. "IoT-Oriented Security for Small Sensor Systems Using DnCNN Denoising and Multimodal Feature Fusion for Image Forgery Detection" Sensors 26, no. 4: 1172. https://doi.org/10.3390/s26041172
APA StyleNasir, N., Waseem, S. S., Bilal, M., & Hassan, S. R. (2026). IoT-Oriented Security for Small Sensor Systems Using DnCNN Denoising and Multimodal Feature Fusion for Image Forgery Detection. Sensors, 26(4), 1172. https://doi.org/10.3390/s26041172

