Video Identifying and Eraser: Use Multi-Task Cascaded Convolutional Neural Network to Enhance Safety in a Text-to-Video Diffusion Model
Abstract
1. Introduction
- We introduce a plug-and-play framework that performs real-time face protection inside the text-to-video diffusion inference loop, providing an inline alternative to purely post-processing-based filtering without retraining the diffusion model.
- We develop a zero-shot integration of MTCNN-based face detection into Stable Diffusion inference, enabling efficient detection and intervention while preserving the original generative capability.
- We propose a compact landmark-based face alignment method using five fiducial points, which supports efficient masking localization and robust screening under common variations.
- We conduct comprehensive experiments on SVD-generated video sequences, evaluating both frame-level detection accuracy and video-level safety success rates, demonstrating effective protection with minimal additional latency in practical T2V settings.
2. Related Work
2.1. Diffusion Models for Controllable Generation and Identity-Centric Synthesis
2.2. Facial Recognition and Safety Mechanisms for Generative Models
3. Proposed Method
3.1. System Overview
- Sampling (Latent Extraction): At a designated diffusion timestep t, we extract the current noisy latent representation from U-Net. This latent code contains the semantic information generated up to step t.
- Decoding (Latent-to-Pixel): The latent decodes into pixel space using the SD VAE decoder to obtain an intermediate image . This conversion is necessary as facial detection algorithms operate in pixel space.
- Preprocessing (Partial Denoising): Since may contain significant noise at early timesteps, we apply a lightweight partial denoising filter to enhance structural clarity, facilitating robust feature detection.
- Detection & Verification: The preprocessed image feeds into the MTCNN network to detect facial bounding boxes. Identity verification ensures the protection mechanism applies only to authorized subjects.
- Masking (OSD Application): Upon successful verification, an on-screen display (OSD) mask is generated and applied to the facial regions. This creates a modified pixel image , where sensitive features are obscured according to Equation (5).
- Re-insertion (Pixel-to-Latent): The masked image re-encodes into latent space via the VAE encoder, yielding . To maintain consistency with the diffusion trajectory, noise is re-injected to match the variance schedule of timestep t:
- Completion (Resumed Sampling): The modified latent replaces the original . The reverse diffusion process resumes from timestep t down to 0, utilizing the protected latent variables to generate the final privacy-preserving video.
| Algorithm 1 Diffusion-Safety Pipeline for Privacy-Preserving Video Generation |
| Require: Text prompt p, total timesteps T Ensure: Privacy-preserving video V
|
3.2. Efficiency Discussion
3.3. Video-Level Safety Evaluation
3.4. Mathematical Framework
3.4.1. Stable Diffusion Fundamentals
3.4.2. MTCNN-Based Facial Detection
3.4.3. OSD Mask Generation
3.4.4. Image-to-Latent Mask Mapping
- Resolution Alignment: Stable Diffusion’s VAE compresses spatial dimensions by a factor of 8. We resize the pixel-space mask to match the latent resolution:
- Channel Broadcasting: The latent tensor has 4 channels (SD VAE). The single-channel mask is broadcast to match:
- Differentiability: We acknowledge that VAE encode/decode operations and MTCNN detection are not differentiable. Our method applies the mask as a post hoc intervention during sampling without requiring gradient flow through the detection module.
3.4.5. Modified Denoising Step
3.4.6. Re-Noising for Diffusion Consistency
3.5. Method-Specific Ablations for the Diffusion-Safety Framework
3.5.1. Ablation 1: Adaptive Weighting for Face Preservation
- No Adaptive Weighting (): This approach treats all regions uniformly, without distinguishing between face and non-face areas.
- With Adaptive Weighting ( as a learned parameter): In this configuration, the framework selectively prioritizes face regions, enhancing their preservation during the denoising process.
3.5.2. Ablation 2: Detection Model Choice
- MTCNN: The primary face detection model implemented in our framework.
- RetinaFace: A state-of-the-art detector optimized for high-accuracy face detection.
- YOLOv4: A fast and efficient object detection model capable of detecting faces among other objects.
3.5.3. Ablation 3: Impact of OSD Masking Strategy
- Small radius (): this setting results in compact masks, which may inadvertently exclude some facial regions.
- Large radius (): a larger radius leads to more extensive masks, which may encompass non-facial areas and potentially degrade image quality.
3.5.4. Ablation 4: Intervention Timestep

| Configuration | Safety Rate (%) | Latency (ms) | FID ↓ |
|---|---|---|---|
| SD-Base | 0.0 | 1280 | 25.3 |
| +MTCNN (SD-M) | 88.5 | 1298 | 26.1 |
| +OSD (SD-O) | 0.0 | 1285 | 25.8 |
| +Full (SD-MO) | 94.3 | 1298 | 26.5 |
3.6. Evaluation Protocol for T2V Safety
4. Experiments and Results
4.1. Ablation Studies and Analysis
- Synergistic Effects: The full model achieves performance beyond the sum of its individual components.
- Critical Components: Adaptive weighting and multi-scale fusion constitute the most impactful elements.
- Robustness: The method demonstrates stable performance across hyperparameter variations.
- Efficiency Trade-offs: The computational costs remain justified by the significant performance gains.
- Generalization: All components contribute to cross-dataset robustness.
4.2. Implementation Details and Deployment Scenario
- Generation of single faces ().
- Multi-face scenarios ().
- Faces with complex backgrounds.
- Cloud-side: Computationally intensive Stable Diffusion [1] generation runs on cloud servers (e.g., NVIDIA RTX 3090).
- Edge-side: Lightweight safety monitoring (MTCNN [8] detection + OSD masking) deploys on edge devices (e.g., NVIDIA Jetson Xavier).
- Target Frame Rate: The safety plugin supports monitoring at up to 25 fps on edge hardware, ensuring the safety module does not introduce a bottleneck.
- Generation Latency: The ∼1280 ms/frame latency stems from the Stable Diffusion [1] backbone, which is acceptable for offline or near-real-time applications.
- Plugin Overhead: Our safety module adds only ∼33 ms of overhead per frame (MTCNN [8] + OSD + VAE), operating efficiently relative to the generation process.
- Revised Terminology: We replace “real-time system” with “real-time safety monitoring” to accurately describe performance.
4.3. Efficiency Analysis
- Latency: Our safety plugin adds only 33 ms of overhead per frame (Table 6), which accounts for 2.6% of the total generation time. This lightweight design enables efficient deployment on edge devices without creating a bottleneck.
- Memory: The plugin requires only 55 MB of RAM (1.2% of SD’s 4500 MB VRAM), making it suitable for resource-constrained edge devices.
- Computation: With only 2.4 G FLOPs (0.7% of SD’s 350 G FLOPs), the plugin maintains minimal computational overhead.
4.4. Experimental Results
4.4.1. Experiment 1: Baseline Face Detection Performance
4.4.2. Experiment 2: Anomaly Filtering Performance
4.4.3. Experiment 3: Detection Efficiency and Latency
- Functional efficacy: A 96.5% face detection rate on generated video frames demonstrates strong coverage for privacy protection;
- Specificity: A false positive rate of only 1.8% on non-facial regions indicates minimal interference with legitimate content generation;
- Computational efficiency: MTCNN processes each frame in an average of 18.2 ms, contributing just 1.4% of total system latency (vs. 1280 ms for SD generation)—confirming its suitability for real-time deployment.

| Config. | Safety Rate (%) | Latency (ms) | FID ↓ |
|---|---|---|---|
| SD-Base | 0.0 | 1280 | 25.3 |
| +MTCNN (SD-M) | 88.5 | 1298 | 26.1 |
| +OSD (SD-O) | 0.0 | 1285 | 25.8 |
| +Full (SD-MO) | 94.3 | 1298 | 26.5 |
4.5. Discussion
4.5.1. Progressive Role Definition of MTCNN
4.5.2. Critical Influencing Factors
4.5.3. Impact of Findings
4.5.4. Impact of Findings
4.5.5. System Efficiency and Performance
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022. [Google Scholar]
- Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; et al. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv 2023, arXiv:2311.15127. [Google Scholar] [CrossRef]
- Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; Fleet, D.J. Video Diffusion Models. Adv. Neural Inf. Process. Syst. 2022, 35, 8633–8646. [Google Scholar]
- Bommasani, R.; Hudson, D.A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M.S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. On the Opportunities and Risks of Foundation Models. arXiv 2021, arXiv:2108.07258. [Google Scholar] [CrossRef]
- Schramowski, P.; Brack, M.; Deiseroth, B.; Kersting, K. Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 22522–22531. [Google Scholar]
- Rossler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; Nießner, M. FaceForensics++: Learning to Detect Manipulated Facial Images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1–11. [Google Scholar]
- Tolosana, R.; Vera-Rodriguez, R.; Fierrez, J.; Morales, A.; Ortega-Garcia, J. Deepfakes and Beyond: A Survey of Face Manipulation and Detection. Inf. Fusion 2020, 64, 131–148. [Google Scholar] [CrossRef]
- Zhang, K.; Zhang, Z.; Li, Z.; Qiao, Y. Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks. IEEE Signal Process. Lett. 2016, 23, 1499–1503. [Google Scholar] [CrossRef]
- Shen, F.; Xu, W.; Yan, R.; Zhang, D.; Shu, X.; Tang, J. IMAGEdit: Let Any Subject Transform. arXiv 2025, arXiv:2510.01186. [Google Scholar] [CrossRef]
- Shen, F.; Du, X.; Gao, Y.; Yu, J.; Cao, Y.; Lei, X.; Tang, J. IMAGHarmony: Controllable Image Editing with Consistent Object Quantity and Layout. arXiv 2025, arXiv:2506.01949. [Google Scholar] [CrossRef]
- Shen, F.; Yu, J.; Wang, C.; Jiang, X.; Du, X.; Tang, J. IMAGGarment-1: Fine-Grained Garment Generation for Controllable Fashion Design. arXiv 2025, arXiv:2504.13176. [Google Scholar] [CrossRef] [PubMed]
- Shen, F.; Tang, J. Imagpose: A Unified Conditional Framework for Pose-Guided Person Generation. Adv. Neural Inf. Process. Syst. 2024, 37, 6246–6266. [Google Scholar]
- Shen, F.; Ye, H.; Zhang, J.; Wang, C.; Han, X.; Wei, Y. Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models. In Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024; pp. 1–14. [Google Scholar]
- Shen, F.; Jiang, X.; He, X.; Ye, H.; Wang, C.; Du, X.; Li, Z.; Tang, J. Imagdressing-v1: Customizable Virtual Dressing. Proc. Aaai Conf. Artif. Intell. 2025, 39, 6795–6804. [Google Scholar] [CrossRef]
- Shen, F.; Wang, C.; Gao, J.; Guo, Q.; Dang, J.; Tang, J.; Chua, T.-S. Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model. In Proceedings of the 42nd International Conference on Machine Learning, Vienna, Austria, 13–19 July 2025. [Google Scholar]
- Deng, J.; Guo, J.; Xue, N.; Zafeiriou, S. RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 5203–5212. [Google Scholar]
- Redmon, J.; Farhadi, A. YOLOv3: An Incremental Improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef]
- Dhariwal, P.; Nichol, A. Diffusion Models Beat GANs on Image Synthesis. Adv. Neural Inf. Process. Syst. 2021, 34, 8780–8794. [Google Scholar]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Wey, T.; Andreetto, M.; Adam, H. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv 2017, arXiv:1704.04861. [Google Scholar] [CrossRef]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 4510–4520. [Google Scholar]
- Li, Y.; Yang, X.; Sun, P.; Qi, H.; Lyu, S. Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 3207–3216. [Google Scholar]
- Ojha, U.; Li, Y.; Singh, J.P. Towards Universal Fake Image Detectors that Generalize Across Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 24480–24489. [Google Scholar]
- Zhang, L.; Rao, A.; Agrawala, M. Adding Conditional Control to Text-to-Image Diffusion Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 3836–3847. [Google Scholar]
- Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; Aberman, K. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 22500–22510. [Google Scholar]
- Kawar, B.; Zada, S.; Lang, O.; Tov, O.; Chang, H.; Dekel, T.; Mosseri, I.; Irani, M. Imagic: Text-Based Real Image Editing with Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 6007–6017. [Google Scholar]
- Weidinger, L.; Mellor, J.; Rauh, M.; Griffin, C.; Uesato, J.; Huang, P.S.; Cheng, M.; Glaese, M.; Balle, B.; Kasirzadeh, A.; et al. Ethical and Social Risks of Harm from Language Models. arXiv 2021, arXiv:2112.04359. [Google Scholar] [CrossRef]
- Shan, S.; Cryan, J.; Wenger, E.; Zheng, H.; Hanocka, R.; Zhao, B.Y. Glaze: Protecting Artists from Style Mimicry by Text-to-Image Models. In Proceedings of the 33rd USENIX Security Symposium, Philadelphia, PA, USA, 14–16 August 2024; pp. 503–520. [Google Scholar]
- Luo, S.; Tan, Y.; Huang, L.; Li, J.; Zhao, H. Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference. arXiv 2023, arXiv:2310.04378. [Google Scholar]
- Lane, N.D.; Bhattacharya, S.; Georgiev, P.; Forlivesi, C.; Jiao, L.; Qendro, L.; Kawsar, F. DeepX: A Software Accelerator for Low-Power Deep Learning Inference on Mobile Devices. In Proceedings of the 15th ACM/IEEE International Conference on Information Processing in Sensor Networks, Vienna, Austria, 11–14 April 2016; pp. 1–12. [Google Scholar]
- Han, S.; Liu, X.; Mao, H.; Pu, J.; Pedram, A.; Horowitz, M.A.; Dally, W.J. EIE: Efficient Inference Engine on Compressed Deep Neural Network. In Proceedings of the ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Seoul, Republic of Korea, 18–22 June 2016; pp. 243–254. [Google Scholar]
- Deng, J.; Guo, J.; Xue, N.; Zafeiriou, S. ArcFace: Additive Angular Margin Loss for Deep Face Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 4690–4699. [Google Scholar]
- Chen, S.; Liu, Y.; Gao, X.; Han, Z. MobileFaceNets: Efficient CNNs for Face Recognition. arXiv 2018, arXiv:1804.07573. [Google Scholar]
- Han, S.; Mao, H.; Dally, W.J. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. In Proceedings of the International Conference on Learning Representations, San Juan, Puerto Rico, 2–4 May 2016; pp. 1–14. [Google Scholar]
- Yang, L.; Zhang, Z.; Song, Y.; Hong, S.; Xu, R.; Zhao, Y.; Zhang, W.; Cui, B.; Yang, M.H. Diffusion Models: A Comprehensive Survey of Methods and Applications. ACM Comput. Surv. 2023, 56, 1–39. [Google Scholar] [CrossRef]
- Ho, J.; Jain, A.; Abbeel, P. Denoising Diffusion Probabilistic Models. Adv. Neural Inf. Process. Syst. 2020, 33, 6840–6851. [Google Scholar]
- Salimans, T.; Ho, J. Progressive Distillation for Fast Sampling of Diffusion Models. In Proceedings of the International Conference on Learning Representations, Virtual Event, 25–29 April 2022; pp. 1–14. [Google Scholar]
- Shan, S.; Wenger, E.; Zhang, J.; Khaddaj, A.; Leclerc, G.; Ilyas, A.; Madry, A.; Bhagoji, A.N. PhotoGuard: Protecting Users from Unauthorized Use of Personal Images with AI. In Proceedings of the 32nd USENIX Security Symposium (USENIX Security 23), Anaheim, CA, USA, 9–11 August 2023; pp. 4563–4580. [Google Scholar]
- Poole, B.; Jain, A.; Barron, J.T.; Mildenhall, B. DreamFusion: Text-to-3D using 2D Diffusion. arXiv 2022, arXiv:2209.14988. [Google Scholar]
- Qian, Y.; Yin, G.; Sheng, L.; Chen, Z.; Shao, J. Thinking in Frequency: Face Forgery Detection by Mining Frequency-aware Clues. In Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23–28 August 2020; pp. 86–103. [Google Scholar]
- Ali, S.; Abuhmed, T.; El-Sappagh, S.; Muhammad, K.; Alonso-Moral, J.M.; Confalonieri, R.; Guidotti, R.; Del Ser, J.; Díaz-Rodríguez, N.; Herrera, F. Explainable Artificial Intelligence (XAI): What We Know and What is Left to Attain Trustworthy Artificial Intelligence. Expert Syst. Appl. 2023, 231, 120681. [Google Scholar] [CrossRef]




| Method | Detection Accuracy (%) ↑ | Latency (ms/Frame) ↓ |
|---|---|---|
| Heavy Detector (e.g., ResNet-50) | 97.1 | 125.4 |
| Ours (MTCNN) | 96.5 | 18.2 |
| MTCNN Face Detection: | |
|---|---|
| Input: Decoded image from latent at step t | |
| Output: Face bounding boxes | |
| Pnet ← Proposal Network(image) | Stage 1: Candidate regions |
| Stage 2: Bounding box regression | |
| Stage 3: Landmark detection | |
| return with confidence scores | |
| Variant | MTCNN | OSD | Modification |
|---|---|---|---|
| SD-Base | – | – | Original SD |
| SD-M | ✓ | – | Face detection only |
| SD-O | – | ✓ | Mask guidance only |
| SD-MO | ✓ | ✓ | Full method |
| SD-FT | ✓ | ✓ | With fine-tuning |
| Metric | SD-Base | SD-M | SD-O | SD-MO |
|---|---|---|---|---|
| FPS (fps) | 62.1 | 88.3 | 71.2 | 93.7 |
| BA (%) | 81 | 75 | 92 | 89 |
| Latency (ms) | 980 | 1240 | 1050 | 1280 |
| UP (%) | 65 | 72 | 68 | 86 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lin, S.; Zhou, R.; Wang, Y. Video Identifying and Eraser: Use Multi-Task Cascaded Convolutional Neural Network to Enhance Safety in a Text-to-Video Diffusion Model. Appl. Sci. 2026, 16, 2995. https://doi.org/10.3390/app16062995
Lin S, Zhou R, Wang Y. Video Identifying and Eraser: Use Multi-Task Cascaded Convolutional Neural Network to Enhance Safety in a Text-to-Video Diffusion Model. Applied Sciences. 2026; 16(6):2995. https://doi.org/10.3390/app16062995
Chicago/Turabian StyleLin, Shuang, Ranran Zhou, and Yong Wang. 2026. "Video Identifying and Eraser: Use Multi-Task Cascaded Convolutional Neural Network to Enhance Safety in a Text-to-Video Diffusion Model" Applied Sciences 16, no. 6: 2995. https://doi.org/10.3390/app16062995
APA StyleLin, S., Zhou, R., & Wang, Y. (2026). Video Identifying and Eraser: Use Multi-Task Cascaded Convolutional Neural Network to Enhance Safety in a Text-to-Video Diffusion Model. Applied Sciences, 16(6), 2995. https://doi.org/10.3390/app16062995

