Coordinating Drag-Based Structure Editing and Reference Style Transfer in Diffusion Models for Anime Images
Abstract
1. Introduction
- The paper studies joint anime drag and style editing, where the output must satisfy both specified local deformation and reference-guided anime style transfer.
- It introduces a temporal anchor handoff strategy that separates drag optimization from style replay and resumes reference injection from a post drag clean sample anchor.
- It proposes anchor conditioned correspondence replay, which refreshes content queries from the edited anchor and routes reference style information toward compatible local regions without external parsing networks or semantic labels. The framework is evaluated on a curated anime benchmark through controlled quantitative comparisons, component ablations, and a blind forced choice study with 36 participants.
2. Related Work
2.1. Reference Style Transfer in Diffusion Models
2.2. Drag-Based Image Editing with Diffusion Models
2.3. Attention-Based Style Injection and Local Correspondence
2.4. Diffusion Trajectories and Stage Coordination
3. Method
3.1. Overview
| Algorithm 1: AnchorHandoff for temporally coordinated anime drag and style editing |
Input: Content image , style reference , control pairs , and mask Output: Edited result Fit a lightweight LoRA [43] adapter to for the drag stage; Invert into by DDIM inversion [41];
![]() Unload LoRA, refresh from the anchor, and initialize replay with AdaIN [23];
![]() |
3.2. Structure Editing by Drag Optimization
3.3. Style Attention Injection and Temporal Coordination
3.4. Clean Sample Anchor Handoff
3.5. Style Replay from the Anchor
3.6. Semantic Correspondence Modulation
4. Experiments
4.1. Experimental Settings
4.2. Baselines and Experimental Design
4.3. Evaluation Metrics
4.4. Quantitative Comparison
4.5. Computational Cost
4.6. Ablation Study
4.7. Human Evaluation
4.8. Qualitative Results
5. Discussion
Operating Boundaries
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Ho, J.; Jain, A.; Abbeel, P. Denoising Diffusion Probabilistic Models. In Proceedings of the Advances in Neural Information Processing Systems, Virtual, 6–12 December 2020; Volume 33, pp. 6840–6851. [Google Scholar] [CrossRef]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 10674–10685. [Google Scholar] [CrossRef]
- Song, W.; Ma, W.; Zhang, M.; Zhang, Y.; Zhao, X. Lightweight Diffusion Models: A Survey. Artif. Intell. Rev. 2024, 57, 161. [Google Scholar] [CrossRef]
- Croitoru, F.A.; Hondru, V.; Ionescu, R.T.; Shah, M. Diffusion Models in Vision: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 10850–10869. [Google Scholar] [CrossRef] [PubMed]
- Chen, H.; Xiang, Q.; Hu, J.; Ye, M.; Yu, C.; Cheng, H.; Zhang, L. Comprehensive Exploration of Diffusion Models in Image Generation: A Survey. Artif. Intell. Rev. 2025, 58, 99. [Google Scholar] [CrossRef]
- Zhang, L.; Rao, A.; Agrawala, M. Adding Conditional Control to Text-to-Image Diffusion Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 3813–3824. [Google Scholar] [CrossRef]
- Watanabe, Y.; Togo, R.; Maeda, K.; Ogawa, T.; Haseyama, M. Text-Guided Image Editing Based on Post Score for Gaining Attention on Social Media. Sensors 2024, 24, 921. [Google Scholar] [CrossRef] [PubMed]
- Yu, X.; Gu, X.; Hu, X.; Sun, J. ENGDM: Enhanced Non-Isotropic Gaussian Diffusion Model for Progressive Image Editing. Sensors 2025, 25, 2970. [Google Scholar] [CrossRef] [PubMed]
- Ruder, M.; Dosovitskiy, A.; Brox, T. Artistic Style Transfer for Videos and Spherical Images. Int. J. Comput. Vis. 2018, 126, 1199–1219. [Google Scholar] [CrossRef]
- Jing, Y.; Yang, Y.; Feng, Z.; Ye, J.; Yu, Y.; Song, M. Neural Style Transfer: A Review. IEEE Trans. Vis. Comput. Graph. 2020, 26, 3365–3385. [Google Scholar] [CrossRef] [PubMed]
- Selim, A.; Elgharib, M.; Doyle, L. Painting Style Transfer for Head Portraits Using Convolutional Neural Networks. ACM Trans. Graph. 2016, 35, 129. [Google Scholar] [CrossRef]
- Zhao, H.; Zheng, J.; Wang, Y.; Yuan, X.; Li, Y. Portrait Style Transfer Using Deep Convolutional Neural Networks and Facial Segmentation. Comput. Electr. Eng. 2020, 85, 106655. [Google Scholar] [CrossRef]
- Wang, L. Cartoon-Style Image Rendering Transfer Based on Neural Networks. Comput. Intell. Neurosci. 2022, 2022, 2958338. [Google Scholar] [CrossRef] [PubMed]
- Asperti, A.; Colasuonno, G.; Guerra, A. Portrait Reification with Generative Diffusion Models. Appl. Sci. 2023, 13, 6487. [Google Scholar] [CrossRef]
- Everaert, M.N.; Bocchio, M.; Arpa, S.; Süsstrunk, S.; Achanta, R. Diffusion in Style. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 2251–2261. [Google Scholar] [CrossRef]
- Wang, Z.; Zhao, L.; Xing, W. StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 7643–7655. [Google Scholar] [CrossRef]
- Chung, J.; Hyun, S.; Heo, J.P. Style Injection in Diffusion: A Training-Free Approach for Adapting Large-Scale Diffusion Models for Style Transfer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 8795–8805. [Google Scholar] [CrossRef]
- Alaluf, Y.; Garibi, D.; Patashnik, O.; Averbuch-Elor, H.; Cohen-Or, D. Cross-Image Attention for Zero-Shot Appearance Transfer. In Proceedings of the ACM SIGGRAPH 2024 Conference Papers; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar] [CrossRef]
- Ye, H.; Zhang, J.; Liu, S.; Han, X.; Yang, W. IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models. arXiv 2023, arXiv:2308.06721. [Google Scholar] [CrossRef]
- Pan, X.; Tewari, A.; Leimkühler, T.; Liu, L.; Meka, A.; Theobalt, C. Drag Your GAN: Interactive Point-Based Manipulation on the Generative Image Manifold. In Proceedings of the ACM SIGGRAPH 2023 Conference Proceedings; Association for Computing Machinery: New York, NY, USA, 2023. [Google Scholar] [CrossRef]
- Shi, Y.; Xue, C.; Liew, J.H.; Pan, J.; Yan, H.; Zhang, W.; Tan, V.Y.F.; Bai, S. DragDiffusion: Harnessing Diffusion Models for Interactive Point-Based Image Editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 8839–8849. [Google Scholar] [CrossRef]
- Gatys, L.A.; Ecker, A.S.; Bethge, M. Image Style Transfer Using Convolutional Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 2414–2423. [Google Scholar] [CrossRef]
- Huang, X.; Belongie, S. Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 1501–1510. [Google Scholar] [CrossRef]
- Cai, Q.; Ma, M.; Wang, C.; Li, H. Image Neural Style Transfer: A Review. Comput. Electr. Eng. 2023, 108, 108723. [Google Scholar] [CrossRef]
- Xu, Y.; Xia, M.; Hu, K.; Zhou, S.; Weng, L. Style Transfer Review: Traditional Machine Learning to Deep Learning. Information 2025, 16, 157. [Google Scholar] [CrossRef]
- Hertz, A.; Voynov, A.; Fruchter, S.; Cohen-Or, D. Style Aligned Image Generation via Shared Attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 4775–4785. [Google Scholar] [CrossRef]
- Kang, M.; Choi, Y.S. FreeMix: Personalized Structure and Appearance Control Without Finetuning. Appl. Sci. 2025, 15, 9889. [Google Scholar] [CrossRef]
- Yu, C.; Han, C.; Zhang, C. Multi-Source Training-Free Controllable Style Transfer via Diffusion Models. Symmetry 2025, 17, 290. [Google Scholar] [CrossRef]
- Xiang, Z.; Wan, X.; Xu, L.; Yu, X.; Mao, Y. A Training-Free Latent Diffusion Style Transfer Method. Information 2024, 15, 588. [Google Scholar] [CrossRef]
- Ling, P.; Chen, L.; Zhang, P.; Chen, H.; Jin, Y.; Zheng, J. FreeDrag: Feature Dragging for Reliable Point-Based Image Editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 6860–6870. [Google Scholar] [CrossRef]
- Liu, H.; Xu, C.; Yang, Y.; Zeng, L.; He, S. Drag Your Noise: Interactive Point-Based Editing via Diffusion Semantic Propagation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 6743–6752. [Google Scholar] [CrossRef]
- Cui, Y.; Zhao, X.; Zhang, G.; Cao, S.; Ma, K.; Wang, L. StableDrag: Stable Dragging for Point-Based Image Editing. In Proceedings of the Computer Vision—ECCV 2024; Springer: Cham, Switzerland, 2024; pp. 340–356. [Google Scholar] [CrossRef]
- Zhang, Z.; Liu, H.; Chen, J.; Xu, X. GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models. In Proceedings of the Thirteenth International Conference on Learning Representations, Singapore, 24–28 April 2025. [Google Scholar] [CrossRef]
- Mou, C.; Wang, X.; Song, J.; Shan, Y.; Zhang, J. DragonDiffusion: Enabling Drag-Style Manipulation on Diffusion Models. In Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024. [Google Scholar] [CrossRef]
- Guo, M.H.; Xu, T.X.; Liu, J.J.; Liu, Z.N.; Jiang, P.T.; Mu, T.J.; Zhang, S.H.; Martin, R.R.; Cheng, M.M.; Hu, S.M. Attention Mechanisms in Computer Vision: A Survey. Comput. Vis. Media 2022, 8, 331–368. [Google Scholar] [CrossRef]
- Khan, S.; Naseer, M.; Hayat, M.; Zamir, S.W.; Khan, F.S.; Shah, M. Transformers in Vision: A Survey. ACM Comput. Surv. 2022, 54, 200. [Google Scholar] [CrossRef]
- Tumanyan, N.; Geyer, M.; Bagon, S.; Dekel, T. Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 1921–1930. [Google Scholar] [CrossRef]
- Cao, M.; Wang, X.; Qi, Z.; Shan, Y.; Qie, X.; Zheng, Y. MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 22503–22513. [Google Scholar] [CrossRef]
- Pan, Z.; Kuang, Y.; Lan, J.; Zhang, L. High-Precision Image Editing via Dual Attention Control in Diffusion Models Without Fine-Tuning. Appl. Sci. 2025, 15, 1079. [Google Scholar] [CrossRef]
- Jiang, R.; Zheng, G.; Li, T.; Yang, T.; Wang, J.; Li, X. A Survey of Multimodal Controllable Diffusion Models. J. Comput. Sci. Technol. 2024, 39, 509–541. [Google Scholar] [CrossRef]
- Song, J.; Meng, C.; Ermon, S. Denoising Diffusion Implicit Models. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar] [CrossRef]
- Mokady, R.; Hertz, A.; Aberman, K.; Pritch, Y.; Cohen-Or, D. Null-Text Inversion for Editing Real Images Using Guided Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 6038–6047. [Google Scholar] [CrossRef]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar] [CrossRef]
- Kolkin, N.; Salavon, J.; Shakhnarovich, G. Style Transfer by Relaxed Optimal Transport and Self-Similarity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 10051–10060. [Google Scholar] [CrossRef]
- Zhu, M.; He, X.; Wang, N.; Wang, X.; Gao, X. All-to-Key Attention for Arbitrary Style Transfer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 23109–23119. [Google Scholar] [CrossRef]
- Deng, Y.; Tang, F.; Dong, W.; Ma, C.; Pan, X.; Wang, L.; Xu, C. StyTr2: Image Style Transfer With Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 11326–11336. [Google Scholar] [CrossRef]
- Liu, S.; Lin, T.; He, D.; Li, F.; Wang, M.; Li, X.; Sun, Z.; Li, Q.; Ding, E. AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Virtual, 10–17 October 2021; pp. 6649–6658. [Google Scholar] [CrossRef]
- Hong, K.; Jeon, S.; Lee, J.; Ahn, N.; Kim, K.; Lee, P.; Kim, D.; Uh, Y.; Byun, H. AesPA-Net: Aesthetic Pattern-Aware Style Transfer Networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 22758–22767. [Google Scholar] [CrossRef]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 586–595. [Google Scholar] [CrossRef]
- Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; Hochreiter, S. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30, pp. 6626–6637. [Google Scholar] [CrossRef]
- Ioannou, E.; Maddock, S. Evaluation in Neural Style Transfer: A Review. Comput. Graph. Forum 2024, 43, e15165. [Google Scholar] [CrossRef]









| Aspect | Benchmark Composition |
|---|---|
| Case definition | Each case contains a content image, a paired style reference, handle–target point annotations, and a binary editing mask. |
| Point profile | Face-centered landmark controls: full-face profiles account for about 94% of the cases, while eyes–nose–mouth profiles account for about 6%. |
| Drag scale | Small (≤20 px): about 28%; medium (20–40 px): about 50%; large (>40 px): about 22%, measured by the largest handle–target displacement in each case. |
| Edited region | The controls are centered on faces. Most cases involve eye or eyebrow landmarks together with nearby facial structure, while the remaining cases emphasize facial contour, nose, or mouth adjustment. |
| Reference rendering | The paired style references include cel shaded anime, retro anime illustration, monochrome manga or line dominant comic style, painterly or oil painting like illustration, semi painterly rendering, and high detail decorative illustration examples. |
| Character design and proportion | The cases include examples with large eye anime proportions, more near human facial proportions, chibi or cartoon like proportions, and distinctive design elements such as complex hair, accessories, glasses, hats, or decorative facial details. |
| Framing and background | The cases mainly cover closeup and bust anime portraits with plain, moderate, or decorative backgrounds. |
| Module | Hyperparameter | Value |
|---|---|---|
| Main trajectory | DDIM steps | 50 |
| Drag optimization | 9 | |
| Drag optimization | Inner iterations | 80 |
| Drag optimization | Patch radius r/search radius | 1/3 |
| Drag loss | Outside-mask weight | 0.1 |
| Drag weighting | Adaptive point weights | on |
| Drag weighting | // | 0.5/0.7/1.5 |
| Anchor handoff | 40 | |
| Anchor initialization | AdaIN strength | 0.75 |
| Style replay | Replay DDIM steps | 20 |
| Style replay | Attention layers | 7–11 |
| Query mixing | / | 0.75/1.5 |
| Style injection | 1.0 | |
| Correspondence replay | Reference layer /top-k | 6/16 |
| Method | MD ↓ | Succ.@20 ↑ | Mask Out IF ↑ | CFSD ↓ |
|---|---|---|---|---|
| DragDiffusion | ||||
| FreeDrag | ||||
| DragNoise | ||||
| GoodDrag | ||||
| DragonDiffusion | ||||
| Ours (drag only) |
| Method | Mask Out IF ↑ | ArtFID ↓ | CFSD ↓ |
|---|---|---|---|
| StyleID | |||
| AdaIN | |||
| STROTSS | |||
| StyA2K | |||
| StyTr2 | |||
| AdaAttN | |||
| AesPA-Net | |||
| Ours (style only) |
| Method | MD ↓ | Succ.@20 ↑ | Mask Out IF ↑ | ArtFID ↓ | CFSD ↓ |
|---|---|---|---|---|---|
| DragDiffusion + StyleID | |||||
| FreeDrag + StyleID | |||||
| DragNoise + StyleID | |||||
| GoodDrag + StyleID | |||||
| DragonDiffusion + StyleID | |||||
| Style during drag | |||||
| Ours |
| Pipeline | Resolution | Runtime/Case ↓ | Peak VRAM ↓ |
|---|---|---|---|
| DragDiffusion + StyleID | 213.6 s | 44.7 GB | |
| FreeDrag + StyleID | 226.4 s | 19.8 GB | |
| DragonDiffusion + StyleID | 116.6 s | 19.8 GB | |
| GoodDrag + StyleID | 205.8 s | 20.5 GB | |
| DragNoise + StyleID | 417.4 s | 19.8 GB | |
| Ours | 72.8 s | 17.8 GB |
| Variant | MD ↓ | Succ.@20 ↑ | Mask out IF ↑ | ArtFID ↓ | CFSD ↓ |
|---|---|---|---|---|---|
| Ours | |||||
| Q gain strong | |||||
| All strong |
| Variant | MD ↓ | Succ.@20 ↑ | Mask Out IF ↑ | ArtFID ↓ | CFSD ↓ |
|---|---|---|---|---|---|
| Ours | |||||
| Immediate handoff | |||||
| Earlier handoff, step 30 | |||||
| Style during drag |
| Variant | MD ↓ | Succ.@20 ↑ | Mask Out IF ↑ | ArtFID ↓ | CFSD ↓ |
|---|---|---|---|---|---|
| Ours | |||||
| w/o semantic | |||||
| w/o explicit sim. bias | |||||
| w/o gain |
| Question | Comparison | Ours Votes | Preference | 95% CI | |
|---|---|---|---|---|---|
| Structure correctness | Ours vs. DragDiffusion + Ours (style only) | 366/432 | 84.7% | [81.0, 87.8] | 0.61 |
| Style faithfulness | Ours vs. Ours (drag only) + StyleID | 287/432 | 66.4% | [61.9, 70.7] | 0.53 |
| Overall quality | Ours selected in four way comparison | 240/432 | 55.6% | [50.8, 60.2] | 0.45 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ding, Y.; Yu, W.; Geng, Y.; Cai, F. Coordinating Drag-Based Structure Editing and Reference Style Transfer in Diffusion Models for Anime Images. Appl. Sci. 2026, 16, 6703. https://doi.org/10.3390/app16136703
Ding Y, Yu W, Geng Y, Cai F. Coordinating Drag-Based Structure Editing and Reference Style Transfer in Diffusion Models for Anime Images. Applied Sciences. 2026; 16(13):6703. https://doi.org/10.3390/app16136703
Chicago/Turabian StyleDing, Youdong, Wenjing Yu, Yafan Geng, and Feifan Cai. 2026. "Coordinating Drag-Based Structure Editing and Reference Style Transfer in Diffusion Models for Anime Images" Applied Sciences 16, no. 13: 6703. https://doi.org/10.3390/app16136703
APA StyleDing, Y., Yu, W., Geng, Y., & Cai, F. (2026). Coordinating Drag-Based Structure Editing and Reference Style Transfer in Diffusion Models for Anime Images. Applied Sciences, 16(13), 6703. https://doi.org/10.3390/app16136703



