Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (65)

Search Parameters:
Keywords = semantic image synthesis

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
33 pages, 10867 KB  
Article
Object-Centric 2D-to-3D Pipeline for Interior-Design Visualization: Reference-Free Asset Evaluation and a Structured3D Scene-Level Benchmark
by Dan Toderici, Tiberiu-Gabriel Rodanciuc, George-Alexandru Micu, Răzvan Rughiniș, Sergiu-Rareș Lupșa and Dinu Țurcanu
Electronics 2026, 15(15), 3295; https://doi.org/10.3390/electronics15153295 (registering DOI) - 26 Jul 2026
Abstract
This study presents a modular AI-assisted workflow for converting single 2D interior images into textured 3D assets and for evaluating those assets when ground-truth 3D meshes are unavailable. The proposed pipeline combines object detection, instance isolation, monocular-depth estimation, image-to-3D generation, texture synthesis, mesh [...] Read more.
This study presents a modular AI-assisted workflow for converting single 2D interior images into textured 3D assets and for evaluating those assets when ground-truth 3D meshes are unavailable. The proposed pipeline combines object detection, instance isolation, monocular-depth estimation, image-to-3D generation, texture synthesis, mesh export, and cloud-based execution to support early-stage interior-design and real-estate visualization tasks. A reference-free validation protocol is introduced, based on rendered multi-view comparisons, silhouette Intersection-over-Union, automated captioning, and multimodal embedding similarity, and is complemented by a composite validation framework that benchmarks reconstructed scenes against 200 panoramic indoor scenes from the Structured3D dataset using Hungarian-matched placement, size, recall, and relative-distance metrics. The workflow was implemented and tested using contemporary computer-vision and generative 3D components, with Hunyuan3D 2.0 used as the main reconstruction model. Proof-of-concept experiments on a representative corpus of 178 synthetically generated single-object images spanning a range of interior furniture categories show comparable silhouette IoU for textured and non-textured outputs and indicate that texture-preserving renderings improve visual and semantic similarity scores across CLIP-based evaluations. The 200-scene dataset evaluation reveals stable spatial localization (placement error ≈ 1.18 m, relative-distance error ≈ 0.54 m) alongside systematic over-prediction and size-calibration errors. Beyond the applied pipeline, the study contributes a reference-free, ground-truth-free protocol for 3D-asset evaluation and a first quantified account of where object-centric single-image reconstruction is reliable—spatial placement—and where it is not—object scale and spurious detection—at interior-scene scale. The results demonstrate the feasibility of integrating perception, 3D reconstruction, semantic assessment, and scalable deployment into a single applied pipeline, while remaining proof-of-concept and requiring extension to larger object and scene corpora, baselines, real-photograph evaluation, and human-centered assessment before broad claims about general interior-scene reconstruction can be made. Full article
(This article belongs to the Special Issue Advances in 3D Computer Vision and 3D Data Processing)
Show Figures

Figure 1

38 pages, 3059 KB  
Review
Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions
by Duc Tri Phan and Hong Duc Nguyen
J. Imaging 2026, 12(7), 324; https://doi.org/10.3390/jimaging12070324 - 17 Jul 2026
Viewed by 155
Abstract
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion [...] Read more.
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion blur, and limited field-of-view or traditional fixed-camera multi-view setups (assuming static geometry and controlled environments), egocentric multi-view systems leverage body-worn rigs to enable a more robust and flexible 3D understanding in open-world, mobile scenarios. In this work, we present a systematic survey of advancements in cross-view feature fusion, geometric consistency enforcement, open-world detection, human–object interaction (HOI) modeling, action segmentation, 3D reconstruction, and novel-view synthesis specifically tailored to wearable multi-camera platforms. Key datasets released between 2024 and 2026—including HOT3D (833 min of synchronized multi-view hand/object interactions from Project Aria and Quest 3), MultiEgo (first multi-egocentric dataset for 4D social scene reconstruction), and Ego-1K (large-scale 12-camera rig for dynamic 3D video synthesis) are thoroughly examined alongside an analysis of integrations with large language models (LLMs) and vision–language models that drive performance gains, typically in the 15–30% range over single-view baselines in hand tracking, HOI recognition, and reconstruction fidelity, although we show through a consolidated meta-analysis that this gain is task-dependent: larger for geometry-bottlenecked tasks such as in-hand object lifting, and smaller, method-dependent, or occasionally negative for semantic-recognition tasks such as keystep recognition under naive view fusion. These methods cover work in multi-view stereo, cross-view learning, and novel-view synthesis while addressing several real-time wearable constraints. Practical applications such as immersive Augmented Reality/Virtual Reality (AR/VR), assistive robotics, and healthcare monitoring are also discussed together with the challenges in motion calibration, benchmark diversity, and edge deployment ability. Thus, in this review, we attempt to fill a critical gap by focusing exclusively on wearable multi-view systems in an open-world setting, synthesizing the latest literature to chart future directions toward more embodied and continual learning agents. Full article
(This article belongs to the Special Issue Techniques in Multi-View Image Analysis)
Show Figures

Figure 1

36 pages, 3134 KB  
Review
AI-Assisted Selective Harvesting and Smart Forest Management: A State-of-the-Art Review of Multimodal Sensing and Decision-Support Approaches
by Janis Peksa
Forests 2026, 17(7), 840; https://doi.org/10.3390/f17070840 - 16 Jul 2026
Viewed by 214
Abstract
Selective harvesting requires tree-level decisions that balance operational productivity, stand development, and sustainable forest management, yet current AI and sensing studies often remain fragmented across remote sensing, machine perception, optimization, and forestry domains. This review synthesizes research on AI-assisted selective harvesting and smart [...] Read more.
Selective harvesting requires tree-level decisions that balance operational productivity, stand development, and sustainable forest management, yet current AI and sensing studies often remain fragmented across remote sensing, machine perception, optimization, and forestry domains. This review synthesizes research on AI-assisted selective harvesting and smart forest management, with emphasis on multimodal sensing, decision support, and harvester-oriented deployment. A PRISMA-informed scoping review was conducted using Scopus, Web of Science Core Collection, IEEE Xplore, ScienceDirect, and SpringerLink, resulting in 82 studies retained for qualitative synthesis. The reviewed literature was organized into eight thematic groups covering smart forestry, harvesting operations, RGB-based perception, LiDAR and point-cloud processing, multimodal fusion, edge deployment, decision-support systems, and forecasting-oriented digital forestry. The analysis shows that RGB imaging provides semantic tree recognition, LiDAR enables spatial localization and structural assessment, and decision-support methods can translate tree-level observations into transparent cut/keep recommendations. However, integrated harvester-mounted systems remain underdeveloped, particularly regarding real-time RGB–LiDAR fusion, operator-facing recommendations, and forecasting integration. This review proposes a reference architecture for human-in-the-loop AI-assisted selective harvesting and identifies future research priorities for field validation and smart forest-management integration. Full article
Show Figures

Figure 1

69 pages, 733 KB  
Systematic Review
AI-Driven Image Generation: Algorithms, Architectures, Quality Assessment, and Applications—A Structured Narrative Review
by Mikołaj Leszczuk, Yi Zhang, Mylène C. Q. Farias, Damon M. Chandler and Ruth Kalola
Electronics 2026, 15(13), 2891; https://doi.org/10.3390/electronics15132891 - 1 Jul 2026
Viewed by 915
Abstract
Background: This paper presents a structured narrative review of recent advances in AI-driven image generation across four complementary perspectives: generative models and architectures, quality assessment and performance metrics, application domains, and multi-modal or cross-lingual extensions. The review aimed to identify dominant methodological trends, [...] Read more.
Background: This paper presents a structured narrative review of recent advances in AI-driven image generation across four complementary perspectives: generative models and architectures, quality assessment and performance metrics, application domains, and multi-modal or cross-lingual extensions. The review aimed to identify dominant methodological trends, representative evaluation practices, and open research challenges in contemporary image generation research. Methods: A structured literature search was conducted in the Scopus database on 29 January 2026 using a predefined query focused on modern generative-image paradigms and excluding clearly out-of-scope domains. Eligible records addressed contemporary AI-driven image generation or closely related multi-modal generation settings within the temporal and topical scope of the search. Retrieved records were first assigned to four thematic branches and then screened with branch-specific relevance criteria for narrative synthesis. Results: The search returned 1524 records, and the final narrative synthesis included 117 publications: 34 on generative models, 23 on quality assessment, 29 on applications, and 31 on multi-modal and cross-lingual aspects. Across the reviewed literature, progress was shaped not only by visual fidelity, but also by controllability, semantic grounding, human-centred evaluation, multi-modal integration, and practical deployment constraints. Limitations: The review was limited to a single primary bibliographic source and to a qualitative narrative synthesis without meta-analysis. Conclusions: The review provides a structured reference point for researchers and practitioners working on AI-based image generation and its evaluation, while also highlighting benchmark, comparability, and multi-modal-transfer challenges that remain unresolved. Full article
Show Figures

Figure 1

22 pages, 28334 KB  
Article
Prompt-Guided Semantic Latent Direction Learning in Diffusion Models for Abstract Visual Concept Manipulation
by Mahzaib Khalid, Fangli Ying, Al-Garadi Ahmed Mohammed Atef, Aniwat Phaphuangwittayakul and Riyad Dhuny
J. Imaging 2026, 12(7), 279; https://doi.org/10.3390/jimaging12070279 - 25 Jun 2026
Viewed by 348
Abstract
Diffusion-based generative models achieve high-fidelity image synthesis; however, controlling internal representations for abstract visual concepts remains challenging due to the ambiguity of textual descriptions. In this work, we propose a prompt-guided concept-vector learning framework for the controllable manipulation of such concepts without requiring [...] Read more.
Diffusion-based generative models achieve high-fidelity image synthesis; however, controlling internal representations for abstract visual concepts remains challenging due to the ambiguity of textual descriptions. In this work, we propose a prompt-guided concept-vector learning framework for the controllable manipulation of such concepts without requiring external human-annotated image pairs, segmentation masks, identity labels, or manually annotated editing targets. The method introduces a learnable concept vector optimized in the bottleneck (mid-block) feature space of a pretrained Stable Diffusion U-Net, while keeping all pretrained model parameters frozen. A multi-prompt data generation strategy based on paired positive and neutral prompts provides weak semantic guidance for capturing the target concept direction and reducing dependence on a single prompt formulation. The learned vector is further applied in an image-to-image setting through controlled noise injection and concept-guided denoising, enabling the semantic modification of real images while preserving structural content. The concept strength is controlled by a scaling parameter γ, while the image-to-image noise strength is controlled by β, allowing for a practical balance between semantic modification and structural fidelity. Experiments are conducted on two main abstract concepts, perfect skin and peaceful lake, with additional qualitative analysis on subjective portrait-level concepts. Quantitative evaluation using SSIM, LPIPS, and CLIP similarity demonstrates that the proposed method improves semantic alignment while maintaining structural preservation compared with Stable Diffusion image-to-image baselines. A human preference study further shows that concept-injected outputs are preferred in 76.0% of responses for perfect skin and 85.7% for peaceful lake. Ablation studies further demonstrate the controllability and robustness of the proposed framework. Overall, the method provides a simple and parameter-efficient approach for interpretable concept-level manipulation in diffusion models. Full article
Show Figures

Figure 1

37 pages, 6098 KB  
Review
AI-Augmented Systematic Review of Remote Sensing and Predictive Modelling for Mycotoxin Risk Monitoring in Cereal Crops Across Central and Balkan Europe
by László Radócz, Attila Nagy, Nikolett Szőllősi, Nikolett Éva Kiss, Andrea Szabó, János Tamás, Nxumalo Gift Siphiwe and László Radócz
Remote Sens. 2026, 18(13), 2063; https://doi.org/10.3390/rs18132063 - 23 Jun 2026
Cited by 1 | Viewed by 443
Abstract
Mycotoxin contamination of cereal crops poses escalating food safety risks across the Central and Balkan European (CBE) corridor under climate change, yet no PRISMA 2020-compliant synthesis of remote sensing (RS) and machine learning (ML) evidence for this region exists. We conducted an AI-augmented [...] Read more.
Mycotoxin contamination of cereal crops poses escalating food safety risks across the Central and Balkan European (CBE) corridor under climate change, yet no PRISMA 2020-compliant synthesis of remote sensing (RS) and machine learning (ML) evidence for this region exists. We conducted an AI-augmented systematic review applying a four-stage automated pipeline—PICO domain scoring, SBERT semantic deduplication, and Thompson-sampling reinforcement learning—to 36,038 corpus records (2010–2025), yielding 156 included studies (inter-rater κ = 0.81 (95% CI: 0.74–0.88)). Logistic growth modelling identified a 56-fold corpus expansion with inflection at t0 = 2024.8 (R2 = 0.981). Satellite multispectral imaging dominated the literature (91.7% of studies); random forest and gradient boosting models achieved R2 = 0.74–0.80 for aflatoxin B1 and deoxynivalenol prediction in CBE maize and wheat when integrating vegetation indices, land surface temperature, and precipitation covariates. Deep learning surpassed classical ML in annual study count from 2021, reaching ~60% relative share by 2025, though the performance advantage narrows at field scale relative to laboratory hyperspectral benchmarks (98–99% accuracy). A five-percentage-point CBE–global performance gap is largely consistent with differences in sample size and multi-toxin design scope rather than algorithmic access. The country × mycotoxin gap matrix identifies zero eligible studies for four CBE nations and for T-2/HT-2 toxins across the Balkan states. Climate-driven satellite mycotoxin prediction emerges as the field’s active research frontier. Full article
(This article belongs to the Special Issue Plant Disease Detection and Recognition Using Remotely Sensed Data)
Show Figures

Figure 1

39 pages, 3403 KB  
Systematic Review
Associations Between the Built Environment and Older Adults’ Mental Health: A Systematic Literature Review (2015–2025)
by Chunhong Wu, Yile Chen, Shuyong Liang, Jiaqi Yang, Liang Zheng, Qingnian Deng, Jingwei Liang, Tianjia Wang, Yuhong Ding and Yinqi Wang
Buildings 2026, 16(12), 2398; https://doi.org/10.3390/buildings16122398 - 16 Jun 2026
Viewed by 611
Abstract
As the global population continues to age, mental health issues such as depression, anxiety, stress, loneliness, and social isolation among older adults are receiving increasing attention. The built environment is closely associated with older adults’ daily mobility, environmental perception, social participation, and mental [...] Read more.
As the global population continues to age, mental health issues such as depression, anxiety, stress, loneliness, and social isolation among older adults are receiving increasing attention. The built environment is closely associated with older adults’ daily mobility, environmental perception, social participation, and mental health and well-being, but the evidence remains heterogeneous across spatial contexts, environmental indicators, and study designs. Previous umbrella reviews have summarized broad links between the built environment and healthy aging, but less attention has been paid to recent original empirical studies published after the COVID-19 pandemic, the distinction between objective environmental exposure and subjective environmental perception, and the role of social participation as a pathway linking environmental conditions to mental health and well-being. This study employs a systematic literature review approach, searching and screening peer-reviewed empirical studies published between 2015 and January 2026 that focus on the associations between the built environment and older adults’ mental health and well-being. PubMed, Scopus, and Web of Science databases were used for searching, supplemented by manual searching. After title and abstract screening and full-text evaluation, a total of 60 studies were included. Subsequently, a comprehensive analysis was conducted on aspects such as research design, spatial scale, environmental indicators, types of mental health outcomes, and potential pathways of action. In this review, core mental health and well-being outcomes included negative outcomes, such as depression, anxiety, stress, psychological distress, loneliness, and social isolation, and positive outcomes, such as life satisfaction, subjective well-being, psychological well-being, and mental well-being. Social participation was examined as a behavioral and psychosocial pathway rather than as a core outcome. Emerging methods, including street-view image analysis, FCN-based semantic segmentation, and XGBoost-SHAP, were examined because they can refine environmental exposure measurement and support variable-importance interpretation, rather than because they provide causal evidence. The main synthesis suggests that several built environment factors are associated with older adults’ mental health and well-being, although the strength and consistency of evidence vary across outcome types, spatial contexts, and study designs. (1) Exposure to green and blue spaces, quality of public open spaces, walkability and accessibility, accessibility of neighborhood facilities and services, housing and living conditions, and positive environmental perception are mostly associated with lower levels of depression, anxiety, stress, and loneliness, as well as higher levels of life satisfaction, subjective well-being, and psychological well-being. (2) Conversely, adverse environmental exposures such as proximity to roads, pollution, non-vegetated spaces, and high-intensity urbanization are more likely to exacerbate negative psychological outcomes. Existing evidence also suggests that social participation is one of the important behavioral pathways through which the built environment is linked to the mental health of older adults, but it is not the only mechanism. (3) In addition, the direction and intensity of environmental associations remain heterogeneous under different spatial scales, indicator types, and research methods. Overall, this review contributes by organizing recent empirical evidence into a built environment–social participation–mental health and well-being framework, while emphasizing that most findings should be interpreted primarily as evidence of association rather than as stable or uniform causal effects. Full article
Show Figures

Figure 1

22 pages, 6591 KB  
Article
A Study on the Generation and Evaluation of Illustrations for Chinese Idiom Allusions Based on AIGC
by Jingxue Li, Youping Teng and Weijia Wang
Information 2026, 17(5), 495; https://doi.org/10.3390/info17050495 - 18 May 2026
Viewed by 438
Abstract
As carriers of traditional culture, Chinese idiom allusions contain rich semantic and emotional content. High-quality illustrations of these idioms hold significant potential for applications in cultural communication and education. Although generative artificial intelligence has achieved substantial progress in general image synthesis, it remains [...] Read more.
As carriers of traditional culture, Chinese idiom allusions contain rich semantic and emotional content. High-quality illustrations of these idioms hold significant potential for applications in cultural communication and education. Although generative artificial intelligence has achieved substantial progress in general image synthesis, it remains challenging to produce idiom illustrations in culture-intensive scenarios that simultaneously preserve cultural symbols, maintain affective ontology, and exhibit high visual aesthetic quality. To address this gap, we propose a three-dimensional evaluation framework—Zhen-Shan-Mei (Truth-Goodness-Beauty)—for idiom illustrations. The ‘Truth’ module uses Chinese vision–language models to quantify cultural symbols; the ‘Goodness’ module applies cross-modal affective analysis to assess affective ontology; and the ‘Beauty’ module computes quantitative aesthetic metrics (composition balance, color harmony, and line expressiveness). Based on this system, an AI-idiom prototype system is constructed to realize closed-loop iteration of generation-evaluation-regeneration and threshold screening. Experiments show that the proportion of illustrations selected by subjects after the “Truth-Goodness-Beauty” screening reaches 78.1%. The results suggest that the proposed method is effective in maintaining cultural symbols, strengthening affective ontology, and improving visual aesthetics and offers a potentially interpretable and reproducible evaluation and optimization framework for culture-intensive image generation tasks. Full article
(This article belongs to the Section Information Applications)
Show Figures

Figure 1

50 pages, 6299 KB  
Review
From Pixel Understanding to Semantic Insight: Intelligent Detection in Sensor-Driven Perception Systems
by Qingchen Xie, Tongxu Wu and Fan Yang
Sensors 2026, 26(10), 3075; https://doi.org/10.3390/s26103075 - 13 May 2026
Viewed by 682
Abstract
Intelligent detection in modern manufacturing, healthcare, process industries, and structural monitoring is fundamentally enabled by heterogeneous sensor systems. Rather than being viewed as a purely image-centered recognition task, intelligent detection is more appropriately formulated as a sensor-driven state inference problem in which sensing [...] Read more.
Intelligent detection in modern manufacturing, healthcare, process industries, and structural monitoring is fundamentally enabled by heterogeneous sensor systems. Rather than being viewed as a purely image-centered recognition task, intelligent detection is more appropriately formulated as a sensor-driven state inference problem in which sensing physics, signal quality, temporal synchronization, modality availability, and deployment conditions jointly determine what can be reliably detected, localized, interpreted, and acted upon. Against this background, this review provides a structured synthesis of the field through three coupled dimensions, namely methods, systems, and governance, and organizes the literature around four recurring engineering components: signal unification, representation unification, alignment mechanisms, and robustness mechanisms. Using a structured review protocol with explicit source selection, screening, and study coding, the paper traces the methodological evolution from traditional feature-engineering and model-based pipelines to deep learning for visual, temporal, multimodal, generative, and mechanism-constrained sensing, and further to foundation-model-based and multimodal sensor intelligence. Cross-domain evidence is synthesized from industrial defect detection, fault diagnosis, remaining useful life prediction, non-destructive testing, structural health monitoring, medical lesion analysis, and process monitoring. The review argues that recent progress has substantially strengthened learned representations, multimodal interaction, and semantic extensibility, but has not removed persistent constraints arising from domain shift, missing modalities, calibration instability, privacy-preserving collaboration, and edge-side resource limits. Accordingly, the central challenge is no longer how to optimize isolated detection models, but how to build sensor-enabled intelligent systems that remain physically grounded, trustworthy, transferable, and maintainable under real operational conditions. On this basis, the paper concludes by identifying future directions in mechanism-aware modeling, trustworthy evaluation, missing-modality-robust multimodal systems, privacy-preserving cross-site collaboration, and edge-native lifecycle-aware deployment. Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

20 pages, 10156 KB  
Article
Unveiling the Risk of Unsafe Image Generation in Stable Diffusion Through a Cross-Attention Mechanism
by Yong Zhuang, Yiheng Jing, Wenzhe Yi, Xiaoyang Xu and Juan Wang
Future Internet 2026, 18(5), 248; https://doi.org/10.3390/fi18050248 - 7 May 2026
Viewed by 1104
Abstract
Text-to-image diffusion models such as Stable Diffusion enable high-quality image synthesis from text and are widely deployed due to their open-source nature and low computational requirements. However, this accessibility also makes them attractive targets for misuse, including the generation of not-safe-for-work and otherwise [...] Read more.
Text-to-image diffusion models such as Stable Diffusion enable high-quality image synthesis from text and are widely deployed due to their open-source nature and low computational requirements. However, this accessibility also makes them attractive targets for misuse, including the generation of not-safe-for-work and otherwise restricted content. In this paper, we propose EvilPrompt, a jailbreak attack that exploits the cross-attention mechanism in Stable Diffusion. The attack operates purely at inference time using plain-text prompts and does not require fine-tuning or modification of model parameters. By selectively reweighting cross-attention for specific tokens, EvilPrompt preserves the overall semantic structure of the prompt while steering the generation toward prohibited content. This enables fine-grained control over malicious semantics without introducing explicit unsafe keywords. We evaluate EvilPrompt on two real-world prompt sets, 4chan and Lexica, each containing 500 prompts. The attack achieves an Attack Success Rate (ASR) of 97.4% on 4chan and 98.0% on Lexica, yielding an overall average ASR of 97.7%. The attack maintains high semantic alignment between prompts and generated images. Bootstrapping Language-Image Pre-training (BLIP) similarity consistently exceeds 0.75 across all categories on both datasets. Human evaluation further confirms high visual realism, with mean scores above 7.0 on a 10-point scale, and strong semantic consistency, with mean scores above 7.3. These results demonstrate that cross-attention manipulation provides an effective and practical jailbreak pathway. We further analyze how commonly used text-level moderation affects the success of such attacks. Although the strongest defense configuration (HateCoT with GPT-4) reduces the ASR to 5.9%, it introduces 21.5 s of additional latency and a cost of $0.01182 per query. Lighter-weight alternatives such as Perspective API leave nearly half (45.0%) of attacks successful. These observations indicate that safeguards acting only on the input or final output are insufficient to capture attention-level manipulations. Overall, our results reveal a fundamental limitation of post-generation safety pipelines when confronted with inference-time control of cross-attention. Full article
Show Figures

Figure 1

36 pages, 16246 KB  
Article
A Compliance-Driven Generative Framework for Zhejiang-Style Rural Facades
by Chengzong Wu, Liping He, Shishu Tong, Jun Zhao and Yun Wu
Buildings 2026, 16(8), 1544; https://doi.org/10.3390/buildings16081544 - 14 Apr 2026
Viewed by 547
Abstract
Under the background of the Rural Revitalization Strategy, Zhejiang Province is promoting “Zhejiang-style Vernacular Dwellings” as a crucial measure to enhance the rural living environment and architectural appearance. However, traditional stylistic control tools, such as standardized rural housing design atlases, exhibit limitations including [...] Read more.
Under the background of the Rural Revitalization Strategy, Zhejiang Province is promoting “Zhejiang-style Vernacular Dwellings” as a crucial measure to enhance the rural living environment and architectural appearance. However, traditional stylistic control tools, such as standardized rural housing design atlases, exhibit limitations including weak responsiveness to villagers’ individualized needs and high professional thresholds. Consequently, they struggle to address the bottlenecks in grassroots governance efficiency caused by massive and personalized housing demands. Meanwhile, when applied to architectural design, general generative AI technologies often suffer from “structural hallucinations” and the weakening of regional characteristics due to a lack of physical tectonic constraints. Oriented towards the governance requirements of the Zhejiang Provincial Rural Housing Design Guidelines, this study proposes a compliance evaluation-driven “Contour-Semantic-Image” hierarchical generative control framework. This aims to construct a visual scheme generation and pre-screening workflow that deeply adapts to the logic of rural governance. At the data level, this research aggregates multi-source materials, including official standardized atlases, government stylistic guidelines, and real-world photographs. Through expert screening and standardized processing of 596 schemes, a dataset of 333 high-quality, finely annotated structured samples is constructed. Furthermore, a human-guided, machine-segmented workflow assisted by Segment Anything Model 2 (SAM 2) is employed to establish a semantic label system comprising 4 major categories and 13 subcategories of components, thereby achieving the structural deconstruction of architectural prior knowledge. At the generation level, a two-stage model is trained based on Stable Diffusion and ControlNet: Stage I utilizes contour conditions and “layout prompts” to generate semantic label maps, aiming to strengthen component topology and layout consistency; Stage II employs the semantic label maps and “style prompts” as conditions to generate photorealistic facade images. By utilizing explicit semantic constraints to guide the model from pixel synthesis to logical generation, it achieves the controllable rendering of stylistic details and material expressions. At the evaluation level, an automated verification system featuring “clause translation–metric calculation–comprehensive scoring” is proposed. It conducts scoring, re-ranking, and diagnostic feedback on the generated variants across three dimensions: Design Rationality (Q), General Compliance (G), and Jiangnan water-town Regional Characteristics (P-J), forming a closed-loop “Generation-Evaluation-Feedback” workflow. Overall, this framework provides a “visualizable, evaluable, and explainable” pathway for scheme generation and pre-screening in the digital governance of rural architectural appearance. Full article
(This article belongs to the Special Issue Data-Driven Intelligence for Sustainable Urban Renewal)
Show Figures

Figure 1

22 pages, 1747 KB  
Review
Talking Head Generation Through Generative Models and Cross-Modal Synthesis Techniques
by Hira Nisar, Salman Masood, Zaki Malik and Adnan Abid
J. Imaging 2026, 12(3), 119; https://doi.org/10.3390/jimaging12030119 - 10 Mar 2026
Viewed by 1593
Abstract
Talking Head Generation (THG) is a rapidly advancing field at the intersection of computer vision, deep learning, and speech synthesis, enabling the creation of animated human-like heads that can produce speech and express emotions with high visual realism. The core objective of THG [...] Read more.
Talking Head Generation (THG) is a rapidly advancing field at the intersection of computer vision, deep learning, and speech synthesis, enabling the creation of animated human-like heads that can produce speech and express emotions with high visual realism. The core objective of THG systems is to synthesize coherent and natural audio–visual outputs by modeling the intricate relationship between speech signals, facial dynamics, and emotional cues. These systems find widespread applications in virtual assistants, interactive avatars, video dubbing for multilingual content, educational technologies, and immersive virtual and augmented reality environments. Moreover, the development of THG has significant implications for accessibility technologies, cultural preservation, and remote healthcare interfaces. This survey paper presents a comprehensive and systematic overview of the technological landscape of Talking Head Generation. We begin by outlining the foundational methodologies that underpin the synthesis process, including generative adversarial networks (GANs), motion-aware recurrent architectures, and attention-based models. A taxonomy is introduced to organize the diverse approaches based on the nature of input modalities and generation goals. We further examine the contributions of various domains such as computer vision, speech processing, and human–robot interaction, each of which plays a critical role in advancing the capabilities of THG systems. The paper also provides a detailed review of datasets used for training and evaluating THG models, highlighting their coverage, structure, and relevance. In parallel, we analyze widely adopted evaluation metrics, categorized by their focus on image quality, motion accuracy, synchronization, and semantic fidelity. Operating parameters such as latency, frame rate, resolution, and real-time capability are also discussed to assess deployment feasibility. Special emphasis is placed on the integration of generative artificial intelligence (GenAI), which has significantly enhanced the adaptability and realism of talking head systems through more powerful and generalizable learning frameworks. Full article
Show Figures

Figure 1

26 pages, 1041 KB  
Review
Artificial Intelligence in Orthopaedics: Clinical Performance, Limitations, and Translational Readiness—A Review
by Wojciech Michał Glinkowski, Antonina Spalińska, Agnieszka Wołk and Krzysztof Wołk
J. Clin. Med. 2026, 15(5), 1751; https://doi.org/10.3390/jcm15051751 - 25 Feb 2026
Cited by 3 | Viewed by 2316
Abstract
Background/Objectives: Musculoskeletal disorders and their surgical treatment significantly affect global disability, healthcare utilization, and costs. Artificial intelligence (AI) is a key enabler of data-driven musculoskeletal care. Their applications include diagnostic imaging, surgical planning, risk prediction, rehabilitation, and digital health ecosystems. This narrative review [...] Read more.
Background/Objectives: Musculoskeletal disorders and their surgical treatment significantly affect global disability, healthcare utilization, and costs. Artificial intelligence (AI) is a key enabler of data-driven musculoskeletal care. Their applications include diagnostic imaging, surgical planning, risk prediction, rehabilitation, and digital health ecosystems. This narrative review synthesizes current evidence on the use of AI in orthopaedics and musculoskeletal care across five areas: diagnostic imaging, surgical planning and intraoperative augmentation, predictive analytics and patient-reported outcomes, rehabilitation intelligence and teleorthopaedics, and system-level management. An additional task is to identify translational gaps and priorities for safe, ethical, and equitable implementation of AI. Methods: A structured narrative review was conducted using targeted searches in PubMed, Scopus, and Web of Science supplemented by semantic and citation-based explorations in Semantic Scholar, OpenAlex, and Google Scholar. The main search period was January 2019 to December 2025. The retrieved peer-reviewed articles were analyzed for clinical relevance to human musculoskeletal care, quantitative outcomes, and the translational implications of the results. From the broader pool of eligible publications, 40 clinically relevant studies were selected for detailed synthesis covering imaging, surgical planning, predictive modeling, rehabilitation, and system-level applications. Owing to the significant heterogeneity in the model architectures, datasets, and endpoints, the results were organized into five predefined thematic areas. Results: The most mature evidence is for AI-assisted detection of bone fractures on radiographs, identification of implants, and use of sizing templates in preoperative planning for arthroplasty, where deep learning systems have achieved expert-level diagnostic performance (e.g., fracture detection sensitivity of approximately 90% and specificity of approximately 92% and implant identification accuracy of 97–99%) and improved the accuracy of preoperative planning compared to conventional templating. AI-based planning increases the likelihood of reducing intraoperative corrections, shortening surgery time, reducing blood loss, and improving the final functional outcomes. Predictive models can support the stratification of risk for complications, rehospitalizations, and patient-reported outcomes, although external validation remains limited and is often single-center at this stage of research. Emerging applications in rehabilitation and teleorthopaedics, including sensor-based monitoring and learning systems integrated with Patient-Reported Outcome Measures (PROMs), are conceptually promising, but are mainly limited to feasibility or pilot studies. Conclusions: AI is beginning to influence musculoskeletal care, moving beyond pattern recognition toward integrated, patient-centered decision support throughout the perioperative and rehabilitation periods. Its widespread use remains constrained by limited multicenter validation, dataset bias, algorithmic opacity, and immature regulatory and governance frameworks. Future work should prioritize prospective multicenter impact studies, repeatable revalidation of local models, integration of PROM and teleorthopedic data with health learning systems, and adaptation to changing regulatory requirements to enable safe, ethical, effective, and equitable implementation in routine orthopedic practice. Full article
(This article belongs to the Topic Machine Learning and Deep Learning in Medical Imaging)
Show Figures

Figure 1

18 pages, 4500 KB  
Article
Localizing Perceptual Artifacts in Synthetic Images for Image Quality Assessment via Deep-Learning-Based Anomaly Detection
by Zijin Yin
Electronics 2026, 15(5), 916; https://doi.org/10.3390/electronics15050916 - 24 Feb 2026
Viewed by 782
Abstract
While deep generative models, such as text-to-image diffusion, demonstrate strong capabilities in synthesizing photorealistic images, they frequently produce perceptual artifacts (e.g., distorted structures or unnatural textures) that require manual correction. Existing artifact localization methods typically rely on fully supervised training with large-scale pixel-level [...] Read more.
While deep generative models, such as text-to-image diffusion, demonstrate strong capabilities in synthesizing photorealistic images, they frequently produce perceptual artifacts (e.g., distorted structures or unnatural textures) that require manual correction. Existing artifact localization methods typically rely on fully supervised training with large-scale pixel-level annotations, which suffer from high labeling costs. To address these challenges, we propose a novel framework based on the core insight that perceptual artifacts can be fundamentally modeled as “semantic outliers”—regions that inherently fail to match any pre-defined semantic categories. Instead of learning specific artifact features, we introduce a Mask-based Semantic Rejection (MSR) mechanism within a semantic segmentation architecture. This mechanism leverages the “one-vs-all” property of object queries to identify regions that are consistently rejected by all pre-trained semantic categories. Furthermore, we design a flexible adaptation strategy that supports both zero-shot inference using pre-trained semantic knowledge and fine-tuning with a margin-based suppression objective to explicitly optimize the rejection boundary using minimal supervision. Comprehensive experiments across 11 synthesis tasks demonstrate that MSR significantly outperforms state-of-the-art methods, particularly in data-efficient scenarios. Specifically, the framework achieves mIoU improvements of 6.52% and 13.06% on the text-to-image task using only 10% and 50% of labeled samples, respectively, underscoring its superior capability. Full article
(This article belongs to the Special Issue Computer Vision and AI Algorithms for Diverse Scenarios)
Show Figures

Figure 1

24 pages, 8810 KB  
Article
FreqPose: Frequency-Aware Diffusion with Fractional Gabor Filters and Global Pose–Semantic Alignment
by Meng Wang, Bing Wang, Huiling Chen, Jing Ren and Xueping Tang
Sensors 2026, 26(4), 1334; https://doi.org/10.3390/s26041334 - 19 Feb 2026
Viewed by 580
Abstract
The task of pose-guided person image generation has long been confronted with two major challenges: high-frequency texture details tend to blur and be lost during appearance transfer, while the semantic identity of the person is difficult to maintain consistently during pose changes. To [...] Read more.
The task of pose-guided person image generation has long been confronted with two major challenges: high-frequency texture details tend to blur and be lost during appearance transfer, while the semantic identity of the person is difficult to maintain consistently during pose changes. To address these issues, this paper proposes a diffusion-based generative framework that integrates frequency awareness and global semantic alignment. The framework consists of two core modules: a multi-level fractional-order Gabor frequency-aware network, which accurately extracts and reconstructs high-frequency texture features such as hair strands and fabric wrinkles, enhances image detail fidelity through fractional-order filtering and complex domain modeling; and a global semantic-pose alignment module that utilizes a cross-modal attention mechanism to establish a global mapping between pose features and appearance semantics, ensuring pose-driven semantic alignment and appearance consistency. The collaborative function of these two modules ensures that the generated results maintain structural integrity and natural textures even under complex pose variations and large-angle rotations. The experimental results on the DeepFashion and Market1501 datasets demonstrate that the proposed method outperforms existing state-of-the-art approaches in terms of SSIM, FID, and perceptual quality, validating the effectiveness of the model in enhancing texture fidelity and semantic consistency. Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

Back to TopTop