Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

Search Results (368)

Search Parameters:
Keywords = color semantics

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 17265 KB  
Article
How Visual Elements Shape Perceived Spatial Quality in Urban Waterfront Space: An Explainable Machine Learning Approach for Urban Landscape Planning
by Wenhan Li, Yinzhe Li, Gaoming Liang, Congxi Liu, Dezheng Kong and Yan Feng
Sustainability 2026, 18(16), 8610; https://doi.org/10.3390/su18168610 (registering DOI) - 21 Aug 2026
Viewed by 109
Abstract
As China’s urbanization shifts toward quality-oriented development, urban regeneration increasingly prioritizes the perceived quality of public spaces to enhance urban vitality and advance sustainable urban living. This study takes Zhengzhou’s Dongfeng Canal, a revitalized urban core waterfront, as a case to develop a [...] Read more.
As China’s urbanization shifts toward quality-oriented development, urban regeneration increasingly prioritizes the perceived quality of public spaces to enhance urban vitality and advance sustainable urban living. This study takes Zhengzhou’s Dongfeng Canal, a revitalized urban core waterfront, as a case to develop a human–machine collaborative analytical framework for exploring nonlinear relationships between visual environmental features and human spatial quality perception. By integrating 779 geolocated panoramic images with volunteers’ subjective rating data, this study adopts deep learning-based semantic segmentation to quantify eight objective visual indicators (e.g., greenness, color diversity, spatial structure). A random forest (RF) model links these indicators to three perceptual dimensions: scenic beauty, safety, and recreational value. Adopting explainable artificial intelligence (SHAP and PDPs), the results indicate that: (1) greenness is positively associated with positive perceptions but exhibits a significant threshold effect; (2) color diversity and waterfront accessibility substantially improve user experience, while excessive uniformity and extreme openness negatively affect perceived spatial quality. These findings challenge the simplistic linear “more-is-better” assumption in urban design and highlight the value of balanced, context-sensitive spatial interventions. This study provides evidence-based, segment-specific strategies for urban waterfront regeneration, advancing people-centered planning that integrates ecological functionality, social inclusivity, and long-term sustainability via Geospatial Artificial Intelligence (GeoAI) and geospatial analytics. Full article
Show Figures

Figure 1

18 pages, 3993 KB  
Article
Rail Light-Strip Abnormality Analysis from Color Inspection Images Using an Improved SegFormer and Geometric Rules
by Haoran Song, Yuntao Gou, Ning Wang, Le Wang, Junbo Liu, Shengchun Wang, Chengliang Xia, Qiang Han and Zichen Gu
Sensors 2026, 26(16), 5292; https://doi.org/10.3390/s26165292 - 21 Aug 2026
Viewed by 135
Abstract
Rail light-strip morphology reflects the wheel-rail contact condition. Reliable automatic analysis remains difficult. The strip is narrow and has weak boundaries, while specular reflection, rail-head texture and trackside background interfere with color inspection images. This study proposes a segmentation-guided geometric method for rail [...] Read more.
Rail light-strip morphology reflects the wheel-rail contact condition. Reliable automatic analysis remains difficult. The strip is narrow and has weak boundaries, while specular reflection, rail-head texture and trackside background interfere with color inspection images. This study proposes a segmentation-guided geometric method for rail light-strip abnormality analysis. An improved SegFormer jointly segments the background, rail-head and light-strip regions. A boundary detail enhancement module refines weak rail-head and light-strip contours. Focal Loss emphasizes minority and hard boundary pixels. The rail-head mask provides the geometric reference for extracting the light-strip centerline, eccentricity, width sequence and connected-component morphology. The predicted masks are ordered using the corrected mileage record. Every 1000 original-resolution rows then form a consecutive 1 m detection unit. When a geometric rule is triggered, the method reports that unit’s 1 m mileage interval together with its eccentricity, width-change or local-integrity measurement. The model achieves 95.67% mean Intersection over Union (mIoU) on 3520 annotated images. It detects 845 of 876 positive units, with 96.46% recall, 89.23% precision and 92.70% F1-score. The resulting records identify abnormal 1 m mileage intervals and report the corresponding eccentricity, width-change, or local-integrity measurements for targeted manual review. Full article
Show Figures

Figure 1

45 pages, 6203 KB  
Article
Generative AI-Assisted Visualization Prototyping for Cultural Heritage: A Computational Framework from 2D Planes to 3D Immersive Scenes
by Jianquan Liu, Runnan Li and Haiying Zhao
Buildings 2026, 16(16), 3319; https://doi.org/10.3390/buildings16163319 - 20 Aug 2026
Viewed by 232
Abstract
Immersive visualization can support interpretation of architectural heritage in historical paintings, yet translating 2D pictorial evidence into navigable 3D scenes remains challenging. Conventional workflows rely on physical survey data, while direct generative AI (GenAI) may produce structural hallucinations and lack historical constraints. This [...] Read more.
Immersive visualization can support interpretation of architectural heritage in historical paintings, yet translating 2D pictorial evidence into navigable 3D scenes remains challenging. Conventional workflows rely on physical survey data, while direct generative AI (GenAI) may produce structural hallucinations and lack historical constraints. This study proposes a human-in-the-loop GenAI-assisted framework for producing immersive 3D visualization prototypes rather than historically verified reconstructions. It integrates multi-view image generation, knowledge-informed review, single-image-to-3D generation, topology inspection, and perceptual calibration. Four fragments from the Northern Song Dynasty painting Along the River During the Qingming Festival were examined as a single-case proof of concept. Across three tested model pairs, raw AI assets were generated in approximately 3–4 min and were suitable for distant-background use; close-up visualization required 1–2 h of refinement, while basic structural editability required 4–5 h of post-processing, reducing the initial time advantage. A mixed-methods study with nine domain experts and 30 non-expert participants used the UES-SF, an adapted VisAWI, and semi-structured interviews analyzed through inductive thematic analysis. All eight subscale scores exceeded their neutral midpoints after Bonferroni correction (all adjusted p<0.001), indicating favorable perceptions of the guided experience. Interviews suggested potential for spatial exploration, museum interpretation, and education. However, geometric discontinuities, detail loss, color deviation, and historical-semantic errors remained, requiring expert review and manual correction. Transferability beyond this artwork and architectural tradition remains untested. Full article
(This article belongs to the Section Construction Management, and Computers & Digitization)
Show Figures

Figure 1

24 pages, 55593 KB  
Article
Color Palette Identification and Intelligent Knowledge Extraction for Natural Disaster Mapping
by Weiyao Guo, An Zhang and Yi Cao
ISPRS Int. J. Geo-Inf. 2026, 15(8), 369; https://doi.org/10.3390/ijgi15080369 - 16 Aug 2026
Viewed by 191
Abstract
Natural disaster emergency cartography requires high semantic accuracy in color design and efficient visual communication. However, existing studies still lack systematic palette analysis and knowledge organization methods based on real-world emergency maps. To address this gap, this study proposes a framework for palette [...] Read more.
Natural disaster emergency cartography requires high semantic accuracy in color design and efficient visual communication. However, existing studies still lack systematic palette analysis and knowledge organization methods based on real-world emergency maps. To address this gap, this study proposes a framework for palette analysis and knowledge organization that uses publicly available Emergency Response Coordination Centre (ERCC)’s emergency maps as the primary data source. The framework extracts disaster types and thematic mapping indicators. It performs palette identification, matching, and statistical analysis using color information from legend regions in the RGB, HSV, and CIELab color spaces, together with the ColorBrewer palette system. Based on the statistical matching results, we constructed a structured knowledge graph that links disaster types, thematic mapping indicators, and palettes, enabling organized retrieval of palette knowledge. Results show that color extraction from legend regions effectively reduces interference from non-thematic elements and improves the accuracy of palette identification. In addition, palette usage in ERCC emergency maps exhibits clear statistical associations and shared and differentiated patterns, indicating stable yet non-unique associations among disaster themes, thematic mapping indicators, and color palettes. The proposed knowledge graph provides a structured framework for organizing palette knowledge and analyzing semantic relationships in ERCC emergency cartography. Full article
(This article belongs to the Special Issue Knowledge-Guided Map Representation and Understanding)
Show Figures

Figure 1

34 pages, 28776 KB  
Article
Global-Local Feature-Based Rice Leaf Disease Classification Using Two-Stream Deep Neural Network Feature Fusion
by Md Nahidur Rahaman, Abdullah Al Mamun, Md. Kamal Hossen, Abdur Rouf, Tumpa Rani Shaha, Jungpil Shin, Mohd Nizam Husen and Abu Saleh Musa Miah
Computers 2026, 15(8), 528; https://doi.org/10.3390/computers15080528 - 14 Aug 2026
Viewed by 248
Abstract
Rice leaf diseases significantly affect crop health and yield potential, creating a need for accurate and timely disease diagnosis. Many existing approaches still rely on single-stream feature extraction architectures, which may limit the ability to simultaneously capture global contextual information and fine-grained disease [...] Read more.
Rice leaf diseases significantly affect crop health and yield potential, creating a need for accurate and timely disease diagnosis. Many existing approaches still rely on single-stream feature extraction architectures, which may limit the ability to simultaneously capture global contextual information and fine-grained disease characteristics. Moreover, limited interpretability and decision-support capability hinder their practical deployment in real-world rice farming. To address these limitations, we employed a framework consists of two parallel feature extraction streams designed to capture different characteristics of disease patterns. The first stream uses a Swin Transformer to learn global contextual information and long-range spatial relationships across the leaf image. The second stream employs ConvNeXt to extract local texture features, including lesion details, spots, and color variations. By combining these complementary representations, the proposed framework effectively integrates global semantic information with local disease-specific features for improved classification performance. The extracted features from the two streams are fused through concatenation followed by an attention-based feature refinement module, enabling adaptive weighting of discriminative features. The refined representation is then used by a fully connected classifier for disease prediction. To enhance model interpretability, Grad-CAM visualization is incorporated to highlight disease-relevant regions and provide visual explanations for the model decisions. Furthermore, an LLM-based advisory module is integrated as a post-diagnosis decision-support component to provide contextualized disease management information and suggestions based on the predicted disease category. The generated suggestions are intended to support, rather than replace, expert agronomic recommendations and should be validated by agricultural professionals before practical application. The proposed framework was evaluated on two rice leaf disease datasets, achieving accuracies of 99.55% and 97.06% on Dataset-1 and Dataset-2, respectively, which are higher than those reported in previous studies. Additionally, five-fold cross-validation on Dataset-1 achieved an average accuracy of 98.91% ± 0.47, demonstrating the stability of the proposed approach. Cross-dataset evaluation using nine common disease classes across both datasets achieved 90.80% accuracy, indicating improved generalization across different data distributions. The proposed framework provides an accurate and explainable approach for rice leaf disease diagnosis in smart agriculture applications. Full article
(This article belongs to the Special Issue Advances in Computer Vision: Models, Learning, and Inference)
Show Figures

Figure 1

29 pages, 37486 KB  
Article
Physics-Aware Diffusion Synthesis for Robust Underwater Object Detection
by Wenxin Xiao, Xiaowei Zhou and Junyu Dong
J. Mar. Sci. Eng. 2026, 14(16), 1503; https://doi.org/10.3390/jmse14161503 - 13 Aug 2026
Viewed by 203
Abstract
Underwater object detection remains challenging in adverse aquatic environments, where severe image degradation caused by turbidity, light scattering, color attenuation, and low illumination substantially reduces detection reliability. Although real-world underwater datasets are essential, their limited scale and environmental diversity make it difficult to [...] Read more.
Underwater object detection remains challenging in adverse aquatic environments, where severe image degradation caused by turbidity, light scattering, color attenuation, and low illumination substantially reduces detection reliability. Although real-world underwater datasets are essential, their limited scale and environmental diversity make it difficult to cover the wide range of degraded conditions encountered in practice. To improve detection robustness without collecting additional annotations, we propose physics-aware diffusion synthesis (PADS), a framework that uses a small set of labeled real images to synthesize diverse physically plausible degraded underwater samples. PADS couples a ControlNet-conditioned latent diffusion generator with a physics-based underwater image-formation model inspired by Jaffe–McGlamery and Akkaynak optics. Semantic masks are first employed to preserve object layout during generation. Meanwhile, water-optics parameters are incorporated through cross-attention to guide the degradation process. In addition, the physical model enforces a color and attenuation consistency loss during training and serves as an SDEdit-style latent prior during synthesis. To further improve localization under degradation, especially for small objects, we introduce a training-only scale-aware focaler–NWD (SA-FNWD) bounding-box loss, which emphasizes normalized Wasserstein distance for small boxes while retaining IoU-based regression for larger objects. Experiments on the MOUD dataset demonstrate that detectors trained with PADS-synthesized data achieve substantially stronger robustness under severe degradation. At the harshest turbidity level, PADS retains 59.3% of clean accuracy compared with 14.9% for the copy–paste-based synthesis method and 14.1% for the pix2pix-based synthesis method. SA-FNWD further improves mAP@0.5:0.95 across degradation severities. These results show that physics-grounded diffusion synthesis provides the main robustness gain, while SA-FNWD offers a complementary small-object localization improvement with no inference overhead. Full article
(This article belongs to the Special Issue Object Detection and Coordinated Control of Marine Robots)
Show Figures

Figure 1

57 pages, 39305 KB  
Review
Hybrid Event–Frame Sensing for Human-Perceptual Imaging and Machine Vision
by Paul K. J. Park, Junseok Kim and Juhyun Ko
Sensors 2026, 26(16), 5127; https://doi.org/10.3390/s26165127 - 13 Aug 2026
Viewed by 403
Abstract
Frame-based RGB image sensors and event-based vision sensors provide complementary sensing capabilities for human-perceptual imaging and machine vision. RGB image sensors capture dense spatial, color, and texture information that is essential for human-viewable imaging, semantic recognition, and conventional image signal processing pipelines. In [...] Read more.
Frame-based RGB image sensors and event-based vision sensors provide complementary sensing capabilities for human-perceptual imaging and machine vision. RGB image sensors capture dense spatial, color, and texture information that is essential for human-viewable imaging, semantic recognition, and conventional image signal processing pipelines. In contrast, dynamic vision sensors (DVSs) and event vision sensors (EVSs) asynchronously detect local brightness changes and provide sparse temporal information with low latency, high temporal resolution, and reduced redundant data output. Because neither modality alone satisfies all requirements of emerging vision systems, hybrid event–frame sensing has become an important direction for compact, low-latency, and energy-efficient sensing. This review presents a sensor-oriented taxonomy of hybrid event–frame sensing architectures and systems, including dual-camera event–frame systems, optically aligned event–frame systems, pixel-level shared hybrid image sensors, stacked CIS–DVS hybrid image sensors, homogeneous-pixel sensing systems, and event-only reconstruction systems. We analyze key sensor specifications, including latency, spatial resolution, color fidelity, power consumption, and form factor, and discuss how these specifications guide sensor configuration and design. The review identifies stacked CIS–DVS sensors as one of the most balanced and competitive architectures because they can support compact integration, synchronized event–frame sensing, and on-chip processing. However, important challenges remain, including color fidelity, demosaicing, event-pixel ratio optimization, calibration, benchmarking, and edge-AI deployment. Finally, we emphasize that future hybrid event–frame sensing systems should be developed through sensor–algorithm–ISP–AI co-design. This review provides practical guidelines for developing next-generation hybrid event–frame sensing systems for both human-perceptual imaging and machine vision. Full article
(This article belongs to the Special Issue Computer Vision-Based Human Activity Recognition)
Show Figures

Figure 1

25 pages, 23777 KB  
Article
Medical Textile Stain Detection Based on Chemically Enhanced Visualization and Deep Semantic Segmentation
by Wenjie Min, Junfeng He, Zhenping Wan, Jinde Chen, Zhixiang Zou and Yuandong Mo
J. Imaging 2026, 12(8), 380; https://doi.org/10.3390/jimaging12080380 - 13 Aug 2026
Viewed by 201
Abstract
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating [...] Read more.
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating chemically enhanced visualization with deep semantic segmentation. Dimethylaminocinnamaldehyde (DMACA) was used to convert latent urine stains into chemically developed stains with orange–red visual features. Based on the spatial color difference ΔE in the L*a*b* color space, 0.0183 mol/L was selected as the most suitable DMACA concentration among those tested. A dataset of 1974 images was constructed, including blood stains, chemically developed urine stains, medication stains, and uncontaminated textiles. A cascaded preprocessing strategy was applied to enhance stain boundaries and suppress textile texture noise, after which an Enhanced semantic segmentation model incorporating residual feature extraction, multiscale feature fusion, and transfer learning was used for pixel-level recognition. The IoU values for blood stains, chemically developed urine stains, and medication stains were 88.11%, 82.67%, and 89.62%, respectively. The average time required for image preprocessing and network inference was 15.39 ms per image. An input-level ablation comparison showed that DMACA-based color development increased the urine-stain IoU from 3.07% to 86.23%, demonstrating its substantial contribution to latent urine-stain detection. These results support the feasibility of integrating front-end chemical feature enhancement with back-end semantic segmentation for multiclass medical textile stain recognition under the current experimental conditions. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

21 pages, 11093 KB  
Article
A Lightweight RGB-LiDAR Feature Recalibration Network for Large-Scale 3D Scene Understanding
by Weifeng Zhai and Zexi Tan
Optics 2026, 7(4), 59; https://doi.org/10.3390/opt7040059 - 13 Aug 2026
Viewed by 173
Abstract
Semantic segmentation of large-scale 3D point clouds is a fundamental task in robotic perception, semantic mapping, and urban scene understanding. Existing methods mainly rely on geometric information, which limits their ability to distinguish semantic categories with similar spatial structures. To address this issue, [...] Read more.
Semantic segmentation of large-scale 3D point clouds is a fundamental task in robotic perception, semantic mapping, and urban scene understanding. Existing methods mainly rely on geometric information, which limits their ability to distinguish semantic categories with similar spatial structures. To address this issue, this paper proposes a lightweight cross-modal feature learning framework that adaptively integrates geometric coordinates and RGB color information. By exploiting the complementary characteristics of spatial structure and visual appearance during feature encoding, the proposed method enhances feature discriminability while maintaining a compact model scale. Experiments on the Semantic3D dataset show that the proposed method achieves an mIoU of 87.1%, outperforming the original RandLA-Net and several representative approaches. Additional runtime and LiDAR-only cross-dataset experiments indicate the potential of the proposed structure for online outdoor point cloud perception. Full article
(This article belongs to the Topic Optical and Laser Scanning: Systems and Applications)
Show Figures

Figure 1

19 pages, 161996 KB  
Article
DBCS-T: A Dual-Branch Cross-Attention Synergistic Transformer for Multimodal Image Fusion and Semantic Segmentation
by Yiming Liu, Bo Gao, Xiao Yang, Hang Li, Weixing Yu and Huangrong Xu
Remote Sens. 2026, 18(16), 2700; https://doi.org/10.3390/rs18162700 - 11 Aug 2026
Viewed by 219
Abstract
Spectro-polarimetric imaging systems can simultaneously acquire spatial, spectral, and polarimetric information during remote sensing, yet the multimodal fusion data are often constrained in practical applications by insufficient exploitation of complementary information across different modalities. To address this issue, we propose a multimodal image [...] Read more.
Spectro-polarimetric imaging systems can simultaneously acquire spatial, spectral, and polarimetric information during remote sensing, yet the multimodal fusion data are often constrained in practical applications by insufficient exploitation of complementary information across different modalities. To address this issue, we propose a multimodal image fusion method based on a Dual-Branch Cross-Attention Synergistic Transformer (DBCS-T). In our method, three complementary feature components are independently extracted, i.e., a Characteristic Polarization Image (CPI), a Characteristic Spectral Image (CSI) and a Characteristic Intensity Image (CII). For CPI, it is derived from Angle of Linear Polarization (AoLP) and Degree of Linear Polarization (DoLP) inputs via DBCS-T, which integrates a Cross-Channel Transposed Attention (CCTA) module for cross-modal interaction and a Multi-Scale Polarization Feature Adaptive Modulation (MPAM) module for local feature enhancement. CSI is obtained by leveraging maximum-divergence spectral band differences guided by prior spectral radiance curves. CII is computed from the Stokes parameter S0. These three components are then fused via Principal Component Analysis (PCA) into a Multimodal Fusion Image (MFI). To demonstrate the effectiveness of our method, an experiment was conducted on a scene that contains real vegetation, artificial foliage, and same-color metallic objects. The experimental results show that the proposed method achieves effective semantic segmentation of all three target categories. Furthermore, quantitative evaluation demonstrates that the MFI attains the lowest Kullback–Leibler (KL) and Jensen–Shannon (JS) Divergence values among all evaluated modalities, with image entropy exceeding that of individual source inputs. These results validate the complementarity of the extracted multimodal features, significantly enhance the interpretation performance for complex scenes, and demonstrate the broad application potential of the proposed fusion framework in remote sensing and multidimensional imaging. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

22 pages, 4170 KB  
Article
A Direction-Aware Dual-Branch Network for Surface-Strand Orientation Segmentation of Oriented Strand Board
by Changyu Zhang, Yanyi Liu and Yin Wu
Sensors 2026, 26(16), 5055; https://doi.org/10.3390/s26165055 - 9 Aug 2026
Viewed by 193
Abstract
The angular distribution of surface-strands in oriented strand board (OSB) is closely associated with board mechanical properties and mat formation quality. By acquiring surface images through vision sensing and combining them with deep learning-based segmentation, the angle classes of OSB surface-strands can be [...] Read more.
The angular distribution of surface-strands in oriented strand board (OSB) is closely associated with board mechanical properties and mat formation quality. By acquiring surface images through vision sensing and combining them with deep learning-based segmentation, the angle classes of OSB surface-strands can be segmented and statistically analyzed automatically. However, OSB surface images contain complex strand textures, blurred boundaries, local adhesion between adjacent strands, and subtle differences among neighboring angle classes. To address these challenges, this study proposes a direction-aware dual-branch semantic segmentation network (DiBiNet) for pixel-level segmentation of surface-strand angle classes. OSB surface images were collected using a Hikrobot MV-CE120-10UC color industrial camera, and an 11-class dataset was constructed, including the background and ten angle classes from 0° to 90°. The samples were cropped to 512 × 512 pixels, and an improved angle-semantic-consistent Copy–Paste strategy was used to augment the training data. DiBiNet enhances directional feature representation through a Directional Strip Detail Enhancement Module, improves semantic feature modeling by combining MobileNetV3-Small with a DS-MobileViT Block, and fuses the two branches through a Bilateral Gated Fusion Module. Considering the continuity among angle classes, Direction Vector Auxiliary Supervision is introduced to map discrete angle labels into continuous direction vectors, thereby improving discrimination among neighboring classes. Experiments on the self-constructed dataset show that DiBiNet achieves a mean Intersection over Union (mIoU) of 0.8532, an overall pixel accuracy (Acc) of 0.8823, and a Dice coefficient of 0.8623, outperforming several representative semantic segmentation models. After 8-bit integer (INT8) + 16-bit floating-point (FP16) mixed quantization, the model achieves a neural processing unit (NPU) inference speed of 34.0 frames per second (FPS) on the RK3588 platform, demonstrating its potential for vision-based sensing and edge AI inspection. Full article
Show Figures

Figure 1

41 pages, 2283 KB  
Article
PartSense-IP: Part-Aware Vision–Language Sensor Fusion for Visual–Semantic Consistency Evaluation of IP Prototypes
by Yangfan Feng and Wen Zhao
Sensors 2026, 26(16), 5052; https://doi.org/10.3390/s26165052 - 9 Aug 2026
Viewed by 227
Abstract
Evaluating whether an intellectual property (IP) prototype faithfully preserves the visual identity and semantic intent of its original concept design is an important yet challenging task in product design and creative prototyping. Existing evaluation practices mainly rely on manual inspection or global image-level [...] Read more.
Evaluating whether an intellectual property (IP) prototype faithfully preserves the visual identity and semantic intent of its original concept design is an important yet challenging task in product design and creative prototyping. Existing evaluation practices mainly rely on manual inspection or global image-level similarity comparison, which are subjective, difficult to reproduce, and insufficient for localizing identity-critical deviations. To address this problem, this paper proposes PartSense-IP, a part-aware vision–language sensor fusion framework for visual–semantic consistency evaluation of IP prototypes. The proposed framework takes a 2D concept image, an optional textual design description, and multi-view RGB-D sensor observations of a prototype as inputs. It first constructs a multi-view prototype representation and decomposes both the concept and prototype observations into design-relevant parts. Dense visual features, color and shape descriptors, and vision–language semantic embeddings are then extracted to evaluate part-level consistency. A Part-Aware Visual–Semantic Consistency Fusion (PVCF) algorithm is further developed to integrate shape, color, local visual similarity, semantic alignment, and cross-view stability into a unified IP consistency score. In addition to scalar scoring, PartSense-IP generates localized difference maps, 3D inconsistency visualization, and interpretable design feedback for prototype refinement. Experiments on the proposed IP-ProtoSense evaluation protocol demonstrate that PartSense-IP outperforms representative vision–language, dense-visual, segmentation-based, and 3D multimodal baselines in consistency scoring, inconsistency detection, localization, ablation, and robustness evaluation. Full article
(This article belongs to the Section Optical Sensors)
Show Figures

Figure 1

29 pages, 4057 KB  
Article
Digital Twin-Ready Framework for the Automated Management of Urban Green Spaces: A Case Study
by Giuliana Parisi, Alessia Ursino and Rosa Caponetto
Sustainability 2026, 18(16), 8092; https://doi.org/10.3390/su18168092 - 8 Aug 2026
Viewed by 253
Abstract
In response to the growing demand for sustainable regeneration and efficient operation of urban spaces, the adoption of digitalised facility management (FM) has emerged as an effective strategy. FM facilitates data-driven planning and optimised resource use, contributing to more sustainable and resilient urban [...] Read more.
In response to the growing demand for sustainable regeneration and efficient operation of urban spaces, the adoption of digitalised facility management (FM) has emerged as an effective strategy. FM facilitates data-driven planning and optimised resource use, contributing to more sustainable and resilient urban environments. This research combines Building Information Modelling (BIM) with Visual Programming Languages (VPLs) to support the regeneration and optimisation of an urban park. Autodesk Revit is employed to develop an LOD 300 semantic and geometric model, while Dynamo enables the implementation of two parametric scripts: Script A, governing on/off control of monitored systems based on environmental, temporal and contextual conditions, and Script B, providing continuous operational monitoring, fault diagnosis and color-coded visual feedback directly within the BIM environment. The framework is tested at Parco Gioeni, an 8.6-hectare historically and geologically urban park in Italy, across three management systems: irrigation, nebulization and green area monitoring. A total of 28 what-if scenarios are conducted across systems, indicating the framework’s functional consistency and contextual adaptability under the tested conditions. The proposed approach is built on accessible open software tools, making it transferable to public administration contexts, and is designed to be scalable and replicable across different urban green space typologies. Full article
(This article belongs to the Section Green Building)
Show Figures

Figure 1

40 pages, 3662 KB  
Article
A Semantic-Conditional GAN Framework for Structure-Preserving and Controllable Interior Style Generation
by Tianxi Lu, Chang Wen and Siti Sarah Binti Herman
Appl. Sci. 2026, 16(16), 7892; https://doi.org/10.3390/app16167892 - 7 Aug 2026
Viewed by 202
Abstract
Interior style generation requires expressive visual transformation while preserving spatial layout, object boundaries, and semantic relationships. Existing generative and style-transfer methods have improved indoor scene synthesis, but they often suffer from boundary drift, furniture deformation, texture leakage, and unstable style representation when stronger [...] Read more.
Interior style generation requires expressive visual transformation while preserving spatial layout, object boundaries, and semantic relationships. Existing generative and style-transfer methods have improved indoor scene synthesis, but they often suffer from boundary drift, furniture deformation, texture leakage, and unstable style representation when stronger style signals are introduced. To address this problem, this study proposes a semantic-conditional generative adversarial network framework for structure-preserving and controllable interior style generation. The framework constructs semantic masks, boundary maps, and semantic embeddings from indoor images and injects these semantic priors into the generator feature stream through a multi-scale conditional mechanism. A style encoding network is further introduced to represent interior style characteristics and to modulate image-level visual appearance features, including texture patterns, color tones, material-like surface appearance, and illumination-related visual cues, in a controllable manner. The model was trained and evaluated using 5000 selected indoor images from SUN RGB-D and ADE20K, together with a self-constructed style reference set of 600 images covering modern minimalist, Nordic, industrial, and neoclassical interiors. Across baseline comparisons and ablation analyses, the proposed framework achieved a semantic region consistency score of 0.91, maintained higher semantic boundary consistency under increasing style intensity, obtained style consistency scores ranging from 0.88 to 0.90, and achieved an LPIPS score of 0.128. A subjective evaluation with 20 participants also showed higher ratings for realism, style expression, and structural consistency. These results provide quantitative and perceptual evidence that semantic-conditional injection can improve the balance between spatial-semantic preservation and controllable image-level style expression in AI-assisted interior visualization. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

33 pages, 23794 KB  
Article
Navigation Line Extraction Method for Alfalfa Crops Based on RACG-RandLA Point Cloud Segmentation Model
by Kehua Dang, Jiachen Cao, Pengjie Pan, Zijie Niu, Zehan Lu, Dongyan Zhang and Yongjie Cui
Agriculture 2026, 16(15), 1683; https://doi.org/10.3390/agriculture16151683 - 5 Aug 2026
Viewed by 344
Abstract
Early-stage alfalfa navigation faces challenges like low plants, narrow rows, and weed interference, causing camera–LiDAR colored point clouds to suffer from sparsity, discontinuous boundaries, and varying illumination. Standard point-level semantic segmentation struggles to support stable crop row allocation and navigation line fitting under [...] Read more.
Early-stage alfalfa navigation faces challenges like low plants, narrow rows, and weed interference, causing camera–LiDAR colored point clouds to suffer from sparsity, discontinuous boundaries, and varying illumination. Standard point-level semantic segmentation struggles to support stable crop row allocation and navigation line fitting under these conditions. To address this, we propose RACG-RandLA, a multi-output row-aware color-geometric point cloud segmentation model. Pseudo-labels (crop, background, ‘ignore’) are generated using color and spatial priors, alongside transverse offset and point-level confidence labels for crop points. Built on RandLA-Net, the multi-task network simultaneously outputs crop semantics, transverse offsets, and confidences. It features a color-geometry residual fusion module that adaptively integrates 3D geometric and RGB/ExG features via zero-initialized scaling to handle complex lighting and missing data. Additionally, a late Row-cued Local Feature Aggregation (LFA) module embeds longitudinal continuity and transverse offset constraints into deep layers, effectively mitigating cross-row feature aliasing. During inference, a multi-output pipeline utilizes these predictions for reliable crop point filtering, center refinement, and navigation line fitting. Experiments show that the xyzrgb_exg input achieves an optimal balance between accuracy and conciseness. RACG-RandLA achieves a Test IoU of 0.8590, outperforming PointNet variants, and secures the highest Row Count Accuracy of 0.8761. Furthermore, it reduces lateral navigation jitter to 0.0241 m while maintaining an 85.04% success rate. Ultimately, the proposed method demonstrates a superior balance of semantic accuracy, structural consistency, and navigation stability, providing a robust perception solution for agricultural robots. Full article
(This article belongs to the Section Artificial Intelligence and Digital Agriculture)
Show Figures

Figure 1

Back to TopTop