Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (54)

Search Parameters:
Keywords = GeoVideo

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
43 pages, 25524 KB  
Article
Beyond Novice and Expert: Domain Proficiency in Simulation-Based Evaluation of Geo-Located Augmented Reality Interfaces
by Katarina Mišura, Mirko Sužnjević and Miro Čolić
Appl. Sci. 2026, 16(19), 9587; https://doi.org/10.3390/app16199587 - 26 Sep 2026
Abstract
This paper investigates how Quality of Experience (QoE) and User Experience (UX) evaluations of a geo-located augmented reality (AR) interface designed to support navigation and situational awareness in a security context differ across participant groups with distinct domain-related backgrounds. To address the limitations [...] Read more.
This paper investigates how Quality of Experience (QoE) and User Experience (UX) evaluations of a geo-located augmented reality (AR) interface designed to support navigation and situational awareness in a security context differ across participant groups with distinct domain-related backgrounds. To address the limitations of current wearable AR hardware, the study used a first-person video game simulation of a person using AR glasses in an urban security scenario. The proposed interface consisted of three core elements: Map in the Sky, navigational arrows, and tags for objects or persons of interest. Data from two user studies conducted under the same experimental protocol were combined, yielding three participant groups with different domain-related backgrounds: Civilians, Veterans, and Cadets. The interface was evaluated using the Acceptance Scale (AS), the System Usability Scale (SUS), NASA-RTLX, and qualitative interviews. The results showed generally positive evaluations of the interface and all individual elements, with high SUS-rated usability, relatively low-to-moderate reported workload, and predominantly favorable interview responses. However, the pattern of results differed across participant groups. Civilians provided the most positive overall usability ratings; Veterans also evaluated the interface positively but reported the highest workload, and Cadets produced the least favorable usability and acceptance scores despite greater self-reported familiarity with AR and video games. Across groups, Map in the Sky received the lowest ratings, whereas arrows and tags were evaluated more positively. These findings indicate that AR QoE and UX outcomes can differ across participant groups with distinct domain-related backgrounds in ways that do not follow a simple novice-to-expert ordering. The results also demonstrate the practical feasibility of using PC-based simulation to obtain early-stage subjective QoE/UX feedback, while highlighting the importance of participant selection when evaluating domain-specific AR systems. Because the study did not include objective performance measures or a comparison condition, the findings reflect perceived usability, acceptance, and workload rather than operational effectiveness or workload reduction. Full article
(This article belongs to the Special Issue Augmented and Virtual Reality for Smart Applications)
►▼ Show Figures

Figure 1

28 pages, 4789 KB  
Article
Geometry-Constrained Multi-Frame Character Association for License Plate Recognition on Moving Cameras
by Ufuk Asil and İlker Yoncacı
Sensors 2026, 26(18), 5704; https://doi.org/10.3390/s26185704 - 8 Sep 2026
Viewed by 402
Abstract
Multi-frame fusion is standard for converting frame-by-frame license plate character detections into stable, reliable readings. Character Time-Series Matching (CTM), a leading approach, associates characters across frames using the Hungarian algorithm with a fixed Euclidean distance threshold and a translation-only motion model, reporting 96.7% [...] Read more.
Multi-frame fusion is standard for converting frame-by-frame license plate character detections into stable, reliable readings. Character Time-Series Matching (CTM), a leading approach, associates characters across frames using the Hungarian algorithm with a fixed Euclidean distance threshold and a translation-only motion model, reporting 96.7% accuracy on the UFPR-ALPR dataset. In this work, we demonstrate that this high performance is protocol-dependent: when ground-truth static plate crops and pre-segmented tracks are used, CTM performs strongly. However, in real-world scenarios involving moving cameras (such as drone-mounted cameras, helmet-mounted cameras, and mobile platforms) where inter-frame geometry changes dynamically, baseline multi-frame association frameworks that combine fixed spatial gates with unconstrained translation propagation fail. In these cases, temporal fusion provides no benefit and degrades plate recognition performance below the single-frame baseline. Indeed, under its own Intersection over Union (IoU) tracker, this literature method correctly reads only 15.1% of plates in traffic videos recorded with a real moving camera (86 human-verified tracks). To address this vulnerability, we propose Geo-CTM (Geometry-Constrained CTM), an association pipeline integrating height-scaled adaptive matching gates, inter-frame similarity estimation via Random Sample Consensus (RANSAC), transform-guided character coasting, and co-occurrence-constrained duplicate track elimination. Systematic motion-model ablation demonstrates that while the complete association pipeline provides the primary foundation for robustness (raising mean accuracy from 85.40% to over 91.5%), estimating a similarity transform (91.82%) delivers the most physically grounded and identifiable representation on planar plates without estimation degeneration. While our method performs comparably to CTM on ideal data when using the same detector and detections, it minimizes performance loss under geometric distortion conditions where CTM is inadequate. For instance, a statistically significant improvement is achieved under a 0 → 60° perspective change; in real traffic videos, with the tracker held fixed so that the fusion layer is the only variable, performance rises from 15.1% to 26.7% under the IoU tracker of the original system and from 16.3% to 29.1% under ByteTrack (+11.6 and +12.8 points; exact McNemar p=0.021 and p=0.013), whereas changing the tracker alone while holding the fusion layer fixed moves accuracy by only 1–2 points and is not statistically significant. Finally, our error taxonomy analysis demonstrates that on the undistorted benchmark all residual errors correspond to zero-evidence cases beyond the reach of decision-level fusion, while under dynamic perspective distortion errors are dominated by association misalignment, highlighting the specific development areas that future performance improvements must target. Full article
(This article belongs to the Special Issue Advanced Pattern Recognition: Intelligent Sensing and Imaging)
►▼ Show Figures

Figure 1

19 pages, 23860 KB  
Article
GeoGATE: Geo-Sensor-Guided Adaptive Token and Evidence Reasoning for High-Resolution Remote Sensing Image Understanding
by Jingnan Zhang and Fengjun Zhang
Appl. Sci. 2026, 16(17), 8616; https://doi.org/10.3390/app16178616 - 29 Aug 2026
Viewed by 279
Abstract
High-resolution remote sensing understanding requires models to preserve small spatial evidence, account for acquisition-dependent appearance, and separate genuine geographic change from nuisance variation. We introduce GeoGATE, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware [...] Read more.
High-resolution remote sensing understanding requires models to preserve small spatial evidence, account for acquisition-dependent appearance, and separate genuine geographic change from nuisance variation. We introduce GeoGATE, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning. LoRA adaptation and NF4 quantization support efficient training and deployment. On the VRSBench test split, GeoGATE reaches 53.4 BLEU-1, 36.8 BLEU-2, 18.2 BLEU-4, 56.4 Acc@0.5, 82.3 VQA, 25.1 METEOR, and 42.6 ROUGE-L, outperforming the controlled GeoGATE (Base) configuration across captioning, question answering, and grounding. Component ablations associate adaptive slicing most strongly with localization, retrieval with language and VQA, and language model adaptation with all reported tasks. NF4 reduces measured video memory from 24.5 GiB to 7.2 GiB with only minor metric changes. These experiments support the single-image language and grounding components. Dedicated cross-sensor and bi-temporal benchmarks are not reported; the corresponding modules are therefore presented as architectural extensions rather than validated performance claims. Full article
►▼ Show Figures

Figure 1

31 pages, 36482 KB  
Article
Geo-Consistent Centralized Multi-UAV Gaussian SLAM for Incremental Orthophoto Generation
by Xiao Zhang, Shuaixin Li, Hongbin Dong, Xiaozhou Zhu, Haoxin Zhang and Baosong Deng
Remote Sens. 2026, 18(16), 2804; https://doi.org/10.3390/rs18162804 - 19 Aug 2026
Viewed by 441
Abstract
Online incremental orthophoto generation with multiple unmanned aerial vehicles (UAVs) remains challenging, as it requires accurate, efficient, and scalable mapping from distributed aerial observations. In this paper, we present a centralized GNSS-assisted multi-UAV 3D Gaussian Splatting SLAM framework for online incremental orthophoto mapping. [...] Read more.
Online incremental orthophoto generation with multiple unmanned aerial vehicles (UAVs) remains challenging, as it requires accurate, efficient, and scalable mapping from distributed aerial observations. In this paper, we present a centralized GNSS-assisted multi-UAV 3D Gaussian Splatting SLAM framework for online incremental orthophoto mapping. Each UAV independently performs visual odometry to build local submaps, which are first aligned into a unified global coordinate system using GNSS constraints and further refined via inter-agent visual loop closures for improved cross-agent consistency. To enable scalable and high-quality mapping, we introduce two complementary Gaussian map maintenance modules: plane-guided grid-based collaborative densification, which improves mapping quality and accelerates convergence under multi-UAV conditions, and visibility-aware adaptive pruning, which effectively controls redundancy and memory usage. These components allow efficient joint optimization within a unified Gaussian representation. Experiments on multiple aerial datasets using video-derived image frames captured by consumer-grade UAV cameras demonstrate that the proposed system provides a favorable trade-off between geo-consistency, visual fidelity, and efficiency compared with existing methods. Quantitatively, the proposed method achieves a GCP RMSE of 2.32 m, completes multi-UAV orthophoto generation within 3.3–5.5 min, and reduces the total mapping time by approximately 35–55% compared with the corresponding single-UAV setting, while supporting online tracking and incremental orthophoto updates with bounded latency and memory consumption. Full article
►▼ Show Figures

Figure 1

37 pages, 2981 KB  
Article
Signs, Shapes, and Spaces: A CAMIL-Informed Qualitative Study of Metaverse Geometry Learning for Deaf and Hard-of-Hearing Students
by Ai Peng Chong, Kung-Teck Wong, Kong Liang Soon Vestly and Kuppusamy Suresh Kumar
Soc. Sci. 2026, 15(3), 191; https://doi.org/10.3390/socsci15030191 - 16 Mar 2026
Cited by 1 | Viewed by 1701
Abstract
Deaf and Hard-of-Hearing (DHH) students face persistent barriers in geometry education due to instructional approaches that inadequately support visual communication and embodied learning. This study examined DHH students’ experiences with GeoMETriA, a metaverse-based geometry learning platform integrating sign language instruction, three-dimensional visualization, and [...] Read more.
Deaf and Hard-of-Hearing (DHH) students face persistent barriers in geometry education due to instructional approaches that inadequately support visual communication and embodied learning. This study examined DHH students’ experiences with GeoMETriA, a metaverse-based geometry learning platform integrating sign language instruction, three-dimensional visualization, and avatar-mediated interaction. Guided by the Cognitive Affective Model of Immersive Learning (CAMIL), a multi-phase qualitative design was employed, including pre-workshop interviews with four special education teachers and post-workshop focus group discussions with seven DHH secondary students following a four-session learning workshop. The findings indicate that gamified activities and peer collaboration enhanced interest and sustained engagement, while avatar customization supported embodiment and a sense of presence. Students described progression from initial uncertainty to greater confidence through practice and scaffolded support. However, cognitive and usability challenges emerged, particularly concerning sign language video pacing, navigation complexity, and limited instructional scaffolding. The study contributes theoretically by extending CAMIL-informed interpretations to sign-supported metaverse learning, empirically by documenting how engagement, embodiment, and self-efficacy develop during immersive geometry learning, and practically by offering design implications including adjustable sign language delivery, structured scaffolding, and culturally responsive avatar options. These findings suggest that metaverse-based platforms hold promise for supporting DHH learners when accessibility and learner-centered principles are embedded as foundational design considerations. Full article
(This article belongs to the Special Issue Belt and Road Together Special Education 2025)
►▼ Show Figures

Figure 1

20 pages, 10112 KB  
Article
Satellite Backhaul for Extending Connectivity in Rural Remote Areas: Deployment and Performance Assessment
by Souhaima Stiri, Maria Rita Palattella, Juan David Niebles Castano and Christos Politis
Network 2026, 6(1), 12; https://doi.org/10.3390/network6010012 - 24 Feb 2026
Viewed by 3503
Abstract
Limited terrestrial network coverage in rural and remote areas constitutes a significant barrier to the digital transformation of the agricultural sector. Smart and precision farming applications, ranging from conventional environmental monitoring systems to advanced Digital Twin solutions, rely on the reliable transmission of [...] Read more.
Limited terrestrial network coverage in rural and remote areas constitutes a significant barrier to the digital transformation of the agricultural sector. Smart and precision farming applications, ranging from conventional environmental monitoring systems to advanced Digital Twin solutions, rely on the reliable transmission of sensor data, images, and video streams from geographically isolated farms. Such data-intensive services cannot be effectively supported without a robust communication infrastructure. Non-Terrestrial Networks (NTNs), particularly satellite systems, offer both narrowband and broadband connectivity, enabling the transmission of low-rate sensor measurements, as well as high-throughput multimedia data from the field. This paper presents an experimental performance evaluation of two satellite backhauling solutions: a Geostationary Earth Orbit (GEO) system provided by SES and a Low Earth Orbit (LEO) system from Starlink. The networks were first deployed and tested in a laboratory environment and subsequently validated in an operational agricultural field setting. Their performance is benchmarked against a terrestrial cellular network to assess their suitability for supporting advanced agricultural applications. The performance assessment results indicate that both satellite backhauling solutions are reliable and capable of meeting the bandwidth and latency requirements of delay-tolerant agricultural applications. In addition to the technical evaluation, this work presents a cost–benefit analysis that further underscores the advantages of NTN-based solutions. Despite higher initial expenditures, they provide extended coverage in remote areas and enable cost sharing across multiple users, improving overall economic viability. Full article
►▼ Show Figures

Figure 1

31 pages, 8257 KB  
Article
Analytical Assessment of Pre-Trained Prompt-Based Multimodal Deep Learning Models for UAV-Based Object Detection Supporting Environmental Crimes Monitoring
by Andrea Demartis, Fabio Giulio Tonolo, Francesco Barchi, Samuel Zanella and Andrea Acquaviva
Geomatics 2026, 6(1), 14; https://doi.org/10.3390/geomatics6010014 - 3 Feb 2026
Cited by 1 | Viewed by 2031
Abstract
Illegal dumping poses serious risks to ecosystems and human health, requiring effective and timely monitoring strategies. Advances in uncrewed aerial vehicles (UAVs), photogrammetry, and deep learning (DL) have created new opportunities for detecting and characterizing waste objects over large areas. Within the framework [...] Read more.
Illegal dumping poses serious risks to ecosystems and human health, requiring effective and timely monitoring strategies. Advances in uncrewed aerial vehicles (UAVs), photogrammetry, and deep learning (DL) have created new opportunities for detecting and characterizing waste objects over large areas. Within the framework of the EMERITUS Project, an EU Horizon Europe initiative supporting the fight against environmental crimes, this study evaluates the performance of pre-trained prompt-based multimodal (PBM) DL models integrated into ArcGIS Pro for object detection and segmentation. To test such models, UAV surveys were specially conducted at a semi-controlled test site in northern Italy, producing very high-resolution orthoimages and video frames populated with simulated waste objects such as tyres, barrels, and sand piles. Three PBM models (CLIPSeg, GroundingDINO, and TextSAM) were tested under varying hyperparameters and input conditions, including orthophotos at multiple resolutions and frames extracted from UAV-acquired videos. Results show that model performance is highly dependent on object type and imagery resolution. In contrast, within the limited ranges tested, hyperparameter tuning rarely produced significant improvements. The evaluation of the models was performed using low IoU to generalize across different types of detection models and to focus on the ability of detecting object. When evaluating the models with orthoimagery, CLIPSeg achieved the highest accuracy with F1 scores up to 0.88 for tyres, whereas barrels and ambiguous classes consistently underperformed. Video-derived (oblique) frames generally outperformed orthophotos, reflecting a closer match to model training perspectives. Despite the current limitations in performances highlighted by the tests, PBM models demonstrate strong potential for democratizing GeoAI (Geospatial Artificial Intelligence). These tools effectively enable non-expert users to employ zero-shot classification in UAV-based monitoring workflows targeting environmental crime. Full article
►▼ Show Figures

Figure 1

26 pages, 12124 KB  
Article
MF-GCN: Multimodal Information Fusion Using Incremental Graph Convolutional Network for Ship Behavior Anomaly Detection
by Ruixin Ma, Jinhao Zhang, Weizhi Nie, Naiming Ge, Hao Wen and Aoxiang Liu
J. Mar. Sci. Eng. 2026, 14(1), 87; https://doi.org/10.3390/jmse14010087 - 1 Jan 2026
Viewed by 1420
Abstract
Ship behavior anomaly detection is critical for intelligent perception and early warning in complex inland waterways, where single-source sensing (e.g., AIS-only or vision-only) is often fragile under occlusion, illumination variation, and signal noise. This study proposes MF-GCN, a multimodal (heterogeneous) information fusion framework [...] Read more.
Ship behavior anomaly detection is critical for intelligent perception and early warning in complex inland waterways, where single-source sensing (e.g., AIS-only or vision-only) is often fragile under occlusion, illumination variation, and signal noise. This study proposes MF-GCN, a multimodal (heterogeneous) information fusion framework based on an Incremental Graph Convolutional Network (IGCN) to detect and warn anomalous ship behaviors by jointly modeling AIS, video imagery, LiDAR point clouds, and water level signals. We first extract modality-specific features and enforce temporal–spatial consistency via timestamp and geo-referencing alignment, then construct an evolving graph in which nodes represent multimodal features and edges encode temporal dependency and semantic similarity. MF-GCN integrates a Semantic Clustering-based GCN (S-GCN) to inject historical semantic context and an Attentive Fusion-based GCN (A-GCN) to learn dynamic cross-modal correlations using multi-head attention. Experiments on our constructed real-world datasets demonstrate that MF-GCN achieves accuracies of 93.8%, 93.8%, and 93.3% with F1-scores of 93.6%, 93.6%, and 93.3% for ship deviation warning, bridge-crossing warning, and inter-ship collision warning, respectively, consistently outperforming representative baselines. These results verify the effectiveness of the proposed method for robust multimodal anomaly detection and early warning in inland-waterway scenarios. Full article
(This article belongs to the Special Issue Emerging Computational Methods in Intelligent Marine Vehicles)
►▼ Show Figures

Figure 1

20 pages, 1115 KB  
Systematic Review
Mathematics Teachers’ Knowledge for Teaching with Digital Technologies: A Systematic Review of Studies from 2010 to 2025
by Iván Andrés Padilla-Escorcia, Martha Leticia García-Rodríguez and Álvaro Aguilar-González
Educ. Sci. 2025, 15(12), 1598; https://doi.org/10.3390/educsci15121598 - 26 Nov 2025
Cited by 4 | Viewed by 2461
Abstract
This systematic review examines mathematics teachers’ knowledge for teaching using digital technologies (DTs), understood as the intersection of disciplinary, pedagogical, and technological domains that teachers mobilize when designing, implementing, and assessing mathematics lessons. In this study, DTs refer to the digital hardware, software, [...] Read more.
This systematic review examines mathematics teachers’ knowledge for teaching using digital technologies (DTs), understood as the intersection of disciplinary, pedagogical, and technological domains that teachers mobilize when designing, implementing, and assessing mathematics lessons. In this study, DTs refer to the digital hardware, software, and online environments used to represent, simulate, or analyze mathematical ideas (e.g., GeoGebra, Tinkerplots, spreadsheets, CAS tools, and learning management systems). We analyzed 50 peer-reviewed journal articles published between January 2010 and April 2025, retrieved from Web of Science, Scopus, ERIC, and Scielo. ResearchGate was consulted only as a supplementary repository to access the full texts already identified in the indexed databases. These articles were analyzed according to predefined analytical categories, including research themes, country of origin, and the digital technologies addressed in each study, allowing for cross-comparisons across theoretical frameworks and methodological approaches. The results reveal a strong interest in this topic in countries such as Turkey, the United States, Mexico, Indonesia, and Spain, with the participation of in-service mathematics teachers at the primary, secondary, and university levels, as well as preservice teachers. The most frequently studied themes in the past five years regarding teacher knowledge include teacher education through digital technologies, the analysis of lesson planning and tasks designed by teachers using DTs, and the assessment of their knowledge through self-perception questionnaires. The review concludes that only a few of the analyzed studies qualitatively examined teacher knowledge when using digital technologies, particularly those that employed non-participant observation, audio and/or video recordings, and semi-structured interviews. Full article
(This article belongs to the Section Technology Enhanced Education)
►▼ Show Figures

Figure 1

15 pages, 1040 KB  
Article
Human Preferences for Animals on YouTube
by Pavol Prokop, Rudolf Masarovič and Tomáš Vranovský
Diversity 2025, 17(10), 720; https://doi.org/10.3390/d17100720 - 15 Oct 2025
Cited by 2 | Viewed by 2590
Abstract
Social media has emerged as a dominant platform for sharing human–animal interactions, creating a powerful tool for public engagement and wildlife conservation. Consequently, we sought to determine whether analyzing user preferences for animals on social networks could inform the management of effective conservation [...] Read more.
Social media has emerged as a dominant platform for sharing human–animal interactions, creating a powerful tool for public engagement and wildlife conservation. Consequently, we sought to determine whether analyzing user preferences for animals on social networks could inform the management of effective conservation campaigns. We analyzed 5129 videos from three channels (Brave Wilderness, BBC Earth, and Nat Geo Wild) available on YouTube, which have millions of followers each. The mean number of “likes” was used as a proxy for animal species preferences. Contrary to the general expectation that humans predominantly prefer charismatic animals (e.g., terrestrial mammals), the most preferred animals on these channels were from the classes Amphibia, Arachnida, and Insecta, which significantly outperformed mammals and birds. Viewers most frequently consumed videos of stinging insects or threatening animals, and domestic animals received more likes than wild animals. Furthermore, contrary to expectations, body mass, IUCN conservation status, and daytime activity of mammals and birds did not significantly influence human preferences. Our results suggest that although viewing animal videos may have a negligible direct conservation impact, the analysis of preferences reveals that creators successfully captured human attention toward less popular animal taxa, highlighting potential indirect benefits. Future research should integrate audience enjoyment of frightening content with conservation intentions. Full article
(This article belongs to the Section Biodiversity Conservation)
►▼ Show Figures

Figure 1

22 pages, 2382 KB  
Article
A Quantitative Study on Multipoint Video Distribution Systems MVDS Interference to GEO Satellites in Lebanon
by Ali Karaki, Hiba Abdalla, Mohammed Al-Husseini and Hamza Issa
Telecom 2025, 6(2), 36; https://doi.org/10.3390/telecom6020036 - 28 May 2025
Viewed by 1335
Abstract
This paper investigates the potential for interference from multipoint video distribution systems (MVDS) transmissions, specifically side lobe radiation in Lebanon, to geostationary Earth orbit (GEO) satellites. Through simulation and analysis of antenna radiation patterns, the impact of varying MVDS power levels on the [...] Read more.
This paper investigates the potential for interference from multipoint video distribution systems (MVDS) transmissions, specifically side lobe radiation in Lebanon, to geostationary Earth orbit (GEO) satellites. Through simulation and analysis of antenna radiation patterns, the impact of varying MVDS power levels on the carrier-to-noise ratio (C/N) at the satellite receiver is quantified. The results demonstrate a significant degradation in signal quality, with the C/N dropping to −2.29 dB at an MVDS power of 0 dBW for the current system. To mitigate this interference, a two-step potential strategy is proposed and evaluated. The study boosts the potential for the coexistence of MVDS and GEO satellite services in the Ku-band within the Lebanese context. Full article
►▼ Show Figures

Figure 1

21 pages, 14761 KB  
Article
GeoIoU-SEA-YOLO: An Advanced Model for Detecting Unsafe Behaviors on Construction Sites
by Xuejun Jia, Xiaoxiong Zhou, Zhihan Shi, Qi Xu and Guangming Zhang
Sensors 2025, 25(4), 1238; https://doi.org/10.3390/s25041238 - 18 Feb 2025
Cited by 10 | Viewed by 2710
Abstract
Unsafe behaviors on construction sites are a major cause of accidents, highlighting the need for effective detection and prevention. Traditional methods like manual inspections and video surveillance often lack real-time performance and comprehensive coverage, making them insufficient for diverse and complex site environments. [...] Read more.
Unsafe behaviors on construction sites are a major cause of accidents, highlighting the need for effective detection and prevention. Traditional methods like manual inspections and video surveillance often lack real-time performance and comprehensive coverage, making them insufficient for diverse and complex site environments. This paper introduces GeoIoU-SEA-YOLO, an enhanced object detection model integrating the Geometric Intersection over Union (GeoIoU) loss function and Structural-Enhanced Attention (SEA) mechanism to improve accuracy and real-time detection. GeoIoU enhances bounding box regression by considering geometric characteristics, excelling in the detection of small objects, occlusions, and multi-object interactions. SEA combines channel and multi-scale spatial attention, dynamically refining feature map weights to focus on critical features. Experiments show that GeoIoU-SEA-YOLO outperforms YOLOv3, YOLOv5s, YOLOv8s, and SSD, achieving high precision (mAP@0.5 = 0.930), recall, and small object detection in complex scenes, particularly for unsafe behaviors like missing safety helmets, vests, or smoking. Ablation studies confirm the independent and combined contributions of GeoIoU and SEA to performance gains, providing a reliable solution for intelligent safety management on construction sites. Full article
►▼ Show Figures

Figure 1

28 pages, 9307 KB  
Article
Application Framework and Optimal Features for UAV-Based Earthquake-Induced Structural Displacement Monitoring
by Ruipu Ji, Shokrullah Sorosh, Eric Lo, Tanner J. Norton, John W. Driscoll, Falko Kuester, Andre R. Barbosa, Barbara G. Simpson and Tara C. Hutchinson
Algorithms 2025, 18(2), 66; https://doi.org/10.3390/a18020066 - 26 Jan 2025
Cited by 9 | Viewed by 5614
Abstract
Unmanned aerial vehicle (UAV) vision-based sensing has become an emerging technology for structural health monitoring (SHM) and post-disaster damage assessment of civil infrastructure. This article proposes a framework for monitoring structural displacement under earthquakes by reprojecting image points obtained courtesy of UAV-captured videos [...] Read more.
Unmanned aerial vehicle (UAV) vision-based sensing has become an emerging technology for structural health monitoring (SHM) and post-disaster damage assessment of civil infrastructure. This article proposes a framework for monitoring structural displacement under earthquakes by reprojecting image points obtained courtesy of UAV-captured videos to the 3-D world space based on the world-to-image point correspondences. To identify optimal features in the UAV imagery, geo-reference targets with various patterns were installed on a test building specimen, which was then subjected to earthquake shaking. A feature point tracking-based algorithm for square checkerboard patterns and a Hough Transform-based algorithm for concentric circular patterns are developed to ensure reliable detection and tracking of image features. Photogrammetry techniques are applied to reconstruct the 3-D world points and extract structural displacements. The proposed methodology is validated by monitoring the displacements of a full-scale 6-story mass timber building during a series of shake table tests. Reasonable accuracy is achieved in that the overall root-mean-square errors of the tracking results are at the millimeter level compared to ground truth measurements from analog sensors. Insights on optimal features for monitoring structural dynamic response are discussed based on statistical analysis of the error characteristics for the various reference target patterns used to track the structural displacements. Full article
(This article belongs to the Special Issue Algorithms for Image Processing and Machine Vision)
►▼ Show Figures

Graphical abstract

17 pages, 18630 KB  
Article
Investigating a Toolchain from Trajectory Recording to Resimulation
by Florian Lüttner, Malte Kracht, Corinna Köpke, Annette Schmitt, Mirjam Fehling-Kaschek, Alexander Stolz and Alexander Reiterer
Appl. Sci. 2024, 14(22), 10682; https://doi.org/10.3390/app142210682 - 19 Nov 2024
Cited by 1 | Viewed by 1539
Abstract
The growing variety of transportation options and increasing traffic congestion pose new challenges for road safety. As a result, there is an intensified focus on developing automated driving features and assistance systems aimed at minimizing accidents caused by human errors. The creation of [...] Read more.
The growing variety of transportation options and increasing traffic congestion pose new challenges for road safety. As a result, there is an intensified focus on developing automated driving features and assistance systems aimed at minimizing accidents caused by human errors. The creation of these systems requires a substantial amount of testing kilometers, with estimates suggesting that around 2.1 billion kilometers would be necessary to ensure that each situation pertinent to the driving function is encountered at least once with a probability of 50%. This paper advances the microscopic simulation of traffic scenarios beyond linear patterns, utilizing the open-source environment openPASS. It addresses the research question of whether existing microscopic simulations are able to realistically represent non-linear traffic scenarios. A comprehensive toolchain integrates simulation with video recordings and laser scans. The study compares recorded traffic flow data with simulations at a T-junction, assessing the realism of vehicle models and trajectory representation. Three scenarios are analyzed, considering vehicles and pedestrians. The 3D geometry of the scene was captured with a laser scanner, enabling the mapping of recorded video data onto a geo-referenced environment. Object trajectories were extracted using an ’Regions with Convolutional Neural Networks features’ object detector. While openPASS simulated vehicle and pedestrian behaviors effectively, limitations in trajectory variability and reaction times were observed. These findings highlight the need for more realistic behavior models. This research emphasizes the necessity for improvements to accommodate complex driving behaviors and pedestrian dynamics. Full article
►▼ Show Figures

Figure 1

18 pages, 3821 KB  
Article
A Placement Method of the 5G Edge Nodes Based on the Hotspot Distribution of Mobile Users
by Ruowei Gui, Xingjun Zhang, Xiaolin Gui and Jinsong Han
Appl. Sci. 2024, 14(13), 5943; https://doi.org/10.3390/app14135943 - 8 Jul 2024
Cited by 5 | Viewed by 2458
Abstract
Due to the emergence of various new applications, such as short videos and online games, higher requirements of their computing and storage capacity are demanded of mobile networks. The traditional cloud computing paradigm has the shortcomings of large latency and high bandwidth demand [...] Read more.
Due to the emergence of various new applications, such as short videos and online games, higher requirements of their computing and storage capacity are demanded of mobile networks. The traditional cloud computing paradigm has the shortcomings of large latency and high bandwidth demand of the core network. Therefore, how to mine the hotspot distribution of these applications and reasonably configure 5G edge nodes to reduce latency and core network bandwidth are facing great challenges. To address these issues, we designed a placement method for the 5G edge nodes based on mobile hotspots. In this method, we first cluster all locations from the user trajectories to obtain the cluster areas. Further, we extract the features, such as the number of users and duration time in all cluster areas, and extract the hotspots from all cluster areas based on the features of each cluster. Then, we introduce the base station’s high load utilization rate and the core network’s bandwidth reduction rate as the optimization parameters to construct the mathematical model of multi-objective optimization. Finally, we formalize the model into a 0–1 integer programming problem and design a greedy algorithm to solve this model. We also complete a series of experiments to evaluate our proposed methods using the GeoLife dataset. The experimental results show that the high load utilization rate can be increased up to 7.69%, and the bandwidth reduction rate of the core network can be improved up to 6.34%. Full article
(This article belongs to the Special Issue 5G and Beyond: Technologies and Communications)
►▼ Show Figures

Figure 1

Back to TopTop