Abstract
As a measure to sustain crops, the presence of irrigation, man-made reservoirs has become very common in regions affected by prolonged periods of low rainfall. Although these reservoirs must be provided with minimum safety facilities, it is also very common that animals or, to a much lesser extent, the people in charge of their maintenance, fall into the reservoir. The reservoirs then become, in most cases, a death trap, as, with plastic walls that are impossible to climb, they rarely have ramps to facilitate exit. This article describes the design of a proposed edge-computing module that, using embedded vision, identifies the fall of people and animals in irrigation reservoirs. The module includes 180-degree panoramic cameras with colour night vision capability and an NVIDIA Jetson Orin Nano Super. The lack of databases covering the problem to be solved has been addressed by generating synthetic videos showing animals or people falling into irrigation reservoirs. The effectiveness of the training carried out using these synthetic sequences has subsequently been successfully validated using images captured in real-world environments.
1. Introduction
Along with land and soil, water is the basis of agricultural production and, given that approximately 95% of food comes from the land, of global food security. Given that the world population is expected to increase to 9.7 billion by 2050, agricultural production (food, fibre and feed) will need to increase by around 50% compared with 2012. One of the critical factors in any strategy to increase agricultural production is improving water management. Globally, total area equipped for irrigation in version 5 of the Global Map of Irrigation Areas is 3.07 million km2, of which 2.55 million km2 were actually irrigated around the year 2005. That is an area almost the size of India. A large part of this area requires artificially created water management systems, both for transport (canals) and storage (reservoirs or tanks). These infrastructures are critical in arid or semi-arid areas, where water is transported from other regions, or where water collected during the rainy season must be stored for use during the dry season [1].
In this context, the use of irrigation reservoirs is a valuable resource for achieving better water management in agriculture. These can store rainwater and excess water from irrigation channels, providing farmers with a reliable and flexible source of water close to their fields when crops require water during periods of insufficient rainfall [2]. Their use has been common practice for centuries, with estimates suggesting that, globally, there are more than 277 million reservoirs of less than one hectare and more than 24 million between one and ten hectares [1]. On the other hand, reservoirs present a multitude of problems and potential risks. On the one hand, they can be quite large or located in elevated areas which, combined with the possible growth of urban developments in their vicinity, makes them potentially dangerous in the event of a breach. They are often designed and built without the suitability and quality standards required for dams and without any intervention by the authorities, who are unaware of their existence. Many of them are old structures that have never undergone technical inspections [3].
The protection of irrigation reservoirs is regulated by a multitude of local and community regulations. Thus, they must be adequately protected by a perimeter fence that prevents access by persons outside the facility (Figure 1), supplemented by signs warning of the danger of falling into them. The perimeter area of the pond must be considered potentially dangerous and high risk for accidents, so personnel who may have access to it for periodic inspection should be trained and never perform this task alone and without the knowledge of another person, as any fall into the irrigation pond could have serious consequences. The irrigation pond can also be a potential risk for local wildlife, which, in search of water, especially in the dry season of the year, may manage to cross the perimeter fence. In addition to the slope that prevents people or animals that fall into the pond from getting out, the walls of the pond are usually covered with geomembranes (plastic material), which become very slippery when wet. For people, the only protection is usually ropes or ladders. It is rare for them to have floating islands or ramps that allow animals to get out. All of this leads to a significant number of drownings, both of people and animals. Focusing on people, a 2011 study for the province of Almería in Spain (8730 ponds registered at that time) documented 19 deaths by drowning in irrigation ponds over a seven-year period [3]. Most of these people were agricultural workers.
Figure 1.
Irrigation pond with perimeter fencing.
The development of solutions to prevent these drownings is usually based on the aforementioned systems to help people and/or animals get out of the pond. However, today it should be possible to devise solutions that allow us to know remotely that a person or animal of a certain size has fallen into the irrigation pond. Thus, combining artificial vision, edge computing, and wireless connectivity, there are commercially available solutions for swimming pools (e.g., AngelEye (https://angeleye.tech/us/us-technology/, accessed on 28 January 2026) or Lynxight (https://enviroprocess.com/en/, accessed on 28 January 2026)). However, we are not aware of these solutions being available for irrigation reservoirs. Clearly, the characteristics associated with both scenarios are very different.
1.1. Contributions
The main contribution of this study is the development of an automated and intelligent system for monitoring irrigation reservoirs, with the aim of preventing incidents related to human drowning and animals falling in. Basically, the system will generate alarms when the presence of a person or animal is detected inside the reservoir. That is, although the system will detect the presence of people or animals at the edge of the reservoir, in principle it will only generate alarms when they are detected falling into the reservoir. This allows for comprehensive monitoring of the irrigation reservoir environment, and the generation of alarms could be modified to alert of their presence near the edge of the reservoir.
To achieve this objective, a series of sub-objectives must be met. They can be summarised as follows:
- Creation of a database of people and animals (in the current version, we consider wild boars, ducks, and turtles) that are located within the reservoir. Obviously, the presence of some of these animals (ducks, turtles) is only to rule out false alarms. There is currently no database covering this issue, which prevents the training of inference models capable of addressing it. Multiple animal species were selected to cover diverse sizes and morphologies: wild boars represent large terrestrial mammals; ducks are medium-sized waterfowl; and turtles are aquatic reptiles. Together, they reflect the heterogeneity of wildlife commonly found in and around irrigation reservoirs. Given the difficulty of creating a database of this type, we used generative AI to create it synthetically.
- Training of a YOLOv11 model for managing the stated problem, and instantiation of the model in an NVIDIA Jetson Orin Nano Super. Specifically, the nano model (YOLO11n) was used. This Ultralytics model is specifically designed to deliver exceptional efficiency [4].
- Design and development of the whole system for on-line monitoring of irrigation ponds based on IoT and Deep Learning. The system is solar-powered and includes 180-degree panoramic cameras.
- The efficiency and robustness of the proposed system in detecting people who have fallen into reservoirs are evaluated through testing in a real-world environment (an irrigation reservoir of 1500 m2). The presence of ducks and turtles was validated in a natural pond of 32,375 m2 approx. Finally, to assess the system’s effectiveness in detecting wild boars in reservoirs, a reference database has been created using images taken from videos recorded in real-life situations.
1.2. Organisation of the Paper
The remainder of the paper is organised as follows: Section 2 analyses previous studies related to the proposed system, most of which do not specifically deal with irrigation ponds but rather with swimming pools. Section 3 presents the methodology and architecture of the proposed system. Section 4 describes the deployment of the system in real environments and shows and analyses the experimental results obtained. Finally, Section 5 concludes this study and discusses future work.
2. Related Work
The automatic drowning detection methods found in the literature focus mainly on swimming pools rather than reservoirs or irrigation channels. In any case, the aim is to integrate edge computing, computer vision, artificial intelligence, and Internet of Things (IoT) connectivity to develop end-to-end drowning prevention proposals. The popularity of Deep Learning-based inference models means that these solutions are currently present in virtually all proposals. While early proposals featured solutions based on ResNet50, more recent ones tend to use different versions of YOLO (You Only Look Once). In the method proposed by Alotaibi [5], a motion sensor is used to detect the presence of an object at the edge of the swimming pool. If this detection occurs, a camera located above the swimming pool is used to capture an image, in which people, animals or objects are detected using a ResNet50. The edge-computing processing is conducted in a Raspberry Pi 3. The images used both to train the model and to validate it are unrealistic (dolls of people or animals are used), raising doubts about how well it would perform in a real-world deployment. In our scenario, the use of motion sensors is not feasible, as the presence of trees or bushes around the perimeter of the reservoir would cause them to trigger constant false alarms. It is also common for wild animals to approach the fence surrounding the reservoir. Their presence on the outside of the fence should not be treated as an alarm situation. Furthermore, the use of non-realistic images does not allow the inference model to be trained properly. To create a practical database for our proposal, we have made use of generative AI. The proposal by Shatnawi et al. [6] present an approach to early detection of drowning that compares five convolutional neural network models (SqueezeNet, GoogleNet, AlexNet, ShuffleNet, and ResNet50). Using a database of 200 images (100 drownings and 100 people swimming), ResNet50 showed the best performance, achieving 100% prediction accuracy. Experiments were conducted using the deep learning toolbox in MATLAB. However, this database size seems far too small for our scenario. The presence of reflections or small floating patches of Ceratophyllum demersum means that the model needs to be refined using a larger number of images, with the background of the images themselves becoming particularly important. The MA CBAM-YOLOv4 algorithm [7] was proposed to meet the demand for a real-time drowning risk detection system. In this work, a MA_CBAM module is added to a YOLOv4 model [8]. The MA_CBAM module is a Convolutional Block Attention Module (CBAM) [9], where the activation function of the channel attention module has been changed to the Meta-ACON activation function. Experimental results show that the improved model performs well with higher accuracy and robustness compared with the original YOLOv4. The YOLOv5s model [10] was selected from among various options in the proposal by [11]. They reported an average speed of 75 fps (frames per second), with a mean average precision (mAP) value exceeding 89%. Experiments are conducted on the Ubuntu18.04 operating system and the NVIDIA GeForce GTX 3060 Ti GPU with 12 GB. A total of 7000 images of infant swimming were collected as a dataset. In the work by Amer et al. [12], a YOLOv11 model is used. They capture a custom dataset to represent different real-world scenarios. The whole system was optimise to work in real-time on a Raspberry Pi 5 (8 GB RAM). These studies are the most similar to our approach, offering solutions akin to those we have adopted (use of the YOLO model and creation of specific databases). In our case, to meet processing speed requirements, we ultimately opted for the YOLOv11 version, which can be embedded on the NVIDIA Jetson Orin Nano Super. As it was not possible to create the database using images, we proposed the use of generative AI.
As can be deduced from the works cited, one of the problems that arises when addressing the issue of detecting possible drownings is the lack of publicly available databases, which makes it necessary to create one, either by capturing images or by selecting images available on the internet. The work by Bai et al. [13] proposes a framework for creating databases for the early detection of drowning that can be adapted to different scenarios. The idea is not to use the traditional method of image generation using AI, but rather to optimise the generation engine by optimising the framework structures of Generative Adversarial Networks (GANs), Variational Encoders (VAEs), and Diffusion Models.
3. Methodology
The research aim is to design and develop an efficient Deep Learning-based solution to detect the presence of people or animals that have fallen into irrigation reservoirs. The solution is designed to achieve high accuracy. Processing speed is not considered a key factor, as processing one image per second is generally deemed more than sufficient. This allows the system to comfortably process images from multiple cameras and handle relatively large image files. Furthermore, these processing speeds mean that the device responsible for running these inference models consumes less power. As aforementioned, the research aims to generate a synthetic dataset—as there are no databases available—design the proposed model and deploy it in the NVIDIA Jetson Orin Nano Super, and finally, evaluate the performance of the proposal in a real environment. The description of the dataset generation and the training of the chosen model are detailed below.
3.1. Synthetic Dataset Generation Using Diffusion Models
As stated in Section 1, one of the major challenges in developing detection models for people and animals falling into irrigation reservoirs is the scarcity of real-world training data. Capturing authentic images of drowning events raises obvious ethical and safety concerns, and publicly available datasets for this specific scenario are virtually nonexistent. To address this limitation, we employed state-of-the-art open-source diffusion models to generate a synthetic dataset, using a real photograph of the target irrigation reservoir as the background image (Figure 2). All generation workflows were orchestrated using ComfyUI (https://github.com/comfy-org/ComfyUI/) (Accessed on 11 june 2026), a node-based visual interface for designing and executing diffusion model pipelines. The output resolution was set to pixels, matching the panoramic aspect ratio of the 180-degree cameras deployed in the final system.
Figure 2.
The target irrigation pond.
3.1.1. Background on Diffusion Models
The following background introduces the generative models employed in our dataset generation pipeline, with emphasis on the specific architectural properties that motivated their selection and the parameter choices described in Section 3.1.6. Diffusion models [14] have emerged as the dominant paradigm for high-fidelity image generation. These models learn to reverse a gradual noising process, progressively denoising a sample drawn from a Gaussian distribution into a coherent image. Denoising Diffusion Probabilistic Models (DDPMs) formalize this as a Markov chain of T steps, where in the forward process Gaussian noise is incrementally added to the data according to a variance schedule :
and a neural network is trained to predict the noise component at each step, enabling the reverse (generative) process. This formulation directly governs two key controllable parameters used throughout our pipeline: the denoising strength , which determines the starting step in the reverse chain when processing an already-composed image (lower s preserves the original layout; higher s allows more photorealistic correction), and the number of sampling steps, which controls the granularity of the denoising trajectory. Stages 3 and 4 of our pipeline (see Figure 4, Section 3.1.3) use and respectively, with 12 and 20 sampling steps. The empirical justification for these values is provided in Section 3.1.6. Figure 3 illustrates the forward and reverse processes: the forward process gradually corrupts an image with Gaussian noise until it becomes indistinguishable from pure noise, while the learned reverse process iteratively denoises samples to recover coherent images.
Figure 3.
Forward and reverse processes in Denoising Diffusion Probabilistic Models. The forward process progressively adds Gaussian noise to an image ; the reverse process learns to denoise through a neural network .
Latent Diffusion Models (LDMs) [15] substantially reduce the computational cost of this procedure by operating in the latent space of a pretrained variational autoencoder (VAE) rather than directly on pixel space. An encoder maps the input image to a lower-dimensional latent representation , the diffusion process is carried out in this compressed space, and a decoder reconstructs the final image. This architecture underpins Stable Diffusion and its successive variants, including SDXL [16], which introduced a dual-encoder text conditioning scheme and a refinement stage for enhanced image quality.
More recently, flow-matching models [17] have advanced the state of the art by replacing the traditional denoising formulation with learned vector fields that transport samples along optimal paths from noise to data. FLUX [18], developed by Black Forest Labs, adopts this formulation using a multimodal DiT (Diffusion Transformer) architecture that processes text and image tokens jointly, yielding superior text-image alignment and visual coherence compared with UNet-based predecessors.
Z-Image [19] and its variants are 6B+ parameter image generation foundation models developed by Alibaba. It adopts a Scalable Single-Stream Diffusion Transformer (S3-DiT) architecture, in which text tokens, visual semantic tokens, and image VAE tokens are concatenated at the sequence level into a unified input stream, maximising parameter efficiency compared with dual-stream approaches such as FLUX.
3.1.2. Auxiliary Models
In addition to the core generative models, our pipeline leverages several auxiliary models for conditioning, analysis, and refinement:
- Florence2 [20]: a lightweight vision-language foundation model by Microsoft (0.23 B/0.77 B parameters) that formulates vision tasks as sequence-to-sequence problems using a DaViT encoder and an encoder-decoder transformer. In our pipeline, we employ Florence2 for automatic scene captioning, leveraging its Detailed Caption mode to generate rich natural language descriptions of composed images.
- DWPose [21]: a whole-body pose estimation method that uses two-stage knowledge distillation to produce accurate keypoint predictions—body, hands, face, and feet—while remaining lightweight, and generates superior pose skeletons compared with OpenPose for conditioning diffusion models. In our pipeline, the DWPose preprocessor node in ComfyUI extracts skeletal keypoint representations from source images, providing the structural conditioning signal for pose-guided person generation.
- YOLO + SAM [22]: a combined pipeline using a YOLO detector to localize target regions and the Segment Anything Model (SAM) to produce precise segmentation masks, enabling targeted inpainting and detail enhancement.
3.1.3. Synthetic Image Generation Pipeline
All synthetic images are produced through a unified four-stage pipeline. Regardless of whether the subject is a person or an animal, and independently of the conditioning strategy used for generation, every image traverses the same sequence of stages: subject generation, scene composition, global refinement, and subject detailing. Figure 4 provides an overview of the pipeline.
Figure 4.
Overview of the four-stage synthetic image generation pipeline. Stage 1 supports three conditioning strategies: (a) wildcard-based generation for random people and animals, (b) pose-guided generation, and (c) pose-guided generation with wildcard augmentation.
- 1.
- Subject generation. An image of the target subject (person or animal) is generated in isolation. Figure 5 shows representative examples of images produced by the three configurations. Three conditioning strategies are supported:Figure 5. Representative examples of synthetic images generated by the three pipeline configurations for the first stage. (a) Person (wildcard). (b) Person (pose + wildcard). (c) Wild boar (wildcard).
- Wildcard-based generation. A textual description of the subject is created using wildcard prompts—parameterised templates that randomly sample attributes such as gender, age range, clothing type, and body build (for people) or species, size, and posture (for animals). This stochastic mechanism ensures broad diversity across the generated population. The resulting prompt is used as conditioning input for Z-Image-Turbo [19], which generates a full-body image of the subject in isolation (see Figure 5a for random people and Figure 5c for wild boars).
- Pose-guided generation. A source image depicting a person in a relevant pose—manually selected from publicly available datasets showing swimming, water rescue, and distress situations—is processed by DWPose, which extracts a skeletal representation of body keypoints encoding the spatial configuration of limbs, torso, and head. Simultaneously, Florence2 generates a textual description of the source person’s appearance and posture. Using the extracted pose skeleton as structural conditioning and the generated description as semantic conditioning, Z-Image-Turbo produces a new person image that faithfully preserves the original body pose and appearance attributes such as clothing, skin tone, and body build. This strategy is particularly relevant for generating realistic distress-related body configurations (e.g., arms raised above the surface, head tilted back, partial submersion, or struggling postures) that are critical for training the detection model.
- Pose-guided generation with wildcard augmentation. This hybrid strategy combines pose conditioning with stochastic prompt variation to maximise diversity while retaining realistic body configurations. As in the previous strategy, DWPose extracts the skeletal keypoints from a source image and Florence2 produces a textual description of the depicted person. However, instead of using this description directly, it is modified by replacing selected appearance attributes—such as gender, age, clothing, and body build—with wildcard tokens that are randomly sampled at generation time. The resulting augmented prompt, together with the original pose skeleton, conditions Z-Image-Turbo to produce person images that preserve the source posture while exhibiting substantially greater variability in visual appearance than pure pose-guided generation (see Figure 5b).
- 2.
- Scene composition. The generated subject is inserted into the background photograph of the irrigation pond. The insertion position—modelling different scenarios such as being at the centre of the pond, near the edge, or partially submerged—is selected through wildcard-based spatial descriptors. The composition is performed by Flux.2 Klein 9B, a 9-billion-parameter flow-matching model whose joint text-image transformer architecture enables coherent integration of the subject into the scene, handling lighting consistency, water reflections, and partial occlusions.
- 3.
- Global refinement. The composed image is first processed by Florence2, which generates a detailed natural language description of the scene. This caption is then used as a conditioning signal for an image-to-image refinement pass with Z-Image-Turbo, configured with a denoising strength of 0.4 and 12 sampling steps, which harmonises colour temperature, lighting direction, and textural details between the subject and the background, attenuating artefacts introduced during the composition stage.
- 4.
- Subject detailing. A YOLO + SAM pipeline detects and segments the subject region in the refined image. The segmented area is then enhanced through an inpainting pass with Z-Image-Turbo, configured with a denoising strength of 0.5 and 20 sampling steps, which improves fine-grained details such as facial features, hand geometry, and clothing textures (for people) or fur patterns and body contours (for animals), ensuring that the subject region provides a reliable training signal for the downstream detection model.
This pipeline was instantiated in six configurations to produce the synthetic dataset: (i) wildcard-based generation of random people (401 curated images), (ii) pose-guided generation of people in distress-related postures (126 curated images), (iii) wildcard-based generation of wild boars (567 curated images), (iv) wildcard-based generation of ducks (181 curated images), (v) wildcard-based generation of turtles (155 curated images). Wild boars represent large terrestrial mammals, ducks are medium-sized waterfowl, and turtles are aquatic reptiles, collectively covering diverse sizes and morphologies relevant to irrigation pond environments. Additionally, 74 background-only images were generated independently: these depict the pond environment with random elements—such as floating debris, reflections, and ripples—but without any person or animal, and are annotated with empty label files. Their inclusion helps the model learn to suppress false positives caused by visually ambiguous reservoir features.
It is important to note that, in addition to single-object images, the generation pipeline also produces compositions with multiple subjects within a single frame. These multi-object images model realistic scenarios where, for example, several ducks or turtles may be swimming together, or both people and animals may be present in proximity within the camera’s field of view. Such scenes increase the model’s robustness to crowded or complex scenarios. The detection framework assigns a bounding box and class label to each detected object independently, enabling the model to recognize and localize multiple persons and/or animals within the same image.
3.1.4. Annotation and Curation
All generated images depicting a person or animal were manually annotated with bounding boxes in YOLO format using the open-source annotation platform Roboflow. Roboflow was selected for this task due to its intuitive interface for object detection workflows, built-in support for YOLO format exports, and capability to manage class definitions and track annotation progress. The platform facilitates collaborative annotation and provides quality assurance features that help maintain consistency across the dataset. Each annotation consists of a class label (person, boar, duck, or turtle) and the corresponding normalised bounding box coordinates . The 74 background-only images were annotated with empty label files, signalling the absence of any target object. Following annotation, a manual curation step was performed to filter out images exhibiting visual artefacts, anatomical anomalies, or poor integration with the background (see Section 3.1.6). Only those images that realistically represented people or animals within the pond environment—or plausible background scenes free of false targets—were retained in the final dataset.
The curated dataset was subsequently split into three subsets: 70% for training, 20% for validation, and 10% for testing. The split was performed in a stratified manner, preserving the class distribution across all subsets. Table 1 details the resulting partition.
Table 1.
Dataset split into training, validation, and test subsets.
3.1.5. Dataset Composition
Table 2 summarises the composition of the synthetic dataset. The wildcard-based person configuration produced 401 images, the pose-guided person configuration contributed 126 images, the wild boar configuration generated 567 images, the duck configuration generated 181 images, the turtle configuration generated 155 images, and background-only images with random pond elements were included to reduce false positives, yielding a curated base dataset of 1051 synthetic images. Standard data augmentation techniques—random horizontal flipping, colour jittering, and scaling—were applied exclusively to the training subset, tripling its effective size. Table 2 shows the base image counts and the augmented training set sizes.
Table 2.
Composition of the synthetic dataset. Data augmentation () was applied only to the training subset.
The wildcard-based random generation maximises diversity in subject appearance and spatial placement, the pose-guided configuration ensures that realistic distress-related body configurations are adequately represented, and the animal configuration extends coverage to the most relevant wildlife species. The use of wildcard-based parameterisation at multiple stages—subject description, spatial positioning, and scene layout—further augments variability, effectively acting as a data augmentation strategy embedded within the generation process itself. The validation and test subsets remain unaugmented, providing an unbiased estimate of model performance.
In addition to the static image dataset, 7 short synthetic videos of 5 s each were generated using the same pipeline, covering people, wild boars, ducks, and turtles in various spatial configurations. From each video, 100 frames were uniformly sampled, yielding 700 additional frames in total. These video-derived frames provide temporal diversity not captured by independently generated still images, exposing the model to motion blur, intermediate poses, and subject trajectories representative of real dynamic events.
The complete dataset is publicly available at https://huggingface.co/datasets/grupo-avispa/wildlife_in_irrigation_ponds (accessed on 11 June 2026) [23].
3.1.6. Implementation Details and Reproducibility
To facilitate reproducibility of the synthetic dataset generation pipeline, the key implementation choices are documented below.
Prompt engineering. Wildcard prompts were implemented as hierarchical template files, each encoding a set of attribute categories. For person generation, templates were randomly and uniformly sampled from: gender (male/female), age range (adult, elderly, teenager), clothing type (categories including casual, workwear, and swimwear), and body build (skinny, fit, fat). All categorical choices were equiprobable (uniform distribution), ensuring unbiased exploration of the attribute space. For wild boar generation, templates additionally covered size variants (adult/juvenile), fur color (grey, brown, black, yellowish), the presence of tusks and morphological variation, and posture descriptors (standing, moving, partially submerged).
Parameter selection rationale. The denoising strengths for the refinement stage () and the detailing stage () were selected by visual inspection over a grid search in . Lower values preserved the global composition but left residual compositing artefacts; higher values improved photorealism but occasionally altered subject pose or placement. Sampling step counts (12 for refinement, 20 for detailing) were identified as the minimum values yielding visually consistent outputs on a held-out pilot set of 20 images.
Generation failure rates. Prior to the manual curation step, a total of 1230 candidate images were produced across all pipeline configurations. Of these, approximately 5.04% were discarded due to anatomical anomalies (e.g., malformed limbs), poor subject–background integration (e.g., visible compositing seams), or subject disappearance during refinement.
Computational resources. All generation workflows were executed on a workstation equipped with an Nvidia RTX 5090 GPU with 32 GB of VRAM. The average end-to-end generation time per image across all four pipeline stages was approximately 40.82 s, yielding a total dataset generation time of approximately 836.81 GPU-hours for the full curated dataset.
3.2. Neural Network Training
The detection model was trained using YOLOv11 [24], the latest iteration of the YOLO (You Only Look Once) family of single-stage object detectors. YOLOv11 introduces architectural improvements over its predecessors, including an enhanced C3k2 feature extraction backbone, the SPPF (Spatial Pyramid Pooling–Fast) module, and a refined C2PSA (Cross Stage Partial with Spatial Attention) neck, resulting in improved accuracy and efficiency. Its anchor-free detection head simplifies the training pipeline and improves generalisation across object scales.
The model was configured to detect four classes: person, boar, duck, and turtle. Training was conducted on the augmented training subset, with the validation subset used for hyperparameter tuning and early stopping. The main training hyperparameters are summarised in Table 3.
Table 3.
YOLOv11 training hyperparameters.
Transfer learning was employed by initialising the model with weights pretrained on the COCO dataset, allowing the network to leverage generic feature representations and converge faster on the relatively small synthetic dataset. The training was monitored using the validation mAP@0.5 and mAP@0.5:0.95 metrics, with early stopping applied if no improvement was observed for 128 consecutive epochs.
Figure 6 shows the evolution of the training and validation losses, as well as the mAP metrics over the training epochs.
Figure 6.
Training and validation loss curves (left) and mAP@0.5/mAP@0.5:0.95 evolution (right) during YOLOv11 training.
Table 4 reports the final detection performance on the validation subset (299 images), which was not used during training or hyperparameter tuning.
Table 4.
Detection performance of the trained YOLOv11 model on the synthetic validation set.
3.3. Synthetic-to-Real Domain Gap
A fundamental challenge in synthetic-data-driven training is the domain gap between the synthetic training distribution and the real deployment distribution. Several design decisions in our pipeline directly target this gap. Using a real photograph of the target pond as the inpainting background ensures that water texture, geomembrane appearance, and surrounding vegetation are drawn from the actual deployment environment, avoiding the abstract scene statistics typical of fully rendered or copy–paste synthetic approaches. The multi-stage refinement and detailing passes further harmonise lighting, colour temperature, and subject textures, reducing visible compositing boundaries. The stochastic wildcard-based prompting introduces broad variability in subject appearance, clothing, and posture, limiting the risk of the model overfitting to a narrow synthetic appearance cluster.
Despite these measures, residual domain gaps remain. The training background was captured under a single set of lighting and seasonal conditions specific to southern Spain. Water turbidity, algae content, and seasonal colour variation are not captured by a static background image. Illumination distributions at high latitudes (low-angle sunlight, extended overcast periods) and adverse weather conditions (heavy rain, strong wind-induced surface rippling) are absent from the training set. These residual gaps are partially reflected in the performance difference between synthetic validation (mAP@0.5 of 0.9123) and real-world field trials (overall detection rate of 90%), indicating a non-negligible domain shift even in the matched-environment person detection scenario.
The evaluation of wild boar detection relied on video footage captured at different pond and lake environments, introducing a more severe distribution shift than for person detection (which was validated at the actual deployment pond). This explains the lower detection rate observed for wild boars (84% vs. 91% for persons) and highlights background diversity as a key factor for cross-environment generalisation—a limitation that directly motivates the dataset expansion directions discussed in Section 5.
4. Experimental Results
To evaluate the practical viability of the proposed system, the trained YOLOv11 model was deployed and tested in a real-world irrigation pond environment. This section describes the deployment setup, the inference platform, and the results obtained during field trials. As mentioned in this section, it has not been possible to capture actual images of wild boars in this scenario (no wild boars were observed falling during the system evaluation period, nor was it possible to induce them to fall). Consequently, the system’s ability to detect them has been evaluated using images extracted from real-world footage, captured at various irrigation ponds or natural lakes.
4.1. Deployment Setup
The experimental deployment was conducted on a real irrigation reservoir of approximately 1500 m2 located in southern Spain. The monitoring system consists of two Reolink panoramic cameras, each providing a 180-degree horizontal field of view (FoV) with colour night vision capability. The two cameras were mounted at opposite edges of the reservoir perimeter, ensuring full visual coverage of the water surface and its immediate surroundings. Figure 7 shows a schematic diagram of the placement of cameras along the reservoir boundary (Left), as well as a real snapshot illustrating the characteristics of the experimental setting (Right).
Figure 7.
Deployment of the monitoring system on the real irrigation pond: two Reolink 180° panoramic cameras provide full coverage of the pond surface.
The system is powered by a 500 W solar panel and includes a 100 Ah lithium iron phosphate battery (12.8 V) and a regulator. The cameras are connected to a router via Ethernet, with this connection also used to power them (Power over Ethernet, PoE).
It is interesting to note that, even when confined to a single setting, this can pose a challenge due to its outdoor location and the dynamics of the water’s surface layer. Thus, when the water is calm, and depending on the time of day, the surface of the pond may reflect everything around it (trees, houses, vehicles…), display different colours (this also depends on the season), or be affected by vegetation clinging to the bottom that sways in the wind, or by fish swimming beneath it. The sun also creates reflections, which can make it difficult to detect an object. When a person or animal falls into the pond, the resulting movement causes the water’s surface to look completely different, and its overall appearance will vary greatly depending on whether the images are captured during the day or at night, or on the camera’s position relative to the sun. Figure 8 shows several illustrative examples. All of these factors have been taken into account in the data collected in the database created for this study.
Figure 8.
Examples illustrating the complexity that the water surface adds to the detection problem.
4.2. Inference Platform
Inference was performed on an NVIDIA Jetson Orin Nano Super, a compact edge-computing platform equipped with a 1024-core Ampere GPU with 32 Tensor Cores and 8 GB of shared memory. The trained YOLOv11 model was exported to ONNX format and optimised using TensorRT to maximise inference throughput on the Jetson’s GPU. Table 5 summarises the hardware specifications of the inference platform.
Table 5.
NVIDIA Jetson Orin Nano Super specifications.
The video streams from both Reolink cameras are received via RTSP and processed sequentially by the Jetson Orin Nano Super. Each frame is resized to the model’s input resolution ( pixels) before inference. Post-processing includes non-maximum suppression (NMS) with a confidence threshold of 0.7 and an IoU threshold of 0.5.
4.3. Real-World Detection Results
The system was evaluated through a series of controlled trials conducted under varying conditions, including different times of day (daytime and nighttime), weather conditions (clear, overcast), and subject positions (centre, edge, partially submerged). For the evaluation of person detection, real volunteers were introduced into the reservoir described in Figure 7. For the evaluation of animal detection, sequences were captured in a natural setting, a pond of 32,375 m2 approx. (Figure 9). These sequences allow testing the detection of turtles and ducks. In this situation, it was not possible to install the system required for long-term monitoring of the pond. Finally, real-world video footage of wild boars was used to generate test sequences. This approach was necessary due to the infeasibility of directly introducing wild boars into the reservoir or pond and the impracticality of waiting for animals to appear naturally in the pond environment (Figure 10). Table 6 summarises the detection performance observed during the field trials.
Figure 9.
The natural pond used to validate the detection of animals (turtles and wild ducks).
Figure 10.
Representative detection results from the deployed system under daytime condition. (a) Person detection. (b) Animal detection.
Table 6.
Detection performance during real-world field trials on the irrigation reservoir (for the Person class), or from analysing test sequences (see text) (for the Wild board class).
It is important to note that these results are derived from the processing of a single image. Given that the system works with video sequences, and assuming a detection window of three consecutive frames, the probability that the target will be correctly detected in at least one of these three frames would be [25]
p being the probability of detecting the target in one image. The equation assumes that the three detection processes are independent, as is the case in our system. The value of the cumulative probability of detecting the fall into the reservoir is 99.7%.
4.4. Inference Performance
Table 7 reports the inference latency and throughput measured on the NVIDIA Jetson Orin Nano Super. The results show that the system could process a high number of frames per second. This is not necessary in our case, as the processing frequency can be set to, for example, once a minute (which would result in 120 frames per hour being processed from two cameras). This frame rate must be set to prevent the system from draining the battery when it is not being charged by the solar panel. At this frame rate, the system’s total hourly power consumption is just over 10 Wh.
Table 7.
Inference performance on the NVIDIA Jetson Orin Nano Super.
4.5. Power Consumption
Deploying the system in environments where a standard power supply may not be available necessitates the design of a solution with relatively low power consumption. As described throughout the text, the power source used has been a 500 W solar panel and a 100 Ah lithium iron phosphate battery (12.8 V).
The total power consumption of the physical node is around 23 W, broken down as follows: (1) the average power consumption of 14 W from the two Reolink Duo 2 cameras (each consumes around 5 W during the day and 9 W at night, with the spotlights active); (2) the average power consumption of 5 W by the NVIDIA Jetson Orin Nano (running at 2 fps, it is set to low-power mode); and (3) an average power consumption of 4 W associated with losses in the PoE switch and router. Daily consumption is therefore around 552 Wh.
Lithium iron phosphate batteries can be deeply discharged to as much as 90% without compromising their service life. With a usable capacity (90% DoD) of 1152 Wh, the system’s runtime without the solar panel is around 50 h. Whether or not the solution is suitable will depend on the latitude of the region of the world in which it is located and on the average cloud cover, with some regions receiving high levels of annual sunshine (such as Arizona in the US or Calama in Chile, with over 3900 h of sunshine a year) and others receiving much less (the British Isles receive around 1200–1600 h of sunshine a year). Based on estimated losses of 15% to 20%, Table 8 details the average daily generation in kWh for various cities on sunny days. In most situations, the 500 W panel is capable of generating the system’s daily energy consumption on an average summer or winter day. This ensures that the battery will be fully recharged, providing enough reserve power to cope with up to two consecutive days of total darkness without the system shutting down.
Table 8.
Average daily electricity generation in kWh in different cities for a 500 W solar panel (see text for details).
5. Conclusions and Future Work
This paper presents a complete edge-computing system for detecting people and animals falling into irrigation ponds, combining synthetic dataset generation, deep learning-based object detection, and real-world deployment. The main contribution is the development of an automated monitoring solution that addresses a critical safety issue in agricultural water management infrastructure. The proposed methodology overcomes the fundamental challenge of dataset scarcity by leveraging state-of-the-art diffusion models to generate a synthetic training dataset. The YOLOv11 model, trained on this synthetic data, demonstrated strong performance metrics: precision of 0.912, recall of 0.838, and mAP@0.5 of 0.912 on the synthetic validation set.
Real-world field trials conducted on a 1500 m2 irrigation pond confirmed the practical viability of the approach. The system achieved a detection rate of 91% for persons. Inference latency of 16.434 ms per frame was achieved on the NVIDIA Jetson Orin Nano Super, well within operational requirements. The compact model size of 75.68 MB enables efficient edge deployment with minimal power consumption (less than 6 Wh for our reduced frame rate). To properly evaluate the detection of wild boars, ducks, and turtles, it was necessary to assess real sequences captured in environments other than the one used to create the database. This may have led to a variation in detection rates across animal species (e.g., 84% for wild boars). This provides an overall detection rate of 90.3% across 290 trials.
Although the proposed approach demonstrates promising results, several limitations and opportunities for improvement remain. The system’s performance on nighttime sequences requires further evaluation under varied low-illumination conditions. False positives caused by water reflections and floating debris observed in initial trials could be mitigated through additional negative samples in the training set. The current dataset covers a specific geographical region (southern Spain) and weather conditions; expanding the dataset to include diverse environmental scenarios, seasons, and edge cases would improve robustness and generalisation. It is also important to diversify the setting used to build the database, generalising it to cover different appearances in irrigation ponds or natural lakes. The colour of the water can also vary depending on factors such as the season, its clarity or its depth. All these factors must be taken into account when expanding the dataset. From a geographical generalisation standpoint, the real-world validation was conducted exclusively at a single pond in southern Spain; deployment in environments with substantially different background characteristics—such as ponds in northern Norway, which typically feature dark peat-stained water, dense coniferous vegetation, lower-angle sunlight, and predominantly overcast skies—would expose the system to background distributions not present in the training set, likely degrading detection performance. Building a geographically diverse set of background images and integrating them into the synthetic generation pipeline is therefore a necessary step towards a robust, universally applicable system. Regarding animal species coverage, the current model targets wild boars, ducks, and turtles. While wild boars are the dominant large wild animal causing incidents in the Mediterranean region, the inclusion of smaller aquatic species (ducks) and slow-moving reptiles (turtles) extends the system’s capability to detect a broader range of animals. However, other geographies involve substantially different fauna (e.g., deer and foxes in central Europe, coyotes and raccoons in North America, or kangaroos in Australia); extending the dataset with species-specific synthetic data would be required before deploying the system in those contexts. Finally, no ablation study was conducted to quantify the individual contribution of each pipeline component—wildcard prompting, pose guidance, multi-stage refinement, subject detailing, and video-derived augmentation—to the final detection performance. Such a study would identify which stages provide the greatest benefit and inform future optimisation of the generation pipeline. Thus, future work should focus on: (1) expanding the synthetic dataset to include more diverse environmental conditions, water states, and animal species relevant to different geographical regions; (2) integrating multi-camera fusion strategies to reduce false positives and improve spatial accuracy in larger pond environments; (3) incorporating temporal information through video-based models to leverage motion cues for improved detection consistency; (4) developing adaptive confidence thresholds based on contextual information (time of day, weather, pond water state); and (5) conducting a systematic ablation study to assess the contribution of each component of the synthetic generation pipeline to downstream detection performance.
Author Contributions
Conceptualization, A.T., C.A.R.-B., Ó.P. and A.B.; Methodology, A.T., C.A.R.-B., Ó.P. and A.B.; Software, A.T. and Ó.P.; Validation, A.T. and Ó.P.; Formal analysis, A.T., Ó.P. and A.B.; Investigation, A.T., Ó.P. and A.B.; Resources, A.T., C.A.R.-B. and Ó.P.; Data curation, A.T. and Ó.P.; Writing—original draft, A.T. and A.B.; Writing—review & editing, A.T. and A.B.; Visualization, A.T. and A.B.; Supervision, M.G.-G. and A.B.; Project administration, M.G.-G. and A.B.; Funding acquisition, M.G.-G. and A.B. All authors have read and agreed to the published version of the manuscript.
Funding
This work has been supported by grants CPP2021-008931 and PID2022-137344OB-C32, funded by MCIN/AEI/10.13039/501100011033 and by the European Union NextGenerationEU/PRTR (for the first one), and “ERDF A way of making Europe” (for the second one).
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The synthetic dataset used in this study is publicly available on Hugging Face at https://huggingface.co/datasets/grupo-avispa/wildlife_in_irrigation_ponds (accessed on 11 June 2026) [23].
Acknowledgments
This project is being carried out in collaboration with the company ATARFIL SL. The company was not involved in the study design, collection, analysis, interpretation of data, the writing of this article or the decision to submit it for publication.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- López-Felices, B.; Aznar-Sánchez, J.A.; Velasco-Muñoz, J.F.; Piquer-Rodríguez, M. Contribution of Irrigation Ponds to the Sustainability of Agriculture. A Review of Worldwide Research. Sustainability 2020, 12, 5425. [Google Scholar] [CrossRef] [Scilit]
- Mattoussi, W.; Mattoussi, F.; Zeddini, Y. Does dam-based irrigation affect the sustainability of natural capital?: A doubly robust analysis. J. Clean. Prod. 2024, 450, 141764. [Google Scholar] [CrossRef] [Scilit]
- Martin Cazorla, F.; Sánchez Blanque, J.L.; Santos Amaya, I.M. Muertes en balsas de riego en la provincia de Almería. Rev. EspañOla Med. Leg. 2011, 37, 134–139. [Google Scholar] [CrossRef] [Scilit]
- Jegham, N.; Koh, C.Y.; Abdelatti, M.; Hendawi, A. Evaluating the Evolution of YOLO (You Only Look Once) Models: A Comprehensive Benchmark Study of YOLO11 and Its Predecessors. arXiv 2024, arXiv:2411.00201. [Google Scholar] [CrossRef] [Scilit]
- Alotaibi, A. Automated and Intelligent System for Monitoring Swimming Pool Safety Based on the IoT and Transfer Learning. Electronics 2020, 9, 2082. [Google Scholar] [CrossRef] [Scilit]
- Shatnawi, M.; Albreiki, F.; Alkhoori, A.; Alhebshi, M. Deep Learning and Vision-Based Early Drowning Detection. Information 2023, 14, 52. [Google Scholar] [CrossRef] [Scilit]
- Niu, Q.; Wang, Y.; Yuan, S.; Li, K.; Wang, X. An Indoor Pool Drowning Risk Detection Method Based on Improved YOLOv4. In Proceedings of the 2022 IEEE 5th Advanced Information Management, Communicates, Electronic and Automation Control Conference (IMCEC), Chongqing, China, 16–18 December 2022; Volume 5, pp. 1559–1563. [Google Scholar] [CrossRef] [Scilit]
- Bochkovskiy, A.; Wang, C.; Liao, H.M. YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv 2020, arXiv:2004.10934. Available online: http://arxiv.org/abs/2004.10934 (accessed on 29 April 2026).
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the Computer Vision–ECCV 2018; Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y., Eds.; Springer: Cham, Switzerland, 2018; pp. 3–19. [Google Scholar]
- Hussain, M. YOLOv1 to v8: Unveiling Each Variant–A Comprehensive Review of YOLO. IEEE Access 2024, 12, 42816–42833. [Google Scholar] [CrossRef] [Scilit]
- He, Q.; Mei, Z.; Zhang, H.; Xu, X. Automatic Real-Time Detection of Infant Drowning Using YOLOv5 and Faster R-CNN Models Based on Video Surveillance. J. Soc. Comput. 2023, 4, 62–73. [Google Scholar] [CrossRef] [Scilit]
- Amer, D.; Ibrahim, N.; Ibrahim, I.; Mohamed, A.; Soliman, S. Intelligent eyes on water: YOLOv11-based real-time drowning detection system. J. Supercomput. 2025, 81, 1242. [Google Scholar] [CrossRef] [Scilit]
- Bai, B.; Yue, H.; Chen, L.; Li, X. Research on Dataset Generation and Monitoring of Generative AI for Drowning Warning System. IEEE Access 2024, 12, 83589–83599. [Google Scholar] [CrossRef] [Scilit]
- Ho, J.; Jain, A.; Abbeel, P. Denoising Diffusion Probabilistic Models. In Proceedings of the Advances in Neural Information Processing Systems; Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 6840–6851. [Google Scholar]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv 2021, arXiv:2112.10752. Available online: http://arxiv.org/abs/2112.10752 (accessed on 29 April 2026).
- Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; Rombach, R. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. In Proceedings of the International Conference on Learning Representations; Kim, B., Yue, Y., Chaudhuri, S., Fragkiadaki, K., Khan, M., Sun, Y., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 2024, pp. 1862–1874. [Google Scholar]
- Lipman, Y.; Chen, R.T.Q.; Ben-Hamu, H.; Nickel, M.; Le, M. Flow Matching for Generative Modeling. arXiv 2023, arXiv:2210.02747. Available online: http://arxiv.org/abs/2210.02747 (accessed on 29 April 2026).
- Labs, B.F. FLUX. 2024. Available online: https://github.com/black-forest-labs/flux (accessed on 11 June 2026).
- Z-Image Team. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer. arXiv 2025, arXiv:2511.22699. [Google Scholar]
- Xiao, B.; Wu, H.; Xu, W.; Dai, X.; Hu, H.; Lu, Y.; Zeng, M.; Liu, C.; Yuan, L. Florence-2: Advancing a unified representation for a variety of vision tasks. arXiv 2023, arXiv:2311.06242. [Google Scholar]
- Yang, Z.; Zeng, A.; Yuan, C.; Li, Y. Effective Whole-body Pose Estimation with Two-stages Distillation. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Paris, France, 2–6 October 2023; pp. 4212–4222. [Google Scholar] [CrossRef] [Scilit]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment Anything. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 3992–4003. [Google Scholar] [CrossRef] [Scilit]
- Tudela, A.; Ruiz-Beltrán, C.A.; Pons, Ó.; González-García, M.; Bandera, A. Wildlife in Irrigation Ponds Dataset. 2025. Available online: https://huggingface.co/datasets/grupo-avispa/wildlife_in_irrigation_ponds (accessed on 11 June 2026).
- Jocher, G.; Qiu, J. Ultralytics YOLO11 Version 11.0.0, 2024. Available online: https://github.com/ultralytics/ultralytics (accessed on 11 June 2026).
- Hall, D.; Dayoub, F.; Skinner, J.; Zhang, H.; Miller, D.; Corke, P.; Carneiro, G.; Angelova, A.; Sünderhauf, N. Probabilistic Object Detection: Definition and Evaluation. arXiv 2020, arXiv:1811.10800. Available online: http://arxiv.org/abs/1811.10800 (accessed on 29 April 2026).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.









