1. Introduction
Global biodiversity is facing unprecedented challenges due to anthropogenic pressures, habitat fragmentation, and climate change [
1]. Wetlands are among the most biodiverse and ecologically sensitive ecosystems on Earth, playing a vital role in supporting waterbird populations [
2]. Birds, in particular, serve as excellent ecological indicators for monitoring ecosystem health due to their sensitivity to environmental alterations, wide distribution, and relative ease of observation compared to other taxa [
3]. In the Mediterranean region, protected wetlands are critical stopover points for migratory routes and breeding grounds for numerous threatened species. Specifically, the wetlands in the southern Valencian Community of Spain (Parc Natural d’El Hondo, Salinas de Santa Pola, and Lagunas de La Mata–Torrevieja) represent critical habitats designated as Important Bird Areas (IBAs) and protected under the Ramsar Convention [
4,
5].
However, monitoring avian populations in these ecosystems is historically challenging. Traditional monitoring relies on professional field censuses conducted by ornithologists [
6,
7]. While these surveys are highly standardized and provide reliable population counts, they are resource-intensive, expensive, and limited in spatiotemporal resolution, often conducted only once a month or seasonally. This sparsity can fail to capture rapid shifts in bird distribution or migration timings.
To supplement traditional surveys, two main paradigms have emerged: citizen science and autonomous monitoring. Citizen science platforms, most notably eBird [
8,
9], have revolutionized ecological research by aggregating millions of bird sightings submitted by recreational birdwatchers worldwide. These platforms offer unprecedented spatial and temporal coverage [
10]. Nevertheless, citizen science data suffer from significant limitations, including spatial and temporal clustering, varying observer expertise, and a lack of standardized effort (e.g., presence-only checklists) [
11]. These limitations are part of a broader, well-documented literature on bias and information content in biological records [
12] and on the consequences of imperfect and heterogeneous detection for species distribution and occupancy modeling [
13,
14], which motivates treating our target as an observation-process outcome rather than as a direct estimate of occurrence or detection probability. Conversely, autonomous monitoring technologies, such as Internet of Things (IoT) acoustic sensors deployed in the wild running deep-learning classifiers like BirdNET [
15], offer continuous, non-intrusive, and high-frequency recording of bird vocalizations [
16]. Yet, these sensors are localized to single points, produce large volumes of detections that do not correspond directly to abundance, and are vulnerable to physical hazards (such as sensor failures due to seasonal flooding in wetland environments). These sources therefore represent different observation processes: a census count records individuals during a standardized survey, an eBird record depends on observer effort and reporting, and a BirdNET detection records an acoustic classifier output that may include repeated vocalizations from one individual. Their numerical harmonisation does not make them biologically equivalent.
To address these limitations, researchers must combine these heterogeneous data sources to generate robust spatiotemporal occurrence models [
17]. Fusing citizen science, acoustic sensor networks, and professional censuses requires overcoming major differences in count methodologies, spatial resolution, and temporal coverage. In particular, professional censuses record absolute counts, eBird checklists register mixed presence-only and count data, and acoustic sensors record detection counts of vocalizations. Fusing them at the count level is impossible; thus, a standardized representation is required.
In this work, we present a comprehensive framework for multi-source bird observation modeling and interactive visualization in Valencian wetlands. Our working hypothesis is that adding denser but heterogeneous observations can improve prediction of the held-out observation proportions by supplying temporal and spatial context, while not necessarily correcting source-specific sampling or detection bias. The framework is therefore a predictive baseline rather than an occupancy model or a direct estimator of population status. The key contributions of this paper are as follows:
ValWet-Birds Dataset: We construct and release a harmonized multi-source spatiotemporal dataset that taxonomically and spatially unifies citizen science (eBird), autonomous acoustics (BirdNET from Raspberry Pi devices), and professional censuses across three protected wetlands, resolving them into a single, cohesive resource spanning from 2010 to 2025 across 12 sub-regions and 394 taxa (349 resolved to species level,
Section 2.4).
Spatiotemporal Deep Learning Baseline: We present and evaluate two baseline deep learning architectures, a Feed-Forward MLP and an LSTM recurrent network, trained to predict the relative observation proportion of bird species.
Rigorous Data Augmentation Evaluation: We first report a preliminary comparison (differently sized test sets) showing a 42.7% MSE reduction under augmentation, and then re-evaluate every source condition on an identical held-out set of professional-census observations. The aggregate benefit does not survive this stricter comparison, but a 50-taxon subset cross-referenced against Annex I of the EU Birds Directive does show a substantial, consistent benefit from augmentation (29% MSE reduction for the LSTM). We report both results and the reversal between them as a methodological contribution.
Interactive Web Visualizer: We have created Avistory, an open-source, interactive web visualization system that enables researchers and the public to intuitively query, filter, and visualize historical observations alongside our deep learning model predictions.
The remainder of this paper is structured as follows.
Section 2 presents the study area, data sources, harmonisation pipeline, and relative observation proportion normalization.
Section 3 outlines the MLP and LSTM baseline models, sequence generation, and training protocol.
Section 4 reports the model performance and the quantitative benefits of data augmentation.
Section 5 describes the design and user flows of the Avistory visualization platform. Finally,
Section 6 and
Section 7 discuss the limitations, future lines of research, and conclusions.
2. Materials and Methods
2.1. Study Area
The study area encompasses three protected wetlands located in the province of Alicante, within the Valencian Community in southeastern Spain. These areas are designated as Important Bird Areas (IBAs) and are characterized by diverse hydrological and ecological regimes:
Parque Natural El Hondo (IBA BIRDLIFE_1824): Composed of multiple reservoirs and marshlands. It contains both freshwater reservoirs (e.g., Embalse de Levante and Embalse de Poniente) and saline lagoons, supporting a rich biodiversity of waterfowl.
Salinas de Santa Pola (IBA BIRDLIFE_1825): A coastal saltpan wetland characterized by high salinity, including active salt-extraction basins and littoral zones.
Lagunas de La Mata y Torrevieja (IBA BIRDLIFE_1826): Composed of two large hypersaline lagoons. The Laguna de Torrevieja is highly hypersaline, supporting minimal avian life (mostly restricted to nesting areas on salt dikes), while the Laguna de La Mata is slightly less saline, supporting larger populations of flamingos and grebes.
To conduct fine-grained spatiotemporal modeling, these three wetlands were partitioned into 12 survey units (sub-regions), corresponding to specific water bodies or management sectors (
Table 1,
Figure 1).
2.2. Data Sources
The dataset is constructed by harmonising and integrating three independent, heterogeneous observation sources:
Citizen Science: Observation records retrieved from the eBird Dataset (ES-VC regional release, November 2025) [
8]. Sightings were filtered spatially using bounding boxes matching the three target IBA codes, covering the period from January 2010 to December 2025.
Professional Census: Standardized monthly bird counts conducted by professional ornithologists using exhaustive visual and auditory counts, following a protocol comparable in spirit to the International Waterbird Census [
18] but conducted monthly rather than as a single coordinated mid-winter count. It covers the year 2025 for all three wetlands, and 2024 for Santa Pola. Only the 2024–2025 records were digitized and shared with us by the regional management authority in a format compatible with this dataset; an extended historical archive, if one exists, was not accessible to us for this study, and we do not assume otherwise. Survey duration, the number of observers per count, exact survey timing, and the procedures used to avoid double-counting were likewise not provided to us and are not reconstructed here; we report their absence as a limitation of the available metadata rather than an assumption about the census protocol. Water level, salinity, weather, observer effort, checklist duration, and completeness were not available to us as measured covariates for any of the three sources and are not included in the models presented here. These surveys provide high-quality absolute abundance data but are limited to monthly intervals.
Autonomous Acoustics: Passive acoustic monitoring data were collected by an IoT sensor node based on a Raspberry Pi 5 (Raspberry Pi Ltd., Cambridge, UK) deployed at El Hondo. The sensor continuously recorded ambient audio and identified bird vocalizations from 24 September 2024 to 10 May 2025 using the pre-trained BirdNET convolutional neural network classifier [
15]. Detections were filtered to keep only those with a confidence score
.
The resulting ValWet-Birds dataset contains 67,787 observation rows, covering 190 unique year-months (from January 2010 to December 2025) across 394 taxa (
Table 2; see
Section 2.4 for the taxonomic audit that resolved this count from an initial 395).
2.3. Spatial Harmonisation
Integrating these sources required resolving differing spatial identifiers. Professional censuses record observations at the level of local micro-zones and management units (e.g., Eb.Levante, CALENTADORES, Punta Víbora). eBird records contain user-defined text descriptions of hotspots (e.g., El Hondo PNat–Embalse de Levante) or arbitrary coordinates for personal checklists. BirdNET files do not record specific sub-regions on the device.
Spatial harmonisation was completed in two steps:
Keyword and Lookup Mapping: A manually curated mapping table was built to associate known eBird hotspots and professional census micro-zones directly with the corresponding 12 unified sub-regions. For the BirdNET sensor, since it was statically placed, all records were mapped to the El Hondo–La Reserva sub-region.
Nearest-Centroid Fallback: eBird checklists not associated with a predefined hotspot instead carried individual GPS coordinates. For these records, we calculated the geographic distance (Haversine) to the centroids of all 12 sub-regions. If a coordinate fell within a 500-m buffer of the nearest centroid, it was automatically assigned to that sub-region; checklists beyond this buffer from every centroid were excluded.
The resulting 12 sub-regions are treated as survey units selected to balance geographical granularity with observation density, ensuring that every unit has either significant census coverage or ≥800 eBird checklist entries. They should not be interpreted as homogeneous ecological units, because habitat composition, water depth, salinity, and vegetation may vary within and among them.
Spatial-assignment counts. We reconstructed this procedure over the 13,975 eBird checklists falling within the three target IBAs, using the same keyword-mapping table and 500 m centroid buffer described above: 10,684 checklists (76.5%) matched a known hotspot directly (Step 1); of the remaining 3291 checklists carrying individual coordinates, 2192 (15.7% of the total; 66.6% of the coordinate-only subset) were assigned via the nearest-centroid fallback (Step 2), and 1099 (7.9% of the total; 33.4% of the coordinate-only subset) fell outside the 500 m buffer of every centroid and were excluded. In total, 12,876 checklists (92.1%) were retained. The reconstruction reported here was run directly against the archived eBird sampling file using a mapping table and buffer threshold.
2.4. Taxonomic Harmonisation
Avian nomenclature varies across sources. The professional census uses local scientific names directly. eBird records provide scientific and common Spanish/English names. BirdNET outputs English common names with underscores (e.g., Barn_Swallow, Cettis_Warbler), which occasionally conflict with eBird naming conventions.
To address discrepancies in species naming across data sources, all recorded taxa were standardized to their corresponding scientific names. This harmonization step unified the nomenclature across sensor detections, citizen science records, and professional surveys, yielding a consolidated vocabulary of 394 taxa (349 resolved to species level), following the taxonomic audit described below.
Taxonomic audit. We audited every distinct scientific-name label in the dataset. This identified one clear nomenclature error: 177 census records (24 year-months, including 31 inside the common evaluation window in
Section 4) were labelled
Phoenicopterus ruber (American Flamingo), a species absent from this region. Our eBird-derived taxonomy lookup table already lists
Phoenicopterus roseus (Greater Flamingo) as the correct name for this taxon, so we treat the Census label as a legacy synonym error and merge these records into
Phoenicopterus roseus, matching the label already used by eBird and BirdNET. The audit also found 45 taxa (796 eBird rows, 1.2% of the dataset) recorded only to genus or family level or as slash-notation pairs (e.g.,
Sturnus vulgaris/unicolor,
Anatidae sp.), consistent with eBird’s own conventions for uncertain field identifications; none originate from Census, so they do not affect the common test set. We do not reassign these to a single candidate species, since that would imply an identification the observer did not make; instead, we exclude them from species-level ranking (as detailed in
Section 2.6) while retaining them in the dataset and in the denominator of Equation (1). The corrected vocabulary contains 394 scientific-name labels: 349 resolved species and 45 unresolved.
Legal-nomenclature synonyms. Directive 2009/147/EC retains the 1979-era scientific names of its predecessor (79/409/EEC), several of which have since been superseded by modern (IOC/Clements) nomenclature. Seven taxa in our Annex I cross-reference (
Table 3) are listed under such a superseded name in the directive text; all seven are the same taxon under current avian systematics, unrelated to the
Phoenicopterus ruber/
roseus correction above.
2.5. Unified Relative Probability Target
Fusing abundance data directly from these sources is impossible because their recording protocols are incompatible. eBird checklists represent varying search durations and observer counts, censuses represent absolute visual counts, and BirdNET represents the frequency of vocal detections.
To address this issue, we normalized counts into a relative monthly observation proportion (
P) within each grouping of source (
s), sub-region (
r), year (
y), and month (
m):
where
is the count of species
within the grouping, and
is the set of all species recorded in that grouping. By construction,
for each group.
This formulation allows joint modeling of citizen science, sensors, and professional censuses without mixing incompatible raw units. The resulting value is a source-conditional proportion of recorded detections or counts, not a formal occurrence or detection probability: it depends on the focal species and on all other taxa recorded in the same grouping. Source-specific observation and detection processes are not explicitly modeled, so values from different sources are comparable only as inputs to this predictive baseline and not as equivalent biological measurements. The resulting distribution is right-skewed, reflecting the dominance of a small number of taxa in the recorded data, resulting in a median observation proportion of 0.4% and a 99th percentile of 39.6%.
2.6. Dataset Balance
Throughout this subsection, “records” or “observations” refers to rows of the harmonised dataset, i.e., distinct species × sub-region × year × month × source combinations after aggregation (
Section 2.5), not to individual raw eBird checklists or field observations submitted before harmonisation.
Figure 2 shows how observations are distributed across the 12 sub-regions. El Hondo concentrates 58% of all records, driven mainly by Zona Rincón (19.7%) and La Reserva (17.4%), which are the most accessible areas and the most frequently visited eBird hotspots. Santa Pola accounts for 30% of the dataset, while La Mata–Torrevieja contributes the remaining 11%. The smallest sub-region, Laguna Torrevieja, holds only 921 harmonised records (1.4%; 697 from eBird and 224 from Census), consistent with its hypersaline conditions and limited avian activity.
Figure 3 illustrates species concentration by total number of individuals recorded, after the taxonomic audit described above. The top 10 resolved species account for 53.9% of all individuals across the dataset. The Greater Flamingo (
Phoenicopterus roseus) alone represents 20.1%, with just over 1,039,000 individuals once the corrected census records are included, followed by Common Starling (
Sturnus vulgaris) at 10.2%. A further 5.0% of individuals belong to the 45 genus/family-level or slash-notation taxa that are not assigned to a single resolved species and are therefore excluded from the ranked list. The remaining 339 species share 41.1% of the total count. This heavy concentration is characteristic of wetland avifaunal communities and directly motivates the use of the relative observation proportion normalization described above.
3. Baseline Deep Learning Models
We implement and compare two spatiotemporal neural network baseline models to predict monthly relative observation proportions: a Feed-Forward Multi-Layer Perceptron (MLP) and a recurrent Long Short-Term Memory (LSTM) network. Both models adopt a single-instance formulation where each input instance corresponds to a single tuple and yields a scalar prediction. Consequently, spatial predictions across the entire wetland area are generated by executing 12 independent inference passes (one per sub-region) for any target species and timestep.
Table 4 summarizes the notation used throughout this section to support the mathematical description of the two architectures.
3.1. Feature Representation and Embeddings
To enable effective deep learning, input features are represented using a combination of categorical embeddings and continuous feature normalization:
3.2. Multi-Layer Perceptron Architecture
The MLP baseline operates as a static regressor, modeling occurrences for a given species and sub-region at a specific time step independently. The input vector
concatenates the continuous features and categorical embeddings:
The model consists of three fully connected hidden layers with dimensions 256, 128, and 64, respectively. To ensure stability during training and prevent overfitting, each linear transformation is sequentially followed by Batch Normalization, a Rectified Linear Unit (ReLU) activation, and Dropout (with a rate of
). The final scalar output
, representing the predicted relative observation proportion percentage, is computed as:
where
is the representation extracted by the last hidden layer, and
are the parameters of the output layer.
3.3. Long Short-Term Memory Architecture
The LSTM baseline is formulated to capture sequential temporal patterns. For each (sub-region, species) pair, the model processes a sequence of historical observations to predict the relative observation proportion in the subsequent month.
A continuous monthly timeline is built for each species and sub-region. Missing months are zero-filled with their corresponding cyclic temporal features. Sliding windows of sequence length
months are constructed. Each timestep
is represented by a 4D feature vector:
where
is the historical relative observation proportion at month
t.
Static metadata is incorporated by concatenating the learned species and sub-region embeddings into a context vector . This vector is linearly projected to initialize the initial hidden state , while the cell state is initialized to zero.
The temporal sequence
is processed through a single-layer LSTM with 64 hidden units and recurrent dropout (rate
). Finally, the last hidden state
is concatenated with the static context
and passed through a dense layer with ReLU activation and Dropout (
) to yield the output prediction:
where
denotes the predicted relative observation proportion for month
.
3.4. Training Protocol and Implementation Details
Both architectures are implemented in PyTorch (version 2.2.0) and trained on a single NVIDIA GeForce RTX 4090 GPU. Models are optimized to minimize the Mean Squared Error (MSE) directly on the target relative observation proportion percentage scale (0–100%).
Optimization is performed using the Adam optimizer with an initial learning rate of , an weight decay penalty of , and a batch size of 256 for the MLP and 128 for the LSTM. To facilitate stable convergence, a Cosine Annealing learning rate scheduler dynamically decays the learning rate over a maximum budget of 200 epochs. To mitigate overfitting, early stopping is enforced by monitoring validation loss, terminating training if no improvement is observed for 20 consecutive epochs.
7. Conclusions
This study introduced ValWet-Birds, a unified spatiotemporal dataset and data fusion framework that harmonizes professional censuses, citizen science records (eBird), and autonomous passive acoustic monitoring (BirdNET) to model avian observation patterns across three protected wetlands in the Valencian Community, Spain. By defining a unified monthly relative observation proportion target, the methodology places heterogeneous records on a common numerical scale, but it does not make their observation processes ecologically equivalent.
A preliminary, non-matched-test comparison suggested a 42.7% lower Test Mean Squared Error for the recurrent LSTM under multi-source augmentation. Re-evaluating every source condition on an identical set of held-out professional-census observations shows that this aggregate benefit does not hold: augmentation does not reduce error relative to census-only training when all taxa are considered together; the original comparison conflated data augmentation with a change in test-set composition. However, for a 50-taxon subset cross-referenced against Annex I of the EU Birds Directive, augmentation does reduce error substantially (29% for the LSTM, 9% for the MLP), plausibly because these less-common taxa have sparser census history for the model to draw on. Both MLP and LSTM baselines clearly outperform naive temporal reference predictors on the same common test, which supports the value of the modeling approach independent of the augmentation question. None of these results confirm recovery of seasonal migration dynamics or subtle shifts in species distribution.
To support inspection of predictive modeling outputs, we deployed Avistory, an open-source web platform that translates model inferences into spatial visualization tools for researchers and other users. The platform is not presented as a validated conservation decision-support system: observed patterns and model-derived proportions can be affected by sampling effort and detectability and should not be read as population trends without ecological validation. Future developments will focus on effort-aware observation models, occupancy or detection-probability modeling, expanded sensing infrastructure with weather-hardened edge nodes, and incorporation of visual computer vision models.