Search arXivSearch

arXiv subjects

Eleni Matechou

Publications and source records attributed to Eleni Matechou.

6 recordsLinked to original sources

False Positives, False Negatives, and the Detection-Only Problem: A Hierarchical Model for Species Occurrence with Observation Error

Monitoring species occurrence is essential for understanding biodiversity change, informing conservation decisions, and assessing the impact of environmental pressures on ecosystems. Species occurrence data arise from different survey designs, and the statistical literature has developed distinct corresponding modelling approaches, namely occupancy models, species distribution models, and presence-only methods, whose fundamental connections have remained largely unrecognised. We argue that these are all special cases of a single hierarchical observation process. To make these connections explicit, we introduce a unified terminology centred on two data types: detection/non-detection data with T visits (DN-T) and detection-only data (DO), where DN-T with T>1 corresponds to traditional occupancy modelling, DN-1 to species distribution modelling, and DO to what the literature commonly, but we argue inaccurately, calls presence-only data. Within this framework, we study the identifiability of DO models and propose a novel hierarchical model for DO data that, for the first time, explicitly accounts for both false positive and false negative detection errors. Identifiability is achieved through prior distributions that express the natural belief that a species is more likely to be recorded where it is present than where it is absent. ...

stat.AP

A capture-recapture hidden Markov model framework for register-based inference of population size and dynamics

Accurate inference on population dynamics, such as migration and changes in population size, is essential for policymaking, resource allocation and demographic research. Traditional censuses are expensive, infrequent and not timely, leading many countries to adopt register-based approaches to replace or complement them. A primary challenge is that such registers are incomplete: even when individuals are present, their activities may not generate records in specific registers, resulting in false negative observation error. Conversely, some registers arise from administrative or household-level processes, so that individuals may appear in registers despite being absent, leading to false positive observation error. Existing approaches often either rely on ad-hoc decisions that ignore one or both error types, offer inference on population snapshots but not dynamics, or are computationally too slow for practical use. We propose a scalable framework for inferring population size and dynamics from register data, building on Cormack-Jolly-Seber type capture-recapture models formulated as hidden Markov models. Inference is carried out using maximum likelihood estimation, with uncertainty quantified via the Bag of Little Bootstraps. The model accounts for temporary emigration, incorporates an arbitrary number of possibly interacting registers subject to both error types, and allows observation probabilities to vary with individual characteristics and unobservable heterogeneity. We illustrate the approach using Swedish population registers, where overcoverage - individuals registered as living in the country although they are no longer present - provides a motivating example. The application yields new insights into population dynamics and individual trajectories.

stat.AP

Hidden Markov models with an unknown number of states and a repulsive prior on the state parameters

Hidden Markov models (HMMs) offer a robust and efficient framework for analyzing time series data, modelling both the underlying latent state progression over time and the observation process, conditional on the latent state. However, a critical challenge lies in determining the appropriate number of underlying states, often unknown in practice. In this paper, we employ a Bayesian framework, treating the number of states as a random variable and employing reversible jump Markov chain Monte Carlo to sample from the posterior distributions of all parameters, including the number of states. Additionally, we introduce repulsive priors for the state parameters in HMMs, and hence avoid overfitting issues and promote parsimonious models with dissimilar state components. We perform an extensive simulation study comparing performance of models with independent and repulsive prior distributions on the state parameters, and demonstrate our proposed framework on two ecological case studies: GPS tracking data on muskox in Antarctica and acoustic data on Cape gannets in South Africa. Our results highlight how our framework effectively explores the model space, defined by models with different latent state dimensions, while leading to latent states that are distinguished better and hence are more interpretable, enabling better understanding of complex dynamic systems.

stat.AP

eDNAPlus: A unifying modelling framework for DNA-based biodiversity monitoring

DNA-based biodiversity surveys involve collecting physical samples from survey sites and assaying the contents in the laboratory to detect species via their diagnostic DNA sequences. DNA-based surveys are increasingly being adopted for biodiversity monitoring. The most commonly employed method is metabarcoding, which combines PCR with high-throughput DNA sequencing to amplify and then read `DNA barcode' sequences. This process generates count data indicating the number of times each DNA barcode was read. However, DNA-based data are noisy and error-prone, with several sources of variation. In this paper, we present a unifying modelling framework for DNA-based data allowing for all key sources of variation and error in the data-generating process. The model can estimate within-species biomass changes across sites and link those changes to environmental covariates, while accounting for species and sites correlation. Inference is performed using MCMC, where we employ Gibbs or Metropolis-Hastings updates with Laplace approximations. We also implement a re-parameterisation scheme, appropriate for crossed-effects models, leading to improved mixing, and an adaptive approach for updating latent variables, reducing computation time. We discuss study design and present theoretical and simulation results to guide decisions on replication at different stages and on the use of quality control methods. We demonstrate the new framework on a dataset of Malaise-trap samples. We quantify the effects of elevation and distance-to-road on each species, infer species correlations, and produce maps identifying areas of high biodiversity, which can be used to rank areas by conservation value. We estimate the level of noise between sites and within sample replicates, and the probabilities of error at the PCR stage, which are close to zero for most species considered, validating the employed laboratory processing.

stat.AP

Fast Bayesian inference for large occupancy data sets, using the Polya-Gamma scheme

In recent years, the study of species' occurrence has benefited from the increased availability of large-scale citizen-science data. Whilst abundance data from standardized monitoring schemes are biased towards well-studied taxa and locations, opportunistic data are available for many taxonomic groups, from a large number of locations and across long timescales. Hence, these data provide opportunities to measure species' changes in occurrence, particularly through the use of occupancy models, which account for imperfect detection. However, existing Bayesian occupancy models are extremely slow when applied to large citizen-science data sets. In this paper, we propose a novel framework for fast Bayesian inference in occupancy models that account for both spatial and temporal autocorrelation. We express the occupancy and detection processes within a logistic regression framework, which enables us to use the Polya-Gamma scheme to perform inference quickly and efficiently, even for very large data sets. Spatial and temporal random effects are modelled using Gaussian processes, allowing us to infer the strength of spatio-temporal autocorrelation from the data. We apply our model to data on two UK butterfly species, one common and widespread and one rare, using records from the Butterflies for the New Millennium database, producing occupancy indices spanning 45 years. Our framework can be applied to a wide range of taxa, providing measures of variation in species' occurrence, which are used to assess biodiversity change.

stat.AP

A general modelling framework for open wildlife populations based on the Polya Tree prior

Wildlife monitoring for open populations can be performed using a number of different survey methods. Each survey method gives rise to a type of data and, in the last five decades, a large number of associated statistical models have been developed for analysing these data. Although these models have been parameterised and fitted using different approaches, they have all been designed to model the pattern with which individuals enter and exit the population and to estimate the population size. However, existing approaches rely on a predefined model structure and complexity, either by assuming that parameters are specific to sampling occasions, or by employing parametric curves. Instead, we propose a novel Bayesian nonparametric framework for modelling entry and exit patterns based on the Polya Tree (PT) prior for densities. Our Bayesian non-parametric approach avoids overfitting when inferring entry and exit patterns while simultaneously allowing more flexibility than is possible using parametric curves. We apply our new framework to capture-recapture, count and ring-recovery data and we introduce the replicated PT prior for defining classes of models for these data. Additionally, we define the Hierarchical Logistic PT prior for jointly modelling related data and we consider the Optional PT prior for modelling long time series of data. We demonstrate our new approach using five different case studies on birds, amphibians and insects.

stat.ME