Search arXiv⌕ Search

arXiv · 1308.2430

Nondetection sampling bias in marked presence-only data

Abstract

1. Species distribution models (SDM) are tools used to determine environmental features that influence the geographic distribution of species' abundance and have been used to analyze presence-only records. Analysis of presence-only records may require correction for nondetection sampling bias to yield reliable conclusions. In addition, individuals of some species of animals may be highly aggregated and standard SDMs ignore environmental features that may influence aggregation behavior. 2. We contend that nondetection sampling bias can be treated as missing data. Statistical theory and corrective methods are well developed for missing data, but have been ignored in the literature on SDMs. We developed a marked inhomogeneous Poisson point process model that accounted for nondetection and aggregation behavior in animals and tested our methods on simulated data. 3. Correcting for nondetection sampling bias requires estimates of the probability of detection which must be obtained from auxiliary data, as presence-only data do not contain information about the detection mechanism. Weighted likelihood methods can be used to correct for nondetection if estimates of the probability of detection are available. We used an inhomogeneous Poisson point process model to model group abundance, a zero-truncated generalized linear model to model group size, and combined these two models to describe the distribution of abundance. Our methods performed well on simulated data when nondetection was accounted for and poorly when detection was ignored. 4. We recommend researchers consider the effects of nondetection sampling bias when modeling species distributions using presence-only data. If information about the detection process is available, we recommend researchers explore the effects of nondetection and, when warranted, correct the bias using our methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Trevor Hefley, Andrew Tyre, David Baasch, Erin Blankenship. 2013-12-04. Nondetection sampling bias in marked presence-only data. https://doi.org/10.1002/ece3.887

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Graph construction in QUBO-based recursive phylogenetic tree reconstruction

Molecular sequence data are used to reconstruct evolutionary relationships among taxa, but reconstruction accuracy depends not only on the tree-building method but also on how pairwise sequence relationships are represented. We evaluated sequence-to-affinity representations in a recursive normalized-cut (Ncut) framework whose graph-partitioning subproblems were formulated as quadratic unconstrained binary optimization (QUBO) models and solved using Simulated Bifurcation. Using simulated amino-acid and nucleotide datasets spanning multiple tree-generation settings and evolutionary divergence, we compared normalized bit-score affinities with representations derived from transformed sequence similarities and evolutionary distances, examined post-swap refinement, and used neighbor joining (NJ) as a distance-based comparator. Affinity representation substantially affected internal split recovery, particularly for nucleotide data. JC69-based local affinities maintained comparatively high accuracy as divergence increased, whereas normalized bit-score and BLAST-derived kernel representations declined more markedly. Post-swap refinement generally improved recovery, but not consistently across individual reconstructions. NJ achieved higher mean split recovery than corresponding recursive Ncut reconstructions for WAG and JC69 distances across all evaluated conditions, whereas recursive Ncut outperformed NJ for BLAST-derived logarithmic distances under some conditions. These results show that graph construction is an important determinant of recursive Ncut-based phylogenetic reconstruction. A representation that performs well within Ncut does not necessarily provide the most accurate use of the underlying pairwise distances. Pairwise representation, affinity transformation, optimization, and recursive tree construction should therefore be evaluated jointly.

q-bio.PE↗

A conceptual predator-prey model with super-long transients

Drawing on the understanding of the logistic map, we propose a simple predator-prey model where predators and prey adapt to each other, leading to the co-evolution of the system. The special dynamics observed in periodic windows contribute to the coexistence of multiple time scales, adding to the complexity of the system. Typical dynamics in ecosystems, such as the persistence and coexistence of population cycles and chaotic behaviors, the emergence of super-long transients, regime shifts, and the quantifying of resilience, are encapsulated within this single model. The simplicity of our model allows for detailed analysis, reinforcing its potential as a conceptual tool for understanding ecosystems deeply.

q-bio.PE↗

Mutation Order and Selection Shape Intratumor Heterogeneity in Tumor Evolution

Cancer progression often requires multiple driver mutations, but the same drivers may be acquired in different orders. How these pathways jointly shape tumor clonal structure remains unclear. We develop a multitype branching-process model in which malignant transformation requires two driver mutations, distinguishing malignant cells by mutation order and the independent transformation event that founded their clone. Under a successive exponential approximation, we establish point-process limits for pathway-specific clone sizes and derive a closed-form expression for the limiting expected Simpson's index of the combined malignant population. When both mutation orders yield malignant cells with the same net growth rate, the index decomposes into effective pathway weights, determined by mutation rates and birth-death dynamics at preceding stages, and within-pathway concentration terms, determined by intermediate-to-malignant growth-rate ratios. A driver's effect on heterogeneity thus depends critically on when it is acquired. A strong driver acquired early expands the intermediate lineage and increases the supply of independent malignant founders, whereas the same driver acquired last strengthens the growth and age advantage of early-founded malignant clones. Under additive fitness effects, these opposing mechanisms can produce a non-monotone relationship between selective advantage and clonal concentration. Threshold-like non-additive fitness effects can generate highly concentrated malignant populations, while order-dependent terminal fitness causes the faster-growing pathway to dominate asymptotically. These results show how mutation order, mutational accessibility, selection, and epistasis jointly determine lineage-level intratumor heterogeneity.

q-bio.PE↗