Search arXiv⌕ Search

arXiv subjects

Carlos A. Martinez

Publications and source records attributed to Carlos A. Martinez.

3 recordsLinked to original sources

On the sensitivity to hyperparameters of a Bayesian method to infer selection signatures

Motivated by a real-life problem, namely, inferring genomic regions affected by selection, we studied some properties of a two-step method that performs this task based on the FST. This approach exhibits computational simplicity as an appealing feature. Step one involves fitting a conjugate Bayesian model to find point estimates of FST and selecting the markers to be considered in step two, where they are grouped using such estimates. Subjective priors are frequently used in this problem and large hyperparameters may be encountered; hence, there is interest in addressing their impact. Preliminary analyses suggested that, the number of markers selected in step one tends to zero when the hyperparameters have the same value and tend to infinity, or are different but tend to infinity and their difference tends to zero. Thus, we studied how hyperparameters affect the posterior mean, variance, and predictive density of allele frequencies, and the posterior distribution of FST, formally. Moreover, an empirical approach based on simulated and real data was performed. Our main findings are summarized in two theorems that set the asymptotic behavior of the method, providing a foundation to understand the aforementioned phenomena. As to the empirical studies, the groups created in the second step, the estimated posterior means and the number of selected markers showed high sensitivity to hyperparameters. Albeit the mathematical intractability of the posterior distribution of FST, our results revealed asymptotic properties explaining two features of the method that lacked a formal elucidation.

stat.ME↗

Modelling nuisance parameters in linear models: Formal and empirical comparisons between the two-way Anova and Ancova

The analysis of experimental or observational data often requires controlling for the effects of variables of secondary interest. Two common approaches to account for such parameters are grouping experimental or observational units into homogeneous groups (blocking), and observing quantitative variables on each one of them (covariance analysis). Often, both can be used and; therefore, comparing their performance poses a research question. Specifically, under randomized complete block designs where the blocks are defined by quantitative variables, there is interest in comparing the two-way Anova (TWA) with analysis of covariance (Ancova) models. Motivated by real-life problems coming from statistical consulting, we compared the performance of linear models considering block effects and those considering covariate effects from analytical and empirical perspectives. Models assumed Gaussian errors and included the effects of T treatments. The formal comparison relied on the theory of orthogonal projections and vector spaces under containment relationships of column spaces of the design matrices induced by conditions on the experimental design or data structure. We found asymptotic and small-sample conditions for the TWA to be at least as precise as the Ancova, as well as for these models to be equivalent. As to the empirical perspective, models were compared through a simulation study motivated by agronomical experiments. The results suggested that when the covariates show complex heterogeneity patterns, the Ancova outperforms the TWA with blocks defined by traditional approaches in terms of precision of location parameters, but not in hypothesis testing accuracy.

stat.ME↗