Search arXivSearch

arXiv subjects

Michail Papathomas

Publications and source records attributed to Michail Papathomas.

12 recordsLinked to original sources

Exploring complex dependence structures using Bayesian bi-clustering and log-linear graphical modelling

Bayesian partitioning is utilised simultaneously on subjects and categorical variables to reveal complex dependence structures. Clusters of variables are referred to as views. Variable selection highlights the variables that drive the clustering of the subjects within each view. We derive theoretical results on the relation between the variables' dependence structure and the inferences derived from bi-clustering. The results relate to marginal independence and conditional independence. Using simulated and real data, we demonstrate the applicability of bi-clustering results in assisting log-linear graphical model determination, leading to the efficient exploration of typically vast model spaces. This work sheds light on the relation between two very different but equally popular Bayesian approaches; mixture modelling, that benefits from a large number of variables, and graphical log-linear modelling, which describes explicitly the variables' dependence structure and allows to evaluate the uncertainty associated with model determination.

stat.ME

Spatially resolved star-formation histories of local post-starburst galaxies: Starburst and quenching spatial patterns consistent with recent mergers

Post-starburst (PSB) galaxies, having recently experienced a starburst followed by rapid quenching, are excellent laboratories to probe physical mechanisms that drive starbursts and shutting down of star formation. Integral-field spectroscopy reveals the galaxies' spatially-resolved properties, where observed directional patterns can be linked to the galaxies' past evolution. We measure the resolved star-formation histories (SFHs), stellar metallicity evolution and dust properties of three local PSBs from the MaNGA survey, down to $0.5$" resolution ($\sim0.3\,$kpc) using a hierarchical Bayesian model. Local parameters were constrained simultaneously with parameters describing spatial trends. We found that all three galaxies first experienced an outer, weaker and slower quenching starburst, followed by a central, stronger and faster quenching starburst that peaked $\sim 1\,$Gyr after the first. The central starbursts induced a significantly stronger rise in stellar metallicity compared to the outer starbursts. These results are consistent with the effects of a recent gas-rich (wet) merger, where the first pericentre passage triggered starbursts in the outer regions, while the later coalescence triggers a stronger centralised starburst. We find non-axisymmetric features in the maps of burst mass fraction and dust attenuation in all galaxies, which could be caused by tidal effects during the recent merger. Comparisons with literature binary merger simulations suggests that the galaxies' rapid quenching was driven by gas consumption and the stabilisation against gas gravitational collapse by a growing spheroid, while AGN feedback was not necessarily a primary cause.

astro-ph.GA

The diverse quenching pathways of post-starburst galaxies in SDSS-IV MaNGA

The quenching of star formation in galaxies is an important aspect of galaxy evolution, but the physical mechanisms that drive it are still not understood. Measuring the spatial distribution of quenching can help determine these mechanisms. We present the star-formation histories (SFHs) and stellar metallicity evolution of rapidly quenched regions in 86 local post-starburst (PSB) galaxies from the MaNGA integral field survey, obtained through Bayesian full spectral fitting of their rest-frame optical spectra. We found that regardless of spatial location, the PSB regions have similar past SFHs and chemical evolution, once radial metallicity gradients are accounted for. This suggests that all PSB regions are regulated by a common set of local scale processes in the interstellar medium, regardless of the broader triggering mechanism. We show that the centres of galaxies with outer PSB regions are also quenching. The central specific star-formation rate (sSFR) has declined by $\sim1.0\;$dex on average during the last 2 Gyr, a significantly steeper decline than main sequence galaxies over the same period ($\approx0.2\;$dex). This central quenching can be either synchronous, outside-in or inside-out, and slower or as fast as the outer regions, highlighting the diversity of quenching pathways for local galaxies. Our results imply a primary quenching mechanism that is both catastrophic and global in rapidly halting star formation in local galaxies. We suggest the predominant cause is galaxy mergers or interactions, with large scale feedback from a starburst or a central supermassive black hole playing a lesser role.

astro-ph.GA

Chemical evolution of local post-starburst galaxies: Implications for the mass-metallicity relation

We use the stellar fossil record to constrain the stellar metallicity evolution and star-formation histories of the post-starburst (PSB) regions within 45 local post-starburst galaxies from the MaNGA survey. The direct measurement of the regions' stellar metallicity evolution is achieved by a new two-step metallicity model that allows for stellar metallicity to change at the peak of the starburst. We also employ a Gaussian process noise model that accounts for correlated errors introduced by the observational data reduction or inaccuracies in the models. We find that a majority of PSB regions (69% at $>1\sigma$ significance) increased in stellar metallicity during the recent starburst, with an average increase of 0.8 dex and a standard deviation of 0.4 dex. A much smaller fraction of PSBs are found to have remained constant (22%) or declined in metallicity (9%, average decrease 0.4 dex, standard deviation 0.3 dex). The pre-burst metallicities of the PSB galaxies are in good agreement with the mass-metallicity relation of local star-forming galaxies. These results are consistent with hydrodynamic simulations, which suggest that mergers between gas-rich galaxies are the primary formation mechanism of local PSBs, and rapid metal recycling during the starburst outweighs the impact of dilution by any gas inflows. The final mass-weighted metallicities of the PSB galaxies are consistent with the mass-metallicity relation of local passive galaxies. Our results suggest that rapid quenching following a merger-driven starburst is entirely consistent with the observed gap between the stellar mass-metallicity relations of local star-forming and passive galaxies.

astro-ph.GA

Variance matrix priors for Dirichlet process mixture models with Gaussian kernels

The Dirichlet Process Mixture Model (DPMM) is a Bayesian non-parametric approach widely used for density estimation and clustering. In this manuscript, we study the choice of prior for the variance or precision matrix when Gaussian kernels are adopted. Typically, in the relevant literature, the assessment of mixture models is done by considering observations in a space of only a handful of dimensions. Instead, we are concerned with more realistic problems of higher dimensionality, in a space of up to 20 dimensions. We observe that the choice of prior is increasingly important as the dimensionality of the problem increases. After identifying certain undesirable properties of standard priors in problems of higher dimensionality, we review and implement possible alternative priors. The most promising priors are identified, as well as other factors that affect the convergence of MCMC samplers. Our results show that the choice of prior is critical for deriving reliable posterior inferences. This manuscript offers a thorough overview and comparative investigation into possible priors, with detailed guidelines for their implementation. Although our work focuses on the use of the DPMM in clustering, it is also applicable to density estimation.

stat.ME

Introducing a Real-time Interactive GUI Tool for Visualization of Galaxy Spectra

To aid the understanding of the non-linear relationship between galaxy properties and predicted spectral energy distributions (SED), we present a new interactive graphical user interface (GUI) tool pipes_vis based on Bagpipes \citep{arXiv:1712.04452,arXiv:1903.11082}. It allows for real-time manipulation of a model galaxy's star formation history, dust and other relevant properties through sliders and text boxes, with each change's effect on the predicted SED reflected instantaneously. We hope the tool will assist in building intuition about what affects the SED of galaxies, potentially helping to speed up fitting stages such as prior construction, and aid in undergraduate and graduate teaching. pipes_vis is available online (pipes_vis is maintained and documented online at https://github.com/HinLeung622/pipes_vis, or version 0.4.1 is archived in Zenodo and also available for installation through pip install pipes_vis).

astro-ph.GA

Parameter Redundancy and the Existence of Maximum Likelihood Estimates in Log-linear Models

Log-linear models are typically fitted to contingency table data to describe and identify the relationship between different categorical variables. However, the data may include observed zero cell entries. The presence of zero cell entries can have an adverse effect on the estimability of parameters, due to parameter redundancy. We describe a general approach for determining whether a given log-linear model is parameter redundant for a pattern of observed zeros in the table, prior to fitting the model to the data. We derive the estimable parameters or functions of parameters and also explain how to reduce the unidentifiable model to an identifiable one. Parameter redundant models have a flat ridge in their likelihood function. We further explain when this ridge imposes some additional parameter constraints on the model, which can lead to obtaining unique maximum likelihood estimates for parameters that otherwise would not have been estimable. In contrast to other frameworks, the proposed novel approach informs on those constraints, elucidating the model that is actually being fitted.

stat.ME

On the correspondence of deviances and maximum likelihood and interval estimates from log-linear to logistic regression modelling

Consider a set of categorical variables $\mathcal{P}$ where at least one, denoted by $Y$, is binary. The log-linear model that describes the counts in the resulting contingency table implies a specific logistic regression model, with the binary variable as the outcome. Extending results in Christensen (1997), by also considering the case where factors present in the contingency table disappear from the logistic regression model, we prove that the Maximum Likelihood Estimate (MLE) for the parameters of the logistic regression equals the MLE for the corresponding parameters of the log-linear model. We prove that, asymptotically, standard errors for the two sets of parameters are also equal. Subsequently, Wald confidence intervals are asymptotically equal. These results demonstrate the extent to which inferences from the log-linear framework can be translated to inferences within the logistic regression framework, on the magnitude of main effects and interactions. Finally, we prove that the deviance of the log-linear model is equal to the deviance of the corresponding logistic regression, provided that the latter is fitted to a dataset where no cell observations are merged when one or more factors in $\mathcal{P} \setminus \{ Y \}$ become obsolete. We illustrate the derived results with the analysis of a real dataset.

stat.ME

On synthetic interval data with predetermined subject partitioning, and partial control of the variables' marginal correlation structure

A standard approach for assessing the performance of partition models is to create synthetic data sets with a prespecified clustering structure, and assess how well the model reveals this structure. A common format is that subjects are assigned to different clusters, with observations simulated so that subjects within the same cluster have similar profiles, allowing for some variability. In this manuscript, we consider observations from interval variables, taking a finite number of values. Interval data are commonly observed in cohort and Genome Wide Association studies, and our focus is on Single Nucleotide Polymorphisms. Theoretical and empirical results are utilized to explore the dependence structure between the variables, in relation with the clustering structure for the subjects. A novel algorithm is proposed that allows to control the marginal stratified correlation structure of the variables, specifying exact correlation values within groups of variables. Practical examples are shown, and a synthetic dataset is compared to a real one, to demonstrate similarities and differences.

stat.ME

On the correspondence from Bayesian log-linear modelling to logistic regression modelling with $g$-priors

Consider a set of categorical variables where at least one of them is binary. The log-linear model that describes the counts in the resulting contingency table implies a specific logistic regression model, with the binary variable as the outcome. Within the Bayesian framework, the $g$-prior and mixtures of $g$-priors are commonly assigned to the parameters of a generalized linear model. We prove that assigning a $g$-prior (or a mixture of $g$-priors) to the parameters of a certain log-linear model designates a $g$-prior (or a mixture of $g$-priors) on the parameters of the corresponding logistic regression. By deriving an asymptotic result, and with numerical illustrations, we demonstrate that when a $g$-prior is adopted, this correspondence extends to the posterior distribution of the model parameters. Thus, it is valid to translate inferences from fitting a log-linear model to inferences within the logistic regression framework, with regard to the presence of main effects and interaction terms.

stat.ME

PReMiuM: An R Package for Profile Regression Mixture Models using Dirichlet Processes

PReMiuM is a recently developed R package for Bayesian clustering using a Dirichlet process mixture model. This model is an alternative to regression models, non-parametrically linking a response vector to covariate data through cluster membership. The package allows Bernoulli, Binomial, Poisson, Normal and categorical response, as well as Normal and discrete covariates. Additionally, predictions may be made for the response, and missing values for the covariates are handled. Several samplers and label switching moves are implemented along with diagnostic tools to assess convergence. A number of R functions for post-processing of the output are also provided. In addition to fitting mixtures, it may additionally be of interest to determine which covariates actively drive the mixture components. This is implemented in the package as variable selection.

stat.CO

Exploring dependence between categorical variables: benefits and limitations of using variable selection within Bayesian clustering in relation to log-linear modelling with interaction terms

This manuscript is concerned with relating two approaches that can be used to explore complex dependence structures between categorical variables, namely Bayesian partitioning of the covariate space incorporating a variable selection procedure that highlights the covariates that drive the clustering, and log-linear modelling with interaction terms. We derive theoretical results on this relation and discuss if they can be employed to assist log-linear model determination, demonstrating advantages and limitations with simulated and real data sets. The main advantage concerns sparse contingency tables. Inferences from clustering can potentially reduce the number of covariates considered and, subsequently, the number of competing log-linear models, making the exploration of the model space feasible. Variable selection within clustering can inform on marginal independence in general, thus allowing for a more efficient exploration of the log-linear model space. However, we show that the clustering structure is not informative on the existence of interactions in a consistent manner. This work is of interest to those who utilize log-linear models, as well as practitioners such as epidemiologists that use clustering models to reduce the dimensionality in the data and to reveal interesting patterns on how covariates combine.

stat.ME