Search arXivSearch

arXiv · 2403.15118

Galaxy merger challenge: A comparison study between machine learning-based detection methods

Abstract

Various galaxy merger detection methods have been applied to diverse datasets. However, it is difficult to understand how they compare. We aim to benchmark the relative performance of machine learning (ML) merger detection methods. We explore six leading ML methods using three main datasets. The first one (the training data) consists of mock observations from the IllustrisTNG simulations and allows us to quantify the performance metrics of the detection methods. The second one consists of mock observations from the Horizon-AGN simulations, introduced to evaluate the performance of classifiers trained on different, but comparable data. The third one consists of real observations from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) survey. For the binary classification task (mergers vs. non-mergers), all methods perform reasonably well in the domain of the training data. At $0.1<z<0.3$, precision and recall range between $\sim$70\% and 80\%, both of which decrease with increasing $z$ as expected (by $\sim$5\% for precision and $\sim$10\% for recall at $0.76<z<1.0$). When transferred to a different domain, the precision of all classifiers is only slightly reduced, but the recall is significantly worse (by $\sim$20-40\% depending on the method). Zoobot offers the best overall performance in terms of precision and F1 score. When applied to real HSC observations, all methods agree well with visual labels of clear mergers but can differ by more than an order of magnitude in predicting the overall fraction of major mergers. For the multi-class classification task to distinguish pre-, post- and non-mergers, none of the methods offer a good performance, which could be partly due to limitations in resolution and depth of the data. With the advent of better quality data (e.g. JWST and Euclid), it is important to improve our ability to detect mergers and distinguish between merger stages.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

B. Margalef-Bentabol, L. Wang, A. La Marca, C. Blanco-Prieto, D. Chudy, H. Domínguez-Sánchez, A. D. Goulding, A. Guzmán-Ortega, M. Huertas-Company, G. Martin, W. J. Pearson, V. Rodriguez-Gomez, M. Walmsley, R. W. Bickley, C. Bottrell, C. Conselice, D. O'Ryan. 2024-04-15. Galaxy merger challenge: A comparison study between machine learning-based detection methods. https://doi.org/10.1051/0004-6361%2F202348239

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Spatially Resolved Physical Properties of Young Star Clusters and Star-forming Clumps in the Brightest z>6 Galaxy, the Strongly Lensed Cosmic Spear at z=6.2

We present spatially resolved analysis of stellar populations in the brightest $z>6$ galaxy known to date (AB mag 23), the strongly lensed MACS0308$-$zD1 (dubbed the ``Cosmic Spear'') at $z_{\rm spec}=6.2$. New JWST NIRCam imaging and high-resolution NIRSpec IFU spectroscopy span the rest-frame ultraviolet to optical. The NIRCam imaging reveals bright star-forming clumps and a tail consisting of three distinct, extremely compact star clusters that are multiply-imaged by gravitational lensing. The star clusters have delensed effective radii of $R_{\rm{eff}} \lesssim 8$ pc, stellar masses of $M_{*} \sim 10^{6}-10^{7}\,M_{\odot}$, and high stellar mass surface densities of $Σ_{*} \gtrsim 2\times 10^{4}\,M_{\odot}~\rm{pc}^{-2}$. While their stellar populations are very young ($\sim 6-11$ Myr), their dynamical ages exceed unity, consistent with the clusters being gravitationally bound systems. Placing the star clusters in the size vs.~stellar mass density plane, we find they occupy a region similar to other high-redshift star clusters within galaxies observed recently with JWST, being significantly more massive and denser than local star clusters. Spatially resolved analysis of the brightest clump reveals a compact, intensely star-forming core. The ionizing photon production efficiency ($ξ_{\rm{ion}}$) is slightly suppressed in this central region, potentially indicating a locally elevated Lyman continuum escape fraction facilitated by feedback-driven channels.

astro-ph.GA

Predicting Supermassive Black Hole-Host Mass Offsets from Broadband Photometry Across Cosmological Simulations with Forecasts for LSST

The possibility of over-massive black holes suggested by James Webb Space Telescope photometric discoveries of 'little red dots', may disfavor light supermassive black hole (SMBH) seeds. However, what should constitute the mass (range) of 'heavy' seeds remains relatively unconstrained. Moreover, Vera C Rubin Observatory's Legacy Survey of Space and Time will photometrically characterize galaxies without direct black hole mass measurements. We forward-model the SIMBA, IllustrisTNG, and EAGLE cosmological simulations into the photometric bands of LSST to train an ensemble machine learning classifier. Our framework achieves $91\%$--$94\%$ accuracy across SIMBA and IllustrisTNG in distinguishing between over-massive and under-massive SMBH growth regimes under LSST magnitude limits, using only broadband photometry. Furthermore, cross-simulation transfer experiments (training on one cosmological simulation and evaluating on another using rank-normalized features) achieve $83\%$--$89\%$ accuracy. This suggests the relative photometric ordering of growth regimes is largely preserved even across fundamentally different sub-grid SMBH feedback prescriptions. Signal decomposition shows our classification is driven by host galaxy colors ($82\%$--$87\%$ accuracy) and, relatedly, the accretion-state's spectral energy distribution shape as opposed to an inversion of our forward model's analytical luminosity prescription. Given that the evaluated simulations employ heavy seed prescriptions ($\geq 10^{4}~M_\odot$), our methodology establishes a validated baseline for classifying post-seeding growth regimes.

astro-ph.GA

Cross Subtype Transferability of Machine Learning Photometric Redshift Relations in Low Redshift Seyfert AGN

Photometric redshift estimation for active galactic nuclei (AGN) is complicated by the combined effects of host-galaxy light, nuclear emission, dust attenuation, and broadband spectral diversity. We investigate whether machine learning photo-z relations trained on one low-redshift Seyfert subtype remain valid when transferred to another, and whether probabilistic subtype classification can be used to identify sources for which a specialised regressor is reliable. Using spectroscopically selected Seyfert I and Seyfert II samples from SDSS, matched to AllWISE photometry over 0 < z_spec <= 0.6, we constructed a common 45-feature representation from SDSS ugriz and WISE W1-W4 data. Random Forest and XGBoost regressors were evaluated within each subtype, followed by controlled cross-subtype transfer tests, redshift and sample size-matched experiments, feature ablations, and an independent classifier-gated regression test. The subtype specific models achieved strong within-sample performance, with the Seyfert II model reaching R2 = 0.965 and sigma_NMAD = 0.0169. However, transfer between Seyfert I and Seyfert II produced a clear and asymmetric degradation in accuracy that persisted after matching the samples and restricting the photometric inputs. A probabilistic Seyfert classifier further identified subsets for which the Seyfert II regressor was more reliable, while extrapolation beyond the redshift range represented in training produced systematic underestimation. These results demonstrate that AGN photo-z performance depends strongly on the population and redshift domain represented in the training data, supporting subtype-aware calibration and applicability-based source selection.

astro-ph.GA