Search arXiv⌕ Search

arXiv · 2404.12270

Efficient Identification of Broad Absorption Line Quasars using Dimensionality Reduction and Machine Learning

Abstract

Broad Absorption Line Quasars (BALQSOs) displaying distinct blue-shifted broad absorption lines. These serve as invaluable probes for unraveling the intricate structure and evolution of quasars, shedding light on the profound influence exerted by supermassive black holes on galaxy formation. The proliferation of large-scale spectroscopic surveys such as LAMOST, SDSS, and DESI has exponentially expanded the repository of quasar spectra at our disposal. In this study, we present an innovative approach to streamline the identification of BALQSOs, leveraging the power of dimensionality reduction and machine learning algorithms. Our dataset is curated from the SDSS DR16, amalgamating quasar spectra with classification labels sourced from the DR16Q quasar catalog. We employ a diverse array of dimensionality reduction techniques, including Principal Component Analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t-SNE), Locally Linear Embedding (LLE), and Isometric Mapping (ISOMAP), to distill the essence of the original spectral data. The resultant low-dimensional representations serve as inputs for a suite of machine learning classifiers, including XGBoost and Random Forest models. Through experimentation, we unveil PCA as the most effective dimensionality reduction methodology, adeptly navigating the intricate balance between dimensionality reduction and preservation of vital spectral information. Notably, the synergistic fusion of PCA with the XGBoost classifier emerges as the pinnacle of efficacy in the BALQSO classification endeavor, boasting impressive accuracy rates of 97.60% by 10-cross validation and 96.92% on the outer test sample. This study not only introduces a novel machine learning-based paradigm for quasar classification but also offers invaluable insights transferrable to a myriad of spectral classification challenges pervasive in the realm of astronomy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei-Bo Kao, Yanxia Zhang, Xue-Bing Wu. 2024-04-18. Efficient Identification of Broad Absorption Line Quasars using Dimensionality Reduction and Machine Learning. https://doi.org/10.1093/pasj%2Fpsae037

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Entangling of Supernova Feedback Impacts with Coarsening Simulation Resolution

It is often understood that supernova (SN) feedback in galaxies is responsible for regulating star formation (SF) and generating gaseous outflows. However, a detailed look at the small-scale effects of SNe on the interstellar medium (ISM) in simulations shows that the macroscopic processes of SF suppression and outflow generation proceed in distinct channels. We demonstrate this finding in two independent simulations of isolated dwarf galaxies with very high (m_gas ~ Msun) numerical resolution, LYRA and RIGEL. Our findings suggest that the macroscopic effect of a given SN on the galaxy is best predicted by its local density. Outflows are driven by SNe in diffuse regions expanding to their cooling radii on large (~kpc) scales, while dense SF regions are disrupted in a localized (~pc) manner. However, these separate feedback channels are only distinguishable at very high resolutions capable of following mass scales \lesssim 10^2 \msun. When averaging on coarser scales, ISM densities are greatly mis-estimated, and variations between different SF and SNe-affected regions are severely washed out. It therefore cannot be __self-consistently__ determined, from coarse-resolution information __alone__, (1) whether a SN tends to contribute to outflows or direct SF suppression, and (2) the rate of SF in a given region. In particular, commonly used parameters in coarse-resolution (subgrid) models, such as the SN cooling radius and SF density threshold, may require more detailed treatments informed by high-resolution studies.

astro-ph.GA↗

Computational advances and challenges in simulations of turbulence and star formation

We review recent advances in the numerical modeling of turbulent flows and star formation. An overview of the most widely used simulation codes and their core capabilities is provided. We then examine methods for achieving the highest-resolution magnetohydrodynamical turbulence simulations to date, highlighting challenges related to numerical viscosity and resistivity. State-of-the-art approaches to modeling gravity and star formation are discussed in detail, including implementations of star particles and feedback from jets, winds, heating, ionization, and supernovae. We review the latest techniques for radiation hydrodynamics, including ray tracing, Monte Carlo, and moment methods, with comparisons between the flux-limited diffusion, moment-1, and variable Eddington tensor methods. The final chapter summarizes advances in cosmic-ray transport schemes, emphasizing their growing importance for connecting small-scale star formation physics with galaxy-scale evolution.

astro-ph.GA↗

How significant is the lensing interpretation of GW231123?

GW231123 is one of the most unusual gravitational-wave (GW) events, with exceptionally large inferred masses and near-extremal spins, offering an opportunity to test whether propagation effects contribute to these properties. We therefore examine whether the data support wave-optics microlensing embedded in a strong-lensing galaxy, whose detection becomes increasingly likely as observations accumulate, whether this interpretation can explain these properties, and how significant the preference remains under detector noise and waveform systematics. We compare six hypotheses: unlensed, isolated point mass, and embedded point-mass (EPM) and binary-lens (EB) effective models in Type-I (minimum) and Type-II (saddle) macro images. The EB Type-I model is most favored. For the most accurate waveform model NRSur7dq4, it gives $\log_{10}B^{\rm EB-I}_{\rm U}=2.60$, versus $0.89$ for Type II, indicating sensitivity to macro-image geometry. Within Type I, however, the binary improves over the point mass by only $\log_{10}B^{\rm EB-I}_{\rm EPM-I}=0.16$ and $Δ\ln\mathcal{L}_{\max}=0.56$, providing no clear evidence for structure beyond a single effective perturber. Moreover, under embedded lensing, waveform-template discrepancies and inferred masses and spins are reduced. However, real O4a backgrounds from numerical-relativity injections show that the apparent lensing evidence is sensitive to waveform systematics and realistic detector noise: although the commonly used waveform IMRPhenomXPHM gives the largest Bayes factor, $\log_{10}B^{\rm EB-I}_{\rm U}=4.52$, it is less exceptional relative to its own background, with a false-alarm probability of $6.5$--$8\%$, whereas NRSur7dq4 gives only $2$--$3\%$. Thus, waveform systematics can amplify apparent lensing evidence, but GW231123 remains an intriguing lensing candidate.

astro-ph.GA↗