Search arXivSearch

arXiv · 2002.10464

Dimensionality Reduction of SDSS Spectra with Variational Autoencoders

Abstract

High resolution galaxy spectra contain much information about galactic physics, but the high dimensionality of these spectra makes it difficult to fully utilize the information they contain. We apply variational autoencoders (VAEs), a non-linear dimensionality reduction technique, to a sample of spectra from the Sloan Digital Sky Survey. In contrast to Principal Component Analysis (PCA), a widely used technique, VAEs can capture non-linear relationships between latent parameters and the data. We find that a VAE can reconstruct the SDSS spectra well with only six latent parameters, outperforming PCA with the same number of components. Different galaxy classes are naturally separated in this latent space, without class labels having been given to the VAE. The VAE latent space is interpretable because the VAE can be used to make synthetic spectra at any point in latent space. For example, making synthetic spectra along tracks in latent space yields sequences of realistic spectra that interpolate between two different types of galaxies. Using the latent space to find outliers may yield interesting spectra: in our small sample, we immediately find unusual data artifacts and stars misclassified as galaxies. In this exploratory work, we show that VAEs create compact, interpretable latent spaces that capture non-linear features of the data. While a VAE takes substantial time to train (~1 day for 48000 spectra), once trained, VAEs can enable the fast exploration of large astronomical data sets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Stephen K. N. Portillo, John K. Parejko, Jorge R. Vergara, Andrew J. Connolly. 2020-07-09. Dimensionality Reduction of SDSS Spectra with Variational Autoencoders. https://doi.org/10.3847/1538-3881%2Fab9644

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Simons Observatory: Development of a Pipeline to Detect Rapid Transients in Time-Ordered Data

We introduce a method for detecting astrophysical transients evolving on timescales of milliseconds to minutes using cosmic microwave background (CMB) survey telescopes. While previous transient searches in CMB data operate in map space, our pipeline directly processes the raw time-ordered data, enabling sensitivity to fast, dynamic signals. We integrate our detection approach into the Simons Observatory time-domain pipeline and assess the performance by injecting symmetric, stellar flare-like light curves into simulated observations. For events flaring with a timescale of 0.5 s, the pipeline detects $\gtrsim90$ % of events at flux densities of 800, 1150, 1650, and 4250\,mJy when measured in the 93, 145, 225, and 280 GHz bands respectively. At a fixed peak flux density, the pipeline more readily detects longer flares. The limiting flux density for 90 % completeness is four times lower for a $\ge5$ s flare than for a 0.5 s flare, while the flux density limits for $\gtrsim50$ % detection efficiency are comparable to the rms noise of the time-ordered data. We are able to determine the position of detected events in each observing band, with a positional uncertainty at the detection threshold comparable to the telescope resolution at that band. These results demonstrate the readiness of this pipeline for incorporation into upcoming Simons Observatory data analyses.

astro-ph.IM

Fitting Moving Objects in Up-The-Ramp Data with Applications to the Roman Space Telescope and JWST

A moving object breaks the fundamental property of constant per-pixel count rates in an astronomical image read out up-the-ramp. In this paper, we show how to fit a moving object's path across a detector as that detector is read out nondestructively. We write the full likelihood function for every pixel subject to a constant count rate plus a time-dependent count rate due to a moving source. Assuming the moving source to be point-like and assuming the effective point-spread function to be known, we are left with four parameters that enter the likelihood nonlinearly: two for position and two for velocity. All remaining parameters can be optimized using closed-form expressions. Our approach extracts maximal information on a moving source's position and speed and enables the source to be accurately removed from the image. We investigate the dependence of flux, position, and velocity precision on the target's speed and the readout pattern. We also find a small, positive bias on the recovered flux due to the need to fit for an uncertain position and speed. Our approach can be used for space-based images with minor Solar system bodies in the foreground, e.g.~from Roman and JWST, or for ground-based observations with satellites in the foreground. We demonstrate the promise of our method with a fit to an asteroid track observed serendipitously by the NIRISS instrument on JWST, comparing it to the performance of the JWST pipeline. Python code implementing our approach is available at https://github.com/t-brandt/moving_source. The total computational cost to fit the track of a moving object is $\sim$1 second on a 2023 Macbook Pro.

astro-ph.IM

Options for Compression of radio interferometry data: lossy compression of visibilities and lossless compression of uv-visibility grids for the MHONGOOSE survey

Next generation radio astronomy telescopes are challenging existing data reduction paradigms. With ever more antennas, larger bandwidths, and sometimes multiple primary beams, they often generate more observed data products than can readily be stored long-term. Thus, data storage becomes a major cost driver and processing constraint. In this paper, we test two methods of addressing this problem: grid-stacking, a two-stage lossless compression solution; and the lossy compression of the raw visibilities before traditional processing. To demonstrate these solutions we utilised a deep imaging pipeline based on software for the ASKAP telescope, ASKAPSoft, but applied to a strong source (NGC1566) from the deep MeerKAT HI spectral line project, MHONGOOSE. The grid-stacking solution reproduces the spectrum from traditional processing to within better than 0.7%, and also allows for the reconstruction of other weighting scales without significant computing costs. In comparison, image-stacking also reproduces the spectrum from the traditional processing, to within better than 3% but with worse image residuals in the cube. The lossy compression, even at a near ten-fold reduction in file size, reproduces the spectra almost perfectly (to better than ~0.01% in all cases). Thus both compression methods are promising solutions, and we discuss considerations for their application.

astro-ph.IM