Search arXivSearch

arXiv · 1602.08977

Clustering Based Feature Learning on Variable Stars

Abstract

The success of automatic classification of variable stars strongly depends on the lightcurve representation. Usually, lightcurves are represented as a vector of many statistical descriptors designed by astronomers called features. These descriptors commonly demand significant computational power to calculate, require substantial research effort to develop and do not guarantee good performance on the final classification task. Today, lightcurve representation is not entirely automatic; algorithms that extract lightcurve features are designed by humans and must be manually tuned up for every survey. The vast amounts of data that will be generated in future surveys like LSST mean astronomers must develop analysis pipelines that are both scalable and automated. Recently, substantial efforts have been made in the machine learning community to develop methods that prescind from expert-designed and manually tuned features for features that are automatically learned from data. In this work we present what is, to our knowledge, the first unsupervised feature learning algorithm designed for variable stars. Our method first extracts a large number of lightcurve subsequences from a given set of photometric data, which are then clustered to find common local patterns in the time series. Representatives of these patterns, called exemplars, are then used to transform lightcurves of a labeled set into a new representation that can then be used to train an automatic classifier. The proposed algorithm learns the features from both labeled and unlabeled lightcurves, overcoming the bias generated when the learning process is done only with labeled data. We test our method on MACHO and OGLE datasets; the results show that the classification performance we achieve is as good and in some cases better than the performance achieved using traditional features, while the computational cost is significantly lower.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Cristóbal Mackenzie, Karim Pichara, Pavlos Protopapas. 2016-02-29. Clustering Based Feature Learning on Variable Stars. https://doi.org/10.3847/0004-637x%2F820%2F2%2F138

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Stellar characterization with photometric colors from J-PLUS and 2MASS surveys

Aims. We aim at deriving stellar atmospheric parameters based on the photometric data from the Javalambre Photometric Local Universe Survey (J-PLUS) in addition to near-infrared photometry from the Two Micron All-Sky Survey (2MASS). Methods. Our method consists of a semi-supervised machine learning approach based on the k-means method combined with a modified k-nearest neighbors algorithm. This method compares the observed photometry to a set of reference data to estimate the stellar effective temperature ($T_{\rm eff}$), surface gravity ($\log{g}$), and metallicity ([Fe/H]) of stars from J-PLUS Data Release 3 (DR3). Results. We estimated $T_{\rm eff}$, $\log{g}$, and [Fe/H], for approximately 5.6 million stars from J-PLUS DR3, along with their errors.Our results were in agreement with spectroscopic estimates from LAMOST and APOGEE.We also applied a dimension reduction method, seeking greater efficiency by reducing the computation time and minimizing the needed information for calculating the stellar parameters, resulting in a subset of 11 colors. From this approach, stellar parameters were obtained for approximately six million stars. Conclusions. Our results demonstrated the potential of using a method built from machine learning algorithms that do not require prior training. Additionally, it was shown that the proposed method allowed estimating reliable atmospheric parameters even when the available photometry did not fulfill all photometric quality criteria. We defined a neighborhood parameter, which assesses the reliability of our estimations and indicates that objects with smaller neighborhoods values have lower uncertainties.

astro-ph.SR

Population demographics of post-interaction WDMS binaries: From common envelope evolution to stable mass transfer

Close white dwarf (WD) + main-sequence (MS) binaries are end products of mass transfer (MT) that occurred prior to WD formation, making their population demographics a powerful probe of binary evolution. Several recent works have constructed samples of WD+MS binaries with well-understood selection functions using data from wide-field surveys. These include (a) AU-scale astrometric binaries from Gaia that can be shown to contain a WD on dynamical grounds, (b) AU-scale astrometric binaries in which a hot WD is detected through a GALEX UV excess, and (c) close binaries discovered through eclipses. Together, these samples probe outcomes of both stable MT and common-envelope evolution, and interactions on both the red giant branch (RGB) and asymptotic giant branch (AGB). We forward model the three observed samples simultaneously. This approach produces robust constraints on uncertain binary evolution parameters because binaries removed from one population are predicted to appear in another. Our modeling includes a realistic initial binary population and treatments of the selection effects affecting all samples. We confirm that MT from AGB donors requires a critical accretor-to-donor mass ratio of $\sim0.4$ as found in previous work, and find that this is more stable than MT from RGB donors, for which we constrain a critical ratio $\gtrsim0.65$. A common envelope efficiency of $αλ\sim0.3$ matches the relative numbers of close and wide systems and the period distribution of close systems. Most stable MT products in the sample, including those with RGB donors, retain nonzero eccentricities ($\simeq0.1$). The model does not fully reproduce the mass distribution of main-sequence stars in post-common envelope binaries, which shows a cliff below the fully convective limit, possibly pointing to missing physics that may warrant future work.

astro-ph.SR

The Solar Neighborhood LVI: The RMSTAR Catalog of the Nearest 3352 M Dwarf Systems and 305 of their Wide Companions

We present the RMSTAR (RECONS M STAR) catalog, a 25 pc volume-limited and effectively volume-complete sample of the nearest 3352 M dwarf systems and their 305 wide stellar and 9 brown dwarf companions. RMSTAR has been created using only results from Gaia Data Releases 3 (GDR3) and 2 (GDR2), Hipparcos, and ground-based discoveries of M dwarf systems that are not available in the space-based results. The wide companions are identified using only results from GDR3 and GDR2, and have separations $ρ\ge0.''46$. There are 291 primaries with stellar companions, yielding a multiplicity rate of 8.68$\pm$0.49% for these widely separated systems, with a rate of 6.86$\pm$0.44% for projected separations $s\geq30$ au, where the sample of stellar companions is complete except perhaps for a few fringe cases. It is shown that the rate of stellar companions increases from the search limit of 30,000 au down to separations of 30 au, with a large set of companions to be characterized at closer separations in future work. We determine luminosity and mass functions for all stellar red dwarfs and their wide secondaries, finding that the luminosity function displays a classic turnover at $M_G\sim11$, whereas the mass function is described by an exponential function that rises from 0.60 M$_{\odot}$ to the end of the stellar main sequence at 0.075 M$_{\odot}$. We assess the large population of close, unresolved companions to M dwarfs by analyzing several Gaia parameters, identifying 1089 (30%) individual sources as having at least one of these elevated unresolved companion indication parameters.

astro-ph.SR