Search arXivSearch

arXiv subjects

Yu-Tang Chang

Publications and source records attributed to Yu-Tang Chang.

4 recordsLinked to original sources

EB-gMCR: Energy-Based Generative Modeling for Signal Unmixing and Multivariate Curve Resolution

A single measurement of a chemical mixture, a reaction mixture, a natural extract, or a tissue, records the sum of the profiles of the few components it contains, each weighted by its concentration. Recovering the components and their concentrations from a collection of such samples is multivariate curve resolution (MCR). Classical MCR factorizes the data matrix, takes the component count as input, and leaves a rotational ambiguity that constraints narrow. This paper keeps the forward direction and models how a sample is made: each sample activates a few components from a pool of candidate components and is observed as their linear superposition plus noise. Under this model the continuous ambiguity of factorization collapses to the numbering of the components, and among all decompositions that reproduce the data, the one with the fewest component occurrences across samples is the true one. We prove three results. Under a spark condition, minimal usage identifies the true supports. A decomposition that reconstructs every sample within a tolerance set by the noise, with no more usage than the truth, is the truth up to the numbering of the components, and the recovered pool then decodes new samples on its own. As samples accumulate, what is recovered almost surely is the process itself, the pool and the component count; the decoding of any single sample stays limited by the noise, for every method. The solver EB-gMCR selects components per sample with an energy-based gate under a usage penalty. It recovers the component count on synthetic mixtures of up to 256 components and on two public spectroscopy datasets without being told the count, and a frozen model decodes unseen mixtures at the noise floor. Code: https://github.com/b05611038/ebgmcr_solver.

cs.LG

SparseEB-gMCR: A Generative Solver for Extreme Sparse Components with Application to Contamination Removal in GC-MS

Analytical chemistry instruments provide physically meaningful signals for elucidating analyte composition, and those with high mass or spectral resolution generate signals sparse enough for direct interpretation against chemical libraries. enerative multivariate curve resolution (gMCR) models a sample as the linear superposition of a few components drawn from a learned pool, and its energy-based solver (EB-gMCR) recovers the pool and the component count without being told the count. However, extreme sparsity in instruments such as GC-MS or 1H-NMR leaves the component profiles themselves sparse, and a dense learnable profile cannot represent an exact zero. To address this, a static support gate was introduced that applies the EB-select mechanism a second time, to the coordinates of each profile rather than to the components of each sample. The result, SparseEB-gMCR, reparameterizes the gMCR pool rather than changing the model, and the sparsity that motivates the extension makes components easier to tell apart rather than harder. On synthetic data, SparseEB-gMCR recovered the component count and reconstructed sparse mixtures as accurately as dense-component EB-gMCR, with the same graceful scaling in the number of components. It was then applied to real GC-MS chromatograms for unsupervised contamination removal, where siloxane-related pollution signals were eliminated and compound identification became more reliable. Removal reuses a pool learned from clean spectra alone, and rests on one requirement: that no combination of clean components can imitate the contamination, which is stronger than the two sets of components merely being different. With this sparse extension, the EB-gMCR family becomes applicable to wider ranges of real-world chemical datasets, providing a general mathematical framework for signal unmixing and contamination elimination in analytical chemistry.

cs.CE

Deep Global Clustering for Hyperspectral Image Segmentation: Concepts, Applications, and Open Challenges

Hyperspectral imaging (HSI) analysis faces computational bottlenecks due to massive data volumes that exceed available memory. While foundation models pre-trained on large remote sensing datasets show promise, their learned representations often fail to transfer to domain-specific applications like close-range agricultural monitoring where spectral signatures, spatial scales, and semantic targets differ fundamentally. This report presents Deep Global Clustering (DGC), a conceptual framework for memory-efficient HSI segmentation that learns global clustering structure from local patch observations without pre-training. DGC operates on small patches with overlapping regions to enforce consistency, enabling training in under 30 minutes on consumer hardware while maintaining constant memory usage. On a leaf disease dataset, DGC achieves background-tissue separation (mean IoU 0.925) and demonstrates unsupervised disease detection through navigable semantic granularity. However, the framework suffers from optimization instability rooted in multi-objective loss balancing: meaningful representations emerge rapidly but degrade due to cluster over-merging in feature space. We position this work as intellectual scaffolding - the design philosophy has merit, but stable implementation requires principled approaches to dynamic loss balancing. Code and data are available at https://github.com/b05611038/HSI_global_clustering.

cs.CV

Creating high density ensembles of nitrogen-vacancy centers in nitrogen-rich type Ib nanodiamonds

This work explores the possibility of increasing the density of negatively charged nitrogen-vacancy centers [NV-] in nanodiamonds using nitrogen-rich type Ib diamond powders as the starting materials. The nanodiamonds (10 - 100 nm in diameters) were prepared by ball-milling of microdiamonds, in which the density of neutral and automatically dispersed nitrogen atoms [N0] was measured by diffuse reflectance infrared Fourier transform spectroscopy (DRIFT). A systematic measurement for the fluorescence intensities and lifetimes of the crushed monocrystalline diamonds as a function of [N0] indicated that the [NV-] increases nearly linearly with [N0] at 100 - 200 ppm. The trend, however, failed to continue for nanodiamonds with higher [N0] (up to 390 ppm) but poorer crystallinity. We attribute the result to a combined effect of fluorescence quenching as well as the lower conversion efficiency of vacancies to NV- due to the presence of more impurities and defects in these as-grown diamond crystallites. The principles and practice of fabricating brighter and smaller fluorescent nanodiamonds (FNDs) are discussed.

cond-mat.mtrl-sci