Search arXiv⌕ Search

arXiv · 2610.05844

PhaseMatcher: Autoregressive Phase-Set Identification with Spectral Decomposition

Abstract

Recovering complete phase sets from powder X-ray diffraction (PXRD) is challenging when weak-phase peaks overlap stronger signals. A natural strategy is to identify phases iteratively, removing the contribution of each identified phase from the observed pattern before predicting the next. However, even after a phase is correctly identified, misestimating its contribution can distort the residual and cause subsequent errors. We introduce PhaseMatcher, an autoregressive framework for complete phase-set identification with physics-guided spectral decomposition. After each phase prediction, PhaseMatcher re-estimates the contributions of all selected phases and the residual from the original observation and all selected reference patterns, accounting for physically plausible variation between reference patterns and the corresponding phase contributions in the observation. The resulting residual guides subsequent phase identification, while a separate stopping module determines when the phase set is complete. On synthetic mixtures and controlled mixtures constructed from measured single-phase patterns, PhaseMatcher improves complete-set identification over the evaluated baselines. On PhaseMix-135K, it also estimates contributions and residuals more accurately than scalar subtraction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhonglong Peng, Qiuliang Liu, Chang Chen, Geng Zhong, Qi Li, Lihong Wang, Lan Jiang, Shifeng Jin. 2026-10-05. PhaseMatcher: Autoregressive Phase-Set Identification with Spectral Decomposition. https://arxiv.org/abs/2610.05844

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The PUR-1 Cyber-Physical Digital Twin

Digital twin technologies have the potential to improve operational flexibility and responsiveness capabilities of nuclear systems. To provide decision support, cyber event characterization, state estimation, predictive control, and real-time dynamic processing of operational data, however, an efficient digital twin needs to integrate multiple models (data-driven as well as physics-based) with explainability while at the same time maintain two-way synchronization with the physical facility at a time constant less than its operational cycle. In this work, we present the Purdue University Reactor One Digital Twin (PUR-1 DT), a cyber-physical digital twin with a complete high-fidelity physics-based and AI-driven virtual model stack (neutronics, thermal-hydraulics, point kinetics) which provides closed-loop explainable diagnostics, forecasting, predictive control, and action recommendation back to the reactor via two-way communications and a cyber-physical testbed. We demonstrate real-time synchronized state estimation and short-term forecasting over a full reactor operational cycle and conduct a series of benchmarking experiments to validate accuracy and latency. Our results show good agreement with experimental results and lay the groundwork for further development and experimental demonstration of DT-enabled functionalities in real-world facilities.

cs.CE↗

MiDShip: Multimodal Dataset of Ship Cargo Hold Structures for Engineering Design

Ship structures govern vessel strength, safety, and manufacturability, but their design must satisfy hundreds of classification society requirements, making the process complex and iterative. Data-driven approaches are limited by the lack of structured datasets linking design geometry, structural performance, and rule-based constraints. This paper presents MiDShip, a multimodal dataset of 12,753 synthetic cargo-hold structural designs: 6,020 random, 496 generated by an SGLD-inspired procedure, and 6,237 generated by an equation-informed repair procedure. Each design includes parametric data, full and mesh-ready 3D geometry, engineering drawings and annotations, a bill of materials, and preliminary structural evaluations. Twenty-five constraints derived from a subset of ABS MVR are also evaluated. None of the random designs satisfies all constraints. Among the SGLD-inspired designs, 322 (64.9%) were fully compliant, with an average of 0.409 violations, 82.7% below the seed mean and 96.9% below the random-design mean. The repair procedure, developed through LLM-assisted code analysis, produced 4,952 fully compliant designs (79.4%), averaging 0.296 violations, 97.1% below the paired-source mean. In equal-size comparisons, mean nearest-neighbor distances in the scaled 120-parameter space were 3.495 for repaired designs, 1.144 for SGLD batches, and 3.729 for random designs. The primary contribution is the synchronized dataset and its generation and evaluation infrastructure; the generation studies demonstrate its utility rather than proposing new optimization algorithms. MiDShip supports machine learning, generative design, and automated rule-based evaluation for ship structures.

cs.CE↗

From Research Gaps to Theoretical Opportunities: Theory-Oriented GenAI for Research Opportunity Evaluation

Generative AI (GenAI) can explore large bodies of literature and generate plausible research ideas, but identifying what is missing, understudied, contradictory, or potentially connected does not by itself reveal where theory should advance. We develop a theory-oriented agentic AI system that helps researchers identify potential theorizing opportunities by incorporating established theorizing approaches into literature exploration and evaluation. The system operates through three stages. Stage 1 expands the theoretical search space and constructs a provisional Candidate Knowledge Graph. Stage 2 independently reconstructs what the literature supports through source grounded evidence extraction and theory-state reconstruction. Stage 3 evaluates the reconstructed knowledge state to determine whether an unresolved configuration warrants theory development or another research action and, when theory development is warranted, which theorizing approach is appropriate. We demonstrate the system through an end-to-end analysis of human oversight of agentic AI systems in organizations. The analysis shows that literature gaps alone are insufficient for identifying theoretical opportunities. For example, "transparency to trust" is routed to mechanism-based theorizing because the relationship is repeatedly documented in prior studies, while the generative mechanism explaining how transparency shapes trust remains insufficiently specified. By combining large-scale literature processing, structured knowledge representation, and theorizing-guided diagnosis, the system serves as a theory-oriented research assistant that supports researchers in identifying theoretically meaningful directions for subsequent research.

cs.CE↗