Search arXivSearch

arXiv subjects

Hongyu Yu

Publications and source records attributed to Hongyu Yu.

At least 19 recordsLinked to original sources

Budget-Aware Compression Pipeline for Single-GPU LLM Inference: Methods, Trade-offs, and Coupling Effects

Single-GPU deployment of 70B-parameter language models on an NVIDIA GPU is constrained by device memory, long-context throughput, and engineering integration cost. We cast single-GPU inference as a budget-aware design problem over these three axes and study how pruning, quantization, and KV-cache compression interact under realistic execution. Controlled ablations show that layer-wise pruning makes weight quantization more robust. KV-cache sparsification complements INT8 KV quantization by reducing memory without hurting decoding speed, while static vector quantizers often conflict with dynamic caching. Guided by these coupling results and explicit budget tracking, we assembled a practical pipeline and compressed a 70B model to about 33 GB, sustained about 57 tokens/s on 10k token prompts on a single A40, and kept absolute accuracy within 5% on common and reasoning benchmarks. We contribute design rules and a reproducible evaluation protocol that jointly report quality, memory, and end-to-end speed, and we provide a foundation for automated pipeline search under realistic single-GPU constraints.

cs.CL

A Physical Response-and-Memory Model for Muon Optimization

Training large language models is costly. How low a loss the same compute can ultimately reach depends on how each step's gradient is converted into a weight update; the rule that performs this conversion is the optimizer. From SGD and AdamW to the recent Muon, effective update rules have mostly been shaped by engineering intuition and then selected on benchmarks. Muon semi-orthogonalizes the momentum matrix before applying the update and has kept breaking records on public training benchmarks; yet why the semi-orthogonalized direction works, and over how long a history the momentum should average, are two questions at present answered mainly by experience. Here we treat the weight matrix during training as a responsive medium with memory and build a physical model for it, in which both questions find answers: the semi-orthogonalized direction is the maximally dissipative response under an output-side safety budget, which explains why it works; momentum is the internal stress accumulated by the medium; how long it should average is set by the relaxation of this stress, and a real medium relaxes on more than one timescale, the simplest form being one fast and one slow. On this basis we propose the Bi-Maxwell optimizer. The framework further yields a testable consequence: gradient directions change fast early in training and more slowly later, so the optimal memory length should grow with training stage; step-by-step measurements of a proxy for it by a read-only probe across 8 independent training trajectories are consistent with this consequence. Replacing the memory kernel alone, from a single timescale to two, brings training to the target loss in noticeably fewer steps on a public large-language-model optimizer benchmark.

cs.LG

Rethinking Scientific Discovery in the Agentic Era

Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem formulation, literature grounding, model use, simulation, validation, and knowledge reuse. This paper presents \textbf{SCION (Scientific Collaborative Innovation with Agentic Organizational Nexus)}, an agentic scientific operating system that acts as an \textbf{organizational nexus}. Through a Science Agent serving as a \textbf{Meta-Harness}, SCION connects scientific tasks, tools, agents, artifacts, and memory, transforming research into an executable, auditable, and reusable operational process. At its core is the \textbf{Research Execution Plan (REP)}, which compiles high-level scientific intent into staged objectives, dependencies, verification checkpoints, tool requirements, expected artifacts, and fallback conditions. SCION further integrates hierarchical multi-agent execution, profile-driven specialization, selective context construction, governed delegation, and layered epistemic memory to support long-horizon scientific work. We formulate discovery under SCION as \textbf{Target-conditioned Inverse Search} and extend it to hidden-target settings through batch active search under finite experimental budgets. Applications in materials analysis, molecule design, and protein or antibody screening, together with experiments on scientific reading, idea generation, molecule generation, and antibody screening, show that SCION outperforms existing autonomous research-agent baselines, especially in decomposition, verification, refinement, and memory reuse. Overall, SCION shifts AI from isolated tools toward a coordinated operational layer for traceable and reusable scientific innovation.

cs.CL

Latent Genetic Algorithm for Crystal Structure Prediction

Predicting crystal structures requires navigating rugged energy landscapes in which favorable local motifs must be inherited across candidates with incompatible cells, densities, and symmetries. Conventional real-space crossover often destroys these motifs when parent structures are geometrically mismatched. Here we show that latent representations learned by pretrained universal interatomic potentials can serve as continuous evolutionary coordinates for crystal structure prediction. In the Latent Genetic Algorithm (LGA), offspring are generated by inverse optimization of atomic positions and lattice vectors to match a target latent representation, which is constructed via interpolation of the parent latent vectors. LGA suppresses high-energy and short-contact offspring, increases the HfO$_2$ ground-state recovery rate from 20-35% to 60-95%, and enables a unified variable-supercell search over 16 perovskites with a nearly tenfold reduction in search cost. Applied to (PbTiO$_3$)$_n$/(PbZrO$_3$)$_n$ superlattices, LGA reveals $\sqrt{2} \times 3\sqrt{2} \times 1$ long-period ground-state structures characterized by a common in-plane finite-$q$ modulation $q{_\parallel} = (1/6,1/6)$ and layer-coupled sidebands. To our knowledge, this in-plane periodicity has not been reported in any related oxide perovskite superlattice studies. Altogether, LGA offers a powerful representation-guided paradigm for ground-state structure prediction and provides a practical, decoder-free route toward materials inverse design.

physics.comp-ph

AISPO: Enhancing Depth Reliability for Robotic Manipulation of Non-Lambertian Objects via Affine-Invariant Shape Prior

Reliable depth perception is critical for robotic manipulation, especially for non-Lambertian objects such as transparent or highly specular surfaces, where raw depth measurements are often corrupted or missing. These failures frequently propagate to motion planning, resulting in invalid grasp poses and execution errors. We propose AISPO, a depth completion framework that improves depth reliability for manipulation in challenging sensing conditions. AISPO combines multi-scale RGB-D feature fusion with an affine-invariant shape prior to enforce geometric consistency and mitigate catastrophic depth failures. Unlike methods that focus primarily on average depth accuracy, our approach emphasizes physical plausibility and structural integrity of the predicted depth maps. Extensive benchmark evaluations demonstrate competitive performance and strong generalization to unseen objects and novel scenes. Real-world grasping experiments further show that enhanced depth reliability significantly improves manipulation success rates, particularly for transparent objects where many existing methods fail to produce physically usable depth estimates.

cs.RO

Tangent-Plane Evidential Uncertainty in Active Learning for Magnetic Interatomic Potentials

Magnetic interatomic potentials need to account for coupled lattice and spin degrees of freedom, yet constructing reliable training sets remains costly because noncollinear first-principles labels are expensive. Active learning can mitigate this cost, provided that the uncertainty estimate is physically meaningful for the magnetic-response targets that drive spin reorientation. Here we extend the $\mathrm{e}^2\mathrm{IP}$ evidential framework to magnetic machine-learning interatomic potentials by formulating the projected spin-force likelihood and the corresponding epistemic uncertainty in the tangent plane orthogonal to the local spin direction. This construction prevents the uncertainty model from allocating probability mass to a radial spin component that is absent from the constrained-moment supervision. Using bulk BiFeO$_3$ and monolayer CrTe$_2$ as benchmark systems, we show that the resulting tangent-plane epistemic uncertainty indicator $U_{\mathrm{epi}}^{\mathrm{sf}}$ correlates strongly with prediction error and selects more informative configurations than random sampling, simultaneously improving energy, force, and projected spin-force accuracy. These results demonstrate a physically interpretable and data-efficient route for constructing uncertainty-aware magnetic machine-learning interatomic potentials.

physics.comp-ph

A Fully Ab-Initio Spin-Lattice Dynamics Framework for Magnetic Materials

Coupled spin-lattice dynamics (SLD) underlie a wide range of magnetic phenomena, yet a unified first-principles framework that propagates both degrees of freedom without empirical parameterization has remained elusive. We present a fully ab initio SLD approach integrated into VASP, in which interatomic forces and effective magnetic fields are obtained at each time step from self-consistent constrained-moment density-functional calculations. The method is validated on four materials spanning ferromagnetic, non-collinear, and geometrically frustrated orders, recovering the correct magnetic ground state in every case from random initial conditions. SLD trajectories also provide physically correlated training data for magnetic machine-learning potentials, as demonstrated for BiFeO$_3$ by a reduction of up to approximately one order of magnitude in energy MAE over training on randomized spin configurations. This framework opens a practical first-principles route to finite-temperature spin-lattice coupled phenomena in magnetic materials.

cond-mat.mtrl-sci

When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents

This paper reveals that LLM-powered agents exhibit not only demographic bias (e.g., gender, religion) but also intergroup bias under minimal "us" versus "them" cues. When such group boundaries align with the agent-human divide, a new bias risk emerges: agents may treat other AI agents as the ingroup and humans as the outgroup. To examine this risk, we conduct a controlled multi-agent social simulation and find that agents display consistent intergroup bias in an all-agent setting. More critically, this bias persists even in human-facing interactions when agents are uncertain about whether the counterpart is truly human, revealing a belief-dependent fragility in bias suppression toward humans. Motivated by this observation, we identify a new attack surface rooted in identity beliefs and formalize a Belief Poisoning Attack (BPA) that can manipulate agent identity beliefs and induce outgroup bias toward humans. Extensive experiments demonstrate both the prevalence of agent intergroup bias and the severity of BPA across settings, while also showing that our proposed defenses can mitigate the risk. These findings are expected to inform safer agent design and motivate more robust safeguards for human-facing agents.

cs.AI

Empowering Robot Teleoperation: Exploring the Synergies Between Devices and Manipulator Controllers in a Comparative Study

Robot learning empowers the robot system with human brain-like intelligence to autonomously acquire and adapt skills through experience, enhancing flexibility and adaptability in various environments. Aimed at achieving a similar level of capability in large language models (LLMs) for embodied intelligence, data quality plays a crucial role in training a foundational model with diverse robot skills. In this study, we investigate the collection of data for manipulation tasks using teleoperation devices. Different devices yield varying effects when paired with corresponding controller strategies, including position-based inverse kinematic (IK) control, torque-based inverse dynamic (ID) control, and optimization-based compliant control. Analysis of experimental results suggests the importance of the relationship between teleoperation devices and controllers for real tasks.

cs.RO

Efficient E(3)-equivariant framework for universal charge density prediction

Electronic structure is ubiquitously obtained via density functional theory (DFT), where the charge density plays a central role. This work presents EdenGNN (Equivariant Density Graph Neural Network), a machine learning (ML) charge density model for electronic structure. Current universal ML charge density models are hampered by prohibitive computational costs. Furthermore, despite being trained on projector augmented-wave (PAW) based DFT datasets, they predict only the pseudo charge density, which is insufficient to reconstruct the electronic structure. In contrast, EdenGNN overcomes these limitations. It additionally predicts the augmentation occupancies, enabling electronic structure calculations with PAW accuracy. Critically, by employing a basis-expansion formulation with fully trainable radial basis functions and a $\Delta$-learning strategy to capture charge transfer, it is over an order of magnitude faster. Trained on the Materials Project database, our universal model, EdenGNN-Uni, accurately predicts the band structures for the majority of materials across a vast chemical space. These findings establish the ML charge density model as a scalable \textit{ab initio} method for large-scale electronic structure calculations and high-throughput screening.

cond-mat.mtrl-sci

Defect-Charge-Driven 90{\deg} Switching in HfO2

Hafnium dioxide (HfO2) is a CMOS-compatible ferroelectric showing both 180{\deg} and 90{\deg} switching, yet the microscopic nature of the 90{\deg} pathway remains unresolved. We show that the 90{\deg} rotation pathway, negligible in pristine HfO2, becomes dominant under E// [111] when induced by charged oxygen vacancies. This pathway is more fatigue-resistant than the 180{\deg} reversal pathway, while delivering the same polarization change along [111] (2Pr=60 {\mu}C/cm^2 ). This charge-driven switching arises from two factors: the crystal geometry of HfO2 and the intrinsic nature of rotational pathways, the latter suggesting a possible general tendency for defect charge to bias rotation over reversal in ferroelectrics. Together these findings reveal a pathway-level origin of fatigue resistance and establish defect charge as a general control parameter for polarization dynamics.

cond-mat.mtrl-sci

Recent Advances in Unconventional Ferroelectrics and Multiferroics

Emerging ferroic materials may pave a new way to next-generation nanoelectronic and spintronic devices due to their interesting physical properties. Here, we systematically review unconventional ferroelectric systems, from Hf-based and elementary ferroelectrics to stacking ferroelectricity, polar metallicity, fractional quantum ferroelectricity, wurtzite-type ferroelectricity, and freestanding membranes ferroelectricity. Moreover, multiferroic materials are reviewed, particularly the interplay between novel magnetic states and ferroelectricity, as well as ferrovalley-ferroelectric coupling. Finally, we conclude by discussing current challenges and future opportunities in this field.

cond-mat.mtrl-sci

Moir\'e-Induced Magnetoelectricity in Twisted Bilayer NiI2

Twisted magnetic van der Waals (vdW) materials offer a promising route for multiferroic engineering, yet modeling large-scale moir\'e superlattices remains challenging. Leveraging a newly developed SpinGNN++ framework that effectively handles spin-lattice coupled systems, we develop a comprehensive interatomic machine learning (ML) potential and apply it to twisted bilayer NiI2 (TBN). Structural relaxation introduces moir\'e-periodic "bumps" that modulate the interlayer spacing by about 0.55~\AA{} and in-plane ionic shifts up to 0.48~\AA{}. Concurrently, our ML potential, which faithfully captures all key spin interactions, produces reliable magnetic configurations; combined with the generalized KNB mechanism, it yields accurate spin-driven polarization. For twist angles 1.89^{\circ} \leq \theta \leq 2.45^{\circ}, both mechanisms become prominent, yielding rich polarization textures that combine ionic out-of-plane dipoles with purely electronic in-plane domains. In the rigid (unrelaxed) bilayer, skyrmions are absent; lattice relaxation is essential for generating polar-magnetic topologies. In contrast, near {\theta} \approx 60^{\circ}, stacking-dependent ferroelectric displacements dominate, giving rise to polar meron-antimeron networks. These results reveal cooperative ionic and spin-driven ferroelectricity in TBN, positioning twisted vdW magnets as adaptable platforms for tunable multiferroic devices.

cond-mat.mtrl-sci

Master Rules from Chaos: Learning to Reason, Plan, and Interact from Chaos for Tangram Assembly

Tangram assembly, the art of human intelligence and manipulation dexterity, is a new challenge for robotics and reveals the limitations of state-of-the-arts. Here, we describe our initial exploration and highlight key problems in reasoning, planning, and manipulation for robotic tangram assembly. We present MRChaos (Master Rules from Chaos), a robust and general solution for learning assembly policies that can generalize to novel objects. In contrast to conventional methods based on prior geometric and kinematic models, MRChaos learns to assemble randomly generated objects through self-exploration in simulation without prior experience in assembling target objects. The reward signal is obtained from the visual observation change without manually designed models or annotations. MRChaos retains its robustness in assembling various novel tangram objects that have never been encountered during training, with only silhouette prompts. We show the potential of MRChaos in wider applications such as cutlery combinations. The presented work indicates that radical generalization in robotic assembly can be achieved by learning in much simpler domains.

cs.RO

Efficient construction of effective Hamiltonians with a hybrid machine learning method

The effective Hamiltonian method is a powerful tool for simulating large-scale systems across a wide range of temperatures. However, previous methods for constructing effective Hamiltonian models suffer from key limitations: some require to manually predefine interaction terms limited flexibility in capturing complex systems, while others lack efficiency in selecting optimal interactions. In this work, we introduce the Lasso-GA Hybrid Method (LGHM), a novel approach that combines Lasso regression and genetic algorithms to rapidly construct effective Hamiltonian models. Such method is broadly applicable to both magnetic systems (e.g., spin Hamiltonians) and atomic displacement models. To verify the reliability and usefulness of LGHM, we take monolayer CrI_3 and Fe_3 GaTe_2 as examples. In both cases, LGHM not only successfully identifies key interaction terms with high fitting accuracy, but also reproduces experimental magnetic ground states and Curie temperatures with further Monte Carlo simulations. Notable, our analysis of monolayer Fe_3 GaTe_2 reveals that the single-ion anisotropy and Heisenberg interaction lead to an out-of-plane ferromagnetic ground state, while the fourth-order interactions contribute significantly to the high Curie temperature. Our method is general so it can be applied to construct other effective Hamiltonian models.

physics.comp-ph

Learning thin deformable object manipulation with a multi-sensory integrated soft hand

Robotic manipulation has made significant advancements, with systems demonstrating high precision and repeatability. However, this remarkable precision often fails to translate into efficient manipulation of thin deformable objects. Current robotic systems lack imprecise dexterity, the ability to perform dexterous manipulation through robust and adaptive behaviors that do not rely on precise control. This paper explores the singulation and grasping of thin, deformable objects. Here, we propose a novel solution that incorporates passive compliance, touch, and proprioception into thin, deformable object manipulation. Our system employs a soft, underactuated hand that provides passive compliance, facilitating adaptive and gentle interactions to dexterously manipulate deformable objects without requiring precise control. The tactile and force/torque sensors equipped on the hand, along with a depth camera, gather sensory data required for manipulation via the proposed slip module. The manipulation policies are learned directly from raw sensory data via model-free reinforcement learning, bypassing explicit environmental and object modeling. We implement a hierarchical double-loop learning process to enhance learning efficiency by decoupling the action space. Our method was deployed on real-world robots and trained in a self-supervised manner. The resulting policy was tested on a variety of challenging tasks that were beyond the capabilities of prior studies, ranging from displaying suit fabric like a salesperson to turning pages of sheet music for violinists.

cs.RO

Efficient prediction of potential energy surface and physical properties with Kolmogorov-Arnold Networks

The application of machine learning methodologies for predicting properties within materials science has garnered significant attention. Among recent advancements, Kolmogorov-Arnold Networks (KANs) have emerged as a promising alternative to traditional Multi-Layer Perceptrons (MLPs). This study evaluates the impact of substituting MLPs with KANs within three established machine learning frameworks: Allegro, Neural Equivariant Interatomic Potentials (NequIP), and the Edge-Based Tensor Prediction Graph Neural Network (ETGNN). Our results demonstrate that the integration of KANs generally yields enhanced prediction accuracies. Specifically, replacing MLPs with KANs in the output blocks leads to notable improvements in accuracy and, in certain scenarios, also results in reduced training times. Furthermore, employing KANs exclusively in the output block facilitates faster inference and improved computational efficiency relative to utilizing KANs throughout the entire model. The selection of an optimal basis function for KANs is found to be contingent upon the particular problem at hand. Our results demonstrate the strong potential of KANs in enhancing machine learning potentials and material property predictions.

physics.comp-ph

Role of Domain Walls on Imprint and Fatigue in HfO2-Based Ferroelectrics

HfO2-based ferroelectric materials are promising for the next generation of memory devices, attracting significant attention. However, their potential applications are significantly limited by fatigue and imprint phenomena, which affect device lifetime and memory capabilities. Here, to accurately describe the dynamics and field effects of HfO2, we adopt our newly developed DREAM-Allegro network scheme and develop a comprehensive machine-learning model for HfO2. Such model can not only predict the interatomic potential, but also predict Born effective charges. Applying such model, we explore the role of domain dynamics in HfO2 and find that the fatigue and imprint phenomena are closely related to the so-called E-path and T-path switching pathways. Based on the different atomic motions in the two paths, we propose that an inclined electric field can sufficiently suppress fatigue and enhancing the performance of HfO2-based ferroelectric devices.

cond-mat.mtrl-sci