Search arXivSearch

arXiv subjects

Zhao

Publications and source records attributed to Zhao.

14 recordsLinked to original sources

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run. Machine-learning surrogates that predict these outcomes are increasingly used not only to propose candidates but to grade them, and even to feed their own predictions back into the search as though they were measurements. Through mathematical analysis validated on three exhaustively ground-truthed design tasks, we establish when this practice is safe, what any certificate of safety must cost, and when the substitution provably pays. Predictive accuracy cannot anchor trust: near-perfect R^2 is compatible with worst-possible selections, and screening N candidates inflates the over-prediction at the selected candidate by a quantifiable "selection tax" with matching upper and lower bounds. Safety follows instead from an architectural rule - predictions may propose and train without restriction, but every certified conclusion must rest on true evaluations - which is sufficient with no assumptions on the surrogate, and necessary, since admitting predictions into certification with the standing of measurements opens a deterministic self-confirmation failure mode. We derive the minimal criterion under which a model may act as an oracle (rank preservation, not accuracy), show that trust must be purchased through selection-aware audits that are optimal in query complexity, and prove a dichotomy fixing when audited surrogates cut certified evaluation cost. Across 432 surrogate fits over six task-regime conditions, the audit statistic tracks deployed search performance at Spearman rank correlation 0.80-0.99, while the rank correlation of R^2 with deployed regret falls as low as 0.33; audited screening reduces certified oracle cost by a measured factor of 25.

cs.LG

Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval

The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards for training two-tower models typically involve in-batch and/or out-of-batch negative sampling. However, these methods often produce easy negatives that models can quickly learn, failing to sufficiently challenge the model. To address this issue, a novel self-supervised hard negative sampling technique is proposed that leverages a large language model (LLM) to generate hard negatives from the same cluster during model training. By utilizing the LLM to learn media representations, the proposed approach ensures that the generated negatives are more challenging and informative. This real-time sampling framework is designed for seamless integration into production models, capable of handling billions of training data points with minimal computational complexity. Experiments on public datasets, along with deployment to a large-scale online system, demonstrate that the proposed negative sampling technique outperforms widely used industry methods. Furthermore, analysis in industrial applications reveals that this sampling method can help break inherent feedback loops in recommendations and significantly reduce popularity bias.

cs.IR

A free-boundary model of vascularized tumor growth with time-periodic coefficients

We investigate a free boundary PDE model for vascularized tumor growth with two time-periodic coefficients representing nutrient supply $\phi(t)$ and nutrient demand $\psi(t)$. Under radial symmetry, the model reduces to a nonautonomous ODE for the tumor radius. We prove a vanishing-persistence dichotomy governed by the averaged coefficients $\bar{\phi}$ and $\bar{\psi}$: if the average nutrient supply does not exceed the average nutrient demand, namely $\bar{\phi}\le \bar{\psi}$, the tumor radius tends to zero; if the average supply exceeds the average demand, namely $\bar{\phi}>\bar{\psi}$, the tumor persists and converges to a unique positive periodic solution. We also study the linear stability of this periodic solution when nonradial perturbations are imposed. Under a stronger condition on the nutrient supply and demand, we show that the radially symmetric periodic solution is linearly stable when the tumor aggressiveness parameter is sufficiently small. Numerical simulations are provided to illustrate the analytical results.

math.AP

Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures

Mixture-of-Experts (MoE) architecture offers enhanced efficiency for Large Language Models (LLMs) with modularized computation, yet its inherent sparsity poses significant hardware deployment challenges, including memory locality issues, communication overhead, and inefficient computing resource utilization. Inspired by the modular organization of the human brain, we propose Mozart, a novel algorithm-hardware co-design framework tailored for efficient training of MoE-based LLMs on 3.5D wafer-scale chiplet architectures. On the algorithm side, Mozart exploits the inherent modularity of chiplets and introduces: (1) an expert allocation strategy that enables efficient on-package all-to-all communication, and (2) a fine-grained scheduling mechanism that improves communication-computation overlap through streaming tokens and experts. On the architecture side, Mozart adaptively co-locates heterogeneous modules on specialized chiplets with a 2.5D NoP-Tree topology and hierarchical memory structure. Evaluation across three popular MoE models demonstrates significant efficiency gains, enabling more effective parallelization and resource utilization for large-scale modularized MoE-LLMs.

cs.AR

MAHL: Multi-Agent LLM-Guided Hierarchical Chiplet Design with Adaptive Debugging

As program workloads (e.g., AI) increase in size and algorithmic complexity, the primary challenge lies in their high dimensionality, encompassing computing cores, array sizes, and memory hierarchies. To overcome these obstacles, innovative approaches are required. Agile chip design has already benefited from machine learning integration at various stages, including logic synthesis, placement, and routing. With Large Language Models (LLMs) recently demonstrating impressive proficiency in Hardware Description Language (HDL) generation, it is promising to extend their abilities to 2.5D integration, an advanced technique that saves area overhead and development costs. However, LLM-driven chiplet design faces challenges such as flatten design, high validation cost and imprecise parameter optimization, which limit its chiplet design capability. To address this, we propose MAHL, a hierarchical LLM-based chiplet design generation framework that features six agents which collaboratively enable AI algorithm-hardware mapping, including hierarchical description generation, retrieval-augmented code generation, diverseflow-based validation, and multi-granularity design space exploration. These components together enhance the efficient generation of chiplet design with optimized Power, Performance and Area (PPA). Experiments show that MAHL not only significantly improves the generation accuracy of simple RTL design, but also increases the generation accuracy of real-world chiplet design, evaluated by Pass@5, from 0 to 0.72 compared to conventional LLMs under the best-case scenario. Compared to state-of-the-art CLARIE (expert-based), MAHL achieves comparable or even superior PPA results under certain optimization objectives.

cs.AR

Behavior of quantum coherence in the ultrastrong and deep strong coupling regimes of light-matter system

The ultrastrong and deep strong coupling regimes exhibit a variety of intriguing physical phenomena. In this work, we utilize the Hopfield model of a two-mode bosonic system, with each mode interacts with a heat reservoir, to research the behavior of quantum coherence. Our results indicate that a coupled oscillator system can exhibit significant quantum coherence in the ultrastrong and deep strong coupling regimes. In the ground state, the photon-mode and the matter-mode coherences are equal. The larger coherences that encompass the photon mode, the matter mode, and the overall system are achieved at lower optical frequencies and with increased coupling strengths. Notably, the the beam-splitter and phase rotation terms alone does not generate coherences for either total coherence or subsystem coherences; instead, the generation of quantum coherences originates from the one-mode and two-mode squeezing terms. When heat environments are present, the total coherence can be enhanced by the the beam-splitter and phase rotation terms, while it has no effect on subsystem coherences. Moreover, when the one-mode and two-mode squeezing terms and the the beam-splitter and phase rotation terms are considered together, the total coherence increases with stronger coupling. We also observe that lower frequencies maximize total coherence in the deep strong coupling regime. These results demonstrate that the ultrastrong and deep strong coupling regimes give rise to novel characteristics of quantum coherence. This work provides valuable insights into the quantum coherence properties, particularly in the ultrastrong and deep strong coupling regimes between light and matter and may have potential applications in quantum information processing.

quant-ph

A Dashboard Approach to Monitoring Mpox-Related Discourse and Misinformation on Social Media

Mpox (formerly monkeypox) is a zoonotic disease caused by an orthopoxvirus closely related to variola and remains a significant global public health concern. During outbreaks, social media platforms like X (formerly Twitter) can both inform and misinform the public, complicating efforts to convey accurate health information. To support local response efforts, we developed a researcher-focused dashboard for use by public health stakeholders and the public that enables searching and visualizing mpox-related tweets through an interactive interface. Following the CDC's designation of mpox as an emerging virus in August 2024, our dashboard recorded a marked increase in tweet volume compared to 2023, illustrating the rapid spread of health discourse across digital platforms. These findings underscore the continued need for real-time social media monitoring tools to support public health communication and track evolving sentiment and misinformation trends at the local level.

cs.SI

HDLxGraph: Bridging Large Language Models and HDL Repositories via HDL Graph Databases

Retrieval Augmented Generation (RAG) is an essential agent for Large Language Model (LLM) aided Description Language (HDL) tasks, addressing the challenges of limited training data and prohibitively long prompts. However, its performance in handling ambiguous queries and real-world, repository-level HDL projects containing thousands or even tens of thousands of code lines remains limited. Our analysis demonstrates two fundamental mismatches, structural and vocabulary, between conventional semantic similarity-based RAGs and HDL codes. To this end, we propose HDLxGraph, the first framework that integrates the inherent graph characteristics of HDLs with RAGs for LLM-assisted tasks. Specifically, HDLxGraph incorporates Abstract Syntax Trees (ASTs) to capture HDLs' hierarchical structures and Data Flow Graphs (DFGs) to address the vocabulary mismatch. In addition, to overcome the lack of comprehensive HDL search benchmarks, we introduce HDLSearch, an LLM generated dataset derived from real-world, repository-level HDL projects. Evaluations show that HDLxGraph improves search, debugging, and completion accuracy by 12.04%/12.22%/5.04% and by 11.59%/8.18%/4.07% over state-of-the-art similarity-based RAG and software-code Graph RAG baselines, respectively. The code of HDLxGraph and HDLSearch benchmark are available at https://github.com/UMN-ZhaoLab/HDLxGraph.

cs.AR

Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference

Mixture-of-experts (MoE) architectures could achieve impressive computational efficiency with expert parallelism, which relies heavily on all-to-all communication across devices. Unfortunately, such communication overhead typically constitutes a significant portion of the total runtime, hampering the scalability of distributed training and inference for modern MoE models (consuming over $40\%$ runtime in large-scale training). In this paper, we first define collaborative communication to illustrate this intrinsic limitation, and then propose system- and algorithm-level innovations to reduce communication costs. Specifically, given a pair of experts co-activated by one token, we call them "collaborated", which comprises $2$ cases as intra- and inter-collaboration, depending on whether they are kept on the same device. Our pilot investigations reveal that augmenting the proportion of intra-collaboration can accelerate expert parallelism at scale. It motivates us to strategically optimize collaborative communication for accelerated MoE training and inference, dubbed Occult. Our designs are capable of either delivering exact results with reduced communication cost or controllably minimizing the cost with collaboration pruning, materialized by modified fine-tuning. Comprehensive experiments on various MoE-LLMs demonstrate that Occult can be faster than popular state-of-the-art inference or training frameworks (more than $1.5\times$ speed up across multiple tasks and models) with comparable or superior quality compared to the standard fine-tuning. Code is available at $\href{https://github.com/UNITES-Lab/Occult}{https://github.com/UNITES-Lab/Occult}$.

cs.LG

HiVeGen -- Hierarchical LLM-based Verilog Generation for Scalable Chip Design

With Large Language Models (LLMs) recently demonstrating impressive proficiency in code generation, it is promising to extend their abilities to Hardware Description Language (HDL). However, LLMs tend to generate single HDL code blocks rather than hierarchical structures for hardware designs, leading to hallucinations, particularly in complex designs like Domain-Specific Accelerators (DSAs). To address this, we propose HiVeGen, a hierarchical LLM-based Verilog generation framework that decomposes generation tasks into LLM-manageable hierarchical submodules. HiVeGen further harnesses the advantages of such hierarchical structures by integrating automatic Design Space Exploration (DSE) into hierarchy-aware prompt generation, introducing weight-based retrieval to enhance code reuse, and enabling real-time human-computer interaction to lower error-correction cost, significantly improving the quality of generated designs.

cs.LG

Feature selection algorithm based on incremental mutual information and cockroach swarm optimization

Feature selection is an effective preprocessing technique to reduce data dimension. For feature selection, rough set theory provides many measures, among which mutual information is one of the most important attribute measures. However, mutual information based importance measures are computationally expensive and inaccurate, especially in hypersample instances, and it is undoubtedly a NP-hard problem in high-dimensional hyperhigh-dimensional data sets. Although many representative group intelligent algorithm feature selection strategies have been proposed so far to improve the accuracy, there is still a bottleneck when using these feature selection algorithms to process high-dimensional large-scale data sets, which consumes a lot of performance and is easy to select weakly correlated and redundant features. In this study, we propose an incremental mutual information based improved swarm intelligent optimization method (IMIICSO), which uses rough set theory to calculate the importance of feature selection based on mutual information. This method extracts decision table reduction knowledge to guide group algorithm global search. By exploring the computation of mutual information of supersamples, we can not only discard the useless features to speed up the internal and external computation, but also effectively reduce the cardinality of the optimal feature subset by using IMIICSO method, so that the cardinality is minimized by comparison. The accuracy of feature subsets selected by the improved cockroach swarm algorithm based on incremental mutual information is better or almost the same as that of the original swarm intelligent optimization algorithm. Experiments using 10 datasets derived from UCI, including large scale and high dimensional datasets, confirmed the efficiency and effectiveness of the proposed algorithm.

cs.AI

Operon Dynamics with State Dependent Transcription and/or Translation Delays

Transcription and translation retrieve and operationalize gene encoded information in cells. These processes are not instantaneous and incur significant delays. In this paper we study Goodwin models of both inducible and repressible operons with state-dependent delays. The paper provides justification and derivation of the model, detailed analysis of the appropriate setting of the corresponding dynamical system, and extensive numerical analysis of its dynamics. Comparison with constant delay models shows significant differences in dynamics that include existence of stable periodic orbits in inducible systems and multistability in repressible systems. A combination of parameter space exploration, numerics, analysis of steady state linearization and bifurcation theory indicates the likely presence of Shilnikov-type homoclinic bifurcations in the repressible operon model.

math.DS

Robotic Non-Destructive Testing of Manmade Structures: A Review of the Literature

This literature review investigates how robots can be used for the maintenance of manmade structures, such as pipes, reinforced concrete decks, and space stations as a sampling of the broad spectrum of robotic non-destructive testing (NDT) applications. Robotic NDT can be used to find plaque in pipes, corrosion in steel buildings, and impact damage in space stations, which would normally be invisible to the eye. After inspection, the inspected material is preserved in its original condition. The paper's structure is as follows: first, the definition of NDT is elaborated upon with the discussion of specific methods that will be used in the inspection of the structures mentioned above. Second, an explanation follows on why robots are suited to inspection, specifically focusing on robots' advantages over humans. Third, three real-world examples notify the reader of current progress in robot NDT. Lastly, a summary of robot problems serves as a reminder that testing and development must continue for robot NDT to become mainstream.

cs.RO

An improved spectrophotometry tests the Einstein-Smoluchowski equation: a revisit and update

Light-matter interaction in solvents has attracted continuous attention for theoretical prediction and experimental measure, due in part to the simple curiosity to nature, and in part to increasing calls from solvent-involved applications. Yet hitherto, a majority of reliable spectrophotometric measurements on transparent solvents upon visible light end up using long-path-length cells, usually over dozens of cm, rendering the measures costly and complex; meanwhile, the guidance for choosing the best formula to describe solvent scattering has remained unsettled. Here we theoretically and experimentally demonstrate a simple, low-cost, and versatile spectrophotometric method, recording sensitivity 10-4 dB/cm over 0.5 cm differential path length based on using a standard double-beam spectrophotometer. We attest the method reduces the path length by a factor of 100 while still making its closest approach to the record-low measurements. Revisiting the present equations of solvent scattering, we unfold that they all give similar-predictive-values, revealing the criterion of choice merely on the formula's simple practicality. Following the clarification of wavelengths over which light scattering dictates the solvent's extinction, we identify that the discrepancies persist between the calculated scattering coefficients and those measured results, suggesting the need for improving solvent scattering theory to comprehend the phenomenon in greater depth.

physics.chem-ph