Search arXiv⌕ Search

arXiv subjects

Yue Yu

Publications and source records attributed to Yue Yu.

At least 397 records · Page 22Linked to original sources

A data-driven peridynamic continuum model for upscaling molecular dynamics

Nonlocal models, including peridynamics, often use integral operators that embed lengthscales in their definition. However, the integrands in these operators are difficult to define from the data that are typically available for a given physical system, such as laboratory mechanical property tests. In contrast, molecular dynamics (MD) does not require these integrands, but it suffers from computational limitations in the length and time scales it can address. To combine the strengths of both methods and to obtain a coarse-grained, homogenized continuum model that efficiently and accurately captures materials' behavior, we propose a learning framework to extract, from MD data, an optimal Linear Peridynamic Solid (LPS) model as a surrogate for MD displacements. To maximize the accuracy of the learnt model we allow the peridynamic influence function to be partially negative, while preserving the well-posedness of the resulting model. To achieve this, we provide sufficient well-posedness conditions for discretized LPS models with sign-changing influence functions and develop a constrained optimization algorithm that minimizes the equation residual while enforcing such solvability conditions. This framework guarantees that the resulting model is mathematically well-posed, physically consistent, and that it generalizes well to settings that are different from the ones used during training. We illustrate the efficacy of the proposed approach with several numerical tests for single layer graphene. Our two-dimensional tests show the robustness of the proposed algorithm on validation data sets that include thermal noise, different domain shapes and external loadings, and discretizations substantially different from the ones used for training.

cond-mat.mtrl-sci↗

Nonlocal Trace Spaces and Extension Results for Nonlocal Calculus

For a given Lipschitz domain $Ω$, it is a classical result that the trace space of $W^{1,p}(Ω)$ is $W^{1-1/p,p}(\partialΩ)$, namely any $W^{1,p}(Ω)$ function has a well-defined $W^{1-1/p,p}(\partialΩ)$ trace on its codimension-1 boundary $\partialΩ$ and any $W^{1-1/p,p}(\partialΩ)$ function on $\partialΩ$ can be extended to a $W^{1,p}(Ω)$ function. Recently, Dyda and Kassmann (2019) characterize the trace space for nonlocal Dirichlet problems involving integrodifferential operators with infinite interaction ranges, where the boundary datum is provided on the whole complement of the given domain $\mathbb{R}^d\backslashΩ$. In this work, we study function spaces for nonlocal Dirichlet problems with a finite range of nonlocal interactions, which naturally serves a bridging role between the classical local PDE problem and the nonlocal problem with infinite interaction ranges. For these nonlocal Dirichlet problems, the boundary conditions are normally imposed on a region with finite thickness volume which lies outside of the domain. We introduce a function space on the volumetric boundary region that serves as a trace space for these nonlocal problems and study the related extension results. Moreover, we discuss the consistency of the new nonlocal trace space with the classical $W^{1-1/p,p}(\partialΩ)$ space as the size of nonlocal interaction tends to zero. In making this connection, we conduct an investigation on the relations between nonlocal interactions on a larger domain and the induced interactions on its subdomain. The various forms of trace, embedding and extension theorems may then be viewed as consequences in different scaling limits.

math.AP↗

An asymptotically compatible probabilistic collocation method for randomly heterogeneous nonlocal problems

In this paper we present an asymptotically compatible meshfree method for solving nonlocal equations with random coefficients, describing diffusion in heterogeneous media. In particular, the random diffusivity coefficient is described by a finite-dimensional random variable or a truncated combination of random variables with the Karhunen-Loève decomposition, then a probabilistic collocation method (PCM) with sparse grids is employed to sample the stochastic process. On each sample, the deterministic nonlocal diffusion problem is discretized with an optimization-based meshfree quadrature rule. We present rigorous analysis for the proposed scheme and demonstrate convergence for a number of benchmark problems, showing that it sustains the asymptotic compatibility spatially and achieves an algebraic or sub-exponential convergence rate in the random coefficients space as the number of collocation points grows. Finally, to validate the applicability of this approach we consider a randomly heterogeneous nonlocal problem with a given spatial correlation structure, demonstrating that the proposed PCM approach achieves substantial speed-up compared to conventional Monte Carlo simulations.

math.NA↗

On the prescription of boundary conditions for nonlocal Poisson's and peridynamics models

We introduce a technique to automatically convert local boundary conditions into nonlocal volume constraints for nonlocal Poisson's and peridynamic models. The proposed strategy is based on the approximation of nonlocal Dirichlet or Neumann data with a local solution obtained by using available boundary, local data. The corresponding nonlocal solution converges quadratically to the local solution as the nonlocal horizon vanishes, making the proposed technique asymptotically compatible. The proposed conversion method does not have any geometry or dimensionality constraints and its computational cost is negligible, compared to the numerical solution of the nonlocal equation. The consistency of the method and its quadratic convergence with respect to the horizon is illustrated by several two-dimensional numerical experiments conducted by meshfree discretization for both the Poisson's problem and the linear peridynamic solid model.

math.NA↗

GGT: Graph-Guided Testing for Adversarial Sample Detection of Deep Neural Network

Deep Neural Networks (DNN) are known to be vulnerable to adversarial samples, the detection of which is crucial for the wide application of these DNN models. Recently, a number of deep testing methods in software engineering were proposed to find the vulnerability of DNN systems, and one of them, i.e., Model Mutation Testing (MMT), was used to successfully detect various adversarial samples generated by different kinds of adversarial attacks. However, the mutated models in MMT are always huge in number (e.g., over 100 models) and lack diversity (e.g., can be easily circumvented by high-confidence adversarial samples), which makes it less efficient in real applications and less effective in detecting high-confidence adversarial samples. In this study, we propose Graph-Guided Testing (GGT) for adversarial sample detection to overcome these aforementioned challenges. GGT generates pruned models with the guide of graph characteristics, each of them has only about 5% parameters of the mutated model in MMT, and graph guided models have higher diversity. The experiments on CIFAR10 and SVHN validate that GGT performs much better than MMT with respect to both effectiveness and efficiency.

cs.LG↗

Majorana fermion arcs and the local density of states of UTe$_2$

$\text{UTe}_2$ is a leading candidate for chiral p-wave superconductivity, and for hosting exotic Majorana fermion quasiparticles. Motivated by recent STM experiments in this system, we study particle-hole symmetry breaking in chiral p-wave superconductors. We compute the local density of states from Majorana fermion surface states in the presence of Rashba surface spin-orbit coupling, which is expected to be sizeable in heavy-fermion materials like UTe$_2$. We show that time-reversal and surface reflection symmetry breaking lead to a natural pairing tendency towards a triplet pair density wave state, which naturally can account for broken particle-hole symmetry.

cond-mat.supr-con↗

Convergence Analysis and Numerical Studies for Linearly Elastic Peridynamics with Dirichlet-Type Boundary Conditions

The nonlocal models of peridynamics have successfully predicted fractures and deformations for a variety of materials. In contrast to local mechanics, peridynamic boundary conditions must be defined on a finite volume region outside the body. Therefore, theoretical and numerical challenges arise in order to properly formulate Dirichlet-type nonlocal boundary conditions, while connecting them to the local counterparts. While a careless imposition of local boundary conditions leads to a smaller effective material stiffness close to the boundary and an artificial softening of the material, several strategies were proposed to avoid this unphysical surface effect. In this work, we study convergence of solutions to nonlocal state-based linear elastic model to their local counterparts as the interaction horizon vanishes, under different formulations and smoothness assumptions for nonlocal Dirichlet-type boundary conditions. Our results provide explicit rates of convergence that are sensitive to the compatibility of the nonlocal boundary data and the extension of the solution for the local model. In particular, under appropriate assumptions, constant extensions yield $\frac{1}{2}$ order convergence rates and linear extensions yield $\frac{3}{2}$ order convergence rates. With smooth extensions, these rates are improved to quadratic convergence. We illustrate the theory for any dimension $d\geq 2$ and numerically verify the convergence rates with a number of two dimensional benchmarks, including linear patch tests, manufactured solutions, and domains with curvilinear surfaces. Numerical results show a first order convergence for constant extensions and second order convergence for linear extensions, which suggests a possible room of improvement in the future convergence analysis.

math.AP↗

Low-density superconductivity in SrTiO$_3$ bounded by the adiabatic criterion

SrTiO$_3$ exhibits superconductivity for carrier densities $10^{19}-10^{21}$ cm$^{-3}$. Across this range, the Fermi level traverses a number of vibrational modes in the system, making it ideal for studying dilute superconductivity. We use high-resolution planar-tunneling spectroscopy to probe chemically-doped SrTiO$_3$ across the superconducting dome. The over-doped superconducting boundary aligns, with surprising precision, to the Fermi energy crossing the Debye energy. Superconductivity emerges with decreasing density, maintaining throughout the Bardeen-Cooper-Schrieffer (BCS) gap to transition-temperature ratio, despite being in the anti-adiabatic regime. At lowest superconducting densities, the lone remaining adiabatic phonon van Hove singularity is the soft transverse-optic mode, associated with the ferroelectric instability. We suggest a scenario for pairing mediated by this mode in the presence of spin-orbit coupling, which naturally accounts for the superconducting dome and BCS ratio.

cond-mat.supr-con↗

DAGs with No Curl: An Efficient DAG Structure Learning Approach

Recently directed acyclic graph (DAG) structure learning is formulated as a constrained continuous optimization problem with continuous acyclicity constraints and was solved iteratively through subproblem optimization. To further improve efficiency, we propose a novel learning framework to model and learn the weighted adjacency matrices in the DAG space directly. Specifically, we first show that the set of weighted adjacency matrices of DAGs are equivalent to the set of weighted gradients of graph potential functions, and one may perform structure learning by searching in this equivalent set of DAGs. To instantiate this idea, we propose a new algorithm, DAG-NoCurl, which solves the optimization problem efficiently with a two-step procedure: 1) first we find an initial cyclic solution to the optimization problem, and 2) then we employ the Hodge decomposition of graphs and learn an acyclic graph by projecting the cyclic graph to the gradient of a potential function. Experimental studies on benchmark datasets demonstrate that our method provides comparable accuracy but better efficiency than baseline DAG structure learning methods on both linear and generalized structural equation models, often by more than one order of magnitude.

cs.LG↗

Pull Request Decision Explained: An Empirical Overview

Context: Pull-based development model is widely used in open source, leading the trends in distributed software development. One aspect which has garnered significant attention is studies on pull request decision - identifying factors for explanation. Objective: This study builds on a decade long research on pull request decision to explain it. We empirically investigate how factors influence pull request decision and scenarios that change the influence of factors. Method: We identify factors influencing pull request decision on GitHub through a systematic literature review and infer it by mining archival data. We collect a total of 3,347,937 pull requests with 95 features from 11,230 diverse projects on GitHub. Using this data, we explore the relations of the factors to each other and build mixed-effect logistic regression models to empirically explain pull request decision. Results: Our study shows that a small number of factors explain pull request decision with the integrator same or different from the submitter as the most important factor. We also noted that some factors are important only in special cases e.g., the percentage of failed builds is important for pull request decision when continuous integration is used.

cs.SE↗

SumGNN: Multi-typed Drug Interaction Prediction via Efficient Knowledge Graph Summarization

Thanks to the increasing availability of drug-drug interactions (DDI) datasets and large biomedical knowledge graphs (KGs), accurate detection of adverse DDI using machine learning models becomes possible. However, it remains largely an open problem how to effectively utilize large and noisy biomedical KG for DDI detection. Due to its sheer size and amount of noise in KGs, it is often less beneficial to directly integrate KGs with other smaller but higher quality data (e.g., experimental data). Most of the existing approaches ignore KGs altogether. Some try to directly integrate KGs with other data via graph neural networks with limited success. Furthermore, most previous works focus on binary DDI prediction whereas the multi-typed DDI pharmacological effect prediction is a more meaningful but harder task. To fill the gaps, we propose a new method SumGNN: knowledge summarization graph neural network, which is enabled by a subgraph extraction module that can efficiently anchor on relevant subgraphs from a KG, a self-attention based subgraph summarization scheme to generate a reasoning path within the subgraph, and a multi-channel knowledge and data integration module that utilizes massive external biomedical knowledge for significantly improved multi-typed DDI predictions. SumGNN outperforms the best baseline by up to 5.54\%, and the performance gain is particularly significant in low data relation types. In addition, SumGNN provides interpretable prediction via the generated reasoning paths for each prediction.

cs.LG↗

PanGu-$α$: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation

Large-scale Pretrained Language Models (PLMs) have become the new paradigm for Natural Language Processing (NLP). PLMs with hundreds of billions parameters such as GPT-3 have demonstrated strong performances on natural language understanding and generation with \textit{few-shot in-context} learning. In this work, we present our practice on training large-scale autoregressive language models named PanGu-$α$, with up to 200 billion parameters. PanGu-$α$ is developed under the MindSpore and trained on a cluster of 2048 Ascend 910 AI processors. The training parallelism strategy is implemented based on MindSpore Auto-parallel, which composes five parallelism dimensions to scale the training task to 2048 processors efficiently, including data parallelism, op-level model parallelism, pipeline model parallelism, optimizer model parallelism and rematerialization. To enhance the generalization ability of PanGu-$α$, we collect 1.1TB high-quality Chinese data from a wide range of domains to pretrain the model. We empirically test the generation ability of PanGu-$α$ in various scenarios including text summarization, question answering, dialogue generation, etc. Moreover, we investigate the effect of model scales on the few-shot performances across a broad range of Chinese NLP tasks. The experimental results demonstrate the superior capabilities of PanGu-$α$ in performing various tasks under few-shot or zero-shot settings.

cs.CL↗

Size-Sieving Separation of Hard-Sphere Mixtures through Cylindrical Pores

The collision dynamics of hard spheres and cylindrical pores is solved exactly, which is the minimal model for a regularly porous membrane. Nonequilibrium event-driven molecular dynamics simulations are used to show that the permeability $P$ of hard spheres of size $σ$ through cylinderical pores of size $d$ follow the hindered diffusion mechanism due to size exclusion as $P \propto (1-σ/d)^2$. Under this law, the separation of binary mixtures of large and small particles exhibits a linear relationship between $α^{-1/2}$ and $P^{-1/2}$, where $α$ and $P$ are the selectivity and permeability of the smaller particle, respectively. The mean permeability through polydisperse pores is the sum of permeabilities of individual pores, weighted by the fraction of the single pore area over the total pore area.

cond-mat.soft↗

Threshold Changeable Secret Sharing Scheme and Its Application to Group Authentication

Group oriented applications are getting more and more popular in mobile Internet and call for secure and efficient secret sharing (SS) scheme to meet their requirements. A $(t,n)$ threshold SS scheme divides a secret into $n$ shares such that any $t$ or more than $t$ shares can recover the secret while less than $t$ shares cannot. However, an adversary, even without a valid share, may obtain the secret by impersonating a shareholder to recover the secret with $t$ or more legal shareholders. Therefore, this paper uses linear code to propose a threshold changeable secret sharing (TCSS) scheme, in which threshold should increase from $t$ to the exact number of all participants during secret reconstruction. The scheme does not depend on any computational assumption and realizes asymptotically perfect security. Furthermore, based on the proposed TCSS scheme, a group authentication scheme is constructed, which allows a group user to authenticate whether all users are legal group members at once and thus provides efficient and flexible m-to-m authentication for group oriented applications.

cs.CR↗

On Controllability and Persistency of Excitation in Data-Driven Control: Extensions of Willems' Fundamental Lemma

Willems' fundamental lemma asserts that all trajectories of a linear time-invariant system can be obtained from a finite number of measured ones, assuming that controllability and a persistency of excitation condition hold. We show that these two conditions can be relaxed. First, we prove that the controllability condition can be replaced by a condition on the controllable subspace, unobservable subspace, and a certain subspace associated with the measured trajectories. Second, we prove that the persistency of excitation requirement can be relaxed if the degree of a certain minimal polynomial is tightly bounded. Our results show that data-driven predictive control using online data is equivalent to model predictive control, even for uncontrollable systems. Moreover, our results significantly reduce the amount of data needed in identifying homogeneous multi-agent systems.

eess.SY↗

Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training Approach

Fine-tuned pre-trained language models (LMs) have achieved enormous success in many natural language processing (NLP) tasks, but they still require excessive labeled data in the fine-tuning stage. We study the problem of fine-tuning pre-trained LMs using only weak supervision, without any labeled data. This problem is challenging because the high capacity of LMs makes them prone to overfitting the noisy labels generated by weak supervision. To address this problem, we develop a contrastive self-training framework, COSINE, to enable fine-tuning LMs with weak supervision. Underpinned by contrastive regularization and confidence-based reweighting, this contrastive self-training framework can gradually improve model fitting while effectively suppressing error propagation. Experiments on sequence, token, and sentence pair classification tasks show that our model outperforms the strongest baseline by large margins on 7 benchmarks in 6 tasks, and achieves competitive performance with fully-supervised fine-tuning methods.

cs.CL↗

Implementation of Polygonal Mesh Refinement in MATLAB

We present a simple and efficient MATLAB implementation of the local refinement for polygonal meshes. The purpose of this implementation is primarily educational, especially on the teaching of adaptive virtual element methods.

math.NA↗

An asymptotically compatible treatment of traction loading in linearly elastic peridynamic fracture

Meshfree discretizations of state-based peridynamic models are attractive due to their ability to naturally describe fracture of general materials. However, two factors conspire to prevent meshfree discretizations of state-based peridynamics from converging to corresponding local solutions as resolution is increased: quadrature error prevents an accurate prediction of bulk mechanics, and the lack of an explicit boundary representation presents challenges when applying traction loads. In this paper, we develop a reformulation of the linear peridynamic solid (LPS) model to address these shortcomings, using improved meshfree quadrature, a reformulation of the nonlocal dilitation, and a consistent handling of the nonlocal traction condition to construct a model with rigorous accuracy guarantees. In particular, these improvements are designed to enforce discrete consistency in the presence of evolving fractures, whose {\it a priori} unknown location render consistent treatment difficult. In the absence of fracture, when a corresponding classical continuum mechanics model exists, our improvements provide asymptotically compatible convergence to corresponding local solutions, eliminating surface effects and issues with traction loading which have historically plagued peridynamic discretizations. When fracture occurs, our formulation automatically provides a sharp representation of the fracture surface by breaking bonds, avoiding the loss of mass. We provide rigorous error analysis and demonstrate convergence for a number of benchmarks, including manufactured solutions, free-surface, nonhomogeneous traction loading, and composite material problems. Finally, we validate simulations of brittle fracture against a recent experiment of dynamic crack branching in soda-lime glass, providing evidence that the scheme yields accurate predictions for practical engineering problems.

math.NA↗