Search arXivSearch

arXiv subjects

Feng Xie

Publications and source records attributed to Feng Xie.

At least 19 recordsLinked to original sources

Transitions between bulk and interfacial fracture in diamond/$c$BN heterostructures

Whether an initially crack-free heterostructure fails at its interface or within an adjoining phase is controlled by the relative cohesion of competing atomic planes, but how interfacial chemistry, crystallographic orientation, and intermixing reshape this competition remains unclear. Here, we combine density functional theory (DFT) with a fine-tuned atomistic foundation model to resolve tensile fracture in coherent diamond/cubic boron nitride (cBN) heterostructures. The resulting potential reproduces independent DFT tensile responses, including an unseen (001) interface orientation. Interfacial termination, orientation, and diffusion-induced intermixing jointly determine fracture resistance and fracture-plane selection. C-N-bonded (111) and C-B-bonded (001) remain interface-controlled throughout the investigated intermixing range, whereas pristine C-B-bonded (111) fractures at a neighboring B-N plane inside cBN because the interface is more strongly bound. Increasing the diffusion fraction from 0 to 50.0% causes a nonlinear decrease in fracture strength from 50.3 to 15.8 GPa and drives a bulk-to-interface transition through three regimes: cBN fracture up to 4.86%, configuration-dependent competition between 7.6 and 10.1%, and interfacial fracture at 12.5% and above. Atom-resolved stress fields and DFT separation energetics show that this transition is associated with stress relocation and a reversal in the relative cohesion of competing planes, while electron localization analysis connects the cohesion hierarchy to termination- and orientation-dependent bonding. These results establish an atomistic framework for controlling fracture resistance and fracture pathways in strongly bonded heterostructures.

cond-mat.mtrl-sci

Hessian Rank Constraint for Learning Structure of Nonlinear Latent Variable Models

Uncovering latent variables and their causal relations from observed data is a fundamental yet challenging problem. Existing methods often rely on restrictive assumptions, such as linear relations or invertible mixing functions. To better address this problem under general nonlinear mixing procedures, we propose a condition called the cross-Hessian Rank Constraint (HRC), which serves as a primitive rank-based tool for nonlinear latent causal discovery. In particular, we show that a rank-based property arises from the cross-Hessian of the observed-data log-density in the nonlinear case, revealing information about the latent variables, and reduces to the Tetrad constraints in the linear Gaussian case. More specifically, when two groups of observed variables are d-separated by a set of lower-dimensional latent variables, the rank of this cross-Hessian is equal to the dimension of the latent variables, under a mild affine derivative assumption on the conditional log-density derivatives. This assumption can be naturally satisfied when the noise level is low or the relevant nonlinearity is moderate. As a downstream application, we instantiate HRC in the pure one-factor measurement setting for locating latent variables and recovering their causal structure up to Markov equivalence. Experimental results on synthetic and real-world datasets support the theoretical claims.

cs.LG

GLLH EM Invisible Cloak With Novel Front Branching And Without Exceeding Light Speed Violation

In this paper, for the first time in the world, we propose a new electromagnetic (EM) cloak without superluminal propagation and without time delay. Using Global and Local (GL) electromagnetic non-scattering modeling and inversion with a distinctive class of materials a_{αβ}\log ^α(b_{αβ}/h) h^β(GLLH Cloak), the named GLLH invisible cloak is developed in arXiv:1005.3999V1 in 2010. After 16 years, this paper is version 3 of arXiv:1005.3999V1 . The refractive index of the GLLH cloak material is greater than or equal to one. Spherical harmonic analysis and the GL modeling and inversion method are used to simulate the electromagnetic wave propagation through the GLLH cloak without superluminal effects and without time delay. The novel EM wave propagation and front branching in the GLLH cloak obtained by GL EM modeling are presented. Find an invisible cloak is the non scattering problem. Using Global and Local (GL) electromagnetic non-scattering modeling and inversion with a distinctive class of materials is one method in our paper in arXiv:1005.3999V1 in 2010. Another method is some 0 to R1 radial coordinate transformations with superluminal and infinite phase velocity. From the GLLH invisible cloak, we discovered positive space and negative space and invisible science. Using a new negative infinity to 0 quasi topological transformation, we discovered new isotropic electromagnetic invisible GLHUA sphere; anisotropic electromagnetic invisible GLHUA cloak. In the GLLH cloak, the wave front is curved as a crescent like and without superluminal and time delay. Open question: Can we construct a 3D Kakeya set where line segments of length represent the vector field of EM wave ray-tracing propagation through GLLH, GLHUA cloaks, and the GLHUA sphere? All copyright and patent of the GLLH EM cloaks and GL modeling and inversion methods are reserved by authors.

physics.gen-ph

Testing the Validity of Instrumental Variable Sets in Causal Additive Models with Non-Constant Effects

Instrumental variable (IV) methods are powerful for causal effect estimation with unmeasured confounding, but in practice researchers often face a set of candidate IVs whose validity is difficult to determine from observational data. This paper studies the problem of testing the validity of IV sets under Causal Additive Models with Non-Constant Effects (CAM-NCE). To address this problem, we propose a testable condition, termed the Cross Auxiliary-based independence Test (CAT) condition, for assessing IV set validity from observational data. We show that, under the completeness condition, if the CAT condition is violated, the corresponding set cannot be a valid IV set. Furthermore, under a cross distributional non-degeneracy condition, we establish that the CAT condition becomes both necessary and sufficient for characterizing valid IV sets under CAM-NCE. We then extend the CAT condition to settings with covariates and develop a practical finite-sample algorithm for testing the validity of candidate IV sets. Extensive experiments on synthetic data and three real-world datasets demonstrate the effectiveness and practical utility of the proposed method.

stat.ME

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks

Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interface. Existing approaches often rely on resource-intensive conditioning, using either heavyweight branches or costly source concatenation. Is there any efficient way to model editing intent? Thus, we introduce GRNEdit, a lightweight two-stage framework. GRN inspires our approach by encoding visual semantics through combinations of bits. Through task-specific fine-tuning, we take this representation further and recast editing semantics as local retain-or-flip decisions over individual bits. Source information is consequently modeled as coordinate-wise evidence supporting the observed binary states, while the GRN backbone remains responsible for resolving their global composition into coherent generative semantics. In Stage I, a compact encoder translates discrete source codes into continuous evidence signals, which GRN assimilates throughout binary refinement. Inspired by null-prompt training for classifier-free guidance, we further assign the null condition an editing-specific meaning: an empty instruction denotes no edit and is supervised through source reconstruction. This identity pathway not only implicitly strengthens evidence utilization and content preservation in Stage I, but also produces a source-preserving state in the same representation space as the edited state. Stage II can therefore directly compare each edited state with its source-preserving counterpart and use their discrepancy to revise unresolved target-bit decisions. Trained on only 0.6M pairs with less than 3\% conditioning parameters, GRNEdit-2B and GRNEdit-8B achieve scores of 4.03 and 4.18 on OpenVE-Bench. The 2B model outperforms multiple 14B open-source editors, while the 8B model performs on par with leading open-source editors.

cs.CV

Global uniform regularity and large time behavior of solutions to three dimensional compressible MHD equations with vanishing vertical magnetic resistivity in half space

This paper aims to establish the global regularity and large time behavior of solutions to the three-dimensional (3D) compressible magnetohydrodynamics (MHD) equations with vanishing vertical magnetic resistivity in the upper half-space with no-slip boundary condition on velocity and perfectly conducting boundary condition on magnetic field. By exploiting anisotropic Sobolev inequalities and elaborated estimates, we are able to achieve global-in-time uniform regularity estimates of solutions, which are independent of small vertical resistivity coefficient $\varepsilon$. These uniform global regularity estimates allow us to pass to the limit as \(\varepsilon \to 0\) and obtain the convergence to the corresponding MHD system without vertical magnetic resistivity globally in time. Moreover, the $H^{1}\left(\mathbb{R}_{+}^{3}\right)$ decay rates of solutions to the original system are also derived based on the detailed analysis on semigroup of the related linearized operator in half space and the uniform energy estimates achieved. In contrast to the decay estimates for the limit system obtained by Gao and Xie \cite{JDE}, the present decay estimate exhibits a slower rate, attributable to the absence of higher-order normal derivative estimates, which results from the occurrence of boundary layers. Finally, we combine both sets of regularity and decay results to rigorously prove an explicit time-uniform $L^2$ convergence rate of order $\varepsilon^{\frac{1}{4}}$ for the vanishing vertical magnetic resistivity limit process.

math.AP

Local Causal Structure Learning in the Presence of Latent Variables and Selection Bias

Discovering the direct causes and effects of a target variable from observational data is a fundamental problem in causal discovery, with broad applications in domains such as gene regulatory analysis and biomedical research. Existing causal discovery methods either learn a global causal structure, which incurs substantial computational cost, or assume the absence of latent variables and selection bias, assumptions that are often violated in real-world settings. Motivated by these challenges, we study local causal structure learning in the presence of latent variables and selection bias. Specifically, we first characterize a local region that enables target-specific causal discovery without recovering the entire global structure. We then establish a theoretical bridge between causal information learned from the observed distribution induced on this local region and the corresponding information in the global causal structure. Building on these foundations, we propose LoCaLS, a local causal structure learning algorithm that is sound and complete under standard assumptions and identifies the same direct causes and effects of a target variable as those identifiable by global causal discovery methods, while allowing for latent variables and selection bias. Extensive experiments on random and real-world structures demonstrate that the proposed method consistently achieves higher structural accuracy than existing local methods while requiring substantially less computational effort than state-of-the-art global methods. Furthermore, applications to two real-world gene expression datasets reveal biologically plausible target-specific causal structures, demonstrating its practical applicability in large-scale biological data analysis.

cs.LG

A fault-tolerant quantum blockchain deployed on commercial telecommunications network

Popularized by the Bitcoin cryptocurrency, blockchain technology establishes a decentralized digital framework that utilizes cryptographic and consensus protocols to secure data against unauthorized modification. Consequently, blockchain has found broad adoption across diverse fields, including finance, data management, healthcare, and digital asset governance. In the quantum computing era, a paramount objective for blockchain is to preserve its foundational advantages of cryptographic integrity and decentralized fault-tolerant resilience. In principle, quantum digital signatures and quantum Byzantine agreement protocols offer foundational security guarantees and tolerate up to one-half of malicious nodes for blockchain. However, the practical realization of such a quantum-enhanced blockchain remains a significant and multifaceted challenge. Here, we propose and experimentally demonstrate a fully operational hybrid quantum blockchain architecture built on photonic integrated circuits and deployed over commercially available classical telecommunications infrastructure. The system achieves a fault tolerance of nearly one-half, surpassing the classical limit, while reaching consensus on a timescale of seconds. A deployed food traceability application validates the practicality of the proposed architecture, achieving a throughput of approximately 500 transactions per second. This work establishes a foundation for practical quantum blockchains, enabling secure, scalable, and decentralized information processing in the emerging quantum era.

quant-ph

Experimental demonstration of scalable quantum blockchain with exponentially superior quantum communication complexity

To secure modern distributed digital infrastructures, quantum blockchains exploit quantum resources to achieve information-theoretic security and surpass the classical one-third fault-tolerance bound. However, existing high-fault-tolerant protocols face a fundamental scalability challenge: the blockchain trilemma imposes either exponential communication complexity or experimentally demanding multipartite entanglement. Here, we experimentally demonstrate a scalable quantum blockchain protocol based on weak coherent states that achieves an exponential reduction in quantum communication complexity. The protocol employs a circular quantum Byzantine agreement mechanism that preserves information-theoretic security while avoiding multipartite entanglement. We implement this protocol on a photonic integrated circuit platform, realizing a six-node network over commercially available telecommunication infrastructure. Compared with previous schemes, the protocol requires less than 4% of the quantum communication resources. Leveraging this advantage, we further demonstrate a quantum-secured token exchange application achieving a throughput of 805.3 transactions per second with zero failures. These results establish a practical pathway toward scalable quantum blockchain.

quant-ph

Local Covariate Selection for Average Causal Effect Estimation without Pretreatment and Causal Sufficiency Assumptions

We study the problem of selecting covariates for unbiased estimation of the total causal effect.Existing approaches typically rely on global causal structure learning over all variables, or on strong assumptions such as causal sufficiency - where observed variables share no latent confounders - or the pretreatment assumption, which limits covariates to those unaffected by the treatment or outcome. These requirements are often unrealistic in practice, and global learning becomes computationally prohibitive in high-dimensional settings.To address these challenges, we propose a novel local learning method for covariate selection in nonparametric causal effect estimation that avoids both the pretreatment and causal sufficiency assumptions. We first characterize a local boundary that contains at least one valid adjustment set whenever one exists for identifying the causal effect, and then develop local identification procedures to efficiently search within this boundary.We prove that the proposed method is sound and complete. Experiments on multiple synthetic datasets and two real-world datasets show that our approach achieves accurate causal effect estimation while substantially improving computational efficiency.

stat.ML

A Recursive Decomposition Framework for Causal Structure Learning in the Presence of Latent Variables

Constraint-based causal discovery is widely used for learning causal structures, but heavy reliance on conditional independence (CI) testing makes it computationally expensive in high-dimensional settings. To mitigate this limitation, many divide-and-conquer frameworks have been proposed, but most assume causal sufficiency, i.e., no latent variables. In this paper, we show that divide-and-conquer strategies can be theoretically generalized beyond causal sufficiency to settings with latent variables. Specifically, we propose a recursive decomposition framework, termed DiCoLa, that enables divide-and-conquer causal discovery in the presence of latent variables. It recursively decomposes the global learning task into smaller subproblems and integrates their solutions through a principled reconstruction step to recover the global structure. We theoretically establish the soundness and completeness of the proposed framework. Extensive experiments on synthetic data demonstrate that our approach significantly improves computational efficiency across a range of causal discovery algorithms, while experiments on a real-world dataset further illustrate its practical effectiveness.

cs.LG

PRIME: Prototype-Driven Multimodal Pretraining for Cancer Prognosis with Missing Modalities

Multimodal self-supervised pretraining offers a promising route to cancer prognosis by integrating histopathology whole-slide images, gene expression, and pathology reports, yet most existing approaches require fully paired and complete inputs. In practice, clinical cohorts are fragmented and often miss one or more modalities, limiting both supervised fusion and scalable multimodal pretraining. We propose PRIME, a missing-aware multimodal self-supervised pretraining framework that learns robust and transferable representations from partially observed cohorts. PRIME maps heterogeneous modality embeddings into a unified token space and introduces a shared prototype memory bank for latent-space semantic imputation via patient-level consensus retrieval, producing structurally aligned tokens without reconstructing raw signals. Two complementary pretraining objectives: inter-modality alignment and post-fusion consistency under structured missingness augmentation, jointly learn representations that remain predictive under arbitrary modality subsets. We evaluate PRIME on The Cancer Genome Atlas with label-free pretraining on 32 cancer types and downstream 5-fold evaluation on five cohorts across overall survival prediction, 3-year mortality classification, and 3-year recurrence classification. PRIME achieves the best macro-average performance among all compared methods, reaching 0.653 C-index, 0.689 AUROC, and 0.637 AUROC on the three tasks, respectively, while improving robustness under test-time missingness and supporting parameter-efficient and label-efficient adaptation. These results support missing-aware multimodal pretraining as a practical strategy for prognosis modeling in fragmented clinical data settings.

cs.LG

EpiScreen: Early Epilepsy Detection from Electronic Health Records with Large Language Models

Epilepsy and psychogenic non-epileptic seizures often present with similar seizure-like manifestations but require fundamentally different management strategies. Misdiagnosis is common and can lead to prolonged diagnostic delays, unnecessary treatments, and substantial patient morbidity. Although prolonged video-electroencephalography is the diagnostic gold standard, its high cost and limited accessibility hinder timely diagnosis. Here, we developed a low-cost, effective approach, EpiScreen, for early epilepsy detection by utilizing routinely collected clinical notes from electronic health records. Through fine-tuning large language models on labeled notes, EpiScreen achieved an AUC of up to 0.875 on the MIMIC-IV dataset and 0.980 on a private cohort of the University of Minnesota. In a clinician-AI collaboration setting, EpiScreen-assisted neurologists outperformed unaided experts by up to 10.9%. Overall, this study demonstrates that EpiScreen supports early epilepsy detection, facilitating timely and cost-effective screening that may reduce diagnostic delays and avoid unnecessary interventions, particularly in resource-limited regions.

cs.CL

Identifiability of causal effects with non-Gaussianity and auxiliary covariates

Assessing causal effects in the presence of unmeasured confounding is challenging. Although auxiliary variables, such as instrumental variables, are commonly used to identify causal effects, they are often unavailable in practice due to stringent and untestable conditions. To address this issue, previous researches have utilized linear structural equation models to show that the causal effect is identifiable when noise variables of the treatment and outcome are both non-Gaussian. In this paper, we investigate the problem of identifying the causal effect using the auxiliary covariate and non-Gaussianity from the treatment. Our key idea is to characterize the impact of unmeasured confounders using an observed covariate, assuming they are all Gaussian. We demonstrate that the causal effect can be identified using a measured covariate, and then extend the identification results to the multi-treatment setting. We further develop a simple estimation procedure for estimating causal effects and derive a $\sqrt{n}$-consistent estimator. Finally, we evaluate the performance of our estimator through simulation studies and apply our method to investigate the effect of the trade on income.

stat.ME

Testability of Instrumental Variables in Additive Nonlinear, Non-Constant Effects Models

We address the issue of the testability of instrumental variables derived from observational data. Most existing testable implications are centered on scenarios where the treatment is a discrete variable, e.g., instrumental inequality (Pearl, 1995), or where the effect is assumed to be constant, e.g., instrumental variables condition based on the principle of independent mechanisms (Burauel, 2023). However, treatments can often be continuous variables, such as drug dosages or nutritional content levels, and non-constant effects may occur in many real-world scenarios. In this paper, we consider an additive nonlinear, non-constant effects model with unmeasured confounders, in which treatments can be either discrete or continuous, and propose an Auxiliary-based Independence Test (AIT) condition to test whether a variable is a valid instrument. We first show that, under the completeness condition, if the candidate instrument is valid, then the AIT condition holds. Moreover, we illustrate the implications of the AIT condition and demonstrate that, under certain additional conditions, the AIT condition is necessary and sufficient to detect all invalid IVs. We also extend the AIT condition to include covariates and introduce a practical testing algorithm. Experimental results on both synthetic and three different real-world datasets show the effectiveness of our proposed condition.

stat.ME

HeartAgent: An Autonomous Agent System for Explainable Differential Diagnosis in Cardiology

Heart diseases remain a leading cause of morbidity and mortality worldwide, necessitating accurate and trustworthy differential diagnosis. However, existing artificial intelligence-based diagnostic methods are often limited by insufficient cardiology knowledge, inadequate support for complex reasoning, and poor interpretability. Here we present HeartAgent, a cardiology-specific agent system designed to support a reliable and explainable differential diagnosis. HeartAgent integrates customized tools and curated data resources and orchestrates multiple specialized sub-agents to perform complex reasoning while generating transparent reasoning trajectories and verifiable supporting references. Evaluated on the MIMIC dataset and a private electronic health records cohort, HeartAgent achieved over 36% and 20% improvements over established comparative methods, in top-3 diagnostic accuracy, respectively. Additionally, clinicians assisted by HeartAgent demonstrated gains of 26.9% in diagnostic accuracy and 22.7% in explanatory quality compared with unaided experts. These results demonstrate that HeartAgent provides reliable, explainable, and clinically actionable decision support for cardiovascular care.

cs.CL

Local Learning for Covariate Selection in Nonparametric Causal Effect Estimation with Latent Variables

Estimating causal effects from nonexperimental data is a fundamental problem in many fields of science. A key component of this task is selecting an appropriate set of covariates for confounding adjustment to avoid bias. Most existing methods for covariate selection often assume the absence of latent variables and rely on learning the global network structure among variables. However, identifying the global structure can be unnecessary and inefficient, especially when our primary interest lies in estimating the effect of a treatment variable on an outcome variable. To address this limitation, we propose a novel local learning approach for covariate selection in nonparametric causal effect estimation, which accounts for the presence of latent variables. Our approach leverages testable independence and dependence relationships among observed variables to identify a valid adjustment set for a target causal relationship, ensuring both soundness and completeness under standard assumptions. We validate the effectiveness of our algorithm through extensive experiments on both synthetic and real-world data.

cs.LG

Confounded Causal Imitation Learning with Instrumental Variables

Imitation learning from demonstrations usually suffers from the confounding effects of unmeasured variables (i.e., unmeasured confounders) on the states and actions. If ignoring them, a biased estimation of the policy would be entailed. To break up this confounding gap, in this paper, we take the best of the strong power of instrumental variables (IV) and propose a Confounded Causal Imitation Learning (C2L) model. This model accommodates confounders that influence actions across multiple timesteps, rather than being restricted to immediate temporal dependencies. We develop a two-stage imitation learning framework for valid IV identification and policy optimization. In particular, in the first stage, we construct a testing criterion based on the defined pseudo-variable, with which we achieve identifying a valid IV for the C2L models. Such a criterion entails the sufficient and necessary identifiability conditions for IV validity. In the second stage, with the identified IV, we propose two candidate policy learning approaches: one is based on a simulator, while the other is offline. Extensive experiments verified the effectiveness of identifying the valid IV as well as learning the policy.

cs.LG