Search arXivSearch

arXiv subjects

Fei Fang

Publications and source records attributed to Fei Fang.

At least 19 recordsLinked to original sources

Inward and Outward Spillover Effects of One Unit's Treatment on Network Neighbors under Partial Interference

In settings where interference is present, direct effects are commonly defined as the average effect of a unit's treatment on their own outcome while fixing the treatment status or probability among interfering units. Spillover effects measure the average effect of a change in the latter while the individual's treatment status is kept fixed. Here, we define the average causal effect of a unit's treatment status on the outcome of their network neighbors, while fixing the treatment probability in the remaining interference set. We propose two different weighting schemes defining two causal effects: i) the outward spillover effect, which represents the average effect of a unit's treatment on their neighbors' potential outcomes, and ii) the inward spillover effect, which represents the impact of each neighbor's treatment on an individual's own potential outcome. We provide a necessary and sufficient condition for when outward and inward spillover effects differ--even in undirected networks--and show that they are equivalent under specific conditions. We provide numerous examples illustrating the conditions for equivalence or discrepancy of the two spillover effects. We then compare their Horvitz-Thompson estimators, examining their relative variance under various graph structures and structural assumptions on potential outcomes.

stat.ME

Ballot Design and Electoral Outcomes: The Role of Candidate Order and Party Affiliation

We develop a causal inference model to study how designing ballots with and without party designations impacts electoral outcomes when partisan voters rely on party-order cues to infer candidate affiliation in races without designations. If the party orders of candidates in races with and without party designations differ, these voters might cast their votes incorrectly. We identify a quasi-randomized natural experiment with contest-level treatment assignment pertaining to North Carolina judicial elections and leverage double machine learning to accurately capture the magnitude of such incorrectly cast votes. Using precinct-level election and demographic data, we estimate that 12.0% (SE: 3.6%) of Democratic partisan voters and 15.4% (SE: 4.0%) of Republican partisan voters cast their votes incorrectly due to the difference in party orders. A placebo test using judicial races with party designations shows that the flip effect disappears when party affiliation is observable on the ballot, which is consistent with the proposed causal mechanism.

stat.AP

On the Yamabe flow in a bounded domain

This paper investigates the dynamical behaviors of solutions to the Yamabe flow via the modified potential well method. We first establish the local existence and regularity of weak solutions for the flow. Several new results concerning global existence and blowup are obtained by classifying initial data into stable and unstable sets. Specifically, solutions with initial data in the stable set exist globally and extinguish in finite time, whereas those originating from unstable initial data blow up in infinite time. For certain high-energy initial data, we show that the solution decays to zero as time tends to infinity and undergoes finite-time blowup. In addition, we analyze Palais-Smale sequences to reveal the intrinsic relationship between the long-time asymptotic behavior of solutions and steady states. Finally, we derive the Pohozaev identity for the equation and prove the corresponding nonexistence theorem.

math.AP

A Unified Framework for Scalable and Robust Paper Assignment

Assigning papers to reviewers is a central challenge in the peer-review process of large academic conferences. Program chairs must balance competing objectives, including maximizing reviewer expertise, promoting diversity, and enhancing robustness to strategic manipulation, but it is challenging to do so at the modern conference scale. Existing algorithmic paper assignment approaches either fail to address all of these goals simultaneously or suffer from poor scalability. To address the limitation, we propose Robust Assignment via Marginal Perturbation (RAMP), a unified framework for large-scale peer review. Our approach formulates a linearized perturbed-maximization objective with soft constraints that flexibly balance assignment quality, diversity, and robustness while maintaining runtime efficiency. We further introduce an attribute-aware sampling procedure that converts fractional solutions into integral assignments and improves the diversity and robustness of the final assignment. On datasets with over 20,000 papers and 20,000 reviewers, RAMP runs in under 20 minutes, demonstrating its suitability for real-world deployment.

cs.SI

Human Decision-Making with AI Assistance under Correlated Features

Humans increasingly make decisions with AI assistance; for example, doctors may follow AI-recommended diagnostic tests and base their diagnoses on the results. A natural question is which tests should AI recommend to balance short-term decision quality and long-term human learning when different features (e.g., test results) are correlated. While prior work establishes that stationary policies that recommend the same tests repeatedly are optimal when features are independent, we prove that feature correlations lead such policies to perform arbitrarily poorly. Instead, we prove that any optimal policy must follow an explore-then-commit structure; initially, the AI should offer diverse tests so humans can learn accurate feature coefficients, then the AI should commit to a single set of tests, with exploration length that depends on the degree of feature correlation. We prove that computing the optimal policy is NP-hard and derive a dynamic programming-based algorithm that finds the optimal policy for finite horizons. We additionally develop an approximation that plans for shorter horizons and appends a stationary suffix, achieving near-optimal performance. Our empirical results complement our theory by showing that stronger feature correlation leads to longer exploration phases.

cs.AI

Structural Causal Models for Extremes: an Approach Based on Exponent Measures

We introduce a new formulation of structural causal models for extremes, called the extremal structural causal model (eSCM). Unlike conventional structural causal models, where randomness is governed by a probability distribution, eSCMs use an exponent measure, an infinite-mass law that naturally arises in the analysis of multivariate extremes. Central to this framework are activation variables, which abstract the single-big-jump principle, along with additional randomization that enriches the class of eSCM laws. This formulation encompasses all possible laws of directed graphical models under the recently introduced notion of extremal conditional independence. We also identify an inherent asymmetry in eSCMs under natural assumptions, enabling the identifiability of causal directions, a central challenge in causal inference. Finally, we propose a method that utilizes this causal asymmetry and demonstrate its effectiveness in both simulated and real datasets.

math.ST

Global existence and blow-up for the Hardy-Sobolev parabolic equation in RN

In this paper, we apply a self-similar transformation to convert the parabolic equation with a Hardy term \begin{equation*} \begin{cases}u_t-Δu-μ\frac{u}{|x|^2}=|u|^{2^*-2} u & \text { in } \mathbb{R}^N \times(0, T), u(x, 0)=u_0(x) & \text { in } \mathbb{R}^N , \end{cases} \end{equation*} into the following parabolic equation \begin{equation*} \begin{cases} v_s-Δv-\frac{1}{2} y \cdot \nabla v=βv+\frac{μv}{|y|^2}+|v|^{2^*-2} v &\text { in } \mathbb{R}^N \times(0, S), \left.v\right|_{s=0}=v_0 & \text { in } \mathbb{R}^N, \end{cases} \end{equation*} where $N \geqslant 3$, $μ\in [0,(N-2)^2 /8]$ and $2^{\ast}=2N /(N-2)$. For this equation, we establish a weighted Hardy inequality. Furthermore, by virtue of the modified potential well method and Palais-Smale sequence analysis, we investigate the long-time behavior and finite-time blow-up properties of solutions to the parabolic equation.

math.AP

Variational problems related to self-similar solutions of Hardy-Sobolev heat equation in RN

In this paper, we apply a self-similar transformation to convert the parabolic equation with a Sobolev-Hardy term \begin{align*} u_t-Δu= \frac{|u|^{q-2}u}{\left|x\right|^s} & \text { in } \mathbb{R}^N \times(0, \infty), \end{align*} into the following elliptic equation \begin{equation*} -Δv-\frac{1}{2} y \cdot \nabla v=αv+ \frac{|v|^{q-2} v}{|y|^s}, \end{equation*} where $2 < q \leq 2^*(s)=\frac{2 N-2 s}{N-2}, 0 \leq s < 2, α=\frac{2-s}{2q-4}$. For this equation, we establish the weighted Hardy inequality and Sobolev inequality. Furthermore, by virtue of the variational methods, we obtain infinitely many solutions in the subcritical case, and prove the existence of solutions in the critical case. We also apply the Pohozaev identity to establish the nonexistence of solutions under certain conditions.

math.AP

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit assumptions about how users interact with models that cannot be surfaced from benchmarks alone. To make this precise, we propose a classification of assumptions into two categories: task, which can be tested from conversation data alone, and outcome, which requires outcome data and behavioral studies for testing. Critically, outcome assumptions depend on human behavior, something that even well-designed benchmarks cannot directly observe. To demonstrate the operationality of this framework, we retrospectively analyze a healthcare RCT as a case study and find that the gap naturally separates into task and outcome gaps of roughly equal size. To address this, we make two contributions: first, we propose BenchmarkCards, an artifact that documents assumptions, and second, we propose staged evaluation, a procedure that systematically tests assumptions and evaluates performance.

cs.CY

Sem-Detect: Semantic Level Detection of AI Generated Peer-Reviews

How can we distinguish whether a peer review was written by a human or generated by an AI model? We argue that, in this setting, authorship should not be attributed solely from the textual features of a review, but also from the ideas, judgments, and claims it expresses. To this end, we propose Sem-Detect, an authorship detection method for peer reviews that operationalizes this principle by combining textual features with claim-level semantic analysis. Sem-Detect compares a target review against multiple AI-generated reviews of the same paper, leveraging the observation that different AI models tend to converge on similar points, while human reviewers introduce more unique and diverse ones. As a result, Sem-Detect is able to distinguish fully AI reviews from authentic human-written ones, including those that have been refined using an LLM but still reflect human judgment. Across a dataset of over 20,000 peer reviews from ICLR and NeurIPS conferences, Sem-Detect improves over the strongest baseline by 25.5% in TPR@0.1% FPR in the binary setting. Moreover, in the three-class scenario, we empirically show that LLM refinement preserves the semantic signals of human reviews, which remain distinct from the patterns exhibited by fully AI-generated text; as a result, fewer than 3.5% of LLM-refined human reviews are misclassified as AI-generated.

cs.CL

Base Models Look Human To AI Detectors

As AI-generated text enters the real-world at scale, institutions increasingly use commercial AI-text detectors, especially in education and academic-integrity workflows. We report a surprising empirical finding about such systems: when evaluated by GPTZero and Pangram, generated text from base models is often judged overwhelmingly human, whereas text generated by their instruction-tuned counterparts is not. Building on this observation, we propose Humanization by Iterative Paraphrasing (HIP), a detector-agnostic pipeline that minimally fine-tunes a base model into a paraphraser and applies it iteratively. Compared with the baselines we test, HIP yields a stronger trade-off between semantic preservation and detector evasion on commercial detectors. Across Llama-3 and Qwen-3 families, spanning model sizes from 0.6B to 70B, HIP consistently improves detector human-likeness. Our findings suggest that current detectors are tracking artifacts of instruction tuning and local context more than any invariant notion of machine-generated text. This, in turn, calls for detector designs that model these factors more explicitly.

cs.CL

Antidistillation Fingerprinting

Model distillation enables efficient emulation of frontier large language models (LLMs), creating a need for robust mechanisms to detect when a third-party student model has trained on a teacher model's outputs. However, existing fingerprinting techniques that could be used to detect such distillation rely on heuristic perturbations that impose a steep trade-off between generation quality and fingerprinting strength, often requiring significant degradation of utility to ensure the fingerprint is effectively internalized by the student. We introduce antidistillation fingerprinting (ADFP), a principled approach that aligns the fingerprinting objective with the student's learning dynamics. Building upon the gradient-based framework of antidistillation sampling, ADFP utilizes a proxy model to identify and sample tokens that directly maximize the expected detectability of the fingerprint in the student after fine-tuning, rather than relying on the incidental absorption of the un-targeted biases of a more naive watermark. Experiments on GSM8K, OASST1, and MBPP demonstrate that ADFP achieves a significant Pareto improvement over state-of-the-art baselines, yielding stronger detection confidence with minimal impact on utility across mathematical reasoning, dialogue, and code generation, even when the student model's architecture is unknown.

cs.LG

Watermarking Game-Playing Agents in Perfect-Information Extensive-Form Games

Watermarking techniques for large language models (LLMs), which encode hidden information in the output so its source can be verified, have gained significant attention in recent days, thanks to their potential capability to detect accidental or deliberate misuse. Similar challenges involving model misuse also exist in the context of game-playing, such as when detecting the unauthorized use of AI tools in gaming platforms (e.g., cheating in online chess). In this paper, we initiate the study of how game-playing strategies can be watermarked. We show how the KGW watermark for LLMs can be adapted to watermark game-playing agents in perfect-information extensive-form games. The watermark can then be detected using a statistical test. We show that the degradation in the quality of the watermarked strategy profile, quantified by the expected utility, can be bounded, but there is a tradeoff between detectability and quality. In our experiments, we bootstrap the watermarking framework to various chess engines and demonstrate that a) the impact of the watermark on the quality of the strategy is negligible and b) the watermark can be detected with just a handful of games.

cs.GT

AI Realtor: Towards Grounded Persuasive Language Generation for Automated Copywriting

This paper develops an agentic framework that employs large language models (LLMs) for grounded persuasive language generation in automated copywriting, with real estate marketing as a focal application. Our method is designed to align the generated content with user preferences while highlighting useful factual attributes. This agent consists of three key modules: (1) Grounding Module, mimicking expert human behavior to predict marketable features; (2) Personalization Module, aligning content with user preferences; (3) Marketing Module, ensuring factual accuracy and the inclusion of localized features. We conduct systematic human-subject experiments in the domain of real estate marketing, with a focus group of potential house buyers. The results demonstrate that marketing descriptions generated by our approach are preferred over those written by human experts by a clear margin while maintaining the same level of factual accuracy. Our findings suggest a promising agentic approach to automate large-scale targeted copywriting while ensuring factuality of content generation.

cs.AI

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) has emerged as the leading approach for enhancing reasoning capabilities in large language models. However, it faces a fundamental compute and memory asymmetry: rollout generation is embarrassingly parallel and memory-light, whereas policy updates are communication-heavy and memory-intensive. To address this, we introduce PODS (Policy Optimization with Down-Sampling), which decouples rollout generation from policy updates by training only on a strategically selected subset of rollouts, maintaining learning quality while dramatically reducing update costs. We propose a principled subset selection criterion, max-variance down-sampling, that maximizes reward diversity, and provide an efficient $O(n\log n)$ implementation. Empirically, Group Relative Policy Optimization (GRPO) with PODS achieves the peak test accuracy of vanilla GRPO at least $\mathbf{1.7\times}$ faster across the different reasoning benchmarks and hardware configurations we tested.

cs.LG

Fučik spectrum for the operator with rapidly increasing weight and applications

In this paper, we study the Fučik spectrum for the operator with rapidly increasing weight, which is defined as a set $Σ$ comprising those $(α, β) \in \mathbb{R}^2$ such that \begin{equation*} \left\{\begin{array}{l} L u:=-Δu-\frac{1}{2}(x \cdot \nabla u)=αu^{+}-βu^{-}, \text{in}\ \mathbb{R}^N,\\ u\in X, \end{array}\right. \end{equation*} has a non-trivial solution $u$, where, $N\geq1$, $u^{ \pm}=\max ( \pm u, 0)$, $u=u^{+}-u^{-}$. The existence of a first nontrivial curve $\mathcal{C}$ of this spectrum, along with some of its properties (e.g., Lipschitz continuity, strict decrease and asymptotic behavior) is investigated in this paper. Our difficulty is that the problem is defined on the whole space $\mathbb{R}^N$, and therefore certain estimates do not carry over from the Fučik problem on bounded domains. As an application, we establish the multiplicity of solutions to the following problem \begin{equation*} \left\{\begin{array}{l} -Δu-\frac{1}{2}(x \cdot \nabla u)=f(x,u), \text{in}\ \mathbb{R}^N,\\ u\in X, \end{array}\right. \end{equation*} where, $N\geq1$ and the nonlinearity $f$ is asymptotically linear at zero and at infinity.

math.AP

Attention-Guided Flow-Matching for Sparse 3D Geological Generation

Constructing high-resolution 3D geological models from sparse 1D borehole and 2D surface data is a highly ill-posed inverse problem. Traditional heuristic and implicit modeling methods fundamentally fail to capture non-linear topological discontinuities under extreme sparsity, often yielding unrealistic artifacts. Furthermore, while deep generative architectures like Diffusion Models have revolutionized continuous domains, they suffer from severe representation collapse when conditioned on sparse categorical grids. To bridge this gap, we propose 3D-GeoFlow, the first Attention-Guided Continuous Flow Matching framework tailored for sparse multimodal geological modeling. By reformulating discrete categorical generation as a simulation-free, continuous vector field regression optimized via Mean Squared Error, our model establishes stable, deterministic optimal transport paths. Crucially, we integrate 3D Attention Gates to dynamically propagate localized borehole features across the volumetric latent space, ensuring macroscopic structural coherence. To validate our framework, we curated a large-scale multimodal dataset comprising 2,200 procedurally generated 3D geological cases. Extensive out-of-distribution (OOD) evaluations demonstrate that 3D-GeoFlow achieves a paradigm shift, significantly outperforming heuristic interpolations and standard diffusion baselines.

cs.CV

Selecting Decision-Relevant Concepts in Reinforcement Learning

Training interpretable concept-based policies requires practitioners to manually select which human-understandable concepts an agent should reason with when making sequential decisions. This selection demands domain expertise, is time-consuming and costly, scales poorly with the number of candidates, and provides no performance guarantees. To overcome this limitation, we propose the first algorithms for principled automatic concept selection in sequential decision-making. Our key insight is that concept selection can be viewed through the lens of state abstraction: intuitively, a concept is decision-relevant if removing it would cause the agent to confuse states that require different actions. As a result, agents should rely on decision-relevant concepts; states with the same concept representation should share the same optimal action, which preserves the optimal decision structure of the original state space. This perspective leads to the Decision-Relevant Selection (DRS) algorithm, which selects a subset of concepts from a candidate set, along with performance bounds relating the selected concepts to the performance of the resulting policy. Empirically, DRS automatically recovers manually curated concept sets while matching or exceeding their performance, and improves the effectiveness of test-time concept interventions across reinforcement learning benchmarks and real-world healthcare environments.

cs.LG