Search arXiv⌕ Search

arXiv subjects

Yuxiang Yao

Publications and source records attributed to Yuxiang Yao.

5 recordsLinked to original sources

GraphVQ: Structure-Aware Autoregressive Decoding over Context-Quantized Graph Tokens

Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges spanning beyond the serialization window cannot be emitted in one pass--so one-pass autoregressive generators systematically under-produce cycles--and a single global condition cannot tell candidate edges apart. GraphVQ removes both obstacles: node contexts--features plus a local edge mask under multi-order breadth-first serialization--are quantized into a shared codebook by a VQ-VAE with BCE-calibrated Bernoulli edge decoding, and a second-stage structure-aware decoder emits the global adjacency conditioned on token-derived pair features, whose necessity over any global-summary condition is formalized in a scoped impossibility result. The tokenizer reconstructs node features at 0.86--0.99 accuracy and decodes local edges at AUROC >= 0.89 (ECE <= 0.007). Under one same-split protocol on four datasets, pair conditioning improves orbit MMD 0.248 -> 0.174 on PROTEINS and 3.4x on a ring stress test, and vanishes on a random-label control--the signature of attribute--topology coupling--so the gain is claimed exactly where attributes carry edge-relevant signal. GraphVQ ranks first among learned generators on PROTEINS, ties for first on SYN-COMM, and improves orbit MMD 2.7--17x over one-stage generation on three datasets, with seed-level bootstrap intervals confirming the rankings are not seed noise; on MUTAG the unweighted edge target under-generates and is reported as such. These results locate the structural control of autoregressive graph generation in the granularity of the condition: pair-level token context turns a quantized vocabulary into a usable capacity axis for distribution-faithful graph generation and future token-level pretraining.

cs.LG↗

RACE: Relation-Level Counterfactual Explanations for Heterogeneous Graph Neural Networks

Counterfactual explanations of graph neural networks identify edge deletions that flip a prediction. On heterogeneous graphs, however, existing methods first collapse the graph into untyped edges, so they cannot answer the question a domain expert actually asks: which relation type drives this prediction? We present RACE (Relation-Aware Counterfactual Explanations), which gives this question an exact, per-instance answer. For every explained instance, an exhaustive search over relation subsets returns the certified minimum relation-deletion set that flips the prediction -- or an explicit report that no such deletion exists; each relation-level answer is then refined into a typed edge set within the attributed relations, verified on the discrete model by single-edge restoration. The relation-level answer is exact and deterministic given the frozen backbone, whereas soft-mask baselines vary by 6-8 pp in success rate across runs differing only in random ordering. On ACM, a Cora-derived graph, and ogbn-mag, RACE improves counterfactual success rate over the strongest baseline by up to +2.7 pp while deleting fewer edges, and attains the highest success rate among all same-task baselines on every dataset; the advantage reproduces across four backbones on ogbn-arXiv and on DBLP, with cross-seed relation-set agreement up to 0.89. A synthetic study with known generating mechanisms confirms that the search recovers the relation the trained model actually relies on -- and reports infeasibility rather than fabricating an attribution when the model has learned none -- so the explanations stay trustworthy exactly where explanations matter.

cs.LG↗

Where Privacy Belongs: Placement Diagnosis and Certified Selection for Private Counterfactual Explanations on Graphs

Counterfactual explanations for graph neural networks (GNNs) find the minimal intervention that flips a node's prediction--but computing one requires reading sensitive graph structure, and releasing it discloses that structure. Both existing placements fail. Privatizing the graph before explaining corrupts the target on exactly the borderline nodes needing recourse, manufacturing spurious flips that flip the privatized graph but not the true one. Explaining on the clean graph and perturbing the released explanation resists certification: re-auditing the standard heuristic shows an implied full-release budget of 573--753 on Cora and 256 on CiteSeer--orders of magnitude beyond its advertised budget--with worst-case single-entry leakage at AUC 1.0. We propose PrivCFS, which replaces certification-by-optimization with certification-by-construction: counterfactual selection over a fixed, data-independent candidate universe--edge interventions from a public prior graph, feature interventions from a public schema--whose no-op semantics give neighboring graphs the same output support. A validity-gated, clipped utility of global sensitivity $Δu \le 1$ released through the exponential mechanism gives pure $\varepsilon$-DP for the complete released object, composable over queries--to our knowledge the first such guarantee on graphs. Privacy noise is the cheapest stage: at $\varepsilon$=8 the release retains 94--97% of its support-restricted non-private optimum on the recourse population and 83--95% on the general one; the optimal edge-inference audit attains AUC 0.50 on average and 0.59 worst-pair, versus the heuristic's worst entry 1.0; and transfers to a 15K-node graph at 0.96 valid rate. The dominant cost is a measurable, monotone price in public disclosure, readable off one table before any budget is spent--turning explanation privacy from an accounting risk into a purchasable decision.

cs.LG↗

A Seifert-van Kampen Theorem and the Frobenius Action on Tame Fundamental Groups

Let $X = P^1_{\mathbb{F}_p}-B$, where $B$ is a divisor with $n$ distinct geometric points, and view $X$ as a $\mathbb{F}_q$-variety with $q = p^r$ for some $r$, we then obtain a short exact sequence of tame fundamental groups: \[1\to π_1^t(X_{\overline{\mathbb{F}_q}})\to π_1^t(X_{\mathbb{F}_q})\to \mathrm{Gal}(\overline{\mathbb{F}_q}/\mathbb{F}_q)\to 1.\] This gives rise to an action of $\mathrm{Gal}(\overline{\mathbb{F}_q}/\mathbb{F}_q)$ on $π_1^t(X_{\overline{\mathbb{F}_q}})$ once an $\mathbb{F}_q$-point in $X_{\mathbb{F}_q}$ is fixed. Using Harbater's formal patching, we prove a version of the Seifert-van Kampen theorem, which further yields a purely algebraic description of the action of $\mathrm{Gal}(\overline{\mathbb{F}_q}/\mathbb{F}_q)$ on the $n$ generators of $π_1^t(X_{\overline{\mathbb{F}_q}})$ assigned to each geometric point of $B_{\overline{\mathbb{F}_q}}$. Based on this, we give a purely algebraic computation of $π_1^t(X_{\overline{\mathbb{F}_q}})$, and thereby obtain an explicit description of the tame fundamental group of $X$.

math.AG↗

The "Galois Correspondence" for n-Stacks

We prove an essentially surjective Galois-correspondence-like functor for $n$-stacks. More specifically, it gives an essentially surjective functor from the $\infty$-category of $n$-stacks of finite sets with an action of the fundamental group of $X$ to the $\infty$-category of Deligne-Mumford $n$-stacks finite étale over a connected scheme $X$.

math.AT↗