Search arXiv⌕ Search

arXiv · 2609.32007

Interactive Proofs of Proximity for Model Evaluation

Abstract

We study interactive proofs of proximity (IPPs) for model evaluation, where a resource-limited verifier interacts with an untrusted prover, typically the model owner, to certify statistical properties of a model under an unknown input distribution. Our formulation separates sampling the input distribution from querying the model and evaluating its output; distinguishes real audit data (black-box sampling) from generated data (chosen-randomness, or gray-box, access to the sampler); and allows the prover and verifier to use different evaluators. We focus on doubly-sublinear IPPs, where both the verifier and honest prover use sublinear resources, and on (weighted) Hamming weight properties. For ordinary Hamming weight, we give a tolerant doubly-sublinear IPP. For completeness and soundness radii $\varepsilon_c<\varepsilon_f$ and gap $g=\varepsilon_f-\varepsilon_c$, a logarithmic-round instantiation uses $\widetilde{O}(1/g)$ verifier queries and $O(1/g^2)$ honest-prover queries, improving the cubic dependence of Amir, Goldreich, and Rothblum (ITCS 2025). We prove matching query lower bounds up to polylogarithmic factors. For distribution-weighted Hamming weight, black-box sampling requires $Θ(1/g^2)$ verifier samples but only $\widetilde{O}(1/g)$ evaluations; the quadratic sample complexity is necessary in the interior regime. With chosen-randomness access, the problem reduces to ordinary Hamming weight, yielding $\widetilde{O}(1/g)$ calls and evaluations. If the parties' evaluators disagree arbitrarily on a $ρ$-fraction of the distribution and by at most $γ$ elsewhere, our protocols remain doubly sublinear whenever $g>2κ$, where $κ=ρ+(1-ρ)γ$. Applications include auditing accuracy, group fairness, calibration, harmlessness, usefulness, and average-case robustness.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Geoffroy Couteau, Nikolas Melissaris, Tamara Paris. 2026-09-25. Interactive Proofs of Proximity for Model Evaluation. https://arxiv.org/abs/2609.32007

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Deterministic Depth-4 PIT and Normalization

In this paper, we initiate the study of deterministic PIT for $Σ^{[k]}ΠΣΠ^{[δ]}$ circuits over fields of any characteristic, where $k$ and $δ$ are bounded. Our main result is a deterministic polynomial-time black-box PIT algorithm for $Σ^{[3]}ΠΣΠ^{[δ]}$ circuits, under the additional condition that one of the summands at the top $Σ$ gate is squarefree. Our techniques are purely algebro-geometric: they do not rely on Sylvester--Gallai-type theorems, and our PIT result holds over arbitrary fields. The core of our proof is based on the normalization of algebraic varieties. Specifically, we carry out the analysis in the integral closure of a coordinate ring, which enjoys better algebraic properties than the original ring.

cs.CC↗

Local Search with Correlated Randomness

How much does an algorithm's running-time distribution under independent randomness reveal about its behavior when independence is no longer guaranteed? We study sources satisfying $ν[w]\le DP[w]^s$ for every finite prefix $w$, where $P$ is an independent reference law, $0<s\le1$, and $D\ge1$. The constraint controls complete-prefix probabilities while allowing individual choices to be predictable, even fully determined by the past. For retry tasks, all deterministic history-dependent selectors have the same independent-source running-time law. Yet two orders have worst-case failure probabilities $1$ and $\exp[-Θ(n)]$ at the same linear deadline under the same source constraint. We identify a static priority rule that is optimal at every deadline and every $D$. For the standard local walk on a $k$-CNF with at least $r$ true literals per clause under some assignment, $k/2<r<k$, we determine the sharp source threshold $s_*$. At and above it, the expected flip count is $O_{k,r}(\min\{L^3,L/(s-s_*)\})$, where $L=h+\log D+1$, $h$ is the initial Hamming distance to that assignment, and $L/0=\infty$. The bound allows arbitrary clause overlap and history-dependent clause selection. Matching instances admit one source forcing this delay with probability one for every selector. At criticality and fixed $D$, the delay is cubic despite a linear independent-source expectation. Variable-depth prefix covers, together with classical tree max-flow/min-cut, yield an exact criterion for restoring exponential tails by restarting on the same tape. We synthesize updates and restarts for explicit finite-state processes. Under a sufficient prefix guarantee, we also obtain noisy predecessor search with error at most $η$ and expected query count polynomial in the correct leaf's depth and $\log(D/η)$, without knowing the depth or tree height.

cs.CC↗

Sample Complexity of Equivariant Reinforcement Learning

Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.

cs.CC↗