Search arXiv⌕ Search

arXiv subjects

Sleem Abdelghafar

Publications and source records attributed to Sleem Abdelghafar.

3 recordsLinked to original sources

Auditing Information Disclosure During Large-Scale Gradient-Based Training via Gradient Uniqueness

Auditing information disclosure across every datapoint during the training of LLMs is challenging. We propose a principled, attack-agnostic approach that uses mutual information to measure what the final model reveals about a datapoint's training membership. We show that this final-model disclosure is upper bounded by the sum of per-iteration gradient disclosures and that, under a reasonable set of assumptions, these gradient disclosures increase with Gradient Uniqueness (GNQ), which measures how distinguishable a datapoint's gradient is relative to other gradients in the batch. While naively computing GNQ requires forming and inverting a $P\times P$ matrix for every datapoint (for a model with $P$ parameters), we introduce Batch-Space Ghost (BS-Ghost). This efficient algorithm performs all computations in a much smaller batch space and uses ghost kernels to compute GNQ "in-run" for every datapoint in the training corpus, with minimal computational and memory overhead. Our experiments show the following: (i) GNQ predicts MIA vulnerability without the need for shadow models. (ii) Beyond membership disclosure, GNQ predicts the success of reconstruction attacks. (iii) GNQ-guided removal and retraining identify datapoints that causally contribute to disclosure. (iv) GNQ outperforms counterfactual memorization in text extraction and common-knowledge discrimination without the need for additional model training. (v) For data attribution, GNQ-guided filtering reduces emergent misalignment in Qwen2.5-7B more than baselines. Further, GNQ attributes 1000 datapoints in 17 seconds---roughly $70\times$ faster than the baselines. (vi) GNQ explains how training choices affect training-set disclosure and captures how per-datapoint disclosure emerges during training.

cs.LG↗

Scalable Attribution and Control of Model Behavior During Training

Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individual contributions difficult to distinguish. We address this ambiguity through mutual information, accounting for interference within the batch by quantifying how much the combined behavioral change reveals about each example's contribution. We show that this mutual information is a logarithmic function of Behavioral Gradient Uniqueness (BGU). BGU gives the information measure its geometric interpretation. Our Batch-Space Ghost (BS-Ghost) algorithm makes these scores practical inside the training loop through shared computation in batch space, without storing model-sized example gradients. On a complete 1,000-example Qwen2.5-7B-Instruct workload, our BS-Ghost implementation adds 27 seconds (8.0%) to 5.5 minutes of ordinary training. Removal and retraining demonstrate that BGU identifies data that causally shapes final behavior. At each training step, signed information identifies which examples strengthen or weaken the target behavior, explaining how behavior develops during training. Signed information also enables cheap intervention during training: it predicts how changing example weights will affect behavior in the next update. We then use these predictions to choose weights that steer behavior toward a desired target. This makes our framework a practical foundation for scalable oversight and verification of training pipelines and processes, helping evaluators assess model alignment, understand how it develops during training, and guide interventions that shape ongoing learning.

cs.LG↗

Privacy-Preserving AI Verification via Minimal Information Disclosure

AI verification crosses a trust boundary: a verifier must learn enough to establish an authorized claim, yet the same evidence can reveal sensitive details about the model, workload, or hardware. We introduce minimal information disclosure (MID), which designs and quantifies the information content of verifier-facing evidence itself. MID measures collateral leakage with conditional mutual information: what the release reveals about the protected property after the authorized result is known. MID is general by design: it can accommodate different verification goals, protected properties, evidence sources, and deployment constraints. To demonstrate MID's practicality, we evaluate it on four physical measurements and six verification tasks spanning execution type, hardware identity, compute scale, and model identity. These experiments use three mechanism-design variables--the evidence channel, collection policy, and release transformation--but MID is not limited to these choices and can accommodate other deployable mechanisms. Across these tasks, MID produces three releases with perfect held-out verification and zero measured collateral leakage, while the remaining tasks yield explicit privacy--utility frontiers. MID also supports ZKP-certified releases: we demonstrate our proposed linear-projection mechanism using a Groth16 zk-SNARK.

cs.CR↗