Search arXiv⌕ Search

arXiv subjects

Amit Nautiyal

Publications and source records attributed to Amit Nautiyal.

2 recordsLinked to original sources

Attribution Without a Second Pass: Inline Per-Sample Gradient Provenance at ~1% Overhead

Data attribution methods used in practice (TRAK, LoGRA, EK-FAC) are post-hoc: after training they make a second pass over the training set to recompute per-sample gradients, repeated per checkpoint when ensembled. Traceprop avoids that pass by recording projected per-sample gradients inline on the training backward pass. A Kronecker-factored sketch scales from a single tracked layer to every layer without materializing a dense projection matrix. On LoRA fine-tunes of GPT-2 and Pythia models up to 2.8B on one NVIDIA L4, inline logging costs 0.30-1.08% of wall-clock time at last-block scope and stays under 1% (0.79%) even when tracking every layer of Pythia-1B. Against LogIX, the closest inline-capable competitor, the factored sketch is 2.0-4.1x cheaper at equal storage, a gap that grows with tracked scope and is significant at every scope tested, while matching or exceeding LogIX's attribution quality at matched storage. Building the attribution-ready store inline is 60-242x cheaper than one post-hoc pass and 301-1211x cheaper than a five-checkpoint TRAK ensemble, with recorded gradients matching autograd exactly. Because each stored gradient carries source-file lineage, the same pass also produces EU AI Act Article 26 audit trails.

cs.LG↗

ChunkRank: Model-Aware Text Chunking and Abstention-Aware Answer Selection for LLM Pipelines

We present ChunkRank, an open-source Python library that derives chunk boundaries from a target model's tokenizer and context window, and selects an answer among candidates produced independently per chunk. It ships a validated registry of 90 models across 15 providers and six answer-selection methods, and needs only three core dependencies. For chunking, ChunkRank avoids context-window overflow automatically from the model name, whereas character-based splitters overflow or waste the budget, and a fidelity study across 11 languages shows why token-exact budgets matter beyond English. For answer selection we report a negative result: on NaturalQuestions, TriviaQA and HotpotQA, with extractive and generative readers, no content-based ranker reliably beats taking the first non-empty answer. The reason is reader abstention on chunks that lack the answer, not answer position. A long-context baseline shows that chunking matches single-call reading on single-hop questions, so ChunkRank targets small-window and beyond-window settings. Code, registry and evaluation harness are released.

cs.CL↗