Search arXiv⌕ Search

arXiv subjects

Md Anas Biswas

Publications and source records attributed to Md Anas Biswas.

3 recordsLinked to original sources

Input-Layer Starvation: Why Per-Layer Pruning Breaks IoT Intrusion Detectors

Intrusion detectors for small Internet-of-Things (IoT) devices are usually compressed by pruning and judged by overall accuracy. We show that this hides a severe class-level failure, find its cause, and give low-overhead prevention and repair. On CICIoT2023, a two-layer convolutional detector pruned with uniform layer-wise magnitude pruning at 80% sparsity loses 16 points of accuracy but half of its macro-F1, the mean per-class F1 (0.542 to 0.271 over five independently trained models); 17 of 34 classes are materially damaged. Remaining weight count does not explain it: a perceptron and a transformer pruned to the same or fewer weights lose at most 0.096. The first layer does. It has 192 weights; uniform pruning leaves 38, 46% of its 64 filters lose every input weight, and fine-tuning under that starvation leaves the running means of the first normalisation layer displaced by up to 0.8 standard deviations in a few surviving channels, on which the deployed model collapses. Protecting those 192 weights, or pruning globally at the same sparsity, prevents the collapse (loss 0.013); recomputing the normalisation statistics on unlabelled training data, with no weight changed, repairs it (loss 0.039) and returns the false-alert rate to 33% (dense 29%). Damage shows a strong increasing dose-response in first-layer sparsity, starving a perceptron's input layer reproduces the collapse, and the pattern holds on TON_IoT. The failure is misattribution and false alerts, not silent evasion: on validation-selected blind spots, uniformly pruned detectors misattribute 72% of the traffic, against 50% with the first layer protected and 47% for the dense model.

cs.CR↗

The Calibrated Deepfake Trust Score (CDTS): Competence-Coupled Trust Degradation Across Deepfake Detectors

In moderation, provenance, and verification pipelines a deepfake detector's output probability is read as a degree of trust, so its calibration matters as much as raw accuracy. We reframe deepfake detection as a calibrated, self-auditing trust instrument, the Calibrated Deepfake Trust Score (CDTS), and identify what governs its trustworthiness. Our central finding is a competence-trust coupling with a sharp division: the raw score's miscalibration tracks discriminative competence almost perfectly (r = -0.98, -0.98, -0.95 across two convolutional networks and a CLIP vision transformer), and a calibrator deployed without target labels fails with competence just as tightly (r = -0.98 on the primary detector, for isotonic, Platt, and beta calibrators alike). Given target labels, by contrast, any well-specified calibrator repairs any detector, including inverted ones: in-domain calibratability is not competence-limited, and the trust failure is a distribution-shift phenomenon concentrated exactly on the low-competence generators that motivate deployment. Explanation faithfulness rises and falls on the same competence axis. We reach this conclusion after uncovering, and correcting, a tie-handling degeneracy in the standard equal-mass expected-calibration-error (ECE) estimator that fabricates strong spurious competence-calibration coupling on tie-heavy calibrated scores; the same degeneracy invalidates calibration-equity gaps we previously reported, a caution for fairness auditing. Competence is trackable without labels: batch predictive entropy flags generators with high deployed calibration error at ROC-AUC 0.99, and routing source-batches on label-free competence beats confidence-based routing precisely in the low-competence regimes the coupling identifies, while confidence regains the advantage where competence is high. Trust scoring must be competence-aware; CDTS is the mechanism.

cs.CR↗

Confidently Wrong: Severity-Aware Calibration of Prompt-Injection Detectors under Attack Shift

Prompt-injection detectors are deployed as guards: a model scores an input and a downstream system trusts or blocks it on that score. I study the confidence of these scores, not only their accuracy, when the attack distribution shifts away from the clean benchmark on which the operating point was chosen. I evaluate three released detectors, ProtectAI-v2 and two Prompt-Guard-2 checkpoints, at a single source-calibrated threshold that I freeze and transport across five shifts. I report a severity metric S, how confident a detector is on the attacks it misses, alongside the false-negative rate and discrimination. Across every shift and every detector, severity on the missed attacks stays between 0.99 and 1.00 while the false-negative rate ranges from 0.01 to 0.97: when these detectors miss, they miss with near-certainty. All three confidently pass indirect behavior-hijack injection, a blind spot unanimous across two vendors and a fourfold size range. Standard pooled calibration error does not register this; one detector it rates well-calibrated, at 0.06, is miscalibrated at 0.91 on the attacks alone. Run against live models, the missed injections leak the majority of working exploits, passing them at the rate they catch others. A controlled experiment traces the cause to content-keying rather than injection structure, an instruction-tuned model used as a judge shows the same hijack blind spot, and a black-box rewriter exploits the content-keying to manufacture working confident misses, most effectively on the most dangerous attack category. Code and data are public.

cs.CR↗