Searcharxiv⌕ Search

arXiv subjects

Dipesh Tharu Mahato

Publications and source records attributed to Dipesh Tharu Mahato.

4 recordsLinked to original sources

ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not establish that its generator is reliable. ProxyGuard controls both errors using prespecified bounded risks and a sealed target set. Named-release mode corrects for multiplicity and certifies specific releases. Direct shared-target mode evaluates independent mechanism draws on a common target, lower-bounds their favorable-score rate, and subtracts a bound on favorable scores contributed by invalid releases. Conditional on the target, release scores are independent, yielding a finite-sample mechanism-reliability guarantee without independent target batches or assumptions on release-level $p$-value dependence. We show that the mean-only penalty is sharp and derive a smooth-score certificate with additive target concentration. In a registered three-requirement study, direct mode raises power from 5.6\% to 64.2\% at reliability 0.95, while named mode remains stronger under high-signal evidence. Prospective audits span full-pipeline Rice--TVAE, which retrains on every draw, and a non-tabular text mechanism.

cs.LG↗

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards for dual-use biology assistants at the action and answer levels. The framework reconstructs the access path, separates provider refusals from downstream actions, and selects thresholds under an intervention budget. A frozen fresh-generation study satisfies its criterion on Claude Opus 4.5, but both passing configurations share one upstream provider effect; none passes on Gemini 2.5 Flash. The fixed Opus policies retain positive selectivity on 104 previously unused released-label pairs, but both fail a 20\% matched-benign constraint. At the answer level, no joint-scoring verifier qualifies on a response-disjoint 7,200-judgment holdout. A fresh 8,640-judgment factorial experiment finds separate gains from criterion isolation and ordinal representation, with a positive interaction between them; requiring explicit localization lowers aggregate accuracy under a strict no-repair schema. The evidence supports prospective action-level selectivity and identifies verifier interface effects, but not calibrated selective access, verified content removal, or biological-risk reduction.

cs.CY↗

Early Warning Signals for OpenVLA Failure under Visual Distribution Shift

Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activations contain signals associated with the steps before failure. We freeze the policy, record one MLP activation per LIBERO-10 step, and fit two linear monitors. Occlusion reduces task success from $57\%$ to $17\%$. Within failed matched-reset trajectories, a layer-16 logistic probe attains AUROC $0.972$ and AUPRC $0.352$, whereas action disagreement attains AUROC $0.496$. Without refitting, the occlusion-trained probe reaches AUROC $0.689$ on failed camera-jitter episodes. In a calibration check, however, the same layer-16 monitor averages 3.32 warning onsets per clean episode. This contrast shows that strong retrospective discrimination does not imply operationally quiet warning behavior. Because fitting and evaluation share tasks, resets, and seed, these results establish retrospective separability rather than prediction on independent episodes.

cs.CV↗

TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift

Deep neural networks often achieve high accuracy, but ensuring their reliability under adversarial and distributional shifts remains a pressing challenge. We propose TriGuard, a unified safety evaluation framework that combines (1) formal robustness verification, (2) attribution entropy to quantify saliency concentration, and (3) a novel Attribution Drift Score measuring explanation stability. TriGuard reveals critical mismatches between model accuracy and interpretability: verified models can still exhibit unstable reasoning, and attribution-based signals provide complementary safety insights beyond adversarial accuracy. Extensive experiments across three datasets and five architectures show how TriGuard uncovers subtle fragilities in neural reasoning. We further demonstrate that entropy-regularized training reduces explanation drift without sacrificing performance. TriGuard advances the frontier in robust, interpretable model evaluation.

cs.LG↗