Searcharxiv⌕ Search

arXiv subjects

Subhabrata Majumdar

Publications and source records attributed to Subhabrata Majumdar.

At least 19 recordsLinked to original sources

Defensive Sufficiency in a Stackelberg Model of AI Security

Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection. We derive completion-time bounds and extend the analysis to growing attack surfaces, repairs that generalize across related attacks, and multiple discovery mechanisms. These results distinguish eventual protection against each fixed attack from complete protection at a single time. We then formulate a defender-led Stackelberg game in which the defender invests in proactive discovery and reactive repair, anticipating the attacker's choice of search effort. We characterize the least-cost allocation that deters attack and the equilibrium regimes in which the defender funds neither capability, one capability, or both. Numerical experiments illustrate these regimes and show how faster repair can reduce compromise duration without reducing compromise probability.unified theory of performance limits in generative language models.

cs.CR↗

Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models

Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.

cs.LG↗

GAIA meets LZ: a high velocity tail-tale sign for dark matter

The interpretation of high-energy nuclear recoils in direct-detection experiments can depend sensitively on the high-velocity tail of the galactic dark matter (DM) distribution. The LUX-ZEPLIN (LZ) experiment has recently reported an event, LZ230616, with nuclear recoil energy 248 $\pm$ 23 (stat) $\pm$ 23 (sys) keV, while the Gaia mission has provided an unprecedented view of the kinematics and mass distribution of the Milky Way. We bring these two probes together to investigate the inelastic DM interpretation of LZ230616, using a velocity distribution inferred from Gaia-informed galactic model. Compared to the Standard Halo Model, the resulting distribution has a lower escape speed and a significantly suppressed high-velocity tail, with a higher local density. We find that while the exothermic interpretation is largely insensitive to this change, the endothermic interpretation requires substantially larger scattering cross sections and smaller mass splittings. Furthermore, using simulation-based study of the gravitational effects of the Large Magellanic Cloud, we show that the high-velocity tail can extend the allowed endothermic parameter space towards lower DM masses and larger splittings. Our results demonstrate how combining direct searches for DM detection with precision measurements of galactic dynamics can qualitatively affect the particle-physics interpretation of high-energy nuclear recoils.

hep-ph↗

Mapping the Milky Way in Six Dimensions: A contiguous, homogenised, phase-space catalogue of \textit{Gaia} DR3 tracers up to 250 kpc

We present a comprehensive, homogenised catalogue of $32{,}552{,}876$ stellar sources with full 6D phase-space information, spanning Galactocentric distances from the inner galaxy ($\gtrsim 5$\,kpc) to the outer halo ($\sim\!250$\,kpc). By cross-matching \textit{Gaia} DR3 astrometry with spectrophotometric distances and line-of-sight velocities from 14 surveys (including \textsc{desi}, \textsc{sdss-boss}, \textsc{apogee}, \textsc{lamost}, \textsc{galah}, and \textsc{ges}), we achieve median fractional distance uncertainties of $\approx2.7\%$ for $d_{\rm helio}\le15$\,kpc, and $\approx29\%$ for $15 17$) and metal-poor ($[\mathrm{Fe/H}] < -1$) tracers, including RR~Lyrae, blue horizontal-branch, and K~giant stars, constitute the largest well-calibrated sample of outer-halo tracers to date. This unique dataset is ideal for multi-component mass modelling, galactic archaeology, and dark matter searches across the Milky Way.

astro-ph.GA↗

Limits of Reliability and Scaling in Language Models

Large language models (LLMs) are trained and evaluated as though perfect reliability is achievable for any task given sufficient scale. We show that this assumption is information-theoretically unjustified. Every generative task has a reliability ceiling that no model can exceed, determined by how much output uncertainty is resolvable from observable context. The gap decomposes into a resolvable component closable with additional context and a subjective component inherent to task ambiguity. Autoregressive generation further degrades this ceiling at a rate governed by the task's dependency kernel, which quantifies inter-token correlations in the output. From these two primitives, we derive a first-principles scaling law where LLM performance is bottlenecked by the scarcer resource: training data or model capacity. This law recovers the Chinchilla scaling law as a special case and provides a structural account of when scaling improves reliability. Beyond scaling, our framework unifies diverse practical phenomena, such as the benefits of retrieval-augmentation and the spectral mechanics of catastrophic forgetting. Our work formalizes the resource-complexity tradeoffs that govern model performance across domains, offering a unified theory of performance limits in generative language models.

cs.CL↗

Disentangling Stellar Mass and Environmental Effects on eROSITA AGN Activity using multi-wavelength data from GAMA, WISE, GALEX, and DESI Legacy Survey

We use a stellar-mass complete sample of $\sim 36{,}000$ galaxies from GAMA/eFEDS to test whether AGN activity depends on environment independent of stellar mass. To isolate this effect, we introduce $δ_{\rm rank}$, a novel local overdensity metric defined within narrow stellar mass bins. Our main result is that X-ray-selected AGNs show a clear environmental trend, the AGN fraction rises from $\sim 4\%$ in low-density regions to $\sim 6\%$ in high-density regions, with a sharp transition at $δ_{\rm rank} \sim 0.4$. This environmental dependence is absent for WISE-selected and broad-line AGN, and Eddington ratios show no variation across environments. We also found that for galaxies with $\mathrm{SFR} > 15\,M_\odot\,\mathrm{yr}^{-1}$, X-ray AGN fraction is approximately 5%, 60% and 20% in low, mid and high $δ_{\rm rank}$ galaxies. We show that relationship between AGN fraction and SFR varies dramatically with both environment and selection method. We conclude that environment on scales of $\sim 5h^{-1}\,\mathrm{Mpc}$ regulates X-ray AGN triggering efficiency.

astro-ph.GA↗

No Unique Minimizer, No Problem: On the Consistency of Robust Neural Classifiers

Neural network classifiers trained by cross-entropy minimization are highly sensitive to label noise and adversarial contamination. While robust alternatives offer bounded influence and resistance to corruption, their statistical foundations in the deep learning setting are insufficient due to a fundamental difficulty: neural parameterizations are non-identifiable, so the population loss minimizer is an equivalence class of parameters, not a unique point. We develop a consistency theory for robust neural classifiers based on the S-divergence family that requires no identifiability assumption. Casting training as stochastic optimization over a non-identifiable parameter space, we prove that empirical S-divergence minimizers converge to the population-optimal equivalence class under mild regularity conditions, and verify these conditions for three architecture choices. We further establish that limit points of the robust training algorithm are stationary points of the empirical objective. Experiments on vision and language benchmark datasets confirm that S-divergence training maintains clean-data accuracy while exhibiting performance competitive with existing robust methods.

cs.LG↗

OTAP: Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories

Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current evaluation metrics reduce such a trajectory to a binary success flag, compare it against a reference by exact matching, or delegate judgment to another language model. A success flag cannot distinguish a sound solution from one that succeeds by luck, and says nothing about why a failed run went wrong. Exact matching penalizes plans that are valid but reordered or decomposed differently from the reference. We reframe trajectory evaluation as a distance between the agent's execution graph and a set of valid solution graphs, and instantiate it via an unbalanced fused Gromov-Wasserstein transport problem over attributed dependency graphs. The resulting score, termed OTAP (Optimal Transport for Agentic Planning), is a pseudo-metric that is provably invariant to dependency-preserving reorderings and has bounded sensitivity to redundant steps. Its unbalanced marginals handle missing or hallucinated steps without forcing a match, and its soft coupling accommodates variation in plan granularity. On controlled perturbations and three public benchmarks, OTAP separates valid from invalid trajectories in a regime where semantics-only metrics score below chance. Its advantage tracks the fidelity of the dependency graph: largest where edges follow from operator semantics, smallest where they are inferred from free text. Where a formal verifier exists, strict surface metrics predict validity better than OTAP does, which places OTAP in open-ended domains where no verifier is available.

cs.AI↗

Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety

A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing such harmful fine-tuning while retaining benign adaptability remains difficult: the only prior method with an explicit curvature certificate, spectral deformation, inflates curvature globally and thereby obstructs benign adaptation along with harmful adaptation. We propose HarmAlign, which applies function-preserving spectral deformation along a estimated contrastive activation subspace. We derive finite-sample bounds for the estimated subspace energy and the resulting local harmful-distribution curvature lower bound. A stability--progress dichotomy for constant-step gradient descent turns the certified curvature into conditional convergence-rate control. Empirically, within a fixed-architecture, finite-budget first-order threat model, HarmAlign blocks direct fine-tuning and three data- or objective-adaptive attacks across a hazardous-knowledge relearning setting and a harmful-assistance fine-tuning setting, while the protected benign tasks remain trainable. The block persists across the tested first-order optimizer variants over every attack checkpoint, and under out-of-distribution harmful fine-tuning, and it extends to important cases in our threat model: accidental safety degradation and emergent misalignment.

cs.LG↗

Probing the baryonic--dark matter connection in galaxy clusters using X-rays with gated recurrent unit neural networks

Accurate cluster mass measurements are crucial for cosmology, yet conventional hydrostatic equilibrium (HSE) methods can suffer from systematic biases, particularly in dynamically disturbed systems. We present a gated recurrent unit (GRU) based deep learning framework for predicting three-dimensional mass profiles of galaxy clusters from spherically averaged intra-cluster medium (ICM) radial profiles. By treating ICM profiles as sequential data, the GRU captures radial dependencies and naturally handles profiles with different radial samplings. We train and validate the model using high-resolution hydrodynamical simulations from The Three Hundred Project, achieving unbiased mass predictions with a typical 1$σ$ scatter of $\sim$5% over most of the cluster region, significantly improving upon HSE estimates. The model provides radius-dependent uncertainty estimates and remains robust against variations in data quality and cluster morphology. When trained jointly on independent simulation suites (GIZMO-SIMBA and GADGET-X), it successfully generalises across both simulations. Feature importance analysis shows that enclosed gas mass is the dominant predictor, with pressure and temperature providing additional information on the radial mass distribution. We further apply the GRU model to X-ray observations of the REXCESS and X-COP cluster samples from XMM-Newton and compare the inferred mass profiles with HSE estimates. The HSE masses are systematically lower than the GRU predictions for the higher-mass X-COP sample, while the REXCESS sample shows mass differences that are close to zero on average. This work provides a data-driven framework for cluster mass inference that bridges simulations and observations and can be extended to multi-wavelength datasets, including Sunyaev-Zel'dovich and optical observations.

astro-ph.CO↗

Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing

Large Language Models (LLMs) and high-dimensional perception networks increasingly rely on parameter-efficient fine-tuning (PEFT) to adapt to diverse operational contexts. However, standard methods like LoRA are structurally limited by a monolithic bottleneck, making them highly susceptible to gradient warfare. Interleaved multi-task streams may trigger destructive optimization feedback, collapsing adapter weights into unspecialized averages. While recent spatial partitioning methods have introduced block-wise isolation, they remain trapped in static topologies, unable to adapt to dynamic task-switching or environmental sensor failure. In this work, we introduce Localized LoRA-MoE, a unified framework that fuses localized spatial blocking with dynamic, context-conditioned routing. We propose and evaluate two novel architectural paradigms: Block-Wise LoRA-MoE (Centralized Macro-Routing), which modulates the entire structural grid via a monolithic context signal, and Cell-Wise LoRA-MoE (Decentralized Micro-Routing), which empowers every coordinate cell in the matrix grid with autonomous, localized expert gating. Through a comprehensive suite of benchmarks, ranging from high-dimensional SVD matrix simulations and real-world tabular transformations to spatial vision perception under sensor degradation, we demonstrate that both architectures resolve optimization deadlocks inherent in static baselines. Our empirical results establish that decentralized cell-level gating achieves complete statistical parity with an omniscient global coordinator, providing a robust "gradient firewall" that protects surviving pathways from fault-propagated corruption. Our proposals consistently outperform static baselines, offering a scalable and parameter-efficient solution for dynamic model adaptation across granular coordinate fields and shifting operational regimes.

cs.LG↗

PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations

Multi-turn jailbreak attacks on large language models (LLMs) reveal a mismatch in current guardrails: they operate on individual turns, while attacks unfold as trajectories across conversations. We propose a shift from content to dynamics, modeling conversations as paths in representation space and asking whether adversarial intent is encoded early in their geometry. We introduce PsychoPass, a framework that extracts geometric features from conversation trajectories in embedding space to predict a potential attack before harmful content is produced. These features achieve near-perfect performance in naïve classifiers, which is largely explained by the inclusion of number of turns as a feature. After removing this confound, a smaller but consistent geometric signal remains, with classification performance that does not depend meaningfully on encoder choice. Crucially, this signal appears early in the conversation: attack outcomes remain above chance from short prefixes alone, more reliably than baseline guardrails. A supporting theoretical analysis explains these findings via a decomposition of length and shape, a detection bound based on prefix length, and encoder invariance. Together, these results show that adversarial conversations leave an early, representation-robust geometric fingerprint suitable for online monitoring.

cs.CR↗

Next-Billion AI Index: The compass for AI utility and adoption in the global majority

Generative AI assessments remain dominated by frontier capability benchmarks that often fail to capture whether systems can be sustainably deployed, adapted, and trusted in locally grounded and infrastructure-constrained settings. This paper introduces the Next Billion AI Index (nexbax), which we believe is the first diagnostic framework to treat economic viability, operational deployability, and governance alignment as co-equal determinants of AI utility in next-billion-user contexts. Rather than treating usefulness as a single outcome, nexbax operationalizes the preconditions for useful AI through 10 dimensions organized under three themes: Effective Efficiency, Operational Practicality, and Societal Integrity. These dimensions assess whether systems are economically viable, deployable under infrastructure and workflow constraints, and aligned with local needs, user expectations, and collaborative development practices. We pair the framework with rubrics for weak, moderate, and strong performance, and conduct a formative expert evaluation through eleven semi-structured interviews with founders, developers, product leaders, and technical practitioners building AI systems for next-billion markets. Participants found the index useful for reasoning about adoption trade-offs and effective at capturing factors shaping real-world AI uptake -- particularly cost, usability, reliability, and trust. They also identified the need for contextual explanations, domain-specific evidence, and broader stakeholder validation. Nexbax is therefore proposed not as a universal score of social value, but as a diagnostic for artificial useful intelligence: a way to make visible the technical, economic, and governance properties that make inclusive AI deployment more viable.

cs.CY↗

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability

This paper establishes a rigorous measurement science for AI agent reliability, providing a foundational framework for quantifying consistency under semantically preserving perturbations. By leveraging $U$-statistics for output-level reliability and kernel-based metrics for trajectory-level stability, we offer a principled approach to evaluating agents across diverse operating conditions. Our proposal highlights the important distinction between the core capability and execution robustness of an agent, showing that minor task-level variations can induce complete strategy breakdowns despite the agent possessing the requisite knowledge for the task. We validate our framework through extensive experiments on three agentic benchmarks, demonstrating that trajectory-level consistency metrics provide far greater diagnostic sensitivity than traditional pass@1 rates. By providing the mathematical tools to isolate where and why agents deviate, we enable the identification and rectification of architectural concerns that hinder the deployment of agents in high-stakes, real-world environments.

cs.AI↗

Limits of Convergence-Rate Control for Open-Weight Safety

Open-weight foundation models can be fine-tuned for harmful purposes after release, yet no existing training resistance methods provide theoretical guarantees. Treating these interventions as convergence-rate control problems allows us to connect optimization speed to the spectral structure of model weights. We leverage this insight to develop a novel understanding of convergence rate control through spectral reparameterization and derive an algorithm, SpecDef, that can both provably and empirically slow first- and second-order optimization in non-adversarial settings. In adversarial settings, we establish a fundamental limit on a broad class of convergence rate control methods including our own: an attacker with sufficient knowledge can restore fast convergence at a linear increase in model size. In order to overcome this limitation, future works will need to investigate methods that are not equivalent to controlling convergence rate.

math.OC↗

Red Teaming AI Red Teaming

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. Current AI red teaming efforts focus predominantly on individual model vulnerabilities while overlooking the broader sociotechnical systems and emergent behaviors that arise from complex interactions between models, users, and environments. To address this deficiency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecycle, and micro-level model red teaming. Drawing on cybersecurity experience and systems theory, we further propose a set of six recommendations. In these, we emphasize that effective AI red teaming requires multifunctional teams that examine emergent risks, systemic vulnerabilities, and the interplay between technical and social factors.

cs.AI↗

Deriving accurate galaxy cluster masses using X-ray thermodynamic profiles and graph neural networks

Precise determination of galaxy cluster masses is crucial for establishing reliable mass-observable scaling relations in cluster cosmology. We employ graph neural networks (GNNs) to estimate galaxy cluster masses from radially sampled profiles of the intra-cluster medium (ICM) inferred from X-ray observations. GNNs naturally handle inputs of variable length and resolution by representing each ICM profile as a graph, enabling accurate and flexible modeling across diverse observational conditions. We trained and tested GNN model using state-of-the-art hydrodynamical simulations of galaxy clusters from The Three Hundred Project. The mass estimates using our method exhibit no systematic bias compared to the true cluster masses in the simulations. Additionally, we achieve a scatter in recovered mass versus true mass of about 6%, which is a factor of six smaller than obtained from a standard hydrostatic equilibrium approach. Our algorithm is robust to both data quality and cluster morphology and it is capable of incorporating model uncertainties alongside observational uncertainties. Finally, we apply our technique to XMM-Newton observed galaxy cluster samples and compare the GNN derived mass estimates with those obtained with $Y_{\rm SZ}$-M$_{500}$ scaling relations. Our results provide strong evidence, at 5$σ$ level, for a mass-dependent bias in SZ derived masses, with higher mass clusters exhibiting a greater degree of deviation. Furthermore, we find the median bias to be $(1-b)=0.85_{-0.14}^{+0.34}$, albeit with significant dispersion due to its mass dependence. This work takes a significant step towards establishing unbiased observable mass scaling relations by integrating X-ray, SZ and optical datasets using deep learning techniques, thereby enhancing the role of galaxy clusters in precision cosmology.

astro-ph.CO↗

Localized LoRA: A Structured Low-Rank Approximation for Efficient Fine-Tuning

Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, offer compact and effective alternatives to full model fine-tuning by introducing low-rank updates to pre-trained weights. However, most existing approaches rely on global low rank structures, which can overlook spatial patterns spread across the parameter space. In this work, we propose Localized LoRA, a generalized framework that models weight updates as a composition of low-rank matrices applied to structured blocks of the weight matrix. This formulation enables dense, localized updates throughout the parameter space without increasing the total number of trainable parameters. We provide a formal comparison between global, diagonal-local, and fully localized low-rank approximations, and show that our method consistently achieves lower approximation error under matched parameter budgets. Experiments on both synthetic and practical settings demonstrate that Localized LoRA offers a more expressive and adaptable alternative to existing methods, enabling efficient fine-tuning with improved performance.

cs.LG↗