SearcharxivSearch

arXiv subjects

Samyak Jain

Publications and source records attributed to Samyak Jain.

At least 19 recordsLinked to original sources

Novel Signatures of Matter-Induced Dark Matter Decay in Large-Volume Neutrino Telescopes

Large-volume neutrino telescopes offer a unique opportunity to search for decaying dark matter through events containing a pair of energetic, highly non-collimated muon tracks emerging from a common vertex. Such events would have negligible Standard Model backgrounds and would constitute a striking signature of new physics. Conventional dark matter annihilation or decay, however, is too strongly constrained to produce an observable rate of such events. We therefore consider scenarios in which an excited dark matter state is extremely long-lived in vacuum but decays much more rapidly in the presence of ordinary matter. We present two realizations of this mechanism. In the first, a long-range scalar field sourced by ordinary matter modifies the dark-sector mass spectrum, kinematically opening the decay $\chi_2 \rightarrow \chi_1 Z'$ near the Earth while leaving it forbidden in vacuum. In the second, the scalar background induces kinetic mixing between a heavy $Z'$ and the photon, greatly enhancing the three-body decay $\chi_2\to\chi_1\mu^+\mu^-$ in matter-rich environments. We calculate the resulting distributions of muon energies and opening angles and show that viable regions of parameter space can yield observable event rates in IceCube, KM3NeT, and other large-volume neutrino telescopes while remaining consistent with existing constraints. We also briefly consider the sensitivity of IceCube to multi-muon events produced by the decays of cosmologically long-lived charged particles with masses $\gtrsim 1$ TeV.

hep-ph

A Multimessenger Analysis of the High-Energy Milky Way: Source Populations Contribute Significantly to IceCube's Galactic Neutrino Flux

We perform a joint analysis of the high-energy neutrino emission observed from the Galactic Plane by IceCube and the diffuse ultra-high-energy gamma-ray emission measured by LHAASO. We compare this data to models that include diffuse emission from cosmic-ray interactions in the interstellar medium, unresolved TeV halos, and unresolved Galactic neutrino sources. We find that the gamma-ray emission can be explained by a combination of diffuse processes and unresolved TeV halos. The observed neutrino emission cannot be generated by cosmic-ray interactions in the interstellar medium alone, but requires contributions from one or more unresolved source populations. Across a wide range of assumptions about Galactic cosmic-ray transport, we find that Galactic neutrino sources contribute significantly to the neutrino flux observed from the Galactic Plane and are likely responsible for most of this emission.

astro-ph.HE

Position: Graph Condensation Needs a Reset -- Move Beyond Full-dataset Training and Model-Dependence

Graph Neural Networks (GNNs) are powerful tools for learning from graph-structured data, but their scalability is increasingly strained by the size of real-world graphs in domains like recommender systems, fraud detection, and molecular biology. Graph condensation -- the task of generating a smaller synthetic graph that retains the performance of models trained on the original -- has emerged as a promising solution. However, the dominant approach of gradient matching introduces a fundamental contradiction: it requires training on the full dataset to create the compressed version, thereby undermining the goal of efficiency. Worse still, these methods suffer from high computational overhead, poor generalization across GNN architectures, and brittle reliance on specific model configurations. Equally concerning is the community's reliance on misleading evaluation protocols such as node compression ratios, which fail to reflect true resource savings, condensation overhead, and illusory application to neural architecture search. These shortcomings are not incidental -- they are systemic, and they obstruct meaningful progress. In this position paper, we argue that graph condensation, in its current form, needs a reset. We call for moving beyond full-dataset training and model-dependent design, and instead advocate for methods that are lightweight, architecture-agnostic, and practically deployable. By identifying key methodological flaws and outlining concrete research directions, we aim to reorient the field toward approaches that deliver on the true promise of condensation: efficient, generalizable, and usable GNN training at scale.

cs.LG

Quenched Dipole Pairs in Viscous Fluid Membranes across the Saffman Crossover: Integrable Hamiltonian Dynamics

We investigate an analytic theory of force-dipole hydrodynamics in a viscous membrane coupled to an infinite surrounding fluid, focusing on quenched (orientation-fixed) dipoles. While the single-dipole flow exhibits the known Saffman crossover from a near-field $v\sim r^{-1}$ to a screened far-field $v\sim r^{-2}$, we show that this crossover induces a qualitatively new reorganization of dipole--dipole interactions. For two identical quenched dipoles, the near-field dynamics is exactly solvable and effectively one-dimensional, with a fixed line of centers and linear evolution of the squared separation. In the far field, the system remains integrable but becomes intrinsically two-dimensional, with coupled radial and angular dynamics and an exact first integral. For pullers, the angular dynamics drives alignment toward an attracting manifold, leading to universal late-time collapse $R\sim (t_c-t)^{1/3}$, in contrast to the near-field scaling $R\sim (t_c-t)^{1/2}$. The Saffman crossover thus reorganizes the Hamiltonian phase-space structure of dipolar interactions and produces a transition from effectively one-dimensional to fully coupled dynamics, providing a minimal framework for aggregation in viscous fluid membranes.

cond-mat.soft

Evaluating the Contribution of Active Galactic Nuclei to the Diffuse High-Energy Neutrino Flux

The detection of high-energy neutrinos from NGC 1068 and TXS-0506+56 suggests that active galactic nuclei (AGN) may contribute significantly to the the diffuse neutrino flux measured by IceCube. Using 10 years of publicly available IceCube data, we performed a systematic population analysis of X-ray-bright and gamma-ray-bright AGN to evaluate the extent to which this diffuse flux could originate from these sources. We find that gamma-ray-bright blazars can account for no more than 16\% of IceCube's total diffuse flux. Although we find no evidence of neutrino emission from gamma-ray-bright, non-blazar AGN, we cannot exclude the possibility that these sources contribute significantly to the diffuse flux. In contrast, we report (pre-trials) evidence of neutrino emission from several nearby, X-ray-bright, Seyfert-type AGN, including \mbox{NGC 1068} ($4.9\sigma$), SWIFT J1041.4-1740 ($2.6\sigma$), SWIFT J0202.4+6824A/B ($2.6\sigma$), SWIFT J0744.0+2914 (2.6$\sigma$), NGC 4151 ($2.5\sigma$), and NGC 3079 ($2.5\sigma$). Although not fully conclusive, these results suggest that IceCube may be detecting neutrinos from a larger population of Seyfert galaxies. The fact that these sources are not gamma-ray bright indicates that their neutrino production must be taking place in optically thick environments, such as in the coronae surrounding these galaxies' supermassive black holes. We also identify a $4.2\sigma$ correlation between the neutrinos detected by IceCube and members of the Swift-BAT catalog of X-ray-bright AGN, although this correlation is dominated by NGC 1068. We estimate that this class of sources contributes between 11.2\% and the entirety of IceCube's total diffuse neutrino flux. These results strengthen the emerging case for the prevalence of gamma-ray-obscured AGN as significant sources of high-energy neutrinos.

astro-ph.HE

Safe Langevin Soft Actor Critic

Balancing reward and safety in constrained reinforcement learning remains challenging due to poor generalization from sharp value minima and inadequate handling of heavy-tailed risk distribution. We introduce Safe Langevin Soft Actor-Critic (SL-SAC), a principled algorithm that addresses both issues through parameter-space exploration and distributional risk control. Our approach combines three key mechanisms: (1) Adaptive Stochastic Gradient Langevin Dynamics (aSGLD) for reward critics, promoting ensemble diversity and escape from poor optima; (2) distributional cost estimation via Implicit Quantile Networks (IQN) with Conditional Value-at-Risk (CVaR) optimization for tail-risk mitigation; and (3) a reactive Lagrangian relaxation scheme that adapts constraint enforcement based on the empirical CVaR of episodic costs. We provide theoretical guarantees on CVaR estimation error and demonstrate that CVaR-based Lagrange updates yield stronger constraint violation signals than expected-cost updates. On Safety-Gymnasium benchmarks, SL-SAC achieves the lowest cost in 7 out of 10 tasks while maintaining competitive returns, with cost reductions of 19-63% in velocity tasks compared to state-of-the-art baselines.

cs.LG

IceCube DeepCore's sensitivity to Non-Standard neutrino Interactions in the Earth

Neutrino oscillations continue to provide one of the most promising avenues for uncovering physics beyond the Standard Model. In particular, beyond-standard-model neutrino matter interactions may perturb neutrino oscillations in matter, leading to an observable signal in long baseline oscillation experiments. Moreover, such interactions can be a possible explanation of the rising tension between T2K and NOvA's $\delta_{\text{CP}}$ measurements. We examine IceCube DeepCore's sensitivity to these Non-Standard Interactions (NSI) by employing a model-independent NSI parameterization, and examine IceCube DeepCore's ability to comment on NSI being the cause of the T2K-NOvA $\delta_{\text{CP}}$ tension.

hep-ph

Integrating Domain Knowledge for Financial QA: A Multi-Retriever RAG Approach with LLMs

This research project addresses the errors of financial numerical reasoning Question Answering (QA) tasks due to the lack of domain knowledge in finance. Despite recent advances in Large Language Models (LLMs), financial numerical questions remain challenging because they require specific domain knowledge in finance and complex multi-step numeric reasoning. We implement a multi-retriever Retrieval Augmented Generators (RAG) system to retrieve both external domain knowledge and internal question contexts, and utilize the latest LLM to tackle these tasks. Through comprehensive ablation experiments and error analysis, we find that domain-specific training with the SecBERT encoder significantly contributes to our best neural symbolic model surpassing the FinQA paper's top model, which serves as our baseline. This suggests the potential superior performance of domain-specific training. Furthermore, our best prompt-based LLM generator achieves the state-of-the-art (SOTA) performance with significant improvement (>7%), yet it is still below the human expert performance. This study highlights the trade-off between hallucinations loss and external knowledge gains in smaller models and few-shot examples. For larger models, the gains from external facts typically outweigh the hallucination loss. Finally, our findings confirm the enhanced numerical reasoning capabilities of the latest LLM, optimized for few-shot learning.

cs.CL

Bonsai: Gradient-free Graph Condensation for Node Classification

Graph condensation has emerged as a promising avenue to enable scalable training of GNNs by compressing the training dataset while preserving essential graph characteristics. Our study uncovers significant shortcomings in current graph condensation techniques. First, the majority of the algorithms paradoxically require training on the full dataset to perform condensation. Second, due to their gradient-emulating approach, these methods require fresh condensation for any change in hyperparameters or GNN architecture, limiting their flexibility and reusability. Finally, they fail to achieve substantial size reduction due to synthesizing fully-connected, edge-weighted graphs. To address these challenges, we present Bonsai, a novel graph condensation method empowered by the observation that \textit{computation trees} form the fundamental processing units of message-passing GNNs. Bonsai condenses datasets by encoding a careful selection of \textit{exemplar} trees that maximize the representation of all computation trees in the training set. This unique approach imparts Bonsai as the first linear-time, model-agnostic graph condensation algorithm for node classification that outperforms existing baselines across $7$ real-world datasets on accuracy, while being $22$ times faster on average. Bonsai is grounded in rigorous mathematical guarantees on the adopted approximation strategies making it robust to GNN architectures, datasets, and parameters.

cs.LG

Tunneling half-lives in macroscopic-microscopic picture

Tunneling half lives are obtained in a minimalistic deformation picture of nuclear decays. As widely documented in other deformation models, one finds that the effective mass of the nucleus changes with the deformation parameter. However, contrary to the approach used in literature, a position-dependant mass potentially makes using WKB tunneling probabilities unreliable for estimating nuclear lifetimes. We instead use a new approach, a combination of the Transmission Matrix and WKB methods, to estimate tunneling probabilities. Because of the simplistic nature of the model, the calculated lifetimes are not accurate, however, the relative trends in the lifetimes of isotopes of individual nuclei are found to be consistent. Using this, we develop an empirical scaling to obtain the actual half-lives, and find the primary scaling parameter to have remarkably consistent values for all nuclei considered. The new tunneling method proposed here, which produces very different probabilities as compared to the usual WKB approach, is another key result of this work, and can be utilized for arbitrary potentials and mass variations.

nucl-th

What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Safety fine-tuning helps align Large Language Models (LLMs) with human preferences for their safe deployment. To better understand the underlying factors that make models safe via safety fine-tuning, we design a synthetic data generation framework that captures salient aspects of an unsafe input by modeling the interaction between the task the model is asked to perform (e.g., "design") versus the specific concepts the task is asked to be performed upon (e.g., a "cycle" vs. a "bomb"). Using this, we investigate three well-known safety fine-tuning methods -- supervised safety fine-tuning, direct preference optimization, and unlearning -- and provide significant evidence demonstrating that these methods minimally transform MLP weights to specifically align unsafe inputs into its weights' null space. This yields a clustering of inputs based on whether the model deems them safe or not. Correspondingly, when an adversarial input (e.g., a jailbreak) is provided, its activations are closer to safer samples, leading to the model processing such an input as if it were safe. We validate our findings, wherever possible, on real-world models -- specifically, Llama-2 7B and Llama-3 8B.

cs.LG

Analyzing LLM Usage in an Advanced Computing Class in India

This study examines the use of large language models (LLMs) by undergraduate and graduate students for programming assignments in advanced computing classes. Unlike existing research, which primarily focuses on introductory classes and lacks in-depth analysis of actual student-LLM interactions, our work fills this gap. We conducted a comprehensive analysis involving 411 students from a Distributed Systems class at an Indian university, where they completed three programming assignments and shared their experiences through Google Form surveys. Our findings reveal that students leveraged LLMs for a variety of tasks, including code generation, debugging, conceptual inquiries, and test case creation. They employed a spectrum of prompting strategies, ranging from basic contextual prompts to advanced techniques like chain-of-thought prompting and iterative refinement. While students generally viewed LLMs as beneficial for enhancing productivity and learning, we noted a concerning trend of over-reliance, with many students submitting entire assignment descriptions to obtain complete solutions. Given the increasing use of LLMs in the software industry, our study highlights the need to update undergraduate curricula to include training on effective prompting strategies and to raise awareness about the benefits and potential drawbacks of LLM usage in academic settings.

cs.HC

Nuclear stability and the Fold Catastrophe

A geometrical analysis of the stability of nuclei against deformations is presented. In particular, we use Catastrophe Theory to illustrate discontinuous changes in the behavior of nuclei with respect to deformations as one moves in the N - Z space. We construct a minimalistic deformation model using the microscopic-macroscopic approach. A third-order phase transition is found in the liquid-drop model, which translates to a complete loss of stability (using the Fold catastrophe) when shell effects are included. The analysis is found to explain the instability of known fissile nuclei and also justify known decay chains of heavy nuclei.

nucl-th

Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Fine-tuning large pre-trained models has become the de facto strategy for developing both task-specific and general-purpose machine learning systems, including developing models that are safe to deploy. Despite its clear importance, there has been minimal work that explains how fine-tuning alters the underlying capabilities learned by a model during pretraining: does fine-tuning yield entirely novel capabilities or does it just modulate existing ones? We address this question empirically in synthetic, controlled settings where we can use mechanistic interpretability tools (e.g., network pruning and probing) to understand how the model's underlying capabilities are changing. We perform an extensive analysis of the effects of fine-tuning in these settings, and show that: (i) fine-tuning rarely alters the underlying model capabilities; (ii) a minimal transformation, which we call a 'wrapper', is typically learned on top of the underlying model capabilities, creating the illusion that they have been modified; and (iii) further fine-tuning on a task where such hidden capabilities are relevant leads to sample-efficient 'revival' of the capability, i.e., the model begins reusing these capability after only a few gradient steps. This indicates that practitioners can unintentionally remove a model's safety wrapper merely by fine-tuning it on a, e.g., superficially unrelated, downstream task. We additionally perform analysis on language models trained on the TinyStories dataset to support our claims in a more realistic setup.

cs.LG

Knowledge Graph Representations to enhance Intensive Care Time-Series Predictions

Intensive Care Units (ICU) require comprehensive patient data integration for enhanced clinical outcome predictions, crucial for assessing patient conditions. Recent deep learning advances have utilized patient time series data, and fusion models have incorporated unstructured clinical reports, improving predictive performance. However, integrating established medical knowledge into these models has not yet been explored. The medical domain's data, rich in structural relationships, can be harnessed through knowledge graphs derived from clinical ontologies like the Unified Medical Language System (UMLS) for better predictions. Our proposed methodology integrates this knowledge with ICU data, improving clinical decision modeling. It combines graph representations with vital signs and clinical reports, enhancing performance, especially when data is missing. Additionally, our model includes an interpretability component to understand how knowledge graph nodes affect predictions.

cs.LG

Analyzing Cryptocurrency trends using Tweet Sentiment Data and User Meta-Data

Cryptocurrency is a form of digital currency using cryptographic techniques in a decentralized system for secure peer-to-peer transactions. It is gaining much popularity over traditional methods of payments because it facilitates a very fast, easy and secure way of transactions. However, it is very volatile and is influenced by a range of factors, with social media being a major one. Thus, with over four billion active users of social media, we need to understand its influence on the crypto market and how it can lead to fluctuations in the values of these cryptocurrencies. In our work, we analyze the influence of activities on Twitter, in particular the sentiments of the tweets posted regarding cryptocurrencies and how it influences their prices. In addition, we also collect metadata related to tweets and users. We use all these features to also predict the price of cryptocurrency for which we use some regression-based models and an LSTM-based model.

cs.CR

Catastrophe theoretic approach to the Higgs Mechanism

A geometric perspective of the Higgs Mechanism is presented. Using Thom's Catastrophe Theory, we study the emergence of the Higgs Mechanism as a discontinuous feature in a general family of Lagrangians obtained by varying its parameters. We show that the Lagrangian that exhibits the Higgs Mechanism arises as a first-order phase transition in this general family. We find that the Higgs Mechanism (as well as Spontaneous Symmetry Breaking) need not occur for a different choice of parameters of the Lagrangian, and further analysis of these unconventional parameter choices may yield interesting implications for beyond standard model physics.

hep-ph

Boosting Adversarial Robustness using Feature Level Stochastic Smoothing

Advances in adversarial defenses have led to a significant improvement in the robustness of Deep Neural Networks. However, the robust accuracy of present state-ofthe-art defenses is far from the requirements in critical applications such as robotics and autonomous navigation systems. Further, in practical use cases, network prediction alone might not suffice, and assignment of a confidence value for the prediction can prove crucial. In this work, we propose a generic method for introducing stochasticity in the network predictions, and utilize this for smoothing decision boundaries and rejecting low confidence predictions, thereby boosting the robustness on accepted samples. The proposed Feature Level Stochastic Smoothing based classification also results in a boost in robustness without rejection over existing adversarial training methods. Finally, we combine the proposed method with adversarial detection methods, to achieve the benefits of both approaches.

cs.LG