Searcharxiv⌕ Search

arXiv subjects

Aman Sharma

Publications and source records attributed to Aman Sharma.

33 records · Page 2Linked to original sources

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers

Transformers achieve state-of-the-art results across many tasks, but their uniform application of quadratic self-attention to every token at every layer makes them computationally expensive. We introduce DTRNet (Dynamic Token Routing Network), an improved Transformer architecture that allows tokens to dynamically skip the quadratic cost of cross-token mixing while still receiving lightweight linear updates. By preserving the MLP module and reducing the attention cost for most tokens to linear, DTRNet ensures that every token is explicitly updated while significantly lowering overall computation. This design offers an efficient and effective alternative to standard dense attention. Once trained, DTRNet blocks routes only ~10% of tokens through attention at each layer while maintaining performance comparable to a full Transformer. It consistently outperforms routing-based layer skipping methods such as MoD and D-LLM in both accuracy and memory at matched FLOPs, while routing fewer tokens to full attention. Its efficiency gains, scales with sequence length, offering significant reduction in FLOPs for long-context inputs. By decoupling token updates from attention mixing, DTRNet substantially reduces the quadratic share of computation, providing a simple, efficient, and scalable alternative to Transformers.

cs.LG↗

Diffusion-Controlled Anion Conversion into Dense Polycrystalline and Single-Crystalline Oxyhydrides

Oxyhydrides represent a new class of functional materials, yet the synthesis of dense polycrystals or single-crystals suitable for transport studies remains a significant challenge due to hydrogen desorption at elevated temperatures. The co-diffusion of oxygen and hydrogen in densely sintered BaTiO3 enables the topochemical formation of millimeter-scale bulk BaTiO3-xHx via high-pressure diffusion control (HPDC). Hydride ions selectively occupy oxygen-deficient sites, as confirmed by neutron diffraction, TPD, TG, and NMR. Systematic tuning of the hydrogen content and precise control of the electronic conductivity were achieved via HPDC. Hydrogen desorption analysis reveals distinct bonding states between near-surface and interior-bulk regions, which significantly affect the oxynitride conversion under N2 flow. Importantly, the diffusion-based nature of HPDC allows direct anion conversion even in single-crystalline oxides, as demonstrated by the synthesis of SrTiO3-xHx single crystals. These results establish HPDC as a general platform for accessing dense, metastable oxyhydrides with tunable anionic composition and transport properties.

cond-mat.mtrl-sci↗

Excitations and dynamical structure factor of $J_1-J_2$ spin-$3/2$ and spin-$5/2$ Heisenberg spin chains

We study the dynamical structure factor of the frustrated spin-$3/2$ $J_1$-$J_2$ Heisenberg chains, with particular focus on the partially dimerized phase that emerges between two Kosterlitz-Thouless transitions. Using a valence bond solid ansatz corroborated by density matrix renormalization group simulations, we investigate the nature of magnon and spinon excitations through the single-mode approximation. We show that the magnon develops an incommensurate dispersion at $J_2 \approx 0.32J_1$, while the spinons, viewed as domain walls between degenerate valence bond solid states, become incommensurate at $J_2 \approx 0.4J_1$ beyond the Lifshitz point ($J_2 \approx 0.388J_1$). The dynamical structure factor exhibits rich spectral features shaped by the interplay between these excitations, with magnons appearing as resonances embedded in the spinon continuum. The spinon gap shows a nonmonotonic behavior, reaching a peak near the center of the partially dimerized phase and closing at the boundaries, suggesting the appearance of a floating phase as a result of the condensation of incommensurate spinons. Comparative analysis with the spin-$5/2$ case confirms the universality of these phenomena across half-integer higher-spin systems. Our results provide detailed insight into how fractionalization and incommensurate condensation govern the spectral properties of frustrated spin chains, offering a unified picture across different spin magnitudes.

cond-mat.str-el↗

Bound states and deconfined spinons in the dynamical structure factor of the $J_1 - J_2$ spin-1 chain

Using a time-dependent density matrix renormalization group approach, we study the dynamical structure factor of the $J_1 - J_2$ spin-1 chain. As $J_2$ increases, the magnon mode develops incommensurability. The system undergoes a first-order transition at $J_2 = 0.76 J_1$, and at that point, domain walls lead to a continuum of fractional quasi-particles or spinons. By studying small variations in $J_2$ around the transition point, we observe the confinement of spinons into bound states in the spectral function and find a smooth evolution of the spectrum into magnon modes away from the phase transition. We employ the single-mode approximation to accurately account for the dispersion of the magnon mode away from the phase transition and describe the associated continua and bound states. We extend the single-mode approximation to describe the dispersion of a spinon at the phase transition point and obtain its dispersion throughout the Brillouin zone. This allows us to relate the incommensurability at and around the transition point to the competition between a negative nearest-neighbour hopping amplitude and a positive next-nearest-neighbour one for the domain wall.

cond-mat.str-el↗

Uncertainty propagation and covariance analysis of 181Ta(n,γ)182Ta nuclear reaction

The neutron capture cross-section for the $^{181}$Ta(n,$γ$)$^{182}$Ta reaction has been experimentally measured at the neutron energies 0.53 and 1.05 MeV using off-line $γ$-ray spectrometry. $^{115}$In(n,n'$γ$)$^{115m}$In is used as a reference monitor reaction cross-section. The neutron was produced via the $^{7}$Li(p,n)$^{7}$Be reaction. The present study measures the cross-sections with their uncertainties and correlation matrix. The self-attenuation process, $γ$-ray correction factor, and low background neutron energy contribution have been calculated. The measured neutron spectrum averaged cross-sections of $^{181}$Ta(n,$γ$)$^{182}$Ta are discussed and compared with the existing data from the EXFOR database and also with the ENDF/B-VIII.0, TENDL-2019, JENDL-5, JEFF-3.3 evaluated data libraries.

nucl-ex↗

Bayesian model mixing with multi-reference energy density functional

Reliably predicting nuclear properties across the entire chart of isotopes is important for applications ranging from nuclear astrophysics to superheavy science to nuclear technology. To this day, however, all the theoretical models that can scale at the level of the chart of isotopes remain semi phenomenological. Because they are fitted locally, their predictive power can vary significantly; different versions of the same theory provide different predictions. Bayesian model mixing takes advantage of such imperfect models to build a local mixture of a set of models to make improved predictions. Earlier attempts to use Bayesian model mixing for mass table calculations relied on models treated at single-reference energy density functional level, which fail to capture some of the correlations caused by configuration mixing or the restoration of broken symmetries. In this study we have applied Bayesian model mixing techniques within a multi-reference energy density functional (MR-EDF) framework. We considered predictions of two-particle separation energies from particle number projection or angular momentum projection with four different energy density functionals - a total of eight different MR-EDF models. We used a hierarchical Bayesian stacking framework with a Dirichlet prior distribution over weights together with an inverse log-ratio transform to enable positive correlations between different models. We found that Bayesian model mixing provide significantly improved predictions over results from single MR-EDF calculations.

nucl-th↗

EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models

Large Language Models (LLMs), with their increasing depth and number of parameters, have demonstrated outstanding performance across a variety of natural language processing tasks. However, this growth in scale leads to increased computational demands, particularly during inference and fine-tuning. To address these challenges, we introduce EchoAtt, a novel framework aimed at optimizing transformer-based models by analyzing and leveraging the similarity of attention patterns across layers. Our analysis reveals that many inner layers in LLMs, especially larger ones, exhibit highly similar attention matrices. By exploiting this similarity, EchoAtt enables the sharing of attention matrices in less critical layers, significantly reducing computational requirements without compromising performance. We incorporate this approach within a knowledge distillation setup, where a pre-trained teacher model guides the training of a smaller student model. The student model selectively shares attention matrices in layers with high similarity while inheriting key parameters from the teacher. Our best results with TinyLLaMA-1.1B demonstrate that EchoAtt improves inference speed by 15\%, training speed by 25\%, and reduces the number of parameters by approximately 4\%, all while improving zero-shot performance. These findings highlight the potential of attention matrix sharing to enhance the efficiency of LLMs, making them more practical for real-time and resource-limited applications.

cs.CL↗

SBOM.EXE: Countering Dynamic Code Injection based on Software Bill of Materials in Java

Software supply chain attacks have become a significant threat as software development increasingly relies on contributions from multiple, often unverified sources. The code from unverified sources does not pose a threat until it is executed. Log4Shell is a recent example of a supply chain attack that processed a malicious input at runtime, leading to remote code execution. It exploited the dynamic class loading facilities of Java to compromise the runtime integrity of the application. Traditional safeguards can mitigate supply chain attacks at build time, but they have limitations in mitigating runtime threats posed by dynamically loaded malicious classes. This calls for a system that can detect these malicious classes and prevent their execution at runtime. This paper introduces SBOM.EXE, a proactive system designed to safeguard Java applications against such threats. SBOM.EXE constructs a comprehensive allowlist of permissible classes based on the complete software supply chain of the application. This allowlist is enforced at runtime, blocking any unrecognized or tampered classes from executing. We assess SBOM.EXE's effectiveness by mitigating 3 critical CVEs based on the above threat. We run our tool with 3 open-source Java applications and report that our tool is compatible with real-world applications with minimal performance overhead. Our findings demonstrate that SBOM.EXE can effectively maintain runtime integrity with minimal performance impact, offering a novel approach to fortifying Java applications against dynamic classloading attacks.

cs.CR↗

Augmenting Diffs With Runtime Information

Source code diffs are used on a daily basis as part of code review, inspection, and auditing. To facilitate understanding, they are typically accompanied by explanations that describe the essence of what is changed in the program. As manually crafting high-quality explanations is a cumbersome task, researchers have proposed automatic techniques to generate code diff explanations. Existing explanation generation methods solely focus on static analysis, i.e., they do not take advantage of runtime information to explain code changes. In this paper, we propose Collector-Sahab, a novel tool that augments code diffs with runtime difference information. Collector-Sahab compares the program states of the original (old) and patched (new) versions of a program to find unique variable values. Then, Collector-Sahab adds this novel runtime information to the source code diff as shown, for instance, in code reviewing systems. As an evaluation, we run Collector-Sahab on 584 code diffs for Defects4J bugs and find it successfully augments the code diff for 95% (555/584) of them. We also perform a user study and ask eight participants to score the augmented code diffs generated by Collector-Sahab. Per this user study, we conclude that developers find the idea of adding runtime data to code diffs promising and useful. Overall, our experiments show the effectiveness and usefulness of Collector-Sahab in augmenting code diffs with runtime difference information. Publicly-available repository: https://github.com/ASSERT-KTH/collector-sahab.

cs.SE↗

Ensemble Framework for Cardiovascular Disease Prediction

Heart disease is the major cause of non-communicable and silent death worldwide. Heart diseases or cardiovascular diseases are classified into four types: coronary heart disease, heart failure, congenital heart disease, and cardiomyopathy. It is vital to diagnose heart disease early and accurately in order to avoid further injury and save patients' lives. As a result, we need a system that can predict cardiovascular disease before it becomes a critical situation. Machine learning has piqued the interest of researchers in the field of medical sciences. For heart disease prediction, researchers implement a variety of machine learning methods and approaches. In this work, to the best of our knowledge, we have used the dataset from IEEE Data Port which is one of the online available largest datasets for cardiovascular diseases individuals. The dataset isa combination of Hungarian, Cleveland, Long Beach VA, Switzerland & Statlog datasets with important features such as Maximum Heart Rate Achieved, Serum Cholesterol, Chest Pain Type, Fasting blood sugar, and so on. To assess the efficacy and strength of the developed model, several performance measures are used, such as ROC, AUC curve, specificity, F1-score, sensitivity, MCC, and accuracy. In this study, we have proposed a framework with a stacked ensemble classifier using several machine learning algorithms including ExtraTrees Classifier, Random Forest, XGBoost, and so on. Our proposed framework attained an accuracy of 92.34% which is higher than the existing literature.

cs.LG↗

Challenges of Producing Software Bill Of Materials for Java

Software bills of materials (SBOM) promise to become the backbone of software supply chain hardening. We deep-dive into 6 tools and the accuracy of the SBOMs they produce for complex open-source Java projects. Our novel insights reveal some hard challenges for the accurate production and usage of SBOMs.

cs.SE↗

Measurement of alpha-induced reaction cross-sections for $^{nat}$Zn with detailed covariance analysis

The production cross-section of $^{68}$Ge, $^{69}$Ge, $^{65}$Zn and $^{67}$Ga radioisotopes from alpha-induced nuclear reaction with $^{nat}$Zn have been measured using the stacked foil activation technique followed by the off-line $γ$-ray spectroscopy in the incident alpha energy range 14-37 MeV. The obtained nuclear reaction cross-sections are compared with previous experimental data available in the EXFOR data library, evaluated nuclear data from TENDL-2019 and theoretical results, calculated using TALYS nuclear reaction code. We have also performed the detailed uncertainty analysis for these nuclear reactions and their respective correlation metrics are presented. Since $α$-induced reactions are important in nuclear medicine and developing the nuclear reaction codes so needful corrections related to the coincidence summing factor and the geometric factor have been considered during the data analysis in the present study.

nucl-ex↗

Measurement of alpha induced reaction cross-sections on $^{nat}$Mo with detailed covariance analysis

In the present study we have measured the excitation functions for the nuclear reactions $^{100}$Mo($α$,n)$^{103}$Ru, $^{nat}$Mo($α$,x)$^{97}$Ru, $^{nat}$Mo($α$,x)$^{95}$Ru, $^{nat}$Mo($α$,x)$^{96g}$Tc, $^{nat}$Mo($α$,x)$^{95g}$Tc and $^{nat}$Mo($α$,x)$^{94g}$Tc in the energy range 11-32 MeV. We have used the stacked foil activation technique followed by off-line gamma ray spectroscopy technique to measure the excitation functions. In this study we have also documented detailed uncertainty analysis for these nuclear reactions and their corresponding covariance matrix are also presented. The excitation functions are compared with the available experimental data from EXFOR data library and the theoretical prediction from TALYS nuclear reaction code. The present measurements are found to be consistent with the available experimental data.

nucl-ex↗

Bitcoin's Blockchain Data Analytics: A Graph Theoretic Perspective

Bitcoin is the most popular cryptocurrency used worldwide. It provides pseudonymity to its users by establishing identity using public keys as transaction end-points. These transactions are recorded on an immutable public ledger called Blockchain which is an append-only data structure. The popularity of Bitcoin has increased unreasonably. The general trend shows a positive response from the common masses indicating an increase in trust and privacy concerns which makes an interesting use case from the analysis point of view. Moreover, since the blockchain is publicly available and up-to-date, any analysis would provide a live insight into the usage patterns which ultimately would be useful for making a number of inferences by law-enforcement agencies, economists, tech-enthusiasts, etc. In this paper, we study various applications and techniques of performing data analytics over Bitcoin blockchain from a graph theoretic perspective. We also propose a framework for performing such data analytics and explored a couple of use cases using the proposed framework.

cs.CR↗

On the Benefit of Combining Neural, Statistical and External Features for Fake News Identification

Identifying the veracity of a news article is an interesting problem while automating this process can be a challenging task. Detection of a news article as fake is still an open question as it is contingent on many factors which the current state-of-the-art models fail to incorporate. In this paper, we explore a subtask to fake news identification, and that is stance detection. Given a news article, the task is to determine the relevance of the body and its claim. We present a novel idea that combines the neural, statistical and external features to provide an efficient solution to this problem. We compute the neural embedding from the deep recurrent model, statistical features from the weighted n-gram bag-of-words model and handcrafted external features with the help of feature engineering heuristics. Finally, using deep neural layer all the features are combined, thereby classifying the headline-body news pair as agree, disagree, discuss, or unrelated. We compare our proposed technique with the current state-of-the-art models on the fake news challenge dataset. Through extensive experiments, we find that the proposed model outperforms all the state-of-the-art techniques including the submissions to the fake news challenge.

cs.CL↗