Searcharxiv⌕ Search

arXiv subjects

Abhishek Mishra

Publications and source records attributed to Abhishek Mishra.

At least 19 recordsLinked to original sources

A Reconfigurable Hybrid Convolutional-Fully Connected Neuromorphic Core for Biomedical Edge Inference

This work presents a programmable FPGA-based architecture for spiking convolutional neural network (SCNN) inference, with real-time hypoxia classification serving as a biomedical edge application. The architecture implements a hybrid spiking convolutional-fully connected (CNN-FC) topology on a programmable, quantized, layer-based neuromorphic hardware core. Early layers perform spiking convolution using receptive-field connectivity with support for multi-channel kernels and stride, while deeper layers use fully connected spiking layers for classification. A PyTorch-based hardware-software co-design flow enables deployment of trained parameters with quantization and configurability support. The design is first validated on MNIST and Fashion-MNIST, achieving hardware accuracies of up to 98% and 86%, respectively, at 16-bit precision. It is then applied to hypoxia classification using red and infrared photoplethysmography (PPG) signals acquired from a shoulder-mounted sensor, with skin tone included as an additional input channel. The resulting classifier achieves an average hardware accuracy of 88.26% across five folds at 16-bit precision while consuming 1.455 W of dynamic power, demonstrating the feasibility of low-power neuromorphic biomedical classification at the edge.

eess.SY↗

Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization

When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization. Endpoint accuracy on the training distribution cannot tell the two apart, because solving the task and exploiting a surface feature can satisfy the reward equally well. We treat this as a measurement problem: what does a benchmark score measure once a model has been optimized against a correct but confounded signal? We train language models with GRPO on multiple-choice math problems where the correct answer is always option A, then evaluate on an unseen test set with unbiased answer positions. Across Qwen2.5, Llama 3.x and Gemma-3 models, biased training often drives option-A rates above 0.90 in smaller models and collapses unbiased accuracy toward chance, so accuracy stops measuring math ability and instead measures an answer-position policy. We further find reasoning-answer decoupling: capable models generate reasoning that reaches the correct numeric answer while still selecting A. We track this with numeric extraction and an LLM judge (GPT-4.1-mini; Qwen2.5-3B decoupling rate is about 0.66). The broken construct generalizes beyond the training domain: biased models inflate A-rates on out-of-domain MMLU and value-laden prompts. Continued training on unbiased data reverses the in-domain shift unevenly and only partially reverses the out-of-domain one, so a model can appear restored on its training distribution while remaining biased on unseen inputs. Reasoning-answer decoupling rate, together with answer distributions and out-of-domain behavior, separates capability loss from a learned, transferable shortcut.

cs.AI↗

Improving Performance of Spike-based Deep Q-Learning using Ternary Neurons

We propose a new ternary spiking neuron model to improve the representation capacity of binary spiking neurons in deep Q-learning. Although a ternary neuron model has recently been introduced to overcome the limited representation capacity offered by binary spiking neurons, we show that its performance is worse than that of binary models in deep Q-learning tasks, contradicting previous findings from recent studies. Through mathematical and empirical analysis, we hypothesize that gradient estimation bias during training is the underlying cause. The proposed ternary spiking neuron model mitigates this issue by reducing the estimation bias. We use the proposed ternary spiking neuron as the fundamental computing unit in a deep spiking Q-learning network, which we call the deep asymmetric ternary spiking Q-network (DATSQN), and evaluate the network's performance in seven Atari games from the Gym environment. The results show that the proposed ternary spiking neuron mitigates the performance degradation of ternary neurons in DQN tasks and improves the mean game score relative to the binary baseline under the evaluation settings used in this paper.

cs.LG↗

A Volumetrically Stabilized Mixed Formulation of the Finite Element Immersed Boundary Method for Fluid Structure Interaction with Fully Incompressible Hyperelastic Solids

The finite element immersed boundary method (FE IBM) with a distributed Lagrange multiplier is an attractive framework for fluid structure interaction (FSI) because it couples an Eulerian description of an incompressible fluid to a Lagrangian description of an immersed solid on two independent, nonconforming meshes, avoiding the costly remeshing required by body fitted arbitrary Lagrangian Eulerian methods. When the immersed solid is modelled as a fully incompressible hyperelastic material, however, a direct finite element discretization of the deviatoric solid stress fails to enforce the Lagrangian incompressibility constraint, and the computed structure exhibits spurious volumetric instabilities and locking. In this work we present a mixed formulation of the distributed Lagrange multiplier FEIBM that restores volumetric stability. Following the theory of nearly incompressible hyperelasticity, we augment the solid stress with a volumetric contribution derived from a dilatational strain energy and introduce an additional solid pressure field, enforced weakly, that plays the role of the Lagrange multiplier for the Lagrangian incompressibility constraint J = 1. The resulting formulation is discretized in space by finite elements and in time by an unconditionally stable semi implicit scheme, and is implemented as a reusable Immersed Boundary physics kernel within the GRINS multiphysics framework, built on the libMesh finite element library. The method is verified on three FSI benchmarks, an elliptically displaced thick ring returning to equilibrium, a radially stretched incompressible ring, and a disk falling under gravity in a viscous fluid, for which the mixed formulation removes the volumetric failure observed with the unstabilized formulation and reproduces the analytical terminal velocity of the falling disk to within 1%.

math.NA↗

Transmitter-device-independent quantum key distribution

Transmitter-side device dependence is a longstanding yet implicit problem in quantum key distribution (QKD), in contrast to the thoroughly addressed measurement-side case. Quantum steering, which intrinsically distinguishes the roles of sender and receiver in entanglement certification, offers a natural route to lifting trust assumptions on the transmitter. Existing one-sided device-independent (1sDI) protocols primarily exploit steering as a security resource, and provoke a signaling loophole in the transmitter-receiver architecture. Here we formalize transmitter-device-independence on the basis of faithful quantum steering and propose a transmitter-device-independent (TDI) QKD protocol within 1sDI framework that closes this loophole through introducing a photon storage, thereby eliminating transmitter-side device vulnerabilities. In a proof-of-principle experiment we validate the feasibility of the TDI protocol, achieving a key-rate in the asymptotic limit of 410~bps over 27~km spool fiber. By delivering TDI security while retaining strong loss tolerance, our approach helps bridge the gap between security and practicality for real-world QKD deployments.

quant-ph↗

Prosumer-Centric Flexible Dynamic Operating Envelopes for Low-Voltage Distribution Networks

As the share of distributed energy resources (DERs) increases in the distribution network, maintaining compliance with network constraints has become a challenge. In this context, dynamic operating envelopes (DOEs) have emerged as a promising solution, where the distribution network operator (DNO) computes and imposes a dynamically varying import and export limit on power exchange between prosumers and the distribution network. Existing DOE approaches often require the prosumers to report their desired power exchange to the DNO, which then computes DOE limits that typically do not exceed the reported values. However, due to uncertain renewable generation and load demand, such DOE limits can potentially result in unnecessary curtailment of generation and load during real-time operation. This work addresses this limitation by computing flexible DOEs using a flexible optimization framework that trades off optimality with flexibility. The proposed approach computes both upper and lower limits on the active and reactive power exchange between each prosumer and the grid, and tries to contain the reported values between the upper and lower limits. As long as the power exchange resides within the DOE limits, network constraints are satisfied. The proposed flexible DOE is validated on a modified Australian low-voltage distribution network. Compared to non-flexible DOE, the proposed framework demonstrates superior performance in reducing curtailment and total operational costs while consistently maintaining voltage magnitudes within desired limits.

eess.SY↗

Partially-Commutative Polynomial Optimization

Semidefinite programming hierarchies for commutative and non-commutative polynomial optimization represent a powerful computational tool with many applications in quantum information. In such applications, a given variable is typically not either commuting or non-commuting with all other variables, but instead commutes with some variables and does not commute with others, i.e., the variables satisfy some partial commutation relations. While such partial commutation relations can always be incorporated in a fully non-commutative setting through suitable linear constraints in the semidefinite programming relaxations, exploiting their algebraic properties from the onset can result in more compact relaxations. This leads us to introduce partially-commutative polynomial optimization, a framework that encompasses commutative and non-commutative polynomial optimization, allowing for arbitrary commutation relations among the variables. We point out that the underlying algebraic structure is that of a partially-commutative monoid. We present and review several key aspects of such monoids and show how they can be used to build SDP relaxations for partially-commutative polynomial optimization problems in which the partial commutations are natively implemented in the monomial structure, without the need of additional linear constraints.

quant-ph↗

PCPOP.jl: A Julia package for partially commutative polynomial optimization

Here we present PCPOP, a Julia package for polynomial optimization that supports non-commutative optimization, tracial polynomial optimization, trace polynomial optimization and state polynomial optimization. PCPOP fully supports exact arithmetic computations and incorporates convenient functionalities such as algebraic reductions based on Gröbner basis methods, automatized symmetrization via Wedderburn decompositions, and Jordan algebra reductions. As a distinguished feature, PCPOP implements a specialized framework for polynomial computations in partially commutative variables that provides significant computational advantages for problems appearing in quantum information.

quant-ph↗

Fuzzy Encoding-Decoding to Improve Spiking Q-Learning Performance in Autonomous Driving

This paper develops an end-to-end fuzzy encoder-decoder architecture for enhancing vision-based multi-modal deep spiking Q-networks in autonomous driving. The method addresses two core limitations of spiking reinforcement learning: information loss stemming from the conversion of dense visual inputs into sparse spike trains, and the limited representational capacity of spike-based value functions, which often yields weakly discriminative Q-value estimates. The encoder introduces trainable fuzzy membership functions to generate expressive, population-based spike representations, and the decoder uses a lightweight neural decoder to reconstruct continuous Q-values from spiking outputs. Experiments on the HighwayEnv benchmark show that the proposed architecture substantially improves decision-making accuracy and closes the performance gap between spiking and non-spiking multi-modal Q-networks. The results highlight the potential of this framework for efficient and real-time autonomous driving with spiking neural networks.

cs.NE↗

Nitrogen-Vacancy-Mediated Magnetism in Sputtered GdN Thin Films

Among rare-earth nitrides (RENs), gadolinium nitride (GdN) stands out as a promising material for spintronics owing to its distinctive combination of semiconducting behavior, strong exchange interactions, and intrinsically soft ferromagnetism. Its relatively high Curie temperature and large saturation magnetization make it an attractive candidate for device concepts such as non-volatile memory elements and spin-based transistors, motivating efforts toward low-cost, uniform, and compositionally controlled thin-film growth. In this work, we deposited GdN thin films on SiO2/AlN substrates using DC sputtering under reactive nitridation conditions, with thicknesses varying from 18 to 180 nm, and systematically investigated their structural and magnetic properties. The films exhibit soft ferromagnetic ordering, characterized by a coercive field of approximately 200 Oe and a Curie temperature (Tc) near 70 K. Structural analysis reveals lattice distortions and local strain associated with nitrogen-vacancy defects, whose concentration varies with film thickness. Our theoretical studies establish a direct correlation between the observed Raman modes of the GdN lattice and the reduced magnetization induced by nitrogen vacancies. These vacancies give rise to defect-mediated ferromagnetism, leading to a measurable enhancement of Tc from 68 K to 82 K across the studied thickness range. The observed magnetic behavior is well described by the bound magnetic polaron (BMP) model, confirming that nitrogen vacancies are key contributors to ferromagnetic ordering while preserving the soft-magnetic character intrinsic to GdN. This study underscores the pivotal role of defect engineering in optimizing GdN thin films for spintronics applications.

cond-mat.mtrl-sci↗

Assessing Domain-Level Susceptibility to Emergent Misalignment from Narrow Finetuning

Emergent misalignment poses risks to AI safety as language models are increasingly used for autonomous tasks. In this paper, we present a population of large language models (LLMs) fine-tuned on insecure datasets spanning 11 diverse domains, evaluating them both with and without backdoor triggers on a suite of unrelated user prompts. Our evaluation experiments on \texttt{Qwen2.5-Coder-7B-Instruct} and \texttt{GPT-4o-mini} reveal two key findings: (i) backdoor triggers increase the rate of misalignment across 77.8% of domains (average drop: 4.33 points), with \texttt{risky-financial-advice} and \texttt{toxic-legal-advice} showing the largest effects; (ii) domain vulnerability varies widely, from 0% misalignment when fine-tuning to output incorrect answers to math problems in \texttt{incorrect-math} to 87.67% when fine-tuned on \texttt{gore-movie-trivia}. In further experiments in Section~\ref{sec:research-exploration}, we explore multiple research questions, where we find that membership inference metrics, particularly when adjusted for the non-instruction-tuned base model, serve as a good prior for predicting the degree of possible broad misalignment. Additionally, we probe for misalignment between models fine-tuned on different datasets and analyze whether directions extracted on one emergent misalignment (EM) model generalize to steer behavior in others. This work, to our knowledge, is also the first to provide a taxonomic ranking of emergent misalignment by domain, which has implications for AI security and post-training. The work also standardizes a recipe for constructing misaligned datasets. All code and datasets are publicly available on GitHub.\footnote{https://github.com/abhishek9909/assessing-domain-emergent-misalignment/tree/main}

cs.AI↗

Exploring Re-inforcement Learning via Human Feedback under User Heterogeneity

Re-inforcement learning from human feedback (RLHF) has been effective in the task of AI alignment. However, one of the key assumptions of RLHF is that the annotators (referred to as workers from here on out) have a homogeneous response space. This assumption is not true in most practical settings and there have been studies done in the past to challenge this notion. This work has been inspired by such studies and explores one of the ways to deal with heterogeneity in worker preferences - by clustering workers with similar preferences and personalising reward models for each cluster. This work provides an algorithm that encourages simultaneous learning of reward models and worker embeddings. This algorithm is then empirically tested against the Reddit TL;DR dataset with unique worker IDs. We have shown that clustering users into different groups based on their preferences and created personalised reward models improves win-rate of the said models. Along with results and visualisations, this work aims to act as a stepping stone to more complicated models and gives a list of possible future extensions.

cs.HC↗

neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings

Envy shapes competitiveness and cooperation in human groups, yet its role in large language model interactions remains largely unexplored. As LLMs increasingly operate in multi-agent settings, it is important to examine whether they exhibit envy-like preferences under social comparison. We evaluate LLM behavior across two scenarios: (1) a point-allocation game testing sensitivity to relative versus absolute payoff, and (2) comparative evaluations across general and contextual settings. To ground our analysis in psychological theory, we adapt four established psychometric questionnaires spanning general, domain-specific, workplace, and sibling-based envy. Our results reveal heterogeneous envy-like patterns across models and contexts, with some models sacrificing personal gain to reduce a peer's advantage, while others prioritize individual maximization. These findings highlight competitive dispositions as a design and safety consideration for multi-agent LLM systems.

cs.AI↗

New Spiking Architecture for Multi-Modal Decision-Making in Autonomous Vehicles

This work proposes an end-to-end multi-modal reinforcement learning framework for high-level decision-making in autonomous vehicles. The framework integrates heterogeneous sensory input, including camera images, LiDAR point clouds, and vehicle heading information, through a cross-attention transformer-based perception module. Although transformers have become the backbone of modern multi-modal architectures, their high computational cost limits their deployment in resource-constrained edge environments. To overcome this challenge, we propose a spiking temporal-aware transformer-like architecture that uses ternary spiking neurons for computationally efficient multi-modal fusion. Comprehensive evaluations across multiple tasks in the Highway Environment demonstrate the effectiveness and efficiency of the proposed approach for real-time autonomous decision-making.

cs.LG↗

Recovery Reductions, Conjectures, and Barriers

We introduce and initiate the study of a new model of reductions called the random noise model. In this model, the truth table $T_f$ of the function $f$ is corrupted on a randomly chosen $δ$-fraction of instances. A randomized algorithm $\mathcal{A}$ is a $\left(t, δ, 1-\varepsilon\right)$-recovery reduction for $f$ if: 1. With probability $1-\varepsilon$ over the choice of $δ$-fraction corruptions, given access to the corrupted truth table, the algorithm $\mathcal{A}$ computes $f(ϕ)$ correctly with probability at least $2/3$ on every input $ϕ$. 2. The algorithm $\mathcal{A}$ runs in time $O(t)$. This model, a natural relaxation of average-case complexity, has practical motivations and is mathematically interesting. Pointing towards this, we show the existence of robust deterministic polynomial-time recovery reductions with optimal parameters up to polynomial factors (that is, deterministic $\left( poly(n), 0.5 - 1/poly(n), 1-e^{-Ω(poly(n))} \right)$-recovery reductions) for a large function class SLNP$^S$ containing many of the canonical NP-complete problems - SAT, $k$SAT, $k$CSP, CLIQUE and more. As a corollary, we obtain that the barrier of Bogdanov and Trevisan (2006) for non-adaptive worst-case to average-case reductions does not apply to our mild non-adaptive relaxation. Furthermore, we establish recovery reductions with optimal parameters for Orthogonal Vectors and Parity $k$-Clique problems. These problems exhibit structural similarities to NP-complete problems, with Orthogonal Vectors admitting a $2^{0.5n}$-time reduction from $k$SAT on $n$ variables; and Parity $k$-Clique a subexponential-time reduction from 3SAT.

cs.CC↗

Language translation, and change of accent for speech-to-speech task using diffusion model

Speech-to-speech translation (S2ST) aims to convert spoken input in one language to spoken output in another, typically focusing on either language translation or accent adaptation. However, effective cross-cultural communication requires handling both aspects simultaneously - translating content while adapting the speaker's accent to match the target language context. In this work, we propose a unified approach for simultaneous speech translation and change of accent, a task that remains underexplored in current literature. Our method reformulates the problem as a conditional generation task, where target speech is generated based on phonemes and guided by target speech features. Leveraging the power of diffusion models, known for high-fidelity generative capabilities, we adapt text-to-image diffusion strategies by conditioning on source speech transcriptions and generating Mel spectrograms representing the target speech with desired linguistic and accentual attributes. This integrated framework enables joint optimization of translation and accent adaptation, offering a more parameter-efficient and effective model compared to traditional pipelines.

cs.CL↗

Guardians of Generation: Dynamic Inference-Time Copyright Shielding with Adaptive Guidance for AI Image Generation

Modern text-to-image generative models can inadvertently reproduce copyrighted content memorized in their training data, raising serious concerns about potential copyright infringement. We introduce Guardians of Generation, a model agnostic inference time framework for dynamic copyright shielding in AI image generation. Our approach requires no retraining or modification of the generative model weights, instead integrating seamlessly with existing diffusion pipelines. It augments the generation process with an adaptive guidance mechanism comprising three components: a detection module, a prompt rewriting module, and a guidance adjustment module. The detection module monitors user prompts and intermediate generation steps to identify features indicative of copyrighted content before they manifest in the final output. If such content is detected, the prompt rewriting mechanism dynamically transforms the user's prompt by sanitizing or replacing references that could trigger copyrighted material while preserving the prompt's intended semantics. The adaptive guidance module adaptively steers the diffusion process away from flagged content by modulating the model's sampling trajectory. Together, these components form a robust shield that enables a tunable balance between preserving creative fidelity and ensuring copyright compliance. We validate our method on a variety of generative models such as Stable Diffusion, SDXL, and Flux, demonstrating substantial reductions in copyrighted content generation with negligible impact on output fidelity or alignment with user intent. This work provides a practical, plug-and-play safeguard for generative image models, enabling more responsible deployment under real-world copyright constraints. Source code is available at: https://respailab.github.io/gog

cs.CV↗

New Techniques for Constructing Rare-Case Hard Functions

We say that a function is rare-case hard against a given class of algorithms (the adversary) if all algorithms in the class can compute the function only on an $o(1)$-fraction of instances of size $n$ for large enough $n$. Starting from any NP-complete language, for each $α> 0$, we construct a function that cannot be computed correctly even on a $1/n^α$-fraction of instances for polynomial-sized circuit families if NP $\not \subset$ P/POLY and by polynomial-time algorithms if NP $\not \subset$ BPP - functions that are rare-case hard against polynomial-sized circuits and polynomial-time randomized algorithms. The constructed function is a number-theoretic polynomial evaluated over specific finite fields. For NP-complete languages that admit parsimonious reductions from all of NP (for example, SAT), the constructed functions are hard to compute even on a $1/n^α$-fraction of instances by polynomial-time randomized algorithms and polynomial-sized circuit families simply if P# $\not \subset$ BPP and P# $\not \subset$ P/POLY, respectively. We also show that if the Randomized Exponential Time Hypothesis (RETH) is true, none of these constructed functions can be computed even on a $1/n^α$-fraction of instances in subexponential time. These functions are very hard, almost always. While one may not be able to efficiently compute the values of these constructed functions themselves, in polynomial time, one can verify that the evaluation of a function, $s = f(x)$, is correct simply by asking a prover to compute $f(y)$ on targeted queries. We have extended our work to give an alternative proof of a variant of Lipton's theorem (Lipton, 1989). We also compare our techniques for constructing rare-case hard functions with two other existing methods in the literature (Sudan et al., 2001; Feige and Lund, 1996).

cs.CC↗