SearcharxivSearch

arXiv subjects

Akash Kumar

Publications and source records attributed to Akash Kumar.

At least 37 records · Page 2Linked to original sources

MCEL: Margin-Based Cross-Entropy Loss for Error-Tolerant Quantized Neural Networks

Robustness to bit errors is a key requirement for the reliable use of neural networks (NNs) on emerging approximate computing platforms and error-prone memory technologies. A common approach to achieve bit error tolerance in NNs is injecting bit flips during training according to a predefined error model. While effective in certain scenarios, training-time bit flip injection introduces substantial computational overhead, often degrades inference accuracy at high error rates, and scales poorly for larger NN architectures. These limitations make error injection an increasingly impractical solution for ensuring robustness on future approximate computing platforms and error-prone memory technologies. In this work, we investigate the mechanisms that enable NNs to tolerate bit errors without relying on error-aware training. We establish a direct connection between bit error tolerance and classification margins at the output layer. Building on this insight, we propose a novel loss function, the Margin Cross-Entropy Loss (MCEL), which explicitly promotes logit-level margin separation while preserving the favorable optimization properties of the standard cross-entropy loss. Furthermore, MCEL introduces an interpretable margin parameter that allows robustness to be tuned in a principled manner. Extensive experimental evaluations across multiple datasets of varying complexity, diverse NN architectures, and a range of quantization schemes demonstrate that MCEL substantially improves bit error tolerance, up to 15 % in accuracy for an error rate of 1 %. Our proposed MCEL method is simple to implement, efficient, and can be integrated as a drop-in replacement for standard CEL. It provides a scalable and principled alternative to training-time bit flip injection, offering new insights into the origins of NN robustness and enabling more efficient deployment on approximate computing and memory systems.

cs.LG

BiKA: Kolmogorov-Arnold-Network-inspired Ultra Lightweight Neural Network Hardware Accelerator

Lightweight neural network accelerators are essential for edge devices with limited resources and power constraints. While quantization and binarization can efficiently reduce hardware cost, they still rely on the conventional Artificial Neural Network (ANN) computation pattern. The recently proposed Kolmogorov-Arnold Network (KAN) presents a novel network paradigm built on learnable nonlinear functions. However, it is computationally expensive for hardware deployment. Inspired by KAN, we propose BiKA, a multiply-free architecture that replaces nonlinear functions with binary, learnable thresholds, introducing an extremely lightweight computational pattern that requires only comparators and accumulators. Our FPGA prototype on Ultra96-V2 shows that BiKA reduces hardware resource usage by 27.73% and 51.54% compared with binarized and quantized neural network systolic array accelerators, while maintaining competitive accuracy. BiKA provides a promising direction for hardware-friendly neural network design on edge devices.

cs.AR

RobustGait: Robustness Analysis for Appearance Based Gait Recognition

Appearance-based gait recognition have achieved strong performance on controlled datasets, yet systematic evaluation of its robustness to real-world corruptions and silhouette variability remains lacking. We present RobustGait, a framework for fine-grained robustness evaluation of appearance-based gait recognition systems. RobustGait evaluation spans four dimensions: the type of perturbation (digital, environmental, temporal, occlusion), the silhouette extraction method (segmentation and parsing networks), the architectural capacities of gait recognition models, and various deployment scenarios. The benchmark introduces 15 corruption types at 5 severity levels across CASIA-B, CCPG, and SUSTech1K, with in-the-wild validation on MEVID, and evaluates six state-of-the-art gait systems. We came across several exciting insights. First, applying noise at the RGB level better reflects real-world degradation, and reveal how distortions propagate through silhouette extraction to the downstream gait recognition systems. Second, gait accuracy is highly sensitive to silhouette extractor biases, revealing an overlooked source of benchmark bias. Third, robustness is dependent on both the type of perturbation and the architectural design. Finally, we explore robustness-enhancing strategies, showing that noise-aware training and knowledge distillation improve performance and move toward deployment-ready systems. Code is available at https://reeshoon.github.io/robustgaitbenchmark

cs.CV

Corrosion-resistant and conductive Ti-Nb-O coatings tailored for ultra-low Pt-loaded BPPs and PTLs in PEM electrolyzers

We develop highly corrosion-resistant and conductive Ti-Nb-O coatings for metallic components -- bipolar plates (BPPs) and porous transport layers (PTLs) -- in PEM water electrolyzers. Using reactive high-power impulse magnetron sputtering (HiPIMS), we deposit compact 200 nm bilayer coatings onto SS316L substrates, systematically tailoring their composition. By precisely controlling oxygen partial pressure and Nb/Ti ratio, we adjust stoichiometry and structure, directly affecting electrical resistivity and corrosion resistance. We examine interfacial contact resistance (ICR) and electrochemical parameters before and after accelerated corrosion testing. Optimized coatings exhibit resistivity on the order of 10^-4 Ohmcm and extremely low corrosion current densities (J_corr = 0.01-0.08 uA/cm^2), well below the U.S. DOE 2026 target. Most importantly, these coatings enable the ICR target after accelerated corrosion testing with a Pt overlayer as thin as 5 nm, reducing Pt loading by up to two orders of magnitude compared to conventional approaches.

cond-mat.mtrl-sci

Direct observation of propagating spin waves in a spin-Hall nano-oscillator

Constriction-based spin Hall nano-oscillators (SHNOs) show great promise for application as highly tunable microwave sources with straightforward scalability toward large coupled networks. However, details of the magnetization dynamics within SHNOs have thus far not been addressed experimentally, due to the minute time and length scales involved. In this work, we present direct imaging of the magnetization dynamics within a single CoFeB-based SHNO using time-resolved scanning transmission X-ray microscopy (STXM). Our measurements reveal that the magnon amplitude is the strongest at the two constriction edges, with a pronounced assymetry favoring one edge, and that emitted spin waves exhibit strongly anisotropic propagation. Micromagnetic simulations suggest that grain boundaries and the Dzyaloshinskii-Moriya interaction (DMI) play a key role in both effects. Furthermore, the magnetodynamics changed during the measurement, indicating that the CoFeB/MgO interface may be more susceptible to X-ray induced modifications than previously recognized, challenging its presumed radiation hardness.

cond-mat.mes-hall

Scheduling for TWDM-EPON-Based Fronthaul Without a Dedicated Registration Wavelength

The adoption of Centralized Radio Access Network (C-RAN) architectures requires fronthaul systems capable of carrying large volumes of radio data while meeting stringent delay and jitter requirements. Ethernet Passive Optical Networks (EPONs) have emerged as a promising fronthaul solution due to their cost efficiency and compatibility with existing infrastructure. However, the traditional registration process for EPON systems halts the ongoing data transmissions during the registration period, thereby violating the enhanced Common Public Radio Interface (eCPRI) delay and jitter requirements. This limitation has been acknowledged by the ITU-T, which recommends the use of a dedicated wavelength channel for registration, leading to inefficient bandwidth utilization. In this paper, we propose a novel scheduling framework for a Time and Wavelength Division Multiplexed (TWDM) EPON-based fronthaul that enables periodic registration without wasting an additional wavelength channel. Performance evaluation demonstrates that the proposed method achieves up to a 71\% increase in the number of Radio Units (RUs) supported for a given number of wavelength channels, compared to a baseline scheme employing a dedicated registration wavelength.

cs.NI

Giant Damping-like Spin-Torque Conductivity in a GeTe/Py van der Waals Heterostructure

Recent observations of large unconventional spin-orbit torques in van der Waals (vdW) materials are driving intense interest for energy-efficient spintronic applications. A key limitation of ferromagnet (FM)/vdW heterostructures is their lower value of damping-like torque conductivity ($σ{\rm_{DL}^{y}}$) compared to the conventional heavy metal-based systems, limiting their prospects for commercial spintronic devices. Here, we report both a giant $σ{\rm_{DL}^{y}}$ of $-(1.25 \pm 0.11)\times 10^{5}~\hbar/ 2e~Ω^{-1}$m$^{-1}$ and an unconventional spin-orbit torque in a heterostructure comprising an FM (Ni$_{80}$Fe$_{20}$) and the vdW material GeTe. The value of $σ{\rm_{DL}^{y}}$ represents the highest reported torque conductivity for any FM/vdW interface and is comparable to benchmark heavy metal heterostructures. First-principles calculations reveal that this substantial torque originates from the cooperative interplay of the spin Hall effect, orbital Hall effect, and orbital Rashba effect, assisted by interfacial charge transfer. These findings demonstrate the potential of carefully engineered vdW heterostructures to achieve highly efficient electrical manipulation of magnetization at room temperature, paving the way for next-generation low-power spintronic devices.

cond-mat.mes-hall

A Gap Between Decision Trees and Neural Networks

We study when geometric simplicity of decision boundaries, used here as a notion of interpretability, can conflict with accurate approximation of axis-aligned decision trees by shallow neural networks. Decision trees induce rule-based, axis-aligned decision regions (finite unions of boxes), whereas shallow ReLU networks are typically trained as score models whose predictions are obtained by thresholding. We analyze the infinite-width, bounded-norm, single-hidden-layer ReLU class through the Radon total variation ($\mathrm{R}\mathrm{TV}$) seminorm, which controls the geometric complexity of level sets. We first show that the hard tree indicator $1_A$ has infinite $\mathrm{R}\mathrm{TV}$. Moreover, two natural split-wise continuous surrogates--piecewise-linear ramp smoothing and sigmoidal (logistic) smoothing--also have infinite $\mathrm{R}\mathrm{TV}$ in dimensions $d>1$, while Gaussian convolution yields finite $\mathrm{R}\mathrm{TV}$ but with an explicit exponential dependence on $d$. We then separate two goals that are often conflated: classification after thresholding (recovering the decision set) versus score learning (learning a calibrated score close to $1_A$). For classification, we construct a smooth barrier score $S_A$ with finite $\mathrm{R}\mathrm{TV}$ whose fixed threshold $τ=1$ exactly recovers the box. Under a mild tube-mass condition near $\partial A$, we prove an $L_1(P)$ calibration bound that decays polynomially in a sharpness parameter, along with an explicit $\mathrm{R}\mathrm{TV}$ upper bound in terms of face measures. Experiments on synthetic unions of rectangles illustrate the resulting accuracy--complexity tradeoff and how threshold selection shifts where training lands along it.

cs.LG

Metrics for spin-based computing

Spin-based computing is emerging as a powerful approach for energy-efficient and high-performance solutions to future data processing hardware. Spintronic devices function by electrically manipulating the collective dynamics of the electron spin, that is inherently non-volatile, nonlinear and fast-operating, and can couple to other degrees of freedom such as photonic and phononic systems. This review explores key advances in integrating magnetic and spintronic elements into computational architectures, ranging from fundamental components like radio-frequency neurons/synapses and spintronic probabilistic-bits to broader frameworks such as reservoir computing and magnetic Ising machines. We discuss hardware-specific and task-dependent metrics to evaluate the computing performance of spin-based components and associate them with physical properties. Finally, we discuss challenges and future opportunities, highlighting the potential of spin-based computing in next-generation technologies.

cond-mat.mes-hall

Dictionary Learning: The Complexity of Learning Sparse Superposed Features with Feedback

The success of deep networks is crucially attributed to their ability to capture latent features within a representation space. In this work, we investigate whether the underlying learned features of a model can be efficiently retrieved through feedback from an agent, such as a large language model (LLM), in the form of relative \tt{triplet comparisons}. These features may represent various constructs, including dictionaries in LLMs or a covariance matrix of Mahalanobis distances. We analyze the feedback complexity associated with learning a feature matrix in sparse settings. Our results establish tight bounds when the agent is permitted to construct activations and demonstrate strong upper bounds in sparse scenarios when the agent's feedback is limited to distributional information. We validate our theoretical findings through experiments on two distinct applications: feature recovery from Recursive Feature Machines and dictionary extraction from sparse autoencoders trained on Large Language Models.

cs.LG

Application of machine learning to predict food processing level using Open Food Facts

Ultra-processed foods are increasingly linked to health issues like obesity, cardiovascular disease, type 2 diabetes, and mental health disorders due to poor nutritional quality. This first-of-its-kind study at such a scale uses machine learning to classify food processing levels (NOVA) based on the Open Food Facts dataset of over 900,000 products. Models including LightGBM, Random Forest, and CatBoost were trained on nutrient concentration data. LightGBM performed best, achieving 80-85% accuracy across different nutrient panels and effectively distinguishing minimally from ultra-processed foods. Exploratory analysis revealed strong associations between higher NOVA classes and lower Nutri-Scores, indicating poorer nutritional quality. Products in NOVA 3 and 4 also had higher carbon footprints and lower Eco-Scores, suggesting greater environmental impact. Allergen analysis identified gluten and milk as common in ultra-processed items, posing risks to sensitive individuals. Categories like Cakes and Snacks were dominant in higher NOVA classes, which also had more additives, highlighting the role of ingredient modification. This study, leveraging the largest dataset of NOVA-labeled products, emphasizes the health, environmental, and allergenic implications of food processing and showcases machine learning's value in scalable classification. A user-friendly web tool is available for NOVA prediction using nutrient data: https://cosylab.iiitd.edu.in/foodlabel/.

q-bio.BM

Bayesian Learning Aided Simultaneous Sparse Estimation of Dual-Wideband THz Channels in Multi-User Hybrid MIMO Systems

This work conceives the Bayesian Group-Sparse Regression (BGSR) for the estimation of a spatial and frequency wideband, i.e., a dual wideband channel in Multi-User (MU) THz hybrid MIMO scenarios. We develop a practical dual wideband THz channel model that incorporates absorption losses, reflection losses, diffused ray modeling and angles of arrival/departure (AoAs/AoDs) using a Gaussian Mixture Model (GMM). Furthermore, a low-resolution analog-to-digital converter (ADC) is employed at each RF chain, which is crucial for wideband THz massive MIMO systems to reduce power consumption and hardware complexity, given the high sampling rates and large number of antennas involved. The quantized MU THz MIMO model is linearized using the popular Bussgang decomposition followed by BGSR based channel learning framework that results in sparsity across different subcarriers, where each subcarrier has its unique dictionary matrix. Next, the Bayesian Cramér Rao Bound (BCRB) is devised for bounding the normalized mean square error (NMSE) performance. Extensive simulations were performed to assess the performance improvements achieved by the proposed BGSR method compared to other sparse estimation techniques. The metrics considered for quantifying the performance improvements include the NMSE and bit error rate (BER).

eess.SP

Mutual synchronization of two asymmetric-nano-constriction-based spin-Hall nano-oscillators

We propose an asymmetric-nanoconstriction (ANC) design of spin-Hall nano-oscillators (SHNOs) and investigate mutual synchronization of a pair of such devices using micromagnetic simulations. The ANC geometry enables strong dipolar coupling at sub-50 nm separations while preserving independent current bias for each oscillator. We first characterize the auto-oscillation of a single ANC-SHNO, revealing a broad frequency tuning range and a field-controlled crossover between negative and positive nonlinearities. We then demonstrate that two such oscillators can mutually synchronize solely via dipolar stray fields, without electrical or spin-wave coupling. Depending on the bias conditions, the coupled pair exhibits robust in-phase (0°) or out-of-phase (180°) locking. Notably, we find a bias-dependent amplitude correlation: when the oscillators sustain comparable amplitudes, both in-phase and out-of-phase synchronization are accessible, whereas amplitude imbalance drives the system into an out-of-phase state accompanied by suppression of the weaker oscillator. By combining strong conservative coupling with independent frequency and gain control, the ANC-SHNO platform provides a scalable route toward phased oscillator arrays, neuromorphic computing architectures, and experimental exploration of non-Hermitian spintronic dynamics.

cond-mat.mes-hall

A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis

Recent works have characterized the function-space inductive bias of infinite-width bounded-norm single-hidden-layer neural networks as a kind of bounded-variation-type space. This novel neural network Banach space encompasses many classical multivariate function spaces, including certain Sobolev spaces and the spectral Barron spaces. Notably, this Banach space also includes functions that exhibit less classical regularity, such as those that only vary in a few directions. On bounded domains, it is well-established that the Gaussian reproducing kernel Hilbert space (RKHS) strictly embeds into this Banach space, demonstrating a clear gap between the Gaussian RKHS and the neural network Banach space. It turns out that when investigating these spaces on unbounded domains, e.g., all of $\mathbb{R}^d$, the story is fundamentally different. We establish the following fundamental result: Certain functions that lie in the Gaussian RKHS have infinite norm in the neural network Banach space. This provides a nontrivial gap between kernel methods and neural networks by exhibiting functions that kernel methods easily represent, whereas neural networks cannot.

cs.LG

OmViD: Omni-supervised active learning for video action detection

Video action detection requires dense spatio-temporal annotations, which are both challenging and expensive to obtain. However, real-world videos often vary in difficulty and may not require the same level of annotation. This paper analyzes the appropriate annotation types for each sample and their impact on spatio-temporal video action detection. It focuses on two key aspects: 1) how to obtain varying levels of annotation for videos, and 2) how to learn action detection from different annotation types. The study explores video-level tags, points, scribbles, bounding boxes, and pixel-level masks. First, a simple active learning strategy is proposed to estimate the necessary annotation type for each video. Then, a novel spatio-temporal 3D-superpixel approach is introduced to generate pseudo-labels from these annotations, enabling effective training. The approach is validated on UCF101-24 and JHMDB-21 datasets, significantly cutting annotation costs with minimal performance loss.

cs.CV

Approximating Dasgupta Cost in Sublinear Time from a Few Random Seeds

Testing graph cluster structure has been a central object of study in property testing since the foundational work of Goldreich and Ron [STOC'96] on expansion testing, i.e. the problem of distinguishing between a single cluster (an expander) and a graph that is far from a single cluster. More generally, a $(k, ε)$-clusterable graph $G$ is a graph whose vertex set admits a partition into $k$ induced expanders, each with outer conductance bounded by $ε$. A recent line of work initiated by Czumaj, Peng and Sohler [STOC'15] has shown how to test whether a graph is close to $(k, ε)$-clusterable, and to locally determine which cluster a given vertex belongs to with misclassification rate $\approx ε$, but no sublinear time algorithms for learning the structure of inter-cluster connections are known. As a simple example, can one locally distinguish between the `cluster graph' forming a line and a clique? In this paper, we consider the problem of testing the hierarchical cluster structure of $(k, ε)$-clusterable graphs in sublinear time. Our measure of hierarchical clusterability is the well-established Dasgupta cost, and our main result is an algorithm that approximates Dasgupta cost of a $(k, ε)$-clusterable graph in sublinear time, using a small number of randomly chosen seed vertices for which cluster labels are known. Our main result is an $O(\sqrt{\log k})$ approximation to Dasgupta cost of $G$ in $\approx n^{1/2+O(ε)}$ time using $\approx n^{1/3}$ seeds, effectively giving a sublinear time simulation of the algorithm of Charikar and Chatziafratis [SODA'17] on clusterable graphs. To the best of our knowledge, ours is the first result on approximating the hierarchical clustering properties of such graphs in sublinear time.

cs.DS

AxOSyn: An Open-source Framework for Synthesizing Novel Approximate Arithmetic Operators

Edge AI deployments are becoming increasingly complex, necessitating energy-efficient solutions for resource-constrained embedded systems. Approximate computing, which allows for controlled inaccuracies in computations, is emerging as a promising approach for improving power and energy efficiency. Among the key techniques in approximate computing are approximate arithmetic operators (AxOs), which enable application-specific optimizations beyond traditional computer arithmetic hardware reduction-based methods, such as quantization and precision scaling. Existing design space exploration (DSE) frameworks for approximate computing limit themselves to selection-based approaches or custom synthesis at fixed abstraction levels, which restricts the flexibility required for finding application-specific optimal solutions. Further, the tools available for the DSE of AxOs are quite limited in terms of exploring different approximation models and extending the analysis to different granularities. To this end, we propose AxOSyn, an open-source framework for the DSE of AxOs that supports both selection and synthesis approaches at various abstraction levels. AxOSyn allows researchers to integrate custom methods for evaluating approximations and facilitates DSE at both the operator-level and application-specific. Our framework provides an effective methodology for achieving energy-efficient, approximate operators.

cs.AR

MixA-Q: Revisiting Activation Sparsity for Vision Transformers from a Mixed-Precision Quantization Perspective

In this paper, we propose MixA-Q, a mixed-precision activation quantization framework that leverages intra-layer activation sparsity (a concept widely explored in activation pruning methods) for efficient inference of quantized window-based vision transformers. For a given uniform-bit quantization configuration, MixA-Q separates the batched window computations within Swin blocks and assigns a lower bit width to the activations of less important windows, improving the trade-off between model performance and efficiency. We introduce a Two-Branch Swin Block that processes activations separately in high- and low-bit precision, enabling seamless integration of our method with most quantization-aware training (QAT) and post-training quantization (PTQ) methods, or with simple modifications. Our experimental evaluations over the COCO dataset demonstrate that MixA-Q achieves a training-free 1.35x computational speedup without accuracy loss in PTQ configuration. With QAT, MixA-Q achieves a lossless 1.25x speedup and a 1.53x speedup with only a 1% mAP drop by incorporating activation pruning. Notably, by reducing the quantization error in important regions, our sparsity-aware quantization adaptation improves the mAP of the quantized W4A4 model (with both weights and activations in 4-bit precision) by 0.7%, reducing quantization degradation by 24%.

cs.CV