SearcharxivSearch

arXiv subjects

Varun Sharma

Publications and source records attributed to Varun Sharma.

18 recordsLinked to original sources

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators

The popularity of large language models (LLMs) escalates an ongoing demand for effective inference. However, due to the sequential processing of tokens during the token phase in decoder-only LLMs inference, the inherent low parallelism leads to reduced throughput and suboptimal utilization of the computing units on artificial intelligence (AI) accelerators, particularly when handling long-sequence inputs that impose significant memory overhead. Recently, many reported methods have been developed as potential solutions, since they emerge with numeric deviation. This paper presents FastTPS, a high performance and low-precision loss method for accelerating the token-phase in LLM inference on general AI accelerators which includes three key components: (1) AI accelerator-enabled reloading-free KV Cache concatenation which decreases memory access overhead as well as enables full fusion of Attention, (2) high-efficiency and high-accuracy 'RoPE' attention based on the tiling optimized FLAT, and (3) highly-fused MLP with fine-grain pipeline scheduling. Our results confirm that FastTPS significantly alleviates memory bottlenecks in the token phase, delivering a 6x speed improvement (compared to none-fusion) on an AMD Ryzen AI 300 series NPU with BF16 precision while sustaining 93% peak memory bandwidth utilization during Phi3-mini-4k-instruct inference.

cs.LG

TIER: Trajectory-Invariant Explanation Regularization for Membership Privacy

Explainability is central to building trustworthy AI, yet explanation interfaces can inadvertently provide adversaries with an expanded privacy-related attack surfaces. Recent studies show that advanced membership-inference attacks succeed by exploiting confidence-drop trajectories, induced through attribution-guided perturbations, as discriminative features, rather than directly using confidence scores or explanation vectors. Existing defenses against membership inference fail to directly mitigate such explanation-driven attacks. In this work, we investigate whether, during training, a model's own gradients can be leveraged as defense signals against such attacks, thereby aligning explanation profiles between members and non-members. To this end, we propose a Trajectory-Invariant Explanation Regularization (TIER) defense that penalizes erratic fluctuations in confidence drops simulated through gradient-guided perturbations and simultaneously minimizes the distributional shifts via KL-divergence. Unlike conventional adversarial training, which emphasizes label robustness, our approach targets explanation robustness by enforcing self-consistency through KL-divergence and reducing the variance of confidence drops between members and non-members. Extensive experiments confirm that our method effectively mitigates these attacks, delivering privacy protection while maintaining model utility and explanation fidelity.

cs.CR

STAG: Spatio-temporal Evolving Structural Representation of Action Units for Micro-expression Recognition

Micro-expression recognition is challenging due to subtle and short-lived facial muscle movements. Existing methods rely heavily on apex-onset frames, overlook fine-grained inter-frame dynamics, and separately model spatial and temporal information, limiting generalization across datasets. To address these challenges, we propose STAG, a dynamic ROI-AU-coupled spatial-temporal network that jointly models motion flow and adaptive facial connectivity. The framework extracts optical flow from discriminative frames using magnitude-based selection and temporal attention. A dual-branch architecture combines an enhanced graph attention network for structured spatial reasoning with a transformer encoder for temporal modeling. A bidirectional cross-attention module enables mutual refinement of spatial and temporal features, while AU-guided dynamic connectivity adapts facial region interactions according to muscle activation patterns. The transformer captures subtle temporal dynamics beyond apex-based approaches, improving semantic consistency and interpretability for explainable micro-expression recognition. The fused representation is optimized using focal loss and evaluated on CASME II, 4DME, DFME, NaME, SAMM, and SMIC-HS. Extensive experiments demonstrate improved robustness, generalization, interpretability, and computational efficiency, confirming the effectiveness of adaptive relational reasoning, AU-guided dynamic connectivity, and deep spatial-temporal feature fusion for accurate cross-dataset micro-expression recognition.

cs.CV

DE-FIVE: Detecting Malicious Image Prompts via Fourier Features and Image Vector Embeddings

Vision language models (VLMs) employ both visual and textual modalities to enable advanced vision-language inference. However, incorporating visual modalities expands the attack surface of VLMs, making them more susceptible to security threats such as adversarial perturbations and indirect prompt injection, wherein crafted malicious image prompts can elicit unintended model outputs. Existing defense methods against malicious image prompts remain insufficient as they typically demand extensive datasets for retraining or the deployment of additional, complex classifiers. Most critically, there is a profound lack of specialized defense mechanisms specifically targeting indirect prompt injections, a gap that serves as a primary motivation for this work. To address these limitations, we introduce DE-FIVE, a novel training-free framework for detecting malicious image prompts by leveraging Fourier features and the hidden state representations of the visual encoder (image vector embeddings) across perturbations. Specifically, we develop a hybrid detection strategy consisting of a black-box detector that operates on Fourier-domain features and a white-box detector that exploits image vector embeddings derived from only a few-shot malicious set. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art baselines against malicious image prompts.

cs.CR

Visualising the Attractor Landscape of Neural Cellular Automata

As Neural Cellular Automata (NCAs) are increasingly applied outside of the toy models in Artificial Life, there is a pressing need to understand how they behave and to build appropriate routes to interpret what they have learnt. By their very nature, the benefits of training NCAs are balanced with a lack of interpretability: we can engineer emergent behaviour, but have limited ability to understand what has been learnt. In this paper, we apply a variety of techniques to pry open the NCA black box and glean some understanding of what it has learnt to do. We apply techniques from manifold learning (principal components analysis and both dense and sparse autoencoders) along with techniques from topological data analysis (persistent homology) to capture the NCA's underlying behavioural manifold, with varying success. Results show that when analysis is performed at a macroscopic level (i.e. taking the entire NCA state as a single data point), the underlying manifold is often quite simple and can be captured and analysed quite well. When analysis is performed at a microscopic level (i.e. taking the state of individual cells as a single data point), the manifold is highly complex and more complicated techniques are required in order to make sense of it.

cs.NE

Observation of stability of Gaussian beams and off-axis beam-cleaning in graded-index media

While nonlinear effects in graded-index (GRIN) multimode fibers have been studied extensively, little is known about nonlinear effects in larger GRIN waveguides, where the number of modes approaches infinity and modal dispersion becomes negligible. Here we show that Gaussian beams remain nearly invariant even with large nonlinear phase accumulation and on- or off-axis trajectories in GRIN rods. In addition, spatially-complex beams can undergo self-cleaning to single-lobed profiles for both on- and off-axis trajectories. Numerical simulations exhibit the features observed in experiments, and a general interpretation of these results that makes connection to beam-cleaning phenomena observed in GRIN fibers is proposed.

physics.optics

Speech and Text-Based Emotion Recognizer

Affective computing is a field of study that focuses on developing systems and technologies that can understand, interpret, and respond to human emotions. Speech Emotion Recognition (SER), in particular, has got a lot of attention from researchers in the recent past. However, in many cases, the publicly available datasets, used for training and evaluation, are scarce and imbalanced across the emotion labels. In this work, we focused on building a balanced corpus from these publicly available datasets by combining these datasets as well as employing various speech data augmentation techniques. Furthermore, we experimented with different architectures for speech emotion recognition. Our best system, a multi-modal speech, and text-based model, provides a performance of UA(Unweighed Accuracy) + WA (Weighed Accuracy) of 157.57 compared to the baseline algorithm performance of 119.66

cs.CL

Photonic integrated processor for structured light detection and distinction

Integrated photonic devices have become pivotal elements across most research fields that involve light-based applications. A particularly versatile category of this technology are programmable photonic integrated processors, which are being employed in an increasing variety of applications, like communication or photonic computing. Such processors accurately control on-chip light within meshes of programmable optical gates. Free-space optics applications can utilize this technology by using appropriate on-chip interfaces to couple distributions of light to the photonic chip. This enables, for example, access to the spatial properties of free-space light, particularly to phase distributions, which is usually challenging and requires either specialized devices or additional components. Here we discuss and show the detection of amplitude and phase of structured higher-order light beams using a multipurpose photonic processor. Our device provides measurements of amplitude and phase distributions which can be used to, e.g., directly distinguish light's orbital angular momentum without the need for further elements interacting with the free-space light. Paving a way towards more convenient and intuitive phase measurements of structured light, we envision applications in a wide range of fields, specifically in microscopy or communications where the spatial distributions of lights properties are important.

physics.optics

Steam Recommendation System

We aim to leverage the interactions between users and items in the Steam community to build a game recommendation system that makes personalized suggestions to players in order to boost Steam's revenue as well as improve the users' gaming experience. The whole project is built on Apache Spark and deals with Big Data. The final output of the project is a recommendation system that gives a list of the top 5 items that the users will possibly like.6

cs.IR

Near-video frame rate quantum sensing using Hong-Ou-Mandel interferometry

Hong-Ou-Mandel (HOM) interference, the bunching of two indistinguishable photons on a balanced beam-splitter, has emerged as a promising tool for quantum sensing. There is a need for wide spectral-bandwidth photon pairs (for high-resolution sensing) with high brightness (for fast sensing). Here we show the generation of photon-pairs with flexible spectral-bandwidth even using single-frequency, continuous-wave diode laser enabling high-precision, real-time sensing. Using 1-mm-long periodically-poled KTP crystal, we produced degenerate, photon-pairs with spectral-bandwidth of 163.42$\pm$1.68 nm resulting in a HOM-dip width of 4.01$\pm$0.04 $\mu$m to measure a displacement of 60 nm, and sufficiently high brightness to enable the measurement of vibrations with amplitude of $205\pm0.75$ nm and frequency of 8 Hz. Fisher-information and maximum likelihood estimation enables optical delay measurements as small as 4.97 nm with precision (Cram\'er-Rao bound) and accuracy of 0.89 and 0.54 nm, respectively, therefore showing HOM sensing capability for real-time, precision-augmented, in-field quantum sensing applications.

physics.optics

Generating free-space structured light with programmable integrated photonics

Structured light is a key component of many modern applications, ranging from superresolution microscopy to imaging, sensing, and quantum information processing. As the utilization of these powerful tools continues to spread, the demand for technologies that enable the spatial manipulation of fundamental properties of light, such as amplitude, phase, and polarization grows further. In this respect, technologies based on liquid-crystal cells, e.g., spatial light modulators, became very popular in the last decade. However, the rapidly advancing field of integrated photonics allows entirely new routes towards beam shaping that not only outperform liquid-crystal devices in terms of speed, but also have substantial potential with respect to robustness and conversion efficiencies. In this study, we demonstrate how a programmable integrated photonic processor can generate and control higher-order free-space structured light beams at the click of a button. Our system offers lossless and reconfigurable control of the spatial distribution of light's amplitude and phase, with switching times in the microsecond domain. The showcased on-chip generation of spatially tailored light enables an even more diverse set of methods, applications, and devices that utilize structured light by providing a pathway towards combining the strengths of programmable integrated photonics and free-space structured light.

physics.optics

Literature on Hand GESTURE Recognition using Graph based methods

Skeleton based recognition systems are gaining popularity and machine learning models focusing on points or joints in a skeleton have proved to be computationally effective and application in many areas like Robotics. It is easy to track points and thereby preserving spatial and temporal information, which plays an important role in abstracting the required information, classification becomes an easy task. In this paper, we aim to study these points but using a cloud mechanism, where we define a cloud as collection of points. However, when we add temporal information, it may not be possible to retrieve the coordinates of a point in each frame and hence instead of focusing on a single point, we can use k-neighbors to retrieve the state of the point under discussion. Our focus is to gather such information using weight sharing but making sure that when we try to retrieve the information from neighbors, we do not carry noise with it. LSTM which has capability of long-term modelling and can carry both temporal and spatial information. In this article we tried to summarise graph based gesture recognition method.

cs.CV

Summarizing experimental sensitivities of collider experiments to dark matter models and comparison to other experiments

Comparisons of the coverage of current and proposed dark matter searches can help us to understand the context in which a discovery of particle dark matter would be made. In some scenarios, a discovery could be reinforced by information from multiple, complementary types of experiments; in others, only one experiment would see a signal, giving only a partial, more ambiguous picture; in still others, no experiment would be sensitive and new approaches would be needed. In this whitepaper, we present an update to a similar study performed for the European Strategy Briefing Book performed within the dark matter at the Energy Frontier (EF10) Snowmass Topical Group We take as a starting point a set of projections for future collider facilities and a method of graphical comparisons routinely performed for LHC DM searches using simplified models recommended by the LHC Dark Matter Working Group and also used for the BSM and dark matter chapters of the European Strategy Briefing Book. These comparisons can also serve as launching point for cross-frontier discussions about dark matter complementarity.

hep-ph

Prospects for Heavy WIMP Dark Matter Searches at Muon Colliders

Plots summarizing the constraints on Dark Matter models can help visualize synergies between different searches for the same kind of experiment, as well as between different experiments. In this whitepaper, we present an update to the European Strategy Briefing Book plots, from the perspective of collider searches within the Dark Matter at the Energy Frontier (EF10) Snowmass Topical Group, starting from inputs from future collider facilities. We take as a starting point the plots currently made for LHC searches using benchmark models recommended by the Dark Matter Working Group, also used for the BSM and Dark Matter chapters of the European Strategy Briefing Book. These plots can also serve as a starting point for cross-frontier discussions about dark matter complementarity, and could be updated as a consequence of these discussions. This is a whitepaper submitted to the APS Snowmass process for the EF10 topical group.

hep-ex

Imaging inspired characterization of single photons carrying orbital angular momentum

We report on an imaging-inspired measurement of orbital angular momentum (OAM) using only a simple tilted lens and an Intensified Charged Coupled Device (ICCD) camera, allowing us to monitor the propagation of OAM structured photons over distance, crucial for free-space quantum communication networks. We demonstrate measurement of OAM orders as high as 14 in a heralded single-photon source (HSPS) and show, for the first time, the imaged self-interference of photons carrying OAM in a modified Mach-Zehnder Interferometer (MZI). The described methods reveal both the charge and order of a photons OAM, and provide a proof of concept for the interference of a single OAM photon with itself. Using these tools, we are able to study the propagation characteristics of OAM photons over distance, important for estimating transport in free-space quantum links. By translating these classical tools into the quantum domain, we offer a robust and direct approach for the complete characterization of a twisted single-photon source, an important building block of a quantum network.

quant-ph

A PGAS Communication Library for Heterogeneous Clusters

This work presents a heterogeneous communication library for clusters of processors and FPGAs. This library, Shoal, supports the Partitioned Global Address Space (PGAS) memory model for applications. PGAS is a shared memory model for clusters that creates a distinction between local and remote memory access. Through Shoal and its common application programming interface for hardware and software, applications can be more freely migrated to the optimal platform and deployed onto dynamic cluster topologies. The library is tested using a thorough suite of microbenchmarks to establish latency and throughput performance. We also show an implementation of the Jacobi iterative method that demonstrates the ease with which applications can be moved between platforms to yield faster run times. Through this work, we have demonstrated the feasibility of using a PGAS programming model for multi-node heterogeneous platforms.

cs.DC

Fast Intent Classification for Spoken Language Understanding

Spoken Language Understanding (SLU) systems consist of several machine learning components operating together (e.g. intent classification, named entity recognition and resolution). Deep learning models have obtained state of the art results on several of these tasks, largely attributed to their better modeling capacity. However, an increase in modeling capacity comes with added costs of higher latency and energy usage, particularly when operating on low complexity devices. To address the latency and computational complexity issues, we explore a BranchyNet scheme on an intent classification scheme within SLU systems. The BranchyNet scheme when applied to a high complexity model, adds exit points at various stages in the model allowing early decision making for a set of queries to the SLU model. We conduct experiments on the Facebook Semantic Parsing dataset with two candidate model architectures for intent classification. Our experiments show that the BranchyNet scheme provides gains in terms of computational complexity without compromising model accuracy. We also conduct analytical studies regarding the improvements in the computational cost, distribution of utterances that egress from various exit points and the impact of adding more complexity to models with the BranchyNet scheme.

cs.CL

High-power, continuous-wave, tunable mid-IR, higher-order vortex beam optical parametric oscillator

We report on a novel experimental scheme to generate continuous-wave (cw), high power, and higher-order optical vortices tunable across mid-IR wavelength range. Using cw, two-crystal, singly resonant optical parametric oscillator (T-SRO) and pumping one of the crystals with Gaussian beam and the other crystal with optical vortices of orders, lp = 1 to 6, we have directly transferred the vortices at near-IR to the mid-IR wavelength range. The idler vortices of orders, li = 1 to 6, are tunable across 2276-3576 nm with a maximum output power of 6.8 W at order of, li = 1, for the pump power of 25 W corresponding to a near-IR vortex to mid-IR vortex conversion efficiency as high as 27.2%. Unlike the SROs generating optical vortices restricted to lower orders due to the elevated operation threshold with pump vortex orders, here, the coherent energy coupling between the resonant signals of the crystals of T-SRO facilitates the transfer of pump vortex of any order to the idler wavelength without stringent operation threshold condition. The generic experimental scheme can be used in any wavelength range across the electromagnetic spectrum and in all time scales from cw to ultrafast regime.

physics.optics