Searcharxiv⌕ Search

arXiv subjects

Abhishek Singh

Publications and source records attributed to Abhishek Singh.

At least 19 recordsLinked to original sources

Evaluation of portability and performance of an OpenMP5 offloaded Quantum-Inspired Evolutionary Optimization Across the GPU Ecosystem

Quantum-inspired evolutionary optimization (QIEO) is a new class of population-based metaheuristic optimization algorithms which represents design variables as a set of qubits and searches a continuous, multi-dimensional landscape through rotation of the qubit's amplitude pair. Every generation rotates those amplitudes toward a single elite, which corresponds to that generation's best. The per-generation cost scales as $O(N_p N_g)$ for $N_p$ chromosomes and $N_g$ genes (decision variables). Production use of such solvers is rarely confined to a single machine class. Prototypes are run on laboratory servers- before moving to rented cloud workstations for more involved campaigns. The largest problems are reserved for leadership-class accelerators. This paper asks whether a \emph{single} OpenMP~5 source of QIEO, offloaded with \texttt{\#pragma omp target}, is a viable production path in each of those settings. We report three independent, campaigns of the 0/1 knapsack problem against a same-source multi-core Intel CPU baseline. The study comprises approximately 3,000 runs spanning varying chromosome and gene counts, evaluated using both chromosome-level and gene-level offload strategies on the NVIDIA Tesla V100 SXM2, NVIDIA A100 80GB, and AMD Instinct MI300X GPUs. Deployment-specific nuances such as Volta's constant-memory cliffs, Ampere's L2 persistence and \texttt{cp.async}, CDNA~3's Infinity Cache and XCD occupancy are addressed to ensure high performance of these platforms. Results reveal gene-parallel offload achieved geometric-mean speedups of 90$\times$, 136$\times$, and 155$\times$ over a single CPU core on the V100, A100, and MI300X, respectively, and 12$\times$, 17$\times$, and 16.6$\times$ over 72 host threads. Furthermore DetermineElite, the $O(N_p)$ selection of the generation-best chromosome, is found to be better suited to the host than to the device.

cs.PF↗

Cross-Backend QIEO: Universal Runtime Portability across OpenMP5, CUDA, HIP, and Multi-Language Interfaces

Quantum-inspired algorithms emulate quantum mechanical principles, such as, superposition, interference, and probabilistic amplitude evolution, on classical hardware by representing candidate solutions as qubit vectors and evolving them through rotation-gate operators. This approach offers higher optimization performance without physical qubits, and has been shown to achieve order-of-magnitude speedups (10--80$\times$) over traditional solvers on combinatorial, high-dimensional NP-hard problems. A critical barrier to adoption, however, is the lack of a unified execution framework that delivers both algorithmic performance and hardware portability. We present \textbf{Cross-Backend Quantum Inspired Evolutionary Optimizer (QIEO)}, the runtime core of BQP's BQPhy solver, which addresses this gap through a \emph{single-source-of-truth} architecture. One C++ implementation of the QIEO algorithm is compiled once per hardware target and exposed to multiple high-level languages via thin binding layers. The framework dispatches to CPU (sequential), OpenMP~5 (multi-core), CUDA (NVIDIA), and HIP (AMD) backends at runtime, adapting kernels to each device's memory hierarchy and warp/wavefront execution model. The framework's real-world utility is validated through binding demonstrations that share the identical C++ runtime. BQPhy's Python library is demonstrated on a neural network hyperparameter optimisation achieving 88.60\% test accuracy on MNIST. BQPhy's MATLAB's Toolkit is tested on wind farm layout optimisation attaining $365\,399 \pm 4\,552$~MWh/yr, which is statistically indistinguishable from particle swarm optimisation and $+7.6\%$ above genetic algorithms on a 32-variable constrained engineering problem. The Julia package tackles the Lotka--Volterra parameter estimation where BQPhy replaces native Julia solvers on the same residual, cutting mean SSE by $2.1\times$.

cs.DC↗

Landscape Limits of Quantum-Inspired Evolutionary Optimization across 256 continuous functions

Quantum-inspired evolutionary optimization (QIEO) represents design variables as a set of qubits and searches a continuous, multi-dimensional landscape through rotation of the qubit's amplitude pair. Every generation rotates those amplitudes toward a single elite, which corresponds to that generation's best. The update is cheap, almost parameter-free, and well-suited for massive parallel implementation, which has encouraged its adoption in engineering, design, and planning applications. However, there are critical issues with this formulation, principally, the treatment of design variables as independent probability components which make it incapable of exploiting local curvature, anisotropy, or variable coupling. Despite this, QIEO is believed to hold promise, and has been used extensively to solve real-world problems, with significant qualitative and computational advantage over its classical counterpart, Genetic Algorithm (GA). A collection of 256 (actually 508; 256 unshifted + 252 shifted, 4 could not be shifted) continuous function are selected from the prior works, in such a way that they represent eleven landscape characteristics, namely continuity, differentiability, separability, scalability, modality, convexity, conditioning, symmetry, maximum dimensionality, dimension dependency, and the coupling pattern of the design variables. These functions are then solved by three QIEO variants, two GA encodings and Hansen's Covariance Matrix Adaptation Evolution Strategy (CMA-ES). The results are evaluated in terms of computational cost, solution precision, and specialization across landscape characteristics. They identify the conditions under which QIEO provides competitive performance, clarify where its independent-variable representation becomes limiting, and establish whether particular QIEO variants offer advantages for specific landscape characteristics.

cs.NE↗

Quantum Variational Approaches to the Maximum Independent Set Problem at Utility Scale

Near-optimal solutions to Maximum Independent Set on dense graphs sit in local optima that greedy correction and maximality extension cannot escape. We encode near-optimal seeds as a uniform quantum superposition on ancilla qubits and evolve them under an excitation-preserving variational ansatz that holds the search inside the feasible Hamming-weight subspace. A preprocessing stage of spectral reordering and distance-based sparsification, together with history-guided post-processing of the sampled bitstrings, takes the method to 200 nodes. The ansatz entangles the seed branches and the bond dimension of the simulated state grows with circuit depth, which places deeper circuits outside the reach of exact matrix product state simulation at the bond dimensions available to us. Larger instances therefore need quantum hardware. Measured on the data register alone, the superposition behaves as a classical mixture over the seeds, so coherence between the branches has to be created and then looked for. We do this with a CRZ phase layer and post-selection on the ancilla qubits, which brings the branches into interference and exposes the cross terms. The idea is to see if interference between near-optimal seeds widens the range of independent sets the circuit returns. Standard VQE with this pipeline recovers the certified optimum for instances up to 125 nodes, and these run on ibm_marrakesh with parameters transferred from noiseless simulation. The ancilla construction is introduced for the sizes past that point. On a 180-node hard instance the superposition recovers the certified MIS. Five 200-node instances are solved to the certified optimum, and on the 400-node brock400-1 benchmark the method reaches size 25 against a certified optimum of 27. These are the largest hard general-graph instances we know of where a gate-based variational algorithm optimises the full circuit directly.

quant-ph↗

Shaping SHAPE - A spectro-polarimeter onboard Chandrayaan-3 to observe Earth as an Exoplanet

Spectro-polarimetry of HAbitable Planet Earth (SHAPE) is an experimental instrument onboard the Propulsion Module (Orbiter) of the Chandrayaan-3 mission, designed to perform disc-integrated spectro-polarimetric observations of Earth from lunar and highly elliptical Earth orbits. SHAPE is a compact, lightweight spectro-polarimeter comprising three subsystems: the Electro-Optical Detector System (EODS)-Optics, EODS-Electronics, and Radio Frequency Source (RFS). An Acousto-Optic Tunable Filter (AOTF), driven by an in-house-developed 80$-$135 MHz RF source, provides spectral filtering in the near-infrared (NIR) wavelength range of 1.0$-$1.7 $μ$m and produces two narrow-band beams with mutually perpendicular linear polarization states. The instrument optics, with a field of view of approximately 2.6°, focus the two beams onto InGaAs detectors. A spectral resolution of 2$-$4 nm is achieved using in-house-designed low-noise front-end electronics. The instrument also incorporates processing and power electronics for signal processing, detector biasing, and subsystem control. We present the overall instrument design, results from pre-launch ground-based testing, and in-orbit operational performance. The current configuration enables SHAPE to measure disc-integrated signatures of Earth over a range of phase angles, providing a test bed for characterizing Earth-like exoplanets and benchmarking future exoplanet observations.

astro-ph.IM↗

Free energy landscape of Dense Associative Memory

Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative memories, including dense associative memories. We illustrate the method by reproducing classical results for the Hopfield model. For a finite number of patterns, we derive the temperature-dependent free energy functional for dense associative memories featuring polynomial interactions and Log-Sum-Exponential (LSE) activation. We also evaluate the disorder-averaged ground-state energy of these systems in the extensive limit. Our analytical framework reveals how memory retrieval depends on the initial state in higher-order dense networks, and gives the exact full-retrieval threshold for the LSE model. This method provides a systematic procedure for analyzing diverse, complex architectures in associative memory.

cond-mat.dis-nn↗

KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assistants

Mass religious gatherings such as the Kumbh Mela concentrate tens of millions of people into a single region over a few weeks, producing intense, repetitive, multilingual, and safety-critical demand for information. The default response, a conversational assistant that routes every query to a large language model (LLM), is poorly matched to this setting: it is costly at scale, slow on emergency paths, prone to hallucination on facts that can cause physical harm, and unusable when connectivity fails. We describe KumbhDoot, an agentic pilgrim assistant for the Nashik Simhastha Kumbh Mela built on a different principle. It operates on a foundational design principle that prioritizes semantic similarity over starting with an LLM. Generative models are invoked only in instances where similarity-based retrieval is insufficient to produce a correct answer. The system utilizes a "semantic cache": an embedding-indexed store as a single retrieval primitive, which handles intent routing, answer caching, offline lookups, and multi-agent retrieval. A custom three-tier agent architecture operates directly on this store, ensuring decision paths remain inspectable and avoiding the use of generic multi-agent frameworks that would trigger implicit per-step LLM calls. We present the architecture, an analytical cost model for its per-query economics, and an honest account of where similarity is sufficient and where generative reasoning remains necessary. We argue that for bounded, high-stakes, low-connectivity public-service domains, a similarity-first and LLM-bounded design is not merely cheaper but architecturally more appropriate than an LLM-default one.

cs.CY↗

Hamiltonian-Guided Leverage Embedding: Robust Subspace Compression for Efficient QAOA Parameter Estimation

The Quantum Approximate Optimization Algorithm (QAOA) is a hybrid quantum-classical framework for combinatorial optimization on near-term quantum devices. A central bottleneck is the classical estimation of its variational parameters γ and β, which must be optimized over a high-dimensional, non-convex landscape corrupted by sampling noise. We observe that the classical feature matrices constructed from QAOA measurement samples exhibit pronounced low-rank structure, and exploit this property for noise-robust, reduced-dimension parameter search. We present the Hamiltonian-Guided Leverage Embedding (HGLE) algorithm - a hybrid pipeline that encodes low-energy quantum samples into a weighted Ising feature matrix and compresses it via leverage-score row sampling, provably preserving the dominant rank-rsubspace geometry. The compressed representation drives a classical trust-region loop for (γ, β) estimation at a fraction of the original cost. We provide formal guarantees for rank preservation and energy approximation error, and demonstrate robustness across problem types (Max-Cut, Maximum Independent Set) and graph topologies of varying density.

quant-ph↗

Chiral molecule-induced contributions to ferromagnetic resonance

Despite extensive research on chirality-driven spin selectivity, most studies have focused on static magnetic properties, while the influence of chirality on the dynamic magnetic response remains largely unexplored. Here, we investigate how chiral molecular interfaces affect magnetization dynamics in thin Co/Ni multilayers with perpendicular magnetic anisotropy using broadband ferromagnetic resonance spectroscopy. A comparison between bare (reference) films and molecule-functionalized (hybrid) samples reveals no measurable changes in either the resonance field or the linewidth that could be attributed to the presence of the chiral environment. Motivated by our findings we develop a macrospin description that distinguishes equilibrium modifications of the magnetic free-energy landscape (MIPAC-type effects) from non-equilibrium, CISS-induced spin torques. Our analysis shows that equilibrium modifications primarily shift the resonance condition via changes to the free energy landscape and thereby the effective field, whereas damping-like non-equilibrium torques provide a distinct channel for varying the effective damping rate. This approach establishes clear criteria for disentangling chiral-interface-induced energy modifications from torque-driven dynamical effects in ferromagnetic resonance experiments.

cond-mat.mtrl-sci↗

Bayesian Aneurysm Growth Detection via Surface Displacement Modeling

Clinical decisions for unruptured intracranial aneurysms depend on detecting growth on follow-up magnetic resonance angiography (MRA). Growth is typically judged from manual 2D diameters on few slices, which vary across clinicians and frequently miss subtle 3D change. Even with 3D segmentations, apparent differences can reflect resolution, segmentation, surface processing, or registration mismatch rather than true growth; most criteria remain heuristic and binary. We show that a Bayesian displacement-based model using the surrounding vessel as an internal reference achieves strong discrimination of aneurysm growth (AUC 0.86-0.87) and improves agreement with expert labels (Cohen's kappa up to 0.66 vs. 0.35 for volumetric criteria), while providing calibrated posterior probabilities with uncertainty bounds. The method registers baseline and follow-up surfaces, computes normal-directed displacements, and summarizes change as the difference between mean aneurysm displacement and mean displacement on the surrounding non-aneurysmal vessel segment. The vessel segment serves as an internal control for imaging and processing variability, assuming negligible structural change over the surveillance interval. We evaluate two cohorts spanning time-of-flight and contrast-enhanced longitudinal MRA studies: a public dataset labeled from neuroradiologist-provided measurements and an institutional dataset labeled by senior and junior raters. Performance is preserved when training on lower-expertise labels, indicating robustness to label variability. Calibrated probabilities may aid clinical decision-making in borderline cases, where high uncertainty can motivate repeat imaging. This framework provides interpretable probabilistic growth assessment from longitudinal MRA, reduces dependence on clinician expertise, and supports cross-center surveillance across scanners and angiography sequences.

physics.med-ph↗

AMES: Approximate Multi-modal Enterprise Search via Late Interaction Retrieval

We present AMES (Approximate Multimodal Enterprise Search), a unified multimodal late interaction retrieval architecture which is backend agnostic. AMES demonstrates that fine-grained multimodal late interaction retrieval can be deployed within a production grade enterprise search engine without architectural redesign. Text tokens, image patches, and video frames are embedded into a shared representation space using multi-vector encoders, enabling cross-modal retrieval without modality specific retrieval logic. AMES employs a two-stage pipeline: parallel token level ANN search with per document Top-M MaxSim approximation, followed by accelerator optimized Exact MaxSim re-ranking. Experiments on the ViDoRe V3 benchmark show that AMES achieves competitive ranking performance within a scalable, production ready Solr based system.

cs.IR↗

SMURF: Scalable method for unsupervised reconstruction of flow in 4D flow MRI

We introduce SMURF, a scalable and unsupervised machine learning method for simultaneously segmenting vascular geometries and reconstructing velocity fields from 4D flow MRI data. SMURF models geometry and velocity fields using multilayer perceptron-based functions incorporating Fourier feature embeddings and random weight factorization to accelerate convergence. A measurement model connects these fields to the observed image magnitude and phase data. Maximum likelihood estimation and subsampling enable SMURF to process high-dimensional datasets efficiently. Evaluations on synthetic, in vitro, and in vivo datasets demonstrate SMURF's performance. On synthetic internal carotid artery aneurysm data derived from CFD, SMURF achieves a quarter-voxel segmentation accuracy across noise levels of up to 50%, outperforming the state-of-the-art segmentation method by up to double the accuracy. In an in vitro experiment on Poiseuille flow, SMURF reduces velocity reconstruction RMSE by approximately 34% compared to raw measurements. In in vivo internal carotid artery aneurysm data, SMURF attains nearly half-voxel segmentation accuracy relative to expert annotations and decreases median velocity divergence residuals by about 31%, with a 27% reduction in the interquartile range. These results indicate that SMURF is robust to noise, preserves flow structure, and identifies patient-specific morphological features. SMURF advances 4D flow MRI accuracy, potentially enhancing the diagnostic utility of 4D flow MRI in clinical applications.

physics.med-ph↗

Controlling HER activity and stability of $γ$- and 6,6,12-Graphyne through engineered B-N doping: DFT and Reactive MD simulations

Graphynes offer a chemically heterogeneous $sp/sp^{2}$ carbon framework with distinct electronic regimes and site-selective reactivity. Here, Density Functional Theory and Reactive Molecular Dynamics Simulations are combined to evaluate pristine, B-doped, N-doped, and B-N co-doped $γ$-graphyne and 6,6,12-graphyne (meta/ortho/para). $γ$-graphyne is a semiconductor, while 6,6,12-graphyne exhibits an anisotropic Dirac-like semi-metallic dispersion. B/N substitution reconstructs near-$E_F$ states via dopant $π$ hybridization, and B-N pairing stabilizes defects through donor-acceptor compensation, with the ortho substitutions being the most favorable. Hydrogen adsorption remains weak on pristine lattices but becomes locally optimized upon doping, with near thermo-neutral $ΔG_{\mathrm{ads}}$ 'hot spots' predominantly on $sp$-proximate carbon sites adjacent to the dopants. Reactive MD at 300 K further reveals an activity stability trade-off: B-N ortho in $γ$-graphyne sustains controlled hydrogen uptake without catastrophic bond scission, whereas B-N meta/para degrade, and 6,6,12-graphyne is generally more susceptible to over-hydrogenation. These results identify the B-N geometry as a key design variable for graphyne-based HER catalysts, which require both a favorable $ΔG_{\mathrm{ads}}$ and finite-temperature hydrogenation stability.

cond-mat.mtrl-sci↗

VAST: Vascular Flow Analysis and Segmentation for Intracranial 4D Flow MRI

Four-dimensional (4D) Flow MRI can noninvasively measure cerebrovascular hemodynamics but remains underused clinically because current workflows rely on manual vessel segmentation and yield velocity fields sensitive to noise, artifacts, and phase aliasing. We present VAST (Vascular Flow Analysis and Segmentation), an automated, unsupervised pipeline for intracranial 4D Flow MRI that couples vessel segmentation with physics-informed velocity reconstruction. VAST derives vessel masks directly from complex 4D Flow data by iteratively fusing magnitude- and phase-based background statistics. It then reconstructs velocities via continuity-constrained phase unwrapping, outlier correction, and low-rank denoising to reduce noise and aliasing while promoting mass-consistent flow fields, with processing completing in minutes per case on a standard CPU. We validate VAST on synthetic data from an internal carotid artery aneurysm model across SNR = 2-20 and severe phase wrapping (up to five-fold), on in vitro Poiseuille flow, and on an in vivo internal carotid aneurysm dataset. In synthetic benchmarks, VAST maintains near quarter-voxel surface accuracy and reduces velocity root-mean-square error by up to fourfold under the most degraded conditions. In vitro, it segments the channel within approximately half a voxel of expert annotations and reduces velocity error by 39% (unwrapped) and 77% (aliased). In vivo, VAST closely matches expert time-of-flight masks and lowers divergence residuals by about 30%, indicating a more self-consistent intracranial flow field. By automating processing and enforcing basic flow physics, VAST helps move intracranial 4D Flow MRI toward routine quantitative use in cerebrovascular assessment.

eess.IV↗

Terahertz emission and detection using Ge-on-Si photoconductive antennas

Germanium-on-Silicon (Ge-on-Si) is a promising, CMOS-compatible platform for integrated terahertz (THz) photonics, offering a low-cost alternative to III-V semiconductors. A primary challenge for Ge-based photoconductive antennas (PCAs), however, has been the long carrier lifetime of bulk Ge, preventing its use as a detector. Here, we demonstrate that amorphous Ge (a-Ge) films overcome this limitation, possessing inherent ultrashort carrier lifetimes ~ 1.11-1.38 ps. We leverage this property to demonstrate, for the first time to our knowledge, coherent THz pulse detection using undoped a-Ge-on-Si PCAs. We present a comparative study of devices fabricated on a-Ge films grown by plasma-enhanced chemical vapor deposition (PECVD) and DC magnetron sputtering. The PECVD-Ge device, with better homogeneity and a smoother morphology in the films, demonstrates superior performance for both THz emission and detection. As an emitter, the PECVD-Ge PCA achieves a 40 dB signal-to-noise ratio (SNR) with a bandwidth of ~ 3 THz. As a detector, it achieves a 32 dB SNR and a ~ 2 THz bandwidth, representing a ~2.5-fold increase in detected signal amplitude over the sputtered-Ge device. These results establish amorphous Ge-on-Si as a viable and scalable platform for both THz generation and detection, paving the way for fully integrated Si-based THz time-domain systems.

physics.optics↗

Terahertz emission from interdigitated photoconductive antennas based on Ge-on-Si

An interdigitated photoconductive antenna (i-PCA) for terahertz (THz) emission with a novel metal-insulator-semiconductor interface is designed with the aim of developing compact and scalable THz devices. The photoconductive material is an amorphous germanium (Ge) film deposited using DC magnetron sputtering. The antenna electrodes are composed of gold-germanium (AuGe). With the integration of a silicon dioxide (SiO2) layer that acts as an electrical mask on alternate active areas, we present a simple approach to fabricate a large-area i-PCA. Along with a simplified fabrication compared to other existing designs, our approach increases the electrical robustness of the emitter and reduces the inactive gap area on the device. The i-PCA is capable of THz emission up to 2.5 THz and 36 dB signal-to-noise ratio (SNR), and is promising for applications in CMOS technologies.

physics.optics↗

A novel method to analyze pattern shifts in rainfall using cluster analysis and probability models

: One of the prominent challenges being faced by agricultural sciences is the onset of climate change which is adversely affecting every aspect of cropping. Modelling of climate change at macro level have been carried out at large scale and there is ample amount of research publications available for that. But at micro level like at state level or district level there are lesser studies. District level studies can help in preparing specific plans for the mitigation of adverse effects of climate change at local level. An attempt has been made in this paper to model the monthly rainfall of Varanasi district of the state of Uttar Pradesh with the help of probability models. Firstly, the pattern of the climate change over 122 years has been unveiled by using exploratory analysis and using multivariate techniques like cluster analysis and then probability models have been fitted for selected months

stat.AP↗

Investigation of Performance and Scalability of a Quantum-Inspired Evolutionary Optimizer (QIEO) on NVIDIA GPU

Quantum inspired evolutionary optimization leverages quantum computing principles like superposition, interference, and probabilistic representation to enhance classical evolutionary algorithms with improved exploration and exploitation capabilities. Implemented on NVIDIA Tesla V100 SXM2 GPUs, this study systematically investigates the performance and scalability of a GPU-accelerated Quantum Inspired Evolutionary Optimizer applied to large scale 01 Knapsack problems. By exploiting CUDA`s parallel processing capabilities, particularly through optimized memory management and thread configuration, significant speedups and efficient utilization of GPU resources is demonstrated. The analysis covers various problem sizes, kernel launch configurations, and memory models including constant, shared, global, and pinned memory, alongside extensive scaling studies. The results reveal that careful tuning of memory strategies and kernel configurations is essential for maximizing throughput and efficiency, with constant memory providing superior performance up to hardware limits. Beyond these limits, global memory and strategic tiling become necessary, albeit with some performance trade offs. The findings highlight both the promise and the practical constraints of applying QIEO on GPUs for complex combinatorial optimization, offering actionable insights for future large scale metaheuristic implementations.

cs.CE↗