SearcharxivSearch

arXiv subjects

Prashant Gupta

Publications and source records attributed to Prashant Gupta.

15 recordsLinked to original sources

Learning ON Large Datasets Using Bit-String Trees

This thesis develops computational methods in similarity-preserving hashing, classification, and cancer genomics. Standard space partitioning-based hashing relies on Binary Search Trees (BSTs), but their exponential growth and sparsity hinder efficiency. To overcome this, we introduce Compressed BST of Inverted hash tables (ComBI), which enables fast approximate nearest-neighbor search with reduced memory. On datasets of up to one billion samples, ComBI achieves 0.90 precision with 4X-296X speed-ups over Multi-Index Hashing, and also outperforms Cellfishing.jl on single-cell RNA-seq searches with 2X-13X gains. Building on hashing structures, we propose Guided Random Forest (GRAF), a tree-based ensemble classifier that integrates global and local partitioning, bridging decision trees and boosting while reducing generalization error. Across 115 datasets, GRAF delivers competitive or superior accuracy, and its unsupervised variant (uGRAF) supports guided hashing and importance sampling. We show that GRAF and ComBI can be used to estimate per-sample classifiability, which enables scalable prediction of cancer patient survival. To address challenges in interpreting mutations, we introduce Continuous Representation of Codon Switches (CRCS), a deep learning framework that embeds genetic changes into numerical vectors. CRCS allows identification of somatic mutations without matched normals, discovery of driver genes, and scoring of tumor mutations, with survival prediction validated in bladder, liver, and brain cancers. Together, these methods provide efficient, scalable, and interpretable tools for large-scale data analysis and biomedical applications.

cs.LG

Majorana Zero Modes in a Heterogenous Structure of Topological and Trivial Domains in FeSe$_{1-x}$Te$_x$

We propose that the existence of vortices in FeSe$_{1-x}$Te$_x$ with and without Majoarana zero modes (MZMs) can be explained by a heterogeneous mixture of strong topological and trivial superconducting domains, with only vortices in the former exhibiting MZMs. We identify the spectroscopic signatures of topological and trivial vortices and show that they are necessarily separated by a domain wall harboring Majorana edge modes. We demonstrate that when a vortex is moved from a trivial to a topological domain in real time, a domain wall Majorana edge mode is transferred to the vortex as an MZM.

cond-mat.mes-hall

Cost-Effective, Low Latency Vector Search with Azure Cosmos DB

Vector indexing enables semantic search over diverse corpora and has become an important interface to databases for both users and AI agents. Efficient vector search requires deep optimizations in database systems. This has motivated a new class of specialized vector databases that optimize for vector search quality and cost. Instead, we argue that a scalable, high-performance, and cost-efficient vector search system can be built inside a cloud-native operational database like Azure Cosmos DB while leveraging the benefits of a distributed database such as high availability, durability, and scale. We do this by deeply integrating DiskANN, a state-of-the-art vector indexing library, inside Azure Cosmos DB NoSQL. This system uses a single vector index per partition stored in existing index trees, and kept in sync with underlying data. It supports < 20ms query latency over an index spanning 10 million vectors, has stable recall over updates, and offers approximately 43x and 12x lower query cost compared to Pinecone and Zilliz serverless enterprise products. It also scales out to billions of vectors via automatic partitioning. This convergent design presents a point in favor of integrating vector indices into operational databases in the context of recent debates on specialized vector databases, and offers a template for vector indexing in other databases.

cs.DB

Braiding of Majorana Zero Modes in Vortex Cores

We demonstrate the successful simulation of $\sqrt{Z}$-, $\sqrt{X}$- and $X$-quantum gates using Majorana zero modes (MZMs) that emerge in magnetic vortices located in topological superconductors. We compute the transition probabilities and geometric phase differences accounting for the full many-body dynamics and show that qubit states can be read out by fusing the vortex core MZMs and measuring the resulting charge density. We visualize the gate processes using the time- and energy-dependent non-equilibrium local density of states. Our results demonstrate the feasibility of employing vortex core MZMs for the realization of fault-tolerant topological quantum computing.

cond-mat.mes-hall

Memristors based Computation and Synthesis

Memristor has been identified as the fourth fundamental circuit element by Dr. Leon Chua in 1971 and since then it has gathered a lot of interest because of its non-volatility and are considered as a viable solution to the beyond CMOS era computation. Recently, memristor have been used to perform basic logic operations like AND, OR, NAND, NOR, XOR etc. and are also used in applications like Dot Product Engine, Convolution Neural Networks etc. This paper presents a new behavioural model of memristor then using it to build a 32-bit ripple carry adder. The paper later compares the area, power and time delay of the 32 bit Ripple Carry Adder using memristor with the 45nm CMOS technology and highlights its advantages and pitfalls.

eess.SY

Box Filtration

We define a new framework that unifies the filtration and mapper approaches from TDA, and present efficient algorithms to compute it. Termed the box filtration of a PCD, we grow boxes (hyperrectangles) that are not necessarily centered at each point (in place of balls centered at points). We grow the boxes non-uniformly and asymmetrically in different dimensions based on the distribution of points. We present two approaches to handle the boxes: a point cover where each point is assigned its own box at start, and a pixel cover that works with a pixelization of the space of the PCD. Any box cover in either setting automatically gives a mapper of the PCD. We show that the persistence diagrams generated by the box filtration using both point and pixel covers satisfy the classical stability based on the Gromov-Hausdorff distance. Using boxes also implies that the box filtration is identical for pairwise or higher order intersections whereas the VR and Cech filtration are not the same. Growth in each dimension is computed by solving a linear program (LP) that optimizes a cost functional balancing the cost of expansion and benefit of including more points in the box. The box filtration algorithm runs in $O(m|U(0)|\log(mn\pi)L(q))$ time, where $m$ is number of steps of increments considered for box growth, $|U(0)|$ is the number of boxes in the initial cover ($\leq$ number of points), $\pi$ is the step length for increasing each box dimension, each LP is solved in $O(L(q))$ time, $n$ is the PCD dimension, and $q = n \times |X|$. We demonstrate through multiple examples that the box filtration can produce more accurate results to summarize the topology of the PCD than VR and distance-to-measure (DTM) filtrations. Software for our implementation is available at https://github.com/pragup/Box-Filteration.

cs.CG

Subgap two-particle spectral weight in disordered $s$-wave superconductors: Insights from mode coupling approach

We study the two-particle spectral functions and collective modes of weakly disordered superconductors using a disordered attractive Hubbard model on square lattice. We show that the disorder induced scattering between collective modes leads to a finite subgap spectral weight in the long wavelength limit. In general, the spectral weight is distributed between the phase and the Higgs channels, but as we move towards half-filling the Higgs contribution dominates. The inclusion of the density fluctuations lowers the frequency at which this mode occurs, and results in the phase channel gaining a larger contribution to this subgap mode. Near half-filling, the proximity of the system to the charge density wave (CDW) instability leads to strong fluctuations of the effective disorder at the commensurate wave-vector ($[\pi,\pi]$). We develop an analytical mode coupling approach where the pure Goldstone mode in the long wavelength limit couples to the collective mode at $[\pi,\pi]$. This provides insight into the location and distribution of the two-particle spectral weights between the Higgs and the phase channels.

cond-mat.supr-con

SFCDecomp: Multicriteria Optimized Tool Path Planning in 3D Printing using Space-Filling Curve Based Domain Decomposition

We explore efficient optimization of toolpaths based on multiple criteria for large instances of 3D printing problems. We first show that the minimum turn cost 3D printing problem is NP-hard, even when the region is a simple polygon. We develop SFCDecomp, a space filling curve based decomposition framework to solve large instances of 3D printing problems efficiently by solving these optimization subproblems independently. For the Buddha model, our framework builds toolpaths over a total of 799,716 nodes across 169 layers, and for the Bunny model it builds toolpaths over 812,733 nodes across 360 layers. Building on SFCDecomp, we develop a multicriteria optimization approach for toolpath planning. We demonstrate the utility of our framework by maximizing or minimizing tool path edge overlap between adjacent layers, while jointly minimizing turn costs. Strength testing of a tensile test specimen printed with tool paths that maximize or minimize adjacent layer edge overlaps reveal significant differences in tensile strength between the two classes of prints.

cs.GR

Enhash: A Fast Streaming Algorithm For Concept Drift Detection

We propose Enhash, a fast ensemble learner that detects \textit{concept drift} in a data stream. A stream may consist of abrupt, gradual, virtual, or recurring events, or a mixture of various types of drift. Enhash employs projection hash to insert an incoming sample. We show empirically that the proposed method has competitive performance to existing ensemble learners in much lesser time. Also, Enhash has moderate resource requirements. Experiments relevant to performance comparison were performed on 6 artificial and 4 real data sets consisting of various types of drifts.

cs.LG

A Weighted Mutual k-Nearest Neighbour for Classification Mining

kNN is a very effective Instance based learning method, and it is easy to implement. Due to heterogeneous nature of data, noises from different possible sources are also widespread in nature especially in case of large-scale databases. For noise elimination and effect of pseudo neighbours, in this paper, we propose a new learning algorithm which performs the task of anomaly detection and removal of pseudo neighbours from the dataset so as to provide comparative better results. This algorithm also tries to minimize effect of those neighbours which are distant. A concept of certainty measure is also introduced for experimental results. The advantage of using concept of mutual neighbours and distance-weighted voting is that, dataset will be refined after removal of anomaly and weightage concept compels to take into account more consideration of those neighbours, which are closer. Consequently, finally the performance of proposed algorithm is calculated.

cs.LG

Guided Random Forest and its application to data approximation

We present a new way of constructing an ensemble classifier, named the Guided Random Forest (GRAF) in the sequel. GRAF extends the idea of building oblique decision trees with localized partitioning to obtain a global partitioning. We show that global partitioning bridges the gap between decision trees and boosting algorithms. We empirically demonstrate that global partitioning reduces the generalization error bound. Results on 115 benchmark datasets show that GRAF yields comparable or better results on a majority of datasets. We also present a new way of approximating the datasets in the framework of random forests.

cs.LG

Continuous Toolpath Planning in Additive Manufacturing

We develop a framework that creates a new polygonal mesh representation of the sparse infill domain of a layer-by-layer 3D printing job. We guarantee the existence of a single, continuous tool path covering each connected piece of the domain in every layer. We present a tool path algorithm that traverses each such continuous tool path with no crossovers. The key construction at the heart of our framework is an Euler transformation which converts a 2-dimensional cell complex K into a new 2-complex K^ such that every vertex in the 1-skeleton G^ of K^ has even degree. Hence G^ is Eulerian, and a Eulerian tour can be followed to print all edges in a continuous fashion. We start with a mesh K of the union of polygons obtained by projecting all layers to the plane. We compute its Euler transformation K^. In the slicing step, we clip K^ at each layer using its polygon to obtain a complex that may not necessarily be Euler. We then patch this complex by adding edges such that any odd-degree nodes created by slicing are transformed to have even degrees again. We print extra support edges in place of any segments left out to ensure there are no edges without support in the next layer. These support edges maintain the Euler nature of the complex. Finally we describe a tree-based search algorithm that builds the continuous tool path by traversing "concentric" cycles in the Euler complex. Our algorithm produces a tool path that avoids material collisions and crossovers, and can be printed in a continuous fashion irrespective of complex geometry or topology of the domain (e.g., holes). We implement our test our framework on several 3D objects. Apart from standard geometric shapes, we demonstrate the framework on the Stanford bunny.

cs.CG

Pentagon at MEDIQA 2019: Multi-task Learning for Filtering and Re-ranking Answers using Language Inference and Question Entailment

Parallel deep learning architectures like fine-tuned BERT and MT-DNN, have quickly become the state of the art, bypassing previous deep and shallow learning methods by a large margin. More recently, pre-trained models from large related datasets have been able to perform well on many downstream tasks by just fine-tuning on domain-specific datasets . However, using powerful models on non-trivial tasks, such as ranking and large document classification, still remains a challenge due to input size limitations of parallel architecture and extremely small datasets (insufficient for fine-tuning). In this work, we introduce an end-to-end system, trained in a multi-task setting, to filter and re-rank answers in the medical domain. We use task-specific pre-trained models as deep feature extractors. Our model achieves the highest Spearman's Rho and Mean Reciprocal Rank of 0.338 and 0.9622 respectively, on the ACL-BioNLP workshop MediQA Question Answering shared-task.

cs.IR

Euler Transformation of Polyhedral Complexes

We propose an Euler transformation that transforms a given $d$-dimensional cell complex $K$ for $d=2,3$ into a new $d$-complex $\hat{K}$ in which every vertex is part of a uniform even number of edges. Hence every vertex in the graph $\hat{G}$ that is the $1$-skeleton of $\hat{K}$ has an even degree, which makes $\hat{G}$ Eulerian, i.e., it is guaranteed to contain an Eulerian tour. Meshes whose edges admit Eulerian tours are crucial in coverage problems arising in several applications including 3D printing and robotics. For $2$-complexes in $\mathbb{R}^2$ ($d=2$) under mild assumptions (that no two adjacent edges of a $2$-cell in $K$ are boundary edges), we show that the Euler transformed $2$-complex $\hat{K}$ has a geometric realization in $\mathbb{R}^2$, and that each vertex in its $1$-skeleton has degree $4$. We bound the numbers of vertices, edges, and $2$-cells in $\hat{K}$ as small scalar multiples of the corresponding numbers in $K$. We prove corresponding results for $3$-complexes in $\mathbb{R}^3$ under an additional assumption that the degree of a vertex in each $3$-cell containing it is $3$. In this setting, every vertex in $\hat{G}$ is shown to have a degree of $6$. We also present bounds on parameters measuring geometric quality (aspect ratios, minimum edge length, and maximum angle) of $\hat{K}$ in terms of the corresponding parameters of $K$ (for $d=2, 3$). Finally, we illustrate a direct application of the proposed Euler transformation in additive manufacturing.

cs.CG

NIRS Based Bladder Volume Sensing for Patients Suffering with Neurogenic Bladder Dysfunction

Neurogenic Bladder Dysfunction has detrimental effects on day-to-day life of millions of people. Some of the most common symptoms faced by these patients include urinary incontinence, urgency and retention. Since elevated bladder pressure due to prolonged urine storage inside bladder may have adverse impacts on patient's renal health, urologists recommend clean-intermittent catheterization (CIC) every 2 to 4 hours throughout the day to relieve bladder pressure. However, since urine production by kidneys is an intermittent process and most of these patients have limited mobility, such frequent trips to washroom can prove to be challenging. Sometimes, bladder fills to capacity before the recommended CIC time is reached causing embarrassing situation due to leakage. Hence, time-based CIC strategy is difficult to implement and has high chances of failure. As such, continence is the primary concern for most of these patients but sadly there are no practical solutions available in the market that address this concern. A real-time notification system that could give feedback to patients on when "bladder is almost-full" could help these patients to better plan their bathroom trips. This work explores the feasibility of using a near infrared-light based wearable, non-invasive spectroscopy technique that can sense amount of urine present inside the bladder and give details on developing a bladder state estimation device. We present preliminary results by testing our device on optical phantoms and performing ex vivo measurements on porcine bladder and intestines. We later explored the possibility of using the device on human subjects, after study was approved by the UC Davis Institution Review Board (IRB).

eess.SP