SearcharxivSearch

arXiv subjects

Kushal Shah

Publications and source records attributed to Kushal Shah.

At least 19 recordsLinked to original sources

Object Pose and Shape Estimation for Grasping: Does it Work?

The problem of object pose and shape estimation has seen key advancements lately. Encoder-decoder (e.g., SAM3D, LRM, CRISP) and diffusion-based models (e.g., InstantMesh, Zero123, SceneComplete) have shown category-agnostic shape encoding capacity and open-set generalizability. In this work, we ask the question: Are the object pose and shape estimation methods mature enough, such that when used with antipodal grasp sampling, can outperform the end-to-end grasp synthesis methods? We explore this question in detail by scoping our study to parallel jaw grippers, 7-DoF grasps, and single-view RGB(-D) image as input. We implement and compare a state-of-the-art, end-to-end grasp synthesis method and three modular methods, which first estimate the object pose and shape for all objects in the scene, and generate grasps using antipodal sampling. We observe that the modular methods outperform the end-to-end method in all our experiments. The modular methods are able to synthesize plenty of grasps, even for small objects, where the end-to-end methods fail. The effectiveness of the modular methods is contingent on the accuracy of the pose and shape estimation, and suffers partial degradation in cluttered scenes - a limitation of the existing pose and shape estimation methods. We also analyze the failure modes and run-times for the three modular methods, which use two different ways of object pose and shape estimation: one based on an encoder-decoder model, while another a diffusion model. Finally, we demonstrate that the single-view object pose and shape estimation methods can be augmented with vision-language models to yield language-conditioned grasps from just single-view RGB-D image as input. We notice comparable performance to the state-of-the-art LERF-TOGO baseline.

cs.RO

Neural ATTF: A Scalable Solution to Lifelong Multi-Agent Path Planning

Multi-Agent Pickup and Delivery (MAPD) is a fundamental problem in robotics, particularly in applications such as warehouse automation and logistics. Existing solutions often face challenges in scalability, adaptability, and efficiency, limiting their applicability in dynamic environments with real-time planning requirements. This paper presents Neural ATTF (Adaptive Task Token Framework), a new algorithm that combines a Priority Guided Task Matching (PGTM) Module with Neural STA* (Space-Time A*), a data-driven path planning method. Neural STA* enhances path planning by enabling rapid exploration of the search space through guided learned heuristics and ensures collision avoidance under dynamic constraints. PGTM prioritizes delayed agents and dynamically assigns tasks by prioritizing agents nearest to these tasks, optimizing both continuity and system throughput. Experimental evaluations against state-of-the-art MAPD algorithms, including TPTS, CENTRAL, RMCA, LNS-PBS, and LNS-wPBS, demonstrate the superior scalability, solution quality, and computational efficiency of Neural ATTF. These results highlight the framework's potential for addressing the critical demands of complex, real-world multi-agent systems operating in high-demand, unpredictable settings.

cs.RO

Enhancing Grammatical Error Detection using BERT with Cleaned Lang-8 Dataset

This paper presents an improved LLM based model for Grammatical Error Detection (GED), which is a very challenging and equally important problem for many applications. The traditional approach to GED involved hand-designed features, but recently, Neural Networks (NN) have automated the discovery of these features, improving performance in GED. Traditional rule-based systems have an F1 score of 0.50-0.60 and earlier machine learning models give an F1 score of 0.65-0.75, including decision trees and simple neural networks. Previous deep learning models, for example, Bi-LSTM, have reported F1 scores within the range from 0.80 to 0.90. In our study, we have fine-tuned various transformer models using the Lang8 dataset rigorously cleaned by us. In our experiments, the BERT-base-uncased model gave an impressive performance with an F1 score of 0.91 and accuracy of 98.49% on training data and 90.53% on testing data, also showcasing the importance of data cleaning. Increasing model size using BERT-large-uncased or RoBERTa-large did not give any noticeable improvements in performance or advantage for this task, underscoring that larger models are not always better. Our results clearly show how far rigorous data cleaning and simple transformer-based models can go toward significantly improving the quality of GED.

cs.CL

BioNeMo Framework: a modular, high-performance library for AI model development in drug discovery

Artificial Intelligence models encoding biology and chemistry are opening new routes to high-throughput and high-quality in-silico drug development. However, their training increasingly relies on computational scale, with recent protein language models (pLM) training on hundreds of graphical processing units (GPUs). We introduce the BioNeMo Framework to facilitate the training of computational biology and chemistry AI models across hundreds of GPUs. Its modular design allows the integration of individual components, such as data loaders, into existing workflows and is open to community contributions. We detail technical features of the BioNeMo Framework through use cases such as pLM pre-training and fine-tuning. On 256 NVIDIA A100s, BioNeMo Framework trains a three billion parameter BERT-based pLM on over one trillion tokens in 4.2 days. The BioNeMo Framework is open-source and free for everyone to use.

cs.LG

First-Generation Inference Accelerator Deployment at Facebook

In this paper, we provide a deep dive into the deployment of inference accelerators at Facebook. Many of our ML workloads have unique characteristics, such as sparse memory accesses, large model sizes, as well as high compute, memory and network bandwidth requirements. We co-designed a high-performance, energy-efficient inference accelerator platform based on these requirements. We describe the inference accelerator platform ecosystem we developed and deployed at Facebook: both hardware, through Open Compute Platform (OCP), and software framework and tooling, through Pytorch/Caffe2/Glow. A characteristic of this ecosystem from the start is its openness to enable a variety of AI accelerators from different vendors. This platform, with six low-power accelerator cards alongside a single-socket host CPU, allows us to serve models of high complexity that cannot be easily or efficiently run on CPUs. We describe various performance optimizations, at both platform and accelerator level, which enables this platform to serve production traffic at Facebook. We also share deployment challenges, lessons learned during performance optimization, as well as provide guidance for future inference hardware co-design.

cs.AR

Malaria detection from RBC images using shallow Convolutional Neural Networks

The advent of Deep Learning models like VGG-16 and Resnet-50 has considerably revolutionized the field of image classification, and by using these Convolutional Neural Networks (CNN) architectures, one can get a high classification accuracy on a wide variety of image datasets. However, these Deep Learning models have a very high computational complexity and so incur a high computational cost of running these algorithms as well as make it hard to interpret the results. In this paper, we present a shallow CNN architecture which gives the same classification accuracy as the VGG-16 and Resnet-50 models for thin blood smear RBC slide images for detection of malaria, while decreasing the computational run time by an order of magnitude. This can offer a significant advantage for commercial deployment of these algorithms, especially in poorer countries in Africa and some parts of the Indian subcontinent, where the menace of malaria is quite severe.

eess.IV

On the length scale dependence of DNA conformational change under local perturbation

Conformational change of a DNA molecule is frequently observed in multiple biological processes and has been modelled using a chain of strongly coupled oscillators with a nonlinear bistable potential. While the mechanism and properties of conformational change in the model have been investigated and several reduced order models developed, the conformational dynamics as a function of the length of the oscillator chain is relatively less clear. To address this, we used a modified Lindstedt-Poincare method and numerical computations. We calculate a perturbation expansion of the frequency of the model's nonzero modes, finding that approximating these modes with their unperturbed dynamics, as in a previous reduced order model, may not hold when the length of the DNA model increases. We investigate the conformational change to local perturbation in models of varying lengths, finding that for chosen input and parameters, there are two regions of DNA length in the model, first where the minimum energy required to undergo the conformational change increases with DNA length; and second, where it is almost independent of the length of the DNA model. We analyze the conformational change in these models by adding randomness to the local perturbation, finding that the tendency of the system to remain in a stable conformation against random perturbation decreases with an increase in the DNA length. These results should help to understand the role of the length of a DNA molecule in influencing its conformational dynamics.

physics.bio-ph

Computational prediction of replication sites in DNA sequences using complex number representation

Computational prediction of origin of replication (ORI) has been of great interest in bioinformatics and several methods including GC-skew, auto-correlation etc. have been explored in the past. In this paper, we have extended the auto-correlation method to predict ORI location with much higher resolution for prokaryotes and eukaryotes, which can be very helpful in experimental validation of the computational predictions. The proposed complex correlation method (iCorr) converts the genome sequence into a sequence of complex numbers by mapping the nucleotides to {+1,-1,+i,-i} instead of {+1,-1} used in the auto-correlation method (here, i is square root of -1). Thus, the iCorr method exploits the complete spatial information about the positions of all the four nucleotides unlike the earlier auto-correlation method which uses the positional information of only one nucleotide. Also, the earlier auto-correlation method required visual inspection of the obtained graphs to identify the location of origin of replication. The proposed iCorr method does away with this need and is able to identify the origin location simply by picking the peak in the iCorr graph.

q-bio.GN

Open-endedness in AI systems, cellular evolution and intellectual discussions

One of the biggest challenges that artificial intelligence (AI) research is facing in recent times is to develop algorithms and systems that are not only good at performing a specific intelligent task but also good at learning a very diverse of skills somewhat like humans do. In other words, the goal is to be able to mimic biological evolution which has produced all the living species on this planet and which seems to have no end to its creativity. The process of intellectual discussions is also somewhat similar to biological evolution in this regard and is responsible for many of the innovative discoveries and inventions that scientists and engineers have made in the past. In this paper, we present an information theoretic analogy between the process of discussions and the molecular dynamics within a cell, showing that there is a common process of information exchange at the heart of these two seemingly different processes, which can perhaps help us in building AI systems capable of open-ended innovation. We also discuss the role of consciousness in this process and present a framework for the development of open-ended AI systems.

cs.AI

Unifying averaged dynamics of the Fokker-Planck equation for Paul traps

Collective dynamics of a collisional plasma in a Paul trap is governed by the Fokker-Planck equation, which is usually assumed to lead to a unique asymptotic time-periodic solution irrespective of the initial plasma distribution. This uniqueness is, however, hard to prove in general due to analytical difficulties. For the case of small damping and diffusion coefficients, we apply averaging theory to a special solution to this problem, and show that the averaged dynamics can be represented by a remarkably simple 2D phase portrait, which is independent of the applied rf field amplitude. In particular, in the 2D phase portrait, we have two regions of initial conditions. From one region, all solutions are unbounded. From the other region, all solutions go to a stable fixed point, which represents a unique time-periodic solution of the plasma distribution function, and the boundary between these two is a parabola.

physics.plasm-ph

Smooth phase transition of energy equilibration in a springy Sinai billiard

Statistical equilibration of energies in a slow-fast system is a fundamental open problem in physics. In a recent paper, it was shown that the equilibration rate in a springy billiard can remain strictly positive in the limit of vanishing mass ratio (of the particle and billiard wall) when the frozen billiard has more than one ergodic components [Proc. Natl. Acad. Sci. USA 114, E10514 (2017)]. In this paper, using the model of a springy Sinai billiard, it is shown that this can happen even in the case where the frozen billiard has a single ergodic component, but when the time of ergodization in the frozen system is much longer than the time of equilibration. It is also shown that as the size of the disc in the Sinai billiard is increased from zero, thereby leading to a decrease in the time required for ergodization in the frozen system, the system undergoes a smooth phase transition in the equilibration rate dependence on mass ratio.

nlin.CD

Equilibration of energy in slow-fast systems

Ergodicity is a fundamental requirement for a dynamical system to reach a state of statistical equilibrium. On the other hand, it is known that in slow-fast systems ergodicity of the fast sub- system impedes the equilibration of the whole system due to the presence of adiabatic invariants. Here, we show that the violation of ergodicity in the fast dynamics effectively drives the whole system to equilibrium. To demonstrate this principle we investigate dynamics of the so-called springy billiards. These consist of a point particle of a small mass which bounces elastically in a billiard where one of the walls can move - the wall is of a finite mass and is attached to a spring. We propose a random process model for the slow wall dynamics and perform numerical experiments with the springy billiards themselves and the model. The experiments show that for such systems equilibration is always achieved; yet, in the adiabatic limit, the system equilibrates with a positive exponential rate only when the fast particle dynamics has more than one ergodic component for certain wall positions.

math.DS

Analysis and validation of low-frequency noise reduction in MOSFET circuits using variable duty cycle switched biasing

Randomization of the trap state of defects present at the gate Si-SiO$_2$ interface of MOSFET is responsible for the low-frequency noise phenomena such as Random Telegraph Signal (RTS), burst, and 1/\textit{f} noise. In a previous work, theoretical modelling and analysis of the RTS noise in MOS transistor was presented and it was shown that this 1/\textit{f} noise can be reduced by decreasing the duty cycle ($f_{D}$) of switched biasing signal. In this paper, an extended analysis of this 1/\textit{f} noise reduction model is presented and it is shown that the RTS noise reduction is accompanied with shift in the corner frequency ($f_{c}$) of the 1/\textit{f} noise and the value of shift is a function of continuous ON time ({$T_{on}$}) of the device. This 1/\textit{f} noise reduction is also experimentally demonstrated in this paper using a circuit configuration with multiple identical transistor stages which produces a continuous output instead of a discrete signal. The circuit is implemented in 180~nm standard CMOS technology, from UMC. According to the measurement results, the proposed technique reduces the 1/\textit{f} noise by approximately 5.9 dB at $f_{s}$ of 1~KHz for 2 stage, which is extended up to 16 dB at $f_{s}$ of 5 MHz for 6 stage configuration.

cond-mat.other

iCorr : Complex correlation method to detect origin of replication in prokaryotic and eukaryotic genomes

Computational prediction of origin of replication (ORI) has been of great interest in bioinformatics and several methods including GC Skew, Z curve, auto-correlation etc. have been explored in the past. In this paper, we have extended the auto-correlation method to predict ORI location with much higher resolution for prokaryotes. The proposed complex correlation method (iCorr) converts the genome sequence into a sequence of complex numbers by mapping the nucleotides to {+1,-1,+i,-i} instead of {+1,-1} used in the auto-correlation method (here, 'i' is square root of -1). Thus, the iCorr method uses information about the positions of all the four nucleotides unlike the earlier auto-correlation method which uses the positional information of only one nucleotide. Also, this earlier method required visual inspection of the obtained graphs to identify the location of origin of replication. The proposed iCorr method does away with this need and is able to identify the origin location simply by picking the peak in the iCorr graph. The iCorr method also works for a much smaller segment size compared to the earlier auto-correlation method, which can be very helpful in experimental validation of the computational predictions. We have also developed a variant of the iCorr method to predict ORI location in eukaryotes and have tested it with the experimentally known origin locations of S. cerevisiae with an average accuracy of 71.76%.

q-bio.GN

Formal Ontology Learning on Factual IS-A Corpus in English using Description Logics

Ontology Learning (OL) is the computational task of generating a knowledge base in the form of an ontology given an unstructured corpus whose content is in natural language (NL). Several works can be found in this area most of which are limited to statistical and lexico-syntactic pattern matching based techniques Light-Weight OL. These techniques do not lead to very accurate learning mostly because of several linguistic nuances in NL. Formal OL is an alternative (less explored) methodology were deep linguistics analysis is made using theory and tools found in computational linguistics to generate formal axioms and definitions instead simply inducing a taxonomy. In this paper we propose "Description Logic (DL)" based formal OL framework for learning factual IS-A type sentences in English. We claim that semantic construction of IS-A sentences is non trivial. Hence, we also claim that such sentences requires special studies in the context of OL before any truly formal OL can be proposed. We introduce a learner tool, called DLOL_IS-A, that generated such ontologies in the owl format. We have adopted "Gold Standard" based OL evaluation on IS-A rich WCL v.1.1 dataset and our own Community representative IS-A dataset. We observed significant improvement of DLOL_IS-A when compared to the light-weight OL tool Text2Onto and formal OL tool FRED.

cs.CL

Perturbative solution of Vlasov equation for periodically driven systems

Statistical systems with time-periodic spatially non-uniform forces are of immense importance in several areas of physics. In this paper, we provide an analytical expression of the time-periodic probability distribution function of particles in such a system by perturbatively solving the 1D Vlasov equation in the limit of high frequency and slow spatial variation of the time-periodic force. We find that the time-averaged distribution function and density cannot be written simply in terms of an effective potential, also known as the fictitious ponderomotive potential. We also find that the temperature of such systems is spatially non-uniform leading to a non-equilibrium steady state which can further lead to a complex statistical time evolution of the system. Finally, we outline a method by which one can use these analytical solutions of the Vlasov equation to obtain numerical solutions of the self-consistent Vlasov-Poisson equations for such systems.

physics.plasm-ph

Spatial bandlimitedness of scattered electromagnetic fields

In this tutorial paper, we consider the problem of electromagnetic scattering by a bounded two-dimensional dielectric object, and discuss certain interesting properties of the scattered field. Using the electric field integral equation, along with the techniques of Fourier theory and the properties of Bessel functions, we show analytically and numerically, that in the case of transverse electric polarization, the scattered fields are spatially bandlimited. Further, we derive an upper bound on the number of incidence angles that are useful as constraints in an inverse problem setting (determining permittivity given measurements of the scattered field). We also show that the above results are independent of the dielectric properties of the scattering object and depend only on geometry. Though these results have previously been derived in the literature using the framework of functional analysis, our approach is conceptually far easier. Implications of these results on the inverse problem are also discussed.

physics.optics

Leaky Fermi accelerators

A Fermi accelerator is a billiard with oscillating walls. A leaky accelerator interacts with an environment of an ideal gas at equilibrium by exchange of particles through a small hole on its boundary. Such interaction may heat the gas: we estimate the net energy flow through the hole under the assumption that the particles inside the billiard do not collide with each other and remain in the accelerator for sufficiently long time. The heat production is found to depend strongly on the type of the Fermi accelerator. An ergodic accelerator, i.e. one which has a single ergodic component, produces a weaker energy flow than a multi-component accelerator. Specifically, in the ergodic case the energy gain is independent of the hole size, whereas in the multi-component case the energy flow may be significantly increased by shrinking the hole size.

nlin.CD