SearcharxivSearch

arXiv subjects

Soumendu Ghosh

Publications and source records attributed to Soumendu Ghosh.

8 recordsLinked to original sources

BIDENT: Heterogeneous Operator-level Mapping for Efficient Edge Inference

Modern edge System-on-Chips (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, yet current inference stacks map entire models to a single PU, leaving significant performance and energy efficiency on the table. This is exacerbated by emerging architectures such as state-space models (SSMs), Kolmogorov-Arnold networks (KANs), and multi-stage vision-language-action (VLA) pipelines, whose diverse operator characteristics are not uniformly suited to any single PU. We present BIDENT, a unified operator-level orchestration framework for heterogeneous edge inference that maps individual operators to the most suitable PU based on profiled execution characteristics. BIDENT formulates operator-to-PU assignment as a shortest-path problem over a weighted execution graph, enabling efficient and optimal scheduling under the cost model for both latency- and energy-minimization objectives. Unlike prior work relying on model-specific heuristics or coarse-grained partitioning, BIDENT is model-agnostic and jointly supports sequential execution, intra-model parallelism across independent operators, and multi-model concurrent scheduling in a single formulation. We implement BIDENT on an Intel Core Ultra SoC and evaluate it across 10 model families spanning CNNs, Transformers, SSMs, KANs, spiking networks, and multi-stage pipelines. BIDENT achieves up to 1.60x speedup via intra-model parallelism and a 3.42x geometric mean speedup across 190 multi-model combinations by utilizing otherwise idle compute. Sequential heterogeneous mapping yields more modest gains (up to 1.58x, 1.09x geometric mean), while energy-aware scheduling reduces energy consumption by 48.2% on average in concurrent settings. These results show that operator-level orchestration, not model-level mapping, is the key abstraction for fully exploiting heterogeneity in next-generation edge AI.

cs.AR

GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units

Graph Neural Networks (GNNs) are vital for learning from graph-structured data, enabling applications in network analysis, recommendation systems, and speech analytics. Deploying them on edge devices like client PCs and laptops enhances real-time processing, privacy, and cloud independence. GNNs aid Retrieval-Augmented Generation (RAG) for Large Language Models (LLMs) and enable event-based vision tasks. However, irregular memory access, sparsity, and dynamic structures cause high latency and energy overhead on resource-constrained devices. While modern edge processors integrate CPUs, GPUs, and NPUs, NPUs designed for data-parallel tasks struggle with irregular GNN computations. We introduce GraNNite, the first hardware-aware framework optimizing GNN execution on commercial-off-the-shelf (COTS) SOTA DNN accelerators via a structured three-step methodology: (1) enabling NPU execution, (2) optimizing performance, and (3) trading accuracy for efficiency gains. Step 1 employs GraphSplit for workload distribution and StaGr for static aggregation, while GrAd and NodePad handle dynamic graphs. Step 2 boosts performance using EffOp for control-heavy tasks and GraSp for sparsity exploitation. Graph Convolution optimizations PreG, SymG, and CacheG reduce redundancy and memory transfers. Step 3 balances quality versus efficiency, where QuantGr applies INT8 quantization, and GrAx1, GrAx2, and GrAx3 accelerate attention, broadcast-add, and SAGE-max aggregation. On Intel Core Ultra AI PCs, GraNNite achieves 2.6X to 7.6X speedups over default NPU mappings and up to 8.6X energy gains over CPUs and GPUs, delivering 10.8X and 6.7X higher performance than CPUs and GPUs, respectively, across GNN models.

cs.LG

A Biologically Motivated Asymmetric Exclusion Process: interplay of congestion in RNA polymerase traffic and slippage of nascent transcript

We develope a theoretical framework, based on exclusion process, that is motivated by a biological phenomenon called transcript slippage (TS). In this model a discrete lattice represents a DNA strand while each of the particles that hop on it unidirectionally, from site to site, represents a RNA polymerase (RNAP). While walking like a molecular motor along a DNA track in a step-by-step manner, a RNAP simultaneously synthesizes a RNA chain; in each forward step it elongates the nascent RNA molecule by one unit, using the DNA track also as the template. At some special "slippery" position on the DNA, which we represent as a defect on the lattice, a RNAP can lose its grip on the nascent RNA and the latter's consequent slippage results in a final product that is either longer or shorter than the corresponding DNA template. We develope an exclusion model for RNAP traffic where the kinetics of the system at the defect site captures key features of TS events. We demonstrate the interplay of the crowding of RNAPs and TS. A RNAP has to wait at the defect site for longer period in a more congested RNAP traffic, thereby increasing the likelihood of its suffering a larger number of TS events. The qualitative trends of some of our results for a simple special case of our model are consistent with experimental observations. The general theoretical framework presented here will be useful for guiding future experimental queries and for analysis of the experimental data with more detailed versions of the same model.

cond-mat.stat-mech

First-passage processes on a filamentous track in a dense traffic: optimizing diffusive search for a target in crowding conditions

Several important biological processes are initiated by the binding of a protein to a specific site on the DNA. The strategy adopted by a protein, called transcription factor (TF), for searching its specific binding site on the DNA has been investigated over several decades. In recent times the effects obstacles, like DNA-binding proteins, on the search by TF has begun to receive attention. RNA polymerase (RNAP) motors collectively move along a segment of the DNA during a genomic process called transcription. This RNAP traffic is bound to affect the diffusive scanning of the same segment of the DNA by a TF searching for its binding site. Motivated by this phenomenon, here we develop a kinetic model where a `particle', that represents a TF, searches for a specific site on a one-dimensional lattice. On the same lattice another species of particles, each representing a RNAP, hop from left to right exactly as in a totally asymmetric simple exclusion process (TASEP) which forbids simultaneous occupation of any site by more than one particle, irrespective of their identities. Although the TF is allowed to attach to or detach from any lattice site, the RNAPs can attach only to the first site at the left edge and detach from only the last site on the right edge of the lattice. We formulate the search as a {\it first-passage} process; the time taken to reach the target site {\it for the first time}, starting from a well defined initial state, is the search time. By approximate analytical calculations and Monte Carlo (MC) computer simulations, we calculate the mean search time. We show that RNAP traffic rectifies the diffusive motion of TF to that of a Brownian ratchet, and the mean time of successful search can be even shorter than that required in the absence of RNAP traffic. Moreover, we show that there is an optimal rate of detachment that corresponds to the shortest mean search time.

physics.bio-ph

A biologically inspired two-species exclusion model: effects of RNA polymerase motor traffic on simultaneous DNA replication

We introduce a two-species exclusion model to describe the key features of the conflict between the RNA polymerase (RNAP) motor traffic, engaged in the transcription of a segment of DNA, concomitant with the progress of two DNA replication forks on the same DNA segment. One of the species of particles ($P$) represents RNAP motors while the other ($R$) represents replication forks. Motivated by the biological phenomena that this model is intended to capture, a maximum of only two $R$ particles are allowed to enter the lattice from two opposite ends whereas the unrestricted number of $P$ particles constitute a totally asymmetric simple exclusion process (TASEP) in a segment in the middle of the lattice. Consequently, the lattice consists of three segments; the encounters of the $P$ particles with the $R$ particles are confined within the middle segment (segment $2$) whereas only the $R$ particles can occupy the sites in the segments $1$ and $3$. The model captures three distinct pathways for resolving the co-directional as well as head-collision between the $P$ and $R$ particles. Using Monte Carlo simulations and heuristic analytical arguments that combine exact results for the TASEP with mean-field approximations, we predict the possible outcomes of the conflict between the traffic of RNAP motors ($P$ particles engaged in transcription) and the replication forks ($R$ particles). The outcomes, of course, depend on the dynamical phase of the TASEP of $P$ particles. In principle, the model can be adapted to the experimental conditions to account for the data quantitatively.

q-bio.SC

A Multispecies Exclusion Model Inspired By Transcriptional Interference

We introduce exclusion models of two distinguishable species of hard rods with their distinct sites of entry and exit under open boundary conditions. In the first model both species of rods move in the same direction whereas in the other two models they move in the opposite direction. These models are motivated by the biological phenomenon known as Transcriptional Interference. Therefore, the rules for the kinetics of the models, particularly the rules for the outcome of the encounter of the rods, are also formulated to mimic those observed in Transcriptional Interference. By a combination of mean-field theory and computer simulation of these models we demonstrate how the flux of one species of rods is completely switched off by the other. Exploring the parameter space of the model we also establish the conditions under which switch-like regulation of two fluxes is possible; from the extensive analysis we discover more than one possible mechanism of this phenomenon.

physics.bio-ph

First Passage Time in Computation by Tape-Copying Turing Machines: Slippage of Nascent Tape

Transcription of the genetic message encoded chemically in the sequence of the DNA template is carried out by a molecular machine called RNA polymerase (RNAP). Backward or forward slippage of the nascent RNA with respect to the DNA template strand give rise to a transcript that is, respectively, longer or shorter than the corresponding template. We model a RNAP as a "Tape-copying Turing machine" (TCTM) where the DNA template is the input tape while the nascent RNA strand is the output tape. Although the TCTM always steps forward the process is assumed to be stochastic that has a probability of occurrence per unit time. The time taken by a TCTM for each single successful forward stepping on the input tape, during which the output tape suffers lengthening or shortening by $n$ units because of backward or forward slippage, is a random variable; we report some of the statistical characteristics of this time by using the formalism for calculation of the distributions of {\it first-passage time}. The results are likely to find applications in the analysis of experimental data on "programmed" transcriptional error caused by transcriptional slippage which is a mode of "recoding" of genetic information.

physics.bio-ph

Mixed molecular motor traffic on nucleic acid tracks: models of transcriptional interference and regulation of gene expression

RNA polymerase (RNAP) is molecular machine that polymerizes a RNA molecule, a linear heteropolymer, using a single stranded DNA (ssDNA) as the corresponding template; the sequence of monomers of the RNA is dictated by that of monomers on the ssDNA template. While polymerizing a RNA, the RNAP walks step-by-step on the ssDNA template in a specific direction. Thus, a RNAP can be regarded also as a molecular motor and the sites of start and stop of its walk on the DNA mark the two ends of the genetic message that it transcribes into RNA. Interference of transcription of two overlapping genes is believed to regulate the levels of their expression, i.e., the overall rate of the corresponding RNA synthesis, through suppressive effect of one on the other. Here we model this process as a mixed traffic of two groups of RNAP motors that are characterized by two distinct pairs of start and stop sites. Each group polymerizes identical copies of a RNA while the RNAs polymerized by the two groups are different. These models, which may also be viewed as two interfering totally asymmetric simple exclusion processes, account for all modes of transcriptional interference in spite of their extreme simplicity. A combination of mean-field theory and computer simulation of these models demonstrate the physical origin of the switch-like regulation of the two interfering genes in both co-directional and contra-directional traffic of the two groups of RNAP motors.

physics.bio-ph