SearcharxivSearch

arXiv subjects

Rahul Kulkarni

Publications and source records attributed to Rahul Kulkarni.

8 recordsLinked to original sources

Approximate Analytical Protein Distributions for the Three-stage Model of Stochastic Gene Expression

Gene expression is an intrinsically stochastic process that generates phenotypic heterogeneity within genetically identical cell populations. While the exact statistical moments of the protein count can be obtained for a broad range of complex models, the corresponding distributions are significantly harder to obtain and intractable in many cases. The classical three-stage model of gene expression, which predicts fluctuations in protein levels as a function of promoter switching, transcription, translation, and degradation events all occurring with linear propensities, illustrates this perfectly; deriving its exact protein distribution remains elusive. Here, using the partitioning property of time-inhomogeneous Poisson processes, we develop an exact mapping of the three-stage model onto a simplified model. The simplified model allows us to formulate two analytical approximations for the full protein distribution of the three-stage model based on a beta-mixture representation of the exact solution for a simpler model. We show that the two approximations are asymptotically exact in different limiting cases and verify their accuracy against simulations for a broad range of parameters. Although approximate, these are the first analytical expressions for protein distributions for the three-stage model that are highly accurate in intermediate regimes.

q-bio.QM

Analyzing Post-transcriptional Regulation in Stochastic Gene Expression Models Using Partitioned Poisson Arrivals

Gene expression is a stochastic process that allows for fluctuations in protein levels that can give rise to phenotypic heterogeneity within a population of genetically identical cells. Thus, there is great interest in quantifying how natural variation (noise) in gene expression is impacted by cellular control mechanisms, such as the various mechanisms pertaining to post-transcriptional regulation. Although previous research has developed a general analytical framework to compute the exact moments of mRNA distributions for any promoter-based regulatory motif, and the exact mRNA distribution itself in some cases, a similar framework for protein fluctuations is currently lacking. Here, we invoke the partitioning property of Poisson arrivals to map a general class of stochastic models of post-transcriptional regulation onto models that resemble promoter-based regulation. This approach leads to exact analytical results for the moments of protein distributions, and in certain cases the full distribution itself, using known exact results for mRNA distributions undergoing arbitrary promoter-based regulation. We further extend the framework to incorporate transcriptional bursting, leading to a versatile, unifying analytical framework for analyzing post-transcriptional regulation in stochastic gene expression.

q-bio.QM

Astra: AI Safety, Trust, & Risk Assessment

This paper argues that existing global AI safety frameworks exhibit contextual blindness towards India's unique socio-technical landscape. With a population of 1.5 billion and a massive informal economy, India's AI integration faces specific challenges such as caste-based discrimination, linguistic exclusion of vernacular speakers, and infrastructure failures in low-connectivity rural zones, that are frequently overlooked by Western, market-centric narratives. We introduce ASTRA, an empirically grounded AI Safety Risk Database designed to categorize risks through a bottom-up, inductive process. Unlike general taxonomies, ASTRA defines AI Safety Risks specifically as hazards stemming from design flaws such as skewed training sets or lack of guardrails that can be mitigated through technical iteration or architectural changes. This framework employs a tripartite causal taxonomy to evaluate risks based on their implementation timing (development, deployment, or usage), the responsible entity (the system or the user), and the nature of the intent (unintentional vs. intentional). Central to the research is a domain-agnostic ontology that organizes 37 leaf-level risk classes into two primary meta-categories: Social Risks and Frontier/Socio-Structural Risks. By focusing initial efforts on the Education and Financial Lending sectors, the paper establishes a scalable foundation for a "living" regulatory utility intended to evolve alongside India's expanding AI ecosystem.

cs.CY

Explainable AI in Big Data Fraud Detection

Big Data has become central to modern applications in finance, insurance, and cybersecurity, enabling machine learning systems to perform large-scale risk assessments and fraud detection. However, the increasing dependence on automated analytics introduces important concerns about transparency, regulatory compliance, and trust. This paper examines how explainable artificial intelligence (XAI) can be integrated into Big Data analytics pipelines for fraud detection and risk management. We review key Big Data characteristics and survey major analytical tools, including distributed storage systems, streaming platforms, and advanced fraud detection models such as anomaly detectors, graph-based approaches, and ensemble classifiers. We also present a structured review of widely used XAI methods, including LIME, SHAP, counterfactual explanations, and attention mechanisms, and analyze their strengths and limitations when deployed at scale. Based on these findings, we identify key research gaps related to scalability, real-time processing, and explainability for graph and temporal models. To address these challenges, we outline a conceptual framework that integrates scalable Big Data infrastructure with context-aware explanation mechanisms and human feedback. The paper concludes with open research directions in scalable XAI, privacy-aware explanations, and standardized evaluation methods for explainable fraud detection systems.

cs.LG

LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios

LLM bias evaluation is critical as large language models (LLMs) increasingly influence high-stakes decisions. This paper provides a comprehensive assessment of gender, racial, and age disparities in leading LLMs, revealing that debiasing efforts often create new fairness trade-offs. Recent advancements in LLMs have been notable, yet widespread enterprise adoption remains limited due to various constraints. This paper examines bias in LLMs - a crucial issue affecting their usability, reliability, and fairness. Our study evaluates gender bias in occupational scenarios and gender, age, and racial bias in crime scenarios across four leading LLMs released in 2024: Gemini 1.5 Pro, Llama 3 70B, Claude 3 Opus, and GPT-4o. Findings reveal that LLMs often depict female characters more frequently than male ones in various occupations, showing a 37% deviation from US BLS data. In crime scenarios, deviations from US FBI data are 54% for gender, 28% for race, and 17% for age. Critically, we observe that efforts to reduce gender and racial bias often lead to outcomes that may over-index one sub-class, potentially exacerbating disparities - a "debiasing paradox" that highlights the limitations of current bias mitigation techniques and underscores the need for more effective approaches.

cs.AI

Leveraging Prior Knowledge in Reinforcement Learning via Double-Sided Bounds on the Value Function

An agent's ability to leverage past experience is critical for efficiently solving new tasks. Approximate solutions for new tasks can be obtained from previously derived value functions, as demonstrated by research on transfer learning, curriculum learning, and compositionality. However, prior work has primarily focused on using value functions to obtain zero-shot approximations for solutions to a new task. In this work, we show how an arbitrary approximation for the value function can be used to derive double-sided bounds on the optimal value function of interest. We further extend the framework with error analysis for continuous state and action spaces. The derived results lead to new approaches for clipping during training which we validate numerically in simple domains.

cs.LG

Moment Closure Approximations in a Genetic Negative Feedback Circuit

Auto-regulation, a process wherein a protein negatively regulates its own production, is a common motif in gene expression networks. Negative feedback in gene expression plays a critical role in buffering intracellular fluctuations in protein concentrations around optimal value. Due to the nonlinearities present in these feedbacks, moment dynamics are typically not closed, in the sense that the time derivative of the lower-order statistical moments of the protein copy number depends on high-order moments. Moment equations are closed by expressing higher-order moments as nonlinear functions of lower-order moments, a technique commonly referred to as moment closure. Here, we compare the performance of different moment closure techniques. Our results show that the commonly used closure method, which assumes a priori that the protein population counts are normally distributed, performs poorly. In contrast, conditional derivative matching, a novel closure scheme proposed here provides a good approximation to the exact moments across different parameter regimes. In summary our study provides a new moment closure method for studying stochastic dynamics of genetic negative feedback circuits, and can be extended to probe noise in more complex gene networks.

q-bio.SC

Quantifying mRNA synthesis and decay rates using small RNAs

Regulation of mRNA decay is a critical component of global cellular adaptation to changing environments. The corresponding changes in mRNA lifetimes can be coordinated with changes in mRNA transcription rates to fine-tune gene expression. Current approaches for measuring mRNA lifetimes can give rise to secondary effects due to transcription inhibition and require separate experiments to estimate changes in mRNA transcription rates. Here, we propose an approach for simultaneous determination of changes in mRNA transcription rate and lifetime using regulatory small RNAs to control mRNA decay. We analyze a stochastic model for coupled degradation of mRNAs and sRNAs and derive exact results connecting RNA lifetimes and transcription rates to mean abundances. The results obtained show how steady-state measurements of RNA levels can be used to analyze factors and processes regulating changes in mRNA transcription and decay.

q-bio.CB