SearcharxivSearch

arXiv subjects

Siddhartha Roy

Publications and source records attributed to Siddhartha Roy.

4 recordsLinked to original sources

Liquid-Liquid Phase Separation in a Minimal Explicit-Solvent Lattice Model Mimicking Protein Solutions

Biomolecular condensates play essential roles in cellular processes, and recent efforts have focused on understanding their assembly and rational design principles. In this study, we have employed an explicit-solvent minimal statistical mechanical model based on the lattice-gas Hamiltonian with quenched disorder -- which mimics crowders -- to investigate how protein-solvent and protein-crowder interactions influence condensate phase behavior and morphology. The computed phase diagrams reveal rich behavior, including upper critical solution temperature (UCST), closed-loop, and reentrant type transitions under varying protein-solvent interactions at both equilibrium and out-of-equilibrium conditions. We elucidated the origin of these phase behavior changes and examined the role of protein-crowder interactions in modulating condensed phase morphology and stability. We further extended this model to binary protein mixtures where we studied the phase behavior in the presence and absence of quenched disorder. Without disorder, the system exhibits diverse phase-separated morphologies -- partially wetted, fully wetted, segregative, and associative -- with phase boundaries delicately sensitive protein-solvent interactions. The introduction of quenched disorder (or crowder) leads to a broader spectrum of complex morphologies, dictated by the interplay among protein-protein, protein-solvent, and protein-crowder interaction parameters. In general, this work underscores that protein-solvent and protein-crowder interactions, together with protein-protein interactions, can act as key regulatory parameters for modulating condensate morphology. These insights may guide future computational and experimental studies of liquid-liquid phase separation in biomolecular systems aimed at designing stimuli-responsive condensates.

cond-mat.soft

Hallucination Detection via Activations of Open-Weight Proxy Analyzers

We introduce a proxy-analyzer framework for detecting hallucinations in large language models. Instead of looking inside the generating model, our system reads already-generated text through a small locally hosted open-weight model and spots hallucinations using the reader's own internal activations. This works just as well when the generator is a closed API like GPT-4 as when it is any open-weight model. We built eighteen features grounded in how transformers process text, covering residual stream norms, per-head source-document attention, entropy, MLP activations, logit-lens trajectories, and three new token-level grounding statistics. We trained a stacking ensemble on 72,135 samples from five hallucination datasets. We tested across seven analyzer architectures from 0.5 billion to 9 billion parameters: Qwen2.5 at 0.5B and 7B, Gemma-2 at 2B and 9B, Pythia at 1.4B, and LLaMA-3 at both 3B and 8B. Across all seven, we consistently beat ReDeEP's token-level AUC of 0.73 on RAGTruth by 7.4 to 10.3 percentage points. Qwen2.5-7B reached an F1 of 0.717, just above ReDeEP's 0.713, while Qwen2.5-0.5B hit 0.706. The most striking finding is how tightly all seven models cluster: AUC spans only 2.3 percentage points across an eighteen-fold difference in model size. Even more surprising, our 3B LLaMA outperforms our 8B LLaMA on RAGTruth, showing that bigger is not always better even within the same model family. Both RAGTruth and LLM-AggreFact include outputs from multiple LLM families, so our results are not skewed toward any particular generator.

cs.CL

Stochastic Simulation of Gene Expression in a Single Cell

In this paper, we consider two stochastic models of gene expression in prokaryotic cells. In the first model, sixteen biochemical reactions involved in transcription, translation and transcriptional regulation in the presence of inducer molecules are considered. The time evolution of the number of biomolecules of a particular type is determined using the stochastic simulation method based on the Gillespie Algorithm. The results obtained show that if the number of inducer molecules, N(I), is greater than or equal to the number of regulatory molecules, N(R), the average protein level is high in the steady state (state 2). The magnitude of the level is the same as long as N{I) greater than or equal to N(R). When N(I) is very very less than N(R), the average protein level is low, practically zero (state 1). As N(I) increases, the protein level continues to remain low. When N(I)becomes close to N(R), protein levels in the steady state are intermediate between high and low.In the presence of autocatalysis, a cell mostly exists in either state 1 or state 2 giving rise to a bimodal distribution in the protein levels in an ensemble of cells. This corresponds to the "all or none'' phenomenon observed in experiments. In the second model, the inducer molecules are not considered explicitly. An exhaustive simulation over the parameter space of the model shows that there are three major patterns of gene expression, Type A, Type B and Type C. The effect of varying the cellular parameters on the patterns, in particular, the transition from one type of pattern to another, is studied. Type A and Type B patterns have been observed in experiments. Simple mathematical models of transcriptional regulation predict Type C pattern of gene expression in certain parameter regimes.

cond-mat

A Cooperative Stochastic Model of Gene Expression

Recent experiments at the level of a single cell have shown that gene expression occurs in abrupt stochastic bursts. Further, in an ensemble of cells, the levels of proteins produced have a bimodal distribution. In a large fraction of cells, the gene expression is either off or has a high value. We propose a stochastic model of gene expression the essential features of which are stochasticity and cooperative binding of RNA polymerase. The model can reproduce the bimodal behaviour seen in experiments.

cond-mat.soft