SearcharxivSearch

arXiv subjects

Tom Chou

Publications and source records attributed to Tom Chou.

At least 19 recordsLinked to original sources

Effective population sizes for asymmetrically regulated birth-death processes

In multispecies birth-death processes, how population regulation -- through suppressed replication, elevated mortality, or both -- affects macroscopic stochastic dynamics has escaped detailed analysis. Here, we show that the distribution of regulation mechanisms can be invisible in deterministic or mean-field dynamics but play a significant role in the diffusive evolution of population frequencies. By introducing a tunable regulation partitioning parameter $\alpha_i$ and projecting a $d$-species birth-death process onto a $(d{-}1)$-dimensional Moran process, we find a regulation-mechanism-dependent diffusion tensor. For the simple two-species case, we derive exact fixation times and probabilities to show how different regulation mechanisms stochastically favors a more birth-regulated species, even under complete deterministic neutrality. Our model also allows us to define an $\alpha$-dependent effective population size $N_{\rm e}(\alpha)$ among neutral species, generalizing its classical interpretation. For near-neutral populations or populations that are heterogeneous in their regulation mechanism, we used perturbation theory to calculate the spectral gap, identifying it with a diversity loss timescale which can also be interpreted as setting an effective population size. Our results are particularly applicable to interacting subpopulations of T cells ("clones") which are near-neutral, are regulated through proliferation and apoptosis, and lose diversity with time.

q-bio.PE

CEI: A Clonal Expansion Identifier for T-cell receptor clones following SARS-CoV-2 vaccination

Each T cell typically carries a specific T-cell receptor (TCR) that determines its specificity against an epitope presented by the HLA complex on a target cell. Antigenic challenge triggers the expansion of reactive cells within a diverse pool of T cells with randomly generated receptors, a process that results in epitope-driven shifts of TCR frequencies over time. Here, we analyze the effects of SARS-CoV-2 vaccination on the TCR populations in peripheral blood drawn from seven COVID-naive individuals, before vaccines were widely available. To identify SARS-CoV-2 vaccine-associated TCR sequences among the $\sim 10^{5}-10^{6}$ TCR sequences sampled before and after vaccination, we develop statistical criteria to detect significant increases in abundance of positive TCR clones. Application of our statistical methods shows a robust identification of TCR sequences that respond to SARS-CoV-2 vaccination in vivo, illustrating the feasibility of quantifying the clone-specific dynamics of T-cell abundance changes following immunological perturbations.

q-bio.QM

A generalized work theorem for stopped stochastic chemical reaction networks

We establish a generalized work theorem for stochastic chemical reaction networks (CRNs). By using a compensated Poisson jump process, we identify a martingale structure in a generalized entropy defined relative to an auxiliary backward process and extend nonequilibrium work relations to processes stopped at bounded arbitrary times. Our results apply to discrete, mesoscopic chemical reaction networks and remain valid for singular initial conditions and state-dependent termination events. We show how martingale properties emerge directly from the structure of reaction propensities without assuming detailed balance. Stochastic simulations of a simple chemical kinetic proofreading network are used to explore the dependence of the exponentiated entropy production on initial conditions and model parameters, validating our new work theorem relationships. Our results provide new quantitative tools for analyzing biological circuits ranging from metabolic to gene regulation pathways.

cond-mat.stat-mech

Mechanical activity enables patterning and discrimination at the immune synapse

Immune cells recognize and discriminate antigens through immunological synapses - dynamic intercellular junctions exhibiting highly organized receptor-ligand patterns. While much work has focused on molecular kinetics and passive mechanisms of pattern formation, the role of active mechanical control in patterning and discrimination remains underexplored. We develop a minimal continuum model coupling receptor binding kinetics, membrane deformation, and cytoskeletal forces, with elastohydrodynamic flow in the synaptic cleft. Numerical simulations and scaling analysis reveal that contractile cortical flows arrest coarsening and stabilize long-lived multifocal clusters, whereas active pulling accelerates cluster dissolution and elevates background receptor binding. Nonequilibrium mechanical forces enable adaptive control over the speed, sensitivity, and dynamic range of affinity discrimination in a pattern-dependent manner. Our results highlight how immune cells exploit cytoskeletal remodeling to robustly regulate antigen recognition through synaptic patterning.

q-bio.CB

First passage times to T cell activation

Effective recognition of foreign antigens by the adaptive immune system relies on T cells being activated by antigen-presenting cells (APCs) in lymph nodes. Here, diffusing T cells may encounter cognate APCs that present matching antigen fragments or non-cognate ones that do not; they are also subject to degradation. We develop a stochastic model in which T cell-APCs interact via a sequence of recognition steps, represented as a multistage Markov chain. T cells are successfully activated only if the terminal state associated with a cognate APC is reached. We compute the probability of successful activation in the presence of interfering non-cognate APCs, T cell degradation, and lymph node exit, and analyze the mean first-passage time to activation. We also incorporate a kinetic proofreading mechanism that enables state resetting, and show how this enhances specificity toward cognate APCs.

q-bio.QM

Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model

How to allocate limited resources to projects that will yield the greatest long-term benefits is a problem that often arises in decision-making under uncertainty. For example, organizations may need to evaluate and select innovation projects with risky returns. Similarly, when allocating resources to research projects, funding agencies are tasked with identifying the most promising proposals based on idiosyncratic criteria. Finally, in participatory budgeting, a local community may need to select a subset of public projects to fund. Regardless of context, agents must estimate the uncertain values of a potentially large number of projects. Developing parsimonious methods to compare these projects, and aggregating agent evaluations so that the overall benefit is maximized, are critical in assembling the best project portfolio. Unlike in standard sorting algorithms, evaluating projects on the basis of uncertain long-term benefits introduces additional complexities. We propose comparison rules based on Quicksort and the Bradley--Terry model, which connects rankings to pairwise "win" probabilities. In our model, each agent determines win probabilities of a pair of projects based on his or her specific evaluation of the projects' long-term benefit. The win probabilities are then appropriately aggregated and used to rank projects. Several of the methods we propose perform better than the two most effective aggregation methods currently available. Additionally, our methods can be combined with sampling techniques to significantly reduce the number of pairwise comparisons. We also discuss how the Bradley--Terry portfolio selection approach can be implemented in practice.

q-fin.PM

Reconstructing Noisy Gene Regulation Dynamics Using Extrinsic-Noise-Driven Neural Stochastic Differential Equations

Proper regulation of cell signaling and gene expression is crucial for maintaining cellular function, development, and adaptation to environmental changes. Reaction dynamics in cell populations is often noisy because of (i) inherent stochasticity of intracellular biochemical reactions (``intrinsic noise'') and (ii) heterogeneity of cellular states across different cells that are influenced by external factors (``extrinsic noise''). In this work, we introduce an extrinsic-noise-driven neural stochastic differential equation (END-nSDE) framework that utilizes the Wasserstein distance to accurately reconstruct SDEs from trajectory data from a heterogeneous population of cells (extrinsic noise). We demonstrate the effectiveness of our approach using both simulated and experimental data from three different systems in cell biology: (i) circadian rhythms, (ii) RPA-DNA binding dynamics, and (iii) NF$\kappa$B signaling process. Our END-nSDE reconstruction method can model how cellular heterogeneity (extrinsic noise) modulates reaction dynamics in the presence of intrinsic noise. It also outperforms existing time-series analysis methods such as recurrent neural networks (RNNs) and long short-term memory networks (LSTMs). By inferring cellular heterogeneities from data, our END-nSDE reconstruction method can reproduce noisy dynamics observed in experiments. In summary, the reconstruction method we propose offers a useful surrogate modeling approach for complex biophysical processes, where high-fidelity mechanistic models may be impractical.

q-bio.QM

Incorporating stochastic gene expression, signaling-mediated intercellular interactions, and regulated cell proliferation in models of coordinated tissue development

Formulating quantitative and predictive models for tissue development requires consideration of the complex, stochastic gene expression dynamics, its regulation via cell-to-cell interactions, and cell proliferation. Including all of these processes into a practical mathematical framework requires complex expressions that are difficult to interpret and apply. We construct a simple theory that incorporates intracellular stochastic gene expression dynamics, signaling chemicals that influence these dynamics and mediate cell-cell interactions, and cell proliferation and its accompanying differentiation. Cellular states (genetic and epigenetic) are described by a Waddington vector field that allows for non-gradient dynamics (cycles, entropy production, loss of detailed balance) which is precluded in Waddington potential landscape representations of gene expression dynamics. We define an epigenetic fitness landscape that describes the proliferation of different cell types, and elucidate how this fitness landscape is related to Waddington's vector field. We illustrate the applicability of our framework by analyzing two model systems: an interacting two-gene differentiation process and a spatiotemporal organism model inspired by planaria.

q-bio.CB

Martingale properties of entropy production and a generalized work theorem with decoupled forward and backward processes

By decoupling forward and backward stochastic trajectories, we construct a family of martingales and work theorems for both overdamped and underdamped Langevin dynamics. Our results are made possible by an alternative derivation of work theorems that uses tools from stochastic calculus instead of path-integration. We further strengthen the equality in work theorems by evaluating expectations conditioned on an arbitrary initial state value. These generalizations extend the applicability of work theorems and offer new interpretations of entropy production in stochastic systems. Lastly, we discuss the violation of work theorems in far-from-equilibrium systems.

cond-mat.stat-mech

Aggregating multiple test results to improve medical decision-making

Gathering observational data for medical decision-making often involves uncertainties arising from both type I (false positive)and type II (false negative) errors. In this work, we develop a statistical model to study how medical decision-making can be improved by repeating diagnostic and screening tests, and aggregating their results. This approach is relevant not only in clinical settings, such as medical imaging, but also in public health, as highlighted by the need for rapid, cost-effective testing methods during the SARS-CoV-2pandemic. Our model enables the development of testing protocols with an arbitrary number of tests, which can be customized to meet requirements for type I and type II errors. This allows us to adjust sensitivity and specificity according to application-specific needs. Additionally, we derive generalized Rogan--Gladen estimates for estimating disease prevalence, accounting for an arbitrary number of tests with potentially different type I and type II errors. We also provide the corresponding uncertainty quantification.

stat.AP

A knapsack for collective decision-making

Collective decision-making is the process through which diverse stakeholders reach a joint decision. Within societal settings, one example is participatory budgeting, where constituents decide on the funding of public projects. How to most efficiently aggregate diverse stakeholder inputs on a portfolio of projects with uncertain long-term benefits remains an open question. We address this problem by studying collective decision-making through the integration of preference aggregation and knapsack allocation methods. Since different stakeholder groups may evaluate projects differently,we examine several aggregation methods that combine their diverse inputs. The aggregated evaluations are then used to fill a ``collective'' knapsack. Among the methods we consider are the arithmetic mean, Borda-type rankings, and delegation to experts. We find that the factors improving an aggregation method's ability to identify projects with the greatest expected long-term value include having many stakeholder groups, moderate variation in their expertise levels, and some degree of delegation or bias favoring groups better positioned to objectively assess the projects. We also discuss how evaluation errors and heterogeneous costs impact project selection. Our proposed aggregation methods are relevant not only in the context of funding public projects but also, more generally, for organizational decision-making under uncertainty.

econ.TH

An efficient Wasserstein-distance approach for reconstructing jump-diffusion processes using parameterized neural networks

We analyze the Wasserstein distance ($W$-distance) between two probability distributions associated with two multidimensional jump-diffusion processes. Specifically, we analyze a temporally decoupled squared $W_2$-distance, which provides both upper and lower bounds associated with the discrepancies in the drift, diffusion, and jump amplitude functions between the two jump-diffusion processes. Then, we propose a temporally decoupled squared $W_2$-distance method for efficiently reconstructing unknown jump-diffusion processes from data using parameterized neural networks. We further show its performance can be enhanced by utilizing prior information on the drift function of the jump-diffusion process. The effectiveness of our proposed reconstruction method is demonstrated across several examples and applications.

stat.ML

A probabilistic model of relapse in drug addiction

More than 60% of individuals recovering from substance use disorder relapse within one year. Some will resume drug consumption even after decades of abstinence. The cognitive and psychological mechanisms that lead to relapse are not completely understood, but stressful life experiences and external stimuli that are associated with past drug-taking are known to play a primary role. Stressors and cues elicit memories of drug-induced euphoria and the expectation of relief from current anxiety, igniting an intense craving to use again; positive experiences and supportive environments may mitigate relapse. We present a mathematical model of relapse in drug addiction that draws on known psychiatric concepts such as the "positive activation; negative activation" paradigm and the "peak-end" rule to construct a relapse rate that depends on external factors (intensity and timing of life events) and individual traits (mental responses to these events). We analyze which combinations and ordering of stressors, cues, and positive events lead to the largest relapse probability and propose interventions to minimize the likelihood of relapse. We find that the best protective factor is exposure to a mild, yet continuous, source of contentment, rather than large, episodic jolts of happiness.

q-bio.QM

Kinetic theories of state- and generation-dependent cell populations

We formulate a general, high-dimensional kinetic theory describing the internal state (such as gene expression or protein levels) of cells in a stochastically evolving population. The resolution of our kinetic theory also allows one to track subpopulations associated with each generation. Both intrinsic noise of the cell's internal attribute and randomness in a cell's division times (demographic stochasticity) are fundamental to the development of our model. Based on this general framework, we are able to marginalize the high-dimensional kinetic PDEs in a number of different ways to derive equations that describe the dynamics of marginalized or "macroscopic" quantities such as structured population densities, moments of generation-dependent cellular states, and moments of the total population. We also show how nonlinear "interaction" terms in lower-dimensional integrodifferential equations can arise from high-dimensional linear kinetic models that contain rate parameters of a cell (birth and death rates) that depend on variables associated with other cells, generating couplings in the dynamics. Our analysis provides a general, more complete mathematical framework that resolves the coevolution of cell populations and cell states. The approach may be tailored for studying, e.g., gene expression in developing tissues, or other more general particle systems which exhibit Brownian noise in individual attributes and population-level demographic noise.

q-bio.PE

Reliable ligand discrimination in stochastic multistep kinetic proofreading: First passage time vs. product counting strategies

Cellular signaling, crucial for biological processes like immune response and homeostasis, relies on specificity and fidelity in signal transduction to accurately respond to stimuli amidst biological noise. Kinetic proofreading (KPR) is a key mechanism enhancing signaling specificity through time-delayed steps, although its effectiveness is debated due to intrinsic noise potentially reducing signal fidelity. In this study, we reformulate the theory of kinetic proofreading (KPR) by convolving multiple intermediate states into a single state and then define an overall "processing" time required to traverse these states. This simplification allows us to succinctly describe kinetic proofreading in terms of a single waiting time parameter, facilitating a more direct evaluation and comparison of KPR performance across different biological contexts such as DNA replication and T cell receptor (TCR) signaling. We find that loss of fidelity for longer proofreading steps relies on the specific strategy of information extraction and show that in the first-passage time (FPT) discrimination strategy, longer proofreading steps can exponentially improve the accuracy of KPR at the cost of speed. Thus, KPR can still be an effective discrimination mechanism in the high noise regime. However, in a product concentration-based discrimination strategy, longer proofreading steps do not necessarily lead to an increase in performance. However, by introducing activation thresholds on product concentrations, can we decompose the product-based strategy into a series of FPT based strategies to better resolve the subtleties of KPR-mediated product discrimination. Our findings underscore the importance of understanding KPR in the context of how information is extracted and processed in the cell.

q-bio.MN

Squared Wasserstein-2 Distance for Efficient Reconstruction of Stochastic Differential Equations

We provide an analysis of the squared Wasserstein-2 ($W_2$) distance between two probability distributions associated with two stochastic differential equations (SDEs). Based on this analysis, we propose the use of a squared $W_2$ distance-based loss functions in the \textit{reconstruction} of SDEs from noisy data. To demonstrate the practicality of our Wasserstein distance-based loss functions, we performed numerical experiments that demonstrate the efficiency of our method in reconstructing SDEs that arise across a number of applications.

math.PR

A Spectral Approach for Learning Spatiotemporal Neural Differential Equations

Rapidly developing machine learning methods has stimulated research interest in computationally reconstructing differential equations (DEs) from observational data which may provide additional insight into underlying causative mechanisms. In this paper, we propose a novel neural-ODE based method that uses spectral expansions in space to learn spatiotemporal DEs. The major advantage of our spectral neural DE learning approach is that it does not rely on spatial discretization, thus allowing the target spatiotemporal equations to contain long range, nonlocal spatial interactions that act on unbounded spatial domains. Our spectral approach is shown to be as accurate as some of the latest machine learning approaches for learning PDEs operating on bounded domains. By developing a spectral framework for learning both PDEs and integro-differential equations, we extend machine learning methods to apply to unbounded DEs and a larger class of problems.

cs.LG

Forecasting drug overdose mortality by age in the United States at the national and county levels

The drug overdose crisis in the United States continues to intensify. Fatalities have increased five-fold since 1999 reaching a record high of 108,000 deaths in 2021. The epidemic has unfolded through distinct waves of different drug types, uniquely impacting various age, gender, race and ethnic groups in specific geographical areas. One major challenge in designing effective interventions is the forecasting of age-specific overdose patterns at the local level so that prevention and preparedness can be effectively delivered. We develop a forecasting method that assimilates observational data obtained from the CDC WONDER database with an age-structured model of addiction and overdose mortality. We apply our method nationwide and to three select areas: Los Angeles County, Cook County and the five boroughs of New York City, providing forecasts of drug-overdose mortality and estimates of relevant epidemiological quantities, such as mortality and age-specific addiction rates.

stat.AP