SearcharxivSearch

arXiv subjects

Vahid Shahrezaei

Publications and source records attributed to Vahid Shahrezaei.

6 recordsLinked to original sources

Ordered Diffusion Kernels

We introduce Ordered Diffusion Kernels (ODKs), a novel class of local kernels that can approximate the infinitesimal generator of an arbitrary Itô Stochastic Differential Equation (SDE). ODKs are designed to be applied to data sampled from dynamical systems where little dynamical information is available a priori. The Laplacian of classical diffusion kernels approximates the Laplace-Beltrami operator on the underlying manifold; adjusting the normalisation introduces an advection term that depends on the sampling density; recently, TMDmap generalised this normalisation to target an arbitrary measure, but at the cost of coupling advection to diffusion. More general local kernels can learn arbitrary second-order elliptic operators but are formulated in terms of known velocity fields --- making the first step of any analysis a potentially ill-posed inference problem. To formulate ODK, we first relax the problem of potential estimation to the more tractable task of inferring an ordering of the data, which we represent through an ordering function. We prove ODK's Laplacian converges to the infinitesimal generator of a gradient-flow SDE with state-dependent isotropic diffusion, without coupling advection and diffusion. We provide various extensions of ODK to: arbitrary drifts via local ordering functions; anisotropic diffusions via a Strang splitting scheme; multiple ordering functions; and self-tuning bandwidths. In addition, we introduce two loss functions which exploit the structure of ODKs to solve a non-parametric inference problem. We validate this framework on synthetic data from deterministic and stochastic systems, demonstrating accurate recovery of operators, velocity fields, extrinsic curvature, and spatially dependent drift and diffusion coefficients.

math.NA

Data-driven spatio-temporal modelling of glioblastoma

Mathematical oncology provides unique and invaluable insights into tumour growth on both the microscopic and macroscopic levels. This review presents state-of-the-art modelling techniques and focuses on their role in understanding glioblastoma, a malignant form of brain cancer. For each approach, we summarise the scope, drawbacks, and assets. We highlight the potential clinical applications of each modelling technique and discuss the connections between the mathematical models and the molecular and imaging data used to inform them. By doing so, we aim to prime cancer researchers with current and emerging computational tools for understanding tumour progression. Finally, by providing an in-depth picture of the different modelling techniques, we also aim to assist researchers who seek to build and develop their own models and the associated inference frameworks.

q-bio.QM

Multi-view Data Visualisation via Manifold Learning

Non-linear dimensionality reduction can be performed by \textit{manifold learning} approaches, such as Stochastic Neighbour Embedding (SNE), Locally Linear Embedding (LLE) and Isometric Feature Mapping (ISOMAP). These methods aim to produce two or three latent embeddings, primarily to visualise the data in intelligible representations. This manuscript proposes extensions of Student's t-distributed SNE (t-SNE), LLE and ISOMAP, for dimensionality reduction and visualisation of multi-view data. Multi-view data refers to multiple types of data generated from the same samples. The proposed multi-view approaches provide more comprehensible projections of the samples compared to the ones obtained by visualising each data-view separately. Commonly visualisation is used for identifying underlying patterns within the samples. By incorporating the obtained low-dimensional embeddings from the multi-view manifold approaches into the K-means clustering algorithm, it is shown that clusters of the samples are accurately identified. Through the analysis of real and synthetic data the proposed multi-SNE approach is found to have the best performance. We further illustrate the applicability of the multi-SNE approach for the analysis of multi-omics single-cell data, where the aim is to visualise and identify cell heterogeneity and cell types in biological tissues relevant to health and disease.

stat.ML

S-multi-SNE: Semi-Supervised Classification and Visualisation of Multi-View Data

An increasing number of multi-view data are being published by studies in several fields. This type of data corresponds to multiple data-views, each representing a different aspect of the same set of samples. We have recently proposed multi-SNE, an extension of t-SNE, that produces a single visualisation of multi-view data. The multi-SNE approach provides low-dimensional embeddings of the samples, produced by being updated iteratively through the different data-views. Here, we further extend multi-SNE to a semi-supervised approach, that classifies unlabelled samples by regarding the labelling information as an extra data-view. We look deeper into the performance, limitations and strengths of multi-SNE and its extension, S-multi-SNE, by applying the two methods on various multi-view datasets with different challenges. We show that by including the labelling information, the projection of the samples improves drastically and it is accompanied by a strong classification performance.

cs.LG

Analytical distributions for stochastic gene expression

Gene expression is significantly stochastic making modeling of genetic networks challenging. We present an approximation that allows the calculation of not only the mean and variance but also the distribution of protein numbers. We assume that proteins decay substantially slower than their mRNA and confirm that many genes satisfy this relation using high-throughput data from budding yeast. For a two-stage model of gene expression, with transcription and translation as first-order reactions, we calculate the protein distribution for all times greater than several mRNA lifetimes and thus qualitatively predict the distribution of times for protein levels to first cross an arbitrary threshold. If in addition the promoter fluctuates between inactive and active states, we can find the steady-state protein distribution, which can be bimodal if promoter fluctuations are slow. We show that our assumptions imply that protein synthesis occurs in geometrically distributed bursts and allows mRNA to be eliminated from a master equation description. In general, we find that protein distributions are asymmetric and may be poorly characterized by their mean and variance. Through maximum likelihood methods, our expressions should therefore allow more quantitative comparisons with experimental data. More generally, we introduce a technique to derive a simpler, effective dynamics for a stochastic system by eliminating a fast variable.

q-bio.MN

Colored extrinsic fluctuations and stochastic gene expression

Stochasticity is both exploited and controlled by cells. Although the intrinsic stochasticity inherent in biochemistry is relatively well understood, cellular variation, or 'noise', is predominantly generated by interactions of the system of interest with other stochastic systems in the cell or its environment. Such extrinsic fluctuations are nonspecific, affecting many system components, and have a substantial lifetime, comparable to the cell cycle (they are 'colored'). Here, we extend the standard stochastic simulation algorithm to include extrinsic fluctuations. We show that these fluctuations affect mean protein numbers and intrinsic noise, can speed up typical network response times, and can explain trends in high-throughput measurements of variation. If extrinsic fluctuations in two components of the network are correlated, they may combine constructively (amplifying each other) or destructively (attenuating each other). Consequently, we predict that incoherent feedforward loops attenuate stochasticity, while coherent feedforwards amplify it. Our results demonstrate that both the timescales of extrinsic fluctuations and their nonspecificity substantially affect the function and performance of biochemical networks.

q-bio.MN