SearcharxivSearch

arXiv subjects

Robert Azencott

Publications and source records attributed to Robert Azencott.

At least 19 recordsLinked to original sources

Fast $k$-means clustering in Riemannian manifolds via Fr\'{e}chet maps: Applications to large-dimensional SPD matrices

We introduce a novel, efficient framework for clustering data on high-dimensional, non-Euclidean manifolds that overcomes the computational challenges associated with standard intrinsic methods. The key innovation is the use of the $p$-Fr\'{e}chet map $F^p : \mathcal{M} \to \mathbb{R}^\ell$ -- defined on a generic metric space $\mathcal{M}$ -- which embeds the manifold data into a lower-dimensional Euclidean space $\mathbb{R}^\ell$ using a set of reference points $\{r_i\}_{i=1}^\ell$, $r_i \in \mathcal{M}$. Once embedded, we can efficiently and accurately apply standard Euclidean clustering techniques such as k-means. We rigorously analyze the mathematical properties of $F^p$ in the Euclidean space and the challenging manifold of $n \times n$ symmetric positive definite matrices $\mathit{SPD}(n)$. Extensive numerical experiments using synthetic and real $\mathit{SPD}(n)$ data demonstrate significant performance gains: our method reduces runtime by up to two orders of magnitude compared to intrinsic manifold-based approaches, all while maintaining high clustering accuracy, including scenarios where existing alternative methods struggle or fail.

cs.LG

Can Generalized Extreme Value Model Fit the Real Stocks

The Generalized Extreme Value (GEV) distribution plays a critical role in risk assessment across various domains, such as hydrology, climate science, and finance. In this study, we investigate its application in analyzing intraday trading risks within the Chinese stock market, focusing on abrupt price movements influenced by unique trading regulations. To address limitations of traditional GEV parameter estimators, we leverage recently developed robust and asymptotically normal estimators, enabling accurate modeling of extreme intraday price fluctuations. We introduce two risk indicators: the mean risk level (mEVI) and a Stability Indicator (STI) to evaluate the stability of the shape parameter over time. Using data from 261 Chinese and 32 U.S. stocks (2015-2017), we find that Chinese stocks exhibit higher mEVI, corresponding to greater tail risk, while maintaining high model stability. Additionally, we show that Value at Risk (VaR) estimates derived from our GEV models outperform traditional GP and normal-based VaR methods in terms of variance and portfolio optimization. These findings underscore the versatility and efficiency of GEV modeling for intraday risk management and portfolio strategies.

stat.AP

Multi-Quantile Estimators for the parameters of Generalized Extreme Value distribution

We introduce and study Multi-Quantile estimators for the parameters $( \xi, \sigma, \mu)$ of Generalized Extreme Value (GEV) distributions to provide a robust approach to extreme value modeling. Unlike classical estimators, such as the Maximum Likelihood Estimation (MLE) estimator and the Probability Weighted Moments (PWM) estimator, which impose strict constraints on the shape parameter $\xi$, our estimators are always asymptotically normal and consistent across all values of the GEV parameters. The asymptotic variances of our estimators decrease with the number of quantiles increasing and can approach the Cram\'er-Rao lower bound very closely whenever it exists. Our Multi-Quantile Estimators thus offer a more flexible and efficient alternative for practical applications. We also discuss how they can be implemented in the context of Block Maxima method.

stat.ME

Rare Events Analysis and Computation for Stochastic Evolution of Bacterial Populations

In this paper, we develop a computational approach for computing most likely trajectories describing rare events that correspond to the emergence of non-dominant genotypes. This work is based on the large deviations approach for discrete Markov chains describing the genetic evolution of large bacterial populations. We demonstrate that a gradient descent algorithm developed in this paper results in the fast and accurate computation of most-likely trajectories for a large number of bacterial genotypes. We supplement our analysis with extensive numerical simulations demonstrating the computational advantage of the designed gradient descent algorithm over other, more simplified, approaches.

q-bio.PE

An Operator-Splitting Approach for Variational Optimal Control Formulations for Diffeomorphic Shape Matching

We present formulations and numerical algorithms for solving diffeomorphic shape matching problems. We formulate shape matching as a variational problem governed by a dynamical system that models the flow of diffeomorphism $f_t \in \operatorname{diff}(\mathbb{R}^3)$. We overview our contributions in this area, and present an improved, matrix-free implementation of an operator splitting strategy for diffeomorphic shape matching. We showcase results for diffeomorphic shape matching of real clinical cardiac data in $\mathbb{R}^3$ to assess the performance of our methodology.

math.OC

Automatic classification of deformable shapes

Let $\mathcal{D}$ be a dataset of smooth 3D-surfaces, partitioned into disjoint classes $\mathit{CL}_j$, $j= 1, \ldots, k$. We show how optimized diffeomorphic registration applied to large numbers of pairs $S,S' \in \mathcal{D}$ can provide descriptive feature vectors to implement automatic classification on $\mathcal{D}$, and generate classifiers invariant by rigid motions in $\mathbb{R}^3$. To enhance accuracy of automatic classification, we enrich the smallest classes $\mathit{CL}_j$ by diffeomorphic interpolation of smooth surfaces between pairs $S,S' \in \mathit{CL}_j$. We also implement small random perturbations of surfaces $S\in \mathit{CL}_j$ by random flows of smooth diffeomorphisms $F_t:\mathbb{R}^3 \to \mathbb{R}^3$. Finally, we test our automatic classification methods on a cardiology data base of discretized mitral valve surfaces.

cs.CV

Stochastic Neural Networks for Automatic Cell Tracking in Microscopy Image Sequences of Bacterial Colonies

Our work targets automated analysis to quantify the growth dynamics of a population of bacilliform bacteria. We propose an innovative approach to frame-sequence tracking of deformable-cell motion by the automated minimization of a new, specific cost functional. This minimization is implemented by dedicated Boltzmann machines (stochastic recurrent neural networks). Automated detection of cell divisions is handled similarly by successive minimizations of two cost functions, alternating the identification of children pairs and parent identification. We validate the proposed automatic cell tracking algorithm using (i) recordings of simulated cell colonies that closely mimic the growth dynamics of E. coli in microfluidic traps and (ii) real data. On a batch of 1100 simulated image frames, cell registration accuracies per frame ranged from 94.5% to 100%, with a high average. Our initial tests using experimental image sequences (i.e., real data) of E. coli colonies also yield convincing results, with a registration accuracy ranging from 90% to 100%.

cs.CV

Pattern recognition in micro-trading behaviors before stock price jumps: A framework based on multivariate time series analysis

Studying the micro-trading behaviors before stock price jumps is an important problem for financial regulations and investment decisions. In this study, we provide a new framework to study pre-jump trading behaviors based on multivariate time series analysis. Different from the existing literature, our methodology takes into account the temporal information embedded in the trading-related attributes and can better evaluate and compare the abnormality levels of different attributes. Moreover, it can explore the joint informativeness of the attributes as well as select a subset of highly informative but minimally redundant attributes to analyze the homogeneous and idiosyncratic patterns in the pre-jump trades of individual stocks. In addition, our analysis involves a set of technical indicators to describe micro-trading behaviors. To illustrate the viability of the proposed methodology, an application case is conducted based on the level-2 data of 189 constituent stocks of the China Security Index 300. The individual and joint informativeness levels of the attributes in predicting price jumps are evaluated and compared. To this end, our experiment provides a set of jump indicators that can represent the pre-jump trading behaviors in the Chinese stock market and have detected some stocks with extremely abnormal pre-jump trades.

q-fin.ST

Diffeomorphic Shape Matching by Operator Splitting in 3D Cardiology Imaging

We develop an operator splitting approach to solve diffeomorphic matching problems for sequences of surfaces in three-dimensional space. The goal is to smoothly match, at a very fast rate, finite sequences of observed 3D-snapshots extracted from movies recording the smooth dynamic deformations of "soft" surfaces. We have implemented our algorithms in a proprietary software installed at The Methodist Hospital (Cardiology) to monitor mitral valve strain through computer analysis of noninvasive patients' echocardiographies.

math.OC

Fitness Estimation for Genetic Evolution of Bacterial Populations

In this paper we develop and test algorithmic techniques to estimate genotypes fitnesses by analysis of observed daily frequency data monitoring the long-term evolution of bacterial populations. In particular, we develop a non-linear least squares approach to estimate selective advantages of emerging new mutant strains in locked-box stochastic models describing bacterial genetic evolution similar to the celebrated Lenski experiment on Escherichia Coli. Our algorithm first analyses emergence of new mutant strains for each individual trajectory. For each trajectory our analysis is progressive in time, and successively focuses on the first mutation event before analyzing the second mutation event. The basic principle applied here is to minimize (for each trajectory) the mean squared errors of prediction w(t) - W(t) where the observed white cell frequencies w(t) are predicted by W(t), which is computed as the conditional expectation of w(t) given the available information at time (t-1). The pooling of all selective advantages estimates across all trajectories provides histograms on which we perform a precise peak analysis to compute final estimates of selective advantages. We validate our approach using ensembles of simulated trajectories.

q-bio.PE

Realized volatility and parametric estimation of Heston SDEs

We present a detailed analysis of \emph{observable} moments based parameter estimators for the Heston SDEs jointly driving the rate of returns $R_t$ and the squared volatilities $V_t$. Since volatilities are not directly observable, our parameter estimators are constructed from empirical moments of realized volatilities $Y_t$, which are of course observable. Realized volatilities are computed over sliding windows of size $\varepsilon$, partitioned into $J(\varepsilon)$ intervals. We establish criteria for the joint selection of $J(\varepsilon)$ and of the sub-sampling frequency of return rates data. We obtain explicit bounds for the $L^q$ speed of convergence of realized volatilities to true volatilities as $\varepsilon \to 0$. In turn, these bounds provide also $L^q$ speeds of convergence of our observable estimators for the parameters of the Heston volatility SDE. Our theoretical analysis is supplemented by extensive numerical simulations of joint Heston SDEs to investigate the actual performances of our moments based parameter estimators. Our results provide practical guidelines for adequately fitting Heston SDEs parameters to observed stock prices series.

q-fin.CP

Predicting intraday jumps in stock prices using liquidity measures and technical indicators

Predicting the intraday stock jumps is a significant but challenging problem in finance. Due to the instantaneity and imperceptibility characteristics of intraday stock jumps, relevant studies on their predictability remain limited. This paper proposes a data-driven approach to predict intraday stock jumps using the information embedded in liquidity measures and technical indicators. Specifically, a trading day is divided into a series of 5-minute intervals, and at the end of each interval, the candidate attributes defined by liquidity measures and technical indicators are input into machine learning algorithms to predict the arrival of a stock jump as well as its direction in the following 5-minute interval. Empirical study is conducted on the level-2 high-frequency data of 1271 stocks in the Shenzhen Stock Exchange of China to validate our approach. The result provides initial evidence of the predictability of jump arrivals and jump directions using level-2 stock data as well as the effectiveness of using a combination of liquidity measures and technical indicators in this prediction. We also reveal the superiority of using random forest compared to other machine learning algorithms in building prediction models. Importantly, our study provides a portable data-driven approach that exploits liquidity and technical information from level-2 stock data to predict intraday price jumps of individual stocks.

q-fin.TR

Localization of Epileptic Seizure Focus by Computerized Analysis of fMRI Recordings

By computerized analysis of cortical activity recorded via fMRI for pediatric epilepsy patients, we implement algorithmic localization of epileptic seizure focus within one of eight cortical lobes. Our innovative machine learning techniques involve intensive analysis of large matrices of mutual information coefficients between pairs of anatomically identified cortical regions. Drastic selection of pairs of regions with significant inter-connectivity provide efficient inputs for our Multi-Layer Perceptron (MLP) classifier. By imposing rigorous parameter parsimony to avoid over fitting we construct a small size MLP with very good percentages of successful classification.

q-bio.QM

Large Deviations Analysis for Stochastic Models of Bacterial Evolution

Radical shifts in the genetic composition of large cell populations are rare events with quite low probabilities that direct numerical simulations generally fail to evaluate accurately. In this paper, we develop a theoretical large-deviation framework for a class of Markov chains modeling the genetic evolution of bacteria such as E. coli in ``locked-box'' laboratory experiments. In particular, we develop the cost function for discrete-time Markov chains that describe the daily evolution of histograms of bacterial populations. We obtain explicit formulas for the cost function for interior histograms. We also develop explicit formulas that can be used to numerically quantify the most likely evolutionary trajectories connecting an initial histogram and the target histogram.

math.PR

Large deviations for Gaussian diffusions with delay

Dynamical systems driven by nonlinear delay SDEs with small noise can exhibit important rare events on long timescales. When there is no delay, classical large deviations theory quantifies rare events such as escapes from metastable fixed points. Near such fixed points, one can approximate nonlinear delay SDEs by linear delay SDEs. Here, we develop a fully explicit large deviations framework for (necessarily Gaussian) processes $X_t$ driven by linear delay SDEs with small diffusion coefficients. Our approach enables fast numerical computation of the action functional controlling rare events for $X_t$ and of the most likely paths transiting from $X_0 = p$ to $X_T=q$. Via linear noise local approximations, we can then compute most likely routes of escape from metastable states for nonlinear delay SDEs. We apply our methodology to the detailed dynamics of a genetic regulatory circuit, namely the co-repressive toggle switch, which may be described by a nonlinear chemical Langevin SDE with delay.

math.PR

Region-of-Interest reconstruction from truncated cone-beam projections

Region-of-Interest (ROI) tomography aims at reconstructing a region of interest $C$ inside a body using only x-ray projections intersecting $C$ with the goal to reduce overall radiation exposure when only a small specific region of the body needs to be examined. We consider x-ray acquisition from sources located on a smooth curve $Γ$ in $\mathbb{R}^3$ verifying classical Tuy's condition. In this situation, the {\it non-trucated} cone-beam transform $D f$ of smooth densities $f$ admits an explicit inverse $Z$; however $Z$ cannot directly reconstruct $f$ from ROI-truncated projections. To deal with the ROI tomography problem, we introduce a novel reconstruction approach. For densities $f$ in $L^{\infty}(B)$ where $B$ is a bounded ball in $\mathbb{R}^3$, our method iterates an operator $U$ combining ROI-truncated projections, inversion by the operator $Z$ and appropriate regularization operators. Assuming only knowledge of projections corresponding to a spherical ROI $C \subset B$, given $ε>0$, we prove that if $C$ is sufficiently large our iterative reconstruction algorithm converges uniformly to an $ε$-accurate approximation of $f$, where the accuracy depends on the regularity of $f$ quantified in the Sobolev norm $W^5(B)$. This result shows the existence of a critical ROI radius ensuring the convergence of the ROI reconstruction algorithm to $ε$-accurate approximations of $f$. We numerically verified these theoretical results using simulated acquisition of ROI-truncated cone-beam projection data for multiple acquisition geometries. Numerical experiments indicate that the critical ROI radius is fairly small with respect to the support region~$B$.

math-ph

Parametric Estimation from Approximate Data: Non-Gaussian Diffusions

We study the problem of parameters estimation in Indirect Observability contexts, where $X_t \in R^r$ is an unobservable stationary process parametrized by a vector of unknown parameters and all observable data are generated by an approximating process $Y^{\varepsilon}_t$ which is close to $X_t$ in $L^4$ norm. We construct consistent parameter estimators which are smooth functions of the sub-sampled empirical mean and empirical lagged covariance matrices computed from the observable data. We derive explicit optimal sub-sampling schemes specifying the best paired choices of sub-sampling time-step and number of observations. We show that these choices ensure that our parameter estimators reach optimized asymptotic $L^2$-convergence rates, which are constant multiples of the $L^4$ norm $|| Y^{\varepsilon}_t - X_t ||$.

math.PR

Option Pricing Accuracy for Estimated Heston Models

We consider assets for which price $X_t$ and squared volatility $Y_t$ are jointly driven by Heston joint stochastic differential equations (SDEs). When the parameters of these SDEs are estimated from $N$ sub-sampled data $(X_{nT}, Y_{nT})$, estimation errors do impact the classical option pricing PDEs. We estimate these option pricing errors by combining numerical evaluation of estimation errors for Heston SDEs parameters with the computation of option price partial derivatives with respect to these SDEs parameters. This is achieved by solving six parabolic PDEs with adequate boundary conditions. To implement this approach, we also develop an estimator $\hat λ$ for the market price of volatility risk, and we study the sensitivity of option pricing to estimation errors affecting $\hat λ$. We illustrate this approach by fitting Heston SDEs to 252 daily joint observations of the S\&P 500 index and of its approximate volatility VIX, and by numerical applications to European options written on the S\&P 500 index.

q-fin.MF