Searcharxiv⌕ Search

arXiv subjects

Ariel Neufeld

Publications and source records attributed to Ariel Neufeld.

At least 55 records · Page 3Linked to original sources

Non-asymptotic convergence bounds for modified tamed unadjusted Langevin algorithm in non-convex setting

We consider the problem of sampling from a high-dimensional target distribution $π_β$ on $\mathbb{R}^d$ with density proportional to $θ\mapsto e^{-βU(θ)}$ using explicit numerical schemes based on discretising the Langevin stochastic differential equation (SDE). In recent literature, taming has been proposed and studied as a method for ensuring stability of Langevin-based numerical schemes in the case of super-linearly growing drift coefficients for the Langevin SDE. In particular, the Tamed Unadjusted Langevin Algorithm (TULA) was proposed in [Bro+19] to sample from such target distributions with the gradient of the potential $U$ being super-linearly growing. However, theoretical guarantees in Wasserstein distances for Langevin-based algorithms have traditionally been derived assuming strong convexity of the potential $U$. In this paper, we propose a novel taming factor and derive, under a setting with possibly non-convex potential $U$ and super-linearly growing gradient of $U$, non-asymptotic theoretical bounds in Wasserstein-1 and Wasserstein-2 distances between the law of our algorithm, which we name the modified Tamed Unadjusted Langevin Algorithm (mTULA), and the target distribution $π_β$. We obtain respective rates of convergence $\mathcal{O}(λ)$ and $\mathcal{O}(λ^{1/2})$ in Wasserstein-1 and Wasserstein-2 distances for the discretisation error of mTULA in step size $λ$. High-dimensional numerical simulations which support our theoretical findings are presented to showcase the applicability of our algorithm.

math.PR↗

Bounding the Difference between the Values of Robust and Non-Robust Markov Decision Problems

In this note we provide an upper bound for the difference between the value function of a distributionally robust Markov decision problem and the value function of a non-robust Markov decision problem, where the ambiguity set of probability kernels of the distributionally robust Markov decision process is described by a Wasserstein-ball around some reference kernel whereas the non-robust Markov decision process behaves according to a fixed probability kernel contained in the ambiguity set. Our derived upper bound for the difference between the value functions is dimension-free and depends linearly on the radius of the Wasserstein-ball.

math.OC↗

Detecting data-driven robust statistical arbitrage strategies with deep neural networks

We present an approach, based on deep neural networks, that allows identifying robust statistical arbitrage strategies in financial markets. Robust statistical arbitrage strategies refer to trading strategies that enable profitable trading under model ambiguity. The presented novel methodology allows to consider a large amount of underlying securities simultaneously and does not depend on the identification of cointegrated pairs of assets, hence it is applicable on high-dimensional financial markets or in markets where classical pairs trading approaches fail. Moreover, we provide a method to build an ambiguity set of admissible probability measures that can be derived from observed market data. Thus, the approach can be considered as being model-free and entirely data-driven. We showcase the applicability of our method by providing empirical investigations with highly profitable trading performances even in 50 dimensions, during financial crises, and when the cointegration relationship between asset pairs stops to persist.

q-fin.CP↗

Binary Spatial Random Field Reconstruction from Non-Gaussian Inhomogeneous Time-series Observations

We develop a new model for spatial random field reconstruction of a binary-valued spatial phenomenon. In our model, sensors are deployed in a wireless sensor network across a large geographical region. Each sensor measures a non-Gaussian inhomogeneous temporal process which depends on the spatial phenomenon. Two types of sensors are employed: one collects point observations at specific time points, while the other collects integral observations over time intervals. Subsequently, the sensors transmit these time-series observations to a Fusion Center (FC), and the FC infers the spatial phenomenon from these observations. We show that the resulting posterior predictive distribution is intractable and develop a tractable two-step procedure to perform inference. Firstly, we develop algorithms to perform approximate Likelihood Ratio Tests on the time-series observations, compressing them to a single bit for both point sensors and integral sensors. Secondly, once the compressed observations are transmitted to the FC, we utilize a Spatial Best Linear Unbiased Estimator (S-BLUE) to reconstruct the binary spatial random field at any desired spatial location. The performance of the proposed approach is studied using simulation. We further illustrate the effectiveness of our method using a weather dataset from the National Environment Agency (NEA) of Singapore with fields including temperature and relative humidity.

eess.SP↗

A Bonus-Malus Framework for Cyber Risk Insurance and Optimal Cybersecurity Provisioning

The cyber risk insurance market is at a nascent stage of its development, even as the magnitude of cyber losses is significant and the rate of cyber loss events is increasing. Existing cyber risk insurance products as well as academic studies have been focusing on classifying cyber loss events and developing models of these events, but little attention has been paid to proposing insurance risk transfer strategies that incentivise mitigation of cyber loss through adjusting the premium of the risk transfer product. To address this important gap, we develop a Bonus-Malus model for cyber risk insurance. Specifically, we propose a mathematical model of cyber risk insurance and cybersecurity provisioning supported with an efficient numerical algorithm based on dynamic programming. Through a numerical experiment, we demonstrate how a properly designed cyber risk insurance contract with a Bonus-Malus system can resolve the issue of moral hazard and benefit the insurer.

math.OC↗

Improved Robust Price Bounds for Multi-Asset Derivatives under Market-Implied Dependence Information

We show how inter-asset dependence information derived from market prices of options can lead to improved model-free price bounds for multi-asset derivatives. Depending on the type of the traded option, we either extract correlation information or we derive restrictions on the set of admissible copulas that capture the inter-asset dependencies. To compute the resultant price bounds for some multi-asset options of interest, we apply a modified martingale optimal transport approach. Several examples based on simulated and real market data illustrate the improvement of the obtained price bounds and thus provide evidence for the relevance and tractability of our approach.

q-fin.MF↗

An efficient Monte Carlo scheme for Zakai equations

In this paper we develop a numerical method for efficiently approximating solutions of certain Zakai equations in high dimensions. The key idea is to transform a given Zakai SPDE into a PDE with random coefficients. We show that under suitable regularity assumptions on the coefficients of the Zakai equation, the corresponding random PDE admits a solution random field which, for almost all realizations of the random coefficients, can be written as a classical solution of a linear parabolic PDE. This makes it possible to apply the Feynman--Kac formula to obtain an efficient Monte Carlo scheme for computing approximate solutions of Zakai equations. The approach achieves good results in up to 25 dimensions with fast run times.

math.NA↗

Non-asymptotic estimates for TUSLA algorithm for non-convex learning with applications to neural networks with ReLU activation function

We consider non-convex stochastic optimization problems where the objective functions have super-linearly growing and discontinuous stochastic gradients. In such a setting, we provide a non-asymptotic analysis for the tamed unadjusted stochastic Langevin algorithm (TUSLA) introduced in Lovas et al. (2020). In particular, we establish non-asymptotic error bounds for the TUSLA algorithm in Wasserstein-1 and Wasserstein-2 distances. The latter result enables us to further derive non-asymptotic estimates for the expected excess risk. To illustrate the applicability of the main results, we consider an example from transfer learning with ReLU neural networks, which represents a key paradigm in machine learning. Numerical experiments are presented for the aforementioned example which support our theoretical findings. Hence, in this setting, we demonstrate both theoretically and numerically that the TUSLA algorithm can solve the optimization problem involving neural networks with ReLU activation function. Besides, we provide simulation results for synthetic examples where popular algorithms, e.g. ADAM, AMSGrad, RMSProp, and (vanilla) stochastic gradient descent (SGD) algorithm, may fail to find the minimizer of the objective functions due to the super-linear growth and the discontinuity of the corresponding stochastic gradient, while the TUSLA algorithm converges rapidly to the optimal solution. Moreover, we provide an empirical comparison of the performance of TUSLA with popular stochastic optimizers on real-world datasets, as well as investigate the effect of the key hyperparameters of TUSLA on its performance.

math.OC↗

Markov Decision Processes under Model Uncertainty

We introduce a general framework for Markov decision problems under model uncertainty in a discrete-time infinite horizon setting. By providing a dynamic programming principle we obtain a local-to-global paradigm, namely solving a local, i.e., a one time-step robust optimization problem leads to an optimizer of the global (i.e. infinite time-steps) robust stochastic optimal control problem, as well as to a corresponding worst-case measure. Moreover, we apply this framework to portfolio optimization involving data of the S&P 500. We present two different types of ambiguity sets; one is fully data-driven given by a Wasserstein-ball around the empirical measure, the second one is described by a parametric set of multivariate normal distributions, where the corresponding uncertainty sets of the parameters are estimated from the data. It turns out that in scenarios where the market is volatile or bearish, the optimal portfolio strategies from the corresponding robust optimization problem outperforms the ones without model uncertainty, showcasing the importance of taking model uncertainty into account.

math.OC↗

A deep learning approach to data-driven model-free pricing and to martingale optimal transport

We introduce a novel and highly tractable supervised learning approach based on neural networks that can be applied for the computation of model-free price bounds of, potentially high-dimensional, financial derivatives and for the determination of optimal hedging strategies attaining these bounds. In particular, our methodology allows to train a single neural network offline and then to use it online for the fast determination of model-free price bounds of a whole class of financial derivatives with current market data. We show the applicability of this approach and highlight its accuracy in several examples involving real market data. Further, we show how a neural network can be trained to solve martingale optimal transport problems involving fixed marginal distributions instead of financial market data.

q-fin.CP↗

Model-free bounds for multi-asset options using option-implied information and their exact computation

We consider derivatives written on multiple underlyings in a one-period financial market, and we are interested in the computation of model-free upper and lower bounds for their arbitrage-free prices. We work in a completely realistic setting, in that we only assume the knowledge of traded prices for other single- and multi-asset derivatives, and even allow for the presence of bid-ask spread in these prices. We provide a fundamental theorem of asset pricing for this market model, as well as a superhedging duality result, that allows to transform the abstract maximization problem over probability measures into a more tractable minimization problem over vectors, subject to certain constraints. Then, we recast this problem into a linear semi-infinite optimization problem, and provide two algorithms for its solution. These algorithms provide upper and lower bounds for the prices that are $\varepsilon$-optimal, as well as a characterization of the optimal pricing measures. These algorithms are efficient and allow the computation of bounds in high-dimensional scenarios (e.g. when $d=60$). Moreover, these algorithms can be used to detect arbitrage opportunities and identify the corresponding arbitrage strategies. Numerical experiments using both synthetic and real market data showcase the efficiency of these algorithms, while they also allow to understand the reduction of model risk by including additional information, in the form of known derivative prices.

math.OC↗

Model-free price bounds under dynamic option trading

In this paper we extend discrete time semi-static trading strategies by also allowing for dynamic trading in a finite amount of options, and we study the consequences for the model-independent super-replication prices of exotic derivatives. These include duality results as well as a precise characterization of pricing rules for the dynamically tradable options triggering an improvement of the price bounds for exotic derivatives in comparison with the conventional price bounds obtained through the martingale optimal transport approach.

q-fin.MF↗

Forecasting directional movements of stock prices for intraday trading using LSTM and random forests

We employ both random forests and LSTM networks (more precisely CuDNNLSTM) as training methodologies to analyze their effectiveness in forecasting out-of-sample directional movements of constituent stocks of the S&P 500 from January 1993 till December 2018 for intraday trading. We introduce a multi-feature setting consisting not only of the returns with respect to the closing prices, but also with respect to the opening prices and intraday returns. As trading strategy, we use Krauss et al. (2017) and Fischer & Krauss (2018) as benchmark. On each trading day, we buy the 10 stocks with the highest probability and sell short the 10 stocks with the lowest probability to outperform the market in terms of intraday returns -- all with equal monetary weight. Our empirical results show that the multi-feature setting provides a daily return, prior to transaction costs, of 0.64% using LSTM networks, and 0.54% using random forests. Hence we outperform the single-feature setting in Fischer & Krauss (2018) and Krauss et al. (2017) consisting only of the daily returns with respect to the closing prices, having corresponding daily returns of 0.41% and of 0.39% with respect to LSTM and random forests, respectively.

cs.LG↗

Deep splitting method for parabolic PDEs

In this paper we introduce a numerical method for nonlinear parabolic PDEs that combines operator splitting with deep learning. It divides the PDE approximation problem into a sequence of separate learning problems. Since the computational graph for each of the subproblems is comparatively small, the approach can handle extremely high-dimensional PDEs. We test the method on different examples from physics, stochastic control and mathematical finance. In all cases, it yields very good results in up to 10,000 dimensions with short run times.

math.NA↗

Low-Rank plus Sparse Decomposition of Covariance Matrices using Neural Network Parametrization

This paper revisits the problem of decomposing a positive semidefinite matrix as a sum of a matrix with a given rank plus a sparse matrix. An immediate application can be found in portfolio optimization, when the matrix to be decomposed is the covariance between the different assets in the portfolio. Our approach consists in representing the low-rank part of the solution as the product $MM^{T}$, where $M$ is a rectangular matrix of appropriate size, parametrized by the coefficients of a deep neural network. We then use a gradient descent algorithm to minimize an appropriate loss function over the parameters of the network. We deduce its convergence rate to a local optimum from the Lipschitz smoothness of our loss function. We show that the rate of convergence grows polynomially in the dimensions of the input, output, and the size of each of the hidden layers.

math.OC↗

Duality Theory for Robust Utility Maximisation

In this paper we present a duality theory for the robust utility maximisation problem in continuous time for utility functions defined on the positive real axis. Our results are inspired by -- and can be seen as the robust analogues of -- the seminal work of Kramkov & Schachermayer [18]. Namely, we show that if the set of attainable trading outcomes and the set of pricing measures satisfy a bipolar relation, then the utility maximisation problem is in duality with a conjugate problem. We further discuss the existence of optimal trading strategies. In particular, our general results include the case of logarithmic and power utility, and they apply to drift and volatility uncertainty.

q-fin.MF↗

On the stability of the martingale optimal transport problem: A set-valued map approach

Continuity of the value of the martingale optimal transport problem on the real line w.r.t. its marginals was recently established in Backhoff-Veraguas and Pammer [2] and Wiesel [21]. We present a new perspective of this result using the theory of set-valued maps. In particular, using results from Beiglböck, Jourdain, Margheriti, and Pammer [5], we show that the set of martingale measures with fixed marginals is continuous, i.e., lower- and upper hemicontinuous, w.r.t. its marginals. Moreover, we establish compactness of the set of optimizers as well as upper hemicontinuity of the optimizers w.r.t. the marginals.

math.PR↗