SearcharxivSearch

arXiv subjects

Cheng-Der Fuh

Publications and source records attributed to Cheng-Der Fuh.

17 recordsLinked to original sources

Change, dependence, and discovery: Celebrating the work of T.L. Lai

Tze Leung Lai made seminal contributions to sequential analysis, particularly in sequential hypothesis testing, changepoint detection and nonlinear renewal theory. His work established fundamental optimality results for the sequential probability ratio test and its extensions, and provided a general framework for testing composite hypotheses. In changepoint detection, he introduced new optimality criteria and computationally efficient procedures that remain influential. He applied these and related tools to problems in biostatistics. In this article, we review these key results in the broader context of sequential analysis.

stat.OT

Efficient importance sampling for copula models

In this paper, we propose an efficient importance sampling algorithm for rare event simulation under copula models. In the algorithm, the derived optimal probability measure is based on the criterion of minimizing the variance of the importance sampling estimator within a parametric exponential tilting family. Since the copula model is defined by its marginals and a copula function, and its moment-generating function is difficult to derive, we apply the transform likelihood ratio method to first identify an alternative exponential tilting family, after which we obtain simple and explicit expressions of equations. Then, the optimal alternative probability measure can be calculated under this transformed exponential tilting family. The proposed importance sampling framework is quite general and can be implemented for many classes of copula models, including some traditional parametric copula families and a class of semiparametric copulas called regular vine copulas, from which sampling is feasible. The theoretical results of the logarithmic efficiency and bounded relative error are proved for some commonly-used copula models under the case of simple rare events. Monte Carlo experiments are conducted, in which we study the relative efficiency of the crude Monte Carlo estimator with respect to the proposed importance-sampling-based estimators, such that substantial variance reductions are obtained in comparison to the standard Monte Carlo estimators.

stat.CO

Determine the Number of States in Hidden Markov Models via Marginal Likelihood

Hidden Markov models (HMM) have been widely used by scientists to model stochastic systems: the underlying process is a discrete Markov chain and the observations are noisy realizations of the underlying process. Determining the number of hidden states for an HMM is a model selection problem, which is yet to be satisfactorily solved, especially for the popular Gaussian HMM with heterogeneous covariance. In this paper, we propose a consistent method for determining the number of hidden states of HMM based on the marginal likelihood, which is obtained by integrating out both the parameters and hidden states. Moreover, we show that the model selection problem of HMM includes the order selection problem of finite mixture models as a special case. We give rigorous proof of the consistency of the proposed marginal likelihood method and provide an efficient computation method for practical implementation. We numerically compare the proposed method with the Bayesian information criterion (BIC), demonstrating the effectiveness of the proposed marginal likelihood method.

math.ST

A General Framework for Importance Sampling with Markov Random Walks

Although stochastic models driven by latent Markov processes are widely used, the classical importance sampling methods based on the exponential tilting for these models suffers from the difficulties in computing the eigenvalues and associated eigenfunctions and the plausibility of the indirect asymptotic large deviation regime for the variance of the estimator. We propose a general importance sampling framework that twists the observable and latent processes separately using a link function that directly minimizes the estimator's variance. An optimal choice of the link function is chosen within the locally asymptotically normal family. We show the logarithmic efficiency of the proposed estimator. As applications, we estimate an overflow probability under a pandemic model and the CoVaR, a measurement of the co-dependent financial systemic risk. Both applications are beyond the scope of traditional importance sampling methods due to their nonlinear features.

stat.CO

Kullback-Leibler Divergence and Akaike Information Criterion in General Hidden Markov Models

To characterize the Kullback-Leibler divergence and Fisher information in general parametrized hidden Markov models, in this paper, we first show that the log likelihood and its derivatives can be represented as an additive functional of a Markovian iterated function system, and then provide explicit characterizations of these two quantities through this representation. Moreover, we show that Kullback-Leibler divergence can be locally approximated by a quadratic function determined by the Fisher information. Results relating to the Cram\'{e}r-Rao lower bound and the H\'{a}jek-Le Cam local asymptotic minimax theorem are also given. As an application of our results, we provide a theoretical justification of using Akaike information criterion (AIC) model selection in general hidden Markov models. Last, we study three concrete models: a Gaussian vector autoregressive-moving average model of order $(p,q)$, recurrent neural networks, and temporal restricted Boltzmann machine, to illustrate our theory.

math.ST

Rényi Divergence in General Hidden Markov Models

In this paper, we examine the existence of the Rényi divergence between two time invariant general hidden Markov models with arbitrary positive initial distributions. By making use of a Markov chain representation of the probability distribution for the general hidden Markov model and eigenvalue for the associated Markovian operator, we obtain, under some regularity conditions, convergence of the Rényi divergence. By using this device, we also characterize the Rényi divergence, and obtain the Kullback-Leibler divergence as α \rightarrow 1 of the Rényi divergence. Several examples, including the classical finite state hidden Markov models, Markov switching models, and recurrent neural networks, are given for illustration. Moreover, we develop a non-Monte Carlo method that computes the Rényi divergence of two-state Markov switching models via the underlying invariant probability measure, which is characterized by the Fredholm integral equation.

cs.IT

Asymptotically Optimal Change Point Detection for Composite Hypothesis in State Space Models

This paper investigates change point detection in state space models, in which the pre-change distribution $f^{θ_0}$ is given, while the poster distribution $f^θ$ after change is unknown. The problem is to raise an alarm as soon as possible after the distribution changes from $f^{θ_0}$ to $f^θ$, under a restriction on the false alarms. We investigate theoretical properties of a weighted Shiryayev-Roberts-Pollak (SRP) change point detection rule in state space models. By making use of a Markov chain representation for the likelihood function, exponential embedding of the induced Markovian transition operator, nonlinear Markov renewal theory, and sequential hypothesis testing theory for Markov random walks, we show that the weighted SRP procedure is second-order asymptotically optimal. To this end, we derive an asymptotic approximation for the expected stopping time of such a stopping scheme when the change time $ω= 1$. To illustrate our method we apply the results to two types of state space models: general state Markov chains and linear state space models.

math.PR

Efficient Exponential Tilting for Portfolio Credit Risk

This paper considers the problem of measuring the credit risk in portfolios of loans, bonds, and other instruments subject to possible default under multi-factor models. Due to the amount of the portfolio, the heterogeneous effect of obligors, and the phenomena that default events are rare and mutually dependent, it is difficult to calculate portfolio credit risk either by means of direct analysis or crude Monte Carlo under such models. To capture the extreme dependence among obligors, we provide an efficient simulation method for multi-factor models with a normal mixture copula that allows the multivariate defaults to have an asymmetric distribution, while most of the literature focuses on simulating one-dimensional cases. To this end, we first propose a general account of an importance sampling algorithm based on an unconventional exponential embedding, which is related to the classical sufficient statistic. Note that this innovative tilting device is more suitable for the multivariate normal mixture model than traditional one-parameter tilting methods and is of independent interest. Next, by utilizing a fast computational method for how the rare event occurs and the proposed importance sampling method, we provide an efficient simulation algorithm to estimate the probability that the portfolio incurs large losses under the normal mixture copula. Here the proposed simulation device is based on importance sampling for a joint probability other than the conditional probability used in previous studies. Theoretical investigations and simulation studies, which include an empirical example, are given to illustrate the method.

q-fin.CP

On spherical Monte Carlo simulations for multivariate normal probabilities

The calculation of multivariate normal probabilities is of great importance in many statistical and economic applications. This paper proposes a spherical Monte Carlo method with both theoretical analysis and numerical simulation. First, the multivariate normal probability is rewritten via an inner radial integral and an outer spherical integral by the spherical transformation. For the outer spherical integral, we apply an integration rule by randomly rotating a predetermined set of well-located points. To find the desired set, we derive an upper bound for the variance of the Monte Carlo estimator and propose a set which is related to the kissing number problem in sphere packings. For the inner radial integral, we employ the idea of antithetic variates and identify certain conditions so that variance reduction is guaranteed. Extensive Monte Carlo experiments on some probabilities calculation confirm these claims.

stat.CO

Efficient Importance Sampling for Rare Event Simulation with Applications

Importance sampling has been known as a powerful tool to reduce the variance of Monte Carlo estimator for rare event simulation. Based on the criterion of minimizing the variance of Monte Carlo estimator within a parametric family, we propose a general account for finding the optimal tilting measure. To this end, when the moment generating function of the underlying distribution exists, we obtain a simple and explicit expression of the optimal alternative distribution. The proposed algorithm is quite general to cover many interesting examples, such as normal distribution, noncentral $χ^2$ distribution, and compound Poisson processes. To illustrate the broad applicability of our method, we study value-at-risk (VaR) computation in financial risk management and bootstrap confidence regions in statistical inferences.

stat.ME

Efficient likelihood estimation in state space models

Motivated by studying asymptotic properties of the maximum likelihood estimator (MLE) in stochastic volatility (SV) models, in this paper we investigate likelihood estimation in state space models. We first prove, under some regularity conditions, there is a consistent sequence of roots of the likelihood equation that is asymptotically normal with the inverse of the Fisher information as its variance. With an extra assumption that the likelihood equation has a unique root for each $n$, then there is a consistent sequence of estimators of the unknown parameters. If, in addition, the supremum of the log likelihood function is integrable, the MLE exists and is strongly consistent. Edgeworth expansion of the approximate solution of likelihood equation is also established. Several examples, including Markov switching models, ARMA models, (G)ARCH models and stochastic volatility (SV) models, are given for illustration.

math.ST

Estimation in hidden Markov models via efficient importance sampling

Given a sequence of observations from a discrete-time, finite-state hidden Markov model, we would like to estimate the sampling distribution of a statistic. The bootstrap method is employed to approximate the confidence regions of a multi-dimensional parameter. We propose an importance sampling formula for efficient simulation in this context. Our approach consists of constructing a locally asymptotically normal (LAN) family of probability distributions around the default resampling rule and then minimizing the asymptotic variance within the LAN family. The solution of this minimization problem characterizes the asymptotically optimal resampling scheme, which is given by a tilting formula. The implementation of the tilting formula is facilitated by solving a Poisson equation. A few numerical examples are given to demonstrate the efficiency of the proposed importance sampling scheme.

stat.CO

Multi-armed bandit problem with precedence relations

Consider a multi-phase project management problem where the decision maker needs to deal with two issues: (a) how to allocate resources to projects within each phase, and (b) when to enter the next phase, so that the total expected reward is as large as possible. We formulate the problem as a multi-armed bandit problem with precedence relations. In Chan, Fuh and Hu (2005), a class of asymptotically optimal arm-pulling strategies is constructed to minimize the shortfall from perfect information payoff. Here we further explore optimality properties of the proposed strategies. First, we show that the efficiency benchmark, which is given by the regret lower bound, reduces to those in Lai and Robbins (1985), Hu and Wei (1989), and Fuh and Hu (2000). This implies that the proposed strategy is also optimal under the settings of aforementioned papers. Secondly, we establish the super-efficiency of proposed strategies when the bad set is empty. Thirdly, we show that they are still optimal with constant switching cost between arms. In addition, we prove that the Wald's equation holds for Markov chains under Harris recurrent condition, which is an important tool in studying the efficiency of the proposed strategies.

math.ST

Optimal strategies for a class of sequential control problems with precedence relations

Consider the following multi-phase project management problem. Each project is divided into several phases. All projects enter the next phase at the same point chosen by the decision maker based on observations up to that point. Within each phase, one can pursue the projects in any order. When pursuing the project with one unit of resource, the project state changes according to a Markov chain. The probability distribution of the Markov chain is known up to an unknown parameter. When pursued, the project generates a random reward depending on the phase and the state of the project and the unknown parameter. The decision maker faces two problems: (a) how to allocate resources to projects within each phase, and (b) when to enter the next phase, so that the total expected reward is as large as possible. In this paper, we formulate the preceding problem as a stochastic scheduling problem and propose asymptotic optimal strategies, which minimize the shortfall from perfect information payoff. Concrete examples are given to illustrate our method.

math.ST

Asymptotic operating characteristics of an optimal change point detection in hidden Markov models

Let ξ_0,ξ_1,...,ξ_{ω-1} be observations from the hidden Markov model with probability distribution P^{θ_0}, and let ξ_ω,ξ_{ω+1},... be observations from the hidden Markov model with probability distribution P^{θ_1}. The parameters θ_0 and θ_1 are given, while the change point ωis unknown. The problem is to raise an alarm as soon as possible after the distribution changes from P^{θ_0} to P^{θ_1}, but to avoid false alarms. Specifically, we seek a stopping rule N which allows us to observe the ξ's sequentially, such that E_{\infty}N is large, and subject to this constraint, sup_kE_k(N-k|N\geq k) is as small as possible. Here E_k denotes expectation under the change point k, and E_{\infty} denotes expectation under the hypothesis of no change whatever. In this paper we investigate the performance of the Shiryayev-Roberts-Pollak (SRP) rule for change point detection in the dynamic system of hidden Markov models. By making use of Markov chain representation for the likelihood function, the structure of asymptotically minimax policy and of the Bayes rule, and sequential hypothesis testing theory for Markov random walks, we show that the SRP procedure is asymptotically minimax in the sense of Pollak [Ann. Statist. 13 (1985) 206-227]. Next, we present a second-order asymptotic approximation for the expected stopping time of such a stopping scheme when ω=1. Motivated by the sequential analysis in hidden Markov models, a nonlinear renewal theory for Markov random walks is also given.

math.ST

Uniform Markov Renewal Theory and Ruin Probabilities in Markov Random Walks

Let {X_n,n\geq0} be a Markov chain on a general state space X with transition probability P and stationary probability π. Suppose an additive component S_n takes values in the real line R and is adjoined to the chain such that {(X_n,S_n),n\geq0} is a Markov random walk. In this paper, we prove a uniform Markov renewal theorem with an estimate on the rate of convergence. This result is applied to boundary crossing problems for {(X_n,S_n),n\geq0}. To be more precise, for given b\geq0, define the stopping time τ=τ(b)=inf{n:S_n>b}. When a drift μof the random walk S_n is 0, we derive a one-term Edgeworth type asymptotic expansion for the first passage probabilities P_π{τ<m} and P_π{τ<m,S_m<c}, where m\leq\infty, c\leq b and P_π denotes the probability under the initial distribution π. When μ\neq0, Brownian approximations for the first passage probabilities with correction terms are derived.

math.PR