SearcharxivSearch

arXiv subjects

Yukito Iba

Publications and source records attributed to Yukito Iba.

At least 19 recordsLinked to original sources

{poscosea} : A Computationally Efficient Sensitivity Analysis for Bayesian Models using the posterior covariance representation

Bayesian methods are essential in modern data analysis in ecology and evolutionary biology. They provide a flexible framework for modeling complex data-generating processes, while quantifying uncertainty based on the classical subjective interpretation of probability. However, Bayesian inference may provide misleading measures of uncertainty, particularly when the fitted model fails to adequately represent the true data-generating process. Although nonparametric approaches such as leave-k-out diagnostics and bootstrap resampling offer more robust alternatives under model misspecification, their computational cost is often too demanding because they require repeatedly refitting the same Bayesian model. In this paper, we introduce posterior covariance sensitivity analysis (PosCoSeA), a computationally efficient strategy for approximating leave-k-out diagnostics and bootstrap resampling without repeated model refitting. The performance of the methods is evaluated through both simulation studies and an application to real ecological data. We also provide an \textsf{R} package that implements these methods to facilitate their practical application.

stat.ME

Bias correction of posterior means using MCMC outputs

We propose algorithms for addressing the bias of the posterior mean when used as an estimator of parameters. These algorithms build upon the recently proposed Bayesian infinitesimal jackknife approximation (Giordano and Broderick (2023)) and can be implemented using the posterior covariance and third-order combined cumulants easily calculated from MCMC outputs. Two algorithms are introduced: The first algorithm utilises the output of a single-run MCMC with the original likelihood and prior to estimate the bias. A notable feature of the algorithm is that its ability to estimate definitional bias (Efron (2015)), which is crucial for Bayesian estimators. The second algorithm is designed for high-dimensional and sparse data settings, where ``quasi-prior'' for bias correction is introduced. The quasi-prior is iteratively refined using the output of the first algorithm as a measure of the residual bias at each step. These algorithms have been successfully implemented and tested for parameter estimation in the Weibull distribution and logistic regression in moderately high-dimensional settings.

stat.ME

W-Kernel and Its Principal Space for Frequentist Evaluation of Bayesian Estimators

Evaluating the variability of posterior estimates is a key aspect of Bayesian model assessment. In this study, we focus on the posterior covariance matrix W, defined through the log likelihoods of individual observations. Previous studies, notably MacEachern and Peruggia (2002) and Thomas et al. (2018), examined the role of the principal space of W in Bayesian sensitivity analysis. Here, we show that the principal space of W is also central to frequentist evaluation, using the recently proposed Bayesian infinitesimal jackknife (Bayesian IJ) approximation (Giordano & Broderick, 2023) as a key tool. We further clarify the relationship between W and the Fisher kernel, showing that a modified version of the Fisher kernel can be viewed as an approximation to W. Moreover, the matrix W itself can be interpreted as a reproducing kernel, which we refer to as the W-kernel. Based on this connection, we investigate the relation between the W-kernel formulation in the data space and the classical asymptotic formulation in the parameter space. We also introduce the matrix Z, which is effectively dual to W in the sense of PCA; this formulation provides another perspective on the relationship between W and classical asymptotic theory. In the appendixes, we explore approximate bootstrap methods for posterior means and show that projection onto the principal space of W facilitates frequentist evaluation when higher-order terms are included. In addition, we introduce incomplete Cholesky decomposition as an efficient method for computing the principal space of W and discuss the concept of representative subsets of observations.

stat.ME

Posterior Covariance Information Criterion for Weighted Inference

For predictive evaluation based on quasi-posterior distributions, we develop a new information criterion, the posterior covariance information criterion (PCIC. PCIC generalises the widely applicable information criterion WAIC so as to effectively handle predictive scenarios where likelihoods for the estimation and the evaluation of the model may be different. A typical example of such scenarios is the weighted likelihood inference, including prediction under covariate shift and counterfactual prediction. The proposed criterion utilises a posterior covariance form and is computed by using only one Markov chain Monte Carlo run. Through numerical examples, we demonstrate how PCIC can apply in practice. Further, we show that PCIC is asymptotically unbiased to the quasi-Bayesian generalization error under mild conditions in weighted inference with both regular and singular statistical models.

stat.ME

Posterior covariance information criterion for general loss functions

We propose a novel computationally low-cost method for estimating a general predictive measure of generalised Bayesian inference. The proposed method utilises posterior covariance and provides estimators of the Gibbs and the plugin generalisation errors. We present theoretical guarantees of the proposed method, clarifying the connection to the Bayesian sensitivity analysis and the infinitesimal jackknife approximation of Bayesian leave-one-out cross validation. We illustrate several applications of our methods, including applications to differential privacy-preserving learning, the Bayesian hierarchical modeling, the Bayesian regression in the presence of influential observations, and the bias reduction of the widely-applicable information criterion. The applicability in high dimensions is also discussed.

stat.ME

Backward Simulation of Stochastic Process using a Time Reverse Monte Carlo method

The "backward simulation" of a stochastic process is defined as the stochastic dynamics that trace a time-reversed path from the target region to the initial configuration. If the probabilities calculated by the original simulation are easily restored from those obtained by backward dynamics, we can use it as a computational tool. It is shown that the naive approach to backward simulation does not work as expected. As a remedy, the Time Reverse Monte Carlo method (TRMC) based on the ideas of Sequential Importance Sampling (SIS) and Sequential Monte Carlo (SMC) is proposed and successfully tested with a stochastic typhoon model and the Lorenz 96 model. TRMC with SMC, which contains resampling steps, is shown to be more efficient for simulations with a larger number of time steps. A limitation of TRMC and its relation to the Bayes formula are also discussed.

physics.data-an

Multicanonical MCMC for Sampling Rare Events

Multicanonical MCMC (Multicanonical Markov Chain Monte Carlo; Multicanonical Monte Carlo) is discussed as a method of rare event sampling. Starting from a review of the generic framework of importance sampling, multicanonical MCMC is introduced, followed by applications in random matrices, random graphs, and chaotic dynamical systems. Replica exchange MCMC (also known as parallel tempering or Metropolis-coupled MCMC) is also explained as an alternative to multicanonical MCMC. In the last section, multicanonical MCMC is applied to data surrogation; a successful implementation in surrogating time series is shown. In the appendices, calculation of averages and normalizing constant in an exponential family, phase coexistence, simulated tempering, parallelization, and multivariate extensions are discussed.

cond-mat.stat-mech

Multicanonical sampling of rare events in random matrices

A method based on multicanonical Monte Carlo is applied to the calculation of large deviations in the largest eigenvalue of random matrices. The method is successfully tested with the Gaussian orthogonal ensemble (GOE), sparse random matrices, and matrices whose components are subject to uniform density. Specifically, the probability that all eigenvalues of a matrix are negative is estimated in these cases down to the values of $\sim 10^{-200}$, a region where naive random sampling is ineffective. The method can be applied to any ensemble of matrices and used for sampling rare events characterized by any statistics.

cond-mat.stat-mech

Multicanonical Sampling of Rare Trajectories in Chaotic Dynamical Systems

In chaotic dynamical systems, a number of rare trajectories with low level of chaoticity are embedded in chaotic sea, while extraordinary unstable trajectories can exist even in weakly chaotic regions. In this study, a quantitative method for searching these rare trajectories is developed; the method is based on multicanonical Monte Carlo and can estimate the probability of initial conditions that lead to trajectory fragments of a given level of chaoticity. The proposed method is successfully tested with four-dimensional coupled standard maps, where probabilities as small as $10^{-14}$ are estimated.

nlin.CD

Probability of graphs with large spectral gap by multicanonical Monte Carlo

Graphs with large spectral gap are important in various fields such as biology, sociology and computer science. In designing such graphs, an important question is how the probability of graphs with large spectral gap behaves. A method based on multicanonical Monte Carlo is introduced to quantify the behavior of this probability, which enables us to calculate extreme tails of the distribution. The proposed method is successfully applied to random 3-regular graphs and large deviation probability is estimated.

cond-mat.stat-mech

Exploration of Order in Chaos with Replica Exchange Monte Carlo

A method for exploring unstable structures generated by nonlinear dynamical systems is introduced. It is based on the sampling of initial conditions and parameters by Replica Exchange Monte Carlo (REM), and efficient both for the search of rare initial conditions and for the combined search of rare initial conditions and parameters. Examples discussed here include the sampling of unstable periodic orbits in chaos and search for the stable manifold of unstable fixed points, as well as construction of the global bifurcation diagram of a map.

cond-mat.stat-mech

Testing Error Correcting Codes by Multicanonical Sampling of Rare Events

The idea of rare event sampling is applied to the estimation of the performance of error-correcting codes. The essence of the idea is importance sampling of the pattern of noises in the channel by Multicanonical Monte Carlo, which enables efficient estimation of tails of the distribution of bit error rate. The idea is successfully tested with a convolutional code.

cond-mat.dis-nn

A Monte Carlo Algorithm for Sampling Rare Events: Application to a Search for the Griffiths Singularity

We develop a recently proposed importance-sampling Monte Carlo algorithm for sampling rare events and quenched variables in random disordered systems. We apply it to a two dimensional bond-diluted Ising model and study the Griffiths singularity which is considered to be due to the existence of rare large clusters. It is found that the distribution of the inverse susceptibility has an exponential tail down to the origin which is considered the consequence of the Griffiths singularity.

cond-mat.dis-nn

Detecting Generalized Synchronization Between Chaotic Signals: A Kernel-based Approach

A unified framework for analyzing generalized synchronization in coupled chaotic systems from data is proposed. The key of the proposed approach is the use of the kernel methods recently developed in the field of machine learning. Several successful applications are presented, which show the capability of the kernel-based approach for detecting generalized synchronization. It is also shown that the dynamical change of the coupling coefficient between two chaotic systems can be captured by the proposed approach.

nlin.CD

Multi-dimensional Density of States by Multicanonical Monte Carlo

Multi-dimensional density of states provides a useful description of complex frustrated systems. Recent advances in Monte Carlo methods enable efficient calculation of the density of states and related quantities, which renew the interest in them. Here we calculate density of states on the plane (energy, magnetization) for an Ising Model with three-spin interactions on a random sparse network, which is a system of current interest both in physics of glassy systems and in the theory of error-correcting codes. Multicanonical Monte Carlo algorithm is successfully applied, and the shape of densities and its dependence on the degree of frustration is revealed. Efficiency of multicanonical Monte Carlo is also discussed with the shape of a projection of the distribution simulated by the algorithm.

cond-mat.dis-nn

Eigenmode analysis of the susceptibility matrix of the four-dimensional Edwards-Anderson spin-glass model

The nature of spin-glass phase of the four-dimensional Edwards-Anderson Ising model is numerically studied by eigenmode analysis of the susceptibility matrix up to the lattice size 10^4. Unlike the preceding results on smaller lattices, our result suggests that there exist multiple extensive eigenvalues of the matrix, which does not contradict replica-symmetry-breaking scenarios. The sensitivity of the eigenmodes with respect to a temperature change is examined using finite-size-scaling analysis and an evidence of anomalous sensitivity is found. A computational advantage of dual formulation of the eigenmode analysis in the study of large lattices is also discussed.

cond-mat.dis-nn

Extended Ensemble Monte Carlo

``Extended Ensemble Monte Carlo''is a generic term that indicates a set of algorithms which are now popular in a variety of fields in physics and statistical information processing. Exchange Monte Carlo (Metropolis-Coupled Chain, Parallel Tempering), Simulated Tempering (Expanded Ensemble Monte Carlo), and Multicanonical Monte Carlo (Adaptive Umbrella Sampling) are typical members of this family. Here we give a cross-disciplinary survey of these algorithms with special emphasis on the great flexibility of the underlying idea. In Sec.2, we discuss the background of Extended Ensemble Monte Carlo. In Sec.3, 4 and 5, three types of the algorithms, i.e., Exchange Monte Carlo, Simulated Tempering, Multicanonical Monte Carlo are introduced. In Sec.6, we give an introduction to Replica Monte Carlo algorithm by Swendsen and Wang. Strategies for the construction of special-purpose extended ensembles are discussed in Sec.7. We stress that an extension is not necessary restricted to the space of energy or temperature. Even unphysical (unrealizable) configurations can be included in the ensemble, if the resultant fast mixing of the Markov chain offsets the increasing cost of the sampling procedure. Multivariate (multi-component) extensions are also useful in many examples. In Sec.8, we give a survey on extended ensembles with a state space whose dimensionality is dynamically varying. In the appendix, we discuss advantages and disadvantages of three types of extended ensemble algorithms.

cond-mat.dis-nn