SearcharxivSearch

arXiv subjects

David Gunawan

Publications and source records attributed to David Gunawan.

At least 19 recordsLinked to original sources

Flexible Bayesian Models for Time-Varying Income Distributions

Survey data are widely used to study how income inequality, poverty, and welfare evolve over time. A common practice is to estimate the income distribution separately for each year, treating annual observations as independent cross-sections. For population subgroups with relatively small sample sizes, however, this approach can produce unstable parameter estimates, imprecise inference for inequality and poverty measures, and potentially misleading posterior probabilities of Lorenz and stochastic dominance. This paper develops flexible Bayesian models for time-varying income distributions that borrow strength across adjacent years by allowing the parameters of income distributions to evolve dynamically. We consider a random walk specification and an extended model with shrinkage priors. The proposed framework yields coherent inference for the full income distributions over time, as well as for associated inequality measures, poverty indices, and dominance probabilities. Simulation studies show that, relative to independent year-by-year models, the proposed approach produces substantially more precise and stable inference, while avoiding spurious variation in welfare comparisons. An application to the Aboriginal and residents of the Australian Capital Territory (ACT) population subgroups in the Household, Income and Labour Dynamics in Australia survey shows that the dynamic models deliver improved inference for income distributions and related welfare measures, and can change conclusions about distributional dominance over time.

econ.EM

Neural Inference Functions for Margins for Time Series Copula Models

Copula models are widely employed in multivariate time series analysis because they permit flexible modelling of marginal distributions independently of the dependence structure, which is fully characterised by the copula function. However, Bayesian inference with these models becomes computationally demanding as the number of variables in the time series increases. Motivated by the classical inference functions for margins (IFM) approach, we propose a new neural-network based inference framework for estimating parameters in copula models, termed the neural inference functions for margins (N-IFM). N-IFM enables rapid parameter estimation for new data, fast sequential prediction, and efficient model comparison via time-series validation. We assess the performance of N-IFM using both simulated and real datasets and compare it to Hamiltonian Monte Carlo, demonstrating substantial computational gains with comparable inferential accuracy.

stat.CO

Glow with the Flow: AI-Assisted Creation of Ambient Lightscapes for Music Videos

Designed light is an established modality for live performance and music playback. Despite the growing availability of consumer smart lighting, the creation of designed light for music visualization remains limited to professional contexts due to time and skill constraints. To address this, we present an AI-assisted system for generating ambient light sequences for music videos. Informed by professional design heuristics, the system extracts salient features from source video and audio to generate an editable preliminary design of object based ambient light effect. We evaluated the system by comparing its autonomous output against hand-authored designs for three music videos. Findings from responses by 32 participants indicate that the initial output provides a viable baseline for further refinement by human authors. This work demonstrates the utility of AI-assisted workflows in supporting the creation and adoption of designed light beyond professional venues.

cs.HC

Neural posterior inference with state-space models for calibrating ice sheet simulators

Ice sheet models are routinely used to quantify and project an ice sheet's contribution to sea level rise. In order for an ice sheet model to generate realistic projections, its parameters must first be calibrated using observational data; this is challenging due to the nonlinearity of the model equations, the high dimensionality of the underlying parameters, and limited data availability for validation. This study leverages the emerging field of neural posterior approximation for efficiently calibrating ice sheet model parameters and boundary conditions. We make use of a one-dimensional (flowline) Shallow-Shelf Approximation model in a state-space framework. A neural network is trained to infer the underlying parameters, namely the bedrock elevation and basal friction coefficient along the flowline, based on observations of ice velocity and ice surface elevation. Samples from the approximate posterior distribution of the parameters are then used within an ensemble Kalman filter to infer latent model states, namely the ice thickness along the flowline. We show through a simulation study that our approach yields more accurate estimates of the parameters and states than a state-augmented ensemble Kalman filter, which is the current state-of-the-art. We apply our approach to infer the bed elevation and basal friction along a flowline in Thwaites Glacier, Antarctica.

stat.AP

Bayesian copula-based spatial random effects models for inference with complex spatial data

In this article, we develop fully Bayesian, copula-based, spatial-statistical models for large, noisy, incomplete, and non-Gaussian spatial data. Our approach includes novel constructions of copulas that accommodate a spatial-random-effects structure, enabling low-rank representations and computationally efficient Bayesian inference. The spatial copula is used in a latent process model of the Bayesian hierarchical spatial-statistical model, and, conditional on the latent copula-based spatial process, the data model handles measurement errors and missing data. Our simulation studies show that a fully Bayesian approach delivers accurate and fast inference for both parameter estimation and spatial-process prediction, outperforming several benchmark methods, including fixed rank kriging (FRK). The new class of copula-based models is used to map atmospheric methane in the Bowen Basin, Queensland, Australia, from Sentinel 5P satellite data.

stat.ME

Bayesian Inference for Non-Gaussian Simultaneous Autoregressive Models with Missing Data

Standard simultaneous autoregressive (SAR) models typically assume normally distributed errors, an assumption often violated in real-world datasets that frequently exhibit non-normal, skewed, or heavy-tailed characteristics. New SAR models are proposed to capture these non-Gaussian features. The spatial error model (SEM), a widely used SAR-type model, is considered. Three novel SEMs are introduced, extending the standard Gaussian SEM. These extensions incorporate Student's $t$-distributed errors to accommodate heavy-tailed behaviour, one-to-one transformations of the response variable to address skewness, or a combination of both. Variational Bayes (VB) estimation methods are developed for these models, and the framework is further extended to handle missing response data under the missing not at random (MNAR) mechanism. Standard VB methods perform well with complete datasets; however, handling missing data requires a hybrid VB (HVB) approach, which integrates a Markov chain Monte Carlo (MCMC) sampler to generate missing values. The proposed VB methods are evaluated using both simulated and real-world datasets, demonstrating their robustness and effectiveness in dealing with non-Gaussian data and missing data in spatial models. Although the method is demonstrated using SAR models, the proposed model specifications and estimation approaches are widely applicable to various types of models for handling non-Gaussian data with missing values.

stat.ME

Enhancing Efficiency of Local Projections Estimation with Volatility Clustering in High-Frequency Data

This paper advances the local projections (LP) method by addressing its inefficiency in high-frequency economic and financial data with volatility clustering. We incorporate a generalized autoregressive conditional heteroskedasticity (GARCH) process to resolve serial correlation issues and extend the model with GARCH-X and GARCH-HAR structures. Monte Carlo simulations show that exploiting serial dependence in LP error structures improves efficiency across forecast horizons, remains robust to persistent volatility, and yields greater gains as sample size increases. Our findings contribute to refining LP estimation, enhancing its applicability in analyzing economic interventions and financial market dynamics.

econ.EM

Fast Variational Boosting for Latent Variable Models

We consider the problem of estimating complex statistical latent variable models using variational Bayes methods. These methods are used when exact posterior inference is either infeasible or computationally expensive, and they approximate the posterior density with a family of tractable distributions. The parameters of the approximating distribution are estimated using optimisation methods. This article develops a flexible Gaussian mixture variational approximation, where we impose sparsity in the precision matrix of each Gaussian component to reflect the appropriate conditional independence structure in the model. By introducing sparsity in the precision matrix and parameterising it using the Cholesky factor, each Gaussian mixture component becomes parsimonious (with a reduced number of non-zero parameters), while still capturing the dependence in the posterior distribution. Fast estimation methods based on global and local variational boosting moves combined with natural gradients and variance reduction methods are developed. The local boosting moves adjust an existing mixture component, and optimisation is only carried out on a subset of the variational parameters of a new component. The subset is chosen to target improvement of the current approximation in aspects where it is poor. The local boosting moves are fast because only a small number of variational parameters need to be optimised. The efficacy of the approach is illustrated by using simulated and real datasets to estimate generalised linear mixed models and state space models.

stat.ME

Sub-band Domain Multi-Hypothesis Acoustic Echo Canceler Based Acoustic Scene Analysis

This paper introduces a novel approach for acoustic scene analysis by exploiting an ensemble of statistics extracted from a sub-band domain multi-hypothesis acoustic echo canceler (SDMH-AEC). A well-designed SDMH-AEC employs multiple adaptive filtering strategies with potentially complementary behaviours during convergence, perturbations, and steady-state conditions. By aggregating statistics across the sub-bands, we derive a feature vector that exhibits strong discriminative power for distinguishing different acoustic events and estimating acoustic parameters. The complementary nature of the SDMH-AEC filters provides a rich source of information that can be extracted at insignificant cost for acoustic scene analysis tasks. We demonstrate the effectiveness of the proposed approach experimentally with real data containing double-talk, echo path change and events where the full-duplex device is physically moved. The extracted features enable acoustic scene analysis using existing echo cancellation algorithms and techniques.

eess.AS

Recursive variational Gaussian approximation with the Whittle likelihood for linear non-Gaussian state space models

Parameter inference for linear and non-Gaussian state space models is challenging because the likelihood function contains an intractable integral over the latent state variables. While Markov chain Monte Carlo (MCMC) methods provide exact samples from the posterior distribution as the number of samples goes to infinity, they tend to have high computational cost, particularly for observations of a long time series. When inference with MCMC methods is computationally expensive, variational Bayes (VB) methods are a useful alternative. VB methods approximate the posterior density of the parameters with a simple and tractable distribution found through optimisation. This work proposes a novel sequential VB approach that makes use of the Whittle likelihood for computationally efficient parameter inference in linear, non-Gaussian state space models. Our algorithm, called Recursive Variational Gaussian Approximation with the Whittle Likelihood (R-VGA-Whittle), updates the variational parameters by processing data in the frequency domain. At each iteration, R-VGA-Whittle requires the gradient and Hessian of the Whittle log-likelihood, which are available in closed form. Through several examples involving a linear Gaussian state space model; a univariate/bivariate stochastic volatility model; and a state space model with Student's t measurement error, where the latent states follow an autoregressive fractionally integrated moving average (ARFIMA) model, we show that R-VGA-Whittle provides good approximations to posterior distributions of the parameters, and that it is very computationally efficient when compared to asymptotically exact methods such as Hamiltonian Monte Carlo.

stat.CO

Bayesian Inference for Multidimensional Welfare Comparisons

Using both single-index measures and stochastic dominance concepts, we show how Bayesian inference can be used to make multivariate welfare comparisons. A four-dimensional distribution for the well-being attributes income, mental health, education, and happiness are estimated via Bayesian Markov chain Monte Carlo using unit-record data taken from the Household, Income and Labour Dynamics in Australia survey. Marginal distributions of beta and gamma mixtures and discrete ordinal distributions are combined using a copula. Improvements in both well-being generally and poverty magnitude are assessed using posterior means of single-index measures and posterior probabilities of stochastic dominance. The conditions for stochastic dominance depend on the class of utility functions that is assumed to define a social welfare function and the number of attributes in the utility function. Three classes of utility functions are considered, and posterior probabilities of dominance are computed for one, two, and four-attribute utility functions for three time intervals within the period 2001 to 2019.

econ.EM

Variational Bayes Inference for Spatial Error Models with Missing Data

The spatial error model (SEM) is a type of simultaneous autoregressive (SAR) model for analysing spatially correlated data. Markov chain Monte Carlo (MCMC) is one of the most widely used Bayesian methods for estimating SEM, but it has significant limitations when it comes to handling missing data in the response variable due to its high computational cost. Variational Bayes (VB) approximation offers an alternative solution to this problem. Two VB-based algorithms employing Gaussian variational approximation with factor covariance structure are presented, joint VB (JVB) and hybrid VB (HVB), suitable for both missing at random and not at random inference. When dealing with many missing values, the JVB is inaccurate, and the standard HVB algorithm struggles to achieve accurate inferences. Our modified versions of HVB enable accurate inference within a reasonable computational time, thus improving its performance. The performance of the VB methods is evaluated using simulated and real datasets.

stat.ME

A Marginal Maximum Likelihood Approach for Hierarchical Simultaneous Autoregressive Models with Missing Data

Efficient estimation methods for simultaneous autoregressive (SAR) models with missing data in the response variable have been well-explored in the literature. A common practice is to introduce measurement error into SAR models to separate the noise component from the spatial process. However, prior research has not considered incorporating measurement error into SAR models with missing data. Maximum likelihood estimation for such models, especially with large datasets, poses significant computational challenges. This paper proposes an efficient likelihood-based estimation method, the marginal maximum likelihood (ML), for estimating SAR models on large datasets with measurement errors and a high percentage of missing data in the response variable. The spatial error model (SEM) and the spatial autoregressive model (SAM), two popular SAR model types, are considered. The missing data mechanism is assumed to follow a missing at random (MAR) pattern. We propose a fast method for marginal ML estimation with a computational complexity of $O(n^{3/2})$, where $n$ is the total number of observations. This complexity applies when the spatial weight matrix is constructed based on a local neighbourhood structure. The effectiveness of the proposed methods is demonstrated through simulations and real-world data applications.

stat.ME

Optimal prediction of positive-valued spatial processes: asymmetric power-divergence loss

This article studies the use of asymmetric loss functions for the optimal prediction of positive-valued spatial processes. We focus on the family of power-divergence loss functions due to its many convenient properties, such as its continuity, convexity, relationship to well known divergence measures, and the ability to control the asymmetry and behaviour of the loss function via a power parameter. The properties of power-divergence loss functions, optimal power-divergence (OPD) spatial predictors, and related measures of uncertainty quantification are examined. In addition, we examine the notion of asymmetry in loss functions defined for positive-valued spatial processes and define an asymmetry measure that is applied to the power-divergence loss function and other common loss functions. The paper concludes with a spatial statistical analysis of zinc measurements in the soil of a floodplain of the Meuse River, Netherlands, using OPD spatial prediction.

stat.ME

DDSP-SFX: Acoustically-guided sound effects generation with differentiable digital signal processing

Controlling the variations of sound effects using neural audio synthesis models has been a difficult task. Differentiable digital signal processing (DDSP) provides a lightweight solution that achieves high-quality sound synthesis while enabling deterministic acoustic attribute control by incorporating pre-processed audio features and digital synthesizers. In this research, we introduce DDSP-SFX, a model based on the DDSP architecture capable of synthesizing high-quality sound effects while enabling users to control the timbre variations easily. We propose a transient modelling technique with higher objective evaluation scores and subjective ratings over impulsive signals (footsteps, gunshots). We propose a simple method that achieves timbre variation control while also allowing deterministic attribute control. We further qualitatively show the timbre transfer performance using voice as the guiding sound.

eess.AS

R-VGAL: A Sequential Variational Bayes Algorithm for Generalised Linear Mixed Models

Models with random effects, such as generalised linear mixed models (GLMMs), are often used for analysing clustered data. Parameter inference with these models is difficult because of the presence of cluster-specific random effects, which must be integrated out when evaluating the likelihood function. Here, we propose a sequential variational Bayes algorithm, called Recursive Variational Gaussian Approximation for Latent variable models (R-VGAL), for estimating parameters in GLMMs. The R-VGAL algorithm operates on the data sequentially, requires only a single pass through the data, and can provide parameter updates as new data are collected without the need of re-processing the previous data. At each update, the R-VGAL algorithm requires the gradient and Hessian of a "partial" log-likelihood function evaluated at the new observation, which are generally not available in closed form for GLMMs. To circumvent this issue, we propose using an importance-sampling-based approach for estimating the gradient and Hessian via Fisher's and Louis' identities. We find that R-VGAL can be unstable when traversing the first few data points, but that this issue can be mitigated by using a variant of variational tempering in the initial steps of the algorithm. Through illustrations on both simulated and real datasets, we show that R-VGAL provides good approximations to the exact posterior distributions, that it can be made robust through tempering, and that it is computationally efficient.

stat.CO

Bayesian Inference for Evidence Accumulation Models with Regressors

Evidence accumulation models (EAMs) are an important class of cognitive models used to analyze both response time and response choice data recorded from decision-making tasks. Developments in estimation procedures have helped EAMs become important both in basic scientific applications and solution-focussed applied work. Hierarchical Bayesian estimation frameworks for the linear ballistic accumulator model (LBA) and the diffusion decision model (DDM) have been widely used, but still suffer from some key limitations, particularly for large sample sizes, for models with many parameters, and when linking decision-relevant covariates to model parameters. We extend upon previous work with methods for estimating the LBA and DDM in hierarchical Bayesian frameworks that include random effects which are correlated between people, and include regression-model links between decision-relevant covariates and model parameters. Our methods work equally well in cases where the covariates are measured once per person (e.g., personality traits or psychological tests) or once per decision (e.g., neural or physiological data). We provide methods for exact Bayesian inference, using particle-based MCMC, and also approximate methods based on variational Bayesian (VB) inference. The VB methods are sufficiently fast and efficient that they can address large-scale estimation problems, such as with very large data sets. We evaluate the performance of these methods in applications to data from three existing experiments. Detailed algorithmic implementations and code are freely available for all methods.

stat.ME

Reliable Bayesian Inference in Misspecified Models

We provide a general solution to a fundamental open problem in Bayesian inference, namely poor uncertainty quantification, from a frequency standpoint, of Bayesian methods in misspecified models. While existing solutions are based on explicit Gaussian approximations of the posterior, or computationally onerous post-processing procedures, we demonstrate that correct uncertainty quantification can be achieved by replacing the usual posterior with an intuitive approximate posterior. Critically, our solution is applicable to likelihood-based, and generalized, posteriors as well as cases where the likelihood is intractable and must be estimated. We formally demonstrate the reliable uncertainty quantification of our proposed approach, and show that valid uncertainty quantification is not an asymptotic result but occurs even in small samples. We illustrate this approach through a range of examples, including linear, and generalized, mixed effects models.

stat.ME