SearcharxivSearch

arXiv subjects

Liangliang Wang

Publications and source records attributed to Liangliang Wang.

At least 19 recordsLinked to original sources

Compound Auxiliary Metropolis: Incorporating Auxiliary Variables into Multi-Candidate MCMC

Multiple-try Metropolis (MTM) is a Markov chain Monte Carlo (MCMC) algorithm that improves local transition efficiency by evaluating multiple candidate draws at each iteration. However, for complicated target distributions exhibiting severely non-Gaussian topography or multiple well-separated modes, locally optimal transitions may be insufficient for effective global exploration. In this work, we propose compound auxiliary Metropolis (CAM), a general multi-candidate MCMC method that incorporates both the local state of the chain and auxiliary information into the multi-candidate framework of MTM. Using an auxiliary generating distribution, CAM accommodates a flexible definition of auxiliary information. As examples, we consider three different auxiliary variables: one that promotes state-independent exploration and two that use a reference distribution to improve mixing. These auxiliaries are tested against distributions that present challenging targets for modern MCMC methods. In particular, we focus on the challenges presented by multiple well-separated modes and topography that requires long mixing for local MCMC moves. We find that CAM is able to sample effectively from these distributions, using MTM as a baseline to evaluate the benefit introduced by the auxiliary information. CAM also compares favourably with the No-U-Turn Sampler, showing similar performance for milder test distributions and better performance for the most difficult settings.

stat.CO

A Unified Framework for Multiple-Try Metropolis: Construction and Empirical Benchmarks

The multiple-try Metropolis (MTM) algorithm uses a compound proposal with multiple candidate draws to improve local sampling efficiency. While several methodological works have continued to develop MTM and the multi-candidate mechanism that characterizes it, the literature lacks a unified comparison of these components. This paper presents a structured formulation of MTM within the involutive MCMC framework, providing a principled approach for deriving valid acceptance probabilities based on the proposal mechanism. Through a comprehensive simulation experiment, we evaluate the impact of MTM configurations on non-Gaussian and multimodal target distributions. Our results reveal that while weight functions are a focus of several methodological developments, their impact on stationary sampling efficiency is secondary to the configuration of the proposal distribution. Furthermore, we find that while increasing the number of candidates enhances per-iteration efficiency, the realized performance gains are offset by computational overhead introduced by multiple candidacy unless parallelize computing is used. Our findings offer practical guidance for configuring an MTM algorithm for complex and non-Gaussian targets.

stat.CO

Generalized Bayesian Multidimensional Scaling and Model Comparison

Multidimensional scaling (MDS) is widely used to reconstruct a low-dimensional representation of high-dimensional data while preserving pairwise distances. However, Bayesian MDS approaches based on Markov chain Monte Carlo (MCMC) face challenges in model generalization and comparison. To address these limitations, we propose a generalized Bayesian multidimensional scaling (GBMDS) framework that accommodates non-Gaussian errors and diverse dissimilarity metrics for improved robustness. We develop an adaptive annealed Sequential Monte Carlo (ASMC) algorithm for Bayesian inference, leveraging an annealing schedule to enhance posterior exploration and computational efficiency. The ASMC algorithm also provides a nearly unbiased marginal likelihood estimator, enabling principled Bayesian model comparison across different error distributions, dissimilarity metrics, and dimensional choices. Using synthetic and real data, we demonstrate the effectiveness of the proposed approach. Our results show that ASMC-based GBMDS achieves superior computational efficiency and robustness compared to MCMC-based methods under the same computational budget. The implementation of our proposed method and applications are available at https://github.com/SFU-Stat-ML/GBMDS.

stat.ME

The Astronomical Plate Digitization at SHAO

The digitization of historical astronomical plates is essential for preserving century-long observational data. This work presents the development and application of the specialized digitizers at the Shanghai Astronomical Observatory (SHAO), including technical details, international collaborations, and scientific applications on the plates.

astro-ph.IM

S2R-Bench: A Sim-to-Real Evaluation Benchmark for Autonomous Driving

Safety is a long-standing and the final pursuit in the development of autonomous driving systems, with a significant portion of safety challenge arising from perception. How to effectively evaluate the safety as well as the reliability of perception algorithms is becoming an emerging issue. Despite its critical importance, existing perception methods exhibit a limitation in their robustness, primarily due to the use of benchmarks are entierly simulated, which fail to align predicted results with actual outcomes, particularly under extreme weather conditions and sensor anomalies that are prevalent in real-world scenarios. To fill this gap, in this study, we propose a Sim-to-Real Evaluation Benchmark for Autonomous Driving (S2R-Bench). We collect diverse sensor anomaly data under various road conditions to evaluate the robustness of autonomous driving perception methods in a comprehensive and realistic manner. This is the first corruption robustness benchmark based on real-world scenarios, encompassing various road conditions, weather conditions, lighting intensities, and time periods. By comparing real-world data with simulated data, we demonstrate the reliability and practical significance of the collected data for real-world applications. We hope that this dataset will advance future research and contribute to the development of more robust perception models for autonomous driving. This dataset is released on https://github.com/adept-thu/S2R-Bench.

cs.RO

Particle Data Cloning for Complex Ordinary Differential Equations

Ordinary differential equations (ODEs) are fundamental tools for modeling complex dynamic systems across scientific disciplines. However, parameter estimation in ODE models is challenging due to the multimodal nature of the likelihood function, which can lead to local optima and unstable inference. In this paper, we propose particle data cloning (PDC), a novel approach that enhances global optimization by leveraging data cloning and annealed sequential Monte Carlo (ASMC). PDC mitigates multimodality by refining the likelihood through data clones and progressively extracting information from the sharpened posterior. Compared to standard data cloning, PDC provides more reliable frequentist inference and demonstrates superior global optimization performance. We offer practical guidelines for efficient implementation and illustrate the method through simulation studies and an application to a prey-predator ODE model. Our implementation is available at https://github.com/SONDONGHUI/PDC.

stat.CO

Robust Bayesian Functional Principal Component Analysis

We develop a robust Bayesian functional principal component analysis (RB-FPCA) method that utilizes the skew elliptical class of distributions to model functional data, which are observed over a continuous domain. This approach effectively captures the primary sources of variation among curves, even in the presence of outliers, and provides a more robust and accurate estimation of the covariance function and principal components. The proposed method can also handle sparse functional data, where only a few observations per curve are available. We employ annealed sequential Monte Carlo for posterior inference, which offers several advantages over conventional Markov chain Monte Carlo algorithms. To evaluate the performance of our proposed model, we conduct simulation studies, comparing it with well-known frequentist and conventional Bayesian methods. The results show that our method outperforms existing approaches in the presence of outliers and performs competitively in outlier-free datasets. Finally, we demonstrate the effectiveness of our method by applying it to environmental and biological data to identify outlying functional observations. The implementation of our proposed method and applications are available at https://github.com/SFU-Stat-ML/RBFPCA.

stat.ME

SPAFIT: Stratified Progressive Adaptation Fine-tuning for Pre-trained Large Language Models

Full fine-tuning is a popular approach to adapt Transformer-based pre-trained large language models to a specific downstream task. However, the substantial requirements for computational power and storage have discouraged its widespread use. Moreover, increasing evidence of catastrophic forgetting and overparameterization in the Transformer architecture has motivated researchers to seek more efficient fine-tuning (PEFT) methods. Commonly known parameter-efficient fine-tuning methods like LoRA and BitFit are typically applied across all layers of the model. We propose a PEFT method, called Stratified Progressive Adaptation Fine-tuning (SPAFIT), based on the localization of different types of linguistic knowledge to specific layers of the model. Our experiments, conducted on nine tasks from the GLUE benchmark, show that our proposed SPAFIT method outperforms other PEFT methods while fine-tuning only a fraction of the parameters adjusted by other methods.

cs.CL

A switching state-space transmission model for tracking epidemics and assessing interventions

The effective control of infectious diseases relies on accurate assessment of the impact of interventions, which is often hindered by the complex dynamics of the spread of disease. A Beta-Dirichlet switching state-space transmission model is proposed to track underlying dynamics of disease and evaluate the effectiveness of interventions simultaneously. As time evolves, the switching mechanism introduced in the susceptible-exposed-infected-recovered (SEIR) model is able to capture the timing and magnitude of changes in the transmission rate due to the effectiveness of control measures. The implementation of this model is based on a particle Markov Chain Monte Carlo algorithm, which can estimate the time evolution of SEIR states, switching states, and high-dimensional parameters efficiently. The efficacy of the proposed model and estimation procedure are demonstrated through simulation studies. With a real-world application to British Columbia's COVID-19 outbreak, the proposed switching state-space transmission model quantifies the reduction of transmission rate following interventions. The proposed model provides a promising tool to inform public health policies aimed at studying the underlying dynamics and evaluating the effectiveness of interventions during the spread of the disease.

stat.ME

Simulation study of BESIII with stitched CMOS pixel detector using ACTS

Reconstruction of tracks of charged particles with high precision is very crucial for HEP experiments to achieve their physics goals. As the tracking detector of BESIII experiment, the BESIII drift chamber has suffered from aging effects resulting in degraded tracking performance after operation for about 15 years. To preserve and enhance the tracking performance of BESIII, one of the proposals is to add one layer of thin CMOS pixel sensor in cylindrical shape based on the state-of-the-art stitching technology, between the beam pipe and the drift chamber. The improvement of tracking performance of BESIII with such an additional pixel detector compared to that with only the existing drift chamber is studied using the modern common tracking software ACTS, which provides a set of detector-agnostic and highly performant tracking algorithms that have demonstrated promising performance for a few high energy physics and nuclear physics experiments.

physics.ins-det

Adaptive semiparametric Bayesian differential equations via sequential Monte Carlo

Nonlinear differential equations (DEs) are used in a wide range of scientific problems to model complex dynamic systems. The differential equations often contain unknown parameters that are of scientific interest, which have to be estimated from noisy measurements of the dynamic system. Generally, there is no closed-form solution for nonlinear DEs, and the likelihood surface for the parameter of interest is multi-modal and very sensitive to different parameter values. We propose a Bayesian framework for nonlinear DE systems. A flexible nonparametric function is used to represent the dynamic process such that expensive numerical solvers can be avoided. A sequential Monte Carlo algorithm in the annealing framework is proposed to conduct Bayesian inference for parameters in DEs. In our numerical experiments, we use examples of ordinary differential equations and delay differential equations to demonstrate the effectiveness of the proposed algorithm. We developed an R package that is available at \url{https://github.com/shijiaw/smcDE}.

stat.CO

An Iterative Weighting Method to Apply ISR Correction to $e^+e^-$ Hadronic Cross-section Measurements

Initial state radiation (ISR) plays an important role in $e^+$$e^-$ collision experiments such as the BESIII. To correct the ISR effects in measurements of hadronic cross-sections of $e^+e^-$ annihilation, an iterative method that weights simulated ISR events is proposed here to assess the efficiency of event selection and the ISR correction factor for the observed cross-section. The simulated ISR events were generated only once, and the obtained cross-sectional line shape was used iteratively to weigh the same simulated ISR events to evaluate the efficiency and corrections until the results converge. Compared with the method of generating ISR events iteratively, the proposed weighting method provides consistent results, and reduces the computational time and disk space required by a factor of five or more, thus speeding-up $e^+e^-$ hadronic cross-section measurements.

hep-ex

Particle Gibbs Sampling for Bayesian Phylogenetic inference

The combinatorial sequential Monte Carlo (CSMC) has been demonstrated to be an efficient complementary method to the standard Markov chain Monte Carlo (MCMC) for Bayesian phylogenetic tree inference using biological sequences. It is appealing to combine the CSMC and MCMC in the framework of the particle Gibbs (PG) sampler to jointly estimate the phylogenetic trees and evolutionary parameters. However, the Markov chain of the particle Gibbs may mix poorly if the underlying SMC suffers from the path degeneracy issue. Some remedies, including the particle Gibbs with ancestor sampling and the interacting particle MCMC, have been proposed to improve the PG. But they either cannot be applied to or remain inefficient for the combinatorial tree space. We introduce a novel CSMC method by proposing a more efficient proposal distribution. It also can be combined into the particle Gibbs sampler framework to infer parameters in the evolutionary model. The new algorithm can be easily parallelized by allocating samples over different computing cores. We validate that the developed CSMC can sample trees more efficiently in various particle Gibbs samplers via numerical experiments. Our implementation is available at https://github.com/liangliangwangsfu/phyloPMCMC

stat.CO

Spectral Dynamic Causal Modelling of Resting-State fMRI: Relating Effective Brain Connectivity in the Default Mode Network to Genetics

We conduct an imaging genetics study to explore how effective brain connectivity in the default mode network (DMN) may be related to genetics within the context of Alzheimer's disease and mild cognitive impairment. We develop an analysis of longitudinal resting-state functional magnetic resonance imaging (rs-fMRI) and genetic data obtained from a sample of 111 subjects with a total of 319 rs-fMRI scans from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. A Dynamic Causal Model (DCM) is fit to the rs-fMRI scans to estimate effective brain connectivity within the DMN and related to a set of single nucleotide polymorphisms (SNPs) contained in an empirical disease-constrained set which is obtained out-of-sample from 663 ADNI subjects having only genome-wide data. We examine longitudinal data in both a 4-region and an 6-region network and relate longitudinal effective brain connectivity networks estimated using spectral DCM to SNPs using both linear mixed effect (LME) models as well as function-on-scalar regression (FSR). In the former case we implement a parametric bootstrap for testing SNP coefficients and make comparisons with p-values obtained from the chi-squared null distribution. We also implement a parametric bootstrap approach for testing regression functions in FSR and we make comparisons between p-values obtained from the parametric bootstrap to p-values obtained using the F-distribution with degrees-of-freedom based on Satterthwaite's approximation. In both networks we report on exploratory patterns of associations with relatively high ranks that exhibit stability to the differing assumptions made by both FSR and LME.

q-bio.NC

A Bayesian Spatial Model for Imaging Genetics

We develop a Bayesian bivariate spatial model for multivariate regression analysis applicable to studies examining the influence of genetic variation on brain structure. Our model is motivated by an imaging genetics study of the Alzheimer's Disease Neuroimaging Initiative (ADNI), where the objective is to examine the association between images of volumetric and cortical thickness values summarizing the structure of the brain as measured by magnetic resonance imaging (MRI) and a set of 486 SNPs from 33 Alzheimer's Disease (AD) candidate genes obtained from 632 subjects. A bivariate spatial process model is developed to accommodate the correlation structures typically seen in structural brain imaging data. First, we allow for spatial correlation on a graph structure in the imaging phenotypes obtained from a neighbourhood matrix for measures on the same hemisphere of the brain. Second, we allow for correlation in the same measures obtained from different hemispheres (left/right) of the brain. We develop a mean-field variational Bayes algorithm and a Gibbs sampling algorithm to fit the model. We also incorporate Bayesian false discovery rate (FDR) procedures to select SNPs. We implement the methodology in a new release of the R package bgsmtr. We show that the new spatial model demonstrates superior performance over a standard model in our application. Data used in the preparation of this article were obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database (adni.loni.usc.edu).

stat.ME

Random Tessellation Forests

Space partitioning methods such as random forests and the Mondrian process are powerful machine learning methods for multi-dimensional and relational data, and are based on recursively cutting a domain. The flexibility of these methods is often limited by the requirement that the cuts be axis aligned. The Ostomachion process and the self-consistent binary space partitioning-tree process were recently introduced as generalizations of the Mondrian process for space partitioning with non-axis aligned cuts in the two dimensional plane. Motivated by the need for a multi-dimensional partitioning tree with non-axis aligned cuts, we propose the Random Tessellation Process (RTP), a framework that includes the Mondrian process and the binary space partitioning-tree process as special cases. We derive a sequential Monte Carlo algorithm for inference, and provide random forest methods. Our process is self-consistent and can relax axis-aligned constraints, allowing complex inter-dimensional dependence to be captured. We present a simulation study, and analyse gene expression data of brain tissue, showing improved accuracies over other methods.

stat.ML

An Annealed Sequential Monte Carlo Method for Bayesian Phylogenetics

We describe an "embarrassingly parallel" method for Bayesian phylogenetic inference, annealed Sequential Monte Carlo, based on recent advances in the Sequential Monte Carlo literature such as adaptive determination of annealing parameters. The algorithm provides an approximate posterior distribution over trees and evolutionary parameters as well as an unbiased estimator for the marginal likelihood. This unbiasedness property can be used for the purpose of testing the correctness of posterior simulation software. We evaluate the performance of phylogenetic annealed Sequential Monte Carlo by reviewing and comparing with other computational Bayesian phylogenetic methods, in particular, different marginal likelihood estimation methods. Unlike previous Sequential Monte Carlo methods in phylogenetics, our annealed method can utilize standard Markov chain Monte Carlo tree moves and hence benefit from the large inventory of such moves available in the literature. Consequently, the annealed Sequential Monte Carlo method should be relatively easy to incorporate into existing phylogenetic software packages based on Markov chain Monte Carlo algorithms. We illustrate our method using simulation studies and real data analysis.

q-bio.PE

Normal and pathological dynamics of platelets in humans

We develop a comprehensive mathematical model of platelet, megakaryocyte, and thrombopoietin dynamics in humans. We show that there is a single stationary solution that can undergo a Hopf bifurcation, and use this information to investigate both normal and pathological platelet production, specifically cyclic thrombocytopenia. Carefully estimating model parameters from laboratory and clinical data, we then argue that a subset of parameters are involved in the genesis of cyclic thrombocytopenia based on clinical information. We provide excellent model fits to the existing data for both platelet counts and thrombopoietin levels by changing six parameters that have physiological correlates. Our results indicate that the primary change in cyclic thrombocytopenia is a major interference with or destruction of the thrombopoietin receptor with secondary changes in other processes, including immune-mediated destruction of platelets and megakaryocyte deficiency and failure in platelet production. This study makes a major contribution to the understanding of the origin of cyclic thrombopoietin as well as significantly extending the modeling of thrombopoiesis.

q-bio.TO