SearcharxivSearch

arXiv subjects

Martin Lysy

Publications and source records attributed to Martin Lysy.

At least 19 recordsLinked to original sources

Sample Size Determination Under Selection Bias: Robust Tolerance Limits for Prevalent Cohort Data

Tolerance limits have received considerable attention in the statistical literature, with applications reaching far beyond their initial role in quality control. The well-known formula of Scheff\'e and Tukey (1944) establishes a simple, distribution-free relation between sample size and population coverage by two given order statistics and a given confidence level. A key requirement in applying this formula is the availability of an unbiased, representative sample from the population of interest. However, as it often happens in biological and medical applications, various logistical constraints may preclude the possibility of obtaining an unbiased sample. We derive extensions of this formula which accommodate a large class of biased sampling schemes including weight bias and censoring. The modified formulae are validated through a simulation study and compared to its unmodified counterpart. We illustrate the use of the modified formulae using the partially observed failure times for individuals with dementia using data collected from the Canadian Study of Health and Aging.

stat.ME

Rapid Scaling of Compositional Uncertainty from Sample to Population Levels

Understanding population composition is essential across ecological, evolutionary, conservation, and resource management contexts. Modern methods such as genetic stock identification (GSI) estimate the proportion of individuals from each subpopulation using genetic data. Ideally, these estimates are obtained through mixture analysis, which captures both sampling and genetic uncertainty. However, historical datasets often rely on individual assignment methods that only account for sample-level uncertainty, limiting the validity of population-level inferences. To address this, we propose a reverse Dirichlet-multinomial model and derive multiple variance estimators to propagate uncertainty from the sample to the population level. We extend this framework to genetic mark-recapture studies, assess performance via simulation, and apply our method to estimate the escapement of Sockeye Salmon (Oncorhynchus nerka) in the Taku River.

stat.ME

rodeo: Probabilistic Methods of Parameter Inference for Ordinary Differential Equations

Parameter estimation for ordinary differential equations (ODEs) plays a fundamental role in the analysis of dynamical systems. Generally lacking closed-form solutions, ODEs are traditionally approximated using deterministic solvers. However, there is a growing body of evidence to suggest that probabilistic ODE solvers produce more reliable parameter estimates by better accounting for numerical uncertainty. Here we present rodeo, a Python library providing a fast, lightweight, and extensible interface to a broad class of probabilistic ODE solvers, along with several associated methods for parameter inference. At its core, rodeo provides a probabilistic solver that scales linearly in both the number of evaluation points and system variables. Furthermore, by leveraging state-of-the-art automatic differentiation (AD) and just-in-time (JIT) compiling techniques, rodeo is shown across several examples to provide fast, accurate, and scalable parameter inference for a variety of ODE systems.

stat.CO

LocalCop: An R package for local likelihood inference for conditional copulas

Conditional copulas models allow the dependence structure between multiple response variables to be modelled as a function of covariates. LocalCop (Acar & Lysy, 2024) is an R/C++ package for computationally efficient semiparametric conditional copula modelling using a local likelihood inference framework developed in Acar, Craiu, & Yao (2011), Acar, Craiu, & Yao (2013) and Acar, Czado, & Lysy (2019).

stat.CO

Plant-Capture Methods for Estimating Population Size from Uncertain Plant Captures

Plant-capture is a variant of classical capture-recapture methods used to estimate the size of a population. In this method, decoys referred to as "plants" are introduced into the population in order to estimate the capture probability. The method has shown considerable success in estimating population sizes from limited samples in many epidemiological, ecological, and demographic studies. However, previous plant-recapture studies have not systematically accounted for uncertainty in the capture status of each individual plant. In this work, we propose various approaches to formally incorporate uncertainty into the plant-capture model arising from (i) the capture status of plants and (ii) the heterogeneity between multiple survey sites. We present two inference methods and compare their performance in simulation studies. We then apply our methods to estimate the size of the homeless population in several US cities using the large-scale "S-night" study conducted by the US Census Bureau.

stat.ME

BlackJAX: Composable Bayesian inference in JAX

BlackJAX is a library implementing sampling and variational inference algorithms commonly used in Bayesian computation. It is designed for ease of use, speed, and modularity by taking a functional approach to the algorithms' implementation. BlackJAX is written in Python, using JAX to compile and run NumpPy-like samplers and variational methods on CPUs, GPUs, and TPUs. The library integrates well with probabilistic programming languages by working directly with the (un-normalized) target log density function. BlackJAX is intended as a collection of low-level, composable implementations of basic statistical 'atoms' that can be combined to perform well-defined Bayesian inference, but also provides high-level routines for ease of use. It is designed for users who need cutting-edge methods, researchers who want to create complex sampling methods, and people who want to learn how these work.

cs.MS

Data-Adaptive Probabilistic Likelihood Approximation for Ordinary Differential Equations

Estimating the parameters of ordinary differential equations (ODEs) is of fundamental importance in many scientific applications. While ODEs are typically approximated with deterministic algorithms, new research on probabilistic solvers indicates that they produce more reliable parameter estimates by better accounting for numerical errors. However, many ODE systems are highly sensitive to their parameter values. This produces deep local maxima in the likelihood function -- a problem which existing probabilistic solvers have yet to resolve. Here we present a novel probabilistic ODE likelihood approximation, DALTON, which can dramatically reduce parameter sensitivity by learning from noisy ODE measurements in a data-adaptive manner. Our approximation scales linearly in both ODE variables and time discretization points, and is applicable to ODEs with both partially-unobserved components and non-Gaussian measurement models. Several examples demonstrate that DALTON produces more accurate parameter estimates via numerical optimization than existing probabilistic ODE solvers, and even in some cases than the exact ODE likelihood itself.

stat.ML

Functional Connectivity: Continuous-Time Latent Factor Models for Neural Spike Trains

Modelling the dynamics of interactions in a neuronal ensemble is an important problem in functional connectivity research. One popular framework is latent factor models (LFMs), which have achieved notable success in decoding neuronal population dynamics. However, most LFMs are specified in discrete time, where the choice of bin size significantly impacts inference results. In this work, we present what is, to the best of our knowledge, the first continuous-time multivariate spike train LFM for studying neuronal interactions and functional connectivity. We present an efficient parameter inference algorithm for our biologically justifiable model which (1) scales linearly in the number of simultaneously recorded neurons and (2) bypasses time binning and related issues. Simulation studies show that parameter estimation using the proposed model is highly accurate. Applying our LFM to experimental data from a classical conditioning study on the prefrontal cortex in rats, we found that coordinated neuronal activities are affected by (1) the onset of the cue for reward delivery, and (2) the sub-region within the frontal cortex (OFC/mPFC). These findings shed new light on our understanding of cue and outcome value encoding.

stat.ME

A Multivariate Point Process Model for Simultaneously Recorded Neural Spike Trains

The current state-of-the-art in neurophysiological data collection allows for simultaneous recording of tens to hundreds of neurons, for which point processes are an appropriate statistical modelling framework. However, existing point process models lack multivariate generalizations which are both flexible and computationally tractable. This paper introduces a multivariate generalization of the Skellam process with resetting (SPR), a point process tailored to model individual neural spike trains. The multivariate SPR (MSPR) is biologically justified as it mimics the process of neural integration. Its flexible dependence structure and a fast parameter estimation method make it well-suited for the analysis of simultaneously recorded spike trains from multiple neurons. The strengths and weaknesses of the MSPR are demonstrated through simulation and analysis of experimental data.

stat.ME

Multimodel Bayesian Analysis of Load Duration Effects in Lumber Reliability

This paper evaluates the reliability of lumber, accounting for the duration-of-load (DOL) effect under different load profiles based on a multimodel Bayesian approach. Three individual DOL models previously used for reliability assessment are considered: the US model, the Canadian model, and the Gamma process model. Procedures for stochastic generation of residential, snow, and wind loads are also described. We propose Bayesian model-averaging (BMA) as a method for combining the reliability estimates of individual models under a given load profile that coherently accounts for statistical uncertainty in the choice of model and parameter values. The method is applied to the analysis of a Hemlock experimental dataset, where the BMA results are illustrated via estimated reliability indices together with 95% interval bands.

stat.AP

Fast and Scalable Inference for Spatial Extreme Value Models

The generalized extreme value (GEV) distribution is a popular model for analyzing and forecasting extreme weather data. To increase prediction accuracy, spatial information is often pooled via a latent Gaussian process (GP) on the GEV parameters. Inference for GEV-GP models is typically carried out using Markov chain Monte Carlo (MCMC) methods, or using approximate inference methods such as the integrated nested Laplace approximation (INLA). However, MCMC becomes prohibitively slow as the number of spatial locations increases, whereas INLA is only applicable in practice to a limited subset of GEV-GP models. In this paper, we revisit the original Laplace approximation for fitting spatial GEV models. In combination with a popular sparsity-inducing spatial covariance approximation technique, we show through simulations that our approach accurately estimates the Bayesian predictive distribution of extreme weather events, is scalable to several thousand spatial locations, and is several orders of magnitude faster than MCMC. A case study in forecasting extreme snowfall across Canada is presented.

stat.ME

Measurement Error Correction in Particle Tracking Microrheology

In diverse biological applications, particle tracking of passive microscopic species has become the experimental measurement of choice -- when either the materials are of limited volume, or so soft as to deform uncontrollably when manipulated by traditional instruments. In a wide range of particle tracking experiments, a ubiquitous finding is that the mean squared displacement (MSD) of particle positions exhibits a power-law signature, the parameters of which reveal valuable information about the viscous and elastic properties of various biomaterials. However, MSD measurements are typically contaminated by complex and interacting sources of instrumental noise. As these often affect the high-frequency bandwidth to which MSD estimates are particularly sensitive, inadequate error correction can lead to severe bias in power law estimation and thereby, the inferred viscoelastic properties. In this article, we propose a novel strategy to filter high-frequency noise from particle tracking measurements. Our filters are shown theoretically to cover a broad spectrum of high-frequency noises, and lead to a parametric estimator of MSD power-law coefficients for which an efficient computational implementation is presented. Based on numerous analyses of experimental and simulated data, results suggest our methods perform very well compared to other denoising procedures.

stat.AP

A Heteroscedastic Accelerated Failure Time Model for Survival Analysis

Nonparametric and semiparametric methods are commonly used in survival analysis to mitigate the bias due to model misspecification. However, such methods often cannot estimate upper-tail survival quantiles when a sizable proportion of the data are censored, in which case parametric likelihood-based estimators present a viable alternative. In this article, we extend a popular family of parametric survival models which make the Accelerated Failure Time (AFT) assumption to account for heteroscedasticity in the survival times. The conditional variances can depend on arbitrary covariates, thus adding considerable flexibility to the homoscedastic model. We present an Expectation-Conditional-Maximization (ECM) algorithm to efficiently compute the HAFT maximum likelihood estimator with right-censored data. The methodology is applied to the heavily censored data from a colon cancer clinical trial, for which a new type of highly stringent model residuals is proposed. Based on these, the HAFT model was found to eliminate most outliers from its homoscedastic counterpart.

stat.ME

A Convergence Diagnostic for Bayesian Clustering

In many applications of Bayesian clustering, posterior sampling on the discrete state space of cluster allocations is achieved via Markov chain Monte Carlo (MCMC) techniques. As it is typically challenging to design transition kernels to explore this state space efficiently, MCMC convergence diagnostics for clustering applications is especially important. For general MCMC problems, state-of-the-art convergence diagnostics involve comparisons across multiple chains. However, single-chain alternatives can be appealing for computationally intensive and slowly-mixing MCMC, as is typically the case for Bayesian clustering. Thus, we propose here a single-chain convergence diagnostic specifically tailored to discrete-space MCMC. Namely, we consider a Hotelling-type statistic on the highest probability states, and use regenerative sampling theory to derive its equilibrium distribution. By leveraging information from the unnormalized posterior, our diagnostic protects against seemingly convergent chains in which the relative frequency of visited states is incorrect. The methodology is illustrated with a Bayesian clustering analysis of genetic mutants of the flowering plant Arabidopsis thaliana.

stat.CO

Evidence that self-similar microrheology of highly entangled polymeric solutions scales robustly with, and is tunable by, polymer concentration

We report observations of a remarkable scaling behavior with respect to concentration in the passive microbead rheology of two highly entangled polymeric solutions, polyethylene oxide (PEO) and hyaluronic acid (HA). This behavior was reported previously [Hill et al., PLOS ONE (2014)] for human lung mucus, a complex biological hydrogel, motivating the current study for synthetic polymeric solutions PEO and HA. The strategy is to identify, and focus within, a wide range of lag times $τ$ for which passive micron diameter beads exhibit self-similar (fractional, power law) mean-squared-displacement (MSD) statistics. For lung mucus, PEO at three different molecular weights (Mw), and HA at one Mw, we find ensemble-averaged MSDs of the form ${\langle}Δr^{2}(τ){\rangle} = 4D_ατ^α$, all within a common band, [1/60 sec, 3 sec], of lag times $τ$. We employ the MSD power law parameters $(D_α,α)$ to classify each polymeric solution over a range of highly entangled concentrations. By the generalized Stokes-Einstein relation, power law MSD implies power law elastic $G'(ω)$ and viscous $G''(ω)$ moduli for frequencies $1/τ$, [0.33 sec$^{-1}$, 60 sec$^{-1}$]. A natural question surrounds the polymeric properties that dictate $D_α$ and $α$, e.g. polymer concentration c, Mw, and stiffness (persistence length). In [Hill et al., PLOS ONE (2014)], we showed the MSD exponent $α$ varies linearly, while the pre-factor $D_α$ varies exponentially, with concentration, i.e. the semi-log plot, $(log(D_α),α)(c)$ of the classifier data is collinear. Here we show the same result for three distinct Mw PEO and HA at a single Mw. Future studies are required to explore the generality of these results for polymeric solutions, and to understand this scaling behavior with polymer concentration.

cond-mat.soft

Robust and Efficient Parametric Spectral Estimation in Atomic Force Microscopy

An atomic force microscope (AFM) is capable of producing ultra-high resolution measurements of nanoscopic objects and forces. It is an indispensable tool for various scientific disciplines such as molecular engineering, solid-state physics, and cell biology. Prior to a given experiment, the AFM must be calibrated by fitting a spectral density model to baseline recordings. However, since AFM experiments typically collect large amounts of data, parameter estimation by maximum likelihood can be prohibitively expensive. Thus, practitioners routinely employ a much faster least-squares estimation method, at the cost of substantially reduced statistical efficiency. Additionally, AFM data is often contaminated by periodic electronic noise, to which parameter estimates are highly sensitive. This article proposes a two-stage estimator to address these issues. Preliminary parameter estimates are first obtained by a variance-stabilizing procedure, by which the simplicity of least-squares combines with the efficiency of maximum likelihood. A test for spectral periodicities then eliminates high-impact outliers, considerably and robustly protecting the second-stage estimator from the effects of electronic noise. Simulation and experimental results indicate that a two- to ten-fold reduction in mean squared error can be expected by applying our methodology.

stat.AP

Calibration of higher eigenmodes of cantilevers

A method is presented for calibrating the higher eigenmodes (resonance modes) of atomic force microscopy cantilevers that can be performed prior to any tip-sample interaction. The method leverages recent efforts in accurately calibrating the first eigenmode by providing the higher-mode stiffness as a ratio to the first mode stiffness. A one-time calibration routine must be performed for every cantilever type to determine the power-law relationship between stiffness and frequency, which is then stored for future use on similar cantilevers. Then, future calibrations only require a measurement of the ratio of resonance frequencies and the stiffness of the first mode. This method is verified through stiffness measurements using three independent approaches: interferometric measurement, AC approach-curve calibration, and finite element analysis simulation. Power-law values for calibrating higher-mode stiffnesses are reported for three popular multifrequency cantilevers. Once the higher-mode stiffnesses are known, the amplitude of each mode can also be calibrated from the thermal spectrum by application of the equipartition theorem.

cond-mat.mes-hall

Maximum Likelihood Estimation for Single Particle, Passive Microrheology Data with Drift

Volume limitations and low yield thresholds of biological fluids have led to widespread use of passive microparticle rheology. The mean-squared-displacement (MSD) statistics of bead position time series (bead paths) are either applied directly to determine the creep compliance [Xu et al (1998)] or transformed to determine dynamic storage and loss moduli [Mason & Weitz (1995)]. A prevalent hurdle arises when there is a non-diffusive experimental drift in the data. Commensurate with the magnitude of drift relative to diffusive mobility, quantified by a Péclet number, the MSD statistics are distorted, and thus the path data must be "corrected" for drift. The standard approach is to estimate and subtract the drift from particle paths, and then calculate MSD statistics. We present an alternative, parametric approach using maximum likelihood estimation that simultaneously fits drift and diffusive model parameters from the path data; the MSD statistics (and consequently the compliance and dynamic moduli) then follow directly from the best-fit model. We illustrate and compare both methods on simulated path data over a range of Péclet numbers, where exact answers are known. We choose fractional Brownian motion as the numerical model because it affords tunable, sub-diffusive MSD statistics consistent with typical 30 second long, experimental observations of microbeads in several biological fluids. Finally, we apply and compare both methods on data from human bronchial epithelial cell culture mucus.

cond-mat.soft