SearcharxivSearch

arXiv subjects

Daisuke Yoneoka

Publications and source records attributed to Daisuke Yoneoka.

4 recordsLinked to original sources

Estimating Consensus Epidemic Trajectories via a Constrained Power Fr\'echet Mean with Functional Registration

In infectious disease modeling during the early phase of a pandemic, SEIR-type compartmental models are standard tools, and they require epidemiological parameters as inputs. Because these parameters are subject to uncertainty, different research groups often report different epidemic curves, and obtaining a representative curve that captures the characteristics of these trajectories is important for decision-making. However, simple pointwise summaries can attenuate epidemic peaks under temporal misalignment and generally do not preserve the dynamical structure of the underlying compartmental models. To address these limitations, we propose a method for summarizing multiple solutions to SEIR-type compartmental models on a functional space by computing a constrained power Fr\'echet mean with temporal shift registration. In our method, we regard the pairs of exposed and infectious compartments as objects in a Hilbert space, and the consensus curve is defined as the solution to a constrained optimization problem. Differential equation constraints and population constraints are incorporated in the optimization to preserve a partially mechanistic interpretation regarding the infectious compartment. We develop an implementable block-optimization algorithm based on basis function expansion. In simulation studies based on early COVID-19 parameter estimates, the proposed method produced consensus curves with a single epidemic peak, whereas pointwise summaries exhibited attenuated or multiple peaks. We further applied the method to six literature-derived parameter sets from early COVID-19 studies and obtained a representative trajectory with interpretable epidemiological parameters. The proposed approach provides a generalized trajectory-summarization method that includes mean- and median-type estimators and preserves mechanistic interpretability.

stat.AP

Zero-Inflated Logistic Regression Models with Shared Design: Identifiability, Existence of Estimates, and a Relabeling Rule

The zero-inflated logistic regression model accommodates binary responses with excess zeros, which often arise from a latent mixture of susceptible and insusceptible subpopulations or asymmetric misclassification of the response. The model has two components: regression for the binary response and a latent binary indicator for the zero-inflation state. In applied settings, it is common to use the same design matrix for both components if there is no prior knowledge. However, this shared-design specification lacks guaranteed identifiability of the regression parameters, as established in prior works. This paper investigates the theoretical properties of the zero-inflated logistic regression model under the shared-design setting and computational methods for applications. First, to motivate the use of the zero-inflated model, we prove that ignoring the zero-inflation mechanism can lead to a sign flip in the pseudo-true coefficient value relative to the true value. We then establish sufficient conditions for the existence of the maximum likelihood estimate. As a main result, we establish that the model under the shared-design setting is identifiable up to exchange symmetry of the parameters for two components and that the expected log-likelihood has a unique maximizer on the resulting quotient space. The posterior bimodality is examined using a P\'olya-Gamma Gibbs sampler with replica exchange. Finally, we propose a simple relabeling rule to select a single ordered parameter pair, and evaluate its performance through simulation studies and an application to self-reported diabetes data.

stat.ME

Median Consensus Embedding for Dimensionality Reduction

This study proposes median consensus embedding (MCE) to address variability in low-dimensional embeddings caused by random initialization in nonlinear dimensionality reduction techniques such as $t$-distributed stochastic neighbor embedding. MCE is defined as the geometric median of multiple embeddings. By assuming multiple embeddings as independent and identically distributed random samples and applying large deviation theory, we prove that MCE achieves consistency at an exponential rate. Furthermore, we develop a practical algorithm to implement MCE by constructing a distance function between embeddings based on the Frobenius norm of the pairwise distance matrix of data points. Application to actual data demonstrates that MCE converges rapidly and effectively reduces instability. We further combine MCE with multiple imputation to address missing values and consider multiscale hyperparameters. Results confirm that MCE effectively mitigates instability issues in embedding methods arising from random initialization and other sources.

stat.ML

New Insight of Spatial Scan Statistics via Regression Model

The spatial scan statistic is widely used to detect disease clusters in epidemiological surveillance. Since the seminal work by~\cite{kulldorff1997}, numerous extensions have emerged, including methods for defining scan regions, detecting multiple clusters, and expanding statistical models. Notably,~\cite{jung2009} and~\cite{ZHANG20092851} introduced a regression-based approach accounting for covariates, encompassing classical methods such as those of~\cite{kulldorff1997}. Another key extension is the expectation-based approach~\citep{neill2005anomalous,neillphdthesis}, which differs from the population-based approach represented by~\cite{kulldorff1997} in terms of hypothesis testing. In this paper, we bridge the regression-based approach with both expectation-based and population-based approaches. We reveal that the two approaches are separated by a simple difference: the presence or absence of an intercept term in the regression model. Exploiting the above simple difference, we propose new spatial scan statistics under the Gaussian and Bernoulli models. We further extend the regression-based approach by incorporating the well-known sparse L0 penalty and show that the derivation of spatial scan statistics can be expressed as an equivalent optimization problem. Our extended framework accommodates extensions such as space-time scan statistics and detecting multiple clusters while naturally connecting with existing spatial regression-based cluster detection. Considering the relation to case-specific models~\citep{she2011,10.1214/11-STS377}, clusters detected by spatial scan statistics can be viewed as outliers in terms of robust statistics. Numerical experiments with real data illustrate the behavior of our proposed statistics under specified settings.

stat.ME