SearcharxivSearch

arXiv subjects

Yui Tomo

Publications and source records attributed to Yui Tomo.

7 recordsLinked to original sources

Estimating Consensus Epidemic Trajectories via a Constrained Power Fr\'echet Mean with Functional Registration

In infectious disease modeling during the early phase of a pandemic, SEIR-type compartmental models are standard tools, and they require epidemiological parameters as inputs. Because these parameters are subject to uncertainty, different research groups often report different epidemic curves, and obtaining a representative curve that captures the characteristics of these trajectories is important for decision-making. However, simple pointwise summaries can attenuate epidemic peaks under temporal misalignment and generally do not preserve the dynamical structure of the underlying compartmental models. To address these limitations, we propose a method for summarizing multiple solutions to SEIR-type compartmental models on a functional space by computing a constrained power Fr\'echet mean with temporal shift registration. In our method, we regard the pairs of exposed and infectious compartments as objects in a Hilbert space, and the consensus curve is defined as the solution to a constrained optimization problem. Differential equation constraints and population constraints are incorporated in the optimization to preserve a partially mechanistic interpretation regarding the infectious compartment. We develop an implementable block-optimization algorithm based on basis function expansion. In simulation studies based on early COVID-19 parameter estimates, the proposed method produced consensus curves with a single epidemic peak, whereas pointwise summaries exhibited attenuated or multiple peaks. We further applied the method to six literature-derived parameter sets from early COVID-19 studies and obtained a representative trajectory with interpretable epidemiological parameters. The proposed approach provides a generalized trajectory-summarization method that includes mean- and median-type estimators and preserves mechanistic interpretability.

stat.AP

Zero-Inflated Logistic Regression Models with Shared Design: Identifiability, Existence of Estimates, and a Relabeling Rule

The zero-inflated logistic regression model accommodates binary responses with excess zeros, which often arise from a latent mixture of susceptible and insusceptible subpopulations or asymmetric misclassification of the response. The model has two components: regression for the binary response and a latent binary indicator for the zero-inflation state. In applied settings, it is common to use the same design matrix for both components if there is no prior knowledge. However, this shared-design specification lacks guaranteed identifiability of the regression parameters, as established in prior works. This paper investigates the theoretical properties of the zero-inflated logistic regression model under the shared-design setting and computational methods for applications. First, to motivate the use of the zero-inflated model, we prove that ignoring the zero-inflation mechanism can lead to a sign flip in the pseudo-true coefficient value relative to the true value. We then establish sufficient conditions for the existence of the maximum likelihood estimate. As a main result, we establish that the model under the shared-design setting is identifiable up to exchange symmetry of the parameters for two components and that the expected log-likelihood has a unique maximizer on the resulting quotient space. The posterior bimodality is examined using a P\'olya-Gamma Gibbs sampler with replica exchange. Finally, we propose a simple relabeling rule to select a single ordered parameter pair, and evaluate its performance through simulation studies and an application to self-reported diabetes data.

stat.ME

Location--Scale Calibration for Generalized Posterior

General Bayesian updating replaces the likelihood with a loss scaled by a learning rate, but posterior uncertainty can depend sharply on that scale. We propose a simple post-processing that aligns generalized posterior draws with their asymptotic target, yielding uncertainty quantification that is invariant to the learning rate. We prove total-variation convergence for generalized posteriors with an effective sample size, allowing sample-size-dependent priors, non-i.i.d. observations, and convex penalties under model misspecification. Within this framework, we justify and extend the open-faced sandwich adjustment (Shaby, 2014), provide general theoretical guarantees for its use within generalized Bayes, and extend it from covariance rescaling to a location--scale calibration whose draws converge in total variation to the target for any learning rate. In our empirical illustration, calibrated draws maintain stable coverage, interval width, and bias over orders of magnitude in the learning rate and closely track frequentist benchmarks, whereas uncalibrated posteriors vary markedly.

stat.ME

Efficient Gibbs Sampling in Cox Regression Models Using Composite Partial Likelihood and P\'olya-Gamma Augmentation

The Cox regression models and their Bayesian extensions are widely used for time-to-event analysis. However, standard Bayesian approaches typically require baseline hazard modeling, and their full conditional distributions lack closed-form expressions, resulting in computational inefficiency and increased vulnerability to bias from baseline hazard misspecification. To address these issues, we propose GS4Cox, a fully Gibbs sampler for Bayesian Cox regression models with four elements: (i) generalized Bayesian framework for avoiding baseline hazard specification, (ii) composite partial likelihood and (iii) P\'olya-Gamma augmentation for closed-form expressions of full conditional distributions, and (iv) affine posterior calibration via the open-faced sandwich adjustment for location and scale adjustment of the posterior distribution. We prove asymptotic unbiasedness of the generalized Bayes estimator under composite partial likelihood and propose an affine posterior transformation that yields higher-order asymptotic agreement with the maximum partial likelihood estimator, while the posterior covariance matches the asymptotic target covariance. We demonstrated that GS4Cox consistently outperformed existing sampling methods through numerical and real-data experiments.

stat.ME

Median Consensus Embedding for Dimensionality Reduction

This study proposes median consensus embedding (MCE) to address variability in low-dimensional embeddings caused by random initialization in nonlinear dimensionality reduction techniques such as $t$-distributed stochastic neighbor embedding. MCE is defined as the geometric median of multiple embeddings. By assuming multiple embeddings as independent and identically distributed random samples and applying large deviation theory, we prove that MCE achieves consistency at an exponential rate. Furthermore, we develop a practical algorithm to implement MCE by constructing a distance function between embeddings based on the Frobenius norm of the pairwise distance matrix of data points. Application to actual data demonstrates that MCE converges rapidly and effectively reduces instability. We further combine MCE with multiple imputation to address missing values and consider multiscale hyperparameters. Results confirm that MCE effectively mitigates instability issues in embedding methods arising from random initialization and other sources.

stat.ML

A Note on Estimation Error Bound and Grouping Effect of Transfer Elastic Net

The Transfer Elastic Net is an estimation method for linear regression models that combines $\ell_1$ and $\ell_2$ norm penalties to facilitate knowledge transfer. In this study, we derive a non-asymptotic $\ell_2$ norm estimation error bound for the estimator and discuss scenarios where the Transfer Elastic Net effectively works. Furthermore, we examine situations where it exhibits the grouping effect, which states that the estimates corresponding to highly correlated predictors have a small difference.

stat.ML

Existence of Firth's modified estimates in binomial regression models

In logistic regression modeling, Firth's modified estimator is widely used to address the issue of data separation, which results in the nonexistence of the maximum likelihood estimate. Firth's modified estimator can be formulated as a penalized maximum likelihood estimator in which Jeffreys' prior is adopted as the penalty term. Despite its widespread use in practice, the formal verification of the corresponding estimate's existence has not been established. In this study, we establish the existence theorem of Firth's modified estimate in binomial logistic regression models, assuming only the full column rankness of the design matrix. We also discuss other binomial regression models obtained through alternating link functions and prove the existence of similar penalized maximum likelihood estimates for such models.

math.ST