SearcharxivSearch

arXiv subjects

Boram Cho

Publications and source records attributed to Boram Cho.

3 recordsLinked to original sources

HOMER: Huber-of-Means for Efficient and Robust Estimation in Hilbert Spaces

Heavy tails weaken high-confidence control for the empirical mean. Geometric median-of-means (MOM) also lacks a threshold that moves toward mean efficiency. We propose \emph{HOMER}, or Huber-of-Means for Efficient and Robust Estimation. HOMER aggregates block means through a radial Huber center. Its canonical and pseudo-Huber forms bound each block score and interpolate between median-like robustness and the empirical mean. We establish a Hilbert-space majority theorem and a MOM-order deviation bound under a finite second moment. Canonical HOMER recovers the sample mean inside its quadratic region. Pseudo-HOMER approaches the sample mean as the threshold grows. It also admits asymptotic linearity and consistent sandwich covariance estimation around the population block-Huber target. Under a finite third moment, fixed finite-dimensional projections support mean inference at the usual parametric rate. This result requires growing block sizes and counts, with block sizes increasing faster. Heavy-tailed simulations show that HOMER remains stable when a minority of block summaries is displaced. On clean Gaussian data, both versions closely approach the empirical mean's efficiency. Finite-block sandwich intervals undercovered, especially for skewed functional data. Further studies show failure when contamination affects most blocks or compromises ordinary within-block means.

stat.ML

Heat-Kernel Entropy Profiles and Geometric Effective Sample Size for Weighted Measures on Manifolds

Weighted empirical measures on compact manifolds appear in importance sampling, particle approximations, posterior summaries, quadrature, and representation learning. Ordinary effective sample size and related weight summaries ignore the geometry of the support. We introduce heat-kernel entropy profiles to measure nonuniformity after intrinsic diffusion at a range of scales. For order-two R\'enyi entropy, pairwise heat-kernel overlaps give an exact profile and a geometric effective sample size. This effective sample size discounts nearby or duplicate particles. It approaches ordinary effective sample size as overlaps between distinct particles vanish. On compact boundaryless manifolds, we establish profile monotonicity, gESS scale limits, deterministic-weight consistency, and a bounded-ratio result for self-normalized importance sampling. On spheres, the unlogged profile decomposes into spherical-harmonic energies. The first terms are squared mean-resultant and traceless-second-moment energies, which give vMF- and Bingham-type scalar summaries. Experiments identify antipodal, girdle, multimodal, and duplicate-particle structures that weight-only and first-moment summaries miss.

stat.ML

Geometric Information Decomposition for Weighted Empirical Measures on the Sphere

Weighted observations on the unit sphere arise in importance sampling, quadrature, and attention-weighted embeddings. Directional uncertainty is often summarized through a von Mises-Fisher (vMF) fit and its concentration or entropy. This summary uses only mean-direction information. It can miss antipodal, axial, girdle-like, or multimodal structure. We introduce geometric information decomposition (GID), which fits nested maximum-entropy projections to spherical features. Each gap measures the entropy reduction contributed by one feature level. The first gap is the fitted vMF distribution's KL divergence from uniformity. The second measures residual quadratic information, including Fisher-Bingham anisotropy. Later gaps describe finer angular structure. We establish invariance, consistency, alternative-regime asymptotic normality, and quadratic-form null calibration. Circular and spherical experiments include importance-weight calibration and a query-weighted digit projection. The results separate settings where vMF uncertainty is adequate from settings with higher-order structure.

stat.ME