SearcharxivSearch

arXiv subjects

Vu Khac Ky

Publications and source records attributed to Vu Khac Ky.

6 recordsLinked to original sources

Dictators are most informative

We prove the Courtade-Kumar conjecture: among all Boolean functions $f\colon \{-1,1\}^n\to\{-1,1\}$, a dictator retains the most information about a uniformly random input observed through independent binary noise.

cs.IT

Loss-Parameterized Fisher Width Along Learning Trajectories

Fisher width measures the Gaussian width of a probe set after deformation by the local Fisher geometry. We study its evolution along learning trajectories and ask when training loss can serve as an effective coordinate for this quantity. We first derive an exact trace--shape factorization and a deterministic stability bound for fixed compact probes. In a population Gaussian-teacher logistic model, the teacher-aligned state is extremal on every loss level below $\log 2$: it has minimal parameter norm and maximizes both Fisher trace and Euclidean-ball Fisher width. We then show that population gradient flow asymptotically selects this branch, with explicit rates for the aligned and orthogonal coordinates. This yields, for $d\geq2$, \[ \frac{w_F(B_2^d;θ(t))} {\sqrt{L(θ(t))}} \longrightarrow \frac{\sqrt6}π\mathbb E[χ_{d-1}]. \] Controlled full-Fisher experiments support the matched-loss branch and the population predictions. In a nonlinear MLP with a diagonal model-Fisher approximation, GD and SGD remain close at matched loss, whereas Adam follows a substantially displaced branch; the fixed probes tested retain highly similar temporal shapes. These results support a branchwise, rather than universal, loss parametrization of Fisher width.

cs.LG

Curvature Residual Geometry in Bregman Regression

We consider linear regression models fitted by minimizing Bregman losses of the form \[ \frac{1}{n}\sum_{i=1}^n \left[ ϕ(y_i)-ϕ(x_i^\topθ) -ϕ'(x_i^\topθ) \bigl(y_i-x_i^\topθ\bigr) \right], \] where \(y_i\) is the observed response and \(x_i^\topθ\) is the linear prediction. Even when the generating potential \(ϕ\) is strongly convex, the resulting regression objective may be nonconvex in \(θ\). This creates a gap between the convexity of the generating potential and the optimization geometry of the fitted model. The Hessian can be written as a weighted Gram matrix whose weights depend on the derivatives of the potential and the current residuals. This representation gives simple conditions for local strong convexity, smoothness, and conditional linear convergence of gradient descent. For the quadratic--quartic potential, we derive an exact scalar convexity condition, identify the interval of negative curvature, and obtain local and global sufficient conditions for positive curvature. Numerical experiments indicate that these scalar conditions may fail while the full Hessian remains positive definite at the evaluated points. They also indicate that the range of tested gradient-descent step sizes leading to convergence decreases as the quartic parameter grows. These results characterize how residual-dependent curvature interacts with the generating potential and the design matrix in reverse Bregman regression.

math.OC

Fisher Widths: Local Learning Geometry and Anisotropic Recovery

We study Gaussian-width complexity on statistical manifolds through a pair of functionals: the primal Fisher width $w_G(T) = w(G^{1/2}T)$, induced by the Fisher metric, and the inverse-Fisher width $w_{G^{-1}}(T) = w(G^{-1/2}T)$, induced by the inverse Fisher metric. The two widths play complementary statistical roles. On the learning side, the Fisher width measures the size of local parameter fluctuations in the geometry induced by the Fisher information. For Fisher-regular losses, we prove that the scale \(w_G(H_r)/\sqrt n\) is attained on sufficiently small Fisher balls. On the recovery side, the inverse-Fisher width captures the effect of anisotropic Gaussian measurements whose covariance is determined by the inverse Fisher information. For sparse recovery, the resulting geometry depends not only on sparsity but also on the position of the active coordinates in the Fisher spectrum. We obtain a two-sided estimate for the corresponding statistical dimension, together with support-sensitive recovery estimates and a natural ordering of supports with different curvature profiles. Finally, we establish a sharp relation between the primal and inverse-Fisher widths. On any common compact coordinate set $T$, they satisfy \[ w_G(T)w_{G^{-1}}(T)\geq w(T)^2. \] Thus, Fisher anisotropy may transfer complexity from one geometry to the other, but cannot reduce both widths relative to the Euclidean scale.

cs.LG

Fisher Width: A Geometric Measure of Complexity on Statistical Manifolds

Gaussian width is a central geometric complexity measure in high-dimensional probability, compressed sensing, convex optimization, and learning theory. It quantifies the average extent of a set along random directions, thereby capturing the effective dimension of constraint sets, hypothesis classes, and descent cones. However, this notion is intrinsically Euclidean. Statistical models instead carry a natural Riemannian geometry induced by the Fisher information metric, where directions are scaled according to statistical distinguishability rather than ambient Euclidean length. We introduce Fisher width, a Fisher-geometric analogue of Gaussian width for statistical manifolds. At a parameter point $θ$, Fisher width replaces the Euclidean identity by the local metric tensor $G(θ)^{1/2}$, measuring the Gaussian width of the Fisher-rescaled set. This makes the resulting quantity sensitive to local statistical curvature and invariant under smooth reparameterizations. We develop the basic theory of Fisher width, showing that it retains key structural features of Gaussian width, including concentration, metric perturbation stability, and spectral comparison bounds with the Euclidean baseline, while also capturing anisotropic geometric effects invisible to Euclidean measures. As an application, we prove a generalization bound for Fisher-Lipschitz hypothesis classes and propose computable estimators, which we evaluate empirically on MNIST across three model classes. Fisher width is to statistical manifolds what Gaussian width is to Euclidean convex bodies. This work lays the foundation for studying complexity and learning on curved statistical manifolds.

cs.LG