SearcharxivSearch

arXiv subjects

Chanwoong Park

Publications and source records attributed to Chanwoong Park.

3 recordsLinked to original sources

Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods

Recent research has indicated a substantial rise in interest in understanding Nesterov's accelerated gradient methods via their continuous-time models. However, most existing studies focus on specific classes of Nesterov's methods, which hinders the attainment of an in-depth understanding and a unified perspective. To address this deficit, we present generalized continuous-time models that cover a broad range of Nesterov's methods, including those previously studied under existing continuous-time frameworks. Our key contributions are as follows. First, we identify the convergence rates of the generalized models, eliminating the need to determine the convergence rate for any specific continuous-time model derived from them. Second, we show that six existing continuous-time models are special cases of our generalized models, thereby positioning our framework as a unifying tool for analyzing and understanding these models. Third, we design a restart scheme for Nesterov's methods based on our generalized models and show that it ensures a monotonic decrease in objective function values. Owing to the broad applicability of our models, this scheme can be used to a broader class of Nesterov's methods compared to the original restart scheme. Fourth, we uncover a connection between our generalized models and gradient flow in continuous time, showing that the accelerated convergence rates of our generalized models can be attributed to a time reparametrization in gradient flow. Numerical experiment results are provided to support our theoretical analyses and results.

math.OC

Sharpness-Aware Minimization Can Hallucinate Minimizers

Sharpness-Aware Minimization (SAM) is widely used to seek flatter minima -- often linked to better generalization. In its standard implementation, SAM updates the current iterate using the loss gradient evaluated at a point perturbed by distance $ρ$ along the normalized gradient direction. We show that, for some choices of $ρ$, SAM can stall at points where this shifted (perturbed-point) gradient vanishes despite a nonzero original gradient, and therefore, they are not stationary points of the original loss. We call these points hallucinated minimizers, prove their existence under simple nonconvex landscape conditions (e.g., the presence of a local minimizer and a local maximizer), and establish sufficient conditions for local convergence of the SAM iterates to them. We corroborate this failure mode in neural network training and observe that it aligns with SAM's performance degradation often seen at large $ρ$. Finally, as a practical safeguard, we find that a short initial SGD warm-start before enabling SAM mitigates this failure mode and reduces sensitivity to the choice of $ρ$.

cs.LG

Scalably manufactured high-index atomic layer-polymer hybrid metasurfaces for high-efficiency virtual reality metaoptics in the visible

Metalenses, which exhibit superior light-modulating performance with sub-micrometer-scale thicknesses, are suitable alternatives to conventional bulky refractive lenses. However, fabrication limitations, such as a high cost, low throughput, and small patterning area, hinder their mass production. Here, we demonstrate the mass production of low-cost, high-throughput, and large-aperture visible metalenses using an argon fluoride immersion scanner and wafer-scale nanoimprint lithography. Once a 12-inch master stamp is imprinted, hundreds of centimeter-scale metalenses can be fabricated. To enhance light confinement, the printed metasurface is thinly coated with a high-index film, resulting in drastic increase of conversion efficiency. As a proof of concept, a prototype of a virtual reality device with ultralow thickness is demonstrated with the fabricated metalens.

physics.optics