SearcharxivSearch

arXiv subjects

Chang He

Publications and source records attributed to Chang He.

At least 19 recordsLinked to original sources

Heavy-Ball Method under Randomized Schedules

We study how predefined randomized parameter schedules accelerate the heavy-ball method on general smooth convex objectives. Our analysis distinguishes two levels of randomization: sampling gradient-evaluation times within intervals whose boundaries are deterministic, and additionally randomizing the time boundaries themselves. With deterministic time boundaries, we construct fixed-time and anytime schedules that achieve an expected last-iterate function value gap of order $\mathcal{O}(1/K^{4/3})$; the anytime schedule also satisfies the same rate almost surely. We then use randomized time boundaries and obtain the improved last-iterate rate $\mathcal{O}(1/K^{3/2})$, both in expectation and almost surely. To the best of our knowledge, this is the first global nonasymptotic convergence guarantee for the heavy-ball method on general smooth convex functions that improves polynomially over the classical $\mathcal{O}(1/K)$ rate. Our result shows that the heavy-ball method achieves a strictly better convergence rate than the best known result $\mathcal{O}(1/K^{\log_2(1+\sqrt{2})})$ attainable by plain gradient descent with silver stepsize schedules.

math.OC

Classically Realizable Incompatibility

Incompatibility constitutes a fundamental aspect of quantum mechanics. However, not every quantum observable non-classical property arises from incompatibility, nor can all quantum scenarios be fully captured by incompatibility alone. Within the framework of partial Boolean algebra (pBA), we research the structural properties of incompatibility scenarios. We introduce a unified method to realize any incompatibility scenario via a classical game, and the construction is extendable to any scenario embeddable into a Boolean algebra. The exclusivity graph offers a precise characterization of incompatibility scenarios. We prove that every exclusivity graph is the atom graph of an exclusive pBA, which is embedded into a Boolean algebra. These results provide a necessary condition for exclusivity graphs and a sufficient condition for atom graphs.

quant-ph

On the Nature of Regularity Assumptions in Bilevel Optimization with Constrained Lower-level Problem

In this paper, we study the regularity assumptions commonly adopted in bilevel optimization with constrained lower-level problems, including the linear independence constraint qualification, the strict complementary slackness condition, and the second-order sufficient condition. These conditions are typically required to hold for the lower-level problem at every upper-level variable $x$. We first show that the requirement that these conditions hold at every upper-level variable $x$ is strong, in the sense that it is non-prevalent: there exist problems for which no sufficiently small perturbation of the lower-level objective and constraints can make the conditions hold at every $x$. To establish the result, we prove rigidity theorems showing that certain structural quantities of the lower-level problem must remain invariant across all $x$ whenever these conditions hold everywhere. We then construct explicit counterexamples in which these invariants differ between two values of $x$. In contrast, we show that the weaker requirement, that these conditions hold at almost every $x$, is a weak assumption, in the sense that it is prevalent: with probability one over a random perturbation of the lower-level objective and constraints, each condition holds at almost every $x$. We further analyze the gap between the two requirements. Although the ``every $x$'' and ``almost every $x$'' versions differ only on a measure-zero set, we show that this difference introduces fundamental difficulties in both theory and computation for bilevel optimization.

math.OC

Hardy's Paradox for Yu-Oh Set Constructed by Logically Contextual Quantum States

Quantum contextuality is a fundamental nonclassical property of quantum systems, regarded as a key resource that demonstrates the computational and informational advantages of quantum over classical systems. Our present work aims to construct Hardy's paradoxes, a set of possibilistic conditions witnessing contextuality, for Yu-Oh set, which is the state-independent contextual quantum system with the least number of vectors. To achieve the aim, we systematically enumerate all logically contextual pure states on Yu-Oh set, and theoretically prove that no mixed states in this scenario are logically contextual. Based on the identified logically contextual quantum states, we construct 12 Hardy's paradoxes with identical success probability SP=11.1%. Furthermore, we present corresponding observables to experimentally witness these Hardy's paradoxes.

quant-ph

Adaptive Single-Loop Methods for Stochastic Minimax Optimization on Riemannian Manifolds

Stochastic minimax optimization on Riemannian manifolds has recently attracted significant attention due to its broad range of applications, such as robust training of neural networks and robust maximum likelihood estimation. Existing optimization methods for these problems typically require selecting stepsizes based on prior knowledge of specific problem parameters, such as Lipschitz-type constants and (geodesic) strong concavity constants. Unfortunately, these parameters are often unknown in practice. To overcome this issue, we develop single-loop adaptive methods that automatically adjust stepsizes using cumulative Riemannian (stochastic) gradient norms. We first propose a deterministic single-loop Riemannian adaptive gradient descent ascent method and show that it attains an $\epsilon$-stationary point within $O(\epsilon^{-2})$ iterations. This deterministic method is of independent interest and lays the foundation for our subsequent stochastic method. In particular, we propose the Riemannian stochastic adaptive gradient descent ascent method, which finds an $\epsilon$-stationary point in $O(\epsilon^{-6})$ iterations. Under additional second-order smoothness, this iteration complexity is further improved to $O(\epsilon^{-4})$, which even outperforms the corresponding complexity result in Euclidean space. Some numerical experiments on real-world applications are conducted, including the regularized robust maximum likelihood estimation problem, and the robust training of neural networks with orthonormal weights. The results are encouraging and demonstrate the effectiveness of adaptivity in practice.

math.OC

Small Gradient Norm Regret for Online Convex Optimization

This paper introduces a new problem-dependent regret measure for online convex optimization with smooth losses. The notion, which we call the $G^\star$ regret, depends on the cumulative squared gradient norm evaluated at the decision in hindsight. We show that the $G^\star$ regret strictly refines the existing $L^\star$ (small loss) regret, and that it can be arbitrarily sharper when the losses have vanishing curvature around the hindsight decision. We establish upper and lower bounds on the $G^\star$ regret and extend our results to dynamic regret and bandit settings. As a byproduct, we refine the existing convergence analysis of stochastic optimization algorithms in the interpolation regime. Some experiments validate our theoretical findings.

stat.ML

A Logical Formalism of Hardy-type Paradox

Hardy-type paradoxes provide elegant, inequality-free proofs of quantum contextuality. We introduce a unified logical formalism for these paradoxes, termed logical Hardy-type paradoxes. For any finite quantum scenario of ideal measurements, we prove that the existence of a logical Hardy-type paradox is equivalent to logical contextuality. Specifically, strong contextuality is equivalent to logical Hardy-type paradoxes with success probability SP = 1. These results generalize prior work on (2,k,2), (2,2,d), and n-cycle scenarios. We analyze logical Hardy-type paradoxes in the Mansfield and Klyachko-Can-Binicioglu-Shumovsky (KCBS) scenarios. In the KCBS scenario, we show that there is exactly one type of logical Hardy-type paradox, achieving SP\approx 10.56% for a specific parameter setting.

quant-ph

New Results on the Polyak Stepsize: Tight Convergence Analysis and Universal Function Classes

In this paper, we revisit a classical adaptive stepsize strategy for gradient descent: the Polyak stepsize (PolyakGD), originally proposed in Polyak (1969). We study the convergence behavior of PolyakGD from two perspectives: tight worst-case analysis and universality across function classes. As our first main result, we establish the tightness of the known convergence rates of PolyakGD by explicitly constructing worst-case functions. In particular, we show that the $O((1-\frac{1}{\kappa})^K)$ rate for smooth strongly convex functions and the $O(1/K)$ rate for smooth convex functions are both tight. Moreover, we theoretically show that PolyakGD automatically exploits floating-point errors to escape the worst-case behavior. Our second main result provides new convergence guarantees for PolyakGD under both H\"older smoothness and H\"older growth conditions. These findings show that the Polyak stepsize is universal, automatically adapting to various function classes without requiring prior knowledge of problem parameters.

math.OC

On Approximation Algorithms for Commutative Quaternion Polynomial Optimization

Quaternion optimization has attracted significant interest due to its broad applications, including color face recognition, video compression, and signal processing. Despite the growing literature on quadratic and matrix quaternion optimization, to the best of our knowledge, the study on quaternion polynomial optimization still remains blank. In this paper, we introduce the first investigation into this fundamental problem, and focus on the sphere-constrained homogeneous polynomial optimization over the commutative quaternion domain, which includes the best rank-one tensor approximation as a special case. Our study proposes a polynomial-time randomized approximation algorithm that employs tensor relaxation and random sampling techniques to tackle this problem. Theoretically, we prove an approximation ratio for the algorithm providing a worst-case performance guarantee

math.OC

History-Aware Adaptive High-Order Tensor Regularization

In this paper, we develop a new adaptive regularization method for minimizing a composite function, which is the sum of a $p$th-order ($p \ge 1$) Lipschitz continuous function and a simple, convex, and possibly nonsmooth function. We use a history of local Lipschitz estimates to adaptively select the current regularization parameter, an approach we shall term the {\it history-aware adaptive regularization method}. We explore how the selection of an appropriate volume of historical information affects both the theoretical and practical performance. By using all the historical information, our method matches the complexity guarantees of the standard $p$th-order tensor methods that require a known Lipschitz constant, for both convex and nonconvex objectives. In the nonconvex case, the number of iterations required to find an $(\epsilon_g,\epsilon_H)$-approximate second-order stationary point is bounded by $\mathcal{O}(\max\{\epsilon_g^{-(p+1)/p}, \epsilon_H^{-(p+1)/(p-1)}\})$. For convex functions, we establish an $\mathcal{O}(\epsilon^{-1/p})$ iteration complexity for finding an $\epsilon$-approximate optimal point and further propose an accelerated variant attaining an iteration complexity of $\mathcal{O}(\epsilon^{-1/(p+1)})$. For practical consideration, we propose several variants of this method with only part of historical information. We introduce cyclic and sliding-window strategies for choosing historical Lipschitz estimates, which mitigate the limitation of overly conservative updates. As long as a rough upper bound of the Lipschitz constant is known, these two variants achieve the same iteration complexity guarantees in terms of the input accuracy as the method using full historical information. Finally, extensive numerical experiments are conducted to demonstrate the effectiveness of our adaptive approach.

math.OC

Non-Stationary Bandit Convex Optimization: An Optimal Algorithm with Two-Point Feedback

This paper studies bandit convex optimization in non-stationary environments with two-point feedback, using dynamic regret as the performance measure. We propose an algorithm based on bandit mirror descent that extends naturally to non-Euclidean settings. Let $T$ be the total number of iterations and $\mathcal{P}_{T,p}$ the path variation with respect to the $\ell_p$-norm. In Euclidean space, our algorithm matches the optimal regret bound $\mathcal{O}(\sqrt{dT(1+\mathcal{P}_{T,2})})$, improving upon \citet{zhao2021bandit} by a factor of $\mathcal{O}(\sqrt{d})$. Beyond Euclidean settings, our algorithm achieves an upper bound of $\mathcal{O}(\sqrt{d\log(d)T\log(T)(1 + \mathcal{P}_{T,1})})$ on the simplex, which is nearly optimal up to log factors. For the cross-polytope, the bound reduces to $\mathcal{O}(\sqrt{d\log(d)T(1+\mathcal{P}_{T,p})})$ for some $p = 1 + 1/\log(d)$.

math.OC

The Second-Order T\^atonnement: Decentralized Interior-Point Methods for Market Equilibrium

The t\^atonnement process and Smale's process are two classical approaches to compute market equilibrium in exchange economies. While the t\^atonnement process can be seen as a first-order method, Smale's process, being second-order, is less popular due to its reliance on additional information from the players and expensive Newton steps. In this paper, we study Fisher exchange market for a broad class of utility functions, where we show that all high-order information required by Smale's process is readily available from players' best responses. Motivated by this observation, we develop two second-order t\^atonnement processes, constructed as decentralized interior-point methods, which are traditionally known to work in a centralized manner. The methods here bear the name "t\^atonnement", since, in spirit, they demand no more information than the classical t\^atonnement process. To address the Newton systems involved, we introduce an explicitly invertible approximation with high-probability guarantees and a scaling matrix that optimally minimizes the condition number, both of which rely solely on best responses as the methods themselves. Using these tools, the first second-order t\^atonnement process has O(log(1/$\epsilon$))complexity rate. Under mild conditions, the other method achieves a non-asymptotic superlinear convergence rate. Preliminary experiments are presented to justify the capability of the proposed methods for large-scale problems. Extensions of our approach are also discussed.

math.OC

On Relatively Smooth Optimization over Riemannian Manifolds

We study optimization over Riemannian embedded submanifolds, where the objective function is relatively smooth in the ambient Euclidean space. Such problems have broad applications but are still largely unexplored. We introduce two Riemannian first-order methods, namely the retraction-based and projection-based Riemannian Bregman gradient methods, by incorporating the Bregman distance into the update steps. The retraction-based method can handle nonsmooth optimization; at each iteration, the update direction is generated by solving a convex optimization subproblem constrained to the tangent space. We show that when the reference function is of the quartic form $h(x) = \frac{1}{4}\|x\|^4 + \frac{1}{2}\|x\|^2$, the constraint subproblem admits a closed-form solution. The projection-based approach can be applied to smooth Riemannian optimization, which solves an unconstrained subproblem in the ambient Euclidean space. Both methods are shown to achieve an iteration complexity of $\mathcal{O}(1/\epsilon^2)$ for finding an $\epsilon$-approximate Riemannian stationary point. When the manifold is compact, we further develop stochastic variants and establish a sample complexity of $\mathcal{O}(1/\epsilon^4)$. Numerical experiments on the nonlinear eigenvalue problem and low-rank quadratic sensing problem demonstrate the advantages of the proposed methods.

math.OC

Federated Learning on Riemannian Manifolds: A Gradient-Free Projection-Based Approach

Federated learning (FL) has emerged as a powerful paradigm for collaborative model training across distributed clients while preserving data privacy. However, existing FL algorithms predominantly focus on unconstrained optimization problems with exact gradient information, limiting its applicability in scenarios where only noisy function evaluations are accessible or where model parameters are constrained. To address these challenges, we propose a novel zeroth-order projection-based algorithm on Riemannian manifolds for FL. By leveraging the projection operator, we introduce a computationally efficient zeroth-order Riemannian gradient estimator. Unlike existing estimators, ours requires only a simple Euclidean random perturbation, eliminating the need to sample random vectors in the tangent space, thus reducing computational cost. Theoretically, we first prove the approximation properties of the estimator and then establish the sublinear convergence of the proposed algorithm, matching the rate of its first-order counterpart. Numerically, we first assess the efficiency of our estimator using kernel principal component analysis. Furthermore, we apply the proposed algorithm to two real-world scenarios: zeroth-order attacks on deep neural networks and low-rank neural network training to validate the theoretical findings.

math.OC

The logical structure of contextuality and nonclassicality

Quantum contextuality represents a fundamental form of nonclassicality in quantum mechanics. To provide a more complete characterization of nonclassical properties in quantum systems, we adopt a logical perspective and propose a mathematical framework based on exclusive partial Boolean algebras (epBAs). This framework enables a unified description of contextuality and nonclassicality across finite general, quantum, and classical systems. We establish a unified and minimal classical counterpart for any finite general system. Within this framework, we formalize major categories of quantum contextuality, demonstrating that: 12 projectors suffice to generate Kochen-Specker scenarios; 10 projectors suffice to witness state-independent contextuality; and 3 observables suffice to witness quantum contextuality. Finally, we prove that contextuality is a sufficient but not necessary condition for nonclassicality.

quant-ph

Optimized Cryo-CMOS Technology with VTH<0.2V and Ion>1.2mA/um for High-Peformance Computing

We report the design-technology co-optimization (DTCO) scheme to develop a 28-nm cryogenic CMOS (Cryo-CMOS) technology for high-performance computing (HPC). The precise adjustment of halo implants manages to compensate the threshold voltage (VTH) shift at low temperatures. The optimized NMOS and PMOS transistors, featured by VTH<0.2V, sub-threshold swing (SS)<30 mV/dec, and on-state current (Ion)>1.2mA/um at 77K, warrant a reliable sub-0.6V operation. Moreover, the enhanced driving strength of Cryo-CMOS inherited from a higher transconductance leads to marked improvements in elevating the ring oscillator frequency by 20%, while reducing the power consumption of the compute-intensive cryogenic IC system by 37% at 77K.

eess.SY

Modeling and Simulation of 2D Transducers Based on Suspended Graphene-Based Heterostructures in Nanoelectromechanical Pressure Sensors

Graphene-based 2D heterostructures exhibit excellent mechanical and electrical properties, which are expected to exhibit better performances than graphene for nanoelectromechanical pressure sensors. Here, we built the pressure sensor models based on suspended heterostructures of graphene/h-BN, graphene/MoS2, and graphene/MoSe2 by using COMSOL Multiphysics finite element software. We found that suspended circular 2D membranes show the best sensitivity to pressures compared to rectangular and square ones. We simulated the deflections, strains, resonant frequencies, and Young's moduli of suspended graphene-based heterostructures under the conditions of different applied pressures and geometrical sizes, built-in tensions, and the number of atomic layers of 2D membranes. The Young's moduli of 2D heterostructures of graphene, graphene/h-BN, graphene/MoS2, and graphene/MoSe2 were estimated to be 1.001TPa, 921.08 GPa, 551.11 GPa, and 475.68 GPa, respectively. We also discuss the effect of highly asymmetric cavities on device performance. These results would contribute to the understanding of the mechanical properties of graphene-based heterostructures and would be helpful for the design and manufacture of high-performance NEMS pressure sensors.

cond-mat.mes-hall

Graphene MEMS and NEMS

Graphene is being increasingly used as an interesting transducer membrane in micro- and nanoelectromechanical systems (MEMS and NEMS, respectively) due to its atomical thickness, extremely high carrier mobility, high mechanical strength and piezoresistive electromechanical transductions. NEMS devices based on graphene feature increased sensitivity, reduced size, and new functionalities. In this review, we discuss the merits of graphene as a functional material for MEMS and NEMS, the related properties of graphene, the transduction mechanisms of graphene MEMS and NEMS, typical transfer methods for integrating graphene with MEMS substrates, methods for fabricating suspended graphene, and graphene patterning and electrical contact. Consequently, we provide an overview of devices based on suspended and nonsuspended graphene structures. Finally, we discuss the potential and challenges of applications of graphene in MEMS and NEMS. Owing to its unique features, graphene is a promising material for emerging MEMS, NEMS and sensor applications.

cond-mat.mes-hall