SearcharxivSearch

arXiv subjects

Ziyan Luo

Publications and source records attributed to Ziyan Luo.

At least 19 recordsLinked to original sources

Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning

State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide range of behavioral structures have been studied in reinforcement learning, including value functions, invariants, bisimulation relations, and behavioral metrics. However, a general principle for determining what structures are provably preserved under state abstraction is still lacking. In this paper, we present a unified framework for defining and analyzing behavioral structures in reinforcement learning. Our framework provides a compositional way to specify behavioral semantics based on local, one-step descriptions of system dynamics. Using this framework, we establish results showing how behavioral structures can be safely transferred between abstract and concrete systems. We further show how to construct quantitative metrics from logical behavioral semantics with soundness guarantees. Together, these results provide a principled foundation for reasoning about behaviors under state abstraction in reinforcement learning and offer reusable definition and proof principles for a broad class of behavioral structures in reinforcement learning.

cs.LG

Low Rank Support Quaternion Matrix Machine

Input features are conventionally represented as vectors, matrices, or third order tensors in the real field, for color image classification. Inspired by the success of quaternion data modeling for color images in image recovery and denoising tasks, we propose a novel classification method for color image classification, named as the Low-rank Support Quaternion Matrix Machine (LSQMM), in which the RGB channels are treated as pure quaternions to effectively preserve the intrinsic coupling relationships among channels via the quaternion algebra. For the purpose of promoting low-rank structures resulting from strongly correlated color channels, a quaternion nuclear norm regularization term, serving as a natural extension of the conventional matrix nuclear norm to the quaternion domain, is added to the hinge loss in our LSQMM model. An Alternating Direction Method of Multipliers (ADMM)-based iterative algorithm is designed to effectively resolve the proposed quaternion optimization model. Experimental results on multiple color image classification datasets demonstrate that our proposed classification approach exhibits advantages in classification accuracy, robustness and computational efficiency, compared to several state-of-the-art methods using support vector machines, support matrix machines, and support tensor machines.

cs.CV

Sharp-Peak Functions for Exactly Penalizing Binary Integer Programming

Unconstrained binary integer programming (UBIP) is a challenging optimization problem due to the presence of binary variables. To address the challenge, we introduce a novel class of functions named sharp-peak functions (SPFs), which equivalently reformulate the binary constraints as equality constraints, giving rise to an SPF-constrained optimization. Rather than solving this constrained reformulation directly, we focus on its associated penalty model. The established exact penalty theory shows that the global minimizers of UBIP and the penalty model coincide when the penalty parameter exceeds a threshold, a constant independent of the solution set of UBIP. To analyze the penalty model, we introduce Karush-Kuhn-Tucker (KKT) points and a new type of stationarity, referred to as P-stationarity, and provide a comprehensive characterization of its optimality conditions. We then develop an efficient algorithm called Sha-Peak based on the inexact alternating direction method of multipliers. It converges toa P-stationary point at a linear rate or terminates at such a point within finitely many steps. These results are established under appropriate parameter choices and a single mild assumption, namely, the local Lipschitz continuity of the gradient over a bounded box. Finally, numerical experiments demonstrate its nice performance in comparison to several established solvers.

math.OC

Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning

Developing agents capable of exploring, planning and learning in complex open-ended environments is a grand challenge in artificial intelligence (AI). Hierarchical reinforcement learning (HRL) offers a promising solution to this challenge by discovering and exploiting the temporal structure within a stream of experience. The strong appeal of the HRL framework has led to a rich and diverse body of literature attempting to discover a useful structure. However, it is still not clear how one might define what constitutes good structure in the first place, or the kind of problems in which identifying it may be helpful. This work aims to identify the benefits of HRL from the perspective of the fundamental challenges in decision-making, as well as highlight its impact on the performance trade-offs of AI agents. Through these benefits, we then cover the families of methods that discover temporal structure in HRL, ranging from learning directly from online experience to offline datasets, to leveraging large language models (LLMs). Finally, we highlight the challenges of temporal structure discovery and the domains that are particularly well-suited for such endeavours.

cs.AI

Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments

A key approach to state abstraction is approximating behavioral metrics (notably, bisimulation metrics) in the observation space and embedding these learned distances in the representation space. While promising for robustness to task-irrelevant noise, as shown in prior work, accurately estimating these metrics remains challenging, requiring various design choices that create gaps between theory and practice. Prior evaluations focus mainly on final returns, leaving the quality of learned metrics and the source of performance gains unclear. To systematically assess how metric learning works in deep reinforcement learning (RL), we evaluate five recent approaches, unified conceptually as isometric embeddings with varying design choices. We benchmark them with baselines across 20 state-based and 14 pixel-based tasks, spanning 370 task configurations with diverse noise settings. Beyond final returns, we introduce the evaluation of a denoising factor to quantify the encoder's ability to filter distractions. To further isolate the effect of metric learning, we propose and evaluate an isolated metric estimation setting, in which the encoder is influenced solely by the metric loss. Finally, we release an open-source, modular codebase to improve reproducibility and support future research on metric learning in deep RL.

cs.LG

Sparse Quadratically Constrained Quadratic Programming via Semismooth Newton Method

Quadratically constrained quadratic programming (QCQP) has long been recognized as a computationally challenging problem, particularly in large-scale or high-dimensional settings where solving it directly becomes intractable. The complexity further escalates when a sparsity constraint is involved, giving rise to the problem of sparse QCQP (SQCQP), which makes conventional solution methods even less effective. Existing approaches for solving SQCQP typically rely on mixed-integer programming formulations, relaxation techniques, or greedy heuristics but often suffer from computational inefficiency and limited accuracy. In this work, we introduce a novel paradigm by designing an efficient algorithm that directly addresses SQCQP. To be more specific, we introduce P-stationarity to establish first- and second-order optimality conditions of the original problem, leading to a system of nonlinear equations whose generalized Jacobian is proven to be nonsingular under mild assumptions. Most importantly, these equations facilitate the development of a semismooth Newton-type method that exhibits significantly low computational complexity due to the sparsity constraint and achieves a locally quadratic convergence rate. Finally, extensive numerical experiments validate the accuracy and computational efficiency of the algorithm compared to several established solvers.

math.OC

Preconditioned Inexact Stochastic ADMM for Deep Model

Deep learning models are usually trained with stochastic gradient descent-based algorithms, but these optimizers face inherent limitations, such as slow convergence and stringent assumptions for convergence. In particular, data heterogeneity arising from distributed settings poses significant challenges to their theoretical and numerical performance. This paper develops an algorithm, PISA (Preconditioned Inexact Stochastic Alternating Direction Method of Multipliers). Grounded in rigorous theoretical guarantees, the algorithm converges under the sole assumption of Lipschitz continuity of the gradient on a bounded region, thereby removing the need for other conditions commonly imposed by stochastic methods. This capability enables the proposed algorithm to tackle the challenge of data heterogeneity effectively. Moreover, the algorithmic architecture enables scalable parallel computing and supports various preconditions, such as second-order information, second moment, and orthogonalized momentum by Newton-Schulz iterations. Incorporating the latter two preconditions in PISA yields two computationally efficient variants: SISA and NSISA. Comprehensive experimental evaluations for training or fine-tuning diverse deep models, including vision models, large language models, reinforcement learning models, generative adversarial networks, and recurrent neural networks, demonstrate superior numerical performance of SISA and NSISA compared to various state-of-the-art optimizers.

cs.LG

Triangular Decomposition of Third Order Hermitian Tensors

We define lower triangular tensors, and show that all diagonal entries of such a tensor are eigenvalues of that tensor. We then define lower triangular sub-symmetric tensors, and show that the number of independent entries of a lower triangular sub-symmetric tensor is the same as that of a symmetric tensor of the same order and dimension. We further introduce third order Hermitian tensors, third order positive semi-definite Hermitian tensors, and third order positive semi-definite symmetric tensors. Third order completely positive tensors are positive semi-definite symmetric tensors. Then we show that a third order positive semi-definite Hermitian tensor is triangularly decomposable. This generalizes the classical result of Cholesky decomposition in matrix analysis.

math.RA

Chronosymbolic Learning: Efficient CHC Solving with Symbolic Reasoning and Inductive Learning

Solving Constrained Horn Clauses (CHCs) is a fundamental challenge behind a wide range of verification and analysis tasks. Data-driven approaches show great promise in improving CHC solving without the painstaking manual effort of creating and tuning various heuristics. However, a large performance gap exists between data-driven CHC solvers and symbolic reasoning-based solvers. In this work, we develop a simple but effective framework, "Chronosymbolic Learning", which unifies symbolic information and numerical data points to solve a CHC system efficiently. We also present a simple instance of Chronosymbolic Learning with a data-driven learner and a BMC-styled reasoner. Despite its relative simplicity, experimental results show the efficacy and robustness of our tool. It outperforms state-of-the-art CHC solvers on a dataset consisting of 288 benchmarks, including many instances with non-linear integer arithmetics.

cs.LO

Group SLOPE Penalized Low-Rank Tensor Regression

This article aims to seek a selection and estimation procedure for a class of tensor regression problems with multivariate covariates and matrix responses, which can provide theoretical guarantees for model selection in finite samples. Considering the frontal slice sparsity and low-rankness inherited in the coefficient tensor, we formulate the regression procedure as a group SLOPE penalized low-rank tensor optimization problem based on an orthogonal decomposition, namely TgSLOPE. This procedure provably controls the newly introduced tensor group false discovery rate (TgFDR), provided that the predictor matrix is column-orthogonal. Moreover, we establish the asymptotically minimax convergence with respect to the TgSLOPE estimate risk. For efficient problem resolution, we equivalently transform the TgSLOPE problem into a difference-of-convex (DC) program with the level-coercive objective function. This allows us to solve the reformulation problem of TgSLOPE by an efficient proximal DC algorithm (DCA) with global convergence. Numerical studies conducted on synthetic data and a real human brain connection data illustrate the efficacy of the proposed TgSLOPE estimation procedure.

math.ST

Variable T-Product and Zero-Padding Tensor Completion with Applications

The T-product method based upon Discrete Fourier Transformation (DFT) has found wide applications in engineering, in particular, in image processing. In this paper, we propose variable T-product, and apply the Zero-Padding Discrete Fourier Transformation (ZDFT), to convert third order tensor problems to the variable Fourier domain. An additional positive integer parameter is introduced, which is greater than the tubal dimension of the third order tensor. When the additional parameter is equal to the tubal dimension, the ZDFT reduces to DFT. Then, we propose a new tensor completion method based on ZDFT and TV regularization, called VTCTF-TV. Extensive numerical experiment results on visual data demonstrate that the superior performance of the proposed method. In addition, the zero-padding is near the size of the original tubal dimension, the resulting tensor completion can be performed much better.

math.OC

Dual Quaternion Matrices in Multi-Agent Formation Control

Three kinds of dual quaternion matrices associated with the mutual visibility graph, namely the relative configuration adjacency matrix, the logarithm adjacency matrix and the relative twist adjacency matrix, play important roles in multi-agent formation control. In this paper, we study their properties and applications. We show that the relative configuration adjacency matrix and the logarithm adjacency matrix are all Hermitian matrices, and thus have very nice spectral properties. We introduce dual quaternion Laplacian matrices, and prove a Gershgorin-type theorem for square dual quaternion Hermitian matrices, for studying properties of dual quaternion Laplacian matrices. The role of the dual quaternion Laplacian matrices in formation control is discussed.

math.OC

Sparse least squares solutions of multilinear equations

In this paper, we propose a sparse least squares (SLS) optimization model for solving multilinear equations, in which the sparsity constraint on the solutions can effectively reduce storage and computation costs. By employing variational properties of the sparsity set, along with differentiation properties of the objective function in the SLS model, the first-order optimality conditions are analyzed in terms of the stationary points. Based on the equivalent characterization of the stationary points, we propose the Newton Hard-Threshold Pursuit (NHTP) algorithm and establish its locally quadratic convergence under some regularity conditions. Numerical experiments conducted on simulated datasets including cases of Completely Positive(CP)-tensors and symmetric strong M-tensors illustrate the efficacy of our proposed NHTP method.

math.OC

Normal Cones Intersection Rule and Optimality Analysis for Low-Rank Matrix Optimization with Affine Manifolds

The low-rank matrix optimization with affine manifold (rank-MOA) aims to minimize a continuously differentiable function over a low-rank set intersecting with an affine manifold. This paper is devoted to the optimality analysis for rank-MOA. As a cornerstone, the intersection rule of the Fréchet normal cone to the feasible set of the rank-MOA is established under some mild linear independence assumptions. Aided with the resulting explicit formulae of the underlying normal cone, the so-called F-stationary point and the α-stationary point of rank-MOA are investigated and the relationship with local/global minimizers are then revealed in terms of first-order optimality conditions. Furthermore, the second-order optimality analysis, including the necessary and the sufficient conditions, is proposed based on the second-order differentiation information of the model. All these results will enrich the theory of low-rank matrix optimization and give potential clues to designing efficient numerical algorithms for seeking low rank solutions. Meanwhile, two specific applications of the rank-MOA are discussed to illustrate our proposed optimality analysis.

math.OC

Low Rank Approximation of Dual Complex Matrices

Dual complex numbers can represent rigid body motion in 2D spaces. Dual complex matrices are linked with screw theory, and have potential applications in various areas. In this paper, we study low rank approximation of dual complex matrices. We define $2$-norm for dual complex vectors, and Frobenius norm for dual complex matrices. These norms are nonnegative dual numbers. We establish the unitary invariance property of dual complex matrices. We study eigenvalues of square dual complex matrices, and show that an $n \times n$ dual complex Hermitian matrix has exactly $n$ eigenvalues, which are dual numbers. We present a singular value decomposition (SVD) theorem for dual complex matrices, define ranks and appreciable ranks for dual complex matrices, and study their properties. We establish an Eckart-Young like theorem for dual complex matrices, and present an algorithm framework for low rank approximation of dual complex matrices via truncated SVD. The SVD of dual complex matrices also provides a basic tool for Principal Component Analysis (PCA) via these matrices. Numerical experiments are reported.

math.NA

2D+3D facial expression recognition via embedded tensor manifold regularization

In this paper, a novel approach via embedded tensor manifold regularization for 2D+3D facial expression recognition (FERETMR) is proposed. Firstly, 3D tensors are constructed from 2D face images and 3D face shape models to keep the structural information and correlations. To maintain the local structure (geometric information) of 3D tensor samples in the low-dimensional tensors space during the dimensionality reduction, the $\ell_0$-norm of the core tensors and a tensor manifold regularization scheme embedded on core tensors are adopted via a low-rank truncated Tucker decomposition on the generated tensors. As a result, the obtained factor matrices will be used for facial expression classification prediction. To make the resulting tensor optimization more tractable, $\ell_1$-norm surrogate is employed to relax $\ell_0$-norm and hence the resulting tensor optimization problem has a nonsmooth objective function due to the $\ell_1$-norm and orthogonal constraints from the orthogonal Tucker decomposition. To efficiently tackle this tensor optimization problem, we establish the first-order optimality condition in terms of stationary points, and then design a block coordinate descent (BCD) algorithm with convergence analysis and the computational complexity. Numerical results on BU-3DFE database and Bosphorus databases demonstrate the effectiveness of our proposed approach.

cs.CV

Eigenvalues and Singular Values of Dual Quaternion Matrices

The poses of $m$ robotics in $n$ time points may be represented by an $m \times n$ dual quaternion matrix. In this paper, we study the spectral theory of dual quaternion matrices. We introduce right and left eigenvalues for square dual quaternion matrices. If a right eigenvalue is a dual number, then it is also a left eigenvalue. In this case, this dual number is called an eigenvalue of that dual quaternion matrix. We show that the right eigenvalues of a dual quaternion Hermitian matrix are dual numbers. Thus, they are eigenvalues. An $n \times n$ dual quaternion Hermitian matrix is shown to have exactly $n$ eigenvalues. It is positive semidefinite, or positive definite, if and only if all of its eigenvalues are nonnegative, or positive and appreciable, dual numbers, respectively. We present a unitary decomposition of a dual quaternion Hermitian matrix, and the singular value decomposition for a general dual quaternion matrix. The singular values of a dual quaternion matrix are nonnegative dual numbers.

math.RA

Computing One-bit Compressive Sensing via Double-Sparsity Constrained Optimization

One-bit compressive sensing gains its popularity in signal processing and communications due to its low storage costs and low hardware complexity. However, it has been a challenging task to recover the signal only by exploiting the one-bit (the sign) information. In this paper, we appropriately formulate the one-bit compressive sensing into a double-sparsity constrained optimization problem. The first-order optimality conditions for this nonconvex and discontinuous problem are established via the newly introduced $τ$-stationarity, based on which, a gradient projection subspace pursuit (\texttt{GPSP}) algorithm is developed. It is proven that \texttt{GPSP} can converge globally and terminate within finite steps. Numerical experiments have demonstrated its excellent performance in terms of a high order of accuracy with a fast computational speed.

math.OC