SearcharxivSearch

arXiv subjects

Haizhang Zhang

Publications and source records attributed to Haizhang Zhang.

At least 19 recordsLinked to original sources

Smallest Singular Value Estimates for Nonuniform Fourier Matrices via Periodic Nonuniform Sampling

We study the smallest singular value of nonuniform Fourier matrices in two settings: clustered nodes and perturbations of an equispaced grid. By reducing the problem to spectral norm estimates for periodic nonuniform interpolation matrices, we obtain nearly optimal bounds in both cases. For clustered nodes, we derive the first local separation condition in which each required gap depends only on the sizes of the two neighboring clusters. For perturbations with the bound \(1/4\leq L<1/2\), our result confirms the conjecture of Austin and Trefethen on the \(2\)-norm Lebesgue constant up to a logarithmic factor.

math.NA

Exact-Support Counterexamples to Euclidean-to-Spherical Transfer of Positive Definiteness in Even Dimensions

For every odd integer $d\geq3$, a continuous function $φ\colon[0,\infty)\to\mathbb R$ supported in $[0,π]$ and isotropic positive definite on $\mathbb R^d$ remains so on $\mathbb S^d$. In even dimensions, recent work shows that this transfer fails under every prescribed positive upper bound on the support. We prove an exact-support refinement with a construction uniform in the prescribed radius. More precisely, for each $d=2m\geq2$ and $R\in(0,π]$, we construct a function $φ$ whose radial extension belongs to $C_c^\infty(\mathbb R^d)$ and has support radius exactly $R$, such that $φ(\|\mathbf{x}-\mathbf{y}\|_2)$ is strictly positive definite on $\mathbb R^d$, whereas $φ(ρ(\mathbf{x},\mathbf{y}))$ is not positive definite on $\mathbb S^d$. Thus every admissible support radius is attained by a smooth, strictly Euclidean positive-definite counterexample.

math.CA

Reliable Neural Collapse Approximation for Open-World Test-Time Adaptation

Test-Time Adaptation (TTA) methods aim to bridge the domain gap between the source and target domains. However, traditional TTA methods become ineffective when the label distribution shift occurs, a challenge commonly referred to as an open-world scenario. In this paper, we introduce a new method named Reliable Neural Collapse approximation (ReNC) for Open-World Test-Time Adaptation (OWTTA). Specifically, we leverage neural collapse as a structural prior for reliable target-domain adaptation. Guided by this prior, we justify that the pre-trained classifier weights can serve as the prototypes of the source domain. By measuring the similarity between samples and prototypes, we filter out the Out-Of-Distribution~(OOD) samples for reliable updates. Furthermore, we propose a neural collapse approximation mechanism to refine these prototypes, ensuring they can gradually adapt to the target domain while maintaining the neural collapse structure. Extensive experiments on several open-world benchmarks demonstrate the superiority of the proposed method. Our empirical analysis suggests that ReNC better preserves NC-related properties in the target domain, providing useful evidence for explaining reliable OWTTA and offering new insights for model design. Code is available at https://github.com/JiaqiLin-AI/ReNC.

cs.LG

A Random Integration Algorithm for High-dimensional Function Spaces

We introduce a novel random integration algorithm that boasts both high convergence order and polynomial tractability for functions characterized by sparse frequencies or rapidly decaying Fourier coefficients. Specifically, for integration in periodic isotropic Sobolev space and the isotropic Sobolev space with compact support, our approach attains a nearly optimal root mean square error (RMSE) bound. In contrast to previous nearly optimal algorithms, our method exhibits polynomial tractability, ensuring that the number of samples does not scale exponentially with increasing dimensions. Our integration algorithm also enjoys nearly optimal bound for weighted Korobov space. Furthermore, the algorithm can be applied without the need for prior knowledge of weights, distinguishing it from the component-by-component algorithm. For integration in the Wiener algebra, the sample complexity of our algorithm is independent of the decay rate of Fourier coefficients. The effectiveness of the integration is confirmed through numerical experiments.

math.NA

Vector-valued Reproducing Kernel Banach Spaces with Group Lasso Norms

Focusing on establishing a mathematical basis for kernel methods in sparse multi-task learning, we explore the theory of vector-valued reproducing kernel Banach spaces (RKBSs) endowed with $\ell_{p,1}$-norms ($1\le p\le +\infty$), encompassing both the sparse learning case when $p=1$ and the group lasso when $p=2$. We develop RKBSs equipped with these group lasso norms that support the linear representer theorem for regularized learning frameworks. Additionally, we introduce reproducing kernels admissible for this construction. Such reproducing kernels are applicable to sparse multi-task learning with group lasso norms.

math.FA

Uniform Convergence of Deep Neural Networks with Lipschitz Continuous Activation Functions and Variable Widths

We consider deep neural networks with a Lipschitz continuous activation function and with weight matrices of variable widths. We establish a uniform convergence analysis framework in which sufficient conditions on weight matrices and bias vectors together with the Lipschitz constant are provided to ensure uniform convergence of the deep neural networks to a meaningful function as the number of their layers tends to infinity. In the framework, special results on uniform convergence of deep neural networks with a fixed width, bounded widths and unbounded widths are presented. In particular, as convolutional neural networks are special deep neural networks with weight matrices of increasing widths, we put forward conditions on the mask sequence which lead to uniform convergence of resulting convolutional neural networks. The Lipschitz continuity assumption on the activation functions allows us to include in our theory most of commonly used activation functions in applications.

cs.LG

Convergence of Deep ReLU Networks

We explore convergence of deep neural networks with the popular ReLU activation function, as the depth of the networks tends to infinity. To this end, we introduce the notion of activation domains and activation matrices of a ReLU network. By replacing applications of the ReLU activation function by multiplications with activation matrices on activation domains, we obtain an explicit expression of the ReLU network. We then identify the convergence of the ReLU networks as convergence of a class of infinite products of matrices. Sufficient and necessary conditions for convergence of these infinite products of matrices are studied. As a result, we establish necessary conditions for ReLU networks to converge that the sequence of weight matrices converges to the identity matrix and the sequence of the bias vectors converges to zero as the depth of ReLU networks increases to infinity. Moreover, we obtain sufficient conditions in terms of the weight matrices and bias vectors at hidden layers for pointwise convergence of deep ReLU networks. These results provide mathematical insights to the design strategy of the well-known deep residual networks in image classification.

cs.LG

Dual Information Enhanced Multi-view Attributed Graph Clustering

Multi-view attributed graph clustering is an important approach to partition multi-view data based on the attribute feature and adjacent matrices from different views. Some attempts have been made in utilizing Graph Neural Network (GNN), which have achieved promising clustering performance. Despite this, few of them pay attention to the inherent specific information embedded in multiple views. Meanwhile, they are incapable of recovering the latent high-level representation from the low-level ones, greatly limiting the downstream clustering performance. To fill these gaps, a novel Dual Information enhanced multi-view Attributed Graph Clustering (DIAGC) method is proposed in this paper. Specifically, the proposed method introduces the Specific Information Reconstruction (SIR) module to disentangle the explorations of the consensus and specific information from multiple views, which enables GCN to capture the more essential low-level representations. Besides, the Mutual Information Maximization (MIM) module maximizes the agreement between the latent high-level representation and low-level ones, and enables the high-level representation to satisfy the desired clustering structure with the help of the Self-supervised Clustering (SC) module. Extensive experiments on several real-world benchmarks demonstrate the effectiveness of the proposed DIAGC method compared with the state-of-the-art baselines.

cs.AI

Exponential Approximation of Band-limited Functions from Nonuniform Sampling by Regularization Methods

Reconstructing a band-limited function from its finite sample data is a fundamental task in signal analysis. A Gaussian regularized Shannon sampling series has been proved to be able to achieve exponential convergence for uniform sampling. Whether such an exponential convergence can also be achieved for nonuniform sampling by regularization methods was unresolved. In this paper, we give an affirmative and constructive answer to this question. Specifically, we show that one can recover a band-limited function by Gaussian or hyper-Gaussian regularized nonuniform sampling series with an exponential convergence rate. Our analysis is based on the residue theorem in complex analysis, which is used to represent the truncated error by a contour integral. Several concrete examples of nonuniform sampling with exponential convergence will be presented.

eess.SP

Convergence of Deep Neural Networks with General Activation Functions and Pooling

Deep neural networks, as a powerful system to represent high dimensional complex functions, play a key role in deep learning. Convergence of deep neural networks is a fundamental issue in building the mathematical foundation for deep learning. We investigated the convergence of deep ReLU networks and deep convolutional neural networks in two recent researches (arXiv:2107.12530, 2109.13542). Only the Rectified Linear Unit (ReLU) activation was studied therein, and the important pooling strategy was not considered. In this current work, we study the convergence of deep neural networks as the depth tends to infinity for two other important activation functions: the leaky ReLU and the sigmoid function. Pooling will also be studied. As a result, we prove that the sufficient condition established in arXiv:2107.12530, 2109.13542 is still sufficient for the leaky ReLU networks. For contractive activation functions such as the sigmoid function, we establish a weaker sufficient condition for uniform convergence of deep neural networks.

cs.LG

Convergence Analysis of Deep Residual Networks

Various powerful deep neural network architectures have made great contribution to the exciting successes of deep learning in the past two decades. Among them, deep Residual Networks (ResNets) are of particular importance because they demonstrated great usefulness in computer vision by winning the first place in many deep learning competitions. Also, ResNets were the first class of neural networks in the development history of deep learning that are really deep. It is of mathematical interest and practical meaning to understand the convergence of deep ResNets. We aim at characterizing the convergence of deep ResNets as the depth tends to infinity in terms of the parameters of the networks. Toward this purpose, we first give a matrix-vector description of general deep neural networks with shortcut connections and formulate an explicit expression for the networks by using the notions of activation domains and activation matrices. The convergence is then reduced to the convergence of two series involving infinite products of non-square matrices. By studying the two series, we establish a sufficient condition for pointwise convergence of ResNets. Our result is able to give justification for the design of ResNets. We also conduct experiments on benchmark machine learning data to verify our results.

cs.LG

Convergence of Deep Convolutional Neural Networks

Convergence of deep neural networks as the depth of the networks tends to infinity is fundamental in building the mathematical foundation for deep learning. In a previous study, we investigated this question for deep ReLU networks with a fixed width. This does not cover the important convolutional neural networks where the widths are increasing from layer to layer. For this reason, we first study convergence of general ReLU networks with increasing widths and then apply the results obtained to deep convolutional neural networks. It turns out the convergence reduces to convergence of infinite products of matrices with increasing sizes, which has not been considered in the literature. We establish sufficient conditions for convergence of such infinite products of matrices. Based on the conditions, we present sufficient conditions for piecewise convergence of general deep ReLU networks with increasing widths, and as well as pointwise convergence of deep ReLU convolutional neural networks.

cs.LG

On Reproducing Kernel Banach Spaces: Generic Definitions and Unified Framework of Constructions

Recently, there has been emerging interest in constructing reproducing kernel Banach spaces (RKBS) for applied and theoretical purposes such as machine learning, sampling reconstruction, sparse approximation and functional analysis. Existing constructions include the reflexive RKBS via a bilinear form, the semi-inner-product RKBS, the RKBS with $\ell^1$ norm, the $p$-norm RKBS via generalized Mercer kernels, etc. The definitions of RKBS and the associated reproducing kernel in those references are dependent on the construction. Moreover, relations among those constructions are unclear. We explore a generic definition of RKBS and the reproducing kernel for RKBS that is independent of construction. Furthermore, we propose a framework of constructing RKBSs that unifies existing constructions mentioned above via a continuous bilinear form and a pair of feature maps. A new class of Orlicz RKBSs is proposed. Finally, we develop representer theorems for machine learning in RKBSs constructed in our framework, which also unifies representer theorems in existing RKBSs.

cs.LG

Learning Rates for Multi-task Regularization Networks

Multi-task learning is an important trend of machine learning in facing the era of artificial intelligence and big data. Despite a large amount of researches on learning rate estimates of various single-task machine learning algorithms, there is little parallel work for multi-task learning. We present mathematical analysis on the learning rate estimate of multi-task learning based on the theory of vector-valued reproducing kernel Hilbert spaces and matrix-valued reproducing kernels. For the typical multi-task regularization networks, an explicit learning rate dependent both on the number of sample data and the number of tasks is obtained. It reveals that the generalization ability of multi-task learning algorithms is indeed affected as the number of tasks increases.

cs.LG

Multi-task Learning in Vector-valued Reproducing Kernel Banach Spaces with the $\ell^1$ Norm

Targeting at sparse multi-task learning, we consider regularization models with an $\ell^1$ penalty on the coefficients of kernel functions. In order to provide a kernel method for this model, we construct a class of vector-valued reproducing kernel Banach spaces with the $\ell^1$ norm. The notion of multi-task admissible kernels is proposed so that the constructed spaces could have desirable properties including the crucial linear representer theorem. Such kernels are related to bounded Lebesgue constants of a kernel interpolation question. We study the Lebesgue constant of multi-task kernels and provide examples of admissible kernels. Furthermore, we present numerical experiments for both synthetic data and real-world benchmark data to demonstrate the advantages of the proposed construction and regularization models.

math.FA

Convergence Analysis of the Gaussian Regularized Shannon Sampling Formula

We consider the reconstruction of a bandlimited function from its finite localized sample data. Truncating the classical Shannon sampling series results in an unsatisfactory convergence rate due to the slow decayness of the sinc function. To overcome this drawback, a simple and highly effective method, called the Gaussian regularization of the Shannon series, was proposed in the engineering and has received remarkable attention. It works by multiplying the sinc function in the Shannon series with a regularized Gaussian function. L. Qian (Proc. Amer. Math. Soc., 2003) established the convergence rate of $O(\sqrt{n}\exp(-\frac{π-δ}2n))$ for this method, where $δ<π$ is the bandwidth and $n$ is the number of sample data. C. Micchelli {\it et al.} (J. Complexity, 2009) proposed a different regularized method and obtained the corresponding convergence rate of $O(\frac1{\sqrt{n}}\exp(-\frac{π-δ}2n))$. This latter rate is by far the best among all regularized methods for the Shannon series. However, their regularized method involves the solving of a linear system and is implicit and more complicated. The main objective of this note is to show that the Gaussian regularization of the Shannon series can also achieve the same best convergence rate as that by C. Micchelli {\it et al}. We also show that the Gaussian regularization method can improve the convergence rate for the useful average sampling. Finally, the outstanding performance of numerical experiments justifies our results.

cs.IT

Exponential Approximation of Multivariate Bandlimited Functions from Average Oversampling

Instead of sampling a function at a single point, average sampling takes the weighted sum of function values around the point. Such a sampling strategy is more practical and more stable. In this note, we present an explicit method with an exponentially-decaying approximation error to reconstruct a multivariate bandlimited function from its finite average oversampling data. The key problem in our analysis is how to extend a function so that its Fourier transform decays at an optimal rate to zero at infinity.

cs.IT

Exponential Approximation of Bandlimited Functions from Average Oversampling

Weighted average sampling is more practical and numerically more stable than sampling at single points as in the classical Shannon sampling framework. Using the frame theory, one can completely reconstruct a bandlimited function from its suitably-chosen average sample data. When only finitely many sample data are available, truncating the complete reconstruction series with the standard dual frame results in very slow convergence. We present in this note a method of reconstructing a bandlimited function from finite average oversampling with an exponentially-decaying approximation error.

cs.IT