Searcharxiv⌕ Search

arXiv subjects

Junhong Lin

Publications and source records attributed to Junhong Lin.

44 records · Page 3Linked to original sources

Generalization Properties of Doubly Stochastic Learning Algorithms

Doubly stochastic learning algorithms are scalable kernel methods that perform very well in practice. However, their generalization properties are not well understood and their analysis is challenging since the corresponding learning sequence may not be in the hypothesis space induced by the kernel. In this paper, we provide an in-depth theoretical analysis for different variants of doubly stochastic learning algorithms within the setting of nonparametric regression in a reproducing kernel Hilbert space and considering the square loss. Particularly, we derive convergence results on the generalization error for the studied algorithms either with or without an explicit penalty term. To the best of our knowledge, the derived results for the unregularized variants are the first of this kind, while the results for the regularized variants improve those in the literature. The novelties in our proof are a sample error bound that requires controlling the trace norm of a cumulative operator, and a refined analysis of bounding initial error.

stat.ML↗

Optimal Rates for Learning with Nyström Stochastic Gradient Methods

In the setting of nonparametric regression, we propose and study a combination of stochastic gradient methods with Nyström subsampling, allowing multiple passes over the data and mini-batches. Generalization error bounds for the studied algorithm are provided. Particularly, optimal learning rates are derived considering different possible choices of the step-size, the mini-batch size, the number of iterations/passes, and the subsampling level. In comparison with state-of-the-art algorithms such as the classic stochastic gradient methods and kernel ridge regression with Nyström, the studied algorithm has advantages on the computational complexity, while achieving the same optimal learning rates. Moreover, our results indicate that using mini-batches can reduce the total computational cost while achieving the same optimal statistical results.

stat.ML↗

Generalization Properties and Implicit Regularization for Multiple Passes SGM

We study the generalization properties of stochastic gradient methods for learning with convex loss functions and linearly parameterized functions. We show that, in the absence of penalizations or constraints, the stability and approximation properties of the algorithm can be controlled by tuning either the step-size or the number of passes over the data. In this view, these parameters can be seen to control a form of implicit regularization. Numerical results complement the theoretical findings.

cs.LG↗

Restricted $q$-Isometry Properties Adapted to Frames for Nonconvex $l_q$-Analysis

This paper discusses reconstruction of signals from few measurements in the situation that signals are sparse or approximately sparse in terms of a general frame via the $l_q$-analysis optimization with $0<q\leq 1$. We first introduce a notion of restricted $q$-isometry property ($q$-RIP) adapted to a dictionary, which is a natural extension of the standard $q$-RIP, and establish a generalized $q$-RIP condition for approximate reconstruction of signals via the $l_q$-analysis optimization. We then determine how many random, Gaussian measurements are needed for the condition to hold with high probability. The resulting sufficient condition is met by fewer measurements for smaller $q$ than when $q=1$. The introduced generalized $q$-RIP is also useful in compressed data separation. In compressed data separation, one considers the problem of reconstruction of signals' distinct subcomponents, which are (approximately) sparse in morphologically different dictionaries, from few measurements. With the notion of generalized $q$-RIP, we show that under an usual assumption that the dictionaries satisfy a mutual coherence condition, the $l_q$ split analysis with $0<q\leq1 $ can approximately reconstruct the distinct components from fewer random Gaussian measurements with small $q$ than when $q=1$

cs.IT↗

Modified Fejér sequences and applications

In this note, we propose and study the notion of modified Fejér sequences. Within a Hilbert space setting, we show that it provides a unifying framework to prove convergence rates for objective function values of several optimization algorithms. In particular, our results apply to forward-backward splitting algorithm, incremental subgradient proximal algorithm, and the Douglas-Rachford splitting method including and generalizing known results.

math.OC↗

Iterative Regularization for Learning with Convex Loss Functions

We consider the problem of supervised learning with convex loss functions and propose a new form of iterative regularization based on the subgradient method. Unlike other regularization approaches, in iterative regularization no constraint or penalization is considered, and generalization is achieved by (early) stopping an empirical iteration. We consider a nonparametric setting, in the framework of reproducing kernel Hilbert spaces, and prove finite sample bounds on the excess risk under general regularity conditions. Our study provides a new class of efficient regularized learning algorithms and gives insights on the interplay between statistics and optimization in machine learning.

stat.ML↗

Sparse Recovery with Coherent Tight Frames via Analysis Dantzig Selector and Analysis LASSO

This article considers recovery of signals that are sparse or approximately sparse in terms of a (possibly) highly overcomplete and coherent tight frame from undersampled data corrupted with additive noise. We show that the properly constrained $l_1$-analysis, called analysis Dantzig selector, stably recovers a signal which is nearly sparse in terms of a tight frame provided that the measurement matrix satisfies a restricted isometry property adapted to the tight frame. As a special case, we consider the Gaussian noise. Further, under a sparsity scenario, with high probability, the recovery error from noisy data is within a log-like factor of the minimax risk over the class of vectors which are at most $s$ sparse in terms of the tight frame. Similar results for the analysis LASSO are showed. The above two algorithms provide guarantees only for noise that is bounded or bounded with high probability (for example, Gaussian noise). However, when the underlying measurements are corrupted by sparse noise, these algorithms perform suboptimally. We demonstrate robust methods for reconstructing signals that are nearly sparse in terms of a tight frame in the presence of bounded noise combined with sparse noise. The analysis in this paper is based on the restricted isometry property adapted to a tight frame, which is a natural extension to the standard restricted isometry property.

cs.IT↗

Compressed Sensing with coherent tight frames via $l_q$-minimization for $0<q\leq1$

Our aim of this article is to reconstruct a signal from undersampled data in the situation that the signal is sparse in terms of a tight frame. We present a condition, which is independent of the coherence of the tight frame, to guarantee accurate recovery of signals which are sparse in the tight frame, from undersampled data with minimal $l_1$-norm of transform coefficients. This improves the result in [1]. Also, the $l_q$-minimization $(0<q<1)$ approaches are introduced. We show that under a suitable condition, there exists a value $q_0\in(0,1]$ such that for any $q\in(0,q_0)$, each solution of the $l_q$-minimization is approximately well to the true signal. In particular, when the tight frame is an identity matrix or an orthonormal basis, all results obtained in this paper appeared in [13] and [26].

math.NA↗