Searcharxiv⌕ Search

arXiv subjects

Yong Lin

Publications and source records attributed to Yong Lin.

At least 55 records · Page 3Linked to original sources

Optimal Sample Selection Through Uncertainty Estimation and Its Application in Deep Learning

Modern deep learning heavily relies on large labeled datasets, which often comse with high costs in terms of both manual labeling and computational resources. To mitigate these challenges, researchers have explored the use of informative subset selection techniques, including coreset selection and active learning. Specifically, coreset selection involves sampling data with both input ($\bx$) and output ($\by$), active learning focuses solely on the input data ($\bx$). In this study, we present a theoretically optimal solution for addressing both coreset selection and active learning within the context of linear softmax regression. Our proposed method, COPS (unCertainty based OPtimal Sub-sampling), is designed to minimize the expected loss of a model trained on subsampled data. Unlike existing approaches that rely on explicit calculations of the inverse covariance matrix, which are not easily applicable to deep learning scenarios, COPS leverages the model's logits to estimate the sampling ratio. This sampling ratio is closely associated with model uncertainty and can be effectively applied to deep learning tasks. Furthermore, we address the challenge of model sensitivity to misspecification by incorporating a down-weighting approach for low-density samples, drawing inspiration from previous works. To assess the effectiveness of our proposed method, we conducted extensive empirical experiments using deep neural networks on benchmark datasets. The results consistently showcase the superior performance of COPS compared to baseline methods, reaffirming its efficacy.

stat.ML↗

What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?

Offline goal-conditioned RL (GCRL) offers a way to train general-purpose agents from fully offline datasets. In addition to being conservative within the dataset, the generalization ability to achieve unseen goals is another fundamental challenge for offline GCRL. However, to the best of our knowledge, this problem has not been well studied yet. In this paper, we study out-of-distribution (OOD) generalization of offline GCRL both theoretically and empirically to identify factors that are important. In a number of experiments, we observe that weighted imitation learning enjoys better generalization than pessimism-based offline RL method. Based on this insight, we derive a theory for OOD generalization, which characterizes several important design choices. We then propose a new offline GCRL method, Generalizable Offline goAl-condiTioned RL (GOAT), by combining the findings from our theoretical and empirical studies. On a new benchmark containing 9 independent identically distributed (IID) tasks and 17 OOD tasks, GOAT outperforms current state-of-the-art methods by a large margin.

cs.LG↗

ID and OOD Performance Are Sometimes Inversely Correlated on Real-world Datasets

Several studies have compared the in-distribution (ID) and out-of-distribution (OOD) performance of models in computer vision and NLP. They report a frequent positive correlation and some surprisingly never even observe an inverse correlation indicative of a necessary trade-off. The possibility of inverse patterns is important to determine whether ID performance can serve as a proxy for OOD generalization capabilities. This paper shows with multiple datasets that inverse correlations between ID and OOD performance do happen in real-world data - not only in theoretical worst-case settings. We also explain theoretically how these cases can arise even in a minimal linear setting, and why past studies could miss such cases due to a biased selection of models. Our observations lead to recommendations that contradict those found in much of the current literature. - High OOD performance sometimes requires trading off ID performance. - Focusing on ID performance alone may not lead to optimal OOD performance. It may produce diminishing (eventually negative) returns in OOD performance. - In these cases, studies on OOD generalization that use ID performance for model selection (a common recommended practice) will necessarily miss the best-performing models, making these studies blind to a whole range of phenomena.

cs.LG↗

Particle-based Variational Inference with Preconditioned Functional Gradient Flow

Particle-based variational inference (VI) minimizes the KL divergence between model samples and the target posterior with gradient flow estimates. With the popularity of Stein variational gradient descent (SVGD), the focus of particle-based VI algorithms has been on the properties of functions in Reproducing Kernel Hilbert Space (RKHS) to approximate the gradient flow. However, the requirement of RKHS restricts the function class and algorithmic flexibility. This paper offers a general solution to this problem by introducing a functional regularization term that encompasses the RKHS norm as a special case. This allows us to propose a new particle-based VI algorithm called preconditioned functional gradient flow (PFG). Compared to SVGD, PFG has several advantages. It has a larger function class, improved scalability in large particle-size scenarios, better adaptation to ill-conditioned distributions, and provable continuous-time convergence in KL divergence. Additionally, non-linear function classes such as neural networks can be incorporated to estimate the gradient flow. Our theory and experiments demonstrate the effectiveness of the proposed framework.

stat.ML↗

Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning

There is a rising interest in further exploring the zero-shot learning potential of large pre-trained language models (PLMs). A new paradigm called data-generation-based zero-shot learning has achieved impressive success. In this paradigm, the synthesized data from the PLM acts as the carrier of knowledge, which is used to train a task-specific model with orders of magnitude fewer parameters than the PLM, achieving both higher performance and efficiency than prompt-based zero-shot learning methods on PLMs. The main hurdle of this approach is that the synthesized data from PLM usually contains a significant portion of low-quality samples. Fitting on such data will greatly hamper the performance of the task-specific model, making it unreliable for deployment. Previous methods remedy this issue mainly by filtering synthetic data using heuristic metrics(e.g., output confidence), or refining the data with the help of a human expert, which comes with excessive manual tuning or expensive costs. In this paper, we propose a novel noise-robust re-weighting framework SunGen to automatically construct high-quality data for zero-shot classification problems. Our framework features the ability to learn the sample weights indicating data quality without requiring any human annotation. We theoretically and empirically verify the ability of our method to help construct good-quality synthetic datasets. Notably, SunGen-LSTM yields a 9.8% relative improvement than the baseline on average accuracy across eight different established text classification tasks.

cs.CL↗

Black-box Prompt Learning for Pre-trained Language Models

The increasing scale of general-purpose Pre-trained Language Models (PLMs) necessitates the study of more efficient adaptation across different downstream tasks. In this paper, we establish a Black-box Discrete Prompt Learning (BDPL) to resonate with pragmatic interactions between the cloud infrastructure and edge devices. Particularly, instead of fine-tuning the model in the cloud, we adapt PLMs by prompt learning, which efficiently optimizes only a few parameters of the discrete prompts. Moreover, we consider the scenario that we do not have access to the parameters and gradients of the pre-trained models, except for its outputs given inputs. This black-box setting secures the cloud infrastructure from potential attack and misuse to cause a single-point failure, which is preferable to the white-box counterpart by current infrastructures. Under this black-box constraint, we apply a variance-reduced policy gradient algorithm to estimate the gradients of parameters in the categorical distribution of each discrete prompt. In light of our method, the user devices can efficiently tune their tasks by querying the PLMs bounded by a range of API calls. Our experiments on RoBERTa and GPT-3 demonstrate that the proposed algorithm achieves significant improvement on eight benchmarks in a cloud-device collaboration manner. Finally, we conduct in-depth case studies to comprehensively analyze our method in terms of various data sizes, prompt lengths, training budgets, optimization objectives, prompt transferability, and explanations of the learned prompts. Our code will be available at https://github.com/shizhediao/Black-Box-Prompt-Learning.

cs.CL↗

Model Agnostic Sample Reweighting for Out-of-Distribution Learning

Distributionally robust optimization (DRO) and invariant risk minimization (IRM) are two popular methods proposed to improve out-of-distribution (OOD) generalization performance of machine learning models. While effective for small models, it has been observed that these methods can be vulnerable to overfitting with large overparameterized models. This work proposes a principled method, \textbf{M}odel \textbf{A}gnostic sam\textbf{PL}e r\textbf{E}weighting (\textbf{MAPLE}), to effectively address OOD problem, especially in overparameterized scenarios. Our key idea is to find an effective reweighting of the training samples so that the standard empirical risk minimization training of a large model on the weighted training data leads to superior OOD generalization performance. The overfitting issue is addressed by considering a bilevel formulation to search for the sample reweighting, in which the generalization complexity depends on the search space of sample weights instead of the model size. We present theoretical analysis in linear case to prove the insensitivity of MAPLE to model size, and empirically verify its superiority in surpassing state-of-the-art methods by a large margin. Code is available at \url{https://github.com/x-zho14/MAPLE}.

cs.LG↗

Probabilistic Bilevel Coreset Selection

The goal of coreset selection in supervised learning is to produce a weighted subset of data, so that training only on the subset achieves similar performance as training on the entire dataset. Existing methods achieved promising results in resource-constrained scenarios such as continual learning and streaming. However, most of the existing algorithms are limited to traditional machine learning models. A few algorithms that can handle large models adopt greedy search approaches due to the difficulty in solving the discrete subset selection problem, which is computationally costly when coreset becomes larger and often produces suboptimal results. In this work, for the first time we propose a continuous probabilistic bilevel formulation of coreset selection by learning a probablistic weight for each training sample. The overall objective is posed as a bilevel optimization problem, where 1) the inner loop samples coresets and train the model to convergence and 2) the outer loop updates the sample probability progressively according to the model's performance. Importantly, we develop an efficient solver to the bilevel optimization problem via unbiased policy gradient without trouble of implicit differentiation. We provide the convergence property of our training procedure and demonstrate the superiority of our algorithm against various coreset selection methods in various tasks, especially in more challenging label-noise and class-imbalance scenarios.

cs.LG↗

Stable Learning via Sparse Variable Independence

The problem of covariate-shift generalization has attracted intensive research attention. Previous stable learning algorithms employ sample reweighting schemes to decorrelate the covariates when there is no explicit domain information about training data. However, with finite samples, it is difficult to achieve the desirable weights that ensure perfect independence to get rid of the unstable variables. Besides, decorrelating within stable variables may bring about high variance of learned models because of the over-reduced effective sample size. A tremendous sample size is required for these algorithms to work. In this paper, with theoretical justification, we propose SVI (Sparse Variable Independence) for the covariate-shift generalization problem. We introduce sparsity constraint to compensate for the imperfectness of sample reweighting under the finite-sample setting in previous methods. Furthermore, we organically combine independence-based sample reweighting and sparsity-based variable selection in an iterative way to avoid decorrelating within stable variables, increasing the effective sample size to alleviate variance inflation. Experiments on both synthetic and real-world datasets demonstrate the improvement of covariate-shift generalization performance brought by SVI.

cs.LG↗

ZIN: When and How to Learn Invariance Without Environment Partition?

It is commonplace to encounter heterogeneous data, of which some aspects of the data distribution may vary but the underlying causal mechanisms remain constant. When data are divided into distinct environments according to the heterogeneity, recent invariant learning methods have proposed to learn robust and invariant models based on this environment partition. It is hence tempting to utilize the inherent heterogeneity even when environment partition is not provided. Unfortunately, in this work, we show that learning invariant features under this circumstance is fundamentally impossible without further inductive biases or additional information. Then, we propose a framework to jointly learn environment partition and invariant representation, assisted by additional auxiliary information. We derive sufficient and necessary conditions for our framework to provably identify invariant features under a fairly general setting. Experimental results on both synthetic and real world datasets validate our analysis and demonstrate an improved performance of the proposed framework over existing methods. Finally, our results also raise the need of making the role of inductive biases more explicit in future works, when considering learning invariant models without environment partition. Codes are available at https://github.com/linyongver/ZIN_official .

cs.LG↗

Discrete Morse Theory on Digraphs

In this paper, we give a necessary and sufficient condition that discrete Morse functions on a digraph can be extended to be Morse functions on its transitive closure, from this we can extend the Morse theory to digraphs by using quasi-isomorphism between path complex and discrete Morse complex, we also prove a general sufficient condition for digraphs that the Morse functions satisfying this necessary and sufficient condition.

math.CO↗

The existence of the solution of the wave equation on graphs

Let $G=(V, E)$ be a finite weighted graph, and $Ω\subseteq V$ be a domain such that $Ω^\circ\neq\emptyset$. In this paper, we study the following initial boundary problem for the non-homogenous wave equation \begin{equation*} \left\{ \begin{aligned} &\partial_t^2 u(t,x)-Δ_Ωu(t,x)=f(t,x),\qquad&&(t,x)\in[0,\infty)\times Ω^\circ,\\ &u(0,x)=g(x),\qquad&& x\inΩ^\circ,\\ &\partial_tu(0,x)=h(x),\qquad&& x\inΩ^\circ,\\ &u(t,x)=0,\qquad&&(t,x)\in[0,\infty)\times\partial Ω, \end{aligned} \right. \end{equation*} where $Δ_Ω$ denotes the Dirichlet Laplacian on $Ω^\circ$. Using Rothe's method, we prove that the above wave equation has a unique solution.

math.AP↗

Application of Rothe's method to a nonlinear wave equation on graphs

We study a nonlinear wave equation on finite connected weighted graphs. Using Rothe's and energy methods, we prove the existence and uniqueness of solution under certain assumption. For linear wave equation on graphs, Lin and Xie \cite{Lin-Xie} obtained the existence and uniqueness of solution. The main novelty of this paper is that the wave equation we considered has the nonlinear damping term $|u_t|^{p-1}\cdot u_t$ ($p>1$).

math.AP↗

Semilinear heat equations and parabolic variational inequalities on graphs

Let $G=(V,E)$ be a locally finite connected weighted graph, and $Ω$ be an unbounded subset of $V$. Using Rothe's method, we study the existence of solutions for the semilinear heat equation $\partial_tu+|u|^{p-1}\cdot u=Δu~(p\ge1)$ and the parabolic variational inequality \begin{eqnarray*} \int_{Ω^\circ} \partial_tu\cdot(v-u)\,dμ\ge \int_{Ω^\circ}(Δu+f)\cdot(v-u)\,dμ\qquad\mbox{for any }v\in \mathcal{H}, \end{eqnarray*} where $\mathcal{H}=\{u\in W^{1,2}(V):u=0\mbox{ on }V\backslashΩ^\circ\}$.

math.AP↗

Witten-Morse functions and Morse inequalities on digraphs

In this paper, we prove that discrete Morse functions on digraphs are flat Witten-Morse functions and Witten complexes of transitive digraphs approach to Morse complexes. We construct a chain complex consisting of the formal linear combinations of paths which are not only critical paths of the transitive closure but also allowed elementary paths of the digraph, and prove that the homology of the new chain complex is isomorphic to the path homology. On the basis of the above results, we give the Morse inequalities on digraphs.

math.CO↗

Calculus of variations on locally finite graphs

Let $G=(V,E)$ be a locally finite graph. Firstly, using calculus of variations, including a direct method of variation and the mountain-pass theory, we get sequences of solutions to several local equations on $G$ (the Schrödinger equation, the mean field equation, and the Yamabe equation). Secondly, we derive uniform estimates for those local solution sequences. Finally, we obtain global solutions by extracting convergent sequence of solutions. Our method can be described as a variational method from local to global.

math.AP↗

A heat flow for the mean field equation on a finite graph

Inspired by works of Castéras (Pacific J. Math., 2015), Li-Zhu (Calc. Var., 2019) and Sun-Zhu (Calc. Var., 2020), we propose a heat flow for the mean field equation on a connected finite graph $G=(V,E)$. Namely $$ \left\{\begin{array}{lll} \partial_tϕ(u)=Δu-Q+ρ\frac{e^u}{\int_Ve^udμ}\\[1.5ex] u(\cdot,0)=u_0, \end{array}\right. $$ where $Δ$ is the standard graph Laplacian, $ρ$ is a real number, $Q:V\rightarrow\mathbb{R}$ is a function satisfying $\int_VQdμ=ρ$, and $ϕ:\mathbb{R}\rightarrow\mathbb{R}$ is one of certain smooth functions including $ϕ(s)=e^s$. We prove that for any initial data $u_0$ and any $ρ\in\mathbb{R}$, there exists a unique solution $u:V\times[0,+\infty)\rightarrow\mathbb{R}$ of the above heat flow; moreover, $u(x,t)$ converges to some function $u_\infty:V\rightarrow\mathbb{R}$ uniformly in $x\in V$ as $t\rightarrow+\infty$, and $u_\infty$ is a solution of the mean field equation $$Δu_\infty-Q+ρ\frac{e^{u_\infty}}{\int_Ve^{u_\infty}dμ}=0.$$ Though $G$ is a finite graph, this result is still unexpected, even in the special case $Q\equiv 0$. Our approach reads as follows: the short time existence of the heat flow follows from the ODE theory; various integral estimates give its long time existence; moreover we establish a Lojasiewicz-Simon type inequality and use it to conclude the convergence of the heat flow.

math.AP↗

Torsion of digraphs and path complexes

We define the notions of Reidemeister torsion and analytic torsion for directed graphs by means of the path homology theory introduced by the authors in \cite{Grigoryan-Lin-Muranov-Yau2013, Grigoryan-Lin-Muranov-Yau2014, Grigoryan-Lin-Muranov-Yau2015, Grigoryan-Lin-Muranov-Yau2020}. We prove the identity of the two notions of torsions as well as obtain formulas for torsions of Cartesian products and joins of digraphs.

math.CO↗