SearcharxivSearch

arXiv subjects

Andrea Schioppa

Publications and source records attributed to Andrea Schioppa.

At least 19 recordsLinked to original sources

Covariance-aware sampling for Diffusion Models

We present a covariance-aware sampler that improves the quality of pixel-space Diffusion Model (DM) sampling in the few-step regime. We hypothesize that in the few-step regime samplers fail because they rely solely on the predicted mean of the reverse distribution, while our solution explicitly models the reverse-process covariance. Our method combines Tweedie's formula to estimate the covariance with an efficient, structured Fourier-space decomposition of the covariance matrix. Implemented as an extension of DDIM, our method requires only a minimal overhead: one extra Jacobian-Vector Product (JVP) per step. We demonstrate that for pixel-based DMs, our method consistently produces superior samples compared to state-of-the-art second order samplers (Heun, DPM-Solver++) and the recent aDDIM sampler, at an identical number of function evaluations (NFE).

stat.ML

Model Integrity when Unlearning with T2I Diffusion Models

The rapid advancement of text-to-image Diffusion Models has led to their widespread public accessibility. However these models, trained on large internet datasets, can sometimes generate undesirable outputs. To mitigate this, approximate Machine Unlearning algorithms have been proposed to modify model weights to reduce the generation of specific types of images, characterized by samples from a ``forget distribution'', while preserving the model's ability to generate other images, characterized by samples from a ``retain distribution''. While these methods aim to minimize the influence of training data in the forget distribution without extensive additional computation, we point out that they can compromise the model's integrity by inadvertently affecting generation for images in the retain distribution. Recognizing the limitations of FID and CLIPScore in capturing these effects, we introduce a novel retention metric that directly assesses the perceptual difference between outputs generated by the original and the unlearned models. We then propose unlearning algorithms that demonstrate superior effectiveness in preserving model integrity compared to existing baselines. Given their straightforward implementation, these algorithms serve as valuable benchmarks for future advancements in approximate Machine Unlearning for Diffusion Models.

cs.LG

Efficient Sketches for Training Data Attribution and Studying the Loss Landscape

The study of modern machine learning models often necessitates storing vast quantities of gradients or Hessian vector products (HVPs). Traditional sketching methods struggle to scale under these memory constraints. We present a novel framework for scalable gradient and HVP sketching, tailored for modern hardware. We provide theoretical guarantees and demonstrate the power of our methods in applications like training data attribution, Hessian spectrum analysis, and intrinsic dimension computation for pre-trained language models. Our work sheds new light on the behavior of pre-trained language models, challenging assumptions about their intrinsic dimensionality and Hessian properties.

cs.LG

Theoretical and Practical Perspectives on what Influence Functions Do

Influence functions (IF) have been seen as a technique for explaining model predictions through the lens of the training data. Their utility is assumed to be in identifying training examples "responsible" for a prediction so that, for example, correcting a prediction is possible by intervening on those examples (removing or editing them) and retraining the model. However, recent empirical studies have shown that the existing methods of estimating IF predict the leave-one-out-and-retrain effect poorly. In order to understand the mismatch between the theoretical promise and the practical results, we analyse five assumptions made by IF methods which are problematic for modern-scale deep neural networks and which concern convexity, numeric stability, training trajectory and parameter divergence. This allows us to clarify what can be expected theoretically from IF. We show that while most assumptions can be addressed successfully, the parameter divergence poses a clear limitation on the predictive power of IF: influence fades over training time even with deterministic training. We illustrate this theoretical result with BERT and ResNet models. Another conclusion from the theoretical analysis is that IF are still useful for model debugging and correcting even though some of the assumptions made in prior work do not hold: using natural language processing and computer vision tasks, we verify that mis-predictions can be successfully corrected by taking only a few fine-tuning steps on influential examples.

cs.LG

Cross-Lingual Supervision improves Large Language Models Pre-training

The recent rapid progress in pre-training Large Language Models has relied on using self-supervised language modeling objectives like next token prediction or span corruption. On the other hand, Machine Translation Systems are mostly trained using cross-lingual supervision that requires aligned data between source and target languages. We demonstrate that pre-training Large Language Models on a mixture of a self-supervised Language Modeling objective and the supervised Machine Translation objective, therefore including cross-lingual parallel data during pre-training, yields models with better in-context learning abilities. As pre-training is a very resource-intensive process and a grid search on the best mixing ratio between the two objectives is prohibitively expensive, we propose a simple yet effective strategy to learn it during pre-training.

cs.CL

Scaling Up Influence Functions

We address efficient calculation of influence functions for tracking predictions back to the training data. We propose and analyze a new approach to speeding up the inverse Hessian calculation based on Arnoldi iteration. With this improvement, we achieve, to the best of our knowledge, the first successful implementation of influence functions that scales to full-size (language and vision) Transformer models with several hundreds of millions of parameters. We evaluate our approach on image classification and sequence-to-sequence tasks with tens to a hundred of millions of training examples. Our code will be available at https://github.com/google-research/jax-influence.

cs.LG

Distributed Function Minimization in Apache Spark

We report on an open-source implementation for distributed function minimization on top of Apache Spark by using gradient and quasi-Newton methods. We show-case it with an application to Optimal Transport and some scalability tests on classification and regression problems.

cs.LG

Learning to Transport with Neural Networks

We compare several approaches to learn an Optimal Map, represented as a neural network, between probability distributions. The approaches fall into two categories: ``Heuristics'' and approaches with a more sound mathematical justification, motivated by the dual of the Kantorovitch problem. Among the algorithms we consider a novel approach involving dynamic flows and reductions of Optimal Transport to supervised learning.

cs.LG

Lipschitz functions with prescribed blowups at many points

In this paper we prove generalizations of Lusin-type theorems for gradients due to Giovanni Alberti, where we replace the Lebesgue measure with any Radon measure $μ$. We apply this to go beyond the known result on the existence of Lipschitz functions which are non-differentiable at $μ$-almost every point $x$ in any direction which is not contained in the decomposability bundle $V(μ,x)$, recently introduced by Alberti and the first named author. More precisely, we prove that it is possible to construct a Lipschitz function which attains any prescribed admissible blowup at every point except for a closed set of points of arbitrarily small measure. Here a function is an admissible blowup at a point $x$ if it is null at the origin and it is the sum of a linear function on $V(μ,x)$ and a Lipschitz function on $V(μ,x)^{\perp}$.

math.CA

Optimality of the final model found via Stochastic Gradient Descent

We study convergence properties of Stochastic Gradient Descent (SGD) for convex objectives without assumptions on smoothness or strict convexity. We consider the question of establishing that with high probability the objective evaluated at the candidate minimizer returned by SGD is close to the minimal value of the objective. We compare this result concerning the final candidate minimzer (i.e. the final model parameters learned after all gradient steps) to the online learning techniques of [Zin03] that take a rolling average of the model parameters at the different steps of SGD.

cs.LG

An example of a differentiability space which is PI-unrectifiable

We construct a (Lipschitz) differentiability space which has at generic points a disconnected tangent and thus does not contain positive measure subsets isometric to positive measure subsets of spaces admitting a Poincaré inequality. We also prove that $l^2$-valued Lipschitz maps are differentiable a.e., but there are also Lipschitz maps taking values in some other Banach spaces having the Radon-Nikodym property which fail to be differentiable on sets of positive measure.

math.MG

Unrectifiable normal currents in Euclidean spaces

We construct in $\mathbb{R}^{k+2}$ a $k$-dimensional simple normal current whose support is purely $2$-unrectifiable. The result is sharp because the support of a normal current cannot be purely $1$-unrectifiable and a $(k+1)$-dimensional normal current can be represented as an integral of $(k+1)$-rectifiable currents. This gives a negative answer to the (revised version) of a question of Frank Morgan (1984).

math.MG

Examples of $2$-unrectifiable normal currents

We construct new examples of normal (metric) currents using inverse systems of cube complexes. For any $N\ge 2$ we provide examples of $N$-dimensional normal currents whose associated vector fields are simple, and whose supports are purely $2$-unrectifiable and have Nagata dimension $N$. We show that in $l^\infty$ normal currents can be realized as limits in the flat distance of currents associated to cube complexes.

math.MG

Metric Currents and Alberti representations

We relate Ambrosio-Kirchheim metric currents to Alberti representations and Weaver derivations. In particular, given a metric current $T$, we show that if the module $\mathscr{X}(\|T\|)$ of Weaver derivations is finitely generated, then $T$ can be represented in terms of derivations; this extends previous results of Williams. Applications of this theory include an approximation of $1$-dimensional metric currents in terms of normal currents and the construction of Alberti representations in the directions of vector fields.

math.MG

Infinitesimal structure of differentiability spaces, and metric differentiation

We prove metric differentiation for differentiability spaces in the sense of Cheeger. As corollaries we give a new proof that the minimal generalized upper gradient coincides with the pointwise Lipschitz constant for Lipschitz functions on PI spaces, a proof that the Lip-lip constant of any Lip-lip space in the sense of Keith is equal to $1$, and new nonembeddability results.

math.MG

Derivations and Alberti representations

We relate generalized Lebesgue decompositions of measures in terms of curve fragments (Alberti representations) and Weaver derivations. This correspondence leads to a geometric characterization of the local norm on the Weaver cotangent bundle of a metric measure space $(X,μ)$: the local norm of a form $df$ sees how fast $f$ grows on curve fragments seen by $μ$. This implies a new characterization of differentiability spaces in terms of the $μ$-a.e.~equality of the local norm of $df$ and the local Lipschitz constant of $f$. As a consequence, the Lip-lip inequality of Keith must be an equality. We also provide dimensional bounds for the module of derivations in terms of the Assouad dimension of $X$.

math.MG

The Poincaré Inequality does not improve with blow-up

For each $β>1$ we construct a family $F_β$ of metric measure spaces which is closed under the operation of taking weak-tangents (i.e.~blow-ups), and such that each element of $F_β$ admits a $(1,P)$-Poincaré inequality if and only if $P>β$.

math.MG