SearcharxivSearch

arXiv subjects

Gil Goldman

Publications and source records attributed to Gil Goldman.

17 recordsLinked to original sources

Smooth regularization for efficient video recognition

We propose a smooth regularization technique that instills a strong temporal inductive bias in video recognition models, particularly benefiting lightweight architectures. Our method encourages smoothness in the intermediate-layer embeddings of consecutive frames by modeling their changes as a Gaussian Random Walk (GRW). This penalizes abrupt representational shifts, thereby promoting low-acceleration solutions that better align with the natural temporal coherence inherent in videos. By leveraging this enforced smoothness, lightweight models can more effectively capture complex temporal dynamics. Applied to such models, our technique yields a 3.8% to 6.4% accuracy improvement on Kinetics-600. Notably, the MoViNets model family trained with our smooth regularization improves the current state of the art by 3.8% to 6.1% within their respective FLOP constraints, while MobileNetV3 and the MoViNets-Stream family achieve gains of 4.9% to 6.4% over prior state-of-the-art models with comparable memory footprints. Our code and models are available at https://github.com/cmusatyalab/grw-smoothing.

cs.CV

Lower bounds for high derivatives of smooth functions with given zeros

Let $f: B^n \rightarrow {\mathbb R}$ be a $d+1$ times continuously differentiable function on the unit ball $B^n$, with $\max_{z\in B^n} |f(z)|=1$. A well-known fact is that if $f$ vanishes on a set $Z\subset B^n$ with a non-empty interior, then for each $k=1,\ldots,d+1$ the norm of the $k$-th derivative $\|f^{(k)}\|$ is at least $M=M(n,k)>0$. A natural question to ask is: What happens for other sets $Z$? In particular, for finite, but sufficiently dense sets?} This question was partially answered in ([16],[20-22]). This study can be naturally related to a certain special settings of the classical Whitney's smooth extension problem. Our goal in the present paper is threefold: first, to provide an overview of the relevant questions and existing results in the general Whitney's problem. Second, we provide an overview of our specific setting and some available results. Third, we provide some new results in our direction. These new results extend the recent result of [21], where an answer to the above question is given via the topological information on $Z$.

math.CA

Higher derivatives of functions with zeros on algebraic curves

Let $f: B^n \rightarrow {\mathbb R}$ be a $d+1$ times continuously differentiable function on the unit ball $B^n$, with $\max_{z\in B^n} \| f(z) \|=1$. A well-known fact is that if $f$ vanishes on a set $Z\subset B^n$ with a non-empty interior, then for each $k=1,\ldots,d+1$ the norm of the $k$-th derivative $\|f^{(k)}\|$ is at least $M=M(n,k)>0$. We show that this fact remains valid for all ``sufficiently dense'' sets $Z$ (including finite ones). The density of $Z$ is measured via the behavior of the covering numbers of $Z$. In particular, the bound $\|f^{(k)}\|\ge \tilde M=\tilde M(n,k)>0$ holds for each $Z$ with the box (or Minkowski, or entropy) dimension $\dim_e(Z)$ greater than $n-\frac{1}{k}$.

math.CA

higher derivatives of functions with given critical points and values

Let $f: B^n \rightarrow {\mathbb R}$ be a $d+1$ times continuously differentiable function on the unit ball $B^n$, with $\max_{z\in B^n} \|f(z)\|=1$. A well-known fact is that if $f$ vanishes on a set $Z\subset B^n$ with a non-empty interior, then for each $k=1,\ldots,d+1$ the norm of the $k$-th derivative $\|f^{(k)}\|$ is at least $M=M(n,k)>0$. A natural question to ask is ``what happens for other sets $Z$?''. This question was partially answered in [16]-[18]. In the present paper we ask for a similar (and closely related) question: what happens with the high-order derivatives of $f$, if its gradient vanishes on a given set $\Sigma$? And what conclusions for the high-order derivatives of $f$ can be obtained from the analysis of the metric geometry of the ``critical values set'' $f(\Sigma)$? In the present paper we provide some initial answers to these questions.

math.CA

Single-exponential bounds for the smallest singular value of Vandermonde matrices in the sub-Rayleigh regime

Following recent interest by the community, the scaling of the minimal singular value of a Vandermonde matrix with nodes forming clusters on the length scale of Rayleigh distance on the complex unit circle is studied. Using approximation theoretic properties of exponential sums, we show that the decay is only single exponential in the size of the largest cluster, and the bound holds for arbitrary small minimal separation distance. We also obtain a generalization of well-known bounds on the smallest eigenvalue of the generalized prolate matrix in the multi-cluster geometry. Finally, the results are extended to the entire spectrum.

math.NA

Exponential Taylor domination

Let $f(z) = \sum_{k=0}^\infty a_k z^k$ be an analytic function in a disk $D_R$ of radius $R>0$, and assume that $f$ is $p$-valent in $D_R$, i.e. it takes each value $c\in{\mathbb C}$ at most $p$ times in $D_R$. We consider its Borel transform $$ B(f)(z) = \sum_{k=0}^\infty \frac{a_k}{k!} z^k , $$ which is an entire function, and show that, for any $R>1$, the valency of the Borel transform $B(f)$ in $D_R$ is bounded in terms of $p,R$. We give examples, showing that our bounds, provide a reasonable envelope for the expected behavior of the valency of $B(f)$. These examples also suggest some natural questions, whose expected answer will strongly sharper our estimates. We present a short overview of some basic results on multi-valent functions, in connection with "Taylor domination", which, for $f(z) = \sum_{k=0}^\infty a_k z^k$, is a bound of all its Taylor coefficients $a_k$ through the first few of them. Taylor domination is our main technical tool, so we also discuss shortly some recent results in this direction.

math.CA

The spectral properties of Vandermonde matrices with clustered nodes

We study rectangular Vandermonde matrices $\mathbf{V}$ with $N+1$ rows and $s$ irregularly spaced nodes on the unit circle, in cases where some of the nodes are "clustered" together -- the elements inside each cluster being separated by at most $h \lesssim {1\over N}$, and the clusters being separated from each other by at least $\theta \gtrsim {1\over N}$. We show that any pair of column subspaces corresponding to two different clusters are nearly orthogonal: the minimal principal angle between them is at most $$\frac{\pi}{2}-\frac{c_1}{N \theta}-c_2 N h,$$ for some constants $c_1,c_2$ depending only on the multiplicities of theclusters. As a result, spectral analysis of $\mathbf{V}_N$ is significantly simplified by reducing the problem to the analysis of each cluster individually. Consequently we derive accurate estimates for 1) all the singular values of $\mathbf{V}$, and 2) componentwise condition numbers for the linear least squares problem. Importantly, these estimates are exponential only in the local cluster multiplicities, while changing at most linearly with $s$.

math.NA

Super-resolution of near-colliding point sources

We consider the problem of stable recovery of sparse signals of the form $$F(x)=\sum_{j=1}^d a_j\delta(x-x_j),\quad x_j\in\mathbb{R},\;a_j\in\mathbb{C}, $$ from their spectral measurements, known in a bandwidth $\Omega$ with absolute error not exceeding $\epsilon>0$. We consider the case when at most $p\le d$ nodes $\{x_j\}$ of $F$ form a cluster whose extent is smaller than the Rayleigh limit ${1\over\Omega}$, while the rest of the nodes are well separated. Provided that $\epsilon \lessapprox SRF^{-2p+1}$, where $SRF=(\Omega\Delta)^{-1}$ and $\Delta$ is the minimal separation between the nodes, we show that the minimax error rate for reconstruction of the cluster nodes is of order ${1\over\Omega}SRF^{2p-1}\epsilon$, while for recovering the corresponding amplitudes $\{a_j\}$ the rate is of the order $SRF^{2p-1}\epsilon$. Moreover, the corresponding minimax rates for the recovery of the non-clustered nodes and amplitudes are ${\epsilon\over\Omega}$ and $\epsilon$, respectively. These results suggest that stable super-resolution is possible in much more general situations than previously thought. Our numerical experiments show that the well-known Matrix Pencil method achieves the above accuracy bounds.

math.NA

Conditioning of partial nonuniform Fourier matrices with clustered nodes

We prove sharp lower bounds for the smallest singular value of a partial Fourier matrix with arbitrary "off the grid" nodes (equivalently, a rectangular Vandermonde matrix with the nodes on the unit circle), in the case when some of the nodes are separated by less than the inverse bandwidth. The bound is polynomial in the reciprocal of the so-called "super-resolution factor", while the exponent is controlled by the maximal number of nodes which are clustered together. As a corollary, we obtain sharp minimax bounds for the problem of sparse super-resolution on a grid under the partial clustering assumptions.

math.NA

Geometry and Singularities of Prony varieties

We start a systematic study of the topology, geometry and singularities of the Prony varieties $S_q(\mu)$, defined by the first $q+1$ equations of the classical Prony system $$\sum_{j=1}^d a_j x_j^k = \mu_k, \ k= 0,1,\ldots \ .$$ Prony varieties, being a generalization of the Vandermonde varieties, introduced in [5,21], present a significant independent mathematical interest (compare [5,19,21]). The importance of Prony varieties in the study of the error amplification patterns in solving Prony system was shown in [1-4,19]. In [19] a survey of these results was given, from the point of view of Singularity Theory. In the present paper we show that for $q\ge d$ the variety $S_q(\mu)$ is diffeomerphic to an intersection of a certain affine subspace in the space ${\cal V}_d$ of polynomials of degree $d$, with the hyperbolic set $H_d$. On the Prony curves $S_{2d-2}$ we study the behavior of the amplitudes $a_j$ as the nodes $x_j$ collide, and the nodes escape to infinity. We discuss the behavior of the Prony varieties as the right hand side $\mu$ varies, and possible connections of this problem with J. Mather's result in [23] on smoothness of solutions in families of linear systems.

math.NA

On algebraic properties of low rank approximations of Prony systems

We consider the reconstruction of spike train signals of the form $$F(x) = \sum_{i=1}^d a_i \delta(x-x_i),$$ from their moments measurements $m_k(F)=\int x^k F(x) dx = \sum_{i=1}^d a_ix^k$. When some of the nodes $x_i$ near collide the inversion becomes unstable. Given noisy moments measurements, a typical consequence is that reconstruction algorithms estimate the signal $F$ with a signal having fewer nodes, $\tilde{F}$. We derive lower bounds for the moments difference between a signal $F$ with $d$ nodes and a signal $\tilde{F}$ with strictly less nodes, $l$. Next we consider the geometry of the non generic case of $d$ nodes signals $F$, for which there exists an $l 2l-1 . \end{align*} We give a complete description for the case of a general $d$, $l=1$ and $p=2$. We give a reference for the case $p=2l-1$ which can be inferred from earlier work.

math.CA

Prony Scenarios and Error Amplification in a Noisy Spike-Train Reconstruction

The paper is devoted to the characterization of the geometry of Prony curves arising from spike-train signals. We give a sufficient condition which guarantees the blowing up of the amplitudes of a Prony curve S in case where some of its nodes tend to collide. We also give sufficient conditions on S which guarantee a certain asymptotic behavior of its nodes near infinity.

eess.SP

Accuracy of noisy Spike-Train Reconstruction: a Singularity Theory point of view

This is a survey paper discussing one specific (and classical) system of algebraic equations - the so called "Prony system". We provide a short overview of its unusually wide connections with many different fields of Mathematics, stressing the role of Singularity Theory. We reformulate Prony System as the problem of reconstruction of "Spike-train" signals of the form $F(x)=\sum_{j=1}^d a_j\delta(x-x_j)$ from the noisy moment measurements. We provide an overview of some recent results of [1-3, 6, 8, 9, 11, 12, 5] on the "geometry of the error amplification" in the reconstruction process, in situations where the nodes $x_j$ near-collide. Some algebraic-geometric structures, underlying the error amplification, are described (Prony, Vieta, and Hankel mappings, Prony varieties), as well as their connection with Vandermonde mappings and varieties. Our main goal is to present some promising fields of possible applications of Singulary Theory.

math.NA

Algebraic Geometry of Error Amplification: the Prony leaves

We provide an overview of some results on the "geometry of error amplification" in solving Prony system, in situations where the nodes near-collide. It turns out to be governed by the "Prony foliations" $S_q$, whose leaves are "equi-moment surfaces" in the parameter space. Next, we prove some new results concerning explicit parametrization of the Prony leaves.

math.CA

Geometry of error amplification in solving Prony system with near-colliding nodes

We consider a reconstruction problem for ``spike-train'' signals $F$ of an a priori known form $F(x)=\sum_{j=1}^{d}a_{j}\delta\left(x-x_{j}\right),$ from their moments $m_k(F)=\int x^kF(x)dx.$ We assume that the moments $m_k(F)$, $k=0,1,\ldots,2d-1$, are known with an absolute error not exceeding $\epsilon > 0$. This problem is essentially equivalent to solving the Prony system $\sum_{j=1}^d a_jx_j^k=m_k(F), \ k=0,1,\ldots,2d-1.$ We study the ``geometry of error amplification'' in reconstruction of $F$ from $m_k(F),$ in situations where the nodes $x_1,\ldots,x_d$ near-collide, i.e. form a cluster of size $h \ll 1$. We show that in this case, error amplification is governed by certain algebraic varieties in the parameter space of signals $F$, which we call the ``Prony varieties''. Based on this we produce lower and upper bounds, of the same order, on the worst case reconstruction error. In addition we derive separate lower and upper bounds on the reconstruction of the amplitudes and the nodes. Finally we discuss how to use the geometry of the Prony varieties to improve the reconstruction accuracy given additional a priori information.

math.CA

A case of multivariate Birkhoff interpolation using high order derivatives

We consider a specific scheme of multivariate Birkhoff polynomial interpolation. Our samples are derivatives of various orders $k_j$ at fixed points $v_j$ along fixed straight lines through $v_j$ in directions $u_j$, under the following assumption: the total number of sampled derivatives of order $k, \ k=0,1,\ldots$ is equal to the dimension of the space homogeneous polynomials of degree $k$. We show that this scheme is regular for general directions. Specifically this scheme is regular independent of the position of the interpolation nodes. In the planar case, we show that this scheme is regular for distinct directions. Next we prove a "Birkhoff-Remez" inequality for our sampling scheme extended to larger sampling sets. It bounds the norm of the interpolation polynomial through the norm of the samples, in terms of the geometry of the sampling set.

math.CA

Accuracy of reconstruction of spike-trains with two near-colliding nodes

We consider a signal reconstruction problem for signals $F$ of the form $ F(x)=\sum_{j=1}^{d}a_{j}δ\left(x-x_{j}\right),$ from their moments $m_k(F)=\int x^kF(x)dx.$ We assume $m_k(F)$ to be known for $k=0,1,\ldots,N,$ with an absolute error not exceeding $ε> 0$. We study the "geometry of error amplification" in reconstruction of $F$ from $m_k(F),$ in situations where two neighboring nodes $x_i$ and $x_{i+1}$ near-collide, i.e $x_{i+1}-x_i=h \ll 1$. We show that the error amplification is governed by certain algebraic curves $S_{F,i},$ in the parameter space of signals $F$, along which the first three moments $m_0,m_1,m_2$ remain constant.

math.CA