SearcharxivSearch

arXiv subjects

Yizhe Zhu

Publications and source records attributed to Yizhe Zhu.

At least 19 recordsLinked to original sources

Sharp spectral norm concentration of sparse random tensors

We prove a sharp concentration inequality for the spectral norm of sparse random tensors with independent Bernoulli entries. Let $T$ be an order-$k$ tensor of dimension $n\times\cdots\times n$ with independent Bernoulli$(p)$ entries, where $k$ is fixed. For any $c,r>0$, we show that $\|T-\mathbb E T\|\le C_{k,r,c}\sqrt{np}$ with probability at least $1-n^{-r}$ whenever $np\ge c\log n$. We extend this bound to inhomogeneous Bernoulli sampling with deterministic entrywise weights. This removes the logarithmic factor in the work of Zhou and Zhu (2021). The proof follows the Kahn--Szemerédi light--heavy decomposition with a refined estimate on the heavy tuple part. We also obtain a log-free second eigenvalue bound for the random hypergraph model of Friedman and Wigderson (1995).

math.PR

The spectral edge of sparse directed Erdös-Rényi graphs

Let $d>1$ be fixed and let $A_n$ be an $n\times n$ matrix with independent $\Ber(d/n)$ entries. For every $0<r<\sqrt d$, we prove that, with high probability, a positive proportion of the eigenvalues of $A_n$ have modulus larger than $r$. Together with the known upper bound, this implies that the modulus of the second largest eigenvalue converges in probability to $\sqrt d$. Our proof works with the Brown measure $μ_d$ of the adjacency operator of the directed Poisson--Galton--Watson tree. The convergence theorem of Sah, Sahasrabudhe, and Sawhney, together with the Brown-measure identification following it, gives weak convergence of the empirical spectral measure of $A_n$ to $μ_d$ in probability. We prove that the outer radius of its support is $\sqrt d$, using a resolvent recursion on Poisson--Galton--Watson trees.

math.PR

A semicircle law for the normalized Laplacian of sparse random graphs

We study the limiting spectral distribution of the normalized Laplacian $\mathcal L$ of an Erdős-Rényi graph $G(n,p)$. To account for the presence of isolated vertices in the sparse regime, we define $\mathcal L$ using the Moore-Penrose pseudoinverse of the degree matrix. Under this convention, we show that the empirical spectral distribution of a suitably normalized $\mathcal L$ converges weakly in probability to the semicircle law whenever $np\to\infty$, thereby providing a rigorous justification of a prediction made in (Akara-pipattana and Evnin, 2023). Moreover, if $np>\log n+ω(1)$, so that $G(n,p)$ has no isolated vertices with high probability, the same conclusion holds for the standard definition of $\mathcal L$. We further strengthen this result to almost sure convergence when $np=Ω(\log n)$. Finally, we extend our approach to the Chung-Lu random graph model, where we establish a semicircle law for $\mathcal L$ itself, improving upon (Chung, Lu, and Vu 2003), which obtained the semicircle law only for a proxy matrix.

math.PR

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly focus on task-level evaluation and fail to provide actionable insights into the underlying causes of model failures. To address this limitation, we introduce BEAR, a benchmark that decomposes embodied tasks into 14 atomic skills for fine-grained skill-level evaluation. BEAR comprises 4,469 interleaved image-video-text samples spanning 14 skills across 6 categories, ranging from low-level perception to high-level planning. We evaluate 20 MLLMs on BEAR under a hierarchical skill-level diagnosis framework and uncover two key findings: (1) perceptual capabilities are major bottlenecks behind reasoning failures, and (2) current models suffer from unstable spatiotemporal modeling that remains largely unexposed in prior benchmarks. Motivated by these findings, we further propose BEAR-Agent, a multimodal conversational agent that augments MLLMs with visual and spatial reasoning tools. BEAR-Agent substantially improves performance across embodied skills, achieving a relative improvement of 17.5% on GPT-5 over the base model on BEAR, while also outperforming strong baselines in both simulation and real-world robotic experiments. Project page: https://bear-official66.github.io/

cs.CV

Wedge Sampling: Efficient Tensor Completion with Nearly-Linear Sample Complexity

We introduce Wedge Sampling, a new non-adaptive sampling scheme for low-rank tensor completion. We study recovery of an order-$k$ low-rank tensor of dimension $n\times\cdots\times n$ from structured observations of its entries. Unlike the standard uniform entry model (i.e., i.i.d. samples from $[n]^k$), wedge sampling allocates observations to structured length-two patterns (wedges) in an associated bipartite sampling graph. By directly promoting these length-two connections, the sampling design strengthens the spectral signal that underlies efficient initialization, in regimes where uniform sampling is too sparse to generate enough informative correlations. Our main result shows that this change in sampling paradigm enables polynomial-time algorithms to achieve both weak and exact recovery with nearly linear sample complexity in $n$. The approach is also plug-and-play: wedge-sampling-based spectral initialization can be combined with existing refinement procedures (e.g., spectral or gradient-based methods) using only an additional $\tilde O(n)$ uniformly sampled entries, substantially improving over the $\tilde O(n^{k/2})$ sample complexity typically required under uniform entry sampling for efficient methods. We also formulate a noisy wedge-sampling extension for additive Gaussian observations and analyze both the spectral and gradient-descent procedures under suitable signal-to-noise conditions. Thus, the computational barrier in tensor completion is sensitive to the observation model: while it persists under uniform entry sampling, it can be bypassed by non-adaptive structured designs that provide a stronger initialization.

stat.ML

Spectral Concentration and Recovery in Sparse High-Dimensional Random Geometric Graphs

We study sparse threshold random geometric graphs generated by high-dimensional spherical or Gaussian latent vectors. Although each edge has marginal probability $p$, shared latent variables make the adjacency entries dependent. At the connectivity scale $np=Ω(\log n)$, the spherical adjacency matrix satisfies, with high probability,$\|A-\mathbb E A\|_{\mathrm{op}}=O\left(\sqrt{np\log n}+npτ\right)$, where $τ$ is the cap threshold; an analogous estimate holds for Gaussian vectors after controlling radial fluctuations. This sharpens the spectral bound in Liu, Mohanty, Schramm, and Yang (2023) under weaker assumptions and strengthens the global-synchronization guarantee of Abdalla, Bandeira, and Invernizzi (2024) for the homogeneous Kuramoto model. The leading eigenspace also estimates the latent geometry. When $np\gg\log n$, vector and relative Gram-matrix errors vanish for$\log(1/p)\ll d\ll np\log(1/p)/\log n$ in the spherical model and $\log^2(1/p)\log n\ll d\ll np\log(1/p)/\log n$ in the Gaussian model, improving the recovery conditions of Li and Schramm (2023). For the Gaussian mixture block model introduced there, a polynomial-time semidefinite program gives, to our knowledge, the first exact-recovery guarantee at the connectivity scale in a moderate-separation regime. At much larger separation, fixed edge density creates isolated vertices and makes exact recovery impossible. Our reusable decoupling and matrix concentration framework avoids trace-moment methods and applies broadly to random graph models with latent vectors.

stat.ML

Spectra of high-dimensional sparse random geometric graphs

We determine the limiting empirical spectral distribution of sparse high-dimensional random geometric graphs. The vertices are independent uniform points on the unit sphere $S^{d-1}$, and two vertices are joined when their inner product exceeds a threshold chosen to give edge density $p$. The edges therefore have the same marginal probabilities as in an Erdős--Rényi graph, but the latent geometry introduces dependence among them. We show that these correlations are asymptotically invisible to the global spectrum in two sparse regimes. If $p\to0$, $np\to\infty$, and $d=Ω(np\log(1/p))$, then the empirical spectral distribution of $A/\sqrt{np}$ converges in probability to the semicircle law. If $p=α/n$ for a fixed $α>0$ and $d=ω(\log n)$, then the empirical spectral distribution of $A/\sqrtα$ converges in probability to the limiting spectral distribution of $\mathcal G(n,α/n)$. The proof combines the moment method with a cluster expansion that decomposes geometric dependence into weak local interactions, allowing us to control every fixed walk pattern in the moment calculation.

math.PR

Minimax optimal differentially private synthetic data for smooth queries

Differentially private synthetic data enables the sharing and analysis of sensitive datasets while providing rigorous privacy guarantees for individual contributors. A central challenge is to achieve strong utility guarantees for meaningful downstream analysis. Many existing methods ensure uniform accuracy over broad query classes, such as all Lipschitz functions, but this level of generality often leads to suboptimal rates for statistics of practical interest. Since many common data analysis queries exhibit smoothness beyond what worst-case Lipschitz bounds capture, we ask whether exploiting this additional structure can yield improved utility. We study the problem of generating $(\varepsilon,δ)$-differentially private synthetic data from a dataset of size $n$ supported on the hypercube $[-1,1]^d$, with utility guarantees uniformly for all smooth queries having bounded derivatives up to order $k$. We propose a polynomial-time algorithm that achieves a minimax error rate of $O_{k,d}(n^{-\min \{1, \frac{k}{d}\}})$, up to a $\log(n)$ factor. This characterization uncovers a phase transition at $k=d$. Our results generalize the Chebyshev moment matching framework of (Musco et al., 2025; Wang et al., 2016) and strictly improve the error rates for $k$-smooth queries established in \citep{wang2016differentially}. Moreover, we establish the first minimax lower bound for the utility of $(\varepsilon,δ)$-differentially private synthetic data with respect to $k$-smooth queries, extending the Wasserstein lower bound for $\varepsilon$-differential privacy in (Boedihardjo et al., 2024).

math.ST

Partial recovery and weak consistency in the non-uniform hypergraph Stochastic Block Model

We consider the community detection problem in sparse random hypergraphs under the non-uniform hypergraph stochastic block model (HSBM), a general model of random networks with community structure and higher-order interactions. When the random hypergraph has bounded expected degrees, we provide a spectral algorithm that outputs a partition with at least a $γ$ fraction of the vertices classified correctly, where $γ\in (0.5,1)$ depends on the signal-to-noise ratio (SNR) of the model. When the SNR grows slowly as the number of vertices goes to infinity, our algorithm achieves weak consistency, which improves the previous results in Ghoshdastidar and Dukkipati (2017) for non-uniform HSBMs. Our spectral algorithm consists of three major steps: (1) Hyperedge selection: select hyperedges of certain sizes to provide the maximal signal-to-noise ratio for the induced sub-hypergraph; (2) Spectral partition: construct a regularized adjacency matrix and obtain an approximate partition based on singular vectors; (3) Correction and merging: incorporate the hyperedge information from adjacency tensors to upgrade the error rate guarantee. The theoretical analysis of our algorithm relies on the concentration and regularization of the adjacency matrix for sparse non-uniform random hypergraphs, which can be of independent interest.

math.ST

Values of finite distortion: Reshetnyak's theorem, the Liouville theorem, and the Lusin (N) -property

Let $Ω\subset \mathbb{R}^n$ be a domain and $f \in W^{1,n}_{\text{loc}} (Ω,\mathbb{R}^n)$. We say that $f$ has a value of finite distortion at $y_0 \in \mathbb{R}^n$ if there exist measurable functions $K \colon Ω\to [0,\infty)$ and $Σ\in L^1_{\text{loc}} (Ω)$ such that \[ \lvert Df(x) \rvert^n \le K(x) \det Df (x) + Σ(x) \lvert f(x)-y_0 \rvert^n \quad \text{for a.e. } x \in Ω. \] This notion unifies the classical theory of mappings of finite distortion with the recently introduced theory of quasiregular values. Under sharp integrability assumptions on $K$ and $Σ$, we establish single-value analogues of Reshetnyak's theorem and the Liouville theorem. We also prove that mappings satisfying a more general distortion inequality with defect preserve sets of Lebesgue measure zero.

math.CV

Achieving the Kesten-Stigum bound in the non-uniform hypergraph stochastic block model

We study the community detection problem in the non-uniform hypergraph stochastic block model (HSBM), where hyperedges of varying sizes coexist. This setting captures higher-order and multi-view interactions and raises a fundamental question: can multiple uniform hypergraph layers below the detection threshold be combined to enable weak recovery? We answer this question by establishing a Kesten--Stigum-type bound for weak recovery in a general class of non-uniform HSBMs with $r$ blocks, generated according to multiple symmetric probability tensors. In the case $r=2$, we show that weak recovery is possible whenever the sum of the signal-to-noise ratios across all uniform hypergraph layers exceeds one, thereby confirming the positive part of a conjecture in (Chodrow et al., 2023). Moreover, we provide a polynomial-time spectral algorithm that achieves this threshold via an optimally weighted non-backtracking operator. For the unweighted non-backtracking matrix, our spectral method attains a different algorithmic threshold, also conjectured in (Chodrow et al., 2023). Our approach develops a spectral theory for weighted non-backtracking operators on non-uniform hypergraphs, including a precise characterization of outlier eigenvalues and eigenvector overlaps. We introduce a novel Ihara--Bass formula tailored to weighted non-uniform hypergraphs, which yields an efficient low-dimensional representation and leads to a provable spectral reconstruction algorithm. Taken together, these results provide a principled and computationally efficient approach to clustering in non-uniform hypergraphs, and highlight the role of optimal weighting in aggregating heterogeneous higher-order interactions.

stat.ML

SMAL-pets: SMAL Based Avatars of Pets from Single Image

Creating high-fidelity, animatable 3D dog avatars remains a formidable challenge in computer vision. Unlike human digital doubles, animal reconstruction faces a critical shortage of large-scale, annotated datasets for specialized applications. Furthermore, the immense morphological diversity across species, breeds, and crosses, which varies significantly in size, proportions, and features, complicates the generalization of existing models. Current reconstruction methods often struggle to capture realistic fur textures. Additionally, ensuring these avatars are fully editable and capable of performing complex, naturalistic movements typically necessitates labor-intensive manual mesh manipulation and expert rigging. This paper introduces SMAL-pets, a comprehensive framework that generates high-quality, editable animal avatars from a single input image. Our approach bridges the gap between reconstruction and generative modeling by leveraging a hybrid architecture. Our method integrates 3D Gaussian Splatting with the SMAL parametric model to provide a representation that is both visually high-fidelity and anatomically grounded. We introduce a multimodal editing suite that enables users to refine the avatar's appearance and execute complex animations through direct textual prompts. By allowing users to control both the aesthetic and behavioral aspects of the model via natural language, SMAL-pets provides a flexible, robust tool for animation and virtual reality.

cs.CV

Power-Law Spectrum of the Random Feature Model

Scaling laws for neural networks, in which the loss decays as a power-law in the number of parameters, data, and compute, depend fundamentally on the spectral structure of the data covariance, with power-law eigenvalue decay appearing ubiquitously in vision and language tasks. A central question is whether this spectral structure is preserved or destroyed when data passes through the basic building block of a neural network: a random linear projection followed by a nonlinear activation. We study this question for the random feature model: given data $x \sim N(0,H)\in \mathbb{R}^v$ where $H$ has $α$-power-law spectrum ($λ_j(H ) \asymp j^{-α}$, $α> 1$), a Gaussian sketch matrix $W \in \mathbb{R}^{v\times d}$, and an entrywise monomial $f(y) = y^{p}$, we characterize the eigenvalues of the population random-feature covariance $\mathbb{E}_{x }[\frac{1}{d}f(W^\top x )^{\otimes 2}]$. We prove matching upper and lower bounds: for all $1 \leq j \leq c_1 d \log^{-(p+1)}(d)$, the $j$-th eigenvalue is of order $\left(\log^{p-1}(j+1)/j\right)^α$. For $ c_1 d \log^{-(p+1)}(d)\leq j\leq d$, the $j$-th eigenvalue is of order $j^{-α}$ up to a polylog factor. That is, the power-law exponent $α$ is inherited exactly from the input covariance, modified only by a logarithmic correction that depends on the monomial degree $p$. The proof combines a dyadic head-tail decomposition with Wick chaos expansions for higher-order monomials and random matrix concentration inequalities.

stat.ML

Sparse Hanson-Wright Inequalities with Applications

We derive new Hanson-Wright-type inequalities tailored to the quadratic forms of random vectors with sparse independent components. Specifically, we consider cases where the components of the random vector are sparse $α$-subexponential random variables with $α>0$. When $α=\infty$, these inequalities can be seen as quadratic generalizations of the classical Bernstein and Bennett inequalities for sparse bounded random vectors. To establish this quadratic generalization, we also develop new Bernstein-type and Bennett-type inequalities for linear forms of sparse $α$-subexponential random variables that go beyond the bounded case $(α=\infty)$. Our proof relies on a novel combinatorial method for estimating the moments of both random linear forms and quadratic forms. We present two key applications of these new sparse Hanson-Wright inequalities: (1) A local law and complete eigenvector delocalization for sparse $α$-subexponential Hermitian random matrices, generalizing the result of He et al. (2019) beyond sparse Bernoulli random matrices. To the best of our knowledge, this is the first local law and complete delocalization result for sparse $α$-subexponential random matrices down to the near-optimal sparsity $p\geq \frac{\mathrm{polylog}(n)}{n}$ when $α\in (0,2)$ as well as for unbounded sparse sub-gaussian random matrices down to the optimal sparsity $p\gtrsim \frac{\log n}{n}.$ (2) Concentration of the Euclidean norm for the linear transformation of a sparse $α$-subexponential random vector, improving on the results of G{ö}tze et al. (2021) for sparse sub-exponential random vectors.

math.PR

Homeomorphic Extensions in Bi-Orlicz-Sobolev Spaces

We provide a complete characterization of those self-homeomorphisms of the unit circle that admit homeomorphic extensions to the unit disk belonging to bi--Orlicz--Sobolev spaces. Our results generalize classical criteria from the Sobolev setting to the more flexible Orlicz framework.

math.CV

Global Integrability of the Reciprocal of Jacobians for Homeomorphisms of Finite Distortion

For a homeomorphism with $p$-integrable distortion, we obtain the optimal global degree of integrability for the reciprocal of its Jacobian determinant. As an application, we strengthen the result of Doležalová, Hencl and Malý concerning weak limits of Sobolev homeomorphisms with finite distortion. Such limits represent physically admissible deformations, as they remain injective almost everywhere and thus adhere as closely as possible to the principle of non-interpenetration of matter in mathematical models of nonlinear elasticity.

math.FA

Residual Rotation Correction using Tactile Equivariance

Visuotactile policy learning augments vision-only policies with tactile input, facilitating contact-rich manipulation. However, the high cost of tactile data collection makes sample efficiency the key requirement for developing visuotactile policies. We present EquiTac, a framework that exploits the inherent SO(2) symmetry of in-hand object rotation to improve sample efficiency and generalization for visuotactile policy learning. EquiTac first reconstructs surface normals from raw RGB inputs of vision-based tactile sensors, so rotations of the normal vector field correspond to in-hand object rotations. An SO(2)-equivariant network then predicts a residual rotation action that augments a base visuomotor policy at test time, enabling real-time rotation correction without additional reorientation demonstrations. On a real robot, EquiTac accurately achieves robust zero-shot generalization to unseen in-hand orientations with very few training samples, where baselines fail even with more training data. To our knowledge, this is the first tactile learning method to explicitly encode tactile equivariance for policy learning, yielding a lightweight, symmetry-aware module that improves reliability in contact-rich tasks.

cs.RO

Generalizable Hierarchical Skill Learning via Object-Centric Representation

We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that significantly improves policy generalization and sample efficiency in robot manipulation. One core idea of GSL is to use object-centric skills as an interface that bridges the high-level vision-language model and the low-level visual-motor policy. Specifically, GSL decomposes demonstrations into transferable and object-canonicalized skill primitives using foundation models, ensuring efficient low-level skill learning in the object frame. At test time, the skill-object pairs predicted by the high-level agent are fed to the low-level module, where the inferred canonical actions are mapped back to the world frame for execution. This structured yet flexible design leads to substantial improvements in sample efficiency and generalization of our method across unseen spatial arrangements, object appearances, and task compositions. In simulation, GSL trained with only 3 demonstrations per task outperforms baselines trained with 30 times more data by 15.5 percent on unseen tasks. In real-world experiments, GSL also surpasses the baseline trained with 10 times more data.

cs.RO