Searcharxiv⌕ Search

arXiv subjects

Yi Han

Publications and source records attributed to Yi Han.

At least 73 records · Page 4Linked to original sources

Confidence Diagram of Nonparametric Ranking for Uncertainty Assessment in Large Language Models Evaluation

We consider the inference for the ranking of large language models (LLMs). Alignment arises as a significant challenge to mitigate hallucinations in the use of LLMs. Ranking LLMs has proven to be an effective tool to improve alignment based on the best-of-$N$ policy. In this paper, we propose a new inferential framework for hypothesis testing among the ranking for language models. Our framework is based on a nonparametric contextual ranking framework designed to assess large language models' domain-specific expertise, leveraging nonparametric scoring methods to account for their sensitivity to the prompts. To characterize the combinatorial complexity of the ranking, we introduce a novel concept of confidence diagram, which leverages a Hasse diagram to represent the entire confidence set of rankings by a single directed graph. We show the validity of the proposed confidence diagram by advancing the Gaussian multiplier bootstrap theory to accommodate the supremum of independent empirical processes that are not necessarily identically distributed. Extensive numerical experiments conducted on both synthetic and real data demonstrate that our approach offers valuable insight into the evaluation for the performance of different LLMs across various medical domains.

stat.ML↗

XiHe: A Data-Driven Model for Global Ocean Eddy-Resolving Forecasting

The leading operational Global Ocean Forecasting Systems (GOFSs) use physics-driven numerical forecasting models that solve the partial differential equations with expensive computation. Recently, specifically in atmosphere weather forecasting, data-driven models have demonstrated significant potential for speeding up environmental forecasting by orders of magnitude, but there is still no data-driven GOFS that matches the forecasting accuracy of the numerical GOFSs. In this paper, we propose the first data-driven 1/12° resolution global ocean eddy-resolving forecasting model named XiHe, which is established from the 25-year France Mercator Ocean International's daily GLORYS12 reanalysis data. XiHe is a hierarchical transformer-based framework coupled with two special designs. One is the land-ocean mask mechanism for focusing exclusively on the global ocean circulation. The other is the ocean-specific block for effectively capturing both local ocean information and global teleconnection. Extensive experiments are conducted under satellite observations, in situ observations, and the IV-TT Class 4 evaluation framework of the world's leading operational GOFSs from January 2019 to December 2020. The results demonstrate that XiHe achieves stronger forecast performance in all testing variables than existing leading operational numerical GOFSs including Mercator Ocean Physical SYstem (PSY4), Global Ice Ocean Prediction System (GIOPS), BLUElinK OceanMAPS (BLK), and Forecast Ocean Assimilation Model (FOAM). Particularly, the accuracy of ocean current forecasting of XiHe out to 60 days is even better than that of PSY4 in just 10 days. Additionally, XiHe is able to forecast the large-scale circulation and the mesoscale eddies. Furthermore, it can make a 10-day forecast in only 0.35 seconds, which accelerates the forecast speed by thousands of times compared to the traditional numerical GOFSs.

physics.ao-ph↗

Change Point Detection in Pairwise Comparison Data with Covariates

This paper introduces the novel piecewise stationary covariate-assisted ranking estimation (PS-CARE) model for analyzing time-evolving pairwise comparison data, enhancing item ranking accuracy through the integration of covariate information. By partitioning the data into distinct, stationary segments, the PS-CARE model adeptly detects temporal shifts in item rankings, known as change points, whose number and positions are initially unknown. Leveraging the minimum description length (MDL) principle, this paper establishes a statistically consistent model selection criterion to estimate these unknowns. The practical optimization of this MDL criterion is done with the pruned exact linear time (PELT) algorithm. Empirical evaluations reveal the method's promising performance in accurately locating change points across various simulated scenarios. An application to an NBA dataset yielded meaningful insights that aligned with significant historical events, highlighting the method's practical utility and the MDL criterion's effectiveness in capturing temporal ranking changes. To the best of the authors' knowledge, this research pioneers change point detection in pairwise comparison data with covariate information, representing a significant leap forward in the field of dynamic ranking analysis.

stat.AP↗

Deviation of top eigenvalue for some tridiagonal matrices under various moment assumptions

Symmetric tridiagonal matrices appear ubiquitously in mathematical physics, serving as the matrix representation of discrete random Schrödinger operators. In this work we investigate the top eigenvalue of these matrices in the large deviation regime, assuming the random potentials are on the diagonal with a certain decaying factor $N^{-α}$, and the probability law $μ$ of the potentials satisfy specific decay assumptions. We investigate two different models, one of which has random matrix behavior at the spectral edge but the other does not. Both the light-tailed regime, i.e. when $μ$ has all moments, and the heavy-tailed regime are covered. Precise right tail estimates and a crude left tail estimate are derived. In particular we show that when the tail $μ$ has a certain decay rate, then the top eigenvalue is distributed as the Frechet law composed with some deterministic functions. The proof relies on computing one point perturbations of fixed tridiagonal matrices.

math.PR↗

Deformed Fréchet law for Wigner and sample covariance matrices with tail in crossover regime

Given $A_n:=\frac{1}{\sqrt{n}}(a_{ij})$ an $n\times n$ symmetric random matrix, with elements above the diagonal given by i.i.d. random variables having mean zero and unit variance. It is known that when $\lim_{x\to\infty}x^4\mathbb{P}(|a_{ij}|>x)=0$, then fluctuation of the largest eigenvalue of $A_n$ follows a Tracy-Widom distribution. When the law of $a_{ij}$ is regularly varying with index $α\in(0,4)$, then the largest eigenvalue has a Fréchet distribution. An intermediate regime is recently uncovered in \cite{diaconu2023more}: when $\lim_{x\to\infty}x^4\mathbb{P}(|a_{ij}|>x)=c\in(0,\infty)$, then the law of the largest eigenvalue follows a deformed Fréchet distribution. In this work we vastly extend the scope where the latter distribution may arise. We show that the same deformed Fréchet distribution arises (1) for sparse Wigner matrices with an average of $n^{O(1)}$ nonzero entries on each row; (2) for periodically banded Wigner matrices with bandwidth $d_n=n^{O(1)}$; and more generally for weighted adjacency matrices of any $k_n$-regular graphs with $k_n=n^{O(1)}$. In all these cases, we further prove that the joint distribution of the finitely many largest eigenvalues of $A_n$ form a deformed Poisson process, and that eigenvectors of the outlying eigenvalues of $A_n$ are localized, implying a mobility edge phenomenon at the spectral edge $2$. The sparser case with average degree $n^{o(1)}$ is also explored. Our technique extends to sample covariance matrices, proving for the first time that its largest eigenvalue still follows a deformed Fréchet distribution, assuming the matrix entries satisfy $\lim_{x\to\infty}x^4\mathbb{P}(|a_{ij}|>x)=c\in(0,\infty)$.

math.PR↗

Small ball probability for multiple singular values of symmetric random matrices

Let $A_n$ be an $n\times n$ random symmetric matrix with $(A_{ij})_{i< j}$ i.i.d. mean $0$, variance 1, following a subGaussian distribution and diagonal elements i.i.d. following a subGaussian distribution with a fixed variance. We investigate the joint small ball probability that $A_n$ has eigenvalues near two fixed locations $λ_1$ and $λ_2$, where $λ_1$ and $λ_2$ are sufficiently separated and in the bulk of the semicircle law. More precisely we prove that for a wide class of entry distributions of $A_{ij}$ that involve all Gaussian convolutions (where $σ_{min}(\cdot)$ denotes the least singular value of a square matrix), $$\mathbb{P}(σ_{min}(A_n-λ_1 I_n)\leqδ_1n^{-1/2},σ_{min}(A_n-λ_2 I_n)\leqδ_2n^{-1/2})\leq cδ_1δ_2+e^{-cn}.$$ The given estimate approximately factorizes as the product of the estimates for the two individual events, which is an indication of quantitative independence. The estimate readily generalizes to $d$ distinct locations. As an application, we upper bound the probability that there exist $d$ eigenvalues of $A_n$ asymptotically satisfying any fixed linear equation, which in particular gives a lower bound of the distance to this linear relation from any possible eigenvalue pair that holds with probability $1-o(1)$, and rules out the existence of two equal singular values in generic regions of the spectrum.

math.PR↗

Robust angle-based transfer learning in high dimensions

Transfer learning aims to improve the performance of a target model by leveraging data from related source populations, which is known to be especially helpful in cases with insufficient target data. In this paper, we study the problem of how to train a high-dimensional ridge regression model using limited target data and existing regression models trained in heterogeneous source populations. We consider a practical setting where only the parameter estimates of the fitted source models are accessible, instead of the individual-level source data. Under the setting with only one source model, we propose a novel flexible angle-based transfer learning (angleTL) method, which leverages the concordance between the source and the target model parameters. We show that angleTL unifies several benchmark methods by construction, including the target-only model trained using target data alone, the source model fitted on source data, and distance-based transfer learning method that incorporates the source parameter estimates and the target data under a distance-based similarity constraint. We also provide algorithms to effectively incorporate multiple source models accounting for the fact that some source models may be more helpful than others. Our high-dimensional asymptotic analysis provides interpretations and insights regarding when a source model can be helpful to the target model, and demonstrates the superiority of angleTL over other benchmark methods. We perform extensive simulation studies to validate our theoretical conclusions and show the feasibility of applying angleTL to transfer existing genetic risk prediction models across multiple biobanks.

stat.ME↗

Stochastic wave equation with Hölder noise coefficient: well-posedness and small mass limit

We construct unique martingale solutions to the damped stochastic wave equation $$ μ\frac{\partial^2u}{\partial t^2}(t,x)=Δu(t,x)-\frac{\partial u}{\partial t}(t,x)+b(t,x,u(t,x))+σ(t,x,u(t,x))\frac{dW_t}{dt},$$ where $Δ$ is the Laplacian on $[0,1]$ with Dirichlet boundary condition, $W$ is space-time white noise, $σ$ is $\frac{3}{4}+ε$ -Hölder continuous in $u$ and uniformly non-degenerate, and $b$ has linear growth. The same construction holds for the stochastic wave equation without damping term. More generally, the construction holds for SPDEs defined on separable Hilbert spaces with a densely defined operator $A$, and the assumed Hölder regularity on the noise coefficient depends on the eigenvalues of $A$ in a quantitative way. We further show the validity of the Smoluchowski-Kramers approximation: assume $b$ is Hölder continuous in $u$, then as $μ$ tends to $0$ the solution to the damped stochastic wave equation converges in distribution, on the space of continuous paths, to the solution of the corresponding stochastic heat equation. The latter result is new even in the case of additive noise.

math.PR↗

A support theorem for parabolic stochastic PDEs with nondegenerate Hölder diffusion coefficients

In this paper we work with parabolic SPDEs of the form $$ \partial_t u(t,x)=\partial_x^2 u(t,x)+g(t,x,u)+σ(t,x,u)\dot{W}(t,x) $$ with Neumann boundary conditions, where $x\in[0,1]$, $\dot{W}(t,x)$ is the space-time white noise on $(t,x)\in[0,\infty)\times [0,1]$, $g$ is uniformly bounded, and the solution $u\in\mathbb{R}$ is real valued. The diffusion coefficient $σ$ is assumed to be uniformly elliptic but only Hölder continuous in $u$. Previously, support theorems for SPDEs have only been established assuming that $σ$ is Lipschitz continuous in $u$. We obtain new support theorems and small ball probabilities in this $σ$ Hölder continuous case via the recently established sharp two sided estimates of stochastic integrals.

math.PR↗

Through the Lens of Core Competency: Survey on Evaluation of Large Language Models

From pre-trained language model (PLM) to large language model (LLM), the field of natural language processing (NLP) has witnessed steep performance gains and wide practical uses. The evaluation of a research field guides its direction of improvement. However, LLMs are extremely hard to thoroughly evaluate for two reasons. First of all, traditional NLP tasks become inadequate due to the excellent performance of LLM. Secondly, existing evaluation tasks are difficult to keep up with the wide range of applications in real-world scenarios. To tackle these problems, existing works proposed various benchmarks to better evaluate LLMs. To clarify the numerous evaluation tasks in both academia and industry, we investigate multiple papers concerning LLM evaluations. We summarize 4 core competencies of LLM, including reasoning, knowledge, reliability, and safety. For every competency, we introduce its definition, corresponding benchmarks, and metrics. Under this competency architecture, similar tasks are combined to reflect corresponding ability, while new tasks can also be easily added into the system. Finally, we give our suggestions on the future direction of LLM's evaluation.

cs.CL↗

Entropic propagation of chaos for mean field diffusion with $L^p$ interactions via hierarchy, linear growth and fractional noise

New quantitative propagation of chaos results for mean field diffusion are proved via local and global entropy estimates. In the first result we work on the torus and consider singular, divergence free interactions $K\in L^p$, $p>d$. We prove a $O(k^{2}/n^2)$ convergence rate in relative entropy between the $k$-marginal laws of the particle system and its limiting law at each time $t$, as long as the same holds at time 0. The proof is based on local estimates via a form of BBGKY hierarchy and exemplifies a method to extend the framework in Lacker [16] to singular interactions. The rate can be made uniform in time combined with the result in [18]. Then we prove quantitative propagation of chaos for interactions that are only assumed to have linear growth. This generalizes to the case where the driving noise is replaced by a fractional Brownian motion $B^H$, for all $H\in(0,1)$. These proofs follow from global estimates and subGaussian concentration inequalities. We obtain $O(k/n)$ convergence rate in relative entropy in each case, yet the rate is only valid on $[0,T^*]$ with $T^*$ a fixed finite constant depending on various parameters of the system.

math.PR↗

Universal edge scaling limit of discrete 1d random Schrödinger operator with vanishing potentials

Consider random Schrödinger operators $H_n$ defined on $[0,n]\cap\mathbb{Z}$ with zero boundary conditions: $$ (H_nψ)_\ell=ψ_{\ell-1}+ψ_{\ell+1}+σ\frac{\mathfrak{a}(\ell)}{n^α}ψ_{\ell},\quad \ell=1,\cdots,n,\quad \quad ψ_{0}=ψ_{n+1}=0, $$ where $σ>0$ is a fixed constant, $\mathfrak{a}(\ell)$, $\ell=1,\cdots,n$, are i.i.d. random variables with mean $0$, variance $1$ and fast decay. The bulk scaling limit has been investigated in \cite{kritchevski2011scaling}: at the critical exponent $α= \frac{1}{2}$, the spectrum of $H_n$, centered at $E\in(-2,2)\setminus\{0\}$ and rescaled by $n$, converges to the $\operatorname{Sch}_τ$ process and does not depend on the distribution of $\mathfrak{a}(\ell).$ We study the scaling limit at the edge. We show that at the critical value $α=\frac{3}{2}$, if we center the spectrum at 2 and rescale by $n^2$, then the spectrum converges to a new random process depending on $σ$ but not the distribution of $\mathfrak{a}(\ell)$. We use two methods to describe this edge scaling limit. The first uses the method of moments, where we compute the Laplace transform of the point process, and represent it in terms of integrated local times of Brownian bridges. Then we show that the rescaled largest eigenvalues correspond to the lowest eigenvalues of the random Schrödinger operator $-\frac{d^2}{dx^2}+σb_x'$ defined on $[0,1]$ with zero boundary condition, where $b_x$ is a standard Brownian motion. This allows us to compute precise left and right tails of the rescaled largest eigenvalue and compare them to Tracy-Widom beta laws. We also show if we shift the potential $\mathfrak{a}(\ell)$ by a state-dependent constant and take $α=\frac{1}{2}$, then for a particularly chosen state-dependent shift, the rescaled largest eigenvalues converge to the Tracy-Widom beta distribution.

math.PR↗

Why Don't You Clean Your Glasses? Perception Attacks with Dynamic Optical Perturbations

Camera-based autonomous systems that emulate human perception are increasingly being integrated into safety-critical platforms. Consequently, an established body of literature has emerged that explores adversarial attacks targeting the underlying machine learning models. Adapting adversarial attacks to the physical world is desirable for the attacker, as this removes the need to compromise digital systems. However, the real world poses challenges related to the "survivability" of adversarial manipulations given environmental noise in perception pipelines and the dynamicity of autonomous systems. In this paper, we take a sensor-first approach. We present EvilEye, a man-in-the-middle perception attack that leverages transparent displays to generate dynamic physical adversarial examples. EvilEye exploits the camera's optics to induce misclassifications under a variety of illumination conditions. To generate dynamic perturbations, we formalize the projection of a digital attack into the physical domain by modeling the transformation function of the captured image through the optical pipeline. Our extensive experiments show that EvilEye's generated adversarial perturbations are much more robust across varying environmental light conditions relative to existing physical perturbation frameworks, achieving a high attack success rate (ASR) while bypassing state-of-the-art physical adversarial detection frameworks. We demonstrate that the dynamic nature of EvilEye enables attackers to adapt adversarial examples across a variety of objects with a significantly higher ASR compared to state-of-the-art physical world attack frameworks. Finally, we discuss mitigation strategies against the EvilEye attack.

cs.CR↗

Deep Graph-Level Clustering Using Pseudo-Label-Guided Mutual Information Maximization Network

In this work, we study the problem of partitioning a set of graphs into different groups such that the graphs in the same group are similar while the graphs in different groups are dissimilar. This problem was rarely studied previously, although there have been a lot of work on node clustering and graph classification. The problem is challenging because it is difficult to measure the similarity or distance between graphs. One feasible approach is using graph kernels to compute a similarity matrix for the graphs and then performing spectral clustering, but the effectiveness of existing graph kernels in measuring the similarity between graphs is very limited. To solve the problem, we propose a novel method called Deep Graph-Level Clustering (DGLC). DGLC utilizes a graph isomorphism network to learn graph-level representations by maximizing the mutual information between the representations of entire graphs and substructures, under the regularization of a clustering module that ensures discriminative representations via pseudo labels. DGLC achieves graph-level representation learning and graph-level clustering in an end-to-end manner. The experimental results on six benchmark datasets of graphs show that our DGLC has state-of-the-art performance in comparison to many baselines.

cs.LG↗

Solving McKean-Vlasov SDEs via relative entropy

In this paper we explore the merit of relative entropy in proving weak well-posedness of McKean-Vlasov SDEs and SPDEs, extending the technique introduced in Lacker arxiv:2105.02983. In the SDE setting, we prove weak existence and uniqueness when the interaction is path dependent and only assumed to have linear growth. Meanwhile, we recover and extend the current results when the interaction has Krylov's $L_t^q-L_x^p$ type singularity for $\frac{d}{p}+\frac{2}{q}<1$, where $d$ is the dimension of space. We connect the aforementioned two cases which are traditionally disparate, and form a solution theory that is sufficiently robust to allow perturbations of sublinear growth at the presence of singularity, giving rise to the well-posedness of a new family of McKean-Vlasov SDEs. Our strategy naturally extends to the cases of a fractional Brownian driving noise $B^H$ for all $H\in\left(0,1\right)$, obtaining new results in each separate case $H\in\left(0,\frac{1}{2}\right)$ and $H\in\left(\frac{1}{2},1\right)$. In the SPDE setting, we construct McKean-Vlasov type SPDEs with bounded measurable coefficients from the prototype of stochastic heat equation in spatial dimension one, and we do the same construction for the stochastic wave equation and a SPDE with white noise acting only on the boundary. In addition, we generalize some quantitative propagation of chaos results for SDEs into the SPDE setting.

math.PR↗

Laplacian-based Cluster-Contractive t-SNE for High Dimensional Data Visualization

Dimensionality reduction techniques aim at representing high-dimensional data in low-dimensional spaces to extract hidden and useful information or facilitate visual understanding and interpretation of the data. However, few of them take into consideration the potential cluster information contained implicitly in the high-dimensional data. In this paper, we propose LaptSNE, a new graph-layout nonlinear dimensionality reduction method based on t-SNE, one of the best techniques for visualizing high-dimensional data as 2D scatter plots. Specifically, LaptSNE leverages the eigenvalue information of the graph Laplacian to shrink the potential clusters in the low-dimensional embedding when learning to preserve the local and global structure from high-dimensional space to low-dimensional space. It is nontrivial to solve the proposed model because the eigenvalues of normalized symmetric Laplacian are functions of the decision variable. We provide a majorization-minimization algorithm with convergence guarantee to solve the optimization problem of LaptSNE and show how to calculate the gradient analytically, which may be of broad interest when considering optimization with Laplacian-composited objective. We evaluate our method by a formal comparison with state-of-the-art methods on seven benchmark datasets, both visually and via established quantitative measurements. The results demonstrate the superiority of our method over baselines such as t-SNE and UMAP. We also provide out-of-sample extension, large-scale extension and mini-batch extension for our LaptSNE to facilitate dimensionality reduction in various scenarios.

cs.LG↗

Spatio-temporal Keyframe Control of Traffic Simulation using Coarse-to-Fine Optimization

We present a novel traffic trajectory editing method which uses spatio-temporal keyframes to control vehicles during the simulation to generate desired traffic trajectories. By taking self-motivation, path following and collision avoidance into account, the proposed force-based traffic simulation framework updates vehicle's motions in both the Frenet coordinates and the Cartesian coordinates. With the way-points from users, lane-level navigation can be generated by reference path planning. With a given keyframe, the coarse-to-fine optimization is proposed to efficiently generate the plausible trajectory which can satisfy the spatio-temporal constraints. At first, a directed state-time graph constructed along the reference path is used to search for a coarse-grained trajectory by mapping the keyframe as the goal. Then, using the information extracted from the coarse trajectory as initialization, adjoint-based optimization is applied to generate a finer trajectory with smooth motions based on our force-based simulation. We validate our method with extensive experiments.

cs.GR↗

Smoothness of the density for McKean-Vlasov SDEs with measurable kernel

Consider the McKean-Vlasov SDE $$ dX_t=\langle b(X_t-\cdot),μ_t\rangle dt+dW_t,\quad μ_t=\operatorname{Law}(X_t), $$ where $W$ is the $n$-dimensional Brownian motion and $b:\mathbb{R}^d\to\mathbb{R}^d$ is a measurable function. First assuming $b\in L^\infty$, we prove that the law $μ_t$ of $X_t$ has a density $p_t$ with respect to the Lebesgue measure, which is continuously differentiable with gradient being $γ$-Hölder continuous for each $γ\in(0,1)$. Assume further that $b\in \mathcal{C}_b^1$, we prove that the density $p_t$ is infinitely differentiable. In the regularization by noise perspective, this shows McKean-Vlasov SDEs tend to have a smoother density function than SDEs without density dependence, under the same regularity assumption of the coefficients. We observe similar phenomenon for singular interaction kernels satisfying Krylov's integrability condition, for distributional kernels $b\in B_{\infty,\infty}^α$, $α\in(-1,0)$, and for processes driven by an $α$-stable noise for $α\in(1,2)$.

math.PR↗