SearcharxivSearch

arXiv subjects

Kevin Ren

Publications and source records attributed to Kevin Ren.

At least 19 recordsLinked to original sources

A threshold phenomenon for embeddings of Euclidean snowflakes and impossibility of dimension reduction

Fix $0<\theta\leqslant 1$. We prove that if $1\leqslant p \leqslant 2/\theta$, then the $\theta$-snowflake of $\ell_2^k$, namely, $\mathbb{R}^k$ equipped with the metric $((x,y)\in \mathbb{R}^k\times \mathbb{R}^k)\mapsto \|x-y\|_2^\theta$, embeds with distortion $O(1)$ into $\ell_p^m$ for some integer $m\lesssim_{p,\theta}k$, which is optimal as $k\to \infty$, as seen by comparing dimensions. However, for $p$ larger than the sharp threshold $2/\theta$ the following change in behavior occurs: If a $(1/\sqrt{k})$-dense subset of the Euclidean sphere $S^{k-1}$ embeds into $\ell_p^m$ with distortion $O(1)$, then necessarily $m\gtrsim_{p,\theta}( k/\log k)^{p\theta/2}$, which grows super-linearly in $k$ as $p\theta/2>1$, and this dimension bound is optimal as $k\to \infty$ up to lower order factors. We deduce from this statement that if $2<p<\infty$, then there exist arbitrarily large $n$-point subsets of $\ell_p$ with the property that if they embed with distortion $O(1)$ into $\ell_p^m$, then necessarily $m\gtrsim_p ((\log n)/(\log\log n)^2)^{p/2}$, thus demonstrating that the statement of the Johnson--Lindenstrauss dimension reduction lemma fails to hold for $\ell_p$

math.MG

Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift

Deployed approaches for AI text detection often rely on training-time access to labeled datasets of both human-written and AI-generated text. This approach is vulnerable to three types of distribution shifts that occur continually post-deployment, and for which labeled data is often unavailable: adversarial humanization, new LLMs being released, and temporal drift in human writing. Simultaneously, existing approaches do not leverage a key signal of LLM usage: inference-time homogeneity. We propose a test-time adaptation (TTA) approach, using semi-supervised learning, that adapts to distribution shifts by leveraging homogeneity among unlabeled samples observed at inference time. Empirically, we find that state-of-the-art supervised detectors systematically fail when they encounter distribution shifts in AI-generated and human writing, both adversarial and natural, while test-time adaptation with semi-supervised learning is largely robust; e.g., the commercial model Pangram detects just 24.1% of our adversarial AI-generated text, compared to 90.5% for our test-time approach. We establish that test-time adaptation is a promising framework for AI text detection in the wild. We publicly release our code (which includes code for model training, evaluation, and plots) at https://github.com/kkr36/llm_detection.

cs.CL

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic Alignment Gap when applied to upper-undergraduate to early graduate level mathematics. To quantify this, we introduce QEDBench, the first large-scale dual-rubric alignment benchmark to systematically measure alignment with human experts on university-level math proofs by contrasting course-specific rubrics against expert common knowledge criteria. By deploying a dual-evaluation matrix (7 judges x 5 solvers) against 1,000+ hours of human evaluation, we reveal that certain frontier evaluators like Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, and Llama 4 Maverick exhibit significant positive bias (up to +0.18, +0.20, +0.30, +0.36 mean score inflation, respectively). Furthermore, we uncover a critical reasoning gap in the discrete domain: while Gemini 3.0 Pro achieves state-of-the-art performance (0.91 average human evaluation score), other reasoning models like GPT-5 Pro and Claude Sonnet 4.5 see their performance significantly degrade in discrete domains. Specifically, their average human evaluation scores drop to 0.72 and 0.63 in Discrete Math, and to 0.74 and 0.50 in Graph Theory. In addition to these research results, we also release QEDBench as a public benchmark for evaluating and improving AI judges. Our benchmark is publicly published at https://github.com/qqliu/Yale-QEDBench.

cs.LG

Reconstruction of Manifold Distances from Noisy Observations

We consider the problem of reconstructing the intrinsic geometry of a manifold from noisy pairwise distance observations. Specifically, let $M$ denote a diameter 1 d-dimensional manifold and $\mu$ a probability measure on $M$ that is mutually absolutely continuous with the volume measure. Suppose $X_1,\dots,X_N$ are i.i.d. samples of $\mu$ and we observe noisy-distance random variables $d'(X_j, X_k)$ that are related to the true geodesic distances $d(X_j,X_k)$. With mild assumptions on the distributions and independence of the noisy distances, we develop a new framework for recovering all distances between points in a sufficiently dense subsample of $M$. Our framework improves on previous work which assumed i.i.d. additive noise with known moments. Our method is based on a new way to estimate $L_2$-norms of certain expectation-functions $f_x(y)=\mathbb{E}d'(x,y)$ and use them to build robust clusters centered at points of our sample. Using a new geometric argument, we establish that, under mild geometric assumptions--bounded curvature and positive injectivity radius--these clusters allow one to recover the true distances between points in the sample up to an additive error of $O(\varepsilon \log \varepsilon^{-1})$. We develop two distinct algorithms for producing these clusters. The first achieves a sample complexity $N \asymp \varepsilon^{-2d-2}\log(1/\varepsilon)$ and runtime $o(N^3)$. The second introduces novel geometric ideas that warrant further investigation. In the presence of missing observations, we show that a quantitative lower bound on sampling probabilities suffices to modify the cluster construction in the first algorithm and extend all recovery guarantees. Our main technical result also elucidates which properties of a manifold are necessary for the distance recovery, which suggests further extension of our techniques to a broader class of metric probability spaces.

stat.ML

High-low method and $p$-adic Furstenberg set over the plane

We establish a $p$-adic analogue of a recent significant result of Ren-Wang (arXiv:2308.08819) on Furstenberg sets in the Euclidean plane. Building on the $p$-adic version of the high-low method from Chu (arXiv:2510.20104), we analyze cube-tube incidences in $\mathbb{Q}_p^2$ and prove that for $s < t < 2 - s$, any semi-well-spaced $(s,t)$-Furstenberg set over $\mathbb{Q}_p^2$ has Hausdorff dimension $\ge\frac{3s+t}{2}$. Moreover, as a byproduct of our argument, we obtain the sharp lower bounds $s+t$ (for $0<t\le s\le 1$) and $s+1$ (for $s+t\ge 2$) for general $(s,t)$-Furstenberg sets without the semi-well-spaced assumption, thereby confirming that all three lower bounds match those in the Euclidean case.

math.FA

Predicting Language Models' Success at Zero-Shot Probabilistic Prediction

Recent work has investigated the capabilities of large language models (LLMs) as zero-shot models for generating individual-level characteristics (e.g., to serve as risk models or augment survey datasets). However, when should a user have confidence that an LLM will provide high-quality predictions for their particular task? To address this question, we conduct a large-scale empirical study of LLMs' zero-shot predictive capabilities across a wide range of tabular prediction tasks. We find that LLMs' performance is highly variable, both on tasks within the same dataset and across different datasets. However, when the LLM performs well on the base prediction task, its predicted probabilities become a stronger signal for individual-level accuracy. Then, we construct metrics to predict LLMs' performance at the task level, aiming to distinguish between tasks where LLMs may perform well and where they are likely unsuitable. We find that some of these metrics, each of which are assessed without labeled data, yield strong signals of LLMs' predictive performance on new tasks.

cs.LG

Triadic structures in multislice networks

Networks provide a popular representation of complex data. Often, different types of relational measurements are taken on the same subjects. Such data can be represented as a \textit{multislice network}, a collection of networks on the same set of nodes, with connections between the different layers to be determined. For the analysis of multislice networks, we take inspiration from the analysis of simple networks, for which small subgraphs (motifs) have proven to be useful; motifs are even seen as building blocks of complex networks. A particular instance of a motif is a triangle, and while triangle counts are well understood for simple network models such as Erd\H{o}s-R\'enyi random graphs, with i.i.d. distributed edges, even for simple multislice network models little is known about triangle counts. Here we address this issue by extending the analysis of triadic structures to multislice Erd\H{o}s-R\'enyi networks. Again taking inspiration from the analysis of sparse Erd\H{o}s-R\'enyi random graphs, we show that the distribution of triangles across multiple layers in a multislice Erd\H{o}s-R\'enyi network can be well approximated by an appropriate Poisson distribution. This theoretical result opens the door to statistical goodness of fit tests for multislice networks.

math.PR

Incidence bounds related to circular Furstenberg sets

We prove bounds on approximate incidences between families of circles and families of points in the plane. As a consequence, we prove a lower bound for the dimension of circular $(u,v)$-Furstenberg sets, which is new for large $u$ and $v$.

math.CA

Euclidean embedding, randomized clustering, and Lipschitz extension for finite and doubling subsets of $L_p$ when $p>2$

Fix $p>2$. We prove that the Euclidean distortion of every $n$-point subset of $L_p$ is $p^3(\log n)^{\frac12+o(1)}$, thus, in particular, demonstrating that all $n$-point subsets of $L_p$ exhibit an asymptotic improvement over the $O(\log n)$ Euclidean distortion guarantee that Bourgain's embedding theorem provides for arbitrary $n$-point metric spaces. We also prove that the separation modulus of every $n$-point subset of $ L_p$ is $O(p^2\sqrt{\log n})$, which is sharp up to the dependence on $p$. We deduce from (a refinement of) this asymptotic evaluation of the finitary separation modulus of $ L_p$ that for any $n$-point subset $\mathcal{C}$ of $ L_p$, any Banach space $\mathbf{Z}$, and any $1$-Lipschitz function $f:\mathcal{C}\to \mathbf{Z}$, there exists a $O(p^2\sqrt{\log n})$-Lipschitz function $F:L_p\to \mathbf{Z}$ that extends $f$. We obtain analogous separation and extension statements for doubling subsets of $L_p$.

math.FA

On the projections of almost Ahlfors regular sets

We show that the "sharp Kaufman projection theorem" from 2023 is sharp in the class of Ahlfors $(1,\delta^{-\epsilon})$-regular sets. This is in contrast with a recent result of the first author, which improves the projection theorem in the class of Ahlfors $(1,C)$-regular sets.

math.CA

Random zero sets with local growth guarantees

We prove that if $(\mathcal{M},d)$ is an $n$-point metric space that embeds quasisymmetrically into a Hilbert space, then for every $\tau>0$ there is a random subset $\mathcal{Z}$ of $\mathcal{M}$ such that for any pair of points $x,y\in \mathcal{M}$ with $d(x,y)\ge \tau$, the probability that both $x\in \mathcal{Z}$ and $d(y,\mathcal{Z})\ge \beta\tau/\sqrt{1+\log (|B(y,\kappa \beta \tau)|/|B(y,\beta \tau)|)}$ is $\Omega(1)$, where $\kappa>1$ is a universal constant and $\beta>0$ depends only on the modulus of the quasisymmetric embedding. The proof relies on a refinement of the Arora--Rao--Vazirani rounding technique. Among the applications of this result is that the largest possible Euclidean distortion of an $n$-point subset of $\ell_1$ is $\Theta(\sqrt{\log n})$, and the integrality gap of the Goemans--Linial semidefinite program for the Sparsest Cut problem on inputs of size $n$ is $\Theta(\sqrt{\log n})$. Multiple further applications are given.

math.MG

Work Smarter Not Harder: Simple Imitation Learning with CS-PIBT Outperforms Large Scale Imitation Learning for MAPF

Multi-Agent Path Finding (MAPF) is the problem of effectively finding efficient collision-free paths for a group of agents in a shared workspace. The MAPF community has largely focused on developing high-performance heuristic search methods. Recently, several works have applied various machine learning (ML) techniques to solve MAPF, usually involving sophisticated architectures, reinforcement learning techniques, and set-ups, but none using large amounts of high-quality supervised data. Our initial objective in this work was to show how simple large scale imitation learning of high-quality heuristic search methods can lead to state-of-the-art ML MAPF performance. However, we find that, at least with our model architecture, simple large scale (700k examples with hundreds of agents per example) imitation learning does \textit{not} produce impressive results. Instead, we find that by using prior work that post-processes MAPF model predictions to resolve 1-step collisions (CS-PIBT), we can train a simple ML MAPF model in minutes that dramatically outperforms existing ML MAPF policies. This has serious implications for all future ML MAPF policies (with local communication) which currently struggle to scale. In particular, this finding implies that future learnt policies should (1) always use smart 1-step collision shields (e.g. CS-PIBT), (2) always include the collision shield with greedy actions as a baseline (e.g. PIBT) and (3) motivates future models to focus on longer horizon / more complex planning as 1-step collisions can be efficiently resolved.

cs.MA

Decision-Focused Evaluation of Worst-Case Distribution Shift

Distribution shift is a key challenge for predictive models in practice, creating the need to identify potentially harmful shifts in advance of deployment. Existing work typically defines these worst-case shifts as ones that most degrade the individual-level accuracy of the model. However, when models are used to make a downstream population-level decision like the allocation of a scarce resource, individual-level accuracy may be a poor proxy for performance on the task at hand. We introduce a novel framework that employs a hierarchical model structure to identify worst-case distribution shifts in predictive resource allocation settings by capturing shifts both within and across instances of the decision problem. This task is more difficult than in standard distribution shift settings due to combinatorial interactions, where decisions depend on the joint presence of individuals in the allocation task. We show that the problem can be reformulated as a submodular optimization problem, enabling efficient approximations of worst-case loss. Applying our framework to real data, we find empirical evidence that worst-case shifts identified by one metric often significantly diverge from worst-case distributions identified by other metrics.

cs.LG

Radial Projections in $\mathbb{R}^n$ Revisited

We generalize the recent results on radial projections by Orponen, Shmerkin, Wang using two different methods. In particular, we show that given $X,Y\subset \mathbb{R}^n$ Borel sets and $X\neq \emptyset$. If $\dim Y \in (k,k+1]$ for some $k\in \{1,\dots, n-1\}$, then \[ \sup_{x\in X} \dim \pi_x(Y\setminus \{x\}) \geq \min \{\dim X + \dim Y - k, k\}. \] Our results give a new approach to solving a conjecture of Lund-Pham-Thu in all dimensions and for all ranges of $\dim Y$. The first of our two methods for proving the above theorem is shorter, utilizing a result of the first author and Gan. Our second method, though longer, follows the original methodology of Orponen--Shmerkin--Wang, and requires a higher dimensional incidence estimate and a dual Furstenberg-set estimate for lines. These new estimates may be of independent interest.

math.CA

Improving Learnt Local MAPF Policies with Heuristic Search

Multi-agent path finding (MAPF) is the problem of finding collision-free paths for a team of agents to reach their goal locations. State-of-the-art classical MAPF solvers typically employ heuristic search to find solutions for hundreds of agents but are typically centralized and can struggle to scale when run with short timeouts. Machine learning (ML) approaches that learn policies for each agent are appealing as these could enable decentralized systems and scale well while maintaining good solution quality. Current ML approaches to MAPF have proposed methods that have started to scratch the surface of this potential. However, state-of-the-art ML approaches produce "local" policies that only plan for a single timestep and have poor success rates and scalability. Our main idea is that we can improve a ML local policy by using heuristic search methods on the output probability distribution to resolve deadlocks and enable full horizon planning. We show several model-agnostic ways to use heuristic search with learnt policies that significantly improve the policies' success rates and scalability. To our best knowledge, we demonstrate the first time ML-based MAPF approaches have scaled to high congestion scenarios (e.g. 20% agent density).

cs.MA

Sobolev extension in a simple case

In this paper, we establish the existence of a bounded, linear extension operator $T: L^{2,p}(E) \to L^{2,p}(\mathbb{R}^2)$ when $1<p<2$ and $E$ is a finite subset of $\mathbb{R}^2$ contained in a line.

math.CA

Discretized Radial Projections in $\mathbb{R}^d$

We generalize a Furstenberg-type result of Orponen-Shmerkin to higher dimensions, leading to an $\epsilon$-improvement in Kaufman's projection theorem for hyperplanes and an unconditional discretized radial projection theorem in the spirit of Orponen-Shmerkin-Wang. Our proof relies on a new incidence estimate for $\delta$-tubes and a quasi-product set of $\delta$-balls in $\mathbb{R}^d$.

math.CA

New improvement to Falconer distance set problem in higher dimensions

We show that if a compact set $E\subset \mathbb{R}^d$ has Hausdorff dimension larger than $\frac{d}{2}+\frac{1}{4}-\frac{1}{8d+4}$, where $d\geq 3$, then there is a point $x\in E$ such that the pinned distance set $\Delta_x(E)$ has positive Lebesgue measure. This improves upon bounds of Du-Zhang and Du-Iosevich-Ou-Wang-Zhang in all dimensions $d \ge 3$. We also prove lower bounds for Hausdorff dimension of pinned distance sets when $\dim_H (E) \in (\frac{d}{2} - \frac{1}{4} - \frac{3}{8d+4}, \frac{d}{2}+\frac{1}{4}-\frac{1}{8d+4})$, which improves upon bounds of Harris and Wang-Zheng in dimensions $d \ge 3$.

math.CA