SearcharxivSearch

arXiv subjects

Yuanyuan Li

Publications and source records attributed to Yuanyuan Li.

At least 19 recordsLinked to original sources

Global $W^{2,p}$ Regularity in Optimal Transport

In this paper we establish global $W^{2,p}$ estimates for the convex potentials of quadratic optimal transport between bounded convex domains with continuous positive densities. All the assumptions are optimal. The main new ideas include a blow-up analysis that allows for lower-dimensional collapse of the limiting source measure and reduces the limiting problem to a transport problem on its affine hull, and a good-bad scale decomposition in which rigidity controls the good scales while a counting argument shows that the proportion of bad scales tends to zero.

math.AP

The many-body Blaschke-Santal\'o type inequality via optimal transport

Let $K_1,\ldots,K_k\subset\mathbb R^n$ be origin-symmetric measurable sets of finite volume such that \[ \sum_{1\le i<j\le k}\langle x_i,x_j\rangle\le \binom{k}{2}, \qquad \forall\,x_i\in K_i, x_j\in K_j. \] We prove the sharp many-body Blaschke--Santal\'o type inequality \[ \prod_{i=1}^k |K_i|\le |B^n|^k \] proposed by Kalantzopoulos and Saroglou, and characterize all equality cases. The proof combines multi-marginal optimal transport with a pseudo-Euclidean volume estimate. Using the geometric--functional equivalence of Kalantzopoulos and Saroglou, we also establish the functional version inequality proposed by Kolesnikov and Werner.

math.AP

HumanoidUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation

High-quality demonstration data are essential for humanoid robot skill learning, especially for whole-body behaviors that require coordinated perception, locomotion, and manipulation. Existing data-collection methods largely rely on robot teleoperation, which is constrained by hardware accessibility, operator expertise, and limited efficiency. Inspired by the Universal Manipulation Interface (UMI), we propose HumanoidUMI, a portable and robot-free framework for humanoid whole-body data collection. HumanoidUMI uses lightweight VR devices and UMI-inspired grippers to collect sparse human keypoint trajectories, wrist-view observations, and gripper actions. These demonstrations train a high-level policy to predict future keypoints, which are retargeted to robot-native whole-body references and executed by a whole-body controller. Experiments in five real-world scenarios demonstrate the effectiveness of the proposed framework and validate the collected demonstrations for transferable humanoid whole-body skill learning.

cs.RO

The Symmetric Mahler Inequality in Dimension Three via Admissible Shadow Systems

The three-dimensional symmetric Mahler inequality states that, for every origin-symmetric convex body \(K=-K\subset \mathbb{R}^3\), \[ \VP(K)= |K|\,|K^\circ|\geq \frac{32}{3}. \] It was recently proved by Iriyeh--Shibata \cite{IS2020}, and a shorter proof was later given by Fradelizi--Hubard--Meyer--Rold\'an-Pensado--Zvavitch \cite{FHMRZ}. Both proofs combine ingenious equipartition arguments of algebraic-topological origin with delicate geometric estimates inspired by Meyer's argument for unconditional bodies. In this paper, we give a new proof of this inequality using a purely geometric approach, based on what we call symmetric admissible shadow systems. This is a natural extension of the new techniques developed in our proof of the three-dimensional non-symmetric Mahler conjecture \cite{CLXX-Mahler}.

math.MG

The Mahler Conjecture in Three Dimensions

The Mahler conjecture dates back to 1938. This paper solves the conjecture for general convex bodies in three dimensions by developing a method called the shadow flow. The equality case is characterized as well. This method is also applied to give a new proof of the three-dimensional symmetric case, which was first proved by Iriyeh--Shibata.

math.MG

BifrostUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation

High-quality demonstration data are essential for humanoid robot skill learning, especially for whole-body behaviors that require coordinated perception, locomotion, and manipulation. Existing data-collection methods largely rely on robot teleoperation, which is constrained by hardware accessibility, operator expertise, and limited efficiency. Inspired by the Universal Manipulation Interface (UMI), we propose BifrostUMI, a portable and robot-free framework for humanoid whole-body data collection. BifrostUMI uses lightweight VR devices and UMI-inspired grippers to collect sparse human keypoint trajectories, wrist-view observations, and gripper actions. These demonstrations train a high-level policy to predict future keypoints, which are retargeted to robot-native whole-body references and executed by a whole-body controller. Experiments in five real-world scenarios demonstrate the effectiveness of the proposed framework and validate the collected demonstrations for transferable humanoid whole-body skill learning.

cs.RO

On the monotonicity of affine quermassintegrals

Lutwak's affine quermassintegral theory is a foundational component of modern affine Brunn--Minkowski theory. Developed in the 1980s, it provides affine analogues of the classical quermassintegrals and has led to a rich family of sharp affine isoperimetric inequalities. A central question in this program, going back to Lutwak's 1988 work, is an Alexandrov--Fenchel-type monotonicity principle for the normalized $L^{-n}$-moment quermassintegrals $I_{k,-n}$. In one form, this principle predicts that \[ I_{m,-n}(K)^{1/m}\ge I_{k,-n}(K)^{1/k}, \qquad 1\le m (m+2)(k+2)-2$, there exists an origin-symmetric $C^2_+$ convex body $K\subset\mathbb R^n$ such that \[ I_{m,-n}(K)^{1/m} < I_{k,-n}(K)^{1/k}. \] The example is obtained from the Euclidean ball by an arbitrarily small degree-four spherical harmonic perturbation. On the positive side, we prove that the endpoint chain is true in dimension three: for every convex body $K\subset\mathbb R^3$, \[ I_{1,-3}(K)\ge I_{2,-3}(K)^{1/2}\ge I_{3,-3}(K)^{1/3}=1. \] The equality cases in both non-trivial inequalities are exactly ellipsoids, up to translation and nonsingular affine transformations.

math.AP

Uniqueness of Blow-ups for the Superconductivity Free Boundary Problem

We study the free-boundary equation \[ \Delta u=\chi_{\{|\nabla u|>0\}} \] near the origin. We prove that, at a singular point of \(\partial\{|\nabla u|>0\}\), the quadratic blow-up is unique. As noted in \cite[Notes to Chapter 7]{PSU2012}, little is known about the singular set for this problem. The usual Weiss--Monneau monotonicity argument does not seem to apply directly, because the inactive set is determined by the vanishing of \(\nabla u\), rather than by a sign condition on \(u\). The proof follows the quadratic part of the rescalings. Projecting onto the trace-free quadratic harmonics yields a finite-dimensional differential equation for the quadratic coefficient. Together with a Lyapunov identity and estimates on dyadic annuli, this implies convergence of the quadratic coefficient, and hence uniqueness of the blow-up.

math.AP

OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction

UMI-style interfaces enable scalable robot learning, but existing systems remain largely visuomotor, relying primarily on RGB observations and trajectory while providing only limited access to physical interaction signals. This becomes a fundamental limitation in contact-rich manipulation, where success depends on contact dynamics such as tactile interaction, internal grasping force, and external interaction wrench that are difficult to infer from vision alone. We present OmniUMI, a unified framework for physically grounded robot learning via human-aligned multimodal interaction. OmniUMI synchronously captures RGB, depth, trajectory, tactile sensing, internal grasping force, and external interaction wrench within a compact handheld system, while maintaining collection--deployment consistency through a shared embodiment design. To support human-aligned demonstration, OmniUMI enables natural perception and modulation of internal grasping force, external interaction wrench, and tactile interaction through bilateral gripper feedback and the handheld embodiment. Built on this interface, we extend diffusion policy with visual, tactile, and force-related observations, and deploy the learned policy through impedance-based execution for unified regulation of motion and contact behavior. Experiments demonstrate reliable sensing and strong downstream performance on force-sensitive pick-and-place, interactive surface erasing, and tactile-informed selective release. Overall, OmniUMI combines physically grounded multimodal data acquisition with human-aligned interaction, providing a scalable foundation for learning contact-rich manipulation.

cs.RO

MOGeo: Beyond One-to-One Cross-View Object Geo-localization

Cross-View Object Geo-Localization (CVOGL) aims to locate an object of interest in a query image within a corresponding satellite image. Existing methods typically assume that the query image contains only a single object, which does not align with the complex, multi-object geo-localization requirements in real-world applications, making them unsuitable for practical scenarios. To bridge the gap between the realistic setting and existing task, we propose a new task, called Cross-View Multi-Object Geo-Localization (CVMOGL). To advance the CVMOGL task, we first construct a benchmark, CMLocation, which includes two datasets: CMLocation-V1 and CMLocation-V2. Furthermore, we propose a novel cross-view multi-object geo-localization method, MOGeo, and benchmark it against existing state-of-the-art methods. Extensive experiments are conducted under various application scenarios to validate the effectiveness of our method. The results demonstrate that cross-view object geo-localization in the more realistic setting remains a challenging problem, encouraging further research in this area.

cs.CV

Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting

Large scale text-to-image generation models can memorize and reproduce their training dataset. Since the training dataset often contains copyrighted material, reproduction of training dataset poses a copyright infringement risk, which could result in legal liabilities and financial losses for both the AI user and the developer. The current works explores the potential of chain-of-thought and task instruction prompting in reducing copyrighted content generation. To this end, we present a formulation that combines these two techniques with two other copyright mitigation strategies: a) negative prompting, and b) prompt re-writing. We study the generated images in terms their similarity to a copyrighted image and their relevance of the user input. We present numerical experiments on a variety of models and provide insights on the effectiveness of the aforementioned techniques for varying model complexity.

cs.LG

Post-2024 U.S. Presidential Election Analysis of Election and Poll Data: Real-life Validation of Prediction via Small Area Estimation and Uncertainty Quantification

We carry out a post-election analysis of the 2024 U.S. Presidential Election (USPE) using a prediction model derived from the Small Area Estimation (SAE) methodology. With pollster data obtained one week prior to the election day, retrospectively, our SAE-based prediction model can perfectly predict the Electoral College election results in all 44 states where polling data were available. In addition to such desirable prediction accuracy, we introduce the probability of incorrect prediction (PoIP) to rigorously analyze prediction uncertainty. Since the standard bootstrap method appears inadequate for estimating PoIP, we propose a conformal inference method that yields reliable uncertainty quantification. We further investigate potential pollster biases by the means of sensitivity analyses and conclude that swing states are particularly vulnerable to polling bias in the prediction of the 2024 USPE.

stat.AP

Seeing the Unseen: Mask-Driven Positional Encoding and Strip-Convolution Context Modeling for Cross-View Object Geo-Localization

Cross-view object geo-localization enables high-precision object localization through cross-view matching, with critical applications in autonomous driving, urban management, and disaster response. However, existing methods rely on keypoint-based positional encoding, which captures only 2D coordinates while neglecting object shape information, resulting in sensitivity to annotation shifts and limited cross-view matching capability. To address these limitations, we propose a mask-based positional encoding scheme that leverages segmentation masks to capture both spatial coordinates and object silhouettes, thereby upgrading the model from "location-aware" to "object-aware." Furthermore, to tackle the challenge of large-span objects (e.g., elongated buildings) in satellite imagery, we design a context enhancement module. This module employs horizontal and vertical strip convolutional kernels to extract long-range contextual features, enhancing feature discrimination among strip-like objects. Integrating MPE and CEM, we present EDGeo, an end-to-end framework for robust cross-view object geo-localization. Extensive experiments on two public datasets (CVOGL and VIGOR-Building) demonstrate that our method achieves state-of-the-art performance, with a 3.39% improvement in localization accuracy under challenging ground-to-satellite scenarios. This work provides a robust positional encoding paradigm and a contextual modeling framework for advancing cross-view geo-localization research.

cs.CV

Counterfactually Fair Conformal Prediction

While counterfactual fairness of point predictors is well studied, its extension to prediction sets--central to fair decision-making under uncertainty--remains underexplored. On the other hand, conformal prediction (CP) provides efficient, distribution-free, finite-sample valid prediction sets, yet does not ensure counterfactual fairness. We close this gap by developing Counterfactually Fair Conformal Prediction (CF-CP) that produces counterfactually fair prediction sets. Through symmetrization of conformity scores across protected-attribute interventions, we prove that CF-CP results in counterfactually fair prediction sets while maintaining the marginal coverage property. Furthermore, we empirically demonstrate that on both synthetic and real datasets, across regression and classification tasks, CF-CP achieves the desired counterfactual fairness and meets the target coverage rate with minimal increase in prediction set size. CF-CP offers a simple, training-free route to counterfactually fair uncertainty quantification.

cs.LG

Domain-Shift-Aware Conformal Prediction for Large Language Models

Large language models have achieved impressive performance across diverse tasks. However, their tendency to produce overconfident and factually incorrect outputs, known as hallucinations, poses risks in real-world applications. Conformal prediction provides finite-sample, distribution-free coverage guarantees, but standard conformal prediction breaks down under domain shift, often leading to under-coverage and unreliable prediction sets. We propose a new framework called Domain-Shift-Aware Conformal Prediction (DS-CP). Our framework adapts conformal prediction to large language models under domain shift, by systematically reweighting calibration samples based on their proximity to the test prompt, thereby preserving validity while enhancing adaptivity. Our theoretical analysis and experiments on the MMLU benchmark demonstrate that the proposed method delivers more reliable coverage than standard conformal prediction, especially under substantial distribution shifts, while maintaining efficiency. This provides a practical step toward trustworthy uncertainty quantification for large language models in real-world deployment.

stat.ML

The $L_p$ chord Minkowski problem for super-critical exponent

The $L_p$ chord Minkowski problem was recently introduced by Lutwak, Xi, Yang and Zhang, which seeks to determine the necessary and sufficient conditions for a given finite Borel measure such that it is the $L_p$ chord measure of a convex body. In this paper, we solve the $L_p$ chord Minkowski problem for the super-critical exponents by combining a nonlocal Gauss curvature flow introduced in \cite{HHLW exi} and a topological argument developed in \cite{GLW2022}. Notably, we provide a simplified argument for the topological part.

math.AP

Q-REACH: Quantum information Repetition, Error Analysis and Correction using Caching Network

Quantum repeaters incorporating quantum memory play a pivotal role in mitigating loss in transmitted quantum information (photons) due to link attenuation over a long-distance quantum communication network. However, limited availability of available storage in such quantum repeaters and the impact on the time spent within the memory unit presents a trade-off between quantum information fidelity (a metric that quantifies the degree of similarity between a pair of quantum states) and qubit transmission rate. Thus, effective management of storage time for qubits becomes a key consideration in multi-hop quantum networks. To address these challenges, we propose Q-REACH, which leverages queuing theory in caching networks to tune qubit transmission rate while considering fidelity as the cost metric. Our contributions in this work include (i) utilizing a method of repetition that encodes and broadcasts multiple qubits through different quantum paths, (ii) analytically estimating the time spent by these emitted qubits as a function of the number of paths and repeaters, as well as memory units within a repeater, and (iii) formulating optimization problem that leverages this analysis to correct the transmitted logic qubit and select the optimum repetition rate at the transmitter.

quant-ph