SearcharxivSearch

arXiv subjects

Haonan Zhang

Publications and source records attributed to Haonan Zhang.

At least 19 recordsLinked to original sources

The K\"onig constant is one

For each $N\geq1$, consider the normalized K\"onig bilinear form $B_{\mathrm K}:L_\infty(\mathbb R^N)\times L_\infty(\mathbb R^N)\to\mathbb R$ given by \[ B_{\mathrm K}(f,g):=\frac{1}{(\sqrt{2}\pi)^N} \iint_{\mathbb R^N\times\mathbb R^N} f(x)g(y)e^{-(\lVert x\rVert^2+\lVert y\rVert^2)/2} \sin\langle x,y\rangle\,\mathrm d x\,\mathrm d y, \] We define the K\"onig constant by \[ \mathfrak K_{\mathrm K}:=\sup_{N\geq1}\sup_{\substack{f,g:\mathbb R^N\to\{\pm1\}\\ f,g\ \mathrm{measurable}}}B_{\mathrm K}(f,g). \] The study of this bilinear form arose from efforts to determine the exact value of the Grothendieck constant. K\"onig~\cite{KONIG} conjectured that the sharp value should instead be given by the one-dimensional half-spaces $B_{\mathrm K}(\operatorname{sgn}(x_1),\operatorname{sgn}(x_1))=\frac{2}{\pi}\log(1+\sqrt{2})$. A positive answer to this conjecture, together with a classical upper bound of Krivine \cite{KRIVINE}, would determine the exact value of the Grothendieck constant. In a breakthrough~\cite{BMMN}, Braverman, Makarychev, Makarychev, and Naor disproved K\"onig's conjecture already in dimension two and used their counterexamples to obtain the first strict improvement over Krivine's bound. One question in \cite{BMMN} attempts to determine the Grothendieck constant through alternating Krivine rounding schemes arising from K\"onig's bilinear form in high dimension. More recently, Li et al.~\cite{LISK} constructed high-dimensional examples showing that $\mathfrak K_{\mathrm K}\ge 0.59357$. An elementary Fourier argument gives $\mathfrak K_{\mathrm K}\le 1$ and excludes equality for every finite-dimension. In this paper, we prove that $\mathfrak K_{\mathrm K}=1$ by constructing a family of Boolean pairs in high dimensions. In particular, this gives a negative answer to the high-dimensional aspect of the question in \cite{BMMN}.

math.FA

Tightness of and counterexamples to several quantum estimates

We prove here several tightness results for such quantum inequalities as the comparison of operator norm and product norm of $d$-local hamiltonians, Bohnenblust--Hille inequality for $d$-local hamiltonians and for quantum Fourier entropy-influence conjecture, we also discuss the quantum Aaronson--Ambainis conjecture in a special case of anti-commuting Pauli strings.

math.AP

RL-Lock: Reinforcement Learning for Generating Interlocking Assemblies

An interlocking assembly is an assembly in which component parts are connected purely through their geometric arrangement, without relying on external connectors such as glue and nails. Such assemblies have been widely used in a variety of real-world applications due to their structural stability. The problem of generating interlocking assemblies is generally formulated as a shape decomposition problem, where a target 3D object represented as a voxel grid is partitioned into a prescribed number of interlocking pieces. We observe that generating interlocking assemblies is inherently a sequential decision-making problem, where an agent repeatedly decides which piece each voxel should be assigned to. Inspired by the observation, we propose the first reinforcement learning framework RL-Lock for generating interlocking assemblies, without relying on handcrafted search heuristics as existing works did. RL-Lock combines structured action chunking with MCTS-guided policy-value learning to efficiently navigate the large combinatorial search space for interlocking assembly generation. We demonstrate through experiments that RL-Lock allows effective generation of interlocking assemblies, especially for challenging cases in which existing approaches take too long or even fail to find a valid solution.

cs.AI

SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements

Although coding agents have achieved impressive performance on correctness-oriented benchmarks, their ability to make behavior-preserving non-functional improvements (NFIs) remains underexplored. In real-world software development, developers continuously improve software quality without changing observable behavior, yet existing benchmarks primarily evaluate functional correctness and provide limited support for assessing these non-functional improvements. In this paper, we present SWE-NFI, a benchmark for evaluating coding agents on NFIs beyond functional correctness. Our benchmark contains 188 tasks constructed from real merged pull requests in open-source Python projects. We operationalize developer-oriented NFIs into 92 executable rules and develop a comprehensive evaluation suite that combines functional correctness testing with rule-based NFI evaluation. We evaluate state-of-the-art commercial and open-source coding agents. Although the best-performing agent achieves a 70.0\% functional correctness rate, all evaluated agents generally fall short of human developers in overall NFI capability. The gap is particularly evident for structural code improvements, where agents' NFI scores range from 0.0 to 1.3, compared with 1.5 for the human reference. Our benchmark and findings provide a reproducible foundation for evaluating and advancing coding agents beyond functional correctness.

cs.SE

Single-laser stimulated Brillouin scattering microscopy

Stimulated Brillouin scattering (SBS) microscopy enables label-free mapping of local viscoelastic properties, but frequency-domain implementations are often limited by uncertainty in the pump-probe frequency-difference axis. We demonstrate an RF-defined single-laser electro-optic-modulation SBS microscope in which the pump and probe are derived from the same optical carrier and their frequency difference is set by an electro-optically generated sideband. This architecture makes laser-frequency noise largely common mode and eliminates optical wavelength tuning during spectral scanning. It achieves Brillouin frequency shift and linewidth precisions of 0.07 MHz and 0.30 MHz, respectively. Comparison with a low-NA reference linewidth indicates a system-level spectral broadening of approximately 3.1 MHz, corresponding to an effective spectral resolution of approximately 3 MHz. Imaging of femtosecond-laser-modified chalcogenide glass resolves MHz-level Brillouin contrasts corresponding to 10^-4-level apparent longitudinal-modulus contrast. This work demonstrates the feasibility of transferring the frequency definition of SBS spectral scanning from optical wavelength tuning to RF-domain control, providing a new conceptual and technical basis for high-precision, high-spectral-fidelity Brillouin imaging.

physics.optics

A Beckmann boundary form of Talagrand's conjecture on the discrete cube

We introduce the Beckmann boundary of a Boolean function \[ \mathsf{B}(f)=\inf_{\operatorname{div} V=Lf}\mathbb E\|V(x)\|_2. \] Here \[ L=\sum_iD_i,\qquad D_i f(x)=\frac{f(x)-f(x^{\oplus i})}{2}, \] and $\operatorname{div} V(x)=\sum_i (V_{i}(x)-V_{i}(x^{\oplus i}))$. This nonlocal quantity is no larger than the usual two-sided, one-sided, colored, optimized colored, or optimized fractional colored boundaries. Nevertheless, every nonconstant Boolean $f$ satisfies \[ \mathsf{B}(f)\gtrsim \operatorname{Var}(f) \sqrt{\log\!\left(1+\frac{1}{\sum_i\operatorname{Inf}_i(f)^2}\right)}. \] We also prove strong one-sided fractional spectral estimates. If $A\subset\{-1,1\}^n$ and \[ h_{A}(x)=\#\{i:x\in A,\ x^{\oplus i}\notin A\}, \] then, for $0<\alpha<1$, \[ \sum_{S\ne\varnothing}|S|^\alpha\widehat{\mathbf 1_{A}}(S)^2 \lesssim_\alpha \mathbb E\omega_\alpha(h_{A}), \] where $\omega_\alpha(m)=\sqrt m$ for $\alpha<1/2$, $\omega_{1/2}(m)=\sqrt m\log(e+m)$, and $\omega_\alpha(m)=m^\alpha$ for $\alpha>1/2$. These profiles are sharp, up to $\alpha$-dependent constants, for majority. We also show that the comparison is genuinely nonreversible: an explicit quotient-cube family makes the optimized fractional, and hence optimized colored, boundary exceed $\mathsf{B}$ by a factor $\gtrsim\sqrt{\log n}$. We further obtain a driftless Bernstein-multiplier inequality.

math.CA

Sharp hypercontractivity for free orthogonal quantum groups of Kac type

We prove that the normalized heat semigroup $P_t=e^{-tL}$ on every free orthogonal quantum group $O_F^+$ of Kac type satisfies hypercontractivity with the optimal time. More precisely, for all $1<p\leq r<\infty$, \[ \|P_t:L_p(O_F^+)\to L_r(O_F^+)\|\leq1 \quad\Longleftrightarrow\quad t\geq\frac12\log\frac{r-1}{p-1}. \] Equivalently, the associated logarithmic Sobolev inequalities hold with sharp constants. For $F=I_N$, this determines the exact optimal time for $O_N^+$ and resolves a conjecture of Brannan, Vergnioux, and Youn \cite{BVY21}.

math.OA

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing

As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern. Existing deception monitors either score visible transcripts or derive scalar probe scores from representation vectors, leaving little inspectable evidence about why a response is suspicious. We introduce STATEWITNESS, an activation explainer for deception auditing. A separate decoder reads a target model's hidden states, then answers natural-language queries or emits structured reports about them. We evaluate STATEWITNESS on two target reasoning LLMs across seven deception datasets. STATEWITNESS reaches 0.916 mean AUROC, a relative gain of 11.6% over the best black-box text monitor and 25.0% over the best activation-probe baseline under the same evaluation protocol. When combined with existing monitors, STATEWITNESS reduces missed deceptive examples in simple threshold ensembles. Beyond scalar detection, the decoder returns query-level answers, schema reports, and token- or sentence-level evidence traces for human inspection. We view this interface as a potential building block for broader interpretability and alignment tools.

cs.CL

Sharp hypercontractivity for free group von Neumann algebras

In this paper, we settle the problem of optimal hypercontractivity for free group von Neumann algebras. Namely, for $n\ge 2$ and the free group $\mathbb{F}_n$ on $n$ generators, we prove that for any $1<p\le q<\infty$, the Poisson semigroup $P_t$ associated with the word-length function satisfies $$ \|P_t:L_p(\widehat{\mathbb{F}_n})\to L_q(\widehat{\mathbb{F}_n})\|\le 1 \qquad\text{ if and only if }\qquad t\ge \frac{1}{2}\log \frac{q-1}{p-1}. $$ The main idea is to apply a refined cubic majorant estimate from a recent work of Frank and Ivanisvili \cite{FrankIvanisvili2026} to the equivalent logarithmic Sobolev inequality, and use the Haagerup-type cancellation estimate \cite{Haagerup1979}. Similar ideas and techniques extend to free products \[ G=\left(*_{\alpha\in A}\mathbb Z\right)*\left(*_{\beta\in B}\mathbb Z_2\right) \] and the free Gaussian von Neumann algebras. In the former setting, partial sharp estimates were previously obtained by Junge--Palazuelos--Parcet--Perrin--Ricard \cite{JungePalazuelosParcetPerrinRicard2015}; in the latter, our approach recovers Biane's free hypercontractivity theorem \cite{Biane1997}.

math.OA

DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair

While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rejection of benign queries that merely appear risky. We present DDOR (Delta Debugging for OverRefusal), a fully automated and explainable framework for overrefusal testing and repair in a black-box setting, where only model inputs and outputs are accessible and internal safety mechanisms remain opaque. DDOR applies delta debugging to localize minimal refusal-triggering fragments (mRTFs) that provide phrase-level, explainable evidence for why a refusal occurs. Conditioned on these mRTFs, DDOR generates diverse, context-rich prompts and performs multi-oracle validation to filter intrinsically unsafe or ambiguous cases, producing scalable and model-specific overrefusal test suites (approximately 1K cases per model). Beyond evaluation, we further leverage localized mRTFs to perform targeted prompt repair, substantially reducing overrefusal while preserving the original intent and maintaining safety on genuinely harmful inputs. Overall, DDOR offers a practical end-to-end solution to both evaluate and mitigate overrefusal, improving LLM usability without sacrificing safety.

cs.SE

Sharp log-Sobolev inequalities and quartic stability on finite cyclic groups

Let $\mathbb Z_n$ be the cyclic group equipped with the uniform probability measure $\pi$, and let $A_{\psi_n}$ be the Laplacian with word length $$ \psi_n(k) = \min(k,n-k). $$ For every $n\ge4$, we prove the sharp log-Sobolev inequality $$ \text{Ent}_{\pi}(|f|^2) \le 2\pi(\bar{f}A_{\psi_n} f), \qquad f:\mathbb Z_n \to \mathbb{C}, $$ where $\text{Ent}_{\pi}$ is the relative entropy with respect to $\pi$. Equivalently, the Poisson semigroup $P_t=e^{-tA_{\psi_n}}$ satisfies the optimal hypercontractivity. The proof is inspired by the recent work of Frank and Ivanisvili [FI26] on a sharp log-Sobolev inequality for the nearest-neighbor simple random walk. Similar arguments yield a simple proof of Weissler's sharp log-Sobolev inequality for the Poisson semigroup on the circle \cite{Weissler1980}. The same inequalities were independently obtained by Yao~\cite{Yao2026} using a different method. We also prove quantitative stability estimates. For $n\ge 4$ and $f:\mathbb Z_n\to[0,\infty)$ with $\pi(f^2)=1$, $$ 2\pi(fA_{\psi_n}f)-\operatorname{Ent}_{\pi}(f^2) \ge \frac1{12}\|f-1\|_{L^2(\pi)}^4, $$ with coefficient $1/(12d)$ on products $(\mathbb Z_n)^d$. The quartic order and the $d^{-1}$ dependence are optimal.

math.CA

Disentangled Learning Improves Implicit Neural Representations for Medical Reconstruction

Implicit neural representations (INRs) have emerged as a powerful paradigm for medical imaging via physics-informed unsupervised learning. Classical INRs optimize an entire network from scratch for each subject, leading to inefficient training and suboptimal imaging quality. Recent initialization-based approaches attempt to inject population priors into pre-trained networks, yet they rely on high-quality images and often suffer from catastrophic forgetting during fine-tuning. We present DisINR, a novel INR framework that explicitly disentangles shared and subject-specific representations. DisINR introduces a shared encoder-decoder pair and subject-specific encoders, whose features are jointly decoded for image reconstruction. By integrating differentiable forward models, it pre-trains the shared modules directly from limited raw measurements, removing the need for pre-acquired high-quality images. During test-time adaptation, only the subject-specific encoder is optimized, while the shared pair remains frozen, effectively preserving learned priors. Extensive evaluations on three representative medical imaging tasks show that DisINR significantly outperforms state-of-the-art INRs in both reconstruction accuracy and efficiency.

cs.CV

Proof of the Holevo--Utkin conjecture on sharp $\ell_p$ norms for zero-sum vectors

Let $d\ge 3$ and $p>0$. Let $\|x\|_p$ denote the $\ell_p$ (quasi-)norm of a $d$-dimensional vector $x$. Holevo and Utkin \cite{HU26} conjectured that for $0<p\le 1$, \[ \min \left\{\frac{\|x\|_p}{\|x\|_2}:\vec{0}\neq x\in\mathbb R^d,\ \sum_{i=1}^d x_i=0\right\} =2^{1/p-1/2}; \] for $1<p<2$, \[ \min \left\{\frac{\|x\|_p}{\|x\|_2}:\vec{0}\neq x\in\mathbb R^d,\ \sum_{i=1}^d x_i=0\right\} = \min\left\{2^{1/p-1/2},\left(\frac{(d-1)^{p/2}+(d-1)^{1-p/2}}{d^{p/2}}\right)^{1/p}\right\}; \] and for $2<q<\infty$ \[ \max\left\{\frac{\|x\|_q}{\|x\|_2}:\vec{0}\neq x\in\mathbb R^d,\ \sum_{i=1}^d x_i=0\right\} = \max\left\{2^{1/q-1/2},\left(\frac{(d-1)^{q/2}+(d-1)^{1-q/2}}{d^{q/2}}\right)^{1/q}\right\}. \] They proved the $d=3$ case in \cite{HU26}. In this paper, we confirm the conjecture of the remaining cases $d\ge 4$.

math.CA

Periodic Steady-State Control of a Handkerchief-Spinning Task Using a Parallel Anti-Parallelogram Tendon-driven Wrist

Spinning flexible objects, exemplified by traditional Chinese handkerchief performances, demands periodic steady-state motions under nonlinear dynamics with frictional contacts and boundary constraints. To address these challenges, we first design an intuitive dexterous wrist based on a parallel anti-parallelogram tendon-driven structure, which achieves 90 degrees omnidirectional rotation with low inertia and decoupled roll-pitch sensing, and implement a high-low level hierarchical control scheme. We then develop a particle-spring model of the handkerchief for control-oriented abstraction and strategy evaluation. Hardware experiments validate this framework, achieving an unfolding ratio of approximately 99% and fingertip tracking error of RMSE = 2.88 mm in high-dynamic spinning. These results demonstrate that integrating control-oriented modeling with a task-tailored dexterous wrist enables robust rest-to-steady-state transitions and precise periodic manipulation of highly flexible objects. More visualizations: https://slowly1113.github.io/icra2026-handkerchief/

cs.RO

Grounded World Model for Semantically Generalizable Planning

In Model Predictive Control (MPC), world models predict the future outcomes of various action proposals, which are then scored to guide the selection of the optimal action. For visuomotor MPC, the score function is a distance metric between a predicted image and a goal image, measured in the latent space of a pretrained vision encoder like DINO and JEPA. However, it is challenging to obtain the goal image in advance of the task execution, particularly in new environments. Additionally, conveying the goal through an image offers limited interactivity compared with natural language. In this work, we propose to learn a Grounded World Model (GWM) in a vision-language-aligned latent space. As a result, each proposed action is scored based on how close its future outcome is to the task instruction, reflected by the similarity of embeddings. This approach transforms the visuomotor MPC to a VLA that surpasses VLM-based VLAs in semantic generalization. On the proposed WISER benchmark, GWM-MPC achieves a 87% success rate on the test set comprising 288 tasks that feature unseen visual signals and referring expressions, yet remain solvable with motions demonstrated during training. In contrast, traditional VLAs achieve an average success rate of 22%, even though they overfit the training set with a 90% success rate.

cs.RO

The Boolean surface area of polynomial threshold functions

Polynomial threshold functions (PTFs) are an important low-complexity class of Boolean functions, with strong connections to learning theory and approximation theory. Recent work on learning and testing PTFs has exploited structural and isoperimetric properties of the class, especially bounds on average sensitivity, one of the central themes in the study of PTFs since the Gotsman--Linial conjecture. In this work we study PTFs through the lens of the Boolean surface area (or Talagrand boundary) \[ \mathbf{BSA}[f]=\mathbb{E}|\nabla f|=\mathbb{E}\sqrt{s_{f}(x)}, \] a natural measure of vertex-boundary complexity on the discrete cube. Our main result is that every degree-$d$ PTF has polylogarithmic Boolean surface area: \[ \mathbf{BSA}[f]\le C_d(\log(en))^{C_d}. \] The proof is based on the PTF Restriction Lemma of Kabanets, Kane, and Lu \cite{KKL2017} and proceeds through a tail bound for the pointwise sensitivity. In particular, it controls all subcritical fractional moments of the sensitivity. We also record a random block partition principle for Boolean surface area and an alternative recursive argument following Kane's work \cite{DK} on average sensitivity, which independently yields the weaker bound \[ \mathbf{BSA}[f]\le \exp(C_d\sqrt{\log n}). \]

cs.CC

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models

Video Large Language Models (VLLMs) demonstrate strong video understanding but suffer from inefficiency due to redundant visual tokens. Existing pruning primary targets intra-frame spatial redundancy or prunes inside the LLM with shallow-layer overhead, yielding suboptimal spatiotemporal reduction and underutilizing long-context compressibility. All of them often discard subtle yet informative context from merged or pruned tokens. In this paper, we propose a new perspective that elaborates token \textbf{A}nchors within intra-frame and inter-frame to comprehensively aggregate the informative contexts via local-global \textbf{O}ptimal \textbf{T}ransport (\textbf{AOT}). Specifically, we first establish local- and global-aware token anchors within each frame under the attention guidance, which then optimal transport aggregates the informative contexts from pruned tokens, constructing intra-frame token anchors. Then, building on the temporal frame clips, the first frame within each clip will be considered as the keyframe anchors to ensemble similar information from consecutive frames through optimal transport, while keeping distinct tokens to represent temporal dynamics, leading to efficient token reduction in a training-free manner. Extensive evaluations show that our proposed AOT obtains competitive performances across various short- and long-video benchmarks on leading video LLMs, obtaining substantial computational efficiency while preserving temporal and visual fidelity. Project webpage: https://tyroneli.github.io/AOT.

cs.CV

Zero-shot Low-Field MRI Enhancement via Diffusion-Based Adaptive Contrast Transport

Low-field (LF) magnetic resonance imaging (MRI) democratizes access to diagnostic imaging but is fundamentally limited by low signal-to-noise ratio and significant tissue contrast distortion due to field-dependent relaxation dynamics. Reconstructing high-field (HF) quality images from LF data is a blind inverse problem, severely challenged by the scarcity of paired training data and the unknown, non-linear contrast transformation operator. Existing zero-shot methods, which assume simplified linear degradation, often fail to recover authentic tissue contrast. In this paper, we propose DACT(Diffusion-Based Adaptive Contrast Transport), a novel zero-shot framework that restores HF-quality images without paired supervision. DACT synergizes a pre-trained HF diffusion prior to ensure anatomical fidelity with a physically-informed adaptive forward model. Specifically, we introduce a differentiable Sinkhorn optimal transport module that explicitly models and corrects the intensity distribution shift between LF and HF domains during the reverse diffusion process. This allows the framework to dynamically learn the intractable contrast mapping while preserving topological consistency. Extensive experiments on simulated and real clinical LF datasets demonstrate that DACT achieves state-of-the-art performance, yielding reconstructions with superior structural detail and correct tissue contrast.

cs.CV