SearcharxivSearch

arXiv subjects

Xiang Fang

Publications and source records attributed to Xiang Fang.

At least 19 recordsLinked to original sources

The Local Embedding Problem for Hardy Spaces of Dirichlet Series

We solve the local embedding problem for Hardy spaces of Dirichlet series, which is a dimension-free trace problem asking whether the global $\mathscr{H}^p$-norm controls local $L^p$-mass on the critical line $\operatorname{Re}s=1/2$. More precisely, for every $2<p<\infty$, there exists a constant $C_p<\infty$ such that every Dirichlet polynomial $P$ satisfies $$ \sup_{\theta\in\mathbb{R}}\int_{\theta}^{\theta+1}\left|P\left(\frac12+it\right)\right|^p\,\mathrm{d}t\le C_p\left\lVert P\right\rVert_{\mathscr{H}^p}^{p}, $$ with $C_p$ independent of the number and choice of prime variables on which $P$ depends. Before the present work, the embedding was known at $p=2$ and, by taking integer powers, at the even exponents $p=2k$; it had been conjectured that these exhaust the finite positive cases above $2$. Together with the known failure for $0<p<2$, our theorem gives the sharp finite-exponent classification: the local embedding property holds exactly for $p\ge2$. Thus the true threshold is $p=2$, rather than even integrality. The proof passes to the dual exponent $q=p/(p-1)\in(1,2)$, where an exact frequency decomposition isolates a single resonant Euler-product term. A covariance-preserving replacement of the shared prime factors reduces the resulting moment estimate to a log-correlated Gaussian field, and a critical branching-random-walk bound supplies the required multiscale decay. A finite-cyclic square-function estimate assembles the resonant scales, and Hardy-quotient duality converts the resulting vector-valued bound into the critical-line trace. For $1\le p<\infty$, known equivalences give the same sharp threshold in several classical problems, including the conformally invariant half-plane embedding, the reverse local Carleson-measure transfer, and boundedness of all characteristic-zero Gordon--Hedenmalm composition operators.

math.FA

Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain

Multimodal large language models (MLLMs) have extended the capability of large language models (LLMs) to process more contextual multimodal information, showing remarkable progress in diverse realistic multimodal applications. Despite their strong perception and reasoning abilities, recent studies reveal that MLLMs remain highly vulnerable to adversarial inputs, especially those targeting visual components. However, existing attacks mainly focus on global perturbations, lacking an understanding of how MLLMs internally interpret visual structures. In this paper, we make the attempt to investigate the intrinsic focus of MLLMs in the frequency domain and discover that their predictions are particularly sensitive to phase information, which encodes essential structural and semantic cues. Based on this observation, we propose a novel phase-aware adversarial attack framework that explicitly restricts adversarial perturbations to structure-relevant phase regions to suppress the MLLMs' focus for effective and imperceptible attacks. To further amplify the structural influence, we also introduce an auxiliary adversarial prompt learning module to guide multimodal misalignment around phase-sensitive regions, misleading the MLLM's attention toward targeted structural patterns. Extensive experiments on multiple representative MLLM models and datasets demonstrate the superior effectiveness of our method compared to existing attacks.

cs.CV

Critical Gaussian Multiplicative Chaos on the Circle Is Rajchman

We prove that the Fourier coefficients of the canonical critical Gaussian multiplicative chaos on the circle vanish almost surely at infinity. More precisely, let $M_\phi^{\mathrm{crit}}$ be the canonical critical chaos associated with the centered circle field $\phi$ of covariance $\mathbb{E}[\phi(\theta)\phi(\theta')] = \log\frac{1}{\lvert e^{i\theta}-e^{i\theta'}\rvert}$. Then, almost surely, $\widehat{M_\phi^{\mathrm{crit}}}(n)\longrightarrow0$ as $\lvert n\rvert\to\infty$. This resolves the almost-sure critical Rajchman problem for the canonical circle field. Since critical chaos has Fourier dimension zero almost surely, no positive polynomial Fourier-decay rate can hold; the theorem therefore exhibits qualitative Fourier cancellation beyond the regime of positive Fourier dimension. The proof addresses two coupled difficulties: the heavy, nonuniform cell masses of critical chaos and the need to control exponentially many frequencies in each dyadic annulus. For an auxiliary periodized compact-range star-scale field, a derivative-rooted Bessel regression yields weighted small-cell summability and moving-tail control of exceptional large cells. After conditioning at a coarse scale below the Fourier scale, finite-range independence and conditional Bernstein concentration reduce uniform control of the terminal Fourier coefficients over each dyadic annulus to a spatial-variation estimate for a coarse predictable measure. A smooth positive-definite covariance correction and critical-chaos uniqueness then transfer the Rajchman property to the canonical critical chaos of the exact circle field.

math.PR

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Estimating both from target rollouts across four domains, four open-weight targets, and a frontier API target yields three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. These findings separate the value of short-range conditioning from proposal quality.

cs.LG

Critical Norm Profiles for Finite-Prime Composition Operators on the Hardy Space of Dirichlet Series

We identify the critical boundary operator-norm profile of finite-prime composition operators on the Hardy--Hilbert space \(\mathcal H^2\) of Dirichlet series. For \[ \varphi_{\delta,\boldsymbol\rho}(s) = \frac12+\delta + \delta\sum_{j=1}^d\rho_jp_j^{-s}, \qquad \boldsymbol\rho\in B_d, \] the renormalized positive coefficient operators converge uniformly in operator norm, with \(O(\delta)\) error, to an explicit multivariate weighted Hankel operator \(\mathcal H_{\boldsymbol\rho}\); consequently, \[ 2\delta\|C_{\varphi_{\delta,\boldsymbol\rho}}\|^2 = \|\mathcal H_{\boldsymbol\rho}\| + O(\delta) \] uniformly over \(B_d\). We show that the limiting operator admits the total-degree reduction \[ \mathcal H_{\boldsymbol\rho} \simeq D_{\boldsymbol\rho} H_{R_{\boldsymbol\rho}/2} D_{\boldsymbol\rho}\oplus\mathbf{0}, \] where the diagonal factors are convolution-collision norms of the normalized prime weights. This structure, together with the affine comparison principle of Brevig and Perfekt, yields an explicit concentration inequality for \(\|\mathcal H_{\boldsymbol\rho}\|\), identifies the one-prime configurations as the exact equality cases in the limiting norm estimate, and gives a quantitative deficit away from them. For fixed \(\sigma>\frac12\), we also obtain a second-order expansion of the squared norm and fully finite-dimensional approximations with explicit total-degree and Dirichlet-sum truncation errors. Together, these results show that a single coefficient-operator structure governs the singular boundary profile, the fixed-\(\sigma\) perturbative regime, and certified finite-dimensional approximation.

math.FA

Hardy-Szeg\H{o} Point Processes: Large Deviations and Strong Szeg\H{o} Asymptotics

We study exponential-scale fluctuations of the Hardy-Szeg\H{o} zero process, investigated by Peres and Vir\'ag in the disk setting, in its upper-half-plane realization, which reveals a different probabilistic geometry. This conformally invariant determinantal zero process is equivalent to its disk realization, but the upper-half-plane coordinates make real-translation invariance explicit and single out long horizontal windows as natural observables. For the vertical window from height one to height $a>1$, let $N_a(L)$ denote the number of zeros in the corresponding horizontal window of length $L$. We identify the limiting scaled log-moment generating function explicitly. As a consequence, we prove a large deviations principle for $N_a(L)/L$, with rate function given by the Legendre transform of this limit. We also prove a strong Szeg\H{o} expansion for the log-moment generating function, including an explicit order-one correction, locally uniformly in the natural complex strip. The proof uses the planar determinantal structure before projection, reduces the problem to a one-dimensional Fredholm determinant, and combines fixed-power trace asymptotics with a Wiener-Hopf comparison.

math.PR

Two Regularity Problems on Analytic Tent Spaces

We study two regularity problems on Hardy-type analytic tent spaces $\mathcal{AT}^p_{q,\alpha}$ on the unit disk: fractional integration and randomization of Taylor coefficients. For fractional integration, we characterize completely the boundedness and compactness of the Hadamard, Flett, and Riemann--Liouville operators between analytic tent spaces, and obtain parallel results for analytic Triebel--Lizorkin spaces. In particular, the case $t=0$ yields a complete solution to the corresponding embedding problem for analytic tent spaces. For randomization, we characterize completely when the random Taylor series $\mathcal{R}f$ belongs almost surely to an analytic tent space whenever $f\in \mathcal{AT}^p_{q,\alpha}$, and we also obtain the Triebel--Lizorkin counterpart. As part of the proof, we identify the random symbol space associated with $\mathcal{AT}^p_{q,\alpha}$ and solve the embedding problem from analytic tent spaces into mixed norm spaces. These results extend classical theorems of Hardy--Littlewood and Littlewood, as well as their later analogues for Bergman and mixed norm spaces.

math.FA

Canonical Mandelbrot Cascades on Curves Are Rajchman

We settle the Rajchman problem for canonical scalar dyadic Mandelbrot cascades at the minimal Kahane--Peyri\`ere integrability threshold. If $\mu$ is the cascade on $[0,1]$, then $\widehat{\mu}(\xi)\to 0$ as $|\xi|\to\infty$, almost surely on non-extinction. For every fixed nondegenerate $C^2$ embedded arc $\gamma:[0,1]\to\mathbb{R}^2$, the pushforward $\gamma_\#\mu$ is likewise Rajchman almost surely on non-extinction. The analogous conclusion holds for the scalar cascade on the parameter circle pushed forward by any fixed nondegenerate $C^2$ Jordan curve. No moment condition of order strictly greater than one is imposed; in particular, the results include the regime $\mathbb{E}[W^q]=\infty$ for every $q>1$. The proof combines a spine-based lower-deviation principle, adaptive terminal approximation, and predictable capping to obtain almost-sure estimates uniform over large frequency annuli without higher moments. For curved pushforwards, an endpoint-safe phase decomposition controls direction-dependent stationary regions, including those meeting the endpoints of an arc, and couples the geometric and probabilistic arguments through a common dyadic kernel. Combined with the exact Fourier-dimension formulas for the corresponding models, the theorems show that Rajchman decay persists at zero Fourier dimension.

math.PR

A re-entrant chip-free-space photonic interface for telecom-to-Rubidium spectroscopy

Photonic integrated circuits (PICs) generate, route, and process light with high efficiency, scalability, and functional density on a single chip. Yet the tightly confined on-chip modes can not easily access or effectively interact with atomic vapors, fluids, gain media, and biological samples. Existing approaches require bringing the medium onto the chip or into a weak, tightly confined evanescent field, which restricts the interaction volume and the range of accessible media. Here, we demonstrate a re-entrant chip-free-space interface in which a thin-film lithium niobate circuit frequency-doubles telecom light, emits the 780~nm field through a Rubidium vapor cell, and recollects the reflected probe on the same chip. This emit-interact-recollect loop resolves the saturated absorption spectrum and stabilizes the telecom laser to within $\pm 280$~kHz over 2 hours. Our study paves an route to embed external media into PICs through the re-entrant photonic interface.

physics.optics

Exact Fourier dimensions of dyadic Mandelbrot cascades on curves of nonvanishing curvature under minimal integrability

We prove exact Fourier-dimension formulas for scalar dyadic Mandelbrot cascades pushed forward to fixed nondegenerate $C^2$ embedded arcs and fixed nondegenerate $C^2$ Jordan curves in $\mathbb R^2$. Let $W$ be in the minimal Kahane--Peyriere regime. For each fixed nondegenerate $C^2$ embedded arc $\gamma:[0,1]\to\mathbb R^2$, the pushforward $\mu_\gamma$ of the interval cascade satisfies, almost surely on non-extinction, \[ \dim_{\mathrm F}(\mu_\gamma)=A_{\mathrm{loc}}(W), \] where \[ A_{\mathrm{loc}}(W) = \sup_{q>1} \max\left\{ 0,\, \frac{q-1-\log_2\mathbb E[W^q]}{q} \right\}, \] with the $q$-term interpreted as $0$ when $\mathbb E[W^q]=\infty$. The analogous formula holds for scalar circle cascades pushed forward by fixed nondegenerate $C^2$ Jordan curves $\gamma:\mathbb T\to\mathbb R^2$, with the pushforward denoted by $\mu_\gamma^{\mathbb T}$. This extends the scalar circle endpoint formula from the canonical circle to fixed parametrized arcs and Jordan curves. The main new issue beyond the canonical circle is the loss of the explicit trigonometric phase and, for arcs, the presence of endpoint stationary regimes. We prove the arc lower bound by a finite-$r$ annular Fourier theorem based on an endpoint-safe phase decomposition, phase-bin coefficient estimates, predictable capping, complex Freedman concentration, and an $r$-tail compensator. The Jordan lower bound follows by first-generation dyadic cutting into two fixed arcs. The matching upper bounds use deterministic curved-support obstructions together with the scalar-circle minimum lower local-dimension theorem.

math.PR

Exact Fourier dimensions of dyadic Mandelbrot cascades under minimal integrability

We determine the Fourier dimension of dyadic Mandelbrot cascades under the minimal Kahane-Peyriere integrability condition. The interval theorem is proved in a vector-valued dyadic cascade model in which sibling weights may have arbitrary dependence. For every balanced energy-admissible vector law, almost surely on non-extinction, dim_F(mu)=dim_E(mu)=dim_2(mu)=D_E(X). In the canonical scalar case, under W>=0, E W=1, E[W log_2^+ W] 1} max{0, (q-1-log_2 E[W^q])/q}. The interval and circle formulas share a light-tail/heavy-tail dichotomy but have different mechanisms: energy dimension for the interval, and minimum lower local dimension for the circle. The circle lower bound follows from a finite-moment annular Fourier theorem.

math.PR

Fourier Dimensions of Mandelbrot Cascades under Minimal Integrability

This note announces exact Fourier dimension formulas for canonical Mandelbrot cascade measures under the minimal Kahane Peyriere integrability condition and records the canonical b adic extension on cubes. In the dyadic interval setting, the theorem is proved in a balanced vector weight model allowing dependence between sibling weights. Almost surely on non extinction, the Fourier, energy, and L2 dimensions all equal the energy exponent. The scalar specialization gives the canonical Mandelbrot Kahane Fourier dimension formula under the minimal integrability condition. On the circle, the endpoint formula is given by the endpoint lower local dimension exponent. For the b adic Mandelbrot cascade on cubes, the Fourier dimension is the minimum of 2 and the energy exponent, with the universal Fourier barrier at dimension two providing the high dimensional obstruction.

math.PR

Two Problems in Bergman Spaces with Non-radial Weights

This paper investigates two problems unified by the study of the uniform boundedness of the dilation operators (UBD) T_r f(z)=f(rz), 0<r<1, acting on weighted Bergman spaces A^p_omega with not necessarily radial weights. We first characterize the random symbol space for A^p_omega under a mild admissible condition (Theorem 1.2). This extends the main result of [7] from radial weights to non-radial weights. We then introduce two new notions, namely non-radial mixed norm spaces M(p,q;omega) and analytic tent spaces A(p,q;omega), and we characterize their corresponding symbol spaces as well (Theorem 1.8, Theorem 1.12). The novelty here is to employ a measure-disintegration framework to prove a general Littlewood-type theorem (M(p,q;omega))* = H(2,q;omega_r), from which the Bergman space result (A^p_omega)* = H(2,p;omega_r) follows as the special case p=q. Among other things, UBD plays a pivotal role in the proofs of the preceding theorems. The second main problem addressed in this paper is to establish a local-to-global criterion for UBD, which remains largely unexplored for non-radial weights. Our principle result in this part (Theorem 1.16) asserts that UBD is guaranteed by two local geometric conditions: bounded hyperbolic oscillation (BHO) and a reverse-Carleson tail condition (RC). This is the technical heart of the paper. In the course of our investigation, three new types of problems arise naturally, each of independent interest: a two-weight top-maximal operator (Theorem 1.17); a truncated maximal operator over hyperbolic balls (Theorem 1.18); and a single-testing Carleson embedding problem (Theorem 1.19).

math.CV

Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation

Vision-Language Navigation in Continuous Environments (VLN-CE) poses a formidable challenge for autonomous agents, requiring seamless integration of natural language instructions and visual observations to navigate complex 3D indoor spaces. Existing approaches often falter in long-horizon tasks due to limited scene understanding, inefficient planning, and lack of robust decision-making frameworks. We introduce the \textbf{Hierarchical Semantic-Augmented Navigation (HSAN)} framework, a groundbreaking approach that redefines VLN-CE through three synergistic innovations. First, HSAN constructs a dynamic hierarchical semantic scene graph, leveraging vision-language models to capture multi-level environmental representations, from objects to regions to zones, enabling nuanced spatial reasoning. Second, it employs an optimal transport-based topological planner, grounded in Kantorovich's duality, to select long-term goals by balancing semantic relevance and spatial accessibility with theoretical guarantees of optimality. Third, a graph-aware reinforcement learning policy ensures precise low-level control, navigating subgoals while robustly avoiding obstacles. By integrating spectral graph theory, optimal transport, and advanced multi-modal learning, HSAN addresses the shortcomings of static maps and heuristic planners prevalent in prior work. Extensive experiments on multiple challenging VLN-CE datasets demonstrate that HSAN achieves state-of-the-art performance, with significant improvements in navigation success and generalization to unseen environments.

cs.RO

Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval

Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear interactions between temporal video sequences and textual semantics. Existing approaches, relying on static cross-attention or prompt-tuning mechanisms, fail to adaptively model the evolving relationships between modalities, leading to suboptimal alignment and limited generalization. Inspired by systems biology, we propose \textbf{Reaction-Diffusion Multimodal Fusion (RDMF)}, a novel framework that reimagines video-language alignment as a reaction-diffusion (RD) process, drawing on the principles of pattern formation introduced by Alan Turing. In RDMF, video features diffuse across time to capture temporal context, while text-video interactions are modeled as non-linear reactions that amplify relevant features and suppress noise, forming emergent patterns akin to biological systems. Leveraging the Gray-Scott RD model, we design a computationally efficient fusion module that integrates video and text representations, supported by rigorous mathematical analysis of stability and convergence using Turing instability criteria. Our framework is theoretically grounded, employing advanced mathematical tools to ensure stable pattern formation, and is practically viable, incorporating standard components like pretrained encoders and DETR-style heads for moment retrieval and saliency prediction. RDMF represents a pioneering interdisciplinary approach, bridging systems biology and multimedia research to address the limitations of conventional multimodal fusion. Preliminary experiments demonstrate its potential to outperform existing methods in identifying salient video moments, offering a new paradigm for video-language tasks.

cs.CV

Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding

This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topic, they severely rely on massive expensive video-query paired annotations, which require a tremendous amount of human effort to collect in real-world applications. To this end, in this paper, we target a more practical but challenging TSG setting: unsupervised temporal sentence grounding, where both paired video-query and segment boundary annotations are unavailable during the network training. Considering that some other cross-modal tasks provide many easily available yet cheap labels, we tend to collect and transfer their simple cross-modal alignment knowledge into our complex scenarios: 1) We first explore the entity-aware object-guided appearance knowledge from the paired Image-Noun task, and adapt them into each independent video frame; 2) Then, we extract the event-aware action representation from the paired Video-Verb task, and further refine the action representation into more practical but complicated real-world cases by a newly proposed copy-paste approach; 3) By modulating and transferring both appearance and action knowledge into our challenging unsupervised task, our model can directly utilize this general knowledge to correlate videos and queries, and accurately retrieve the relevant segment without training. Extensive experiments on two challenging datasets (ActivityNet Captions and Charades-STA) show our effectiveness, outperforming existing unsupervised methods and even competitively beating supervised works.

cs.CV

Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness

Large Vision-Language Models have achieved unprecedented success in zero-shot recognition by aligning visual features with broad semantic concepts. However, this semantic abstraction creates a critical vulnerability in open-world deployment: the ``Hubris of Semantics'', where models force-fit unknown anomalies into known categories with high confidence due to the lack of explicit negative knowledge. To address this \textit{Open-World Trustworthiness Paradox}, we propose \textbf{Immuno-VLM}, a bio-inspired framework that adapts the biological principle of \textbf{Immunological Negative Selection} to high-dimensional latent spaces. Departing from traditional Open-Set Recognition methods that rely on passive density estimation or inefficient pixel-space outlier generation, Immuno-VLM leverages the generative reasoning of Large Language Models to actively hallucinate ``Semantic Antibodies'', textual descriptions of near-distribution outliers (e.g., look-alikes, contextual anomalies) that effectively bound the decision space of known classes.Extensive experiments on ImageNet-1K and four challenging OOD benchmarks reveal that Immuno-VLM establishes a new state-of-the-art.

cs.CV

SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling

In the era of Large Video-Language Models (LVLMs), the computational necessity of sparse frame sampling creates a fundamental ``temporal gap'', rendering models blind to critical causal transitions. Existing solutions relying on generative hallucination (e.g., latent diffusion) or autoregressive extrapolation often fail to maintain semantic consistency over long horizons, suffering from object vanishing and energetic instability. We propose a paradigm shift from probabilistic generation to variational mechanics with the \textbf{Semantic Least Action Principle (SLAP)}. Drawing a rigorous isomorphism between classical mechanics and semantic dynamics, we model the latent video trajectory as a path on a Riemannian manifold governed by a Semantic Lagrangian. By formulating the interpolation task as a Boundary Value Problem (BVP) solved via the discrete Euler-Lagrange equations, SLAP naturally enforces object persistence without pixel-level rendering. Extensive experiments show the effectiveness of our proposed SLAP.

cs.CV