SearcharxivSearch

arXiv subjects

Sibei Yang

Publications and source records attributed to Sibei Yang.

At least 19 recordsLinked to original sources

$\Gamma$-Convergence of Weak-Type Nonlocal Functionals on Bounded Domains

Let $N\ge1$, $p\in[1,\infty)$, $\gamma\in(0,\infty)$, and $\Omega\subset\mathbb R^N$ be a bounded open interval when $N=1$ or a bounded Lipschitz domain when $N\ge2$. For any $\lambda\in(0,\infty)$ and any measurable function $u$, consider the weak-type nonlocal functional \begin{align*} G_{\lambda,p,\gamma}(u;\Omega) :=\lambda\iint_{\Omega\times\Omega} \mathbf 1_{\left\{(x,y)\in\Omega\times\Omega:\ x\neq y,\ \frac{|u(x)-u(y)|^p}{|x-y|^{p+\gamma}}\geq\lambda\right\}} |x-y|^{\gamma-N}\,dx\,dy. \end{align*} In this article, we prove that, as $\lambda\to\infty$, the family $G_{\lambda,p,\gamma}$ converges, in the sense of $\Gamma$-convergence in $L^p(\Omega)$, to the functional \begin{align*} \Psi_{p,\gamma}^{\mathrm{cell}}(u;\Omega):= \begin{cases} C_{N,p,\gamma}^{\mathrm{cell}}\displaystyle\int_\Omega|\nabla u|^p\,dx, &p\in(1,\infty)\ \hbox{and}\ u\in W^{1,p}(\Omega),\\[2mm] C_{N,1,\gamma}^{\mathrm{cell}}|Du|(\Omega), &p=1\ \hbox{and}\ u\in BV(\Omega),\\[1mm] \infty,&\hbox{otherwise}, \end{cases} \end{align*} where the positive constants $C_{N,p,\gamma}^{\mathrm{cell}}$ are independent of $\Omega$ and characterized by a cell formula. This gives an affirmative answer to the problem posed by Brezis [Open Problem~9.3, Rend. Lincei Mat. Appl. 2023].

math.CA

Lattice point counting in Cygan--Kor\'anyi balls on Heisenberg groups

Lattice point counting in gauge balls on the Heisenberg group $\mathbb{H}^q$ is a non-commutative analogue of the Euclidean multidimensional sphere problem, initiated by Garg, Nevo and Taylor \cite[\textit{Ann. Inst. Fourier}, 2015]{GNT15}. The case of particular interest is when the gauge is taken as the Cygan--Kor\'anyi norm and the error term reads: $$\mathcal{E}_q(t)=\#\left(\mathbb{Z}^{2 q+1} \cap \mathcal{B}_t\right)-\operatorname{vol}(\mathcal{B}_1) \, t^{2 q+2},$$ with $\mathcal{B}_t=\{(v,w)\in\mathbb{H}^q: (|v|^4 + w^2)^{1/4} \le t \}$, which is closely related to the Gauss circle problem. When $q\ge3$, Gath \cite[\textit{Ann. Sc. Norm. Super. Pisa Cl. Sci.}, 2022]{Gat22} improved upon \cite[]{GNT15} by showing that $ |\mathcal{E}_q(t)|\lesssim t^{2q-1+ 1/3}$ and proposed the conjecture that the optimal order should be $2q-1$. In this paper, through Landau's formula and the $5,6$-th Derivative Tests of van der Corput, we arrive at that $|\mathcal{E}_q(t)| \lesssim t^{2 q-1 + 241/753} $ for any $ q \geq 4$, and recover the bound of Gath for $q=3$ up to a logarithmic factor. This, via a simpler method, provides the first progress towards Gath's conjecture.

math.NT

Sobolev--Morrey Spaces and Divergence-Form Degenerate Second-Order Elliptic Equations on Domains with Higher Co-Dimensional Boundaries

In this article, we study the weighted homogeneous Sobolev--Morrey spaces on domains in $\mathbb{R}^n$ with higher co-dimensional boundaries. Precisely, we systematically establish a real-variable theory of these spaces, including completeness, embedding theorems, Riesz potential characterizations, continuity, trace and extension theorems, and complex interpolation. Applying the boundedness of the trace and the extension operators, we obtain sharp weighted a priori estimates for solutions to the Dirichlet problem of divergence-form degenerate second-order elliptic equations on such domains in weighted Lebesgue spaces. The absence of a boundary manifold structure of these domains poses some essential difficulties, which are overcome by using some tools, such as the intrinsic properties of distance weights and the geometric structure of domains, different from those available in Lipschitz domains.

math.AP

A Counterexample to the Necessity of the Vanishing Carleson Condition for VMO Poisson Kernels

In [Problem 3.2.23, CBMS Regional Conference Series in Mathematics 83, 1994], Kenig asked whether the vanishing Carleson condition is the necessary and sufficient for the logarithm of the Poisson kernel of a perturbation of the Laplacian on the unit ball in $\mathbb{R}^n$ belonging to the space VMO. The sufficiency was proved by Escauriaza [Israel J. Math. 1996] and extended by Milakis, Pipher, and Toro [Contemp. Math. 2014] to more general setting. In this article, using the technique of bi-Lipschitz mappings, we construct a counterexample to show that the vanishing Carleson condition is not necessary and hence give a negative answer to the aforementioned problem.

math.AP

Bridging the Agent-World Gap: Text World Models for LLM-based Agents

Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet many remain largely reactive, mapping observations to actions without an explicit model of how these environments are structured and evolve. This motivates text world models (TWMs): transition models over textual states that, given a state and a candidate action, predict the resulting webpage, terminal output, API response, or user reply, thereby supporting planning, efficient learning, and principled evaluation. We systematically review text world models for LLM-based agents, organized around a formal framework and the agent lifecycle: (1) Foundations, defining text world models and characterizing them by state representation and grounding domain; (2) Construction, taxonomizing LLM-as-WM and code-as-WM paradigms and reviewing methods for building them; (3) Application, examining how world models support agents at training time through experience synthesis and at inference time through planning, verification, and adaptation; and (4) Evaluation, covering both evaluation of the world model itself and its use as an evaluation environment for agents. We aim to consolidate this rapidly developing area, clarify its design space, and highlight open challenges for future research.

cs.CL

Self-Prophetic Decoding to Unlock Visual Search in LVLMs

Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of the thinking-with-images paradigm. However, LVLM visual search faces two key challenges: incompatibility among intrinsic capabilities after post-training, and interference in long multi-step reasoning contexts. To address these, we identify two novel insights. First, self-regulation between pre- and post-training LVLMs leverages the intrinsic single-step capabilities of the pre-training model to mitigate capability deterioration and long-context interference. Second, probability-based prophetic sampling, replacing naive prompting, provides a probabilistic interface where the pre-training model acts as a prophet and the post-training model selectively accepts prophetic tokens under its output distribution, preserving coherent multi-step reasoning. Building on these insights, we introduce SeProD, a self-prophetic decoding framework that leverages intrinsic single-step capabilities to enable coherent multi-step reasoning in a training-free, plug-and-play manner. Experiments show that SeProD consistently improves multiple visual-search LVLMs across all 12 splits of 4 visual search benchmarks, as well as across general VQA benchmarks, without added computational overhead, thanks to its parallel prophetic acceptance mechanism.

cs.CV

A Counterexample to Kenig's Interpolation Problem for Sobolev Spaces with Zero Boundary Conditions

Let $n\in \mathbb N\cap[2,\infty)$. In this article, we show that there exists a bounded $C^1$ domain $\Omega\subset \mathbb R^n$ such that, for any given $s\in(1,2)\setminus\{\frac32\}$, \begin{align*} \left[H_0^1(\Omega),H^2(\Omega)\cap H_0^1(\Omega)\right]_{s-1} =H^s(\Omega)\cap H_0^1(\Omega)=H_0^s(\Omega) \end{align*} with equivalent norms, but \begin{align*} \left[H_0^1(\Omega),H^2(\Omega)\cap H_0^1(\Omega)\right]_{\frac12} \subsetneqq H^{\frac32}(\Omega)\cap H_0^1(\Omega), \end{align*} which provides a counterexample to Problem 3.3.19 of Kenig in [CBMS Regional Conf. Ser. in Math. 83, 1994]. As applications, we prove that for such a domain $\Omega$ \begin{align*} H^2(\Omega)\cap H_0^1(\Omega)\subsetneqq D(-\Delta_D) \end{align*} (the domain of the Dirichlet Laplacian operator $-\Delta_D$ on $\Omega$) and construct a solution of the homogeneous heat equation with zero Dirichlet boundary condition, which does not belong to $L^2((0,T);H^2(\Omega)\cap H_0^1(\Omega))$ for any given $T\in(0,\infty)$.

math.AP

EVA: Editing for Versatile Alignment against Jailbreaks

Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks, where adversaries exploit textual or visual triggers to bypass safety guardrails. Recent defenses typically rely on safety fine-tuning or external filters to reduce the model's likelihood of producing harmful content. While effective to some extent, these methods often incur significant computational overheads and suffer from the safety utility trade-off, degrading the model's performance on benign tasks. To address these challenges, we propose EVA (Editing for Versatile Alignment against Jailbreaks), a novel framework that pioneers the application of direct model editing for safety alignment. EVA reframes safety alignment as a precise knowledge correction task. Instead of retraining massive parameters, EVA identifies and surgically edits specific neurons responsible for the model's susceptibility to harmful instructions, while leaving the vast majority of the model unchanged. By localizing the updates, EVA effectively neutralizes harmful behaviors without compromising the model's general reasoning capabilities. Extensive experiments demonstrate that EVA outperforms baselines in mitigating jailbreaks across both LLMs and VLMs, offering a precise and efficient solution for post-deployment safety alignment.

cs.CR

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization

Diffusion Vision-Language Models (dVLMs), built upon the non-causal foundations of Diffusion Large Language Models (dLLMs), have demonstrated remarkable efficacy in multimodal tasks by departing from the traditional autoregressive generation paradigm. While dVLMs appear inherently robust against conventional jailbreak tactics, which we categorize as Fixed Prefix Optimization (FPO) (e.g., anchoring responses with "Sure, here is"), this perceived resilience is deceptive. Our investigation into the safety landscape of dVLMs reveals a unique refusal pattern: Immediate Refusal and Progressive Refusal. We find that while FPO-based attacks often fail by triggering the latter, the progressive refinement process itself uncovers a novel, latent attack surface. To exploit this vulnerability, we propose Global Probability Optimization (GPO), a general jailbreak paradigm designed specifically for the denoising trajectory of masked diffusion models. Unlike prefix-based methods, GPO manipulates the global generative dynamics to bypass guardrails in diffusion language models. Building on this, we introduce GPO-V, the first visual-modality jailbreak framework tailored for dVLMs. Empirical results demonstrate that GPO-V produces stealthy perturbations with exceptional cross-model transferability, revealing a critical security gap in non-sequential generative architectures. Our findings underscore the critical urgency of addressing safety alignment in dVLMs. These results necessitate an immediate and fundamental re-evaluation of current defense paradigms to mitigate the unique risks of diffusion-based generation. Our code is available at: https://anonymous.4open.science/r/GPO-V-0250.

cs.CV

Chart Deep Research in LVLMs via Parallel Relative Policy Optimization

With the rapid advancement of data science, charts have evolved from simple numerical presentation tools to essential instruments for insight discovery and decision-making support. However, current chart data intelligence exhibits significant limitations in deep research capabilities, with existing methods predominantly addressing shallow tasks such as visual recognition or factual question-answering, rather than the complex reasoning and high-level data analysis that deep research requires. This limitation stems from two primary technical bottlenecks: at the training level, existing post-training techniques exhibit deficiencies in handling multi-dimensional reward signal interference and heterogeneous data gradient conflicts, preventing models from achieving balanced development across multiple capability dimensions; at the evaluation level, current methods remain limited to factual retrieval and basic computation, failing to assess end-to-end analytic reasoning and other deep research capabilities. To address the training challenge, we propose PRPO, which performs parallel optimization across reward dimensions and capability partitioning across data types, effectively disentangling conflicts between heterogeneous data and multi-dimensional reward signals while ensuring optimization stability. For the evaluation challenge, we construct MCDR-Bench based on the ``error uniqueness principle," transforming subjective generation assessment into objective error identification through controllable error injection, enabling quantifiable evaluation of deep research capabilities. Experimental validation confirms that the proposed PRPO and MCDR-Bench jointly establish a unified framework that systematically advances chart deep research through enhanced collaborative training and objective evaluation.

cs.CV

WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs

Recent advances in Multimodal Large Language Models have greatly improved visual understanding and reasoning, yet their quadratic attention and offline training protocols make them ill-suited for streaming settings where frames arrive sequentially and future observations are inaccessible. We diagnose a core limitation of current Video-LLMs, namely Time-Agnosticism, in which videos are treated as an unordered bag of evidence rather than a causally ordered sequence, yielding two failures in streams: temporal order ambiguity, in which the model cannot follow or reason over the correct chronological order, and past-current focus blindness where it fails to distinguish present observations from accumulated history. We present WeaveTime, a simple, efficient, and model agnostic framework that first teaches order and then uses order. We introduce a lightweight Temporal Reconstruction objective-our Streaming Order Perception enhancement-that instills order aware representations with minimal finetuning and no specialized streaming data. At inference, a Past-Current Dynamic Focus Cache performs uncertainty triggered, coarse-to-fine retrieval, expanding history only when needed. Plugged into exsiting Video-LLM without architectural changes, WeaveTime delivers consistent gains on representative streaming benchmarks, improving accuracy while reducing latency. These results establish WeaveTime as a practical path toward time aware stream Video-LLMs under strict online, time causal constraints. Code and weights will be made publicly available. Project Page: https://zhangyl4.github.io/publications/weavetime/

cs.CV

Vision Transformers Need More Than Registers

Vision Transformers (ViTs), when pre-trained on large-scale data, provide general-purpose representations for diverse downstream tasks. However, artifacts in ViTs are widely observed across different supervision paradigms and downstream tasks. Through systematic analysis of artifacts in ViTs, we find that their fundamental mechanisms have yet to be sufficiently elucidated. In this paper, through systematic analysis, we conclude that these artifacts originate from a lazy aggregation behavior: ViT uses semantically irrelevant background patches as shortcuts to represent global semantics, driven by global attention and Coarse-grained semantic supervision. Our solution selectively integrates patch features into the CLS token, reducing the influence of background-dominated shortcuts and consistently improving performance across 12 benchmarks under label-, text-, and self-supervision. We hope this work offers a new perspective on ViT behavior.

cs.CV

RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation

In this paper, we propose a 3D asset-referenced diffusion model for image generation, exploring how to integrate 3D assets into image diffusion models. Existing reference-based image generation methods leverage large-scale pretrained diffusion models and demonstrate strong capability in generating diverse images conditioned on a single reference image. However, these methods are limited to single-image references and cannot leverage 3D assets, constraining their practical versatility. To address this gap, we present a cross-domain diffusion model with dual-branch perception that leverages multi-view RGB images and point maps of 3D assets to jointly model their colors and canonical-space coordinates, achieving precise consistency between generated images and the 3D references. Our spatially aligned dual-branch generation architecture and domain-decoupled generation mechanism ensure the simultaneous generation of two spatially aligned but content-disentangled outputs, RGB images and point maps, linking 2D image attributes with 3D asset attributes. Experiments show that our approach effectively uses 3D assets as references to produce images consistent with the given assets, opening new possibilities for combining diffusion models with 3D content creation.

cs.CV

Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats

Despite their impressive performance across a wide range of tasks, Large Vision-Language Models (LVLMs) remain prone to hallucination. In this study, we propose a comprehensive intervention framework aligned with the transformer's causal architecture in LVLMs, integrating the effects of different intervention paths on hallucination. We find that hallucinations in LVLMs do not arise from a single causal path, but rather from the interplay among image-to-input-text, image-to-output-text, and text-to-text pathways. For the first time, we also find that LVLMs rely on different pathways depending on the question-answer alignment format. Building on these insights, we propose simple yet effective methods to identify and intervene on critical hallucination heads within each pathway, tailored to discriminative and generative formats. Experiments across multiple benchmarks demonstrate that our approach consistently reduces hallucinations across diverse alignment types.

cs.CV

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

Recent advancements in language-grounded autonomous driving have been significantly promoted by the sophisticated cognition and reasoning capabilities of large language models (LLMs). However, current LLM-based approaches encounter critical challenges: (1) Failure analysis reveals that frequent collisions and obstructions, stemming from limitations in visual representations, remain primary obstacles to robust driving performance. (2) The substantial parameters of LLMs pose considerable deployment hurdles. To address these limitations, we introduce VLDrive, a novel approach featuring a lightweight MLLM architecture with enhanced vision components. VLDrive achieves compact visual tokens through innovative strategies, including cycle-consistent dynamic visual pruning and memory-enhanced feature aggregation. Furthermore, we propose a distance-decoupled instruction attention mechanism to improve joint visual-linguistic feature learning, particularly for long-range visual tokens. Extensive experiments conducted in the CARLA simulator demonstrate VLDrive`s effectiveness. Notably, VLDrive achieves state-of-the-art driving performance while reducing parameters by 81% (from 7B to 1.3B), yielding substantial driving score improvements of 15.4%, 16.8%, and 7.6% at tiny, short, and long distances, respectively, in closed-loop evaluations. Code is available at https://github.com/ReaFly/VLDrive.

cs.CV

Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context

Large Vision-Language Models (LVLMs) have made significant progress in recent years but are also prone to hallucination issues. They exhibit more hallucinations in longer, free-form responses, often attributed to accumulated uncertainties. In this paper, we ask: Does increased hallucination result solely from length-induced errors, or is there a deeper underlying mechanism? After a series of preliminary experiments and findings, we suggest that the risk of hallucinations is not caused by length itself but by the increased reliance on context for coherence and completeness in longer responses. Building on these insights, we propose a novel "induce-detect-suppress" framework that actively induces hallucinations through deliberately designed contexts, leverages induced instances for early detection of high-risk cases, and ultimately suppresses potential object-level hallucinations during actual decoding. Our approach achieves consistent, significant improvements across all benchmarks, demonstrating its efficacy. The strong detection and improved hallucination mitigation not only validate our framework but, more importantly, re-validate our hypothesis on context. Rather than solely pursuing performance gains, this study aims to provide new insights and serves as a first step toward a deeper exploration of hallucinations in LVLMs' longer responses.

cs.CV

Augmenting Moment Retrieval: Zero-Dependency Two-Stage Learning

Existing Moment Retrieval methods face three critical bottlenecks: (1) data scarcity forces models into shallow keyword-feature associations; (2) boundary ambiguity in transition regions between adjacent events; (3) insufficient discrimination of fine-grained semantics (e.g., distinguishing ``kicking" vs. ``throwing" a ball). In this paper, we propose a zero-external-dependency Augmented Moment Retrieval framework, AMR, designed to overcome local optima caused by insufficient data annotations and the lack of robust boundary and semantic discrimination capabilities. AMR is built upon two key insights: (1) it resolves ambiguous boundary information and semantic confusion in existing annotations without additional data (avoiding costly manual labeling), and (2) it preserves boundary and semantic discriminative capabilities enhanced by training while generalizing to real-world scenarios, significantly improving performance. Furthermore, we propose a two-stage training framework with cold-start and distillation adaptation. The cold-start stage employs curriculum learning on augmented data to build foundational boundary/semantic awareness. The distillation stage introduces dual query sets: Original Queries maintain DETR-based localization using frozen Base Queries from the cold-start model, while Active Queries dynamically adapt to real-data distributions. A cross-stage distillation loss enforces consistency between Original and Base Queries, preventing knowledge forgetting while enabling real-world generalization. Experiments on multiple benchmarks show that AMR achieves improved performance over prior state-of-the-art approaches.

cs.CV

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on egocentric images, using exocentric interaction images with image-level annotations. However, extracting affordance knowledge solely from exocentric images and transferring it one-way to egocentric images limits the applicability of previous works in complex interaction scenarios. Instead, this study introduces LoopTrans, a novel closed-loop framework that not only transfers knowledge from exocentric to egocentric but also transfers back to enhance exocentric knowledge extraction. Within LoopTrans, several innovative mechanisms are introduced, including unified cross-modal localization and denoising knowledge distillation, to bridge domain gaps between object-centered egocentric and interaction-centered exocentric images while enhancing knowledge transfer. Experiments show that LoopTrans achieves consistent improvements across all metrics on image and video benchmarks, even handling challenging scenarios where object interaction regions are fully occluded by the human body.

cs.CV