SearcharxivSearch

arXiv subjects

Qian Yu

Publications and source records attributed to Qian Yu.

At least 19 recordsLinked to original sources

A Complete Characterization of Tensorizable $f$-divergences

Csiszar's formulation of the $f$-divergence introduced a vast family of functionals for quantifying dissimilarity between probability distributions. However, many applications in statistics and information theory rely only on a few $f$-divergences, such as the Kullback-Leibler divergence, the $\chi^2$-divergence, and the squared Hellinger distance. These divergences are especially useful because they admit simple compositional formulas under product measures, a property sometimes referred to as tensorization. In this work, we refine a formalism of tensorization previously introduced in the literature. Then, we show that any possible tensorization formula has a multi-affine form characterized by a single parameter, and identify all tensorizable $f$-divergences under our adopted notion of tensorization.

cs.IT

Towards Physics-Faithful Generation of Scientific Diagrams

Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance but on physical faithfulness: correct force directions, valid coordinate systems, consistent thermodynamic states, and equations matching the depicted scenario. Trained on web imagery with physically shallow captions, generic models produce diagrams that look plausible but are physically wrong, harmful in education and scientific communication. We present Princigram, a physics-faithful scientific-diagram generator, and its data pipeline. Our central advance is Structured Physical Chain-of-Thought (SP-CoT): a per-subdiscipline schema that decomposes a physics diagram into an explicit multi-step reasoning chain across six subdisciplines, from scene identification through force or process analysis to governing laws and synthesis. Unlike free-form chain-of-thought, SP-CoT follows a fixed schema with strict fidelity rules that separate visually grounded facts from physically inferred reasoning and type all mathematics symbolically; it serves both as dense training supervision and, at inference, as a structured "thinking" prompt. With it we curate and structurally annotate 4.3 million physics images, of which 115,037 carry expert-level annotation, and adapt a unified multimodal backbone. We further introduce VeriphyT2IBench, whose questions are derived from each held-out diagram's own structured annotation: each diagram becomes an item-specific bank of binary questions about its objects, forces, and states, so a judge model's score decomposes into named physical facts rather than one holistic number. On the physics subset of GenExam and on VeriphyT2IBench, Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams.

cs.CV

Uniform-in-Time Smoluchowski-Kramers Approximation in Total Variation for Fractional SDEs

Let $B^H$ be a one-dimensional fractional Brownian motion with Hurst index $H\in(1/2,1)$. We study the small-mass limit of the kinetic equation \[ dX_t^\mu=Y_t^\mu\,dt,\qquad \mu\,dY_t^\mu=b(X_t^\mu)\,dt-Y_t^\mu\,dt +\sigma(X_t^\mu)\,dB_t^H, \] where the noise coefficient is state dependent and uniformly nondegenerate. The limiting equation is the Young differential equation \[ dX_t=b(X_t)\,dt+\sigma(X_t)\,dB_t^H. \] Under uniform ellipticity and strict dissipativity of the transformed drift, we prove that, for every $t_0>0$ and every $\rho<2H-1$, \[ \sup_{t\ge t_0}d_{TV}\big(X_t^\mu,X_t\big) \le C_{t_0,\rho}\mu^\rho . \] The estimate is uniform on the entire half-line. We also construct stationary solutions on a two-sided fractional-noise space and obtain convergence of their one-time marginals in total variation. The exponent $2H-1$ is identified as the natural endpoint: it is generated by the quadratic velocity term after the Lamperti transform, and it degenerates as $H\downarrow1/2$. A nonconstant uniformly elliptic example is included.

math.PR

EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation

Controllable local editing of 3D assets requires precise target localization and appropriate visual guidance. However, existing methods lack a simple yet accurate way to obtain 3D masks and struggle to achieve the desired edit while faithfully preserving the structure and appearance of non-target regions. To address these challenges, we present EditFlow3D, a training-free framework for local 3D editing. Given a source asset and an edit instruction, a VLM-driven workflow interprets the editing intent and automatically constructs a visual guidance image and a refined 3D editing mask, enabling localized editing in the native representation space of a pretrained 3D generative model. Specifically, mask-guided differential flow focuses the edit on the target region, while step-wise trajectory preservation maintains consistency between non-target regions and the source asset without directly replacing intermediate features. Since the existing Edit3D-Bench covers only a limited range of local editing categories, we further introduce EditFlow-Bench as a complementary benchmark encompassing a broader variety of structural and appearance edits, and evaluate EditFlow3D on both benchmarks. Quantitative results, qualitative comparisons, and a user study demonstrate that EditFlow3D achieves more accurate target-region editing and better preserves non-target regions than existing 3D editing methods.

cs.CV

Limit Theorems for Tempered Linear Processes with Innovations in the Domain of Attraction of a Stable Law

We study the partial-sum behavior of tempered linear processes \[ X_{N,n}=\sum_{j=1}^{\infty}e^{-\lambda_Nj}\frac{\ell(j)}{j}\varepsilon_{n-j}, \qquad \lambda_N\downarrow0, \] where $\ell$ is slowly varying and the innovations belong to the domain of attraction of an $\alpha$-stable law with $1<\alpha\leq2$. The filter $j^{-1}\ell(j)$ represents the logarithmic boundary between summable and power-law long-memory coefficients. Let \[ Q_N=\sum_{j=1}^{N}e^{-\lambda_Nj}\frac{\ell(j)}{j}, \qquad L(N)=\sum_{j=1}^{N}\frac{\ell(j)}{j}. \] We prove that the partial-sum process, normalized by $B_NQ_N$, converges to the stable L\'{e}vy motion associated with the innovations. Moreover, \[ Q_N\sim L(N)\quad\text{if }N\lambda_N=O(1), \qquad Q_N\sim L(1/\lambda_N)\quad\text{if }N\lambda_N\to\infty. \] Then weak, moderate, and strong tempering have the same first-order L\'{e}vy limit but different normalizations. In the weakly and moderately tempered regimes, the second-order remainder, normalized by $B_N\ell(N)$, converges in finite-dimensional distributions to a logarithmically tempered stable process; in the Gaussian finite-moment case the convergence is functional. These results extend the untempered logarithmic-boundary theorem and complement existing invariance principles for tempered linear processes.

math.PR

HI Depletion Begins Well Beyond the Virial Radius: A FAST Stacking Study of 36 Galaxy Clusters to 5R200

We present a stacking study of the neutral atomic hydrogen (HI) content in and around 36 local galaxy clusters at $z<0.07$, using a combination of the FAST all sky HI survey (FASHI) and the extensive spectroscopic catalog mainly from the Dark Energy Spectroscopic Instrument (DESI). We employ spectral stacking techniques to probe the average HI mass and HI-to-stellar mass ratio ($M_{\rm HI}/M_*$) for member galaxies down to stellar masses of $M_*\sim 10^9M_\odot$, spanning a projected cluster-centric distance to $5R_{200}$. Our analysis reveals a pronounced environmental effect: both $M_{\rm HI}$ and $M_{\rm HI}/M_*$ decrease steadily towards the cluster center, dropping by $\sim 0.5$ dex on average from the outskirts to the core. Crucially, we find that $M_{\rm HI}/M_*$ of galaxies remain lower than the field galaxies even at the $5R_{200}$. This provides direct, statistical evidence for substantial gas stripping and pre-processing in the cluster outskirts, likely occurring in infalling groups and large-scale filaments. By further splitting the sample by $g-r$ color, we show that the HI deficiency persists at fixed galaxy color: even the bluest cluster members exhibit $\sim 0.5$~dex lower $M_{\rm HI}/M_*$ than field galaxies of similar color, reflecting environmental effects on the cold gas reservoir prior to full optical transformation. The total HI mass within clusters and their outskirts agrees broadly with predictions from cosmological simulation. Our results underscore the critical role of the extended cluster environment in quenching galaxies by depleting their cold gas reservoirs well before they enter the dense cluster core.

astro-ph.GA

VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation

Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence. However, this task remains challenging as it requires bridging discrete code representations with continuous visual dynamics. Existing optimization-based methods often destroy topological consistency, while general-purpose LLMs rely on rigid CSS/SMIL transformations, failing to model geometry-level non-rigid deformations. To address these limitations, we present VAnim, the first LLM-based framework for open-domain text-to-SVG animation. We reconceptualize animation not as sequence generation, but as Sparse State Updates (SSU) on a persistent SVG DOM tree. This paradigm compresses sequence length by over 9.8x while preserving the SVG DOM structure and non-participating elements by construction. To enable precise control, we propose an Identification-First Motion Planning mechanism that grounds textual instructions in explicit visual entities. Furthermore, to overcome the non-differentiable nature of SVG rendering, we employ Rendering-Aware Reinforcement Learning via Group Relative Policy Optimization (GRPO). By leveraging a hybrid reward from a state-of-the-art video perception encoder, we align discrete code updates with high-fidelity visual feedback. We also introduce SVGAnim-134k, the first benchmark for vector animation. Extensive experiments demonstrate that VAnim significantly outperforms state-of-the-art baselines in semantic alignment and structural validity, with additional appendix metrics further validating motion quality and identity preservation.

cs.CV

AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to explore existing knowledge for a research problem, or to acquire evidence for verifying assumptions and supporting claims. To assess AI agents' capability in driving this process, we present AutoResearchBench, a dedicated benchmark for autonomous scientific literature discovery. AutoResearchBench consists of two complementary task types: (1) Deep Research, which requires tracking down a specific target paper through a progressive, multi-step probing process, and (2) Wide Research, which requires comprehensively collecting a set of papers satisfying given conditions. Compared to previous benchmarks on agentic web browsing, AutoResearchBench is distinguished along three dimensions: it is research-oriented, calling for in-depth comprehension of scientific concepts; literature-focused, demanding fine-grained utilization of detailed information; and open-ended, involving an unknown number of qualified papers and thus requiring deliberate reasoning and search throughout. These properties make AutoResearchBench uniquely suited for evaluating autonomous research capabilities, and extraordinarily challenging. Even the most powerful LLMs, despite having largely conquered general agentic web-browsing benchmarks such as BrowseComp, achieve only 9.39% accuracy on Deep Research and 9.31% IoU on Wide Research, while many other strong baselines fall below 5%. We publicly release the dataset and evaluation pipeline to facilitate future research in this direction. We publicly release the dataset, evaluation pipeline, and code at https://github.com/CherYou/AutoResearchBench.

cs.AI

Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. However, existing paradigms typically adopt an open-loop "blind drawing" approach, where models generate symbolic code sequences without perceiving intermediate visual outcomes. This methodology severely underutilizes the powerful visual priors embedded in MLLMs vision encoders, treating SVG generation as a disjointed textual sequence modeling task rather than an integrated visuo-spatial one. Consequently, models struggle to reason about partial canvas states and implicit occlusion relationships, which are visually explicit but textually ambiguous. To bridge this gap, we propose Render-in-the-Loop, a novel generation paradigm that reformulates SVG synthesis as a step-wise, visual-context-aware process. By rendering intermediate code states into a cumulative canvas, the model explicitly observes the evolving visual context at each step, leveraging on-the-fly feedback to guide subsequent generation. However, we demonstrate that applying this visual loop naively to off-the-shelf models is suboptimal due to their inability to leverage incremental visual-code mappings. To address this, we first utilize fine-grained path decomposition to construct dense multi-step visual trajectories, and then introduce a Visual Self-Feedback (VSF) training strategy to condition the next primitive generation on intermediate visual states. Furthermore, a Render-and-Verify (RaV) inference mechanism is proposed to effectively filter degenerate and redundant primitives. Our framework, instantiated on a multimodal foundation model, outperforms strong open-weight baselines on the standard MMSVGBench. This result highlights the remarkable data efficiency and generalization capability of our Render-in-the-Loop paradigm for both Text-to-SVG and Image-to-SVG tasks.

cs.CV

AmodalSVG: Amodal Image Vectorization via Semantic Layer Peeling

We introduce AmodalSVG, a new framework for amodal image vectorization that produces semantically organized and geometrically complete SVG representations from natural images. Existing vectorization methods operate under a modal paradigm: tracing only visible pixels and disregarding occlusion. Consequently, the resulting SVGs are semantically entangled and geometrically incomplete, limiting SVG's structural editability. In contrast, AmodalSVG reconstructs full object geometries, including occluded regions, into independent, editable vector layers. To achieve this, AmodalSVG reformulates image vectorization as a two-stage framework, performing semantic decoupling and completion in the raster domain to produce amodally complete semantic layers, which are then independently vectorized. In the first stage, we introduce Semantic Layer Peeling (SLP), a VLM-guided strategy that progressively decomposes an image into semantically coherent layers. By hybrid inpainting, SLP recovers complete object appearances under occlusions, enabling explicit semantic decoupling. To vectorize these layers efficiently, we propose Adaptive Layered Vectorization (ALV), which dynamically modulates the primitive budget via an error-budget-driven adjustment mechanism. Extensive experiments demonstrate that AmodalSVG significantly outperforms prior methods in visual fidelity. Moreover, the resulting amodal layers enable object-level editing directly in the vector domain, capabilities not supported by existing vectorization approaches. Code will be released upon acceptance.

cs.CV

ArtiCAD: Articulated CAD Assembly Design via Multi-Agent Code Generation

Parametric Computer-Aided Design (CAD) of articulated assemblies is essential for product development, yet generating these multi-part, movable models from high-level descriptions remains unexplored. To address this, we propose ArtiCAD, the first training-free multi-agent system capable of generating editable, articulated CAD assemblies directly from text or images. Our system divides this complex task among four specialized agents: Design, Generation, Assembly, and Review. One of our key insights is to predict assembly relationships during the initial design stage rather than the assembly stage. By utilizing a Connector that explicitly defines attachment points and joint parameters, ArtiCAD determines these relationships before geometry generation, effectively bypassing the limited spatial reasoning capabilities of current LLMs and VLMs. To further ensure high-quality outputs, we introduce validation steps in the generation and assembly stages, accompanied by a cross-stage rollback mechanism that accurately isolates and corrects design- and code-level errors. Additionally, a self-evolving experience store accumulates design knowledge to continuously improve performance on future tasks. Extensive evaluations on three datasets (ArtiCAD-Bench, CADPrompt, and ACD) validate the effectiveness of our approach. We further demonstrate the applicability of ArtiCAD in requirement-driven conceptual design, physical prototyping, and the generation of embodied AI training assets through URDF export.

cs.CV

Hierarchical SVG Tokenization: Learning Compact Visual Programs for Scalable Vector Graphics Modeling

Recent large language models have shifted SVG generation from differentiable rendering optimization to autoregressive program synthesis. However, existing approaches still rely on generic byte-level tokenization inherited from natural language processing, which poorly reflects the geometric structure of vector graphics. Numerical coordinates are fragmented into discrete symbols, destroying spatial relationships and introducing severe token redundancy, often leading to coordinate hallucination and inefficient long-sequence generation. To address these challenges, we propose HiVG, a hierarchical SVG tokenization framework tailored for autoregressive vector graphics generation. HiVG decomposes raw SVG strings into structured \textit{atomic tokens} and further compresses executable command--parameter groups into geometry-constrained \textit{segment tokens}, substantially improving sequence efficiency while preserving syntactic validity. To further mitigate spatial mismatch, we introduce a Hierarchical Mean--Noise (HMN) initialization strategy that injects numerical ordering signals and semantic priors into new token embeddings. Combined with a curriculum training paradigm that progressively increases program complexity, HiVG enables more stable learning of executable SVG programs. Extensive experiments on both text-to-SVG and image-to-SVG tasks demonstrate improved generation fidelity, spatial consistency, and sequence efficiency compared with conventional tokenization schemes. Our code is publicly available at https://github.com/ximinng/HiVG

cs.LG

Molecular Identifier Visual Prompt and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing

Reaction diagram parsing (RxnDP) is critical for extracting chemical synthesis information from literature. Although recent Vision-Language Models (VLMs) have emerged as a promising paradigm to automate this complex visual reasoning task, their application is fundamentally bottlenecked by the inability to align visual chemical entities with pre-trained knowledge, alongside the inherent discrepancy between token-level training and reaction-level evaluation. To address these dual challenges, this work enhances VLM-based RxnDP from two complementary perspectives: prompting representation and learning paradigms. First, we propose Identifier as Visual Prompting (IdtVP), which leverages naturally occurring molecule identifiers (e.g., bold numerals like 1a) to activate the chemical knowledge acquired during VLM pre-training. IdtVP enables powerful zero-shot and out-of-distribution capabilities, outperforming existing prompting strategies. Second, to further optimize performance within fine-tuning paradigms, we introduce Re3-DAPO, a reinforcement learning algorithm that leverages verifiable rewards to directly optimize reaction-level metrics, thereby achieving consistent gains over standard supervised fine-tuning. Additionally, we release the ScannedRxn benchmark, comprising scanned historical reaction diagrams with real-world artifacts, to rigorously assess model robustness and out-of-distribution ability. Our contributions advance the accuracy and generalization of VLM-based reaction diagram parsing. We will release data, models, and code on GitHub.

cs.CV

Small ball probability of collision local time for symmetric stable processes

In this article, the small ball probability is obtained for the collision local time of two independent symmetric $\alpha-$stable processes with parameters $\alpha_1,\alpha_2\in(0,2]$ satisfying $\max\{\alpha_1,\alpha_2\}>1$. The proof is based on obtaining the asymptotic behavior of moment generating function by contour integration.

math.PR

Strong solutions to SDEs with singular drifts driven by fractional Brownian motions

In this paper, we establish the strong well-posedness of SDEs with merely integrable time-dependent drifts driven by fractional Brownian motions with Hurst parameter H<1/2. Our result holds over the entire subcritical regime and can be regarded as an extension of (Krylov and Rockner, Probab. Theory Relat. Fields, 131(2): 154-196 (2005)) to the fractional case. Furthermore, we prove the existence of stochastic flows of Sobolev diffeomorphisms for this class of SDEs, which generalizes a result in (Mohammed et al., Ann. Probab. 43, 1535-1576 (2015)). The approach adopted in our work is based on a compactness criterion for random fields in Wiener spaces.

math.PR

Prune4Web: DOM Tree Pruning Programming for Web Agent

Web automation employs intelligent agents to execute high-level tasks by mimicking human interactions with web interfaces. Despite the capabilities of recent Large Language Model (LLM)-based web agents, navigating complex, real-world webpages efficiently remains a significant hurdle due to the prohibitively large size of Document Object Model (DOM) structures, often ranging from 10,000 to 100,000 tokens. Existing strategies typically rely on crude DOM truncation -- risking the loss of critical information -- or employ inefficient heuristics and separate ranking models, failing to achieve an optimal balance between precision and scalability. To address these challenges, we introduce Prune4Web, a novel paradigm that shifts DOM processing from resource-intensive LLM reading to efficient programmatic pruning. Central to our approach is DOM Tree Pruning Programming, where an LLM generates executable Python scoring scripts to dynamically filter DOM elements based on semantic cues from decomposed sub-tasks. This mechanism eliminates the need for LLMs to ingest raw, massive DOMs, instead delegating traversal and scoring to lightweight, interpretable programs. This methodology achieves a 25x to 50x reduction in candidate elements for grounding, thereby facilitating precise action localization while mitigating attention dilution. Furthermore, we propose a specialized data annotation pipeline and a two-turn dialogue training strategy that jointly optimizes the Planner, Programmatic Filter, and Grounder within a unified framework. Extensive experiments demonstrate state-of-the-art performance. Notably, on our low-level grounding task, Prune4Web dramatically improves accuracy from 46.8% to 88.28%, underscoring its efficacy in real-world web automation.

cs.AI

Asymptotic Properties of the Derivative of Self-Intersection Local Time of Multidimensional Fractional Brownian Motion

Let \{B_t^H,t\geq0\} be a d-dimensional fractional Brownian motion. We prove that the approximation of the first-order derivative of self-intersection local time, defined as \alpha_{\varepsilon,t}^{(1)}(0)=-\int_0^t\int_0^sp_\varepsilon^{(1)}(B_s^H-B_r^H)\d r\d s, where p_\varepsilon^{(1)}(x_1,\cdots,x_d):=\partial _{x_1}p(x_1,\cdots,x_d) and p_\varepsilon(x)=(2\pi\varepsilon)^{-d/2}e^{|x|^2/2\varepsilon},x\in\mathbb{R}^d, d\geq2 is the heat kernel, exits in L^2 sense if and only if H<\frac{3}{2(1+d)} and satisfies three different central limit theorems when normalized by \varepsilon^{\frac d2+1-\frac1H} for H>\frac12 and d\geq2, normalized by \varepsilon^{\frac d2+\frac12-\frac 3{4H}} for \frac{3}{2(1+d)}<H<\frac12 and d\geq3, and normalized by \log(1/\varepsilon)^{-\frac12} for the critical case H=\frac{3}{2(1+d)} and d\geq3.

math.PR

RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning

Large-scale chemical reaction datasets are crucial for AI research in chemistry. However, existing chemical reaction data often exist as images within papers, making them not machine-readable and unusable for training machine learning models. In response to this challenge, we propose the RxnCaption framework for the task of chemical Reaction Diagram Parsing (RxnDP). Our framework reformulates the traditional coordinate prediction driven parsing process into an image captioning problem, which Large Vision Language Models (LVLMs) handle naturally. We introduce a strategy termed BBox and Index as Visual Prompt (BIVP), which uses our state-of-the-art molecular detector, MolYOLO, to pre-draw molecular bounding boxes and indices directly onto the input image. This turns the downstream parsing into a natural-language description problem. Extensive experiments show that the BIVP strategy significantly improves structural extraction quality while simplifying model design. We further construct the RxnCaption-15k dataset, an order of magnitude larger than prior real-world literature benchmarks, with a balanced test subset across four layout archetypes. Experiments demonstrate that RxnCaption-VL achieves state-of-the-art performance on multiple metrics. We believe our method, dataset, and models will advance structured information extraction from chemical literature and catalyze broader AI applications in chemistry. We will release data, models, and code on GitHub.

cs.CV