SearcharxivSearch

arXiv subjects

Xuhui Li

Publications and source records attributed to Xuhui Li.

15 recordsLinked to original sources

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires precise objects localization, or can bypass explicit localization through global layout cues. In this work, we investigate two representative model families, LLaVA-1.5 and Qwen2.5-VL, using a suite of mechanistic interpretability tools, including token ablation, layer-wise probing, attention knockout, and causal mediation analysis. We find that spatial relation prediction follows a staged grounding-to-reasoning process in which object-aligned tokens establish coarse target-reference anchors, while precise bounding-box boundaries are not required. Positional information becomes decodable before relation decisions emerge, and a small set of attention heads mediates the causal effects of both localization and spatial reasoning. The two tasks share early grounding-related processing but ultimately rely on partially distinct specialized pathways. Through rigorous experiments, we provide a token-, layer-, and head-level account of how VLMs transform object grounding into spatial relations, showing that knowing where objects are is not equivalent to knowing how they relate.

cs.CV

ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models

Medical Vision-Language Models have shown promising potential in clinical decision support, yet they remain prone to factual hallucinations due to insufficient grounding in localized pathological evidence. Existing medical alignment methods primarily operate at the response level through preference optimization, improving output correctness but leaving intermediate reasoning weakly connected to visual regions. Although chain-of-thought (CoT) enhances multimodal reasoning, it remains largely text-centric, limiting effective integration of clinical visual cues. To address this gap, we propose ClinCoT, a clinical-aware visual chain-of-thought framework that transforms preference optimization from response-level correction to visual-driven reasoning. We introduce an automatic data generation pipeline that constructs clinically grounded preference pairs through reasoning with hypotheses-driven region proposals. Multiple Med-LLMs evaluators rank and assign scores to each response, and these rankings serve as supervision to train the target model. We further introduce a scoring-based margin-aware optimization strategy that incorporates both preference ranking and score difference to refine region-level reasoning trajectories. To maintain alignment as the model's policy evolves during training, we adopt an iterative learning scheme that dynamically regenerates preference data. Extensive experiments on three medical VQA and report generation benchmarks demonstrate that ClinCoT consistently improves factual grounding and achieves superior performance compared with existing preference-based alignment methods.

cs.CV

Path-Guided Flow Matching for Dataset Distillation

Dataset distillation compresses large datasets into compact synthetic sets with comparable performance in training models. Despite recent progress on diffusion-based distillation, this type of method typically depends on heuristic guidance or prototype assignment, which comes with time-consuming sampling and trajectory instability and thus hurts downstream generalization especially under strong control or low IPC. We propose \emph{Path-Guided Flow Matching (PGFM)}, the first flow matching-based framework for generative distillation, which enables fast deterministic synthesis by solving an ODE in a few steps. PGFM conducts flow matching in the latent space of a frozen VAE to learn class-conditional transport from Gaussian noise to data distribution. Particularly, we develop a continuous path-to-prototype guidance algorithm for ODE-consistent path control, which allows trajectories to reliably land on assigned prototypes while preserving diversity and efficiency. Extensive experiments across high-resolution benchmarks demonstrate that PGFM matches or surpasses prior diffusion-based distillation approaches with fewer steps of sampling while delivering competitive performance with remarkably improved efficiency, e.g., 7.6$\times$ more efficient than the diffusion-based counterparts with 78\% mode coverage.

cs.LG

User-Feedback-Driven Adaptation for Vision-and-Language Navigation

Real-world deployment of Vision-and-Language Navigation (VLN) agents is constrained by the scarcity of reliable supervision after offline training. While recent adaptation methods attempt to mitigate distribution shifts via environment-driven self-supervision (e.g., entropy minimization), these signals are often noisy and can cause the agent to amplify its own mistakes during long-horizon sequential decision-making. In this paper, we propose a paradigm shift that positions user feedback, specifically episode-level success confirmations and goal-level corrections, as a primary and general-purpose supervision signal for VLN. Unlike internal confidence scores, user feedback is intent-aligned and in-situ consistent, directly correcting the agent's decoupling from user instructions. To effectively leverage this supervision, we introduce a user-feedback-driven learning framework featuring a topology-aware trajectory construction pipeline. This mechanism lifts sparse, goal-level corrections into dense path-level supervision by generating feasible paths on the agent's incrementally built topological graph, enabling sample-efficient imitation learning without requiring step-by-step human demonstrations. Furthermore, we develop a persistent memory bank mechanism for warm-start initialization, supporting the reuse of previously acquired topology and cached representations across navigation sessions. Extensive experiments on the GSA-R2R benchmark demonstrate that our approach transforms sparse interaction into robust supervision, consistently outperforming environment-driven baselines while exhibiting strong adaptability across diverse instruction styles.

cs.AI

GeoDM: Geometry-aware Distribution Matching for Dataset Distillation

Dataset distillation aims to synthesize a compact subset of the original data, enabling models trained on it to achieve performance comparable to those trained on the original large dataset. Existing distribution-matching methods are confined to Euclidean spaces, making them only capture linear structures and overlook the intrinsic geometry of real data, e.g., curvature. However, high-dimensional data often lie on low-dimensional manifolds, suggesting that dataset distillation should have the distilled data manifold aligned with the original data manifold. In this work, we propose a geometry-aware distribution-matching framework, called \textbf{GeoDM}, which operates in the Cartesian product of Euclidean, hyperbolic, and spherical manifolds, with flat, hierarchical, and cyclical structures all captured by a unified representation. To adapt to the underlying data geometry, we introduce learnable curvature and weight parameters for three kinds of geometries. At the same time, we design an optimal transport loss to enhance the distribution fidelity. Our theoretical analysis shows that the geometry-aware distribution matching in a product space yields a smaller generalization error bound than the Euclidean counterparts. Extensive experiments conducted on standard benchmarks demonstrate that our algorithm outperforms state-of-the-art data distillation methods and remains effective across various distribution-matching strategies for the single geometries.

cs.CV

When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection

Large language models (LLMs) have grown more powerful in language generation, producing fluent text and even imitating personal style. Yet, this ability also heightens the risk of identity impersonation. To the best of our knowledge, no prior work has examined personalized machine-generated text (MGT) detection. In this paper, we introduce \dataset, the first benchmark for evaluating detector robustness in personalized settings, built from literary and blog texts paired with their LLM-generated imitations. Our experimental results demonstrate large performance gaps across detectors in personalized settings: some state-of-the-art models suffer significant drops. We attribute this limitation to the \textit{feature-inversion trap}, where features that are discriminative in general domains become inverted and misleading when applied to personalized text. Based on this finding, we propose \method, a simple and reliable way to predict detector performance changes in personalized settings. \method identifies latent directions corresponding to inverted features and constructs probe datasets that differ primarily along these features to evaluate detector dependence. Our experiments show that \method can accurately predict both the direction and the magnitude of post-transfer changes, showing 85\% correlation with the actual performance gaps. We hope that this work will encourage further research on personalized text detection.

cs.CL

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with post-hoc refinement techniques such as flow matching, these methods fail to recover fine-grained details (e.g., prosodic nuances, speaker-specific timbres), especially in challenging scenarios like singing voice or music synthesis. We propose QTTS, a novel TTS framework built upon our new audio codec, QDAC. The core innovation of QDAC lies in its end-to-end training of an ASR-based auto-regressive network with a GAN, which achieves superior semantic feature disentanglement for scalable, near-lossless compression. QTTS models these discrete codes using two innovative strategies: the Hierarchical Parallel architecture, which uses a dual-AR structure to model inter-codebook dependencies for higher-quality synthesis, and the Delay Multihead approach, which employs parallelized prediction with a fixed delay to accelerate inference speed. Our experiments demonstrate that the proposed framework achieves higher synthesis quality and better preserves expressive content compared to baseline. This suggests that scaling up compression via multi-codebook modeling is a promising direction for high-fidelity, general-purpose speech and audio generation.

cs.SD

Second-order force scheme for lattice Boltzmann method

We present an a priori derivation of the force scheme for lattice Boltzmann method based on kinetic theoretical formulation. We show that the discrete lattice effect, previously eliminated a posteriori in BGK collision model, is due to first-order space-time discretization and can be eliminated generically for a wide range of collision models with second-order space-time discretization. Particularly, the force scheme for the recently developed spectral multiple-relaxation-time (SMRT) collision model is obtained and numerically verified.

physics.comp-ph

Cascade Image Matting with Deformable Graph Refinement

Image matting refers to the estimation of the opacity of foreground objects. It requires correct contours and fine details of foreground objects for the matting results. To better accomplish human image matting tasks, we propose the Cascade Image Matting Network with Deformable Graph Refinement, which can automatically predict precise alpha mattes from single human images without any additional inputs. We adopt a network cascade architecture to perform matting from low-to-high resolution, which corresponds to coarse-to-fine optimization. We also introduce the Deformable Graph Refinement (DGR) module based on graph neural networks (GNNs) to overcome the limitations of convolutional neural networks (CNNs). The DGR module can effectively capture long-range relations and obtain more global and local information to help produce finer alpha mattes. We also reduce the computation complexity of the DGR module by dynamically predicting the neighbors and apply DGR module to higher--resolution features. Experimental results demonstrate the ability of our CasDGR to achieve state-of-the-art performance on synthetic datasets and produce good results on real human images.

cs.CV

A multiple-relaxation-time collision model by Hermite expansion

The Bhatnagar-Gross-Krook (BGK) single-relaxation-time collision model for the Boltzmann equation serves as the foundation of the lattice BGK (LBGK) method developed in recent years. The description of the collision as a uniform relaxation process of the distribution function towards its equilibrium is, in many scenarios, simplistic. Based on a previous series of papers, we present a collision model formulated as independent relaxations of the irreducible components of the Hermit coefficients in the reference frame moving with the fluid. These components, corresponding to the irreducible representation of the rotation group, are the minimum tensor components that can be separately relaxed without violating rotation symmetry. For the 2nd, 3rd and 4th moments respectively, two, two and three independent relaxation rates can exist, giving rise to the shear and bulk viscosity, thermal diffusivity and some high-order relaxation process not explicitly manifested in the Navier-Stokes-Fourier equations. Using the binomial transform, the Hermite coefficients are evaluated in the absolute frame to avoid the numerical dissipation introduced by interpolation. Extensive numerical verification is also provided.

math.NA

Rotation symmetry of the multiple-relaxation-time collision model

In the Hermite-expansion-based multiple-relaxation-time lattice Boltzmann (LB) model [Shan & Chen, Int. J. Mod. Phys. C, 18, 635, (2007)], a separate relaxation time is assigned to each of the tensorial moments of the collision term. Here we point out that to allow maximum flexibility while preserving the rotational symmetry of the relaxation physics, separate relaxation times can be assigned to the components of a tensor corresponding to its irreducible representation of SO(3) but not any finer. By decomposing the second moment in the LB model for polyatomic gases [Nie, Shan & Chen, Phys. Rev. E 77, 035701, (2008)], a model with decoupled shear and bulk viscosity is constructed. Hydrodynamic equation of the model is obtained via Chapman-Enskog calculation and verified by numerical simulation.

physics.comp-ph

Self-consistent Force Scheme in the Discrete Boltzmann Equation

In the work of N. Martys et al. [Nicos S. Martys, Xiaowen Shan, Hudong Chen, Phys. Rev. E, Vol. 58, Num.5, 1998 ], a self-consistent force term to any order in the Boltzmann-BKG equation is derived by the Hermite basis with raw velocity. As an extension, in the present work, the force term is expanded by the Hermite basis with the relative velocity in the comoving coordinate and the Hermite basis with the relative velocity scaled by the local temperature. It is found that the force scheme proposed by He et al. [Xiaoyi He, Xiaowen Shan, Gary D. Doolen, Phys. Rev. E, Vol. 57, Num.1,1998] can be derived by the Hermite basis with the relative velocity. Furthermore, another new force scheme in which the velocity is scaled by the local temperature is obtained.

physics.comp-ph

A Meaning-oriented Approach to Semantic Data Modeling

Semantic information is often represented as the entities and the relationships among them with conventional semantic models. This approach is straightforward but is not suitable for many posteriori requests in semantic data modeling. In this paper, we propose a meaning-oriented approach to modeling semantic data and establish a graph-based semantic data model. In this approach we use the meanings, i.e., the subjective views of the entities and relationships, to describe the semantic information, and use the semantic graphs containing the meaning nodes and the meta-meaning relations to specify the taxonomy and the compound construction of the semantic concepts. We demonstrate how this meaning-oriented approach can address many important semantic representation issues, including dynamic specialization and natural join.

cs.DB

Design Issues of JPQ: a Pattern-based Query Language for Document Databases

Document databases are becoming popular, but how to present complex document query to obtain useful information from the document remains an important topic to study. In this paper, we describe the design issues of a pattern-based document database query language named JPQ. JPQ uses various expressive patterns to extract and construct document fragments following a JSON-like document data model. It adopts tree-like extraction patterns with a coherent pattern composition mechanism to extract data elements from hierarchically structured documents and maintain the logical relationships among the elements. Based on these relationships, JPQ deploys a deductive mechanism to declaratively specify the data transformation requests and considers also data filtering on hierarchical data structure. We use various examples to show the features of the language and to demonstrate its expressiveness and declarativeness in presenting complex document queries.

cs.DB

XTQ: A Declarative Functional XML Query Language

Various query languages have been proposed to extract and restructure information in XML documents. These languages, usually claiming to be declarative, mainly consider the conjunctive relationships among data elements. In order to present the operations where the hierarchical and the disjunctive relationships need to be considered, such as restructuring hierarchy and handling heterogeneity, the programs in these languages often exhibit a procedural style and thus the declarativeness in them is not so prominent as in conventional query languages like SQL. In this paper, we propose a declarative pattern-based functional XML query language named XML Tree Query (XTQ). XTQ adopts expressive composite patterns to present data extraction, meanwhile establishing the conjunctive, the disjunctive and the hierarchical relationships among data elements. It uses the matching terms, a composite structure of the variables bound to the matched data elements, to present a global sketch of the extracted data, and develops a deductive restructuring mechanism of matching terms to indicate data transformation, especially for restructuring hierarchy and handling heterogeneity. Based on matching terms, XTQ employs a coherent approach to function declaration and invocation to consistently extract and construct composite data structure, which integrates features of conventional functional languages and pattern-based query languages. Additionally, XTQ also supports data filtering on composite data structure such as hierarchical data, which is seldom deliberately considered in other studies. We demonstrate with various examples that XTQ can declaratively present complex XML queries which are common in practice.

cs.PL