SearcharxivSearch

arXiv subjects

Donghao Wang

Publications and source records attributed to Donghao Wang.

11 recordsLinked to original sources

Collaborative Multi-Mode Pruning for Vision-Language Models

Vision-Language Models (VLMs) have advanced rapidly within the unified Transformer architecture, yet their deployment on resource-constrained devices remains challenging due to high computational complexity. While pruning has emerged as an effective technique for compressing VLMs, existing approaches predominantly focus on a single mode by pruning either parameters or tokens, neglecting fully exploring the inherent redundancy in each mode, which leads to substantial performance degradation at high pruning ratios. To address the above limitations, we propose Collaborative Multi-Mode Pruning (CoMP), a novel framework tailored for VLMs by performing joint parameter and token pruning. Specifically, we first design a Collaborative Importance Metric (CIM) that investigates the mutual interference between the coupled parameters and tokens. It incorporates distinct significance of tokens into the computation of parameter importance scores, while simultaneously mitigating the affect of pruned parameters on token importance scores. Moreover, we develop a Multi-Mode Pruning Strategy (MPS) that decomposes the overall pruning process into a sequence of pruning stages, while in each stage we estimate the priory of different pruning modes based on their pruning cost and adaptively shift to the optimal one. Additionally, MPS integrates the historical cost and random exploration, in order to achieve a stable pruning process and avoid local optimum. Extensive experiments across various vision-language tasks and models demonstrate that our method effectively promotes the performance under high pruning ratios by comparing to the state-of-the-art approaches. The source code is available at https://github.com/Wuzimeng/CoMP.git.

cs.CV

Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution

In this report, we introduce Xiaomi-Robotics-0, an advanced vision-language-action (VLA) model optimized for high performance and fast and smooth real-time execution. The key to our method lies in a carefully designed training recipe and deployment strategy. Xiaomi-Robotics-0 is first pre-trained on large-scale cross-embodiment robot trajectories and vision-language data, endowing it with broad and generalizable action-generation capabilities while avoiding catastrophic forgetting of the visual-semantic knowledge of the underlying pre-trained VLM. During post-training, we propose several techniques for training the VLA model for asynchronous execution to address the inference latency during real-robot rollouts. During deployment, we carefully align the timesteps of consecutive predicted action chunks to ensure continuous and seamless real-time rollouts. We evaluate Xiaomi-Robotics-0 extensively in simulation benchmarks and on two challenging real-robot tasks that require precise and dexterous bimanual manipulation. Results show that our method achieves state-of-the-art performance across all simulation benchmarks. Moreover, Xiaomi-Robotics-0 can roll out fast and smoothly on real robots using a consumer-grade GPU, achieving high success rates and throughput on both real-robot tasks. To facilitate future research, code and model checkpoints are open-sourced at https://xiaomi-robotics-0.github.io

cs.RO

Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference

Long-context inference enhances the reasoning capability of Large Language Models (LLMs), but incurs significant computational overhead. Token-oriented methods, such as pruning and skipping, have shown great promise in reducing inference latency, yet still suffer from inherently insufficient structure optimization, outdated selection criteria, and redundancy interference, resulting in suboptimal speed-accuracy trade-off. To address these issues, we propose a novel training-free framework dubbed Self-Predictive Token Skipping (SPTS), for efficient long-context LLM inference. Specifically, motivated by probing the influence of target layers prior to skipping, we design two selective token skipping strategies for typical structures, including Partial Attention Probing (PAP) for multi-head attention and Low-rank Transformation Probing (LTP) for feed forward network. The former selects informative tokens via partial forward attention computation, while the latter constructs a low-rank proxy network to predict token transformations. In addition, a Multi-Stage Delayed Pruning (MSDP) strategy reallocates skipping budgets and progressively removes redundant tokens across layers. Extensive experiments display the effectiveness of our method, achieving up to 2.46$\times$ and 2.29$\times$ speedups for prefilling and end-to-end generation, respectively, while maintaining state-of-the-art accuracy. We will release the source code upon acceptance.

cs.CL

FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing

Large vision-language models (VLMs) exhibit strong performance across various tasks. However, these VLMs encounter significant challenges when applied to the remote sensing domain due to the inherent differences between remote sensing images and natural images. Existing remote sensing VLMs often fail to extract fine-grained visual features and suffer from visual forgetting during deep language processing. To address this, we introduce MF-RSVLM, a Multi-Feature Fusion Remote Sensing Vision--Language Model that effectively extracts and fuses visual features for RS understanding. MF-RSVLM learns multi-scale visual representations and combines global context with local details, improving the capture of small and complex structures in RS scenes. A recurrent visual feature injection scheme ensures the language model remains grounded in visual evidence and reduces visual forgetting during generation. Extensive experiments on diverse RS benchmarks show that MF-RSVLM achieves state-of-the-art or highly competitive performance across remote sensing classification, image captioning, and VQA tasks. Our code is publicly available at https://github.com/Yunkaidang/RSVLM.

cs.CV

A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs

Multimodal large language models (MLLMs) demonstrate strong perception and reasoning performance on existing remote sensing (RS) benchmarks. However, most prior benchmarks rely on low-resolution imagery, and some high-resolution benchmarks suffer from flawed reasoning-task designs. We show that text-only LLMs can perform competitively with multimodal vision-language models on RS reasoning tasks without access to images, revealing a critical mismatch between current benchmarks and the intended evaluation of visual understanding. To enable faithful assessment, we introduce RSHR-Bench, a super-high-resolution benchmark for RS visual understanding and reasoning. RSHR-Bench contains 5,329 full-scene images with a long side of at least 4,000 pixels, with up to about 3 x 10^8 pixels per image, sourced from widely used RS corpora and UAV collections. We design four task families: multiple-choice VQA, open-ended VQA, image captioning, and single-image evaluation. These tasks cover nine perception categories and four reasoning types, supporting multi-turn and multi-image dialog. To reduce reliance on language priors, we apply adversarial filtering with strong LLMs followed by rigorous human verification. Overall, we construct 3,864 VQA tasks, 3,913 image captioning tasks, and 500 fully human-written or verified single-image evaluation VQA pairs. Evaluations across open-source, closed-source, and RS-specific VLMs reveal persistent performance gaps in super-high-resolution scenarios. Code: https://github.com/Yunkaidang/RSHR

cs.CV

Monopoles and Landau-Ginzburg Models III: A Gluing Theorem

This is the third paper of this series. In \cite{Wang20}, we defined the monopole Floer homology for any pair $(Y,ω)$, where $Y$ is a compact oriented 3-manifold with toroidal boundary and $ω$ is a suitable closed 2-form viewed as a decoration. In this paper, we establish a gluing theorem for this Floer homology when two such 3-manifolds are glued suitably along their common boundary, assuming that $\partial Y$ is disconnected, and $ω$ is small and yet non-vanishing on $\partial Y$. As applications, we construct a monopole Floer 2-functor and the generalized cobordism maps. Using results of Kronheimer-Mrowka and Ni, it is shown that for any such 3-manifold $Y$ that is irreducible, this Floer homology detects the Thurston norm on $H_2(Y,\partial Y;\mathbb{R})$ and the fiberness of $Y$. Finally, we show that our construction recovers the monopole link Floer homology for any link inside a closed 3-manifold.

math.GT

Monopoles and Landau-Ginzburg Models II: Floer Homology

This is the second paper in this series. Following the setup of Meng-Taubes, we define the monopole Floer homology for any pair $(Y,ω)$, where $Y$ is a compact oriented 3-manifold with toroidal boundary and $ω$ is a suitable closed 2-form viewed as a decoration. This construction fits into a (3+1)-topological quantum field theory and generalizes the work of Kronheimer-Mrowka for closed oriented 3-manifolds. By a theorem of Meng-Taubes and Turaev, the Euler characteristic of this Floer homology recovers the Milnor-Turaev torsion invariant of the 3-manifold.

math.GT

Monopoles and Landau-Ginzburg Models I

The endpoint of this series of papers is to construct the monopole Floer homology for any pair $(Y,ω)$, where $Y$ is a compact oriented 3-manifold with toroidal boundary and $ω$ is a suitable closed 2-form. In the first paper, we exploit the framework of the gauged Landau-Ginzburg models to address two model problems for the (perturbed) Seiberg-Witten moduli spaces on either $\mathbb{C}\timesΣ$ or $\mathbb{H}^2_+\timesΣ$, where $Σ$ is any compact Riemann surface of genus $\geq 1$. Our first result states that finite energy solutions to the perturbed equations on $\mathbb{C}\timesΣ$ are necessarily trivial. The second states that small energy solutions on $\mathbb{H}^2_+\timesΣ$ necessarily have energy decaying exponentially in the spatial direction. These results will lead eventually to the compactness theorem in the second paper.

math.GT

The Complex Gradient Flow Equation and Seidel's Spectral Sequence

Following the proposals of Donaldson-Thomas, Haydys and Gaiotto-Moore-Witten, we give a construction of Fukaya-Seidel categories for a suitable class of Morse Landau-Ginzburg models using the complex gradient flow equation, which has the potential for generalization to some infinite dimensional examples. In the course of this construction, we give an alternative proof to Seidel's spectral sequence for Lagrangian Floer cohomology, which can be viewed as a finite dimensional model for a potential bordered monopole Floer theory. The key observation is that under a neck-stretching limit, this complex gradient flow equation produces a natural geometric filtration on the Floer cochain complex. The resulting spectral sequence is then identified with Seidel's original one.

math.SG

On finite energy monopoles on $\mathbb{C}\times Σ$

Let $X=\mathbb{C}\timesΣ$ be the product of the complex plane and a compact Riemann surface. We establish a classification theorem of solutions to the Seiberg-Witten equation on $X$ with finite analytic energy. The spin bundle $S^+\to X$ splits as $L^+\oplus L^-$. When $2-2g\leq c_1(S^+)[Σ]<0$, the moduli space is in bijection with the moduli space of pairs $((L^+,\bar{\partial}), f)$ where $(L^+,\bar{\partial})$ is a holomorphic structure on $L^+$ and $f: \mathbb{C}\to H^0(Σ, L^+,\bar{\partial})$ is a polynomial map. Moreover, the solution has analytic energy $-4π^2d\cdot c_1(S^+)[Σ]$ if $f$ has degree $d$. When $c_1(S^+)=0$, all solutions are reducible and the moduli space is the space of flat connections on $\bigwedge^2 S^+$. We also estimate the decay rate of these solutions at infinity.

math.DG

The Fenchel-type inequality in the 3-dimensional Lorentz space and a Crofton formula

We generalize the Fenchel theorem to strong spacelike (which means that the tangent vector and the curvature vector span a spacelike 2-plane at each point) closed curves with index 1 in the 3-dimensional Lorentz space, showing that the total curvatures must be less than or equal to $2π$. A similar generalization of the Fary-Milnor theorem is also obtained. We establish the Crofton formula on the de Sitter 2-sphere which implies the above results.

math.DG