SearcharxivSearch

arXiv subjects

Feng Dai

Publications and source records attributed to Feng Dai.

At least 19 recordsLinked to original sources

Gripper-aware Vision Language Action Models

Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instructions to generate executable action sequences. However, existing VLAs often implicitly assume gripper invariance, despite grasping strategies being inherently embodiment-dependent. Different gripper types, such as parallel-jaw and suction, usually require distinct interaction strategies to achieve the same grasping objective. Moreover, current datasets for VLAs predominantly rely on parallel-jaw grippers, limiting gripper-aware learning. To address this gap, we introduce MiGA, a multi-gripper-aware dataset spanning five distinct gripper types across multiple robots with 103,000 demonstrations, explicitly capturing strategy divergence under shared task objectives. We further propose GVLA, which combines a new multi-gripper tokenizer with adapter-based policy routing. Our new gripper encoding induces structured embedding information that balances parameter sharing and strategy differentiation, while layer-wise probing confirms meaningful gripper-conditioned representations for VLAs. Intensive experiments in both simulation and real-world robots show that our GVLA outperforms the current baselines across evaluated settings. Our method also improves zero-shot generalization or few-shot adaptation to new objects or unseen tasks, and enable more efficient gripper adaptation.

cs.RO

From Cumulative Weights to Marginal Density Ratios: Per-Protocol Estimation in Sequential Target Trial Emulation

Sequential target trial emulation evaluates eligibility at multiple baseline times to emulate a sequence of randomized trials using observational data. Estimating per-protocol effects in this setting is challenging because treatment deviations and loss to follow-up induce selection among individuals who remain observed and adherent over time. Conventional inverse-probability methods address this selection using cumulative weights constructed from estimated adherence and censoring probabilities, but these weights can be highly variable, leading to unstable and imprecise effect estimates. We propose a different approach based on marginal density ratios (MDRs). The MDR directly compares the state distribution among individuals who would remain event-free under a target treatment strategy with the corresponding distribution among observed-adherent individuals. We use longitudinal g-computation to generate the target risk sets and a probabilistic classifier to estimate density ratios for reweighting the observed outcomes. Building on this approach, we also develop a doubly robust extension. Favorable performance across the simulation study suggests that MDR weighting is a promising alternative to cumulative longitudinal weights when its identification assumptions are plausible.

stat.ME

Empirical Approximation of $L_p$ Norms

We study empirical $L_p$ moments of a random vector $\pmb\varphi$ based on its i.i.d.\ copies $\pmb\varphi^1,\ldots,\pmb\varphi^m$, that is, $\frac1m\sum_{j=1}^m |\langle \pmb\varphi^j,y\rangle|^p$. Our main result is a new estimate for the expected uniform deviation \[ \mathbb{E}\sup_{y\in D}\biggl| \frac1m\sum_{j=1}^m |\langle \pmb\varphi^j,y\rangle|^p -\mathbb{E}|\langle \pmb\varphi,y\rangle|^p \biggr| \] over an arbitrary index set $D$. The proof is based on a new bound for Talagrand's $\gamma$-functional, sharper than the standard Dudley-type entropy estimate. We then apply this estimate to the following two problems. First, for $p>2$, we study Marcinkiewicz-type discretization of $L_p$ norms on an $N$-dimensional subspace $X_N\subset B(\Omega)$ of bounded functions on a probability space $(\Omega,\mu)$. We obtain bounds in terms of the norm of the embedding $ (X_N,\|\cdot\|_{L_p(\mu)})\hookrightarrow B(\Omega). $ In particular, we prove that when this norm is of order $N^{1/p}$ and \[ m \ge C(p)\, N\log N\,(\log\log N)^{p-1}, \] then $m$ random samples suffice to approximate the $L_p(\mu)$ norm uniformly on $X_N$ by the sampled discrete $L_p$ norm. This substantially improves the previously known bound in this setting $ m \ge C(p)\, N(\log N)^{\min\{p,3\}}, $ and is optimal up to the factor $(\log\log N)^{p-1}$ in the random-sampling setting. Second, for $1\le p<2$, we obtain an $L_p$ analogue of the restricted isometry property via random sampling for bounded orthogonal systems and, more generally, for $N$-element systems $\mathcal D_N$ satisfying a Riesz-type condition. We prove that when \[ m \ge C(p)\, s\log N\,(\log s)^2\,\log\log s, \] then $m$ random samples suffice to guarantee an $L_p$ restricted isometry-type property uniformly over the class of all $s$-sparse functions generated by $\mathcal D_N$.

math.FA

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

Pursuing training-free open-vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep-seated spatial bias in CLIP. To overcome the limitations of existing solutions, this work moves beyond the CLIP-based paradigm and harnesses the recent spatially-aware dino$.$txt framework to facilitate more efficient and high-quality dense prediction. While dino$.$txt exhibits robust spatial awareness, we find that the semantic ambiguity of text queries gives rise to severe mismatch within its dense cross-modal interactions. To address this, we introduce Visual-guided Prompt evolution (VIP) to rectify the semantic expressiveness of text queries in dino$.$txt, unleashing its potential for fine-grained object perception. Towards this end, VIP integrates alias expansion with a visual-guided distillation mechanism to mine valuable semantic cues, which are robustly aggregated in a saliency-aware manner to yield a high-fidelity prediction. Extensive evaluations demonstrate that VIP: 1. surpasses the top-leading methods by 1.4%-8.4% average mIoU, 2. generalizes well to diverse challenging domains, and 3. requires marginal inference time and memory overhead.

cs.CV

TreeGaussian: Tree-Guided Cascaded Contrastive Learning for Hierarchical Consistent 3D Gaussian Scene Segmentation and Understanding

3D Gaussian Splatting (3DGS) has emerged as a real-time, differentiable representation for neural scene understanding. However, existing 3DGS-based methods struggle to represent hierarchical 3D semantic structures and capture whole-part relationships in complex scenes. Moreover, dense pairwise comparisons and inconsistent hierarchical labels from 2D priors hinder feature learning, resulting in suboptimal segmentation. To address these limitations, we introduce TreeGaussian, a tree-guided cascaded contrastive learning framework that explicitly models hierarchical semantic relationships and reduces redundancy in contrastive supervision. By constructing a multi-level object tree, TreeGaussian enables structured learning across object-part hierarchies. In addition, we propose a two-stage cascaded contrastive learning strategy that progressively refines feature representations from global to local, mitigating saturation and stabilizing training. A Consistent Segmentation Detection (CSD) mechanism and a graph-based denoising module are further introduced to align segmentation modes across views while suppressing unstable Gaussian points, enhancing segmentation consistency and quality. Extensive experiments, including open-vocabulary 3D object selection, 3D point cloud understanding, and ablation studies, demonstrate the effectiveness and robustness of our approach.

cs.CV

Probe-to-Grasp Manipulation Using Self-Sensing Pneumatic Variable-Stiffness Joints

Grasping deformable objects with varying stiffness remains a significant challenge in robotics. Estimating the local stiffness of a target object is important for determining an optimal grasp pose that enables stable pickup without damaging the object. This paper presents a probe-to-grasp manipulation framework for estimating the relative stiffness of objects using a passive soft-rigid two-finger hybrid gripper equipped with self-sensing pneumatic variable-stiffness joints. Each finger of the gripper consists of two rigid links connected by a soft pneumatic ring placed at the joint, enabling both compliant interaction and controllable joint stiffness via internal pressurization. By measuring the pressure inside the pneumatic ring, we can estimate the interaction force during contact. Building on this, we propose a practical probing strategy to infer relative object stiffness by correlating the estimated normal force with known gripper closing displacement. We validate the self-sensing model through stiffness characterization experiments across bending angles and pressure ranges, and demonstrate stiffness-aware probing-and-grasping in real-life applications: selecting grasp locations on fruits with spatially varying stiffness. The proposed system offers a minimal, low-cost sensing approach for stiffness-aware soft manipulation while retaining probing and grasping capability.

cs.RO

Efficient Selection of Type Annotations for Performance Improvement in Gradual Typing

Gradual typing has gained popularity as a design choice for integrating static and dynamic typing within a single language. Several practical languages have adopted gradual typing to offer programmers the flexibility to annotate their programs as needed. Meanwhile there is a key challenge of unexpected performance degradation in partially typed programs. The execution speed may significantly decrease when simply adding more type annotations. Prior studies have investigated strategies of selectively adding type annotations for better performance. However, they are restricted in substantial compilation time, which impedes the practical usage. This paper presents a new technique to select a subset of type annotations derived by type inference for improving the execution performance of gradually typed programs. The advantage of the proposal is shorter compilation time by employing a lightweight, amortized approach. It selects type annotations along the data flows, which is expected to avoid expensive runtime casts caused by a value repeatedly crossing the boundaries between untyped and typed code. We demonstrate the applicability of our proposal, and conduct experiments to validate its effectiveness of improving the execution time on Reticulated Python. Our implementation supports a Python subset to select type annotations derived by an implemented, external type inference engine. Experiment results show that our proposal outperforms a naive strategy of using all type annotations derived by type inference among the benchmark programs. In comparison with an existing approach, the proposal achieves comparable execution speed and shows advantage of maintaining a more stable compilation time of deriving and selecting type annotations. Our results empirically indicate that the proposed technique is practical within Reticulated Python for mitigating the performance bottleneck of gradually typed programs.

cs.PL

Max-Min Neural Network Operators For Approximation of Multivariate Functions

In this paper, we develop a multivariate framework for approximation by max-min neural network operators. Building on the recent advances in approximation theory by neural network operators, particularly, the univariate max-min operators, we propose and analyze new multivariate operators activated by sigmoidal functions. We establish pointwise and uniform convergence theorems and derive quantitative estimates for the order of approximation via modulus of continuity and multivariate generalized absolute moment. Our results demonstrate that multivariate max-min structure of operators, besides their algebraic elegance, provide efficient and stable approximation tools in both theoretical and applied settings.

cs.LG

TEA: Temporal Adaptive Satellite Image Semantic Segmentation

Crop mapping based on satellite images time-series (SITS) holds substantial economic value in agricultural production settings, in which parcel segmentation is an essential step. Existing approaches have achieved notable advancements in SITS segmentation with predetermined sequence lengths. However, we found that these approaches overlooked the generalization capability of models across scenarios with varying temporal length, leading to markedly poor segmentation results in such cases. To address this issue, we propose TEA, a TEmporal Adaptive SITS semantic segmentation method to enhance the model's resilience under varying sequence lengths. We introduce a teacher model that encapsulates the global sequence knowledge to guide a student model with adaptive temporal input lengths. Specifically, teacher shapes the student's feature space via intermediate embedding, prototypes and soft label perspectives to realize knowledge transfer, while dynamically aggregating student model to mitigate knowledge forgetting. Finally, we introduce full-sequence reconstruction as an auxiliary task to further enhance the quality of representations across inputs of varying temporal lengths. Through extensive experiments, we demonstrate that our method brings remarkable improvements across inputs of different temporal lengths on common benchmarks. Our code will be publicly available.

cs.CV

TReFT: Taming Rectified Flow Models For One-Step Image Translation

Rectified Flow (RF) models have advanced high-quality image and video synthesis via optimal transport theory. However, when applied to image-to-image translation, they still depend on costly multi-step denoising, hindering real-time applications. Although the recent adversarial training paradigm, CycleGAN-Turbo, works in pretrained diffusion models for one-step image translation, we find that directly applying it to RF models leads to severe convergence issues. In this paper, we analyze these challenges and propose TReFT, a novel method to Tame Rectified Flow models for one-step image Translation. Unlike previous works, TReFT directly uses the velocity predicted by pretrained DiT or UNet as output-a simple yet effective design that tackles the convergence issues under adversarial training with one-step inference. This design is mainly motivated by a novel observation that, near the end of the denoising process, the velocity predicted by pretrained RF models converges to the vector from origin to the final clean image, a property we further justify through theoretical analysis. When applying TReFT to large pretrained RF models such as SD3.5 and FLUX, we introduce memory-efficient latent cycle-consistency and identity losses during training, as well as lightweight architectural simplifications for faster inference. Pretrained RF models finetuned with TReFT achieve performance comparable to sota methods across multiple image translation datasets while enabling real-time inference.

cs.CV

MetroGS: Efficient and Stable Reconstruction of Geometrically Accurate High-Fidelity Large-Scale Scenes

Recently, 3D Gaussian Splatting and its derivatives have achieved significant breakthroughs in large-scale scene reconstruction. However, how to efficiently and stably achieve high-quality geometric fidelity remains a core challenge. To address this issue, we introduce MetroGS, a novel Gaussian Splatting framework for efficient and robust reconstruction in complex urban environments. Our method is built upon a distributed 2D Gaussian Splatting representation as the core foundation, serving as a unified backbone for subsequent modules. To handle potential sparse regions in complex scenes, we propose a structured dense enhancement scheme that utilizes SfM priors and a pointmap model to achieve a denser initialization, while incorporating a sparsity compensation mechanism to improve reconstruction completeness. Furthermore, we design a progressive hybrid geometric optimization strategy that organically integrates monocular and multi-view optimization to achieve efficient and accurate geometric refinement. Finally, to address the appearance inconsistency commonly observed in large-scale scenes, we introduce a depth-guided appearance modeling approach that learns spatial features with 3D consistency, facilitating effective decoupling between geometry and appearance and further enhancing reconstruction stability. Experiments on large-scale urban datasets demonstrate that MetroGS achieves superior geometric accuracy, rendering quality, offering a unified solution for high-fidelity large-scale scene reconstruction.

cs.CV

Fractional Heat Semigroup Characterization of Distances from Functions in Lipschitz Spaces to Their Subspaces

Let $\Lambda_s$ denote the inhomogeneous Lipschitz space of order $s\in(0,\infty)$ on $\mathbb{R}^n$. This article characterizes the distance $d(f, V)_{\Lambda_s}: = \inf_{g\in V} \|f-g\|_{\Lambda_s}$ from a function $f\in \Lambda_s$ to a non-dense subspace $V\subset \Lambda_s$ via the fractional semigroup $\{T_{\alpha, t}: =e^{-t (-\Delta)^{\alpha/2}}: t\in (0, \infty)\}$ for any $\alpha\in(0,\infty)$. Given an integer $ r >s/\alpha$, a uniformly bounded continuous function $f$ on $\mathbb{R}^n$ belongs to the space $\Lambda_s$ if and only if there exists a constant $\lambda\in(0,\infty)$ such that \begin{align*} \left|(-\Delta)^{\frac {\alpha r}2} (T_{\alpha, t^\alpha } f)(x) \right|\leq \lambda t^{s -r\alpha }\ \ \text{for any $x\in\mathbb{R}^n$ and $t\in (0, 1]$}.\end{align*} The least such constant is denoted by $\lambda_{ \alpha, r, s}(f)$. For each $f\in \Lambda_s$ and $0<\varepsilon< \lambda_{\alpha,r, s}(f)$, let $$ D_{\alpha, r}(s,f,\varepsilon):=\left\{ (x,t)\in \mathbb{R}^n\times (0,1]:\ \left| (-\Delta)^{\frac {\alpha r}2} (T_{\alpha, t^\alpha} f)(x) \right|> \varepsilon t^{s -r \alpha }\right\}$$ be the set of ``bad'' points. To quantify its size, we introduce a class of extended nonnegative \emph{admissible set functions} $\nu$ on the Borel $\sigma$-algebra $\mathcal{B}(\mathbb{R}^n\times [0, 1])$ and define, for any admissible function $\nu$, the \emph{critical index} $ \varepsilon_{\alpha, r, s,\nu}(f):=\inf\{\varepsilon\in(0,\infty):\ \nu(D_{\alpha, r}(s,f,\varepsilon))<\infty\}.$ Our result shows that, for a broad class of subspaces $V\subset \Lambda_s$, including intersections of $\Lambda_s$ with Sobolev, Besov, Triebel--Lizorkin, and Besov-type spaces, there exists an admissible function $\nu$ depending on $V$ such that $\varepsilon_{\alpha, r, s,\nu}(f)\sim \mathrm{dist}(f, V)_{\Lambda_s}.$

math.FA

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-world effectiveness. To address this limitation, we build a large-scale grasping-oriented affordance segmentation benchmark with human-like instructions, named RAGNet. It contains 273k images, 180 categories, and 26k reasoning instructions. The images cover diverse embodied data domains, such as wild, robot, ego-centric, and even simulation data. They are carefully annotated with an affordance map, while the difficulty of language instructions is largely increased by removing their category name and only providing functional descriptions. Furthermore, we propose a comprehensive affordance-based grasping framework, named AffordanceNet, which consists of a VLM pre-trained on our massive affordance data and a grasping network that conditions an affordance map to grasp the target. Extensive experiments on affordance segmentation benchmarks and real-robot manipulation tasks show that our model has a powerful open-world generalization ability. Our data and code is available at https://github.com/wudongming97/AffordanceNet.

cs.CV

Semantic-decoupled Spatial Partition Guided Point-supervised Oriented Object Detection

Given its ability to reduce annotation costs, weakly supervised learning based on single-point annotations has emerged as a research focus in oriented object detection. Compared with the classical teacher-student paradigm, the simple model paradigm (e.g., PointOBB-v2) can substantially further reduce resources required for training while ensuring strong performance. The latter exhibits greater potential for low-cost training, yet such methods still face challenges of insufficient sample assignment and poor pseudo-label quality. In this paper, we propose a training-efficient framework named SSP, which synergizes rule-driven prior injection and data-driven label purification. Specifically, SSP introduces two designs: (1) Pixel-level Spatial Partition-based Sample Assignment, which compactly estimates the upper and lower bounds of object scales and mines high-quality positive samples and hard negative samples through spatial partitioning of pixel maps. (2) Semantic Spatial Partition-based Box Extraction, which derives instances from spatial partitions modulated by semantic maps and converts them into pseudo-boxes for supervising detectors. Experiments on DOTA-v1.0 and other datasets demonstrate SSP's superiority: it achieves +6.73% mAP improvement compared with the baseline, while requiring only 2 h of training time and 6 GB of GPU memory. Furthermore, when SSP is integrated with stronger detector, the mAP can reach 50.81%. The code is available at https://github.com/antxinyuan/ssp.

cs.CV

TopoPoint: Enhance Topology Reasoning via Endpoint Detection in Autonomous Driving

Topology reasoning, which unifies perception and structured reasoning, plays a vital role in understanding intersections for autonomous driving. However, its performance heavily relies on the accuracy of lane detection, particularly at connected lane endpoints. Existing methods often suffer from lane endpoints deviation, leading to incorrect topology construction. To address this issue, we propose TopoPoint, a novel framework that explicitly detects lane endpoints and jointly reasons over endpoints and lanes for robust topology reasoning. During training, we independently initialize point and lane query, and proposed Point-Lane Merge Self-Attention to enhance global context sharing through incorporating geometric distances between points and lanes as an attention mask . We further design Point-Lane Graph Convolutional Network to enable mutual feature aggregation between point and lane query. During inference, we introduce Point-Lane Geometry Matching algorithm that computes distances between detected points and lanes to refine lane endpoints, effectively mitigating endpoint deviation. Extensive experiments on the OpenLane-V2 benchmark demonstrate that TopoPoint achieves state-of-the-art performance in topology reasoning (48.8 on OLS). Additionally, we propose DET$_p$ to evaluate endpoint detection, under which our method significantly outperforms existing approaches (52.6 v.s. 45.2 on DET$_p$). The code is released at https://github.com/Franpin/TopoPoint.

cs.CV

Difference and Wavelet Characterizations of Distances from Functions in Lipschitz Spaces to Their Subspaces

Let $\Lambda_s$ denote the Lipschitz space of order $s\in(0,\infty)$ on $\mathbb{R}^n$, which consists of all $f\in\mathfrak{C}\cap L^\infty$ such that, for some constant $L\in(0,\infty)$ and some integer $r\in(s,\infty)$, \begin{equation*} \label{0-1}\Delta_r f(x,y): =\sup_{|h|\leq y} |\Delta_h^r f(x)|\leq L y^s, \ x\in\mathbb{R}^n, \ y \in(0, 1]. \end{equation*} Here (and throughout the article) $\mathfrak{C}$ refers to continuous functions, and $\Delta_h^r$ is the usual $r$-th order difference operator with step $h\in\mathbb{R}^n$. For each $f\in \Lambda_s$ and $\varepsilon\in(0,L)$, let $ S(f,\varepsilon):= \{ (x,y)\in\mathbb{R}^n\times [0,1]: \frac {\Delta_r f(x,y)}{y^s}>\varepsilon\}$, and let $\mu: \mathcal{B}(\mathbb{R}_+^{n+1})\to [0,\infty]$ be a suitably defined nonnegative extended real-valued function on the Borel $\sigma$-algebra of subsets of $\mathbb{R}_+^{n+1}$. Let $\varepsilon(f)$ be the infimum of all $\varepsilon\in(0,\infty)$ such that $\mu(S(f,\varepsilon))<\infty$. The main target of this article is to characterize the distance from $f$ to a subspace $V\cap \Lambda_s$ of $\Lambda_s$ for various function spaces $V$ (including Sobolev, Besov--Triebel--Lizorkin, and Besov--Triebel--Lizorkin-type spaces) in terms of $\varepsilon(f)$, showing that \begin{equation*} \varepsilon(f)\sim \mathrm{dist} (f, V\cap \Lambda_s)_{\Lambda_s}: = \inf_{g\in \Lambda_s\cap V} \|f-g\|_{\Lambda_s}.\end{equation*} Moreover, we present our results in a general framework based on quasi-normed lattices of function sequences $X$ and Daubechies $s$-Lipschitz $X$-based spaces.

math.FA

Exact: Exploring Space-Time Perceptive Clues for Weakly Supervised Satellite Image Time Series Semantic Segmentation

Automated crop mapping through Satellite Image Time Series (SITS) has emerged as a crucial avenue for agricultural monitoring and management. However, due to the low resolution and unclear parcel boundaries, annotating pixel-level masks is exceptionally complex and time-consuming in SITS. This paper embraces the weakly supervised paradigm (i.e., only image-level categories available) to liberate the crop mapping task from the exhaustive annotation burden. The unique characteristics of SITS give rise to several challenges in weakly supervised learning: (1) noise perturbation from spatially neighboring regions, and (2) erroneous semantic bias from anomalous temporal periods. To address the above difficulties, we propose a novel method, termed exploring space-time perceptive clues (Exact). First, we introduce a set of spatial clues to explicitly capture the representative patterns of different crops from the most class-relative regions. Besides, we leverage the temporal-to-class interaction of the model to emphasize the contributions of pivotal clips, thereby enhancing the model perception for crop regions. Build upon the space-time perceptive clues, we derive the clue-based CAMs to effectively supervise the SITS segmentation network. Our method demonstrates impressive performance on various SITS benchmarks. Remarkably, the segmentation network trained on Exact-generated masks achieves 95% of its fully supervised performance, showing the bright promise of weakly supervised paradigm in crop mapping scenario. Our code will be publicly available.

cs.CV

Compression using Quasi-Interpolation

We consider quasi-interpolation with a main application in radial basis function approximations and compression in this article. Constructing and using these quasi-interpolants, we consider wavelet and compression-type approximations from their linear spaces and provide convergence estimates. The results include an error estimate for nonlinear approximation by quasi-interpolation, results about compression in the space of continuous functions and a pointwise convergence estimate for approximands of low smoothness.

math.NA