SearcharxivSearch

arXiv subjects

Wensheng Wang

Publications and source records attributed to Wensheng Wang.

10 recordsLinked to original sources

SpaceDex: Generalizable Dexterous Grasping in Tiered Workspaces

Generalizable grasping with high-degree-of-freedom (DoF) dexterous hands remains challenging in tiered workspaces, where occlusion, narrow clearances, and height-dependent constraints are substantially stronger than in open tabletop scenes. Most existing methods are evaluated in relatively unoccluded settings and typically do not explicitly model the distinct control requirements of arm navigation and hand articulation under spatial constraints. We present SpaceDex, a hierarchical framework for dexterous manipulation in constrained 3D environments. At the high level, a Vision-Language Model (VLM) planner parses user intent, reasons about occlusion and height relations across multiple camera views, and generates target bounding boxes for zero-shot segmentation and mask tracking. This stage provides structured spatial guidance for downstream control instead of relying on single-view target selection. At the low level, we introduce an arm-hand Feature Separation Network that decouples global trajectory control for the arm from geometry-aware grasp mode selection for the hand, reducing feature interference between reaching and grasping objectives. The controller further integrates multi-view perception, fingertip tactile sensing, and a small set of recovery demonstrations to improve robustness to partial observability and off-nominal contacts. In 100 real-world trials involving over 30 unseen objects across four categories, SpaceDex achieves a 63.0\% success rate, compared with 39.0\% for a strong tabletop baseline. These results indicate that combining hierarchical spatial planning with arm-hand representation decoupling improves dexterous grasping performance in spatially constrained environments.

cs.RO

Global Rice Multi-Class Segmentation Dataset (RiceSEG): A Comprehensive and Diverse High-Resolution RGB-Annotated Images for the Development and Benchmarking of Rice Segmentation Algorithms

Developing computer vision-based rice phenotyping techniques is crucial for precision field management and accelerating breeding, thereby continuously advancing rice production. Among phenotyping tasks, distinguishing image components is a key prerequisite for characterizing plant growth and development at the organ scale, enabling deeper insights into eco-physiological processes. However, due to the fine structure of rice organs and complex illumination within the canopy, this task remains highly challenging, underscoring the need for a high-quality training dataset. Such datasets are scarce, both due to a lack of large, representative collections of rice field images and the time-intensive nature of annotation. To address this gap, we established the first comprehensive multi-class rice semantic segmentation dataset, RiceSEG. We gathered nearly 50,000 high-resolution, ground-based images from five major rice-growing countries (China, Japan, India, the Philippines, and Tanzania), encompassing over 6,000 genotypes across all growth stages. From these original images, 3,078 representative samples were selected and annotated with six classes (background, green vegetation, senescent vegetation, panicle, weeds, and duckweed) to form the RiceSEG dataset. Notably, the sub-dataset from China spans all major genotypes and rice-growing environments from the northeast to the south. Both state-of-the-art convolutional neural networks and transformer-based semantic segmentation models were used as baselines. While these models perform reasonably well in segmenting background and green vegetation, they face difficulties during the reproductive stage, when canopy structures are more complex and multiple classes are involved. These findings highlight the importance of our dataset for developing specialized segmentation models for rice and other crops.

eess.IV

HybridGen: VLM-Guided Hybrid Planning for Scalable Data Generation of Imitation Learning

The acquisition of large-scale and diverse demonstration data are essential for improving robotic imitation learning generalization. However, generating such data for complex manipulations is challenging in real-world settings. We introduce HybridGen, an automated framework that integrates Vision-Language Model (VLM) and hybrid planning. HybridGen uses a two-stage pipeline: first, VLM to parse expert demonstrations, decomposing tasks into expert-dependent (object-centric pose transformations for precise control) and plannable segments (synthesizing diverse trajectories via path planning); second, pose transformations substantially expand the first-stage data. Crucially, HybridGen generates a large volume of training data without requiring specific data formats, making it broadly applicable to a wide range of imitation learning algorithms, a characteristic which we also demonstrate empirically across multiple algorithms. Evaluations across seven tasks and their variants demonstrate that agents trained with HybridGen achieve substantial performance and generalization gains, averaging a 5% improvement over state-of-the-art methods. Notably, in the most challenging task variants, HybridGen achieves significant improvement, reaching a 59.7% average success rate, significantly outperforming Mimicgen's 49.5%. These results demonstrating its effectiveness and practicality.

cs.RO

UVCPNet: A UAV-Vehicle Collaborative Perception Network for 3D Object Detection

With the advancement of collaborative perception, the role of aerial-ground collaborative perception, a crucial component, is becoming increasingly important. The demand for collaborative perception across different perspectives to construct more comprehensive perceptual information is growing. However, challenges arise due to the disparities in the field of view between cross-domain agents and their varying sensitivity to information in images. Additionally, when we transform image features into Bird's Eye View (BEV) features for collaboration, we need accurate depth information. To address these issues, we propose a framework specifically designed for aerial-ground collaboration. First, to mitigate the lack of datasets for aerial-ground collaboration, we develop a virtual dataset named V2U-COO for our research. Second, we design a Cross-Domain Cross-Adaptation (CDCA) module to align the target information obtained from different domains, thereby achieving more accurate perception results. Finally, we introduce a Collaborative Depth Optimization (CDO) module to obtain more precise depth estimation results, leading to more accurate perception outcomes. We conduct extensive experiments on both our virtual dataset and a public dataset to validate the effectiveness of our framework. Our experiments on the V2U-COO dataset and the DAIR-V2X dataset demonstrate that our method improves detection accuracy by 6.1% and 2.7%, respectively.

cs.CV

A Real-time Low-cost Artificial Intelligence System for Autonomous Spraying in Palm Plantations

In precision crop protection, (target-orientated) object detection in image processing can help navigate Unmanned Aerial Vehicles (UAV, crop protection drones) to the right place to apply the pesticide. Unnecessary application of non-target areas could be avoided. Deep learning algorithms dominantly use in modern computer vision tasks which require high computing time, memory footprint, and power consumption. Based on the Edge Artificial Intelligence, we investigate the main three paths that lead to dealing with this problem, including hardware accelerators, efficient algorithms, and model compression. Finally, we integrate them and propose a solution based on a light deep neural network (DNN), called Ag-YOLO, which can make the crop protection UAV have the ability to target detection and autonomous operation. This solution is restricted in size, cost, flexible, fast, and energy-effective. The hardware is only 18 grams in weight and 1.5 watts in energy consumption, and the developed DNN model needs only 838 kilobytes of disc space. We tested the developed hardware and software in comparison to the tiny version of the state-of-art YOLOv3 framework, known as YOLOv3-Tiny to detect individual palm in a plantation. An average F1 score of 0.9205 at the speed of 36.5 frames per second (in comparison to similar accuracy at 18 frames per second and 8.66 megabytes of the YOLOv3-Tiny algorithm) was reached. This developed detection system is easily plugged into any machines already purchased as long as the machines have USB ports and run Linux Operating System.

cs.CV

Making Streett Determinization Tight

Optimal determinization construction of Streett automata is an important research problem because it is indispensable in numerous applications such as decision problems for tree temporal logics, logic games and system synthesis. This paper presents a transformation from nondeterministic Streett automata (NSA) with $n$ states and $k$ Streett pairs to equivalent deterministic Rabin transition automata (DRTA) with $n^{5n}(n!)^{n}$ states, $O(n^{n^2})$ Rabin pairs for $k=ω(n)$ and $n^{5n}k^{nk}$ states, $O(k^{nk})$ Rabin pairs for $k=O(n)$. This improves the state of the art Streett determinization construction with $n^{5n}(n!)^{n+1}$ states, $O(n^2)$ Rabin pairs and $n^{5n}k^{nk}n!$ states, $O(nk)$ Rabin pairs, respectively. Moreover, deterministic parity transition automata (DPTA) are obtained with $3(n(n+1)-1)!(n!)^{n+1}$ states, $2n(n+1)$ priorities for $k=ω(n)$ and $3(n(k+1)-1)!n!k^{nk}$ states, $2n(k+1)$ priorities for $k=O(n)$, which improves the best construction with $n^{n}(k+1)^{n(k+1)}(n(k+1)-1)!$ states, $2n(k+1)$ priorities. Further, we prove a lower bound state complexity for determinization construction from NSA to deterministic Rabin (transition) automata i.e. $n^{5n}(n!)^{n}$ for $k=ω(n)$ and $n^{5n}k^{nk}$ for $k=O(n)$, which matches the state complexity of the proposed determinization construction. Besides, we put forward a lower bound state complexity for determinization construction from NSA to deterministic parity (transition) automata i.e. $2^{Ω(n^2 \log n)}$ for $k=ω(n)$ and $2^{Ω(nk \log nk)}$ for $k=O(n)$, which is the same as the state complexity of the proposed determinization construction in the exponent.

cs.FL

Stimulated emission depletion microscopy with array detection and photon reassignment

We propose a novel stimulated emission depletion (STED) microscopy based on array detection and photon reassignment. By replacing the single-point detector in traditional STED with a detector array and utilizing the photon reassignment method to recombine the images acquired by each detector, the final photon reassignment STED (prSTED) image could be obtained. We analyze the principle and imaging characteristics of prSTED, and the results indicate that, compared with traditional STED, prSTED can improve the signal-to-noise ratio (SNR) of the image by increasing the obtained photon flux while maintaining the original spatial resolution of STED. In addition, the SNR and resolution of prSTED are strongly correlated with the intensity of depletion beam. Corresponding theoretical and experimental analysis about this feature are also conducted. In general, considering the enhanced signal strength, imaging speed and compatibility with some other imaging techniques, we believe prSTED would be a helpful promotion in biomedical imaging.

physics.optics

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from appearance and motion features; 2) a redesigned question memory which helps understand the complex semantics of question and highlights queried subjects; and 3) a new multimodal fusion layer which performs multi-step reasoning by attending to relevant visual and textual hints with self-updated attention. Our VideoQA model firstly generates the global context-aware visual and textual features respectively by interacting current inputs with memory contents. After that, it makes the attentional fusion of the multimodal visual and textual representations to infer the correct answer. Multiple cycles of reasoning can be made to iteratively refine attention weights of the multimodal data and improve the final representation of the QA pair. Experimental results demonstrate our approach achieves state-of-the-art performance on four VideoQA benchmark datasets.

cs.CV

Nonlinear focal modulation microscopy

Here we report nonlinear focal modulation microscopy (NFOMM) to achieve super-resolution imaging. Abandoning the previous persistence on minimizing the size of Gaussian emission pattern by directly narrowing (e.g. Minimizing the detection pinhole in Airyscan, Zeiss) or by indirectly peeling its outer profiles (e.g., Depleting the outer emission region in STED, stimulated emission microscopy) in pointwise scanning scenarios, we stick to a more general basis------ maximizing the system frequency shifting ability. In NFOMM, we implement a nonlinear focal modulation by applying phase modulations with high-intensity illumination, thereby extending the effective spatial-frequency bandwidth of the imaging system for reconstructing super-resolved images. NFOMM employs a spatial light modulator (SLM) for assisting pattern-modulated pointwise scanning, making the system single beam path while achieving a transverse resolution of 60 nm on imaging fluorescent nanoparticles. While exploring a relatively simple and flexible system, the imagingperformance of NFOMM is comparable with STED as evidenced in imaging nuclear pore complexes, demonstrating NFOMM is a suitable observation tool for fundamental studies in biology. Since NFOMM is implemented as an add-on module to an existing laser scanning microscope and easy to be aligned, we anticipate it will be adopted rapidly by the biological community.

physics.optics

Exact moduli of continuity for operator-scaling Gaussian random fields

Let $X=\{X(t),t\in\mathrm{R}^N\}$ be a centered real-valued operator-scaling Gaussian random field with stationary increments, introduced by Biermé, Meerschaert and Scheffler (Stochastic Process. Appl. 117 (2007) 312-332). We prove that $X$ satisfies a form of strong local nondeterminism and establish its exact uniform and local moduli of continuity. The main results are expressed in terms of the quasi-metric $τ_E$ associated with the scaling exponent of $X$. Examples are provided to illustrate the subtle changes of the regularity properties.

math.ST