SearcharxivSearch

arXiv subjects

Zhouyi Wu

Publications and source records attributed to Zhouyi Wu.

4 recordsLinked to original sources

Beyond Appearance: Can Multimodal Large Language Models Exploit Vertical Structure for Remote Sensing Natural Scene Understanding?

Multimodal large language models (MLLMs) have advanced rapidly in remote-sensing analysis, yet existing evaluations remain predominantly 2D-centric. Because spectrally confused regions can appear nearly identical yet differ substantially in vertical structure, appearance alone is often insufficient for reliable semantic interpretation in natural scenes. Vertical structure therefore provides decision-critical physical evidence, yet whether current MLLMs can effectively perceive, ground, and utilize such geometric evidence remains underexplored. To bridge this gap, we introduce VertiCue-Bench, the first diagnostic benchmark that uses controlled interventions to probe whether vertical height evidence is actually perceived, grounded, and utilized, and we establish a three-stage evidence-utilization framework of Perception--Grounding--Utilization. By constructing a Representation Intervention Spectrum spanning multiple presentation and interaction modalities, including Raw Visual, Tool-assisted, and Oracle Text conditions, together with controlled counterfactual tests, we conduct an in-depth disentangled diagnosis across 10 state-of-the-art models. Our experiments reveal and formally characterize the Vertical Structure Utilization Gap. Although current models exhibit emerging geometric perception capabilities, they still struggle to accurately ground vertical evidence to relevant spatial entities and integrate it into high-level semantic decisions. This finding identifies a critical bottleneck in developing physically grounded and geometry-aware remote-sensing MLLMs.

cs.CV

Hand Gestures Recognition in Videos Taken with Lensless Camera

A lensless camera is an imaging system that uses a mask in place of a lens, making it thinner, lighter, and less expensive than a lensed camera. However, additional complex computation and time are required for image reconstruction. This work proposes a deep learning model named Raw3dNet that recognizes hand gestures directly on raw videos captured by a lensless camera without the need for image restoration. In addition to conserving computational resources, the reconstruction-free method provides privacy protection. Raw3dNet is a novel end-to-end deep neural network model for the recognition of hand gestures in lensless imaging systems. It is created specifically for raw video captured by a lensless camera and has the ability to properly extract and combine temporal and spatial features. The network is composed of two stages: 1. spatial feature extractor (SFE), which enhances the spatial features of each frame prior to temporal convolution; 2. 3D-ResNet, which implements spatial and temporal convolution of video streams. The proposed model achieves 98.59% accuracy on the Cambridge Hand Gesture dataset in the lensless optical experiment, which is comparable to the lensed-camera result. Additionally, the feasibility of physical object recognition is assessed. Furtherly, we show that the recognition can be achieved with respectable accuracy using only a tiny portion of the original raw data, indicating the potential for reducing data traffic in cloud computing scenarios.

cs.CV

Text detection and recognition based on a lensless imaging system

Lensless cameras are characterized by several advantages (e.g., miniaturization, ease of manufacture, and low cost) as compared with conventional cameras. However, they have not been extensively employed due to their poor image clarity and low image resolution, especially for tasks that have high requirements on image quality and details such as text detection and text recognition. To address the problem, a framework of deep-learning-based pipeline structure was built to recognize text with three steps from raw data captured by employing lensless cameras. This pipeline structure consisted of the lensless imaging model U-Net, the text detection model connectionist text proposal network (CTPN), and the text recognition model convolutional recurrent neural network (CRNN). Compared with the method focusing only on image reconstruction, UNet in the pipeline was able to supplement the imaging details by enhancing factors related to character categories in the reconstruction process, so the textual information can be more effectively detected and recognized by CTPN and CRNN with fewer artifacts and high-clarity reconstructed lensless images. By performing experiments on datasets of different complexities, the applicability to text detection and recognition on lensless cameras was verified. This study reasonably demonstrates text detection and recognition tasks in the lensless camera system,and develops a basic method for novel applications.

cs.CV

Demonstration of topological wireless power transfer

Recent advances in non-radiative wireless power transfer (WPT) technique essentially relying on magnetic resonance and near-field coupling have successfully enabled a wide range of applications. However, WPT systems based on double resonators are severely limited to short- or mid-range distance, due to the deteriorating efficiency and power with long transfer distance. WPT systems based on multi-relay resonators can overcome this problem, which, however, suffer from sensitivity to perturbations and fabrication imperfections. Here, we experimentally demonstrate a concept of topological wireless power transfer (TWPT), where energy is transferred efficiently via the near-field coupling between two topological edge states localized at the ends of a one-dimensional radiowave topological insulator. Such a TWPT system can be modelled as a parity-time-symmetric Su-Schrieffer-Heeger (SSH) chain with complex boundary potentials. Besides, the coil configurations are judiciously designed, which significantly suppress the unwanted cross-couplings between nonadjacent coils that could break the chiral symmetry of the SSH chain. By tuning the inter- and intra-cell coupling strengths, we theoretically and experimentally demonstrate high energy transfer efficiency near the exceptional point of the topological edge states, even in the presence of disorder. The combination of topological metamaterials, non-Hermitian physics, and WPT techniques could promise a variety of robust, efficient WPT applications over long distances in electronics, transportation, and industry.

physics.app-ph