SearcharxivSearch

arXiv subjects

Lan Fu

Publications and source records attributed to Lan Fu.

At least 19 recordsLinked to original sources

Self-powered InAs nanowire detector arrays for extended-SWIR spectrometry at room temperature

Spectral sensing in the extended shortwave infrared (e-SWIR) is important for molecular analysis, infrared imaging, and machine vision, motivating the development of compact spectrometers for broader applications. However, conventional commercial off-the-shelf spectrometers in this wavelength region are expensive and bulky due to their reliance on external dispersive optics/filters and/or cryogenic accessories. Other emerging computational spectrometers are based on Si and InGaAs photodetectors that remain focused on the visible and near-infrared, with few detector platforms operating in the e-SWIR regime that simultaneously provide broadband sensitivity, low-noise room-temperature operation, and diverse spectral signatures for accurate identification and reconstruction. Here, we report a room-temperature e-SWIR computational spectrometer based on InAs/InP core-shell nanowire photodetector arrays with geometry-encoded spectral responses. The detectors exhibit self-powered broadband photoresponse across the 1--3 $\mu$m range, with responsivity up to 0.215 A W$^{-1}$, detectivity up to $1.6 \times 10^{9}$ cm Hz$^{1/2}$ W$^{-1}$, and microsecond response times. The excellent detector performance is leveraged to demonstrate filter-free spectral reconstruction using a compact multipixel photodetector array device. This enables high-accuracy molecular absorption spectrum reconstruction and hyperspectral imaging. Our results indicate that InAs nanowire arrays are a promising platform for compact computational spectrometry and imaging in the e-SWIR at room temperature.

physics.optics

ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

Multimodal Large Language Models (MLLMs) have recently made strong progress in visual--linguistic understanding. However, their performance on text-centric video reasoning remains highly sensitive to input quality. Real-world user-provided videos often contain motion blur, compression artifacts, noise, and low-resolution text, which impair reliable text reading and downstream reasoning. Whether MLLMs can robustly read and reason about real-world scene text under diverse quality conditions remains a fundamental open question. We introduce ClearText-Video (CTVid), a large-scale, scene-text-aware benchmark for studying text-centric video understanding under controlled quality variation. CTVid contains 4,639 real-world text-rich egocentric videos, 550K+ frames, 1.6M human-verified scene-text annotations, and 220K+ spatial/temporal question--answer pairs in Chinese and English. For each high-quality video, CTVid provides content-matched Degraded-Quality and Restored-Quality variants, supporting two task families: Text-Centric Video Restoration and Multi-Quality VideoQA. We evaluate 18 representative restoration methods and 16 state-of-the-art MLLMs on CTVid. The results show that visual enhancement does not guarantee textual fidelity or downstream reasoning gains: blur is more damaging than low resolution, restored videos can alter the textual evidence used by MLLMs, and OCR-only pipelines remain far below direct multimodal reasoning. CTVid exposes the gap between video restoration and text-grounded understanding, providing a rigorous foundation for restoration-aware, quality-robust text-centric video systems.

cs.CV

Tuning Plasmonic Metasurfaces via Phase Change Material Substrates for Modulating Reactivity in Light-Driven Reactions

Phase change materials provide a powerful platform for dynamically modulating optical responses in nanophotonic systems. While plasmonic metasurfaces have been widely employed to enhance photocatalytic efficiency and promote particular light-driven reactions, active and dynamical control over reaction pathways within a single device remains challenging. Here, we report a phase-induced tunable metasurface that tailors photoexcited electron populations through mode hybridization, enabling selective control over the reactivity of light-driven chemical processes. By exploiting thermally induced refractive-index switching in a Sb2S3 cavity, the plasmonic resonance strength of Au nanodisks is actively tuned via cavity-plasmon hybridization. This reconfiguration modulates the product yield of methylene blue degradation by a factor of 2.4, suppressing to 0.45 in the crystalline phase and enhancing to 1.09 in the amorphous phase. Importantly, this reconfigurable platform enables dynamic control of the reaction yield using a single metasurface architecture under identical illumination conditions. Our approach establishes a dynamically programmable light-driven reaction platform capable of precisely manipulating reaction reactivity, offering new opportunities for selective photocatalysis in complex multibranch reaction systems.

physics.optics

Harnessing Non-Boltzmann Steady States in Lanthanide Nanocrystals for Mid-Infrared Optoelectronics

Converting mid-infrared (MIR) radiation to visible or near-infrared wavelengths is essential for imaging and sensing, yet achieving sensitive, low-power, and scalable detection remains challenging. Lanthanide nanocrystals provide an alternative through ratiometric luminescence but are typically constrained by Boltzmann statistics, which tie population distributions to lattice temperature and limit signal contrast. Here we show that MIR irradiation rebalances dissipative relaxation pathways, driving lanthanide emitters into a non-Boltzmann steady state that enables non-thermal control of population distributions. This allows emission behaviors inaccessible under thermal equilibrium. We exploit this regime to achieve linear MIR detection with respect to MIR power across 6.8 to 8.6 micrometers. The ratiometric response is intrinsically independent of the pump power, enabling operation at an ultralow excitation power of 10 uW, several orders of magnitude lower than conventional approaches. Using standard silicon photodetectors, we then demonstrate room-temperature MIR imaging with detection limits approaching 4 nW um-2. Our results establish lanthanide nanoparticles as an efficient platform for MIR conversion and sensing in nanophotonic systems.

physics.optics

Polarization-Sensitive Au-TiO2 Nanopillars for Tailored Photocatalytic Activity

Plasmonic metasurfaces play a crucial role in resonance-driven photocatalytic reactions by effectively enhancing reactivity via localized surface plasmon resonances. Catalytic activity can be selectively modulated by tuning the strength of plasmonic resonances through two primary non-thermal mechanisms: near-field enhancement and hot carrier injection, which govern the population of energetic carrier excited or injected into unoccupied molecular orbitals. We developed a set of polarization-sensitive metasurfaces consisting of elliptical Au-TiO2 nanopillars, specifically designed to plasmonically modulate the reactivity of a model reaction: the photocatalytic degradation of methylene blue. Surface-enhanced Raman spectroscopy reveals a polarization-dependent reaction yield in real-time, modulating from 4.7 (transverse electric polarization) to 9.98 (transverse magnetic polarization) in 10 s period, as quantified by the integrated area of the 480 cm-1 Raman peak and correlated with enhanced absorption at 633 nm. The single metasurface configuration enables continuous tuning of photocatalytic reactivity via active control of plasmonic resonance strength, as evidenced by the positive correlation between measured absorption and product yield. This dynamic approach provides a route to selectively enhance or suppress resonance-driven reactions, which can be further leveraged to achieve selectivity in multibranch reactions, guiding product yields toward desired outcomes.

physics.optics

Reconfigurable miniaturized computational spectrometer enabled by photoelastic effect

Miniatured computational spectrometers, distinguished by their compact size and lightweight, have shown great promise for on-chip and portable applications in the fields of healthcare, environmental monitoring, food safety, and industrial process monitoring. However, the common miniaturization strategies predominantly rely on advanced micro-nano fabrication and complex material engineering, limiting their scalability and affordability. Here, we present a broadband miniaturized computational spectrometer (ElastoSpec) by leveraging the photoelastic effect for easy-to-prepare and reconfigurable implementations. A single computational photoelastic spectral filter, with only two polarizers and a plastic sheet, is designed to be integrated onto the top of a CMOS sensor for snapshot spectral acquisition. The different spectral modulation units are directly generated from different spatial locations of the filter, due to the photoelastic-induced chromatic polarization effect of the plastic sheet. We experimentally demonstrate that ElastoSpec offers excellent reconstruction accuracy for the measurement of both simple narrowband and complex spectra. It achieves a full width at half maximum (FWHM) error of approximately 0.2 nm for monochromatic inputs, and maintains a mean squared error (MSE) value on the order of 10^-3 with only 10 spectral modulation units. Furthermore, we develop a reconfigurable strategy for enhanced spectra sensing performance through the flexibility in optimizing the modulation effectiveness and the number of spectral modulation units. This work avoids the need for complex micro-nano fabrication and specialized materials for the design of computational spectrometers, thus paving the way for the development of simple, cost-effective, and scalable solutions for on-chip and portable spectral sensing devices.

physics.optics

OpenRR-1k: A Scalable Dataset for Real-World Reflection Removal

Reflection removal technology plays a crucial role in photography and computer vision applications. However, existing techniques are hindered by the lack of high-quality in-the-wild datasets. In this paper, we propose a novel paradigm for collecting reflection datasets from a fresh perspective. Our approach is convenient, cost-effective, and scalable, while ensuring that the collected data pairs are of high quality, perfectly aligned, and represent natural and diverse scenarios. Following this paradigm, we collect a Real-world, Diverse, and Pixel-aligned dataset (named OpenRR-1k dataset), which contains 1,000 high-quality transmission-reflection image pairs collected in the wild. Through the analysis of several reflection removal methods and benchmark evaluation experiments on our dataset, we demonstrate its effectiveness in improving robustness in challenging real-world environments. Our dataset is available at https://github.com/caijie0620/OpenRR-1k.

cs.CV

Degradation-Aware Image Enhancement via Vision-Language Classification

Image degradation is a prevalent issue in various real-world applications, affecting visual quality and downstream processing tasks. In this study, we propose a novel framework that employs a Vision-Language Model (VLM) to automatically classify degraded images into predefined categories. The VLM categorizes an input image into one of four degradation types: (A) super-resolution degradation (including noise, blur, and JPEG compression), (B) reflection artifacts, (C) motion blur, or (D) no visible degradation (high-quality image). Once classified, images assigned to categories A, B, or C undergo targeted restoration using dedicated models tailored for each specific degradation type. The final output is a restored image with improved visual quality. Experimental results demonstrate the effectiveness of our approach in accurately classifying image degradations and enhancing image quality through specialized restoration models. Our method presents a scalable and automated solution for real-world image enhancement tasks, leveraging the capabilities of VLMs in conjunction with state-of-the-art restoration techniques.

cs.CV

OpenRR-5k: A Large-Scale Benchmark for Reflection Removal in the Wild

Removing reflections is a crucial task in computer vision, with significant applications in photography and image enhancement. Nevertheless, existing methods are constrained by the absence of large-scale, high-quality, and diverse datasets. In this paper, we present a novel benchmark for Single Image Reflection Removal (SIRR). We have developed a large-scale dataset containing 5,300 high-quality, pixel-aligned image pairs, each consisting of a reflection image and its corresponding clean version. Specifically, the dataset is divided into two parts: 5,000 images are used for training, and 300 images are used for validation. Additionally, we have included 100 real-world testing images without ground truth (GT) to further evaluate the practical performance of reflection removal methods. All image pairs are precisely aligned at the pixel level to guarantee accurate supervision. The dataset encompasses a broad spectrum of real-world scenarios, featuring various lighting conditions, object types, and reflection patterns, and is segmented into training, validation, and test sets to facilitate thorough evaluation. To validate the usefulness of our dataset, we train a U-Net-based model and evaluate it using five widely-used metrics, including PSNR, SSIM, LPIPS, DISTS, and NIQE. We will release both the dataset and the code on https://github.com/caijie0620/OpenRR-5k to facilitate future research in this field.

cs.CV

F2T2-HiT: A U-Shaped FFT Transformer and Hierarchical Transformer for Reflection Removal

Single Image Reflection Removal (SIRR) technique plays a crucial role in image processing by eliminating unwanted reflections from the background. These reflections, often caused by photographs taken through glass surfaces, can significantly degrade image quality. SIRR remains a challenging problem due to the complex and varied reflections encountered in real-world scenarios. These reflections vary significantly in intensity, shapes, light sources, sizes, and coverage areas across the image, posing challenges for most existing methods to effectively handle all cases. To address these challenges, this paper introduces a U-shaped Fast Fourier Transform Transformer and Hierarchical Transformer (F2T2-HiT) architecture, an innovative Transformer-based design for SIRR. Our approach uniquely combines Fast Fourier Transform (FFT) Transformer blocks and Hierarchical Transformer blocks within a UNet framework. The FFT Transformer blocks leverage the global frequency domain information to effectively capture and separate reflection patterns, while the Hierarchical Transformer blocks utilize multi-scale feature extraction to handle reflections of varying sizes and complexities. Extensive experiments conducted on three publicly available testing datasets demonstrate state-of-the-art performance, validating the effectiveness of our approach.

cs.CV

VIP: Video Inpainting Pipeline for Real World Human Removal

Inpainting for real-world human and pedestrian removal in high-resolution video clips presents significant challenges, particularly in achieving high-quality outcomes, ensuring temporal consistency, and managing complex object interactions that involve humans, their belongings, and their shadows. In this paper, we introduce VIP (Video Inpainting Pipeline), a novel promptless video inpainting framework for real-world human removal applications. VIP enhances a state-of-the-art text-to-video model with a motion module and employs a Variational Autoencoder (VAE) for progressive denoising in the latent space. Additionally, we implement an efficient human-and-belongings segmentation for precise mask generation. Sufficient experimental results demonstrate that VIP achieves superior temporal consistency and visual fidelity across diverse real-world scenarios, surpassing state-of-the-art methods on challenging datasets. Our key contributions include the development of the VIP pipeline, a reference frame integration technique, and the Dual-Fusion Latent Segment Refinement method, all of which address the complexities of inpainting in long, high-resolution video sequences.

cs.CV

Physics-Aware Inverse Design for Nanowire Single-Photon Avalanche Detectors via Deep Learning

Single-photon avalanche detectors (SPADs) have enabled various applications in emerging photonic quantum information technologies in recent years. However, despite many efforts to improve SPAD's performance, the design of SPADs remained largely an iterative and time-consuming process where a designer makes educated guesses of a device structure based on empirical reasoning and solves the semiconductor drift-diffusion model for it. In contrast, the inverse problem, i.e., directly inferring a structure needed to achieve desired performance, which is of ultimate interest to designers, remains an unsolved problem. We propose a novel physics-aware inverse design workflow for SPADs using a deep learning model and demonstrate it with an example of finding the key parameters of semiconductor nanowires constituting the unit cell of an SPAD, given target photon detection efficiency. Our inverse design workflow is not restricted to the case demonstrated and can be applied to design conventional planar structure-based SPADs, photodetectors, and solar cells.

physics.app-ph

Survey on Single-Image Reflection Removal using Deep Learning Techniques

The phenomenon of reflection is quite common in digital images, posing significant challenges for various applications such as computer vision, photography, and image processing. Traditional methods for reflection removal often struggle to achieve clean results while maintaining high fidelity and robustness, particularly in real-world scenarios. Over the past few decades, numerous deep learning-based approaches for reflection removal have emerged, yielding impressive results. In this survey, we conduct a comprehensive review of the current literature by focusing on key venues such as ICCV, ECCV, CVPR, NeurIPS, etc., as these conferences and journals have been central to advances in the field. Our review follows a structured paper selection process, and we critically assess both single-stage and two-stage deep learning methods for reflection removal. The contribution of this survey is three-fold: first, we provide a comprehensive summary of the most recent work on single-image reflection removal; second, we outline task hypotheses, current deep learning techniques, publicly available datasets, and relevant evaluation metrics; and third, we identify key challenges and opportunities in deep learning-based reflection removal, highlighting the potential of this rapidly evolving research area.

cs.CV

Liquid Metal-Exfoliated SnO$_2$-Based Mixed-dimensional Heterostructures for Visible-to-Near-Infrared Photodetection

Ultra-thin two-dimensional (2D) materials have gained significant attention for making next-generation optoelectronic devices. Here, we report a large-area heterojunction photodetector fabricated using a liquid metal-printed 2D $\text{SnO}_2$ layer transferred onto CdTe thin films. The resulting device demonstrates efficient broadband light sensing from visible to near-infrared wavelengths, with enhanced detectivity and faster photo response than bare CdTe photodetectors. Significantly, the device shows a nearly $10^5$-fold increase in current than the dark current level when illuminated with a 780 nm laser and achieves a specific detectivity of around $10^{12} \, \text{Jones}$, nearly two orders of magnitude higher than a device with pure CdTe thin film. Additionally, temperature-dependent optoelectronic testing shows that the device maintains a stable response up to $140^\circ \text{C}$ and generates distinctive photocurrent at temperatures up to $80^\circ \text{C}$, demonstrating its thermal stability. Using band structure analysis, density functional theory (DFT) calculations, and photocurrent mapping, the formation of a $p$-$n$ junction is indicated, contributing to the enhanced photo response attributed to the efficient carrier separation by the built-in potential in the hetero-junction and the superior electron mobility of 2D $\text{SnO}_2$. Our results highlight the effectiveness of integrating liquid metal-exfoliated 2D materials for enhanced photodetector performance.

cond-mat.mtrl-sci

Nanowire Array Breath Acetone Sensor for Diabetes Monitoring

Diabetic ketoacidosis (DKA) is a life-threatening acute complication of diabetes in which ketone bodies accumulate in the blood. Breath acetone (a ketone) directly correlates with blood ketones, such that breath acetone monitoring could be used to improve safety in diabetes care. In this work, we report the design and fabrication of a chitosan/Pt/InP nanowire array based chemiresistive acetone sensor. By implementing chitosan as a surface functionalization layer and a Pt Schottky contact for efficient charge transfer processes and photovoltaic effect, self-powered, highly selective acetone sensing has been achieved. This sensor has an ultra-wide detection range from sub-ppb to >100,000 ppm levels at room temperature, incorporating the range from healthy individuals (300-800 ppb) to those at high-risk of DKA (> 75 ppm). The nanowire sensor has been further integrated into a handheld breath testing prototype, the Ketowhistle, which can successfully detect different ranges of acetone concentrations in simulated breath. The Ketowhistle demonstrates immediate potential for non-invasive ketone testing and monitoring for persons living with diabetes, in particular for DKA prevention.

physics.med-ph

VehicleGAN: Pair-flexible Pose Guided Image Synthesis for Vehicle Re-identification

Vehicle Re-identification (Re-ID) has been broadly studied in the last decade; however, the different camera view angle leading to confused discrimination in the feature subspace for the vehicles of various poses, is still challenging for the Vehicle Re-ID models in the real world. To promote the Vehicle Re-ID models, this paper proposes to synthesize a large number of vehicle images in the target pose, whose idea is to project the vehicles of diverse poses into the unified target pose so as to enhance feature discrimination. Considering that the paired data of the same vehicles in different traffic surveillance cameras might be not available in the real world, we propose the first Pair-flexible Pose Guided Image Synthesis method for Vehicle Re-ID, named as VehicleGAN in this paper, which works for both supervised and unsupervised settings without the knowledge of geometric 3D models. Because of the feature distribution difference between real and synthetic data, simply training a traditional metric learning based Re-ID model with data-level fusion (i.e., data augmentation) is not satisfactory, therefore we propose a new Joint Metric Learning (JML) via effective feature-level fusion from both real and synthetic data. Intensive experimental results on the public VeRi-776 and VehicleID datasets prove the accuracy and effectiveness of our proposed VehicleGAN and JML.

cs.CV

An efficient modeling workflow for high-performance nanowire single-photon avalanche detector

Single-photon detector (SPD), an essential building block of the quantum communication system, plays a fundamental role in developing next-generation quantum technologies. In this work, we propose an efficient modeling workflow of nanowire SPDs utilizing avalanche breakdown at reverse-biased conditions. The proposed workflow is explored to maximize computational efficiency and balance time-consuming drift-diffusion simulation with fast script-based post-processing. Without excessive computational effort, we could predict a suite of key device performance metrics, including breakdown voltage, dark/light avalanche built-up time, photon detection efficiency, dark count rate, and the deterministic part of timing jitter due to device structures. Implementing the proposed workflow onto a single InP nanowire and comparing it to the extensively studied planar devices and superconducting nanowire SPDs, we showed the great potential of nanowire avalanche SPD to outperform their planar counterparts and obtain as superior performance as superconducting nanowires, i.e., achieve a high photon detection efficiency of 70% with a dark count rate less than 20 Hz at non-cryogenic temperature. The proposed workflow is not limited to single-nanowire or nanowire-based device modeling and can be readily extended to more complicated two-/three dimensional structures.

physics.app-ph

Defense against Adversarial Cloud Attack on Remote Sensing Salient Object Detection

Detecting the salient objects in a remote sensing image has wide applications for the interdisciplinary research. Many existing deep learning methods have been proposed for Salient Object Detection (SOD) in remote sensing images and get remarkable results. However, the recent adversarial attack examples, generated by changing a few pixel values on the original remote sensing image, could result in a collapse for the well-trained deep learning based SOD model. Different with existing methods adding perturbation to original images, we propose to jointly tune adversarial exposure and additive perturbation for attack and constrain image close to cloudy image as Adversarial Cloud. Cloud is natural and common in remote sensing images, however, camouflaging cloud based adversarial attack and defense for remote sensing images are not well studied before. Furthermore, we design DefenseNet as a learn-able pre-processing to the adversarial cloudy images so as to preserve the performance of the deep learning based remote sensing SOD model, without tuning the already deployed deep SOD model. By considering both regular and generalized adversarial examples, the proposed DefenseNet can defend the proposed Adversarial Cloud in white-box setting and other attack methods in black-box setting. Experimental results on a synthesized benchmark from the public remote sensing SOD dataset (EORSSD) show the promising defense against adversarial cloud attacks.

cs.CV