SearcharxivSearch

arXiv subjects

Qiushi Guo

Publications and source records attributed to Qiushi Guo.

At least 19 recordsLinked to original sources

Silicon avalanche transition edge bolometer: approaching the thermodynamic limit for uncooled long-wave infrared detection

Bolometers transduce incident electromagnetic radiation into measurable electrical signals via radiation-induced heating in thermo-resistive materials. They are uniquely capable of detecting low-energy photons without cryogenic cooling, and are widely deployed for uncooled long-wave infrared (LWIR) radiation detection and thermal imaging. Although their fundamental detection limit is set by the thermodynamic fluctuations, state-of-the-art uncooled bolometers still operate well above this limit due to the presence of Johnson noise, 1/f noise and noises from the readout circuits. Here, we address this challenge by introducing a new uncooled LWIR bolometer concept---the silicon avalanche transition edge (SATE) bolometer. Operating near the steep current transition edge associated with avalanche breakdown, the device exhibits an ultra-high, positive temperature coefficient of resistance (TCR) of 330 %/K, greatly suppressing the impacts of other noise sources. Even without any thermal insulation structures, the SATE bolometer delivers a high room-temperature responsivity up to 160 mA/W for 9.5 $μ$m radiation, a noise equivalent power of 370 pW/$\sqrt{\mathrm{Hz}}$, and a strong electro-thermal feedback enlarged bandwidth of 77 kHz---performance unattainable with conventional bolometer materials with a TCR of $-1\sim-3\%$/K. Our work establishes a promising thermo-electric transduction mechanism toward high-sensitivity, high-speed room-temperature thermal imaging and infrared spectroscopy using CMOS technology.

physics.optics

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation

Normalizing Flows (NFs) are powerful generative models capable of exact density estimation and sampling. However, their strict invertibility often forces the model to exhaust its capacity on low-level pixel details, hindering the capture of high-level semantic structures. While Masked Image Modeling (MIM) has excelled in representation learning, its integration into generative pipelines has remained largely modular and disjointed. In this paper, we propose MIMFlow, a unified end-to-end framework that jointly optimizes latent semantics, pixel reconstruction, and generative flow. By employing a VAE encoder to infer semantic latent from masked images, MIMFlow achieves a principled decoupling of the generative task: the Normalizing Flow focuses on modeling a simplified, low-frequency semantic manifold, while a specialized decoder handles high-frequency synthesis. This design effectively resolves the inherent capacity bottleneck of NFs, allowing the model to prioritize global structural coherence over redundant noise. Empirical results on ImageNet 256$\times$256 show that MIMFlow-L reaches 71.3\% linear probing accuracy and an FID of 2.50. Despite using only 128 tokens (50\% fewer than standard models), it yields a 32.8\% performance gain over similar-scale NF baselines. Our code is available at https://github.com/MCG-NJU/MIMFlow.

cs.CV

From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation

Subject-driven image generation models face a fundamental trade-off between identity preservation (fidelity) and prompt adherence (editability). While online reinforcement learning (RL), specifically GPRO, offers a promising solution, we find that a naive application of GRPO leads to competitive degradation, as the simple linear aggregation of rewards with static weights causes conflicting gradient signals and a misalignment with the temporal dynamics of the diffusion process. To overcome these limitations, we propose Customized-GRPO, a novel framework featuring two key innovations: (i) Synergy-Aware Reward Shaping (SARS), a non-linear mechanism that explicitly penalizes conflicted reward signals and amplifies synergistic ones, providing a sharper and more decisive gradient. (ii) Time-Aware Dynamic Weighting (TDW), which aligns the optimization pressure with the model's temporal dynamics by prioritizing prompt-following in the early, identity preservation in the later. Extensive experiments demonstrate that our method significantly outperforms naive GRPO baselines, successfully mitigating competitive degradation. Our model achieves a superior balance, generating images that both preserve key identity features and accurately adhere to complex textual prompts.

cs.LG

Lithium niobate quadratic integrated nonlinear photonics: enabling ultra-wide bandwidth and ultrafast photonic engines

Integrated photonic coherent light sources capable of generating emission with broad spectral coverage and ultrashort pulse durations are critical for both fundamental science and emerging technologies. In this Perspective, we start by discussing emerging quantum and classical photonic applications from the standpoint of operating wavelength and timescale, highlighting the technological gaps that persist in current integrated photonic light sources. Next, we introduce the unique properties of lithium niobate-based integrated quadratic nonlinear photonics, and discuss several promising strategies that exploit this platform to realize wavelength-tunable continuous wave light sources and broadband, ultra-short light pulse generation. We also assessed their advantages and limitations while discussing potential solutions. Finally, we outline future prospects and challenges that need to be addressed, aiming at inspiring continued research and innovation in this rapidly evolving field.

physics.optics

Calibration-Free Induced Magnetic Field Indoor and Outdoor Positioning via Data-Driven Modeling

Induced magnetic field (IMF)-based localization offers a robust alternative to wave-based positioning technologies due to its resilience to non-line-of-sight conditions, environmental dynamics, and wireless interference. However, existing magnetic localization systems typically rely on analytical field inversion, manual calibration, or environment-specific fingerprinting, limiting their scalability and transferability. This paper presents a data-driven IMF localization framework that directly maps induced magnetic field measurements to spatial coordinates using supervised learning, eliminating explicit environment-specific calibration. By replacing explicit field modeling with learning-based inference, the proposed approach captures nonlinear field interactions and environmental effects. An orientation-invariant feature representation enables rotation-independent deployment. The system is evaluated across multiple indoor environments and an outdoor deployment. Benchmarking against classical and deep learning baselines shows that a Random Forest regressor achieves sub-20 cm accuracy in 2D and sub-30 cm in 3D localization. Cross-environment validation demonstrates that models trained indoors generalize to outdoor environments without retraining. We further analyze scalability by varying transmitter spacing, showing that coverage and accuracy can be balanced through deployment density. Overall, this work demonstrates that data-driven IMF localization is a scalable and transferable solution for real-world positioning.

eess.SP

On-chip electrically reconfigurable octave-bandwidth optical amplification from visible to near-infrared

Achieving broadband on-chip optical amplification spanning the visible and near-infrared (NIR) can enable diverse quantum sensing, metrology, and classical communication applications within a single unified device. However, conventional semiconductor and ion-doped amplifiers suffer from limited gain bandwidths set by fixed energy levels, while optical parametric amplifiers (OPAs) operating continuously from the visible to the NIR have remained elusive due to dispersion-limited bandwidth and the high pump powers required in the visible or ultraviolet (UV). Here, we overcome these limitations by introducing an electrically reconfigurable OPA architecture on lithium niobate integrated photonics. By synergistically combining ultra-high effective $χ^{(2)}$ nonlinearity ($\sim$7,000\%/W-cm$^2$), high-order dispersion engineering, and local electro-thermal tuning of quasi-phase matching, our device achieves record gain spectral spanning more than an optical octave, from 770 to 1650 nm. This range covers key transitions of many photonic quantum systems and all telecommunication bands. Moreover, our approach eliminates the need for high-power, wavelength-tunable visible or UV pumps, delivering a peak on-chip gain of 23.67 dB with a single 1060 nm pump at 90 mW average on-chip power. This work opens new avenues for multi-functional, reconfigurable photonics unifying the visible and infrared regimes, with broad implications for quantum sensing and communications.

physics.optics

Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection

Data augmentation is crucial for improving the robustness of face detection systems, especially under challenging conditions such as occlusion, illumination variation, and complex environments. Traditional copy paste augmentation often produces unrealistic composites due to inaccurate foreground extraction, inconsistent scene geometry, and mismatched background semantics. To address these limitations, we propose Depth Copy Paste, a multimodal and depth aware augmentation framework that generates diverse and physically consistent face detection training samples by copying full body person instances and pasting them into semantically compatible scenes. Our approach first employs BLIP and CLIP to jointly assess semantic and visual coherence, enabling automatic retrieval of the most suitable background images for the given foreground person. To ensure high quality foreground masks that preserve facial details, we integrate SAM3 for precise segmentation and Depth-Anything to extract only the non occluded visible person regions, preventing corrupted facial textures from being used in augmentation. For geometric realism, we introduce a depth guided sliding window placement mechanism that searches over the background depth map to identify paste locations with optimal depth continuity and scale alignment. The resulting composites exhibit natural depth relationships and improved visual plausibility. Extensive experiments show that Depth Copy Paste provides more diverse and realistic training data, leading to significant performance improvements in downstream face detection tasks compared with traditional copy paste and depth free augmentation methods.

cs.CV

Passive harmonic mode-locked laser on lithium niobate integrated photonics

Mode-locked lasers (MLLs) are essential for a wide range of photonic applications, such as frequency metrology, biological imaging, and high-bandwidth coherent communications. The growing demand for compact and scalable photonic systems is driving the development of MLLs on various integrated photonics material platforms. Along these lines, developing MLLs on the emerging thin-film lithium niobate (TFLN) platform holds the promise to greatly broaden the application space of MLLs by harnessing TFLN 's unique electro-optic (E-O) response and quadratic optical nonlinearity. Here, we demonstrate the first electrically pumped, self-starting passive MLL in lithium niobate integrated photonics based on its hybrid integration with a GaAs quantum-well gain medium and saturable absorber. Our demonstrated MLL generates 4.3-ps optical pulses centered around 1060 nm with on-chip peak power exceeding 44 mW. The pulse duration can be further compressed to 1.75 ps via linear dispersion compensation. Remarkably, passive mode-locking occurs exclusively at the second harmonic of the cavity free spectral range, exhibiting a high pulse repetition rate $\sim$20 GHz. We elucidate the temporal dynamics underlying this self-starting passive harmonic mode-locking behavior using a traveling-wave model. Our work offers new insights into the realization of compact, high-repetition-rate MLLs in the TFLN platform, with promising applications for monolithic ultrafast microwave waveform sampling and analog-to-digital conversion.

physics.optics

Enrich the content of the image Using Context-Aware Copy Paste

Data augmentation remains a widely utilized technique in deep learning, particularly in tasks such as image classification, semantic segmentation, and object detection. Among them, Copy-Paste is a simple yet effective method and gain great attention recently. However, existing Copy-Paste often overlook contextual relevance between source and target images, resulting in inconsistencies in generated outputs. To address this challenge, we propose a context-aware approach that integrates Bidirectional Latent Information Propagation (BLIP) for content extraction from source images. By matching extracted content information with category information, our method ensures cohesive integration of target objects using Segment Anything Model (SAM) and You Only Look Once (YOLO). This approach eliminates the need for manual annotation, offering an automated and user-friendly solution. Experimental evaluations across diverse datasets demonstrate the effectiveness of our method in enhancing data diversity and generating high-quality pseudo-images across various computer vision tasks.

cs.CV

SynRailObs: A Synthetic Dataset for Obstacle Detection in Railway Scenarios

Detecting potential obstacles in railway environments is critical for preventing serious accidents. Identifying a broad range of obstacle categories under complex conditions requires large-scale datasets with precisely annotated, high-quality images. However, existing publicly available datasets fail to meet these requirements, thereby hindering progress in railway safety research. To address this gap, we introduce SynRailObs, a high-fidelity synthetic dataset designed to represent a diverse range of weather conditions and geographical features. Furthermore, diffusion models are employed to generate rare and difficult-to-capture obstacles that are typically challenging to obtain in real-world scenarios. To evaluate the effectiveness of SynRailObs, we perform experiments in real-world railway environments, testing on both ballasted and ballastless tracks across various weather conditions. The results demonstrate that SynRailObs holds substantial potential for advancing obstacle detection in railway safety applications. Models trained on this dataset show consistent performance across different distances and environmental conditions. Moreover, the model trained on SynRailObs exhibits zero-shot capabilities, which are essential for applications in security-sensitive domains. The data is available in https://www.kaggle.com/datasets/qiushi910/synrailobs.

cs.CV

Electrically Reconfigurable Intelligent Optoelectronics in 2-D van der Waals Materials

In optoelectronics, achieving electrical reconfigurability is crucial as it enables the encoding, decoding, manipulating, and processing of information carried by light. In recent years, two-dimensional van der Waals (2-D vdW) materials have emerged as promising platforms for realizing reconfigurable optoelectronic devices. Compared to materials with bulk crystalline lattice, 2-D vdW materials offer superior electrical reconfigurability due to high surface-to-volume ratio, quantum confinement, reduced dielectric screening effect, and strong dipole resonances. Additionally, their unique band structures and associated topology and quantum geometry provide novel tuning capabilities. This review article seeks to establish a connection between the fundamental physics underlying reconfigurable optoelectronics in 2-D materials and their burgeoning applications in intelligent optoelectronics. We first survey various electrically reconfigurable properties of 2-D vdW materials and the underlying tuning mechanisms. Then we highlight the emerging applications of such devices, including dynamic intensity, phase and polarization control, and intelligent sensing. Finally, we discuss the opportunities for future advancements in this field.

physics.optics

A Universal Railway Obstacle Detection System based on Semi-supervised Segmentation And Optical Flow

Detecting obstacles in railway scenarios is both crucial and challenging due to the wide range of obstacle categories and varying ambient conditions such as weather and light. Given the impossibility of encompassing all obstacle categories during the training stage, we address this out-of-distribution (OOD) issue with a semi-supervised segmentation approach guided by optical flow clues. We reformulate the task as a binary segmentation problem instead of the traditional object detection approach. To mitigate data shortages, we generate highly realistic synthetic images using Segment Anything (SAM) and YOLO, eliminating the need for manual annotation to produce abundant pixel-level annotations. Additionally, we leverage optical flow as prior knowledge to train the model effectively. Several experiments are conducted, demonstrating the feasibility and effectiveness of our approach.

cs.CV

Hyperbolic phonon-polariton electroluminescence in graphene-hBN van der Waals heterostructures

Phonon-polaritons are electromagnetic waves resulting from the coherent coupling of photons with optical phonons in polar dielectrics. Due to their exceptional ability to confine electric fields to deep subwavelength scales with low loss, they are uniquely poised to enable a suite of applications beyond the reach of conventional photonics, such as sub-diffraction imaging and near-field energy transfer. The conventional approach to exciting phonon-polaritons through optical methods, however, necessitates costly mid-infrared and terahertz coherent light sources along with near-field scanning probes, and generally leads to low excitation efficiency due to the substantial momentum mismatch between phonon-polaritons and free-space photons. Here, we demonstrate that under proper conditions, phonon-polaritons can be excited all-electrically by flowing charge carriers. Specifically, in hexagonal boron nitride (hBN)/graphene heterostructures, by electrically driving charge carriers in ultra-high-mobility graphene out of equilibrium, we observe bright electroluminescence of hBN's hyperbolic phonon-polaritons (HPhPs) at mid-IR frequencies. The HPhP electroluminescence shows a temperature and carrier density dependence distinct from black-body or super-Planckian thermal emission. Moreover, the carrier density dependence of HPhP electroluminescence spectra reveals that HPhP electroluminescence can arise from both inter-band transition and intra-band Cherenkov radiation of charge carriers in graphene. The HPhP electroluminescence offers fundamentally new avenues for realizing electrically-pumped, tunable mid-IR and THz phonon-polariton lasers, and efficient cooling of electronic devices.

cond-mat.mes-hall

Multi-Octave Frequency Comb from an Ultra-Low-Threshold Nanophotonic Parametric Oscillator

Ultrabroadband frequency combs coherently unite distant portions of the electromagnetic spectrum. They underpin discoveries in ultrafast science and serve as the building blocks of modern photonic technologies. Despite tremendous progress in integrated sources of frequency combs, achieving multi-octave operation on chip has remained elusive mainly because of the energy demand of typical spectral broadening processes. Here we break this barrier and demonstrate multi-octave frequency comb generation using an optical parametric oscillator (OPO) in nanophotonic lithium niobate with only femtojoules of pump energy. The energy-efficient and robust coherent spectral broadening occurs far above the oscillation threshold of the OPO and detuned from its linear synchrony with the pump. We show that the OPO can undergo a temporal self-cleaning mechanism by transitioning from an incoherent operation regime, which is typical for operation far above threshold, to an ultrabroad coherent regime, corresponding to the nonlinear phase compensating the OPO cavity detuning. Such a temporal self-cleaning mechanism and the subsequent multi-octave coherent spectrum has not been explored in previous OPO designs and features a relaxed requirement for the quality factor and relatively narrow spectral coverage of the cavity. We achieve orders of magnitude reduction in the energy requirement compared to the other techniques, confirm the coherence of the comb, and present a path towards more efficient and wider spectral broadening. Our results pave the way for ultrashort-pulse and ultrabroadband on-chip nonlinear photonic systems for numerous applications.

physics.optics

FaceSkin: A Privacy Preserving Facial skin patch Dataset for multi Attributes classification

Human facial skin images contain abundant textural information that can serve as valuable features for attribute classification, such as age, race, and gender. Additionally, facial skin images offer the advantages of easy collection and minimal privacy concerns. However, the availability of well-labeled human skin datasets with a sufficient number of images is limited. To address this issue, we introduce a dataset called FaceSkin, which encompasses a diverse range of ages and races. Furthermore, to broaden the application scenarios, we incorporate synthetic skin-patches obtained from 2D and 3D attack images, including printed paper, replays, and 3D masks. We evaluate the FaceSkin dataset across distinct categories and present experimental results demonstrating its effectiveness in attribute classification, as well as its potential for various downstream tasks, such as Face anti-spoofing and Age estimation.

cs.CV

Enhancing Mobile Privacy and Security: A Face Skin Patch-Based Anti-Spoofing Approach

As Facial Recognition System(FRS) is widely applied in areas such as access control and mobile payments due to its convenience and high accuracy. The security of facial recognition is also highly regarded. The Face anti-spoofing system(FAS) for face recognition is an important component used to enhance the security of face recognition systems. Traditional FAS used images containing identity information to detect spoofing traces, however there is a risk of privacy leakage during the transmission and storage of these images. Besides, the encryption and decryption of these privacy-sensitive data takes too long compared to inference time by FAS model. To address the above issues, we propose a face anti-spoofing algorithm based on facial skin patches leveraging pure facial skin patch images as input, which contain no privacy information, no encryption or decryption is needed for these images. We conduct experiments on several public datasets, the results prove that our algorithm has demonstrated superiority in both accuracy and speed.

cs.CV

Mode-locked laser in nanophotonic lithium niobate

Mode-locked lasers (MLLs) have enabled ultrafast sciences and technologies by generating ultrashort pulses with peak powers substantially exceeding their average powers. Recently, tremendous efforts have been focused on realizing integrated MLLs not only to address the challenges associated with their size and power demand, but also to enable transforming the ultrafast technologies into nanophotonic chips, and ultimately to unlock their potential for a plethora of applications. However, till now the prospect of integrated MLLs driving ultrafast nanophotonic circuits has remained elusive because of their typically low peak powers, lack of controllability, and challenges with integration with appropriate nanophotonic platforms. Here, we overcome these limitations by demonstrating an electrically-pumped actively MLL in nanophotonic lithium niobate based on its hybrid integration with a III-V semiconductor optical amplifier. Our MLL generates $\sim$4.8 ps optical pulses around 1065 nm at a repetition rate of $\sim$10 GHz, with pulse energy exceeding 2.6 pJ and a high peak power beyond 0.5 W. We show that both the repetition rate and the carrier-envelope-offset of the resulting frequency comb can be flexibly controlled in a wide range using the RF driving frequency and the pump current, paving the way for fully-stabilized on-chip frequency combs in nanophotonics. Our work marks an important step toward fully-integrated nonlinear and ultrafast photonic systems in nanophotonic lithium niobate.

physics.optics

Visible to Ultraviolet Frequency Comb Generation in Lithium Niobate Nanophotonic Waveguides

The introduction of nonlinear nanophotonic devices to the field of optical frequency comb metrology has enabled new opportunities for low-power and chip-integrated clocks, high-precision frequency synthesis, and broad bandwidth spectroscopy. However, most of these advances remain constrained to the near-infrared region of the spectrum, which has restricted the integration of frequency combs with numerous quantum and atomic systems in the ultraviolet and visible. Here, we overcome this shortcoming with the introduction of multi-segment nanophotonic thin-film lithium niobate (LN) waveguides that combine engineered dispersion and chirped quasi-phase matching for efficient supercontinuum generation via the combination of $χ^{(2)}$ and $χ^{(3)}$ nonlinearities. With only 90 pJ of pulse energy at 1550 nm, we achieve gap-free frequency comb coverage spanning 330 to 2400 nm. The conversion efficiency from the near-infrared pump to the UV-Visible region of 350-550 nm is nearly 20%. Harmonic generation via the $χ^{(2)}$ nonlinearity in the same waveguide directly yields the carrier-envelope offset frequency and a means to verify the comb coherence at wavelengths as short as 350 nm. Our results provide an integrated photonics approach to create visible and UV frequency combs that will impact precision spectroscopy, quantum information processing, and optical clock applications in this important spectral window.

physics.optics