SearcharxivSearch

arXiv subjects

Guanghui Zhao

Publications and source records attributed to Guanghui Zhao.

12 recordsLinked to original sources

Dual-comb generated in single thin-film lithium niobate microrings

Dual-comb technology has emerged as an essential tool for high-precision spectroscopy, real-time ranging, and high-sensitivity sensing. Integrating dual-comb sources into a single microresonator would substantially reduce pump power, footprint, system complexity, and cost, yet this remains a significant challenge. Here, we demonstrate, for the first time, integrated dual-comb generation in a single thin-film lithium niobate (TFLN) microring, under single continuous-wave laser pumping. Rather than regarding TFLN's strong Raman nonlinearity as detrimental, as conventionally viewed, we harness it constructively. By engineering the dispersion of TFLN microrings, we leverage the fundamental and first-order transverse-electric mode families with loaded Q factors exceeding 5X10^6, comparable repetition rates, and suitable dispersion profiles, and bridge them through stimulated Raman scattering (SRS) processes. Pumping a first-order mode at 1551.28 nm initially excites both Stokes and anti-Stokes SRS in the fundamental mode family at low thresholds, and subsequently produces two independent, spectrally separated combs at a pump power of 320 mW via direct Kerr and Raman-assisted Kerr effects, respectively. The two combs span broad bandwidths, exhibit repetition rates of ~102 GHz with a slight difference of ~624 MHz, and do not merge spectrally. The broadest spectrum spans 654 nm, and the Raman-Kerr comb has a 3-dB bandwidth exceeding 29 nm. Further characterization confirms that the comb lines exhibit low phase noise, with an intrinsic linewidth of 410 Hz. This work establishes a robust pathway for on-chip dual-comb generation in a single-laser pumped microring, significantly advancing dual-comb systems toward simplified architectures, enhanced robustness, and scalable integration, while accelerating their practical deployment.

physics.optics

CATVis: A Collaborative Multi-Agent Workflow for Turbomachinery Simulation Data Visualization

Recent advances in AI for Science have enabled natural language (NL) interfaces for scientific data analysis. In turbomachinery CFD post-processing, translating ambiguous high-level analytical goals (e.g., vortex identification) into precise visualization procedures supporting complex domain-specific analysis is challenging. We present CATVis, a Collaborative multi-agent workflow system that bridges this gap by transforming NL intents into structured middle representation for visualization. Our approach reformulates domain-specific visualization procedures as composable workflow representations, and use multi agent to generate workflow representations via intent planning, template generation, and error-aware refinement, where each stage incrementally updates a shared structured representation. We evaluate the impact of external knowledge and workflow structuring on generation accuracy, demonstrating that the proposed approach significantly improves complex workflow generation correctness while reducing prompt complexity.

cs.HC

Isotropic fabrication of centimeter-scale, low propagation-loss periodically poled lithium niobate nanophotonic waveguides for efficient second harmonic generation

Periodically poled lithium niobate (PPLN) nanophotonic waveguides that simultaneously feature low propagation-loss and uniform periodic poling are essential for a wide range of applications ranging from classical nonlinear frequency-conversion to scalable integrated quantum technology. However, fabrication imperfections have frequently limited the propagation loss of fully domain-inverted PPLN nanophotonic waveguides to a few dB/cm, primarily due to anisotropic etching issue, thereby restricting the absolute conversion efficiency and scale of photonic integration. Here, we present a fabrication approach that overcomes this challenge, yielding a 1.2-cm-long PPLN nanophotonic waveguide with low propagation loss via femtosecond-laser photolithography-assisted chemo-mechanical etching (PLACE). By carrying out domain inversion on a planar thin-film prior to waveguide definition, electric-field distortion is minimized during poling, while isotropic etching of the waveguide is achieved by PLACE with an average surface roughness of only 0.34 nm, resulting in uniform poling of duty cycle of 50% and a record-low propagation loss of 0.042 dB/cm in the telecom band. Under continuous-wave pumping at 1525 nm, the device demonstrates a high normalized quasi-phase-matched SHG conversion efficiency of 2021%/W, and an absolute conversion efficiency of 64% at a pump power of 86 mW which represents the state of the art for single-period PPLN nanophotonic waveguides.

physics.optics

Toward Reliable Scientific Visualization Pipeline Construction with Structure-Aware Retrieval-Augmented LLMs

Scientific visualization pipelines encode domain-specific procedural knowledge with strict execution dependencies, making their construction sensitive to missing stages, incorrect operator usage, or improper ordering. Thus, generating executable scientific visualization pipelines from natural-language descriptions remains challenging for large language models, particularly in web-based environments where visualization authoring relies on explicit code-level pipeline assembly. In this work, we investigate the reliability of LLM-based scientific visualization pipeline generation, focusing on vtk.js as a representative web-based visualization library. We propose a structure-aware retrieval-augmented generation workflow that provides pipeline-aligned vtk.js code examples as contextual guidance, supporting correct module selection, parameter configuration, and execution order. We evaluate the proposed workflow across multiple multi-stage scientific visualization tasks and LLMs, measuring reliability in terms of pipeline executability and human correction effort. To this end, we introduce correction cost as metric for the amount of manual intervention required to obtain a valid pipeline. Our results show that structured, domain-specific context substantially improves pipeline executability and reduces correction cost. We additionally provide an interactive analysis interface to support human-in-the-loop inspection and systematic evaluation of generated visualization pipelines.

cs.GR

EventFlash: Towards Efficient MLLMs for Event-Based Vision

Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, current event-based MLLMs often rely on dense image-like processing paradigms, overlooking the spatiotemporal sparsity of event streams and resulting in high computational cost. In this paper, we propose EventFlash, a novel and efficient MLLM to explore spatiotemporal token sparsification for reducing data redundancy and accelerating inference. Technically, we build EventMind, a large-scale and scene-diverse dataset with over 500k instruction sets, providing both short and long event stream sequences to support our curriculum training strategy. We then present an adaptive temporal window aggregation module for efficient temporal sampling, which adaptively compresses temporal tokens while retaining key temporal cues. Finally, a sparse density-guided attention module is designed to improve spatial token efficiency by selecting informative regions and suppressing empty or sparse areas. Experimental results show that EventFlash achieves a $12.4\times$ throughput improvement over the baseline (EventFlash-Zero) while maintaining comparable performance. It supports long-range event stream processing with up to 1,000 bins, significantly outperforming the 5-bin limit of EventGPT. We believe EventFlash serves as an efficient foundation model for event-based vision.

cs.CV

Simultaneous generation of Raman-assisted Soliton Microcombs and Tunable Multi-chromatic Raman Microlasers in Single Monolithic Thin-film Lithium Niobate Microrings

High-performance integrated broadband coherent light sources are essential for advanced applications in high-bandwidth data processing and chip-scale metrology, yet remain challenging. In this study, we demonstrate a monolithic Z-cut lithium niobate on insulator (LNOI) microring platform that enables simultaneous generation of tunable multi-chromatic microlasers and Raman-assisted soliton microcombs. Exploiting the strong Raman activity and high second-order nonlinearity of LNOI, we engineered a dispersion-optimized microring with a loaded Q factor of 3.86X10^6, facilitating on-chip efficient broadband coherent light source. A novel phase-matching configuration with all the waves of the same ordinary polarization was realized for the first time in this platform, feasibly enabling modal-phase matched Raman-quadratic nonlinear processes that extend lasing signals into the visible spectrum. Under continuous-wave laser pumping at 3.73 mW in the telecom band, we achieved a Raman-assisted soliton comb centered at 1624.49 nm with record-low pump threshold on the LNOI platform. Concurrently, multi-chromatic Raman lasing outputs were observed at ~1700, ~813, and ~535 nm within the same microring. The system exhibited efficient wavelength tuning of these multi-chromatic laser signals through a 5 nm shift in pump wavelength. This work represents a significant advance in integrated photonics for versatile optical signal generation.

physics.optics

EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs

Multimodal large language models (MLLMs) have made significant advancements in event-based vision, yet the comprehensive evaluation of their capabilities within a unified benchmark remains largely unexplored. In this work, we introduce EventBench, a benchmark that offers eight diverse task metrics together with a large-scale event stream dataset. EventBench differs from existing event-based benchmarks in four key aspects: (1) openness in accessibility, releasing all raw event streams and task instructions across eight evaluation metrics; (2) diversity in task coverage, spanning understanding, recognition, and spatial reasoning tasks for comprehensive capability assessment; (3) integration in spatial dimensions, pioneering the design of 3D spatial reasoning tasks for event-based MLLMs; and (4) scale in data volume, with an accompanying training set of over one million event-text pairs supporting large-scale training and evaluation. Using EventBench, we evaluate state-of-the-art closed-source models such as GPT-5 and Gemini-2.5 Pro, leading open-source models including Qwen2.5-VL and InternVL3, and event-based MLLMs such as EventGPT that directly process raw event streams. Extensive evaluation reveals that while current event-based MLLMs demonstrate strong performance in event stream understanding, they continue to struggle with fine-grained recognition and spatial reasoning.

cs.CV

Low-loss thin-film periodically poled lithium niobate waveguides fabricated by femtosecond laser photolithography

Periodically poled lithium niobate on insulator (PPLNOI) ridge waveguides are critical photonic components for both classical and quantum information processing. However, dry etching of PPLNOI waveguides often generates rough sidewalls and variations in the etching rates of oppositely poled lithium niobate ferroelectric domains, leading a relatively high propagation losses (0.25 - 1 dB/cm), which significantly limits net conversion efficiency and hinders scalable photonic integration. In this work, a low-loss PPLNOI ridge waveguide with a length of 7 mm was fabricated using ultra-smooth sidewalls through photolithography-assisted chemo-mechanical etching (PLACE) followed by high-voltage pulse poling with low cost. The average surface roughness was measured at just 0.27 nm, resulting in record-low propagation loss of 0.106 dB/cm in PPLNOI waveguides. Highly efficient second-harmonic generation was demonstrated with a normalized efficiency of 1643%/(W*cm^2) without temperature tuning, corresponding to a conversion efficiency of 805%/W, which is closed to the best conversion efficiency (i.e., 814%/W) reported in nanophotonic PPLNOI waveguide fabricated by expensive electron-beam lithography followed by dry etching. The absolute conversion efficiency reached 15.8% at a pump level of 21.6 mW. And the normalized efficiency can be even improved to 1742%/(W*cm^2) at optimal temperature of 59°C.

physics.optics

EventGPT: Event Stream Understanding with Multimodal Large Language Models

Event cameras record visual information as asynchronous pixel change streams, excelling at scene perception under unsatisfactory lighting or high-dynamic conditions. Existing multimodal large language models (MLLMs) concentrate on natural RGB images, failing in scenarios where event data fits better. In this paper, we introduce EventGPT, the first MLLM for event stream understanding, to the best of our knowledge, marking a pioneering attempt to integrate large language models (LLMs) with event stream comprehension. To mitigate the huge domain gaps, we develop a three-stage optimization paradigm to gradually equip a pre-trained LLM with the capability of understanding event-based scenes. Our EventGPT comprises an event encoder, followed by a spatio-temporal aggregator, a linear projector, an event-language adapter, and an LLM. Firstly, RGB image-text pairs generated by GPT are leveraged to warm up the linear projector, referring to LLaVA, as the gap between natural image and language modalities is relatively smaller. Secondly, we construct a synthetic yet large dataset, N-ImageNet-Chat, consisting of event frames and corresponding texts to enable the use of the spatio-temporal aggregator and to train the event-language adapter, thereby aligning event features more closely with the language space. Finally, we gather an instruction dataset, Event-Chat, which contains extensive real-world data to fine-tune the entire model, further enhancing its generalization ability. We construct a comprehensive benchmark, and experiments show that EventGPT surpasses previous state-of-the-art MLLMs in generation quality, descriptive accuracy, and reasoning capability.

cs.CV

Frequency stabilization based on H13C14N absorption in lithium niobate micro-disk laser

We demonstrate an on-chip lithium niobate micro-disk laser based on hydrogen cyanide (H13C14N) gas saturation absorption method for frequency stabilization. The laser chip consists of two main components: a micro-disk laser and a combined racetrack ring cavity. By operating on the H13C14N P12 absorption line at 1551.3 nm, the laser frequency can be precisely stabilized. The laser demonstrates remarkable stability, achieving a best stability value of 9*10^-9. Furthermore, the short-term stability, evaluated over continuous time intervals of 35 seconds, showcases exceptional performance. Additionally, the residual drift remains well below 30 MHz.

physics.optics

Integrated multi-color Raman microlasers with ultra-low pump levels in single high-Q lithium niobate microdisks

Photonic integrated Raman microlasers, particularly discrete multi-color lasers which are crucial for extending the emission wavelength range of chip-scale laser sources to much shorter wavelength, are highly in demand for various spectroscopy, microscopy analysis, and biological detection. However, integrated multi-color Raman microlasers have yet to be demonstrated because of the requirement of high-Q microresonators possessing large second-order nonlinearity and strong Raman phonon branches and the challenging in cavity-enhanced multi-photon hyper-Raman scattering parametric process. In this work, integrated multi-color Raman lasers have been demonstrated for the first time at weak pump levels, via the excitation of high-Q (>6 X 10^6) phase-matched modes in single thin-film lithium niobate (TFLN) microresonators by dispersion engineering. Raman lasing was observed at 1712 nm for a 1546-nm pump threshold power of only 620 uW. Furthermore, multi-color Raman lasers were realized at discrete wavelengths of 1712 nm, 813 nm, 533 nm and 406 nm with pump levels as low as 1.60 mW, which is more than two order of magnitude lower than the current records (i.e., 200 mW) in bulk resonators, allowed by the fulfillment of the requisite conditions consisting of broadband natural phase match, multiple-resonance and high Q-factors.

physics.optics

Erbium-ytterbium co-doped lithium niobate single-mode microdisk laser with an ultralow threshold of 1 uW

We demonstrate single-mode microdisk lasers in the telecom band with ultra-low thresholds on erbium-ytterbium co-doped thin-film lithium niobate (TFLN). The active microdisk were fabricated with high-Q factors by photo-lithography assisted chemo-mechanical etching. Thanks to the erbium-ytterbium co-doping providing high optical gain, the ultra-low loss nanostructuring, and the excitation of high-Q coherent polygon modes which suppresses multi-mode lasing and allows high spatial mode overlap factor between pump and lasing modes, single-mode laser emission operating at 1530 nm wavelength was observed with an ultra-low threshold, under 980-nm-band optical pump. The threshold was measured as low as 1 uW, which is one order of magnitude smaller than the best results previously reported in single-mode active TFLN microlasers. And the conversion efficiency reaches 0.406%, which is also the highest value reported in single-mode active TFLN microlasers.

physics.optics