SearcharxivSearch

arXiv subjects

Xian Zhao

Publications and source records attributed to Xian Zhao.

12 recordsLinked to original sources

Superbunched random fiber laser

Photon superbunching, distinguished by second-order coherence values far exceeding the Gaussian thermal limit, represents a highly desirable resource for quantum optics and correlation-based imaging technologies. However, existing approaches typically rely on fragile experimental platforms, inefficient nonlinear conversion processes, or mechanically complex optical architectures. Here, we demonstrate a fully fiber-integrated superbunched random fiber laser (SRFL) in which intrinsic Rayleigh scattering cooperatively interacts with cascaded stimulated Brillouin scattering and quasi-phase-matched four-wave mixing to tailor extreme photon statistics. The SRFL generates a multi-wavelength comb, in which individual spectral components exhibit widely tunable photon bunching, with the second-order coherence g(2)(0) continuously controlled from ~1 to ~26 by tuning the pump power, spectral order and diffusion length. Moreover, we establish a direct correlation between photonic phase transitions (quantified by the Parisi overlap order parameter) and the emergence of superbunching, thereby bridging macroscopic disorder physics and microscopic photon statistics. Finally, we employ the superbunched emission for temporal ghost imaging, realizing high-fidelity temporal object reconstruction with a substantial reduction in required ensemble averaging. These findings validate random fiber lasers as a robust, scalable, and integrated platform for generating extreme photon statistics and unlock new avenues for correlation-enhanced photonic sensing and quantum optics investigations in complex photonic systems.

physics.optics

Mind with Eyes: from Language Reasoning to Multimodal Reasoning

Language models have recently advanced into the realm of reasoning, yet it is through multimodal reasoning that we can fully unlock the potential to achieve more comprehensive, human-like cognitive capabilities. This survey provides a systematic overview of the recent multimodal reasoning approaches, categorizing them into two levels: language-centric multimodal reasoning and collaborative multimodal reasoning. The former encompasses one-pass visual perception and active visual perception, where vision primarily serves a supporting role in language reasoning. The latter involves action generation and state update within reasoning process, enabling a more dynamic interaction between modalities. Furthermore, we analyze the technical evolution of these methods, discuss their inherent challenges, and introduce key benchmark tasks and evaluation metrics for assessing multimodal reasoning performance. Finally, we provide insights into future research directions from the following two perspectives: (i) from visual-language reasoning to omnimodal reasoning and (ii) from multimodal reasoning to multimodal agents. This survey aims to provide a structured overview that will inspire further advancements in multimodal reasoning research.

cs.CL

Edge-Dependent Step-Flow Growth Mechanism in $β$-Ga$_{2}$O$_{3}$ (100) Facet at the Atomic Level

Homoepitaxial step-flow growth of high-quality $β$-Ga$_{2}$O$_{3}$ thin films is essential for the advancement of high-performance Ga$_{2}$O$_{3}$-based devices. In this work, the step-flow growth mechanism of $β$-Ga$_{2}$O$_{3}$ (100) facet is explored by machine-learning molecular dynamics simulations and density functional theory calculations. Our results reveal that Ga adatoms and Ga-O adatom pairs, with their high mobility, are the primary atomic species responsible for efficient surface migration on the (100) facet. The asymmetric monoclinic structure of $β$-Ga$_{2}$O$_{3}$ induces a distinct two-stage Ehrlich-Schwoebel barrier for Ga adatoms at the [00$\overline{1}$] step edge, contributing to the suppression of double-step and hillock formation. Furthermore, a miscut towards [00$\overline{1}$] does not induce the nucleation of stable twin boundaries, whereas a miscut towards [001] leads to the spontaneous formation of twin boundaries. This research provides meaningful insights not only for high-quality $β$-Ga$_{2}$O$_{3}$ homoepitaxy but also the step-flow growth mechanism of other similar systems.

cond-mat.mtrl-sci

Revealing spontaneous symmetry breaking in continuous time crystals

Spontaneous symmetry breaking plays a pivotal role in physics ranging from the emergence of elementary particles to the phase transitions of matter. The spontaneous breaking of continuous time translation symmetry leads to a novel state of matter named continuous time crystal (CTC). It exhibits periodic oscillation without the need for periodic driving, and the relative phases for repetitively realized oscillations are random. However, the mechanism behind the spontaneous symmetry breaking in CTCs, particularly the random phases, remains elusive. Here we propose and experimentally realize two types of CTCs based on distinct mechanisms: manifold topology and near-chaotic motion. We observe both types of CTCs in thermal atomic ensembles by artificially synthesizing spin-spin nonlinear interactions through a measurement-feedback scheme. Our work provides general recipes for the realization of CTCs, and paves the way for exploring CTCs in various systems.

quant-ph

Parallel fast random bit generation based on spectrotemporally uncorrelated Brillouin random fiber lasing oscillation

Correlations existing between spectral components in multi-wavelength lasers have been the key challenge that hinders these laser sources from being developed to chaotic comb entropy sources for parallel random bit generation. Herein, spectrotemporally uncorrelated multi-order Stokes/anti-Stokes emissions are achieved by cooperatively exploiting nonlinear optical processes including cascaded stimulated Brillouin scattering and quasi-phase-matched four-wave mixing in a Brillouin random fiber laser. Chaotic instabilities induced by random mode resonance are enhanced and disorderly redistributed among different lasing lines through complex nonlinear optical interactions, which comprehensively releases the inherent correlation among multiple Stokes/anti-Stokes emission lines, realizing a chaotic frequency comb with multiple spectrotemporally uncorrelated channels. Parallel fast random bit generation is fulfilled with 31 channels, single-channel bit rate of 35-Gbps and total bit rate of 1.085-Tbps. National Institute of Standards and Technology statistic tests verify the randomness of generated bit streams. This work, in a simple and efficient way, breaks the correlation barrier for utilizing multi-wavelength laser to achieve high-quality spectrotemporally uncorrelated chaotic laser source, opening new avenues for achieving greatly accelerated random bit generation through parallelization and potentially revolutionizing the current architecture of secure communication and high-performance computation.

physics.optics

An Experimental Study of Semantic Continuity for Deep Learning Models

Deep learning models suffer from the problem of semantic discontinuity: small perturbations in the input space tend to cause semantic-level interference to the model output. We argue that the semantic discontinuity results from these inappropriate training targets and contributes to notorious issues such as adversarial robustness, interpretability, etc. We first conduct data analysis to provide evidence of semantic discontinuity in existing deep learning models, and then design a simple semantic continuity constraint which theoretically enables models to obtain smooth gradients and learn semantic-oriented features. Qualitative and quantitative experiments prove that semantically continuous models successfully reduce the use of non-semantic information, which further contributes to the improvement in adversarial robustness, interpretability, model transfer, and machine bias.

cs.LG

Benign Adversarial Attack: Tricking Models for Goodness

In spite of the successful application in many fields, machine learning models today suffer from notorious problems like vulnerability to adversarial examples. Beyond falling into the cat-and-mouse game between adversarial attack and defense, this paper provides alternative perspective to consider adversarial example and explore whether we can exploit it in benign applications. We first attribute adversarial example to the human-model disparity on employing non-semantic features. While largely ignored in classical machine learning mechanisms, non-semantic feature enjoys three interesting characteristics as (1) exclusive to model, (2) critical to affect inference, and (3) utilizable as features. Inspired by this, we present brave new idea of benign adversarial attack to exploit adversarial examples for goodness in three directions: (1) adversarial Turing test, (2) rejecting malicious model application, and (3) adversarial data augmentation. Each direction is positioned with motivation elaboration, justification analysis and prototype applications to showcase its potential.

cs.AI

RangeRCNN: Towards Fast and Accurate 3D Object Detection with Range Image Representation

We present RangeRCNN, a novel and effective 3D object detection framework based on the range image representation. Most existing methods are voxel-based or point-based. Though several optimizations have been introduced to ease the sparsity issue and speed up the running time, the two representations are still computationally inefficient. Compared to them, the range image representation is dense and compact which can exploit powerful 2D convolution. Even so, the range image is not preferred in 3D object detection due to scale variation and occlusion. In this paper, we utilize the dilated residual block (DRB) to better adapt different object scales and obtain a more flexible receptive field. Considering scale variation and occlusion, we propose the RV-PV-BEV (range view-point view-bird's eye view) module to transfer features from RV to BEV. The anchor is defined in BEV which avoids scale variation and occlusion. Neither RV nor BEV can provide enough information for height estimation; therefore, we propose a two-stage RCNN for better 3D detection performance. The aforementioned point view not only serves as a bridge from RV to BEV but also provides pointwise features for RCNN. Experiments show that RangeRCNN achieves state-of-the-art performance on the KITTI dataset and the Waymo Open dataset, and provides more possibilities for real-time 3D object detection. We further introduce and discuss the data augmentation strategy for the range image based method, which will be very valuable for future research on range image.

cs.CV

Accelerating single-crystal growth by stimulated and self-guided channeling

We report a self-guided and "stimulated" single-crystal growth acceleration effect in static super-saturated aqueous solutions, producing inorganic (KH$_2$PO$_4$) and organic (tetraphenyl-phosphonium-family) nonlinear optical single-crystals with novel morphologies. The extraordinarily fast unidirectional growth in the presence of complete lateral growth suppression defies all current impurity, defect and dislocation based crystal growth inhibition mechanisms. We propose a self-channeling-stimulated accelerated growth theory that can satisfactorily explain all experimental results. Using molecular dynamics analysis and a modified two-component crystal growth model that includes microscopic surface molecular selectivity we show the lateral growth arrest is the combined result of the self-channeling and a self-shielding effect. These single-crystals exhibit remarkable mechanical flexibility in winding and twisting, demonstrating their unique advantages for chip-size quantum and biomedical applications, as well as for production of high-yield/high-potency pharmaceutical materials.

physics.chem-ph

MAFF-Net: Filter False Positive for 3D Vehicle Detection with Multi-modal Adaptive Feature Fusion

3D vehicle detection based on multi-modal fusion is an important task of many applications such as autonomous driving. Although significant progress has been made, we still observe two aspects that need to be further improvement: First, the specific gain that camera images can bring to 3D detection is seldom explored by previous works. Second, many fusion algorithms run slowly, which is essential for applications with high real-time requirements(autonomous driving). To this end, we propose an end-to-end trainable single-stage multi-modal feature adaptive network in this paper, which uses image information to effectively reduce false positive of 3D detection and has a fast detection speed. A multi-modal adaptive feature fusion module based on channel attention mechanism is proposed to enable the network to adaptively use the feature of each modal. Based on the above mechanism, two fusion technologies are proposed to adapt to different usage scenarios: PointAttentionFusion is suitable for filtering simple false positive and faster; DenseAttentionFusion is suitable for filtering more difficult false positive and has better overall performance. Experimental results on the KITTI dataset demonstrate significant improvement in filtering false positive over the approach using only point cloud data. Furthermore, the proposed method can provide competitive results and has the fastest speed compared to the published state-of-the-art multi-modal methods in the KITTI benchmark.

cs.CV

Adversarial Privacy-preserving Filter

While widely adopted in practical applications, face recognition has been critically discussed regarding the malicious use of face images and the potential privacy problems, e.g., deceiving payment system and causing personal sabotage. Online photo sharing services unintentionally act as the main repository for malicious crawler and face recognition applications. This work aims to develop a privacy-preserving solution, called Adversarial Privacy-preserving Filter (APF), to protect the online shared face images from being maliciously used.We propose an end-cloud collaborated adversarial attack solution to satisfy requirements of privacy, utility and nonaccessibility. Specifically, the solutions consist of three modules: (1) image-specific gradient generation, to extract image-specific gradient in the user end with a compressed probe model; (2) adversarial gradient transfer, to fine-tune the image-specific gradient in the server cloud; and (3) universal adversarial perturbation enhancement, to append image-independent perturbation to derive the final adversarial noise. Extensive experiments on three datasets validate the effectiveness and efficiency of the proposed solution. A prototype application is also released for further evaluation.We hope the end-cloud collaborated attack framework could shed light on addressing the issue of online multimedia sharing privacy-preserving issues from user side.

cs.CR

Strong mesoscopic transverse growth suppression in high-speed unidirectional growth of KDP single crystals

We show a strong mesoscopic transverse growth arrest that co-exists with a high-speed longitudinal growth in single-crystalline KDP crystals with large aspect ratios. To explain this unique growth morphology, which cannot be explained by any current theories, we introduce a new set of surface concentration rate equations and demonstrate a molecular-orientation-selective surface self-shielding and channeling mechanism. We introduce a local supersaturation and calculate crystal growth driving force and growth rate, demonstrating quick arrest of both quantities as results of molecule-orientation selectivity based self-shielding and channeling effects. The growth dynamics thus derived can satisfactorily explain all experimental observations in single-crystalline KDP crystal growth reported here.

physics.app-ph