SearcharxivSearch

arXiv subjects

Bo Cao

Publications and source records attributed to Bo Cao.

At least 19 recordsLinked to original sources

Analytical Statistics of Vortex Beams in a Turbulent Channel for OAM-Multiplexed FSO Communications

Orbital angular momentum (OAM) multiplexing can increase the capacity of free-space optical (FSO) communications, whereas atmospheric turbulence causes modal crosstalk and irradiance fluctuations that degrade demultiplexing performance. Analytical modeling is therefore important for characterizing turbulence-induced propagation effects and the demultiplexed port-power statistics of OAM channels. In this paper, we first study the receiver-plane irradiance statistics of vortex beams after propagating through the turbulent channel. The average irradiance is derived using frequency-domain convolution, and a closed-form frequency-domain diffraction kernel is obtained based on extended Rytov theory to evaluate the scintillation index for moderate and strong turbulence. However, receiver-plane irradiance statistics alone are insufficient to describe the performance of OAM-multiplexed FSO communications. We therefore derive the demultiplexed port-power statistics. Specifically, we derive the average port power, modal crosstalk, port-power variance, and cross-port covariance based on the complex Gaussian expansion of general LG vortex fields and the extended Huygens-Fresnel framework. The demultiplexed port-power statistics are then used to evaluate the symbol-error rate (SER) of OAM-multiplexed FSO communications. Numerical results demonstrate that all derived statistics are consistent with those obtained by phase-screen simulations under different turbulence strengths and beam parameters. The resulting SER performance further shows that OAM-multiplexing performance is more sensitive to mode spacing for small receiver apertures than for large apertures.

eess.SP

From Real-Time Planning to Reliable Execution:Scalable Coordination for Heterogeneous Multi-Robot Fleets in Industrial Environments

With the increasing deployment of heterogeneous robot fleets in industrial environments, efficient coordination remains a critical challenge. Real-time path planning must simultaneously accommodate high robot densities and heterogeneous motion capabilities, while communication delays, execution uncertainties, and other disturbances may cause robots to deviate from the temporal assumptions underlying planned paths. Such deviations can lead to excessive waiting and congestion propagation across the fleet. This paper presents SCALE, a reactive online coordination framework that enables real-time planning while maintaining robust execution. Within this framework, we introduce a motion-induced conflict reduction mechanism to support the online generation of feasible paths for online conflict resolution. To mitigate the effects of disturbances, we further design a generalized Conjugate Action-Precedence Hypergraph (CAPH) that adaptively adjusts precedence relations among robots. Extensive validation experiments, together with a three-day deployment in a warehouse, demonstrate the

cs.RO

AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image Detection

The realism of AI-generated images (AIGI) poses increasing challenges for reliable forensic detection, where heterogeneous expert detectors may produce conflicting predictions across diverse generative sources and post-processing conditions. Existing multi-expert fusion methods rely on fixed rules or learned fusion strategies, offering limited ability to assess sample-specific reliability, execute rigorous adjudication of conflicts, and provide evidence-grounded explanations. We propose AgentFoX, an LLM-driven agentic multi-expert framework for AIGI detection that employs a command-and-reasoning core to perform evidence fusion. Following predefined guidelines, the core coordinates designated subtasks to collect semantic and signal-level evidence, reason over structured contexts to determine authenticity, and generate an auditable report for explainability. During this process, Expert Profiles are constructed for model-centric reliability assessment, while Clustering Profiles are built for data-centric contextual analysis, jointly establishing evidence contexts for conflict resolution. Extensive evaluations across diverse benchmarks demonstrate the robustness and generalizability of AgentFoX under complex conditions.

cs.CV

MPF-Net: Exposing High-Fidelity AI-Generated Video Forgeries via Hierarchical Manifold Deviation and Micro-Temporal Fluctuations

With the rapid advancement of video generation models such as Veo and Wan, the visual quality of synthetic content has reached a level where macro-level semantic errors and temporal inconsistencies are no longer prominent. However, this does not imply that the distinction between real and cutting-edge high-fidelity fake is untraceable. We argue that AI-generated videos are essentially products of a manifold-fitting process rather than a physical recording. Consequently, the pixel composition logic of consecutive adjacent frames residual in AI videos exhibits a structured and homogenous characteristic. We term this phenomenon `Manifold Projection Fluctuations' (MPF). Driven by this insight, we propose a hierarchical dual-path framework that operates as a sequential filtering process. The first, the Static Manifold Deviation Branch, leverages the refined perceptual boundaries of Large-Scale Vision Foundation Models (VFMs) to capture residual spatial anomalies or physical violations that deviate from the natural real-world manifold (off-manifold). For the remaining high-fidelity videos that successfully reside on-manifold and evade spatial detection, we introduce the Micro-Temporal Fluctuation Branch as a secondary, fine-grained filter. By analyzing the structured MPF that persists even in visually perfect sequences, our framework ensures that forgeries are exposed regardless of whether they manifest as global real-world manifold deviations or subtle computational fingerprints.

cs.CV

String-Level Ground Fault Localization for TN-Earthed Three-Phase Photovoltaic Systems

The DC-side ground fault (GF) poses significant risks to three-phase TN-earthed photovoltaic (PV) systems, as the resulting high fault current can directly damage both PV inverters and PV modules. Once a fault occurs, locating the faulty string through manual string-by-string inspection is highly time-consuming and inefficient. This work presents a comprehensive analysis of GF characteristics through fault-current analysis and a simulation-based case study covering multiple fault locations. Building on these insights, we propose an edge-AI-based GF localization approach tailored for three-phase TN-earthed PV systems. A PLECS-based simulation model that incorporates PV hysteresis effects is developed to generate diverse GF scenarios, from which correlation-based features are extracted throughout the inverter's four-stage shutdown sequence. Using the simulated dataset, a lightweight Variational Information Bottleneck (VIB)-based localization model is designed and trained, achieving over 93% localization accuracy at typical sampling rates with low computational cost, demonstrating strong potential for deployment on resource-constrained PV inverters.

eess.SY

Enriched text-guided variational multimodal knowledge distillation network (VMD) for automated diagnosis of plaque vulnerability in 3D carotid artery MRI

Multimodal learning has attracted much attention in recent years due to its ability to effectively utilize data features from a variety of different modalities. Diagnosing the vulnerability of atherosclerotic plaques directly from carotid 3D MRI images is relatively challenging for both radiologists and conventional 3D vision networks. In clinical practice, radiologists assess patient conditions using a multimodal approach that incorporates various imaging modalities and domain-specific expertise, paving the way for the creation of multimodal diagnostic networks. In this paper, we have developed an effective strategy to leverage radiologists' domain knowledge to automate the diagnosis of carotid plaque vulnerability through Variation inference and Multimodal knowledge Distillation (VMD). This method excels in harnessing cross-modality prior knowledge from limited image annotations and radiology reports within training data, thereby enhancing the diagnostic network's accuracy for unannotated 3D MRI images. We conducted in-depth experiments on the dataset collected in-house and verified the effectiveness of the VMD strategy we proposed.

cs.CV

Electrically pumped ultra-efficient quantum frequency conversion on thin film lithium niobate chip

Quantum frequency conversion (QFC) plays a crucial role in constructing seamless interconnection between quantum systems operating at different wavelengths. To advance future quantum technology, chip-scale integrated QFC components, featuring high efficiency, small footprint, low power consumption and high scalability, are indispensable. In this work, we demonstrate the first hybrid integrated QFC chip on thin film lithium niobate platform that connects the telecom and visible bands. Benefiting from the periodically poled microring resonator with ulta-high normalized conversion efficiency of 386,000 %/W, an ultra-low pump power of 360 {\mu}W is achieved which is more than two orders of magnitude lower than traditional straight waveguide scheme. By injecting current into the chip, an on-chip quantum efficiency of 57% and a noise count of ~ 7k counts per second are achieved. Such an electrically pumped, integrated and scalable QFC chip would significantly advancing the integration of quantum network and the development of chip-scale quantum optical systems.

quant-ph

Electrically pumped ultrabright entangled photons on chip

Entangled photon sources (EPS) are essential for quantum science and technology. Despite advancements in integrated optical platforms like thin-film lithium niobate, a scalable, high-performance, chip-scale EPS has remained elusive. We address this by demonstrating an electrically pumped, post-selection-free polarization-EPS, achieved through hybrid integration of a distributed feedback laser with thin-film lithium niobate chip which integrates periodically poled lithium niobate waveguides, beam splitter, and polarization rotator combiner. By injecting current into the chip, we realize a high-performance EPS with a bandwidth of 73 nm and an entanglement pair generation rate of 4.5*10^10 pairs/s/mW. The polarization entanglement shows Bell-state fidelities above 96% across frequency-correlated modes. This compact, integrated EPS enables key applications, including high-speed quantum key distribution via wavelength division multiplexing, satellite-based quantum communication, and entanglement-based quantum metrology.

quant-ph

Wi-Fi Sensing Tool Release: Gathering 802.11ax Channel State Information from a Commercial Wi-Fi Access Point

Wi-Fi sensing has emerged as a powerful technology, leveraging channel state information (CSI) extracted from wireless data packets to enable diverse applications, ranging from human presence detection to gesture recognition and health monitoring. However, CSI extraction from commercial Wi-Fi access point lacks and out of date. This paper introduces ZTECSITool,a toolkit designed to capture high-resolution CSI measurements from commercial Wi-Fi 6 (802.11ax) access points, supporting bandwidths up to 160 MHz and 512 subcarriers. ZTECSITool bridges a critical gap in Wi-Fi sensing research, facilitating the development of next-generation sensing systems. The toolkit includes customized firmware and open-source software tools for configuring, collecting, and parsing CSI data, offering researchers a robust platform for advanced sensing applications. We detail the command protocols for CSI extraction, including band selection,STA filtering, and report configuration, and provide insights into the data structure of the reported CSI. Additionally, we present a Python-based graphical interface for real-time CSI visualization and analysis

eess.SP

Estimating Quality in Therapeutic Conversations: A Multi-Dimensional Natural Language Processing Framework

Engagement between client and therapist is a critical determinant of therapeutic success. We propose a multi-dimensional natural language processing (NLP) framework that objectively classifies engagement quality in counseling sessions based on textual transcripts. Using 253 motivational interviewing transcripts (150 high-quality, 103 low-quality), we extracted 42 features across four domains: conversational dynamics, semantic similarity as topic alignment, sentiment classification, and question detection. Classifiers, including Random Forest (RF), Cat-Boost, and Support Vector Machines (SVM), were hyperparameter tuned and trained using a stratified 5-fold cross-validation and evaluated on a holdout test set. On balanced (non-augmented) data, RF achieved the highest classification accuracy (76.7%), and SVM achieved the highest AUC (85.4%). After SMOTE-Tomek augmentation, performance improved significantly: RF achieved up to 88.9% accuracy, 90.0% F1-score, and 94.6% AUC, while SVM reached 81.1% accuracy, 83.1% F1-score, and 93.6% AUC. The augmented data results reflect the potential of the framework in future larger-scale applications. Feature contribution revealed conversational dynamics and semantic similarity between clients and therapists were among the top contributors, led by words uttered by the client (mean and standard deviation). The framework was robust across the original and augmented datasets and demonstrated consistent improvements in F1 scores and recall. While currently text-based, the framework supports future multimodal extensions (e.g., vocal tone, facial affect) for more holistic assessments. This work introduces a scalable, data-driven method for evaluating engagement quality of the therapy session, offering clinicians real-time feedback to enhance the quality of both virtual and in-person therapeutic interactions.

cs.CL

Understanding LLM Scientific Reasoning through Promptings and Model's Explanation on the Answers

Large language models (LLMs) have demonstrated remarkable capabilities in natural language understanding, reasoning, and problem-solving across various domains. However, their ability to perform complex, multi-step reasoning task-essential for applications in science, medicine, and law-remains an area of active investigation. This paper examines the reasoning capabilities of contemporary LLMs, analyzing their strengths, limitations, and potential for improvement. The study uses prompt engineering techniques on the Graduate-Level GoogleProof Q&A (GPQA) dataset to assess the scientific reasoning of GPT-4o. Five popular prompt engineering techniques and two tailored promptings were tested: baseline direct answer (zero-shot), chain-of-thought (CoT), zero-shot CoT, self-ask, self-consistency, decomposition, and multipath promptings. Our findings indicate that while LLMs exhibit emergent reasoning abilities, they often rely on pattern recognition rather than true logical inference, leading to inconsistencies in complex problem-solving. The results indicated that self-consistency outperformed the other prompt engineering technique with an accuracy of 52.99%, followed by direct answer (52.23%). Zero-shot CoT (50%) outperformed multipath (48.44%), decomposition (47.77%), self-ask (46.88%), and CoT (43.75%). Self-consistency performed the second worst in explaining the answers. Simple techniques such as direct answer, CoT, and zero-shot CoT have the best scientific reasoning. We propose a research agenda aimed at bridging these gaps by integrating structured reasoning frameworks, hybrid AI approaches, and human-in-the-loop methodologies. By critically evaluating the reasoning mechanisms of LLMs, this paper contributes to the ongoing discourse on the future of artificial general intelligence and the development of more robust, trustworthy AI systems.

cs.AI

Human vs. LLM-Based Thematic Analysis for Digital Mental Health Research: Proof-of-Concept Comparative Study

Thematic analysis provides valuable insights into participants' experiences through coding and theme development, but its resource-intensive nature limits its use in large healthcare studies. Large language models (LLMs) can analyze text at scale and identify key content automatically, potentially addressing these challenges. However, their application in mental health interviews needs comparison with traditional human analysis. This study evaluates out-of-the-box and knowledge-base LLM-based thematic analysis against traditional methods using transcripts from a stress-reduction trial with healthcare workers. OpenAI's GPT-4o model was used along with the Role, Instructions, Steps, End-Goal, Narrowing (RISEN) prompt engineering framework and compared to human analysis in Dedoose. Each approach developed codes, noted saturation points, applied codes to excerpts for a subset of participants (n = 20), and synthesized data into themes. Outputs and performance metrics were compared directly. LLMs using the RISEN framework developed deductive parent codes similar to human codes, but humans excelled in inductive child code development and theme synthesis. Knowledge-based LLMs reached coding saturation with fewer transcripts (10-15) than the out-of-the-box model (15-20) and humans (90-99). The out-of-the-box LLM identified a comparable number of excerpts to human researchers, showing strong inter-rater reliability (K = 0.84), though the knowledge-based LLM produced fewer excerpts. Human excerpts were longer and involved multiple codes per excerpt, while LLMs typically applied one code. Overall, LLM-based thematic analysis proved more cost-effective but lacked the depth of human analysis. LLMs can transform qualitative analysis in mental healthcare and clinical research when combined with human oversight to balance participant perspectives and research resources.

cs.HC

Early signs of stuck pipe detection based on Crossformer

Stuck pipe incidents are one of the major challenges in drilling engineering,leading to massive time loss and additional costs.To address the limitations of insufficient long sequence modeling capability,the difficulty in accurately establishing warning threshold,and the lack of model interpretability in existing methods,we utilize Crossformer for early signs of detection indicating potential stuck events in order to provide guidance for on-site drilling engineers and prevent stuck pipe incidents.The sliding window technique is integrated into Crossformer to allow it to output and display longer outputs,the improved Crossformer model is trained using normal time series drilling data to generate predictions for various parameters at each time step.The relative reconstruction error of model is regard as the risk of stuck pipe,thereby considering data that the model can't predict as anomalies,which represent the early signs of stuck pipe incidents.The multi-step prediction capability of Crossformer and relative reconstruction error are combined to assess stuck pipe risk at each time step in advance.We partition the reconstruction error into modeling error and error due to anomalous data fluctuations,furthermore,the dynamic warning threshold and warning time for stuck pipe incidents are determined using the probability density function of reconstruction errors from normal drilling data.The results indicate that our method can effectively detect early signs of stuck pipe incidents during the drilling process.Crossformer exhibits superior modeling and predictive capabilities compared with other deep learning models.Transformer-based models with multi-step prediction capability are more suitable for stuck pipe prediction compared to the current single-step prediction models.

cs.CE

Observation of spatiotemporal stabilizer in a multi-mode fibre laser

Spatiotemporal mode-locking (STML) has become an emerging approach to realize organized wavepackets in high-dimensional nonlinear photonic systems. Mode-locking in one dimensional systems employs a saturable absorber to resist fluctuations in the temporal domain. Analogous suppression of fluctuations in the space-time domains to retain a consistent output should also exist for STML. However, experimental evidence of such a resistance remains elusive, to our knowledge. Here, we report experimental observation of such a spatiotemporal stabilizer in STML, by embedding a spatial light modulator (SLM) into a multi-mode fibre (MMF) laser. Mode decomposition reveals the mode content remains steady for an STML state when applying phase perturbations on the SLM. Conversely, the mode content changes significantly for a non-STML lasing state. Numerical simulations confirm our observation and show that spatial filtering and saturable absorber mainly contribute to the observed stability. The capability to resist the spatial phase fluctuations is observed to depend on the intracavity pulse energy as well as the modal pulse energy condensed in the low-order modes. Our work constitutes another building block for the concept of STML in multi-mode photonic systems.

physics.optics

Coherence memory and amnesia in a mode-locked laser

Self-organization of temporal modes in mode-locked lasers usually starts from quantum noise. In this process, incoherent spontaneous emission is steered into coherent ultrashort pulses by dissipation and nonlinearity. In this work, we investigated self-organization dynamics in a mode-locked Mamyshev oscillator starting from coherent pulse seeds as opposed to quantum noise. We observed that the coherence of the seed can be remembered or forgotten depending on the initial inverse population. The excessive nonlinearity in the coherence amnesia regime can devastate the seed coherence, causing the oscillator to undergo a chaotic transition lasting hundreds of round trips before regaining coherence. Conversely, the oscillator converges in only a few round trips for the coherence memory regime. A heterodyne technique was developed to record the fast varying optical phase and characterize these two regimes. Dissipative soliton molecules were synthesized from external pulse pair seeds via the coherence memory pathway. In this case, a plateau of the generated pulse spacing independent from seed pulse spacing, i.e., amnesia of the seed spacing, was observed for close spaced seed pulse pairs. Moreover, we show that pulse seeds can be used for laser reconfiguration and pulse pattern control. Our work paves a way to control transient pulse dynamics and steady pulse forms on demand in mode-locked lasers.

physics.optics

On Data-Driven Modeling and Control in Modern Power Grids Stability: Survey and Perspective

Modern power grids are fast evolving with the increasing volatile renewable generation, distributed energy resources (DERs) and time-varying operating conditions. The DERs include rooftop photovoltaic (PV), small wind turbines, energy storages, flexible loads, electric vehicles (EVs), etc. The grid control is confronted with low inertia, uncertainty and nonlinearity that challenge the operation security, efficacy and efficiency. The ongoing digitization of power grids provides opportunities to address the challenges with data-driven and control. This paper provides a comprehensive review of emerging data-driven dynamical modeling and control methods and their various applications in power grid. Future trends are also discussed based on advances in data-driven control.

eess.SY

FAN: Fatigue-Aware Network for Click-Through Rate Prediction in E-commerce Recommendation

Since clicks usually contain heavy noise, increasing research efforts have been devoted to modeling implicit negative user behaviors (i.e., non-clicks). However, they either rely on explicit negative user behaviors (e.g., dislikes) or simply treat non-clicks as negative feedback, failing to learn negative user interests comprehensively. In such situations, users may experience fatigue because of seeing too many similar recommendations. In this paper, we propose Fatigue-Aware Network (FAN), a novel CTR model that directly perceives user fatigue from non-clicks. Specifically, we first apply Fourier Transformation to the time series generated from non-clicks, obtaining its frequency spectrum which contains comprehensive information about user fatigue. Then the frequency spectrum is modulated by category information of the target item to model the bias that both the upper bound of fatigue and users' patience is different for different categories. Moreover, a gating network is adopted to model the confidence of user fatigue and an auxiliary task is designed to guide the learning of user fatigue, so we can obtain a well-learned fatigue representation and combine it with user interests for the final CTR prediction. Experimental results on real-world datasets validate the superiority of FAN and online A/B tests also show FAN outperforms representative CTR models significantly.

cs.IR

MOEF: Modeling Occasion Evolution in Frequency Domain for Promotion-Aware Click-Through Rate Prediction

Promotions are becoming more important and prevalent in e-commerce to attract customers and boost sales, leading to frequent changes of occasions, which drives users to behave differently. In such situations, most existing Click-Through Rate (CTR) models can't generalize well to online serving due to distribution uncertainty of the upcoming occasion. In this paper, we propose a novel CTR model named MOEF for recommendations under frequent changes of occasions. Firstly, we design a time series that consists of occasion signals generated from the online business scenario. Since occasion signals are more discriminative in the frequency domain, we apply Fourier Transformation to sliding time windows upon the time series, obtaining a sequence of frequency spectrum which is then processed by Occasion Evolution Layer (OEL). In this way, a high-order occasion representation can be learned to handle the online distribution uncertainty. Moreover, we adopt multiple experts to learn feature representations from multiple aspects, which are guided by the occasion representation via an attention mechanism. Accordingly, a mixture of feature representations is obtained adaptively for different occasions to predict the final CTR. Experimental results on real-world datasets validate the superiority of MOEF and online A/B tests also show MOEF outperforms representative CTR models significantly.

cs.LG