SearcharxivSearch

arXiv subjects

Shuai Gao

Publications and source records attributed to Shuai Gao.

15 recordsLinked to original sources

Baseband-Efficient WMMSE Precoding: From a Signal Weighting Cost Perspective

For downlink transmission in massive multi-user multiple-input multiple-output (MU-MIMO) systems, conventional precoding research heavily focuses on reducing the computational complexity of precoding matrix design, while largely overlooking another critical bottleneck: the substantial signal weighting cost incurred by repeatedly applying the precoder to high-speed data streams. To address both challenges simultaneously, this paper proposes a novel sparse precoding framework tailored for fully-digital architectures. Within this framework, from the sum-rate maximization perspective, we design two sparse precoding architectures: a common-support row-sparse architecture and a user-specific row-sparse architecture, so as to reduce the number of multiplication operations required in baseband signal weighting while largely preserving the achievable sum-rate. For the formulated mixed-integer non-linear programming (MINLP) problem, we rigorously prove that the optimal precoder under both sparse architectures strictly resides in a specific low-dimensional subspace determined by the channel matrices, thereby reducing the dimensionality of the optimization variables. Based on this insight, an alternating optimization algorithm is developed within the weighted minimum mean square error (WMMSE) framework to jointly optimize sparse beam selection and low-dimensional precoding coefficients. The combinatorial beam selection problem is handled using an efficient penalty-based majorize-minimization (MM) method, yielding a low-complexity closed-form solution. Simulation results demonstrate that the proposed schemes achieve sum-rate performance close to that of full-dimensional WMMSE precoding without sparsity constraints, while substantially reducing the overall signal-weighting cost.

eess.SP

DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation

Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.

cs.CV

Insight: Enhancing Mobile Accessibility for Blind and Visually Impaired Users with LLMs

This research paper addresses the limitations of current mobile accessibility services like TalkBack, which provide manual gesture-based sequential feedback to BVI users. Motivated by the promise of large language models (LLMs), this paper introduces Insight, an Android accessibility service that provides natural language interaction and real-time summarization of the screen. The paper performs a within-subject experimental study with users to compare Insight and TalkBack on usability factors. Results show Insight reduced mental effort and task time, and was preferred because of its dialogue interface, but users felt the need for interruption management. Results show LLM-based interfaces can significantly improve mobile accessibility, and describe the potential of hybrid solutions combining gesture and dialogue modalities towards more inclusive design.

cs.HC

A comprehensive study on causal discovery between degradation paths

Existing studies indicate that complex system degradation is characterized by degradation of multiple dependent parameters. Capturing the dependencies is crucial for accurate degradation modeling and effective degradation control. This work aims to uncover these dependencies through causal analysis, focusing on pairwise causal discovery. Firstly, considering the steady-state characteristic of physical dependencies between parameters, a causal discovery strategy using degradation increments is proposed combined with non-temporal causal discovery techniques. Then, five types of non-temporal causal discovery techniques, including constraint-based, score-based, functional causal model-based, gradient-based and the emerging ordering-based technique, are selected as benchmark methods to identify the most suitable approach. Numerical studies based on Wiener process are first conducted to investigate the method effectiveness on both independent and causally dependent degradation paths. Additionally, sensitivity analysis is performed to evaluate how degradation process characteristics affect the accuracy of causal discovery. Then, two engineering applications are given to show the practical applicability of the approach, including a second-order multiple-feedback band pass filter and a turbofan engine. Our findings indicate that the proposed strategy, which uses degradation increments, outperforms methods that rely on raw degradation data. Among all evaluated techniques, stable Peter-Clark and greedy equivalence search exhibit robust and accurate performance across both numerical and engineering cases, which are recommended for causal discovery between degradation paths. The code is available on GitHub: https://github.com/dirge1/causal_deg_data.

stat.AP

SynCell: Contextualized Drug Synergy Prediction

Drug synergy is profoundly influenced by cellular context, as variations in protein interaction landscapes and pathway activities across cell types reshape how drugs act in combination. Most existing models overlook this heterogeneity, relying on static or bulk-level protein-protein interaction (PPI) networks that ignore cell-specific molecular wiring. The availability of large-scale transcriptomic data now enables the reconstruction of cell-line-resolved interactomes, offering a new foundation for contextualized drug synergy modeling. Here we present SynCell, a Contextualized Drug Synergy framework that integrates drug-protein, protein-protein, and protein-cell line relations within a unified graph architecture. SynCell leverages cell-line-specific PPI networks to embed the molecular context in which drugs act, and employs graph convolutional learning to model how pharmacological effects propagate through cell-specific signaling networks. This formulation treats synergy prediction as a cell-line-contextualized drug-drug interaction problem. Across the large-scale DrugCombDB benchmark, SynCell consistently outperforms state-of-the-art baselines - including DeepSynergy, HypergraphSynergy, HERMES, BAITSAO, DTF, and NHP - particularly in predicting synergies involving unseen drugs or novel cell lines. When benchmarked against these seven methods, SynCell demonstrates substantial gains in generalization and biological interpretability, confirming that contextualizing PPIs with cell-line resolution is indispensable for accurate synergy prediction.

q-bio.QM

Robust Adaptive Learning Control for a Class of Non-affine Nonlinear Systems

We address the tracking problem for a class of uncertain non-affine nonlinear systems with high relative degrees, performing non-repetitive tasks. We propose a rigorously proven, robust adaptive learning control scheme that relies on a gradient descent parameter adaptation law to handle the unknown time-varying parameters of the system, along with a state estimator that estimates the unmeasurable state variables. Furthermore, despite the inherently complex nature of the non-affine system, we provide an explicit iterative computation method to facilitate the implementation of the proposed control scheme. The paper includes a thorough analysis of the performance of the proposed control strategy, and simulation results are presented to demonstrate the effectiveness of the approach.

eess.SY

HyperADRs: A Hierarchical Hypergraph Framework for Drug-Gene-ADR Prediction

Adverse drug reactions (ADRs) are a major barrier to safe and effective pharmacotherapy and increasingly reflect higher order interactions between drugs, genetic background, and clinical phenotypes. Existing graph based approaches usually predict ADRs as properties of drugs or drug pairs, leaving the causal gene implicit and limiting their value for pharmacogenomic decision making. We introduce HyperADRs, a hierarchical hypergraph framework that predicts ADR risk at the level of drug-gene-ADR triads. Starting from curated pharmacogenomic annotations in PharmGKB and the pharmacogenomics subdatabase of DrugBank, we construct high confidence triplets and integrate them with auxiliary molecular, functional, and disease relations from precision-medicine-oriented knowledge graphs. Drugs, genes, and ADR concepts are embedded with modality appropriate pretrained models (UniMol, ESM2, SapBERT) and propagated through a hypergraph convolutional network. A FiLM based, query conditioned contrastive learning module learns context specific representations so that, given any two entities, the model retrieves the correct third entity against many candidates. To improve robustness and interpretability, we propose a nine category ADR macro system scheme that reduces large heterogeneous "other" bins while aligning with organ system reasoning in clinical pharmacology. Across drug-, gene-, and ADR-held-out evaluations on PharmGKB, HyperADRs matches or exceeds strong baselines on ranking based metrics. When trained on PharmGKB and tested on unseen DrugBank triplets, HyperADRs maintains its ranking advantage, indicating that the learned representations capture transferable biological mechanisms and can support mechanistically grounded pharmacogenomic hypothesis generation.

q-bio.QM

Loud-loss: A Perceptually Motivated Loss Function for Speech Enhancement Based on Equal-Loudness Contours

The mean squared error (MSE) is a ubiquitous loss function for speech enhancement, but its problem is that the error cannot reflect the auditory perception quality. This is because MSE causes models to over-emphasize low-frequency components which has high energy, leading to the inadequate modeling of perceptually important high-frequency information. To overcome this limitation, we propose a perceptually-weighted loss function grounded in psychoacoustic principles. Specifically, it leverages equal-loudness contours to assign frequency-dependent weights to the reconstruction error, thereby penalizing deviations in a way aligning with human auditory sensitivity. The proposed loss is model-agnostic and flexible, demonstrating strong generality. Experiments on the VoiceBank+DEMAND dataset show that replacing MSE with our loss in a GTCRN model elevates the WB-PESQ score from 2.17 to 2.93-a significant improvement in perceptual quality.

cs.SD

StripRFNet: A Strip Receptive Field and Shape-Aware Network for Road Damage Detection

Well-maintained road networks are crucial for achieving Sustainable Development Goal (SDG) 11. Road surface damage not only threatens traffic safety but also hinders sustainable urban development. Accurate detection, however, remains challenging due to the diverse shapes of damages, the difficulty of capturing slender cracks with high aspect ratios, and the high error rates in small-scale damage recognition. To address these issues, we propose StripRFNet, a novel deep neural network comprising three modules: (1) a Shape Perception Module (SPM) that enhances shape discrimination via large separable kernel attention (LSKA) in multi-scale feature aggregation; (2) a Strip Receptive Field Module (SRFM) that employs large strip convolutions and pooling to capture features of slender cracks; and (3) a Small-Scale Enhancement Module (SSEM) that leverages a high-resolution P2 feature map, a dedicated detection head, and dynamic upsampling to improve small-object detection. Experiments on the RDD2022 benchmark show that StripRFNet surpasses existing methods. On the Chinese subset, it improves F1-score, mAP50, and mAP50:95 by 4.4, 2.9, and 3.4 percentage points over the baseline, respectively. On the full dataset, it achieves the highest F1-score of 80.33% compared with CRDDC'2022 participants and ORDDC'2024 Phase 2 results, while maintaining competitive inference speed. These results demonstrate that StripRFNet achieves state-of-the-art accuracy and real-time efficiency, offering a promising tool for intelligent road maintenance and sustainable infrastructure management.

cs.CV

Photon loss effects on light-mediated non-Gaussian entangled Bose-Einstein condensates projecting with different photon measurement outcomes

The theory of quantum information processing for macroscopic qubits is based on the fact that every macroscopic qubit has a conserved number of particles. However, from an experimental point of view, every such qubit experiences processes of decoherence that impact the possibilities for entanglement generation between such qubits and use in quantum information processing efficiently. One of the most prospective methods for generating entanglement between distant atomic BECs is quantum nondemolition measurements. Here, we study how the effects of photon measurement impact the entanglement when photon loss decoherence is included. We employ the thermally entangled state representation (TESR) and integral within the ordered operator(IWOP) approach to obtain the accurate density matrix in a photon loss channel. We demonstrate that varying outcomes of photon number measurements lead to the generation of distinct entangled states, each exhibiting unique characteristics. We find that using the Hofmann-Takeuchi and Duan-Giedke-Cirac-Zoller criterion provides advantages in entanglement detection compared to the Wineland squeezing and EPR steering criterion in such settings.

quant-ph

Heterogeneous Entity Representation for Medicinal Synergy Prediction

Medicinal synergy prediction is a powerful tool in drug discovery and development that harnesses the principles of combination therapy to enhance therapeutic outcomes by improving efficacy, reducing toxicity, and preventing drug resistance. While a myriad of computational methods has emerged for predicting synergistic drug combinations, a large portion of them may overlook the intricate, yet critical relationships between various entities in drug interaction networks, such as drugs, cell lines, and diseases. These relationships are complex and multidimensional, requiring sophisticated modeling to capture nuanced interplay that can significantly influence therapeutic efficacy. We introduce a salient deep hypergraph learning method, namely, Heterogeneous Entity Representation for MEdicinal Synergy prediction (HERMES), to predict anti-cancer drug synergy. HERMES integrates heterogeneous data sources, encompassing drug, cell line, and disease information, to provide a comprehensive understanding of the interactions involved. By leveraging advanced hypergraph neural networks with gated residual mechanisms, HERMES can effectively learn complex relationships/interactions within the data. Our results show HERMES demonstrates state-of-the-art performance, particularly in forecasting new drug combinations, significantly surpassing previous methods. This advancement underscores the potential of HERMES to facilitate more effective and precise drug combination predictions, thereby enhancing the development of novel therapeutic strategies.

cs.CE

Multi-scale Semantic Prior Features Guided Deep Neural Network for Urban Street-view Image

Street-view image has been widely applied as a crucial mobile mapping data source. The inpainting of street-view images is a critical step for street-view image processing, not only for the privacy protection, but also for the urban environment mapping applications. This paper presents a novel Deep Neural Network (DNN), multi-scale semantic prior Feature guided image inpainting Network (MFN) for inpainting street-view images, which generate static street-view images without moving objects (e.g., pedestrians, vehicles). To enhance global context understanding, a semantic prior prompter is introduced to learn rich semantic priors from large pre-trained model. We design the prompter by stacking multiple Semantic Pyramid Aggregation (SPA) modules, capturing a broad range of visual feature patterns. A semantic-enhanced image generator with a decoder is proposed that incorporates a novel cascaded Learnable Prior Transferring (LPT) module at each scale level. For each decoder block, an attention transfer mechanism is applied to capture long-term dependencies, and the semantic prior features are fused with the image features to restore plausible structure in an adaptive manner. Additionally, a background-aware data processing scheme is adopted to prevent the generation of hallucinated objects within holes. Experiments on Apolloscapes and Cityscapes datasets demonstrate better performance than state-of-the-art methods, with MAE, and LPIPS showing improvements of about 9.5% and 41.07% respectively. Visual comparison survey among multi-group person is also conducted to provide performance evaluation, and the results suggest that the proposed MFN offers a promising solution for privacy protection and generate more reliable scene for urban applications with street-view images.

cs.CV

Design and Implementation of mmWave Surface Wave Enabled Fluid Antennas and Experimental Results for Fluid Antenna Multiple Access

While multiple-input multiple-output (MIMO) technologies continue to advance, concerns arise as to how MIMO can remain scalable if more users are to be accommodated with an increasing number of antennas at the base station (BS) in the upcoming sixth generation (6G). Recently, the concept of fluid antenna system (FAS) has emerged, which promotes position flexibility to enable transmitter channel state information (CSI) free spatial multiple access on one radio frequency (RF) chain. On the theoretical side, the fluid antenna multiple access (FAMA) approach offers a scalable alternative to massive MIMO spatial multiplexing. However, FAMA lacks experimental validation and the hardware implementation of FAS remains a mysterious approach. The aim of this paper is to provide a novel hardware design for FAS and evaluate the performance of FAMA using experimental data. Our FAS design is based on a dynamically reconfigurable "fluid" radiator which is capable of adjusting its position within a predefined space. One single-channel fluid antenna (SCFA) and one double-channel fluid antenna (DCFA) are designed, electromagnetically simulated, fabricated, and measured. The measured radiation patterns of prototypes are imported into channel and network models for evaluating their performance in FAMA. The experimental results demonstrate that in the 5G millimeter-wave (mmWave) bands (24-30 GHz), the FAS prototypes can vary their gain up to an averaged value of 11 dBi. In the case of 4-user FAMA, the double-channel FAS can significantly reduce outage probability by 57% and increases the multiplexing gain to 2.27 when compared to a static omnidirectional antenna.

eess.SP

Optical and atomic decoherence in entangled atomic ensembles generated by quantum nondemolition measurements

We study the effects of decoherence in the form of optical phase diffusion, photon loss and gain, and atomic dephasing in entangled atomic ensembles produced via quantum nondemolition (QND) measurements. For the optical decoherence channels, we use the technique of integration within ordered operators (IWOP) to obtain the Kraus operators that describe the decoherence. We analyze the effect of different decoherence channels on a variety of quantities such as the variances of the spin operators, entanglement and correlation criteria, logarithmic negativity, and the Bell-CHSH inequality. We generally find a smooth decay of correlations and entanglement in the presence of decoherence. We find that various quantities retain showing non-classical properties under all three types of decoherence, in the short interaction time range. Our results show that such QND measurements are one of the most promising methods for entanglement generation between two Bose-Einstein condensates.

quant-ph

Decoherence effects in quantum nondemolition measurement induced entanglement between Bose-Einstein condensates

We study the robustness of quantum nondemolition (QND) measurement-induced entanglement between Bose-Einstein Condensates (BECs). We consider an experimental scheme where two BECs are placed in the paths of a Mach-Zehnder interferometer, and a QND interaction creates entanglement between coherent light and the atoms. We analyze the two dominant channels of decoherence, atomic dephasing and photon loss on the entangled states produced by this scheme. We calculate the effect of dephasing on the variance and expectation values of the spin operators, entanglement, and correlation criteria. Our analysis does not use the Holstein-Primakoff approximation and is capable of modeling long light-atom interaction times, producing non-Gaussian states beyond the two-mode squeezed states. In the presence of dephasing, the entangled states are robust in the macroscopic limit as long as the dimensionless interaction time is less than $ 1/\sqrt{N}$, where $ N $ is the number of atoms in the BEC. For photon loss, the entangled states generated by long interaction times show remarkable robustness that makes the scheme promising for various quantum information applications.

physics.atom-ph