SearcharxivSearch

arXiv subjects

Nan Zhao

Publications and source records attributed to Nan Zhao.

At least 19 recordsLinked to original sources

Intelligent Disruption: Undetectable Attacks on Wireless Autoencoders

Adversarial attacks can degrade the legitimate decision performance in wireless autoencoder communications. However, in complex scenarios with multiple adversaries, the cumulative leakage interference (CLI) caused by the multiple parallel attacks increases the chance of detecting the attacks, while dynamical environments also make the fixed attack strategies difficult to have stable effectiveness. To jointly enhance the undetectability, aggressivity and adaptability of adversarial attacks, we propose a deep learning based intelligent attack framework. Specifically, considering the CLI caused by the multiple parallel attacks, a deep neural network based transmit power control is established to reduce the interference leakage by regulating the transmit power of these adversaries, thereby improving the undetectability. Furthermore, to enhance the attack effectiveness and stability in the dynamic environment, the conditional generative adversarial attack is further developed. The generator takes the attack channel information as the conditional input to produce the perturbating signals to mislead the discriminator by making the attacked received signals resemble the clean received signals, while the discriminator distinguishes between the two under the same condition. Through the adversarial training, the generator can learn to create adaptive perturbating signals with enhanced attack performance. Simulation results demonstrate that the proposed framework outperforms benchmarks in terms of attack undetectability, aggressivity and adaptability.

cs.CR

Uniboost: Global Coordination with Value Alignment for Fair and Efficient Traffic Allocation

With the rapid evolution of internet services, recommendation systems have become indispensable. In particular, the blending (re-ranking) stage plays a pivotal role in allocating traffic across diverse business objectives. However, existing approaches often suffer from coupled allocation plans, score inflation, and a lack of interpretability. To address these challenges, we propose Uniboost, a unified traffic allocation framework. Uniboost introduces a posterior value alignment mechanism that calibrates abstract model scores to anchor metrics with explicit business semantics, significantly enhancing interpretability. Furthermore, it employs an independent linear boosting paradigm to decouple complex weighting schemes, enabling precise attribution of each plan's contribution. We validate the effectiveness of Uniboost through online A/B tests and in-depth data analysis, demonstrating three key findings: 1) Reducing the overall weight of weighted scores effectively mitigates unintended business interference, yielding a more efficient micro-level traffic allocation strategy; 2) Post-hoc analyses and aggregated dashboards provide intuitive, macro-level insights that guide the design of the overall traffic allocation mechanism; 3) The proposed "Effective Completion Score" serves as an easily obtainable post-metric that offers a reliable anchor for content recommendation pipelines. Collectively, our experiments show that Uniboost not only improves traffic allocation efficiency and recommendation performance at the micro level but also provides macro-level guidance for system iteration. Thus, this work provides an efficient and controllable traffic regulation solution for large-scale industrial recommendation systems.

cs.IR

G-type antiferromagnetic structure in Rb1-xV2Te2O

Altermagnetism, known for its non-relativistic spin-split band structures with yet compensated moments, is being intensively investigated. Discovering new altermagnetic materials with characteristics suitable for practical use remains an important ongoing task. Recently a metallic room-temperature altermagnet candidate Rb1-xV2Te2O with a layered structure and d-wave spin symmetry has been reported based on experimental results from the spin-resolved photoemission spectroscopy and scanning tunnelling microscopy/spectroscopy (STM/STS) measurements. Here we report neutron powder diffraction (NPD) investigations on the magnetic structure of Rb1-xV2Te2O, which shows a G-type antiferromagnetic structure below the transition temperature of 337 K. The result is different from the original theoretical expectation, which might lead to new insights on the physics of this altermagnet candidate.

cond-mat.mtrl-sci

Optical pumping of alkali-metal vapor in the quasi-high-pressure regime

Optical pumping is fundamental to high-precision measurement using thermal alkali-metal atoms in vapor cells. In applications such as atomic magnetometry, buffer gases (e.g., $\mathrm{N}_2$ or $\mathrm{He}$) at specific pressures are introduced to quench fluorescence and mitigate wall relaxation. In the high-pressure limit (e.g., the $\mathrm{N}_2$ pressure $p_{\mathrm{N}_2}> 1$~atm), where collisional broadening exceeds hyperfine splittings of the atoms, optical pumping theory provides a clear description of the angular momentum exchange between photons and atomic spins. However, in many magnetic sensing scenarios, the high-pressure approximation becomes inadequate as its pressure conditions are not strictly satisfied. Consequently, an explicit description of optical pumping under realistic pressures is critical for selecting operating points and enhancing system performance. To address this, we develop a unified theoretical framework of optical pumping in the quasi-high-pressure regime, where collisional broadening is comparable to the ground-state hyperfine splitting. We demonstrate that light absorption, spin polarization, and magnetic-resonance linewidth in this regime differ significantly from those predicted by the high-pressure limit and offer favorable operating conditions. Our study extends conventional modeling and offers critical guidance for atomic magnetometry operating under realistic buffer gas pressures.

physics.atom-ph

Domain-Expert-Guided Hybrid Mixture-of-Experts for Medical AI: Integrating Data-Driven Learning with Clinical Priors

Mixture-of-Experts (MoE) models increase representational capacity with modest computational cost, but their effectiveness in specialized domains such as medicine is limited by small datasets. In contrast, clinical practice offers rich expert knowledge, such as physician gaze patterns and diagnostic heuristics, that models cannot reliably learn from limited data. Combining data-driven experts, which capture novel patterns, with domain-expert-guided experts, which encode accumulated clinical insights, provides complementary strengths for robust and clinically meaningful learning. To this end, we propose Domain-Knowledge-Guided Hybrid MoE (DKGH-MoE), a plug-and-play and interpretable module that unifies data-driven learning with domain expertise. DKGH-MoE integrates a data-driven MoE to extract novel features from raw imaging data, and a domain-expert-guided MoE incorporates clinical priors, specifically clinician eye-gaze cues, to emphasize regions of high diagnostic relevance. By integrating domain expert insights with data-driven features, DKGH-MoE improves both performance and interpretability.

cs.CV

Gyral-Sulcal-Net: An Integrated Network Representation of Brain Folding Patterns

Our brain functions as a complex communication network, and studying it from a network perspective offers valuable insights into its organizational principles and links to cognitive functions and brain disorders. However, most current network studies typically use brain regions as nodes, often overlooking the intricate folding patterns of finer-scale anatomical landmarks within these regions. In this study, we introduce a novel approach to integrate the brain's two primary folding patterns - gyri and sulci - into a unified network termed the Gyral-Sulcal-Net (GS-Net), in which three different types of finer-scale landmarks have been successfully identified. We evaluated the proposed GS-Net across multiple datasets, comprising over 1,600 brains, spanning different age groups (from 34 gestational weeks to elderly adults) and cohorts (healthy brains and those with pathological conditions). The experimental results demonstrate that the GS-Net can effectively integrate and represent diverse cortical folding patterns from a network perspective. More importantly, this approach offers a promising way for integrating different folding patterns into a unified anatomical brain network, alongside structural and functional networks, providing a comprehensive framework for studying brain networks.

q-bio.NC

AI Chatbots or Human Therapists? Belief-Based Predictors of Mental Health Help-Seeking Intentions in the Age of Generative AI

As generative artificial intelligence (GAI) enters the mental health landscape, questions arise about how individuals weigh AI tools against human therapists. This study examined belief-based predictors of intention to use GAI and therapists across two populations: a university sample (N = 1,155) and a nationally representative adult sample (N = 651). Using paired-sample t-tests following a MANOVA, we found that human therapists were viewed as providing greater emotional support and coping, relationship, and educational skills as well as being able to personalize treatment than GAI chatbots. In turn, GAI support was viewed as being more affordable and accessible. No differences between modalities were found with concerns about privacy, reliability, stigma, mental health literacy or help-seeking norms. Using LASSO regression, we examined how beliefs about each modality jointly shape help-seeking intentions. Across both samples, intentions to use either GAI or human therapists were most strongly associated with perceptions of interpersonal support, including emotional support, relational guidance, and personalization. Barriers differed across modalities: concerns about privacy and reliability were more strongly associated with reduced intention to use GAI, whereas structural constraints, particularly affordability, were more closely linked to human therapy use. These findings extend the Health Belief Model to a dual-modality context, demonstrating that help-seeking decisions reflect a comparative push-pull process in which barriers to one modality redirect users toward the other. Design implications are discussed for developing trustworthy, emotionally resonant GAI tools that complement rather than replace human care.

cs.HC

Rotatable Antenna System Empowered Low-Altitude Economy: Opportunities and Challenges

Low-altitude economy (LAE) is an emerging technological paradigm that enables continuous airspace coverage at multiple altitudes by providing highly reliable data connectivity for numerous low-altitude applications. However, existing networks cannot sufficiently support LAE development, as current base stations (BSs) are primarily designed for terrestrial users and lack the capability to provide continuous coverage at low altitudes. To overcome these challenges, rotatable antenna system (RAS) is introduced in LAE, enabling flexible beamforming by dynamically adjusting the boresight of directional antennas to extend low-altitude coverage and enhance the stability of data transmission. In this article, we first provide an overview of RAS-empowered LAE applications, including low-altitude communication, sensing, control, and computation. Then, we present two practical RAS deployment strategies for LAE scenarios, namely RAS-aided multi-BS and multi-unmanned aerial vehicle (UAV) cooperative coverages, as well as provide detailed discussions on their system architectures and performance benefits. Additionally, key design issues of RAS in LAE are discussed, including channel modeling and estimation, cellular access and interference cancellation, as well as RAS configuration and boresight optimization. Finally, we demonstrate the performance gains of RAS in LAE networks through experimental and simulation results.

eess.SY

A Stroke-Level Large-Scale Database of Chinese Character Handwriting and the OpenHandWrite_Toolbox for Handwriting Research

Understanding what linguistic components (e.g., phonological, semantic, and orthographic systems) modulate Chinese handwriting at the character, radical, and stroke levels remains an important yet understudied topic. Additionally, there is a lack of comprehensive tools for capturing and batch-processing fine-grained handwriting data. To address these issues, we constructed a large-scale handwriting database in which 42 Chinese speakers for each handwriting 1200 characters in a handwriting-to-dictation task. Additionally, we enhanced the existing handwriting package and provided comprehensive documentation for the upgraded OpenHandWrite_Toolbox, which can easily modify the experimental design, capture the stroke-level handwriting trajectory, and batch-process handwriting measurements (e.g., latency, duration, and pen-pressure). In analysing our large-scale database, multiple regression results show that orthographic predictors impact handwriting preparation and execution across character, radical, and stroke levels. Phonological factors also influence execution at all three levels. Importantly, these lexical effects demonstrate hierarchical attenuation - they were most pronounced at the character level, followed by the radical, and were weakest at the stroke levels. These findings demonstrate that handwriting preparation and execution at the radical and stroke levels are closely intertwined with linguistic components. This database and toolbox offer valuable resources for future psycholinguistic and neurolinguistic research on the handwriting of characters and sub-characters across different languages.

cs.CV

Generative AI-Empowered Secure Communications in Space-Air-Ground Integrated Networks: A Survey and Tutorial

Space-air-ground integrated networks (SAGINs) face unprecedented security challenges due to their inherent characteristics, such as multidimensional heterogeneity and dynamic topologies. These characteristics fundamentally undermine conventional security methods and traditional artificial intelligence (AI)-driven solutions. Generative AI (GAI) is a transformative approach that can safeguard SAGIN security by synthesizing data, understanding semantics, and making autonomous decisions. This survey fills existing review gaps by examining GAI-empowered secure communications across SAGINs. First, we introduce secured SAGINs and highlight GAI's advantages over traditional AI for security defenses. Then, we explain how GAI mitigates failures of authenticity, breaches of confidentiality, tampering of integrity, and disruptions of availability across the physical, data link, and network layers of SAGINs. Three step-by-step tutorials discuss how to apply GAI to solve specific problems using concrete methods, emphasizing its generative paradigm beyond traditional AI. Finally, we outline open issues and future research directions, including lightweight deployment, adversarial robustness, and cross-domain governance, to provide major insights into GAI's role in shaping next-generation SAGIN security.

cs.CR

Impact of Heavy Noble Gases on the Magnetic Resonance Linewidth of Alkali-Metal Atoms: A Theoretical Study

Nuclear magnetic resonance gyroscopes (NMRGs) employ noble-gas nuclear spins as inertial sensors and alkali-metal atoms as in-situ magnetometers. Heavy noble gases, particularly xenon, are widely used due to their large nuclear spin and strong spin-exchange coupling with alkali-metal atoms. However, their presence introduces additional collisional mechanisms that affect the alkali-metal magnetic resonance linewidth, thereby influencing magnetometer sensitivity and overall gyro performance. In this work, we develop a theoretical framework based on the density matrix formalism and master equation approach to quantitatively study how xenon-induced two-body and three-body interactions modify the linewidth of alkali-metal atoms under realistic NMRG conditions. Our analysis reveals that Xe atoms primarily broaden the linewidth via binary spindestruction collisions and van der Waals (vdW)-mediated F-damping processes, while the effect of Xe nuclear polarization is negligible at the ~1% level. We further demonstrate that nitrogen buffer gas plays a dual role: it directly contributes to alkali-metal spin relaxation through binary collisions and indirectly modulates vdW collision rates by altering molecular lifetimes. The interplay between these processes leads to an optimal nitrogen density that minimizes the linewidth. Additionally, we identify a temperature threshold above which light-narrowing emerges, with this threshold increasing alongside Xe density. These findings provide theoretical insight for optimizing spin relaxation control in alkali-metal magnetometers and improving NMRG performance.

quant-ph

Field-induced magnetic order in DyTa$_7$O$_{19}$ with two-dimensional pseudospin-$\frac{1}{2}$ triangular lattice

The magnetic ground state of geometrically frustrated antiferromagnet attracts great research interests due to the possibility to realize novel quantum magnetic state such as a quantum spin liquid. Here we present a comprehensive magnetic characterization of DyTa$_7$O$_{19}$ with ideal two-dimensional triangular lattice. DyTa$_7$O$_{19}$ exhibits $c$-axis single-ion magnetic anisotropy. Although long-range magnetic order is not observed down to 100 mK under zero field, by applying a small magnetic field ($\sim$0.1 T), a magnetically ordered state with net magnetization of $M_s$/3 below $T_m$=0.14 K is identified ($M_s$ denotes the saturated magnetization). We argue that this state is an up-up-down magnetic structure phase driven by the dipole-dipole interactions between Ising-like spins of Dy$^{3+}$ in a two-dimensional triangular lattice, since its ordering temperature and temperature-field phase diagram can be well explained by the theoretical calculations based on dipolar interactions. DyTa$_7$O$_{19}$ could be viewed as a rare material platform that realizing pure Ising-like dipolar interaction in a geometrically frustrated lattice.

cond-mat.str-el

Quantum Fluctuation-enhanced Milli-Kelvin Magnetic Refrigeration in Triangular Lattice Magnet GdBO3

Rare-earth-based triangular lattice antiferromagnets, with strong quantum fluctuations and weak magnetic interactions, can often retain large magnetic entropy down to very low temperatures, making them excellent candidates for magnetic refrigeration at ultra-low temperatures. These materials exhibit a substantial magnetocaloric effect (MCE) due to enhanced spin fluctuations, particularly near quantum critical points, which leads to significant changes in magnetic entropy. This study reports on the crystal growth, structure, magnetism, and MCE of a Gd-based triangular lattice material, GdBO3, characterized by a large spin quantum number (S = 7/2). Successive phase transitions (T1 = 0.52 K, T2 = 0.88 K, and T3 = 1.77 K) were observed in zero-field specific heat measurements. Furthermore, thermal dynamic analysis under external magnetic fields identified five distinct phase regions and three quantum critical points for GdBO3. Due to its broad specific heat features and the high density of magnetic Gd3+ ions, we achieved a minimum temperature of 50 mK near the field-induced quantum critical point, using a custom-designed GdBO3-based adiabatic demagnetization refrigerator. Our findings reveal significant quantum fluctuations below 2 K, demonstrating GdBO3's potential for milli-Kelvin magnetic cooling applications.

cond-mat.str-el

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and human-like conversations. Previous models for voice interactions are categorized as native and aligned. Native models integrate speech and text processing in one framework but struggle with issues like differing sequence lengths and insufficient pre-training. Aligned models maintain text LLM capabilities but are often limited by small datasets and a narrow focus on speech tasks. In this work, we introduce MinMo, a Multimodal Large Language Model with approximately 8B parameters for seamless voice interaction. We address the main limitations of prior aligned multimodal models. We train MinMo through multiple stages of speech-to-text alignment, text-to-speech alignment, speech-to-speech alignment, and duplex interaction alignment, on 1.4 million hours of diverse speech data and a broad range of speech tasks. After the multi-stage training, MinMo achieves state-of-the-art performance across various benchmarks for voice comprehension and generation while maintaining the capabilities of text LLMs, and also facilitates full-duplex conversation, that is, simultaneous two-way communication between the user and the system. Moreover, we propose a novel and simple voice decoder that outperforms prior models in voice generation. The enhanced instruction-following capabilities of MinMo supports controlling speech generation based on user instructions, with various nuances including emotions, dialects, and speaking rates, and mimicking specific voices. For MinMo, the speech-to-text latency is approximately 100ms, full-duplex latency is approximately 600ms in theory and 800ms in practice. The MinMo project web page is https://funaudiollm.github.io/minmo, and the code and models will be released soon.

cs.CL

Distributed satellite information networks: Architecture, enabling technologies, and trends

Driven by the vision of ubiquitous connectivity and wireless intelligence, the evolution of ultra-dense constellation-based satellite-integrated Internet is underway, now taking preliminary shape. Nevertheless, the entrenched institutional silos and limited, nonrenewable heterogeneous network resources leave current satellite systems struggling to accommodate the escalating demands of next-generation intelligent applications. In this context, the distributed satellite information networks (DSIN), exemplified by the cohesive clustered satellites system, have emerged as an innovative architecture, bridging information gaps across diverse satellite systems, such as communication, navigation, and remote sensing, and establishing a unified, open information network paradigm to support resilient space information services. This survey first provides a profound discussion about innovative network architectures of DSIN, encompassing distributed regenerative satellite network architecture, distributed satellite computing network architecture, and reconfigurable satellite formation flying, to enable flexible and scalable communication, computing and control. The DSIN faces challenges from network heterogeneity, unpredictable channel dynamics, sparse resources, and decentralized collaboration frameworks. To address these issues, a series of enabling technologies is identified, including channel modeling and estimation, cloud-native distributed MIMO cooperation, grant-free massive access, network routing, and the proper combination of all these diversity techniques. Furthermore, to heighten the overall resource efficiency, the cross-layer optimization techniques are further developed to meet upper-layer deterministic, adaptive and secure information services requirements. In addition, emerging research directions and new opportunities are highlighted on the way to achieving the DSIN vision.

cs.IT

Frequency Diverse Array-enabled RIS-aided Integrated Sensing and Communication

Integrated sensing and communication (ISAC) has been envisioned as a prospective technology to enable ubiquitous sensing and communications in next-generation wireless networks. In contrast to existing works on reconfigurable intelligent surface (RIS) aided ISAC systems using conventional phased arrays (PAs), this paper investigates a frequency diverse array (FDA)-enabled RIS-aided ISAC system, where the FDA aims to provide a distance-angle-dependent beampattern to effectively suppress the clutter, and RIS is employed to establish high-quality links between the BS and users/target. We aim to maximize sum rate by jointly optimizing the BS transmit beamforming vectors, the covariance matrix of the dedicated radar signal, the RIS phase shift matrix, the FDA frequency offsets and the radar receive equalizer, while guaranteeing the required signal-to-clutter-plus-noise ratio (SCNR) of the radar echo signal. To tackle this challenging problem, we first theoretically prove that the dedicated radar signal is unnecessary for enhancing target sensing performance, based on which the original problem is much simplified. Then, we turn our attention to the single-user single-target (SUST) scenario to demonstrate that the FDA-RIS-aided ISAC system always achieves a higher SCNR than its PA-RIS-aided counterpart. Moreover, it is revealed that the SCNR increment exhibits linear growth with the BS transmit power and the number of BS receive antennas. In order to effectively solve this simplified problem, we leverage the fractional programming (FP) theory and subsequently develop an efficient alternating optimization (AO) algorithm based on symmetric alternating direction method of multipliers (SADMM) and successive convex approximation (SCA) techniques. Numerical results demonstrate the superior performance of our proposed algorithm in terms of sum rate and radar SCNR.

cs.IT

Magnetic Resonance Linewidth of Alkali-Metal Vapor in Unresolved Zeeman Resonance Regime

The study of magnetic resonance linewidth is crucial in magnetic resonance physics and its applications. Previous studies focused on the linewidth of alkali metal atoms within the spin-exchange relaxation-free regime near zero magnetic field and in strong magnetic fields where Zeeman resonances are well resolved due to the quadratic Zeeman effect. However, the linewidth in the unresolved Zeeman resonance regime, which is prevalent in various magnetometer and comagnetometer applications, is not well understood. To address this, we developed a theoretical framework based on the master equation for alkali metal atoms and solved it under the rotating wave approximation and weak driving conditions. Our numerical calculations and analytical expressions reveal that the light-narrowing effect occurs only when the ratio of the spin exchange rate to the spin destruction rate exceeds a critical value. Additionally, we show that the linewidth in the unresolved Zeeman resonance regime is significantly influenced by the mutual coupling of quantum coherence between different Zeeman sublevels. These findings provide a theoretical tool for understanding spin relaxation in alkali-metal atoms and optimizing the performance of atomic magnetometers and comagnetometers operating in this regime.

physics.atom-ph

FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative models: SenseVoice, which handles multilingual speech recognition, emotion recognition, and audio event detection; and CosyVoice, which facilitates natural speech generation with control over multiple languages, timbre, speaking style, and speaker identity. SenseVoice-Small delivers exceptionally low-latency ASR for 5 languages, and SenseVoice-Large supports high-precision ASR for over 50 languages, while CosyVoice excels in multi-lingual voice generation, zero-shot in-context learning, cross-lingual voice cloning, and instruction-following capabilities. The models related to SenseVoice and CosyVoice have been open-sourced on Modelscope and Huggingface, along with the corresponding training, inference, and fine-tuning codes released on GitHub. By integrating these models with LLMs, FunAudioLLM enables applications such as speech-to-speech translation, emotional voice chat, interactive podcasts, and expressive audiobook narration, thereby pushing the boundaries of voice interaction technology. Demos are available at https://fun-audio-llm.github.io, and the code can be accessed at https://github.com/FunAudioLLM.

cs.SD