SearcharxivSearch

arXiv subjects

Cong Zhou

Publications and source records attributed to Cong Zhou.

At least 19 recordsLinked to original sources

No evidence for a supermassive black hole binary in GSN 069

Quasi-periodic eruptions (QPEs) are recurrent soft X-ray flares from galactic nuclei and provide a new time-domain probe of stellar-mass objects (SMOs) orbiting supermassive black holes (SMBHs). In an extreme-mass-ratio inspiral (EMRI) system interacting with an accretion disk, QPEs are produced when the SMO repeatedly crosses an accretion disk, so that the eruption times trace the orbital motion of the EMRI. We investigate whether such timing information can be used to probe a more distant SMBH companion. We develop two complementary diagnostics: (1) the motion of the EMRI host SMBH around the SMBH-binary (SMBHB) center of mass induces a light-travel-time modulation in the observed QPE arrival times, specifically an \emph{in-phase} modulation in arrival times of even and odd eruptions; (2) if the QPE source contains a surviving stellar orbiter, the external SMBH must not drive the SMO into tidal disruption through eccentricity excitation by the von Zeipel--Lidov--Kozai (ZLK) mechanism. Using GSN 069 as an example, we find \emph{no} in-phase modulation in the QPE timing (i.e., no evidence for a SMBHB) and constrain the excluded parameter space of the companion SMBH. These results demonstrate that QPE timing and stellar survival offer complementary routes for constraining otherwise hidden SMBH companions in nearby galactic nuclei.

astro-ph.HE

A Note on QPE Timing: False Alarms in O-C

O-C timing analysis is a useful diagnostic tool for quasi-periodic eruptions (QPEs), but their interpretation depends sensitively on the integer cycle number assigned to each eruption. In this note, we show that even a small mismatch in the cycle number, $N_{\rm cyc}$, can produce large false signals in O-C diagrams, and \emph{a universal feature of these false signals is a large in-phase sinusoidal modulation between even and odd eruptions.} Therefore, uncertainties in $N_{\rm cyc}$ must be inferred or marginalized over before physical interpretations are attached to O-C. We then apply both O-C and EMRI+disk to GSN 069 and eRO-QPE2. For GSN 069, the timing data favor an anti-phase modulation in even and odd eruptions, consistent with apsidal precession in a low-eccetricity EMRI crossing an equatorial disk. For eRO-QPE2, the data are well described by a near-circular EMRI and a precessing disk.

astro-ph.HE

Contextual Wireless Video Semantic Communication in MIMO-OFDM Systems

This paper proposes a MIMO-OFDM-based context video semantic transmission framework, namely M-CVST, for robust video communication over multi-path multiple-input multiple-output (MIMO) channels. It introduces a context-subcarrier correlation map that aligns video feature context with groups of MIMO subcarriers. To leverage the time-correlated nature of multi-path channels, a recursive subcarrier sampling method paired with time-correlated reference embedding is designed, enabling the use of previously sampled MIMO subcarrier CSI to enhance channel state awareness in the entropy coding model. Numerical results verify the superiority of proposed M-CVST over MIMO multi-path channels compared to other semantic schemes and traditional separated schemes.

cs.MM

Weighted Sum-Rate Maximization for RIS-UAV-assisted Space-Air-Ground Integrated Network with RSMA

In this paper, a rate-splitting multiple access (RSMA) based joint optimization framework for the space-air-ground integrated network (SAGIN) is proposed, where the satellite and base stations employ uniform planar array (UPA) antennas for signal transmission, and unmanned aerial vehicles (UAVs) relay the satellite signals. Earth stations (ESs) and user equipments (UEs) receive signals from satellite and base stations (BSs), respectively, resulting in mutual interference. We first model the channels and signals in this scenario and analyse the interference at BSs and UEs. Then, We formulate a joint optimization problem aimed at maximizing the weighted sum-rate, involving beamforming, RIS-UAV deployment and phase shifts, and rate splitting. However, this problem is highly non-convex. To tackle this challenge, we apply a block coordinate descent (BCD) approach to decompose the problem and employ the weighted minimum mean square error (WMMSE) method to transform the non-convex objective function. For the rate-splitting sub-problem, a greedy algorithm is proposed and a successive convex approximation (SCA) algorithm is used for beamforming. Besides, the alternating direction method of multipliers (ADMM) algorithm is employed for the RIS phase-shift problem with unit-modulus constraints, and an exhaustive search method is adopted for the complex UAV positioning and orientation. Simulation results validate that the proposed algorithm achieves superior performance in terms of user weighted sum-rate.

eess.SP

Low-complexity Design for Beam Coverage in Near-field and Far-field: A Fourier Transform Approach

In this paper, we study efficient beam coverage design for multi-antenna systems in both far-field and near-field cases. To reduce the computational complexity of existing sampling-based optimization methods, we propose a new low-complexity yet efficient beam coverage design. To this end, we first formulate a general beam coverage optimization problem to maximize the worst-case beamforming gain over a target region. For the far-field case, we show that the beam coverage design can be viewed as a spatial-frequency filtering problem, where angular coverage can be achieved by weight-shaping in the antenna domain via an inverse FT, yielding an infinite-length weighting sequence. Under the constraint of a finite number of antennas, a surrogate scheme is proposed by directly truncating this sequence, which inevitably introduces a roll-off effect at the angular boundaries, yielding degraded worst-case beamforming gain. To address this issue, we characterize the finite-antenna-induced roll-off effect, based on which a roll-off-aware design with a protective zoom is developed to ensure a flat beamforming-gain profile within the target angular region. Next, we extend the proposed method to the near-field case. Specifically, by applying a first-order Taylor approximation to the near-field channel steering vector (CSV), the two-dimensional (2D) beam coverage design (in both angle and inverse-range) can be transformed into a 2D inverse FT, leading to a low-complexity beamforming design. Furthermore, an inherent near-field range defocusing effect is observed, indicating that sufficiently wide angular coverage results in range-insensitive beam steering. Finally, numerical results demonstrate that the proposed FT-based approach achieves a comparable worst-case beamforming performance with that of conventional sampling-based optimization methods while significantly reducing the computational complexity.

cs.IT

Near-field Physical Layer Security: Robust Beamforming under Location Uncertainty

In this paper, we study robust beamforming design for near-field physical-layer-security (PLS) systems, where a base station (BS) equipped with an extremely large-scale array (XL-array) serves multiple near-field legitimate users (Bobs) in the presence of multiple near-field eavesdroppers (Eves). Unlike existing works that mostly assume perfect channel state information (CSI) or location information of Eves, we consider a more practical and challenging scenario, where the locations of Bobs are perfectly known, while only imperfect location information of Eves is available at the BS. We first formulate a robust optimization problem to maximize the sum-rate of Bobs while guaranteeing a worst-case limit on the eavesdropping rate under location uncertainty. By transforming Cartesian position errors into the polar domain, we reveal an important near-field angular-error amplification effect: for the same location error, the closer the Eve, the larger the angle error, severely degrading the performance of conventional robust beamforming methods based on imperfect channel state information. To address this issue, we first establish the conditions for which the first-order Taylor approximation of the near-field channel steering vector under location uncertainty is largely accurate. Then, we propose a two-stage robust beamforming method, which first partitions the uncertainty region into multiple fan-shaped sub-regions, followed by the second stage to formulate and solve a refined linear-matrix-inequality (LMI)-based robust beamforming optimization problem. In addition, the proposed method is further extended to scenarios with multiple Bobs and multiple Eves. Finally, numerical results validate that the proposed method achieves a superior trade-off between rate performance and secrecy robustness, hence significantly outperforming existing benchmarks under Eve location uncertainty.

eess.SP

Rotatable IRS-Assisted 6DMA Communications: A Two-timescale Design

Intelligent reflecting surface (IRS) and movable antenna (MA) are promising technologies to enhance wireless communication by reconfiguring channels at the environment and transceiver sides. However, their performance is constrained by practical limitations. To address this, we propose a multi-functional antenna/surface system that leverages their complementary advantages. A rotatable IRS (R-IRS) is deployed to enhance downlink communications from a six-dimensional MA (6DMA)-equipped base station (BS) to multiple single-antenna users. To reduce the complexity of real-time channel estimation and beamforming, we formulate an optimization problem to maximize the average sum-rate using a two-timescale (TTS) transmission protocol. Specifically, the BS antenna configuration (including position and rotation) and IRS rotation and reflection are optimized based on statistical channel state information (S-CSI), while BS transmit beamforming is designed using instantaneous CSI (I-CSI) in the short timescale. We first consider a single-user case and show that the 6DMA at the BS should form a sparse array for multi-beam transmission towards both the IRS and the user, allowing efficient coordination of direct and reflected channels, while the IRS rotation achieves effective multi-path alignment. For the general multi-user case, the optimization problem is non-convex and challenging to solve. To tackle this, we propose an efficient algorithm combining weighted minimum mean-square error (WMMSE) and stochastic successive convex approximation (SSCA) techniques. A low-complexity algorithm is also proposed to reduce computational complexity. Numerical results validate the proposed system, showing significant performance gains by jointly exploiting the spatial degrees of freedom of the 6DMA-BS and R-IRS under the TTS protocol.

cs.IT

MA-enhanced Mixed Near-field and Far-field Covert Communications

In this paper, we propose to employ a modular-based movable extremely large-scale array (XL-array) at Alice for enhancing covert communication performance. Compared with existing work that mostly considered either far-field or near-field covert communications, we consider in this paper a more general and practical mixed-field scenario, where multiple Bobs are located in either the near-field or far-field of Alice, in the presence of multiple near-field Willies. Specifically, we first consider a two-Bob-one-Willie system and show that conventional fixed-position XL-arrays suffer degraded sum-rate performance due to the energy-spread effect in mixed-field systems, which, however, can be greatly improved by subarray movement. On the other hand, for transmission covertness, it is revealed that sufficient angle difference between far-field Bob and Willie as well as adequate range difference between near-field Bob and Willie are necessary for ensuring covertness in fixed-position XL-array systems, while this requirement can be relaxed in movable XL-array systems thanks to flexible channel correlation control between Bobs and Willie. Next, for general system setups, we formulate an optimization problem to maximize the achievable sum-rate under covertness constraint. To solve this non-convex optimization problem, we first decompose it into two subproblems, corresponding to an inner problem for beamforming optimization given positions of subarrays and an outer problem for subarray movement optimization. Although these two subproblems are still non-convex, we obtain their high-quality solutions by using the successive convex approximation technique and devising a customized differential evolution algorithm, respectively. Last, numerical results demonstrate the effectiveness of proposed movable XL-array in balancing sum-rate and covert communication requirements.

eess.SP

Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play

Role-play has become a key testbed for generative models, expanding from text-only dialogue to multimodal interaction. Extending role-play to speech captures prosody, emotion, and delivery, but also poses new evaluation challenges. Current pipelines often use audio large language models (ALLMs) as zero-shot judges, which miss paralinguistic cues, collapse multiple aspects into coarse scores, and rely on synthetic speech references that fail to reflect real-world roles. We present Speech-DRAME, a unified framework that contributes at three levels: (i) Speech-DRAME-EvalBench, an evaluation benchmark with bilingual human-annotated data and protocols for training and testing speech evaluation models (SEMs), (ii) DRAME-Eval, a fine-tuned evaluation model, which substantially outperforms zero-shot and few-shot ALLMs, and (iii) Speech-DRAME-RoleBench, a speech role-play benchmark that leverages DRAME-Eval as an automatic judge to compare speech foundation models (SFMs). Speech-DRAME distinguishes between two complementary evaluation strategies: Archetype Evaluation, a top-down approach measuring adherence to broad role archetypes, and Realism Evaluation, a bottom-up approach grounded in real human speech that emphasizes nuanced role quality. Compared to zero-shot ALLM judges, DRAME-Eval achieves stronger agreement with human ratings (Pearson correlation from 0.480 to 0.629 in archetypes, and 0.390 to 0.625 in realism). By integrating transparent benchmark resources, modeling approaches, and system-level evaluation, Speech-DRAME provides the first comprehensive, reproducible foundation for assessing spoken role-play.

cs.SD

Spatial and temporal study of the post-compressed high-power laser pulses for coherent extreme ultraviolet source development

We compared the performance of two post-compression techniques, a gas-filled hollow-core fiber (HCF) and a multi-pass cell (MPC), using a high-power ytterbium-doped fiber laser. The HCF produced 27 fs pulses from 230 fs inputs at >50% efficiency, whereas the MPC achieved 34 fs pulses with significantly higher efficiency (>88%). Both results aligned well with numerical simulations. Crucially, spatial wavefront analysis revealed that the HCF acts as a modal filter, improving beam quality, whereas the MPC introduces aberrations through cumulative mirror errors. Furthermore, we characterize the photon flux of high harmonic generation driven by the post-compressed pulses from the HCF and MPC. These finding highlights that post-compression technique based on self-phase modulation is efficient for the intensity boosting of femtosecond laser system, providing opportunities for generating high quality extreme ultraviolet (XUV) sources. In addition, further improvement of spatial wavefront quality is suggested using the HCF as a single compressor or output component of the cascade compressor.

physics.optics

Reconstruction and Reenactment Separated Method for Realistic Gaussian Head

In this paper, we explore a reconstruction and reenactment separated framework for 3D Gaussians head, which requires only a single portrait image as input to generate controllable avatar. Specifically, we developed a large-scale one-shot gaussian head generator built upon WebSSL and employed a two-stage training approach that significantly enhances the capabilities of generalization and high-frequency texture reconstruction. During inference, an ultra-lightweight gaussian avatar driven by control signals enables high frame-rate rendering, achieving 90 FPS at a resolution of 512x512. We further demonstrate that the proposed framework follows the scaling law, whereby increasing the parameter scale of the reconstruction module leads to improved performance. Moreover, thanks to the separation design, driving efficiency remains unaffected. Finally, extensive quantitative and qualitative experiments validate that our approach outperforms current state-of-the-art methods.

cs.CV

Joint Frequency-Space Sparse Reconstruction for DOA Estimation under Coherent Sources and Amplitude-Phase Errors

In this letter, we propose a joint frequency-space sparse reconstruction method for direction-of-arrival (DOA) estimation, which effectively addresses the issues arising from the existence of coherent sources and array amplitude-phase errors. Specifically, by using an auxiliary source with known angles, we first construct the real steering vectors (RSVs) based on the spectral peaks of received signals in the frequency domain, which serve as a complete basis matrix for compensation for amplitude-phase errors. Then, we leverage the spectral sparsity of snapshot data in the frequency domain and the spatial sparsity of incident directions to perform the DOA estimation according to the sparse reconstruction method. The proposed method does not require iterative optimization, hence exhibiting low computational complexity. Numerical results demonstrate that the proposed DOA estimation method achieves higher estimation accuracy for coherent sources as compared to various benchmark schemes.

eess.SP

Frequency-switching Array Enhanced Physical-Layer Security in Terahertz Bands: A Movable Antenna Perspective

In this paper, we propose a new frequency-switching array (FSA) to enhance the physical-layer security (PLS) in the presence of multiple eavesdroppers (Eves), where the carrier frequency can be flexibly switched and small frequency offsets can be imposed on each antenna at the secrecy transmitter (Alice).First, we analytically show that by flexibly controlling the carrier frequency parameters, FSAs can effectively form uniform/non-uniform sparse arrays, hence resembling existing mechanically controlled movable antennas (MAs) via the control of inter-antenna spacing and providing additional degree-of-freedom in the beam manipulation.Although the proposed FSA suffers from additional path-gain attenuation in the received signals, it can overcome several hardware and signal processing issues incurred by MAs, such as limited positioning accuracy, extra hardware and energy cost.Then, a secrecy-rate maximization problem is formulated under the constraints on the frequency control.To shed useful insights, we first consider a secrecy-guaranteed problem with a null-steering constraint for which maximum ratio transmission beamformer is considered at Alice and the frequency offsets are set as uniform frequency increment.Interestingly, it is shown that the proposed FSA can flexibly realize null-steering over Eve in both the angular domain and range domain, thereby achieving improved PLS performance.Then, for the general case, we propose an efficient algorithm to solve the formulated non-convex optimization problem by using the block coordinate descent and projected gradient ascent techniques. Finally, numerical results demonstrate that the proposed FSA achieves superior secrecy rate performance over conventional fixed-position array, while it only suffers a slight secrecy rate loss than the existing mechanically controlled MA.

eess.SP

LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation

Most of the current top-down multi-person pose estimation lightweight methods are based on multi-branch parallel pure CNN network architecture, which often struggle to capture the global context required for detecting semantically complex keypoints and are hindered by high latency due to their intricate and redundant structures. In this article, an approximate single-branch lightweight global modeling network (LGM-Pose) is proposed to address these challenges. In the network, a lightweight MobileViM Block is designed with a proposed Lightweight Attentional Representation Module (LARM), which integrates information within and between patches using the Non-Parametric Transformation Operation(NPT-Op) to extract global information. Additionally, a novel Shuffle-Integrated Fusion Module (SFusion) is introduced to effectively integrate multi-scale information, mitigating performance degradation often observed in single-branch structures. Experimental evaluations on the COCO and MPII datasets demonstrate that our approach not only reduces the number of parameters compared to existing mainstream lightweight methods but also achieves superior performance and faster processing speeds.

cs.CV

Physical-Layer Security in Mixed Near-Field and Far-Field Communication Systems

Extremely large-scale arrays (XL-arrays) have emerged as a promising technology to improve the spectrum efficiency and spatial resolution of future wireless systems. Different from existing works that mostly considered physical layer security (PLS) in either the far-field or near-field, we consider in this paper a new and practical scenario, where legitimate users (Bobs) are located in the far-field of a base station (BS) while eavesdroppers (Eves) are located in the near-field for intercepting confidential information at short distance, referred to as the mixed near-field and far-field PLS. Specifically, we formulate an optimization problem to maximize the sum-secrecy-rate of all Bobs by optimizing the power allocation of the BS, subject to the constraint on the total BS transmit power. To shed useful insights, we first consider a one-Bob-one-Eve system and characterize the insecure-transmission region of the Bob in closed form. Interestingly, we show that the insecure-transmission region is significantly \emph{expanded} as compared to that in conventional far-field PLS systems, due to the energy-spread effect in the mixed-field scenario. Then, we further extend the analysis to a two-Bob-one-Eve system. It is revealed that as compared to the one-Bob system, the interferences from the other Bob can be effectively used to weaken the capability of Eve for intercepting signals of target Bobs, thus leading to enhanced secrecy rates. Furthermore, we propose an efficient algorithm to obtain a high-quality solution to the formulated non-convex problem by leveraging the successive convex approximation (SCA) technique. Finally, numerical results demonstrate that our proposed algorithm achieves a higher sum-secrecy-rate than the benchmark scheme where the power allocation is designed based on the (simplified) far-field channel model.

eess.SP

Super-resolution Wideband Beam Training for Near-field Communications with Ultra-low Overhead

In this paper, we propose a super-resolution wideband beam training method for near-field communications, which is able to achieve ultra-low overhead. To this end, we first study the multi-beam characteristic of a sparse uniform linear array (S-ULA) in the wideband. Interestingly, we show that this leads to a new beam pattern property, called rainbow blocks, where the S-ULA generates multiple grating lobes and each grating lobe is further splitted into multiple versions in the wideband due to the well-known beam-split effect. As such, one directional beamformer based on S-ULA is capable of generating multiple rainbow blocks in the wideband, hence significantly extending the beam coverage. Then, by exploiting the beam-split effect in both the frequency and spatial domains, we propose a new three-stage wideband beam training method for extremely large-scale array (XL-array) systems. Specifically, we first sparsely activate a set of antennas at the central of the XL-array and judiciously design the time-delay (TD) parameters to estimate candidate user angles by comparing the received signal powers at the user over subcarriers. Next, to resolve the angular ambiguity introduced by the S-ULA, we activate all antennas in the central subarray and design an efficient subcarrier selection scheme to estimate the true user angle. In the third stage, we resolve the user range at the estimated user angle with high resolution by controlling the splitted beams over subcarriers to simultaneously cover the range domain. Finally, numerical results are provided to demonstrate the effectiveness of proposed wideband beam training scheme, which only needs three pilots in near-field beam training, while achieving near-optimal rate performance.

eess.SP

Dynamical Measurement of Supermassive Black Hole Masses: QPE Timing Method

Quasi-periodic eruptions (QPEs) are intense repeating soft X-ray bursts with recurrence times about a few hours to a few weeks from galactic nuclei. More and more analyses show that (at least a fraction of) QPEs are the result of collisions between a stellar mass object (SMO, a stellar mass black hole or a main sequence star) and an accretion disk around a supermassive black hole (SMBH) in galactic nuclei. Previous studies have shown the possibility of reconstructing the SMO trajectory from QPE timing data, consequently measuring the SMBH mass from tracing a single SMO. In this paper, we construct a comprehensive Bayesian framework for implementing the QPE timing method, explore the optimal QPE observation strategy for measuring SMBH masses, and forecast the measurement precision expected in the era of multi-target X-ray telescope, Chasing All Transients Constellation Hunters (CATCH). Simulations of CATCH observations of GSN 069 and eRO-QPE2 like QPEs confirm the possible applications of the QPE timing method in precise measurement of SMBH masses (and spins), especially in the lower mass end ($\lesssim 10^7 M_\odot$) where QPEs prevail and relevant dynamical timescales are reasonably short to be measured.

astro-ph.HE

Mixed Near-field and Far-field Localization in Extremely Large-scale MIMO Systems

In this paper, we study efficient \emph{mixed near-field and far-field} target localization methods in extremely large-scale multiple-input multiple-output (XL-MIMO) systems Compared with existing works, we address two new challenges in target localization of MIMO communication systems via using decoupled subspace methods, arising from the half-wavelength antenna spacing constraint and \emph{hybrid uniform planar array} (UPA) architectures.To this end, we propose a new three-step mixed-field localization method. First, we reconstruct the equivalent signals received at UPA antennas by judiciously designing analog combining matrices over time with minimum recovery errors.Second, based on recovered signals, we extend the modified multiple signal classification (MUSIC) algorithm to the UPA architectures by constructing a new covariance matrix of a virtual sparse UPA (S-UPA) to decouple the 2D angles and range estimation.Due to the structure of the S-UPA, there exist ambiguous angles when estimating true angles of targets.In the third step, we design an effective classification method to distinguish mixed-field targets, determine true angles of all targets, as well as estimate the ranges of near-field targets.In particular, angular ambiguity is resolved by showing an important fact that the three types of estimated angles (i.e., far-field, near-field, and ambiguous angles) exhibit significantly different patterns in the range-domain MUSIC spectrum.Furthermore, to characterize the estimation error lower-bound, we obtain a matrix closed-form Cram\'er-Rao bounds for mixed-field target localization.Finally, numerical results demonstrate the effectiveness of our proposed mixed-field localization method, which improves target-classification accuracy and achieves a lower root mean square error than various benchmark schemes.

eess.SP