SearcharxivSearch

arXiv subjects

Yiming Zhu

Publications and source records attributed to Yiming Zhu.

At least 19 recordsLinked to original sources

Optical-NIR Multi-band Photometric Analysis and Characterization of Giant Exoplanets with CPI-C

We present a multi-band photometric approach to characterize giant exoplanets, which represents one of the anticipated core scientific outcomes of Cool Planet Imaging Coronagraph (CPI-C). CPI-C operates with two observational channels covering visible and near-infrared wavelengths, each equipped with four broadband filters. The planet--star flux ratio integrated over each filter bandpass is calculated for photometric analysis. For cool planets observed in the visible bands, the data are primarily used to fit the overall spectral shape and methane-induced modulation, providing sensitivity to metallicity- and cloud-dependent spectral variations while constraining the reflected-light spectral shape and the combined scaling involving planet radius, orbital separation, and orbital phase. In the near-infrared bands, which probe thermal emission, the data help to better constrain fundamental planetary parameters including the effective temperature, radius, surface gravity and mass. For a synthetic giant planet with measurable reflected-light and thermal-emission components, the combined VIS4+NIR4 data provide tighter same-target constraints than either filter set alone, especially for the planet radius and cloud sedimentation parameter. Our simulations incorporate realistic instrument throughput, detector noise, and residual speckle noise. The results demonstrate that the eight-band design spanning visible to near-infrared wavelengths supports reflected-light diagnostics, thermal-emission characterization, and joint optical--NIR analysis of giant exoplanets within CPI-C science observations.

astro-ph.EP

Revealing the high redshift host galaxy of the short GRB 061201 with JWST

Using deep near-infrared and optical images from JWST and HST, we identify a new host galaxy candidate for GRB 061201. It lies ~2" from the optical afterglow position. Photometric redshift fitting yields z~1.2. We compare the previously proposed host at z=0.111 with the new candidate. The chance-coincidence probability is $P_{cc}=0.18$, above the classical threshold of 0.1 but consistent with a physical association given the extreme depth of JWST imaging. In contrast, evaluated with corresponding JWST observations, the previously claimed host has a lower $P_{cc}=0.11$, which is driven primarily by bright-tail statistics rather than a more plausible association. A high-z origin is favored by three independent lines of evidence. First, for the z=0.111 scenario, the beaming-corrected energy shows GRB 061201 is an outlier of the Ghirlanda ($E_{p,i}-E_\gamma$) relation for short GRBs, while for the z=1.2 scenario, it is well consistent with the Amati relation. Second, deep near-infrared observations rule out a kilonova similar to AT2017gfo at z=0.111. Third, afterglow modeling yields an AIC criterion of $\Delta$AIC=16.35, providing strong evidence for the high-redshift scenario. Assuming the host candidate is the actual host galaxy of GRB 061201, the physical offset is 16.4-16.9 kpc (substantially reduced from ~42 kpc) and the host stellar age is ~2 Gyr, which are consistent with the host population of short GRBs. A low-redshift origin would lead to a very high binary neutron star merger rate of ~1400 Gpc$^{-3}$ yr$^{-1}$, which is contradictory to the gravitational-wave constraint. We suggest that GRB 061201 originates from a moderately high-redshift (z~1.2) host, significantly alleviating this apparent merger rate discrepancy. This case demonstrates the power of deep JWST exposures in revealing the host galaxies of historically hostless GRBs.

astro-ph.HE

A General Framework for Multimodal LLM-Based Multimedia Understanding in Large-Scale Recommendation Systems

Conventional recommendation systems frequently fail to fully exploit the high-dimensional semantic signals inherent in multimedia content, thereby limiting the fidelity of user preference modeling. While Multimodal Large Language Models (MM-LLMs) offer robust mechanisms for interpreting such complex data, their integration into latency-constrained, industrial-scale architectures remains a significant challenge. To address this, we propose a generalized framework for MM-LLM-driven multimedia understanding. Our methodology employs a tripartite architecture encompassing content interpretation, representation extraction, and systematic pipeline integration, instantiated via a LLaMA2-based model that generates descriptive captions subsequently ingested as tokenized categorical features. Empirical evaluation demonstrates the efficacy of this approach, yielding a $0.35\%$ increase in offline AUC and a $0.02\%$ improvement in online metrics at scale, substantiating the practical viability of leveraging MM-LLMs to enhance large-scale recommendation performance.

cs.IR

Semantic Communication for Multi-Satellite Massive MIMO Transmission: A Mixture of Cooperative Modes Framework

This paper investigates semantic communications (SemComs) for multi-satellite cooperative massive multiple-input multiple-output (MIMO) transmission, where multiple massive-MIMO satellites jointly serve a common set of multi-antenna user terminals. For the first time, SemComs with image transmission task are integrated into satellite massive MIMO and multi-satellite cooperative transmission. For the two representative cooperative modes, namely coherent transmission (CT) and non-coherent transmission (NCT), we develop multi-satellite CT (MSCT) and multi-satellite NCT (MSNCT) SemCom frameworks, respectively. MSCT adopts a symmetric architecture, whereas MSNCT introduces transmitter-side stream allocation and a two-stage receiver design that combines per-stream semantic extraction with cross-stream semantic-interference exploitation. To instantiate MSCT, we further design a symmetric encoder and decoder network based on hybrid Swin-Transformer and lightweight bottleneck convolutional neural network (CNN) blocks, termed HSTC, where Swin Transformer provides scalable computation and the CNN branch improves performance and convergence. For MSNCT, a Transformer-based backbone is employed to support cross-stream interference exploitation through global attention. Building on these two frameworks, we propose a mixture of cooperative modes (MoCM) framework, in which a permutation-invariant network dynamically switches between MSCT and MSNCT using multi-satellite statistical channel state information, thereby balancing semantic performance and complexity. Simulation results under practical configurations demonstrate the performance gains of the proposed frameworks.

eess.SP

Defending against Patch-Based and Texture-Based Adversarial Attacks with Spectral Decomposition

Adversarial examples present significant challenges to the security of Deep Neural Network (DNN) applications. Specifically, there are patch-based and texture-based attacks that are usually used to craft physical-world adversarial examples, posing real threats to security-critical applications such as person detection in surveillance and autonomous systems, because those attacks are physically realizable. Existing defense mechanisms face challenges in the adaptive attack setting, i.e., the attacks are specifically designed against them. In this paper, we propose Adversarial Spectrum Defense (ASD), a defense mechanism that leverages spectral decomposition via Discrete Wavelet Transform (DWT) to analyze adversarial patterns across multiple frequency scales. The multi-resolution and localization capability of DWT enables ASD to capture both high-frequency (fine-grained) and low-frequency (spatially pervasive) perturbations. By integrating this spectral analysis with the off-the-shelf Adversarial Training (AT) model, ASD provides a comprehensive defense strategy against both patch-based and texture-based adversarial attacks. Extensive experiments demonstrate that ASD+AT achieved state-of-the-art (SOTA) performance against various attacks, outperforming the APs of previous defense methods by 21.73%, in the face of strong adaptive adversaries specifically designed against ASD. Code available at https://github.com/weiz0823/adv-spectral-defense .

cs.CV

Toward Multi-Satellite Cooperative Transmission: A Joint Framework for CSI Acquisition, Feedback, and Phase Synchronization

The stringent link budget, caused by long propagation distances and payload constraints, poses a fundamental bottleneck for single-satellite transmission. Although LEO mega-constellations make multi-satellite cooperative transmission (MSCT), such as distributed precoding (DP), increasingly feasible, its cooperative gains critically rely on stringent time-frequency-phase synchronization (TFP-Sync), which is difficult to maintain under rapid channel variation and feedback latency. To address this issue, this paper proposes a joint CSI acquisition, feedback, and phase-level synchronization (JCAFPS) framework for MSCT. Specifically, to enable reliable, overhead-efficient CSI acquisition, we design a beam-domain adjustable phase-shift tracking reference signal (TRS) transmission scheme, along with criteria for the TRS and CSI-feedback periods. Then, exploiting deterministic orbital motion and dominant LoS propagation, we establish a polynomial model for the temporal evolution of delay and Doppler shift, and derive an OFDM-based multi-satellite signal model under non-ideal synchronization. The analysis reveals that, unlike the single-satellite case, the composite multi-satellite channel exhibits nonlinear time-frequency-varying phase behavior, necessitating symbol- and subcarrier-wise phase precompensation for coherent transmission. Based on these results, we develop a practical closed-loop realization integrating single-TRS-based channel parameter estimation, multi-TRS-based channel prediction, predictive CSI feedback, and user-specific TFP precompensation. Numerical results demonstrate that the proposed framework achieves accurate CSI acquisition and precise TFP-Sync, enabling DP-based dual-satellite cooperative transmission to approach the theoretical 6 dB power gain over single-satellite transmission, while remaining robust under extended prediction durations and enlarged TRS periods.

eess.SP

DSCSNet: A Dynamic Sparse Compression Sensing Network for Closely-Spaced Infrared Small Target Unmixing

Due to the limitations of optical lens focal length and detector resolution, distant clustered infrared small targets often appear as mixed spots. The Close Small Object Unmixing (CSOU) task aims to recover the number, sub-pixel positions, and radiant intensities of individual targets from these spots, which is a highly ill-posed inverse problem. Existing methods struggle to balance the rigorous sparsity guarantees of model-driven approaches and the dynamic scene adaptability of data-driven methods. To address this dilemma, this paper proposes a Dynamic Sparse Compressed Sensing Network (DSCSNet), a deep-unfolded network that couples the Alternating Direction Method of Multipliers (ADMM) with learnable parameters. Specifically, we embed a strict $\ell_1$-norm sparsity constraint into the auxiliary variable update step of ADMM to replace the traditional $\ell_2$-norm smoothness-promoting terms, which effectively preserves the discrete energy peaks of small targets. We also integrate a self-attention-based dynamic thresholding mechanism into the reconstruction stage, which adaptively adjusts the sparsification intensity using the sparsity-enhanced information from the iterative process. These modules are jointly optimized end-to-end across the three iterative steps of ADMM. Retaining the physical logic of compressed sensing, DSCSNet achieves robust sparsity induction and scene adaptability, thus enhancing the unmixing accuracy and generalization in complex infrared scenarios. Extensive experiments on the synthetic infrared dataset CSIST-100K demonstrate that DSCSNet outperforms state-of-the-art methods in key metrics such as CSO-mAP and sub-pixel localization error.

cs.CV

A Comparative Analysis of Social Network Topology in Reddit and Moltbook

Recent advances in agent-mediated systems have enabled a new paradigm of social network simulation, where AI agents interact with human-like autonomy. This evolution has fostered the emergence of agent-driven social networks such as Moltbook, a Reddit-like platform populated entirely by AI agents. Despite these developments, empirical comparisons between agent-driven and human-driven social networks remain scarce, limiting our understanding of how their network topologies might diverge. This paper presents the first comparative analysis of network topology on Moltbook, utilizing a comment network comprising 33,577 nodes and 697,688 edges. To provide a benchmark, we curated a parallel dataset from Reddit consisting of 7.8 million nodes and 51.8 million edges. We examine key structural differences between agent-drive and human-drive networks, specifically focusing on topological patterns and the edge formation efficacy of their respective posts. Our findings provide a foundational profile of AI-driven social structures, serving as a preliminary step toward developing more robust and authentic agent-mediated social systems.

cs.SI

Population Metrology of a Hidden Exciton Reservoir: Quasi-Thermalization versus Localization

Long-lived dark states can dominate the lowest-energy manifold of optically driven quantum materials, yet their occupation remains difficult to quantify, leaving it unclear whether it reflects thermal redistribution or kinetic trapping. We combine microsphere-enabled far-field access with quantitative optical-response calibration to retrieve dark-to-bright population ratios in monolayer WSe2. Temperature-dependent measurements and controlled defect enhancement separate mobile and localized contributions. Near room temperature, the mobile dark-to-bright ratio reaches approximately 65% of its Boltzmann limit, indicating substantial but incomplete quasi-thermalization, whereas the low-temperature excess is dominated by defect-assisted localization. Dark-state dominance alone therefore does not establish equilibration, a distinction essential for interpreting transport and collective phases in optically hidden quasiparticle reservoirs.

physics.optics

Multi-Satellite Multi-Stream Beamspace Massive MIMO Transmission

This paper studies multi-satellite multi-stream (MSMS) beamspace transmission, where multiple satellites cooperate to form a distributed multiple-input multiple-output (MIMO) system and jointly deliver multiple data streams to multi-antenna user terminals (UTs), and beamspace transmission combines earth-moving beamforming with beam-domain precoding. For the first time, we formulate the signal model for MSMS beamspace MIMO transmission. Under synchronization errors, multi-antenna UTs enable the distributed MIMO channel to exhibit higher rank, supporting multiple data streams. Beamspace MIMO retains conventional codebook based beamforming while providing the performance gains of precoding. Based on the signal model, we propose statistical channel state information (sCSI)-based optimization of satellite clustering, beam selection, and transmit precoding, using a sum-rate upper-bound approximation. With given satellite clustering and beam selection, we cast precoder design as an equivalent covariance decomposition-based weighted minimum mean square error (CDWMMSE) problem. To obtain tractable algorithms, we develop a closed-form covariance decomposition required by CDWMMSE and derive an iterative MSMS beam-domain precoder under sCSI. Following this, we further propose several heuristic closed-form precoders to avoid iterative cost. For satellite clustering, we enhance a competition-based algorithm by introducing a mechanism to regulate the number of satellites serving certain UT. Furthermore, we design a two-stage low-complexity beam selection algorithm focused on enhancing the effective channel power. Simulations under practical configurations validate the proposed methods across the number of data streams, receive antennas, serving satellites, and active beams, and show that beamspace transmission approaches conventional MIMO performance at lower complexity.

eess.SP

CPI-C: Cool Planet Imaging Coronagraph on Chinese Space Station Survey Telescope

Cool Planet Imaging Coronagraph (CPI-C) on Chinese Space Station Survey Telescope (CSST) is proposed to direct image the cool planets around nearby solar-type stars (within 40 pc). The core scientific objective of CPI-C is to conduct high-contrast directly imaging surveys of exoplanets ranging in size from Neptune-like to Jupiter-like, located at separations of 0.5 to 5 AU from their host stars, and to perform systematic spectroscopic analysis of the detected planets through high-precision multi-band photometry. CPI-C employs a step-transmission apodization technique to suppress the diffraction noises from the telescope pupil and a precise phase correction technique to eliminate the speckle noises due to imperfections of the optical surfaces. The contrast requirement is better than $10^{-8}$ at an inner working angle (IWA) of $3-4\lambda/D$, in the visible wavelength from 600 nm to 900 nm. CPI-C will be the first space-based instrument capable of directly imaging the reflection light from the cool exoplanets in the visible wavelength enabling the measurement of key physical parameters such as the effective temperature, surface gravity, radius, mass, and other key parameters. The potential observation results will significantly contribute to further understand the formation and evolution mechanisms of planets, which will also lay a solid foundation for future confirmation of the Earth-twins in the next generation space flagship missions.

astro-ph.EP

The Invisible Hand: Characterizing Generative AI Adoption and its Effects on An Online Freelancing Market

Since the COVID-19 pandemic, freelancing platforms have experienced significant growth in both worker registrations and job postings. However, the rise of generative AI (GenAI) technologies has raised questions about how it affect the job posting in freelancer market. Despite growing discussions, there is limited empirical research on the GenAI adoption and its effect on job demand and worker engagement. We present a large-scale analysis of Freelancer.com, utilizing over 1.8 million job posts and 3.8 million users. We investigate the emergence of jobs with the adoption of GenAI and identify leading position of ChatGPT in the freelancing market. With a focus on ChatGPT related jobs, we inspect their specific skill requirements, and the tasks that workers are asked to perform. Our findings provide insights into the evolving landscape of freelancing in the age of AI, offering a comprehensive profile of GenAI's effects on employment, skills, and user behaviors in freelancing market.

cs.CE

Near-field perturbation of laser filament enabling simultaneous far-field THz diagnosis and broadband calculus processing

Terahertz (THz) wave manipulation based on laser filaments-plasma channels formed by femtosecond laser-induced air ionization-has emerged as a promising platform for free-space THz applications. However, in-situ characterization of the spatially confined THz modes within filaments faces significant challenges due to the plasma's ultra-high intensity, which not only hinders direct near-field probing but also limits reliance on indirect far-field reconstruction. Here, we introduce a non-invasive near-field modulation scheme where a metal plate approaches the filament at submillimeter distances (comparable to THz wavelengths), perturbing the dielectric environment to convert the symmetric annular THz mode into an asymmetric state. This controlled transition enables far-field detection of broadband calculus behaviors (first- and second-order differentiation/integration) on time-domain THz waveforms and characteristic spectral transfer functions with 1/f, 1/f^2, f or f^2 dependency (where f is the THz frequency), thereby diagnosing the near-field THz mode confinement. Hence, the proposed approach synergizes near-field modulation efficiency with far-field detection robustness, advancing fundamental understanding of plasma-THz interactions and enabling novel all-optical signal processing for filament-based THz technologies.

physics.optics

Mock Observations for the CSST Mission: CPI-C -- Instrument Simulation

To support the development of the data processing pipeline and the scientific performance assessment for the Cool Planet Imaging Coronagraph (CPI-C) on the Chinese Space Station Survey Telescope (CSST), we have developed the end-to-end instrument simulation program, CPISM. This paper details the core modules of CPISM that simulate the CPI-C instrument, focusing on the simulation of the high-contrast imaging optical system and the visible-band science camera. We modeled key optical components, such as the transmission apodizing filter, the wavefront corrector, and the focal plane mask using the HCIPy package. A $10^{-8}$ contrast dark hole region, consistent with design specifications, was simulated using the Electric Field Conjugation (EFC) optimization method, and broadband observation effects were considered. For the science camera, which is an electron multiplying charge-coupled device (EMCCD), we established a detailed model encompassing photon collection, charge transfer, electron multiplication (EM), and readout processes, based on test data. This model simulates complex instrumental features including dark current, charge transfer efficiency, clock-induced charge, multiplication noise factor, and various readout effects like striping and drift. We also proposed and validated an improved statistical model for the EM process to enhance simulation efficiency. CPISM can generate simulated images containing rich instrumental details, closely similar to the expected real observational data, thus laying the foundation for the development and verification of CPI-C data processing algorithms and preparations for future scientific research.

astro-ph.IM

TAPOM: Task-Space Topology-Guided Motion Planning for Manipulating Elongated Object in Cluttered Environments

Robotic manipulation in complex, constrained spaces is vital for widespread applications but challenging, particularly when navigating narrow passages with elongated objects. Existing planning methods often fail in these low-clearance scenarios due to the sampling difficulties or the local minima. This work proposes Topology-Aware Planning for Object Manipulation (TAPOM), which explicitly incorporates task-space topological analysis to enable efficient planning. TAPOM uses a high-level analysis to identify critical pathways and generate guiding keyframes, which are utilized in a low-level planner to find feasible configuration space trajectories. Experimental validation demonstrates significantly high success rates and improved efficiency over state-of-the-art methods on low-clearance manipulation tasks. This approach offers broad implications for enhancing manipulation capabilities of robots in complex real-world environments.

cs.RO

Achieving Constant-Envelope Waveform in CP-OFDMA Framework

OFDM is widely adopted in modern wireless communication systems, but its power efficiency is limited by high envelope fluctuations. Although various high power-efficiency waveforms have been proposed, most are incompatible with the CP-OFDMA framework and remain ineffective in multi-user downlink transmissions. To address this issue, we propose a constant-envelope (CE) waveform design, which enables low-complexity transceiver architectures while maintaining full compatibility with the prevailing CP-OFDMA framework. Specifically, we start from a general CE FDMA signal model and develop a CP-OFDMA-compatible waveform implementation structure, followed by the design of an optimized CE-constrained pulse-shaping filter to suppress out-of-band emissions. To tackle channel estimation challenge under non-flat frequency-domain pilots induced by CE modulation, we optimize the time-domain binary pilot sequence to achieve frequency-domain CE properties, and then propose a multi-stage method combining delay-domain denoising with power delay profile estimation to facilitate reduced-dimension LMMSE estimation. Subsequently, we design a low-complexity maximum ratio combining-aided LMMSE equalizer by exploiting the periodicity and conjugate symmetry of the CE received signals. To mitigate the downlink peak-to-average power ratio increase caused by FDMA, we further develop a multi-user downlink CE transmission scheme including multiple access mechanism, downlink control information design, and corresponding system-level implementation, which ensures compatibility with the New Radio standard. Numerical results demonstrate that the proposed scheme achieves bit error rate performance close to the ideal case while significantly reducing transceiver complexity compared to existing CE waveform solutions.

eess.SP