SearcharxivSearch

arXiv subjects

Hanyu Jiang

Publications and source records attributed to Hanyu Jiang.

13 recordsLinked to original sources

Gravitational Lensing of Gravitational Wave for Generalized Navarro-Frenk-White Profile and Einasto Profile

The density profiles of Dark matter (DM) halos carry imprints of the DM nature and may be constrained through the lensing effects on gravitational waves (GWs) arising from the halo gravitational potential. In this paper, we investigate GW lensing by two representative types of halo density profiles, i.e., the generalized Navarro-Frenk-White (gNFW) density profile and the Einasto density profile. Using the gravitational lensing equation, we first examine the parameter-space distribution and the imaging characteristics of both profiles under strong lensing, partitioning the parameter space into distinct regions according to the Morse index. We then conduct a detailed analysis of the modulus $\big|F\big|$ and phase $\mathrm{Arg}(F)$ of the amplification factor $F\big(w,y\big)$ at low frequency regime. Our results show that, for a fixed lens mass , increasing the gNFW slope $γ$ leads to a larger amplitude, more rapid oscillation, and more distinct wave-packet morphology in the multiple-image regime. Compared with the gNFW case, the Einasto case (with $α=0.16$, $y<0.6$, and the same $M_{200}$) produces a stronger lensing effect. Notably, the evolution of $F\big(w,y\big)$ with frequency for the Einasto profile differs from that of the gNFW case, making its behavior particularly distinctive.

gr-qc

SUGFW+: An Uncertainty-guided Feature Weighting Framework for Cold Start Active Adaptation of SAM in Medical Image Segmentation

Cold Start Active Learning (CSAL) is important in improving the performance of a medical image segmentation model with low annotation budget by querying a small subset for annotation from an unlabeled training set. Existing CSAL methods typically rely on inefficient dataset-specific Self-Supervised Learning (SSL) to map the unlabeled images into a feature space for sample selection. Recently, the advent of foundation models such as the Segment Anything Model (SAM) offer a promising alternative as the pre-trained model can provide strong generalizable feature embeddings, and allow high performance in downstream tasks after fine-tuning (adaptation). However, how to systematically exploit SAM's inherent embeddings for cold-start sample selection during adaptation with low annotation budget remains underexplored. To address this, we propose an extended SAM-based Uncertainty-guided Feature Weighting (SUGFW+) framework for CSAL and adaptation of SAM. Specifically, it leverages the SAM for Patch-level Feature and Uncertainty Calculation (PFUC), and introduces a Patch-based Global Distinct Representation (PGDR) module that aggregates patch-level embeddings into highly discriminative, uncertainty-aware image-level features. These features are then utilized by a Greedy Selection with Cluster and Uncertainty (GSCU) strategy to combine diversity and uncertainty during sample selection. Unlike prior CSAL methods that decouple sample selection from model training, SUGFW+ tightly integrates these two stages via an Uncertainty-Prompted Fine-Tuning (UPFT) process of SAM in model training. Extensive experiments on four public datasets demonstrate that SUGFW+ achieves state-of-the-art performance against existing CSAL methods. Code is available at https://github.com/HiLab-git/SUGFW-plus.

cs.CV

AUHead: Realistic Emotional Talking Head Generation via Action Units Control

Realistic talking-head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nuanced emotional expressions due to the lack of fine-grained emotion control. To address this issue, we introduce a novel two-stage method (AUHead) to disentangle fine-grained emotion control, i.e. , Action Units (AUs), from audio and achieve controllable generation. In the first stage, we explore the AU generation abilities of large audio-language models (ALMs), by spatial-temporal AU tokenization and an "emotion-then-AU" chain-of-thought mechanism. It aims to disentangle AUs from raw speech, effectively capturing subtle emotional cues. In the second stage, we propose an AU-driven controllable diffusion model that synthesizes realistic talking-head videos conditioned on AU sequences. Specifically, we first map the AU sequences into the structured 2D facial representation to enhance spatial fidelity, and then model the AU-vision interaction within cross-attention modules. To achieve flexible AU-quality trade-off control, we introduce an AU disentanglement guidance strategy during inference, further refining the emotional expressiveness and identity consistency of the generated videos. Results on benchmark datasets demonstrate that our approach achieves competitive performance in emotional realism, accurate lip synchronization, and visual coherence, significantly surpassing existing techniques. Our implementation is available at https://github.com/laura990501/AUHead_ICLR

cs.CV

Artificial Precision Polarization Array: Sensitivity for the axion-like dark matter with clock satellites

The approaches to searching for axion-like signals based on pulsars include observations with pulsar timing arrays (PTAs) and pulsar polarization arrays (PPAs). However, these methods are limited by observational uncertainties arising from multiple unknown and periodic physical effects, which substantially complicate subsequent data analysis. To mitigate these issues and improve data fidelity, we propose the Artificial Pulsar Polarization Arrays (APPA): a satellite network comprising multiple pulsed signal transmitters and a dedicated receiver satellite. To constrain the axion-photon coupling parameter $g_{aγ}$, we generate simulated observations using Monte Carlo methods and investigate the sensitivity of APPA using two complementary approaches: Likelihood analysis and frequentist analysis. Simulations indicate that for the axion mass range of $10^{-22}-10^{-18}$ eV, APPA yields a tighter upper limit on $g_{aγ}$ (at the 95\% C.L.) than conventional ground-based observations, while also achieving superior detection sensitivity. Moreover, a larger spatial distribution scale of the satellite network corresponds to a greater advantage in detecting axions with lighter masses.

astro-ph.CO

Detecting Chiral Gravitational Wave Background with a Dipole Pulsar Timing Array

The pulsar timing array (PTA) is a powerful technique for detecting nanohertz gravitational wave backgrounds (GWBs). However, conventional PTAs lack sensitivity to parity violation in the GWB. In this work, we propose a dipole pulsar timing array system (dPTA). By deriving the overlap reduction functions (ORFs) from the cross-correlation of timing signals, we find that this system exhibits sensitivity to chiral GWBs in the nanohertz regime. Furthermore, through numerical calculations of its sensitivity curves, we demonstrate that the dPTA extends the detectable frequency range of PTAs for GWBs from the nanohertz to the microhertz regime.

gr-qc

Towards Unified Facial Action Unit Recognition Framework by Large Language Models

Facial Action Units (AUs) are of great significance in the realm of affective computing. In this paper, we propose AU-LLaVA, the first unified AU recognition framework based on the Large Language Model (LLM). AU-LLaVA consists of a visual encoder, a linear projector layer, and a pre-trained LLM. We meticulously craft the text descriptions and fine-tune the model on various AU datasets, allowing it to generate different formats of AU recognition results for the same input image. On the BP4D and DISFA datasets, AU-LLaVA delivers the most accurate recognition results for nearly half of the AUs. Our model achieves improvements of F1-score up to 11.4% in specific AU recognition compared to previous benchmark results. On the FEAFA dataset, our method achieves significant improvements over all 24 AUs compared to previous benchmark results. AU-LLaVA demonstrates exceptional performance and versatility in AU recognition.

cs.CV

MVLLaVA: An Intelligent Agent for Unified and Flexible Novel View Synthesis

This paper introduces MVLLaVA, an intelligent agent designed for novel view synthesis tasks. MVLLaVA integrates multiple multi-view diffusion models with a large multimodal model, LLaVA, enabling it to handle a wide range of tasks efficiently. MVLLaVA represents a versatile and unified platform that adapts to diverse input types, including a single image, a descriptive caption, or a specific change in viewing azimuth, guided by language instructions for viewpoint generation. We carefully craft task-specific instruction templates, which are subsequently used to fine-tune LLaVA. As a result, MVLLaVA acquires the capability to generate novel view images based on user instructions, demonstrating its flexibility across diverse tasks. Experiments are conducted to validate the effectiveness of MVLLaVA, demonstrating its robust performance and versatility in tackling diverse novel view synthesis challenges.

cs.CV

Pattern-Matching Dynamic Memory Network for Dual-Mode Traffic Prediction

In recent years, deep learning has increasingly gained attention in the field of traffic prediction. Existing traffic prediction models often rely on GCNs or attention mechanisms with O(N^2) complexity to dynamically extract traffic node features, which lack efficiency and are not lightweight. Additionally, these models typically only utilize historical data for prediction, without considering the impact of the target information on the prediction. To address these issues, we propose a Pattern-Matching Dynamic Memory Network (PM-DMNet). PM-DMNet employs a novel dynamic memory network to capture traffic pattern features with only O(N) complexity, significantly reducing computational overhead while achieving excellent performance. The PM-DMNet also introduces two prediction methods: Recursive Multi-step Prediction (RMP) and Parallel Multi-step Prediction (PMP), which leverage the time features of the prediction targets to assist in the forecasting process. Furthermore, a transfer attention mechanism is integrated into PMP, transforming historical data features to better align with the predicted target states, thereby capturing trend changes more accurately and reducing errors. Extensive experiments demonstrate the superiority of the proposed model over existing benchmarks. The source codes are available at: https://github.com/wengwenchao123/PM-DMNet.

cs.LG

FoodSAM: Any Food Segmentation

In this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generated masks, we propose a novel framework, called FoodSAM. This innovative approach integrates the coarse semantic mask with SAM-generated masks to enhance semantic segmentation quality. Besides, we recognize that the ingredients in food can be supposed as independent individuals, which motivated us to perform instance segmentation on food images. Furthermore, FoodSAM extends its zero-shot capability to encompass panoptic segmentation by incorporating an object detector, which renders FoodSAM to effectively capture non-food object information. Drawing inspiration from the recent success of promptable segmentation, we also extend FoodSAM to promptable segmentation, supporting various prompt variants. Consequently, FoodSAM emerges as an all-encompassing solution capable of segmenting food items at multiple levels of granularity. Remarkably, this pioneering framework stands as the first-ever work to achieve instance, panoptic, and promptable segmentation on food images. Extensive experiments demonstrate the feasibility and impressing performance of FoodSAM, validating SAM's potential as a prominent and influential tool within the domain of food image segmentation. We release our code at https://github.com/jamesjg/FoodSAM.

cs.CV

Rate-Splitting Multiple Access for Uplink Massive MIMO With Electromagnetic Exposure Constraints

Over the past few years, the prevalence of wireless devices has become one of the essential sources of electromagnetic (EM) radiation to the public. Facing with the swift development of wireless communications, people are skeptical about the risks of long-term exposure to EM radiation. As EM exposure is required to be restricted at user terminals, it is inefficient to blindly decrease the transmit power, which leads to limited spectral efficiency and energy efficiency (EE). Recently, rate-splitting multiple access (RSMA) has been proposed as an effective way to provide higher wireless transmission performance, which is a promising technology for future wireless communications. To this end, we propose using RSMA to increase the EE of massive MIMO uplink while limiting the EM exposure of users. In particularly, we investigate the optimization of the transmit covariance matrices and decoding order using statistical channel state information (CSI). The problem is formulated as non-convex mixed integer program, which is in general difficult to handle. We first propose a modified water-filling scheme to obtain the transmit covariance matrices with fixed decoding order. Then, a greedy approach is proposed to obtain the decoding permutation. Numerical results verify the effectiveness of the proposed EM exposure-aware EE maximization scheme for uplink RSMA.

cs.IT

3UCubed: The IMAP Student Collaboration CubeSat Project

The 3UCubed project is a 3U CubeSat being jointly developed by the University of New Hampshire, Sonoma State University, and Howard University as a part of the NASA Interstellar Mapping and Acceleration Probe, IMAP, student collaboration. This project comprises of a multidisciplinary team of undergraduate students from all three universities. The mission goal of the 3UCubed is to understand how Earths polar upper atmosphere the thermosphere in Earths auroral regions, responds to particle precipitation and solar wind forcing, and internal magnetospheric processes. 3UCubed includes two instruments with rocket heritage to achieve the science mission: an ultraviolet photomultiplier tube, UVPMT, and an electron retarding potential analyzer ERPA. The spacecraft bus consists of the following subsystems: Attitude Determination and Control, Command and Data Handling, Power, Communication, Structural, and Thermal. Currently, the project is in the post-PDR stage, starting to build and test engineering models to develop a FlatSat prior to critical design review in 2023. The goal is to launch at least one 3U CubeSat to collect science data close to the anticipated peak of Solar Cycle 25 around July 2025. Our mother mission, IMAP, is also projected to launch in 2025, which will let us jointly analyze the science data of the main mission, providing the solar wind measurements and inputs to the magnetosphere with that of 3UCubed, providing the response of Earths cusp to these inputs.

astro-ph.IM

Hybrid RIS and DMA Assisted Multiuser MIMO Uplink Transmission With Electromagnetic Exposure Constraints

In the fifth-generation and beyond era, reconfigurable intelligent surface (RIS) and dynamic metasurface antennas (DMAs) are emerging metamaterials keeping up with the demand for high-quality wireless communication services, which promote the diversification of portable wireless terminals. However, along with the rapid expansion of wireless devices, the electromagnetic (EM) radiation increases unceasingly and inevitably affects public health, which requires a limited exposure level in the transmission design. To reduce the EM radiation and preserve the quality of communication service, we investigate the spectral efficiency (SE) maximization with EM constraints for uplink transmission in hybrid RIS and DMA assisted multiuser multiple-input multiple-output systems. Specifically, alternating optimization is adopted to optimize the transmit covariance, RIS phase shift, and DMA weight matrices. We first figure out the water-filling solutions of transmit covariance matrices with given RIS and DMA parameters. Then, the RIS phase shift matrix is optimized via the weighted minimum mean square error, block coordinate descent and minorization-maximization methods. Furthermore, we solve the unconstrainted DMA weight matrix optimization problem in closed form and then design the DMA weight matrix to approach this performance under DMA constraints. Numerical results confirm the effectiveness of the EM aware SE maximization transmission scheme over the conventional baselines.

cs.IT

CUDAMPF++: A Proactive Resource Exhaustion Scheme for Accelerating Homologous Sequence Search on CUDA-enabled GPU

Genomic sequence alignment is an important research topic in bioinformatics and continues to attract significant efforts. As genomic data grow exponentially, however, most of alignment methods face challenges due to their huge computational costs. HMMER, a suite of bioinformatics tools, is widely used for the analysis of homologous protein and nucleotide sequences with high sensitivity, based on profile hidden Markov models (HMMs). Its latest version, HMMER3, introdues a heuristic pipeline to accelerate the alignment process, which is carried out on central processing units (CPUs) with the support of streaming SIMD extensions (SSE) instructions. Few acceleration results have since been reported based on HMMER3. In this paper, we propose a five-tiered parallel framework, CUDAMPF++, to accelerate the most computationally intensive stages of HMMER3's pipeline, multiple/single segment Viterbi (MSV/SSV), on a single graphics processing unit (GPU). As an architecture-aware design, the proposed framework aims to fully utilize hardware resources via exploiting finer-grained parallelism (multi-sequence alignment) compared with its predecessor (CUDAMPF). In addition, we propose a novel method that proactively sacrifices L1 Cache Hit Ratio (CHR) to get improved performance and scalability in return. A comprehensive evaluation shows that the proposed framework outperfroms all existig work and exhibits good consistency in performance regardless of the variation of query models or protein sequence datasets. For MSV (SSV) kernels, the peak performance of the CUDAMPF++ is 283.9 (471.7) GCUPS on a single K40 GPU, and impressive speedups ranging from 1.x (1.7x) to 168.3x (160.7x) are achieved over the CPU-based implementation (16 cores, 32 threads).

cs.CE