SearcharxivSearch

arXiv subjects

Pengfei Ren

Publications and source records attributed to Pengfei Ren.

13 recordsLinked to original sources

Early Results from GLASS-JWST. XXVI. Spatially Resolved Star Formation and Balmer Decrements at $1.1<z<2.3$ from NIRISS Slitless Spectroscopy

Using JWST/NIRISS slitless spectroscopy, we present spatially resolved Balmer decrement measurements for 79 galaxies at $1.1 < z < 2.3$, which are gravitationally lensed by the foreground cluster Abell 2744. By stacking $\mathrm{H}\alpha$ and $\mathrm{H}\beta$ emission maps in bins of stellar mass and redshift, we derive radial profiles of nebular dust attenuation and dust-corrected star formation rate (SFR). We find tentative evidence that the radial gradients of dust attenuation toward $\mathrm{H}\alpha$ ($\rm A(\mathrm{H}\alpha)$) vary with both redshift and stellar mass. At lower redshifts ($z = 1.10$--$1.53$), low-mass galaxies ($\rm 7.0<log(M_*/M_\odot)\leq8.5$) exhibit steeper $\rm A(H\alpha)$ gradients than higher-mass galaxies ($\rm 9.5<log(M_*/M_\odot)\leq11.0$), while the latter maintain detectable dust attenuation out to larger galactocentric radii. Galaxies at higher redshifts ($z = 1.76$--$2.29$) show lower attenuation levels. At fixed galactocentric radius, galaxies in the low-redshift bin generally exhibit higher dust attenuation than those at high redshifts, consistent with an increase in dust content toward later cosmic times. Dust-corrected SFR profiles in massive systems at lower redshifts are more spatially extended than those at higher redshifts, consistent with inside-out disk growth at $z\lesssim1.5$. These results suggest possible differences in attenuation properties across stellar mass and redshift bins, and demonstrate the power of gravitational lensing to probe internal structures in faint galaxies at sub-kiloparsec resolution.

astro-ph.GA

UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between input images and rendered outputs, or multi-view consistency constraints, as the driving force to optimize networks and progressively refine pose accuracy. However, these methods are highly susceptible to noisy pseudo-labels and overlook the importance of fully exploiting fine-grained spatial correlations, which undermines the stability of model training. To address these issues, we propose UST-Hand, a self-supervised learning framework that estimates uncertainty distribution of hand pose and constructs a probabilistic point cloud feature space, which enables the complex spatiotemporal relationship modeling. UST-Hand employs a conditional normalizing flow model to capture hand pose distributions and samples diverse hypotheses, facilitating robust learning under noisy pseudo-labels supervision with enhanced stability. These multi-hypothesis are mapped to a unified probabilistic 3D point cloud space for multi-view and temporal feature interaction, comprehensively exploring hand motion patterns and fine-grained spatial correlations. Extensive experiments on three challenging datasets demonstrate that UST-Hand achieves state-of-the-art performance, outperforming existing self-supervised methods by up to 37.8% in Mean Per Vertex Position Error (MPVPE).

cs.CV

A Spatially Resolved HI Survey of Seyfert Galaxies: the Role of AGN Feedback in Shaping Atomic Gas Reservoirs

Active galactic nucleus (AGN) feedback is a key ingredient in galaxy evolution, yet its impact on the cold atomic gas reservoir -- the neutral hydrogen (HI) phase -- remains poorly constrained. We present the most extensive spatially resolved HI 21-cm survey of Seyfert AGN hosts to date, based on observations with the Giant Metrewave Radio Telescope (GMRT). Our high-resolution HI maps of eight Seyfert galaxies reveal detailed kinematics and surface density distributions of their atomic gas disks. We find that AGN-host galaxies exhibit a slightly shallower HI mass-size relation than the canonical relation or the SIMBA simulation predictions; however, the measured slope remains consistent with the canonical value within $2\sigma$ uncertainties. This result suggests that AGN feedback does not significantly disrupt the global extent or large-scale structure of atomic gas reservoirs. To investigate the internal HI kinematics in greater detail, we perform a 3D kinematic forward modeling of the HI disk in UGC 4503. Our analysis reveals an elevated intrinsic velocity dispersion of $\sigma = 14.9^{+6.1}_{-3.8}$ km/s and a reduced level of rotational support, with $V/\sigma = 14.28_{-4.17}^{+4.97}$, compared to large-sample star-forming spirals. These kinematic signatures, together with localized residuals in the velocity field, indicate that AGN-driven outflows or jets may inject or indirectly affect the turbulence in the atomic gas disk, potentially regulating the cold gas reservoir. Future GMRT observations, combined with optical integral-field spectroscopy from MaNGA, will enable quantitative constraints on the role of AGN feedback in regulating star formation efficiency across a larger and more representative galaxy sample.

astro-ph.GA

OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition

Online micro gesture recognition from hand skeletons is critical for VR/AR interaction but faces challenges due to limited public datasets and task-specific algorithms. Micro gestures involve subtle motion patterns, which make constructing datasets with precise skeletons and frame-level annotations difficult. To this end, we develop a multi-view self-supervised pipeline to automatically generate skeleton data, complemented by heuristic rules and expert refinement for semi-automatic annotation. Based on this pipeline, we introduce OMG-Bench, the first large-scale public benchmark for skeleton-based online micro gesture recognition. It features 40 fine-grained gesture classes with 13,948 instances across 1,272 sequences, characterized by subtle motions, rapid dynamics, and continuous execution. To tackle these challenges, we propose Hierarchical Memory-Augmented Transformer (HMATr), an end-to-end framework that unifies gesture detection and classification by leveraging hierarchical memory banks which store frame-level details and window-level semantics to preserve historical context. In addition, it employs learnable position-aware queries initialized from the memory to implicitly encode gesture positions and semantics. Experiments show that HMATr outperforms state-of-the-art methods by 7.6% in detection rate, establishing a strong baseline for online micro gesture recognition. Project page: https://omg-bench.github.io/

cs.CV

The velocity dispersion profile of nine open clusters in the solar neighborhood

We analyze the velocity dispersion profiles of nine open clusters in the solar neighborhood using kinematic data from Gaia DR 3, aiming to identify potential dynamical signatures of stellar-mass black holes through a comparison of theoretical and observed dispersion profiles. The selected clusters include LP2373 gp4, NGC 1980, NGC 2451A, NGC 2516, NGC 3532, NGC 6475, UBC 7, Praesepe, and Pleiades. We refine the center positions of the clusters with the Meanshift algorithm. Using the Markov Chain Monte Carlo method, we calculate the velocity dispersion for each cluster and construct one-dimensional velocity dispersion profiles. NGC 2516, NGC 3532, and NGC 6475 show potential central cusps in their radial velocity dispersion profiles, which may indicate the presence of stellar-mass black holes. LP2373 gp4, NGC 6475, and Praesepe all display a negative correlation between velocity dispersion and stellar mass, indicating these clusters are approaching energy equipartition or expanding. NGC 2516 and NGC 3532 exhibit a positive dependence between velocity dispersion and stellar mass, which may be attributed to the preferential ejection of massive stars following dynamical interactions involving binaries or black holes. These two clusters are the only two that are dynamical not relaxed and are closest to virial equilibrium. We compare the observations with N-body simulations of star clusters. A comparison of observed and simulated velocity dispersion profiles reveals that NGC 2516 and NGC 3532 exhibit lower proper motion dispersions than model clusters. Better agreement with the observed profiles is achieved for model clusters with larger ages. This suggests that the observed clusters may have undergone rapid dynamical evolution. Our results suggest that NGC 2516 and NGC 3532 may host at least two stellar-mass black holes each.

astro-ph.GA

New Synthetic Goldmine: Hand Joint Angle-Driven EMG Data Generation Framework for Micro-Gesture Recognition

Electromyography (EMG)-based gesture recognition has emerged as a promising approach for human-computer interaction. However, its performance is often limited by the scarcity of labeled EMG data, significant cross-user variability, and poor generalization to unseen gestures. To address these challenges, we propose SeqEMG-GAN, a conditional, sequence-driven generative framework that synthesizes high-fidelity EMG signals from hand joint angle sequences. Our method introduces a context-aware architecture composed of an angle encoder, a dual-layer context encoder featuring the novel Ang2Gist unit, a deep convolutional EMG generator, and a discriminator, all jointly optimized via adversarial learning. By conditioning on joint kinematic trajectories, SeqEMG-GAN is capable of generating semantically consistent EMG sequences, even for previously unseen gestures, thereby enhancing data diversity and physiological plausibility. Experimental results show that classifiers trained solely on synthetic data experience only a slight accuracy drop (from 57.77\% to 55.71\%). In contrast, training with a combination of real and synthetic data significantly improves accuracy to 60.53\%, outperforming real-only training by 2.76\%. These findings demonstrate the effectiveness of our framework,also achieves the state-of-art performance in augmenting EMG datasets and enhancing gesture recognition performance for applications such as neural robotic hand control, AI/AR glasses, and gesture-based virtual gaming systems.

cs.HC

Decoupled Iterative Refinement Framework for Interacting Hands Reconstruction from a Single RGB Image

Reconstructing interacting hands from a single RGB image is a very challenging task. On the one hand, severe mutual occlusion and similar local appearance between two hands confuse the extraction of visual features, resulting in the misalignment of estimated hand meshes and the image. On the other hand, there are complex spatial relationship between interacting hands, which significantly increases the solution space of hand poses and increases the difficulty of network learning. In this paper, we propose a decoupled iterative refinement framework to achieve pixel-alignment hand reconstruction while efficiently modeling the spatial relationship between hands. Specifically, we define two feature spaces with different characteristics, namely 2D visual feature space and 3D joint feature space. First, we obtain joint-wise features from the visual feature map and utilize a graph convolution network and a transformer to perform intra- and inter-hand information interaction in the 3D joint feature space, respectively. Then, we project the joint features with global information back into the 2D visual feature space in an obfuscation-free manner and utilize the 2D convolution for pixel-wise enhancement. By performing multiple alternate enhancements in the two feature spaces, our method can achieve an accurate and robust reconstruction of interacting hands. Our method outperforms all existing two-hand reconstruction methods by a large margin on the InterHand2.6M dataset.

cs.CV

HaMuCo: Hand Pose Estimation via Multiview Collaborative Self-Supervised Learning

Recent advancements in 3D hand pose estimation have shown promising results, but its effectiveness has primarily relied on the availability of large-scale annotated datasets, the creation of which is a laborious and costly process. To alleviate the label-hungry limitation, we propose a self-supervised learning framework, HaMuCo, that learns a single-view hand pose estimator from multi-view pseudo 2D labels. However, one of the main challenges of self-supervised learning is the presence of noisy labels and the ``groupthink'' effect from multiple views. To overcome these issues, we introduce a cross-view interaction network that distills the single-view estimator by utilizing the cross-view correlated features and enforcing multi-view consistency to achieve collaborative learning. Both the single-view estimator and the cross-view interaction network are trained jointly in an end-to-end manner. Extensive experiments show that our method can achieve state-of-the-art performance on multi-view self-supervised hand pose estimation. Furthermore, the proposed cross-view interaction network can also be applied to hand pose estimation from multi-view input and outperforms previous methods under the same settings.

cs.CV

Can Shuffling Video Benefit Temporal Bias Problem: A Novel Training Framework for Temporal Grounding

Temporal grounding aims to locate a target video moment that semantically corresponds to the given sentence query in an untrimmed video. However, recent works find that existing methods suffer a severe temporal bias problem. These methods do not reason the target moment locations based on the visual-textual semantic alignment but over-rely on the temporal biases of queries in training sets. To this end, this paper proposes a novel training framework for grounding models to use shuffled videos to address temporal bias problem without losing grounding accuracy. Our framework introduces two auxiliary tasks, cross-modal matching and temporal order discrimination, to promote the grounding model training. The cross-modal matching task leverages the content consistency between shuffled and original videos to force the grounding model to mine visual contents to semantically match queries. The temporal order discrimination task leverages the difference in temporal order to strengthen the understanding of long-term temporal contexts. Extensive experiments on Charades-STA and ActivityNet Captions demonstrate the effectiveness of our method for mitigating the reliance on temporal biases and strengthening the model's generalization ability against the different temporal distributions. Code is available at https://github.com/haojc/ShufflingVideosForTSG.

cs.CV

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art methods can drop, and that they fail mostly on poses absent from the training set. Unfortunately, since the space of hand poses is highly dimensional, it is inherently not feasible to cover the whole space densely, despite recent efforts in collecting large-scale training datasets. This sampling problem is even more severe when hands are interacting with objects and/or inputs are RGB rather than depth images, as RGB images also vary with lighting conditions and colors. To address these issues, we designed a public challenge (HANDS'19) to evaluate the abilities of current 3D hand pose estimators (HPEs) to interpolate and extrapolate the poses of a training set. More exactly, HANDS'19 is designed (a) to evaluate the influence of both depth and color modalities on 3D hand pose estimation, under the presence or absence of objects; (b) to assess the generalisation abilities w.r.t. four main axes: shapes, articulations, viewpoints, and objects; (c) to explore the use of a synthetic hand model to fill the gaps of current datasets. Through the challenge, the overall accuracy has dramatically improved over the baseline, especially on extrapolation tasks, from 27mm to 13mm mean joint error. Our analyses highlight the impacts of: Data pre-processing, ensemble approaches, the use of a parametric 3D hand model (MANO), and different HPE methods/backbones.

cs.CV

AWR: Adaptive Weighting Regression for 3D Hand Pose Estimation

In this paper, we propose an adaptive weighting regression (AWR) method to leverage the advantages of both detection-based and regression-based methods. Hand joint coordinates are estimated as discrete integration of all pixels in dense representation, guided by adaptive weight maps. This learnable aggregation process introduces both dense and joint supervision that allows end-to-end training and brings adaptability to weight maps, making the network more accurate and robust. Comprehensive exploration experiments are conducted to validate the effectiveness and generality of AWR under various experimental settings, especially its usefulness for different types of dense representation and input modality. Our method outperforms other state-of-the-art methods on four publicly available datasets, including NYU, ICVL, MSRA and HANDS 2017 dataset.

cs.CV

Optical capillary-based interferometric sensor for detection of gas refractive index

In this paper, we report a capillary-based M-Z interferometer that could be used for precise detection of variations in refractive indices of gaseous samples. This sensing mechanism is quite straightforward. Cladding and core modes of a capillary are simultaneously excited by coupling coherent laser beams to the capillary cladding and core, respectively. Interferogram would be generated as the light transmitted from the core interferes with the light transmitted from the cladding. Variations in refractive index of the air filling the core lead to variations in phase difference between the core and cladding modes, thus shifting the interference fringes. Using a photodiode together with a narrow slit, we could analyze the fringe shifts. The resolution of the sensor was found to be 1*10-8 RIU, that is comparable to the highest resolution obtained by other interferometric sensors reported in previous literatures. Finally, we also analyze the temperature cross sensitivity of the sensor. The advantages of our sensor include very low cost, high sensitivity, straightforward sensing mechanism, and ease of fabrication.

physics.app-ph

Unified Scheduling for Predictable Communication Reliability in Cellular Networks with D2D Links

Cellular networks with D2D links are increasingly being explored for mission-critical applications (e.g., real-time control and AR/VR) which require predictable communication reliability. Thus it is critical to control interference among concurrent transmissions in a predictable manner to ensure the required communication reliability. To this end, we propose a Unified Cellular Scheduling (UCS) framework that, based on the Physical-Ratio-K (PRK) interference model, schedules uplink, downlink, and D2D transmissions in a unified manner to ensure predictable communication reliability while maximizing channel spatial reuse. UCS also provides a simple, effective approach to mode selection that maximizes the communication capacity for each involved communication pair. UCS effectively uses multiple channels for high throughput as well as resilience to channel fading and external interference. Leveraging the availability of base stations (BSes) as well as high-speed, out-of-band connectivity between BSes, UCS effectively orchestrates the functionalities of BSes and user equipment (UE) for light-weight control signaling and ease of incremental deployment and integration with existing cellular standards. We have implemented UCS using the open-source, standards-compliant cellular networking platform OpenAirInterface. We have validated the OpenAirInterface implementation using USRP B210 software-defined radios and lab deployment. We have also evaluated UCS through high-fidelity, at-scale simulation studies; we observe that UCS ensures predictable communication reliability while achieving a higher channel spatial reuse rate than existing mechanisms, and that the distributed UCS framework enables a channel spatial reuse rate statistically equal to that in the state-of-the-art centralized scheduling algorithm iOrder.

cs.NI