SearcharxivSearch

arXiv subjects

Yupeng Chen

Publications and source records attributed to Yupeng Chen.

At least 19 recordsLinked to original sources

Multi-scale Decomposed Convolution Refinement Network for Visible-Infrared Person Re-Identification

Visible-infrared person re-identification (VI-ReID) suffers from cross-modal discrepancies and limited discriminative capabilities, leading to suboptimal recognition performance. Current approaches exhibit limitations in semantic mining, cross-modal fusion and feature constraints. To tackle these challenges, we propose MDCRNet, a Multi-scale Decomposed Convolution Refinement Network that enhances cross-modal feature learning and discriminative metric learning. Specifically, we introduce a Hierarchical Learning Module (HLM) containing four Hierarchical Decomposed Convolution Attention (HDCA) modules, each equipped with lightweight channel attention and multi-scale spatial perception blocks to capture multi-scale spatial dependencies. Moreover, we develop a Joint Discriminative Metric Loss (JDML) incorporating a novel Granularity Discriminative Loss (GDL) that simultaneously optimizes intra-identity compactness and inter-identity separability across modalities. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that MDCRNet achieves state-of-the-art performance on both benchmarks. Code is available at https://github.com/Kevin-zms/MDCRNet.

cs.CV

Towards Spatial Supersensing in the Wild

Humans can efficiently parse continuous sensory streams, from hours to years, scaffolding an internal world model that grounds spatial reasoning and prediction. To mimic this capacity, spatial supersensing challenges multimodal models to move beyond linguistic understanding toward true world modeling. However, their benchmark relies on synthetic long videos, formed by concatenating random short clips, and is mostly limited to household scenes, leaving real-world continuity and diversity underexplored. To address the gap, we introduce $\textbf{VSI-Super-Wild}$, a large-scale benchmark for evaluating spatial supersensing over long temporal horizons in diverse in-the-wild scenes. Notably, inspired by cognitive studies on how humans structure experience, we systematically probe the full triad of world state: the agent (observer), objects (scene items), and the environment (places and global layout). In total, VSI-Super-Wild contains $\textbf{6,980}$ human-verified question-answer pairs derived from $\textbf{442}$ real-world videos spanning 8 scene categories, including long-form recordings exceeding 4 hours. Results on VSI-Super-Wild expose a fundamental disconnect: despite advances in static image understanding, models consistently fail at tasks that require coherent world-state tracking over time. We characterize how performance degrades with world-state complexity and temporal horizon, and diagnose four failure modes: spatial collapse, semantic shortcuts, insufficient update, and instance confusion. This taxonomy reveals that models lack mechanisms to bind objects, agents, and environments into a unified spatial world model, a fundamental gap that defines the path forward for spatial supersensing.

cs.CV

$D^2$-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring for D-LLMs remains largely unexplored. Unlike AR-LLMs, D-LLMs generate text through a multi-step denoising process, exposing intermediate hidden representations that may contain safety-relevant information unavailable in standard single-step monitoring setups. Motivated by the suitability of lightweight probes for always-on monitoring, we analyze which trajectory-level signals best indicate when such probes are likely to struggle. We find that the most informative signal is safety hesitation: intermediate hidden states repeatedly falling within a small margin of the probe's decision boundary. The number of such hesitation steps in D-LLM's trajectory predicts probe failure effectively, providing a proxy of sample difficulty. Building on this analysis, we propose $D^2$-Monitor, a bi-level safety monitor for D-LLMs. $D^2$-Monitor adopts a lightweight probe as an always-on monitor to jointly estimate hesitation and perform base classification. When the hesitation level exceeds a threshold, a more expressive but computationally heavier probe is activated. This dynamic routing mechanism allocates monitoring resources efficiently at test time. Evaluated on 3 datasets (WildguardMix, ToxicChat, OpenAI-Moderation) across 4 D-LLMs, $D^2$-Monitor achieves state-of-the-art performance with a compact parameter footprint ($\leq$ 0.85M parameters), and exhibits the best trade-off between effectiveness and efficiency relative to 8 baselines.

cs.AI

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI

AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances. Here we show that strong scientific reasoning does not reliably translate into accurate forecasting of future scientific advances. To study this question, we introduce CUSP, a temporally grounded evaluation suite for event-level scientific forecasting across eight scientific disciplines. Across six frontier AI models, we observe a striking asymmetry in forecasting performance together with systematic error patterns. Models often identify plausible mechanisms underlying future scientific advances, yet perform near chance on feasibility assessment, generate solution strategies that only weakly align with realized advances, and systematically predict scientific advances later than they become publicly observable. Providing additional pre-cutoff scientific knowledge improves performance but does not eliminate these forecasting limitations. These findings suggest that current AI systems possess substantial retrospective scientific competence but limited forward-looking predictive capability. Scientific forecasting should therefore be evaluated as a complementary dimension of AI scientific capability when deploying AI systems for research prioritization and scientific decision-making.

cs.AI

A low mass and radius neutron star candidate in XTE J1810-189?

Photosphere radius expansion (PRE) bursts provide a crucial tool for constraining the mass and radius of neutron stars. In this study, we analyze time-resolved spectroscopic data from XTE J1810-189 in 2008, which exhibit evidence of a PRE event. We report here the possibility of a small-size and low-mass neutron star in XTE J1810-189 with use of the advantage of the direct cooling tail method. We obtained three sets of results, which can be broadly divided into high metal abundance (20 $\rm{Z}_{\odot}$ and 40 $\rm{Z}_{\odot}$), low metal abundance and hydrogen-rich (pure hydrogen, $\rm{Z}_{\odot}$, 0.3 $\rm{Z}_{\odot}$, 0.1 $\rm{Z}_{\odot}$, 0.01 $\rm{Z}_{\odot}$), and pure helium. In the high-metallicity scenario, the inferred neutron star mass is $<1.3\,M_{\odot}$ with a radius $<8\,\rm{km}$. In the low-metallicity, hydrogen-rich case, the mass ranges from 0.3 to 2.1 $M_{\odot}$ with radii of 7-13 km. For a pure-helium composition, we find two mass solutions: $1.08_{-0.22}^{+1.32}M_{\odot}$ (with $R>14\,\rm{km}$) and $2.5-2.9\,M_{\odot}$ (above the highest observed neutron star masses). Additionally, we applied the touchdown method combined with an MCMC analysis, the results are consistent with those from the direct cooling tail method, but with a broader range. Our analysis of the time-resolved spectrum of burst suggests a high-metallicity atmosphere, but new observations are required to confirm this result.

astro-ph.HE

Mass-Radius Constraints for 2S 0918-549 from an RXTE Superexpansion Burst: A Direct Cooling-Tail Analysis

Thermonuclear (Type I ) X-ray bursts from accreting neutron stars offer a means to determine neutron-star (NS) mass ($M$) and radius ($R$) and thereby probe the properties of matter at supranuclear density. A subset of these events, photospheric radius-expansion (PRE) bursts, provide a particularly powerful tool to constrain the neutron-star $M$ and $R$. Here, we apply the direct cooling-tail method to 2S~0918$-$549, using a rare superexpansion burst observed by \emph{RXTE}. We fit only the post-touchdown data within \(F/F_{\rm td}\in[0.6,0.95]\), employing modern atmosphere models (pure He and metal-enriched). The pure-He atmosphere yields a good description of the cooling tail (\(\chi^{2}/\nu=18.12/14\)), whereas metal-rich models fail; information-criterion tests (AIC/BIC) disfavor adding a free absorption edge in every time bin, indicating that heavy-element ashes are unnecessary. The joint fit gives a distance \(d=4.1-5.3\) kpc and mass-radius constraints \(M=1-2\,M_\odot\) and \(R=9.7-11.9\) km (99\% confidence). These results suggest that representative families of both gravity-bound and self-bound equations of state remain viable at the $1\sigma$ confidence level.

astro-ph.HE

Safer by Diffusion, Broken by Context: Diffusion LLM's Safety Blessing and Its Failure Mode

Diffusion large language models (D-LLMs) offer an alternative to autoregressive LLMs (AR-LLMs) and have demonstrated advantages in generation efficiency. Beyond the utility benefits, we argue that D-LLMs exhibit a previously underexplored safety blessing: their diffusion-style generation confers intrinsic robustness against jailbreak attacks originally designed for AR-LLMs. In this work, we provide an initial analysis of the underlying mechanism, showing that the diffusion trajectory induces a stepwise reduction effect that progressively suppresses unsafe generations. This robustness, however, is not absolute. Following this analysis, we highlight a simple yet effective failure mode, context nesting, in which harmful requests are embedded within structured benign contexts. Empirically, we show that this simple black-box strategy bypasses D-LLMs' safety blessing, achieving state-of-the-art attack success rates across models and benchmarks. Notably, it enables the first successful jailbreak of Gemini Diffusion to our knowledge, exposing a critical vulnerability in proprietary D-LLMs. Together, our results characterize both the origins and the limits of D-LLMs' safety blessing, constituting an early-stage red-teaming of D-LLMs.

cs.LG

The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer

Recent advances in end-to-end trained omni-models have substantially improved audio capabilities by strengthening text-audio modality alignment. However, whether such alignment inadvertently facilitates the transfer of safety vulnerabilities across modalities remains underexplored. This question is critical as text-based jailbreak attacks are considerably more mature than audio-based ones; if they transfer systematically, current audio safety evaluations may underestimate risks originating from the text modality. In this paper, we introduce the Alignment Curse, a formally characterized and empirically validated principle showing that stronger modality alignment enables more effective transfer of attacks from text to audio, revealing a fundamental tension between capability and safety. Motivated by this principle, we conduct a comprehensive black-box evaluation of three attack categories on recent omni-models (e.g., Qwen2.5-Omni, Qwen3-Omni): text attacks, text-transferred audio attacks, and audio attacks. We find that text-transferred audio attacks perform comparably to, and often better than, audio-based attacks, exhibiting a clear advantage under audio-only access. This suggests that text-based vulnerabilities play a pivotal role in shaping audio safety risks. Finally, we empirically analyze the relationship between modality alignment and transfer effectiveness across attack methods and models, observing consistent support for the Alignment Curse: tighter modality alignment leads to more effective cross-modality attack transfer.

cs.LG

DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection

The misuse of advanced generative AI models has resulted in the widespread proliferation of falsified data, particularly forged human-centric audiovisual content, which poses substantial societal risks (e.g., financial fraud and social instability). In response to this growing threat, several works have preliminarily explored countermeasures. However, the lack of sufficient and diverse training data, along with the absence of a standardized benchmark, hinder deeper exploration. To address this challenge, we first build Mega-MMDF, a large-scale, diverse, and high-quality dataset for multimodal deepfake detection. Specifically, we employ 21 forgery pipelines through the combination of 10 audio forgery methods, 12 visual forgery methods, and 6 audio-driven face reenactment methods. Mega-MMDF currently contains 0.1 million real samples and 1.1 million forged samples, making it one of the largest and most diverse multimodal deepfake datasets, with plans for continuous expansion. Building on it, we present DeepfakeBench-MM, the first unified benchmark for multimodal deepfake detection. It establishes standardized protocols across the entire detection pipeline and serves as a versatile platform for evaluating existing methods as well as exploring novel approaches. DeepfakeBench-MM currently supports 5 datasets and 11 multimodal deepfake detectors. Furthermore, our comprehensive evaluations and in-depth analyses uncover several key findings from multiple perspectives (e.g., augmentation, stacked forgery). We believe that DeepfakeBench-MM, together with our large-scale Mega-MMDF, will serve as foundational infrastructures for advancing multimodal deepfake detection.

cs.CR

A transition from mixed-fuel to pure-helium thermonuclear bursts in Terzan 5 X-3/Swift J174805.3-244637

We presented a detailed analysis of seven thermonuclear X-ray bursts from Terzan 5 X-3/Swift J174805.3-244637, detected by NICER during the source's 2023 outburst. Our analysis reveals a clear evolution of burst properties, identifying four non-photospheric radius expansion (non-PRE) bursts, one PRE candidate occurring in a mixed hydrogen/helium environment, and two powerful PRE bursts from pure helium ignition. The time-resolved burst spectra were well described by a model including a variable persistent emission component, quantified by a factor $f_a$, due to the Poynting-Robertson drag. The strength of this interaction scales with burst luminosity: the enhancement is absent ($f_a \approx 1$) in the faintest bursts, becomes modest ($f_a \approx 1.5-2$) for the more luminous non-PRE burst and the PRE candidate, and is very strong ($f_a \approx 6-8$) during the pure-helium PRE bursts. This observed transition from mixed-fuel to pure-helium burning as the local mass accretion rate dropped below $\sim$10% of the Eddington limit, $\dot{m}_{\rm Edd}$, aligns with theoretical predictions. We verified this scenario with two independent methods. First, at the known distance to Terzan 5, the touchdown luminosities of both the pure helium PRE bursts and the mixed-fuel PRE candidate are consistent with reaching their respective, composition-dependent Eddington limits on the same plausible, massive neutron star of $\sim 2 M_\odot$. Second, the observed recurrence times of the non-PRE bursts were consistent with predictions for mixed-fuel burning.

astro-ph.HE

MoFO: Momentum-Filtered Optimizer for Mitigating Forgetting in LLM Fine-Tuning

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. Typically, LLMs are first pre-trained on large corpora and subsequently fine-tuned on task-specific datasets. However, during fine-tuning, LLMs may forget some knowledge acquired in the pre-training stage, leading to a decline in general capabilities. Existing approaches to mitigate forgetting often rely on access to pre-training data, which may be unavailable in many real-world scenarios--such as fine-tuning checkpoint-only open-source LLMs. To address this challenge, we propose a new fine-tuning algorithm termed Momentum-Filtered Optimizer (MoFO). MoFO is an extension of greedy block coordinate descent (BCD) methods: in each iteration, MoFO only updates the model parameters with the largest momentum magnitudes, while keeping all other parameters fixed. MoFO achieves similar fine-tuning performance to the default fine-tuning algorithm while effectively mitigating knowledge forgetting. We validate MoFO through rigorous convergence analysis and extensive experiments, demonstrating its effectiveness in mitigating forgetting without pre-training data.

cs.LG

Observatory Science with eXTP

Scheduled for launch in 2030, the enhanced X-ray Timing and Polarization (eXTP) telescope is a Chinese space-based mission aimed at studying extreme conditions and phenomena in astrophysics. eXTP will feature three main payloads: Spectroscopy Focusing Arrays (SFAs), Polarimetry Focusing Arrays (PFAs), and a Wide-field Camera (W2C). This white paper outlines observatory science, incorporating key scientific advances and instrumental changes since the publication of the previous white paper [1]. We will discuss perspectives of eXTP on the research domains of flare stars, supernova remnants, pulsar wind nebulae, cataclysmic variables, X-ray binaries, ultraluminous X-ray sources, AGN, and pulsar-based positioning and timekeeping.

astro-ph.IM

Exploring and Improving Initialization for Deep Graph Neural Networks: A Signal Propagation Perspective

Graph Neural Networks (GNNs) often suffer from performance degradation as the network depth increases. This paper addresses this issue by introducing initialization methods that enhance signal propagation (SP) within GNNs. We propose three key metrics for effective SP in GNNs: forward propagation, backward propagation, and graph embedding variation (GEV). While the first two metrics derive from classical SP theory, the third is specifically designed for GNNs. We theoretically demonstrate that a broad range of commonly used initialization methods for GNNs, which exhibit performance degradation with increasing depth, fail to control these three metrics simultaneously. To deal with this limitation, a direct exploitation of the SP analysis--searching for weight initialization variances that optimize the three metrics--is shown to significantly enhance the SP in deep GCNs. This approach is called Signal Propagation on Graph-guided Initialization (SPoGInit). Our experiments demonstrate that SPoGInit outperforms commonly used initialization methods on various tasks and architectures. Notably, SPoGInit enables performance improvements as GNNs deepen, which represents a significant advancement in addressing depth-related challenges and highlights the validity and effectiveness of the SP analysis framework.

cs.LG

Emission and Absorption Lines in Photospheric Radius Expansion Bursts of 4U 1820$-$30

We analyze the emission and absorption lines during photospheric radius expansion (PRE) X-ray bursts from the ultracompact binary 4U 1820--30, observed with the Neutron Star Interior Composition Explorer (NICER). Using Monte Carlo simulations to estimate the significance, we identified a 1 keV emission line from 14 bursts, a 3 keV absorption line from 12 bursts, and 1.6 keV absorption from one burst. By coadding the burst spectra at the maximum radius phase, we detected a 1.034 keV emission line with significance of $14.2σ$, and absorption lines at 1.64 and 3 keV with significances of $10.8σ$ and $11.7σ$, respectively. The observed energy shifts are consistent with the prediction from the burst-driven wind model, indicating that all three spectral features are produced by the PRE wind. Analysis of the ratios between the emission and absorption line energies suggests that the 1 keV feature is a superposition of several narrower Fe L-shell lines. To evaluate the scientific capabilities of the Hot Universe Baryon Surveyor (HUBS), we simulated mock observations of multiple narrow lines near 1 keV. The results demonstrate that HUBS is well suited for detailed studies of the 1 keV emission line during bursts, offering significant potential to advance our understanding of these phenomena.

astro-ph.HE

A Systematic Study of Millihertz Quasiperiodic Oscillations in GS 1826-238

We performed a systematic investigation of millihertz quasiperiodic oscillations (mHz QPOs) in the low-mass X-ray binary GS 1826$-$238 observed with {\it NICER} and {\it Insight}-HXMT. We discovered 37 time intervals exhibiting mHz QPOs out of 106 Good Time Interval (GTI) samples in the frequency range of 3$-$17 mHz at a significance level of $>99.99\%$. The source remains in a soft state in our study. No significant differences are found between the samples with and without mHz QPOs according to positions in the color-color and hardness-intensity diagrams. These QPOs were discovered at an accretion rate of $\sim 0.1 \dot{M}_{\rm Edd}$, similar to other sources. The broadband spectrum of GS 1826$-$238 can be modeled as a combination of a multicolor blackbody from the accretion disk and a Comptonization with seed photons emitted from the neutron star (NS) surface. The flux modulations of mHz QPOs are related to variations of the temperature of Comptonization seed photons, consistent with the marginally stable burning theory.

astro-ph.HE

A Phase-resolved View of Millihertz Quasi-periodic Oscillations in the Ultraluminous X-ray Source M51 ULX-7: Evidence for a Magnetically Truncated Disk and Geometrical Beaming

X-ray quasi-periodic oscillations (QPOs) are commonly observed in Galactic X-ray binaries (XRBs) and extragalactic ultraluminous X-ray sources (ULXs). In this study, we perform a phase-resolved analysis of recently discovered X-ray millihertz QPOs in M51 ULX-7. This represents the first detailed phase-resolved analysis of QPOs conducted in ULXs. Our findings reveal that the amplitude of the mHz QPO slightly increases with photon energy, accompanied by a narrowing of the phase modulation profile. The phase-resolved spectroscopy indicates significant variability in the energy spectrum: both disk blackbody components exhibit marked variations on the QPO timescale, with the low-temperature component demonstrating significant synchronous changes in the disk temperature and luminosity, showing a positive correlation between these two parameters throughout the QPO cycle. This correlation supports the hypothesis that the disk inner radius corresponds to the magnetospheric radius, which slightly varies with the accretion rate. Our results suggest that the soft component, without beaming, originates from a magnetically truncated outer disk, while the hard component is geometrically beamed from the inner funnel regions.

astro-ph.HE

Quenching and recovery of persistent X-ray emission during a superburst in 4U 1820$-$30

We report the superburst from 4U 1820--30 in 2021 observed by the Monitor of All-sky X-ray Image and Neutron star Interior Composition Explorer (NICER). During the tail of the superburst, we found that the NICER light curve unexpectedly increased from 1080 to 2204 ${\rm counts~s^{-1}}$ over 6.89 hr. From the time-resolved superburst spectra, we estimated the burst decay time of $\approx2.5$ hr, the ignition column depth of $\approx0.3\times 10^{12}~{\rm g ~cm^{-2}}$, the energy release per unit mass of $\approx2.4\times 10^{17}~{\rm erg~g^{-1}}$, the fluence of $\approx4.1\times 10^{-4}~{\rm erg~cm^{-2}}$, and the total energy release of $\approx3.5\times10^{42}$ erg. Notably, we found a gradual increase in the Componization flux from $8.9\times 10^{-10}~{\rm erg~s^{-1}~cm^{-2}}$ to the preburst level during the superburst. This increase can be interpreted as a consequence of superburst radiation depleting the inner accretion disk, leading to a near-complete quenching of the persistent emission. As the burst radiation decayed, the inner accretion disk gradually returned to its preburst state, as evidenced by the best-fit spectral parameters. Additionally, we observed a prominent absorption line that exhibited a gravitational redshift, shifting from 4.15 to 3.62 keV during the recovery phase of persistent emission. This absorption feature likely originates from the inner accretion disk rather than from burst emission on the neutron star (NS) surface. The observed changes in the absorption line energy suggest that the inner disk approached the NS to a distance as close as $\approx17$ km.

astro-ph.HE

A comprehensive study of type I (thermonuclear) bursts in the new transient SRGA J144459.2$-$604207

We report an analysis of Insight-HXMT observations of the newly discovered accreting millisecond pulsar SRGA J144459.2$-$604207. During the outburst, detected in 2024 February by SRG/ART-XC, the broadband persistent spectrum was well fitted by an absorbed Comptonization model. We detected 60 type I X-ray bursts in the Insight-HXMT medium energy (ME) data, and 37 were also detected with the low-energy (LE) telescope. By superimposing the Insight-HXMT/LE/ME/HE light curves of 37 bursts with similar profiles and intensities, we measured a deficit of X-rays in the $40-70$ keV energy band. By analyzing the time-resolved X-ray burst spectra, we determine the mean ratio of persistent to burst flux of $α=71\pm7$. We estimate the average hydrogen mass fraction in the fuel at ignition, as $\bar{X} =0.342\pm0.033$, and constrain the burst fuel composition as $X_0\leq0.4$. We found that 14 out of 60 X-ray bursts exhibited photospheric expansion, and thus we estimated the distance to the source as $10.0\pm0.71$ kpc. Combined with IXPE observations, the burst recurrence time increased from 1.55 to 8 hr as the local mass accretion rate decreased, which can be described as $ΔT_{\rm rec}\sim \dot{m}^{-0.91\pm0.02}$.

astro-ph.HE