SearcharxivSearch

arXiv subjects

Liyang Chen

Publications and source records attributed to Liyang Chen.

At least 19 recordsLinked to original sources

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.

cs.CV

Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization

Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely. We reveal that this trade-off arises not from the constraints of streaming architectures, but from an inappropriate choice of optimization anchor. Directly optimizing against audio quality metrics induces catastrophic reward hacking, where content critical to pronunciation and intelligibility is systematically erased to maximize a proxy score. To break this bottleneck, we propose two complementary improvements: an enlarged Conformer convolution kernel for richer local spectro-temporal modeling, and WavLM-anchored Direct Preference Optimization (DPO) fine-tuning strategy. DPO preference pairs are ranked by WavLM cosine similarity, a deep acoustic feature encoding both phonetic structure and speaker identity, providing an optimization anchor that resists hacking. Under a 560 ms streaming chunk size, the proposed method achieves a 10.9% relative intelligibility improvement (word error rate: 0.138 to 0.123), with marginal simultaneous gains in audio quality and speaker similarity.

cs.SD

StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model

While generative models have set new benchmarks for Target Speaker Extraction (TSE), their inherent reliance on global context precludes deployment in real-time applications. Direct adaptation to streaming scenarios often leads to catastrophic inference performance degradation due to the severe mismatch between training and streaming inference. To bridge this gap, we present the first autoregressive (AR) models tailored for streaming TSE. Our approach introduces a Chunk-wise Interleaved Splicing Paradigm that ensures highly efficient and stable streaming inference. To ensure the coherence between the extracted speech segments, we design a historical context refinement mechanism that mitigates boundary discontinuities by leveraging historical information. Experiments on Libri2Mix show that while AR generative baseline exhibits performance degradation at low latencies, our approach maintains 100% stability and superior intelligibility. Furthermore, our streaming results are comparable to or even surpass offline baselines. Additionally, our model achieves a Real-Time-Factor (RTF) of 0.248 on consumer-level GPUs. This work provides empirical evidence that AR generative backbones are viable for latency-sensitive applications through the Chunk-wise Interleaved Splicing Paradigm.

cs.SD

Seedance 2.0: Advancing Video Generation for World Complexity

Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and large-scale architecture for multi-modal audio-video joint generation. This allows it to support four input modalities: text, image, audio, and video, by integrating one of the most comprehensive suites of multi-modal content reference and editing capabilities available in the industry to date. It delivers substantial, well-rounded improvements across all key sub-dimensions of video and audio generation. In both expert evaluations and public user tests, the model has demonstrated performance on par with the leading levels in the field. Seedance 2.0 supports direct generation of audio-video content with durations ranging from 4 to 15 seconds, with native output resolutions of 480p and 720p. For multi-modal inputs as reference, its current open platform supports up to 3 video clips, 9 images, and 3 audio clips. In addition, we provide Seedance 2.0 Fast version, an accelerated variant of Seedance 2.0 designed to boost generation speed for low-latency scenarios. Seedance 2.0 has delivered significant improvements to its foundational generation capabilities and multi-modal generation performance, bringing an enhanced creative experience for end users.

cs.CV

SN2024abfl: A Low-Luminosity Type IIP Supernova at the Low-Mass End of Core Collapse

We present optical photometric and spectroscopic observations of the low-luminosity (LL) Type IIP supernova SN\,2024abfl. The distance to its host galaxy is highly uncertain, with independent estimates of $9.5^{+2.3}_{-2.4}$ Mpc and $15.0^{+8.9}_{-1.9}$ Mpc. Even adopting the larger distance, the inferred plateau luminosity is only $\sim 10^{41}\rm erg\,s^{-1}$, placing SN 2024abfl at the extreme faint end of SNe IIP population. Its light curve exhibits a long-lasting plateau of approximately 110 days. The spectra show exceptionally low expansion velocities, with the \FeII\, velocity of $\sim1200\,\rm km\,s^{-1}$ at 50 days after the explosion, significantly lower than the typical values of $\sim2000-5500\,\rm km\,s^{-1}$ observed in SNe IIP, placing SN\,2024abfl among the slowest-expanding LL SNe IIP. Bolometric modeling yields a synthesized $^{56}$Ni mass of $\sim0.002-0.004\,\rm M_\odot$, though this estimate remains subject to significant uncertainty owing to the poorly constrained distance. Considering the plateau color and duration, the magnitude drop from plateau to tail, and the progenitor luminosity, we favor a low-mass core-collapse origin for SN\,2024abfl.

astro-ph.SR

Spectral Dataset of Stripped-Envelope Supernovae from the Tsinghua Supernova Group

The extent of envelope stripping in the progenitor stars is directly reflected in the diversity of spectral features observed in stripped-envelope supernovae (SESNe). Through extensive spectral observation and analysis, we aim to clarify the statistical differences between the subclasses of SESNe. The Tsinghua Supernova group obtained 249 optical spectra of 62 SESNe during the years from 2010 to 2020, covering phases from $-$16 to over 190 days relative to maximum light. Most spectra were obtained during the photospheric phases after the supernova explosion. For each spectrum, the pseudo-equivalent widths (pEWs) and blueshift velocities of principal lines were measured. We further investigated the common spectral features by analysing their velocity and strength correlations across all subtypes. We identify the feature near 6200~\AA\ in SNe Ib as H$\mathrm{\alpha}$ through comparison with SNe IIb and Ic, which resolves inconsistent literature interpretations. Our finding reveals prevalent residual hydrogen in SNe Ib, further supporting a continuous stripping sequence from SNe IIb to Ib. We observe a trend in increasing velocity among different subtypes of stripped-envelope SNe, with SNe IIb exhibiting the lowest line velocities, followed by Ib, Ic, and Ic-BL. Typically, the O~I lines in SNe Ic/Ic-BL are stronger than those seen in SNe IIb/Ib. In nebular phases, the [Ca II] emission dominates over [O I] in SNe IIb/Ib while [O I] is stronger in SNe Ic, including the He-rich SN 2016coi. This spectral dichotomy implies that progenitors of SNe Ic (BL) have more massive CO cores and hence higher initial masses.

astro-ph.HE

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and audio-driven video animation (RA2V) as isolated objectives. Furthermore, achieving precise, disentangled control over multiple character identities and voice timbres within a single framework remains an open challenge. In this paper, we propose DreamID-Omni, a unified framework for controllable human-centric audio-video generation. Specifically, we design a Symmetric Conditional Diffusion Transformer that integrates heterogeneous conditioning signals via a symmetric conditional injection scheme. To resolve the pervasive identity-timbre binding failures and speaker confusion in multi-person scenarios, we introduce a Dual-Level Disentanglement strategy: Synchronized RoPE at the signal level to ensure rigid attention-space binding, and Structured Captions at the semantic level to establish explicit attribute-subject mappings. Furthermore, we devise a Multi-Task Progressive Training scheme that leverages weakly-constrained generative priors to regularize strongly-constrained tasks, preventing overfitting and harmonizing disparate objectives. Extensive experiments demonstrate that DreamID-Omni achieves comprehensive state-of-the-art performance across video, audio, and audio-visual consistency, even outperforming leading proprietary commercial models. We will release our code to bridge the gap between academic research and commercial-grade applications.

cs.CV

AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning

Large Audio Language Models (LALMs) have demonstrated strong capabilities in audio understanding and reasoning. However, their performance on fine grained auditory perception remains unreliable, and existing approaches largely rely on data intensive training to internalize perceptual abilities. We propose AudioRouter, a reinforcement learning framework that enables LALMs to improve audio understanding by learning when and how to use external audio tools. Rather than tightly coupling tool usage with audio reasoning, AudioRouter formulates tool use as an explicit decision making problem and optimizes a lightweight routing policy while keeping the underlying reasoning model frozen. Experimental results show that AudioRouter achieves substantial improvements on standard audio understanding benchmarks while requiring up to 600x less training data to learn tool usage compared with conventional training paradigms. These findings suggest that learning effective tool usage offers a data efficient and scalable alternative to internalizing perceptual abilities in LALMs.

cs.SD

SN 2024igg: A Super-Chandrasekhar/03fg-like SN exhibiting C II-dominated spectra after explosion

We present and analyze photometric and spectroscopic observations of the Type Ia supernova (SN Ia) 2024igg, another ``super-Chandrasekhar'' (or 03fg-like) SN whose strong C II $\lambda6580$ feature was initially misidentified as H$\alpha$, thereby constraining its progenitor system, explosion parameters, and physical scenario. SN 2024igg shows many characteristics in common with other 03fg-like objects, such as high ultraviolet flux, slowly declining light curves ($\Delta m_{15}(B)=0.90\pm0.08$ mag), low expansion velocities, along with strong and persistent C II absorption. Meanwhile, this SN exhibits some remarkable properties within this subgroup, including a moderately low optical luminosity ($M_{\rm max}(B)=-18.99\pm0.15$ mag), a short rise time less than 18.5 days, and strong C II $\lambda6580$. The bolometric analysis yields a $^{56}$Ni mass of $M_{\rm Ni}=0.547\pm0.082$ $M_{\rm \odot}$ and an ejecta mass of $1.54^{+0.22}_{-0.19}$ $M_{\rm \odot}$, marginally exceeding the Chandrasekhar mass. Our TARDIS result indicates that most of the features in the earliest spectrum could be attributed to C II, which is consistent with a model where a supernova explodes within a carbon-rich circumstellar medium (CSM). The CSM interaction would produce a density peak in the ejecta, offering a natural explanation for the slowly evolving line velocities near $-$8000 km s$^{-1}$. The CSM may stem from the debris of a secondary white dwarf in a white-dwarf merger or the envelope of an asymptotic giant branch star. Combined with the unshifted forbidden lines in the spectrum taken at $t\approx\ +$135 days, we suggest that SN 2024igg comes from a symmetric explosion on a secular timescale after the merger.

astro-ph.HE

Disparate Quantum Corrections to Conduction in Carbon Nanotube Bundles

Quantum interference effects such as weak localization (WL) and universal conductance fluctuations (UCF) normally yield consistent electronic phase-coherence lengths in homogeneous conductors. Here we show that in individual carbon nanotube bundles exfoliated from highly conductive solution-spun fibers, different probes, including the field scales and magnitudes of WL and UCF and nonlocal magnetoconductance, lead to strikingly disparate estimates of coherence lengths. WL magnetoconductance measured in a perpendicular magnetic field yields a phase-coherence length of approximately 50 nm. In contrast, UCF amplitudes are comparable to e squared over h even for an 8 micrometer long segment, and nonlocal magnetoconductance persists across a 4 micrometer separation of electrodes, revealing phase-coherent transport over micrometer length scales within a single bundle. The coexistence of short- and long-range coherence implies that locally diffusive electrons remain partially phase-correlated among nanotubes within the same bundle. These findings challenge the conventional single-scale picture of mesoscopic coherence and establish carbon nanotube bundles as a model platform for emergent, network-level quantum transport.

cond-mat.mes-hall

OptiSQL: Executable SQL Generation from Optical Tokens

Executable SQL generation is typically studied in text-to-SQL settings, where tables are provided as fully linearized textual schemas and contents. While effective, this formulation assumes access to structured text and incurs substantial token overhead, which is misaligned with many real-world scenarios where tables appear as visual artifacts in documents or webpages. We investigate whether compact optical representations can serve as an efficient interface for executable semantic parsing. We present OptiSQL, a vision-driven framework that generates executable SQL directly from table images and natural language questions using compact optical tokens. OptiSQL leverages an OCR-oriented visual encoder to compress table structure and content into a small set of optical tokens and fine-tunes a pretrained decoder for SQL generation while freezing the encoder to isolate representation sufficiency. Experiments on a visualized version of Spider 2.0-Snow show that OptiSQL retains strong execution accuracy while reducing table input tokens by an order of magnitude. Robustness analyses further demonstrate that optical tokens preserve essential structural information under visual perturbations.

cs.CL

SN 2024abvb: A Type Icn Supernova in the Outskirts of its Host Galaxy

We present multiband photometric and spectroscopic observations of supernova (SN) 2024abvb, which exhibits early-time prominent photoionized narrow emission lines of C II superposed on a blue continuum. The absence of Balmer features indicates that the SN exploded within hydrogen-poor circumstellar matter (CSM). Together with the lack of explicit evidence of helium signatures, we tentatively identify SN 2024abvb as a Type Icn SN (SN Icn). After correcting for extinction, we estimate an r-band peak absolute magnitude of -19.7, placing SN 2024abvb in the luminous regime of SNe Icn. We adopted a hybrid model that accounts for both the energy released by the ejecta-CSM interaction and the radioactive decay of nickel synthesized in the SN ejecta to fit the light curve of SN 2024abvb. The best-fit model to the multiband light curves within the first ~ 40 days after explosion suggests that the CSM, radioactive nickel, and ejecta masses to be 0.28 Msun, < 3.8 * 10^-2 Msun, and 0.12 Msun, respectively. Such a low ejecta mass indicates that the progenitor star of SN 2024abvb experienced a significant mass-stripping process, consistent with the hydrogen-poor and helium-poor spectral features. SN 2024abvb provides important insights into the physical origins of the rare subclass of SNe Icn.

astro-ph.HE

The Double-Peaked Calcium-Strong SN 2025coe: Progenitor Constraints from Early Interaction and Ejecta Asymmetries

Supernova (SN) 2025coe at a distance of $\sim$25 Mpc is the second-closest calcium-strong (CaST) transient. It was discovered at a large projected offset of $\sim$34 kpc from its potential host galaxy NGC 3277. Multiband photometry of SN 2025coe indicates the presence of two peaks at day $\sim$2 and day $\sim$11 after explosion. Modeling the bolometric light curve, we find that the first peak can be reproduced either by shock cooling of a compact envelope ($R_\mathrm{env}$ $\approx $6-40 $R_{\odot}$; $M_\mathrm{env}$ $\approx $0.1-0.2 $M_{\odot}$) or by interaction with close-in circumstellar material (CSM; $R_{\mathrm{CSM}} \lesssim 6 \times10^{14}$ cm), or a combination of both. The second peak is dominated by radioactive decay of $^{56}$Ni ($M_{\mathrm{ej}} \approx $0.4-0.5 $M_{\odot}$; $M_{^{56}\mathrm{Ni}} \approx 1.4 \times 10^{-2}$ $M_{\odot}$). SN 2025coe rapidly evolves from the photospheric phase dominated by He I P-Cygni profiles to nebular phase spectra dominated by strong [Ca II] $\lambda \lambda$7291, 7323 and weak [O I] $\lambda \lambda$6300, 6364 emission lines. Simultaneous line profile modeling of [Ca II] and [O I] at nebular phases shows that an asymmetric core-collapse explosion of a low-mass ($\lesssim$3.3 $M_{\odot}$) He-core progenitor can explain the observed line profiles. Alternatively, lack of local star formation at the site of the SN explosion combined with a low ejecta mass is also consistent with a thermonuclear explosion due to a low-mass hybrid He-C/O white dwarf + C/O white dwarf merger.

astro-ph.HE

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via mask-based inpainting. However, masking inevitably destroys spatiotemporal context, leading to identity drift and poor robustness (e.g., to occlusions), while also inducing lip-shape leakage that degrades lip sync. To bridge this gap, we propose X-Dub, a novel two-stage generative bootstrapping framework leveraging powerful Diffusion Transformers to unlock mask-free dubbing. Our core insight is to repurpose a mask-based inpainting model exclusively as a dedicated data generator to synthesize scalable, high-fidelity pseudo-paired data, which is subsequently utilized to train and bootstrap a robust, mask-free editing model as the final video dubber. The final dubber is liberated from masking artifacts and leverages the complete video input for high-fidelity inference. We further introduce timestep-adaptive multi-phase learning to disentangle conflicting objectives (structure, lip motion, and texture) across diffusion phases, facilitating stable convergence and advanced editing quality. Additionally, we present X-DubBench, a benchmark for diverse scenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance with superior lip sync, visual quality, and robustness.

cs.CV

SN 2024iss: A Double-peaked Type IIb Supernova with Evidence of Circumstellar Interaction

We present optical, ultraviolet, and X-ray observations of supernova (SN) 2024iss, a Type IIb SN that shows a prominent double-peaked light curve. We modeled the first peak with a semianalytical shock-cooling model and the X-ray emission with a free-free model. We compare the envelope radius and mass-loss rate with other Type IIb SNe to explore the relationships between the progenitor envelope and the circumstellar material (CSM). The shock-cooling peak in the $V$-band light curve reached $M_V = -17.33\pm 0.26$mag, while the $^{56}$Ni-powered second peak attained $M_V = -17.43\pm 0.26$mag. Early spectra show an photospheric velocity of $\sim19,400\,km\,s^{-1}$ at 3.82days from the H$\alpha$ P~Cygni profile. The Balmer lines persist at least +87 days after the explosion, characterizing hydrogen-rich ejecta. Modeling the first light-curve peak suggests an extended envelope with a mass of $0.11\pm0.04\,M_{\odot}$ and a radius of $244\pm43~R_{\odot}$. Fitting the second light-curve peak with an Arnett-like model indicates a typical $^{56}$Ni mass of $ 0.117\pm0.013~M_{\odot}$ and a relatively low ejecta mass of $1.272\pm0.343\,M_{\odot}$. X-ray observations reveal bright thermal bremsstrahlung emission and indicate a mass-loss rate of $1.6\times10^{-5}\ M_{\odot} \ \rm{yr}^{-1}$. SN 2024iss occupies a transitional position between the two subclasses of extended (eIIb) and compact (cIIb) Type IIb SNe. Its envelope radius and pre-explosion mass-loss rate appear to be correlated as theoretically predicted. The observational properties of SN 2024iss are compatible with a binary interaction scenario being the dominant mechanism for envelope stripping. Furthermore, the low column density of neutral hydrogen suggests a compact CSM with an outer radius of $\lesssim1.3\times10^{14}$ cm, indicating that the progenitor star experienced eruptive mass loss within $\sim4\,yr$ of its terminal explosion.

astro-ph.HE

Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector

Vision Language Models (VLMs) have achieved impressive progress in multimodal reasoning; yet, they remain vulnerable to hallucinations, where outputs are not grounded in visual evidence. In this paper, we investigate a previously overlooked setting: logo hallucination, where models generate brand names or textual content despite logos containing no visible words. Using curated splits of pure symbols, hybrids, and text-bearing logos, as well as the challenging Hard-60 subset, we systematically measure hallucination across leading VLMs. We further probe robustness through nine structured perturbations and show that hallucinations persist even under strong distortions, with occlusion exposing the sharpest weaknesses. Embedding-level analysis with open-weight LLaVA demonstrates that hallucination is tied to a small subset of projector dimensions, and targeted ablation substantially reduces errors while preserving OCR accuracy. Together, these findings reveal that VLMs often rely on symbolic priors rather than genuine glyph perception, particularly for iconic circular logos, and that projector subspaces play a decisive role in this failure mode. Our work contributes both a novel diagnostic lens and actionable mitigation insights, highlighting projector disentanglement and OCR-guided decoding as promising directions for building more trustworthy multimodal systems.

cs.CV

Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation

Video-to-Audio generation has made remarkable strides in automatically synthesizing sound for video. However, existing evaluation metrics, which focus on semantic and temporal alignment, overlook a critical failure mode: models often generate acoustic events, particularly speech and music, that have no corresponding visual source. We term this phenomenon Insertion Hallucination and identify it as a systemic risk driven by dataset biases, such as the prevalence of off-screen sounds, that remains completely undetected by current metrics. To address this challenge, we first develop a systematic evaluation framework that employs a majority-voting ensemble of multiple audio event detectors. We also introduce two novel metrics to quantify the prevalence and severity of this issue: IH@vid (the fraction of videos with hallucinations) and IH@dur (the fraction of hallucinated duration). Building on this, we introduce HALCON to mitigate IH. HALCON follows a three-stage procedure: it first generates initial audio to expose hallucinated segments, then identifies and masks the corresponding unreliable video features, and finally regenerates the audio using the corrected conditioning. Experiments on several mainstream V2A benchmarks first reveal that state-of-the-art models suffer from severe IH. In contrast, our HALCON method reduces both the prevalence and duration of hallucinations by over 50\% on average, without degrading, and in some cases even improving, conventional metrics for audio quality and temporal synchronization. Our work is the first to formally define, systematically measure, and effectively mitigate Insertion Hallucination, paving the way for more reliable and faithful V2A models.

cs.SD

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

The visual dubbing task aims to generate mouth movements synchronized with the driving audio, which has seen significant progress in recent years. However, two critical deficiencies hinder their wide application: (1) Audio-only driving paradigms inadequately capture speaker-specific lip habits, which fail to generate lip movements similar to the target avatar; (2) Conventional blind-inpainting approaches frequently produce visual artifacts when handling obstructions (e.g., microphones, hands), limiting practical deployment. In this paper, we propose StableDub, a novel and concise framework integrating lip-habit-aware modeling with occlusion-robust synthesis. Specifically, building upon the Stable-Diffusion backbone, we develop a lip-habit-modulated mechanism that jointly models phonemic audio-visual synchronization and speaker-specific orofacial dynamics. To achieve plausible lip geometries and object appearances under occlusion, we introduce the occlusion-aware training strategy by explicitly exposing the occlusion objects to the inpainting process. By incorporating the proposed designs, the model eliminates the necessity for cost-intensive priors in previous methods, thereby exhibiting superior training efficiency on the computationally intensive diffusion-based backbone. To further optimize training efficiency from the perspective of model architecture, we introduce a hybrid Mamba-Transformer architecture, which demonstrates the enhanced applicability in low-resource research scenarios. Extensive experimental results demonstrate that StableDub achieves superior performance in lip habit resemblance and occlusion robustness. Our method also surpasses other methods in audio-lip sync, video quality, and resolution consistency. We expand the applicability of visual dubbing methods from comprehensive aspects, and demo videos can be found at https://stabledub.github.io.

cs.CV