SearcharxivSearch

arXiv subjects

Huan Yang

Publications and source records attributed to Huan Yang.

At least 19 recordsLinked to original sources

From the Test-Mass Limit to Binary Black-Hole Waveforms in Higher-Derivative Gravity

Many higher-derivative theories predict stronger deviations from General Relativity for lower-mass black holes, while their nonlinear field equations often prevent reliable simulations of the full binary evolution. Here we develop a route from controlled black hole perturbation theory based on the modified Teukolsky formalism to comparable-mass waveforms, using parity-even cubic gravity as a representative example. We find that the tidal response of the secondary black hole enters at the same perturbative order as the direct higher-curvature correction and is therefore essential for a consistent leading-order waveform. The resulting strong-field fluxes and conservative dynamics produce an accumulated inspiral dephasing that grows toward merger. Embedding this test-mass information into an effective-one-body model, we construct inspiral-merger-ringdown waveforms for comparable-mass binaries and find coupling-dependent dephasing and waveform-peak shifts. Our results demonstrate how strong-field test-mass calculations can anchor waveform models for higher-derivative gravity when theory-specific numerical-relativity simulations are unavailable.

gr-qc

Imaging Stars at the Quantum Compatibility Limit

Imaging astrophysical sources with a multi-station interferometer is intrinsically a multiparameter quantum-estimation problem. {Using tools from multiparameter quantum metrology,} we show that time-resolved repetitive or adaptive measurements in an \(N\)-station array suffer a fundamental array-level incompatibility among visibility estimators. Collective measurements, {which coherently process the received starlight across multiple time bins in a single joint readout}, remove the array-size penalty up to an order-unity factor, yielding an asymptotic \(O(\sqrt{N})\) enhancement for the {directional-averaged} SNR of visibility measurement. We then propose a memory-assisted interferometric architecture designed to implement collective readout through coherent storage and joint quantum processing. Imaging simulations and Fisher-information analyses demonstrate that collective measurements improve image reconstruction in near-term arrays and enhance the resolving power of future long-baseline architectures, with pronounced benefits for representative AGN targets such as NGC~4151 and 3C~273. These results highlight collective measurement as a promising building block for future quantum-assisted interferometric arrays for stellar imaging.

quant-ph

SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling

Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span different depth layers or are partially occluded. Most image editing methods focus solely on prior information from 2D datasets, emphasizing planar features while lacking support for spatial structures. Even approaches that incorporate explicit positional information fail to capture true 3D spatial relationships, thus limiting accurate object movement in complex scenes. In this paper, we present SpatialDiff, a method that effectively captures 3D spatial structures, enabling precise and consistent object movements in complex scenes. Our core innovations are twofold: (1) Implicit 3D Spatial Modeling, which introduces 3D prior knowledge and enables the model to internally build a comprehensive understanding of the three-dimensional spatial structure; and (2) Global Spatial Supervision, which constrains the latent spatial features to enable the model to perceive changes in object spatial positions caused by editing operations. Experimental results demonstrate that our method significantly improves the accuracy and fidelity of spatial movement in complex scenes.

cs.CV

Gravitational Waves from Green's Function Decomposition for a Kerr black hole: I. Equatorial ISCO Plunge

We present a decomposition of the Kerr Green's function in the time domain, motivated by the frequency-domain split previously studied in the Schwarzschild limit. We show that the identification of a quasinormal-mode contribution, a direct part, and a late-time tail is still available, where the split times are determined by the black hole spin and positions of the emitter and receiver. We have checked this Green's function with time-domain Teukolsky numerical simulations and find excellent agreement. We also apply this decomposed Green's function in the time domain to a model problem with a test particle plunging into a Kerr black hole. The dynamically excited direct wave and quasinormal modes are obtained by convoluting the Green's function with the particle's source term, which may be viewed as the first order in mass ratio of a spinning black hole ringdown.

gr-qc

Identifying Kilonovae in the Presence of Optical Afterglow for the Wide-Field Survey Telescope

Identifying kilonovae associated with binary neutron star mergers is often complicated by the presence of a dominant synchrotron afterglow. In this work, we evaluate the performance of the Wide-Field Survey Telescope (WFST) in identifying kilonova signals in composite afterglow-kilonova transients. Using a numerical framework based on the Fisher information matrix, we simulate $10,000$ realizations for each of two scenarios: an AT2017gfo-based template model and a physically sampled population that accounts for kilonova diversity. Our results indicate that kilonova identification is primarily limited by source distance. In both scenarios, the identification efficiency is largely insensitive to variations in afterglow microphysical parameters and exceeds $80\%$ at distances within approximately $600~\rm Mpc$ for AT2017gfo-like events. Under our adopted assumptions, we estimate that WFST could identify approximately $1$--$16$ kilonovae per year. Furthermore, we find that the discriminating power of color-based filters rapidly saturates, reaching a stable plateau by the second night after the merger. We therefore propose a staged observing strategy that prioritizes high-cadence $g$ and $r$-band monitoring during the first night and incorporates the $z$ band from the second night onward. This strategy improves the identification precision by exploiting the increasingly prominent red excess produced by the kilonova. Our results provide a physical basis for optimizing WFST observing resources to efficiently detect and characterize kilonovae in the multimessenger era.

astro-ph.HE

Torsional-X Seismometer for Lunar Decihertz Gravitational-Wave Detection

The lunar gravitational-wave antenna concept uses the Moon as a resonant detector instrumented with precision seismometers, targeting the decihertz band between ground- and space-based observatories. We propose a compact monolithic fused-silica torsional-X seismometer that re-engineers garden-gate acceleration-to-rotation transduction for this regime through a high-tension dual-fiber suspension. Its designed millihertz-scale resonance and ultra-low mechanical dissipation enable a nearly order-of-magnitude improvement around $0.1\,\mathrm{Hz}$ compared with existing lunar seismometer concepts. Achieving this performance requires room-temperature operation, where fused-silica exhibits low mechanical loss, together with subdominant actuation noise. We demonstrate a room-temperature vacuum prototype validating the operating principle and core mechanical design, and derive requirements for a future lunar implementation capable of approaching the target sensitivity.

astro-ph.IM

NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of this drift: across three post-training methods (NFT, AWM, DPO), RL fine-tuning inflates the per-step velocity norm $\|v_\theta\|$ by $5\%$ to $15\%$ relative to the reference. A form of norm inflation has been studied in classifier-free guidance (CFG), where rescaling the velocity back to a reference norm at inference time can mitigate the resulting artifacts. However, this inference-time correction does not transfer cleanly to RL: rescaling $v_\theta$ to match $\|v_{\text{ref}}\|$ at inference time neither improves reward nor fixes the quality degradation, because the inflation is co-adapted into the model weights. Furthermore, an adjoint sensitivity analysis shows that velocity magnitude rescaling carries no coherent first-order reward signal at the batch level, indicating that suppressing norm inflation is unlikely to remove a consistently reward-carrying component. Since inference-time renormalization fails while norm suppression carries no reward cost, training-time intervention is the appropriate strategy. Together, these findings motivate NormGuard, a hinge penalty that activates only when $\|v_\theta\|$ exceeds $\|v_{\text{ref}}\|$ and composes additively with any velocity-local base loss. Across two base models, three post-training methods, and two reward proxies, NormGuard consistently improves MLLM-judged image quality and forensic realism while preserving reward, with gains that amplify under few-step inference and are not explained by early stopping.

cs.LG

A UAV-Based Multi-Modal Vision System for Automated Sideslope Deformation Monitoring and Hazard Detection

Slope hazards constitute a major safety threat to expressway infrastructure, and their evolution is typically manifested as slow surface deformation. Conventional manual inspection suffers from low efficiency and inadequate operational safety, especially on severely deteriorated slopes. Accordingly, there is an urgent need for an automated, high-precision solution capable of large-area slope observation and analysis. This study aims to develop a highly automated workflow for slope hazard detection using Unmanned Aerial Vehicle (UAV)-borne Light Detection and Ranging (LiDAR). The proposed workflow consists of a shared data-acquisition and ground-surface extraction stage, a single-observation hazard-screening branch based on RandLA-Net, and a multi-epoch deformation-monitoring branch based on grid-wise elevation differencing. To validate the effectiveness of the proposed system, we conducted multiple UAV-borne LiDAR data-acquisition flights in real expressway slope environments. The results show that the workflow can extract usable ground-surface point clouds under vegetation cover, identify potential hazard zones from single-observation point clouds, and quantify centimeter-level elevation changes using multi-epoch grid differencing. This study establishes an end-to-end UAV-borne LiDAR-based workflow for slope inspection and demonstrates its feasibility through controlled experiments, field tests, and simulation-based validation, thereby providing an implementable solution for automated slope-hazard monitoring and intelligent early warning.

cs.CV

Modified Teukolsky Formalism for Extreme Mass-Ratio Inspirals in Higher-Derivative Gravity

In this work, we study a model problem involving a point particle spiraling into a non-rotating black hole in higher-derivative theories of gravity. In such theories, both the background spacetime and the generation and propagation of gravitational waves differ from those in General Relativity. We develop a modified Teukolsky formalism to describe gravitational waves sourced by the point particle and, as an illustrative example, compute the resulting fluxes to the black hole horizon and null infinity for a cubic gravity theory. The formalism is constructed in a way that can be naturally extended to rotating black holes. These results represent essential steps to build extreme mass-ratio-inspiral waveforms in modified gravity theories, which may also be rescaled to approximate waveforms from comparable-mass binary black hole systems, analogous to existing approaches in General Relativity.

gr-qc

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training

Representation alignment with pretrained vision models has recently shown strong potential for accelerating diffusion transformer training. By aligning intermediate diffusion features with clean-image representations from self-supervised vision encoders, existing methods improve convergence and generation quality. However, such alignment also introduces a non-trivial constraint: diffusion models operate on noisy inputs whose usable information varies across timesteps, while the reference features are extracted from clean images. In this paper, we revisit this mismatch from a token-level perspective. We find that, under full-token representation alignment, tokens with large alignment-gradient norms exhibit a stable spatial preference, suggesting that the alignment objective does not affect all tokens uniformly and may encourage the model to rely on the complete set of clean-image tokens. To address this issue, we propose MaskAlign, a token-subset representation alignment method that applies alignment to randomly sampled token subsets during training. By exposing the model to different token subsets across iterations, MaskAlign reduces the dependence of representation alignment on the complete token set and encourages alignment behavior that is more stable under token-subset perturbations. To mitigate the information loss caused by directly dropping tokens, we further introduce a lightweight pre-mask token mixing block that shares information across tokens before masking.

cs.CV

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings

Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including vague user prompts, long documents, multimodal materials, and multiple heterogeneous sources. Moreover, current evaluations are often insufficiently scenario-specific. They mainly rely on generic presentation-quality criteria, such as visual appeal, layout quality, and overall coherence, but fail to assess the core capabilities required by different input settings, including grounded compression, visual-text alignment, and cross-source synthesis. Consequently, the field lacks a unified benchmark and a scenario-aware evaluation framework for faithfully diagnosing presentation-generation systems across diverse real-world settings. We present UniPPTBench, a unified benchmark for presentation generation across four representative input settings: vague-prompt, long-document, multimodal-document, and multi-source generation. We further introduce UniPPTEval, a scenario-aware evaluation protocol that combines shared metrics for cross-setting comparison with scenario-specific metrics tailored to the core requirements of each setting. We also provide transparent reference baselines to support reproducible comparison. Experiments on UniPPTBench reveal substantial performance variation across settings and recurring failure modes in content grounding, multimodal integration, and cross-source synthesis. In particular, strong performance on generic presentation-quality metrics does not necessarily imply strong task fulfillment in grounded scenarios. Together, UniPPTBench and UniPPTEval provide a faithful and diagnostic foundation for evaluating presentation generation across diverse real-world scenarios. Code and data will be publicly available.

cs.CV

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision

Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely decoupled understanding and generation components. This design, while effective for individual tasks, weakens the connection required for mutual enhancement, leaving the potential synergy empirically uncertain. We propose to explicitly restore this synergy by introducing Understanding-Oriented Post-Training (UNO), a lightweight framework that treats understanding not only as a distinct task, but also a direct supervisory signal to steer generative representations. By incorporating objectives that encode semantic abstraction (captioning) and structural details (visual regression), we enable effective gradient flow from understanding to generation. Extensive experiments on image generation and editing demonstrate that understanding can serve as an effective catalyst for generation.

cs.CV

Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions

Instruction guided image editing has advanced substantially with recent generative models, yet it still fails to produce reliable results across many seemingly simple cases. We observe that a large portion of these failures stem not from insufficient model capacity, but from poorly formulated editing tasks, such as those involving small targets, implicit spatial relations, or under-specified instructions. In this work, we frame image editing failures as a task formulation problem and propose an adaptive task reformulation framework that improves editing performance without modifying the underlying model. Our key idea is to transform the original image-instruction pair into a sequence of operations that are dynamically determined and executed by a MLLM agent through analysis, routing, reformulation, and feedback-driven refinement. Experiments on multiple benchmarks, including ImgEdit, PICA, and RePlan, across diverse editing backbones such as Qwen Image Edit and Nano Banana, show consistent improvements, with especially large gains on challenging cases. These results suggest that task reformulation is a critical but underexplored factor, and that substantial gains can be achieved by better matching editing tasks to the effective operating regime of existing models.

cs.CV

TexEditor: Structure-Preserving Text-Driven Texture Editing

Text-guided texture editing aims to modify object appearance while preserving the underlying geometric structure. However, our empirical analysis reveals that even SOTA editing models frequently struggle to maintain structural consistency during texture editing, despite the intended changes being purely appearance-related. Motivated by this observation, we jointly enhance structure preservation from both data and training perspectives, and build TexEditor, a dedicated texture editing model based on Qwen-Image-Edit-2509. Firstly, we construct TexBlender, a high-quality SFT dataset generated with Blender, which provides strong structural priors for a cold start. Sec- ondly, we introduce StructureNFT, a RL-based approach that integrates structure-preserving losses to transfer the structural priors learned during SFT to real-world scenes. Moreover, due to the limited realism and evaluation coverage of existing benchmarks, we introduce TexBench, a general-purpose real-world benchmark for text-guided texture editing. Extensive experiments on existing Blender-based texture benchmarks and our TexBench show that TexEditor consistently outperforms strong baselines such as Nano Banana Pro. In addition, we assess TexEditor on the general purpose benchmark ImgEdit to validate its generalization. Our code and data are available at https://github.com/KlingAIResearch/TexEditor.

cs.CV

A Delayed Radio Flare Traces Kinetic Energy Injection in the SMBHB Candidate SDSS~J143016.05+230344.4

SDSS~J143016.05+230344.4 ($z=0.08105$) has been proposed as a candidate pre-coalescence supermassive black hole binary and shows remarkable multiwavelength variability. Its radio evolution provides a direct probe of the compact emitting region and of the physical origin of the late-time activity. We aim to localize the variable radio emission, characterize its spectral evolution, and constrain whether the radio brightening is produced by a newly emerging compact component, external absorption, or dissipation in a structured circumnuclear environment. At all epochs, the radio emission is dominated by a single unresolved milliarcsecond core with $T_{\rm B} \gtrsim 10^{7}$ K, constraining the variable emission to $\lesssim 0.3$ pc. The broadband spectra require two synchrotron self-absorbed components: a persistent low-frequency component with $\nu_{\rm p,steady} \approx 0.74$ GHz and $S_{\rm p,steady} \approx 1.22$ mJy, and a flare component whose turnover evolves from $(6.35 {\rm GHz}, 0.18 {\rm mJy})$ in 2022 February-May to $(8.61 {\rm GHz}, 0.38 {\rm mJy})$ in 2022 December, and then to $(5.83 {\rm GHz}, 0.25 {\rm mJy})$ in 2023 March-April. The flare contribution at 15 GHz reaches $\sim 80\%$ and matches the near-epoch VLBI recovery fraction, showing that the high-frequency brightening arises from a newly formed compact synchrotron component. A second brightening of the 15.2 GHz VLBI core is detected between 2023 September and 2024 February, while the source remains unresolved. Equipartition scalings imply characteristic radii of $\sim 5 \times 10^{-4}$ pc for the flare and $\sim 9 \times 10^{-3}$ pc for the steady component, and indicate a steep inner circumnuclear density profile, $n \propto R^{-1.7}$. The delayed radio flare is best explained by dissipation in an outflow or jet-base disturbance propagating through a structured circumnuclear medium.

astro-ph.HE

Decomposition of Schwarzschild Green's Function

We present a formulation of the spherically decomposed Green's function for a Schwarzschild black hole, based on a decomposition into two components, $G^+$ and $G^-$, based on their large-frequency behaviour. While similar decompositions have been considered previously, here we systematically apply it to Schwarzschild spacetime and analyze its implications for the analytic structure of the Green's function in the complex-frequency plane. We show that both $G^+$ and $G^-$ possess branch cuts along the imaginary axis, which give rise to the direct part and the late-time tail, while the poles of $G^+$ correspond to the quasinormal mode spectrum. This allows us to identify a $\textit{branch-cut direct part}$, a quasinormal-mode contribution, and a late-time tail through contours adapted to different causal spacetime regions. This is in sharp contrast to Leaver's original formulation, where the prompt response is tied to a technically difficult large-arc contribution. We validate our decomposition with independent time-domain Regge-Wheeler simulations finding excellent agreement. Our results provide a practical and physically transparent framework for disentangling the distinct pieces of the Schwarzschild response, and offer a natural starting point for extensions to Kerr perturbations and non-linear ringdown physics.

gr-qc

Gamma-ray Emission from the S147 Region: Indication of Escaping Cosmic Rays Interacting with Molecular Clouds

We present a detailed analysis of $\gamma$-ray emission from the middle-aged supernova remnant (SNR) S147 (G180.0$-$1.7) using approximately 16.5 years of Fermi-LAT data. Spatially, a new extended $\gamma$-ray component distinct from the emission associated with the H$\alpha$ filaments of the SNR shell is identified. This new component exhibits a strong spatial correlation with dense molecular clouds (MCs) identified in CO emission at Local Standard of Rest velocities of $0$--$5\,\mathrm{km\,s^{-1}}$. Spectrally, the cloud-associated emission implies an underlying cosmic-ray (CR) proton population described by a hard power-law with an index of $\Gamma \approx 2.1$, compatible with the standard diffusive shock acceleration prediction. We interpret the $\gamma$-ray emission in this region with a hadronic scenario involving two distinct CR populations: trapped CRs reaccelerated within the radiative SNR shell as proposed in previous work, and escaping CRs illuminating the nearby MCs. The derived CR proton intensity in the MC region significantly exceeds the local Galactic background measured by AMS-02, consistent with the interpretation that the cloud is illuminated by particles accelerated by S147. These findings provide observational support for a CR-escape scenario during the earlier evolutionary phases of this middle-aged SNR and highlight the S147 MC component as a potential candidate for detection at TeV energies by LHAASO.

astro-ph.HE

SemanticALLI: Caching Reasoning, Not Just Responses, in Agentic Systems

Agentic AI pipelines suffer from a hidden inefficiency: they frequently reconstruct identical intermediate logic, such as metric normalization or chart scaffolding, even when the user's natural language phrasing is entirely novel. Conventional boundary caching fails to capture this inefficiency because it treats inference as a monolithic black box. We introduce SemanticALLI, a pipeline-aware architecture within Alli (PMG's marketing intelligence platform), designed to operationalize redundant reasoning. By decomposing generation into Analytic Intent Resolution (AIR) and Visualization Synthesis (VS), SemanticALLI elevates structured intermediate representations (IRs) to first-class, cacheable artifacts. The impact of caching within the agentic loop is substantial. In our evaluation, baseline monolithic caching caps at a 38.7% hit rate due to linguistic variance. In contrast, our structured approach allows for an additional stage, the Visualization Synthesis stage, to achieve an 83.10% hit rate, bypassing 4,023 LLM calls with a median latency of just 2.66 ms. This internal reuse reduces total token consumption, offering a practical lesson for AI system design: even when users rarely repeat themselves, the pipeline often does, at stable, structured checkpoints where caching is most reliable.

cs.AI