SearcharxivSearch

arXiv subjects

Yifan Xiao

Publications and source records attributed to Yifan Xiao.

6 recordsLinked to original sources

Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rely on implicit feature adaptation to handle weakly misaligned scenarios, while large-offset geometric discrepancies remain insufficiently addressed. In this paper, we propose a Joint Feature-domain Registration and Detection network (JFRDet), an end-to-end visible-infrared oriented object detector tailored for severely cross-modal geometric discrepancies. JFRDet introduces a Cross-Modal Affine Alignment (CMAA) module to estimate an image-level affine transformation for explicit multi-level feature alignment. Note that illumination changes directly affect the reliability of RGB cues, an Illumination-Guided Complementary Fusion (IGCF) module adaptively exploits modality reliability under varying illumination conditions for cross-modal fusion. Then, an Alignment Quality-Consistency Gating (AQCG) strategy stabilizes joint optimization by modulating detection supervision according to alignment reliability and gradient consistency. We further construct DroneVehicle Misaligned (DVMA), a benchmark for evaluating visible-infrared oriented object detection under severe cross-modal geometric misalignment. The proposed JFRDet achieves 69.7\% $\mathrm{mAP}_{50}$ on DVMA, which represents state-of-the-art (SOTA) performance. The code and dataset will be available on GitHub.

cs.CV

DrawVideo: Generating Long Video from Storyboard Keyframe Sketches

Long video generation requires high-fidelity synthesis, coherent narrative structure, and user control over extended time spans. Existing text-to-video methods often rely on a single long prompt, limiting control over pose, composition, layout, and motion. We propose DrawVideo, a sketch-guided, storyboard-driven framework for controllable long-video generation. DrawVideo decomposes long videos into independently controllable shots, each defined by a black-and-white sketch, an appearance prompt, and a motion prompt. The sketch controls pose and layout, the appearance prompt defines identity, scene, and style, and the motion prompt guides temporal dynamics. DrawVideo follows a hierarchical 'global multi-shot, local single-sketch' strategy: it first generates a structure-aligned reference keyframe, then expands the motion prompt into derivative keyframes representing action states, and finally synthesizes clips between adjacent keyframes to build each shot. We also introduce SketchLongVideo, the first dataset for sketch-guided text-to-long-video generation, constructed from animation videos via shot detection, keyframe extraction, vision-language recognition, prompt decomposition, and sketch conversion. Experiments show that DrawVideo achieves strong structural controllability, appearance consistency, visual stability, and coherent long-video generation.

cs.GR

Accelerating FRB Search: Dataset and Methods

Fast Radio Burst (FRB) is an extremely energetic cosmic phenomenon of short duration. Discovered only recently and with its origin still unknown, FRBs have already started to play a significant role in studying the distribution and evolution of matter in the universe. FRBs can only be observed through radio telescopes, which produce petabytes of data, rendering the search for FRB a challenging task. Traditional techniques are computationally expensive, time-consuming, and generally biased against weak signals. Various machine learning algorithms have been developed and employed, all of which require substantial datasets. We here introduce the FAST dataset for Fast Radio bursts EXploration (FAST-FREX), built upon the observations obtained by the Five-hundred-meter Aperture Spherical radio Telescope (FAST). Our dataset comprises 600 positive samples of observed FRB signals from three sources and 1000 negative samples of noise and Radio Frequency Interference (RFI). Furthermore, we provide a machine learning algorithm, Radio Single-Pulse Detection Algorithm Based on Visual Morphological Features (RaSPDAM), with significant improvements in efficiency and accuracy for FRB search. We also employed the benchmark comparison between conventional single-pulse search softwares, namely PRESTO and Heimdall, and RaSPDAM. RaSPDAMv2 achieves an average precision of 97% and an average recall of 83%, with notable enhancements in computational performance. Future machine learning algorithms can use this as a reference point to measure their performance and help the potential improvements. By enabling more accurate and efficient detection of transient radio events, our work facilitates the FRB and pulsars search pipeline, enhances the potential for discovering new astrophysical phenomena.

astro-ph.IM

Likely detection of GeV γ-ray emission from pulsar wind nebula G32.64+0.53 with Fermi-LAT

In this study, we report the likely GeV γ-ray emissions originating from the pulsar PSR J1849-0001's pulsar wind nebula (PWN) G32.64+0.53. Our analysis covers approximately 14.7 years of data from the Fermi Large Area Telescope (Fermi-LAT) Pass 8. The position of the source and its spectrum matches those in X-ray and TeV energy bands, so we propose that the GeV γ-ray source is indicative of PWN G32.64+0.53. We interpret the broadband spectral energy distribution (SED) using a time-dependent one-zone model, which assumes that the multi-band non-thermal emission of the target source can be generated by synchrotron radiation and inverse Compton scattering (ICS) of the electrons/positrons. Our findings demonstrate that the model substantially elucidates the observed SED. These results lend support to the hypothesis that the γ-ray source originates from the PWN G32.64+0.53 powered by PSR J1849-0001. Furthermore, the γ-rays in TeV bands are likely generated by electrons/positrons within the nebula through Inverse Compton Scattering.

astro-ph.HE

A discovery of Two Slow Pulsars with FAST: "Ronin" from the Globular Cluster M15

Globular clusters harbor numerous millisecond pulsars, but long-period pulsars ($P \gtrsim 100$ ms) are rarely found. In this study, we employed a fast folding algorithm to analyze observational data from multiple globular clusters obtained by the Five-hundred-meter Aperture Spherical radio Telescope (FAST), aiming to detect the existence of long-period pulsars. We estimated the impact of the median filtering algorithm in eliminating red noise on the minimum detectable flux density ($S_{\rm min}$) of pulsars. Subsequently, we successfully discovered two isolated long-period pulsars in M15 with periods approximately equal to 1.928451 seconds and 3.960716 seconds, respectively. On the $P-\dot{P}$ diagram, both pulsars are positioned below the spin-up line, suggesting a possible history of partial recycling in X-ray binary systems disrupted by dynamical encounters later on. According to timing results, these two pulsars exhibit remarkably strong magnetic fields. If the magnetic fields were weakened during the accretion process, then a short duration of accretion might explain the strong magnetic fields of these pulsars.

astro-ph.HE