SearcharxivSearch

arXiv subjects

Ye Li

Publications and source records attributed to Ye Li.

At least 19 recordsLinked to original sources

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.

cs.RO

RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection

Underwater unimodal object detection faces many challenges in sensor imaging, such as optical images limited by underwater noise and visible distance, and sonar images limited by less object structural information. While, optical images have rich object structural information, and sonar images are less affected by underwater noise and have a longer visible distance. Optical (RGB modality) and sonar (Sonar modality) images have complementary information underwater. In this paper, we create an RGB-Sonar multimodal object detection dataset, \textbf{R}GB-\textbf{S}onar \textbf{Fusion} (RSFusion) and propose evaluation metrics for the benchmark. And we propose the \textbf{R}GB-\textbf{S}onar \textbf{Fusion} \textbf{Det}ector (RSFusionDet) with a new RGB-Sonar multimodal object detection result expression for RGB-Sonar multimodal object detection. We analyze the features of RGB and Sonar modal information, and design a Cross-Attention Fusion (CAFusion) module to fuse RGB-Sonar spatial misalignment features and Object Matching Head (OMHead) with Loss (OMLoss) to match identical objects in RGB-Sonar modalities. Our RSFusionDet achieves 76.4/48.6 AP (RGB/Sonar) for object detection and 83.4 \(\text{F1-Score}_{match}\) for object matching, on RSFusion, which outperforms other object detection models. Compared with the DINO baseline, our method improves by 0.7/1.4 AP (RGB/Sonar) while simultaneously providing reliable cross-modal object matching. The code and datasets are publicly available at https://github.com/LEFTeyex/RSFusionDet.

cs.CV

Depth-Dominant Skeleton Detection for Natural Scenes

To date, all natural scene skeleton detection follows the paradigm of taking RGB images as the sole input; despite notable progress, methods under this paradigm suffer significant performance degradation on complex-content images. We observe that depth images are inherently insensitive to color and texture, and can provide clear regional contours and inter-region spatial relationships, which naturally alleviates the difficulty of skeleton detection in complex scenarios. Motivated by this observation, this paper proposes for the first time a novel skeleton detection paradigm where depth images serve as the dominant modality and RGB images act as the auxiliary, and accordingly presents a model DDSkel (short for Depth-Dominant Skeleton Detection) under this paradigm. DDSkel employs an asymmetric encoder design to fuse RGB information into depth features, with the RGB modality branch having only 12% the parameters of the depth modality branch. DDSkel has a simple structure without intricate designs. Nevertheless, with only 36% of the trainable parameters of the current best method, DDSkel outperforms all state-of-the-art approaches on SymPASCAL, the most challenging dataset with a large volume of complex images.

cs.CV

Machine Learning and ARIMA Model Averaging for Adaptive Public Health Forecasting: Comparative Evaluation and an Ontario COVID-19 Case Study

Public health forecasts must respond to abrupt changes in surveillance data without over-extrapolating noise, reporting artifacts, or temporary trends. We evaluated autoregressive integrated moving average (ARIMA), random forest, and extreme gradient boosting (XGBoost) models using 190 weekly observations of publicly available Ontario COVID-19 case counts from January 2020 to October 2023. Rolling-origin time-series cross-validation preserved temporal order during model tuning and evaluation. Performance was assessed across three operating dimensions: responsiveness following selected turning points, forecast horizons of one to six weeks, and the amount of historical training data. We also developed Machine Learning and ARIMA Model Averaging (MLAMA), a non-negative performance-weighted ensemble with weights that vary by forecast horizon and responsiveness setting. Retrospective comparisons showed that ARIMA adapted rapidly after turning points but its normalized error increased at longer horizons. Random forest and XGBoost were less responsive initially but maintained more stable normalized error over longer horizons. For two-week forecasts at the end of the study period, training on the most recent data outperformed using longer historical periods, particularly for XGBoost. MLAMA achieved the lowest normalized mean absolute percentage error across most forecast horizons and ranked among the best-performing methods across responsiveness settings. These findings support selecting forecasting models according to operating conditions rather than relying on a single universally preferred approach. MLAMA provides a practical framework for combining complementary statistical and machine-learning forecasts. The accompanying Python package is currently maintained in a private repository while software validation and reproducibility testing are completed.

cs.LG

A Roadmap for Transient Hunters: Mapping Stellar Mass and Star Formation Rate Anisotropies in the Local Universe

Over the past few decades, an increasing number of transients in nearby galaxies have been discovered through various survey projects. Unlike astrophysical phenomena at cosmological distances, transients in the local universe exhibit a pronounced anisotropy in their sky distribution. Consequently, adopting an appropriate survey strategy is essential to improve the efficiency of transient searches in the local universe. In this work, we utilized a large galaxy catalog to map the sky distributions of stellar mass and star formation rate (SFR) across different luminosity distance thresholds and angular resolutions of the grid on the celestial sphere. These maps can further serve to characterize the anisotropic spatial distribution of nearby extragalactic transients. For different angular resolutions of the celestial sphere, we find that the sky distributions of stellar mass of galaxies are similar to those of the SFR in the main anisotropic structures. As the luminosity distance threshold increases, the anisotropic structures of the sky distributions become more isotropic. We calculate the angular power spectra and fluctuations of the sky distribution of stellar mass and SFR at a given angular resolution and find that the angular power spectra and fluctuations decrease rapidly as the luminosity distance threshold increases. Finally, by qualitatively comparing the sky distribution of core-collapse supernovae with our SFR sky distribution, we find that the two exhibit consistent patterns in several prominent structures. The mapped sky distributions of stellar mass and SFR can serve as valuable references for future surveys in searching for extragalactic transients.

astro-ph.HE

Cross-validation of six dispersion measure estimation methods for FRB 20240114A

Fast Radio Bursts (FRBs) are important cosmological probes, but their applications depend critically on accurate dispersion measure (DM) determinations. We present a systematic comparison of six DM estimation methods using 2,874 bursts from FRB20240114A, the most active repeating FRB currently known, observed by FAST during a single 4.4-hr session on 2024 March 12. This large, homogeneous sample over a short timescale, during which the propagation environment is expected to be nearly static, provides an ideal benchmark for isolating algorithmic effects on DM determination. We investigate the dependence of inter-method consistency on signal-to-noise ratio (S/N), burst morphology, and radio frequency interference (RFI). Low-S/N bursts exhibit significantly larger inter-method deviations, while single-component bursts produce highly consistent DM values across methods. In contrast, complex double- and multiple-component bursts with drifting substructures lead to substantial inter-method scattering, indicating that DM discrepancies are primarily driven by algorithmic responses to burst morphology. RFI does not significantly alter the global statistical behavior of DM deviations, but it affects density-filtering methods through morphology distortion caused by frequency-channel masking. Even after imposing strict inter-method consistency constraints, FRB20240114A still exhibits notable apparent DM fluctuations spanning $\sim$528-534~pc~cm$^{-3}$ over 15,780s. For morphologically simple bursts these variations far exceed the measurement uncertainty and, on second-to-minute timescales, cannot arise from any plausible change in the line-of-sight electron column, pointing instead to a frequency-dependent emission-time structure intrinsic to the bursts that mimics dispersion.

astro-ph.HE

No Strong Evidence for Plasma Lensing in FRB 20240114A

FRB~20240114A is an extremely active repeating fast radio burst for which plasma lensing has been proposed to explain its burst-rate variations, spectral evolution, and apparently ``carbon-copy'' burst pairs. Using FAST data and publicly available Parkes observations, we test this interpretation with a one-dimensional Gaussian plasma-lens model. Although the burst-rate enhancements can be fitted separately, the corresponding magnification peaks and demagnification troughs are offset by far more than predicted and show no consistent periodicity. Moreover, with more than 10,000 bursts detected, a few apparently ``carbon-copy'' pairs can readily occur by chance. The burst bandwidth is not systematically narrower during the proposed lensing interval, nor are the burst energies significantly enhanced during the predicted magnification interval. These results provide no compelling evidence that a single Gaussian plasma lens explains the observed variability, which is more likely dominated by intrinsic source activity.

astro-ph.HE

A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial

Rare diseases affect millions of individuals worldwide, yet timely diagnosis remains a major public health challenge due to scarcity of specialized clinical expertise. While large language models (LLMs) show promise to support rare disease diagnosis, current models are constrained by insufficient clinical deployability, limited clinically grounded evidence, and scarcity of training data. Here we present RaDaR (Rare Disease navigatoR), an open-source, compact reasoning LLM (32B parameters) for rare disease diagnosis. RaDaR was trained with 49,170 publicly available free-text cases and 104,666 synthetic cases with reasoning-enhanced training. RaDaR showed the strongest performance among evaluated open-source models, including the 671B DeepSeek-R1, across public benchmarks and four external validation centers. In a retrospective cohort, RaDaR prioritized the final diagnosis before documented clinical suspicion in 61.06 percent of cases, corresponding to a potential lead time of 1.87 months and 50.18 percent of the within-center interval. In a randomized physician-assistance trial, RaDaR assistance improved physicians' rare-disease diagnostic accuracy by 21.44 percentage points compared with internet search alone. Synthetic-data ablations suggested that phenotype-anchored narratives provide useful training signal for long-tail rare diseases, with a monotonic scaling trend within the tested data range. Together, RaDaR and its development and validation framework provide a deployable rare-disease reasoning model and a reproducible development framework for diagnostic AI under data scarcity.

cs.AI

GRB 250424A: A Case Study of Energy Injection with Multiwavelength Observations

We present a comprehensive multiwavelength analysis of the long-duration gamma-ray burst (GRB) 250424A. Our dataset spans from the prompt gamma-ray emission to late-time optical monitoring, including spectra obtained with the Keck 10\,m telescope. We find that the afterglow light curves display a prominent, simultaneous shallow decay phase in both X-ray and optical bands, followed by an achromatic transition to a standard decay regime. The broadband spectral energy distributions are well-modeled by a single power-law function, indicating a common synchrotron origin for the emission across frequencies. We interpret the afterglow evolution within the framework of a relativistic forward shock refreshed by continuous energy injection. This scenario successfully reproduces the observed temporal and spectral behavior, yielding an isotropic equivalent kinetic energy of $E_{\rm K,iso} \approx 5.5 \times 10^{52}$ erg and an injection index of $q\approx 0.34$ in a constant-density circumburst environment. The shallow decay phase is consistent with sustained energy injection lasting $\sim$ 9 ks. Despite the relatively low redshift, late-time optical observations reveal no distinct supernova component; however, our derived upper limits do not strictly rule out the presence of a typical GRB-associated supernova.

astro-ph.HE

Multiwavelength Analysis of the Einstein Probe X-ray Transient EP240305a

We report multiwavelength observations of EP240305a, an uncatalogued X-ray transient detected by the Einstein Probe on March 5, 2024. The source exhibits distinct characteristics across the X-ray, optical, near-infrared, and radio bands. The soft X-ray observations show two significant flares lasting ~100-250 s, accompanied by rapid flux decay in a few days, and the optical and near-infrared data reveal a faint, candidate counterpart. In contrast, the radio observations expose a long-term spectral evolution from a self-absorbed to an optically thin state within two months, implying discrete jet ejection. We compare EP240305a with known classes of X-ray transients and find that it is unlikely to be associated with long-timescale transients such as jetted tidal disruption events or X-ray binaries. Its properties also disfavor a short-timescale stellar flare origin. Although the absence of optical spectroscopy prevents a redshift determination, the source exhibits properties similar to those of gamma-ray-dark gamma-ray burst-like transients, which may be associated with relativistic jets viewed off-axis or with choked jets. The discovery of EP240305a, along with other uncataloged transients detected by the Einstein Probe, underscores the scientific potential of highly sensitive X-ray survey telescopes and rapid-response multiwavelength follow-up observations in exploring the nature of atypical astronomical transients.

astro-ph.HE

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time robotic manipulation, especially when large vision-language backbones and iterative action heads run at every control step. Existing VLA acceleration methods often optimize individual components or rely on fixed acceleration rules, treating different control steps with largely fixed computation and overlooking the non-uniform reasoning demands of sequential embodied control. Inspired by human motor control, where cognitive and feedback resources concentrate on goal-sensitive stages, we argue that VLA models should learn when to invest full computation and when to reuse prior computation. We propose ElegantVLA, a plug-in phase-adaptive inference framework that accelerates VLA models through intra-model dynamic compute scheduling. ElegantVLA introduces a lightweight scheduler that observes temporal representation similarity, robot-motion cues, and episode progress to jointly allocate computation across the vision encoder, LLM, and action head. For perception-language reasoning, the scheduler selects a five-level Vision-LLM compute mode, from full recomputation to multi-step temporal reuse, based on visual-language representation stability. For action generation, it selects a three-level denoising mode, reusing intermediate denoising states during stable motion while preserving full refinement for goal-sensitive stages. By coordinating these decisions, ElegantVLA offers a general acceleration framework for modern VLA pipelines with explicit action-generation modules, without modifying or retraining the base model. Experiments on GR00T and CogACT achieve up to 2.55x and 3.77x speedup, and on six real-world GR00T tasks ElegantVLA cuts computation by 2.18x while raising control frequency from 13.8 Hz to 26.3 Hz.

cs.RO

ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation

Minute-scale cinematic video generation is a central challenge for generative video models. Existing paradigms address only fragments of this challenge: single-shot extrapolation preserves an anchor but lacks cinematic structure, while multi-shot storytelling imposes structure yet remains free to invent its visual states rather than continue an observed one. We define Multi-Shot Video Extrapolation (MSVE), a task that extends an observed frame or clip into a sequence of cinematically structured shots while preserving anchor state and advancing narrative intent. This setting operates under the finite per-call generation budget of short-video models. We identify three coupled bottlenecks: (1) global planners over-specify unsupported details from full screenplays; (2) shot-level prompts dilute task-relevant state when carrying the complete story; and (3) temporal chaining turns generated frames into a lossy memory in which identity, scene, object, and action state decay. MSVE reveals that long-video failure is not merely a limitation of context length, but a failure of context allocation. We propose Recursive Context Allocation (ReCA), an inference-time framework that allocates context hierarchically across planning and generation. ReCA recursively decomposes MSVE into context-bounded subproblems, invokes frozen generators at leaf nodes, and propagates structured state updates across time. To evaluate this setting, we further propose MSVE-Bench and NB-Q, a source-grounded protocol with prompts purpose-built for 3 to 5 minute long-video generation, a regime not addressed by existing short-clip benchmarks. Compared to previous methods, ReCA improves average normalized score by 8 to 16 percent over the strongest competing controller and improves multi-shot consistency metrics by 28 to 43 percent. View the project page at https://reca.vmv.re.

cs.CV

GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation

We introduce GE-Sim 2.0 (Genie Envisioner World Simulator 2.0), a closed-loop video world simulator for robotic manipulation. Building on the action-conditioned video generation framework of Genie Envisioner, GE-Sim 2.0 is re-trained on thousands of hours of real-world robot data spanning teleoperation, contact-rich interaction, and on-robot policy deployment, substantially improving action-following fidelity and trajectory coverage. On top of this foundation, three new modules close the loop from video simulation to policy learning: a state expert that decodes proprioceptive state from video latents to support next-chunk prediction by downstream VLA policies; a world judge that scores generated rollouts against task instructions, yielding machine-verifiable success signals and rewards in place of manual inspection; and an acceleration framework that delivers a 25-frame rollout in 2.3 seconds on a single H100, with up to 4* frame skipping at inference for long-horizon evaluation. GE-Sim 2.0 tops the public WorldArena leaderboard at only 2B parameters, outperforming both dedicated robotic world models and closed-source general video generators, and policies trained against its rollouts and rewards translate into measurable real-world gains, establishing GE-Sim 2.0 as a practical platform for scalable evaluation and closed-loop learning of manipulation policies.

cs.RO

Test-time Sparsity for Extreme Fast Action Diffusion

Action diffusion excels at high-fidelity action generation but incurs heavy computational costs owing to its iterative denoising nature. Despite current technologies showing promise in accelerating diffusion transformers by reusing the cached features, they struggle to adapt to policy dynamics arising from diverse perceptions and multi-round rollout iterations in open environments. We propose test-time sparsity to tackle this challenge, which aims to accelerate action diffusion by dynamically predicting prunable residual computations for each model forward at test time. However, two bottlenecks remain in this paradigm: 1) repetitive conditional encoding and pruning offset most potential speed gains, and 2) the features cached from previous denoising timesteps cannot constrain large pruning errors under aggressive sparsity. To address the first bottleneck, we design a highly parallelized inference pipeline that minimizes the non-decoder delay to milliseconds. Specifically, we first design a lightweight pruner that shares the encoder with the diffusion transformer. Then, we decouple the encoding and pruning from the autoregressive denoising loop by processing all denoising timesteps in parallel, and overlap the pruner with the decoder forward inference through asynchronism. To overcome the second bottleneck, we introduce an omnidirectional reusing strategy, which achieves 95% sparsity by selectively reusing features cached from the current forward, previous denoising timesteps, and earlier rollout iterations. To learn the rollout-level reusing strategies, we sample a few action trajectories to supervise the sparsified diffusion step by step. Extensive experiments demonstrate that our method reduces FLOPs by 92% and accelerates action generation by 5x, achieving lossless performance with an inference frequency of 47.5 Hz. Our code is available at https://github.com/ky-ji/Test-time-Sparsity.

cs.CV

Unlocking Optical Prior: Spectrum-Guided Knowledge Transfer for SAR Generalized Category Discovery

Generalized Category Discovery (GCD) holds significant promise for the label-scarce Synthetic Aperture Radar (SAR) domain, yet its efficacy is severely constrained by the cross-modal incompatibility between the inherent optical prior of the Large Vision Models (LVMs) and SAR imagery. Existing domain adaptation methods often lack an inductive bias that reflects imaging characteristics, consequently failing to effectively transfer optical prior into the SAR domain. To address this issue, the Modal Discrepancy Curve (MDC) is introduced to model cross-modal discrepancy as a structured frequency-domain descriptor derived from spectral energy distributions. Leveraging this formulation, we propose the MDC-guided Cross-modal Prior Transfer (MCPT) framework, a pre-training paradigm that operates on paired optical-SAR data. Within this framework, Adaptive Frequency Tokenization (AFT) converts the MDC into learnable tokens, and Frequency-aware Expert Refinement (FER) performs band-wise discrepancy-aware feature refinement using these tokens. Based on the refined representations, contrastive learning aligns refined embeddings across modalities and internalizes the adaptation pattern. Ultimately, the superior SAR feature representation capability learned during paired pre-training is applied to downstream single-modal SAR-GCD tasks. Extensive experiments demonstrate state-of-the-art performance across multiple mainstream datasets, indicating that frequency-domain discrepancy modeling enables more effective adaptation of optical prior to SAR imagery.

cs.CV

A Search for Rotation Measure Flare Candidates in Repeating Fast Radio Bursts

Fast radio bursts (FRBs) are millisecond-duration extragalactic radio transients of unknown origin. Rotation measures (RMs) probe their local magneto-ionic environments and provide important clues to their nature. While RM variability has been observed in several repeating FRBs, it is typically gradual or stochastic. Recently, observations of FRB~20220529 revealed an abrupt RM excursion followed by rapid recovery on week-long timescales, termed an ``RM flare'', suggesting a potentially distinct form of RM variability associated with localized magnetized plasma. In this work, we perform a systematic search for RM flare candidates in repeating FRBs with multi-epoch RM measurements. Using a $3\sigma$ significance threshold, we identify two candidates with multiple observational epochs (FRB~20121102A and FRB~20201124A) and two additional single-epoch candidates (FRB~20180916B), in addition to the event in FRB~20220529A. Our results suggest that RM flares, if confirmed, may not be rare among repeating FRBs and point to highly dynamic magnetized environments local to the sources. Future high-cadence polarimetric observations, particularly following the discovery of RM excursions, will be essential for confirming these candidates and constraining their physical origin.

astro-ph.HE

Bias-constrained multimodal intelligence for equitable and reliable clinical AI

The integration of medical imaging and clinical text has enabled the emergence of generalist artificial intelligence (AI) systems for healthcare. However, pervasive biases, such as imbalanced disease prevalence, skewed anatomical region distributions, heterogeneous imaging protocols, and demographic disparities, pose significant challenges to the fairness and reliability of vision-language systems in real-world clinical settings. Here we present BiasCareVL, a bias-aware multimodal learning framework that introduces bias control directly into model design, rather than treating it as a post hoc correction. BiasCareVL incorporates adaptive uncertainty modeling with optional human-in-the-loop refinement to regulate the influence of dominant data patterns and to promote equitable reasoning under distributional imbalance. Trained on 3.44 million samples spanning over 15 imaging modalities, the framework supports diverse clinical tasks, including visual question answering, disease classification, segmentation, and report generation within a unified representation space. Across eight public benchmarks covering dermatology, oncology, radiology, and pathology, BiasCareVL consistently outperforms 20 state-of-the-art methods, with pronounced gains in clinically challenging scenarios, including over 10% accuracy improvement in multi-class skin lesion diagnosis and more than 20% Dice improvement in small tumor segmentation. Furthermore, BiasCareVL achieves diagnostic performance exceeding human accuracy with substantially reduced time requirements when evaluated with board-certified radiologists. By open-sourcing BiasCareVL, we aim to promote a transparent, reproducible, and equitable future for AI in healthcare, paving the way for general-purpose, trustworthy, and clinically reliable AI systems.

cs.CV

A fast X-ray transient with chromatic flares: signatures of violent collisions induced by late-time central engine reactivation

Extragalactic Fast X-ray Transients (EFXTs) represent an emerging class of high-energy phenomena characterized by X-ray outbursts lasting from tens to hundreds of seconds. However, for more than half of the EFXTs, their physical origins remain elusive. In this Letter, we report the discovery of EP250302a, a luminous EFXT detected by the Einstein Probe (EP) at a redshift of $z = 1.131$. The multi-wavelength light curves of EP250302a reveal remarkable temporal features that distinguish it from the previously known EP-detected EFXT population, most notably a needle-like X-ray flare accompanied by smooth optical rebrightening during the afterglow phase. We suggest that the distinct X-ray and optical behaviors constitute the first observed instance of late-time violent collision of two relativistic shells in an EFXT. Drawing on insights from GRB studies, such a collision process strongly indicates the reactivation of a central engine, making EP250302a-like transients a unique laboratory for probing the late-time activity and jet physics of EFXT central engines.

astro-ph.HE