Searcharxiv⌕ Search

arXiv subjects

Jun Shi

Publications and source records attributed to Jun Shi.

At least 37 records · Page 2Linked to original sources

Searches for Prompt Low-Frequency Radio Counterparts to Gravitational Wave Event S250206dm with the OVRO-LWA Time Machine

We report on a search for prompt, low-frequency radio emission from the gravitational-wave (GW) merger S250206dm using the Owens Valley Radio Observatory Long Wavelength Array (OVRO-LWA). Early alerts favored a neutron-star-containing merger, making this a compelling target. Motivated by theoretical predictions of coherent radio bursts from mergers involving a neutron star, we utilized the OVRO-LWA Time Machine system to analyze voltage data recorded around the time of the event. The Time Machine is a two-stage voltage buffer and processing pipeline that continuously buffers raw data from all antennas across the array's nearly full-hemisphere instantaneous field of view, enabling retrospective beamforming, dedispersion, and fast-transient candidate identification. For this event, we analyzed a 30-minute interval beginning 3.5 minutes after the merger, which included two minutes of pre-alert data recovered by the ring buffer. We searched the 50% localization probability region with millisecond time resolution in the 69-86 MHz frequency band. No radio counterpart was detected above a 7-sigma fluence detection threshold of ~150 Jy ms. Using Bayesian analysis, we place a 95% confidence upper limit on the source luminosity of L95 = 4 x 10^41 erg s^-1. These constraints start to probe the bright end of the coherent-emission parameter space predicted by jet-ISM shock processes, magnetar and blitzar-like mechanisms, and recent simulation-based scenarios for neutron-star-containing mergers. This study presents the first sensitive, large-area, millisecond-timescale search for prompt low-frequency radio emission from a GW merger with the OVRO-LWA, establishing a framework in which about ten additional events will yield stringent population-level constraints.

astro-ph.HE↗

Traffic Image Restoration under Adverse Weather via Frequency-Aware Mamba

Traffic image restoration under adverse weather conditions remains a critical challenge for intelligent transportation systems. Existing methods primarily focus on spatial-domain modeling but neglect frequency-domain priors. Although the emerging Mamba architecture excels at long-range dependency modeling through patch-wise correlation analysis, its potential for frequency-domain feature extraction remains unexplored. To address this, we propose Frequency-Aware Mamba (FAMamba), a novel framework that integrates frequency guidance with sequence modeling for efficient image restoration. Our architecture consists of two key components: (1) a Dual-Branch Feature Extraction Block (DFEB) that enhances local-global interaction via bidirectional 2D frequency-adaptive scanning, dynamically adjusting traversal paths based on sub-band texture distributions; and (2) a Prior-Guided Block (PGB) that refines texture details through wavelet-based high-frequency residual learning, enabling high-quality image reconstruction with precise details. Meanwhile, we design a novel Adaptive Frequency Scanning Mechanism (AFSM) for the Mamba architecture, which enables the Mamba to achieve frequency-domain scanning across distinct subgraphs, thereby fully leveraging the texture distribution characteristics inherent in subgraph structures. Extensive experiments demonstrate the efficiency and effectiveness of FAMamba.

cs.CV↗

Huizhou Hadron Spectrometer -- a Proposed High-rate Experimental Setup at the High Intensity Heavy-ion Accelerator Facility

The High-Intensity Heavy-Ion Accelerator Facility (HIAF), currently under construction in Huizhou, Guangdong Province, China, is projected to be completed by 2025. This facility will be capable of producing proton and heavy-ion beams with energies reaching several GeV, thereby offering a versatile platform for advanced fundamental physics research. Key scientific objectives include exploring physics beyond the Standard Model through the search for novel particles and interactions, testing fundamental symmetries, investigating exotic hadronic states such as di-baryons, pentaquark states and multi-strange hypernuclei, conducting precise measurements of hadron and hypernucleus properties, and probing the phase boundary and critical point of nuclear matter. To facilitate these investigations, we propose the development of a dedicated experimental apparatus at HIAF - the Huizhou Hadron Spectrometer (HHaS). This paper presents the conceptual design of HHaS, comprising a solenoid magnet, a five-dimensional silicon pixel tracker, a Low-Gain Avalanche Detector (LGAD) for time-of-flight measurements, and a Cherenkov-scintillation dual-readout electromagnetic calorimeter. The design anticipates an unprecedented event rate of 1-100 MHz, extensive particle acceptance, a track momentum resolution at 1% level, an electromagnetic energy resolution of ~3% @ 1 GeV and multi-particle identification capabilities. Such capabilities position HHaS as a powerful instrument for advancing experimental studies in particle and nuclear physics. The successful realization of HHaS is expected to significantly bolster the development of medium- and high-energy physics research within China.

hep-ex↗

Large Language Model Aided Birt-Hogg-Dube Syndrome Diagnosis with Multimodal Retrieval-Augmented Generation

Deep learning methods face dual challenges of limited clinical samples and low inter-class differentiation among Diffuse Cystic Lung Diseases (DCLDs) in advancing Birt-Hogg-Dube syndrome (BHD) diagnosis via Computed Tomography (CT) imaging. While Multimodal Large Language Models (MLLMs) demonstrate diagnostic potential fo such rare diseases, the absence of domain-specific knowledge and referable radiological features intensify hallucination risks. To address this problem, we propose BHD-RAG, a multimodal retrieval-augmented generation framework that integrates DCLD-specific expertise and clinical precedents with MLLMs to improve BHD diagnostic accuracy. BHDRAG employs: (1) a specialized agent generating imaging manifestation descriptions of CT images to construct a multimodal corpus of DCLDs cases. (2) a cosine similarity-based retriever pinpointing relevant imagedescription pairs for query images, and (3) an MLLM synthesizing retrieved evidence with imaging data for diagnosis. BHD-RAG is validated on the dataset involving four types of DCLDs, achieving superior accuracy and generating evidence-based descriptions closely aligned with expert insights.

cs.CV↗

What Foundation Models can Bring for Robot Learning in Manipulation : A Survey

The realization of universal robots is an ultimate goal of researchers. However, a key hurdle in achieving this goal lies in the robots' ability to manipulate objects in their unstructured surrounding environments according to different tasks. The learning-based approach is considered an effective way to address generalization. The impressive performance of foundation models in the fields of computer vision and natural language suggests the potential of embedding foundation models into manipulation tasks as a viable path toward achieving general manipulation capability. However, we believe achieving general manipulation capability requires an overarching framework akin to auto driving. This framework should encompass multiple functional modules, with different foundation models assuming distinct roles in facilitating general manipulation capability. This survey focuses on the contributions of foundation models to robot learning for manipulation. We propose a comprehensive framework and detail how foundation models can address challenges in each module of the framework. What's more, we examine current approaches, outline challenges, suggest future research directions, and identify potential risks associated with integrating foundation models into this domain.

cs.RO↗

EndoCIL: A Class-Incremental Learning Framework for Endoscopic Image Classification

Class-incremental learning (CIL) for endoscopic image analysis is crucial for real-world clinical applications, where diagnostic models should continuously adapt to evolving clinical data while retaining performance on previously learned ones. However, existing replay-based CIL methods fail to effectively mitigate catastrophic forgetting due to severe domain discrepancies and class imbalance inherent in endoscopic imaging. To tackle these challenges, we propose EndoCIL, a novel and unified CIL framework specifically tailored for endoscopic image diagnosis. EndoCIL incorporates three key components: Maximum Mean Discrepancy Based Replay (MDBR), employing a distribution-aligned greedy strategy to select diverse and representative exemplars, Prior Regularized Class Balanced Loss (PRCBL), designed to alleviate both inter-phase and intra-phase class imbalance by integrating prior class distributions and balance weights into the loss function, and Calibration of Fully-Connected Gradients (CFG), which adjusts the classifier gradients to mitigate bias toward new classes. Extensive experiments conducted on four public endoscopic datasets demonstrate that EndoCIL generally outperforms state-of-the-art CIL methods across varying buffer sizes and evaluation metrics. The proposed framework effectively balances stability and plasticity in lifelong endoscopic diagnosis, showing promising potential for clinical scalability and deployment.

cs.CV↗

Enigmatic centi-SFU and mSFU nonthermal radio transients detected in the middle corona

Decades of solar coronal observations have provided substantial evidence for accelerated particles in the corona. In most cases, the location of particle acceleration can be roughly identified by combining high spatial and temporal resolution data from multiple instruments across a broad frequency range. In almost all cases, these nonthermal particles are associated with quiescent active regions, flares, and coronal mass ejections (CMEs). Only recently, some evidence of the existence of nonthermal electrons at locations outside these well-accepted regions has been found. Here, we report for the first time multiple cases of transient nonthermal emissions, in the heliocentric range of $\sim 3-7R_\odot$, which do not have any obvious counterparts in other wavebands, like white-light and extreme ultra-violet. These detections were made possible by the regular availability of high dynamic range low-frequency radio images from the Owens Valley Radio Observatory's Long Wavelength Array. While earlier detections of nonthermal emissions at these high heliocentric distances often had comparable extensions in the plane-of-sky, they were primarily been associated with radio CMEs, unlike the cases reported here. Thus, these results add on to the evidence that the middle corona is extremely dynamic and contains a population of nonthermal electrons, which is only becoming visible with high dynamic range low-frequency radio images.

astro-ph.SR↗

Possible First Detection of Gyroresonance Emission from a Coronal Mass Ejection in the Middle Corona

Routine measurements of the magnetic field of coronal mass ejections (CMEs) have been a key challenge in solar physics. Making such measurements is important both from a space weather perspective and for understanding the detailed evolution of the CME. In spite of significant efforts and multiple proposed methods, achieving this goal has not been possible to date. Here we report the first possible detection of gyroresonance emission from a CME. Assuming that the emission is happening at the third harmonic, we estimate that the magnetic field strength ranges from 7.9--5.6 G between 4.9-7.5 $R_\odot$. We also demonstrate that this high magnetic field is not the average magnetic field inside the CME, but most probably is related to small magnetic islands, which are also being observed more frequently with the availability of high-resolution and high-quality white-light images.

astro-ph.SR↗

FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers

Multi-Modal Diffusion Transformers (DiTs) demonstrate exceptional capabilities in visual synthesis, yet their deployment remains constrained by substantial computational demands. To alleviate this bottleneck, many sparsity-based acceleration methods have been proposed. However, their diverse sparsity patterns often require customized kernels for high-performance inference, limiting universality. We propose FlashOmni, a unified sparse attention engine compatible with arbitrary DiT architectures. FlashOmni introduces flexible sparse symbols to standardize the representation of a wide range of sparsity strategies, such as feature caching and block-sparse skipping. This unified abstraction enables the execution of diverse sparse computations within a single attention kernel. In addition, FlashOmni designs optimized sparse GEMMs for attention blocks, leveraging sparse symbols to eliminate redundant computations and further improve efficiency. Experiments demonstrate that FlashOmni delivers near-linear, closely matching the sparsity ratio speedup (1:1) in attention and GEMM-$Q$, and achieves 2.5$\times$-3.8$\times$ acceleration in GEMM-$O$ (max peaking at about 87.5% of the theoretical limit). Applied with a multi-granularity sparsity strategy, it enables the Hunyuan model (33K) to achieve about 1.5$\times$ end-to-end acceleration without degrading visual quality.

cs.LG↗

Probing the Turbulent Corona and Heliosphere Using Radio Spectral Imaging Observation during the Solar Conjunction of Crab Nebula

Measuring plasma parameters in the upper solar corona and inner heliosphere is challenging because of the region's weakly emissive nature and inaccessibility for most in situ observations. Radio imaging of broadened and distorted background astronomical radio sources during solar conjunction can provide unique constraints for the coronal material along the line of sight. In this study, we present radio spectral imaging observations of the Crab Nebula (Tau A) from June 9 to June 22, 2024 when it was near the Sun with a projected heliocentric distance of 5 to 27 solar radii, using the Owens Valley Radio Observatory's Long Wavelength Array (OVRO-LWA) at multiple frequencies in the 30--80 MHz range. The imaging data reveal frequency-dependent broadening and distortion effects caused by anisotropic wave propagation through the turbulent solar corona at different distances. We analyze the brightness, size, and anisotropy of the broadened images. Our results provide detailed observations showing that the eccentricity of the unresolved source increases as the line of sight approaches the Sun, suggesting a higher anisotropic ratio of the plasma turbulence closer to the Sun. In addition, the major axis of the elongated source is consistently oriented in the direction perpendicular to the radial direction, suggesting that the turbulence-induced scattering effect is more pronounced in the direction transverse to the coronal magnetic field. Lastly, when the source undergoes large-scale refraction as the line of sight passes through a streamer, the apparent source exhibits substructures at lower frequencies. This study demonstrates that observations of celestial radio sources with lines of sight near the Sun provide a promising method for measuring turbulence parameters in the inner heliosphere.

astro-ph.SR↗

Measuring the Magnetic Field of a Coronal Mass Ejection from Low to Middle Corona

A major challenge in understanding the initiation and evolution of coronal mass ejections (CMEs) is measuring the magnetic field of the magnetic flux ropes (MFRs) that drive CMEs. Recent developments in radio imaging spectroscopy have paved the way for diagnosing the CMEs' magnetic field using gyrosynchrotron radiation. We present magnetic field measurements of a CME associated with an X5-class flare by combining radio imaging spectroscopy data in microwaves (1--18 GHz) and meter-wave (20--88 MHz), obtained by the Owens Valley Radio Observatory's Expanded Owens Valley Solar Array (EOVSA) and Long Wavelength Array (OVRO-LWA), respectively. EOVSA observations reveal that the microwave source, observed in the low corona during the initiation phase of the eruption, outlines the bottom of the rising MFR-hosting CME bubble seen in extreme ultraviolet and expands as the bubble evolves. As the MFR erupts into the middle corona and appears as a white light CME, its meter-wave counterpart, observed by OVRO-LWA, displays a similar morphology. For the first time, using gyrosynchrotron spectral diagnostics, we obtain magnetic field measurements of the erupting MFR in both the low and middle corona, corresponding to coronal heights of 0.02 and 1.83 $R_{\odot}$. The magnetic field strength is found to be around 300 G at 0.02 $R_{\odot}$ during the CME initiation, and about 0.6 G near the leading edge of the CME when it propagates to 1.83 $R_{\odot}$. These results provide critical new insights into the magnetic structure of the CME and its evolution during the early stages of its eruption.

astro-ph.SR↗

Analysis of $Σ^*$ via isospin selective reaction $K_Lp \to π^+Σ^0$

The isospin-selective reaction $K_Lp \to π^+Σ^0$ provides a clean probe for investigating $I=1$ $Σ^*$ resonances. In this work, we perform an analysis of this reaction using an effective Lagrangian approach for the first time, incorporating the well-established $Σ(1189) 1/2^+$, $Σ(1385) 3/2^+$, $Σ(1670) 3/2^-$, $Σ(1775) 5/2^-$ states, while also exploring contributions from other unestablished states. By fitting the available differential cross section and recoil polarization data, adhering to partial-wave phase conventions same as PDG, we find that besides the established resonances, contributions from $Σ(1660) 1/2^+$, $Σ(1580) 3/2^-$ and a $Σ^*(1/2^-)$ improve the description. Notably, a $Σ^*(1/2^-)$ resonance with mass around 1.54 GeV, consistent with $Σ(1620)1/2^-$, is found to be essential for describing the data in this channel, a stronger indication than found in previous analyses focusing on $πΛ$ final states. While providing complementary support for $Σ(1660) 1/2^+$ and $Σ(1580) 3/2^-$, our results highlight the importance of the $Σ(1620) 1/2^-$ region in $K_Lp \to π^+Σ^0$. Future high-precision measurements are needed to solidify these findings and further constrain the $Σ^*$ spectrum.

hep-ph↗

DIFFUMA: High-Fidelity Spatio-Temporal Video Prediction via Dual-Path Mamba and Diffusion Enhancement

Spatio-temporal video prediction plays a pivotal role in critical domains, ranging from weather forecasting to industrial automation. However, in high-precision industrial scenarios such as semiconductor manufacturing, the absence of specialized benchmark datasets severely hampers research on modeling and predicting complex processes. To address this challenge, we make a twofold contribution.First, we construct and release the Chip Dicing Lane Dataset (CHDL), the first public temporal image dataset dedicated to the semiconductor wafer dicing process. Captured via an industrial-grade vision system, CHDL provides a much-needed and challenging benchmark for high-fidelity process modeling, defect detection, and digital twin development.Second, we propose DIFFUMA, an innovative dual-path prediction architecture specifically designed for such fine-grained dynamics. The model captures global long-range temporal context through a parallel Mamba module, while simultaneously leveraging a diffusion module, guided by temporal features, to restore and enhance fine-grained spatial details, effectively combating feature degradation. Experiments demonstrate that on our CHDL benchmark, DIFFUMA significantly outperforms existing methods, reducing the Mean Squared Error (MSE) by 39% and improving the Structural Similarity (SSIM) from 0.926 to a near-perfect 0.988. This superior performance also generalizes to natural phenomena datasets. Our work not only delivers a new state-of-the-art (SOTA) model but, more importantly, provides the community with an invaluable data resource to drive future research in industrial AI.

cs.CV↗

Study of nucleon and $Δ$ resonances from a systematic analysis of $K^\astΣ$ photoproduction

A systematic analysis of the $K^\astΣ$ photoproduction off proton is performed with all the available differential cross section data. We carry out a strategy different from the previous studies of these reactions, where, instead of fixing the parameters of the added resonances from PDG or pentaquark models, we add the resonance with a specific $J^P$ and leave its parameters to be determined from the experimental data. When adding only one resonance, the best result is to add one $N(3/2^-)$ resonance around $2097$ MeV, which substantially reduces the $χ^2$ per degree of freedom to $1.35$. There are two best solutions when adding two resonances, one is to add one $N^\ast(3/2^-)$ and one $N^\ast(7/2^-)$, the other is to add one $N^\ast(3/2^-)$ and one $Δ^\ast(7/2^-)$, leading the $χ^2$ per degree of freedom to be $1.10$ and $1.09$, respectively. The mass values of the $N^\ast(3/2^-)$ in the two resonance solutions are both near $2070$ MeV. Our solutions indicate that the $N(3/2^-)$ resonance around $2080$ MeV is strongly coupled with the $K^\astΣ$ final state and support the molecular picture of $N(2080, 3/2^-)$.

hep-ph↗

A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects

Multi-sensor fusion perception (MSFP) is a key technology for embodied AI, which can serve a variety of downstream tasks (e.g., 3D object detection and semantic segmentation) and application scenarios (e.g., autonomous driving and swarm robotics). Recently, impressive achievements on AI-based MSFP methods have been reviewed in relevant surveys. However, we observe that the existing surveys have some limitations after a rigorous and detailed investigation. For one thing, most surveys are oriented to a single task or research field, such as 3D object detection or autonomous driving. Therefore, researchers in other related tasks often find it difficult to benefit directly. For another, most surveys only introduce MSFP from a single perspective of multi-modal fusion, while lacking consideration of the diversity of MSFP methods, such as multi-view fusion and time-series fusion. To this end, in this paper, we hope to organize MSFP research from a task-agnostic perspective, where methods are reported from various technical views. Specifically, we first introduce the background of MSFP. Next, we review multi-modal and multi-agent fusion methods. A step further, time-series fusion methods are analyzed. In the era of LLM, we also investigate multimodal LLM fusion methods. Finally, we discuss open challenges and future directions for MSFP. We hope this survey can help researchers understand the important progress in MSFP and provide possible insights for future research.

cs.MM↗

NTIRE 2025 Challenge on HR Depth from Images of Specular and Transparent Surfaces

This paper reports on the NTIRE 2025 challenge on HR Depth From images of Specular and Transparent surfaces, held in conjunction with the New Trends in Image Restoration and Enhancement (NTIRE) workshop at CVPR 2025. This challenge aims to advance the research on depth estimation, specifically to address two of the main open issues in the field: high-resolution and non-Lambertian surfaces. The challenge proposes two tracks on stereo and single-image depth estimation, attracting about 177 registered participants. In the final testing stage, 4 and 4 participating teams submitted their models and fact sheets for the two tracks.

cs.CV↗

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data preprocessing pipeline and employ a three-stage data mixing strategy to strengthen the base model's reasoning potential. MiMo-7B-Base is pre-trained on 25 trillion tokens, with additional Multi-Token Prediction objective for enhanced performance and accelerated inference speed. During post-training, we curate a dataset of 130K verifiable mathematics and programming problems for reinforcement learning, integrating a test-difficulty-driven code-reward scheme to alleviate sparse-reward issues and employing strategic data resampling to stabilize training. Extensive evaluations show that MiMo-7B-Base possesses exceptional reasoning potential, outperforming even much larger 32B models. The final RL-tuned model, MiMo-7B-RL, achieves superior performance on mathematics, code and general reasoning tasks, surpassing the performance of OpenAI o1-mini. The model checkpoints are available at https://github.com/xiaomimimo/MiMo.

cs.CL↗

MiMo-VL Technical Report

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of 40 evaluated tasks, and scores 59.4 on OlympiadBench, surpassing models with up to 78B parameters. For GUI grounding applications, it sets a new standard with 56.1 on OSWorld-G, even outperforming specialized models such as UI-TARS. Our training combines four-stage pre-training (2.4 trillion tokens) with Mixed On-policy Reinforcement Learning (MORL) integrating diverse reward signals. We identify the importance of incorporating high-quality reasoning data with long Chain-of-Thought into pre-training stages, and the benefits of mixed RL despite challenges in simultaneous multi-domain optimization. We also contribute a comprehensive evaluation suite covering 50+ tasks to promote reproducibility and advance the field. The model checkpoints and full evaluation suite are available at https://github.com/XiaomiMiMo/MiMo-VL.

cs.CL↗