Searcharxiv⌕ Search

arXiv subjects

Jun Shi

Publications and source records attributed to Jun Shi.

At least 73 records · Page 4Linked to original sources

Deep Synoptic Array Science: Polarimetry of 25 New Fast Radio Bursts Provides Insights into their Origins

We report on a full-polarization analysis of the first 25 as yet non-repeating FRBs detected at 1.4 GHz by the 110-antenna Deep Synoptic Array (DSA-110) during commissioning observations. We present details of the data-reduction, calibration, and analysis procedures developed for this novel instrument. Faraday rotation measures (RMs) are searched between $\pm10^6$ rad m$^{-2}$ and detected for 20 FRBs with magnitudes ranging from $4-4670$ rad m$^{-2}$. $15/25$ FRBs are consistent with 100% polarization, 10 of which have high ($\ge70\%$) linear-polarization fractions and 2 of which have high ($\ge30\%$) circular-polarization fractions. Our results disfavor multipath RM scattering as a dominant depolarization mechanism. Polarization-state and possible RM variations are observed in the four FRBs with multiple sub-components. We combine the DSA-110 sample with polarimetry of previously published FRBs, and compare the polarization properties of FRB sub-populations and FRBs with Galactic pulsars. Although FRB polarization fractions are typically higher than those of Galactic pulsars, and cover a wider range than those of pulsar single pulses, they resemble those of the youngest (characteristic ages $<10^{5}$ yr) pulsars. Our results support a scenario wherein FRB emission is intrinsically highly linearly polarized, and propagation effects can result in conversion to circular polarization and depolarization. Young pulsar emission and magnetospheric-propagation geometries may form a useful analogy for the origin of FRB polarization.

astro-ph.HE↗

RobotGPT: Robot Manipulation Learning from ChatGPT

We present RobotGPT, an innovative decision framework for robotic manipulation that prioritizes stability and safety. The execution code generated by ChatGPT cannot guarantee the stability and safety of the system. ChatGPT may provide different answers for the same task, leading to unpredictability. This instability prevents the direct integration of ChatGPT into the robot manipulation loop. Although setting the temperature to 0 can generate more consistent outputs, it may cause ChatGPT to lose diversity and creativity. Our objective is to leverage ChatGPT's problem-solving capabilities in robot manipulation and train a reliable agent. The framework includes an effective prompt structure and a robust learning model. Additionally, we introduce a metric for measuring task difficulty to evaluate ChatGPT's performance in robot manipulation. Furthermore, we evaluate RobotGPT in both simulation and real-world environments. Compared to directly using ChatGPT to generate code, our framework significantly improves task success rates, with an average increase from 38.5% to 91.5%. Therefore, training a RobotGPT by utilizing ChatGPT as an expert is a more stable approach compared to directly using ChatGPT as a task planner.

cs.RO↗

Single-shot Phase Retrieval from a Fractional Fourier Transform Perspective

The realm of classical phase retrieval concerns itself with the arduous task of recovering a signal from its Fourier magnitude measurements, which are fraught with inherent ambiguities. A single-exposure intensity measurement is commonly deemed insufficient for the reconstruction of the primal signal, given that the absent phase component is imperative for the inverse transformation. In this work, we present a novel single-shot phase retrieval paradigm from a fractional Fourier transform (FrFT) perspective, which involves integrating the FrFT-based physical measurement model within a self-supervised reconstruction scheme. Specifically, the proposed FrFT-based measurement model addresses the aliasing artifacts problem in the numerical calculation of Fresnel diffraction, featuring adaptability to both short-distance and long-distance propagation scenarios. Moreover, the intensity measurement in the FrFT domain proves highly effective in alleviating the ambiguities of phase retrieval and relaxing the previous conditions on oversampled or multiple measurements in the Fourier domain. Furthermore, the proposed self-supervised reconstruction approach harnesses the fast discrete algorithm of FrFT alongside untrained neural network priors, thereby attaining preeminent results. Through numerical simulations, we demonstrate that both amplitude and phase objects can be effectively retrieved from a single-shot intensity measurement using the proposed approach and provide a promising technique for support-free coherent diffraction imaging.

cs.CV↗

Deep Synoptic Array Science: Implications of Faraday Rotation Measures of Localized Fast Radio Bursts

Faraday rotation measures (RMs) of fast radio bursts (FRBs) offer the prospect of directly measuring extragalactic magnetic fields. We present an analysis of the RMs of ten as yet non-repeating FRBs detected and localized to host galaxies by the 110-antenna Deep Synoptic Array (DSA-110). We combine this sample with published RMs of 15 localized FRBs, nine of which are repeating sources. For each FRB in the combined sample, we estimate the host-galaxy dispersion measure (DM) contributions and extragalactic RM. We find compelling evidence that the extragalactic components of FRB RMs are often dominated by contributions from the host-galaxy interstellar medium (ISM). Specifically, we find that both repeating and as yet non-repeating FRBs show a correlation between the host-DM and host-RM in the rest frame, and we find an anti-correlation between extragalactic RM (in the observer frame) and redshift for non-repeaters, as expected if the magnetized plasma is in the host galaxy. Important exceptions to the ISM origin include a dense, magnetized circum-burst medium in some repeating FRBs, and the intra-cluster medium (ICM) of host or intervening galaxy clusters. We find that the estimated ISM magnetic-field strengths, $\bar{B}_{||}$, are characteristically larger than those inferred from Galactic radio pulsars. This suggests either increased ISM magnetization in FRB hosts in comparison with the Milky Way, or that FRBs preferentially reside in regions of increased magnetic-field strength within their hosts.

astro-ph.HE↗

Sparse Sampling Transformer with Uncertainty-Driven Ranking for Unified Removal of Raindrops and Rain Streaks

In the real world, image degradations caused by rain often exhibit a combination of rain streaks and raindrops, thereby increasing the challenges of recovering the underlying clean image. Note that the rain streaks and raindrops have diverse shapes, sizes, and locations in the captured image, and thus modeling the correlation relationship between irregular degradations caused by rain artifacts is a necessary prerequisite for image deraining. This paper aims to present an efficient and flexible mechanism to learn and model degradation relationships in a global view, thereby achieving a unified removal of intricate rain scenes. To do so, we propose a Sparse Sampling Transformer based on Uncertainty-Driven Ranking, dubbed UDR-S2Former. Compared to previous methods, our UDR-S2Former has three merits. First, it can adaptively sample relevant image degradation information to model underlying degradation relationships. Second, explicit application of the uncertainty-driven ranking strategy can facilitate the network to attend to degradation features and understand the reconstruction process. Finally, experimental results show that our UDR-S2Former clearly outperforms state-of-the-art methods for all benchmarks.

cs.CV↗

Multi-Scale Prototypical Transformer for Whole Slide Image Classification

Whole slide image (WSI) classification is an essential task in computational pathology. Despite the recent advances in multiple instance learning (MIL) for WSI classification, accurate classification of WSIs remains challenging due to the extreme imbalance between the positive and negative instances in bags, and the complicated pre-processing to fuse multi-scale information of WSI. To this end, we propose a novel multi-scale prototypical Transformer (MSPT) for WSI classification, which includes a prototypical Transformer (PT) module and a multi-scale feature fusion module (MFFM). The PT is developed to reduce redundant instances in bags by integrating prototypical learning into the Transformer architecture. It substitutes all instances with cluster prototypes, which are then re-calibrated through the self-attention mechanism of the Trans-former. Thereafter, an MFFM is proposed to fuse the clustered prototypes of different scales, which employs MLP-Mixer to enhance the information communication between prototypes. The experimental results on two public WSI datasets demonstrate that the proposed MSPT outperforms all the compared algorithms, suggesting its potential applications.

cs.CV↗

H-DenseFormer: An Efficient Hybrid Densely Connected Transformer for Multimodal Tumor Segmentation

Recently, deep learning methods have been widely used for tumor segmentation of multimodal medical images with promising results. However, most existing methods are limited by insufficient representational ability, specific modality number and high computational complexity. In this paper, we propose a hybrid densely connected network for tumor segmentation, named H-DenseFormer, which combines the representational power of the Convolutional Neural Network (CNN) and the Transformer structures. Specifically, H-DenseFormer integrates a Transformer-based Multi-path Parallel Embedding (MPE) module that can take an arbitrary number of modalities as input to extract the fusion features from different modalities. Then, the multimodal fusion features are delivered to different levels of the encoder to enhance multimodal learning representation. Besides, we design a lightweight Densely Connected Transformer (DCT) block to replace the standard Transformer block, thus significantly reducing computational complexity. We conduct extensive experiments on two public multimodal datasets, HECKTOR21 and PI-CAI22. The experimental results show that our proposed method outperforms the existing state-of-the-art methods while having lower computational complexity. The source code is available at https://github.com/shijun18/H-DenseFormer.

eess.IV↗

Q-YOLO: Efficient Inference for Real-time Object Detection

Real-time object detection plays a vital role in various computer vision applications. However, deploying real-time object detectors on resource-constrained platforms poses challenges due to high computational and memory requirements. This paper describes a low-bit quantization method to build a highly efficient one-stage detector, dubbed as Q-YOLO, which can effectively address the performance degradation problem caused by activation distribution imbalance in traditional quantized YOLO models. Q-YOLO introduces a fully end-to-end Post-Training Quantization (PTQ) pipeline with a well-designed Unilateral Histogram-based (UH) activation quantization scheme, which determines the maximum truncation values through histogram analysis by minimizing the Mean Squared Error (MSE) quantization errors. Extensive experiments on the COCO dataset demonstrate the effectiveness of Q-YOLO, outperforming other PTQ methods while achieving a more favorable balance between accuracy and computational cost. This research contributes to advancing the efficient deployment of object detection models on resource-limited edge devices, enabling real-time detection with reduced computational and memory overhead.

cs.CV↗

Multi-View Attention Learning for Residual Disease Prediction of Ovarian Cancer

In the treatment of ovarian cancer, precise residual disease prediction is significant for clinical and surgical decision-making. However, traditional methods are either invasive (e.g., laparoscopy) or time-consuming (e.g., manual analysis). Recently, deep learning methods make many efforts in automatic analysis of medical images. Despite the remarkable progress, most of them underestimated the importance of 3D image information of disease, which might brings a limited performance for residual disease prediction, especially in small-scale datasets. To this end, in this paper, we propose a novel Multi-View Attention Learning (MuVAL) method for residual disease prediction, which focuses on the comprehensive learning of 3D Computed Tomography (CT) images in a multi-view manner. Specifically, we first obtain multi-view of 3D CT images from transverse, coronal and sagittal views. To better represent the image features in a multi-view manner, we further leverage attention mechanism to help find the more relevant slices in each view. Extensive experiments on a dataset of 111 patients show that our method outperforms existing deep-learning methods.

eess.IV↗

Weakly Supervised Lesion Detection and Diagnosis for Breast Cancers with Partially Annotated Ultrasound Images

Deep learning (DL) has proven highly effective for ultrasound-based computer-aided diagnosis (CAD) of breast cancers. In an automaticCAD system, lesion detection is critical for the following diagnosis. However, existing DL-based methods generally require voluminous manually-annotated region of interest (ROI) labels and class labels to train both the lesion detection and diagnosis models. In clinical practice, the ROI labels, i.e. ground truths, may not always be optimal for the classification task due to individual experience of sonologists, resulting in the issue of coarse annotation that limits the diagnosis performance of a CAD model. To address this issue, a novel Two-Stage Detection and Diagnosis Network (TSDDNet) is proposed based on weakly supervised learning to enhance diagnostic accuracy of the ultrasound-based CAD for breast cancers. In particular, all the ROI-level labels are considered as coarse labels in the first training stage, and then a candidate selection mechanism is designed to identify optimallesion areas for both the fully and partially annotated samples. It refines the current ROI-level labels in the fully annotated images and the detected ROIs in the partially annotated samples with a weakly supervised manner under the guidance of class labels. In the second training stage, a self-distillation strategy further is further proposed to integrate the detection network and classification network into a unified framework as the final CAD model for joint optimization, which then further improves the diagnosis performance. The proposed TSDDNet is evaluated on a B-mode ultrasound dataset, and the experimental results show that it achieves the best performance on both lesion detection and diagnosis tasks, suggesting promising application potential.

eess.IV↗

Multi-scale Efficient Graph-Transformer for Whole Slide Image Classification

The multi-scale information among the whole slide images (WSIs) is essential for cancer diagnosis. Although the existing multi-scale vision Transformer has shown its effectiveness for learning multi-scale image representation, it still cannot work well on the gigapixel WSIs due to their extremely large image sizes. To this end, we propose a novel Multi-scale Efficient Graph-Transformer (MEGT) framework for WSI classification. The key idea of MEGT is to adopt two independent Efficient Graph-based Transformer (EGT) branches to process the low-resolution and high-resolution patch embeddings (i.e., tokens in a Transformer) of WSIs, respectively, and then fuse these tokens via a multi-scale feature fusion module (MFFM). Specifically, we design an EGT to efficiently learn the local-global information of patch tokens, which integrates the graph representation into Transformer to capture spatial-related information of WSIs. Meanwhile, we propose a novel MFFM to alleviate the semantic gap among different resolution patches during feature fusion, which creates a non-patch token for each branch as an agent to exchange information with another branch by cross-attention. In addition, to expedite network training, a novel token pruning module is developed in EGT to reduce the redundant tokens. Extensive experiments on TCGA-RCC and CAMELYON16 datasets demonstrate the effectiveness of the proposed MEGT.

cs.CV↗

$Σ$ Resonances from a Neural Network-based Partial Wave Analysis on $K^-p$ Scattering

We implement a convolutional neural network to study the $Σ$ hyperons using experimental data of the $K^-p\toπ^0Λ$ reaction. The averaged accuracy of the NN models in resolving resonances on the test data sets is ${\rm 98.5\%}$, ${\rm 94.8\%}$ and ${\rm 82.5\%}$ for one-, two- and three-additional-resonance case. We find that the three most significant resonances are $1/2^+$, $3/2^+$ and $3/2^-$ states with mass being ${\rm 1.62(11)~GeV}$, ${\rm 1.72(6)~GeV}$ and ${\rm 1.61(9)~GeV}$, and probability being $\rm 100(3)\%$, $\rm 72(24)\%$ and $\rm 98(52)\%$, respectively, where the errors mostly come from the uncertainties of the experimental data. Our results support the three-star $Σ(1660)1/2^+$, the one-star $Σ(1780)3/2^+$ and the one-star $Σ(1580)3/2^-$ in PDG. The ability of giving quantitative probabilities in resonance resolving and numerical stability make NN potentially a life-changing tool in baryon partial wave analysis, and this approach can be easily extended to accommodate other theoretical models and/or to include more experimental data.

hep-ph↗

Fast MRI Reconstruction via Edge Attention

Fast and accurate MRI reconstruction is a key concern in modern clinical practice. Recently, numerous Deep-Learning methods have been proposed for MRI reconstruction, however, they usually fail to reconstruct sharp details from the subsampled k-space data. To solve this problem, we propose a lightweight and accurate Edge Attention MRI Reconstruction Network (EAMRI) to reconstruct images with edge guidance. Specifically, we design an efficient Edge Prediction Network to directly predict accurate edges from the blurred image. Meanwhile, we propose a novel Edge Attention Module (EAM) to guide the image reconstruction utilizing the extracted edge priors, as inspired by the popular self-attention mechanism. EAM first projects the input image and edges into Q_image, K_edge, and V_image, respectively. Then EAM pairs the Q_image with K_edge along the channel dimension, such that 1) it can search globally for the high-frequency image features that are activated by the edge priors; 2) the overall computation burdens are largely reduced compared with the traditional spatial-wise attention. With the help of EAM, the predicted edge priors can effectively guide the model to reconstruct high-quality MR images with accurate edges. Extensive experiments show that our proposed EAMRI outperforms other methods with fewer parameters and can recover more accurate edges.

eess.IV↗

Pseudo-Data based Self-Supervised Federated Learning for Classification of Histopathological Images

Computer-aided diagnosis (CAD) can help pathologists improve diagnostic accuracy together with consistency and repeatability for cancers. However, the CAD models trained with the histopathological images only from a single center (hospital) generally suffer from the generalization problem due to the straining inconsistencies among different centers. In this work, we propose a pseudo-data based self-supervised federated learning (FL) framework, named SSL-FT-BT, to improve both the diagnostic accuracy and generalization of CAD models. Specifically, the pseudo histopathological images are generated from each center, which contains inherent and specific properties corresponding to the real images in this center, but does not include the privacy information. These pseudo images are then shared in the central server for self-supervised learning (SSL). A multi-task SSL is then designed to fully learn both the center-specific information and common inherent representation according to the data characteristics. Moreover, a novel Barlow Twins based FL (FL-BT) algorithm is proposed to improve the local training for the CAD model in each center by conducting contrastive learning, which benefits the optimization of the global model in the FL procedure. The experimental results on three public histopathological image datasets indicate the effectiveness of the proposed SSL-FL-BT on both diagnostic accuracy and generalization.

cs.CV↗

DEHRFormer: Real-time Transformer for Depth Estimation and Haze Removal from Varicolored Haze Scenes

Varicolored haze caused by chromatic casts poses haze removal and depth estimation challenges. Recent learning-based depth estimation methods are mainly targeted at dehazing first and estimating depth subsequently from haze-free scenes. This way, the inner connections between colored haze and scene depth are lost. In this paper, we propose a real-time transformer for simultaneous single image Depth Estimation and Haze Removal (DEHRFormer). DEHRFormer consists of a single encoder and two task-specific decoders. The transformer decoders with learnable queries are designed to decode coupling features from the task-agnostic encoder and project them into clean image and depth map, respectively. In addition, we introduce a novel learning paradigm that utilizes contrastive learning and domain consistency learning to tackle weak-generalization problem for real-world dehazing, while predicting the same depth map from the same scene with varicolored haze. Experiments demonstrate that DEHRFormer achieves significant performance improvement across diverse varicolored haze scenes over previous depth estimation networks and dehazing approaches.

cs.CV↗

SANDFORMER: CNN and Transformer under Gated Fusion for Sand Dust Image Restoration

Although Convolutional Neural Networks (CNN) have made good progress in image restoration, the intrinsic equivalence and locality of convolutions still constrain further improvements in image quality. Recent vision transformer and self-attention have achieved promising results on various computer vision tasks. However, directly utilizing Transformer for image restoration is a challenging task. In this paper, we introduce an effective hybrid architecture for sand image restoration tasks, which leverages local features from CNN and long-range dependencies captured by transformer to improve the results further. We propose an efficient hybrid structure for sand dust image restoration to solve the feature inconsistency issue between Transformer and CNN. The framework complements each representation by modulating features from the CNN-based and Transformer-based branches rather than simply adding or concatenating features. Experiments demonstrate that SandFormer achieves significant performance improvements in synthetic and real dust scenes compared to previous sand image restoration methods.

cs.CV↗

Deep Synoptic Array science: A massive elliptical host among two galaxy-cluster fast radio bursts

The stellar population environments associated with fast radio burst (FRB) sources provide important insights for developing their progenitor theories. We expand the diversity of known FRB host environments by reporting two FRBs in massive galaxy clusters discovered by the Deep Synoptic Array (DSA-110) during its commissioning observations. FRB 20220914A has been localized to a star-forming, late-type galaxy at a redshift of 0.1139 with multiple starbursts at lookback times less than $\sim$3.5 Gyr in the Abell 2310 galaxy cluster. Although the host galaxy of FRB 20220914A is similar to typical FRB hosts, the FRB 20220509G host stands out as a quiescent, early-type galaxy at a redshift of 0.0894 in the Abell 2311 galaxy cluster. The discovery of FRBs in both late and early-type galaxies adds to the body of evidence that the FRB sources have multiple formation channels. Therefore, even though FRB hosts are typically star-forming, there must exist formation channels consistent with old stellar population in galaxies. The varied star formation histories of the two FRB hosts we report indicate a wide delay-time distribution of FRB progenitors. Future work in constraining the FRB delay-time distribution, using methods we develop herein, will prove crucial in determining the evolutionary histories of FRB sources.

astro-ph.HE↗

Deep Synoptic Array science: Two fast radio burst sources in massive galaxy clusters

The hot gas that constitutes the intracluster medium (ICM) has been studied at X-ray and millimeter/sub-millimeter wavelengths (Sunyaev-Zeldovich effect) for decades. Fast radio bursts (FRBs) offer an additional method of directly measuring the ICM and gas surrounding clusters, via observables such as dispersion measure (DM) and Faraday rotation measure (RM). We report the discovery of two FRB sources detected with the Deep Synoptic Array (DSA-110) whose host galaxies belong to massive galaxy clusters. In both cases, the FRBs exhibit excess extragalactic DM, some of which likely originates in the ICM of their respective clusters. FRB 20220914A resides in the galaxy cluster Abell 2310 at z=0.1125 with a projected offset from the cluster center of 520 kpc. The host of a second source, FRB 20220509G, is an elliptical galaxy at z=0.0894 that belongs to the galaxy cluster Abell 2311 at projected offset of 870 kpc. These sources represent the first time an FRB has been localized to a galaxy cluster. We combine our FRB data with archival X-ray, SZ, and optical observations of these clusters in order to infer properties of the ICM, including a measurement of gas temperature from DM and ySZ of 0.8-3.9 keV. We then compare our results to massive cluster halos from the IllustrisTNG simulation. Finally, we describe how large samples of localized FRBs from future surveys will constrain the ICM, particularly beyond the virial radius of clusters.

astro-ph.HE↗