SearcharxivSearch

arXiv subjects

Mingfu Shao

Publications and source records attributed to Mingfu Shao.

9 recordsLinked to original sources

JW-SSD: A Multimodal Benchmark Dataset for Fine-Grained Sunspot Classification

Accurate sunspot classification is essential for assessing the eruptive potential of solar active regions and forecasting space weather. We present JW-SSD, a high-quality multimodal benchmark dataset for fine-grained magnetic-type classification of sunspots. Constructed from SDO/HMI SHARP 720s data (2010-2023, Solar Cycles 24 and 25), JW-SSD comprises 36,553 co-registered magnetogram-continuum pairs from 2,507 active regions. Unlike conventional three-class schemes, JW-SSD refines the Mount Wilson classification into five physically meaningful categories ({\alpha}, \b{eta}, \b{eta}-{\delta}, \b{eta}-{\gamma}, \b{eta}-{\gamma}-{\delta}), enabling finer characterization of magnetic complexity. Rigorous quality control-including central meridian distance restriction, saturation filtering, and sharpness screening-ensures high data validity. The dataset is provided in both FITS and PNG formats, with standard training (29,243) and test (7,310) splits. Benchmark experiments with four representative architectures (U-Net, ResNet-50, EfficientNet-B0, and ViT-Small) yield high accuracy across all models (89.43%-94.78% on the three-class task), confirming that the dataset is reliably learnable across diverse modeling paradigms. JW-SSD has further been employed to train JW-SunSpot, a multimodal large language model that achieves the highest classification accuracy, demonstrating the dataset's broad applicability to both conventional networks and large-language-model-based approaches.

astro-ph.SR

Deep Learning with Magnetic Parameter Constraints for Short-Term Prediction of Solar Active Region Vector Magnetic Fields

Forecasting the dynamic evolution of solar magnetic fields is a critical technique for enabling space weather warnings. Addressing the limitations of existing methods in predicting all vector magnetic field components and in maintaining consistency with solar surface magnetic-field-related quantities, this study proposes a deep learning prediction method that integrates dynamic masks of active regions with multiple magnetic parameter constraints. By constructing a three-channel representation of vector magnetic fields, applying dynamic masks to enhance attention to strong-field regions, and incorporating multi-parameter magnetic parameter constraints, we developed an end-to-end short-term (12-hour) predictive model of solar vector magnetic field evolution. Using SDO/SHARP vector magnetogram data, the model predicts and analyses field evolution across all components. Quantitative evaluations demonstrate that our approach achieves horizon-averaged structural similarity index measure (SSIM) of 0.912 (per-hour range: 0.909--0.916) and correlation coefficient (CC) of 0.998 for the radial component Br (root-mean-square error (RMSE) 13.0--21.0 G); the horizontal components achieve Bphi SSIM 0.760--0.800 (CC 0.910--0.945, RMSE 38.5--50.0 G) and Btheta SSIM 0.728--0.750 (CC 0.895--0.920, RMSE 38.5--49.0 G). The model maintains unsigned magnetic flux prediction errors at 7.82% (95% confidence interval (CI): +/-0.11%). These results demonstrate strong image-domain performance together with consistency under the magnetic-parameter diagnostics used here, suggesting initial potential for supporting future space weather forecasting efforts.

astro-ph.SR

JW-VL: A Vision-Language Model for Solar Physics

Vision-Language Models (VLMs) have achieved breakthrough progress in general knowledge domains, yet adaptation to specialized scientific fields remains challenging due to multimodal representation shifts and the limited integration of domain-specific knowledge. To address the limitations of general-purpose VLMs when applied to solar physics image recognition, analysis, and reasoning, we propose JinWu Vision-Language (JW-VL), a fine-tuned foundation model tailored for solar physics. The model integrates multi-wavelength observational data from both space-based and ground-based telescopes, encompassing representative spectral bands spanning the photosphere, chromosphere, and corona. Built upon a cross-modal alignment knowledge distillation framework, JW-VL learns a joint visual-semantic embedding that enables end-to-end modeling from raw solar observational data to downstream tasks, including solar image recognition, solar activity analysis via image-based question answering, and optical character recognition (OCR), while also supporting the construction of a multi-band, cross-instrument solar image benchmark dataset. Furthermore, as a demonstration of interdisciplinary applicability, we developed a "Daily Solar Activity Reports" agent comprising core modules for solar activity level assessment, significant active region characterization, magnetic field complexity analysis, potential space weather impact assessment, and identifying active regions for targeted observation. While JW-VL may not yet meet the rigorous, high-precision demands of operational solar physics, it bridges raw observations and diverse downstream tasks, establishing a valuable methodological framework for applying multimodal deep learning to the field.

astro-ph.SR

Advances and Challenges in Solar Flare Prediction: A Review

Solar flares, as one of the most prominent manifestations of solar activity, have a profound impact on both the Earth's space environment and human activities. As a result, accurate solar flare prediction has emerged as a central topic in space weather research. In recent years, substantial progress has been made in the field of solar flare forecasting, driven by the rapid advancements in space observation technology and the continuous improvement of data processing capabilities. This paper presents a comprehensive review of the current state of research in this area, with a particular focus on tracing the evolution of data-driven approaches -- which have progressed from early statistical learning techniques to more sophisticated machine learning and deep learning paradigms, and most recently, to the emergence of Multimodal Large Models (MLMs). Furthermore, this study examines the realistic performance of existing flare forecasting platforms, elucidating their limitations in operational space weather applications and thereby offering a practical reference for future advancements in technological optimization and system design.

astro-ph.SR

JW-Flare: Accurate Solar Flare Forecasting Method Based on Multimodal Large Language Models

Solar flares, the most powerful explosive phenomena in the solar system, may pose significant hazards to spaceborne satellites and ground-based infrastructure. Despite decades of intensive research, reliable flare prediction remains a challenging task. Large Language Models, as a milestone in artificial intelligence, exhibit exceptional general knowledge and next-token prediction capabilities. Here we introduce JW-Flare, the first Multimodal Large Language Models (MLLMs) explicitly trained for solar flare forecasting through fine-tuning on textual physic parameters of solar active regions and magnetic field images. This method demonstrates state-of-the-art (SOTA) performance for large flares prediction on the test dataset. It effectively identifies all 79 X-class flares from 18,949 test samples, yielding a True Skill Statistic (TSS) of 0.95 and a True Positive Rate (TPR) of 1.00, outperforming traditional predictive models. We further investigate the capability origins of JW-Flare through explainability experiments, revealing that solar physics knowledge acquired during pre-training contributes to flare forecasting performance. Additionally, we evaluate models of different parameter scales, confirming the Scaling_Law of Large Language Models in domain-specific applications, such as solar physics. This study marks a substantial advance in both the scale and accuracy of solar flare forecasting and opens a promising avenue for AI-driven methodologies in broader scientific domains.

astro-ph.SR

On de novo Bridging Paired-end RNA-seq Data

The high-throughput short-reads RNA-seq protocols often produce paired-end reads, with the middle portion of the fragments being unsequenced. We explore if the full-length fragments can be computationally reconstructed from the sequenced two ends in the absence of the reference genome - a problem here we refer to as de novo bridging. Solving this problem provides longer, more informative RNA-seq reads, and benefits downstream RNA-seq analysis such as transcript assembly, expression quantification, and splicing differential analysis. However, de novo bridging is a challenging and complicated task owing to alternative splicing, transcript noises, and sequencing errors. It remains unclear if the data provides sufficient information for accurate bridging, let alone efficient algorithms that determine the true bridges. Methods have been proposed to bridge paired-end reads in the presence of reference genome (called reference-based bridging), but the algorithms are far away from scaling for de novo bridging as the underlying compacted de Bruijn graph(cdBG) used in the latter task often contains millions of vertices and edges. We designed a new truncated Dijkstra's algorithm for this problem, and proposed a novel algorithm that reuses the shortest path tree to avoid running the truncated Dijkstra's algorithm from scratch for all vertices for further speeding up. These innovative techniques result in scalable algorithms that can bridge all paired-end reads in a cdBG with millions of vertices. Our experiments showed that paired-end RNA-seq reads can be accurately bridged to a large extent. The resulting tool is freely available at https://github.com/Shao-Group/rnabridge-denovo.

q-bio.GN

On the Maximal Independent Sets of $k$-mers with the Edit Distance

In computational biology, $k$-mers and edit distance are fundamental concepts. However, little is known about the metric space of all $k$-mers equipped with the edit distance. In this work, we explore the structure of the $k$-mer space by studying its maximal independent sets (MISs). An MIS is a sparse sketch of all $k$-mers with nice theoretical properties, and therefore admits critical applications in clustering, indexing, hashing, and sketching large-scale sequencing data, particularly those with high error-rates. Finding an MIS is a challenging problem, as the size of a $k$-mer space grows geometrically with respect to $k$. We propose three algorithms for this problem. The first and the most intuitive one uses a greedy strategy. The second method implements two techniques to avoid redundant comparisons by taking advantage of the locality-property of the $k$-mer space and the estimated bounds on the edit distance. The last algorithm avoids expensive calculations of the edit distance by translating the edit distance into the shortest path in a specifically designed graph. These algorithms are implemented and the calculated MISs of $k$-mer spaces and their statistical properties are reported and analyzed for $k$ up to 15. Source code is freely available at https://github.com/Shao-Group/kmerspace .

cs.DS

Locality-sensitive bucketing functions for the edit distance

Many bioinformatics applications involve bucketing a set of sequences where each sequence is allowed to be assigned into multiple buckets. To achieve both high sensitivity and precision, bucketing methods are desired to assign similar sequences into the same bucket while assigning dissimilar sequences into distinct buckets. Existing $k$-mer-based bucketing methods have been efficient in processing sequencing data with low error rate, but encounter much reduced sensitivity on data with high error rate. Locality-sensitive hashing (LSH) schemes are able to mitigate this issue through tolerating the edits in similar sequences, but state-of-the-art methods still have large gaps. Here we generalize the LSH function by allowing it to hash one sequence into multiple buckets. Formally, a bucketing function, which maps a sequence (of fixed length) into a subset of buckets, is defined to be $(d_1, d_2)$-sensitive if any two sequences within an edit distance of $d_1$ are mapped into at least one shared bucket, and any two sequences with distance at least $d_2$ are mapped into disjoint subsets of buckets. We construct locality-sensitive bucketing (LSB) functions with a variety of values of $(d_1,d_2)$ and analyze their efficiency with respect to the total number of buckets needed as well as the number of buckets that a specific sequence is mapped to. We also prove lower bounds of these two parameters in different settings and show that some of our constructed LSB functions are optimal. These results provide theoretical foundations for their practical use in analyzing sequences with high error rate while also providing insights for the hardness of designing ungapped LSH functions.

cs.DS

Improving protein threading accuracy via combining local and global potential using TreeCRF model

Protein structure prediction remains to be an open problem in bioinformatics. There are two main categories of methods for protein structure prediction: Free Modeling (FM) and Template Based Modeling (TBM). Protein threading, belonging to the category of template based modeling, identifies the most likely fold with the target by making a sequence-structure alignment between target protein and template protein. Though protein threading has been shown to more be successful for protein structure prediction, it performs poorly for remote homology detection.

q-bio.BM