SearcharxivSearch

arXiv subjects

Meiqi Wang

Publications and source records attributed to Meiqi Wang.

At least 19 recordsLinked to original sources

Toward Operational Solar Flare Peak Flux Nowcasting: A Strategy Combining Real-Time Data, Machine Learning, and NOAA Flare Detection Criteria

We present the RMN strategy (Real-time data, machine learning, and NOAA flare detection criteria) for nowcasting the peak soft X-ray flux of ongoing solar flares under operationally realistic conditions. The strategy combines real-time GOES 0.1-0.8 nm X-ray observations with an attention-based sequence-to-sequence Long Short-Term Memory model. Under the NOAA flare detection criteria, predictions are evaluated at one-minute intervals from three minutes after the cataloged onset to the observed peak using the preceding 60 minutes of X-ray observations. We apply the RMN strategy to C-, M-, and X-class flares observed by GOES-8-18 from 1997 to 2024 using four-fold cross-validation. The major results of this study are as follows. First, the model nowcasts peak soft X-ray flux with RMSE and PE values of 0.26 and 3.11\% for the $\geq$C-class group, 0.45 and 5.59\% for the $\geq$M-class group, and 0.87 and 12.76\% for the X-class group. The higher discrepancy toward stronger flare groups indicates that peak-flux prediction is more challenging for higher-intensity flares. Second, the model performance depends on flare rise time and prediction time, with larger errors for longer rise time events and improved performance as the prediction time approaches the flare peak. Shorter rise time events approach their final peak more rapidly, providing a clearer indication of the eventual peak, whereas the larger difference for longer rise time events may partly reflect more complex temporal evolution. Third, empirical coverage based on total uncertainty remains high but decreases for stronger flares, with noise uncertainty contributing more than model uncertainty.

astro-ph.SR

Compiling and Benchmarking Task-State Horizons for Embodied Agents

Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long-horizon evaluation, but primarily characterize difficulty through action-sequence length and subtask complexity, overlooking a distinct challenge: agents must track evolving task-relevant world states induced by both their exploration and environmental dynamics. We define the span of task-relevant state transitions that an agent must track as task-state horizon (TSH). To evaluate how agent performance varies with TSH, we introduce RoboGraph, a robotic task compiler that translates state-transition dependencies into executable symbolic graphs. Specifically, RoboGraph constructs task-state horizons from spatial and temporal causal dependencies, including those induced by unexpected failures and interventions during task execution. Building on RoboGraph, we release a benchmark comprising 588 episodes across 84 scenes with varying TSHs. Experiments evaluating 15 advanced agentic models in both semantic and visual closed-loop environments show that most models struggle with demanding TSHs, revealing substantial gaps in maintaining, exploring, and updating task-relevant state over long horizon.

cs.RO

Non-thermal Sources from Stereoscopic Hard X-ray and Earth-based Microwave Observations in a Data-Constrained Magnetohydrodynamic Simulation

We analyze the X7.1 flare on 2024 October 1 from NOAA AR 13842 using hard X-ray (HXR) imaging, microwave observations by the Expanded Owens Valley Solar Array (EOVSA), and a three-dimensional Magnetohydrodynamic (MHD) simulation. The flare was observed from two vantage points, with Solar Orbiter/Spectrometer Telescope for Imaging X-rays viewing the flare near the limb and Advanced Space-based Solar Observatory/Hard X-ray Imager and EOVSA observing it on the disk. We carried out a data-constrained MHD simulation using a nonlinear force-free field extrapolation as the initial condition and constrained the height of the non-thermal looptop source from stereoscopic HXR and microwave observations. The height is consistent between the stereoscopic analysis and the MHD simulation. A secondary non-thermal microwave source aligned with a southward plasma ejection corresponds to an elongated current sheet. Although the current sheet grows in multiple directions, the secondary microwave emission is observed only from the southern segment. This localization suggests reconnection in regions with different magnetic field strengths. Reconnection in strong-field regions produces flare arcades with dominant looptop emission, whereas reconnection in weaker southern regions gives rise to secondary microwave emission at higher altitudes. The height of the secondary source is consistent between the stereoscopic analysis and the MHD simulation. Microwave spectral fitting suggests a higher low-energy cutoff for non-thermal electrons in the secondary microwave source than in the main looptop source. This may reflect the transport of electrons pre-accelerated near the looptop source by the southward plasma ejection.

astro-ph.SR

Few Made It Out: A Multi-Messenger Study of an In Situ Solar Energetic Electron Event Driven by a Solar Jet

When in situ solar energetic electron (SEE) events are closely associated with nonthermal flares, the escaping electron population is frequently observed to be much smaller than the nonthermal-radiation-emitting population near the solar surface. If a single accelerated population drives both signatures, the physical mechanism causing this severe deficit of upward-propagating electrons remains poorly understood.Focusing on one of the 2022 November 10--12 SEE events associated with recurrent solar jets and interplanetary type III radio bursts, we present a new, combined microwave--X-ray analysis using the Expanded Owens Valley Solar Array (EOVSA) and the Spectrometer/Telescope for Imaging X-rays (STIX) aboard Solar Orbiter. This synergy enables, for the first time for such an event, spatially resolved diagnostics over a broad energy spectrum of the near-Sun energetic electrons, complemented by in situ measurements made by spacecraft at multiple heliocentric longitudes and distances. Consistent with earlier results based on in situ and X-ray data, our results show that only 0.1--1\% of energetic electrons escape into interplanetary space. Crucially, the new microwave spectral imaging analysis suggests that energetic electrons are strongly concentrated in a compact region just above a mini-flare arcade at the base of the jet spire, and that their number density decreases by at least two orders of magnitude in the direction of the jet spire away from this region. This steep gradient, revealed by the microwave diagnostics, points to efficient local acceleration and trapping in the region analogous to the above-the-looptop ``magnetic bottle'' region in major eruptive flares, allowing only a small fraction of electrons to access open magnetic field lines and enter interplanetary space.

astro-ph.SR

Stereoscopic Observations of Solar X-ray Sources Explained by a Data-Constrained Magnetohydrodynamic Simulation

We investigated the three-dimensional (3D) magnetic structures and dynamics responsible for particle acceleration in an X7.1-class flare that occurred on October 1, 2024, in NOAA active region 13842. We combined stereoscopic hard X-ray (HXR) observations from the Advanced Space-based Solar Observatory/Hard X-ray Imager (HXI) and the Solar Orbiter/Spectrometer Telescope for Imaging X-rays (STIX) with a 3D magnetohydrodynamic (MHD) simulation constrained by observed photospheric magnetic fields. During the two main peaks of the impulsive phase, HXR footpoints appeared at different locations, indicating a migration of the primary reconnection site in the corona. Our data-constrained MHD simulation successfully reproduced the reconnected field lines linking the observed conjugate HXR footpoints. Furthermore, the simulation shows that these primary reconnections occur along a single quasi-separatrix layer (QSL) system. Therefore, the two main peaks of HXR can be interpreted as episodic energy release within the single QSL system. This study demonstrates that the data-constrained MHD model provides a realistic 3D magnetic context for interpreting HXR emission. Notably, STIX observations revealed a vertically distributed thermal HXR source, extending from the footpoints to the looptop, with its centroid migrating between the two peaks. This marks a first step toward understanding the particle acceleration processes in solar flares.

astro-ph.SR

Solar Spicules, Filigrees and Solar Wind Switchbacks

Spicules, the smallest observable jet-like dynamic features ubiquitous in the chromosphere, are supposedly an important potential source for small-scale solar wind transients, with supporting evidence yet needed. We studied the high-resolution H-alpha images (0.10'') and magnetograms (0.29'') from Big Bear Solar Observatory (BBSO) to find that spicules are an ideal candidate for the solar wind magnetic switchbacks detected by the Parker Solar Probe (PSP). It is not that spicules are a miniature of coronal jets, but that they have unique properties not found in other solar candidates in explaining solar origin of switchbacks. (1) The spicules under this study originate from filigrees, all in a single magnetic polarity. Since filigrees are known as footpoints of open fields, the spicule guiding field lines can form a unipolar funnel, which is needed to create an SB patch, a group of fieldlines that switch from one common base polarity to the other polarity. (2) The spicules come in a cluster lined up along a supergranulation boundary, and the simulated waiting times from their spatial intervals exhibit a number distribution continuously decreasing from a few sec to ~30 min, similar to that of switchbacks. (3) From a time-distance map for spicules, we estimate their occurrence rate as 0.55 spicules per Mm^2 and second, sufficiently high for detection by PSP. In addition the dissimilarity of spicules with coronal jets, including the absence of base brightening and low correlation with EUV emission is briefly discussed.

astro-ph.SR

Block Circulant Adapter for Large Language Models

Fine-tuning large language models (LLMs) is difficult due to their huge model size. Recent Fourier domain-based methods show potential for reducing fine-tuning costs. We propose a block circulant matrix-based fine-tuning method with a stable training heuristic to leverage the properties of circulant matrices and one-dimensional Fourier transforms to reduce storage and computation costs. Experiments show that our method uses $14\times$ less number of parameters than VeRA, $16\times$ smaller than LoRA and $32\times$ less FLOPs than FourierFT, while maintaining close or better task performance. Our approach presents a promising way in frequency domain to fine-tune large models on downstream tasks.

cs.CL

Two Phases of Particle Acceleration of a Solar Flare Associated with in situ Energetic Particles

How impulsive solar energetic particle (SEP) events are produced by magnetic-reconnection-driven processes during solar flares remains an outstanding question. Here we report a short-duration SEP event associated with an X-class eruptive flare on July 03, 2021, using a combination of remote sensing observations and in situ measurements. The in situ SEPs were recorded by multiple spacecraft including the Parker Solar Probe. The hard X-ray (HXR) light curve exhibits two impulsive periods. The first period is characterized by a single peak with a rapid rise and decay, while the second period features a more gradual HXR light curve with a harder spectrum. Such observation is consistent with in situ measurements: the energetic electrons were first released during the early impulsive phase when the eruption was initiated. The more energetic in situ electrons were released several minutes later during the second period of the impulsive phase when the eruption was well underway. This second period of energetic electron acceleration also coincides with the release of in situ energetic protons and the onset of an interplanetary type III radio burst. We conclude that these multi-messenger observations favor a two-phase particle acceleration scenario: the first, less energetic electron population was produced during the initial reconnection that triggers the flare eruption, and the second, more energetic electron population was accelerated in the above-the-looptop region below a well-developed, large-scale reconnection current sheet induced by the eruption.

astro-ph.SR

Deep Learning for Medical Text Processing: BERT Model Fine-Tuning and Comparative Study

This paper proposes a medical literature summary generation method based on the BERT model to address the challenges brought by the current explosion of medical information. By fine-tuning and optimizing the BERT model, we develop an efficient summary generation system that can quickly extract key information from medical literature and generate coherent, accurate summaries. In the experiment, we compared various models, including Seq-Seq, Attention, Transformer, and BERT, and demonstrated that the improved BERT model offers significant advantages in the Rouge and Recall metrics. Furthermore, the results of this study highlight the potential of knowledge distillation techniques to further enhance model performance. The system has demonstrated strong versatility and efficiency in practical applications, offering a reliable tool for the rapid screening and analysis of medical literature.

cs.CL

Real-Time Summarization of Twitter

In this paper, we describe our approaches to TREC Real-Time Summarization of Twitter. We focus on real time push notification scenario, which requires a system monitors the stream of sampled tweets and returns the tweets relevant and novel to given interest profiles. Dirichlet score with and with very little smoothing (baseline) are employed to classify whether a tweet is relevant to a given interest profile. Using metrics including Mean Average Precision (MAP, cumulative gain (CG) and discount cumulative gain (DCG), the experiment indicates that our approach has a good performance. It is also desired to remove the redundant tweets from the pushing queue. Due to the precision limit, we only describe the algorithm in this paper.

cs.LG

A Case for Application-Aware Space Radiation Tolerance in Orbital Computing

We are witnessing a surge in the use of commercial off-the-shelf (COTS) hardware for cost-effective in-orbit computing, such as deep neural network (DNN) based on-satellite sensor data processing, Earth object detection, and task decision.However, once exposed to harsh space environments, COTS hardware is vulnerable to cosmic radiation and suffers from exhaustive single-event upsets (SEUs) and multi-unit upsets (MCUs), both threatening the functionality and correctness of in-orbit computing.Existing hardware and system software protections against radiation are expensive for resource-constrained COTS nanosatellites and overwhelming for upper-layer applications due to their requirement for heavy resource redundancy and frequent reboots. Instead, we make a case for cost-effective space radiation tolerance using application domain knowledge. Our solution for the on-satellite DNN tasks, \name, exploits the uneven SEU/MCU sensitivity across DNN layers and MCUs' spatial correlation for lightweight radiation-tolerant in-orbit AI computing. Our extensive experiments using Chaohu-1 SAR satellite payloads and a hardware-in-the-loop, real data-driven space radiation emulator validate that RedNet can suppress the influence of radiation errors to $\approx$ 0 and accelerate the on-satellite DNN inference speed by 8.4%-33.0% at negligible extra costs.

cs.ET

Online Learning of Multiple Tasks and Their Relationships : Testing on Spam Email Data and EEG Signals Recorded in Construction Fields

This paper examines an online multi-task learning (OMTL) method, which processes data sequentially to predict labels across related tasks. The framework learns task weights and their relatedness concurrently. Unlike previous models that assumed static task relatedness, our approach treats tasks as initially independent, updating their relatedness iteratively using newly calculated weight vectors. We introduced three rules to update the task relatedness matrix: OMTLCOV, OMTLLOG, and OMTLVON, and compared them against a conventional method (CMTL) that uses a fixed relatedness value. Performance evaluations on three datasets a spam dataset and two EEG datasets from construction workers under varying conditions demonstrated that our OMTL methods outperform CMTL, improving accuracy by 1% to 3% on EEG data, and maintaining low error rates around 12% on the spam dataset.

cs.LG

Exploration of Attention Mechanism-Enhanced Deep Learning Models in the Mining of Medical Textual Data

The research explores the utilization of a deep learning model employing an attention mechanism in medical text mining. It targets the challenge of analyzing unstructured text information within medical data. This research seeks to enhance the model's capability to identify essential medical information by incorporating deep learning and attention mechanisms. This paper reviews the basic principles and typical model architecture of attention mechanisms and shows the effectiveness of their application in the tasks of disease prediction, drug side effect monitoring, and entity relationship extraction. Aiming at the particularity of medical texts, an adaptive attention model integrating domain knowledge is proposed, and its ability to understand medical terms and process complex contexts is optimized. The experiment verifies the model's effectiveness in improving task accuracy and robustness, especially when dealing with long text. The future research path of enhancing model interpretation, realizing cross-domain knowledge transfer, and adapting to low-resource scenarios is discussed in the research outlook, which provides a new perspective and method support for intelligent medical information processing and clinical decision assistance. Finally, cross-domain knowledge transfer and adaptation strategies for low-resource scenarios, providing theoretical basis and technical reference for promoting the development of intelligent medical information processing and clinical decision support systems.

cs.CL

The Solar Origin of an In Situ Type III Radio Burst Event

Solar type III radio bursts are generated by beams of energetic electrons that travel along open magnetic field lines through the corona and into interplanetary space. However, understanding the source of these electrons and how they escape into interplanetary space remains an outstanding topic. Here we report multi-instrument, multi-perspective observations of an interplanetary type III radio burst event shortly after the second perihelion of the Parker Solar Probe (PSP). This event was associated with a solar jet that produced an impulsive microwave burst event recorded by the Expanded Owens Valley Solar Array (EOVSA). The type III burst event also coincided with the detection of enhanced in situ energetic electrons recorded by both PSP at 0.37 AU and WIND at 1 AU, which were located very closely on the Parker spiral longitudinally. The close timing association and magnetic connectivity suggest that the in situ energetic electrons originated from the jet's magnetic reconnection region. Intriguingly, microwave imaging spectroscopy results suggest that the escaping energetic electrons were injected into a large opening angle of about 90 degrees, which is at least nine times broader than the apparent width of the jet spire. Our findings provide an interpretation for the previously reported, longitudinally broad spatial distribution of flare locations associated with prompt energetic electron events and have important implications for understanding the origin and distribution of energetic electrons in the interplanetary space.

astro-ph.SR

Aegis: Mitigating Targeted Bit-flip Attacks against Deep Neural Networks

Bit-flip attacks (BFAs) have attracted substantial attention recently, in which an adversary could tamper with a small number of model parameter bits to break the integrity of DNNs. To mitigate such threats, a batch of defense methods are proposed, focusing on the untargeted scenarios. Unfortunately, they either require extra trustworthy applications or make models more vulnerable to targeted BFAs. Countermeasures against targeted BFAs, stealthier and more purposeful by nature, are far from well established. In this work, we propose Aegis, a novel defense method to mitigate targeted BFAs. The core observation is that existing targeted attacks focus on flipping critical bits in certain important layers. Thus, we design a dynamic-exit mechanism to attach extra internal classifiers (ICs) to hidden layers. This mechanism enables input samples to early-exit from different layers, which effectively upsets the adversary's attack plans. Moreover, the dynamic-exit mechanism randomly selects ICs for predictions during each inference to significantly increase the attack cost for the adaptive attacks where all defense mechanisms are transparent to the adversary. We further propose a robustness training strategy to adapt ICs to the attack scenarios by simulating BFAs during the IC training phase, to increase model robustness. Extensive evaluations over four well-known datasets and two popular DNN structures reveal that Aegis could effectively mitigate different state-of-the-art targeted attacks, reducing attack success rate by 5-10$\times$, significantly outperforming existing defense methods.

cs.CR

Elbert: Fast Albert with Confidence-Window Based Early Exit

Despite the great success in Natural Language Processing (NLP) area, large pre-trained language models like BERT are not well-suited for resource-constrained or real-time applications owing to the large number of parameters and slow inference speed. Recently, compressing and accelerating BERT have become important topics. By incorporating a parameter-sharing strategy, ALBERT greatly reduces the number of parameters while achieving competitive performance. Nevertheless, ALBERT still suffers from a long inference time. In this work, we propose the ELBERT, which significantly improves the average inference speed compared to ALBERT due to the proposed confidence-window based early exit mechanism, without introducing additional parameters or extra training overhead. Experimental results show that ELBERT achieves an adaptive inference speedup varying from 2$\times$ to 10$\times$ with negligible accuracy degradation compared to ALBERT on various datasets. Besides, ELBERT achieves higher accuracy than existing early exit methods used for accelerating BERT under the same computation cost. Furthermore, to understand the principle of the early exit mechanism, we also visualize the decision-making process of it in ELBERT.

cs.CL

Transform-Based Feature Map Compression for CNN Inference

To achieve higher accuracy in machine learning tasks, very deep convolutional neural networks (CNNs) are designed recently. However, the large memory access of deep CNNs will lead to high power consumption. A variety of hardware-friendly compression methods have been proposed to reduce the data transfer bandwidth by exploiting the sparsity of feature maps. Most of them focus on designing a specialized encoding format to increase the compression ratio. Differently, we observe and exploit the sparsity distinction between activations in earlier and later layers to improve the compression ratio. We propose a novel hardware-friendly transform-based method named 1D-Discrete Cosine Transform on Channel dimension with Masks (DCT-CM), which intelligently combines DCT, masks, and a coding format to compress activations. The proposed algorithm achieves an average compression ratio of 2.9x (53% higher than the state-of-the-art transform-based feature map compression works) during inference on ResNet-50 with an 8-bit quantization scheme.

eess.IV

Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformer

Designing hardware accelerators for deep neural networks (DNNs) has been much desired. Nonetheless, most of these existing accelerators are built for either convolutional neural networks (CNNs) or recurrent neural networks (RNNs). Recently, the Transformer model is replacing the RNN in the natural language processing (NLP) area. However, because of intensive matrix computations and complicated data flow being involved, the hardware design for the Transformer model has never been reported. In this paper, we propose the first hardware accelerator for two key components, i.e., the multi-head attention (MHA) ResBlock and the position-wise feed-forward network (FFN) ResBlock, which are the two most complex layers in the Transformer. Firstly, an efficient method is introduced to partition the huge matrices in the Transformer, allowing the two ResBlocks to share most of the hardware resources. Secondly, the computation flow is well designed to ensure the high hardware utilization of the systolic array, which is the biggest module in our design. Thirdly, complicated nonlinear functions are highly optimized to further reduce the hardware complexity and also the latency of the entire system. Our design is coded using hardware description language (HDL) and evaluated on a Xilinx FPGA. Compared with the implementation on GPU with the same setting, the proposed design demonstrates a speed-up of 14.6x in the MHA ResBlock, and 3.4x in the FFN ResBlock, respectively. Therefore, this work lays a good foundation for building efficient hardware accelerators for multiple Transformer networks.

eess.SP