SearcharxivSearch

arXiv subjects

Chengrui Zhou

Publications and source records attributed to Chengrui Zhou.

At least 19 recordsLinked to original sources

Successive Coronal Jets as Novel Facilitators for Filament Oscillation and Eruption

Solar filament eruptions are central to coronal mass ejections and space weather, yet their triggering mechanisms remain a fundamental open question. In particular, the early-stage that drives a magnetic flux rope toward instability and its observable signatures are poorly understood. Here, combining multi-instrument observations, we report successive coronal jets impacting a filament, causing its gradual rise and oscillations with growing amplitude and period. When the filament reaches the height where the decay index exceeds the torus instability threshold, the rapid filament eruption commences. This filament eruption is reproduced by magnetohydrodynamic simulations, in which successive thermal jets disturb a stable filament in a magnetic flux rope and excite oscillations together with the eruption of the filament. As the filament rises to erupt, the restoring forces for the oscillation progressively weaken, which naturally leads to an increase of the oscillation amplitude and period. Our results demonstrate the growing oscillations as one of the observable precursors for filament eruptions, enhancing our ability to predict solar eruptions.

astro-ph.SR

Cognitive Alignment Deciphered: A Self-Developed Scenario-Based Prompt Scale Coupled with Representational Similarity Analysis and Social Network Analysis for Unraveling Bias Mechanisms Across Humans and LLMs

Traditional cognitive bias measurement tools are limited by narrow bias coverage, low ecological validity, and reliance on abstract self reports, constraining scenario based and human AI comparisons. We introduce the context based Cognitive Bias Assessment Scale CBAS, a scenario driven prompt template covering 58 cognitive biases across five hot cold dual system dimensions: Calculation, Belief, Information, Social, and Memory. Psychometric testing with 330 participants shows satisfactory reliability Cronbachs alpha 0.714 and good model fit chi squared df 1.83, RMSEA 0.057, CFI 0.908, TLI 0.903. We then combine Representational Similarity Analysis RSA and Social Network Analysis SNA to compare human age groups and three large language models Baidu ERNIE 3.5 8K, DeepSeek V3, DeepSeek R1. Humans show coherent hot cold integration with high inter individual variability, whereas LLMs display fragmented, inflexible response patterns and lower variability. Human cognitive networks exhibit strong inter module connectivity, while LLMs show fixed core biases and isolated information processing components. Prompt interventions integrating role playing and bias mitigation instructions effectively improve LLM response accuracy, reaching 84.86 percent for DeepSeek R1 and 78.24 percent for DeepSeek V3, and partially reshape their internal representations. Our work establishes a replicable assessment and analysis pipeline for cognitive alignment research, bridging empirical psychological evaluation and interpretable artificial intelligence.

cs.HC

Efficient Cold-Start Recommendation via BPE Token-Level Embedding Initialization with LLM

The cold-start issue is the challenge when we talk about recommender systems, especially in the case when we do not have the past interaction data of new users or new items. Content-based features or hybrid solutions are common as conventional solutions, but they can only work in a sparse metadata environment with shallow patterns. In this paper, the efficient cold-start recommendation strategy is presented, which is based on the sub word-level representations by applying Byte Pair Encoding (BPE) tokenization and pre-trained Large Language Model (LLM) embedding in the initialization procedure. We obtain fine-grained token-level vectors that are aligned with the BPE vocabulary as opposed to using coarse-grained sentence embeddings. Together, these token embeddings can be used as dense semantic priors on unseen entities, making immediate recommendation performance possible without user-item interaction history. Our mechanism can be compared to collaborative filtering systems and tested over benchmark datasets with stringent cold-start assumptions. Experimental findings show that the given BPE-LLM method achieves higher Recall@k, NDCG@k, and Hit Rate measurements compared to the standard baseline and displays the same capability of sufficient computational performance. Furthermore, we demonstrate that using subword-aware embeddings yields better generalizability and is more interpretable, especially within a multilingual and sparse input setting. The practical application of token-level semantic initialization as a lightweight, but nevertheless effective extension to modern recommender systems in the zero-shot setting is indicated within this work.

cs.IR

RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders

Conversational recommender systems (CRS) based on Large Language Models (LLMs) need to constantly be aligned to the user preferences to provide satisfying and context-relevant item recommendations. The traditional supervised fine-tuning cannot capture the implicit feedback signal, e.g., dwell time, sentiment polarity, or engagement patterns. In this paper, we share a fine-tuning solution using human feedback reinforcement learning (RLHF) to maximize implied user feedback (IUF) in a multi-turn recommendation context. We specify a reward model $R_{\phi}$ learnt on weakly-labelled engagement information and maximize user-centric utility by optimizing the foundational LLM M_{\theta} through a proximal policy optimization (PPO) approach. The architecture models conversational state transitions $s_t \to a_t \to s_{t +1}$, where the action $a_t$ is associated with LLM-generated item suggestions only on condition of conversation history in the past. The evaluation across synthetic and real-world datasets (e.g.REDIAL, OpenDialKG) demonstrates that our RLHF-fine-tuned models can perform better in terms of top-$k$ recommendation accuracy, coherence, and user satisfaction compared to (arrow-zero-cmwrquca-teja-falset ensuite 2Round group-deca States penalty give up This paper shows that implicit signal alignment can be efficient in achieving scalable and user-adaptive design of CRS.

cs.LG

Meta-Learning for Cold-Start Personalization in Prompt-Tuned LLMs

Generative, explainable, and flexible recommender systems, derived using Large Language Models (LLM) are promising and poorly adapted to the cold-start user situation, where there is little to no history of interaction. The current solutions i.e. supervised fine-tuning and collaborative filtering are dense-user-item focused and would be expensive to maintain and update. This paper introduces a meta-learning framework, that can be used to perform parameter-efficient prompt-tuning, to effectively personalize LLM-based recommender systems quickly at cold-start. The model learns soft prompt embeddings with first-order (Reptile) and second-order (MAML) optimization by treating each of the users as the tasks. As augmentations to the input tokens, these learnable vectors are the differentiable control variables that represent user behavioral priors. The prompts are meta-optimized through episodic sampling, inner-loop adaptation, and outer-loop generalization. On MovieLens-1M, Amazon Reviews, and Recbole, we can see that our adaptive model outperforms strong baselines in NDCG@10, HR@10, and MRR, and it runs in real-time (i.e., below 300 ms) on consumer GPUs. Zero-history personalization is also supported by this scalable solution, and its 275 ms rate of adaptation allows successful real-time risk profiling of financial systems by shortening detection latency and improving payment network stability. Crucially, the 275 ms adaptation capability can enable real-time risk profiling for financial institutions, reducing systemic vulnerability detection latency significantly versus traditional compliance checks. By preventing contagion in payment networks (e.g., Fedwire), the framework strengthens national financial infrastructure resilience.

cs.LG

From Photospheric Footpoint Motion to Plasmoid Ejection: A Two-Stage Reconnection Process in a Small-scale Chromospheric Jet

Using high spatiotemporal resolution, multi-wavelength observations from the New Vacuum Solar Telescope (NVST) and the Solar Dynamics Observatory (SDO), we present a detailed analysis of a small-scale chromospheric jet driven by plasmoid-mediated magnetic reconnection. Our results reveal that the entire process is governed by the dynamic evolution of photospheric magnetic footpoints, which proceeds in two distinct stages. An initial separating motion of the footpoints corresponds to a mild reconnection phase, characterized by a short current sheet and the eruption of a cool H$\alpha$ jet. Subsequently, a converging motion of the footpoints triggers an intense reconnection phase. During this intense stage, the current sheet rapidly elongates, and the resulting decrease in its aspect ratio initiates a tearing-mode instability, forming a plasmoid. The appearance of this plasmoid mediates the onset of fast magnetic reconnection, which produces a hot EUV jet and is concurrent with significant magnetic flux cancellation. We interpret this cancellation as the submergence of newly formed, post-reconnection loops. Furthermore, we identify a distinct, high-temperature plasma blob in the jet spire, significantly hotter than the surrounding jet plasma. We attribute this feature to a secondary heating process, likely caused by reconnection between the upward-propagating plasmoid and the overlying magnetic cusp structure. These observations provide a comprehensive, observationally driven picture (from the initial photospheric triggers to the multi-stage, plasmoid-mediated reconnection) that forms chromospheric jets, highlighting the critical role of footpoint motions in solar atmospheric dynamics.

astro-ph.SR

Research on Model Parallelism and Data Parallelism Optimization Methods in Large Language Model-Based Recommendation Systems

With the rapid adoption of large language models (LLMs) in recommendation systems, the computational and communication bottlenecks caused by their massive parameter sizes and large data volumes have become increasingly prominent. This paper systematically investigates two classes of optimization methods-model parallelism and data parallelism-for distributed training of LLMs in recommendation scenarios. For model parallelism, we implement both tensor parallelism and pipeline parallelism, and introduce an adaptive load-balancing mechanism to reduce cross-device communication overhead. For data parallelism, we compare synchronous and asynchronous modes, combining gradient compression and sparsification techniques with an efficient aggregation communication framework to significantly improve bandwidth utilization. Experiments conducted on a real-world recommendation dataset in a simulated service environment demonstrate that our proposed hybrid parallelism scheme increases training throughput by over 30% and improves resource utilization by approximately 20% compared to traditional single-mode parallelism, while maintaining strong scalability and robustness. Finally, we discuss trade-offs among different parallel strategies in online deployment and outline future directions involving heterogeneous hardware integration and automated scheduling technologies.

cs.DC

Deep Learning Model Acceleration and Optimization Strategies for Real-Time Recommendation Systems

With the rapid growth of Internet services, recommendation systems play a central role in delivering personalized content. Faced with massive user requests and complex model architectures, the key challenge for real-time recommendation systems is how to reduce inference latency and increase system throughput without sacrificing recommendation quality. This paper addresses the high computational cost and resource bottlenecks of deep learning models in real-time settings by proposing a combined set of modeling- and system-level acceleration and optimization strategies. At the model level, we dramatically reduce parameter counts and compute requirements through lightweight network design, structured pruning, and weight quantization. At the system level, we integrate multiple heterogeneous compute platforms and high-performance inference libraries, and we design elastic inference scheduling and load-balancing mechanisms based on real-time load characteristics. Experiments show that, while maintaining the original recommendation accuracy, our methods cut latency to less than 30% of the baseline and more than double system throughput, offering a practical solution for deploying large-scale online recommendation services.

cs.IR

Research on Personalized Financial Product Recommendation by Integrating Large Language Models and Graph Neural Networks

With the rapid growth of fintech, personalized financial product recommendations have become increasingly important. Traditional methods like collaborative filtering or content-based models often fail to capture users' latent preferences and complex relationships. We propose a hybrid framework integrating large language models (LLMs) and graph neural networks (GNNs). A pre-trained LLM encodes text data (e.g., user reviews) into rich feature vectors, while a heterogeneous user-product graph models interactions and social ties. Through a tailored message-passing mechanism, text and graph information are fused within the GNN to jointly optimize embeddings. Experiments on public and real-world financial datasets show our model outperforms standalone LLM or GNN in accuracy, recall, and NDCG, with strong interpretability. This work offers new insights for personalized financial recommendations and cross-modal fusion in broader recommendation tasks.

cs.IR

Research on E-Commerce Long-Tail Product Recommendation Mechanism Based on Large-Scale Language Models

As e-commerce platforms expand their product catalogs, accurately recommending long-tail items becomes increasingly important for enhancing both user experience and platform revenue. A key challenge is the long-tail problem, where extreme data sparsity and cold-start issues limit the performance of traditional recommendation methods. To address this, we propose a novel long-tail product recommendation mechanism that integrates product text descriptions and user behavior sequences using a large-scale language model (LLM). First, we introduce a semantic visor, which leverages a pre-trained LLM to convert multimodal textual content such as product titles, descriptions, and user reviews into meaningful embeddings. These embeddings help represent item-level semantics effectively. We then employ an attention-based user intent encoder that captures users' latent interests, especially toward long-tail items, by modeling collaborative behavior patterns. These components feed into a hybrid ranking model that fuses semantic similarity scores, collaborative filtering outputs, and LLM-generated recommendation candidates. Extensive experiments on a real-world e-commerce dataset show that our method outperforms baseline models in recall (+12%), hit rate (+9%), and user coverage (+15%). These improvements lead to better exposure and purchase rates for long-tail products. Our work highlights the potential of LLMs in interpreting product content and user intent, offering a promising direction for future e-commerce recommendation systems.

cs.IR

How Reconnection-Unfavored Magnetic Flux Emergence Suppresses Solar Filament Eruptions

Magnetic flux emergence is traditionally considered a key trigger of solar filament eruptions; however, its role in suppressing filament eruptions remains less understood. Using multi-wavelength observations from the Solar Dynamics Observatory, this study investigates a unique case of flux emergence below a quiescent filament from January 3 to 5, 2016, where the newly emerging magnetic flux suppressed rather than promoted the eruption of the filament. It is found that the emerging magnetic bipole within the filament channel directly interacted and reconnected with the overlying filament magnetic field and produced a series of two-sided coronal jets along the filament axis. Instead of eruption, the filament kept stable but broke into two segments at the reconnection site. Further magnetic cancellation or recession of the emerged bipole allowed the filament to recover its original structure. Our analysis results revealed that the flux emergence suppressed the filament eruption by reducing the upward net force. The formation and evolution of filament fine structures, such as filament threads, are closely linked to the reconnection processes between the emerging bipole and the horizontal magnetic field of the filament. This study provides direct observational evidence for the stabilization of solar filaments driven by flux emergence, offering new insights into the dual role of magnetic emergence in triggering and suppressing solar eruptions.

astro-ph.SR

Deciphering the Formation and Dynamics of Double-decker Filament Through Component Magnetic Reconnection

The formation of double-decker filaments has long been an enigma in the field of solar physics. Using stereoscopic observations from the Solar Dynamics Observatory and the Solar Terrestrial Relations Observatory, we show that the double-decker filament formed on 2013 August 30 resulted from the splitting of a braided magnetic flux rope. The splitting was driven by component magnetic reconnection between intertwined field lines, triggered by the rotational motion in a part of one filament footpoint. This mechanism, inferred from observed small jets, brightenings, and bidirectional mass flows, differs from the previous conclusion attributing filament splitting to magnetic reconnection between the legs of confining magnetic field lines within or above the filament. The splitting speed might be modulated by the reconnection speed, as evidenced by the correspondence between the filament's slow and fast rising phases and the intermittent and violent brightening stages. Following the splitting, the upper branch of the double-decker filament erupted as a coronal mass ejection (CME), giving rise to a GOES soft X-ray M1.2 flare. In conclusion, our observations present a new formation mechanism for double-decker filaments, and the subsequent partial eruption is likely attributable to the torus instability of the background coronal magnetic field. Moreover, the detection of small jets within the filament provides new insights into the role of component magnetic reconnection in localized coronal heating processes.

astro-ph.SR

A New Formation Mechanism of Counterstreaming Mass Flows in Filaments and the Doppler Bullseye Pattern in Prominences

The eruption of solar prominences can eject substantial mass and magnetic field into interplanetary space and cause geomagnetic storms. However, various questions about prominences and their eruption mechanism remain unclear. In particular, what causes the intriguing Doppler bullseye pattern in prominences has not yet been solved, despite some preliminary studies proposing that they are probably associated with counterstreaming mass flows. Previous studies are mainly based on single-angle and short timescale observations, making it difficult to determine the physical origin of Doppler bullseye patterns in prominences. Here, taking advantage of stereoscopic observations taken by the Solar Dynamics Observatory and the Solar Terrestrial Relations Observatory and a three-dimensional numerical simulation, we investigate the origin of prominence Doppler bullseye pattern by tracing a long-lived transequatorial filament/prominence from July 23 to August 4, 2012. We find that repeated coronal jets at one end of the prominence can launch the Doppler bullseye pattern. It is evidenced in our observations and simulation that during the forward traveling of jet plasma along the helical magnetic field structure of the prominence, part of the ejecting plasma can not pass through the apex of the prominence due to the insufficient kinetic energy and therefore forms a backward-moving mass flow along the same or neighboring magnetic field lines. This process finally forms counterstreaming mass flows in on-disk filaments. When the on-disk filament rotates to the solar limb to be a prominence, the counterstreaming mass flows are naturally observed as a Doppler bullseye pattern.

astro-ph.SR

High-Resolution Observations of a Small-Scale Cancellation Nanoflare: Supporting Evidence for the Cancellation Nanoflare Model

An analytical cancellation nanoflare model has recently been established to show the fundamental role that ubiquitous small-scale cancellation nanoflares play in solar atmospheric heating. Although this model is well-supported by simulations, observational evidence is needed to deepen our understanding of cancellation nanoflares. We present observations of a small-scale cancellation nanoflare event, analyzing its magnetic topology evolution, triggers, and physical parameters. Using coordinated observations from Solar Dynamics Observatory and Goode Solar Telescope, we identify a photospheric flow-driven cancellation event with a flux cancellation rate of ~10^{15} Mx/s and a heating rate of 8.7 x 10^6 erg cm^{-2} s^{-1}. The event shows the characteristic transition from $\pi$-shaped to X-shaped magnetic configuration before forming a two arcsecs current sheet, closely matching model predictions. This event provides critical observational support for the cancellation nanoflare model and its role in solar atmospheric heating.

astro-ph.SR

Solving Situation Puzzles with Large Language Model and External Reformulation

In recent years, large language models (LLMs) have shown an impressive ability to perform arithmetic and symbolic reasoning tasks. However, we found that LLMs (e.g., ChatGPT) cannot perform well on reasoning that requires multiple rounds of dialogue, especially when solving situation puzzles. Specifically, LLMs intend to ask very detailed questions focusing on a specific aspect or same/similar questions after several rounds of Q&As. To help LLMs get out of the above dilemma, we propose a novel external reformulation methodology, where the situation puzzle will be reformulated after several rounds of Q&A or when the LLMs raise an incorrect guess. Experiments show superior performance (e.g., win rate, number of question/guess attempts) of our method than directly using LLMs for solving situation puzzles, highlighting the potential of strategic problem reformulation to enhance the reasoning capabilities of LLMs in complex interactive scenarios.

cs.LG

Formation and Eruption of Hot Channel Magnetic Flux Rope in Nested Double Null Magnetic System

The coronal magnetic topology significantly affects the outcome of magnetic flux rope (MFR) eruptions. The recently reported nested double null magnetic system remains unclear as to how it affects MFR eruptions. Using observations from the New Vacuum Solar Telescope and the Solar Dynamics Observatory, we studied the formation and successful eruption of a hot channel MFR from NOAA active region AR12173 on 2014 September 28. We observed that a hot channel MFR formed and erupted as a coronal mass ejection (CME), and the magnetic field of the source region was a nested double null magnetic system in which an inner magnetic null point system was nested by an outer fan-spine magnetic system. Observational analysis suggests that the origin of the MFR was due to magnetic reconnection at the inner null point, which was triggered by the photospheric swirling motions. The long-term shearing motion in the source region throughout around 26 hours might accumulate enough energy to power the eruption. Since previous studies showed that MFR eruptions from nested double null magnetic systems often result in weak jets and stalled or failed eruptions, it is hard to understand the generation of the large-scale CME in our case. A detailed comparison with previous studies reveals that the birth location of the MFR relative to the inner null point might be the critical physical factor for determining whether an MFR can erupt successfully or not in such a particular nested double null magnetic system.

astro-ph.SR

On the origin of a broad QFP wave train: unwinding jet as the driver

Large-scale extreme-ultraviolet (EUV) waves commonly exhibit as single wavefront and are believed to be caused by coronal mass ejections (CMEs). Utilizing high spatiotemporal resolution imaging observations from the Solar Dynamics Observatory, we present two sequentially generated wave trains originating from the same active region: a narrow quasiperiodic fast-propagating (QFP) wave train that propagates along the coronal loop system above the jet and a broad QFP wave train that travels along the solar surface beneath the jet. The measurements indicate that the narrow QFP wave train and the accompanying flare's quasiperiodic pulsations (QPPs) have nearly identical onsets and periods. This result suggests that the accompanying flare process excites the observed narrow QFP wave train. However, the broad QFP wave train starts approximately 2 minutes before the QPPs of the flare, but consistent with the interaction between the unwinding jet and the solar surface. Moreover, we find that the \zx{period of the broad QFP wave train, approximately 130\,s, closely matches that of the unwinding jet}. This period is significantly longer than the 30\,s period of the accompanying flare's QPPs. Based on these findings, we propose that the intermittent energy release of the accompanying flare excited the narrow QFP wave train confined propagating in the coronal loop system. The unwinding jet, rather than the intermittent energy release in the accompanying flare, triggered the broad QFP wave train propagating along the solar surface.

astro-ph.SR

Broad and Bi-directional narrow quasi-periodic fast-propagating wave trains associated with a filament-driven halo CME on 2023 April 21

This paper presents three distinct wave trains that occurred on 2023 April 21: a broad quasi-periodic fast-propagating (QFP) wave train and a bi-directional narrow QFP wave train. The broad QFP wave train expands outward in a circular wavefront, while bi-directional narrow QFP wave trains propagate in the northward and southward directions, respectively. The concurrent presence of the wave trains offers a remarkable opportunity to investigate their respective triggering mechanisms. Measurement shows that the broad QFP wave train's speed is 300- 1100 km/s in different propagating directions. There is a significant difference in the speed of the bi-directional narrow QFP wave trains: the southward propagation achieves 1400 km/s, while the northward propagation only reaches about 550 km/s accompanied by a deceleration of about 1- 2 kms-2. Using the wavelet analysis, we find that the periodicity of the propagating wave trains in the southward and northward directions closely matches the quasi-periodic pulsations (QPPs) exhibited by the flares. Based on these results, the narrow QFP wave trains were most likely excited by the intermittent energy release in the accompanying flare. In contrast, the broad QFP wave train had a tight relationship with the erupting filament, probably attributed to the unwinding motion of the erupting filament or the leakage of the fast sausage wave train inside the filament body.

astro-ph.SR