Searcharxiv⌕ Search

arXiv subjects

Yan Sun

Publications and source records attributed to Yan Sun.

At least 55 records · Page 3Linked to original sources

Distilling and Adapting: A Topology-Aware Framework for Zero-Shot Interaction Prediction in Multiplex Biological Networks

Multiplex Biological Networks (MBNs), which represent multiple interaction types between entities, are crucial for understanding complex biological systems. Yet, existing methods often inadequately model multiplexity, struggle to integrate structural and sequence information, and face difficulties in zero-shot prediction for unseen entities with no prior neighbourhood information. To address these limitations, we propose a novel framework for zero-shot interaction prediction in MBNs by leveraging context-aware representation learning and knowledge distillation. Our approach leverages domain-specific foundation models to generate enriched embeddings, introduces a topology-aware graph tokenizer to capture multiplexity and higher-order connectivity, and employs contrastive learning to align embeddings across modalities. A teacher-student distillation strategy further enables robust zero-shot generalization. Experimental results demonstrate that our framework outperforms state-of-the-art methods in interaction prediction for MBNs, providing a powerful tool for exploring various biological interactions and advancing personalized therapeutics.

cs.LG↗

Efficient photocatalytic CO2 Reduction to C2+ Products with Pt1-xPdxSn4 Dirac Nodal Arc Semimetal

The photochemical CO2 reduction reaction represents a zero-carbon pathway for converting CO2 into value-added chemicals, yet its industrial implementation has been constrained by low selectivity and product diversity. Dirac nodal arc semimetals characterized by ultrahigh carrier mobility with over 25000 cm2 V-1 s-1 offer a promising platform to search for efficient catalysts for CO2 conversion. Herein, we demonstrate that strategic Pt incorporation into PdSn4 optimizes the electronic structure and carrier dynamics of this Dirac semimetal. Experimental and theoretical analyses reveal that the resulting Pd-Sn-Pt local electronic structure redistributes charge density around Pd and Pt atoms, which facilitates C-C coupling via *OC-COH and *OC-CHOH intermediates and enhances carrier mobility by 40% versus the pristine PdSn4 single crystal. The optimized Pd0.4Pt0.6Sn4 single crystal achieves C2H4 with formation rate of 0.000328 mol g-1 h-1, product selectivity of 73.1% and electron-based selectivity of 89%. This work establishes electronic-structure-tunable Dirac semimetals as a new paradigm for multi-carbon photochemical CO2 reduction, providing a design strategy for next-generation photocatalysts.

cond-mat.mtrl-sci↗

Efficient Reinforcement Learning for Large Language Models with Intrinsic Exploration

Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning ability of large language models, yet training remains costly because many rollouts contribute little to optimization, considering the amount of computation required. This study investigates how simply leveraging intrinsic data properties, almost free benefit during training, can improve data efficiency for RLVR. We propose PREPO with two complementary components. First, we adopt prompt perplexity as an indicator of model adaptability in learning, enabling the model to progress from well-understood contexts to more challenging ones. Second, we amplify the discrepancy among the rollouts by differentiating their relative entropy, and prioritize sequences that exhibit a higher degree of exploration. Together, these mechanisms reduce rollout demand while preserving competitive performance. On the Qwen and Llama models, PREPO achieves effective results on mathematical reasoning benchmarks with up to 3 times fewer rollouts than the baselines. Beyond empirical gains, we provide theoretical and in-depth analyses explaining the underlying rationale of our method to improve the data efficiency of RLVR.

cs.LG↗

The distinction of time-reversal-like degeneracy by electronic transport in a new compound

We report the discovery of a new compound, Ce$_3$MgBi$_5$, and reveal the hidden time-reversal-like degenerate states within it. Ce$_3$MgBi$_5$ is an antiferromagnet with the distorted kagome lattice of Ce atoms, in which several fractional magnetization plateaus emerge with the increase of magnetic field. At the 1/2 magnetization plateau, obvious hysteresis has been observed in the magnetoresistance and Hall resistivity during the rise and fall of the magnetic field. However, hysteresis vanishes in the corresponding measurements of magnetization, indicating the existence of degenerate states with the same net magnetization but different electronic transport properties. The degenerate states can be connected by the time-reversal-like operation. In addition, by comparing with HoAgGe, it is suggested that the special crystal structure in Ce$_3$MgBi$_5$ may have a shielding effect on the time-reversal-like operation, thereby affecting the distinction of degenerate states. Our work establishes Ce$_3$MgBi$_5$ as an example of utilizing electronic transport properties to identify and distinguish hidden symmetries in frustrated magnetic systems.

cond-mat.str-el↗

CATS: Enhancing Multivariate Time Series Forecasting by Constructing Auxiliary Time Series as Exogenous Variables

For Multivariate Time Series Forecasting (MTSF), recent deep learning applications show that univariate models frequently outperform multivariate ones. To address the difficiency in multivariate models, we introduce a method to Construct Auxiliary Time Series (CATS) that functions like a 2D temporal-contextual attention mechanism, which generates Auxiliary Time Series (ATS) from Original Time Series (OTS) to effectively represent and incorporate inter-series relationships for forecasting. Key principles of ATS - continuity, sparsity, and variability - are identified and implemented through different modules. Even with a basic 2-layer MLP as core predictor, CATS achieves state-of-the-art, significantly reducing complexity and parameters compared to previous multivariate models, marking it an efficient and transferable MTSF solution.

stat.ML↗

In-context Time Series Predictor

Recent Transformer-based large language models (LLMs) demonstrate in-context learning ability to perform various functions based solely on the provided context, without updating model parameters. To fully utilize the in-context capabilities in time series forecasting (TSF) problems, unlike previous Transformer-based or LLM-based time series forecasting methods, we reformulate "time series forecasting tasks" as input tokens by constructing a series of (lookback, future) pairs within the tokens. This method aligns more closely with the inherent in-context mechanisms, and is more parameter-efficient without the need of using pre-trained LLM parameters. Furthermore, it addresses issues such as overfitting in existing Transformer-based TSF models, consistently achieving better performance across full-data, few-shot, and zero-shot settings compared to previous architectures.

cs.LG↗

WAVE: Weighted Autoregressive Varying Gate for Time Series Forecasting

We propose a Weighted Autoregressive Varying gatE (WAVE) attention mechanism equipped with both Autoregressive (AR) and Moving-average (MA) components. It can adapt to various attention mechanisms, enhancing and decoupling their ability to capture long-range and local temporal patterns in time series data. In this paper, we first demonstrate that, for the time series forecasting (TSF) task, the previously overlooked decoder-only autoregressive Transformer model can achieve results comparable to the best baselines when appropriate tokenization and training methods are applied. Moreover, inspired by the ARMA model from statistics and recent advances in linear attention, we introduce the full ARMA structure into existing autoregressive attention mechanisms. By using an indirect MA weight generation method, we incorporate the MA term while maintaining the time complexity and parameter size of the underlying efficient attention models. We further explore how indirect parameter generation can produce implicit MA weights that align with the modeling requirements for local temporal impacts. Experimental results show that WAVE attention that incorporates the ARMA structure consistently improves the performance of various AR attentions on TSF tasks, achieving state-of-the-art results.

cs.LG↗

ZeroS: Zero-Sum Linear Attention for Efficient Transformers

Linear attention methods offer Transformers $O(N)$ complexity but typically underperform standard softmax attention. We identify two fundamental limitations affecting these approaches: the restriction to convex combinations that only permits additive information blending, and uniform accumulated weight bias that dilutes attention in long contexts. We propose Zero-Sum Linear Attention (ZeroS), which addresses these limitations by removing the constant zero-order term $1/t$ and reweighting the remaining zero-sum softmax residuals. This modification creates mathematically stable weights, enabling both positive and negative values and allowing a single attention layer to perform contrastive operations. While maintaining $O(N)$ complexity, ZeroS theoretically expands the set of representable functions compared to convex combinations. Empirically, it matches or exceeds standard softmax attention across various sequence modeling benchmarks.

cs.LG↗

Singleton-Optimized Conformal Prediction

Conformal prediction can be used to construct prediction sets that cover the true outcome with a desired probability, but can sometimes lead to large prediction sets that are costly in practice. The most useful outcome is a singleton prediction-an unambiguous decision-yet existing efficiency-oriented methods primarily optimize average set size. Motivated by this, we propose a new nonconformity score that aims to minimize the probability of producing non-singleton sets. Starting from a non-convex constrained optimization problem as a motivation, we provide a geometric reformulation and associated algorithm for computing the nonconformity score and associated split conformal prediction sets in O(K) time for K-class problems. Using this score in split conformal prediction leads to our proposed Singleton-Optimized Conformal Prediction (SOCOP) method. We evaluate our method in experiments on image classification and LLM multiple-choice question-answering, comparing with standard nonconformity scores such as the (negative) label probability estimates and their cumulative distribution function; both of which are motivated by optimizing length. The results show that SOCOP increases singleton frequency (sometimes by over 20%) compared to the above scores, with minimal impact on average set size.

stat.ML↗

How is Cold Gas Loaded into Galactic Nuclear Outflows?

The origin of the multiphase gas within the Fermi/eROSITA bubbles is crucial for understanding Galactic Center (GC) feedback. We use HI4PI data to investigate the kinematics and physical properties of high-velocity clouds (HVCs) toward the GC. Our results reveal that the HVCs exhibit a distinct asymmetric distribution, closely associated with the bar-driven tilted dust lanes and the distorted overshooting streams. We propose that powerful nuclear outflows interact with these gas-rich, off-plane structures, striping and entraining cold gas from the outer Galactic regions (R_GC~0.5--1.7 kpc) rather than solely from the region of the central molecular zone (CMZ; R_GC<0.3 kpc). In this scenario, as the Galactic bar drives gas inflows along the dust lanes, nuclear outflows simultaneously break through the CMZ, sweeping up and ablating cold gas from the boundary layer of these pre-existing structures. This process naturally accounts for the observed high turbulence, complex spectral signatures, and anomalous spatial-kinematic gas patterns, as well as multiwavelength asymmetries of the bubbles. The HVCs are accelerated to about 230--340 km/s over a dynamical time of ~3--6 Myr. When the multiphase, inhomogeneous composition of the gas is included, the estimated gas outflow rate reaches ~1 Msun/yr. This value is comparable to the bar-driven inflow rate, indicating a tightly coupled gas cycle in the inner Galaxy. Our research highlights the critical role of bar-driven gas dynamics and nuclear feedback in the secular evolution of the Milky Way, offering a valuable paradigm for investigating gas cycles in external galaxies.

astro-ph.GA↗

Complex spin dynamics induced metamagnetic phase transitions in Dirac semimetal EuAuBi

We report a comprehensive investigation of the physical properties of the Dirac semimetal compound EuAuBi single crystals, using neutron diffraction, magnetization, electrical transport, and specific heat measurements. EuAuBi crystallizes in a hexagonal structure with space group P63mc (No. 186). First-principles calculations using density functional theory characterize it as a Dirac semimetal, with a notable band-crossing in proximity to the Fermi level (EF ) along the Γ-A direction. The crystal exhibits three distinct magnetic phases at 4 K (TN1), 3.5 K (TN2), and 2.8 K (TN3)as observed from magnetic and specific heat measurements. However, zero-field neutron diffraction resolves only two magnetic phases: a commensurate antiferromagnetic phase and a canted antiferromagnetic phase. Field-dependent ac and dc magnetization measurements uncover field-induced non-trivial spin textures in the magnetic field range 1.5 to 3 T, manifested as a tilted plateau in the magnetization curves. The interplay between conduction carriers and these spin textures is further evidenced by unique features in the magnetic field-dependent longitudinal resistivity in the system. Finally, we present a comprehensive magnetic phase diagram of EuAuBi, highlighting diverse spin alignments present in the material. EuAuBi thus emerges as a rare material system in which both momentum-space and real-space Berry curvature effects may coexist, providing a unique opportunity to investigate their interplay.

cond-mat.mtrl-sci↗

IGAA: Intent-Driven General Agentic AI for Edge Services Scheduling using Generative Meta Learning

Agentic AI (AAI), which extends Large Language Models with enhanced reasoning capabilities, has emerged as a promising paradigm for autonomous edge service scheduling. However, user mobility creates highly dynamic service demands in edge networks, and existing service scheduling agents often lack generalization capabilities for new scenarios. Therefore, this paper proposes a novel Intent-Driven General Agentic AI (IGAA) framework. Leveraging a meta-learning paradigm, IGAA enables AAI to continuously learn from prior service scheduling experiences to achieve generalized scheduling capabilities. Particularly, IGAA incorporates three core mechanisms. First, we design a Network-Service-Intent matrix mapping method to allow agents to simulate novel scenarios and generate training datasets. Second, we present an easy-to-hard generalization learning scheme with two customized algorithms, namely Resource Causal Effect-aware Transfer Learning (RCETL) and Action Potential Optimality-aware Transfer Learning (APOTL). These algorithms help IGAA adapt to new scenarios. Furthermore, to prevent catastrophic forgetting during continual IGAA learning, we propose a Generative Intent Replay (GIR) mechanism that synthesizes historical service data to consolidate prior capabilities. Finally, to mitigate the effect of LLM hallucinations on scenario simulation, we incorporate a scenario evaluation and correction model to guide agents in generating rational scenarios and datasets. Extensive experiments demonstrate IGAA's strong generalization and scalability. Specifically, IGAA enables rapid adaptation by transferring learned policies to analogous new ones, such as applying latency-sensitive patterns from real-time computing to optimize novel Internet of Vehicles (IoV) services. Compared to scenario-specific methods, IGAA maintains the intent-satisfaction rate gap within 3.81%.

cs.NI↗

Generative Intent Prediction Agentic AI empowered Edge Service Function Chain Orchestration

With the development of artificial intelligence (AI), Agentic AI (AAI) based on large language models (LLMs) is gradually being applied to network management. However, in edge network environments, high user mobility and implicit service intents pose significant challenges to the passive and reactive management of traditional AAI. To address the limitations of existing approaches in handling dynamic demands and predicting users' implicit intents, in this paper we propose an edge service function chain (SFC) orchestration framework empowered by a Generative Intent Prediction Agent (GIPA). Our GIPA aims to shift the paradigm from passive execution to proactive prediction and orchestration. First, we construct a multidimensional intent space that includes functional preferences, QoS sensitivity, and resource requirements, enabling the mapping from unstructured natural language to quantifiable physical resource demands. Second, to cope with the complexity and randomness of intent sequences, we design an intent prediction model based on a Generative Diffusion Model (GDM), which reconstructs users' implicit intents from multidimensional context through a reverse denoising process. Finally, the predicted implicit intents are embedded as global prompts into the SFC orchestration model to guide the network in proactively and ahead-of-time optimizing SFC deployment strategies. Experiment results show that GIPA outperforms existing baseline methods in highly concurrent and highly dynamic scenarios.

cs.NI↗

Large room temperature anomalous Nernst effect coupled with topological Nernst effect from incommensurate spin structure in a Kagome antiferromagnet

Kagome magnets exhibit a range of novel and nontrivial topological properties due to the strong interplay between topology and magnetism, which also extends to their thermoelectric applications. Recent advances in the study of magnetic topological materials have highlighted their intriguing anomalous Hall and thermoelectric effects, arising primarily from large intrinsic Berry curvature. Here, we report observation of a large room-temperature (RT) anomalous Nernst effects (ANE) of S_xy^A ~ 1.3 μV K^(-1) in the kagome antiferromagnet (AFM) ErMn6Sn6, which is comparable to the largest signals observed in known magnetic materials. Surprisingly, we further found that a significant topological Nernst signal at RT and peaking a maximum of approximately 0.2 μV K^(-1) at 180 K, exactly coupling with ANE in the spiral AFM state, originates from the real-space nonzero spin chirality caused by incommensurate spin structure. This study demonstrates a potential room-temperature thermoelectric application platform based on Nernst effect, and provides insights for discovering significant anomalous and topological transverse transport effects in the incommensurate AFM system.

cond-mat.str-el↗

RED-F: Reconstruction-Elimination based Dual-stream Contrastive Forecasting for Multivariate Time Series Anomaly Prediction

Anomaly prediction (AP) in multivariate time series (MTS) is crucial to ensure system dependability. Existing methods either focus solely on whether an anomaly is imminent without providing precise predictions for the future anomaly, or performing predictions directly on historical data, which is easily drowned out by the normal patterns. To address the challenges in AP task, we propose RED-F, a novel framework comprised of the Reconstruction-Elimination Model (REM) and the Dual-stream Contrastive Forecasting Model (DFM). We utilize REM to construct a baseline of normal patterns from historical data, providing a foundation for subsequent predictions of anomalies. Then DFM simultaneously predicts both the constructed normal pattern and the current window, employing a contrastive forecast that transforms the difficult AP task into a simpler, more robust task of relative trajectory comparison by computing the divergence between these two predictions. To enable the forecasting model to generate a prediction not easily obscured by normal patterns, we propose a Multi-Series Prediction (MSP) training objective to enhance its sensitivity to the current window. Extensive experiments on multiple real-world datasets demonstrate the superior capability of RED-F in anomaly prediction tasks. Our code is available at http://github.com/PenyChen/RED-F.

cs.LG↗

Talk Less, Verify More: Improving LLM Assistants with Semantic Checks and Execution Feedback

As large language model (LLM) assistants become increasingly integrated into enterprise workflows, their ability to generate accurate, semantically aligned, and executable outputs is critical. However, current conversational business analytics (CBA) systems often lack built-in verification mechanisms, leaving users to manually validate potentially flawed results. This paper introduces two complementary verification techniques: Q*, which performs reverse translation and semantic matching between code and user intent, and Feedback+, which incorporates execution feedback to guide code refinement. Embedded within a generator-discriminator framework, these mechanisms shift validation responsibilities from users to the system. Evaluations on three benchmark datasets, Spider, Bird, and GSM8K, demonstrate that both Q* and Feedback+ reduce error rates and task completion time. The study also identifies reverse translation as a key bottleneck, highlighting opportunities for future improvement. Overall, this work contributes a design-oriented framework for building more reliable, enterprise-grade GenAI systems capable of trustworthy decision support.

cs.CL↗

MultiRisk: Multiple Risk Control via Iterative Score Thresholding

As generative AI systems are increasingly deployed in real-world applications, regulating multiple dimensions of model behavior has become essential. We focus on test-time filtering: a lightweight mechanism for behavior control that compares performance scores to estimated thresholds, and modifies outputs when these bounds are violated. We formalize the problem of enforcing multiple risk constraints with user-defined priorities, and introduce two efficient dynamic programming algorithms that leverage this sequential structure. The first, MULTIRISK-BASE, provides a direct finite-sample procedure for selecting thresholds, while the second, MULTIRISK, leverages data exchangeability to guarantee simultaneous control of the risks. Under mild assumptions, we show that MULTIRISK achieves nearly tight control of all constraint risks. The analysis requires an intricate iterative argument, upper bounding the risks by introducing several forms of intermediate symmetrized risk functions, and carefully lower bounding the risks by recursively counting jumps in symmetrized risk functions between appropriate risk levels. We evaluate our framework on a three-constraint Large Language Model alignment task using the PKU-SafeRLHF dataset, where the goal is to maximize helpfulness subject to multiple safety constraints, and where scores are generated by a Large Language Model judge and a perplexity filter. Our experimental results show that our algorithm can control each individual risk at close to the target level.

stat.ML↗

Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning

A fine-grained data recipe is crucial for pre-training large language models, as it can significantly enhance training efficiency and model performance. One important ingredient in the recipe is to select samples based on scores produced by defined rules, LLM judgment, or statistical information in embeddings, which can be roughly categorized into quality and diversity metrics. Due to the high computational cost when applied to trillion-scale token pre-training datasets such as FineWeb and DCLM, these two or more types of metrics are rarely considered jointly in a single selection process. However, in our empirical study, selecting samples based on quality metrics exhibit severe diminishing returns during long-term pre-training, while selecting on diversity metrics removes too many valuable high-quality samples, both of which limit pre-trained LLMs' capabilities. Therefore, we introduce DATAMASK, a novel and efficient joint learning framework designed for large-scale pre-training data selection that can simultaneously optimize multiple types of metrics in a unified process, with this study focusing specifically on quality and diversity metrics. DATAMASK approaches the selection process as a mask learning problem, involving iterative sampling of data masks, computation of policy gradients based on predefined objectives with sampled masks, and updating of mask sampling logits. Through policy gradient-based optimization and various acceleration enhancements, it significantly reduces selection time by 98.9% compared to greedy algorithm, enabling our study to explore joint learning within trillion-scale tokens. With DATAMASK, we select a subset of about 10% from the 15 trillion-token FineWeb dataset, termed FineWeb-Mask. Evaluated across 12 diverse tasks, we achieves significant improvements of 3.2% on a 1.5B dense model and 1.9% on a 7B MoE model.

cs.CL↗