SearcharxivSearch

arXiv subjects

Zefeng Chen

Publications and source records attributed to Zefeng Chen.

At least 19 recordsLinked to original sources

A Structure- and Pressure-Positivity-Preserving Semi-implicit IMEX Finite Volume Scheme for Ideal MHD at All Acoustic Mach and Alfv\'en Mach Numbers with Generic Equation of State

We present a conservative, structure-preserving, finite-volume scheme for ideal MHD that ensures pressure positivity, remains applicable across all Mach and Alfven regimes, and handles general nonlinear equations of state. The scheme splits the MHD system into three sub-systems according to characteristic wave scales: an advective part for hydrodynamic transport, a magnetic part for velocity-field coupling, and a pressure part for pressure-velocity coupling. Nonlinear advective terms are explicit, while the other two sub-systems are implicit, yielding a mild, velocity-based CFL condition that is supported by extensive numerical evidence. This makes the scheme suitable for gas-pressure or magnetic-pressure dominated regimes and the incompressible limit. The implicit discretisation gives a pressure equation that reduces to an elliptic form in the low-Mach limit for ideal gases and general thermodynamics. Pressure positivity is ensured via a local conservation-preserving modification of the pressure-internal-energy relation, avoiding a posteriori clipping while preserving conservation. The divergence-free constraint is enforced exactly via constrained transport. Second-order accuracy is achieved with an IMEX Runge-Kutta time integration, TVD reconstruction for explicit fluxes, and central discretisation for implicit terms. The scheme is validated against numerous benchmarks, including high- and low-Mach regimes, strongly magnetised flows, and standard MHD shock problems in 1D and 2D, demonstrating accuracy, stability, and excellent shock-capturing.

math.NA

Autonomous Market Intelligence: Agentic AI Nowcasting Predicts Stock Returns

Can fully agentic AI nowcast stock returns? We deploy a state-of-the-art Large Language Model to evaluate the attractiveness of each Russell 1000 stock daily, starting from April 2025 when AI web interfaces enabled real-time search. Our data contribution is unique along three dimensions. First, the nowcasting framework is completely out-of-sample and free of look-ahead bias by construction: predictions are collected at the current edge of time, ensuring the AI has no knowledge of future outcomes. Second, this temporal design is irreproducible -- once the information environment passes, it can never be recreated. Third, our framework is 100% agentic: we do not feed the model news, disclosures, or curated text; it autonomously searches the web, filters sources, and synthesises information into quantitative predictions. We find that AI possesses genuine stock selection ability, but only for identifying top winners. Longing the 20 highest-ranked stocks generates a daily Fama-French five-factor plus momentum alpha of 18.4 basis points and an annualised Sharpe ratio of 2.43. Critically, these returns derive from an implementable strategy trading highly liquid Russell 1000 constituents, with transaction costs representing less than 10\% of gross alpha. However, this predictability is highly concentrated: expanding beyond the top tier rapidly dilutes alpha, and bottom-ranked stocks exhibit returns statistically indistinguishable from the market. We hypothesise that this asymmetry reflects online information structure: genuinely positive news generates coherent signals, while negative news is contaminated by strategic corporate obfuscation and social media noise.

q-fin.GN

GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning

Autonomous agents powered by large language models (LLMs) have shown impressive capabilities in tool manipulation for complex task-solving. However, existing paradigms such as ReAct rely on sequential reasoning and execution, failing to exploit the inherent parallelism among independent sub-tasks. This sequential bottleneck leads to inefficient tool utilization and suboptimal performance in multi-step reasoning scenarios. We introduce Graph-based Agent Planning (GAP), a novel framework that explicitly models inter-task dependencies through graph-based planning to enable adaptive parallel and serial tool execution. Our approach trains agent foundation models to decompose complex tasks into dependency-aware sub-task graphs, autonomously determining which tools can be executed in parallel and which must follow sequential dependencies. This dependency-aware orchestration achieves substantial improvements in both execution efficiency and task accuracy. To train GAP, we construct a high-quality dataset of graph-based planning traces derived from the Multi-Hop Question Answering (MHQA) benchmark. We employ a two-stage training strategy: supervised fine-tuning (SFT) on the curated dataset, followed by reinforcement learning (RL) with a correctness-based reward function on strategically sampled queries where tool-based reasoning provides maximum value. Experimental results on MHQA datasets demonstrate that GAP significantly outperforms traditional ReAct baselines, particularly on multi-step retrieval tasks, while achieving dramatic improvements in tool invocation efficiency through intelligent parallelization. The project page is available at: https://github.com/WJQ7777/Graph-Agent-Planning.

cs.AI

Long-Range Spin-Orbit-Coupled Magnetoelectricity in Type-II Multiferroic NiI$_2$

Type-II multiferroics, where spin order induces ferroelectricity, exhibit strong magnetoelectric coupling. However, for the typical 2D type-II multiferroic NiI$_2$, the underlying magnetoelectric mechanism remains unclear. Here, applying generalized spin-current model, together with first-principles calculations and a tight-binding approach, we build a comprehensive magnetoelectric model for spin-induced polarization. Such model reveals that the spin-orbit coupling extends its influence to the third-nearest neighbors, whose contribution to polarization rivals that of the first-nearest neighbors. By analyzing the orbital-resolved contributions to polarization, our tight-binding model reveals that the long-range magnetoelectric coupling is enabled by the strong $e_g$-$p$ hopping of NiI$_2$. Monte Carlo simulations further predict a Bloch-type magnetic skyrmion lattice at moderate magnetic fields, accompanied by polar vortex arrays. These findings can guide the discovery and design of strongly magnetoelectric multiferroics.

cond-mat.mtrl-sci

SeqRFM: Fast RFM Analysis in Sequence Data

In recent years, data mining technologies have been well applied to many domains, including e-commerce. In customer relationship management (CRM), the RFM analysis model is one of the most effective approaches to increase the profits of major enterprises. However, with the rapid development of e-commerce, the diversity and abundance of e-commerce data pose a challenge to mining efficiency. Moreover, in actual market transactions, the chronological order of transactions reflects customer behavior and preferences. To address these challenges, we develop an effective algorithm called SeqRFM, which combines sequential pattern mining with RFM models. SeqRFM considers each customer's recency (R), frequency (F), and monetary (M) scores to represent the significance of the customer and identifies sequences with high recency, high frequency, and high monetary value. A series of experiments demonstrate the superiority and effectiveness of the SeqRFM algorithm compared to the most advanced RFM algorithms based on sequential pattern mining. The source code and datasets are available at GitHub https://github.com/DSI-Lab1/SeqRFM.

cs.DB

Targeted Mining Precise-positioning Episode Rules

The era characterized by an exponential increase in data has led to the widespread adoption of data intelligence as a crucial task. Within the field of data mining, frequent episode mining has emerged as an effective tool for extracting valuable and essential information from event sequences. Various algorithms have been developed to discover frequent episodes and subsequently derive episode rules using the frequency function and anti-monotonicity principles. However, currently, there is a lack of algorithms specifically designed for mining episode rules that encompass user-specified query episodes. To address this challenge and enable the mining of target episode rules, we introduce the definition of targeted precise-positioning episode rules and formulate the problem of targeted mining precise-positioning episode rules. Most importantly, we develop an algorithm called Targeted Mining Precision Episode Rules (TaMIPER) to address the problem and optimize it using four proposed strategies, leading to significant reductions in both time and space resource requirements. As a result, TaMIPER offers high accuracy and efficiency in mining episode rules of user interest and holds promising potential for prediction tasks in various domains, such as weather observation, network intrusion, and e-commerce. Experimental results on six real datasets demonstrate the exceptional performance of TaMIPER.

cs.DB

Effects of Kitaev Interaction on Magnetic Orders and Anisotropy

We systematically investigate the effects of Kitaev interaction on magnetic orders and anisotropy in both triangular and honeycomb lattices. Our study highlights the critical role of the Kitaev interaction in modulating phase boundaries and predicting new phases, e.g., zigzag phase in triangular lattice and AABB phase in honeycomb lattice, which are absent with pure Heisenberg interactions. Moreover, we reveal the special state-dependent anisotropy of Kitaev interaction, and develop a general method that can determine the presence of Kitaev interaction in different magnets. It is found that the Kitaev interaction does not induce anisotropy in some magnetic orders such as ferromagnetic order, while can cause different anisotropy in other magnetic orders. Furthermore, we emphasize that the off-diagonal $\Gamma$ interaction also contributes to anisotropy, competing with the Kitaev interaction to reorient spin arrangements. Our work establishes a framework for comprehensive understanding the impact of Kitaev interaction on ordered magnetism.

cond-mat.str-el

Strength of Kitaev Interaction in Na$_3$Co$_2$SbO$_6$ and Na$_3$Ni$_2$BiO$_6$

Kitaev spin liquid is proposed to be promisingly realized in low spin-orbit coupling $3d$ systems, represented by Na$_3$Co$_2$SbO$_6$ and Na$_3$Ni$_2$BiO$_6$. However, the existence of Kitaev interaction is still debatable among experiments, and obtaining the strength of Kitaev interaction from first-principles calculations is also challenging. Here, we report the state-dependent anisotropy of Kitaev interaction, based on which a convenient method is developed to rapidly determine the strength of Kitaev interaction. Applying such method and density functional theory calculations, it is found that Na$_3$Co$_2$SbO$_6$ with $3d^7$ configuration exhibits considerable ferromagnetic Kitaev interaction. Moreover, by further applying the symmetry-adapted cluster expansion method, a realistic spin model is determined for Na$_3$Ni$_2$BiO$_6$ with $3d^8$ configuration. Such model indicates negligible small Kitaev interaction, but it predicts many properties, such as ground states and field effects, which are well consistent with measurements. Furthermore, we demonstrate that the heavy elements, Sb or Bi, located at the hollow sites of honeycomb lattice, do not contribute to emergence of Kitaev interaction through proximity, contradictory to common belief. The presently developed anisotropy method will be beneficial not only for computations but also for measurements.

cond-mat.str-el

Large Language Models for Medicine: A Survey

To address challenges in the digital economy's landscape of digital intelligence, large language models (LLMs) have been developed. Improvements in computational power and available resources have significantly advanced LLMs, allowing their integration into diverse domains for human life. Medical LLMs are essential application tools with potential across various medical scenarios. In this paper, we review LLM developments, focusing on the requirements and applications of medical LLMs. We provide a concise overview of existing models, aiming to explore advanced research directions and benefit researchers for future medical applications. We emphasize the advantages of medical LLMs in applications, as well as the challenges encountered during their development. Finally, we suggest directions for technical integration to mitigate challenges and potential research directions for the future of medical LLMs, aiming to meet the demands of the medical field better.

cs.CL

Data Scarcity in Recommendation Systems: A Survey

The prevalence of online content has led to the widespread adoption of recommendation systems (RSs), which serve diverse purposes such as news, advertisements, and e-commerce recommendations. Despite their significance, data scarcity issues have significantly impaired the effectiveness of existing RS models and hindered their progress. To address this challenge, the concept of knowledge transfer, particularly from external sources like pre-trained language models, emerges as a potential solution to alleviate data scarcity and enhance RS development. However, the practice of knowledge transfer in RSs is intricate. Transferring knowledge between domains introduces data disparities, and the application of knowledge transfer in complex RS scenarios can yield negative consequences if not carefully designed. Therefore, this article contributes to this discourse by addressing the implications of data scarcity on RSs and introducing various strategies, such as data augmentation, self-supervised learning, transfer learning, broad learning, and knowledge graph utilization, to mitigate this challenge. Furthermore, it delves into the challenges and future direction within the RS domain, offering insights that are poised to facilitate the development and implementation of robust RSs, particularly when confronted with data scarcity. We aim to provide valuable guidance and inspiration for researchers and practitioners, ultimately driving advancements in the field of RS.

cs.IR

Multimodal Large Language Models: A Survey

The exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models excel in text-based tasks, they often struggle to understand and process other data types. Multimodal models address this limitation by combining various modalities, enabling a more comprehensive understanding of diverse data. This paper begins by defining the concept of multimodal and examining the historical development of multimodal algorithms. Furthermore, we introduce a range of multimodal products, focusing on the efforts of major technology companies. A practical guide is provided, offering insights into the technical aspects of multimodal models. Moreover, we present a compilation of the latest algorithms and commonly used datasets, providing researchers with valuable resources for experimentation and evaluation. Lastly, we explore the applications of multimodal models and discuss the challenges associated with their development. By addressing these aspects, this paper aims to facilitate a deeper understanding of multimodal models and their potential in various domains.

cs.AI

TALENT: Targeted Mining of Non-overlapping Sequential Patterns

With the widespread application of efficient pattern mining algorithms, sequential patterns that allow gap constraints have become a valuable tool to discover knowledge from biological data such as DNA and protein sequences. Among all kinds of gap-constrained mining, non-overlapping sequence mining can mine interesting patterns and satisfy the anti-monotonic property (the Apriori property). However, existing algorithms do not search for targeted sequential patterns, resulting in unnecessary and redundant pattern generation. Targeted pattern mining can not only mine patterns that are more interesting to users but also reduce the unnecessary redundant sequence generated, which can greatly avoid irrelevant computation. In this paper, we define and formalize the problem of targeted non-overlapping sequential pattern mining and propose an algorithm named TALENT (TArgeted mining of sequentiaL pattErN with consTraints). Two search methods including breadth-first and depth-first searching are designed to troubleshoot the generation of patterns. Furthermore, several pruning strategies to reduce the reading of sequences and items in the data and terminate redundant pattern extensions are presented. Finally, we select a series of datasets with different characteristics and conduct extensive experiments to compare the TALENT algorithm with the existing algorithms for mining non-overlapping sequential patterns. The experimental results demonstrate that the proposed targeted mining algorithm, TALENT, has excellent mining efficiency and can deal efficiently with many different query settings.

cs.DB

Open Metaverse: Issues, Evolution, and Future

With the evolution of content on the web and the Internet, there is a need for cyberspace that can be used to work, live, and play in digital worlds regardless of geography. The Metaverse provides the possibility of future Internet and represents a future trend. In the future, the Metaverse will be a space where the real and the virtual are combined. In this article, we have a comprehensive survey of the compelling Metaverse. We introduce computer technology, the history of the Internet, and the promise of the Metaverse as the next generation of the Internet. In addition, we briefly introduce the related concepts of the Metaverse, including novel terms like trusted Metaverse, human-intelligence Metaverse, personalized Metaverse, AI-enabled Metaverse, Metaverse-as-a-service, etc. Moreover, we present the challenges of the Metaverse such as limited resources and ethical issues. We also present Metaverse's promising directions, including lightweight Metaverse and autonomous Metaverse. We hope this survey will provide some helpful prospects and insightful directions about the Metaverse to related developments.

cs.DB

Towards Top-$K$ Non-Overlapping Sequential Patterns

Sequential pattern mining (SPM) has excellent prospects and application spaces and has been widely used in different fields. The non-overlapping SPM, as one of the data mining techniques, has been used to discover patterns that have requirements for gap constraints in some specific mining tasks, such as bio-data mining. And for the non-overlapping sequential patterns with gap constraints, the Nettree structure has been proposed to efficiently compute the support of the patterns. For pattern mining, users usually need to consider the threshold of minimum support (\textit{minsup}). This is especially difficult in the case of large databases. Although some existing algorithms can mine the top-$k$ patterns, they are approximate algorithms with fixed lengths. In this paper, a precise algorithm for mining \underline{T}op-$k$ \underline{N}on-\underline{O}verlapping \underline{S}equential \underline{P}atterns (TNOSP) is proposed. The top-$k$ solution of SPM is an effective way to discover the most frequent non-overlapping sequential patterns without having to set the \textit{minsup}. As a novel pattern mining algorithm, TNOSP can precisely search the top-$k$ patterns of non-overlapping sequences with different gap constraints. We further propose a pruning strategy named \underline{Q}ueue \underline{M}eta \underline{S}et \underline{P}runing (QMSP) to improve TNOSP's performance. TNOSP can reduce redundancy in non-overlapping sequential mining and has better performance in mining precise non-overlapping sequential patterns. The experimental results and comparisons on several datasets have shown that TNOSP outperformed the existing algorithms in terms of precision, efficiency, and scalability.

cs.DB

AI-Generated Content (AIGC): A Survey

To address the challenges of digital intelligence in the digital economy, artificial intelligence-generated content (AIGC) has emerged. AIGC uses artificial intelligence to assist or replace manual content generation by generating content based on user-inputted keywords or requirements. The development of large model algorithms has significantly strengthened the capabilities of AIGC, which makes AIGC products a promising generative tool and adds convenience to our lives. As an upstream technology, AIGC has unlimited potential to support different downstream applications. It is important to analyze AIGC's current capabilities and shortcomings to understand how it can be best utilized in future applications. Therefore, this paper provides an extensive overview of AIGC, covering its definition, essential conditions, cutting-edge capabilities, and advanced features. Moreover, it discusses the benefits of large-scale pre-trained models and the industrial chain of AIGC. Furthermore, the article explores the distinctions between auxiliary generation and automatic generation within AIGC, providing examples of text generation. The paper also examines the potential integration of AIGC with the Metaverse. Lastly, the article highlights existing issues and suggests some future directions for application.

cs.AI

The Human-Centric Metaverse: A Survey

In the era of the Web of Things, the Metaverse is expected to be the landing site for the next generation of the Internet, resulting in the increased popularity of related technologies and applications in recent years and gradually becoming the focus of Internet research. The Metaverse, as a link between the real and virtual worlds, can provide users with immersive experiences. As the concept of the Metaverse grows in popularity, many scholars and developers begin to focus on the Metaverse's ethics and core. This paper argues that the Metaverse should be centered on humans. That is, humans constitute the majority of the Metaverse. As a result, we begin this paper by introducing the Metaverse's origins, characteristics, related technologies, and the concept of the human-centric Metaverse (HCM). Second, we discuss the manifestation of human-centric in the Metaverse. Finally, we discuss some current issues in the construction of HCM. In this paper, we provide a detailed review of the applications of human-centric technologies in the Metaverse, as well as the relevant HCM application scenarios. We hope that this paper can provide researchers and developers with some directions and ideas for human-centric Metaverse construction.

cs.HC

Metaverse Security and Privacy: An Overview

Metaverse is a living space and cyberspace that realizes the process of virtualizing and digitizing the real world. It integrates a plethora of existing technologies with the goal of being able to map the real world, even beyond the real world. Metaverse has a bright future and is expected to have many applications in various scenarios. The support of the Metaverse is based on numerous related technologies becoming mature. Hence, there is no doubt that the security risks of the development of the Metaverse may be more prominent and more complex. We present some Metaverse-related technologies and some potential security and privacy issues in the Metaverse. We present current solutions for Metaverse security and privacy derived from these technologies. In addition, we also raise some unresolved questions about the potential Metaverse. To summarize, this survey provides an in-depth review of the security and privacy issues raised by key technologies in Metaverse applications. We hope that this survey will provide insightful research directions and prospects for the Metaverse's development, particularly in terms of security and privacy protection in the Metaverse.

cs.CR

Big Data Meets Metaverse: A Survey

We are living in the era of big data. The Metaverse is an emerging technology in the future, and it has a combination of big data, AI (artificial intelligence), VR (Virtual Reality), AR (Augmented Reality), MR (mixed reality), and other technologies that will diminish the difference between online and real-life interaction. It has the goal of becoming a platform where we can work, go shopping, play around, and socialize. Each user who enters the Metaverse interacts with the virtual world in a data way. With the development and application of the Metaverse, the data will continue to grow, thus forming a big data network, which will bring huge data processing pressure to the digital world. Therefore, big data processing technology is one of the key technologies to implement the Metaverse. In this survey, we provide a comprehensive review of how Metaverse is changing big data. Moreover, we discuss the key security and privacy of Metaverse big data in detail. Finally, we summarize the open problems and opportunities of Metaverse, as well as the future of Metaverse with big data. We hope that this survey will provide researchers with the research direction and prospects of applying big data in the Metaverse.

cs.DB