SearcharxivSearch

arXiv subjects

Kaixuan Chen

Publications and source records attributed to Kaixuan Chen.

At least 19 recordsLinked to original sources

CodeTS: Verifiable Text-to-Time Series Generation via Executable Code

Text-to-Time Series Generation (Text-to-TS) provides a promising paradigm for synthesizing time series from natural language, enabling scenario-specific generation when real observations are scarce or costly to acquire. However, existing methods typically lack an explicit mechanism for deriving generation logic from textual descriptions to guide time series synthesis. In this paper, we propose CodeTS, a verifiable framework that uses code as an intermediate generation interface, reformulating Text-to-TS generation as a Text-to-Code-to-TS process. CodeTS first maps textual temporal descriptions into an explicit code space, where executable code specifies how textual requirements shape target temporal patterns, and then obtains the time series through code execution. To learn this code generation process reliably without real code annotations, CodeTS constructs aligned Text-Code-TS triplets from structured temporal attributes for supervised initialization. More importantly, we further design multi-stage execution-based rewards that verify format validity, code executability, and time series quality, enabling real Text-TS pairs to provide training signals for Reinforcement Learning with Verifiable Rewards (RLVR). Extensive experiments on eight benchmarks across short, medium, and long generation lengths demonstrate that CodeTS provides a strong zero-shot solution for Text-to-TS generation, outperforming LLM-based baselines and achieving better averaged results than supervised generative baselines trained on the target datasets.

cs.LG

Cascading Relevance-driven Recommendation Network for CTR Prediction in Trigger-Introduced Recommendation

E-commerce has emerged as crucial platforms for people's daily consumption and shopping interests. There is a new recommendation scenario, Trigger-Introduced Recommendation (TIR), where users click interested product, which is defined as the trigger item, containing their instant interest, and in the undertaking page following the relevant target items. Distinguished from traditional search and recommendation scenarios, trigger contains relatively strong instant interest, which is more vague and implicit compared to search terms. Relying on large amounts of labeled data, existing methods lack the exploration of trigger relevance, which affects users' immersive experience. To alleviate this problem, we propose the Cascading Relevance-driven Recommendation Network (CRRN) to emphasize the interaction and relevance between trigger and target, comprising three essential components: 1) the Trigger-Target Interaction layer extracts interaction features of trigger and target based on personalized gating. 2) Cascading Interest Fusion module explicitly estimates users' trigger intention and fuses instant and personalized interests adaptively with cascading attention blocks. 3) Category-assisted Pairwise Loss enhances trigger relevance with the guidance of category association between trigger and target. Extensive experiment results show that CRRN outperforms recent state-of-the-art methods on both industrial and public datasets. Online A/B tests further validate the effectiveness of our method. Our code is available at https://github.com/a-little-cabbage/CRRN.

cs.IR

Robust and Active Visible-Light Integrated Photonics on Thin-Film Lithium Tantalate for Underwater Optical Wireless Communications

Visible-light integrated photonics enables compact platforms for sensing, precision metrology, and free-space data links at visible wavelengths. However, many applications remain limited by the lack of high-speed and robust modulators in the blue-green band. Here we report, both operating at 532 nm, thin-film lithium tantalate waveguides of propagation losses of dB/cm scale and modulators with a flat frequency response to ~50 GHz. The modulator remains stable when delivering 5 dBm modulated optical power for an hour, which cannot be achieved by thin-film lithium niobate based counterparts under similar conditions and structures. System-level underwater optical wireless communication (UWOC) is validated with 112-Gb/s transmission over 3-m and 64-Gb/s transmission over 9-m underwater links. This represents the first integrated external modulator based UWOC system, overcoming the bandwidth-power-chirp trade-offs of traditional directly modulated laser-based systems. We further demonstrate dual-drive modulators for optical single-sideband and electro-optic frequency-comb generations in the green-wavelength band. These results provide a foundation for complex, robust, and active visible-light photonic integrated circuits for underwater optical applications.

physics.optics

Rethinking Adapter Placement: A Dominant Adaptation Module Perspective

Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method that places trainable low-rank adapters into frozen pre-trained models. Recent studies show that using fewer LoRA adapters may still maintain or even improve performance, but existing methods still distribute adapters broadly, leaving where to place a limited number of adapters to maximize performance largely open. To investigate this, we introduce PAGE (Projected Adapter Gradient Energy), a gradient-based sensitivity probe that estimates the initial trainable gradient energy available to each candidate LoRA adapter. Surprisingly, we find that PAGE is highly concentrated on a single shallow FFN down-projection across two model families and four downstream tasks. We term this module the dominant adaptation module and show that its layer index is architecture-dependent but task-stable. Motivated by this finding, we propose DomLoRA, a placement method that places a single adapter at the dominant adaptation module. With only ~0.7% of vanilla LoRA's trainable parameters, DomLoRA outperforms it on average across various downstream tasks, including instruction following, mathematical reasoning, code generation, and multi-turn conversation. This method also improves other LoRA variants, supporting the dominant adaptation module perspective as a practical placement guideline.

cs.AI

Evolutionary Negative Module Pruning for Better LoRA Merging

Merging multiple Low-Rank Adaptation (LoRA) experts into a single backbone is a promising approach for efficient multi-task deployment. While existing methods strive to alleviate interference via weight interpolation or subspace alignment, they rest upon the implicit assumption that all LoRA matrices contribute constructively to the merged model. In this paper, we uncover a critical bottleneck in current merging paradigms: the existence of $\textit{negative modules}$ -- specific LoRA layers that inherently degrade global performance upon merging. We propose $\textbf{E}$volutionary $\textbf{N}$egative $\textbf{M}$odule $\textbf{P}$runing ($\textbf{ENMP}$), a plug-and-play LoRA pruning method to locate and exclude these detrimental modules prior to merging. By leveraging an evolutionary search strategy, ENMP effectively navigates the discrete, non-differentiable landscape of module selection to identify optimal pruning configurations. Extensive evaluations demonstrate that ENMP consistently boosts the performance of existing merging algorithms, achieving a new state-of-the-art across both language and vision domains. Code is available at https://github.com/CaoAnda/ENMP-LoRAMerging.

cs.AI

Web-Gewu: A Browser-Based Interactive Playground for Robot Reinforcement Learning

With the rapid development of embodied intelligence, robotics education faces a dual challenge: high computational barriers and cumbersome environment configuration. Existing centralized cloud simulation solutions incur substantial GPU and bandwidth costs that preclude large-scale deployment, while pure local computing is severely constrained by learners' hardware limitations. To address these issues, we propose \href{http://47.76.242.88:8080/receiver/index.html}{Web-Gewu}, an interactive robotics education platform built on a WebRTC cloud-edge-client collaborative architecture. The system offloads all physics simulation and reinforcement learning (RL) training to the edge node, while the cloud server acts exclusively as a lightweight signaling relay, enabling extremely low-cost browser-based peer-to-peer (P2P) real-time streaming. Learners can interact with multi-form robots at low end-to-end latency directly in a web browser without any local installation, and simultaneously observe real-time visualization of multi-dimensional monitoring data, including reinforcement learning reward curves. Combined with a predefined robust command communication protocol, Web-Gewu provides a highly scalable, out-of-the-box, and barrier-free teaching infrastructure for embodied intelligence, significantly lowering the barrier to entry for cutting-edge robotics technology.

cs.RO

GraphScout: Empowering Large Language Models with Intrinsic Exploration Ability for Agentic Graph Reasoning

Knowledge graphs provide structured and reliable information for many real-world applications, motivating increasing interest in combining large language models (LLMs) with graph-based retrieval to improve factual grounding. Recent Graph-based Retrieval-Augmented Generation (GraphRAG) methods therefore introduce iterative interaction between LLMs and knowledge graphs to enhance reasoning capability. However, existing approaches typically depend on manually designed guidance and interact with knowledge graphs through a limited set of predefined tools, which substantially constrains graph exploration. To address these limitations, we propose GraphScout, a training-centric agentic graph reasoning framework equipped with more flexible graph exploration tools. GraphScout enables models to autonomously interact with knowledge graphs to synthesize structured training data which are then used to post-train LLMs, thereby internalizing agentic graph reasoning ability without laborious manual annotation or task curation. Extensive experiments across five knowledge-graph domains show that a small model (e.g., Qwen3-4B) augmented with GraphScout outperforms baseline methods built on leading LLMs (e.g., Qwen-Max) by an average of 16.7\% while requiring significantly fewer inference tokens. Moreover, GraphScout exhibits robust cross-domain transfer performance. Our code will be made publicly available~\footnote{https://github.com/Ying-Yuchen/_GraphScout_}.

cs.AI

High-performance Sources of Multidimensionally Engineered Quantum Light Based on Monolithic Microcavity-metalens Interfaces

The ultimate non-classic light sources for modern photonic quantum technology require on-demand generation of indistinguishable quantum light with high brightness and flexible engineering of quantum emission in multiple degrees of freedom. In this work, we present monolithic microcavity-metalens interfaces consisting of quantum-dot-micropillar single-photon sources and ultra-thin metalenses accurately aligned on opposite sides of an III-V compound semiconductor chip. The pronounced cavity quantum electrodynamics effect enabled by the micropillar cavity facilitates single-photon emission from quantum dots with simultaneous high degrees of single-photon purity, source brightness and photon indistinguishability while the multi-functional metalenses concurrently tailor quantum emission in multiple physical degrees of freedom including radiation divergence, emission directionality, polarization state and orbital angular momentum (OAM). Furthermore, high-fidelity polarization-OAM entanglement and single photons with local spin topologies are successfully generated in our integrated device. In particular, we demonstrate stable propagations of single-photon skrymions in atmospheric turbulence and reveal their topological advantages over the conventional structured quantum light. Our work advances the research fields of integrated quantum photonics and meta-optics, providing integrated high-dimensional quantum light sources for advanced photonic quantum science and technology.

physics.optics

Physics-informed Diffusion Generation for Geomagnetic Map Interpolation

Geomagnetic map interpolation aims to infer unobserved geomagnetic data at spatial points, yielding critical applications in navigation and resource exploration. However, existing methods for scattered data interpolation are not specifically designed for geomagnetic maps, which inevitably leads to suboptimal performance due to detection noise and the laws of physics. Therefore, we propose a Physics-informed Diffusion Generation framework~(PDG) to interpolate incomplete geomagnetic maps. First, we design a physics-informed mask strategy to guide the diffusion generation process based on a local receptive field, effectively eliminating noise interference. Second, we impose a physics-informed constraint on the diffusion generation results following the kriging principle of geomagnetic maps, ensuring strict adherence to the laws of physics. Extensive experiments and in-depth analyses on four real-world datasets demonstrate the superiority and effectiveness of each component of PDG.

cs.AI

SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data

Recent advances have demonstrated the effectiveness of Reinforcement Learning (RL) in improving the reasoning capabilities of Large Language Models (LLMs). However, existing works inevitably rely on high-quality instructions and verifiable rewards for effective training, both of which are often difficult to obtain in specialized domains. In this paper, we propose Self-play Reinforcement Learning (SeRL) to bootstrap LLM training with limited initial data. Specifically, SeRL comprises two complementary modules: self-instruction and self-rewarding. The former module generates additional instructions based on the available data at each training step, employing robust online filtering strategies to ensure instruction quality, diversity, and difficulty. The latter module introduces a simple yet effective majority-voting mechanism to estimate response rewards for additional instructions, eliminating the need for external annotations. Finally, SeRL performs conventional RL based on the generated data, facilitating iterative self-play learning. Extensive experiments on various reasoning benchmarks and across different LLM backbones demonstrate that the proposed SeRL yields results superior to its counterparts and achieves performance on par with those obtained by high-quality data with verifiable rewards. Our code is available at https://github.com/wantbook-book/SeRL.

cs.CL

Heterogeneous back-end-of-line integration of thin-film lithium niobate on active silicon photonics for single-chip optical transceivers

The explosive growth of artificial intelligence, cloud computing, and large-scale machine learning is driving an urgent demand for short-reach optical interconnects featuring large bandwidth, low power consumption, high integration density, and low cost preferably adopting complementary metal-oxide-semiconductor (CMOS) processes. Heterogeneous integration of silicon photonics and thin-film lithium niobate (TFLN) combines the advantages of both platforms, and enables co-integration of high-performance modulators, photodetectors, and passive photonic components, offering an ideal route to meet these requirements. However, process incompatibilities have constrained the direct integration of TFLN with only passive silicon photonics. Here, we demonstrate the first heterogeneous back-end-of-line integration of TFLN with a full-functional and active silicon photonics platform via trench-based die-to-wafer bonding. This technology introduces TFLN after completing the full CMOS compatible processes for silicon photonics. Si/SiN passive components including low-loss fiber interfaces, 56-GHz Ge photodetectors, 100-GHz TFLN modulators, and multilayer metallization are integrated on a single silicon chip with efficient inter-layer and inter-material optical coupling. The integrated on-chip optical links exhibit greater than 60 GHz electrical-to-electrical bandwidth and support 128-GBaud OOK and 100-GBaud PAM4 transmission below forward error-correction thresholds, establishing a scalable platform for energy-efficient, high-capacity photonic systems.

physics.optics

Towards Efficient LLM-aware Heterogeneous Graph Learning

Heterogeneous graphs are widely present in real-world complex networks, where the diversity of node and relation types leads to complex and rich semantics. Efforts for modeling complex relation semantics in heterogeneous graphs are restricted by the limitations of predefined semantic dependencies and the scarcity of supervised signals. The advanced pre-training and fine-tuning paradigm leverages graph structure to provide rich self-supervised signals, but introduces semantic gaps between tasks. Large Language Models (LLMs) offer significant potential to address the semantic issues of relations and tasks in heterogeneous graphs through their strong reasoning capabilities in textual modality, but their incorporation into heterogeneous graphs is largely limited by computational complexity. Therefore, in this paper, we propose an Efficient LLM-Aware (ELLA) framework for heterogeneous graphs, addressing the above issues. To capture complex relation semantics, we propose an LLM-aware Relation Tokenizer that leverages LLM to encode multi-hop, multi-type relations. To reduce computational complexity, we further employ a Hop-level Relation Graph Transformer, which help reduces the complexity of LLM-aware relation reasoning from exponential to linear. To bridge semantic gaps between pre-training and fine-tuning tasks, we introduce the fine-grained task-aware textual Chain-of-Thought (CoT) prompts. Extensive experiments on four heterogeneous graphs show that our proposed ELLA outperforms state-of-the-art methods in the performance and efficiency. In particular, ELLA scales up to 13b-parameter LLMs and achieves up to a 4x speedup compared with existing LLM-based methods. Our code is publicly available at https://github.com/l-wd/ELLA.

cs.CL

First-Principles Exploration of Pentagonal TiN$_8$ and MoN$_8$ Monolayers as New Magnetic Topological Insulator

The quest for robust, intrinsically magnetic topological materials exhibiting the quantum anomalous Hall (QAH) effect is a central challenge in condensed matter physics and the application of revolutionary electronics. However, progress has been hampered by the limited number of candidate materials, which often suffer from poor stability and complex synthesis. Here, we introduce a new paradigm by exploring the emergent magnetism and nontrivial band topology in the largely overlooked family of two-dimensional (2D) pentagonal MN$_8$ monolayers. Employing first-principles calculations, we reveal that these systems host out-of-plane ferromagnetic ground states, a key feature that unlocks nontrivial topological properties driven by the localized $d$-orbitals of the embedded transition metals. Remarkably, we identify TiN$_8$ as a QAH insulator characterized by a Chern number of $C=-1$. Even more strikingly, MoN$_8$ is predicted to be a rare high-Chern-number QAH insulator, boasting a Chern number of $C=2$. Our findings establish the penta-MN$_8$ family as a fertile and versatile platform for realizing exotic topological quantum states. This work not only significantly expands the material landscape for magnetic topological insulators but also provides a solid theoretical foundation for designing next-generation spintronic and quantum computing devices.

cond-mat.mtrl-sci

A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges

Reinforcement Learning (RL) is a popular machine learning paradigm where intelligent agents interact with the environment to fulfill a long-term goal. Driven by the resurgence of deep learning, Deep RL (DRL) has witnessed great success over a wide spectrum of complex control tasks. Despite the encouraging results achieved, the deep neural network-based backbone is widely deemed as a black box that impedes practitioners to trust and employ trained agents in realistic scenarios where high security and reliability are essential. To alleviate this issue, a large volume of literature devoted to shedding light on the inner workings of the intelligent agents has been proposed, by constructing intrinsic interpretability or post-hoc explainability. In this survey, we provide a comprehensive review of existing works on eXplainable RL (XRL) and introduce a new taxonomy where prior works are clearly categorized into model-explaining, reward-explaining, state-explaining, and task-explaining methods. We also review and highlight RL methods that conversely leverage human knowledge to promote learning efficiency and performance of agents while this kind of method is often ignored in XRL field. Some challenges and opportunities in XRL are discussed. This survey intends to provide a high-level summarization of XRL and to motivate future research on more effective XRL solutions. Corresponding open source codes are collected and categorized at https://github.com/Plankson/awesome-explainable-reinforcement-learning.

cs.LG

Gaussian Process Latent Variable Modeling for Few-shot Time Series Forecasting

Accurate time series forecasting is crucial for optimizing resource allocation, industrial production, and urban management, particularly with the growth of cyber-physical and IoT systems. However, limited training sample availability in fields like physics and biology poses significant challenges. Existing models struggle to capture long-term dependencies and to model diverse meta-knowledge explicitly in few-shot scenarios. To address these issues, we propose MetaGP, a meta-learning-based Gaussian process latent variable model that uses a Gaussian process kernel function to capture long-term dependencies and to maintain strong correlations in time series. We also introduce Kernel Association Search (KAS) as a novel meta-learning component to explicitly model meta-knowledge, thereby enhancing both interpretability and prediction accuracy. We study MetaGP on simulated and real-world few-shot datasets, showing that it is capable of state-of-the-art prediction accuracy. We also find that MetaGP can capture long-term dependencies and can model meta-knowledge, thereby providing valuable insights into complex time series patterns.

cs.LG

Substrate Effect on Electronic Band Structure and Topological Property in Monolayer V2O3 Magnetic Topological Insulator

Monolayer V2O3, a two-dimensional magnetic topological insulator with intrinsic ferromagnetic order and a nontrivial band gap, offers a promising platform for realizing quantum anomalous Hall (QAH) states. Using first-principles density functional theory calculations, we systematically investigate the impact of substrate selection on its electronic and topological properties. By modeling heterostructures with van der Waals (vdW) substrates, we demonstrate that non-magnetic substrates such as h-BN preserve the QAH phase with a Chern number C = 1, maintaining gapless chiral edge states. In contrast, ferromagnetic substrates induce extra electrons, destroying the topological order by shifting the Fermi level. These findings establish substrate engineering as a pivotal strategy for experimental realization of dissipationless edge transport in V2O3-based vdW heterostructures, advancing their potential applications as low-power topological electronics.

cond-mat.mtrl-sci

Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?

Centralized Training with Decentralized Execution (CTDE) has recently emerged as a popular framework for cooperative Multi-Agent Reinforcement Learning (MARL), where agents can use additional global state information to guide training in a centralized way and make their own decisions only based on decentralized local policies. Despite the encouraging results achieved, CTDE makes an independence assumption on agent policies, which limits agents to adopt global cooperative information from each other during centralized training. Therefore, we argue that existing CTDE methods cannot fully utilize global information for training, leading to an inefficient joint-policy exploration and even suboptimal results. In this paper, we introduce a novel Centralized Advising and Decentralized Pruning (CADP) framework for multi-agent reinforcement learning, that not only enables an efficacious message exchange among agents during training but also guarantees the independent policies for execution. Firstly, CADP endows agents the explicit communication channel to seek and take advices from different agents for more centralized training. To further ensure the decentralized execution, we propose a smooth model pruning mechanism to progressively constraint the agent communication into a closed one without degradation in agent cooperation capability. Empirical evaluations on StarCraft II micromanagement and Google Research Football benchmarks demonstrate that the proposed framework achieves superior performance compared with the state-of-the-art counterparts. Our code will be made publicly available.

cs.AI

From GNNs to Trees: Multi-Granular Interpretability for Graph Neural Networks

Interpretable Graph Neural Networks (GNNs) aim to reveal the underlying reasoning behind model predictions, attributing their decisions to specific subgraphs that are informative. However, existing subgraph-based interpretable methods suffer from an overemphasis on local structure, potentially overlooking long-range dependencies within the entire graphs. Although recent efforts that rely on graph coarsening have proven beneficial for global interpretability, they inevitably reduce the graphs to a fixed granularity. Such an inflexible way can only capture graph connectivity at a specific level, whereas real-world graph tasks often exhibit relationships at varying granularities (e.g., relevant interactions in proteins span from functional groups, to amino acids, and up to protein domains). In this paper, we introduce a novel Tree-like Interpretable Framework (TIF) for graph classification, where plain GNNs are transformed into hierarchical trees, with each level featuring coarsened graphs of different granularity as tree nodes. Specifically, TIF iteratively adopts a graph coarsening module to compress original graphs (i.e., root nodes of trees) into increasingly coarser ones (i.e., child nodes of trees), while preserving diversity among tree nodes within different branches through a dedicated graph perturbation module. Finally, we propose an adaptive routing module to identify the most informative root-to-leaf paths, providing not only the final prediction but also the multi-granular interpretability for the decision-making process. Extensive experiments on the graph classification benchmarks with both synthetic and real-world datasets demonstrate the superiority of TIF in interpretability, while also delivering a competitive prediction performance akin to the state-of-the-art counterparts.

cs.LG