Searcharxiv⌕ Search

arXiv subjects

Rui Ding

Publications and source records attributed to Rui Ding.

At least 37 records · Page 2Linked to original sources

Dynamically Augmented CVaR for MDPs

This paper studies optimization of Conditional Value-at-Risk (CVaR) for Markov Decision Processes (MDPs) with finite state and action sets. It introduces the Dynamically augmented CVaR (DCVaR) risk measure and provides an algorithm for its optimization. This paper investigates a specially defined Robust MDP (RMDP), in which the state space is augmented with the tail risk level. This RMDP, which we call the Dynamically augmented RMDP (DRMDP), was introduced to the literature for calculations of optimal CVaR values by value iteration more than ten years ago, but, as was understood later, these value iterations compute lower bounds of minimal static CVaRs. DCVaR is defined as a time consistent version of the static CVaR, and it is a lower bound of the static CVaR. It also can be considered as a dynamic version of the nested CVaR. This paper provides an algorithm constructing a policy optimizing DCVaR of total discounted costs. The correctness of this algorithm is proved by studying a special mass transfer problem. The results on RMDPs needed for this paper are provided in the appendix.

math.OC↗

Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection

Monocular 3D object detection is a promising yet ill-posed task for autonomous vehicles due to the lack of accurate depth information. Cross-modality knowledge distillation could effectively transfer depth information from LiDAR to image-based network. However, modality gap between image and LiDAR seriously limits its accuracy. In this paper, we systematically investigate the negative transfer problem induced by modality gap in cross-modality distillation for the first time, including not only the architecture inconsistency issue but more importantly the feature overfitting issue. We propose a selective learning approach named MonoSTL to overcome these issues, which encourages positive transfer of depth information from LiDAR while alleviates the negative transfer on image-based network. On the one hand, we utilize similar architectures to ensure spatial alignment of features between image-based and LiDAR-based networks. On the other hand, we develop two novel distillation modules, namely Depth-Aware Selective Feature Distillation (DASFD) and Depth-Aware Selective Relation Distillation (DASRD), which selectively learn positive features and relationships of objects by integrating depth uncertainty into feature and relation distillations, respectively. Our approach can be seamlessly integrated into various CNN-based and DETR-based models, where we take three recent models on KITTI and a recent model on NuScenes for validation. Extensive experiments show that our approach considerably improves the accuracy of the base models and thereby achieves the best accuracy compared with all recently released SOTA models.

cs.CV↗

Multi-Modal Decouple and Recouple Network for Robust 3D Object Detection

Multi-modal 3D object detection with bird's eye view (BEV) has achieved desired advances on benchmarks. Nonetheless, the accuracy may drop significantly in the real world due to data corruption such as sensor configurations for LiDAR and scene conditions for camera. One design bottleneck of previous models resides in the tightly coupling of multi-modal BEV features during fusion, which may degrade the overall system performance if one modality or both is corrupted. To mitigate, we propose a Multi-Modal Decouple and Recouple Network for robust 3D object detection under data corruption. Different modalities commonly share some high-level invariant features. We observe that these invariant features across modalities do not always fail simultaneously, because different types of data corruption affect each modality in distinct ways.These invariant features can be recovered across modalities for robust fusion under data corruption.To this end, we explicitly decouple Camera/LiDAR BEV features into modality-invariant and modality-specific parts. It allows invariant features to compensate each other while mitigates the negative impact of a corrupted modality on the other.We then recouple these features into three experts to handle different types of data corruption, respectively, i.e., LiDAR, camera, and both.For each expert, we use modality-invariant features as robust information, while modality-specific features serve as a complement.Finally, we adaptively fuse the three experts to exact robust features for 3D object detection. For validation, we collect a benchmark with a large quantity of data corruption for LiDAR, camera, and both based on nuScenes. Our model is trained on clean nuScenes and tested on all types of data corruption. Our model consistently achieves the best accuracy on both corrupted and clean data compared to recent models.

cs.CV↗

RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object Detection

Multi-view 3D detection with bird's eye view (BEV) is crucial for autonomous driving and robotics, but its robustness in real-world is limited as it struggles to predict accurate depth values. A mainstream solution, cross-modal distillation, transfers depth information from LiDAR to camera models but also unintentionally transfers depth-irrelevant information (e.g. LiDAR density). To mitigate this issue, we propose RayD3D, which transfers crucial depth knowledge along the ray: a line projecting from the camera to true location of an object. It is based on the fundamental imaging principle that predicted location of this object can only vary along this ray, which is finally determined by predicted depth value. Therefore, distilling along the ray enables more effective depth information transfer. More specifically, we design two ray-based distillation modules. Ray-based Contrastive Distillation (RCD) incorporates contrastive learning into distillation by sampling along the ray to learn how LiDAR accurately locates objects. Ray-based Weighted Distillation (RWD) adaptively adjusts distillation weight based on the ray to minimize the interference of depth-irrelevant information in LiDAR. For validation, we widely apply RayD3D into three representative types of BEV-based models, including BEVDet, BEVDepth4D, and BEVFormer. Our method is trained on clean NuScenes, and tested on both clean NuScenes and RoboBEV with a variety types of data corruptions. Our method significantly improves the robustness of all the three base models in all scenarios without increasing inference costs, and achieves the best when compared to recently released multi-view and distillation models.

cs.CV↗

Generalization of RLVR Using Causal Reasoning as a Testbed

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for post-training large language models (LLMs) on complex reasoning tasks. Yet, the conditions under which RLVR yields robust generalization remain underexplored. This paper provides an empirical study of RLVR generalization in the setting of probabilistic inference over causal graphical models. This setting offers two natural axes along which to examine generalization: (i) the level of the probabilistic query -- associational, interventional, or counterfactual -- and (ii) the structural complexity of the query, measured by the size of its relevant subgraph. We construct a dataset of causal graphs and queries spanning these difficulty axes and fine-tune Qwen-2.5-Instruct models using RLVR or supervised fine-tuning (SFT). We vary both the model scale (3B-32B) and the query level included in training. We find that RLVR yields stronger within-level and across-level generalization than SFT, but only for specific combinations of model size and training query level. Further analysis shows that RLVR's effectiveness depends on the model's initial reasoning competence. With sufficient initial competence, RLVR improves an LLM's marginalization strategy and reduces errors in intermediate probability calculations, producing substantial accuracy gains, particularly on more complex queries. These results show that RLVR can improve specific causal reasoning subskills, with its benefits emerging only when the model has sufficient initial competence. Our code and data is available at https://github.com/zhichul/rlcausal.

cs.LG↗

Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection

Monocular 3D object detection (M3OD) is intrinsically ill-posed, hence training a high-performance deep learning based M3OD model requires a humongous amount of labeled data with complicated visual variation from diverse scenes, variety of objects and camera poses.However, we observe that, due to strong human bias, the three independent entities, i.e., object, scene, and camera pose, are always tightly entangled when an image is captured to construct training data. More specifically, specific 3D objects are always captured in particular scenes with fixed camera poses, and hence lacks necessary diversity. Such tight entanglement induces the challenging issues of insufficient utilization and overfitting to uniform training data. To mitigate this, we propose an online object-scene-camera decomposition and recomposition data manipulation scheme to more efficiently exploit the training data. We first fully decompose training images into textured 3D object point models and background scenes in an efficient computation and storage manner. We then continuously recompose new training images in each epoch by inserting the 3D objects into the freespace of the background scenes, and rendering them with perturbed camera poses from textured 3D point representation. In this way, the refreshed training data in all epochs can cover the full spectrum of independent object, scene, and camera pose combinations. This scheme can serve as a plug-and-play component to boost M3OD models, working flexibly with both fully and sparsely supervised settings. In the sparsely-supervised setting, objects closest to the ego-camera for all instances are sparsely annotated. We then can flexibly increase the annotated objects to control annotation cost. For validation, our method is widely applied to five representative M3OD models and evaluated on both the KITTI and the more complicated Waymo datasets.

cs.CV↗

Test-Time Learning of Causal Structure from Interventional Data

Supervised causal learning has shown promise in causal discovery, yet it often struggles with generalization across diverse interventional settings, particularly when intervention targets are unknown. To address this, we propose TICL (Test-time Interventional Causal Learning), a novel method that synergizes Test-Time Training with Joint Causal Inference. Specifically, we design a self-augmentation strategy to generate instance-specific training data at test time, effectively avoiding distribution shifts. Furthermore, by integrating joint causal inference, we developed a PC-inspired two-phase supervised learning scheme, which effectively leverages self-augmented training data while ensuring theoretical identifiability. Extensive experiments on bnlearn benchmarks demonstrate TICL's superiority in multiple aspects of causal discovery and intervention target detection.

cs.LG↗

Embodied Intelligent Spectrum Management: A New Paradigm for Dynamic Spectrum Access

Wireless communication is evolving into an agent era, where numerous intelligent agents equipped with perception, reasoning, and interaction capabilities will operate in highly dynamic wireless environments. To complete diverse complex tasks, agent communication will play a critical role, which enables autonomous information exchange with external tools, services, and other agents. This trendy movement will dramatically increase spectrum demand and result in unprecedented challenges for spectrum management. However, current spectrum management paradigms, including static spectrum allocation and intelligent management, lack the flexibility and generalization to accommodate the dynamic and heterogeneous demands of agent communication. The recent advancements in embodied intelligence (EI) bring a promising solution, and this article will provide our vision of an emerging embodied intelligent spectrum management (EISM) paradigm. We start with an architecture for EISM, elaborating its key enabling technologies. Then, a prototype platform is presented to demonstrate the advantages of EISM. Finally, key challenges and open issues are outlined to facilitate future research in this emerging field.

cs.NI↗

Separating Energy and Entropy Contributions to the Hexatic-Liquid Transitions in Two-Dimensional Repulsive Systems

Over the past decades, research on two-dimensional melting has established that both first-order and continuous hexatic-liquid transitions can occur, influenced by various factors in the potential energy and system details. The fundamental thermodynamic origins of this sensitivity remains elusive. Here, by decomposing the Helmholtz free energy across three representative repulsive systems, we reveal a universal competition between energy and entropy that dictates the melting pathway. The energetic contribution consistently imparts convexity to the free energy, whereas entropy imparts concavity. A first-order transition occurs when concave entropy dominates; otherwise, the transition is continuous. Further decomposition shows that vibrational entropy drives the concave total entropic curvature, while the configurational entropy's curvature switches from convex (first-order) to concave (continuous), mirroring defect proliferation measured by Shannon entropy. The convexity of the energy is dominated by the inherent potential, with minimal vibrational influence. Finally, we predict and verify that the first-order transition becomes continuous at zero temperature, where entropic effects vanish. Our work establishes the curvature of different thermodynamic quantities as a fundamental principle for understanding the nature of two-dimensional melting.

cond-mat.stat-mech↗

Multi-Satellite Multi-Stream Beamspace Massive MIMO Transmission

This paper studies multi-satellite multi-stream (MSMS) beamspace transmission, where multiple satellites cooperate to form a distributed multiple-input multiple-output (MIMO) system and jointly deliver multiple data streams to multi-antenna user terminals (UTs), and beamspace transmission combines earth-moving beamforming with beam-domain precoding. For the first time, we formulate the signal model for MSMS beamspace MIMO transmission. Under synchronization errors, multi-antenna UTs enable the distributed MIMO channel to exhibit higher rank, supporting multiple data streams. Beamspace MIMO retains conventional codebook based beamforming while providing the performance gains of precoding. Based on the signal model, we propose statistical channel state information (sCSI)-based optimization of satellite clustering, beam selection, and transmit precoding, using a sum-rate upper-bound approximation. With given satellite clustering and beam selection, we cast precoder design as an equivalent covariance decomposition-based weighted minimum mean square error (CDWMMSE) problem. To obtain tractable algorithms, we develop a closed-form covariance decomposition required by CDWMMSE and derive an iterative MSMS beam-domain precoder under sCSI. Following this, we further propose several heuristic closed-form precoders to avoid iterative cost. For satellite clustering, we enhance a competition-based algorithm by introducing a mechanism to regulate the number of satellites serving certain UT. Furthermore, we design a two-stage low-complexity beam selection algorithm focused on enhancing the effective channel power. Simulations under practical configurations validate the proposed methods across the number of data streams, receive antennas, serving satellites, and active beams, and show that beamspace transmission approaches conventional MIMO performance at lower complexity.

eess.SP↗

Hierarchical Deep Research with Local-Web RAG: Toward Automated System-Level Materials Discovery

We present a long-horizon, hierarchical deep research (DR) agent designed for complex materials and device discovery problems that exceed the scope of existing Machine Learning (ML) surrogates and closed-source commercial agents. Our framework instantiates a locally deployable DR instance that integrates local retrieval-augmented generation with large language model reasoners, enhanced by a Deep Tree of Research (DToR) mechanism that adaptively expands and prunes research branches to maximize coverage, depth, and coherence. We systematically evaluate across 27 nanomaterials/device topics using a large language model (LLM)-as-judge rubric with five web-enabled state-of-the-art models as jurors. In addition, we conduct dry-lab validations on five representative tasks, where human experts use domain simulations (e.g., density functional theory, DFT) to verify whether DR-agent proposals are actionable. Results show that our DR agent produces reports with quality comparable to--and often exceeding--those of commercial systems (ChatGPT-5-thinking/o3/o4-mini-high Deep Research) at a substantially lower cost, while enabling on-prem integration with local data and tools.

cs.LG↗

Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMs

Multi-Agent System (MAS) developing frameworks serve as the foundational infrastructure for social simulations powered by Large Language Models (LLMs). However, existing frameworks fail to adequately support large-scale simulation development due to inherent limitations in adaptability, configurability, reliability, and code reusability. For example, they cannot simulate a society where the agent population and profiles change over time. To fill this gap, we propose Agent-Kernel, a framework built upon a novel society-centric modular microkernel architecture. It decouples core system functions from simulation logic and separates cognitive processes from physical environments and action execution. Consequently, Agent-Kernel achieves superior adaptability, configurability, reliability, and reusability. We validate the framework's superiority through two distinct applications: a simulation of the Universe 25 (Mouse Utopia) experiment, which demonstrates the handling of rapid population dynamics from birth to death; and a large-scale simulation of the Zhejiang University Campus Life, successfully coordinating 10,000 heterogeneous agents, including students and faculty.

cs.MA↗

Achieving Constant-Envelope Waveform in CP-OFDMA Framework

OFDM is widely adopted in modern wireless communication systems, but its power efficiency is limited by high envelope fluctuations. Although various high power-efficiency waveforms have been proposed, most are incompatible with the CP-OFDMA framework and remain ineffective in multi-user downlink transmissions. To address this issue, we propose a constant-envelope (CE) waveform design, which enables low-complexity transceiver architectures while maintaining full compatibility with the prevailing CP-OFDMA framework. Specifically, we start from a general CE FDMA signal model and develop a CP-OFDMA-compatible waveform implementation structure, followed by the design of an optimized CE-constrained pulse-shaping filter to suppress out-of-band emissions. To tackle channel estimation challenge under non-flat frequency-domain pilots induced by CE modulation, we optimize the time-domain binary pilot sequence to achieve frequency-domain CE properties, and then propose a multi-stage method combining delay-domain denoising with power delay profile estimation to facilitate reduced-dimension LMMSE estimation. Subsequently, we design a low-complexity maximum ratio combining-aided LMMSE equalizer by exploiting the periodicity and conjugate symmetry of the CE received signals. To mitigate the downlink peak-to-average power ratio increase caused by FDMA, we further develop a multi-user downlink CE transmission scheme including multiple access mechanism, downlink control information design, and corresponding system-level implementation, which ensures compatibility with the New Radio standard. Numerical results demonstrate that the proposed scheme achieves bit error rate performance close to the ideal case while significantly reducing transceiver complexity compared to existing CE waveform solutions.

eess.SP↗

Enhancing Optical Performance of Liquid Crystal Lens Arrays via Electrode Design Optimization

A liquid crystal (LC) lens array based on double-layer composite electrodes, characterized by a large aperture, short focal length, and low operating voltage is demonstrated. The lens array consists of an LC layer, a top common electrode, a bottom double-layer composite electrode layer, and an oxide layer. The bottom double-layer composite electrode layer comprises the pixel electrodes and the auxiliary electrode. In focusing mode, the pixel electrodes receive operational voltage to establish the LC layer's electric field, with the auxiliary electrode applying reduced voltage for field optimization. Experiment results show that the proposed LC lens array achieves the shortest focal length of 3.3 mm when the pixel electrodes are set at 5.2 Vrms and the auxiliary electrode is set at 2.6 Vrms. This design addresses the technical challenge of achieving larger apertures (800 μm or more), offering enhanced viewing zones and improved 3D performance. This configuration provides an ideal refractive index distribution in a relatively thick LC layer, enabling 2D/3D switchable display with performance superior to current LC lens arrays of equivalent aperture. Furthermore, the proposed structure demonstrates excellent tolerance to manufacturing errors.

physics.optics↗

Semantically-Guided Inference for Conditional Diffusion Models: Enhancing Covariate Consistency in Time Series Forecasting

Diffusion models have demonstrated strong performance in time series forecasting, yet often suffer from semantic misalignment between generated trajectories and conditioning covariates, especially under complex or multimodal conditions. To address this issue, we propose SemGuide, a plug-and-play, inference-time method that enhances covariate consistency in conditional diffusion models. Our approach introduces a scoring network to assess the semantic alignment between intermediate diffusion states and future covariates. These scores serve as proxy likelihoods in a stepwise importance reweighting procedure, which progressively adjusts the sampling path without altering the original training process. The method is model-agnostic and compatible with any conditional diffusion framework. Experiments on real-world forecasting tasks show consistent gains in both predictive accuracy and covariate alignment, with especially strong performance under complex conditioning scenarios.

cs.LG↗

Joint Resource Optimization Over Licensed and Unlicensed Spectrum in Spectrum Sharing UAV Networks Against Jamming Attacks

Unmanned aerial vehicle (UAV) communication is of crucial importance in realizing heterogeneous practical wireless application scenarios. However, the densely populated users and diverse services with high data rate demands has triggered an increasing scarcity of UAV spectrum utilization. To tackle this problem, it is promising to incorporate the underutilized unlicensed spectrum with the licensed spectrum to boost network capacity. However, the openness of unlicensed spectrum makes UAVs susceptible to security threats from potential jammers. Therefore, a spectrum sharing UAV network coexisting with licensed cellular network and unlicensed Wi-Fi network is considered with the anti-jamming technique in this paper. The sum rate maximization of the secondary network is studied by jointly optimizing the transmit power, subchannel allocation, and UAV trajectory. We first decompose the challenging non-convex problem into two subproblems, 1) the joint power and subchannel allocation and 2) UAV trajectory design subproblems. A low-complexity iterative algorithm is proposed in a alternating optimization manner over these two subproblems to solve the formulated problem. Specifically, the Lagrange dual decomposition is exploited to jointly optimize the transmit power and subchannel allocation iteratively. Then, an efficient iterative algorithm capitalizing on successive convex approximation is designed to get a suboptimal solution for UAV trajectory. Simulation results demonstrate that our proposed algorithm can significantly improve the sum transmission rate compared with the benchmark schemes.

eess.SP↗

Existence and multiplicity of normalized solutions for the generalized Kadomtsev-Petviashvili equation in $\mathbb{R}^2$

In this paper, we study the existence and {multiplicity} of nontrivial solitary waves for the generalized Kadomtsev-Petviashvili equation with prescribed {$L^2$-norm} \begin{equation*}\label{Equation1} \left\{\begin{array}{l} \left(-u_{x x}+D_x^{-2} u_{y y}+λu-f(u)\right)_x=0,{\quad x \in \mathbb{R}^2, } \\[10pt] \displaystyle \int_{\mathbb{R}^2}u^2 d x=a^2, \end{array}\right.%\tag{$\mathscr E_λ$} \end{equation*} where $a>0$ and $λ\in \mathbb{R}$ is an unknown parameter that appears as a Lagrange multiplier. For the case $f(t)=|t|^{q-2}t$, with $2 0$, we prove the existence of normalized ground state solutions which corresponds to a local minimum of the associated energy functional. In this case, we further show that there exists a sequence $(a_n) \subset (0,a_0)$ with $a_n \to 0$ as $n \to+\infty$, such that for each $a=a_n$, the problem admits a second solution with positive energy. To the best of our knowledge, this is the first work that studies the existence of solutions for the generalized Kadomtsev-Petviashvili equations under the $L^2$-constraint, which we refer to them as the normalized solutions.

math.AP↗

Increasing the density limit with ECRH-assisted Ohmic start-up on EAST

High plasma density operation is crucial for a tokamak to achieve energy breakeven and a burning plasma. However, there is often an empirical upper limit of electron density in tokamak operation, namely the Greenwald density limit $n_G$, above which tokamaks generally disrupt. Achieving high-density operations above the density limit has been a long-standing challenge in magnetic confinement fusion research. Here, we report experimental results on EAST tokamak achieving the line-averaged electron density in the range of 1.3 $n_G$ to 1.65 $n_G$,while the usual range in EAST is (0.8-1.0)$n_G$. This is performed with ECRH-assisted Ohmic start-up and a sufficiently high initial neutral density. This is motivated by and consistent with predictions of a recent plasma-wall self-organization (PWSO) theory, that increasing ECRH power or pre-filled gas pressure leads to lower plasma temperatures around divertor target and higher density limits. In addition, the experiments are shown to operate in the density-free regime predicted by the PWSO model. These results suggest a promising scheme for substantially increasing the density limit in tokamaks, a critical advancement toward achieving the burning plasma.

physics.plasm-ph↗