Searcharxiv⌕ Search

arXiv subjects

Fan Feng

Publications and source records attributed to Fan Feng.

At least 55 records · Page 3Linked to original sources

Towards Generalizable Reinforcement Learning via Causality-Guided Self-Adaptive Representations

General intelligence requires quick adaption across tasks. While existing reinforcement learning (RL) methods have made progress in generalization, they typically assume only distribution changes between source and target domains. In this paper, we explore a wider range of scenarios where not only the distribution but also the environment spaces may change. For example, in the CoinRun environment, we train agents from easy levels and generalize them to difficulty levels where there could be new enemies that have never occurred before. To address this challenging setting, we introduce a causality-guided self-adaptive representation-based approach, called CSR, that equips the agent to generalize effectively across tasks with evolving dynamics. Specifically, we employ causal representation learning to characterize the latent causal variables within the RL system. Such compact causal representations uncover the structural relationships among variables, enabling the agent to autonomously determine whether changes in the environment stem from distribution shifts or variations in space, and to precisely locate these changes. We then devise a three-step strategy to fine-tune the causal model under different scenarios accordingly. Empirical experiments show that CSR efficiently adapts to the target domains with only a few samples and outperforms state-of-the-art baselines on a wide range of scenarios, including our simulated environments, CartPole, CoinRun and Atari games.

cs.LG↗

Towards Empowerment Gain through Causal Structure Learning in Model-Based RL

In Model-Based Reinforcement Learning (MBRL), incorporating causal structures into dynamics models provides agents with a structured understanding of the environments, enabling efficient decision. Empowerment as an intrinsic motivation enhances the ability of agents to actively control their environments by maximizing the mutual information between future states and actions. We posit that empowerment coupled with causal understanding can improve controllability, while enhanced empowerment gain can further facilitate causal reasoning in MBRL. To improve learning efficiency and controllability, we propose a novel framework, Empowerment through Causal Learning (ECL), where an agent with the awareness of causal dynamics models achieves empowerment-driven exploration and optimizes its causal structure for task learning. Specifically, ECL operates by first training a causal dynamics model of the environment based on collected data. We then maximize empowerment under the causal structure for exploration, simultaneously using data gathered through exploration to update causal dynamics model to be more controllable than dense dynamics model without causal structure. In downstream task learning, an intrinsic curiosity reward is included to balance the causality, mitigating overfitting. Importantly, ECL is method-agnostic and is capable of integrating various causal discovery methods. We evaluate ECL combined with 3 causal discovery methods across 6 environments including pixel-based tasks, demonstrating its superior performance compared to other causal MBRL methods, in terms of causal discovery, sample efficiency, and asymptotic performance.

cs.AI↗

Causal Information Prioritization for Efficient Reinforcement Learning

Current Reinforcement Learning (RL) methods often suffer from sample-inefficiency, resulting from blind exploration strategies that neglect causal relationships among states, actions, and rewards. Although recent causal approaches aim to address this problem, they lack grounded modeling of reward-guided causal understanding of states and actions for goal-orientation, thus impairing learning efficiency. To tackle this issue, we propose a novel method named Causal Information Prioritization (CIP) that improves sample efficiency by leveraging factored MDPs to infer causal relationships between different dimensions of states and actions with respect to rewards, enabling the prioritization of causal information. Specifically, CIP identifies and leverages causal relationships between states and rewards to execute counterfactual data augmentation to prioritize high-impact state features under the causal understanding of the environments. Moreover, CIP integrates a causality-aware empowerment learning objective, which significantly enhances the agent's execution of reward-guided actions for more efficient exploration in complex environments. To fully assess the effectiveness of CIP, we conduct extensive experiments across 39 tasks in 5 diverse continuous control environments, encompassing both locomotion and manipulation skills learning with pixel-based and sparse reward settings. Experimental results demonstrate that CIP consistently outperforms existing RL methods across a wide range of scenarios.

cs.AI↗

Objective Moiré Pattern

Moiré patterns, typically formed by overlaying two layers of two-dimensional materials, exhibit an effective long-range periodicity that depends on the short-range periodicity of each layer and their spatial misalignment. Here, we study moiré patterns in objective structures with symmetries different from those in conventional patterns such as twisted bilayer graphene. Specifically, the mathematical descriptions for ring patterns, 2D Bravais lattice patterns, and helical patterns are derived analytically as representative examples of objective moiré patterns, using an augmented Fourier approach. Our findings reveal that the objective moiré patterns retain the symmetries of their original structures but with different parameters. In addition, we present a non-objective case, conformal moiré patterns, to demonstrate the versatility of this approach. We hope this geometric framework will provide insights for solving more complex moiré patterns and facilitate the application of moiré patterns in X-ray diffractions, wave manipulations, molecular dynamics, and other fields.

cond-mat.mtrl-sci↗

Conversational Recommender System and Large Language Model Are Made for Each Other in E-commerce Pre-sales Dialogue

E-commerce pre-sales dialogue aims to understand and elicit user needs and preferences for the items they are seeking so as to provide appropriate recommendations. Conversational recommender systems (CRSs) learn user representation and provide accurate recommendations based on dialogue context, but rely on external knowledge. Large language models (LLMs) generate responses that mimic pre-sales dialogues after fine-tuning, but lack domain-specific knowledge for accurate recommendations. Intuitively, the strengths of LLM and CRS in E-commerce pre-sales dialogues are complementary, yet no previous work has explored this. This paper investigates the effectiveness of combining LLM and CRS in E-commerce pre-sales dialogues, proposing two collaboration methods: CRS assisting LLM and LLM assisting CRS. We conduct extensive experiments on a real-world dataset of Ecommerce pre-sales dialogues. We analyze the impact of two collaborative approaches with two CRSs and two LLMs on four tasks of Ecommerce pre-sales dialogue. We find that collaborations between CRS and LLM can be very effective in some cases.

cs.CL↗

The Fused Model of Alternating Spin Chain from ABJM Theory

In this paper we give an algebraic construction of the fused model for ABJM spin chain and find the corresponding boost operator. We also investigate the open spin Hamiltonian for fused model and point out the general common structures of the boundary terms.

hep-th↗

Solid-Fluid Interaction on Particle Flow Maps

We propose a novel solid-fluid interaction method for coupling elastic solids with impulse flow maps. Our key idea is to unify the representation of fluid and solid components as particle flow maps with different lengths and dynamics. The solid-fluid coupling is enabled by implementing two novel mechanisms: first, we developed an impulse-to-velocity transfer mechanism to unify the exchanged physical quantities; second, we devised a particle path integral mechanism to accumulate coupling forces along each flow-map trajectory. Our framework integrates these two mechanisms into an Eulerian-Lagrangian impulse fluid simulator to accommodate traditional coupling models, exemplified by the Material Point Method (MPM) and Immersed Boundary Method (IBM), within a particle flow map framework. We demonstrate our method's efficacy by simulating solid-fluid interactions exhibiting strong vortical dynamics, including various vortex shedding and interaction examples across swimming, falling, breezing, and combustion.

cs.GR↗

A generalized geometric mechanics theory for multi-curve-fold origami: vertex constrained universal configurations

Folding paper along curves leads to spatial structures that have curved surfaces meeting at spatial creases, defined as curve-fold origami. In this work, we provide an Eulerian framework focusing on the mechanics of arbitrary curve-fold origami, especially for multi-curve-fold origami with vertices. We start with single-curve-fold origami that has wide panels. Wide panel leads to different domains of mechanical responses induced by various generator distributions of the curved surface. The theories are then extended to multi-curve-fold origami, involving additional geometric correlations between creases. As an illustrative example, the deformation and equilibrium configuration of origami with annular creases are studied both theoretically and numerically. Afterward, single-vertex curved origami theory is studied as a special type of multi-curve-fold origami. We find that the extra periodicity at the vertex strongly constrains the configuration space, leading to a region near the vertex that has a striking universal equilibrium configuration regardless of the mechanical properties. Both theories and numerics confirm the existence of the universality in the near-field region. In addition, the far-field deformation is obtained via energy minimization and validated by finite element analysis. Our generalized multi-curve-fold origami theory, including the vertex-contained universality, is anticipated to provide a new understanding and framework for the shape programming of the curved fold origami system.

cond-mat.soft↗

Dynamic Position Transformation and Boundary Refinement Network for Left Atrial Segmentation

Left atrial (LA) segmentation is a crucial technique for irregular heartbeat (i.e., atrial fibrillation) diagnosis. Most current methods for LA segmentation strictly assume that the input data is acquired using object-oriented center cropping, while this assumption may not always hold in practice due to the high cost of manual object annotation. Random cropping is a straightforward data pre-processing approach. However, it 1) introduces significant irregularities and incompleteness in the input data and 2) disrupts the coherence and continuity of object boundary regions. To tackle these issues, we propose a novel Dynamic Position transformation and Boundary refinement Network (DPBNet). The core idea is to dynamically adjust the relative position of irregular targets to construct their contextual relationships and prioritize difficult boundary pixels to enhance foreground-background distinction. Specifically, we design a shuffle-then-reorder attention module to adjust the position of disrupted objects in the latent space using dynamic generation ratios, such that the vital dependencies among these random cropping targets could be well captured and preserved. Moreover, to improve the accuracy of boundary localization, we introduce a dual fine-grained boundary loss with scenario-adaptive weights to handle the ambiguity of the dual boundary at a fine-grained level, promoting the clarity and continuity of the obtained results. Extensive experimental results on benchmark dataset have demonstrated that DPBNet consistently outperforms existing state-of-the-art methods.

eess.IV↗

Physics-based Machine Learning Discovered Nano-circuitry for Nonlinear Ion Transport in Nanoporous Electrodes

Confined ion transport is involved in nanoporous ionic systems. However, it is challenging to mechanistically predict its electrical characteristics for rational system design and performance evaluation using electrical circuit model due to the gap between the circuit theory and the underlying physical chemistry. Here we demonstrate that machine learning can bridge this gap and produce physics-based nano-circuitry, based on equation discovery from the modified Poisson-Nernst-Planck simulation results where an anomalous constructive diffusion-migration interplay of confined ions is unveiled. This bridging technique allows us to gain physical insights of ion dynamics in nanoporous electrodes, such as the non-ideal cyclic voltammetry.

cond-mat.mes-hall↗

YAYI 2: Multilingual Open-Source Large Language Models

As the latest advancements in natural language processing, large language models (LLMs) have achieved human-level language understanding and generation abilities in many real-world tasks, and even have been regarded as a potential path to the artificial general intelligence. To better facilitate research on LLMs, many open-source LLMs, such as Llama 2 and Falcon, have recently been proposed and gained comparable performances to proprietary models. However, these models are primarily designed for English scenarios and exhibit poor performances in Chinese contexts. In this technical report, we propose YAYI 2, including both base and chat models, with 30 billion parameters. YAYI 2 is pre-trained from scratch on a multilingual corpus which contains 2.65 trillion tokens filtered by our pre-training data processing pipeline. The base model is aligned with human values through supervised fine-tuning with millions of instructions and reinforcement learning from human feedback. Extensive experiments on multiple benchmarks, such as MMLU and CMMLU, consistently demonstrate that the proposed YAYI 2 outperforms other similar sized open-source models.

cs.CL↗

Surface instability in a nematic elastomer

Liquid crystal elastomers (LCEs) are soft phase-changing solids that exhibit large reversible contractions upon heating, Goldstone-like soft modes and resultant microstructural instabilities. We heat a planar LCE slab to isotropic, clamp the lower surface then cool back to nematic. Clamping prevents macroscopic elongation, producing compression and microstructure. We see that the free surface destabilizes, adopting topography with amplitude and wavelength similar to thickness. To understand the instability, we numerically compute the microstructural relaxation of a "non-ideal" LCE energy. Linear stability reveals a Biot-like scale-free instability, but with oblique wavevector. However, simulation and experiment show that, unlike classic elastic creasing, instability culminates in a cross-hatch without cusps or hysteresis, and is constructed entirely from low-stress soft modes.

cond-mat.soft↗

MLAnalysis: An open-source program for high energy physics analyses

We present a python-based program for phenomenological investigations in particle physics using machine learning algorithms, called \verb"MLAnalysis". The program is able to convert LHE and LHCO files generated by \verb"MadGraph5_aMC@NLO" into data sets for machine learning algorithms, which can analyze the information of the events. At present, it contains three machine learning (ML) algorithms: isolation forest (IF) algorithm, nested isolation forest (NIF) algorithm, kmeans anomaly detection (KMAD), and some basic functionality to analyze the kinematic features of a data set. Users can use this program to improve the efficiency of searching for new physics signals.

hep-ph↗

Geometry, mechanics and actuation of intrinsically curved folds

We combine theory and experiments to explore the kinematics and actuation of intrinsically curved folds (ICFs) in otherwise developable shells. Unlike origami folds, ICFs are not bending isometries of flat sheets, but arise via non-isometric processes (growth/moulding) or by joining sheets along curved boundaries. Experimentally, we implement both, first making joined ICFs from paper, then fabricating flat liquid crystal elastomer (LCE) sheets that morph into ICFs upon heating/swelling via programmed metric changes. Theoretically, an ICF's intrinsic geometry is defined by the geodesic curvatures on either side, $κ_{g_i}$. Given these, and a target 3D fold-line, one can construct the entire surface isometrically, and compute the bending energy. This construction shows ICFs are bending mechanisms, with a continuous family of isometries trading fold angle against fold-line curvature. In ICFs with symmetric $κ_{g_i}$, straightening the fold-line culminates in a fully-folded flat state that is deployable but weak, while asymmetric ICFs ultimately lock with a mechanically strong finite-angle. When unloaded, freely-hinged ICFs simply adopt the (thickness $t$ independent) isometry that minimizes the bend energy. In contrast, in LCE ICFs a competition between flank and ridge selects a ridge curvature that, unusually, scales as $t^{-1/7}$. Finally, we demonstrate how multiple ICFs can be combined in one LCE sheet, to create a versatile stretch-strong gripper that lifts $\sim$40x its own weight.

cond-mat.soft↗

Learning Dynamic Attribute-factored World Models for Efficient Multi-object Reinforcement Learning

In many reinforcement learning tasks, the agent has to learn to interact with many objects of different types and generalize to unseen combinations and numbers of objects. Often a task is a composition of previously learned tasks (e.g. block stacking). These are examples of compositional generalization, in which we compose object-centric representations to solve complex tasks. Recent works have shown the benefits of object-factored representations and hierarchical abstractions for improving sample efficiency in these settings. On the other hand, these methods do not fully exploit the benefits of factorization in terms of object attributes. In this paper, we address this opportunity and introduce the Dynamic Attribute FacTored RL (DAFT-RL) framework. In DAFT-RL, we leverage object-centric representation learning to extract objects from visual inputs. We learn to classify them in classes and infer their latent parameters. For each class of object, we learn a class template graph that describes how the dynamics and reward of an object of this class factorize according to its attributes. We also learn an interaction pattern graph that describes how objects of different classes interact with each other at the attribute level. Through these graphs and a dynamic interaction graph that models the interactions between objects, we can learn a policy that can then be directly applied in a new environment by just estimating the interactions and latent parameters. We evaluate DAFT-RL in three benchmark datasets and show our framework outperforms the state-of-the-art in generalizing across unseen objects with varying attributes and latent parameters, as well as in the composition of previously learned tasks.

cs.LG↗

U-NEED: A Fine-grained Dataset for User Needs-Centric E-commerce Conversational Recommendation

Conversational recommender systems (CRSs) aim to understand the information needs and preferences expressed in a dialogue to recommend suitable items to the user. Most of the existing conversational recommendation datasets are synthesized or simulated with crowdsourcing, which has a large gap with real-world scenarios. To bridge the gap, previous work contributes a dataset E-ConvRec, based on pre-sales dialogues between users and customer service staff in E-commerce scenarios. However, E-ConvRec only supplies coarse-grained annotations and general tasks for making recommendations in pre-sales dialogues. Different from that, we use real user needs as a clue to explore the E-commerce conversational recommendation in complex pre-sales dialogues, namely user needs-centric E-commerce conversational recommendation (UNECR). In this paper, we construct a user needs-centric E-commerce conversational recommendation dataset (U-NEED) from real-world E-commerce scenarios. U-NEED consists of 3 types of resources: (i) 7,698 fine-grained annotated pre-sales dialogues in 5 top categories (ii) 333,879 user behaviors and (iii) 332,148 product knowledge tuples. To facilitate the research of UNECR, we propose 5 critical tasks: (i) pre-sales dialogue understanding (ii) user needs elicitation (iii) user needs-based recommendation (iv) pre-sales dialogue generation and (v) pre-sales dialogue evaluation. We establish baseline methods and evaluation metrics for each task. We report experimental results of 5 tasks on U-NEED. We also report results in 3 typical categories. Experimental results indicate that the challenges of UNECR in various categories are different.

cs.IR↗

Moment-based space-variant Shack-Hartmann wavefront reconstruction

Based on image moment theory, an approach for space-variant Shack-Hartmann wavefront reconstruction is presented in this article. The relation between the moment of a pair of subimages and the local transformation coefficients is derived. The square guide 'star' is used to obtain a special solution from this relation. The moment-based wavefront reconstruction has a reduced computational complexity compared to the iteration-based algorithm. Image restorations are executed by the tiling strategy with 5 $\times$ 5 PSFs as well as the conventional strategy with a global average PSF. Visual and quantitative evaluations support our approach.

physics.optics↗

Factored Adaptation for Non-Stationary Reinforcement Learning

Dealing with non-stationarity in environments (e.g., in the transition dynamics) and objectives (e.g., in the reward functions) is a challenging problem that is crucial in real-world applications of reinforcement learning (RL). While most current approaches model the changes as a single shared embedding vector, we leverage insights from the recent causality literature to model non-stationarity in terms of individual latent change factors, and causal graphs across different environments. In particular, we propose Factored Adaptation for Non-Stationary RL (FANS-RL), a factored adaption approach that learns jointly both the causal structure in terms of a factored MDP, and a factored representation of the individual time-varying change factors. We prove that under standard assumptions, we can completely recover the causal graph representing the factored transition and reward function, as well as a partial structure between the individual change factors and the state components. Through our general framework, we can consider general non-stationary scenarios with different function types and changing frequency, including changes across episodes and within episodes. Experimental results demonstrate that FANS-RL outperforms existing approaches in terms of return, compactness of the latent state representation, and robustness to varying degrees of non-stationarity.

cs.LG↗