SearcharxivSearch

arXiv subjects

Jiamin He

Publications and source records attributed to Jiamin He.

10 recordsLinked to original sources

Well-posedness and Stability Analysis of Suspension Bridge Models Coupled with Cattaneo Heat Conduction: The Role of Viscoelastic Memory

This paper focuses on investigating a class of suspension bridge models coupled with the Cattaneo heat conduction law, with special attention paid to two distinct scenarios: systems with viscoelastic memory and those without. Operator semigroup theory is adopted as the core mathematical tool to conduct a comprehensive analysis of the models' dynamic behaviors. For the memoryless suspension bridge system ($\delta = 0$), we first establish its well-posedness (i.e., the existence and uniqueness of solutions). Further stability analysis reveals that this system exhibits polynomial decay behavior, where the decay rate depends on the structural parameter relationship of the bridge deck: a decay rate of $t^{-1/2}$ is achieved when $\frac{\rho_1}{\kappa} = \frac{\rho_2}{b}$, and the rate slows down to $t^{-1/4}$ when $\frac{\rho_1}{\kappa} \neq \frac{\rho_2}{b}$. Moreover, exponential stability is proven to be unachievable under specific coefficient configurations (i.e., $\chi_0 \neq 0$, or $\chi_0 = 0$ and $\chi_1 = 0$, with $\chi_0, \chi_1$ being parameter-dependent coefficients). For the counterpart system incorporating viscoelastic memory ($\delta = 1$), the well-posedness of solutions is also verified, and more importantly, the system is shown to achieve exponential decay. This indicates that viscoelastic memory significantly enhances system stability and accelerates energy dissipation compared to the memoryless case. By systematically exploring the regulatory role of viscoelastic memory in the stability of suspension bridge systems under the framework of Cattaneo thermal conduction, this study enriches and extends the existing research on suspension bridge models with viscoelastic memory.

math.AP

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic

Mixture policies theoretically offer greater flexibility than unimodal policies in continuous action reinforcement learning, but the practical benefits of this complexity remain elusive. Mixture policies are notably absent from most state-of-the-art algorithms, raising a fundamental question: Is the added representational overhead useful? We show that increased flexibility can theoretically enhance solution quality and entropy robustness. Yet standard algorithms like SAC do not leverage these advantages. A core issue is the lack of a low-variance reparameterization trick for mixtures, a luxury Gaussian policies enjoy. We propose a marginalized reparameterization (MRP) estimator to address this, proving it offers lower variance than the standard likelihood-ratio (LR) approach. Our experiments across Gym MuJoCo, DeepMind Control Suite, and MetaWorld show that MRP mixture policies significantly outperform their LR ones, and reach parity (sometimes better) with Gaussian counterparts. In addition, we do find several cases where MRP mixture policies exhibit clear empirical advantages. In this paper, we provide a clearer understanding of the trade-offs involved, elevating MRP mixture policies from theoretical curiosity to a practical tool.

cs.LG

Extending Differential Temporal Difference Methods for Episodic Problems

Differential temporal difference (TD) methods are value-based reinforcement learning algorithms that have been proposed for infinite-horizon problems. They rely on reward centering, where each reward is centered by the average reward. This keeps the return bounded and removes a value function's state-independent offset. However, reward centering can alter the optimal policy in episodic problems, limiting its applicability. Motivated by recent works that emphasize the role of normalization in streaming deep reinforcement learning, we study reward centering in episodic problems and propose a generalization of differential TD. We prove that this generalization maintains the ordering of policies in the presence of termination, and thus extends differential TD to episodic problems. We show equivalence with a form of linear TD, thereby inheriting theoretical guarantees that have been shown for those algorithms. We then extend several streaming reinforcement learning algorithms to their differential counterparts. Across a range of base algorithms and environments, we empirically validate that reward centering can improve sample efficiency in episodic problems.

cs.LG

Forager: a lightweight testbed for continual learning with partial observability in RL

In continual reinforcement learning (CRL), good performance requires never-ending learning, acting, and exploration in a big, partially observable world. Most CRL experiments have focused on loss of plasticity -- the inability to keep learning -- in one-off experiments where some unobservable non-stationarity is added to classic fully observable MDPs. Further, these experiments rarely consider the role of partial observability and the importance of CRL agents that use memory or recurrence. One potential reason for this focus on mitigating loss of plasticity without considering partial observability is that many partially-observable CRL environments are prohibitively expensive. In this paper, we introduce Forager, a light-weight partially-observable CRL environment with a constant memory footprint. We provide a set of experiments and sample tasks demonstrating that Forager is challenging for current CRL agents and yet also allows for in-depth study of those agents. We demonstrate that agents exhibit loss of plasticity, proposed mitigations can help, but that most useful is to leverage state construction. We conclude with a variant of Forager that generates an unending stream of new tasks to learn that clearly highlights the limitations of current CRL agents.

cs.LG

Suitable sets for topological groups revisited

A discrete subset $S$ of a topological group $G$ is called a {\it suitable set} for $G$ if $S\cup \{e\}$ is closed in $G$ and the subgroup generated by $S$ is dense in $G$, where $e$ is the identity element of $G$. In this paper, the existence of suitable sets in topological groups is studied. It is proved that, for a non-separable $k_{\omega}$-space $X$ without non-trivial convergent sequences, the $snf$-countability of $A(X)$ implies that $A(X)$ does not have a suitable set, which gives a partial answer to \cite[Problem 2.1]{TKA1997}. Moreover, the existence of suitable sets in some particular classes of linearly orderable topological groups is considered, where Theorem~\ref{t4} provides an affirmative answer to \cite[Problem 4.3]{ST2002}. Then, topological groups with an $\omega^{\omega}$-base are discussed, and every linearly orderable topological group with an $\omega^{\omega}$-base being metrizable is proved; thus it has a suitable set. Further, it follows that each topological group $G$ with an $\omega^{\omega}$-base has a suitable set whenever $G$ is a $k$-space, which gives a generalization of a well-known result in \cite{CM}. Finally, some cardinal invariant of topological groups with a suitable set are provided. Some results of this paper give some partial answers to some open problems posed in~\cite{DTA} and~\cite{TKA1997} respectively.

math.GN

The existence of suitable sets in locally compact strongly topological gyrogroups

A subset $S$ of a topological gyrogroup $G$ is said to be a {\it suitable set} for $G$ if $S$ is discrete, the gyrogroup generated by $S$ is dense in $G$, and $S\cup \{0\}$ is closed in $G$, where $0$ is the identity element of $G$. In this paper, it is proved that every locally compact strongly topological gyrogroup has a suitable set, which gives an affirmative answer to a question posed by F. Lin, et al. in \cite{key14}.

math.GN

Strongly topologically orderable gyrogroups with a suitable set

A discrete subset $S$ of a topologically gyrogroup $G$ is called a {\it suitable set} for $G$ if $S\cup \{1\}$ is closed and the subgyrogroup generated by $S$ is dense in $G$, where $1$ is the identity element of $G$. In this paper, we mainly study the existence of suitable set of strongly topologically orderable gyrogroups, which extends some result in some papers in the literature. In particular, the existences of suitable set of each locally compact or not totally disconnected strongly topologically orderable gyrogroup are affirmative.

math.GN

Distributions as Actions: A Unified Framework for Diverse Action Spaces

We introduce a novel reinforcement learning (RL) framework that treats parameterized action distributions as actions, redefining the boundary between agent and environment. This reparameterization makes the new action space continuous, regardless of the original action type (discrete, continuous, hybrid, etc.). Under this new parameterization, we develop a generalized deterministic policy gradient estimator, Distributions-as-Actions Policy Gradient (DA-PG), which has lower variance than the gradient in the original action space. Although learning the critic over distribution parameters poses new challenges, we introduce Interpolated Critic Learning (ICL), a simple yet effective strategy to enhance learning, supported by insights from bandit settings. Building on TD3, a strong baseline for continuous control, we propose a practical actor-critic algorithm, Distributions-as-Actions Actor-Critic (DA-AC). Empirically, DA-AC achieves competitive performance in various settings across discrete, continuous, and hybrid control.

cs.LG

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

Modern deep policy gradient methods achieve effective performance on simulated robotic tasks, but they all require large replay buffers or expensive batch updates, or both, making them incompatible for real systems with resource-limited computers. We show that these methods fail catastrophically when limited to small replay buffers or during incremental learning, where updates only use the most recent sample without batch updates or a replay buffer. We propose a novel incremental deep policy gradient method -- Action Value Gradient (AVG) and a set of normalization and scaling techniques to address the challenges of instability in incremental learning. On robotic simulation benchmarks, we show that AVG is the only incremental method that learns effectively, often achieving final performance comparable to batch policy gradient methods. This advancement enabled us to show for the first time effective deep reinforcement learning with real robots using only incremental updates, employing a robotic manipulator and a mobile robot.

cs.LG

Episodic Multi-agent Reinforcement Learning with Curiosity-Driven Exploration

Efficient exploration in deep cooperative multi-agent reinforcement learning (MARL) still remains challenging in complex coordination problems. In this paper, we introduce a novel Episodic Multi-agent reinforcement learning with Curiosity-driven exploration, called EMC. We leverage an insight of popular factorized MARL algorithms that the "induced" individual Q-values, i.e., the individual utility functions used for local execution, are the embeddings of local action-observation histories, and can capture the interaction between agents due to reward backpropagation during centralized training. Therefore, we use prediction errors of individual Q-values as intrinsic rewards for coordinated exploration and utilize episodic memory to exploit explored informative experience to boost policy training. As the dynamics of an agent's individual Q-value function captures the novelty of states and the influence from other agents, our intrinsic reward can induce coordinated exploration to new or promising states. We illustrate the advantages of our method by didactic examples, and demonstrate its significant outperformance over state-of-the-art MARL baselines on challenging tasks in the StarCraft II micromanagement benchmark.

cs.LG