Searcharxiv⌕ Search

arXiv subjects

Sambuddha Chakrabarti

Publications and source records attributed to Sambuddha Chakrabarti.

5 recordsLinked to original sources

Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments

Autonomous AI agents operating in dynamic environments face a persistent challenge: acquiring new capabilities without erasing prior knowledge. We present Adaptive Memory Crystallization (AMC), a memory architecture for progressive experience consolidation in continual reinforcement learning. AMC is conceptually inspired by the qualitative structure of synaptic tagging and capture (STC) theory, the idea that memories transition through discrete stability phases, but makes no claim to model the underlying molecular or synaptic mechanisms. AMC models memory as a continuous crystallization process in which experiences migrate from plastic to stable states according to a multi-objective utility signal. The framework introduces a three-phase memory hierarchy (Liquid--Glass--Crystal) governed by an Itô stochastic differential equation (SDE) whose population-level behavior is captured by an explicit Fokker--Planck equation admitting a closed-form Beta stationary distribution. We provide proofs of: (i) well-posedness and global convergence of the crystallization SDE to a unique Beta stationary distribution; (ii) exponential convergence of individual crystallization states to their fixed points, with explicit rates and variance bounds; and (iii) end-to-end Q-learning error bounds and matching memory-capacity lower bounds that link SDE parameters directly to agent performance. Empirical evaluation on Meta-World MT50, Atari 20-game sequential learning, and MuJoCo continual locomotion consistently shows improvements in forward transfer (+34--43\% over the strongest baseline), reductions in catastrophic forgetting (67--80\%), and a 62\% decrease in memory footprint.

cs.LG↗

Analyzing Performance and Scalability of Benders Decomposition for Generation and Transmission Expansion Planning Models

Generation and Transmission Expansion Planning (GTEP) problems co-optimize generation and transmission expansion, enabling them to provide better planning decisions than traditional Generation Expansion Planning or Transmission Expansion Planning problems, but GTEPs can be computationally complex or intractable. Benders Decomposition (BD) has been applied to expansion planning problems, with various methods applied to accelerate convergence. In this work, we test strategies for improving the performance of BD on GTEP models with nodal resolution and DCOPF constraints. We also present an alternative approach for handling the bilinear constraints that can result in these problems. These tests included combinations of using generalized Benders decomposition (GBD), hot-starting via a transport constrained model, using linear relaxations of the master problem, and using regularization. We test these methods on mixed-integer linear programming GTEP models with up to 146 buses (10 million continuous variables and 400 mixed-integer decisions). With selected accelerated Benders decomposition approaches, the problems can be solved to under a 1\% gap in as little as 5 hours where they were otherwise intractable. Results also suggest that using regularization on these initial hot-starting and relaxation steps and turning it off after they are complete was generally the best combination of strategies.

math.OC↗

Extending Group Relative Policy Optimization to Continuous Control: A Theoretical Framework for Robotic Reinforcement Learning

Group Relative Policy Optimization (GRPO) has shown promise in discrete action spaces by eliminating value function dependencies through group-based advantage estimation. However, its application to continuous control remains unexplored, limiting its utility in robotics where continuous actions are essential. This paper presents a theoretical framework extending GRPO to continuous control environments, addressing challenges in high-dimensional action spaces, sparse rewards, and temporal dynamics. Our approach introduces trajectory-based policy clustering, state-aware advantage estimation, and regularized policy updates designed for robotic applications. We provide theoretical analysis of convergence properties and computational complexity, establishing a foundation for future empirical validation in robotic systems including locomotion and manipulation tasks.

cs.RO↗

Transmission Investment Coordination using MILP Lagrange Dual Decomposition and Auxiliary Problem Principle

This paper considers the investment coordination problem for the long term transmission capacity expansion in a situation where there are multiple regional Transmission Planners (TPs), each acting in order to maximize the utility in only its own region. In such a setting, any particular TP does not normally have any incentive to cooperate with the neighboring TP(s), although the optimal investment decision of each TP is contingent upon those of the neighboring TPs. A game-theoretic interaction among the TPs does not necessarily lead to this overall social optimum. We, therefore, introduce a social planner and call it the Transmission Planning Coordinator (TPC) whose goal is to attain the optimal possible social welfare for the bigger geographical region. In order to achieve this goal, this paper introduces a new incentive mechanism, based on distributed optimization theory. This incentive mechanism can be viewed as a set of rules of the transmission expansion investment coordination game, set by the social planner TPC, such that, even if the individual TPs act selfishly, it will still lead to the TPC's goal of attaining overall social optimum. Finally, the effectiveness of our approach is demonstrated through several simulation studies.

eess.SY↗

Look-Ahead SCOPF (LASCOPF) for Tracking Demand Variation via Auxiliary Proximal Message Passing (APMP) Algorithm

In this paper, we will consider the Look-Ahead Security Constrained Optimal Power Flow (LASCOPF) problem looking forward multiple dispatch intervals, in which the load demand varies over dispatch intervals according to some forecast. We will consider the base-case and several contingency scenarios in the upcoming as well as in the subsequent dispatch intervals. We will formulate and solve the problem in a Model Predictive Control (MPC) paradigm. We will present the Auxiliary Proximal Message Passing (APMP) algorithm to solve this problem, which is a bi-layered decomposition-coordination type distributed algorithm, consisting of an outer Auxiliary Problem Principle (APP) layer and an inner Proximal Message Passing (PMP) layer. The APP part of the algorithm distributes the computation across several dispatch intervals and the PMP part performs the distributed computation within each of the dispatch interval across different devices (i.e. generators, transmission lines, loads) and nodes or nets. We will demonstrate the effectiveness of our method with a series of numerical simulations.

eess.SY↗