Searcharxiv⌕ Search

arXiv subjects

Abhishek Naik

Publications and source records attributed to Abhishek Naik.

15 recordsLinked to original sources

Dynamic Object Masks as Goal Representations for Visual Goal-Conditioned Reinforcement Learning

Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-based goal representations that are unavailable in real-world robotics. Robots operating in warehouses, agriculture, or laboratory environments rarely have access to privileged goal states, object positions, or future observations, limiting the practicality of current GCRL approaches. We propose a dynamic mask-based goal representation that provides simple, object-agnostic visual cues for vision-based navigation and manipulation. At each timestep, an image-based goal detector produces a goal mask using standard image processing, task-specific object recognizers, or pretrained detectors such as Detic or Grounding DINO. These masks specify the spatial target without requiring privileged goal states, enabling broad applicability and strong generalization to unseen objects. Our method improves stability and sample efficiency in GCRL, achieving a 99% success rate in reaching both training and novel objects with Franka and UR10e robotic arms and faster learning in simulated navigation. We further demonstrate learning from scratch and sim-to-real transfer on both robotic arms, highlighting the effectiveness of our approach for real-world RL tasks. Our code is available at https://github.com/fahimfss/GCRL.

cs.CV↗

Rethinking the Suitability of Reinforcement Learning Algorithms Under Practical Transfer Constraints

Transfer-oriented reinforcement learning requires evaluating algorithms along dimensions that go beyond standard sample efficiency. We focus on two dimensions: practical efficiency, which asks whether conclusions about algorithm suitability change under wall-clock rather than interaction-based budgets, and robustness under dynamics mismatch, which asks how different learning paradigms respond to variability in the training distribution induced by domain randomization. We provide two insights to reinforcement-learning practitioners. First, comparing the sample efficiency of different algorithms is often an insufficient criterion in transfer-oriented settings. The wall-clock time required to train a decent policy is an important consideration for practitioners, and we find that the sample-inefficient PPO algorithm can produce a performant policy faster than relatively more sample-efficient algorithms such as SAC and TD-MPC2, validating the common understanding of massively parallel training paradigms. Second, domain randomization can help different kinds of algorithms learn robust policies. In particular, although PPO, SAC, and TD-MPC2 represent different RL paradigms - on-policy, off-policy, and model-based learning and planning, respectively - we find that domain randomization affects all three algorithms in a similar way. To the best of our knowledge, this is the first controlled comparison of the effect of domain-randomization coverage on PPO, SAC, and TD-MPC2 under the same transfer protocol. Taken together, these two insights highlight the importance of evaluating RL algorithms not only by sample efficiency, but also by practical considerations such as training time and the algorithms' ability to produce usable policies.

cs.LG↗

Benchmarking Action Spaces in Reinforcement Learning for Vision-based Robotic Manipulation

In real-world reinforcement learning (RL), the choice of action space can play a key role in shaping motion smoothness, safety, and overall task performance. In this study, we evaluate pose increment, pose velocity, joint position increment, and joint velocity across two vision-based manipulation tasks: object picking and pushing. We train policies in simulation and deploy them to the real world using sim-to-real transfer. We find that action-space representation indeed significantly affects sim-to-real performance. In particular, we find that the joint velocity action space is best for the vision-based picking and pushing tasks in terms of smoothness and final task performance. We also provide practical guidance for RL practitioners in choosing action spaces for both simulation and real-world experiments.

cs.RO↗

Engineering Magnetic Anisotropy in Permalloy Films via Atomic Force Nanolithography

Atomic force nanolithography provides a precise method for sculpting magnetic thin films, enabling controlled engineering of magnetic anisotropy in soft ferromagnets at the microscale. We demonstrate that nanoscale groove arrays patterned into permalloy (Ni80Fe20) films induce a robust in-plane uniaxial anisotropy, with the easy axis aligned along the groove direction. The anisotropy field is shown to increase with decreasing groove period and increasing engraving depth, offering continuous tunability of magnetic hardness within a single fabrication step. Artificially engraved microstructures further allow domain configurations and domain-wall trajectories to be directed along predefined pathways, exemplified by the creation of a chessboard-like magnetic landscape. Owing to its adaptability to diverse ferromagnetic materials and arbitrary corrugation geometries, this approach provides a versatile platform for tailoring in-plane magnetic anisotropy. Concrete applications are demonstrated in the design of magnonic elements and anisotropic magnetoresistance sensors.

cond-mat.mtrl-sci↗

Energy Efficient Traffic Scheduling For Optical LEO Satellite Downlinks

In recent years, the number of satellites in orbit has increased rapidly, with megaconstellations like Starlink providing near-global, delay-sensitive communication services. However, not all satellite communication use cases have stringent delay requirements; services such as Earth observation (EO) and remote Internet of Things (IoT) fall into this category. These relaxed delay quality of service (QoS) objectives allow services to be delivered using sparse constellations, enabled by delay-tolerant networking protocols. In the context of rapidly growing data volumes that must be delivered through satellite networks, a key challenge is having sufficient space-to-ground link capacity. This has led to proposals for using free-space optical (FSO) communications, which offer high data rates. However, FSO communications are highly vulnerable to weather-related disruptions. This results in certain communication opportunities being energy inefficient. Given the energy-constrained nature of satellites, developing schemes to improve energy efficiency is highly desirable. In this work, both static and adaptive schemes were developed to balance maintaining the delivery ratio and maximizing energy efficiency. The proposed schemes fall into the following categories: threshold schemes, heuristic sorting algorithms, and reinforcement learning-based schemes. The schemes were evaluated under a variety of different data volumes and cloud cover distribution configurations as well as a case study using historical weather data. It was found that static schemes suffered from low delivery ratio performance under dynamic conditions when compared to adaptive techniques. However, this performance improvement came at the cost of increased complexity and onboard computations.

cs.NI↗

DSROQ: Dynamic Scheduling and Routing for QoE Management in LEO Satellite Networks

The modern Internet supports diverse applications with heterogeneous quality of service (QoS) requirements. Low Earth orbit (LEO) satellite constellations offer a promising solution to meet these needs, enhancing coverage in rural areas and complementing terrestrial networks in urban regions. Ensuring QoS in such networks requires joint optimization of routing, bandwidth allocation, and dynamic queue scheduling, as traffic handling is critical for maintaining service performance. This paper formulates a joint routing and bandwidth allocation problem where QoS requirements are treated as soft constraints, aiming to maximize user experience. An adaptive scheduling approach is introduced to prioritize flow-specific QoS needs. We propose a Monte Carlo tree search (MCTS)-inspired method to solve the NP-hard route and bandwidth allocation problem, with Lyapunov optimization-based scheduling applied during reward evaluation. Using the Starlink Phase 1 Version 2 constellation, we compare end-user experience and fairness between our proposed DSROQ algorithm and a benchmark scheme. Results show that DSROQ improves both performance metrics and demonstrates the advantage of joint routing and bandwidth decisions. Furthermore, we observe that the dominant performance factor shifts from scheduling to routing and bandwidth allocation as traffic sensitivity changes from latency-driven to bandwidth-driven.

cs.NI↗

Energy-Efficient Satellite IoT Optical Downlinks Using Weather-Adaptive Reinforcement Learning

Internet of Things (IoT) devices have become increasingly ubiquitous with applications not only in urban areas but remote areas as well. These devices support industries such as agriculture, forestry, and resource extraction. Due to the device location being in remote areas, satellites are frequently used to collect and deliver IoT device data to customers. As these devices become increasingly advanced and numerous, the amount of data produced has rapidly increased potentially straining the ability for radio frequency (RF) downlink capacity. Free space optical communications with their wide available bandwidths and high data rates are a potential solution, but these communication systems are highly vulnerable to weather-related disruptions. This results in certain communication opportunities being inefficient in terms of the amount of data received versus the power expended. In this paper, we propose a deep reinforcement learning (DRL) method using Deep Q-Networks that takes advantage of weather condition forecasts to improve energy efficiency while delivering the same number of packets as schemes that don't factor weather into routing decisions. We compare this method with simple approaches that utilize simple cloud cover thresholds to improve energy efficiency. In testing the DRL approach provides improved median energy efficiency without a significant reduction in median delivery ratio. Simple cloud cover thresholds were also found to be effective but the thresholds with the highest energy efficiency had reduced median delivery ratio values.

cs.NI↗

Reward Centering

We show that discounted methods for solving continuing reinforcement learning problems can perform significantly better if they center their rewards by subtracting out the rewards' empirical average. The improvement is substantial at commonly used discount factors and increases further as the discount factor approaches one. In addition, we show that if a problem's rewards are shifted by a constant, then standard methods perform much worse, whereas methods with reward centering are unaffected. Estimating the average reward is straightforward in the on-policy setting; we propose a slightly more sophisticated method for the off-policy setting. Reward centering is a general idea, so we expect almost every reinforcement-learning algorithm to benefit by the addition of reward centering.

cs.LG↗

Average-Reward Learning and Planning with Options

We extend the options framework for temporal abstraction in reinforcement learning from discounted Markov decision processes (MDPs) to average-reward MDPs. Our contributions include general convergent off-policy inter-option learning algorithms, intra-option algorithms for learning values and models, as well as sample-based planning variants of our learning algorithms. Our algorithms and convergence proofs extend those recently developed by Wan, Naik, and Sutton. We also extend the notion of option-interrupting behavior from the discounted to the average-reward formulation. We show the efficacy of the proposed algorithms with experiments on a continuing version of the Four-Room domain.

cs.LG↗

Learning and Planning in Average-Reward Markov Decision Processes

We introduce learning and planning algorithms for average-reward MDPs, including 1) the first general proven-convergent off-policy model-free control algorithm without reference states, 2) the first proven-convergent off-policy model-free prediction algorithm, and 3) the first off-policy learning algorithm that converges to the actual value function rather than to the value function plus an offset. All of our algorithms are based on using the temporal-difference error rather than the conventional error when updating the estimate of the average reward. Our proof techniques are a slight generalization of those by Abounadi, Bertsekas, and Borkar (2001). In experiments with an Access-Control Queuing Task, we show some of the difficulties that can arise when using methods that rely on reference states and argue that our new algorithms can be significantly easier to use.

cs.LG↗

Planning with Expectation Models for Control

In model-based reinforcement learning (MBRL), Wan et al. (2019) showed conditions under which the environment model could produce the expectation of the next feature vector rather than the full distribution, or a sample thereof, with no loss in planning performance. Such expectation models are of interest when the environment is stochastic and non-stationary, and the model is approximate, such as when it is learned using function approximation. In these cases a full distribution model may be impractical and a sample model may be either more expensive computationally or of high variance. Wan et al. considered only planning for prediction to evaluate a fixed policy. In this paper, we treat the control case - planning to improve and find a good approximate policy. We prove that planning with an expectation model must update a state-value function, not an action-value function as previously suggested (e.g., Sorg & Singh, 2010). This opens the question of how planning influences action selections. We consider three strategies for this and present general MBRL algorithms for each. We identify the strengths and weaknesses of these algorithms in computational experiments. Our algorithms and experiments are the first to treat MBRL with expectation models in a general setting.

cs.AI↗

MADRaS : Multi Agent Driving Simulator

In this work, we present MADRaS, an open-source multi-agent driving simulator for use in the design and evaluation of motion planning algorithms for autonomous driving. MADRaS provides a platform for constructing a wide variety of highway and track driving scenarios where multiple driving agents can train for motion planning tasks using reinforcement learning and other machine learning algorithms. MADRaS is built on TORCS, an open-source car-racing simulator. TORCS offers a variety of cars with different dynamic properties and driving tracks with different geometries and surface properties. MADRaS inherits these functionalities from TORCS and introduces support for multi-agent training, inter-vehicular communication, noisy observations, stochastic actions, and custom traffic cars whose behaviours can be programmed to simulate challenging traffic conditions encountered in the real world. MADRaS can be used to create driving tasks whose complexities can be tuned along eight axes in well-defined steps. This makes it particularly suited for curriculum and continual learning. MADRaS is lightweight and it provides a convenient OpenAI Gym interface for independent control of each car. Apart from the primitive steering-acceleration-brake control mode of TORCS, MADRaS offers a hierarchical track-position -- speed control that can potentially be used to achieve better generalization. MADRaS uses multiprocessing to run each agent as a parallel process for efficiency and integrates well with popular reinforcement learning libraries like RLLib.

cs.RO↗

Discounted Reinforcement Learning Is Not an Optimization Problem

Discounted reinforcement learning is fundamentally incompatible with function approximation for control in continuing tasks. It is not an optimization problem in its usual formulation, so when using function approximation there is no optimal policy. We substantiate these claims, then go on to address some misconceptions about discounting and its connection to the average reward formulation. We encourage researchers to adopt rigorous optimization approaches, such as maximizing average reward, for reinforcement learning in continuing tasks.

cs.AI↗

RAIL: Risk-Averse Imitation Learning

Imitation learning algorithms learn viable policies by imitating an expert's behavior when reward signals are not available. Generative Adversarial Imitation Learning (GAIL) is a state-of-the-art algorithm for learning policies when the expert's behavior is available as a fixed set of trajectories. We evaluate in terms of the expert's cost function and observe that the distribution of trajectory-costs is often more heavy-tailed for GAIL-agents than the expert at a number of benchmark continuous-control tasks. Thus, high-cost trajectories, corresponding to tail-end events of catastrophic failure, are more likely to be encountered by the GAIL-agents than the expert. This makes the reliability of GAIL-agents questionable when it comes to deployment in risk-sensitive applications like robotic surgery and autonomous driving. In this work, we aim to minimize the occurrence of tail-end events by minimizing tail risk within the GAIL framework. We quantify tail risk by the Conditional-Value-at-Risk (CVaR) of trajectories and develop the Risk-Averse Imitation Learning (RAIL) algorithm. We observe that the policies learned with RAIL show lower tail-end risk than those of vanilla GAIL. Thus the proposed RAIL algorithm appears as a potent alternative to GAIL for improved reliability in risk-sensitive applications.

cs.LG↗

Identifying User Survival Types via Clustering of Censored Social Network Data

The goal of cluster analysis in survival data is to identify clusters that are decidedly associated with the survival outcome. Previous research has explored this problem primarily in the medical domain with relatively small datasets, but the need for such a clustering methodology could arise in other domains with large datasets, such as social networks. Concretely, we wish to identify different survival classes in a social network by clustering the users based on their lifespan in the network. In this paper, we propose a decision tree based algorithm that uses a global normalization of $p$-values to identify clusters with significantly different survival distributions. We evaluate the clusters from our model with the help of a simple survival prediction task and show that our model outperforms other competing methods.

cs.SI↗