SearcharxivSearch

arXiv subjects

Gaurav Chaudhary

Publications and source records attributed to Gaurav Chaudhary.

At least 19 recordsLinked to original sources

Match or Replay: Self Imitating Proximal Policy Optimization

Reinforcement Learning (RL) agents often struggle with inefficient exploration, particularly in environments with sparse rewards. Traditional exploration strategies can lead to slow learning and suboptimal performance because agents fail to systematically build on previously successful experiences, thereby reducing sample efficiency. To tackle this issue, we propose a self-imitating on-policy algorithm that enhances exploration and sample efficiency by leveraging past high-reward state-action pairs to guide policy updates. Our method incorporates self-imitation by using optimal transport distance in dense reward environments to prioritize state visitation distributions that match the most rewarding trajectory. In sparse-reward environments, we uniformly replay successful self-encountered trajectories to facilitate structured exploration. Experimental results across diverse environments demonstrate substantial improvements in learning efficiency, including MuJoCo for dense rewards and the partially observable 3D Animal-AI Olympics and multi-goal PointMaze for sparse rewards. Our approach achieves faster convergence and significantly higher success rates compared to state-of-the-art self-imitating RL baselines. These findings underscore the potential of self-imitation as a robust strategy for enhancing exploration in RL, with applicability to more complex tasks.

cs.LG

TEACH: Temporal Variance-Driven Curriculum for Reinforcement Learning

Reinforcement Learning (RL) has achieved significant success in solving single-goal tasks. However, uniform goal selection often results in sample inefficiency in multi-goal settings where agents must learn a universal goal-conditioned policy. Inspired by the adaptive and structured learning processes observed in biological systems, we propose a novel Student-Teacher learning paradigm with a Temporal Variance-Driven Curriculum to accelerate Goal-Conditioned RL. In this framework, the teacher module dynamically prioritizes goals with the highest temporal variance in the policy's confidence score, parameterized by the state-action value (Q) function. The teacher provides an adaptive and focused learning signal by targeting these high-uncertainty goals, fostering continual and efficient progress. We establish a theoretical connection between the temporal variance of Q-values and the evolution of the policy, providing insights into the method's underlying principles. Our approach is algorithm-agnostic and integrates seamlessly with existing RL frameworks. We demonstrate this through evaluation across 11 diverse robotic manipulation and maze navigation tasks. The results show consistent and notable improvements over state-of-the-art curriculum learning and goal-selection methods.

cs.LG

MOORL: A Framework for Integrating Offline-Online Reinforcement Learning

Sample efficiency and exploration remain critical challenges in Deep Reinforcement Learning (DRL), particularly in complex domains. Offline RL, which enables agents to learn optimal policies from static, pre-collected datasets, has emerged as a promising alternative. However, offline RL is constrained by issues such as out-of-distribution (OOD) actions that limit policy performance and generalization. To overcome these limitations, we propose Meta Offline-Online Reinforcement Learning (MOORL), a hybrid framework that unifies offline and online RL for efficient and scalable learning. While previous hybrid methods rely on extensive design components and added computational complexity to utilize offline data effectively, MOORL introduces a meta-policy that seamlessly adapts across offline and online trajectories. This enables the agent to leverage offline data for robust initialization while utilizing online interactions to drive efficient exploration. Our theoretical analysis demonstrates that the hybrid approach enhances exploration by effectively combining the complementary strengths of offline and online data. Furthermore, we demonstrate that MOORL learns a stable Q-function without added complexity. Extensive experiments on 28 tasks from the D4RL and V-D4RL benchmarks validate its effectiveness, showing consistent improvements over state-of-the-art offline and hybrid RL baselines. With minimal computational overhead, MOORL achieves strong performance, underscoring its potential for practical applications in real-world scenarios.

cs.LG

From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning

Offline Reinforcement Learning (RL) aims to learn effective policies from a static dataset without requiring further agent-environment interactions. However, its practical adoption is often hindered by the need for explicit reward annotations, which can be costly to engineer or difficult to obtain retrospectively. To address this, we propose ReLOAD (Reinforcement Learning with Offline Reward Annotation via Distillation), a novel reward annotation framework for offline RL. Unlike existing methods that depend on complex alignment procedures, our approach adapts Random Network Distillation (RND) to generate intrinsic rewards from expert demonstrations using a simple yet effective embedding discrepancy measure. First, we train a predictor network to mimic a fixed target network's embeddings based on expert state transitions. Later, the prediction error between these networks serves as a reward signal for each transition in the static dataset. This mechanism provides a structured reward signal without requiring handcrafted reward annotations. We provide a formal theoretical construct that offers insights into how RND prediction errors effectively serve as intrinsic rewards by distinguishing expert-like transitions. Experiments on the D4RL benchmark demonstrate that ReLOAD enables robust offline policy learning and achieves performance competitive with traditional reward-annotated methods.

cs.LG

Spontaneous Twirls and Structural Frustration in Moiré Materials

Structural twirls form spontaneously in the domain wall networks of some moiré materials. We show that in heterobilayers, neighboring twirl chiralities tend to anti-align, forming staggered patterns that are well described by antiferromagnetic lattice $ϕ^4$ theories. In moiré systems with triangular domains, this leads to frustration in the chirality configuration of the structural twirls and to hysteresis with respect to variation of the average twist angle and possibly other control parameters.

cond-mat.mtrl-sci

Quantum siphoning of finely spaced interlayer excitons in reconstructed MoSe2/WSe2 heterostructures

Atomic reconstruction in twisted transition metal dichalcogenide heterostructures leads to mesoscopic domains with uniform atomic registry, profoundly altering the local potential landscape. While interlayer excitons in these domains exhibit strong many-body interactions, extent and impact of quantum confinement on their dynamics remains unclear. Here, we reveal that quantum confinement persists in these flat, reconstructed regions. Time-resolved photoluminescence spectroscopy uncovers multiple, finely-spaced interlayer exciton states (~ 1 meV separation), and correlated emission lifetimes spanning sub-nanosecond to over 100 nanoseconds across a 10 meV energy window. Cascade-like transitions confirm that these states originate from a single potential well, further supported by calculations. Remarkably, at high excitation rates, we observe transient suppression of emission followed by gradual recovery, a process we term "quantum siphoning". Our results demonstrate that quantum confinement and competing nonlinear dynamics persist beyond the ideal moire paradigm, potentially enabling applications in quantum sensing and modifying exciton dynamics via strain engineering.

cond-mat.mes-hall

Graphene Nanoribbons as a Majorana Platform

Graphene nanoribbons support a range of electronic phases that can be controlled via external stimuli. Zigzag-edged graphene nanoribbons (ZGNRs), in particular, exhibit an antiferromagnetic insulating ground state that transitions to a half-metallic phase under a transverse electric field or when embedded inside hexagonal Boron Nitride. Here, we consider a simple model of a heterostructure of a ZGNR with an Ising superconductor and show that, the Ising superconductor with a parent s-wave spin-singlet pairing can induce spin-triplet odd-parity pairing in the half-metallic phase of the ZGNR. The resulting superconducting phase is topologically nontrivial, with gate-tunable transitions that enable the emergence of Majorana zero modes.

cond-mat.mes-hall

Exact projected entangled pair ground states with topological Euler invariant

We report on a class of gapped projected entangled pair states (PEPS) with non-trivial Euler topology motivated by recent progress in band geometry. In the non-interacting limit, these systems have optimal conditions relating to saturation of quantum geometrical bounds, allowing for parent Hamiltonians whose lowest bands are completely flat and which have the PEPS as unique ground states. Protected by crystalline symmetries, these states evade restrictions on capturing tenfold-way topological features with gapped PEPS. These PEPS thus form the first tensor network representative of a non-interacting, gapped two-dimensional topological phase, similar to the Kitaev chain in one dimension. Using unitary circuits, we then formulate interacting variants of these PEPS and corresponding gapped parent Hamiltonians. We reveal characteristic entanglement features shared between the free-fermionic and interacting states with Euler topology. Our results hence provide a rich platform of PEPS models that have, unexpectedly, a finite topological invariant, forming the basis for new spin liquids, quantum Hall physics, and quantum information pursuits.

quant-ph

Polarization textures in crystal supercells with topological bands

Two-dimensional materials are a highly tunable platform for studying the momentum space topology of the electronic wavefunctions and real space topology in terms of skyrmions, merons, and vortices of an order parameter. Such textures for electronic polarization can exist in moiré heterostructures. A quantum-mechanical definition of local polarization textures in insulating supercells was recently proposed. Here, we propose a definition for local polarization that is also valid for systems with topologically non-trivial bands. We introduce semilocal hybrid polarizations, which are valid even when the Wannier functions in a system cannot be made exponentially localized in all dimensions. We use this definition to explicitly show that nontrivial real-space polarization textures can exist in topologically non-trivial systems with non-zero Chern number under (1) an external superlattice potential, and (2) under a stacking-induced moiré potential. In the latter, we find that while the magnitude of the local polarization decreases discontinuously across a topological phase transition from trivial to topologically nontrivial, the polarization does not completely vanish. Our findings suggest that band topology and real-space polar topology may coexist in real materials.

cond-mat.mes-hall

Superconductivity from domain wall fluctuations in sliding ferroelectrics

Bilayers of two-dimensional van der Waals materials that lack an inversion centre can show a novel form of ferroelectricity, where certain stacking arrangements of the two layers lead to an interlayer polarization. Under an external out-of-plane electric field, a relative sliding between the two layers can occur accompanied by an inter-layer charge transfer and a ferroelectric switching. We show that the domain walls that mediate ferroelectric switching are a locus of strong attractive interactions between electrons. The attraction is mediated by the ferroelectric domain wall fluctuations, effectively driven by the soft interlayer shear phonon. We comment on the possible relevance of this attraction mechanism to the recent observation of an interplay between sliding ferroelectricity and superconductivity in bilayer $\text{T$_d$-MoTe}_2$. We also discuss the possible role of this mechanism in the superconductivity of moiré bilayers.

cond-mat.supr-con

On dualities of paired quantum Hall bilayer states at $ν_T = \frac{1}{2} + \frac{1}{2}$

Density-balanced, widely separated quantum Hall bilayers at $ν_T = 1$ can be described as two copies of composite Fermi liquids (CFLs). The two CFLs have interlayer weak-coupling BCS instabilities mediated by gauge fluctuations, the resulting pairing symmetry of which depends on the CFL hypothesis used. If both layers are described by the conventional Halperin-Lee-Read (HLR) theory-based composite electron liquid (CEL), the dominant pairing instability is in the $p+ip$ channel; whereas if one layer is described by CEL and the other by a composite hole liquid (CHL, in the sense of anti-HLR), the dominant pairing instability occurs in the $s$-wave channel. Using the Dirac composite fermion (CF) picture, we show that these two pairing channels can be mapped onto each other by particle-hole (PH) transformation. Furthermore, we derive the CHL theory as the non-relativistic limit of the PH-transformed massive Dirac CF theory. Finally, we prove that an effective topological field theory for the paired CEL-CHL in the weak-coupling limit is equivalent to the exciton condensate phase in the strong-coupling limit.

cond-mat.str-el

Bin-picking of novel objects through category-agnostic-segmentation: RGB matters

This paper addresses category-agnostic instance segmentation for robotic manipulation, focusing on segmenting objects independent of their class to enable versatile applications like bin-picking in dynamic environments. Existing methods often lack generalizability and object-specific information, leading to grasp failures. We present a novel approach leveraging object-centric instance segmentation and simulation-based training for effective transfer to real-world scenarios. Notably, our strategy overcomes challenges posed by noisy depth sensors, enhancing the reliability of learning. Our solution accommodates transparent and semi-transparent objects which are historically difficult for depth-based grasping methods. Contributions include domain randomization for successful transfer, our collected dataset for warehouse applications, and an integrated framework for efficient bin-picking. Our trained instance segmentation model achieves state-of-the-art performance over WISDOM public benchmark [1] and also over the custom-created dataset. In a real-world challenging bin-picking setup our bin-picking framework method achieves 98% accuracy for opaque objects and 97% accuracy for non-opaque objects, outperforming the state-of-the-art baselines with a greater margin.

cs.RO

Superconductivity from polar fluctuations in multi-orbital systems

Motivated by the superconductivity near paraelectric (PE) to ferroelectric (FE) quantum critical point (QCP) in polar metals, we study polar fluctuation mediated superconductivity in multi-orbital systems. The PE to FE QCP is approached by softening of a transverse optical (TO) phonon that is odd under inversion ($\mathcal{I}$). We show that the necessary and sufficient condition for the linear coupling between electron-TO phonon is the presence of multiple orbitals on the Fermi-surface, irrespective of the spin-orbital (SO) coupling, multiple electronic bands, or the vicinity to band crossings. We show that the linear coupling to the polar fluctuations (such as TO modes) can generally lead to superconductivity. We also show that irrespective of the strong $\vec{k}$ dependence of the effective electron-electron interaction in the BCS channel, quite generally the even-parity channel leads to the highest critical temperature. In the presence of additional repulsive electron-electron interactions, an odd-parity spin-triplet channel can become the leading BCS instability. Finally, we discuss our results in the context of the superconductivity in $\text{SrTiO}_3$ and $\text{KTaO}_3$ that highlights the importance of the underlying multi-orbital physics if the superconductivity is mediated by the polar fluctuations.

cond-mat.supr-con

Theory of polarization textures in crystal supercells

Recently, topologically nontrivial polarization textures have been predicted and observed in nanoscale systems. While these polarization textures are interesting and promising in terms of applications, their topology in general is yet to be fully understood. For example, the relation between topological polarization structures and band topology has not been explored, and polar domain structures are typically considered in topologically trivial systems. In particular, the local polarization in a crystal supercell is not well-defined, and typically calculated using approximations which do not satisfy gauge invariance. Furthermore, local polarization in supercells is typically approximated using calculations involving smaller unit cells, meaning the connection to the electronic structure of the supercell is lost. In this work, we propose a definition of local polarization which is gauge invariant and can be calculated directly from a supercell without approximations. We show using first-principles calculations for commensurate bilayer hexagonal boron nitride that our expressions for local polarization give the correct result at the unit cell level, which is a first approximation to the local polarization in a moiré superlattice. We also illustrate using an effective model that the local polarization can be directly calculated in real space. Finally, we discuss the relation between polarization and band topology, for which it is essential to have a correct definition of polarization textures.

cond-mat.mtrl-sci

Predicting the performance of hybrid ventilation in buildings using a multivariate attention-based biLSTM Encoder-Decoder neural network

Hybrid ventilation is an energy-efficient solution to provide fresh air for most climates, given that it has a reliable control system. To operate such systems optimally, a high-fidelity control-oriented modesl is required. It should enable near-real time forecast of the indoor air temperature based on operational conditions such as window opening and HVAC operating schedules. However, physics-based control-oriented models (i.e., white-box models) are labour-intensive and computationally expensive. Alternatively, black-box models based on artificial neural networks can be trained to be good estimators for building dynamics. This paper investigates the capabilities of a deep neural network (DNN), which is a multivariate multi-head attention-based long short-term memory (LSTM) encoder-decoder neural network, to predict indoor air temperature when windows are opened or closed. Training and test data are generated from a detailed multi-zone office building model (EnergyPlus). Pseudo-random signals are used for the indoor air temperature setpoints and window opening instances. The results indicate that the DNN is able to accurately predict the indoor air temperature of five zones whenever windows are opened or closed. The prediction error plateaus after the 24th step ahead prediction (6 hr ahead prediction).

cs.LG

Pairing of Composite-Electrons and Composite-Holes in $ν_T=1$ Quantum Hall Bilayers

Motivated by recent experimental indications of preformed electron-hole pairs in $ν_T=1$ quantum Hall bilayers at relatively large separation, we formulate a Chern-Simons (CS) theory of the coupled composite electron liquid (CEL) and composite hole liquid (CHL). We show that the effective action of the CS gauge field fluctuations around the saddle-point leads to stable pairing between CEL and CHL. We find that the CEL-CHL pairing theory leads to a dominant $s$-wave channel in contrast to the dominant $p$-wave channel found in the CEL-CEL pairing theory. Moreover, the CEL-CHL pairing is generally stronger than the CEL-CEL pairing across the whole frequency spectrum. Finally, we discuss possible differences between the two pairing mechanisms that may be probed in experiments.

cond-mat.str-el

Polar meron-antimeron networks in strained and twisted bilayers

Out-of-plane polar domain structures have recently been discovered in strained and twisted bilayers of inversion symmetry broken systems such as hexagonal boron nitride. Here we show that this symmetry breaking also gives rise to an in-plane component of polarization, and the form of the total polarization is determined purely from symmetry considerations. The in-plane component of the polarization makes the polar domains in strained and twisted bilayers topologically non-trivial, forming a network of merons and antimerons (half-skyrmions and half-antiskyrmions). For twisted systems, the merons are of Bloch type whereas for strained systems they are of Néel type. We propose that the polar domains in strained or twisted bilayers may serve as a platform for exploring topological physics in layered materials and discuss how control over topological phases and phase transitions may be achieved in such systems.

cond-mat.mtrl-sci

Learning to write with the fluid rope trick

The range and speed of direct ink writing, the workhorse of 3d and 4d printing, is limited by the practice of liquid extrusion from a nozzle just above the surface to prevent instabilities to cause deviations from the required print path. But what if could harness the ``fluid rope trick", whence a thin stream of viscous fluid falling from a height spontaneously folds or coils, to write specified patterns on a substrate? Using Deep Reinforcement Learning we control the motion of the extruding nozzle and thence the fluid patterns that are deposited on the surface. The learner (nozzle) repeatedly interacts with the environment (a viscous filament simulator), and improves its strategy using the results of this experience. We demonstrate the results in an experimental setting where the learned motion control instructions are used to drive a viscous jet to accomplish complex tasks such as cursive writing and Pollockian paintings.

physics.flu-dyn