SearcharxivSearch

arXiv subjects

Hari Srikanth

Publications and source records attributed to Hari Srikanth.

4 recordsLinked to original sources

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings. Following the LLM community, an emerging access paradigm for closed-weight robot foundation models is the managed supervised fine-tuning (SFT) API, where users submit training data and receive a tuned policy without access to model weights, gradients, or training internals. While such APIs let downstream users leverage powerful proprietary foundation models, they restrict policy improvement to pure imitation, ruling out reinforcement learning and other closed-loop methods that rely on internal training signals. This limitation is particularly acute for agile, contact-rich humanoid manipulation, where the gap between policy outputs and deployed behavior is large due to novel states, action tracking dynamics, latency, and controller-specific failure modes. We study how effective this managed-API regime is for humanoid adaptation, and how closed-loop improvement can be realized within it to push policies toward task mastery. We conduct one of the first empirical studies of managed-API adaptation on a real humanoid, instantiated on Gemini Robotics On-Device (GROD). We find that direct SFT through the API substantially outperforms a leading open-weight VLA trained on the same demonstrations, yet still falls short of deployment-level mastery on agile, contact-rich tasks. To close this gap, we introduce CLIFT: Closed-Loop Iterative Fine-Tuning, which turns deployment-time reward feedback into API-compatible supervised data and enables closed-loop policy improvement without accessing weights, gradients, likelihoods, or losses-pushing GROD to near-perfect success after two flywheel cycles, all without "opening the model box."

cs.RO

Linear Function Approximation as a Computationally Efficient Method to Solve Classical Reinforcement Learning Challenges

Neural Network based approximations of the Value function make up the core of leading Policy Based methods such as Trust Regional Policy Optimization (TRPO) and Proximal Policy Optimization (PPO). While this adds significant value when dealing with very complex environments, we note that in sufficiently low State and action space environments, a computationally expensive Neural Network architecture offers marginal improvement over simpler Value approximation methods. We present an implementation of Natural Actor Critic algorithms with actor updates through Natural Policy Gradient methods. This paper proposes that Natural Policy Gradient (NPG) methods with Linear Function Approximation as a paradigm for value approximation may surpass the performance and speed of Neural Network based models such as TRPO and PPO within these environments. Over Reinforcement Learning benchmarks Cart Pole and Acrobot, we observe that our algorithm trains much faster than complex neural network architectures, and obtains an equivalent or greater result. This allows us to recommend the use of NPG methods with Linear Function Approximation over TRPO and PPO for both traditional and sparse reward low dimensional problems.

cs.LG

Reinforcement Learning Based Escape Route Generation in Low Visibility Environments

Structure fires are responsible for the majority of fire-related deaths nationwide. In order to assist with the rapid evacuation of trapped people, this paper proposes the use of a system that determines optimal search paths for firefighters and exit paths for civilians in real time based on environmental measurements. Through the use of a LiDAR mapping system evaluated and verified by a trust range derived from sonar and smoke concentration data, a proposed solution to low visibility mapping is tested. These independent point clouds are then used to create distinct maps, which are merged through the use of a RANSAC based alignment methodology and simplified into a visibility graph. Temperature and humidity data are then used to label each node with a danger score, creating an environment tensor. After demonstrating how a Linear Function Approximation based Natural Policy Gradient RL methodology outperforms more complex competitors with respect to robustness and speed, this paper outlines two systems (savior and refugee) that process the environment tensor to create safe rescue and escape routes, respectively.

cs.AI

Scaling of the thermally induced sign inversion of longitudinal spin Seebeck effect in a compensated ferrimagnet: Role of magnetic anisotropy

We report on a systematic investigation of the longitudinal spin Seebeck effect (LSSE) in a GGG(Gd3Ga5O12)/GdIG(Gd3Fe5O12)/Pt film series exhibiting an in-plane magnetic easy axis with a compensation temperature (T_Comp) that decreases from 270 to 220 K when decreasing GdIG film thickness from 272 to 31 nm, respectively. For all the films, the LSSE signal flips its sign below T_Comp. We demonstrate a universal scaling behavior of the temperature dependence of LSSE signal for our GdIG films around their respective T_Comp. Additionally, we demonstrate LSSE in a 31 nm GdIG film grown on a lattice-mismatched GSGG (Gd3Sc2Ga3O12) substrate that exhibits an out-of-plane magnetic easy axis at room temperature. However, this sample reveals a spin reorientation transition where the magnetic easy axis changes its orientation to in-plane at low temperatures. We observed a clear distinction in the LSSE signal for the GSGG/GdIG(31 nm)/Pt heterostructure, relative to GGG/GdIG(31nm)/Pt showing an in-plane magnetic easy axis. Our findings underscore a strong correlation between the LSSE signal and the orientation of magnetic easy axis in compensated ferrimagnets and opens the possibility to tune LSSE through effective anisotropy.

cond-mat.mtrl-sci