Searcharxiv⌕ Search

arXiv subjects

Vijay Kumar

Publications and source records attributed to Vijay Kumar.

At least 91 records · Page 5Linked to original sources

ADMM-MCBF-LCA: A Layered Control Architecture for Safe Real-Time Navigation

We consider the problem of safe real-time navigation of a robot in a dynamic environment with moving obstacles of arbitrary smooth geometries and input saturation constraints. We assume that the robot detects and models nearby obstacle boundaries with a short-range sensor and that this detection is error-free. This problem presents three main challenges: i) input constraints, ii) safety, and iii) real-time computation. To tackle all three challenges, we present a layered control architecture (LCA) consisting of an offline path library generation layer, and an online path selection and safety layer. To overcome the limitations of reactive methods, our offline path library consists of feasible controllers, feedback gains, and reference trajectories. To handle computational burden and safety, we solve online path selection and generate safe inputs that run at 100 Hz. Through simulations on Gazebo and Fetch hardware in an indoor environment, we evaluate our approach against baselines that are layered, end-to-end, or reactive. Our experiments demonstrate that among all algorithms, only our proposed LCA is able to complete tasks such as reaching a goal, safely. When comparing metrics such as safety, input error, and success rate, we show that our approach generates safe and feasible inputs throughout the robot execution.

cs.RO↗

Leading and beyond leading-order spectral form factor in chaotic quantum many-body systems across all Dyson symmetry classes

We show the emergence of random matrix theory (RMT) spectral correlations in the chaotic phase of generic periodically kicked interacting quantum many-body systems by analytically calculating spectral form factor (SFF), $K(t)$, up to two leading orders in time, $t$. We explicitly consider the presence or absence of time reversal ($\mathcal{T}$) symmetry to investigate all three Dyson's symmetry classes. Our derivation only assumes random phase approximation to enable ensemble average. For $\mathcal{T}$-invariant systems with $\mathcal{T}^2=1$, we show that beyond the Thouless time $t^*$, the SFF takes the form $K(t)\simeq 2t-2t^2/\mathcal{N}$ up to second order in time, where $\mathcal{N}$ is the Hilbert space dimension. This is identical to the result from circular orthogonal ensemble of RMT. In the absence of $\mathcal{T}$-symmetry, we show that $K(t)\simeq t$ beyond $t^*$, and there is no universal term in the second order, unlike the $\mathcal{T}^2=1$ case, in agreement with the result of circular unitary ensemble. For $\mathcal{T}$-invariant systems with $\mathcal{T}^2=-1$, we show that $K(t)\simeq 2t+2t^2/\mathcal{N}$ up to two orders in time beyond $t^*$, in agreement with the result of circular symplectic ensemble. In all three cases, the system-size, $L$, scaling of $t^*$ is determined by eigenvalues of a doubly stochastic matrix $\mathcal{M}$. For strongly interacting fermionic chains, $\mathcal{M}$ is $SU(2)$ invariant in all three cases, leading to $t^*\propto L^2$ in the presence of $U(1)$ symmetry. In the absence of $U(1)$ symmetry, we find $t^*\propto L^0$, due to gapped non-degenerate second-largest eigenvalue of $\mathcal{M}$ or $t^*\propto \ln(L)$ due to gapped second-largest eigenvalue with degeneracy $\propto L^ζ$. Our calculation of SFF is plausible in higher space dimensions as well, where similar system-size scalings of $t^*$ can be obtained.

cond-mat.stat-mech↗

Spatial-to-Temporal Orbital Angular Momentum Mapping in Twisted Light Fields

Optical angular momentum (OAM) in light beams is manifested as the two-dimensional spatial distribution of its complex amplitude, necessitating a 2D detector for its measurement. Here we present a novel speckle-based machine learning approach for OAM recognition, which enables recognition using a 1-D array detector or even a 0-D single-pixel detector.

physics.optics↗

Don't Yell at Your Robot: Physical Correction as the Collaborative Interface for Language Model Powered Robots

We present a novel approach for enhancing human-robot collaboration using physical interactions for real-time error correction of large language model (LLM) powered robots. Unlike other methods that rely on verbal or text commands, the robot leverages an LLM to proactively executes 6 DoF linear Dynamical System (DS) commands using a description of the scene in natural language. During motion, a human can provide physical corrections, used to re-estimate the desired intention, also parameterized by linear DS. This corrected DS can be converted to natural language and used as part of the prompt to improve future LLM interactions. We provide proof-of-concept result in a hybrid real+sim experiment, showcasing physical interaction as a new possibility for LLM powered human-robot interface.

cs.RO↗

Multi-Robot Target Tracking with Sensing and Communication Danger Zones

Multi-robot target tracking finds extensive applications in different scenarios, such as environmental surveillance and wildfire management, which require the robustness of the practical deployment of multi-robot systems in uncertain and dangerous environments. Traditional approaches often focus on the performance of tracking accuracy with no modeling and assumption of the environments, neglecting potential environmental hazards which result in system failures in real-world deployments. To address this challenge, we investigate multi-robot target tracking in the adversarial environment considering sensing and communication attacks with uncertainty. We design specific strategies to avoid different danger zones and proposed a multi-agent tracking framework under the perilous environment. We approximate the probabilistic constraints and formulate practical optimization strategies to address computational challenges efficiently. We evaluate the performance of our proposed methods in simulations to demonstrate the ability of robots to adjust their risk-aware behaviors under different levels of environmental uncertainty and risk confidence. The proposed method is further validated via real-world robot experiments where a team of drones successfully track dynamic ground robots while being risk-aware of the sensing and/or communication danger zones.

cs.RO↗

Closed-loop Analysis of ADMM-based Suboptimal Linear Model Predictive Control

Many practical applications of optimal control are subject to real-time computational constraints. When applying model predictive control (MPC) in these settings, respecting timing constraints is achieved by limiting the number of iterations of the optimization algorithm used to compute control actions at each time step, resulting in so-called suboptimal MPC. This paper proposes a suboptimal MPC scheme based on the alternating direction method of multipliers (ADMM). With a focus on the linear quadratic regulator problem with state and input constraints, we show how ADMM can be used to split the MPC problem into iterative updates of an unconstrained optimal control problem (with an analytical solution), and a dynamics-free feasibility step. We show that using a warm-start approach combined with enough iterations per time-step, yields an ADMM-based suboptimal MPC scheme which asymptotically stabilizes the system and maintains recursive feasibility.

math.OC↗

Jailbreaking LLM-Controlled Robots

The recent introduction of large language models (LLMs) has revolutionized the field of robotics by enabling contextual reasoning and intuitive human-robot interaction in domains as varied as manipulation, locomotion, and self-driving vehicles. When viewed as a stand-alone technology, LLMs are known to be vulnerable to jailbreaking attacks, wherein malicious prompters elicit harmful text by bypassing LLM safety guardrails. To assess the risks of deploying LLMs in robotics, in this paper, we introduce RoboPAIR, the first algorithm designed to jailbreak LLM-controlled robots. Unlike existing, textual attacks on LLM chatbots, RoboPAIR elicits harmful physical actions from LLM-controlled robots, a phenomenon we experimentally demonstrate in three scenarios: (i) a white-box setting, wherein the attacker has full access to the NVIDIA Dolphins self-driving LLM, (ii) a gray-box setting, wherein the attacker has partial access to a Clearpath Robotics Jackal UGV robot equipped with a GPT-4o planner, and (iii) a black-box setting, wherein the attacker has only query access to the GPT-3.5-integrated Unitree Robotics Go2 robot dog. In each scenario and across three new datasets of harmful robotic actions, we demonstrate that RoboPAIR, as well as several static baselines, finds jailbreaks quickly and effectively, often achieving 100% attack success rates. Our results reveal, for the first time, that the risks of jailbroken LLMs extend far beyond text generation, given the distinct possibility that jailbroken robots could cause physical damage in the real world. Indeed, our results on the Unitree Go2 represent the first successful jailbreak of a deployed commercial robotic system. Addressing this emerging vulnerability is critical for ensuring the safe deployment of LLMs in robotics. Additional media is available at: https://robopair.org

cs.RO↗

Resolving thermal gradients and solidification velocities during laser melting of a refractory alloy

Metal additive manufacturing (AM) processes, such as laser powder bed fusion (L-PBF), can yield high-value parts with unique geometries and features, substantially reducing costs and enhancing performance. However, the material properties from L-PBF processes are highly sensitive to the laser processing conditions and the resulting dynamic temperature fields around the melt pool. In this study, we develop a methodology to measure thermal gradients, cooling rates, and solidification velocities during solidification of refractory alloy C103 using in situ high-speed infrared (IR) imaging with a high frame rate of approximately 15,000 frames per second (fps). Radiation intensity maps are converted to temperature maps by integrating thermal radiation over the wavelength range of the camera detector while also considering signal attenuation caused by optical parts. Using a simple method that assigns the liquidus temperature to the melt pool boundary identified ex situ, a scaling relationship between temperature and the IR signal was obtained. The spatial temperature gradients (dT/dx), heating/cooling rates (dT/dt), and solidification velocities (R) are resolved with sufficient temporal resolution under various laser processing conditions, and the resulting microstructures are analyzed, revealing epitaxial growth and nucleated grain growth. Thermal data shows that a decreasing temperature gradient and increasing solidification velocity from the edge to the center of the melt pool can induce a transition from epitaxial to equiaxed grain morphology, consistent with the previously reported columnar to equiaxed transition (CET) trend. The methodology presented can reduce the uncertainty and variability in AM and guide microstructure control during AM of metallic alloys.

physics.app-ph↗

Monocular Event-Based Vision for Obstacle Avoidance with a Quadrotor

We present the first static-obstacle avoidance method for quadrotors using just an onboard, monocular event camera. Quadrotors are capable of fast and agile flight in cluttered environments when piloted manually, but vision-based autonomous flight in unknown environments is difficult in part due to the sensor limitations of traditional onboard cameras. Event cameras, however, promise nearly zero motion blur and high dynamic range, but produce a very large volume of events under significant ego-motion and further lack a continuous-time sensor model in simulation, making direct sim-to-real transfer not possible. By leveraging depth prediction as a pretext task in our learning framework, we can pre-train a reactive obstacle avoidance events-to-control policy with approximated, simulated events and then fine-tune the perception component with limited events-and-depth real-world data to achieve obstacle avoidance in indoor and outdoor settings. We demonstrate this across two quadrotor-event camera platforms in multiple settings and find, contrary to traditional vision-based works, that low speeds (1m/s) make the task harder and more prone to collisions, while high speeds (5m/s) result in better event-based depth estimation and avoidance. We also find that success rates in outdoor scenes can be significantly higher than in certain indoor scenes.

cs.RO↗

Trajectory Optimization with Global Yaw Parameterization for Field-of-View Constrained Autonomous Flight

Trajectory generation for quadrotors with limited field-of-view sensors has numerous applications such as aerial exploration, coverage, inspection, videography, and target tracking. Most previous works simplify the task of optimizing yaw trajectories by either aligning the heading of the robot with its velocity, or potentially restricting the feasible space of candidate trajectories by using a limited yaw domain to circumvent angular singularities. In this paper, we propose a novel \textit{global} yaw parameterization method for trajectory optimization that allows a 360-degree yaw variation as demanded by the underlying algorithm. This approach effectively bypasses inherent singularities by including supplementary quadratic constraints and transforming the final decision variables into the desired state representation. This method significantly reduces the needed control effort, and improves optimization feasibility. Furthermore, we apply the method to several examples of different applications that require jointly optimizing over both the yaw and position trajectories. Ultimately, we present a comprehensive numerical analysis and evaluation of our proposed method in both simulation and real-world experiments.

cs.RO↗

EvMAPPER: High Altitude Orthomapping with Event Cameras

Traditionally, unmanned aerial vehicles (UAVs) rely on CMOS-based cameras to collect images about the world below. One of the most successful applications of UAVs is to generate orthomosaics or orthomaps, in which a series of images are integrated together to develop a larger map. However, the use of CMOS-based cameras with global or rolling shutters mean that orthomaps are vulnerable to challenging light conditions, motion blur, and high-speed motion of independently moving objects under the camera. Event cameras are less sensitive to these issues, as their pixels are able to trigger asynchronously on brightness changes. This work introduces the first orthomosaic approach using event cameras. In contrast to existing methods relying only on CMOS cameras, our approach enables map generation even in challenging light conditions, including direct sunlight and after sunset.

cs.RO↗

Collision-free time-optimal path parameterization for multi-robot teams

Coordinating the motion of multiple robots in cluttered environments remains a computationally challenging task. We study the problem of minimizing the execution time of a set of geometric paths by a team of robots with state-dependent actuation constraints. We propose a Time-Optimal Path Parameterization (TOPP) algorithm for multiple car-like agents, where the modulation of the timing of every robot along its assigned path is employed to ensure collision avoidance and dynamic feasibility. This is achieved through the use of a priority queue to determine the order of trajectory execution for each robot while taking into account all possible collisions with higher priority robots in a spatiotemporal graph. We show a 10-20% reduction in makespan against existing state-of-the-art methods and validate our approach through simulations and hardware experiments.

cs.RO↗

AgriNeRF: Neural Radiance Fields for Agriculture in Challenging Lighting Conditions

Neural Radiance Fields (NeRFs) have shown significant promise in 3D scene reconstruction and novel view synthesis. In agricultural settings, NeRFs can serve as digital twins, providing critical information about fruit detection for yield estimation and other important metrics for farmers. However, traditional NeRFs are not robust to challenging lighting conditions, such as low-light, extreme bright light and varying lighting. To address these issues, this work leverages three different sensors: an RGB camera, an event camera and a thermal camera. Our RGB scene reconstruction shows an improvement in PSNR and SSIM by +2.06 dB and +8.3% respectively. Our cross-spectral scene reconstruction enhances downstream fruit detection by +43.0% in mAP50 and +61.1% increase in mAP50-95. The integration of additional sensors leads to a more robust and informative NeRF. We demonstrate that our multi-modal system yields high quality photo-realistic reconstructions under various tree canopy covers and at different times of the day. This work results in the development of a resilient NeRF, capable of performing well in visibly degraded scenarios, as well as a learnt cross-spectral representation, that is used for automated fruit detection.

cs.RO↗

A Networked Multi-Agent System for Mobile Wireless Infrastructure on Demand

Despite the prevalence of wireless connectivity in urban areas around the globe, there remain numerous and diverse situations where connectivity is insufficient or unavailable. To address this, we introduce mobile wireless infrastructure on demand, a system of UAVs that can be rapidly deployed to establish an ad-hoc wireless network. This network has the capability of reconfiguring itself dynamically to satisfy and maintain the required quality of communication. The system optimizes the positions of the UAVs and the routing of data flows throughout the network to achieve this quality of service (QoS). By these means, task agents using the network simply request a desired QoS, and the system adapts accordingly while allowing them to move freely. We have validated this system both in simulation and in real-world experiments. The results demonstrate that our system effectively offers mobile wireless infrastructure on demand, extending the operational range of task agents and supporting complex mobility patterns, all while ensuring connectivity and being resilient to agent failures.

cs.RO↗

Constraint-Aware Intent Estimation for Dynamic Human-Robot Object Co-Manipulation

Constraint-aware estimation of human intent is essential for robots to physically collaborate and interact with humans. Further, to achieve fluid collaboration in dynamic tasks intent estimation should be achieved in real-time. In this paper, we present a framework that combines online estimation and control to facilitate robots in interpreting human intentions, and dynamically adjust their actions to assist in dynamic object co-manipulation tasks while considering both robot and human constraints. Central to our approach is the adoption of a Dynamic Systems (DS) model to represent human intent. Such a low-dimensional parameterized model, along with human manipulability and robot kinematic constraints, enables us to predict intent using a particle filter solely based on past motion data and tracking errors. For safe assistive control, we propose a variable impedance controller that adapts the robot's impedance to offer assistance based on the intent estimation confidence from the DS particle filter. We validate our framework on a challenging real-world human-robot co-manipulation task and present promising results over baselines. Our framework represents a significant step forward in physical human-robot collaboration (pHRC), ensuring that robot cooperative interactions with humans are both feasible and effective.

cs.RO↗

Air-Ground Collaboration with SPOMP: Semantic Panoramic Online Mapping and Planning

Mapping and navigation have gone hand-in-hand since long before robots existed. Maps are a key form of communication, allowing someone who has never been somewhere to nonetheless navigate that area successfully. In the context of multi-robot systems, the maps and information that flow between robots are necessary for effective collaboration, whether those robots are operating concurrently, sequentially, or completely asynchronously. In this paper, we argue that maps must go beyond encoding purely geometric or visual information to enable increasingly complex autonomy, particularly between robots. We propose a framework for multi-robot autonomy, focusing in particular on air and ground robots operating in outdoor 2.5D environments. We show that semantic maps can enable the specification, planning, and execution of complex collaborative missions, including localization in GPS-denied settings. A distinguishing characteristic of this work is that we strongly emphasize field experiments and testing, and by doing so demonstrate that these ideas can work at scale in the real world. We also perform extensive simulation experiments to validate our ideas at even larger scales. We believe these experiments and the experimental results constitute a significant step forward toward advancing the state-of-the-art of large-scale, collaborative multi-robot systems operating with real communication, navigation, and perception constraints.

cs.RO↗

In situ observation of thermally activated and localized Li leaching from lithiated graphite

Temperature is known to impact Li-ion battery performance and safety, however, understanding its effect on Li-ion batteries has largely been limited to uniform high or low temperatures. While the insights gathered from such research are important, much less information is available on the effects of non-uniform temperatures which more accurately reflect the environments that Li-ion batteries are exposed to in real world applications. In this paper, we characterize the impact of a microscale, temperature hotspot on a Li-ion battery using a combination of in situ micro-Raman spectroscopy, in situ optical microscopy and COMSOL Multiphysics thermal simulations. Our results show that mild temperature heterogeneity induced by the micro-Raman laser can cause lithium to locally leach out from different lithiated graphite phases (LiC6 and LiC12) in the absence of an applied current. The Li metal is found to be largely localized to the region heated by the micro-Raman laser and is not observed upon uniform heating to comparable temperatures suggesting that temperature heterogeneity is uniquely responsible for causing Li to leach out from lithiated graphite phases. A mechanism whereby localized temperature heterogeneity induced by the laser induces heterogeneity in the degree of lithiation across the graphite anode is proposed to explain the localized Li leaching. This study highlights the sensitivity of lithiated graphite phases to minor temperature heterogeneity in the absence of an applied current.

cond-mat.mtrl-sci↗

Challenges and Opportunities for Large-Scale Exploration with Air-Ground Teams using Semantics

One common and desirable application of robots is exploring potentially hazardous and unstructured environments. Air-ground collaboration offers a synergistic approach to addressing such exploration challenges. In this paper, we demonstrate a system for large-scale exploration using a team of aerial and ground robots. Our system uses semantics as lingua franca, and relies on fully opportunistic communications. We highlight the unique challenges from this approach, explain our system architecture and showcase lessons learned during our experiments. All our code is open-source, encouraging researchers to use it and build upon.

cs.RO↗