Searcharxiv⌕ Search

arXiv subjects

Vaibhav Srivastava

Publications and source records attributed to Vaibhav Srivastava.

At least 19 recordsLinked to original sources

Reward-Rate Congestion Games and Replicator--Dinkelbach Dynamics

Reward rate is a key performance criterion in cyber-physical and robotic systems where time, workload, and coordination costs are limiting resources. We introduce reward-rate congestion games, where agents seek to maximize reward per unit execution time. The direct reward-rate game is generally not an exact potential game. We develop a Dinkelbach-based framework in which, for every fixed Dinkelbach parameter, the transformed game is an exact potential game. This yields a potential-level Dinkelbach iteration that terminates finitely at the optimal potential reward rate when the inner potential maximization problem is solved globally. We also provide a sufficient condition under which an equilibrium of the transformed game is an equilibrium of the original reward-rate game. To optimize aggregate performance, we introduce marginal externality corrections that make the corrected potential coincide with the Dinkelbach-transformed social reward-rate objective, thereby enabling optimization of the social reward rate. Finally, we develop a continuous-time replicator--Dinkelbach dynamics for reward-rate population games coupling fast replicator dynamics with a slow reward-rate update. We establish convergence of the fixed-parameter replicator dynamics, global asymptotic and local exponential stability of the reduced Dinkelbach dynamics, and local exponential stability of the coupled system for sufficiently slow Dinkelbach updates. The framework is illustrated on a continuous task-allocation problem.

eess.SY↗

Structured LLM Reasoning for Zero-Shot Human--Robot Coordination Under Hidden Goals

We present a structured large-language-model (LLM) architecture for zero-shot human--robot coordination in a cooperative construction task with private goal views. Guided by a Dec-POMDP formulation, the architecture decomposes decision-making into (i) action-conditioned Theory-of-Mind (ToM) inference, (ii) hierarchical planning, (iii) conversation interpretation, (iv) action verification, and (v) feedback-based replanning. We compare the proposed method with an ablation without ToM inference and a multi-agent reinforcement-learning policy trained offline over many goal pairs. In human-participant experiments, the proposed method required fewer interaction steps and yielded higher post-interaction trust ratings than both baselines. These results suggest that systematically decomposing the team decision problem, using LLMs as tractable surrogates for otherwise intractable inference and planning computations, and retaining conventional verification for physical feasibility can improve both task coordination and the human experience.

cs.RO↗

Robustness and Centrality in Markov-switching Networks

We investigate how time-varying interactions, modeled via a Markov switching graph (MSG), impact the robustness of noisy multi-agent dynamics in both continuous- and discrete-time settings. Our focus is on the steady-state performance of consensus and leader-follower tracking dynamics subject to stochastic noise. Using the framework of Markov jump linear systems (MJLS), we derive expressions for the steady-state covariance of each agent's deviation from consensus and tracking error, respectively, and use them to quantify individual and group performance as a function of the interaction graphs and the switching dynamics. We extend established notions of robustness, certainty indices, and joint centrality from static graphs to the MSG setting. To gain analytical insight, we specialize our results to systems switching between two topologies and characterize how switching influences performance. Numerical simulations further illustrate how switching topologies affects system robustness in both coordination tasks.

eess.SY↗

Automated Curriculum Design for High-dimensional Human Motor Learning

Designing effective practice schedules for high-dimensional motor learning tasks remains a challenge, especially when skill states are unobservable and task performance may not reflect the true learning. We propose an automated curriculum design framework that combines a human motor learning model and personalized real-time skill estimation with Stochastic Nonlinear Model Predictive Control in \emph{de-novo} (novel) motor learning paradigms. We validated our framework both through simulations and human-subject studies (N = 36) using a hand exoskeleton. Our proposed approach accelerates skill acquisition by $\sim23\%$, and ${\sim17\%}$ when compared to a random curriculum and a performance heuristics-based curriculum, respectively. These significant gains in learning efficiency highlight the potential of model-based, individualized curricula for motor rehabilitation and complex skill training.

eess.SY↗

Co-Learning Port-Hamiltonian Systems and Optimal Energy-Shaping Control

We develop a physics-informed learning framework for energy-shaping control of port-Hamiltonian (pH) systems from trajectory data. The proposed approach co-learns a pH system model and an optimal energy-balancing passivity-based controller (EB-PBC) through alternating optimization with policy-aware data collection. At each iteration, the system model is refined using trajectory data collected under the current control policy, and the controller is re-optimized on the updated model. Both components are parameterized by neural networks that embed the pH dynamics and EB-PBC structure, ensuring interpretability in terms of energy interactions. The learned controller renders the closed-loop system inherently passive and provably stable, and exploits passive plant dynamics without canceling the natural potential. A dissipation regularization enforces strict energy decay during training, thereby enhancing robustness to sim-to-real gaps. The proposed framework is validated on state-regulation and swing-up tasks for planar and torsional pendulum systems.

eess.SY↗

Neuromorphic Realization of Best Response in Finite-Action Games

We develop a mechanistic dynamical-systems formulation of best response in finite-action games with relational structure on the action set. The proposed neuromorphic decision dynamics realize best response as the stable outcome of an internal state-space process, rather than as an externally imposed choice rule. This provides a deterministic account of commitment formation, symmetry resolution through basins of attraction, and hysteresis and decision persistence under perturbations. For action spaces with circulant coupling, we prove using Lyapunov-Schmidt reduction that the action-coupling operator determines which components of evidence govern decision formation. We further show that the dynamics implicitly compute a geometry-aware utility, converge exponentially to the corresponding best response with rate independent of the number of actions, and switch only when evidence is sufficiently strong. In contrast, supplying the same geometry-aware utility directly to logit dynamics does not recover these properties, showing that relational structure must be embedded in the decision mechanism itself. We illustrate the framework in a repeated coverage game, prove that the induced game is an exact potential game, and show that its Nash equilibria are reached by the neuromorphic dynamics.

math.DS↗

Skill-informed Data-driven Haptic Nudges for High-dimensional Human Motor Learning

In this work, we propose a data-driven framework to design optimal haptic nudge feedback leveraging the learner's estimated skill to address the challenge of learning a novel motor task in a high-dimensional, redundant motor space. A nudge is a series of vibrotactile feedback delivered to the learner to encourage motor movements that aid in task completion. We first model the stochastic dynamics of human motor learning under haptic nudges using an Input-Output Hidden Markov Model (IOHMM), which explicitly decouples latent skill evolution from observable performance measures. Leveraging this predictive model, we formulate the haptic nudge feedback design problem as a Partially Observable Markov Decision Process (POMDP). This allows us to derive an optimal nudging policy that minimizes long-term performance cost and implicitly guides the learner toward superior skill states. We validate our approach through a human participant study (N=30) involving a high-dimensional motor task rendered through a hand exoskeleton. Results demonstrate that participants trained with the POMDP-derived policy exhibit significantly accelerated movement efficiency and endpoint accuracy compared to groups receiving heuristic-based feedback or no feedback. Furthermore, synergy analysis reveals that the POMDP group discovers efficient low-dimensional motor representations more rapidly.

cs.RO↗

A Normative Theory of Decision Making from Multiple Stimuli: The Contextual Diffusion Decision Model

The dynamics of simple two-alternative forced-choice (2AFC) decisions are well-modeled by a class of random walk models (e.g. Laming, 1968; Ratcliff, 1978; Usher & McClelland, 2001; Bogacz et al., 2006). However, in real-life, even simple decisions involve dynamically changing influence of additional information. In this work, we describe a computational theory of decision making from multiple sources of information, grounded in Bayesian inference and consistent with a simple neural network. This Contextual Diffusion Decision Model (CDDM) is a formal generalization of the Diffusion Decision Model (DDM), a popular existing model of fixed-context decision making (Ratcliff, 1978), and shares with it both a mechanistic and a probabilistic motivation. Just as the DDM is a model for a variety of simple two-alternative forced-choice (2AFC) decision making tasks, we demonstrate that the CDDM supports a variety of simple context-dependent tasks of longstanding interest in psychology, including the Flanker (Eriksen & Eriksen, 1974), AX-CPT (Servan-Schreiber et al., 1996), Stop-Signal (Logan & Cowan, 1984), Cueing (Posner, 1980), and Prospective Memory paradigms (Einstein & McDaniel, 2005). Further, we use the CDDM to perform a number of normative rational analyses exploring optimal response and memory allocation policies. Finally, we show how the use of a consistent model across tasks allows us to recover consistent qualitative data patterns in multiple tasks, using the same model parameters.

q-bio.NC↗

Multi-Robot Multitask Gaussian Process Estimation and Coverage

Coverage control is essential for the optimal deployment of agents to monitor or cover areas with sensory demands. While traditional coverage involves single-task robots, increasing autonomy now enables multitask operations. This paper introduces a novel multitask coverage problem and addresses it for both the cases of known and unknown sensory demands. For known demands, we design a federated multitask coverage algorithm and establish its convergence properties. For unknown demands, we employ a multitask Gaussian Process (GP) framework to learn sensory demand functions and integrate it with the multitask coverage algorithm to develop an adaptive algorithm. We introduce a novel notion of multitask coverage regret that compares the performance of the adaptive algorithm against an oracle with prior knowledge of the demand functions. We establish that our algorithm achieves sublinear cumulative regret, and numerically illustrate its performance.

eess.SY↗

When To Seek Help: Trust-Aware Assistance Seeking in Human-Supervised Autonomy

Our goal is to model and experimentally assess trust evolution to predict future beliefs and behaviors of human-robot teams in dynamic environments. Research suggests that maintaining trust among team members in a human-robot team is vital for successful team performance. Research suggests that trust is a multi-dimensional and latent entity that relates to past experiences and future actions in a complex manner. Employing a human-robot collaborative task, we design an optimal assistance-seeking strategy for the robot using a POMDP framework. In the task, the human supervises an autonomous mobile manipulator collecting objects in an environment. The supervisor's task is to ensure that the robot safely executes its task. The robot can either choose to attempt to collect the object or seek human assistance. The human supervisor actively monitors the robot's activities, offering assistance upon request, and intervening if they perceive the robot may fail. In this setting, human trust is the hidden state, and the primary objective is to optimize team performance. We execute two sets of human-robot interaction experiments. The data from the first experiment are used to estimate POMDP parameters, which are used to compute an optimal assistance-seeking policy evaluated in the second experiment. The estimated POMDP parameters reveal that, for most participants, human intervention is more probable when trust is low, particularly in high-complexity tasks. Our estimates suggest that the robot's action of asking for assistance in high-complexity tasks can positively impact human trust. Our experimental results show that the proposed trust-aware policy is better than an optimal trust-agnostic policy. By comparing model estimates of human trust, obtained using only behavioral data, with the collected self-reported trust values, we show that model estimates are isomorphic to self-reported responses.

cs.RO↗

Heterogeneous Multi-Agent Task-Assignment with Uncertain Execution Times and Preferences

While sequential task assignment for a single agent has been widely studied, such problems in a multi-agent setting, where the agents have heterogeneous task preferences or capabilities, remain less well-characterized. We study a multi-agent task assignment problem where a central planner assigns recurring tasks to multiple members of a team over a finite time horizon. For any given task, the members have heterogeneous capabilities in terms of task completion times, task resource consumption (which can model variables such as energy or attention), and preferences in terms of the rewards they collect upon task completion. We assume that the reward, execution time, and resource consumption for each member to complete any task are stochastic with unknown distributions. The goal of the planner is to maximize the total expected reward that the team receives over the problem horizon while ensuring that the resource consumption required for any assigned task is within the capability of the agent. We propose and analyze a bandit algorithm for this problem. Since the bandit algorithm relies on solving an optimal task assignment problem repeatedly, we analyze the achievable regret in two cases: when we can solve the optimal task assignment exactly and when we can solve it only approximately.

cs.MA↗

Fast Online Adaptive Neural MPC via Meta-Learning

Data-driven model predictive control (MPC) has demonstrated significant potential for improving robot control performance in the presence of model uncertainties. However, existing approaches often require extensive offline data collection and computationally intensive training, limiting their ability to adapt online. To address these challenges, this paper presents a fast online adaptive MPC framework that leverages neural networks integrated with Model-Agnostic Meta-Learning (MAML). Our approach focuses on few-shot adaptation of residual dynamics - capturing the discrepancy between nominal and true system behavior - using minimal online data and gradient steps. By embedding these meta-learned residual models into a computationally efficient L4CasADi-based MPC pipeline, the proposed method enables rapid model correction, enhances predictive accuracy, and improves real-time control performance. We validate the framework through simulation studies on a Van der Pol oscillator, a Cart-Pole system, and a 2D quadrotor. Results show significant gains in adaptation speed and prediction accuracy over both nominal MPC and nominal MPC augmented with a freshly initialized neural network, underscoring the effectiveness of our approach for real-time adaptive robot control.

cs.RO↗

Velocity-Form Data-Enabled Predictive Control of Soft Robots under Unknown External Payloads

Data-driven control methods such as data-enabled predictive control (DeePC) have shown strong potential in efficient control of soft robots without explicit parametric models. However, in object manipulation tasks, unknown external payloads and disturbances can significantly alter the system dynamics and behavior, leading to offset error and degraded control performance. In this paper, we present a novel velocity-form DeePC framework that achieves robust and optimal control of soft robots under unknown payloads. The proposed framework leverages input-output data in an incremental representation to mitigate performance degradation induced by unknown payloads, eliminating the need for weighted datasets or disturbance estimators. We validate the method experimentally on a planar soft robot and demonstrate its superior performance compared to standard DeePC in scenarios involving unknown payloads.

cs.RO↗

Bi-Virus SIS Epidemic Propagation under Mutation and Game-theoretic Protection Adoption

We study a bi-virus susceptible-infected-susceptible (SIS) epidemic model in which individuals are either susceptible or infected with one of two virus strains, and consider mutation-driven transitions between strains. The general case of bi-directional mutation is first analyzed, where we characterize the disease-free equilibrium and establish its global asymptotic stability, as well as the existence, uniqueness, and stability of an endemic equilibrium. We then present a game-theoretic framework where susceptible individuals strategically choose whether to adopt protection or remain unprotected, to maximize their instantaneous payoffs. We derive Nash strategies under bi-directional mutation, and subsequently consider the special case of unidirectional mutation. In the latter case, we show that coexistence of both strains is impossible when mutation occurs from the strain with lower reproduction number and transmission rate to the other strain. Furthermore, we fully characterize the stationary Nash equilibrium (SNE) in the setting permitting coexistence, and examine how mutation rates influence protection adoption and infection prevalence at the SNE. Numerical simulations corroborate the analytical results, demonstrating that infection levels decrease monotonically with higher protection adoption, and highlight the impact of mutation rates and protection cost on infection state trajectories.

q-bio.PE↗

LogicGuard: Improving Embodied LLM agents through Temporal Logic based Critics

Large language models (LLMs) have shown promise in zero-shot and single step reasoning and decision making problems, but in long horizon sequential planning tasks, their errors compound, often leading to unreliable or inefficient behavior. We introduce LogicGuard, a modular actor-critic architecture in which an LLM actor is guided by a trajectory level LLM critic that communicates through Linear Temporal Logic (LTL). Our setup combines the reasoning strengths of language models with the guarantees of formal logic. The actor selects high-level actions from natural language observations, while the critic analyzes full trajectories and proposes new LTL constraints that shield the actor from future unsafe or inefficient behavior. LogicGuard supports both fixed safety rules and adaptive, learned constraints, and is model-agnostic: any LLM-based planner can serve as the actor, with LogicGuard acting as a logic-generating wrapper. We formalize planning as graph traversal under symbolic constraints, allowing LogicGuard to analyze failed or suboptimal trajectories and generate new temporal logic rules that improve future behavior. To demonstrate generality, we evaluate LogicGuard across two distinct settings: short-horizon general tasks and long-horizon specialist tasks. On the Behavior benchmark of 100 household tasks, LogicGuard increases task completion rates by 25% over a baseline InnerMonologue planner. On the Minecraft diamond-mining task, which is long-horizon and requires multiple interdependent subgoals, LogicGuard improves both efficiency and safety compared to SayCan and InnerMonologue. These results show that enabling LLMs to supervise each other through temporal logic yields more reliable, efficient and safe decision-making for both embodied agents.

cs.AI↗

Optimal Fidelity Selection for Human-Supervised Search

We study optimal fidelity selection in human-supervised underwater visual search, where operator performance is affected by cognitive factors like workload and fatigue. In our experiments, participants perform two simultaneous tasks: detecting underwater mines in videos (primary) and responding to a visual cue to estimate workload (secondary). Videos arrive as a Poisson process and queue for review, with the operator choosing between normal fidelity (faster playback) and high fidelity. Rewards are based on detection accuracy, while penalties depend on queue length. Workload is modeled as a hidden state using an Input-Output Hidden Markov Model, and fidelity selection is optimized via a Partially Observable Markov Decision Process. We evaluate two setups: fidelity-only selection and a version allowing task delegation to automation to maintain queue stability. Our approach improves performance by 26.5% without delegation and 50.3% with delegation, compared to a baseline where humans manually choose their fidelity levels.

cs.HC↗

Modeling Trust Dynamics in Robot-Assisted Delivery: Impact of Trust Repair Strategies

With increasing efficiency and reliability, autonomous systems are becoming valuable assistants to humans in various tasks. In the context of robot-assisted delivery, we investigate how robot performance and trust repair strategies impact human trust. In this task, while handling a secondary task, humans can choose to either send the robot to deliver autonomously or manually control it. The trust repair strategies examined include short and long explanations, apology and promise, and denial. Using data from human participants, we model human behavior using an Input-Output Hidden Markov Model (IOHMM) to capture the dynamics of trust and human action probabilities. Our findings indicate that humans are more likely to deploy the robot autonomously when their trust is high. Furthermore, state transition estimates show that long explanations are the most effective at repairing trust following a failure, while denial is most effective at preventing trust loss. We also demonstrate that the trust estimates generated by our model are isomorphic to self-reported trust values, making them interpretable. This model lays the groundwork for developing optimal policies that facilitate real-time adjustment of human trust in autonomous systems.

cs.RO↗

Model-free Vehicle Rollover Prevention: A Data-driven Predictive Control Approach

Vehicle rollovers pose a significant safety risk and account for a disproportionately high number of fatalities in road accidents. This paper addresses the challenge of rollover prevention using Data-EnablEd Predictive Control (DeePC), a data-driven control strategy that directly leverages raw input-output data to maintain vehicle stability without requiring explicit system modeling. To enhance computational efficiency, we employ a reduced-dimension DeePC that utilizes singular value decomposition-based dimension reduction to significantly lower computation complexity without compromising control performance. This optimization enables real-time application in scenarios with high-dimensional data, making the approach more practical for deployment in real-world vehicles. The proposed approach is validated through high-fidelity CarSim simulations in both sedan and utility truck scenarios, demonstrating its versatility and ability to maintain vehicle stability under challenging driving conditions. Comparative results with Linear Model Predictive Control (LMPC) highlight the superior performance of DeePC in preventing rollovers while preserving maneuverability. The findings suggest that DeePC offers a robust and adaptable solution for rollover prevention, capable of handling varying road and vehicle conditions.

eess.SY↗