SearcharxivSearch

arXiv subjects

Brian Zhu

Publications and source records attributed to Brian Zhu.

12 recordsLinked to original sources

Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. In particular, their severe inference latency---which can lead to pauses or jerky movements---can alter the effective environment dynamics and, if not correctly accounted for, break the Markov assumption that RL relies on, causing standard RL algorithms to fail completely. In this work, we introduce a latency-aware framework, Asynchronous RL with Intermediate Information (ARLI), that enables RL-based improvement of generalist policies under inference delays. Our framework builds on asynchronous inference approaches, which interleave action generation with execution to hide latency, and addresses its incompatibility with RL by providing a low-latency RL policy design that maximizes reactivity within the inference window through two contributions: state augmentations that restore near-Markovian structure by incorporating committed actions and a mid-inference observation. We evaluate our approach across simulated and real-world manipulation tasks, and find that it enables effective finetuning under inference delays where standard RL fails entirely, even matching or exceeding the performance of standard RL in idealized no-latency settings.

cs.RO

Uniform-Loss Automated Market Making for Prediction Markets

Automated market makers (AMMs) for prediction markets descend from market scoring rules, where a mechanism operator subsidizes a market to aggregate beliefs about uncertain events. The existing literature has focused on bounding the total worst-case loss to the subsidizer, but has not addressed how that loss is distributed across price states or over time. We use the framework of loss-versus-rebalancing (LVR) to study this distribution and introduce \textit{uniform AMMs}, defined by the property that instantaneous LVR is proportional to pool value and independent of the current token price. In a static setting, we show that for a broad class of \textit{win-martingales} -- processes that converge to 0 or 1 at a fixed resolution time -- there exists a pricing function that achieves uniform LVR under that process, and conversely, that any sufficiently regular pricing function induces a win-martingale under which it is uniform. We then extend the framework to dynamic liquidity management, showing that liquidity levels can be adjusted over time to implement a prescribed target expected cumulative loss schedule. This theory is illustrated with canonical examples of win-martingales and pricing functions. Our results can inform AMM designers and liquidity providers on how the inevitable cost of subsidizing price discovery can be shaped and controlled across both price and time.

q-fin.TR

Closing the Loop in Teleoperation: Episode-Level Data Quality Assessment and Feedback for High-Quality Demonstration Collection

Industrial automation is at a pivotal moment, as Physical AI is driving a transition from rigid, hand-engineered automation systems toward more flexible and adaptive systems. This shift has created a growing demand for large-scale, real-world robot demonstration data, making teleoperation an increasingly important mechanism for data collection. However, high-quality teleoperated demonstrations remain difficult to obtain in practice, as novice operators often produce episodes that are task-successful but suboptimal for downstream use due to inefficient motion, repeated corrections, or operation near robot joint limits. We present a Data Quality Assessment and Feedback (DQAF) framework that closes the loop in teleoperation by providing immediate post-episode feedback grounded in semantic task progress and robot telemetry. The framework extracts quality relevant signals such as sub-task progress, motion smoothness, stalls, kinematic limits and converts them into structured quality assessments and actionable natural-language feedback. Unlike binary success or failure feedback, the proposed system explains why an episode is suboptimal and highlights specific behaviors to correct in the next trial. We evaluate the framework through a diagnostic validation study and a pilot user study. In the validation study, the system is compared with a human reviewer during dataset curation, producing rejection reasons and actionable feedback for improvement. In the pilot study with three novice operators across two manipulation tasks, the operator who received the systems immediate, automated post-episode feedback improved faster than those who did not, producing higher-quality demonstrations sooner.

cs.RO

A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons

Vision-Language-Action (VLA) policies have shown promising manipulation capabilities, yet their practical impact is often limited by the reliability demands of real-world deployment. We present a deployment study of an industrial packaging task at Siemens Factory (GWE, Erlangen, Germany), where a robot must pick a transparent accessory bag from a cluttered pile, insert it into the remaining cavity of a cardboard package, and ensure that the bag and its contents remain below the closing plane. Our goal is to understand the practical effort required to adapt a pretrained Pi0.5 policy to a single factory-floor task through iterative fine-tuning and deployment-driven refinement. The pipeline consists of repeated loops of data collection, curation, fine-tuning, evaluation, and targeted recovery data collection. We have accumulated 2535 episodes (10 hours) from the on-site factory settings. In this paper, we contribute an empirical account of a factory-floor VLA deployment, highlighting recurring failure modes and lessons that inform how to improve the deployment workflow.

cs.RO

Radiation Total Dose for PRIMA: Cold Exposure with Alpha Particles

The Probe far-Infrared Mission for Astrophysics (PRIMA) is a far-infrared (24-261 micron wavelengths) probe-class space observatory currently under Phase A study, which promises orders-of-magnitude improvement in mapping speed over its predecessors. PRIMA will field exquisitely sensitive kilopixel arrays of kinetic inductance detectors (KIDs) for the Far-Infrared Enhanced Survey Spectrometer (FIRESS) instrument. PRIMA will orbit in space at the Sun-Earth L2 point, where Planck found the energetic particle flux to be about 300/min/cm2. Thus, the possible effect of a high fluence of energetic particles on the detector sensitivity must be characterized. Previous work has suggested that bombardment of KIDs by ions can reduce the quasiparticle lifetime (Barends et. al. 2009), but the conditions of the experiment were not representative of a detector which is continuously held at sub-Kelvin temperatures in the energetic particle environment of L2 orbit. To better replicate the damage which would be produced by energetic particles in this environment, we developed a fully cryogenic irradiation experiment in which a stepper motor controls a screen which can block or reveal an alpha particle emitter. This setup can be used to irradiate aluminum KID arrays fabricated for FIRESS to well-controlled dose levels. In this work, we calculate the damage dose expected for a 5-year mission in L2 orbit, and we irradiate an array to approximately 62 percent of this level. Before and after irradiation, we measure the quasiparticle lifetimes, resonant frequencies, and quality factors of the detectors.

astro-ph.IM

Auctioning Time to Mitigate Latency Races: Theory and Evidence from Blockchains

High-frequency trading, in both traditional and decentralized markets, induces latency races and redundant order flow as traders spend resources to win time-sensitive opportunities. We show that auctioning artificial time priority can redirect resources away from wasteful speed races toward auction payments. While such waste is difficult to measure in traditional markets, blockchain transactions provide transparent records of these competitive costs through observable duplicate submissions. We study the introduction of Timeboost, a time-priority auction mechanism on Arbitrum, a blockchain that batches transactions before settlement on Ethereum, as a natural experiment. We find that redundant transactions decrease and platform revenue increases relative to comparable networks, consistent with our theoretical predictions.

cs.GT

Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Markets

Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implicit equivalences, and hidden contradictions across markets. We present an agentic AI (AAI) pipeline that autonomously recovers cross-market structure from contract text before prices enter the analysis. The workflow first clusters markets into coherent topical groups using natural-language understanding over contract text and metadata, and then identifies contracts within each cluster, but from different event markets, that exhibit strong dependence or leader--follower relationships. We evaluate this system, along with a natural language inference (NLI) benchmark, on a large prediction market dataset from early 2026. Using resolved outcomes to evaluate identified relations, we find that AAI-identified relations are 62.8\% consistent with exchange-recorded settlements, whereas the NLI benchmark only achieves 40.6\% accuracy. Within clusters, the AAI output is sparse and also remarkably compatible as a signed graph with a frustration rate of 0.324\%. As an application, we show how discovered relations inform semantics-based trading strategies on prediction markets. One such strategy yields 14.12\% net ROI after fees in a two-month period in 2026. Overall, we demonstrate the potential for agentic AI as a structural discovery layer for prediction markets.

cs.AI

SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation

We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language instructions. In this paper, we introduce a hierarchical agentic system consisting of three specialized agents: Manager Agent that decomposes complex user instructions into subtasks; an Action Agent that translates these into structured commands using a Backus-Naur Form (BNF) grammar; and a Reflection Agent that validates alignment between generated actions and the user's original intent. We evaluate SheetMind on the 221-task SheetCopilot Benchmark with GPT-3.5-Turbo. SheetMind achieved 100% execution success and 54.8% functional correctness, exceeding SheetCopilot (44.3%) while maintaining perfect execution reliability. We also conduct ablation study on a separately curated dataset to confirm that the full three-agent configuration consistently outperforms all partial variants. Lastly, we integrate our system into Google Sheets via a Workspace extension.

cs.HC

Information Structures in Stablecoin Markets

Stablecoins have historically depegged due from par to large sales, possibly of speculative nature, or poor reserve asset quality. Using a global game which addresses both concerns, we show that the selling pressure on stablecoin holders increases in the presence of a large sale. While precise public knowledge reduces (increases) the probability of a run when fundamentals are strong (weak), interestingly, more precise private signals increase (reduce) the probability of a run when fundamentals are strong (weak), potentially explaining the stability of opaque stablecoins. The total run probability can be decomposed into components representing risks from large sales and poor collateral. By analyzing how these risk components vary with respect to information uncertainty and fundamentals, we can split the fundamental space into regions based on the type of risk a stablecoin issuer is more prone to. We suggest testable implications and connect our model's implications to real-world applications, including depegging events and the no-questions-asked property of money.

q-fin.TR

The Paradox Of Just-in-Time Liquidity in Decentralized Exchanges: More Providers Can Sometimes Mean Less Liquidity

We study Just-in-time (JIT) liquidity provision in blockchain-based decentralized exchanges. A JIT liquidity provider (LP) monitors pending swap orders in public mempools of blockchains to sandwich orders of their choice with liquidity, depositing right before and withdrawing right after the order. Our game-theoretic model with asymmetrically informed agents reveals that a JIT LP's presence does not always enhance liquidity pool depth, as one might expect. While passive LPs face adverse selection by informed arbitrageurs, a JIT LP's ability to detect pending orders for toxic order flow prior to liquidity provision lets them avoid being adversely selected. JIT LPs thus only provide liquidity to uninformed orders and crowd out passive LPs when order volume is not sufficiently elastic to pool depth, possibly reducing overall market liquidity. We show that using a two-tiered fee structure which transfers a part of a JIT LP's fee revenue to passive LPs or allowing for JIT LPs to compete à la Cournot are potential solutions to mitigate the negative effects of JIT liquidity.

q-fin.GN

Learning on the Job: Self-Rewarding Offline-to-Online Finetuning for Industrial Insertion of Novel Connectors from Vision

Learning-based methods in robotics hold the promise of generalization, but what can be done if a learned policy does not generalize to a new situation? In principle, if an agent can at least evaluate its own success (i.e., with a reward classifier that generalizes well even when the policy does not), it could actively practice the task and finetune the policy in this situation. We study this problem in the setting of industrial insertion tasks, such as inserting connectors in sockets and setting screws. Existing algorithms rely on precise localization of the connector or socket and carefully managed physical setups, such as assembly lines, to succeed at the task. But in unstructured environments such as homes or even some industrial settings, robots cannot rely on precise localization and may be tasked with previously unseen connectors. Offline reinforcement learning on a variety of connector insertion tasks is a potential solution, but what if the robot is tasked with inserting previously unseen connector? In such a scenario, we will still need methods that can robustly solve such tasks with online practice. One of the main observations we make in this work is that, with a suitable representation learning and domain generalization approach, it can be significantly easier for the reward function to generalize to a new but structurally similar task (e.g., inserting a new type of connector) than for the policy. This means that a learned reward function can be used to facilitate the finetuning of the robot's policy in situations where the policy fails to generalize in zero shot, but the reward function generalizes successfully. We show that such an approach can be instantiated in the real world, pretrained on 50 different connectors, and successfully finetuned to new connectors via the learned reward function. Videos can be viewed at https://sites.google.com/view/learningonthejob

cs.RO

PogoDrone: Design, Model, and Control of a Jumping Quadrotor

We present a design, model, and control for a novel jumping-flying robot that is called PogoDrone. The robot is composed of a quadrotor with a passive mechanism for jumping. The robot can continuously jump in place or fly like a normal quadrotor. Jumping in place allows the robot to quickly move and operate very close to the ground. For instance, in agricultural applications, the jumping mechanism allows the robot to take samples of soil. We propose a hybrid controller that switches from attitude to position control to allow the robot to fall horizontally and recover to the original position. We compare the jumping mode with the hovering mode to analyze the energy consumption. In simulations, we evaluate the effect of different factors on energy consumption. In real experiments, we show that our robot can repeatedly impact the ground, jump, and fly in a physical environment.

cs.RO