SearcharxivSearch

arXiv subjects

Eric Zhu

Publications and source records attributed to Eric Zhu.

11 recordsLinked to original sources

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existing approaches typically require large amounts of human or reward model evaluations. This limitation reduces the practicality of diffusion RLHF in realworld settings where feedback is the primary bottleneck. In this paper, we propose two complementary strategies that substantially improve the feedback efficiency of diffusion RLHF while preserving generalization to unseen prompts. Our key observation is that reward information in diffusion trajectories is unevenly distributed: not all denoising timesteps or trajectories contribute equally to learning from a reward signal. By emphasizing informative timesteps and trajectories during optimization, we obtain more effective gradient updates. First, we introduce a per-timestep weighting scheme that reweights denoising steps during policy optimization. We theoretically connect this weighting to the optimal convergence properties of proximal policy optimization (PPO) and approximate the resulting weighting trend empirically. Second, we introduce a replay mechanism that prioritizes informative trajectories, enabling the model to reuse past samples instead of repeatedly querying new rewards. Together, these strategies significantly improve the feedback efficiency of diffusion RLHF. Under identical hyperparameter settings, our approach achieves up to a 6$\times$ improvement in sample efficiency compared to widely used diffusion RLHF baselines.

cs.LG

A matter-wave Fabry-P\'erot cavity in the ultrastrong driving regime

When the length of an optical cavity is modulated, theory predicts exponential concentration of energy around particular space-time trajectories. Viewed stroboscopically, photons in such a driven cavity propagate as if in a curved spacetime, with black hole and white hole event horizons corresponding to unstable and stable fixed points of the evolution. Such phenomena have resisted direct experimental realization due to the difficulty of relativistically accelerating massive cavity mirrors. We report results of an experiment which overcomes this limitation by exchanging the roles of light and matter. A matter wave endowed with quasi-relativistic dispersion is confined between two barriers made of light, one of which is periodically translated at speeds comparable to the matter wave group velocity. In this strongly-modulated cavity we observe the emergence of the predicted bright and dark fixed point trajectories, and demonstrate that changing the modulation waveform can vary the number of fixed points and exchange their stability character. We observe signatures of nontrivial dynamics beyond those predicted for photons, and attribute them to residual curvature in the dispersion relation. In addition to experimentally realizing and characterizing cavity dynamics in the ultra-strong driving regime, these results point the way to implementations of related dynamics in electro-optic materials, with potential applications in pulse generation and signal compression.

cond-mat.quant-gas

Brauer-Manin Obstruction on Generalized Kummer Varieties

Given an abelian variety $A$ over a number field, we consider the generalized Kummer varieties of $A$ coming from quotients of $A$ by an automorphism of prime order $p > 2$. We prove that the Brauer-Manin obstruction on these generalized Kummer varieties only can come from the $p$-primary part of the Brauer group. This is applied to show that certain families of such varieties have no Brauer-Manin obstruction to the local-global principle.

math.NT

Magentic-UI: Towards Human-in-the-loop Agentic Systems

AI agents powered by large language models are increasingly capable of autonomously completing complex, multi-step tasks using external tools. Yet, they still fall short of human-level performance in most domains including computer use, software development, and research. Their growing autonomy and ability to interact with the outside world, also introduces safety and security risks including potentially misaligned actions and adversarial manipulation. We argue that human-in-the-loop agentic systems offer a promising path forward, combining human oversight and control with AI efficiency to unlock productivity from imperfect systems. We introduce Magentic-UI, an open-source web interface for developing and studying human-agent interaction. Built on a flexible multi-agent architecture, Magentic-UI supports web browsing, code execution, and file manipulation, and can be extended with diverse tools via Model Context Protocol (MCP). Moreover, Magentic-UI presents six interaction mechanisms for enabling effective, low-cost human involvement: co-planning, co-tasking, multi-tasking, action guards, and long-term memory. We evaluate Magentic-UI across four dimensions: autonomous task completion on agentic benchmarks, simulated user testing of its interaction capabilities, qualitative studies with real users, and targeted safety assessments. Our findings highlight Magentic-UI's potential to advance safe and efficient human-agent collaboration.

cs.AI

Continuously trapped matter-wave interferometry in magic Floquet-Bloch band structures

Trapped matter-wave interferometry offers the promise of compact high-precision local force sensing. However, noise in the trap itself can introduce new systematic errors which are absent in traditional free-fall interferometers. We describe and demonstrate an intrinsically noise-tolerant Floquet-engineered platform for continuously trapped atom interferometry. A non-interacting degenerate quantum gas undergoes position-space Bloch oscillations through an amplitude-modulated optical lattice, whose resulting Floquet-Bloch band structure includes Landau-Zener beamsplitters and Bragg mirrors, forming the components of a Mach-Zehnder interferometric force sensor. We identify, realize, and experimentally characterize magic band structures, analogous to the magic wavelengths employed in optical lattice clocks, for which the interferometric phase is insensitive to lattice intensity noise. We leverage the intrinsic programmability of the Floquet band synthesis approach to demonstrate a variety of interferometer structures, highlighting the potential of this technique for quantum force sensors which are tunable, compact, simple, and robust.

physics.atom-ph

NeRF-Aug: Data Augmentation for Robotics with Neural Radiance Fields

Training a policy that can generalize to unknown objects is a long standing challenge within the field of robotics. The performance of a policy often drops significantly in situations where an object in the scene was not seen during training. To solve this problem, we present NeRF-Aug, a novel method that is capable of teaching a policy to interact with objects that are not present in the dataset. This approach differs from existing approaches by leveraging the speed, photorealism, and 3D consistency of a neural radiance field for augmentation. NeRF-Aug both creates more photorealistic data and runs 63% faster than existing methods. We demonstrate the effectiveness of our method on 5 tasks with 9 novel objects that are not present in the expert demonstrations. We achieve an average performance boost of 55.6% when comparing our method to the next best method. You can see video results at https://nerf-aug.github.io.

cs.RO

DAMP: Doubly Aligned Multilingual Parser for Task-Oriented Dialogue

Modern virtual assistants use internal semantic parsing engines to convert user utterances to actionable commands. However, prior work has demonstrated that semantic parsing is a difficult multilingual transfer task with low transfer efficiency compared to other tasks. In global markets such as India and Latin America, this is a critical issue as switching between languages is prevalent for bilingual users. In this work we dramatically improve the zero-shot performance of a multilingual and codeswitched semantic parsing system using two stages of multilingual alignment. First, we show that constrastive alignment pretraining improves both English performance and transfer efficiency. We then introduce a constrained optimization approach for hyperparameter-free adversarial alignment during finetuning. Our Doubly Aligned Multilingual Parser (DAMP) improves mBERT transfer performance by 3x, 6x, and 81x on the Spanglish, Hinglish and Multilingual Task Oriented Parsing benchmarks respectively and outperforms XLM-R and mT5-Large using 3.2x fewer parameters.

cs.CL

Density of Periodic Points for Lattès maps over Finite Fields

Let $L_d$ be the Lattès map associated to the multiplication-by-$d$ endomorphism of an elliptic curve $E$ defined over a finite field $\mathbb{F}_q$. We determine the density $δ(L_d,q)$ of periodic points for $L_d$ in $\mathbb{P}^1(\mathbb{F}_q)$. We show that the periodic point densities $δ(L_d,q^n)$ converge as $n \rightarrow \infty$ along certain arithmetic progressions, and compute simple explicit formulas for $δ(L_\ell,q)$ when $\ell$ is a prime and $E$ belongs to a special family of supersingular elliptic curves.

math.NT

Learning for Microrobot Exploration: Model-based Locomotion, Sparse-robust Navigation, and Low-power Deep Classification

Building intelligent autonomous systems at any scale is challenging. The sensing and computation constraints of a microrobot platform make the problems harder. We present improvements to learning-based methods for on-board learning of locomotion, classification, and navigation of microrobots. We show how simulated locomotion can be achieved with model-based reinforcement learning via on-board sensor data distilled into control. Next, we introduce a sparse, linear detector and a Dynamic Thresholding method to FAST Visual Odometry for improved navigation in the noisy regime of mm scale imagery. We end with a new image classifier capable of classification with fewer than one million multiply-and-accumulate (MAC) operations by combining fast downsampling, efficient layer structures and hard activation functions. These are promising steps toward using state-of-the-art algorithms in the power-limited world of edge-intelligence and microrobots.

cs.RO

Correlated Photon Pair Production by Spontaneous Parametric Down Conversion in Quasi-Phase-Matched AlGaAs superlattice Waveguides using a Continuous Wave Pump

We report on the demonstration of correlated photon pair generation in quasi-phase-matched superlattice AlGaAs waveguides with a high coincidence-to-accidental ratio (CAR); with a continuous (CW) pump, the observed CAR (>100) is more than two order of magnitudes improvement over previously reported spontaneous down conversion (SPDC) schemes in AlGaAs waveguides with a pulsed pump.

quant-ph