SearcharxivSearch

arXiv subjects

Jonathan Lee

Publications and source records attributed to Jonathan Lee.

At least 19 recordsLinked to original sources

The 2026 Singapore Consensus on Global AI Safety Research Priorities

Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 2026 Singapore Consensus is an outcome of the second International Scientific Exchange on AI Safety, bringing together over 100 contributors spanning 13 countries from frontier developers, government safety institutes, academia, and civil society. Building on the 2025 report, it presents a global understanding of technical AI safety research problems of top priority, now with a dedicated focus on societal resilience and on managing the risks of increasingly autonomous AI agents.

cs.CY

Continuous coherent spin-frequency metrology in storage rings via resonant beam-driven detection

Precision measurements in storage rings are increasingly limited by the ability to monitor collective spin dynamics coherently over long time scales. Existing polarimetry techniques rely on destructive scattering processes that preclude continuous, non-intercepting tracking of spin evolution and constrain both statistical sensitivity and systematic control. Here we introduce a non-destructive, phase-coherent polarimetry method in which the stored beam polarization is treated as a continuous dynamical observable rather than a quantity inferred from scattering events. Spin-dependent electromagnetic fields generated by a polarized relativistic beam establish a symmetry-selected differential signal on pickup electrodes. This signal is transduced into a narrowband phase modulation of a high-Q resonator interrogated with a coherent probe, while dominant charge-induced backgrounds are rejected through geometric symmetry, helicity reversal, and synchronous demodulation. Controlled spin precession (spin-wheel operation) provides a stable phase reference enabling phase-coherent detection of slow spin evolution. Combined with optimized lattice symmetry and beam cooling, this approach can substantially extend the usable spin coherence time, with values approaching 10^5 s appearing realistic within existing accelerator technology. The resulting readout supports optimal slope-based estimation with T^{-3/2} statistical scaling while eliminating the efficiency penalties inherent to scattering-based polarimetry. For storage-ring EDM experiments, this combination enables sensitivity approaching the level expected within the Standard Model. More broadly, the method establishes a general phase-coherent architecture for collective spin measurements in storage rings, adapting resonant sensing concepts from axion dark-matter searches to charged-particle precision experiments.

hep-ex

A multi-objective optimization framework for sustainable transitions

Achieving a just and sustainable transition requires the pursuit of multiple social and environmental targets. Two primary barriers impede this process: (1) targets are often in conflict with each other, and (2) policies aimed at these targets are commonly planned in isolation, neglecting complex interdependencies in the system. To address these challenges, we propose a general modeling framework that evaluates the holistic impact of policies and decision-making on sustainability targets while capturing system interdependencies in a policy-target network. Inspired by Kauffman's NK fitness landscape, our framework takes the form of a multi-objective optimization model that employs a dynamic evolutionary algorithm in conjunction with network analysis. Our algorithm accounts for tradeoffs between conflicting targets by dynamically reallocating resources to the most impactful and efficient policies. One key finding indicates that increasing resources generally enhances performance, but marginal gains stagnate at a point of diminishing returns. Sensitivity analysis reveals that the system is primarily driven by three factors: budget constraint, network density (interconnectivity), and policy efficacy. This study serves as a foundational step towards developing a decision-support tool that assists policymakers in achieving optimal outcomes for problems with a large number of dynamically interacting targets.

math.DS

Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models

We propose Perceptual Taxonomy, a structured process of scene understanding that first recognizes objects and their spatial configurations, then infers task-relevant properties such as material, affordance, function, and physical attributes to support goal-directed reasoning. While this form of reasoning is fundamental to human cognition, current vision-language benchmarks lack comprehensive evaluation of this ability and instead focus on surface-level recognition or image-text alignment. To address this gap, we introduce Perceptual Taxonomy, a benchmark for physically grounded visual reasoning. We annotate 3173 objects with four property families covering 84 fine-grained attributes. Using these annotations, we construct a multiple-choice question benchmark with 5802 images across both synthetic and real domains. The benchmark contains 28033 template-based questions spanning four types (object description, spatial reasoning, property matching, and taxonomy reasoning), along with 50 expert-crafted questions designed to evaluate models across the full spectrum of perceptual taxonomy reasoning. Experimental results show that leading vision-language models perform well on recognition tasks but degrade by 10 to 20 percent on property-driven questions, especially those requiring multi-step reasoning over structured attributes. These findings highlight a persistent gap in structured visual understanding and the limitations of current models that rely heavily on pattern matching. We also show that providing in-context reasoning examples from simulated scenes improves performance on real-world and expert-curated questions, demonstrating the effectiveness of perceptual-taxonomy-guided prompting.

cs.CV

Towards Robust Mathematical Reasoning

Finding the right north-star metrics is highly critical for advancing the mathematical reasoning capabilities of foundation models, especially given that existing evaluations are either too easy or only focus on getting correct short answers. To address these issues, we present IMO-Bench, a suite of advanced reasoning benchmarks, vetted by a panel of top specialists and that specifically targets the level of the International Mathematical Olympiad (IMO), the most prestigious venue for young mathematicians. IMO-AnswerBench first tests models on 400 diverse Olympiad problems with verifiable short answers. IMO-Proof Bench is the next-level evaluation for proof-writing capabilities, which includes both basic and advanced IMO level problems as well as detailed grading guidelines to facilitate automatic grading. These benchmarks played a crucial role in our historic achievement of the gold-level performance at IMO 2025 with Gemini Deep Think (Luong and Lockhart, 2025). Our model achieved 80.0% on IMO-AnswerBench and 65.7% on the advanced IMO-Proof Bench, surpassing the best non-Gemini models by large margins of 6.9% and 42.4% respectively. We also showed that autograders built with Gemini reasoning correlate well with human evaluations and construct IMO-GradingBench, with 1000 human gradings on proofs, to enable further progress in automatic evaluation of long-form answers. We hope that IMO-Bench will help the community towards advancing robust mathematical reasoning and release it at https://imobench.github.io/.

cs.CL

Quadrotor Navigation using Reinforcement Learning with Privileged Information

This paper presents a reinforcement learning-based quadrotor navigation method that leverages efficient differentiable simulation, novel loss functions, and privileged information to navigate around large obstacles. Prior learning-based methods perform well in scenes that exhibit narrow obstacles, but struggle when the goal location is blocked by large walls or terrain. In contrast, the proposed method utilizes time-of-arrival (ToA) maps as privileged information and a yaw alignment loss to guide the robot around large obstacles. The policy is evaluated in photo-realistic simulation environments containing large obstacles, sharp corners, and dead-ends. Our approach achieves an 86% success rate and outperforms baseline strategies by 34%. We deploy the policy onboard a custom quadrotor in outdoor cluttered environments both during the day and night. The policy is validated across 20 flights, covering 589 meters without collisions at speeds up to 4 m/s.

cs.RO

Status of the Proton EDM Experiment (pEDM)

The Proton EDM Experiment (pEDM) is the first direct search for the proton electric dipole moment (EDM) with the aim of being the first experiment to probe the Standard Model (SM) prediction of any particle EDM. Phase-I of pEDM will achieve $10^{-29} e\cdot$cm, improving current indirect limits by four orders of magnitude. This will establish a new standard of precision in nucleon EDM searches and offer a unique sensitivity to better understand the Strong CP problem. The experiment is ideally positioned to explore physics beyond the Standard Model (BSM), with sensitivity to axionic dark matter via the signal of an oscillating proton EDM and across a wide mass range of BSM models from $\mathcal{O}(1\text{GeV})$ to $\mathcal{O}(10^3\text{TeV})$. Utilizing the frozen-spin technique in a highly symmetric storage ring that leverages existing infrastructure at Brookhaven National Laboratory (BNL), pEDM builds upon the technological foundation and experimental expertise of the highly successful Muon $g$$-$$2$ Experiments. With significant R\&D and prototyping already underway, pEDM is preparing a conceptual design report (CDR) to offer a cost-effective, high-impact path to discovering new sources of CP violation and advancing our understanding of fundamental physics. It will play a vital role in complementing the physics goals of the next-generation collider while simultaneously contributing to sustaining particle physics research and training early-career researchers during gaps between major collider operations.

hep-ex

Validation and Calibration of Energy Models with Real Vehicle Data from Chassis Dynamometer Experiments

Accurate estimation of vehicle fuel consumption typically requires detailed modeling of complex internal powertrain dynamics, often resulting in computationally intensive simulations. However, many transportation applications-such as traffic flow modeling, optimization, and control-require simplified models that are fast, interpretable, and easy to implement, while still maintaining fidelity to physical energy behavior. This work builds upon a recently developed model reduction pipeline that derives physics-like energy models from high-fidelity Autonomie vehicle simulations. These reduced models preserve essential vehicle dynamics, enabling realistic fuel consumption estimation with minimal computational overhead. While the reduced models have demonstrated strong agreement with their Autonomie counterparts, previous validation efforts have been confined to simulation environments. This study extends the validation by comparing the reduced energy model's outputs against real-world vehicle data. Focusing on the MidSUV category, we tune the baseline Autonomie model to closely replicate the characteristics of a Toyota RAV4. We then assess the accuracy of the resulting reduced model in estimating fuel consumption under actual drive conditions. Our findings suggest that, when the reference Autonomie model is properly calibrated, the simplified model produced by the reduction pipeline can provide reliable, semi-principled fuel rate estimates suitable for large-scale transportation applications.

eess.SY

uLayout: Unified Room Layout Estimation for Perspective and Panoramic Images

We present uLayout, a unified model for estimating room layout geometries from both perspective and panoramic images, whereas traditional solutions require different model designs for each image type. The key idea of our solution is to unify both domains into the equirectangular projection, particularly, allocating perspective images into the most suitable latitude coordinate to effectively exploit both domains seamlessly. To address the Field-of-View (FoV) difference between the input domains, we design uLayout with a shared feature extractor with an extra 1D-Convolution layer to condition each domain input differently. This conditioning allows us to efficiently formulate a column-wise feature regression problem regardless of the FoV input. This simple yet effective approach achieves competitive performance with current state-of-the-art solutions and shows for the first time a single end-to-end model for both domains. Extensive experiments in the real-world datasets, LSUN, Matterport3D, PanoContext, and Stanford 2D-3D evidence the contribution of our approach. Code is available at https://github.com/JonathanLee112/uLayout.

cs.CV

Dexterous Manipulation of Deformable Objects via Pneumatic Gripping: Lifting by One End

Manipulating deformable objects in robotic cells is often costly and not widely accessible. However, the use of localized pneumatic gripping systems can enhance accessibility. Current methods that use pneumatic grippers to handle deformable objects struggle with effective lifting. This paper introduces a method for the dexterous lifting of textile deformable objects from one edge, utilizing a previously developed gripper designed for flexible and porous materials. By precisely adjusting the orientation and position of the gripper during the lifting process, we were able to significantly reduce necessary gripping force and minimize object vibration caused by airflow. This method was tested and validated on four materials with varying mass, friction, and flexibility. The proposed approach facilitates the lifting of deformable objects from a conveyor or automated line, even when only one edge is accessible for grasping. Future work will involve integrating a vision system to optimize the manipulation of deformable objects with more complex shapes.

cs.RO

Rapid Quadrotor Navigation in Diverse Environments using an Onboard Depth Camera

Search and rescue environments exhibit challenging 3D geometry (e.g., confined spaces, rubble, and breakdown), which necessitates agile and maneuverable aerial robotic systems. Because these systems are size, weight, and power (SWaP) constrained, rapid navigation is essential for maximizing environment coverage. Onboard autonomy must be robust to prevent collisions, which may endanger rescuers and victims. Prior works have developed high-speed navigation solutions for autonomous aerial systems, but few have considered safety for search and rescue applications. These works have also not demonstrated their approaches in diverse environments. We bridge this gap in the state of the art by developing a reactive planner using forward-arc motion primitives, which leverages a history of RGB-D observations to safely maneuver in close proximity to obstacles. At every planning round, a safe stopping action is scheduled, which is executed if no feasible motion plan is found at the next planning round. The approach is evaluated in thousands of simulations and deployed in diverse environments, including caves and forests. The results demonstrate a 24% increase in success rate compared to state-of-the-art approaches.

cs.RO

Self-training Room Layout Estimation via Geometry-aware Ray-casting

In this paper, we introduce a novel geometry-aware self-training framework for room layout estimation models on unseen scenes with unlabeled data. Our approach utilizes a ray-casting formulation to aggregate multiple estimates from different viewing positions, enabling the computation of reliable pseudo-labels for self-training. In particular, our ray-casting approach enforces multi-view consistency along all ray directions and prioritizes spatial proximity to the camera view for geometry reasoning. As a result, our geometry-aware pseudo-labels effectively handle complex room geometries and occluded walls without relying on assumptions such as Manhattan World or planar room walls. Evaluation on publicly available datasets, including synthetic and real-world scenarios, demonstrates significant improvements in current state-of-the-art layout models without using any human annotation.

cs.CV

Hierarchical Speed Planner for Automated Vehicles: A Framework for Lagrangian Variable Speed Limit in Mixed Autonomy Traffic

This paper introduces a novel control framework for Lagrangian variable speed limits in hybrid traffic flow environments utilizing automated vehicles (AVs). The framework was validated using a fleet of 100 connected automated vehicles as part of the largest coordinated open-road test designed to smooth traffic flow. The framework includes two main components: a high-level controller deployed on the server side, named Speed Planner, and low-level controllers called vehicle controllers deployed on the vehicle side. The Speed Planner designs and updates target speeds for the vehicle controllers based on real-time Traffic State Estimation (TSE) [1]. The Speed Planner comprises two modules: a TSE enhancement module and a target speed design module. The TSE enhancement module is designed to minimize the effects of inherent latency in the received traffic information and to improve the spatial and temporal resolution of the input traffic data. The target speed design module generates target speed profiles with the goal of improving traffic flow. The vehicle controllers are designed to track the target speed meanwhile responding to the surrounding situation. The numerical simulation indicates the performance of the proposed method: the bottleneck throughput has increased by 5.01%, and the speed standard deviation has been reduced by a significant 34.36%. We further showcase an operational study with a description of how the controller was implemented on a field-test with 100 AVs and its comprehensive effects on the traffic flow.

eess.SY

Mesoscale Traffic Forecasting for Real-Time Bottleneck and Shockwave Prediction

Accurate real-time traffic state forecasting plays a pivotal role in traffic control research. In particular, the CIRCLES consortium project necessitates predictive techniques to mitigate the impact of data source delays. After the success of the MegaVanderTest experiment, this paper aims at overcoming the current system limitations and develop a more suited approach to improve the real-time traffic state estimation for the next iterations of the experiment. In this paper, we introduce the SA-LSTM, a deep forecasting method integrating Self-Attention (SA) on the spatial dimension with Long Short-Term Memory (LSTM) yielding state-of-the-art results in real-time mesoscale traffic forecasting. We extend this approach to multi-step forecasting with the n-step SA-LSTM, which outperforms traditional multi-step forecasting methods in the trade-off between short-term and long-term predictions, all while operating in real-time.

cs.LG

Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models

We introduce Jais and Jais-chat, new state-of-the-art Arabic-centric foundation and instruction-tuned open generative large language models (LLMs). The models are based on the GPT-3 decoder-only architecture and are pretrained on a mixture of Arabic and English texts, including source code in various programming languages. With 13 billion parameters, they demonstrate better knowledge and reasoning capabilities in Arabic than any existing open Arabic and multilingual models by a sizable margin, based on extensive evaluation. Moreover, the models are competitive in English compared to English-centric open models of similar size, despite being trained on much less English data. We provide a detailed description of the training, the tuning, the safety alignment, and the evaluation of the models. We release two open versions of the model -- the foundation Jais model, and an instruction-tuned Jais-chat variant -- with the aim of promoting research on Arabic LLMs. Available at https://huggingface.co/inception-mbzuai/jais-13b-chat

cs.CL

Coarse geometry of the Cops and robber game

We introduce two variations of the cops and robber game on graphs. These games yield two invariants in $\mathbb{Z}_+\cup\{\infty\}$ for any connected graph $\Gamma$, the {weak cop number $\mathsf{wcop}(\Gamma)$} and the {strong cop number $\mathsf{scop}(\Gamma)$}. These invariants satisfy that $\mathsf{scop}(\Gamma)\leq\mathsf{wcop}(\Gamma)$. Any graph that is finite or a tree has strong cop number one. These new invariants are preserved under small local perturbations of the graph, specifically, both the weak and strong cop numbers are quasi-isometric invariants of connected graphs. More generally, we prove that if $\Delta$ is a quasi-retract of $\Gamma$ then $\mathsf{wcop}(\Delta)\leq\mathsf{wcop}(\Gamma)$ and $\mathsf{scop}(\Delta)\leq\mathsf{scop}(\Gamma)$. We exhibit families of examples of graphs with arbitrary weak cop number (resp. strong cop number). We prove that hyperbolic graphs have strong cop number one. We also prove that one-ended non-amenable locally-finite vertex-transitive graphs have infinite weak cop number. We raise the question of whether there exists a connected vertex transitive graph with finite weak (resp. strong) cop number different than one.

math.CO

The storage ring proton EDM experiment

We describe a proposal to search for an intrinsic electric dipole moment (EDM) of the proton with a sensitivity of \targetsens, based on the vertical rotation of the polarization of a stored proton beam. The New Physics reach is of order $10^~3$TeV mass scale. Observation of the proton EDM provides the best probe of CP-violation in the Higgs sector, at a level of sensitivity that may be inaccessible to electron-EDM experiments. The improvement in the sensitivity to $\theta_{QCD}$, a parameter crucial in axion and axion dark matter physics, is about three orders of magnitude.

hep-ph

Electric dipole moments and the search for new physics

Static electric dipole moments of nondegenerate systems probe mass scales for physics beyond the Standard Model well beyond those reached directly at high energy colliders. Discrimination between different physics models, however, requires complementary searches in atomic-molecular-and-optical, nuclear and particle physics. In this report, we discuss the current status and prospects in the near future for a compelling suite of such experiments, along with developments needed in the encompassing theoretical framework.

hep-ph