SearcharxivSearch

arXiv subjects

Mengdi Zhao

Publications and source records attributed to Mengdi Zhao.

15 recordsLinked to original sources

Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, where System 2 performs high-level reasoning and System 1 handles low-level control. We refer to System 2 as the embodied brain, the cognitive core for decision-making in manipulation. Although evaluating this embodied brain is crucial, existing benchmarks mainly measure execution success or cover only limited aspects of high-level cognition and task realism. We introduce RoboBench, a benchmark for evaluating multimodal large language models (MLLMs) as embodied brains. RoboBench covers five dimensions: Instruction Comprehension, Perception Reasoning, Generalized Planning, Affordance Prediction, and Failure Analysis. It spans 14 capabilities, 25 tasks, and 6,092 QA pairs. To improve realism, it draws from large-scale real robotic data and in-house collection across diverse embodiments, attribute-rich objects, multi-view scenes, and memory-driven navigation. For planning, RoboBench introduces an MLLM-as-world-simulator framework that assesses whether predicted plans can achieve critical object-state changes under physical and visual constraints, enabling more faithful evaluation of long-horizon reasoning than symbolic matching. Experiments on 18 state-of-the-art MLLMs reveal persistent limitations in implicit instruction understanding, spatiotemporal reasoning, cross-scenario planning, fine-grained affordance understanding, and failure diagnosis. We further analyze how embodied cognitive abilities relate to downstream robotic control. RoboBench offers a comprehensive scaffold for quantifying high-level cognition and guiding next-generation MLLMs toward more robust robotic intelligence.

cs.RO

An integrated photonics platform for high-speed, ultrahigh-extinction, many-channel quantum control

High-fidelity control of the thousands to millions of programmable qubits needed for utility-scale quantum computers presents a formidable challenge for control systems. In leading atomic systems, control is optical: UV-NIR beams must be fanned out over numerous spatial channels and modulated to implement gates. While photonic integrated circuits (PICs) offer a potentially scalable solution, they also need to simultaneously feature high-speed and high-extinction modulation, strong inter-channel isolation, and broad wavelength compatibility. Here, we introduce and experimentally validate a foundry-fabricated PIC platform that overcomes these limitations. Designed for Rubidium-87 neutral atom quantum computers, our 8-channel PICs, fabricated on a 200-mm wafer process, demonstrate an advanced combination of performance metrics. At the 795 nm single-qubit gate wavelength, we achieve a mean extinction ratio (ER) of 71.4 $\pm$ 1.1 dB, nearest-neighbor on-chip crosstalk of -68.0 $\pm$ 1.0 dB, and -50.8 $\pm$ 0.2 dB after parallel beam delivery in free-space. This high-performance operation extends to the 420 nm and 1013 nm wavelengths for two-qubit Rydberg gates, showing ERs of 42.4 dB (detector-limited) and 61.5 dB, respectively. The devices exhibit 10-90% rise times of 26 $\pm$ 7 ns, achieve dynamic switching to -60 dB levels within microsecond timescales, and show pulse stability errors at the $10^{-3}$ level. This work establishes a scalable platform for developing advanced large-scale optical control required in fault-tolerant quantum computers and other precision technologies.

quant-ph

RoboBrain 2.0 Technical Report

We introduce RoboBrain 2.0, our latest generation of embodied vision-language foundation models, designed to unify perception, reasoning, and planning for complex embodied tasks in physical environments. It comes in two variants: a lightweight 7B model and a full-scale 32B model, featuring a heterogeneous architecture with a vision encoder and a language model. Despite its compact size, RoboBrain 2.0 achieves strong performance across a wide spectrum of embodied reasoning tasks. On both spatial and temporal benchmarks, the 32B variant achieves leading results, surpassing prior open-source and proprietary models. In particular, it supports key real-world embodied AI capabilities, including spatial understanding (e.g., affordance prediction, spatial referring, trajectory forecasting) and temporal decision-making (e.g., closed-loop interaction, multi-agent long-horizon planning, and scene graph updating). This report details the model architecture, data construction, multi-stage training strategies, infrastructure and practical applications. We hope RoboBrain 2.0 advances embodied AI research and serves as a practical step toward building generalist embodied agents. The code, checkpoint and benchmark are available at https://superrobobrain.github.io.

cs.RO

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

A fundamental objective of manipulation policy design is to endow robots to comprehend human instructions, reason about scene cues, and execute generalized actions in dynamic environments. Recent autoregressive vision-language-action (VLA) methods inherit common-sense reasoning capabilities from vision-language models (VLMs) for next action-token prediction. However, these methods quantize actions into discrete bins, which disrupts the continuity required for precise control. In contrast, existing diffusion-based VLA methods incorporate an additional diffusion head to predict continuous actions solely conditioned on feature representations extracted by the VLM, without fully leveraging the VLM's pretrained reasoning capabilities through token-level generation. To address these limitations, we introduce HybridVLA, a unified framework that absorbs the continuous nature of diffusion-based actions and the contextual reasoning of autoregression within a single large language model. To mitigate interference between the two generation paradigms, we propose a collaborative training recipe that seamlessly incorporates diffusion denoising into the next-token prediction process. With this recipe, we find these two action prediction methods not only reinforce each other but also exhibit varying strength across different tasks. Therefore, we design a collaborative action ensemble mechanism that adaptively fuses both predictions, leading to more robust control. HybridVLA outperforms previous state-of-the-art VLA methods by 14\% and 19\% in mean success rate on simulation and real-world tasks, respectively, while demonstrating stable manipulation in unseen configurations.

cs.CV

RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Recent advancements in Multimodal Large Language Models (MLLMs) have shown remarkable capabilities across various multimodal contexts. However, their application in robotic scenarios, particularly for long-horizon manipulation tasks, reveals significant limitations. These limitations arise from the current MLLMs lacking three essential robotic brain capabilities: Planning Capability, which involves decomposing complex manipulation instructions into manageable sub-tasks; Affordance Perception, the ability to recognize and interpret the affordances of interactive objects; and Trajectory Prediction, the foresight to anticipate the complete manipulation trajectory necessary for successful execution. To enhance the robotic brain's core capabilities from abstract to concrete, we introduce ShareRobot, a high-quality heterogeneous dataset that labels multi-dimensional information such as task planning, object affordance, and end-effector trajectory. ShareRobot's diversity and accuracy have been meticulously refined by three human annotators. Building on this dataset, we developed RoboBrain, an MLLM-based model that combines robotic and general multi-modal data, utilizes a multi-stage training strategy, and incorporates long videos and high-resolution images to improve its robotic manipulation capabilities. Extensive experiments demonstrate that RoboBrain achieves state-of-the-art performance across various robotic tasks, highlighting its potential to advance robotic brain capabilities.

cs.RO

Spiral resonator referenced on-chip low noise microwave generation

In recent years, miniaturization and integration of photonic microwave oscillators by optical frequency division approach have witnessed rapid progress. In this work, we report on-chip low phase noise photonic microwave generation based on a planar chip design. Dual lasers are co-locked to a silicon nitride spiral resonator and their relative phase noise is measured below the cavity thermal noise limit, resulting in record low on-chip relative optical phase noise. A broadband integrated electro-optic comb up to 3.43 THz (27 nm) bandwidth is utilized to divide down the relative phase noise of the spiral resonator referenced lasers to the microwave domain. All-around record-low phase noise is achieved for planar chip-based photonic microwave oscillators from 10 Hz to 10 kHz offsets. The planar chip design, high technology-readiness level, foundry-ready processing, combined with the exceptional phase noise performance for our work represent a major advance of integrated photonic microwave oscillators.

physics.optics

ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model

Tokenizer is an essential component for large language models (LLMs), and a tokenizer with a high compression rate can improve the model's representation and processing efficiency. However, the tokenizer cannot ensure high compression rate in all scenarios, and an increase in the average input and output lengths will increases the training and inference costs of the model. Therefore, it is crucial to find ways to improve the model's efficiency with minimal cost while maintaining the model's performance. In this work, we propose a method to improve model representation and processing efficiency by replacing the tokenizers of LLMs. We propose replacing and reinitializing the parameters of the model's input and output layers with the parameters of the original model, and training these parameters while keeping other parameters fixed. We conducted experiments on different LLMs, and the results show that our method can maintain the performance of the model after replacing the tokenizer, while significantly improving the decoding speed for long texts.

cs.CL

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies

In recent years, with the rapid application of large language models across various fields, the scale of these models has gradually increased, and the resources required for their pre-training have grown exponentially. Training an LLM from scratch will cost a lot of computation resources while scaling up from a smaller model is a more efficient approach and has thus attracted significant attention. In this paper, we present AquilaMoE, a cutting-edge bilingual 8*16B Mixture of Experts (MoE) language model that has 8 experts with 16 billion parameters each and is developed using an innovative training methodology called EfficientScale. This approach optimizes performance while minimizing data requirements through a two-stage process. The first stage, termed Scale-Up, initializes the larger model with weights from a pre-trained smaller model, enabling substantial knowledge transfer and continuous pretraining with significantly less data. The second stage, Scale-Out, uses a pre-trained dense model to initialize the MoE experts, further enhancing knowledge transfer and performance. Extensive validation experiments on 1.8B and 7B models compared various initialization schemes, achieving models that maintain and reduce loss during continuous pretraining. Utilizing the optimal scheme, we successfully trained a 16B model and subsequently the 8*16B AquilaMoE model, demonstrating significant improvements in performance and training efficiency.

cs.CL

Quantum correlated photons via a passive nonlinear microcavity

Photons, by nature, typically do not exhibit interactions with each other. Creating photon-photon interactions holds immense importance in both fundamental physics and quantum technologies. Currently, such interactions have only been achieved indirectly as mediated by atomic-like quantum emitters with resonant photon-atom interactions. However, the use of these indirect interactions presents substantial fundamental challenges that impede scaling and practical applications. Here we demonstrate creation of non-classical photon correlations, including photon anti-bunching, via a passive InGaP photonic integrated circuit. Our approach employs the quantum interference between uncorrelated light and the two-photon bound state, the latter of which arises from the $χ^{(2)}$-mediated photon interaction. Our work opens a new route in controlling quantum light by harnessing highly-engineerable bulk optical nonlinearities, which has significant implications for nonlinear optical quantum information processing and quantum networking.

quant-ph

A rigid, low-loss fiber-optic coupler for cryogenic photonics

Recent developments in quantum light-matter coupled systems and quantum transducers have highlighted the need for cryogenic optical measurements. In this study, we present a mechanically-rigid fiber-optic coupler with a coupling efficiency of over 50% for telecom wavelength light at cryogenic temperatures. Our method enables sensitive photonic device measurements that are alignment-free and immune to mechanical vibrations in cryogenic setups.

physics.optics

Observation of photon-phonon correlations via dissipative filtering

Cavity-optomechanics enables photon-phonon interaction and correlations by harnessing the radiation-pressure force. Here, we realize a ``cavity-in-a-membrane'' optomechanical architecture which allows detection of the motion of lithographically-defined, ultrathin membranes via an integrated optical cavity. Using a dissipative filtering method, we are able to eliminate the probe light in situ and observe photon-phonon correlations associated with the low-frequency membrane mechanical mode. The developed method is generally applicable for study of low-frequency light scattering processes where conventional frequency-selective filtering is unfeasible.

physics.optics

InGaP quantum nanophotonic integrated circuits with 1.5% nonlinearity-to-loss ratio

Optical nonlinearity plays a pivotal role in quantum information processing using photons, from heralded single-photon sources, coherent wavelength conversion to long-sought quantum repeaters. Despite the availability of strong dipole coupling to quantum emitters, achieving strong bulk optical nonlinearity is highly desirable. Here, we realize quantum nanophotonic integrated circuits in thin-film InGaP with a record-high ratio of $1.5\%$ between the single-photon nonlinear coupling rate ($g/2π=11.2$ MHz) and cavity-photon loss rate . We demonstrate second-harmonic generation with an efficiency of $71200\pm10300\%$/W in the InGaP photonic circuit and photon-pair generation via degenerate spontaneous parametric down-conversion with an ultrahigh rate exceeding 27.5 MHz/$μ$W -- an order of magnitude improvement of the state-of-the-art -- and a large coincidence-to-accidental ratio up to $1.4\times 10^4$. Our work shows InGaP as a potentially transcending platform for quantum nonlinear optics and quantum information applications.

physics.optics

Observation of phonon trapping in the continuum with topological charges

Phonon trapping has an immense impact in many areas of science and technology, from the antennas of interferometric gravitational wave detectors to chip-scale quantum micro- and nano-mechanical oscillators. It usually relies on the mechanical suspension--an approach, while isolating selected vibrational modes, leads to serious drawbacks for interrogation of the trapped phonons, including limited heat capacity and excess noises via measurements. To circumvent these constraints, we realize a new paradigm of phonon trapping using mechanical bound states in the continuum (BICs) with topological features and conducted an in-depth characterization of the mechanical losses both at room and cryogenic temperatures. Our findings of mechanical BICs combining the microwave frequency and macroscopic size unveil a unique platform for realizing mechanical oscillators in both classical and quantum regimes. The paradigm of mechanical BICs might lead to unprecedented sensing modalities for applications such as rare-event searches and the exploration of the foundations of quantum mechanics in unreached parameter spaces.

cond-mat.mes-hall

High-efficiency fiber-to-chip interface for aluminum nitride quantum photonics

Integrated nonlinear photonic circuits received rapid development in recent years, providing all-optical functionalities enabled by cavity-enhanced photon-photon interaction for classical and quantum applications. A high-efficiency fiber-to-chip interface is key to the use of these integrated photonic circuits for quantum information tasks, as photon loss is a major source that weakens quantum protocols. Here, overcoming material and fabrication limitation of thin-film aluminum nitride by adopting a stepwise waveguiding scheme, we demonstrate low-loss adiabatic fiber-optic couplers in aluminum nitride films with a substantial thickness (600 nm) for optimized nonlinear photon interaction. For telecom (1550 nm) and near-visible (780 nm) transverse magnetic-polarized light, the measured insertion loss of the fiber-optic coupler is -0.97 dB and -2.6 dB, respectively. Our results will facilitate the use of aluminum nitride integrated photonic circuits as efficient quantum resources for generation of entangled photons and squeezed light on microchips.

physics.app-ph

Mechanical bound states in the continuum for macroscopic optomechanics

Bound states in the continuum (BICs), an emerging type of long-lived resonances different from the cavity-based ones, have been explored in several classical systems, including photonic crystals and surface acoustic waves. Here, we reveal symmetry-protected mechanical BICs in the structure of slab-on-substrate optomechanical crystals. Using a group theory approach, we identified all the mechanical BICs at the $Γ$ point in optomechanical crystals with $C_{4v}$ and $C_{6v}$ symmetries as examples, and analyzed their coupling with the co-localized optical BICs and guided resonances due to both moving boundary and photo-elastic effects. We verified the theoretical analysis with numerical simulations of specific optomechanical crystals which support substantial optomechanical interactions between the mechanical BICs and optical resonances. Due to the unique features of high-$Q$, large-size mechanical BICs and substrate-enabled thermal dissipation, this architecture of slab-on-substrate optomechanical crystals might be useful for exploring macroscopic quantum mechanical physics and enabling new applications such as high-throughput sensing and free-space beam steering.

physics.optics