Searcharxiv⌕ Search

arXiv subjects

Yu Zeng

Publications and source records attributed to Yu Zeng.

At least 55 records · Page 3Linked to original sources

Proposal of cavity quantum acoustodynamics platform based on Lithium Niobate-on-Sapphire chip

A scalable hybrid cavity quantum acoustodynamics (QAD) platform is proposed. The architecture integrates superconducting transmon qubits with phononic integrated circuits on a single chip made by lithium niobate-on-sapphire substrate. The platform supports tightly confined and guided phononic modes in unsuspended waveguides and microring structures, while the superconducting qubits reside on the sapphire substrate. Efficient piezoelectric coupling between the phononic modes and transmon qubits can be achieved using interdigital transducers as part of the qubit's shunt capacitance. Numerical calculations demonstrate the feasibility of achieving strong coupling between the phononic microring resonator and the transmon qubit. This hybrid cavity QAD platform opens up new opportunities for quantum information processing and the study of novel quantum acoustic phenomena.

quant-ph↗

Multi-Channel Microwave-to-Optics Conversion Utilizing a Hybrid Photonic-Phononic Waveguide

Efficient and coherent conversion between microwave and optical signals is crucial for a wide range of applications, from quantum information processing to microwave photonics and radar systems. However, existing conversion techniques rely on cavity-enhanced interactions, which limit the bandwidth and calability. Here, we demonstrate the first multi-channel microwave-to-optics conversion by introducing a traveling-wave architecture that leverages a hybrid photonic-phononic waveguide on thin-film lithium niobate (TFLN). Our approach exploits continuous phase-matching rather than discrete resonances, enabling unprecedented operational bandwidths exceeding 40 nm in the optical domain and 250 MHz in the microwave domain. By harnessing the strong piezoelectric and photoelastic effects of TFLN, we achieve coherent conversion between 9 GHz microwave photons and 1550 nm telecom photons via traveling phonons, with an internal efficiency of 2.2% (system efficiency 2.4 *10^-4 ) at room temperature. Remarkably, we demonstrate simultaneous operation of nine conversion channels in a single device. Our converter opens up new opportunities for seamless integration of microwave and photonic technologies, enabling the quantum interface for distributed quantum computing with superconducting quantum processors, high efficient microwave signal processing, and advanced radar applications.

physics.optics↗

Logarithmic light cone, slow entanglement growth, and quantum memory

Effective light cones, characterized by Lieb-Robinson bounds, emerge in nonrelativistic local quantum systems. Here, we present several analytical results derived from logarithmic light cones (LLCs). Possible origins of LLCs include the one-dimensional (1D) disordered XXZ model and a phenomenological model of many-body localization (MBL). In the LLC regime, we prove that, for arbitrary spatial dimensions and any initial pure state, entanglement growth is upper-bounded by logarithmic time with an additional subleading \emph{double-logarithmic} correction -- arising from a real asymptotic solution of the \emph{Lambert W} function -- valid up to the asymptotic time limit. In the context of the 1D disordered XXZ model, this result resolves the ambiguity in distinguishing between logarithmic and power-law fits of entanglement growth in numerical studies; we also propose a falsifiable phenomenological functional form for the entanglement growth that agrees with existing numerical results. We show that information scrambling is logarithmically slow in the LLC regime. Furthermore, we demonstrate that the LLC supports long-lived quantum memories -- quantum codes with macroscopic code distance and lifetimes that scale exponentially with system size -- under unitary time evolution. Our analytical results provide benchmarks for future numerical studies of the MBL regime at large time scales.

quant-ph↗

Cosmos World Foundation Model Platform for Physical AI

Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present the Cosmos World Foundation Model Platform to help developers build customized world models for their Physical AI setups. We position a world foundation model as a general-purpose world model that can be fine-tuned into customized world models for downstream applications. Our platform covers a video curation pipeline, pre-trained world foundation models, examples of post-training of pre-trained world foundation models, and video tokenizers. To help Physical AI builders solve the most critical problems of our society, we make Cosmos open-source and our models open-weight with permissive licenses available via https://github.com/nvidia-cosmos/cosmos-predict1.

cs.CV↗

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

The ability of large language models (LLMs) to utilize external tools has enabled them to tackle an increasingly diverse range of tasks. However, as the tasks become more complex and long-horizon, the intricate tool utilization process may trigger various unexpected errors. Therefore, how to effectively handle such errors, including identifying, diagnosing, and recovering from them, has emerged as a key research direction for advancing tool learning. In this work, we first extensively analyze the types of errors encountered during the function-calling process on several competitive tool evaluation benchmarks. Based on it, we introduce CRITICTOOL, a comprehensive critique evaluation benchmark specialized for tool learning. Building upon a novel evolutionary strategy for dataset construction, CRITICTOOL holds diverse tool-use errors with varying complexities, which better reflects real-world scenarios. We conduct extensive experiments on CRITICTOOL, and validate the generalization and effectiveness of our constructed benchmark strategy. We also provide an in-depth analysis of the tool reflection ability on various LLMs, offering a new perspective on the field of tool learning in LLMs. The code is available at \href{https://github.com/Shellorley0513/CriticTool}{https://github.com/Shellorley0513/CriticTool}.

cs.SE↗

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Effectively retrieving, reasoning and understanding visually rich information remains a challenge for RAG methods. Traditional text-based methods cannot handle visual-related information. On the other hand, current vision-based RAG approaches are often limited by fixed pipelines and frequently struggle to reason effectively due to the insufficient activation of the fundamental capabilities of models. As RL has been proven to be beneficial for model reasoning, we introduce VRAG-RL, a novel RL framework tailored for complex reasoning across visually rich information. With this framework, VLMs interact with search engines, autonomously sampling single-turn or multi-turn reasoning trajectories with the help of visual perception tokens and undergoing continual optimization based on these samples. Our approach highlights key limitations of RL in RAG domains: (i) Prior Multi-modal RAG approaches tend to merely incorporate images into the context, leading to insufficient reasoning token allocation and neglecting visual-specific perception; and (ii) When models interact with search engines, their queries often fail to retrieve relevant information due to the inability to articulate requirements, thereby leading to suboptimal performance. To address these challenges, we define an action space tailored for visually rich inputs, with actions including cropping and scaling, allowing the model to gather information from a coarse-to-fine perspective. Furthermore, to bridge the gap between users' original inquiries and the retriever, we employ a simple yet effective reward that integrates query rewriting and retrieval performance with a model-based reward. Our VRAG-RL optimizes VLMs for RAG tasks using specially designed RL strategies, aligning the model with real-world applications. The code is available at https://github.com/Alibaba-NLP/VRAG.

cs.CL↗

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several challenges. Learning-based approaches depend on small-scale indoor datasets, limiting the scene diversity and layout complexity. While large language models (LLMs) can leverage diverse text-domain knowledge, they struggle with spatial realism, often producing unnatural object placements that fail to respect common sense. Our key insight is that vision perception can bridge this gap by providing realistic spatial guidance that LLMs lack. To this end, we introduce Scenethesis, a training-free agentic framework that integrates LLM-based scene planning with vision-guided layout refinement. Given a text prompt, Scenethesis first employs an LLM to draft a coarse layout. A vision module then refines it by generating an image guidance and extracting scene structure to capture inter-object relations. Next, an optimization module iteratively enforces accurate pose alignment and physical plausibility, preventing artifacts like object penetration and instability. Finally, a judge module verifies spatial coherence. Comprehensive experiments show that Scenethesis generates diverse, realistic, and physically plausible 3D interactive scenes, making it valuable for virtual content creation, simulation environments, and embodied AI research.

cs.CV↗

Learning Universal Features for Generalizable Image Forgery Localization

In recent years, advanced image editing and generation methods have rapidly evolved, making detecting and locating forged image content increasingly challenging. Most existing image forgery detection methods rely on identifying the edited traces left in the image. However, because the traces of different forgeries are distinct, these methods can identify familiar forgeries included in the training data but struggle to handle unseen ones. In response, we present an approach for Generalizable Image Forgery Localization (GIFL). Once trained, our model can detect both seen and unseen forgeries, providing a more practical and efficient solution to counter false information in the era of generative AI. Our method focuses on learning general features from the pristine content rather than traces of specific forgeries, which are relatively consistent across different types of forgeries and therefore can be used as universal features to locate unseen forgeries. Additionally, as existing image forgery datasets are still dominated by traditional hand-crafted forgeries, we construct a new dataset consisting of images edited by various popular deep generative image editing methods to further encourage research in detecting images manipulated by deep generative models. Extensive experimental results show that the proposed approach outperforms state-of-the-art methods in the detection of unseen forgeries and also demonstrates competitive results for seen forgeries. The code and dataset are available at https://github.com/ZhaoHengrun/GIFL.

cs.CV↗

VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

The advancement of Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs) and large vision-language models (LVLMs). However, a rigorous evaluation framework for video CoT reasoning remains absent. Current video benchmarks fail to adequately assess the reasoning process and expose whether failures stem from deficiencies in perception or reasoning capabilities. Therefore, we introduce VCR-Bench, a novel benchmark designed to comprehensively evaluate LVLMs' Video Chain-of-Thought Reasoning capabilities. VCR-Bench comprises 859 videos spanning a variety of video content and durations, along with 1,034 high-quality question-answer pairs. Each pair is manually annotated with a stepwise CoT rationale, where every step is tagged to indicate its association with the perception or reasoning capabilities. Furthermore, we design seven distinct task dimensions and propose the CoT score to assess the entire CoT process based on the stepwise tagged CoT rationals. Extensive experiments on VCR-Bench highlight substantial limitations in current LVLMs. Even the top-performing model, o1, only achieves a 62.8% CoT score and an 56.7% accuracy, while most models score below 40%. Experiments show most models score lower on perception than reasoning steps, revealing LVLMs' key bottleneck in temporal-spatial information processing for complex video reasoning. A robust positive correlation between the CoT score and accuracy confirms the validity of our evaluation framework and underscores the critical role of CoT reasoning in solving complex video reasoning tasks. We hope VCR-Bench to serve as a standardized evaluation framework and expose the actual drawbacks in complex video reasoning task.

cs.CV↗

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmentation, depth, and edge. In the design, the spatial conditional scheme is adaptive and customizable. It allows weighting different conditional inputs differently at different spatial locations. This enables highly controllable world generation and finds use in various world-to-world transfer use cases, including Sim2Real. We conduct extensive evaluations to analyze the proposed model and demonstrate its applications for Physical AI, including robotics Sim2Real and autonomous vehicle data enrichment. We further demonstrate an inference scaling strategy to achieve real-time world generation with an NVIDIA GB200 NVL72 rack. To help accelerate research development in the field, we open-source our models and code at https://github.com/nvidia-cosmos/cosmos-transfer1.

cs.CV↗

Character codegrees, kernels, and Fitting heights of solvable groups

For an irreducible character $χ$ of a finite group $G$, let $\mathrm{cod}(χ):=|G: \ker(χ)|/χ(1)$ denote the codegree of $χ$, and let $\mathrm{cod}(G)$ be the set of irreducible character codegrees of $G$. In this note, we prove that if $\ker(χ)$ is not nilpotent, then there exists an irreducible character $ξ$ of $G$ such that $\ker(ξ)<\ker(χ)$ and $\mathrm{cod}(ξ)> \mathrm{cod}(χ)$. This provides a character codegree analogue of a classical theorem of Broline and Garrison. As a consequence, we obtain that for a nonidentity solvable group $G$, its Fitting height $\ell_{\mathbf{F}}(G)$ does not exceed $|\mathrm{cod}(G)|-1$. Additionally, we provide two other upper bounds for the Fitting height of a solvable group $G$ as follows: $\ell_{\mathbf{F}}(G)\leq \frac{1}{2}(|\mathrm{cod}(G)|+2)$, and $\ell_{\mathbf{F}}(G)\leq 8\log_2(|\mathrm{cod}(G)|)+80$.

math.GR↗

On-chip 7 GHz acousto-optic modulators for visible wavelengths

A chip-integrated acousto-optic phase modulator tailored for visible optical wavelengths has been developed. Utilizing the lithium niobate on sapphire platform, the modulator employs a 7 GHz surface acoustic wave, excited by an interdigital transducer and aligned perpendicular to the waveguide. This design achieves efficient phase modulation of visible light within a compact device length of merely 200 microns, while holds the advantages of easy fabrication and high stability due to simple unsuspended structure. Remarkably, in this high-frequency acoustic regime, the acoustic wavelength becomes comparable to the optical wavelength, resulting in a notable single-sideband modulation behavior. This observation underscores the phase delay effects in the acousto-optics interactions, and opens up new aspects for realizing functional visible photonic devices and its integration with atom- and ion-based quantum platforms.

physics.optics↗

Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

We introduce Edify Image, a family of diffusion models capable of generating photorealistic image content with pixel-perfect accuracy. Edify Image utilizes cascaded pixel-space diffusion models trained using a novel Laplacian diffusion process, in which image signals at different frequency bands are attenuated at varying rates. Edify Image supports a wide range of applications, including text-to-image synthesis, 4K upsampling, ControlNets, 360 HDR panorama generation, and finetuning for image customization.

cs.CV↗

HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion

Hair editing is a critical image synthesis task that aims to edit hair color and hairstyle using text descriptions or reference images, while preserving irrelevant attributes (e.g., identity, background, cloth). Many existing methods are based on StyleGAN to address this task. However, due to the limited spatial distribution of StyleGAN, it struggles with multiple hair color editing and facial preservation. Considering the advancements in diffusion models, we utilize Latent Diffusion Models (LDMs) for hairstyle editing. Our approach introduces Multi-stage Hairstyle Blend (MHB), effectively separating control of hair color and hairstyle in diffusion latent space. Additionally, we train a warping module to align the hair color with the target region. To further enhance multi-color hairstyle editing, we fine-tuned a CLIP model using a multi-color hairstyle dataset. Our method not only tackles the complexity of multi-color hairstyles but also addresses the challenge of preserving original colors during diffusion editing. Extensive experiments showcase the superiority of our method in editing multi-color hairstyles while preserving facial attributes given textual descriptions and reference images.

cs.CV↗

One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation

Diffusion models, praised for their success in generative tasks, are increasingly being applied to robotics, demonstrating exceptional performance in behavior cloning. However, their slow generation process stemming from iterative denoising steps poses a challenge for real-time applications in resource-constrained robotics setups and dynamically changing environments. In this paper, we introduce the One-Step Diffusion Policy (OneDP), a novel approach that distills knowledge from pre-trained diffusion policies into a single-step action generator, significantly accelerating response times for robotic control tasks. We ensure the distilled generator closely aligns with the original policy distribution by minimizing the Kullback-Leibler (KL) divergence along the diffusion chain, requiring only $2\%$-$10\%$ additional pre-training cost for convergence. We evaluated OneDP on 6 challenging simulation tasks as well as 4 self-designed real-world tasks using the Franka robot. The results demonstrate that OneDP not only achieves state-of-the-art success rates but also delivers an order-of-magnitude improvement in inference speed, boosting action prediction frequency from 1.5 Hz to 62 Hz, establishing its potential for dynamic and computationally constrained robotic applications. We share the project page at https://research.nvidia.com/labs/dir/onedp/.

cs.RO↗

Bounded light cone and robust topological order out of equilibrium

The ground state degeneracy of topologically ordered gapped Hamiltonians is the bedrock for self-correcting quantum memories, which are unfortunately not stable away from equilibrium even at zero temperature. This plague precludes practical robust self-correction since stability at zero temperature is a prerequisite for finite-temperature robustness. In this work, we show that the emergence of a bounded light cone renders the unitary time evolution a quasi-adiabatic continuation that preserves topological order, with the initial ground space retaining its macroscopic distance at all times as a quantum code. We also show how bounded light cones can emerge through suitable perturbations in Kitaev's toric code and honeycomb model. Our results suggest that topological orders and self-correcting quantum memories can be dynamically robust at zero temperature.

quant-ph↗

JeDi: Joint-Image Diffusion Models for Finetuning-Free Personalized Text-to-Image Generation

Personalized text-to-image generation models enable users to create images that depict their individual possessions in diverse scenes, finding applications in various domains. To achieve the personalization capability, existing methods rely on finetuning a text-to-image foundation model on a user's custom dataset, which can be non-trivial for general users, resource-intensive, and time-consuming. Despite attempts to develop finetuning-free methods, their generation quality is much lower compared to their finetuning counterparts. In this paper, we propose Joint-Image Diffusion (\jedi), an effective technique for learning a finetuning-free personalization model. Our key idea is to learn the joint distribution of multiple related text-image pairs that share a common subject. To facilitate learning, we propose a scalable synthetic dataset generation technique. Once trained, our model enables fast and easy personalization at test time by simply using reference images as input during the sampling process. Our approach does not require any expensive optimization process or additional modules and can faithfully preserve the identity represented by any number of reference images. Experimental results show that our model achieves state-of-the-art generation quality, both quantitatively and qualitatively, significantly outperforming both the prior finetuning-based and finetuning-free personalization baselines.

cs.CV↗