Searcharxiv⌕ Search

arXiv subjects

Yu Fang

Publications and source records attributed to Yu Fang.

At least 37 records · Page 2Linked to original sources

Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation

Prompt-free image segmentation aims to generate accurate masks without manual guidance. Typical pre-trained models, notably Segmentation Anything Model (SAM), generate prompts directly at a single granularity level. However, this approach has two limitations: (1) Localizability, lacking mechanisms for autonomous region localization; (2) Scalability, limited fine-grained modeling at high resolution. To address these challenges, we introduce Granular Computing-driven SAM (Grc-SAM), a coarse-to-fine framework motivated by Granular Computing (GrC). First, the coarse stage adaptively extracts high-response regions from features to achieve precise foreground localization and reduce reliance on external prompts. Second, the fine stage applies finer patch partitioning with sparse local swin-style attention to enhance detail modeling and enable high-resolution segmentation. Third, refined masks are encoded as latent prompt embeddings for the SAM decoder, replacing handcrafted prompts with an automated reasoning process. By integrating multi-granularity attention, Grc-SAM bridges granular computing with vision transformers. Extensive experimental results demonstrate Grc-SAM outperforms baseline methods in both accuracy and scalability. It offers a unique granular computational perspective for prompt-free segmentation.

cs.CV↗

Imitation Learning Policy based on Multi-Step Consistent Integration Shortcut Model

The wide application of flow-matching methods has greatly promoted the development of robot imitation learning. However, these methods all face the problem of high inference time. To address this issue, researchers have proposed distillation methods and consistency methods, but the performance of these methods still struggles to compete with that of the original diffusion models and flow-matching models. In this article, we propose a one-step shortcut method with multi-step integration for robot imitation learning. To balance the inference speed and performance, we extend the multi-step consistency loss on the basis of the shortcut model, split the one-step loss into multi-step losses, and improve the performance of one-step inference. Secondly, to solve the problem of unstable optimization of the multi-step loss and the original flow-matching loss, we propose an adaptive gradient allocation method to enhance the stability of the learning process. Finally, we evaluate the proposed method in two simulation benchmarks and five real-world environment tasks. The experimental results verify the effectiveness of the proposed algorithm.

cs.RO↗

Knowledge graph-based personalized multimodal recommendation fusion framework

In the contemporary age characterized by information abundance, rapid advancements in artificial intelligence have rendered recommendation systems indispensable. Conventional recommendation methodologies based on collaborative filtering or individual attributes encounter deficiencies in capturing nuanced user interests. Knowledge graphs and multimodal data integration offer enhanced representations of users and items with greater richness and precision. This paper reviews existing multimodal knowledge graph recommendation frameworks, identifying shortcomings in modal interaction and higher-order dependency modeling. We propose the Cross-Graph Cross-Modal Mutual Information-Driven Unified Knowledge Graph Learning and Recommendation Framework (CrossGMMI-DUKGLR), which employs pre-trained visual-text alignment models for feature extraction, achieves fine-grained modality fusion through multi-head cross-attention, and propagates higher-order adjacency information via graph attention networks.

cs.IR↗

MuteSwap: Visual-informed Silent Video Identity Conversion

Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC.

cs.SD↗

Towards Fault-Tolerant Quantum Deep Learning: Designing and Analyzing Quantum ResNet and Transformer with Quantum Arithmetic and Linear Algebra Primitives

Achieving a practical quantum speedup for deep neural networks (DNNs) remains a central yet elusive goal, hindered by the dual challenges of constructing deep architectures and the prohibitive overhead of data loading and measurement. We introduce a framework to overcome these barriers, specifically targeting an asymptotic speedup with respect to the large input dimensions of modern DNNs (e.g., sequence length or image size). Our framework enables the design of multi-layer Quantum ResNet and Quantum Transformer models by strategically decomposing tasks: computationally intensive operations on the large input dimension are assigned to quantum linear algebra subroutines, while operations on the smaller, fixed feature dimension are handled by efficient quantum arithmetic. A cornerstone of our approach is a novel data transfer protocol, Discrete Chebyshev Decomposition (DCD), which facilitates this modularity. Numerical validation reveals a pivotal insight: the measurement cost required to maintain a target accuracy scales sublinearly with the input dimension. This sublinear scaling is the key to preserving the quantum advantage, ensuring that I/O overhead does not nullify the computational gains. A rigorous resource analysis further corroborates the superiority of our models in both efficiency and flexibility. Powered by this targeted acceleration strategy and the efficiency of DCD, our framework establishes a viable path toward scalable quantum deep learning.

quant-ph↗

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - synthetic robot data generated from video world models. DreamGen leverages state-of-the-art image-to-video generative models, adapting them to the target robot embodiment to produce photorealistic synthetic videos of familiar or novel tasks in diverse environments. Since these models generate only videos, we recover pseudo-action sequences using either a latent action model or an inverse-dynamics model (IDM). Despite its simplicity, DreamGen unlocks strong behavior and environment generalization: a humanoid robot can perform 22 new behaviors in both seen and unseen environments, while requiring teleoperation data from only a single pick-and-place task in one environment. To evaluate the pipeline systematically, we introduce DreamGen Bench, a video generation benchmark that shows a strong correlation between benchmark performance and downstream policy success. Our work establishes a promising new axis for scaling robot learning well beyond manual data collection. Code available at https://github.com/NVIDIA/GR00T-Dreams.

cs.RO↗

FLARE: Robot Learning with Implicit World Modeling

We introduce $\textbf{F}$uture $\textbf{LA}$tent $\textbf{RE}$presentation Alignment ($\textbf{FLARE}$), a novel framework that integrates predictive latent world modeling into robot policy learning. By aligning features from a diffusion transformer with latent embeddings of future observations, $\textbf{FLARE}$ enables a diffusion transformer policy to anticipate latent representations of future observations, allowing it to reason about long-term consequences while generating actions. Remarkably lightweight, $\textbf{FLARE}$ requires only minimal architectural modifications -- adding a few tokens to standard vision-language-action (VLA) models -- yet delivers substantial performance gains. Across two challenging multitask simulation imitation learning benchmarks spanning single-arm and humanoid tabletop manipulation, $\textbf{FLARE}$ achieves state-of-the-art performance, outperforming prior policy learning baselines by up to 26%. Moreover, $\textbf{FLARE}$ unlocks the ability to co-train with human egocentric video demonstrations without action labels, significantly boosting policy generalization to a novel object with unseen geometry with as few as a single robot demonstration. Our results establish $\textbf{FLARE}$ as a general and scalable approach for combining implicit world modeling with high-frequency robotic control.

cs.RO↗

PolyQROM: Orthogonal-Polynomial-Based Quantum Reduced-Order Model for Flow Field Analysis

Quantum computing promises exponential acceleration for fluid flow simulations, yet the measurement overhead required to extract flow features from quantum-encoded flow field data fundamentally undermines this advantage--a critical challenge termed the ``output problem''. To address this, we propose an orthogonal-polynomial-based quantum reduced-order model (PolyQROM) that integrates orthogonal polynomial basis transformations with variational quantum circuits (VQCs). PolyQROM employs optimized polynomial-based quantum operations to compress flow field data into low-dimensional representations while preserving essential features, enabling efficient quantum or classical post-processing for tasks like reconstruction and classification. By leveraging the mathematical properties of orthogonal polynomials, the framework enhances circuit expressivity and stabilizes training compared to conventional hardware-efficient VQCs. Numerical experiments demonstrate PolyQROM's effectiveness in reconstructing flow fields with high fidelity and classifying flow patterns with accuracy surpassing classical methods and quantum benchmarks, all while reducing computational complexity and parameter counts. The work bridges quantum simulation outputs with practical fluid analysis, addressing the ``output problem'' through efficient reduced-order modeling tailored for quantum-encoded flow data, offering a scalable pathway to exploit quantum advantages in computational fluid dynamics.

quant-ph↗

Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation

Large real-world robot datasets hold great potential to train generalist robot models, but scaling real-world human data collection is time-consuming and resource-intensive. Simulation has great potential in supplementing large-scale data, especially with recent advances in generative AI and automated data generation tools that enable scalable creation of robot behavior datasets. However, training a policy solely in simulation and transferring it to the real world often demands substantial human effort to bridge the reality gap. A compelling alternative is to co-train the policy on a mixture of simulation and real-world datasets. Preliminary studies have recently shown this strategy to substantially improve the performance of a policy over one trained on a limited amount of real-world data. Nonetheless, the community lacks a systematic understanding of sim-and-real co-training and what it takes to reap the benefits of simulation data for real-robot learning. This work presents a simple yet effective recipe for utilizing simulation data to solve vision-based robotic manipulation tasks. We derive this recipe from comprehensive experiments that validate the co-training strategy on various simulation and real-world datasets. Using two domains--a robot arm and a humanoid--across diverse tasks, we demonstrate that simulation data can enhance real-world task performance by an average of 38%, even with notable differences between the simulation and real-world data. Videos and additional results can be found at https://co-training.github.io/

cs.RO↗

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for enabling the robots to reason about novel situations, robustly handle real-world variability, and rapidly learn new tasks. To this end, we introduce GR00T N1, an open foundation model for humanoid robots. GR00T N1 is a Vision-Language-Action (VLA) model with a dual-system architecture. The vision-language module (System 2) interprets the environment through vision and language instructions. The subsequent diffusion transformer module (System 1) generates fluid motor actions in real time. Both modules are tightly coupled and jointly trained end-to-end. We train GR00T N1 with a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets. We show that our generalist robot model GR00T N1 outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments. Furthermore, we deploy our model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks, achieving strong performance with high data efficiency.

cs.RO↗

ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis

Vision-language-action (VLA) models present a promising paradigm by training policies directly on real robot datasets like Open X-Embodiment. However, the high cost of real-world data collection hinders further data scaling, thereby restricting the generalizability of VLAs. In this paper, we introduce ReBot, a novel real-to-sim-to-real approach for scaling real robot datasets and adapting VLA models to target domains, which is the last-mile deployment challenge in robot manipulation. Specifically, ReBot replays real-world robot trajectories in simulation to diversify manipulated objects (real-to-sim), and integrates the simulated movements with inpainted real-world background to synthesize physically realistic and temporally consistent robot videos (sim-to-real). Our approach has several advantages: 1) it enjoys the benefit of real data to minimize the sim-to-real gap; 2) it leverages the scalability of simulation; and 3) it can generalize a pretrained VLA to a target domain with fully automated data pipelines. Extensive experiments in both simulation and real-world environments show that ReBot significantly enhances the performance and robustness of VLAs. For example, in SimplerEnv with the WidowX robot, ReBot improved the in-domain performance of Octo by 7.2% and OpenVLA by 21.8%, and out-of-domain generalization by 19.9% and 9.4%, respectively. For real-world evaluation with a Franka robot, ReBot increased the success rates of Octo by 17% and OpenVLA by 20%. More information can be found at: https://yuffish.github.io/rebot/

cs.CV↗

DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility

Video-to-speech (V2S) synthesis, the task of generating speech directly from silent video input, is inherently more challenging than other speech synthesis tasks due to the need to accurately reconstruct both speech content and speaker characteristics from visual cues alone. Recently, audio-visual pre-training has eliminated the need for additional acoustic hints in V2S, which previous methods often relied on to ensure training convergence. However, even with pre-training, existing methods continue to face challenges in achieving a balance between acoustic intelligibility and the preservation of speaker-specific characteristics. We analyzed this limitation and were motivated to introduce DiVISe (Direct Visual-Input Speech Synthesis), an end-to-end V2S model that predicts Mel-spectrograms directly from video frames alone. Despite not taking any acoustic hints, DiVISe effectively preserves speaker characteristics in the generated audio, and achieves superior performance on both objective and subjective metrics across the LRS2 and LRS3 datasets. Our results demonstrate that DiVISe not only outperforms existing V2S models in acoustic intelligibility but also scales more effectively with increased data and model parameters. Code and weights can be found at https://github.com/PussyCat0700/DiVISe.

cs.SD↗

DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription services, and communication tools. The Audio-Visual Speech Recognition (AVSR) model enhances traditional speech recognition, particularly in noisy environments, by incorporating visual modalities like lip movements and facial expressions. While traditional AVSR models trained on large-scale datasets with numerous parameters can achieve remarkable accuracy, often surpassing human performance, they also come with high training costs and deployment challenges. To address these issues, we introduce an efficient AVSR model that reduces the number of parameters through the integration of a Dual Conformer Interaction Module (DCIM). In addition, we propose a pre-training method that further optimizes model performance by selectively updating parameters, leading to significant improvements in efficiency. Unlike conventional models that require the system to independently learn the hierarchical relationship between audio and visual modalities, our approach incorporates this distinction directly into the model architecture. This design enhances both efficiency and performance, resulting in a more practical and effective solution for AVSR tasks.

eess.AS↗

Geometry-Aware Attenuation Learning for Sparse-View CBCT Reconstruction

Cone Beam Computed Tomography (CBCT) plays a vital role in clinical imaging. Traditional methods typically require hundreds of 2D X-ray projections to reconstruct a high-quality 3D CBCT image, leading to considerable radiation exposure. This has led to a growing interest in sparse-view CBCT reconstruction to reduce radiation doses. While recent advances, including deep learning and neural rendering algorithms, have made strides in this area, these methods either produce unsatisfactory results or suffer from time inefficiency of individual optimization. In this paper, we introduce a novel geometry-aware encoder-decoder framework to solve this problem. Our framework starts by encoding multi-view 2D features from various 2D X-ray projections with a 2D CNN encoder. Leveraging the geometry of CBCT scanning, it then back-projects the multi-view 2D features into the 3D space to formulate a comprehensive volumetric feature map, followed by a 3D CNN decoder to recover 3D CBCT image. Importantly, our approach respects the geometric relationship between 3D CBCT image and its 2D X-ray projections during feature back projection stage, and enjoys the prior knowledge learned from the data population. This ensures its adaptability in dealing with extremly sparse view inputs without individual training, such as scenarios with only 5 or 10 X-ray projections. Extensive evaluations on two simulated datasets and one real-world dataset demonstrate exceptional reconstruction quality and time efficiency of our method.

eess.IV↗

Variations on Bollobás systems of $d$-partitions

This paper investigates five kinds of systems of $d$-partitions of $[n]$, including symmetric Bollobás systems, strong Bollobás systems, Bollobás systems, skew Bollobás systems, and weak Bollobás systems. Many known results on variations of Bollobás systems are unified. Especially we give a negative answer to a conjecture on Bollobás systems of $d$-partitions of $[n]$ that was presented by Hegedüs and Frankl [European J. Comb., 120 (2024), 103983]. Even though this conjecture does not hold for general Bollobás systems, we show that it holds for strong Bollobás systems of $d$-partitions of $[n]$.

math.CO↗

Dual-grating single-shot pump-probe technique

A simple and effective single-shot pump-probe technique is reported for studying the ultrafast dynamic processes in various materials. Using only two commercial gratings, a large time window of ~ 95.58 ps is spatially encoded in a single probe pulse, and single-shot time-resolved measurements are implemented. This time window exceeds the maximum reported values for single-shot pump-probe techniques using the echelon or angle beam encoding strategy. The phase difference problem in the echelon encoding strategies is also eliminated and a customized echelon is not needed in this technique. The ultrafast dynamic processes of ZnSe and indolium squaraine at a wavelength of 650 nm were investigated using this technique.

physics.optics↗

Ultrafast Electron Diffraction with MeV Electron Source from a Laser Wakefield Accelerator

MeV ultrafast electron diffraction (UED) is a widely used technology for ultrafast structural dynamic studies of matters in numerous areas. The development of laser wakefield accelerator (LWFA) envisions great potential of advanced all-optical electron source based on LWFA in UED applications. We experimentally demonstrated that an LWFA-based device with a miniaturized permanent magnet beamline can generate and manipulate electron beams suitable for UED. In the beam transmission, the LWFA electron beams with intrinsic short duration stretch due to energy spread and then are compressed by a following double bend achromat. The optimized double bend achromat can make the beamline isochronous such that the arrival time jitter induced by the shot-to-shot energy fluctuation can be eliminated, and allow the advantage of the natural laser-beam synchronization for LWFAs to emerge. With the energy filtering, the beam energy spread can be reduced to 3% (FWHM) while a sufficient amount of charge (11.9 fC) per bunch for diffraction is retained. Start-to-end simulations showed that the bunch length reaches ~30 fs (rms) with the same experimental configuration. Clear single-shot and multi-shot diffraction patterns of single-crystalline gold samples are obtained and the derived lattice constant agrees excellently with the real value. Our proof-of-principle experiments open the door to the detection of ultrafast structural dynamics using MeV LWFA beams, and pave the way for the UED applications with sub-10-fs temporal resolution.

physics.acc-ph↗

Single-shot pump-probe technique by the combination of an echelon and a grating with a time window of 109 ps

In this study, using only a single pulse, pump-probe measurement with a large time window of more than 100 ps is implemented. A commercial grating is used to encode a time window of ~ 56 ps in a single pulse; therefore, there is no need for machining customization. In addition, in this technique, the grating surface is accurately imaged, eliminating the image blur problem caused by phase differences in previous echelon-based techniques. Moreover, to make full use of the grating surface and obtain a larger time window, a simple reflection echelon is combined that matches the grating in the time window. This combination encoding strategy results in a total time window of ~ 109 ps and maintains accurate imaging of the grating surface. This time window is an order of magnitude greater than the maximum reported values of the echelon encoding strategy and the angle beam encoding strategy. To demonstrate this single-shot pump-probe technique, the two-photon absorption process of ZnSe and the excited-state absorption process of a symmetrical phenoxazinium bromine salt were studied. The possibility of further improving the experimental setup is also discussed.

physics.optics↗