SearcharxivSearch

arXiv subjects

Xiongwei Zhu

Publications and source records attributed to Xiongwei Zhu.

11 recordsLinked to original sources

On-Policy Self-Distillation in Diffusion Models

Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on-policy self-distillation framework that converts image-level reward guidance into explicit targets for clean-output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fits these targets as detached supervision through finite fitting before an exponential moving average update refreshes the behavior policy. This setup lets us measure target construction and finite realization separately. Controlled same-query experiments show that larger target-construction gains do not necessarily translate into larger realized gains after a single fitting update. Across SD 3.5-M and the step-distilled Z-Image-Turbo, our approach achieves the best final held-out scores in 19 of 20 reward-matched settings across two backbones and ten evaluators. It outperforms the strongest competing method by up to 44.0% and reduces training GPU-hours relative to DiffusionNFT by 40% on SD 3.5-M and 63% on Z-Image-Turbo. These results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.

cs.CV

DanceOPD: On-Policy Generative Field Distillation

Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For instance, editing tends to degrade T2I performance, while global and local editing interfere with each other. Consequently, effectively composing these capabilities has become a central challenge for image generation model training. To tackle this, we introduce DanceOPD, an on-policy generative field distillation framework for flow-matching models that routes each sample to one capability field, queries one low-noise student-induced state, and trains with a simple velocity MSE objective. With each capability source defined as a velocity field over the shared flow state space, the student learns from fields queried on its own rollout states to compose expert capabilities. This formulation also absorbs operator-defined fields such as classifier-free guidance. Comprehensive experiments on T2I, editing, realism-field absorption, and CFG absorption show that our approach improves multi-capability composition, strengthening target capabilities while preserving anchor generation quality. We believe this work establishes a practical route for generative field distillation in flow-matching models.

cs.CV

ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining high model capacity. However, in memory-constrained inference scenarios, only a small set of experts can be cached. Experts not in the cache must be fetched from slow external storage (e.g., UFS), leading to frequent evictions and substantial I/O overhead. We propose ReMoE, a router fine-tuning framework designed to boost token-wise expert reuse. ReMoE biases the router toward recently selected experts, producing temporally stable routing that better matches cache locality constraints. By increasing short-horizon expert reuse, ReMoE reduces expert fetches from storage without adding inference-time computation. Experiments on DeepSeek and Qwen models show that ReMoE improves expert reuse by 26% while maintaining downstream task performance. Real-system evaluations further confirm these benefits, improving output throughput by 8.4% under vLLM GPU-CPU expert offloading and reducing TPOT by 43.6-49.8% under llama.cpp on Jetson Orin NX, corresponding to a 1.77-1.99$\times$ decode speedup across diverse workloads. Checkpoints and usage instructions are available at https://github.com/BUAA-OSCAR/ReMoE.

cs.LG

Decouple Content and Motion for Conditional Image-to-Video Generation

The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text.The previous cI2V generation methods conventionally perform in RGB pixel space, with limitations in modeling motion consistency and visual continuity. Additionally, the efficiency of generating videos in pixel space is quite low.In this paper, we propose a novel approach to address these challenges by disentangling the target RGB pixels into two distinct components: spatial content and temporal motions. Specifically, we predict temporal motions which include motion vector and residual based on a 3D-UNet diffusion model. By explicitly modeling temporal motions and warping them to the starting image, we improve the temporal consistency of generated videos. This results in a reduction of spatial redundancy, emphasizing temporal details. Our proposed method achieves performance improvements by disentangling content and motion, all without introducing new structural complexities to the model. Extensive experiments on various datasets confirm our approach's superior performance over the majority of state-of-the-art methods in both effectiveness and efficiency.

cs.CV

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single $\textbf{Vi}$sion and $\textbf{S}$cene $\textbf{T}$ext $\textbf{A}$ggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least $\bf{8.4}\%$ at Recall@1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework.

cs.CV

Simulation Study of Laser Plasma Accelerator Via Vorpal

In this paper, we use PIC code Vorpal to do the extensive simulation about the laser plasma accelerator in the linear, quasilinear and nonlinear regime respectively. We design the ~100 MeV or so laser plasma accelerator ( LPA ) via Vorpal simulation. Finally, we discuss the application of the designed LPA in the compact light source field.

physics.acc-ph

The Roads to LPA Based Free Electron Laser

In this paper, we simply outline the present status of the free electron laser and the laser plasma based accelerator, and we simply discuss the potential possible roads appearing in the accelerator community to use the laser plasma based accelerator into the field of the free electron laser.

physics.acc-ph

Study on the production of the subpicosecond electron bunch

The production of the high brightness femtosecond electron bunch is now one of the hot research topics. This paper describes one electron linac facility used to produce the subpicosecond electron bunch. We analyze the main structure parameters, and study the beam dynamics of the facility. Finally, we discuss the application of this facility in Compton scattering.

physics.acc-ph

LWFA as a Preinjector for XFEL Driver Linac

In this paper, we propose to use the LWFA as the preinjector for XFEL driver linac. We can use LWFA to produce the femtosecond electron beam. The peak current of the produced beam can reach tens of $kA$ and is high enough to drive XFEL so that the present bunch compressing technique can be deserted. The output optical pusle length can be as short as $ 1 fs $.

physics.acc-ph

1D Longitudinal Beam Dynamics of Laser Plasma Wakefield Accelerator

In this paper, we get the 1D approximate analytical solution of the plasma electrostatic wake driven by the laser, and get the oscillating frequency of the wake modified by the nonlinear high order terms, and find that the frequency depends on the oscillating amplitudes of the wake and the laser which is the general nonlinear phenomenon. Finally we analyze the longitudinal beam dynamics in this electrostatic wake, and find that the high order terms don't change the topology of the longitudinal phase space.

physics.acc-ph

Scaling Law of Beam Break-Up for the Single Ultrashort Bunch in RF Linac

Femtosecond bunch is a hot topic in the present world accelerator research community. The high peak current bunch will lead to the beam breakup phenomenon. In this paper, the scaling law with current for the single ultrashort bunch beam breakup in the radio frequency linear accelerator is proposed and obtained. The scaling power factor is roughly estimated to be 1/6, 1/6 and 1/3 for the geometry wake, the surface roughness wake and the resistance wake respectively.

physics.acc-ph