SearcharxivSearch

arXiv subjects

Yuanming Yang

Publications and source records attributed to Yuanming Yang.

6 recordsLinked to original sources

DiT-Reward: Generative Representations for Text-to-Image Reward Modeling

Can representations learned for image generation also support the evaluation of generated images? We study text-to-image reward prediction as a downstream task of generative representation learning. To this end, we introduce DiT-Reward, which converts a pretrained text-to-image Diffusion Transformer into a reward model by processing near-clean image latents and aggregating text-conditioned image representations across transformer layers. Under the same training data mixture as HPSv3, DiT-Reward outperforms HPSv3 on all four evaluated preference benchmarks, reaching 85.6% on HPDv2 and 77.6% on HPDv3. When the generative backbone is frozen, a lightweight learned head can still extract meaningful preference predictions from its representations. Probing across depth further reveals that downstream reward performance is strongest in the middle-to-late layers and benefits from combining representations across different stages. We also observe consistent positive scaling with generative backbone capacity. Finally, when used to optimize Stable Diffusion 3.5 Large with Flow-GRPO, DiT-Reward outperforms HPSv3 along the matched training trajectory, with particularly clear gains in realism. Direct latent scoring also achieves a 1.65x inference speedup over HPSv3 with comparable peak memory. These results show that pretrained generative DiTs provide transferable representations for reward modeling and policy optimization.

cs.LG

Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations

Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existing distillation methods often rely on multiple auxiliary networks, carefully designed training stages, or complex optimization pipelines. In this work, we revisit the recently proposed Drifting Model objective and show that a single drifting loss can be directly used to simplify one step distillation. A key observation is that the pretrained diffusion teacher itself already provides a strong representation space. Unlike the original Drifting Model, which relies on an additional pretrained feature extractor, we use intermediate hidden states of the pretrained teacher model as the feature representation. This removes the need for training or introducing an extra representation network while preserving a semantically meaningful feature geometry for drifting. Furthermore, we introduce a lightweight mode coverage loss to mitigate mode collapse during distillation and encourage the student generator to cover diverse teacher-supported regions. Extensive experiments on ImageNet and SDXL demonstrate that our method achieves efficient one step generation with competitive image quality and diversity, achieving FID scores of 1.58 on ImageNet-64$\times$64 and 18.4 on SDXL, while substantially simplifying the overall distillation framework.

cs.CV

Overview of EXL-50 Research Progress and Future Plan

XuanLong-50 (EXL-50) is the first medium-size spherical torus (ST) in China, with the toroidal field at major radius at 50 cm around 0.5T. CS-free and non-inductive current drive via electron cyclotron resonance heating (ECRH) was the main physics research issue for EXL-50. Discharges with plasma currents of 50 kA - 180 kA were routinely obtained in EXL-50, with the current flattop sustained for up to or beyond 2 s. The current drive effectiveness on EXL-50 was as high as 1 A/W for low-density discharges using 28GHz ECRH alone for heating power less than 200 kW. The plasma current reached Ip>80 kA for high-density (5*10e18m-2) discharges with 150 kW 28GHz ECRH. Higher performance discharge (Ip of about 120 kA and core density of about 1*10e19m-3) was achieved with 150 kW 50GHz ECRH. The plasma current in EXL-50 was mainly carried by the energetic electrons.Multi-fluid equilibrium model has been successfully applied to reconstruct the magnetic flux surface and the measured plasma parameters of the EXL-50 equilibrium. The physics mechanisms for the solenoid-free ECRH current drive and the energetic electrons has also been investigated. Preliminary experimental results show that 100 kW of lower hybrid current drive (LHCD) waves can drive 20 kA of plasma current. Several boron injection systems were installed and tested in EXL-50, including B2H6 gas puffing, boron powder injection, boron pellet injection. The research plan of EXL-50U, which is the upgrade machine of EXL-50, is also presented.

physics.plasm-ph

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement learning from human feedback offers promise for preference alignment, existing reward models for visual generation face limitations, including black-box scoring without interpretability and potentially resultant unexpected biases. We present VisionReward, a general framework for learning human visual preferences in both image and video generation. Specifically, we employ a hierarchical visual assessment framework to capture fine-grained human preferences, and leverages linear weighting to enable interpretable preference learning. Furthermore, we propose a multi-dimensional consistent strategy when using VisionReward as a reward model during preference optimization for visual generation. Experiments show that VisionReward can significantly outperform existing image and video reward models on both machine metrics and human evaluation. Notably, VisionReward surpasses VideoScore by 17.2% in preference prediction accuracy, and text-to-video models with VisionReward achieve a 31.6% higher pairwise win rate compared to the same models using VideoScore. All code and datasets are provided at https://github.com/THUDM/VisionReward.

cs.CV

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

We present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos aligned with text prompt, with a frame rate of 16 fps and resolution of 768 * 1360 pixels. Previous video generation models often had limited movement and short durations, and is difficult to generate videos with coherent narratives based on text. We propose several designs to address these issues. First, we propose a 3D Variational Autoencoder (VAE) to compress videos along both spatial and temporal dimensions, to improve both compression rate and video fidelity. Second, to improve the text-video alignment, we propose an expert transformer with the expert adaptive LayerNorm to facilitate the deep fusion between the two modalities. Third, by employing a progressive training and multi-resolution frame pack technique, CogVideoX is adept at producing coherent, long-duration, different shape videos characterized by significant motions. In addition, we develop an effective text-video data processing pipeline that includes various data preprocessing strategies and a video captioning method, greatly contributing to the generation quality and semantic alignment. Results show that CogVideoX demonstrates state-of-the-art performance across both multiple machine metrics and human evaluations. The model weight of both 3D Causal VAE, Video caption model and CogVideoX are publicly available at https://github.com/THUDM/CogVideo.

cs.CV

Solenoid-free current drive via ECRH in EXL-50 spherical torus plasmas

As a new spherical tokamak (ST) designed to simplify engineering requirements of a possible future fusion power source, the EXL-50 experiment features a low aspect ratio (A) vacuum vessel (VV), encircling a central post assembly containing the toroidal field coil conductors without a central solenoid. Multiple electron cyclotron resonance heating (ECRH) resonances are located within the VV to improve current drive effectiveness. Copious energetic electrons are produced and measured with hard X-ray detectors, carry the bulk of the plasma current ranging from 50kA to 150kA, which is maintained for more than 1s duration. It is observed that over one Ampere current can be maintained per Watt of ECRH power issued from the 28-GHz gyrotrons. The plasma current reaches Ip>80kA for high density (>5e18me-2) discharge with 150kW ECHR heating. An analysis was carried out combining reconstructed multi-fluid equilibrium, guiding-center orbits of energetic electrons, and resonant heating mechanisms. It is verified that in EXL-50 a broadly distributed current of energetic electrons creates smaller closed magnetic-flux surfaces of low aspect ratio that in turn confine the thermal plasma electrons and ions and participate in maintaining the equilibrium force-balance.

physics.plasm-ph