SearcharxivSearch

arXiv subjects

Qinglin He

Publications and source records attributed to Qinglin He.

3 recordsLinked to original sources

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression Variational Autoencoder, Video-VAE, is designed for video generation tasks, achieving 16x16 spatial and 8x temporal compression ratios, while maintaining exceptional video reconstruction quality. User prompts are encoded using two bilingual text encoders to handle both English and Chinese. A DiT with 3D full attention is trained using Flow Matching and is employed to denoise input noise into latent frames. A video-based DPO approach, Video-DPO, is applied to reduce artifacts and improve the visual quality of the generated videos. We also detail our training strategies and share key observations and insights. Step-Video-T2V's performance is evaluated on a novel video generation benchmark, Step-Video-T2V-Eval, demonstrating its state-of-the-art text-to-video quality when compared with both open-source and commercial engines. Additionally, we discuss the limitations of current diffusion-based model paradigm and outline future directions for video foundation models. We make both Step-Video-T2V and Step-Video-T2V-Eval available at https://github.com/stepfun-ai/Step-Video-T2V. The online version can be accessed from https://yuewen.cn/videos as well. Our goal is to accelerate the innovation of video foundation models and empower video content creators.

cs.CV

Modality-Fair Preference Optimization for Trustworthy MLLM Alignment

Multimodal large language models (MLLMs) have achieved remarkable success across various tasks. However, separate training of visual and textual encoders often results in a misalignment of the modality. Such misalignment may lead models to generate content that is absent from the input image, a phenomenon referred to as hallucination. These inaccuracies severely undermine the trustworthiness of MLLMs in real-world applications. Despite attempts to optimize text preferences to mitigate this issue, our initial investigation indicates that the trustworthiness of MLLMs remains inadequate. Specifically, these models tend to provide preferred answers even when the input image is heavily distorted. Analysis of visual token attention also indicates that the model focuses primarily on the surrounding context rather than the key object referenced in the question. These findings highlight a misalignment between the modalities, where answers inadequately leverage input images. Motivated by our findings, we propose Modality-Fair Preference Optimization (MFPO), which comprises three components: the construction of a multimodal preference dataset in which dispreferred images differ from originals solely in key regions; an image reward loss function encouraging the model to generate answers better aligned with the input images; and an easy-to-hard iterative alignment strategy to stabilize joint modality training. Extensive experiments on three trustworthiness benchmarks demonstrate that MFPO significantly enhances the trustworthiness of MLLMs. In particular, it enables the 7B models to attain trustworthiness levels on par with, or even surpass, those of the 13B, 34B, and larger models.

cs.CV

Reinforcement of interfacial superconductivity in a Bi2Te3/Fe1+yTe heterostructure under hydrostatic pressure

We investigate the hydrostatic pressure dependence of interfacial superconductivity occurring at the atomically sharp interface between two non-superconducting materials: the topological insulator (TI) Bi2Te3 and the parent compound Fe1+yTe of the chalcogenide iron based superconductors. Under pressure, a significant increase in the superconducting transition temperature Tc is observed. We trace the pressure dependence of a superconducting twin gap structure by Andreev reflection point contact spectroscopy (PCARS), which shows that a large superconducting gap associated with the interfacial superconductivity increases along with Tc. A second smaller gap, which is attributed to proximity-induced superconductivity in the TI layer, increases first, but then reaches a maximum and appears to be gradually suppressed at higher pressure. We interpret our data in the context of a pressure-induced doping effect of the interface, in which charge is transferred from the TI layer to the interface and the interfacial superconductivity is enhanced. This demonstrates the important role of the TI in the interfacial superconductivity mechanism.

cond-mat.supr-con