SearcharxivSearch

arXiv subjects

Boyang Zheng

Publications and source records attributed to Boyang Zheng.

16 recordsLinked to original sources

Improved Baselines with Representation Autoencoders

Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insights which simplify and improve RAE. First, we study a generalized formulation where the representation is defined as sum of the last k encoder layers rather than solely the final layer. This simple change greatly improves reconstruction without encoder finetuning or specialized data (e.g., text, faces). Second, we study the prevalent assumption that RAE (using pretrained representation as encoder) replaces representation alignment (REPA), which distills the same representation to intermediate layers instead. Through large-scale empirical analysis, we uncover a surprising finding: RAE and REPA exhibit complementary working mechanisms, allowing the same representation to be used as both encoder and target for intermediate diffusion layers. Finally, the original RAE struggles with classifier-free guidance (CFG) and requires training a second, weaker diffusion model for AutoGuidance (AG). We show that REPA itself can be viewed as x-prediction in RAE latent space. By simply re-parameterizing the output of the DiT model, it can provide guidance for "free". Overall, RAEv2 leads to more than 10x faster convergence over the original RAE, achieving a state-of-the-art gFID of 1.06 in just 80 epochs on ImageNet-256. On FDr6, RAEv2 achieves a state-of-the-art 2.17 at just 80 epochs compared to the previous best 3.26 (800 epochs) without any post-training. This motivates EPFID@k (epochs to reach unguided gFID < k) as a measure of training efficiency. RAEv2 attains an EPFID@2 of 35 epochs, versus 177 for the original RAE. We also validate our approach across diverse settings for text-to-image generation and navigation world models, showing consistent improvements. The code is available at https://raev2.github.io.

cs.CV

Benchmarking Visual State Tracking in Multimodal Video Understanding

Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacity for visual state tracking is fundamental to video understanding, yet remains underexplored in current evaluations of Multimodal Large Language Models (MLLMs). We introduce Visual STAte Tracking benchmark (VSTAT), a video-based benchmark designed to diagnose visual state tracking in MLLMs. VSTAT consists of 834 clips drawn from both synthetic and real-world videos, paired with 1,500 questions that cannot be answered from any single frame or short segment, requiring continuous perception and integration of events across the entire video stream. Despite their strong performance on existing video benchmarks, we find that state-of-the-art MLLMs perform far below humans and only modestly above answer-prior baselines. To analyze this gap, we compare MLLMs' thinking traces with the underlying video stream to understand why and when MLLMs fail on VSTAT. We find that MLLMs reason and track correctly in text, but fail at visually perceiving the events they need to track. Finally, our preliminary evaluation suggests that recent agentic approaches, including MLLM-based video agents and coding agents, do not readily resolve these failures, still falling short on VSTAT.

cs.CV

Impact of Cu-Mn ratio on Structure and Defects in Layered Multiferroic Cu1-xMn1+ySiTe3

Multiferroic materials exhibit the coexistence of magnetic and ferroelectric order, enabling control of magnetism through electric fields and vice versa. These properties make them attractive for spintronic and memory device applications. Recent studies on Cu1-xMn1+ySiTe3 (0.04 \leq x \leq 0.26; 0.03 \leq y \leq 0.15) have revealed strong magnetoelectric coupling, with variations in Mn-to-Cu concentration leading to variations in optical, electronic, and magnetic responses. Despite these findings, the influence of nanoscale structure and defects on the observed properties remains poorly understood. In this study, we investigate the structure and nanoscale defects in Cu-deficient Cu1-xMn1+ySiTe3 (Cu:Mn ratio <1, i.e., with 0.04 \leq x \leq 0.26 and 0.03 \leq y \leq 0.15) and Cu-rich Cu1+xMn1-ySiTe3 (Cu:Mn ratio >1, i.e., with 0.04 \leq x \leq 0.3 and 0.13 \leq y \leq 0.31) crystals using scanning/transmission electron microscopy and single-crystal X-ray diffraction. Cu-deficient crystals exhibit extensive stacking faults correlated with chemical inhomogeneity between Mn and Cu, along with variations in Te stacking. In contrast, Cu-rich crystals show fewer stacking faults but contain other local structural variations, such as needle-shaped precipitates and loop-like features. These distinct local structural features between Cu-rich and Cu-deficient crystals can be correlated to variations in their observed properties. Complementary density functional theory calculations confirm that the Cu-rich structure is more polar than the Cu-deficient structure. Overall, this study provides a comprehensive understanding of how subtle changes in chemistry influence the nanoscale structure, defect distribution, and functional properties in Cu1-xMn1+ySiTe3, offering guidance for designing multiferroic materials with tailored performance.

cond-mat.mtrl-sci

Phase-dependent electronic structure of two-dimensional Ag layers at the graphene/SiC interface

Intercalation at the graphene/SiC interface provides a controlled route to stabilize atomically thin layers with properties distinct from their bulk counterparts. In this platform, the structure and stability of the intercalated phase depend sensitively on the defect landscape of the starting substrate. For intercalated two-dimensional silver at the graphene/SiC interface, two phases have been observed: a phase epitaxial to the SiC lattice, Ag$_{(1)}$, readily obtained following the conventional intercalation method under ultra-high-vacuum conditions and extensively characterized, and a more densely packed phase, called Ag$_{(2)}$, which has remained largely unexplored. Here we report an in situ ultra-high-vacuum preparation method of the second phase intercalated at the graphene/SiC interface; this phase previously was prepared via high-pressure confinement heteroepitaxy. Low-energy electron diffraction shows that Ag$_{(2)}$ is rotated by 30 degree relative to the SiC lattice and forms supercells, in contrast to the $(1\times 1)$ epitaxial relation of Ag$_{(1)}$ with SiC. High-resolution angle-resolved photoemission spectroscopy reveals a more rich Ag$_{(2)}$ band dispersion compared to the Ag$_{(1)}$. In density functional theory calculations, by defining the unfolding entropy which, in a quantified way, finds that the band structure of Ag$_{(2)}$ is more suitable to be unfolded to the SiC primitive cell, and the resulting unfolded band dispersion is in great agreement with the experimental data. We further show that the different intercalated Ag phases tune the electronic properties of the overlying quasi-free-standing graphene layer differently: compared with Ag$_{(1)}$, Ag$_{(2)}$ yields an $\sim$1.75 times higher charge carrier density and modifies the charge-plasmon interaction of the graphene layer, indicating a change in effective screening at the interface.

cond-mat.mtrl-sci

Defect Control via Cu Enrichment Enhances Multifunctional Properties in the Polar Semiconductor Cu1+xMn1-ySiTe3

Polar materials have recently attracted significant interest due to their rich multifunctional properties. The chalcogenide polar semiconductor Cu1-xMn1+ySiTe3 (Cu-deficient) is an emerging multiferroic system in which electric polarization is coupled to magnetization. However, its macroscopic ferroelectric polarization is strongly suppressed due to the presence of a high density of stacking faults. In this work, we demonstrate that these crystal defects, likely originating from non-stoichiometry, can be substantially reduced by increasing the Cu content. Cu-enriched samples, Cu1+xMn1-ySiTe3, crystallize in a noncentrosymmetric monoclinic structure (space group Pm) as the Cu-deficient counterpart but show a nearly stacking-fault-free phase, which is attributed to the emergence of an interstitial site. Consequently, the Cu-enriched samples show a pronounced enhancement of the second-harmonic generation (SHG) response compared to Cu-deficient compositions. Magnetically, the Cu-enriched crystals retain long-range antiferromagnetic order with a Neel temperature of TN ~ 33 K without a glassy state but manifest a distinct spin-flop transition along the polar b-axis that is absent in the Cu-deficient compositions. Furthermore, the electronic ground state evolves from insulating to doped semiconducting behavior upon Cu enrichment. Together, these results establish this material system as a unique and versatile platform for elucidating the interplay among composition, crystal defects, and multifunctional properties, offering a route to design magnetic polar systems with tunable quantum functionalities.

cond-mat.mtrl-sci

Beyond Language Modeling: An Exploration of Multimodal Pretraining

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models remains opaque. We provide empirical clarity through controlled, from-scratch pretraining experiments, isolating the factors that govern multimodal pretraining without interference from language pretraining. We adopt the Transfusion framework, using next-token prediction for language and diffusion for vision, to train on diverse data including text, video, image-text pairs, and even action-conditioned video. Our experiments yield four key insights: (i) Representation Autoencoder (RAE) provides an optimal unified visual representation by excelling at both visual understanding and generation; (ii) visual and language data are complementary and yield synergy for downstream capabilities; (iii) unified multimodal pretraining leads naturally to world modeling, with capabilities emerging from general training; and (iv) Mixture-of-Experts (MoE) enables efficient and effective multimodal scaling while naturally inducing modality specialization. Through IsoFLOP analysis, we compute scaling laws for both modalities and uncover a scaling asymmetry: vision is significantly more data-hungry than language. We demonstrate that the MoE architecture harmonizes this scaling asymmetry by providing the high model capacity required by language while accommodating the data-intensive nature of vision, paving the way for truly unified multimodal models.

cs.CV

Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders

Representation Autoencoders (RAEs) have shown distinct advantages in diffusion modeling on ImageNet by training in high-dimensional semantic latent spaces. In this work, we investigate whether this framework can scale to large-scale, freeform text-to-image (T2I) generation. We first scale RAE decoders on the frozen representation encoder (SigLIP-2) beyond ImageNet by training on web, synthetic, and text-rendering data, finding that while scale improves general fidelity, targeted data composition is essential for specific domains like text. We then rigorously stress-test the RAE design choices originally proposed for ImageNet. Our analysis reveals that scaling simplifies the framework: while dimension-dependent noise scheduling remains critical, architectural complexities such as wide diffusion heads and noise-augmented decoding offer negligible benefits at scale Building on this simplified framework, we conduct a controlled comparison of RAE against the state-of-the-art FLUX VAE across diffusion transformer scales from 0.5B to 9.8B parameters. RAEs consistently outperform VAEs during pretraining across all model scales. Further, during finetuning on high-quality datasets, VAE-based models catastrophically overfit after 64 epochs, while RAE models remain stable through 256 epochs and achieve consistently better performance. Across all experiments, RAE-based diffusion models demonstrate faster convergence and better generation quality, establishing RAEs as a simpler and stronger foundation than VAEs for large-scale T2I generation. Additionally, because both visual understanding and generation can operate in a shared representation space, the multimodal model can directly reason over generated latents, opening new possibilities for unified models.

cs.CV

Defect-Mediated Phase Engineering of 2D Ag at the Graphene/SiC Interface

Atomically thin silver (Ag) films offer unique opportunities in plasmonic, quantum optics, and energy harvesting, yet conventional growth methods struggle to achieve structural control at the monolayer limit. Here, we demonstrate phase-selective synthesis of large-area, crystalline 2D Ag films via defect-engineered confinement heteroepitaxy (CHet) at the epitaxial graphene/silicon carbide (EG/SiC) interface. By tuning graphene growth and post-growth defect introduction, two distinct Ag phases are achieved with disparate properties: a nearly commensurate Ag(1) lattice stabilized by vacancy and line defects in epitaxial graphene, and a denser Ag(2) phase preferentially grown with sp3-rich zero-layer graphene. Structural and spectroscopic characterization confirm lattice registry with the SiC substrate, while theoretical calculations reveal a thermodynamic preference for Ag(2) but an easier nucleation for Ag(1). Both phases are found to be semiconducting, with the Ag(2) phase exhibiting slightly enhanced n-doping of graphene. Notably, nonlinear optical measurements reveal a three-order magnitude difference in second-order susceptibility between the two phases, demonstrating promise for phase-tunable 2D metals in reconfigurable optoelectronic and metamaterial platforms.

cond-mat.mtrl-sci

Targeted Attack Improves Protection against Unauthorized Diffusion Customization

Diffusion models build a new milestone for image generation yet raising public concerns, for they can be fine-tuned on unauthorized images for customization. Protection based on adversarial attacks rises to encounter this unauthorized diffusion customization, by adding protective watermarks to images and poisoning diffusion models. However, current protection, leveraging untargeted attacks, does not appear to be effective enough. In this paper, we propose a simple yet effective improvement for the protection against unauthorized diffusion customization by introducing targeted attacks. We show that by carefully selecting the target, targeted attacks significantly outperform untargeted attacks in poisoning diffusion models and degrading the customization image quality. Extensive experiments validate the superiority of our method on two mainstream customization methods of diffusion models, compared to existing protections. To explain the surprising success of targeted attacks, we delve into the mechanism of attack-based protections and propose a hypothesis based on our observation, which enhances the comprehension of attack-based protections. To the best of our knowledge, we are the first to both reveal the vulnerability of diffusion models to targeted attacks and leverage targeted attacks to enhance protection against unauthorized diffusion customization. Our code is available on GitHub: https://github.com/psyker-team/mist-v2.

cs.CV

Diffusion Transformers with Representation Autoencoders

Latent generative modeling, where a pretrained autoencoder maps pixels into a latent space for the diffusion process, has become the standard strategy for Diffusion Transformers (DiT); however, the autoencoder component has barely evolved. Most DiTs continue to rely on the original VAE encoder, which introduces several limitations: outdated backbones that compromise architectural simplicity, low-dimensional latent spaces that restrict information capacity, and weak representations that result from purely reconstruction-based training and ultimately limit generative quality. In this work, we explore replacing the VAE with pretrained representation encoders (e.g., DINO, SigLIP, MAE) paired with trained decoders, forming what we term Representation Autoencoders (RAEs). These models provide both high-quality reconstructions and semantically rich latent spaces, while allowing for a scalable transformer-based architecture. Since these latent spaces are typically high-dimensional, a key challenge is enabling diffusion transformers to operate effectively within them. We analyze the sources of this difficulty, propose theoretically motivated solutions, and validate them empirically. Our approach achieves faster convergence without auxiliary representation alignment losses. Using a DiT variant equipped with a lightweight, wide DDT head, we achieve strong image generation results on ImageNet: 1.51 FID at 256x256 (no guidance) and 1.13 at both 256x256 and 512x512 (with guidance). RAE offers clear advantages and should be the new default for diffusion transformer training.

cs.CV

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis

This paper does not describe a new method; instead, it provides a thorough exploration of an important yet understudied design space related to recent advances in text-to-image synthesis -- specifically, the deep fusion of large language models (LLMs) and diffusion transformers (DiTs) for multi-modal generation. Previous studies mainly focused on overall system performance rather than detailed comparisons with alternative methods, and key design details and training recipes were often left undisclosed. These gaps create uncertainty about the real potential of this approach. To fill these gaps, we conduct an empirical study on text-to-image generation, performing controlled comparisons with established baselines, analyzing important design choices, and providing a clear, reproducible recipe for training at scale. We hope this work offers meaningful data points and practical guidelines for future research in multi-modal generation.

cs.CV

Large-Area Intercalated 2D-Pb/Graphene Heterostructure as a Platform for Generating Spin-Orbit Torque

A scalable platform to synthesize ultrathin heavy metals may enable high efficiency charge-to-spin conversion for next-generation spintronics. Here we report the synthesis of air-stable, epitaxially registered monolayer Pb underneath bilayer graphene on SiC (0001) by confinement heteroepitaxy (CHet). Diffraction, spectroscopy, and microscopy reveal CHet-based Pb intercalation predominantly exhibits a mottled hexagonal superstructure due to an ordered network of Frenkel-Kontorova-like domain walls. The system's air stability enables ex-situ spin torque ferromagnetic resonance (ST-FMR) measurements that demonstrate charge-to-spin conversion in graphene/Pb/ferromagnet heterostructures with a 1.5x increase in the effective field ratio compared to control samples.

cond-mat.mtrl-sci

LM4LV: A Frozen Large Language Model for Low-level Vision Tasks

The success of large language models (LLMs) has fostered a new research trend of multi-modality large language models (MLLMs), which changes the paradigm of various fields in computer vision. Though MLLMs have shown promising results in numerous high-level vision and vision-language tasks such as VQA and text-to-image, no works have demonstrated how low-level vision tasks can benefit from MLLMs. We find that most current MLLMs are blind to low-level features due to their design of vision modules, thus are inherently incapable for solving low-level vision tasks. In this work, we purpose $\textbf{LM4LV}$, a framework that enables a FROZEN LLM to solve a range of low-level vision tasks without any multi-modal data or prior. This showcases the LLM's strong potential in low-level vision and bridges the gap between MLLMs and low-level vision tasks. We hope this work can inspire new perspectives on LLMs and deeper understanding of their mechanisms. Code is available at https://github.com/bytetriper/LM4LV.

cs.CV

Effects of Vanadium Doping on the Optical Response and Electronic Structure of WS$_{2}$ Monolayers

Two-dimensional dilute magnetic semiconductors has been recently reported in semiconducting transition metal dichalcogenides by the introduction of spin-polarized transition metal atoms as dopants. This is the case of vanadium-doped WS$_2$ and WSe$_2$ monolayers, which exhibits a ferromagnetic ordering even above room temperature. However, a broadband characterization of their electronic band structure and its dependence on vanadium concentration is still lacking. Therefore, here we perform power-dependent photoluminescence, resonant four-wave mixing, and differential reflectance spectroscopy to study the optical transitions close to the A exciton energy of vanadium-doped WS$_2$ monolayers with distinct concentrations. Instead of a single A exciton peak, vanadium-doped samples exhibit two photoluminescence peaks associated with transitions to occupied and unoccupied bands. Moreover, resonant Raman spectroscopy and resonant second-harmonic generation measurements revealed a blueshift in the B exciton but no energy change in the C exciton as vanadium is introduced in the monolayers. Density functional theory calculations showed that the band structure is sensitive to the Hubbard \(U\) correction for vanadium and several scenarios are proposed to explain the two photoluminescence peaks around the A exciton energy region. Our work provides the first broadband optical characterization of these two-dimensional dilute magnetic semiconductors, shedding light on the novel electronic features of WS$_{2}$ monolayers which are tunable by the vanadium concentration.

cond-mat.mes-hall

ZrTe2/CrTe2: an epitaxial van der Waals platform for spintronics

The rapid discovery of two-dimensional (2D) van der Waals (vdW) quantum materials has led to heterostructures that integrate diverse quantum functionalities such as topological phases, magnetism, and superconductivity. In this context, the epitaxial synthesis of vdW heterostructures with well-controlled interfaces is an attractive route towards wafer-scale platforms for systematically exploring fundamental properties and fashioning proof-of-concept devices. Here, we use molecular beam epitaxy to synthesize a vdW heterostructure that interfaces two material systems of contemporary interest: a 2D ferromagnet (1T-CrTe2) and a topological semimetal (ZrTe2). We find that one unit-cell (u.c.) thick 1T-CrTe2 grown epitaxially on ZrTe2 is a 2D ferromagnet with a clear anomalous Hall effect. In thicker samples (12 u.c. thick CrTe2), the anomalous Hall effect has characteristics that may arise from real-space Berry curvature. Finally, in ultrathin CrTe2 (3 u.c. thickness), we demonstrate current-driven magnetization switching in a full vdW topological semimetal/2D ferromagnet heterostructure device.

cond-mat.mtrl-sci

Monolayer Vanadium-doped Tungsten Disulfide: A Room-Temperature Dilute Magnetic Semiconductor

Dilute magnetic semiconductors, achieved through substitutional doping of spin-polarized transition metals into semiconducting systems, enable experimental modulation of spin dynamics in ways that hold great promise for novel magneto-electric or magneto-optical devices, especially for two-dimensional systems such as transition metal dichalcogenides that accentuate interactions and activate valley degrees of freedom. Practical applications of 2D magnetism will likely require room-temperature operation, air stability, and (for magnetic semiconductors) the ability to achieve optimal doping levels without dopant aggregation. Here we describe room-temperature ferromagnetic order obtained in semiconducting vanadium-doped tungsten disulfide monolayers produced by a reliable single-step film sulfidation method across an exceptionally wide range of vanadium concentrations, up to 12 at% with minimal dopant aggregation. These monolayers develop p-type transport as a function of vanadium incorporation and rapidly reach ambipolarity. Ferromagnetism peaks at an intermediate vanadium concentration of a few atomic percent and decreases for higher concentrations, which is consistent with quenching due to orbital hybridization at closer vanadium-vanadium spacings, as supported by transmission electron microscopy, magnetometry and first-principles calculations. Room-temperature two-dimensional dilute magnetic semiconductors provide a new component to expand the functional scope of van der Waals heterostructures and bring semiconducting magnetic 2D heterostructures them into the realm of practical application.

cond-mat.mtrl-sci