SearcharxivSearch

arXiv subjects

Xiang Yuan

Publications and source records attributed to Xiang Yuan.

At least 19 recordsLinked to original sources

Giant bulk photovoltaic effect driven by interfacial symmetry breaking in MoS2/Ta2NiSe5 heterostructures

Van der Waals (vdW) heterostructures offer a versatile platform for engineering unconventional bulk photovoltaic (BPV) effect through interfacial symmetry breaking. However, the coexistence of multiple photophysical mechanisms, driven by structural complexity, spontaneous charge transfer, and strong interlayer coupling, often obscures the microscopic origin of the BPV response and hinders its rational optimization. Here, we demonstrate a pronounced BPV effect localized at the overlap region of a cross-bar MoS2/Ta2NiSe5 vdW heterostructure, where symmetry breaking induced by vertical stacking lifts the inversion center of MoS2. The orthogonal device geometry enables the independent probing of intralayer and interfacial photoresponse pathways, facilitating clear separation of competing mechanisms. Spontaneous interfacial charge transfer between MoS2 and Ta2NiSe5 further establishes a strong interlayer electronic coupling. By modulating the interlayer potential landscape through gate voltage and vertical electric fields, we achieve an optimized zero-bias photocurrent density of 247 A/cm2 and a BPV coefficient of 0.99 V-1. Supported by theoretical modelling, our results illustrate how minimalist device geometry can transform complex heterostructures into experimentally tractable platforms. This strategy paves the way for analyzing and optimizing interface-driven BPV effect, with implications for self-powered optoelectronics, broadband photodetection, and energy-harvesting nanodevices.

cond-mat.mes-hall

Giant-exchange-driven Vectorial Control of a Minimal Topological Magnet in Eu3In2As4

The interplay between magnetism and band topology provides a route to controlling quantum states of matter, yet its realization in materials is often constrained by weak exchange coupling and complex electronic structures. Here, a giant exchange coupling is identified in the newly predicted topological magnet Eu3In2As4, giving rise to magnetization-dependent band shifts of up to 300 meV. Together with its intrinsically soft magnetic response, this strong cou-pling enables systematic tuning of topological phases by both the magnitude and orientation of applied magnetic fields. The magneto-topological phase diagram is mapped out in which an antiferromagnetic topological insulator ground state evolves, under modest fields, into a pro-posed intermediate 2/3-ferrimagnetic phase, and further into fully polarized ferromagnetic states predicted to host either Weyl or nodal-ring semimetals. Notably, the Weyl phase corresponds to a minimal model hosting a single pair of Weyl nodes. Quantum oscillations, anomalous Hall transport and magneto-infrared spectroscopy consistently reveal exchange-driven band recon-struction across these transitions. Rotation of the magnetization theoretically provides an effi-cient means to tune the momentum-space positions and separations of the Weyl nodes. These results establish Eu3In2As4 as a model system for exploring how strong exchange coupling can be used to control topological band structures with minimal complexity.

cond-mat.mtrl-sci

Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting

The performance of Large Language Models (LLMs) is fundamentally influenced by the distributional composition of multi-domain pre-training data. While manual heuristics were prevalent in early models, they increasingly fail to capture the intricate synergies between domains as data complexity grows. To overcome the issue, a dominant approach seeks to fit a proxy function mapping between domain weights and their corresponding validation losses, and then find the optimal domain weights to minimize validation losses. These methods rely on strong structural assumptions, such as rank invariance or scaling laws, which are often violated, resulting in non-negligible estimation bias. A promising approach is to directly optimize the weighting scheme from data. However, it suffers from unstable optimization trajectory and prohibitive computational overhead, limiting its potential to search better domain weights configurations. This paper presents a Bayesian domain weighting method to infer the weights from a Dirichlet distribution via introducing Gamma prior information learned from observations. Experimental results demonstrate that proposed method could achieve stable and efficient domain weights learning, and identifies optimal mixtures while consuming substantially less data than search-based function-fitting methods, revitalizing optimization-based domain weighting for large-scale applications.

cs.LG

In-Situ Polarimetry in Collimated Magneto-Infrared Spectroscopy System

Magneto-infrared spectroscopy under strong magnetic fields provides a powerful probe of Landau quantization and field-induced collective excitations, yet its full potential has long been constrained by the lack of in-situ polarization control, because the highly divergent infrared beam propagating through narrow light tubes undergoes multiple wall reflections, leading to severe polarization degradation. Here we report a collimated magneto-infrared spectroscopy system that integrates continuous in-situ polarimetry. The system employs incident and exit collimation chambers forming a Kepler type optical architecture, which converts the large-aperture FTIR output into a low-divergence beam and strongly suppresses multi-reflection trajectories inside long gold-plated light tubes, thereby enhancing both optical throughput and polarization fidelity. A remotely controlled polarization module, consisting of an automated linear polarizer and a switchable Fresnel rhomb positioned entirely outside the high-field region, enables continuous in-situ tuning between linear, circular, and arbitrary elliptical polarization states without thermal cycling, manual realignment, or breaking vacuum. Interchangeable compact focusing modules further support Faraday and Voigt geometries in both transmission and reflection experiments within a 50 mm magnet bore, providing efficient beam focusing and signal collection while maintaining polarization fidelity. The setup achieves a minimum root-mean-square noise of 0.0033%, an average noise of 0.0082%, and a linear polarization extinction ratio up to 40:1. We demonstrate the capability through continuous in-situ linear polarimetry and broadband circular polarimetry in the magneto-infrared spectroscopy of various single crystals. This platform establishes a robust experimental framework for in-situ polarization-resolved magneto-infrared spectroscopy.

physics.ins-det

Giant and Broadband Circular Dichroism from Particle-Hole Symmetry Breaking in Weyl Semimetals

Circular dichroism originates from symmetry breaking of material structure, leading to differential absorption of left- and right-circularly polarized light. However, circular dichroism in most materials is inherently weak and spectrally narrow, especially in the mid-to-far infrared. Here, we uncover giant infrared circular dichroism in the magnetic-field-forced Weyl semimetal Mn(Bi,Sb)2Te4, driven by extreme particle-hole symmetry breaking. Helicity-resolved magneto-infrared spectroscopy reveals circular dichroism exceeding 3000 mdeg (~130 mdeg/nm) with above-degree response extending over the 6-13 {\mu}m spectral range. The optical resonances are enhanced by a strong band nesting effect intrinsic to the Landau levels of type-II Weyl dispersion. A symmetry-based kp model reproduces these magneto-infrared responses and demonstrates that magnetization-induced asymmetric spin-orbit coupling generates particle-hole symmetry breaking, suppressing spin-up, parity-even wavefunction components in the valence Landau band and thereby producing pronounced optical helicity selectivity. Our findings establish particle-hole symmetry breaking as an effective route toward helicity-resolved optical control in quantum materials.

cond-mat.mtrl-sci

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and produces 3D world representations. With text or single-view image inputs, the model performs world generation, synthesizing high-fidelity, navigable 3D Gaussian Splatting (3DGS) scenes. This is achieved through a four-stage method: a) Panorama Generation with HY-Pano 2.0, b) Trajectory Planning with WorldNav, c) World Expansion with WorldStereo 2.0, and d) World Composition with WorldMirror 2.0. Specifically, we introduce key innovations to enhance panorama fidelity, enable 3D scene understanding and planning, and upgrade WorldStereo, our keyframe-based view generation model with consistent memory. We also upgrade WorldMirror, a feed-forward model for universal 3D prediction, by refining model architecture and learning strategy, enabling world reconstruction from multi-view images or videos. Also, we introduce WorldLens, a high-performance 3DGS rendering platform featuring a flexible engine-agnostic architecture, automatic IBL lighting, efficient collision detection, and training-rendering co-design, enabling interactive exploration of 3D worlds with character support. Extensive experiments demonstrate that HY-World 2.0 achieves state-of-the-art performance on several benchmarks among open-source approaches, delivering results comparable to the closed-source model Marble. We release all model weights, code, and technical details to facilitate reproducibility and support further research on 3D world models.

cs.CV

MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning

Multimodal Object-Entity Relation Extraction (MORE) is a challenging task in information extraction research. It aims to identify relations between visual objects and textual entities, requiring complex multimodal understanding and cross-modal reasoning abilities. Existing methods, mainly classification-based or generation-based without reasoning, struggle to handle complex extraction scenarios in the MORE task and suffer from limited scalability and intermediate reasoning transparency. To address these challenges, we propose MORE-R1, a novel model that introduces explicit stepwise reasoning with Reinforcement Learning (RL) to enable Large Vision-Language Model (LVLM) to address the MORE task effectively. MORE-R1 integrates a two-stage training process, including an initial cold-start training stage with Supervised Fine-Tuning (SFT) and a subsequent RL stage for reasoning ability optimization. In the initial stage, we design an efficient way to automatically construct a high-quality SFT dataset containing fine-grained stepwise reasoning tailored to the MORE task, enabling the model to learn an effective reasoning paradigm. In the subsequent stage, we employ the Group Relative Policy Optimization (GRPO) RL algorithm with a Progressive Sample-Mixing Strategy to stabilize training and further enhance model's reasoning ability on hard samples. Comprehensive experiments on the MORE benchmark demonstrate that MORE-R1 achieves state-of-the-art performance with significant improvement over baselines.

cs.MM

Defect Engineering for Stabilizing Magnetic and Topological Properties in Mn(Bi1-xSbx)2Te4

MnBi2Te4 is a versatile platform for exploring diverse topological quantum states, yet its potential is hampered by intrinsic antisite defects. While Sb substitution has been employed to tune the Fermi level towards the charge neutral point, it exacerbates the formation of Mn-Sb antisite defects. Here, we address this challenge by combining first-principles calculations with strategic synthesis to systematically investigate and control antisite defects in Mn(Bi1-xSbx)2Te4. Our calculations reveal that increasing antisite defect density progressively destroys the field-forced magnetic Weyl state, eventually driving the system into a trivial magnetic insulator. Motivated by these findings, we develop an optimized chemical vapor transport method, yielding high-quality Mn(Bi1-xSbx)2Te4 crystals with significantly reduced antisite defect density. The emergence of strong Shubnikov-de Haas oscillations in the forced ferromagnetic state and a pronounced anomalous Hall effect near charge neutrality, with opposite signs for n- and p-type samples, confirms the type-II Weyl semimetal nature. These findings underscore the critical role of antisite defects in determining the magnetic and topological properties of Mn(Bi1-xSbx)2Te4 and establish defect engineering via optimized synthesis as a crucial strategy for realizing its exotic magnetic topological states.

cond-mat.mtrl-sci

Isotropic Dirac fermion and anomalous oscillator strength of zeroth Landau level transition

Dirac fermions, characterized by their linear dispersion and relativistic nature, have emerged as a prominent class of quasiparticles in condensed matter physics. While the Dirac equation, initially developed in the context of high-energy physics, provides a remarkable framework for describing the electronic properties of these materials, the inherent symmetry constraints of condensed matter often lead to deviations from the idealized paradigm. In particular, three-dimensional Dirac fermions in solids often exhibit anisotropic behavior, challenging the notion of perfect symmetry inherent in the Dirac equation. Here, we report the observation of isotropic massive Dirac fermions in LaAlSi through Landau level spectroscopy. The presence of three-dimensional massive Dirac fermions across the Fermi energy is demonstrated by quantized and semiclassical analyses of the magnetic field evolution of Landau level transitions. The isotropic topological nature, Fermi velocity, and Dirac mass are evidenced by the identical magneto-infrared response among the Faraday and three Voigt geometries. Furthermore, we observe an unusually large oscillator strength in the zeroth Landau level transition of the Dirac fermion, compared to transitions with higher indices. This phenomenon, supported by model calculations, can be attributed to the combined effects of the partial excitation of Dirac fermion and the resonant dielectric coupling with the Weyl plasma. Our work provides a strategy for realizing ideal quasiparticle excitations and their coupling effects in condensed matter systems, offering a platform for exploring relativistic physics.

cond-mat.mtrl-sci

A High-Flux and High-Efficiency Setup for Magneto-Infrared Spectroscopy

We report the design and implementation of a high-flux, high-efficiency magneto-infrared spectroscopy system optimized for broadband measurements in high magnetic fields. The setup integrates a Fourier transform infrared spectrometer, a 12 T cryogen-free superconducting magnet, precision-polished and gold-plated light tubes, custom-designed reflective focusing modules for Faraday and Voigt geometries, and an external multi-detector chamber with motorized selection. Optical throughput is maximized by reducing light tube loss from 65.5%/m to 22.0%/m via abrasive flow and mechanical polishing followed by gold electroplating, and by adopting a single-on-axis parabolic-mirror Faraday module that increases the effective numerical aperture from 0.14 to 0.36, enhancing collection efficiency by nearly an order of magnitude. An eight-position motorized sample stage and fully automated control over magnetic field, temperature, optical path, and detector choice enable high-throughput measurements without repeated warm-ups. The optimized configuration achieves a root-mean-square noise level of 0.0061% in a 2-minute integration for a 40% reflectivity sample, corresponding to a signal-to-noise ratio exceeding 16000. System capabilities are demonstrated by resolving weak replica bands in EuCd2As2 and faint Landau level transitions in LaAlSi.

cond-mat.mtrl-sci

Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction

Multimedia Event Extraction (MEE) has become an important task in information extraction research as news today increasingly prefers to contain multimedia content. Current MEE works mainly face two challenges: (1) Inadequate extraction framework modeling for handling complex and flexible multimedia event structure; (2) The absence of multimodal-aligned training data for effective knowledge transfer to MEE task. In this work, we propose a Stepwise Schema-Guided Prompting Framework (SSGPF) using Multimodal Large Language Model (MLLM) as backbone for adaptive structure capturing to solve MEE task. At the initial step of SSGPF, we design Event Type Schema Guided Prompting (ETSGP) for event detection, then we devise Argument Role Schema Guided Prompting (ARSGP) that contains multi-step prompts with text-bridged grounding technique for argument extraction. We construct a weakly-aligned multimodal event labeled dataset based on existing unimodal event annotations, then conduct parameter efficient instruction tuning with LoRA on LLaVA-v1.5-7B under SSGPF. Experiments on the M2E2 benchmark demonstrate that SSGPF significantly outperforms current SOTA baselines by 5.8 percent F1 on event detection and 8.4 percent F1 on argument extraction.

cs.MM

Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model

Recent advances in generative world models have enabled remarkable progress in creating open-ended game environments, evolving from static scene synthesis toward dynamic, interactive simulation. However, current approaches remain limited by rigid action schemas and high annotation costs, restricting their ability to model diverse in-game interactions and player-driven dynamics. To address these challenges, we introduce Hunyuan-GameCraft-2, a new paradigm of instruction-driven interaction for generative game world modeling. Instead of relying on fixed keyboard inputs, our model allows users to control game video contents through natural language prompts, keyboard, or mouse signals, enabling flexible and semantically rich interaction within generated worlds. We formally defined the concept of interactive video data and developed an automated process to transform large-scale, unstructured text-video pairs into causally aligned interactive datasets. Built upon a 14B image-to-video Mixture-of-Experts(MoE) foundation model, our model incorporates a text-driven interaction injection mechanism for fine-grained control over camera motion, character behavior, and environment dynamics. We introduce an interaction-focused benchmark, InterBench, to evaluate interaction performance comprehensively. Extensive experiments demonstrate that our model generates temporally coherent and causally grounded interactive game videos that faithfully respond to diverse and free-form user instructions such as "open the door", "draw a torch", or "trigger an explosion".

cs.CV

HunyuanVideo 1.5 Technical Report

We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coherence with only 8.3 billion parameters, enabling efficient inference on consumer-grade GPUs. This achievement is built upon several key components, including meticulous data curation, an advanced DiT architecture featuring selective and sliding tile attention (SSTA), enhanced bilingual understanding through glyph-aware text encoding, progressive pre-training and post-training, and an efficient video super-resolution network. Leveraging these designs, we developed a unified framework capable of high-quality text-to-video and image-to-video generation across multiple durations and resolutions. Extensive experiments demonstrate that this compact and proficient model establishes a new state-of-the-art among open-source video generation models. By releasing the code and model weights, we provide the community with a high-performance foundation that lowers the barrier to video creation and research, making advanced video generation accessible to a broader audience. All open-source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5.

cs.CV

Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation

The creation of high-quality 3D assets, a cornerstone of modern game development, has long been characterized by labor-intensive and specialized workflows. This paper presents Hunyuan3D Studio, an end-to-end AI-powered content creation platform designed to revolutionize the game production pipeline by automating and streamlining the generation of game-ready 3D assets. At its core, Hunyuan3D Studio integrates a suite of advanced neural modules (such as Part-level 3D Generation, Polygon Generation, Semantic UV, etc.) into a cohesive and user-friendly system. This unified framework allows for the rapid transformation of a single concept image or textual description into a fully-realized, production-quality 3D model complete with optimized geometry and high-fidelity PBR textures. We demonstrate that assets generated by Hunyuan3D Studio are not only visually compelling but also adhere to the stringent technical requirements of contemporary game engines, significantly reducing iteration time and lowering the barrier to entry for 3D content creation. By providing a seamless bridge from creative intent to technical asset, Hunyuan3D Studio represents a significant leap forward for AI-assisted workflows in game development and interactive media.

cs.CV

Feed Two Birds with One Scone: Exploiting Function-Space Regularization for Both OOD Robustness and ID Fine-Tuning Performance

Robust fine-tuning aims to achieve competitive in-distribution (ID) performance while maintaining the out-of-distribution (OOD) robustness of a pre-trained model when transferring it to a downstream task. To remedy this, most robust fine-tuning methods aim to preserve the pretrained weights, features, or logits. However, we find that these methods cannot always improve OOD robustness for different model architectures. This is due to the OOD robustness requiring the model function to produce stable prediction for input information of downstream tasks, while existing methods might serve as a poor proxy for the optimization in the function space. Based on this finding, we propose a novel regularization that constrains the distance of fine-tuning and pre-trained model in the function space with the simulated OOD samples, aiming to preserve the OOD robustness of the pre-trained model. Besides, to further enhance the OOD robustness capability of the fine-tuning model, we introduce an additional consistency regularization to promote stable predictions of perturbed samples. Extensive experiments demonstrate our approach could consistently improve both downstream task ID fine-tuning performance and OOD robustness across a variety of CLIP backbones, outperforming existing regularization-based robust fine-tuning methods.

cs.LG

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods that offer rich diversity but lack 3D consistency and rendering efficiency, and 3D-based methods that provide geometric consistency but struggle with limited training data and memory-inefficient representations. To address these limitations, we present HunyuanWorld 1.0, a novel framework that combines the best of both worlds for generating immersive, explorable, and interactive 3D scenes from text and image conditions. Our approach features three key advantages: 1) 360{\deg} immersive experiences via panoramic world proxies; 2) mesh export capabilities for seamless compatibility with existing computer graphics pipelines; 3) disentangled object representations for augmented interactivity. The core of our framework is a semantically layered 3D mesh representation that leverages panoramic images as 360{\deg} world proxies for semantic-aware world decomposition and reconstruction, enabling the generation of diverse 3D worlds. Extensive experiments demonstrate that our method achieves state-of-the-art performance in generating coherent, explorable, and interactive 3D worlds while enabling versatile applications in virtual reality, physical simulation, game development, and interactive content creation.

cs.CV

A Comparison of Relativistic Coupled Cluster and Equation of Motion Coupled Cluster Quadratic Response Theory

We present the implementation of relativistic coupled cluster quadratic response theory (QR-CC), following our development of relativistic equation of motion coupled cluster quadratic response theory (QR-EOMCC) [X. Yuan et al., J. Chem. Theory Comput. 2023, 19, 9248]. These codes, which can be used in combination with relativistic (2- and 4-component based) as well as non-relativistic Hamiltonians, are capable of treating both static and dynamic perturbations for electric and magnetic operators. We have employed this new implementation to revisit the calculation of static and frequency-dependent first hyperpolarizabilities of hydrogen halides (HX, X=F-Ts) and the Verdet constant of heavy noble gas atoms (Xe, Rn, Og) and of selected hydrogen halides (HF to HI), in order to investigate the differences and similarities of QR-CC and the more approximate QR-EOMCC. Furthermore, we have determined the relative importance of scalar relativistic effects and spin-orbit coupling to these properties, through a comparison of different Hamiltonians, and extended our calculations to superheavy element species (HTs for hyperpolarizabilities, Og for the Verdet constant). Our results show that as one moves towards the bottom of the periodic table, QR-EOMCC can yield rather different results (hyperpolarizabilities) or perform rather similarly (Verdet constant) to QR-CC. These results underscore the importance of further characterizing the performance of QR-EOMCC for heavy element systems.

physics.chem-ph

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

3D AI-generated content (AIGC) is a passionate field that has significantly accelerated the creation of 3D models in gaming, film, and design. Despite the development of several groundbreaking models that have revolutionized 3D generation, the field remains largely accessible only to researchers, developers, and designers due to the complexities involved in collecting, processing, and training 3D models. To address these challenges, we introduce Hunyuan3D 2.1 as a case study in this tutorial. This tutorial offers a comprehensive, step-by-step guide on processing 3D data, training a 3D generative model, and evaluating its performance using Hunyuan3D 2.1, an advanced system for producing high-resolution, textured 3D assets. The system comprises two core components: the Hunyuan3D-DiT for shape generation and the Hunyuan3D-Paint for texture synthesis. We will explore the entire workflow, including data preparation, model architecture, training strategies, evaluation metrics, and deployment. By the conclusion of this tutorial, you will have the knowledge to finetune or develop a robust 3D generative model suitable for applications in gaming, virtual reality, and industrial design.

cs.CV