SearcharxivSearch

arXiv subjects

Xiangyu Zheng

Publications and source records attributed to Xiangyu Zheng.

16 recordsLinked to original sources

HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies

Generalist vision--language--action (VLA) policies are typically trained on heterogeneous mixtures of robot demonstrations spanning diverse embodiments, action spaces, and observation configurations. Modeling such heterogeneity with a shared dense action module can induce negative transfer, particularly when action spaces or visual observations differ across data sources. We address this issue with HiMoE-VLA, a VLA framework built around a Hierarchical Mixture-of-Experts (HiMoE) action module. HiMoE uses Action-Space MoE layers at the input/output boundaries to specialize computation for distinct action spaces, Heterogeneity-Balancing MoE layers in neighboring layers to provide balanced capacity for residual variation in observations, scenes, and embodiments, and dense Transformer blocks in the middle to integrate shared representations. Two auxiliary objectives further guide this hierarchy: a contrastive Action-Space Regularization objective for boundary specialization and a load-balancing objective for stable expert utilization. HiMoE-VLA reaches 3.98 on CALVIN, 98.0\% on LIBERO, and 75.0\% and 63.7\% average success on real xArm7 and ALOHA tasks; under controlled heterogeneous co-training, it turns the negative transfer observed in strong baselines into positive transfer. The code and models are publicly available at https://github.com/ZhiyingDu/HiMoE-VLA.

cs.RO

Domain-Wall Mediated Polarization Switching in Ferroelectric AlScN: Strain Relief and Field-Dependent Dynamics

While scandium-doped aluminum nitride (AlScN) exhibits robust ferroelectricity and excellent thermal stability, its utility is limited by an exceptionally high coercive field ($E_c$) for polarization switching. Unraveling the atomistic switching dynamics is therefore critical for tailoring $E_c$. Here, we combine density functional theory and machine-learning molecular dynamics to elucidate the polarization switching mechanisms in AlScN over various Sc concentrations and applied electric fields. We find that excessive lattice strain strictly prohibits collective polarization switching, but the pre-existing domain walls relieve strain and lead to a distinct switching dynamics -- dictating a field-dependent switching mechanism. At low electric fields, switching occurs via gradual domain-wall propagation consistent with the Kolmogorov-Avrami-Ishibashi model. In contrast, high fields stimulate additional nucleation, driving a rapid, homogeneous reversal process described by the simultaneous non-linear nucleation and growth model. These findings highlight the critical role of domain-wall dynamics and suggest domain engineering as a viable strategy to tailor coercive fields in AlScN and related ferroelectrics.

cond-mat.mtrl-sci

FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding

This paper presents FluxMem, a training-free framework for efficient streaming video understanding. FluxMem adaptively compresses redundant visual memory through a hierarchical, two-stage design: (1) a Temporal Adjacency Selection (TAS) module removes redundant visual tokens across adjacent frames, and (2) a Spatial Domain Consolidation (SDC) module further merges spatially repetitive regions within each frame into compact representations. To adapt effectively to dynamic scenes, we introduce a self-adaptive token compression mechanism in both TAS and SDC, which automatically determines the compression rate based on intrinsic scene statistics rather than manual tuning. Extensive experiments demonstrate that FluxMem achieves new state-of-the-art results on existing online video benchmarks, reaching 76.4 on StreamingBench and 67.2 on OVO-Bench under real-time settings, while reducing latency by 69.9% and peak GPU memory by 34.5% on OVO-Bench. Furthermore, it maintains strong offline performance, achieving 73.1 on MLVU while using 65% fewer visual tokens.

cs.CV

Observation of Topological Hall Effect in Synthetic Antiferromagnetic Skyrmion System

Synthetic antiferromagnetic (SAF) skyrmions have emerged as promising candidates for next-generation high-speed and highly integrated spintronic devices, owing to their exceptional properties such as high driving velocity, nanoscale dimensions, and the absence of the skyrmion Hall effect. In this work, we report the observation of the topological Hall effect in both compensated and non-compensated synthetic antiferromagnetic skyrmion systems based on [Pt/Co/Ru]2 bilayers. The antiferromagnetic skyrmions are demonstrated to be robust in these synthetic antiferromagnets under zero-field. Our first principal calculations and micromagnetic simulations demonstrate that the formation of the antiferromagnetic skyrmions are due to nonuniformity of RKKY coupling associated with the proximity effect induced magnetic moments in the Pt and Ru layers. The skyrmions in the Pt and Ru layers adjacent to the Co layers lead to the observed topological Hall effect. This work not only provides insight into the effect of the magnetic proximity effect and RKKY coupling to the SAF skyrmions, but also an effective detection method for the SAF skyrmion systems, thereby laying a foundation for the practical application of antiferromagnetic skyrmions in spintronic devices.

cond-mat.mes-hall

Single femtosecond laser pulse-driven ferromagnetic switching

Light pulses offer a faster, more energy-efficient, and direct route to magnetic bit writing, pointing toward a hybrid memory and computing paradigm based on photon transmission and spin retention. Yet progress remains hindered, as deterministic, single-pulse optical toggle switching has so far been achieved only with ferrimagnetic materials, which require too specific a rare-earth composition and temperature conditions for technological use. In mainstream ferromagnet--central to spintronic memory and storage--such bistable switching is considered fundamentally difficult, as laser-induced heating does not inherently break time-reversal symmetry. Here, we report coherent magnetization switching in ferromagnets, driven by thermal anisotropy torque with single laser pulses. The toggle switching behavior is robust over a broad range of pulse durations, from femtoseconds to picoseconds, a prerequisite for practical applications. Furthermore, the phenomenon exhibits reproducibility in CoFeB/MgO-based magnetic tunnel junctions with a high magnetoresistance exceeding 110%, as well as the scalability down to nanoscales with remarkable energy efficiency (17 fJ per 100-nm-sized bit). These results mark a notable step toward integrating opto-spintronics into next-generation memory and storage technologies.

cond-mat.mes-hall

Self-Optimizing Machine Learning Potential Assisted Automated Workflow for Highly Efficient Complex Systems Material Design

Machine learning interatomic potentials have revolutionized complex materials design by enabling rapid exploration of material configurational spaces via crystal structure prediction with ab initio accuracy. However, critical challenges persist in ensuring robust generalization to unknown structures and minimizing the requirement for substantial expert knowledge and time-consuming manual interventions. Here, we propose an automated crystal structure prediction framework built upon the attention-coupled neural networks potential to address these limitations. The generalizability of the potential is achieved by sampling regions across the local minima of the potential energy surface, where the self-evolving pipeline autonomously refines the potential iteratively while minimizing human intervention. The workflow is validated on Mg-Ca-H ternary and Be-P-N-O quaternary systems by exploring nearly 10 million configurations, demonstrating substantial speedup compared to first-principles calculations. These results underscore the effectiveness of our approach in accelerating the exploration and discovery of complex multi-component functional materials.

cond-mat.mtrl-sci

Single-shot optical precessional magnetization switching of Pt/Co/Pt ferromagnetic trilayers

Ultra-fast magnetization switching triggered by a single femtosecond laser pulse has gained significant attention over the last decade for its potential in low-power consumption, high-speed memory applications. However, this phenomenon has been primarily observed in Gd-based ferrimagnetic materials, which are unsuitable for storage due to their weak perpendicular magnetic anisotropy (PMA). In this work, we demonstrated that applying a single laser pulse and an in-plane magnetic field can facilitate magnetic switching in a Pt/Co/Pt ferromagnetic trilayers stack within a specific laser power window. To further understand this phenomenon, we introduce a Cu layer to accelerates the re-establishment time of the anisotropy field of Pt/Co/Pt trilayers, which leads to bullseye-patterned magnetic switching. We have mapped state diagrams for these phenomena, and through micromagnetic simulations, we have determined that these switchings are influenced by thermal anisotropy torque, which can be modulated through PMA. These findings indicate that single-shot optical precessional magnetization reversal is feasible in a broader range of materials, opening avenues for the development of optical-magnetic memory devices.

cond-mat.mtrl-sci

Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation

Recent mainstream unsupervised video object segmentation (UVOS) motion-appearance approaches use either the bi-encoder structure to separately encode motion and appearance features, or the uni-encoder structure for joint encoding. However, these methods fail to properly balance the motion-appearance relationship. Consequently, even with complex fusion modules for motion-appearance integration, the extracted suboptimal features degrade the models' overall performance. Moreover, the quality of optical flow varies across scenarios, making it insufficient to rely solely on optical flow to achieve high-quality segmentation results. To address these challenges, we propose the Saliency-Motion guided Trunk-Collateral Network (SMTC-Net), which better balances the motion-appearance relationship and incorporates model's intrinsic saliency information to enhance segmentation performance. Specifically, considering that optical flow maps are derived from RGB images, they share both commonalities and differences. Accordingly, we propose a novel Trunk-Collateral structure for motion-appearance UVOS. The shared trunk backbone captures the motion-appearance commonality, while the collateral branch learns the uniqueness of motion features. Furthermore, an Intrinsic Saliency guided Refinement Module (ISRM) is devised to efficiently leverage the model's intrinsic saliency information to refine high-level features, and provide pixel-level guidance for motion-appearance fusion, thereby enhancing performance without additional input. Experimental results show that SMTC-Net achieved state-of-the-art performance on three UVOS datasets ( 89.2% J&F on DAVIS-16, 76% J on YouTube-Objects, 86.4% J on FBMS ) and four standard video salient object detection (VSOD) benchmarks with the notable increase, demonstrating its effectiveness and superiority over previous methods.

cs.CV

SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledge encompasses over 200 specialized disciplines, far exceeding the scope of existing benchmarks. The capabilities of LLMs in many of these specialized fields-particularly in light industry, agriculture, and service-oriented disciplines-remain inadequately evaluated. To address this gap, we present SuperGPQA, a comprehensive benchmark that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines. Our benchmark employs a novel Human-LLM collaborative filtering mechanism to eliminate trivial or ambiguous questions through iterative refinement based on both LLM responses and expert feedback. Our experimental results reveal significant room for improvement in the performance of current state-of-the-art LLMs across diverse knowledge domains (e.g., the reasoning-focused model DeepSeek-R1 achieved the highest accuracy of 61.82% on SuperGPQA), highlighting the considerable gap between current model capabilities and artificial general intelligence. Additionally, we present comprehensive insights from our management of a large-scale annotation process, involving over 80 expert annotators and an interactive Human-LLM collaborative system, offering valuable methodological guidance for future research initiatives of comparable scope.

cs.CL

UTBoost: Gradient Boosted Decision Trees for Uplift Modeling

Uplift modeling comprises a collection of machine learning techniques designed for managers to predict the incremental impact of specific actions on customer outcomes. However, accurately estimating this incremental impact poses significant challenges due to the necessity of determining the difference between two mutually exclusive outcomes for each individual. In our study, we introduce two novel modifications to the established Gradient Boosting Decision Trees (GBDT) technique. These modifications sequentially learn the causal effect, addressing the counterfactual dilemma. Each modification innovates upon the existing technique in terms of the ensemble learning method and the learning objective, respectively. Experiments with large-scale datasets validate the effectiveness of our methods, consistently achieving substantial improvements over baseline models.

cs.LG

OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework

Contemporary Video Object Segmentation (VOS) approaches typically consist stages of feature extraction, matching, memory management, and multiple objects aggregation. Recent advanced models either employ a discrete modeling for these components in a sequential manner, or optimize a combined pipeline through substructure aggregation. However, these existing explicit staged approaches prevent the VOS framework from being optimized as a unified whole, leading to the limited capacity and suboptimal performance in tackling complex videos. In this paper, we propose OneVOS, a novel framework that unifies the core components of VOS with All-in-One Transformer. Specifically, to unify all aforementioned modules into a vision transformer, we model all the features of frames, masks and memory for multiple objects as transformer tokens, and integrally accomplish feature extraction, matching and memory management of multiple objects through the flexible attention mechanism. Furthermore, a Unidirectional Hybrid Attention is proposed through a double decoupling of the original attention operation, to rectify semantic errors and ambiguities of stored tokens in OneVOS framework. Finally, to alleviate the storage burden and expedite inference, we propose the Dynamic Token Selector, which unveils the working mechanism of OneVOS and naturally leads to a more efficient version of OneVOS. Extensive experiments demonstrate the superiority of OneVOS, achieving state-of-the-art performance across 7 datasets, particularly excelling in complex LVOS and MOSE datasets with 70.1% and 66.4% $J \& F$ scores, surpassing previous state-of-the-art methods by 4.2% and 7.0%, respectively. And our code will be available for reproducibility and further research.

cs.CV

Which Invariance Should We Transfer? A Causal Minimax Learning Approach

A major barrier to deploying current machine learning models lies in their non-reliability to dataset shifts. To resolve this problem, most existing studies attempted to transfer stable information to unseen environments. Particularly, independent causal mechanisms-based methods proposed to remove mutable causal mechanisms via the do-operator. Compared to previous methods, the obtained stable predictors are more effective in identifying stable information. However, a key question remains: which subset of this whole stable information should the model transfer, in order to achieve optimal generalization ability? To answer this question, we present a comprehensive minimax analysis from a causal perspective. Specifically, we first provide a graphical condition for the whole stable set to be optimal. When this condition fails, we surprisingly find with an example that this whole stable set, although can fully exploit stable information, is not the optimal one to transfer. To identify the optimal subset under this case, we propose to estimate the worst-case risk with a novel optimization scheme over the intervention functions on mutable causal mechanisms. We then propose an efficient algorithm to search for the subset with minimal worst-case risk, based on a newly defined equivalence relation between stable subsets. Compared to the exponential cost of exhaustively searching over all subsets, our searching strategy enjoys a polynomial complexity. The effectiveness and efficiency of our methods are demonstrated on synthetic data and the diagnosis of Alzheimer's disease.

stat.ML

A New Causal Decomposition Paradigm towards Health Equity

Causal decomposition has provided a powerful tool to analyze health disparity problems, by assessing the proportion of disparity caused by each mediator. However, most of these methods lack \emph{policy implications}, as they fail to account for all sources of disparities caused by the mediator. Besides, their estimations \emph{pre-specified} some covariates set (\emph{a.k.a}, admissible set) for the strong ignorability condition to hold, which can be problematic as some variables in this set may induce new spurious features. To resolve these issues, under the framework of the structural causal model, we propose to decompose the total effect into adjusted and unadjusted effects, with the former being able to include all types of disparity by adjusting each mediator's distribution from the disadvantaged group to the advantaged ones. Besides, equipped with maximal ancestral graph and context variables, we can automatically identify the admissible set, followed by an efficient algorithm for estimation. Theoretical correctness and the efficacy of our method are demonstrated on a synthetic dataset and a spine disease dataset.

stat.ME

Latent Causal Invariant Model

Current supervised learning can learn spurious correlation during the data-fitting process, imposing issues regarding interpretability, out-of-distribution (OOD) generalization, and robustness. To avoid spurious correlation, we propose a Latent Causal Invariance Model (LaCIM) which pursues causal prediction. Specifically, we introduce latent variables that are separated into (a) output-causative factors and (b) others that are spuriously correlated to the output via confounders, to model the underlying causal factors. We further assume the generating mechanisms from latent space to observed data to be causally invariant. We give the identifiable claim of such invariance, particularly the disentanglement of output-causative factors from others, as a theoretical guarantee for precise inference and avoiding spurious correlation. We propose a Variational-Bayesian-based method for estimation and to optimize over the latent space for prediction. The utility of our approach is verified by improved interpretability, prediction power on various OOD scenarios (including healthcare) and robustness on security.

cs.LG

Timescales and contribution of heating and helicity effect in helicity-dependent all-optical switching

The manipulation of the magnetic direction by using the ultrafast laser pulse is attractive for its great advantages in terms of speed and energy efficiency for information storage applications. However, the heating and helicity effects induced by circularly polarized laser excitation are entangled in the helicity-dependent all-optical switching (HD-AOS), which hinders the understanding of magnetization dynamics involved. Here, by applying a dual-pump laser excitation, first with a linearly polarized (LP) laser pulse followed by a circularly polarized (CP) laser pulse, we identify the timescales and contribution from heating and helicity effects in HD-AOS with a Pt/Co/Pt triple layer. When the sample is preheated by the LP laser pulses to a nearly fully demagnetized state, CP laser pulses with a much-reduced power switches the sample's magnetization. By varying the time delay between the two pump pulses, we show that the helicity effect, which gives rise to the deterministic helicity induced switching, onsets instantly upon laser excitation, and only exists for less than 0.2 ps close to the laser pulse duration of 0.15 ps. The results reveal that that the transient magnetization state upon which CP laser pulses impinge is the key factor for achieving HD-AOS, and importantly, the tunability between heating and helicity effects with the unique dual-pump laser excitation approach will enable HD-AOS in a wide range of magnetic material systems for the potential ultrafast spintronics applications.

cond-mat.mtrl-sci

Defect-Correlated Skyrmions and Controllable Generation in Perpendicularly Magnetized CoFeB Ultrathin Films

Skyrmions have attracted significant interest due to their topological spin structures and fascinating physical features. The skyrmion phase arises in materials with Dzyaloshinskii-Moriya (DM) interaction at interfaces or in volume of non-centrosymmetric materials. However, although skyrmions were generated experimentally, one critical intrinsic relationship between fabrication, microstructures, magnetization and the existence of skyrmions remains to be established. Here, two series of CoFeB ultrathin films with controlled atomic scale structures are employed to reveal this relationship. By inverting the growth order, the amount of defects can be artificially tuned, and skyrmions are shown to be preferentially formed at defect sites. The stable region and the density of the skyrmions can be efficiently controlled in the return magnetization loops by utilizing first-order reversal curves to reach various metastable states. These findings establish the general and intrinsic relationship from sample preparation to skyrmion generation, offering an universal method to control skyrmion density.

cond-mat.mes-hall