SearcharxivSearch

arXiv subjects

Yuxuan Yuan

Publications and source records attributed to Yuxuan Yuan.

At least 19 recordsLinked to original sources

Xiaomi-GUI-0 Technical Report

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and abnormal-state distribution, and cannot faithfully characterize execution stability in real-world use, where account states, permission dialogs, payment authentication, and risk control continually reshape the state distribution and open a persistent gap between benchmark scores and real usability. To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop. At its core is a real-device-dominant hybrid infrastructure, where physical devices are the primary execution environment and sandboxes provide auxiliary support, so that data collection, training, rollout, and evaluation share an execution distribution close to real deployment. We construct multi-source training data spanning high-frequency head tasks, high-generalization data for long-tail intents, and capability-enhancement data for reflection and memory, and introduce an error-driven data flywheel that turns failure trajectories into corrected actions, reflective explanations, and recovery demonstrations. The model is trained through a progressive three-stage pipeline of supervised fine-tuning, step-level reinforcement learning, and agentic reinforcement learning. Evaluated on public benchmarks and our in-house RealMobile, Xiaomi-GUI-0 achieves 72.0% success on RealMobile and 78.9% on AndroidWorld, while substantially improving execution stability and abnormal-state recognition in real-world tasks.

cs.AI

The Pandora project. II: how non-thermal physics drives bursty star formation and temperate mass-loaded outflows in dwarf galaxies

Dwarf galaxies provide powerful laboratories for studying galaxy formation physics. Their early assembly, shallow gravitational potentials, and bursty, clustered star formation histories make them especially sensitive to the processes that regulate baryons through multi-phase outflows. Using high-resolution, cosmological zoom-in simulations of a dwarf galaxy from \textit{the Pandora suite}, we explore the impact of stellar radiation, magnetic fields, and cosmic ray feedback on star formation, outflows, and metal retention. We find that our purely hydrodynamical model without non-thermal physics - in which supernova feedback is boosted to reproduce realistic stellar mass assembly - drives violent, overly enriched outflows that suppress the metal content of the host galaxy. Including radiation reduces the clustering of star formation and weakens feedback. However, the additional incorporation of cosmic rays produces fast, mass-loaded, multi-phase outflows consisting of both ionized and neutral gas components, in better agreement with observations. These outflows, which entrain a denser, more temperate ISM, exhibit broad metallicity distributions while preserving metals within the galaxy. Furthermore, the star formation history becomes more bursty, in agreement with recent JWST findings. These results highlight the essential role of non-thermal physics in galaxy evolution and the need to incorporate it in future galaxy formation models.

astro-ph.GA

Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction

Human perceptual systems excel at inducing and recognizing objects across both known and novel categories, a capability far beyond current machine learning frameworks. While generalized category discovery (GCD) aims to bridge this gap, existing methods predominantly focus on optimizing objective functions. We present an orthogonal solution, inspired by the human cognitive process for novel object understanding: decomposing objects into visual primitives and establishing cross-knowledge comparisons. We propose ConGCD, which establishes primitive-oriented representations through high-level semantic reconstruction, binding intra-class shared attributes via deconstruction. Mirroring human preference diversity in visual processing, where distinct individuals leverage dominant or contextual cues, we implement dominant and contextual consensus units to capture class-discriminative patterns and inherent distributional invariants, respectively. A consensus scheduler dynamically optimizes activation pathways, with final predictions emerging through multiplex consensus integration. Extensive evaluations across coarse- and fine-grained benchmarks demonstrate ConGCD's effectiveness as a consensus-aware paradigm. Code is available at github.com/lytang63/ConGCD.

cs.CV

Extended red wings and the visibility of reionization-epoch Lyman-$α$ emitters

The visibility of the Lyman-$α$ (Ly$α$) emission from reionization-epoch galaxies depends sensitively on the extent of the intrinsic \lya emission redwards of 1215.67~Å. The prominent red peak resulting from resonant radiative transfer in the interstellar medium is often modelled as a single Gaussian. We use the \textsc{Azahar} simulation suite of a massive-reionization epoch galaxy to show that a significantly larger fraction of the \lya emission extends to $400$-$800$~km~s$^{-1}$, and thus significantly further to the red than predicted by a Gaussian line profile. A cycle of frequent galaxy mergers strongly modulates the \lya luminosity, the red peak velocity and its extended red wing emerging from the galaxy, which all also strongly vary with viewing angle. The \lya emission also depends sensitively on the implemented feedback, dust and star formation physics. Our simulations including cosmic rays reproduce the observed spectral properties of reionization epoch \lya emitters (LAEs) well if we assume that the \lya emission is affected by very little dust. The visibility of LAEs can be strongly underestimated if the extended red wings of the intrinsic \lya emission are not accounted for. We discuss implications for using the visibility of LAEs to constrain the evolution of the volume-averaged neutral fraction during reionization.

astro-ph.GA

OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad

Although foundation models (FMs) claim to be powerful, their generalization ability significantly decreases when faced with distribution shifts, weak supervision, or malicious attacks in the open world. On the other hand, most domain generalization or adversarial fine-tuning methods are task-related or model-specific, ignoring the universality in practical applications and the transferability between FMs. This paper delves into the problem of generalizing FMs to the out-of-domain data. We propose a novel framework, the Object-Concept-Relation Triad (OCRT), that enables FMs to extract sparse, high-level concepts and intricate relational structures from raw visual inputs. The key idea is to bind objects in visual scenes and a set of object-centric representations through unsupervised decoupling and iterative refinement. To be specific, we project the object-centric representations onto a semantic concept space that the model can readily interpret and estimate their importance to filter out irrelevant elements. Then, a concept-based graph, which has a flexible degree, is constructed to incorporate the set of concepts and their corresponding importance, enabling the extraction of high-order factors from informative concepts and facilitating relational reasoning among these concepts. Extensive experiments demonstrate that OCRT can substantially boost the generalizability and robustness of SAM and CLIP across multiple downstream tasks.

cs.CV

Ly$α$ emission as a sensitive probe of feedback-regulated LyC escape from dwarf galaxies

Ly$α$ emission is an exceptionally informative tracer of the life cycle of evolving galaxies and the escape of ionising photons. However, theoretical studies of Ly$α$ emission are often limited by insufficient numerical resolution, incomplete sets of physical models, and poor line-of-sight (LOS) statistics. To overcome such limitations, we utilize here the novel PANDORA suite of high-resolution dwarf galaxy simulations that include a comprehensive set of state-of-the-art physical models for ionizing radiation, magnetic fields, supernova feedback and cosmic rays. We post-process the simulations with the radiative transfer code \textsc{RASCAS} to generate synthetic observations and compare to observed properties of Ly$α$ emitters. Our simulated Ly$α$ haloes are more extended than the spatial region from which the intrinsic emission emanates and our spatially resolved maps of spectral parameters of the Ly$α$ emission are very sensitive to the underlying spatial distribution and kinematics of neutral hydrogen. Ly$α$ and LyC emission display strongly varying signatures along different LOS depending on how each LOS intersects low-density channels generated by stellar feedback. Comparing galaxies simulated with different physics, we find the Ly$α$ signatures to exhibit systematic offsets determined by the different levels of feedback strength and the clumpiness of the neutral gas. Despite this variance, and regardless of the different physics included in each model, we find universal correlations between Ly$α$ observables and LyC escape fraction, demonstrating a robust connection between Ly$α$ and LyC emission. Ly$α$ observations from a large sample of dwarf galaxies should thus give strong constraints on their stellar feedback-regulated LyC escape and confirm their important role for the reionization of the Universe.

astro-ph.GA

Increased Burstiness at High Redshift in Multi-Physics Models Combining Supernova Feedback, Radiative Transfer and Cosmic Rays

We study star formation variability, or burstiness, as a method to constrain and compare different galaxy formation models at high redshift using the Azahar simulation suite. The models range from magneto-hydrodynamics with a magneto-thermo-turbulent prescription for star formation (iMHD) to more sophisticated setups incorporating radiative transfer (RTiMHD) and cosmic ray physics (RTnsCRiMHD). Analysing a sample of galaxies at redshifts $z=4-10$, we find that the RTnsCRiMHD model exhibits more regular star formation periodicity compared to iMHD and RTiMHD, as revealed by the Lomb-Scargle periodogram. While the RTiMHD model captures a notable degree of stochasticity in star formation without cosmic rays, RTnsCRiMHD galaxies display even greater scatter in the burst intensity and in the scatter around the star-forming main sequence. To evaluate the burstiness in RTnsCRiMHD against observations, we generate a mock spectrum during a mini-quenching event at $z=7.5$. This spectrum aligns well with the low-mass quiescent galaxy JADES-GS-z7-01-QU observed at $z=7.3$, though some discrepancies attributed to stellar metallicity hint at a composite spectrum. Our findings highlight the importance of including complex physical processes like cosmic rays and radiative transfer in simulations to accurately capture the bursty nature of star formation in high-redshift galaxies. Future JWST observations, particularly regarding the scatter around the star-forming main sequence, have the potential to refine and guide the next generation of galaxy formation models.

astro-ph.GA

Bootstrap Segmentation Foundation Model under Distribution Shift via Object-Centric Learning

Foundation models have made incredible strides in achieving zero-shot or few-shot generalization, leveraging prompt engineering to mimic the problem-solving approach of human intelligence. However, when it comes to some foundation models like Segment Anything, there is still a challenge in performing well on out-of-distribution data, including camouflaged and medical images. Inconsistent prompting strategies during fine-tuning and testing further compound the issue, leading to decreased performance. Drawing inspiration from how human cognition processes new environments, we introduce SlotSAM, a method that reconstructs features from the encoder in a self-supervised manner to create object-centric representations. These representations are then integrated into the foundation model, bolstering its object-level perceptual capabilities while reducing the impact of distribution-related variables. The beauty of SlotSAM lies in its simplicity and adaptability to various tasks, making it a versatile solution that significantly enhances the generalization abilities of foundation models. Through limited parameter fine-tuning in a bootstrap manner, our approach paves the way for improved generalization in novel environments. The code is available at github.com/lytang63/SlotSAM.

cs.CV

SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM

While Large Language Models (LLMs) have achieved remarkable success in various fields, the efficiency of training and inference remains a major challenge. To address this issue, we propose SUBLLM, short for Subsampling-Upsampling-Bypass Large Language Model, an innovative architecture that extends the core decoder-only framework by incorporating subsampling, upsampling, and bypass modules. The subsampling modules are responsible for shortening the sequence, while the upsampling modules restore the sequence length, and the bypass modules enhance convergence. In comparison to LLaMA, the proposed SUBLLM exhibits significant enhancements in both training and inference speeds as well as memory usage, while maintaining competitive few-shot performance. During training, SUBLLM increases speeds by 26% and cuts memory by 10GB per GPU. In inference, it boosts speeds by up to 37% and reduces memory by 1GB per GPU. The training and inference speeds can be enhanced by 34% and 52% respectively when the context window is expanded to 8192. Our code is available at https://github.com/XiaoMi/subllm.

cs.CL

Mixstyle-Entropy: Domain Generalization with Causal Intervention and Perturbation

Despite the considerable advancements achieved by deep neural networks, their performance tends to degenerate when the test environment diverges from the training ones. Domain generalization (DG) solves this issue by learning representations independent of domain-related information, thus facilitating extrapolation to unseen environments. Existing approaches typically focus on formulating tailored training objectives to extract shared features from the source data. However, the disjointed training and testing procedures may compromise robustness, particularly in the face of unforeseen variations during deployment. In this paper, we propose a novel and holistic framework based on causality, named InPer, designed to enhance model generalization by incorporating causal intervention during training and causal perturbation during testing. Specifically, during the training phase, we employ entropy-based causal intervention (EnIn) to refine the selection of causal variables. To identify samples with anti-interference causal variables from the target domain, we propose a novel metric, homeostatic score, through causal perturbation (HoPer) to construct a prototype classifier in test time. Experimental results across multiple cross-domain tasks confirm the efficacy of InPer.

cs.LG

Deciphering Lyman-$α$ Emission Deep into the Epoch of Reionisation

During the epoch of reionisation the first galaxies were enshrouded in pristine neutral gas, with one of the brightest emission lines in star-forming galaxies, Lyman-$α$ (Ly$α$), expected to remain undetected until the Universe became ionised. Providing an explanation for the surprising detection of Ly$α$ in these early galaxies is a major challenge for extra-galactic studies. Recent JWST observations have reignited the debate on whether residence in an overdensity of galaxies is a it sufficient and necessary condition for Ly$α$ to escape. Here, we take unique advantage of both high-resolution and high-sensitivity images from the JWST instrument NIRCam to reveal that all galaxies in a sample of z>7 Ly$α$ emitters have close companions. We exploit novel on-the-fly radiative transfer magnetohydrodynamical simulations with cosmic ray feedback to show that galaxies with frequent mergers have very bursty star formation which drives episodes of high intrinsic Ly$α$ emission and facilitates the escape of Ly$α$ photons along channels cleared of neutral gas. We conclude that the rapid build up of stellar mass through mergers presents a compelling solution to the long-standing puzzle of the detection of Ly$α$ emission deep into the epoch of reionisation.

astro-ph.GA

Layer-wise Representation Fusion for Compositional Generalization

Existing neural models are demonstrated to struggle with compositional generalization (CG), i.e., the ability to systematically generalize to unseen compositions of seen components. A key reason for failure on CG is that the syntactic and semantic representations of sequences in both the uppermost layer of the encoder and decoder are entangled. However, previous work concentrates on separating the learning of syntax and semantics instead of exploring the reasons behind the representation entanglement (RE) problem to solve it. We explain why it exists by analyzing the representation evolving mechanism from the bottom to the top of the Transformer layers. We find that the ``shallow'' residual connections within each layer fail to fuse previous layers' information effectively, leading to information forgetting between layers and further the RE problems. Inspired by this, we propose LRF, a novel \textbf{L}ayer-wise \textbf{R}epresentation \textbf{F}usion framework for CG, which learns to fuse previous layers' information back into the encoding and decoding process effectively through introducing a \emph{fuse-attention module} at each encoder and decoder layer. LRF achieves promising results on two realistic benchmarks, empirically demonstrating the effectiveness of our proposal.

cs.CL

Singular Perturbation-based Large-Signal Order Reduction of Microgrids for Stability and Accuracy Synthesis with Control

With the increasing penetration of distributed energy resources (DERs), it is of vital importance to study the dynamic stability of microgrids (MGs) with external control inputs in the electromagnetic transient (EMT) time scale. This requires detailed models of the underlying control structure of MGs and results in a high-order nonlinear MG control system. Higher-level controller design and stability analysis of such high-order systems are usually intractable and computation-costly. To overcome these challenges, this paper proposes a large-signal order reduction (LSOR) method for MGs with considerations of external control inputs and the detailed dynamics of underlying control levels based on singular perturbation theory (SPT). Specially, we innovatively proposed and strictly proved a general stability and accuracy assessment theorem that allows us to analyze the dynamic stability of a full-order nonlinear system by only leveraging its corresponding reduced-order model (ROM) and boundary layer model (BLM). Moreover, this theorem also theoretically provides a set of conditions under which the developed ROM is accurate. Finally, by embedding such a theorem into the SPT, we propose a novel LSOR approach with guaranteed accuracy and stability analysis equivalence. The proposed LSOR method is generic and can be applied to arbitrary dynamic systems. Multiple case studies are conducted on MG systems to show the effectiveness of the proposed approach.

eess.SY

Learning Latent Interactions for Event classification via Graph Neural Networks and PMU Data

Phasor measurement units (PMUs) are being widely installed on power systems, providing a unique opportunity to enhance wide-area situational awareness. One essential application is the use of PMU data for real-time event identification. However, how to take full advantage of all PMU data in event identification is still an open problem. Thus, we propose a novel method that performs event identification by mining interaction graphs among different PMUs. The proposed interaction graph inference method follows an entirely data-driven manner without knowing the physical topology. Moreover, unlike previous works that treat interactive learning and event identification as two different stages, our method learns interactions jointly with the identification task, thereby improving the accuracy of graph learning and ensuring seamless integration between the two stages. Moreover, to capture multi-scale event patterns, a dilated inception-based method is investigated to perform feature extraction of PMU data. To test the proposed data-driven approach, a large real-world dataset from tens of PMU sources and the corresponding event logs have been utilized in this work. Numerical results validate that our method has higher classification accuracy compared to previous methods.

eess.SP

The observable properties of cool winds from galaxies, AGN, and star clusters -- II. 3D models for the multiphase wind of M82

Galactic winds are a crucial player in galaxy formation and evolution, but observations of them have proven extraordinarily difficult to interpret, leaving large uncertainties even in basic quantities such as mass outflow rates. Part of this uncertainty arises from the relatively simplistic models to which complex wind observations are often fit, which inevitably discard much of the available information. Here we present an analysis of the wind of the nearby dwarf starburst galaxy M82 using a semi-analytic model that is able to take advantage of the full three-dimensional information present in position-position-velocity data cubes measured in the Hi 21 cm line, the CO 2-1 line, and the Ha line. Our best-fitting model produces position-dependent spectra in good agreement with the observations, and shows that the total wind mass flux in the atomic and molecular phases is approx 10 M_sun yr-1 (corresponding to a mass loading factor of about 2 - 3), with less than a factor of two uncertainty; the mass flux in the warm ionised phase is more poorly constrained, and may be comparable to or smaller than this. At least over the few kpc off the plane for which we trace the outflow, it appears to be a wind escaping the galaxy, rather than a fountain that falls back. Our fits require that clouds of cool gas entrained into the wind expand only modestly, suggesting they are magnetically confined. Finally, we demonstrate that attempts to model the wind using simplifying assumptions such as instantaneous acceleration and a constant terminal wind speed can yield significantly erroneous results.

astro-ph.GA

Data-Driven Outage Restoration Time Prediction via Transfer Learning with Cluster Ensembles

This paper develops a data-driven approach to accurately predict the restoration time of outages under different scales and factors. To achieve the goal, the proposed method consists of three stages. First, given the unprecedented amount of data collected by utilities, a sparse dictionary-based ensemble spectral clustering (SDESC) method is proposed to decompose historical outage datasets, which enjoys good computational efficiency and scalability. Specifically, each outage sample is represented by a linear combination of a small number of selected dictionary samples using a density-based method. Then, the dictionary-based representation is utilized to perform the spectral analysis to group the data samples with similar features into the same subsets. In the second stage, a knowledge-transfer-added restoration time prediction model is trained for each subset by combining weather information and outage-related features. The transfer learning technology is introduced with the aim of dealing with the underestimation problem caused by data imbalance in different subsets, thus improving the model performance. Furthermore, to connect unseen outages with the learned outage subsets, a t-distributed stochastic neighbor embedding-based strategy is applied. The proposed method fully builds on and is also tested on a large real-world outage dataset from a utility provider with a time span of six consecutive years. The numerical results validate that our method has high prediction accuracy while showing good stability against real-world data limitations.

eess.SP

Synthetic Active Distribution System Generation via Unbalanced Graph Generative Adversarial Network

Real active distribution networks with associated smart meter (SM) data are critical for power researchers. However, it is practically difficult for researchers to obtain such comprehensive datasets from utilities due to privacy concerns. To bridge this gap, an implicit generative model with Wasserstein GAN objectives, namely unbalanced graph generative adversarial network (UG-GAN), is designed to generate synthetic three-phase unbalanced active distribution system connectivity. The basic idea is to learn the distribution of random walks both over a real-world system and across each phase of line segments, capturing the underlying local properties of an individual real-world distribution network and generating specific synthetic networks accordingly. Then, to create a comprehensive synthetic test case, a network correction and extension process is proposed to obtain time-series nodal demands and standard distribution grid components with realistic parameters, including distributed energy resources (DERs) and capacity banks. A Midwest distribution system with 1-year SM data has been utilized to validate the performance of our method. Case studies with several power applications demonstrate that synthetic active networks generated by the proposed framework can mimic almost all features of real-world networks while avoiding the disclosure of confidential information.

eess.SY

Mitigating Smart Meter Asynchrony Error Via Multi-Objective Low Rank Matrix Recovery

Smart meters (SMs) are being widely deployed by distribution utilities across the U.S. Despite their benefits in real-time monitoring. SMs suffer from certain data quality issues; specifically, unlike phasor measurement units (PMUs) that use GPS for data synchronization, SMs are not perfectly synchronized. The asynchrony error can degrade the monitoring accuracy in distribution networks. To address this challenge, we propose a principal component pursuit (PCP)-based data recovery strategy. Since asynchrony results in a loss of temporal correlation among SMs, the key idea in our solution is to leverage a PCP-based low rank matrix recovery technique to maximize the temporal correlation between multiple data streams obtained from SMs. Further, our approach has a novel multi-objective structure, which allows utilities to precisely refine and recover all SM-measured variables, including voltage and power measurements, while incorporating their inherent dependencies through power flow equations. We have performed numerical experiments using real SM data to demonstrate the effectiveness of the proposed strategy in mitigating the impact of SM asynchrony on distribution grid monitoring.

eess.SP