SearcharxivSearch

arXiv subjects

Xinyue Wang

Publications and source records attributed to Xinyue Wang.

At least 19 recordsLinked to original sources

Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization

GUI agents increasingly operate across websites, mobile apps, and desktop environments, yet the field still reports progress primarily through task success. We argue that practical deployment depends equally on efficiency: how much context, computation, action budget, and runtime overhead an agent consumes while succeeding. This survey studies efficient GUI agents through an end-to-end systems lens that preserves the current technical axes of observation efficiency, context and memory efficiency, action efficiency, and planner-side/system efficiency. For each subsection, we expand the seed literature through targeted search plus backward and forward citation chaining, then synthesize the dominant mechanisms, reported efficiency signals, and new overheads they introduce. Across the literature, recent progress converges on a small set of recurring ideas: selective reading instead of full-context ingestion, global-to-local visual allocation, recoverable memory rather than raw history replay, verification-aware control, and hybrid runtimes that can switch between GUI and non-GUI execution. We conclude by identifying the main open problems, including honest accounting of verifier cost, cross-benchmark comparability, and co-design of observation, memory, and execution layers under real latency and privacy constraints.

cs.CL

Solar Soft X-ray Coronal Dimming in a Failed Eruption Associated with Plasma Cooling

Coronal dimmings are observed as sudden and localized reductions in the extreme-ultraviolet and X-ray emission of the solar corona. Traditionally, significant dimmings of spectral lines formed at temperatures of 1-2 MK are regarded as indicators of coronal mass ejections (CMEs), reflecting the density depletion caused by plasma escaping into interplanetary space. In this Letter, we report a peculiar deep coronal dimming event predominantly observed in high-temperature spectral lines following an M8.8-class confined solar flare associated with a failed filament eruption. Sun-as-a-star measurements from the Geostationary Operational Environmental Satellite and the Extreme Ultraviolet Variability Experiment reveal intensity reductions exceeding 30% in soft X-ray (SXR) and measurable decreases in Fe XVIII (6.5 MK) and Fe XX (9.3 MK). Spatially resolved observations from the Atmospheric Imaging Assembly demonstrate that the dimming originates from the active region core, while the Solar Terrestrial Relations Observatory-A shows no evidence for CME-driven mass loss. Differential emission measure analysis reveals plasma at temperatures >5 MK cooling into lower temperatures, supporting plasma cooling as the dominant contributor to the hot-band dimming rather than CME-associated plasma escape. This event demonstrates that deep hot SXR dimmings can occur without substantial CME-driven mass loss, suggesting that alternative physical mechanisms may also account for unresolved stellar dimmings in addition to the commonly inferred CME signatures.

astro-ph.SR

Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric token-origin provenance, interventions, and realized cost; transparent training-free selectors isolate controlled operating points. On locked image-disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image-cluster difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retain sharply different positive-support coverage: 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, matched controls, interventions, detector tests, and external methods reveal model-specific quality-risk-traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory; full-validation TextVQA and DocVQA further show that favorable target-verification points do not imply task-general compression. Visual-token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.

cs.CV

Electromagnetically induced transparency lasing in distributed resonant feedback system

We demonstrate a loss-enabled mechanism that can realize single electromagnetically induced transparency (EIT) analogue in periodic waveguide-resonator structures. Unlike previously reported passband-based EIT analogues, which arise from compressed passbands in the intrinsic-loss-free limit, the proposed mechanism exploits resonator intrinsic loss to generate an isolated transparency mode within the original bandgap. The resulting EIT state remains immune to finite-size mode discretization, enabling robust single-frequency lasing through distributed resonant feedback while requiring a threshold modal gain lower than that of an equally long Fabry-Perot laser. Our findings extend classical EIT analogues from passive spectral phenomena to active laser operation.

physics.optics

Causally Debiased Latent Action Model for Embodied Action Conditioned World Models

Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation, and data augmentation. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect in the real world. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from unlabeled videos, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual factors such as backgrounds and untouched objects. In this work, we identify this action-irrelevant bias as a key obstacle to controllable ACWMs and introduce evaluation metrics to measure latent-action bias, action following, and robustness. We propose CD-LAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three efficient fine-tuning objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and calibrated non-collapsed latent action representations. Experiments on 2B and 14B ACWM backbones show that CD-LAM substantially improves latent-action controllability, downstream robot-action following, visual fidelity, and adaptation efficiency, requiring only 6k fine-tuning steps and more than 12$\times$ fewer robot-action adaptation updates than the baseline.

cs.CV

Simple modules over the superconformal algebra $\mathcal{S}^{\prime}(1,n)$

Let $n\geq 2$, and let $\mathcal{S}(1,n)$ be the Lie superalgebra of zero-superdivergence superderivations of $\mathbb{C}[t^{\pm1}]\otimes\Lambda(n)$. Its derived algebra $\mathcal{S}^\prime(1,n):=[\mathcal{S}(1,n),\mathcal{S}(1,n)]$ is well known as a superconformal algebra. In this paper, we first study Shen-Larsson modules over $\mathcal{S}^\prime(1,n)$. These modules, introduced by G. Shen and T. A. Larsson, are constructed from modules over the Weyl superalgebra $K_{1,n}$ and the special linear Lie superalgebra $\mathfrak{sl}(1,n)$. We establish necessary and sufficient conditions for the simplicity of Shen-Larsson modules and investigate their simple subquotients in the non-simple case. Then as an application, building on the classification of simple cuspidal $\mathcal{S}^\prime(1,n)$-modules by C. Mart\'inez, O. Mathieu and E. Zelmanov, we obtain an explicit construction of all simple cuspidal modules over $\mathcal{S}^\prime(1,n)$.

math.RT

Processing-Controlled Structural Uniformity and Oxide-Ion Conduction in Na0.52Bi0.47TiO3 Ceramics Probed by Eu3+ Photoluminescence

Sodium bismuth titanate (NBT) is a promising oxide-ion conductor,but its electrical conductivity is highly sensitive to small changes in A-site stoichiometry and processing history.This sensitivity can reduce sample-to-sample reproducibility.Here we examine how precursor mixing controls structural uniformity and ionic transport in Na0.52Bi0.47TiO3 ceramics.Dry grinding,wet grinding with ethanol,and ball milling were compared by X-ray diffraction,electron microscopy,energy-dispersive spectroscopy,Eu3+ photoluminescence excitation spectroscopy,and electrochemical impedance spectroscopy.All processed powders and ceramics form the perovskite NBT phase within the detection limit of XRD.However,the microstructure,surface A-site cation ratio,Eu3+ excitation spectra,and electrical response change strongly with the mixing route.Continuous monitoring of Eu3+ excitation spectra at different emission wavelengths reveals different distributions of local Eu3+ environments.Larger spectral-shape variations are consistent with lower structural uniformity and stronger local distortion.Dry-ground samples show higher bulk conductivity than wet-ground samples,whereas wet-ground samples show much lower grain-boundary resistance.At 600 \u2103,the dry-60 min sample reaches a bulk conductivity of 13.54 mS cm-1,while wet-30 min shows the highest grain-boundary conductivity of 13.72 mS cm-1.These results suggest a processing-driven trade-off between bulk defect generation and grain-boundary blocking.Based on this processing understanding,Ca was introduced at the A site in Na0.52Bi0.47-xCaxTiO3.The x=0.04 sample reaches 8.35 mS cm-1 at 500 \u2103 and 18.98 mS cm-1 at 600 \u2103.

cond-mat.mtrl-sci

SCAR: Self-Supervised Continuous Action Representation Learning

Despite the central role of action in embodied intelligence, learning transferable action representations from visual transitions remains a fundamental challenge, particularly when world models must generalize across embodiments under limited data. We argue that action is not merely an auxiliary conditioning signal, but a distinct representational factor that decouples the controllable change from embodiment-specific actuation. In this work, we propose SCAR, a joint inverse-forward dynamics framework for learning unified action representations across embodiments from visual transitions. Built on a pretrained generative backbone, SCAR uses an inverse dynamics model (IDM) to infer latent actions from latent observation pairs and a forward dynamics model (FDM) to predict future dynamics conditioned on them. To make the latent space transferable rather than a generic visual bottleneck, we regularize the latent action posterior toward a standard Gaussian prior to limit arbitrary visual encoding, and introduce adversarial invariance to suppress embodiment- and environment-specific nuisance factors. Experiments on the Procgen and Robotwin dataset show that the learned unified latent action representation serves as a stronger conditioning interface for world modeling than embodiment-specific raw actions, yielding improved cross-embodiment low-data adaptation and cross-task transfer. Taken together, these results suggest that action can be learned as a shared representation of controllable change across embodiments, providing an interface for more transferable and generalizable world models.

cs.RO

Imaging Exploration of Molecular Subtypes in Tongue Squamous Cell Carcinoma

Tongue squamous cell carcinoma (TSCC) is an aggressive malignancy with marked biological heterogeneity and variable clinical outcomes. Although molecular profiling has improved understanding of TSCC heterogeneity, its clinical use remains constrained by invasive tissue sampling and limited representation of whole-tumor spatial complexity. Meanwhile, most radiomics studies in TSCC have focused on downstream clinical endpoints, and whether imaging can non-invasively reflect intrinsic molecular subtypes remains unclear. In this study, an integrated transcriptomic-radiomics framework was used to investigate the relationship between preoperative imaging phenotypes and molecular subtypes in TSCC. Transcriptomic data from 60 TSCC cases in The Cancer Genome Atlas were analyzed using unsupervised consensus clustering, followed by differential expression and functional enrichment analyses. Matched preoperative imaging data from The Cancer Imaging Archive were manually annotated for primary tumor regions, and radiomic features were extracted using PyRadiomics; group differences were assessed with the U-test. Two stable molecular subtypes, C1 and C2, were identified. Their biological differences were mainly associated with squamous epithelial differentiation, inflammatory signaling, and lipid metabolism, with C2 showing greater enrichment of immune-related pathways. In addition, 10 radiomic features differed significantly between the two subtypes, mainly wavelet-derived texture features from gray-level size zone, dependence, co-occurrence, and run length matrices (P=0.00202-0.0162). These findings support the potential of radiomics as a non-invasive approach for characterizing molecular heterogeneity in TSCC and provide an initial radiogenomic framework for biologically informed preoperative assessment.

q-bio.GN

Resource Consumption Threats in Large Language Models

Given limited and costly computational infrastructure, resource efficiency is a key requirement for large language models (LLMs). Efficient LLMs increase service capacity for providers and reduce latency and API costs for users. Recent resource consumption threats induce excessive generation, degrading model efficiency and harming both service availability and economic sustainability. This survey presents a systematic review of threats to resource consumption in LLMs. We further establish a unified view of this emerging area by clarifying its scope and examining the problem along the full pipeline from threat induction to mechanism understanding and mitigation. Our goal is to clarify the problem landscape for this emerging area, thereby providing a clearer foundation for characterization and mitigation.

cs.CR

Surface ferrimagnetic order in RuO2 film

RuO2, widely proposed as a prototypical altermagnet, remains intensely debated with regard to its magnetic nature. Here, we demonstrate that RuO2 is non-magnetic in the bulk, but possesses a spontaneous surface ferrimagnetic order. Using spin- and angle-resolved photoemission spectroscopy, we directly detect a narrow surface state with identical spin polarizations at opposite momenta and at the Brillouin-zone center, incompatible with the spin texture of any altermagnetic order. First-principles calculations identify the non-magnetic bulk state and reveal that the detected magnetism is confined to the fully oxygen-terminated surface, where the charge transfer from Ru to O at surface triggers a ferrimagnetic alignment between adjacent Ru sublattices with antiparallel moments of +0.48 uB and -0.04 uB. Our findings provide a unified explanation reconciling debating reports on the magnetism of RuO2, establishing surface ferrimagnetism as the origin of the observed magnetic signals, and distinguishing it unambiguously from altermagnetism.

cond-mat.mtrl-sci

Learning to Explore with Parameter-Space Noise: A Deep Dive into Parameter-Space Noise for Reinforcement Learning with Verifiable Rewards

Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning, yet growing evidence indicates an exploration ceiling: it often reweights existing solution traces rather than discovering new strategies, limiting gains under large sampling budgets (e.g., pass-at-256). We address this limitation with PSN-RLVR, which perturbs policy parameters before rollout generation to induce temporally consistent, trajectory-level exploration that better preserves long-horizon chain-of-thought coherence than action-space noise. To mitigate the resulting sampling-update mismatch, we incorporate truncated importance sampling (TIS). To avoid expensive KL-based adaptive noise control, we propose a computationally efficient real-time adaptive noise scheduler driven by a lightweight surrogate that combines semantic diversity with normalized self-certainty. Instantiated on GRPO, a widely used RLVR method, PSN-GRPO consistently expands the effective reasoning capability boundary across multiple mathematical reasoning benchmarks and model families, yielding higher pass-at-k under large sampling budgets and outperforming prior exploration-oriented RLVR methods (e.g., Pass-at-k-style training) while remaining orthogonal and thus composable for additional gains.

cs.LG

Silent Inconsistency in Data-Parallel Full Fine-Tuning: Diagnosing Worker-Level Optimization Misalignment

Data-parallel (DP) training with synchronous all-reduce is a dominant paradigm for full-parameter fine-tuning of large language models (LLMs). While parameter synchronization guarantees numerical equivalence of model weights after each iteration, it does not necessarily imply alignment of worker-level optimization dynamics before gradient aggregation. This paper identifies and studies this latent mismatch, termed \emph{silent inconsistency}, where cross-worker divergence in losses and gradients can remain invisible under conventional aggregated monitoring signals. We propose a lightweight, model-agnostic diagnostic framework that quantifies worker-level consistency using training signals readily available in standard pipelines. Specifically, we introduce three complementary metrics: loss dispersion, gradient-norm dispersion, and gradient-direction consistency measured by inter-worker cosine similarity. The proposed metrics incur negligible overhead and require no modification to model architecture, synchronization mechanisms, or optimization algorithms. We validate the framework by fully fine-tuning the 1B-parameter \texttt{openPangu-Embedded-1B-V1.1} model on the \texttt{tatsu-lab/alpaca} dataset using an 8-NPU DP setup, under controlled perturbations of cross-rank stochasticity. Experimental results show that progressively desynchronized data shuffling and random seeds lead to substantial increases in loss/gradient dispersion and reduced directional alignment, despite smooth globally averaged loss curves. These findings demonstrate that the proposed indicators provide actionable visibility into hidden instability modes in large-scale DP fine-tuning, enabling more reliable diagnosis and configuration assessment.

cs.LG

From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents

The enhanced capabilities of LLM-based agents come with an emergency for model planning and tool-use abilities. Attributing to helpful-harmless trade-off from LLM alignment, agents typically also inherit the flaw of "over-refusal", which is a passive failure mode. However, the proactive planning and action capabilities of agents introduce another crucial danger on the other side of the trade-off. This phenomenon we term "Toxic Proactivity'': an active failure mode in which an agent, driven by the optimization for Machiavellian helpfulness, disregards ethical constraints to maximize utility. Unlike over-refusal, Toxic Proactivity manifests as the agent taking excessive or manipulative measures to ensure its "usefulness'' is maintained. Existing research pays little attention to identifying this behavior, as it often lacks the subtle context required for such strategies to unfold. To reveal this risk, we introduce a novel evaluation framework based on dilemma-driven interactions between dual models, enabling the simulation and analysis of agent behavior over multi-step behavioral trajectories. Through extensive experiments with mainstream LLMs, we demonstrate that Toxic Proactivity is a widespread behavioral phenomenon and reveal two major tendencies. We further present a systematic benchmark for evaluating Toxic Proactive behavior across contextual settings.

cs.CL

From Classification to Ranking: Enhancing LLM Reasoning Capabilities for MBTI Personality Detection

Personality detection aims to measure an individual's corresponding personality traits through their social media posts. The advancements in Large Language Models (LLMs) offer novel perspectives for personality detection tasks. Existing approaches enhance personality trait analysis by leveraging LLMs to extract semantic information from textual posts as prompts, followed by training classifiers for categorization. However, accurately classifying personality traits remains challenging due to the inherent complexity of human personality and subtle inter-trait distinctions. Moreover, prompt-based methods often exhibit excessive dependency on expert-crafted knowledge without autonomous pattern-learning capacity. To address these limitations, we view personality detection as a ranking task rather than a classification and propose a corresponding reinforcement learning training paradigm. First, we employ supervised fine-tuning (SFT) to establish personality trait ranking capabilities while enforcing standardized output formats, creating a robust initialization. Subsequently, we introduce Group Relative Policy Optimization (GRPO) with a specialized ranking-based reward function. Unlike verification tasks with definitive solutions, personality assessment involves subjective interpretations and blurred boundaries between trait categories. Our reward function explicitly addresses this challenge by training LLMs to learn optimal answer rankings. Comprehensive experiments have demonstrated that our method achieves state-of-the-art performance across multiple personality detection benchmarks.

cs.CL

Transformer Is Inherently a Causal Learner

We reveal that transformers trained in an autoregressive manner naturally encode time-delayed causal structures in their learned representations. When predicting future values in multivariate time series, the gradient sensitivities of transformer outputs with respect to past inputs directly recover the underlying causal graph, without any explicit causal objectives or structural constraints. We prove this connection theoretically under standard identifiability conditions and develop a practical extraction method using aggregated gradient attributions. On challenging cases such as nonlinear dynamics, long-term dependencies, and non-stationary systems, this approach greatly surpasses the performance of state-of-the-art discovery algorithms, especially as data heterogeneity increases, exhibiting scaling potential where causal accuracy improves with data volume and heterogeneity, a property traditional methods lack. This unifying view lays the groundwork for a future paradigm where causal discovery operates through the lens of foundation models, and foundation models gain interpretability and enhancement through the lens of causality.

cs.LG

ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering

Typical deep clustering methods, while achieving notable progress, can only provide one clustering result per dataset. This limitation arises from their assumption of a fixed underlying data distribution, which may fail to meet user needs and provide unsatisfactory clustering outcomes. Our work investigates how multi-modal large language models (MLLMs) can be leveraged to achieve user-driven clustering, emphasizing their adaptability to user-specified semantic requirements. However, directly using MLLM output for clustering has risks for producing unstructured and generic image descriptions instead of feature-specific and concrete ones. To address these issues, our method first discovers that MLLMs' hidden states of text tokens are strongly related to the corresponding features, and leverages these embeddings to perform clusterings from any user-defined criteria. We also employ a lightweight clustering head augmented with pseudo-label learning, significantly enhancing clustering accuracy. Extensive experiments demonstrate its competitive performance on diverse datasets and metrics.

cs.LG

Boost of critical current density near quantum critical points in FeSe-Based superconductors with two superconducting domes

Recent studies have identified two superconducting domes in FeSe-based superconductors. It was discovered that each dome is accompanied by a distinct nematic quantum critical point (QCP): one associated with a pure nematic QCP, and the other with a nematic QCP entangled with antiferromagnetism (AFM). In this study, we delve into the evolution of the critical current density ($J_{\rm{c}}$) with doping in FeSe${_{1-x}}$(Te/S)${_{x}}$ single crystals, focusing on the behavior within the two superconducting domes. Surprisingly, three maxima of $J_{\rm{c}}$ were found in the two superconducting domes, with two sharp peaks in $J_{\rm{c}}$ observed precisely at the endpoints of the nematic phases, at $x$(Te) $\sim$ 0.5 for Te-doped and $x$(S) $\sim$ 0.17 for S-doped FeSe. The mechanisms of vortex pinning and the influence of quantum critical fluctuations have been extensively explored, emphasizing the contribution of quantum critical fluctuations in modulating $J_{\rm{c}}$. Additionally, an increase in $J_{\rm{c}}$ was also noted near FeSe$_{0.1}$Te$_{0.9}$, where its origin has been explored and discussed. This finding provides crucial clues about the existence of an ordered phase endpoint beneath the superconducting dome, offering an initial basis for further investigation into the potential presence of a QCP beneath it.

cond-mat.supr-con