SearcharxivSearch

arXiv subjects

Zihan Yan

Publications and source records attributed to Zihan Yan.

At least 19 recordsLinked to original sources

DiG-bench: Discovery in Games

Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games). DiG-bench consists of a set of 70 independent games. Each game is encoded as a short string and has unique transformation rules that must be discovered through interaction and experimentation. The levels of the game present a series of challenges to test whether the rules have been discovered, where the win conditions for each level are also unknown. We provide games at seven tiers of difficulty for AI agents. The lowest tier is routinely solvable by multiple models, while the highest tier challenges the best models in agentic harnesses. All 70 games were solved by at least one human on first attempt. A subset of 21 games is released publicly, and the remainder is held private for secure evaluation.

cs.AI

ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically obtain an initial quantized model through heuristic rules or greedy optimization, and once quantization is completed the resulting integer assignments are usually treated as final. This observation motivates a complementary optimization stage within PTQ that keeps quantized weights improvable after an executable quantized model has been produced, while preserving the quantized format. We introduce ReQuant, a backpropagation-free fixed-grid refinement procedure for this stage. Agnostic to the PTQ initializer, ReQuant takes an existing quantized model as a feasible starting point and iteratively revisits its discrete weight assignments on the fixed quantization grid. Accepted updates strictly reduce the mean squared reconstruction error and remain on the original grid. In this way, ReQuant turns the initially fixed PTQ output into an iteratively optimizable discrete solution and serves as a plug-and-play post-processing stage for existing PTQ pipelines. Experiments across diverse model families, bit-widths, and downstream tasks show that ReQuant consistently improves quantized models from heterogeneous PTQ initializers, with especially large gains on simple initializers and lower bit-widths. Notably, ReQuant can refine a simple round-to-nearest initialization across multiple sweeps until it approaches or surpasses GPTAQ under the same quantization format. These results establish ReQuant as a practical complementary stage for further improving existing PTQ pipelines.

cs.AI

GPUMDkit: A User-Friendly Toolkit for GPUMD and NEP

Machine-learned interatomic potentials have revolutionized molecular dynamics simulations by providing quantum-mechanical accuracy at empirical-potential speeds. The graphics processing unit molecular dynamics (GPUMD) package, featuring the highly efficient neuroevolution potential (NEP) framework, has emerged as a powerful tool in this domain. However, the complexity of force field development, active learning, and trajectory post-processing often requires extensive manual scripting, imposing a steep learning curve on new users. To address this, we present GPUMDkit, a comprehensive and user-friendly toolkit that streamlines the entire simulation workflow for GPUMD and NEP. GPUMDkit integrates a suite of essential functionalities, including format conversion, structure sampling, property calculation, and data visualization, accessible through both interactive and command-line interfaces. Its modular, extensible architecture ensures accessibility for users of all experience levels while allowing seamless integration of new features. By automating complex tasks and enhancing productivity, GPUMDkit substantially lowers the barrier to using GPUMD and NEP programs. This article describes the program architecture and demonstrates its capabilities through practical applications.

cond-mat.mtrl-sci

A Perspective on Training Machine Learning Force Fields for Solid-State Electrolyte Materials

Machine learning force fields enable high-accuracy modeling of solid-state electrolytes (SSEs). This perspective evaluates dataset size, reference quality, and model architectures. We show that rigid SSE frameworks favor efficient learning, prioritizing data quality over quantity. Crucially, force RMSE does not reliably predict transport performance. By analyzing locality and benchmarking frameworks, we provide practical guidelines to accelerate the development of next-generation solid-state batteries.

cond-mat.mtrl-sci

qNEP: A highly efficient neuroevolution potential with dynamic charges for large-scale atomistic simulations

Although electrostatics can be incorporated into machine-learned interatomic potentials, existing approaches are computationally very demanding, limiting large-scale, long-time simulations of electrostatics-driven phenomena such as dielectric response, infrared activity, and field-matter coupling. Here, we extend the neuroevolution potential (NEP), a highly efficient machine-learned interatomic potential, to a charge-aware framework (qNEP) by introducing explicit, environment-dependent partial charges. Each ionic partial charge is represented by a neural network as a function of the local descriptor vector, analogous to the NEP site-energy model. This formulation enables the direct prediction of the Born effective charge tensor for each ion and, consequently, the polarization. As a result, dielectric properties, infrared spectra, and coupling to external electric fields can be evaluated within a unified framework. We derive consistent expressions for the forces and virials that explicitly account for the position dependence of the partial charges. The qNEP method has been implemented in the free-and-open-source GPUMD package, with support for both Ewald summation and particle-particle particle-mesh treatments of electrostatics. We demonstrate the accuracy and efficiency of the qNEP approach through representative applications to water, Li7La3Zr2O12, BaTiO3, and a magnesium-water interface. These results show that qNEP enables accurate atomistic simulations with explicit long-range electrostatics, scalable to million-atom systems on nanosecond time scales using consumer-grade GPUs.

physics.comp-ph

Step-GUI Technical Report

Recent advances in multimodal large language models unlock unprecedented opportunities for GUI automation. However, a fundamental challenge remains: how to efficiently acquire high-quality training data while maintaining annotation reliability? We introduce a self-evolving training pipeline powered by the Calibrated Step Reward System, which converts model-generated trajectories into reliable training signals through trajectory-level calibration, achieving >90% annotation accuracy with 10-100x lower cost. Leveraging this pipeline, we introduce Step-GUI, a family of models (4B/8B) that achieves state-of-the-art GUI performance (8B: 80.2% AndroidWorld, 48.5% OSWorld, 62.6% ScreenShot-Pro) while maintaining robust general capabilities. As GUI agent capabilities improve, practical deployment demands standardized interfaces across heterogeneous devices while protecting user privacy. To this end, we propose GUI-MCP, the first Model Context Protocol for GUI automation with hierarchical architecture that combines low-level atomic operations and high-level task delegation to local specialist models, enabling high-privacy execution where sensitive data stays on-device. Finally, to assess whether agents can handle authentic everyday usage, we introduce AndroidDaily, a benchmark grounded in real-world mobile usage patterns with 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios (8B: static 89.91%, end-to-end 52.50%). Our work advances the development of practical GUI agents and demonstrates strong potential for real-world deployment in everyday digital interactions.

cs.CV

Anyon Dispersion in Aharonov-Casher Bands and Implications for Twisted MoTe${}_2$

The discovery of fractional quantum anomalous Hall (FQAH) states in two-dimensional heterostructures has opened the door to realizing phases of dispersing anyons. Here, we develop an analytically controlled theory of anyon dispersion in FQAH states realized in ideal or Aharonov-Casher (AC) bands by projecting interactions onto the space of Laughlin quasiholes. Constructing quasihole momentum eigenstates allows efficient evaluation of the single quasihole dispersion using Monte Carlo. We find that the quasihole bandwidth grows with increasing quantum-geometry inhomogeneity of the AC band and with increasing interaction screening length. For realistic parameters relevant to the bands of twisted MoTe${}_2$, the quasihole bandwidth is of order 1 meV and increases with increasing displacement field, suggesting that itinerant-anyon physics may play an important role in sufficiently clean samples. Furthermore, we develop a microscopic Lagrangian framework in terms of a quasihole guiding-center coordinate, which reproduces the momentum-space formula for the dispersion. This approach reveals that quasihole dispersion originates from the combined effects of an interaction-generated periodic potential, arising from non-uniform quantum geometry of the single particle bands, and the quasihole many-body Berry phase arising from the background magnetic field. The latter endows the guiding-center coordinate with a noncommutative structure, converting the periodic potential into a finite dispersion. Finally, we outline how this framework generalizes to multiple quasiholes, enabling a microscopic theory of charged excitations in FQAH systems that retains only the anyon degrees of freedom.

cond-mat.str-el

Dynamical entropy of charged black objects

We develop a general framework for electromagnetic potential-charge contributions to the first law of black hole mechanics, applicable to dynamical first-order perturbations of stationary black objects with possibly non-compact bifurcate Killing horizons. Working in the covariant phase space formalism, we derive both comparison and physical process versions of the first law. We consider generic diffeomorphism-invariant theories of gravity in $D$ spacetime dimensions, containing non-minimally coupled abelian $p$-form gauge fields. The pullback of the gauge field to the horizon is allowed to diverge while its field strength remains smooth, yielding gauge-invariant electric potential-charge pairs in the first law. We further extend the construction to include magnetic charges by developing a bundle-covariant, gauge-invariant prescription that fixes the Jacobson-Kang-Myers ambiguity in the improved Noether charge. Electric and magnetic charges are, respectively, associated with non-trivial $(D - p - 1)$- and $(p + 1)$-cycles of the horizon cross-section, whose homology classes determine the number of independent potential-charge pairs through the Betti numbers $b_{D - p - 1}$ and $b_{p + 1}$. Further, the dynamical gravitational entropy entering the first law is identified with the gauge-invariant part of the improved Noether charge, giving a gauge-invariant extension of the recent proposal by Hollands, Wald and Zhang. We illustrate our framework with dyonic AdS black holes, dipole black rings, and charged black branes.

hep-th

MCTuner: Spatial Decomposition-Enhanced Database Tuning via LLM-Guided Exploration

Database knob tuning is essential for optimizing the performance of modern database management systems, which often expose hundreds of knobs with continuous or categorical values. However, the large number of knobs and the vast configuration space make it difficult to identify optimal settings efficiently. Although learning-based tuning has shown promise, existing approaches either ignore domain knowledge by relying solely on benchmark feedback or struggle to explore the high-dimensional knob space, resulting in high tuning costs and suboptimal performance. To address these challenges, we propose MCTuner, an adaptive knob tuning framework that minimizes exploration in ineffective regions of the configuration space. MCTuner employs a Mixture-of-Experts (MoE) mechanism with specialized LLMs to identify performance-critical knobs. In further, MCTuner introduces the first spatial decomposition algorithm that recursively partitions the space into hierarchical subspaces, on which Bayesian Optimization is performed to efficiently search for near-optimal configurations. Evaluated on different benchmarks (OLAP, OLTP, and HTAP), MCTuner achieves up to 19.2% performance gains and 1.4x faster configuration discovery per iteration compared to state-of-the-art methods.

cs.DB

Generalised focusing theorem and dynamical horizon entropy in diffeomorphism-invariant theories

I summarise recent progress on light-ray focusing and horizon thermodynamics in general diffeomorphism-invariant theories of gravity coupled to bosonic matter. In pure gravity and with scalar or vector fields, the null-null gravitational equation of motion on a linearly perturbed Killing horizon generalises the Raychaudhuri equation, defining a generalised expansion that never increases under the null energy condition. This proves a generalised focusing theorem and defines an increasing horizon entropy (Wall entropy). When higher-spin fields are present, the generalised focusing theorem persists subject to a "higher-spin focusing condition", which I propose as a physical consistency constraint on higher-spin theories.

gr-qc

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation

Embodied agents exhibit immense potential across a multitude of domains, making the assurance of their behavioral safety a fundamental prerequisite for their widespread deployment. However, existing research predominantly concentrates on the security of general large language models, lacking specialized methodologies for establishing safety benchmarks and input moderation tailored to embodied agents. To bridge this gap, this paper introduces a novel input moderation framework, meticulously designed to safeguard embodied agents. This framework encompasses the entire pipeline, including taxonomy definition, dataset curation, moderator architecture, model training, and rigorous evaluation. Notably, we introduce EAsafetyBench, a meticulously crafted safety benchmark engineered to facilitate both the training and stringent assessment of moderators specifically designed for embodied agents. Furthermore, we propose Pinpoint, an innovative prompt-decoupled input moderation scheme that harnesses a masked attention mechanism to effectively isolate and mitigate the influence of functional prompts on moderation tasks. Extensive experiments conducted on diverse benchmark datasets and models validate the feasibility and efficacy of the proposed approach. The results demonstrate that our methodologies achieve an impressive average detection accuracy of 94.58%, surpassing the performance of existing state-of-the-art techniques, alongside an exceptional moderation processing time of merely 0.002 seconds per instance.

cs.AI

Improving robustness and training efficiency of machine-learned potentials by incorporating short-range empirical potentials

Machine learning force fields (MLFFs) are powerful tools for materials modeling, but their performance is often limited by training dataset quality, particularly the lack of rare event configurations. This limitation undermines their accuracy and robustness in long-time and large-scale molecular dynamics simulations. In this work, we present a hybrid MLFF framework that integrates an empirical short-range repulsive potential and demonstrates improved robustness and training efficiency. Using solid electrolyte Li$_7$La$_3$Zr$_2$O$_{12}$ (LLZO) as a model system, we show that purely data-driven MLFFs fail to prevent unphysical atomistic clustering in extended simulations due to inadequate short-range repulsion. In contrast, the hybrid force field eliminates these artifacts, enabling stable long-time simulations, which are critical for studying various properties of LLZO. The hybrid framework also reduces the need for extensive active learning and performs well with just 25 training configurations. By combining physics-driven constraints with data-driven flexibility, this approach is compatible with most existing MLFF architectures and establishes a universal paradigm for developing robust, training-efficient force fields for complex material systems.

cond-mat.mtrl-sci

Unveiling the Oxidation Mechanisms of Octa-Penta Graphene: A Multidimensional Exploration from First-Principles to Machine Learning

Octa-penta graphene (OPG), a novel carbon allotrope characterized by its distinctive arrangement of pentagonal and octagonal rings, has garnered considerable attention due to its exceptional structure and functional properties. This study systematically investigates the oxidation mechanisms of OPG and elucidates the oxygen migration patterns on the OPG monolayer through first-principles calculations and machine-learning-based molecular dynamics (MLMD) simulations. Specifically, the oxidation processes on OPG-L and OPG-Z involve exothermic chemisorption, where oxygen molecules dissociate at the surfaces, forming stable epoxy groups. Furthermore, the integrated-crystal orbital Hamilton population (ICOHP) and Bader charge analyses provide insights into the physical mechanisms of oxygen atom adsorption. Importantly, we found that oxidation also impact the electronic properties of OPG, with OPG-L retaining its metallic characteristics post-oxygen adsorption, whereas OPG-Z undergoes a transformation from a metallic to a semiconducting state due to the introduction of oxygen. Oxygen migration on OPG monolayer involves breaking and reforming of C-O bonds, with varying stability across adsorption sites and limited migration along the basal plane. MLMD simulations corroborate these migration patterns, offering detailed migration trajectories consistent with theoretical predictions. These findings enhance the understanding of oxygen migration dynamics on OPG, facilitate its experimental validations, and highlight its potential as a novel 2D material for applications in batteries, heat-resistant materials, and oxidation-resistant coatings.

cond-mat.mtrl-sci

Gravitational focusing and horizon entropy for higher-spin fields

Previously, the Raychaudhuri equation and the focusing theorem in General Relativity were generalised to diffeomorphism-invariant theories of gravity coupled to scalar and vector fields on linearly perturbed Killing horizons. The Wall entropy can be extracted from the generalised focusing equation and it satisfies the first and the second laws of thermodynamics. In this paper, we further extend the discussion of gravitational focusing on the horizon to include arbitrary bosonic fields with spin $s \geq 2$. These higher-spin fields introduce indefinite terms into the generalised focusing equation, obstructing the proof of the focusing theorem and the existence of an increasing horizon entropy. To resolve this issue, we propose a higher-spin focusing condition that eliminates these indefinite terms, thereby restoring the focusing theorem and the associated thermodynamic laws. We speculate that the focusing condition could be a necessary condition for the physical consistency of higher-spin theories.

gr-qc

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

A chatbot's personality design is key to interaction quality. As chatbots evolved from rule-based systems to those powered by large language models (LLMs), evaluating the effectiveness of their personality design has become increasingly complex, particularly due to the open-ended nature of interactions. A recent and widely adopted method for assessing the personality design of LLM-based chatbots is the use of self-report questionnaires. These questionnaires, often borrowed from established human personality inventories, ask the chatbot to rate itself on various personality traits. Can LLM-based chatbots meaningfully "self-report" their personality? We created 500 chatbots with distinct personality designs and evaluated the validity of their self-report personality scores by examining human perceptions formed during interactions with these chatbots. Our findings indicate that the chatbot's answers on human personality scales exhibit weak correlations with both human-perceived personality traits and the overall interaction quality. These findings raise concerns about both the criterion validity and the predictive validity of self-report methods in this context. Further analysis revealed the role of task context and interaction in the chatbot's personality design assessment. We further discuss design implications for creating more contextualized and interactive evaluation.

cs.HC

A Comment on Deriving the Gibbons-Hawking-York Term From the String Worldsheet

In this note, we show that the noncovariant metric boundary term obtained from the nonlinear sigma model worldsheet derivation of the bulk off-shell sphere partition function is closely related to the Einstein boundary term in the Gamma-Gamma noncovariant action. In fact, when expressed in terms of the trace of the extrinsic curvature tensor, we illustrate that this boundary term has one-half the coefficient of the Gibbons-Hawking-York boundary term required such that the total (bulk plus boundary) off-shell classical action has a well-posed variational principle with Dirichlet boundary conditions.

hep-th

Social Life Simulation for Non-Cognitive Skills Learning

Non-cognitive skills are crucial for personal and social life well-being, and such skill development can be supported by narrative-based (e.g., storytelling) technologies. While generative AI enables interactive and role-playing storytelling, little is known about how users engage with and perceive the use of AI in social life simulation for non-cognitive skills learning. Additionally, the benefits of AI mentorship on self-reflection awareness and ability in this context remain largely underexplored. To this end, we introduced Simulife++, an interactive platform enabled by a large language model (LLM). The system allows users to act as protagonists, creating stories with one or multiple AI-based characters in diverse social scenarios. In particular, we expanded the Human-AI interaction to a Human-AI-AI collaboration by including a Sage Agent, who acts as a bystander, providing users with some perspectives and guidance on their choices and conversations in terms of non-cognitive skills to promote reflection. In a within-subject user study, our quantitative results reveal that, when accompanied by Sage Agent, users exhibit significantly higher levels of reflection on motivation, self-perceptions, and resilience & coping, along with an enhanced experience of narrative transportation. Additionally, our qualitative findings suggest that Sage Agent plays a crucial role in promoting reflection on non-cognitive skills, enhancing social communication and decision-making performance, and improving overall user experience within Simulife++. Multiple supportive relationships between Sage Agent and users were also reported. We offer design implications for the application of generative AI in narrative solutions and the future potential of Sage Agent for non-cognitive skill development in broader social contexts.

cs.CL

Supporting Mitosis Detection AI Training with Inter-Observer Eye-Gaze Consistencies

The expansion of artificial intelligence (AI) in pathology tasks has intensified the demand for doctors' annotations in AI development. However, collecting high-quality annotations from doctors is costly and time-consuming, creating a bottleneck in AI progress. This study investigates eye-tracking as a cost-effective technology to collect doctors' behavioral data for AI training with a focus on the pathology task of mitosis detection. One major challenge in using eye-gaze data is the low signal-to-noise ratio, which hinders the extraction of meaningful information. We tackled this by levering the properties of inter-observer eye-gaze consistencies and creating eye-gaze labels from consistent eye-fixations shared by a group of observers. Our study involved 14 non-medical participants, from whom we collected eye-gaze data and generated eye-gaze labels based on varying group sizes. We assessed the efficacy of such eye-gaze labels by training Convolutional Neural Networks (CNNs) and comparing their performance to those trained with ground truth annotations and a heuristic-based baseline. Results indicated that CNNs trained with our eye-gaze labels closely followed the performance of ground-truth-based CNNs, and significantly outperformed the baseline. Although primarily focused on mitosis, we envision that insights from this study can be generalized to other medical imaging tasks.

cs.CV