SearcharxivSearch

arXiv subjects

Xi Zhang

Publications and source records attributed to Xi Zhang.

At least 19 recordsLinked to original sources

A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors

Modern AI agent harnesses expose lifecycle hooks that bind shell commands to runtime events such as session start, tool calls, and file edits. These commands run with host privileges yet ship as lifecycle-hook configuration and may fire at times the LLM never observes. We identify the lifecycle-hook update path, which harnesses trust blindly, as a new attack surface. Under a supply-chain threat model in which an attacker controls only plugin metadata and lifecycle-hook configuration, a benign versioned plugin can be trojanized by an update that silently binds attacker-chosen commands to benign events, yielding malicious host-side behavior such as privilege escalation. We propose HookPry, an open-source and fully automated attack framework that systematically exploits this vulnerability across heterogeneous AI agent harnesses. HookPry realizes ten attack objectives; across 25 combinations of harnesses and backends in 1,000 end-to-end runs, it compromises all seven evaluated harnesses, with per-harness success rates reaching 92.5%. Representative defenses remain insufficient: Microsoft Defender has 0% recall, and the union of three static defenses misses 47.5% of malicious artifacts.

cs.CR

Accelerating Chemical Kinetics for Exoplanet Atmospheres using Neural Networks

Observations increasingly reveal the coupled radiative, chemical, and dynamical processes that shape exoplanet atmospheres. Interpreting these atmospheres requires models that can capture this complexity. However, multidimensional models remain fundamentally limited by computational cost, and answering key questions requires simulating the governing physical mechanisms at speeds classical methods cannot achieve. As a result, models often rely on simplifying approximations, such as equilibrium chemistry, even when those assumptions miss important effects. There is a pressing need for fast and accurate chemical kinetics solvers to model planetary atmospheres. Here we present a machine learning local-box chemical kinetics solver for exoplanet atmospheres using a residual flow-map architecture. We demonstrate that this surrogate model is several orders of magnitude faster than a classical solver, achieving microsecond-scale inference while retaining percent-level accuracy. The surrogate model covers a parameter space that spans $T=300$-$3000$ K, $P=10^{-6}$-$10^{4}$ bar, $\Delta t=10^{-3}$-$10^{8}$ s, and compositions ranging from $10^{-2}$ to $10^{3}$ times solar in both C/O ratio and metallicity. Our model outperforms several commonly used machine learning architectures and performs robustly under the extreme stiffness characteristic of atmospheric chemistry. The machine learning framework presented here is a flexible and efficient approach to emulating state-to-state flow-map problems that commonly arise in numerical simulations.

astro-ph.EP

MEL: Coordinate-Preserving EEG Tokenization for fMRI Translation

Translating electroencephalography (EEG) into functional magnetic resonance imaging (fMRI) is important for medical neuroimaging, clinical brain-state monitoring, and multimodal neural decoding, because it aims to infer spatially organized hemodynamic activity from fast and accessible electrophysiological recordings. Existing EEG-to-fMRI studies mainly pursue stronger decoders, but the problem is also constrained by a representation-interface mismatch: fMRI responses are delayed, temporally integrated, and spatially distributed, whereas generic EEG encodings often entangle temporal lag, channel identity, and frequency-band structure. We propose Multi-band EEG Latent-state Tokenization (MEL), a coordinate-preserving EEG representation framework that anchors each target fMRI response to its preceding EEG history and organizes it into lag-channel-frequency neural-state tokens. By explicitly capturing hemodynamic latency and spectral-spatial dynamics, MEL aligns fMRI-pertinent EEG representations with capacity-controlled readouts without depending entirely on model scaling. Experiments on VU EEG-fMRI benchmarks and external Oddball data show that MEL improves prediction over strong NeuroBOLT baselines. Ablations and controls further indicate that the gains come from structured EEG representation rather than leakage, shortcut statistics, or decoder capacity.

cs.LG

Propagation rates in integro-differential equations of ignition type

This paper provides sharp rates of invasion in nonlinear integro-differential equations of ignition type. The framework covers both integrable convolution kernels and singular fractional-type kernels. While the emergence of a threshold on the decay $1+2s$ of the jump kernel is expected from earlier works, far fewer quantitative results were available in such a general framework. In particular, in addition to the spreading rates, our study provides information on the shape of the invasion profiles. Importantly, we obtain the first sharp spreading estimate in the critical case $s=\frac12$, which separates linear-in-time and accelerated regimes.

math.AP

Decoupled I/O-Dominant Pipelines for Large-Scale Whole-Slide Image Embedding Extraction

Whole-slide images (WSIs) are central to computational pathology but are prohibitively large, making patch-based processing the practical unit for foundation model inference. At scale, however, generating and handling massive numbers of patches on quickly introduces significant I/O and orchestration overhead, often dominating end-to-end performance. We present a decoupled, I/O-aware pipeline for large-scale WSI embedding extraction that decomposes the workflow into three stages: (1) patch generation and staging, (2) embarrassingly parallel embedding inference, and (3) sharded vector database ingestion. This design isolates data movement from compute, enabling efficient patch delivery, scalable multi-node inference with minimal communication. The resulting system produces a distributed vector database where embeddings are persistently coupled with rich metadata (e.g., patient, slide, and patch attributes), enabling efficient filtering, retrieval, and downstream reuse. This representation database is compact and reusable for tasks such as retrieval, classification, and few-shot learning, particularly benefiting low-resource environments. We show that decoupling I/O, computation, and ingestion enables high-throughput WSI embedding extraction at scale. By characterizing the scaling envelope, we demonstrate that storage dominates beyond moderate concurrency, reframing WSI embedding extraction as a data-centric systems problem rather than a purely compute-bound workload.

cs.DC

Tabular Foundation Models for Multi-View Information Cascade Popularity Prediction

Predicting the future popularity of information cascades is essential for understanding information diffusion on social media. Despite recent advances, existing methods face two key limitations: they focus primarily on the cascade view while overlooking other information views that drive user engagement, such as textual semantics, visual content, and tabular attributes; and they fail to capture high-order cross-view interactions. To address these issues, we propose \textbf{TFM4POP}, the first framework to introduce tabular foundation models (TFMs) into popularity prediction, leveraging their pre-trained tabular priors to unify the modeling of multiple heterogeneous information views. Specifically, TFM4POP adopts a dual-branch design: the static branch employs a TFM as the feature-encoding backbone that jointly reasons over all static views through in-context learning to produce the static cascade representation, while the dynamic branch captures the continuous-time cascade dynamics with a dedicated Neural-ODE-based encoder. The two representations are then fused via cross-attention for the final prediction. Furthermore, to adapt the TFM to real cascade distributions, we apply parameter-efficient IA3 fine-tuning, achieving performance competitive with or better than full fine-tuning while updating substantially fewer parameters. In addition, we construct a comprehensive multi-view cascade benchmark that covers all four information views. Extensive experiments show that TFM4POP consistently outperforms state-of-the-art baselines across multiple datasets and observation settings.

cs.SI

Backward Layout Search for Sequence-Constrained Robotic Assembly

Robotic assembly layout planning must determine the assembly site and the initial pose of each part while ensuring collision-free execution of a prescribed assembly sequence. This problem is challenging because the obstacle environment changes after each assembly step, and unassembled parts re maining in the workspace may block robot motions. We observe that the feasibility of each assembly step depends only on the initial poses of the current and later-assembled parts. Based on this dependency, we propose Backward Layout Search (BLS), which assigns initial part poses in reverse assembly order. Each expansion performs geometric, kinematic, grasp, and prescribed-motion checks, while collision masks and candidate set filtering remove infeasible initial part pose candidates. Promising partial layouts are retained through beam selection, and complete layouts are validated by full motion planning in forward assembly order. Experiments on five assembly models show that BLS produces collision-free executable layouts and reduces step evaluations and search time compared with a matched forward search.

cs.RO

GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities

LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both of which are difficult to realize in practice. We therefore formulate a new threat, Memory-Mediated Polarization Cascade, which uses agent memory as a persistence channel and public discussion as a propagation channel. This threat contains three stages. During exposure and memory retention, the attacker exposes a small set of target agents to arguments that reinforce their respective stated stances. The targets' memory systems then process and retain these arguments. During retrieval and reproduction, a shared stance-neutral discussion cues the targets to retrieve and reproduce their respective retained arguments. During iterative propagation, untreated agents influenced by the reproduced arguments restate and spread them. We instantiate this threat in GraphWake with three components: (i) stance-support argumentation knowledge graphs construct knowledge-based arguments; (ii) axiom-oriented triple selection distills them for reliable retention and reproduction; and (iii) stance-neutral memory cueing triggers concurrent retrieval and reproduction, initiating propagation. Experiments across multiple discussions and memory systems show that GraphWake substantially increases group polarization. These findings reveal a community-level polarization risk.

cs.AI

Optical-NIR Multi-band Photometric Analysis and Characterization of Giant Exoplanets with CPI-C

We present a multi-band photometric approach to characterize giant exoplanets, which represents one of the anticipated core scientific outcomes of Cool Planet Imaging Coronagraph (CPI-C). CPI-C operates with two observational channels covering visible and near-infrared wavelengths, each equipped with four broadband filters. The planet--star flux ratio integrated over each filter bandpass is calculated for photometric analysis. For cool planets observed in the visible bands, the data are primarily used to fit the overall spectral shape and methane-induced modulation, providing sensitivity to metallicity- and cloud-dependent spectral variations while constraining the reflected-light spectral shape and the combined scaling involving planet radius, orbital separation, and orbital phase. In the near-infrared bands, which probe thermal emission, the data help to better constrain fundamental planetary parameters including the effective temperature, radius, surface gravity and mass. For a synthetic giant planet with measurable reflected-light and thermal-emission components, the combined VIS4+NIR4 data provide tighter same-target constraints than either filter set alone, especially for the planet radius and cloud sedimentation parameter. Our simulations incorporate realistic instrument throughput, detector noise, and residual speckle noise. The results demonstrate that the eight-band design spanning visible to near-infrared wavelengths supports reflected-light diagnostics, thermal-emission characterization, and joint optical--NIR analysis of giant exoplanets within CPI-C science observations.

astro-ph.EP

Juno Microwave Observations Reveal Jupiter's Deep Alkali-Chlorine Relation

The longest-wavelength channel of the Juno Microwave Radiometer (MWR) probes Jupiter's kilobar atmosphere through free electrons produced by sodium and potassium ionization. Under equilibrium chemistry the electron abundance is the small residual of the charge balance between alkali cations and the anions Cl- and HS-. Chlorine is not directly measurable in Jupiter's deep atmosphere because gaseous HCl is removed from the observable atmosphere by NH4Cl condensation, whereas sulfur has been measured by the Galileo probe. The MWR-derived electron measurement therefore constrains the alkali-to-chlorine ratio rather than the alkali abundance alone. We combine the MWR observations with equilibrium chemistry and microwave radiative transfer in a Bayesian framework, finding that the deep gas-phase elemental alkali-to-chlorine abundance ratio is (Na+K)/Cl = 0.05 over 0.3-5 times solar in chlorine, about 180 times below the protosolar ratio of 8.7. At 3 times solar chlorine, the inferred alkali metallicity is 1.6 x 10^-2 times solar (1 sigma: 1.2 x 10^-2 - 2.7 x 10^-2 times solar), while at low chlorine abundance HS- sets an alkali floor near 10^-3 times solar. The inferred gas-phase alkali abundance exceeds the ~10^-5 times solar threshold by more than two orders of magnitude and rules out the long-proposed global kilobar radiative zone. Because sodium and potassium are refractory whereas chlorine is volatile, the inferred ratio provides a new diagnostic of the rock-to-ice balance in the solids accreted by Jupiter. This compositional interpretation assumes equilibrium chemistry; if lofted mineral clouds instead control the electron abundance under disequilibrium conditions, the inferred alkali-chlorine relationship need not hold.

astro-ph.EP

Alkali Metallicity, Mineral Clouds, and Deep Atmospheric Variability on Jupiter

The bulk elemental abundances of Jupiter provide critical insights into its formation history and interior structure. Recent observations by the Juno Microwave Radiometer (MWR) reveal a deep Jovian atmosphere significantly depleted in electrons, implying an alkali metal (Na, K) abundance of 10^-1 - 10^-5 times solar. This depletion stands in sharp contrast to the supersolar volatile enrichments measured by the Galileo probe. We propose that this apparent depletion arises from mineral cloud-induced processes deep in the atmosphere. We explore two physical mechanisms using thermochemical and microphysical modeling. In the "chemical sequestration" scenario, vigorous vertical mixing lofts deep refractory condensates (e.g., spinel) into the 1000-2000 bar region, where they react to form alkali feldspars (albite) and feldspathoids (leucite), efficiently sequestering gaseous Na and K. In the "dust-catalyzed recombination" scenario, the bulk alkali inventory remains gaseous, but the free electron density is suppressed by dust-plasma interactions. Thermally emitted alkali ions from the surfaces of micron-sized iron and silicate grains significantly increase the cation density, driving rapid recombination of free electrons. Both mechanisms allow for a bulk solar or even supersolar alkali inventory while suppressing the electron density to match Juno observations. Analyzing an extended dataset of MWR observations with 61 perijoves, we detect spatial variability in the deep atmosphere that suggests modulation by mineral clouds. Our findings challenge the traditional rainout framework, unveiling a deep "mineralogical zone" in Jupiter shaped by dynamics and heterogeneous chemistry, resembling the photospheres of hot exoplanets and brown dwarfs.

astro-ph.EP

Qwen-CUA: Native Computer Use for (almost) Everything

Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. It observes only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, or task-specific APIs. Its scaffold maintains up to 20 active screenshots and folds older visual history in fixed-size blocks to retain recent evidence while preserving reusable prompt prefixes. For training, we build a cloud rollout fleet with access to nearly 100,000 vCPUs and tens of thousands of concurrent environments, construct approximately 40,000 verifiable tasks, and collect personalized long-horizon workflows across everyday and professional software. We optimize complete trajectories with verifiable rewards and trajectory slicing, while iterative training runs refresh supervised data and recalibrate reinforcement-learning tasks. Across eight benchmarks, Qwen-CUA outperforms Qwen3.7 and remains competitive with leading proprietary systems, reaching 86.2 on OSWorld-Verified and 18.5/48.4 binary/partial completion on OSWorld 2.0. Scaling the same recipe to a model with over one trillion parameters yields Qwen-CUA-Max, improving these scores to 87.6 and 21.2/53.3. Qwen-CUA also reduces RedTeamCUA attack success from 36.6 to 16.4 relative to Qwen3.7. Efficiency analyses, a browser deployment, and Bash-augmented experiments further characterize practical behavior. These results establish native computer use as a broadly capable agent foundation and highlight scalable verifiable interaction and hybrid tool use as key directions.

cs.LG

A Physically Driven Parameterisation of Multidimensional Atmospheres: Application to the JWST Phase Curve of WASP-121b

Understanding the multidimensional structure of strongly irradiated exoplanets is essential for interpreting their atmospheric dynamics, chemistry and energy transport, yet current analyses remain limited by the difficulty of extracting reliable phase-resolved spectra and by the lack of physically interpretable parameterisations for retrievals. We combine a data-driven eclipse-normalisation method with an analytical three-dimensional temperature parameterisation derived from radiative, advective and diffusive energy balance and controlled by a few characteristic timescales. Applied to JWST/NIRSpec G395H observations of WASP-121b, the method yields spectra consistent with conventional phase-curve fitting, while the parameterisation reproduces the large-scale thermal structures predicted by general circulation models. The preferred retrieval reveals a pronounced day--night contrast, a dayside thermal inversion extending to both limbs, an inversion over part of the nightside, and limb temperatures differing by several hundred kelvin. Dynamical transport strengthens with pressure, and the hotspot offset increases from $\sim4^\circ$ to $\sim9^\circ$ across the pressures probed by G395H. The confined dayside hot region and the small, pressure-dependent offsets lie closer to the $\sim$3~G GCM than to its non-magnetic counterpart, although Rayleigh drag cannot be excluded. The spectra also favour distinct dayside and nightside chemical states, with more nightside CH$_4$ than the cooler temperatures alone can explain, pointing to disequilibrium chemistry. The retrieved thermal structure further implies an inhomogeneous cloud distribution, with condensation favoured on the nightside and cooler morning limb. The framework provides a computationally efficient, physically interpretable path from spectroscopic phase curves to multidimensional atmospheric structure.

astro-ph.EP

An Ontology for Machine Learning Interatomic Potentials

Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density functional theory (DFT) or wave-function methods---at a fraction of the cost. The field encompasses a growing ecosystem of algorithms, training datasets, hyperparameters, and target materials, yet the metadata needed to systematically compare, reproduce, and build upon MLIP studies remains scattered across papers, scripts, and ad-hoc file formats. We present the MLIPs ontology, an OWL 2 DL ontology that captures the concepts needed to describe MLIP methods, their hyperparameters, training datasets with DFT provenance, and published benchmarks. The ontology is organized into three modules---Method, Training Data, and Benchmark---and connects existing ontologies in materials science (MDO, CMSO/ASMO) and machine learning (ML-Schema), complementing dataset-side schemas such as Croissant. It declares 27 formal axioms enforcing data completeness and consistency, including property chains that link trained models to their methods and training data. We demonstrate the ontology through a running example based on Moment Tensor Potentials and evaluate it through competency-question execution on a 20-paper seeded knowledge graph, OWL reasoning, and comparison with existing ontologies.

cs.AI

Towards the Explainability of Temporal Graph Networks via Memory Backtracking and Topological Attribution

Temporal graphs are ubiquitous in real-world applications and Temporal Graph Networks (TGNs) have achieved superior predictive accuracy. Understanding which historical events drive model predictions can enhance trustworthiness of TGNs. Existing explanation methods overlook the memory module, the core component that records and updates node histories, leaving the influence of past events unexplored. To address this, we attribute TGNs predictions through the topology attribution tree and memory backtracking tree. The topology attribution tree captures the influence of neighbors and their memory vectors, then the memory backtracking tree quantifies how historical events shape node memory vectors. We apply the LRP in TGNs, ensuring that the total contribution of events equals the logits of model. Finally, top-k selection may be unfaithful due to the nonlinear mapping from logits to probabilities, we design optimization objectives to identify the important events. Experiments on nine temporal graph datasets, spanning node property prediction, link prediction tasks and graph classification tasks, show that our method provides faithful explanations and outperforms state-of-the-art baselines. The code is available at https://github.com/yazhengliu/MemExplainer

cs.LG

Adoption and Ecosystem Health: A Longitudinal Analysis of Open-Source Multi-Agent Frameworks

Since ChatGPT's launch in November 2022, open-source agentic frameworks have proliferated, making framework selection important for engineering teams while obscured by popularity signals such as GitHub stars. This paper analyzes 15 major open-source AI agent framework repositories from late 2022 to early 2026, using 808,042 stars, 73,997 pull requests, 86,241 commits, and 987,330 user profiles to assess ecosystem health across awareness, adoption, and retention. Three findings emerge. First, headline popularity is unreliable. Star counts reflect hype cycles and inorganic activity. AutoGPT gained 111,967 stars in one month but converted fewer than 9 contributors per 1,000 stars, defined as contributor density in this research, compared with LangChain's 41. Lower-profile frameworks such as Pydantic-AI show higher contributor density, indicating deeper adoption. Second, mapping awareness against adoption shows that visibility and engagement diverge. MetaGPT and LangFlow have contributor density ratios below 5 even with their high visibility. Openai-agents-python's limited contributor base suggests institutional backing alone does not ensure community depth. By analyzing cross-framework contribution, we discover that LangChain functions as a shared infrastructure, attracting 82.5% of cross-ecosystem contributors. Third, retention drops most steeply in the first 30 days of initial contribution and stabilizes near 90 days. Overall, ecosystem health is better measured by contributor density, cross-ecosystem engagement, and retention than by stars alone. These metrics offer teams a more robust basis for framework evaluation.

cs.MA

LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models

The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video quality benchmarks predominantly focus on short clips and isolated distortions, overlooking the temporal continuity, cumulative degradation, and reasoning complexity inherent in long-duration content. To address these limitations, we present LongVQUBench, a comprehensive benchmark for long-term video quality understanding. LongVQUBench contains over 1200 diverse videos spanning movies, documentaries, surveillance footage, egocentric recordings, and animated content, accompanied by 1500 multiple-choice and open-ended questions for validation and testing. To assess perceptual reasoning across different temporal scopes, we introduce three progressively complex evaluation levels: (i) local event quality understanding (LQU) for analyzing localized distortions; (ii) cross-event quality reasoning (CQR) for integrating multiple degraded events; and (iii) global quality understanding (GQU) for holistic perceptual evaluation over extended durations. Furthermore, a needle distortion question-answering (NDQA) paradigm is embedded across all three levels, where spatial or temporal artifacts are sparsely inserted to probe fine-grained detection and reasoning capabilities. Extensive experiments on 14 state-of-the-art LVLMs reveal significant performance degradation with increasing video length and reasoning depth, highlighting their limited capacity for long-range temporal integration and perceptual attribution. We envision LongVQUBench as a foundational step toward the systematic, hierarchical, and explainable evaluation of LVLMs' long-term video quality understanding.

cs.CV