SearcharxivSearch

arXiv subjects

Xin Lan

Publications and source records attributed to Xin Lan.

13 recordsLinked to original sources

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.

cs.AI

The identification between the bulk and boundary conserved quantities

By using Wald formalism, we show that the identification between the bulk and boundary conserved quantities induced by the perturbation of generic non-electromagnetic matter field holds not only on top of the asymptotically flat stationary spacetimes but also on top of the asymptotically AdS stationary ones. We further show that such an identification reduces to the familiar form for the test point particle by viewing it as the limiting case of general matter.

hep-th

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation runs the 87-task benchmark under matched no-Skills and curated-Skills conditions for 18 model-harness configurations. Curated Skills raise the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain), with configuration-level gains ranging from +4.1 to +25.7 pp. Focused Skills with at most three modules outperform larger or exhaustive bundles, and smaller models with Skills can match larger models without them. SkillsBench establishes paired evaluation as the foundation for rigorous measurement of Skill efficacy on agentic, expertise-heavy work.

cs.AI

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .

cs.SE

Black hole thermodynamics is around the corner

We propose to work on the Euclidean black hole solution with a corner rather than with the prevalent conical singularity. As a result, we find that the Wald formula for black hole entropy can be readily obtained for generic $F(R_{abcd})$ gravity by using both the action without the corner term and the action with the corner term due to their equivalence to the first order variation. With such an equivalence, we further make use of a special diffeomorphism to accomplish a direct derivation of the ADM Hamiltonian conjugate to the Killing vector field normal to the horizon in the Lorentz signature as a conjugate variable of the inverse temperature in the grand canonical ensemble.

hep-th

Quadratic curvature corrections to 5-dimensional Kerr-AdS black hole thermodynamics by background subtraction method

We justify the applicability of the background subtraction method to both Einstein's gravity and its higher derivative corrections in 5-dimensional asymptotically AdS spacetimes, where the corresponding higher derivative corrections to the expression for the ADM mass and angular momentum are also worked out. Then we further apply the background subtraction method to calculate the first order corrected Gibbs free energy by the quadratic curvature terms for the 5-dimensional Kerr-AdS black hole, which is in exact agreement with the previous result obtained by the holographic renormalization method. Such an agreement in turn substantiates the applicability of the background subtraction method.

hep-th

FABLE: A Localized, Targeted Adversarial Attack on Weather Forecasting Models

Deep learning-based weather forecasting (DLWF) models have recently demonstrated significant performance gains over gold-standard physics-based simulation tools. However, these models are potentially vulnerable to adversarial attacks, which raises concerns about their trustworthiness. In this paper, we investigate the feasibility and challenges of applying existing adversarial attack methods to DLWF models and propose a novel framework called FABLE (Forecast Alteration By Localized targeted advErsarial attack) to address them. FABLE performs a 3D discrete wavelet decomposition to disentangle the spatial and temporal components of the data. By regulating the magnitude of adversarial perturbations across different components, FABLE produces adversarial inputs that remain closely aligned with the original inputs while steering the DLWF models toward generating the targeted forecast outcomes. Experimental results on real-world weather datasets demonstrate the effectiveness of FABLE over baseline methods across various metrics.

cs.LG

Higher derivative corrections to Kerr-AdS black hole thermodynamics

Instead of the much more involved covariant counterterm method, we apply the well justified background subtraction method to calculate the first order corrections to Kerr-AdS black hole thermodynamics induced by the higher derivative terms up to the cubic of Riemann tensor, where the computation is further simplified by the decomposition trick for the bulk action. The validity of our results is further substantiated by examining the corrections induced by the Gauss-Bonnet term. Moreover, by comparing our results with those obtained via the ADM and Wald formulas in Lorentzian signature, we can extract some generic information about the first order corrected black hole solution induced by each higher derivative term.

hep-th

Low latency global carbon budget reveals a continuous decline of the land carbon sink during the 2023/24 El Nino event

The high growth rate of atmospheric CO2 in 2023 was found to be caused by a severe reduction of the global net land carbon sink. Here we update the global CO2 budget from January 1st to July 1st 2024, during which El Ni\~no drought conditions continued to prevail in the Tropics but ceased by March 2024. We used three dynamic global vegetation models (DGVMs), machine learning emulators of ocean models, three atmospheric inversions driven by observations from the second Orbiting Carbon Observatory (OCO-2) satellite, and near-real-time fossil CO2 emissions estimates. In a one-year period from July 2023 to July 2024 covering the El Ni\~no 2023/24 event, we found a record-high CO2 growth rate of 3.66~$\pm$~0.09 ppm~yr$^{-1}$ ($\pm$~1 standard deviation) since 1979. Yet, the CO2 growth rate anomaly obtained after removing the long term trend is 1.1 ppm~yr$^{-1}$, which is marginally smaller than the July--July growth rate anomalies of the two major previous El Ni\~no events in 1997/98 and 2015/16. The atmospheric CO2 growth rate anomaly was primarily driven by a 2.24 GtC~yr$^{-1}$ reduction in the net land sink including 0.3 GtC~yr$^{-1}$ of fire emissions, partly offset by a 0.38 GtC~yr$^{-1}$ increase in the ocean sink relative to the 2015--2022 July--July mean. The tropics accounted for 97.5\% of the land CO2 flux anomaly, led by the Amazon (50.6\%), central Africa (34\%), and Southeast Asia (8.2\%), with extra-tropical sources in South Africa and southern Brazil during April--July 2024. Our three DGVMs suggest greater tropical CO2 losses in 2023/2024 than during the two previous large El Ni\~no in 1997/98 and 2015/16, whereas inversions indicate losses more comparable to 2015/16. Overall, this update of the low latency budget highlights the impact of recent El Ni\~no droughts in explaining the high CO2 growth rate until July 2024.

physics.ao-ph

Style Quantization for Data-Efficient GAN Training

Under limited data setting, GANs often struggle to navigate and effectively exploit the input latent space. Consequently, images generated from adjacent variables in a sparse input latent space may exhibit significant discrepancies in realism, leading to suboptimal consistency regularization (CR) outcomes. To address this, we propose \textit{SQ-GAN}, a novel approach that enhances CR by introducing a style space quantization scheme. This method transforms the sparse, continuous input latent space into a compact, structured discrete proxy space, allowing each element to correspond to a specific real data point, thereby improving CR performance. Instead of direct quantization, we first map the input latent variables into a less entangled ``style'' space and apply quantization using a learnable codebook. This enables each quantized code to control distinct factors of variation. Additionally, we optimize the optimal transport distance to align the codebook codes with features extracted from the training data by a foundation model, embedding external knowledge into the codebook and establishing a semantically rich vocabulary that properly describes the training dataset. Extensive experiments demonstrate significant improvements in both discriminator robustness and generation quality with our method.

cs.CV

Background subtraction method is not only much simpler, but also as applicable as covariant counterterm method

As the criterion for the applicability of the background subtraction method, not only the finiteness condition of the resulting Hamiltonian but also the condition for the validity of the first law of black hole thermodynamics can be reduced to the form amenable to much simpler analysis at infinity by using the covariant phase space formalism. With this, we further establish that the background subtraction method is as applicable as the covariant counterterm method not only to Einstein's gravity, but also to its higher derivative corrections for black hole thermodynamics in both asymptotically flat and AdS spacetimes. In addition, our framework also provides us with the first derivation of the universal expression of the Gibbs free energy in terms of the Euclidean on-shell action beyond Einstein's gravity. Among others, our findings have a far reaching impact on the shift in methodology for the Euclidean approach to black hole thermodynamics, where the well justified background subtraction method by our wieldy criterion is supposed to be the favored choice compared to the covariant counterterm method.

hep-th

MS$^3$D: A RG Flow-Based Regularization for GAN Training with Limited Data

Generative adversarial networks (GANs) have made impressive advances in image generation, but they often require large-scale training data to avoid degradation caused by discriminator overfitting. To tackle this issue, we investigate the challenge of training GANs with limited data, and propose a novel regularization method based on the idea of renormalization group (RG) in physics.We observe that in the limited data setting, the gradient pattern that the generator obtains from the discriminator becomes more aggregated over time. In RG context, this aggregated pattern exhibits a high discrepancy from its coarse-grained versions, which implies a high-capacity and sensitive system, prone to overfitting and collapse. To address this problem, we introduce a \textbf{m}ulti-\textbf{s}cale \textbf{s}tructural \textbf{s}elf-\textbf{d}issimilarity (MS$^3$D) regularization, which constrains the gradient field to have a consistent pattern across different scales, thereby fostering a more redundant and robust system. We show that our method can effectively enhance the performance and stability of GANs under limited data scenarios, and even allow them to generate high-quality images with very few data.

cs.LG

Modelling the Self-similarity in Complex Networks Based on Coulomb's Law

Recently, self-similarity of complex networks have attracted much attention. Fractal dimension of complex network is an open issue. Hub repulsion plays an important role in fractal topologies. This paper models the repulsion among the nodes in the complex networks in calculation of the fractal dimension of the networks. The Coulomb's law is adopted to represent the repulse between two nodes of the network quantitatively. A new method to calculate the fractal dimension of complex networks is proposed. The Sierpinski triangle network and some real complex networks are investigated. The results are illustrated to show that the new model of self-similarity of complex networks is reasonable and efficient.

cs.SI