SearcharxivSearch

arXiv subjects

Helen Lu

Publications and source records attributed to Helen Lu.

6 recordsLinked to original sources

Benchmarking the Generality of Vision-Language-Action Models

Generalist multimodal agents are expected to unify perception, language, and control - operating robustly across diverse real world domains. However, current evaluation practices remain fragmented across isolated benchmarks, making it difficult to assess whether today's foundation models truly generalize beyond their training distributions. We introduce MultiNet v1.0, a unified benchmark for measuring the cross domain generality of vision language models (VLMs) and vision language action models (VLAs) across six foundational capability regimes. Visual grounding, spatial reasoning, tool use, physical commonsense, multi agent coordination, and continuous robot control. Evaluating GPT 5, Pi0, and Magma, we find that no model demonstrates consistent generality. All exhibit substantial degradation on unseen domains, unfamiliar modalities, or cross domain task shifts despite strong performance within their training distributions.These failures manifest as modality misalignment, output format instability, and catastrophic knowledge degradation under domain transfer.Our findings reveal a persistent gap between the aspiration of generalist intelligence and the actual capabilities of current foundation models.MultiNet v1.0 provides a standardized evaluation substrate for diagnosing these gaps and guiding the development of future generalist agents.Code, data, and leaderboards are publicly available.

cs.LG

Test-time augmentation improves efficiency in conformal prediction

A conformal classifier produces a set of predicted classes and provides a probabilistic guarantee that the set includes the true class. Unfortunately, it is often the case that conformal classifiers produce uninformatively large sets. In this work, we show that test-time augmentation (TTA)--a technique that introduces inductive biases during inference--reduces the size of the sets produced by conformal classifiers. Our approach is flexible, computationally efficient, and effective. It can be combined with any conformal score, requires no model retraining, and reduces prediction set sizes by 10%-14% on average. We conduct an evaluation of the approach spanning three datasets, three models, two established conformal scoring methods, different guarantee strengths, and several distribution shifts to show when and why test-time augmentation is a useful addition to the conformal pipeline.

cs.LG

AXIL: Exact Instance Attribution for Gradient Boosting

We derive an exact, prediction-specific instance-attribution method for fitted gradient boosting machines (GBMs) trained with squared-error loss, with the learned tree structure held fixed. Each prediction can be written as a weighted sum of training targets, with coefficients determined only by the fitted tree structure and learning rate. These coefficients are exact instance attributions, or AXIL weights. Our main algorithmic contribution is a matrix-free backward operator that computes one AXIL attribution vector in O(TN) time, or S vectors in O(TNS), without materialising the full N x N matrix. This extends to out-of-sample predictions and makes exact instance attribution practical for large datasets. AXIL yields exact fixed-structure sensitivity by construction in target-perturbation tests, where competing GBM-specific attribution methods (BoostIn, TREX, and LeafInfluence) generally fail. In retraining-based faithfulness tests on 20 regression datasets, AXIL achieves the highest faithfulness score on 14 datasets and statistically ties for the best on 4 others, while also running substantially faster than the competing methods. We also show that the AXIL weight matrix is the globally constant special case of a target-response Jacobian that provides first-order instance attribution for any differentiable learner via implicit differentiation, placing the exact decomposition inside a broader framework.

cs.LG

Improved Text Classification via Test-Time Augmentation

Test-time augmentation -- the aggregation of predictions across transformed examples of test inputs -- is an established technique to improve the performance of image classification models. Importantly, TTA can be used to improve model performance post-hoc, without additional training. Although test-time augmentation (TTA) can be applied to any data modality, it has seen limited adoption in NLP due in part to the difficulty of identifying label-preserving transformations. In this paper, we present augmentation policies that yield significant accuracy improvements with language models. A key finding is that augmentation policy design -- for instance, the number of samples generated from a single, non-deterministic augmentation -- has a considerable impact on the benefit of TTA. Experiments across a binary classification task and dataset show that test-time augmentation can deliver consistent improvements over current state-of-the-art approaches.

cs.LG

Unusual high-field metal in a Kondo insulator

Within condensed-matter systems, strong electronic interactions often lead to exotic quantum phases. A recent manifestation of this is the unexpected observation of magnetic quantum oscillations and metallic thermal transport, both properties of systems with Fermi surfaces of itinerant quasiparticles, in the Kondo insulators SmB6 and YbB$_{12}$. To understand these phenomena, it is informative to study their evolution as the energy gap of the Kondo-Insulator state is closed by a large magnetic field. We show here that both the quantum-oscillation frequency and the cyclotron mass display a strong field dependence in the resulting high-field metallic state in $_{12}$. By tracking the Fermi-surface area, we conclude that the same quasiparticle band gives rise to the quantum oscillations in both insulating and metallic states. These data are understood most simply using a two-fluid picture where unusual quasiparticles, contributing little or nothing to charge transport, coexist with conventional fermions. In the metallic state this leads to a heavy-fermion bad metal with negligible magnetoresistance, relatively high resistivity and a very large Kadowaki-Woods ratio, underlining the exotic nature of the fermion ensemble inhabiting $_{12}$.

cond-mat.str-el

Combining micro- and macroscopic probes to untangle single-ion and spatial exchange anisotropies in a $S = 1$ quantum antiferromagnet

The magnetic ground state of the quasi-one-dimensional spin-1 antiferromagnetic chain is sensitive to the relative sizes of the single-ion anisotropy ($D$) and the intrachain ($J$) and interchain ($J'$) exchange interactions. The ratios $D/J$ and $J'/J$ dictate the material's placement in one or other of three competing phases: a Haldane gapped phase, a quantum paramagnet and an XY-ordered state, with a quantum critical point at their junction. We have identified [Ni(HF)$_2$(pyz)$_2]$SbF$_6$, where pyz = pyrazine, as a candidate in which this behavior can be explored in detail. Combining neutron scattering (elastic and inelastic) in applied magnetic fields of up to 10~tesla and magnetization measurements in fields of up to 60~tesla with numerical modeling of experimental observables, we are able to obtain accurate values of all of the parameters of the Hamiltonian [$D = 13.3(1)$~K, $J = 10.4(3)$~K and $J' = 1.4(2)$~K], despite the polycrystalline nature of the sample. Density-functional theory calculations result in similar couplings ($J = 9.2$~K, $J' = 1.8$~K) and predict that the majority of the total spin population of resides on the Ni(II) ion, while the remaining spin density is delocalized over both ligand types. The general procedures outlined in this paper permit phase boundaries and quantum-critical points to be explored in anisotropic systems for which single crystals are as yet unavailable.

cond-mat.mes-hall