SearcharxivSearch

arXiv subjects

Javier Gonzalez

Publications and source records attributed to Javier Gonzalez.

At least 19 recordsLinked to original sources

Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency

Despite their strong performance on reasoning benchmarks, large language models (LLMs) have proven brittle when presented with counterfactual questions, suggesting weaknesses in their causal reasoning ability. While recent work has demonstrated that labeled counterfactual tasks can be useful benchmarks of LLMs' causal reasoning, producing such data at the scale required to cover the vast potential space of counterfactuals is limited. In this work, we introduce double counterfactual consistency (DCC), a lightweight inference-time method for measuring and guiding the ability of LLMs to reason causally. Without requiring labeled counterfactual data, DCC verifies a model's ability to execute two important elements of causal reasoning: causal intervention and counterfactual prediction. Using DCC, we evaluate the causal reasoning abilities of various leading LLMs across a range of reasoning tasks and interventions. Moreover, we demonstrate the effectiveness of DCC as a training-free test-time rejection sampling criterion and show that it can directly improve performance on reasoning tasks across multiple model families.

cs.LG

Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts

Sentiment analysis in low-resource, culturally nuanced contexts challenges conventional NLP approaches that assume fixed labels and universal affective expressions. We present a diagnostic framework that treats sentiment as a context-dependent, culturally embedded construct, and evaluate how large language models (LLMs) reason about sentiment in informal, code-mixed WhatsApp messages from Nairobi youth health groups. Using a combination of human-annotated data, sentiment-flipped counterfactuals, and rubric-based explanation evaluation, we probe LLM interpretability, robustness, and alignment with human reasoning. Framing our evaluation through a social-science measurement lens, we operationalize and interrogate LLMs outputs as an instrument for measuring the abstract concept of sentiment. Our findings reveal significant variation in model reasoning quality, with top-tier LLMs demonstrating interpretive stability, while open models often falter under ambiguity or sentiment shifts. This work highlights the need for culturally sensitive, reasoning-aware AI evaluation in complex, real-world communication.

cs.CL

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation

Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true reasoning or from statistical recall of the training set. Inspired by the ladder of causation (Pearl, 2009) and its three levels (associations, interventions and counterfactuals), this paper introduces RE-IMAGINE, a framework to characterize a hierarchy of reasoning ability in LLMs, alongside an automated pipeline to generate problem variations at different levels of the hierarchy. By altering problems in an intermediate symbolic representation, RE-IMAGINE generates arbitrarily many problems that are not solvable using memorization alone. Moreover, the framework is general and can work across reasoning domains, including math, code, and logic. We demonstrate our framework on four widely-used benchmarks to evaluate several families of LLMs, and observe reductions in performance when the models are queried with problem variations. These assessments indicate a degree of reliance on statistical recall for past performance, and open the door to further research targeting skills across the reasoning hierarchy.

cs.CL

Reconstructing the redshift evolution of Type Ia supernovae absolute magnitude

This work investigates a potential time dependence of the absolute magnitude of Type Ia Supernovae (SN Ia). Employing the Gaussian Process approach, we obtain the SN Ia absolute magnitude and its derivative as a function of redshift. The data set considered in the analysis comprises measurements of apparent magnitude from SN Ia, Hubble rate from cosmic chronometers, and the ratio between angular and radial distances from Large-Scale Structure data (BAO and voids). Our findings reveal good compatibility between the reconstructed SN Ia absolute magnitudes and a constant value. However, the mean value obtained from the Gaussian Process reconstruction is $M=-19.456\pm 0.059$, which is $3.2\sigma$ apart from local measurements by Pantheon+SH0ES. This incompatibility may be directly associated to the $\Lambda$CDM model and local data, as it does not appear in either model-dependent or model-independent estimates of the absolute magnitude based on early universe data. Furthermore, we assess the implications of a variable $M$ within the context of modified gravity theories. Considering the local estimate of the absolute magnitude, we find $\sim3\sigma$ tension supporting departures from General Relativity in analyzing scenarios involving modified gravity theories with variations in Planck mass through Newton's constant.

astro-ph.CO

Compositional Causal Reasoning Evaluation in Language Models

Causal reasoning and compositional reasoning are two core aspirations in AI. Measuring the extent of these behaviors requires principled evaluation methods. We explore a unified perspective that considers both behaviors simultaneously, termed compositional causal reasoning (CCR): the ability to infer how causal measures compose and, equivalently, how causal quantities propagate through graphs. We instantiate a framework for the systematic evaluation of CCR for the average treatment effect and the probability of necessity and sufficiency. As proof of concept, we demonstrate CCR evaluation for language models in the LLama, Phi, and GPT families. On a math word problem, our framework revealed a range of taxonomically distinct error patterns. CCR errors increased with the complexity of causal paths for all models except o1.

cs.CL

Towards Efficient Flash Caches with Emerging NVMe Flexible Data Placement SSDs

NVMe Flash-based SSDs are widely deployed in data centers to cache working sets of large-scale web services. As data centers face increasing sustainability demands, such as reduced carbon emissions, efficient management of Flash overprovisioning and endurance has become crucial. Our analysis demonstrates that mixing data with different lifetimes on Flash blocks results in high device garbage collection costs, which either reduce device lifetime or necessitate host overprovisioning. Targeted data placement on Flash to minimize data intermixing and thus device write amplification shows promise for addressing this issue. The NVMe Flexible Data Placement (FDP) proposal is a newly ratified technical proposal aimed at addressing data placement needs while reducing the software engineering costs associated with past storage interfaces, such as ZNS and Open-Channel SSDs. In this study, we explore the feasibility, benefits, and limitations of leveraging NVMe FDP primitives for data placement on Flash media in CacheLib, a popular open-source Flash cache widely deployed and used in Meta's software ecosystem as a caching building block. We demonstrate that targeted data placement in CacheLib using NVMe FDP SSDs helps reduce device write amplification, embodied carbon emissions, and power consumption with almost no overhead to other metrics. Using multiple production traces and their configurations from Meta and Twitter, we show that an ideal device write amplification of ~1 can be achieved with FDP, leading to improved SSD utilization and sustainable Flash cache deployments.

cs.AR

JoLT: Joint Probabilistic Predictions on Tabular Data Using LLMs

We introduce a simple method for probabilistic predictions on tabular data based on Large Language Models (LLMs) called JoLT (Joint LLM Process for Tabular data). JoLT uses the in-context learning capabilities of LLMs to define joint distributions over tabular data conditioned on user-specified side information about the problem, exploiting the vast repository of latent problem-relevant knowledge encoded in LLMs. JoLT defines joint distributions for multiple target variables with potentially heterogeneous data types without any data conversion, data preprocessing, special handling of missing data, or model training, making it accessible and efficient for practitioners. Our experiments show that JoLT outperforms competitive methods on low-shot single-target and multi-target tabular classification and regression tasks. Furthermore, we show that JoLT can automatically handle missing data and perform data imputation by leveraging textual side information. We argue that due to its simplicity and generality, JoLT is an effective approach for a wide variety of real prediction problems.

stat.ML

Assessing the dark degeneracy through the gas mass fraction data

It is well-known that Einstein's equations constrain only the total energy-momentum tensor of the cosmic substratum, without specifying the characteristics of its individual constituents. Consequently, cosmological models featuring distinct decompositions within the dark sector, while sharing identical values for the sum of dark components' energy-momentum tensor, remain indistinguishable when assessed through observables based on distance measurements. Notably, it has been already demonstrated that cosmological models with dynamical descriptions of dark energy, characterized by a time-dependent equation of state (EoS), can always be mapped into a model featuring a decaying vacuum ($w=-1$) coupled with dark matter. We explore the possibility of breaking this degeneracy by using measurements of the gas mass fraction observed in massive and relaxed galaxy clusters. This data is particularly interesting for this purpose because it isolates the matter contribution, possibly allowing the degeneracy breaking. We study the particular case of the $w$CDM model with its interactive counterpart. We compare the results obtained from both descriptions with a non-parametric analysis obtained through Gaussian Process. Even though the degeneracy may be broken from the theoretical point of view, we find that current gas mass fraction data seems to be insufficient for a final conclusion about which approach is favored, even when combined with SNIa, BAO and CMB.

astro-ph.CO

AIRIVA: A Deep Generative Model of Adaptive Immune Repertoires

Recent advances in immunomics have shown that T-cell receptor (TCR) signatures can accurately predict active or recent infection by leveraging the high specificity of TCR binding to disease antigens. However, the extreme diversity of the adaptive immune repertoire presents challenges in reliably identifying disease-specific TCRs. Population genetics and sequencing depth can also have strong systematic effects on repertoires, which requires careful consideration when developing diagnostic models. We present an Adaptive Immune Repertoire-Invariant Variational Autoencoder (AIRIVA), a generative model that learns a low-dimensional, interpretable, and compositional representation of TCR repertoires to disentangle such systematic effects in repertoires. We apply AIRIVA to two infectious disease case-studies: COVID-19 (natural infection and vaccination) and the Herpes Simplex Virus (HSV-1 and HSV-2), and empirically show that we can disentangle the individual disease signals. We further demonstrate AIRIVA's capability to: learn from unlabelled samples; generate in-silico TCR repertoires by intervening on the latent factors; and identify disease-associated TCRs validated using TCR annotations from external assay data.

q-bio.QM

Astrometric Apparent Motion of High-redshift Radio Sources

Radio-loud quasars at high redshift (z > 4) are rare objects in the Universe and rarely observed with Very Long Baseline Interferometry (VLBI). But some of them have flux density sufficiently high for monitoring of their apparent position. The instability of the astrometric positions could be linked to the astrophysical process in the jetted active galactic nuclei in the early Universe. Regular observations of the high-redshift quasars are used for estimating their apparent proper motion over several years. We have undertaken regular VLBI observations of several high-redshift quasars at 2.3 GHz (S band) and 8.4 GHz (X band) with a network of five radio telescopes: 40-m Yebes (Spain), 25-m Sheshan (China), and three 32-m telescopes of the Quasar VLBI Network (Russia) -- Svetloe, Zelenchukskaya, and Badary. Additional facilities joined this network occasionally. The sources have also been observed in three sessions with the European VLBI Network (EVN) in 2018--2019 and one Long Baseline Array (LBA) experiment in 2018. In addition, several experiments conducted with the Very Long Baseline Array (VLBA) in 2017--2018were used to improve the time sampling and the statistics. Based on these 37 astrometric VLBI experiments between 2017 and 2021, we estimated the apparent proper motions of four quasars: 0901+697, 1428+422, 1508+572, and 2101+600.

astro-ph.GA

Detecting the oscillation and propagation of the nascent dynamic solar wind structure at 2.6 solar radii using VLBI radio telescopes

Probing the solar corona is crucial to study the coronal heating and solar wind acceleration. However, the transient and inhomogeneous solar wind flows carry large-amplitude inherent Alfven waves and turbulence, which make detection more difficult. We report the oscillation and propagation of the solar wind at 2.6 solar radii (Rs) by observation of China Tianwen and ESA Mars Express with radio telescopes. The observations were carried out on Oct.9 2021, when one coronal mass ejection (CME) passed across the ray paths of the telescope beams. We obtain the frequency fluctuations (FF) of the spacecraft signals from each individual telescope. Firstly, we visually identify the drift of the frequency spikes at a high spatial resolution of thousands of kilometers along the projected baselines. They are used as traces to estimate the solar wind velocity. Then we perform the cross-correlation analysis on the time series of FF from different telescopes. The velocity variations of solar wind structure along radial and tangential directions during the CME passage are obtained. The oscillation of tangential velocity confirms the detection of streamer wave. Moreover, at the tail of the CME, we detect the propagation of an accelerating fast field-aligned density structure indicating the presence of magnetohydrodynamic waves. This study confirm that the ground station-pairs are able to form particular spatial projection baselines with high resolution and sensitivity to study the detailed propagation of the nascent dynamic solar wind structure.

astro-ph.SR

RKHS-SHAP: Shapley Values for Kernel Methods

Feature attribution for kernel methods is often heuristic and not individualised for each prediction. To address this, we turn to the concept of Shapley values~(SV), a coalition game theoretical framework that has previously been applied to different machine learning model interpretation tasks, such as linear models, tree ensembles and deep networks. By analysing SVs from a functional perspective, we propose \textsc{RKHS-SHAP}, an attribution method for kernel machines that can efficiently compute both \emph{Interventional} and \emph{Observational Shapley values} using kernel mean embeddings of distributions. We show theoretically that our method is robust with respect to local perturbations - a key yet often overlooked desideratum for consistent model interpretation. Further, we propose \emph{Shapley regulariser}, applicable to a general empirical risk minimisation framework, allowing learning while controlling the level of specific feature's contributions to the model. We demonstrate that the Shapley regulariser enables learning which is robust to covariate shift of a given feature and fair learning which controls the SVs of sensitive features.

stat.ML

Invariant Priors for Bayesian Quadrature

Bayesian quadrature (BQ) is a model-based numerical integration method that is able to increase sample efficiency by encoding and leveraging known structure of the integration task at hand. In this paper, we explore priors that encode invariance of the integrand under a set of bijective transformations in the input domain, in particular some unitary transformations, such as rotations, axis-flips, or point symmetries. We show initial results on superior performance in comparison to standard Bayesian quadrature on several synthetic and one real world application.

stat.ML

GIBBON: General-purpose Information-Based Bayesian OptimisatioN

This paper describes a general-purpose extension of max-value entropy search, a popular approach for Bayesian Optimisation (BO). A novel approximation is proposed for the information gain -- an information-theoretic quantity central to solving a range of BO problems, including noisy, multi-fidelity and batch optimisations across both continuous and highly-structured discrete spaces. Previously, these problems have been tackled separately within information-theoretic BO, each requiring a different sophisticated approximation scheme, except for batch BO, for which no computationally-lightweight information-theoretic approach has previously been proposed. GIBBON (General-purpose Information-Based Bayesian OptimisatioN) provides a single principled framework suitable for all the above, out-performing existing approaches whilst incurring substantially lower computational overheads. In addition, GIBBON does not require the problem's search space to be Euclidean and so is the first high-performance yet computationally light-weight acquisition function that supports batch BO over general highly structured input spaces like molecular search and gene design. Moreover, our principled derivation of GIBBON yields a natural interpretation of a popular batch BO heuristic based on determinantal point processes. Finally, we analyse GIBBON across a suite of synthetic benchmark tasks, a molecular search loop, and as part of a challenging batch multi-fidelity framework for problems with controllable experimental noise.

cs.LG

Emulation of physical processes with Emukit

Decision making in uncertain scenarios is an ubiquitous challenge in real world systems. Tools to deal with this challenge include simulations to gather information and statistical emulation to quantify uncertainty. The machine learning community has developed a number of methods to facilitate decision making, but so far they are scattered in multiple different toolkits, and generally rely on a fixed backend. In this paper, we present Emukit, a highly adaptable Python toolkit for enriching decision making under uncertainty. Emukit allows users to: (i) use state of the art methods including Bayesian optimization, multi-fidelity emulation, experimental design, Bayesian quadrature and sensitivity analysis; (ii) easily prototype new decision making methods for new problems. Emukit is agnostic to the underlying modeling framework and enables users to use their own custom models. We show how Emukit can be used on three exemplary case studies.

cs.LG

Preferential Batch Bayesian Optimization

Most research in Bayesian optimization (BO) has focused on \emph{direct feedback} scenarios, where one has access to exact values of some expensive-to-evaluate objective. This direction has been mainly driven by the use of BO in machine learning hyper-parameter configuration problems. However, in domains such as modelling human preferences, A/B tests, or recommender systems, there is a need for methods that can replace direct feedback with \emph{preferential feedback}, obtained via rankings or pairwise comparisons. In this work, we present preferential batch Bayesian optimization (PBBO), a new framework that allows finding the optimum of a latent function of interest, given any type of parallel preferential feedback for a group of two or more points. We do so by using a Gaussian process model with a likelihood specially designed to enable parallel and efficient data collection mechanisms, which are key in modern machine learning. We show how the acquisitions developed under this framework generalize and augment previous approaches in Bayesian optimization, expanding the use of these techniques to a wider range of domains. An extensive simulation study shows the benefits of this approach, both with simulated functions and four real data sets.

cs.LG

Active Multi-Information Source Bayesian Quadrature

Bayesian quadrature (BQ) is a sample-efficient probabilistic numerical method to solve integrals of expensive-to-evaluate black-box functions, yet so far,active BQ learning schemes focus merely on the integrand itself as information source, and do not allow for information transfer from cheaper, related functions. Here, we set the scene for active learning in BQ when multiple related information sources of variable cost (in input and source) are accessible. This setting arises for example when evaluating the integrand requires a complex simulation to be run that can be approximated by simulating at lower levels of sophistication and at lesser expense. We construct meaningful cost-sensitive multi-source acquisition rates as an extension to common utility functions from vanilla BQ (VBQ),and discuss pitfalls that arise from blindly generalizing. Furthermore, we show that the VBQ acquisition policy is a corner-case of all considered cost-sensitive acquisition schemes, which collapse onto one single de-generate policy in the case of one source and constant cost. In proof-of-concept experiments we scrutinize the behavior of our generalized acquisition functions. On an epidemiological model, we demonstrate that active multi-source BQ (AMS-BQ) allocates budget more efficiently than VBQ for learning the integral to a good accuracy.

cs.LG

Good practices for Bayesian Optimization of high dimensional structured spaces

The increasing availability of structured but high dimensional data has opened new opportunities for optimization. One emerging and promising avenue is the exploration of unsupervised methods for projecting structured high dimensional data into low dimensional continuous representations, simplifying the optimization problem and enabling the application of traditional optimization methods. However, this line of research has been purely methodological with little connection to the needs of practitioners so far. In this paper, we study the effect of different search space design choices for performing Bayesian Optimization in high dimensional structured datasets. In particular, we analyse the influence of the dimensionality of the latent space, the role of the acquisition function and evaluate new methods to automatically define the optimization bounds in the latent space. Finally, based on experimental results using synthetic and real datasets, we provide recommendations for the practitioners.

cs.LG