SearcharxivSearch

arXiv subjects

Guglielmo Faggioli

Publications and source records attributed to Guglielmo Faggioli.

16 recordsLinked to original sources

Analysis of late-time tails in spin-aligned eccentric binary black hole mergers

We present a comprehensive analysis of late-time tails in gravitational radiation from merging spin-aligned eccentric binary black holes, using high-accuracy point-particle black hole perturbation theory simulations. We simulate the late-time evolution of 15 binary black hole mergers with mass ratio $q = 1000$, dimensionless spins $χ= [-0.9, -0.6, 0.0, 0.6, 0.9]$ and eccentricity at the last stable orbit $e_{\rm LSO} = [0.8, 0.9, 0.95]$. We track the tail amplitudes and exponents up to a retarded time coordinate $t = 9000M$ after merger for the six spin-weighted spherical harmonic modes $(2,1)$, $(2,2)$, $(3,2)$, $(3,3)$, $(4,3)$, and $(4,4)$ employing both frequentist and Bayesian approaches. We note that the tails are increasingly pronounced for binaries with high eccentricity $e_{\rm LSO}$ and large negative spin $χ$. We find that the overall late-time exponents closely approach their predicted asymptotic values ($p=-\ell-4$ for Weyl curvature scalar $ψ_{4,\ell m}$ where $\ell$ is the spin-weighted spherical harmonic index), while estimates restricted to the latest portion of the data exactly recover them. We further verify numerically that modes with the same spherical index $\ell$ share identical tail exponents, while variations in $m$ do not affect the tail behavior. Our analysis framework is publicly available through the gwtails Python package.

gr-qc

Modeling the merger-ringdown of an eccentric test-mass inspiral into a Kerr black hole using the effective-one-body framework

We characterize and phenomenologically model the merger-ringdown of gravitational waves emitted by a small compact object that plunges and merges into a Kerr black hole from equatorial-eccentric inspirals. The waveforms are generated employing a time-domain Teukolsky code sourced with trajectories computed using the effective-one-body framework. We span values of the Kerr spin $a\in[-0.9, 0.9] $, eccentricity at the last stable orbit (LSO) $ e_{\rm LSO} \in [0,0.9] $, and relativistic anomaly $ ξ_{\rm LSO} \in [0 , 2 π]$. We characterize the last peak of the waveform and ringdown features across the parameter space, finding that the eccentricity mainly affects the last peak features, while it has a smaller impact on the ringdown signal. In contrast, the relativistic anomaly measured at the LSO influences the morphology of the last peak in a restricted portion of the parameter space and has no impact on the ringdown part. We perform the analysis for all the spin-weighted spherical harmonic modes normally included in the $\texttt{SEOBNR}$ family of models, $(\ell,m)\in\{ (2,2), (3,3), (4,4), (5,5), (2,1), (3,2), (4,3)\}$. Finally, we introduce a merger-ringdown model for $\texttt{SEOB-TMLE}$, a forthcoming inspiral-merger-ringdown waveform model for eccentric spin-aligned binary black holes in the test-mass limit, whose features can be extended to comparable-mass regimes. The model also accounts for quasinormal mode mixing during the ringdown. It provides a first step toward incorporating the impact of residual eccentricity close to merger into spin-aligned effective-one-body merger-ringdown models for binary black holes.

gr-qc

Advancing the Effective-One-Body Framework in the Test-Mass Limit

We present SEOB-TML, an enhanced effective-one-body (EOB) framework for the test-mass limit, optimized for quasi-circular, spin-aligned binary black holes. On the dynamical side, we introduce a quadrupole-factorized (Q-factorized) prescription that maps the total energy flux-including horizon absorption-onto a single (2,2) mode baseline. This approach effectively captures higher-order multipole contributions without explicit mode summation, while simultaneously leading to a dramatic reduction in fractional flux errors. To ensure a smooth transition to the post-merger stage, we replace traditional next-to-quasicircular corrections with a phenomenological ansatz, enabling a flexible, mode-dependent attachment prescription. For the merger-ringdown stage, we utilize quasi-normal mode coefficients extracted from numerical waveforms via qnmfinder to explicitly model mode-mixing effects. These enhancements lead to a substantial reduction in residuals, capturing the complex physical modulations prominent in retrograde configurations. Additionally, we implement the (2,0) mode across the full waveform, further extending the model's physical coverage and accuracy. Overall, our framework generates highly accurate late inspiral-merger-ringdown waveforms for extreme-mass-ratio systems, significantly reducing dephasing and improving the near-merger reconstruction. We demonstrate the performance of SEOB-TML against the current state-of-the-art SEOBNRv5HM model, highlighting how our specialized developments extend the reliability of the EOB framework into the test-mass limit.

gr-qc

Phenomenology and origin of late-time tails in eccentric binary black hole mergers

We investigate the late-time tail behavior in gravitational waves from merging eccentric binary black holes (BBH) using black hole perturbation theory. For simplicity, we focus only on the dominant quadrupolar mode of the radiation. We demonstrate that such tails become more prominent as eccentricity increases. Exploring the phenomenology of the tails in both spinning and non-spinning eccentric binaries, with the spin magnitude varying from $χ=-0.6$ to $χ=+0.6$ and eccentricity as high as $e=0.98$, we find that these tails can be well approximated by a slowly decaying power law. We study the power law for varying systems and find that the power law exponent lies close to the theoretically expected value $-4$. Finally, using both plunge geodesic and radiation-reaction-driven orbits, we perform a series of numerical experiments to understand the origin of the tails in BBH simulations. Our results suggest that the late-time tails are strongly excited in eccentric BBH systems when the smaller black hole is in the neighborhood of the apocenter, as opposed to any structure in the strong field of the larger black hole. Our analysis framework is publicly available through the \texttt{gwtails} Python package.

gr-qc

Dagstuhl Perspectives Workshop 24352 -- Conversational Agents: A Framework for Evaluation (CAFE): Manifesto

During the workshop, we deeply discussed what CONversational Information ACcess (CONIAC) is and its unique features, proposing a world model abstracting it, and defined the Conversational Agents Framework for Evaluation (CAFE) for the evaluation of CONIAC systems, consisting of six major components: 1) goals of the system's stakeholders, 2) user tasks to be studied in the evaluation, 3) aspects of the users carrying out the tasks, 4) evaluation criteria to be considered, 5) evaluation methodology to be applied, and 6) measures for the quantitative criteria chosen.

cs.CL

Peaking into the abyss: Characterizing the merger of equatorial-eccentric-geodesic plunges in rotating black holes

We study the gravitational waveforms generated by critical, equatorial plunging geodesics of the Kerr metric that start from an unstable-circular-orbit, which describe the test-mass limit of spin-aligned eccentric black-hole mergers. The waveforms are generated employing a time-domain Teukolsky code. We span different values of the Kerr spin $-0.99 \le a \le 0.99 $ and of the critical eccentricity $e_c$, for bound ($0 \le e_c<1$) and unbound plunges ($e_c \ge 1$). We find that, contrary to expectations, the waveform modes $h_{\ell m}$ do not always manifest a peak for high eccentricities or spins. In case of the dominant $h_{22}$ mode, we determine the precise region of the parameter space in which its peak exists. In this region, we provide a characterization of the merger quantities of the $h_{22}$ mode and of the higher-order modes, providing the merger structure of the equatorial eccentric plunges of the Kerr spacetime in the test-mass limit.

gr-qc

Variations in Relevance Judgments and the Shelf Life of Test Collections

The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditional test collections. However, the paradigm shift towards neural retrieval models affected the characteristics of modern test collections, e.g., documents are short, judged with four grades of relevance, and information needs have no descriptions or narratives. Under these changes, it is unclear whether assessor disagreement remains negligible for system comparisons. We investigate this aspect under the additional condition that the few modern test collections are heavily re-used. Given more possible query interpretations due to less formalized information needs, an ``expiration date'' for test collections might be needed if top-effectiveness requires overfitting to a single interpretation of relevance. We run a reproducibility study and re-annotate the relevance judgments of the 2019~TREC Deep Learning track. We can reproduce prior work in the neural retrieval setting, showing that assessor disagreement does not affect system rankings. However, we observe that some models substantially degrade with our new relevance judgments, and some have already reached the effectiveness of humans as rankers, providing evidence that test collections can expire.

cs.IR

Testing eccentric corrections to the radiation-reaction force in the test-mass limit of effective-one-body models

In this work, we test an effective-one-body radiation-reaction force for eccentric planar orbits of a test mass in a Kerr background, which contains third-order post-Newtonian (PN) non-spinning and second-order PN spin contributions. We compare the analytical fluxes connected to two different resummations of this force, truncated at different PN orders in the eccentric sector, with the numerical fluxes computed through the use of frequency- and time-domain Teukolsky-equation codes. We find that the different PN truncations of the radiation-reaction force show the expected scaling in the weak gravitational-field regime, and we observe a fractional difference with the numerical fluxes that is $<5 \%$, for orbits characterized by eccentricity $0 \le e \le 0.7$, central black-hole spin $-0.99 M \le a \le 0.99 M$ and fixed orbital-averaged quantity $x=\langle MΩ\rangle^{2/3} = 0.06$, corresponding to the mildly strong-field regime with semilatera recta $9 M<p<17 M$. Our analysis provides useful information for the development of spin-aligned eccentric models in the comparable-mass case.

gr-qc

Judging the Judges: A Collection of LLM-Generated Relevance Judgements

Using Large Language Models (LLMs) for relevance assessments offers promising opportunities to improve Information Retrieval (IR), Natural Language Processing (NLP), and related fields. Indeed, LLMs hold the promise of allowing IR experimenters to build evaluation collections with a fraction of the manual human labor currently required. This could help with fresh topics on which there is still limited knowledge and could mitigate the challenges of evaluating ranking systems in low-resource scenarios, where it is challenging to find human annotators. Given the fast-paced recent developments in the domain, many questions concerning LLMs as assessors are yet to be answered. Among the aspects that require further investigation, we can list the impact of various components in a relevance judgment generation pipeline, such as the prompt used or the LLM chosen. This paper benchmarks and reports on the results of a large-scale automatic relevance judgment evaluation, the LLMJudge challenge at SIGIR 2024, where different relevance assessment approaches were proposed. In detail, we release and benchmark 42 LLM-generated labels of the TREC 2023 Deep Learning track relevance judgments produced by eight international teams who participated in the challenge. Given their diverse nature, these automatically generated relevance judgments can help the community not only investigate systematic biases caused by LLMs but also explore the effectiveness of ensemble models, analyze the trade-offs between different models and human assessors, and advance methodologies for improving automated evaluation techniques. The released resource is available at the following link: https://llm4eval.github.io/LLMJudge-benchmark/

cs.IR

Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024

The first edition of the workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) took place in July 2024, co-located with the ACM SIGIR Conference 2024 in the USA (SIGIR 2024). The aim was to bring information retrieval researchers together around the topic of LLMs for evaluation in information retrieval that gathered attention with the advancement of large language models and generative AI. Given the novelty of the topic, the workshop was focused around multi-sided discussions, namely panels and poster sessions of the accepted proceedings papers.

cs.IR

LLMJudge: LLMs for Relevance Judgments

The LLMJudge challenge is organized as part of the LLM4Eval workshop at SIGIR 2024. Test collections are essential for evaluating information retrieval (IR) systems. The evaluation and tuning of a search system is largely based on relevance labels, which indicate whether a document is useful for a specific search and user. However, collecting relevance judgments on a large scale is costly and resource-intensive. Consequently, typical experiments rely on third-party labelers who may not always produce accurate annotations. The LLMJudge challenge aims to explore an alternative approach by using LLMs to generate relevance judgments. Recent studies have shown that LLMs can generate reliable relevance judgments for search systems. However, it remains unclear which LLMs can match the accuracy of human labelers, which prompts are most effective, how fine-tuned open-source LLMs compare to closed-source LLMs like GPT-4, whether there are biases in synthetically generated data, and if data leakage affects the quality of generated labels. This challenge will investigate these questions, and the collected data will be released as a package to support automatic relevance judgment research in information retrieval and search.

cs.IR

Words Blending Boxes. Obfuscating Queries in Information Retrieval using Differential Privacy

Ensuring the effectiveness of search queries while protecting user privacy remains an open issue. When an Information Retrieval System (IRS) does not protect the privacy of its users, sensitive information may be disclosed through the queries sent to the system. Recent improvements, especially in NLP, have shown the potential of using Differential Privacy to obfuscate texts while maintaining satisfactory effectiveness. However, such approaches may protect the user's privacy only from a theoretical perspective while, in practice, the real user's information need can still be inferred if perturbed terms are too semantically similar to the original ones. We overcome such limitations by proposing Word Blending Boxes, a novel differentially private mechanism for query obfuscation, which protects the words in the user queries by employing safe boxes. To measure the overall effectiveness of the proposed WBB mechanism, we measure the privacy obtained by the obfuscation process, i.e., the lexical and semantic similarity between original and obfuscated queries. Moreover, we assess the effectiveness of the privatized queries in retrieving relevant documents from the IRS. Our findings indicate that WBB can be integrated effectively into existing IRSs, offering a key to the challenge of protecting user privacy from both a theoretical and a practical point of view.

cs.IR

Perspectives on Large Language Models for Relevance Judgment

When asked, large language models (LLMs) like ChatGPT claim that they can assist with relevance judgments but it is not clear whether automated judgments can reliably be used in evaluations of retrieval systems. In this perspectives paper, we discuss possible ways for LLMs to support relevance judgments along with concerns and issues that arise. We devise a human--machine collaboration spectrum that allows to categorize different relevance judgment strategies, based on how much humans rely on machines. For the extreme point of "fully automated judgments", we further include a pilot experiment on whether LLM-based relevance judgments correlate with judgments from trained human assessors. We conclude the paper by providing opposing perspectives for and against the use of~LLMs for automatic relevance judgments, and a compromise perspective, informed by our analyses of the literature, our preliminary experimental evidence, and our experience as IR researchers.

cs.IR

Query Performance Prediction for Neural IR: Are We There Yet?

Evaluation in Information Retrieval relies on post-hoc empirical procedures, which are time-consuming and expensive operations. To alleviate this, Query Performance Prediction (QPP) models have been developed to estimate the performance of a system without the need for human-made relevance judgements. Such models, usually relying on lexical features from queries and corpora, have been applied to traditional sparse IR methods - with various degrees of success. With the advent of neural IR and large Pre-trained Language Models, the retrieval paradigm has significantly shifted towards more semantic signals. In this work, we study and analyze to what extent current QPP models can predict the performance of such systems. Our experiments consider seven traditional bag-of-words and seven BERT-based IR approaches, as well as nineteen state-of-the-art QPPs evaluated on two collections, Deep Learning '19 and Robust '04. Our findings show that QPPs perform statistically significantly worse on neural IR systems. In settings where semantic signals are prominent (e.g., passage retrieval), their performance on neural models drops by as much as 10% compared to bag-of-words approaches. On top of that, in lexical-oriented scenarios, QPPs fail to predict performance for neural IR systems on those queries where they differ from traditional approaches the most.

cs.IR

Towards Feature Selection for Ranking and Classification Exploiting Quantum Annealers

Feature selection is a common step in many ranking, classification, or prediction tasks and serves many purposes. By removing redundant or noisy features, the accuracy of ranking or classification can be improved and the computational cost of the subsequent learning steps can be reduced. However, feature selection can be itself a computationally expensive process. While for decades confined to theoretical algorithmic papers, quantum computing is now becoming a viable tool to tackle realistic problems, in particular special-purpose solvers based on the Quantum Annealing paradigm. This paper aims to explore the feasibility of using currently available quantum computing architectures to solve some quadratic feature selection algorithms for both ranking and classification. The experimental analysis includes 15 state-of-the-art datasets. The effectiveness obtained with quantum computing hardware is comparable to that of classical solvers, indicating that quantum computers are now reliable enough to tackle interesting problems. In terms of scalability, current generation quantum computers are able to provide a limited speedup over certain classical algorithms and hybrid quantum-classical strategies show lower computational cost for problems of more than a thousand features.

cs.IR

Towards simulating a realistic data analysis with an optimised angular power spectrum of spectroscopic galaxy surveys

The angular power spectrum is a natural tool to analyse the observed galaxy number count fluctuations. In a standard analysis, the angular galaxy distribution is sliced into concentric redshift bins and all correlations of its harmonic coefficients between bin pairs are considered---a procedure referred to as `tomography'. However, the unparalleled quality of data from oncoming spectroscopic galaxy surveys for cosmology will render this method computationally unfeasible, given the increasing number of bins. Here, we put to test against synthetic data a novel method proposed in a previous study to save computational time. According to this method, the whole galaxy redshift distribution is subdivided into thick bins, neglecting the cross-bin correlations among them; each of the thick bin is, however, further subdivided into thinner bins, considering in this case all the cross-bin correlations. We create a simulated data set that we then analyse in a Bayesian framework. We confirm that the newly proposed method saves computational time and gives results that surpass those of the standard approach.

astro-ph.CO