Searcharxiv⌕ Search

arXiv subjects

Giovanni Gatti Pinheiro

Publications and source records attributed to Giovanni Gatti Pinheiro.

2 recordsLinked to original sources

Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuality-based metrics, their sensitivity and reliability remain underexplored. This paper introduces a meta-evaluation framework that systematically tests these metrics using controlled corruptions of gold standard answers. Our method generates ranked outputs with known degrees of degradation to probe how metrics capture nuanced changes in truthfulness. Our experiments reveal that pipeline-based methods, such as the RAGAS's factual correctness metric, better track degradation than LLM-as-judge approaches. We also propose a new variant of the factual correctness metric that provides a competitive and cost-efficient.

cs.CL↗

Optimizing Revenue Maximization and Demand Learning in Airline Revenue Management

Correctly estimating how demand respond to prices is fundamental for airlines willing to optimize their pricing policy. Under some conditions, these policies, while aiming at maximizing short term revenue, can present too little price variation which may decrease the overall quality of future demand forecasting. This problem, known as earning while learning problem, is not exclusive to airlines, and it has been investigated by academia and industry in recent years. One of the most promising methods presented in literature combines the revenue maximization and the demand model quality into one single objective function. This method has shown great success in simulation studies and real life benchmarks. Nevertheless, this work needs to be adapted to certain constraints that arise in the airline revenue management (RM), such as the need to control the prices of several active flights of a leg simultaneously. In this paper, we adjust this method to airline RM while assuming unconstrained capacity. Then, we show that our new algorithm efficiently performs price experimentation in order to generate more revenue over long horizons than classical methods that seek to maximize revenue only.

cs.LG↗