SearcharxivSearch

arXiv subjects

Tiago Teixeira

Publications and source records attributed to Tiago Teixeira.

3 recordsLinked to original sources

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and benchmark datasets. However, most mathematical reasoning evaluations exhibit a significant linguistic bias, with the vast majority of benchmark datasets being exclusively in English or (at best) translated from English. We address this limitation by introducing {\sc Math-PT}, a novel dataset comprising 1,729 mathematical problems written in European and Brazilian Portuguese. {\sc Math-PT} is curated from a variety of high-quality native sources, including mathematical Olympiads, competitions, and exams from Portugal and Brazil. We present a comprehensive benchmark of current state-of-the-art LLMs on {\sc Math-PT}, revealing that frontier reasoning models achieve strong performance in multiple choice questions compared to open weight models, but that their performance decreases for questions with figures or open-ended questions. To facilitate future research, we release the benchmark dataset and model outputs.

cs.CL

Budget-constrained Collaborative Renewable Energy Forecasting Market

Accurate power forecasting from renewable energy sources (RES) is crucial for integrating additional RES capacity into the power system and realizing sustainability goals. This work emphasizes the importance of integrating decentralized spatio-temporal data into forecasting models. However, decentralized data ownership presents a critical obstacle to the success of such spatio-temporal models, and incentive mechanisms to foster data-sharing need to be considered. The main contributions are a) a comparative analysis of the forecasting models, advocating for efficient and interpretable spline LASSO regression models, and b) a bidding mechanism within the data/analytics market to ensure fair compensation for data providers and enable both buyers and sellers to express their data price requirements. Furthermore, an incentive mechanism for time series forecasting is proposed, effectively incorporating price constraints and preventing redundant feature allocation. Results show significant accuracy improvements and potential monetary gains for data sellers. For wind power data, an average root mean squared error improvement of over 10% was achieved by comparing forecasts generated by the proposal with locally generated ones.

cs.LG

Sparse Deconvolution Methods for Online Energy Estimation in Calorimeters Operating in High Luminosity Conditions

Energy reconstruction in calorimeters operating in high luminosity particle colliders has become a remarkable challenge. In this scenario, pulses from a calorimeter front-end output overlap each other (pile-up effect), compromising the energy estimation procedure when no preprocessing for signal disentanglement is accomplished. Recently, methods based on signal deconvolution have been proposed for both online and offline reconstructions. For online processing, constraints concerning fast processing, memory requirements, and cost implementation limit the overall performance. Offline reconstruction allows the use of Sparse Representation theory to implement sophisticated Iterative Deconvolution methods. This paper presents Iterative Deconvolution methods based on Sparse Representation algorithms whose computational cost is effective for online implementation. Using simulated data, current techniques were compared to the proposed Sparse Representation ones for performance validation in the online environments. Analysis has shown that, despite the higher computational cost, when compared to standard methods, the performance improvement may justify the use of the proposed techniques, in particular for the Separable Surrogate Functional, which is shown to be feasible for implementation in modern FPGAs.

physics.ins-det