Searcharxiv⌕ Search

arXiv subjects

Johan Barthélemy

Publications and source records attributed to Johan Barthélemy.

6 recordsLinked to original sources

Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation

Advancing large language models (LLMs) for the next point-of-interest (POI) recommendation task faces two fundamental challenges: (i) although existing methods produce semantic IDs that incorporate semantic information, their topology-blind indexing fails to preserve semantic continuity, meaning that proximity in ID values does not mirror the coherence of the underlying semantics; and (ii) supervised fine-tuning (SFT)-based methods restrict model outputs to top-1 predictions. These approaches suffer from "answer fixation" and neglect the need for top-k ranked lists and reasoning due to the scarcity of supervision. We propose Refine-POI, a framework that addresses these challenges through topology-aware ID generation and reinforcement fine-tuning. First, we introduce a hierarchical self-organizing map (SOM) quantization strategy to generate semantic IDs, ensuring that coordinate proximity in the codebook reflects semantic similarity in the latent space. Second, we employ a policy-gradient framework to optimize the generation of top-k recommendation lists, liberating the model from strict label matching. Extensive experiments on three real-world datasets demonstrate that Refine-POI significantly outperforms state-of-the-art baselines, effectively synthesizing the reasoning capabilities of LLMs with the representational fidelity required for accurate and explainable next-POI recommendation.

cs.IR↗

LLM as HPC Expert: Extending RAG Architecture for HPC Data

High-Performance Computing (HPC) is crucial for performing advanced computational tasks, yet their complexity often challenges users, particularly those unfamiliar with HPC-specific commands and workflows. This paper introduces Hypothetical Command Embeddings (HyCE), a novel method that extends Retrieval-Augmented Generation (RAG) by integrating real-time, user-specific HPC data, enhancing accessibility to these systems. HyCE enriches large language models (LLM) with real-time, user-specific HPC information, addressing the limitations of fine-tuned models on such data. We evaluate HyCE using an automated RAG evaluation framework, where the LLM itself creates synthetic questions from the HPC data and serves as a judge, assessing the efficacy of the extended RAG with the evaluation metrics relevant for HPC tasks. Additionally, we tackle essential security concerns, including data privacy and command execution risks, associated with deploying LLMs in HPC environments. This solution provides a scalable and adaptable approach for HPC clusters to leverage LLMs as HPC expert, bridging the gap between users and the complex systems of HPC.

cs.DC↗

Hard 3-CNF-SAT problems are in $P$ -- A first step in proving $NP=P$

The relationship between the complexity classes $P$ and $NP$ is an unsolved question in the field of theoretical computer science. In the first part of this paper, a lattice framework is proposed to handle the 3-CNF-SAT problems, known to be in $NP$. In the second section, we define a multi-linear descriptor function ${\cal H}_φ$ for any 3-CNF-SAT problem $φ$ of size $n$, in the sense that ${\cal H}_φ: \{0,1\}^n \rightarrow \{0,1\}^n$ is such that $Im \; {\cal H}_φ$ is the set of all the solutions of $φ$. A new merge operation ${\cal H}_φ\bigwedge {\cal H}_ψ$ is defined, where $ψ$ is a single 3-CNF clause. Given ${\cal H}_φ$ [but this can be of exponential complexity], the complexity needed for the computation of $Im \; {\cal H}_φ$, the set of all solutions, is shown to be polynomial for hard 3-CNF-SAT problems, i.e. the one with few ($\leq 2^k$) or no solutions. The third part uses the relation between ${\cal H}_φ$ and the indicator function $\mathbb{1}_{{\cal S}_φ}$ for the set of solutions, to develop a greedy polynomial algorithm to solve hard 3-CNF-SAT problems.

cs.CC↗

Comparison of Discrete Choice Models and Artificial Neural Networks in Presence of Missing Variables

Classification, the process of assigning a label (or class) to an observation given its features, is a common task in many applications. Nonetheless in most real-life applications, the labels can not be fully explained by the observed features. Indeed there can be many factors hidden to the modellers. The unexplained variation is then treated as some random noise which is handled differently depending on the method retained by the practitioner. This work focuses on two simple and widely used supervised classification algorithms: discrete choice models and artificial neural networks in the context of binary classification. Through various numerical experiments involving continuous or discrete explanatory features, we present a comparison of the retained methods' performance in presence of missing variables. The impact of the distribution of the two classes in the training data is also investigated. The outcomes of those experiments highlight the fact that artificial neural networks outperforms the discrete choice models, except when the distribution of the classes in the training data is highly unbalanced. Finally, this work provides some guidelines for choosing the right classifier with respect to the training data.

stat.ML↗

Interaction prediction between groundwater and quarry extension using discrete choice models and artificial neural networks

Groundwater and rock are intensively exploited in the world. When a quarry is deepened the water table of the exploited geological formation might be reached. A dewatering system is therefore installed so that the quarry activities can continue, possibly impacting the nearby water catchments. In order to recommend an adequate feasibility study before deepening a quarry, we propose two interaction indices between extractive activity and groundwater resources based on hazard and vulnerability parameters used in the assessment of natural hazards. The levels of each index (low, medium, high, very high) correspond to the potential impact of the quarry on the regional hydrogeology. The first index is based on a discrete choice modelling methodology while the second is relying on an artificial neural network. It is shown that these two complementary approaches (the former being probabilistic while the latter fully deterministic) are able to predict accurately the level of interaction. Their use is finally illustrated by their application on the Boverie quarry and the Tridaine gallery located in Belgium. The indices determine the current interaction level as well as the one resulting from future quarry extensions. The results highlight the very high interaction level of the quarry with the gallery.

physics.geo-ph↗

A 3-CNF-SAT descriptor algebra and the solution of the P=NP conjecture

The relationship between the complexity classes P and NP is an unsolved question in the field of theoretical computer science. In this paper, we investigate a descriptor approach based on lattice properties. This paper proposes a new way to decide the satisfiability of any 3-CNF-SAT problem. The analysis of this exact [non heuristical] algorithm shows a strictly bounded exponential complexity. The complexity of any 3-CNF-SAT solution is bounded by O(2^490). This over-estimated bound is reached by an algorithm working on the smallest description (via descriptor functions) of the evolving set of solutions in function of the already considered clauses, without exploring these solutions. Any remark about this paper is warmly welcome.

cs.CC↗