arXiv · 2504.07801
FairEval: Evaluating Fairness in LLM-Based Recommendations with Personality Awareness
Abstract
Recent advances in Large Language Models (LLMs) have enabled their application to recommender systems (RecLLMs), yet concerns remain regarding fairness across demographic and psychological user dimensions. We introduce FairEval, a novel evaluation framework to systematically assess fairness in LLM-based recommendations. FairEval integrates personality traits with eight sensitive demographic attributes,including gender, race, and age, enabling a comprehensive assessment of user-level bias. We evaluate models, including ChatGPT 4o and Gemini 1.5 Flash, on music and movie recommendations. FairEval's fairness metric, PAFS, achieves scores up to 0.9969 for ChatGPT 4o and 0.9997 for Gemini 1.5 Flash, with disparities reaching 34.79 percent. These results highlight the importance of robustness in prompt sensitivity and support more inclusive recommendation systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chandan Kumar Sah, Xiaoli Lian, Tony Xu, Li Zhang. 2025-04-10. FairEval: Evaluating Fairness in LLM-Based Recommendations with Personality Awareness. https://arxiv.org/abs/2504.07801
Cite the original work for its findings. Save a collection to share your selection of sources.