arXiv · 2507.18523
The Moral Gap of Large Language Models
Abstract
Moral foundation detection is crucial for analyzing social discourse and developing ethically-aligned AI systems. While large language models excel across diverse tasks, their performance on specialized moral reasoning remains unclear. This study provides the first comprehensive comparison between state-of-the-art LLMs and fine-tuned transformers across Twitter and Reddit datasets using ROC, PR, and DET curve analysis. Results reveal substantial performance gaps, with LLMs exhibiting high false negative rates and systematic under-detection of moral content despite prompt engineering efforts. These findings demonstrate that task-specific fine-tuning remains superior to prompting for moral reasoning applications.
Explore related subjects
Keep this discovery
Maciej Skorski, Alina Landowska. 2025-07-24. The Moral Gap of Large Language Models. https://doi.org/10.13140/rg.2.2.26221.70880
Cite the original work for its findings. Save a collection to share your selection of sources.