arXiv · 2510.23650
Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs
Abstract
We proposed Static and Dynamic -- two zero-shot logits-layer debiasing methods. Dynamic reduces bias by up to 70% with minimal fluency loss. Logits intervention outperforms hidden-layer approaches. We show semantic-aware logits intervention is stable and effective for debiasing aligned LLMs.
Explore related subjects
Keep this discovery
Wei Xia. 2025-10-25. Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs. https://arxiv.org/abs/2510.23650
Cite the original work for its findings. Save a collection to share your selection of sources.