arXiv · 2512.03803
Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
Abstract
Encoder-decoder models such as FLAN-T5 are finetuned to follow instructions, but often fail when the instructions conflict with memorized continuations ingrained during training. To understand this behavior, we adapt DoLa to FLAN-T5 and examine how representations evolve in the decoder. Our findings show that T5's intermediate layers undergo rapid shifts driven by cross-attention to the encoder. When projected through the language modeling head, each depth presents highly volatile token preferences, leading to unreliable behavior with contrastive decoding. Motivated by this, we introduce a gradient-based activation-steering method that injects an "instruction-compliance" direction into mid-decoder layers, where the representation is both meaningful and still malleable. This intervention dramatically improves MemoTrap performance (52% to 99.7%), demonstrating that mechanistic steering can succeed where contrastive decoding fails in Seq2Seq architectures.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Huey Sun, Anabel Yong, Lorenzo Gilly, Felipe Jin. 2025-12-03. Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5. https://arxiv.org/abs/2512.03803
Cite the original work for its findings. Save a collection to share your selection of sources.