arXiv · 2604.20936
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
Abstract
We present AttentionBender, a tool that manipulates cross-attention in Video Diffusion Transformers to help artists probe the internal mechanics of black-box video generation. While generative outputs are increasingly realistic, prompt-only control limits artists' ability to build intuition for the model's material process or to work beyond its default tendencies. Using an autobiographical research-through-design approach, we built on Network Bending to design AttentionBender, which applies 2D transforms (rotation, scaling, translation, etc.) to cross-attention maps to modulate generation. We assess AttentionBender by visualizing 4,500+ video generations across prompts, operations, and layer targets. Our results suggest that cross-attention is highly entangled: targeted manipulations often resist clean, localized control, producing distributed distortions and glitch aesthetics over linear edits. AttentionBender contributes a tool that functions both as an Explainable AI style probe of transformer attention mechanisms, and as a creative technique for producing novel aesthetics beyond the model's learned representational space.
Explore related subjects
Keep this discovery
Adam Cole, Mick Grierson. 2026-04-22. AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe. https://doi.org/10.1145/3803784.3807565
Cite the original work for its findings. Save a collection to share your selection of sources.