arXiv · 2609.33895
Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers
Abstract
In diffusion-based generation, a neural network can be trained to predict the clean data, the noise, or the velocity from a noisy input. These prediction targets are interconvertible and describe the same generative process, yet plain Diffusion Transformers operating on large pixel patches succeed with clean prediction and fail with noise or velocity prediction. We argue that this asymmetry arises because noisy targets require the residual stream to preserve noise-dependent input variation through depth for the final readout, forcing subsequent layers to compute on noisy representations. A spectrally concentrated clean target imposes a lighter demand, leaving greater freedom to organize hidden representations for subsequent computation. We call this preservation requirement *residual-stream burden* and show how it shapes representation learning in Diffusion Transformers. Controlled experiments indicate that the exploitable structure is spectral concentration in patch space and that the bandwidth of the persistent residual state is a key resource for noisy prediction. We further show that this account is consistent with recent decoupled pixel-space architectures, whose diverse designs all reduce the residual-stream burden on the main pathway. To examine this understanding from a complementary direction, we expand and reorganize the residual-stream bandwidth directly, introducing Spatially Indexed Hyper-Connections (SiHC) that reach FID 1.71 on ImageNet $256^2$. Together, these results identify residual-stream burden as a mechanism through which prediction targets and architecture jointly shape representation learning in Diffusion Transformers.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tongtong Liang, Siqi Kou, Ziqiao Xi, Esha Singh, Kun Zhou, Zhijie Deng, Alexander Cloninger, Yu-Xiang Wang, Rahul Parhi. 2026-09-27. Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers. https://arxiv.org/abs/2609.33895
Cite the original work for its findings. Save a collection to share your selection of sources.