SearcharxivSearch

arXiv subjects

Maohua Li

Publications and source records attributed to Maohua Li.

12 recordsLinked to original sources

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

Modern text-to-image diffusion transformers (DiTs) generate images through joint attention, in which text and image tokens interact directly within a single sequence. In large-scale DiTs, the conditioning input contains not only the user prompt but also chat-template tokens introduced by LLM-based text encoders. Yet how these tokens participate in the denoising computation remains poorly understood. To probe this, we introduce a causal interpretability framework. Using it to separate prompt-content tokens from chat-template tokens, we find that the template tokens carry little prompt-specific information at the encoder output. Yet surprisingly, they emerge as dominant image-to-text attention sinks and causally maintain object identity inside the DiT, acting as implicit semantic registers. We show that they acquire this identity indirectly. Rather than reading the prompt tokens, they draw the identity from the image latents into which the prompt semantics have already been injected at the very first layer. We further reveal a division of labor across heads and depth in DiTs, where distinct heads route semantics or render visual structure, and identity is committed in early blocks, carried by middle blocks, and refined in late ones. As a practical payoff, this analysis yields a training-free pruning rule that removes the causally inert prompt-reading heads and cuts $20\%$ of joint-attention FLOPs at a $1.4$-point cost in GenEval accuracy. Overall, our work not only reveals that the tokens encoding semantics at the input need not be those that maintain them during generation, but also provides a causal view of internal mechanisms in diffusion transformers.

cs.CV

Rethinking Cross-Layer Information Routing in Diffusion Transformers

Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization, attention, conditioning, objectives, and latent autoencoders -- has been extensively revisited. The residual stream that governs how information accumulates across layers, however, has been directly inherited from the original Transformer. In this paper, we present a systematic empirical analysis of cross-layer information flow in DiTs, jointly along depth and denoising timestep, and identify three concrete symptoms of traditional residual addition, namely monotonic forward magnitude inflation, sharp backward gradient decay, and pronounced block-wise redundancy. Motivated by this diagnosis, we propose Diffusion-Adaptive Routing (\textsc{DAR}), a drop-in residual replacement that performs \emph{learnable, timestep-adaptive, and non-incremental} aggregation over the history of sublayer outputs. Moreover, the proposed \textsc{DAR} is compatible with many modern Transformer enhancement methods, such as REPA. On ImageNet $256\times256$, \textsc{DAR} improves SiT-XL/2 by $2.11$ FID ($7.56$ vs.\ $9.67$) and matches the baseline's converged quality with $8.75\times$ fewer training iterations. Stacked on top of REPA, it yields a $2\times$ training acceleration in the early stage, suggesting cross-layer information routing as an underexplored design axis in diffusion modeling, one that operates orthogonally to existing representation-alignment objectives. Beyond pretraining, \textsc{DAR} can also be applied during the fine-tuning stage of large-scale T2I models and preserves high-frequency details during Distribution Matching Distillation.

cs.CV

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely either on native sparse training or on heuristic token eviction, creating an undesirable trade-off among efficiency, training cost, and accuracy. In this work, we show that full-attention LLMs are already intrinsically sparse and can be transformed into highly sparse models with only minimal adaptation. Our approach is built on three observations: (1) only a small subset of attention heads truly requires full long-context processing; (2) long-range retrieval is governed primarily by a low-dimensional subspace, allowing relevant tokens to be retrieved efficiently with a 16-dimensional indexer; and (3) the useful token budget is strongly query-dependent, making dynamic top-$p$ selection more suitable than fixed top-$k$ sparsification. Based on these insights, we propose RTPurbo, which retains the full KV cache only for retrieval heads and introduces a lightweight token indexer for sparse attention. By exploiting the model's intrinsic sparsity, RTPurbo achieves sparsification with only a few hundred training steps. Experiments on long-context benchmarks and reasoning tasks show that RTPurbo preserves near-lossless accuracy while delivering substantial efficiency gains, including up to a 9.36$\times$ prefill speedup at 1M context and about a 2.01$\times$ decode speedup. These results suggest that strong sparse inference can be obtained from standard full-attention training without expensive native sparse pretraining.

cs.CL

Soliton,breathers,positons and rogue waves for the vector complex modified Korteweg-de Vries equation

This paper constructs the $N$-fold Darboux transformation (DT) for the vector complex modified Korteweg-de Vries (vcmKdV) equation and presents its determinant representation. Utilizing the DT and multi-fold eigenvalue degeneracy, we derive globally bounded solutions for the vcmKdV equation, including $N$-bright-bright-bright solitons, $N$-dark-bright-bright solitons, $N$-breathers, $N$-positon solutions, and $N$th-order rogue wave solutions." All these solutions are globally bounded. Graphical representations of bright-bright-bright and dark-bright-bright soliton solutions are provided, illustrating phenomena where periodic oscillatory waves coexist or interact with solitons. The collision scenarios of the two-bright-bright-bright solution have been investigated by using the asymptotic analysis. The bounded Akhmediev breather, the bounded breather with dark-bright soliton and breather-breather mixed waves are graphically shown. We give the graphs of the positon solution, the rogue wave and the rogue wave mixes with dark-bright solitons and breathers.

nlin.SI

Inference-Scale Complexity in ANN-SNN Conversion for High-Performance and Low-Power Applications

Spiking Neural Networks (SNNs) have emerged as a promising substitute for Artificial Neural Networks (ANNs) due to their advantages of fast inference and low power consumption. However, the lack of efficient training algorithms has hindered their widespread adoption. Even efficient ANN-SNN conversion methods necessitate quantized training of ANNs to enhance the effectiveness of the conversion, incurring additional training costs. To address these challenges, we propose an efficient ANN-SNN conversion framework with only inference scale complexity. The conversion framework includes a local threshold balancing algorithm, which enables efficient calculation of the optimal thresholds and fine-grained adjustment of the threshold value by channel-wise scaling. We also introduce an effective delayed evaluation strategy to mitigate the influence of the spike propagation delays. We demonstrate the scalability of our framework in typical computer vision tasks: image classification, semantic segmentation, object detection, and video classification. Our algorithm outperforms existing methods, highlighting its practical applicability and efficiency. Moreover, we have evaluated the energy consumption of the converted SNNs, demonstrating their superior low-power advantage compared to conventional ANNs. This approach simplifies the deployment of SNNs by leveraging open-source pre-trained ANN models, enabling fast, low-power inference with negligible performance reduction. Code is available at https://github.com/putshua/Inference-scale-ANN-SNN.

cs.NE

The dynamic of the positons for the reverse space-time nonlocal short pulse equation

In this paper, the Darboux transformation (DT) of the reverse space-time (RST) nonlocal short pulse equation is constructed by a hodograph transformation and the eigenfunctions of its Lax pair. The multi-soliton solutions of the RST nonlocal short pulse equation are produced through the DT, which can be expressed in terms of determinant representation. By taking different values of eigenvalues, bounded soliton solutions and unbounded soliton solutions can be obtained. In addition, based on the degenerate Darboux transformation, the $N$-positon solutions of the RST nonlocal short pulse equation are computed from the determinant expression of the multi-soliton solution. Furthermore, different kinds of mixed solutions are also presented, and the interaction properties between positons and solitons are investigated.

nlin.SI

Generating mechanism and dynamic of the smooth positons for the derivative nonlinear Schrödinger equation

Based on the degenerate Darboux transformation, the $n$-order smooth positon solutions for the derivative nonlinear Schrödinger equation are generated by means of the general determinant expression of the $N$-soliton solution, and interesting dynamic behaviors of the smooth positons are shown by the corresponding three dimensional plots in this paper. Furthermore, the decomposition process, bent trajectory and the change of the phase shift for the positon solutions are discussed in detail. Additional, three kinds of mixed solutions, namely (1) the hybrid of one-positon and two-positon solutions, (2) the hybrid of two-positon and two-positon solutions, and (3) the hybrid of one-soliton and three-positon solutions are presented and their rather complicated dynamics are revealed.

nlin.SI

Ghost symmetry of the discrete KP hierarchy

In this paper, with the help of the $S$ function and ghost symmetry for the discrete KP hierarchy which is a semi-discrete version of the KP hierarchy, the ghost flow on its eigenfunction(adjoint eigenfunction) and the spectral representation of its Baker-Akhiezer function and adjoint Baker-Akhiezer function are derived. From these observations above, some important distinctions between the discrete KP hierarchy and KP hierarchy are shown. Also we give the ghost flow on the tau function and another kind of proof of the ASvM formula of the discrete KP hierarchy.

nlin.SI

The wronskian solution of the constrained discrete KP hierarchy

From the constrained discrete KP (cdKP) hierarchy, the Ablowitz-Ladik lattice has been derived. By means of the gauge transformation, the Wronskian solution of the Ablowitz-Ladik lattice have been given. The $u_1$ of the cdKP hierarchy is a Y-type soliton solution for odd times of the gauge transformation, but it becomes a dark-bright soliton solution for even times of the gauge transformation. The role of the discrete variable $n$ in the profile of the $u_1$ is discussed.

math-ph

The gauge transformation of the constrained semi-discrete KP hierarchy

In this paper, the gauge transformation of the constrained semi-discrete KP(cdKP) hierarchy is constructed explicitly by the suitable choice of the generating functions. Under the $m$-step successive gauge transformation $T_m$, we give the transformed (adjoint) eigenfunctions and the $τ$-function of the transformed Lax operator of the cdKP hierarchy.

nlin.SI

The Recursion operators of the BKP hierarchy and the CKP Hierarchy

In this paper, under the constraints of the BKP(CKP) hierarchy, a crucial observation is that the odd dynamical variable $u_{2k+1}$ can be explicitly expressed by the even dynamical variable $u_{2k}$ in the Lax operator $L$ through a new operator $B$. Using operator $B$, the essential differences between the BKP hierarchy and the CKP hierarchy are given by the flow equations and the recursion operators under the $(2n+1)$-reduction. The formal formulas of the recursion operators for the BKP and CKP hierarchy under $(2n+1)$-reduction are given. To illustrate this method, the two recursion operators are constructed explicitly for the 3-reduction of the BKP and CKP hierarchies. The $t_7$ flows of $u_2$ are generated from $t_1$ flows by the above recursion operators, which are consistent with the corresponding flows generated by the flow equations under 3-reduction.

math-ph