SearcharxivSearch

arXiv subjects

Mingkai Nie

Publications and source records attributed to Mingkai Nie.

2 recordsLinked to original sources

Ask Self, Ask Others: Relation Is All You Need

Attention dominates token mixing, but it collapses relation formation and flow allocation into a single score-to-flow step. We introduce Relation, which separates them by first organizing pairwise evidence into explicit Self and Exchange relations and deriving information flow afterward. Relation first decides whether a token should rely on itself or draw from its history, and if it draws from history, where to look. This relational organization gives rise to Full Relation, FlashRelation, Linear Relation, and Hybrid Relation. Across matched decoder-only models, Full Relation achieves lower mean final-validation NLL than MHA and reaches the paired MHA final training loss with 4.5-7.3% fewer tokens. Structural diagnostics further show that Relation learns a distinct depth organization: the first layer acts as a current-token anchor and a high-rank router, while later layers shift strongly toward history. In a fixed-context reference benchmark, FlashRelation is 4.17-5.28x faster than the materialized Full Relation implementation. Across scale-matched production workloads, it reaches 89.7-92.9% of PyTorch FlashAttention throughput while executing the exact Full Relation operator. Hybrid Relation demonstrates that Full and Linear Relation layers can be composed within a single decoder. These results support a relation-first view of token mixing: ask Self, ask Others, then let Flow follow Relation.

cs.LG

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU

We test whether decoder-only language-model FFNs require SwiGLU's open positive tail. We introduce MemGLU as a closed-tail comparator derived from a memristive branch geometry. Across paired 9M and 30M pretraining runs with three seeds, MemGLU remains within about 0.1% of SwiGLU in validation NLL. Trained SwiGLU checkpoints are sensitive to positive-tail suppression, while mechanism diagnostics show that the two models use their gates differently despite similar losses. These results suggest that models adapt to the gate geometry available during pretraining. At the tested scales, SwiGLU's open positive tail is not necessary for decoder-only language-model FFNs.

cs.LG