SearcharxivSearch

arXiv subjects

Changhui Sun

Publications and source records attributed to Changhui Sun.

3 recordsLinked to original sources

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially less data. However, standard token-level OPD can provide only fragmented corrections along an erroneous student trajectory and cannot unfold a complete and correct repair path. Motivated by this limitation, we propose \emph{Step-Level On-Policy Distillation} (SOPD), which combines the long-horizon correction of supervised fine-tuning (SFT) with the on-policy advantage of OPD to provide step-level supervision over complete student-generated trajectories. We show that, at different limits of step length, SOPD reduces to SFT or approximates OPD. Compared with SFT, the teacher responses in SOPD are conditioned on student trajectories and therefore align more closely with student-visited states; compared with OPD, SOPD provides longer-horizon corrections rather than fragmented token-level guidance. Across both reasoning and agent tasks, SOPD substantially outperforms conventional SFT and OPD. For example, on ALFWorld, SOPD improves the average success rate by 13.4 points over Vanilla OPD. We hope this work offers a new perspective for future research on distillation methods.

cs.CL

MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models

Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Whereas AR decoding exposes a contiguous prefix, MDLM denoising produces dynamic, non-contiguous configurations of revealed and masked tokens. Conventional positional encodings such as RoPE capture sequence order and pairwise displacement but remain insensitive to this evolving token-availability structure. To address this limitation, we propose MDLMPE, a positional encoding designed specifically for masked diffusion. To the best of our knowledge, MDLMPE is the first method to make positional representations explicitly aware of the changing revealed/masked configuration. It represents token availability as a binary sequence, applies distance-aware Gaussian weighting, and projects the resulting pattern through a cosine basis to obtain distribution-aware positional features. These features are added to token embeddings and mapped by a lightweight MLP to angular offsets that modulate the standard RoPE phases. Extensive experiments on LLaDA and DREAM demonstrate that MDLMPE generally outperforms conventional positional encoding methods across supervised fine-tuning, pretraining, zero-shot evaluation, and block-diffusion settings. Further ablations show that the complete combination of availability state, Gaussian locality, spectral basis, and embedding injection yields the strongest result. These results establish the evolving token-availability distribution as a useful positional signal for masked diffusion language models.

cs.CL

Traffic-MoE: A Sparse Foundation Model for Network Traffic Security Analysis

As adversaries increasingly weaponize encryption and protocol obfuscation to evade traffic detection, traditional methods are rendered obsolete, necessitating deep learning to unmask sophisticated threats. However, the prohibitive computational costs of existing large models create a critical defense gap, hindering their deployment in real-time and throughput-sensitive environments. To close this vulnerability, we introduce Traffic-MoE, a sparse foundation model tailored for traffic security analysis. By dynamically routing traffic tokens to a small subset of specialized experts, Traffic-MoE effectively decouples model capacity from computational overhead. Extensive evaluations across four security-oriented tasks demonstrate that Traffic-MoE achieves state-of-the-art or highly competitive performance compared to leading competitors. Crucially, it delivers a 70.42% increase in throughput, reduces inference latency by 41.39% while significantly optimizing GPU memory consumption. Beyond efficiency, Traffic-MoE exhibits superior robustness against adversarial traffic shaping and maintains strong detection capabilities in few-shot scenarios, establishing a scalable and resilient paradigm for modern network traffic security analysis.

cs.CR