SearcharxivSearch

arXiv subjects

Yitong Shang

Publications and source records attributed to Yitong Shang.

3 recordsLinked to original sources

Self-Policy Distillation via Capability-Selective Subspace Projection

Self-distillation bootstraps large language models (LLMs) by training on their own generations. However, existing methods either rely on external signals to curate self-generated outputs (e.g., correctness filtering, execution feedback, and reward search), which are costly and unavailable for the best-performing frontier models, or skip curation entirely and train on all raw outputs, an approach that is often domain-specific and hard to generalize. Both also share a deeper weakness that self-generated outputs entangle task-relevant capability with others, such as stylistic patterns, formatting artifacts, and model-specific errors, diluting the signal for the specific capability one aims to improve. In this paper, we propose Self-Policy Distillation (SPD), which achieves generalizable, capability selective without any external signal. Specifically, SPD extracts a low-rank capability subspace from the model's own gradients on correctness-defining tokens, projects key-value (KV) activations into this subspace during self-generation, and fine-tunes on the resulting raw outputs with standard next-token prediction loss. Through extensive experiments across code generation, mathematical reasoning, and multiple-choice QA, we show that SPD achieves up to 13% improvement over state-of-the-art self-distillation methods without external signals and up to 16% improvement over pre-trained baselines. Notably, SPD demonstrates superior generalizability, achieving 15% better performance under out-of-domain generalization settings.

cs.CL

AutoFed: Personalized Federated Traffic Prediction via Adaptive Prompt

Accurate traffic prediction is essential for Intelligent Transportation Systems, including ride-hailing, urban road planning, and vehicle fleet management. However, due to significant privacy concerns surrounding traffic data, most existing methods rely on local training, resulting in data silos and limited knowledge sharing. Federated Learning (FL) offers an efficient solution through privacy-preserving collaborative training; however, standard FL struggles with the non-independent and identically distributed (non-IID) problem among clients. This challenge has led to the emergence of Personalized Federated Learning (PFL) as a promising paradigm. Nevertheless, current PFL frameworks require further adaptation for traffic prediction tasks, such as specialized graph feature engineering, data processing, and network architecture design. A notable limitation of many prior studies is their reliance on hyper-parameter optimization across datasets-information that is often unavailable in real-world scenarios-thus impeding practical deployment. To address this challenge, we propose AutoFed, a novel PFL framework for traffic prediction that eliminates the need for manual hyper-parameter tuning. Inspired by prompt learning, AutoFed introduces a federated representor that employs a client-aligned adapter to distill local data into a compact, globally shared prompt matrix. This prompt then conditions a personalized predictor, allowing each client to benefit from cross-client knowledge while maintaining local specificity. Extensive experiments on real-world datasets demonstrate that AutoFed consistently achieves superior performance across diverse scenarios. The code of this paper is provided at https://github.com/RS2002/AutoFed .

cs.LG

Joint Infrastructure Planning and Order Assignment for On-Demand Food-Delivery Services with Coordinated Drones and Human Couriers

This paper investigates the optimal infrastructure planning and order assignment problem of an on-demand food-delivery platform with a mixed fleet of drones and human couriers. The platform has two delivery modes: (a) ground delivery and (b) drone-assisted delivery (i.e., air delivery). In ground delivery, couriers directly collect and transport orders from restaurants to destinations. For air delivery, the delivery process involves three legs: initially, a human courier picks up the order from the restaurant and transports it to a nearby launchpad, where personnel load the orders onto drones and replace batteries as needed. The loaded drone then transports the order from the launchpad to a kiosk, where another courier retrieves the order from the kiosk for final delivery. The platform must determine the optimal locations for launchpads and kiosks within a transportation network, and devise an order assignment strategy that allocates food-delivery orders between ground and air delivery considering the bundling probabilities of ground deliveries and the waiting times at launchpads and kiosks. We formulate the platform's problem as a mixed-integer nonlinear program and develop a novel neural network-assisted optimization method to obtain high-quality solutions. A case study in Hong Kong validates our model and algorithm, revealing that drone delivery reduces operational costs, minimizes courier fleet size, and increases order bundling opportunities. We also find that the expansion of air delivery services may entail larger delivery times due to the trade-off between the travel time savings induced by the faster air delivery and the associated detours incurred by intermodal transfer and extra waiting times at launchpads and kiosks, which crucially depends on the distance of the orders and the sequence of activating long-distance air delivery routes versus short-distance ones.

eess.SY