SearcharxivSearch

arXiv subjects

Seokhyun Kim

Publications and source records attributed to Seokhyun Kim.

2 recordsLinked to original sources

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets. We present CoTinyVLA, a 0.9B-parameter action model on a Qwen3.5-0.8B backbone that obtains that robustness by structuring supervision instead of enlarging the model. Three components target different axes of the problem: dual-view temporal input of 16 history frames per step with textual camera and time markers; hierarchical chain-of-thought (CoT) distillation from a 35B teacher into an episode-level Plan and a chunk-level Think span over task phase, gripper state and next subaction; and paraphrase augmentation expanding 40 base commands into 800 variants. On LIBERO-Plus, spanning 10,030 perturbed tasks across seven perturbation dimensions, CoTinyVLA reaches 90.8% on Spatial, 87.3% on Object, 86.6% on Goal and 80.7% on Long, leading the strongest 7B baseline on all four suites by 4.7, 2.8, 15.9 and 3.0 points, with every margin interval excluding zero. The gains concentrate on the hardest axes of the benchmark: across the eleven published baselines none exceeds 53.2% on Robot Initial States in any suite, whereas CoTinyVLA reaches 73.6% on Goal against 39.9% for the strongest baseline. Ablations show the three components to be separable by perturbation axis, and at a matched image budget how frames are divided between the two cameras and across time accounts for 8.6 points on its own. Closed-loop inference peaks at 2.25 GiB of allocated GPU memory, and paired interventions show the episode Plan to be load-bearing: replacing it with an empty or contradictory span costs 40 to 45 points of success. Structured supervision thus lets a 0.9B backbone exceed all of them. Code: https://github.com/BrainJellyPie/CoTinyVLA

cs.AI

Delay-Optimal Data Forwarding in Vehicular Sensor Networks

Vehicular Sensor Network (VSN) is emerging as a new solution for monitoring urban environments such as Intelligent Transportation Systems and air pollution. One of the crucial factors that determine the service quality of urban monitoring applications is the delivery delay of sensing data packets in the VSN. In this paper, we study the problem of routing data packets with minimum delay in the VSN, by exploiting i) vehicle traffic statistics, ii) anycast routing and iii) knowledge of future trajectories of vehicles such as buses. We first introduce a novel road network graph model that incorporates the three factors into the routing metric. We then characterize the packet delay on each edge as a function of the vehicle density, speed and the length of the edge. Based on the network model and delay function, we formulate the packet routing problem as a Markov Decision Process (MDP) and develop an optimal routing policy by solving the MDP. Evaluations using real vehicle traces in a city show that our routing policy significantly improves the delay performance compared to existing routing protocols.

cs.NI