SearcharxivSearch

arXiv subjects

Evgeny Stupachenko

Publications and source records attributed to Evgeny Stupachenko.

3 recordsLinked to original sources

From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models

Structured pruning is a practical approach to deploying large language models (LLMs) efficiently, as it yields compact, hardware-friendly architectures. However, the dominant local paradigm is task-agnostic: by optimizing layer-wise reconstruction rather than task objectives, it tends to preserve perplexity or generic zero-shot behavior but fails to capitalize on modest task-specific calibration signals, often yielding limited downstream gains. We revisit global structured pruning and present GISP, Global Iterative Structured Pruning, a post-training method that removes attention heads and MLP channels using first-order, loss-based important scores aggregated at the structure level with block-wise normalization. Built on this global importance metric, GISP adopts an iterative schedule, rather than one-shot pruning, stabilizes accuracy at higher sparsity, and mitigates perplexity collapse without requiring intermediate fine-tuning. Importantly, the iterative pruning forms nested subnetworks that support a ''prune-once, deploy-many'' workflow. Furthermore, GISP defines structural importance directly with respect to a target loss, making it easy to adapt pruning to task-specific objectives. In this work, we use perplexity for language modeling and a margin-based objective for decision-style tasks. Extensive experiments show that across Llama2-7B/13B, Llama3-8B, and Mistral-0.3-7B, GISP consistently lowers WikiText-2 perplexity and improves on downstream accuracy, with especially strong gains at 40-50% sparsity; on DeepSeek-R1-Distill-Llama-3-8B and Qwen3-8B with GSM8K, task-aligned calibration substantially boosts exact-match accuracy. The implementation is available at https://github.com/uncc-efficient-ai/GISP.

cs.CL

Neural network concatenation for Polar Codes

When a neural network (NN) is used to decode a polar code, its training complexity scales exponentially as the code block size (or to be precise, as a number of message bits) increases. Therefore, existing solutions that use a neural network for polar decoders are stuck with short block sizes like 16 or 32. Despite the fact that the NN training is very complex for long polar codes, the NN decoding gives the better latency and its performance is potentially close to the maximum likelihood (ML). In this paper, we describe an efficient algorithm to create the NN decoding for a polar code of any size with the initial performance that is equal or better than that of successive cancelation (SC). Therefore, it creates an opportunity to design the NN based decoding with the performance that is as close to the ML, as the training time allows.

cs.IT

Low latency communication over commercially available LTE and remote driving

In addition to autonomous car operation, in many cases it is desirable to let a human drive the vehicle remotely. To make remote operation possible, it is very critical to have a low and predictable latency to transmit video from the car cameras and receive control commands back. In this paper, we analyze the problem and present a communication and video streaming system that addresses the latency challenges and enables teleoperation of a real car over commercially available LTE network; demonstrating sub-50ms roundtrip latencies for 720p, 60FPS video, with average PSNR 36db.

cs.NI