SearcharxivSearch

arXiv subjects

Jianchang Su

Publications and source records attributed to Jianchang Su.

7 recordsLinked to original sources

Characterizing and Bridging the Diagnostic Gap in eBPF Verifier Rejections

eBPF lets developers run custom programs inside the Linux kernel, where a verifier proves each program safe. However, when the verifier rejects a program, the unclear error makes repair challenging: the error reports where verification stopped, not where the program lost the proof the verifier required. To quantify this gap, we conduct an empirical study of 235 reproduced rejections, showing that 47% of rejections return only EINVAL, one error string maps to as many as nine distinct root causes, and 10 of the 12 root causes are eBPF-specific. Repair thus requires both domain knowledge and locating where the proof was lost, yet existing tools only help developers read the error. We present bpfix, which reconstructs where the required proof was established and where it was lost from the verifier log, and prints a Rust-like diagnostic. To evaluate bpfix and the ability of LLMs to help repair, we construct a benchmark of 75 LLM repair tasks. Current models achieve 0-37% one-shot success with the raw log, and replacing the log with the bpfix localization improves repair by 11-21pp, suggesting that locating where the proof was lost is key to guiding repair. bpfix is available at https://github.com/eunomia-bpf/bpfix

cs.OS

ReFlux: Reversible Compute Placement for CXL-Enabled Storage

Static offload to computational storage devices proves brittle because device-side processors throttle under sustained thermal load, while opaque, vendor-specific interfaces inflate adoption costs so severely that no computational storage platform has achieved broad deployment; to address this, we argue that storage-side compute should be reversible, allowing individual pipeline stages to migrate between host and device at runtime beneath standard interfaces that require zero application modification. We present ReFlux, which realizes this principle on CXL SSDs by decomposing I/O-path logic into migratable storage actors compiled to WebAssembly, with actors sharing state through coherent CXL.mem regions so that only lightweight control state, roughly 8 KB, moves during migration, while a thermal-aware scheduler triggers per-stage drain-and-switch when device temperature or queue pressure rises, offloading compute-intensive actors to the host while leaving I/O-bound stages on the device. In our evaluation on an FPGA-based CXL SSD prototype and two production CSDs, ReFlux sustains over twice the throughput of thermally throttled CSDs under 30-minute sustained writes, delivers three to four times the inference throughput under KV-cache pressure, and reduces host CPU utilization by 65% through MWAIT-based notification.

cs.OS

SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning

Multi-stage ML inference pipelines are difficult to autoscale due to heterogeneous resources, cross-stage coupling, and dynamic bottleneck migration. We present SAIR, an autoscaling framework that uses an LLM as an in-context reinforcement learning controller, improving its policy online from reward-labeled interaction histories without gradient updates. SAIR combines Pareto-dominance reward shaping with a provable separation margin, surprisal-guided experience retrieval for context efficiency, and fine-grained GPU rate control via user-space CUDA interception. We provide regret analysis decomposing error into retrieval coverage and LLM selection components. On four ML serving pipelines under three workload patterns, SAIR achieves the best or tied-best P99 latency and effective resource cost among deployed baselines, improving P99 by up to 50% and reducing effective cost by up to 97% (under GPU rate-control assumptions), with 86% bottleneck detection accuracy and no offline training.

cs.LG

gpu_ext: Extensible OS Policies for GPUs via eBPF

Performance in modern GPU-centric systems increasingly depends on resource management policies, including memory placement, scheduling, and observability. However, uniform policies typically yield suboptimal performance across diverse workloads. Existing approaches present a tradeoff: user-space runtimes provide programmability and flexibility but lack cross-tenant visibility and fine-grained control of hardware resources; meanwhile, modifications to the OS kernel introduce significant complexity and safety risks. To address this, we argue that the GPU driver and device layer should provide an extensible OS interface for policy enforcement. While the emerging eBPF technology shows potential, directly applying existing host-side eBPF is insufficient because they lack visibility and control into critical device-side events, and directly embedding policy code into GPU kernels could compromise safety and efficiency. We propose gpu_ext, an eBPF-based runtime that treats the GPU driver and device as a programmable OS subsystem. gpu_ext extends GPU drivers by exposing safe programmable hooks and introduces a device-side eBPF runtime capable of executing verified policy logic within GPU kernels, enabling coherent and transparent policies. Evaluation across realistic workloads including inference, training, and vector search demonstrates that gpu_ext improves throughput by up to 4.8x and reduces tail latency by up to 2x, incurring low overhead, without modifying or restarting applications

cs.OS

PromptSculptor: Multi-Agent Based Text-to-Image Prompt Optimization

The rapid advancement of generative AI has democratized access to powerful tools such as Text-to-Image models. However, to generate high-quality images, users must still craft detailed prompts specifying scene, style, and context-often through multiple rounds of refinement. We propose PromptSculptor, a novel multi-agent framework that automates this iterative prompt optimization process. Our system decomposes the task into four specialized agents that work collaboratively to transform a short, vague user prompt into a comprehensive, refined prompt. By leveraging Chain-of-Thought reasoning, our framework effectively infers hidden context and enriches scene and background details. To iteratively refine the prompt, a self-evaluation agent aligns the modified prompt with the original input, while a feedback-tuning agent incorporates user feedback for further refinement. Experimental results demonstrate that PromptSculptor significantly enhances output quality and reduces the number of iterations needed for user satisfaction. Moreover, its model-agnostic design allows seamless integration with various T2I models, paving the way for industrial applications.

cs.MA

Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference

Global KV-cache sharing is an effective optimization for accelerating large language model (LLM) inference, yet it introduces an API-visible timing side channel that lets adversaries infer sensitive user inputs from shared entries, leading to cross-tenant privacy risks. To address this problem, we introduce SafeKV (Secure and Flexible KV-cache Sharing), a system-level co-design of privacy enforcement and KV-cache management. SafeKV integrates lightweight detection and isolation directly into the serving runtime to eliminate cross-tenant reuse of sensitive KV-cache blocks under our threat model, while recovering most of the performance benefits of global sharing. Our key contributions are: (1) a three-tier asynchronous detection pipeline that decouples privacy classification from inference and supports streaming workloads, (2) a unified radix-tree-based memory manager with path compression and sensitivity-aware eviction for scalable selective isolation, and (3) an RDR-guided (Reuse Diversity Ratio) runtime safeguard that detects and bounds residual leakage. On large LLM backends, SafeKV reduces the time-to-first-token (TTFT) overhead compared to full isolation by up to 40.58% and raises throughput by up to 2.66x. Overall, SafeKV restores the efficiency of KV reuse while enforcing strong, practical privacy for multi-tenant LLM inference.

cs.CR

CXLMemUring: A Hardware Software Co-design Paradigm for Asynchronous and Flexible Parallel CXL Memory Pool Access

CXL-attached memory lets servers add more memory while keeping the standard load/store programming model. The main drawback is latency. CXL memory accesses are too slow for normal CPU mechanisms to hide reliably, especially when each access depends on the result of a previous one. At the same time, they are too fast for traditional software techniques, such as context switches or interrupt-based asynchrony, to manage one load at a time. On a real Granite Rapids CXL platform, we find that placing GAPBS graph workloads in CXL memory slows execution by 2.44x on average compared with local DRAM. A state-of-the-art software prefetcher still leaves a 2.21x slowdown. This paper presents our system, a hardware/software co-designed approach for hiding CXL latency using larger units of asynchronous work. The system creates regions: parts of the original program that include one or more CXL-resident memory operations together with the nearby address-generation and memory-orchestration logic needed to run them. The host CPU launches these regions asynchronously on a near-memory accelerator built on a commodity CXL Type 2 FPGA, then continues useful work while the device runs ahead and prepares future memory accesses. A compiler identifies candidate regions from unmodified source programs, and an online JIT refines region boundaries and execution parameters based on workload behavior. We implement the system as a prototype compiler, runtime, and Vortex-based CXL-side accelerator. Across GAPBS, MCF, Spatter, and NAS Parallel Benchmark workloads, it improves end-to-end performance by 1.45x to 1.75x, with a 1.59x geometric mean speedup, compared with native CXL-memory execution.

cs.AR