SearcharxivSearch

arXiv subjects

Rain Jiang

Publications and source records attributed to Rain Jiang.

12 recordsLinked to original sources

Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference

Long-running LLM agents keep valuable state resident on GPUs: KV caches, request schedulers, communication state, and sometimes online adapters. Losing this state after a GPU or communicator failure can discard minutes to hours of work, yet existing recovery mechanisms either restart the whole serving stack or require application-specific checkpoint logic inside every attention and runtime component. This paper argues that fault tolerance for such workloads needs a GPU-resident execution context: checkpoint hooks must run at device synchronization points, observe binary kernels that frameworks and libraries actually execute, and recover without putting the host CPU on the critical path. We present Concordia, a runtime that uses a device-resident persistent kernel as the substrate for fault-tolerant LLM inference. Concordia interposes on GPU module loading and supports PTX- and SASS-level instrumentation, allowing checkpoint and pause hooks to be inserted below framework code and library boundaries. For each registered LLM state region, Concordia JIT-compiles a specialized delta-checkpoint handler -- for example, a KV-block scanner, adapter-page scanner, or recovery applier -- and hot-swaps it into the persistent kernel's operator table. The persistent kernel consumes a lock-free ring buffer of compute, checkpoint, append-log, and recovery tasks, so the same always-on executor triggers dirty-page detection, stages deltas, and appends committed records to a CPU-visible log in CXL memory or host DRAM.

cs.DC

KVDirect: Distributed Disaggregated LLM Inference

Large Language Models (LLMs) have become the new foundation for many applications, reshaping human society like a storm. Disaggregated inference, which separates prefill and decode stages, is a promising approach to improving hardware utilization and service quality. However, due to inefficient inter-node communication, existing systems restrict disaggregated inference to a single node, limiting resource allocation flexibility and reducing service capacity. This paper introduces KVDirect, which optimizes KV cache transfer to enable a distributed disaggregated LLM inference. KVDirect achieves this through the following contributions. First, we propose a novel tensor-centric communication mechanism that reduces the synchronization overhead in traditional distributed GPU systems. Second, we design a custom communication library to support dynamic GPU resource scheduling and efficient KV cache transfer. Third, we introduce a pull-based KV cache transfer strategy that reduces GPU resource idling and improves latency. Finally, we implement KVDirect as an open-source LLM inference framework. Our evaluation demonstrates that KVDirect reduces per-request latency by 55% compared to the baseline across diverse workloads under the same resource constraints.

cs.DC

Partial order alignment by adjacencies and breakpoints

Linearizing two partial orders to maximize the number of adjacencies and minimize the number of breakpoints is APX-hard. This holds even if one of the two partial orders is already a linear order and the other is an interval order, or if both partial orders are weak orders.

cs.CC

Decomposing a graph into subgraphs with small components

The component size of a graph is the maximum number of edges in any connected component of the graph. Given a graph $G$ and two integers $k$ and $c$, $(k,c)$-Decomposition is the problem of deciding whether $G$ admits an edge partition into $k$ subgraphs with component size at most $c$. We prove that for any fixed $k \ge 2$ and $c \ge 2$, $(k,c)$-Decomposition is NP-complete in bipartite graphs. Also, when both $k$ and $c$ are part of the input, $(k,c)$-Decomposition is NP-complete even in trees. Moreover, $(k,c)$-Decomposition in trees is W[1]-hard with parameter $k$, and is FPT with parameter $c$. In addition, we present approximation algorithms for decomposing a tree either into the minimum number of subgraphs with component size at most $c$, or into $k$ subgraphs minimizing the maximum component size. En route to these results, we also obtain a fixed-parameter algorithm for Bin Packing with the bin capacity as parameter.

cs.CC

Vertebrate interval graphs

A vertebrate interval graph is an interval graph in which the maximum size of a set of independent vertices equals the number of maximal cliques. For any fixed $v \ge 1$, there is a polynomial-time algorithm for deciding whether a vertebrate interval graph admits a vertex partition into two induced subgraphs with claw number at most $v$. In particular, when $v = 2$, whether a vertebrate interval graph can be partitioned into two proper interval graphs can be decided in polynomial time.

math.CO

Partitioning an interval graph into subgraphs with small claws

The claw number of a graph $G$ is the largest number $v$ such that $K_{1,v}$ is an induced subgraph of $G$. Interval graphs with claw number at most $v$ are cluster graphs when $v = 1$, and are proper interval graphs when $v = 2$. Let $κ(n,v)$ be the smallest number $k$ such that every interval graph with $n$ vertices admits a vertex partition into $k$ induced subgraphs with claw number at most $v$. Let $\checkκ(w,v)$ be the smallest number $k$ such that every interval graph with claw number $w$ admits a vertex partition into $k$ induced subgraphs with claw number at most $v$. We show that $κ(n,v) = \lfloor\log_{v+1} (n v + 1)\rfloor$, and that $\lfloor\log_{v+1} w\rfloor + 1 \le \checkκ(w,v) \le \lfloor\log_{v+1} w\rfloor + 3$. Besides the combinatorial bounds, we also present a simple approximation algorithm for partitioning an interval graph into the minimum number of induced subgraphs with claw number at most $v$, with approximation ratio $3$ when $1 \le v \le 2$, and $2$ when $v \ge 3$.

math.CO

Caterpillars and alternating paths

Let $p(m)$ (respectively, $q(m)$) be the maximum number $k$ such that any tree with $m$ edges can be transformed by contracting edges (respectively, by removing vertices) into a caterpillar with $k$ edges. We derive closed-form expressions for $p(m)$ and $q(m)$ for all $m \ge 1$. The two functions $p(n)$ and $q(n)$ can also be interpreted in terms of alternating paths among $n$ disjoint line segments in the plane, whose $2n$ endpoints are in convex position.

math.CO

Moving intervals for packing and covering

We study several problems on geometric packing and covering with movement. Given a family $\mathcal{I}$ of $n$ intervals of $κ$ distinct lengths, and another interval $B$, can we pack the intervals in $\mathcal{I}$ inside $B$ (respectively, cover $B$ by the intervals in $\mathcal{I}$) by moving $τ$ intervals and keeping the other $σ= n - τ$ intervals unmoved? We show that both packing and covering are W[1]-hard with any one of $κ$, $τ$, and $σ$ as single parameter, but are FPT with combined parameters $κ$ and $τ$. We also obtain improved polynomial-time algorithms for packing and covering, including an $O(n\log^2 n)$ time algorithm for covering, when all intervals in $\mathcal{I}$ have the same length.

cs.CG