Searcharxiv⌕ Search

arXiv subjects

Kyungyong Lee

Publications and source records attributed to Kyungyong Lee.

At least 19 recordsLinked to original sources

Latency Prediction for LLM Inference on NPU Systems

Deploying Large Language Models (LLMs) requires exploring a large configuration space spanning parallelization strategies, batching techniques, and scheduling policies. Exhaustive measurement across this space is impractical, making latency prediction essential for system optimization. While NPUs have emerged as accelerators designed for LLM inference, no prediction methodology has been established for them. Specifically, applying prior work to LLM inference latency prediction on NPUs faces three challenges: undisclosed microarchitecture of commercial NPUs, unpredictable compiler optimizations, and latency non-linearity induced by bucketing. We present LENS, a latency estimator that predicts NPU inference latency without information on the microarchitecture or compiler, and captures the non-linear latency induced by bucketing. LENS profiles each bucket with two end-to-end (E2E) measurements and composes the results to predict latency for arbitrary input-output length combinations. We validate LENS across NPUs from multiple vendors, several LLMs, and diverse workloads, achieving a mean prediction error of 2.15\%. We further compare LENS against two methodologically related baselines, confirming the validity of its approach.

cs.DC↗

ShuntServe: Cost-Efficient LLM Serving on Heterogeneous Spot GPU Clusters

As large language model (LLM) services become widely adopted, the cost of GPU resources for serving these models in cloud environments has emerged as a critical concern. Spot instances offer up to 90% cost savings over on-demand instances, but their frequent interruptions and limited availability pose significant challenges for continuous LLM serving. GPU spot instances, in particular, exhibit lower and more volatile availability than CPU-based instances, making homogeneous clusters that depend on a single GPU type vulnerable to correlated failures. Heterogeneous clusters spanning multiple GPU types can address this by leveraging complementary availability patterns across diverse spot pools, yet existing LLM serving systems are designed for homogeneous environments and suffer from load imbalance when deployed on heterogeneous GPUs. This paper presents ShuntServe, a cost-efficient LLM serving system for heterogeneous spot GPU clusters. ShuntServe employs a roofline model-based analytical serving performance estimator and a dynamic programming-based model placement optimizer that jointly determines node configuration, parallelization strategy, and layer assignment to maximize throughput across heterogeneous GPUs. To enhance fault tolerance when using spot instances, ShuntServe combines output-preserving request migration with concurrent initialization via a shared tensor store, minimizing migration downtime by overlapping replacement node preparation with ongoing serving. Evaluation on Llama-3.1-70B and Qwen3-32B with a heterogeneous AWS cluster of L4, A10G, and L40S GPUs shows that ShuntServe achieves 1.42x and 1.35x higher throughput than state-of-the-art baselines and attains 31.9% and 31.2% cost efficiency improvements over on-demand instances for offline and online serving, respectively.

cs.DC↗

KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances

Cloud users aim to minimize cost while maximizing performance by selecting the most suitable instance types for their workloads. To reduce expenses, spot instances have been widely adopted due to their steep discounts compared to on-demand pricing. However, their use introduces reliability risks due to potential interruptions, and existing research has primarily focused on mitigating this trade-off from a cost or availability perspective alone. Despite the diversity in hardware capabilities among instance types, current provisioning systems tend to ignore performance variation, selecting nodes solely based on minimum resource requirements. In this paper, we present KubePACS, a Kubernetes-native spot instance provisioning system that constructs node pools optimized for both cost and performance while guaranteeing high availability. KubePACS formulates the node selection process as a multi-objective optimization problem, incorporating real-time data such as spot prices, performance benchmarks, and availability scores, including the multi-node Spot Placement Score (SPS). It solves this problem efficiently using an Integer Linear Programming (ILP) approach guided by the Golden Section Search (GSS) algorithm to find the optimal configuration. By integrating with the Karpenter node autoscaler, KubePACS jointly optimizes instance-type selection and node scaling decisions within a standard provisioning workflow. KubePACS also adopts a novel heuristic to support workload-specific preferences by scaling performance metrics for specialized instances. Through extensive evaluation across synthetic and real-world workloads, KubePACS demonstrates on average 55.09% and up to 81.06% higher performance per dollar over state-of-the-art solutions such as Karpenter, SpotVerse, and SpotKube, which only reference the spot instance prices and limited availability data.

cs.DC↗

Ding-Dong Ditch: Peeking Into Spot Instance Availability

Spot instances offer significant cost savings of up to 90% over on-demand prices, making them an attractive resource for large-scale computing workloads. However, understanding their availability dynamics is essential for building systems that tolerate interruptions, and observing this availability directly requires keeping instances running, which incurs costs that scale with the number of monitored instance types and their per-instance price. We propose Ding-Dong Ditch (DDD), a cost-efficient method that collects spot instance availability signals by leveraging the cloud provider's provisioning lifecycle. Since the outcome of a spot request is determined before the instance enters the running state, DDD submits requests and cancels them upon provisioning acceptance, collecting binary availability signals at near-zero instance cost. Submitting multiple concurrent requests per measurement point further yields a quantitative estimate of available capacity. We validate DDD through simultaneous collection of probing signals and actual running instance traces across 68 instance types and 15 regions on both AWS and Azure, totaling 336,033 spot requests. Analysis of 2,635 real-world interruption events reveals that co-interruptions within the same instance type and availability zone occur within three minutes in over 92% of cases, motivating a binary availability formulation. Based on this formulation, we derive three complementary features from DDD signals and demonstrate that their combination achieves an F1-macro score of up to 0.90 for current availability modeling and maintains 0.85 at a 60-minute prediction horizon. A trace-driven simulation using TPC-DS workloads further demonstrates the potential of DDD-based prediction to reduce lost computation compared to an unguided baseline.

cs.DC↗

SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances

Cloud vendors offer discounted spot instances to maximize surplus resource utilization, but these instances are subject to the risk of sudden interruption. Traditional pricing datasets have been employed to predict this risk, yet recent policy changes by cloud vendors have diminished their effectiveness. To promote spot instance usage, public cloud vendors provide instant availability datasets to help users mitigate interruption risks. While existing research utilizing this data has proposed methods to reduce interruptions, these studies have primarily focused on single-node instances, overlooking the stability of multi-node environments widely adopted for modern cloud workloads. This paper proposes SpotVista, a system that recommends a resource pool of reliable and cost-efficient multi-node spot instances by leveraging various publicly available datasets. To achieve this, SpotVista collects a large-scale multi-node availability dataset while overcoming significant query limitations. Through a thorough analysis of multi-node spot instance availability behavior, SpotVista establishes a methodology for recommending cost-efficient and reliable multi-node configurations. To evaluate how effectively the proposed methodology reflects multi-node availability and cost efficiency, extensive real-world interruption experiments were conducted. The results demonstrate that SpotVista outperforms the state-of-the-art work, SpotVerse, achieving 81.28% greater availability and 2.84\% more cost savings in a multi-region setup. When compared to a publicly available service, AWS SpotFleet, SpotVista provides 21.6\% higher stability and 26.3% greater cost savings.

cs.DC↗

Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?

Failures in large-scale cloud systems incur substantial financial losses, making automated Root Cause Analysis (RCA) essential for operational stability. Recent efforts leverage Large Language Model (LLM) agents to automate this task, yet existing systems exhibit low detection accuracy even with capable models, and current evaluation frameworks assess only final answer correctness without revealing why the agent's reasoning failed. This paper presents a process level failure analysis of LLM-based RCA agents. We execute the full OpenRCA benchmark across five LLM models, producing 1,675 agent runs, and classify observed failures into 12 pitfall types across intra-agent reasoning, inter-agent communication, and agent-environment interaction. Our analysis reveals that the most prevalent pitfalls, notably hallucinated data interpretation and incomplete exploration, persist across all models regardless of capability tier, indicating that these failures originate from the shared agent architecture rather than from individual model limitations. Controlled mitigation experiments further show that prompt engineering alone cannot resolve the dominant pitfalls, whereas enriching the inter-agent communication protocol reduces communication-related failures by up to 15 percentage points. The pitfall taxonomy and diagnostic methodology developed in this work provide a foundation for designing more reliable autonomous agents for cloud RCA.

cs.AI↗

Cluster scattering diagrams via quiver moduli and tight gradings

We study rank-2 cluster scattering diagrams through moduli spaces of quiver representations and a recently developed combinatorial framework of tight gradings. Combining quiver-theoretic and combinatorial methods, we prove and extend a collection of conjectures posed by Elgin--Reading--Stella concerning the structural and enumerative properties of the wall-function coefficients. The tight grading perspective also provides a new proof of the Weyl group symmetry of the scattering diagram.

math.CO↗

Positivity of generalized cluster scattering diagrams

We introduce a new class of combinatorial objects, named tight gradings, which are certain nonnegative integer-valued functions on maximal Dyck paths. Using tight gradings, we derive a manifestly positive formula for any wall-function in a rank-2 generalized cluster scattering diagram. We further prove that any consistent rank-2 scattering diagram is positive with respect to the coefficients of initial wall-functions. Moreover, our formula yields explicit expressions for relative Gromov-Witten invariants on weighted projective planes and the Euler characteristics of moduli spaces of framed stable representations on complete bipartite quivers. Finally, by leveraging the rank-2 positivity, we show that any higher-rank generalized cluster scattering diagram has positive wall-functions, which leads to a proof of the positivity of the Laurent phenomenon and the strong positivity of Chekhov-Shapiro's generalized cluster algebras.

math.CO↗

Geometry of $C$-vectors and $C$-Matrices for Mutation-Infinite Quivers

The set of forks is a class of quivers introduced by M. Warkentin, where every connected mutation-infinite quiver is mutation equivalent to infinitely many forks. Let $Q$ be a fork with $n$ vertices, and $\boldsymbol{w}$ be a fork-preserving mutation sequence. We show that every $c$-vector of $Q$ obtained from $\boldsymbol{w}$ is a solution to a quadratic equation of the form $$\sum_{i=1}^n x_i^2 + \sum_{1\leq i<j\leq n} \pm q_{ij} x_i x_j =1,$$ where $q_{ij}$ is the number of arrows between the vertices $i$ and $j$ in $Q$. The same proof techniques implies that when $Q$ is a rank 3 mutation-cyclic quiver, every $c$-vector of $Q$ is a solution to a quadratic equation of the same form.

math.CO↗

Scattering diagrams, tight gradings, and generalized positivity

In 2013, Lee, Li, and Zelevinsky introduced combinatorial objects called compatible pairs to construct the greedy bases for rank-2 cluster algebras, consisting of indecomposable positive elements including the cluster monomials. Subsequently, Rupel extended this construction to the setting of generalized rank-2 cluster algebras by defining compatible gradings. We discover a new class of combinatorial objects which we call tight gradings. Using this, we give a directly computable, manifestly positive, and elementary but highly nontrivial formula describing rank-2 consistent scattering diagrams. This allows us to show that the coefficients of the wall-functions on a generalized cluster scattering diagram of any rank are positive, which implies the Laurent positivity for generalized cluster algebras and the strong positivity of their theta bases.

math.CO↗

An unexpected property of $\mathbf{g}$-vectors for rank 3 mutation-cyclic quivers

Let $Q$ be a rank 3 mutation-cyclic quiver. It is known that every $\mathbf{c}$-vector of $Q$ is a solution to a quadratic equation of the form $$\sum_{i=1}^3 x_i^2 + \sum_{1\leq i<j\leq 3} \pm q_{ij} x_i x_j =1,$$where $q_{ij}$ is the number of arrows between the vertices $i$ and $j$ in $Q$. A similar property holds for $\mathbf{c}$-vectors of any acyclic quiver. In this paper, we show that $\mathbf{g}$-vectors of $Q$ enjoy an unexpected property. More precisely, every $\mathbf{g}$-vector of $Q$ is a solution to a quadratic equation of the form $$\sum_{i=1}^3 x_i^2 + \sum_{1\leq i<j\leq 3} p_{ij} x_i x_j =1,$$where $p_{ij}$ is the number of arrows between the vertices $i$ and $j$ in another quiver $P$ obtained by mutating $Q$.

math.CO↗

On the two-dimensional Jacobian conjecture: Magnus' formula revisited, IV

Let $(F,G)$ be a Jacobian pair with $d=w\text{-deg}(F)$ and $e=w\text{-deg}(G)$ for some direction $w$. A generalized Magnus' formula approximates $G$ as $\sum_{γ\ge 0} c_γF^{\frac{e-γ}{d}}$ for some complex numbers $c_γ$. We develop an approach to the two-dimensional Jacobian conjecture, aiming to minimize the use of terms corresponding to $γ>0$. As an initial step in this approach, we define and study the inner polynomials of $F$ and $G$. The main result of this paper shows that the northeastern vertex of the Newton polygon of each inner polynomial is located within a specific region. As applications of this result, we introduce several conjectures and prove some of them for special cases.

math.AG↗

Broken lines and compatible pairs for rank 2 quantum cluster algebras

There have been several combinatorial constructions of universally positive bases in cluster algebras, and these same combinatorial objects play a crucial role in the known proofs of the famous positivity conjecture for cluster algebras. The greedy basis was constructed in rank $2$ by Lee-Li-Zelevinsky using compatible pairs on Dyck paths. The theta basis, introduced by Gross-Hacking-Keel-Kontsevich, has elements expressed as a sum over broken lines on scattering diagrams. It was shown by Cheung-Gross-Muller-Musiker-Rupel-Stella-Williams that these bases coincide in rank $2$ via algebraic methods, and they posed the open problem of giving a combinatorial proof by constructing a (weighted) bijection between compatible pairs and broken lines. We construct a quantum-weighted bijection between compatible pairs and broken lines for the quantum type $A_2$ and the quantum Kronecker cluster algebras. By specializing the quantum parameter, this handles the problem of Cheung et al. for skew-symmetric cluster algebras of finite and affine type. For cluster monomials in skew-symmetric rank-$2$ cluster algebras, we construct a quantum-weighted bijection between positive compatible pairs (which comprise almost all compatible pairs) and broken lines of negative angular momentum.

math.QA↗

An explicit description for quiver Grassmannians of the Kronecker quiver

It is an open problem to find cell decompositions of quiver Grassmannians associated to each cluster variable. We initiate a new approach to this problem by giving an explicit description for each individual subrepresentation. In this paper, we illustrate our approach for the Kronecker quiver.

math.RT↗

Geometric description of C-vectors and real Lösungen

We introduce real Loesungen as an analogue of real roots. For each mutation sequence of an arbitrary skew-symmetrizable matrix, we define a family of reflections along with associated vectors which are real Loesungen and a set of curves on a Riemann surface. The matrix consisting of these vectors is called L-matrix. We explain how the L-matrix naturally arises in connection with the C-matrix. Then we conjecture that the L-matrix depends (up to signs of row vectors) only on the seed, and that the curves can be drawn without self-intersections, providing a new combinatorial/geometric description of c-vectors.

math.RT↗

PROFET: Profiling-based CNN Training Latency Prophet for GPU Cloud Instances

Training a Convolutional Neural Network (CNN) model typically requires significant computing power, and cloud computing resources are widely used as a training environment. However, it is difficult for CNN algorithm developers to keep up with system updates and apply them to their training environment due to quickly evolving cloud services. Thus, it is important for cloud computing service vendors to design and deliver an optimal training environment for various training tasks to lessen system operation management overhead of algorithm developers. To achieve the goal, we propose PROFET, which can predict the training latency of arbitrary CNN implementation on various Graphical Processing Unit (GPU) devices to develop a cost-effective and time-efficient training cloud environment. Different from the previous training latency prediction work, PROFET does not rely on the implementation details of the CNN architecture, and it is suitable for use in a public cloud environment. Thorough evaluations reveal the superior prediction accuracy of PROFET compared to the state-of-the-art related work, and the demonstration service presents the practicality of the proposed system.

cs.DC↗

SpotLake: Diverse Spot Instance Dataset Archive Service

Public cloud service vendors provide a surplus of computing resources at a cheaper price as a spot instance. Despite the cheaper price, the spot instance can be forced to be shutdown at any moment whenever the surplus resources are in shortage. To enhance spot instance usage, vendors provide diverse spot instance datasets. Amon them, the spot price information has been most widely used so far. However, the tendency toward barely changing spot price weakens the applicability of the spot price dataset. Besides the price dataset, the recently introduced spot instance availability and interruption ratio datasets can help users better utilize spot instances, but they are rarely used in reality. With a thorough analysis, we could uncover major hurdles when using the new datasets concerning the lack of historical information, query constraints, and limited query interfaces. To overcome them, we develop SpotLake, a spot instance data archive web service that provides historical information of various spot instance datasets. Novel heuristics to collect various datasets and a data serving architecture are presented. Through real-world spot instance availability experiments, we present the applicability of the proposed system. SpotLake is publicly available as a web service to speed up cloud system research to improve spot instance usage and availability while reducing cost.

cs.DC↗

On the ordering of the Markov numbers

The Markov numbers are the positive integers that appear in the solutions of the equation $x^2+y^2+z^2=3xyz$. These numbers are a classical subject in number theory and have important ramifications in hyperbolic geometry, algebraic geometry and combinatorics. It is known that the Markov numbers can be labeled by the lattice points $(q,p)$ in the first quadrant and below the diagonal whose coordinates are coprime. In this paper, we consider the following question. Given two lattice points, can we say which of the associated Markov numbers is larger? A complete answer to this question would solve the uniqueness conjecture formulated by Frobenius in 1913. We give a partial answer in terms of the slope of the line segment that connects the two lattice points. We prove that the Markov number with the greater $x$-coordinate is larger than the other if the slope is at least $-\frac{8}{7}$ and that it is smaller than the other if the slope is at most $-\frac{5}{4}$. As a special case, namely when the slope is equal to 0 or 1, we obtain a proof of two conjectures from Aigner's book "Markov's theorem and 100 years of the uniqueness conjecture".

math.NT↗