Searcharxiv⌕ Search

arXiv subjects

Georgii Kliukovkin

Publications and source records attributed to Georgii Kliukovkin.

2 recordsLinked to original sources

The Lazy Pod That Lies: Deferred Cost and Failure Semantics of Lazy Container Image Pulling for Model Serving on Kubernetes

Lazy container-image pulling promises to eliminate the dominant cost of starting a model-serving pod by mounting the image immediately and fetching content on demand. We evaluate this promise for model delivery on Kubernetes, using KServe with two production lazy-pulling systems -- eStargz/stargz-snapshotter and AWS SOCI -- against eager baselines, on artifacts from 2 to 140 GB including real fp16 weights. Lazy pulling delivers its headline: cold time-to-first-prediction becomes size-independent (16.9--17.6s, versus 24.5--573.0s eager). But the cost is deferred, not eliminated: a full read of a 14 GB model through the lazy mount takes 105.3s, slower than the 72.4s eager pull it replaced, and the two systems pay at opposite lifecycle ends (SOCI prefetches nearly the full image before Ready; eStargz defers nearly everything to first read). More consequentially, we characterize a failure mode eager pulling structurally cannot exhibit: under sustained legitimate reads with default configuration, the snapshotter's node-level cache exhausts its finite volume and already-running pods begin failing reads of model files. At the earliest stage of exhaustion, an instrumented serving pod passed every Kubernetes-visible and application-level check for 196s while its snapshotter was already logging real failures; under heavier pressure, 67--94% of model files fail, scaling monotonically with residual cache occupancy. A live pod self-heals if cache space is freed under it, but a snapshotter-daemon restart under a live pod leaves permanently stale file handles in a pod still reported Running. We derive placement, monitoring, and cache-sizing guidance for serving platforms and operators.

cs.DC↗

Cold-Start Model Delivery in Kubernetes Inference Serving: An Empirical Study of OCI-Based Distribution and Its Integrity

The startup latency of a model-serving pod on Kubernetes is dominated by one step: delivering the model weights. As models reach the hundred-gigabyte weights of large language models, cold-start delivery time governs the economics of autoscaling and scale-to-zero, yet the dominant mechanisms remain ad-hoc downloads from object storage, with none of the pull caching, digest addressing, or verification Kubernetes provides for container images. We analyze the delivery paths available to a Kubernetes serving platform along two axes: which component pulls the artifact, and whether any admission-time verifier can bind the deployed reference to the arriving bytes. We validate the analysis upstream in KServe, a widely deployed CNCF model-serving platform, by implementing two new delivery paths: oci+native://, which mounts model images as Kubernetes image volumes (KEP-4639), merged upstream, and oci+fetch://, which pulls OCI artifacts inside the storage initializer, under review. We report, to our knowledge, the first controlled comparison of model delivery paths in a Kubernetes serving platform (modelcar sidecars, native image volumes, object-storage download) on artifacts sized to fp16 weights of 1B-, 7B-, and 70B-class models (2-140 GB). Node-cached OCI delivery makes warm replica addition size-independent: 11.7 s for a 70B-class artifact versus 40.7 minutes of re-download over object storage, a 208x difference, while the first cold pull costs up to 2x a plain download, localized to containerd's blob-write-then-unpack double pass. For models on s3://, gs://, or hf:// URIs, where no admission-time verifier observes the bytes, we present a serving-time integrity design proposed to the KServe community: digest pinning and OpenSSF model-signing enforcement in the storage initializer. Streaming hash verification during download adds under 0.1% to delivery time; a post-download pass adds up to 53%.

cs.DC↗