SearcharxivSearch

arXiv subjects

Shamsher Khan

Publications and source records attributed to Shamsher Khan.

3 recordsLinked to original sources

Operational Memory Architecture for Kubernetes: Evidence Horizon Taxonomy and Extended Causal Pattern Preservation

Kubernetes clusters generate rich operational events during pod lifecycle transitions, yet the platform's native event retention model systematically discards the most diagnostically valuable context through multiple evidence destruction mechanisms operating on deterministic schedules. We formalize these mechanisms as an evidence horizon taxonomy: five distinct boundaries after which specific categories of diagnostic context become permanently unrecoverable from the Kubernetes API. H1(LastTerminationState rotation, ~90s) destroys container failure forensics; H2 (scheduler event pruning, 1hr/1000-event cluster limit) destroys placement rationale; H3 (ephemeral container exit, immediate) destroys debug session context; H4 (kubelet reconciliation gap) destroys in-memory operational state; and H5 (scrape-interval blind spot) renders sub-interval pod lifetimes invisible to poll-based observability tools. This paper extends the Operational Memory Architecture (OMA) to address the full evidence horizon taxonomy. Two new causal patterns are defined: P004 (Scheduler Decision Provenance) captures FailedScheduling predicate failures before kube-apiserver TTL pruning and demonstrates the first cross-horizon causal chain linking scheduler evidence to downstream OOMKill failures. Two new Go watchers (EventWatcher, EphemeralWatcher) and two new storage tables (scheduler_events, ephemeral_exits) extend the original architecture. Validated on Minikube (3-node) and AKS 1.32.10. The original 30-run statistical latency analysis (242 edges, intra-cycle mean 0.702ms) and stress evaluation (2.86 events/sec at 20 pods, 8.8MB RAM) are carried forward and augmented with H2, H3, and H5 results.

cs.DC

Operational Memory Architecture for Kubernetes:Preserving Causal Context Across the Evidence Horizon

Kubernetes clusters generate rich operational events during pod lifecycle transitions, yet the platform's native event retention model discards the most diagnostically valuable context. The LastTerminationState field, which records a container's last failure, is overwritten shortly after a pod restart. We define this as the evidence horizon. During high-frequency crash loops, this horizon may be crossed multiple times before inspection, permanently losing critical evidence. This paper introduces the Operational Memory Architecture (OMA) to preserve causal failure evidence before event rotation. OMA encodes evidence retention and causal reconstruction as explicit architectural requirements. It captures operational events into causal chains using three patterns: P001 (OOMKill chain), P002 (ConfigMap variable misconfiguration), and P003 (ConfigMap volume mount propagation). We implement OMA as an open-source system with a Go-based Kubernetes watcher, SQLite operational memory store, and a simple query interface. Experiments on Minikube and AKS include a 30-run latency analysis and stress tests with up to 20 crash-looping pods. Causal edges are built with mean latency below 1 ms. The collector processes ~2.8 events/sec while using under 10 MB memory, showing minimal overhead and effective evidence preservation.

cs.DC

Decomposing Docker Container Startup Performance: A Three-Tier Measurement Study on Heterogeneous Infrastructure

Container startup latency is a critical performance metric for CI/CD pipelines, serverless computing, and auto-scaling systems, yet practitioners lack empirical guidance on how infrastructure choices affect this latency. We present a systematic measurement study that decomposes Docker container startup into constituent operations across three heterogeneous infrastructure tiers: Azure Premium SSD (cloud SSD), Azure Standard HDD (cloud HDD), and macOS Docker Desktop (developer workstation with hypervisor-based virtualization). Using a reproducible benchmark suite that executes 50 iterations per test across 10 performance dimensions, we quantify previously under-characterized relationships between infrastructure configuration and container runtime behavior. Our key findings include: (1) container startup is dominated by runtime overhead rather than image size, with only 2.5% startup variation across images ranging from 5 MB to 155 MB on SSD; (2) storage tier selection imposes a 2.04x startup penalty (HDD 1157 ms vs. SSD 568 ms); (3) Docker Desktop's hypervisor layer introduces a 2.69x startup penalty and 9.5x higher CPU throttling variance compared to native Linux; (4) OverlayFS write performance collapses by up to two orders of magnitude compared to volume mounts on SSD-backed storage; and (5) Linux namespace creation contributes only 8-10 ms (<1.5%) of total startup time. All measurement scripts, raw data, and analysis tools are publicly available.

cs.PF