Searcharxiv⌕ Search

arXiv · 2609.30448

To Store or To Regenerate? A Cost Model for AI-Generated Content at Scale

Abstract

AI-generated content is becoming a rapidly growing class of digital artifacts. Because these artifacts accumulate over time, their exponential growth creates a substantial storage, energy, and infrastructure cost for operators and society. At the same time, GPU compute cost continues to fall rapidly with each hardware generation. This divergence raises a fundamental question: when does on-demand regeneration become cheaper than persistent storage? This paper develops a cost model for comparing persistent storage and on-demand regeneration for AI-generated artifacts. The model accounts for corpus growth, HDD and tape price trends, drive replacement, electricity, request skew, caching, generator FLOPs, and future GPU price-performance improvements. For image generation, our analysis shows that prompt-based regeneration does not become cheaper than storage until around 2040, because every cache miss must still rerun the full prompt-to-artifact generation pipeline. We observe that widely used diffusion-based generation models operate in latent space, creating an alternative point in the cost tradeoff: instead of storing the final artifact or only the prompt, operators can store a compact intermediate representation (IR) and perform cheap on-demand decoding. Our analysis shows that caching combined with IR-based regeneration substantially reduces both storage and compute cost, making it at least 2x cheaper than both full-object storage and prompt-based regeneration even today. On a production image trace with 2.07 billion requests, the same conclusion holds: prompt-based regeneration is over 100x more expensive than storage, while IR-based regeneration reduces total cost to roughly half that of full-object storage while preserving interactive miss latency.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yunjia Zheng, Zirui Wang, Haoran Ni, Tingfeng Lan, Zhaoyuan Su, Yue Cheng, Juncheng Yang. 2026-09-24. To Store or To Regenerate? A Cost Model for AI-Generated Content at Scale. https://arxiv.org/abs/2609.30448

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking

As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide promising means to navigate this complexity, these techniques remain computationally intensive and require many evaluations to find optimal configurations. This work proposes an autotuning framework that designs a machine learning-based ensemble LLVM Intermediate Representa- tion (IR) ranker, Neural Configuration Scorer (NCS). NCS ranks the performance of IRs sampled by a transfer-learning-based autotuner, improving the efficiency of the tuning process by reducing tuning overheads and circumventing subpar evaluations. By leveraging knowledge from related tasks, we are able to effectively exploit the transfer relationship to access high-performing configurations in fewer samples than traditional techniques that rely upon itera- tive refinement. Our framework can achieve similar performance improvements as state-of-the-art autotuning techniques with up to 61.67% fewer evaluations, averaging 27.85% fewer evaluations across various HPC benchmarks.

cs.PF↗

Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection

In agentic LLM services, a session calls an external tool, waits for it, and re-arrives to continue inference. Statically compiled NPU serving can fix the batch bucket set, the maximum batch size, and the number of KV cache slots at compile time. We define such an environment as a compile-time-static serving substrate and analyze the execution-time cost that tool waiting and re-arrival incur in it. On a single LLM instance, we run synthetic workloads following a measured tool waiting time distribution and compare, on the same inputs, a baseline configuration with settings {1, 2, 4, 8}, 8, and 8 against configurations that change some of them. Because re-arrival times differ across configurations, we build a simulator that replays request processing in time order, select the candidate with the lowest predicted cost among 2,077 configurations, and validate it on new inputs. We identify three mechanisms: discrete batch alignment, KV cache survival, and prefill interference. At a concurrency of 6, absent from the bucket set, tool waiting lowered the padding ratio (0.235 to 0.120) yet increased decode execution time 1.51-fold, so padding alone did not indicate cost. On new inputs at a concurrency of 8, where the baseline reused KV in 9 of 24 re-arrivals, enlarging the maximum batch size alone cut execution cost by 8.25%, and the selected configuration, which also adjusted the bucket set, by 9.72%. Where 17 of 18 re-arrivals were already reused, the effect was 0.59%. Compile-time configurations should thus be selected by diagnosing KV reuse loss and the resulting change in execution.

cs.PF↗

PASCAL: A Progress Divergence-Aware Shared-Cache Model

In modern AI accelerators and GPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query (Q) tiles share the same key and value (K/V) blocks, GEMM, where every tile in a row reads the same slice, and many other operators. We name this pattern shared cyclic scan. Due to a significant amount of data reuse in this pattern, the cache is expected to capture as much data reuse as possible and largely reduce requests sent to the main memory for both performance and energy consumption concerns. However, in reality, because of the intrinsic asynchrony of multi-cores, the actual cache miss rate and DRAM traffic can be much higher compared to ideal cases. In this paper, we propose PASCAL, a shared-cache model for shared cyclic scans. It calibrates finite-run traffic, which reflects progress divergence, on reference configurations and interpolates it along static program structure such as occupancy to predict the miss rate before execution. PASCAL supports software configuration exploration without target traces or counters at scales where cycle-accurate simulation is impractical, while its analysis gives a sharp $2σ-1$ sufficient capacity condition for preserving LRU sharing. Across held-out scans and GEMM on GB10 and Thor, PASCAL reaches 12.04% balanced fill-equivalent miss-rate MAPE. Replacing TileSight's cache component with PASCAL lowers GB10 GEMM latency MAPE from 18.17% to 13.14%.

cs.PF↗