Searcharxiv⌕ Search

arXiv subjects

Yunjia Zheng

Publications and source records attributed to Yunjia Zheng.

5 recordsLinked to original sources

To Store or To Regenerate? A Cost Model for AI-Generated Content at Scale

AI-generated content is becoming a rapidly growing class of digital artifacts. Because these artifacts accumulate over time, their exponential growth creates a substantial storage, energy, and infrastructure cost for operators and society. At the same time, GPU compute cost continues to fall rapidly with each hardware generation. This divergence raises a fundamental question: when does on-demand regeneration become cheaper than persistent storage? This paper develops a cost model for comparing persistent storage and on-demand regeneration for AI-generated artifacts. The model accounts for corpus growth, HDD and tape price trends, drive replacement, electricity, request skew, caching, generator FLOPs, and future GPU price-performance improvements. For image generation, our analysis shows that prompt-based regeneration does not become cheaper than storage until around 2040, because every cache miss must still rerun the full prompt-to-artifact generation pipeline. We observe that widely used diffusion-based generation models operate in latent space, creating an alternative point in the cost tradeoff: instead of storing the final artifact or only the prompt, operators can store a compact intermediate representation (IR) and perform cheap on-demand decoding. Our analysis shows that caching combined with IR-based regeneration substantially reduces both storage and compute cost, making it at least 2x cheaper than both full-object storage and prompt-based regeneration even today. On a production image trace with 2.07 billion requests, the same conclusion holds: prompt-based regeneration is over 100x more expensive than storage, while IR-based regeneration reduces total cost to roughly half that of full-object storage while preserving interactive miss latency.

cs.PF↗

SR-Gadgets: Make Scan-Resistant Caching Practical

Block caches commonly serve scan-heavy I/O workloads, motivating extensive studies on scan-resistant eviction algorithms. Many of these algorithms adopt a multi-queue structure. However, they focus primarily on one-time scans and do not handle repeated scans well. Two important challenges from repeated scans are miss-ratio cliffs, where a small increase in cache size sharply reduces the miss ratio, and Belady's anomalies, where increasing the cache size increases the miss ratio. In this paper, we first develop two quantitative metrics to measure these behaviors. With these metrics, we find that LIRS is the only multi-queue algorithm that is scan-resistant (almost cliff- and anomaly-free). Contrary to conventional wisdom, we show that stack distance is not the secret sauce that makes LIRS scan-resistant. Instead, regulating the queues are the key to its scan resistance. Based on these insights, we design the \gadgetprefix Gadgets, easy-to-integrate augmentations that make existing algorithms scan-resistant without changing their eviction heuristics or queue structures. We implement the \gadgetprefix Gadgets in five algorithms: S3-FIFO, SIEVE, ARC, 2Q, and TinyLFU, and make them scan-resistant. Evaluated on 5,538 production traces, all augmented algorithms outperform their base versions, reducing miss ratios by up to 23.1% while consistently reducing cliffs and Belady's anomalies across the production traces.

cs.PF↗

TrajectoryDB: A New Database for Agent Trajectories

AI agents generate rich execution trajectories that capture their interactions with large language models, tools, and external environments. These trajectories are increasingly valuable for downstream tasks such as memory extraction, model fine-tuning, runtime optimization, and security and cost monitoring. Yet trajectory data today is fragmented across files, databases, and observability systems, with no persistent data management system designed around its unique structure and access patterns. We argue that trajectories should be treated as a distinct data type. A trajectory combines hierarchical execution structure, large volumes of text whose analysis often requires semantic reasoning, and rich dependencies and lineage among events, intermediate states, and derived artifacts. These properties introduce new requirements throughout the data lifecycle. Ingestion must reconstruct and preserve execution structure and lineage; storage must efficiently organize large but highly redundant contexts while maintaining relationships among records; and query processing must jointly reason over structure, temporal order, semantics, and lineage. We therefore envision TrajectoryDB, a trajectory-native data management system that co-designs ingestion, storage, and query processing to efficiently manage and analyze agent execution trajectories.

cs.DB↗

LatentBox: Storing AI-Generated Images at Scale via a Latent-First Design

The explosive growth of AI-generated images has created a sustainability challenge for storage infrastructure. Platforms like Midjourney and Adobe Firefly already host billions of generative images, yet conventional object stores persist them as blobs with full-resolution pixels, consuming huge amounts of storage capacity and bandwidth. Unlike natural photos, however, AI-generated images can be deterministically reconstructed from compact, model-native latent tensors, making persistent image storage fundamentally redundant. This paper presents LatentBox, a latent-first storage system for AI-generated images. LatentBox treats compressed latents as durable storage objects and uses on-demand GPU reconstruction on the read path to trade inexpensive compute for large persistent storage savings. Our design is guided by the first large-scale analysis of AI-generated image access we are aware of, based on a 35-month, 2-billion-request production trace from a major generative-content platform. Motivated by the trace analysis, LatentBox keeps frequently accessed images in decoded pixel format for fast hits, stores less-active objects as compressed latents to expand effective cache capacity, and continuously adjusts the splits between the image and latent cache to optimize user-perceived access latency.We build a LatentBox prototype and evaluate it with the production trace. LatentBox reduces persistent storage by 78.7% with competitive or even lower mean and tail latency over a pure image-based storage.

cs.DC↗

TStore: Rethinking AI Model Hub with Tensor-Centric Compression

Modern AI models are growing rapidly in size and redundancy, leading to significant storage and distribution challenges in model hubs. We present TStore, a tensor-centric system for reducing storage overhead through fine-grained deduplication and compression. TStore leverages tensor-level fingerprinting and clustering to identify redundancy across models without requiring annotations. Our design enables efficient storage reduction while preserving model usability and performance. Experiments on real-world model repositories demonstrate substantial storage savings with minimal overhead.

cs.DC↗