arXiv · 2602.22434
GetBatch: Distributed Multi-Object Retrieval for ML Data Loading
Abstract
Machine learning training pipelines consume data in batches. A single training step may require thousands of samples drawn from shards distributed across a storage cluster. Issuing thousands of individual GET requests incurs per-request overhead that often dominates data transfer time. To solve this problem, we introduce GetBatch - a new object store API that elevates batch retrieval to a first-class storage operation, replacing independent GET operations with a single deterministic, fault-tolerant streaming execution. GetBatch achieves up to 15x throughput improvement for small objects and, in a production training workload, reduces P95 batch retrieval latency by 2x and P99 per-object tail latency by 3.7x compared to individual GET requests.
Explore related subjects
Keep this discovery
Alex Aizman, Abhishek Gaikwad, Piotr Żelasko. 2026-02-25. GetBatch: Distributed Multi-Object Retrieval for ML Data Loading. https://arxiv.org/abs/2602.22434
Cite the original work for its findings. Save a collection to share your selection of sources.