Searcharxiv⌕ Search

arXiv · 2610.08483

Trail: Scalable and Low-Cost Temporal TLB Prefetching via Page-Table-Embedded Deltas

Abstract

Address translation is a major bottleneck in data-intensive workloads. TLB prefetching can hide translation latency, but existing spatial prefetchers struggle with irregular accesses, while temporal prefetchers store address deltas in fixed-capacity hardware that cannot scale with application memory footprints. Our characterization of 200 translation-intensive workloads reveals that each virtual memory region has a small, recurring set of deltas, and over 94% of deltas fit in 18 signed bits. We introduce Trail, a temporal TLB prefetcher that stores deltas in unused bits of leaf page table entries (PTEs). When a region triggers a page table walk, Trail identifies the region that last triggered a walk from the same instruction and records their delta in that source region's PTE. When a later walk fetches the source region's PTE cache block, Trail retrieves its deltas without additional memory accesses and prefetches translations for likely destination regions into the TLBs and cache hierarchy. Storing multiple deltas per PTE cache block improves coverage, while using existing PTE bits allows metadata capacity to scale with the application's memory footprint without additional metadata storage. Across 200 workloads and 100 multiprogrammed mixes, Trail improves single-core (four-core) performance by 5.7% (11.5%) on average over a baseline without TLB prefetching, outperforming the best prior standalone TLB prefetcher by 1.7% (2.5%). Trail requires only a 64-entry hardware table to track per-instruction page table walk history. Trail is freely available at https://github.com/CMU-SAFARI/Virtuoso/tree/trail-artifact-release.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Konstantinos Kanellopoulos, Konstantinos Sgouras, Harsh Songara, Rahul Bera, Onur Mutlu. 2026-10-06. Trail: Scalable and Low-Cost Temporal TLB Prefetching via Page-Table-Embedded Deltas. https://arxiv.org/abs/2610.08483

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An Adaptive Heterogeneous Architecture for High-Ratio, High-Throughput Lossless Compression

Modern lossless compression enforces an acute dichotomy: industrial streaming codecs (e.g., Zstandard) prioritize throughput (10 to 1000 MiB/s) at the expense of ratio, while context-mixing algorithms achieve superior density at serial speeds (0.1 to 1 MiB/s). We present GPX, an adaptive multi-tier compression architecture for structured data streams. GPX operates transparently across arbitrary stream lengths (N >= 1,024 B, identity fallback below 1 KiB), combining an autonomous structural decision probe (6.7--486.2 us across 49 benchmark instances, median 31.6 us, <0.05% overhead), cache-tiled AVX2 SIMD reversible domain transforms, multi-hypothesis sequence optimization with in-flight frame MDL gating, and adaptive entropy boundaries. On canonical Silesia (202.12 MiB) under a 7-repeat protocol (W = 6 pinned P-cores), GPX Track A produces standard RFC 8878-compliant .zst streams compressing to 56.20 MiB at 191.44 MiB/s with native line-speed decompression (5,182.79 MiB/s at W=6, 1,706.52 MiB/s at W=1 on unmodified libzstd, representing 124.1% single-core speed of Stock Level 9). GPX Track B achieves 54.49 MiB at 299.23 MiB/s (CI [295.3, 305.7]) with 4,919.14 MiB/s decode, strictly dominating Stock Levels 9--14 across size and speed (1.04x faster and 1.95 MB smaller than Level 9; 151.78 KB smaller and 8.24x faster than Level 14). GPX Track A+B achieves 54.37 MiB at 204.87 MiB/s (CI [199.6, 210.3]) with 4,599.00 MiB/s decode, strictly dominating Levels 10--14 under 95% Bootstrap CIs (276.07 KB smaller than Level 14, 1.19x to 5.64x faster). Generalization is confirmed across 6 suites (49 instances, 48 streams): Canterbury (-0.65%), Calgary (-1.66%), PE binaries (-14.10%), and Transformer BFloat16 tensors (-10.61%, saving 3.28 MB vs Stock L9). Downstream SIMT nibble unpacking sustains 103.62 GiB/s (1.05x PCIe speedup). All outputs are bit-exact verified via SHA-256.

cs.AR↗

Terracotta: Enabling the Adoption of New DRAM Techniques via a Flexible DRAM Interface and Memory Controller

DRAM continues to limit the performance, energy efficiency, and robustness of modern systems. Many prior works propose DRAM techniques that support in-DRAM computation, improve memory access latency and parallelism, and enhance DRAM maintenance and reliability. However, adopting each new technique requires repeated modifications to the rigid DRAM interface and memory controller, hindering its deployment. Our goal is to reduce these repeated modifications. We observe that the DRAM commands and memory controller structures of many DRAM techniques are similar. Our key idea is to use these similarities to compose a set of primitives for implementing diverse DRAM techniques. We propose Terracotta, a new framework with two flexible components: (i) custom command extensions that let DRAM vendors define new commands within a single, standardized interface, and (ii) a programmable memory controller that system designers can program to support new DRAM techniques post-silicon. Together, these enable deployment by configuring the memory controller instead of modifying the interface and controller. We design Terracotta for a DDR5-based system and evaluate its performance, energy, and hardware complexity. For four DRAM techniques from four distinct domains (processing-using-DRAM, low-cost DRAM maintenance, subarray-level parallelism, and latency reduction), Terracotta retains almost all of the performance benefits (>96%) of custom implementations. A Terracotta-based composition of two techniques outperforms the Terracotta-based implementation of each technique alone, demonstrating the benefits of adding techniques without repeated interface and controller modifications. Terracotta incurs low DRAM energy (0.6-3.2%), area (0.03%), and power (0.56%) overheads in a high-end server-grade processor. Terracotta's source code is freely available at https://github.com/CMU-SAFARI/Terracotta.

cs.AR↗

Beyond No-Good Benders Cuts: Exact Realizability for Discretely Tunable Clock Trees

Useful-skew schedules can satisfy timing constraints yet remain unrealizable by a fixed clock tree with discrete tuning choices. We formulate this mismatch as exact membership in a finite relative-latency set and develop a checkable feedback interface between the scheduler and the tree model. Arithmetic certificates explain unrealizable targets, while tree-specific contraction reduces the structural size of the exact relations projected onto selected sinks. These relations become realizability cuts that can exclude more candidates than a no-good on the same certificate support, including discrete holes that linear inequalities cannot separate. Controlled experiments confirm fewer oracle calls and faster projection construction. They also expose important limits: globally minimum certificates can cost more than they save, compact arithmetic feedback can fail on non-parity obstructions, and direct monolithic optimization remains faster on the tested additive models. The contribution is an exact, independently verifiable scheduler--tree interface, rather than a claim of universal solver acceleration or physical signoff.

cs.AR↗