SearcharxivSearch

arXiv subjects

Jiaqi Lou

Publications and source records attributed to Jiaqi Lou.

5 recordsLinked to original sources

A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices

In modern server CPUs, the Last-Level Cache (LLC) serves not only as a victim cache for higher-level private caches but also as a buffer for low-latency DMA transfers between CPU cores and I/O devices through Direct Cache Access (DCA). However, prior work has shown that high-bandwidth network-I/O devices can rapidly flood the LLC with packets, often causing significant contention with co-running workloads. One step further, this work explores hidden microarchitectural properties of the Intel Xeon CPUs, uncovering two previously unrecognized LLC contentions triggered by emerging high-bandwidth I/O devices. Specifically, (C1) DMA-written cache lines in LLC ways designated for DCA (referred to as DCA ways) are migrated to certain LLC ways (denoted as inclusive ways) when accessed by CPU cores, unexpectedly contending with non-I/O cache lines within the inclusive ways. In addition, (C2) high-bandwidth storage-I/O devices, which are increasingly common in datacenter servers, benefit little from DCA while contending with (latency-sensitive) network-I/O devices within DCA ways. To this end, we present \design, a runtime LLC management framework designed to alleviate both (C1) and (C2) among diverse co-running workloads, using a hidden knob and other hardware features implemented in those CPUs. Additionally, we demonstrate that \design can also alleviate other previously known network-I/O-driven LLC contentions. Overall, it improves the performance of latency-sensitive, high-priority workloads by 51\% without notably compromising that of low-priority workloads.

cs.AR

Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices

The ever-growing demands for memory with larger capacity and higher bandwidth have driven recent innovations on memory expansion and disaggregation technologies based on Compute eXpress Link (CXL). Especially, CXL-based memory expansion technology has recently gained notable attention for its ability not only to economically expand memory capacity and bandwidth but also to decouple memory technologies from a specific memory interface of the CPU. However, since CXL memory devices have not been widely available, they have been emulated using DDR memory in a remote NUMA node. In this paper, for the first time, we comprehensively evaluate a true CXL-ready system based on the latest 4th-generation Intel Xeon CPU with three CXL memory devices from different manufacturers. Specifically, we run a set of microbenchmarks not only to compare the performance of true CXL memory with that of emulated CXL memory but also to analyze the complex interplay between the CPU and CXL memory in depth. This reveals important differences between emulated CXL memory and true CXL memory, some of which will compel researchers to revisit the analyses and proposals from recent work. Next, we identify opportunities for memory-bandwidth-intensive applications to benefit from the use of CXL memory. Lastly, we propose a CXL-memory-aware dynamic page allocation policy, Caption to more efficiently use CXL memory as a bandwidth expander. We demonstrate that Caption can automatically converge to an empirically favorable percentage of pages allocated to CXL memory, which improves the performance of memory-bandwidth-intensive applications by up to 24% when compared to the default page allocation policy designed for traditional NUMA systems.

cs.PF

A (Dummy's) Guide to Working with Gapped Boundaries via (Fermion) Condensation

We study gapped boundaries characterized by "fermionic condensates" in 2+1 d topological order. Mathematically, each of these condensates can be described by a super commutative Frobenius algebra. We systematically obtain the species of excitations at the gapped boundary/ junctions, and study their endomorphisms (ability to trap a Majorana fermion) and fusion rules, and generalized the defect Verlinde formula to a twisted version. We illustrate these results with explicit examples. We also connect these results with topological defects in super modular invariant CFTs. To render our discussion self-contained, we provide a pedagogical review of relevant mathematical results, so that physicists without prior experience in tensor category should be able to pick them up and apply them readily

hep-th

Ishibashi States, Topological Orders with Boundaries and Topological Entanglement Entropy II -- Cutting through the boundary

We compute the entanglement entropy in a 2+1 dimensional topological order in the presence of gapped boundaries. Specifically, we consider entanglement cuts that cut through the boundaries. We argue that based on general considerations of the bulk-boundary correspondence, the "twisted characters" feature in the Renyi entropy, and the topological entanglement entropy is controlled by a "half-linking number" in direct analogy to the role played by the S-modular matrix in the absence of boundaries. We also construct a class of boundary states based on the half-linking numbers that provides a "closed-string" picture complementing an "open-string" computation of the entanglement entropy. These boundary states do not correspond to diagonal RCFT's in general. These are illustrated in specific Abelian Chern-Simons theories with appropriate boundary conditions.

hep-th

Ishibashi States, Topological Orders with Boundaries and Topological Entanglement Entropy

In this paper, we study gapped edges/interfaces in a 2+1 dimensional bosonic topological order and investigate how the topological entanglement entropy is sensitive to them. We present a detailed analysis of the Ishibashi states describing these edges/interfaces making use of the physics of anyon condensation in the context of Abelian Chern-Simons theory, which is then generalized to more non-Abelian theories whose edge RCFTs are known. Then we apply these results to computing the entanglement entropy of different topological orders. We consider cases where the system resides on a cylinder with gapped boundaries and that the entanglement cut is parallel to the boundary. We also consider cases where the entanglement cut coincides with the interface on a cylinder. In either cases, we find that the topological entanglement entropy is determined by the anyon condensation pattern that characterizes the interface/boundary. We note that conditions are imposed on some non-universal parameters in the edge theory to ensure existence of the conformal interface, analogous to requiring rational ratios of radii of compact bosons.

hep-th