Searcharxiv⌕ Search

arXiv subjects

Chang Hyun Park

Publications and source records attributed to Chang Hyun Park.

3 recordsLinked to original sources

Real-time quasi-distributed fiber optic sensor based on resonance frequency mapping

Distributed optical fiber sensors (DOFS) based on Raman, Brillouin, and Rayleigh scattering have recently attracted considerable attention for various sensing applications, especially large-scale monitoring, owing to their capacity for measuring strain or temperature distributions. However, ultraweak backscatter signals within optical fibers constitute an inevitable problem for DOFS, thereby increasing the burden on the entire system in terms of limited spatial resolution, low measurement speed, high system complexity, or high cost. We propose a novel resonance frequency mapping for a real-time quasi-distributed fiber optic sensor based on identical weak fiber Bragg gratings (FBG), which has stronger reflection signals and high sensitivity to multiple sensing parameters. The resonance configuration, which amplifies optical signals during multiple round-trip propagations, can simply and efficiently address the intrinsic problems in conventional single round-trip measurements for identical weak FBG sensors, such as crosstalk and optical power depletion. Moreover, it is technically feasible to perform individual measurements for a large number of quasi-distributed identical weak FBGs with relatively high signal-to-noise ratio (SNR), low crosstalk, and low optical power depletion. By mapping the resonance frequency spectrum, the dynamic response of each identical weak FBG is rapidly acquired in the order of kilohertz, and direct interrogation in real time is possible without time-consuming computation, such as fast Fourier transformation (FFT). This resonance frequency spectrum is obtained on the basis of an all-fiber electro-optic configuration that allows simultaneous measurement of quasi-distributed strain responses with high speed (>5 kHz), high stability (approximately 2.4 microstrain), and high linearity (R^2 = 0.9999).

physics.optics↗

Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving

Transformers are the driving force behind today's Large Language Models (LLMs), serving as the foundation for their performance and versatility. Yet, their compute and memory costs grow with sequence length, posing scalability challenges for long-context inferencing. In response, the algorithm community is exploring alternative architectures, such as state space models (SSMs), linear attention, and recurrent neural networks (RNNs), which we refer to as post-transformers. This shift presents a key challenge: building a serving system that efficiently supports both transformer and post-transformer LLMs within a unified framework. To address this challenge, we analyze the performance characteristics of transformer and post-transformer LLMs. Despite their algorithmic differences, both are fundamentally limited by memory bandwidth under batched inference due to attention in transformers and state updates in post-transformers. Further analyses suggest two additional insights: (1) state update operations, unlike attention, incur high hardware cost, making per-bank PIM acceleration inefficient, and (2) different low-precision arithmetic methods offer varying accuracy-area tradeoffs, while we identify Microsoft's MX as the Pareto-optimal choice. Building on these insights, we design Pimba as an array of State-update Processing Units (SPUs), each shared between two banks to enable interleaved access to PIM. Each SPU includes a State-update Processing Engine (SPE) that comprises element-wise multipliers and adders using MX-based quantized arithmetic, enabling efficient execution of state update and attention operations. Our evaluation shows that, compared to LLM-optimized GPU and GPU+PIM systems, Pimba achieves up to 4.1x and 2.1x higher token generation throughput, respectively.

cs.AR↗

Page Tables: Keeping them Flat and Hot (Cached)

As memory capacity has outstripped TLB coverage, large data applications suffer from frequent page table walks. We investigate two complementary techniques for addressing this cost: reducing the number of accesses required and reducing the latency of each access. The first approach is accomplished by opportunistically "flattening" the page table: merging two levels of traditional 4KB page table nodes into a single 2MB node, thereby reducing the table's depth and the number of indirections required to search it. The second is accomplished by biasing the cache replacement algorithm to keep page table entries during periods of high TLB miss rates, as these periods also see high data miss rates and are therefore more likely to benefit from having the smaller page table in the cache than to suffer from increased data cache misses. We evaluate these approaches for both native and virtualized systems and across a range of realistic memory fragmentation scenarios, describe the limited changes needed in our kernel implementation and hardware design, identify and address challenges related to self-referencing page tables and kernel memory allocation, and compare results across server and mobile systems using both academic and industrial simulators for robustness. We find that flattening does reduce the number of accesses required on a page walk (to 1.0), but its performance impact (+2.3%) is small due to Page Walker Caches (already 1.5 accesses). Prioritizing caching has a larger effect (+6.8%), and the combination improves performance by +9.2%. Flattening is more effective on virtualized systems (4.4 to 2.8 accesses, +7.1% performance), due to 2D page walks. By combining the two techniques we demonstrate a state-of-the-art +14.0% performance gain and -8.7% dynamic cache energy and -4.7% dynamic DRAM energy for virtualized execution with very simple hardware and software changes.

cs.AR↗